large scale pre-training VLMs
☆25Sep 14, 2026Updated this week
Alternatives and similar repositories for vlm-training
Users that are interested in vlm-training are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official evaluation scripts and baseline prompts for the DocVQA 2026 (ICDAR 2026) Competition on Multimodal Reasoning over Documents.☆18Mar 16, 2026Updated 6 months ago
- [ECCV'24] Official Implementation of Autoregressive Visual Entity Recognizer.☆14Mar 2, 2024Updated 2 years ago
- ☆17Feb 20, 2025Updated last year
- Deep learning Framework from scratch.☆11Jul 23, 2025Updated last year
- This is the official repository for the paper "Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction". ICCV …☆28May 13, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A curated list of resources on Document Layout Analysis☆12Aug 7, 2025Updated last year
- A kid friendly AI chat demo☆19Apr 14, 2026Updated 5 months ago
- Recurrence Meets Transformers for Universal Multimodal Retrieval☆15Dec 15, 2025Updated 9 months ago
- [CVPR 2025 Highlight] Official implementation of HySAC, a hyperbolic safety-aware vision-language model for safer multimodal retrieval an…☆31Apr 8, 2025Updated last year
- [CVPR2025] Official implementation of the paper "Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practi…☆48Oct 29, 2025Updated 10 months ago
- This repository is part of a high school research work about generating images with a simulated quantum computer.☆10Feb 13, 2026Updated 7 months ago
- [CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering☆57Jul 14, 2025Updated last year
- ☆11Apr 10, 2024Updated 2 years ago
- [ECCV'24] Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities☆52Jul 2, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆26Mar 13, 2021Updated 5 years ago
- [ICLR 2026] "Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals"☆50Mar 6, 2026Updated 6 months ago
- An implementation from scratch of major Graph Neural Network (GNN) architectures using Numpy☆17Mar 16, 2026Updated 6 months ago
- [ICLR 2025] Causal Graphical Models for Vision-Language Compositional Understanding☆10Apr 15, 2025Updated last year
- Chesster.io: My personal project to recreate the classic chess game online, offering a seamless experience with a focus on responsive des…☆24Feb 21, 2024Updated 2 years ago
- [ECCV 2024] Official implementation of Safe-CLIP, a safety-aligned vision-language model for mitigating NSFW content in multimodal retrie…☆68Aug 10, 2024Updated 2 years ago
- Standalone Gemma 4 PyTorch Model using Claude Code☆15Jul 26, 2026Updated last month
- Simple JupyterHub like setup using only JupyterLab☆25Sep 1, 2022Updated 4 years ago
- [ICCV 2023] Going Beyond Nouns With Vision & Language Models Using Synthetic Data☆13Sep 30, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- An open source implementation of CLIP.☆33Nov 7, 2022Updated 3 years ago
- ☆30Jan 8, 2021Updated 5 years ago
- This is an implementation of the paper "Are We Done with Object-Centric Learning?"☆14Jun 21, 2026Updated 2 months ago
- An advanced Omniverse Isaac Sim extension for dynamic management and real-time animation of scene lighting, enabling more flexible, inter…☆15Aug 29, 2025Updated last year
- [IJCAI 2025] Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives☆37Nov 25, 2025Updated 9 months ago
- timm, evolved☆61May 28, 2026Updated 3 months ago
- [ICCVW 2025] LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning☆160Aug 8, 2025Updated last year
- Multimodal Agentic Document QA benchmark (MADQA)☆42Mar 13, 2026Updated 6 months ago
- TUI application for viewing the status of GPU allocations on a Slurm cluster☆11Dec 11, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆15Feb 13, 2025Updated last year
- ☆14Mar 22, 2024Updated 2 years ago
- A simple wrapper library for binding timm models as detectron2 backbones☆45May 31, 2023Updated 3 years ago
- An Educational Framework Based on PyTorch for Deep Learning Education and Exploration☆11Dec 24, 2023Updated 2 years ago
- The code for the paper "Pre-trained Vision-Language Models Learn Discoverable Concepts"☆21Jun 5, 2024Updated 2 years ago
- Generative neural networks for 3D terrain.☆33Dec 18, 2024Updated last year
- ☆17Apr 29, 2025Updated last year