☆119Feb 4, 2026Updated 5 months ago
Alternatives and similar repositories for WorldVQA
Users that are interested in WorldVQA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMs☆57Mar 9, 2025Updated last year
- Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities☆1,206Jul 15, 2025Updated last year
- ☆219Dec 19, 2025Updated 7 months ago
- ☆15Feb 25, 2026Updated 4 months ago
- Open Visual Agentic Intelligence☆2,218Jan 31, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for "SCL-RAI: Span-based Contrastive Learning with Retrieval Augmented Inference for Unlabeled Entity Problem in NER" @COLING-2022☆11Aug 20, 2022Updated 3 years ago
- Ideas for projects related to Tinker☆191Nov 6, 2025Updated 8 months ago
- Vision Large Language Models trained on M3IT instruction tuning dataset☆17Aug 16, 2023Updated 2 years ago
- Step3-VL-10B: A compact yet frontier multimodal model achieving SOTA performance at the 10B scale, matching open-source models 10-20x its…☆407Jan 21, 2026Updated 6 months ago
- Implementation for "The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer"☆85Oct 29, 2025Updated 8 months ago
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving stat…☆1,583Jun 14, 2025Updated last year
- Source code for paper "ATP: AMRize Than Parse! Enhancing AMR Parsing with PseudoAMRs" @NAACL-2022☆15Mar 31, 2023Updated 3 years ago
- High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)☆30Jan 22, 2026Updated 6 months ago
- Muon is Scalable for LLM Training☆1,509Aug 3, 2025Updated 11 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- We introduce BabyVision, a benchmark revealing the infancy of AI vision.☆232Jan 13, 2026Updated 6 months ago
- Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)☆726Sep 24, 2025Updated 9 months ago
- ☆453Aug 10, 2025Updated 11 months ago
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence☆385Jun 20, 2026Updated last month
- ☆1,481Nov 17, 2025Updated 8 months ago
- FlashKDA: high-performance Kimi Delta Attention kernels☆466May 26, 2026Updated last month
- BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution☆61Oct 13, 2025Updated 9 months ago
- ☆80Jun 20, 2025Updated last year
- ☆32Jul 2, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning☆103May 20, 2025Updated last year
- Vero: An Open RL Recipe for General Visual Reasoning☆134Jun 19, 2026Updated last month
- MiMo-VL☆642Aug 21, 2025Updated 11 months ago
- ☆102Dec 22, 2023Updated 2 years ago
- LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations☆30May 21, 2025Updated last year
- Official Repository of VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents☆114May 3, 2026Updated 2 months ago
- Official PyTorch implementation of the paper "Enhancing Vision-Language Pre-Training with Jointly Learned Questioner and Dense Captioner"☆15Aug 9, 2023Updated 2 years ago
- The official repo for the DanQing dataset.☆36Mar 25, 2026Updated 3 months ago
- [ACL 2023] Delving into the Openness of CLIP☆24Jan 11, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Accelerating the development of large multimodal models (LMMs) with one-click evaluation module - lmms-eval.☆70Aug 8, 2025Updated 11 months ago
- Muon in Int8 Precision Made Possible☆20Jun 18, 2026Updated last month
- Checkpoint-engine is a simple middleware to update model weights in LLM inference engines☆975Jul 4, 2026Updated 2 weeks ago
- Source code of our paper "Focus on the Target’s Vocabulary: Masked Label Smoothing for Machine Translation" @ACL-2022☆18May 19, 2022Updated 4 years ago
- [ICLR 2025] Source code for paper "A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegr…☆80Dec 10, 2024Updated last year
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo☆2,102Updated this week
- Extending context length of visual language models☆12Dec 18, 2024Updated last year