☆54Oct 20, 2025Updated 10 months ago
Alternatives and similar repositories for verl-internvl
Users that are interested in verl-internvl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMs☆57Mar 9, 2025Updated last year
- 北航“冯如杯”论文模板 (2022年)☆12Apr 24, 2022Updated 4 years ago
- ☆19Jul 8, 2026Updated 2 months ago
- Official code repo of Video-Browser: Towards Agentic Open-web Video Browsing☆28Jan 19, 2026Updated 8 months ago
- 【2024 ECAI】First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending☆14Jun 16, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads☆15Feb 11, 2026Updated 7 months ago
- [preprint] Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning☆19Feb 18, 2026Updated 7 months ago
- [NAACL 2024] Part-based, explainable and editable fine-grained image classifier that allows users to define a species in text☆13Sep 19, 2025Updated 11 months ago
- ☆28Mar 27, 2025Updated last year
- Contains some of the ML codes which I made while learning ML☆14Dec 21, 2017Updated 8 years ago
- Comprehensive benchmark for video text understanding☆29Jun 4, 2025Updated last year
- [SCIS 2024] The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Di…☆64Nov 7, 2024Updated last year
- Paper List for Dialogue and Interactive Systems☆16Jun 5, 2020Updated 6 years ago
- [WACV 2026] MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval☆15Sep 18, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official code repo for our work "Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models"☆54Jun 17, 2025Updated last year
- Python package to accelerate research on generalized out-of-distribution (OOD) detection.☆15Jun 19, 2024Updated 2 years ago
- [ICML 2026] Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions☆58Jun 29, 2026Updated 2 months ago
- Codes of Paper "Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding"☆21Jul 14, 2026Updated 2 months ago
- [Accepted By EMNLP 2026 Main Conference] Sequential Diffusion Language Model (SDLM) enhances pre-trained autoregressive language models b…☆99Dec 27, 2025Updated 8 months ago
- ☆16Aug 28, 2024Updated 2 years ago
- Repository for the Paper: Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Inj…☆21Apr 17, 2026Updated 5 months ago
- Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning☆44Jul 2, 2025Updated last year
- [ACL'26] Official Repository for The Paper: What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time☆21Apr 7, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆20Sep 3, 2025Updated last year
- [ICLR 2025] DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models☆20Mar 25, 2025Updated last year
- ☆29Apr 18, 2026Updated 5 months ago
- Fleming-VL: Towards Universal Medical Visual Understanding with Multimodal LLMs☆18Nov 6, 2025Updated 10 months ago
- DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning☆19Jun 14, 2026Updated 3 months ago
- An official Project related to Paper "Perceiving Ambiguity and Semantics without Recognition: An Efficient and Effective Ambiguous Scene …☆22Dec 3, 2023Updated 2 years ago
- Approximating the joint distribution of language models via MCTS☆22Nov 3, 2024Updated last year
- Collection of PhD Advice Links☆23Oct 14, 2022Updated 3 years ago
- Official code for ICML 2024 paper, "Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models"☆19Jun 12, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆15May 26, 2025Updated last year
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation (ICLR 2026)☆22Apr 27, 2026Updated 4 months ago
- ☆17Jul 15, 2022Updated 4 years ago
- CVPR 2025 - R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning☆22Aug 28, 2025Updated last year
- g2o_frontend: a simple structure for handling slam with unknown data association by using the g2o backend optimization and its extensions☆15Jun 12, 2014Updated 12 years ago
- [ICLR 2025] TRACE: Temporal Grounding Video LLM via Casual Event Modeling☆157Aug 22, 2025Updated last year
- Bridge Megatron-Core to Hugging Face/Reinforcement Learning☆232Jun 15, 2026Updated 3 months ago