Self-evolving vision language models from zero data
☆77Mar 14, 2026Updated 4 months ago
Alternatives and similar repositories for MM-Zero
Users that are interested in MM-Zero are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Synthetic Video hallucination and Mitigation☆23Sep 21, 2025Updated 10 months ago
- ☆19Jul 1, 2026Updated 2 weeks ago
- Reinforcement Learning of Vision Language Models with Self Visual Perception Reward☆175Mar 14, 2026Updated 4 months ago
- The offical repo for "Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing"☆19Feb 3, 2026Updated 5 months ago
- [ECCV 2026] Official repository of "Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning".☆22Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [CVPR'26] VisPlay: Self-Evolving Vision-Language Models☆63Feb 25, 2026Updated 4 months ago
- The code implementation for TTCS: Test-Time Curriculum Synthesis for Self-Evolving.☆50Apr 22, 2026Updated 2 months ago
- Video Content Customization Using First Frame☆193Mar 17, 2026Updated 4 months ago
- Official implementation for the paper "Video-Based Reward Modeling for Computer-Use Agents"☆16Mar 14, 2026Updated 4 months ago
- This is the codes of "DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrieval"☆16Mar 6, 2026Updated 4 months ago
- [ICML'26] VideoGPA is a self-supervised framework that enhances 3D consistency in Video Diffusion Models.☆70Jun 6, 2026Updated last month
- MemSyco-Bench: Benchmarking Sycophancy in Agent Memory☆17Jul 7, 2026Updated 2 weeks ago
- [CVPR 2026] Official repo for "VideoSSR: Video Self-Supervised Reinforcement Learning"☆41Nov 11, 2025Updated 8 months ago
- The released data for paper "Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models".☆34Sep 16, 2023Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- OpenClaw-style theorem proving☆26Jun 11, 2026Updated last month
- [ICML 2026] Official implementation for paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation☆16Jun 4, 2026Updated last month
- A Survey of Self-Evolving Agents | A curated list of resources (surveys, papers, benchmarks, and opensource projects) on Self-Evolving Ag…☆339Updated this week
- A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.☆669Jul 8, 2026Updated last week
- Build coherent and visually polished multimodal webpages with hierarchical planning, AIGC tools, and iterative reflection.☆15May 17, 2026Updated 2 months ago
- ☆15Mar 10, 2026Updated 4 months ago
- MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models☆27May 23, 2026Updated last month
- [ICLR2026] codes for R-Zero: Self-Evolving Reasoning LLM from Zero Data (https://www.arxiv.org/pdf/2508.05004)☆824Feb 4, 2026Updated 5 months ago
- [ICLR 2026] Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play.☆136Feb 6, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Dependency-Aware Structural Retrieval for Massive Agent Skills☆188May 4, 2026Updated 2 months ago
- Official project page and code repository for WiT, a pixel space diffusion☆17May 31, 2026Updated last month
- [CVPR 2026] IOMM: Fast Pre-training of Unified Multimodal Models without Text-Image Pairs☆26Apr 11, 2026Updated 3 months ago
- ☆35Apr 3, 2026Updated 3 months ago
- Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models☆45Jun 14, 2024Updated 2 years ago
- A critical analysis of the Cambrian-S model and VSI-Super benchmarks☆16Nov 20, 2025Updated 8 months ago
- AutoHallusion Codebase (EMNLP 2024)☆23Dec 6, 2024Updated last year
- [ICML 2026] Residual Context Diffusion (RCD): Repurposing discarded signals as structured priors for high-performance reasoning in dLLMs.☆58Jun 28, 2026Updated 3 weeks ago
- SuperDebug,debug如此简单!☆17Jul 19, 2022Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [ICLR 2026] Official repo for "FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting"☆50Oct 9, 2025Updated 9 months ago
- We revisit the Platonic Representation Hypothesis using calibrated representational similarity metrics with statistical guarantees.☆36Jun 24, 2026Updated 3 weeks ago
- ☆25May 14, 2026Updated 2 months ago
- A unified framework for vision-language environments with Gymnasium-compatible interface☆35Mar 17, 2026Updated 4 months ago
- ☆16Jun 17, 2026Updated last month
- [ECCV 2026] An official implementation of "EndoCoT". Scaling endogenous Chain-of-Thought (CoT) reasoning in diffusion models for complex …☆43Jun 26, 2026Updated 3 weeks ago
- AcademiClaw: When Students Set Challenges for AI Agents — a bilingual benchmark of 80 university student-sourced academic tasks.☆18Jun 26, 2026Updated 3 weeks ago