A unified framework for controllable caption generation across images, videos, and audio. Supports multi-modal inputs and customizable caption styles.
☆54Jul 24, 2025Updated last year
Alternatives and similar repositories for AnyCap
Users that are interested in AnyCap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Codebase of 'From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model'☆45Jun 27, 2026Updated last month
- ☆35Jan 20, 2026Updated 6 months ago
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 4 months ago
- [ICML 2026] Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO☆19Jun 15, 2026Updated last month
- [CVPR 2026] See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning☆22Jun 28, 2026Updated last month
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer☆49Jan 6, 2026Updated 6 months ago
- [ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasoner☆64May 29, 2026Updated 2 months ago
- ☆35Feb 12, 2026Updated 5 months ago
- Large Language Models Can Self-Improve in Long-context Reasoning☆72Nov 24, 2024Updated last year
- [ICLR 2026] High-Fidelity Visual Reasoning on Structured Images☆30Jul 17, 2026Updated 2 weeks ago
- [ICML 2026] The offical code of Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis☆88Jun 2, 2026Updated 2 months ago
- [2025-TMLR] A Survey on the Honesty of Large Language Models☆66Dec 8, 2024Updated last year
- [NeurIPS 2025] Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance☆649Jan 5, 2026Updated 6 months ago
- [ICLR 2025] ChartMimic: Evaluating LMM’s Cross-Modal Reasoning Capability via Chart-to-Code Generation☆132Dec 19, 2025Updated 7 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests☆61Feb 28, 2026Updated 5 months ago
- ☆13May 17, 2025Updated last year
- Unified Language-driven Zero-shot Domain Adaptation (CVPR 2024)☆17Nov 28, 2024Updated last year
- [ACL 2024 Findings] CriticBench: Benchmarking LLMs for Critique-Correct Reasoning☆31Mar 5, 2024Updated 2 years ago
- RePO: Replay-Enhanced Policy Optimization☆24Jun 12, 2025Updated last year
- ☆16Jul 17, 2026Updated 2 weeks ago
- Description for MV-MATH☆15Jul 20, 2025Updated last year
- This is an official code for the paper: TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation☆28Mar 26, 2026Updated 4 months ago
- [NeurIPS 2025] Implementation for the paper "The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning"☆166Mar 2, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- official repo for "VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation" [EMNLP2024]☆122Dec 4, 2025Updated 8 months ago
- Official PyTorch Implementation of "Giving a Hand to Diffusion Models: a Two-Stage Approach to Improving Conditional Human Image Generati…☆26Mar 19, 2024Updated 2 years ago
- The Source Code for IF-VidCap @ICLR 2026☆19Oct 22, 2025Updated 9 months ago
- CVPR2025: Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning☆39Mar 21, 2025Updated last year
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 5 months ago
- Collection of awesome Continual Test-Time Adaptation methods☆24Jun 4, 2024Updated 2 years ago
- [ICML 2026] Official implementation of "Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence"☆158May 1, 2026Updated 3 months ago
- This is a repository contains the implementation of our NeurIPS'24 paper "Temporal Sentence Grounding with Relevance Feedback in Videos"☆13Aug 22, 2025Updated 11 months ago
- Official implementation of "Opt-In Art: Learning Art Styles Only from Few Examples" (Accepted by NeurIPS 2025)☆33Nov 30, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’☆60Jan 26, 2026Updated 6 months ago
- Interface Design for Self-Supervised Speech Models, Accepted to Interspeech2024☆16Nov 19, 2024Updated last year
- Code of StyleCrafter on SDXL☆20Jun 25, 2024Updated 2 years ago
- ☆14Jan 22, 2025Updated last year
- [ICLR2023] Video Scene Graph Generation from Single-Frame Weak Supervision☆12Sep 17, 2023Updated 2 years ago
- Test-time Scaling for VAR models☆33Sep 19, 2025Updated 10 months ago
- 3D generation code for iFusion☆15Dec 29, 2023Updated 2 years ago