A unified framework for controllable caption generation across images, videos, and audio. Supports multi-modal inputs and customizable caption styles.
☆54Jul 24, 2025Updated last year
Alternatives and similar repositories for AnyCap
Users that are interested in AnyCap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Codebase of 'From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model'☆45Jun 27, 2026Updated last month
- ☆35Jan 20, 2026Updated 7 months ago
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 4 months ago
- [ICML 2026] Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO☆19Jun 15, 2026Updated 2 months ago
- [CVPR 2026] See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning☆23Jun 28, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer☆49Jan 6, 2026Updated 7 months ago
- [ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasoner☆65May 29, 2026Updated 2 months ago
- ☆34Feb 12, 2026Updated 6 months ago
- Large Language Models Can Self-Improve in Long-context Reasoning☆72Nov 24, 2024Updated last year
- [ICML 2026] The offical code of Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis☆90Jun 2, 2026Updated 2 months ago
- [2025-TMLR] A Survey on the Honesty of Large Language Models☆66Dec 8, 2024Updated last year
- [NeurIPS 2025] Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance☆654Jan 5, 2026Updated 7 months ago
- [ICLR 2025] ChartMimic: Evaluating LMM’s Cross-Modal Reasoning Capability via Chart-to-Code Generation☆134Dec 19, 2025Updated 8 months ago
- ☆13May 17, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Unified Language-driven Zero-shot Domain Adaptation (CVPR 2024)☆17Nov 28, 2024Updated last year
- The code for paper "LLM-Neo: Parameter Efficient Knowledge Distillation for Large Language Models"☆15Mar 2, 2025Updated last year
- [ACL 2024 Findings] CriticBench: Benchmarking LLMs for Critique-Correct Reasoning☆31Mar 5, 2024Updated 2 years ago
- RePO: Replay-Enhanced Policy Optimization☆24Jun 12, 2025Updated last year
- ☆16Jul 17, 2026Updated last month
- Description for MV-MATH☆15Jul 20, 2025Updated last year
- [NeurIPS 2025] Implementation for the paper "The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning"☆167Mar 2, 2026Updated 5 months ago
- official repo for "VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation" [EMNLP2024]☆125Dec 4, 2025Updated 8 months ago
- Official PyTorch Implementation of "Giving a Hand to Diffusion Models: a Two-Stage Approach to Improving Conditional Human Image Generati…☆26Mar 19, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The Source Code for IF-VidCap @ICLR 2026☆18Oct 22, 2025Updated 10 months ago
- CVPR2025: Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning☆40Mar 21, 2025Updated last year
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 5 months ago
- [ICLR 2026] Official Implementation of ProxyThinker: Test-Time Guidance through Small Visual Reasoners.☆22Sep 24, 2025Updated 10 months ago
- [ICML 2026] Official implementation of "Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence"☆162May 1, 2026Updated 3 months ago
- This is a repository contains the implementation of our NeurIPS'24 paper "Temporal Sentence Grounding with Relevance Feedback in Videos"☆13Aug 22, 2025Updated last year
- Official implementation of "Opt-In Art: Learning Art Styles Only from Few Examples" (Accepted by NeurIPS 2025)☆33Nov 30, 2025Updated 8 months ago
- (ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’☆60Jan 26, 2026Updated 6 months ago
- Interface Design for Self-Supervised Speech Models, Accepted to Interspeech2024☆16Nov 19, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code of StyleCrafter on SDXL☆20Jun 25, 2024Updated 2 years ago
- ☆14Jan 22, 2025Updated last year
- [ICLR2023] Video Scene Graph Generation from Single-Frame Weak Supervision☆12Sep 17, 2023Updated 2 years ago
- Test-time Scaling for VAR models☆33Sep 19, 2025Updated 11 months ago
- 3D generation code for iFusion☆15Dec 29, 2023Updated 2 years ago
- Repo for paper "MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding".☆39Jun 9, 2025Updated last year
- ☆94May 8, 2026Updated 3 months ago