☆94Oct 21, 2025Updated 10 months ago
Alternatives and similar repositories for HunyuanVision
Users that are interested in HunyuanVision are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆43Aug 5, 2025Updated last year
- Native Multimodal Models are World Learners☆1,555Dec 30, 2025Updated 8 months ago
- [TPAMI 2023] Object Affinity Learning: Towards Annotation-free Instance Segmentation☆14Sep 14, 2023Updated 3 years ago
- Initial code for computer vision experiments☆11Jan 1, 2023Updated 3 years ago
- [TPAMI 26/ NeurIPS 24] Official PyTorch Implementation of "FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity Refiner…☆75Oct 21, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation☆3,265Jun 23, 2026Updated 2 months ago
- ☆19May 17, 2025Updated last year
- [ICML 2026] Elastic Diffusion Transformer: Accelerating SOTA generation models (e.g., Qwen-Image, Hunyuan3d ) through adaptive computatio…☆50May 1, 2026Updated 4 months ago
- HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation☆674Oct 14, 2025Updated 11 months ago
- Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities☆1,227Jul 15, 2025Updated last year
- Learning Remote Sensing Object Detection with Single Point Supervision☆18Dec 18, 2023Updated 2 years ago
- [Preprint] UCGM: Unified Continuous Generative Models☆188May 27, 2025Updated last year
- [ICCV 2025 Highlight] The official repository for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining"☆197Mar 17, 2025Updated last year
- [CVPR2025 Highlight] Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models☆242Nov 7, 2025Updated 10 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ECCV 2024] Efficient Inference of Vision Instruction-Following Models with Elastic Cache☆43Jul 26, 2024Updated 2 years ago
- Official Implementation of ICCV 2023 Paper - SegPrompt: Boosting Open-World Segmentation via Category-level Prompt Learning☆110May 28, 2025Updated last year
- Chain-of-Spot: Interactive Reasoning Improves Large Vision-language Models☆100Mar 22, 2024Updated 2 years ago
- Paper List for In-context Learning 🌷☆19Jan 3, 2023Updated 3 years ago
- ☆32Aug 18, 2026Updated last month
- CODA: Repurposing Continuous VAEs for Discrete Tokenization☆37Jul 4, 2025Updated last year
- SccovNet for remote sensing scene image classification which accepted by TNNLS☆13Jun 24, 2019Updated 7 years ago
- Code release for Ming-UniVision: Joint Image Understanding and Geneation with a Continuous Unified Tokenizer☆143Oct 14, 2025Updated 11 months ago
- [CVPR 2024] Official implementation of "ViTamin: Designing Scalable Vision Models in the Vision-language Era"☆211Jun 9, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official inference code and LongText-Bench benchmark for our paper X-Omni (https://arxiv.org/pdf/2507.22058).☆428Aug 26, 2025Updated last year
- ☆20Jul 11, 2023Updated 3 years ago
- [Arxiv'25] MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization☆55Sep 16, 2025Updated last year
- ☆29Aug 21, 2025Updated last year
- This is the official repo for the paper "LongCat-Flash-Omni Technical Report"☆507May 9, 2026Updated 4 months ago
- [IGARSS 2022] CapFormer: Pure transformer for remote sensing image caption☆21Oct 6, 2022Updated 3 years ago
- [NeurIPS 2025 Spotlight] Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning☆40May 20, 2025Updated last year
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving stat…☆1,588Jun 14, 2025Updated last year
- ☆190Mar 13, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆18Jul 8, 2026Updated 2 months ago
- GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning☆2,388Sep 2, 2026Updated 2 weeks ago
- The SAIL-VL2 series model developed by the BytedanceDouyinContent Group☆79Sep 18, 2025Updated last year
- ☆220Dec 19, 2025Updated 9 months ago
- [AAAI'24] Official dataset & demo code for MID-FiLD: MIDI Dataset for Fine-Level Dynamics☆21Mar 31, 2024Updated 2 years ago
- LIGHTVOC AN UPSAMPLING-FREE GAN VOCODER BASED ON CONFORMER AND INVERSE SHORT-TIME FOURIER TRANSFORM☆18May 17, 2024Updated 2 years ago
- ☆20Jul 22, 2025Updated last year