☆94Oct 21, 2025Updated 11 months ago
Alternatives and similar repositories for HunyuanVision
Users that are interested in HunyuanVision are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆43Aug 5, 2025Updated last year
- Native Multimodal Models are World Learners☆1,561Dec 30, 2025Updated 9 months ago
- [TPAMI 2023] Object Affinity Learning: Towards Annotation-free Instance Segmentation☆14Sep 14, 2023Updated 3 years ago
- Initial code for computer vision experiments☆11Jan 1, 2023Updated 3 years ago
- [TPAMI 26/ NeurIPS 24] Official PyTorch Implementation of "FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity Refiner…☆75Oct 21, 2025Updated 11 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation☆3,284Jun 23, 2026Updated 3 months ago
- ☆19May 17, 2025Updated last year
- [ICML 2026] Elastic Diffusion Transformer: Accelerating SOTA generation models (e.g., Qwen-Image, Hunyuan3d ) through adaptive computatio…☆50Sep 30, 2026Updated last week
- ☆16Aug 10, 2025Updated last year
- HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation☆674Oct 14, 2025Updated 11 months ago
- Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities☆1,226Jul 15, 2025Updated last year
- Learning Remote Sensing Object Detection with Single Point Supervision☆18Dec 18, 2023Updated 2 years ago
- [Preprint] UCGM: Unified Continuous Generative Models☆188May 27, 2025Updated last year
- [ICCV 2025 Highlight] The official repository for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining"☆197Mar 17, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [CVPR2025 Highlight] Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models☆242Nov 7, 2025Updated 11 months ago
- [ECCV 2024] Efficient Inference of Vision Instruction-Following Models with Elastic Cache☆43Jul 26, 2024Updated 2 years ago
- Official Implementation of ICCV 2023 Paper - SegPrompt: Boosting Open-World Segmentation via Category-level Prompt Learning☆110May 28, 2025Updated last year
- Chain-of-Spot: Interactive Reasoning Improves Large Vision-language Models☆100Mar 22, 2024Updated 2 years ago
- Paper List for In-context Learning 🌷☆19Jan 3, 2023Updated 3 years ago
- CODA: Repurposing Continuous VAEs for Discrete Tokenization☆37Jul 4, 2025Updated last year
- Code release for Ming-UniVision: Joint Image Understanding and Geneation with a Continuous Unified Tokenizer☆144Oct 14, 2025Updated 11 months ago
- [CVPR 2024] Official implementation of "ViTamin: Designing Scalable Vision Models in the Vision-language Era"☆210Jun 9, 2024Updated 2 years ago
- This is the official code for paper [RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward N…☆18Jun 20, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Official inference code and LongText-Bench benchmark for our paper X-Omni (https://arxiv.org/pdf/2507.22058).☆428Aug 26, 2025Updated last year
- ☆20Jul 11, 2023Updated 3 years ago
- [Arxiv'25] MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization☆55Sep 16, 2025Updated last year
- ☆29Aug 21, 2025Updated last year
- This is the official repo for the paper "LongCat-Flash-Omni Technical Report"☆508May 9, 2026Updated 5 months ago
- PSGCNet: A Pyramidal Scale and Global Context Guided Network for Dense Object Counting in Remote-Sensing Images☆20Jun 13, 2022Updated 4 years ago
- FormulaOne: A dataset of algorithmic problems based on MSO formulas.☆26Mar 1, 2026Updated 7 months ago
- [IGARSS 2022] CapFormer: Pure transformer for remote sensing image caption☆21Oct 6, 2022Updated 4 years ago
- [NeurIPS 2025 Spotlight] Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning☆40May 20, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆190Mar 13, 2026Updated 6 months ago
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving stat…☆1,593Jun 14, 2025Updated last year
- ☆19Jul 8, 2026Updated 3 months ago
- Code for the paper "Toward Fully Self-Supervised Multi-Pitch Estimation".☆25Sep 27, 2025Updated last year
- GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning☆2,396Oct 2, 2026Updated last week
- The SAIL-VL2 series model developed by the BytedanceDouyinContent Group☆79Sep 18, 2025Updated last year
- ☆220Dec 19, 2025Updated 9 months ago