☆94Oct 21, 2025Updated 9 months ago
Alternatives and similar repositories for HunyuanVision
Users that are interested in HunyuanVision are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆43Aug 5, 2025Updated 11 months ago
- Native Multimodal Models are World Learners☆1,538Dec 30, 2025Updated 6 months ago
- [TPAMI 2023] Object Affinity Learning: Towards Annotation-free Instance Segmentation☆14Sep 14, 2023Updated 2 years ago
- Initial code for computer vision experiments☆11Jan 1, 2023Updated 3 years ago
- [TPAMI 26/ NeurIPS 24] Official PyTorch Implementation of "FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity Refiner…☆75Oct 21, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation☆3,203Jun 23, 2026Updated last month
- ☆19May 17, 2025Updated last year
- [ICML 2026] Elastic Diffusion Transformer: Accelerating SOTA generation models (e.g., Qwen-Image, Hunyuan3d ) through adaptive computatio…☆49May 1, 2026Updated 2 months ago
- HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation☆673Oct 14, 2025Updated 9 months ago
- Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities☆1,214Jul 15, 2025Updated last year
- [Preprint] UCGM: Unified Continuous Generative Models☆185May 27, 2025Updated last year
- Learning Remote Sensing Object Detection with Single Point Supervision☆18Dec 18, 2023Updated 2 years ago
- [ICCV 2025 Highlight] The official repository for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining"☆196Mar 17, 2025Updated last year
- [CVPR2025 Highlight] Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models☆240Nov 7, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ECCV 2024] Efficient Inference of Vision Instruction-Following Models with Elastic Cache☆43Jul 26, 2024Updated 2 years ago
- Official Implementation of ICCV 2023 Paper - SegPrompt: Boosting Open-World Segmentation via Category-level Prompt Learning☆112May 28, 2025Updated last year
- Paper List for In-context Learning 🌷☆19Jan 3, 2023Updated 3 years ago
- ☆32Dec 17, 2025Updated 7 months ago
- CODA: Repurposing Continuous VAEs for Discrete Tokenization☆37Jul 4, 2025Updated last year
- SccovNet for remote sensing scene image classification which accepted by TNNLS☆13Jun 24, 2019Updated 7 years ago
- Code release for Ming-UniVision: Joint Image Understanding and Geneation with a Continuous Unified Tokenizer☆143Oct 14, 2025Updated 9 months ago
- [CVPR 2024] Official implementation of "ViTamin: Designing Scalable Vision Models in the Vision-language Era"☆211Jun 9, 2024Updated 2 years ago
- This is the official code for paper [RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward N…☆18Jun 20, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official inference code and LongText-Bench benchmark for our paper X-Omni (https://arxiv.org/pdf/2507.22058).☆426Aug 26, 2025Updated 11 months ago
- ☆20Jul 11, 2023Updated 3 years ago
- [Arxiv'25] MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization☆55Sep 16, 2025Updated 10 months ago
- ☆29Aug 21, 2025Updated 11 months ago
- This is the official repo for the paper "LongCat-Flash-Omni Technical Report"☆501May 9, 2026Updated 2 months ago
- [IGARSS 2022] CapFormer: Pure transformer for remote sensing image caption☆21Oct 6, 2022Updated 3 years ago
- [NeurIPS 2025 Spotlight] Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning☆40May 20, 2025Updated last year
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving stat…☆1,583Jun 14, 2025Updated last year
- ☆190Mar 13, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆17Jul 8, 2026Updated 3 weeks ago
- Code for the paper "Toward Fully Self-Supervised Multi-Pitch Estimation".☆25Sep 27, 2025Updated 10 months ago
- GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning☆2,361Jul 21, 2026Updated last week
- The SAIL-VL2 series model developed by the BytedanceDouyinContent Group☆79Sep 18, 2025Updated 10 months ago
- ☆219Dec 19, 2025Updated 7 months ago
- [AAAI'24] Official dataset & demo code for MID-FiLD: MIDI Dataset for Fine-Level Dynamics☆21Mar 31, 2024Updated 2 years ago
- LIGHTVOC AN UPSAMPLING-FREE GAN VOCODER BASED ON CONFORMER AND INVERSE SHORT-TIME FOURIER TRANSFORM☆18May 17, 2024Updated 2 years ago