Ola: Pushing the Frontiers of Omni-Modal Language Model
☆395Jun 13, 2025Updated last year
Alternatives and similar repositories for Ola
Users that are interested in Ola are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] MLLM for On-Demand Spatial-Temporal Understanding at Arbitrary Resolution☆329Jul 4, 2025Updated last year
- ✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction☆2,520Mar 28, 2025Updated last year
- [CVPR2025 Highlight] Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models☆240Nov 7, 2025Updated 8 months ago
- Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos☆72Sep 5, 2025Updated 10 months ago
- ☆193Feb 8, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Baichuan-Omni: Towards Capable Open-source Omni-modal LLM 🌊☆273Jan 27, 2025Updated last year
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception☆159Dec 6, 2024Updated last year
- Align Anything: Training All-modality Model with Feedback☆4,663Nov 27, 2025Updated 7 months ago
- Chain-of-Spot: Interactive Reasoning Improves Large Vision-language Models☆100Mar 22, 2024Updated 2 years ago
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving stat…☆1,582Jun 14, 2025Updated last year
- Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities☆1,205Jul 15, 2025Updated last year
- WavReward: Spoken Dialogue Models With Generalist Reward Evaluators☆56May 15, 2025Updated last year
- A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.☆1,468Updated this week
- [ICLR 2026] Data Pipeline, Models, and Benchmark for Omni-Captioner.☆142Apr 7, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [ECCV 2024] Efficient Inference of Vision Instruction-Following Models with Elastic Cache☆43Jul 26, 2024Updated last year
- [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.☆1,963Jan 8, 2026Updated 6 months ago
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs☆88Jan 17, 2026Updated 6 months ago
- Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and pe…☆4,042Jun 12, 2025Updated last year
- [ICCV 2025] Official repo for "GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation"☆204Jan 7, 2026Updated 6 months ago
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction☆223Feb 28, 2025Updated last year
- This repo contains evaluation code for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?"☆31Dec 23, 2024Updated last year
- A fork to add multimodal model training to open-r1☆1,591Feb 8, 2025Updated last year
- [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding☆529Nov 14, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fully Open Framework for Democratized Multimodal Training☆1,146Updated this week
- Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/…☆309Updated this week
- A simple implementation for improving CosyVoice2 by GRPO method☆39May 5, 2026Updated 2 months ago
- [🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s …☆689Feb 27, 2026Updated 4 months ago
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs☆1,303Jan 23, 2025Updated last year
- Next-Token Prediction is All You Need☆2,432Jan 12, 2026Updated 6 months ago
- Valley is a cutting-edge multimodal large model designed to handle a variety of tasks involving text, images, video, and audio data.☆287May 8, 2026Updated 2 months ago
- [CVPR 2025] PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models☆54Jun 12, 2025Updated last year
- Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities。☆1,905Jan 16, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [arXiv: 2502.05178] QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation☆97Mar 1, 2025Updated last year
- [NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix☆214Feb 25, 2026Updated 4 months ago
- Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)☆726Sep 24, 2025Updated 9 months ago
- ☆4,710Jun 15, 2026Updated last month
- Streamable Text-to-Speech model using a language modeling approach, without vector quantization☆108May 20, 2025Updated last year
- Code associated with the paper: CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition.☆17May 16, 2025Updated last year
- Frontier Multimodal Foundation Models for Image and Video Understanding☆1,172Aug 14, 2025Updated 11 months ago