Code release for "LLMs can see and hear without any training"
☆460May 8, 2025Updated last year
Alternatives and similar repositories for MILS
Users that are interested in MILS are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- GitHub repository for AudioToolAgent☆20Feb 13, 2026Updated 5 months ago
- ☆197May 5, 2025Updated last year
- [ACL 2025 🔥] Rethinking Step-by-step Visual Reasoning in LLMs☆307May 21, 2025Updated last year
- ☆13Jul 10, 2024Updated 2 years ago
- High-Performance Implementation of OpenAI's TikToken.☆475Jul 3, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- (Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators☆642Jun 1, 2026Updated last month
- Fully neural approach for text chunking☆416Oct 23, 2025Updated 9 months ago
- [ICLR 2026] TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching☆877Jan 28, 2026Updated 5 months ago
- A browser-based, WebGL2 implementation of GPT-2 with transform block and attention matrix visualization☆346Oct 24, 2025Updated 9 months ago
- Calibrating LLM Confidence by Probing Perturbed Representation Stability☆19Jul 5, 2025Updated last year
- Code to train and evaluate Neural Attention Memory Models to obtain universally-applicable memory systems for transformers.☆360Oct 22, 2024Updated last year
- Rewriting Principia Mathematica in Lean☆139Feb 5, 2026Updated 5 months ago
- The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM", IJCV2025☆279May 26, 2025Updated last year
- \infty-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation☆21Feb 14, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, im…☆3,904Apr 23, 2026Updated 3 months ago
- Code for "Goal-Guided Neural Cellular Automata: Learning to Control Self-Organising Systems"☆56May 31, 2022Updated 4 years ago
- ☆14Jan 22, 2025Updated last year
- LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve spee…☆3,143May 19, 2025Updated last year
- ☆16Jun 23, 2026Updated last month
- Next-Token Prediction is All You Need☆2,433Jan 12, 2026Updated 6 months ago
- ☆279Mar 6, 2025Updated last year
- Everything about the SmolLM and SmolVLM family of models☆3,853May 26, 2026Updated last month
- Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audi…☆10,717May 16, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI☆1,384Jan 27, 2026Updated 5 months ago
- 🍃 MINT-1T: A one trillion token multimodal interleaved dataset.☆833Jul 31, 2024Updated last year
- Janus-Series: Unified Multimodal Understanding and Generation Models☆17,752Feb 1, 2025Updated last year
- An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image editing within a single powerfu…☆450Dec 2, 2025Updated 7 months ago
- Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models☆205May 29, 2024Updated 2 years ago
- Open-source watermark segmentation by DiffusionDynamics.ai and clear.photo. Harness deep learning plus synthetic data augmentation in PyT…☆77Apr 24, 2025Updated last year
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspace☆19Oct 21, 2024Updated last year
- The official implementation of V-AURA: Temporally Aligned Audio for Video with Autoregression (ICASSP 2025) (Oral)☆35Feb 11, 2026Updated 5 months ago
- 🔥ICLR 2025 (Spotlight) One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt☆318Oct 20, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Transductive regular expressions☆257Sep 25, 2025Updated 10 months ago
- [CVPR 2025] Magma: A Foundation Model for Multimodal AI Agents☆1,936Mar 3, 2026Updated 4 months ago
- Plugin Marketplace for Claude Code☆19Feb 8, 2026Updated 5 months ago
- Official repository for our work on micro-budget training of large-scale diffusion models.☆1,589Jan 12, 2025Updated last year
- Things you can do with the token embeddings of an LLM☆1,450Dec 1, 2025Updated 7 months ago
- code for "TVG: A Training-free Transition Video Generation Method with Diffusion Models"☆50Aug 19, 2024Updated last year
- Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation☆1,960Aug 15, 2024Updated last year