An open-weight 11B model series for long-form and real-time video understanding
☆562Sep 3, 2026Updated this week
Alternatives and similar repositories for MOSS-VL
Users that are interested in MOSS-VL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An 11B model for real-time video understanding with gated cross-attention and streaming inference☆169Jul 16, 2026Updated last month
- Official repository for the EMNLP 2025 paper “UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets”.☆16Sep 19, 2025Updated 11 months ago
- [EMNLP Findings'25] Official PyTorch Implementation of Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Align…☆16Sep 19, 2025Updated 11 months ago
- ☆146Updated this week
- A tool for better use of Inspire platform (Beta: Codeberg version is more up-to-date)☆30Apr 2, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICML 2026] Prism: Spectral-Aware Block-Sparse Attention☆27May 22, 2026Updated 3 months ago
- An asynchronous file-based handoff protocol for Claude Code and Codex agents sharing one project workspace☆37Jul 4, 2026Updated 2 months ago
- An open-source model for understanding speech, environmental sounds, and music through captioning, question answering, and reasoning☆654Updated this week
- We introduce 'Thinking with Video', a new paradigm leveraging video generation for multimodal reasoning. Our VideoThinkBench shows that S…☆320Aug 23, 2026Updated last week
- A Python library for building simple, modular, multifunctional, and efficient large model training data synthesis/augmentation pipelines.☆35May 29, 2026Updated 3 months ago
- OpenMOSS presents a collection of our research on LLMs, supported by SII, Fudan and Mosi.☆31Updated this week
- A foundation model that generates synchronized video and audio in a single model☆1,108Updated this week
- 一个面向启智平台(Inspire)的 awesome list☆38Mar 29, 2026Updated 5 months ago
- A multi-user OpenClaw deployment with isolated workspaces and configurable model backends☆25Mar 30, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness☆1,820Updated this week
- We propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervis…☆432Jul 22, 2026Updated last month
- An open-source personal academic homepage template characterized by its user-friendly design and extensive scalability.☆36Oct 6, 2025Updated 10 months ago
- A benchmark for evaluating future-event forecasting from audio and video context in multimodal language models☆29Jan 22, 2026Updated 7 months ago
- An end-to-end speech-to-speech language model that generates spoken responses without text guidance☆140Feb 13, 2026Updated 6 months ago
- A curated list of awesome resources about reward construction for AI agents. This repository covers cutting-edge research, and practical …☆61Sep 1, 2025Updated last year
- 启智平台的 Agent 驾驶舱:Skill + CLI,一条命令直达。Agent cockpit for the Inspire ML platform: one command, every operation, straight from chat.☆202Updated this week
- [ICML 2026] Sparser Block-Sparse Attention via Token Permutation☆33May 22, 2026Updated 3 months ago
- ☆25Jan 29, 2026Updated 7 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A 1.6B causal Transformer audio tokenizer with streaming, variable bitrates, and semantic alignment across speech, sound, and music☆254Jun 16, 2026Updated 2 months ago
- An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS☆4,061Updated this week
- A multi-device environment for evaluating intelligent vehicle interaction across connected cockpit systems☆25Sep 16, 2025Updated 11 months ago
- ☆23Mar 2, 2026Updated 6 months ago
- 启智平台任务管理 CLI:资源查询、任 务提交、日志查看和 MCP/agent workflow☆123Aug 27, 2026Updated last week
- Band-constrained Policy Optimization with probability-aware clipping for LLM reinforcement learning☆50Apr 8, 2026Updated 4 months ago
- Create SSH and TCP Proxy to your company container.☆30Jun 10, 2026Updated 2 months ago
- A multilingual model for long-form, multi-speaker dialogue synthesis with flexible speaker control and zero-shot voice cloning☆1,393Updated this week
- A curated list of models, benchmarks, tools and guides for audio editing☆44Aug 26, 2026Updated last week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Inference-time alignment for harmlessness through cross-model guidance (ACL 2024). Code + MM-Harmful Bench.☆38Oct 2, 2024Updated last year
- SocialClaw is a screen-aware social copilot that watches live chat windows, builds personalized memory and profile context, and suggests …☆41Apr 9, 2026Updated 4 months ago
- FamilyTool benchmark☆14Sep 10, 2025Updated 11 months ago
- A curated collection of papers, explainers, and resources on World Action Models for embodied AI☆1,375Updated this week
- A survey of long-context language models covering architecture, infrastructure, training, and evaluation☆63Mar 31, 2025Updated last year
- A proactive robot manipulation model for multimodal physical environments☆119Mar 28, 2026Updated 5 months ago
- A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation☆4,282Updated this week