Implementation of Qformer from BLIP2 in Zeta Lego blocks.
β51Nov 11, 2024Updated last year
Alternatives and similar repositories for qformer
Users that are interested in qformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- a suite of finetuned LLMs for atomically precise function calling π§ͺβ16Updated this week
- Implementation of a Hierarchical Mamba as described in the paper: "Hierarchical State Space Models for Continuous Sequence-to-Sequence Moβ¦β16Nov 11, 2024Updated last year
- Multi-threading, Concurrency, Asynchrony, and various Execution Methods implemented in a Rust backend for bleeding edge performance.β20Nov 11, 2024Updated last year
- The open source implementation of the cross attention mechanism from the paper: "JOINTLY TRAINING LARGE AUTOREGRESSIVE MULTIMODAL MODELS"β37Mar 11, 2024Updated 2 years ago
- Implementation of the model: "(MC-ViT)" from the paper: "Memory Consolidation Enables Long-Context Video Understanding"β27Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- The Swarm Ecosystemβ29Aug 1, 2024Updated last year
- Implementation of the LDP module block in PyTorch and Zeta from the paper: "MobileVLM: A Fast, Strong and Open Vision Language Assistant β¦β15Mar 11, 2024Updated 2 years ago
- An open source replication of the stawberry method that leverages Monte Carlo Search with PPO and or DPOβ30Updated this week
- Pytorch Implementation of Deepmind's SIMA: "Scaling Instructable Agents Across Many Simulated Worlds"β35Jun 17, 2024Updated 2 years ago
- Implementation of SoundtStream from the paper: "SoundStream: An End-to-End Neural Audio Codec"β13Jan 27, 2025Updated last year
- Repository of the IJCV'26 & WACV'24 paperβ34Apr 27, 2026Updated 2 months ago
- QuickSplat: Fast 3D Surface Reconstruction via Learned Gaussian Initializationβ25Nov 11, 2025Updated 8 months ago
- Implementation of "PaLM2-VAdapter:" from the multi-modal model paper: "PaLM2-VAdapter: Progressively Aligned Language Model Makes a Stronβ¦β17Nov 11, 2024Updated last year
- Implementation of the Pairformer model used in AlphaFold 3β14Jul 13, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Deploy your autonomous agents to production grade environments with 99% Uptime Guarantee, Infinite Scalability, and self-healing.β54Jul 13, 2026Updated last week
- Community Implementation of the paper: "Multi-Head Mixture-of-Experts" In PyTorchβ30Updated this week
- β48Jul 7, 2025Updated last year
- Generative Expressive Conversational Speech Synthesis (Accepted by MM'2024)β61Nov 1, 2024Updated last year
- An unofficial implementation of "UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding".β26Nov 4, 2023Updated 2 years ago
- π Official pytorch implementation of "D2ADA: Dynamic Density-aware Active Domain Adaptation for Semantic Segmentation. Wu et al. ECCV 20β¦β25Feb 2, 2023Updated 3 years ago
- β15Updated this week
- β10Sep 25, 2019Updated 6 years ago
- [NeurIPS 2023] Bootstrapping Vision-Language Learning with Decoupled Language Pre-trainingβ26Dec 5, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Implementation of VisionLLaMA from the paper: "VisionLLaMA: A Unified LLaMA Interface for Vision Tasks" in PyTorch and Zetaβ15Nov 11, 2024Updated last year
- Ultra-low bitrate neural audio codec (0.31~1.40 kbps) with a better semantic in the latent space.β254Mar 7, 2025Updated last year
- Compute WER and SER for speech recognition evaluationβ27Jun 6, 2026Updated last month
- Official respository for ReasonGen-R1β75Jun 23, 2025Updated last year
- An neural full-band audio codec for general audio sampled at 48 kHz with 7.5 kps or 4.5 kbps.β212Jun 22, 2026Updated 3 weeks ago
- Simple Implementation of a Transformer in the new framework MLX by Appleβ19Nov 18, 2024Updated last year
- β17Dec 18, 2023Updated 2 years ago
- β44May 20, 2025Updated last year
- Source code for the EMNLP 2025 paper βDM-Codec: Distilling Multimodal Representations for Speech Tokenizationββ57Jun 1, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Speech-To-Text forced-alignment Speech processing Universal PERformance Benchmarkβ39May 7, 2025Updated last year
- β40Aug 26, 2025Updated 10 months ago
- JATTS: A modern, research-oriented Japanese Text-to-speech Open-sourced Toolkitβ43Mar 13, 2026Updated 4 months ago
- (ICCV2025) Official repository of paper "ViSpeak: Visual Instruction Feedback in Streaming Videos"β53Jul 1, 2025Updated last year
- β10Mar 18, 2025Updated last year
- A single-layer, streaming codec model providing SOTA audio quality and discrete tokens designed for superior downstream modelability.β124Jun 4, 2025Updated last year
- A python algorithm to change the pitch of the voice in real timeβ13Dec 13, 2020Updated 5 years ago