Lightweight toolkit package to train and fine-tune 1.58bit Language models
☆146Apr 30, 2026Updated 2 months ago
Alternatives and similar repositories for onebitllms
Users that are interested in onebitllms are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Thin wrapper around GGML to make life easier☆48Updated this week
- All information and news with respect to Falcon-H1 series☆122Oct 9, 2025Updated 9 months ago
- Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking☆25Apr 4, 2025Updated last year
- YASEM - Yet Another Splade|Sparse Embedder - A simple and efficient library for SPLADE embeddings☆13May 22, 2025Updated last year
- An LLM Client for the PS Vita☆13Jun 23, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Evalution: evolve your LLMs with better evals.☆16Updated this week
- BitLinear implementation☆38Jul 8, 2026Updated 3 weeks ago
- Starbucks: Improved Training for 2D Matryoshka Embeddings☆25Jun 30, 2025Updated last year
- Latent Large Language Models☆19Aug 24, 2024Updated last year
- An implementation is provided here for the NeurIPS2024 paper "MemoryFormer : Minimize Transformer Computation by Removing Fully-Connected…☆16Mar 24, 2026Updated 4 months ago
- Efficient non-uniform quantization with GPTQ for GGUF☆64Sep 17, 2025Updated 10 months ago
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- Sparse Embedding Compression for Scalable Retrieval in Recommender Systems☆39Nov 21, 2025Updated 8 months ago
- Trully flash implementation of DeBERTa disentangled attention mechanism.☆90Feb 10, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Linear Relational Embeddings (LREs) and Linear Relational Concepts (LRCs) for LLMs in PyTorch☆11Aug 7, 2024Updated last year
- Minimalistic large language model 3D-parallelism training☆2,768May 26, 2026Updated 2 months ago
- ☆16Dec 11, 2025Updated 7 months ago
- We release Open Meditron, a fully open, clinician-audited medical training corpus and evaluation protocol that closes the open-vs-closed …☆15May 15, 2026Updated 2 months ago
- Pre-train Static Word Embeddings☆109Jun 9, 2026Updated last month
- Official implementation of Half-Quadratic Quantization (HQQ)☆948Feb 26, 2026Updated 5 months ago
- Official Repository for "Hypencoder: Hypernetworks for Information Retrieval"☆41Sep 20, 2025Updated 10 months ago
- DPO, but faster 🚀☆52Dec 6, 2024Updated last year
- Python library to use Pleias-RAG models☆72Jul 1, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This repository contains code for the MicroAdam paper.☆21Dec 14, 2024Updated last year
- An Open Source Toolkit For LLM Distillation☆992May 12, 2026Updated 2 months ago
- Repo hosting codes and materials related to speeding LLMs' inference using token merging.☆37Oct 9, 2025Updated 9 months ago
- A pipeline parallel training script for LLMs.☆167Apr 30, 2025Updated last year
- An efficent implementation of the method proposed in "The Era of 1-bit LLMs"☆155Oct 15, 2024Updated last year
- The codebase and database of KuaiSearch: A Large-Scale E-Commerce Search Dataset for Recall, Ranking, and Relevance☆24Updated this week
- AI Edge Quantizer: flexible post training quantization for LiteRT models.☆185Updated this week
- Estimate MFU for DeepSeekV3☆26Jan 5, 2025Updated last year
- EnriCo: Enriched Representation and Globally Constrained Inference for Entity and Relation Extraction☆26May 22, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆63Jul 10, 2025Updated last year
- UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation☆48Aug 26, 2025Updated 11 months ago
- Fast, Modern, and Low Precision PyTorch Optimizers☆129May 16, 2026Updated 2 months ago
- Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients.☆206Jul 17, 2024Updated 2 years ago
- Build LLM Application with Local Documents☆20Jun 13, 2025Updated last year
- ☆12May 20, 2025Updated last year
- [ICLR 2025 & COLM 2025] Official PyTorch implementation of the Forgetting Transformer and Adaptive Computation Pruning☆150Feb 25, 2026Updated 5 months ago