⛔ DEPRECATED -- use flash-head instead (pip install flash-head)
☆29Apr 10, 2026Updated 4 months ago
Alternatives and similar repositories for embedl-models
Users that are interested in embedl-models are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- QLoRA: Efficient Finetuning of Quantized LLMs☆11Jul 22, 2023Updated 3 years ago
- ☆15Dec 4, 2025Updated 8 months ago
- ☆35Feb 6, 2026Updated 6 months ago
- ☆18Jul 1, 2025Updated last year
- A PyTorch implementation of gradient-free optimization for directly optimizing NDCG (Normalized Discounted Cumulative Gain) in neural inf…☆19Dec 21, 2025Updated 7 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks☆31May 22, 2024Updated 2 years ago
- A repository to store helpful information and emerging insights in regard to LLMs☆21Oct 27, 2023Updated 2 years ago
- ☆18Jul 12, 2025Updated last year
- [ICLR 2026] StableToken: A state-of-the-art noise-robust semantic speech tokenizer featuring Voting-LFQ for resilient SpeechLLMs.☆33Feb 27, 2026Updated 5 months ago
- Compose, manage, and run MCP servers as Docker containers. With a Unified API gateway built in.☆57Oct 9, 2025Updated 10 months ago
- ☆11Nov 12, 2024Updated last year
- Separation of planning concerns in ReAct-style LLM agents. Planner fine-tuning on synthetic trajectories.☆19Jul 28, 2024Updated 2 years ago
- This is the official repository for the paper "Flora: Low-Rank Adapters Are Secretly Gradient Compressors" in ICML 2024.☆108Jul 1, 2024Updated 2 years ago
- Python library to compress LitGPT models for resource efficient inference.☆16Jul 31, 2026Updated last week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Evalution: evolve your LLMs with better evals.☆16Updated this week
- Code for Fast as CHITA: Neural Network Pruning with Combinatorial Optimization☆14Aug 2, 2023Updated 3 years ago
- Learnable Semi-structured Sparsity for Vision Transformers and Diffusion Transformers☆15Feb 7, 2025Updated last year
- Incremental Learning with Adaptive Resonance Theory (ART) & Developmental Resonance networks☆12Dec 18, 2019Updated 6 years ago
- [deprecated] reference code for string segmentation using LSTM(tensorflow)☆19Feb 19, 2020Updated 6 years ago
- This project is a versatile and powerful search tool that leverages state-of-the-art natural language processing models to provide releva…☆12Apr 3, 2023Updated 3 years ago
- ☆34Mar 24, 2026Updated 4 months ago
- Precision Knowledge Editing (PKE): A novel method to reduce toxicity in LLMs while preserving performance, with robust evaluations and ha…☆12Nov 26, 2024Updated last year
- MNIST in the browser☆21Apr 27, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Second Brain is a desktop application that acts as a personal knowledge base, using retrieval-augmented generation (RAG), multimodal AI m…☆27Jan 30, 2026Updated 6 months ago
- [ICML 2026 Spotlight] Official implementation of TetraJet-v2: Accurate NVFP4 Training for LLMs, with fully-NVFP4 linear layer with unbias…☆16Jul 3, 2026Updated last month
- ☆14Mar 19, 2025Updated last year
- hacky fixes for reSOLume☆13Dec 20, 2021Updated 4 years ago
- Multilingual Entity Linking model by BELA model☆12Jul 20, 2023Updated 3 years ago
- 5X faster 60% less memory QLoRA finetuning☆21May 28, 2024Updated 2 years ago
- Successor to Annoy https://github.com/spotify/annoy☆13Oct 28, 2015Updated 10 years ago
- A very simple launcher for a certain factory building game☆17Mar 17, 2025Updated last year
- Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks (EMNLP'24)☆144Sep 20, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Source codes for our paper "Neural Temporality Adaptation for Document Classification: Diachronic Word Embeddings and Domain Adaptation M…☆12Apr 20, 2021Updated 5 years ago
- Terminal 3D Model Viewer Made In Typescript.☆24Nov 12, 2024Updated last year
- Super simple python connectors for llama.cpp, including vision models (Gemma 3, Qwen2-VL). Compile llama.cpp and run!☆31Dec 11, 2025Updated 7 months ago
- Code repository for the paper on "Predicting the Performance of Black-Box LLMs through Self-Queries".☆12Jan 9, 2025Updated last year
- 33B Chinese LLM, DPO QLORA, 100K context, AirLLM 70B inference with single 4GB GPU☆14May 5, 2024Updated 2 years ago
- Official implementation of BitMamba-2. A scalable 1.58-bit State Space Model (Mamba-2 + BitNet) trained from scratch on 150B tokens. Incl…☆26Feb 4, 2026Updated 6 months ago
- AI Intelligence Platform -- knowledge graph + semantic search + reasoning + multi-agent debate in a single binary☆18Apr 10, 2026Updated 4 months ago