3x Faster Inference; Unofficial implementation of EAGLE Speculative Decoding
☆85Jul 3, 2025Updated last year
Alternatives and similar repositories for BaldEagle
Users that are interested in BaldEagle are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)☆56Mar 14, 2025Updated last year
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,537Feb 20, 2026Updated 7 months ago
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.☆1,178Updated this week
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆413Apr 22, 2025Updated last year
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆54May 12, 2026Updated 4 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ScribePal is an Open Source intelligent browser extension that leverages AI to empower your web experience by providing contextual insigh…☆22Sep 13, 2026Updated last week
- Fine-tuned LLaMa2 13B model designed for ReAct-style and Tree-Of-Thoughts style prompting.☆16Jul 23, 2023Updated 3 years ago
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- A fast, local, and secure approach for training LLMs for coding tasks using GRPO with WebAssembly and interpreter feedback.☆42Apr 4, 2025Updated last year
- A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM☆847Updated this week
- This repository is the official implementation of "Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE" [ACL 2026 Mai…☆38Oct 5, 2025Updated 11 months ago
- QQQ is an innovative and hardware-optimized W4A8 quantization solution for LLMs.☆160Aug 21, 2025Updated last year
- an auto-sleeping and -waking framework around llama.cpp☆13Feb 8, 2025Updated last year
- Code for "Domain Adaptive Meta-learning for Dialogue State Tracking"(TASLP2021)☆10Sep 14, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- GUI for LLaDA Diffusion LLM with Quantization for low end GPU and CPU options.☆24Mar 7, 2025Updated last year
- official code for GliDe with a CaPE☆21Aug 13, 2024Updated 2 years ago
- Single-file, pure CUDA C implementation for running inference on Qwen3 0.6B GGUF. No Dependencies.☆25Nov 26, 2025Updated 9 months ago
- Python package for compressing floating-point PyTorch tensors☆14Jul 22, 2024Updated 2 years ago
- RestoreChitra AI. An open-source photo restoration tool. Restore your old, blurry face photos with AI Magic. 100% free, no catch.☆10Nov 7, 2023Updated 2 years ago
- Automated Identification of Redundant Layer Blocks for Pruning in Large Language Models☆271Apr 23, 2024Updated 2 years ago
- win32 native frontend for llama-cli☆14Nov 2, 2024Updated last year
- Code for "Retaining Key Information under High Compression Rates: Query-Guided Compressor for LLMs" (ACL 2024)☆19Jun 12, 2024Updated 2 years ago
- 🕹️ Performance Comparison of MLOps Engines, Frameworks, and Languages on Mainstream AI Models.☆142Jul 25, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- BitTorrent Data Set☆12Jan 2, 2025Updated last year
- Official evaluation models and configuration for Stage 1 of the SAIR Mathematics Distillation Challenge: Equational Theories.☆19Apr 19, 2026Updated 5 months ago
- ☆171Feb 15, 2025Updated last year
- Official code for the paper "Examining Post-Training Quantization for Mixture-of-Experts: A Benchmark"☆31Jun 30, 2025Updated last year
- python-v4l2 fork☆14Sep 20, 2021Updated 5 years ago
- Distributed Reinforcement Learning for LLM Fine-Tuning with multi-GPU utilization☆22Mar 12, 2025Updated last year
- A curated collection of resources for building “AI Scientist” systems: AI that assists scientific discovery through literature intelligen…☆18Aug 4, 2026Updated last month
- "a towel is about the most massively useful thing an interstellar AI hitchhiker can have"☆48Oct 9, 2024Updated last year
- Code for "Context-Aware Recurrent Encoder for Neural Machine Translation" (TASLP 2017)☆12Oct 29, 2018Updated 7 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A single repo with all scripts and utils to train / fine-tune the Mamba model with or without FIM☆63Apr 8, 2024Updated 2 years ago
- ☆16Dec 16, 2024Updated last year
- Statically configured API Gateway framework☆14Jun 25, 2022Updated 4 years ago
- TKDE'20 paper, "CONNA: Addressing Name Disambiguation on the Fly".☆15May 31, 2021Updated 5 years ago
- Resources and code for paper "Multiplingual Knowledge Graph Completion via Ensemble Knowledge Transfer"☆15Sep 15, 2021Updated 5 years ago
- Autonomous GPU kernel optimization system driven by AI agents.☆31Mar 29, 2026Updated 5 months ago
- Benchmarking Social Intelligence of Language Agents through Interactive Scenarios☆12Jan 4, 2025Updated last year