My reasearch of losslessly compressing LLM weights.
☆63Jul 23, 2026Updated last month
Alternatives and similar repositories for weight-compression
Users that are interested in weight-compression are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Submodule of evalverse forked from [google-research/instruction_following_eval](https://github.com/google-research/google-research/tree/m…☆15May 4, 2024Updated 2 years ago
- Evalution: evolve your LLMs with better evals.☆16Updated this week
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish (EACL 2026, Offic…☆30Feb 27, 2026Updated 6 months ago
- Official Pytorch Implementation of MOCHI: Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence (CVPR …☆23Aug 24, 2026Updated 2 weeks ago
- Allows Codex on macOS to patch and extend its own desktop app.☆18Mar 24, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code to the paper: The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence☆36Jul 31, 2025Updated last year
- Demo of knowledge graph creation and Graph RAG with Dspy and Kuzu☆22Jun 30, 2025Updated last year
- A Framework For Intelligence Farming☆16Apr 3, 2025Updated last year
- FLK: a Low-Latency Biomechanics-Aware Filter for Real-time 3D Human Pose Estimation☆25Apr 7, 2025Updated last year
- ☆27Jul 13, 2026Updated 2 months ago
- ☆58Apr 25, 2026Updated 4 months ago
- Compile programs directly into transformer weights. Includes a 2D convex-hull KV cache with O(log n) inference.☆220Jun 1, 2026Updated 3 months ago
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 5 months ago
- Project code for training LLMs to write better unit tests + code☆22May 19, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Benchmarking long-horizon chain-of-thought reasoning.☆44Apr 20, 2026Updated 4 months ago
- ESLTTS dataset☆16Feb 6, 2025Updated last year
- A tiny ~10K-parameter LLM router that learns which open-source model (deepseek-v4-pro / glm-5p2 / kimi-k2p6 via Fireworks) should answer …☆315Jul 6, 2026Updated 2 months ago
- Iterative specification refinement tool: feeds your docs through GPT Pro Extended Reasoning via Oracle for multiple revision rounds until…☆68Aug 25, 2026Updated 2 weeks ago
- Official implementation for "MOVIN: Real-time Motion Capture using a Single LiDAR"☆20Oct 24, 2024Updated last year
- Heterogeneous prefill/decode for DeepSeek-V4-Flash: CUDA prefill (DGX Spark, vLLM) -> Metal decode (Mac Studio, oMLX) over plain 10GbE☆110Updated this week
- SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB G…☆212Sep 1, 2026Updated last week
- ☆11Oct 31, 2021Updated 4 years ago
- A repo on resources for training agents.☆103Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- KAN (Kolmogorov–Arnold Networks) in the MLX framework for Apple Silicon☆32Jun 18, 2025Updated last year
- Self-contained Python lib with zero-dependencies that give you a unified device properties for gpu, cpu, and npu. No more calling separat…☆16Aug 23, 2026Updated 3 weeks ago
- MCP server for intelligent handling of large files — smart chunking, search, navigation, and streaming.☆19Aug 4, 2026Updated last month
- A python engine for playing dnd 5e☆24Updated this week
- ELECTRA MODEL NLP☆13Apr 8, 2020Updated 6 years ago
- PyTorch - CHIMLE, an IMLE-based general-purpose multimodal conditional image synthesis method [NeurIPS 2022]☆12Jul 8, 2026Updated 2 months ago
- Evaluation repository of wikipedia index with Dria☆10Mar 14, 2024Updated 2 years ago
- ☆20Jun 17, 2024Updated 2 years ago
- Official Code Release for Pix2NPHM☆68Dec 22, 2025Updated 8 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- 记录Transformer升级的论文笔记☆19Jun 25, 2023Updated 3 years ago
- ☆17May 21, 2020Updated 6 years ago
- All-in-one environment to use Dria, the collective knowledge for AI.☆14Mar 15, 2024Updated 2 years ago
- ☆37Mar 30, 2026Updated 5 months ago
- This is the CUDA GPU implementation + Python interface (using PyTorch) of DCI. The paper can be found at https://arxiv.org/abs/1512.00442…☆13Dec 20, 2023Updated 2 years ago
- ☆21Apr 17, 2026Updated 4 months ago
- The official implementation of DMEL the method presented in the paper "DMEL: The differentiable log-Mel spectrogram as a trainable layer …☆24Dec 21, 2024Updated last year