Benchmarking Intelligence Efficiency of LM Inference
☆87Sep 9, 2026Updated this week
Alternatives and similar repositories for intelligence-per-watt
Users that are interested in intelligence-per-watt are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 5 months ago
- Tinker ↔ KernelBench Integration enabling RL for GPU Kernel Generation☆31Mar 5, 2026Updated 6 months ago
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- ☆19Jan 19, 2026Updated 7 months ago
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆23May 25, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code for the paper "Greed is All You Need: An Evaluation of Tokenizer Inference Methods"☆15Sep 2, 2026Updated last week
- FlashMemory DS-V4 Retriever: a lightweight retriever that sparsifies DeepSeek-V4 CSA KV-cache. Weights available on Hugging Face.☆110Jul 18, 2026Updated last month
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated last year
- Official implementation of "Data Mixture Inference: What do BPE tokenizers reveal about their training data?"☆23May 15, 2025Updated last year
- Gallery repo for the pangeo-tutorial landsat-8 notebook on Pangeo Gallery http://gallery.pangeo.io/index.html☆13Nov 17, 2020Updated 5 years ago
- Information about the CodedotAI reading group sessions.☆13Aug 16, 2021Updated 5 years ago
- Code for ICML 2025 paper | Joint Localization and Activation Editing for Low-Resource Fine-Tuning☆28Jun 18, 2025Updated last year
- ☆49Jan 3, 2026Updated 8 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆82Feb 18, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code repository for "Eliciting Secret Knowledge from Language Models"☆24Mar 30, 2026Updated 5 months ago
- ☆19Mar 29, 2026Updated 5 months ago
- ☆26Mar 26, 2026Updated 5 months ago
- ☆16Feb 4, 2026Updated 7 months ago
- Implementation of ModernBERT in MLX☆21Jan 7, 2026Updated 8 months ago
- learning & making kernels in cuda / triton☆22Aug 24, 2025Updated last year
- ☆32Sep 4, 2025Updated last year
- ☆15Jan 24, 2025Updated last year
- a fast and lightweight distributed background task processing framework with seamless scheduling.☆15Mar 30, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ScaleRL Curve Fitting☆17Oct 13, 2025Updated 10 months ago
- Optimized GPU compiler for LLM inference. Choose from a list of optimized recipes or optimize your own model via kernel fusion, autotunin…☆80Updated this week
- Shiny based data explorer with report templates based on field selection☆11Oct 27, 2015Updated 10 years ago
- ☆80Apr 26, 2026Updated 4 months ago
- Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200☆123Feb 28, 2026Updated 6 months ago
- ☆57Mar 18, 2026Updated 5 months ago
- Data Parallel Programming, the 2022 edition☆12Jan 18, 2024Updated 2 years ago
- Pixels, Patterns, but no Poetry: To See the World like Humans☆18Aug 11, 2025Updated last year
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆19Jul 4, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 🏆 A decentralized layer to support NFT on Mixin Messenger and Kernel.☆17Feb 19, 2026Updated 6 months ago
- Cascading filter modules for Shiny☆12Jul 16, 2020Updated 6 years ago
- Archon provides a modular framework for combining different inference-time techniques and LMs with just a JSON config file.☆216Mar 7, 2025Updated last year
- ☆15Nov 19, 2025Updated 9 months ago
- [ICLR '26 W] Diagnosing Retrieval vs. Utilization Bottlenecks in LLM Agent Memory https://arxiv.org/abs/2603.02473☆17Updated this week
- ☆15Jan 27, 2025Updated last year
- 🕶 Cross-platform network interface command-line utility.☆18Jan 23, 2023Updated 3 years ago