Benchmarking Intelligence Efficiency of LM Inference
☆85Aug 18, 2026Updated this week
Alternatives and similar repositories for intelligence-per-watt
Users that are interested in intelligence-per-watt are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 5 months ago
- Tinker ↔ KernelBench Integration enabling RL for GPU Kernel Generation☆31Mar 5, 2026Updated 5 months ago
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- ☆18Jan 19, 2026Updated 7 months ago
- Code for the paper "Greed is All You Need: An Evaluation of Tokenizer Inference Methods"☆15Nov 26, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆17Mar 21, 2026Updated 5 months ago
- FlashMemory DS-V4 Retriever: a lightweight retriever that sparsifies DeepSeek-V4 CSA KV-cache. Weights available on Hugging Face.☆107Jul 18, 2026Updated last month
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated last year
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 4 months ago
- Information about the CodedotAI reading group sessions.☆13Aug 16, 2021Updated 5 years ago
- ☆13Mar 28, 2022Updated 4 years ago
- Code for ICML 2025 paper | Joint Localization and Activation Editing for Low-Resource Fine-Tuning☆28Jun 18, 2025Updated last year
- ☆49Jan 3, 2026Updated 7 months ago
- Gallery repo for the pangeo-tutorial landsat-8 notebook on Pangeo Gallery http://gallery.pangeo.io/index.html☆13Nov 17, 2020Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆81Feb 18, 2026Updated 6 months ago
- Big & Small LLMs working together☆1,348Mar 12, 2026Updated 5 months ago
- ☆19Mar 29, 2026Updated 4 months ago
- Sequential Monte Carlo Speculative Decoding☆52Updated this week
- ☆17Feb 4, 2026Updated 6 months ago
- ☆29Sep 4, 2025Updated 11 months ago
- learning & making kernels in cuda / triton☆22Aug 24, 2025Updated last year
- ☆15Jan 24, 2025Updated last year
- Optimized GPU compiler for LLM inference. Choose from a list of optimized recipes or optimize your own model via kernel fusion, autotunin…☆79Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Shiny based data explorer with report templates based on field selection☆11Oct 27, 2015Updated 10 years ago
- Embedding Recycling for Language models☆38Jul 11, 2023Updated 3 years ago
- Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200☆121Feb 28, 2026Updated 5 months ago
- Burstable Cloud Scheduler☆17Jun 6, 2024Updated 2 years ago
- ☆57Mar 18, 2026Updated 5 months ago
- Data Parallel Programming, the 2022 edition☆12Jan 18, 2024Updated 2 years ago
- A curated biochemical database that integrates and refines data from KEGG and ATLAS databases to support precise analyses of biochemical …☆16May 13, 2026Updated 3 months ago
- Cascading filter modules for Shiny☆12Jul 16, 2020Updated 6 years ago
- Archon provides a modular framework for combining different inference-time techniques and LMs with just a JSON config file.☆212Mar 7, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆15Nov 19, 2025Updated 9 months ago
- [ICLR '26 W] Diagnosing Retrieval vs. Utilization Bottlenecks in LLM Agent Memory https://arxiv.org/abs/2603.02473☆17Mar 15, 2026Updated 5 months ago
- ☆15Jan 27, 2025Updated last year
- minimal Energy-based transformer☆44Dec 11, 2025Updated 8 months ago
- A benchmark to measure AI progress on unsolved research problems in mathematics.☆32Updated this week
- Asynchronous pipeline parallel optimization☆23Feb 2, 2026Updated 6 months ago
- Planning and collaboration hub for the CHAOSS AI Alignment Working Group. We develop metrics to evaluate how effectively AI systems respe…☆19Updated this week