Benchmarking Intelligence Efficiency of LM Inference
☆72Jul 17, 2026Updated last week
Alternatives and similar repositories for intelligence-per-watt
Users that are interested in intelligence-per-watt are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Dynamic per-token early exit for LLM inference. Skip layers tokens don't need☆33Mar 18, 2026Updated 4 months ago
- Tinker ↔ KernelBench Integration enabling RL for GPU Kernel Generation☆29Mar 5, 2026Updated 4 months ago
- ☆18Jan 19, 2026Updated 6 months ago
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆22May 25, 2026Updated 2 months ago
- ☆20Jul 5, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆17Mar 21, 2026Updated 4 months ago
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 3 months ago
- Code for ICML 2025 paper | Joint Localization and Activation Editing for Low-Resource Fine-Tuning☆28Jun 18, 2025Updated last year
- ☆49Jan 3, 2026Updated 6 months ago
- Gallery repo for the pangeo-tutorial landsat-8 notebook on Pangeo Gallery http://gallery.pangeo.io/index.html☆13Nov 17, 2020Updated 5 years ago
- GPU-accelerated Schulze voting method in Python, Numba, CUDA, and Mojo 🔥, using ideas from Algebraic Graph Theory☆19Oct 28, 2025Updated 9 months ago
- A comprehensive toolkit for streamlining data editing, search, and inspection for large-scale language model training and interpretabilit…☆21Oct 30, 2025Updated 8 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆79Feb 18, 2026Updated 5 months ago
- Code repository for "Eliciting Secret Knowledge from Language Models"☆24Mar 30, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Big & Small LLMs working together☆1,342Mar 12, 2026Updated 4 months ago
- ☆19Mar 29, 2026Updated 3 months ago
- ☆25Mar 26, 2026Updated 4 months ago
- Sequential Monte Carlo Speculative Decoding☆52Updated this week
- ☆16Feb 4, 2026Updated 5 months ago
- Implementation of ModernBERT in MLX☆21Jan 7, 2026Updated 6 months ago
- ☆26Sep 4, 2025Updated 10 months ago
- learning & making kernels in cuda / triton☆22Aug 24, 2025Updated 11 months ago
- a fast and lightweight distributed background task processing framework with seamless scheduling.☆15Mar 30, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ScaleRL Curve Fitting☆17Oct 13, 2025Updated 9 months ago
- Benchmark and deploy optimized LLM models on GPU servers with vLLM or SGLang. Chose from a list of optimized recipes for popular models o…☆67Updated this week
- Learning stochastic dynamics from snapshots through regularized unbalanced optimal transport (ICLR'25 oral)☆35Dec 10, 2025Updated 7 months ago
- Shiny based data explorer with report templates based on field selection☆11Oct 27, 2015Updated 10 years ago
- ☆76Apr 26, 2026Updated 3 months ago
- Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200☆121Feb 28, 2026Updated 5 months ago
- Burstable Cloud Scheduler☆17Jun 6, 2024Updated 2 years ago
- ☆57Mar 18, 2026Updated 4 months ago
- A CLI in Rust to generate synthetic data for MLX friendly training☆26Jan 13, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A curated biochemical database that integrates and refines data from KEGG and ATLAS databases to support precise analyses of biochemical …☆16May 13, 2026Updated 2 months ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆18Jul 4, 2025Updated last year
- Cascading filter modules for Shiny☆12Jul 16, 2020Updated 6 years ago
- Archon provides a modular framework for combining different inference-time techniques and LMs with just a JSON config file.☆207Mar 7, 2025Updated last year
- ☆15Nov 19, 2025Updated 8 months ago
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics☆26Jan 30, 2026Updated 5 months ago
- An OpenAI API Compatible Honeypot Gateway☆26Mar 17, 2025Updated last year