Dynamic per-token early exit for LLM inference. Skip layers tokens don't need
☆33Mar 18, 2026Updated 4 months ago
Alternatives and similar repositories for TIDE
Users that are interested in TIDE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Run Qwen3.5-35B-A3B with 1 million token context on a single NVIDIA L4☆28Apr 9, 2026Updated 3 months ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆17Jul 23, 2026Updated last week
- Benchmarking Intelligence Efficiency of LM Inference☆73Updated this week
- SGLang kernel library for Intel XPU☆27Updated this week
- Code for TR2-D2: Tree Search Guided Trajectory-Aware Fine-Tuning for Discrete Diffusion☆22Feb 12, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Rust TUI for the Lelit Mara espresso machine — runs on ESP32, streams telemetry to MQTT, built with Ratatui☆26Jun 28, 2026Updated last month
- Jupyterlab extension containing a UI for debugging☆10Dec 2, 2019Updated 6 years ago
- Sequential Monte Carlo Speculative Decoding☆52Jul 25, 2026Updated last week
- merge two sorted lists fast☆12Nov 15, 2023Updated 2 years ago
- Full-stack Amazon-inspired e-commerce platform built with Next.js, TypeScript, Prisma, MongoDB, and Stripe.☆20Jun 2, 2026Updated 2 months ago
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention☆23May 25, 2026Updated 2 months ago
- ☆10Nov 22, 2023Updated 2 years ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,496Mar 19, 2026Updated 4 months ago
- A simple script and howto for censoring photos for publication without faces☆15May 23, 2019Updated 7 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 9 months ago
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆836Jul 14, 2026Updated 3 weeks ago
- Self-contained Python lib with zero-dependencies that give you a unified device properties for gpu, cpu, and npu. No more calling separat…☆16Jul 21, 2026Updated 2 weeks ago
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆370Jul 9, 2026Updated 3 weeks ago
- A collection of GPU experiments and benchmarks for my personal understanding and research.☆36Jul 22, 2026Updated last week
- Example implementation of Iteration of Tought - Gives a star if you like the project☆41Dec 24, 2024Updated last year
- Large language model of Medical AI, General Medical AI (GMAI)☆17Jan 30, 2024Updated 2 years ago
- This project shows how to build a simple handwriting recognizer in Keras with the IAM dataset.☆13Aug 15, 2021Updated 4 years ago
- [WACV 2024] Enhancing Multimodal Compositional Reasoning of Visual Language Models with Generative Negative Mining, WACV 2024☆13Jan 3, 2024Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 3 months ago
- minimal Energy-based transformer☆44Dec 11, 2025Updated 7 months ago
- Text Classification Dataset for Turkish Language☆10Nov 16, 2021Updated 4 years ago
- Alleycat plugin by devttys0, ported to IDA 8☆10Jan 15, 2025Updated last year
- ✅ Agentic OS-aware Intention Programming Technology☆22Jul 28, 2026Updated last week
- Pokémon damage calculator☆14Feb 7, 2024Updated 2 years ago
- A quick way to get started with Transformer Lens☆14Dec 13, 2023Updated 2 years ago
- Inference code for LLaMA models☆42Mar 13, 2023Updated 3 years ago
- An unofficial implementation of SOLAR-10.7B model and the newly proposed interlocked-DUS(iDUS) implementation and experiment details.☆14Mar 20, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 5 months ago
- Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.☆18Jul 27, 2026Updated last week
- Example OCaml library written using Rust and ocaml-rs☆17Mar 10, 2021Updated 5 years ago
- ☆18Dec 8, 2025Updated 7 months ago
- My reasearch of losslessly compressing LLM weights.☆62Jul 23, 2026Updated last week
- 💧 Query mode for agents☆104May 11, 2026Updated 2 months ago
- Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment☆62Aug 30, 2024Updated last year