Dynamic per-token early exit for LLM inference. Skip layers tokens don't need
☆33Sep 18, 2026Updated 2 weeks ago
Alternatives and similar repositories for TIDE
Users that are interested in TIDE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Run Qwen3.5-35B-A3B with 1 million token context on a single NVIDIA L4☆29Apr 9, 2026Updated 5 months ago
- Ultra-low-latency, high-throughput multiprocess transport over SHM and mmap. LMAX-Disruptor-style cross-process ring substrate.☆20Aug 6, 2026Updated last month
- Benchmarking Intelligence Efficiency of LM Inference☆95Sep 9, 2026Updated 3 weeks ago
- Evalution: evolve your LLMs with better evals.☆16Sep 26, 2026Updated last week
- OAuth Login for Gradio. Supports multiple identity providers.☆16Jul 6, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for TR2-D2: Tree Search Guided Trajectory-Aware Fine-Tuning for Discrete Diffusion☆24Sep 11, 2026Updated 3 weeks ago
- Rust TUI for the Lelit Mara espresso machine — runs on ESP32, streams telemetry to MQTT, built with Ratatui☆27Jun 28, 2026Updated 3 months ago
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆48Jan 13, 2026Updated 8 months ago
- llama2 inference engine in Rust☆13Apr 12, 2024Updated 2 years ago
- [ICASSP'22] Integer-only Zero-shot Quantization for Efficient Speech Recognition☆34Oct 11, 2021Updated 4 years ago
- AnyDSL traversal code☆16Feb 18, 2019Updated 7 years ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,577Mar 19, 2026Updated 6 months ago
- Sparse Autoencoders (SAE) vs CLIP fine-tuning fun.☆18Dec 19, 2024Updated last year
- A library of speech gadgets.☆16Oct 15, 2022Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Yet another useless cat here..☆23Jul 6, 2026Updated 2 months ago
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 11 months ago
- Kakao Mobility MCP Server for directions and transit information☆11Sep 14, 2025Updated last year
- TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained …☆854Aug 4, 2026Updated last month
- ☆307Apr 5, 2026Updated 5 months ago
- ☆15Jul 13, 2026Updated 2 months ago
- Self-contained Python lib with zero-dependencies that give you a unified device properties for gpu, cpu, and npu. No more calling separat…☆17Aug 23, 2026Updated last month
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆383Jul 9, 2026Updated 2 months ago
- Complete solution to enable RDMA (on both InfiniBand and RoCE) and accelerate TCP to bare metal performance on Kubernetes☆11Aug 1, 2018Updated 8 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A collection of GPU experiments and benchmarks for my personal understanding and research.☆43Aug 19, 2026Updated last month
- Example implementation of Iteration of Tought - Gives a star if you like the project☆41Dec 24, 2024Updated last year
- 삼각형의 실전! Triton☆16Feb 15, 2024Updated 2 years ago
- This project shows how to build a simple handwriting recognizer in Keras with the IAM dataset.☆13Aug 15, 2021Updated 5 years ago
- [WACV 2024] Enhancing Multimodal Compositional Reasoning of Visual Language Models with Generative Negative Mining, WACV 2024☆13Jan 3, 2024Updated 2 years ago
- Alphabot: a screen-less interactive spelling primer powered by computer vision☆14Sep 11, 2018Updated 8 years ago
- minimal Energy-based transformer☆44Dec 11, 2025Updated 9 months ago
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 5 months ago
- Alleycat plugin by devttys0, ported to IDA 8☆10Jan 15, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intellige…☆115Aug 18, 2026Updated last month
- bddem is a SWI-Prolog pack for using Binary Decision Diagrams☆13Sep 25, 2026Updated last week
- Inference code for LLaMA models☆42Mar 13, 2023Updated 3 years ago
- An unofficial implementation of SOLAR-10.7B model and the newly proposed interlocked-DUS(iDUS) implementation and experiment details.☆14Mar 20, 2024Updated 2 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 7 months ago
- Example OCaml library written using Rust and ocaml-rs☆17Mar 10, 2021Updated 5 years ago
- A relational logic programming language embedded in Rust.☆12Aug 15, 2025Updated last year