First Latency-Aware Competitive LLM Agent Benchmark
☆32Jun 3, 2025Updated last year
Alternatives and similar repositories for LatencySensitiveBench
Users that are interested in LatencySensitiveBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Coco is a proactive co-assistant that connects user workspace with a broader ecosystem of AI agents.☆38Sep 23, 2026Updated last week
- Beyond KV Caching: Shared Attention for Efficient LLMs☆20Jul 19, 2024Updated 2 years ago
- PyTorch compilation tutorial covering TorchScript, torch.fx, and Slapo☆17Mar 13, 2023Updated 3 years ago
- Multi-agent system for resolving Site Reliability Engineering tasks.☆27Oct 23, 2025Updated 11 months ago
- Official code, models, and dataset for "Evolution Fine-Tuning (EFT): Learning to Discover Across 371 Optimization Tasks"☆30Jun 30, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Streaming Diffusion Policy: Fast Policy Synthesis with Variable Noise Diffusion Models☆80May 14, 2025Updated last year
- PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training☆22Jun 12, 2024Updated 2 years ago
- World Model & VLA Survey - Interactive Research Page☆21May 26, 2026Updated 4 months ago
- NeurIPS 2024: Few-Shot Task Learning through Inverse Generative Modeling☆17Dec 6, 2024Updated last year
- Simple example of how to write an Implicit GEMM Convolution in CUDA using the tensor core WMMA API and bindings for PyTorch.☆19Jun 29, 2023Updated 3 years ago
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization (ICLR 2021)☆41Jan 12, 2021Updated 5 years ago
- Private Adaptive Optimization with Side Information (ICML '22)☆16Jun 23, 2022Updated 4 years ago
- The official code for [ECCV2020] "HALO: Hardware-aware Learning to Optimize"☆10Mar 22, 2023Updated 3 years ago
- Code for paper Geometry-aware Policy Imitation☆44Oct 13, 2025Updated 11 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers☆33Mar 1, 2025Updated last year
- ☆15Jul 13, 2021Updated 5 years ago
- [NeurIPS 2023] ShiftAddViT: Mixture of Multiplication Primitives Towards Efficient Vision Transformer☆31Dec 6, 2023Updated 2 years ago
- [NSDI'26] PolyRL is a reinforcement learning framework for LLM that harvest spot instances on the cloud to reduce cost.☆19Mar 30, 2026Updated 6 months ago
- ☆31Sep 26, 2025Updated last year
- ☆11Sep 7, 2023Updated 3 years ago
- Scripts to prepare OXFORD VGG Face dataset☆12Mar 29, 2016Updated 10 years ago
- The NYU Systems Seminar☆25Feb 26, 2024Updated 2 years ago
- FROM $f(x)$ AND $g(x)$ TO $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones☆73Jan 26, 2026Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆235Jun 11, 2024Updated 2 years ago
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. 🚀 The official implementation of https://arx…☆33Feb 17, 2025Updated last year
- [ICML 2022] ShiftAddNAS: Hardware-Inspired Search for More Accurate and Efficient Neural Networks☆15May 18, 2022Updated 4 years ago
- A Unreal Engine 5 (UE5) based plugin aiming to provide real-time visulization, management, editing, and scalable hybrid rendering of Guas…☆14Feb 13, 2026Updated 7 months ago
- CUDA 8-bit Tensor Core Matrix Multiplication based on m16n16k16 WMMA API☆38Sep 15, 2023Updated 3 years ago
- The tool facilitates debugging convergence issues and testing new algorithms and recipes for training LLMs using Nvidia libraries such as…☆23Sep 17, 2025Updated last year
- [NeurIPS 2022] "Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Spee…☆17Sep 19, 2023Updated 3 years ago
- A curated list for Efficient Large Language Models☆11Mar 25, 2024Updated 2 years ago
- [NeurIPS2025]VideoVLA: Video Generators Can Be Generalizable Robot Manipulators☆32Jun 26, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Dev repo for power measurement for the MLPerf™ benchmarks☆28Sep 11, 2025Updated last year
- ☆29Jul 2, 2026Updated 3 months ago
- Systolic array based hardware for Image processing on the SPARTAN-6 FPGA☆13May 26, 2016Updated 10 years ago
- [ICML 2024] Sparse Model Inversion: Efficient Inversion of Vision Transformers with Less Hallucination☆15Apr 29, 2025Updated last year
- Initial commit☆13Aug 14, 2023Updated 3 years ago
- My solution code to parallel architecture and programming Spring 2016☆12Aug 15, 2016Updated 10 years ago
- A holistic framework to enable the design, development, and evaluation of autonomous AIOps agents.☆11May 21, 2025Updated last year