Benchmarking Open-Ended Inference Optimization by AI Agents
β41Jul 6, 2026Updated last month
Alternatives and similar repositories for InferenceBench
Users that are interested in InferenceBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β13Jun 18, 2024Updated 2 years ago
- Python package for serving a local search engine. One command to download and serve a datastore---that's it π.β26Jun 6, 2025Updated last year
- β48Aug 6, 2026Updated 2 weeks ago
- Estimate MFU for DeepSeekV3β26Jan 5, 2025Updated last year
- A Top-Down Profiler for GPU Applicationsβ24Feb 29, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hoursβ528Updated this week
- β23Jun 18, 2026Updated 2 months ago
- A Difficulty-Calibrated Benchmark for Building Terminal Agentsβ30Feb 20, 2026Updated 6 months ago
- Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.β19Jun 1, 2026Updated 2 months ago
- Terminal-Bench 2.1β82Updated this week
- Simple template for quick prototyping and standardization of deep learning projectsβ11Dec 29, 2023Updated 2 years ago
- β14Aug 9, 2023Updated 3 years ago
- β48Oct 1, 2024Updated last year
- Official eval scripts for JobBenchβ38Aug 5, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Benchmark SGLang on SLURMβ24Apr 20, 2026Updated 4 months ago
- An evaluation suite for Retrieval-Augmented Generation (RAG).β25Apr 26, 2025Updated last year
- β39Apr 14, 2026Updated 4 months ago
- Code for safety test in "Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates"β22Sep 21, 2025Updated 11 months ago
- Official implementation of the ΞBelief-RL method.β31Feb 28, 2026Updated 5 months ago
- β21Jun 12, 2026Updated 2 months ago
- A research workbench for developing and testing attacks against large language models, with a focus on prompt injection vulnerabilities aβ¦β60Jul 24, 2026Updated 3 weeks ago
- A collection of judges for evaluating LLM model output for safety & toxicity with a standardized API.β16Jan 7, 2026Updated 7 months ago
- β10Dec 15, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official Repo of CudaForgeβ87Dec 2, 2025Updated 8 months ago
- Measuring and evolving with the frontier of agent workβ533Updated this week
- Code for "Evaluating Spatial Understanding of Large Language Models" TMLR 2024.β16Feb 22, 2024Updated 2 years ago
- Convert GitHub PRs into Harbor tasksβ77Jul 13, 2026Updated last month
- context-efficient terminal agent powered by an RLMβ60Feb 7, 2026Updated 6 months ago
- [CHIL 2024] Interpretation of Intracardiac Electrograms Through Textual Representationsβ12Sep 4, 2024Updated last year
- [NeurIPS 2025 spotlight] Mitigating the AI Safety Impact of Multi-Agent Scaffoldsβ19Sep 22, 2025Updated 11 months ago
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacksβ92Aug 3, 2026Updated 2 weeks ago
- Using PCA, Autoencoder and Fisher linear discriminant to extract the effective representations from the face images. Do the reconstructioβ¦β12Apr 23, 2019Updated 7 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ACL24β11Jun 7, 2024Updated 2 years ago
- NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across speciβ¦β43Updated this week
- A curated list of awesome Harbor ecosystem projectsβ52May 29, 2026Updated 2 months ago
- Comparison of gradient estimation techniques for black-box adversarial examplesβ11Oct 31, 2018Updated 7 years ago
- Tinkering RLβ28Updated this week
- Pioneer Evolutionary Agent for solving long horizon tasksβ24Jan 21, 2026Updated 7 months ago
- Practically and gracefully stop your K8s node on (termination|scale down|maintenance)β13Jul 15, 2020Updated 6 years ago