☆310Jul 1, 2026Updated 3 months ago
Alternatives and similar repositories for hal-harness
Users that are interested in hal-harness are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆80Nov 23, 2025Updated 10 months ago
- [ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery☆176Jul 18, 2026Updated 2 months ago
- Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike stat…☆561Sep 30, 2026Updated last week
- Inspect: A framework for large language model evaluations☆2,971Updated this week
- Inspect Flow is a workflow stack built on Inspect AI that enables research organisations to run AI evaluations at scale.☆22Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- SkyRL: A Modular Full-stack RL Library for LLMs☆2,401Updated this week
- 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource…☆530Sep 4, 2026Updated last month
- An agent benchmark with tasks in a simulated software company.☆792Nov 17, 2025Updated 10 months ago
- Collection of evals for Inspect AI☆697Updated this week
- ☆30Mar 11, 2025Updated last year
- ☆197Aug 17, 2026Updated last month
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆20Jul 4, 2025Updated last year
- Code for T-MARS data filtering☆35Aug 23, 2023Updated 3 years ago
- ☆43Jan 8, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Framework for evaluating and improving agents☆5,979Updated this week
- Code and Data for Tau-Bench☆1,459Mar 18, 2026Updated 6 months ago
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,769Apr 24, 2026Updated 5 months ago
- OpenAI Frontier Evals☆1,311Apr 21, 2026Updated 5 months ago
- An Ultra-Long Output Reinforcement Learning Approach☆23Jul 31, 2025Updated last year
- Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]☆749Jul 29, 2025Updated last year
- ☆28Jun 2, 2026Updated 4 months ago
- Code for the paper "Coding Agents with Multimodal Browsing are Generalist Problem Solvers"☆105Oct 27, 2025Updated 11 months ago
- τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains☆2,200Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"☆1,622Nov 26, 2025Updated 10 months ago
- ☆130Sep 25, 2024Updated 2 years ago
- [NeurIPS'25] Official codebase for "SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution"☆719Mar 16, 2025Updated last year
- BenchBench is a Python package to evaluate multi-task benchmarks.☆24Oct 12, 2025Updated 11 months ago
- ☆17Jan 27, 2026Updated 8 months ago
- A Sober Look at Language Model Reasoning☆92Nov 18, 2025Updated 10 months ago
- ☆57Mar 18, 2026Updated 6 months ago
- ☆138Oct 2, 2026Updated last week
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Agentic RL Training at Scale☆2,148Updated this week
- SWE-bench: Can Language Models Resolve Real-world Github Issues?☆5,995Sep 18, 2026Updated 3 weeks ago
- A benchmark for LLMs on complicated tasks in the terminal☆2,600Jul 11, 2026Updated 3 months ago
- Functional Benchmarks and the Reasoning Gap☆90Jul 23, 2026Updated 2 months ago
- Reproducible, flexible LLM evaluations☆399Mar 24, 2026Updated 6 months ago
- slime is an LLM post-training framework for RL Scaling.☆8,626Updated this week
- Code and data for Koo et al's ACL 2024 paper "Benchmarking Cognitive Biases in Large Language Models as Evaluators"☆23Feb 16, 2024Updated 2 years ago