Agentic Research and Evaluation Suite
☆115Aug 12, 2026Updated last month
Alternatives and similar repositories for ares
Users that are interested in ares are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆23Jun 18, 2026Updated 3 months ago
- ☆148Mar 31, 2026Updated 5 months ago
- A curated list of awesome Harbor ecosystem projects☆54May 29, 2026Updated 3 months ago
- A compact high-signal benchmark for evaluating frontier agents☆39Aug 3, 2026Updated last month
- Well documented examples of running distributed training jobs on Modal☆34Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 8 months ago
- Framework for evaluating and improving agents☆5,637Updated this week
- Samaya AI's FrontierFinance Benchmark Grader☆23Updated this week
- benchmarks for LLM tokenizers☆21Mar 25, 2026Updated 6 months ago
- ☆33Nov 14, 2025Updated 10 months ago
- Coco is a proactive co-assistant that connects user workspace with a broader ecosystem of AI agents.☆36Updated this week
- ☆15Dec 12, 2024Updated last year
- ☆19Jun 19, 2026Updated 3 months ago
- Repository for results and data (coming soon!) for ClawsBench☆35Aug 28, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Data recipes and robust infrastructure for training AI agents☆295Updated this week
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆3,008Updated this week
- Ongoing research training transformer language models at scale, including: BERT & GPT-2☆17Sep 18, 2026Updated last week
- Our library for RL environments + evals☆4,652Updated this week
- Atropos is a Language Model Reinforcement Learning Environments framework for collecting and evaluating LLM trajectories through diverse …☆1,351Jul 4, 2026Updated 2 months ago
- Agentic RL Training at Scale☆2,092Updated this week
- Convert GitHub PRs into Harbor tasks☆88Jul 13, 2026Updated 2 months ago
- Small, simple agent task environments for training and evaluation☆20Nov 1, 2024Updated last year
- Lexical semantic change detection shared task at SemEval 2020: UiO-UVA team☆16Jan 10, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆30Jan 7, 2026Updated 8 months ago
- OpenTinker is an RL-as-a-Service infrastructure for foundation models☆678Mar 21, 2026Updated 6 months ago
- An LLM agent framework for automated AI interpretability research☆19Apr 17, 2026Updated 5 months ago
- Minimal Decision Transformer Implementation written in Jax (Flax).☆18Aug 8, 2022Updated 4 years ago
- context-efficient terminal agent powered by an RLM☆61Feb 7, 2026Updated 7 months ago
- Applying SAEs for fine-grained control☆28Dec 15, 2024Updated last year
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆570Updated this week
- Package for calling different models with same interface☆32Jul 21, 2025Updated last year
- Accompanying Code for "Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement Learning", ICML 2023☆26Dec 29, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Agen is a minimalist language for agent loops and state machines.☆49Mar 30, 2026Updated 5 months ago
- ☆17Jun 15, 2026Updated 3 months ago
- Simple AI CLI that generates docs, unit tests and README.md files☆15Mar 8, 2026Updated 6 months ago
- Continual Learning Bench☆226Jul 19, 2026Updated 2 months ago
- Solidity contracts for the decentralized Prime Network protocol☆26Jul 6, 2025Updated last year
- AI agent benchmark hackability scanner — find evaluation vulnerabilities before they undermine your results☆47May 25, 2026Updated 4 months ago
- Statistical inference for AI evaluations, from model comparisons to corrections for LLM judge bias, optimized for small sample sizes. All…☆127Updated this week