Agentic Research and Evaluation Suite
☆111Aug 12, 2026Updated 3 weeks ago
Alternatives and similar repositories for ares
Users that are interested in ares are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆23Jun 18, 2026Updated 2 months ago
- ☆142Mar 31, 2026Updated 5 months ago
- A curated list of awesome Harbor ecosystem projects☆52May 29, 2026Updated 3 months ago
- A compact high-signal benchmark for evaluating frontier agents☆27Aug 3, 2026Updated last month
- Well documented examples of running distributed training jobs on Modal☆33Aug 11, 2026Updated 3 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 7 months ago
- Framework for evaluating and improving agents☆4,952Updated this week
- Samaya AI's FrontierFinance Benchmark Grader☆19Jul 16, 2026Updated last month
- benchmarks for LLM tokenizers☆21Mar 25, 2026Updated 5 months ago
- ☆32Nov 14, 2025Updated 9 months ago
- Coco is a proactive co-assistant that connects user workspace with a broader ecosystem of AI agents.☆34Updated this week
- A System for Morphology-Task Generalization via Unified Representation and Behavior Distillation (ICLR2023)☆14Feb 3, 2023Updated 3 years ago
- ☆15Dec 12, 2024Updated last year
- ☆17Jun 19, 2026Updated 2 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆2,551Updated this week
- Ongoing research training transformer language models at scale, including: BERT & GPT-2☆17Mar 11, 2026Updated 5 months ago
- Navi-Bench: benchmarking web agents on everyday tasks directly on real websites☆19Updated this week
- Our library for RL environments + evals☆4,589Updated this week
- Atropos is a Language Model Reinforcement Learning Environments framework for collecting and evaluating LLM trajectories through diverse …☆1,351Jul 4, 2026Updated 2 months ago
- Agentic RL Training at Scale☆2,012Updated this week
- Convert GitHub PRs into Harbor tasks☆82Jul 13, 2026Updated last month
- Small, simple agent task environments for training and evaluation☆20Nov 1, 2024Updated last year
- ☆26Jan 7, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- OpenTinker is an RL-as-a-Service infrastructure for foundation models☆677Mar 21, 2026Updated 5 months ago
- An LLM agent framework for automated AI interpretability research☆19Apr 17, 2026Updated 4 months ago
- Accompanying Code for "Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement Learning", ICML 2023☆25Dec 29, 2023Updated 2 years ago
- Minimal Decision Transformer Implementation written in Jax (Flax).☆18Aug 8, 2022Updated 4 years ago
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆546Updated this week
- Applying SAEs for fine-grained control☆27Dec 15, 2024Updated last year
- ☆38May 4, 2026Updated 4 months ago
- Agen is a minimalist language for agent loops and state machines.☆49Mar 30, 2026Updated 5 months ago
- ☆17Jun 15, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Continual Learning Bench☆216Jul 19, 2026Updated last month
- Solidity contracts for the decentralized Prime Network protocol☆26Jul 6, 2025Updated last year
- AI agent benchmark hackability scanner — find evaluation vulnerabilities before they undermine your results☆45May 25, 2026Updated 3 months ago
- Statistical inference for AI evaluations, from model comparisons to corrections for LLM judge bias, optimized for small sample sizes. All…☆127Aug 28, 2026Updated last week
- ☆81Feb 18, 2026Updated 6 months ago
- Measuring and evolving with the frontier of agent work☆616Updated this week
- A videogame made with PyGame turned into an Open AI Gym Learning Environment for Reinforcement Learning agents.☆14Jan 3, 2023Updated 3 years ago