Agentic Research and Evaluation Suite
☆107Jul 24, 2026Updated this week
Alternatives and similar repositories for ares
Users that are interested in ares are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Jun 18, 2026Updated last month
- ☆134Mar 31, 2026Updated 3 months ago
- A curated list of awesome Harbor ecosystem projects☆48May 29, 2026Updated last month
- Well documented examples of running distributed training jobs on Modal☆29Jul 19, 2026Updated last week
- Trajectory Recording and Capture Environments☆19Jan 24, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Framework for evaluating and improving agents☆3,504Updated this week
- Samaya AI's FrontierFinance Benchmark Grader☆17Jul 16, 2026Updated last week
- benchmarks for LLM tokenizers☆20Mar 25, 2026Updated 4 months ago
- ☆31Nov 14, 2025Updated 8 months ago
- Coco is a proactive co-assistant that connects user workspace with a broader ecosystem of AI agents.☆23Updated this week
- A System for Morphology-Task Generalization via Unified Representation and Behavior Distillation (ICLR2023)☆14Feb 3, 2023Updated 3 years ago
- ☆14Dec 12, 2024Updated last year
- ☆15Jun 19, 2026Updated last month
- Data recipes and robust infrastructure for training AI agents☆265Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.☆1,789Updated this week
- Ongoing research training transformer language models at scale, including: BERT & GPT-2☆17Mar 11, 2026Updated 4 months ago
- Navi-Bench: benchmarking web agents on everyday tasks directly on real websites☆19Updated this week
- Our library for RL environments + evals☆4,400Updated this week
- Atropos is a Language Model Reinforcement Learning Environments framework for collecting and evaluating LLM trajectories through diverse …☆1,340Jul 4, 2026Updated 3 weeks ago
- Agentic RL Training at Scale☆1,724Updated this week
- Convert GitHub PRs into Harbor tasks☆72Jul 13, 2026Updated last week
- Small, simple agent task environments for training and evaluation☆20Nov 1, 2024Updated last year
- Lexical semantic change detection shared task at SemEval 2020: UiO-UVA team☆16Jan 10, 2023Updated 3 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆20Jan 7, 2026Updated 6 months ago
- OpenTinker is an RL-as-a-Service infrastructure for foundation models☆676Mar 21, 2026Updated 4 months ago
- Minimal Decision Transformer Implementation written in Jax (Flax).☆18Aug 8, 2022Updated 3 years ago
- context-efficient terminal agent powered by an RLM☆60Feb 7, 2026Updated 5 months ago
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆467Updated this week
- Applying SAEs for fine-grained control☆27Dec 15, 2024Updated last year
- ☆38May 4, 2026Updated 2 months ago
- Package for calling different models with same interface☆34Jul 21, 2025Updated last year
- Agen is a minimalist language for agent loops and state machines.☆50Mar 30, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Simple AI CLI that generates docs, unit tests and README.md files☆15Mar 8, 2026Updated 4 months ago
- ☆16Jun 15, 2026Updated last month
- Continual Learning Bench☆189Jul 19, 2026Updated last week
- Solidity contracts for the decentralized Prime Network protocol☆26Jul 6, 2025Updated last year
- AI agent benchmark hackability scanner — find evaluation vulnerabilities before they undermine your results☆40May 25, 2026Updated 2 months ago
- Statistical analysis methods for comparing prompt and model performance in LLM evaluations.☆109Updated this week
- ☆81Feb 18, 2026Updated 5 months ago