Meta-Harness: 76.4% on Terminal-Bench 2.0 (Claude Opus 4.6)
☆1,215Mar 26, 2026Updated 5 months ago
Alternatives and similar repositories for meta-harness-tbench2-artifact
Users that are interested in meta-harness-tbench2-artifact are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Reference code for the Meta-Harness paper.☆1,584Sep 11, 2026Updated last week
- KIRA☆930May 29, 2026Updated 3 months ago
- Meta Harness Implementation☆167Sep 2, 2026Updated 2 weeks ago
- Framework for evaluating and improving agents☆5,386Updated this week
- The official repository of "Position: Agentic Evolution is the Path to Evolving LLMs".☆796Aug 22, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A benchmark for LLMs on complicated tasks in the terminal☆2,588Jul 11, 2026Updated 2 months ago
- Self-referential self-improving agents that can optimize for any computable task☆2,746Jul 31, 2026Updated last month
- Official AHE code — Agentic Harness Engineering: observability-driven automatic evolution of coding-agent harnesses (concurrent w/ meta-h…☆893Aug 3, 2026Updated last month
- OpenClaw-RL: Train any agent simply by talking☆5,688May 23, 2026Updated 3 months ago
- autonomous harness engineering☆4,575Apr 3, 2026Updated 5 months ago
- LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training.…☆3,240Aug 20, 2026Updated 3 weeks ago
- AutoHarness: Automated Harness Engineering for AI Agents☆377Apr 2, 2026Updated 5 months ago
- AI agents running research on single-GPU nanochat training automatically☆96,264Mar 26, 2026Updated 5 months ago
- Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and mu…☆1,002Sep 8, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [EMNLP 2026] Official code for "SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization"☆375Aug 12, 2026Updated last month
- Optimize prompts, code, and more with AI-powered Reflective Optimization☆6,631Updated this week
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆5,633Aug 26, 2026Updated 3 weeks ago
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,796Jul 23, 2026Updated last month
- ALMA (Automated meta-Learning of Memory designs for Agentic systems) is a framework that meta-learns memory designs to replace human-engi…☆295Apr 8, 2026Updated 5 months ago
- Open-source implementation of AlphaEvolve☆7,399Jul 18, 2026Updated 2 months ago
- Agentic RL on Any Harness at Scale☆841Aug 13, 2026Updated last month
- Evolve your language agent with Agentic Context Engineering (ACE)☆1,320Aug 24, 2026Updated 3 weeks ago
- Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding ag…☆27,281Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ThetaEvolve: Test-time Learning on Open Problems, enabling RL training on AlphaEvolve/OpenEvolve and emphasizing scaling test-time comput…☆179Feb 27, 2026Updated 6 months ago
- AI-Driven Scientific, Algorithmic, and Systems Discovery☆681Updated this week
- Reinforcement Learning via Self-Distillation (SDPO)☆1,100Jul 1, 2026Updated 2 months ago
- Hierarchal Agent Loop Optimizer☆1,173Sep 8, 2026Updated last week
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆565Sep 10, 2026Updated last week
- Bring your own agent and build a self-improving agentic system. Automatically mine failures, optimize the agent harness, and gate against…☆535Updated this week
- [ICML'26] MemEvolve & EvolveLab☆269May 5, 2026Updated 4 months ago
- ☆637May 24, 2026Updated 3 months ago
- EdgeBench: Unveiling scaling laws of learning from real-world environments☆450Sep 9, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning☆978May 17, 2026Updated 4 months ago
- The first distributed AGI system. Thousands of autonomous AI agents collaboratively train models, share experiments via P2P gossip, and p…☆2,055Updated this week
- ⚒ Evolutionary self-improvement for Hermes Agent — optimize skills, prompts, and code using DSPy + GEPA☆5,367Jun 17, 2026Updated 3 months ago
- Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement…☆10,751Updated this week
- ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution 🧬☆1,397Aug 21, 2026Updated 3 weeks ago
- [ICML 2026] Meta Context Engineering via Agentic Skill Evolution☆169May 4, 2026Updated 4 months ago
- AgentFlow: In-the-Flow Agentic System Optimization☆2,045Feb 8, 2026Updated 7 months ago