Meta-Harness: 76.4% on Terminal-Bench 2.0 (Claude Opus 4.6)
☆1,189Mar 26, 2026Updated 5 months ago
Alternatives and similar repositories for meta-harness-tbench2-artifact
Users that are interested in meta-harness-tbench2-artifact are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Reference code for the Meta-Harness paper.☆1,477Jul 11, 2026Updated last month
- KIRA☆929May 29, 2026Updated 3 months ago
- Meta Harness Implementation☆156Updated this week
- Framework for evaluating and improving agents☆4,759Updated this week
- The official repository of "Position: Agentic Evolution is the Path to Evolving LLMs".☆762Aug 22, 2026Updated last week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A benchmark for LLMs on complicated tasks in the terminal☆2,557Jul 11, 2026Updated last month
- Self-referential self-improving agents that can optimize for any computable task☆2,702Jul 31, 2026Updated 3 weeks ago
- OpenClaw-RL: Train any agent simply by talking☆5,658May 23, 2026Updated 3 months ago
- Official AHE code — Agentic Harness Engineering: observability-driven automatic evolution of coding-agent harnesses (concurrent w/ meta-h…☆860Aug 3, 2026Updated 3 weeks ago
- LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training.…☆3,009Aug 20, 2026Updated last week
- autonomous harness engineering☆4,568Apr 3, 2026Updated 4 months ago
- AutoHarness: Automated Harness Engineering for AI Agents☆370Apr 2, 2026Updated 4 months ago
- AI agents running research on single-GPU nanochat training automatically☆94,901Mar 26, 2026Updated 5 months ago
- Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and mu…☆936Aug 23, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official code for "SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization"☆369Aug 12, 2026Updated 2 weeks ago
- Optimize prompts, code, and more with AI-powered Reflective Optimization☆6,286Updated this week
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆5,555Updated this week
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,732Jul 23, 2026Updated last month
- Open-source implementation of AlphaEvolve☆7,285Jul 18, 2026Updated last month
- Evolve your language agent with Agentic Context Engineering (ACE)☆1,283Updated this week
- Agentic RL on Any Harness at Scale☆819Aug 13, 2026Updated 2 weeks ago
- ALMA (Automated meta-Learning of Memory designs for Agentic systems) is a framework that meta-learns memory designs to replace human-engi…☆291Apr 8, 2026Updated 4 months ago
- Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding ag…☆26,922Aug 19, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ThetaEvolve: Test-time Learning on Open Problems, enabling RL training on AlphaEvolve/OpenEvolve and emphasizing scaling test-time comput…☆176Feb 27, 2026Updated 6 months ago
- AI-Driven Scientific, Algorithmic, and Systems Discovery☆626Updated this week
- Reinforcement Learning via Self-Distillation (SDPO)☆1,077Jul 1, 2026Updated last month
- Hierarchal Agent Loop Optimizer☆1,152Aug 19, 2026Updated last week
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆537Aug 21, 2026Updated last week
- Bring your own agent and build a self-improving agentic system. Automatically mine failures, optimize the agent harness, and gate against…☆535Jul 8, 2026Updated last month
- EdgeBench: Unveiling scaling laws of learning from real-world environments☆432Jul 17, 2026Updated last month
- [ICML'26] MemEvolve & EvolveLab☆260May 5, 2026Updated 3 months ago
- ☆632May 24, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning☆956May 17, 2026Updated 3 months ago
- The first distributed AGI system. Thousands of autonomous AI agents collaboratively train models, share experiments via P2P gossip, and p…☆2,032Updated this week
- Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement…☆10,679Updated this week
- ⚒ Evolutionary self-improvement for Hermes Agent — optimize skills, prompts, and code using DSPy + GEPA☆5,190Jun 17, 2026Updated 2 months ago
- ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution 🧬☆1,359Aug 21, 2026Updated last week
- [ICML 2026] Meta Context Engineering via Agentic Skill Evolution☆162May 4, 2026Updated 3 months ago
- AgentFlow: In-the-Flow Agentic System Optimization☆2,021Feb 8, 2026Updated 6 months ago