Evaluation harness for Apodex-1.0 on public deep-research benchmarks.
☆387Jun 8, 2026Updated 2 months ago
Alternatives and similar repositories for AgentHarness
Users that are interested in AgentHarness are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis qu…☆47Jul 6, 2026Updated last month
- MiroTrain is an efficient and algorithm-first framework research agent.☆142Aug 27, 2025Updated 11 months ago
- MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning.☆285Aug 12, 2025Updated last year
- DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation.☆142Feb 10, 2026Updated 6 months ago
- MiroRL is an MCP-first reinforcement learning framework for deep research agent.☆248Aug 27, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MARS, a framework optimized for autonomous AI research☆40May 19, 2026Updated 3 months ago
- [ICML 2026] <MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier>☆94May 18, 2026Updated 3 months ago
- 🏆 Top-1 on 5+ benchmarks | Web UI | Supports MiroThinker, Claude, Kimi, OpenAI☆3,099Jul 6, 2026Updated last month
- MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74…☆8,359Jul 6, 2026Updated last month
- Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside emb…☆43Updated this week
- Official Implementation of Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution☆91Dec 8, 2025Updated 8 months ago
- ☆18Mar 3, 2025Updated last year
- This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents wit…☆73Apr 8, 2026Updated 4 months ago
- EdgeBench: Unveiling scaling laws of learning from real-world environments☆431Jul 17, 2026Updated last month
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- SSH auth via browser passkeys (Touch ID / Face ID). No passwords, no SSH keys — just biometrics.☆25Mar 22, 2026Updated 5 months ago
- DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research sys…☆83Aug 14, 2026Updated last week
- ☆94May 8, 2026Updated 3 months ago
- The Source Code for DR3-Eval☆39Aug 12, 2026Updated last week
- REDSearch: A scalable, cost-efficient framework for long-horizon search agents. Features complex task synthesis, optimized mid-training, …☆141Feb 26, 2026Updated 5 months ago
- Implementation of "Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation"☆21Jul 31, 2023Updated 3 years ago
- ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery☆40Jul 20, 2026Updated last month
- The official repo of "WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents"☆121Sep 29, 2025Updated 10 months ago
- Scaling the Horizon, Not the Parameters☆548Jul 16, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆27Jun 22, 2026Updated 2 months ago
- (ACL 2025 Main) Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification - Offici…☆21Dec 26, 2025Updated 7 months ago
- Tmux for claude code☆29Jun 30, 2026Updated last month
- [SIGIR 2025] The official repo for "Scaling Sparse and Dense Retrieval in Decoder-Only LLMs"☆22Mar 31, 2025Updated last year
- Agentic RL on Any Harness at Scale☆802Aug 13, 2026Updated last week
- A Python library that solves context window degradation in long-running LLM agents by moving memory management out of the model layer and…☆18Jul 23, 2026Updated last month
- The most RAM efficient harness☆18,354Updated this week
- Code for "Key Fact as Pivot: A Two-Stage Model for Low Resource Table-to-Text Generation" (ACL 2019)☆17Aug 7, 2019Updated 7 years ago
- Data Synthesis for Deep Research Based on Semi-Structured Data☆216Jul 14, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Lean formalizations of IMO problem statements☆36Apr 23, 2026Updated 4 months ago
- ☆26May 29, 2026Updated 2 months ago
- UniScientist is designed to advance universal scientific research intelligence through a unified paradigm☆170Mar 14, 2026Updated 5 months ago
- SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, va…☆16,276Updated this week
- ☆26Jul 2, 2026Updated last month
- A self-hosted, zero-knowledge, secure PHP notebook. Your notes, in your browser, on your server - every byte encrypted with AES-256 using…☆26Jun 20, 2026Updated 2 months ago
- A private, open-source alternative to Windows Recall — for Linux. Captures your screen, OCRs it, and makes everything you've seen instant…☆50Jun 24, 2026Updated last month