Evaluation harness for Apodex-1.0 on public deep-research benchmarks.
☆435Aug 25, 2026Updated 2 weeks ago
Alternatives and similar repositories for AgentHarness
Users that are interested in AgentHarness are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis qu…☆49Jul 6, 2026Updated 2 months ago
- MiroTrain is an efficient and algorithm-first framework research agent.☆142Aug 27, 2025Updated last year
- MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning.☆287Aug 12, 2025Updated last year
- DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation.☆143Feb 10, 2026Updated 7 months ago
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification☆22Oct 8, 2025Updated 11 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- MiroRL is an MCP-first reinforcement learning framework for deep research agent.☆249Aug 27, 2025Updated last year
- MARS, a framework optimized for autonomous AI research☆41May 19, 2026Updated 3 months ago
- [ICML 2026] <MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier>☆96May 18, 2026Updated 3 months ago
- 🏆 Top-1 on 5+ benchmarks | Web UI | Supports MiroThinker, Claude, Kimi, OpenAI☆3,106Jul 6, 2026Updated 2 months ago
- MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74…☆8,379Jul 6, 2026Updated 2 months ago
- Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside emb…☆48Updated this week
- Official Implementation of Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution☆91Dec 8, 2025Updated 9 months ago
- ☆18Mar 3, 2025Updated last year
- This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents wit…☆73Apr 8, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- SSH auth via browser passkeys (Touch ID / Face ID). No passwords, no SSH keys — just biometrics.☆28Mar 22, 2026Updated 5 months ago
- EdgeBench: Unveiling scaling laws of learning from real-world environments☆443Updated this week
- DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research sys…☆86Updated this week
- ☆95May 8, 2026Updated 4 months ago
- The Source Code for DR3-Eval☆40Aug 12, 2026Updated last month
- REDSearch: A scalable, cost-efficient framework for long-horizon search agents. Features complex task synthesis, optimized mid-training, …☆142Feb 26, 2026Updated 6 months ago
- 机构级 A股量化系统 - Hermes 多智能体 + Barra 中性化 + Level2 微结构 + 15 个圈内 tricks 完整实现☆25Aug 4, 2026Updated last month
- Implementation of "Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation"☆21Jul 31, 2023Updated 3 years ago
- ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery☆40Jul 20, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The official repo of "WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents"☆122Sep 29, 2025Updated 11 months ago
- ☆27Jun 22, 2026Updated 2 months ago
- (ACL 2025 Main) Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification - Offici…☆21Dec 26, 2025Updated 8 months ago
- Tmux for claude code☆30Jun 30, 2026Updated 2 months ago
- Agentic RL on Any Harness at Scale☆837Aug 13, 2026Updated last month
- A Python library that solves context window degradation in long-running LLM agents by moving memory management out of the model layer and…☆20Jul 23, 2026Updated last month
- Code for "Key Fact as Pivot: A Two-Stage Model for Low Resource Table-to-Text Generation" (ACL 2019)☆17Aug 7, 2019Updated 7 years ago
- Data Synthesis for Deep Research Based on Semi-Structured Data☆216Jul 14, 2026Updated last month
- The most RAM efficient harness☆19,544Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- UniScientist is designed to advance universal scientific research intelligence through a unified paradigm☆172Mar 14, 2026Updated 5 months ago
- SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, va…☆16,940Sep 5, 2026Updated last week
- Turn PDFs, books and papers into interactive learning webpages|将复杂材料转化为可追溯、可测验、可做笔记的学习网页☆115Updated this week
- ☆27Jul 2, 2026Updated 2 months ago
- A private, open-source alternative to Windows Recall — for Linux. Captures your screen, OCRs it, and makes everything you've seen instant…☆51Jun 24, 2026Updated 2 months ago
- A tiny, fully-reproducible JEPA world model that learns the physics of a bouncing DVD logo in representation space, dreams its future, an…☆34Jun 13, 2026Updated 3 months ago
- TokenSpeed is a speed-of-light LLM inference engine.☆2,116Updated this week