Evaluation harness for Apodex-1.0 on public deep-research benchmarks.
☆377Jun 8, 2026Updated last month
Alternatives and similar repositories for AgentHarness
Users that are interested in AgentHarness are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis qu…☆46Jul 6, 2026Updated 3 weeks ago
- MiroTrain is an efficient and algorithm-first framework research agent.☆142Aug 27, 2025Updated 11 months ago
- MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning.☆282Aug 12, 2025Updated 11 months ago
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification☆21Oct 8, 2025Updated 9 months ago
- MiroRL is an MCP-first reinforcement learning framework for deep research agent.☆246Aug 27, 2025Updated 11 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- MARS, a framework optimized for autonomous AI research☆39May 19, 2026Updated 2 months ago
- 🏆 Top-1 on 5+ benchmarks | Web UI | Supports MiroThinker, Claude, Kimi, OpenAI☆3,086Jul 6, 2026Updated 3 weeks ago
- MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74…☆8,362Jul 6, 2026Updated 3 weeks ago
- Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside emb…☆37Updated this week
- Official Implementation of Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution☆89Dec 8, 2025Updated 7 months ago
- ☆18Mar 3, 2025Updated last year
- This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents wit…☆72Apr 8, 2026Updated 3 months ago
- EdgeBench: Unveiling scaling laws of learning from real-world environments☆407Jul 17, 2026Updated 2 weeks ago
- SSH auth via browser passkeys (Touch ID / Face ID). No passwords, no SSH keys — just biometrics.☆25Mar 22, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research sys…☆75May 14, 2026Updated 2 months ago
- ☆95May 8, 2026Updated 2 months ago
- ☆39May 7, 2026Updated 2 months ago
- REDSearch: A scalable, cost-efficient framework for long-horizon search agents. Features complex task synthesis, optimized mid-training, …☆134Feb 26, 2026Updated 5 months ago
- Implementation of "Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation"☆21Jul 31, 2023Updated 3 years ago
- ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery☆40Jul 20, 2026Updated 2 weeks ago
- The official repo of "WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents"☆120Sep 29, 2025Updated 10 months ago
- ☆27Jun 22, 2026Updated last month
- (ACL 2025 Main) Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification - Offici…☆21Dec 26, 2025Updated 7 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Tmux for claude code☆29Jun 30, 2026Updated last month
- [SIGIR 2025] The official repo for "Scaling Sparse and Dense Retrieval in Decoder-Only LLMs"☆22Mar 31, 2025Updated last year
- Agentic RL on Any Harness at Scale☆737Jul 15, 2026Updated 2 weeks ago
- A Python library that solves context window degradation in long-running LLM agents by moving memory management out of the model layer and…☆18Jul 23, 2026Updated last week
- The most RAM efficient harness☆15,581Updated this week
- Data Synthesis for Deep Research Based on Semi-Structured Data☆216Jul 14, 2026Updated 3 weeks ago
- ☆24May 29, 2026Updated 2 months ago
- UniScientist is designed to advance universal scientific research intelligence through a unified paradigm☆169Mar 14, 2026Updated 4 months ago
- SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, va…☆15,539Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆26Jul 2, 2026Updated last month
- A private, open-source alternative to Windows Recall — for Linux. Captures your screen, OCRs it, and makes everything you've seen instant…☆45Jun 24, 2026Updated last month
- TokenSpeed is a speed-of-light LLM inference engine.☆1,795Updated this week
- The implementation for SIGIR 2026: Learning to Retrieve from Agent Trajectories.☆56Jul 14, 2026Updated 2 weeks ago
- A vector index built on TurboQuant, written in Rust with Python bindings☆14,599Updated this week
- Qwen-AgentWorld: Language World Models for General Agents☆919Jul 20, 2026Updated 2 weeks ago
- Tongyi Deep Research, the Leading Open-source Deep Research Agent☆19,779Feb 27, 2026Updated 5 months ago