A benchmark for evaluating LLM × harness performance.
☆109Aug 3, 2026Updated last month
Alternatives and similar repositories for PawBench
Users that are interested in PawBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM-powered MCP server for building financial deep-research agents, integrating web search, Crawl4AI scraping, and entity extraction into…☆25Feb 11, 2026Updated 7 months ago
- A platform for running AI agents on realtime sport data and make predictions about game outcomes.☆50Jul 27, 2026Updated last month
- OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards☆846Sep 11, 2026Updated last week
- FlowLLM: Build LLM applications with ease.☆35Jun 27, 2026Updated 2 months ago
- A Multi-Agent Approach Integrating Socratic Guidance for Automated Prompt Optimization☆18Dec 15, 2025Updated 9 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- An in-the-wild benchmark for AI agents in the production harness.☆523Updated this week
- Cutting-edge platform for LLM agent tuning. Deliver RL tuning with flexibility, reliability, speed, multi-agent optimization and realtime…☆237Aug 25, 2026Updated 3 weeks ago
- Real-world browser-agent benchmark: 210 tasks across 107 websites, multi-agent/multi-browser evaluation, reproducible leaderboard and res…☆19Jul 31, 2026Updated last month
- Harness for running and evaluating AI agents against RL environments☆276Updated this week
- The implementation of <Factual Consistency Evaluation for Text Summarization via Counterfactual Estimation> in PyTorch.☆17Nov 11, 2021Updated 4 years ago
- UIUC CS 225 SP20 ZJUI Environment☆12May 30, 2020Updated 6 years ago
- Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (…☆702Updated this week
- An Efficent BPE Algorithm Faster then Hugging Face Tokenizer's Implementation☆13Sep 9, 2024Updated 2 years ago
- Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding.☆13Nov 19, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Jan 31, 2024Updated 2 years ago
- A collection of ready-to-use Python sample agents built with AgentScope and AgentScope Runtime, covering use cases from CLI tools to full…☆345Apr 10, 2026Updated 5 months ago
- 🤖 A list of latest AGI-related repos, resources and courses including LLMs and AI Agents.☆13Sep 24, 2024Updated last year
- Agent-based implementation of RAG, incorporating AI agents into the RAG pipeline to orchestrate its components and perform additional act…☆20Feb 20, 2025Updated last year
- Dynamic dual-granularity skill bank for agentic RL, jointly evolving policy and skills to improve long-horizon decision making in agentic…☆70Apr 1, 2026Updated 5 months ago
- 一个强调工程化、可观测、可测试、可扩展的 RAG 项目。TraceRAG 的目标不是只把答案“生成出来”,而是把文档导入、切块、向量化、检索、带来源回答、评估与后续 tracing 拆成可独立验证的阶段,逐步演进成一个可维护、可解释、可复盘的生产级 RAG。☆15Apr 2, 2026Updated 5 months ago
- [NeurIPS 2025] Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models☆17Nov 2, 2025Updated 10 months ago
- support kubernetes feature for autogen(https://github.com/microsoft/autogen)☆11Sep 15, 2025Updated last year
- Academic papers and works related to SWE-bench and SWE-agents☆17Dec 8, 2025Updated 9 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Here is the repo for public scripts.☆12Jul 16, 2022Updated 4 years ago
- 【PyTorch】Easy-to-use package of Cognitive Evolutionary Search (CELS) for Click-Through Rate Prediction☆21Oct 29, 2023Updated 2 years ago
- ☆15Feb 26, 2026Updated 6 months ago
- General Agent Benchmark for OpenClaw, made by Qwen Team, Alibaba Group.☆61Jun 10, 2026Updated 3 months ago
- [NAACL 2025] Source code for MMEvalPro, a more trustworthy and efficient benchmark for evaluating LMMs☆25Sep 26, 2024Updated last year
- [KDD'25] Combinatorial Optimization Perspective based Framework for Multi-behavior Recommendation☆20Jan 7, 2025Updated last year
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,804Jul 23, 2026Updated last month
- Accelerated in CUDA☆11Oct 28, 2022Updated 3 years ago
- MUA-RL: MULTI-TURN USER-INTERACTING AGENT REINFORCEMENT LEARNING FOR AGENTIC TOOL USE☆68Nov 5, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A new ground-based cloud data set.☆10Jun 29, 2024Updated 2 years ago
- [ECCV 2024] Versatile Incremental Learning: Towards Class and Domain-Agnostic Incremental Learning☆18Sep 23, 2024Updated last year