A benchmark for evaluating LLM × harness performance.
☆96Aug 3, 2026Updated this week
Alternatives and similar repositories for PawBench
Users that are interested in PawBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM-powered MCP server for building financial deep-research agents, integrating web search, Crawl4AI scraping, and entity extraction into…☆24Feb 11, 2026Updated 5 months ago
- A platform for running AI agents on realtime sport data and make predictions about game outcomes.☆47Jul 27, 2026Updated last week
- OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards☆763Updated this week
- FlowLLM: Build LLM applications with ease.☆34Jun 27, 2026Updated last month
- A Multi-Agent Approach Integrating Socratic Guidance for Automated Prompt Optimization☆18Dec 15, 2025Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆38Jul 17, 2026Updated 2 weeks ago
- An in-the-wild benchmark for AI agents in the OpenClaw Environment.☆498Updated this week
- EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions☆45Jun 26, 2026Updated last month
- Cutting-edge platform for LLM agent tuning. Deliver RL tuning with flexibility, reliability, speed, multi-agent optimization and realtime…☆231Updated this week
- Harness for running and evaluating AI agents against RL environments☆227Updated this week
- Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding☆18May 6, 2026Updated 2 months ago
- The implementation of <Factual Consistency Evaluation for Text Summarization via Counterfactual Estimation> in PyTorch.☆17Nov 11, 2021Updated 4 years ago
- UIUC CS 225 SP20 ZJUI Environment☆12May 30, 2020Updated 6 years ago
- Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (…☆674Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- An Efficent BPE Algorithm Faster then Hugging Face Tokenizer's Implementation☆13Sep 9, 2024Updated last year
- Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding.☆13Nov 19, 2024Updated last year
- A collection of ready-to-use Python sample agents built with AgentScope and AgentScope Runtime, covering use cases from CLI tools to full…☆335Apr 10, 2026Updated 3 months ago
- 中文金融大模型测评基准,六大类二十五任务、等级化评价,国内模型获得A级☆10May 6, 2024Updated 2 years ago
- 🤖 A list of latest AGI-related repos, resources and courses including LLMs and AI Agents.☆13Sep 24, 2024Updated last year
- Dynamic dual-granularity skill bank for agentic RL, jointly evolving policy and skills to improve long-horizon decision making in agentic…☆71Apr 1, 2026Updated 4 months ago
- Agent-based implementation of RAG, incorporating AI agents into the RAG pipeline to orchestrate its components and perform additional act…☆20Feb 20, 2025Updated last year
- 一个强调工程化、可观测、可测试、可扩展的 RAG 项目。TraceRAG 的目标不是只把答案“生成出来”,而是把文档导入、切块、向量化、检索、带来源回答、评估与后续 tracing 拆成可独立验证的阶段,逐步演进成一个可维护、可解释、可复盘的生产级 RAG。☆15Apr 2, 2026Updated 4 months ago
- [NeurIPS 2025] Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models☆17Nov 2, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆36May 31, 2026Updated 2 months ago
- ☆30Mar 30, 2026Updated 4 months ago
- Code for the paper AdvST: Revisiting Data Augmentations for Single Domain Generalization (AAAI 2024)☆15May 6, 2024Updated 2 years ago
- A curated collection of skills around AgentScope ecosystem and CoPaw applications.☆119Apr 15, 2026Updated 3 months ago
- [VLDB2026] Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards☆15Jul 8, 2026Updated 3 weeks ago
- 【PyTorch】Easy-to-use package of Cognitive Evolutionary Search (CELS) for Click-Through Rate Prediction☆21Oct 29, 2023Updated 2 years ago
- [NAACL 2025] Source code for MMEvalPro, a more trustworthy and efficient benchmark for evaluating LMMs☆25Sep 26, 2024Updated last year
- [KDD'25] Combinatorial Optimization Perspective based Framework for Multi-behavior Recommendation☆20Jan 7, 2025Updated last year
- ☆15May 15, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- SkillsBench evaluates how well skills work and how effective agents are at using them.☆1,626Jul 23, 2026Updated last week
- MUA-RL: MULTI-TURN USER-INTERACTING AGENT REINFORCEMENT LEARNING FOR AGENTIC TOOL USE☆67Nov 5, 2025Updated 8 months ago
- A new ground-based cloud data set.☆10Jun 29, 2024Updated 2 years ago
- ☆35May 16, 2025Updated last year
- 提供公益寻亲平台的网站☆11Feb 17, 2017Updated 9 years ago
- AutoLibra: Metric Induction for Agents from Open-Ended Human Feedback☆19Apr 23, 2026Updated 3 months ago
- The implementation of RAGSynth: Synthetic Data for Robust and Faithful RAG Component Optimization☆21May 26, 2025Updated last year