A iterative feedback driven benchmark on LLM's instruction following ability
☆58May 25, 2026Updated 3 months ago
Alternatives and similar repositories for Meeseeks
Users that are interested in Meeseeks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated last year
- ☆15Dec 29, 2025Updated 8 months ago
- [ACL 2026] A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models☆23Jul 10, 2026Updated 2 months ago
- [EMNLP 2025] Verification Engineering for RL in Instruction Following☆61Mar 30, 2026Updated 5 months ago
- ☆40Nov 20, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR'26] R-HORIZON: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?☆25May 9, 2026Updated 4 months ago
- CATArena is an engineering-level tournament evaluation platform for Large Language Model-driven code agents (LLM-driven code agents), bas…☆68Dec 25, 2025Updated 8 months ago
- ☆287May 13, 2026Updated 4 months ago
- A Recipe for Building LLM Reasoners to Solve Complex Instructions☆32Oct 9, 2025Updated 11 months ago
- ☆13Apr 15, 2024Updated 2 years ago
- ☆173Sep 9, 2026Updated last week
- ☆1,367Jun 23, 2026Updated 2 months ago
- Official implementation of the paper "Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following"☆40Jan 11, 2026Updated 8 months ago
- [NeurIPS 2024 D&B Track] DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation☆14Mar 5, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Nexusflow function call, tool use, and agent benchmarks.☆28Dec 13, 2024Updated last year
- 基于pytorch的不平衡数据的文本分类☆12Dec 26, 2021Updated 4 years ago
- [NIPS 2025 DB Spotlight] AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios☆41Dec 1, 2025Updated 9 months ago
- ☆26Jul 25, 2023Updated 3 years ago
- ☆44Mar 31, 2026Updated 5 months ago
- ☆24Feb 3, 2024Updated 2 years ago
- [ICLR 2026] Code, benchmark and environment for "ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows"☆137Feb 2, 2026Updated 7 months ago
- ToMBench: Benchmarking Theory of Mind in Large Language Models, ACL 2024.☆69Jun 24, 2024Updated 2 years ago
- ☆34Jan 26, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆30Jun 30, 2025Updated last year
- [ICLR 2026] VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications☆176Feb 22, 2026Updated 6 months ago
- Lean4 Code Editor☆19Aug 18, 2026Updated last month
- 《Python编程 从入门到实践》原书配套源代码☆21Oct 18, 2021Updated 4 years ago
- Code and models for EMNLP 2024 paper "WPO: Enhancing RLHF with Weighted Preference Optimization"☆41Sep 24, 2024Updated last year
- ☆10Sep 10, 2023Updated 3 years ago
- Data and codes for EMNLP 2022 paper "CDConv: A Benchmark for Contradiction Detection in Chinese Conversations"☆13May 8, 2023Updated 3 years ago
- ☆259May 9, 2026Updated 4 months ago
- OpenART is an open-source framework designed to evaluate the safety and robustness of autonomous AI agents in dynamic, long-horizon, and …☆189Sep 8, 2026Updated last week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆10Apr 5, 2025Updated last year
- The official implementation of the paper "Deep Reinforcement Learning with Task-Adaptive Retrieval via Hypernetwork".☆12Feb 27, 2024Updated 2 years ago
- ☆11Mar 3, 2026Updated 6 months ago
- ☆13May 23, 2021Updated 5 years ago
- code for progressive gsl☆12Jan 15, 2026Updated 8 months ago
- ☆15Oct 27, 2025Updated 10 months ago
- ☆15Apr 14, 2025Updated last year