A light-weight tool for evaluating LLMs in rule-based ways.
☆87Jun 19, 2025Updated last year
Alternatives and similar repositories for MathRuler
Users that are interested in MathRuler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation☆112Sep 18, 2025Updated 10 months ago
- ☆1,172Jan 10, 2026Updated 6 months ago
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL☆5,083Updated this week
- ☆22Sep 19, 2023Updated 2 years ago
- xVerify: Efficient Answer Verifier for Reasoning Model Evaluations☆149Nov 13, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning.☆24Oct 7, 2025Updated 9 months ago
- ☆12Mar 22, 2025Updated last year
- mcp wrapper for openai built-in tools☆12Mar 13, 2025Updated last year
- ☆24Jun 18, 2025Updated last year
- MetricEval: A framework that conceptualizes and operationalizes four main components of metric evaluation, in terms of reliability and va…☆12Nov 6, 2023Updated 2 years ago
- (ACL 2025) MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale☆50Jun 4, 2025Updated last year
- Code for paper: Reinforced Vision Perception with Tools☆74Oct 3, 2025Updated 9 months ago
- ☆187Dec 5, 2025Updated 7 months ago
- OpenVLThinker: An Early Exploration to Vision-Language Reasoning via Iterative Self-Improvement☆155May 25, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward☆94Aug 8, 2025Updated 11 months ago
- ☆42Jul 15, 2025Updated last year
- A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning☆297Sep 25, 2025Updated 10 months ago
- RACE is a multi-dimensional benchmark for code generation that focuses on Readability, mAintainability, Correctness, and Efficiency.☆14Oct 12, 2024Updated last year
- ☆47Apr 9, 2025Updated last year
- An interactive daily life logging tool☆10Oct 15, 2024Updated last year
- Official Repo for Open-Reasoner-Zero☆2,096Jun 2, 2025Updated last year
- ☆10Feb 16, 2025Updated last year
- Extend OpenRLHF to support LMM RL training for reproduction of DeepSeek-R1 on multimodal tasks.☆848May 14, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [AAAI 2025 oral] Evaluating Mathematical Reasoning Beyond Accuracy☆80Oct 9, 2025Updated 9 months ago
- A fork to add multimodal model training to open-r1☆1,594Feb 8, 2025Updated last year
- Code execution sandbox(support Open-R1), Supports multiple languages(Python/Java/C/Kotlin/Swift/OC/GO/...)☆23Mar 6, 2025Updated last year
- Understanding R1-Zero-Like Training: A Critical Perspective☆1,269Aug 27, 2025Updated 11 months ago
- [NeurIPS 2025] Reasoning Models Better Express Their Confidence"☆23Nov 19, 2025Updated 8 months ago
- [ICLR 2026]🚀ReVisual-R1 is a 7B open-source multimodal language model that follows a three-stage curriculum—cold-start pre-training, mul…☆212Dec 10, 2025Updated 7 months ago
- Repository for "Scaling Evaluation-time Compute with Reasoning Models as Process Evaluators"☆12Mar 25, 2025Updated last year
- Unleashing the Power of Reinforcement Learning for Math and Code Reasoners☆740Jun 6, 2025Updated last year
- Text Adventure Learning Environment Suite - Benchmark to evaluate language models on interactive text environments.☆30Jul 17, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆39Aug 1, 2025Updated 11 months ago
- ☆21Jul 9, 2025Updated last year
- ☆13Jan 12, 2023Updated 3 years ago
- Toy implementation of Strawberry☆33Sep 24, 2024Updated last year
- Bridge Megatron-Core to Hugging Face/Reinforcement Learning☆228Jun 15, 2026Updated last month
- Async pipelined version of Verl☆124Apr 8, 2025Updated last year
- Official implementation of "Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation" (CVPR 202…☆40May 26, 2025Updated last year