MathEval is a benchmark dedicated to the holistic evaluation on mathematical capacities of LLMs.
☆87Nov 15, 2024Updated last year
Alternatives and similar repositories for MathEval
Users that are interested in MathEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆12Dec 8, 2021Updated 4 years ago
- [ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset☆115May 22, 2025Updated last year
- AAAI2024 Global Competition on Math Problem Solving and Reasoning☆14Oct 4, 2023Updated 2 years ago
- The official repository of the Omni-MATH benchmark.☆94Dec 22, 2024Updated last year
- A simple toolkit for benchmarking LLMs on mathematical reasoning tasks. 🧮✨☆277Apr 26, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Mix of Minimal Optimal Sets (MMOS) of dataset has two advantages for two aspects, higher performance and lower construction costs on math…☆73Jul 27, 2024Updated last year
- [NeurIPS 2024] OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI☆106Mar 6, 2025Updated last year
- ☆30Dec 27, 2024Updated last year
- Code for ICML21 paper "Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation"☆12Feb 8, 2023Updated 3 years ago
- Code for the ACL2022 paper "Synthetic Question Value Estimation for Domain Adaptation of Question Answering"☆18Mar 21, 2022Updated 4 years ago
- [ICLR 2025] This is the code repo for our ICLR’25 paper "RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rew…☆55Feb 10, 2025Updated last year
- Rudi, A., Camoriano, R. and Rosasco, L., Less is more: Nyström computational regularization. In Advances in Neural Information Processing…☆13Sep 17, 2019Updated 6 years ago
- PULSE-EVAL☆24Jan 12, 2024Updated 2 years ago
- [ACL '26] source code for the paper: "Long-Chain Reasoning Distillation via Adaptive Prefix Alignment"☆16Jan 21, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Pytorch implementation for "Particle Flow Bayes' Rule"☆16Jun 22, 2019Updated 7 years ago
- [AAAI 2025] Augmenting Math Word Problems via Iterative Question Composing (https://arxiv.org/abs/2401.09003)☆23Oct 2, 2025Updated 9 months ago
- ☆10Dec 28, 2023Updated 2 years ago
- An Experiment on Dynamic NTK Scaling RoPE☆65Nov 26, 2023Updated 2 years ago
- Graph4Tree is a simple example code for our EMNLP'20 Findings paper idea.☆26Nov 18, 2020Updated 5 years ago
- The dataset and code for paper: TheoremQA: A Theorem-driven Question Answering dataset☆161Apr 23, 2024Updated 2 years ago
- GSM-Plus: Data, Code, and Evaluation for Enhancing Robust Mathematical Reasoning in Math Word Problems.☆66Jul 8, 2024Updated 2 years ago
- [ACL 2024 Demo] Official GitHub repo for UltraEval: An open source framework for evaluating foundation models.☆256Oct 30, 2024Updated last year
- We systematically studied the influencing factors when LLM generates benchmarks,By using our code, you can generate high-quality QA datas…☆20May 20, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories☆41Sep 4, 2024Updated last year
- ☆27Nov 1, 2021Updated 4 years ago
- [ICML'24] TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks☆33Sep 20, 2024Updated last year
- the instructions and demonstrations for building a formal logical reasoning capable GLM☆54Sep 3, 2024Updated last year
- A custom line wrap layout ,support set max lines.(自定 义流式布局,支持设置最大行数)☆10Apr 13, 2018Updated 8 years ago
- [NAACL 2025] Representing Rule-based Chatbots with Transformers☆23Feb 9, 2025Updated last year
- Technical Report: Is ChatGPT a Good NLG Evaluator? A Preliminary Study☆43Mar 8, 2023Updated 3 years ago
- Builds a WMT18-like corpus for word-level QE with annotations in the source and target words.☆10Sep 19, 2022Updated 3 years ago
- [MathCoder, MathCoder-VL] Family of LLMs/LMMs for mathematical reasoning.☆340Oct 18, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Dec 24, 2019Updated 6 years ago
- Spherical random features for polynomial kernels☆10Dec 1, 2015Updated 10 years ago
- The official code of "Towards Long-horizon Agentic Multimodal Search"☆27Apr 17, 2026Updated 3 months ago
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- NAACL 2021: Are NLP Models really able to Solve Simple Math Word Problems?☆142Jun 30, 2022Updated 4 years ago
- Code and data to support Bamman et al. (2020), "A Dataset of Literary Coreference" (LREC)☆11Dec 8, 2022Updated 3 years ago
- The code corresponding to paper RKT : Relation-Aware Self-Attention for Knowledge Tracing☆45Mar 12, 2021Updated 5 years ago