A new dataset of difficult graduate-level applied mathematics problems; evaluations demonstrate that leading LLMs currently exhibit low accuracy in solving these problems.
☆30Feb 14, 2025Updated last year
Alternatives and similar repositories for HARDMath
Users that are interested in HARDMath are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The rule-based evaluation subset and code implementation of Omni-MATH☆29Dec 23, 2024Updated last year
- The official repository of the Omni-MATH benchmark.☆93Dec 22, 2024Updated last year
- ☆30Dec 27, 2024Updated last year
- Official Code Repository for [AutoScale📈: Scale-Aware Data Mixing for Pre-Training LLMs] Published as a conference paper at **COLM 2025*…☆14Aug 8, 2025Updated last year
- ☆84Jan 25, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models☆74Feb 25, 2025Updated last year
- ☆12Jun 5, 2024Updated 2 years ago
- ☆11Jul 15, 2020Updated 6 years ago
- ☆14Oct 21, 2024Updated last year
- ☆11Jan 2, 2022Updated 4 years ago
- A curated biochemical database that integrates and refines data from KEGG and ATLAS databases to support precise analyses of biochemical …☆16May 13, 2026Updated 2 months ago
- A framework for evolving and testing question-answering datasets with various models.☆26Feb 28, 2024Updated 2 years ago
- [ACL 2024]Official GitHub repo for OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scie…☆195Jun 8, 2025Updated last year
- DGCIT: Double Generative Adversarial Networks for Conditional Independence Testing☆11Nov 22, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- An official implementation of "Catastrophic Failure of LLM Unlearning via Quantization" (ICLR 2025)☆39Feb 22, 2025Updated last year
- The implementation of paper "LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Fee…☆38Jul 25, 2024Updated 2 years ago
- As defined in Lubotzky, Philips and Sarnak☆10Oct 25, 2022Updated 3 years ago
- !!!!(DEMO)!!!! !!! CHECK OUT THE NEW VERSİON !!! Counting Close People with Yolov7☆13Sep 14, 2022Updated 3 years ago
- Original code base for On Pretraining Data Diversity for Self-Supervised Learning☆14Dec 30, 2024Updated last year
- Code and data to support Bamman et al. (2020), "A Dataset of Literary Coreference" (LREC)☆11Dec 8, 2022Updated 3 years ago
- ☆13May 13, 2021Updated 5 years ago
- ☆16Apr 16, 2025Updated last year
- Codes for coreference-aware machine reading comprehension☆13Mar 13, 2022Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).☆21Jun 9, 2026Updated 2 months ago
- Code for paper "Principled feature attribution for unsupervised gene expression analysis"☆13Mar 7, 2023Updated 3 years ago
- ☆80Nov 19, 2024Updated last year
- LaTeX Beamer template crafted for University of Illinois Chicago☆12Dec 7, 2024Updated last year
- Logical Operations On Puzzles: Simple Iterative Reasoning Tests for LLMs first through wordgrids☆18Feb 19, 2025Updated last year
- [ICML 2025] Closed-Loop Long-Horizon Robotic Planning via Equilibrium Sequence Modeling☆13May 5, 2025Updated last year
- [ICML2025] Official Repo for Paper "Optimizing Temperature for Language Models with Multi-Sample Inference"☆23Feb 16, 2025Updated last year
- Repository for the paper: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning☆18Feb 21, 2025Updated last year
- Code for SLT 2016 paper on Grapheme-to-Phoneme conversion using attention based encoder-decoder models☆15Feb 20, 2019Updated 7 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [EMNLP 2024] A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners☆28Dec 11, 2024Updated last year
- Mix of Minimal Optimal Sets (MMOS) of dataset has two advantages for two aspects, higher performance and lower construction costs on math…☆73Jul 27, 2024Updated 2 years ago
- Official Code for Learning to Sample Effective and Diverse Prompts for Text-to-Image Generation (CVPR 2025)☆15Apr 2, 2025Updated last year
- ☆17Sep 2, 2023Updated 2 years ago
- [ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset☆116May 22, 2025Updated last year
- Elastic Workplace Search Official Python Client☆10Aug 8, 2024Updated 2 years ago
- Implementation of the ICML 2024 paper "Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning" pr…☆117Feb 9, 2024Updated 2 years ago