Evaluation utilities based on SymPy.
☆25Dec 12, 2024Updated last year
Alternatives and similar repositories for symeval
Users that are interested in symeval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The rule-based evaluation subset and code implementation of Omni-MATH☆29Dec 23, 2024Updated last year
- ☆12Nov 5, 2024Updated last year
- [AAAI 2025] Augmenting Math Word Problems via Iterative Question Composing (https://arxiv.org/abs/2401.09003)☆23Oct 2, 2025Updated 11 months ago
- A curated list of papers, tools, and benchmarks on LLM-based computer-use agents, covering both terminal/CLI and GUI approaches.☆17Aug 17, 2026Updated 3 weeks ago
- [NeurIPS'24] Official code for *🎯DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving*☆121Dec 10, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆52Mar 5, 2025Updated last year
- [AAAI 2025 oral] Evaluating Mathematical Reasoning Beyond Accuracy☆80Oct 9, 2025Updated 10 months ago
- ☆16Nov 11, 2025Updated 9 months ago
- [ACL25' Findings] SWE-Dev is an SWE agent with a scalable test case construction pipeline.☆66Jul 21, 2025Updated last year
- ☆20May 15, 2026Updated 3 months ago
- [NeurIPS ENLSP Workshop'24] CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios☆16Oct 18, 2024Updated last year
- Potluck with different functions for different purposes that can be shared among C programs☆14May 9, 2026Updated 3 months ago
- ☆34Sep 14, 2024Updated last year
- Collections of RLxLM experiments using minimal codes☆14Feb 17, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Dive-into-LLMs Tutorial for Beginners☆29May 14, 2024Updated 2 years ago
- Official Implementation of Paper "Learning to Jump: Thinning and Thickening Latent Counts for Generative Modeling" (ICML 2023)☆10Jun 6, 2023Updated 3 years ago
- 清华大学第六届人工智能挑战赛电子系赛道(原电子系第 24 届队式程序设计大赛 teamstyle24)☆29May 11, 2024Updated 2 years ago
- ☆34May 31, 2025Updated last year
- JsonTuning: Towards Generalizable, Robust, and Controllable Instruction Tuning☆10Nov 3, 2024Updated last year
- The official repository for the paper Multilingual Mathematical Autoformalization☆39May 20, 2024Updated 2 years ago
- This repo contains my customised style python based plots for NLP papers, and includes my reproduction for my favourite papers' plots☆39Mar 4, 2024Updated 2 years ago
- ☆13Jul 14, 2024Updated 2 years ago
- ☆14Oct 21, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Modern development with Python in 2024☆12Aug 17, 2026Updated 3 weeks ago
- ☆16Apr 23, 2023Updated 3 years ago
- Improving Language Understanding from Screenshots. Paper: https://arxiv.org/abs/2402.14073☆32Jul 9, 2024Updated 2 years ago
- A Model Context Protocol (MCP) server for Google Calendar integration in Cluade Desktop with auto authentication support. This server ena…☆12Mar 11, 2025Updated last year
- Experimental online password and secret locker☆21Feb 20, 2024Updated 2 years ago
- Replicating O1 inference-time scaling laws☆94Dec 1, 2024Updated last year
- SemiDefinite Programming Algorithm (SDPA) for Python☆12Jul 1, 2026Updated 2 months ago
- [AAAI 2025] Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems☆13May 5, 2025Updated last year
- ☆11Oct 11, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICLR 26] The official code repository for the paper "Mirage or Method? How Model–Task Alignment Induces Divergent RL Conclusions".☆19Feb 9, 2026Updated 6 months ago
- [Technical Report] Official PyTorch implementation code for realizing the technical part of Phantom of Latent representing equipped with …☆63Oct 9, 2024Updated last year
- Work in progress save editor for Monster Hunter: World☆11Aug 15, 2018Updated 8 years ago
- Top Picks for Data Science Self-Study: From Newbies to Pros!☆11Apr 2, 2024Updated 2 years ago
- [NeurIPS 2025 D&B Track] Evaluation Code Repo for Paper "PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts"☆43May 22, 2025Updated last year
- Revisiting Mid-training in the Era of Reinforcement Learning Scaling☆188Jul 23, 2025Updated last year
- [ICML 2025] Teaching Language Models to Critique via Reinforcement Learning☆126May 6, 2025Updated last year