[NeurIPS 2025 D&B Track] Evaluation Code Repo for Paper "PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts"
☆43May 22, 2025Updated last year
Alternatives and similar repositories for PolyMath
Users that are interested in PolyMath are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14May 21, 2024Updated 2 years ago
- [NeurIPS 2024] Code and Data Repo for Paper "Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning"☆28May 28, 2024Updated 2 years ago
- CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings☆74Feb 3, 2025Updated last year
- ☆17Aug 2, 2023Updated 3 years ago
- Dive-into-LLMs Tutorial for Beginners☆29May 14, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The rule-based evaluation subset and code implementation of Omni-MATH☆29Dec 23, 2024Updated last year
- A documentation translation tool specifically designed for Qwen Code☆48Updated this week
- Open-source examples and guides for building with the Qwen. Browse a collection of snippets, advanced techniques and walkthroughs.☆39Nov 20, 2024Updated last year
- {DeepL, Google, WMT-Best, davinci-003, turbo, gpt-4} × {En-De, En-Cs, En-Ru, En-Zh, De-Fr, En-Ja, Uk-En, Uk-Cs, En-Hr, En-Ha, En-Is}☆14Jun 18, 2023Updated 3 years ago
- A retrieval augmented sequence modeling toolkit implemented based on Fairseq☆29Mar 3, 2023Updated 3 years ago
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space☆17Apr 16, 2026Updated 4 months ago
- A GitHub Action that integrates Qwen Code into your development workflow.☆38Aug 3, 2026Updated last month
- The geometry of multilingual language model representations (EMNLP 2022).☆22Oct 21, 2022Updated 3 years ago
- LEIA: Facilitating Cross-Lingual Knowledge Transfer in Language Models with Entity-based Data Augmentation☆23Apr 24, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ACL' 25] The official code repository for PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.☆94Feb 15, 2025Updated last year
- ACL Paper Lists(machine translation)☆13Mar 23, 2022Updated 4 years ago
- [ICLR 2025 Oral] Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition☆17Nov 25, 2024Updated last year
- Code for Blog Post: Can Better Cold-Start Strategies Improve RL Training for LLMs?☆20Mar 9, 2025Updated last year
- Evaluation utilities based on SymPy.☆25Dec 12, 2024Updated last year
- ☆15Feb 10, 2026Updated 6 months ago
- [ACL 2024]Official GitHub repo for OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scie…☆196Jun 8, 2025Updated last year
- Official repository for ACL 2025 paper "ProcessBench: Identifying Process Errors in Mathematical Reasoning"☆192May 20, 2025Updated last year
- ☆45Jun 7, 2021Updated 5 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Source codes of ACL 2022-Efficient Cluster-based k-Nearest-Neighbor Machine Translation☆27Sep 30, 2022Updated 3 years ago
- Code for our ACL2021 paper Neural Machine Translation with Monolingual Translation Memory☆81Jun 12, 2023Updated 3 years ago
- Source codes for paper "BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity".☆19Jan 10, 2026Updated 7 months ago
- A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning☆302Sep 25, 2025Updated 11 months ago
- 💻 Terminal-Agent with Human-in-the-Loop Learning☆42Jan 16, 2026Updated 7 months ago
- [WACV 2024 LLVM-AD Challenge] UCU Dataset☆15Sep 9, 2023Updated 2 years ago
- [EMNLP 2023] Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-Thoughts☆28Nov 4, 2023Updated 2 years ago
- [WMT 2022] Implementation of TAL-SJTU's system for WMT22 English-Livonian☆23May 4, 2023Updated 3 years ago
- Feeling confused about super alignment? Here is a reading list☆42Jan 9, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Project OCELoT: an Open, Collaborative Evaluation Leaderboard of Translations☆23Jul 11, 2026Updated last month
- ☆38Jan 23, 2024Updated 2 years ago
- Academic papers and works related to SWE-bench and SWE-agents☆15Dec 8, 2025Updated 8 months ago
- ☆21Apr 16, 2025Updated last year
- This is the official repository of the paper "Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Schedulin…☆15Jul 27, 2025Updated last year
- Code for "Multi-Domain Neural Machine Translation with Word-Level Domain Context Discrimination"(EMNLP2019)☆31Nov 9, 2018Updated 7 years ago
- Lowering PyTorch's Memory Consumption for Selective Differentiation☆12Aug 29, 2024Updated 2 years ago