Repository for NPHardEval, a quantified-dynamic benchmark of LLMs
☆65Mar 26, 2024Updated 2 years ago
Alternatives and similar repositories for NPHardEval
Users that are interested in NPHardEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Nov 25, 2024Updated last year
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- Evaluating LLMs with CommonGen-Lite☆95Mar 21, 2024Updated 2 years ago
- [ACL'24] Chain of Thought (CoT) is significant in improving the reasoning abilities of large language models (LLMs). However, the correla…☆47May 11, 2025Updated last year
- An end-to-end benchmark suite of multi-modal DNN applications for system-architecture co-design☆23Dec 13, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆12Aug 8, 2023Updated 3 years ago
- ☆12Sep 23, 2023Updated 2 years ago
- code for EMNLP2018 paper 'Associative-multichannel-autoencoder for multimodal word representation'☆13Aug 24, 2018Updated 7 years ago
- The evaluation code for the paper "MoreHopQA: More Than Multi-hop Reasoning"☆15Jun 21, 2024Updated 2 years ago
- [ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset☆116May 22, 2025Updated last year
- ☆15Sep 30, 2023Updated 2 years ago
- Minimal implementation of the Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models paper (ArXiv 20232401.01335)☆29Mar 1, 2024Updated 2 years ago
- ☆14Aug 15, 2024Updated 2 years ago
- MVA Course "Algorithms for speech and language processing", 2020, Dupoux, Zeghidour & Sagot☆18Jan 5, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- a benckmark for evaluating logical reasoning of LLMs☆23Jan 25, 2024Updated 2 years ago
- Unifew: Unified Fewshot Learning Model☆18Sep 10, 2021Updated 4 years ago
- ☆17Feb 22, 2024Updated 2 years ago
- A complete guide to evaluate LLMs and RAGs. Both theory and code based approaches covered.☆28Nov 16, 2023Updated 2 years ago
- This the implementation of LeCo☆33Jan 20, 2025Updated last year
- ☆12Oct 23, 2022Updated 3 years ago
- Code repo for "Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers" (ACL 2023)☆22Nov 1, 2023Updated 2 years ago
- This repository contains the code for the paper The Open Proof Corpus: Building a Large-Scale, Human-Validated Dataset of LLM-Generated P…☆18Aug 4, 2025Updated last year
- Metrics for "Beyond neural scaling laws: beating power law scaling via data pruning " (NeurIPS 2022 Outstanding Paper Award)☆58Apr 24, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Scalable Meta-Evaluation of LLMs as Evaluators☆43Feb 15, 2024Updated 2 years ago
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents (EMNLP Findings 2024)☆110Jan 11, 2026Updated 7 months ago
- A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models☆29Nov 25, 2024Updated last year
- AlphaVerus: Formally Verified Code Generation through Self-Improving Translation and Treefinement☆33May 14, 2025Updated last year
- Paper Implementation of Self-Rewarding Language Models☆13Feb 1, 2024Updated 2 years ago
- [ICLR 2025] Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist☆34Oct 23, 2024Updated last year
- [ACL'25] We propose a novel fine-tuning method, Separate Memory and Reasoning, which combines prompt tuning with LoRA.☆89Nov 2, 2025Updated 9 months ago
- Advanced Reasoning Benchmark Dataset for LLMs☆48Nov 19, 2023Updated 2 years ago
- Temporal Knowledge Graph Forecasting Using In-Context Learning (EMNLP 2023)☆33Jun 28, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [AAAI 2024] History Matters: Temporal Knowledge Editing in Large Language Model☆13Dec 17, 2023Updated 2 years ago
- ☆52Jul 4, 2023Updated 3 years ago
- WarAgent: LLM-based Multi-Agent Simulation of World Wars☆451Mar 5, 2024Updated 2 years ago
- Solving Inequality Proofs with Large Language Models.☆61Dec 15, 2025Updated 8 months ago
- AI for Mathematics Paper List☆17Jan 14, 2025Updated last year
- Official github repo for the paper "Compression Represents Intelligence Linearly" [COLM 2024]☆152Sep 20, 2024Updated last year
- The data and implementation for the experiments in the paper "Flows: Building Blocks of Reasoning and Collaborating AI".☆32Feb 12, 2024Updated 2 years ago