Repository for NPHardEval, a quantified-dynamic benchmark of LLMs
☆66Mar 26, 2024Updated 2 years ago
Alternatives and similar repositories for NPHardEval
Users that are interested in NPHardEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Mar 9, 2025Updated last year
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- Implementation of the model: "Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models" in PyTorch☆29Aug 29, 2026Updated last week
- Evaluating LLMs with CommonGen-Lite☆95Mar 21, 2024Updated 2 years ago
- [ACL'24] Chain of Thought (CoT) is significant in improving the reasoning abilities of large language models (LLMs). However, the correla…☆47May 11, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- An end-to-end benchmark suite of multi-modal DNN applications for system-architecture co-design☆23Dec 13, 2024Updated last year
- ☆12Aug 8, 2023Updated 3 years ago
- Codes for the paper "CausalCite: A Causal Formulation of Paper Citations" (2023)☆16Jan 11, 2024Updated 2 years ago
- ☆12Sep 23, 2023Updated 2 years ago
- code for EMNLP2018 paper 'Associative-multichannel-autoencoder for multimodal word representation'☆13Aug 24, 2018Updated 8 years ago
- The evaluation code for the paper "MoreHopQA: More Than Multi-hop Reasoning"☆15Jun 21, 2024Updated 2 years ago
- ☆15Sep 30, 2023Updated 2 years ago
- [ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset☆116May 22, 2025Updated last year
- Minimal implementation of the Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models paper (ArXiv 20232401.01335)☆29Mar 1, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆14Aug 15, 2024Updated 2 years ago
- a benckmark for evaluating logical reasoning of LLMs☆23Jan 25, 2024Updated 2 years ago
- 儿童故事常识推理与寓意理解评测(Commonsense Reasoning and Moral Understanding Evaluation in Children's Stories,CRMU)☆18Oct 22, 2024Updated last year
- Unifew: Unified Fewshot Learning Model☆18Sep 10, 2021Updated 4 years ago
- ☆17Feb 22, 2024Updated 2 years ago
- TrustAgent: Towards Safe and Trustworthy LLM-based Agents☆60Feb 7, 2025Updated last year
- A complete guide to evaluate LLMs and RAGs. Both theory and code based approaches covered.☆27Nov 16, 2023Updated 2 years ago
- EmojiCrypt: Prompt Encryption for Secure Communication with Large Language Models☆27Feb 21, 2024Updated 2 years ago
- This the implementation of LeCo☆33Jan 20, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆12Oct 23, 2022Updated 3 years ago
- Neural Collaborative Reasoning☆30Jun 23, 2021Updated 5 years ago
- Bayes-Adaptive RL for LLM Reasoning☆46May 28, 2025Updated last year
- Official repository for ACL 2025 paper "Model Extrapolation Expedites Alignment"☆75May 20, 2025Updated last year
- Repository for Skill Set Optimization☆14Jul 26, 2024Updated 2 years ago
- Code repo for "Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers" (ACL 2023)☆22Nov 1, 2023Updated 2 years ago
- [ECCV 2024] M3DBench introduces a comprehensive 3D instruction-following dataset with support for interleaved multi-modal prompts.☆61Oct 1, 2024Updated last year
- This repository contains the code for the paper The Open Proof Corpus: Building a Large-Scale, Human-Validated Dataset of LLM-Generated P…☆18Aug 4, 2025Updated last year
- Metrics for "Beyond neural scaling laws: beating power law scaling via data pruning " (NeurIPS 2022 Outstanding Paper Award)☆58Apr 24, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"☆105Jun 15, 2023Updated 3 years ago
- Scalable Meta-Evaluation of LLMs as Evaluators☆43Feb 15, 2024Updated 2 years ago
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents (EMNLP Findings 2024)☆112Jan 11, 2026Updated 7 months ago
- A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models☆29Nov 25, 2024Updated last year
- Paper Implementation of Self-Rewarding Language Models☆13Feb 1, 2024Updated 2 years ago
- [ICLR 2025] Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist☆34Oct 23, 2024Updated last year
- Advanced Reasoning Benchmark Dataset for LLMs☆48Nov 19, 2023Updated 2 years ago