PhD/MBA-level human-annotated rubrics dataset across Physics, Chemistry, Finance and Consulting
☆33Oct 30, 2025Updated 9 months ago
Alternatives and similar repositories for ProfBench
Users that are interested in ProfBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Open source codebase for PRBench☆18Jan 15, 2026Updated 6 months ago
- ☆16Jul 10, 2023Updated 3 years ago
- A Python implementation of word2vec that allows custom sampling strategies☆10Jan 30, 2014Updated 12 years ago
- Materials for paper "Self-improving World Modelling with Latent Actions"☆20Feb 5, 2026Updated 6 months ago
- Accelerate pretraining by pre-pretraining on formal languages!☆22Feb 13, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- (ACL 2025 Main) Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification - Offici…☆21Dec 26, 2025Updated 7 months ago
- ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery☆40Jul 20, 2026Updated 3 weeks ago
- A simple GPT-based evaluation tool for multi-aspect, interpretable assessment of LLMs.☆90Jan 29, 2024Updated 2 years ago
- ACL24☆11Jun 7, 2024Updated 2 years ago
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆290Jul 31, 2026Updated last week
- Probe how GPT-n performs on statutory reasoning☆10Sep 17, 2024Updated last year
- Bayesian Inverse Graphics for Few-Shot Concept Learning☆12Mar 16, 2025Updated last year
- ☆23Jul 5, 2024Updated 2 years ago
- Fine-tuning GPT-2 to generate research paper abstracts☆12Apr 28, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆16Mar 22, 2025Updated last year
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 9 months ago
- Indranet Explorer, a simulated browser☆16Nov 12, 2024Updated last year
- Your finetuned model's back to its original safety standards faster than you can say "SafetyLock"!☆11Oct 16, 2024Updated last year
- ☆32Jul 8, 2024Updated 2 years ago
- Implementation for the paper "Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning"☆11Jan 10, 2025Updated last year
- ☆13Aug 3, 2026Updated last week
- SciQAG is a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific lit…☆34Mar 24, 2025Updated last year
- The backend behind the LLM-Perf Leaderboard☆11May 5, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- KnowMAN: Weakly Supervised Multinomial Adversarial Networks☆12Nov 9, 2021Updated 4 years ago
- ☆11Jan 19, 2025Updated last year
- ☆16Jul 23, 2024Updated 2 years ago
- [ICLR 2025] Code&Data for the paper "Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization"☆15Jun 21, 2024Updated 2 years ago
- Repository for the listwise reranker Rank-K☆16May 23, 2025Updated last year
- ☆15Mar 12, 2024Updated 2 years ago
- ☆13Sep 12, 2024Updated last year
- Smart commit messages☆18Oct 25, 2024Updated last year
- ☆24Jun 18, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆10Nov 29, 2024Updated last year
- [ACL 2022] CLUES: A Benchmark for Learning Classifiers using Natural Language Explanations☆10Jun 5, 2022Updated 4 years ago
- Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consis…☆138Updated this week
- Example for a Monty-enabled RLM in DSPy☆20Feb 16, 2026Updated 5 months ago
- A minimalist benchmarking tool designed to test the routine-generation capabilities of LLMs.☆27Nov 28, 2024Updated last year
- ☆11May 5, 2022Updated 4 years ago
- ☆17Feb 4, 2026Updated 6 months ago