PhD/MBA-level human-annotated rubrics dataset across Physics, Chemistry, Finance and Consulting
☆34Oct 30, 2025Updated 10 months ago
Alternatives and similar repositories for ProfBench
Users that are interested in ProfBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Open source codebase for PRBench☆20Jan 15, 2026Updated 7 months ago
- 毕业设计简单酒店信息管理系统,使用Django框架、MySql数据库。☆10Dec 9, 2021Updated 4 years ago
- ☆16Jul 10, 2023Updated 3 years ago
- Materials for paper "Self-improving World Modelling with Latent Actions"☆20Feb 5, 2026Updated 6 months ago
- (ACL 2025 Main) Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification - Offici…☆21Dec 26, 2025Updated 8 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery☆40Jul 20, 2026Updated last month
- A simple GPT-based evaluation tool for multi-aspect, interpretable assessment of LLMs.☆90Jan 29, 2024Updated 2 years ago
- ACL24☆11Jun 7, 2024Updated 2 years ago
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆307Aug 21, 2026Updated last week
- Probe how GPT-n performs on statutory reasoning☆10Sep 17, 2024Updated last year
- ☆10Mar 5, 2024Updated 2 years ago
- Bayesian Inverse Graphics for Few-Shot Concept Learning☆12Mar 16, 2025Updated last year
- ☆23Jul 5, 2024Updated 2 years ago
- Fine-tuning GPT-2 to generate research paper abstracts☆12Apr 28, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆16Mar 22, 2025Updated last year
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 9 months ago
- Indranet Explorer, a simulated browser☆16Nov 12, 2024Updated last year
- ☆32Jul 8, 2024Updated 2 years ago
- A super simple html + css + js ollama chat interface for hacking on.☆16May 1, 2025Updated last year
- ☆12Aug 6, 2024Updated 2 years ago
- Implementation for the paper "Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning"☆11Jan 10, 2025Updated last year
- ☆13Updated this week
- SciQAG is a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific lit…☆34Mar 24, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The backend behind the LLM-Perf Leaderboard☆11May 5, 2024Updated 2 years ago
- ☆16Jul 23, 2024Updated 2 years ago
- [ICLR 2025] Code&Data for the paper "Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization"☆15Jun 21, 2024Updated 2 years ago
- A toolkit for testing and improving named entity recognition [ESEC/FSE'23]☆11Aug 31, 2023Updated 3 years ago
- Repository for the listwise reranker Rank-K☆16May 23, 2025Updated last year
- Tasks for describing differences between text distributions.☆17Aug 9, 2024Updated 2 years ago
- ☆48Dec 16, 2025Updated 8 months ago
- ☆13Sep 12, 2024Updated last year
- A git repository with hands-on material for Data Science on KSL plattaforms☆14Oct 27, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Smart commit messages☆18Oct 25, 2024Updated last year
- ☆24Jun 18, 2025Updated last year
- ☆10Nov 29, 2024Updated last year
- mReasoner is a unified computational implementation of the model theory of thinking and reasoning☆16Aug 17, 2023Updated 3 years ago
- A minimalist benchmarking tool designed to test the routine-generation capabilities of LLMs.☆27Nov 28, 2024Updated last year
- ☆11May 5, 2022Updated 4 years ago
- implementation of paper "Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners"☆20Aug 17, 2023Updated 3 years ago