ABench is an evolving open-source benchmark suite designed to rigorously evaluate and enhance Large Language Models (LLMs) on complex cross-domain tasks.
☆27Apr 17, 2026Updated 3 months ago
Alternatives and similar repositories for ABench
Users that are interested in ABench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆24Aug 20, 2025Updated 11 months ago
- Unofficial implementation of the Ask-LLM paper 'How to Train Data-Efficient LLMs', arXiv:2402.09668.☆12Jun 19, 2024Updated 2 years ago
- Intro to using DSPy with Kuzu to enrich the data within the Nobel Laureate mentorship network☆16Sep 16, 2025Updated 10 months ago
- ☆44Feb 28, 2026Updated 4 months ago
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆16Aug 5, 2025Updated 11 months ago
- Official Implementation of Video-MA2MBA☆12Dec 3, 2024Updated last year
- ☆17Jul 31, 2025Updated 11 months ago
- The official repo for “Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem” [EMNLP25]☆33Sep 1, 2025Updated 10 months ago
- ☆17Apr 25, 2025Updated last year
- ☆21Jun 12, 2025Updated last year
- ☆24May 23, 2026Updated 2 months ago
- ☆22Jun 16, 2026Updated last month
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://metr.org/blog/2025-07-10-early-2025-ai-e…☆16Feb 23, 2026Updated 5 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Multi-Word Probabilistic based supertokenizer☆15May 15, 2025Updated last year
- Semantic Regex☆18Nov 13, 2025Updated 8 months ago
- Mixture-of-Basis-Experts for Compressing MoE-based LLMs☆37Dec 24, 2025Updated 7 months ago
- The training codes of Jasper-Token-Compression-600M☆20Nov 19, 2025Updated 8 months ago
- ☆26Sep 4, 2025Updated 10 months ago
- [ICLR 2026] Official Implementation of ProxyThinker: Test-Time Guidance through Small Visual Reasoners.☆22Sep 24, 2025Updated 10 months ago
- Triton kernels for dynamic causal short convolutions.☆24Jun 4, 2026Updated last month
- Repository for the paper: "TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining" ACL Oral 2025☆24Apr 19, 2026Updated 3 months ago
- Training hybrid models for dummies.☆31Nov 1, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆30Jan 5, 2026Updated 6 months ago
- M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning☆47Jul 17, 2025Updated last year
- ☆15Oct 27, 2025Updated 8 months ago
- MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models☆27May 23, 2026Updated 2 months ago
- Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI, derived from Ling.☆109Aug 5, 2025Updated 11 months ago
- A unified suite for generating elite reasoning problems and training high-performance LLMs, including pioneering attention-free architect…☆132Jan 31, 2026Updated 5 months ago
- ☆101Nov 22, 2025Updated 8 months ago
- Subliminal learning in LLMs: language models can transmit hidden preferences through seemingly unrelated training data.☆25Nov 9, 2025Updated 8 months ago
- Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI.☆272Oct 4, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- React.js Server Rendering with php-v8js and the Symfony Microkernel☆10Jun 30, 2016Updated 10 years ago
- Code for "DynaGuard: A Dynamic Guardrail Model With User-Defined Policies."☆23Nov 3, 2025Updated 8 months ago
- 自己的学习记录☆17Aug 19, 2020Updated 5 years ago
- Benchmark scripts for comparing different tokenizers and sentence segmenters of German☆12Feb 27, 2023Updated 3 years ago
- [ACL'26 Findings] Steering LLM Thinking with Budget Guidance☆33Feb 19, 2026Updated 5 months ago
- Ring-V2 is a reasoning MoE LLM provided and open-sourced by InclusionAI.☆98Oct 23, 2025Updated 9 months ago
- Measuring the Signal to Noise Ratio in Language Model Evaluation☆31Aug 19, 2025Updated 11 months ago