π€ Benchmark Large Language Models Reliably On Your Data
β457Sep 8, 2026Updated last week
Alternatives and similar repositories for yourbench
Users that are interested in yourbench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Benchmark Large Language Models Reliably On Your Dataβ19Dec 27, 2025Updated 8 months ago
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backendsβ2,545Updated this week
- Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard aβ¦β2,147Dec 3, 2025Updated 9 months ago
- Python library to use Pleias-RAG modelsβ72Jul 1, 2026Updated 2 months ago
- Tool for generating high quality Synthetic datasetsβ1,639Oct 28, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Fast Multimodal Semantic Deduplication & Filteringβ967May 24, 2026Updated 3 months ago
- Build datasets using natural languageβ590Sep 19, 2025Updated last year
- β160Dec 2, 2024Updated last year
- A lightweight, local-first, and free experiment tracking library from Hugging Face π€β1,686Updated this week
- β15Updated this week
- Everything about the SmolLM and SmolVLM family of modelsβ3,897Updated this week
- β15Apr 26, 2025Updated last year
- X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domainsβ48Feb 4, 2026Updated 7 months ago
- Train your own SOTA deductive reasoning modelβ111Mar 6, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β25Jul 15, 2026Updated 2 months ago
- A framework for few-shot evaluation of language models.β14,027Updated this week
- Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.β3,341Updated this week
- Our library for RL environments + evalsβ4,634Updated this week
- β17Oct 21, 2025Updated 10 months ago
- β16Aug 18, 2025Updated last year
- Hugging Face Inference Toolkit used to serve transformers, sentence-transformers, and diffusers models.β97Updated this week
- Let's build better datasets, together!β276Jun 9, 2026Updated 3 months ago
- Collection of scripts and notebooks for OpenAI's latest GPT OSS modelsβ507Aug 25, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A framework for few-shot evaluation of language models.β37Apr 3, 2026Updated 5 months ago
- β29Aug 21, 2025Updated last year
- Synthetic data curation for post-training and structured data extractionβ1,731Updated this week
- β13Apr 16, 2025Updated last year
- Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verifiβ¦β3,395Updated this week
- The most modern LLM evaluation toolkitβ70Apr 30, 2026Updated 4 months ago
- A course on aligning smol models.β6,751Updated this week
- β19Jul 24, 2025Updated last year
- β74Sep 27, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Curated list of datasets and tools for post-training.β4,788Apr 29, 2026Updated 4 months ago
- Lab Cookbookβ42Aug 5, 2026Updated last month
- Code for evaluating with Flow-Judge-v0.1 - an open-source, lightweight (3.8B) language model optimized for LLM system evaluations. Crafteβ¦β86Oct 29, 2024Updated last year
- π¨ NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.β2,245Updated this week
- Robust recipes to align language models with human and AI preferencesβ5,680Updated this week
- π€ smolagents: a barebones library for agents that think in code.β29,400Aug 25, 2026Updated 3 weeks ago
- Automatic evals for LLMsβ610Feb 24, 2026Updated 6 months ago