π€ Benchmark Large Language Models Reliably On Your Data
β458Sep 8, 2026Updated last month
Alternatives and similar repositories for yourbench
Users that are interested in yourbench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Benchmark Large Language Models Reliably On Your Dataβ19Dec 27, 2025Updated 9 months ago
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backendsβ2,553Updated this week
- Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard aβ¦β2,153Dec 3, 2025Updated 10 months ago
- Python library to use Pleias-RAG modelsβ72Jul 1, 2026Updated 3 months ago
- Tool for generating high quality Synthetic datasetsβ1,646Oct 28, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fast Multimodal Semantic Deduplication & Filteringβ974Oct 4, 2026Updated last week
- Build datasets using natural languageβ591Sep 19, 2025Updated last year
- β160Dec 2, 2024Updated last year
- A lightweight, local-first, and free experiment tracking library from Hugging Face π€β1,716Updated this week
- β15Sep 23, 2026Updated 2 weeks ago
- Everything about the SmolLM and SmolVLM family of modelsβ3,918Updated this week
- β15Apr 26, 2025Updated last year
- X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domainsβ48Feb 4, 2026Updated 8 months ago
- Train your own SOTA deductive reasoning modelβ111Mar 6, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- β25Jul 15, 2026Updated 2 months ago
- A framework for few-shot evaluation of language models.β14,176Sep 14, 2026Updated 3 weeks ago
- Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.β3,378Updated this week
- Our library for RL environments + evalsβ4,686Updated this week
- β17Oct 21, 2025Updated 11 months ago
- β16Aug 18, 2025Updated last year
- Hugging Face Inference Toolkit used to serve transformers, sentence-transformers, and diffusers models.β97Updated this week
- Let's build better datasets, together!β276Sep 22, 2026Updated 2 weeks ago
- Collection of scripts and notebooks for OpenAI's latest GPT OSS modelsβ509Aug 25, 2025Updated last year
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- A framework for few-shot evaluation of language models.β37Apr 3, 2026Updated 6 months ago
- β29Aug 21, 2025Updated last year
- Synthetic data curation for post-training and structured data extractionβ1,743Updated this week
- β13Apr 16, 2025Updated last year
- Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verifiβ¦β3,416Updated this week
- The most modern LLM evaluation toolkitβ70Apr 30, 2026Updated 5 months ago
- β19Jul 24, 2025Updated last year
- A course on aligning smol models.β6,770Updated this week
- β74Sep 27, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Curated list of datasets and tools for post-training.β4,801Apr 29, 2026Updated 5 months ago
- Lab Cookbookβ44Aug 5, 2026Updated 2 months ago
- Code for evaluating with Flow-Judge-v0.1 - an open-source, lightweight (3.8B) language model optimized for LLM system evaluations. Crafteβ¦β86Oct 29, 2024Updated last year
- π¨ NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.β2,319Updated this week
- Robust recipes to align language models with human and AI preferencesβ5,692Sep 23, 2026Updated 2 weeks ago
- π€ smolagents: a barebones library for agents that think in code.β29,769Updated this week
- Automatic evals for LLMsβ613Feb 24, 2026Updated 7 months ago