This repository stems from our paper, “Cataloguing LLM Evaluations”, and serves as a living, collaborative catalogue of LLM evaluation frameworks, benchmarks and papers.
☆23Nov 16, 2023Updated 2 years ago
Alternatives and similar repositories for LLM-Evals-Catalogue
Users that are interested in LLM-Evals-Catalogue are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Moonshot - A simple and modular tool to evaluate and red-team any LLM application.☆349Jun 10, 2026Updated 2 months ago
- ☆29Aug 30, 2026Updated last week
- Comparison of Metaflow, MLFlow and DVC☆14Aug 4, 2021Updated 5 years ago
- [SDM22] PyTorch implementation for "Neural Graph Matching for Pre-training Graph Neural Networks".☆18Apr 5, 2022Updated 4 years ago
- [EMNLP 2022] Code for our paper “ZeroGen: Efficient Zero-shot Learning via Dataset Generation”.☆47Feb 18, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- WhatsApp chatbot with Dialogflow and Twilio api☆10May 6, 2024Updated 2 years ago
- Example showing unstructured.io + timescaledb + PGAI☆18Nov 15, 2024Updated last year
- Oscillink — Self‑Optimizing Coherent Memory for Embedding Workflows☆15Nov 24, 2025Updated 9 months ago
- [NeurIPS 2022] Explaining Graph Neural Networks with Structure-Aware Cooperative Games (GStarX)☆15Oct 20, 2022Updated 3 years ago
- Reading comprehension based question-answering model for news articles.☆11Jun 22, 2022Updated 4 years ago
- Sample web applications built using FastAPI and Cloud SQL☆15May 5, 2025Updated last year
- ☆11Feb 3, 2025Updated last year
- Distributed parallel computing on data.table☆35Apr 30, 2016Updated 10 years ago
- TYPO3 Extension ⇢ Integration of sendinblue as finisher of the form extension☆11Jan 23, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆10Jun 5, 2021Updated 5 years ago
- Simulator of a VC Portfolio☆12Apr 23, 2025Updated last year
- [Preprint] On the Effectiveness of Mitigating Data Poisoning Attacks with Gradient Shaping☆10Feb 27, 2020Updated 6 years ago
- ☆21Updated this week
- Code of our paper "Method-Level Bug Severity Prediction using Source Code Metrics and LLMs" which is accepted to ISSRE 2023.☆10Nov 12, 2023Updated 2 years ago
- CLUE: A Clinical Language Understanding Evaluation for LLMs☆24Jan 22, 2025Updated last year
- Abstractive text summarization http://arxiv.org/abs/1509.00685☆24Mar 18, 2016Updated 10 years ago
- ☆17Jun 5, 2024Updated 2 years ago
- This repository accompanies our paper “Do Prompt-Based Models Really Understand the Meaning of Their Prompts?”☆85May 10, 2022Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆15May 5, 2026Updated 4 months ago
- Pre-trained Online Contrastive Learning for Insurance Fraud Detection☆15Jul 12, 2024Updated 2 years ago
- Sample notebooks and prompts for LLM evaluation☆176Nov 2, 2025Updated 10 months ago
- ☆16May 15, 2025Updated last year
- ☆13Oct 20, 2022Updated 3 years ago
- Documenting large text datasets 🖼️ 📚☆14Dec 17, 2024Updated last year
- Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning with LLMs☆41Jan 30, 2024Updated 2 years ago
- CSCW 2023 Best Demo Award: Conversational AI Explanations to Support Human-AI Scientific Writing☆15Jun 25, 2023Updated 3 years ago
- Simple WPF HL7 Fluentpath Tester tool☆20Jun 22, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Agent Skills Collection☆22Aug 1, 2026Updated last month
- Repository for my LLM notebooks☆30Aug 8, 2024Updated 2 years ago
- [EMNLP 2022] Code for our paper “ZeroGen: Efficient Zero-shot Learning via Dataset Generation”.☆16Feb 18, 2022Updated 4 years ago
- ☆15Feb 21, 2024Updated 2 years ago
- An open-source framework for evaluating out-of-knowledge base robustness☆28Feb 11, 2026Updated 6 months ago
- source for llmsec.net☆16Jul 24, 2024Updated 2 years ago
- ☆11Oct 22, 2020Updated 5 years ago