An easy python package to run quick basic QA evaluations. This package includes standardized QA evaluation metrics and semantic evaluation metrics: Black-box and Open-Source large language model prompting and evaluation, exact match, F1 Score, PEDANT semantic match, transformer match. Our package also supports prompting OPENAI and Anthropic API.
☆61Jul 18, 2025Updated last year
Alternatives and similar repositories for qa_metrics
Users that are interested in qa_metrics are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Synthetic Video hallucination and Mitigation☆23Sep 21, 2025Updated 10 months ago
- Reinforcement Learning of Vision Language Models with Self Visual Perception Reward☆175Mar 14, 2026Updated 4 months ago
- Official Implementation of "Learning to Refuse: Towards Mitigating Privacy Risks in LLMs"☆10Dec 13, 2024Updated last year
- Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging. Arxiv, 2024.☆16Oct 28, 2024Updated last year
- Butler 是一个用于自动化服务管理和任务调度的工具项目。☆17Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆19Jun 21, 2025Updated last year
- [NeurIPS 2024 D&B] Evaluating Copyright Takedown Methods for Language Models☆17Jul 17, 2024Updated 2 years ago
- 🌏 UI component library for the future, based on WebComponent.☆23Nov 12, 2024Updated last year
- ☆27Oct 6, 2024Updated last year
- Official repo for EMNLP'24 paper "SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning"☆30Oct 1, 2024Updated last year
- [CVPRW'26] A collection and survey of 3d dataset☆33Jun 4, 2026Updated last month
- Generated geosite.dat based on Antifilter Community List☆29Updated this week
- ☆33Aug 9, 2024Updated last year
- Source code for Jordan Boyd-Graber's academic webpage.☆12Jul 5, 2026Updated 2 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆47Oct 1, 2024Updated last year
- Converter for EN16931 invoices from CII to UBL☆45Updated this week
- Official repository for the paper "ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming"☆60Sep 20, 2024Updated last year
- Code and data for the FACTOR paper☆54Nov 15, 2023Updated 2 years ago
- This is the official code for the paper "Vaccine: Perturbation-aware Alignment for Large Language Models" (NeurIPS2024)☆51Jan 15, 2026Updated 6 months ago
- “悟道”数据☆51Jul 5, 2021Updated 5 years ago
- TCM Lingdan LLM☆51Jun 1, 2026Updated last month
- Notebooks for JHU EN 601.320/420/620☆10May 1, 2019Updated 7 years ago
- ☆70Apr 14, 2023Updated 3 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆76Jul 15, 2024Updated 2 years ago
- Knowledge-Based System'25☆11Dec 15, 2024Updated last year
- Tutorial: Introduction to CrewAI☆84Jul 5, 2024Updated 2 years ago
- ☆15Apr 19, 2021Updated 5 years ago
- Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models☆91Apr 4, 2024Updated 2 years ago
- Programs for Microsoft Academic Graph☆16Jun 8, 2016Updated 10 years ago
- Dataset associated with "BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation" paper☆88Mar 2, 2021Updated 5 years ago
- Corpus-based Set Expansion with Lexical Features and Distributed Representations (SIGIR '19)☆13Jul 18, 2019Updated 7 years ago
- Dataset for TACL 2022 paper: "FeTaQA: Free-form Table Question Answering"☆90May 11, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- SR²AM: Efficient Agentic Reasoning Through Self-Regulated Simulative Planning☆21May 22, 2026Updated 2 months ago
- AutoHallusion Codebase (EMNLP 2024)☆23Dec 6, 2024Updated last year
- Code for paper "Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language?"☆21Oct 13, 2020Updated 5 years ago
- ☆136May 7, 2026Updated 2 months ago
- ☆15May 19, 2026Updated 2 months ago
- ChiMed-GPT is a Chinese medical large language model (LLM) built by continually training Ziya-v2 on Chinese medical data, where pre-train…☆106Dec 29, 2023Updated 2 years ago
- ☆129Jan 22, 2024Updated 2 years ago