β33Jul 11, 2024Updated 2 years ago
Alternatives and similar repositories for AutoBencher
Users that are interested in AutoBencher are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- πΎ Universal, customizable and deployable fine-grained evaluation for text generation.β24Apr 22, 2026Updated 3 months ago
- Code for "FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge". EMNLP 2023.β20Dec 25, 2023Updated 2 years ago
- [ICLR'25] DataGen: Unified Synthetic Dataset Generation via Large Language Modelsβ69Mar 8, 2025Updated last year
- This is a new metric that can be used to evaluate faithfulness of text generated by LLMs. The work behind this repository can be found heβ¦β31Aug 25, 2023Updated 2 years ago
- Implementation of the paper "FactGraph: Evaluating Factuality in Summarization with Semantic Graph Representations (NAACL 2022)"β52Jul 26, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- β13May 17, 2025Updated last year
- β19Mar 16, 2025Updated last year
- [CICAI2026] Efficiently creating diverse multi-turn Text-to-SQL training samples in 3 steps! πβ15Jul 16, 2026Updated last week
- Exploring limitations of LLM-as-a-judgeβ20Aug 17, 2024Updated last year
- β24Mar 8, 2024Updated 2 years ago
- The implementation of <Factual Consistency Evaluation for Text Summarization via Counterfactual Estimation> in PyTorch.β17Nov 11, 2021Updated 4 years ago
- ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind (AAAI2025)β20Apr 16, 2025Updated last year
- playing with gpt4β13Mar 17, 2023Updated 3 years ago
- The repository for papaer "Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs"β14Dec 16, 2024Updated last year
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Codebase, data and models for the Re-Thinking the Shuffle Test paper at ACL2021β10Oct 14, 2022Updated 3 years ago
- Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)β58May 12, 2025Updated last year
- [ICLR'26, NAACL'25 Demo] Toolkit & Benchmark for evaluating the trustworthiness of generative foundation models.β132Aug 22, 2025Updated 11 months ago
- β13May 7, 2023Updated 3 years ago
- β12Apr 29, 2024Updated 2 years ago
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"β10Jun 23, 2024Updated 2 years ago
- β15Oct 7, 2024Updated last year
- Benchmarking of 1D pattern classification networksβ11Jul 19, 2023Updated 3 years ago
- Simple phoenix setup for padded window managementβ13Apr 25, 2018Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Mental state inference from observable behaviorβ15Dec 3, 2021Updated 4 years ago
- Symbolic Regression from Scratch with Pythonβ14Dec 6, 2022Updated 3 years ago
- Code used to run experiments for the ICLR 2023 paper "Computational Language Acquisition with Theory of Mind".β15Apr 27, 2023Updated 3 years ago
- β43Jan 26, 2025Updated last year
- β14Aug 30, 2023Updated 2 years ago
- The example of correspondence between fine classes and superclasses (coarse classes) in ImageNet.β13Dec 4, 2024Updated last year
- [IJCNN2025] MMSQL: Multi-turn Multi-type text-to-SQL test suit. Repository contains scripts, code, datasets in the paper "Evaluating and β¦β21Jul 15, 2026Updated last week
- β12Jan 20, 2025Updated last year
- Code and instructions accompanying ICCV'23 paper Protoype-based Dataset Comparisonβ18Dec 15, 2023Updated 2 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- The code implementation of the paper Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks (Aβ¦β13Jul 16, 2024Updated 2 years ago
- β22Sep 20, 2022Updated 3 years ago
- The official repository of 'Unnatural Language Are Not Bugs but Features for LLMs'β24May 20, 2025Updated last year
- β12May 13, 2023Updated 3 years ago
- Data and code for the paper "The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems"β21Jul 18, 2023Updated 3 years ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"β19May 30, 2025Updated last year
- Fine-grained attention in hierarchical transformers for tabular time-series.β12Dec 24, 2024Updated last year