β33Jul 11, 2024Updated 2 years ago
Alternatives and similar repositories for AutoBencher
Users that are interested in AutoBencher are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- πΎ Universal, customizable and deployable fine-grained evaluation for text generation.β24Apr 22, 2026Updated 3 months ago
- Improving Your Model Ranking on Chatbot Arena by Vote Rigging (ICML 2025)β27Feb 25, 2025Updated last year
- β11Oct 11, 2023Updated 2 years ago
- Code for "FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge". EMNLP 2023.β20Dec 25, 2023Updated 2 years ago
- β21Jun 12, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This is a new metric that can be used to evaluate faithfulness of text generated by LLMs. The work behind this repository can be found heβ¦β31Aug 25, 2023Updated 2 years ago
- Implementation of the paper "FactGraph: Evaluating Factuality in Summarization with Semantic Graph Representations (NAACL 2022)"β52Jul 26, 2023Updated 3 years ago
- β13May 17, 2025Updated last year
- β19Mar 16, 2025Updated last year
- [CICAI2026] Efficiently creating diverse multi-turn Text-to-SQL training samples in 3 steps! πβ15Jul 16, 2026Updated 3 weeks ago
- Exploring limitations of LLM-as-a-judgeβ20Aug 17, 2024Updated last year
- Codebase for Instruction Following without Instruction Tuningβ36Sep 24, 2024Updated last year
- β24Mar 8, 2024Updated 2 years ago
- playing with gpt4β13Mar 17, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The repository for papaer "Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs"β14Dec 16, 2024Updated last year
- Measuring if attention is explanation with ROARβ22Mar 3, 2023Updated 3 years ago
- Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)β59May 12, 2025Updated last year
- β12Apr 29, 2024Updated 2 years ago
- β34Nov 16, 2025Updated 8 months ago
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"β10Jun 23, 2024Updated 2 years ago
- β13Aug 14, 2022Updated 4 years ago
- [ArXiv 2025] Imperceptible Jailbreaking against Large Language Modelsβ25Oct 7, 2025Updated 10 months ago
- Benchmarking of 1D pattern classification networksβ11Jul 19, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Code used to run experiments for the ICLR 2023 paper "Computational Language Acquisition with Theory of Mind".β15Apr 27, 2023Updated 3 years ago
- β44Jan 26, 2025Updated last year
- β14Aug 30, 2023Updated 2 years ago
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generationβ14Aug 19, 2025Updated 11 months ago
- The example of correspondence between fine classes and superclasses (coarse classes) in ImageNet.β13Dec 4, 2024Updated last year
- [IJCNN2025] MMSQL: Multi-turn Multi-type text-to-SQL test suit. Repository contains scripts, code, datasets in the paper "Evaluating and β¦β21Jul 15, 2026Updated 3 weeks ago
- β12Jan 20, 2025Updated last year
- β22Sep 20, 2022Updated 3 years ago
- β12May 13, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"β19May 30, 2025Updated last year
- Code for the paper "Symmetric Machine Theory of Mind", presented at ICML 2022.β12Jul 18, 2022Updated 4 years ago
- Code for Columbia University COMS 3997 β LLM Ethics and Foundationsβ16Jan 7, 2025Updated last year
- This repo contains the ToMnet+ model for preference inference. Developed by Yun-Shiuan, Edwinn, Hsin-Yi, and Elaine.β10Feb 24, 2023Updated 3 years ago
- This repository includes the code implementation of the paper Improving Pacing in Long-Form Story Planning by Yichen Wang, Kevin Yang, Xiβ¦β18Nov 19, 2024Updated last year
- Resolving Knowledge Conflicts in Large Language Models, COLM 2024β18Oct 7, 2025Updated 10 months ago
- Official implementation for "Parameter-Efficient Fine-Tuning Design Spaces"β27Jan 4, 2023Updated 3 years ago