β34Jul 11, 2024Updated 2 years ago
Alternatives and similar repositories for AutoBencher
Users that are interested in AutoBencher are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- πΎ Universal, customizable and deployable fine-grained evaluation for text generation.β24Apr 22, 2026Updated 5 months ago
- β11Oct 11, 2023Updated 2 years ago
- We systematically studied the influencing factors when LLM generates benchmarks,By using our code, you can generate high-quality QA datasβ¦β20May 20, 2025Updated last year
- [ICLR'25] DataGen: Unified Synthetic Dataset Generation via Large Language Modelsβ69Mar 8, 2025Updated last year
- β23Jun 12, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This is a new metric that can be used to evaluate faithfulness of text generated by LLMs. The work behind this repository can be found heβ¦β31Aug 25, 2023Updated 3 years ago
- Implementation of the paper "FactGraph: Evaluating Factuality in Summarization with Semantic Graph Representations (NAACL 2022)"β52Jul 26, 2023Updated 3 years ago
- β13May 17, 2025Updated last year
- β19Mar 16, 2025Updated last year
- The repository for papaer "Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs"β14Dec 16, 2024Updated last year
- Measuring if attention is explanation with ROARβ22Mar 3, 2023Updated 3 years ago
- Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)β59May 12, 2025Updated last year
- [ICLR'26, NAACL'25 Demo] Toolkit & Benchmark for evaluating the trustworthiness of generative foundation models.β137Aug 22, 2025Updated last year
- β12Apr 29, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"β10Jun 23, 2024Updated 2 years ago
- β13Aug 14, 2022Updated 4 years ago
- β15Oct 7, 2024Updated last year
- Explaining neural decisions contrastively to alternative decisions.β24Mar 18, 2021Updated 5 years ago
- [ArXiv 2025] Imperceptible Jailbreaking against Large Language Modelsβ25Oct 7, 2025Updated 11 months ago
- Benchmarking of 1D pattern classification networksβ11Jul 19, 2023Updated 3 years ago
- Mental state inference from observable behaviorβ15Dec 3, 2021Updated 4 years ago
- Symbolic Regression from Scratch with Pythonβ14Dec 6, 2022Updated 3 years ago
- Code used to run experiments for the ICLR 2023 paper "Computational Language Acquisition with Theory of Mind".β15Apr 27, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- β14Aug 30, 2023Updated 3 years ago
- [IJCNN2025] MMSQL: Multi-turn Multi-type text-to-SQL test suit. Repository contains scripts, code, datasets in the paper "Evaluating and β¦β21Jul 15, 2026Updated 2 months ago
- β12Jan 20, 2025Updated last year
- Code and instructions accompanying ICCV'23 paper Protoype-based Dataset Comparisonβ19Dec 15, 2023Updated 2 years ago
- Data for discourse connective prediction.β12May 3, 2018Updated 8 years ago
- The code implementation of the paper Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks (Aβ¦β13Jul 16, 2024Updated 2 years ago
- β22Sep 20, 2022Updated 4 years ago
- The official repository of 'Unnatural Language Are Not Bugs but Features for LLMs'β25May 20, 2025Updated last year
- β12May 13, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Data and code for the paper "The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems"β22Jul 18, 2023Updated 3 years ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"β19May 30, 2025Updated last year
- Fine-grained attention in hierarchical transformers for tabular time-series.β12Dec 24, 2024Updated last year
- Efficient Scaling laws and collaborative pretraining.β24Jul 19, 2026Updated 2 months ago
- Code for the paper "Symmetric Machine Theory of Mind", presented at ICML 2022.β12Jul 18, 2022Updated 4 years ago
- Code for Columbia University COMS 3997 β LLM Ethics and Foundationsβ16Jan 7, 2025Updated last year
- This repository includes the code implementation of the paper Improving Pacing in Long-Form Story Planning by Yichen Wang, Kevin Yang, Xiβ¦β18Nov 19, 2024Updated last year