β34Jul 11, 2024Updated 2 years ago
Alternatives and similar repositories for AutoBencher
Users that are interested in AutoBencher are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- πΎ Universal, customizable and deployable fine-grained evaluation for text generation.β24Apr 22, 2026Updated 4 months ago
- β11Oct 11, 2023Updated 2 years ago
- Code for "FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge". EMNLP 2023.β20Dec 25, 2023Updated 2 years ago
- β23Jun 12, 2025Updated last year
- β13May 17, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- β19Mar 16, 2025Updated last year
- [CICAI2026] Efficiently creating diverse multi-turn Text-to-SQL training samples in 3 steps! πβ15Jul 16, 2026Updated last month
- Exploring limitations of LLM-as-a-judgeβ20Aug 17, 2024Updated 2 years ago
- The implementation of <Factual Consistency Evaluation for Text Summarization via Counterfactual Estimation> in PyTorch.β17Nov 11, 2021Updated 4 years ago
- ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind (AAAI2025)β20Apr 16, 2025Updated last year
- playing with gpt4β13Mar 17, 2023Updated 3 years ago
- The repository for papaer "Distance between Relevant Information Pieces Causes Bias in Long-Context LLMs"β14Dec 16, 2024Updated last year
- Measuring if attention is explanation with ROARβ22Mar 3, 2023Updated 3 years ago
- Codebase, data and models for the Re-Thinking the Shuffle Test paper at ACL2021β10Oct 14, 2022Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)β59May 12, 2025Updated last year
- [ICLR'26, NAACL'25 Demo] Toolkit & Benchmark for evaluating the trustworthiness of generative foundation models.β135Aug 22, 2025Updated last year
- β12Apr 29, 2024Updated 2 years ago
- β36Nov 16, 2025Updated 9 months ago
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"β10Jun 23, 2024Updated 2 years ago
- β15Oct 7, 2024Updated last year
- Explaining neural decisions contrastively to alternative decisions.β24Mar 18, 2021Updated 5 years ago
- [ArXiv 2025] Imperceptible Jailbreaking against Large Language Modelsβ25Oct 7, 2025Updated 10 months ago
- Benchmarking of 1D pattern classification networksβ11Jul 19, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Mental state inference from observable behaviorβ15Dec 3, 2021Updated 4 years ago
- Symbolic Regression from Scratch with Pythonβ14Dec 6, 2022Updated 3 years ago
- A tool for calling (and calling out to) large language models.β16Aug 13, 2024Updated 2 years ago
- β45Jan 26, 2025Updated last year
- β14Aug 30, 2023Updated 3 years ago
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generationβ14Aug 19, 2025Updated last year
- The example of correspondence between fine classes and superclasses (coarse classes) in ImageNet.β13Dec 4, 2024Updated last year
- [IJCNN2025] MMSQL: Multi-turn Multi-type text-to-SQL test suit. Repository contains scripts, code, datasets in the paper "Evaluating and β¦β21Jul 15, 2026Updated last month
- β12Jan 20, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The code implementation of the paper Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks (Aβ¦β13Jul 16, 2024Updated 2 years ago
- β22Sep 20, 2022Updated 3 years ago
- β12May 13, 2023Updated 3 years ago
- Data and code for the paper "The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems"β21Jul 18, 2023Updated 3 years ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"β19May 30, 2025Updated last year
- Fine-grained attention in hierarchical transformers for tabular time-series.β12Dec 24, 2024Updated last year
- Efficient Scaling laws and collaborative pretraining.β24Jul 19, 2026Updated last month