[ACL'24] A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
☆40Jul 19, 2024Updated 2 years ago
Alternatives and similar repositories for KIEval
Users that are interested in KIEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Exploiting Unlabeled Data for Target-Oriented Opinion Words Extraction☆24Sep 30, 2022Updated 3 years ago
- ☆17Feb 28, 2024Updated 2 years ago
- Improving fast adversarial training with prior-guided knowledge (TPAMI2024)☆43Apr 21, 2024Updated 2 years ago
- Code and Data for GlitchBench☆13Feb 27, 2024Updated 2 years ago
- ☆15Jan 27, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 🎉 TrustJudge is accepted to ICLR 2026!☆49Sep 27, 2025Updated 10 months ago
- ☆926May 22, 2024Updated 2 years ago
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation☆14Aug 19, 2025Updated 11 months ago
- Improved techniques for optimization-based jailbreaking on large language models (ICLR2025)☆146Apr 7, 2025Updated last year
- ☆20Feb 3, 2022Updated 4 years ago
- [KDD'22] Partial Label Learning with Discrimination Augmentation☆10May 21, 2024Updated 2 years ago
- ☆23Jan 25, 2023Updated 3 years ago
- playing with gpt4☆13Mar 17, 2023Updated 3 years ago
- Your finetuned model's back to its original safety standards faster than you can say "SafetyLock"!☆11Oct 16, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for COLING 2022 paper "FactMix: Using a Few Labeled In-domain Examples to Generalize to Cross-domain Named Entity Recognition"☆15Jan 15, 2023Updated 3 years ago
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"☆10Jun 23, 2024Updated 2 years ago
- Official implementation of Bootstrapping Language Models via DPO Implicit Rewards☆47Apr 15, 2025Updated last year
- We leverage 14 datasets as OOD test data and conduct evaluations on 8 NLU tasks over 21 popularly used models. Our findings confirm that …☆100Aug 15, 2023Updated 2 years ago
- [ACL'26 Main Conference] Instruction Data Selection via Answer Divergence☆22Apr 14, 2026Updated 3 months ago
- Official repository for Decentralized Arena via Collective LLM Intelligence☆18May 19, 2025Updated last year
- [NDSS'24] Inaudible Adversarial Perturbation: Manipulating the Recognition of User Speech in Real Time☆56Sep 28, 2024Updated last year
- ☆44Jan 26, 2025Updated last year
- ☆14Aug 30, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆52Oct 24, 2023Updated 2 years ago
- The code implementation of the paper Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks (A…☆13Jul 16, 2024Updated 2 years ago
- Code for "FactKB: Generalizable Factuality Evaluation using Language Models Enhanced with Factual Knowledge". EMNLP 2023.☆20Dec 25, 2023Updated 2 years ago
- Conversational Recommender System Evaluation via Simulation☆22Jul 21, 2026Updated last week
- I-SHEEP: Iterative Self-enHancEmEnt Paradigm of LLMs through Self-Instruct and Self-Assessment☆17Jan 16, 2025Updated last year
- Resolving Knowledge Conflicts in Large Language Models, COLM 2024☆18Oct 7, 2025Updated 9 months ago
- This GitHub provides the source code for the paper "Exploring Facial Expression and Action Units in Parkinson Disease"☆10Dec 21, 2022Updated 3 years ago
- [NeurIPS 2024] A Novel Rank-Based Metric for Evaluating Large Language Models☆59May 28, 2025Updated last year
- Cue-CoT: Chain-of-thought Prompting for Responding to In-depth Dialogue Questions with LLMs [EMNLP 2023 Findings]☆24Nov 18, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models, COLM 2026☆17Jun 28, 2025Updated last year
- [AAAI'26, Oral] Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAG☆32Apr 14, 2026Updated 3 months ago
- Repository for the ACL 2024 conference website☆18Feb 3, 2025Updated last year
- ☆10Oct 22, 2024Updated last year
- ☆48Sep 5, 2024Updated last year
- Repository of paper "How Likely Do LLMs with CoT Mimic Human Reasoning?"☆23Feb 19, 2025Updated last year
- JPEG-LM: LLMs as Image Generators with Canonical Codec Representations☆16Sep 29, 2024Updated last year