[ACL'24] A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
☆40Jul 19, 2024Updated 2 years ago
Alternatives and similar repositories for KIEval
Users that are interested in KIEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Aug 3, 2024Updated 2 years ago
- ☆20May 25, 2024Updated 2 years ago
- Exploiting Unlabeled Data for Target-Oriented Opinion Words Extraction☆24Sep 30, 2022Updated 3 years ago
- ☆117Jun 13, 2023Updated 3 years ago
- ☆68Feb 1, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Code and Data for GlitchBench☆13Feb 27, 2024Updated 2 years ago
- ☆15Jan 27, 2025Updated last year
- 🎉 TrustJudge is accepted to ICLR 2026!☆49Sep 27, 2025Updated 10 months ago
- The repository for paper <Evaluating Open-QA Evaluation>☆25Apr 9, 2024Updated 2 years ago
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation☆14Aug 19, 2025Updated 11 months ago
- ☆20Feb 3, 2022Updated 4 years ago
- [KDD'22] Partial Label Learning with Discrimination Augmentation☆10May 21, 2024Updated 2 years ago
- ☆472Feb 7, 2025Updated last year
- ☆23Jan 25, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Your finetuned model's back to its original safety standards faster than you can say "SafetyLock"!☆11Oct 16, 2024Updated last year
- Code for COLING 2022 paper "FactMix: Using a Few Labeled In-domain Examples to Generalize to Cross-domain Named Entity Recognition"☆15Jan 15, 2023Updated 3 years ago
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"☆10Jun 23, 2024Updated 2 years ago
- Official implementation of Bootstrapping Language Models via DPO Implicit Rewards☆48Apr 15, 2025Updated last year
- ☆12Jun 29, 2024Updated 2 years ago
- Official repository for Decentralized Arena via Collective LLM Intelligence☆18May 19, 2025Updated last year
- ☆34Jul 11, 2024Updated 2 years ago
- ☆33Jun 12, 2024Updated 2 years ago
- [NDSS'24] Inaudible Adversarial Perturbation: Manipulating the Recognition of User Speech in Real Time☆57Sep 28, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆45Jan 26, 2025Updated last year
- Distributed Reinforcement Learning for LLM Fine-Tuning with multi-GPU utilization☆22Mar 12, 2025Updated last year
- ☆12Jan 20, 2025Updated last year
- Conversational Recommender System Evaluation via Simulation☆22Aug 11, 2026Updated last week
- This repository includes the code implementation of the paper Improving Pacing in Long-Form Story Planning by Yichen Wang, Kevin Yang, Xi…☆18Nov 19, 2024Updated last year
- I-SHEEP: Iterative Self-enHancEmEnt Paradigm of LLMs through Self-Instruct and Self-Assessment☆17Jan 16, 2025Updated last year
- Resolving Knowledge Conflicts in Large Language Models, COLM 2024☆18Oct 7, 2025Updated 10 months ago
- [NeurIPS 2024] A Novel Rank-Based Metric for Evaluating Large Language Models☆59May 28, 2025Updated last year
- Clean, extensible implementation of MACAW [ICML 2021]☆12Dec 7, 2021Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- GraphDancer: Training LLMs to Explore and Reason over Graphs via Curriculum Reinforcement Learning☆20May 25, 2026Updated 2 months ago
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models, COLM 2026☆17Jun 28, 2025Updated last year
- Repository for the ACL 2024 conference website☆18Feb 3, 2025Updated last year
- ☆10Oct 22, 2024Updated last year
- ☆49Sep 5, 2024Updated last year
- Repository of paper "How Likely Do LLMs with CoT Mimic Human Reasoning?"☆23Feb 19, 2025Updated last year
- NeurIPS 2025☆15Feb 4, 2026Updated 6 months ago