[ACL'24] A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
☆40Jul 19, 2024Updated 2 years ago
Alternatives and similar repositories for KIEval
Users that are interested in KIEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Exploiting Unlabeled Data for Target-Oriented Opinion Words Extraction☆24Sep 30, 2022Updated 3 years ago
- ☆117Jun 13, 2023Updated 3 years ago
- ☆17Feb 28, 2024Updated 2 years ago
- Code and Data for GlitchBench☆13Feb 27, 2024Updated 2 years ago
- ☆15Jan 27, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- 🎉 TrustJudge is accepted to ICLR 2026!☆51Sep 27, 2025Updated last year
- ☆924May 22, 2024Updated 2 years ago
- Exploring CoT-Decoding from Google DeepMind's paper, "Chain-of-Thought Reasoning Without Prompting".☆13Feb 22, 2024Updated 2 years ago
- A Pytorch (support batch and channel) implementation of GoogleBrain's SpecAugment: A Simple Data Augmentation Method for Automatic Speech…☆11Jul 24, 2024Updated 2 years ago
- ☆23Jan 25, 2023Updated 3 years ago
- [CVPR'26] TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding☆26Jan 4, 2026Updated 8 months ago
- playing with gpt4☆13Mar 17, 2023Updated 3 years ago
- Your finetuned model's back to its original safety standards faster than you can say "SafetyLock"!☆11Oct 16, 2024Updated last year
- The offical code for paper "What Constitutes a Faithful Summary? Preserving Author Perspectives in News Summarization"☆10Jun 23, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official implementation of Bootstrapping Language Models via DPO Implicit Rewards☆49Apr 15, 2025Updated last year
- ☆33May 31, 2024Updated 2 years ago
- [ACL'26 Main Conference] Instruction Data Selection via Answer Divergence☆25Apr 14, 2026Updated 5 months ago
- Official repository for Decentralized Arena via Collective LLM Intelligence☆18May 19, 2025Updated last year
- ☆34Jul 11, 2024Updated 2 years ago
- ☆33Jun 12, 2024Updated 2 years ago
- ☆14Aug 30, 2023Updated 3 years ago
- Leveraging ChatGPT for Text Data Augmentation☆54Sep 21, 2024Updated 2 years ago
- ☆12Jan 20, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The code implementation of the paper Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks (A…☆13Jul 16, 2024Updated 2 years ago
- The code for "Past, Present, and Future: Conversational Emotion Recognition through Structural Modeling of Psychological Commonsense Know…☆20May 22, 2022Updated 4 years ago
- Official repository for ICLR 2024 Spotlight paper "Large Language Models Are Not Robust Multiple Choice Selectors"☆43May 20, 2025Updated last year
- This repository includes the code implementation of the paper Improving Pacing in Long-Form Story Planning by Yichen Wang, Kevin Yang, Xi…☆18Nov 19, 2024Updated last year
- I-SHEEP: Iterative Self-enHancEmEnt Paradigm of LLMs through Self-Instruct and Self-Assessment☆17Jan 16, 2025Updated last year
- Resolving Knowledge Conflicts in Large Language Models, COLM 2024☆18Oct 7, 2025Updated 11 months ago
- Conversational Recommender System Evaluation via Simulation☆23Aug 24, 2026Updated last month
- [NeurIPS 2024] A Novel Rank-Based Metric for Evaluating Large Language Models☆58May 28, 2025Updated last year
- Clean, extensible implementation of MACAW [ICML 2021]☆12Dec 7, 2021Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Cue-CoT: Chain-of-thought Prompting for Responding to In-depth Dialogue Questions with LLMs [EMNLP 2023 Findings]☆24Nov 18, 2023Updated 2 years ago
- mPLUG-HalOwl: Multimodal Hallucination Evaluation and Mitigating☆99Jan 29, 2024Updated 2 years ago
- A Japanese G2P tool based on pyopenjtalk☆25Aug 6, 2022Updated 4 years ago
- Repository for the ACL 2024 conference website☆18Feb 3, 2025Updated last year
- ☆10Oct 22, 2024Updated last year
- Repository of paper "How Likely Do LLMs with CoT Mimic Human Reasoning?"☆23Feb 19, 2025Updated last year
- NeurIPS 2025☆16Feb 4, 2026Updated 7 months ago