[EMNLP 2025] Dataset and Code of "PersonaGym: Evaluating Persona Agents and LLMs"
☆43Aug 21, 2025Updated 11 months ago
Alternatives and similar repositories for PersonaGym
Users that are interested in PersonaGym are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2024] Dataset and Code of "ImplicitAVE: An Open-Source Dataset and Multimodal LLMs Benchmark for Implicit Attribute Value Extraction…☆17Jun 10, 2024Updated 2 years ago
- code for paper "Towards Unbiased Training in Federated Open-world Semi-supervised Learning"☆18Aug 15, 2023Updated 3 years ago
- LogicIF: Towards Complex Logic Instruction Following☆18Jul 12, 2026Updated last month
- Android app for cloning MIFARE Ultralight UIDs☆11Apr 10, 2015Updated 11 years ago
- ☆12Dec 14, 2022Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [AAAI'25] CharacterBench: Benchmarking Character Customization of Large Language Models☆23Aug 1, 2025Updated last year
- ☆13Jun 12, 2024Updated 2 years ago
- Depth video data-enabled predictions of longitudinal dairy cow body weight using thresholding and Mask R-CNN algorithms☆11Sep 22, 2024Updated last year
- Kim, J., Evans, J., & Schein, A. (2025). Linear Representations of Political Perspective Emerge in Large Language Models. ICLR.☆25Mar 27, 2025Updated last year
- Vanity address generator for Ethereum for Windows☆16Mar 29, 2024Updated 2 years ago
- Official code repository for the paper "Rethinking Model Prototyping through the MedMNIST+ Dataset Collection" @ Scientific Reports☆13Mar 5, 2025Updated last year
- YOLOv7 Object Cropping Using OpenCV☆19Feb 15, 2025Updated last year
- AbstainQA, ACL 2024☆29Feb 4, 2026Updated 6 months ago
- ☆130Nov 7, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The repository includes evidence that a published paper in TKDE shares surprisingly high similarity to our paper.☆32Dec 18, 2022Updated 3 years ago
- A deep semi-supervised method (UATS) for medical segmentation☆14Jul 18, 2024Updated 2 years ago
- ☆10Sep 17, 2022Updated 3 years ago
- Codebase for Paper Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs☆23Apr 24, 2025Updated last year
- This is a read-only mirror of the CRAN R package repository. AER — Applied Econometrics with R☆11Jul 11, 2026Updated last month
- MetricEval: A framework that conceptualizes and operationalizes four main components of metric evaluation, in terms of reliability and va…☆12Nov 6, 2023Updated 2 years ago
- Debug DeepSpeed-Chat step by step in IDE (在IDE里一步一步调试DeepSpeed-Chat)☆10Apr 17, 2023Updated 3 years ago
- ☆65Jun 2, 2026Updated 2 months ago
- ☆15Feb 27, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- code for ACL2024-main: BatchEval: Towards Human-like Text Evaluation☆19May 20, 2024Updated 2 years ago
- Code associated with Tuning Language Models by Proxy (Liu et al., 2024)☆134Mar 30, 2024Updated 2 years ago
- protein backbone refinement☆15Sep 12, 2024Updated last year
- ☆118Oct 11, 2024Updated last year
- ☆17Oct 22, 2024Updated last year
- ☆12Oct 20, 2020Updated 5 years ago
- Replication code for "The Structure of Toxic Conversations on Twitter" (WWW'21)☆10May 25, 2021Updated 5 years ago
- ☆11Oct 16, 2023Updated 2 years ago
- GOPHI: an AMR-to-English Verbalizer☆12Feb 5, 2020Updated 6 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆30Jun 29, 2026Updated last month
- Code associated with our paper "Estimating Risk and Uncertainty in Reinforcement Learning"☆11Oct 3, 2023Updated 2 years ago
- Multi-Level Adversarial for Cross-lingual Name Tagging☆12Jun 18, 2020Updated 6 years ago
- Official code for the paper: InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews (previo…☆101May 27, 2025Updated last year
- ☆13Jun 4, 2024Updated 2 years ago
- ☆18Jul 6, 2023Updated 3 years ago
- Dataset for Conversation Semantic Role Labeling☆11Aug 26, 2021Updated 4 years ago