☆313Nov 17, 2023Updated 2 years ago
Alternatives and similar repositories for reversal_curse
Users that are interested in reversal_curse are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for paper 'Are We Falling in a Middle-Intelligence Trap? An Analysis and Mitigation of the Reversal Curse'☆14Aug 2, 2024Updated last year
- Measuring the situational awareness of language models☆41Feb 12, 2024Updated 2 years ago
- A text-based game where language models learn to lie and to detect lies.☆12Oct 4, 2023Updated 2 years ago
- ☆13Apr 24, 2024Updated 2 years ago
- Teaching Models to Express Their Uncertainty in Words☆39May 26, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Forcing Diffuse Distributions out of Language Models☆18Sep 10, 2024Updated last year
- ☆16Mar 22, 2025Updated last year
- The Effect of Sampling Temperature on Problem Solving in Large Language Models☆25Nov 25, 2024Updated last year
- Code for EMNLP'24 paper - On Diversified Preferences of Large Language Model Alignment☆16Aug 6, 2024Updated last year
- ☆43Sep 3, 2024Updated last year
- Reversal Curse Experiment☆15Sep 24, 2023Updated 2 years ago
- Code for the paper "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples"☆325Dec 20, 2023Updated 2 years ago
- Scaling Data-Constrained Language Models☆345Jun 28, 2025Updated last year
- Scripts for generating synthetic finetuning data for reducing sycophancy.☆125Aug 16, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Functional Benchmarks and the Reasoning Gap☆90Updated this week
- Code for the ICLR 2024 paper "How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions"☆74Jun 19, 2024Updated 2 years ago
- Models, data, and codes for the paper: MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models☆24Sep 26, 2024Updated last year
- ☆17Dec 21, 2023Updated 2 years ago
- Parkar and Kim et al.'s paper on Can LLMs Select Important Instructions to Annotate?"☆13Jul 4, 2024Updated 2 years ago
- Salesforce open-source LLMs with 8k sequence length.☆727Jun 2, 2026Updated last month
- Editing Models with Task Arithmetic☆548Jan 11, 2024Updated 2 years ago
- Source code for the paper "Positional Attention: Expressivity and Learnability of Algorithmic Computation"☆14May 26, 2025Updated last year
- Adversarial Attack for Pre-trained Code Models☆10Jul 19, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ICML 2024: Improving Factuality and Reasoning in Language Models through Multiagent Debate☆544Apr 24, 2025Updated last year
- Code repository for the paper - "AdANNS: A Framework for Adaptive Semantic Search"☆69Oct 10, 2023Updated 2 years ago
- ☆30Jun 19, 2023Updated 3 years ago
- CoPur: Certifiably Robust Collaborative Inference via Feature Purification (NeurIPS 2022)☆11Dec 7, 2022Updated 3 years ago
- ☆58Jun 15, 2023Updated 3 years ago
- Minimal implementation of multiple PEFT methods for LLaMA fine-tuning☆13May 7, 2023Updated 3 years ago
- Continual Memorization of Factoids in Large Language Models☆12Nov 20, 2024Updated last year
- ☆47Feb 8, 2024Updated 2 years ago
- Few-shot Learning with Auxiliary Data☆31Dec 8, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Locating and editing factual associations in GPT (NeurIPS 2022)☆770Apr 20, 2024Updated 2 years ago
- ☆37May 28, 2023Updated 3 years ago
- Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"☆1,853Jun 17, 2025Updated last year
- This is the repository for our paper: Untying the Reversal Curse via Bidirectional Language Model Editing☆11May 25, 2025Updated last year
- Official github repo for AutoDetect, an automated weakness detection framework for LLMs.☆47Jun 25, 2024Updated 2 years ago
- Dataset Interfaces: Diagnosing Model Failures Using Controllable Counterfactual Generation☆45Feb 27, 2023Updated 3 years ago
- ☆164Nov 23, 2024Updated last year