☆14Dec 16, 2024Updated last year
Alternatives and similar repositories for MRPO
Users that are interested in MRPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Model-based Episodic Control & Complementary Learning Systems☆17Dec 13, 2021Updated 4 years ago
- Episodic Policy Gradient Training☆17Mar 1, 2022Updated 4 years ago
- Source code for Stable Hadamard Memory☆24May 6, 2025Updated last year
- Self-attentive Associative Memory & SAM-based Two-Memory Model☆61May 4, 2022Updated 4 years ago
- Source code for the paper: Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention☆18Aug 14, 2026Updated 2 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official implementation of Bidirectional Diffusion Bridge Models☆27May 26, 2025Updated last year
- ☆17Jun 9, 2025Updated last year
- Implementation of "Decoding-time Realignment of Language Models", ICML 2024.☆21Jun 17, 2024Updated 2 years ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆26Dec 30, 2025Updated 8 months ago
- Benchmarking Generalization to New Tasks from Natural Language Instructions☆28Jul 2, 2021Updated 5 years ago
- [WACV 2024] Domain Generalisation via Risk Distribution Matching☆23Sep 19, 2024Updated last year
- Code for ICML 2024 paper☆34Sep 18, 2025Updated 11 months ago
- The task aims at extracting required fields in receipts captured by mobile devices☆36Nov 4, 2022Updated 3 years ago
- ☆33May 9, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Code for the paper "Spectral Editing of Activations for Large Language Model Alignments"☆31Dec 20, 2024Updated last year
- ☆47Apr 24, 2022Updated 4 years ago
- Learning to Retrieve by Trying - Source code for Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval☆52Oct 31, 2024Updated last year
- The repository contains resources of our paper published at AAAI 2020 ``A Dataset for Low-Resource Stylized Sequence-to-Sequence Generati…☆33Jan 21, 2020Updated 6 years ago
- ☆46Apr 10, 2023Updated 3 years ago
- The official code release for Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization☆39Mar 9, 2025Updated last year
- Code for our paper: "GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models"☆57Apr 23, 2023Updated 3 years ago
- Active Example Selection for In-Context Learning (EMNLP'22)☆48Jul 22, 2024Updated 2 years ago
- ☆37Apr 8, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆47Jul 25, 2024Updated 2 years ago
- Official code for the paper "Attention as a Hypernetwork"☆59Feb 24, 2026Updated 6 months ago
- Directional Preference Alignment☆62Sep 23, 2024Updated last year
- Rewarded soups official implementation☆66Sep 27, 2023Updated 2 years ago
- Released code for our ICLR23 paper.☆66Mar 23, 2023Updated 3 years ago
- Algebraic value editing in pretrained language models☆71Nov 1, 2023Updated 2 years ago
- Codebase for TIME benchmark☆60Updated this week
- Code of ICLR paper: https://openreview.net/forum?id=-cqvvvb-NkI☆95Feb 22, 2023Updated 3 years ago
- Time-R1: Framework and resources for endowing LLMs with comprehensive temporal reasoning (understanding, prediction, creative generation)…☆71Jun 11, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆75Apr 13, 2025Updated last year
- We view Large Language Models as stochastic language layers in a network, where the learnable parameters are the natural language prompts…☆95Jul 25, 2024Updated 2 years ago
- Official code for paper "TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning"☆68Oct 22, 2025Updated 10 months ago
- Code for the ICML 2024 paper "Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment"☆78Jun 10, 2025Updated last year
- COVID-19 Named Entity Recognition for Vietnamese (NAACL 2021)☆73Jul 22, 2024Updated 2 years ago
- [CVPR 2025] h-Edit: Effective and Flexible Diffusion-Based Editing via Doob’s h-Transform☆79Jun 11, 2025Updated last year
- ☆102Aug 24, 2022Updated 4 years ago