Online Preference Alignment for Language Models via Count-based Exploration
☆21Jan 14, 2025Updated last year
Alternatives and similar repositories for COPO
Users that are interested in COPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NIPS'25] Official Implementation of "Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective" in PyTorch.☆18Nov 11, 2025Updated 10 months ago
- [NeurIPS 2025] Official Implementation of "HumanoidGen: Data Generation for Bimanual Dexterous Manipulation via LLM Reasoning"☆95Nov 6, 2025Updated 10 months ago
- [NeurIPS' 24] The PyTorch implementation of our paper: "Kaleidoscope: Learnable Masks for Heterogeneous Multi-agent Reinforcement Learnin…☆22Oct 10, 2024Updated last year
- Code space for L4DC paper "State-wise Safe Reinforcement Learning With Pixel Observations"☆11Apr 5, 2024Updated 2 years ago
- [IROS2024] STAIR: Semantic-Targeted Active Implicit Reconstruction☆16Aug 3, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF☆25Oct 8, 2024Updated last year
- Official implementation of paper: LiNo: Advancing Recursive Residual Decomposition of Linear and Nonlinear Patterns for Robust Time Serie…☆18Dec 19, 2025Updated 9 months ago
- ☆18Jul 31, 2026Updated last month
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 5 months ago
- ☆14May 13, 2025Updated last year
- Official repository for ICLR 2025 paper "Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs"☆20Mar 18, 2025Updated last year
- G-HER algorithm☆18May 24, 2019Updated 7 years ago
- Code and datasets of AAAI'22 paper "TIGGER: Scalable Generative Modelling for Temporal Interaction Graphs" . To be appear in AAAI-2022☆14Nov 17, 2023Updated 2 years ago
- ☆38Aug 24, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [TGRS 2024] CutMix-CD: Advancing Semi-Supervised Change Detection via Mixed Sample Consistency☆22Nov 30, 2025Updated 9 months ago
- [AAAI-25] Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning.☆37May 29, 2025Updated last year
- ☆13Jun 4, 2025Updated last year
- Official Code for "Adversarial Locomotion and Motion Imitation for Humanoid Policy Learning"☆156May 16, 2025Updated last year
- ☆16Jun 12, 2024Updated 2 years ago
- This is the official git for team Kanaloa☆11Dec 9, 2022Updated 3 years ago
- ☆20Nov 3, 2024Updated last year
- [CVPR'24] Code for Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models☆18Jul 22, 2024Updated 2 years ago
- ICML 2024 - Self-Driven Entropy Aggregation for Byzantine-Robust Heterogeneous Federated Learning☆11Jul 16, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Implementation of "Towards Understanding Mixture of Experts in Deep Learning", NeurIPS 2022☆10Jan 6, 2023Updated 3 years ago
- The official implementation of "Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation"☆25Feb 8, 2026Updated 7 months ago
- The repository is for Reinforcement-Learning Uncertainty research, in which we investigate various uncertain factors in RL.☆23Jun 16, 2023Updated 3 years ago
- This is the source code of FUSION, a safety-aware causal representation for generalizable driving agents.☆29Oct 23, 2024Updated last year
- ☆24Dec 30, 2024Updated last year
- LLM-Empowered State Representation for Reinforcement Learning (ICML2024 Accepted paper)☆42Jun 14, 2024Updated 2 years ago
- Code for Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning, AAAI 2025☆15Dec 19, 2024Updated last year
- Feasibility Consistent Representation Learning for Safe Reinforcement Learning (ICML 2024). Current SOTA model-free safe RL algorithm on …☆17Jul 12, 2024Updated 2 years ago
- Code accompanying the paper "Off-Policy Primal-Dual Safe Reinforcement Learning"☆23Mar 29, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆24Jun 16, 2025Updated last year
- Comprehensive Assessment of Trustworthiness in Multimodal Foundation Models☆29Mar 15, 2025Updated last year
- A PyTorch implementation of [VCT](https://github.com/google-research/google-research/tree/master/vct)☆10Nov 25, 2022Updated 3 years ago
- ☆16Oct 26, 2025Updated 11 months ago
- Official PyTorch Implementation of Paper -- "MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains"☆317Nov 11, 2025Updated 10 months ago
- ☆28Aug 25, 2026Updated last month
- Implementation of Negative-aware Finetuning (NFT) algorithm for "Bridging Supervised Learning and Reinforcement Learning in Math Reasonin…☆92Sep 8, 2025Updated last year