[NeurIPS 2023] Large Language Models Are Semi-Parametric Reinforcement Learning Agents
☆40May 2, 2024Updated 2 years ago
Alternatives and similar repositories for Rememberer
Users that are interested in Rememberer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Universal Platform for Training and Evaluation of Mobile Interaction☆63Sep 24, 2025Updated 10 months ago
- ☆15Mar 26, 2024Updated 2 years ago
- ☆234Dec 20, 2024Updated last year
- ☆12Jul 4, 2024Updated 2 years ago
- ☆36Mar 14, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Text-to-Drive: Diverse Driving Behaviors Synthesis via Large Language Models☆11Mar 17, 2024Updated 2 years ago
- Official implementation of our ICSE 2023 paper on Automatic Code Generation.☆27Nov 8, 2023Updated 2 years ago
- [ECCV 2024] The first zero-shot setting for spatio-temporal video grounding.☆11Jul 16, 2024Updated 2 years ago
- Langchain Agent finetuning using 7B - LLAMA 2 , on hotpotQA (Retroformer framework)☆16Sep 5, 2023Updated 2 years ago
- [ICML'24] TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks☆33Sep 20, 2024Updated last year
- ☆14Dec 25, 2024Updated last year
- ☆13Aug 26, 2024Updated last year
- ☆70Dec 15, 2024Updated last year
- SCoRe: Training Language Models to Self-Correct via Reinforcement Learning☆16May 14, 2026Updated 2 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [ACL 2026] A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models☆23Jul 10, 2026Updated last month
- Implementation of "PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier"☆17Jun 27, 2025Updated last year
- Code for the paper "Reactive Exploration to Cope with Non-Stationarity in Lifelong Reinforcement Learning"☆16Jul 4, 2022Updated 4 years ago
- Model-based Hindsight Experience Replay☆10Jun 8, 2022Updated 4 years ago
- [ACL 2025] Official code for ''Learning to Reason from Feedback at Test-Time''.☆13May 16, 2025Updated last year
- A Recipe for Building LLM Reasoners to Solve Complex Instructions☆32Oct 9, 2025Updated 10 months ago
- ☆17Oct 25, 2023Updated 2 years ago
- Official PyTorch implementation of "A Rotated Hyperbolic Wrapped Normal Distribution for Hierarchical Representation Learning"☆28Oct 12, 2022Updated 3 years ago
- Implementation of TWOSOME☆82Jan 11, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- It is about how to load and aggregate pretrained word embeddings in pytorch, e.g., ELMo\BERT\XLNET.☆12Mar 2, 2020Updated 6 years ago
- A full fledged mistral+wandb☆13Aug 16, 2024Updated last year
- Convert CVXPY expressions to PyTorch expressions☆18Jul 8, 2025Updated last year
- A variant of Varibad that is robust to difficult tasks☆11Aug 30, 2023Updated 2 years ago
- Implementation of Mean Field Multi-Agent Reinforcement Learning in Pytorch☆21Apr 27, 2024Updated 2 years ago
- CoRL2024 | Hint-AD: Holistically Aligned Interpretability for End-to-End Autonomous Driving☆74Oct 30, 2024Updated last year
- ☆20Aug 15, 2023Updated 2 years ago
- The official implementation of NOSA☆20Jun 11, 2026Updated 2 months ago
- Official implementation of the DECKARD Agent from the paper "Do Embodied Agents Dream of Pixelated Sheep?"☆94May 23, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- implementation of dualformer☆25Mar 1, 2025Updated last year
- ☆16Jan 12, 2026Updated 7 months ago
- [ACL 2025] "World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning." https://arxiv.org/abs/2503.1…☆18Jul 22, 2025Updated last year
- Data and codes for EMNLP 2022 paper "CDConv: A Benchmark for Contradiction Detection in Chinese Conversations"☆13May 8, 2023Updated 3 years ago
- ☆17Mar 1, 2026Updated 5 months ago
- The repo for using the model https://huggingface.co/thu-coai/Attacker-v0.1☆13Apr 23, 2025Updated last year
- [NeurIPS 2025] Bag of Tricks for Inference-time Computation of LLM Reasoning☆16Sep 20, 2025Updated 10 months ago