The official repository for Trust-Region Adaptive Policy Optimization (TRAPO) – a novel hybrid framework designed to enhance large language models' reasoning abilities by interleaving SFT and RL within each training instance.
☆16Mar 2, 2026Updated 7 months ago
Alternatives and similar repositories for TRAPO
Users that are interested in TRAPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026] G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance☆16Sep 15, 2026Updated 3 weeks ago
- ☆15Feb 10, 2026Updated 7 months ago
- PyTorch implementation of CARE☆16Oct 6, 2023Updated 3 years ago
- 人工智能概论大作业---基于ANN与KNN的图像分类☆11Dec 11, 2020Updated 5 years ago
- UFT: Unifying Supervised and Reinforcement Fine-Tuning☆34Jun 30, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆14Aug 7, 2023Updated 3 years ago
- ☆14Jan 10, 2024Updated 2 years ago
- Official implementation for "Pruning Large Language Models with Semi-Structural Adaptive Sparse Training" (AAAI 2025)☆20Jul 1, 2025Updated last year
- ☆12Nov 17, 2024Updated last year
- UECA-Prompt: Universal Prompt for Emotion Cause Analysis(COLING 2022)☆16Jun 6, 2023Updated 3 years ago
- Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model☆13Dec 29, 2024Updated last year
- Fine tuning of the Retrieval-Augmented Generation (RAG) with a custom knowledge source.☆13Feb 10, 2021Updated 5 years ago
- Web one-click mode full process platform, including train data upload, fine-tuning, model merge, model deploy, gpu monitor etc., no need …☆19Nov 28, 2023Updated 2 years ago
- 基于深度学习的文本分类,实现基于CNN和RNN的文本分类☆11Sep 20, 2021Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [NeurIPS 2024 poster] Cross-model Control: Improving Multiple Large Language Models in One-time Training☆15Oct 25, 2024Updated last year
- Materials for "Multi-property Steering of Large Language Models with Dynamic Activation Composition"☆14Nov 22, 2024Updated last year
- ☆24Sep 12, 2023Updated 3 years ago
- [AAAI2024] Debiasing Multimodal Sarcasm Detection with Contrastive Learning☆17Jan 5, 2024Updated 2 years ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆26Dec 30, 2025Updated 9 months ago
- ☆26Apr 25, 2024Updated 2 years ago
- A mapping tool for vehicle routing problem datasets and solutions☆11Feb 12, 2020Updated 6 years ago
- [COLM 2024] Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation☆15Jul 15, 2024Updated 2 years ago
- Code and Model for NeurIPS 2024 Spotlight Paper "Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training…☆44Oct 16, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official repository for LLaVA-Reward (ICCV 2025): Multimodal LLMs as Customized Reward Models for Text-to-Image Generation☆27Jul 30, 2025Updated last year
- The official implementation of the paper "Self-Updatable Large Language Models by Integrating Context into Model Parameters"☆16May 18, 2025Updated last year
- 蚂蚁金融自然语言处理竞赛。☆10Sep 3, 2018Updated 8 years ago
- The project is focused on creating simple and TorchScript compilable inference interface for the original pretrained models to free them …☆17Dec 24, 2023Updated 2 years ago
- [NeurIPS 2026 Pre-to-Post ] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space☆18Apr 16, 2026Updated 5 months ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆74Updated this week
- 3DSlicer plugin for inpainting lung nodules in 3D chest CT data.☆11Dec 2, 2024Updated last year
- 复旦大学nlp实验室入门小实验nlp-beginner☆27Jan 22, 2022Updated 4 years ago
- [NeurIPS'24] Official code for *🎯DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving*☆120Dec 10, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The codes of our paper "ObjectAdd: Adding Objects into Image via a Training-Free Diffusion Modification Fashion"☆14Jun 29, 2025Updated last year
- Complexity Based Prompting for Multi-Step Reasoning☆17Mar 10, 2023Updated 3 years ago
- Probing task; contextual embeddings -> textual definitions (EMNLP19)☆12Apr 22, 2021Updated 5 years ago
- Official code for paper "Revisiting Model Interpolation for Efficient Reasoning"☆17Jul 14, 2026Updated 2 months ago
- ☆47Jun 2, 2026Updated 4 months ago
- Consolidated Ground Truth (CGT) for Weaknesses of Ethereum Smart Contracts☆27Oct 12, 2024Updated last year
- [ICLR 2025 Oral] Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition☆17Nov 25, 2024Updated last year