The official repository for Trust-Region Adaptive Policy Optimization (TRAPO) – a novel hybrid framework designed to enhance large language models' reasoning abilities by interleaving SFT and RL within each training instance.
☆16Mar 2, 2026Updated 6 months ago
Alternatives and similar repositories for TRAPO
Users that are interested in TRAPO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026] G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance☆16Updated this week
- ☆11Dec 26, 2025Updated 8 months ago
- ☆15Feb 10, 2026Updated 7 months ago
- Code for "Accelerating Transformer Pre-training with 2:4 Sparsity"☆28Dec 8, 2024Updated last year
- PyTorch implementation of CARE☆16Oct 6, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- 人工智能概论大作业---基于ANN与KNN的图像分类☆11Dec 11, 2020Updated 5 years ago
- UFT: Unifying Supervised and Reinforcement Fine-Tuning☆33Jun 30, 2025Updated last year
- ☆14Aug 7, 2023Updated 3 years ago
- Official implementation for "Pruning Large Language Models with Semi-Structural Adaptive Sparse Training" (AAAI 2025)☆20Jul 1, 2025Updated last year
- [COLING2022] A Multi-turn Machine Reading Comprehension Framework with Rethink Mechanism for Emotion-Cause Pair Extraction☆18Oct 13, 2022Updated 3 years ago
- Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model☆13Dec 29, 2024Updated last year
- Official repository for the Findings of ACL 2023 paper "AugESC: Dialogue Augmentation with Large Language Models for Emotional Support Co…☆22May 16, 2023Updated 3 years ago
- 基于深度学习的文本分类,实现基于CNN和RNN的文本分类☆11Sep 20, 2021Updated 4 years ago
- Code for ICML 2021 submission☆35Mar 24, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Implicit Emotion Cause Extraction☆16Aug 2, 2022Updated 4 years ago
- [NeurIPS 2024 poster] Cross-model Control: Improving Multiple Large Language Models in One-time Training☆15Oct 25, 2024Updated last year
- ☆24Sep 12, 2023Updated 3 years ago
- [AAAI2024] Debiasing Multimodal Sarcasm Detection with Contrastive Learning☆17Jan 5, 2024Updated 2 years ago
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆26Dec 30, 2025Updated 8 months ago
- 利用Transformer模型实现的机器翻译☆12Dec 6, 2020Updated 5 years ago
- [COLM 2024] Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation☆15Jul 15, 2024Updated 2 years ago
- Official repository for LLaVA-Reward (ICCV 2025): Multimodal LLMs as Customized Reward Models for Text-to-Image Generation☆26Jul 30, 2025Updated last year
- The official implementation of the paper "Self-Updatable Large Language Models by Integrating Context into Model Parameters"☆16May 18, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 蚂蚁金融自然语言处理竞赛。☆10Sep 3, 2018Updated 8 years ago
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space☆17Apr 16, 2026Updated 5 months ago
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆66Updated this week
- 3DSlicer plugin for inpainting lung nodules in 3D chest CT data.☆11Dec 2, 2024Updated last year
- This repository collects the list of accepted paper from (currently only deep learning) top conferences. All lists are crawled by python …☆12Feb 11, 2023Updated 3 years ago
- End-to-End Neural Event Coreference Resolution☆11Jun 18, 2023Updated 3 years ago
- pytorch版的命名实体识别,LSTM和LSTM_CRF☆25Aug 16, 2019Updated 7 years ago
- [NeurIPS'24] Official code for *🎯DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving*☆121Dec 10, 2024Updated last year
- The codes of our paper "ObjectAdd: Adding Objects into Image via a Training-Free Diffusion Modification Fashion"☆14Jun 29, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Complexity Based Prompting for Multi-Step Reasoning☆17Mar 10, 2023Updated 3 years ago
- Probing task; contextual embeddings -> textual definitions (EMNLP19)☆12Apr 22, 2021Updated 5 years ago
- Official code for paper "Revisiting Model Interpolation for Efficient Reasoning"☆17Jul 14, 2026Updated 2 months ago
- ☆47Jun 2, 2026Updated 3 months ago
- [ICLR 2025 Oral] Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition☆17Nov 25, 2024Updated last year
- UnitEval is a benchmarking and evaluation tools for AutoDev Coder.☆14Jan 2, 2024Updated 2 years ago
- The resources for the paper "User Modeling with Click Preference and Reading Satisfaction for News Recommendation"☆11Jan 17, 2021Updated 5 years ago