☆22Jun 4, 2025Updated last year
Alternatives and similar repositories for RTO
Users that are interested in RTO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is an official implementation of the paper ``Building Math Agents with Multi-Turn Iterative Preference Learning'' with multi-turn DP…☆32Dec 5, 2024Updated last year
- Code the ICML 2024 paper: "Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models"☆12Jun 25, 2024Updated 2 years ago
- ☆22Updated this week
- ☆16Jul 29, 2025Updated last year
- Ancestral Gumbel-Top-k Sampling☆24Apr 11, 2020Updated 6 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The official implementation of Self-Exploring Language Models (SELM)☆63Jun 4, 2024Updated 2 years ago
- This is the code used for the paper "PMGT-VR: A decentralized proximal-gradient algorithmic framework with variance reduction", prepint.☆15Jul 2, 2022Updated 4 years ago
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression☆41May 30, 2025Updated last year
- [VLDB'2025] LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data☆20Nov 3, 2025Updated 9 months ago
- ☆14Jun 24, 2024Updated 2 years ago
- Aligning Agentic World Models via Knowledgeable Experience Learning☆39May 15, 2026Updated 2 months ago
- Source code for NeurIPS 2020 paper "Node Classification on Graphs with Few-Shot Novel Labels via Meta Transformed Network Embedding"☆10Nov 17, 2020Updated 5 years ago
- PyTorch code for "Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training"☆39Mar 4, 2024Updated 2 years ago
- Watermarking LLM papers up-to-date☆12Dec 17, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Directional Preference Alignment☆62Sep 23, 2024Updated last year
- DOMAINEVAL is an auto-constructed benchmark for multi-domain code generation that consists of 2k+ subjects (i.e., description, reference …☆13Dec 12, 2024Updated last year
- Source for the sample efficient tabular RL submission to the 2019 NIPS workshop on Biological and Artificial RL☆25Apr 14, 2022Updated 4 years ago
- Repository for Skill Set Optimization☆14Jul 26, 2024Updated 2 years ago
- ☆41Feb 3, 2026Updated 6 months ago
- Multimodal Federated Learning on IoT Data☆11Dec 17, 2023Updated 2 years ago
- Explore what LLMs are really leanring over SFT☆28Mar 30, 2024Updated 2 years ago
- ☆42Sep 20, 2022Updated 3 years ago
- ☆13Sep 12, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [WACV 2024] Enhancing Multimodal Compositional Reasoning of Visual Language Models with Generative Negative Mining, WACV 2024☆13Jan 3, 2024Updated 2 years ago
- ☆15Nov 26, 2024Updated last year
- Preference Learning for LLaVA☆60Nov 9, 2024Updated last year
- Accompanying code for the paper "Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving"☆17May 7, 2026Updated 3 months ago
- Official repository of the paper: Who Wrote this Code? Watermarking for Code Generation (ACL 2024)☆40May 28, 2024Updated 2 years ago
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization☆19Mar 7, 2025Updated last year
- A RLHF Infrastructure for Vision-Language Models☆201Nov 15, 2024Updated last year
- All you need to get started with the LM Playpen Environment for Learning in Interaction.☆17Jun 22, 2026Updated last month
- IPO: Interpretable Prompt Optimization for Vision-Language Models(NeurIPS 2024)☆15Jun 12, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This repository hosts the source code for the paper "ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Mo…☆16Dec 16, 2025Updated 7 months ago
- 在线课程网站☆11May 5, 2018Updated 8 years ago
- [ICML 2025] Teaching Language Models to Critique via Reinforcement Learning☆127May 6, 2025Updated last year
- Code for NeurIPS 2024 paper "Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs"☆47Feb 20, 2025Updated last year
- [ACL 2024, Main Conference] CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Fol…☆15Aug 7, 2024Updated 2 years ago
- ☆13Nov 5, 2024Updated last year
- ☆17Nov 8, 2024Updated last year