☆70Apr 28, 2026Updated 3 months ago
Alternatives and similar repositories for data_synth_and_rl
Users that are interested in data_synth_and_rl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- a toolkit on knowledge distillation for large language models☆443Mar 10, 2026Updated 4 months ago
- COS-PLAY: Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Game Play☆29Jul 11, 2026Updated 3 weeks ago
- The official paper for EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL.☆86Jun 5, 2026Updated last month
- Fused KL divergence from hidden states for knowledge distillation☆21Apr 28, 2026Updated 3 months ago
- ☆11Jul 21, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- To Trust Or Not To Trust Your Vision-Language Model's Prediction☆15May 30, 2025Updated last year
- verl: Volcano Engine Reinforcement Learning for LLMs☆22Nov 6, 2025Updated 8 months ago
- [AAAI 2026] LungNoduleAgent: A Collaborative Multi-Agent System for Precision Diagnosis of Lung Nodules☆36Jul 28, 2026Updated last week
- ☆23Dec 17, 2024Updated last year
- ☆23Oct 10, 2025Updated 9 months ago
- Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement☆18Jan 16, 2026Updated 6 months ago
- ToolOrchestra is an end-to-end RL training framework for orchestrating tools and agentic workflows.☆753Mar 25, 2026Updated 4 months ago
- 个人毕业设计,一个分别使用U-net与CNN进行肺结节分割与分类的项目。☆13Jun 17, 2025Updated last year
- ☆22May 14, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆25Feb 12, 2026Updated 5 months ago
- CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search☆15Apr 28, 2026Updated 3 months ago
- [CVPR 2026] An official implementation of "Think Visually, Reason Textually: Vision-Language Synergy in ARC"☆46Nov 26, 2025Updated 8 months ago
- ☆17Mar 26, 2026Updated 4 months ago
- Project for EE609 Convex Optimization☆12May 6, 2021Updated 5 years ago
- ☆14Oct 17, 2024Updated last year
- ☆21Apr 21, 2026Updated 3 months ago
- A research framework for evaluating proactive AI assistants through active user simulation☆37May 23, 2026Updated 2 months ago
- [NeurIPS 2025] The implementation of paper "On Reasoning Strength Planning in Large Reasoning Models"☆31Jul 6, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Training terminal-agents☆260Jul 22, 2026Updated last week
- ☆29Jan 31, 2026Updated 6 months ago
- ☆11May 17, 2024Updated 2 years ago
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"☆17Jun 30, 2025Updated last year
- A replication of Google's VideoPoet model☆12Feb 18, 2024Updated 2 years ago
- Code for SafeMERGE (ICLR 2025).☆15Apr 1, 2025Updated last year
- A toolkit for automated alignment research.☆15Jul 3, 2026Updated last month
- AgentIR is a retriever specialized for Deep Research agents.☆62Apr 16, 2026Updated 3 months ago
- A Workbench for Autograding Retrieve/Generate Systems☆15Jun 30, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official code for the paper: "Multi-User Large Language Model Agents"☆27May 11, 2026Updated 2 months ago
- [ICML'26] Official implementation of paper "Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models"☆76Jul 17, 2026Updated 2 weeks ago
- ☆17Mar 1, 2026Updated 5 months ago
- Official code for "SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization"☆358Jul 22, 2026Updated last week
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning☆423May 28, 2026Updated 2 months ago
- ☆12Sep 1, 2023Updated 2 years ago
- LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning☆38Apr 4, 2024Updated 2 years ago