Repo for paper "Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability"
☆109Apr 23, 2026Updated 3 months ago
Alternatives and similar repositories for rethink_sft_generalization
Users that are interested in rethink_sft_generalization are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Repo of Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents☆91Jun 2, 2026Updated 2 months ago
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis☆35Jul 10, 2026Updated last month
- A Diagnostic Guardrail Framework for AI Agent Safety and Security☆682Jun 8, 2026Updated 2 months ago
- All-in-One Safety Evaluation Framwork☆53Jul 15, 2026Updated 3 weeks ago
- Diagnostic Framework for LLMs and MLLMs☆39Mar 2, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆68Jul 14, 2025Updated last year
- 🧨 TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?☆81Nov 27, 2025Updated 8 months ago
- [ACL 2025] Data and Code for Paper VLSBench: Unveiling Visual Leakage in Multimodal Safety☆62Jul 21, 2025Updated last year
- Jointly Optimizing Large Language Models for Reasoning and Self-Refinement☆15Apr 22, 2026Updated 3 months ago
- This repo is the official implementation of “Are Your Agents Upward Deceivers?”. The paper is accepted by ICML 2026.☆24Dec 15, 2025Updated 7 months ago
- Official Repository of "Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Ste…☆28Mar 9, 2026Updated 5 months ago
- [ACL 2026 Main] Official Repo for Paper "Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Ali…☆17Jul 1, 2026Updated last month
- ☆45Mar 30, 2026Updated 4 months ago
- ☆11Oct 25, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Trace origins, shared sources, and contamination risk☆27May 27, 2026Updated 2 months ago
- On Policy Distillation Build on top of Verl☆94May 25, 2026Updated 2 months ago
- The official code of FineRMoE.☆21Mar 17, 2026Updated 4 months ago
- [NeurIPS 2025] Official repository of RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents☆123Dec 2, 2025Updated 8 months ago
- [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repositor…☆24Apr 24, 2026Updated 3 months ago
- [AAAI 2026] Data and Code for Paper IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks☆49Nov 24, 2025Updated 8 months ago
- ☆26Feb 20, 2026Updated 5 months ago
- ReplayCode — first open-source rebuild of Claude Code that actually runs. Built from decompiled source with Node.js/esbuild☆20Apr 1, 2026Updated 4 months ago
- Socratic-Zero is a fully autonomous framework that generates high-quality training data for mathematical reasoning☆37Oct 26, 2025Updated 9 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT Data☆35May 1, 2026Updated 3 months ago
- Official implementation of Selective Entropy Regularization (SIREN), proposed by paper 'Rethinking Entropy Regularization in Large Reason…☆32Dec 10, 2025Updated 8 months ago
- The repository of the paper "REEF: Representation Encoding Fingerprints for Large Language Models," aims to protect the IP of open-source…☆79Jan 16, 2025Updated last year
- JoinAI是一个开源仓库,专注于算法工程能力的培养,包括工程和数学原理的整理☆11Apr 20, 2025Updated last year
- ☆36Apr 13, 2026Updated 3 months ago
- PyTorch implementation of the paper "Discovering and Explaining the Representation Bottleneck of DNNs" (ICLR 2022 Oral)☆37Oct 30, 2024Updated last year
- We introduce BabyVision, a benchmark revealing the infancy of AI vision.☆237Jan 13, 2026Updated 6 months ago
- Some thoughts about writing scientific papers☆23Nov 8, 2024Updated last year
- A novel approach to improve the safety of large language models, enabling them to transition effectively from unsafe to safe state.☆72May 22, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- (ICLR 2026 🔥) Code for "The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs"☆80Feb 9, 2026Updated 6 months ago
- [ICML 2024] Code release for "On the Emergence of Cross-Task Linearity in Pretraining-Finetuning Paradigm"☆11Feb 20, 2025Updated last year
- ☆30May 22, 2024Updated 2 years ago
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe☆912Jun 29, 2026Updated last month
- PaperPub is an academic arena where diverse AI Agents read papers daily, pick apart each other's arguments, and fiercely debate.☆43Jun 12, 2026Updated 2 months ago
- instruction-following benchmark for large reasoning models☆49Apr 19, 2026Updated 3 months ago
- Official Implementation of MARS☆30Apr 21, 2026Updated 3 months ago