☆101May 23, 2025Updated last year
Alternatives and similar repositories for LLM-Post-Training
Users that are interested in LLM-Post-Training are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 基于Llama3,通过进一步CPT,SFT,ORPO得到的中文版Llama3☆16Apr 24, 2024Updated 2 years ago
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"☆17Jun 30, 2025Updated last year
- This repo is reproduction resources for linear alignment paper, still working☆17May 19, 2024Updated 2 years ago
- Code and models for EMNLP 2024 paper "WPO: Enhancing RLHF with Weighted Preference Optimization"☆41Sep 24, 2024Updated 2 years ago
- Official PyTorch implementation of `[ACMMM 2023]Relational Contrastive Learning for Scene Text Recognition`☆17Sep 22, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- a family of highly capabale yet efficient large multimodal models☆194Aug 23, 2024Updated 2 years ago
- Awesome Reasoning LLM Tutorial/Survey/Guide☆2,564Sep 5, 2026Updated 3 weeks ago
- Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition☆15Oct 26, 2025Updated 11 months ago
- trending repositories and news related to AI☆11Mar 22, 2019Updated 7 years ago
- [ACL 24 Findings] Implementation of Resonance RoPE and the PosGen synthetic dataset.☆24Mar 5, 2024Updated 2 years ago
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism☆18Nov 18, 2025Updated 10 months ago
- [ICLR 2026] RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling☆42Feb 25, 2026Updated 7 months ago
- [ACL2026 MainConference]Code Repo for paper "Scaling Behaviors of LLM Reinforcement Learning Post-Training"☆26Jul 1, 2026Updated 2 months ago
- KnowLA: Enhancing Parameter-efficient Finetuning with Knowledgeable Adaptation, NAACL 2024☆16Jul 29, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR 2024] Adapting Short-Term Transformers for Action Detection in Untrimmed Videos☆11Jun 11, 2024Updated 2 years ago
- This is the official code for OThink-R1 project.☆21Jun 19, 2025Updated last year
- Mixture-of-Basis-Experts for Compressing MoE-based LLMs☆40Dec 24, 2025Updated 9 months ago
- CS172 Final project: Text Image Super-Resolution Reconstruction☆14Jun 15, 2020Updated 6 years ago
- [TOIS 2024] Target-constrained Bidirectional Planning for Generation of Target-oriented Proactive Dialogue☆14Oct 18, 2025Updated 11 months ago
- ☆55Feb 11, 2025Updated last year
- [ACL 2026] G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance☆16Sep 15, 2026Updated last week
- UFT: Unifying Supervised and Reinforcement Fine-Tuning☆33Jun 30, 2025Updated last year
- The official repository for Trust-Region Adaptive Policy Optimization (TRAPO) – a novel hybrid framework designed to enhance large langua…☆16Mar 2, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A repository for awesome resources in mechanistic interpretability☆17Jan 18, 2023Updated 3 years ago
- Code repository of "ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs" (ACL 2025).☆23Mar 22, 2026Updated 6 months ago
- [ICLR 2023] Soft Neighbors are Positive Supporters in Contrastive Visual Representation Learning☆15Aug 2, 2023Updated 3 years ago
- ☆27Aug 28, 2023Updated 3 years ago
- PatchPilot: A Stable and Cost-Efficient Agentic Patching Framework☆24Jun 5, 2025Updated last year
- Pytorch implementation of Centered Kernel Alignment(CKA) and its minibatch version.☆11May 11, 2022Updated 4 years ago
- A step-by-step tutorial about how to use Distributed Data Parallel feature of PyTorch☆16Nov 20, 2020Updated 5 years ago
- 嵌入式作業系統分析與實作 ANALYSIS AND IMPLEMENTATION OF EMBEDDED OPERATING SYSTEMS, 張大緯☆16Jun 22, 2026Updated 3 months ago
- Repository having the code and models from the paper: data2vec-aqc: Search for the right Teaching Assistant in the Teacher-Student traini…☆12Mar 18, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- WisdoMentor - Series: A LLM for undergraduates | 博导智言(辅助大学生 学习)☆13May 9, 2024Updated 2 years ago
- This is the codes of "DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrieval"☆15Aug 11, 2026Updated last month
- LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards☆40Jun 1, 2026Updated 3 months ago
- Course materials for introduction to web-based application development, fall 2017.☆14Dec 14, 2017Updated 8 years ago
- Agent-RRM: Exploring Reasoning Reward Model for Agents☆73Mar 17, 2026Updated 6 months ago
- PyTorch Implementation of "NDDR-CNN: Layerwise Feature Fusing in Multi-Task CNNs by Neural Discriminative Dimensionality Reduction"☆14Jun 29, 2019Updated 7 years ago
- ☆18Jul 8, 2025Updated last year