Learning and research after DeepSeek-R1, around test-time computing, resurgence of RL, and new LLM learning/application paradigms.
☆24Apr 23, 2026Updated 4 months ago
Alternatives and similar repositories for Post-DeepSeek-R1_LLM-RL
Users that are interested in Post-DeepSeek-R1_LLM-RL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- XWikisCorpus, cross-lingual summarisation, multi-lingual summarisation, pre-trained language models, zero-shot and few-shot summarisation…☆10Nov 4, 2022Updated 3 years ago
- Official Implementation of Papar CM2☆26Apr 21, 2026Updated 4 months ago
- AbstainQA, ACL 2024☆30Feb 4, 2026Updated 7 months ago
- Scaffold for NLP researcher to quickly set up the codebase☆17Mar 25, 2025Updated last year
- Learning to Prune: Exploring the Frontier of Fast and Accurate Parsing☆22Sep 24, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR 2024] This is the official implementation for the paper: "Beyond imitation: Leveraging fine-grained quality signals for alignment"☆10May 5, 2024Updated 2 years ago
- Code for experiments on self-prediction as a way to measure introspection in LLMs☆17Dec 10, 2024Updated last year
- This repository contains the dataset and code for "WiCE: Real-World Entailment for Claims in Wikipedia" in EMNLP 2023.☆44Dec 15, 2023Updated 2 years ago
- Jointly Optimizing Large Language Models for Reasoning and Self-Refinement☆14Apr 22, 2026Updated 4 months ago
- [KDD'23] This is the code repo for our KDD'23 paper "DyGen: Learning from Noisy Labels via Dynamics-Enhanced Generative Modeling".☆11Jun 14, 2023Updated 3 years ago
- Forcing Diffuse Distributions out of Language Models☆18Sep 10, 2024Updated last year
- Code and Results for "Universals of word order reflect optimization of grammars for efficient communication"☆14Aug 5, 2022Updated 4 years ago
- [ACL 2026 Main] Official Repo for Paper "Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Ali…☆16Jul 1, 2026Updated 2 months ago
- The code implementation of the paper Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks (A…☆13Jul 16, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ICLR 2025] On Evluating the Durability of Safegurads for Open-Weight LLMs☆13Jun 20, 2025Updated last year
- Preparing for ML Interviews.☆53Jan 12, 2026Updated 7 months ago
- Code for the paper "Data Feedback Loops: Model-driven Amplification of Dataset Biases"☆18Sep 9, 2022Updated 3 years ago
- [ICLR 2025] FLAT: LLM Unlearning via Loss Adjustment with Only Forget Data☆14Feb 26, 2025Updated last year
- This repository includes the code implementation of the paper Improving Pacing in Long-Form Story Planning by Yichen Wang, Kevin Yang, Xi…☆18Nov 19, 2024Updated last year
- ☆15May 28, 2024Updated 2 years ago
- Codebase for LLM story generation; updated version of https//github.com/yangkevin2/doc-story-generation☆93May 13, 2026Updated 3 months ago
- RHO: Evolving Agents in the Dark — Retrospective Harness Optimization via Self-Preference. Improving LLM agents from unlabeled past traje…☆54Jun 12, 2026Updated 2 months ago
- A curated collection of research papers exploring diversity in Large Language Model text generation. This repository tracks cutting-edge …☆16Jun 19, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Model for multi-turn tool calling of bash functions☆34Jan 26, 2026Updated 7 months ago
- Some Pwn Challenges from winesap.☆14Aug 15, 2019Updated 7 years ago
- Pytorch implementation of QAnet☆13May 6, 2018Updated 8 years ago
- [ACL2025 Best Paper] Language Models Resist Alignment☆52Jun 11, 2025Updated last year
- An experimental custom seq-2-seq model with both layer-wise (inter-layer), and intra-layer attention (attention to previous hidden states…☆10Nov 30, 2017Updated 8 years ago
- Code and Hummingbird dataset for EMNLP 2021 paper "Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica"☆14Apr 13, 2022Updated 4 years ago
- Repo for paper: Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge☆14Feb 20, 2024Updated 2 years ago
- Competitive Programming Code Template☆10Nov 6, 2022Updated 3 years ago
- ☆20Nov 24, 2020Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code for SafeMERGE (ICLR 2025).☆15Apr 1, 2025Updated last year
- [TMLR, Survey Certification] A technical and progressive review of self-improvement of LLMs for the future.☆44Aug 20, 2026Updated 2 weeks ago
- Pytorch-Lightning Seq2seq model with the use of recurrent neural network☆10Mar 29, 2021Updated 5 years ago
- MoCo: A One-Stop Shop for Model Collaboration Research☆63Updated this week
- BERT Baseline for the Natural Questions☆11Jan 24, 2019Updated 7 years ago
- personalized-llms with allen institute☆13Jun 22, 2023Updated 3 years ago
- Accompanying repo for the DP2O paper accepted by AAAI 2024 main conference☆17Mar 28, 2024Updated 2 years ago