Long CoT Fine-Tuning and Reinforcement Learning for LLMs in the Context of the 24-Point Game: A Toy Project
☆27Feb 22, 2025Updated last year
Alternatives and similar repositories for LLM4Game24
Users that are interested in LLM4Game24 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 超简单复现Deepseek-R1-Zero和Deepseek-R1,以「24点游戏」为例。通过zero-RL、SFT以及SFT+RL,以激发LLM的自主验证反思能力。 About Clean, minimal, accessible reproduction of Dee…☆35Apr 5, 2025Updated last year
- [WWW 2024] Inductive Cognitive Diagnosis for Fast Student Learning in Web-Based Online Intelligent Education Systems☆11Apr 18, 2024Updated 2 years ago
- Anti exploration in offline reinforcement learning☆11May 17, 2021Updated 5 years ago
- The code and data for paper "Large Language Models are few(1)-shot Table Reasoners" [EACL2023]☆47Apr 30, 2024Updated 2 years ago
- [ECAI 2023] QCCDM: A Q-Augmented Causal Cognitive Diagnosis Model for Student Learning☆12Aug 4, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code of Scaling Multi-Objective Security Games Provably via Space Discretization Based Evolutionary Search (SDES).☆11Nov 26, 2024Updated last year
- [ EMNLP 2025 Main ] Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs☆18Nov 7, 2025Updated 9 months ago
- This repository includes code and materials for the paper "Efficient PRM Training Data Synthesis via Formal Verification" (ACL 2026 Findi…☆19Apr 7, 2026Updated 4 months ago
- Capturing Homogeneous Influence among Students: Hypergraph Cognitive Diagnosis for Intelligent Education Systems. This paper has been pub…☆19Jun 17, 2026Updated 2 months ago
- 📰 Must-read papers on Diffusion Models for Text Generation 🔥☆20Jun 21, 2024Updated 2 years ago
- [KDD 2024] ORCDF: An Oversmoothing-Resistant Cognitive Diagnosis Framework for Student Learning in Online Education System☆16Dec 9, 2024Updated last year
- ☆19Jul 18, 2021Updated 5 years ago
- [KDD 2025] Fine-tuning Multimodal Large Language Models for Product Bundling☆16Sep 20, 2025Updated 10 months ago
- The paper "Symbolic Cognitive Diagnosis via Hybrid Optimization for Intelligent Education Systems" published in proceedings of the 38th A…☆15Apr 18, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Author's implementation of ReBRAC, a minimalist improvement upon TD3+BC☆19Oct 22, 2023Updated 2 years ago
- A lightweight reimplementation of Adversarially Trained Actor Critic☆19Mar 19, 2026Updated 4 months ago
- The implementation of paper "Strategy-aware Bundle Recommender System", SIGIR'23.☆17Sep 4, 2023Updated 2 years ago
- ☆19Jun 14, 2024Updated 2 years ago
- Author's PyTorch implementation of ICML'23 paper "Policy Regularization with Dataset Constraint for Offline Reinforcement Learning" for D…☆17Nov 8, 2024Updated last year
- 🚀 Engram-PEFT: An unofficial implementation of DeepSeek Engram. Inject high-capacity conditional memory into LLMs via sparse retrieval P…☆41Jul 17, 2026Updated last month
- ☆19Sep 7, 2025Updated 11 months ago
- Adapted source code of Niklaus Wirth's "Compiler Construction" book☆21May 11, 2023Updated 3 years ago
- Diffusion-Based ECG Noise Quantification via Anomaly Detection☆22Mar 31, 2026Updated 4 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- 订餐系统☆14Mar 5, 2016Updated 10 years ago
- ☆20Jun 25, 2023Updated 3 years ago
- The implementation of paper "EliMRec: Eliminating single-modal bias in multimedia recommendation", MM'22.☆24Dec 7, 2023Updated 2 years ago
- [MM 2025] Towards Modality Generalization: A Benchmark and Prospective Analysis☆31May 22, 2025Updated last year
- A controlled benchmark on evaluating and studying the dynamics of Long Context Language Models☆26Oct 17, 2025Updated 10 months ago
- LLM手撕代码合集☆23Mar 25, 2025Updated last year
- Deep Learning with EEG.☆27Oct 11, 2025Updated 10 months ago
- Evaluation on Logical Reasoning and Abstract Reasoning Challenges☆31Apr 21, 2025Updated last year
- 2022中山大学编译原理☆23Jul 27, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Towards Domain Free Transformer for Generalized EEG Pre-training☆30Apr 22, 2025Updated last year
- ☆16Jul 10, 2025Updated last year
- [ICLR 2025] The implementation of paper "Preference Diffusion for Recommendation"☆29Apr 21, 2025Updated last year
- Pretraining codebase for Apertus models, based on Megatron-LM☆21Sep 25, 2025Updated 10 months ago
- ☆29Mar 13, 2026Updated 5 months ago
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning☆29Feb 21, 2022Updated 4 years ago
- A comprehensive paper list of Table-based Question Answering.☆40Sep 1, 2023Updated 2 years ago