RL significantly the reasoning capability of Qwen2.5-1.5B-Instruct
☆31Feb 23, 2025Updated last year
Alternatives and similar repositories for RL4LLM
Users that are interested in RL4LLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Composition of Multimodal Language Models From Scratch☆15Aug 16, 2024Updated 2 years ago
- Custom triton kernels for training Karpathy's nanoGPT.☆19Oct 21, 2024Updated last year
- ☆16Mar 18, 2025Updated last year
- built a 124M param GPT☆23Jan 28, 2025Updated last year
- A distilled DeepSeek-R1 variant built on Qwen2.5-32B, fine-tuned with curated data for enhanced performance and efficiency. <metadata> gp…☆15Mar 11, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A Python reimplementation + extension of "Planning with Large Language Models for Code Generation" (https://arxiv.org/abs/2303.05510)☆17Dec 1, 2023Updated 2 years ago
- Open-source transpiler for CUDA Tile (13.1) migration☆19Dec 9, 2025Updated 8 months ago
- manipulating cointegrated pairs to achieve a market-neutral strategy that outperforms indices☆11Jan 12, 2021Updated 5 years ago
- Example of multi-process, multi-GPU training using Torch-parallel, nVidia-nccl, and nVidia-MPS☆17Sep 22, 2016Updated 9 years ago
- Pytorch Implementation of the paper: "Learning to (Learn at Test Time): RNNs with Expressive Hidden States"☆24Updated this week
- RuCLIP-SB (Russian Contrastive Language–Image Pretraining SWIN-BERT) is a multimodal model for obtaining images and text similarities and…☆15Jan 25, 2022Updated 4 years ago
- Community Open Source Implementation of GPT4o in PyTorch☆32Aug 3, 2026Updated 2 weeks ago
- Our solution to ML Talent Match hackathon☆11Mar 22, 2024Updated 2 years ago
- Classify documents using Python based on SVM and TF-IDF.☆15Nov 19, 2019Updated 6 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- PyTorch implementation of GRPO.☆16Apr 21, 2025Updated last year
- coloring terminal text with intensities (used for plotting probability, entropy with tokens)☆12Oct 11, 2024Updated last year
- A few models converted from caffe to CoreMLs format.☆15Jun 6, 2017Updated 9 years ago
- ☆25Feb 13, 2026Updated 6 months ago
- 机器人人工智能,优达学城cs373作业。 Artificial Intelligence for Robotics, this repository contains all the homework…☆12Nov 12, 2017Updated 8 years ago
- Scratchpad/Chain-of-Thought Prompts☆12Jun 6, 2022Updated 4 years ago
- Few-Shot Prompting - Chain-of-Thought (CoT) Prompting - Hallucinations - Self-Consistency - Generated Knowledge Prompting - Tree of …☆30Nov 15, 2023Updated 2 years ago
- ☆10May 19, 2022Updated 4 years ago
- A Statistical Arbitrage Strategy to trade Cryptocurrency Pairs☆13Nov 6, 2020Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"☆13Jul 18, 2024Updated 2 years ago
- This is the official code for the paper "Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning" (NeurIPS2024)☆29Sep 10, 2024Updated last year
- A Qwen .5B reasoning model trained on OpenR1-Math-220k☆14Jul 29, 2026Updated 2 weeks ago
- ☆14Oct 30, 2024Updated last year
- ☆11Feb 9, 2024Updated 2 years ago
- ☆14Mar 2, 2025Updated last year
- This is the repository containing the solution of the homework for the CS224W course at Stanford: Machine Learning with Graphs☆11Jul 19, 2020Updated 6 years ago
- ☆10Mar 28, 2022Updated 4 years ago
- ML from scratch in Jax☆12Aug 20, 2025Updated 11 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Skoltech NLA 2025 course.☆17Nov 30, 2025Updated 8 months ago
- Source code for the paper "Positional Attention: Expressivity and Learnability of Algorithmic Computation"☆14May 26, 2025Updated last year
- Geographical Graph Attention Networks: Spatial Deep Learning Models for Spatial Prediction and Exploratory Spatial Data Analysis☆19Jul 28, 2025Updated last year
- Stable Diffusion in TensorRT 8.5+☆15Mar 19, 2023Updated 3 years ago
- Code for reproducing the Stanford Alpaca InstructLLaMA result on consumer hardware☆19Mar 16, 2023Updated 3 years ago
- ☆69Mar 21, 2025Updated last year
- Teknofest 2023 Türkçe Doğal Dil İşleme yarışması için gerçekleştirilen bu çalışma, Shap Analizi yöntemi kullanılarak modelin tahminlerini…☆27Mar 31, 2023Updated 3 years ago