Qwen2.5 0.5B GRPO
☆86Feb 16, 2025Updated last year
Alternatives and similar repositories for qwen2.5-0.5b-grpo
Users that are interested in qwen2.5-0.5b-grpo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Train a Language Model with GRPO to create a schedule from a list of events and priorities☆272Apr 8, 2026Updated 4 months ago
- 这是一个open-r1的复现项目,对0.5B、1.5B、3B、7B的qwen模型进行GRPO训练,观察到一些有趣的现象。☆64Apr 13, 2025Updated last year
- Huggingface PPO Demo☆30Sep 7, 2025Updated 11 months ago
- Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style☆15Aug 18, 2025Updated 11 months ago
- Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models☆55Sep 2, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Our(team: winter) solutions for MICCAI Learn2Reg 2021.☆20May 8, 2022Updated 4 years ago
- This is a repository for fine-tuning Qwen2-Audio, currently supporting Distributed Data Parallel (DDP) and DeepSpeed.☆50Jul 28, 2025Updated last year
- Penn CIS 5650 (GPU Programming and Architecture) Final Project☆46Dec 11, 2023Updated 2 years ago
- ☆13Nov 8, 2022Updated 3 years ago
- This is the repo of our work titled “Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception”☆36Mar 31, 2026Updated 4 months ago
- [NeurIPS 2024 Oral] "Bayesian-Guided Label Mapping for Visual Reprogramming"☆12Dec 20, 2024Updated last year
- Cross Visual Prompt Tuning [ICCV 2025]☆13Aug 3, 2025Updated last year
- Simplistic Implementation of Zipformer:A faster and better encoder for automatic speech recognition in PyTorch☆22Jun 3, 2024Updated 2 years ago
- A toolkit for benchmarking on a wide variety of audio deepfake datasets.☆36May 22, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 这是一个从头训练大语言模型的项目,包括预训练、微调和直接偏好优化,模型拥有1B参数,支持中英文。☆862Feb 18, 2025Updated last year
- [ICLR 2024 Spotlight] Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization☆34Mar 12, 2026Updated 4 months ago
- This repository is the official Pytorch implementation of Balanced Product of Calibrated Experts for Long-Tailed Recognition (CVPR 2023).☆17Mar 13, 2025Updated last year
- A comparison of deepseek grpo and qwen gspo on Qwen2.5-1.5B-Instruct fine tunning.☆170Mar 28, 2026Updated 4 months ago
- This is an official implementation of "Self-supervised Image Denoising with Downsampled Invariance Loss and Conditional Blind-Spot Networ…☆18Sep 4, 2023Updated 2 years ago
- [ICLR 2025] EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing☆27Apr 1, 2025Updated last year
- Implement Code for UniMix and Bayias Compensated Loss☆19Mar 7, 2023Updated 3 years ago
- LIGHTVOC AN UPSAMPLING-FREE GAN VOCODER BASED ON CONFORMER AND INVERSE SHORT-TIME FOURIER TRANSFORM☆18May 17, 2024Updated 2 years ago
- A very simple GRPO implement for reproducing r1-like LLM thinking.☆1,701Nov 21, 2025Updated 8 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- a Neural Vocoder supporting Ring Attention, Conformer and NSF.☆25Aug 1, 2025Updated last year
- This repository presents FSD dataset for song deepfake detection.☆24Aug 18, 2025Updated 11 months ago
- LayoutFlow: Flow Matching for Layout Generation [Andrade Guerreiro et al., ECCV 2024]☆40Sep 29, 2025Updated 10 months ago
- SASV2 baseline, a track on ASVspoof5 phase2 challenge☆28Nov 12, 2025Updated 8 months ago
- OpenVLA Lightweight Version(0.5B). It uses qwen2-0.5B and fine-tunes using mllm format, without occupying LLM's inherent tokens. It repre…☆19Jan 7, 2026Updated 7 months ago
- [NeurIPS2023] Neural-Logic Human-Object Interaction Detection☆14Aug 24, 2024Updated last year
- A simple Transformer where the softmax has been replaced with normalization☆20Sep 11, 2020Updated 5 years ago
- ☆14Jul 17, 2024Updated 2 years ago
- D3PE (Deep Data-Driven Policy Evaluation) aims to evaluation a large set of candidate policies from a fixed dataset to select best ones.☆10Jun 2, 2022Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆13Nov 12, 2024Updated last year
- 语音合成VITS 纯中文微调☆12Mar 15, 2023Updated 3 years ago
- Collection of works for evaluating (and analyzing) large audio-language models (LALMs)☆41Aug 11, 2025Updated last year
- The offical implement of ImbSAM (Imbalanced-SAM)☆28Mar 4, 2024Updated 2 years ago
- ☆11Jul 27, 2021Updated 5 years ago
- Diff-SFCT: A Diffusion Model with Spatial-Frequency Cross Transformer for Medical Image Segmentation☆10Apr 15, 2024Updated 2 years ago
- Compute WER and SER for speech recognition evaluation☆28Jun 6, 2026Updated 2 months ago