☆44Mar 6, 2025Updated last year
Alternatives and similar repositories for grpo-loss
Users that are interested in grpo-loss are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MLLM @ Game☆17May 12, 2025Updated last year
- ☆17Jan 14, 2026Updated 6 months ago
- Awesome_CV的中文版本,clone本项目到overleaf即可轻松愉快编写自己的CV☆18May 24, 2024Updated 2 years ago
- ☆10Jan 1, 2022Updated 4 years ago
- WisdoMentor - Series: A LLM for undergraduates | 博导智言(辅助大学生 学习)☆13May 9, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- differentiable top-k operator☆23Dec 30, 2024Updated last year
- 基于Llama3,通过进一步CPT,SFT,ORPO得到的中文版Llama3☆16Apr 24, 2024Updated 2 years ago
- ☆20Apr 16, 2025Updated last year
- Robust Self-augmentation for NER with Meta-reweighting☆29Nov 8, 2022Updated 3 years ago
- ☆11Jan 29, 2026Updated 6 months ago
- LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild☆16Oct 31, 2024Updated last year
- Cross-domain word representation learning☆10May 23, 2015Updated 11 years ago
- minimal-cost for training 0.5B R1-Zero☆817May 14, 2025Updated last year
- ☆15May 27, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This announcement is used in the ATMHUFK's video. The original is from the another up,Which is called 原无奇变in Chinese.You can use it to av…☆10Jan 26, 2025Updated last year
- ☆11Oct 11, 2023Updated 2 years ago
- this is based on the paper Chain-of-Retrieval Augmented Generation☆15Mar 29, 2025Updated last year
- [ICLR 2025] SuperCorrect: Advancing Small LLM Reasoning with Thought Template Distillation and Self-Correction☆90Mar 23, 2025Updated last year
- This is the official implementation of TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data☆13Jul 21, 2024Updated 2 years ago
- ☆12Jun 30, 2024Updated 2 years ago
- Analyzing Latent Concept in Pre-trained Transformer Models☆12Jul 18, 2022Updated 4 years ago
- Weakly-supervised road-lane markings detection for autonomous driving, mitigating the lack of training data☆14Oct 8, 2024Updated last year
- ☆11Apr 29, 2019Updated 7 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ACM MM2026] This is the official implementation of MedCCO☆17Jul 12, 2026Updated 3 weeks ago
- Akcio is a demonstration project for Retrieval Augmented Generation (RAG). It leverages the power of LLM to generate responses and uses v…☆12Oct 30, 2023Updated 2 years ago
- ☆10Apr 6, 2022Updated 4 years ago
- [ICCV2019] Attract or Distract: Exploit the Margin of Open Set☆35Dec 19, 2020Updated 5 years ago
- Code for ICCV2023 paper: Homography Guided Temporal Fusion for Road Line and Marking Segmentation☆14Oct 13, 2024Updated last year
- This is the implementation of CounterCurate, the data curation pipeline of both physical and semantic counterfactual image-caption pairs.☆19Jun 27, 2024Updated 2 years ago
- ☆12May 23, 2020Updated 6 years ago
- Code for SafeMERGE (ICLR 2025).☆15Apr 1, 2025Updated last year
- COVID-19 Risk Estimation for L.A. County using a Bayesian Time-varying SIR-model☆12Feb 17, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- HiSV: a computational pipeline for structural variation detection from Hi-C data☆18Jun 20, 2025Updated last year
- Vision-Language Pre-Training for Boosting Scene Text Detectors (CVPR2022)☆12Mar 21, 2022Updated 4 years ago
- [ACL 2025] We introduce ScaleQuest, a scalable, novel and cost-effective data synthesis method to unleash the reasoning capability of LLM…☆69Oct 27, 2024Updated last year
- Eden Flux LoRA trainer and full-finetuning☆23Mar 21, 2025Updated last year
- ☆17Dec 11, 2024Updated last year
- Dataset of hackable TerminalBench-style tasks and exploit trajectories☆35Apr 18, 2026Updated 3 months ago
- Official repository for AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning☆17Jul 24, 2025Updated last year