Curriculum-RLAIF is a data-centric curriculum learning framework for reward model training in RLAIF-based LLM alignment
☆23Apr 18, 2026Updated 4 months ago
Alternatives and similar repositories for Curriculum-RLAIF
Users that are interested in Curriculum-RLAIF are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Data and Code Repository for “STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems”☆17Apr 17, 2026Updated 4 months ago
- MemoryDial☆15Mar 10, 2026Updated 5 months ago
- [ACL2026] UCAS: Uncertainty-aware Advantage Shaping for RLVR☆32Apr 14, 2026Updated 4 months ago
- [ACL 2026] Dissecting Failure Dynamics in Large Language Model Reasoning☆18Apr 17, 2026Updated 4 months ago
- This is the official repository for JailExpert☆23Sep 9, 2025Updated 11 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing - ACL Findings 2026☆25Jul 15, 2026Updated last month
- [NeurIPS 2025] SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly☆17Oct 22, 2025Updated 10 months ago
- [ACL26 Findings] TopoDIM: One-shot Topology Generation of Diverse Interaction Modes for Multi-Agent Systems☆19Jan 19, 2026Updated 7 months ago
- Official Codebase of the ACL 2026 Oral paper "Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contra…☆28Jun 25, 2026Updated 2 months ago
- videoPro: Adaptive Program Reasoning for Long Video Understanding☆45Apr 15, 2026Updated 4 months ago
- [🏆CVPR'26] Official Repo for IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding☆33Jun 2, 2026Updated 2 months ago
- ☆17Feb 22, 2026Updated 6 months ago
- Implementation for paper Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs, which is accepted by ACL 2026 (main con…☆16Oct 10, 2025Updated 10 months ago
- This is the official repo for the paper "General365: Benchmarking General Reasoning in LLMs under High Difficulty and Diversity".☆88Apr 14, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆79Apr 12, 2026Updated 4 months ago
- Sparse Adapter Fusion for Continual Learning in NLP - EACL 2026☆16Apr 9, 2026Updated 4 months ago
- This is the official repo for the paper "AMO-Bench: Large Language Models Still Struggle in High School Math Competitions".☆171Feb 6, 2026Updated 6 months ago
- ACL26 Long Paper☆19Jul 4, 2026Updated last month
- [ACL 2026] Context-Agent: Dynamic Discourse Trees for Non-Linear Dialogue☆24Apr 14, 2026Updated 4 months ago
- The implementation of ACL 2026 paper "Rethinking entropy interventions in rlvr: An entropy change perspective"☆27Jul 19, 2026Updated last month
- [🏆AAAI'25] Official Repo for ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area.☆91Apr 14, 2026Updated 4 months ago
- [AAAI2024] Debiasing Multimodal Sarcasm Detection with Contrastive Learning☆17Jan 5, 2024Updated 2 years ago
- ☆13May 31, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- 本项目综合运用d3、echarts来完成可视化工作,实现了对nba两场比赛的可视化数据分析,包括球员运动轨迹、个人数据、传球次数以及得分位置等多种可交互式图表。通过可视化方法,我们能够进一步深入分析球队的具体情况,便于制定更佳的战术。☆15Dec 19, 2022Updated 3 years ago
- [REALM25 @ ACL25] - "StateAct" Official Paper Repo (SOTA LLM Agent)☆19Aug 7, 2026Updated 2 weeks ago
- This is the GPT2 baseline for ProtoQA☆12Jan 3, 2022Updated 4 years ago
- [🏆ECCV'26] Official Repo for SlowBA: An efficiency backdoor attack towards VLM-based GUI agents☆18Jul 1, 2026Updated last month
- Domain Adaptation through Synthesis☆11Dec 15, 2018Updated 7 years ago
- MinT-2M: Long-context training system for resident-prefix GRPO☆45Jul 24, 2026Updated last month
- The implementation of ACL main 2026 paper "ReCreate: Reasoning and Creating Domain Agents Driven by Experience"☆165Apr 29, 2026Updated 3 months ago
- Official code for the paper "FairerCLIP: Debiasing CLIP’s Zero-Shot Predictions using Functions in RKHSs".☆16Oct 14, 2025Updated 10 months ago
- Structured Domain Adaptation with Online Relation Regularization for Unsupervised Person Re-ID☆18Jun 9, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆23Aug 9, 2026Updated 2 weeks ago
- uPU, nnPU and PN learning with Extra Trees classifier.☆20Dec 2, 2024Updated last year
- ☆26Jun 3, 2026Updated 2 months ago
- Official repo for PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks☆17Feb 17, 2026Updated 6 months ago
- TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs☆124Jul 27, 2026Updated last month
- ☆33Aug 21, 2025Updated last year
- ☆16Jan 16, 2025Updated last year