Official implementation of the paper "Contextual Counterfactual Credit Assignment for Multi-Agent Reinforcement Learning in LLM Collaboration". (by Yanjun Chen)
☆36Mar 10, 2026Updated 5 months ago
Alternatives and similar repositories for C3
Users that are interested in C3 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AT2PO: Agentic Turn-based Policy Optimization via Tree Search☆22May 21, 2026Updated 3 months ago
- Real-time AI governance platform implementing a Cognitive Digital Twin framework. Monitors AI decisions, traces reasoning steps, detect…☆29Jul 29, 2026Updated last month
- [EMNLP 2024 Main] Official implementation of the paper "Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mech…☆15Oct 8, 2024Updated last year
- The code for paper FLDCF, with various forgery detection and localization methods.☆19Mar 16, 2026Updated 5 months ago
- [arXiv] "Linear Dynamics in the RLVR Training of Large Language Models"☆19May 25, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆23Jul 10, 2025Updated last year
- ☆27Jun 5, 2025Updated last year
- Offline Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits☆11Oct 21, 2024Updated last year
- Official implementation of "Diffusion Language Models Know the Answer Before Decoding"☆61Apr 28, 2026Updated 4 months ago
- This repository contains the implementation of reinforcement learning algorithms like PPO and A2C, to solve the problem: Dynamic Obstacle…☆19Jan 17, 2022Updated 4 years ago
- ☆17Jan 24, 2024Updated 2 years ago
- ☆19Apr 9, 2026Updated 4 months ago
- pytorch implementation for "Mutual Information Neural Estimation"☆11Dec 13, 2019Updated 6 years ago
- Official Code For EMNLP2025 Findings: {DLPO : Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Le…☆10Dec 25, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆46Jun 27, 2025Updated last year
- ☆12Aug 15, 2020Updated 6 years ago
- [ICLR 2023] The official code for paper "Guarded Policy Optimization with Imperfect Online Demonstrations"☆14Apr 30, 2023Updated 3 years ago
- ☆25Feb 24, 2023Updated 3 years ago
- This repository provides a summarization of recent empirical studies/human studies that measure human understanding with machine explanat…☆14Jul 24, 2024Updated 2 years ago
- Recovery-Bench is a benchmark for evaluating the capability of LLM agents to recover from mistakes☆28Jun 17, 2026Updated 2 months ago
- AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems☆17May 12, 2026Updated 3 months ago
- ☆14Nov 22, 2025Updated 9 months ago
- ☆14Nov 19, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A dual-agent framework leveraging left-brain logic (structured coding, syntax validation, debugging) and right-brain intuition (macro arc…☆20Jun 3, 2026Updated 2 months ago
- A message bus that lets AI assistants talk to each other. Works with Claude, ChatGPT, Gemini, Perplexity, and any AI that supports MCP or…☆17Updated this week
- Agent-Omit: Training Efficient LLM Agents for Adaptive Thought and Observation Omission via Reinforcement Learning☆33May 11, 2026Updated 3 months ago
- ☆23Updated this week
- The implementation for ACL 2026: MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching.☆20Jul 27, 2026Updated last month
- ☆169Dec 15, 2025Updated 8 months ago
- The open-source Claude Code GUI — parallel multi-project sessions, multi-engine BYOK (Codex / DeepSeek / Kimi / Ollama), terminal, browse…☆34Updated this week
- ☆117Oct 21, 2025Updated 10 months ago
- RETROAGENT: From Solving to Evolving via Retrospective Dual Intrinsic Feedback☆30Mar 30, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICLR 2025] Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization☆32Jan 7, 2026Updated 7 months ago
- Official Implementation of ReALFRED (ECCV'24)☆48Oct 11, 2024Updated last year
- Cooperative Graph-based Networked Agent Challenges for Multi-Agent Reinforcement Learning☆16Jan 26, 2026Updated 7 months ago
- Official code for the paper Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception. The code is based on t…☆21Aug 5, 2025Updated last year
- The code for Adaptive Semantic-Enhanced Denoising Diffusion Probabilistic Model for Remote Sensing Image Super-Resolution☆45Apr 11, 2025Updated last year
- ☆16Dec 13, 2022Updated 3 years ago
- Curated list of AI + GRC resources: AI governance frameworks and AI-powered compliance tools☆22Feb 16, 2026Updated 6 months ago