Diagnosing policy grounding in LLM judges via causal erasure. Companion code to the AIES 2026 submission READ.
☆21Aug 1, 2026Updated 2 months ago
Alternatives and similar repositories for READ
Users that are interested in READ are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Demo setups for ai-atlas-nexus☆16Updated this week
- Long-form factuality assessor for large language models☆43Updated this week
- The repo consists of a Python package that works with functional data. In particular, it includes two distinct methodologies: Functional …☆14Sep 18, 2025Updated last year
- FairCVtest: Testbed for Fair Automatic Recruitment and Multimodal Bias Analysis☆23Jul 24, 2023Updated 3 years ago
- An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems☆31Mar 13, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Offline-first sync engine for Flutter & Solid.js☆38Updated this week
- The Steerability toolkit is an extensible library for general purpose steering of language models.☆123Updated this week
- Proof Of Concept showcasing composable GPUs in Kubernetes☆30Sep 7, 2026Updated last month
- AI risk ontology☆26Aug 1, 2025Updated last year
- AutoML system for building trustworthy peptide bioactivity predictors☆42Aug 24, 2026Updated last month
- Ontology representing a 360-view of a person (or cohort) that spans across multiple domains, from health to social.☆34Sep 17, 2025Updated last year
- A collection of MCP servers for connecting real-world tools and data to LLMs.☆25Mar 16, 2026Updated 6 months ago
- Responsible Prompting is an LLM-agnostic tool that aims at dynamically supporting users in crafting prompts that embed responsible intent…☆50Updated this week
- Agent specification certification pipeline — from skill specs to governed, certified Mellea pipelines☆52Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The driver for LMCache core to run in vLLM☆69Feb 4, 2025Updated last year
- This is the data associated with the PERSUADE Corpus 2.0 version☆65Feb 3, 2026Updated 8 months ago
- TrustyAI Explainability Toolkit☆64Oct 1, 2026Updated last week
- A framework for designing, executing and analysing experiment campaigns☆65Updated this week
- Knowledge-enhanced learned representation enriches protein sequence and SMILES drug databases with a large Knowledge Graph fused from dif…☆65Apr 24, 2026Updated 5 months ago
- In-Context Explainability 360 toolkit☆72Updated this week
- Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation re…☆133Updated this week
- EvalAssist is an open-source project that simplifies using large language models as evaluators (LLM-as-a-Judge) of the output of other la…☆103Updated this week
- AI Atlas Nexus: tooling to bring together resources related to governance of foundation models.☆141Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM☆163Updated this week
- The Granite Guardian models are designed to detect risks in prompts and responses.☆181Aug 26, 2026Updated last month
- ☆205Jul 15, 2025Updated last year
- A Lossless Compression Library for AI pipelines☆339Apr 11, 2026Updated 5 months ago
- A framework for agentic tool use training with reinforcement learning☆205Apr 8, 2026Updated 6 months ago
- ☆273Jun 25, 2025Updated last year
- Fuse two frontier models into one Fable-tier answer: Opus 4.8 drafts, a second model (Opus 4.8 or GPT-5.5 via codex) checks, Opus fuses. …☆470Jul 20, 2026Updated 2 months ago
- Zero and Few shot named entity & relationships recognition☆403Sep 17, 2025Updated last year
- The official repo of GraphRAG-Bench for evaluating GraphRAG models. "When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrie…☆502Jun 7, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Python CLI for Google Trends analysis with advanced rate limiting, cookie auth, and beautiful terminal output. Built on pytrends with …☆396Apr 8, 2026Updated 6 months ago
- Workshop: Agentic Search for Context Engineering☆328Apr 1, 2026Updated 6 months ago
- A Claude skill that activates Fable-style agentic behavior: explicit multi-stage planning, sub-agent delegation, and self-verification.☆872Sep 2, 2026Updated last month
- Extract structured text from pdfs quickly☆728Jul 8, 2026Updated 3 months ago
- Deploy, and share agents with open infrastructure, free from vendor lock-in.☆1,166Updated this week
- Mellea is a library for writing generative programs.☆1,819Updated this week
- When Philosophy meets AI☆1,550Oct 20, 2025Updated 11 months ago