Diagnosing policy grounding in LLM judges via causal erasure. Companion code to the AIES 2026 submission READ.
☆20Jul 2, 2026Updated 2 weeks ago
Alternatives and similar repositories for READ
Users that are interested in READ are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Demo setups for ai-atlas-nexus☆16Updated this week
- Long-form factuality assessor for large language models☆34Jul 14, 2026Updated last week
- The repo consists of a Python package that works with functional data. In particular, it includes two distinct methodologies: Functional …☆14Sep 18, 2025Updated 10 months ago
- FairCVtest: Testbed for Fair Automatic Recruitment and Multimodal Bias Analysis☆22Jul 24, 2023Updated 2 years ago
- An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems☆29Mar 13, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The AI Steerability 360 toolkit is an extensible library for general purpose steering of LLMs.☆99Jul 8, 2026Updated last week
- Proof Of Concept showcasing composable GPUs in Kubernetes☆24Updated this week
- AI risk ontology☆25Aug 1, 2025Updated 11 months ago
- The driver for LMCache core to run in vLLM☆69Feb 4, 2025Updated last year
- AutoML system for building trustworthy peptide bioactivity predictors☆40Updated this week
- Ontology representing a 360-view of a person (or cohort) that spans across multiple domains, from health to social.☆35Sep 17, 2025Updated 10 months ago
- A collection of MCP servers for connecting real-world tools and data to LLMs.☆24Mar 16, 2026Updated 4 months ago
- Responsible Prompting is an LLM-agnostic tool that aims at dynamically supporting users in crafting prompts that embed responsible intent…☆49Jan 26, 2026Updated 5 months ago
- Agent specification certification pipeline — from skill specs to governed, certified Mellea pipelines☆39Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is the data associated with the PERSUADE Corpus 2.0 version☆61Feb 3, 2026Updated 5 months ago
- TrustyAI Explainability Toolkit☆63Jun 19, 2026Updated last month
- A framework for designing, executing and analysing experiment campaigns☆55Updated this week
- Knowledge-enhanced learned representation enriches protein sequence and SMILES drug databases with a large Knowledge Graph fused from dif…☆64Apr 24, 2026Updated 2 months ago
- In-Context Explainability 360 toolkit☆70Mar 9, 2026Updated 4 months ago
- Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation re…☆94Jul 4, 2026Updated 2 weeks ago
- EvalAssist is an open-source project that simplifies using large language models as evaluators (LLM-as-a-Judge) of the output of other la…☆102Apr 9, 2026Updated 3 months ago
- AI Atlas Nexus: tooling to bring together resources related to governance of foundation models.☆129Updated this week
- ☆199Jul 15, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM☆140Jun 30, 2026Updated 3 weeks ago
- The Granite Guardian models are designed to detect risks in prompts and responses.☆164May 5, 2026Updated 2 months ago
- A Lossless Compression Library for AI pipelines☆322Apr 11, 2026Updated 3 months ago
- A framework for agentic tool use training with reinforcement learning☆195Apr 8, 2026Updated 3 months ago
- ☆271Jun 25, 2025Updated last year
- Zero and Few shot named entity & relationships recognition☆400Sep 17, 2025Updated 10 months ago
- Fuse two frontier models into one Fable-tier answer: Opus 4.8 drafts, a second model (Opus 4.8 or GPT-5.5 via codex) checks, Opus fuses. …☆457Updated this week
- The official repo of GraphRAG-Bench for evaluating GraphRAG models. "When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrie…☆464Jun 7, 2026Updated last month
- Python CLI for Google Trends analysis with advanced rate limiting, cookie auth, and beautiful terminal output. Built on pytrends with …☆370Apr 8, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Workshop: Agentic Search for Context Engineering☆318Apr 1, 2026Updated 3 months ago
- A Claude skill that activates Fable-style agentic behavior: explicit multi-stage planning, sub-agent delegation, and self-verification.☆789Jul 10, 2026Updated last week
- Extract structured text from pdfs quickly☆707Jul 8, 2026Updated last week
- Deploy, and share agents with open infrastructure, free from vendor lock-in.☆1,132Updated this week
- Mellea is a library for writing generative programs.☆1,772Updated this week
- When Philosophy meets AI☆1,525Oct 20, 2025Updated 9 months ago
- Fast and Accurate Code Search for Agents. Uses ~98% fewer tokens than grep+read☆5,665Updated this week