Evaluation Matrices for Explainability Methods
☆15Nov 5, 2025Updated 9 months ago
Alternatives and similar repositories for xai_evals
Users that are interested in xai_evals are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DL Backtrace is a new explainablity technique for deep learning models that works for any modality and model type.☆27May 13, 2026Updated 3 months ago
- Aligntune : A Modular Toolkit for Post Training Alignment of LLMs☆38Updated this week
- ☆48Nov 6, 2025Updated 9 months ago
- TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models☆124Aug 11, 2026Updated 2 weeks ago
- Benchmark to Evaluate EXplainable AI☆21Mar 14, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆11Feb 11, 2025Updated last year
- A curated reading list of research in Sparse Autoencoders, Feature Extraction and related topics in Mechanistic Interpretability☆33Jan 30, 2025Updated last year
- Implementing LRP (Layer-wise Relevance Propagation) for a sequence-to-sequence model with GRU layers.☆12Sep 8, 2023Updated 2 years ago
- Use AI to personify books, so that you can talk to them 🙊☆18Mar 25, 2023Updated 3 years ago
- NYU Tandon Machine Learning and Finance Fall 2022☆11Dec 13, 2022Updated 3 years ago
- A tool for model sparse based on torch.fx☆13Jun 3, 2024Updated 2 years ago
- TabMap for high-performance tabular data analysis - Nature BME☆24Jan 8, 2025Updated last year
- Sifaka is an open-source framework that adds reflection and reliability to large language model (LLM) applications.☆15Dec 4, 2025Updated 8 months ago
- Interpretable ML for TabPFN☆54Jul 13, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- PyTorch compilation tutorial covering TorchScript, torch.fx, and Slapo☆17Mar 13, 2023Updated 3 years ago
- Official code repository for the paper "Internal Activation as the Polar Star for Steering Unsafe LLM Behavior"☆15May 31, 2026Updated 2 months ago
- Pre-trained word embeddings using the text of published clinical case reports.☆14Jan 22, 2020Updated 6 years ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated 11 months ago
- Explain Neural Networks using Layer-Wise Relevance Propagation and evaluate the explanations using Pixel-Flipping and Area Under the Curv…☆17Aug 7, 2022Updated 4 years ago
- A pytorch implemention of the Explainable AI work 'Contrastive layerwise relevance propagation (CLRP)'☆17Jun 24, 2022Updated 4 years ago
- A suite of interpretability tasks to evaluate agents using Scribe for notebook access☆18Oct 2, 2025Updated 10 months ago
- Reproducible benchmarks for LLM output drift and agent replay in financial operations: DFAH-Bench, the installable dfah package, and a ha…☆18Aug 2, 2026Updated 3 weeks ago
- ☆21Jun 13, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [EMNLP 2025] Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment☆16Jul 22, 2025Updated last year
- Contrast-guided Feature Adjustment Module for Visual Information Extraction☆30May 23, 2023Updated 3 years ago
- [NeurIPS 2025]: Personalized Safety in LLMs — A Benchmark and a Planning-Based Agent Approach☆18Oct 30, 2025Updated 10 months ago
- ☆18Updated this week
- [ECCV24] Layer-Wise Relevance Propagation with Conservation Property for ResNet☆15Sep 20, 2024Updated last year
- ☆24Sep 16, 2022Updated 3 years ago
- An XAI library that helps to explain AI models in a really quick & easy way☆18Mar 8, 2024Updated 2 years ago
- Our 1st place solution to finnet challenge☆10May 29, 2020Updated 6 years ago
- Hypothesis testing (Parametric/Non-Parametric)☆12Oct 8, 2019Updated 6 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- TuneTables is a tabular classifier that implements prompt tuning for frozen prior-fitted networks.☆25Mar 31, 2025Updated last year
- B-LRP is the repository for the paper How Much Can I Trust You? — Quantifying Uncertainties in Explaining Neural Networks☆19Jun 22, 2022Updated 4 years ago
- The TABLET benchmark for evaluating instruction learning with LLMs for tabular prediction.☆25Apr 28, 2023Updated 3 years ago
- A set of methods for finding an appropriate number of topics in a text collection☆15Apr 13, 2026Updated 4 months ago
- ☆19Feb 12, 2026Updated 6 months ago
- LUNA: a Framework for Language Understanding and Naturalness Assessment.☆12Sep 9, 2023Updated 2 years ago
- offical implementation of MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming☆18Jun 2, 2025Updated last year