Fiddler Auditor is a tool to evaluate language models.
β193Mar 11, 2024Updated 2 years ago
Alternatives and similar repositories for fiddler-auditor
Users that are interested in fiddler-auditor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Research notes and extra resources for all the work at explodinggradients.comβ26Mar 11, 2025Updated last year
- π LangKit: An open-source toolkit for monitoring Large Language Models (LLMs). π Extracts signals from prompts & responses, ensuring saβ¦β993Nov 22, 2024Updated last year
- Sample notebooks and prompts for LLM evaluationβ175Nov 2, 2025Updated 8 months ago
- LLM evaluation.β16Nov 7, 2023Updated 2 years ago
- Coffee Chat Voice Assistant is a voice-driven ordering system powered by Azure OpenAI GPT-4o Realtime API, simulating the experience of oβ¦β32Jun 22, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A tool for evaluating LLMsβ428Mar 15, 2026Updated 4 months ago
- Just a bunch of benchmark logs for different LLMsβ130Jul 28, 2024Updated last year
- A complete guide to evaluate LLMs and RAGs. Both theory and code based approaches covered.β28Nov 16, 2023Updated 2 years ago
- π Unstructured Data Connectors for Haystack 2.0β18Sep 21, 2023Updated 2 years ago
- Generating Realistic Synthetic Dataβ46Feb 15, 2024Updated 2 years ago
- Sample project to get started with dbt-power-user vscode extension using dev-containerβ12Apr 5, 2024Updated 2 years ago
- Deepchecks: Tests for Continuous Validation of ML Models & Data. Deepchecks is a holistic open-source solution for all of your AI & ML vaβ¦β4,039Dec 28, 2025Updated 6 months ago
- π’ Open-Source Evaluation & Testing library for LLM Agentsβ5,706Updated this week
- IBM Quantum Challenge Fall 2023β10May 23, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Oxygen is a Robot Framework tool that empowers the user to convert the results of any testing tool or framework to Robot Framework's repoβ¦β26Jun 26, 2024Updated 2 years ago
- Python SDK for running evaluations on LLM generated responsesβ301Jun 6, 2025Updated last year
- A library for red-teaming LLM applications with LLMs.β28Oct 11, 2024Updated last year
- LLM Prompt Injection Detectorβ1,514Aug 7, 2024Updated last year
- A content inspecting SMTP proxyβ17Jun 9, 2014Updated 12 years ago
- β13Apr 6, 2025Updated last year
- Evidently is ββan open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. Froβ¦β7,744May 2, 2026Updated 2 months ago
- Mint and purchase NFTs representing time for performing freelance services and other use cases. Time NFTs can represent an on-chain attesβ¦β42Nov 28, 2022Updated 3 years ago
- Awesome material(papers, tools, etc.) about testing machine learning system, including deep learning system.β47Oct 12, 2021Updated 4 years ago
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- This sample shows how to use a Cosmos DB Trigger in Azure Functions Triggers (C# or Python) to automatically generate embeddings on data β¦β21Mar 11, 2025Updated last year
- NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.β6,778Updated this week
- This project involves using llamaindex Multi Agents concierge system and Qdrant vector database to customize the RAG application with useβ¦β56Aug 20, 2024Updated last year
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMsβ337Jun 7, 2024Updated 2 years ago
- A novel BCI-based messaging web application that uses P300 speller and emotion detection to bridge the gap in social media accessibility β¦β12Aug 31, 2021Updated 4 years ago
- Metrics to evaluate the quality of responses of your Retrieval Augmented Generation (RAG) applications.β327Jul 10, 2025Updated last year
- AI Observability & Evaluationβ10,699Updated this week
- The Security Toolkit for LLM Interactionsβ3,191Jul 8, 2026Updated 2 weeks ago
- Neo4j Cybersecurity Demoβ19Mar 16, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Deepmark AI enables a unique testing environment for language models (LLM) assessment on task-specific metrics and on your own data so yoβ¦β104Nov 24, 2023Updated 2 years ago
- This project provides a solution to AWS customers for reporting on what tags exists, the resources they are applied to, and what resourceβ¦β25Feb 28, 2024Updated 2 years ago
- ToolBench, an evaluation suite for LLM tool manipulation capabilities.β180Feb 28, 2024Updated 2 years ago
- Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasetsβ5,046Updated this week
- Adding guardrails to large language models.β7,194Updated this week
- [ECCV 2024] M3DBench introduces a comprehensive 3D instruction-following dataset with support for interleaved multi-modal prompts.β61Oct 1, 2024Updated last year
- [ICLR 2025 Spotlight] An open-sourced LLM judge for evaluating LLM-generated answers.β438Feb 11, 2025Updated last year