[EMNLP 2024] A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models.
☆25Sep 23, 2024Updated last year
Alternatives and similar repositories for ToolBeHonest
Users that are interested in ToolBeHonest are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [2025-TMLR] A Survey on the Honesty of Large Language Models☆66Dec 8, 2024Updated last year
- ☆21Aug 19, 2024Updated 2 years ago
- Implementation of LREC-COLING 2024 paper A Frustratingly Simple Decoding Method for Neural Text Generation☆19Feb 23, 2024Updated 2 years ago
- [ICLR 2025] ChartMimic: Evaluating LMM’s Cross-Modal Reasoning Capability via Chart-to-Code Generation☆134Dec 19, 2025Updated 8 months ago
- ☆29May 24, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- RePO: Replay-Enhanced Policy Optimization☆24Jun 12, 2025Updated last year
- [ICCV 2023 Workshop] The Official Implementation of The First Prize Solution for RVOS Competition☆14Jan 1, 2024Updated 2 years ago
- [ACL 2023] Solving Math Word Problems via Cooperative Reasoning induced Language Models (LLMs + MCTS + Self-Improvement)☆51Dec 15, 2023Updated 2 years ago
- ☆21Nov 26, 2024Updated last year
- Source code for Truth-Aware Context Selection: Mitigating the Hallucinations of Large Language Models Being Misled by Untruthful Contexts☆17Sep 2, 2024Updated last year
- I-SHEEP: Iterative Self-enHancEmEnt Paradigm of LLMs through Self-Instruct and Self-Assessment☆17Jan 16, 2025Updated last year
- SPUQ: Perturbation-Based Uncertainty Quantification for Large Language Models☆17Jun 24, 2024Updated 2 years ago
- Safety-J: Evaluating Safety with Critique☆16Jul 28, 2024Updated 2 years ago
- Codebase for the EMNLP 2021 paper "HittER: Hierarchical Transformers for Knowledge Graph Embeddings".☆12Nov 1, 2021Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Codebase of 'From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model'☆45Jun 27, 2026Updated 2 months ago
- This is small practice to location and detect the fog exponent.☆13May 9, 2021Updated 5 years ago
- Python client library for Cleanlab Trustworthy Language Model☆24Dec 9, 2025Updated 8 months ago
- A simple pytorch implementation of baseline based-on CLIP for Image-text Matching.☆19May 25, 2023Updated 3 years ago
- Code for paper "Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication"☆23Mar 30, 2024Updated 2 years ago
- Implementation for "RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content"☆24Jul 28, 2024Updated 2 years ago
- Fork of https://github.com/getAsterisk/deepclaude, with added features.☆14Feb 25, 2025Updated last year
- ☆26Feb 23, 2026Updated 6 months ago
- [NeurIPS 2024 poster] Cross-model Control: Improving Multiple Large Language Models in One-time Training☆15Oct 25, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for the paper "HALoGEN: Fantastic LLM Hallucinations and Where To Find Them"☆26May 18, 2025Updated last year
- [ACL 2024] ANAH & [NeurIPS 2024] ANAH-v2 & [ICLR 2025] Mask-DPO☆67Apr 30, 2025Updated last year
- This is the repository containing the solution of the homework for the CS224W course at Stanford: Machine Learning with Graphs☆11Jul 19, 2020Updated 6 years ago
- ☆18Jul 21, 2026Updated last month
- Very concise example of integrated gradients (a method to reveal areas of attention in input images)☆10Jun 17, 2019Updated 7 years ago
- [NeurIPS'25] Official Implementation of RISE (Reinforcing Reasoning with Self-Verification)☆33Aug 8, 2025Updated last year
- [ICML 2026] Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO☆19Jun 15, 2026Updated 2 months ago
- Chinese Generation Evaluation☆13Aug 14, 2023Updated 3 years ago
- Official Implementation of "DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucination"☆30Dec 18, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- a jekyll powered blog theme specially designed to take notes, not just blogs☆12Nov 5, 2023Updated 2 years ago
- ☆43Feb 2, 2024Updated 2 years ago
- Official repository for paper "DeepCritic: Deliberate Critique with Large Language Models"☆41Jun 24, 2025Updated last year
- [ACL 2024] Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models. Detect and mitigate object hallucinatio…☆25Jan 31, 2025Updated last year
- ☆31Oct 20, 2025Updated 10 months ago
- ☆47Jul 10, 2024Updated 2 years ago
- ☆18Mar 19, 2023Updated 3 years ago