[EMNLP 2024] RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
☆15May 13, 2025Updated last year
Alternatives and similar repositories for RoTBench
Users that are interested in RoTBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2024] ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages☆15Sep 12, 2024Updated 2 years ago
- [COLING 2025] ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios☆74May 13, 2025Updated last year
- [ACL 2026] A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models☆23Jul 10, 2026Updated 2 months ago
- This repository provides open-source code for sparse continuous distributions and corresponding Fenchel-Young losses.☆15May 10, 2023Updated 3 years ago
- This is an agent (including contextual prompts) that queries your CSV☆10Jun 8, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- PACIFIC: Towards Proactive Conversational Question Answering over Tabular and Textual Data in Finance☆14May 15, 2024Updated 2 years ago
- ☆25Jul 15, 2023Updated 3 years ago
- Repo for our AKBC-2021 paper: Abg-CoQA: Clarifying Ambiguity in Conversational Question Answering☆11Oct 10, 2021Updated 4 years ago
- Papers about the trend of Entity Linking in recent years.☆11Sep 5, 2022Updated 4 years ago
- ☆13Jul 13, 2018Updated 8 years ago
- Code for 'Contrastive Multi-Document Question Generation'☆11Oct 16, 2022Updated 3 years ago
- Code for the paper: Improving Multi-Document Summarization through Referenced Flexible Extraction with Credit-Awareness☆12Oct 22, 2023Updated 2 years ago
- DICE: Detecting In-distribution Data Contamination with LLM's Internal State☆11Sep 21, 2024Updated 2 years ago
- [ACL 2026] Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments☆52Jul 10, 2026Updated 2 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- IPython magic for Telegram notifications☆15Oct 12, 2021Updated 4 years ago
- Official implementation for "ALI-Agent: Assessing LLMs'Alignment with Human Values via Agent-based Evaluation"☆21Jan 31, 2026Updated 7 months ago
- Completing the Puzzle of All-in-One Event Understanding Benchmark with Event Arguments☆16Mar 12, 2024Updated 2 years ago
- [ICLR 2025] Permute-and-Flip: An optimally robust and watermarkable decoder for LLMs☆19Mar 20, 2025Updated last year
- This repository contains data and code used for On the Risk of Misinformation Pollution with Large Language Models (EMNLP 2023 Findings).☆17Dec 14, 2023Updated 2 years ago
- [NAACL 2024 Findings] Evaluation suite for the systematic evaluation of instruction selection methods.☆24Jul 26, 2023Updated 3 years ago
- ☆23Sep 17, 2024Updated 2 years ago
- Source code for EMNLP2022 paper "Finding Skill Neurons in Pre-trained Transformers via Prompt Tuning".☆19Mar 13, 2023Updated 3 years ago
- [ICML 2024] Self-Infilling Code Generation☆18May 5, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [EMNLP2025] Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling☆16Nov 20, 2025Updated 10 months ago
- ☆12Mar 18, 2021Updated 5 years ago
- ☆16Jul 9, 2025Updated last year
- 🌿 DeepPrune: Parallel Scaling without Inter-trace Redundancy☆21Apr 20, 2026Updated 5 months ago
- This is a repo listing some must-read papers on *AI-driven MOOCs* or *Intelligent Education* published in recent years, mainly contribute…☆18Jun 8, 2022Updated 4 years ago
- ☆21Nov 19, 2025Updated 10 months ago
- Let's Qitos! A torch-like agent-native framework for researchers.☆43Sep 9, 2026Updated 2 weeks ago
- ☆18Jun 5, 2024Updated 2 years ago
- ☆17Sep 11, 2026Updated 2 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [ICLR 25 Spotlight] A testbed for agents and environments that can automatically improve models through data generation.☆35Mar 4, 2025Updated last year
- Pytorch implementation of 'Semi-Implicit Methods for Deep Neural Networks'☆25May 13, 2019Updated 7 years ago
- Repository of paper "Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis" (ACL 2025 Main)☆20Jul 19, 2025Updated last year
- 3-day dive into deep learning at csc☆24Feb 3, 2017Updated 9 years ago
- A repository of OpenDecoder framework: Open Large Language Model Decoding to Incorporate Document Quality in RAG (WWW 2026)☆27Jan 27, 2026Updated 8 months ago
- 北航教务辅助脚本工具集☆24Sep 27, 2022Updated 4 years ago
- Code for the paper "Aligning LLM Agents by Learning Latent Preference from User Edits".☆46Nov 23, 2024Updated last year