m&ms: A Benchmark to Evaluate Tool-Use for multi-step multi-modal tasks
☆46Sep 26, 2024Updated last year
Alternatives and similar repositories for mnms
Users that are interested in mnms are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR23 Highlight] CREPE: Can Vision-Language Foundation Models Reason Compositionally?☆36Apr 27, 2023Updated 3 years ago
- A instruction data generation system for multimodal language models.☆37Jan 31, 2025Updated last year
- This is the repository for VFig: Vectorizing Complex Figures with Vision-Language Models☆19Apr 24, 2026Updated 3 months ago
- ☆70Jun 2, 2026Updated 2 months ago
- [NeurIPS 2023] A faithful benchmark for vision-language compositionality☆95Feb 13, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆38Aug 25, 2025Updated 11 months ago
- [ACL2024] Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios☆71Aug 5, 2025Updated last year
- EcoAssistant: using LLM assistant more affordably and accurately☆132Jun 30, 2024Updated 2 years ago
- ☆36Feb 5, 2024Updated 2 years ago
- [ACL 2024] On the Multi-turn Instruction Following for Conversational Web Agents☆16Oct 12, 2024Updated last year
- ☆16Apr 10, 2025Updated last year
- [ICLR'24] MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use☆116Mar 21, 2024Updated 2 years ago
- ☆29Apr 30, 2024Updated 2 years ago
- VOCAL-UDF: Self-Enhancing Video Data Management System for Compositional Events with Large Language Models☆13Dec 12, 2025Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Using conversational games to evaluate powerful LLMs☆18Sep 3, 2023Updated 2 years ago
- DeepEverest: a system for efficient DNN interpretation.☆13Jan 22, 2024Updated 2 years ago
- A Python toolkit for the OmniLabel benchmark providing code for evaluation and visualization☆23Feb 1, 2025Updated last year
- ☆21Oct 10, 2023Updated 2 years ago
- [NeurIPS 2021] WRENCH: Weak supeRvision bENCHmark☆231Feb 13, 2024Updated 2 years ago
- ☆59Aug 30, 2023Updated 2 years ago
- Official implementation of "Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data" (ICLR 2024)☆36Oct 16, 2024Updated last year
- Evaluating tool-augmented LLMs in conversation settings☆89May 31, 2024Updated 2 years ago
- PyTorch code for Improving Commonsense in Vision-Language Models via Knowledge Graph Riddles (DANCE)☆22Nov 29, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 23] Contrastive Aligned of Vision to Language Through Parameter-Efficient Transfer Learning☆40Jul 29, 2023Updated 3 years ago
- MIRIS: Fast Object Track Queries in Video☆17Mar 24, 2023Updated 3 years ago
- Scene Graph Prediction with Limited Labels☆54Oct 3, 2023Updated 2 years ago
- Open-source repository for the OOPSLA'24 paper "CYCLE: Learning to Self-Refine Code Generation"☆10Mar 8, 2024Updated 2 years ago
- Python codes for mathematical modeling.☆13Sep 5, 2021Updated 4 years ago
- ☆12Aug 8, 2024Updated 2 years ago
- [IROS 2026] Implementation of FailSafe Pipeline in Maniskill Simulator☆24Updated this week
- A makeshift python program which relies on nltk and Stanford Core NLP models to expand common contractions in the english language.☆10Nov 8, 2017Updated 8 years ago
- Nemo debugs Distributed Systems by analyzing provenance graphs obtained during fault injection.☆19Dec 10, 2019Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆17Oct 8, 2019Updated 6 years ago
- NLPBench: Evaluating NLP-Related Problem-solving Ability in Large Language Models☆10Oct 27, 2023Updated 2 years ago
- end-to-end dialog system dataset☆13Sep 15, 2019Updated 6 years ago
- The code for NeurIPS 2020 paper: Adversarial Crowdsourcing Through Robust Rank-One Matrix Completion.☆10Oct 26, 2020Updated 5 years ago
- [BMVC 2022] Information Theoretic Representation Distillation☆19Oct 6, 2023Updated 2 years ago
- List of NLP Datasets☆10Mar 12, 2019Updated 7 years ago
- codes for "Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models"☆12Feb 10, 2025Updated last year