[ICLR 2025] DSBench: How Far are Data Science Agents from Becoming Data Science Experts?
☆128Aug 17, 2025Updated last year
Alternatives and similar repositories for DSBench
Users that are interested in DSBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DataSciBench: An LLM Agent Benchmark for Data Science (Findings of ACL 2026)☆66Jan 21, 2026Updated 8 months ago
- [EMNLP 2024 Findings] Benchmarking Language Model Agents for Data-Driven Science☆38Oct 25, 2024Updated last year
- ☆20Nov 28, 2024Updated last year
- InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks (ICML 2024)☆207May 29, 2025Updated last year
- This is the codes of "DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrieval"☆15Aug 11, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- Official repository for paper "TableBench: A Comprehensive and Complex Benchmark for Table Question Answering"☆95May 8, 2025Updated last year
- This repository contains the code for the paper “Neuro-Symbolic Query Compiler”, accepted to the Findings of ACL 2025.☆19Oct 20, 2025Updated 11 months ago
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification☆23Oct 8, 2025Updated 11 months ago
- XmodelLM☆38Nov 19, 2024Updated last year
- ☆34Jun 24, 2024Updated 2 years ago
- [ICLR/AAAI/KDD/EMNLP2026] Open-Source LLM-Based Data Analysis Agents☆142Sep 13, 2026Updated last week
- ☆354Jun 19, 2024Updated 2 years ago
- [NeurIPS 2024 D&B Track] DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation☆14Mar 5, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering☆1,747Apr 24, 2026Updated 4 months ago
- A repository for the Kramabench benchmark☆73Jul 24, 2026Updated last month
- ☆107Oct 30, 2025Updated 10 months ago
- ☆15Jan 27, 2025Updated last year
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated last year
- ☆79Nov 23, 2025Updated 9 months ago
- Run TPC-DS against different databases including Hive, Spark SQL and IBM BigSQL☆14Jan 4, 2022Updated 4 years ago
- ☆16May 18, 2026Updated 4 months ago
- ☆35Jan 12, 2026Updated 8 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆317Dec 4, 2024Updated last year
- FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models☆35Nov 27, 2025Updated 9 months ago
- Llemma formal2formal (tactic prediction) theorem proving experiments☆20Oct 17, 2023Updated 2 years ago
- Continuously updated paper list on advancements in Data Agents. Companion repo to our paper "A Survey of Data Agents: Emerging Paradigm o…☆752Aug 5, 2026Updated last month
- Official Code Repository for paper "HYDRA: Model Factorization Framework for Black-Box LLM Personalization"☆16Oct 7, 2024Updated last year
- ☆16Aug 31, 2023Updated 3 years ago
- [AAAI'25] SPRING: Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models☆26Sep 24, 2025Updated 11 months ago
- NestJS project template, configured with prisma and ejs☆12Dec 1, 2024Updated last year
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆74Jan 28, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code and data for "Medical Dialogue Generation via Dual Flow Modeling" (ACL 2023 Findings)☆14Nov 22, 2023Updated 2 years ago
- Some example codes for drawing figures in research paper☆36Mar 3, 2022Updated 4 years ago
- This is the reading list of Large Language Model-Based Data Science Agent☆42Nov 3, 2025Updated 10 months ago
- Applying Deep Reinforcement Learning for dialogue generation. aka chatbot☆13Apr 30, 2017Updated 9 years ago
- DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use☆30Mar 13, 2026Updated 6 months ago
- ☆48Sep 8, 2025Updated last year
- HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches☆41Oct 9, 2025Updated 11 months ago