DataSciBench: An LLM Agent Benchmark for Data Science (Findings of ACL 2026)
☆67Jan 21, 2026Updated 8 months ago
Alternatives and similar repositories for DataSciBench
Users that are interested in DataSciBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] DSBench: How Far are Data Science Agents from Becoming Data Science Experts?☆130Aug 17, 2025Updated last year
- Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding☆19May 6, 2026Updated 5 months ago
- InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks (ICML 2024)☆207May 29, 2025Updated last year
- Official repository for paper "TableBench: A Comprehensive and Complex Benchmark for Table Question Answering"☆97May 8, 2025Updated last year
- The official repo of paper "Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller"☆18Aug 13, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆16May 18, 2026Updated 4 months ago
- Agent ADA is a comprehensive evaluation and data analytics framework focused on insights generation and skills assessment.☆15Aug 19, 2025Updated last year
- SciGLM: Training Scientific Language Models with Self-Reflective Instruction Annotation and Tuning (NeurIPS D&B Track 2024)☆89Feb 25, 2024Updated 2 years ago
- ☆55Aug 24, 2025Updated last year
- GBM implementation on Legate☆13Jul 10, 2026Updated 2 months ago
- Reproducible Language Agent Research☆36Jun 25, 2025Updated last year
- Q-Probe: A Lightweight Approach to Reward Maximization for Language Models☆41Jun 10, 2024Updated 2 years ago
- Continuously updated paper list on advancements in Data Agents. Companion repo to our paper "A Survey of Data Agents: Emerging Paradigm o…☆771Aug 5, 2026Updated 2 months ago
- ☆34Sep 28, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A repository for the Kramabench benchmark☆75Jul 24, 2026Updated 2 months ago
- This repository contains the code for the paper “Neuro-Symbolic Query Compiler”, accepted to the Findings of ACL 2025.☆19Oct 20, 2025Updated 11 months ago
- ☆34Jun 30, 2026Updated 3 months ago
- Code and Data for "MIRAI: Evaluating LLM Agents for Event Forecasting"☆113Jul 2, 2024Updated 2 years ago
- Code for paper Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding☆94Jun 18, 2024Updated 2 years ago
- 一个开源数学大模型项目,旨在探索大模型是否具有数学创造能力,以及大模型在前沿数学研究中的潜在能力。☆33Mar 19, 2026Updated 6 months ago
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated 2 years ago
- Overflow Prevention Enhances Long-Context Recurrent LLMs (COLM 2025)☆18Jul 8, 2025Updated last year
- Official implementation of TreeSynth: Synthesizing Diverse Data via Tree-Guided Subspace Partitioning (NeurIPS 2025 Spotlight).☆33Oct 3, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Plancraft is a minecraft environment and agent suite to test planning capabilities in LLMs☆31Nov 7, 2025Updated 11 months ago
- [EMNLP 2025] TongSearch-QR☆44Dec 4, 2025Updated 10 months ago
- Official implementation of ECCV24 paper: POA☆24Aug 8, 2024Updated 2 years ago
- This is the official repository for our paper "Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning" pu…☆67Apr 11, 2026Updated 5 months ago
- TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks☆22Jun 7, 2026Updated 4 months ago
- AutoLibra: Metric Induction for Agents from Open-Ended Human Feedback☆19Apr 23, 2026Updated 5 months ago
- A Multi-graph Multi-head Adaptive Temporal Graph Convolutional Network☆11Oct 1, 2026Updated last week
- THUIR website☆10Feb 23, 2026Updated 7 months ago
- [EMNLP 2025] Code for paper "Table-R1: Inference-Time Scaling for Table Reasoning"☆34Jun 3, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆12Feb 27, 2025Updated last year
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- Augmenting Statistical Models with Natural Language Parameters☆28Sep 17, 2024Updated 2 years ago
- 多语言降噪预训练模型MBart的中文生成任务☆11May 27, 2021Updated 5 years ago
- ☆10Mar 13, 2023Updated 3 years ago
- ☆13Jun 14, 2022Updated 4 years ago
- ☆35Aug 11, 2025Updated last year