DataSciBench: An LLM Agent Benchmark for Data Science (Findings of ACL 2026)
☆66Jan 21, 2026Updated 7 months ago
Alternatives and similar repositories for DataSciBench
Users that are interested in DataSciBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] DSBench: How Far are Data Science Agents from Becoming Data Science Experts?☆128Aug 17, 2025Updated last year
- Reinforcing LLM Reasoning through Self-Training and Value-Guided Decoding☆19May 6, 2026Updated 4 months ago
- InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks (ICML 2024)☆207May 29, 2025Updated last year
- Official repository for paper "TableBench: A Comprehensive and Complex Benchmark for Table Question Answering"☆95May 8, 2025Updated last year
- The official repo of paper "Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller"☆18Aug 13, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Agent ADA is a comprehensive evaluation and data analytics framework focused on insights generation and skills assessment.☆15Aug 19, 2025Updated last year
- SciGLM: Training Scientific Language Models with Self-Reflective Instruction Annotation and Tuning (NeurIPS D&B Track 2024)☆89Feb 25, 2024Updated 2 years ago
- GBM implementation on Legate☆14Jul 10, 2026Updated 2 months ago
- Reproducible Language Agent Research☆36Jun 25, 2025Updated last year
- ☆74May 6, 2026Updated 4 months ago
- Q-Probe: A Lightweight Approach to Reward Maximization for Language Models☆41Jun 10, 2024Updated 2 years ago
- Continuously updated paper list on advancements in Data Agents. Companion repo to our paper "A Survey of Data Agents: Emerging Paradigm o…☆748Aug 5, 2026Updated last month
- A repository for the Kramabench benchmark☆72Jul 24, 2026Updated last month
- This repository contains the code for the paper “Neuro-Symbolic Query Compiler”, accepted to the Findings of ACL 2025.☆19Oct 20, 2025Updated 10 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆34Jun 30, 2026Updated 2 months ago
- Code and Data for "MIRAI: Evaluating LLM Agents for Event Forecasting"☆114Jul 2, 2024Updated 2 years ago
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search in ACL'25☆105Jun 16, 2025Updated last year
- Code for paper Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding☆94Jun 18, 2024Updated 2 years ago
- 一个开源数学大模型项目,旨在探索大模型是否具有数学创造能力,以及大模型在前沿数学研究中的潜在能力。☆32Mar 19, 2026Updated 6 months ago
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated last year
- Official implementation of TreeSynth: Synthesizing Diverse Data via Tree-Guided Subspace Partitioning (NeurIPS 2025 Spotlight).☆33Oct 3, 2025Updated 11 months ago
- Official implementation of ECCV24 paper: POA☆24Aug 8, 2024Updated 2 years ago
- This is the official repository for our paper "Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning" pu…☆63Apr 11, 2026Updated 5 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks☆21Jun 7, 2026Updated 3 months ago
- AutoLibra: Metric Induction for Agents from Open-Ended Human Feedback☆19Apr 23, 2026Updated 4 months ago
- A Multi-graph Multi-head Adaptive Temporal Graph Convolutional Network☆11May 21, 2023Updated 3 years ago
- THUIR website☆10Feb 23, 2026Updated 6 months ago
- [EMNLP 2025] Code for paper "Table-R1: Inference-Time Scaling for Table Reasoning"☆34Jun 3, 2025Updated last year
- ☆12Feb 27, 2025Updated last year
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- Augmenting Statistical Models with Natural Language Parameters☆28Sep 17, 2024Updated 2 years ago
- 多语言降噪预训练模型MBart的中文生成任务☆11May 27, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10May 18, 2023Updated 3 years ago
- ☆35Aug 11, 2025Updated last year
- ☆34Oct 2, 2024Updated last year
- ☆34Jun 24, 2024Updated 2 years ago
- Enhancing Large Vision Language Models with Self-Training on Image Comprehension.☆68May 31, 2024Updated 2 years ago
- Code for the examples presented in the talk "Training a Llama in your backyard: fine-tuning very large models on consumer hardware" given…☆15Oct 16, 2023Updated 2 years ago
- Automated Generation of Reproducible Test Cases for Android Apps☆12May 14, 2019Updated 7 years ago