New testbed of interactive SWE tasks for coding agents, set in a realistic multi-turn developer driven environment
☆24Jun 30, 2026Updated last month
Alternatives and similar repositories for SWE-Interact
Users that are interested in SWE-Interact are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆51Jul 7, 2026Updated 3 weeks ago
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space☆17Apr 16, 2026Updated 3 months ago
- ☆31Jun 2, 2026Updated 2 months ago
- [EMNLP 2023] Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-Thoughts☆27Nov 4, 2023Updated 2 years ago
- Toolathlon-Gym for testing AI agents real-world tool-use capabilities across diverse MCP servers.☆142Jul 22, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LLM 时代的 Hot 100 - 大模型面试手撕代码☆30May 4, 2026Updated 3 months ago
- SimKO: Simple Pass@K Policy Optimization☆31Oct 24, 2025Updated 9 months ago
- ☆12Feb 16, 2024Updated 2 years ago
- The repository of CLEME (EMNLP 2023) and CLEME2.0 (ACL 2025)☆12May 17, 2025Updated last year
- Data and code for EACL'24 paper: Over-Reasoning and Redundant Calculation of Large Language Models☆10Jan 23, 2024Updated 2 years ago
- Transfer Learning in Dialogue Benchmarking Toolkit☆14Mar 31, 2023Updated 3 years ago
- ☆13May 18, 2022Updated 4 years ago
- ☆14Jul 17, 2025Updated last year
- Convert GitHub PRs into Harbor tasks☆73Jul 13, 2026Updated 3 weeks ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures☆34Jan 29, 2026Updated 6 months ago
- ☆18Sep 3, 2024Updated last year
- Code Repository for "A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language Models".☆15Oct 14, 2022Updated 3 years ago
- The repo for using the model https://huggingface.co/thu-coai/Attacker-v0.1☆13Apr 23, 2025Updated last year
- This repository open-sources our GEC system submitted by THU KELab (sz) in the CCL2023-CLTC Track 1: Multidimensional Chinese Learner Tex…☆15Nov 25, 2023Updated 2 years ago
- Codes for "EDG-based Question Decomposition for Complex Question Answering over Knowledge Bases"☆13Nov 12, 2021Updated 4 years ago
- [PR 2021] Code for "GraphAIR: Graph Representation Learning with Neighborhood Aggregation and Interaction"☆12Aug 25, 2021Updated 4 years ago
- Towards Real-World Writing Assistance: A Chinese Character Checking Benchmark with Faked and Misspelled Characters☆17May 30, 2024Updated 2 years ago
- ☆11Mar 3, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆16Apr 11, 2022Updated 4 years ago
- Parses, Analyzes and Predicts for the Korean Baseball League☆17Dec 8, 2022Updated 3 years ago
- [NeurIPS 2024 poster] Cross-model Control: Improving Multiple Large Language Models in One-time Training☆14Oct 25, 2024Updated last year
- Code for Navigating Connected Memories with a Task-oriented Dialog System☆18Dec 12, 2022Updated 3 years ago
- [ICLR 2024] COLLIE: Systematic Construction of Constrained Text Generation Tasks☆64Aug 2, 2023Updated 3 years ago
- A Long Short Term Memory neural network for time series prediction. Memory blocks contain one memory cell in each. Weights for the networ…☆15Sep 3, 2018Updated 7 years ago
- ☆13Nov 7, 2023Updated 2 years ago
- A supervised fine-tuning method for controllable reasoning length in large language models (一种通过有监督微调实现大语言模型思考长度可控的方法)☆11May 8, 2025Updated last year
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆449Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆116Jul 2, 2024Updated 2 years ago
- My 1st place solution to the Kaggle Invasive Species Monitoring Competition☆10Aug 17, 2017Updated 8 years ago
- Code for CIKM 2018 Paper "PRRE: Personalized Relation Ranking for Attributed Network Embedding"☆16Jul 6, 2023Updated 3 years ago
- ☆12Jul 10, 2023Updated 3 years ago
- Code for "End-to-End Learning of Flowchart Grounded Task-Oriented Dialogs"☆14Oct 10, 2022Updated 3 years ago
- The project page for "SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables"☆23Dec 21, 2023Updated 2 years ago
- MARL: the model of the IJCAI 2020 paper 'Retrieve, Program, Repeat: Complex Knowledge Base Question Answering via Alternate Meta-learning…☆13Oct 8, 2020Updated 5 years ago