New testbed of interactive SWE tasks for coding agents, set in a realistic multi-turn developer driven environment
☆24Jun 30, 2026Updated last month
Alternatives and similar repositories for SWE-Interact
Users that are interested in SWE-Interact are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions☆56Jul 7, 2026Updated last month
- [arxiv: 2604.14142] From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space☆17Apr 16, 2026Updated 4 months ago
- ☆17Nov 1, 2024Updated last year
- ☆14Jul 19, 2020Updated 6 years ago
- [NeurIPS 2025] Reasoning Models Better Express Their Confidence"☆23Nov 19, 2025Updated 9 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [EMNLP 2023] Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-Thoughts☆28Nov 4, 2023Updated 2 years ago
- [ICSE 2026] [ACM SIGSOFT Distinguished Paper Award] SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Re…☆21Apr 22, 2025Updated last year
- Toolathlon-Gym for testing AI agents real-world tool-use capabilities across diverse MCP servers.☆145Jul 22, 2026Updated last month
- LLM 时代的 Hot 100 - 大模型面试手撕代码☆40May 4, 2026Updated 3 months ago
- SimKO: Simple Pass@K Policy Optimization☆31Oct 24, 2025Updated 9 months ago
- The repository of CLEME (EMNLP 2023) and CLEME2.0 (ACL 2025)☆12May 17, 2025Updated last year
- Data and code for EACL'24 paper: Over-Reasoning and Redundant Calculation of Large Language Models☆10Jan 23, 2024Updated 2 years ago
- ☆12Mar 6, 2026Updated 5 months ago
- Transfer Learning in Dialogue Benchmarking Toolkit☆14Mar 31, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆13May 18, 2022Updated 4 years ago
- An attempt to apply reinforcement learning to graph signal recovery problem☆11Aug 25, 2021Updated 4 years ago
- ☆25Aug 2, 2025Updated last year
- Benchmarking Social Intelligence of Language Agents through Interactive Scenarios☆13Jan 4, 2025Updated last year
- Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures☆35Jan 29, 2026Updated 6 months ago
- ☆18Sep 3, 2024Updated last year
- Data and codes for EMNLP 2022 paper "CDConv: A Benchmark for Contradiction Detection in Chinese Conversations"☆13May 8, 2023Updated 3 years ago
- The repo for using the model https://huggingface.co/thu-coai/Attacker-v0.1☆13Apr 23, 2025Updated last year
- LLMatic is a 2-archive QD algorithm that uses LLMs to mutate the networks. Tested for Neural Architecture search but can easily be used f…☆21Aug 14, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [PR 2021] Code for "GraphAIR: Graph Representation Learning with Neighborhood Aggregation and Interaction"☆12Aug 25, 2021Updated 4 years ago
- ☆11Mar 3, 2026Updated 5 months ago
- ☆16Apr 11, 2022Updated 4 years ago
- GCN-LASE: Towards Adequately Incorporating Link Attributes in Graph Convolutional Networks☆13Aug 13, 2021Updated 5 years ago
- [NeurIPS 2024 poster] Cross-model Control: Improving Multiple Large Language Models in One-time Training☆15Oct 25, 2024Updated last year
- Investigating Cultural Alignment of Large Language Models☆13Aug 14, 2024Updated 2 years ago
- Code for Navigating Connected Memories with a Task-oriented Dialog System☆18Dec 12, 2022Updated 3 years ago
- Malware Classification using Graph Clustering☆14Nov 12, 2012Updated 13 years ago
- Code for the paper "Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching" (COLING 2025)☆20May 27, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICLR 2024] COLLIE: Systematic Construction of Constrained Text Generation Tasks☆64Aug 2, 2023Updated 3 years ago
- a jax benchmark for ad hoc teamwork☆23Updated this week
- ☆13Nov 7, 2023Updated 2 years ago
- A supervised fine-tuning method for controllable reasoning length in large language models (一种通过有监督微调实现大语言模型思考长度可控的方法)☆11May 8, 2025Updated last year
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆461Updated this week
- ☆117Jul 2, 2024Updated 2 years ago
- Code for CIKM 2018 Paper "PRRE: Personalized Relation Ranking for Attributed Network Embedding"☆16Jul 6, 2023Updated 3 years ago