Compare chatbots pairwise via multi‑round evaluations for SE tasks.
☆15Apr 4, 2026Updated 3 months ago
Alternatives and similar repositories for SWE-Chatbot-Arena
Users that are interested in SWE-Chatbot-Arena are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A self-evolving coding agent in Rust: the smallest possible implementation that actually works.☆16Jul 10, 2026Updated last week
- ☆15May 7, 2026Updated 2 months ago
- AutoResearch official style beginner tutorial, from 0 to 1☆18May 8, 2026Updated 2 months ago
- ☆12Jun 19, 2026Updated last month
- A curated list of awesome interactive theorem prover frameworks☆23Jun 19, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- How to Start a Startup — AI Agent Skill☆25Apr 17, 2026Updated 3 months ago
- ☆11Oct 15, 2022Updated 3 years ago
- You install it. Claude drives the CLI tool. At first it might seem like too much. Eventually, nothing less will make sense.☆26Jun 2, 2026Updated last month
- PocketFlow from 0 to 1 | 100 行代码构建所有 LLM 应用 | 首个 PocketFlow 交互式教程 | 光速掌握智能体开发实战☆33May 11, 2026Updated 2 months ago
- A curated list of awesome autonomous researcher frameworks☆141Updated this week
- An autonomous research agent that turns a topic into a peer-reviewed technical report☆120Jun 3, 2026Updated last month
- A curated list of awesome open source libraries to deploy, monitor, version and scale agentic applications and systems☆154Jul 1, 2026Updated 2 weeks ago
- A curated list of awesome curated lists of many topics related to artificial intelligence.☆224Jul 9, 2026Updated last week
- Evaluation tools for Retrieval-augmented Generation (RAG) methods.☆171Nov 18, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Program Transformation Tool for Java Methods☆10Sep 16, 2022Updated 3 years ago
- MODIT: On Multi-Modal Learning of Editing Source Code.☆20Apr 24, 2021Updated 5 years ago
- An official implementation of "Catastrophic Failure of LLM Unlearning via Quantization" (ICLR 2025)☆39Feb 22, 2025Updated last year
- SZZ Algorithm To Detect Fault-Inducing Commits☆51Oct 24, 2023Updated 2 years ago
- A BPMN.js extension to improve working with SpiffWorkflow - the python BPMN engine.☆31Jul 14, 2026Updated last week
- CoditT5: Pretraining for Source Code and Natural Language Editing☆29Jan 16, 2025Updated last year
- IST'21 & SANER'22: Semantic-Preserving Program Transformations☆31Oct 25, 2022Updated 3 years ago
- ☆41Jan 13, 2023Updated 3 years ago
- Code for paper: "Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines"☆11Oct 11, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 《机器学习理论导引》(宝箱书)的证明、案例、概念补充与参考文献讲解。☆1,710Jul 7, 2026Updated 2 weeks ago
- A Tool for Detecting Performance Bugs in Rails Applications☆60Jul 3, 2018Updated 8 years ago
- LLM benchmarks☆13Feb 22, 2024Updated 2 years ago
- 中文金融大模型测评基准,六大类二十五任务、等级化评价,国内模型获得A级☆10May 6, 2024Updated 2 years ago
- Knowledge Graph based Question Answering benchmark.☆10Feb 1, 2020Updated 6 years ago
- ☆49Nov 19, 2025Updated 8 months ago
- Code and data for automatic paraphrase dataset augmentation.☆11Mar 8, 2021Updated 5 years ago
- ☆11Jan 3, 2024Updated 2 years ago
- Align, a general text alignment function☆15Dec 7, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- You Can Set Up Kimi K2 & Launch 80% Cheaper Full Stack | Claude Code, NextJS, Supabase, Moonshot.ai☆37Sep 23, 2025Updated 9 months ago
- [CVPR2024] Learning from Synthetic Human Group Activities☆14Feb 24, 2025Updated last year
- LGEB: Benchmark of Language Generation Evaluation☆16Oct 21, 2022Updated 3 years ago
- Website for release of TellMeWhy dataset for why question answering☆14Nov 11, 2022Updated 3 years ago
- Code for our project CROWN (Conversational Passage Ranking by Reasoning over Word Networks)☆10Jan 11, 2024Updated 2 years ago
- [LREC-Coling 2024] PECC: Problem Extraction and Coding Challenges☆14May 30, 2024Updated 2 years ago
- Math-aware QA system☆18May 8, 2026Updated 2 months ago