LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
☆22Jun 2, 2026Updated 2 months ago
Alternatives and similar repositories for LoCoBench-Agent
Users that are interested in LoCoBench-Agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆46Jun 2, 2026Updated 2 months ago
- [CHIL 2024] Interpretation of Intracardiac Electrograms Through Textual Representations☆12Sep 4, 2024Updated last year
- [ECCV 2022] "Improve Few-Shot Transfer Learning with Low-Rank Decompose and Align" by Ziyu Jiang, Tianlong Chen, Xuxi Chen, Yu Cheng, Luo…☆13Jul 19, 2022Updated 4 years ago
- ☆10Sep 24, 2019Updated 6 years ago
- [ICML 2026] Sparser Block-Sparse Attention via Token Permutation☆31May 22, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Functional Optimal Transport: Map Estimation and Domain Adaptation for Functional data☆28Jun 7, 2021Updated 5 years ago
- [EACL 2023] Transfer Knowledge from Natural Language to Electrocardiography: Can We Detect Cardiovascular Disease Through Language Models…☆18May 7, 2024Updated 2 years ago
- ☆13May 27, 2019Updated 7 years ago
- Ranking LLM-Generated Loop Invariants for Program Verification.☆13Aug 20, 2024Updated last year
- [COLM 2024] SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models☆24Oct 5, 2024Updated last year
- IEEE S&P 2023 - DEVFUZZ: Automatic Device Model-Guided Device Driver Fuzzing☆13Dec 16, 2024Updated last year
- Papers on concurrency vulnerability analysis, including multithreaded programs, multi-tasking programs and interrupt driven programs.☆14Nov 11, 2022Updated 3 years ago
- [NIPS 2025 DB Spotlight] AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios☆40Dec 1, 2025Updated 8 months ago
- Code for Deep learning models for electrocardiograms are susceptible to adversarial attack☆24Feb 4, 2021Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆11Sep 29, 2022Updated 3 years ago
- [EMNLP 2023] An Empirical Exploration of Cross-domain Alignment between Language and Electroencephalogram☆31Nov 9, 2023Updated 2 years ago
- Sotopia-RL: Reward Design for Social Intelligence☆52Apr 1, 2026Updated 4 months ago
- Detecting Concurrency Memory Corruption Vulnerabilities (ESEC/FSE 2019)☆15Dec 5, 2023Updated 2 years ago
- [EMNLP 2025🔥] UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective☆20Jan 7, 2026Updated 6 months ago
- Neural Adversarial Agent Mutation-based Security Evaluator☆17Apr 18, 2026Updated 3 months ago
- [DMLR 2024] Benchmarking Robustness of Multimodal Image-Text Models under Distribution Shift☆39Jan 25, 2024Updated 2 years ago
- RTFM! Automatic Assumption Discovery and VerificationDerivation from Library Document for API Misuse Detection☆19Oct 5, 2021Updated 4 years ago
- Benchmarking Language Agents Under Controllable and Extreme Context Growth☆51Apr 29, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for ICML2020 "Sequence Generation with Mixed Representations"☆12Jun 27, 2020Updated 6 years ago
- Evaluating Visual Fidelity of Image Descriptions☆11Aug 15, 2019Updated 6 years ago
- Convert pretrained RoBerta models to various long-document transformer models☆11Apr 5, 2022Updated 4 years ago
- The first blockchain to support both the EVM and Ewasm virtual machines for current and next-gen Ethereum. Based on the CyberMiles blockc…☆16Dec 23, 2021Updated 4 years ago
- ☆11Jun 11, 2024Updated 2 years ago
- A python implementation for computing the PoR metric for video summarization from "Performance over Random: A Robust Evaluation Protocol …☆10May 4, 2022Updated 4 years ago
- A collection of publications that works on code models but beyond focusing on the accuracies.☆12Jun 30, 2023Updated 3 years ago
- Repository for the paper: "TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining" ACL Oral 2025☆24Apr 19, 2026Updated 3 months ago
- Alignment between clustered datasets via hierarchical Wasserstein distance☆38Sep 26, 2023Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Mixture-of-Basis-Experts for Compressing MoE-based LLMs☆36Dec 24, 2025Updated 7 months ago
- Implementation of Neural Style Transfer on Video☆10Nov 6, 2018Updated 7 years ago
- The Infibench variant of bigcode-evaluation-harness --- a framework for the evaluation of autoregressive code generation language models.☆14Oct 19, 2024Updated last year
- Dataset for Bilingual VLN☆11Dec 5, 2020Updated 5 years ago
- [ICSE 2026] [ACM SIGSOFT Distinguished Paper Award] SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Re…☆20Apr 22, 2025Updated last year
- Keras implementation of Deep Wasserstein Embeddings☆48Apr 15, 2018Updated 8 years ago
- JMLR Cover Letter Template☆10Dec 15, 2021Updated 4 years ago