LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
☆22Jun 2, 2026Updated 2 months ago
Alternatives and similar repositories for LoCoBench-Agent
Users that are interested in LoCoBench-Agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆46Jun 2, 2026Updated 2 months ago
- Evaluating Durability: Benchmark Insights into Multimodal Watermarking☆12Jun 7, 2024Updated 2 years ago
- ☆25Jan 29, 2026Updated 7 months ago
- [ECCV 2022] "Improve Few-Shot Transfer Learning with Low-Rank Decompose and Align" by Ziyu Jiang, Tianlong Chen, Xuxi Chen, Yu Cheng, Luo…☆13Jul 19, 2022Updated 4 years ago
- [ICML 2026] Sparser Block-Sparse Attention via Token Permutation☆32May 22, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Functional Optimal Transport: Map Estimation and Domain Adaptation for Functional data☆28Jun 7, 2021Updated 5 years ago
- [ACL 2023]: Training Trajectories of Language Models Across Scales https://arxiv.org/pdf/2212.09803.pdf☆25Nov 14, 2023Updated 2 years ago
- ☆13May 27, 2019Updated 7 years ago
- [COLM 2024] SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models☆25Oct 5, 2024Updated last year
- [NIPS 2025 DB Spotlight] AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios☆40Dec 1, 2025Updated 8 months ago
- Code for Deep learning models for electrocardiograms are susceptible to adversarial attack☆24Feb 4, 2021Updated 5 years ago
- [EMNLP 2023] An Empirical Exploration of Cross-domain Alignment between Language and Electroencephalogram☆31Nov 9, 2023Updated 2 years ago
- Sotopia-RL: Reward Design for Social Intelligence☆52Apr 1, 2026Updated 4 months ago
- [EMNLP 2025🔥] UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective☆20Jan 7, 2026Updated 7 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Neural Adversarial Agent Mutation-based Security Evaluator☆17Apr 18, 2026Updated 4 months ago
- [DMLR 2024] Benchmarking Robustness of Multimodal Image-Text Models under Distribution Shift☆40Jan 25, 2024Updated 2 years ago
- Preprint: Asymmetry in Low-Rank Adapters of Foundation Models☆40Feb 27, 2024Updated 2 years ago
- Benchmarking Language Agents Under Controllable and Extreme Context Growth☆56Apr 29, 2026Updated 4 months ago
- Convert pretrained RoBerta models to various long-document transformer models☆11Apr 5, 2022Updated 4 years ago
- Code for the NIPS 2016 paper "Single-Image Depth Perception in the Wild"☆11Nov 17, 2017Updated 8 years ago
- ☆12Jun 11, 2024Updated 2 years ago
- Generate Python docstrings automatically with LLM and syntax trees☆20Jun 13, 2025Updated last year
- Repository for the paper: "TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining" ACL Oral 2025☆24Apr 19, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- NeurIPS 2024: SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation☆13May 24, 2025Updated last year
- Mixture-of-Basis-Experts for Compressing MoE-based LLMs☆37Dec 24, 2025Updated 8 months ago
- Implementation of Neural Style Transfer on Video☆10Nov 6, 2018Updated 7 years ago
- MiniGPT-4 :: Updated to Torch 2.0, simple setup, easier API, cut out training code☆15Jun 12, 2023Updated 3 years ago
- The Infibench variant of bigcode-evaluation-harness --- a framework for the evaluation of autoregressive code generation language models.☆14Oct 19, 2024Updated last year
- ☆16May 3, 2024Updated 2 years ago
- Deep hybrid models: bridging discriminative and generative approaches https://cs.stanford.edu/~ermon/papers/uai2017_cr.pdf☆17Dec 3, 2017Updated 8 years ago
- 基于CodeBert预训练模型,微调后/直接对目标数据集进行测试☆14Oct 19, 2021Updated 4 years ago
- The replication package of <Sentiment Analysis for Software Engineering: How Far Can Pre-trained Transformer Models Go?>. Accepted by IC…☆11Nov 29, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LISA for ICML 2022☆52Apr 12, 2023Updated 3 years ago
- Multiple Futures Prediction (MFP) on CARLA data☆12Apr 22, 2021Updated 5 years ago
- ☆17May 25, 2020Updated 6 years ago
- MCP as a Judge is a behavioral MCP that strengthens AI coding assistants by requiring explicit LLM evaluations☆17Dec 15, 2025Updated 8 months ago
- AutoLossGen: Automatic Loss Function Generation for Recommender Systems☆22Apr 29, 2022Updated 4 years ago
- ☆24Jul 1, 2024Updated 2 years ago
- ☆67Jun 2, 2026Updated 2 months ago