DependEval: a hierarchical benchmark for evaluating LLMs on repository-level code understanding across 8 programming languages.
☆16Jul 28, 2025Updated last year
Alternatives and similar repositories for DependEval
Users that are interested in DependEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICSE '25] LLM Based Input Space Partitioning Testing for Library APIs☆13Jul 27, 2025Updated last year
- This is a JLU-SNL-COMPILER project☆13Oct 6, 2023Updated 2 years ago
- ☆23Mar 28, 2026Updated 5 months ago
- Code and resource of "FinEntity: Entity-level Sentiment Classification for Financial Texts" (EMNLP 2023)☆20Jan 30, 2024Updated 2 years ago
- Codebase for EMNLP 2025 Findings paper "Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs"☆19Nov 14, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆15Oct 15, 2025Updated 10 months ago
- The repository of data and codes for SynTeR, an LLM-based approach to repair obsolete test cases caused by syntactic BCs.☆16Aug 25, 2024Updated 2 years ago
- The open-source repository for PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment, which provides a general per…☆17Aug 28, 2025Updated last year
- The implementation of paper "Leveraging Multimodal Features and Item-level User Feedback for Bundle Construction", WSDM'24.☆18Oct 30, 2025Updated 10 months ago
- ☆22Dec 28, 2024Updated last year
- ☆28Nov 13, 2024Updated last year
- [KDD 2025] Fine-tuning Multimodal Large Language Models for Product Bundling☆16Sep 20, 2025Updated 11 months ago
- [NeurIPS 2025] The implementation of paper "The Emergence of Abstract Thought in Large Language Models Beyond Any Language"☆20Jun 9, 2025Updated last year
- Python library for backtranslation (with Google Translate)☆12Jan 11, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution [ICSE 2026]☆33Nov 11, 2025Updated 9 months ago
- ☆15Oct 17, 2023Updated 2 years ago
- The implementation of paper "Strategy-aware Bundle Recommender System", SIGIR'23.☆17Sep 4, 2023Updated 2 years ago
- Code and data for the paper "Understanding Hidden Context in Preference Learning: Consequences for RLHF"☆27Aug 21, 2024Updated 2 years ago
- Offical implementation of our paper "Exploring the Potential of Diffusion Large Language Models in Code Generation".☆22Oct 29, 2025Updated 10 months ago
- TestGenEval A Real World Unit Test Generation and Test Completion Benchmark☆33Mar 12, 2026Updated 5 months ago
- This repository contains the code for applying One-Token Approximation to a pretrained language model using subword-level tokenization.☆12May 7, 2020Updated 6 years ago
- This repository contains the WordNet Language Model Probing (WNLaMPro) dataset introduced in "Rare Words: A Major Problem for Contextuali…☆14Feb 2, 2020Updated 6 years ago
- CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding [ISSTA 2026]☆30Feb 2, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆12Mar 25, 2022Updated 4 years ago
- ☆13Oct 17, 2020Updated 5 years ago
- The implementation for the work "Recommender Systems with Generative Retrieval".☆21Aug 20, 2025Updated last year
- ☆13Jul 6, 2021Updated 5 years ago
- The implementation of paper "EliMRec: Eliminating single-modal bias in multimedia recommendation", MM'22.☆24Dec 7, 2023Updated 2 years ago
- ☆25Apr 3, 2026Updated 4 months ago
- [NAACL 2025] Benchmark for Repository-Level Code Generation, focus on Executability, Correctness from Test Cases and Usage of Contexts fr…☆47Jan 8, 2026Updated 7 months ago
- SocialDial: A Benchmark for Socially-Aware Dialogue Systems (SIGIR'23)☆16Aug 4, 2023Updated 3 years ago
- a small, non-commercial, fair-use subset of the Penn-Treebank, in JSON.☆18Apr 10, 2018Updated 8 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [2025-上海人工智能实验室书生实训营十佳、优秀项目]☆43Sep 22, 2025Updated 11 months ago
- Code and data for "KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark" (LREC-COLING…☆18Apr 15, 2025Updated last year
- ☆29Jun 2, 2026Updated 3 months ago
- [EMNLP 2021] Code and data for our paper "Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers…☆20Jan 17, 2022Updated 4 years ago
- Data and preprocessing scripts for SemEval 2022 Task 2: Multilingual Idiomaticity Detection and Sentence Embedding☆16Feb 3, 2022Updated 4 years ago
- ☆18Jan 17, 2024Updated 2 years ago
- Weakly Supervised Temporal Anomaly Segmentation☆16Nov 24, 2021Updated 4 years ago