A Lightweight Visual Reasoning Benchmark for Evaluating Large Multimodal Models through Complex Diagrams in Coding Tasks
☆17Feb 25, 2025Updated last year
Alternatives and similar repositories for HumanEval-V-Benchmark
Users that are interested in HumanEval-V-Benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Source code for ISSTA'24 paper "AI Coders Are Among Us: Rethinking Programming Language Grammar Towards Efficient Code Generation"☆14Oct 21, 2024Updated last year
- 基于CodeBert预训练模型,微调后/直接对目标数据集进行测试☆14Oct 19, 2021Updated 4 years ago
- EvoEval: Evolving Coding Benchmarks via LLM☆84Apr 6, 2024Updated 2 years ago
- Replication package for ISSTA2023 paper - Towards Efficient Fine-tuning of Pre-trained Code Models: An Experimental Study and Beyond☆23Apr 9, 2023Updated 3 years ago
- Source code embeddings for various programming languages☆17Jul 11, 2018Updated 8 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- This repo is the artifact of FUEL☆17May 19, 2026Updated 3 months ago
- This repo is for our submission for ICSE 2025.☆20Jun 12, 2024Updated 2 years ago
- This is the official implement for the paper 'Domain Adaptive Code Completion via Language Models and Decoupled Domain Databases''☆14Oct 4, 2023Updated 2 years ago
- We introduce FixEval , a dataset for competitive programming bug fixing along with a comprehensive test suite and show the necessity of e…☆26Aug 31, 2022Updated 3 years ago
- Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs☆103Oct 23, 2024Updated last year
- ☆14Jan 22, 2025Updated last year
- UICrit is a dataset containing human-generated natural language design critiques, corresponding bounding boxes for each critique, and des…☆28Nov 19, 2024Updated last year
- The code and datasets of our ACM MM 2024 paper "Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed …☆11Sep 27, 2024Updated last year
- This repo illustrates how to evaluate the artifacts in the paper An Extensive Study on Pre-trained Models for Program Understanding and G…☆26Aug 12, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Exercises for the Dafny Tutorial☆14May 21, 2018Updated 8 years ago
- ☆10Mar 13, 2023Updated 3 years ago
- Official implementation for our paper: Rethinking Video Tokenization: A Conditioned Diffusion-based Approach☆17Apr 2, 2025Updated last year
- 🌿 DeepPrune: Parallel Scaling without Inter-trace Redundancy☆21Apr 20, 2026Updated 3 months ago
- PyTorch使用技巧和教程☆12Apr 17, 2023Updated 3 years ago
- [EACL 2024] ICE-Score: Instructing Large Language Models to Evaluate Code☆80Jun 16, 2024Updated 2 years ago
- [ACL'25 Findings] Official repo for "HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Task"☆40Apr 7, 2025Updated last year
- ☆10May 14, 2024Updated 2 years ago
- Replication Package for "Compressing Pre-trained Models of Code into 3 MB", ASE 2022☆31Oct 10, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor☆23Jun 12, 2025Updated last year
- ☆35Sep 14, 2025Updated 11 months ago
- ☆10Apr 15, 2023Updated 3 years ago
- KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding☆70Apr 5, 2026Updated 4 months ago
- [NeurIPS 2024] TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration☆25Oct 17, 2024Updated last year
- [AAAI 2026] ReCode: Reinforced Code Knowledge Editing for API Updates☆25Jul 1, 2025Updated last year
- [ISSTA 2025] A Large-scale Empirical Study on Fine-tuning Large Language Models for Unit Testing☆13Feb 9, 2025Updated last year
- A Cross-Language Dynamic Information Flow Analysis.☆30Nov 29, 2022Updated 3 years ago
- Repository for Knowledge Enhanced Machine Learning Pipeline (KEMLP)☆10Jun 5, 2021Updated 5 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆13Feb 29, 2024Updated 2 years ago
- RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment☆18Dec 19, 2024Updated last year
- Mutation-based Fault Localization of Deep Neural Networks☆10Jan 25, 2024Updated 2 years ago
- VulTrigger is a tool to for identifying vulnerability-triggering statements across functions and investigating the effectiveness of funct…☆41Dec 29, 2023Updated 2 years ago
- ☆12Jun 8, 2017Updated 9 years ago
- [ICLR 2025] SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models☆20Sep 17, 2025Updated 11 months ago
- Code for ICSE'24 Paper☆14Apr 21, 2024Updated 2 years ago