☆60Jul 10, 2026Updated 2 months ago
Alternatives and similar repositories for WebGen-Bench
Users that are interested in WebGen-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Jul 10, 2026Updated 2 months ago
- ☆23Jul 5, 2024Updated 2 years ago
- A Prompt Learning Framework for Source Code Summarization☆14Dec 26, 2023Updated 2 years ago
- [NeurIPS 2025] UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents☆61Nov 27, 2025Updated 9 months ago
- ☆18May 18, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆26Jul 10, 2026Updated 2 months ago
- ☆21Aug 9, 2024Updated 2 years ago
- [AAAI 2026] Multimodal Deepresearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework☆58Jun 8, 2026Updated 3 months ago
- ⚔️ [ICLR 2026] Official code of "Search Arena: Analyzing Search-Augmented LLMs".☆59Feb 23, 2026Updated 7 months ago
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories☆50Aug 16, 2026Updated last month
- All you need to get started with the LM Playpen Environment for Learning in Interaction.☆17Aug 20, 2026Updated last month
- BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution☆62Oct 13, 2025Updated 11 months ago
- ☆47May 29, 2026Updated 3 months ago
- UICrit is a dataset containing human-generated natural language design critiques, corresponding bounding boxes for each critique, and des…☆28Nov 19, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Automate the build, execution and test of GitHub repositories across programming languages and operating systems.☆146Updated this week
- SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents☆89Aug 14, 2026Updated last month
- ☆18Mar 2, 2026Updated 6 months ago
- Code for the paper: Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping☆41Oct 29, 2024Updated last year
- Assessing Context-Aware Creative Intelligence in MLLMs☆23Jul 22, 2025Updated last year
- ☆162Oct 12, 2022Updated 3 years ago
- Repo for Anonymous purpose, pls don't distribute☆10Oct 2, 2024Updated last year
- Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs☆103Oct 23, 2024Updated last year
- Official implementation of the paper: "Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts"☆32Mar 12, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official implementation of EMNLP'2022 paper "Non-Parametric Domain Adaptation for End-to-End Speech Translation"☆11Oct 26, 2022Updated 3 years ago
- ☆53Updated this week
- [SIGIR 2025] Benchmarking Recommendation, Classification, and Tracing Based on Hugging Face Knowledge Graph☆17Jun 6, 2025Updated last year
- Source code for "An Empirical Study of Code Smells in Transformer-based Code Generation Techniques".☆11Oct 4, 2022Updated 3 years ago
- GPT Table Semantic Parsing with complex & non-intuitive structure.☆17Jul 16, 2025Updated last year
- Model Selection with Large Language Models for Reasoning (EMNLP2023 Findings)☆30Dec 23, 2023Updated 2 years ago
- An experiment to see if chatgpt can improve the output of the stanford alpaca dataset☆12Mar 29, 2023Updated 3 years ago
- WebApp1k benchmark☆14Nov 21, 2025Updated 10 months ago
- Some microbenchmarks and design docs before commencement☆11Feb 1, 2021Updated 5 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A Multi-domain Benchmark for Personalized Search Evaluation☆12Sep 7, 2023Updated 3 years ago
- Light local website for displaying performances from different chat models.☆86Nov 13, 2023Updated 2 years ago
- Official Repo for SvS: A Self-play with Variational Problem Synthesis strategy for RLVR training☆60Updated this week
- Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models☆124Oct 16, 2025Updated 11 months ago
- ☆29Apr 19, 2026Updated 5 months ago
- Code for paper: Unified Text-to-Image Generation and Retrieval☆15Jul 19, 2026Updated 2 months ago
- 中文原生等级化代码能力测试基准☆16Apr 11, 2024Updated 2 years ago