☆80Oct 1, 2026Updated this week
Alternatives and similar repositories for Vision2Web
Users that are interested in Vision2Web are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆87May 2, 2026Updated 5 months ago
- Computer-Use Agents as Judges for Generative UI☆45Nov 27, 2025Updated 10 months ago
- WorldSense benchmark for grounded reasoning in language models☆25Nov 28, 2023Updated 2 years ago
- ☆39Nov 14, 2025Updated 10 months ago
- ☆20Aug 1, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆48Jul 26, 2026Updated 2 months ago
- Accompanying repo for NeurIPSW'23: GPT4GEO: How a Language Model Sees the World's Geography☆27May 24, 2025Updated last year
- ☆11Jun 11, 2025Updated last year
- NeurIPS 2024: SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation☆13May 24, 2025Updated last year
- ☆43Aug 23, 2026Updated last month
- Official implementation for MGN☆20Dec 22, 2022Updated 3 years ago
- Official Implementation for the paper "VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models"☆24Aug 14, 2025Updated last year
- ☆29Sep 15, 2026Updated 2 weeks ago
- Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction (CVPR 2026)☆16Feb 27, 2026Updated 7 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models☆22Apr 14, 2026Updated 5 months ago
- The official repo for "VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search" [EMNLP25]☆39Feb 1, 2026Updated 8 months ago
- ☆12Jan 10, 2025Updated last year
- ☆54Jun 3, 2026Updated 4 months ago
- This is an implementation of the paper "Are We Done with Object-Centric Learning?"☆14Jun 21, 2026Updated 3 months ago
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients☆22Jun 17, 2025Updated last year
- SAM4SS: Tailoring SAM and SAM2 for Semantic Segmentation☆11Jul 31, 2024Updated 2 years ago
- ☆14Jan 22, 2025Updated last year
- Official repo for the TMLR paper "Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners"☆29Apr 27, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Kaleido: Open-sourced multi-subject reference video generation model, enabling controllable, high-fidelity video synthesis from multiple …☆149Mar 2, 2026Updated 7 months ago
- ☆25Jul 10, 2026Updated 2 months ago
- Code and Models for "GeneCIS A Benchmark for General Conditional Image Similarity"☆61Jun 12, 2023Updated 3 years ago
- ☆22Apr 24, 2025Updated last year
- SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner☆40Jun 29, 2025Updated last year
- ☆10Dec 21, 2022Updated 3 years ago
- Official code repo for the paper "ChemToolAgent: The Impact of Tools on Language Agents for Chemistry Problem Solving" (previously "Tooli…☆20Jun 7, 2025Updated last year
- Official repo for "TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders"☆28Apr 9, 2026Updated 5 months ago
- OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models☆30Feb 4, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This is the official code of "Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic Segmentation, NeurIPS 23"☆27Dec 7, 2023Updated 2 years ago
- LMM for VQA, tcsvt version☆10Jul 19, 2024Updated 2 years ago
- COLA: Evaluate how well your vision-language model can Compose Objects Localized with Attributes!☆25May 14, 2026Updated 4 months ago
- Demo page of TAVGBench: Benchmarking Text to Audible-Video Generation☆15Apr 7, 2025Updated last year
- Compress conventional Vision-Language Pre-training data☆52Sep 22, 2023Updated 3 years ago
- ☆12Jul 24, 2023Updated 3 years ago
- ☆565Updated this week