[CVPR 2025] Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
☆16Jun 16, 2025Updated last year
Alternatives and similar repositories for DocMark
Users that are interested in DocMark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents☆61Nov 27, 2025Updated 8 months ago
- PyTorch implementation of "UNIT: Unifying Image and Text Recognition in One Vision Encoder", NeurlPS 2024.☆34Sep 26, 2024Updated last year
- [3DV2026] Official repository for "CamC2V: Context-aware Controllable Video Generation"☆14Nov 11, 2025Updated 8 months ago
- 测试 https://huggingface.co/OFA-Sys/gsm8k-rft-llama7b-u13b 的 GSM8K 分数☆15Aug 10, 2023Updated 3 years ago
- [Paper] Code for the EMNLP2023 (Findings) paper "Global Structure Knowledge-Guided Relation Extraction Method for Visually-Rich Document"☆17Dec 1, 2023Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Create PDF animations from graphics files and inline graphics using LaTeX☆12Jun 8, 2018Updated 8 years ago
- It's the code for the paper Pushing the Performance Limit of Scene Text Recognizer without Human Annotation, CVPR 2022.☆28Jul 6, 2022Updated 4 years ago
- ☆12Mar 8, 2022Updated 4 years ago
- ☆42Sep 2, 2023Updated 2 years ago
- Official implementation of "LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Guided Image Editing☆16May 27, 2025Updated last year
- 清华大学校园网客户端与联网库,适用于命令行环境,Windows、Linux、Mac OS X桌面平台与UWP、iOS、Android移动平台☆12Mar 3, 2020Updated 6 years ago
- AI-powered slide workspace for creating, editing, versioning, and presenting beautiful reveal.js decks from prompts and source files.☆17Apr 14, 2026Updated 3 months ago
- ☆11Jan 27, 2020Updated 6 years ago
- Official repository for the EMNLP 2025 paper “UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets”.☆16Sep 19, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark☆17May 25, 2025Updated last year
- [EMNLP Findings'25] Official PyTorch Implementation of Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Align…☆16Sep 19, 2025Updated 10 months ago
- [ICCV 2025] Official implementation of "Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing"☆28Apr 15, 2025Updated last year
- ☆19Jul 7, 2025Updated last year
- 《카카오 아레나 데이터 경진대회 1등 노하우》 예제 코드☆17Jan 15, 2021Updated 5 years ago
- Repository for ACL2020 paper "Refer360° A Referring Expression Recognition Dataset in 360°Images"☆15Jun 26, 2021Updated 5 years ago
- [TPAMI] Locating and Counting Heads in Crowds With a Depth Prior☆10Jan 7, 2022Updated 4 years ago
- Official PyTorch implementation of “MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation”☆18Dec 5, 2024Updated last year
- DCEN☆13Aug 12, 2021Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Index of URLs to pdf files all over the internet and scripts☆25May 2, 2023Updated 3 years ago
- [WAICA-26 Best Student Paper] Official repository of "Enhancing Vision Foundation Models via Multimodal Continual Pre-Training"☆49Jul 21, 2026Updated 3 weeks ago
- Code release for "Weakly Supervised Open-Vocabulary Object Detection", AAAI2024☆36Sep 9, 2024Updated last year
- A tool for better use of Inspire platform (Beta: Codeberg version is more up-to-date)☆29Apr 2, 2026Updated 4 months ago
- The official repository for paper Evaluating Financial Relational Graphs: Interpretation Before Prediction☆20Jan 2, 2026Updated 7 months ago
- [CVPR 2025] Docopilot: Improving Multimodal Models for Document-Level Understanding☆37Jul 22, 2025Updated last year
- Data and code for ACL 2023 paper "RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations"☆15Feb 8, 2024Updated 2 years ago
- [ICLR 2026] Code for Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model☆30Mar 1, 2026Updated 5 months ago
- The official code for the CVPR 2024 paper: Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer☆55Jun 14, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆16May 20, 2026Updated 2 months ago
- Official implementation for [ICML 2026] Scalable GANs with Transformers☆18Jul 1, 2026Updated last month
- [ICRA 2026] Official implementation of the paper: “EgoTraj-bench: Towards robust trajectory prediction under ego-view noisy observations”☆22Jul 25, 2026Updated 2 weeks ago
- [CVPR 2026] SOTA Chemical Reaction Diagram Parsing Framework☆26Mar 24, 2026Updated 4 months ago
- Trying to kill the "Uncanny Valley" of uniform entropy or something.☆28Apr 29, 2026Updated 3 months ago
- Official implementation for our paper: Rethinking Video Tokenization: A Conditioned Diffusion-based Approach☆17Apr 2, 2025Updated last year
- A PyTorch implementation of "VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics"☆19Mar 9, 2026Updated 5 months ago