[CVPR 2025] Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
☆17Jun 16, 2025Updated last year
Alternatives and similar repositories for DocMark
Users that are interested in DocMark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents☆62Nov 27, 2025Updated 10 months ago
- PyTorch implementation of "UNIT: Unifying Image and Text Recognition in One Vision Encoder", NeurlPS 2024.☆34Sep 26, 2024Updated 2 years ago
- Benchmarking End-to-End Photographed Document Parsing and Translation☆18Dec 4, 2025Updated 10 months ago
- Official repository of "SeGA: Preference-Aware Self-Contrastive Learning with Prompts for Anomalous User Detection on Twitter" @ AAAI 202…☆10Nov 30, 2024Updated last year
- 一个桌面宠物程序,现在似乎发展成为桌面便签了。桌面便签程序见develop-todolist分支。☆11Nov 17, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [3DV2026] Official repository for "CamC2V: Context-aware Controllable Video Generation"☆14Nov 11, 2025Updated 10 months ago
- SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis☆37Jun 13, 2025Updated last year
- 为visinger SVS系统写的展示系统~本质仍然是个音乐播放器☆11Apr 18, 2023Updated 3 years ago
- This project uses Centernet and Conditional Convolutions for Instance Segmentation☆13Sep 17, 2020Updated 6 years ago
- 测试 https://huggingface.co/OFA-Sys/gsm8k-rft-llama7b-u13b 的 GSM8K 分数☆15Aug 10, 2023Updated 3 years ago
- Create PDF animations from graphics files and inline graphics using LaTeX☆12Jun 8, 2018Updated 8 years ago
- ☆12Mar 8, 2022Updated 4 years ago
- [ICCV2025] TextSSR: Diffusion-based Data Synthesis for Scene Text Recognition☆92Mar 6, 2026Updated 7 months ago
- ☆42Sep 2, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation of "LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Guided Image Editing☆16May 27, 2025Updated last year
- 清华大学校园网客户端与联网库,适用于命令行环境,Windows、Linux、Mac OS X桌面平台与UWP、iOS、Android移动平台☆12Mar 3, 2020Updated 6 years ago
- PFCC 社区博客☆14Updated this week
- AI-powered slide workspace for creating, editing, versioning, and presenting beautiful reveal.js decks from prompts and source files.☆17Apr 14, 2026Updated 5 months ago
- ☆11Jan 27, 2020Updated 6 years ago
- [ICCV 2025] Official implementation of "Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing"☆28Apr 15, 2025Updated last year
- pytorch crnn with centerloss to solve the near word problem☆16Jan 27, 2022Updated 4 years ago
- Repository for ACL2020 paper "Refer360° A Referring Expression Recognition Dataset in 360°Images"☆15Jun 26, 2021Updated 5 years ago
- [TPAMI] Locating and Counting Heads in Crowds With a Depth Prior☆10Jan 7, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official PyTorch implementation of “MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation”☆18Dec 5, 2024Updated last year
- [WAICA-26 Best Student Paper] Official repository of "Enhancing Vision Foundation Models via Multimodal Continual Pre-Training"☆50Jul 21, 2026Updated 2 months ago
- Code release for "Weakly Supervised Open-Vocabulary Object Detection", AAAI2024☆36Sep 9, 2024Updated 2 years ago
- Code for paper "Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-based Approach"☆21Jul 9, 2020Updated 6 years ago
- [ICML2025] A Non-isotropic Time Series Diffusion Model with Moving Average Transitions☆17Jun 23, 2025Updated last year
- 记录一些学习生活中的收集☆18Sep 7, 2026Updated last month
- [CVPR 2025] Docopilot: Improving Multimodal Models for Document-Level Understanding☆37Jul 22, 2025Updated last year
- Data and code for ACL 2023 paper "RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations"☆15Feb 8, 2024Updated 2 years ago
- [ICLR 2026] Code for Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model☆30Mar 1, 2026Updated 7 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- The official code for the CVPR 2024 paper: Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer☆55Jun 14, 2024Updated 2 years ago
- 音乐可视化 canvas☆19Jan 4, 2023Updated 3 years ago
- Pedestrian Trajectory Prediction with Missing Data: Datasets, Imputation, and Benchmarking (NeurIPS 2024)☆18Feb 16, 2025Updated last year
- ☆17Nov 12, 2025Updated 10 months ago
- Official implementation for [ICML 2026] Scalable GANs with Transformers☆22Jul 1, 2026Updated 3 months ago
- [ICCV 2025] CHORDS: Diffusion Sampling Accelerator with Multi-core Hierarchical ODE Solvers☆17Mar 3, 2026Updated 7 months ago
- A PyTorch implementation of "VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics"☆20Aug 18, 2026Updated last month