Ray-powered accelerator for MinerU, turning PDF → Markdown into a scalable, cluster-ready data infrastructure. 基于 Ray 的 MinerU 加速层,将 PDF → Markdown 构建为可扩展、面向集群的数据基础设施。
☆69Apr 20, 2026Updated 4 months ago
Alternatives and similar repositories for Flash-MinerU
Users that are interested in Flash-MinerU are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Dataflow-MM, multi-media operators for Dataflow. We aim to prepare data for Multimodal Large Language Models.☆48Apr 13, 2026Updated 5 months ago
- All-in-one intelligent assistant powered by LlamaIndex — RAG, GraphRAG, NL2SQL, Skills & Memory with multimodal support.☆137Updated this week
- The First Unified Agent Data Synthesis Framework for Custom Agentic Task with all-in-one envrionment☆140May 4, 2026Updated 4 months ago
- [SCIS 2024] The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Di…☆64Nov 7, 2024Updated last year
- The source code for “Homophily-Related: Adaptive Hybrid Graph Filter for Multi-View Graph Clustering”☆11Apr 10, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [Accepted By EMNLP 2026 Main Conference] Sequential Diffusion Language Model (SDLM) enhances pre-trained autoregressive language models b…☆99Dec 27, 2025Updated 8 months ago
- codes for Efficient Test-Time Scaling via Self-Calibration☆22Sep 13, 2025Updated last year
- Automated system for LLM evaluation via agents. Doc as below:☆164Aug 31, 2026Updated 2 weeks ago
- A Comprehensive survey on business use cases of AI that help them thrive in the digital economy☆12Oct 7, 2020Updated 5 years ago
- Just a demonstration of some sampling techniques (rejection sampling, importance sampling, sampling importance resampling, Metropolis sam…☆11Aug 24, 2013Updated 13 years ago
- DataFlow Knowledge Graph -- Knowledge graph data preparation with DataFlow style operators and pipelines☆45Jun 30, 2026Updated 2 months ago
- The official code for paper "Can We Leave Deepfake Data Behind in Training Deepfake Detector" (NIPS2024 poster)☆21May 4, 2025Updated last year
- Some useful dockerfiles for DeepLearning and Computer Vision☆28Oct 14, 2021Updated 4 years ago
- A Python package for interacting with the MinerU Vision-Language Model.☆139Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- This is a AUTOSAR documents specific retriever based on LLM and RAG.☆16Nov 12, 2024Updated last year
- ☆17Oct 13, 2025Updated 11 months ago
- ☆19Oct 28, 2025Updated 10 months ago
- An MCP server designed for academic literature research using the OpenAlex free API.☆16Jun 25, 2025Updated last year
- Image2Points: A 3D Point-based Context Clusters GAN for High-Quality PET Image Reconstruction (ICASSP 2024)☆14Jun 16, 2024Updated 2 years ago
- Sliding Convolutional Attention Network for Scene Text Recognition☆11Aug 31, 2018Updated 8 years ago
- Detecting AI-Generated Video via Frame Consistency☆16Feb 28, 2026Updated 6 months ago
- ComfyUI Textual Inversion Training nodes using input images from workflow☆13Jul 21, 2025Updated last year
- [ICLR 2025] Mathematical Visual Instruction Tuning for Multi-modal Large Language Models☆157Dec 5, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A bibliographic reference correction service☆22Dec 8, 2022Updated 3 years ago
- ☆16May 30, 2025Updated last year
- Breaking Semantic Artifacts for Generalized AI-generated Image Detection☆23Mar 3, 2026Updated 6 months ago
- The OpenCitations metadata model: documents and other material.☆19Apr 16, 2026Updated 5 months ago
- ☆33Jun 12, 2025Updated last year
- ☆19Dec 2, 2023Updated 2 years ago
- [PR 2025] The official GitHub page of "MegaHan97K: A Large-Scale Dataset for Mega-Category Chinese Character Recognition with over 97K Ca…☆87May 18, 2026Updated 4 months ago
- Front end ComfyUI nodes for CartoonSegmentation☆17May 22, 2024Updated 2 years ago
- ☆34May 13, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- We introduce BabyVision, a benchmark revealing the infancy of AI vision.☆252Jan 13, 2026Updated 8 months ago
- ☆333Jul 31, 2026Updated last month
- Official repository for "TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving"☆23Sep 1, 2025Updated last year
- AAAI 2024-Controllable Mind Visual Diffusion Model☆16Dec 18, 2023Updated 2 years ago
- R package to analyse Q methodology data☆39Nov 20, 2023Updated 2 years ago
- A mini Photoshop software with c++, OpenCV and Qt☆10Jun 6, 2021Updated 5 years ago
- [NeurIPS'22] Learning from Future: A Novel Self-Training Framework for Semantic Segmentation.☆32Sep 22, 2022Updated 3 years ago