MinerU: Turn PDF and Office Documents into LLM-Ready Markdown and JSON
MinerU: Turn PDF and Office Documents into LLM-Ready Markdown and JSON
In RAG, enterprise knowledge bases, and agent workflows, document parsing is often underestimated. Converting a PDF to plain text may produce a quick demo, but it can lose heading hierarchy, table relationships, mathematical formulas, images, and cross-page context. Chunking, embedding, and retrieval then amplify those errors.
OpenDataLab/MinerU is an open-source document parsing engine that converts PDF, images, DOCX, PPTX, and XLSX into machine-readable formats such as Markdown and JSON. It is not a chatbot. Its role is the layer between raw documents and usable data, turning complex layouts into structured input for downstream models.
At the time of verification, the repository had more than 77,000 GitHub stars and continued to release updates. That makes MinerU worth evaluating as an engineering component for real document pipelines rather than treating it as an OCR showcase alone.
Why plain text extraction is not enough
A conventional PDF text extractor can retrieve characters without understanding their layout. A two-column paper may be read in the wrong order. A table may become a flat string with no relationship between headers and values. Equations and special symbols may be damaged by OCR. Images, captions, and body text may become disconnected. Scanned pages need OCR, but OCR still has to be combined with layout analysis to restore reading order.
Humans can recover much of this information visually. Embedding models and LLMs cannot. For them, the extracted text is the document. Parsing quality therefore affects retrieval recall, citation accuracy, and answer reliability.
MinerU treats document understanding as a complete pipeline rather than a single OCR function. The official README describes capabilities covering layout analysis, text and table recognition, formula processing, image descriptions, and outputs such as Markdown and JSON. Results still depend on document types and model backends, so the project recommends evaluating representative samples with its online demo first.
How MinerU fits into a pipeline
A typical flow looks like this:
PDF / DOCX / PPTX / XLSX / images
|
v
Layout analysis and OCR/VLM
|
v
Headings, paragraphs, tables, formulas, images
|
v
Markdown / JSON / artifacts
|
v
Cleaning, chunking, embeddings
|
v
RAG or agent apps
This separation matters. MinerU does not decide your chunk size, vector database schema, or prompts. It first preserves the document structure so downstream components can work with a more stable representation. For engineering teams, that is easier to test and replace than putting every responsibility inside one end-to-end question-answering service.
The project provides a local CLI, FastAPI, Gradio WebUI, Docker deployment, and mineru-router for multi-service and multi-GPU routing. You can validate one document through the CLI and later evolve the same capability into an API service.
Minimal usage
The official README shows installation with uv:
uv pip install -U "mineru[all]"
To install from source:
git clone https://github.com/opendatalab/MinerU.git
cd MinerU
uv pip install -e .[all]
The basic CLI command specifies an input and output path:
mineru -p ./documents/report.pdf -o ./outputs
You can also explicitly select the pipeline backend:
mineru -p ./documents/report.pdf -o ./outputs -b pipeline
These commands are useful for establishing a reproducible baseline. Start with representative samples: two-column papers, scanned PDFs, table-heavy reports, documents containing formulas, and slide decks. Compare output quality and processing time before running an entire archive.
Three design points that matter for RAG
1. Preserve structure before choosing chunk boundaries
If a document is flattened into unstructured text before splitting, it becomes difficult to identify the chapter or table a passage belongs to. Preserve Markdown headings, tables, and image descriptions, then attach the section path as metadata during chunking:
{
"text": "This section discusses hybrid retrieval...",
"metadata": {
"source": "annual-report.pdf",
"section": "Chapter 3 / Retrieval Architecture",
"page": 18,
"content_type": "paragraph"
}
}
This does not mean every metadata field will be perfect automatically. It means the parsing result retains enough information for your own cleaning layer to add source, page, and section metadata.
2. Do not immediately flatten tables into paragraphs
Tables often contain the most valuable comparison data and are also the easiest content to damage during text conversion. Consider retaining two representations: a Markdown table for human reading and a JSON structure for programmatic processing. Depending on the question, retrieval can use text search, field filtering, or table reasoning.
3. Make parsing observable
Document parsing is not a one-off script. Record the input hash, page count, backend, processing time, output length, table count, OCR pages, and failure reason. When a user reports an incorrect citation, you can then determine whether the problem is in parsing, chunking, retrieval, or generation.
Pipeline, VLM, and service deployment
MinerU presents pipeline and VLM-style backends as part of the same product direction. A practical interpretation is:
- Use the
pipelinebackend to establish a stable, batch-oriented local process and benchmark common formats. - Use a VLM-oriented backend when complex visual understanding is required, while accounting for model size, GPU memory, latency, and cost.
- Once quality is validated on a single machine, consider API, Docker, or router deployment so multiple upstream systems can share the parsing service.
Recent project updates focus on batch processing, model caching, long documents, asynchronous tasks, multithreading, and multi-GPU routing. These capabilities matter more in production than a single demo because real workloads involve long-running jobs, retries, concurrency, and resource isolation.
Licensing and data governance
The README currently identifies the MinerU Open Source License, based on Apache 2.0 with additional conditions. Before commercial use, confidential-document processing, or repackaging the model and service, read the repository's LICENSE.md and obtain a legal review.
Document parsing may involve contracts, financial reports, medical records, or internal policies. Even when the tool runs locally, design for temporary-file cleanup, access isolation, output redaction, and audit records. Before sending documents to a third-party API, confirm that the data is allowed to leave your environment.
Who should evaluate MinerU?
MinerU is a strong candidate when you need to:
- Ingest a large PDF or Office corpus into a RAG knowledge base.
- Preserve tables, formulas, headings, and image context rather than accepting OCR text alone.
- Validate a workflow through CLI or WebUI before deploying an API service.
- Process documents on local or owned GPU infrastructure.
- Separate parsing from chunking, embeddings, vector search, and LLM generation.
It is not automatically suitable for every document. Highly unusual layouts, large amounts of handwriting, poor scans, and domain-specific correction requirements still need evaluation against your own dataset. The official README also notes that complex layouts, scanned pages, and handwriting may produce unexpected results.
Conclusion: Treat document parsing as AI infrastructure
A common RAG mistake is to discuss vector databases and prompts before verifying whether the document still has the correct structure when it enters the system. MinerU has a clear position: it is the parsing layer between documents and machine-readable data. Markdown, JSON, CLI, API, and service routing make that layer testable, observable, and scalable.
If your current workflow is still “PDF to plain text, then chunk,” MinerU is worth using as a benchmark. Compare it on a small set of real documents, focusing on tables, formulas, images, and long pages, and then decide whether it fits your AI Chain. That evaluation is more meaningful than looking only at a demo or star count.