OpenDataLoader PDF: a PDF parser for AI-ready data
opendataloader-project/opendataloader-pdf · 29,370★ · 2,799 forks
Everything worth knowing about opendataloader-project/opendataloader-pdf: an open-source PDF parser, written mostly in Java, that converts PDF documents into Markdown, JSON (with bounding boxes), and HTML for RAG and LLM pipelines, and also automates accessibility by tagging structure-less PDFs into Tagged PDF. It’s a Hancom Inc. (Korea) project developed alongside the PDF Association and Dual Lab (the creators of veraPDF).
Origin
The opendataloader-project organization was created on May 12, 2025 (the repository, on May 13, 2025), and its profile places it in South Korea, with 339 followers and 6 public repositories. The main contributor, bundolee (Bundo Lee), has accumulated 562 commits (versus 75 each for MaximPlusov and LonelyMidoriya), lists Seoul in their profile, and their company field reads “Hancom Inc. @opendataloader-project.” Hancom is a South Korean office and document software company (creator of the HWP format), which gives the project a clear corporate origin: it’s not an individual experiment but a bet placed by a company in the document sector.
Visible public traction started around September 2025, when the reference Hacker News thread (109 points, 28 comments) linked to the repository. The jump to v2.0 — which introduced the hybrid engine, the Apache 2.0 license, and Llama/LangChain integrations — spread in March 2026. The latest version measured in this run is v2.5.2, published August 21, 2026.
The README documents a relevant license change: versions before 2.0 were distributed under Mozilla Public License 2.0, and from 2.0 onward under Apache 2.0 — a decision the FAQ justifies as avoiding MPL’s per-file copyleft and easing enterprise adoption.
Philosophy and principles
The README summarizes the proposal in two blocks that define its principles:
- AI-ready data, not just text: the goal is to produce structure (reading order, heading hierarchy, tables, coordinates) that LLMs and RAG pipelines need, not just plain text. The site’s message is explicit: “If your data isn’t parsed well, your RAG system will never retrieve correct answers. garbage in = garbage out.”
- Local-first and cloud-free: the README and site insist everything runs locally, with no cloud calls or API keys; “your document never leaves your environment.” Version 2.0 frames this as usable in fully air-gapped environments, eliminating data leakage risk.
- Local determinism + smart hybrid: local mode is deterministic (rules, no GPU, no models), and a hybrid mode routes only complex pages (difficult tables, scans, formulas, charts) to a local AI backend.
- Coordinates for every element: each extracted element carries a
[left, bottom, right, top]bounding box in PDF points, which enables source citations and “click-to-source” in RAG. - Accessibility as first-class, not an add-on: the project presents itself as the first open-source tool that generates Tagged PDF end-to-end, built on the PDF Association’s Well-Tagged PDF specification and validated with veraPDF.

How it works
The input is a PDF (digital, scanned, or already tagged), and the outputs are Markdown, JSON (with bounding boxes), HTML, Tagged PDF, and, as a paid add-on, PDF/UA. The core is a deterministic Java parser that applies layout analysis and an XY-Cut++ reading order (correct sequence across multi-column pages, sidebars, and mixed layouts).

Two modes documented in the README:
- Local (default, fast):
opendataloader-pdf file1.pdf folder/. Extracts text with reading order, tables (borders), headings, lists, and images, and computes bounding boxes. The README states 60+ pages/second on CPU; with multi-threaded batch processing, over 100 pages/second on 8+ core machines (measured on Apple M4). - Hybrid: combines the Java core with a local AI backend. Simple pages are processed locally; complex ones are sent to the backend to boost accuracy (the README cites a +90% improvement in table accuracy: from 0.489 to 0.928 TEDS). Requires starting a server (
opendataloader-pdf-hybrid --port 5002) and calling the client with--hybrid docling-fast. Supports OCR (80+ languages) for scanned PDFs, LaTeX formula extraction, and chart/image description (a 256M SmolVLM model).

Other documented mechanisms: Tagged PDF support, where if the PDF already has structure tags, it extracts the author’s exact layout with no heuristics; and auto-tagging into Tagged PDF, which analyzes the layout and generates structure tags for an untagged PDF, under Apache 2.0 (PDF/UA-1/2 conversion and the visual editor are paid add-ons).

AI safety (anti-injection): automatically filters hidden text (transparent or zero-size fonts), off-page content, and suspicious invisible layers. With --sanitize it also replaces sensitive data (emails, URLs, phone numbers) with markers.

Quick-start guide
Installation and first run
Requirements: Java 11+ (check with java -version; if missing, install a JDK 11+ from Adoptium) and Python 3.10+.
pip install -U opendataloader-pdf
import opendataloader_pdf
# Batch all files into a single call: each convert() spawns a JVM process,
# so calling it in a loop is slow
opendataloader_pdf.convert(
input_path=["file1.pdf", "file2.pdf", "folder/"],
output_dir="output/",
format="markdown,json"
)
For hybrid mode (complex tables, scans, formulas): pip install "opendataloader-pdf[hybrid]". There’s also a Node.js SDK (npm install @opendataloader/pdf) and Java (Maven org.opendataloader:opendataloader-pdf-core).
Common workflows
- To extract to Markdown + JSON for RAG:
opendataloader_pdf.convert(input_path=[...], output_dir="output/", format="markdown,json"); the result lands inoutput/with Markdown ready to chunk and JSON with bounding boxes for citations. - For a scanned PDF or another language: start the backend with OCR —
opendataloader-pdf-hybrid --port 5002 --force-ocr --ocr-lang "ko,en"— and process withopendataloader-pdf --hybrid docling-fast file.pdf. - For formulas or chart descriptions: backend with
--enrich-formulaor--enrich-picture-description, and client with--hybrid-mode full(formula and chart enrichment requires this mode). - To access an untagged PDF (accessibility):
opendataloader-pdf --format tagged-pdf file.pdf, or in Pythonformat="tagged-pdf".
Essential configuration
There’s no central configuration file: “configuration” means convert()/CLI parameters. The ones a new user touches first: format (markdown, json, html, pdf, text, or tagged-pdf; combinable), hybrid (enables hybrid mode), output_dir (output folder), use_struct_tree (uses the PDF’s native structure tags), and image_output/image_format (off, embedded, or external, with png/jpeg).
Common pitfalls and fixes
- Every
convert()spawns a JVM process: repeated calls in a loop are slow; batch all files into a singleconvert()call. --use-struct-treetakes priority over--hybrid: if both are set on a tagged PDF, the structure tree is used and the hybrid backend isn’t called; to use the hybrid, drop--use-struct-tree.- Quality depends on tag quality: poorly tagged PDFs produce worse results; with sparse or incorrect tags, heuristic mode or
--hybrid docling-fasttends to give better quality. - Formulas and chart descriptions require
--hybrid-mode fullon the client; without that mode, the enrichment isn’t applied. - Missing Java: if
java -versiondoesn’t respond, the Python install works as a wrapper but the engine needs the JDK; install Java 11+. - Language limitation (community): an HN user (
hermitcrab) noted that being Java/Python, it doesn’t fit those who need to call it from C++; there’s no documented native C++ binding. - Self-reported benchmark figures: the “number 1” scores come from the project’s own harness (
opendataloader-bench); the repository publishes it so others can reproduce it, but it’s not an independent evaluation.
Integrations and migration
- LangChain: officially documented integration —
pip install -U langchain-opendataloader-pdfandOpenDataLoaderPDFLoader(file_path=[...], format="text"). - LlamaIndex: official reader from the organization (
opendataloader-project/opendataloader-pdf-llamaindex), “fast, accurate, local PDF reading.” - Local hybrid backend: requires no cloud or keys; the
opendataloader-pdf-hybridserver runs on the same machine. - Migrating from other parsers: the README positions itself as a replacement for tools like
pdf2docx(one HN user reports migrating “to replace pdf2docx … and it’s much better”). Against docling, marker, or pymupdf4llm, the README highlights default bounding boxes and injection filters, which those projects don’t offer by default. There’s no formal migration process: it’s reconfiguring the conversion call towardconvert().
The ecosystem
Repositories from the same opendataloader-project organization: opendataloader-bench (83 stars, the project’s own evaluation harness, measuring reading order NID, table fidelity TEDS, and heading hierarchy MHS over a corpus of ~200 real PDFs), langchain-opendataloader-pdf (59 stars, LangChain document loader), opendataloader-pdf-llamaindex (3 stars, LlamaIndex reader), opendataloader-pdf-examples (2 stars, usage examples), and .github (1 star).

Community projects and derivatives (GitHub repository search for “opendataloader”): NameetP/pdfmux (81 stars, Python, created March 2026) is a PDF extractor that “audits its own output and certifies other extractors’,” scoring 0.903 on opendataloader-bench (2nd of 8 engines) and exposing a 7-tool MCP server — the most notable related project, since it competes on the same benchmark as OpenDataLoader. muschellij2/opendataloader (0 stars, R, April 2026) is an R interface to OpenDataLoader PDF. coryisakson/opendataloader_docker (4 stars) is a self-hosted Docker package. dpaidev/opendataloader-pdf (2 stars) is a full mirror with history. JamoCA/cfml-OpenDataLoader-demo (1 star) and longshines/pdf-parser-online (0) are demos/web wrappers.
Most other results are mirrors, personal forks, or zero-star repos. Disambiguation note: AstunTechnology/OpenDataLoader (7 stars) is a different project (“a utility for loading OpenData into PostGIS”); its name match is accidental and unrelated to this repository.
Official / semi-official status
There’s no vendor “marketplace” in the style of agent plugins, but the project has solid semi-official status backed by standards bodies and corporate support: collaboration with standards organizations (the README and site state that auto-tagging was built in collaboration with the PDF Association, author of the Well-Tagged PDF specification, and with Dual Lab, developers of veraPDF, the open reference validator for PDF/A and PDF/UA); Hancom Inc. backing (the main contributor lists Hancom as their company; the README announces the enterprise integration Hancom Data Loader — document analysis with custom models, OCR with SLA, and native HWP/HWPX support — as a next step, Q2–Q3 2026 roadmap); “official” third-party integrations (it appears in LangChain’s official integration docs and maintains an official LlamaIndex reader); and an open-core model (the core — extraction, layout analysis, auto-tagging to Tagged PDF — is Apache 2.0 and free; the enterprise add-ons — PDF/UA-1/2 export and the visual accessibility editor/studio — are paid on request).
The project self-describes as “the only open-source parser” that combines local deterministic extraction, per-element bounding boxes, XY-Cut++, AI safety filters, and Tagged PDF support, and positions itself as “number 1 on the benchmark.” That’s a project claim backed by its own harness (opendataloader-bench), not a formal standard designation; no independent external certification was found in this run.

Repo numbers
Measured: August 21–22, 2026, GitHub API.
| Metric | Value |
|---|---|
| Stars | 28,641 |
| Forks | 2,734 |
| Real subscribers | 111 |
| Commits | 861 |
| Open issues + PRs | 77 |
| Primary language | Java |
| License | Apache-2.0 (pre-2.0: MPL 2.0) |
| Created | May 13, 2025 |
| Last push | August 21, 2026 |
| Latest release | v2.5.2, August 21, 2026 |
Top contributors by contributions: bundolee (562), MaximPlusov (75), LonelyMidoriya (75), hyunhee-jo (58), hnc-jglee (19), dependabot[bot] (16), suji-cho (16).
Registry downloads: PyPI opendataloader-pdf (latest version 2.5.2, 69 releases) with 44,408 downloads in the last week and 191,684 in the last month. npm @opendataloader/pdf with 54,297 downloads in the July 22–August 22, 2026 range. Java distribution is via Maven (org.opendataloader:opendataloader-pdf-core).
watchers_count mirrors the star count, so subscribers_count is reported as the real subscriber figure. open_issues_count includes open change requests, not only issues. Registry download figures are delivery metrics, not a count of users or unique deployments.
How the community received it
Verifiable conversation concentrates on one Hacker News thread; the rest of the submissions have very little traction.
Hacker News (109 points, 28 comments, “OpenDataLoader-PDF: An open source tool for structured PDF parsing,” September 23, 2025) gathered concrete opinions. 4d66ba06 (praise and adoption): “just finished migrating to this to replace pdf2docx in a project I was working on and it’s much better. Thanks for open-sourcing OpenDataLoader-PDF.” emilburzo (practical): tested it with bank statements, “PDFs that are surprisingly hard… the JSON extraction looks pretty good and seems to produce something usable in a single pass.” fedeb95 (non-AI use): “Very cool. I’ll probably use it, but not for AI. I have a lot of PDFs for which no epub exists.” agsqwe asked how it compares to docling. hermitcrab (technical objection): “got excited until I read it was Java/Python. Looking for a library that extracts PDF tables and can be called from a C++ program.” trevor-e (deeper critique) argues that “maybe we need a new AI-friendly file format instead of continuing to patch over the complicated PDF spec.” constantinum mentioned Unstract (Zipstack/unstract) as an open structured-extraction + ETL alternative.
Other low-traction submissions don’t constitute a broad review: “Show HN: OpenDataLoader – Safe, Open, High-Performance PDF Loader for AI” (4 points), one about v2.0 (1 point), and “LLM Is Not a PDF Parser: Use OpenDataLoader First” (2 points). No additional verifiable community evidence was recovered outside HN in this run.
OpenDataLoader PDF versus other projects
The table values come from the project’s own harness, opendataloader-bench (reading order NID, table TEDS, heading MHS, over ~200 PDFs); they’re self-reported by the project, not an independent evaluation. Speed is in seconds/page (lower is better).
| Project | Overall score (own bench) | Speed (s/page) | Verifiable difference |
|---|---|---|---|
| opendataloader [hybrid] | 0.907 | 0.463 | Local hybrid engine; per-element bounding boxes; injection filters; Apache 2.0. |
| nutrient | 0.885 | 0.008 | Commercial engine; the fastest in the harness. |
| docling | 0.882 | 0.762 | MIT; per the README, lacks bounding boxes and injection filters by default. |
| marker | 0.861 | 53.932 | GPL-3.0; the README describes it as much slower (≈54 s/page). |
| unstructured [hi_res] | 0.841 | 3.008 | Apache-2.0; standard unstructured also exists (0.686). |
| edgeparse | 0.837 | 0.036 | Apache-2.0. |
| opendataloader (local mode) | 0.831 | 0.015 | The non-AI variant, very fast; worse on tables (0.489). |
| mineru | 0.831 | 5.962 | AGPL-3.0 (opendatalab/MinerU). |
| pymupdf4llm | 0.732 | 0.091 | AGPL-3.0; fast but worse on tables (0.401) and headings (0.412). |
| markitdown | 0.589 | 0.114 | MIT. |
| liteparse | 0.576 | 1.061 | Apache-2.0. |
In practice, the differentiator the project itself claims is the combination of local deterministic extraction + per-element bounding boxes + auto-tagging to Tagged PDF; a commercial engine (nutrient) or a fast parser (pymupdf4llm) can win on speed, and a GPU-based parser (marker, docling) is competitive on quality, but per the project’s own harness, none combine all those capabilities in local mode.
Contributing
CONTRIBUTING.md documents an explicit process: fork the repository and clone the fork; create a working branch (git checkout -b my-feature); build the project (prerequisites: Java 11+, Maven, Python 3.10+, uv, Node.js 24 active LTS, and pnpm via corepack enable pnpm; CI builds against Node 24 and pnpm 11.21.0, and Node must be ≥22.13); for questions, open an issue tagged Question; for bugs, use the Bug Report template; for proposals, the Feature Request template.
Use cases and who this repository can help
- Teams building RAG/LLM pipelines over PDF documents: extraction to Markdown (reading order, tables, headings) and to JSON with bounding boxes enables semantic chunking and “click-to-source” citations. The local hybrid mode covers complex tables and scans without sending data to the cloud — relevant for sectors with sensitive data (legal, healthcare, finance).
- Organizations that must comply with accessibility regulations (EU EAA, US ADA/Section 508, Korea’s Digital Inclusion Act): the audit → auto-tag → Tagged PDF pipeline (free, Apache 2.0) replaces part of the manual remediation work (which the README quantifies at $50–200 USD per document); veraPDF validation makes it suitable for compliance teams.
- Processing scanned or multilingual PDFs: hybrid OCR (80+ languages) serves anyone converting digitized files or documents in Korean/Japanese/Chinese/Arabic; Hancom’s Korean origin explains the announced native HWP support.
- Scientists and technical staff extracting formulas and charts: LaTeX formula extraction and chart/image description cover scientific papers and technical documents where table and equation quality matters.
- Developers integrating into existing frameworks: official LangChain and LlamaIndex loaders let you plug OpenDataLoader in as a document source in already-built pipelines.
- Who is NOT its audience: those needing a native C++ binding, those requiring a fully independent validation benchmark, or those needing PDF/UA without the paid add-on.
Resources
- Repository: https://github.com/opendataloader-project/opendataloader-pdf
- Official site and documentation: https://opendataloader.org
- Official benchmark harness: https://github.com/opendataloader-project/opendataloader-bench
- Official LangChain integration: https://docs.langchain.com/oss/python/integrations/document_loaders/opendataloader_pdf
- Official LlamaIndex reader: https://github.com/opendataloader-project/opendataloader-pdf-llamaindex
- Package registries: PyPI https://pypi.org/project/opendataloader-pdf/ · npm https://www.npmjs.com/package/@opendataloader/pdf
Note: this report combines the repository’s README, CONTRIBUTING.md, documentation, and benchmark harness (main branch), the GitHub API, PyPI/npm, and Hacker News results consulted on August 21–22, 2026. Benchmark scores are self-reported by the project via opendataloader-bench. Figures change over time.
Comments