August 25, 2026 · By YasKad
opendataloader-project/opendataloader-pdf

OpenDataLoader PDF: a PDF parser for AI-ready data

opendataloader-project/opendataloader-pdf · 29,370★ · 2,799 forks

Everything worth knowing about opendataloader-project/opendataloader-pdf: an open-source PDF parser, written mostly in Java, that converts PDF documents into Markdown, JSON (with bounding boxes), and HTML for RAG and LLM pipelines, and also automates accessibility by tagging structure-less PDFs into Tagged PDF. It’s a Hancom Inc. (Korea) project developed alongside the PDF Association and Dual Lab (the creators of veraPDF).


Origin

The opendataloader-project organization was created on May 12, 2025 (the repository, on May 13, 2025), and its profile places it in South Korea, with 339 followers and 6 public repositories. The main contributor, bundolee (Bundo Lee), has accumulated 562 commits (versus 75 each for MaximPlusov and LonelyMidoriya), lists Seoul in their profile, and their company field reads “Hancom Inc. @opendataloader-project.” Hancom is a South Korean office and document software company (creator of the HWP format), which gives the project a clear corporate origin: it’s not an individual experiment but a bet placed by a company in the document sector.

Visible public traction started around September 2025, when the reference Hacker News thread (109 points, 28 comments) linked to the repository. The jump to v2.0 — which introduced the hybrid engine, the Apache 2.0 license, and Llama/LangChain integrations — spread in March 2026. The latest version measured in this run is v2.5.2, published August 21, 2026.

The README documents a relevant license change: versions before 2.0 were distributed under Mozilla Public License 2.0, and from 2.0 onward under Apache 2.0 — a decision the FAQ justifies as avoiding MPL’s per-file copyleft and easing enterprise adoption.

Philosophy and principles

The README summarizes the proposal in two blocks that define its principles:

  • AI-ready data, not just text: the goal is to produce structure (reading order, heading hierarchy, tables, coordinates) that LLMs and RAG pipelines need, not just plain text. The site’s message is explicit: “If your data isn’t parsed well, your RAG system will never retrieve correct answers. garbage in = garbage out.”
  • Local-first and cloud-free: the README and site insist everything runs locally, with no cloud calls or API keys; “your document never leaves your environment.” Version 2.0 frames this as usable in fully air-gapped environments, eliminating data leakage risk.
  • Local determinism + smart hybrid: local mode is deterministic (rules, no GPU, no models), and a hybrid mode routes only complex pages (difficult tables, scans, formulas, charts) to a local AI backend.
  • Coordinates for every element: each extracted element carries a [left, bottom, right, top] bounding box in PDF points, which enables source citations and “click-to-source” in RAG.
  • Accessibility as first-class, not an add-on: the project presents itself as the first open-source tool that generates Tagged PDF end-to-end, built on the PDF Association’s Well-Tagged PDF specification and validated with veraPDF.

Dark-mode cyberpunk visualization of local-first, air-gapped document processing. A secure local workstation sits in the foreground, with a PDF file icon feeding into a glowing Java engine and Python wrapper interface. No cloud icons, no external network cables; instead, all data flows remain inside a translucent neon containment field around the machine. Terminal windows display conversion commands and output formats: Markdown, JSON, HTML, and Tagged PDF. The environment feels private, deterministic, and enterprise-safe, with cool blue lighting, subtle grid lines, and a high-tech laboratory atmosphere. Ultra-detailed, cinematic, cyberpunk aesthetic, neon accents, 8K resolution.

How it works

The input is a PDF (digital, scanned, or already tagged), and the outputs are Markdown, JSON (with bounding boxes), HTML, Tagged PDF, and, as a paid add-on, PDF/UA. The core is a deterministic Java parser that applies layout analysis and an XY-Cut++ reading order (correct sequence across multi-column pages, sidebars, and mixed layouts).

Detailed conceptual image of a deterministic PDF layout parser using an XY-Cut++ reading-order algorithm. A complex multi-column document page is shown split by glowing vertical and horizontal slicing lines, with neon arrows tracing the correct reading order across columns, sidebars, footnotes, and mixed layouts. The page is rendered as a holographic wireframe in a dark tech workspace, with bounding boxes highlighting paragraphs, headings, tables, lists, and images. The visual conveys precision, speed, and rule-based intelligence, with cyan and violet neon outlines, a dark charcoal background, and futuristic diagnostic panels showing layout analysis. Ultra-detailed, cyberpunk tech aesthetic, 8K resolution.

Two modes documented in the README:

  1. Local (default, fast): opendataloader-pdf file1.pdf folder/. Extracts text with reading order, tables (borders), headings, lists, and images, and computes bounding boxes. The README states 60+ pages/second on CPU; with multi-threaded batch processing, over 100 pages/second on 8+ core machines (measured on Apple M4).
  2. Hybrid: combines the Java core with a local AI backend. Simple pages are processed locally; complex ones are sent to the backend to boost accuracy (the README cites a +90% improvement in table accuracy: from 0.489 to 0.928 TEDS). Requires starting a server (opendataloader-pdf-hybrid --port 5002) and calling the client with --hybrid docling-fast. Supports OCR (80+ languages) for scanned PDFs, LaTeX formula extraction, and chart/image description (a 256M SmolVLM model).

Futuristic split-screen image illustrating hybrid parsing mode. On the left, simple PDF pages are processed rapidly by a deterministic local Java engine, shown as clean neon pathways in a dark terminal. On the right, complex pages containing difficult tables, scanned text, mathematical formulas, and charts are routed through a glowing local AI backend server. The backend is represented by a compact local AI node with neural mesh patterns, not a cloud, processing OCR, LaTeX formula extraction, and image descriptions with a small SmolVLM-inspired visual model. Neon cyan lines separate fast local processing from precise AI enhancement, all within a secure offline environment. Dark cyberpunk style, ultra-detailed, 8K resolution.

Other documented mechanisms: Tagged PDF support, where if the PDF already has structure tags, it extracts the author’s exact layout with no heuristics; and auto-tagging into Tagged PDF, which analyzes the layout and generates structure tags for an untagged PDF, under Apache 2.0 (PDF/UA-1/2 conversion and the visual editor are paid add-ons).

High-detail, dark-mode cyberpunk data visualization of PDF extraction with bounding boxes and source citations. A glowing PDF page is overlaid with translucent neon rectangles marking text spans, tables, headings, and images. Each rectangle connects by thin luminous lines to JSON data cards displaying coordinates in the format of left, bottom, right, and top values. In the background, a RAG pipeline visual shows chunks being indexed and retrieved, with a "click-to-source" highlight linking an answer back to the exact PDF location. The scene emphasizes structured data, traceability, and AI-ready retrieval, using dark mode, electric blue and magenta accents, terminal UI elements, and ultra-detailed futuristic graphics. 8K resolution.

AI safety (anti-injection): automatically filters hidden text (transparent or zero-size fonts), off-page content, and suspicious invisible layers. With --sanitize it also replaces sensitive data (emails, URLs, phone numbers) with markers.

Security-themed dark cyberpunk image showing AI anti-injection protection for PDFs. A suspicious PDF contains hidden invisible text, transparent fonts, out-of-page content, and malicious prompt injections rendered as faint red ghost text and corrupted glyphs. A neon filter engine scans the document and removes or masks the threats, replacing sensitive data such as emails, URLs, and phone numbers with glowing redaction markers. The scene includes a secure firewall-like grid, diagnostic warnings, and a clean sanitized output document. Vigilant, precise, enterprise-grade mood, dark mode, red and cyan neon accents, ultra-detailed UI overlays, 8K resolution.

Quick-start guide

Installation and first run

Requirements: Java 11+ (check with java -version; if missing, install a JDK 11+ from Adoptium) and Python 3.10+.

pip install -U opendataloader-pdf
import opendataloader_pdf

# Batch all files into a single call: each convert() spawns a JVM process,
# so calling it in a loop is slow
opendataloader_pdf.convert(
    input_path=["file1.pdf", "file2.pdf", "folder/"],
    output_dir="output/",
    format="markdown,json"
)

For hybrid mode (complex tables, scans, formulas): pip install "opendataloader-pdf[hybrid]". There’s also a Node.js SDK (npm install @opendataloader/pdf) and Java (Maven org.opendataloader:opendataloader-pdf-core).

Common workflows

  • To extract to Markdown + JSON for RAG: opendataloader_pdf.convert(input_path=[...], output_dir="output/", format="markdown,json"); the result lands in output/ with Markdown ready to chunk and JSON with bounding boxes for citations.
  • For a scanned PDF or another language: start the backend with OCR — opendataloader-pdf-hybrid --port 5002 --force-ocr --ocr-lang "ko,en" — and process with opendataloader-pdf --hybrid docling-fast file.pdf.
  • For formulas or chart descriptions: backend with --enrich-formula or --enrich-picture-description, and client with --hybrid-mode full (formula and chart enrichment requires this mode).
  • To access an untagged PDF (accessibility): opendataloader-pdf --format tagged-pdf file.pdf, or in Python format="tagged-pdf".

Essential configuration

There’s no central configuration file: “configuration” means convert()/CLI parameters. The ones a new user touches first: format (markdown, json, html, pdf, text, or tagged-pdf; combinable), hybrid (enables hybrid mode), output_dir (output folder), use_struct_tree (uses the PDF’s native structure tags), and image_output/image_format (off, embedded, or external, with png/jpeg).

Common pitfalls and fixes

  • Every convert() spawns a JVM process: repeated calls in a loop are slow; batch all files into a single convert() call.
  • --use-struct-tree takes priority over --hybrid: if both are set on a tagged PDF, the structure tree is used and the hybrid backend isn’t called; to use the hybrid, drop --use-struct-tree.
  • Quality depends on tag quality: poorly tagged PDFs produce worse results; with sparse or incorrect tags, heuristic mode or --hybrid docling-fast tends to give better quality.
  • Formulas and chart descriptions require --hybrid-mode full on the client; without that mode, the enrichment isn’t applied.
  • Missing Java: if java -version doesn’t respond, the Python install works as a wrapper but the engine needs the JDK; install Java 11+.
  • Language limitation (community): an HN user (hermitcrab) noted that being Java/Python, it doesn’t fit those who need to call it from C++; there’s no documented native C++ binding.
  • Self-reported benchmark figures: the “number 1” scores come from the project’s own harness (opendataloader-bench); the repository publishes it so others can reproduce it, but it’s not an independent evaluation.

Integrations and migration

  • LangChain: officially documented integration — pip install -U langchain-opendataloader-pdf and OpenDataLoaderPDFLoader(file_path=[...], format="text").
  • LlamaIndex: official reader from the organization (opendataloader-project/opendataloader-pdf-llamaindex), “fast, accurate, local PDF reading.”
  • Local hybrid backend: requires no cloud or keys; the opendataloader-pdf-hybrid server runs on the same machine.
  • Migrating from other parsers: the README positions itself as a replacement for tools like pdf2docx (one HN user reports migrating “to replace pdf2docx … and it’s much better”). Against docling, marker, or pymupdf4llm, the README highlights default bounding boxes and injection filters, which those projects don’t offer by default. There’s no formal migration process: it’s reconfiguring the conversion call toward convert().

The ecosystem

Repositories from the same opendataloader-project organization: opendataloader-bench (83 stars, the project’s own evaluation harness, measuring reading order NID, table fidelity TEDS, and heading hierarchy MHS over a corpus of ~200 real PDFs), langchain-opendataloader-pdf (59 stars, LangChain document loader), opendataloader-pdf-llamaindex (3 stars, LlamaIndex reader), opendataloader-pdf-examples (2 stars, usage examples), and .github (1 star).

Broad ecosystem visualization of an open-source PDF parser integrated into modern AI tooling. In the center, a glowing PDF conversion core connects to multiple integration nodes: LangChain document loader, LlamaIndex reader, Node.js SDK, Java Maven artifact, Docker container, R interface, examples repository, and benchmark dashboard. The benchmark panel displays metrics for reading order, table fidelity, and heading hierarchy with neon graphs and leaderboard bars. Community forks and wrappers orbit around the main core as smaller connected nodes. The entire composition feels like a living open-source platform, with dark mode, cyberpunk neon lines, blue and magenta accents, terminal typography, and ultra-detailed futuristic architecture. 8K resolution.

Community projects and derivatives (GitHub repository search for “opendataloader”): NameetP/pdfmux (81 stars, Python, created March 2026) is a PDF extractor that “audits its own output and certifies other extractors’,” scoring 0.903 on opendataloader-bench (2nd of 8 engines) and exposing a 7-tool MCP server — the most notable related project, since it competes on the same benchmark as OpenDataLoader. muschellij2/opendataloader (0 stars, R, April 2026) is an R interface to OpenDataLoader PDF. coryisakson/opendataloader_docker (4 stars) is a self-hosted Docker package. dpaidev/opendataloader-pdf (2 stars) is a full mirror with history. JamoCA/cfml-OpenDataLoader-demo (1 star) and longshines/pdf-parser-online (0) are demos/web wrappers.

Most other results are mirrors, personal forks, or zero-star repos. Disambiguation note: AstunTechnology/OpenDataLoader (7 stars) is a different project (“a utility for loading OpenData into PostGIS”); its name match is accidental and unrelated to this repository.

Official / semi-official status

There’s no vendor “marketplace” in the style of agent plugins, but the project has solid semi-official status backed by standards bodies and corporate support: collaboration with standards organizations (the README and site state that auto-tagging was built in collaboration with the PDF Association, author of the Well-Tagged PDF specification, and with Dual Lab, developers of veraPDF, the open reference validator for PDF/A and PDF/UA); Hancom Inc. backing (the main contributor lists Hancom as their company; the README announces the enterprise integration Hancom Data Loader — document analysis with custom models, OCR with SLA, and native HWP/HWPX support — as a next step, Q2–Q3 2026 roadmap); “official” third-party integrations (it appears in LangChain’s official integration docs and maintains an official LlamaIndex reader); and an open-core model (the core — extraction, layout analysis, auto-tagging to Tagged PDF — is Apache 2.0 and free; the enterprise add-ons — PDF/UA-1/2 export and the visual accessibility editor/studio — are paid on request).

The project self-describes as “the only open-source parser” that combines local deterministic extraction, per-element bounding boxes, XY-Cut++, AI safety filters, and Tagged PDF support, and positions itself as “number 1 on the benchmark.” That’s a project claim backed by its own harness (opendataloader-bench), not a formal standard designation; no independent external certification was found in this run.

Dramatic accessibility-focused cyberpunk image showing an untagged PDF being transformed into a Tagged PDF. A rough, structureless document page enters a glowing semantic tagging engine and emerges as a clean, organized semantic tree with labeled nodes for headings, paragraphs, lists, tables, figures, and reading order. A validation badge inspired by PDF Association and veraPDF appears as a luminous seal of correctness. The visual emphasizes accessibility as a first-class feature, with warm amber highlights for structure labels, cool blue data streams, and a dark high-tech laboratory background. Ultra-detailed, professional tech illustration, neon accents, 8K resolution.

Repo numbers

Measured: August 21–22, 2026, GitHub API.

MetricValue
Stars28,641
Forks2,734
Real subscribers111
Commits861
Open issues + PRs77
Primary languageJava
LicenseApache-2.0 (pre-2.0: MPL 2.0)
CreatedMay 13, 2025
Last pushAugust 21, 2026
Latest releasev2.5.2, August 21, 2026

Top contributors by contributions: bundolee (562), MaximPlusov (75), LonelyMidoriya (75), hyunhee-jo (58), hnc-jglee (19), dependabot[bot] (16), suji-cho (16).

Registry downloads: PyPI opendataloader-pdf (latest version 2.5.2, 69 releases) with 44,408 downloads in the last week and 191,684 in the last month. npm @opendataloader/pdf with 54,297 downloads in the July 22–August 22, 2026 range. Java distribution is via Maven (org.opendataloader:opendataloader-pdf-core).

watchers_count mirrors the star count, so subscribers_count is reported as the real subscriber figure. open_issues_count includes open change requests, not only issues. Registry download figures are delivery metrics, not a count of users or unique deployments.

How the community received it

Verifiable conversation concentrates on one Hacker News thread; the rest of the submissions have very little traction.

Hacker News (109 points, 28 comments, “OpenDataLoader-PDF: An open source tool for structured PDF parsing,” September 23, 2025) gathered concrete opinions. 4d66ba06 (praise and adoption): “just finished migrating to this to replace pdf2docx in a project I was working on and it’s much better. Thanks for open-sourcing OpenDataLoader-PDF.” emilburzo (practical): tested it with bank statements, “PDFs that are surprisingly hard… the JSON extraction looks pretty good and seems to produce something usable in a single pass.” fedeb95 (non-AI use): “Very cool. I’ll probably use it, but not for AI. I have a lot of PDFs for which no epub exists.” agsqwe asked how it compares to docling. hermitcrab (technical objection): “got excited until I read it was Java/Python. Looking for a library that extracts PDF tables and can be called from a C++ program.” trevor-e (deeper critique) argues that “maybe we need a new AI-friendly file format instead of continuing to patch over the complicated PDF spec.” constantinum mentioned Unstract (Zipstack/unstract) as an open structured-extraction + ETL alternative.

Other low-traction submissions don’t constitute a broad review: “Show HN: OpenDataLoader – Safe, Open, High-Performance PDF Loader for AI” (4 points), one about v2.0 (1 point), and “LLM Is Not a PDF Parser: Use OpenDataLoader First” (2 points). No additional verifiable community evidence was recovered outside HN in this run.

OpenDataLoader PDF versus other projects

The table values come from the project’s own harness, opendataloader-bench (reading order NID, table TEDS, heading MHS, over ~200 PDFs); they’re self-reported by the project, not an independent evaluation. Speed is in seconds/page (lower is better).

ProjectOverall score (own bench)Speed (s/page)Verifiable difference
opendataloader [hybrid]0.9070.463Local hybrid engine; per-element bounding boxes; injection filters; Apache 2.0.
nutrient0.8850.008Commercial engine; the fastest in the harness.
docling0.8820.762MIT; per the README, lacks bounding boxes and injection filters by default.
marker0.86153.932GPL-3.0; the README describes it as much slower (≈54 s/page).
unstructured [hi_res]0.8413.008Apache-2.0; standard unstructured also exists (0.686).
edgeparse0.8370.036Apache-2.0.
opendataloader (local mode)0.8310.015The non-AI variant, very fast; worse on tables (0.489).
mineru0.8315.962AGPL-3.0 (opendatalab/MinerU).
pymupdf4llm0.7320.091AGPL-3.0; fast but worse on tables (0.401) and headings (0.412).
markitdown0.5890.114MIT.
liteparse0.5761.061Apache-2.0.

In practice, the differentiator the project itself claims is the combination of local deterministic extraction + per-element bounding boxes + auto-tagging to Tagged PDF; a commercial engine (nutrient) or a fast parser (pymupdf4llm) can win on speed, and a GPU-based parser (marker, docling) is competitive on quality, but per the project’s own harness, none combine all those capabilities in local mode.

Contributing

CONTRIBUTING.md documents an explicit process: fork the repository and clone the fork; create a working branch (git checkout -b my-feature); build the project (prerequisites: Java 11+, Maven, Python 3.10+, uv, Node.js 24 active LTS, and pnpm via corepack enable pnpm; CI builds against Node 24 and pnpm 11.21.0, and Node must be ≥22.13); for questions, open an issue tagged Question; for bugs, use the Bug Report template; for proposals, the Feature Request template.

Use cases and who this repository can help

  • Teams building RAG/LLM pipelines over PDF documents: extraction to Markdown (reading order, tables, headings) and to JSON with bounding boxes enables semantic chunking and “click-to-source” citations. The local hybrid mode covers complex tables and scans without sending data to the cloud — relevant for sectors with sensitive data (legal, healthcare, finance).
  • Organizations that must comply with accessibility regulations (EU EAA, US ADA/Section 508, Korea’s Digital Inclusion Act): the audit → auto-tag → Tagged PDF pipeline (free, Apache 2.0) replaces part of the manual remediation work (which the README quantifies at $50–200 USD per document); veraPDF validation makes it suitable for compliance teams.
  • Processing scanned or multilingual PDFs: hybrid OCR (80+ languages) serves anyone converting digitized files or documents in Korean/Japanese/Chinese/Arabic; Hancom’s Korean origin explains the announced native HWP support.
  • Scientists and technical staff extracting formulas and charts: LaTeX formula extraction and chart/image description cover scientific papers and technical documents where table and equation quality matters.
  • Developers integrating into existing frameworks: official LangChain and LlamaIndex loaders let you plug OpenDataLoader in as a document source in already-built pipelines.
  • Who is NOT its audience: those needing a native C++ binding, those requiring a fully independent validation benchmark, or those needing PDF/UA without the paid add-on.

Resources


Note: this report combines the repository’s README, CONTRIBUTING.md, documentation, and benchmark harness (main branch), the GitHub API, PyPI/npm, and Hacker News results consulted on August 21–22, 2026. Benchmark scores are self-reported by the project via opendataloader-bench. Figures change over time.

Comments