Skip to content
#1 Non-ML PDF Parser Leads the current harness · 8× faster than Docling · No ML stack required for born-digital PDFs; optional in-browser OCR for image tables

The PDF Engine for RAG Pipelines

Best published harness score without ML. 8× faster than Docling and 0.1× faster than OpenDataLoader. Zero GPU, zero JVM — just a 15 MB Rust binary. No ML stack required for born-digital PDFs; optional in-browser OCR for image tables. Official odl-bench hybrid overall 0.900.

pip install edgeparse
0+ docs/sec
0% accuracy
0 ML dependencies
0 SDK languages
Works with
Python Node.js Rust CLI WebAssembly

One Command. Instant PDF Intelligence.

pip install edgeparse — and your AI stack can read any PDF in milliseconds. No GPU warmup, no model downloads, no infrastructure.

# Step 1: Register the EdgeParse agent skill
npx skills add raphaelmansuy/edgeparse --skill edgeparse

# This adds to skills-lock.json:
# {
#   "version": 1,
#   "skills": {
#     "edgeparse": {
#       "source": "raphaelmansuy/edgeparse",
#       "sourceType": "github"
#     }
#   }
# }

# Step 2: Install the Python runtime
pip install edgeparse
# macOS / Linux — one-time setup
brew tap raphaelmansuy/tap
brew install edgeparse

# Verify installation
edgeparse --version

# Parse a PDF to Markdown
edgeparse report.pdf --format markdown

# Parse to JSON with bounding boxes
edgeparse invoice.pdf --format json

# Batch convert a directory
edgeparse docs/*.pdf --format markdown --output-dir results/
import edgeparse, json

# Convert PDF to Markdown
md = edgeparse.convert("report.pdf", format="markdown")
print(md[:500])

# Parse structured JSON with bounding boxes
doc = json.loads(edgeparse.convert("report.pdf", format="json"))
for el in doc["kids"][:3]:
  print(el["type"], el.get("content", "")[:60])

# Save to output file
path = edgeparse.convert_file("report.pdf", output_dir="out/", format="markdown")

# Extract specific pages with table clustering
md = edgeparse.convert("report.pdf", pages="1-5", table_method="cluster")
import { convert, convertFile } from "edgeparse";

// Convert PDF to Markdown
const md = convert("report.pdf", { format: "markdown" });
console.log(md.slice(0, 500));

// Parse structured JSON output
const doc = JSON.parse(convert("invoice.pdf", { format: "json" }));
doc.kids.slice(0, 3).forEach(el => console.log(el.type, el.content ?? ''));

// Extract specific pages
const pages = convert("report.pdf", { format: "markdown", pages: "1-5" });

// Save to output directory
const path = convertFile("report.pdf", { outputDir: "out/", format: "markdown" });
# Install via Homebrew (macOS / Linux)
brew tap raphaelmansuy/tap && brew install edgeparse

# Or via pip
pip install edgeparse

# Extract PDF to Markdown
edgeparse report.pdf --format markdown

# Extract to JSON with bounding boxes
edgeparse invoice.pdf --format json

# Batch convert entire directory
edgeparse docs/*.pdf --format markdown --output-dir results/
import init, { convert_to_string } from 'edgeparse-wasm';

// Load WASM binary (once)
await init();

// Read PDF from user upload
const bytes = new Uint8Array(await file.arrayBuffer());

// Extract Markdown — runs entirely in the browser
const markdown = convert_to_string(bytes, 'markdown');

// Extract structured JSON
const json = convert_to_string(bytes, 'json');

// Extract HTML
const html = convert_to_string(bytes, 'html');

// Try it live: edgeparse.com/demo/
Features

Everything Your AI Stack Needs From a PDF

EdgeParse is the only PDF parser with ML-level accuracy that runs without ML — in Python, Node.js, the browser, and Rust.

8× Faster Than Docling

0.284 s/doc on Apple M4 Max. 2× faster than PyMuPDF4LLM and 0.1× faster than OpenDataLoader. Parallel per-page processing via Rayon — CPU only.

Best-in-Class Table Extraction

TEDS score of 0.562 on the EdgeParse harness. On the official odl-bench board, EdgeParse hybrid leads TEDS at 0.928. Ruling-line + borderless cluster detection with merged cell support.

Multi-Column Reading Order

XY-Cut++ reads multi-column layouts, sidebars, and mixed content in the correct logical order. NID score of 0.880 on the EdgeParse harness.

Full Document Hierarchy

Headings, paragraphs, lists, figures — all classified with nesting. MHS score of 0.478 on the EdgeParse harness; hybrid leads official odl-bench MHS at 0.836.

WebAssembly: Runs in the Browser

The only PDF parser with a WebAssembly build and @edgeparse/web SDK. Full Rust engine in the browser — PDF data never leaves the device. Optional in-browser OCR for image tables.

AI Safety Built-In

Filters hidden text, off-page content, tiny-text, and invisible layers — blocks prompt injection payloads embedded in PDFs before they reach your LLM.

Zero Dependencies

No ML stack required for born-digital PDFs; optional in-browser OCR for image tables. No GPU, no JVM, no Python runtime for the CLI. A single 15 MB binary. Deploy everywhere: Lambda, containers, edge functions, browsers.

5 SDK Languages

Native packages for Python (PyO3), Node.js (NAPI-RS), Rust, CLI binary via Homebrew/Cargo, and WebAssembly (@edgeparse/web). Pre-built wheels and addons — no compilation needed.

Bounding Boxes for Citations

Every element — paragraph, heading, table, image — includes [left, bottom, right, top] coordinates in PDF points. Cite exact sources in your RAG answers.

#1 Non-ML PDF Parser in Independent Benchmarks

Tested on 200 real-world PDFs — academic papers, financial reports, multi-column layouts, and complex tables. Running on Apple M4 Max.

EdgeParse
76.0%
0.284 s/doc
EdgeParse [hybrid]
76.7%
0.810 s/doc
Docling (IBM)
75.8%
2.332 s/doc
OpenDataLoader
73.3%
0.022 s/doc
PyMuPDF4LLM
73.2%
0.640 s/doc
OpenDataLoader [hybrid]
73.1%
2.537 s/doc
MarkItDown
56.4%
0.189 s/doc
Tool NID TEDS MHS Overall Speed
EdgeParse 0.880 0.562 0.478 0.760 0.284 s/doc
EdgeParse [hybrid] 0.863 0.603 0.506 0.767 0.810 s/doc
Docling (IBM) 0.875 0.568 0.450 0.758 2.332 s/doc
OpenDataLoader 0.873 0.320 0.441 0.733 0.022 s/doc
PyMuPDF4LLM 0.860 0.509 0.411 0.732 0.640 s/doc
OpenDataLoader [hybrid] 0.869 0.422 0.411 0.731 2.537 s/doc
MarkItDown 0.807 0.193 0.001 0.564 0.189 s/doc

EdgeParse harness snapshot updated 2026-10-02 (overall 0.760, 0.284 s/doc). Official odl-bench: EdgeParse hybrid leads at 0.900 overall. No ML stack required for born-digital PDFs; optional in-browser OCR for image tables.

Head-to-Head Comparison

Why Engineers Choose EdgeParse

EdgeParse is the only PDF engine that delivers near-ML accuracy without an ML stack — No ML stack required for born-digital PDFs; optional in-browser OCR for image tables. Just a 15 MB Rust binary.

EdgeParse
0.760
Overall (EdgeParse harness)
0.284 s/doc · CPU only
No GPU No ML stack No JVM WebAssembly 5 SDKs
OpenDataLoader
0.733
Fast heuristic pipeline
0.022 s/doc · 0.1× slower
Python only No WASM
IBM Docling
0.758
Requires OCR / ML stack
2.332 s/doc · 8× slower
Needs OCR Heavy setup
Feature EdgeParse This project OpenDataLoader Heuristic Docling IBM PyMuPDF4LLM PyMuPDF
Overall accuracy 0.760 ✅ 0.733 0.758 0.732
Speed (s/doc) 0.284 ✅ 0.022 2.332 0.640
Table extraction (TEDS) 0.562 ✅ 0.320 0.568 0.509
Reading order (NID) 0.880 ✅ 0.873 0.875 0.860
Heading detection (MHS) 0.478 ✅ 0.441 0.450 0.411
Dependencies
GPU required ❌ None ❌ None ⚠️ Optional ❌ None
OCR models required ❌ Optional (image tables) ⚠️ Optional ✅ Required ❌ None
Binary size 15 MB ✅ ~100 MB+ ~500 MB+ ~20 MB
SDK / Deployment
Python SDK ✅ ✅ ✅ ✅
Node.js / JavaScript SDK ✅ ❌ ❌ ❌
WebAssembly (browser) ✅ ❌ ❌ ❌
Rust native library ✅ ❌ ❌ ❌
CLI binary ✅ ❌ ❌ ❌
Safety & Privacy
Prompt injection protection ✅ ✅ ❌ ❌
In-browser (data never uploaded) ✅ WASM ❌ ❌ ❌
Deterministic output ✅ ✅ ❌ ✅
Bounding boxes (JSON) ✅ ✅ ✅ ❌

EdgeParse harness: 200 real-world PDFs on Apple M4 Max (overall 0.760, 0.284 s/doc). Snapshot 2026-10-02. Official odl-bench: EdgeParse hybrid leads at 0.900 overall (NID/TEDS/MHS/overall). Scores: NID = reading order, TEDS = table structure, MHS = heading hierarchy. Full methodology →

AI Integration

One Engine, Every AI Workflow

EdgeParse sits at the foundation of your AI stack — turning messy PDFs into clean, structured data that LLMs, agents, and RAG pipelines actually understand.

Any PDF
EdgeParse
Structured Data

RAG Pipelines

Feed your vector database clean, hierarchically-chunked data with bounding boxes for source citation. No more garbled embeddings from raw PDF text.

# Chunk-ready output for your RAG pipeline
chunks = edgeparse.convert("report.pdf", format="json")
embeddings = embed(chunks) # Clean structured data
Embeddings Citations LangChain LlamaIndex

Copilot Skills

Build custom Copilot Skills and MCP servers that give AI assistants deep PDF understanding. Extract tables, headings, and metadata on demand.

# MCP server tool definition
@server.tool("extract_pdf")
async def extract(uri: str) -> str:
  return edgeparse.convert(uri, format="md")
MCP Copilot ChatGPT Claude
Use Cases

Built for Real Production Workloads

Teams building RAG pipelines, legal tech, financial analysis, and browser apps choose EdgeParse for its speed, accuracy, and zero-dependency deployment.

RAG & Vector Search

Feed your vector database perfectly structured, hierarchical chunks with bounding boxes for source citation. Higher retrieval quality, better LLM answers.

LangChainLlamaIndexEmbeddingsCitations
Learn more

Legal & Compliance

Extract clauses, tables, and signature blocks from contracts and regulatory filings. Deterministic output means no surprises in production.

ContractsComplianceAudit Trail

Financial Reports

Parse earnings reports, balance sheets, and SEC filings with accurate table extraction (TEDS 0.562 on harness; hybrid 0.928 on odl-bench) — columns, merged cells, and nested headers intact.

SEC FilingsEarningsTablesJSON

Research & Academic

Extract papers with correct multi-column reading order (NID 0.880) — figures, citations, and section hierarchy preserved for downstream analysis.

arXivMulti-columnCitations

In-Browser Apps (WASM)

Full extraction in the browser via @edgeparse/web — no server, no uploads, privacy by design. Optional in-browser OCR for image tables. Works offline after first load.

WebAssemblyPrivacyOfflineReact/Vue

Healthcare & Life Sciences

Process clinical notes, drug labels, and research protocols with AI safety filters that block prompt injection attacks embedded in uploaded PDFs.

HIPAASafetyStructured Data

Start Parsing PDFs in 30 Seconds

No API key. No cloud account. No GPU. Just install and parse.

pip install edgeparse

Need enterprise deployment? Visit the Enterprise page or contact us for architecture reviews and production rollouts.