The PDF Engine for RAG Pipelines
Best published harness score without ML. 8× faster than Docling and 0.1× faster than OpenDataLoader. Zero GPU, zero JVM — just a 15 MB Rust binary. No ML stack required for born-digital PDFs; optional in-browser OCR for image tables. Official odl-bench hybrid overall 0.900.
One Command. Instant PDF Intelligence.
pip install edgeparse — and your AI stack can read any PDF in milliseconds. No GPU warmup, no model downloads, no infrastructure.
# Step 1: Register the EdgeParse agent skill
npx skills add raphaelmansuy/edgeparse --skill edgeparse
# This adds to skills-lock.json:
# {
# "version": 1,
# "skills": {
# "edgeparse": {
# "source": "raphaelmansuy/edgeparse",
# "sourceType": "github"
# }
# }
# }
# Step 2: Install the Python runtime
pip install edgeparse # macOS / Linux — one-time setup
brew tap raphaelmansuy/tap
brew install edgeparse
# Verify installation
edgeparse --version
# Parse a PDF to Markdown
edgeparse report.pdf --format markdown
# Parse to JSON with bounding boxes
edgeparse invoice.pdf --format json
# Batch convert a directory
edgeparse docs/*.pdf --format markdown --output-dir results/ import edgeparse, json
# Convert PDF to Markdown
md = edgeparse.convert("report.pdf", format="markdown")
print(md[:500])
# Parse structured JSON with bounding boxes
doc = json.loads(edgeparse.convert("report.pdf", format="json"))
for el in doc["kids"][:3]:
print(el["type"], el.get("content", "")[:60])
# Save to output file
path = edgeparse.convert_file("report.pdf", output_dir="out/", format="markdown")
# Extract specific pages with table clustering
md = edgeparse.convert("report.pdf", pages="1-5", table_method="cluster") import { convert, convertFile } from "edgeparse";
// Convert PDF to Markdown
const md = convert("report.pdf", { format: "markdown" });
console.log(md.slice(0, 500));
// Parse structured JSON output
const doc = JSON.parse(convert("invoice.pdf", { format: "json" }));
doc.kids.slice(0, 3).forEach(el => console.log(el.type, el.content ?? ''));
// Extract specific pages
const pages = convert("report.pdf", { format: "markdown", pages: "1-5" });
// Save to output directory
const path = convertFile("report.pdf", { outputDir: "out/", format: "markdown" }); # Install via Homebrew (macOS / Linux)
brew tap raphaelmansuy/tap && brew install edgeparse
# Or via pip
pip install edgeparse
# Extract PDF to Markdown
edgeparse report.pdf --format markdown
# Extract to JSON with bounding boxes
edgeparse invoice.pdf --format json
# Batch convert entire directory
edgeparse docs/*.pdf --format markdown --output-dir results/ import init, { convert_to_string } from 'edgeparse-wasm';
// Load WASM binary (once)
await init();
// Read PDF from user upload
const bytes = new Uint8Array(await file.arrayBuffer());
// Extract Markdown — runs entirely in the browser
const markdown = convert_to_string(bytes, 'markdown');
// Extract structured JSON
const json = convert_to_string(bytes, 'json');
// Extract HTML
const html = convert_to_string(bytes, 'html');
// Try it live: edgeparse.com/demo/ Everything Your AI Stack Needs From a PDF
EdgeParse is the only PDF parser with ML-level accuracy that runs without ML — in Python, Node.js, the browser, and Rust.
8× Faster Than Docling
0.284 s/doc on Apple M4 Max. 2× faster than PyMuPDF4LLM and 0.1× faster than OpenDataLoader. Parallel per-page processing via Rayon — CPU only.
Best-in-Class Table Extraction
TEDS score of 0.562 on the EdgeParse harness. On the official odl-bench board, EdgeParse hybrid leads TEDS at 0.928. Ruling-line + borderless cluster detection with merged cell support.
Multi-Column Reading Order
XY-Cut++ reads multi-column layouts, sidebars, and mixed content in the correct logical order. NID score of 0.880 on the EdgeParse harness.
Full Document Hierarchy
Headings, paragraphs, lists, figures — all classified with nesting. MHS score of 0.478 on the EdgeParse harness; hybrid leads official odl-bench MHS at 0.836.
WebAssembly: Runs in the Browser
The only PDF parser with a WebAssembly build and @edgeparse/web SDK. Full Rust engine in the browser — PDF data never leaves the device. Optional in-browser OCR for image tables.
AI Safety Built-In
Filters hidden text, off-page content, tiny-text, and invisible layers — blocks prompt injection payloads embedded in PDFs before they reach your LLM.
Zero Dependencies
No ML stack required for born-digital PDFs; optional in-browser OCR for image tables. No GPU, no JVM, no Python runtime for the CLI. A single 15 MB binary. Deploy everywhere: Lambda, containers, edge functions, browsers.
5 SDK Languages
Native packages for Python (PyO3), Node.js (NAPI-RS), Rust, CLI binary via Homebrew/Cargo, and WebAssembly (@edgeparse/web). Pre-built wheels and addons — no compilation needed.
Bounding Boxes for Citations
Every element — paragraph, heading, table, image — includes [left, bottom, right, top] coordinates in PDF points. Cite exact sources in your RAG answers.
#1 Non-ML PDF Parser in Independent Benchmarks
Tested on 200 real-world PDFs — academic papers, financial reports, multi-column layouts, and complex tables. Running on Apple M4 Max.
| Tool | NID | TEDS | MHS | Overall | Speed |
|---|---|---|---|---|---|
| EdgeParse | 0.880 | 0.562 | 0.478 | 0.760 | 0.284 s/doc |
| EdgeParse [hybrid] | 0.863 | 0.603 | 0.506 | 0.767 | 0.810 s/doc |
| Docling (IBM) | 0.875 | 0.568 | 0.450 | 0.758 | 2.332 s/doc |
| OpenDataLoader | 0.873 | 0.320 | 0.441 | 0.733 | 0.022 s/doc |
| PyMuPDF4LLM | 0.860 | 0.509 | 0.411 | 0.732 | 0.640 s/doc |
| OpenDataLoader [hybrid] | 0.869 | 0.422 | 0.411 | 0.731 | 2.537 s/doc |
| MarkItDown | 0.807 | 0.193 | 0.001 | 0.564 | 0.189 s/doc |
EdgeParse harness snapshot updated 2026-10-02 (overall 0.760, 0.284 s/doc). Official odl-bench: EdgeParse hybrid leads at 0.900 overall. No ML stack required for born-digital PDFs; optional in-browser OCR for image tables.
Why Engineers Choose EdgeParse
EdgeParse is the only PDF engine that delivers near-ML accuracy without an ML stack — No ML stack required for born-digital PDFs; optional in-browser OCR for image tables. Just a 15 MB Rust binary.
| Feature | EdgeParse This project | OpenDataLoader Heuristic | Docling IBM | PyMuPDF4LLM PyMuPDF |
|---|---|---|---|---|
| Overall accuracy | 0.760 ✅ | 0.733 | 0.758 | 0.732 |
| Speed (s/doc) | 0.284 ✅ | 0.022 | 2.332 | 0.640 |
| Table extraction (TEDS) | 0.562 ✅ | 0.320 | 0.568 | 0.509 |
| Reading order (NID) | 0.880 ✅ | 0.873 | 0.875 | 0.860 |
| Heading detection (MHS) | 0.478 ✅ | 0.441 | 0.450 | 0.411 |
| Dependencies | ||||
| GPU required | ❌ None | ❌ None | ⚠️ Optional | ❌ None |
| OCR models required | ❌ Optional (image tables) | ⚠️ Optional | ✅ Required | ❌ None |
| Binary size | 15 MB ✅ | ~100 MB+ | ~500 MB+ | ~20 MB |
| SDK / Deployment | ||||
| Python SDK | ✅ | ✅ | ✅ | ✅ |
| Node.js / JavaScript SDK | ✅ | ❌ | ❌ | ❌ |
| WebAssembly (browser) | ✅ | ❌ | ❌ | ❌ |
| Rust native library | ✅ | ❌ | ❌ | ❌ |
| CLI binary | ✅ | ❌ | ❌ | ❌ |
| Safety & Privacy | ||||
| Prompt injection protection | ✅ | ✅ | ❌ | ❌ |
| In-browser (data never uploaded) | ✅ WASM | ❌ | ❌ | ❌ |
| Deterministic output | ✅ | ✅ | ❌ | ✅ |
| Bounding boxes (JSON) | ✅ | ✅ | ✅ | ❌ |
EdgeParse harness: 200 real-world PDFs on Apple M4 Max (overall 0.760, 0.284 s/doc). Snapshot 2026-10-02. Official odl-bench: EdgeParse hybrid leads at 0.900 overall (NID/TEDS/MHS/overall). Scores: NID = reading order, TEDS = table structure, MHS = heading hierarchy. Full methodology →
One Engine, Every AI Workflow
EdgeParse sits at the foundation of your AI stack — turning messy PDFs into clean, structured data that LLMs, agents, and RAG pipelines actually understand.
RAG Pipelines
Feed your vector database clean, hierarchically-chunked data with bounding boxes for source citation. No more garbled embeddings from raw PDF text.
# Chunk-ready output for your RAG pipeline
chunks = edgeparse.convert("report.pdf", format="json")
embeddings = embed(chunks) # Clean structured data AI Agents
Give your AI agents the ability to read, understand, and reason over any PDF document. Structured extraction means reliable tool use — no hallucinations.
# Agent tool: extract PDF intelligence
@tool("read_pdf")
def read_pdf(path: str) -> dict:
return edgeparse.convert(path, format="json") Copilot Skills
Build custom Copilot Skills and MCP servers that give AI assistants deep PDF understanding. Extract tables, headings, and metadata on demand.
# MCP server tool definition
@server.tool("extract_pdf")
async def extract(uri: str) -> str:
return edgeparse.convert(uri, format="md") Built for Real Production Workloads
Teams building RAG pipelines, legal tech, financial analysis, and browser apps choose EdgeParse for its speed, accuracy, and zero-dependency deployment.
RAG & Vector Search
Feed your vector database perfectly structured, hierarchical chunks with bounding boxes for source citation. Higher retrieval quality, better LLM answers.
Learn moreLegal & Compliance
Extract clauses, tables, and signature blocks from contracts and regulatory filings. Deterministic output means no surprises in production.
Financial Reports
Parse earnings reports, balance sheets, and SEC filings with accurate table extraction (TEDS 0.562 on harness; hybrid 0.928 on odl-bench) — columns, merged cells, and nested headers intact.
Research & Academic
Extract papers with correct multi-column reading order (NID 0.880) — figures, citations, and section hierarchy preserved for downstream analysis.
In-Browser Apps (WASM)
Full extraction in the browser via @edgeparse/web — no server, no uploads, privacy by design. Optional in-browser OCR for image tables. Works offline after first load.
Healthcare & Life Sciences
Process clinical notes, drug labels, and research protocols with AI safety filters that block prompt injection attacks embedded in uploaded PDFs.
Start Parsing PDFs in 30 Seconds
No API key. No cloud account. No GPU. Just install and parse.
pip install edgeparse Need enterprise deployment? Visit the Enterprise page or contact us for architecture reviews and production rollouts.