Skip to content

Quick Start: Web SDK

@edgeparse/web is the product Web SDK for EdgeParse 0.3.0. It wraps the WASM engine with:

  • Observable client state (subscribe / getSnapshot for React useSyncExternalStore)
  • Two-phase ParseSession (plan → async OCR → assemble)
  • Lazy, consent-gated PP-OCR model downloads
  • Graceful degradation when OCR is declined or offline

On the official odl-bench board, EdgeParse WASM scores 0.854 overall (TEDS 0.776).

Prefer this path for apps. The low-level edgeparse-wasm package remains available for direct convert_to_string use.

Terminal window
npm install @edgeparse/web edgeparse-wasm
# or
pnpm add @edgeparse/web edgeparse-wasm
import { EdgeParse } from '@edgeparse/web';
const ep = await EdgeParse.create({
models: 'lazy', // download OCR models only when needed
ocr: 'small',
onBeforeDownload: async (model) => {
// Show consent UI; return false to degrade without OCR
return confirm(`Download ${model.id} (~${(model.bytes / 1e6).toFixed(1)} MB)?`);
},
});
ep.subscribe(() => {
const snap = ep.getSnapshot();
// snap.engine / snap.models / snap.jobs — React-friendly
});
const job = ep.parse(file, { format: 'markdown', tableMethod: 'cluster' });
job.on('progress', (p) => console.log(p.fraction, p.label));
const result = await job.result;
// result.quality: 'full' | 'degraded' | 'skipped'
console.log(result.markdown);
PolicyBehavior
lazy (default)Download when a parse needs OCR
preloadFetch the selected OCR tier at create()
manualCall ep.models.preload() / ensure() yourself
offNever download models; image tables degrade to PDF text
  1. Parse worker — ParseSession (Rust/WASM): extract OCR candidates, finish with OCR words
  2. OCR worker — onnxruntime-web (WebGPU → WASM+SIMD), PP-OCRv6 tiers
  3. ModelManager — pinned manifest, OPFS/Cache API, sha256 verify, Web Locks

Missing OCR never fails a parse — cells fall back to PDF text and quality: "degraded".

Try it: edgeparse.com/demo/