Batch extraction, without the cleanup
Drop in documents or a whole folder. DocDistill writes collision-safe output names and one extraction-manifest.json per batch, so every export remains traceable and no file is silently overwritten.
Private document extraction for macOS
Batch-convert PDFs, Word, Excel, PowerPoint, HTML, and email into clean Markdown, tables, JSON, or retrieval-ready JSONL. Entirely offline.
USD 29.99 · one-time purchase
No subscription. No account.
{
"source": "research-deck.pptx",
"sha256": "83f9…b612",
"bytes": 2849134,
"output": "research-deck.jsonl",
"chunks": 42,
"engine": "docdistill-1.0"
}Built for dependable pipelines
Drop in documents or a whole folder. DocDistill writes collision-safe output names and one extraction-manifest.json per batch, so every export remains traceable and no file is silently overwritten.
Chunks break on the document’s own headings—never mid-sentence. Oversized paragraphs split at sentence boundaries, and DocDistill can repeat the full heading path inside every chunk for retrieval context.
Every result records the source file’s SHA-256 and byte count, plus engine revision, app version, and extraction time. You can prove exactly what was processed, and by which build.
A scanned PDF with no text layer is reported as “No readable text.” It is never exported as a plausible-looking empty file. DocDistill tells you when it cannot extract what is not there.
Input → useful output
READS
PDF · Word · Excel · PowerPoint · OpenDocument · HTML · XML · email · plain text
WRITES
Clean Markdown · tables · JSON · retrieval-ready JSONL
One Mac app. One fair price.
No subscription, no account, no cloud dependency. Requires macOS 14 or later. Universal for Apple silicon and Intel.