/// EXTRACT
2 credits

PDF to Markdown API — clean input for LLMs

Headings, lists, tables. Feed straight to an LLM.

/// Quick answer

PDF to Markdown converts a document into clean Markdown while preserving its structure — headings, lists, tables and emphasis. This matters most when feeding documents to language models: plain text extraction discards structure, and losing it measurably degrades retrieval quality because the model can no longer distinguish a heading from a sentence. Scanned PDFs require OCR first, since they contain images rather than text.

Try PDF to Markdown API — free →Read the docs

Free account. No card required.

How does the PDF to Markdown API work?

One HTTP request from any language. Every ASHDOCS endpoint uses the same X-API-Key header.

curl -X POST https://www.ashdocs.com/api/v1/tools/pdf-to-markdown \
  -H "X-API-Key: ash_live_..." \
  -F "files=@whitepaper.pdf"

How do I use PDF to Markdown API in Make, Zapier or n8n?

MAKE

Add an HTTP "Make a request" module and POST the file. Leave Parse response on — the Markdown comes back as JSON, not a binary file.

Make setup guide →
ZAPIER

Use Webhooks by Zapier with a POST action and map the returned Markdown into your vector store, a document field, or a downstream processing step.

Zapier setup guide →
n8n

Use the HTTP Request node with Response Format as JSON, then pass the Markdown into a chunking or embedding node.

n8n setup guide →
AIRTABLE

Trigger on a new attachment, convert it, and store the Markdown in a long-text field for later processing.

Airtable setup guide →

Common problems and fixes

SymptomCauseFix
The output is emptyThe PDF is scanned and has no text layerRun OCR first to create a text layer, then convert. A text extraction returning almost nothing on a visibly text-filled page confirms this.
Columns interleave into nonsenseA multi-column layout read straight acrossUse layout-aware extraction. Academic papers and newsletters are the common cases — test one early, because the output reads plausibly while being wrong.
Headers and footers repeat throughoutPage furniture is extracted as body contentStrip repeating elements before chunking. A company name repeated 200 times pollutes every embedding.
Complex tables become unreadableMerged cells and nested headers do not map to Markdown pipe syntaxExtract complex tables separately as structured data and reference them, rather than forcing them into Markdown.
Searches miss words containing fi or flSome PDFs encode ligatures as single glyphsNormalise Unicode after extraction, or a search for "workflow" will miss the ligature form.
Footnotes interrupt sentencesThey sit at the page bottom but belong to a specific sentenceMove them to the end of their section, or drop them. Leaving them mid-flow corrupts the surrounding chunk.

Common use cases

Frequently asked questions

Why convert PDF to Markdown instead of plain text?+

Markdown preserves headings, lists, tables and emphasis, which plain text discards. That structure gives each chunk semantic context and measurably improves retrieval quality in RAG pipelines, at very small token cost.

Can I convert a scanned PDF to Markdown?+

Not directly — a scanned PDF contains images rather than text. Run OCR first to produce a text layer, then convert.

Is Markdown the right format for RAG?+

For most retrieval pipelines, yes. It preserves structure at low token cost, and language models handle it natively because they were trained on large volumes of it. JSON with full layout data is more precise but consumes far more context.

How should I chunk the Markdown output?+

Split on heading boundaries rather than fixed character counts, aim for roughly 200 to 800 tokens per chunk, overlap by 10 to 15 percent, and prepend the section heading path to each chunk.

Does converting to Markdown lose information?+

It discards visual formatting — fonts, colours, exact positioning — and keeps semantic structure. For retrieval that is the correct trade, since models reason over meaning rather than appearance.

RELATED TOOLS