2 creditsPDF to Markdown API — clean input for LLMs
Headings, lists, tables. Feed straight to an LLM.
PDF to Markdown converts a document into clean Markdown while preserving its structure — headings, lists, tables and emphasis. This matters most when feeding documents to language models: plain text extraction discards structure, and losing it measurably degrades retrieval quality because the model can no longer distinguish a heading from a sentence. Scanned PDFs require OCR first, since they contain images rather than text.
How does the PDF to Markdown API work?
One HTTP request from any language. Every ASHDOCS endpoint uses the same X-API-Key header.
curl -X POST https://www.ashdocs.com/api/v1/tools/pdf-to-markdown \ -H "X-API-Key: ash_live_..." \ -F "files=@whitepaper.pdf"
import requests
r = requests.post(
"https://www.ashdocs.com/api/v1/tools/pdf-to-markdown",
headers={"X-API-Key": "ash_live_..."},
files=[("files", open("whitepaper.pdf", "rb"))],
)
job = r.json()
print(job["status"], job["output_files"][0]["url"])import { readFile } from "node:fs/promises";
const form = new FormData();
form.append("files", new Blob([await readFile("whitepaper.pdf")]), "whitepaper.pdf");
const res = await fetch("https://www.ashdocs.com/api/v1/tools/pdf-to-markdown", {
method: "POST",
headers: { "X-API-Key": "ash_live_..." },
body: form,
});
const job = await res.json();
console.log(job.status, job.output_files[0].url);using var client = new HttpClient();
client.DefaultRequestHeaders.Add("X-API-Key", "ash_live_...");
using var form = new MultipartFormDataContent();
form.Add(new ByteArrayContent(await File.ReadAllBytesAsync("whitepaper.pdf")), "files", "whitepaper.pdf");
var res = await client.PostAsync("https://www.ashdocs.com/api/v1/tools/pdf-to-markdown", form);
var job = await res.Content.ReadAsStringAsync();
Console.WriteLine(job);import java.io.ByteArrayOutputStream;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
public class Sample {
static final String BOUNDARY = "AshdocsBoundary7MA4YWxkTrZu0gW";
public static void main(String[] args) throws Exception {
ByteArrayOutputStream body = new ByteArrayOutputStream();
addFile(body, "files", "whitepaper.pdf");
body.write(("--" + BOUNDARY + "--\r\n").getBytes(StandardCharsets.UTF_8));
HttpRequest req = HttpRequest.newBuilder(URI.create("https://www.ashdocs.com/api/v1/tools/pdf-to-markdown"))
.header("X-API-Key", "ash_live_...")
.header("Content-Type", "multipart/form-data; boundary=" + BOUNDARY)
.POST(HttpRequest.BodyPublishers.ofByteArray(body.toByteArray()))
.build();
HttpResponse<String> res = HttpClient.newHttpClient().send(req, HttpResponse.BodyHandlers.ofString());
System.out.println(res.body());
}
static void addFile(ByteArrayOutputStream body, String name, String path) throws Exception {
body.write(("--" + BOUNDARY + "\r\nContent-Disposition: form-data; name=\"" + name + "\"; filename=\""
+ Path.of(path).getFileName() + "\"\r\nContent-Type: application/octet-stream\r\n\r\n").getBytes(StandardCharsets.UTF_8));
body.write(Files.readAllBytes(Path.of(path)));
body.write("\r\n".getBytes(StandardCharsets.UTF_8));
}
static void addField(ByteArrayOutputStream body, String name, String value) throws Exception {
body.write(("--" + BOUNDARY + "\r\nContent-Disposition: form-data; name=\"" + name + "\"\r\n\r\n"
+ value + "\r\n").getBytes(StandardCharsets.UTF_8));
}
}#include <stdio.h>
#include <curl/curl.h>
int main(void) {
CURL *curl = curl_easy_init();
if (!curl) return 1;
struct curl_slist *headers = NULL;
headers = curl_slist_append(headers, "X-API-Key: ash_live_...");
curl_mime *mime = curl_mime_init(curl);
curl_mimepart *part;
part = curl_mime_addpart(mime);
curl_mime_name(part, "files");
curl_mime_filedata(part, "whitepaper.pdf");
curl_easy_setopt(curl, CURLOPT_MIMEPOST, mime);
curl_easy_setopt(curl, CURLOPT_URL, "https://www.ashdocs.com/api/v1/tools/pdf-to-markdown");
curl_easy_setopt(curl, CURLOPT_HTTPHEADER, headers);
CURLcode rc = curl_easy_perform(curl);
printf("\n");
curl_mime_free(mime);
curl_slist_free_all(headers);
curl_easy_cleanup(curl);
return rc == CURLE_OK ? 0 : 1;
}How do I use PDF to Markdown API in Make, Zapier or n8n?
Add an HTTP "Make a request" module and POST the file. Leave Parse response on — the Markdown comes back as JSON, not a binary file.
Make setup guide →Use Webhooks by Zapier with a POST action and map the returned Markdown into your vector store, a document field, or a downstream processing step.
Zapier setup guide →Use the HTTP Request node with Response Format as JSON, then pass the Markdown into a chunking or embedding node.
n8n setup guide →Trigger on a new attachment, convert it, and store the Markdown in a long-text field for later processing.
Airtable setup guide →Common problems and fixes
| Symptom | Cause | Fix |
|---|---|---|
| The output is empty | The PDF is scanned and has no text layer | Run OCR first to create a text layer, then convert. A text extraction returning almost nothing on a visibly text-filled page confirms this. |
| Columns interleave into nonsense | A multi-column layout read straight across | Use layout-aware extraction. Academic papers and newsletters are the common cases — test one early, because the output reads plausibly while being wrong. |
| Headers and footers repeat throughout | Page furniture is extracted as body content | Strip repeating elements before chunking. A company name repeated 200 times pollutes every embedding. |
| Complex tables become unreadable | Merged cells and nested headers do not map to Markdown pipe syntax | Extract complex tables separately as structured data and reference them, rather than forcing them into Markdown. |
| Searches miss words containing fi or fl | Some PDFs encode ligatures as single glyphs | Normalise Unicode after extraction, or a search for "workflow" will miss the ligature form. |
| Footnotes interrupt sentences | They sit at the page bottom but belong to a specific sentence | Move them to the end of their section, or drop them. Leaving them mid-flow corrupts the surrounding chunk. |
Common use cases
- →RAG ingestion
- →Static-site conversion
- →AI training data
Frequently asked questions
Why convert PDF to Markdown instead of plain text?+
Markdown preserves headings, lists, tables and emphasis, which plain text discards. That structure gives each chunk semantic context and measurably improves retrieval quality in RAG pipelines, at very small token cost.
Can I convert a scanned PDF to Markdown?+
Not directly — a scanned PDF contains images rather than text. Run OCR first to produce a text layer, then convert.
Is Markdown the right format for RAG?+
For most retrieval pipelines, yes. It preserves structure at low token cost, and language models handle it natively because they were trained on large volumes of it. JSON with full layout data is more precise but consumes far more context.
How should I chunk the Markdown output?+
Split on heading boundaries rather than fixed character counts, aim for roughly 200 to 800 tokens per chunk, overlap by 10 to 15 percent, and prepend the section heading path to each chunk.
Does converting to Markdown lose information?+
It discards visual formatting — fonts, colours, exact positioning — and keeps semantic structure. For retrieval that is the correct trade, since models reason over meaning rather than appearance.