Process Documents
Submit documents for OCR, check processing status, stream results, and download output in multiple formats.
| Endpoint | Description |
|---|---|
| POST v3/pdf | Submit a document for async OCR processing |
| GET v3/pdf/{pdf_id}/stream | Stream page results via SSE |
| GET v3/pdf/{pdf_id} | Check processing status |
| GET v3/converter/{pdf_id} | Check conversion format status |
| GET v3/pdf/{pdf_id}.{ext} | Download results in a specific format |
| DELETE v3/pdf/{pdf_id} | Permanently delete output data |
See the document processing guide for step-by-step examples. For data retention and how to keep outputs long-term, see Data Retention.
POST v3/pdf
POST api.mathpix.com/v3/pdf
Process PDFs, ebooks, and documents asynchronously. Returns a pdf_id for polling status and downloading results. Maximum file size: 1 GB.
v3/pdf and Files API requests use document-layout recognition automatically. The enable_document_layout option is only needed when a full-page image is sent directly to v3/text.
Supported inputs (full list):
- Documents: PDF, DOCX, PPTX, DOC, WPD, ODT
- Ebooks: EPUB, AZW/AZW3/KFX, MOBI, DJVU
Supported outputs (full list):
- Text: MMD, MD, HTML, LaTeX (.tex.zip)
- Office: DOCX, XLSX, PPTX
- PDF: HTML-rendered or LaTeX-rendered
- Archives: ZIP variants with embedded images
Example
- cURL
- Python
- JavaScript / TypeScript
- Go
- Java
curl -X POST https://api.mathpix.com/v3/pdf \
-H 'app_id: APP_ID' \
-H 'app_key: APP_KEY' \
-H 'Content-Type: application/json' \
--data '{"url": "https://cdn.mathpix.com/examples/cs229-notes1.pdf", "conversion_formats": {"docx": true, "tex.zip": true}}'
import requests
r = requests.post("https://api.mathpix.com/v3/pdf",
json={
"url": "https://cdn.mathpix.com/examples/cs229-notes1.pdf",
"conversion_formats": {"docx": True, "tex.zip": True}
},
headers={
"app_id": "APP_ID",
"app_key": "APP_KEY",
"Content-type": "application/json"
}
)
print(r.json()) # {"pdf_id": "..."}
const response = await fetch("https://api.mathpix.com/v3/pdf", {
method: "POST",
headers: {
app_id: "APP_ID",
app_key: "APP_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
url: "https://cdn.mathpix.com/examples/cs229-notes1.pdf",
conversion_formats: { docx: true, "tex.zip": true },
}),
});
const { pdf_id } = await response.json();
console.log(`PDF ID: ${pdf_id}`);
body := bytes.NewBufferString(`{
"url": "https://cdn.mathpix.com/examples/cs229-notes1.pdf",
"conversion_formats": {"docx": true, "tex.zip": true}
}`)
req, _ := http.NewRequest("POST", "https://api.mathpix.com/v3/pdf", body)
req.Header.Set("app_id", "APP_ID")
req.Header.Set("app_key", "APP_KEY")
req.Header.Set("Content-Type", "application/json")
resp, _ := http.DefaultClient.Do(req)
defer resp.Body.Close()
result, _ := io.ReadAll(resp.Body)
fmt.Println(string(result)) // {"pdf_id": "..."}
HttpClient client = HttpClient.newHttpClient();
String body = """
{
"url": "https://cdn.mathpix.com/examples/cs229-notes1.pdf",
"conversion_formats": { "docx": true, "tex.zip": true }
}
""";
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.mathpix.com/v3/pdf"))
.header("app_id", "APP_ID")
.header("app_key", "APP_KEY")
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body))
.build();
HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
{
"pdf_id": "2024_01_15_abc123def456"
}
Request parameters
You can either send a file URL in a JSON body, or upload a file via multipart form-data (with parameters in options_json).
url HTTP URL where the file can be downloaded from
streaming Whether streaming should be enabled for this request, see stream pdf pages.
metadata Key-value object. Supports improve_mathpix for extra privacy controls.
alphabets_allowed Specify which alphabets you don't want in the output
rm_spaces Determines whether extra white space is removed from equations in latex_styled and text formats.
See output option examples for the same equation returned with and without the spacing.
rm_fonts Determines whether font commands such as \mathbf and \mathrm are removed from equations in latex_styled and text formats.
See output option examples for the same equation returned with and without the font commands.
idiomatic_eqn_arrays Specifies whether to use aligned, gathered, or cases instead of an array environment for a list of equations.
See output option examples for a piecewise function returned as an array and as cases.
include_equation_tags Specifies whether to include equation number tags inside equations LaTeX. When set to true, it sets "idiomatic_eqn_arrays": true, because equation numbering works better in those environments compared to the array environment.
Example
\tag{eq_number}, where eq_number is an equation number (e.g. 1.12)
See output option examples for a numbered equation returned with and without its tag.
include_smiles Enable experimental chemistry diagram OCR, via RDKIT normalized SMILES with isomericSmiles=False, included in text output format, via MMD SMILES syntax <smiles>...</smiles>.
See output option examples for a chemical diagram returned with and without its SMILES string.
include_chemistry_as_image Returns an image crop containing the SMILES in the alt-text for chemical diagrams.
Example

See output option examples for the same document returned with the diagram as SMILES and as an image crop.
include_page_breaks If true, \pagebreak is appended to each page's MMD in the concatenated output (one marker per page - an N-page PDF produces N markers).
Example
page-1 content\n\pagebreak\npage-2 content\n\pagebreak\n
See output option examples for a two page document returned with and without page break markers.
include_hyperlinks If true, hyperlinks the PDF carries are preserved: the anchor text becomes a markdown link in the text output, and every line carrying a link also reports it in PDF lines data. http, https and mailto links only.
Only links the PDF itself stores are preserved. A scanned page has no link layer, and a link with no text under it (a clickable figure or logo) has no anchor text to place - in both cases nothing is emitted. A printed URL that is not also a link stays inert text.
The link label is the text as recognized on the page, while the target is the URL stored in the PDF, so the two can differ where recognition and the stored link disagree.
Example
A page printing Visit our docs to learn more, where our docs is a link to https://mathpix.com/docs:
Visit [our docs](https://mathpix.com/docs) to learn more
See output option examples for the same document returned with and without its links.
include_page_info Controls whether page info elements are included in the final MMD output. Page info refers to elements like headers, footers, page numbers, and detected QR codes that are not part of the main text (unlike v3/text where it defaults to true). Set this to true to return detected QR codes as cropped images.
Also accepts an array of page info type names to include only a subset: "margin_note" for margin notes (notes and comments beside the main text), "qr_code" for detected QR codes (returned as cropped images). For example, "include_page_info": ["margin_note"] keeps margin notes in the output while page numbers and headers stay excluded. Type names select what the model classifies each element as, not its physical position on the page - so "margin_note" returns everything recognized as margin-note content, wherever it appears.
The selectable names are exactly the page_info subtypes reported in line data. General page info without a subtype (headers, footers, page numbers, stamps, and similar) is only included via the boolean true.
See output option examples for a page processed with this option on and off. That example runs on v3/text, where the default is true; here the default is false.
disable_lstlisting Controls how recognized code and pseudocode are represented. By default, document output uses \begin{lstlisting} so inline math can be rendered inside pseudocode. Set to true to emit standard Markdown triple-backtick fences instead; math inside the fallback fence remains plain text.
If you render MMD yourself, lstlisting requires mathpix-markdown-it 2.0.29 or newer, and version 3.0.0 is recommended. Older versions can omit the entire listing body when rendering. Mathpix-hosted conversion formats are rendered server-side and do not require a client upgrade. Files API requests inherit this option from v3/pdf.
See output option examples for the same input processed with this option off and on.
disable_itemize Controls how recognized lists are represented. By default, lists are emitted as \begin{itemize} / \item environments. Set to true to emit each entry as a plain line with its marker kept inline (for example 1., -, □); lists inside tables are emitted as \\-separated rows instead of a nested itemize. Files API requests inherit this option from v3/pdf.
See output option examples for the same input processed with this option off and on.
numbers_default_to_math Specifies whether numbers are always math.
Example
Answer: \( 17 \) instead of Answer: 17
See output option examples for the same input processed with this option off and on.
math_inline_delimiters Specifies begin inline math and end inline math delimiters for text outputs.
See output option examples for the same text returned with the default delimiters and with dollar signs.
math_display_delimiters Specifies begin display math and end display math delimiters for text outputs.
See output option examples for the same equation returned with the default delimiters and with double dollar signs.
page_ranges Specifies a page range as a comma-separated string.
Examples
2,4-6selects pages [2,4,5,6]2 - -2selects all pages starting with the second page and ending with the next-to-last page (specified by -2)
enable_spell_check Deprecated, has no effect on the output.
auto_number_sections Specifies whether sections and subsections in the output are automatically numbered note
See output option examples for an unnumbered document returned with and without generated numbering.
remove_section_numbering Specifies whether to remove existing numbering for sections and subsections note
See output option examples for a numbered document returned with its numbering kept and removed.
preserve_section_numbering Specifies whether to keep existing section numbering as is note
See output option examples for a numbered document returned with its numbering kept, which is this default, and removed.
enable_tables_fallback Deprecated, accepted for backward compatibility but has no effect. Tables are always recognized with the table segmentation algorithm.
fullwidth_punctuation Controls if punctuation will be fullwidth Unicode (default for east Asian languages like Chinese), or halfwidth Unicode (default for Latin scripts, Cyrillic scripts etc.). When null, fullwidth vs halfwidth will be decided based on image content. Punctuation inside math will always stay halfwidth.
See output option examples for the same Chinese text returned with fullwidth and with halfwidth punctuation.
conversion_formats Specifies formats that the v3/pdf output (Mathpix Markdown) should automatically be converted into on completion.
conversion_options Options for specific output formats (e.g. font, margins, orientation for DOCX).
callback_url HTTPS URL Mathpix sends a webhook notification to when this document
finishes, instead of you polling for it. A request without one is not notified. This is a separate
mechanism from the callback object on v3/text,
v3/latex, and v3/batch.
callback_headers Sent on this document's notification, so you can authenticate the request on your side. Belongs with
callback_url: set both together, because a request that sets callback_url and no callback_headers
is notified with no headers. See
callback parameters at submission for the limits.
callback_events The events to receive for this document. A v3/pdf request is
one document, so it is file.completed or file.error; an empty array turns notifications off for
just this request.
A callback parameter this endpoint refuses fails the whole submission; the refusals and their messages are listed under callback parameter errors.
Section numbering
Only one of auto_number_sections, remove_section_numbering, or preserve_section_numbering can be true at a time. The default behavior is to preserve section numbering (preserve_section_numbering set to true). Setting multiple flags returns error opts_section_numbering.
Response body
pdf_id Tracking ID to get status and result when completed
error US locale error message
error_info Error info object
GET v3/pdf/{pdf_id}/stream
Stream page results via server-sent events (SSE) for lower time to first data. Requires streaming: true in the initial POST request.
GET api.mathpix.com/v3/pdf/{pdf_id}/stream
Example
- cURL
- Python
- JavaScript / TypeScript
- Go
- Java
curl -N https://api.mathpix.com/v3/pdf/PDF_ID/stream \
-H 'app_id: APP_ID' \
-H 'app_key: APP_KEY' \
-H 'Accept: text/event-stream'
import requests
r = requests.get("https://api.mathpix.com/v3/pdf/PDF_ID/stream",
headers={
"app_id": "APP_ID",
"app_key": "APP_KEY",
"Accept": "text/event-stream"
},
stream=True
)
for line in r.iter_lines(decode_unicode=True):
if line.startswith("data: "):
print(line[6:])
const response = await fetch("https://api.mathpix.com/v3/pdf/PDF_ID/stream", {
headers: {
app_id: "APP_ID",
app_key: "APP_KEY",
Accept: "text/event-stream",
},
});
const reader = response.body!.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const chunk = decoder.decode(value);
for (const line of chunk.split("\n")) {
if (line.startsWith("data: ")) console.log(line.slice(6));
}
}
req, _ := http.NewRequest("GET", "https://api.mathpix.com/v3/pdf/PDF_ID/stream", nil)
req.Header.Set("app_id", "APP_ID")
req.Header.Set("app_key", "APP_KEY")
req.Header.Set("Accept", "text/event-stream")
resp, _ := http.DefaultClient.Do(req)
defer resp.Body.Close()
scanner := bufio.NewScanner(resp.Body)
for scanner.Scan() {
line := scanner.Text()
if strings.HasPrefix(line, "data: ") {
fmt.Println(line[6:])
}
}
HttpClient client = HttpClient.newHttpClient();
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.mathpix.com/v3/pdf/PDF_ID/stream"))
.header("app_id", "APP_ID")
.header("app_key", "APP_KEY")
.header("Accept", "text/event-stream")
.GET()
.build();
HttpResponse<Stream<String>> response = client.send(request,
HttpResponse.BodyHandlers.ofLines());
response.body()
.filter(line -> line.startsWith("data: "))
.forEach(line -> System.out.println(line.substring(6)));
text Mathpix Markdown output
page_idx page index from selected page range, starting at 1 and going all the way to pdf_selected_len
pdf_selected_len total number of pages inside selected page range
confidence overall confidence score for the page (0-1)
confidence_rate rate of high-confidence characters on the page (0-1)
version model version used for processing (e.g. SuperNet-200)
Pages are streamed one JSON object at a time. Pages are not guaranteed to be in order, although they generally will be.
GET v3/pdf/{pdf_id}
Check the processing status of a PDF.
GET api.mathpix.com/v3/pdf/{pdf_id}
status, not percent_donestatus is the only field that says results are ready. percent_done and num_pages_completed report OCR progress alone and reach 100 before the output files are assembled, so a client that fetches on a timer instead of on status will sometimes fetch too early and get a 404. Poll this endpoint until status is completed (or error), then download from GET v3/pdf/{pdf_id}.{ext}.
Example
- cURL
- Python
- JavaScript / TypeScript
- Go
- Java
curl https://api.mathpix.com/v3/pdf/PDF_ID \
-H 'app_id: APP_ID' \
-H 'app_key: APP_KEY'
import requests
r = requests.get("https://api.mathpix.com/v3/pdf/PDF_ID",
headers={"app_id": "APP_ID", "app_key": "APP_KEY"}
)
print(r.json()) # {"status": "completed", "num_pages": 10, ...}
const response = await fetch("https://api.mathpix.com/v3/pdf/PDF_ID", {
headers: { app_id: "APP_ID", app_key: "APP_KEY" },
});
const status = await response.json();
console.log(status);
req, _ := http.NewRequest("GET", "https://api.mathpix.com/v3/pdf/PDF_ID", nil)
req.Header.Set("app_id", "APP_ID")
req.Header.Set("app_key", "APP_KEY")
resp, _ := http.DefaultClient.Do(req)
defer resp.Body.Close()
result, _ := io.ReadAll(resp.Body)
fmt.Println(string(result))
HttpClient client = HttpClient.newHttpClient();
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.mathpix.com/v3/pdf/PDF_ID"))
.header("app_id", "APP_ID")
.header("app_key", "APP_KEY")
.GET()
.build();
HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
pdf_id The PDF tracking ID
status | Value | Meaning |
|---|---|
received | Request accepted |
loaded | PDF downloaded onto our servers |
split | Pages split and sent for processing |
completed | Processing finished successfully |
error | A problem occurred during processing |
num_pages Total number of pages in PDF document
num_pages_completed Current number of pages in PDF document that have been OCR-ed. Counts OCR progress only, so it can equal num_pages before results are downloadable.
percent_done Percentage of pages in PDF that have been OCR-ed, computed as 100 * num_pages_completed / num_pages. It reaches 100 when the last page finishes OCR, which is before the output files are assembled; use status to decide when to download.
app_id The app ID that submitted the request
group_id The group ID associated with the request
input_file Original filename of the uploaded document
version Release identifier for the PDF processing pipeline that produced this document. It changes whenever output may have changed, including conversion and formatting releases
conversion_status Status of each requested conversion format.
GET v3/converter/{pdf_id}
Check the status of requested conversion formats.
GET api.mathpix.com/v3/converter/{pdf_id}
Example
- cURL
- Python
- JavaScript / TypeScript
- Go
- Java
curl https://api.mathpix.com/v3/converter/PDF_ID \
-H 'app_id: APP_ID' \
-H 'app_key: APP_KEY'
import requests
r = requests.get("https://api.mathpix.com/v3/converter/PDF_ID",
headers={"app_id": "APP_ID", "app_key": "APP_KEY"}
)
print(r.json()) # {"status": "completed", "conversion_status": {...}}
const response = await fetch("https://api.mathpix.com/v3/converter/PDF_ID", {
headers: { app_id: "APP_ID", app_key: "APP_KEY" },
});
const status = await response.json();
console.log(status);
req, _ := http.NewRequest("GET", "https://api.mathpix.com/v3/converter/PDF_ID", nil)
req.Header.Set("app_id", "APP_ID")
req.Header.Set("app_key", "APP_KEY")
resp, _ := http.DefaultClient.Do(req)
defer resp.Body.Close()
result, _ := io.ReadAll(resp.Body)
fmt.Println(string(result))
HttpClient client = HttpClient.newHttpClient();
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.mathpix.com/v3/converter/PDF_ID"))
.header("app_id", "APP_ID")
.header("app_key", "APP_KEY")
.GET()
.build();
HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
status Always completed once the PDF has finished processing. Individual format progress is tracked in conversion_status.
conversion_status Status of each requested conversion format.
GET v3/pdf/{pdf_id}.{ext}
Download results by appending the format extension to the pdf_id. Results are available once status=completed. Conversion formats (e.g., docx) require the format's conversion status to be completed.
Requesting a result before it exists returns 404 with the status object as the response body, so a 404 carrying a non-completed status means "not ready yet" rather than "no such document". Poll the status endpoint until status is completed instead of waiting a fixed amount of time after submitting: assembly time scales with document size and with how much work is in flight.
Requesting a conversion format whose conversion is still running returns 202 with the same status object as the body and a Retry-After header naming how many seconds to wait before requesting again.
GET api.mathpix.com/v3/pdf/{pdf_id}.{ext}
Example
- cURL
- Python
- JavaScript / TypeScript
- Go
- Java
curl https://api.mathpix.com/v3/pdf/PDF_ID.docx \
-H 'app_id: APP_ID' \
-H 'app_key: APP_KEY' \
-o output.docx
import requests
r = requests.get("https://api.mathpix.com/v3/pdf/PDF_ID.docx",
headers={"app_id": "APP_ID", "app_key": "APP_KEY"}
)
with open("output.docx", "wb") as f:
f.write(r.content)
import { writeFile } from "fs/promises";
const response = await fetch("https://api.mathpix.com/v3/pdf/PDF_ID.docx", {
headers: { app_id: "APP_ID", app_key: "APP_KEY" },
});
const buffer = await response.arrayBuffer();
await writeFile("output.docx", Buffer.from(buffer));
req, _ := http.NewRequest("GET", "https://api.mathpix.com/v3/pdf/PDF_ID.docx", nil)
req.Header.Set("app_id", "APP_ID")
req.Header.Set("app_key", "APP_KEY")
resp, _ := http.DefaultClient.Do(req)
defer resp.Body.Close()
out, _ := os.Create("output.docx")
defer out.Close()
io.Copy(out, resp.Body)
HttpClient client = HttpClient.newHttpClient();
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.mathpix.com/v3/pdf/PDF_ID.docx"))
.header("app_id", "APP_ID")
.header("app_key", "APP_KEY")
.GET()
.build();
HttpResponse<Path> response = client.send(request,
HttpResponse.BodyHandlers.ofFile(Path.of("output.docx")));
System.out.println("Saved to " + response.body());
Accepted extensions: .mmd, .md, .docx, .tex.zip, .html, .pdf, .latex.pdf, .mmd.overlay.pdf, .pptx, .xlsx, .mmd.zip, .md.zip, .html.zip, .lines.json, .lines.mmd.json
See Conversion Formats for availability and descriptions. The lines.json and lines.mmd.json extensions are v3/pdf-only - see PDF lines data and PDF MMD lines data.
PDF lines data
Detailed line-by-line data for PDFs, useful for building custom experiences on top of original PDFs.
Response data object
pages List of PdfPageData objects
PdfPageData object
image-id PDF ID, plus hyphen, plus page number, starting at page 1
page Page number
lines List of LineData objects
page_height Page height (in pixel coordinates)
page_width Page width (in pixel coordinates)
PdfLineData object
id Unique line identifier
parent_id Unique line identifier of the parent.
children_ids List of children unique identifiers.
type See line types and subtypes for details.
subtype See line types and subtypes for details.
line Line number
text Searchable text, empty string for page elements that do not necessarily have associated text (for example individual equations inside block of math equations).
text_display Mathpix Markdown content with contextual elements such as article, section and inline image URLs. Can be empty for page elements that will not render (for example page number, auxiliary text in the page header, etc.).
conversion_output When true, text_display from the line is included in the final MMD output, otherwise excluded.
is_printed True if line contains printed text, false otherwise.
is_handwritten True if line contains handwritten text, false otherwise.
links Hyperlinks the PDF carries on this line, present only when include_hyperlinks is true and the line carries at least one. A link is reported here even when it could not be placed in the text.
region Bounding box of the line in pixel coordinates
cnt Specifies the image area as list of (x,y) pixel coordinate pairs. For axis-aligned bounding boxes, vertices are in [TL, TR, BR, BL] order (clockwise from top-left). This captures handwritten content much better than a region object
confidence Estimated probability 100% correct (product of per token OCR confidence).
confidence_rate Estimated confidence of output quality (geometric mean of per token OCR confidence).
PDF MMD lines data (deprecated)
Deprecated. Use lines.json instead, which contains all this information and more.
Response data object (MMD Lines)
pages List of PdfMMDPageData objects
PdfMMDPageData object
image-id PDF ID, plus hyphen, plus page number, starting at page 1
page Page number
lines List of PageMMDLineData objects
page_height Page height (in pixel coordinates)
page_width Page width (in pixel coordinates)
PdfMMDLineData object
line Line number
text Mathpix Markdown content with contextual elements such as article, section and inline image URLs
is_printed True if line contains printed text, false otherwise.
is_handwritten True if line contains handwritten text, false otherwise.
links Hyperlinks the PDF carries on this line, present only when include_hyperlinks is true and the line carries at least one. A link is reported here even when it could not be placed in the text.
region Bounding box of the line in pixel coordinates
cnt Specifies the image area as list of (x,y) pixel coordinate pairs. For axis-aligned bounding boxes, vertices are in [TL, TR, BR, BL] order (clockwise from top-left). This captures handwritten content much better than a region object
confidence Estimated probability 100% correct (product of per token OCR confidence).
confidence_rate Estimated confidence of output quality (geometric mean of per token OCR confidence).
Link object
A hyperlink stored in the source PDF, reported when include_hyperlinks is true.
url The link target as stored in the PDF. External links keep their scheme (http, https or mailto); an internal jump to another page of the same document is reported as #page=N.
anchor_text The text the PDF stores under the link, taken from its own text layer. This can differ from the line's recognized text where OCR and the stored text disagree.
kind uri for a link to an external address, goto for a jump to another page of the same document. Only uri links are placed in the text; goto links are reported here only.
DELETE v3/pdf/{pdf_id}
Permanently delete a PDF's output data.
DELETE api.mathpix.com/v3/pdf/{pdf_id}
Example
- cURL
- Python
- JavaScript / TypeScript
- Go
- Java
curl -X DELETE https://api.mathpix.com/v3/pdf/PDF_ID \
-H 'app_id: APP_ID' \
-H 'app_key: APP_KEY'
import requests
r = requests.delete("https://api.mathpix.com/v3/pdf/PDF_ID",
headers={"app_id": "APP_ID", "app_key": "APP_KEY"}
)
print(r.json())
const response = await fetch("https://api.mathpix.com/v3/pdf/PDF_ID", {
method: "DELETE",
headers: { app_id: "APP_ID", app_key: "APP_KEY" },
});
console.log(response.status);
req, _ := http.NewRequest("DELETE", "https://api.mathpix.com/v3/pdf/PDF_ID", nil)
req.Header.Set("app_id", "APP_ID")
req.Header.Set("app_key", "APP_KEY")
resp, _ := http.DefaultClient.Do(req)
defer resp.Body.Close()
result, _ := io.ReadAll(resp.Body)
fmt.Println(string(result))
HttpClient client = HttpClient.newHttpClient();
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("https://api.mathpix.com/v3/pdf/PDF_ID"))
.header("app_id", "APP_ID")
.header("app_key", "APP_KEY")
.DELETE()
.build();
HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
System.out.println(response.body());
When a PDF is deleted:
- All output files are permanently removed from our servers (MMD, images, JSON Lines, and all requested formats)
- The original input file is also deleted
- This deletion is permanent and cannot be undone
Download and store files locally before deleting if you need to keep them.
PDF page images and cropped images (figures, diagrams) served via CDN may remain accessible for up to 10 minutes after deletion while cached copies expire.
Minimal metadata is retained for auditing and billing: status, input_file name, num_pages, timestamps, and processing version. No output content is stored.
If privacy is a concern, rename the file before upload to avoid storing identifiable filenames.
Retention
Uploaded source documents and the page image files that back cdn.mathpix.com/cropped/... URLs are retained for up to 30 days. Mathpix Markdown, JSON Lines, and other text outputs are retained for up to 90 days. If you need image URLs to remain accessible beyond 30 days, request a zip output format at processing time (the zip embeds all images inline). See Data Retention for the full retention matrix and long-term storage patterns.
Response body
Returns the PDF status object at the time of deletion.
pdf_id The PDF tracking ID
app_id The app ID that submitted the request
group_id The group ID associated with the request
status Processing status at time of deletion (e.g. completed)
input_file Original filename of the uploaded document
num_pages Total number of pages in PDF document
num_pages_completed Number of pages that were OCR-ed
percent_done Percentage of pages that were OCR-ed
conversion_status Status of each requested conversion format
After deletion, subsequent GET requests to v3/pdf/{pdf_id} return the same status object with an additional deleted_at field:
deleted_at ISO 8601 timestamp of when the PDF was deleted. Only present on GET requests after deletion.
Table layout quality and image preservation
SuperNet-201 preserves rejected table structure as a source image through the existing document image-output pipeline. No new request option is required. Good neighboring text remains usable, while the rejected table's cells are withheld.
Each affected lines.json entry has parse_status and fallback_reason. A successful replacement is image_fallback; a source/storage failure is failed with image_fallback_unavailable. text_display contains the source-image reference, and text does not contain rejected structure. The existing type, id and geometry still describe the source region.
{"page":1,"parse_status":"partial_fallback","fallback_count":1,"lines":[{"id":"table-17","type":"table","text":"","text_display":"","parse_status":"image_fallback","fallback_reason":"invalid_structure"}]}
The image path is illustrative and follows the configured output destination. Page parse_status and fallback_count also appear in worker page results. parsed means structured output was delivered, not that all content is guaranteed correct. Confidence fields are described below.
Unified page and line confidence
Image lines_json and each PDF lines.json page use the same confidence contract. Every region has confidence, text_confidence, and layout_confidence; unknown or inapplicable values are null. The page contains a page_confidence object with the same three fields and method: "prototype_min_v1". Worker page summaries also expose this object.
{
"page_confidence": {
"confidence": 0.8,
"text_confidence": 0.95,
"layout_confidence": 0.8,
"method": "prototype_min_v1"
},
"lines": [
{
"id": "line-1",
"type": "text",
"text": "Example",
"confidence": 0.8,
"text_confidence": 0.95,
"layout_confidence": 0.8
}
]
}
Numbers are illustrative. This first implementation supplies prototype reliability estimates, not validated correctness probabilities. Line text_confidence is the existing OCR confidence value. Layout uses model-matched calibrated confidence when available, otherwise its raw evidence as an explicitly provisional proxy. Combined confidence is the minimum of required layout and text components; it does not assume independent errors. Text-bearing leaves missing OCR evidence have unknown combined confidence. Non-text regions have text_confidence: null and use layout evidence alone. Structural parents do not duplicate their children's OCR evidence.
Page text confidence is the minimum across text-bearing leaves, or unknown if any such leaf lacks evidence. Page layout uses the root and nested segmentation pass assessment, rather than just returned lines. A failed or image-fallback structured extraction has combined/layout confidence zero; preserving pixels does not mean successful structured recognition. An empty page with no evidence has unknown confidence. These heuristics do not guarantee detection of omitted, repeated, or misordered content and are sensitive to page length. No reviewed joint calibration artifact ships with this change. The prototype does not change rejection/fallback thresholds.
Migration: lines_json.lines[*].confidence changes from OCR-only to combined confidence. Read text_confidence for its old meaning. Legacy image line_data, word confidence, top-level image OCR confidence, and confidence_rate retain their existing semantics. layout_score is removed from both the page and every region in lines_json; raw evidence remains in saved-page/internal layout diagnostics (and the legacy line_data contract). There is no duplicate top-level layout_confidence; page components live under page_confidence.