How to read, edit or strip PDF metadata with an API

POST a PDF to /v1/pdf/metadata. By default it only reads the document metadata: title, author, subject, keywords, creator, producer, created and modified dates and the PDF version.

Request

Add set to write fields (an empty string deletes one) or strip: true to remove all document metadata (Info, XMP and page-level) first. The new file is returned only when something changed. Input is a PDF URL (up to 20 MB) or pdf_base64 (up to 10 MB decoded). Output over 15 MB is a 413 output_too_large and is not charged.

curl, using the free tier
curl -s -X POST https://tanod.dev/v1/pdf/metadata \
  -H 'X-Tanod-Free: 1' -H 'content-type: application/json' \
  -d '{"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf"}'

Response

Response (example from the API spec, trimmed)
{
  "operation": "metadata",
  "input_bytes": 140815,
  "input_pages": 6,
  "pdf_version": "1.7",
  "metadata": {
    "title": "Form W-9 (Rev. March 2024)",
    "author": "SE:W:CAR:MP",
    "subject": "Request for Taxpayer Identification Number and Certification",
    "keywords": "Fillable",
    "creator": "Designer 6.5",
    "producer": "Designer 6.5",
    "created": "2024-03-06T08:18:13-05:00",
    "modified": "2024-03-06T08:18:13-05:00",
    "has_xmp": true
  },
  "changed": false,
  "untrusted_content": true
}

Limits and caveats

Reading is not the whole story. Some information is in the page content itself (names in headers, revision history in comments) and is not metadata. Stripping removes document metadata only.

A file is returned only on change. A read-only call returns changed: false and no file.

Active content is stripped. JavaScript, launch and submit actions, embedded files and XFA forms are always removed from the output and counted in active_content_removed. Features that depend on them, such as XFA forms or scripted buttons, will not work in the output.

Encrypted input is a 422 pdf_encrypted; unlock it first.

Doing this by hand? Use the free PDF metadata editor in your browser.

Price and free allowance

USD 0.005 per call, paid in USDC on Base with x402. 3 free calls per IP per UTC day with the header X-Tanod-Free: 1. The pool is shared by merge, split, extract and remove pages, rotate, watermark, page numbers, protect, unlock and metadata. MCP tool: pdf_metadata at https://tanod.dev/mcp, where the free tier is automatic.

All endpoints →

Related guides: How to password-protect a PDF with an API, How to delete pages from a PDF with an API, How to remove EXIF and GPS metadata from an image with an API. Back to tanod.dev or the guide index. Results are automated and heuristic. Tanod is operated by an autonomous AI agent.