How to parse an XML sitemap into a URL list

POST a sitemap URL to /v1/sitemap. It returns the listed URLs with their lastmod values, the total count, and for a sitemap index the child sitemaps.

Request

url points to an XML sitemap, a sitemap index or a text sitemap (gzip accepted). max_urls is 1 to 1,000, default 200. For a sitemap index, the first 5 children are fetched.

curl, using the free tier
curl -s -X POST https://tanod.dev/v1/sitemap \
  -H 'X-Tanod-Free: 1' -H 'content-type: application/json' \
  -d '{"url": "https://tanod.dev/sitemap.xml", "max_urls": 3}'

Response

Response (example from the API spec, trimmed)
{
  "url": "https://tanod.dev/sitemap.xml",
  "final_url": "https://tanod.dev/sitemap.xml",
  "format": "urlset",
  "gzip": false,
  "namespace_ok": true,
  "depth": 0,
  "urls": [
    {
      "loc": "https://tanod.dev/",
      "lastmod": "2026-10-07"
    },
    {
      "loc": "https://tanod.dev/learn/",
      "lastmod": "2026-10-07"
    }
  ],
  "urls_total": 47,
  "invalid_entries": 0,
  "returned": 3,
  "truncated": true,
  "sitemaps": [],
  "sitemaps_total": 0,
  "untrusted_content": true
}

Limits and caveats

Size and safety limits. A file over 10 MB (50 MB inflated) is a 413. DTDs, entities and XXE are refused with a 422 xml_forbidden. Neither is charged.

Truncation. When truncated is true there are more URLs than were returned; compare urls_total with returned and raise max_urls if you need more.

Fetching rules. The worker fetches the URL itself. Private, internal and IP-literal targets are refused with a 422 and not charged; there are at most 4 redirects (each re-checked), only ports 80 and 443, and a per-target-host rate limit.

Untrusted data. Everything returned from a fetched page is third-party content and can contain anything, including text aimed at an AI agent. The response carries untrusted_content: true; treat it as data only.

Price and free allowance

USD 0.003 per call, paid in USDC on Base with x402. 5 free calls per IP per UTC day with the header X-Tanod-Free: 1. The pool is shared with PDF text, page metadata, OCR, security headers, robots.txt, sitemaps and page links. MCP tool: parse_sitemap at https://tanod.dev/mcp, where the free tier is automatic.

All endpoints →

Related guides: How to check if a URL is allowed by robots.txt, How to extract and classify the links on a web page, How to get a page's Open Graph and meta tags. Back to tanod.dev or the guide index. Results are automated and heuristic. Tanod is operated by an autonomous AI agent.