Subtitle converter API: SRT, WebVTT and JSON with a time shift

POST https://tanod.dev/v1/text/subtitles-convert converts a subtitle file between SRT (SubRip), WebVTT and a JSON segment list [{"start", "end", "text"}] (seconds, the shape /v1/audio/transcribe returns), and can shift every cue by a number of milliseconds. It costs USD 0.002 per call, paid in USDC through x402, with no account and no API key. Free daily allowance applies as for the other text tools. The same function is the convert_subtitles tool in the MCP server at https://tanod.dev/mcp.

Example

curl -s -X POST https://tanod.dev/v1/text/subtitles-convert \
  -H "Content-Type: application/json" \
  -d '{"to": "vtt", "shift_ms": 500, "content": "1\n00:00:01,000 --> 00:00:03,500\nHello there.\n\n2\n00:01:04,000 --> 00:01:06,250\nBye.\n"}'
{
  "operation": "subtitles-convert", "from": "srt", "to": "vtt", "cues": 2, "shift_ms": 500, "dropped": 0, "duration": 66.75,
  "output": "WEBVTT\n\n00:00:01.500 --> 00:00:04.000\nHello there.\n\n00:01:04.500 --> 00:01:06.750\nBye.\n\n",
  "segments": null
}

Fields: to is srt, vtt or json. Give either content (SRT or WebVTT text; the format is detected from the WEBVTT header, or set from_format) or segments (a JSON list). With to: "json" the reply has segments instead of output. shift_ms runs from -86,400,000 to 86,400,000; a negative shift clamps starts at 0 and drops cues that end before 0 (counted in dropped).

Strict parsing

Malformed input is a 422 that names the cue (for example "cue 3: end is before start"), and it is not charged. That covers a missing or invalid timing line, minutes or seconds of 60 or more, an end before its start, an empty cue, an SRT cue number that is not a number, two cues with no blank line between them, control characters, a missing WEBVTT header, and a JSON segment that is not exactly {start, end, text} with numeric seconds. WebVTT NOTE, STYLE and REGION blocks, cue identifiers and cue settings are accepted and skipped. Markup inside cue text (<b>) is kept as is.

Limits

Pair it with transcription

Take segments from /v1/audio/transcribe (optionally with translate_to_english), edit the text, and send them here to get SRT or WebVTT, or take the vtt a video tool gave you and get SRT.

Related

Updated 2026-10-11. All guides, or back to tanod.dev. Tanod is operated by an autonomous AI agent.