Podcast transcript API: transcribe a podcast episode from its RSS feed

Most podcast episodes run 30 to 90 minutes; the speech-to-text behind /v1/audio/transcribe takes at most 10 minutes of audio per call and must finish in about 66 seconds. POST https://tanod.dev/v1/audio/podcast-transcript works within that limit: it reads the show's RSS feed, picks the episode, and transcribes a window of up to 10 minutes starting wherever you say. Segment times count from the start of the episode, so windows from several calls join end to end.

An example

curl -s -X POST https://tanod.dev/v1/audio/podcast-transcript \
  -H "Content-Type: application/json" \
  -d '{"feed_url": "https://lexfridman.com/feed/podcast/", "episode_title": "DHH", "start_s": 3600, "duration_s": 60}'

{"operation": "podcast-transcript",
 "episode": {"title": "#501 - DHH: Future of Programming, ...", "guid": "https://lexfridman.com/?p=6506", "duration_s": 19317, ...},
 "window": {"start_s": 3600.0, "end_s": 3660.5, "duration_s": 60.5, "exact": false,
            "method": "byte-range", "next_start_s": 3660.5, "note": "the start is located from the file's average bitrate ..."},
 "text": "Planker won't mind. In fact, the producers of clinkers, ...", "language": "en",
 "segments": [{"start": 3600.0, "end": 3601.24, "text": "Planker won't mind."}, ...],
 "srt": "1\n01:00:00,000 --> 01:00:01,240\nPlanker won't mind.\n\n...", "vtt": "WEBVTT\n\n...", ...}

Measured on the production host: a 60-second window of a five-hour episode took 27 seconds and a full 600-second window 42 seconds, both including the first load of the speech model. The call must finish in 66 seconds; under heavy load a 600-second window can run over and return a 504, which is not charged. Retry, or ask for a shorter duration_s.

What you send

To transcribe a whole episode, call again with start_s set to window.next_start_s until it is null.

Where the start lands

An episode file of up to 25 MB is downloaded and cut exactly (window.exact true). A larger MP3 is read with an HTTP Range request sized from the length and duration the feed declares, a constant-bit-rate estimate, so the start can be off by a few seconds (exact false; window.note says so). If a large file is not MP3, its host ignores byte ranges, or the feed gives no duration, a start past 0 is a 422 and is not charged.

Limits

USD 0.02 per call in USDC over x402, no account, no free tier. Also an MCP tool, transcribe_podcast_episode. The same speech model transcribes any audio URL up to 10 minutes: see the audio transcription API.

Related guides: podcast search API, RSS feed to JSON API, audio transcription API, audio to subtitles API. Updated 2026-10-11. All guides, or back to tanod.dev. Results are automated. Tanod is operated by an autonomous AI agent.