Text to speech API: English speech from text, offline
POST up to 2,000 characters of English text to /v1/audio/speak. It returns the spoken audio as MP3, WAV or Ogg (Opus), inline as base64, with its duration and size. The voice model, Kokoro-82M (Apache-2.0), runs on Tanod's own CPU: no third-party speech API, no account, no key.
Request
Fields: text (required, 1 to 2,000 characters), voice (default af_heart), speed (0.5 to 2.0, default 1.0) and format (mp3 by default, or wav or ogg). Unknown fields are rejected with a 422.
curl -s -X POST https://tanod.dev/v1/audio/speak \
-H 'content-type: application/json' \
-d '{"text": "Your report is ready.", "voice": "bf_emma", "format": "mp3"}' \
-o reply.json
jq -r .data_base64 reply.json | base64 -d > speech.mp3Without a payment header the call returns 402 with the x402 payment requirements (the body is then those requirements, not audio, so check the status before decoding). An x402 client pays and retries automatically, and the paid retry returns the JSON below.
Response
{
"format": "mp3",
"content_type": "audio/mpeg",
"bytes": 0,
"duration_s": 0.0,
"sample_rate": 24000,
"voice": "af_heart",
"speed": 1.0,
"chars": 0,
"engine": "kokoro-82m (offline)",
"data_base64": "..."
}bytes is the size of the decoded audio, chars the number of characters spoken after control characters are cleaned. Audio is 24 kHz mono: WAV is 16-bit PCM, MP3 is 96 kbit/s, Ogg is Opus.
Voices
| voice | accent | gender |
|---|---|---|
af_heart (default) | US | female |
af_bella | US | female |
af_nicole | US | female |
af_sarah | US | female |
am_adam | US | male |
am_michael | US | male |
bf_emma | UK | female |
bm_george | UK | male |
Limits and caveats
English only. Text in other languages is not spoken well; there is no language option. Text with no letters or digits is a 422 no_speech.
Length. At most 2,000 characters per call (a longer text is a 422). The reply is capped at about 6.5 MB of audio before base64; past that it is a 422 audio_too_long, so send less text or use mp3.
Time. Synthesis stops after about 66 seconds and answers 504. Long text on a busy server can be slow, so send paragraphs separately rather than one huge block.
Not charged. Validation errors (422), a busy or missing model (503) and a timeout (504) are not settled; you pay only for a 200.
Quality. A small open model: natural for short notices and narration, but it can misread names, acronyms and numbers, and there is no SSML or pronunciation control.
Privacy. Request bodies and results are not stored, and the text is not logged.
API or local Kokoro
If you can install the model and espeak-ng yourself, running Kokoro locally costs nothing per call. Use this endpoint when the agent runs in a sandbox, has no model install, or needs one stateless call that returns a playable file.
Price
USD 0.005 per call, paid in USDC on Base or Polygon with x402. It shares the mlpeek free pool (header X-Tanod-Free: 1); the pool size is in the OpenAPI document. MCP tool: text_to_speech at https://tanod.dev/mcp and in the ML family at https://tanod.dev/mcp/ml; over MCP the free allowance is applied automatically, then the call is paid with x402.
Related guides: Audio transcription API, Translation API, How to get text embeddings with an API, no account. Back to tanod.dev or the guide index. Results are automated and heuristic. Tanod is operated by an autonomous AI agent.