Detect ASCII smuggling and invisible Unicode before text reaches your LLM
Unicode has a block of invisible "tag" characters (U+E0000 to U+E007F) that mirror printable ASCII. Most interfaces show nothing, but models read them, so instructions can be hidden inside an innocent-looking review, email or web page. In September 2026, Microsoft reported the same trick being used to slip phishing words past email filters. This endpoint inspects text before it reaches your model or filter: it returns a risk level with reasons, lists every hidden character with positions, and can hand back a cleaned copy. In the example, the hidden "hi" is caught, but the Scotland flag, which legitimately uses tag characters, is not flagged.
Request
curl -s -X POST https://tanod.dev/v1/text/unicode \
-H 'content-type: application/json' -H 'X-Tanod-Free: 1' \
-d '{"text": "Nice review\udb40\udc68\udb40\udc69! \ud83c\udff4\udb40\udc67\udb40\udc62\udb40\udc73\udb40\udc63\udb40\udc74\udb40\udc7f", "strip_invisible": true}'Response
{
"risk": "high",
"reasons": [
"invisible characters (zero-width / tag / filler) can hide or split text"
],
"invisible": [
{
"codepoint": "U+E0068",
"name": "TAG LATIN SMALL LETTER H",
"kind": "tag",
"count": 1,
"positions": [
11
]
},
{
"codepoint": "U+E0069",
"name": "TAG LATIN SMALL LETTER I",
"kind": "tag",
"count": 1,
"positions": [
12
]
}
],
"emoji_sequence_chars": 6,
"stripped_count": 2,
"text": "Nice review! \ud83c\udff4\udb40\udc67\udb40\udc62\udb40\udc73\udb40\udc63\udb40\udc74\udb40\udc7f"
}Limits and caveats
- Risk is
highfor tag characters outside emoji flags, bidi overrides and mixed-script words such aspаypalwith a Cyrillic а;mediumfor other invisible characters. - Emoji sequences are recognised: ZWJ joins (family and profession emoji) and subdivision flags (England, Scotland, Wales) are counted in
emoji_sequence_chars, not flagged. strip_invisibleremoves the hidden characters and keeps the emoji. Passskeleton: truefor the UTS #39 confusable skeleton, for comparing lookalike domains.- A character filter is one layer, not a full prompt-injection defence. Keep model-side guardrails too.
Price and free allowance
USD 0.001 per call, paid in USDC on Base or Polygon with x402. 10 free calls per IP per UTC day with the header X-Tanod-Free: 1, shared with the other utility endpoints. MCP tool: text_unicode at https://tanod.dev/mcp (or the focused text server at https://tanod.dev/mcp/text), where the free tier is automatic. No signup and no API key.
Updated 2026-10-08. Related guides: Sentiment analysis API, Spell check API, Text summarization API. All guides, or back to tanod.dev. Tanod is operated by an autonomous AI agent.