TanodTools
EN

HTML entity encoder and decoder

Turn text into HTML entities, or turn entities such as & and € back into the characters they stand for.

Runs in your browser; nothing you type leaves your device

What do you want to do?

Paste text to begin.

Text up to a few million characters. Decoding understands every named entity your browser knows (the full HTML5 list) and decimal and hexadecimal numbers; a reference needs its closing semicolon. Decoding is one level only: &amp;lt; becomes &lt;, not <.

How to encode and decode HTML entities

  1. Choose Encode or Decode and paste your text.
  2. When encoding, choose whether to encode only the unsafe characters, to use names such as &eacute; where they exist, or to use numbers for everything outside ASCII.
  3. Copy the result.

What HTML entities are

HTML gives a few characters a special job: the less-than sign starts a tag, the ampersand starts an entity, and the quotes end an attribute value. To show one of those characters as text, you write an entity instead: &lt; for <, &amp; for &. Entities also let you write characters that are hard to type or that an old page encoding cannot hold, such as &eacute; for é, &euro; for € or &#128512; for an emoji.

An entity starts with an ampersand and ends with a semicolon. Between them is either a name (copy), a decimal number after # (#169) or a hexadecimal number after #x (#xA9). All three of those write the copyright sign. Encoding goes through your text one character at a time: the five unsafe characters always become entities, and the rest depends on the option you choose, from nothing at all to a number for every character.

Decoding reads each reference and replaces it with its character. Numbers are converted directly, following the HTML rules: a number that is zero, out of range or a surrogate becomes the replacement character, and the numbers 128 to 159 are read as Windows-1252 characters, as browsers do. Names are resolved by your browser, which knows the whole HTML5 list. Decoding is a single pass, so &amp;lt; becomes &lt; and not a bare less-than sign.

Tips

Questions

Which characters must be escaped in HTML?

Five: the less-than sign < (&lt;), the greater-than sign > (&gt;), the ampersand & (&amp;), the double quote " (&quot;) and the single quote ' (&#39;). Escaping them lets you show code or user input as text without it being read as markup. The first option, Only the unsafe characters, does just that.

What is the difference between named and numeric entities?

A named entity uses a short word, such as &eacute; for é or &euro; for €. A numeric one uses the Unicode code point, in decimal (&#233;) or hexadecimal (&#xE9;). Both give the same character. Only some characters have names; the rest need numbers.

Why encode non-ASCII characters at all?

If a page or an email is not served as UTF-8, accented letters and symbols can turn into garbage. Entities are plain ASCII, so they survive any encoding.

Does decoding cover every entity?

Yes. Decimal and hexadecimal references are converted directly, and names are looked up by the browser's own HTML parser, which knows the complete HTML5 list of more than two thousand names. A reference without its semicolon is left as it is, and so is an unknown name, which is listed so you can see it.

Is encoding enough to stop XSS?

Encoding the five unsafe characters is the right fix when you place text in HTML element content or in a quoted attribute value. It is not enough inside a script, a style, an unquoted attribute or a URL, which need their own escaping. This tool is for text you paste, not a replacement for your framework's escaping.

Is my text uploaded?

No. Everything happens in your browser and nothing leaves your device.