Online Toolbox
Interface language: English

HTML Entities

Convert text to HTML entities and back. Choose what gets escaped and how it is written; the decoder refuses code points browsers silently replace.

Waiting for input

How to use

  1. Paste text, or a fragment full of entities, into the input box
  2. Choose the direction: text into entities, or entities back into text
  3. If you are encoding, pick the output style (named, hexadecimal, decimal) and how much to escape
  4. Copy the result with the button at the top right of the output

About this tool

An HTML entity is how you write a character that would otherwise be read as markup rather than as text. Five characters have grammatical jobs inside a document: the ampersand opens an entity, the less-than sign opens a tag, the greater-than sign closes it, and the two quote marks delimit attribute values. Written bare they change what your document means, so they get spelled as & < > " ' or as the numeric forms & and &. This tool works in both directions: it turns text you are about to paste into HTML into something that cannot break the markup, and it turns entities scraped from a page or lifted out of an email source back into words a person can read.

On the encoding side you choose how much to escape and how to write it. How much: only the five dangerous characters, which is what escaping HTML normally means and leaves your Chinese, accents and emoji exactly as they were; the five plus everything outside ASCII, which is what you want when a legacy system, an old mail gateway or somebody else's template pipeline will not carry UTF-8; or literally every character, which produces a dense string of entities. How: named, hexadecimal numeric, or decimal numeric. One deliberate limit is worth stating plainly - the named style applies to those five characters only and everything else comes out numeric. A reverse table of more than two thousand names is not something to type out by hand, and when you are ASCII-fying content the numeric form is what you actually need anyway. One detail worth knowing before you paste: escaping walks over code points rather than over UTF-16 units, so an emoji comes out as the single reference 😀 instead of the two surrogate halves - and those halves are precisely what the decoder refuses, so a tool that got this wrong would produce output it could not read back. Numeric output uses uppercase hex digits, which is what web encoders emit.

Decoding follows the HTML specification, which defines over two thousand named references plus the numeric forms. Those tables are data rather than opinion, so they come from a conforming implementation instead of from transcription, and the result was checked against a second independent one: all 2125 fully named references shipped in the Python standard library decoded to exactly the same character here, with no disagreements and no extra refusals. Several counter-intuitive rules are kept because browsers keep them. A numeric reference in the range 0x80 to 0x9F does not give you that code point - the specification maps it elsewhere, so € is a euro sign and – is an en dash. A name without its trailing semicolon still decodes, because the standard preserves that legacy set, and that is also why a plain ampersand sitting in ordinary text, as in AT&T, is left completely alone instead of being reported as a broken entity.

One class of input is refused rather than decoded: a reference to NUL, one landing in the surrogate range D800 to DFFF, one above the Unicode maximum, and the shape &#x with not a single hex digit after it. A browser cannot refuse, so it substitutes U+FFFD, the replacement character, and carries on drawing. Round-tripping shows what that costs you: decode such a reference, encode the result again, and you get something different from what you started with - the data changed while nobody was looking. So the rule here is measurable rather than stylistic: exactly the shapes a browser silently turns into U+FFFD are the ones reported as errors, and the shapes a browser leaves alone are left alone. The boundaries are worth naming too. Entities go into HTML. Percent-encoding goes into URLs. Backslash-u escapes go into JavaScript source. Base64 carries arbitrary bytes. The hex converter shows you the bytes themselves. All five are ways of writing something that cannot go in as-is, and none of them substitutes for another. Everything runs in your browser and nothing is uploaded.

Frequently asked questions

Why was my Chinese left unconverted?
Because the default escape scope only touches the five characters that carry grammatical meaning in markup. Chinese text is safe inside HTML as long as the document declares the right character set, and turning it into entities makes the source unreadable for no benefit. If you need pure ASCII output, switch the scope to the option that also escapes non-ASCII characters, and each Chinese character becomes something like 中.
Is a missing semicolon after an entity name a mistake?
By the strict grammar, yes, but the specification deliberately keeps a set of names that decode without the trailing semicolon, browsers honour it, and so does this tool: &amp gives you an ampersand and &nbsp gives you a non-breaking space. Two practical consequences. Seeing no semicolon in scraped HTML does not mean the data is broken. And a bare ampersand in ordinary prose - AT&T, or Q and A - is not treated as a malformed entity, because the specification does not read it as one either. What does get reported is a name that ends in a semicolon yet appears nowhere in the standard, because that is almost always somebody trying to write an entity and misspelling it.
Why does 128 as a numeric reference give a euro sign instead of U+0080?
The specification carries a replacement table for numeric references in the range 0x80 to 0x9F: those numbers map to other characters, so 128 becomes the euro sign, 150 an en dash and 151 an em dash. It was added in the late 1990s to keep badly encoded pages readable and it is still what conforming parsers do - Chrome, the Python standard library and this tool all return the same character. There is a small regret buried in it: because that range is remapped, there is no way to obtain a genuine U+0080 control character through an entity at all.
A browser displays that surrogate reference happily. Why does this tool reject it?
The browser is not decoding it into a character. It substitutes U+FFFD, the replacement character, and draws a box, because a browser has to render something and cannot stop. The range D800 to DFFF holds halves of a pair in UTF-16, and half a pair is not a character - the usual cause is an emoji cut in the middle and pasted that way. To write a smiling face properly, use the whole code point as 😀. Reporting this rather than returning a box matters because the two look identical on screen, and whoever receives the box normally believes they received the original character.
Should I use named entities or numeric ones?
For the five common characters, named: they read well and everybody recognises them. Beyond those, numeric is the safer choice, because not every parser knows every name. The apostrophe is the classic case - ' was only adopted into the standard with HTML5, so an older HTML4 parser and some XML toolchains treat it as an undefined entity and fail outright, while ' decodes everywhere. When your output crosses a system you do not control, or gets pasted into a template somebody else wrote, reach for the numeric style. That is also why the named style in this tool applies only to those five characters and everything else comes out numeric.

Related tools

Back to all Encoding Conversion tools

Everything is processed inside your browser · nothing is uploaded · no advertising cookies · Updated 2026-09-29