Unicode Escapes
Convert text to Unicode escapes and back. Fixed four-digit or braced form, escape non-ASCII or everything, surrogate pairs handled exactly as JSON writes them.
Waiting for input
How to use
- Paste text, or a string full of escapes, into the input box
- Choose the direction: text into escapes, or escapes back into text
- If you are escaping, pick the form (fixed four digits or the braced code point), how much to escape and the letter case
- Copy the result with the button at the top right of the output
About this tool
A Unicode escape writes any character using only ASCII: a backslash, a u, and the character's code point in hexadecimal, so the Chinese character for middle comes out as \u4e2d. You meet these inside source code more than anywhere else - JSON files, string literals in Java and JavaScript, configuration for systems that only accept ASCII, and the debugging session where some layer between you and the destination quietly dropped everything above 0x7F and you need a way to survive it. This tool runs both directions: it turns text into escapes, and it turns an escaped string back into the words it stands for.
The thing this converter is most often confused with is the hexadecimal one, and the difference is worth stating up front: escapes carry code points, hex carries bytes. The letter A happens to have code point U+0041 and UTF-8 byte 41, so when you only look at English the two seem interchangeable. The character 明 has code point U+660E but occupies three UTF-8 bytes, E6 98 8E, and now the two answers are nothing alike. Pick the wrong tool and nothing complains - you get an answer that looks perfectly plausible and is wrong for every non-ASCII character you own. If you want the bytes, use the hex converter; the results of the two are not substitutes for one another.
There are two spellings and they are not stylistic variants - they differ in reach, which is why you pick rather than guess. The fixed form \uXXXX is exactly four hex digits, which covers the Basic Multilingual Plane up to U+FFFF and nothing beyond it, so emoji and other astral characters have to be written as a pair of surrogates: a smiling face becomes \ud83d\ude00, two escapes for one visible character. That is not a habit this tool invented; it is how UTF-16 and JSON define it, and it is byte for byte what Python writes when you serialise with ensure_ascii enabled. The braced form, \u{1F600} from ES2015 onward, holds one code point per escape and is easier to read, but it only exists in JavaScript string literals and regular expressions - hand it to a JSON parser and it is rejected outright with an invalid-escape error, which was verified rather than remembered. When the output has to travel between languages, use the fixed form.
Decoding here is deliberately narrow: it recognises \uXXXX, the braced code point form, and the two-backslash sequence that stands for one literal backslash. Sequences like \n and \t are left exactly as they are, because they belong to the C and JSON escape families rather than to Unicode escaping, and a converter that claims to do one thing should not silently start doing another. The concrete benefit is that a Windows path survives the round trip: in C:\temp the \t is not quietly turned into a tab, which would shorten your text without any error to explain it. Since the encoder always emits a doubled backslash for every backslash it sees, this narrowness costs nothing - escaping then unescaping returns your input unchanged, character for character, and that invariant is pinned by tests rather than asserted in prose. A lone surrogate, such as \ud800 with no partner, is refused instead of returned: feed that string to a TextEncoder and you will get EF BF BD back in silence, which is the tool telling you it produced a half character only after the damage has travelled downstream. That refusal is the whole point of the options above: each one changes what the output is for, and none of them is allowed to change what the output means.
Frequently asked questions
- What is the difference between \u4e2d and the E4 B8 AD from the hex tool?
- One is a code point, the other is bytes. \u4e2d names the position of that character in the Unicode table and says nothing about how it is stored; E4 B8 AD is what those three bytes actually occupy in a file encoded as UTF-8. You use the first inside source code and JSON, the second when you are checking file contents, hashing or reading a hex dump. English hides the difference entirely: the letter A is code point 41 and UTF-8 byte 41. If you want bytes, go to the hex converter - the two outputs are not interchangeable.
- Why did my emoji turn into two escapes? Is that broken?
- It is correct. The fixed form has four hex digits and therefore reaches only U+FFFF, while emoji live in the supplementary planes above U+10000, where UTF-16 can only express them as a surrogate pair - a smiling face is \ud83d\ude00, two escapes for one character. That is what the JSON specification requires and it is exactly what Python emits when you serialise with ensure_ascii. If you want one escape per character, choose the braced code point form instead, but keep it inside JavaScript: a JSON parser rejects that spelling.
- Can I paste the result straight into JSON or JavaScript?
- Yes for the fixed form - it is valid content inside a JSON string, and the default output (lowercase four digits, only non-ASCII escaped, backslash doubled) matches what Python produces with ensure_ascii turned on, verified character by character. The one thing to watch is that JSON demands exactly four hex digits after \u: neither the braced form nor a short \u4e will do. The braced form belongs to JavaScript literals and regular expressions only.
- Why are \n and \t not turned into a newline and a tab?
- Because they are not Unicode escapes; they belong to the C and JSON escape families, and this converter does one job. The practical gain is that pasting a Windows path does not corrupt it: in C:\temp the \t stays a backslash and a t rather than becoming a tab, which would silently delete a character from your text. If you need whole-JSON-string unescaping, a JSON tool is the right one; if you need Chinese to survive an ASCII-only pipe, that is exactly the \u form handled here.
- My Python program produced garbled output from unicode_escape. Why?
- Because that codec does not work on code points - it treats the text as latin-1 first, so Chinese routed through UTF-8 bytes and then decoded with unicode_escape comes back as pairs of Latin letters instead of \u4e2d\u6587. Use an encoder that works on code points: json.dumps with ensure_ascii set to true gives you the escaped form directly, which is also what this tool produces by default. The tell-tale sign of having taken the wrong path is seeing \u followed by something that is not four hex digits.