Hex to Text
Convert text to hexadecimal and back, byte by byte in UTF-8. Choose letter case, a 0x prefix and the separator; invalid bytes are reported instead of guessed.
Waiting for input
How to use
- Paste text or a hex string into the input box
- Choose the direction: text to hex, or hex back to text
- Set letter case, the 0x prefix and the separator (space, comma, newline or none) if you need them
- Copy the result from the button at the top right of the output
About this tool
This tool does two symmetric jobs. Encoding takes text, turns it into bytes using UTF-8, and writes each byte as two hexadecimal digits. Decoding reads a string of hex digits, groups them in pairs, and interprets those bytes as UTF-8 text. The examples make the shape clear: the letter A is one byte, 41, while a character such as 明 occupies three bytes, E6 98 8E, so the two-character string A明 becomes 41 E6 98 8E, four bytes written as eight digits. Most real uses amount to reading what a machine left behind. You inspect a file signature in a hexdump and recognise a PNG by its first eight bytes, 89 50 4E 47 0D 0A 1A 0A. You look for a byte order mark, EF BB BF, to find out why a script parses fine until it hits the first line. You compare one field of a device protocol against its specification, or you need to paste a handful of raw bytes into a source file as an array literal.
The distinction that matters most, and the one that separates a correct result from a plausible one, is that this tool converts bytes rather than code points. For the ASCII range the two readings coincide, which is why the difference stays hidden for anyone who only tests English: the letter A has code point U+0041 and its UTF-8 byte is also 41, so both interpretations print 41. Past ASCII they diverge completely. A Chinese character such as 中 has code point U+4E2D, which is four hex digits and looks tidy, but it is encoded as three UTF-8 bytes, E4 B8 AD. Neither answer is wrong; they are answers to different questions, and pasting one into a context expecting the other produces text that decodes to the wrong characters. If you want the code-point spelling, that is the Unicode escape tool, and its output is not interchangeable with this one. The same byte-level view explains percent-encoding in URLs: a query value that reads %E4%B8%AD is not an odd notation, it is exactly the UTF-8 bytes of one Chinese character written in hex, which is why hex literacy pays off when debugging links containing non-ASCII text.
The second thing to understand is that hex digits alone do not determine text; the encoding does. The four bytes D6 D0 C4 E4 are perfectly readable Chinese under GBK and are not a valid UTF-8 sequence at all, so the same input is meaningful in one system and garbage in another. This tool decodes UTF-8 and nothing else, and when the bytes do not form a valid UTF-8 sequence it reports that instead of producing output. That choice is deliberate. A decoder configured for tolerance would substitute a replacement character for each bad byte and hand back a string that looks successful while being quietly wrong, and the people affected by that behaviour usually find out later, after the string has been stored, compared or signed. The same reasoning covers the byte order mark: a leading EF BB BF decodes to U+FEFF here rather than being silently removed, because a converter that deletes input it does not find useful is not a converter.
On input the parser is deliberately forgiving, because the sources of hex are varied. It accepts an optional 0x or 0X prefix on every byte, upper or lower case digits, and any mix of spaces, commas and newlines as separators, so a C array copied straight out of source code, a log line and a compact digest all parse without editing. The one thing it will not swallow is genuine garbage: characters such as a dollar sign or a percent sign are rejected, and an odd digit count is reported separately from an illegal character, because those two mistakes have different fixes and the second usually means a byte was split. On output nothing is implicit; you choose the case, whether each byte carries a prefix, and how bytes are separated, including the compact form, so the result can be pasted back where it came from. Finally, a sizing note when choosing an encoding: hexadecimal costs exactly two characters per byte, always double the original size, while Base64 costs about four thirds. If a human will read the value, or compare it byte by byte, hex is worth the space; if the bytes just have to survive a text channel, Base64 carries them more cheaply.
Frequently asked questions
- Why will my hex not decode into text?
- Three possibilities, and the message on the page tells you which one you hit: an odd number of digits (usually a character lost while copying), a character outside 0-9 and a-f (only an optional 0x prefix plus spaces, commas or newlines are accepted), or byte pairs that do not form valid UTF-8, which means a GBK byte stream or data that was never text. That last case cannot be solved by guessing: every workaround gives you wrong data that looks fine.
- Should I use hex or Base64 for binary data?
- Depends on who reads it. Hex always costs two characters per byte, exactly double the size, but it maps one-to-one onto bytes, which is why checksums, hashes and key fingerprints are written in hex: you can see which byte is which at a glance. Base64 costs about 1.33 times the size and carries arbitrary bytes safely through text-only channels, so use it for email attachments, embedded images and MIME.
- Are these bytes or Unicode code points?
- Bytes. A has code point U+0041 and its UTF-8 byte is 41 too, so both readings agree there. A Chinese character such as 中 has code point U+4E2D but UTF-8 bytes E4 B8 AD, and this tool outputs the latter. Reach for the Unicode escape tool when you specifically need the \u4E2D spelling.
- Does byte order, big-endian or little-endian, matter here?
- Not for this tool. It walks the bytes in order and writes each one separately, so there is no multi-byte integer to interpret and no order to get wrong. Byte order only becomes relevant when you read a group of hex digits as a number, for example turning 12 34 AB CD into either 0x1234ABCD or 0xCDAB3412 depending on the convention your format uses. That is a different job, and it belongs to a radix or byte-order converter. What you see here is the data in the order it sits in the file, which is also the order a protocol specification lists its fields.
- How do I turn a hex dump back into a file?
- First ask whether the bytes can be written as text at all, because that is the limit of this tool: anything that is not valid UTF-8 will be reported rather than shown, and a JPEG or a zip archive is not valid UTF-8. For real binary files the practical route is to carry the bytes in Base64, which survives any text channel, and then write them out with a system command or a browser download. Hexadecimal works too, and some formats and tools insist on it, but it costs twice the space and every line must stay byte aligned, so losing one digit while copying shifts the whole file by half a byte and corrupts everything after that point silently.