DevKitHub

Encoding & Conversion

Text to Binary Converter — Binary, Hex and ASCII Translator

Type text to see its bytes in binary, hex, octal or decimal, or paste bytes to turn them back into text. Every character is listed with its code point and its UTF-8 bytes.

1 line
1 line

Bytes

Bytes
5
Characterscode points
4

Each character in UTF-8

Each character, its code point and its UTF-8 bytes
CharacterCode pointUTF-8 bytesBytes
cU+0063011000111
aU+0061011000011
fU+0066011001101
éU+00E911000011 101010012

Text is encoded as UTF-8: é is two bytes, € three and an emoji four. Converters that show one byte per character only do so because they ignore everything outside ASCII.

This tool runs entirely in your browser. Your input is never uploaded, stored or logged.

How it works

Text has no binary form until an encoding is chosen; this tool uses UTF-8, the encoding of the web and of almost every file and API. U+0000 to U+007F take one byte each and match ASCII, which is why A is 01000001 everywhere. Everything else takes more: two bytes up to U+07FF (é, ß, Greek, Cyrillic), three up to U+FFFF (€, most Chinese and Japanese), four beyond that (most emoji). So é, U+00E9, is C3 A9 — sixteen bits, not eight. A converter that shows é as the single byte 11101001 is writing Latin-1 or a bare code point, and its output will not decode as UTF-8 anywhere else. The table under the tool lists each character’s code point and bytes, which is also where a decomposed é — e plus the combining accent U+0301 — shows up as three bytes.

Decoding is strict, because the alternative is silently wrong text. The bytes must be valid UTF-8 as RFC 3629 defines it: no continuation byte without a lead byte, no overlong forms (C0 AF is a two-byte spelling of /, and accepting it was once a way past path checks), no UTF-16 surrogates, nothing above U+10FFFF, and never F5 to FF. The first bad byte is reported by its offset and its place in the input, rather than replaced with U+FFFD. Binary must come in whole bytes: some converters drop the leading zero and write A as 1000001, and a seven-bit group is rejected with that explanation rather than padded on a guess.

In Auto, the format is decided by a fixed precedence, because the same digits can often be read several ways. A prefix decides first: 0b is binary; 0x, \x or % is hex; 0o, or a backslash before digits, is octal. Then only 0s and 1s in groups of seven or more digits is binary — so 10101010 is one binary byte, although it is also four valid hex bytes. Then any letter from A to F, or a single unbroken run of more than three digits, is hex. Then separate groups of digits are decimal, 0 to 255. Octal is never guessed without a prefix, since 110 145 is valid decimal too. When the guess is not valid UTF-8 but another format gives readable text, the error names that format. Spaces, new lines, commas, colons and semicolons all separate bytes. Only UTF-8 is decoded, not UTF-16 or Latin-1, and this converts text, not numbers — the number base converter does that.

Common problems

Every example below is run against this tool in our test suite, so what it says here is what the tool actually does.

Not valid UTF-8 at byte offset 0 (0xE9)

11101001
Why:
The binary came from a converter that writes é as one byte — its Latin-1 value, which is also its code point, 233. In UTF-8 é is two bytes, and E9 on its own starts a three-byte character that never arrives.
Fix:
Convert from the original text here to get the UTF-8 bytes, 11000011 10101001. If the bytes really are Latin-1, they need re-encoding before any UTF-8 tool can read them.

Every group is 7 bits, not 8.

1001000 1101001
Why:
ASCII written without its leading zero. Seven bits are enough for ASCII, so some converters — and Python's bin() — leave the zero off, but a byte is eight bits and UTF-8 needs all of them.
Fix:
Add a 0 to the front of each group: 01001000 01101001 is "Hi".

The input has an odd number of hex digits.

48656c6c6
Why:
Each byte is exactly two hex digits, so an odd count means one was lost — usually the last character of a copy, or the leading zero of a byte such as 0A written as A.
Fix:
Copy the value again, or restore the missing zero. With separators, every group must be two digits as well.

Group 3 is more than 255, the most a byte can hold.

72 105 8364
Why:
Decimal here means bytes, from 0 to 255. 8364 is the code point of €, not a byte: in UTF-8, € is the three bytes 226 130 172.
Fix:
Convert from the text instead, or turn each code point into its UTF-8 bytes first. The character table shows them.

é comes out as three bytes instead of two.

Why:
The é is decomposed: the letter e followed by the combining acute accent U+0301, as file names from macOS often are. It looks identical to the single character U+00E9.
Fix:
Normalise the text to NFC first — text.normalize('NFC') in JavaScript. The tool warns when text is decomposed, and the table shows both code points.

Frequently asked questions

Why is é 16 bits in binary?
Because UTF-8 needs two bytes for it: U+00E9 becomes C3 A9, which is 11000011 10101001. Only the 128 ASCII characters fit in one byte. A converter that shows é as 8 bits is using Latin-1, and that output is not valid UTF-8.
What is the difference between ASCII and UTF-8?
ASCII defines 128 characters, one byte each. UTF-8 encodes every Unicode character in one to four bytes, and its first 128 are byte for byte the same as ASCII. So plain English text gives identical bytes in both; for anything else, only UTF-8 is defined.
How do I convert hex to ASCII?
Switch to Binary → Text and paste the hex. Auto reads it with or without spaces, and with 0x, \x or % prefixes. ASCII is a subset of UTF-8, so ASCII hex decodes the same either way, and anything from 80 upwards is read as UTF-8.
Why does 10101010 decode as binary rather than hex?
It is valid as both, and binary takes precedence when the input is only 0s and 1s in groups of seven or more digits. Choose Hex under Format to read it as the four bytes 10 10 10 10.

Last updated