Text has no binary form until an encoding is chosen; this tool uses UTF-8, the encoding of the web and of almost every file and API. U+0000 to U+007F take one byte each and match ASCII, which is why A is 01000001 everywhere. Everything else takes more: two bytes up to U+07FF (é, ß, Greek, Cyrillic), three up to U+FFFF (€, most Chinese and Japanese), four beyond that (most emoji). So é, U+00E9, is C3 A9 — sixteen bits, not eight. A converter that shows é as the single byte 11101001 is writing Latin-1 or a bare code point, and its output will not decode as UTF-8 anywhere else. The table under the tool lists each character’s code point and bytes, which is also where a decomposed é — e plus the combining accent U+0301 — shows up as three bytes.
Decoding is strict, because the alternative is silently wrong text. The bytes must be valid UTF-8 as RFC 3629 defines it: no continuation byte without a lead byte, no overlong forms (C0 AF is a two-byte spelling of /, and accepting it was once a way past path checks), no UTF-16 surrogates, nothing above U+10FFFF, and never F5 to FF. The first bad byte is reported by its offset and its place in the input, rather than replaced with U+FFFD. Binary must come in whole bytes: some converters drop the leading zero and write A as 1000001, and a seven-bit group is rejected with that explanation rather than padded on a guess.
In Auto, the format is decided by a fixed precedence, because the same digits can often be read several ways. A prefix decides first: 0b is binary; 0x, \x or % is hex; 0o, or a backslash before digits, is octal. Then only 0s and 1s in groups of seven or more digits is binary — so 10101010 is one binary byte, although it is also four valid hex bytes. Then any letter from A to F, or a single unbroken run of more than three digits, is hex. Then separate groups of digits are decimal, 0 to 255. Octal is never guessed without a prefix, since 110 145 is valid decimal too. When the guess is not valid UTF-8 but another format gives readable text, the error names that format. Spaces, new lines, commas, colons and semicolons all separate bytes. Only UTF-8 is decoded, not UTF-16 or Latin-1, and this converts text, not numbers — the number base converter does that.