The input is parsed the way a browser parses it rather than searched and replaced, because HTML lets you leave out things JSX requires. The end tags of <li>, <p>, <td>, <tr> and <option> are optional, a browser inserts a <tbody> into a table that has none, and void elements such as <br> and <input> never close. Each of those is resolved into the tree first — the <tbody> included, since without it React warns and a server-rendered page fails to hydrate. Anything genuinely broken, such as a <div> that is never closed or tags closed in the wrong order, is reported with its line and column instead of being guessed at. The contents of <script>, <style> and <textarea> are read as text, so a "<" in code cannot open a phantom element.
Attribute names come from the table React itself uses when it warns "Invalid DOM property": class becomes className, for becomes htmlFor, tabindex becomes tabIndex and stroke-width becomes strokeWidth, while data-* and aria-* pass through untouched. Values need converting too, and this is where copied HTML breaks quietly. A style string makes React throw, so it becomes an object with camelCased keys whose values stay strings, because React adds px to a bare number. React reads disabled="" as false, so boolean attributes become bare props. A value or checked on an input would make it read-only, so they become defaultValue and defaultChecked, as do a textarea’s text and an option’s selected. Inline handlers such as onclick cannot become functions, so each is removed and named in a warning.
JSX removes a line break and the indentation around it wherever it sits next to a tag, which is why reformatted HTML loses the space between two links. Spaces are kept where a browser renders them — between inline elements and within text — and dropped beside block-level elements and inside tables, where React warns about whitespace text. A kept space that falls at a line break is written as {' '}. Text is escaped for JSX: braces become {'{'} and {'}'}, and entities stay as written, because JSX decodes the HTML 4 set. It does not decode HTML5 names such as ✓, a hex reference with a capital X, or © without its semicolon; each of those is fixed or reported. Inside <pre>, every space and line break is kept.