DevKitHub

Programming

HTML to Markdown Converter — CommonMark and GitHub Markdown

Paste HTML from a web page, a CMS export, Google Docs or an email to get Markdown for a README, a docs site or an LLM prompt. Anything Markdown cannot express is kept as HTML or removed, and every case is listed under the output.

14 lines
17 lines

Text is escaped so it renders back exactly as the HTML did. Markup Markdown has no syntax for, such as <sub> or a table with merged cells, is kept as HTML and listed above.

This tool runs entirely in your browser. Your input is never uploaded, stored or logged.

How it works

The HTML is parsed the way a browser parses it rather than matched with regular expressions. End tags HTML lets you leave out, such as </p>, </li> and </td>, are closed where a browser closes them, entities are decoded from the full HTML5 table, and the contents of <script> and <style> are read as text and then dropped along with comments. Elements a browser does not show are removed too: the hidden attribute, an inline display: none and 1×1 tracking images. Whitespace collapses as it renders, so a run of spaces and line breaks becomes one space, and none at the edge of a block. Each element is then written as Markdown: headings with # or an underline, lists with their start number and nesting, <pre> as a fenced block with the language from class="language-…", and for GitHub Flavored Markdown, tables with column alignment from align or text-align, ~~strikethrough~~ and task lists. Bold and italic written as inline styles, which is how Google Docs exports them, are read from font-weight and font-style.

The step most converters skip is escaping. Text that happens to contain Markdown syntax changes meaning when the Markdown is rendered: a literal * or _ opens emphasis, a backtick opens code, [ and ] start a link, < starts raw HTML, a # or "1." at the start of a line becomes a heading or a list, and & before a name becomes an entity. Each of these is escaped with a backslash, as are | and ~ for GitHub, and bare URLs that GitHub would otherwise turn into links. Emphasis needs more than escaping. CommonMark lets ** open only when the next character is not a space, and not punctuation straight after a letter, so a<b>"x"</b>b cannot be written with asterisks at all. Every bold, italic and strikethrough is checked against the characters that will surround it, and written as <strong>, <em> or <del> where no delimiter would work. The test suite converts a corpus of fragments, renders the Markdown back with this site’s CommonMark renderer and checks that the HTML matches.

The rule for what Markdown cannot express: markup with a meaning of its own — <u>, <sub>, <sup>, <kbd>, <mark>, <abbr title>, <details>, <dl>, <iframe>, <video>, and tables with merged cells or lists inside cells — is kept as raw HTML, which CommonMark passes through; presentational wrappers such as <span>, <div> and <section> are unwrapped; form controls, inline SVG and <canvas> are removed. Tables marked role="presentation", tables that contain tables, and single-column tables without a header are layout, as emails are built, so their cells are unwrapped in order. GitHub strips some raw HTML, such as <iframe>, and links a bare e-mail address whatever the escaping; both are limits of GitHub, not of the conversion.

Common problems

Every example below is run against this tool in our test suite, so what it says here is what the tool actually does.

The <a> tag on line 1 opens an attribute quote that is never closed, so a browser drops the tag and everything after it.

<p>See <a href="https://example.com>the docs</a> for more.</p>
Why:
A missing closing quote makes the attribute value run on to the next quote in the file. With none left, a browser drops the unfinished tag and everything after it, which is why other converters return a page that simply stops.
Fix:
Add the closing quote after the URL: href="https://example.com".

There is no visible content to convert: the HTML holds only scripts, styles, comments, hidden elements or form controls.

<script>window.dataLayer = [];</script><style>body { margin: 0 }</style>
Why:
Pages built by JavaScript send almost no content in their HTML: the text arrives later, so View Source and a saved copy of the page show little but scripts and styles.
Fix:
Copy the rendered HTML instead: in the browser’s developer tools, right-click the article in the Elements panel and choose Copy › Copy outerHTML.

The Markdown shows \<p\>Hello\</p\> instead of a paragraph.

Why:
The HTML was escaped before it was pasted, as &lt;p&gt;, which is how a page shows markup as text. Escaped markup is text, so it is converted as text, with its angle brackets escaped so they stay visible.
Fix:
Decode it first with the HTML entity decoder, then convert the result.

Everything turned bold after pasting from Google Docs.

Why:
Google Docs wraps the whole clipboard in <b style="font-weight:normal" id="docs-internal-guid-…">. A converter that maps <b> to ** without reading the style makes the entire document bold.
Fix:
Nothing to do here: a <b> or <strong> whose style sets font-weight back to normal is not bold, and bold and italic are read from the inline styles of the spans inside.

Bold text shows its asterisks, as a**"x"**b, instead of rendering.

Why:
CommonMark only lets ** open emphasis when it is not followed by a space, and not by punctuation straight after a letter. Converters that wrap every <strong> in ** produce delimiters that never open.
Fix:
Nothing to do here: every delimiter is checked against the characters around it, and where none would work the output uses <strong>, <em> or <del> instead.

Frequently asked questions

Should I choose GitHub Flavored Markdown or CommonMark?
GitHub Flavored Markdown if the result is going to GitHub, GitLab, most static site generators or an LLM prompt: it adds tables, ~~strikethrough~~ and task lists to CommonMark. Choose CommonMark for a renderer that only implements the core specification; tables are then kept as HTML and task-list checkboxes are removed, because CommonMark has no syntax for them.
Why is some HTML left in the Markdown?
Because Markdown has no syntax for it and dropping it would lose meaning: underline, subscript and superscript, keyboard keys, <details>, definition lists, embedded video, and tables with merged cells or lists in their cells. CommonMark renders raw HTML as it is, so the result still displays correctly. Wrappers with no meaning of their own, such as <span> and <div>, are removed and their contents converted.
Why are there backslashes in the output?
Each one stops a character in your text being read as Markdown. A literal *, _ or [ in a sentence would otherwise start emphasis or a link, and a # or "1." at the start of a line would become a heading or a list. The backslashes do not appear when the Markdown is rendered.
Can it convert a whole web page?
Yes. The <head>, scripts, styles, comments and hidden elements are dropped, and navigation links are kept as ordinary links, so you may want to delete the site chrome at the top and bottom afterwards. For a page that builds its content with JavaScript, copy the rendered HTML from the browser’s developer tools rather than the page source.
Is my HTML uploaded anywhere?
No. The conversion runs entirely in your browser; large documents are converted in a background worker on your own machine, and nothing you paste is sent to a server.

Last updated