DevKitHub

JSON & Data

XML to JSON Converter (and JSON to XML)

Paste XML, such as an RSS feed, a SOAP response or a pom.xml, to get JSON, or switch direction to turn JSON into XML. Attributes become @ keys, and every element that became an array is named, so you can keep its shape fixed.

20 lines
30 lines
Document
<rss>, 15 elements
  • <item> became an array because it repeats. A document with only one would give a single value instead, which is how code reading converted XML usually breaks; add item to Always arrays if the count can vary.

The mapping is xmltodict’s: attributes are @ keys, text beside them is #text, and an element that repeats is an array. An element that appears once is a single value, so list any that can repeat in Always arrays for. Every value is a string unless Convert numbers and booleans is on, and comments, processing instructions and the DOCTYPE are dropped.

This tool runs entirely in your browser. Your input is never uploaded, stored or logged.

How it works

XML and JSON do not map onto each other, so every converter picks a convention, and this one follows xmltodict’s, the most widely used. The root element is the single top-level key, an attribute becomes a key with @ in front (@id), an element’s text becomes a string, and when the element also has attributes or children its text goes under #text; the other common shape, attributes grouped under _attributes and text under _text, is one switch away. An element that appears more than once becomes an array, which is where most bugs come from: a feed with two <item> elements gives an array, the same feed with one gives an object, and code that loops over it breaks on the day there is only one. Always arrays takes a list of element names that are arrays however many there are, and the result names every element that became an array, so you know which to list. Empty elements become null, namespace prefixes stay part of the name, as in soap:Envelope, and xmlns declarations are kept as attributes unless you drop them.

Some of XML has no JSON form. Comments and processing instructions are dropped and counted. In mixed content, text sitting between child elements as in HTML-like markup, the text is joined into #text and its position among the elements is lost, and a warning names the element; the order of a repeated element interleaved with others is lost the same way. Every value is a string, because XML has no types. Convert numbers and booleans changes only text that prints back unchanged, so 42 and true convert while 007, 1e5 and 1.50 stay strings, as zip codes and version numbers must. Text is trimmed unless you turn that off, CDATA becomes ordinary text, the five predefined entities and numeric character references are decoded, and line ends and whitespace in attribute values are normalised as XML 1.0 requires.

The parser is strict where lenient ones guess: a bare & in a URL, an HTML entity such as &nbsp;, an attribute with no value or no quotes, a second root element and a mismatched closing tag are each an error with a line and column. A DOCTYPE is skipped, and one that declares entities is refused, because <!ENTITY> is how XML external entity (XXE) attacks read files and how a billion-laughs document grows to gigabytes; nothing here runs on a server, but no data needs them. From JSON to XML, the key of a one-key object becomes the root element and anything else is wrapped in a root you name; @ keys become attributes, arrays become repeated elements, & < and > are escaped, keys that are not XML names are renamed with a warning, and numbers are copied exactly as written.

Common problems

Every example below is run against this tool in our test suite, so what it says here is what the tool actually does.

A bare & must be written as &amp; in XML, including in a URL such as ?a=1&amp;b=2.

<link>https://example.com/search?q=xml&page=2</link>
Why:
A literal & in text or an attribute value. In XML an & always starts an entity or character reference, so an unescaped one in a query string, or in a name like AT&T, is a syntax error. HTML parsers forgive it; XML parsers do not.
Fix:
Write it as &amp;, or wrap the text in <![CDATA[ … ]]>. Whatever produced the XML should have escaped it.

&nbsp; is an HTML entity, and XML predefines only &lt; &gt; &amp; &quot; and &apos;. Write &#160; instead.

<p>Price:&nbsp;10</p>
Why:
XML defines five named entities. &nbsp;, &copy;, &mdash; and the rest come from HTML’s DTD, which an XML parser does not have, so it cannot know what they stand for.
Fix:
Use the numeric reference, &#160; for a non-breaking space, or type the character itself.

<item> is a second root element.

<item>1</item>
<item>2</item>
Why:
Two top-level elements, usually a fragment copied out of a larger document or several records pasted one after another. An XML document has exactly one root element.
Fix:
Wrap them in one element, such as <items>…</items>, and convert that.

Entity declarations (<!ENTITY …>) are refused, not expanded.

<?xml version="1.0"?>
<!DOCTYPE foo [ <!ENTITY xxe SYSTEM "file:///etc/passwd"> ]>
<foo>&xxe;</foo>
Why:
The DOCTYPE declares entities. Declared entities are how an XXE attack makes a parser read a file such as /etc/passwd, and how the billion laughs document nests entities until a few hundred bytes expand to gigabytes. Safe parsers disable them, and this one refuses them.
Fix:
Replace each entity reference with the text it stands for, and delete the <!DOCTYPE … [ … ]> block.

The attribute disabled in <input> has no value.

<input type="checkbox" disabled/>
Why:
An HTML boolean attribute. HTML allows an attribute with no value, but XML requires every attribute to have one, in quotes.
Fix:
Give it a value, as in disabled="disabled", or read the page with an HTML parser rather than an XML one.

A list with one element came out as an object, and the code that loops over it broke.

Why:
An element that repeats becomes an array and one that does not becomes a single value, so the shape of the JSON depends on how many there happen to be. Every convention that maps repeated elements to arrays has this problem.
Fix:
Add the element names to Always arrays, such as item or dependency, or turn on Every element an array.

Every value in the JSON is a string, even 42 and true.

Why:
XML has no data types: <count>42</count> holds the text 42. Guessing would also turn a zip code such as 007 into the number 7.
Fix:
Turn on Convert numbers and booleans. It converts only text that prints back unchanged, so 007, 1e5 and 1.50 stay strings.

Frequently asked questions

Why is one element an object in the JSON but two are an array?
Because the converter only sees this document. When an element repeats it becomes an array; when it appears once there is nothing to say it could repeat, so it becomes a single value. List the element names in Always arrays, or turn on Every element an array, to get the same shape whatever the count.
How are XML attributes represented in JSON?
As keys with @ in front, beside the child elements: <book id="bk101"> gives "@id": "bk101", and the element’s text goes under #text when it has attributes too. This is the convention of Python’s xmltodict. The other style puts all attributes in an _attributes object and the text under _text.
Is converting XML to JSON and back lossless?
Not entirely. Comments, processing instructions and the DOCTYPE are dropped, every value comes back as a string, text mixed between child elements loses its position, and repeated elements with others in between are regrouped. Data-shaped XML, such as feeds, configuration and API responses, survives the round trip.
What happens to XML namespaces?
Prefixes stay part of the names, so soap:Envelope is a key as written, and xmlns declarations are kept as @xmlns attributes so the XML converts back to a valid document. Drop them with Keep xmlns turned off. The converter does not resolve prefixes to namespace URIs.
Why is XML with <!ENTITY> declarations rejected?
Entity declarations are the mechanism behind XML external entity (XXE) attacks and the billion laughs denial of service, and no data needs them. This page runs in your browser, so nothing is at risk, but it refuses them as a safe server-side parser would. A DOCTYPE without entity declarations is skipped with a warning.

Last updated