Skip to content
Browse tools

Data formats guide Β· XML

Your own tags.
Strict rules.
Still everywhere.

XML β€” Extensible Markup Language β€” describes structured data with tags you invent yourself. It lost the web-API race to JSON, but it quietly runs sitemaps, RSS feeds, Office documents, SOAP services and half the config files in the .NET and Java worlds. Here's how to read it, write it and query it. πŸ‘‡

🏷️
Self-describing
Custom tags

You name the elements. <invoice> means what your system says it means.

πŸ“
Contract-first
Schemas

DTD and XSD let a document be checked against an agreed structure before anyone trusts it.

πŸ“„
Document-shaped
Mixed content

Text with markup inside it β€” the job XML was designed for and JSON handles poorly.

πŸ“œ

Where it came from: XML 1.0 became a W3C Recommendation in 1998 as a simplified subset of SGML, the same ancestor HTML came from. XML 1.0 (Fifth Edition) is the version virtually everything uses today.

02 Β· The building blocks

🧱 Elements, attributes & namespaces

An XML document is a tree. Here's a small one with every common part labelled in the comments:

β–Έ catalogue.xml
<?xml version="1.0" encoding="UTF-8"?>          <!-- declaration (prolog) -->
<catalogue xmlns="https://example.com/ns/shop">   <!-- root element + default namespace -->
  <product id="TEA-250" currency="INR">   <!-- element with attributes -->
    <name>Assam Tea 250g</name>             <!-- element with text content -->
    <price>249.00</price>
    <tags>
      <tag>gift</tag>
      <tag>organic</tag>
    </tags>
    <discontinued/>                      <!-- self-closing empty element -->
  </product>
</catalogue>
PartWhat it isRule of thumb
🏷️ ElementA start tag, content and end tag: <name>…</name>Use for the data itself, especially anything that repeats or nests.
πŸ”– AttributeA name/value pair on a start tag: id="TEA-250"Use for metadata β€” IDs, units, flags. An attribute can't repeat or hold structure.
🌳 RootThe single outermost elementExactly one per document.
πŸ“ Declaration<?xml version="1.0" encoding="UTF-8"?>Optional for UTF-8 documents, but if present it must be the very first thing in the file.
πŸ’¬ Comment<!-- … -->Can't contain -- and can't be nested.

🌐 Namespaces in one paragraph

Because anyone can invent tag names, two vocabularies can collide β€” your <title> and someone else's. A namespace ties names to a URI so they stay distinct. xmlns="…" sets the default namespace for an element and its children; xmlns:img="…" declares a prefix you then write as <img:image>. The URI is just an identifier β€” nothing is ever downloaded from it.

πŸͺ€

The number one XPath bug: a document with a default namespace (like the sitemap below) looks un-namespaced, but every element in it is namespaced. Queries for plain //url then match nothing. You must register the namespace with your parser and query with a prefix.

03 Β· Two levels of correct

βœ… Well-formed vs valid

XML has two distinct bars, and people mix them up constantly.

🧩 Well-formed

Syntax Β· always required

  • Exactly one root element
  • Every start tag closed (or self-closed)
  • Properly nested β€” no <b><i></b></i>
  • Names are case-sensitive: <Name> β‰  <name>
  • Attribute values quoted; no duplicate attributes
  • < and & escaped in text

Break any of these and a conforming parser must stop with an error β€” no guessing, unlike HTML.

πŸ“ Valid

Structure Β· optional

A well-formed document that also matches a schema: which elements are allowed, in what order, how many times, and what data types their values have.

  • DTD β€” the original, built into XML 1.0. Limited types.
  • XSD (XML Schema) β€” the common choice today; typed values like dates and decimals.
  • RELAX NG β€” a simpler alternative used by some document formats.

A formatter or parser tells you whether a document is well-formed. Checking validity needs the schema too. Paste a document into the XML Formatter and it will point at the first well-formedness error.

04 Β· Special characters

πŸ”£ Entities, escaping & CDATA

Because < starts a tag and & starts an entity, those characters can't appear raw in text. XML predefines exactly five entities:

CharacterEntityWhen you must escape it
<&lt;Always, in text and attribute values
&&amp;Always β€” the most common real-world error
>&gt;Only strictly required in the sequence ]]>, but escaping it is harmless
"&quot;Inside an attribute value delimited by double quotes
'&apos;Inside an attribute value delimited by single quotes

Any Unicode character can also be written as a numeric reference: &#233; or &#xE9; for Γ©. Note that HTML's named entities such as &nbsp; and &copy; are not defined in XML unless a DTD declares them.

β–Έ escaping vs CDATA
<!-- escaped -->
<query>price &lt; 500 &amp;&amp; stock &gt; 0</query>

<!-- CDATA: everything up to ]]> is literal text -->
<query><![CDATA[price < 500 && stock > 0]]></query>
πŸ’‘

Let the library escape for you. Build XML with a proper API rather than string concatenation and escaping is handled correctly every time. For one-off snippets, the HTML Encoder / Decoder converts the same five characters.

05 Β· Where you'll meet it

πŸ—ΊοΈ XML vs JSON, and where XML still wins

WhatXMLJSON
StructureElements, attributes, text β€” orderedObjects, arrays, six value types
MetadataAttributes and namespaces built inOnly by convention (extra keys)
Mixed contentNatural: <p>Hi <b>there</b></p>Awkward
CommentsYesNo
Validation & queryingXSD, XPath, XSLT β€” mature, standardisedJSON Schema, JSONPath-style tools
VerbosityHigher β€” every element closes by nameLower
Browser-nativeVia DOMParserJSON.parse β€” the default for web APIs

For a fresh web API, JSON is the sensible default β€” see What is JSON?. But you'll keep running into XML here:

WhereExamples
🧭 Sitemaps & feedssitemap.xml, RSS and Atom feeds, podcast feeds
🧼 Enterprise servicesSOAP web services, WSDL contracts, SAML single sign-on, banking and e-invoicing formats
πŸ“Ž Documents.docx, .xlsx and .pptx are ZIP files full of XML (Office Open XML); so are EPUB books. SVG images are XML too.
βš™οΈ Build & config.NET .csproj and web.config, Maven pom.xml, Android layouts and manifests

A sitemap is a good real example of a namespaced document:

β–Έ sitemap.xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/</loc>
    <lastmod>2026-03-14</lastmod>
  </url>
</urlset>

Want to check a real site's sitemap β€” every URL, its status code and whether it resolves? The Sitemap Checker finds and parses it for you.

06 Β· Querying

🎯 XPath basics

XPath is a small path language for selecting nodes in an XML tree. It's built into most XML libraries, browser DevTools and XSLT. Against the catalogue.xml example above (ignoring its namespace for a moment):

ExpressionSelects
/catalogue/productEvery product that is a direct child of the root
//tagEvery tag element anywhere in the document
//product/@idThe id attribute of every product
//product[@currency='INR']Products whose currency attribute is INR
//product[price > 200]/nameNames of products priced above 200
//product[1]The first product within its parent β€” XPath counts from 1, not 0
//name/text()The text nodes inside each name
count(//product)How many products there are
πŸ§ͺ

Try it in the browser: open any XML file or HTML page, go to the DevTools console and run $x("//a") in Chrome, Edge or Firefox. It evaluates an XPath expression against the current document.

07 Β· In your code

πŸ› οΈ Parsing XML safely

Each snippet reads the sitemap above β€” namespace included, because that's what real documents look like.

β–Έ C# Β· LINQ to XML
using System.Xml.Linq;

XNamespace sm = "http://www.sitemaps.org/schemas/sitemap/0.9";
var doc = XDocument.Parse(xml);

var urls = doc.Descendants(sm + "url")
              .Select(u => (string?)u.Element(sm + "loc"))
              .ToList();
β–Έ Python Β· ElementTree
import xml.etree.ElementTree as ET

ns = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
root = ET.fromstring(xml_text)
urls = [loc.text for loc in root.findall("sm:url/sm:loc", ns)]
β–Έ JavaScript Β· browser
const doc = new DOMParser().parseFromString(text, "application/xml");

// DOMParser doesn't throw β€” it returns a document containing <parsererror>
if (doc.querySelector("parsererror")) throw new Error("Not well-formed");

const urls = [...doc.getElementsByTagNameNS(
  "http://www.sitemaps.org/schemas/sitemap/0.9", "loc"
)].map(el => el.textContent);
πŸ›‘οΈ

Untrusted XML needs care. DTDs can define entities that expand exponentially (the "billion laughs" attack) or pull in local files and URLs (XXE, XML External Entity injection). Keep DTD processing disabled unless you genuinely need it β€” modern .NET's XmlReader prohibits DTDs by default β€” and for Python the standard library docs recommend the defusedxml package when parsing data you don't control.

Got a wall of minified XML from an API log? The XML Formatter indents it, checks it's well-formed and lets you explore it as a tree. If the same data also comes as JSON, compare it in the JSON Formatter.

πŸ“Œ Syntax on this page follows the W3C XML 1.0 and Namespaces in XML recommendations. XPath examples use XPath 1.0, which is what browsers and most standard libraries implement.

What is XML?
8 min read