Data formats guide Β· XML
Your own tags.
Strict rules.
Still everywhere.
XML β Extensible Markup Language β describes structured data with tags you invent yourself. It lost the web-API race to JSON, but it quietly runs sitemaps, RSS feeds, Office documents, SOAP services and half the config files in the .NET and Java worlds. Here's how to read it, write it and query it. π
You name the elements. <invoice> means what your system says it means.
DTD and XSD let a document be checked against an agreed structure before anyone trusts it.
Text with markup inside it β the job XML was designed for and JSON handles poorly.
Where it came from: XML 1.0 became a W3C Recommendation in 1998 as a simplified subset of SGML, the same ancestor HTML came from. XML 1.0 (Fifth Edition) is the version virtually everything uses today.
02 Β· The building blocks
π§± Elements, attributes & namespaces
An XML document is a tree. Here's a small one with every common part labelled in the comments:
<?xml version="1.0" encoding="UTF-8"?> <!-- declaration (prolog) --> <catalogue xmlns="https://example.com/ns/shop"> <!-- root element + default namespace --> <product id="TEA-250" currency="INR"> <!-- element with attributes --> <name>Assam Tea 250g</name> <!-- element with text content --> <price>249.00</price> <tags> <tag>gift</tag> <tag>organic</tag> </tags> <discontinued/> <!-- self-closing empty element --> </product> </catalogue>
| Part | What it is | Rule of thumb |
|---|---|---|
| π·οΈ Element | A start tag, content and end tag: <name>β¦</name> | Use for the data itself, especially anything that repeats or nests. |
| π Attribute | A name/value pair on a start tag: id="TEA-250" | Use for metadata β IDs, units, flags. An attribute can't repeat or hold structure. |
| π³ Root | The single outermost element | Exactly one per document. |
| π Declaration | <?xml version="1.0" encoding="UTF-8"?> | Optional for UTF-8 documents, but if present it must be the very first thing in the file. |
| π¬ Comment | <!-- β¦ --> | Can't contain -- and can't be nested. |
π Namespaces in one paragraph
Because anyone can invent tag names, two vocabularies can collide β your <title> and someone else's. A namespace ties names to a URI so they stay distinct. xmlns="β¦" sets the default namespace for an element and its children; xmlns:img="β¦" declares a prefix you then write as <img:image>. The URI is just an identifier β nothing is ever downloaded from it.
The number one XPath bug: a document with a default namespace (like the sitemap below) looks un-namespaced, but every element in it is namespaced. Queries for plain //url then match nothing. You must register the namespace with your parser and query with a prefix.
03 Β· Two levels of correct
β Well-formed vs valid
XML has two distinct bars, and people mix them up constantly.
π§© Well-formed
Syntax Β· always required
- Exactly one root element
- Every start tag closed (or self-closed)
- Properly nested β no
<b><i></b></i> - Names are case-sensitive:
<Name>β<name> - Attribute values quoted; no duplicate attributes
<and&escaped in text
Break any of these and a conforming parser must stop with an error β no guessing, unlike HTML.
π Valid
Structure Β· optional
A well-formed document that also matches a schema: which elements are allowed, in what order, how many times, and what data types their values have.
- DTD β the original, built into XML 1.0. Limited types.
- XSD (XML Schema) β the common choice today; typed values like dates and decimals.
- RELAX NG β a simpler alternative used by some document formats.
A formatter or parser tells you whether a document is well-formed. Checking validity needs the schema too. Paste a document into the XML Formatter and it will point at the first well-formedness error.
04 Β· Special characters
π£ Entities, escaping & CDATA
Because < starts a tag and & starts an entity, those characters can't appear raw in text. XML predefines exactly five entities:
| Character | Entity | When you must escape it |
|---|---|---|
| < | < | Always, in text and attribute values |
| & | & | Always β the most common real-world error |
| > | > | Only strictly required in the sequence ]]>, but escaping it is harmless |
| " | " | Inside an attribute value delimited by double quotes |
| ' | ' | Inside an attribute value delimited by single quotes |
Any Unicode character can also be written as a numeric reference: é or é for Γ©. Note that HTML's named entities such as and © are not defined in XML unless a DTD declares them.
<!-- escaped --> <query>price < 500 && stock > 0</query> <!-- CDATA: everything up to ]]> is literal text --> <query><![CDATA[price < 500 && stock > 0]]></query>
Let the library escape for you. Build XML with a proper API rather than string concatenation and escaping is handled correctly every time. For one-off snippets, the HTML Encoder / Decoder converts the same five characters.
05 Β· Where you'll meet it
πΊοΈ XML vs JSON, and where XML still wins
| What | XML | JSON |
|---|---|---|
| Structure | Elements, attributes, text β ordered | Objects, arrays, six value types |
| Metadata | Attributes and namespaces built in | Only by convention (extra keys) |
| Mixed content | Natural: <p>Hi <b>there</b></p> | Awkward |
| Comments | Yes | No |
| Validation & querying | XSD, XPath, XSLT β mature, standardised | JSON Schema, JSONPath-style tools |
| Verbosity | Higher β every element closes by name | Lower |
| Browser-native | Via DOMParser | JSON.parse β the default for web APIs |
For a fresh web API, JSON is the sensible default β see What is JSON?. But you'll keep running into XML here:
| Where | Examples |
|---|---|
| π§ Sitemaps & feeds | sitemap.xml, RSS and Atom feeds, podcast feeds |
| π§Ό Enterprise services | SOAP web services, WSDL contracts, SAML single sign-on, banking and e-invoicing formats |
| π Documents | .docx, .xlsx and .pptx are ZIP files full of XML (Office Open XML); so are EPUB books. SVG images are XML too. |
| βοΈ Build & config | .NET .csproj and web.config, Maven pom.xml, Android layouts and manifests |
A sitemap is a good real example of a namespaced document:
<?xml version="1.0" encoding="UTF-8"?> <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"> <url> <loc>https://example.com/</loc> <lastmod>2026-03-14</lastmod> </url> </urlset>
Want to check a real site's sitemap β every URL, its status code and whether it resolves? The Sitemap Checker finds and parses it for you.
06 Β· Querying
π― XPath basics
XPath is a small path language for selecting nodes in an XML tree. It's built into most XML libraries, browser DevTools and XSLT. Against the catalogue.xml example above (ignoring its namespace for a moment):
| Expression | Selects |
|---|---|
| /catalogue/product | Every product that is a direct child of the root |
| //tag | Every tag element anywhere in the document |
| //product/@id | The id attribute of every product |
| //product[@currency='INR'] | Products whose currency attribute is INR |
| //product[price > 200]/name | Names of products priced above 200 |
| //product[1] | The first product within its parent β XPath counts from 1, not 0 |
| //name/text() | The text nodes inside each name |
| count(//product) | How many products there are |
Try it in the browser: open any XML file or HTML page, go to the DevTools console and run $x("//a") in Chrome, Edge or Firefox. It evaluates an XPath expression against the current document.
07 Β· In your code
π οΈ Parsing XML safely
Each snippet reads the sitemap above β namespace included, because that's what real documents look like.
using System.Xml.Linq; XNamespace sm = "http://www.sitemaps.org/schemas/sitemap/0.9"; var doc = XDocument.Parse(xml); var urls = doc.Descendants(sm + "url") .Select(u => (string?)u.Element(sm + "loc")) .ToList();
import xml.etree.ElementTree as ET
ns = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
root = ET.fromstring(xml_text)
urls = [loc.text for loc in root.findall("sm:url/sm:loc", ns)]
const doc = new DOMParser().parseFromString(text, "application/xml"); // DOMParser doesn't throw β it returns a document containing <parsererror> if (doc.querySelector("parsererror")) throw new Error("Not well-formed"); const urls = [...doc.getElementsByTagNameNS( "http://www.sitemaps.org/schemas/sitemap/0.9", "loc" )].map(el => el.textContent);
Untrusted XML needs care. DTDs can define entities that expand exponentially (the "billion laughs" attack) or pull in local files and URLs (XXE, XML External Entity injection). Keep DTD processing disabled unless you genuinely need it β modern .NET's XmlReader prohibits DTDs by default β and for Python the standard library docs recommend the defusedxml package when parsing data you don't control.
Got a wall of minified XML from an API log? The XML Formatter indents it, checks it's well-formed and lets you explore it as a tree. If the same data also comes as JSON, compare it in the JSON Formatter.