Diagnosing XML Encoding Mismatches: BOMs, Declarations, and Bad Bytes
XML parse failures and mojibake usually mean the BOM, the encoding declaration, and the actual bytes disagree. A diagnostic guide: detection order, a cause table, ordered checks, and matched fixes.
25 Sept 2025, 11:36 UTC

An XML document that fails to parse with an "invalid byte sequence" error — or that parses but displays accented characters as garbage like é instead of é — almost always has a disagreement between three things: the byte-order mark (BOM), the encoding declared in the XML declaration, and the actual bytes on disk. The fix is usually a one-line change, but only after you identify which of the three is lying. This guide walks through how parsers decide the encoding, how to diagnose the mismatch, and which fix applies to each finding.
How a parser decides the encoding
XML 1.0 defines a strict detection order. The parser looks at the raw bytes before it reads any text:
- Byte-order mark. If the file starts with a BOM (
EF BB BFfor UTF-8,FF FEorFE FFfor UTF-16), that wins. - The encoding pseudo-attribute. If there is no BOM, the parser reads the XML declaration, e.g.
<?xml version="1.0" encoding="ISO-8859-1"?>. - Default. With neither a BOM nor a declaration, the parser assumes UTF-8 (or UTF-16 if the byte pattern implies it).
A declaration that contradicts the actual bytes is a fatal well-formedness error, not a warning. A parser in strict mode must stop; a parser in a lenient or recovery mode may continue and produce mojibake, which is arguably worse because the failure surfaces later and farther from the cause.
Recognizable conditions and likely causes
| Symptom | Likely cause |
|---|---|
| Parser aborts: "invalid byte sequence in UTF-8" at a specific offset | Declared UTF-8, but bytes are Windows-1252 or Latin-1 (smart quotes and accented letters are the usual offenders) |
| Parser aborts immediately, before any content | UTF-16/UTF-32 content with no BOM and no declaration, or whitespace/BOM before the declaration |
| "Content is not allowed in prolog" or similar | A BOM or stray byte before <?xml, often from concatenating files that each carry a BOM |
Parses, but text shows é, ’, etc. | Double-encoded text, or bytes decoded with the wrong charset somewhere upstream |
| Works locally, fails when served over HTTP | Transport charset (Content-Type header) disagrees with the document declaration |
Ordered checks
Work through these in order; each one narrows the cause.
1. Hex-dump the first bytes
Do not trust your editor's rendering — editors silently re-encode and hide the problem. Inspect the raw bytes. On Linux or macOS:
xxd -l 32 document.xmlNo special permissions needed; run it wherever the file lives. Look for:
ef bb bf— UTF-8 BOMff feorfe ff— UTF-16 BOM (little- or big-endian)3c 3f 78 6d 6c— the literal<?xml, meaning no BOM and the declaration is at byte zero- Anything else before
<?xml— a stray byte, whitespace, or a mid-stream BOM from concatenation
2. Compare the declaration to the observed bytes
Read the encoding pseudo-attribute and check it against the dump. If the declaration says encoding="UTF-8" but you see bytes like 93 or e9 standing alone (invalid as UTF-8 continuation patterns), the content is likely Windows-1252 or Latin-1.
3. Isolate with a minimal document
Create a test file containing only the declaration and one non-ASCII character, saved in the suspected encoding, and parse it with a strict parser. This separates encoding problems from schema or namespace issues. With Python (any environment, standard library):
python3 -c "import xml.dom.minidom; xml.dom.minidom.parse('document.xml')"Note the exact error and any reported byte offset. Parser messages vary by implementation and version — treat them as hints, not gospel, and confirm with the hex dump.
4. Check the transport layer
If the file parses locally but fails from a web service, compare the HTTP Content-Type charset with the document declaration. For text/xml, the MIME default is ASCII in some interpretations, which can override the declaration. application/xml defers to the document and is the safer media type. Also confirm no file transfer (FTP in text mode, for instance) re-encoded the bytes in transit.
5. Re-encode a copy and re-parse
As a confirming test, convert a copy to UTF-8 and parse again. With iconv:
iconv -f WINDOWS-1252 -t UTF-8 document.xml > document-utf8.xmlAdjust -f to the encoding you actually found in step 2. If the converted copy parses cleanly and the text content matches expectations, the diagnosis is confirmed. Never run this on the only copy of the file — work on a duplicate.
Fixes matched to findings
- Declaration says UTF-8, bytes are Latin-1/Windows-1252: Either re-encode the file to UTF-8, or correct the declaration to
encoding="windows-1252". Re-encoding is the better long-term fix; changing the declaration is the better immediate fix if the producer will keep emitting Latin-1. - UTF-16 content with no BOM and no declaration: Add the declaration (
encoding="UTF-16") or prepend the correct BOM. Better: re-encode to UTF-8 unless a consumer specifically requires UTF-16. - BOM conflicts with declaration: Remove the BOM or fix the declaration so they agree. Consistency is what matters.
- Bytes before the declaration: Strip everything before
<?xml. If the cause is concatenated files each carrying a BOM, fix the concatenation step to strip BOMs from all but the first input. - Double-encoded text: The corruption happened upstream. Find the conversion step that re-encoded already-UTF-8 text and remove it; repairing the file by hand only masks a bug that will recur.
- Transport mismatch: Align the
Content-Typecharset with the document, or switch toapplication/xmland let the declaration govern.
A decision worth standardizing
If you control the producer, settle this once: emit UTF-8 without a BOM, always include an explicit encoding="UTF-8" declaration, and validate encoding at ingestion boundaries. A cheap check at the point where XML enters your system — parse the first bytes, confirm declaration matches content — surfaces mismatches at the producer instead of deep inside a consumer pipeline, where the same bug presents as an inexplicable parse failure three services downstream.
When to escalate
Stop patching locally and escalate when:
- The mismatch originates in a third-party feed or a library serializer you cannot change. You need a normalization step at your boundary (re-encode on ingest), plus a conversation with the producer — local fixes will drift.
- The document requires XML 1.1 or an unusual encoding such as EBCDIC or UTF-32. XML 1.1 relaxes some character and encoding rules; confirm which version the document declares and whether your parser supports it before assuming a bug.
- The investigation reveals DTD or external entity handling in the same document. That is a security review, not an encoding fix.
Verifying the fix
After any change, re-run the hex dump to confirm the first bytes are what you intend, re-parse with a strict (non-recovering) parser, and spot-check the parsed text content of any non-ASCII strings against the source. If your pipeline has a lenient parser anywhere, test with a strict one — lenient modes mask exactly this class of bug and defer the failure to a later stage.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.