GHSA-g53g-w8rj-fmg7
xmldom PI grammar regex ReDoS: quadratic backtracking on unterminated processing instructions
Quick fix
GHSA-g53g-w8rj-fmg7 — @xmldom/xmldom: upgrade to the fixed version with the command below.
npm install @xmldom/xmldom@0.9.11Details
## Summary
`@xmldom/xmldom`'s processing-instruction (PI) grammar regex exhibits quadratic-time backtracking (ReDoS) when parsing an **unterminated** processing instruction. A single small XML document containing `<?` + a target + a long run of whitespace and no closing `?>` forces the regular expression engine into O(n²) work, stalling the Node.js event loop. The input is parsed with `DOMParser.parseFromString` under **default options**, so it is reachable from unauthenticated, network-delivered XML (SOAP/SAML, webhooks, uploads, XML APIs).
## Details
The PI production in `lib/grammar.js` compiles (flags `mu`) to:
``` ^<\?(NameChars)(?:[\x20\x09\x0D\x0A]+([Char]*?))?\?> ^^^ S+ greedy ^^^ Char*? lazy ```
- `lib/grammar.js` line 261: https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/grammar.js#L261
In the optional tail `(?:S+(Char*?))?`, both the greedy separator `S+` and the lazy data `Char*?` match XML whitespace. When the required trailing `?>` is absent, the engine must ultimately fail — but first it tries every partition of the whitespace run between `S+` and `Char*?`, which is O(n²) in the length of the trailing whitespace.
The regex is executed against the **entire remaining source string** in two places in `lib/sax.js`, so the whole whitespace tail is scanned:
- `parsePI` — https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L680-L691 - `parseProcessingInstruction` — https://github.com/xmldom/xmldom/blob/bb7a085dc5ba1eea3212388509b97bb4b4af32b9/lib/sax.js#L862-L879
## Affected Versions
Only the `0.9.x` line is affected. `lib/grammar.js` (and this PI regex) was introduced in commit `726b471` ("fix!: preserve DOCTYPE internal subset (#498)"), first released in **0.9.0-beta.9**, and is unchanged through **0.9.10**.
The `0.8.x` line (≤ 0.8.13) and the unscoped `xmldom` package (≤ 0.6.0) parse PIs via a different code path bounded by `indexOf('?>')` — they do **not** contain this regex and are **not affected** by this issue. (They were not separately tested for a *different* PI ReDoS; the scope here is the specific `grammar.js` regex.)
| Line | PI code path | Affected? | |---|---|---| | `0.9.x` (0.9.0-beta.9 … 0.9.10) | `grammar.js` `PI` regex over full remaining source | **Yes** | | `0.8.x` (≤ 0.8.13) | `parseInstruction`, bounded by `indexOf('?>')` | No | | unscoped `xmldom` (≤ 0.6.0) | older `indexOf('?>')`-bounded parsing | No |
## Proof of Concept
```js const { DOMParser } = require('@xmldom/xmldom'); const n = 32 * 1024; const payload = '<a><?p' + ' '.repeat(n); // unterminated PI, no `?>` console.time('parse'); new DOMParser().parseFromString(payload, 'text/xml'); console.timeEnd('parse'); ```
Measured (Node 18), trailing whitespace after `<?p`, no `?>` — time quadruples per doubling of input length (canonical O(n²)):
| Trailing whitespace | `g.PI.exec` | `parseFromString` | |---|---|---| | 2 KB | 4.4 ms | 5.1 ms | | 4 KB | 16.9 ms | 17.0 ms | | 8 KB | 111.4 ms | 66.3 ms | | 16 KB | 263.8 ms | 336.5 ms | | 32 KB | 1073.1 ms | — |
## Impact
Availability only: a single parse of a small crafted document blocks the Node.js event loop for the duration of the quadratic scan (≈1 s at 32 KB; multi-second with larger inputs). No memory blow-up, no data exposure, no integrity impact. Because XML is routinely accepted from untrusted sources and parsed with default options, one request can stall a server.
## Fix Applied
Fixed in `@xmldom/xmldom` **0.9.11** (`0.9.x`-only; the `0.8.x` LTS line and the unscoped `xmldom` package use a different, bounded PI code path and are not affected).
PR [#1039](https://github.com/xmldom/xmldom/pull/1039) inserts a fixed-width negative lookahead `(?!\s)` immediately after the greedy `S+`, so the separator can no longer hand whitespace back to the lazy data group:
``` - var PI = reg(/^<\?/, '(', Name, ')', regg(S, '(', Char, '*?)'), '?', /\?>/); + var PI = reg(/^<\?/, '(', Name, ')', regg(S, '(?!', _SChar, ')(', Char, '*?)'), '?', /\?>/); ```
The change is correct, minimal, and behavior-preserving: it produces identical `[target, data]` captures on all valid PIs tested (incl. whitespace-heavy, tab/newline, empty-data, and xml-decl cases) and removes the backtracking blow-up (linear, ~0.4 ms at 128 KB after the fix). The lookahead is fixed-width and cannot itself backtrack — a strict improvement with no new parsing risk.
## Severity note
The complexity is **quadratic**, not exponential, so a multi-second stall requires tens-to-hundreds of KB of input. `VA:H` reflects that xmldom applies **no input-size limit** and the path runs on default-options parsing, so a single unbounded parse can fully stall the event loop.
Are you affected?
Enter the version of the package you're using.
Affected packages
References
- https://github.com/xmldom/xmldom/security/advisories/GHSA-g53g-w8rj-fmg7[WEB]
- https://nvd.nist.gov/vuln/detail/CVE-2026-83606[ADVISORY]
- https://github.com/xmldom/xmldom/pull/1039[WEB]
- https://github.com/xmldom/xmldom/commit/73df6b8bdbd86f904b9e8c3ab9c49aa54ef2802e[WEB]
- https://github.com/xmldom/xmldom[PACKAGE]
- https://github.com/xmldom/xmldom/releases/tag/0.9.11[WEB]