VDB
Sign up
HIGH

GHSA-6mj3-qw4j-hgrw

xmldom: HTML raw-text closing-tag case mismatch causes output amplification

Quick fix

GHSA-6mj3-qw4j-hgrw — @xmldom/xmldom: upgrade to the fixed version with the command below.

npm install @xmldom/xmldom@0.9.12

Details

## Summary

In HTML mode (`text/html`), a raw-text element (`script`, `style`, `textarea`, `title`) whose closing tag differs in case from its opening tag (e.g. `</ScRiPt>` for `<script>`) is mishandled by the parser, producing quadratic (O(n²)) output growth — a small crafted document parses and serializes into output orders of magnitude larger, exhausting CPU and memory. A modest input of tens of KB can therefore cause a denial of service in any service that parses untrusted HTML with xmldom. Only HTML mode is affected.

## Details

The parser calls `parseHtmlSpecialContent` for each raw-text element in HTML mode, matched via `isHTMLRawTextElement` / `isHTMLEscapableRawTextElement` (so all four types — `script`, `style`, `textarea`, `title` — are in scope). It searches for the element's closing tag with `source.indexOf('</' + tagName + '>', elStartEnd)`, a byte-for-byte case-sensitive match. A mixed-case closing tag never matches, so the search returns `-1`, and the following `source.substring(elStartEnd + 1, -1)` extracts text backwards from the start of the document instead of the element's content. The function then returns `-1` to the parse loop, which cannot advance normally and falls back to character-by-character reprocessing. Every raw-text element re-captures all source text preceding it, so output grows as O(n²) in the number of such elements.

### Root Cause

1. **Case-sensitive close-tag search** (`lib/sax.js:549`): `source.indexOf('</' + tagName + '>', elStartEnd)` does not fold case, contrary to the WHATWG HTML RAWTEXT end-tag-name rule. 2. **Unguarded `-1`** (`lib/sax.js:550`): `source.substring(elStartEnd + 1, elEndStart)` runs even when `elEndStart === -1`, extracting text backwards from position 0. 3. **Unstable progression** (`lib/sax.js:556`): the function returns `elEndStart` (`-1`), driving repeated character-by-character fallback in the parse loop.

## Affected Versions

Only the `0.9.x` line is affected — the amplification was introduced in `0.9.0-beta.1` when `parseHtmlSpecialContent` was refactored, and remains through `0.9.11`. The `0.8.x` line is **not** affected: its older `parseHtmlSpecialContent` does not amplify, despite sharing the same case-sensitive `indexOf`.

## Proof of Concept

```js const { DOMParser, XMLSerializer } = require('@xmldom/xmldom');

const n = 1000; const payload = '<html><body>' + '<script>x</ScRiPt>'.repeat(n) + '</body></html>'; const doc = new DOMParser().parseFromString(payload, 'text/html'); const out = new XMLSerializer().serializeToString(doc); console.log(payload.length, out.length, (out.length / payload.length).toFixed(1) + 'x'); // 18026 9037063 501.3x — an 18 KB input yields ~9 MB of output ```

Output size grows quadratically with the number of case-mismatched raw-text elements:

``` Repeats | Input len | Output len | Ratio 1 | 44 | 109 | 2.5x 100 | 1826 | 93763 | 51.3x 500 | 9026 | 2268563 | 251.3x 1000 | 18026 | 9037063 | 501.3x 2000 | 36026 | 36074063 | 1001.3x ```

Proof of Concept from @KarimTantawey (tested with `script`); the same amplification occurs for `style`, `textarea`, and `title`.

## Impact

Small attacker payloads can force disproportionate CPU and memory usage in services that parse and serialize untrusted HTML via xmldom. The quadratic growth means a modest-sized input (tens of kilobytes) can produce output in the tens or hundreds of megabytes, potentially exhausting memory or causing timeouts.

The attack only requires HTML mode (`text/html` MIME type) and mixed-case closing tags for any of the four raw-text element types. No special configuration or error handler setup is needed.

## Severity note

The CVSS 4.0 vector scores availability only (`VA:H`, with `VC:N/VI:N`): the flaw neither discloses nor corrupts data, but a small untrusted HTML input (tens of KB) can force output and memory in the tens to hundreds of MB, enough to exhaust a service's heap or stall its event loop. It is reachable with no authentication, configuration, or error-handler setup — only that the application parses untrusted `text/html` and serializes the result.

## Fix Applied

The raw-text closing tag is now matched case-insensitively in HTML raw-text mode (per the WHATWG HTML [RAWTEXT end-tag rule](https://html.spec.whatwg.org/multipage/parsing.html#rawtext-end-tag-name-state)), and a missing closing tag is handled explicitly, removing the quadratic output amplification. Output for well-formed input is unchanged. Non-breaking; 0.9.x-only.

Are you affected?

Enter the version of the package you're using.

Affected packages

npm/@xmldom/xmldom
Introduced in: 0.9.0-beta.1Fixed in: 0.9.12
Fixnpm install @xmldom/xmldom@0.9.12

References