—
PYSEC-2026-1455
Arbitrary HTML present after sanitization because of unicode normalization
Quick fix
PYSEC-2026-1455 — html-sanitizer: upgrade to the fixed version with the command below.
pip install --upgrade 'html-sanitizer>=2.4.2'Details
### Impact
If using `keep_typographic_whitespace=False` (which is the default), the sanitizer normalizes unicode to the NFKC form at the end. Some unicode characters normalize to chevrons; this allows specially crafted HTML to escape sanitization.
### Patches
The problem has been fixed in 2.4.2.
### Workarounds
Set `keep_typographic_whitespace=True` explicitly, or normalize to NFKC yourself earlier.
Are you affected?
Enter the version of the package you're using.
Affected packages
References
- https://github.com/matthiask/html-sanitizer/security/advisories/GHSA-wvhx-q427-fgh3[WEB]
- https://github.com/matthiask/html-sanitizer/commit/48db42fc5143d0140c32d929c46b802f96913550[WEB]
- https://github.com/matthiask/html-sanitizer[PACKAGE]
- https://pypi.org/project/html-sanitizer[PACKAGE]
- https://github.com/advisories/GHSA-wvhx-q427-fgh3[ADVISORY]
- https://nvd.nist.gov/vuln/detail/CVE-2024-34078[ADVISORY]