VDB
KO
MEDIUM

GHSA-3rcm-vjrc-p45j

JustHTML has a Sanitizer Bypass (in Markdown)

Quick fix

GHSA-3rcm-vjrc-p45j — justhtml: upgrade to the fixed version with the command below.

pip install --upgrade 'justhtml>=1.12.0'

Details

## Summary

`to_markdown()` does not sufficiently escape text content that looks like HTML. As a result, untrusted input that is safe in `to_html()` can become raw HTML in Markdown output.

This is not specific to tokenizer raw-text states like `<title>`, `<noscript>`, or `<plaintext>`, although those states can trigger the behavior. The root cause is broader: Markdown text serialization leaves angle brackets unescaped in text nodes.

## Details

When converting a parsed document to Markdown, text nodes are escaped for a small set of Markdown metacharacters, but HTML-significant characters such as `<` and `>` are preserved. That means content parsed as text, including entity-decoded text or text produced by RCDATA/RAWTEXT-style parsing, can be emitted into Markdown as raw HTML.

Examples of affected input include:

- Text produced from entity-decoded input such as `&lt;script&gt;...&lt;/script&gt;` - Text inside elements like `<title>`, `<textarea>`, `<noscript>` (when parsed as raw text), and `<plaintext>`

This is distinct from actual `<script>` or `<style>` elements in the DOM. Those are already dropped by default in `to_markdown()` unless `html_passthrough=True`.

## Proof of Concept

### General case

```python from justhtml import JustHTML

doc = JustHTML("<p>&lt;img src=x onerror=alert(1)&gt;</p>", fragment=True)

print(doc.to_html()) print() print(doc.to_markdown())

Are you affected?

Enter the version of the package you're using.

Affected packages

PyPI / justhtml
Introduced in: 0 Fixed in: 1.12.0
Fix pip install --upgrade 'justhtml>=1.12.0'

References