VDB
Sign up
MEDIUM6.5

PYSEC-2026-3831

eml_parser has a URL extraction bypass via HTML entities in URLs

Quick fix

PYSEC-2026-3831 — eml-parser: upgrade to the fixed version with the command below.

pip install --upgrade 'eml-parser>=3.0.2'

Details

## Summary

`eml_parser` performs certain validations on potential URL strings to discard bogus values. In versions prior to `3.0.2`, this validation was performed before unescaping any HTML entities that might occur in the string. This caused the library to wrongfully reject valid URLs that use HTML entities for the `:`, `/`, or `.` characters. These URLs would then not be included in the list of extracted URLs. Similarly, the host parts of such URLs would not be extracted.

For example, neither the URL `https://phishing.example.com` nor its host (`phishing.example.com`) would appear in the parsing result.

## Impact

`eml_parser` is used in email security gateways and SOC pipelines to extract URLs as IOCs. Those URLs are then checked against threat-intel feeds, URL reputation services, and sandboxes. A URL that is not extracted is never checked.

## Patches

Since version 3.0.2 the library unescapes all HTML entities in every URL before deciding to accept or reject it. A test was added to prevent regressions.

Are you affected?

Enter the version of the package you're using.

Affected packages

PyPI/eml-parser
Introduced in: 0Fixed in: 3.0.2
Fixpip install --upgrade 'eml-parser>=3.0.2'

References