Summary
A high-severity hardening fix landed in the project character-range settings code: the helper sanitizeUnicodeInput(inputString) was doing its validation before Unicode normalization, not after. The function rewrote U+ prefixes to 0x, ran parseInt(), and — when parsing failed — returned the caller's string untouched. Any compatibility-equivalent spelling of U+ therefore sailed straight through the "sanitizer" and came out the other side as raw, unvalidated text that downstream code treats as a vetted codepoint.
The fix is one added line and one changed return value. The reasoning behind it is worth more than the diff.
Introduction
The settings screen for a project's character ranges lets a user type a Unicode codepoint in the familiar U+0041 notation. Before that value reaches enableCharacterRange() or the sorting logic in sortCharacterRanges(), it passes through a normalizing helper whose job is to turn whatever the human typed into a canonical hex string:
function sanitizeUnicodeInput(inputString) {
let sanString = inputString.replace(/U\+/gi, '0x');
let sanInt = parseInt(sanString);
if (!isNaN(sanInt)) return decToHex(Math.abs(sanInt));
else return inputString;
}
Read that else branch carefully. It is the entire problem. When the input does not parse as a number, the function hands back inputString — the original, unmodified, attacker-controlled string — from a function named sanitizeUnicodeInput. Every caller downstream now holds a value it believes has been through sanitization.
The second problem is what makes the first one reachable. The /U\+/gi regex matches exactly two ASCII codepoints, U (or u) followed by +. Unicode offers many strings that a human reads as "U+" and that NFKC collapses to U+, but which this regex does not match at all. Feed one of those in and you land in the pass-through branch by construction.
This is the classic ordering mistake in Unicode handling: validate after normalizing, never before. If you normalize after your checks, your checks ran against a different string than the one your application ultimately uses.
Affected Versions
| Affected | not applicable (first-party code) — the sanitizeUnicodeInput() helper in the project settings module |
| Fixed in | not applicable (first-party code) — fixed by the hardening commit titled "harden: secure settings_project.js (CWE-178)" |
| Ecosystem | not applicable (first-party application code, JavaScript) |
| CVE / GHSA | not assigned |
| CWE | CWE-178 (Improper Handling of Case Sensitivity), as cited by the fix; the concrete defect is missing compatibility normalization in the same equivalence-handling family |
There is no package version to upgrade here. If your codebase contains a helper with this shape — a regex rewrite plus parseInt(), with a raw pass-through fallback — you have the same pattern regardless of what you call it.
The Vulnerability Explained
The vulnerable code
let sanString = inputString.replace(/U\+/gi, '0x');
let sanInt = parseInt(sanString);
if (!isNaN(sanInt)) return decToHex(Math.abs(sanInt));
else return inputString;
Three distinct assumptions are baked into those four lines:
- That
U+only ever appears as ASCIIU+ ASCII+. Theiflag covers ASCII case folding, sou+0041works. It does nothing for codepoints that are compatibility equivalent toU. - That
parseInt()is a validator. It is a lenient scanner. It succeeds on prefixes and ignores trailing junk, and it fails — returningNaN— on digit forms that are not ASCII0–9. - That the
NaNcase is harmless. It is the opposite of harmless: it is the branch that returns unvalidated input.
Concrete bypasses against this exact code path
Fullwidth prefix. Type U+0041 — U+FF35 FULLWIDTH LATIN CAPITAL LETTER U followed by U+FF0B FULLWIDTH PLUS SIGN. The regex finds no U+ to replace, so sanString is unchanged. parseInt('U+0041') is NaN because the first character is not a digit. The else branch fires and U+0041 is returned as the "sanitized" value. Under NFKC, that same string normalizes to U+0041, which the function would have turned into the hex form of 65. The pre-fix and post-fix results for a canonically identical input differ completely.
Compatibility digits. Type U+⁰⁰⁴¹ using superscript digits (U+2070, U+2074, U+00B9…). The prefix rewrite does fire, producing 0x⁰⁰⁴¹. parseInt() then stops at the first non-ASCII-digit character after 0x and yields NaN. Pass-through again. NFKC maps superscript digits to ASCII digits, so after normalization this is the perfectly ordinary U+0041.
Mathematical alphanumerics. 𝐔+0041 (U+1D414 MATHEMATICAL BOLD CAPITAL U) renders as a bold U to a reviewer reading a diff or a log line, does not match /U\+/gi, and NFKC-folds to plain U.
Why the pass-through branch is the real sink
The dangerous outcome is not "the number is wrong." It is that two different kinds of value exit the same function through the same return type:
- a strictly-formatted hex string produced by
decToHex(Math.abs(sanInt)), or - arbitrary user text of arbitrary length and arbitrary codepoints.
Callers cannot tell which one they got. Any downstream code that stores the result as a character-range bound, renders it as a range label, writes it into persisted project settings, or compares it for equality is now operating on unvalidated input while believing the opposite.
Two practical consequences in this code path:
Range entry spoofing via canonical equivalence. sortCharacterRanges() and the range-enabling logic compare and order range identifiers. Pre-fix, U+0041 and U+0041 are two distinct strings that a user sees as the same range. An attacker-supplied project file can carry a homoglyph twin of a legitimate range that shadows it in a list, defeats a "does this range already exist?" check, or survives a deduplication pass that only compares raw strings.
Unbounded text where a 1–6 character hex string was expected. decToHex() output is short and fixed-shape. The pass-through branch imposes no length or character-class limit at all, so a crafted input is free to carry whatever the attacker wants into whichever renderer or serializer consumes the range name next.
Attack scenario
An attacker distributes a project file containing a character range whose bound is U+D800. On import, the settings code calls sanitizeUnicodeInput() on that bound. The regex misses, parseInt() returns NaN, and the raw string is accepted as a sanitized range bound. The range list now contains an entry that is indistinguishable on screen from the legitimate U+D800 entry but is a different key everywhere in code — so enabling, disabling, sorting, and overwriting operations target a different object than the one the user is looking at. The numeric guard that Math.abs() and decToHex() were supposed to enforce never executed.
It is worth being precise about severity: the fix's own description calls this defence-in-depth and notes it could not be demonstrated as remotely exploitable in this codebase, and the change was not covered by an automated check. The pattern is nonetheless a genuine correctness and trust-boundary defect, and the ordering mistake it encodes is the one that produces real bypasses in authentication and allow-list code.
The Fix
Before:
function sanitizeUnicodeInput(inputString) {
let sanString = inputString.replace(/U\+/gi, '0x');
let sanInt = parseInt(sanString);
if (!isNaN(sanInt)) return decToHex(Math.abs(sanInt));
else return inputString;
}
After:
function sanitizeUnicodeInput(inputString) {
let normalizedString = inputString.normalize('NFKC');
let sanString = normalizedString.replace(/U\+/gi, '0x');
let sanInt = parseInt(sanString);
if (!isNaN(sanInt)) return decToHex(Math.abs(sanInt));
else return normalizedString;
}
Two changes, and both are load-bearing.
1. inputString.normalize('NFKC') runs first. NFKC is Normalization Form Compatibility Composition: it applies compatibility mappings and then recomposes. That is the form that folds U+FF35 to U, U+FF0B to +, fullwidth and superscript digits to ASCII digits, and mathematical alphanumerics to their basic Latin equivalents. Because the normalization now happens before the /U\+/gi replace and before parseInt(), all three of the bypasses above become ordinary inputs: U+0041, U+⁰⁰⁴¹, and 𝐔+0041 all become U+0041 and take the validated decToHex(Math.abs(sanInt)) path.
Note that NFC would not have been sufficient here. NFC handles canonical equivalence — precomposed versus decomposed accents — but leaves fullwidth and superscript forms alone. The compatibility mappings in NFKC are exactly the ones this input format needs.
2. The fallback returns normalizedString, not inputString. This is the half of the fix that is easy to skip and important to keep. Without it, the normalization would only affect the numeric path, and every non-numeric input would still escape in its original, unnormalized form — the function would be inconsistent about which representation it emits. After the change, there is exactly one invariant for every caller: whatever sanitizeUnicodeInput() returns is NFKC-normalized. That single guarantee is what makes downstream string comparison, deduplication, and sorting meaningful, and it is why the homoglyph-twin range entry can no longer shadow the real one.
What the fix deliberately does not do
The function is still not a strict validator, and the change does not pretend otherwise:
parseInt()still accepts a numeric prefix and discards the rest, soU+41junknormalizes, rewrites to0x41junk, and yields65.- The
elsebranch still returns free-form text — now guaranteed normalized, but still unbounded in length and character class. - NFKC can lengthen a string (the ligature
fibecomes two characters), so any length limit applied by a caller must be applied to the normalized output, not the raw input.
If a caller needs a hard guarantee, the right addition is an anchored test against the normalized string, for example /^(U\+)?[0-9A-Fa-f]{1,6}$/, with explicit rejection rather than pass-through on failure.
Key Takeaways
- Normalize before you validate, always.
sanitizeUnicodeInput()ran/U\+/giandparseInt()against the raw argument; any check performed beforenormalize()was performed against a string the application does not ultimately use. - **A regex
iflag is ASCII case folding, not Unicode equ