Back to Blog
high SEVERITY8 min read

sanitizeUnicodeInput(): Fullwidth U+ Bypasses Codepoint Validation

The `sanitizeUnicodeInput()` helper used by the project character-range settings screen rewrote `U+` prefixes to `0x` and called `parseInt()`, but never normalized its argument first. Compatibility-equivalent forms such as fullwidth `U+`, superscript digits, or mathematical alphanumerics never matched the `/U\+/gi` regex, fell through to the `else return inputString` branch, and were handed back to callers verbatim as "sanitized" values. The fix inserts a `String.prototype.normalize('NFKC')` pas

O
By Orbis AppSec
•Technically reviewed by Anupam Mediratta•Published October 2, 2026•Reviewed October 2, 2026

Answer Summary

The affected API is the first-party `sanitizeUnicodeInput(inputString)` helper that backs Unicode character-range entry in the project settings screen; there is no published package or version range. Before the fix, an attacker (or a crafted imported project file) could supply a compatibility-equivalent codepoint reference such as fullwidth `U+0041` or `U+⁰⁰⁴¹`, which never matched the `/U\+/gi` rewrite, produced `NaN` from `parseInt()`, and was returned unchanged — so a value that looks like a validated hex codepoint to every downstream caller is actually raw attacker text, and two visually identical range entries can compare unequal in `sortCharacterRanges()`. The fix normalizes the input with `inputString.normalize('NFKC')` before the regex and returns the normalized form from the fallback branch; it is a first-party commit with no version number. The fix cites CWE-178 (Improper Handling of Case Sensitivity); no CVE or GHSA has been assigned.

Vulnerability at a Glance

cweCWE-178
fixApply `normalize('NFKC')` before the regex and return the normalized string from the fallback branch
riskCompatibility-equivalent `U+` notation skips the regex rewrite and is returned unvalidated to callers that assume a normalized hex codepoint
languageJavaScript
root cause`sanitizeUnicodeInput()` ran `/U\+/gi` and `parseInt()` without NFKC normalization, and its `else` branch returned the raw input
vulnerabilityUnicode equivalence / homoglyph validation bypass in a sanitizer helper

Summary

A high-severity hardening fix landed in the project character-range settings code: the helper sanitizeUnicodeInput(inputString) was doing its validation before Unicode normalization, not after. The function rewrote U+ prefixes to 0x, ran parseInt(), and — when parsing failed — returned the caller's string untouched. Any compatibility-equivalent spelling of U+ therefore sailed straight through the "sanitizer" and came out the other side as raw, unvalidated text that downstream code treats as a vetted codepoint.

The fix is one added line and one changed return value. The reasoning behind it is worth more than the diff.

Introduction

The settings screen for a project's character ranges lets a user type a Unicode codepoint in the familiar U+0041 notation. Before that value reaches enableCharacterRange() or the sorting logic in sortCharacterRanges(), it passes through a normalizing helper whose job is to turn whatever the human typed into a canonical hex string:

function sanitizeUnicodeInput(inputString) {
    let sanString = inputString.replace(/U\+/gi, '0x');
    let sanInt = parseInt(sanString);

    if (!isNaN(sanInt)) return decToHex(Math.abs(sanInt));
    else return inputString;
}

Read that else branch carefully. It is the entire problem. When the input does not parse as a number, the function hands back inputString — the original, unmodified, attacker-controlled string — from a function named sanitizeUnicodeInput. Every caller downstream now holds a value it believes has been through sanitization.

The second problem is what makes the first one reachable. The /U\+/gi regex matches exactly two ASCII codepoints, U (or u) followed by +. Unicode offers many strings that a human reads as "U+" and that NFKC collapses to U+, but which this regex does not match at all. Feed one of those in and you land in the pass-through branch by construction.

This is the classic ordering mistake in Unicode handling: validate after normalizing, never before. If you normalize after your checks, your checks ran against a different string than the one your application ultimately uses.

Affected Versions

Affected not applicable (first-party code) — the sanitizeUnicodeInput() helper in the project settings module
Fixed in not applicable (first-party code) — fixed by the hardening commit titled "harden: secure settings_project.js (CWE-178)"
Ecosystem not applicable (first-party application code, JavaScript)
CVE / GHSA not assigned
CWE CWE-178 (Improper Handling of Case Sensitivity), as cited by the fix; the concrete defect is missing compatibility normalization in the same equivalence-handling family

There is no package version to upgrade here. If your codebase contains a helper with this shape — a regex rewrite plus parseInt(), with a raw pass-through fallback — you have the same pattern regardless of what you call it.

The Vulnerability Explained

The vulnerable code

let sanString = inputString.replace(/U\+/gi, '0x');
let sanInt = parseInt(sanString);

if (!isNaN(sanInt)) return decToHex(Math.abs(sanInt));
else return inputString;

Three distinct assumptions are baked into those four lines:

  1. That U+ only ever appears as ASCII U + ASCII +. The i flag covers ASCII case folding, so u+0041 works. It does nothing for codepoints that are compatibility equivalent to U.
  2. That parseInt() is a validator. It is a lenient scanner. It succeeds on prefixes and ignores trailing junk, and it fails — returning NaN — on digit forms that are not ASCII 0–9.
  3. That the NaN case is harmless. It is the opposite of harmless: it is the branch that returns unvalidated input.

Concrete bypasses against this exact code path

Fullwidth prefix. Type U+0041 — U+FF35 FULLWIDTH LATIN CAPITAL LETTER U followed by U+FF0B FULLWIDTH PLUS SIGN. The regex finds no U+ to replace, so sanString is unchanged. parseInt('U+0041') is NaN because the first character is not a digit. The else branch fires and U+0041 is returned as the "sanitized" value. Under NFKC, that same string normalizes to U+0041, which the function would have turned into the hex form of 65. The pre-fix and post-fix results for a canonically identical input differ completely.

Compatibility digits. Type U+⁰⁰⁴¹ using superscript digits (U+2070, U+2074, U+00B9…). The prefix rewrite does fire, producing 0x⁰⁰⁴¹. parseInt() then stops at the first non-ASCII-digit character after 0x and yields NaN. Pass-through again. NFKC maps superscript digits to ASCII digits, so after normalization this is the perfectly ordinary U+0041.

Mathematical alphanumerics. 𝐔+0041 (U+1D414 MATHEMATICAL BOLD CAPITAL U) renders as a bold U to a reviewer reading a diff or a log line, does not match /U\+/gi, and NFKC-folds to plain U.

Why the pass-through branch is the real sink

The dangerous outcome is not "the number is wrong." It is that two different kinds of value exit the same function through the same return type:

  • a strictly-formatted hex string produced by decToHex(Math.abs(sanInt)), or
  • arbitrary user text of arbitrary length and arbitrary codepoints.

Callers cannot tell which one they got. Any downstream code that stores the result as a character-range bound, renders it as a range label, writes it into persisted project settings, or compares it for equality is now operating on unvalidated input while believing the opposite.

Two practical consequences in this code path:

Range entry spoofing via canonical equivalence. sortCharacterRanges() and the range-enabling logic compare and order range identifiers. Pre-fix, U+0041 and U+0041 are two distinct strings that a user sees as the same range. An attacker-supplied project file can carry a homoglyph twin of a legitimate range that shadows it in a list, defeats a "does this range already exist?" check, or survives a deduplication pass that only compares raw strings.

Unbounded text where a 1–6 character hex string was expected. decToHex() output is short and fixed-shape. The pass-through branch imposes no length or character-class limit at all, so a crafted input is free to carry whatever the attacker wants into whichever renderer or serializer consumes the range name next.

Attack scenario

An attacker distributes a project file containing a character range whose bound is U+D800. On import, the settings code calls sanitizeUnicodeInput() on that bound. The regex misses, parseInt() returns NaN, and the raw string is accepted as a sanitized range bound. The range list now contains an entry that is indistinguishable on screen from the legitimate U+D800 entry but is a different key everywhere in code — so enabling, disabling, sorting, and overwriting operations target a different object than the one the user is looking at. The numeric guard that Math.abs() and decToHex() were supposed to enforce never executed.

It is worth being precise about severity: the fix's own description calls this defence-in-depth and notes it could not be demonstrated as remotely exploitable in this codebase, and the change was not covered by an automated check. The pattern is nonetheless a genuine correctness and trust-boundary defect, and the ordering mistake it encodes is the one that produces real bypasses in authentication and allow-list code.

The Fix

Before:

function sanitizeUnicodeInput(inputString) {
    let sanString = inputString.replace(/U\+/gi, '0x');
    let sanInt = parseInt(sanString);

    if (!isNaN(sanInt)) return decToHex(Math.abs(sanInt));
    else return inputString;
}

After:

function sanitizeUnicodeInput(inputString) {
    let normalizedString = inputString.normalize('NFKC');
    let sanString = normalizedString.replace(/U\+/gi, '0x');
    let sanInt = parseInt(sanString);

    if (!isNaN(sanInt)) return decToHex(Math.abs(sanInt));
    else return normalizedString;
}

Two changes, and both are load-bearing.

1. inputString.normalize('NFKC') runs first. NFKC is Normalization Form Compatibility Composition: it applies compatibility mappings and then recomposes. That is the form that folds U+FF35 to U, U+FF0B to +, fullwidth and superscript digits to ASCII digits, and mathematical alphanumerics to their basic Latin equivalents. Because the normalization now happens before the /U\+/gi replace and before parseInt(), all three of the bypasses above become ordinary inputs: U+0041, U+⁰⁰⁴¹, and 𝐔+0041 all become U+0041 and take the validated decToHex(Math.abs(sanInt)) path.

Note that NFC would not have been sufficient here. NFC handles canonical equivalence — precomposed versus decomposed accents — but leaves fullwidth and superscript forms alone. The compatibility mappings in NFKC are exactly the ones this input format needs.

2. The fallback returns normalizedString, not inputString. This is the half of the fix that is easy to skip and important to keep. Without it, the normalization would only affect the numeric path, and every non-numeric input would still escape in its original, unnormalized form — the function would be inconsistent about which representation it emits. After the change, there is exactly one invariant for every caller: whatever sanitizeUnicodeInput() returns is NFKC-normalized. That single guarantee is what makes downstream string comparison, deduplication, and sorting meaningful, and it is why the homoglyph-twin range entry can no longer shadow the real one.

What the fix deliberately does not do

The function is still not a strict validator, and the change does not pretend otherwise:

  • parseInt() still accepts a numeric prefix and discards the rest, so U+41junk normalizes, rewrites to 0x41junk, and yields 65.
  • The else branch still returns free-form text — now guaranteed normalized, but still unbounded in length and character class.
  • NFKC can lengthen a string (the ligature fi becomes two characters), so any length limit applied by a caller must be applied to the normalized output, not the raw input.

If a caller needs a hard guarantee, the right addition is an anchored test against the normalized string, for example /^(U\+)?[0-9A-Fa-f]{1,6}$/, with explicit rejection rather than pass-through on failure.

Key Takeaways

  • Normalize before you validate, always. sanitizeUnicodeInput() ran /U\+/gi and parseInt() against the raw argument; any check performed before normalize() was performed against a string the application does not ultimately use.
  • **A regex i flag is ASCII case folding, not Unicode equ

Prevention and further reading

Frequently Asked Questions

Why didn't the `/U\+/gi` regex already handle this, given it has the `i` flag?

The `i` flag only covers ASCII case folding, so it matches `u+` as well as `U+`. It does not match U+FF35 FULLWIDTH LATIN CAPITAL LETTER U or U+1D414 MATHEMATICAL BOLD CAPITAL U, which are distinct codepoints that NFKC — not case folding — maps back to `U`.

What did `sanitizeUnicodeInput()` return before the fix when `parseInt()` produced `NaN`?

It returned `inputString` completely unmodified, which is why the bypass mattered: the function's name promises sanitization, but the fallback path was a pure pass-through of attacker-controlled text.

Does adding `normalize('NFKC')` make `sanitizeUnicodeInput()` a complete validator?

No. `parseInt()` still accepts trailing garbage (`0x41abc` yields `65`) and the `else` branch still returns free-form text, now normalized. Callers that need a strict codepoint should additionally test the normalized string against an anchored pattern such as `/^(U\+)?[0-9A-F]{1,6}$/i`.

View the Security Fix

Check out the pull request that fixed this vulnerability

View PR #379

Related Articles

critical

pet-window.js Dynamic Code Evaluation: CWE-94 Hardening via Number

The pet-window module constructed dynamic JavaScript by embedding raw configuration values into code strings. An attacker with local access could inject arbitrary JavaScript by modifying stored configuration. The fix replaces string interpolation with explicit Number() coercion and NaN validation for all numeric configuration parameters.

high

brace-expansion Stack Exhaustion: CVE-2026-102276 Patched

A critical stack exhaustion vulnerability in brace-expansion allows attackers to crash Node.js applications by supplying specially crafted brace patterns that trigger unbounded recursion. The fix upgrades the library across all maintained version lines to enforce depth limits on recursive expansion. This vulnerability affects any service that expands user-controlled brace patterns without input validation.

critical

BFF Proxy QR Code Endpoint Prototype Pollution via Unvalidated JSON

A critical prototype pollution vulnerability in a backend-for-frontend (BFF) proxy endpoint allowed attackers to inject malicious properties into the JavaScript Object prototype by crafting JSON requests with forbidden keys. This could compromise application behavior across all objects. The fix adds explicit key validation to reject payloads containing `__proto__`, `constructor`, or `prototype`.

high

smol-toml 1.7.0 DoS: Malformed TOML Documents Crash Parser

A denial-of-service vulnerability in smol-toml 1.7.0 allows attackers to crash the parser by supplying malformed TOML documents. The vulnerability affects any application that parses untrusted TOML input. The fix, available in smol-toml 1.7.1, hardens input validation and error recovery.

high

JOSMFileHack TransformerFactory XXE: External DTD Processing Enabled

OSM2World's JOSMFileHack utility, which processed OpenStreetMap files generated by the JOSM editor, contained an insecure TransformerFactory configuration that permitted external DTD and stylesheet access. The vulnerability was resolved by completely removing the vulnerable code path rather than hardening it in place.

high

adm-zip 0.6.0 Preserves SUID Bits From ZIPs: CVE-2026-102282

The `adm-zip` dependency resolved to 0.6.0 in this project's dependency tree, a version affected by CVE-2026-102282: during extraction it applies the Unix permission bits stored in each ZIP entry's external file attributes verbatim, including the setuid (`04000`), setgid (`02000`), and sticky bits. An attacker who controls an archive passed to `extractAllTo()` or `extractEntryTo()` can therefore have the extractor create a setuid binary owned by whatever user the extraction process runs as. The