Back to Blog
high SEVERITY5 min read

clean_path() Path Traversal: rel_path Escapes USERDATA Dir

A path traversal flaw in the `clean_path()` helper let a user-controlled `rel_path` value escape the intended USERDATA directory and reach arbitrary files on disk. The function joined path segments without checking the final result, so `../` sequences in route parameters or query strings could be used to read or write files outside the sandboxed storage area. The fix normalizes the path and verifies it still resolves inside the USERDATA root before returning it, raising an error otherwise.

O
By Orbis AppSec
•Technically reviewed by Anupam Mediratta•Published October 6, 2026•Reviewed October 6, 2026

Answer Summary

The affected code is the `clean_path()` function in the application's first-party user-data storage utilities (no package version applies). An attacker who controls the `rel_path` argument — reachable via route parameters such as `:waktu` and `:nomor_surah`, or a `number` query parameter — could supply `../` sequences to escape the USERDATA directory and read or write arbitrary JSON files on the filesystem. The fix normalizes the path with `os.path.normpath()` and checks it against the USERDATA root using `os.path.commonpath()`, raising a `ValueError` if the path escapes the sandbox; no package version is associated with this first-party fix. This is tracked as CWE-22, Path Traversal.

Vulnerability at a Glance

cweCWE-22
fixNormalize the path and reject any result whose common path with USERDATA root does not match the root
riskArbitrary file read/write outside the intended USERDATA sandbox via crafted route/query parameters
languagePython
root cause`clean_path()` joined unsanitized path segments with `os.path.join()` without verifying the final path stayed under the USERDATA root
vulnerabilityPath Traversal

Introduction

This path traversal issue let an attacker turn a routine file-lookup helper into an arbitrary file reader. The affected function, clean_path(), builds a filesystem path by joining a relative path onto a fixed USERDATA root. The problem: the rel_path argument comes from user-controlled input — route parameters like :waktu and :nomor_surah, and a number query parameter — and was split and joined without ever checking whether the final path actually stayed inside USERDATA.

The vulnerable pattern is deceptively simple:

paths = rel_path.split('/')
for path in paths:
  cleaned = os.path.join(cleaned, path)

os.path.join() happily accepts .. as a path component. Splitting on / and joining each segment back on gives an attacker full control over how many directories the resulting path climbs before descending again. This is a classic case where "the code looks like it's building a safe path" and "the code is actually safe" are two different things — the function name promises cleaning, but nothing in the body validates the outcome.

Affected Versions

Affected not applicable (first-party code)
Fixed in not applicable (first-party code)
Ecosystem Python (pypi N/A — internal module)
CVE / GHSA not assigned
CWE CWE-22 (Path Traversal)

Because this is first-party application code rather than a published package, there's no version range to check — any deployment running the pre-fix clean_path() implementation is exposed.

The Vulnerability Explained

Here's the vulnerable function in full:

def clean_path(rel_path: str):
  """Cleans a relative path by splitting on forward slash and os.path.joining."""
  cleaned = USERDATA
  paths = rel_path.split('/')
  for path in paths:
    cleaned = os.path.join(cleaned, path)
  return cleaned

Notice what's missing: there is no check anywhere that cleaned ends up inside USERDATA. The function's docstring describes splitting and joining, not validating. Each segment of rel_path is joined in sequence, so a value like ../../etc/passwd splits into ['..', '..', 'etc', 'passwd'] and is joined one directory at a time — each .. walks back up a level in the real filesystem, exactly as it would from a shell.

Given that rel_path is assembled from route parameters such as :waktu and :nomor_surah plus a number query parameter, an attacker doesn't need any special access — just the ability to send an HTTP request with a crafted path segment. A request where :nomor_surah is set to something like ../../../../etc/passwd (URL-encoded as needed) would cause clean_path() to return a path far outside USERDATA, and whatever code calls it next — a JSON loader, in this case — would happily read (or, via save_userdata_json(), write) that file.

The real-world impact is significant: any JSON-readable file on the host becomes a potential read target, and because clean_path() is also used on the write path, an attacker could potentially overwrite files outside the sandbox too, depending on process permissions. For a service storing per-user data under USERDATA, this collapses the entire isolation model that directory is supposed to provide.

The Fix

The fix adds exactly the check that was missing: normalize the constructed path and confirm it's still rooted under USERDATA before returning it.

paths = os.path.normpath(rel_path).split(os.sep)
for path in paths:
  cleaned = os.path.join(cleaned, path)
cleaned = os.path.normpath(cleaned)
userdata_root = os.path.normpath(USERDATA)
if os.path.commonpath([cleaned, userdata_root]) != userdata_root:
  raise ValueError(f'Invalid path: "{rel_path}" is not under userdata.')

Three things changed, each closing a specific gap:

  1. rel_path is now run through os.path.normpath() before splitting, so sequences like ../.. collapse predictably rather than being treated as opaque segments.
  2. After building cleaned, it's normalized again, since os.path.join() can reintroduce .. segments that weren't resolved during the loop.
  3. The decisive change: os.path.commonpath([cleaned, userdata_root]) is compared against userdata_root. If the normalized target path doesn't share USERDATA as its common ancestor, the function raises ValueError instead of silently returning an escaped path.

This turns clean_path() from a function that describes cleaning into one that enforces a boundary — any caller, whether loading JSON for a :waktu/:nomor_surah lookup or saving user data, now gets either a path guaranteed to be under USERDATA or an exception.

Key Takeaways

  • os.path.join() does not sanitize .. segments — joining a path.split('/') list segment-by-segment is functionally identical to pasting the raw string into the filesystem call.
  • A function named clean_path() is not automatically safe just because it sounds like it validates input; the docstring here described splitting/joining, not boundary enforcement — read the implementation, not the name.
  • When a route parameter (:waktu, :nomor_surah) or query parameter (number) feeds into any path-construction helper, treat it as hostile input, not as a filename.
  • Validate after normalization: checking rel_path for .. before calling os.path.normpath() can be bypassed with encoding tricks; the fix instead checks the final, normalized result against the root using os.path.commonpath().
  • The same unvalidated helper served both reads and writes (save_userdata_json()), so a single missing check doubled as both an information-disclosure and a file-tampering bug.

How Orbis AppSec Detected This

  • Source: route parameters :waktu and :nomor_surah, and the number query parameter, flowing into the rel_path argument of clean_path()
  • Sink: the dynamic file path returned by clean_path() being used to load/save JSON data from disk
  • Missing control: no normalization or root-containment check on the joined path before it was used for file I/O
  • CWE: CWE-22 (Path Traversal)
  • Fix: normalize the input and resulting path, then raise ValueError unless os.path.commonpath() confirms the result is still under the USERDATA root

Orbis AppSec automatically detected this vulnerability and opened a pull request with the fix. Try Orbis AppSec on your repositories to find and fix issues like this automatically.

Conclusion

This finding is a reminder that path-building helpers are only as safe as their weakest validation step — and in clean_path(), there was no validation step at all. Splitting on / and joining each piece back together looks like sanitization but performs none; .. segments pass straight through to the filesystem. By normalizing both the input and the final joined path, and rejecting anything whose common ancestor isn't the USERDATA root, the fix closes the gap for every caller that depends on clean_path() — whether it's reading surah data by route parameter or writing user state back to disk.

Prevention and further reading

Frequently Asked Questions

Which route parameters feed into `clean_path()`'s `rel_path` argument?

The `:waktu` and `:nomor_surah` route parameters, along with a `number` query parameter, were interpolated into paths that eventually passed through `clean_path()`.

Does the fix change the signature or return value of `clean_path()`?

No — it still returns a joined path string, but now it calls `os.path.normpath()` on the input and raises `ValueError` if the resolved path isn't under the USERDATA root.

Would a path traversal payload have let an attacker write files, not just read them?

Yes — `clean_path()` is used by both the read and `save_userdata_json()` write paths, so an unvalidated `rel_path` could be exploited for both arbitrary file read and arbitrary file write within reach of the process's permissions.

View the Security Fix

Check out the pull request that fixed this vulnerability

View PR #786

Related Articles

critical

path.resolve() Path Traversal in Node CLI's Dynamic import()

A Node.js CLI script for validating expression definitions took a file path from `process.argv[2]`, resolved it with `path.resolve()`, and passed the result straight into a dynamic `import()` — with no check that the resolved path stayed inside the working directory. An attacker (or a malicious skill/plugin invocation) could supply traversal sequences to load and execute arbitrary `.js` files from anywhere on disk. The fix adds a boundary check with `path.relative()` and an extension allowlist b

high

bookDir() Path Traversal via Unsanitized bookId Parameter

The `bookDir()` function accepted unsanitized `bookId` values derived from user-created book titles, enabling path traversal attacks through `../` sequences. A fix was applied that validates the identifier using `path.basename()` and throws on mismatch, ensuring all resolved paths remain within `LIBRARY_DIR`.

high

Express `app.get('*')` Wildcard Handler Path Traversal in watch.js

A first-party Express server's wildcard route handler used `req.url.indexOf('font.woff2')` to gate access to a font file, allowing attackers to bypass the substring check with crafted paths. The fix replaces the catch-all handler with explicit route registration.

high

updateCardBg() Follows Unvalidated 302 Location Headers

A background-image updater fetched a configured image URL with manual redirect handling and then re-issued the request to whatever `Location` header came back, with no scheme or host checks. A redirect to `http://169.254.169.254/` or `http://127.0.0.1:<port>/` would have been followed with the original fetch options attached, and the response body written to disk as an image asset. The fix resolves the redirect target against `imgDownloadUrl` and rejects anything that is not HTTPS on the same ho

high

markitdown_bridge.py Path Traversal: Arbitrary File Read via sys.argv

The markitdown_bridge.py script, used by MDView for DOCX-to-Markdown conversion, accepted file paths directly from command-line arguments without validating they stayed within intended directories. An attacker could exploit this to read arbitrary files from the filesystem by passing path traversal sequences in the source_path parameter.

high

picomatch 2.3.1 ReDoS: Extglob Pattern Catastrophic Backtracking

picomatch versions below 2.3.2, 3.0.2, or 4.0.4 contain a Regular Expression Denial of Service vulnerability in extglob pattern parsing. An attacker can cause catastrophic backtracking with patterns containing nested alternations and quantifiers, freezing any Node.js process that evaluates untrusted glob expressions.