Back to Blog
high SEVERITY9 min read

file_md5() MD5 Digest Lets Agent Docs Collide (CWE-328)

A first-party maintenance script used `hashlib.md5()` inside a `file_md5(path)` helper to decide whether the per-vendor agent instruction documents in a repository were byte-equivalent. Because MD5 is collision-broken, two meaningfully different documents could be crafted to produce the same digest and pass the "these files are in sync" gate. The fix replaces the MD5 call with `hashlib.sha256()` over the same `path.read_bytes()` input.

O
By Orbis AppSec
•Technically reviewed by Anupam Mediratta•Published October 8, 2026•Reviewed October 8, 2026

Answer Summary

The affected code is first-party, not a published package: a `file_md5(path)` helper in a repository maintenance script that digests the per-vendor AI agent instruction documents and compares the results for equality. An attacker who can land a file in the repository could exploit MD5's practical chosen-prefix collisions to make a tampered vendor-specific instruction document hash identically to the reviewed canonical document, so the sync check reports everything is consistent while one assistant is served different instructions. The fix swaps `hashlib.md5(path.read_bytes())` for `hashlib.sha256(path.read_bytes())`; there is no package version, so "fixed in" is the fix commit itself. The weakness class is CWE-328 (Use of Weak Hash), as referenced in the fix PR.

Vulnerability at a Glance

cweCWE-328
fixCompute the digest with `hashlib.sha256()` over the same byte input
riskTwo different agent instruction documents can produce the same digest, so a tampered file passes the repository's "docs are in sync" verification
languagePython
root cause`file_md5()` computed `hashlib.md5(path.read_bytes()).hexdigest()` and the result was used as a content-identity token
vulnerabilityUse of a collision-broken hash function (MD5) for content-equality and integrity checks

Summary

A first-party maintenance script used hashlib.md5() inside a file_md5(path) helper to decide whether the per-vendor agent instruction documents in a repository were byte-equivalent. Because MD5 is collision-broken, two meaningfully different documents could be crafted to produce the same digest and pass the "these files are in sync" gate. The fix replaces the MD5 call with hashlib.sha256() over the same path.read_bytes() input.

Introduction

This is a weak-hash finding with an unusually modern blast radius. The affected code is a small repository hygiene script that keeps the AI-assistant instruction documents consistent — the canonical agent guidance document and its per-vendor siblings, one for each coding assistant the project supports. The script reads each document with file_text(path), digests each one with file_md5(path), and uses digest equality as its definition of "these files say the same thing." The entire vulnerable surface is one expression:

def file_md5(path: Path) -> str:
    return hashlib.md5(path.read_bytes()).hexdigest()

The digest here is not a cache key or a log fingerprint. It is a trust decision: if two digests match, the tooling (and by extension a reviewer or CI job reading its output) concludes that the documents are identical and moves on. MD5 has been unsuitable for that role since 2004, and chosen-prefix collisions have been practical on commodity hardware for years. That turns a one-line convenience into a bypass for the only automated control that guarantees every AI assistant working in the repository receives the same instructions.

If you maintain tooling that compares files, deduplicates content, or asserts that a generated artifact matches its source, this is the pattern to look for: a hash used as an identity proof rather than as a hint.

Affected Versions

Affected not applicable (first-party code) — the file_md5() helper in the agent-docs maintenance script, prior to the fix commit
Fixed in not applicable (first-party code) — the fix commit replacing hashlib.md5() with hashlib.sha256()
Ecosystem N/A (first-party Python script, not a published package)
CVE / GHSA not assigned
CWE CWE-328: Use of Weak Hash (as referenced in the fix)

The Vulnerability Explained

The code

The script's digest primitive, before the fix:

def file_md5(path: Path) -> str:
    return hashlib.md5(path.read_bytes()).hexdigest()

path.read_bytes() pulls the entire document into memory and hashlib.md5() reduces it to a 128-bit value. The problematic part is not the I/O, the type hint, or the .hexdigest() call — it is the choice of md5 as the compression function behind a comparison that the rest of the tooling treats as authoritative.

The helper sits alongside file_text(path) and feeds describe(), which produces the human- and CI-readable report of the documents' state. Downstream, matching digests mean "equivalent", and differing digests mean "out of sync, go fix it."

Why MD5 breaks this specific check

MD5 collision resistance is gone, in two flavours that both matter here:

  1. Identical-prefix collisions are essentially free. Published collision block pairs can be pasted into two files to make them hash identically.
  2. Chosen-prefix collisions are practical — minutes to hours on a GPU budget any motivated attacker has. This is the dangerous one, because it lets an attacker pick meaningful content for both documents and then append a short binary-looking suffix to each to force the digests to match.

Markdown is a hospitable format for this. HTML comments, trailing reference-style link definitions, and fenced code blocks all let an attacker park collision bytes somewhere a skim-reading human will not look. Nothing in file_md5() constrains the file to printable characters or a maximum length.

Attack scenario against this code path

Consider a repository where the canonical agent instruction document is the reviewed source of truth and each vendor-specific document must be an exact copy. The script is the thing that enforces "exact copy."

  1. The attacker contributes a routine documentation change. The canonical document gets a benign new section, ending in an innocuous HTML comment holding collision block A.
  2. The vendor-specific copy for one assistant gets the same benign section — plus an extra instruction that the attacker wants only that assistant to see: permission to run a build step that reaches the network, a "load additional context from this URL" directive, or a relaxation of a rule like "never modify dependency manifests." It ends in an HTML comment holding collision block B.
  3. Because A and B are a chosen-prefix collision pair for the two prefixes, file_md5() returns the same 32-hex-character string for both documents.
  4. describe() reports the documents as in sync. The CI gate is green. A reviewer who checked the digests, or who trusted the tool's output instead of diffing both files, sees no discrepancy.
  5. The targeted assistant now operates under instructions that were never reviewed as part of the canonical document.

The real-world impact is a prompt-injection supply-chain hole hidden behind a passing integrity check. The agent instruction documents govern what an autonomous coding assistant is allowed to do inside the repository — which commands it may run, which files it may touch, which external context it may fetch. Divergence between them is exactly the condition the script exists to detect, and MD5 makes that detection forgeable. Secondary impact: the same digest feeds automated reports, so forensic reconstruction after an incident ("were these files ever different?") inherits the same false confidence.

Note also that the collision does not need to be forged to cause harm. MD5's 128-bit output with known collision structure means any accidental or tooling-induced collision — however unlikely in practice — silently suppresses a real divergence. There is no cost to being correct here.

Related pattern worth auditing: unsalted digests as identifiers

The same weak-hash rule class flags a second, structurally different misuse that commonly appears in the same codebases: deriving an owner or tenant identifier by hashing an email address and truncating the result — for example, taking the first 16 hex characters of a plain SHA-256 over the address. Switching MD5 to SHA-256 does not fix that pattern:

  • Email addresses are low-entropy and enumerable. An attacker who obtains such an identifier can hash a wordlist of candidate addresses and recover the original, because no secret is involved.
  • Truncating to 16 hex characters leaves 64 bits, which also shrinks the collision margin far below SHA-256's design strength and raises the chance of two distinct users mapping to one identifier.

The correct construction for that case is a keyed one — hmac.new(server_secret, email.lower().encode(), hashlib.sha256) — or a salted password-grade KDF if the value must resist offline attack. The lesson is that "use SHA-256" is the right fix for integrity comparison and the wrong fix for de-identification; the two uses need different primitives.

The Fix

The change is one line, and it is the whole fix for the integrity-comparison problem.

Before:

def file_md5(path: Path) -> str:
    return hashlib.md5(path.read_bytes()).hexdigest()

After:

def file_md5(path: Path) -> str:
    return hashlib.sha256(path.read_bytes()).hexdigest()

Why this resolves the issue:

  • Collision resistance is restored. SHA-256 has no known practical collision attack. Crafting two different agent instruction documents with the same SHA-256 digest is not something an attacker can do, so digest equality once again implies content equality for all practical purposes.
  • The function contract is unchanged. Both branches return the hex string from .hexdigest(), and every caller only compares digests to one another — nothing persists, transmits, or parses a fixed 32-character value. The output simply grows from 32 to 64 hex characters, which is why no call site needed to change and why file_text() and describe() are untouched.
  • No truncation was introduced. The full 64-character digest is returned. Shortening it for display would have reintroduced a collision margin problem; the fix deliberately does not do that.
  • Performance is a non-issue. The inputs are small markdown documents read with path.read_bytes(), and SHA-256 is hardware-accelerated on modern CPUs. There is no reason to trade correctness for speed at this size.

Recommended follow-up

The function is still named file_md5. That name is now wrong, and a wrong name is a liability: it invites a future maintainer to "fix the inconsistency" by putting hashlib.md5() back. Renaming it to file_digest (or file_sha256) and updating its single-expression body's call sites is a one-commit cleanup. Python 3.11+ also offers hashlib.file_digest(f, "sha256"), which streams the file instead of loading it whole — a reasonable modernization if the script ever digests larger artifacts.

Key Takeaways

  • A digest used to answer "are these two files the same?" is a trust decision, and hashlib.md5() cannot support one — chosen-prefix MD5 collisions are practical, so an attacker can make a tampered document and a reviewed document share a digest.
  • Markdown is a good carrier for collision blocks: HTML comments, trailing reference-link definitions, and fenced blocks hide arbitrary bytes from a skim-reading reviewer, and file_md5() imposed no character or length constraints on path.read_bytes().
  • A sync check over AI agent instruction documents is a security control, not housekeeping — forging it lets one assistant receive unreviewed instructions about which commands it may run and which context it may fetch.
  • hashlib.sha256() is a drop-in for hashlib.md5() when callers only compare .hexdigest() values to each other; the only behavioural change is a 32-to-64-character output, so the fix needed no call-site edits.
  • "Use SHA-256" fixes integrity comparison but not identifier derivation: a truncated, unsalted SHA-256 of an email address is still reversible by enumeration and needs HMAC with a server-side secret instead.
  • After changing the algorithm, rename the helper — leaving a file_md5() that returns SHA-256 is an invitation for someone to regress it back.

How Orbis AppSec Detected This

  • Source: the raw file bytes returned by path.read_bytes() for each agent instruction document, including documents contributed through pull requests.
  • Sink: hashlib.md5() invoked inside the file_md5(path) helper, whose .hexdigest() output is consumed as a content-identity token by the document-comparison and describe() reporting logic.
  • Missing control: no collision-resistant digest algorithm. The equality check relied entirely on a 128-bit MD5 value with publicly practical chosen-prefix collisions, and no secondary byte-for-byte comparison or signature backed it up.
  • CWE: CWE-328 — Use of Weak Hash.
  • Fix: hashlib.md5(path.read_bytes()) was replaced with hashlib.sha256(path.read_bytes()), restoring collision resistance for the digest used to prove the documents are identical.

Orbis AppSec automatically detected this vulnerability and opened a pull request with the fix. Try Orbis AppSec on your repositories to find and fix issues like this automatically.

Conclusion

The vulnerable code was a single expression in a helper nobody thinks about, and that is precisely why it survived: file_md5() looked like plumbing. But its output was the sole automated evidence that every AI coding assistant working in the repository reads the same instructions, and MD5's broken collision resistance made that evidence forgeable by anyone who could land a markdown file. Swapping in hashlib.sha256() over the same path.read_bytes() input restores the guarantee at zero cost to the API, since callers only ever compare digests to one another.

Two things are worth carrying forward. First, audit the helper's name as well as its body — a file_md5() that computes SHA-256 is a regression waiting to happen. Second, resist generalizing "replace MD5 with SHA-256" into a rule: for integrity comparison it is correct, but for turning an email address into an identifier, a plain or truncated SHA-256 is just as reversible as MD5 and needs a keyed HMAC instead. The right primitive depends on what the digest is being asked to prove.

Prevention and further reading

Frequently Asked Questions

Does swapping `hashlib.md5()` for `hashlib.sha256()` in `file_md5()` change the function's return type or callers?

No. Both return a lowercase hex string from `.hexdigest()`, and every caller only compares digests to each other, so no call site needed changes. The only behavioural difference is that the string is now 64 hex characters instead of 32.

Should the helper still be called `file_md5` after the fix?

No — the name is now actively misleading and a future maintainer could "restore" MD5 to match it. Renaming it to `file_digest` or `file_sha256` is the recommended follow-up, and it is cheap because the function has a single-expression body.

Why is SHA-256 the right answer for the document digests but not for deriving an owner identifier from an email address?

Document comparison only needs collision resistance, which SHA-256 provides. Deriving an identifier from a low-entropy, enumerable input like an email address needs a keyed construction such as HMAC-SHA-256 with a server-side secret, because a plain SHA-256 — especially truncated to 16 hex characters — is trivially reversed by enumerating candidate addresses.

View the Security Fix

Check out the pull request that fixed this vulnerability

View PR #4

Related Articles

high

SQLite Auth Database Plaintext Storage in InitAuthDB()

The `InitAuthDB()` function opened SQLite databases without restricting filesystem permissions, leaving bcrypt password hashes, session tokens, and user settings exposed to any local account with read access. The fix explicitly sets `os.Chmod(filepath, 0600)` immediately after database creation, ensuring only the owner can access authentication secrets.

critical

Hybridauth Telegram OAuth Timing Attack in `authenticateCheckError()`

The Telegram OAuth provider in Hybridauth relied on `strcmp()` to validate HMAC-SHA256 signatures, exposing a timing side-channel that could allow attackers to forge authentication tokens character by character. The fix replaces this with PHP's timing-safe `hash_equals()` function, eliminating the information leak.

high

undici 7.29.0 TLS Bypass: BalancedPool Drops Connect Options

A critical flaw in undici's BalancedPool implementation silently discards TLS certificate validation settings when routing requests through certain connection paths. Attackers on the network could intercept HTTPS traffic without triggering validation errors. The fix preserves connect options throughout the connection lifecycle.

critical

Model Fetching Without Integrity Verification in ONNX Loading

A machine learning application was fetching ONNX model files and chunks over the network without verifying their integrity, creating an opening for model poisoning attacks. The fix adds cryptographic integrity verification at the point where downloaded chunks are reassembled and cached, ensuring models have not been modified in transit or at rest.

high

nanoid 3.3.11 Integer Overflow: Predictable ID Generation

An integer overflow in nanoid 3.3.11's internal randomness generation causes the library to fall back to predictable ID sequences, undermining the cryptographic guarantees of its supposedly unguessable identifiers. The fix upgrades the dependency tree to patched versions 3.3.12 or 5.1.11.

high

MapManager.get() Race Condition Duplicates API Requests

The MapManager's `get(mapUid, cache)` method used a check-then-act pattern that permitted multiple concurrent requests to pass the cache miss check simultaneously, triggering redundant API calls and risking cache corruption. The fix introduces a `_pending` promise map to deduplicate in-flight fetches for identical map UIDs.