Summary
A first-party maintenance script used hashlib.md5() inside a file_md5(path) helper to decide whether the per-vendor agent instruction documents in a repository were byte-equivalent. Because MD5 is collision-broken, two meaningfully different documents could be crafted to produce the same digest and pass the "these files are in sync" gate. The fix replaces the MD5 call with hashlib.sha256() over the same path.read_bytes() input.
Introduction
This is a weak-hash finding with an unusually modern blast radius. The affected code is a small repository hygiene script that keeps the AI-assistant instruction documents consistent — the canonical agent guidance document and its per-vendor siblings, one for each coding assistant the project supports. The script reads each document with file_text(path), digests each one with file_md5(path), and uses digest equality as its definition of "these files say the same thing." The entire vulnerable surface is one expression:
def file_md5(path: Path) -> str:
return hashlib.md5(path.read_bytes()).hexdigest()
The digest here is not a cache key or a log fingerprint. It is a trust decision: if two digests match, the tooling (and by extension a reviewer or CI job reading its output) concludes that the documents are identical and moves on. MD5 has been unsuitable for that role since 2004, and chosen-prefix collisions have been practical on commodity hardware for years. That turns a one-line convenience into a bypass for the only automated control that guarantees every AI assistant working in the repository receives the same instructions.
If you maintain tooling that compares files, deduplicates content, or asserts that a generated artifact matches its source, this is the pattern to look for: a hash used as an identity proof rather than as a hint.
Affected Versions
| Affected | not applicable (first-party code) — the file_md5() helper in the agent-docs maintenance script, prior to the fix commit |
| Fixed in | not applicable (first-party code) — the fix commit replacing hashlib.md5() with hashlib.sha256() |
| Ecosystem | N/A (first-party Python script, not a published package) |
| CVE / GHSA | not assigned |
| CWE | CWE-328: Use of Weak Hash (as referenced in the fix) |
The Vulnerability Explained
The code
The script's digest primitive, before the fix:
def file_md5(path: Path) -> str:
return hashlib.md5(path.read_bytes()).hexdigest()
path.read_bytes() pulls the entire document into memory and hashlib.md5() reduces it to a 128-bit value. The problematic part is not the I/O, the type hint, or the .hexdigest() call — it is the choice of md5 as the compression function behind a comparison that the rest of the tooling treats as authoritative.
The helper sits alongside file_text(path) and feeds describe(), which produces the human- and CI-readable report of the documents' state. Downstream, matching digests mean "equivalent", and differing digests mean "out of sync, go fix it."
Why MD5 breaks this specific check
MD5 collision resistance is gone, in two flavours that both matter here:
- Identical-prefix collisions are essentially free. Published collision block pairs can be pasted into two files to make them hash identically.
- Chosen-prefix collisions are practical — minutes to hours on a GPU budget any motivated attacker has. This is the dangerous one, because it lets an attacker pick meaningful content for both documents and then append a short binary-looking suffix to each to force the digests to match.
Markdown is a hospitable format for this. HTML comments, trailing reference-style link definitions, and fenced code blocks all let an attacker park collision bytes somewhere a skim-reading human will not look. Nothing in file_md5() constrains the file to printable characters or a maximum length.
Attack scenario against this code path
Consider a repository where the canonical agent instruction document is the reviewed source of truth and each vendor-specific document must be an exact copy. The script is the thing that enforces "exact copy."
- The attacker contributes a routine documentation change. The canonical document gets a benign new section, ending in an innocuous HTML comment holding collision block A.
- The vendor-specific copy for one assistant gets the same benign section — plus an extra instruction that the attacker wants only that assistant to see: permission to run a build step that reaches the network, a "load additional context from this URL" directive, or a relaxation of a rule like "never modify dependency manifests." It ends in an HTML comment holding collision block B.
- Because A and B are a chosen-prefix collision pair for the two prefixes,
file_md5()returns the same 32-hex-character string for both documents. describe()reports the documents as in sync. The CI gate is green. A reviewer who checked the digests, or who trusted the tool's output instead of diffing both files, sees no discrepancy.- The targeted assistant now operates under instructions that were never reviewed as part of the canonical document.
The real-world impact is a prompt-injection supply-chain hole hidden behind a passing integrity check. The agent instruction documents govern what an autonomous coding assistant is allowed to do inside the repository — which commands it may run, which files it may touch, which external context it may fetch. Divergence between them is exactly the condition the script exists to detect, and MD5 makes that detection forgeable. Secondary impact: the same digest feeds automated reports, so forensic reconstruction after an incident ("were these files ever different?") inherits the same false confidence.
Note also that the collision does not need to be forged to cause harm. MD5's 128-bit output with known collision structure means any accidental or tooling-induced collision — however unlikely in practice — silently suppresses a real divergence. There is no cost to being correct here.
Related pattern worth auditing: unsalted digests as identifiers
The same weak-hash rule class flags a second, structurally different misuse that commonly appears in the same codebases: deriving an owner or tenant identifier by hashing an email address and truncating the result — for example, taking the first 16 hex characters of a plain SHA-256 over the address. Switching MD5 to SHA-256 does not fix that pattern:
- Email addresses are low-entropy and enumerable. An attacker who obtains such an identifier can hash a wordlist of candidate addresses and recover the original, because no secret is involved.
- Truncating to 16 hex characters leaves 64 bits, which also shrinks the collision margin far below SHA-256's design strength and raises the chance of two distinct users mapping to one identifier.
The correct construction for that case is a keyed one — hmac.new(server_secret, email.lower().encode(), hashlib.sha256) — or a salted password-grade KDF if the value must resist offline attack. The lesson is that "use SHA-256" is the right fix for integrity comparison and the wrong fix for de-identification; the two uses need different primitives.
The Fix
The change is one line, and it is the whole fix for the integrity-comparison problem.
Before:
def file_md5(path: Path) -> str:
return hashlib.md5(path.read_bytes()).hexdigest()
After:
def file_md5(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()
Why this resolves the issue:
- Collision resistance is restored. SHA-256 has no known practical collision attack. Crafting two different agent instruction documents with the same SHA-256 digest is not something an attacker can do, so digest equality once again implies content equality for all practical purposes.
- The function contract is unchanged. Both branches return the hex string from
.hexdigest(), and every caller only compares digests to one another — nothing persists, transmits, or parses a fixed 32-character value. The output simply grows from 32 to 64 hex characters, which is why no call site needed to change and whyfile_text()anddescribe()are untouched. - No truncation was introduced. The full 64-character digest is returned. Shortening it for display would have reintroduced a collision margin problem; the fix deliberately does not do that.
- Performance is a non-issue. The inputs are small markdown documents read with
path.read_bytes(), and SHA-256 is hardware-accelerated on modern CPUs. There is no reason to trade correctness for speed at this size.
Recommended follow-up
The function is still named file_md5. That name is now wrong, and a wrong name is a liability: it invites a future maintainer to "fix the inconsistency" by putting hashlib.md5() back. Renaming it to file_digest (or file_sha256) and updating its single-expression body's call sites is a one-commit cleanup. Python 3.11+ also offers hashlib.file_digest(f, "sha256"), which streams the file instead of loading it whole — a reasonable modernization if the script ever digests larger artifacts.
Key Takeaways
- A digest used to answer "are these two files the same?" is a trust decision, and
hashlib.md5()cannot support one — chosen-prefix MD5 collisions are practical, so an attacker can make a tampered document and a reviewed document share a digest. - Markdown is a good carrier for collision blocks: HTML comments, trailing reference-link definitions, and fenced blocks hide arbitrary bytes from a skim-reading reviewer, and
file_md5()imposed no character or length constraints onpath.read_bytes(). - A sync check over AI agent instruction documents is a security control, not housekeeping — forging it lets one assistant receive unreviewed instructions about which commands it may run and which context it may fetch.
hashlib.sha256()is a drop-in forhashlib.md5()when callers only compare.hexdigest()values to each other; the only behavioural change is a 32-to-64-character output, so the fix needed no call-site edits.- "Use SHA-256" fixes integrity comparison but not identifier derivation: a truncated, unsalted SHA-256 of an email address is still reversible by enumeration and needs HMAC with a server-side secret instead.
- After changing the algorithm, rename the helper — leaving a
file_md5()that returns SHA-256 is an invitation for someone to regress it back.
How Orbis AppSec Detected This
- Source: the raw file bytes returned by
path.read_bytes()for each agent instruction document, including documents contributed through pull requests. - Sink:
hashlib.md5()invoked inside thefile_md5(path)helper, whose.hexdigest()output is consumed as a content-identity token by the document-comparison anddescribe()reporting logic. - Missing control: no collision-resistant digest algorithm. The equality check relied entirely on a 128-bit MD5 value with publicly practical chosen-prefix collisions, and no secondary byte-for-byte comparison or signature backed it up.
- CWE: CWE-328 — Use of Weak Hash.
- Fix:
hashlib.md5(path.read_bytes())was replaced withhashlib.sha256(path.read_bytes()), restoring collision resistance for the digest used to prove the documents are identical.
Orbis AppSec automatically detected this vulnerability and opened a pull request with the fix. Try Orbis AppSec on your repositories to find and fix issues like this automatically.
Conclusion
The vulnerable code was a single expression in a helper nobody thinks about, and that is precisely why it survived: file_md5() looked like plumbing. But its output was the sole automated evidence that every AI coding assistant working in the repository reads the same instructions, and MD5's broken collision resistance made that evidence forgeable by anyone who could land a markdown file. Swapping in hashlib.sha256() over the same path.read_bytes() input restores the guarantee at zero cost to the API, since callers only ever compare digests to one another.
Two things are worth carrying forward. First, audit the helper's name as well as its body — a file_md5() that computes SHA-256 is a regression waiting to happen. Second, resist generalizing "replace MD5 with SHA-256" into a rule: for integrity comparison it is correct, but for turning an email address into an identifier, a plain or truncated SHA-256 is just as reversible as MD5 and needs a keyed HMAC instead. The right primitive depends on what the digest is being asked to prove.