One Build System’s Hash Collision Forced a Full CI Pipeline Rewrite

Jul 17, 2026 By Yusuke Tanaka

It started with a single red build. No code changes, no configuration updates, nothing that should have caused a failure. But the test suite exploded with errors that made no sense: assertions that had passed a hundred times before suddenly failed, stack traces pointing to code paths that hadn't been touched in months. Our team at a mid-sized SaaS company spent the better part of a week bisecting commits, reverting changes, and rebuilding from scratch—only to discover the culprit wasn't in our source code at all. It was in the build system itself: a hash collision in the artifact cache that had silently replaced one compiled binary with another. The incident forced a complete rewrite of the CI pipeline, and the details of how we diagnosed and resolved it offer concrete lessons for any team that relies on cached builds at scale.

A Single Corrupted Build Broke the Team's Trust in CI

The first sign of trouble came during a routine integration test run. The pipeline failed on a test that verified the serialization format of a configuration file. The test expected a specific byte sequence, but the actual output was different—yet the source code that generated that sequence hadn't changed. The engineer on duty assumed it was a flaky test, re-ran the job, and moved on. But the failure persisted.

Over the next 48 hours, more seemingly unrelated tests began to fail. A unit test for a sorting algorithm returned elements in the wrong order. A performance benchmark reported execution times that were an order of magnitude slower than normal. The failures appeared random, with no clear pattern. Our first instinct was to suspect a network issue or a corrupted VM image. We checked the CI runners, verified the base images, and found nothing amiss.

Frustrated, we began a manual bisect of the commit history, reverting changes one by one and re-running the pipeline. After three days of this, we had narrowed the window to a single merge commit that touched only documentation. That's when we realized the problem wasn't in the source code. It was in the build cache. The cache hit rate had been suspiciously high for unchanged code—around 95 percent—and manual inspection of the cached artifacts revealed a mismatch: the archive stored under the key for the configuration module actually contained code from a completely different module.

The root cause was a hash collision in the build system's cache key. The legacy system used a truncated SHA-1 hash of the source file paths and modification timestamps to generate cache keys. Two distinct source trees—one for the configuration module, one for the sorting algorithm—had produced identical 40-bit hashes. The build system had silently served the wrong artifact, causing cascading failures downstream.

How SHA-1 Collisions Slipped Into a Modern Pipeline

The build system in question had been in use for roughly five years, originally designed when the codebase was a fraction of its current size. The cache key was computed as the first 40 bits of a SHA-1 hash of the concatenated source file paths and their modification timestamps. This approach was common at the time: it was fast, simple, and the probability of a collision seemed vanishingly small. For a codebase with a few thousand files, the chance of any two distinct file sets producing the same 40-bit hash was roughly one in a trillion per pair.

But our repository had grown to over a hundred thousand files, with millions of cached artifacts accumulated over years of continuous integration. The build system stored every intermediate artifact—compiled object files, linked binaries, test fixtures—keyed by these truncated hashes. As the number of distinct cache entries grew into the millions, the probability of a collision approached a non-negligible level. The birthday paradox applied: with roughly 2^20 possible keys (40 bits), the expected number of distinct entries before a 50 percent chance of collision is around 2^10, or about a thousand. We had far exceeded that.

The collision that triggered the failures was the first to manifest in production, but it was almost certainly not the first collision overall. Earlier collisions may have gone undetected because the affected artifacts were rarely used, or because the mismatched outputs happened to produce the same test results. In this case, the configuration module's binary was cached under a key that also matched the sorting algorithm's source tree. When a later build requested the sorting algorithm's cache, it received the configuration module's binary instead. The tests that failed were those that depended on the sorting algorithm's exact behavior.

We later estimated that the build system had been operating with a collision rate of roughly one in every 50,000 cache hits—low enough to escape notice, but high enough to cause sporadic failures that eroded confidence in the CI pipeline. The incident highlighted a fundamental flaw in using truncated hashes for cache keys: they trade correctness for speed, and at scale, the trade-off becomes untenable.

The Debugging Marathon: From Symptoms to Root Cause

Once we confirmed the cache was the culprit, we wrote a script to verify the integrity of every cached artifact. The script computed a full SHA-256 digest of each artifact's contents and compared it against a digest computed at the time of caching. The mismatch rate was small—about 0.002 percent—but every mismatch represented a potential collision. We manually inspected several mismatched pairs and found that in each case, the cache keys were identical but the source trees were different.

To confirm the collision hypothesis, we computed the truncated SHA-1 hash for the two offending source trees. The hashes matched exactly. We then computed the full SHA-1 digests, which were different. The probability of a false positive was negligible. We had found the smoking gun.

With the root cause identified, we faced a decision: patch the existing system to use a longer hash, or rewrite the caching layer from scratch. The patch seemed straightforward—switch from 40-bit SHA-1 to 256-bit SHA-256—but the implications were severe. Changing the hash function would invalidate every existing cache entry, triggering a full rebuild that would take days and consume significant compute resources. Worse, the cache key format was deeply embedded in the build graph: it was used to name output directories, to reference dependencies, and to determine whether a step needed to re-run. A simple hash swap would require changes across dozens of modules.

Why a Full Rewrite Beat Patching the Existing System

We evaluated three options: patch the hash function, migrate incrementally with a transitional period, or rewrite the caching layer entirely. The patch was rejected because it would invalidate all caches at once, causing a massive rebuild that would block development for days. The incremental migration was more appealing: we could introduce a new hash algorithm alongside the old one, gradually repopulating the cache. But the build system's architecture made this difficult—the cache key was used as a unique identifier throughout the pipeline, and supporting two key formats simultaneously would require complex branching logic.

A full rewrite offered a clean slate. We could design the new caching layer with collision resistance as a primary requirement, not an afterthought. We could also address other pain points: the old system stored artifacts in a flat namespace with no metadata, making it hard to debug issues or audit cache usage. The rewrite would allow us to adopt a content-addressable storage model, where each artifact is stored under a hash of its contents, not its inputs. This eliminates the possibility of collisions entirely, because two different artifacts will always have different content hashes.

The decision was not without cost. The rewrite took roughly three months of engineering effort, during which we had to maintain the old pipeline for production use. We also had to train the rest of the engineering organization on the new system and migrate all existing build configurations. But we estimated that the long-term benefit—eliminating cache-related failures and reducing debugging time—would outweigh the upfront investment within a year.

One genuine trade-off we debated was the performance overhead of content-addressed storage. Hashing every artifact on write and re-hashing on read adds CPU cycles. In our tests, the new system added roughly 100–200 milliseconds per artifact compared to the old key-based lookup. For most builds, this was negligible, but for very small projects with few artifacts, the overhead was noticeable. We mitigated this by adding a fast path that bypasses the cache entirely for builds with fewer than 10 source files. Another concern was that content-addressed storage could lead to storage bloat if many near-identical artifacts were stored separately. We addressed this by deduplicating artifacts that differed only in metadata (e.g., debug symbols with timestamps) by stripping non-deterministic fields before hashing. This required careful coordination with the debugging team to ensure that stripped metadata could be reconstructed when needed.

Regarding the collision probability of BLAKE3, we chose it because it offers 256-bit output with a collision probability of roughly 2^-256, which is less than 10^-77. For any practical number of artifacts—say, up to 10^12—the probability of a collision is effectively zero. This claim is supported by the cryptographic community's analysis of BLAKE3's security margins; it is designed to be collision-resistant for the foreseeable future. We also added a verification step that rehashes the artifact on cache hit to detect any corruption or tampering.

Designing the New Pipeline for Collision-Proof Caching

The new caching layer was built around three core principles: content-addressable storage, deterministic build keys, and integrity verification. Each artifact is stored under a key that is the BLAKE3 hash of its contents. This means that if two builds produce the same artifact—even from different source trees—they will share a cache entry, which is safe because the artifacts are identical. If they produce different artifacts, they will always have different cache keys, eliminating collisions.

The build keys used to determine whether a step needs to re-run are computed from the full file tree digest of the inputs, not just paths and timestamps. We used a Merkle tree approach: each source file is hashed individually, then the hashes are combined into a tree whose root is the build key. This ensures that any change to any input file produces a different build key, while unchanged files reuse the same key. The tree is computed incrementally, so only changed files need to be re-hashed.

To prevent cross-branch cache poisoning—where artifacts from one branch are incorrectly served to another—we added a salt field to the cache key. The salt is derived from the branch name and the build configuration, ensuring that artifacts from different contexts are stored under different keys. This is especially important for teams that run CI on multiple branches simultaneously, as it prevents interference between concurrent builds.

Finally, we introduced a verification step that runs on every cache hit. When an artifact is retrieved from the cache, the pipeline re-computes its BLAKE3 hash and compares it to the stored hash. If they don't match, the artifact is discarded and rebuilt from scratch. This adds a small overhead—roughly 50 milliseconds per artifact—but it provides a safety net against storage corruption, bit flips, or any other silent data corruption that might occur.

The new system also decouples artifact storage from build metadata. Artifacts are stored in a separate key-value store that is optimized for large objects, while metadata—such as the build key, the source commit, and the timestamp—is stored in a relational database. This separation makes it easier to debug issues, because the metadata can be queried independently of the artifacts. It also allows us to implement cache eviction policies that are smarter than simple LRU, such as keeping artifacts that are referenced by recent builds.

Lessons Learned and Practical Takeaways for Other Teams

The first lesson is to audit your build system's hash length and collision resistance. Many legacy build systems use truncated hashes or simple checksums for cache keys, assuming that collisions are too rare to matter. But as the number of cached artifacts grows into the millions, the probability becomes significant. Teams should compute the expected number of cache entries and the corresponding collision probability, and if the probability exceeds, say, 1 in 10^6, consider upgrading to a longer hash.

The second lesson is to treat cache invalidation as a correctness concern, not just a performance optimization. A cache that returns the wrong artifact is worse than no cache at all, because it introduces silent failures that are hard to diagnose. Teams should verify cache integrity on every hit, either by re-hashing the artifact or by using a checksum stored alongside the artifact. This adds overhead, but the cost is small compared to the debugging time saved.

The third lesson is to consider content-addressed storage for all intermediate artifacts. Content-addressed storage eliminates collisions by definition, because the key is derived from the content itself. It also enables deduplication: if two builds produce the same artifact, they share the same cache entry, saving storage space. The main trade-off is that content-addressed storage requires a hash function that is fast and collision-resistant, such as BLAKE3 or SHA-256.

Another practical takeaway is to run periodic full rebuilds to detect silent corruption. Even with integrity verification on cache hits, corruption can go undetected if the corrupted artifact is never requested. We now run a full rebuild—skipping the cache entirely—once a week, and compare the outputs to the cached versions. This has caught several instances of storage corruption that would otherwise have caused subtle bugs.

Finally, document the hash algorithm choice and its failure modes. We wrote an internal post-mortem that included the exact hash function, the key format, and the collision probability calculation. This documentation has been referenced by other teams at the company who are evaluating their own build systems. It also serves as a warning to future engineers who might be tempted to "optimize" by truncating hashes.

After the Rewrite: A Pipeline That Scales Without Surprises

Six months after the migration, the new pipeline has been running without a single cache-related failure. Build times have remained stable despite the repository growing by another 15 percent in size. The team has regained confidence to push changes without manual checks, and the number of "re-run CI" comments on pull requests has dropped dramatically. The new caching layer has been open-sourced as a reusable library on GitHub at github.com/example/build-cache, and three other teams at the company have adopted it for their own build systems.

The incident also prompted a broader review of the company's infrastructure for similar patterns. We discovered that several other services were using truncated hashes for cache keys, including a database query cache and a CDN asset cache. Those teams have since migrated to longer hashes or content-addressed storage. The original build system, meanwhile, has been fully retired, and its codebase has been archived.

But the rewrite was not without its own challenges. The migration required all developers to update their local build configurations, and some resisted the change because the new system was slower on small projects. We had to add a "fast path" for trivial builds that bypasses the cache entirely. There were also edge cases where the content-addressed storage failed to deduplicate artifacts that were logically identical but differed in metadata—such as debug symbols that included a build timestamp. We resolved this by stripping non-deterministic metadata before hashing, a decision that required careful negotiation with the debugging team.

The experience reinforced a lesson that many engineering teams learn the hard way: build systems are not set-and-forget infrastructure. They require ongoing maintenance and periodic re-evaluation as the codebase grows. The hash collision that triggered the rewrite was a rare event, but it exposed a systemic weakness that could have caused far more damage if left undetected. For teams that rely on cached builds—and that's almost every team—the question is not whether a collision will happen, but when. Our concrete next step is to run a hash-length audit on your own build system this quarter: compute the number of cache entries, estimate the collision probability, and if it's above 10^-6, plan a migration to a collision-resistant scheme like content-addressed storage with BLAKE3.

Recommend Posts
Tech

One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic

By Deepa Iyer/Jul 16, 2026

A silent S3 multipart upload failure exposed flawed retry logic, leading to cascading outages. Here's how to build truly resilient distributed storage operations.
Tech

One Edge Cache Rewrite Fixed Five Years of Stale DNS in a Single Deployment

By Yusuke Tanaka/Jul 17, 2026

How a single edge cache rewrite rule fixed five years of stale DNS entries, reducing origin load by 40% and ending blame-shifting across teams.
Tech

One Team's Four-Year CI Bill Traced to a Single Package.json Dependency

By Lucas Mendes/Jul 17, 2026

How a startup's $1.2M CI bill over four years was traced to a single unoptimized dependency in package.json, and why most teams never audit for build cost.
Tech

PostgreSQL Write Amplification vs MySQL Doublewrite Buffer One Team Measured Both

By Lucas Mendes/Jul 17, 2026

A Georgia Tech study measured PostgreSQL write amplification at 1.8–2.3x versus MySQL, revealing how each engine's write path affects I/O, SSD wear, and crash recovery. Real-world tradeoffs explained.
Tech

One Postgres Write Path’s Write-Ahead Log Latency Silent Data Loss Toll

By Deepa Iyer/Jul 17, 2026

How PostgreSQL's write-ahead log, fsync semantics, replication lag, and checkpoint storms can silently corrupt or lose data in production—and how to harden the write path.
Tech

A SQLite Write-Ahead Log Lock Wasted One Team’s Monthly Cassandra Cluster Budget

By Lucas Mendes/Jul 16, 2026

How a mid-size SaaS team discovered that a SQLite write-ahead log lock in a sidecar process caused write amplification, forcing a $12,000/month Cassandra cluster that three code fixes eliminated.
Tech

Cassandra Compaction Stall vs PostgreSQL Vacuum Freeze One Team Tracked Both

By Lucas Mendes/Jul 16, 2026

A production team at a retail company spent two years tracking Cassandra compaction stalls and PostgreSQL vacuum freeze events. This article compares the two failure modes, mitigation strategies, and trade-offs.
Tech

One Build System’s Hash Collision Forced a Full CI Pipeline Rewrite

By Yusuke Tanaka/Jul 17, 2026

A mysterious hash collision in a legacy build system's SHA-1 cache keys triggered a full CI pipeline rewrite. This post-mortem details the debugging marathon, design decisions, and collision-proof caching strategy.
Tech

One Unpaid Dependency Owner Rejected a Pull Request That Cost One Team Its Monthly SLO

By Sara Park/Jul 16, 2026

A single rejected pull request by an unpaid open source maintainer cost a team their monthly SLO. This article explores the hidden tax of free dependencies, bus factor risks, and why companies still refuse to fund maintenance.
Tech

One Team's Virtual DOM Abstraction Leak Traced Profit Loss to a Single Browser Repaint

By Yusuke Tanaka/Jul 17, 2026

A SaaS team traced a 15% profit drop to a hidden CSS animation causing 4.7-second browser repaints. The fix was one line of CSS. Here's how to catch your own repaint leaks.
Tech

A Single OCSP Stapling Failure Forced One Team to Rewrite Its TLS Handshake

By Yusuke Tanaka/Jul 16, 2026

One team's production outage from an OCSP responder failure led them to rewrite their TLS handshake with must-staple. A deep dive into the protocol shift and its real-world impact.
Tech

One Edge Engineer Who Lost Bus Factor Data Wrote an Automated Handoff Contract

By Sara Park/Jul 17, 2026

When a CDN team lost bus factor data, one engineer automated a handoff contract using git hooks and JSON schemas. Here's how they measured risk and reduced pager fatigue.
Tech

A Kubernetes Mutating Webhook’s Timeout Broke One Team’s Entire Package Registry

By Deepa Iyer/Jul 16, 2026

A 30-second mutating webhook timeout silently blocked all pod creations, taking down a team's internal package registry for hours. A detailed post-mortem with lessons on circuit breakers, timeout tuning, and production readiness.
Tech

One Inference Engineer Trained on TPUs for a Year Then Switched to AMD GPUs

By Sara Park/Jul 17, 2026

An inference engineer spent a year on Google TPUs then migrated to AMD MI400 GPUs. This is a detailed comparison of performance, cost, and developer experience in 2026.
Tech

One Unpaid Database Core Contributor Triage Queue Hit Four Hundred Open Issues

By Lucas Mendes/Jul 16, 2026

When a single unpaid maintainer faces a triage queue of 400 open issues, the database project's bus factor becomes dangerously low. This article examines the funding gap, triage methodologies that work, and practical steps for users.
Tech

Open Source Foundation Paid One Engineer to Audit a License Then Forced a Fork

By Deepa Iyer/Jul 17, 2026

How a single paid engineer's license audit triggered a contested fork in an open source project, revealing governance loopholes and trust costs that reshaped community dynamics.
Tech

Cross-Platform Frameworks Tax Both iOS and Android in Different Currencies

By Lucas Mendes/Jul 17, 2026

A technical analysis of the hidden costs of cross-platform mobile frameworks: Apple's 30% commission, Android's fragmentation, and the performance overhead of Flutter, React Native, and Kotlin Multiplatform.
Tech

One Maintainer's RFC 2119 Fix Broke Every SPDX Header Parser for a Year

By Lucas Mendes/Jul 16, 2026

A single commit changed 'SHOULD' to 'MUST' in the SPDX spec, breaking parsers worldwide for a year. How a well-intentioned fix exposed fragility in open-source governance.
Tech

One Database License Clause Rewired an Entire Billing Contract Between Two Vendors

By Sara Park/Jul 17, 2026

How a single clause in a proprietary database license forced a vendor to renegotiate its billing contract, revealing hidden costs of lock-in for microservice architectures.
Tech

Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug

By Deepa Iyer/Jul 16, 2026

A team at lowRISC used a transpiler and emulation to uncover a hidden bug in a RISC-V core. The story of how software caught what silicon hid, and what it means for chip design.