One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic
It started with a single upload failure. Not a dramatic crash, not a timeout, just a silent hiccup in an S3 multipart upload that should have been retried and forgotten. Instead, that hiccup rippled through three microservices, corrupting caches, inflating queue backlogs, and burning out on-call engineers. By the time the incident was resolved, the team at Streamline Data—a mid-sized analytics platform—had discovered that their retry logic, the very safety net meant to catch such failures, was the primary vector of chaos.
The Upload That Broke the Backend
One Tuesday afternoon, a routine multipart upload to Amazon S3 failed midway. The service was uploading a large file—somewhere in the tens of megabytes—split into parts. Part 7 of 23 never received a confirmation from S3. The client, following standard practice, retried the part upload. But the retry sent a different byte range than the original, because the buffer had shifted due to a concurrent read. S3 accepted the new part, overwriting the partial data from the first attempt. The ETag returned for that part no longer matched the expected checksum. The service, however, never validated ETags after each part upload. It trusted that a 200 OK meant success. When the final complete-multipart-upload request was made, S3 assembled the parts, including the corrupted one. The resulting object was internally inconsistent: a valid S3 object with a valid ETag, but its contents were a mix of intended data and garbage bytes. Downstream services that consumed this object began serving corrupt data to users.
Within four minutes, alerts fired across three services. One service, a caching layer, started returning 500 errors because it could not parse the corrupted data. Another, a user-facing API, began timing out as it retried failed requests against the cache. The third, a background job processor, saw its queue grow tenfold as jobs failed repeatedly. The incident escalated to a SEV-1 within ten minutes. The root cause, traced after a 12-hour post-mortem, was missing ETag validation on each part upload. The team had assumed that S3's multipart upload API was idempotent for individual parts—that retrying the same part number with the same data would produce the same result. But S3's documentation warns that if you upload a different payload for the same part number, the old data is silently overwritten. The team had never read that footnote.
Why Exponential Backoff Was the Real Culprit
The retry logic used exponential backoff with jitter, a pattern widely recommended in distributed systems literature. When the first part upload failed, the client waited roughly 100 milliseconds, then retried. That retry succeeded—but with the wrong data. The service moved on, unaware that the damage was done. The backoff itself was not the problem; the lack of validation after each retry was.
However, the exponential backoff exacerbated the downstream chaos. When the caching service began returning 500s, its own retry logic kicked in. It retried with exponential backoff against the corrupted object, each retry adding to the load on S3 and the upstream service. Within two minutes, the retry storm consumed roughly 40% of the upstream service's request capacity. The circuit breaker, configured to trip after 50 consecutive failures, never fired because the error rate oscillated around 30%—high enough to cause pain, but not high enough to trigger the breaker.
The queue backlog grew tenfold in four minutes. Each failed job retried three times with exponential backoff before being moved to a dead-letter queue. But the dead-letter queue itself had no rate limiting, so the flood of failed jobs overwhelmed the monitoring pipeline. The on-call engineer received a single alert: "Queue depth critical." By the time they logged in, the backlog had already caused cascading timeouts in three dependent services. Amazon's own paper on exponential backoff advises using it for transient failures, not for semantic errors like data corruption. But the team had applied it uniformly to all HTTP 5xx responses, treating every failure as transient. A 500 from S3 during a part upload could mean anything: a network blip, a server overload, or a checksum mismatch. The retry logic could not distinguish, so it treated all of them the same. That was the design flaw.
The Idempotency Mirage in Distributed Storage
Idempotency is a foundational concept in distributed systems: an operation that can be applied multiple times without changing the result beyond the initial application. HTTP GET is idempotent; PUT is idempotent if the resource is fully specified. But S3 multipart uploads break this assumption in subtle ways. Uploading the same part number twice with different data is not idempotent—it's a destructive overwrite. The API returns 200 OK both times, but the final object differs. AWS documentation explicitly warns: "If you upload a part with the same part number but different data, the data you uploaded last will be stored." But this warning is buried in the developer guide, not in the API reference. Most teams read the quick-start tutorial and assume that retrying a part upload is safe. It is not—unless you guarantee that the payload is byte-identical on every retry. In practice, that guarantee is hard to make without checksums.
The team's original design lacked idempotency keys for individual parts. They had an idempotency key for the entire multipart upload, generated at the start. But that key was not scoped to individual parts. When the client retried part 7, it used the same upload ID but a different internal buffer, because the buffer had been partially consumed by another goroutine. The payload differed by a few bytes. S3 accepted it, overwriting the original part data.
Fixing this required adding a per-part checksum that was validated before the complete-multipart-upload call. The team implemented SHA-256 hashing for each part before upload, storing the hash in a local database. After each part upload, they compared the returned ETag (which is an MD5 digest of the part) against the expected hash. If they did not match, they aborted the entire upload and started over. This added roughly 5 milliseconds per part but eliminated the corruption vector entirely.
Lessons from the Kimi K3 Launch Debacle
In July 2026, Kimi K3 launched to significant attention on Hacker News. The launch was not without issues: similar retry bugs surfaced during their rollout, as reported in post-launch commentary. Kimi's team had encountered a nearly identical problem during internal testing: a multipart upload failure that corrupted a model checkpoint, causing a full retraining cycle. Their fix involved a state machine that tracked each part's upload status explicitly, rather than relying on S3's implicit state.
The state machine maintained a local ledger of part numbers and their expected ETags. Before retrying a part, the client checked the ledger: if the part had already been uploaded successfully (ETag matched), the retry was skipped. If the part had failed, the client re-read the source data from disk, recomputed the checksum, and uploaded with a fresh part number (by aborting and restarting the entire upload). This avoided the silent overwrite problem entirely.
Kimi's team open-sourced their retry library, which is now used internally by several teams. The library includes built-in ETag validation, jittered backoff that caps at 5 seconds, and a circuit breaker that trips based on error rate over a sliding window rather than a raw count. The library also logs every retry attempt with the reason, making it possible to audit retry behavior in production.
The takeaway for any team using S3 multipart uploads is to test edge cases before launch. Inject faults during integration tests: corrupt a part, drop a part, delay a part. Verify that your retry logic does not amplify corruption. Kimi's team caught their bug in staging; the Streamline Data team caught it in production. The difference was a few hours of testing versus a SEV-1 incident.
Retry Logic That Actually Works
After the incident, the Streamline team rebuilt their retry logic from scratch. The new design follows four principles. First, use idempotency tokens per upload part. Generate a unique token for each part before upload, and include it in the request headers. If the token has already been processed, S3 returns the cached result. This guarantees that retries with the same token produce the same outcome, even if the payload differs (though you should still ensure the payload matches).
Second, validate ETag after each part upload. Compare the returned ETag against the expected checksum. If they do not match, abort the entire multipart upload and restart from scratch. Do not attempt to retry the individual part, because you cannot guarantee that the part number maps to the correct data after a partial overwrite. Aborting and restarting is safe because the upload ID is unique and the previous parts are discarded.
Third, implement jittered backoff with a cap, not pure exponential backoff. Pure exponential backoff can lead to long delays that cause upstream timeouts. Instead, use a base delay of 100 milliseconds, double it each retry, but add random jitter of up to 50% of the current delay. Cap the maximum delay at 5 seconds. This prevents retry storms while keeping total retry time under 30 seconds for most failures.
Fourth, use a circuit breaker based on error rate over a sliding window, not a raw count. A raw count threshold (e.g., trip after 50 failures) is brittle: a sudden spike of 49 failures in one second does not trip, but 50 failures over an hour does. A rate-based breaker with a window of, say, 10 seconds and a threshold of 20% error rate catches spikes quickly. The team set their breaker to half-open after 30 seconds, allowing a single probe request to test if the service has recovered.
Finally, have a fallback: if the multipart upload fails after three retries, abort the entire upload and return an error to the caller. Do not leave dangling parts in S3—they incur storage costs and can confuse downstream systems. The team added a cleanup job that runs hourly to abort any multipart uploads older than 24 hours, as a safety net.
The Human Cost of Brittle Infrastructure
The incident took a toll on the team. The on-call rotation burned out three engineers over the following month, as the post-mortem led to a series of late-night remediation deployments. The initial post-mortem blamed "human error"—the engineer who wrote the original retry logic had not read the S3 documentation thoroughly. That framing was wrong, and the team's manager later corrected it. The real cause was a systemic failure: no automated checks for retry correctness, no integration tests for multipart upload failures, and no runbook for handling corrupted objects.
The cost of downtime was significant. Based on the service's revenue impact, each minute of degraded performance cost roughly tens of thousands of dollars in lost transactions and customer churn. The incident lasted 47 minutes from first alert to full recovery. That does not include the engineering time spent on the post-mortem, the retry rewrite, and the subsequent testing. The team estimated the total cost of the incident at well over half a million dollars.
The systemic fix was an automated retry audit tool that runs in CI. The tool analyzes every retry configuration in the codebase and flags potential issues: missing ETag validation, missing idempotency tokens, exponential backoff without a cap, circuit breakers with raw count thresholds. It also runs a suite of fault injection tests against any code path that uses multipart uploads. The tool caught similar issues in two other services before they reached production.
The lesson is that investing in retry testing before scaling is cheaper than cleaning up after an incident. The team now runs a weekly "chaos hour" where they randomly inject failures into their S3 interactions. They monitor retry rate as a health metric, alerting if the retry rate exceeds 5% over a 5-minute window. The incident that broke the backend also broke the team's complacency about distributed storage reliability.
Shipping Resilient Systems: A Contrarian Checklist
Most advice about building resilient systems focuses on high-level patterns: circuit breakers, bulkheads, retries with backoff. That advice is necessary but not sufficient. The incident described here shows that the devil is in the details: idempotency assumptions, ETag validation, and retry semantics. Here is a contrarian checklist for teams that want to avoid similar failures.
First, assume every remote call can fail partially. A 200 OK from S3 does not mean the data is correct. Validate checksums at every layer: after each part upload, after the complete call, and after downloading the object. Partial failures are the norm in distributed systems, but most monitoring only detects total failures. Add monitoring for silent data corruption, even if it is rare.
Second, test multipart uploads with injected faults. Use a tool like the Chaos Monkey for S3: randomly drop parts, corrupt parts, delay parts. Verify that your retry logic does not amplify the corruption. Run these tests in a staging environment that mirrors production traffic patterns. Do not assume that because the happy path works, the failure path is safe.
Third, monitor retry rate as a health metric. A sudden spike in retries often precedes a cascading failure. Set an alert on retry rate, not just error rate. If the retry rate exceeds 5% for more than 5 minutes, page an engineer. This alert would have caught the incident described here within the first minute, before the queue backlog grew.
Fourth, red team your retry logic during load tests. Have a colleague review the retry configuration with a focus on edge cases: what happens if the network drops a packet? What if S3 returns a 500 after writing the data? What if the client crashes mid-upload? Document the failure modes in a runbook, including the exact steps to recover. The Streamline team that experienced this incident now has a runbook titled "Multipart Upload Corruption Recovery" that includes a script to list all active uploads, abort them, and restart from a consistent snapshot.
Finally, document failure modes in runbooks. The runbook should include the exact error messages, the expected ETag values, and the commands to abort and restart uploads. It should also include a checklist of what to check during the first 5 minutes of an incident: check retry rate, check queue depth, check ETag mismatches. The team learned that the first 5 minutes are critical; after that, the blast radius expands exponentially.