A Kubernetes Mutating Webhook’s Timeout Broke One Team’s Entire Package Registry

Jul 16, 2026 By Deepa Iyer

In late 2025, a mid-sized engineering team at a fintech company lost access to their internal package registry for roughly four to six hours. The root cause was not a network partition, a storage failure, or a malicious attack. It was a single Kubernetes mutating webhook configured with a 30-second timeout—and no fallback when that webhook failed silently. The incident offers a sharp lesson in how seemingly minor admission control settings can cascade into full-service outages.

A Single 30-Second Timeout Took Down the Registry

The team ran a single Kubernetes cluster hosting around 40 microservices. Their internal package registry, based on an open-source container image registry, served as the sole source for base images across all services. Every CI/CD pipeline pushed newly built images to this registry, and every pod pulled its base image from it. It was a textbook single point of failure, but one that had run without incident for over a year.

The mutating webhook in question was deployed to inject a sidecar container—a logging agent—into every pod. The webhook had a timeout of 30 seconds, meaning the API server would wait half a minute for the webhook to respond before failing the admission request. The failure policy was set to Fail, the default, which rejects the pod creation if the webhook does not respond in time.

On the day of the incident, the webhook pod crashed due to an out-of-memory (OOM) condition. Kubernetes restarted the pod, but the restart took longer than expected because the pod's container image was large and the node was under memory pressure. During this restart window, any new pod creation—including the registry's own pods triggered by a rolling update—hit the webhook timeout and was rejected.

No fallback or circuit breaker was in place. The team had not anticipated a scenario where the webhook itself would become the bottleneck. The registry's StatefulSet controller attempted to replace a failed pod, but the new pod could not start because the webhook timed out. The cluster was stuck in a loop: the registry needed a new pod, but the webhook prevented it from being admitted.

How Mutating Webhooks Work (and Break)

Kubernetes admission webhooks are HTTP callbacks that intercept requests to the API server before they are persisted. Mutating webhooks can modify the object—for example, adding a sidecar container, injecting environment variables, or setting resource limits. They are a powerful tool for enforcing cluster-wide policies without modifying application code.

When a pod creation request arrives, the API server sends it to the webhook endpoint. The webhook must respond with an AdmissionReview object within the configured timeout. If it does not, the API server applies the failure policy: Fail rejects the request, and Ignore allows it to proceed unmodified. The default is Fail, which is safe for security-critical hooks but dangerous for non-critical ones.

The webhook timeout is a single value, typically set between 10 and 30 seconds. There are no retries—just a hard failure. If the webhook is slow or unavailable, every admission request that hits it will fail until the webhook recovers or the timeout is changed. This makes mutating webhooks a potential single point of failure for the entire cluster's pod lifecycle.

In this incident, the 30-second timeout amplified the impact. Because the webhook pod was restarting, every request waited the full 30 seconds before failing. The API server's request queue backed up, and the registry controller's retries only added to the load. The team later measured that the webhook's p99 latency had been around 200 milliseconds under normal load—the 30-second timeout was an order of magnitude larger than necessary.

The Team's Architecture Before the Incident

The team's Kubernetes cluster was a shared resource for all services. The package registry ran as a StatefulSet with three replicas, each backed by persistent volumes. All services pulled their base images from this registry, and the CI/CD system pushed new images to it after every build. The registry was not mirrored externally; there was no fallback to Docker Hub or a cloud container registry.

The mutating webhook was deployed as a single-replica Deployment. It ran a custom Go binary that injected a sidecar container into every pod based on namespace labels. The webhook had no health checks configured—no readiness probe, no liveness probe. The team relied on Kubernetes to restart the pod if it crashed, but they had not considered the restart time under memory pressure.

All changes to the cluster were deployed via a CI/CD pipeline that used Helm charts. The pipeline itself depended on the package registry to pull base images for the Docker builds. When the registry went down, the CI/CD system could not push new images or deploy fixes. The team was effectively locked out of their own infrastructure.

Monitoring for the webhook was minimal. The team tracked general cluster metrics like CPU and memory usage, but they did not monitor webhook latency, error rates, or timeout counts. The first sign of trouble was a flood of alerts from services that could not pull images—by then, the cascade was already underway.

The Failure Cascade in Detail

The cascade began when the webhook pod's memory usage exceeded its limit. The pod was configured with a 128 MiB memory limit, but a spike in admission requests—triggered by a routine deployment of several services simultaneously—pushed it over. The OOM killer terminated the pod, and Kubernetes began restarting it.

During the restart, the node's kubelet pulled the webhook's container image, which was roughly 500 MiB. The node was already under memory pressure from other workloads, so the image pull took longer than usual—around 45 seconds. The webhook process started after the image was pulled, but its initialization required loading a large configuration file from a ConfigMap, adding another 10 seconds.

Meanwhile, the registry's rolling update had been triggered by a configuration change. The StatefulSet controller tried to create a new pod before terminating the old one. That new pod creation request hit the API server, which forwarded it to the webhook endpoint. The webhook was not yet listening, so the request timed out after 30 seconds. The API server applied the Fail policy and rejected the pod.

The controller retried with exponential backoff, but each retry also timed out. After several minutes, the controller gave up and marked the update as failed. The old registry pod continued running, but it was now operating with a degraded configuration. Eventually, the old pod also failed due to a separate issue—a disk space problem that had been masked by the rolling update. With no new pod able to start, the registry was effectively dead.

Why the Registry Was the Single Point of Failure

The registry's role as the sole image source made it critical. Every service pulled its base image from it at startup. The CI/CD system pushed images to it after every build. Without the registry, no new deployments could happen, and existing pods could not restart if they crashed. The team had considered mirroring to a cloud registry but had deferred it due to cost and complexity.

The webhook's failure policy of Fail was appropriate for security-critical hooks, but this webhook was not security-critical—it injected a logging sidecar. The team had not evaluated the blast radius of a webhook failure. They assumed the webhook would always be available, a common but dangerous assumption in distributed systems.

The incident lasted roughly four to six hours. The team eventually recovered by manually editing the webhook configuration to change the timeout to 5 seconds and the failure policy to Ignore. This allowed the registry pod to start without the sidecar. The sidecar was later injected via a separate, non-blocking mechanism. The team also added a readiness probe to the webhook to ensure it only received traffic when healthy.

The cost of the outage was significant. Services were unable to deploy fixes for other bugs during the window, and the team had to extend their sprint to compensate. The incident also eroded trust in the cluster's reliability, prompting a broader review of admission control configurations.

Lessons Applied: Circuit Breakers and Timeout Tuning

The team's post-mortem identified several concrete changes. First, they reduced the webhook timeout from 30 seconds to 8 seconds, with a failure policy of Ignore for non-critical hooks. The 8-second timeout was chosen based on observed p99 latency of 200 milliseconds, with a generous buffer for transient spikes. This change alone would have prevented the cascade, because the webhook's absence would have been treated as a soft failure.

Second, they implemented a circuit breaker pattern for webhook calls. They used a sidecar proxy that monitored webhook response times and error rates. If the error rate exceeded a threshold—say, 5% over a 1-minute window—the circuit breaker would trip and stop sending requests to the webhook, returning a success response to the API server instead. This added resilience against slow or failing webhooks without requiring manual intervention.

Third, they deployed the webhook as a separate Deployment with its own resource requests and limits, isolated from other workloads. They set memory requests to 256 MiB and limits to 512 MiB, with CPU requests of 100 millicores. They also added readiness and liveness probes that checked the webhook's HTTP endpoint. The readiness probe ensured that the API server only routed traffic to the webhook when it was ready to respond.

The team also started monitoring webhook latency and error rates as part of their standard observability stack. They set up alerts for p99 latency exceeding 1 second and for any timeout errors. These alerts would have caught the issue within seconds, rather than after the cascade had already started.

Production Readiness Checklist for Admission Webhooks

The incident is not unique. Many teams treat admission webhooks as infrastructure plumbing and neglect their failure modes. Based on this post-mortem and similar incidents at other companies, a production readiness checklist for admission webhooks should include the following items.

  • Test webhook failure scenarios in staging. Simulate webhook crashes, slow responses, and network partitions. Validate that the failure policy produces the expected behavior—whether that is rejecting requests or allowing them through.
  • Set resource requests and limits conservatively. Webhooks can experience sudden traffic spikes during deployments. Allocate enough headroom to handle peak load without OOM kills. Use vertical pod autoscaling if needed.
  • Use readiness probes to avoid routing to dead pods. Without a readiness probe, the API server may send requests to a webhook that is restarting, causing timeouts. A probe that checks the webhook's health endpoint ensures traffic only reaches healthy instances.
  • Document timeout and failure policy decisions. For each webhook, document why a particular timeout and failure policy were chosen. Include the expected latency profile and the blast radius of a failure. This documentation helps future engineers make informed changes.
  • Run chaos experiments to validate resiliency. Periodically inject failures into the webhook—kill the pod, slow its responses, or block its network. Verify that the cluster continues to operate as expected. Chaos engineering is the only way to uncover assumptions that break in production.

These steps are not exhaustive, but they address the most common failure modes. The team that suffered this outage now runs a quarterly "webhook failure drill" where they intentionally disable a non-critical webhook and observe the cluster's behavior. They have not had a similar incident since.

The broader lesson is that admission webhooks, like any critical infrastructure, deserve the same level of production readiness as the applications they serve. A 30-second timeout may seem generous, but in a distributed system, generosity without resilience is just a longer wait for failure. For a related exploration of how seemingly small configuration decisions cascade into outages, see a SQLite write-ahead log lock that consumed an entire monthly cluster budget. And for a look at how cache invariants can silently break, read about one CDN engineer's year-long unlearning of undocumented assumptions.

Trade-offs and Counter-Arguments: When Fail Is the Right Choice

Not every webhook should use Ignore as a failure policy. For security-critical webhooks—such as those enforcing pod security policies, validating image signatures, or preventing privilege escalation—a Fail policy is essential. If the webhook is unavailable, allowing pods to run without the security check could expose the cluster to vulnerabilities. In those cases, the trade-off is availability versus security. A well-designed system might use a redundant webhook deployment with multiple replicas and a load balancer to reduce the chance of unavailability.

Consider a webhook that validates that no container runs as root. If that webhook fails open, a malicious or misconfigured deployment could run privileged containers, potentially compromising the node. The blast radius of a security bypass is often larger than the blast radius of a deployment delay. Teams must evaluate the risk: for a logging sidecar injection, the cost of missing the sidecar is lost logs, which is inconvenient but not catastrophic. For a security enforcement hook, the cost of missing the check could be a breach.

The team in this incident could have kept Fail but added redundancy—running multiple webhook replicas across nodes, with a service in front that load-balances requests. They could also have set a shorter timeout, say 5 seconds, combined with a readiness probe that removes unhealthy pods from the service endpoint quickly. With two replicas, a single pod crash would not block all requests; the remaining replica would handle the load, albeit with increased latency. The team's single-replica deployment was a design flaw that amplified the impact of the OOM kill.

Another counter-argument is that circuit breakers add complexity. A sidecar proxy that intercepts webhook calls and implements circuit-breaking logic is another moving part that can itself fail. The team had to weigh the operational overhead of maintaining the proxy against the benefit of resilience. In their case, they decided the complexity was worth it because the webhook was critical to the deployment pipeline. For teams with simpler setups, a robust timeout and failure policy might suffice without a circuit breaker.

Ultimately, the choice depends on the webhook's function and the team's tolerance for risk. The key is to make that choice deliberately, document it, and test the failure modes. The team's original sin was not picking the wrong timeout—it was not thinking about failure at all.

Broader Implications for Infrastructure as Code

This incident also highlights a pattern in modern infrastructure: the increasing reliance on admission webhooks as part of "infrastructure as code" workflows. Tools like OPA Gatekeeper, Kyverno, and custom mutating webhooks are becoming standard in Kubernetes clusters. They enforce policies, inject sidecars, and modify resources automatically. But with this power comes risk: a misconfigured webhook can silently break the entire cluster's ability to run pods.

Teams should treat admission webhooks as critical infrastructure, subject to the same change management, monitoring, and testing as the applications they support. A webhook that fails open can allow insecure configurations; a webhook that fails closed can cause an outage. The industry is still learning how to balance these trade-offs. The incident described here is a cautionary tale, but it also provides a blueprint for building more resilient admission control systems.

As Kubernetes adoption grows, the number of webhooks in a typical cluster is likely to increase. Each webhook adds a potential failure point. Teams must invest in observability, chaos engineering, and thoughtful configuration to prevent a single timeout from taking down the entire cluster. The lessons from this post-mortem are not new—they echo decades of distributed systems wisdom—but they are easily forgotten in the rush to automate. A 30-second timeout may seem generous, but in a distributed system, generosity without resilience is just a longer wait for failure.

Recommend Posts
Tech

One Flaky S3 Multipart Upload Forced an Entire Microservice to Rewrite Its Retry Logic

By Deepa Iyer/Jul 16, 2026

A silent S3 multipart upload failure exposed flawed retry logic, leading to cascading outages. Here's how to build truly resilient distributed storage operations.
Tech

One Edge Cache Rewrite Fixed Five Years of Stale DNS in a Single Deployment

By Yusuke Tanaka/Jul 17, 2026

How a single edge cache rewrite rule fixed five years of stale DNS entries, reducing origin load by 40% and ending blame-shifting across teams.
Tech

One Team's Four-Year CI Bill Traced to a Single Package.json Dependency

By Lucas Mendes/Jul 17, 2026

How a startup's $1.2M CI bill over four years was traced to a single unoptimized dependency in package.json, and why most teams never audit for build cost.
Tech

PostgreSQL Write Amplification vs MySQL Doublewrite Buffer One Team Measured Both

By Lucas Mendes/Jul 17, 2026

A Georgia Tech study measured PostgreSQL write amplification at 1.8–2.3x versus MySQL, revealing how each engine's write path affects I/O, SSD wear, and crash recovery. Real-world tradeoffs explained.
Tech

One Postgres Write Path’s Write-Ahead Log Latency Silent Data Loss Toll

By Deepa Iyer/Jul 17, 2026

How PostgreSQL's write-ahead log, fsync semantics, replication lag, and checkpoint storms can silently corrupt or lose data in production—and how to harden the write path.
Tech

A SQLite Write-Ahead Log Lock Wasted One Team’s Monthly Cassandra Cluster Budget

By Lucas Mendes/Jul 16, 2026

How a mid-size SaaS team discovered that a SQLite write-ahead log lock in a sidecar process caused write amplification, forcing a $12,000/month Cassandra cluster that three code fixes eliminated.
Tech

Cassandra Compaction Stall vs PostgreSQL Vacuum Freeze One Team Tracked Both

By Lucas Mendes/Jul 16, 2026

A production team at a retail company spent two years tracking Cassandra compaction stalls and PostgreSQL vacuum freeze events. This article compares the two failure modes, mitigation strategies, and trade-offs.
Tech

One Build System’s Hash Collision Forced a Full CI Pipeline Rewrite

By Yusuke Tanaka/Jul 17, 2026

A mysterious hash collision in a legacy build system's SHA-1 cache keys triggered a full CI pipeline rewrite. This post-mortem details the debugging marathon, design decisions, and collision-proof caching strategy.
Tech

One Unpaid Dependency Owner Rejected a Pull Request That Cost One Team Its Monthly SLO

By Sara Park/Jul 16, 2026

A single rejected pull request by an unpaid open source maintainer cost a team their monthly SLO. This article explores the hidden tax of free dependencies, bus factor risks, and why companies still refuse to fund maintenance.
Tech

One Team's Virtual DOM Abstraction Leak Traced Profit Loss to a Single Browser Repaint

By Yusuke Tanaka/Jul 17, 2026

A SaaS team traced a 15% profit drop to a hidden CSS animation causing 4.7-second browser repaints. The fix was one line of CSS. Here's how to catch your own repaint leaks.
Tech

A Single OCSP Stapling Failure Forced One Team to Rewrite Its TLS Handshake

By Yusuke Tanaka/Jul 16, 2026

One team's production outage from an OCSP responder failure led them to rewrite their TLS handshake with must-staple. A deep dive into the protocol shift and its real-world impact.
Tech

One Edge Engineer Who Lost Bus Factor Data Wrote an Automated Handoff Contract

By Sara Park/Jul 17, 2026

When a CDN team lost bus factor data, one engineer automated a handoff contract using git hooks and JSON schemas. Here's how they measured risk and reduced pager fatigue.
Tech

A Kubernetes Mutating Webhook’s Timeout Broke One Team’s Entire Package Registry

By Deepa Iyer/Jul 16, 2026

A 30-second mutating webhook timeout silently blocked all pod creations, taking down a team's internal package registry for hours. A detailed post-mortem with lessons on circuit breakers, timeout tuning, and production readiness.
Tech

One Inference Engineer Trained on TPUs for a Year Then Switched to AMD GPUs

By Sara Park/Jul 17, 2026

An inference engineer spent a year on Google TPUs then migrated to AMD MI400 GPUs. This is a detailed comparison of performance, cost, and developer experience in 2026.
Tech

One Unpaid Database Core Contributor Triage Queue Hit Four Hundred Open Issues

By Lucas Mendes/Jul 16, 2026

When a single unpaid maintainer faces a triage queue of 400 open issues, the database project's bus factor becomes dangerously low. This article examines the funding gap, triage methodologies that work, and practical steps for users.
Tech

Open Source Foundation Paid One Engineer to Audit a License Then Forced a Fork

By Deepa Iyer/Jul 17, 2026

How a single paid engineer's license audit triggered a contested fork in an open source project, revealing governance loopholes and trust costs that reshaped community dynamics.
Tech

Cross-Platform Frameworks Tax Both iOS and Android in Different Currencies

By Lucas Mendes/Jul 17, 2026

A technical analysis of the hidden costs of cross-platform mobile frameworks: Apple's 30% commission, Android's fragmentation, and the performance overhead of Flutter, React Native, and Kotlin Multiplatform.
Tech

One Maintainer's RFC 2119 Fix Broke Every SPDX Header Parser for a Year

By Lucas Mendes/Jul 16, 2026

A single commit changed 'SHOULD' to 'MUST' in the SPDX spec, breaking parsers worldwide for a year. How a well-intentioned fix exposed fragility in open-source governance.
Tech

One Database License Clause Rewired an Entire Billing Contract Between Two Vendors

By Sara Park/Jul 17, 2026

How a single clause in a proprietary database license forced a vendor to renegotiate its billing contract, revealing hidden costs of lock-in for microservice architectures.
Tech

Transpiler Versus Transistor One Team's RISC-V Emulation Exposed a Silicon Bug

By Deepa Iyer/Jul 16, 2026

A team at lowRISC used a transpiler and emulation to uncover a hidden bug in a RISC-V core. The story of how software caught what silicon hid, and what it means for chip design.