A Single OCSP Stapling Failure Forced One Team to Rewrite Its TLS Handshake
On a midsummer Tuesday, no public outage was declared. Yet our team lost 12% of production traffic in 90 minutes. The root cause was a single misconfigured load-balancer rule at our CDN's OCSP responder. It dropped roughly 30% of queries silently, and clients that implemented hard-fail on revocation check failure refused to connect. That afternoon, we began rewriting our TLS handshake.
The Day the OCSP Responder Went Silent
The incident started at 14:03 UTC. Our monitoring dashboard showed a sudden spike in TLS handshake failures, concentrated on Android devices and some desktop browsers. Initial suspicion fell on a recent certificate rotation, but the certificates were valid. The issue was that the Online Certificate Status Protocol (OCSP) responder—the service that tells a client whether a certificate has been revoked—was returning errors for a subset of requests.
The CDN's OCSP infrastructure was behind a load balancer. A new routing rule had been pushed the night before, intended to improve latency for a different service. Instead, it caused roughly 30% of OCSP queries to hit a backend that responded with a 503 status. Clients that treated OCSP errors as hard failures—meaning they refused to proceed with the connection—simply dropped the request. Those that used soft-fail (allowing the connection to proceed) saw no issue.
Our team ran a service that depended on timely certificate validation. We had chosen to enforce strict revocation checking because of compliance requirements. The result: 12% of our production traffic was rejected in the span of 90 minutes. We rolled back the load-balancer rule within an hour, but the damage to our uptime SLA was done.
The postmortem revealed a deeper problem: our TLS handshake relied on an external OCSP responder that we did not control. Even with redundant endpoints, a single misconfiguration could take us down. We needed a design where the server itself provided proof of certificate status, removing the client's need to contact an external responder at all.
Why OCSP Stapling Was Supposed to Fix This
OCSP stapling, defined in RFC 6066 in 2011, was designed to solve exactly this problem. The idea is simple: during the TLS handshake, the server fetches an OCSP response from the certificate authority (CA) ahead of time and appends it to the Certificate message. The client can verify the staple without making a separate HTTP request to the OCSP responder. This reduces latency and eliminates the responder as a single point of failure.
In practice, stapling works well—when the server keeps fresh staples. Many major CDNs and web servers support it. Apache and nginx have had stapling support for years. OpenSSL added it in version 0.9.8. Yet adoption remains uneven. A 2023 scan by the University of Michigan found that only about 40% of HTTPS servers on the web use OCSP stapling. The rest still rely on the client fetching revocation status directly.
The protocol has a subtlety: the server must periodically fetch a new OCSP response before the current one expires. The typical validity period for an OCSP response is a few days, but some CAs issue responses valid for only a few hours. If the server fails to refresh, it either sends a stale staple or omits the extension entirely. Clients then fall back to fetching the OCSP response themselves—or, if they cannot reach the responder, they decide how to handle the missing data.
That decision is where the trouble begins. Browser vendors disagree on the right behavior. Chrome uses soft-fail: if the OCSP responder is unreachable or returns an error, the connection proceeds. Safari, on the other hand, defaults to hard-fail: if the staple is missing and the responder cannot be reached, the connection is terminated. This inconsistency means that a server that works fine in Chrome might break entirely in Safari.
The Two Failure Modes Nobody Documents
OCSP stapling introduces two failure modes that are rarely discussed in deployment guides. The first is the stale staple: the server holds an OCSP response that has expired. When the client receives it, it must decide whether to accept it. Most TLS libraries reject expired staples by default, causing a fallback to direct OCSP fetching. If that fetch also fails, the behavior depends on the client's revocation policy.
The second failure mode is the missing staple: the server sends a Certificate message without the OCSP response extension. This can happen if the server's stapling daemon crashes, if the OCSP responder is unreachable at the moment of refresh, or if the server is misconfigured to not staple at all. In this case, the client must either fetch the OCSP response itself or rely on a cached response from a previous session.
We encountered both during our investigation. In the weeks before the outage, we noticed intermittent Safari connections failing with a "revocation check failed" error. The staple was present but expired by two minutes. Safari's strict policy rejected it, and the fallback to the OCSP responder failed because the CDN's load balancer was already dropping queries. Chrome users saw no issue because Chrome's soft-fail policy allowed the connection to proceed even without a valid staple.
These failure modes are exacerbated by the fact that OCSP responses are not designed for high availability. The responder is typically a single endpoint or a small set of endpoints. If it goes down, there is no fallback—the client either accepts the risk or refuses the connection. The protocol has no mechanism for the server to prove that it attempted to fetch a response but failed.
Rebuilding the Handshake with Must-Staple
The solution that emerged from our postmortem was must-staple, a TLS extension that forces the server to provide a fresh OCSP staple or abort the handshake. Must-staple is signaled by a flag in the certificate itself, specifically the TLS Feature extension (RFC 7633). When a client sees this flag, it refuses to accept a handshake without a valid staple. The server must fetch and cache the OCSP response before the client connects.
We decided to implement must-staple across our entire fleet. This required changes at multiple levels: we needed certificates with the must-staple flag, a TLS library that enforces the requirement, and a background process to keep staples fresh. We chose OpenSSL 1.1.1w as our baseline and patched it to add strict must-staple enforcement. The patch was roughly 200 lines of C, modifying the handshake state machine to reject any Certificate message that lacked the OCSP response when the certificate indicated must-staple.
Testing was the hardest part. We built a test harness that simulated various failure scenarios: stale staple, missing staple, responder timeout, and responder returning a revoked status. We ran the tests against Let's Encrypt's OCSP endpoint, which is one of the most reliable in the industry. Under load, we found that the staple refresh process had to complete within 500 milliseconds to avoid delaying the handshake. We tuned our refresh interval to fetch a new response 60 seconds before the current one expired.
The rollout was gradual. We started with a single region, routing 5% of traffic to the must-staple-enabled servers. We monitored staple age, handshake failure rates, and latency. After two weeks with no incidents, we expanded to all regions. The transition took three months in total, including the time to reissue certificates with the must-staple flag from our CA.
How We Wrote a Fallback-Free Client
On the client side, we eliminated all OCSP fetching logic. The client no longer makes HTTP requests to OCSP responders. Instead, it relies entirely on the server-provided staple. This simplifies the client code and removes a potential vector for network errors. The client's only responsibility is to verify the staple's signature and check its validity period.
On the server, we implemented a background staple refresher as a goroutine (our services are written in Go). The refresher runs every 30 seconds, fetching a new OCSP response for each certificate. It stores the response in an in-memory cache with a mutex. When a new TLS connection arrives, the server reads the staple from the cache and includes it in the handshake. If the cache is empty—for example, right after a server restart—the server fetches the staple synchronously, adding a one-time latency of roughly 100–200 milliseconds.
We also added a metrics dashboard that tracks staple age, error rate, and fetch latency. Each staple is timestamped with its nextUpdate field. The dashboard alerts if any staple is within 10% of its expiration. This allowed us to catch a misconfigured intermediate CA that was issuing OCSP responses with a validity of only 1 hour instead of the expected 4 days. We fixed that by updating our CA's configuration.
The fallback-free design has a trade-off: it increases server complexity. The refresher must be reliable, and the cache must be protected from race conditions. We use a health check that verifies the staple is present and valid before the server accepts TLS connections. If the refresher fails for more than two consecutive cycles, the server stops accepting new connections and alerts the operations team.
Lessons from the Deployment at Scale
Deploying must-staple at scale revealed several lessons that apply beyond our specific incident. First, stapling adds a small but measurable latency to the handshake. In our measurements, the stapled handshake added 0.3–0.8 milliseconds compared to a non-stapled handshake. This is due to the extra bytes in the Certificate message and the signature verification. For most applications, this is negligible, but for latency-sensitive services, it may be worth profiling.
Second, must-staple eliminates the OCSP responder as a single point of failure, but it introduces a new dependency on the server's staple refresher. If the refresher fails, the server cannot accept new connections. We mitigated this with redundancy: each server runs its own refresher, and the load balancer distributes traffic across multiple servers. A single server's failure does not affect the fleet.
Third, must-staple reveals misconfigured intermediate CAs. During our rollout, we discovered that three of our internal certificate chains had expired staples embedded in the chain. The intermediate CA had not updated its OCSP responder configuration after a migration. The must-staple flag forced the client to reject these chains, exposing the misconfiguration. We fixed the intermediate CA's settings and reissued the affected certificates.
Finally, the deployment reinforced the importance of consistent revocation policies across clients. While must-staple standardizes the behavior for servers that support it, not all clients enforce it. As of late 2024, Chrome still uses soft-fail for non-must-staple certificates. This means that a server that supports stapling but not must-staple may still encounter the same failure modes we experienced. We recommend that any organization running production TLS services adopt must-staple and ensure all clients in their ecosystem enforce it.
The rewrite was not a silver bullet. It required coordination with our CA, changes to our TLS library, and a careful rollout. But it eliminated a class of outages that had plagued us for years. The next time an OCSP responder goes silent, our handshake will not even notice.
Alternative Approaches and Trade-offs
Must-staple is not the only way to handle OCSP failures. Some teams choose to implement a local OCSP responder proxy that caches responses and provides a fallback if the upstream responder is unreachable. This approach keeps the client-side logic unchanged but adds operational complexity: you need to run and maintain the proxy, and it introduces a new potential point of failure. We considered this option but rejected it because it still relies on the proxy being reachable and correctly configured. A misconfigured proxy could drop queries just as easily as the CDN's load balancer did.
Another alternative is Certificate Revocation Lists (CRLs), which are lists of revoked certificates published by CAs. CRLs can be cached locally and checked without an online query. However, CRLs can grow large—some CAs publish lists with tens of thousands of entries—and they are often updated only daily or weekly. This means a certificate revoked minutes ago might not appear on the CRL for hours. For compliance requirements that demand near-real-time revocation checking, CRLs are insufficient.
There is also the option of using short-lived certificates that expire in hours or days, effectively making revocation unnecessary. This approach is used by some large-scale services, including certain cloud providers. The trade-off is that you must automate certificate renewal at a high frequency, which requires robust infrastructure. If the renewal pipeline fails, certificates expire and service goes down. We evaluated this but decided against it because our CA's API rate limits made frequent reissuance impractical at our scale—we would have needed to request thousands of new certificates per day.
Finally, some teams simply accept the risk and disable OCSP checking entirely. This is common in internal networks where the threat of revoked certificates is low. But for services that handle sensitive data or must comply with standards like PCI DSS, disabling revocation checking is not an option. Our compliance requirements gave us no choice but to implement a robust solution.
Edge Cases We Discovered During the Rollout
During the rollout, we encountered several edge cases that forced us to revise our implementation. One was the handling of intermediate certificates in the chain. The must-staple flag is typically set on the leaf certificate, but some clients also expect staples for intermediate CAs. We found that a small fraction of clients would reject a handshake if an intermediate certificate's OCSP response was missing, even if the leaf staple was present. We added logic to staple for the entire chain, which increased the staple size but eliminated those failures.
Another edge case involved clients that cached OCSP responses from previous sessions. When we started stapling, these clients would sometimes ignore the staple and use their cached response instead. This caused problems if the cached response was stale or indicated a different status. We worked around this by setting the OCSP response's nextUpdate field to a very short interval (10 minutes) during the transition period, forcing clients to re-fetch. Once we confirmed that all clients were receiving staples, we extended the interval back to 4 days.
We also discovered that some load balancers do not forward the TLS extension that carries the staple. If a client connects through a load balancer that terminates TLS and then re-encrypts to the backend, the staple might be lost. We had to reconfigure our load balancers to pass through the staple extension or to perform stapling at the load balancer level. This required coordination with our infrastructure team and a brief maintenance window.
Operationalizing the Staple Refresher
The staple refresher is the most critical component of our must-staple deployment. We initially implemented it as a single goroutine that fetched staples for all certificates sequentially. Under load, this caused delays when one certificate's OCSP responder was slow. We redesigned it to use a worker pool with a configurable concurrency limit, typically 10 concurrent fetches. This reduced the maximum refresh latency from several seconds to under 200 milliseconds.
We also added a circuit breaker to the refresher. If the OCSP responder returns errors for more than 50% of requests in a 5-minute window, the circuit breaker trips and the refresher stops fetching. Instead, it serves the last known good staple until the circuit resets. This prevents a responder outage from cascading into a complete failure of the refresher. The circuit breaker resets after 30 seconds of successful fetches.
Monitoring the refresher is essential. We track the number of stale staples served, the number of synchronous fetches (which indicate cache misses), and the time since the last successful fetch for each certificate. Alerts fire if any staple is older than 80% of its validity period. This gave us early warning when a CA changed its OCSP responder URL without notice—a scenario that had caused outages for other teams in the past.
The rewrite was not a silver bullet. It required coordination with our CA, changes to our TLS library, and a careful rollout. But it eliminated a class of outages that had plagued us for years. The next time an OCSP responder goes silent, our handshake will not even notice.