One Edge Engineer Who Lost Bus Factor Data Wrote an Automated Handoff Contract
It started with a midnight pager alert that nobody on the team could triage. The engineer who had written the rate-limit configuration had left the company three weeks earlier. The remaining three people stared at a YAML file they had never seen before, containing cryptic thresholds and a comment that read simply: "don't touch this." That night, the bus factor for that edge config became real — and it was exactly one.
When Bus Factor Data Goes Missing on a CDN Team
The team was small: four engineers responsible for over 200 edge configurations spread across a global content delivery network. These configs controlled routing, caching, rate limiting, and authentication for dozens of services. Some were simple redirects; others were complex Lua scripts that rewrote request paths and headers. The bus factor — the number of people who could reconstruct a piece of knowledge if the primary owner were hit by a bus — was dangerously low for the most critical configs.
A quick audit revealed that twelve files had exactly one author in the git history. Those twelve files handled traffic for the company's highest-revenue APIs. The engineer who wrote them had been the sole reviewer, the sole deployer, and the sole mind that understood why certain magic numbers existed. When he left, the knowledge left with him.
Pager fatigue had been rising for months. Each midnight incident triggered a scramble to find someone who vaguely remembered a config change. The team's mean time to acknowledge alerts had crept from five minutes to over twenty. The on-call rotation felt less like a duty and more like a guessing game.
The problem wasn't laziness or bad intentions. It was a systemic failure to treat configuration as code — and code as something that must survive its authors. The team had documentation, but it lived in a wiki that nobody updated. They had code review, but it rarely caught missing context for future maintainers.
The Handoff Contract That Replaced Tribal Knowledge
One engineer, who had been burned by the midnight incident, decided to try something different. Instead of writing more wiki pages or scheduling knowledge-transfer sessions that would inevitably grow stale, she built an automated handoff contract. The idea was simple: every edge config file in the repository must carry a structured metadata block that declares its owner, escalation path, test coverage, and bus factor status.
The contract was enforced via a JSON schema that ran as a pre-commit git hook. Before a config could be merged, the hook validated that the metadata block existed and met minimum requirements. The owner field had to be a valid email alias, not a personal address. The escalation path had to list at least two people. Test coverage had to reference specific integration tests. And, crucially, the bus factor — computed as the number of unique authors who had touched the file in the last six months — had to be at least two.
If the bus factor dropped below two, the merge was blocked. The only way to override was to get a second person to review and explicitly approve, which effectively forced knowledge transfer. The contract lived in the repository alongside the configs, not in a separate wiki. It was versioned, reviewable, and automatically enforced.
The schema was deliberately minimal. The engineer knew that too many required fields would create friction and encourage people to game the system. She started with just four fields: owner, escalation, tests, and bus factor. Over time, the team added optional fields for documentation links and incident history. The contract became a living document that evolved with the team's needs.
One trade-off that emerged early: the bus factor metric relied on git blame, which only counted authors who had committed changes. If a team member reviewed code thoroughly but never made a commit, their knowledge was invisible to the metric. The team addressed this by encouraging reviewers to make minor commits or annotations on critical files, but it remained a blind spot. Another blind spot was that the bus factor check only applied to files that were actively being modified. Configs that were stable and untouched for months could still have a bus factor of one, but the check would never flag them because no merge was attempted. The team added a periodic audit script that scanned all config files regardless of change activity, ensuring that dormant single-author files were surfaced.
How One Team Measured Bus Factor With Git Blame
The bus factor metric came from a simple script that ran git blame on every edge config file and counted unique authors. The script produced a histogram showing how many files had one author, two authors, and so on. The team set a threshold: any file with fewer than two unique authors in the last six months was flagged as a critical risk.
The results were sobering. Of the 200+ configs, 12 were single-author files. Those 12 files handled routing for the company's payment API, authentication redirects, and a custom cache invalidation endpoint. The engineer who had written most of them had left two months earlier. The team had known the bus factor was low, but the data made it impossible to ignore.
The script ran automatically after every merge and posted a summary to the team's chat channel. Over time, the histogram shifted. Single-author files dropped from 12 to 3 within three months. The team adopted a practice of pairing on new configs and rotating review responsibilities so that no file ever belonged to one person.
But the script had a blind spot: it measured authorship, not understanding. A file could have multiple authors who each added a line without understanding the whole. The team supplemented the metric with a manual review process for the most critical configs. They also added a comment convention: any config with complex logic required a rationale comment explaining why the logic existed. For example, a config that set a rate limit of 100 requests per second for a specific endpoint now had a comment like "# Rate limit set to 100 RPS based on load test results from March 2023; average peak traffic is 80 RPS with burst to 120 RPS. See incident #452 for details." This made the rationale discoverable even if the original author was unavailable.
The team also experimented with a "bus factor score" that combined authorship count with code review participation. They ran a script that parsed git log for both authors and reviewers (using the Reviewed-by trailer). Files with only one reviewer were flagged even if they had multiple authors. This caught cases where one person was the sole reviewer for a file, creating a knowledge bottleneck at the review stage. The score was displayed in a dashboard alongside the authorship histogram, giving a more nuanced view of knowledge distribution.
Why Edge Infrastructure Amplifies Bus Factor Risk
Edge configurations are uniquely dangerous because they sit at the intersection of global scale and local control. A single misconfigured rate limit can cause a cascading failure that takes down traffic for an entire region. Unlike backend services, which can be rolled back in seconds, CDN changes often take minutes to propagate and even longer to revert due to stale caches.
The team had experienced this firsthand. A config change that looked safe in a local test caused a 5xx spike across Europe when it hit production. The rollback took 15 minutes because the CDN vendor's API had no atomic rollback — they had to re-deploy the old config and wait for propagation. The incident cost an estimated $50,000 in lost revenue and engineering time.
Edge infrastructure also lacks the staging environments that backend teams take for granted. Many CDN vendors offer a staging environment, but it often doesn't mirror production traffic patterns. The team had to rely on canary deployments and feature flags, which added complexity to an already fragile system.
The combination of global blast radius, slow rollback, and limited staging makes edge configs a high-stakes domain. A bus factor of one is not just a knowledge risk — it's an operational risk that can cause real financial damage. The handoff contract was the team's attempt to reduce that risk without adding bureaucratic overhead.
To further mitigate risk, the team implemented a "pre-flight check" that ran before any config deployment. The check simulated the config against a snapshot of recent traffic patterns, flagging anomalies like sudden spikes in error rates or cache miss ratios. This caught several potential incidents before they reached production. One notable example: a config change that would have disabled caching for a high-traffic image endpoint was blocked by the pre-flight check because it would have increased origin load by an estimated 400%. The engineer who made the change had not realized the impact, and the pre-flight check saved the team from a costly outage.
Funding the Fix: How a Small Team Got Budget for Automation
Getting buy-in for automation required a different kind of contract: a business case. The engineer calculated the cost of a single P0 incident — roughly $50,000 in lost revenue, engineering time, and customer trust. She then estimated that the team had experienced three such incidents in the past year, all of which could have been prevented or mitigated with better bus factor management.
The math was compelling. Three incidents at $50,000 each meant $150,000 in annual losses. The automation project — the git hooks, the schema, the monitoring script — would take about two months of one engineer's time, or roughly $30,000 in salary. The return on investment was five to one, assuming the automation prevented just one incident per year.
The VP of infrastructure approved the budget on the strength of that analysis. The engineer used incident postmortems as evidence, highlighting the pattern of single-owner configs causing prolonged outages. She also pointed out that the team was already spending thousands of hours on pager fatigue and manual knowledge transfer.
The project reused open-source tooling where possible: the JSON schema validator was a standard library, the git blame script was a 50-line Python script, and the chat integration used a webhook. The team contributed the schema back to an internal tooling repository, and later shared it at a conference. The cost was low, but the impact was lasting.
However, the business case had a hidden assumption: that the automation would prevent incidents. In practice, the team found that the contract prevented some incidents but also created new ones. For example, a hotfix that needed to be deployed immediately was blocked by the bus factor check, causing a delay that escalated a minor issue into a major one. The team learned to add an emergency override that bypassed the check for critical fixes, but required a post-hoc review within 24 hours. This trade-off between safety and speed was a recurring theme in the team's retrospective discussions.
The Contract's Hidden Leverage for On-Call Engineers
The handoff contract did more than enforce bus factor — it gave on-call engineers a lifeline. Before the contract, a midnight alert meant scrolling through git logs and hoping to find someone who knew the config. After the contract, the on-call engineer could query the metadata block directly from the repository and see the owner's email alias, escalation path, and a link to relevant tests.
The escalation path was particularly valuable. The team had a policy that the escalation path must include at least two people who had reviewed the last change. That meant the on-call engineer could contact someone who had seen the code, even if they hadn't written it. Mean time to acknowledge dropped from 20 minutes to under 5 within the first month.
New hires also benefited. Instead of reading outdated wiki pages, they could read the contract diffs — the commit history of the metadata block showed who had owned a config over time and why it had changed. The contract became a lightweight onboarding document that stayed accurate because it was enforced by the build pipeline.
The team measured a 30% reduction in pager fatigue, as measured by the number of alerts that required escalation to a second person. Fewer incidents went unresolved because the on-call engineer could quickly find the right person. The contract didn't eliminate late-night pages, but it made them less terrifying.
One unexpected benefit was that the contract surfaced undocumented rate limits. The metadata block required a field for "rate limit rationale", and engineers had to explain why a particular limit existed. In several cases, the rationale revealed that the limit was a guess or a copy-paste from another config. The team then ran load tests to validate or adjust the limits, preventing future incidents caused by misconfigured thresholds. For instance, one config had a rate limit of 50 requests per second for a critical API, but the rationale field said "copied from auth service". After load testing, the team found the actual capacity was 200 RPS, and they updated the limit accordingly, reducing customer-facing errors during peak traffic.
What the Handoff Contract Teaches About Sustainable CDN Ops
The handoff contract was a success, but it wasn't a silver bullet. Automation alone cannot fix the social dynamics that create bus factor problems. Some engineers resist sharing ownership because they fear losing control or being blamed for others' mistakes. The team had to address those concerns through culture, not code.
The contract also introduced friction. Blocking merges because of bus factor felt heavy-handed to some team members. The override mechanism was used more often than expected, mostly for urgent hotfixes. The team eventually added a grace period: the bus factor check only blocked merges if the file had been single-author for more than three months. That gave people time to transfer knowledge without blocking critical fixes.
Bus factor is a metric, not a feeling. The team learned to treat it as one data point among many, not a hard rule. The real value was in making the invisible visible — forcing conversations about ownership and handoff that had been avoided for years. The contract was a tool for those conversations, not a substitute for them.
The engineer who built the contract recommends starting small: pick one critical config file, add the metadata block manually, and see how it feels. Then automate the validation. The template is open-source and has been adopted by several other teams inside the company. The lesson is that sustainable CDN operations require both tooling and trust — and the contract helps build both.
Looking back, the team identified a few key principles that made the contract effective. First, the metadata block had to be lightweight — adding too many fields would have made it a burden. Second, the enforcement had to be automatic, but with a humane escape hatch for emergencies. Third, the bus factor metric had to be supplemented with qualitative reviews, because no algorithm can capture the depth of someone's understanding. The team also learned that the contract was most valuable not when it blocked changes, but when it prompted conversations. Over time, the team's culture shifted from "this is my config" to "this is our config", and that was the real win.