When engineers break production: blame, safety, and resilient systems
In 2016, while interviewing for a technical role at Facebook, I asked what controls would prevent me from breaking production. The answer was memorable: I would break production. The point was not that outages were desirable. It was that, in a system operated and changed by people, pretending mistakes could be eliminated would be more dangerous than preparing for them.
That idea was consistent with Facebook's old “move fast and break things” reputation, but slogans are not operating models. A serious engineering organization still needs reviews, automated safeguards, limited blast radius, observability, rollback procedures, and people accountable for careful work. The real question is what happens after those defenses fail: do we improve the system, or find a person to punish?
The 2021 Facebook outage was not simply “DNS”
Facebook's October 4, 2021 outage is often compressed into “it was DNS” or “someone ran the wrong command.” Both descriptions omit the important part.
According to Meta's published incident account, a command used during routine backbone maintenance unintentionally disconnected its data centers. A system was supposed to audit and stop such a command, but a bug in that safeguard allowed it through. The backbone failure then made DNS locations withdraw their BGP advertisements. Normal internal tools and remote access also failed, so engineers had to reach data centers physically before recovery could begin.
A person issued the command. It is still misleading to treat the person as the root cause. The event required a hazardous command, a defective audit control, tightly coupled backbone and DNS behavior, inaccessible recovery tooling, and difficult emergency access to align. In psychologist James Reason's terms, the “person approach” asks who made the error; the “system approach” asks which working conditions and defenses allowed one error to become a catastrophe.1
There is no reliable public evidence for the claim that a large group of Facebook network engineers left because of this incident or that Facebook subsequently moved substantial production infrastructure to AWS. The stronger argument does not need those claims.
What psychological safety actually predicts
Psychological safety does not mean that errors have no consequences or that standards are optional. It means that people believe they can report a mistake, ask for help, challenge an assumption, or propose an experiment without being humiliated or punished for the interpersonal risk itself.
Amy Edmondson's foundational field study of 51 work teams found that psychological safety was associated with learning behavior, and that learning behavior mediated team performance.2 This matters during incidents because the organization needs bad news early. An engineer who fears punishment has a rational incentive to delay disclosure, minimize uncertainty, or quietly work around a weak control. Every minute of silence increases the blast radius.
The finding is not based on one famous paper. A 2017 meta-analysis combined 136 independent samples, representing more than 22,000 people and nearly 5,000 groups. It found psychological safety related to outcomes including information sharing, learning, voice, creativity, and performance beyond adjacent ideas such as positive leader relationships and work engagement.3
Organizational sociology describes the inverse condition as organizational silence: shared beliefs that speaking about problems is unwise. Morrison and Milliken argued that structures and managerial beliefs can turn individual caution into a collective pattern that blocks change and development.4 In production operations, this appears as “never touch a working system,” warnings softened until they are harmless, and risky manual work that everyone knows about but nobody wants to own.
Error tolerance alone is not enough
The evidence does not justify being casual with production. It supports an error-management culture: expect fallible actions, detect errors quickly, communicate them, contain their consequences, analyze them, and change the conditions that produced them.
Van Dyck, Frese, Baer, and Sonnentag studied 65 Dutch and 47 German organizations. Cultures characterized by communicating about, detecting, analyzing, and rapidly correcting errors were positively associated with goal achievement and objective or longitudinal measures of economic performance.5 Baer and Frese similarly found, across 47 mid-sized German companies, that climates for initiative and psychological safety supported the relationship between process innovation and firm performance.6 The practical implication is not “break things freely.” It is “make small, recoverable changes and extract information from failure.”
There is also a useful boundary condition. A 2023 paper covering five studies found that very high psychological safety could reduce performance on routine, standardized tasks. Collective accountability buffered that effect.7 That is directly relevant to operations: a creative architecture discussion and a routine certificate rotation are not the same kind of work. Teams should invite challenge and experimentation while still requiring a runbook, peer review, maintenance window, tested rollback, and explicit ownership where the task is known and repeatable.
Psychological safety answers “Can I tell the truth and take a thoughtful risk?” Accountability answers “Did I prepare, follow the agreed controls, and learn from the outcome?” Reliable teams require both.
Why punishing the operator is expensive
Removing the person closest to an incident can create the appearance of action. It may also discard the engineer with the freshest understanding of a system's undocumented dependencies, failure modes, and recovery path. A 2023 systematic review synthesized 91 empirical studies on knowledge loss caused by organizational turnover and found effects across organizational and unit levels.8 That literature does not prove that any specific Facebook departure caused a particular cost, but it supports the general warning: operational context is an asset, and turnover can destroy it.
Dismissal can be justified for recklessness, concealment, repeated disregard of controls, or malicious behavior. An honest error made inside an approved process is a different category. Conflating them teaches everyone that transparency is personally dangerous. The next incident then begins with worse information.
What a high-trust production culture looks like
The goal is not to protect applications from engineers by preventing change. An unchanged production system still accumulates expiring certificates, obsolete dependencies, security exposure, capacity limits, and operational knowledge that exists in fewer heads each year. Avoidance converts planned change into emergency work.
A resilient organization makes the safe action the easy action:
- Constrain blast radius. Use staged rollouts, canaries, feature flags, rate limits, and isolation boundaries.
- Automate hazardous checks. Treat a guardrail failure as a system defect, not as proof that people must never make mistakes.
- Preserve a recovery path. Test rollback, out-of-band access, backups, and incident tooling under the same failure modes that can take production down.
- Make changes observable. Connect deployments and infrastructure changes to telemetry so responders can quickly correlate cause and effect.
- Reward early disclosure. Separate rapid incident communication from the later evaluation of preparation and judgment.
- Review conditions, not character. Ask what made the action reasonable at the time, which defenses failed, and how recurrence will be harder.
- Keep accountability specific. Distinguish an honest mistake from negligent repetition, concealment, or deliberate bypass of controls.
Readiness beats the promise of perfection
The research supports the core claim with an important correction. Aggression after mistakes can suppress voice, learning, and experimentation. Unqualified tolerance can also weaken routine execution. The effective position is neither fear nor carelessness: it is psychological safety combined with engineering controls and collective accountability.
Any sufficiently complex application can go down. The useful measure of an engineering organization is not whether it can promise that nobody will ever break production. It is whether a mistake is contained, detected, reported without delay, recovered from quickly, and converted into a stronger system.
References
- Reason, J. (2000). Human error: models and management. BMJ, 320, 768–770.
- Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383.
- Frazier, M. L., Fainshmidt, S., Klinger, R. L., Pezeshkan, A., & Vracheva, V. (2017). Psychological safety: A meta-analytic review and extension. Personnel Psychology, 70(1), 113–165.
- Morrison, E. W., & Milliken, F. J. (2000). Organizational silence: A barrier to change and development in a pluralistic world. Academy of Management Review, 25(4), 706–725.
- van Dyck, C., Frese, M., Baer, M., & Sonnentag, S. (2005). Organizational error management culture and its impact on performance. Journal of Applied Psychology, 90(6), 1228–1240.
- Baer, M., & Frese, M. (2003). Innovation is not enough. Journal of Organizational Behavior, 24(1), 45–68.
- Eldor, L., Hodor, M., & Cappelli, P. (2023). The limits of psychological safety: Nonlinear relationships with performance. Organizational Behavior and Human Decision Processes, 177, 104255.
- Galan, N. (2023). Knowledge loss induced by organizational member turnover: a review of empirical literature, synthesis and future research directions (Part I). The Learning Organization, 30(2), 117–136.