Building a Resilient Tech Organisation: Systems Over Heroics
The call came in on a Sunday evening, which is usually the first sign something has been wrong for longer than anyone admitted. A mid-market fintech’s core payments reconciliation job had failed silently for the third weekend running, and the only person who understood how to diagnose it - a senior engineer everyone in the company referred to, without irony, as “the guy who knows how that works” - was on a flight with no signal for six hours. Nobody else could read the logs meaningfully. Nobody else knew which of the four downstream services depended on the job completing before Monday’s open. The on-call rotation existed on paper, but in practice it routed every non-trivial incident to this one person by informal convention, because he was faster and everyone else was afraid of touching a system they didn’t fully understand. When I was brought in afterward to review what had happened, the CTO’s framing was that they’d been “unlucky” with timing. They hadn’t been unlucky. They had built an organisation whose actual reliability depended entirely on one person’s availability, and had been mistaking his competence for the system’s resilience for roughly three years.
This is the pattern I want to name precisely, because it’s disguised as a virtue almost everywhere it occurs: hero culture. Every tech organisation has, or has had, someone like this - the person who gets paged first because they’re the one who actually knows, who stays late without being asked because they can see the consequence of not staying, who becomes, without anyone deciding it deliberately, the organisation’s actual disaster recovery plan. The uncomfortable finding, across every version of this I’ve reviewed, is that hero culture doesn’t emerge from bad intentions. It emerges from an organisation quietly optimising for the person who solves the immediate problem fastest, at the expense of ever building the thing that would make the problem un-solvable by design. The heroism is real. The resilience it’s mistaken for is not, and the gap between the two is where the actual risk accumulates, invisibly, for years, until a Sunday evening makes it visible at the worst possible time.
Why Competence Gets Mistaken for Resilience
The confusion is understandable because, from inside the organisation, the two look identical for a long time. Incidents get resolved. Deadlines get hit, usually because someone worked the weekend to hit them. Systems that shouldn’t really still be running keep running, because someone who understands their undocumented quirks intervenes before they fail publicly. Leadership sees a track record of problems getting solved and reasonably concludes the organisation is capable - without ever separating the question of whether the organisation is capable from the question of whether one or two specific people are, and whether that distinction has ever been tested.
It gets tested exactly once, involuntarily, when the hero is unavailable - on leave, poached by a competitor, burned out and finally saying no, or simply on a flight at the wrong moment - and what the organisation discovers in that moment is the actual state of its infrastructure, its documentation, and its decision-making capacity, stripped of the one person who had been quietly compensating for all three. This is what I mean by the difference between competence and resilience: competence solves the problem in front of you; resilience means the organisation doesn’t need a specific person to solve it. Most tech organisations I review have invested heavily in the first and have no clear mechanism for building the second, because the first is visible and rewarded continuously, while the second is invisible until the day it’s tested - and by definition, you don’t get advance warning of which day that will be.
Four Patterns That Build Hero Dependency Without Anyone Deciding To
The informal escalation shortcut.
Every organisation has an official incident process - an on-call rotation, a runbook, an escalation ladder - and in organisations with hero dependency, that official process exists mostly as documentation nobody actually follows under pressure, because everyone has learned that the fastest path to resolution is routing directly to the person who actually knows, bypassing the rotation entirely. This isn’t malicious; it’s rational individual behaviour that produces a collectively catastrophic outcome, because every time someone routes around the official process to get a faster fix, the official process gets marginally weaker and the hero dependency gets marginally stronger, and nobody experiences this as a decision because it never gets made explicitly. It accumulates one shortcut at a time until the org chart says one thing and the actual dependency graph says another.
Tribal knowledge as load-bearing infrastructure.
The reconciliation job in the opening example wasn’t undocumented because anyone had decided documentation wasn’t worth doing. It was undocumented because documenting it properly would have taken the one person who understood it away from solving the next fire, and there was always a next fire, and the organisation’s incentive structure rewarded him for solving fires, not for writing the runbook that would make his own presence unnecessary. This is the structural trap: the people best positioned to eliminate a knowledge single-point-of-failure are the same people whose daily incentives push them toward using that knowledge to solve problems quickly rather than toward transferring it, and no organisation I’ve reviewed has solved this by asking people to document more in their spare time. It gets solved, when it gets solved, by making knowledge transfer an explicitly assigned and measured piece of work, not a virtuous afterthought competing with an always-full queue of urgent fixes.
Firefighting rewarded over fire prevention.
Performance review cycles, promotion narratives, and informal reputation inside engineering organisations consistently reward the visible save - the person who stayed up all night and fixed the outage - far more reliably than they reward the invisible absence of an outage that would have happened if someone hadn’t spent a quarter earlier redesigning the failure-prone system. This is a measurement problem with real behavioural consequences: if the organisation’s reward signal only fires when there’s a fire to fight, the rational response from capable people is to become excellent firefighters rather than to invest in fire prevention, because fire prevention is unrewarded, hard to attribute, and actively reduces the number of opportunities to be visibly heroic. I’ve watched organisations promote their best firefighters into leadership roles repeatedly while the engineers quietly doing preventive systems work - the unglamorous kind that means nothing breaks in the first place - got passed over for having a thinner list of dramatic saves.
Post-incident reviews that thank the hero instead of fixing the structure.
The retrospective after most fire-drill incidents follows a predictable shape: relief that it’s resolved, genuine gratitude toward whoever fixed it, a timeline of what happened, and an action item or two about monitoring. What’s consistently missing is the harder question - why did this depend on one specific person being available, and what would need to be true structurally so that the next version of this incident doesn’t depend on anyone’s specific availability at all. Gratitude toward the hero is appropriate and should be expressed. It is not, on its own, a fix, and organisations that stop at gratitude are guaranteeing a repeat performance, with the same person carrying the same disproportionate load until they can’t or won’t anymore.
What Systems-Based Resilience Actually Requires
The fix isn’t “document everything” as a slogan, because that instruction has been given to every engineering organisation at every all-hands for a decade and has produced, on its own, almost no measurable change in hero dependency. What actually shifts the pattern is a small number of structural changes that remove the informal shortcuts hero dependency runs on, rather than appealing to individual discipline to close a gap the incentive structure keeps reopening.
The escalation path needs to be the fastest path, not just the official one - which means investing in the on-call rotation and runbooks enough that routing through them is genuinely quicker than finding the one person who knows, because as long as the informal shortcut remains faster under pressure, it will keep winning regardless of what the org chart says. This is the same discipline I’ve written about in the context of execution accountability more broadly: the gap between the designed process and the process people actually use under pressure needs a visible, monitored channel, because it stays invisible until it’s expensive. (check Execution Accountability: How to Build a Culture That Delivers) for the underlying mechanism - the same principle that applies to status reporting applies directly to whether an incident process is actually being used.)
Knowledge transfer needs to be assigned, scheduled, and measured as real work competing for the same calendar space as feature delivery, not left to compete against an always-full queue of urgent fires where it will lose every time by default. Concretely, this means naming specific systems with a documented single point of knowledge failure, assigning a named second person to reach working competence within a defined window, and tracking that as a deliverable with the same seriousness as a shipped feature - because if it’s not tracked, it’s not real, regardless of how many times it gets mentioned as a priority.
Reward structures need an explicit mechanism for recognising prevention, not only response - which is genuinely harder to measure than a dramatic save, and organisations that don’t solve the measurement problem will keep defaulting to rewarding visible heroics, because visible heroics are easy to see and prevention is, by design, invisible. This is directly connected to the governance discipline that fast-growing organisations need more broadly: growth doesn’t create hero dependency on its own, but it multiplies the cost of every shortcut the organisation has been quietly taking, at exactly the pace where nobody has time to notice. (See [Scaling Without Chaos: Governance Frameworks for Fast-Growth Teams](/scaling-without-chaos-governance-fast-growth-teams) for how this compounds specifically under growth conditions.)
And post-incident reviews need a standing structural question, asked every time regardless of how quickly the incident was resolved: what specific single point of dependency - human or system - did this incident expose, and what is the committed fix with an owner and a date. Without that standing question, retrospectives will keep producing gratitude and monitoring tickets, which feel like progress and change nothing about the underlying dependency graph.
Implementation Trade-offs
None of this is free, and the trade-offs deserve honest treatment rather than being waved away as pure upside. Investing in the official escalation path and runbooks over the informal shortcut takes real engineering time away from feature delivery in the near term, and it will be tempting to defer this work indefinitely because the informal shortcut keeps working - right up until the specific person it depends on is unavailable, at which point the cost of not having done it arrives all at once and considerably larger than the cost of having done it gradually.
Making knowledge transfer an explicit, tracked deliverable creates near-term friction with the people who are best at their jobs, because the instruction is effectively asking your most capable individual contributors to spend time making themselves less uniquely necessary, which runs against both the natural pull toward being the indispensable expert and, in some cases, against a genuine (if often unconscious) instinct toward job security through irreplaceability. This needs to be handled directly and honestly in how the organisation frames career growth - the message needs to be that becoming un-necessary for a specific system is what enables the next level of scope, not a threat to it, and that only works if leadership actually follows through on giving expanded scope to people who do this well.
Rebalancing reward structures toward prevention requires genuine measurement discipline that most organisations don’t currently have - you need some way of attributing “this didn’t happen because of preventive work six months ago” that’s more rigorous than anecdote, or the rebalancing will be perceived, correctly, as unfairly favouring whoever tells the better story in a calibration meeting rather than whoever actually reduced risk.
The Strategic Reflection
Hero culture persists in tech organisations not because leaders fail to notice it, but because it works, continuously and visibly, right up until the specific day it doesn’t - and that day rarely announces itself in advance. The organisations that build genuine resilience are the ones that treat every heroic save as a diagnostic signal rather than a success story: not “great job, crisis averted,” full stop, but “great job, crisis averted - now what does it tell us about a dependency we haven’t fixed, and who owns fixing it by when.” That reframe is uncomfortable in the moment, because it takes something that feels like an unambiguous win and insists on treating it as evidence of unfinished structural work, and it will meet resistance from people who’d rather be thanked than assigned a follow-up.
If your organisation depends, right now, on a small number of specific people being available for things to keep working, the diagnostic question worth sitting with isn’t whether those people are good at their jobs. They almost certainly are - that’s usually not in question. The question is what would actually happen, in concrete operational terms, if the best of them left with four weeks’ notice tomorrow, and whether anyone has mapped that answer honestly or has simply been trusting that it won’t happen on their watch. Most organisations running on heroics can’t answer that question with any confidence. The ones that build real resilience are the ones that make answering it, uncomfortably and specifically, a standing part of how they operate - before a Sunday evening forces the answer on them.


