Every engineering leader remembers their first real production incident. Mine was a payment authorization gateway that started timing out at the worst possible moment, mid-morning on a settlement day, with merchant volume climbing. I was a senior engineer then, not a CTO, and I remember the particular feeling of the room tightening, voices speeding up, and three people typing into the same channel with different theories. We resolved it eventually, but it took far longer than it should have, mostly because we were managing our own adrenaline as much as the system.
Over the years I have come to believe that the single most valuable thing a leader brings to an incident is not technical brilliance. It is calm. Not the absence of urgency, but the presence of structure under pressure. This post is about how I try to handle production incidents calmly, and how I help my teams do the same, in an environment where the systems move money and a mistake has consequences beyond a frustrated user.
Why Calm Is an Engineering Skill, Not a Personality Trait
It is tempting to treat composure as something you either have or do not. Some people are naturally unflappable, the story goes, and the rest of us just have to fake it. I have found that to be mostly wrong. Calm during an incident is a learned, practiced capability, and it is downstream of preparation. The people who appear serene in a war room are almost never improvising serenity. They are executing a process they have rehearsed, and the process is doing the work their nervous system would otherwise have to do.
This matters because panic is contagious and so is composure. When the most senior person in the channel writes in short, clipped sentences and jumps between hypotheses, everyone else mirrors that energy and the investigation fragments. When the same person writes deliberately, states what is known and unknown, and asks one question at a time, the whole group settles. Your demeanor is a load-bearing part of the incident response, whether you intend it to be or not.
In regulated fintech the stakes raise the temperature further. An outage may mean failed settlements, breached service-level agreements, regulatory notification obligations, and merchants who cannot get paid. The instinct to do something, anything, immediately, is strong precisely because the cost of inaction feels enormous. That instinct is the enemy. Calm is what lets you slow down just enough to act correctly rather than quickly.
The First Five Minutes Set the Tone
The opening minutes of an incident are where most of the damage to the response is done. Not damage to the system, that is already happening, but damage to the response itself. This is when people skip declaring an incident because they hope it will resolve itself, when two engineers independently start poking at production without telling each other, and when someone restarts a service that destroys the very evidence needed to understand the failure.
My rule is simple: declare early and declare loudly. It costs almost nothing to open an incident channel and stand it down ten minutes later if it turns out to be noise. It costs a great deal to be thirty minutes into an unstructured scramble before anyone admits this is serious. The act of formally declaring creates a container. It tells everyone where to look, who is in charge, and that the normal rules of casual debugging no longer apply.
The goal of the first five minutes is not to fix the problem. It is to establish who is coordinating, what we are seeing, and a shared place to think. Resolution can wait two minutes for structure; it will be faster for the delay.
Separating the Commander From the Fixer
One of the most reliable ways to lose composure is to ask one person to both coordinate the incident and resolve it. These are different jobs that compete for the same attention. The person elbow-deep in logs cannot simultaneously track the timeline, brief stakeholders, and decide when to escalate. When they try, they do all of it badly, and the cognitive load is exactly what tips an investigation into panic.
So I insist on a clear separation. An incident commander owns the process: they run the channel, hold the timeline, decide on communications, and protect the responders from interruption. The responders own the technical work. The commander does not need to be the most senior engineer or even understand every detail; they need to ask good questions, keep the group focused, and shield the team from the executives and account managers who will inevitably show up wanting updates.
I have learned to be deliberate about who plays which role, and to say it out loud. A few things I now treat as non-negotiable when an incident is in progress:
- One named incident commander at all times, handed off explicitly if they tire or need to dig into the technical work themselves.
- A single source of truth for the current state, usually a pinned message that is updated rather than a stream of contradictory comments.
- A scribe, even an informal one, capturing timestamps and decisions as they happen, because nobody reconstructs an accurate timeline from memory afterward.
- A clear boundary between the people working the problem and the people who need to be informed about it.
Resisting the Urge to Act Without Understanding
The strongest pull during an incident is toward action, because action feels like progress and progress feels like relief. But uninformed action is how a contained outage becomes a compounded one. I have watched a team take a degraded but functioning system fully offline because someone applied a fix to a problem they had not actually diagnosed. The cure was worse than the disease, and we then had two incidents instead of one.
Calm gives you the few seconds needed to ask the questions that change everything. What changed recently? When exactly did the symptoms start? Does the timing correlate with a deployment, a configuration change, a traffic pattern, a downstream dependency? In my experience the overwhelming majority of production incidents in a mature system trace back to a change, and the single most useful artifact in the first ten minutes is an accurate deployment and change log.
This does not mean analysis paralysis. There is a real difference between a degrading situation where every minute of investigation costs real money, and a stable-but-broken one where you can afford to think. The skill is reading which situation you are in, and being honest about it. If the system is actively losing data or money, you may need to stop the bleeding first, even crudely, and understand later. If it is merely down, you usually have more room to diagnose than your adrenaline is telling you.
Enjoying this article?
Get more like it in your inbox — practical engineering leadership, fintech, and AI. No spam, unsubscribe anytime.
Mitigate First, Understand Root Cause Later
A distinction I drill into every team I lead is the difference between mitigation and resolution. Mitigation is making the pain stop for customers. Resolution is understanding and fixing the underlying defect. These are sequential, not simultaneous, and confusing them is a common source of prolonged outages. Engineers, by temperament, want to understand why something broke before they touch it. During an active incident that instinct must be consciously overridden.
If we have a recent deployment that correlates with the symptoms, we roll it back. We do not first prove it was the cause. Rolling back a suspected change is cheap and reversible; debugging a live production failure under load is expensive and stressful. If a particular merchant integration is generating malformed requests that are poisoning a queue, we isolate that traffic before we work out why their payload changed. Restoring service buys back the calm we need to do the careful work properly.
The root cause investigation is real and necessary, but it belongs to a different phase with a different emotional register. Once customers are no longer affected, the clock pressure drops dramatically, and we can afford to be rigorous rather than fast. Conflating the two phases is what keeps teams in fight-or-flight for hours longer than the actual customer impact warranted.
Communicating Clearly While Everything Is On Fire
During an incident, communication is not overhead that distracts from the real work. It is part of the work. Poor communication generates its own secondary incidents: an executive who does not know what is happening will start their own investigation, a support team without talking points will improvise and contradict you, and merchants left in silence assume the worst and escalate. Each of these pulls responders away from the fix.
I aim for communication that is frequent, honest, and bounded. Frequent, meaning updates on a predictable cadence even when there is nothing new, because silence reads as chaos. Honest, meaning I say what we know, what we do not know, and what we are doing, without false reassurance. And bounded, meaning I commit to the next update time rather than to a resolution time. Promising a fix by a deadline you cannot control is how trust evaporates when you miss it.
In a regulated context there is an additional discipline. Some incidents trigger formal obligations, whether contractual notification windows or regulatory reporting. Part of staying calm is knowing in advance which incidents carry those obligations, so that the moment of recognition is a checklist item rather than a fresh source of dread. The commander should be tracking, from early on, whether this incident crosses any of those thresholds, so legal and compliance are engaged on time rather than in a retrospective panic.
Managing Your Own Physiology
It feels slightly unserious to talk about breathing and posture in an engineering essay, but the body is where panic lives, and ignoring it is a mistake. When the adrenaline arrives, your field of attention narrows, your working memory shrinks, and your sense of time distorts. These are precisely the faculties an incident demands. If you do nothing to manage your own state, the situation will manage it for you, and not in your favor.
The practical interventions are unglamorous and effective. I deliberately slow my typing. I read messages twice before responding. I stand up. If I notice I am holding my breath, I take a few slow ones. None of this is mysticism; it is simply refusing to let the body's stress response dictate the pace of my decisions. A leader who is visibly steady gives everyone else permission to be steady, and the inverse is just as true.
I also watch for the moment when I have been in the incident too long to be effective. Fatigue degrades judgment quietly, and the person who has been commanding for three hours is often the worst judge of their own decline. Building in explicit handoffs, and normalizing the idea that stepping back is responsible rather than weak, is one of the most protective things a culture can do. The most dangerous responder is an exhausted one who believes only they can hold it together.
Turning the Incident Into Learning, Not Blame
How an organization treats people after an incident determines how calmly they will behave during the next one. If the postmortem is a search for someone to punish, then every future incident becomes an exercise in self-protection. People will hesitate to declare, hesitate to admit what they touched, and hesitate to volunteer the awkward detail that turns out to be the key clue. Fear makes incidents longer and worse. A blameless culture is not softness; it is operational pragmatism.
A blameless postmortem assumes that people acted reasonably given what they knew at the time, and asks why the system allowed a reasonable action to cause harm. The engineer who ran the deployment that broke production is not the problem; the absence of a safeguard that would have caught it is. This framing keeps the conversation on the systemic improvements that actually prevent recurrence, rather than on individual shame that prevents nothing.
The output of this phase is what compounds over time. Every incident, handled well, should make the next one calmer: a new alert, a clearer runbook, a faster rollback path, a deployment guardrail. The calm I bring to an incident today is largely borrowed from the lessons of incidents I handled poorly years ago. Composure, in the end, is mostly the dividend of having prepared for this exact kind of bad day.

Conclusion
Handling a production incident calmly is not about being unmoved. The stakes are real, and a degree of urgency is appropriate and useful. It is about building enough structure, before and during the event, that your team can channel that urgency into correct action rather than frantic motion. Declare early, separate coordination from fixing, mitigate before you chase root cause, communicate with honesty and cadence, and treat the aftermath as learning rather than blame. Do those things consistently and calm stops being a personality trait you envy in others and becomes a capability your whole organization owns. That, more than any single technical fix, is what keeps the worst days from becoming the defining ones.
Get new posts in your inbox
Occasional, practical notes on engineering leadership, fintech, and building with AI. No spam, unsubscribe anytime.
Comments (10)
Leave a Comment
Toby Barlow
September 8, 2026
Does the "Managing Your Own Physiology" still hold on a 353-service estate? We're at the awkward middle and some of these patterns feel like they need a dedicated ops person to run properly.
Uchenna Ibe
August 31, 2026
Refreshing to read this framed for our market rather than lifted from a Silicon Valley playbook. Specifically the "Turning the Incident Into Learning, Not Blame" piece — CBN have opinions, and that changes the design constraints in ways the US-centric literature never touches.
Uchenna Anyanwu
August 16, 2026
Fractional CTO — done 3 fintech engagements in the last 4 years. Reading this after a rough sprint — the "Communicating Clearly While Everything Is On Fire" bit lands, because I spent this week in the middle of hiring in Lagos market on my engagements. Honestly the framing would have saved me at least a Slack DM I regret sending.
Sadiq Danjuma
August 14, 2026
The framing on "The First Five Minutes Set the Tone" alone is worth the read.
Chuka Umeh
August 8, 2026
Fractional CTO — done 3 fintech engagements in the last 5 years. The "Why Calm Is an Engineering Skill, Not a Personality Trait" is what I wish someone had spelled out for me in my first year as a fractional CTO — learned it the hard way when a skip-level asked a question I could not answer.
Musa Mohammed
August 3, 2026
One more thing worth naming: this only holds when the manager themselves has been coached this way. Copy-pasting the technique without the underlying model tends to fall flat.
Adeola Abiola
August 2, 2026
Refreshing to read this framed for our market rather than lifted from a Silicon Valley playbook. Specifically the "The First Five Minutes Set the Tone" piece — NIBSS have opinions, and that changes the design constraints in ways the US-centric literature never touches.
Akosua Amankwah
August 1, 2026
Does the "Communicating Clearly While Everything Is On Fire" still hold on a 346-service estate? We're at the awkward middle and some of these patterns feel like they need a dedicated ops person to run properly.
Kwesi Amankwah
July 27, 2026
This is why I keep coming back to this blog.
Yemisi Ogunbanjo
July 27, 2026
Would add: the incentives inside the calibration room matter as much as the actual conversation with the engineer.
