- Time
- 11 minutes
- You work in
- Channels
- Connect first
Slack
Before you start
- Slack connected
- An employee you have hired
During an incident the engineering work is usually well handled and the communication is usually not. Not from indifference: the people who know what is happening are the people fixing it, and writing an update costs them exactly the attention you want on the problem.
The result is silence, and silence is what turns an outage into a trust problem.
Decide the cadence before you need it
Agree it now, while nothing is broken. Something like: an update within fifteen minutes of confirming an incident, then every thirty minutes whether or not there is news, and one when it is resolved.
The "whether or not there is news" half is the part people skip and the part that matters. "Still working, no new information" is a real update.
Keep one message current rather than posting many
During an incident, post a status update to the incident channel and keep that same message updated rather than posting a new one each time. Each update: what is affected, what we know, what we are doing, and when we will next say something. Say "no new information" when there is none. Never speculate about cause, never estimate a fix time unless a human has given you one, and never say anything is resolved until a human says so. If thirty minutes pass with no update from anyone, post that we are still working and nothing has changed.
"Never estimate a fix time unless a human gave you one" is not caution, it is the whole credibility of the channel. An estimate that slips does more damage than silence, because it converts a technical problem into a broken promise. The honest version, "we do not know yet and will say more in thirty minutes", is always available and always survives.
The dead-man's update is the real feature
The instruction to post when nobody else has is what fixes the actual failure. Incidents go quiet exactly when they are hardest, which is exactly when customers most want to know somebody is still there.
Automating the "we are still here" message costs the responders nothing and buys most of the goodwill.
Write the resolution update yourself
Everything above can run unattended. The final message should not. It says the problem is over, and that assertion should come from a person who has checked.
What good looks like
Customers stop asking whether you know. The channel has a heartbeat during the worst hour, and the people fixing it were never asked to write anything.
The measure is not update count. It is whether the longest gap between updates during your next incident is under thirty minutes.