Incident management
Things go wrong. It's how we deal with it that counts. The following documentation outlines how to behave in such situations and the processes that should be followed.
General guidance
Understand before acting
Reacting to issues can often make things worse. The first thing you should do is stop, understand the issue, document your findings and discuss them with others.
Seek consensus
You should speak with stakeholders and others supporting on the incident to agree next steps to resolve the incident and who should remain involved.
Be communicative
As you are working through findings and executing on the resolution the decisions, findings, actions and executions should be documented as you progress.
Provide a retrospective
After an incident is resolved the team should be involved in discussing preventative measures, how the incident was handled and any process changes that can be made to help us in the future.
Raising an incident
If you notice or are made aware of a problem, post to the #incidents channel on Slack with a traffic light-coded emoji based on the severity of the incident:
| Level | Definition | Examples |
|---|---|---|
| 🔴 (red) | Critical incident | Customer UX severely degraded, website unavailable, payments not working |
| 🟠 (amber) | Partial degradation | Customer UX significantly degraded, non-critical service outage |
| 🟢 (green) | Resolved | A previous incident is now resolved, fixed or downgraded to a low or non-severe issue |
- Include as much info and context as you can so that people understand the issue
- However, if the reporting of the issue is time sensitive, more detail can be added to the thread as the situation evolves
- Tag any relevant people or teams
- Be clear about the current status and action that is required
- Are you actively investigating the issue?
- Do you need input from somebody else?
Examples
🔴 Emails are not sending
Our emails have stopped sending in the last 48 hours, due to an error that Klaviyo is reporting. @[crm-lead] reached out to support to understand what is causing the problem. We'll update on a thread here with news - but please be aware that we may need dev support.
cc. @[developer] @[product-manager] @[marketing-manager]
🟠 The url bbcmaestro.com is timing out instead of redirecting to www.bbcmaestro.com
This will mostly affect users on browsers like Firefox, that don’t automatically insert thewww.prefix to entered urls that don’t have it already. Cloudflare and Heroku are not reporting any incidents on their status pages.
The engineering team is currently investigating – no action currently required from anyone else, this post is for visibility.
Out of hours proposal
We currently have no official channels of escalation for incidents that happen outside of working hours. What follows is a suggestion.
The same process as raising an incident should be followed, but since the team are not likely to be online, additional steps need to be taken:
- Product, engineers and QA belong to a WhatsApp group called BBC Maestro: First Response
- If a critical incident happens, a member posts the incident in the WhatsApp channel and ensures that the post exists in
- Engineers (if available) will respond on WhatsApp to say that they are looking in to the issue
NOTE
This does not consitute an "on call rota". Engineers are not expected to be available, with access to their work laptop, during unsociable hours. The idea of this is to be better able to reach the engineering team in the case that anyone might be available.
The company respects the fact that WhatsApp is primarily a personal channel. For this reason:
- Avoid posting more than necessary in the WhatsApp channel
- Move the conversation to Slack as soon as possible
- It's ok to mute notifications on this channel on occasions such as annual leave
Any time spent outside of normal working hours to investigate and respond to an incident may be taken back as time in lieu. Report any spent time to the Head of Engineering to have your time added to the HR system.
Retrospective
Once an incident has been resolved, a retrospective should be undertaken. This may result in a post-mortem.
- Small post-mortems with any FYIs should be raised in product x engineering retros.
- Post-mortems with significant recommendations should have their own dedicated review meeting