Skip to content

Incident management ​

Things go wrong. It's how we deal with it that counts. The following documentation outlines how to behave in such situations and the processes that should be followed.

General guidance ​

Understand before acting ​

Reacting to issues can often make things worse. The first thing you should do is stop, understand the issue, document your findings and discuss them with others.

Seek consensus ​

You should speak with stakeholders and others supporting on the incident to agree next steps to resolve the incident and who should remain involved.

Be communicative ​

As you are working through findings and executing on the resolution the decisions, findings, actions and executions should be documented as you progress.

Provide a retrospective ​

After an incident is resolved the team should be involved in discussing preventative measures, how the incident was handled and any process changes that can be made to help us in the future.

Raising an incident ​

If you notice or are made aware of a problem, post to the #incidents channel on Slack with a traffic light-coded emoji based on the severity of the incident:

LevelDefinitionExamples
🔴 (red)Critical incidentCustomer UX severely degraded, website unavailable, payments not working
🟠 (amber)Partial degradationCustomer UX significantly degraded, non-critical service outage
🟢 (green)ResolvedA previous incident is now resolved, fixed or downgraded to a low or non-severe issue
  • Include as much info and context as you can so that people understand the issue
  • However, if the reporting of the issue is time-sensitive, more detail can be added to the thread as the situation evolves
  • Tag any relevant people or teams
  • Be clear about the current status and action that is required
    • Are you actively investigating the issue?
    • Do you need input from somebody else?

Examples ​

🔴 Emails are not sending

Our emails have stopped sending in the last 48 hours, due to an error that Klaviyo is reporting. @[crm-lead] reached out to support to understand what is causing the problem. We'll update on a thread here with news - but please be aware that we may need dev support.

cc. @[developer] @[product-manager] @[marketing-manager]

🟠 The url bbcmaestro.com is timing out instead of redirecting to www.bbcmaestro.com

This will mostly affect users on browsers like Firefox, that don’t automatically insert the www. prefix to entered urls that don’t have it already. Cloudflare and Heroku are not reporting any incidents on their status pages.

The engineering team is currently investigating – no action currently required from anyone else, this post is for visibility.

Out of hours ​

Since the team are not likely to be online, one additional step needs to be taken to raise an incident outside of working hours.

The same process as raising an incident should be followed, but the string [ALERT] should be included in the message body. The presence of [ALERT] will trigger an SMS to be sent to recipients on the "first responders" list.

The company respects the fact that SMS on personal phones is a personal channel and should not be abused. For this reason:

  • Incident reporters. Only include [ALERT] when the situation can't wait until the following working day
  • First responders. It's OK to mute notifications from the sender on occasions such as annual leave

Any time spent outside of normal working hours to investigate and respond to an incident may be taken back as 2x time in lieu. Report any spent time to the Head of Engineering to have this added to the HR system.

NOTE

This out of hours incident policy does not constitute an "on-call rota". Engineers are not expected to be available, with access to their work laptop, during unsociable hours. The aim is to make it easier to reach the engineering team in case anyone might be available.

First responders list ​

The first responder list includes:

  • Engineering team (including QA)
  • Product team

Example ​

[ALERT] 🔴 Sale discount code is not applying correctly

The discount code THIS_IS_A_TEST – which is only live for this weekend – is not auto-applying at checkout. The obvious checks have been done and this is still proving to be a live issue, affecting customers and time/revenue-sensitive due to the limited offer period.

Responding to an incident ​

  • React to the incident report with the 👀 (:eyes:) emoji, so that others know you're looking into the issue
  • Comment on the thread to manage others' expectations. If you're still investigating, let people know. This is particularly important when dealing with an out of hours incident so that we don't unnecessarily duplicate effort.
  • As and when the incident is resolved or downgraded, post a new message with the appropriate status emoji per raising an incident

Retrospective ​

Once an incident has been resolved, a retrospective should be undertaken. This may result in a post-mortem.

  • Small post-mortems with any FYIs should be raised in product x engineering retros.
  • Post-mortems with significant recommendations should have their own dedicated review meeting