Skip to content

45 β€” Incident Response Playbooks for Small Teams

Level: Intermediate Β· Time: ~22 min Β· Prerequisites: Lesson 44 β€” Detecting Intrusion with Open-Source Tools


Why this matters

Lesson 24 gave you the discipline of incident response β€” prepare, detect, contain, eradicate, recover, learn. This lesson gives you the artefacts that make that discipline usable on the night it is needed, by a tired person who has never done it before. A small organisation has no shift rota, no duty manager and nobody watching a screen: the person who notices the problem is usually not the person who is allowed to decide what happens next. A playbook is the bridge between those two people β€” a short, decision-ready, rehearsed document that tells whoever is holding the phone what to do next, in order. Its value is not the paper; it is that somebody read it before the day it mattered.


The mental model: what a playbook is, and what it is not

A playbook is A playbook is not
one to two pages, per incident type a 60-page incident response plan nobody has opened
a decision ladder: if this, do that, then call them a description of your tools and their features
written for the least experienced person on the team written for the person who wrote it
short lines, active verbs, answers not essays background reading, definitions, or theory
a list of thresholds that trigger outside help a promise that you can handle everything
readable when email, chat and the wiki are down a page hosted only on the systems under attack

The test of a good playbook is brutal and simple: hand it to somebody who was not in the room, at an hour when they are not at their best, and see whether they can act. If they have to make a judgement call that the document does not resolve, that is a gap to fix in rehearsal, not during the incident.

Why a small team needs this more than a large one

Small team Large team
nobody is on shift; detection waits for someone to notice a rota watches alerts continuously
the same three people do everything, and one may be on holiday depth of cover, named roles, handovers
the person who notices is often not the decision-maker escalation paths are defined and practised
there is no forensic capability and no retained specialist a retainer and in-house forensics
an outage is felt by customers within the hour buffers, redundancy, communications staff

The conclusion is uncomfortable but true: a small team is more fragile per incident and therefore needs more preparation per head, not less. You are not buying a security operations centre. You are buying the thirty minutes of good decisions that a documented, rehearsed process delivers.


The structure every playbook follows

Use the same eight headings in every playbook. Uniformity is the point: under stress, nobody reads a table of contents.

Heading What goes in it
Trigger the single observation that makes you open this page β€” written so it is unambiguous
First fifteen minutes the three to five actions that must happen before anything else, in order
Scope questions what you must be able to answer to know how bad this is: which accounts, which systems, which data
Containment decision the choices available, who may authorise each, and what each one costs
Communication who must be told, in what order, within what deadline, by whom
Evidence to preserve the specific artefacts to collect before you start changing things (see Lesson 25 β€” Digital Forensics Basics)
Recovery the sequence back to normal, including what must be verified before restoring
Follow-up the review, the dated actions, the owner β€” within two weeks, while memory is fresh

Two rules run through all of it. Preserve before you act, because isolation, reboots and cleanups destroy what you need to understand the incident. Record as you go, because an incident log β€” timestamp, action, who decided, why β€” cannot be reconstructed afterwards and is what the insurer, the regulator and the customer conversation are built from.

[!WARNING] Isolate, do not power off. Pulling the network cable preserves memory, running processes and the attacker's current sessions. Pressing the power button destroys all three, and may trigger the attacker's own cleanup if their tooling notices. The same applies to a "quick reboot to see if it comes back".


The playbooks

Each block below is deliberately compact: it is the shape of a one-page document you finish yourself, with your names, your systems and your numbers in it.

1. Ransomware. Trigger: files will not open, a ransom note appears, or a share suddenly fills with renamed files. - First 15: isolate affected hosts from the network rather than shutting them down; disconnect any mapped drives; stop the backup job that targets them; start the incident log; call the incident lead. - Scope: which systems are encrypted, which are still clean, which accounts are involved, what was the earliest suspicious event, is the backup infrastructure reachable from the affected network. - Contain: disable or restrict the compromised account; block the C2 destinations at the firewall and DNS; segment the affected network; do not let anyone "helpfully" clean a machine. - Communicate: incident lead, decision authority, insurer, legal counsel, and leadership β€” the insurance policy usually requires prompt notification and rarely accepts "we were busy". - Evidence: volatile data first (memory, process list, connections), then ransom note copies, the files themselves, and the logs that survive. - Recover: verify the backup exists, is offline or immutable, and was not deleted or encrypted; find and close the entry path before restoring; rebuild rather than clean; rotate every credential the affected hosts could reach. - Do not: pay without legal and law-enforcement advice, assume nothing was exfiltrated, or restore while the entry path is open.

2. Business email compromise and invoice fraud. Trigger: a customer says they received an invoice you did not send, a supplier's bank details have "changed", or somebody cannot log in to their mailbox. - First 15: do not reply to the email; take the suspicious message out of the recipient's mailbox if it is still there; start the log; call the finance lead before any payment is released. - Scope: whose mailbox is compromised, which messages were sent, which invoices were altered, which payment is pending, how far back the access goes. - Contain: the callback rule is the control β€” verify any change of bank details by telephoning a number you already had, never one supplied in the message, and never by replying; block or freeze the payment; reset the account and revoke its sessions and tokens; check the mailbox rules (next playbook). - Communicate: finance, the customer or supplier affected, the bank immediately if money has moved (a recall may still be possible), the insurer, and counsel if a payment is disputed. - Evidence: the fraudulent message with full headers, the sent items, the mailbox rule that hid replies, the forwarding configuration, the bank details in the message. - Follow-up: remind every customer-facing team of the callback rule; check who else received the thread; review how a bank-detail change is approved internally.

3. A phishing email reported, and a phishing link clicked. Trigger: a user forwards something suspicious, or a user says they entered their password. - Reported, not clicked: thank the reporter, examine the message headers and links, remove the message from other mailboxes, check whether anyone else received it, and treat it as an awareness win. - Clicked, but nothing entered: check the endpoint's browsing and proxy logs, and block the domain; a link alone is often harmless. - Credentials entered: treat it as the start of an incident, not a training matter. Reset the password, revoke active sessions and refresh tokens, verify the registered second factor is still the user's own, check mailbox rules and forwarding, check for new application consents, and check sign-in logs for a session from an unusual location, device or user agent. - First 15: capture the URL and any downloaded file; ask the user exactly what they typed and when; check MFA prompts the user did not initiate. - Procedure: every one of these steps should be a checklist with a name against it, because the two most commonly missed are the mailbox rules and the newly registered second factor. - Culture: the reporting must be safe and unpunished, or the next person says nothing. People who report quickly are the cheapest detection you own.

[!TIP] The most valuable thing a small team can do after a reported phish is to say thank you in public. The alternative β€” a quiet investigation, then a reminder email about carelessness β€” teaches the organisation that reporting is a risk to your reputation.

4. A compromised account sending spam to contacts. Trigger: colleagues, customers or a partner report mail they did not expect from a legitimate address. - First: the account is still authenticated somewhere; get into the account's own audit trail before changing anything. - Check in this order: inbox rules that move replies away (especially rules matching a keyword or a bank name), mailbox-level forwarding and "connected accounts", delegate and mailbox permissions, application (OAuth) consents, SMTP authentication settings, and the sent-items folder against the reported messages. - Contain: change the password, revoke all sessions and tokens (a "sign out everywhere" is not the same as revoking refresh tokens β€” do both), remove and re-register the second factor, delete the rules, remove the consents, then force a sign-out of the mail client on all devices and reauthenticate from scratch. - Scope: the same rules, consents and forwarding patterns may exist across the whole tenant β€” check every mailbox for the same rule pattern, because if one account was compromised this way, others often were too. - Preserve: export the mailbox audit log and the rules before deletion, and record what the attacker sent, to whom, and when. - Communicate: tell the contacts who received the mail, in plain language, with what they should do; this is reputational work as much as technical work.

5. A lost or stolen laptop or phone. Trigger: a device is missing, or left in a taxi. - First 15: report it, start the log, and check whether disk encryption was enabled β€” that single fact decides how serious this is. - Contain: remote lock, then remote wipe when you are confident it will not come back; block its access to email and file shares; check the last sign-ins from that device in the identity logs; if a token or password was stored unencrypted, assume it is compromised. - Rotate: any credential stored outside the managed vault, plus any session for the account used on the device. - The saving control: with disk encryption (BitLocker, FileVault, LUKS) and a screen lock, a stolen laptop is a hardware loss. Without it, it is a possible data breach β€” and that is the question you must then answer. - Data-protection question: was personal data on it, and was it encrypted? If personal data was present and unencrypted, or if you cannot rule that out, treat it as a reportable breach question and take advice on the notification clock (Lesson 24 covers the 72-hour principle under the GDPR). - Follow-up: fix the policy, not just this device β€” full-disk encryption enforced on every endpoint, remote wipe enabled, and nothing sensitive stored outside the vault.

6. Account takeover, or suspected credential compromise. Trigger: you find credentials in an external dump, a sign-in alert fires, or a user's account behaves unusually. - Where the session lives: the attacker does not need to log in again once they hold a valid session, so look in the identity provider's sign-in logs for the session that was created β€” device, location, IP and user agent β€” not just the password change you were expecting. - First 15: reset the password, revoke sessions and refresh tokens, remove any new second-factor method, and add a temporary strong factor (ideally a phishing-resistant one). - Check: mailbox rules, forwarding, application consents, new devices, new API keys, new mail-forwarding addresses, and any credential the account had access to, because takeovers are often the first step towards something else. - Scope: if the account had administrative rights, the blast radius is everything it could reach β€” apply the rotation discipline from Lesson 24 rather than only fixing the login. - Preserve: export the sign-in log and the audit trail before you change things, and record the timeline.

7. A departing employee, or a current one, suspected of data theft. Trigger: bulk downloads, mailbox forwarding to a personal address, USB activity, or a resignation with a suspicious tail of access. - Preserve before confronting. The moment somebody knows they are suspected, evidence disappears and the tone changes. Collect logs, exports and access records first, with counsel involved, and document the chain of custody. - Legal constraints are real. Monitoring staff is constrained by employment law and data-protection rules in most European countries: usually work-related, proportionate, and with employees informed that monitoring may occur. This is not legal advice β€” involve counsel and your data-protection adviser before you watch anybody's activity, not after. - Scope: what they accessed, what they exported, where it went, whether it is still reachable, whether a third party (a new employer, a competitor) received it. - Contain: revoke access in the correct order (accounts, then sessions, then tokens, then physical access, then any credential they knew that others share), and disable forwarding and personal-device access. - Communicate: leadership and counsel only; keep the circle small, and let counsel run the conversation with the individual.

8. Website defacement or web application compromise. Trigger: the site shows content you did not publish, or a scanner finds a file you did not upload. - First 15: decide between taking it offline and patching in place β€” taking it offline is faster and safer, and is usually the right first call; preserve a copy of the current state before you change anything. - Look for web shells: small scripts in upload directories, recently modified files, files with names matching your own conventions, and content management system plugins, themes or modules you did not install. - Review access logs: look for POST requests to unusual paths, requests with parameters that look like commands, uploads, and the first request from the suspicious source β€” that is usually the entry point. - Recover: rebuild from a clean known-good copy, patch the actual cause rather than the symptom, rotate hosting, database, CMS and file-transfer credentials, and check scheduled tasks, cron entries and database users for persistence. - Verify: assume the attacker kept a way back in; a site cleaned without finding the entry point will be compromised again, usually within weeks. - Follow-up: enable a web application firewall where appropriate, and keep the platform patched (Lesson 10 β€” Web Application Attacks).

9. A supplier or service provider breach that affects you. Trigger: a vendor announces an incident, or you read it in the news before they tell you. - Ask them, in writing: which of our data or systems are involved; was our data exfiltrated and what exactly; when did it start and when did you detect it; how did they get in; what have you contained; what credential or access we gave you is now in their hands; who is our contact; and what do you recommend we do. - Assume, until proven otherwise: every credential they held for your environment is compromised, every integration they had can reach further than you think, and the timeline they give you will move. - Do: rotate credentials and integration keys you shared with them, review their access to your systems with least privilege in mind, check your own logs for activity from their integration, and record everything against the contract's breach-notification clause. - Contract matters: notification obligations, audit rights and the right to terminate are only useful if they were agreed before the incident (Lesson 14 β€” Insider Threats and Supply Chain Attacks).

10. Service unavailability, or a denial-of-service attack. Trigger: the site, mail or an application stops responding. - First: is it an outage or an attack? Check the provider's status page, your own monitoring history, whether all users are affected or only some, whether one resource saturated or everything did, and whether traffic volume is abnormal. An outage is fixed by finding the failing component; an attack is mitigated by absorbing or filtering traffic. - Look for: a traffic spike from many sources, a spike on one endpoint or one login page, a pathological query or a slow database query pattern, and the difference between volumetric floods and application-layer requests that look almost legitimate. - What your provider absorbs: a hosting or content delivery provider absorbs volumetric, network-level floods as part of the service. Application-layer attacks against your own code usually arrive as normal requests, and are yours to handle with rate limiting, caching, request validation and authentication. - Contain: enable provider protection, rate limit the affected endpoint, put the site behind the content delivery network or filter, and contact the provider early with evidence. - Do not: assume it will stop on its own, change firewall rules under pressure without recording them, or let the mitigation block your real customers for a week afterwards. - Follow-up: agree a response expectation with your provider in writing now, while nothing is on fire.


When to stop improvising and get professional help

Escalate immediately, rather than after a day of trying, when any of the following is true. This list is the most valuable part of the whole lesson, because it catches the moment where confident amateurs do the most damage.

Escalate when Why
Confirmed or strongly suspected attacker access you cannot scope it alone, and every hour of quiet access widens it
Personal data is or may be involved notification duties have deadlines; the analysis needs advice
Ransomware, or any destructive activity recovery sequencing gets expensive fast when it is wrong
Legal exposure β€” fraud, insider theft, a disputed payment, a regulator preservation, privilege and employment rules apply
A system nobody understands β€” an undocumented application, an old server, a consultant's legacy tool you cannot contain what you cannot describe
The decision is above your pay grade β€” a shutdown, a payment, a public statement these are business decisions, not technical ones

Retaining a professional is not an admission of failure; it is the same decision as calling a lawyer or an accountant. What you should have prepared in advance is who you would call, under what arrangement, and who is allowed to authorise the spend at two in the morning.


The machinery to prepare before you need it

Artefact What it must contain Why it fails in practice
Contact sheet, offline as well as online incident lead, deputy, decision authority, insurer and policy number, legal counsel, data-protection adviser, hosting provider, key suppliers, out-of-hours numbers it lives in the wiki or the mail system you can no longer reach
Decision authority list who may shut down a service, spend money without a further round of approvals, contact the regulator, and speak to customers or the press whoever answers the phone ends up deciding, badly
A single incident log one file or channel, one line per action: timestamp (in UTC), action, decision-maker, reason three people keep three sets of notes and nobody can reconstruct the day
A named owner per playbook one person responsible for it being correct, current and rehearsed a document with no owner is a document nobody updates
A copy that survives the outage printed page, offline export, phone photo of the current version the playbook is on the file share that just got encrypted

Keep the contact sheet on paper in the office and in each senior person's phone, and print the two playbooks you consider most likely. It costs nothing and it is the difference between a managed first hour and a lost one.


Rehearsal: the sixty-minute tabletop

You need no tooling, no budget and no external facilitator. Book an hour, invite four to six people from outside security (finance, operations, a manager, whoever answers the phone), and run this:

  1. Pick one playbook. Ransomware or business email compromise are the best first choices: they are likely and they expose ownership gaps quickly.
  2. Read the trigger aloud, exactly as written, and stop. Do not explain the scenario further β€” a real incident does not arrive with context.
  3. Walk the decision points. At each one ask: who decides, what do they need to know, and how would they get it right now?
  4. Write down where nobody knew. Two of those per session is a good outcome; more than five means you have found the real work.
  5. Do not rehearse the happy path. Inject one complication: the incident lead is unreachable, the backup server is on the affected network, the customer has already paid.
  6. Close with three dated actions and a name against each, then repeat the same playbook in three months β€” not a different one.

How often: at least twice a year for each playbook, once for a new playbook before it goes live, and always within a month after a real incident, using what actually happened instead of a scenario. Record who attended and what came out of it, so the rehearsal itself is evidence that your response is practised.


Keeping playbooks alive

  • Review after every real incident. The incident tells you which step was wrong, missing or unreachable. Update it while the memory is fresh, not at the next annual review.
  • Review on a fixed cycle. Twice a year, plus whenever a named person leaves, a system changes, or a supplier changes. An out-of-date contact sheet is worse than none, because it is trusted.
  • One owner per playbook, named in the document, with the review date written on the page.
  • Store them where they can be reached when systems are down β€” a printed copy, an offline export, a copy in someone's phone. Test this by opening it with the network unplugged.
  • Keep them short. If a playbook grows past two pages, it has become a plan, and nobody reads a plan at 02:00.

A closing thought, and the reason this lesson exists separately from Lesson 24: the value of a playbook is not the document. It is that somebody read it, thought about it and decided what they would do, on a normal Tuesday afternoon, in advance. The document is just the receipt.


Attack it / Defend it

The attack How it works The control that stops it
Ransomware at 02:00 nobody is awake and the first responder has never done this a rehearsed one-page playbook with a named decision-maker
Fake bank-detail change a plausible email changes the payment destination the callback rule: verify by phone on a number you already held
Compromised mailbox hiding replies a rule moves or deletes incoming mail so the victim does not notice a checklist that inspects rules, forwarding and consents first
Stolen unencrypted laptop hardware theft becomes a data breach enforced full-disk encryption, remote wipe, nothing sensitive stored outside the vault
Attacker re-enters after recovery a rule, token or credential nobody revoked revoke sessions and tokens, remove consents, check the whole tenant, not one mailbox
Insider theft during notice period data leaves before anyone is suspicious preserve before confronting, logging and DLP, least privilege, counsel involved early
Web shell in an upload folder an unnoticed file keeps a way back in file integrity checks, review of modified files, rebuild rather than clean
Supplier breach cascading to you their compromise holds your credentials supplier questionnaire and clauses agreed before the incident, prompt rotation
Phone-based pressure to pay deadlines and threats aimed at whoever answers a pre-agreed decision authority and an escalation threshold, not improvisation
The unreachable contact the incident lead is on a plane and the insurer number is in email an offline contact sheet, tested with the network unplugged

Key takeaways

  • A playbook is a decision ladder, not a plan. One to two pages per incident type, written for the least experienced person who might have to use it.
  • A small team needs playbooks more than a large one, because there is no rota, no shift handover and no depth of cover β€” the person who notices is rarely the person who may decide.
  • Isolate, do not power off, and preserve before you act. The evidence you destroy in the first fifteen minutes is the evidence you need in the second week.
  • The callback rule is the cheapest control in the BEC playbook: verify any change of bank details by telephone on a number you already held, never one from the message.
  • Know your escalation threshold and your decision authority in advance. Confirmed access, personal data, ransomware, legal exposure, or a system nobody understands means you stop improvising and call for help.
  • Rehearse with a sixty-minute tabletop and no tooling. Two findings per session β€” places where nobody knew who decides β€” is a good evening's work.

Check yourself

  1. Why is isolating a compromised host better than switching it off, and which three things do you lose by powering it down?
  2. A user reports that they entered their password on a link sent by "their bank". List, in order, the six things you check, and say which two are most often forgotten.
  3. A supplier announces a breach. What five questions do you ask them in writing, and what do you assume about the credentials they held for you?
  4. Name the five conditions that should trigger an immediate call for professional help rather than another day of trying.
  5. Your contact sheet is in the company wiki. Why is that a design flaw, and what would you do about it this week?

Next

Lesson 46 β€” Risk, Governance and Compliance