Safeguard
Best Practices

Deciding a Security Incident's Severity While You Still Know Nothing

Outage rubrics rate impact, which you can measure while it happens. A security incident's impact is unknown at the moment you must classify it, and often stays unknown for days. Rate the observation instead.

Aisha Rahman
Security Analyst
6 min read

Something looks wrong. An unfamiliar login, an alert nobody recognises, a customer describing behaviour that should not be possible. You have to decide, now, whether this is an incident and how serious, on almost no information.

Severity rubrics do not help, because they were written for outages. An outage's impact is measurable while it is happening: requests failing, users affected, revenue per minute. A security incident's impact is unknown at the moment you must classify it, and often stays unknown for days.

This post is how to make that call. For whoever is on call at a company too small to have a separate security rotation.

Classify by what you can observe

The mistake is trying to rate the consequence, which you cannot know yet. Rate the observation instead, because that is available immediately and it maps well enough onto urgency.

Highest: confirmed unauthorised access to production or to customer data. Someone or something is in, or was in. You do not need to know what they did.

High: credible evidence of compromise without confirmation. A credential known to be exposed, an unexplained administrative action, an alert with a plausible story behind it. Also: any report from an external party, because they usually have information you do not.

Medium: a vulnerability with no evidence of exploitation, where exploitation would be serious. This is most of what you handle, and the thing that distinguishes it from the band above is that nothing suggests it has happened.

Low: an issue with limited impact and no urgency, which is a ticket rather than an incident.

The useful property of this scale is that every level can be assigned from information you have in the first five minutes, without waiting for an investigation.

Declare early, downgrade freely

The bias at small companies is to under-declare, for understandable reasons. Waking colleagues is socially expensive, and the reasonable-sounding thought is to look into it a bit first.

Two things make that expensive. The first hour matters more than any other hour, because logs roll off, sessions expire, and attackers finish what they came for. And a single person investigating alone, at night, without a second opinion, makes worse decisions than two people would.

So: declare on suspicion, and make downgrading explicitly normal. If your culture treats a downgraded incident as a false alarm somebody should be embarrassed about, people will hesitate next time, and the hesitation is what costs you. Say out loud, repeatedly, that declaring and standing down is a good outcome.

One role that matters

Large organisations have incident commanders, scribes, communications leads and subject matter experts. At twenty people you have three engineers awake.

The only role worth formalising is that one person is running it, and everyone else is helping. That person does not investigate. They decide what happens next, track what has been tried, and make the calls about who else to wake and when to notify.

The failure mode without it is four people investigating four theories in parallel, nobody writing anything down, and the same log queried three times. Naming a coordinator costs one sentence and it is the highest-value thing you will do in the first ten minutes.

Write it down as you go

A running log with timestamps, in whatever channel you are already using. What was observed, what was checked, what was found, what was decided and by whom.

It serves three purposes and you need all of them: it keeps the team aligned while people join and leave, it is the basis of your notification obligations if this turns out to involve personal data, and it is the only accurate account you will have afterwards. Memory of an incident is reliably wrong about timing.

The practical trick is to have the coordinator do it, since they are the one not investigating.

Decisions that need making early

Four things worth deciding deliberately rather than by default:

Do we cut access? Disabling accounts or revoking sessions may stop the activity and may also destroy the evidence and alert the attacker. There is no universal right answer, and the wrong version is doing it without anyone deciding.

Do we preserve state? Snapshot instances, export logs, capture memory if you are able. Cloud logs have retention limits and the interesting window is often just outside them.

Who needs to know, and when? Leadership, legal, privacy, the customers, and possibly a regulator. If personal data may be involved, the privacy owner joins now rather than at the end, because their clock has already started.

Do we bring in help? External incident response is expensive and slow to arrange under pressure. Knowing in advance who you would call, and having their contact details somewhere reachable when your systems are not, converts a two-day procurement into a phone call.

Prepare the three things that must exist beforehand

None of this works if it has to be improvised.

A way to reach people out of hours that does not depend on the systems that might be compromised. A phone tree is unfashionable and it works.

Access to logs by someone who is awake. If only one engineer can query the audit logs, your incident response has a single point of failure who may be asleep or on a plane.

A short declaration procedure, one page, stored offline. Who to call, how to declare, where the log goes, what to preserve.

The concession

There is a real cost to declaring early and often. Each declaration consumes people, interrupts work, and a team that gets woken repeatedly for things that turn out to be nothing will start resisting, quietly, by taking longer to escalate.

So the calibration matters as much as the bias. Review the declarations monthly, and if most are standing down at the lowest level, tighten what triggers a page rather than asking people to be more judicious in the moment. Judgement under pressure at 03:00 is the wrong place to absorb a tuning problem.

The implication

The severity decision is made with the least information you will ever have about the incident, and it determines how fast everything else happens.

Which is why it should be based on what you can see rather than what you fear, and why the cheapest improvement available to most teams is making it socially free to declare something that turns out to be nothing.

Never miss an update

Weekly insights on software supply chain security, delivered to your inbox.

Self-healing security runs on Safeguard.

Your first fix PR is minutes away.

No sales call required, even your agent can complete the purchase over MCP.