Skip to content

Read the alert you already have

The Response program trains operators to trust or explain a detection, run a playbook under pressure, and prioritize when everything changes at once.

By ALDBRN

Program Response · Canis Major
HISTORIAN TRENDWRITE OUTSIDE WINDOWENG WORKSTATION ADDR · NO WORK ORDERSAME ASSETSAME WINDOWWHO WAS LOGGED ONTRUST OR EXPLAINSTAND DOWNACTOPERATOR (SHIFT)SHIFT LEADCONTROL ENGINEERIT SECURITYT+0T+30 MIN
A historian trend breaks its pattern, and the first thirty minutes after it: the questions that separate an explanation from a decision to act, and the escalation path from the operator on shift to the control engineer and IT security.

Read the detection you already have

Most control systems already produce signal before anyone calls it an incident. A historian trend breaks its usual pattern. A PLC changes from run to program mode without a work order behind it. A firewall at the IT-OT boundary logs a connection it was built to deny. Where a site runs one, an OT monitoring tool raises a flag of its own.

Incident Detection and Response in OT starts with a plain question: would you notice, and who acts first? Most operators can answer the first half after a shift or two on the floor. Fewer can say exactly what each record supports and what it does not.

A firewall deny shows that a connection was attempted, not that it succeeded. A mode change shows that something wrote to the controller, not who did it or why. The program trains operators to read a detection for what it actually proves, because an alert nobody trusts gets muted, and a muted alert is worse than no alert at all.

The operator on shift is the first responder

Plans name an incident response team, a manager, an on-call security contact. In practice, the first responder is whoever is watching the process when the value moves. That is almost always the operator on shift, whether the plan says so or not.

Take a typical case: an engineering workstation address that has no business on the network after hours writes to a PLC outside the scheduled maintenance window. The historian shows it. The monitoring tool, if there is one, flags it. Nobody upstream has seen it yet.

The operator’s first job is not to decide what happened. It is to ask what would tell two explanations apart: is this the same asset the maintenance note refers to, does the timing actually overlap, and who was logged on to make the change. Those three questions separate an observation from an explanation.

What happens next is the decision point the program drills. The operator either trusts the alert and acts, calling the shift lead and starting an isolation conversation, or finds the explanation and stands down. Both are legitimate outcomes. Guessing is not.

Isolate or keep running is a process decision

Once the first questions come back unanswered, or point the wrong way, isolating the segment looks like the obvious move. Whether it is safe is a process question before it is a security one.

A packaging line mid-batch cannot be cut off without a product loss the plant manager will ask about by name. A boiler running under manual water level watch needs the operator who understands that constraint, not just the one with firewall access. The program trains operators to weigh that call themselves, because they are the ones who know what the line can actually tolerate.

A working playbook names a safe state for the process, not only for the network: when to shift to manual operation, when isolating a segment is worth the production it costs, and when the fastest safe move is calling the vendor instead of waiting on an internal escalation. None of this authorizes anyone to change a live system outside the site’s own change process.

Who signed for this risk

Risk, Governance and Compliance asks a blunter question than it sounds: which risks does the business accept, and who signed for them? IEC 62443 zones and conduits, NIST SP 800-82, and NIS2 or a sector regulator where one applies, set out what a site is expected to do. None of that tells an operator, in the moment, who is allowed to make which call.

That is what a risk register is supposed to answer, if anyone can find it. A vendor’s remote access path, accepted with a compensating control. A segment that stayed flat because separating it would have cost a shutdown nobody could schedule. Somewhere, someone signed off on that exposure and on what happens if it is ever used against the plant.

The program trains operators to find that signature, or the gap where none exists, before the incident, not during it. Knowing whose authority backs a decision is what turns a fast call into a defensible one.

Prioritization under fire

Incidents rarely stay tidy. Prioritization Under Fire asks the question directly: when everything changes at once, what comes first? A real morning can bring alarms from two areas at once, a vendor calling about the very system now in question, an IT request to isolate a segment that also carries a safety system, a line that a customer contract says must keep running, and a regulator’s notification clock already counting down.

None of those demands announces its own importance. The program trains operators to rank them by consequence to the process and to people, not by who called first or who has the most seniority. Life safety and process safety outrank containment. Containment outranks convenience, a vendor’s schedule, or anyone’s discomfort with saying not yet.

That ranking is a skill, not a checklist, because the inputs keep changing while you are still deciding. Practicing it before the morning it matters is the point of the drill.

ALL AT ONCERANKED BY CONSEQUENCEALARMS (2 AREAS)VENDOR ON PHONEIT: ISOLATE SEGMENTLINE MUST KEEP RUNNINGREGULATOR DEADLINE1ALARMS (2 AREAS)2IT: ISOLATE SEGMENT3LINE MUST KEEP RUNNING4VENDOR ON PHONE5REGULATOR DEADLINESAFETY FIRST · SCHEDULE LAST
Five simultaneous demands during an incident, pulled into a ranked order by consequence to the process and to safety, with the highest-consequence item highlighted.

Ninety days, and the argument that wins it

Review and Your Path Forward closes the program with a plain question: what do the next ninety days look like? The honest answer is small and specific, not a sweeping plan nobody will finish.

  • One tabletop exercise, run with the actual shift team, not a subset of managers.
  • One playbook rehearsed end to end, including the call to the vendor and the call to the plant manager.
  • One detection rule tuned until the team stops muting it.

What the operator has at the end

Practitioner Q&A carries the program’s other closing question: how have others argued for this, and won? The operators who have made this case inside their own companies describe a similar pattern. Fear does not move a budget. A specific failure tied to a cost leadership already tracks, a shift’s worth of downtime, a batch scrapped, a contract penalty, moves it. A small ask, one drill, one afternoon, gets approved faster than a big program request.

By the end of the program, an operator has practiced telling a real detection from noise, run a playbook that names who decides what, and argued a specific risk in front of the people who can act on it.

Everything produced along the way, incident timelines, rehearsed playbooks, decisions with an owner’s name attached, stays in the company’s own systems. ALDBRN trains the judgment. The record belongs to the site.

See how Response fits with Visibility and Hardening on the Programs page.