TechOneDigital Start a pilot

IT Operations

AI Incident Triage & Dispatch

Every alert becomes an incident with an owner, a risk class and a next step.

Starts with an assessment · Human approval · Complete audit trail

The operational problem

An alert without context is not an incident

Monitoring usually knows what crossed a threshold. The responder still has to discover which services depend on it, whether the alert is duplicate, who owns the response and which action is safe.

The service turns an alert into a structured incident: affected scope, impact, risk class, applicable runbook, proposed owner and SLA clock. Any mutating next step remains behind human approval.

Incident context

What triage has to establish

Alert evidence

The original source, time, object, condition and related alerts remain attached.

Affected scope

Dependencies identify which services, workloads or users may be impacted.

Risk class

Read-only diagnosis is separated from mutating or destructive remediation.

Runbook

The next step comes from an approved procedure with prerequisites and stop conditions.

Ownership and SLA

The incident names an accountable responder and starts the agreed timing rule.

Method

How an alert becomes an owned incident

  1. Receive and deduplicateRelated alerts are grouped while the original evidence is preserved.
  2. Resolve dependenciesThe service identifies the likely affected systems and potential business scope.
  3. Classify impact and action riskIncident priority and the safety class of a proposed response are recorded separately.
  4. Select the approved runbookA next step is proposed only when its prerequisites match the observed state.
  5. Confirm owner and next stepA person accepts responsibility and approves any action that changes a system.
  6. Keep the complete timelineAlerts, decisions, hand-offs, actions and results stay available for the post-mortem.

Example incident

Impact and next action in one record

Illustrative incident record — no customer infrastructure data

Impact
3 VMs on the volume
Class
Mutating (extend volume)
Runbook
Extend volume, then verify
Owner
On-call engineer
SLA clock
Started

Operations discipline

The rules behind reliable dispatch

Priority is not action risk

A critical incident can begin with safe read-only diagnosis; a destructive step still needs its own approval.

No incident without an owner

Automation may propose the responder, but accountable ownership is confirmed and visible.

Runbooks have prerequisites

A familiar alert does not justify a familiar fix until the current conditions match the procedure.

The timeline is evidence

A post-mortem needs what was known and decided at each moment, not a summary written afterwards.

Fixed start

What the three-week design and pilot delivers

  • Incident taxonomyPriority, impact, risk and ownership rules mapped to the selected alert sources.
  • Runbook mapApproved procedures connected to their prerequisites, stop conditions and escalation paths.
  • Triage pilotRepresentative alerts flow through deduplication, impact analysis, ownership and next-step approval.
  • SLA and hand-off designTiming, acknowledgement and escalation rules are made explicit.
  • Post-mortem recordThe pilot shows the evidence and decisions retained for later review.

Fit

When AI triage is—and is not—ready

A good fit

  • Monitoring creates alerts but ownership and impact are resolved manually.
  • Approved runbooks exist for at least part of the incident landscape.
  • The team wants faster triage while keeping change approval with responders.

Not the right fit

  • The expectation is autonomous remediation with unrestricted administrator access.
  • Nobody owns incident priority, runbooks or escalation policy.
  • The alert source lacks stable identifiers or enough evidence to correlate events.

Where the human approves

You confirm the owner and the next step. The AI works through a fixed list of allowed operations; every approval and result is logged.

Incident questions

Questions an operations lead should ask

Does the service replace our monitoring or ticketing system?

No. It consumes agreed alerts and can create or update incidents through controlled operations. Monitoring and the system of record remain in place.

Can it remediate an incident automatically?

The triage pilot does not grant unrestricted remediation. Mutating steps require the defined human approval, and destructive steps require two approvers.

How are duplicate alerts handled?

They can be grouped when evidence points to the same dependency or incident, while the original alerts remain attached for traceability.

What if no runbook matches?

The incident is marked as needing investigation and routed to the accountable person. The service does not invent an operating procedure.

When does the SLA clock start?

The rule is agreed during process design—for example at alert acceptance or incident creation—and then applied consistently.

Experience behind the service

Built from infrastructure incident workflows

The design is based on TechOne operational work with monitoring, infrastructure dependencies, runbooks, risk classes and incident hand-offs. Customer topology, locations, addresses and internal response data remain private.

Technical stewardship: David Máj, Founder & Technology Consultant. Last reviewed 24 September 2026.