IT Service & Asset Management

ITIL Incident Management Process: A Step Guide

Standarity Editorial Team·ITIL 4 and ISO/IEC 20000-1 practitioners
··5 min read

ITIL incident management is the practice responsible for restoring normal service operation as quickly as possible after an unplanned interruption or reduction in service quality, while limiting the impact on users and the business. Its measure of success is speed of recovery, not the elegance of the fix.

This article covers service-management incident management, which is about restoring normal service. It is distinct from security incident response, which is about containing and recovering from a cyber attack. If you handle breaches and security events, see our ISO 27035 security incident management guide, the security counterpart to this process.

The distinction matters because the two disciplines answer different questions. Service-management incident management asks how fast can we get the user working again. Security incident response asks how do we contain the attacker and preserve evidence. They share vocabulary but follow different playbooks, and confusing them leads to slow recovery or mishandled evidence.

The ITIL Incident Management Process Steps

A well-run incident moves through a consistent sequence from the moment it is reported to the moment it is closed. Each step exists to shorten recovery time and keep the right people informed.

  • Identify: detect the incident from a user report, an event alert, or monitoring.
  • Log: record the incident with a unique reference, timestamp, and the details needed to work it.
  • Categorise: assign a category so the ticket routes to the right team and similar incidents can be grouped.
  • Prioritise: set priority from impact multiplied by urgency using the priority matrix.
  • Diagnose: investigate the cause of the disruption and attempt an initial resolution.
  • Escalate: pass the incident to a higher support tier or specialist when it cannot be resolved in place.
  • Resolve: restore service, using a workaround from the known error database where one exists.
  • Close: confirm with the user that service is restored, record the resolution, and close the record.

The Priority Matrix: Impact and Urgency

ITIL calculates priority as a function of two inputs. Impact is the degree of business disruption and the number of users affected. Urgency is how time-sensitive the resolution is and the business risk of delay. Plotting the two on a grid produces a priority, typically P1 to P4 or Critical to Low, and each priority carries a target response and resolution time.

Making the matrix explicit removes argument from triage. A single executive who cannot open one document is high urgency but low impact. An outage of the customer-facing checkout is high on both axes and becomes a P1. Agreeing these rules in advance means the service desk can prioritise consistently under pressure, rather than relying on whoever shouts loudest.

Good prioritisation also protects the team from its own biases. Without an agreed matrix, priority tends to drift toward the most visible or most senior requester rather than the greatest business risk. A published matrix, reviewed periodically, keeps the queue honest and gives analysts something objective to point to when a caller insists their issue is critical. The following inputs are worth defining before an incident ever arrives.

  • Impact levels: how you rate a single user, a department, a site, or the whole organisation.
  • Urgency levels: how quickly the business needs restoration given deadlines and dependencies.
  • The resulting priority grid: which impact and urgency combinations map to P1 through P4.
  • Target response and resolution times for each priority, agreed in the SLA.
  • The trigger and named owner for declaring a major incident.

Illustrative SLA targets by priority: P1 incidents can expect a response within 15 minutes and escalation at 30 minutes; P2 up to 4 hours to resolve; P3 up to 8 hours; P4 up to 24 hours. Source: TOPdesk and ManageEngine ITIL incident priority guidance (2025). Set your own targets in your SLA.

SLAs and Major Incident Handling

Service level agreements turn priority into a commitment. Each priority level is tied to a response and resolution target, and breaching that target triggers escalation. Our service level agreement guide explains how to define, measure, and report against those targets without drowning the team in penalties.

A major incident is a high-impact, high-urgency event that significantly affects business operations, multiple users, or a critical service, and it demands a response beyond the standard flow. Once declared, organisations stand up a dedicated team, usually a major incident manager, technical responders, a business representative, and a communications coordinator. That team drives immediate containment, regular status updates, and service restoration, followed by a post-incident review that often feeds problem management.

The Service Desk: Single Point of Contact

The service desk sits at the centre of incident management as the single point of contact between the IT provider and its users. It logs and manages every incident and service request, gathers diagnostics, resolves what it can on first contact, and provides the interface to all other service operation activities.

A strong first line is the cheapest place to resolve an incident. When the service desk can close a ticket immediately using a known error workaround, it improves both resolution time and user satisfaction while shielding specialist teams from routine work. That first-contact resolution rate is one of the most watched metrics in any service operation, because every ticket that escapes to a specialist queue costs more time and delays the users behind it.

To play that role well, the service desk needs three things: a clear categorisation scheme so tickets route correctly, ready access to the known error database so workarounds are one search away, and defined escalation paths so nothing stalls when first line reaches its limit. Where any of these is weak, incidents queue, users chase updates, and the whole recovery time inflates regardless of how skilled the specialists behind the desk may be.

How Incident Management Links to Problem and Change

Incident management does not work in isolation. When the same incident keeps returning, it becomes the trigger for problem management, which investigates the root cause and drives a permanent fix. Our problem management guide describes how that investigation runs and how the known error database feeds workarounds back to the service desk.

Change management completes the loop. Many permanent resolutions, and some emergency fixes during a major incident, are delivered as changes so the risk is assessed and controlled. Keeping incident, problem, and change connected is what turns reactive firefighting into a system that gradually reduces its own workload. For organisations formalising these practices, ISO/IEC 20000-1 provides the management-system requirements that wrap around them.

The practical takeaway is simple. Restore service first, record everything as you go, and let the recurring and severe cases flow into problem and change management. Do that consistently and incident volume, not just incident speed, starts to improve.

Frequently Asked Questions

What is ITIL incident management?

ITIL incident management is the practice responsible for restoring normal service operation as quickly as possible after an unplanned interruption, while limiting the impact on users and the business. Its priority is speed of recovery rather than a permanent fix.

What are the steps in the ITIL incident management process?

The typical steps are identify, log, categorise, prioritise, diagnose, escalate, resolve, and close. Prioritisation uses the priority matrix, and resolution may rely on a workaround from the known error database.

How is incident priority calculated in ITIL?

Priority is a function of impact multiplied by urgency. Impact reflects how many users and business processes are affected, and urgency reflects how time-sensitive the resolution is. The two combine in a matrix to produce a priority level such as P1 to P4.

What is a major incident in ITIL?

A major incident is a high-impact, high-urgency event that significantly affects business operations, multiple users, or a critical service. It triggers a dedicated response team, frequent communications, and a post-incident review.

Is ITIL incident management the same as security incident response?

No. ITIL incident management restores normal service after any disruption, while security incident response contains and recovers from cyber attacks and preserves evidence. They share terms but follow different playbooks; ISO 27035 covers the security side.

What is the role of the service desk in incident management?

The service desk is the single point of contact for users. It logs and manages incidents, gathers diagnostics, resolves what it can on first contact, and provides the interface to problem, change, and other service operation activities.

Explore Courses on Udemy

Intermediate

IT Service Management (ITSM) Simplified

Intermediate

IT Service Management (ITSM), Processes and Templates

Intermediate

Implement ISO 20000-1:2018 Step By Step With Templates