ITIL incident management is the practice responsible for restoring normal service operation as quickly as possible after an unplanned interruption or reduction in service quality, while limiting the impact on users and the business. Its measure of success is speed of recovery, not the elegance of the fix.
This article covers service-management incident management, which is about restoring normal service. It is distinct from security incident response, which is about containing and recovering from a cyber attack. If you handle breaches and security events, see our ISO 27035 security incident management guide, the security counterpart to this process.
The distinction matters because the two disciplines answer different questions. Service-management incident management asks how fast can we get the user working again. Security incident response asks how do we contain the attacker and preserve evidence. They share vocabulary but follow different playbooks, and confusing them leads to slow recovery or mishandled evidence.
The ITIL Incident Management Process Steps
A well-run incident moves through a consistent sequence from the moment it is reported to the moment it is closed. Each step exists to shorten recovery time and keep the right people informed.
- Identify: detect the incident from a user report, an event alert, or monitoring.
- Log: record the incident with a unique reference, timestamp, and the details needed to work it.
- Categorise: assign a category so the ticket routes to the right team and similar incidents can be grouped.
- Prioritise: set priority from impact multiplied by urgency using the priority matrix.
- Diagnose: investigate the cause of the disruption and attempt an initial resolution.
- Escalate: pass the incident to a higher support tier or specialist when it cannot be resolved in place.
- Resolve: restore service, using a workaround from the known error database where one exists.
- Close: confirm with the user that service is restored, record the resolution, and close the record.
The Priority Matrix: Impact and Urgency
ITIL calculates priority as a function of two inputs. Impact is the degree of business disruption and the number of users affected. Urgency is how time-sensitive the resolution is and the business risk of delay. Plotting the two on a grid produces a priority, typically P1 to P4 or Critical to Low, and each priority carries a target response and resolution time.
Making the matrix explicit removes argument from triage. A single executive who cannot open one document is high urgency but low impact. An outage of the customer-facing checkout is high on both axes and becomes a P1. Agreeing these rules in advance means the service desk can prioritise consistently under pressure, rather than relying on whoever shouts loudest.
Good prioritisation also protects the team from its own biases. Without an agreed matrix, priority tends to drift toward the most visible or most senior requester rather than the greatest business risk. A published matrix, reviewed periodically, keeps the queue honest and gives analysts something objective to point to when a caller insists their issue is critical. The following inputs are worth defining before an incident ever arrives.
- Impact levels: how you rate a single user, a department, a site, or the whole organisation.
- Urgency levels: how quickly the business needs restoration given deadlines and dependencies.
- The resulting priority grid: which impact and urgency combinations map to P1 through P4.
- Target response and resolution times for each priority, agreed in the SLA.
- The trigger and named owner for declaring a major incident.
Illustrative SLA targets by priority: P1 incidents can expect a response within 15 minutes and escalation at 30 minutes; P2 up to 4 hours to resolve; P3 up to 8 hours; P4 up to 24 hours. Source: TOPdesk and ManageEngine ITIL incident priority guidance (2025). Set your own targets in your SLA.
SLAs and Major Incident Handling
Service level agreements turn priority into a commitment. Each priority level is tied to a response and resolution target, and breaching that target triggers escalation. Our service level agreement guide explains how to define, measure, and report against those targets without drowning the team in penalties.
A major incident is a high-impact, high-urgency event that significantly affects business operations, multiple users, or a critical service, and it demands a response beyond the standard flow. Once declared, organisations stand up a dedicated team, usually a major incident manager, technical responders, a business representative, and a communications coordinator. That team drives immediate containment, regular status updates, and service restoration, followed by a post-incident review that often feeds problem management.
The Service Desk: Single Point of Contact
The service desk sits at the centre of incident management as the single point of contact between the IT provider and its users. It logs and manages every incident and service request, gathers diagnostics, resolves what it can on first contact, and provides the interface to all other service operation activities.
A strong first line is the cheapest place to resolve an incident. When the service desk can close a ticket immediately using a known error workaround, it improves both resolution time and user satisfaction while shielding specialist teams from routine work. That first-contact resolution rate is one of the most watched metrics in any service operation, because every ticket that escapes to a specialist queue costs more time and delays the users behind it.
To play that role well, the service desk needs three things: a clear categorisation scheme so tickets route correctly, ready access to the known error database so workarounds are one search away, and defined escalation paths so nothing stalls when first line reaches its limit. Where any of these is weak, incidents queue, users chase updates, and the whole recovery time inflates regardless of how skilled the specialists behind the desk may be.
How Incident Management Links to Problem and Change
Incident management does not work in isolation. When the same incident keeps returning, it becomes the trigger for problem management, which investigates the root cause and drives a permanent fix. Our problem management guide describes how that investigation runs and how the known error database feeds workarounds back to the service desk.
Change management completes the loop. Many permanent resolutions, and some emergency fixes during a major incident, are delivered as changes so the risk is assessed and controlled. Keeping incident, problem, and change connected is what turns reactive firefighting into a system that gradually reduces its own workload. For organisations formalising these practices, ISO/IEC 20000-1 provides the management-system requirements that wrap around them.
The practical takeaway is simple. Restore service first, record everything as you go, and let the recurring and severe cases flow into problem and change management. Do that consistently and incident volume, not just incident speed, starts to improve.