IT Service & Asset Management

Problem Management in ITIL 4: A Practical Guide

Standarity Editorial Team·ITIL 4 and ISO/IEC 20000-1 practitioners
··5 min read

Problem management is the ITIL 4 practice responsible for reducing the likelihood and impact of incidents by identifying their actual and potential causes and by managing workarounds and known errors. Where incident management restores service, problem management removes the underlying reason a service failed in the first place.

Most IT teams are fluent in firefighting. A ticket arrives, the service desk restores the user, and everyone moves on. The trouble is that the same fire keeps starting. Problem management exists to break that loop. It treats a recurring or high-impact fault as a distinct object of study, investigates why it happens, and feeds a permanent correction into change management so the incident does not return.

Problem vs Incident vs Known Error

These three terms are used loosely in day-to-day conversation, but ITIL gives each a precise meaning. Keeping them separate is what lets a service desk and a problem team hand work back and forth without confusion.

  • Incident: an unplanned interruption or reduction in quality of a service. The goal is to restore normal operation quickly.
  • Problem: the cause, or potential cause, of one or more incidents. The goal is to understand and remove it.
  • Known error: a problem that has been analysed, has a documented root cause, and has a recorded workaround. It is ready to be resolved permanently.

A worked example makes the chain obvious. Fifty users cannot reach the payroll app (fifty incidents). Analysis shows a memory leak in a shared library (one problem). The team documents the leak and a scheduled restart that keeps the app alive (a known error with a workaround). A vendor patch finally removes the leak (permanent resolution).

Reactive vs Proactive Problem Management

ITIL splits the practice into two complementary modes. Reactive problem management is triggered by incidents that have already happened, typically a series of linked incidents whose root cause could not be resolved during incident handling. It asks why an outage occurred and how to stop it returning.

Proactive problem management is a continuous activity that hunts for weaknesses before they cause an outage. Teams examine incident records, operational logs, and monitoring trends to spot patterns, such as a disk that fills a little more each week or an error rate creeping upward after every release. Acting on those signals early is far cheaper than absorbing the eventual major incident.

ITIL 4 defines problem management as the practice responsible for reducing the likelihood and impact of incidents by identifying actual and potential causes and managing workarounds and known errors. Source: Axelos ITIL 4 problem management practice guidance, summarised by TOPdesk and Xurrent (2025).

The Problem Management Workflow

A mature problem process moves through a repeatable set of stages. The exact tooling varies, but the logical steps are stable across ITIL implementations.

  • Detect: identify a problem from repeated incidents, a major incident review, monitoring trends, or supplier notification.
  • Log and categorise: record the problem, assign a category and priority based on impact and risk, and link the related incidents.
  • Investigate root cause: analyse the fault using structured techniques rather than guesswork.
  • Record a known error and workaround: document what is known so the service desk can restore users quickly while a fix is pending.
  • Resolve: raise a change to implement the permanent correction and validate that it works.
  • Close and review: confirm no recurrence, update the knowledge base, and capture lessons learned.

Root cause techniques: 5 Whys and Ishikawa

Two techniques dominate ITIL root-cause work. The 5 Whys method drills down a single causal chain by asking why repeatedly until the true origin surfaces, and it is ideal for a focused, linear fault. The Ishikawa or fishbone diagram maps many possible causes across categories such as people, process, technology, and suppliers, and it shines when several factors interact. The two are complementary: use the fishbone to widen the search, then the 5 Whys to drill into the most likely branch. Our companion guide on root cause analysis explains how to run each session and avoid the classic trap of stopping at a symptom.

The known error database

The known error database, or KEDB, is the memory of the problem practice. It stores each problem with its documented root cause and proven workaround so that the next time the fault appears, the service desk resolves it in minutes instead of reopening an investigation. A well-tended KEDB is one of the single biggest accelerators of first-line resolution and a direct link between problem management and knowledge management.

How Problem Management Links to Incident and Change

Problem management does not operate alone. It receives its raw material from incident management, which surfaces the recurring faults worth investigating, and it depends on the service desk to link incidents to the right problem record. Our incident management guide describes how those tickets are triaged and escalated in the first place.

On the output side, almost every permanent fix is delivered through change management. Problem management identifies what must change and why; change management assesses the risk and schedules the deployment. That handoff keeps corrective work controlled rather than ad hoc, which matters because a rushed fix is itself a common cause of the next major incident.

Metrics That Show Problem Management Is Working

You measure problem management by its effect on incidents, not by activity for its own sake. The most telling signals are a falling rate of repeat incidents, a shrinking backlog of open problems, and a growing proportion of incidents that the service desk resolves quickly using a KEDB workaround.

  • Repeat incident rate: recurring incidents as a share of total volume, which should trend downward.
  • Problems resolved versus opened: shows whether the backlog is being cleared or growing.
  • Percentage of incidents resolved via a known error workaround.
  • Mean time to resolve a problem, tracked separately from mean time to restore an incident.
  • Incidents prevented by proactive problem management, estimated from trends acted on before an outage.

Treat these numbers as a conversation starter, not a scorecard. A rising problem backlog is not automatically bad if the team is deliberately tackling the highest-impact faults first. The goal is fewer and less severe incidents over time, and every metric should ladder up to that outcome.

Frequently Asked Questions

What is the difference between incident and problem management?

Incident management focuses on restoring normal service as quickly as possible, often with a temporary fix. Problem management investigates the underlying cause of one or more incidents and drives a permanent resolution so the incidents do not recur.

What is a known error in ITIL?

A known error is a problem that has been analysed, has a documented root cause, and has a recorded workaround. It is stored in the known error database so the service desk can restore users quickly while a permanent fix is prepared.

What is the difference between reactive and proactive problem management?

Reactive problem management is triggered by incidents that have already occurred and seeks to stop them recurring. Proactive problem management continuously analyses logs and trends to find weaknesses and remove them before they cause an incident.

Which root cause analysis techniques does ITIL use?

The most common are the 5 Whys, which drills into a single causal chain, and the Ishikawa or fishbone diagram, which maps many possible causes across categories such as people, process, and technology. They are often used together.

How does problem management link to change management?

Problem management identifies what must change to remove a root cause and why it matters. Change management then assesses the risk and schedules the deployment, so the permanent fix is implemented in a controlled way.

What metrics measure problem management success?

Useful metrics include the repeat incident rate, the ratio of problems resolved to problems opened, the share of incidents resolved with a known error workaround, and the number of incidents prevented by proactive analysis.

Explore Courses on Udemy

Intermediate

IT Service Management (ITSM) Simplified

Intermediate

IT Service Management (ITSM), Processes and Templates

Intermediate

Implement ISO 20000-1:2018 Step By Step With Templates