Support IT Helpdesk

ITIL and process

Incident, problem or change? A decision tree that works

The three records exist to answer three different questions. Here is the decision tree, and the two habits that turn an ITIL implementation into filing for its own sake.

Teams adopting ITIL usually stall at the same place: everything looks like it could be an incident, a problem or a change, so people either raise all three or raise none. The distinction is simpler than the framework makes it sound, because each record answers a different question.

Three records, three questions.
RecordThe question it answersWhen it closes
Incident How do we get service back now? When service is restored — even by a workaround.
Problem Why did this keep happening? When the root cause is understood and dealt with, or accepted.
Change What are we about to alter, and what if it goes wrong? When the change has been made and reviewed.

The decision tree

  1. Is something broken right now for somebody trying to work? → Incident. Restore service first; understanding can wait.
  2. Has that same thing broken more than twice, or once expensively? → Also raise a Problem, linked to the incidents.
  3. Are you about to alter something that other people depend on? → Change, before you touch it.
  4. Is somebody asking for something normal — access, software, a new laptop? → Service request, not an incident. It is not broken; they want something.

The fourth one catches most teams out. "I need access to the finance folder" is not an incident, and filing it as one makes your incident numbers meaningless — which matters, because incident volume is the number people use to argue for headcount.

Restore first, understand later

An incident closes when service is restored, including by a workaround you are not proud of. Rebooting the server is a legitimate incident resolution. It is not a legitimate problem resolution, and that is the whole point of having two records.

Teams that refuse to close an incident until they understand the cause end up with incidents open for weeks, an unusable ageing report, and an engineer who has moved on anyway.

Problems: the useful discipline, and the trap

Problem management is where the value is, and it is also the first thing to be abandoned, because it competes with today's queue and nobody notices when it is skipped.

One rule makes it survivable: only raise a problem when you can name the incidents it explains. A problem record linked to six incidents is a case for spending a day on it. A problem record linked to nothing is a hunch, and it will sit open until somebody tidies it away.

Change: the approval is not the point

Most small teams implement change management as an approval step and stop there. Approval is the least useful part. The parts that repay the effort are the backout plan and the schedule conflict check.

  • A backout plan forces somebody to think about failure before it happens, which is when thinking is cheap.
  • A conflict check catches the two engineers who both planned to touch the same system on Saturday.
  • A review after the fact is where you find out that your changes fail for a reason you could fix.

If you implement only one thing, implement the backout plan. If you implement only the approval, you have added a delay and gained nothing.

Two habits that make the whole thing pointless

The first is raising records after the fact to satisfy an audit. A change raised on Monday for something done on Saturday is a lie with a reference number, and everybody involved knows it.

The second is filing without linking. An incident that does not reference its problem, and a change that does not reference the incident that prompted it, leave you with three piles of paper instead of a story. The links are what make the records worth having.

The priority argument, settled once

Every desk eventually argues about priority, usually while something is on fire. The argument is avoidable because priority is not an opinion: it is impact multiplied by urgency, and both halves have observable answers. Impact is how many people cannot work. Urgency is how fast it is getting worse. A finance system down on the last day of the month is high on both; the same system down on the fifth is high on impact and low on urgency, and that is a real difference worth acting on.

Write the matrix once, put it where the desk can see it, and let the requester state impact rather than priority. People are honest about how many colleagues are affected and optimistic about how urgent their own request is, so ask them the question they can answer accurately.

What a ten-person team actually needs

Start with incidents and service requests, because they cover the day. Add problems when you notice yourself fixing the same thing for the third time. Add changes when something you altered breaks something you did not expect. Adopt the register when you feel its absence, not because a framework diagram has a box for it.

All five ITIL registers, each with its own lifecycle, linked to the tickets that prompted them. See the ITIL registers

Tagged: Itil

Written by Support IT Ventures team

The people who build the product

We build and run Support IT Helpdesk. Everything here comes out of running an IT service desk ourselves before we ever sold one, so it leans towards what actually happens on a Tuesday afternoon rather than what sounds good in a framework diagram.

Everything by this author

Read next

Try it on your own desk

Everything here is what the product does. You can have a workspace in about a minute, with no card.

Start free