Likelihood and impact scales are useful only when people interpret them consistently. A five-by-five matrix can support discussion, but the definitions behind the numbers matter more than the coloured cells.
Define the decision horizon
A likelihood scale should state the relevant period. “Possible” over one month is not the same as “possible” over five years. Use frequency ranges where data exist and observable conditions where they do not.
Tie impact to objectives
A project may focus on schedule, scope, cost, quality, and stakeholder acceptance. An operation may focus on service, safety, compliance, financial loss, and recovery time. The scale should reflect what actually matters.
Do not average away severity
A risk with modest financial impact but serious safety or legal implications should not automatically become medium because categories were averaged. Define rules for the highest credible impact and mandatory escalation.
Matrix limitations
Example scale
| Rating | Likelihood | Impact |
|---|---|---|
| 1 | Rare in the stated period | Minor and easily absorbed |
| 2 | Unlikely but plausible | Limited disruption |
| 3 | Possible under normal conditions | Material management attention |
| 4 | Likely or recurring | Major objective impairment |
| 5 | Almost certain or emerging | Severe or sustained objective failure |
Write descriptions before choosing colours
A scale works best when each level has a plain-language meaning. Likelihood may be described using frequency, probability ranges, or qualitative conditions. Impact may cover service, safety, financial, legal, schedule, reputation, or other objectives. The descriptions should match the organization’s decisions; copying a generic matrix can create false consistency.
Avoid averaging away serious consequences
If a risk has moderate cost impact but severe safety or legal impact, averaging the categories can hide what matters. Some organizations use the highest credible consequence, separate category ratings, or explicit escalation rules. The method should be documented and applied consistently enough for comparison while still allowing judgment.
Scale design checks
- Do adjacent levels have meaningfully different descriptions?
- Is the assessment period stated?
- Can users rate both frequent small events and rare severe events?
- Are control assumptions separated from the raw consequence?
- Will the scale support a decision, or merely produce a colour?