top of page
Search

SLA, KPI or Vanity Metric? How to Measure IT Service Performance Properly

  • ipunton
  • 8 hours ago
  • 3 min read

Boards and IT leaders are not short of dashboards. Red-amber-green status, hundreds of tickets closed, 99.9% uptime—on paper, everything looks healthy. Yet service complaints persist, audit findings repeat, and risk exposure grows. The problem is rarely a lack of data; it is a lack of meaningful measurement.


In regulated UK SMEs, where assurance, resilience and accountability are non‑negotiable, metrics must do more than report activity. They must evidence service quality, demonstrate control effectiveness, and inform decisions on risk and improvement. This is where the distinction between SLAs, KPIs and vanity metrics becomes critical.


Why this matters in a regulated environment


Frameworks such as ISO/IEC 20000‑1 (service management) and ISO/IEC 27001 (information security) both require organisations to define, monitor and review performance using appropriate metrics. The NCSC and Cyber Essentials Plus further emphasise evidence of control operation and continuous monitoring. Under UK GDPR, organisations must be able to demonstrate appropriate technical and organisational measures.


Put simply: if you cannot show that your IT services are performing effectively—beyond superficial reporting—you cannot demonstrate due diligence.



Where organisations go wrong


  1. Confusing SLAs with outcomes


    SLAs are contractual thresholds (e.g. 4-hour response, 8-hour resolution). Meeting them does not necessarily mean users are satisfied or services are reliable. An incident resolved within SLA that recurs weekly still represents poor service.

  2. Overloading dashboards with activity metrics


    Tickets closed, calls answered, patches applied—these show effort, not effectiveness. Activity without context can mask underlying issues.

  3. RAG status without substance


    “Green” often reflects compliance with arbitrary thresholds rather than genuine service health. It can desensitise leadership to emerging risks.

  4. Lack of linkage to risk and business impact


    Metrics rarely map to business services (e.g. payroll, customer portal). Without this, reporting cannot support risk-based decision-making.

  5. No ownership or review discipline


    Metrics are produced but not challenged. Without governance—defined owners, review forums, and actions—they quickly become noise.


Governance implications


Poor measurement design leads to poor assurance. Boards may receive false confidence, auditors raise repeat findings, and management lacks a clear basis for prioritisation. Over time, this erodes trust in reporting and weakens control environments.

A governance-led approach requires that metrics are:

  • Aligned to objectives and risk (what matters to the business)

  • Evidenced and auditable (how the number is derived and validated)

  • Actionable (what decision it informs)

  • Owned (who is accountable for improvement)


What good looks like: practical controls


Below is a pragmatic approach to designing meaningful IT service metrics.


1. Start with service outcomes, not tools

Define what “good” looks like for each critical business service:

  • Availability and stability (e.g. payroll available on processing days)

  • Data integrity and security (aligned to ISO 27001 controls)

  • User experience (measurable satisfaction and usability)

Then derive KPIs that reflect these outcomes.


2. Separate SLAs, KPIs and KRIs

  • SLAs: Minimum service thresholds (contractual/operational)

  • KPIs: Indicators of performance and quality (trend over time)

  • KRIs (Key Risk Indicators): Signals of increasing exposure (e.g. repeated high-severity incidents)

Example set:

  • SLA: 95% of P1 incidents responded to within 15 minutes

  • KPI: Reduction in repeat P1 incidents per quarter

  • KRI: Percentage of critical systems without current security patches


3. Focus on service quality and stability

Introduce metrics that reveal systemic health:

  • First-time fix rate

  • Incident recurrence rate

  • Change success rate (aligned to change management under ISO 20000)

  • Mean time between failures (MTBF) for critical services

These are far more indicative of maturity than volume metrics.


4. Map metrics to business services and risk

Create a simple service catalogue linking IT components to business outcomes. Report performance at this level:

  • “Customer portal availability during peak hours”

  • “Time to restore finance system following failure”

This enables boards to understand impact, not just activity.


5. Define clear data ownership and lineage

For each metric:

  • Data source (ITSM tool, monitoring platform)

  • Calculation method

  • Frequency of reporting

  • Owner responsible for accuracy and improvement

This supports auditability and aligns with ISO 27001 evidence expectations.


6. Embed review and challenge

Establish a monthly service review with:

  • Trend analysis (not snapshots)

  • Root cause of exceptions

  • Agreed actions with owners and deadlines

Avoid static dashboards. Focus on narrative and decision-making.


7. Eliminate vanity metrics

Regularly challenge each metric:

  • Does it drive a decision?

  • Does it reflect user or business impact?

  • Would a failure of this metric trigger action?

If not, remove it.


A governance-led approach to assurance


At cyberISMS, we typically see organisations evolve from activity-heavy reporting to outcome-led assurance by redesigning their measurement framework. This includes aligning metrics to ISO 20000‑1 service management processes, embedding security and risk indicators from ISO 27001, and ensuring reporting supports board-level oversight.

The result is not more metrics—but better ones: fewer, clearer, and directly linked to risk, service quality and continuous improvement.


Call to action


If your dashboards are always green but issues persist, it may be time to challenge your measurement design. Start by identifying one critical service and assess whether your current metrics genuinely reflect its performance and risk.

A governance-led approach to IT service measurement turns reporting from reassurance theatre into real assurance.


 
 
 

Comments


bottom of page