Service level indicators

Run the new indicator beside the old one before switching alerts

A measurement change can alter the reliability story without changing the service. Compare definitions and known failures before moving operational decisions to it.

In this article

Explain what the new measure fixes

State the gap in the current indicator, such as counting accepted requests while missing failed background completion. Define the new population and good outcome.

Keep the old and new definitions visible. The objective target may need review if the new measure covers more of the user journey.

Do not silently retain the same target and imply that historical performance is directly comparable.

Collect both views

Run the new measurement alongside the existing one through representative traffic and scheduled work. Compare event counts, latency and failure classifications.

Investigate differences using known operation identities. Some differences will be the intended improvement, while others may reveal instrumentation defects.

Use synthetic cases for rare outcomes rather than waiting indefinitely for a production incident to validate them.

Test alert behaviour before paging

Evaluate the proposed alert on historical or controlled data, including short spikes, sustained degradation and missing telemetry. Review whether it produces an actionable signal at the intended urgency.

Avoid switching every alert at once if the team cannot distinguish measurement changes from service changes. Keep a clear transition record and owner.

Confirm that the runbook points to the operations and dashboards used by the new indicator.

Retire the old path deliberately

After the new measure is accepted, update operational decisions, reports and documentation. Remove redundant alerts that would page twice for the same impact unless they serve a distinct purpose.

Preserve historical definitions and annotate the transition date. Future reviews should know why the apparent reliability changed.

Monitor the new telemetry pipeline itself. The rollout is complete when the measurement is understood and used correctly, not simply when a new chart appears beside the old one.

Compare both definitions against the same events

Run the old and new indicator definitions over the same time window before replacing the dashboard. If one counts requests and the other counts completed business operations, their percentages can differ even when the system's behaviour is unchanged.

For illustration, ten successful polling requests for one incomplete export should not become ten successful exports. Keep the denominator and event source in the migration evidence so an apparent reliability improvement is not merely a measurement change.

Primary sources

Google SRE: implementing service objectives

References checked 11 September 2026.