Define the metric before the dashboard

Specify the unit, the population being measured, the period, the source, and whether improvement means increasing or decreasing the value. “Hours spent reviewing a comparable intake cohort” is more interpretable than an unlabeled “efficiency” percentage. Record a target separately from the baseline; neither is an observation.

Compare like with like

A lower workload during a quiet month does not establish that AI caused the reduction. Compare periods and cohorts with appropriate context: volume, case complexity, staffing, seasonal variation, and other process changes. State where attribution remains uncertain. Keep the calculation and source records available so a reviewer can reproduce the observation.

Put a value on capacity

Time released may allow people to handle more work, reduce a queue, or focus on a different task. Value released capacity by multiplying hours by a labor rate; measure payroll savings and revenue against their own baselines. A financial outcome needs its own baseline, relevant costs, observation period, source, calculation, and verification. Avoid counting the same benefit in several systems.

Show gaps in the operating picture

Label stale sources, missing periods, incomplete records, and unverified observations. A failed sync should not become a zero or a green status. Record each scenario’s inputs, currency, version, and period alongside the calculation. If there is not enough evidence for a total, show the individual metrics and explain the boundary.