The six‑metric checklist and simple scoring
- Error & exception rate — score 0–5 by frequency of bad outputs: 0 if >10% of runs produce an error/misfire, 2 for 6–10%, 4 for 2–5%, 5 for <2%.
- Human review time per item — score 0–5 by average minutes a person spends checking/fixing a single AI output: 0 if >10 minutes, 2 for 5–10, 4 for 2–5, 5 for <2.
- Token / compute cost per run — score 0–5 on direct run cost: 0 if >£0.50, 2 for £0.20–0.50, 4 for £0.05–0.20, 5 for <£0.05.
- Business value delivered — score 0–5 on clear operational benefit (time saved, revenue enabled, SLA met): 0 if negligible, 3 if modest but useful, 5 if high and repeatable.
- Adoption — score 0–5 by % of eligible users or records that actually use the automation: 0 if <30%, 2 for 30–60%, 4 for 60–80%, 5 for >80%.
- Maintenance effort — score 0–5 by hours per month to keep it running: 0 if >12 hrs/mo, 2 for 6–12, 4 for 2–6, 5 for <2.
Scoring: add the six scores (max 30). Guidance: 0–10 = retire, 11–19 = improve, 20–30 = keep. Use the per‑metric thresholds above to make consistent calls across automations.
One‑week sampling plan any ops manager can run
Run the sample over one calendar week and reserve the following week for scoring and quick fixes — two weeks total. Pick every automation you want to review, then decide a sample size: if the automation runs <300 times/week, capture every run; otherwise sample 50–200 runs (aim for 100 if unsure). Keep the sample simple: date, input summary, AI output, outcome (ok/error), human review minutes, and any exception notes.
You don’t need engineers. Pull workflow history or run logs from your platform: HubSpot workflow history or exported workflow runs, Salesforce Flow interviews or debug logs, Marketo activity exports, Pardot automation logs, or your Zapier/Make run history. If logs are thin, add a lightweight 'sample' flag or manual checkbox for incoming records for the week and have reviewers paste outputs into a shared sheet. Each day capture runs into the sheet and ask reviewers to time themselves (stopwatch on phone) and note problems.
At the end of the week compute the six metrics from the sheet, convert them to 0–5 scores, total the score and apply the retire/improve/keep thresholds. Keep notes on recurring error causes and one or two quick fixes to try in week two.
Quick remediation options for borderline automations and next steps
If an automation lands in the "improve" band, try small, fast fixes first: add gating so the automation only runs on high‑trust records (use a trust score or a required field check); reduce context sent to the model (trim fields, summarise recent activity) to lower hallucinations and tokens; lower run frequency by batching non‑urgent jobs; add a lightweight human‑in‑the‑loop for high‑risk outputs; and add simple fallbacks (a polite ‘needs review’ flag instead of taking irreversible action) and a kill switch so you can pause the job quickly.
Assign an owner, set one‑page maintenance rules (naming, expected error rate, review cadence) and re‑run the checklist quarterly or after any major model/config change. If you want a hands‑on walk‑through to run the checklist and the one‑week sample for one automation, Optira can help run a practical two‑week review and hand over a scored report you can use going forward.