Incident Drill Checklist
Use this checklist to validate on-call readiness for recommendation serving incidents.
Frequency
- Weekly lightweight drill (15-20 minutes).
- Monthly full drill (45-60 minutes).
Pre-Drill Setup
- Confirm dashboards and alert routing are accessible.
- Confirm runbook links are reachable.
- Assign roles:
- Incident commander
- Primary operator
- Observer/scribe
Drill Scenarios
- Bad model rollout:
- Simulate canary fallback/latency breach.
- Execute model rollback runbook.
- Queue lag incident:
- Simulate rising queue depth and pending age.
- Execute queue lag runbook.
- ANN staleness:
- Simulate stale index alert.
- Execute ANN staleness runbook.
Success Criteria
- Detection to acknowledgement < 5 minutes.
- Correct runbook selected without escalation confusion.
- Mitigation command sequence executed correctly.
- Recovery validation metrics checked and documented.
Post-Drill Review
- Record time to detect, mitigate, and validate.
- Record any ambiguous steps or missing commands.
- Create follow-up tasks for documentation or automation gaps.
- Update runbooks and alert annotations if navigation was unclear.