Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Incident Drill Checklist

Use this checklist to validate on-call readiness for recommendation serving incidents.

Frequency

  • Weekly lightweight drill (15-20 minutes).
  • Monthly full drill (45-60 minutes).

Pre-Drill Setup

  1. Confirm dashboards and alert routing are accessible.
  2. Confirm runbook links are reachable.
  3. Assign roles:
    • Incident commander
    • Primary operator
    • Observer/scribe

Drill Scenarios

  1. Bad model rollout:
    • Simulate canary fallback/latency breach.
    • Execute model rollback runbook.
  2. Queue lag incident:
    • Simulate rising queue depth and pending age.
    • Execute queue lag runbook.
  3. ANN staleness:
    • Simulate stale index alert.
    • Execute ANN staleness runbook.

Success Criteria

  1. Detection to acknowledgement < 5 minutes.
  2. Correct runbook selected without escalation confusion.
  3. Mitigation command sequence executed correctly.
  4. Recovery validation metrics checked and documented.

Post-Drill Review

  1. Record time to detect, mitigate, and validate.
  2. Record any ambiguous steps or missing commands.
  3. Create follow-up tasks for documentation or automation gaps.
  4. Update runbooks and alert annotations if navigation was unclear.