Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Grafana Alert Rules Reference

Source: Grafana Alerting Last updated: 2026-07-20 Total: 42 alert rules (all normal)

k8s-pod-dashboard

Group: k8s-pod-dashboard
Rules: 4
Folder: Grafana-managed
Evaluate every: 1m

AlertStateSummary
Current hh-server Replicas in end-userNormalhungryhub-server pod is currently {{ $values.B.Value }} replica
Current hh-server Replicas in synNormalhungryhub-server pod is currently {{ $values.B.Value }} replica
Current hh-server Replicas in vendorNormalhungryhub-server pod is currently {{ $values.B.Value }} replica
Current hh-server Replicas in cosmosNormalhungryhub-server pod is currently {{ $values.B.Value }} replica

Monitors current replica count for hh-server across all 4 namespaces: end-user, syn, vendor, cosmos.

Kubernetes Cluster

Group: Kubernetes Cluster
Rules: 2
Evaluate every: 1m

AlertStateSummary
Cluster CPU usageNormalKubernetes cluster CPU usage currently is {{ $values.B.Value }}
Cluster Pod RestartPausedPod {{ $labels.pod }} has restarted {{ $values.A.Value }} times in the last 10 minutes

Monitors overall cluster health — CPU utilization and pod restart counts.

opensearch-hh-prod — CPU Utilization

Group: opensearch-hh-prod
Rules: 12 (4 severity levels × 3 metric types)
Evaluate every: 1m

4 CPU metric types tracked:

  • System — CPU system utilization
  • IOwait — CPU I/O wait time
  • IRQ — CPU interrupt request utilization
  • User — CPU user-space utilization

3 severity levels per metric:

AlertStateSeveritySummary
CPU System/Normal/High/Highest Utilization *NormalNormal / High / Highest[Severity] CPU {type} usage has reached {{ printf "%.2f" $values.B.Value }}% on host {{ $labels.host }}

Applied to the Aiven OpenSearch opensearch-hh-prod nodes.

opensearch-hh-prod — Memory Usage

Group: opensearch-hh-prod
Rules: 3
Evaluate every: 1m

AlertStateSeveritySummary
Memory Usage NormalNormalNormal[Normal Alert] Memory usage has reached {{ printf "%.2f" $values.B.Value }}% on host {{ $labels.host }}
Memory Usage HighNormalHigh[High Alert] Memory usage has reached {{ printf "%.2f" $values.B.Value }}% on host {{ $labels.host }}
Memory Usage HighestNormalHighest[Highest Alert] Memory usage has reached {{ printf "%.2f" $values.B.Value }}% on host {{ $labels.host }}

Services Metric — health-check-status

Group: Services Metric
Rules: 1
Evaluate every: 1m

AlertStateSummary
Public ServicesPaused🚨 PUBLIC ENDPOINT DOWN — Service: {{ .Labels.name }}. Impact: Endpoint is unreachable. Action: Immediate investigation required. Status Page: https://internal-status.hungryhub.com/

Note: Currently paused. Monitors public-facing service endpoints for availability.

Services Metric — hh-menu

Group: Services Metric
Rules: 2
Evaluate every: 1m

AlertStateSummary
hh-menu High CPU Usage AlertPausedhh-menu deployment has reached its scaling limit (5 replicas), but CPU usage (>500m) remains high
hh-menu Backlog AlertPausedPuma request backlog detected on {{ $labels.namespace }} / {{ $labels.pod }}

Note: Both currently paused. First alert triggers when CPU exceeds 500m at max replicas (5). Second monitors Puma backlog.

Services Metric — imgproxy

Group: Services Metric
Rules: 1
Evaluate every: 1m

AlertStateSummary
imgproxy Timeout ErrorsPausedimgproxy is returning timeout errors

Note: Currently paused. Alerts when imgproxy image-processing service returns timeout errors.

Services Metric — Kafka

Group: Services Metric
Rules: 3
Evaluate every: 1m

AlertStateSummary
Kafka Lag AlertPausedHigh Kafka Consumer Lag detected on topic {{ $labels.topic }}. Current lag is {{ $values.B.Value }} messages, above threshold of 1000
URP AlertPausedOne or more Kafka partitions are under-replicated. Replicas not in sync with leader — risk of data loss if broker fails. Current value: {{ $value }}
Offline Partitions AlertPausedOne or more Kafka partitions have no active leader. Producers and consumers cannot read/write. Current value: {{ $value }}

Note: All paused. Covers consumer lag, under-replicated partitions (URP), and offline partitions.

Sidekiq Monitor — Latency

Group: Sidekiq Monitor
Rules: 14 (7 queues × 2 severity levels)
Evaluate every: 1m

7 Sidekiq queues monitored:

QueueSeverityAlert Name
Kafka ProducerWarningKafka Producer - Warning
Kafka ProducerCriticalKafka Producer - Critical
Kafka Inventory ProducerWarningKafka Inventory Producer - Warning
Kafka Inventory ProducerCriticalKafka Inventory Producer - Critical
Long ProcessWarningLong Process - Warning
Long ProcessCriticalLong Process - Critical
CriticalWarningCritical - Warning
CriticalCriticalCritical - Critical
DefaultWarningDefault - Warning
DefaultCriticalDefault - Critical
Loyalty ProgramWarningLoyalty Program - Warning
Loyalty ProgramCriticalLoyalty Program - Critical
InventoryWarningInventory - Warning
InventoryCriticalInventory - Critical

Summary: 🔥 Current Sidekiq latency is **{{ $values.A.Value }}s** – above threshold!

Note: All currently paused. Warning/Critical thresholds vary per queue. Default threshold applies when no queue-specific threshold is configured.


Legend

  • Normal — Alert rule is healthy, no firing condition met
  • Paused — Rule evaluation suspended (manual or scheduled pause)
  • Firing — Condition met, alert actively firing (not present at time of capture)
  • Pending — Condition met, waiting for duration/frequency before firing

Alerting URL

Grafana: https://grafana.hungryhub.com/alerting/list