#prometheus

4 posts

Pushgateway Heartbeat Gotcha: When ndots and NetworkPolicy Silently Eat Your Alerts

Pushgateway Heartbeat Gotcha: When ndots and NetworkPolicy Silently Eat Your Alerts

How ndots:5, a wildcard DNS record, and a default-deny NetworkPolicy combine to swallow CronJob heartbeats without a single error or alert.

Longhorn Read-Only Mounts: Detection, Recovery, and Closing the Silent Failure Window

Longhorn Read-Only Mounts: Detection, Recovery, and Closing the Silent Failure Window

A Longhorn volume can report Healthy while the filesystem inside your pod has been read-only for hours. How to detect it, recover it, and alert on it.

Prometheus Alerting Rules That Don't Cry Wolf

Prometheus Alerting Rules That Don't Cry Wolf

How to write Prometheus alerts that carry context, tolerate transient scrape blips, and page only when something is actually broken.

Grafana Dashboards: Information Density vs Readability

Grafana Dashboards: Information Density vs Readability

Stop cramming every metric into one screen. A practical look at balancing information density and performance in Grafana dashboards.

← All tags