#monitoring

6 posts

Pushgateway Heartbeat Gotcha: When ndots and NetworkPolicy Silently Eat Your Alerts

Pushgateway Heartbeat Gotcha: When ndots and NetworkPolicy Silently Eat Your Alerts

How ndots:5, a wildcard DNS record, and a default-deny NetworkPolicy combine to swallow CronJob heartbeats without a single error or alert.

Longhorn Read-Only Mounts: Detection, Recovery, and Closing the Silent Failure Window

Longhorn Read-Only Mounts: Detection, Recovery, and Closing the Silent Failure Window

A Longhorn volume can report Healthy while the filesystem inside your pod has been read-only for hours. How to detect it, recover it, and alert on it.

Prometheus Alerting Rules That Don't Cry Wolf

Prometheus Alerting Rules That Don't Cry Wolf

How to write Prometheus alerts that carry context, tolerate transient scrape blips, and page only when something is actually broken.

Grafana Dashboards: Information Density vs Readability

Grafana Dashboards: Information Density vs Readability

Stop cramming every metric into one screen. A practical look at balancing information density and performance in Grafana dashboards.

Condition-Based vs Time-Based Maintenance: Making the Switch

Condition-Based vs Time-Based Maintenance: Making the Switch

Stop replacing parts that aren't broken. A deep dive into moving from calendar-based schedules to real-time condition-based maintenance in IIoT.

Longhorn Volume Health: The Gap Between 'Healthy' and Actually Working

Longhorn Volume Health: The Gap Between 'Healthy' and Actually Working

Stop trusting the Longhorn UI blindly. Learn to monitor replication, fix stale mounts, and manage snapshot bloat in production K8s storage.

← All tags