Monitoring & Alerting Articles

Build monitoring covering external availability, infrastructure resources, application logs and business metrics. Use SLOs, alerts and postmortems for continuous improvement.

What to Monitor: A Three-Layer Metrics Checklist for Servers, Applications, and Business

A monitoring dashboard full of numbers raises one question: what should you actually watch? This guide divides metrics into three layers — server (CPU/memory/disk/bandwidth), application (rate/errors/latency), and business (conversion/registrations/orders) — with the meaning and typical alert thresholds for each, plus a ready-to-use metrics checklist table.

Alert Design Basics: How to Set Thresholds and Severity So Alerts Are Neither Noisy nor Missed

Too many alerts numb the on-call team with "cry wolf" fatigue, so real incidents go unnoticed; too few alerts hide failures until users complain. Good alert design must be both "not noisy" and "not missing anything". This guide explains how to set thresholds, assign P1–P4 severities, and avoid alert fatigue, with copy-ready alert rule examples.