Key Metrics to Monitor

Metric What to Check Alert Threshold Check Frequency
Uptime Site accessible? 1 failed check Every 1-5 minutes
Response Time Server response speed > 3 seconds Every 5 minutes
SSL Certificate Days until expiry < 30 days Daily
Disk Usage Storage remaining > 85% Every 30 minutes
Memory Usage RAM consumption > 90% Every 30 minutes
CPU Load Processor utilization > 80% sustained Every 5 minutes
Database Connection status Connection failure Every 5 minutes

Monitoring Tools Comparison

Tool Type Price Best For
Uptime Kuma Self-hosted Free Small-medium sites
Better Uptime SaaS Free / $20/mo Team with on-call
Checkly SaaS Free / $30/mo API + Browser monitoring
Datadog SaaS $15/host/mo Enterprise full-stack
Grafana + Prometheus Self-hosted Free Advanced infrastructure
Netdata Self-hosted Free Real-time server metrics

Recommended Monitoring Stack

For Small Sites (Free)

Uptime Monitoring: Uptime Kuma (self-hosted)
Server Metrics: Netdata (one-liner install)
SSL Check: Uptime Kuma built-in
Notifications: Telegram / Email

For Medium Sites ($0-30/month)

Uptime Monitoring: Better Uptime or Checkly
Server Metrics: Netdata or basic Prometheus
APM: Sentry (free tier for errors)
Logs: Simple log rotation + manual check
Notifications: Slack / Telegram / Email

For Enterprise ($$$)

Uptime: Datadog / Pingdom
Metrics: Grafana + Prometheus
APM: Datadog APM / New Relic
Logs: ELK Stack / Loki
Notifications: PagerDuty / Opsgenie

Alert Notification Channels

Channel Cost Urgency Best For
Email Free Low Daily summaries, reports
Telegram Free Medium Real-time alerts, team chat
Slack Free Medium Team collaboration
Discord Free Medium Dev community channels
SMS Paid High Critical issues
Phone Call Paid Highest P0 incidents

Alert Configuration Tips

1. Set Up Alert Escalation

First alert → Telegram/Slack (immediate)
No response in 15 min → SMS
No response in 30 min → Phone call

2. Configure Maintenance Windows

  • Exclude scheduled maintenance from alerting
  • Set quiet hours for non-critical alerts (e.g., 10 PM - 8 AM)

3. Avoid Alert Fatigue

  • Don't alert on every small event
  • Use hysteresis (e.g., alert only if 3 consecutive failures)
  • Group related alerts

Monitoring Checklist

  • Uptime monitoring configured (5-min intervals)
  • SSL certificate expiry alert (< 30 days)
  • Disk usage alert (> 85%)
  • Memory usage alert (> 90%)
  • Database connection monitoring
  • Notification channel configured (Telegram/Slack/Email)
  • Maintenance windows defined
  • Alert escalation policy documented
  • Dashboard created for key metrics
  • Incident response plan documented