Volver a la wiki

Why monitor your infrastructure

Why monitor

Without monitoring, you find out about problems when users complain — too late. Good monitoring lets you:

Key metrics

Network

MetricWhat it measuresTypical alert
BandwidthTraffic per interface (bps)> 80% of capacity
LatencyRound-trip time (ms)> 50 ms on LAN, > 200 ms on WAN
Packet loss% of packets that never arrive> 1%
Interface errorsCRC errors, discardsAny sustained increase
Link stateInterface up/downTransition to down

Devices

MetricWhat it measuresTypical alert
CPUProcessor load (%)> 80% sustained
MemoryRAM usage (%)> 90%
TemperatureChassis/CPU degreesVendor-specific, often > 65 °C
Power suppliesStatus (OK/failed)Any failure
FansStatus and RPMFailure or abnormal speed

Monitoring protocols

Alerts without the noise

A good alert strategy avoids both false alarms (alert fatigue) and silence during real incidents:

  1. Set realistic thresholds based on your normal baseline, not arbitrary numbers.
  2. Require persistence: don’t alert on a momentary spike — e.g. CPU > 90% for 5 minutes.
  3. Escalate progressively: warning to the on-duty tech, critical to the manager.
  4. Avoid cascades: if a router dies, don’t fire alerts for everything behind it.
  5. Review regularly: delete alerts everyone ignores, tune the rest.

How CreaRack covers this

CreaRack splits infrastructure monitoring into focused modules, all feeding the same alert system:

Suggested polling intervals

Metric typeRecommended interval
Availability (ping)60 seconds
Interface traffic120 seconds
CPU / memory300 seconds
Temperature / sensors300 seconds

Shorter intervals give finer detail but generate more traffic and stored data. Start conservative and tighten only where you need it. In CreaRack, this is exactly what the per-device Polling cadence (15 s to 5 min) controls — see [[crearack—network—device-page]]; availability pings automatically stay at 30 s or faster regardless of the cadence you pick.

Véase también

Subir