Why monitor your infrastructure
Why monitor
Without monitoring, you find out about problems when users complain — too late. Good monitoring lets you:
- Catch problems early, before they hit the service
- Diagnose faster with historical data instead of guesswork
- Plan capacity by watching growth trends
- Prove SLA compliance with verifiable metrics
- Find bottlenecks before they become outages
Key metrics
Network
| Metric | What it measures | Typical alert |
|---|---|---|
| Bandwidth | Traffic per interface (bps) | > 80% of capacity |
| Latency | Round-trip time (ms) | > 50 ms on LAN, > 200 ms on WAN |
| Packet loss | % of packets that never arrive | > 1% |
| Interface errors | CRC errors, discards | Any sustained increase |
| Link state | Interface up/down | Transition to down |
Devices
| Metric | What it measures | Typical alert |
|---|---|---|
| CPU | Processor load (%) | > 80% sustained |
| Memory | RAM usage (%) | > 90% |
| Temperature | Chassis/CPU degrees | Vendor-specific, often > 65 °C |
| Power supplies | Status (OK/failed) | Any failure |
| Fans | Status and RPM | Failure or abnormal speed |
Monitoring protocols
- SNMP: the main protocol for reading metrics from network gear. Works by polling — the monitor queries each device on a schedule (typically every 60-300 seconds). Universal vendor support. See [[crearack—redes-infra—conceptos-snmp]].
- ICMP (ping): the simplest availability check. Send a packet, wait for the reply. Fast and needs no configuration on the device, but only tells you “reachable or not”.
- HTTP/HTTPS checks: verify a web service responds correctly — status code, response time, expected content.
- Syslog: devices push log messages to a central server in real time. Complements polling with event detail.
Alerts without the noise
A good alert strategy avoids both false alarms (alert fatigue) and silence during real incidents:
- Set realistic thresholds based on your normal baseline, not arbitrary numbers.
- Require persistence: don’t alert on a momentary spike — e.g. CPU > 90% for 5 minutes.
- Escalate progressively: warning to the on-duty tech, critical to the manager.
- Avoid cascades: if a router dies, don’t fire alerts for everything behind it.
- Review regularly: delete alerts everyone ignores, tune the rest.
How CreaRack covers this
CreaRack splits infrastructure monitoring into focused modules, all feeding the same alert system:
- [[crearack—monitoring—que-es-observatory]] — Network Observatory, the core: SNMP polling, ping and HTTP checks on your network devices, with real-time dashboards and historical charts. Metrics are stored in a time-series database with 180-day retention, so you can answer both “what is happening now” and “what changed last quarter”.
- [[crearack—monitoring—que-es-ups]] — UPS Monitor: battery charge, load and runtime of your uninterruptible power supplies. This is how you cover the power layer described in [[crearack—redes-infra—alimentacion]].
- [[crearack—monitoring—que-es-wireless]] — Wireless Monitor: WiFi access points and their controllers — radios, SSIDs, connected clients.
- [[crearack—monitoring—alertas]] — threshold alerts across all of the above: define the metric, the threshold and the persistence, and CreaRack notifies you when it trips.
Suggested polling intervals
| Metric type | Recommended interval |
|---|---|
| Availability (ping) | 60 seconds |
| Interface traffic | 120 seconds |
| CPU / memory | 300 seconds |
| Temperature / sensors | 300 seconds |
Shorter intervals give finer detail but generate more traffic and stored data. Start conservative and tighten only where you need it. In CreaRack, this is exactly what the per-device Polling cadence (15 s to 5 min) controls — see [[crearack—network—device-page]]; availability pings automatically stay at 30 s or faster regardless of the cadence you pick.
Related
- [[crearack—monitoring—que-es-observatory]] — Network Observatory overview
- [[crearack—monitoring—que-es-ups]] — UPS Monitor overview
- [[crearack—monitoring—que-es-wireless]] — Wireless Monitor overview
- [[crearack—monitoring—alertas]] — creating threshold alerts
- [[crearack—monitoring—ping-http]] — ICMP and HTTP checks
- [[crearack—monitoring—configurar-snmp]] — set up SNMP polling
- [[crearack—redes-infra—conceptos-snmp]] — SNMP in depth
- [[crearack—network—device-page]] — the polling cadence dial
Véase también
- [[crearack—monitoring—que-es-observatory]]
- [[crearack—monitoring—que-es-ups]]
- [[crearack—monitoring—que-es-wireless]]
- [[crearack—monitoring—alertas]]
- [[crearack—monitoring—ping-http]]
- [[crearack—monitoring—configurar-snmp]]
- [[crearack—redes-infra—conceptos-snmp]]