Monitoring alerts
What it is
An alert is a rule: a metric, a threshold and a severity. Observatory evaluates the rules continuously against the data your Local Agents push (every metrics batch, roughly twice a minute) and on every manual check (ping, SNMP poll, HTTP check). When a rule matches, it opens an alert event that records the value at trigger time. While the condition persists, the same open event is updated — you get one event per incident, not one per check.
Alerts can be per-device or global (one rule applied to every target in your organization).
How to create an alert
- In Observatory, open the alert manager from a device tab (Alert Configuration modal).
- Fill in:
- Alert Name — e.g. “High Latency Warning”.
- Condition — see the table below.
- Threshold Value — the number that trips the rule.
- Sustained for (s) — how many seconds the condition must hold continuously before the alert fires (default 60; set 0 to fire on the first reading). This keeps a single half-second spike from opening an alert. For Target down for, this field is disabled: the threshold value itself is the duration.
- Severity — Good (green), Info (blue) or Critical (red).
- Tick Apply to All Devices to make it global instead of per-device.
- Click Add Alert.
Conditions
| Group | Condition | Unit |
|---|---|---|
| Heartbeat (Ping) | Latency above / Latency below | ms |
| Heartbeat (Ping) | Packet loss above / Packet loss below | % |
| Bandwidth (SNMP) | Bandwidth above / Bandwidth below | Mbps |
| HTTP | HTTP time above / HTTP time below | ms |
| Status | Target down for | s |
Latency, packet loss and down/up state are evaluated live from Agent data. Bandwidth and HTTP conditions evaluate when their checks run (SNMP poll, HTTP check).
Global alerts show a Global badge in the list and draw a dashed threshold line on every device’s chart. You can also drag a threshold line directly on a chart to adjust its value, and toggle the lines with the AL (alert lines) button.
What happens when an alert fires
- The event appears in the sidebar severity cards (Good / Info / Critical counters). Click a card to list the events of that severity.
- A toast notification pops up in real time (pushed over the Observatory WebSocket).
- If any Critical event is active, the interface switches to red-alert mode (the background asteroids canvas turns red) until all critical events are archived.
- Storm protection for global alerts: if a global rule trips on 10 or more devices in the same cycle (say, a whole LAN going down), Observatory opens a single grouped event labelled All Devices with the affected count, instead of flooding the panel with one event per device.
Be aware of what does not happen: threshold alerts do not send emails or chat messages by themselves. There are two ways to be told by email or chat: the team channels of the ITSM module (email lists, webhook, Slack, Teams), which fire for CNS insights and incidents — see [[crearack—monitoring—itsm]] — and each person’s personal email alerts, described below.
Personal email alerts
Each person with a CreaRack account can choose which device problems they want to hear about by email, without touching the team’s ITSM channels. Everything is off by default: nobody receives anything until they switch it on.
How to set them up
- Open Configuration in the top bar and choose Users → User Settings. The Email alerts section is right there, in the left column, for every user.
- Tick Send me alerts by email.
- Address: leave it empty to use your account email. If you type a different address, CreaRack sends a confirmation email to it, and that address receives nothing until its owner opens the link and presses Confirm this address. While you wait, the section says Waiting for confirmation — check that inbox; Resend confirmation sends it again (at most once every 10 minutes; the link expires after 48 hours).
- Tick the alerts you want:
- Device down (more than 2 min) — a monitored device stops answering for more than two minutes.
- Device unstable (3+ drops in an hour) — a device drops three or more times in an hour. Only drops of 90 seconds or longer count (a real restart takes that long; a single lost ping does not). CreaRack also opens a CNS incident for it, and the email carries the incident’s diagnosis and suggested action.
- CNS incident — high severity / CNS incident — any severity — a new CNS incident in your organization. If you tick both, you get one email per incident, not two.
- Incident response time missed (SLA) — a CNS incident has not been acknowledged or resolved within its SLA time.
- Remind me: Once per problem, or a reminder Every hour, Every 4 hours or Every 24 hours while a device stays down.
- Also tell me when it is resolved: an email when the device answers again, with how long it was down.
- Click Save Email Alerts.
An Admin can also set up someone else’s alerts from that person’s Edit User window (User Settings → Edit next to the user).
What you receive
Each email says which device (name and IP), what happened and when (Madrid time), and has a link that opens the device in Observatory. You only get alerts for what your role can see: device alerts need access to Observatory, CNS and SLA alerts need access to CNS.
To protect your inbox, CreaRack sends at most 10 alert emails per hour to each person. If there are more, you get a single There are more alerts in CreaRack email and the rest are held back until you are under the limit again. Devices that are out of service or inside an active maintenance window send nothing.
Good to know
- When you switch alerts on, CreaRack does not send you the backlog: only problems from the last hour onwards.
- These emails are personal. The team channels (email lists, Slack, Teams, webhooks) are configured separately in [[crearack—monitoring—itsm]].
- In a CNS incident, the detail window now lists the device’s drops in the last 24 hours, so you can see at a glance whether it is a one-off or a pattern.
How to manage events and rules
- Acknowledge an event to tell the team someone owns it. CreaRack records who acknowledged it and when.
- Resolve (archive) an event when the incident is over. Resolving also marks it acknowledged, with the resolver recorded.
- Auto-resolve: when the condition stops being met (the device comes back, latency drops under the threshold), the open event is closed automatically with no resolver, and the Active Alerts list refreshes on its own. A grouped All Devices event closes as soon as fewer than 10 devices still meet the condition; the ones still failing get their own per-device events.
- Rules can be edited (name, condition, threshold, severity), enabled/disabled without deleting, or deleted.
Each event stores the alert, the affected target, the metric value at trigger, and the full acknowledge/resolve trail. Resolved events are kept for 90 days, then removed by the daily cleanup job.
Devices that are out of service
A device you have taken out of service (Take out of Service on its card) raises no alerts and no notifications while it is stored, and it returns to normal alerting by itself as soon as it answers ping again. See [[crearack—monitoring—que-es-observatory]].
Troubleshooting
| Problem | Likely cause | Fix |
|---|---|---|
| Alert never fires | Monitor for that metric is off | Enable ping/SNMP/HTTP on the target — rules only evaluate metrics that are being collected |
| Alert seems delayed | The Sustained for duration is doing its job | The condition must hold continuously for the configured seconds before firing; lower it (or set 0) if you want instant alerts |
| Alert fires constantly | Threshold below the device’s normal baseline | Watch the chart for a day, then set the threshold above the observed baseline — and give it a sustain duration so spikes don’t count |
| Events pile up unresolved | The condition is still true | Events close by themselves when the condition clears; if one stays open, the device is still failing the rule — fix it or adjust the threshold |
Related
- [[crearack—monitoring—que-es-observatory]] — module overview
- [[crearack—monitoring—metricas]] — the metrics rules evaluate
- [[crearack—monitoring—ping-http]] — latency, loss and HTTP time sources
- [[crearack—monitoring—configurar-snmp]] — bandwidth source for SNMP alerts
- [[crearack—monitoring—itsm]] — incidents, escalation and notification channels
Véase también
- [[crearack—monitoring—que-es-observatory]]
- [[crearack—monitoring—metricas]]
- [[crearack—monitoring—ping-http]]
- [[crearack—monitoring—configurar-snmp]]
- [[crearack—monitoring—itsm]]