Volver a la wiki

Monitoring alerts

What it is

An alert is a rule: a metric, a threshold and a severity. Observatory evaluates the rules continuously against the data your Local Agents push (every metrics batch, roughly twice a minute) and on every manual check (ping, SNMP poll, HTTP check). When a rule matches, it opens an alert event that records the value at trigger time. While the condition persists, the same open event is updated — you get one event per incident, not one per check.

Alerts can be per-device or global (one rule applied to every target in your organization).

How to create an alert

  1. In Observatory, open the alert manager from a device tab (Alert Configuration modal).
  2. Fill in:
    • Alert Name — e.g. “High Latency Warning”.
    • Condition — see the table below.
    • Threshold Value — the number that trips the rule.
    • Sustained for (s) — how many seconds the condition must hold continuously before the alert fires (default 60; set 0 to fire on the first reading). This keeps a single half-second spike from opening an alert. For Target down for, this field is disabled: the threshold value itself is the duration.
    • Severity — Good (green), Info (blue) or Critical (red).
  3. Tick Apply to All Devices to make it global instead of per-device.
  4. Click Add Alert.

Conditions

GroupConditionUnit
Heartbeat (Ping)Latency above / Latency belowms
Heartbeat (Ping)Packet loss above / Packet loss below%
Bandwidth (SNMP)Bandwidth above / Bandwidth belowMbps
HTTPHTTP time above / HTTP time belowms
StatusTarget down fors

Latency, packet loss and down/up state are evaluated live from Agent data. Bandwidth and HTTP conditions evaluate when their checks run (SNMP poll, HTTP check).

Global alerts show a Global badge in the list and draw a dashed threshold line on every device’s chart. You can also drag a threshold line directly on a chart to adjust its value, and toggle the lines with the AL (alert lines) button.

What happens when an alert fires

Be aware of what does not happen: threshold alerts do not send emails or chat messages by themselves. There are two ways to be told by email or chat: the team channels of the ITSM module (email lists, webhook, Slack, Teams), which fire for CNS insights and incidents — see [[crearack—monitoring—itsm]] — and each person’s personal email alerts, described below.

Personal email alerts

Each person with a CreaRack account can choose which device problems they want to hear about by email, without touching the team’s ITSM channels. Everything is off by default: nobody receives anything until they switch it on.

How to set them up

  1. Open Configuration in the top bar and choose Users → User Settings. The Email alerts section is right there, in the left column, for every user.
  2. Tick Send me alerts by email.
  3. Address: leave it empty to use your account email. If you type a different address, CreaRack sends a confirmation email to it, and that address receives nothing until its owner opens the link and presses Confirm this address. While you wait, the section says Waiting for confirmation — check that inbox; Resend confirmation sends it again (at most once every 10 minutes; the link expires after 48 hours).
  4. Tick the alerts you want:
    • Device down (more than 2 min) — a monitored device stops answering for more than two minutes.
    • Device unstable (3+ drops in an hour) — a device drops three or more times in an hour. Only drops of 90 seconds or longer count (a real restart takes that long; a single lost ping does not). CreaRack also opens a CNS incident for it, and the email carries the incident’s diagnosis and suggested action.
    • CNS incident — high severity / CNS incident — any severity — a new CNS incident in your organization. If you tick both, you get one email per incident, not two.
    • Incident response time missed (SLA) — a CNS incident has not been acknowledged or resolved within its SLA time.
  5. Remind me: Once per problem, or a reminder Every hour, Every 4 hours or Every 24 hours while a device stays down.
  6. Also tell me when it is resolved: an email when the device answers again, with how long it was down.
  7. Click Save Email Alerts.

An Admin can also set up someone else’s alerts from that person’s Edit User window (User Settings → Edit next to the user).

What you receive

Each email says which device (name and IP), what happened and when (Madrid time), and has a link that opens the device in Observatory. You only get alerts for what your role can see: device alerts need access to Observatory, CNS and SLA alerts need access to CNS.

To protect your inbox, CreaRack sends at most 10 alert emails per hour to each person. If there are more, you get a single There are more alerts in CreaRack email and the rest are held back until you are under the limit again. Devices that are out of service or inside an active maintenance window send nothing.

Good to know

How to manage events and rules

Each event stores the alert, the affected target, the metric value at trigger, and the full acknowledge/resolve trail. Resolved events are kept for 90 days, then removed by the daily cleanup job.

Devices that are out of service

A device you have taken out of service (Take out of Service on its card) raises no alerts and no notifications while it is stored, and it returns to normal alerting by itself as soon as it answers ping again. See [[crearack—monitoring—que-es-observatory]].

Troubleshooting

ProblemLikely causeFix
Alert never firesMonitor for that metric is offEnable ping/SNMP/HTTP on the target — rules only evaluate metrics that are being collected
Alert seems delayedThe Sustained for duration is doing its jobThe condition must hold continuously for the configured seconds before firing; lower it (or set 0) if you want instant alerts
Alert fires constantlyThreshold below the device’s normal baselineWatch the chart for a day, then set the threshold above the observed baseline — and give it a sustain duration so spikes don’t count
Events pile up unresolvedThe condition is still trueEvents close by themselves when the condition clears; if one stays open, the device is still failing the rule — fix it or adjust the threshold

Véase también

Subir