Volver a la wiki

Incident management (ITSM)

What it is

The ITSM tab of Network Observatory turns CNS insights into managed incidents. It tracks how fast your team reacts (SLAs), where alerts get sent (notification channels), what happens when nobody responds (escalation), and what your network keeps doing wrong (incident groups, recurring patterns, known issues).

Open Observatory and click the ITSM tab. The top shows an analytics dashboard; below it, collapsible sections you expand by clicking their headers.

How to read the analytics dashboard

The dashboard covers the last 30 days:

An insight counts as resolved when it is acknowledged (manually or auto-resolved after the device recovers) or its recommended fix is applied successfully. Insights that expire untouched do not count as resolved — they lower the resolution rate on purpose, so the numbers distinguish what was handled from what was left to die.

How to configure SLA policies

  1. Expand SLA Policies.
  2. First time: click Initialize Default Policies. This creates one policy per risk level — HIGH (ack 15 min / resolve 60 min), MEDIUM (30/240), LOW (60/480).
  3. Edit the acknowledge and resolve minutes per risk level and click Save.

Compliance against these targets feeds the dashboard above. When an incident misses a target, ITSM marks the breach and, if the deadline passed within the last hour, sends a notification through your channels (events sla_ack_breached and sla_resolve_breached), with the same rate limit and maintenance-window silencing as any other notification.

How to set up notification channels

  1. Expand Notification Channels and click New Channel.
  2. Pick a type: Email, Webhook, Slack or Microsoft Teams, and fill in the address or URL. Any other type is rejected when saving.
  3. Choose which risk levels the channel receives (leave all unchecked to receive everything) and click Save.
  4. Use Test on a saved channel to send a test notification, Edit to change it, Delete to remove it.

To verify channels actually fire, check the Notification Log section below — see next.

How to read the Notification Log

Expand Notification Log to see the most recent deliveries across every channel: Time, Channel, Event, Status (sent, failed or retrying) and Detail. It shows the last 50 sends.

Failed deliveries record only the error type or a short reason in Detail (for example ConnectionError, HTTPError 503 or maintenance check unavailable) — never the webhook URL or any other secret. Entries older than 90 days are purged nightly.

How escalation works

Expand Escalation Policies to see policies per risk level. Each policy defines levels like “L1: after 15m → NOC Email” — if an incident is not handled within the delay, the next level’s channel is notified. Policies are listed with a Delete action in the UI.

How to use maintenance windows

A maintenance window suppresses insights and/or notifications for planned work, so a firmware upgrade does not page anyone. Expand Maintenance Windows to see scheduled windows with their start/end times, affected targets and what they suppress. Active windows show an Active badge. Windows are created through the API: start and end go in ISO 8601 with a time zone offset (for example 2026-09-04T22:00:00+02:00), values without a zone are rejected, and the end must be after the start. A window with no targets applies to the whole organization and silences both insights and notifications.

Knowledge base: known issues and runbooks

Incident groups and recurring patterns

Troubleshooting

Véase también

Subir