What it is
The ITSM tab of Network Observatory turns CNS insights into managed incidents. It tracks how fast your team reacts (SLAs), where alerts get sent (notification channels), what happens when nobody responds (escalation), and what your network keeps doing wrong (incident groups, recurring patterns, known issues).
Open Observatory and click the ITSM tab. The top shows an analytics dashboard; below it, collapsible sections you expand by clicking their headers.
How to read the analytics dashboard
The dashboard covers the last 30 days:
- MTTA — mean time to acknowledge an insight.
- MTTR — mean time to resolve it.
- SLA Compliance — percentage of incidents handled within your SLA targets, plus an MTTA/MTTR trend chart. If no notification channel is enabled, the card shows “No escalation channel configured” instead of a percentage (a 0% would be misleading, since nobody could be alerted), with a Configure channels button that opens the Notification Channels section.
An insight counts as resolved when it is acknowledged (manually or auto-resolved after the device recovers) or its recommended fix is applied successfully. Insights that expire untouched do not count as resolved — they lower the resolution rate on purpose, so the numbers distinguish what was handled from what was left to die.
How to configure SLA policies
- Expand SLA Policies.
- First time: click Initialize Default Policies. This creates one policy per risk level — HIGH (ack 15 min / resolve 60 min), MEDIUM (30/240), LOW (60/480).
- Edit the acknowledge and resolve minutes per risk level and click Save.
Compliance against these targets feeds the dashboard above. When an incident misses a target, ITSM marks the breach and, if the deadline passed within the last hour, sends a notification through your channels (events sla_ack_breached and sla_resolve_breached), with the same rate limit and maintenance-window silencing as any other notification.
How to set up notification channels
- Expand Notification Channels and click New Channel.
- Pick a type: Email, Webhook, Slack or Microsoft Teams, and fill in the address or URL. Any other type is rejected when saving.
- Choose which risk levels the channel receives (leave all unchecked to receive everything) and click Save.
- Use Test on a saved channel to send a test notification, Edit to change it, Delete to remove it.
To verify channels actually fire, check the Notification Log section below — see next.
How to read the Notification Log
Expand Notification Log to see the most recent deliveries across every channel: Time, Channel, Event, Status (sent, failed or retrying) and Detail. It shows the last 50 sends.
Failed deliveries record only the error type or a short reason in Detail (for example ConnectionError, HTTPError 503 or maintenance check unavailable) — never the webhook URL or any other secret. Entries older than 90 days are purged nightly.
How escalation works
Expand Escalation Policies to see policies per risk level. Each policy defines levels like “L1: after 15m → NOC Email” — if an incident is not handled within the delay, the next level’s channel is notified. Policies are listed with a Delete action in the UI.
How to use maintenance windows
A maintenance window suppresses insights and/or notifications for planned work, so a firmware upgrade does not page anyone. Expand Maintenance Windows to see scheduled windows with their start/end times, affected targets and what they suppress. Active windows show an Active badge. Windows are created through the API: start and end go in ISO 8601 with a time zone offset (for example 2026-09-04T22:00:00+02:00), values without a zone are rejected, and the end must be after the start. A window with no targets applies to the whole organization and silences both insights and notifications.
Knowledge base: known issues and runbooks
- Known Issues — documented recurring problems with their resolution. When CNS generates a new insight, it matches it against known issues to enrich the diagnosis. Each card shows how many times the issue was seen.
- Runbooks — step-by-step procedures, optionally tied to a vendor and to trigger types. When an insight matches a runbook’s triggers, CNS suggests it on the insight. Cards show step count and usage count.
Incident groups and recurring patterns
- Incident Groups — related insights are grouped automatically by the correlation engine (e.g. several devices in one subnet going down together). A group resolves itself as soon as its last insight is closed (acknowledged, fixed or expired) — and reopens automatically if a new related insight is correlated into it. You can still click Resolve to close a whole group by hand at any time. Acknowledge All on a group handles up to 500 pending insights per click and tells you how many remain.
- Recurring Patterns — a daily job (03:00) looks for repeating incidents on the same target and reports them with a confidence score. Acknowledge a pattern once you have seen it, or Delete it. Patterns that have not recurred for over 30 days are removed automatically; if the cycle comes back, the detector recreates them.
Troubleshooting
- Dashboard shows ”—” for MTTA/MTTR — there are not enough acknowledged/resolved insights in the last 30 days yet.
- Resolution rate looks low despite active work — insights that expired without anyone acknowledging them count against the rate. Acknowledge what you triage, even briefly.
- No incident groups listed — normal; groups are only auto-created when the correlation engine links related insights, and resolved groups leave the default view once closed.
- Channel saved but nothing arrives — click Test and check the Notification Log; for webhooks, verify the URL is reachable from the internet.
Related
- [[crearack—monitoring—que-es-observatory]] — the Observatory module ITSM lives in
- [[crearack—monitoring—cns-sentinel]] — the AI insights that ITSM manages as incidents
- [[crearack—monitoring—alertas]] — threshold alerts, upstream of the incident cycle
- [[crearack—monitoring—fleet-manager]] — the Agents that detect anomalies in Sentinel Mode
Véase también
- [[crearack—monitoring—que-es-observatory]]
- [[crearack—monitoring—cns-sentinel]]
- [[crearack—monitoring—alertas]]
- [[crearack—monitoring—fleet-manager]]
Referenciado desde
- AI Insights (CNS) and Sentinel Mode
- Cola de auditoría C (task #286): ITSM, SLA, notificaciones y tareas de fondo — marcar y notificar en una sola transacción, ventanas de mantenimiento con fail-open, secretos en el log
- Guía de Troubleshooting con CNS/ITSM
- Guía ITSM — CreaRack Network Sentinel
- Monitoring alerts
- What is Network Observatory