CreaRack-SL

Agent Fleet Manager

What it is

The Agent Fleet Manager shows every Local Agent registered for your organization and lets you control which one does the monitoring work. You can install the Agent on several machines; exactly one is Primary at a time — it runs Sentinel monitoring and receives jobs like Deep Discovery. The rest are Secondary: connected, idle, and ready to take over.

Because the Primary runs monitoring 24/7, it must live on a machine that stays powered on — a server or an always-on PC, not a laptop that gets shut down at night. If every Agent goes offline, monitoring stops. See [[crearack—terminal—local-agent]] → Where to install it.

You can also see the fleet at a glance in the Agent Fleet widget on the Observatory Overview tab.

How to open it

  1. Open Observatory and go to the Overview tab.
  2. In the Agent Fleet widget, click Agent (admins only).
  3. The Agent Fleet Manager modal lists each agent with its hostname, IP, version, role, Sentinel state, footprint and last-seen time. Refresh Status reloads the table.

What online and offline mean

  • Online (green dot) — the Agent holds a live connection to the cloud and has reported in within the last 10 minutes.
  • Offline (red dot) — the Agent has no connection to the cloud. It is not receiving jobs; if it is still running, it buffers metrics locally and syncs when it reconnects.
  • Sentinel: Active — that agent is currently running 24/7 Sentinel monitoring (normally only the Primary).
  • Footprint — the CPU and RAM the Agent process itself is using, the size of its local metrics database, and whether any metrics are still waiting to sync. Metrics piling up unsynced usually mean that agent is having trouble reaching CreaRack. Agents on older versions simply don’t show these figures. See [[crearack—terminal—local-agent]] for details.
  • Samples 1h — shown for the Primary running Sentinel: how many ping samples reached CreaRack in the last hour versus how many the polling cadence should have produced (for example 240/300). Green is healthy; amber (below half) or red (none) means the Agent is online but not actually producing data — restart it, and check its log if it happens again.
  • SNMP 1h — same idea for the SNMP loops: samples received in the last hour versus this Agent’s own hourly average over the past week (for example 12/40). It only appears once the Agent has some SNMP history, so devices that never answered SNMP do not count against it. Zero with a non-zero average means the SNMP loops are stuck even though ping still works — that is exactly how a “silent” Agent looks.
  • (clone?) next to an agent — that Agent’s credentials were created on a different machine (typically a PC or VM cloned with the Agent already paired). Two machines sharing one identity confuse the fleet. Install the Agent fresh on the copy and re-pair it; never clone a machine with a paired Agent.

Automatic vs manual failover

The Role Assignment selector at the top of the modal sets how the Primary role is assigned:

  • Auto (default) — the first agent to connect becomes Primary. If the Primary disconnects, the longest-connected online Secondary is promoted automatically and Sentinel monitoring moves to it. You get a toast like “HOSTNAME is now Primary” when a role changes.
  • Manual — roles never change on their own. New agents join as Secondary, and if the Primary goes offline, monitoring stops until you promote another agent yourself.

Use Manual when you want monitoring pinned to one specific machine (e.g. the only one with access to a management VLAN).

Either way, failover only helps if another Agent is online to take over. With a single Agent, or if all of them are on machines that get switched off, there is nothing to promote and monitoring simply stops. This is why at least one Agent belongs on an always-on machine.

Every admin gets an email when the Primary moves to another machine — whether you promoted it yourself or the automatic failover did it. The subject reads like [CreaRack] RACK-MINI is now the Primary agent (automatic failover), and the message names the machine that lost the role, the cause and who made the change. This matters most for the silent case: the machine that was watching your network shut down overnight and monitoring quietly moved somewhere else.

Two cases send no email, because monitoring did not move anywhere: the very first role an Agent gets when you install it (that is a new Agent joining, not the monitoring moving), and the same machine dropping its connection and reconnecting a few seconds later — a Wi-Fi blip, a laptop waking up. With a single Agent that is the normal daily rhythm, and it used to produce a stream of “is now the Primary” emails about a move that never happened. The Role changes history below still records every one of these transitions.

How to manage agents

Each row offers actions depending on its state (admins only):

  • Promote — make an online Secondary the new Primary. The current Primary is demoted automatically. A confirmation names both machines — “Sentinel monitoring will move from OFFICE-PC to RACK-MINI. Targets will be monitored from that machine.” — before anything changes, and the confirm button says Promote to Primary.
  • Demote — turn the Primary into a Secondary. Its confirmation spells out the consequence people usually don’t expect: no other agent takes over automatically, so your targets stop being monitored until you promote one yourself.
  • Delete — remove an offline agent from the fleet. You cannot delete an online Primary.
  • Reauth — issue fresh credentials for the agent and push them to it. Use this when an agent shows offline but is actually running.
  • Reset — clear the Primary’s Sentinel state (rate limits, cooldowns, circuit breakers). Shown only when Sentinel is active on that agent.

Agents that have been offline a long time (added 25-09-2026)

By default the table only shows agents that are online or were recently seen — an agent offline for a long time (a decommissioned machine, a test install) drops out of the normal view instead of piling up. It is not deleted for you: click Show inactive agents at the top of the modal to bring those rows back, review them, and Delete the ones you no longer need. Earlier versions of CreaRack deleted long-offline agents automatically; that stopped, because an agent that had been legitimately offline for a while (a machine powered off during a holiday, for example) could disappear from the fleet on its own.

Who changed the Primary, and when

Below the fleet table, the Role changes section keeps the history of every role change in your organization, newest first:

  • When the change happened.
  • Agent — the machine that changed role.
  • Change — from which role to which, and, when a promotion relieved another machine, which one it replaced.
  • Cause — Manual (someone clicked Promote or Demote), Automatic failover (the Primary went offline and another agent took over), or Automatic assignment (an agent connected and the system gave it a role).
  • By — the person who did it, or System when nobody did.

Use it to answer “why is monitoring running from that machine now?” without guessing. Automatic entries with no person behind them are normal — that is failover doing its job.

Troubleshooting

  • Agent is running but shows Offline — its credentials likely expired. Click Reauth; the new tokens are pushed to the agent on the same machine. If the browser cannot reach it, open the agent’s debug page to reconnect.
  • No agents registered — install the Local Agent first. While your organization has no Agent at all, the Observatory shows a banner at the top of every tab (“Monitoring needs the Local Agent”) with a Download Agent button, and the widget offers the same button. The banner disappears on its own once the first Agent is paired. See [[crearack—terminal—local-agent]].
  • Promote fails — you can only promote an agent that is currently online.
  • Two machines, monitoring on the wrong one — switch Role Assignment to Manual and Promote the machine you want.
  • You got an email saying the Primary changed and you didn’t do it — open Role changes: an entry with cause Automatic failover means the previous Primary went offline. Check whether that machine is meant to stay powered on.
  • Samples 1h is amber or red while the Primary shows Online — the Agent is alive but its polling loops are stuck. Restart the Agent on that machine; if it repeats, send us its log (agent.log next to the executable).
  • Monitoring keeps stopping at night — your Primary is on a machine that gets switched off. Move the Agent to an always-on machine, or keep a Secondary running on one.
  • A machine I decommissioned months ago isn’t in the list anymore, but I want to remove it for good — click Show inactive agents; it stopped appearing in the normal view because it has been offline a long time, but it is still there until you delete it.
  • [[crearack—terminal—local-agent]] — installing and running the Local Agent
  • [[crearack—monitoring—cns-sentinel]] — Sentinel Mode, the job the Primary agent runs
  • [[crearack—monitoring—que-es-observatory]] — the Observatory module
  • [[crearack—monitoring—deep-discovery]] — another job dispatched to the Primary agent

Véase también

  • [[crearack—terminal—local-agent]]
  • [[crearack—monitoring—cns-sentinel]]
  • [[crearack—monitoring—que-es-observatory]]
  • [[crearack—monitoring—deep-discovery]]