Exclusive Access · Invitation Only

Monitoring & Observability

What a node can tell you about itself, and which source to reach for. There are six, and picking the wrong one is the usual reason a question feels hard to answer.

SourceAnswersShape
StatusIs this node healthy right now?One snapshot
MetricsHow much of what, right now?Counters and gauges
HistoryHow did it look over the last hour, day, month or year?Graph data the node keeps itself
FlowsWho is actually talking to whom?Live connections
AlertsTell me when something crosses a linePush
Audit logWho changed what, from where, and when?A record of every change

Logs sit underneath all of them — see Operations for reading them.

Status: is this node healthy?

The first command to run, and usually the only one needed:

cenvero-str-ctl status

It reports whether the agent is running, the license state, packet processing, and the bridges. When something is wrong this is where it shows up first.

Health checks run on their own schedule and record their results:

cenvero-str-ctl heal status      # latest result for every check
cenvero-str-ctl heal check       # force a run now, don't wait for the schedule

heal status is the better command when you want to know whether a problem is recurring rather than whether it is happening this second.

Metrics: how much of what

cenvero-str-ctl metrics
cenvero-str-ctl metrics --format json

A snapshot of the node's counters. The JSON form is what you scrape into whatever you already use for graphing — the same data is available over the API at GET /api/v1/metrics for a collector that cannot run commands on the node.

The node itself. Measured every 10 seconds:

MetricWhat it is
stratum_node_cpu_busy_percentShare of all CPU time spent busy over the last 10 seconds, 0–100
stratum_node_cpu_iowait_percentShare spent waiting for storage
stratum_node_load1, …_load5, …_load15Load average over one, five and fifteen minutes
stratum_node_memory_total_bytes, …_memory_used_bytes, …_memory_available_bytesMemory: in use means total less what is available for new work
stratum_node_swap_total_bytes, …_swap_used_bytesSwap space
stratum_node_filesystem_size_bytes, …_filesystem_used_bytesThe filesystems the node depends on (filesystem label: root; data, where the node keeps its state; pool, where machines' disks live, on a node that has one), as df counts them
stratum_node_network_receive_bytes_total, …_transmit_bytes_totalPer interface the node manages (interface label): its bridges, network cards, bonds and overlays

Virtual machines. Every machine on the node has its own series, labelled with its id (vm), name and tenant, so you can graph and bill per machine or per customer. They are refreshed every 15 seconds:

MetricWhat it is
stratum_vm_up1 while the machine runs
stratum_vm_down1 while it should be running and is not (crashed, stopped or in error)
stratum_vm_cpu_seconds_totalCPU time it has used
stratum_vm_vcpusvCPUs it runs with
stratum_vm_memory_bytes, stratum_vm_memory_max_bytesMemory it has now, and may grow to without a restart
stratum_vm_memory_resident_bytesHost memory it actually occupies
stratum_vm_disk_read_bytes_total, …_written_bytes_total, …_reads_total, …_writes_totalPer disk (disk label: vda, vdb…)
stratum_vm_disk_allocation_bytes, stratum_vm_disk_capacity_bytesSpace a disk takes on the node, and its size
stratum_vm_network_receive_bytes_total, …_transmit_bytes_total, …_packets_total, …_errors_total, …_drops_totalPer interface (interface and mac labels), from the machine's side: receive is what reached the machine

A machine that does not run has only stratum_vm_up and stratum_vm_down; its series disappear when it is deleted.

For traffic volume specifically — how much a tenant or endpoint used, rather than how the node is behaving — use accounting instead, which is what billing reads:

cenvero-str-ctl bandwidth list          # configured limits
cenvero-str-ctl quota list              # volume caps and consumption

History: the node's graphs

Metrics are the figures of this moment. The node also keeps their history for graphs — of itself, of each interface it manages, and of each virtual machine — without any collector of yours:

View (timeframe)One point everyCovers
live10 seconds (15 for a machine)The last hour (hour and a half)
hourminuteThe last hour
dayminuteThe last 24 hours
week10 minutesThe last 7 days
monthhourThe last 31 days
year12 hoursThe last year

Each point holds the average of what was measured in it and the maximum (cf=avg or cf=max), so a one-minute spike still shows on the year's graph.

curl -k "$NODE/api/v1/rrd/node?timeframe=day" -H "Authorization: Bearer $TOKEN"
curl -k "$NODE/api/v1/rrd/vms/vm-1a2b3c4d?timeframe=week&cf=max" -H "Authorization: Bearer $TOKEN"
curl -k "$NODE/api/v1/rrd/interfaces/cnv-nic-0?timeframe=live" -H "Authorization: Bearer $TOKEN"
SeriesNodeInterfaceMachine
CPUcpu and iowait, percent of all corescpu, percent of the machine's own vCPUs
Loadload1, load5, load15
Memorymem_used, mem_total, swap_used, swap_totalmem_used (host memory it occupies), mem_total (memory the guest has)
Diskroot_used/root_total, data_used/data_total, pool_used/pool_totaldisk_read, disk_write, bytes per second
Networknet_rx, net_tx: bytes per second on the node's network cardsrx, tx, bytes per secondnet_rx, net_tx, bytes per second from the machine's side

A gap — the node was off, the machine was stopped — is null, never 0. A tenant reads its own machines' graphs through the portal (GET /api/v1/tenant/{id}/vms/{vmid}/rrd); the node's graphs are the operator's.

Disk use is fixed. Each graph's history takes the same space from its first minute to its last: about 0.57 MB for the node, 94 KB per interface and 0.25 MB per virtual machine — about 78 MB for a node with 300 machines. A machine's history is removed with the machine; an interface's, seven days after the interface is gone. It lives in /var/lib/cenvero-str/rrd/.

Flows: who is talking to whom

Metrics tell you a link is busy. Flows tell you what is making it busy.

cenvero-str-ctl flow list      # live connections
cenvero-str-ctl flow stats     # aggregate view

This is the tool for "the network is slow", "is this rule doing anything", and "what is this host actually connecting to". Over the API, flows can also be exported for offline analysis.

Two things to know. Flows show conversations the data plane is currently tracking, so a connection that finished is gone — this is a live view, not a history. And a long-established connection appears even after you have tightened a rule against it, because existing conversations survive a policy change until they end or are flushed. That surprise is covered in Firewall.

Alerts: tell me when something happens

Everything above is you asking. Alerts are the node telling you.

cenvero-str-ctl alert condition list      # what is being watched
cenvero-str-ctl alert list firing         # what is firing now (also: acknowledged, resolved)
cenvero-str-ctl alert history             # every alert kept, newest first
cenvero-str-ctl alert status              # is alerting itself working?
cenvero-str-ctl alert ack <id>            # acknowledge a firing alert

Conditions define what to watch for. Actions define what happens when one fires, so an alert can reach a system you already run rather than waiting to be noticed.

What a condition can watch

metric_typeThe readingTarget
bandwidthBytes per second the node receives, every 5 secondsglobal
ppsPackets per second the node receives, every 5 secondsglobal
vm_down1 while a virtual machine that should run does not, else 0; every 15 seconds and at the moment of a crashthe machine's id
threatHow severe an intrusion-detection hit is (1 means exactly at the detection threshold)the source address

A condition compares the reading with threshold using operator (gt, lt or eq). target is global — every target — unless you name one. Nothing on a node measures connections or quota, so a condition on either is refused; a condition like that from an earlier version is listed with a warning saying it never fires, and can be removed.

A virtual machine that goes down. The vm_down metric is 1 while a machine that should be running is not — it crashed (also when it is restarted automatically), or it is stopped or in error — and 0 otherwise. To be told:

cenvero-str-ctl alert condition add '{"metric_type":"vm_down","operator":"gt","threshold":0}'

That watches every machine; set "target" to a machine's id to watch one. A machine you stop yourself, or that a tenant's suspension stopped, is not down.

How an alert fires and resolves

  • Each target is its own alert. A condition on global that watches every machine raises one alert per machine that goes down, and each resolves on its own.
  • duration_secs: the condition must hold that long. With "duration_secs": 60, the readings must stay past the threshold for a full minute before the alert fires. A reading back within the threshold, or a gap of more than a minute with no reading at all, starts the minute again. A threat hit is a single event, so a threat condition takes no duration.
  • One alert while it lasts. While the condition keeps holding, the alert stays open; it does not fire again.
  • It resolves by itself. The alert resolves when a reading is back within the threshold, and also when nothing has reported the target for a while: 2 minutes for bandwidth, packets and machines (a deleted machine, for example), and 10 minutes without a new detection for a threat. Removing a condition resolves its open alerts.
  • cooldown_secs (default 300) is the shortest time between two alerts for the same condition and target, so a value hovering around the threshold does not alert over and over.
  • Acknowledging (alert ack) records who is on it. The alert stays open and still resolves when the condition clears.

Every alert records metric_type, target, the value that fired it, fired_at, and once resolved resolved_at and resolved_reason:

resolved_reasonMeaning
clearedA reading was back within the threshold
not_reportedNothing reported the target for the time above
condition_removedThe condition was removed
supersededAn alert left open by an earlier version, replaced by a newer one for the same condition and target
manualResolved by hand

History is kept for 30 days. A resolved alert is removed 30 days after it resolved, and at most the 10,000 most recent alerts are kept (the oldest resolved ones go first). Open alerts are never removed.

Where an alert goes

alert action add <condition_id> <type> [config] attaches an action; each type can be attached once per condition, and adding it again replaces it.

  • log writes the fire and the resolution to the agent's log.
  • websocket publishes them on the live event stream as alert_fired and alert_resolved events in the ALERT category.
  • webhook POSTs them to the URL given as config:
cenvero-str-ctl alert action add <condition_id> webhook https://hooks.example.net/stratum

The answer includes a secret. It is shown only this once — store it with your receiver; adding the webhook action again gives it a new one. Every delivery is signed with it exactly like an event webhook: the body is the alert as JSON (its state is firing or resolved), and these headers come with it:

HeaderValue
X-Stratum-Signaturesha256= followed by the hex HMAC-SHA256 of the raw body, keyed by the secret
X-Stratum-EventALERT
X-Stratum-DeliveryThe alert id followed by -firing or -resolved
X-Stratum-TimestampSend time, Unix seconds (UTC)
User-Agentcenvero-stratum-alert/1

To verify one, recompute the HMAC over the exact bytes received and compare it in constant time with X-Stratum-Signature:

printf '%s' "$body" | openssl dgst -sha256 -hmac "$SECRET"   # prepend "sha256="

A delivery that fails (no answer, or an answer other than 2xx) is tried up to three times in a few seconds. X-Stratum-Delivery is the same on every try, so a receiver that sees it twice can ignore the repeat.

The webhook URL must be http or https and reach a public address. The node refuses its own loopback address, private and shared (carrier-grade NAT) ranges, link-local addresses including the cloud metadata address, multicast and unspecified addresses — when the action is added, and again on every delivery. A receiver that answers with a redirect is not followed. Deliveries connect directly: a proxy set in the agent's environment is not used. A receiver on a private network therefore cannot be a webhook target; read alerts over the API or the event stream instead.

A webhook action added before deliveries were signed keeps working unsigned and is listed with a warning; add it again to get a secret.

Check alert status occasionally. It reports whether action dispatch is succeeding, with the last error. Alerting that is configured but silently failing to deliver is worse than no alerting, because it is mistaken for quiet.

Who changed what: the audit log

Every node keeps an audit log of its own. Every change made on it is recorded — through the API, the web console, the tenant portal or cenvero-str-ctl — as it happens:

  • who made it: the credential's name (api token, api key ak-12ab, tenant key tk-… (tenant t-acme), or root for the local command line) — never the credential itself;
  • from where: the client's address, and the web-console session it came through;
  • what: the action (vms.start, tenant.keys.create, apikeys.mint, …), the exact API route or command, and the object (/vms/vm-1a2b3c4d), with the id of anything a create made;
  • how it ended: success, failure, or refused — with what refused it: authentication, tenant scope, licence or plan.

It also records console sign-ins, refused sign-ins and sign-outs, and every machine console opened, refused and closed. A change that became a task names it (task_id), and the task's own start and end are records too: task.started, then task.succeeded, task.failed, task.cancelled, or task.interrupted when the node restarted while it ran, under the name of whoever started it. Reads are not recorded. **No record holds a key, token, password, private key or the body of a request.**

cenvero-str-ctl audit list                                   # newest first
cenvero-str-ctl audit list --action vms. --outcome refused   # filtered
cenvero-str-ctl audit export --output audit.jsonl            # everything, as JSON lines
cenvero-str-ctl audit verify                                 # has anything been changed?

Over the API: GET /api/v1/audit, /api/v1/audit/export and /api/v1/audit/verify (see the Management API Reference). A tenant reads the records of what its own keys did at GET /api/v1/tenant/{id}/audit, without the node's details.

Tamper-evident. Each record carries a fingerprint (SHA-256) of the record before it, so a record changed, removed or slipped in afterwards breaks the chain, and audit verify names the first record where it breaks. Anyone with root on the node could still rewrite the whole log; to catch that too, keep the number and hash audit verify prints somewhere else, and check it later with audit verify --checkpoint <number>:<hash> — or export the log off the node regularly. The export carries every record's fingerprints, so it can be checked on its own.

Kept for 90 days, or the newest 1,000,000 records, whichever is fewer. Removing old records does not break verification: each removal leaves an audit.pruned record saying where the log now starts, and verify checks the start against it.

The agent's log still carries its audit=true lines as before, for anyone who already collects it.

A dashboard in one call

GET /api/v1/dashboard returns one snapshot for a status screen: the host (hostname, uptime, CPU cores, memory), each network card with its byte counters, traffic totals and active flows, the packets dropped by policy, the firewall rules in force right now, intrusion-detection hits, firing alerts, and the cluster summary. Every figure comes from the node itself; a figure the node cannot read is null, never a made-up 0. See the Management API Reference.

Live events

For a continuous feed rather than polling, the agent publishes events over a WebSocket on port 7072, grouped into categories (traffic, bandwidth, security, DHCP, DNS, network, alerts, system, load balancing, and clustering) so a consumer can subscribe to only what it cares about. See the Management API Reference.

What to watch

If you are setting up monitoring for the first time, start here:

  • Agent up, and license not frozen. A frozen license blocks changes silently from a traffic standpoint — everything keeps flowing, so nothing looks wrong until a change fails. Watch the license state, not just the process.
  • Health-check results. Repeated failures of one check are the earliest warning of most problems.
  • Certificate expiry. See TLS/SSL & License Operations.
  • Alert dispatch failures. As above — verify the alerting path works.
  • The audit log verifies. Run cenvero-str-ctl audit verify on a schedule (it exits non-zero when the chain is broken) and keep the hash it prints.

Reaching a node from a collector

The CLI works over a local socket and needs no network, which is why it keeps working when the API is off. A remote collector needs the API turned on, which means a token and an address allowlist — see Security Model before exposing it.

Where to go next

↓ This page as JSON ↓ All documentation as JSON