Monitoring & Observability
What a node can tell you about itself, and which source to reach for. There are six, and picking the wrong one is the usual reason a question feels hard to answer.
| Source | Answers | Shape |
|---|---|---|
| Status | Is this node healthy right now? | One snapshot |
| Metrics | How much of what, right now? | Counters and gauges |
| History | How did it look over the last hour, day, month or year? | Graph data the node keeps itself |
| Flows | Who is actually talking to whom? | Live connections |
| Alerts | Tell me when something crosses a line | Push |
| Audit log | Who changed what, from where, and when? | A record of every change |
Logs sit underneath all of them — see Operations for reading them.
Status: is this node healthy?
The first command to run, and usually the only one needed:
cenvero-str-ctl status
It reports whether the agent is running, the license state, packet processing, and the bridges. When something is wrong this is where it shows up first.
Health checks run on their own schedule and record their results:
cenvero-str-ctl heal status # latest result for every check
cenvero-str-ctl heal check # force a run now, don't wait for the schedule
heal status is the better command when you want to know whether a problem is
recurring rather than whether it is happening this second.
Metrics: how much of what
cenvero-str-ctl metrics
cenvero-str-ctl metrics --format json
A snapshot of the node's counters. The JSON form is what you scrape into whatever
you already use for graphing — the same data is available over the API at
GET /api/v1/metrics for a collector that cannot run commands on the node.
The node itself. Measured every 10 seconds:
| Metric | What it is |
|---|---|
stratum_node_cpu_busy_percent | Share of all CPU time spent busy over the last 10 seconds, 0–100 |
stratum_node_cpu_iowait_percent | Share spent waiting for storage |
stratum_node_load1, …_load5, …_load15 | Load average over one, five and fifteen minutes |
stratum_node_memory_total_bytes, …_memory_used_bytes, …_memory_available_bytes | Memory: in use means total less what is available for new work |
stratum_node_swap_total_bytes, …_swap_used_bytes | Swap space |
stratum_node_filesystem_size_bytes, …_filesystem_used_bytes | The filesystems the node depends on (filesystem label: root; data, where the node keeps its state; pool, where machines' disks live, on a node that has one), as df counts them |
stratum_node_network_receive_bytes_total, …_transmit_bytes_total | Per interface the node manages (interface label): its bridges, network cards, bonds and overlays |
Virtual machines. Every machine on the node has its own series, labelled
with its id (vm), name and tenant, so you can graph and bill per machine or per
customer. They are refreshed every 15 seconds:
| Metric | What it is |
|---|---|
stratum_vm_up | 1 while the machine runs |
stratum_vm_down | 1 while it should be running and is not (crashed, stopped or in error) |
stratum_vm_cpu_seconds_total | CPU time it has used |
stratum_vm_vcpus | vCPUs it runs with |
stratum_vm_memory_bytes, stratum_vm_memory_max_bytes | Memory it has now, and may grow to without a restart |
stratum_vm_memory_resident_bytes | Host memory it actually occupies |
stratum_vm_disk_read_bytes_total, …_written_bytes_total, …_reads_total, …_writes_total | Per disk (disk label: vda, vdb…) |
stratum_vm_disk_allocation_bytes, stratum_vm_disk_capacity_bytes | Space a disk takes on the node, and its size |
stratum_vm_network_receive_bytes_total, …_transmit_bytes_total, …_packets_total, …_errors_total, …_drops_total | Per interface (interface and mac labels), from the machine's side: receive is what reached the machine |
A machine that does not run has only stratum_vm_up and stratum_vm_down; its
series disappear when it is deleted.
For traffic volume specifically — how much a tenant or endpoint used, rather than how the node is behaving — use accounting instead, which is what billing reads:
cenvero-str-ctl bandwidth list # configured limits
cenvero-str-ctl quota list # volume caps and consumption
History: the node's graphs
Metrics are the figures of this moment. The node also keeps their history for graphs — of itself, of each interface it manages, and of each virtual machine — without any collector of yours:
View (timeframe) | One point every | Covers |
|---|---|---|
live | 10 seconds (15 for a machine) | The last hour (hour and a half) |
hour | minute | The last hour |
day | minute | The last 24 hours |
week | 10 minutes | The last 7 days |
month | hour | The last 31 days |
year | 12 hours | The last year |
Each point holds the average of what was measured in it and the maximum
(cf=avg or cf=max), so a one-minute spike still shows on the year's graph.
curl -k "$NODE/api/v1/rrd/node?timeframe=day" -H "Authorization: Bearer $TOKEN"
curl -k "$NODE/api/v1/rrd/vms/vm-1a2b3c4d?timeframe=week&cf=max" -H "Authorization: Bearer $TOKEN"
curl -k "$NODE/api/v1/rrd/interfaces/cnv-nic-0?timeframe=live" -H "Authorization: Bearer $TOKEN"
| Series | Node | Interface | Machine |
|---|---|---|---|
| CPU | cpu and iowait, percent of all cores | cpu, percent of the machine's own vCPUs | |
| Load | load1, load5, load15 | ||
| Memory | mem_used, mem_total, swap_used, swap_total | mem_used (host memory it occupies), mem_total (memory the guest has) | |
| Disk | root_used/root_total, data_used/data_total, pool_used/pool_total | disk_read, disk_write, bytes per second | |
| Network | net_rx, net_tx: bytes per second on the node's network cards | rx, tx, bytes per second | net_rx, net_tx, bytes per second from the machine's side |
A gap — the node was off, the machine was stopped — is null, never 0. A
tenant reads its own machines' graphs through the portal
(GET /api/v1/tenant/{id}/vms/{vmid}/rrd); the node's graphs are the
operator's.
Disk use is fixed. Each graph's history takes the same space from its first
minute to its last: about 0.57 MB for the node, 94 KB per interface and
0.25 MB per virtual machine — about 78 MB for a node with 300 machines. A
machine's history is removed with the machine; an interface's, seven days after
the interface is gone. It lives in /var/lib/cenvero-str/rrd/.
Flows: who is talking to whom
Metrics tell you a link is busy. Flows tell you what is making it busy.
cenvero-str-ctl flow list # live connections
cenvero-str-ctl flow stats # aggregate view
This is the tool for "the network is slow", "is this rule doing anything", and "what is this host actually connecting to". Over the API, flows can also be exported for offline analysis.
Two things to know. Flows show conversations the data plane is currently tracking, so a connection that finished is gone — this is a live view, not a history. And a long-established connection appears even after you have tightened a rule against it, because existing conversations survive a policy change until they end or are flushed. That surprise is covered in Firewall.
Alerts: tell me when something happens
Everything above is you asking. Alerts are the node telling you.
cenvero-str-ctl alert condition list # what is being watched
cenvero-str-ctl alert list firing # what is firing now (also: acknowledged, resolved)
cenvero-str-ctl alert history # every alert kept, newest first
cenvero-str-ctl alert status # is alerting itself working?
cenvero-str-ctl alert ack <id> # acknowledge a firing alert
Conditions define what to watch for. Actions define what happens when one fires, so an alert can reach a system you already run rather than waiting to be noticed.
What a condition can watch
metric_type | The reading | Target |
|---|---|---|
bandwidth | Bytes per second the node receives, every 5 seconds | global |
pps | Packets per second the node receives, every 5 seconds | global |
vm_down | 1 while a virtual machine that should run does not, else 0; every 15 seconds and at the moment of a crash | the machine's id |
threat | How severe an intrusion-detection hit is (1 means exactly at the detection threshold) | the source address |
A condition compares the reading with threshold using operator (gt, lt
or eq). target is global — every target — unless you name one. Nothing on
a node measures connections or quota, so a condition on either is refused; a
condition like that from an earlier version is listed with a warning saying it
never fires, and can be removed.
A virtual machine that goes down. The vm_down metric is 1 while a machine
that should be running is not — it crashed (also when it is restarted
automatically), or it is stopped or in error — and 0 otherwise. To be told:
cenvero-str-ctl alert condition add '{"metric_type":"vm_down","operator":"gt","threshold":0}'
That watches every machine; set "target" to a machine's id to watch one. A
machine you stop yourself, or that a tenant's suspension stopped, is not down.
How an alert fires and resolves
- Each target is its own alert. A condition on
globalthat watches every machine raises one alert per machine that goes down, and each resolves on its own. duration_secs: the condition must hold that long. With"duration_secs": 60, the readings must stay past the threshold for a full minute before the alert fires. A reading back within the threshold, or a gap of more than a minute with no reading at all, starts the minute again. Athreathit is a single event, so a threat condition takes no duration.- One alert while it lasts. While the condition keeps holding, the alert stays open; it does not fire again.
- It resolves by itself. The alert resolves when a reading is back within the threshold, and also when nothing has reported the target for a while: 2 minutes for bandwidth, packets and machines (a deleted machine, for example), and 10 minutes without a new detection for a threat. Removing a condition resolves its open alerts.
cooldown_secs(default 300) is the shortest time between two alerts for the same condition and target, so a value hovering around the threshold does not alert over and over.- Acknowledging (
alert ack) records who is on it. The alert stays open and still resolves when the condition clears.
Every alert records metric_type, target, the value that fired it,
fired_at, and once resolved resolved_at and resolved_reason:
resolved_reason | Meaning |
|---|---|
cleared | A reading was back within the threshold |
not_reported | Nothing reported the target for the time above |
condition_removed | The condition was removed |
superseded | An alert left open by an earlier version, replaced by a newer one for the same condition and target |
manual | Resolved by hand |
History is kept for 30 days. A resolved alert is removed 30 days after it resolved, and at most the 10,000 most recent alerts are kept (the oldest resolved ones go first). Open alerts are never removed.
Where an alert goes
alert action add <condition_id> <type> [config] attaches an action; each type
can be attached once per condition, and adding it again replaces it.
logwrites the fire and the resolution to the agent's log.websocketpublishes them on the live event stream asalert_firedandalert_resolvedevents in theALERTcategory.webhookPOSTs them to the URL given asconfig:
cenvero-str-ctl alert action add <condition_id> webhook https://hooks.example.net/stratum
The answer includes a secret. It is shown only this once — store it with
your receiver; adding the webhook action again gives it a new one. Every delivery
is signed with it exactly like an event webhook: the body is the
alert as JSON (its state is firing or resolved), and these headers come with
it:
| Header | Value |
|---|---|
X-Stratum-Signature | sha256= followed by the hex HMAC-SHA256 of the raw body, keyed by the secret |
X-Stratum-Event | ALERT |
X-Stratum-Delivery | The alert id followed by -firing or -resolved |
X-Stratum-Timestamp | Send time, Unix seconds (UTC) |
User-Agent | cenvero-stratum-alert/1 |
To verify one, recompute the HMAC over the exact bytes received and compare it in
constant time with X-Stratum-Signature:
printf '%s' "$body" | openssl dgst -sha256 -hmac "$SECRET" # prepend "sha256="
A delivery that fails (no answer, or an answer other than 2xx) is tried up to
three times in a few seconds. X-Stratum-Delivery is the same on every try, so a
receiver that sees it twice can ignore the repeat.
The webhook URL must be http or https and reach a public address. The
node refuses its own loopback address, private and shared (carrier-grade NAT)
ranges, link-local addresses including the cloud metadata address, multicast and
unspecified addresses — when the action is added, and again on every delivery.
A receiver that answers with a redirect is not followed. Deliveries connect
directly: a proxy set in the agent's environment is not used. A receiver on a
private network therefore cannot be a webhook target; read alerts over the API or
the event stream instead.
A webhook action added before deliveries were signed keeps working unsigned and
is listed with a warning; add it again to get a secret.
Check alert status occasionally. It reports whether action dispatch is
succeeding, with the last error. Alerting that is configured but silently failing
to deliver is worse than no alerting, because it is mistaken for quiet.
Who changed what: the audit log
Every node keeps an audit log of its own. Every change made on it is recorded —
through the API, the web console, the tenant portal or cenvero-str-ctl — as it
happens:
- who made it: the credential's name (
api token,api key ak-12ab,tenant key tk-… (tenant t-acme), orrootfor the local command line) — never the credential itself; - from where: the client's address, and the web-console session it came through;
- what: the action (
vms.start,tenant.keys.create,apikeys.mint, …), the exact API route or command, and the object (/vms/vm-1a2b3c4d), with the id of anything a create made; - how it ended:
success,failure, orrefused— with what refused it:authentication,tenant scope,licenceorplan.
It also records console sign-ins, refused sign-ins and sign-outs, and every
machine console opened, refused and closed. A change that became a
task names it (task_id), and the task's own start and end
are records too: task.started, then task.succeeded, task.failed,
task.cancelled, or task.interrupted when the node restarted while it ran,
under the name of whoever started it. Reads are not recorded. **No record
holds a key, token, password, private key or the body of a request.**
cenvero-str-ctl audit list # newest first
cenvero-str-ctl audit list --action vms. --outcome refused # filtered
cenvero-str-ctl audit export --output audit.jsonl # everything, as JSON lines
cenvero-str-ctl audit verify # has anything been changed?
Over the API: GET /api/v1/audit, /api/v1/audit/export and
/api/v1/audit/verify (see the Management API Reference). A tenant
reads the records of what its own keys did at GET /api/v1/tenant/{id}/audit,
without the node's details.
Tamper-evident. Each record carries a fingerprint (SHA-256) of the record
before it, so a record changed, removed or slipped in afterwards breaks the
chain, and audit verify names the first record where it breaks. Anyone with
root on the node could still rewrite the whole log; to catch that too, keep the
number and hash audit verify prints somewhere else, and check it later with
audit verify --checkpoint <number>:<hash> — or export the log off the node
regularly. The export carries every record's fingerprints, so it can be checked
on its own.
Kept for 90 days, or the newest 1,000,000 records, whichever is fewer.
Removing old records does not break verification: each removal leaves an
audit.pruned record saying where the log now starts, and verify checks the
start against it.
The agent's log still carries its audit=true lines as before, for anyone who
already collects it.
A dashboard in one call
GET /api/v1/dashboard returns one snapshot for a status screen: the host
(hostname, uptime, CPU cores, memory), each network card with its byte counters,
traffic totals and active flows, the packets dropped by policy, the firewall
rules in force right now, intrusion-detection hits, firing alerts, and the
cluster summary. Every figure comes from the node itself; a figure the node
cannot read is null, never a made-up 0. See the
Management API Reference.
Live events
For a continuous feed rather than polling, the agent publishes events over a WebSocket on port 7072, grouped into categories (traffic, bandwidth, security, DHCP, DNS, network, alerts, system, load balancing, and clustering) so a consumer can subscribe to only what it cares about. See the Management API Reference.
What to watch
If you are setting up monitoring for the first time, start here:
- Agent up, and license not frozen. A frozen license blocks changes silently from a traffic standpoint — everything keeps flowing, so nothing looks wrong until a change fails. Watch the license state, not just the process.
- Health-check results. Repeated failures of one check are the earliest warning of most problems.
- Certificate expiry. See TLS/SSL & License Operations.
- Alert dispatch failures. As above — verify the alerting path works.
- The audit log verifies. Run
cenvero-str-ctl audit verifyon a schedule (it exits non-zero when the chain is broken) and keep the hash it prints.
Reaching a node from a collector
The CLI works over a local socket and needs no network, which is why it keeps working when the API is off. A remote collector needs the API turned on, which means a token and an address allowlist — see Security Model before exposing it.
Where to go next
- Operations — logs, health checks, and troubleshooting.
- Management API Reference — the endpoints behind these commands.
- Security Model — before exposing the API to a collector.