Operations & Troubleshooting
Day-to-day operation of a Stratum node: managing its systemd units, finding the logs, turning individual network services on and off, changing settings, applying updates, and a runbook for the issues you are most likely to hit.
Everything here runs on the node itself as root. The operator tool,
cenvero-str-ctl, talks to the running agent over a local socket — see the
CLI Reference for the full command surface.
Systemd units
A node runs three systemd units, installed and enabled by the installer:
| Unit | What it does |
|---|---|
cenvero-stratum | The agent itself — configuration, APIs, and the node's packet processing. |
cenvero-str-watchdog | A small separate watchdog that monitors the agent and restarts it if it stops responding. It is bound to the agent (BindsTo), so it follows the agent's lifecycle. |
cenvero-str-network | A one-shot boot unit that brings up the host uplinks and the two bridges (cnv-mgmt-br0, cnv-user-br0) before the agent starts, so the interfaces are present when the agent attaches to them. |
Manage them with the usual systemctl verbs:
# Status and recent log lines for the agent
systemctl status cenvero-stratum
# Start / stop / restart the agent (the watchdog follows it)
sudo systemctl restart cenvero-stratum
sudo systemctl stop cenvero-stratum
sudo systemctl start cenvero-stratum
# The watchdog and the boot-time network unit
systemctl status cenvero-str-watchdog
systemctl status cenvero-str-network
Restarting the agent is safe: packet forwarding does not depend on the agent process, and the interfaces stay in place across the restart, so the brief restart does not tear down existing traffic. Stopping the agent leaves the forwarding rules and the bridges in place.
Do not disable cenvero-str-network while the node is in service — it owns the
host networking the agent attaches to.
Logs
The agent logs structured JSON to standard out, which under systemd is captured by the journal. That is the first place to look:
# Follow the agent's live log
journalctl -u cenvero-stratum -f
# Last 200 lines
journalctl -u cenvero-stratum -n 200
# Since a time, or only errors and worse
journalctl -u cenvero-stratum --since "1 hour ago"
journalctl -u cenvero-stratum -p err
# The watchdog has its own unit log
journalctl -u cenvero-str-watchdog -n 100
On-disk files live under /var/log/cenvero-str/ — most notably the installer
log written during install. The agent's self-healing service prunes old rotated
log files (older than 7 days) from this directory when disk space runs low, so it
is also where any rotated logs accumulate.
Turn up the detail when you need it (see Changing settings below):
sudo cenvero-str-ctl config set log_level debug
sudo systemctl restart cenvero-stratum
Valid levels are debug, info, warn, and error. Set it back to info when
you are done — debug is noisy.
Network services: on / off
The agent exposes several network services you can independently enable or
disable. View them and their live state with service status, and flip each one
with service <name> on|off:
# See every service: its switch, its address:port, and whether it is listening
cenvero-str-ctl service status
# Turn services off / on
sudo cenvero-str-ctl service rest off
sudo cenvero-str-ctl service metrics on
The toggleable services are:
| Service | Name | What it is |
|---|---|---|
| REST API | rest | The HTTPS management & operator/billing API. |
| gRPC API | grpc | The gRPC management API. |
| WebSocket API | websocket | The live-events / streaming API. |
| Metrics | metrics | The Prometheus /metrics endpoint. |
| DNS | dns | The built-in DNS server. |
| DHCP | dhcp | The built-in DHCP server. |
The verb-first form works too — service off rest, service on dns — and -h
on the group shows the help.
A few notes:
- A change takes effect on the next agent restart. After toggling a service, apply it with
sudo systemctl restart cenvero-stratum, then re-check withservice status. - The local control socket can never be disabled — it is how
cenvero-str-ctlreaches the agent (including the command you just ran). Trying to turn it off is refused with an explanation. - Disabling the REST API cuts off remote callers. The REST API also serves the operator/billing API (tenant suspend/resume/limit). Turning it off is allowed — it is your call — but the command warns you first; the local socket and the panel are unaffected.
Under the hood, service <name> on|off writes the matching off-switch into the
same operator local-overrides file that config set uses (see below), so a panel
re-sync never clobbers your choice.
Changing settings
config set adjusts a node's operational settings on a running node. It
records your change in an operator local-override file that takes precedence over
the panel-delivered config and survives a re-sync — it never edits the signed
node config itself.
# Inspect the active configuration (paths, ports, API settings, services)
cenvero-str-ctl config show
# Change a setting, then restart to apply it
sudo cenvero-str-ctl config set api_rate_limit 2000
sudo systemctl restart cenvero-stratum
The key names match what you see in config show, so the same name round-trips
between the two. Commonly-set keys include:
| Key | Meaning |
|---|---|
log_level | Agent log level: debug, info, warn, error |
api_bind_address | IP the REST/gRPC/WebSocket APIs listen on (e.g. 0.0.0.0 or a management IP) |
api_rate_limit / api_rate_burst | API requests per minute from each source IP address, and the burst allowance (0 rate disables the limiter) |
api_allowed_ips | Comma-separated IPs/CIDRs allowed to reach the API (empty = allow all) |
port_rest, port_grpc, port_websocket | API listen ports |
metrics_bind_addr | Prometheus metrics bind address host:port |
dns_listen_addr, dns_allowed_clients | Built-in DNS server bind address and allowed clients |
Run cenvero-str-ctl config set -h for the full whitelist and per-key help.
What you cannot set locally. Identity, credential and certificate fields —>node_id,license_server,api_token,gateway_shared_key, and the TLS paths — and thecluster_*keys are not settable withconfig setby design. They are panel-assigned or provisioned; a cluster is formed withcluster createand joined withcluster join --code, and a member's identity is issued by its cluster. Reaching for one returns a clear message telling you where it is managed instead (e.g. TLS material is handled bycenvero-str-ctl tls …, and the API token bycenvero-str-ctl api-token …).
Every config set change applies on the next restart. It is saved
immediately, but the running agent only picks it up when it restarts.
Running alongside Docker (or another host firewall)
Docker switches on a kernel setting that sends traffic crossing a Linux bridge through the host's firewall, and it sets that firewall to drop forwarded traffic by default. On a node that also runs Docker, that would drop traffic between your virtual machines on the workload bridge — even two of the same tenant — and traffic the node routes between the workload bridge and the internet.
The agent handles this by itself. Whenever the host passes bridged traffic through its firewall, it keeps rules that let exactly this traffic through:
- traffic between the workloads on the workload bridge;
- traffic the node routes between the workload bridge and its uplink — the interface named in
gateway_wan_interfaceand the interfaces of the node's default routes — in both directions.
The rules sit at the top of the DOCKER-USER chain (the chain Docker keeps for
your own rules), or at the top of FORWARD where there is no DOCKER-USER. The
agent does this for IPv4, and for IPv6 where that is passed to the firewall too.
Every 30 seconds it checks that the rules are still there and still match the
uplink, and it takes them out again if the setting is switched off.
Nothing else is opened. Traffic between the workload bridge and any other
interface is left to the host's firewall, as you and Docker set it up: Docker's
containers stay behind Docker's own rules, and another hypervisor's bridge or a
network on another card of the server stays closed to your virtual machines
unless you allow it yourself — in DOCKER-USER, for instance. The same applies if the node routes
workload traffic out through an interface other than its uplink (a second
uplink or an overlay): let it through yourself. Stratum's own protections —
tenant separation and address checks at every workload's port, and your
Stratum firewall rules for the traffic the node routes — still apply to what
the rules let through.
The rules carry the comment cenvero-stratum-workload-bridge; no other rule in
the host's firewall is touched. To see them:
sudo iptables -S DOCKER-USER | grep cenvero-stratum
sudo iptables -S FORWARD | grep cenvero-stratum
They stay in place while the agent is stopped, so running virtual machines keep their traffic. If you remove Stratum from a server, stop the agent first and then take the rules out:
sudo systemctl stop cenvero-stratum
sudo /usr/local/bin/cenvero-stratum -remove-host-firewall-rules
Updates
Updates are pull-based: the agent polls a signed manifest from the panel and, when a newer fully-verified release for its channel is available, downloads, verifies, and applies it on its own. You can check or trigger this manually:
# What's installed now
cenvero-str-ctl status
# Ask the agent to check the manifest for a newer release (no download)
cenvero-str-ctl update check
# Apply an available update now (otherwise it applies on the next poll)
sudo cenvero-str-ctl update apply
update check reports the current version, the latest on offer for your license
channel, and whether an update is available — without downloading anything. If the
panel is unreachable it says so plainly rather than failing hard. update apply
performs the pull, checksum + publisher-signature verification, an atomic
in-place swap, and a watchdog-supervised restart, with self-rollback if the
post-restart health check fails. See Upgrades for the full
model (channels, downgrade protection, and rollback).
A node in the Frozen license state keeps running its current version but cannot pull updates until the license is renewed.
Health checks
The agent continuously self-checks core subsystems (disk space, memory, its own state store, the bridges, and packet processing) and attempts an automatic repair when one fails. Inspect or force a run:
# Latest result for every health check
cenvero-str-ctl heal status
# Force an immediate run now
sudo cenvero-str-ctl heal check
Troubleshooting
The agent won't start
1. Read the journal — it almost always names the cause:
journalctl -u cenvero-stratum -n 100 --no-pager
systemctl status cenvero-stratum
2. Confirm you are running as root and the host network came up:
systemctl status cenvero-str-network
The network unit must succeed first — it creates the bridges the agent attaches
to. If it failed, fix the host networking and sudo systemctl restart
cenvero-str-network.
3. Inspect the config the agent is loading:
sudo cenvero-str-ctl config show
A config that fails to decode falls back to defaults with a note in the output. A signed config that fails verification is rejected — re-fetch a fresh panel-signed config rather than hand-editing the file.
4. Check disk space and the data directory (/var/lib/cenvero-str/). A full disk
stops the agent from writing its state; the health check reclaims rotated logs
but a genuinely full volume needs operator action.
cenvero-str-ctl says the agent is unreachable (exit code 3)
The CLI talks to the agent over its local socket. If commands report the agent is unreachable, the agent process is not running — start it and re-check:
sudo systemctl start cenvero-stratum
cenvero-str-ctl status
License is not active
Check the installed license and this machine's hardware identity:
sudo cenvero-str-ctl license status
- Expired / in grace / frozen — renew it. Existing traffic keeps running even when frozen; only new or changing operations are blocked. Renew with
sudo cenvero-str-ctl license renew(or fetch a fresh one after renewing in your account). See Licensing for the warn → grace → freeze model. - Not activated — activate this machine with
sudo cenvero-str-ctl license activate CNVR-XXXX-XXXX-XXXX-XXXX, then confirm it in your account. - Hardware-ID unavailable —
license statusshows the machine's hardware ID. The binding requires firmware/hardware identifiers that are not exposed inside generic virtual machines, so a node must be bare metal. If the ID shows as unavailable, you are likely running in an unsupported virtualized environment. - Wrong release channel — a stable license runs only stable builds and a pre-release license runs only beta/RC builds; a mismatch fails closed. Make sure the build you installed matches your license's channel.
TLS certificate is pending
After install the agent generates its own certificate; if a customer-approval or domain step is outstanding the certificate can show as pending. Confirm what the agent is using and re-issue if needed:
sudo cenvero-str-ctl config show # shows the TLS cert/key paths in use
sudo cenvero-str-ctl tls -h # certificate management commands
TLS material is managed by the certificate manager (cenvero-str-ctl tls …), not
config set. Adding or removing a domain name with tls domain requests a new
certificate, which waits for your approval in your account — see
TLS/SSL & License Operations.
Node is not registered
Registration binds the node to its license and mints its node token. If a node never registered, re-check connectivity to the panel and the license, then let it re-register:
cenvero-str-ctl status # shows license + registration state
sudo cenvero-str-ctl license activate CNVR-XXXX-XXXX-XXXX-XXXX
Registration needs outbound HTTPS to your management/license server and a valid license. If the panel is unreachable, the agent keeps retrying — fix connectivity and it completes on its own.
"REST API disabled" — set an API token
If a client gets a REST API disabled (or 401/unauthorized) response, two settings govern access:
1. Is the service enabled? Confirm with cenvero-str-ctl service status. If
rest is off, turn it back on and restart:
sudo cenvero-str-ctl service rest on
sudo systemctl restart cenvero-stratum
2. Is an API token set? The REST/gRPC/WebSocket APIs require the bearer token
to be configured. cenvero-str-ctl api-token status reports whether one is set
(never its value). The token is a credential, so it is not set with
config set. If none is set, create one on the node and restart:
sudo cenvero-str-ctl api-token generate # prints the token once
# or supply your own, read from stdin:
printf '%s' "$MY_TOKEN" | sudo cenvero-str-ctl api-token set
sudo systemctl restart cenvero-stratum
No reinstall is needed, and the token survives a later panel sync.
Where to look first
| Symptom | First check |
|---|---|
| Agent down / crash-looping | journalctl -u cenvero-stratum -n 100 |
| CLI "agent unreachable" | systemctl status cenvero-stratum |
| License problems | cenvero-str-ctl license status |
| A service not answering | cenvero-str-ctl service status |
| Settings not taking effect | Did you systemctl restart cenvero-stratum? |
| Overall health | cenvero-str-ctl heal status |
Next steps
- CLI Reference — the full command surface.
- Configuration — the node config and what it controls.
- Licensing — the enforcement state machine.
- Upgrades — how updates are pulled and verified.