Lifecycle and Restarts
The node keeps a record of every virtual machine — its size, its interfaces, and whether you want it running or stopped — and that record is what counts. The node starts the machine itself once the machine's networking is in place, and puts things back when something changes them behind its back.
Part of Compute, in early access.
Starting and stopping
sudo cenvero-str-ctl vm start vm-3f9a1c2e
sudo cenvero-str-ctl vm stop vm-3f9a1c2e # ask the guest to shut down
sudo cenvero-str-ctl vm stop vm-3f9a1c2e --timeout 300 # wait longer before forcing it off
sudo cenvero-str-ctl vm force-stop vm-3f9a1c2e # power off at once
sudo cenvero-str-ctl vm restart vm-3f9a1c2e # stop, then start again
sudo cenvero-str-ctl vm restart vm-3f9a1c2e --force # power off, then start again
vm stop presses the machine's power button: the guest shuts down cleanly. If it
has not finished after the timeout (120 seconds unless you say otherwise), the
machine is powered off. A guest that is still booting may not react to the power
button in time. Before every start, the machine's port, its protection and its
first-boot disk are put in place again.
| Method | Path | Body | Answer |
|---|---|---|---|
POST | /api/v1/vms/{id}/start | — | 200 running, or 202 starting while it boots |
POST | /api/v1/vms/{id}/stop | {"timeout_seconds": 300} (optional, up to 3600) | 202 |
POST | /api/v1/vms/{id}/force-stop | — | 202 |
POST | /api/v1/vms/{id}/restart | {"force": true} (optional) | 202 |
Each answer names the task that records the operation ("task", and a
Location header). A stop's task stays running until the guest is off — "the
guest shut down", or "forced off" after the timeout — and a restart's until the
machine runs again, so the task tells you when the operation really finished:
sudo cenvero-str-ctl task log tsk-3f2a0c1e9b7d-0mumq7fmz-a05dee --follow
See Tasks for the whole record, and what happens to a task the agent was carrying out when it restarted.
States
state | Meaning |
|---|---|
creating | Being created |
starting / running | Starting, running |
stopping / stopped | Shutting down, off |
paused | Paused. With state_reason "storage error", the disk pool ran out of space — free space, then restart the machine |
crashed | The guest crashed and is not being restarted (below) |
deleting | Being deleted; running vm delete again carries on where it stopped |
unknown | The node has lost touch with the machines for a moment (for example while the virtualization service restarts); the reason says what was last known |
state_reason says why in plain words. flags marks conditions that need your
attention:
| Flag | Meaning | What to do |
|---|---|---|
pending_restart | A change waits for the next restart: one you made with vm update that could not be applied live, an interface that needs a free slot, or a definition changed outside Stratum and put back while the running machine still has the changed one | Restart it with vm restart |
drifted | The running machine's network or disk wiring was changed outside Stratum. It is never killed for it; a security event is raised | Restart it to restore its protection |
network_degraded | The machine's network interface was removed outside Stratum. It has been recreated and protected again, but the running machine could not be reconnected to it (usually it is reconnected without a restart) | Restart it |
crash_loop | It crashed too often and is left off | Find the cause, then vm start |
What happens when…
…the guest shuts itself down (poweroff inside it). The machine stays off:
Stratum does not restart a machine its owner switched off. The same goes for a
machine stopped by some other tool on the node — and a machine started that way
is adopted as running.
…the guest crashes. It is started again at once, then after 30 seconds, then
after two minutes. A fourth crash within ten minutes leaves it crashed, flagged
crash_loop, until you start it — also across agent restarts and host reboots.
…the agent restarts or updates. Machines keep running: they do not depend on the agent process. While it is down, nothing can be changed and consoles close; when it is back, it picks the running machines up again without restarting them.
…the server reboots. When the server comes back, the node first rebuilds each machine's network port, its lock and its separation from other tenants, then starts the machines that were meant to be running. Machines you had stopped stay stopped. The installer sets the server up so that its own start-up does not start them early and a server shutdown shuts them down cleanly (on a server shared with another panel, those settings stay as that panel has them).
…the virtualization service restarts or is upgraded. Running machines keep running. The node reconnects, checks every machine against its records and accepts changes again; until then changes are refused with 503 and the reason (for example "Compute is connecting to the virtualization service; try again in a moment").
…the disk pool fills up. Guests that try to write pause, with the reason
"storage error". compute status warns at 85 % and 95 % full. Free space and
restart the paused machines.
…someone changes a machine outside Stratum. A changed definition is put back
(pending_restart while it runs); a deleted definition is recreated; a removed
network interface is recreated, locked again and reconnected to the running
machine (network_degraded only when that cannot be done). Changing
the running machine's wiring raises drifted and a security event — the machine
is never killed for it. Machines that Stratum did not create are never touched.
…the licence expires or is frozen. Machines keep running and come back after a reboot; starting, stopping, creating and deleting are refused until it is renewed. Consoles still open (see Consoles).
…the tenant is suspended. Its traffic is cut, and its machines are stopped:
each running one is asked to shut down and is powered off if it has not within
the stop timeout (2 minutes). While the suspension lasts they cannot be started
(vm start says why), and one started outside Stratum is stopped again. The node
remembers which machines were running — across agent restarts and reboots — and
when the tenant is resumed exactly those start again; the ones that were
stopped stay stopped. Stopping a machine yourself during the suspension means it
stays off after it. vm show marks a held machine with "tenant_suspended":
true and "resume_state" (what the resume restores). See
Tenants & Bandwidth.
Events
Machines report to the node's event stream
and webhooks — creation, corrections (compute.drift_corrected), console
sessions and more — next to everything else the node reports. Every change of a
machine's state, desired state or flags is a compute.vm_state event, so a
screen that shows machines can follow them without asking again; every task's
start, progress and end are task.* events.