Exclusive Access · Invitation Only

Lifecycle and Restarts

The node keeps a record of every virtual machine — its size, its interfaces, and whether you want it running or stopped — and that record is what counts. The node starts the machine itself once the machine's networking is in place, and puts things back when something changes them behind its back.

Part of Compute, in early access.

Starting and stopping

sudo cenvero-str-ctl vm start vm-3f9a1c2e
sudo cenvero-str-ctl vm stop vm-3f9a1c2e               # ask the guest to shut down
sudo cenvero-str-ctl vm stop vm-3f9a1c2e --timeout 300 # wait longer before forcing it off
sudo cenvero-str-ctl vm force-stop vm-3f9a1c2e         # power off at once
sudo cenvero-str-ctl vm restart vm-3f9a1c2e            # stop, then start again
sudo cenvero-str-ctl vm restart vm-3f9a1c2e --force    # power off, then start again

vm stop presses the machine's power button: the guest shuts down cleanly. If it has not finished after the timeout (120 seconds unless you say otherwise), the machine is powered off. A guest that is still booting may not react to the power button in time. Before every start, the machine's port, its protection and its first-boot disk are put in place again.

MethodPathBodyAnswer
POST/api/v1/vms/{id}/start—200 running, or 202 starting while it boots
POST/api/v1/vms/{id}/stop{"timeout_seconds": 300} (optional, up to 3600)202
POST/api/v1/vms/{id}/force-stop—202
POST/api/v1/vms/{id}/restart{"force": true} (optional)202

Each answer names the task that records the operation ("task", and a Location header). A stop's task stays running until the guest is off — "the guest shut down", or "forced off" after the timeout — and a restart's until the machine runs again, so the task tells you when the operation really finished:

sudo cenvero-str-ctl task log tsk-3f2a0c1e9b7d-0mumq7fmz-a05dee --follow

See Tasks for the whole record, and what happens to a task the agent was carrying out when it restarted.

States

stateMeaning
creatingBeing created
starting / runningStarting, running
stopping / stoppedShutting down, off
pausedPaused. With state_reason "storage error", the disk pool ran out of space — free space, then restart the machine
crashedThe guest crashed and is not being restarted (below)
deletingBeing deleted; running vm delete again carries on where it stopped
unknownThe node has lost touch with the machines for a moment (for example while the virtualization service restarts); the reason says what was last known

state_reason says why in plain words. flags marks conditions that need your attention:

FlagMeaningWhat to do
pending_restartA change waits for the next restart: one you made with vm update that could not be applied live, an interface that needs a free slot, or a definition changed outside Stratum and put back while the running machine still has the changed oneRestart it with vm restart
driftedThe running machine's network or disk wiring was changed outside Stratum. It is never killed for it; a security event is raisedRestart it to restore its protection
network_degradedThe machine's network interface was removed outside Stratum. It has been recreated and protected again, but the running machine could not be reconnected to it (usually it is reconnected without a restart)Restart it
crash_loopIt crashed too often and is left offFind the cause, then vm start

What happens when…

…the guest shuts itself down (poweroff inside it). The machine stays off: Stratum does not restart a machine its owner switched off. The same goes for a machine stopped by some other tool on the node — and a machine started that way is adopted as running.

…the guest crashes. It is started again at once, then after 30 seconds, then after two minutes. A fourth crash within ten minutes leaves it crashed, flagged crash_loop, until you start it — also across agent restarts and host reboots.

…the agent restarts or updates. Machines keep running: they do not depend on the agent process. While it is down, nothing can be changed and consoles close; when it is back, it picks the running machines up again without restarting them.

…the server reboots. When the server comes back, the node first rebuilds each machine's network port, its lock and its separation from other tenants, then starts the machines that were meant to be running. Machines you had stopped stay stopped. The installer sets the server up so that its own start-up does not start them early and a server shutdown shuts them down cleanly (on a server shared with another panel, those settings stay as that panel has them).

…the virtualization service restarts or is upgraded. Running machines keep running. The node reconnects, checks every machine against its records and accepts changes again; until then changes are refused with 503 and the reason (for example "Compute is connecting to the virtualization service; try again in a moment").

…the disk pool fills up. Guests that try to write pause, with the reason "storage error". compute status warns at 85 % and 95 % full. Free space and restart the paused machines.

…someone changes a machine outside Stratum. A changed definition is put back (pending_restart while it runs); a deleted definition is recreated; a removed network interface is recreated, locked again and reconnected to the running machine (network_degraded only when that cannot be done). Changing the running machine's wiring raises drifted and a security event — the machine is never killed for it. Machines that Stratum did not create are never touched.

…the licence expires or is frozen. Machines keep running and come back after a reboot; starting, stopping, creating and deleting are refused until it is renewed. Consoles still open (see Consoles).

…the tenant is suspended. Its traffic is cut, and its machines are stopped: each running one is asked to shut down and is powered off if it has not within the stop timeout (2 minutes). While the suspension lasts they cannot be started (vm start says why), and one started outside Stratum is stopped again. The node remembers which machines were running — across agent restarts and reboots — and when the tenant is resumed exactly those start again; the ones that were stopped stay stopped. Stopping a machine yourself during the suspension means it stays off after it. vm show marks a held machine with "tenant_suspended": true and "resume_state" (what the resume restores). See Tenants & Bandwidth.

Events

Machines report to the node's event stream and webhooks — creation, corrections (compute.drift_corrected), console sessions and more — next to everything else the node reports. Every change of a machine's state, desired state or flags is a compute.vm_state event, so a screen that shows machines can follow them without asking again; every task's start, progress and end are task.* events.

See also

↓ This page as JSON ↓ All documentation as JSON