Clustering Overview
Clustering needs agent 1.0.0-rc.81 or later on every member, and a licence that includes clustering on every node that forms, joins or manages a cluster. Reading a node's status, withdrawing a join code and taking a node out work without it (see Licences in a cluster). An older node can be updated first (see Upgrades).
A cluster is a group of nodes you manage as one. You form it on one node, give each further node a one-time join code, and from then on you can sign in to any member's web console, or call any member's API, and see and act on every member's machines, networks, volumes, images and tenants, with one list of tasks, one audit trail and one event stream. The command line manages the cluster itself from any member. There is no separate management server: every member can do it, and it does not matter which one you talk to.
What a cluster shares, and what stays with a node
The members agree on a small set of shared settings through a leader they elect among themselves. A change to one of them, made on any member, reaches every member:
- Tenants: each tenant's name, its status (active or suspended) and its bandwidth cap. Suspending a tenant on any member suspends it on every member, and stops its machines wherever they run. Deleting a tenant deletes it on every member and revokes its tenant API keys on every member; it is refused while any member still holds the tenant's machines, networks or volumes, or cannot be asked. A tenant's other quotas — machines, addresses, rules and volumes — are set on each node, for that node.
- The IP blocklist: the addresses that intrusion detection and per-source connection limits block on their own. An address one member blocks is blocked on every member, and the block ends at the same moment on every member, even after a member restarts or a node joins. On the other members it lasts at most 7 days. A block never applies to a cluster member's address (see Members' addresses are never blocked).
- Overlay peers of your VXLAN networks.
- Address allocations, so two members do not hand out the same address from a shared range.
- Floating IP assignments. An address still stays on the node you assign it on: moving it to another node when its node fails comes in a later release.
- The cluster itself: its members, their identities and the join codes.
Everything else belongs to the node it was created on: machines, volumes, images, networks and their endpoints, firewall rules, load balancers, DNS zones, DHCP scopes, routes, BGP sessions, and the node's own configuration, API token, API keys and tenant keys. You can manage all of it from any member, but each object stays on its node, and the web console shows each one with its node.
Before you start
- Agent 1.0.0-rc.81 or later on every node.
- A licence with clustering on every node. Licences stay per node and bound to that node's hardware; a cluster does not pool them. Every Suite plan includes clustering, the free Suite Lab too; Compute and Storage plans do not. The pricing page shows which Fabric plans include it, and
cenvero-str-ctl license statusshows what a node's licence includes. - A fixed address for each node on your management network, and TCP 7073 open between the members, in both directions, a new node included. The members talk to each other on it only over an encrypted connection on which both ends prove who they are, and a new node uses it to join with its code; nothing else is accepted on that port. Never open it to the internet.
- Clocks in step. Keep NTP running on every node (licensing needs it too). A member whose clock is more than 2 seconds off is flagged; more than 5 seconds off, requests to it are refused until the clocks agree.
- Nodes that are not in a cluster. A node that belonged to another cluster is cleared first with
cenvero-str-ctl cluster forget --yes. - Up to 32 nodes in one cluster.
Form a cluster
On the first node, in its web console: Datacenter › Cluster, the **Form a cluster** card. Give the cluster a name, and the node's own address that the other members will reach it on, then press Form the cluster. The same from the command line:
sudo cenvero-str-ctl cluster create --name prod --bind 10.0.0.5 --wait
--nameis how the cluster is shown. It can be up to 64 bytes long: 64 characters for a name of unaccented letters, digits and punctuation, fewer for a name with accented letters or other scripts, which take two to four bytes per character. It holds printable characters only.--bindis an address of this node, written as10.0.0.5or10.0.0.5:7073. It cannot be a wildcard such as0.0.0.0, and the port is always 7073.--waitfollows the work until it ends; without it the command answers at once with the cluster's id and name and a task you can follow.
The node becomes the cluster's first member and its leader, and starts
accepting the other members on port 7073. Over the API it is
POST /api/v1/cluster/create (see the API Reference).
Add a node
1. Make a join code on any member: **Datacenter › Cluster › Make a join code** in its web console (the dialog offers 15 minutes, 1, 4 or 24 hours), or:
sudo cenvero-str-ctl cluster join-code create --expires 30m
The code — a single line starting STRJ1- — is shown once, when it is made.
It works once, and expires after 15 minutes unless you choose otherwise
(from 60 seconds to 24 hours). Treat it like a password until it has been used:
it lets one node join your cluster. A cluster holds at most 20 active codes
(made, and not yet used, revoked or expired); revoke one to make room.
cluster join-code list shows every code with its state (active, used,
revoked or expired), who made it and which node used it, but never the code
itself. cluster join-code revoke <id> withdraws a code that has not been used.
2. Join with it on the new node. In its web console: **Datacenter › Cluster › Join a cluster**. Paste the code, check the address it offers for this node, and press Next: the node asks the cluster what joining would mean, without using up the code. The summary then names the cluster and how many nodes it has, the member the node joins through, the address the others will reach it at, the tenants it will share and any records that stop the join. Nothing changes until you press Join the cluster. From the command line:
sudo cenvero-str-ctl cluster join --code - --bind 10.0.0.6 --wait
--code - reads the code from standard input, so it never lands in your shell
history: paste it and press Enter. --bind is the new node's own address, as
for cluster create.
Before it sends anything secret, the new node checks that it is talking to the cluster the code was made for, so the code's secret never reaches anything else. It then joins without a vote and catches up on the shared settings; once it has caught up it can be given a vote, if the cluster's size calls for one more (see Voting members).
What happens to the new node's tenants. Its tenants join the cluster's:
- A tenant with the same id and name as one of the cluster's is the same tenant. The cluster's status and bandwidth cap win.
- A tenant with the same id as one of the cluster's but another name stops the join, and the answer lists each conflict. Rename one of them, or join with
--adopt-cluster-tenantsto let the cluster's record win. - A tenant only the new node has becomes a tenant of the whole cluster.
- The two blocklists are combined. Each block keeps the end it had, and a block of a member's address, the new node's included, is lifted. An address allocated to different owners on the node and in the cluster stops the join, and is listed.
3. Check on any member:
sudo cenvero-str-ctl cluster status
{
"data": {
"enabled": true,
"state": "leader",
"leader": "10.0.0.5:7073",
"is_leader": true,
"cluster": { "id": "0f1e2d3c4b5a69788796a5b4c3d2e1f0", "name": "prod" },
"voters": 1,
"peers": [
{ "id": "3f2a0c1e-9b7d-4c1a-8e2f-0123456789ab", "address": "10.0.0.5:7073", "state": "leader",
"name": "edge-1", "voter": true, "pinned": true, "online": true },
{ "id": "7c9e6679-7425-40de-944b-e07fc1f90ae7", "address": "10.0.0.6:7073", "state": "non-voter",
"name": "edge-2", "voter": false, "pinned": true, "online": true, "last_seen": "2026-09-30T11:59:58Z" }
]
},
"status": "ok"
}
Each member is listed with its id, which is what the commands and the API name a
node by, its role (leader, follower, or non-voter), whether it votes,
whether it is online and when it was last heard from.
Which node to talk to
Any of them. Sign in to any member's web console and its Datacenter lists every member with its objects. An action on another member's machine, network or volume is carried out by that member, under that member's own rules, and its task, audit record and events appear wherever you look.
- A member that stops answering is shown as not answering, with when it was last heard from, as soon as one attempt to read its objects fails — within a few seconds while you are looking. Its objects stay listed from the last time it answered, marked as such, and the console holds back changes to them until it is back. A request sent to it through the API in the first moments waits until the member you called gives up on it, about 15 to 20 seconds after it stopped answering, and is then refused with 503, naming when it was last seen; from then on, requests to it are refused at once. Everything else keeps working.
- A member on a version the others cannot manage is listed with a link to its own web console: update it, or manage it there.
- The API lists the members at
GET /api/v1/nodes, and reaches another member's objects at/api/v1/nodes/<node id>/…— the same paths as on the node itself. The cluster-wide lists are/api/v1/cluster/resources,…/tasks,…/auditand…/events. See the API Reference. - The command line works on the node it runs on. The
clustercommands manage the whole cluster from any member; for another member's machines and networks, use the web console or the API. - Credentials. Keys and tokens stay per node, but one that a member accepts lets you act on every member through it: the member you call checks it, and asks the others on your behalf. So treat each member's API token and operator keys as access to the whole cluster, and keep them as safe as the most sensitive node's. Tenant keys, and the customer portal they sign in to, stay with the node that issued them.
- Address allowlists. Each member applies its own list of allowed addresses (
api_allowed_ips) to your address, also when you reach it through another member. A member that does not accept your address refuses with 403forbidden: source address not allowed, the same answer it gives you directly, and the member you called passes it back unchanged. The task and audit lists of the whole cluster list that member as not answering. Its machines, networks and other objects still appear in the console's tree and in/api/v1/cluster/resources, and its events in the cluster's event stream: that is a known limitation. The command line on a member is not affected:cluster leaveandcluster removerun there reach the leader whatever the leader's list says.
Voting members and what a cluster survives
Changes to the shared settings are agreed by the voting members. How many vote depends on the size of the cluster, and the leader keeps it that way as nodes join and leave:
| Members | Voting members | Can lose, and still change shared settings |
|---|---|---|
| 1 | 1 | — |
| 2 | 1 | the member that does not vote |
| 3 or 4 | 3 | any one voting member |
| 5 or 6 | 5 | any two voting members |
| 7 to 32 | 7 | any three voting members |
The other members do not vote but receive every change. A node that joins is given a vote only once it has caught up with the others, and only when the cluster's size calls for another voting member; when a voting member leaves, another member takes its place.
A two-node cluster has no spare. Only one of its nodes votes, so while that node is down, shared settings cannot change. Add a third node to survive the loss of any one.
A node cut off from the others
A node that cannot reach the cluster's leader — it is on the smaller side of a network split, it is cut off from the others, or too many voting members are down — refuses changes to tenants and overlay peers: creating, renaming, suspending, resuming or limiting a tenant, setting its quota where that changes its bandwidth cap, deleting it, and adding or removing an overlay peer. The answer is 503 (unavailable), the node changes nothing, and it says why:
the cluster has no leader reachable from this node; changes to shared settings are paused until it has one
Why: these settings are the cluster's, not the node's. A change has to reach the cluster as a whole, so that every member applies the same one; a node that went ahead on its own would disagree with the others once they met again.
What to do: make the change on a member that can reach the others — any member on the larger side of a split — or restore the network between the members, and the node takes changes again. A billing system suspending a tenant should do the same: retry on another member.
What keeps working: reading — tenants, their quotas and usage, every list — and everything that belongs to the node alone: its machines, networks and the rest can still be changed on it. **Traffic and running workloads are never affected.**
A member that was away
A member that was switched off while nodes joined or were removed catches up when it starts again. It asks the members it already knows which members the cluster has now, accepts the nodes that joined in the meantime, and refuses at once any node that was removed. There is nothing to do.
When it cannot take part. If none of the members it knows answers, or its
clustering could not start when the agent started, the node keeps running on
its own: its machines, networks and traffic carry on, and what belongs to it
alone can still be changed. Changes to shared settings are paused on it, as on
a node cut off from the others. It says so in
cluster status (the error field), in its own entry of GET /api/v1/nodes
(the reason field) and in the agent log, which has the cause. Restore the
network between the members, or fix what the log names and restart the agent. A
node that should no longer be a member is made standalone with
cluster forget --yes.
Remove a node, or leave
Remove a member from any member. In the console, open **Datacenter › Cluster, press Remove…** on the member's row, type the node's name to confirm, and press Remove node. Or:
sudo cenvero-str-ctl cluster remove 7c9e6679-7425-40de-944b-e07fc1f90ae7
The node is cut off at once: from that moment the others refuse its connections. Where you can, remove a node before you shut it down for good. If the voting members cannot agree to the removal — a majority is down — the node stays cut off and the answer says so:
removal pending: the cluster could not agree; the node can no longer reach the others
Leave from the node itself, whichever member it is. In its console, open Datacenter › Cluster › Advanced, press Leave the cluster…, type the cluster's name to confirm, and press Leave the cluster. Or:
sudo cenvero-str-ctl cluster leave
A node that was removed, or left, becomes a standalone node again: it keeps its own machines, networks and other objects, and its tenants as they were, and it no longer receives shared changes. Port 7073 closes. The leader of a two-node cluster can leave too — the other node takes the vote over first — and when the last member leaves, the cluster is gone.
On a member that is not the leader, the node asks the leader to take it out, and answers once it is standalone. If the leader's answer is lost and the node has not seen itself removed, the leave answers 503 and says what to run:
the cluster's leader did not answer, and this node has not seen itself removed: if `cluster status` on another member still lists this node, run the leave again; if it does not, run `cenvero-str-ctl cluster forget --yes` here
Removing a member and leaving work whatever the licences say: while a licence is frozen, and on a node whose plan does not include clustering. So you can always take a node out.
A node that cannot leave normally. A node that was removed while it was switched off, or that was away so long that its identity expired, can make itself standalone without the others' agreement:
sudo cenvero-str-ctl cluster forget --yes
It changes nothing on the other members: if the node is still listed there,
remove it with cluster remove <id>. To bring it back, join it again with a new
join code.
A member's identity
Each member has its own identity, issued by the cluster when it joins, so the
members know exactly who they are talking to. An identity is valid for one year,
and each member renews its own automatically, about four months before it
expires. Nothing to do; GET /api/v1/nodes shows each member's
cert_expires_at, and the console warns when less than 30 days are left.
A member whose identity expired while it was switched off cannot reconnect: run
cluster forget --yes on it, cluster remove <id> on the cluster, and join it
again with a new code.
Rekey a member to give it a new identity — for example if its files may have been copied:
sudo cenvero-str-ctl cluster rekey 7c9e6679-7425-40de-944b-e07fc1f90ae7 --wait
Without a node id it rekeys the node you run it on. The member's open connections close, and reopen with its new identity.
The cluster's certificate authority is what issues the identities. It is valid for ten years, and you can replace it whenever you choose:
sudo cenvero-str-ctl cluster ca rotate --wait
Each member moves to the new authority with its identity unchanged; both authorities are trusted until every member has moved, and then the old one is retired. The work is a task that shows how many members have moved. A member that is switched off moves when it is back: if the task ends before that, it says how many have moved, and the old authority stays trusted until the rest have, then is retired. One rotation runs at a time. Join codes made from then on name the new authority; unused codes made before it stop working, so make new ones. In the console, rekeying (Give it a new identity) and rotating (Replace the authority…) are under Datacenter › Cluster › Advanced.
Protect every member as you protect the cluster. Every member holds what it takes to admit new members, so anyone with root on one member can administer the whole cluster. If a member may be compromised: remove it, revoke any join codes that have not been used, rotate the certificate authority, and replace the API tokens and keys it held. See Security.
Members' addresses are never blocked
A block of a member's address would cut that member off from the others, the cluster's own traffic included. So the address a member uses in the cluster, and its API address, are never blocked:
- Intrusion detection and per-source connection limits never block a member's address. A detection is still reported.
- A node that joins has any block of its addresses lifted, on every member.
- A rule you add with
cenvero-str-ctl rules addorrules batchthat drops or rejects a member's address is refused. - A firewall deny rule whose source is a member's address is refused in the same words when it could drop the cluster's own traffic. That holds for
cenvero-str-ctl firewall deny,POST /api/v1/rulesandPUT /api/v1/rules/{id}, which answer409, and each rule ofPOST /api/v1/rules/batch. A deny of a member that leaves the cluster's traffic alone is accepted: one limited to UDP or ICMP, to a destination address that is not a member's, or to destination ports other than 7073 and the ports the node's own outgoing connections use (32768 to 60999 on most systems).
The answer's error reads:
rules: 10.0.0.6 is the address of cluster member 7c9e6679-7425-40de-944b-e07fc1f90ae7. A block of it would drop that member's traffic to this node, the cluster's own included, and cut this node off from it. To stop trusting that machine, take it out of the cluster first (cenvero-str-ctl cluster remove 7c9e6679-7425-40de-944b-e07fc1f90ae7)
A rule made before the address became a member's is kept, but it is not
enforced against that member while it is one; the agent log says so. To stop
trusting a member, take it out with cluster remove <node-id>; its address can
be blocked from then on.
A firewall rule that drops traffic to an address is not checked against the members, nor is the firewall's default action: both are policies for what reaches the node. If one covers this node's own address, keep TCP 7073 open with an allow rule from the other members that is evaluated before it (see Before you start).
Licences in a cluster
- Every member has its own licence. A Fabric node can join a Suite cluster: the console then shows it without Compute or Storage sections, unless it holds machines or volumes.
- Forming, joining, and managing other members from a node need clustering in that node's licence: forming or joining a cluster, making join codes, rekeying a member, replacing the certificate authority, acting on another member's objects through the node, and the summary and cluster-wide views (
GET /api/v1/cluster,/cluster/state,/cluster/resources,/cluster/tasks,/cluster/auditand/cluster/events). - The member that does the work applies its own licence. A change asked of a member whose licence is frozen is refused by that member, wherever you asked from; a member without Compute shows its machines but cannot create one.
- A licence never takes a node out of its cluster. A member whose licence later lacks clustering stays a member. It can still read its status, list and revoke join codes, remove a member and leave. It cannot manage the other members, make join codes or show the cluster-wide lists until its licence includes clustering again.
Five things work whatever the licence says: without clustering in it, and while it is frozen. So you can always see where a node stands, withdraw a code and take a node out.
| Command | API |
|---|---|
cluster status | GET /api/v1/cluster/status |
cluster join-code list | GET /api/v1/cluster/join-codes |
cluster join-code revoke <id> | DELETE /api/v1/cluster/join-codes/{id} |
cluster remove <node-id> | DELETE /api/v1/cluster/members/{id} |
cluster leave | POST /api/v1/cluster/leave |
cluster forget --yes always works too, as root on the node itself. A tenant
key is refused on every /api/v1/cluster/… route.
Ports
| Port | Protocol | Purpose |
|---|---|---|
| 7073 | TCP | Between the members of a cluster, encrypted, each end proving who it is, and for a new node joining with its code. Closed on a node that is not in a cluster |
| 7074 | UDP | The heartbeat between the two nodes of a high-availability pair (see Gateway High Availability) |
Open them only between the nodes that use them, on your management network — the management bridge is kept apart from the workload bridge for exactly this reason (see Networking Overview).
When something is refused
| Answer | When |
|---|---|
the join code was not accepted | The code is wrong, or not one of this cluster's |
this join code has expired, this join code has already been used, this join code was revoked | Make a new one |
this cluster has 32 nodes, the most it supports | The cluster is full |
this node is already a member of a cluster; … | Leave first, or cluster forget a node the cluster no longer knows |
this node still holds the data of an earlier cluster; … | Run cluster forget --yes on it first |
the members this join code lists could not admit this node (…), or … could not be asked (…) when the console checks the code | The new node cannot reach them: open TCP 7073 from it to the members |
the cluster did not add this node: … | The rest of the answer says why; most often the members cannot reach the new node, so open TCP 7073 to its --bind address. The code is used up: make a new one. If the members list the node as joining, remove it there with cluster remove <id> before you join it again |
node <name> is not answering (last seen <time>); try again when it is back | That member is down or cut off |
node <name> has not finished joining the cluster | It was admitted but has not completed its join; one that never completes it is dropped after an hour |
node <name> runs a version that cannot be managed from here; … | Update it, or use its own console |
node <name>'s clock differs from this node's by <n> s; … | Fix the clocks (NTP) |
the cluster has no leader reachable from this node; … | See A node cut off from the others |
tenant <id> still has machines, networks or volumes on node <name> | Delete them on that member first |
node <name> is not answering, so its resources cannot be checked | A tenant is deleted only once every member can say it holds nothing of it |
the cluster's leader did not answer, and this node has not seen itself removed: … (503) | A leave from a member that is not the leader lost the leader's answer. If cluster status on another member still lists the node, run the leave again; if not, run cluster forget --yes on it |
forbidden: source address not allowed (403) | The member that holds the object does not accept your address (its api_allowed_ips); see Address allowlists |
… is not included in your plan …, naming the cluster feature | This node's licence does not include clustering; see Licences in a cluster for what still works |
rules: <address> is the address of cluster member <node id>. … | See Members' addresses are never blocked |
Current limits
- 32 nodes per cluster.
- The customer portal works on the node whose tenant key the customer holds; tenant keys are not shared between members.
- The command line manages the cluster from any member, but another member's machines and networks only through the console or the API.
- Machines stay on their node. There is no placement across nodes and no migration; a network belongs to its node, and an overlay connects networks between nodes (see Networking Overview).
- Gateway failover and floating IPs that move to another node come in a later release (see Gateway High Availability).
- Updates are applied node by node, as on any node (see Upgrades).
- A member that does not accept your address refuses every request and console you send it through another member, but its objects still appear in the console's tree and in
/api/v1/cluster/resources, and its events in the cluster's event stream (see Which node to talk to).
See also
- The Web Console — managing every member from one console.
- Management API Reference — the cluster routes and their answers.
- CLI Reference — the
clustercommands. - Gateway High Availability — a redundant pair sharing a floating address.
- Moving Workloads Between Nodes — keeping an endpoint's IP/MAC when it moves between nodes.