Exclusive Access · Invitation Only

Gateway High Availability

You cannot set up an HA pair yourself yet. A pair's settings — the partner's address, the shared address, each node's priority and the pair's shared key — reach a node only inside the configuration delivered to it, and your account area does not offer them yet. No command on the node sets them: ha set-peer and ha configure exist only to say so. Setting up a pair comes in a later release. This page describes how a pair behaves and what a single node reports today.

Two nodes can be paired in a priority-based active/standby arrangement around a shared virtual IP (VIP). One node is the active VIP owner; the other stands by, monitoring the active node over a dedicated, encrypted health channel. When the active node fails, the standby promotes itself, assumes the VIP, and announces the new owner — without manual intervention. A pair does not need a cluster: the pairing is its own setting.

What a single node reports

A node without a partner is active in a pair of one:

cenvero-str-ctl ha status
{
  "data": {
    "local_state": "active",
    "peer_state": "solo",
    "peer_addr": "",
    "vip": "",
    "last_heartbeat": "0001-01-01T00:00:00Z",
    "active_conns": 0,
    "uptime": "2h44m47s",
    "heartbeat": "not needed"
  },
  "status": "ok"
}

local_state is active (owns the VIP) or standby (monitoring, ready to promote). peer_state is active, standby, or solo — no peer configured. heartbeat is running, not running (a peer is configured but the node cannot hear it — see below), or not needed (no peer).

How a pair works

The two nodes exchange a health signal at a short interval (100 ms) on UDP port 7074. Each node listens on that port on all of its addresses and sends to its partner's address. Each message is authenticated with the pair's shared key, and a forged, stale or replayed signal is rejected.

VIP ownership is decided by priority: while both peers are alive, the higher-priority node owns the VIP and the lower-priority node stays in standby. A standby that stops hearing its partner promotes itself regardless of priority — so the VIP is never orphaned — and yields the VIP back to a higher-priority peer once that peer returns. Equal priorities are broken deterministically by node identity so the pair can never both stay active under normal operation.

On promotion, the new owner brings the VIP up on its workload bridge (cnv-user-br0) and sends a gratuitous ARP, so neighbours update their ARP caches and start sending to it within milliseconds rather than waiting for an entry to age out. On demotion it removes only the VIP — every other address on that interface is left untouched, so a node losing the VIP never loses its own addressing with it.

A node with a partner configured but no working heartbeat — the shared key is missing or is the predictable default, the partner's address does not resolve, or the heartbeat port is already in use — cannot tell whether its partner holds the VIP, so it never takes the VIP on its own: it stays standby, ha status reports heartbeat not running, and the agent log says why. cenvero-str-ctl gateway failover moves the VIP deliberately to the node you run it on.

At start-up a node takes the VIP only once it has heard its partner and wins on priority, or once the partner has stayed silent past the failover threshold — not in the moment before the first heartbeat arrives. If the VIP cannot be installed at the moment the node becomes active — for example because the interface is not up yet during boot — the node keeps retrying about every 2 seconds for as long as it stays active.

Heartbeat timing

The heartbeat interval and the failover threshold are fixed:

ParameterValueDescription
Heartbeat interval100msTime between heartbeat probes
Failover threshold3Consecutive missed probes before declaring the peer dead

Failover detection therefore takes at most 100ms × 3 = 300 ms.

VIP takeover

When the standby detects the active node has failed, it:

  1. Promotes itself to active and assumes the VIP on its own interface.
  2. Sends a gratuitous ARP for the VIP so upstream neighbours redirect traffic to it.
  3. Marks the partner as unavailable in its local state.

The VIP on the surviving node begins accepting new connections immediately. Existing connections that were being handled by the failed node are dropped (the client must reconnect); this is inherent to a stateless L4 failover.

When the failed node returns, arbitration over the heartbeat converges the pair back to a single owner: if the returning node has higher priority it preempts and reclaims the VIP; otherwise it stays in standby.

Split-brain on a full partition

A 2-node pair has no third arbiter, witness, or quorum. If the two nodes stop hearing each other's heartbeats simultaneously — for example a management-network partition where each side is otherwise up — both nodes will promote and assume the VIP, because each believes its partner is dead. When the partition heals, the higher-priority node keeps the VIP and the other releases it. To remove the window entirely, front the pair with a third arbiter at the network layer; the platform does not provide one for two nodes.

Monitoring a pair

The fields worth alerting on are local_state — active on both nodes of a pair indicates the partition case above — heartbeat, which must be running on both nodes, and last_heartbeat, which should stay within a few hundred milliseconds of now while the peer is healthy.

See also

↓ This page as JSON ↓ All documentation as JSON