LINUXOR.SK ... open source notes ...

Vault 03 - High-level design

category: solutionz · date: 2024-12-31 · updated: 2026-10-02 · author: LALA

Vault Solution · Previous: Secret model · Next: Network design and firewall flows

Three Vault clusters, each in its own environment, each with the same shape. This article explains the shape, why there are three, and what the design does when something fails.

Three environments

EnvironmentClusterNodesServesUnsealed by
PRODprod-vault5The production instances of the portal and the orchestratorCOMMON, automatically
NONPRODnonprod-vault5Their test and development instances, both on one cluster to save costCOMMON, automatically
COMMONcommon-vault3Nobody. It holds two Transit keys, one per cluster above.People, with Shamir key shares
mermaid
flowchart TB
  subgraph prodc["Production consumers"]
    op["orchestrator"]
    sp["service-portal"]
  end
  subgraph testc["Test and development consumers"]
    ot["orchestrator"]
    st["service-portal"]
  end
  subgraph cloud["Private cloud"]
    subgraph e1["PROD"]
      v1["prod-vault"]
    end
    subgraph e2["NONPROD"]
      v2["nonprod-vault"]
    end
    subgraph e3["COMMON"]
      v3["common-vault"]
    end
  end
  op -- API --> v1
  sp -- API --> v1
  ot -- API --> v2
  st -- API --> v2
  v1 -- "auto-unseal" --> v3
  v2 -- "auto-unseal" --> v3

COMMON exists because the Community edition cannot use a hardware security module and the private cloud offered no key management service. A Vault that seals itself on every restart and waits for people with key shares does not meet a 99.95 % target, so the unseal key of the two working clusters is wrapped by a key in a third Vault. That third one is small, stores nothing a consumer needs, and is the only one that still needs people after a restart.

One environment

mermaid
flowchart TB
  clients["API clients"] -- "HTTPS 443" --> vip
  admin["Administrator"] -- "SSH 22" --> bastion
  subgraph env["One environment"]
    subgraph lb["Load balancer, active and passive"]
      vip(["Virtual address"])
      lb1["lb1<br/>HAProxy, Keepalived"]
      lb2["lb2<br/>HAProxy, Keepalived"]
      vip --- lb1
      vip -.- lb2
    end
    bastion["bastion1"]
    subgraph cluster["Vault cluster, Raft"]
      n1["node1<br/>standby"]
      n2["node2<br/>standby"]
      n3["node3<br/>active"]
      n4["node4<br/>standby"]
      n5["node5<br/>standby"]
    end
    lb1 -- "TCP 8200" --> n3
    bastion -. "SSH" .-> lb1
    bastion -. "SSH" .-> n1
  end
ComponentCountWhat it is for
Vault node5, or 3 in COMMONA Vault server with its own copy of the data on a local disk. One node is active and answers every request; the others replicate and wait.
Load-balancer node2HAProxy sends clients to the active Vault node; Keepalived keeps one address on whichever of the two is healthy. See Load balancer.
Bastion host1The only server administrators can log in to from outside. Everything else accepts SSH from the bastion's network only.

The load balancer is there for availability, not for load. In the Community edition a standby does not serve reads; it forwards them to the active node. Sending a client to a standby would work and add a hop, so the health check picks the one node that answers "active".

Storage

Each node keeps Vault's data in a BoltDB file on its own disk, and the Raft protocol replicates every write to a majority of nodes before it is acknowledged. There is no shared storage and no second product.

Cluster sizeMajorityNode failures survived
321
532

The storage is disk-bound: a write is as fast as the slowest of the majority can sync it. The sizing below follows the vendor's "small cluster" figures, with provisioned rather than burstable disk performance.

RolevCPUMemorySystem diskData disk
Vault node416 GiB40 GiB100 GiB, mounted at /data
Load-balancer node24 GiB40 GiBnone
Bastion host24 GiB40 GiBnone

Two zones where three are wanted

A Raft cluster is meant to be spread over three failure domains, so that losing one leaves a majority. The private cloud had two availability zones, and its own guidance was to run a workload in one zone and let the platform restart it in the other after a failure.

mermaid
flowchart LR
  subgraph az1["Availability zone 1"]
    p["prod-vault<br/>all 5 nodes, both load balancers"]
    n["nonprod-vault<br/>all 5 nodes, both load balancers"]
  end
  subgraph az2["Availability zone 2"]
    c["common-vault<br/>all 3 nodes, both load balancers"]
  end
  p -. "restarted here if zone 1 fails" .-> az2
  n -. "restarted here if zone 1 fails" .-> az2
  c -. "restarted here if zone 2 fails" .-> az1

The clusters are therefore not spread across zones at all. PROD and NONPROD run whole in zone 1, COMMON runs whole in zone 2, and the storage of every virtual machine is replicated synchronously to the other zone by the platform. Zone failure is handled below Vault: the machines come back in the surviving zone with their disks.

Putting COMMON in the other zone is deliberate. The question is what must already be running when PROD restarts.

EventWhat happensNeeds people?
One or two Vault nodes of PROD failThe cluster keeps its majority. The failed nodes restart, unseal through COMMON and catch up.No
The active node failsThe remaining nodes elect a new one within seconds; HAProxy's health check follows.No
One load-balancer node failsKeepalived moves the address to the other.No
Zone 1 failsPROD and NONPROD restart in zone 2. COMMON is already running there, unsealed, so they unseal themselves.No
Zone 2 failsPROD and NONPROD keep running; they hold their keys in memory. COMMON restarts in zone 1, sealed.Yes: until someone unseals COMMON, a restarted PROD node would stay sealed
COMMON is down or sealed, and nothing else happensNothing. A running Vault does not need its seal.No
COMMON is down and a PROD node restartsThat node stays sealed until COMMON is back.Yes
Both zones failEverything restarts sealed, COMMON first.Yes

Had COMMON run in zone 1 with the others, the single most likely large failure, losing zone 1, would have restarted all three sealed and left production waiting for key holders. In zone 2, COMMON is the one cluster that survives that failure untouched.

What this arrangement does not protect against is a split between the zones with both alive. Raft cannot help there, because each cluster has all of its voters on one side by design; the platform decides which side runs the machines.

Integrations

SystemDirectionProtocol
Orchestration platform, customer portalClients of the Vault APIHTTPS to the load balancer, authentication by AppRole
Security operations centreReceives the audit trail and operating-system security eventsSyslog over TCP
MonitoringPolls agents on every server and probes the client endpointZabbix agent, HTTPS
Object storage of the private cloudReceives snapshots and rotated audit logsS3 API over HTTPS
Internal certificate authorityIssues node certificatesManual request, see TLS, DNS and certificates

The team

The design document also sized the team, and it is worth repeating because it is the part most often skipped. A shared secrets service is run by at least two people, better three, so that on-call does not rest on one. The useful background is system administration and integration work, not security policy: most of the effort goes into keeping the service up and helping other teams connect to it. For operations the list was concrete: Linux services, firewall, SELinux, LVM and upgrades; TLS and how to request a certificate; HAProxy and a floating address; backup and restore.

Reading it today

The reference architecture for Integrated Storage is unchanged: five voters over three zones, or a documented compromise. Vault 2.1 still offers the Community edition no redundancy zones, no non-voting standbys and no seal high availability, so a third cluster as the unsealer remains a reasonable design. Autopilot, which was already reporting cluster health here, can also remove dead servers, but that is off by default and needs a minimum quorum to be set first.

← solutionz