Vault 03 - High-level design
Vault Solution · Previous: Secret model · Next: Network design and firewall flows
Three Vault clusters, each in its own environment, each with the same shape. This article explains the shape, why there are three, and what the design does when something fails.
Three environments
| Environment | Cluster | Nodes | Serves | Unsealed by |
|---|---|---|---|---|
| PROD | prod-vault | 5 | The production instances of the portal and the orchestrator | COMMON, automatically |
| NONPROD | nonprod-vault | 5 | Their test and development instances, both on one cluster to save cost | COMMON, automatically |
| COMMON | common-vault | 3 | Nobody. It holds two Transit keys, one per cluster above. | People, with Shamir key shares |
flowchart TB subgraph prodc["Production consumers"] op["orchestrator"] sp["service-portal"] end subgraph testc["Test and development consumers"] ot["orchestrator"] st["service-portal"] end subgraph cloud["Private cloud"] subgraph e1["PROD"] v1["prod-vault"] end subgraph e2["NONPROD"] v2["nonprod-vault"] end subgraph e3["COMMON"] v3["common-vault"] end end op -- API --> v1 sp -- API --> v1 ot -- API --> v2 st -- API --> v2 v1 -- "auto-unseal" --> v3 v2 -- "auto-unseal" --> v3
COMMON exists because the Community edition cannot use a hardware security module and the private cloud offered no key management service. A Vault that seals itself on every restart and waits for people with key shares does not meet a 99.95 % target, so the unseal key of the two working clusters is wrapped by a key in a third Vault. That third one is small, stores nothing a consumer needs, and is the only one that still needs people after a restart.
One environment
flowchart TB clients["API clients"] -- "HTTPS 443" --> vip admin["Administrator"] -- "SSH 22" --> bastion subgraph env["One environment"] subgraph lb["Load balancer, active and passive"] vip(["Virtual address"]) lb1["lb1<br/>HAProxy, Keepalived"] lb2["lb2<br/>HAProxy, Keepalived"] vip --- lb1 vip -.- lb2 end bastion["bastion1"] subgraph cluster["Vault cluster, Raft"] n1["node1<br/>standby"] n2["node2<br/>standby"] n3["node3<br/>active"] n4["node4<br/>standby"] n5["node5<br/>standby"] end lb1 -- "TCP 8200" --> n3 bastion -. "SSH" .-> lb1 bastion -. "SSH" .-> n1 end
| Component | Count | What it is for |
|---|---|---|
| Vault node | 5, or 3 in COMMON | A Vault server with its own copy of the data on a local disk. One node is active and answers every request; the others replicate and wait. |
| Load-balancer node | 2 | HAProxy sends clients to the active Vault node; Keepalived keeps one address on whichever of the two is healthy. See Load balancer. |
| Bastion host | 1 | The only server administrators can log in to from outside. Everything else accepts SSH from the bastion's network only. |
The load balancer is there for availability, not for load. In the Community edition a standby does not serve reads; it forwards them to the active node. Sending a client to a standby would work and add a hop, so the health check picks the one node that answers "active".
Storage
Each node keeps Vault's data in a BoltDB file on its own disk, and the Raft protocol replicates every write to a majority of nodes before it is acknowledged. There is no shared storage and no second product.
| Cluster size | Majority | Node failures survived |
|---|---|---|
| 3 | 2 | 1 |
| 5 | 3 | 2 |
The storage is disk-bound: a write is as fast as the slowest of the majority can sync it. The sizing below follows the vendor's "small cluster" figures, with provisioned rather than burstable disk performance.
| Role | vCPU | Memory | System disk | Data disk |
|---|---|---|---|---|
| Vault node | 4 | 16 GiB | 40 GiB | 100 GiB, mounted at /data |
| Load-balancer node | 2 | 4 GiB | 40 GiB | none |
| Bastion host | 2 | 4 GiB | 40 GiB | none |
Two zones where three are wanted
A Raft cluster is meant to be spread over three failure domains, so that losing one leaves a majority. The private cloud had two availability zones, and its own guidance was to run a workload in one zone and let the platform restart it in the other after a failure.
flowchart LR subgraph az1["Availability zone 1"] p["prod-vault<br/>all 5 nodes, both load balancers"] n["nonprod-vault<br/>all 5 nodes, both load balancers"] end subgraph az2["Availability zone 2"] c["common-vault<br/>all 3 nodes, both load balancers"] end p -. "restarted here if zone 1 fails" .-> az2 n -. "restarted here if zone 1 fails" .-> az2 c -. "restarted here if zone 2 fails" .-> az1
The clusters are therefore not spread across zones at all. PROD and NONPROD run whole in zone 1, COMMON runs whole in zone 2, and the storage of every virtual machine is replicated synchronously to the other zone by the platform. Zone failure is handled below Vault: the machines come back in the surviving zone with their disks.
Putting COMMON in the other zone is deliberate. The question is what must already be running when PROD restarts.
| Event | What happens | Needs people? |
|---|---|---|
| One or two Vault nodes of PROD fail | The cluster keeps its majority. The failed nodes restart, unseal through COMMON and catch up. | No |
| The active node fails | The remaining nodes elect a new one within seconds; HAProxy's health check follows. | No |
| One load-balancer node fails | Keepalived moves the address to the other. | No |
| Zone 1 fails | PROD and NONPROD restart in zone 2. COMMON is already running there, unsealed, so they unseal themselves. | No |
| Zone 2 fails | PROD and NONPROD keep running; they hold their keys in memory. COMMON restarts in zone 1, sealed. | Yes: until someone unseals COMMON, a restarted PROD node would stay sealed |
| COMMON is down or sealed, and nothing else happens | Nothing. A running Vault does not need its seal. | No |
| COMMON is down and a PROD node restarts | That node stays sealed until COMMON is back. | Yes |
| Both zones fail | Everything restarts sealed, COMMON first. | Yes |
Had COMMON run in zone 1 with the others, the single most likely large failure, losing zone 1, would have restarted all three sealed and left production waiting for key holders. In zone 2, COMMON is the one cluster that survives that failure untouched.
What this arrangement does not protect against is a split between the zones with both alive. Raft cannot help there, because each cluster has all of its voters on one side by design; the platform decides which side runs the machines.
Integrations
| System | Direction | Protocol |
|---|---|---|
| Orchestration platform, customer portal | Clients of the Vault API | HTTPS to the load balancer, authentication by AppRole |
| Security operations centre | Receives the audit trail and operating-system security events | Syslog over TCP |
| Monitoring | Polls agents on every server and probes the client endpoint | Zabbix agent, HTTPS |
| Object storage of the private cloud | Receives snapshots and rotated audit logs | S3 API over HTTPS |
| Internal certificate authority | Issues node certificates | Manual request, see TLS, DNS and certificates |
The team
The design document also sized the team, and it is worth repeating because it is the part most often skipped. A shared secrets service is run by at least two people, better three, so that on-call does not rest on one. The useful background is system administration and integration work, not security policy: most of the effort goes into keeping the service up and helping other teams connect to it. For operations the list was concrete: Linux services, firewall, SELinux, LVM and upgrades; TLS and how to request a certificate; HAProxy and a floating address; backup and restore.
Reading it today
The reference architecture for Integrated Storage is unchanged: five voters over three zones, or a documented compromise. Vault 2.1 still offers the Community edition no redundancy zones, no non-voting standbys and no seal high availability, so a third cluster as the unsealer remains a reasonable design. Autopilot, which was already reporting cluster health here, can also remove dead servers, but that is off by default and needs a minimum quorum to be set first.