Balabit 05 - High availability
Balabit SCB Solution · Previous: Hardware, cabling and networks · Next: IPMI out-of-band management
In non-transparent mode the SCB is the only door to the jump servers. The servers themselves keep running when it breaks, which is why the design calls the mode safe, but nobody can administer them through it. So high availability was a requirement twice over (FR11 and NR2: "Solution must be deployed in High Availability configuration"). This part is about how the pair was meant to work, why DNS was not the answer, the heartbeat and next-hop settings as built in site 1, and what a support bundle of the site 2 pair showed nine months later, including a slave that could not reach its time servers.
Master and slave
The SCB's HA is a master/slave pair with identical configuration. One node went into datacenter A, the other into datacenter B of site 1, so that a whole datacenter can be lost. The master holds the cluster addresses and does all the work; it shares all its data with the slave over the HA interface, port 4 of the appliance. If the master stops, the slave takes over the IP addresses of the master's interfaces and sends gratuitous ARP requests, so that the hosts on the local networks learn the new MAC address behind the same IP address. Users and administrators keep using the same addresses: 10.11.16.81 for management, 10.11.16.65 and 2001:db8:a1:c0e::f:1 for user sessions, 10.11.18.225 towards the NetApp.
flowchart LR subgraph dca["Datacenter A"] a["dc1-a-ablb001"] sa2["DC1-A-SPRO002"] sa1["DC1-A-SPRO001"] end subgraph dcb["Datacenter B"] b["dc1-b-ablb001"] sb2["DC1-B-SPRO002"] sb1["DC1-B-SPRO001"] end vips["Cluster addresses 10.11.16.81, 10.11.16.65, 2001:db8:a1:c0e::f:1, 10.11.18.225"] nh["Next hop 10.11.18.222"] a -- "port 4, 1.2.4.1" --> sa2 b -- "port 4, 1.2.4.2" --> sb2 sa2 -- "VLAN 1016, heartbeat and DRBD" --- sb2 a -- "port 3, 10.11.18.209" --> sa1 b -- "port 3, 10.11.18.210" --> sb1 sa1 -- "VLAN 1018, heartbeat only" --- sb1 vips -. "held by the master" .- a vips -. "or by the other node" .- b a -. "ping" .-> nh b -. "ping" .-> nh
Under the web interface, the support bundle of site 2 shows what the pair is made of: the classic Linux-HA heartbeat for membership and DRBD for the data, one DRBD resource replicated synchronously over the HA link. That is my reading of the files in the bundle; the design describes only the behaviour.
Why not DNS
My pre-design analysis started from a different idea. Its hand-drawn sketch has two appliances, but every role of the first one, management and operation, on two ports with two addresses each, and the second appliance with ports 1 and 2 not connected:
flowchart LR mgmt["MGMT, SCB admins"] ops["OPERATORS"] subgraph scb1["SCB 1"] s1p1["port 1: MGMT_IP1 IPv4, OP_IP1 IPv4 IPv6"] s1p2["port 2: OP_IP2 IPv4 IPv6, MGMT_IP2 IPv4"] s1p3["port 3: 10.50.0.1 IPv4"] s1p4["port 4: 1.2.4.1 IPv4"] end subgraph scb2["SCB 2"] s2p12["ports 1 and 2, unconnected"] s2p3["port 3: 10.50.0.2 IPv4"] s2p4["port 4: 1.2.4.2 IPv4"] end hared["HA Redundant"] hapri["HA Primary"] arch["Archive"] mgmt --> s1p1 mgmt --> s1p2 ops --> s1p1 ops --> s1p2 s1p3 --- hared hared --- s2p3 s1p4 --- hapri hapri --- s2p4 s1p3 -- "NFS, Rsync over SSH, SMB/CIFS" --- arch
As I read it, the reason for two addresses per role was that the SCB cannot bond or team its network cards (constraint C1 of the design; "teaming/bonding" and "configure the prod interface (in HA mode/team/bond)" stood on my TODO list); the analysis itself gives no reason. It asked whether DNS could serve "as a crutch" to choose between MGMT_IP1 and MGMT_IP2, and took the idea apart: plain round-robin DNS gives no reasonable HA (as I would explain it, a client keeps trying the address it got, dead or not); SRV records would work, but most software, all browsers included, does not use them; and the commercial "DNS failover" services work for services exposed to the internet, not for an appliance hidden in a management LAN, and would make the clients use external DNS. The sentence breaks off there in my notes. The design dropped the idea: one address per role, on one port, and the cluster moves the address.
Redundant heartbeat and next-hop monitoring
Two settings protect that decision against the two classic failures of a two-node cluster.
The first is a split brain: the HA link breaks, each node thinks the other is dead, and both take the addresses. Against it, a redundant heartbeat runs on a second path, port 3 in VLAN 1018, between 10.11.18.209 and 10.11.18.210. It carries only heartbeat messages, no data. If the HA link fails but the redundant heartbeat still answers, there is no takeover, and no data is synchronised to the slave until the HA link is back. If only the redundant heartbeat fails, nothing happens either.
The second is a master that is alive but cut off from the world, which a heartbeat alone does not notice. Against it, next-hop monitoring: both nodes ping an address, and if the master cannot reach it while the slave can, the slave forces a takeover even though the master otherwise works.
flowchart TB start["The slave watches the master"] hb{"Heartbeat from the master on the HA link or the redundant link?"} take["Takeover: the slave takes the cluster addresses and sends gratuitous ARP"] nh{"10.11.18.222 unreachable from the master but reachable from the slave?"} force["Forced takeover, although the master works"] link{"HA link up?"} stay["No takeover"] nosync["No takeover; no data synchronised until the HA link is back"] start --> hb hb -- "neither" --> take hb -- "yes" --> nh nh -- "yes" --> force nh -- "no" --> link link -- "yes" --> stay link -- "no, only the redundant link" --> nosync
Chapter 7 of the design records the High availability page as configured: HA link speed auto-negotiated on both nodes, the HA interface with 1.2.4.1 and 1.2.4.2, both marked (FIX) (the fixed addresses of the cluster link, which the analysis sketch already had), the redundant heartbeat on physical interface 3 with 10.11.18.209 (this node, datacenter A) and 10.11.18.210 (other node, datacenter B), nothing on physical interfaces 1 and 2, and next-hop monitoring of 10.11.18.222 on physical interface 3 from both nodes. The page is a Config document: High availability (site 1).
Reading it again, the next hop is in the CLS2 network 10.11.18.208/28, the heartbeat VLAN, so it tests the path of port 3 and nothing else. Port 3 hangs on SPRO001 together with port 1, so losing that switch cuts the next hop too and forces a takeover. But a master whose port 2 lost its switch (SPRO002, which also carries the HA link; the redundant heartbeat would still answer) or whose port-1 cable failed would still reach 10.11.18.222, and would keep the cluster addresses on a dead link. Monitoring the gateways of the management and production networks would have covered that; the design does not say why only port 3 was chosen. I also never recorded what happens to sessions in progress at a takeover. As I understand the product, they are cut and the users connect again to the same address; the documents hold no takeover test at all.
Site 2, nine months later
The site 2 pair has no design document, but a support bundle of 2018-09-17, collected for a problem the folder calls "RDP002" (presumably the RDP jump server dc2-a-vcrdp002), holds the HA state of both nodes. The nodes call themselves scb1 and scb2, the bundle master and slave. The bundle shows 5.0.6 for the core firmware and both boot firmwares. The master holds all resources, the slave none, and DRBD is in the state you want: connected, both disks UpToDate, nothing out of sync. The whole listing is a Config document: HA state of the site 2 cluster.
Two lines of the master's heartbeat configuration, ha.cf (the whole file is in the Config document):
output 2 lines
auto_failback off ucast eth3 1.2.4.2
One heartbeat path only: unicast over eth3, the HA port, to the other node. The redundant-heartbeat settings in the same bundle are empty for every interface, and the exported config.xml has no CLS2 interface on port 3 at all, although the site 2 L3 sheet plans 10.12.18.209 and 10.12.18.210 for it. Site 2 was built without the redundant heartbeat, whether by decision or by omission the material does not say. auto_failback off means, as I understand heartbeat, that a recovered node does not take the resources back, so the cluster stays on whichever node took over last.
The DRBD resource r0 sits on /dev/sda4 and replicates between 1.2.4.1:1111 and 1.2.4.2:1111 with protocol C, that is synchronously; after-sb-1pri discard-secondary resolves a split brain, as I read the option, in favour of the node that was primary. The master's local address is 1.2.4.1, which the site 2 sheet gives to datacenter A, so on that day the master was the node in datacenter A.
The slave that was out of sync
The same bundle has the NTP peers of both nodes. The master synchronises from the two Infoblox appliances of site 1. The slave's peers look like this:
output 5 lines
remote refid st t when poll reach delay offset jitter ============================================================================== 10.11.16.145 .INIT. 16 u - 1024 0 0.000 0.000 0.000 10.11.18.145 .INIT. 16 u - 1024 0 0.000 0.000 0.000 scb1 10.11.18.145 3 s 792 64 0 0.000 0.000 0.000
.INIT. with reach 0 means the slave has never had an answer from either Infoblox, and its peer scb1, the master, has not answered recently either. My guess is that the slave, which holds none of the cluster addresses, has no address from which to reach the Infoblox and depends on the master for time; that is an inference, not something the bundle says. The SCB web interface has a warning for a slave out of time sync, a yellow banner on Basic Settings > Date & Time: "Slave is out of sync with the master". I kept a screenshot of it, and I kept the vendor's support engineer's answer on how to synchronise by hand: from the boot shell over SSH, compare date -R on both nodes (ssh scb-other reaches the other node), look at the peers with ntpq -c peers, stop ntp, run ntpd -q -g once, start ntp again and compare the offsets. The commands are a Config document: Manual NTP synchronisation. Neither the screenshot nor the answer is dated in a way that ties it to this bundle, so I cannot say they are the same event.
Time matters more on an SCB than on most appliances: audit trails are timestamped and signed, and, as I see it, a takeover to a slave with a wrong clock would produce trails with wrong times. My working notes also wanted the SCB and the domain controllers on the same NTP source, for the Active Directory integration (see Active Directory and access control).
Loose ends
- No takeover test, planned or done, is in the material, for either site.
- Next-hop monitoring covers only port 3 in site 1; site 2 shows none either: the "Physical interface" lines of its HA state, which I take to be the next-hop and heartbeat addresses, are empty.
- The redundant heartbeat runs over a switch port whose native VLAN is 1018 while the SCB tags VLAN 1018; see Hardware, cabling and networks.
- "upgrade (second node)" stood on my TODO list: in an HA pair both nodes have to be upgraded; the upgrade history is in Operations, upgrades and troubleshooting.
- The slave's NTP state in site 2 was never explained in the material.
What I would do differently
I would monitor a next hop on the production and the management side, not only on the heartbeat VLAN, and I would test a takeover before go-live and write down what users see. Today's vendor documentation says it plainly: the shift between the nodes "might take 5–10 minutes", and connections "are lost and they are not automatically restored". I would also check the redundant heartbeat on the switch side, and build it in site 2 as planned.