LINUXOR.SK ... open source notes ...

Balabit 05 - High availability

category: solutionz · date: 2018-12-31 · updated: 2026-10-03 · author: LALA

Balabit SCB Solution · Previous: Hardware, cabling and networks · Next: IPMI out-of-band management

In non-transparent mode the SCB is the only door to the jump servers. The servers themselves keep running when it breaks, which is why the design calls the mode safe, but nobody can administer them through it. So high availability was a requirement twice over (FR11 and NR2: "Solution must be deployed in High Availability configuration"). This part is about how the pair was meant to work, why DNS was not the answer, the heartbeat and next-hop settings as built in site 1, and what a support bundle of the site 2 pair showed nine months later, including a slave that could not reach its time servers.

Master and slave

The SCB's HA is a master/slave pair with identical configuration. One node went into datacenter A, the other into datacenter B of site 1, so that a whole datacenter can be lost. The master holds the cluster addresses and does all the work; it shares all its data with the slave over the HA interface, port 4 of the appliance. If the master stops, the slave takes over the IP addresses of the master's interfaces and sends gratuitous ARP requests, so that the hosts on the local networks learn the new MAC address behind the same IP address. Users and administrators keep using the same addresses: 10.11.16.81 for management, 10.11.16.65 and 2001:db8:a1:c0e::f:1 for user sessions, 10.11.18.225 towards the NetApp.

mermaid
flowchart LR
  subgraph dca["Datacenter A"]
    a["dc1-a-ablb001"]
    sa2["DC1-A-SPRO002"]
    sa1["DC1-A-SPRO001"]
  end
  subgraph dcb["Datacenter B"]
    b["dc1-b-ablb001"]
    sb2["DC1-B-SPRO002"]
    sb1["DC1-B-SPRO001"]
  end
  vips["Cluster addresses 10.11.16.81, 10.11.16.65, 2001:db8:a1:c0e::f:1, 10.11.18.225"]
  nh["Next hop 10.11.18.222"]
  a -- "port 4, 1.2.4.1" --> sa2
  b -- "port 4, 1.2.4.2" --> sb2
  sa2 -- "VLAN 1016, heartbeat and DRBD" --- sb2
  a -- "port 3, 10.11.18.209" --> sa1
  b -- "port 3, 10.11.18.210" --> sb1
  sa1 -- "VLAN 1018, heartbeat only" --- sb1
  vips -. "held by the master" .- a
  vips -. "or by the other node" .- b
  a -. "ping" .-> nh
  b -. "ping" .-> nh

Under the web interface, the support bundle of site 2 shows what the pair is made of: the classic Linux-HA heartbeat for membership and DRBD for the data, one DRBD resource replicated synchronously over the HA link. That is my reading of the files in the bundle; the design describes only the behaviour.

Why not DNS

My pre-design analysis started from a different idea. Its hand-drawn sketch has two appliances, but every role of the first one, management and operation, on two ports with two addresses each, and the second appliance with ports 1 and 2 not connected:

mermaid
flowchart LR
  mgmt["MGMT, SCB admins"]
  ops["OPERATORS"]
  subgraph scb1["SCB 1"]
    s1p1["port 1: MGMT_IP1 IPv4, OP_IP1 IPv4 IPv6"]
    s1p2["port 2: OP_IP2 IPv4 IPv6, MGMT_IP2 IPv4"]
    s1p3["port 3: 10.50.0.1 IPv4"]
    s1p4["port 4: 1.2.4.1 IPv4"]
  end
  subgraph scb2["SCB 2"]
    s2p12["ports 1 and 2, unconnected"]
    s2p3["port 3: 10.50.0.2 IPv4"]
    s2p4["port 4: 1.2.4.2 IPv4"]
  end
  hared["HA Redundant"]
  hapri["HA Primary"]
  arch["Archive"]
  mgmt --> s1p1
  mgmt --> s1p2
  ops --> s1p1
  ops --> s1p2
  s1p3 --- hared
  hared --- s2p3
  s1p4 --- hapri
  hapri --- s2p4
  s1p3 -- "NFS, Rsync over SSH, SMB/CIFS" --- arch

As I read it, the reason for two addresses per role was that the SCB cannot bond or team its network cards (constraint C1 of the design; "teaming/bonding" and "configure the prod interface (in HA mode/team/bond)" stood on my TODO list); the analysis itself gives no reason. It asked whether DNS could serve "as a crutch" to choose between MGMT_IP1 and MGMT_IP2, and took the idea apart: plain round-robin DNS gives no reasonable HA (as I would explain it, a client keeps trying the address it got, dead or not); SRV records would work, but most software, all browsers included, does not use them; and the commercial "DNS failover" services work for services exposed to the internet, not for an appliance hidden in a management LAN, and would make the clients use external DNS. The sentence breaks off there in my notes. The design dropped the idea: one address per role, on one port, and the cluster moves the address.

Redundant heartbeat and next-hop monitoring

Two settings protect that decision against the two classic failures of a two-node cluster.

The first is a split brain: the HA link breaks, each node thinks the other is dead, and both take the addresses. Against it, a redundant heartbeat runs on a second path, port 3 in VLAN 1018, between 10.11.18.209 and 10.11.18.210. It carries only heartbeat messages, no data. If the HA link fails but the redundant heartbeat still answers, there is no takeover, and no data is synchronised to the slave until the HA link is back. If only the redundant heartbeat fails, nothing happens either.

The second is a master that is alive but cut off from the world, which a heartbeat alone does not notice. Against it, next-hop monitoring: both nodes ping an address, and if the master cannot reach it while the slave can, the slave forces a takeover even though the master otherwise works.

mermaid
flowchart TB
  start["The slave watches the master"]
  hb{"Heartbeat from the master on the HA link or the redundant link?"}
  take["Takeover: the slave takes the cluster addresses and sends gratuitous ARP"]
  nh{"10.11.18.222 unreachable from the master but reachable from the slave?"}
  force["Forced takeover, although the master works"]
  link{"HA link up?"}
  stay["No takeover"]
  nosync["No takeover; no data synchronised until the HA link is back"]
  start --> hb
  hb -- "neither" --> take
  hb -- "yes" --> nh
  nh -- "yes" --> force
  nh -- "no" --> link
  link -- "yes" --> stay
  link -- "no, only the redundant link" --> nosync

Chapter 7 of the design records the High availability page as configured: HA link speed auto-negotiated on both nodes, the HA interface with 1.2.4.1 and 1.2.4.2, both marked (FIX) (the fixed addresses of the cluster link, which the analysis sketch already had), the redundant heartbeat on physical interface 3 with 10.11.18.209 (this node, datacenter A) and 10.11.18.210 (other node, datacenter B), nothing on physical interfaces 1 and 2, and next-hop monitoring of 10.11.18.222 on physical interface 3 from both nodes. The page is a Config document: High availability (site 1).

Reading it again, the next hop is in the CLS2 network 10.11.18.208/28, the heartbeat VLAN, so it tests the path of port 3 and nothing else. Port 3 hangs on SPRO001 together with port 1, so losing that switch cuts the next hop too and forces a takeover. But a master whose port 2 lost its switch (SPRO002, which also carries the HA link; the redundant heartbeat would still answer) or whose port-1 cable failed would still reach 10.11.18.222, and would keep the cluster addresses on a dead link. Monitoring the gateways of the management and production networks would have covered that; the design does not say why only port 3 was chosen. I also never recorded what happens to sessions in progress at a takeover. As I understand the product, they are cut and the users connect again to the same address; the documents hold no takeover test at all.

Site 2, nine months later

The site 2 pair has no design document, but a support bundle of 2018-09-17, collected for a problem the folder calls "RDP002" (presumably the RDP jump server dc2-a-vcrdp002), holds the HA state of both nodes. The nodes call themselves scb1 and scb2, the bundle master and slave. The bundle shows 5.0.6 for the core firmware and both boot firmwares. The master holds all resources, the slave none, and DRBD is in the state you want: connected, both disks UpToDate, nothing out of sync. The whole listing is a Config document: HA state of the site 2 cluster.

Two lines of the master's heartbeat configuration, ha.cf (the whole file is in the Config document):

output 2 lines
auto_failback off
ucast eth3 1.2.4.2

One heartbeat path only: unicast over eth3, the HA port, to the other node. The redundant-heartbeat settings in the same bundle are empty for every interface, and the exported config.xml has no CLS2 interface on port 3 at all, although the site 2 L3 sheet plans 10.12.18.209 and 10.12.18.210 for it. Site 2 was built without the redundant heartbeat, whether by decision or by omission the material does not say. auto_failback off means, as I understand heartbeat, that a recovered node does not take the resources back, so the cluster stays on whichever node took over last.

The DRBD resource r0 sits on /dev/sda4 and replicates between 1.2.4.1:1111 and 1.2.4.2:1111 with protocol C, that is synchronously; after-sb-1pri discard-secondary resolves a split brain, as I read the option, in favour of the node that was primary. The master's local address is 1.2.4.1, which the site 2 sheet gives to datacenter A, so on that day the master was the node in datacenter A.

The slave that was out of sync

The same bundle has the NTP peers of both nodes. The master synchronises from the two Infoblox appliances of site 1. The slave's peers look like this:

output 5 lines
     remote           refid      st t when poll reach   delay   offset  jitter
==============================================================================
 10.11.16.145 .INIT.          16 u    - 1024    0    0.000    0.000   0.000
 10.11.18.145 .INIT.          16 u    - 1024    0    0.000    0.000   0.000
 scb1            10.11.18.145  3 s  792   64    0    0.000    0.000   0.000

.INIT. with reach 0 means the slave has never had an answer from either Infoblox, and its peer scb1, the master, has not answered recently either. My guess is that the slave, which holds none of the cluster addresses, has no address from which to reach the Infoblox and depends on the master for time; that is an inference, not something the bundle says. The SCB web interface has a warning for a slave out of time sync, a yellow banner on Basic Settings > Date & Time: "Slave is out of sync with the master". I kept a screenshot of it, and I kept the vendor's support engineer's answer on how to synchronise by hand: from the boot shell over SSH, compare date -R on both nodes (ssh scb-other reaches the other node), look at the peers with ntpq -c peers, stop ntp, run ntpd -q -g once, start ntp again and compare the offsets. The commands are a Config document: Manual NTP synchronisation. Neither the screenshot nor the answer is dated in a way that ties it to this bundle, so I cannot say they are the same event.

Time matters more on an SCB than on most appliances: audit trails are timestamped and signed, and, as I see it, a takeover to a slave with a wrong clock would produce trails with wrong times. My working notes also wanted the SCB and the domain controllers on the same NTP source, for the Active Directory integration (see Active Directory and access control).

Loose ends

What I would do differently

I would monitor a next hop on the production and the management side, not only on the heartbeat VLAN, and I would test a takeover before go-live and write down what users see. Today's vendor documentation says it plainly: the shift between the nodes "might take 5–10 minutes", and connections "are lost and they are not automatically restored". I would also check the redundant heartbeat on the switch side, and build it in site 2 as planned.

← solutionz