LINUXOR.SK ... open source notes ...

Balabit - HA state of the site 2 cluster

category: solutionz · date: 2018-12-31 · updated: 2026-10-03 · author: LALA

Balabit SCB Solution · Config document · referenced from High availability

noteThis is a snapshot of one day, 2018-09-17, from a support bundle collected for a problem the folder calls "RDP002". It shows a healthy DRBD pair, no redundant heartbeat, and a slave whose NTP peers do not answer. It is not a configuration to copy.

The high-availability, DRBD, heartbeat and NTP state of both nodes of the site 2 cluster dc2-s-xblb001, as the support bundle collected it: the cluster status, the heartbeat configuration file ha.cf, the DRBD resource and its kernel status, the redundant-heartbeat settings, and the NTP peers of each node.

ItemValue
Clusterdc2-s-xblb001, two SCB T-10 appliances in site 2; the bundle calls the nodes scb1 and scb2 and its folders master and slave
Firmware5.0.6 (core firmware of the bundle's node, boot firmware and "other boot firmware"); the slave's core firmware is not shown
Collected2018-09-17, support bundle of the site 2 cluster (folder "DC2 - RDP002 problem")
SourceSelected files of the support bundle: info/info.txt, info/ha.txt and the info/master/ and info/slave/ files below
Not in itWhich node is in datacenter A and which in B (see below), the heartbeat logs, any takeover history

The listing

output 194 lines
======================================================================
 info/info.txt (host and firmware lines)
======================================================================
Host: dc2-s-xblb001.dc2-s-xblb001.adm.example.net
HA state: ha

======================================================================
 info/ha.txt
======================================================================
cs=Connected
st_this=Primary
ds=UpToDate
st_other=Secondary
cs=Connected
st_this=Secondary
ds=UpToDate
st_other=Primary
internal={'gw': '', 'other_ip': '', 'iface': 'ha2', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
external={'gw': '', 'other_ip': '', 'iface': 'ha0', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
mgmt={'gw': '', 'other_ip': '', 'iface': 'ha1', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
internal={'gw': '', 'other_ip': '', 'iface': 'ha2', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
external={'gw': '', 'other_ip': '', 'iface': 'ha0', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
mgmt={'gw': '', 'other_ip': '', 'iface': 'ha1', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
Physical interface 1 (self):

Physical interface 1 (other):

Physical interface 2 (self):

Physical interface 2 (other):

Physical interface 3 (self):

Physical interface 3 (other):

======================================================================
 info/master/hardware_type.txt
======================================================================
T10

======================================================================
 info/master/cl_status.out.txt
======================================================================
This node is holding all resources.

======================================================================
 info/master/ha.cf.txt
======================================================================
node scb1 scb2
crm no
use_logd on
auto_failback off
hbgenmethod time
respawn hacluster /usr/lib/heartbeat/dopd
apiauth dopd gid=haclient uid=hacluster
ucast eth3 1.2.4.2

======================================================================
 info/master/ha_drbd.txt
======================================================================
{'cs': 'Connected', 'st_this': 'Primary', 'ds': 'UpToDate', 'st_other': 'Secondary'}

======================================================================
 info/master/ha_redundant.txt
======================================================================
{'internal': {'gw': '', 'other_ip': '', 'iface': 'ha2', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}, 'external': {'gw': '', 'other_ip': '', 'iface': 'ha0', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}, 'mgmt': {'gw': '', 'other_ip': '', 'iface': 'ha1', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}}

======================================================================
 info/master/drbdsetup.show.r0.txt
======================================================================
resource r0 {
    options {
    }
    net {
        max-epoch-size          8000;
        max-buffers             8000;
        after-sb-1pri           discard-secondary;
        csums-alg               "sha1";
    }
    _remote_host {
        address                 ipv4 1.2.4.2:1111;
    }
    _this_host {
        address                 ipv4 1.2.4.1:1111;
        volume 0 {
            device                      minor 0;
            disk                        "/dev/sda4";
            meta-disk                   internal;
            disk {
                on-io-error             call-local-io-error;
                resync-rate             1280k; # bytes/second
                al-extents              1801;
                c-max-rate              131072k; # bytes/second
            }
        }
    }
}

======================================================================
 info/master/proc-drbd.out.txt
======================================================================
version: 8.4.5 (api:1/proto:86-101)
built-in
 0: cs:Connected ro:Primary/Secondary ds:UpToDate/UpToDate C r-----
    ns:174129120 nr:0 dw:130929236 dr:1103436753 al:7162 bm:0 lo:0 pe:0 ua:0 ap:0 ep:1 wo:f oos:0

======================================================================
 info/master/ntpq.peers.txt
======================================================================
     remote           refid      st t when poll reach   delay   offset  jitter
==============================================================================
+dc1-a-adnsvip00 10.11.16.161  2 u  617 1024  377    7.668    0.080   0.107
*dc1-b-adnsvip00 10.11.16.162  2 u  245 1024  377    7.725    0.017   0.110
 scb2            127.0.0.1       15 s   27   64    0    0.000    0.000   0.000

======================================================================
 info/slave/hardware_type.txt
======================================================================
T10

======================================================================
 info/slave/cl_status.out.txt
======================================================================
This node is holding none resources.

======================================================================
 info/slave/ha.cf.txt
======================================================================
node scb1 scb2
crm no
use_logd on
auto_failback off
hbgenmethod time
respawn hacluster /usr/lib/heartbeat/dopd
apiauth dopd gid=haclient uid=hacluster
ucast eth3 1.2.4.1

======================================================================
 info/slave/ha_drbd.txt
======================================================================
{'cs': 'Connected', 'st_this': 'Secondary', 'ds': 'UpToDate', 'st_other': 'Primary'}

======================================================================
 info/slave/ha_redundant.txt
======================================================================
{'internal': {'gw': '', 'other_ip': '', 'iface': 'ha2', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}, 'external': {'gw': '', 'other_ip': '', 'iface': 'ha0', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}, 'mgmt': {'gw': '', 'other_ip': '', 'iface': 'ha1', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}}

======================================================================
 info/slave/drbdsetup.show.r0.txt
======================================================================
resource r0 {
    options {
    }
    net {
        max-epoch-size          8000;
        max-buffers             8000;
        after-sb-1pri           discard-secondary;
        csums-alg               "sha1";
    }
    _remote_host {
        address                 ipv4 1.2.4.1:1111;
    }
    _this_host {
        address                 ipv4 1.2.4.2:1111;
        volume 0 {
            device                      minor 0;
            disk                        "/dev/sda4";
            meta-disk                   internal;
            disk {
                on-io-error             call-local-io-error;
                resync-rate             1280k; # bytes/second
                al-extents              1801;
                c-max-rate              131072k; # bytes/second
            }
        }
    }
}

======================================================================
 info/slave/proc-drbd.out.txt
======================================================================
version: 8.4.5 (api:1/proto:86-101)
built-in
 0: cs:Connected ro:Secondary/Primary ds:UpToDate/UpToDate C r-----
    ns:0 nr:174145012 dw:174145012 dr:1102852704 al:0 bm:0 lo:0 pe:0 ua:0 ap:0 ep:1 wo:f oos:0

======================================================================
 info/slave/ntpq.peers.txt
======================================================================
     remote           refid      st t when poll reach   delay   offset  jitter
==============================================================================
 10.11.16.145 .INIT.          16 u    - 1024    0    0.000    0.000   0.000
 10.11.18.145 .INIT.          16 u    - 1024    0    0.000    0.000   0.000
 scb1            10.11.18.145  3 s  792   64    0    0.000    0.000   0.000
LinesWhat they say
Host: dc2-s-xblb001.dc2-s-xblb001.adm.example.netThe cluster's host name appears twice in the FQDN; as I read it, the bundle concatenated host name and a domain that already contained it. The <networking> element of the site 2 config.xml has host name dc2-s-xblb001 and domain adm.example.net
HA state: haThe cluster runs as a pair
info/ha.txtThe DRBD state seen from both nodes, then the redundant-heartbeat settings twice: every field (ip, other_ip, gw, enabled) is empty for the interfaces ha0, ha1 and ha2. The "Physical interface 1 to 3" lines, which would show next-hop or heartbeat addresses, are empty too
cl_status: "holding all resources" / "holding none resources"The master holds the cluster addresses and services; the slave holds nothing and waits
ha.cf: node scb1 scb2, ucast eth3 1.2.4.2The heartbeat configuration (as I understand it, of the classic Linux-HA heartbeat package without a cluster resource manager, crm no). The only heartbeat path is unicast over eth3, the HA port, to the other node's 1.2.4.x address; there is no second ucast line for a redundant heartbeat
auto_failback offAfter a takeover the recovered node does not take the resources back; the cluster stays on the node that took over (general knowledge of heartbeat, not from my notes)
drbdsetup show r0One DRBD resource r0 on /dev/sda4 with internal metadata, replicated between 1.2.4.1:1111 and 1.2.4.2:1111, that is over the HA port. after-sb-1pri discard-secondary tells DRBD, after a split brain where one node was primary, to throw away the changes of the secondary; csums-alg "sha1" makes a resync send only blocks whose checksums differ (my reading of the DRBD 8.4 options)
/proc/drbdDRBD 8.4.5, protocol C (synchronous: a write completes when both nodes have it), cs:Connected, ds:UpToDate/UpToDate, oos:0: the slave was a complete copy. The ns of the master and the nr of the slave are 174 million KiB, about 166 GiB, sent and received since the counters were last reset
Master ntpq peersBoth Infoblox appliances of site 1 (the names are cut to 15 characters by ntpq), reachable (reach 377), stratum 2, one selected (*); scb2 is the other node, at stratum 15 and not reachable
Slave ntpq peers10.11.16.145 and 10.11.18.145 stay in .INIT. with reach 0: the slave never got an answer from them. Its peer scb1 (the master, at stratum 3) also has reach 0, last heard 792 seconds earlier. So at that moment the slave had no reachable time source

The node whose DRBD local address is 1.2.4.1 is the master. The site 2 L3 sheet gives 1.2.4.1 to datacenter A, so the master was the node in datacenter A on that day; that is a conclusion from the addresses, the bundle does not name the datacenters.

Checked against One Identity Safeguard for Privileged Sessions 9.0

As builtToday
SCB 5.0.6, DRBD between master and slaveDRBD status and a DRBD sync rate limit are still on the HA page of SPS 9.0; master and slave are now called primary and secondary. 5.0.x LTS support was discontinued on 2020-05-28
A slave whose time was not synchronisedThe alert xcbTimeSyncLost is still in the 9.0 alert list, as is xcbHaNodeChanged for a takeover

The research did not cover the heartbeat software or the files behind it, so I cannot say whether a 9.0 support bundle would contain ha.cf and drbdsetup output in the same form.

← solutionz