Balabit - HA state of the site 2 cluster
Balabit SCB Solution · Config document · referenced from High availability
The high-availability, DRBD, heartbeat and NTP state of both nodes of the site 2 cluster dc2-s-xblb001, as the support bundle collected it: the cluster status, the heartbeat configuration file ha.cf, the DRBD resource and its kernel status, the redundant-heartbeat settings, and the NTP peers of each node.
| Item | Value |
|---|---|
| Cluster | dc2-s-xblb001, two SCB T-10 appliances in site 2; the bundle calls the nodes scb1 and scb2 and its folders master and slave |
| Firmware | 5.0.6 (core firmware of the bundle's node, boot firmware and "other boot firmware"); the slave's core firmware is not shown |
| Collected | 2018-09-17, support bundle of the site 2 cluster (folder "DC2 - RDP002 problem") |
| Source | Selected files of the support bundle: info/info.txt, info/ha.txt and the info/master/ and info/slave/ files below |
| Not in it | Which node is in datacenter A and which in B (see below), the heartbeat logs, any takeover history |
The listing
output 194 lines
======================================================================
info/info.txt (host and firmware lines)
======================================================================
Host: dc2-s-xblb001.dc2-s-xblb001.adm.example.net
HA state: ha
======================================================================
info/ha.txt
======================================================================
cs=Connected
st_this=Primary
ds=UpToDate
st_other=Secondary
cs=Connected
st_this=Secondary
ds=UpToDate
st_other=Primary
internal={'gw': '', 'other_ip': '', 'iface': 'ha2', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
external={'gw': '', 'other_ip': '', 'iface': 'ha0', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
mgmt={'gw': '', 'other_ip': '', 'iface': 'ha1', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
internal={'gw': '', 'other_ip': '', 'iface': 'ha2', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
external={'gw': '', 'other_ip': '', 'iface': 'ha0', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
mgmt={'gw': '', 'other_ip': '', 'iface': 'ha1', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}
Physical interface 1 (self):
Physical interface 1 (other):
Physical interface 2 (self):
Physical interface 2 (other):
Physical interface 3 (self):
Physical interface 3 (other):
======================================================================
info/master/hardware_type.txt
======================================================================
T10
======================================================================
info/master/cl_status.out.txt
======================================================================
This node is holding all resources.
======================================================================
info/master/ha.cf.txt
======================================================================
node scb1 scb2
crm no
use_logd on
auto_failback off
hbgenmethod time
respawn hacluster /usr/lib/heartbeat/dopd
apiauth dopd gid=haclient uid=hacluster
ucast eth3 1.2.4.2
======================================================================
info/master/ha_drbd.txt
======================================================================
{'cs': 'Connected', 'st_this': 'Primary', 'ds': 'UpToDate', 'st_other': 'Secondary'}
======================================================================
info/master/ha_redundant.txt
======================================================================
{'internal': {'gw': '', 'other_ip': '', 'iface': 'ha2', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}, 'external': {'gw': '', 'other_ip': '', 'iface': 'ha0', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}, 'mgmt': {'gw': '', 'other_ip': '', 'iface': 'ha1', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}}
======================================================================
info/master/drbdsetup.show.r0.txt
======================================================================
resource r0 {
options {
}
net {
max-epoch-size 8000;
max-buffers 8000;
after-sb-1pri discard-secondary;
csums-alg "sha1";
}
_remote_host {
address ipv4 1.2.4.2:1111;
}
_this_host {
address ipv4 1.2.4.1:1111;
volume 0 {
device minor 0;
disk "/dev/sda4";
meta-disk internal;
disk {
on-io-error call-local-io-error;
resync-rate 1280k; # bytes/second
al-extents 1801;
c-max-rate 131072k; # bytes/second
}
}
}
}
======================================================================
info/master/proc-drbd.out.txt
======================================================================
version: 8.4.5 (api:1/proto:86-101)
built-in
0: cs:Connected ro:Primary/Secondary ds:UpToDate/UpToDate C r-----
ns:174129120 nr:0 dw:130929236 dr:1103436753 al:7162 bm:0 lo:0 pe:0 ua:0 ap:0 ep:1 wo:f oos:0
======================================================================
info/master/ntpq.peers.txt
======================================================================
remote refid st t when poll reach delay offset jitter
==============================================================================
+dc1-a-adnsvip00 10.11.16.161 2 u 617 1024 377 7.668 0.080 0.107
*dc1-b-adnsvip00 10.11.16.162 2 u 245 1024 377 7.725 0.017 0.110
scb2 127.0.0.1 15 s 27 64 0 0.000 0.000 0.000
======================================================================
info/slave/hardware_type.txt
======================================================================
T10
======================================================================
info/slave/cl_status.out.txt
======================================================================
This node is holding none resources.
======================================================================
info/slave/ha.cf.txt
======================================================================
node scb1 scb2
crm no
use_logd on
auto_failback off
hbgenmethod time
respawn hacluster /usr/lib/heartbeat/dopd
apiauth dopd gid=haclient uid=hacluster
ucast eth3 1.2.4.1
======================================================================
info/slave/ha_drbd.txt
======================================================================
{'cs': 'Connected', 'st_this': 'Secondary', 'ds': 'UpToDate', 'st_other': 'Primary'}
======================================================================
info/slave/ha_redundant.txt
======================================================================
{'internal': {'gw': '', 'other_ip': '', 'iface': 'ha2', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}, 'external': {'gw': '', 'other_ip': '', 'iface': 'ha0', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}, 'mgmt': {'gw': '', 'other_ip': '', 'iface': 'ha1', 'ip': '', 'prod_mac': '', 'enabled': '', 'ha_mac': ''}}
======================================================================
info/slave/drbdsetup.show.r0.txt
======================================================================
resource r0 {
options {
}
net {
max-epoch-size 8000;
max-buffers 8000;
after-sb-1pri discard-secondary;
csums-alg "sha1";
}
_remote_host {
address ipv4 1.2.4.1:1111;
}
_this_host {
address ipv4 1.2.4.2:1111;
volume 0 {
device minor 0;
disk "/dev/sda4";
meta-disk internal;
disk {
on-io-error call-local-io-error;
resync-rate 1280k; # bytes/second
al-extents 1801;
c-max-rate 131072k; # bytes/second
}
}
}
}
======================================================================
info/slave/proc-drbd.out.txt
======================================================================
version: 8.4.5 (api:1/proto:86-101)
built-in
0: cs:Connected ro:Secondary/Primary ds:UpToDate/UpToDate C r-----
ns:0 nr:174145012 dw:174145012 dr:1102852704 al:0 bm:0 lo:0 pe:0 ua:0 ap:0 ep:1 wo:f oos:0
======================================================================
info/slave/ntpq.peers.txt
======================================================================
remote refid st t when poll reach delay offset jitter
==============================================================================
10.11.16.145 .INIT. 16 u - 1024 0 0.000 0.000 0.000
10.11.18.145 .INIT. 16 u - 1024 0 0.000 0.000 0.000
scb1 10.11.18.145 3 s 792 64 0 0.000 0.000 0.000| Lines | What they say |
|---|---|
Host: dc2-s-xblb001.dc2-s-xblb001.adm.example.net | The cluster's host name appears twice in the FQDN; as I read it, the bundle concatenated host name and a domain that already contained it. The <networking> element of the site 2 config.xml has host name dc2-s-xblb001 and domain adm.example.net |
HA state: ha | The cluster runs as a pair |
info/ha.txt | The DRBD state seen from both nodes, then the redundant-heartbeat settings twice: every field (ip, other_ip, gw, enabled) is empty for the interfaces ha0, ha1 and ha2. The "Physical interface 1 to 3" lines, which would show next-hop or heartbeat addresses, are empty too |
cl_status: "holding all resources" / "holding none resources" | The master holds the cluster addresses and services; the slave holds nothing and waits |
ha.cf: node scb1 scb2, ucast eth3 1.2.4.2 | The heartbeat configuration (as I understand it, of the classic Linux-HA heartbeat package without a cluster resource manager, crm no). The only heartbeat path is unicast over eth3, the HA port, to the other node's 1.2.4.x address; there is no second ucast line for a redundant heartbeat |
auto_failback off | After a takeover the recovered node does not take the resources back; the cluster stays on the node that took over (general knowledge of heartbeat, not from my notes) |
drbdsetup show r0 | One DRBD resource r0 on /dev/sda4 with internal metadata, replicated between 1.2.4.1:1111 and 1.2.4.2:1111, that is over the HA port. after-sb-1pri discard-secondary tells DRBD, after a split brain where one node was primary, to throw away the changes of the secondary; csums-alg "sha1" makes a resync send only blocks whose checksums differ (my reading of the DRBD 8.4 options) |
/proc/drbd | DRBD 8.4.5, protocol C (synchronous: a write completes when both nodes have it), cs:Connected, ds:UpToDate/UpToDate, oos:0: the slave was a complete copy. The ns of the master and the nr of the slave are 174 million KiB, about 166 GiB, sent and received since the counters were last reset |
Master ntpq peers | Both Infoblox appliances of site 1 (the names are cut to 15 characters by ntpq), reachable (reach 377), stratum 2, one selected (*); scb2 is the other node, at stratum 15 and not reachable |
Slave ntpq peers | 10.11.16.145 and 10.11.18.145 stay in .INIT. with reach 0: the slave never got an answer from them. Its peer scb1 (the master, at stratum 3) also has reach 0, last heard 792 seconds earlier. So at that moment the slave had no reachable time source |
The node whose DRBD local address is 1.2.4.1 is the master. The site 2 L3 sheet gives 1.2.4.1 to datacenter A, so the master was the node in datacenter A on that day; that is a conclusion from the addresses, the bundle does not name the datacenters.
Checked against One Identity Safeguard for Privileged Sessions 9.0
| As built | Today |
|---|---|
| SCB 5.0.6, DRBD between master and slave | DRBD status and a DRBD sync rate limit are still on the HA page of SPS 9.0; master and slave are now called primary and secondary. 5.0.x LTS support was discontinued on 2020-05-28 |
| A slave whose time was not synchronised | The alert xcbTimeSyncLost is still in the 9.0 alert list, as is xcbHaNodeChanged for a takeover |
The research did not cover the heartbeat software or the files behind it, so I cannot say whether a 9.0 support bundle would contain ha.cf and drbdsetup output in the same form.