LINUXOR.SK ... open source notes ...

Balabit 14 - Operations, upgrades and troubleshooting

category: solutionz · date: 2018-12-31 · updated: 2026-10-03 · author: LALA

Balabit SCB Solution · Previous: End-user access

An appliance hides its operating system until something breaks, and then the hidden part is all that matters. This last part collects what I needed to run the SCB day to day: how to reach it, the two firmwares and their shells, the support and debug bundles, and the problems that the notes kept with their fixes, most of them from the vendor's support. It also has the upgrade history of the site 2 cluster, the only one recorded, and the TODO lists that were still open. The material for it is the least finished of all: an operation how-to with half of its sections reading "TODO TODO", troubleshooting notes and an IPMI and debug how-to.

Getting to the appliance

Administration started with a VPN connection into the organisation's network, made with the Cisco AnyConnect client and the group "CP-ORG-PROD". The design allows the administrators' VPN segment ORG_VPN to the management address 10.11.16.81 on 443 and 22; the notes add "TODO: 10.11.20.0/25 - VPN segment for admins and operators (ORG)". The section "Import CA certificates to PC" of the how-to was never written, so the browser's trust in the SCB's certificate is not documented; see Certificates and keys.

The names of the cluster went into the workstation's hosts file, all three on the management address: Workstation hosts file. Two of the three lines contradict the design's DNS table, which puts dc1-s-xblb001.adm.example.net and scb.example.net on the production address 10.11.16.65; the Config document explains what that breaks. The web interface then answered on https://dc1-s-xblb001.adm.example.net, https://dc1-s-xblm001.adm.example.net, https://scb.example.net or https://10.11.16.81. Administrators logged in with their AD account as account@ad.example.net; the local admin was kept as the only enabled local web account, for the day AD would not answer. The design supports the current Firefox and Chrome, Edge and Internet Explorer 11; replaying an audit trail in the browser needed Internet Explorer 11 with the WebM plugin, otherwise the Audit Player or the Desktop Player of Audit trails.

The SSH console and the two firmwares

SSH to 10.11.16.81 was meant "exclusively for troubleshooting purposes". The design gives the local SSH service one key exchange (diffie-hellman-group-exchange-sha256), two ciphers (aes256-ctr, aes128-ctr) and two MACs (hmac-sha2-512, hmac-sha2-256), and the vendor's brute-force protection: after 20 failed logins the address is blocked for ten minutes. Logging in to the console menu locks the web interface, and the menu is available only when nobody is logged in to the web interface.

The menu offers "Shells", then "1. Boot shell" or "2. Core shell". My operation how-to explains the difference in one sentence: "The boot firmware boots up SCB, provides high availability support, and starts the core firmware. The core firmware, in turn, handles everything else: provides the web interface, manages the connections, and so on." Most procedures in the notes start by checking which one the shell is in, with cat /etc/firmware-type, because the same files sit at different paths in the two (/opt/scb/var/debug/ in the core is /mnt/firmware/opt/scb/var/debug/ from the boot shell). In 5 LTS, as the troubleshooting notes put it, the core firmware's services run in a systemd-nspawn container of the boot firmware, and the notes keep the one check for it, run as root in the boot shell (no output kept):

bash
$ systemctl status systemd-nspawn@scb-core.service

There are two ways from the boot shell into the core: the core-shell command, which as I understand it enters the running container, and chroot /mnt/firmware, which works on the core firmware's files directly and is what the support used when the core firmware was reported as not running. Which command ran where:

mermaid
flowchart TB
  adm["Administrator, SSH to 10.11.16.81 as root"]
  menu["Console menu, Shells"]
  subgraph bootfw["Boot firmware, firmware-type boot"]
    b1["HA, DRBD and heartbeat"]
    b2["systemctl status systemd-nspawn@scb-core.service"]
    b3["xcbclient self xcb_check_core_files"]
    b4["rm of tainted files in /mnt/drbd/private/root"]
    b5["ntpq, ntpd -q -g, date via ssh scb-other"]
  end
  subgraph corefw["Core firmware, container scb-core, firmware-type core"]
    c1["Web interface, lighttpd"]
    c2["Connections and audit trails"]
    c3["console.php troubleshooting start and stop"]
    c4["rm -rf of old debug data in /opt/scb/var/debug"]
    c5["generate-core-auth-key.sh, inside the chroot"]
  end
  adm --> menu
  menu -- "1. Boot shell" --> bootfw
  menu -- "2. Core shell" --> corefw
  bootfw -- "core-shell" --> corefw
  bootfw -- "chroot /mnt/firmware" --> corefw

The design and the how-to disagree about the account. The design's text says the remote console is "allowed only for local SCB user root"; its own table for the console says "Only enabled account is admin". The IPMI and debug how-to logs in as root; the operation how-to names no account.

Support and debug bundles

In two of the support's answers a debug bundle came first: one asks for it, the other thanks for it. The web interface makes it under Basic Settings – Troubleshooting – System Debug, "Collect and save current system state info". When the web interface itself was the problem, the bundle had to come from the core shell with console.php --troubleshooting-start and --troubleshooting-stop, and when even that failed, the leftovers of earlier debug runs in /opt/scb/var/debug/ had to be removed first. The commands are in Debug bundle from the command line. The sections of the operation how-to that should have described both ways, and the support portal, stayed "TODO TODO".

One bundle survives: that of the site 2 cluster from 2018-09-17. Its folder is named after a problem with RDP002, presumably the RDP jump server dc2-a-vcrdp002; the notes say nothing more about it. It holds the exported config.xml that the other parts of this Solution use as the as-built configuration of site 2, the HA and NTP state of High availability, and the firmware history below.

The upgrade from 4 to 5

The notes are headed "BALABIT notes - 4.0.7.a". As I read that heading, the appliances arrived with SCB 4 LTS and were upgraded to 5 LTS, the release the design is written for. The first item of the SCB TODO list is "upgrade (second node)". The design explains the choice of an LTS release: supported for three years after publication and one year after the next LTS release, whichever is later, with maintenance releases that contain only bug fixes and security updates.

After the upgrade the web interface did not come up. The error log of the web server, read inside the core firmware, contained, among others:

output 2 lines
2017-07-12 09:25:15: (mod_fastcgi.c.1112) the fastcgi-backend /usr/lib/cgi-bin/php5 failed to start:
2017-07-12 09:25:15: (server.c.1022) Configuration of plugins failed. Going down.

The new core firmware had PHP 7.0, the web server's configuration still pointed to PHP 5. The next morning I linked /usr/lib/cgi-bin/php5 to /usr/lib/cgi-bin/php7.0 and the web server started: Web server fix after the upgrade from 4 to 5. The vendor's support, in the answer that explained how to make a debug bundle from the command line, saw the same symptom: "It seems the lighttpd did not start and that is the reason you do not have a web UI." Whether that answer was about this case the notes do not say.

Two more answers of the support are undated. In one, the login said "Core fimware not running", and the support's fix was to run /opt/scb/bin/generate-core-auth-key.sh inside the chroot: Core auth key procedure. In the other, the firmware was reported as "TAINTED".

Tainted firmware

"Tainted" means that files belonging to the firmware differ from what the firmware shipped; System Monitor shows it, and so does Basic Settings > System > Version details. From the boot shell, xcbclient self xcb_check_core_files lists the files and xcbclient self xcb_check_boot_files does the same for the boot firmware. One check found six files under /mnt/drbd/private/root/, among them the web server's lighttpd.conf and the SSH server's configuration and binary, and I removed them; on another occasion the support named two blkid cache files, "possibly" left over from the upgrade. Both are in Finding and removing tainted files. The lighttpd.conf in that list makes me think that the PHP 5 problem was a version 4 configuration left in the persistent area, and that removing it was the real repair of which the symbolic link was only a workaround; the notes do not connect the two.

Clock out of sync

Basic Settings > Date & Time once showed the yellow banner "Slave is out of sync with the master". Separately, the support sent a manual synchronisation from the boot shell: measure the offset to the other node with date -R; ssh scb-other date -R;date -R, look at the peers with ntpq, stop ntpd, run ntpd -q -g once and start it again: Manual NTP synchronisation. The notes do not say which case it answered. The site 2 bundle of September 2018 shows a slave whose time servers did not answer, both Infoblox servers .INIT. and unreachable, while the master synchronised from them. The banner and the procedure are undated, so I cannot say that they belong together or to the same event as the bundle. A possible reason why the slave could not reach the time servers is discussed in High availability.

Upgrade history of site 2

The site 2 cluster kept its upgrade logs, one file per run, with the time in the file name. They tell a short story: upgrade runs from 5.0.0 to 5.0.3b in two days in January 2018, then 5.0.3b run four more times until February, then every maintenance release within weeks. The site 2 SCB's own CA certificate is dated 6 March 2018, so the initial configuration may be later than the first runs.

mermaid
flowchart LR
  u1["2018-01-16, 5.0.0, 5.0.1, 5.0.2"]
  u2["2018-01-17, 5.0.3, 5.0.3a, 5.0.3b twice"]
  u3["2018-01-25 to 02-05, 5.0.3b four more runs"]
  u4["2018-02-14, 5.0.4"]
  u5["2018-04-03, 5.0.4b"]
  u6["2018-04-10, 5.0.5"]
  u7["2018-05-09, 5.0.5a"]
  u8["2018-06-05, 5.0.6"]
  u9["2018-07-24, 5.0.6 again"]
  u10["2018-09-17, support bundle at 5.0.6"]
  u1 --> u2 --> u3 --> u4 --> u5 --> u6 --> u7 --> u8 --> u9 --> u10

Why 5.0.3b needed six runs and 5.0.6 two, the logs that would say it are not in my selection, and the file names do not say which node ran them. The boot firmware directory keeps five slots, 5.0.4 to 5.0.6, and its active, current and previous links all point to 5.0.6. In the bundle's summary the core firmware calls itself "Balabit Privileged Session Management 5.0.6"; the documents do not explain the name. The whole listing is in Firmware versions and upgrade history, site 2. For site 1 there is only the design's version page, 5 LTS (5.0.3) built on 2017-11-11, in System settings.

What was still open

The notes end with TODO lists that were never ticked off in writing:

Loose ends

Checked against One Identity Safeguard for Privileged Sessions 9.0

As builtToday
Balabit SCB 5 LTS, 5.0.3 to 5.0.6The product is now One Identity Safeguard for Privileged Sessions (SPS). 5.0.x LTS lost support on 2020-05-28. The current LTS is 8.0 (8.0.2.1 LTS, June 2026); the newest release is 9.0 (September 2026)
Two pairs of T-10 appliancesThe T-Series reached End of Support on 2024-07-31, and SPS 8.0 is not supported on T-Series hardware. SPS 9.0 installs on generic hardware or on certified Dell and HPE servers, with a documented data migration from one SPS instance to another
Upgrade 4 to 5, then 5.0.x maintenance releasesFrom 5 LTS the path goes through every LTS: the latest 5.0.x, 6.0 LTS, 7.0 LTS, 8.0 LTS, then 9.0. "Downgrading to a previous SPS version is not possible."
Boot firmware and core firmware, console menu "Shells"The same sentence about the two firmwares is in the 9.0 guide, and the console menu still offers the core and boot shells "only required in certain troubleshooting situations". The internals used in my notes (systemd-nspawn, /etc/firmware-type, core-shell) were not documented then and are not now
Local SSH for troubleshooting, 20 failed logins block for ten minutesThe same algorithms and the same brute-force rule. Completing the Welcome Wizard disables SSH access, and from 7.4 the local console accepts no SHA-1 signature algorithms, so a client needs RFC 8332 support (OpenSSH since 7.2, PuTTY since 0.75)
Tainted firmware, alert xcbFirmwareTaintedThe firmware status in the web interface shows "Corrupted" or "Tainted"; the alert is now xcbFirmwareError
Debug bundle in the web interface or with console.phpBasic Settings > Troubleshooting > Create support bundle, with Start, reproduce, Stop for a specific error
Replay in Internet Explorer 11 with the WebM pluginThe plugin is not needed since 6.10, Internet Explorer 11 is not supported since 6.13.0; supported browsers are Chrome, Firefox, Safari and Edge
Master and slaveRenamed primary and secondary node; DRBD, redundant heartbeat and next-hop monitoring remain

Nothing in the procedures of this part can be carried over without checking it against the current documentation, and the T-10 hardware is out of support, with SPS 8.0 and later not supported on it.

What I would do differently

The fixes are in the notes and the procedures they belonged to are not. I would write the upgrade as a runbook before the first upgrade: firmware type, the tainted check, the web interface and the NTP peers on both nodes before and after, and a debug bundle at the end. And I would have looked for the cause of the PHP 5 error before linking around it.

← solutionz