LINUXOR.SK ... open source notes ...

Balabit 01 - Requirements and concept

category: solutionz · date: 2018-12-31 · updated: 2026-10-03 · author: LALA

Balabit SCB Solution · Next: Logical design and integrations

The organisation let several partner companies administer parts of its infrastructure, and the access of all these people had to be controlled and recorded somewhere they could not reach themselves. The tool chosen was a Balabit Shell Control Box (SCB) 5 LTS: an appliance that sits between the administrators and the servers, terminates their SSH and RDP sessions, records everything into audit trails and lets only the allowed channels through. This first part covers what the box is, the questions that were asked before the design was written, the requirements and constraints the design set, the users and their roles, and the conceptual picture that everything else in this Solution fills in. My sources are the design document of the site 1 cluster (version 0.5 of 2017-12-01, still marked as a draft), the pre-design analysis that came before it, and my working notes.

What the SCB is

The SCB inspects traffic on the application layer. It speaks SSH (with forwarded X11), SCP, SFTP, RDP, HTTP, Citrix ICA, Telnet, VMware Horizon View and VNC, and simply forwards everything else. For the protocols it knows it acts as a man in the middle: the connection is cut into two halves, client to SCB and SCB to server, both are decrypted, and nothing passes directly from one side to the other. What passes is written into audit trails, which the SCB can encrypt, timestamp and sign, so that an administrator who has root on the target still cannot touch the record of what he did there. The trails can be replayed like a film, and an indexer (with OCR for graphical sessions) makes the text that appeared on the screen searchable.

What may pass is decided by policies. A connection policy is selected by the source, the destination address and the destination port; it then references the other policies: a channel policy (which channels, for example a shell, SCP, or drawing and clipboard in RDP, and for whom), an audit policy, an authentication policy, an LDAP policy, backup and archive policies, and optionally a usermapping policy. The SCB compares the connection policies in their order and applies the first one that matches. How these hang together in this installation is in Modes of operation and connections.

The design document gives the reason for the box in one sentence that I still find the best summary: administration is done over encrypted protocols, so it is hard to audit, and if the logs are kept on the administered system, a skilled administrator or an attacker can manipulate them. The SCB moves the record to a separate layer that the administrators do not control.

Before the design: the analysis

Before I wrote the design I went through a list of questions with the architect, the security architect, the Windows and Linux administrators and the logging engineer. The answers decided most of what follows.

QuestionAnswer recorded in the analysisWhat the design made of it
Transparent or non-transparent mode?Non-transparent (proxy) modeNon-transparent, see Article 3
Inband destination selection for RDP, SSH, SCP?Yes; SCP only for the file servers; SFTP left as "????"Used for all connections, SFTP included
Gateway authentication on the SCB's web interface?Not wanted: "the operator directly uses the SSH, RDP client and nothing more"; yet the RDP questions want further RDP channels filtered, which, the analysis says, needs gateway authenticationNot used on any connection
Where are the users?Active Directory (a RADIUS server exists too)AD over LDAPS, see Article 8
Backup protocol?Rsync at first; I recommended SMB/CIFS to the NetApp as the target solutionNFS in the end, see Article 11
Sensu, SIEM, NetApp integration?Yes to all three; security alarms must end up in the SIEMSNMPv3 for Sensu, syslog to the central server
Certification authority?Exists, certificate validity 365 days, a CRL existsBy design only for IPMI (C3 below); the certificate folder also has an organisation-CA certificate for the cluster's web name, see Article 9
Infoblox for DNS and NTP?YesYes
Syslog transport?TCP, according to the logging engineerLegacy BSD syslog over TCP, no TLS
Credential store for RDP?"Not considered for now"Not used
Real-time alerting and blocking?Alert, do not block (SNMP trap preferred)System alerts go by e-mail, none as SNMP traps; the one content policy, no-WINSCP, would terminate the session: in site 1 it is defined but no channel policy uses it, in site 2 both shell policies do
Who goes through the SCB?Only partner 1; partner 2's users connect directlyThe design prepares partner 1, the organisation and partner 2

Two notes in the analysis matter later. The first is about granularity: the SCB can apply policies per target server, but if the target is a jump server, it only sees the connection to the jump server, not what the user does from there. The second is about the source address that the target server sees (the SCB's, the client's or a fixed one); the analysis asked the question and left it unanswered, and the built connections use the SCB's own address.

The last row turned out to be the most temporary. In the design, partner 2's users already appear, and the user access guide of November 2018 has connections for partner 1, the cloud team of partner 1, partner 3, the organisation and partner 4; the connection summary has empty sheets for partners 2 and 5.

Requirements

The design lists fifteen functional requirements, two non-functional ones, four constraints and one assumption.

IDRequirement
FR1, FR2Monitor user activity; record sessions and play them back
FR3The user selects the destination server
FR4Support SSH, SCP/SFTP, X11, RDP, VNC, Citrix ICA, Telnet and HTTP/HTTPS
FR5Control the channels inside a protocol (terminal and SCP in SSH, drawing and clipboard in RDP)
FR6Management web interface over HTTPS only
FR7, FR8Users and groups in Active Directory, integrated over LDAPS
FR9All certificates signed by the internal certification authority
FR10Access for third-party administrators over IPv4 and IPv6
FR11, NR2High availability (the same requirement twice)
FR12, FR13Automatic archiving and backup, stored on the NetApp array
FR14Configured for partner 1's users from day one
FR15Prepared for the organisation's users, who need not go through the SCB from day one
NR1"Solution must meet international standards"

The constraints are the limits the product and the environment set, and the only assumption is about licences.

IDConstraint or assumption
C1The SCB does not support NIC teaming or bonding
C2Jump servers are not the optimal deployment model, because inspection per target server is not possible
C3The internal CA cannot issue timestamping (TSA) certificates; only the IPMI certificates are signed by it, everything else by the SCB's own CA
C4XRDP, the open-source RDP server, is not compatible with the SCB
A1All software components correctly licensed

Read against the rest of the design, three of these do not hold as written. FR9 says all certificates come from the internal CA, and C3, in the constraints table, says the opposite for all but the IPMI certificates; C3 is what was built, and Certificates and keys tells why. FR4 lists eight protocols, while chapter 7 of the design disables HTTP, Telnet and VNC and configures only SSH and RDP connections; the requirement describes what the product can do rather than what was needed. And C2 is a constraint the design accepted knowingly: every user connection ends on a jump server, and the SCB records the session there, not the hop from the jump server onwards. NR1 cannot be checked. The second assumption, A2, is an empty row.

C1 had a practical consequence that is easy to miss: without bonding, each logical network sits on one physical port, and resilience has to come from the HA pair, not from the links. My early TODO list, in notes headed "BALABIT notes - 4.0.7.a", still has "teaming/bonding" and "configure the prod interface (in HA mode/team/bond)" on it; the answer was C1. The ports and cabling are in Hardware, cabling and networks, the pair in High availability.

Users and roles

The design divides the people twice: by the company they work for and by what they do with the SCB.

By companyWho
PARTNER1Users of partner 1, the first and, at the start, the only group to go through the SCB
ORGUsers of the organisation itself
PARTNER2Users of partner 2; the design says these are users of the integrator

Across the companies run three roles, and each role maps to one kind of AD role group.

By roleWhat they getAD role group
AdministratorsThe SCB web interface with all rightsSCB_ADM
OperatorsThe SCB web interface read-only, plus the troubleshooting pageSCB_OPS
UsersSessions to the jump servers through the SCB, nothing on the SCB itselfSCB_USERS_ORG, SCB_USERS_PARTNER1, SCB_USERS_PARTNER2, SCB_USERS_GUESTS

Administrators and operators exist only on the organisation's side (SCB_ADM_ORG, SCB_OPS_ORG); every company gets its own group of users. Two local accounts stay as exceptions: admin as the last resort for the web interface when Active Directory does not answer, and root for the console. The groups, their nesting and the access control page are in Active Directory and access control.

mermaid
flowchart LR
  adm["Administrators, SCB_ADM"]
  ops["Operators, SCB_OPS"]
  usr["Users, SCB_USERS_ORG, _PARTNER1, _PARTNER2, _GUESTS"]
  loc["Local admin, emergency only"]
  root["Local root"]
  web["SCB web interface, HTTPS on 10.11.16.81"]
  con["SCB console, SSH on 10.11.16.81"]
  pro["SCB user address 10.11.16.65, SSH, RDP, SCP, SFTP"]
  jmp["Jump servers"]
  tgt["Target servers and devices"]
  adm -- "read and write, all pages" --> web
  ops -- "read, all pages" --> web
  loc -- "if AD fails" --> web
  root --> con
  usr --> pro
  pro -- "per partner connection" --> jmp
  jmp -- "outside the SCB's view" --> tgt

Later, more companies came. The test accounts in my notes cover groups for partner 1, its cloud team, the organisation, partner 3, partner 4 and partner 5 (SCB_USR_CLOUD_PARTNER1, SCB_USR_PARTNER3, SCB_USR_PARTNER4, SCB_USR_PARTNER5), none of which is in the design. Who these partners are beyond "a partner company" the documents do not say, and it does not matter for the design: each new partner meant its own connections and its own AD group, and in the material that shows it, its own backup and archive policies.

The conceptual design

The concept is short. The SCB is the single entry point for management access in the management LAN of site 1. A user connects to the SCB, the SCB connects him to a jump server, and from the jump server he reaches the production infrastructure. Jump servers are grouped by protocol (SSH, RDP, SCP/SFTP), by placement (in the management LAN or in an out-of-band network), by company (the organisation's, partner 1's, or shared) and by site. The out-of-band jump servers outside site 1 are reached over the organisation's backbone. Active Directory holds the users and groups, and the NetApp stores the backups and archives.

The conceptual figure of the design, redrawn:

mermaid
flowchart LR
  p1u["PARTNER1 VPN USERS"]
  adu["admin VPN USERS"]
  subgraph s1["Site 1 - management LAN"]
    scb["Balabit SCB"]
    ds["Directory Services"]
    nas["Network Attached Storage"]
    pi["Production Infrastructure"]
    subgraph orgj["ORG_JUMP_SERVERS"]
      ordp["RDP_JUMP_SERVER"]
      ossh["SSH_JUMP_SERVER"]
    end
    subgraph p1j["PARTNER1_JUMP_SERVERS"]
      prs["RDP+SSH_JUMP_SERVER"]
    end
    subgraph shj["SHARED_JUMP_SERVERS"]
      scp["SCP_JUMP_SERVER"]
    end
  end
  bb["ORG BACKBONE"]
  oob1["Site 1 - OOB"]
  oob2["Site 2 - OOB"]
  oob4["Site 4 - OOB"]
  oobx["Remote Location X - OOB"]
  p1u -- "SSH, RDP, SCP, SFTP" --> scb
  adu -- "SSH, RDP, SCP, SFTP" --> scb
  scb -- "RDP" --> ordp
  scb -- "SSH" --> ossh
  scb -- "SSH" --> prs
  scb -- "RDP" --> prs
  scb -- "SCP, SFTP" --> scp
  scb -- "LDAPS" --> ds
  scb -- "CIFS" --> nas
  scb --> oob1
  scb -.-> bb
  bb -.-> oob2
  bb -.-> oob4
  bb -.-> oobx

Each out-of-band box in the original holds a group of partner 1 jump servers, a group of the organisation's jump servers and its own production infrastructure; I left the groups out to keep the picture readable. The production infrastructure in site 1 has no line in the original either: nothing in the concept connects to it except through a jump server.

The line to the storage says CIFS. That was the plan in December 2017, and it did not survive: the backup over SMB failed and the exports became NFS, as told in SVMs and NFS of the NetApp Solution. The backup and archive tables of the same design document already say NFS, so the figure and the text of the concept are older than the tables.

Loose ends

Checked against One Identity Safeguard for Privileged Sessions 9.0

Everything in this build is out of support. One Identity announced the acquisition of Balabit on 2018-01-17, within weeks of the design; the product is now One Identity Safeguard for Privileged Sessions (SPS); the newest documents are those of SPS 9.0, of September 2026. The individual settings are compared in the Articles and Config documents that hold them; this is the summary for the concept.

As built, 2017 to 2018Today
Balabit Shell Control Box 5 LTS, 5.0.3 to 5.0.65.0.x LTS: full support until 2019-05-28, limited support until 2020-05-28, then discontinued. The way up goes through every LTS: the latest 5.0.x, then 6.0 LTS, 7.0 LTS, 8.0 LTS and 9.0; downgrading is not possible
Two SCB T-10 appliances per siteT-Series hardware: last day sold 2019-06-30, end of support 2024-07-31; SPS 8.0 is not supported on T-Series. SPS 9.0 installs on generic hardware or certified Dell and HPE servers and documents a data migration from one SPS instance to another, which is the practical way off a T-10
Host-based licence, 1000 protected hosts, jump servers and the hosts behind them all countedSPS 9.0 needs no licence file for core functionality; the licence page has been removed
C1: no NIC teaming or bondingNo bonding, teaming or link aggregation is documented in 9.0 either
C3: server and TSA certificates from the same CA, which the organisation's CA could not provideUnchanged: the server and TSA certificates must be issued by the same CA, and the TSA certificate's Extended Key Usage must be critical Time Stamping. The vendor now recommends certificates from one's own PKI over those generated on the box, because those cannot be revoked
Gateway authentication not used; users authenticate to the jump server through the SCBThe vendor's security checklist, in SCB 5 and SPS 9.0 alike: "Always use gateway authentication to authenticate clients. Do not trust the source IP address of a connection, or the result of server authentication."

The architecture is the same in substance: a boot firmware that provides HA and starts the core firmware, a master/slave pair (now called primary and secondary) with DRBD, a redundant heartbeat and next-hop monitoring, and most of the policies of this design under the same names; user lists are announced for removal. What would change most in a rebuild today is the hardware and the path to it, not the concept.

← solutionz