Balabit 04 - Hardware, cabling and networks
Balabit SCB Solution · Previous: Modes of operation and connections · Next: High availability
An appliance that every administrator's session has to pass is only as available as its cables. This part is about the two boxes themselves, their six network ports, where each port went and which VLAN it carried, the five networks the cluster lives in, the names those networks got in DNS, and an earlier network variant that was dropped. It is also the part where the labels printed on the appliance and the names in my documents disagree most, so I say which name means what as I go.
The appliance
Each of the two datacenters of site 1, A and B, got one SCB T-10. The design lists the hardware like this:
| Parameter | Value |
|---|---|
| Product | SCB T-10, dc1-a-ablb001 in datacenter A, dc1-b-ablb001 in datacenter B |
| Power supply | Redundant |
| Processor | 2 × Intel Xeon E5-2630 v2 @ 2.6 GHz |
| Memory | 8 × 4 GB |
| Disks | 13 × 1 TB on an LSI 2208 RAID controller with 1 GB cache |
| Out-of-band | IPMI |
Under the Balabit front it is a SuperMicro server. One of the first questions of my pre-design analysis was what the other ports on the right side of the back panel (beside the network ports) are; the answer was that they are the external SAS ports of the RAID card. They play no part in the design.
Six ports and the labels on them
The back panel has the IPMI port, two USB ports, a serial port and VGA on the left, and the six network ports on the right: four 1 Gb Ethernet ports and two 10 Gb SFP+ ports.
flowchart LR t10["Balabit SCB T10, back panel"] subgraph lft["Left side"] ipmi["IPMI, 100 Mb Ethernet"] usb["USB 1, USB 2"] ser["SERIAL"] vga["VGA"] end subgraph rgt["Right side, 1 Gb Ethernet"] p1["1 EXT, eth0, plugged"] p2["2 MGMT, eth1, plugged"] p3["3 INT, eth2, plugged"] p4["4 HA, eth3, plugged"] end subgraph sfp["Right side, 10 Gb SFP+"] p5["5 A, eth4, unplugged"] p6["6 B, eth5, unplugged"] end t10 --- lft t10 --- rgt t10 --- sfp
The labels EXT, MGMT and INT are printed by the vendor; as I understand it, they come from older SCB versions in which the external, management and internal interfaces had fixed roles. In SCB 5 every port can carry any VLAN-tagged logical interface, and I gave the logical interfaces names after their role. The result crosses over: port 1, labelled EXT, carries the in-band management network (IBM, where administrators reach the web interface and SSH), and port 2, labelled MGMT, carries the production network (PRO, where the users' sessions arrive). Port 3, INT, carries the redundant heartbeat and the backup network; port 4, HA, is reserved for the cluster.
The cabling I found and the one I recommended
When I started the analysis, the first appliance was already in the rack. According to the network engineer's table it was cabled on ports 1, 2 and 3, with port 4 (HA) not connected, and a photo from the site confirmed it, although from the operating system's point of view port 3 had no link. I recommended a minimal cabling of all four 1 Gb ports and listed what each should carry:
| Port | Purpose in the analysis |
|---|---|
| 1, EXT | Management and operation addresses MGMT_IP1 and OP_IP1; input and output in one VLAN, management in another |
| 2, MGMT | The second pair, MGMT_IP2 and OP_IP2, in two further VLANs |
| 3, INT | Redundant HA link (heartbeat only), and backup and archiving to the external storage, "connected temporarily, later moved to Port5 and Port6" |
| 4, HA | Dedicated to HA: heartbeat and synchronisation |
| 5, 6, SFP+ | Not cabled in this phase; sensible later for backup and archiving to the NetApp |
The SFP+ ports raised two questions. Would Cisco Twinax cables work? The vendor's support answered that the Intel 82599 dual-port card "supports 10GBASE-CX4 Operating Mode", so "theoretically" yes, but "We haven't tested". Multimode or single-mode? That depends on the module; the SuperMicro documentation lists both SR and LR modules, and the vendor's packaging checklist warns that SFP transceivers encoded for non-Intel hosts may not work with the 82599. The material records no modules being ordered. The backup traffic stayed on port 3 in VLAN 1020, and ports 5 and 6 stayed empty: the as-built interface page of site 1 shows them without configuration, and so does the site 2 export of September 2018.
The final cabling and the switch ports
The design's cabling pictures and the L1 sheet give every port a switch port. Both datacenters are cabled alike, with their own switches.
flowchart LR subgraph dca["Datacenter A"] a["DC1-A-ABLB001"] aoob["DC1-A-SOOB003 Gi1/0/16"] acon["DC1-A-ROOB002 port 19"] as1["DC1-A-SPRO001"] as2["DC1-A-SPRO002"] end subgraph dcb["Datacenter B"] b["DC1-B-ABLB001"] boob["DC1-B-SOOB003 Gi1/0/16"] bcon["DC1-B-ROOB002 port 19"] bs1["DC1-B-SPRO001"] bs2["DC1-B-SPRO002"] end a -- "IPMI" --> aoob a -- "SERIAL" --> acon a -- "eth0 1/EXT to Eth1/5" --> as1 a -- "eth2 3/INT to Eth1/21" --> as1 a -- "eth1 2/MGMT to Eth1/5" --> as2 a -- "eth3 4/HA to Eth1/21" --> as2 b -- "IPMI" --> boob b -- "SERIAL" --> bcon b -- "eth0 1/EXT to Eth1/5" --> bs1 b -- "eth2 3/INT to Eth1/21" --> bs1 b -- "eth1 2/MGMT to Eth1/5" --> bs2 b -- "eth3 4/HA to Eth1/21" --> bs2
Ports 5 and 6 (eth4, eth5) are unplugged on both appliances. The switch side, as configured:
| Appliance port | Switch port | VLAN and mode |
|---|---|---|
| IPMI | DC1-x-SOOB003 (Catalyst 2960-X) Gi1/0/16 | 12, access |
| 1, EXT | DC1-x-SPRO001 (Nexus 9396PX) Eth 1/5 | 1010, trunk |
| 2, MGMT | DC1-x-SPRO002 (Nexus 9396PX) Eth 1/5 | 1009, trunk |
| 3, INT | DC1-x-SPRO001 Eth 1/21 | 1018 and 1020, trunk, native VLAN 1018 |
| 4, HA | DC1-x-SPRO002 Eth 1/21 | 1016, access |
| Serial | DC1-x-ROOB002 (Cisco 2901) async port 19, patch panel 20 | none |
The full tables of both datacenters are a Config document: Switch ports (site 1).
The spread over two switches looks deliberate, although the design does not explain it: management and redundant heartbeat on one Nexus, production and HA on the other, so a single switch failure leaves one heartbeat path and one network side alive. The SCB does not support NIC teaming or bonding (constraint C1 of the design), so this is all the redundancy a single appliance gets; the rest comes from the second node. The primary HA link is not a cable between the two boxes but an access port in VLAN 1016 on each side, so the cluster's most important link depends on the network between the two datacenters.
One detail I would check today. Port 3 is a trunk with native VLAN 1018, and the SCB sends VLAN 1018 tagged on eth2.1018. As I understand Nexus trunks, the switch sends native-VLAN frames untagged, which a tagged interface does not receive. The documents do not record a test of the redundant heartbeat, and the site 2 cluster was built without one.
Networks and VLANs
The design's logical network table, with the names that the L3 sheet uses for the columns:
| Name | VLAN | Network | Purpose |
|---|---|---|---|
| OOB | 12 | 10.11.15.0/24 (A), 10.11.23.0/24 (B) | IPMI of the appliances |
| IBM | 1010 | 10.11.16.80/29 | In-band management: web interface and SSH of the SCB for administrators |
| PRO | 1009 | 10.11.16.64/29, 2001:db8:a1:c0e::/64 | Production: users' sessions to the jump servers |
| CLS2 | 1018 | 10.11.18.208/28 | Secondary cluster sync, the redundant heartbeat |
| BCK | 1020 | 10.11.18.224/28 | Backup and archive to the NetApp |
| CLS1 | 1016 | 1.2.4.0/24 | Primary cluster sync, the main HA interface |
The design's picture of the logical network of the pair:
flowchart LR nodea["DC1-A-ABLB001"] nodeb["DC1-B-ABLB001"] ooba["Balabit OOB 10.11.15.0/24 VLAN-12"] oobb["Balabit OOB 10.11.23.0/24 VLAN-12"] ibm["Balabit IBM 10.11.16.80/29 VLAN-1010"] pro["Balabit PRO 10.11.16.64/29 VLAN-1009"] cls2["Balabit CLS2 HA Secondary 10.11.18.208/28 VLAN-1018"] bck["Balabit BCK Backup/Archive 10.11.18.224/28 VLAN-1020"] cls1["Balabit CLS1 HA Primary 1.2.4.0/24 VLAN-1016"] nodea -- "IPMI .29" --> ooba nodeb -- "IPMI .29" --> oobb nodea -- "1/1 81 (VIP)" --> ibm nodeb -- "1/1 81 (VIP)" --> ibm nodea -- "2/1 65 (VIP)" --> pro nodeb -- "2/1 65 (VIP)" --> pro nodea -- "3/1 209" --> cls2 nodeb -- "3/1 210" --> cls2 nodea -- "3/2 225 (VIP)" --> bck nodeb -- "3/2 225 (VIP)" --> bck nodea -- "4/1 1" --> cls1 nodeb -- "4/1 2" --> cls1
"VIP" marks the cluster addresses, which the master node holds and the slave takes over; .209/.210 and 1.2.4.1/.2 belong to one node each. The /29 networks hold six hosts: the cluster address, the gateway and room for little else, which is all a bastion needs. The 1.2.4.x addresses are the fixed HA addresses of the cluster link, marked (FIX) on the HA page (the analysis sketch has them too); they are not private address space, which, as I see it, is harmless only if VLAN 1016 carries nothing else and is not routed.
The out-of-band network also has IPv6, and here the design contradicts itself: its network table gives 2001:db8:b1:ff3::/64 to datacenter A and 2001:db8:a1:ff3::/64 to B, while the IPMI pages give node A an a1 address and node B a b1 address. More in IPMI out-of-band management.
Interfaces, routes and names as built
Chapter 7 of the design records the Network page as configured. The interfaces match the sheet: eth0.1010 IBM 10.11.16.81/29, eth1.1009 PRO 10.11.16.65/29 and 2001:db8:a1:c0e::f:1/64, eth2.1018 CLS2 10.11.18.209/28, eth2.1020 BCK 10.11.18.225/28. Host name dc1-s-xblb001, nickname dc1-s-xblb001pro, search domain adm.example.net, DNS servers the two Infoblox appliances. The whole page is a Config document: Network settings (site 1).
The routing table has five specific routes through 10.11.16.86 in the management network (admin VPN 10.11.20.0/25, Sensu, Active Directory, both Infoblox segments), an IPv6 default route through 2001:db8:a1:c0e::1 in the production network, and an IPv4 default gateway 10.11.17.70. That last address is in none of the SCB's networks, so, as I read it, the appliance could not have reached it. The site 2 export has 10.12.16.70, inside its production /29, which makes 10.11.16.70 the likely intended value in site 1; that is my inference, and the material has no site 1 export to confirm it. The design behind it is sensible: administrators and infrastructure services are routed back through the management side, everything else, the users above all, through production.
Names in DNS
The design asked the Infoblox administrators for A and AAAA records, and a PTR record for each; the list is a Config document: DNS records (site 1). Those records use an older naming scheme in which dc1-s-xblm001 was management and dc1-s-xblb001 operation. The L3 sheet, chapters 7 and 8 and the IPMI pages use the newer one: dc1-s-xblb001 for management at 10.11.16.81, dc1-s-xblb001pro for production, dc1-x-ablb001m for the IPMI modules, cl01/cl02 for the heartbeat addresses, bck for backup. The users' name scb.example.net points to the production address in both. The certificates still carry dc1-s-xblm001, and my operation how-to has administrators add both names (and scb.example.net, pointing at the management address) to their workstation's hosts file, which suggests that the records were not there yet when I wrote it.
The variant that was dropped
An earlier variant of each node's networking, drawn once per datacenter (the picture for datacenter A below; the one for B is the same), had two networks per role: primary IBM and PRO on port 1 and "secondary, NAT-ed" ones on port 2, each in its own VLAN, besides the HA networks CLS1 and CLS2.
flowchart LR scb1["Balabit SCB1, management LAN DC1-A"] ibm1["Balabit IBM1 Primary 10.11.16.80/29 VLAN-1010"] pro1["Balabit PRO1 Primary 10.11.16.64/29 VLAN-1009"] ibm2["Balabit IBM2 Secondary/NAT-ed 10.11.16.88/29 VLAN-1030"] pro2["Balabit PRO2 Secondary/NAT-ed 10.11.16.72/29 VLAN-1029"] cls2["Balabit CLS2 HA Secondary 10.11.18.208/28 VLAN-1018"] bck["Balabit BCK Backup/Archive 10.11.18.224/28 VLAN-1020"] cls1["Balabit CLS1 HA Primary 1.2.4.0/24 VLAN-1016"] oob["Balabit OOB 10.11.15.0/24 VLAN-12"] scb1 -- "1/1 81" --> ibm1 scb1 -- "1/2 65 (VIP)" --> pro1 scb1 -- "2/1 89" --> ibm2 scb1 -- "2/2 73 (VIP)" --> pro2 scb1 -- "3/1 209" --> cls2 scb1 -- "3/2 225 (VIP)" --> bck scb1 -- "4/4 1" --> cls1 scb1 -- "IPMI 29" --> oob
As I read it, it was the analysis' way of surviving a port or switch failure without bonding: every role on both ports, with addresses MGMT_IP1/MGMT_IP2 and OP_IP1/OP_IP2, and DNS to choose between them. The first certificate request I wrote still lists all four addresses, 10.11.16.81, 10.11.16.89, 10.11.16.65 and 10.11.16.73, with the names scb1.example.net and scb2.example.net (see Certificates and keys). The picture does not say where the translation would have happened. The analysis itself concluded that DNS gives no usable failover for this, and the final design gives each role one address on one port and leaves failover to the cluster; that discussion is in High availability.
Site 2
Site 2 has no design document, only its own L1–L3 sheet and the exported configuration from a support bundle of 2018-09-17. The sheet is the site 1 sheet with 10.12.x addresses, dc2- names and the IPv6 address 2001:db8:a2:c0e::f:1; the cabling is port for port the same on the DC2-x switches. The copy left traces: the Location column still says DC1-A/DC1-B, and scb.example.net appears as the production name of site 2 too. See Network sheet L1-L3 (site 1) and Network sheet L1-L3 (site 2).
The export shows what was really built in site 2, and it differs from the sheet: port 3 carries only BCK, without the CLS2 interface for the redundant heartbeat; there is no IPv6 route at all; DNS, Active Directory and NTP are those of site 1, reached through explicit routes; there is a route to the site 1 Sensu segment, but the SNMP agent answers only three site 2 addresses, 10.12.16.129 to .131. The element is a Config document: Networking element of the site 2 config.xml.
Loose ends
- The default gateway
10.11.17.70of site 1 is outside every network of the cluster;10.11.16.70is my guess. - The IPv6 networks of the two out-of-band segments are swapped between the design's network table and the IPMI pages.
- The interface-summary picture in the design shows
eth0.1009for the production interface; the sheet and the interface table sayeth1.1009. - Two naming schemes for the same addresses; the heartbeat names are spelt
dc1-a-ablb001cl02anddc1-b-ablb001cl2in the same sheet. - Port 3 as a trunk with native VLAN 1018 while the appliance tags VLAN 1018.
- No switch configuration of site 2 and no firewall rules in the material, so the network side that forces traffic through the SCB cannot be shown.