Balabit 07 - Basic settings, logging and monitoring
Balabit SCB Solution · Previous: IPMI out-of-band management · Next: Active Directory and access control
An appliance that records what everybody else does has to be watched itself: it must tell somebody when its disks fill up, when a backup fails or when the slave loses time, and its own logs must leave the box. This part covers the pages of Basic Settings other than network and high availability: local services, Management, Alerting & Monitoring, Date & Time and System. The values come from chapter 7 of my design document, which shows only what was changed from the defaults, and from the syslog-ng configuration of the site 1 cluster; the exported configuration of the site 2 cluster serves as a check.
What leaves the box
Every service the appliance talks to sits in the management LAN, and every local service listens on one address, the cluster address 10.11.16.81 in the in-band management network (IBM, VLAN 1010). The production address 10.11.16.65, where users arrive, offers nothing of the appliance itself.
flowchart LR subgraph scb["Cluster dc1-s-xblb001, IBM 10.11.16.81"] sng["syslog-ng"] mail["Mail and alerts"] ntpc["NTP client"] res["DNS resolver"] agent["SNMPv3 agent, port 161"] end sys["Central syslog dc1-a-vcsys001, 10.11.17.113"] siem["SIEM"] mx["SMTP dc1-a-vcmsx001"] grp["scb_mail_group"] ib1["Infoblox dc1-a-adnsvip001, 10.11.16.145"] ib2["Infoblox dc1-b-adnsvip001, 10.11.18.145"] sensu["Sensu dc1-a-vcsns001, 10.11.16.129"] sng -- "legacy-TCP 514, no TLS" --> sys sys -. "planned in the analysis" .-> siem mail -- "TCP 25, STARTTLS, as scb_mail" --> mx mx --> grp ntpc -- "NTP" --> ib1 ntpc -- "NTP" --> ib2 res -- "DNS" --> ib1 res -- "DNS" --> ib2 sensu -- "SNMPv3, user sensu" --> agent
| Service | Peer | Port | As built |
|---|---|---|---|
| Syslog | dc1-a-vcsys001, 10.11.17.113 | TCP 514 | legacy-TCP, no TLS, node ID in boot firmware messages |
dc1-a-vcmsx001.adm.example.net | TCP 25 | STARTTLS, server certificate not checked, login as scb_mail@ad.example.net | |
| NTP and DNS | Infoblox 10.11.16.145, 10.11.18.145 | 123, 53 | Both appliances of the Infoblox pair |
| SNMP | Sensu 10.11.16.129 asks | 161 | SNMPv3 only, SHA1 and AES, only this client |
The design's network communication table lists DNS, NTP, LDAPS and SMTP from 10.11.16.81, but it has no row for syslog to 10.11.17.113 and none for SNMP from Sensu. And the routing table of chapter 7 has specific routes for the admin VPN, Sensu, AD and Infoblox through 10.11.16.86, but none for the syslog server or the mail server. If the table is complete, syslog and mail left by the default route 10.11.17.70, which lies in none of the appliance's networks as documented (see Hardware, cabling and networks); nothing in my material shows the messages arriving or not.
Local services
The SSH server for the console and the web login for administrators listen only on 10.11.16.81, with the brute-force protection on and no restriction of clients. My notes' checklist for this page has a line "Restrict Clients (Allowed clients)" under SSH, web and SNMP, and a TODO names the admin VPN 10.11.20.0/25; only SNMP got a client list. For SSH and web, the only filter was the network: the communication table allows ORG_VPN to reach ports 22 and 443. SSH has password authentication on and no authorized keys. The second web login, "User only", has no address. The analysis had recommended a separate login address for users; as I understand it, it was not needed, because without gateway authentication no end user logs in to the web interface.
One consequence of putting the web interface on the management address shows in the naming. The design's access table gives the URL https://dc1-s-xblb001.adm.example.net/ with the address 10.11.16.81, while its DNS records give dc1-s-xblb001 the production address 10.11.16.65 and the older name dc1-s-xblm001 the management address. The operation how-to has every administrator add dc1-s-xblb001, dc1-s-xblm001 and scb.example.net to the workstation's hosts file, all pointing to 10.11.16.81; as I read it, that is how the mismatch was worked around. The lines are in Workstation hosts file.
The indexer runs on the box with 12 audit trails in parallel and no remote indexer. My notes copied the vendor's sizing figures while deciding, 50 to 100 MB per terminal and 150 to 300 MB per graphical trail, so 12 graphical trails take at most about 3.6 GB of the 32 GB (my arithmetic).
The whole page is a Config document: Local services (site 1).
SNMP: queries, not traps
The analysis recorded a disagreement: the security architect preferred Sensu to query the appliance over SNMP, the logging engineer preferred SNMP traps. The as-built configuration took the query. The SNMP agent speaks only version 3, with the single user sensu (SHA1 authentication, AES encryption, passwords not written down) and the single allowed client 10.11.16.129, the Sensu server dc1-a-vcsns001 that also watches the NetApp clusters (Logging, monitoring and AutoSupport). On the Alerting page the SNMP column is No for every alert, so the appliance never pushes anything. The site 2 cluster allows three client addresses, 10.12.16.129 to .131, which the material does not name, and has a trap setting of version 2c, community public, with no receiver. The only traps I can see being sent are local: the generated syslog-ng configuration calls snmptrap towards 127.0.0.1 with the name xcbInitSystemUnitFailed when a systemd unit fails, the name of the "A system service failed" alert that is mailed.
Syslog
The syslog decisions were taken in the analysis and repeated as reasons in the design: TCP, because the logging engineer wanted the server to receive every message; no TLS, because the central syslog server had none; the BSD protocol of RFC 3164 (legacy) rather than the IETF one, because not every device speaks the latter; IPv4, because the appliance did not support IPv6 for syslog; and the node ID in the host name of boot firmware messages, so the two nodes can be told apart. My TODO list for syslog had two more points: set "the right syslog prefixes", and separate operational logs (logins to the appliance) from security logs (logins to the target servers). Neither shows in the configuration.
The generated /etc/syslog-ng/syslog-ng.conf of the core firmware shows how the page is implemented. Messages of the core firmware get the cluster name dc1-s-xblb001.adm.example.net as host name; a TCP listener on port 1514 receives, by its name, the boot firmware messages of the slave with their own host name; everything is written to local files, one per weekday and per proxy (zorp-ssh, zorp-rdp and so on), so the box keeps a week of its own logs. One unfiltered log path sends it all to the central server:
destination remote {
syslog(
"10.11.17.113"
transport("tcp")
port(514)
template("$MSG\n")
template_escape(no)
);
};The page says legacy-TCP; the file uses syslog-ng's syslog() driver, which in the syslog-ng documentation I know is the IETF syslog driver. With the template reduced to $MSG, I cannot tell from the file what the server actually received, and nothing in my material comes from the receiving side. The analysis wanted the central server to forward security events to the SIEM (SPLUNK); that forwarding is not in my material either. An include file listens on localhost TCP 56456 and discards whatever arrives; the site 2 config.xml shows that port as the syslog target of the appliance's notifier, with RabbitMQ off.
The files are Config documents: syslog-ng.conf of the SCB (site 1) and message-queue-client.conf of the SCB.
Mail and alerts
Mail goes to the internal mail server by name, dc1-a-vcmsx001.adm.example.net, on port 25 with STARTTLS, authenticated as the AD account scb_mail, which is also the sender. Administrator mail, alerts and reports all go to the distribution group scb_mail_group, whose members are four named administrators "and …". The server certificate is not checked ("No certificate is required"), so the encryption protects against a passive listener but not against a server that pretends to be the mail server. The page names one node, not the service: in the Email Solution dc1-a-vcmsx001 is one node of the mail pair and 10.11.19.33 the pair's virtual address dc1-s-xcmsx001. The design's integration table gives 10.11.19.33 beside the node's name, and its chapter 4.1.4 gives 10.11.19.34. As I read it, the appliance's mail depended on that one node being up.
On the Alerting & Monitoring page, health monitoring mails when the disks pass 80 % or swap passes 70 %; the three load-average thresholds are 0, which the design does not explain. Of the 19 system alerts, 16 send mail: configuration change, backup and archive failures, HA node state change, timestamping error, lost time synchronisation, RAID, hardware, tainted firmware, too many login attempts, licence and failed services. Only the three login and logout events are off. All 16 traffic alerts are off, among them the real-time audit event. The analysis had said of RDP sessions "Do not block but send an alert"; as built, nothing raises such an alert. The site 2 export lists 23 system alerts by OID, 19 of them mailed, without names.
The pages are Config documents: Management and date and time (site 1) and Alerting and monitoring (site 1).
The rest of the Management page
The web session ends after ten idle minutes; the RPC API is off; debug logging is off; core files older than 14 days are deleted. The configuration backup uses the policy SYSTEM-BACKUP (daily at 00:00, NFS to the NetApp) and is encrypted to the public GPG key SCB-BACKUP; the policies are in Backup, archive and retention and the key in Certificates and keys. Clients are disconnected when the disks are 80 % used, and archiving is not started automatically then. Web gateway authentication is off, like every other part of gateway authentication. The SSL certificate section of the same page is a Config document of its own: SSL certificate page (site 1).
Date, time, version, licence
The appliance takes its time from both Infoblox appliances. My notes add that once the appliance works with the Windows domain, it and the domain controllers should follow the same NTP source, ideally the same server. The site 2 cluster takes its time from the same two site 1 addresses; in its support bundle of September 2018 the master synchronises, while both peers of the slave stand at .INIT., never reached. That and the "Slave is out of sync with the master" banner are told in High availability.
The System page shows core and boot firmware 5.0.3, built on 2017-11-11; my notes are headed "BALABIT notes - 4.0.7.a", which, as I read it, dates them before the upgrade to 5. The licence is host-based with a limit of 1000 protected hosts, which counts the jump servers and every server reached from them. The licence block in the site 2 config.xml says something else: Edition: T10 (Connection based), Limit: 20, Session-Based-License: yes, dated 2017/07/18. Sealed mode is disabled in site 1; the site 2 file has <seal_the_box>yes</seal_the_box> under <features>, and I cannot tell whether that is the switch or only the option. The page is a Config document: System (site 1).
Loose ends
- The network communication table has no rows for syslog and SNMP, and the routing table no route to the syslog or mail server other than the default route.
- The syslog page says
legacy-TCP, the generated configuration uses thesyslog()driver. - The NetApp Solution names the syslog server at
10.11.17.113dc1-s-xcsys001; this design calls itdc1-a-vcsys001. - The mail integration table gives the password row as "Password for scb@ad.example.net" while the account is
scb_mail; the design's chapter 4.1.4 and its integration table give two different addresses for the mail node's name, and the as-built page names a node rather than the virtual address.
What I would do differently
I would restrict the SSH and web logins to the admin VPN on the appliance itself, as my own checklist said, instead of relying on firewall rules I cannot show. I would give the mail settings the organisation's CA so that the mail server's certificate is checked. And I would ask for TLS on the syslog server rather than accept its absence: the analysis already listed TCP with TLS as an option, and the logs of a box whose purpose is accountability crossed the management LAN in clear text.