Proxy 02 - Interfaces and firewalld zones
Proxy Solution · Previous: Overview and design · Next: firewalld rich rules
A host with two interfaces, one towards the clients and one towards the internet, has to treat the two sides differently, and in firewalld the way to do that is a zone per interface. This part is the base configuration that every rich rule of the next part sits on: how each interface was bound to its zone, in the lab by two methods and in the datacenters by one, what the zones' targets were set to, which default services and ports were taken away, why masquerading was switched off, and how denied traffic was made visible. The whole command sets with their outputs are three Config documents: firewalld on the lab host, firewalld on dc1-a-vcprx001 and firewalld on dc2-a-vcprx001.
The starting point
On a fresh RHEL 7.5 every interface lands in the default zone, public. The lab's first firewall-cmd --get-active-zones, as root on proxy.lab.example.net, showed exactly that, and the note under it says why it is wrong: both interfaces are in the same zone.
$ firewall-cmd --get-active-zonesoutput 2 lines
public interfaces: ens32 ens33
The two production hosts did not start clean. On both, the same command listed internal bound to eth0 with sources: 192.168.0.0/16 and external bound to eth1, while the machines have ens3 and ens4 and no eth0 or eth1. The notes say it is wrong for that reason and nothing about where it came from; it was a configuration that was already on the machine. Datacenter 1 and datacenter 2 begin with the identical listing, but only the datacenter 2 note goes on to remove the stale bindings; in the datacenter 1 note they are gone from the next listing without a step that removed them. Whether the first listing of one note was pasted into the other, or something not recorded cleaned datacenter 1, I cannot tell.
Two ways to bind an interface to a zone
In the lab I tried the method that lives in the network configuration first: a ZONE= line in the ifcfg file of each interface, ZONE=external in ifcfg-ens32 and ZONE=internal in ifcfg-ens33, then systemctl restart network. The two files are Config documents: ifcfg-ens32, external and ifcfg-ens33, internal. The note carries two warnings: after these commands you can cut yourself off from the system, and the method works only when NetworkManager is enabled and controls the interface (NM_CONTROLLED=yes). As I understand it, it is NetworkManager that reads ZONE= and tells firewalld which zone the connection belongs to, so without it the line does nothing. The output after the restart is the first contradiction of the notes.
$ firewall-cmd --get-active-zonesoutput 4 lines
internal interfaces: ens32 external interfaces: ens33
The files say the opposite: ens32, the external interface with the default route, should be in external. The note under this listing says it is fine because the zones are now separated by interface, and the reversal was not noticed at the time. The notes do not say why the result came out reversed, and I will not guess; the next step replaced it anyway.
The second method talks to firewalld directly and does not need NetworkManager at all, which the note states (NM_CONTROLLED=no works too). As root on the lab host, then systemctl restart firewalld:
$ firewall-cmd --permanent --change-interface=ens32 --zone=external $ firewall-cmd --permanent --change-interface=ens33 --zone=internal $ systemctl restart firewalld
After that the active zones read internal: ens33 and external: ens32, as intended. This is the method the production hosts use, with ens4 to external and ens3 to internal. In datacenter 2 the first listing after it showed both the old and the new names side by side, interfaces: eth0 ens3 and eth1 ens4, so two --permanent --remove-interface commands took eth1 out of external and eth0 out of internal, followed by another restart. The sources: 192.168.0.0/16 line on internal stayed through the whole datacenter 2 note and is still in the final --list-all. As I understand firewalld, a source binding puts traffic from that range into the zone whichever interface it arrives on; the notes do not comment on it and never remove it.
Targets
A zone's target is what happens to a packet no rule in the zone accepts. All three zones that matter, internal, external and the default zone public, were read with --get-target (both answered default) and set to DROP with --permanent --set-target=DROP. As I understand the difference, default ends in a reject that sends an ICMP message back to the sender, and DROP discards silently; the notes just set it, on all three hosts in the same way. public keeps no interface, but as the default zone it is where any new interface would land, so it gets the same target.
This is the second contradiction. Right after the three --set-target commands each note lists the three zones with firewall-cmd --zone=… --list-all. The lab and datacenter 1 show target: DROP; datacenter 2 shows target: default in all three. Nothing in between differs, and there is no reload or restart between the command and the listing in any of the notes. My inference is that --list-all without --permanent shows the running configuration, which a --permanent change does not touch until a reload, so default is what the command should print, and the two notes that print DROP were tidied up after the fact. The firewalld 0.4.4.4 man page supports half of it: --set-target existed only for the permanent configuration then. Whether the other two listings were edited remains an inference, and the notes do not settle it.
What was removed from each zone
The --list-all outputs are also the inventory of what RHEL's zone definitions bring with them, and that inventory was pruned to what a proxy needs. Nothing was added at this stage; the additions are the rich rules of firewalld rich rules.
| Zone | Found | Removed in the lab and datacenter 1 | Removed in datacenter 2 |
|---|---|---|---|
external | ssh, masquerade: yes | masquerade | masquerade, ssh |
internal | ssh, mdns, samba-client, dhcpv6-client; in datacenter 2 also squid, ports 443/tcp, 3128/tcp, 80/tcp, 22/tcp | mdns, samba-client, dhcpv6-client | the same three services, the four ports and squid |
public | ssh, dhcpv6-client; in datacenter 2 also 3128/tcp | dhcpv6-client | nothing |
mdns, samba-client and dhcpv6-client are what the internal zone definition opens for a workstation; a proxy serving servers answers none of them. The extra services and ports on datacenter 2's internal and public belong to the configuration that was already on the machine, and removing them left the proxy port to be opened again, as a rich rule limited to the client ranges. Two things stayed that a reader will notice. The ssh service stayed on external in the lab and in datacenter 1, so SSH was accepted on the internet-facing interface there; only datacenter 2 removed it. And public in datacenter 2 kept dhcpv6-client, ssh and 3128/tcp, without effect while no interface is in that zone, but they would apply to the next interface that lands there. ssh also stayed on internal everywhere, so the SSH rich rules of the next part limit nothing that this service does not already allow; the notes do not remark on it.
Masquerading is on by default in the external zone, whose definition, as I understand the zone, is meant for a router hiding a private network behind its outside address. The step title in all three notes gives the reason for taking it off: "we are not firewall, but proxy". A proxy does not forward the clients' packets; it ends their connections and opens its own from its own external address, so there is nothing to translate, and leaving masquerading on would make the host a NAT router for anyone who could route through it; that is my reading of the step title, not something the notes spell out. As root on the lab host and on dc1-a-vcprx001; on dc2-a-vcprx001 the note has --remove-service=ssh between the two lines:
$ firewall-cmd --permanent --zone=external --remove-masquerade $ firewall-cmd --reload
flowchart LR inet["Internet"] clients["Servers, Ansible hosts, Sensu hosts"] subgraph ext["zone external, target DROP"] e4["ens4, in the lab ens32"] erm["removed masquerade, in datacenter 2 also ssh"] end subgraph int["zone internal, target DROP"] e3["ens3, in the lab ens33"] irm["removed mdns, samba-client, dhcpv6-client, in datacenter 2 also squid and four ports"] ikeep["kept ssh, then the rich rules"] end subgraph pub["zone public, default zone, target DROP, no interface"] prm["removed dhcpv6-client in the lab and datacenter 1"] pkeep["kept ssh, in datacenter 2 also dhcpv6-client and 3128/tcp"] end inet --- e4 e3 --- clients
Logging denied traffic
With a DROP target a packet that no rule accepts disappears, and when a client cannot reach the proxy the firewall is the first suspect. firewalld can log what it denies, and the setting is global, not per zone. On each host the setting was read (off), set to all, and read again.
$ firewall-cmd --get-log-denied $ firewall-cmd --set-log-denied=all $ firewall-cmd --get-log-denied
The step title explains all as unicast, broadcast and multicast. The command has no --permanent; as I understand it, --set-log-denied writes the value into firewalld.conf and reloads, so it is permanent anyway. Where the log lines went, and whether anybody read them, is not in the notes.
Lab, datacenter 1 and datacenter 2 side by side
| Item | Lab | Datacenter 1 | Datacenter 2 |
|---|---|---|---|
| External and internal interface | ens32, ens33 | ens4, ens3 | ens4, ens3 |
| Found at the start | both in public | eth0 in internal with sources: 192.168.0.0/16, eth1 in external | the same as datacenter 1 |
| Binding method | ZONE= in ifcfg, then --change-interface | --change-interface | --change-interface, then --remove-interface for eth0 and eth1 |
sources: on internal at the end | none | none | 192.168.0.0/16 |
Target of internal, external, public | DROP | DROP | DROP |
--list-all after --set-target shows | DROP | DROP | default |
Removed from internal | mdns, samba-client, dhcpv6-client | the same | the same, plus squid and 443/tcp, 80/tcp, 3128/tcp, 22/tcp |
Removed from external | masquerade | masquerade | masquerade, ssh |
Removed from public | dhcpv6-client | dhcpv6-client | nothing |
| Log denied | all | all | all |
| Rich rules added afterwards | 2 | 19 | 29, plus one direct rule for ICMPv6 |
The IPv6 side exists only in the datacenters: the lab's ifcfg files say IPV6INIT=no, and no lab rule has family=ipv6.
What is missing
The notes hold no ifcfg file of a production host, so how ens3 and ens4 got their addresses, whether NetworkManager ran there, and where the default route and the routes to the three private ranges pointed, is not recorded. There is no firewall-cmd --list-all --permanent anywhere, which would have settled the target: default question, and no test from a client of what the zones drop or log. The zone configuration files under /etc/firewalld/zones/ were never shown.
Today
Checked against firewalld 2.5.2 and RHEL 10. The first binding method is gone: RHEL 9 has no network-scripts package, RHEL 10 removed support for the ifcfg format, and with it ZONE= and systemctl restart network; the zone of an interface is now a property of its NetworkManager profile, nmcli connection modify <profile> connection.zone internal, and firewalld's own documentation says NetworkManager-managed interfaces need no binding in firewalld because NetworkManager binds them itself. --change-interface still exists and, as its man page says today, first asks NetworkManager to change the connection's zone and falls back to a binding in firewalld only when that fails. The zone definitions have not moved: the same nine zones, external still with ssh and masquerading for routers, internal still with the three workstation services, and since 1.0.0 all of them with intra-zone forwarding switched on. Underneath, the backend has been nftables since 0.6.0, and the default target was redefined in 1.0.0 to be exactly REJECT plus accepted ICMP, after the old behaviour had caused what the 1.0.0 announcement calls zone drifting, so setting DROP explicitly, as this build did, is still the way to get a silent zone. And the source-before-interface order that makes datacenter 2's leftover sources: 192.168.0.0/16 line matter is confirmed in the nftables backend's source, where source-based dispatch is deliberately kept ahead of interface-based dispatch.
What I would do differently
I would read the output of each check: the lab's reversed listing went through with a note saying all was well. One --list-all --permanent per zone after the last reload would have caught the reversed binding, the target: default question, the ssh service left on external in two of three places and the leftover sources: line in datacenter 2. And I would remove ssh from external everywhere, as datacenter 2 did: the SSH rules on internal are there so that the management side is the only way in.