Oracle RAC 03 - DNS server: Knot
Oracle RAC Solution · Previous: Network and DNS plan · Next: Operating system preparation
The cluster keeps none of its names in local files. The public names, the VIPs, the SCAN, the interconnect, backup and management names are all resolved by one small server, dns1, running the authoritative DNS server Knot on Debian 8.5. This Article is how that server was set up: one package, two blocks in its configuration, five zone files, a restart and five zone transfers as tests. The tests are the interesting part, because their output shows mistakes in my zone files that I did not notice at the time.
The server
dns1 has the address 10.30.40.13 in the management network. The nodes name it as their only name server and also as their NTP server; of the time service the notes show only the client side on the nodes, so this Article is about DNS alone.
Knot is an authoritative server. It answers for the zones it holds and does not resolve anything else on behalf of a client. For this installation that was all that was needed: one forward zone and a reverse zone for each network in which the cluster has names.
| Zone | Zone file | Network |
|---|---|---|
example.net | /etc/knot/example.net.zone | all names |
10.30.10.in-addr.arpa | /etc/knot/10.30.10.in-addr.arpa | public, with VIPs and SCAN |
20.30.10.in-addr.arpa | /etc/knot/20.30.10.in-addr.arpa | interconnect |
30.30.10.in-addr.arpa | /etc/knot/30.30.10.in-addr.arpa | backup |
40.30.10.in-addr.arpa | /etc/knot/40.30.10.in-addr.arpa | management |
flowchart LR subgraph d["dns1, 10.30.40.13, Knot DNS"] conf["knot.conf, remotes and zones"] fz["example.net"] r10["10.30.10.in-addr.arpa, public"] r20["20.30.10.in-addr.arpa, interconnect"] r30["30.30.10.in-addr.arpa, backup"] r40["40.30.10.in-addr.arpa, management"] conf --> fz conf --> r10 conf --> r20 conf --> r30 conf --> r40 end n1["oradb01, resolv.conf"] -->|"queries, port 53"| d n2["oradb02, resolv.conf"] -->|"queries, port 53"| d dig["dig on dns1 itself"] -->|"axfr, allowed for this-server only"| d bsrv["mng-backupsrv01, finds the virtual RAC node"] -.->|"DNS or its own hosts file"| d
The installation is one command, run as root on dns1.
$ apt-get install knot
The notes do not record which Knot version the package brought. Debian 8 carried Knot DNS 1.6.0, so that is the most likely one, and the configuration below is in the 1.x syntax with blocks in braces, which no later Knot reads.
The configuration: one remote, five zones
Two things were added to /etc/knot/knot.conf. The first is a remotes block that gives a name to an address, and the only remote declared is the server itself:
remotes {
this-server {
address 10.30.40.13@53;
}
}The second is one entry per zone in the zones block, all of the same shape:
example.net {
file "/etc/knot/example.net.zone";
xfr-out this-server;
}file names the zone file and xfr-out names the remote that may transfer the zone. Allowing a zone transfer to the server's own address sounds pointless, since there is no secondary server anywhere in the notes. It is there for testing: with it, a dig ... axfr started on dns1 and sent to 10.30.40.13 returns the complete zone as the server has loaded it, and nobody else can do the same. The added blocks as a whole are a Config document: knot.conf, remotes and zones.
The forward zone
The zone example.net has a $ORIGIN, a TTL of one hour, the SOA and NS records, and then one address record per name, grouped as in the tables of Network and DNS plan. The head of the file and the SCAN records:
$ORIGIN example.net. $TTL 3600 @ SOA dns1.example.net. hostmaster.example.net. ( 2016102609 ; serial 6h ; refresh 1h ; retry 1w ; expire 1d ) ; minimum NS dns1 dns1 A 10.30.40.13 ; Oracle RAC - scan IP oradb-scan A 10.30.10.31 oradb-scan A 10.30.10.32 oradb-scan A 10.30.10.33
The SCAN is nothing special to the DNS server: one name with three address records. Oracle's guide asks for them to be "configured as round robin addresses" in DNS, and I took that for granted. Knot 1.6 has no record rotation at all, so dns1 returned the three addresses in the same order every time. Nothing in the notes shows that this caused a problem, and nothing shows that anybody looked. The whole file is a Config document: Zone example.net.
Near its end the file has two records that do not belong to the cluster itself:
rac-clsopdb-1234567890 A 10.30.30.11 rac-clsopdb-1234567890 A 10.30.30.12
Backup Exec addresses a RAC database as a virtual node whose name is built from the database name and the DBID. The backup server has to resolve that name to the nodes, and here it resolves to the backup addresses of both. The notes offer two ways to get there, the hosts file of the backup server or records on the DNS server, and show the second. The DBID in the name is a made-up value in this write-up. If the serials are dates, as their shape suggests, this zone was last edited on 26 October, while the reverse zones carry serials of 6 and 7 October; the backup records are the likely reason. The story is in Backup Exec server and RAC backup.
The reverse zones
The four reverse zone files are written in another style than the forward zone: no $ORIGIN, every owner name spelled out in full with the final dot, a TTL of one day, and the SOA timers commented in a different wording. The head of the management zone file and its first record:
$TTL 86400 40.30.10.in-addr.arpa. IN SOA dns1.example.net hostmaster.example.net. ( 20161006 ; serial 4h ; slave refresh 2h ; slave retry interval 2w ; slave data expiration 1h ) ; maximum caching time when lookups fail ; 40.30.10.in-addr.arpa. IN NS dns1.example.net. 11.40.30.10.in-addr.arpa. IN PTR oradb01-mng.example.net.
The other three look the same with their own records and without the NS line. Each is a Config document: public, interconnect, backup, management.
Restart and tests
After the files were in place, Knot was restarted, as root on dns1.
$ /etc/init.d/knot restartoutput 1 line
[ ok ] Restarting knot (via systemctl): knot.service.
The test of each zone was a zone transfer from the server to itself. The notes remark at each of them that this is possible only because the transfer was explicitly allowed.
$ dig @10.30.40.13 example.net axfr
output 15 lines
; <<>> DiG 9.9.5-9+deb8u6-Debian <<>> @10.30.40.13 example.net axfr ... example.net. 3600 IN SOA dns1.example.net. hostmaster.example.net. 2016102609 21600 3600 604800 86400 example.net. 3600 IN NS dns1.example.net. dns1.example.net. 3600 IN A 10.30.40.13 ... oradb-scan.example.net. 3600 IN A 10.30.10.31 oradb-scan.example.net. 3600 IN A 10.30.10.32 oradb-scan.example.net. 3600 IN A 10.30.10.33 oradb01.example.net. 3600 IN A 10.30.10.11 ... rac-clsopdb-1234567890.example.net. 3600 IN A 10.30.30.11 rac-clsopdb-1234567890.example.net. 3600 IN A 10.30.30.12 ... ;; XFR size: 22 records (messages 1, bytes 704)
The output is shortened here, and every cut is marked with a line of three dots. Twenty-two records came back, the SOA counted twice as in every transfer, and the forward zone is as it was written. The same command for the public reverse zone, shortened in the same way:
$ dig @10.30.40.13 10.30.10.in-addr.arpa axfr
output 11 lines
... 10.30.10.in-addr.arpa. 86400 IN SOA dns1.example.net.10.30.10.in-addr.arpa. hostmaster.example.net. 20161007 14400 7200 1209600 3600 11.10.30.10.in-addr.arpa. 86400 IN PTR oradb01.example.net. 12.10.30.10.in-addr.arpa. 86400 IN PTR oradb02.example.net. 21.10.30.10.in-addr.arpa. 86400 IN PTR oradb01-vip.example.net. 22.10.30.10.in-addr.arpa. 86400 IN PTR oradb02-vip.example.net. 31.10.30.10.in-addr.arpa. 86400 IN PTR oradb-scan.example.net. 32.10.30.10.in-addr.arpa. 86400 IN PTR oradb-scan.example.net. 33.10.30.10.in-addr.arpa. 86400 IN PTR oradb-scan.example.net. ... ;; XFR size: 9 records (messages 1, bytes 491)
The three remaining transfers were run the same way.
| Zone | Records in the transfer | NS record in the output |
|---|---|---|
example.net | 22 | yes |
10.30.10.in-addr.arpa | 9 | no |
20.30.10.in-addr.arpa | 4 | no |
30.30.10.in-addr.arpa | 4 | no |
40.30.10.in-addr.arpa | 5 | yes |
These are the only tests in the notes. There is no ordinary query for a name or an address in them, neither from dns1 nor from a node.
The mistakes the outputs show
Look at the first record of the reverse transfer above. The SOA primary of the zone is dns1.example.net.10.30.10.in-addr.arpa.
In the zone file the SOA line reads SOA dns1.example.net hostmaster.example.net.: the mailbox name ends with a dot, the server name does not. A name without the final dot is relative, and the server completed it with the origin, which for a file without $ORIGIN is the zone name from knot.conf. All four reverse zones have the same line and the same result, each with its own zone name appended. The forward zone has the dot and is right. I had the proof on the screen four times and wrote it into the notes as the expected output.
It did no harm that I know of. The field names the primary server of the zone for whoever needs to find it, a client sending dynamic updates for example, and this installation had no such client and no secondary server. The PTR records, which are what the reverse zones are for, are written in full and came out as intended.
The second mistake is the missing NS records. Only the management reverse zone file has one; the public, interconnect and backup zones consist of a SOA and PTR records and nothing else. The NS records at the top of a zone are, with the SOA, the mandatory records of every zone (RFC 2181). The server loaded and served all three all the same, which the transfers prove: in Knot 1.6 the check for a missing NS record was one of the semantic checks that are off by default and do not stop a zone from loading. Nothing asked for the NS records of those zones.
The third group is records that are simply not there.
| Missing | Where one would expect it |
|---|---|
oradb02-iscsi | forward zone, next to oradb01-iscsi (10.30.50.11) |
A reverse zone for the iSCSI network 10.30.50.x | knot.conf and a fifth reverse zone file |
PTR for dns1 (10.30.40.13) | management reverse zone |
PTR for mng-backupsrv01 (10.30.40.14) | management reverse zone |
PTR for mng-backupsrv01-bck (10.30.30.14) | backup reverse zone |
The reverse zones hold only the cluster's own names (the two nodes, and in the public zone also the two VIP and three SCAN records), which suggests they were written for the cluster installation only and not touched when the backup server came into the forward zone. The single oradb01-iscsi record also stands under the comment "OS/DB management" in the zone file, a sign that it was added in passing.
One more small thing: the comment above the Backup Exec records calls the database CLSDB, and the record below it is rac-clsopdb-…. That is the naming contradiction described in Overview and design, visible inside a single zone file.
What I would do differently
Everything here follows from the outputs above.
- Write the reverse zones like the forward zone, with
$ORIGINand relative names, or check every name on the right-hand side for its final dot. One reading of the SOA line in thedigoutput would have caught it. - Give every zone its
NSrecord, and run a zone checker before the restart. Knot 1.6 hadknotc checkzone; current Knot haskzonecheck, which reports a missing NS record at the top of a zone as an error. Knot 3.6.0 given my reverse zone files loads them, logs a warning for the missing NS records and reproduces the wrong SOA primary exactly. - Reload instead of restart.
knotc reloadpicks up changed zone files without stopping the server. - Decide about the order of the SCAN addresses instead of assuming it. Record rotation came to Knot in 2.7.3 as
answer-rotation, off by default. - Test with the queries the cluster will send, a lookup of
oradb-scanand of each address, from a node, and not only with a zone transfer on the server. - Keep forward and reverse in step: every address record of a host with its PTR, including the DNS server and the backup server, and both nodes on the iSCSI network or neither.
- Not rely on a single name server for a cluster that has every name in DNS only.
The configuration itself could not be carried over today: the 1.x format was replaced by YAML in Knot 2.0, and the permission to transfer a zone is now an access rule and no longer a remote. The comparison is in the Config documents, under "Checked against Knot DNS 3.6.0".