Oracle RAC 04 - Operating system preparation
Oracle RAC Solution · Previous: DNS server: Knot · Next: Storage and ASM disks
Before the first Oracle installer starts, both nodes have to look the way the installation guide describes: the right packages, two accounts with the same numeric IDs on both machines, kernel parameters, resource limits, a time source and a directory tree. None of it is difficult, all of it is checked by the installer, and every difference between oradb01 and oradb02 comes back later as a failed prerequisite or a cluster that does not form. Everything in this Article was done as root on both nodes unless a section says otherwise.
The requirements come from Oracle's 11.2 installation guide for Linux x86-64, which my working notes cite chapter by chapter: package requirements, groups and users, kernel parameters and resource limits and required directories.
Packages
The guide lists 23 packages for SUSE Linux Enterprise Server 11, each with a minimum version. I installed every one of them by name with zypper, one command per package, and took whatever version SLES 11 SP3 offered.
| Package | Minimum in the guide |
|---|---|
binutils | 2.19 |
gcc, gcc-32bit | 4.3 |
gcc-c++ | 4.3 |
glibc, glibc-32bit | 2.9 |
glibc-devel, glibc-devel-32bit | 2.9 |
ksh | 93t |
libaio, libaio-32bit | 0.3.104 |
libaio-devel, libaio-devel-32bit | 0.3.104 |
libstdc++33, libstdc++33-32bit | 3.3.3 |
libstdc++43, libstdc++43-32bit | 4.3.3_20081022 |
libstdc++43-devel, libstdc++43-devel-32bit | 4.3.3_20081022 |
libgcc43 | 4.3.3_20081022 |
libstdc++-devel | 4.3 |
make | 3.81 |
sysstat | 8.1.5 |
The notes hold the rpm -q command that prints name, version, release and architecture of all 23, but not its output, so the versions that were really installed are not recorded. On top of the list came packages from other parts of the guide and from experience.
| Package | Version | Why | Source |
|---|---|---|---|
unixODBC, unixODBC-devel, unixODBC-32bit | 2.2.12 or later | the guide's section on Oracle ODBC drivers | zypper |
xorg-x11 | not recorded | the installers are graphical; the X server itself ran on my PC | zypper |
libcap1 | not recorded | "the installation needs it", say the notes, with no reference | zypper |
numactl | not recorded | listed under "other packages Oracle RAC + ASM requires" | zypper |
cvuqdisk | 1.0.9-1 | used by the Cluster Verification Utility to look at shared disks | Oracle installation medium |
oracleasm-support | 2.1.8-1 | the oracleasm command and init script | downloaded from Oracle |
oracleasm | not recorded | the kernel driver | zypper |
oracleasmlib | 2.0.4-1 | the library ASM loads to find ORCL: disks | downloaded from Oracle |
cvuqdisk has an order of its own: it must be installed after the group oinstall exists, so it comes after the next section. The three ASMLib packages are from Oracle's ASMLib page for SLES 11 and were installed in the order support tools, driver, library. All commands are a Config document: Package installation.
One thing in the notes is out of order. cvuqdisk is installed from /data/install/database/rpm, the unpacked database medium, in chapter 2 of the notes; the volume behind /data is created in chapter 4 (see Storage and ASM disks). The chapters are sorted by topic, not by time.
Users and groups
Grid Infrastructure and the database have separate owners, which is the "job role separation" layout of the guide: grid owns Clusterware and ASM, oracle owns the database software. The numeric IDs are set by hand so that they are identical on both nodes.
| Name | Kind | ID | Membership or role |
|---|---|---|---|
oinstall | group | 30001 | primary group of both accounts, owner of the Oracle inventory |
asmadmin | group | 30002 | ASM administrators; group of the ASM disks |
asmdba | group | 30003 | database access to ASM |
dba | group | 30004 | database administrators |
grid | user | 30101 | primary oinstall, secondary asmadmin, asmdba |
oracle | user | 30102 | primary oinstall, secondary dba, asmdba |
$ groupadd -g 30001 oinstall $ groupadd -g 30002 asmadmin $ groupadd -g 30003 asmdba $ groupadd -g 30004 dba $ useradd -g oinstall -G asmadmin,asmdba -u 30101 grid $ useradd -g oinstall -G dba,asmdba -u 30102 oracle $ passwd oracle $ passwd grid
<ORACLE_OS_PASSWORD> and <GRID_OS_PASSWORD>. The Grid Infrastructure and database installers ask for them again when they set up SSH between the nodes.useradd is run without -m in the notes. Both accounts nevertheless have a home directory with a .bash_profile a few steps later, so the homes were created, by a default or by hand; the notes do not say which.
Kernel parameters
The guide gives minimum values for eleven kernel parameters. Ten of them went into /etc/sysctl.conf unchanged; shared memory needed more thought.
| Parameter | Guide minimum | As set |
|---|---|---|
fs.aio-max-nr | 1048576 | 1048576 |
fs.file-max | 6815744 | 6815744 |
kernel.shmall | 2097152 | 2097152 |
kernel.shmmax | 536870912 | 4124030976 |
kernel.shmmni | 4096 | 4096 |
kernel.sem | 250 32000 100 128 | 250 32000 100 128 |
net.ipv4.ip_local_port_range | 9000 65500 | 9000 65500 |
net.core.rmem_default | 262144 | 262144 |
net.core.rmem_max | 4194304 | 4194304 |
net.core.wmem_default | 262144 | 262144 |
net.core.wmem_max | 1048576 | 1048576 |
kernel.shmmax is the largest single shared memory segment in bytes, kernel.shmall the total of shared memory in pages. The guide's 512 MB for shmmax is a floor, not a recommendation: as I understand it, an SGA larger than shmmax is cut into several segments. After reading a page on the optimal SHMMAX I let a small script, shmsetup.sh, calculate both values from the memory of the machine, and the Grid Infrastructure installer then disagreed with both results.
| Parameter | shmsetup.sh | The installer wanted |
|---|---|---|
kernel.shmmax | 4124028928 | 4124030976 |
kernel.shmall | 1006843 | 2097152 |
The two shmmax values differ by 2048 bytes, half a memory page: the script's shmall of 1006843 pages times 4096 is exactly its shmmax, so it rounds to whole pages and the installer does not. For shmall the installer simply insisted on the guide's minimum. I took the installer's pair, which is what the table above and the file show. The notes do not state how much memory the virtual machines had; shmmax is the only trace of it.
Below the eleven parameters the block has two more groups. vm.hugetlb_shm_group names the one group that may create shared memory segments from huge pages, a SUSE requirement according to the comment. It is set twice, to 30001 (oinstall) with the comment "installation" and to 30004 (dba) with the comment "operation". sysctl -p reads the file from top to bottom, so the second line wins and the first has no effect; the notes do not say that the first was ever commented out. The last two lines, net.ipv4.conf.eth3.rp_filter = 0 and net.ipv4.icmp_echo_ignore_broadcasts = 0, stand under a comment saying that eth3 with its alias eth3:INT carries the interconnect. That was true only after the interconnect had been moved off eth2, so these lines were added during the trouble described in Grid patches and the multicast problem, not during preparation. The notes give no reason for either line beyond that comment.
The block was activated with sysctl -p. The whole of it is a Config document: sysctl.conf, Oracle block.
Limits, PAM and the profile
The guide's nproc and nofile limits for the software owner are in /etc/security/limits.conf only as comments; the lines in effect below them are values I raised, for both accounts.
| Limit | Commented block, soft / hard | As set for oracle and grid, soft and hard |
|---|---|---|
nproc | 2047 / 16384 | 131072 |
nofile | 1024 / 65536 | 131072 |
as | unlimited | unlimited |
core | not in the commented block | unlimited |
memlock | not in the commented block | 3500000 |
The file is a Config document: limits.conf, Oracle block. Limits from that file are applied by pam_limits, so one line was appended to /etc/pam.d/login, a step that Oracle's older guides print and the 11.2 guides no longer do:
# PAM configuration for running Oracle Database
session required pam_limits.soOlder Oracle instructions also carry a shell snippet that raises the soft limits of the Oracle accounts at login. Mine is /etc/profile.local, mode 644, with one block for oracle and one for grid: profile.local. It carries the remark "I had to comment this out" in front of ulimit -p 131072. Oracle's snippet, which the user profiles of Database creation and user environments repeat with its original numbers, uses ulimit -p in a branch for the Korn shell. My file tests for /bin/bash in that branch, and in bash -p is the pipe buffer size, which cannot be set. That is today's explanation; the notes say only that the line had to be commented out. What remains raises only the number of open files for bash; the process limit comes from limits.conf alone.
The last piece of the user environment at this stage is the file creation mask. As grid and as oracle, each ~/.bash_profile got one line:
umask 022Names and time
Every name the cluster uses is resolved by the DNS server dns1. /etc/resolv.conf of the nodes points at it and nowhere else, and the /etc/hosts file in the notes is marked there as not used "since every record is in DNS". Both are described in Network and DNS plan, the hosts file as the Config document /etc/hosts; I do not repeat them here.
The same server is the time source. NTP was configured with YaST, which leaves one line in /etc/ntp.conf:
server 10.30.40.13 iburst
Oracle Clusterware wants more than a synchronised clock: the clock must never be set backwards while the cluster runs, so ntpd has to slew, not step. That is the -x switch, added to the daemon options in /etc/sysconfig/ntp:
NTPD_OPTIONS="-x -g -u ntp:ntp"
The notes show only this line of the file and no restart of the daemon.
Directories
The software lives under /data/u01, on the local volume built in the next Article.
| Path | Purpose | Owner after these commands |
|---|---|---|
/data/u01/app/grid11204 | Grid Infrastructure home | grid:oinstall |
/data/u01/app/grid | Oracle base of grid | grid:oinstall |
/data/u01/app/oracle | Oracle base of oracle | oracle:oinstall |
$ mkdir -p /data/u01/app/grid11204 $ mkdir -p /data/u01/app/grid $ chown -R grid:oinstall /data/u01 $ mkdir -p /data/u01/app/oracle $ chown oracle:oinstall /data/u01/app/oracle $ chmod -R 775 /data/u01
The Grid home is beside the Oracle base of grid, not below it. The notes give the paths without comment; my understanding, which is not something they say, is that root.sh hands parts of the Grid home to root and that Oracle wants it outside any Oracle base. The inventory /data/u01/app/oraInventory and the database home /data/u01/app/oracle/db11204 are not created here; the installers create them. The header of the notes gives the permissions of both software locations as 755, while the command sets 775. Both are in the notes; the command is what was run.
SSH keys and user equivalence
The installers copy the software from the first node to the second over SSH and need to log in there as grid and as oracle without a password. They can set this up themselves, and I let them: the screen "SSH Connectivity" has a "Setup" button and an option to reuse keys that already exist in the home directory (see Grid Infrastructure installation). At this stage only the keys were created, on both nodes, first as oracle and then the same as grid:
$ mkdir ~/.ssh $ chmod 700 ~/.ssh $ ssh-keygen -t rsa $ ssh-keygen -t dsa
The manual procedure in the 11.2 Grid Infrastructure guide uses DSA keys, the installer's automatic setup uses RSA keys. The notes create both; with the automatic setup only the RSA key was needed. The notes do not show a passphrase, a key size or any authorized_keys file; the last is what the installer writes.
The chapter ends with the equivalence tests. They can only succeed after the installer's "Setup", so in time they belong to the installation, not to the preparation. On oradb01, as oracle and again as grid:
$ ssh ORADB02 date $ ssh ORADB02-int date $ ssh ORADB02.example.net date
Then the same three from oradb02 towards ORADB01, ORADB01-int and ORADB01.example.net. Each must print the date without asking anything, neither for a password nor for confirmation of a host key, because, as I understand it, the installer runs its remote commands without a terminal. Three names are tested, I take it because each is a separate entry in known_hosts (the notes give no reason): the short public name, the interconnect name and the fully qualified name. The notes write the host names in capitals and record no output.
What the notes do not hold
- The output of the package checks, and with it the installed versions.
- How SLES itself was installed, and how the home directories of
gridandoraclecame to be. - The amount of memory and the number of CPUs of the virtual machines.
- A restart or a check of
ntpdafter the change, and the output of the SSH tests.
What I would do differently
- The
stacklimit is missing. The 11.2 guide already asked for a soft and hardstacklimit of at least 10240 KB next tonprocandnofile. My block hasasandcore, which Oracle's table does not contain, and nostack. - One value for
vm.hugetlb_shm_group. Oracle's text says to enter the GID ofoinstalland that only one group can be defined. The effective value here isdba, of whichgridis not a member. It did no harm only because no huge pages were configured. - No limits in shell profiles. With soft and hard limits equal in
limits.confthere is nothing left for a profile to raise, and the 11.2.0.4-era guides contain noulimit -psnippet any more. The same goes for the PAM line. - RSA keys only. The DSA keys were never used, and OpenSSH has since removed DSA altogether (off by default since 7.0, removed in 10.0).
- SLES 11 SP4, not SP3. General support for SP3 had ended on 31 January 2016, months before this build.
On current software most of this Article still applies in spirit and little of it literally. The kernel parameter table of Oracle AI Database 26ai keeps my values, except that kernel.shmall is now "greater than or equal to shmmax in pages" and kernel.panic_on_oops = 1 and kernel.panic = 10 were added. The -x requirement for ntpd is gone from the 19c and 26ai guides, which name chrony; SLES 15 moved ntp to its Legacy Module and SLES 16 no longer ships it. chronyd slews by default, and its own -x option means something else (it stops chronyd from controlling the clock), so the switch must not be carried over. The per-file details are in the "Checked against" sections of the Config documents.