LINUXOR.SK ... open source notes ...

Oracle RAC 04 - Operating system preparation

category: solutionz · date: 2016-12-31 · updated: 2026-10-02 · author: LALA

Oracle RAC Solution · Previous: DNS server: Knot · Next: Storage and ASM disks

Before the first Oracle installer starts, both nodes have to look the way the installation guide describes: the right packages, two accounts with the same numeric IDs on both machines, kernel parameters, resource limits, a time source and a directory tree. None of it is difficult, all of it is checked by the installer, and every difference between oradb01 and oradb02 comes back later as a failed prerequisite or a cluster that does not form. Everything in this Article was done as root on both nodes unless a section says otherwise.

The requirements come from Oracle's 11.2 installation guide for Linux x86-64, which my working notes cite chapter by chapter: package requirements, groups and users, kernel parameters and resource limits and required directories.

Packages

The guide lists 23 packages for SUSE Linux Enterprise Server 11, each with a minimum version. I installed every one of them by name with zypper, one command per package, and took whatever version SLES 11 SP3 offered.

PackageMinimum in the guide
binutils2.19
gcc, gcc-32bit4.3
gcc-c++4.3
glibc, glibc-32bit2.9
glibc-devel, glibc-devel-32bit2.9
ksh93t
libaio, libaio-32bit0.3.104
libaio-devel, libaio-devel-32bit0.3.104
libstdc++33, libstdc++33-32bit3.3.3
libstdc++43, libstdc++43-32bit4.3.3_20081022
libstdc++43-devel, libstdc++43-devel-32bit4.3.3_20081022
libgcc434.3.3_20081022
libstdc++-devel4.3
make3.81
sysstat8.1.5

The notes hold the rpm -q command that prints name, version, release and architecture of all 23, but not its output, so the versions that were really installed are not recorded. On top of the list came packages from other parts of the guide and from experience.

PackageVersionWhySource
unixODBC, unixODBC-devel, unixODBC-32bit2.2.12 or laterthe guide's section on Oracle ODBC driverszypper
xorg-x11not recordedthe installers are graphical; the X server itself ran on my PCzypper
libcap1not recorded"the installation needs it", say the notes, with no referencezypper
numactlnot recordedlisted under "other packages Oracle RAC + ASM requires"zypper
cvuqdisk1.0.9-1used by the Cluster Verification Utility to look at shared disksOracle installation medium
oracleasm-support2.1.8-1the oracleasm command and init scriptdownloaded from Oracle
oracleasmnot recordedthe kernel driverzypper
oracleasmlib2.0.4-1the library ASM loads to find ORCL: disksdownloaded from Oracle

cvuqdisk has an order of its own: it must be installed after the group oinstall exists, so it comes after the next section. The three ASMLib packages are from Oracle's ASMLib page for SLES 11 and were installed in the order support tools, driver, library. All commands are a Config document: Package installation.

One thing in the notes is out of order. cvuqdisk is installed from /data/install/database/rpm, the unpacked database medium, in chapter 2 of the notes; the volume behind /data is created in chapter 4 (see Storage and ASM disks). The chapters are sorted by topic, not by time.

Users and groups

Grid Infrastructure and the database have separate owners, which is the "job role separation" layout of the guide: grid owns Clusterware and ASM, oracle owns the database software. The numeric IDs are set by hand so that they are identical on both nodes.

NameKindIDMembership or role
oinstallgroup30001primary group of both accounts, owner of the Oracle inventory
asmadmingroup30002ASM administrators; group of the ASM disks
asmdbagroup30003database access to ASM
dbagroup30004database administrators
griduser30101primary oinstall, secondary asmadmin, asmdba
oracleuser30102primary oinstall, secondary dba, asmdba
bash
$ groupadd -g 30001 oinstall
$ groupadd -g 30002 asmadmin
$ groupadd -g 30003 asmdba
$ groupadd -g 30004 dba
$ useradd -g oinstall -G asmadmin,asmdba -u 30101 grid
$ useradd -g oinstall -G dba,asmdba -u 30102 oracle
$ passwd oracle
$ passwd grid
noteThe notes record the two passwords typed here as the placeholders <ORACLE_OS_PASSWORD> and <GRID_OS_PASSWORD>. The Grid Infrastructure and database installers ask for them again when they set up SSH between the nodes.

useradd is run without -m in the notes. Both accounts nevertheless have a home directory with a .bash_profile a few steps later, so the homes were created, by a default or by hand; the notes do not say which.

Kernel parameters

The guide gives minimum values for eleven kernel parameters. Ten of them went into /etc/sysctl.conf unchanged; shared memory needed more thought.

ParameterGuide minimumAs set
fs.aio-max-nr10485761048576
fs.file-max68157446815744
kernel.shmall20971522097152
kernel.shmmax5368709124124030976
kernel.shmmni40964096
kernel.sem250 32000 100 128250 32000 100 128
net.ipv4.ip_local_port_range9000 655009000 65500
net.core.rmem_default262144262144
net.core.rmem_max41943044194304
net.core.wmem_default262144262144
net.core.wmem_max10485761048576

kernel.shmmax is the largest single shared memory segment in bytes, kernel.shmall the total of shared memory in pages. The guide's 512 MB for shmmax is a floor, not a recommendation: as I understand it, an SGA larger than shmmax is cut into several segments. After reading a page on the optimal SHMMAX I let a small script, shmsetup.sh, calculate both values from the memory of the machine, and the Grid Infrastructure installer then disagreed with both results.

Parametershmsetup.shThe installer wanted
kernel.shmmax41240289284124030976
kernel.shmall10068432097152

The two shmmax values differ by 2048 bytes, half a memory page: the script's shmall of 1006843 pages times 4096 is exactly its shmmax, so it rounds to whole pages and the installer does not. For shmall the installer simply insisted on the guide's minimum. I took the installer's pair, which is what the table above and the file show. The notes do not state how much memory the virtual machines had; shmmax is the only trace of it.

Below the eleven parameters the block has two more groups. vm.hugetlb_shm_group names the one group that may create shared memory segments from huge pages, a SUSE requirement according to the comment. It is set twice, to 30001 (oinstall) with the comment "installation" and to 30004 (dba) with the comment "operation". sysctl -p reads the file from top to bottom, so the second line wins and the first has no effect; the notes do not say that the first was ever commented out. The last two lines, net.ipv4.conf.eth3.rp_filter = 0 and net.ipv4.icmp_echo_ignore_broadcasts = 0, stand under a comment saying that eth3 with its alias eth3:INT carries the interconnect. That was true only after the interconnect had been moved off eth2, so these lines were added during the trouble described in Grid patches and the multicast problem, not during preparation. The notes give no reason for either line beyond that comment.

The block was activated with sysctl -p. The whole of it is a Config document: sysctl.conf, Oracle block.

Limits, PAM and the profile

The guide's nproc and nofile limits for the software owner are in /etc/security/limits.conf only as comments; the lines in effect below them are values I raised, for both accounts.

LimitCommented block, soft / hardAs set for oracle and grid, soft and hard
nproc2047 / 16384131072
nofile1024 / 65536131072
asunlimitedunlimited
corenot in the commented blockunlimited
memlocknot in the commented block3500000

The file is a Config document: limits.conf, Oracle block. Limits from that file are applied by pam_limits, so one line was appended to /etc/pam.d/login, a step that Oracle's older guides print and the 11.2 guides no longer do:

ini
# PAM configuration for running Oracle Database
session     required       pam_limits.so

Older Oracle instructions also carry a shell snippet that raises the soft limits of the Oracle accounts at login. Mine is /etc/profile.local, mode 644, with one block for oracle and one for grid: profile.local. It carries the remark "I had to comment this out" in front of ulimit -p 131072. Oracle's snippet, which the user profiles of Database creation and user environments repeat with its original numbers, uses ulimit -p in a branch for the Korn shell. My file tests for /bin/bash in that branch, and in bash -p is the pipe buffer size, which cannot be set. That is today's explanation; the notes say only that the line had to be commented out. What remains raises only the number of open files for bash; the process limit comes from limits.conf alone.

The last piece of the user environment at this stage is the file creation mask. As grid and as oracle, each ~/.bash_profile got one line:

bash
umask 022

Names and time

Every name the cluster uses is resolved by the DNS server dns1. /etc/resolv.conf of the nodes points at it and nowhere else, and the /etc/hosts file in the notes is marked there as not used "since every record is in DNS". Both are described in Network and DNS plan, the hosts file as the Config document /etc/hosts; I do not repeat them here.

The same server is the time source. NTP was configured with YaST, which leaves one line in /etc/ntp.conf:

ini
server 10.30.40.13  iburst

Oracle Clusterware wants more than a synchronised clock: the clock must never be set backwards while the cluster runs, so ntpd has to slew, not step. That is the -x switch, added to the daemon options in /etc/sysconfig/ntp:

ini
NTPD_OPTIONS="-x -g -u ntp:ntp"

The notes show only this line of the file and no restart of the daemon.

Directories

The software lives under /data/u01, on the local volume built in the next Article.

PathPurposeOwner after these commands
/data/u01/app/grid11204Grid Infrastructure homegrid:oinstall
/data/u01/app/gridOracle base of gridgrid:oinstall
/data/u01/app/oracleOracle base of oracleoracle:oinstall
bash
$ mkdir -p /data/u01/app/grid11204
$ mkdir -p /data/u01/app/grid
$ chown -R grid:oinstall /data/u01
$ mkdir -p /data/u01/app/oracle
$ chown oracle:oinstall /data/u01/app/oracle
$ chmod -R 775 /data/u01

The Grid home is beside the Oracle base of grid, not below it. The notes give the paths without comment; my understanding, which is not something they say, is that root.sh hands parts of the Grid home to root and that Oracle wants it outside any Oracle base. The inventory /data/u01/app/oraInventory and the database home /data/u01/app/oracle/db11204 are not created here; the installers create them. The header of the notes gives the permissions of both software locations as 755, while the command sets 775. Both are in the notes; the command is what was run.

SSH keys and user equivalence

The installers copy the software from the first node to the second over SSH and need to log in there as grid and as oracle without a password. They can set this up themselves, and I let them: the screen "SSH Connectivity" has a "Setup" button and an option to reuse keys that already exist in the home directory (see Grid Infrastructure installation). At this stage only the keys were created, on both nodes, first as oracle and then the same as grid:

bash
$ mkdir ~/.ssh
$ chmod 700 ~/.ssh
$ ssh-keygen -t rsa
$ ssh-keygen -t dsa

The manual procedure in the 11.2 Grid Infrastructure guide uses DSA keys, the installer's automatic setup uses RSA keys. The notes create both; with the automatic setup only the RSA key was needed. The notes do not show a passphrase, a key size or any authorized_keys file; the last is what the installer writes.

The chapter ends with the equivalence tests. They can only succeed after the installer's "Setup", so in time they belong to the installation, not to the preparation. On oradb01, as oracle and again as grid:

bash
$ ssh ORADB02 date
$ ssh ORADB02-int date
$ ssh ORADB02.example.net date

Then the same three from oradb02 towards ORADB01, ORADB01-int and ORADB01.example.net. Each must print the date without asking anything, neither for a password nor for confirmation of a host key, because, as I understand it, the installer runs its remote commands without a terminal. Three names are tested, I take it because each is a separate entry in known_hosts (the notes give no reason): the short public name, the interconnect name and the fully qualified name. The notes write the host names in capitals and record no output.

What the notes do not hold

What I would do differently

On current software most of this Article still applies in spirit and little of it literally. The kernel parameter table of Oracle AI Database 26ai keeps my values, except that kernel.shmall is now "greater than or equal to shmmax in pages" and kernel.panic_on_oops = 1 and kernel.panic = 10 were added. The -x requirement for ntpd is gone from the 19c and 26ai guides, which name chrony; SLES 15 moved ntp to its Legacy Module and SLES 16 no longer ships it. chronyd slews by default, and its own -x option means something else (it stops chronyd from controlling the clock), so the switch must not be carried over. The per-file details are in the "Checked against" sections of the Config documents.

← solutionz