LINUXOR.SK ... open source notes ...

Oracle RAC 06 - Grid Infrastructure installation

category: solutionz · date: 2016-12-31 · updated: 2026-10-02 · author: LALA

Oracle RAC Solution · Previous: Storage and ASM disks · Next: Grid patches and the multicast problem

Oracle Grid Infrastructure is the layer under the database: Clusterware, which decides which node is a member of the cluster and starts and moves resources, and ASM, which owns the shared disks. Everything prepared so far (names in DNS, kernel parameters, users, shared disks) is consumed by one graphical installer and two scripts run as root. This part describes the installation as it finally worked; the two failures on the way to it have a part of their own, Grid patches and the multicast problem.

What was installed where

The Grid Infrastructure notes open with a header block that fixes the decisions before the installer is started.

ItemValue
SoftwareGrid Infrastructure 11.2.0.4, Clusterware and ASM
Serversall RAC nodes, oradb01 and oradb02
ORACLE_BASE/data/u01/app/grid
ORACLE_HOME/data/u01/app/grid11204
Ownergrid, primary group oinstall, secondary groups asmadmin and asmdba
Permissions755
Storage of the OCR and the voting fileASM
Oracle inventory/data/u01/app/oraInventory
Cluster nameclsdb

The Grid home is not below the Grid base. The notes give the two paths without comment. My understanding, which is not something the notes say, is that this is a rule of the product for a cluster installation: root.sh changes the ownership of the home and of the directories above it to root, so the home must not sit in a path that the software owner has to keep. The directories and the two accounts were created in Operating system preparation.

The header carries one warning that I wrote down because it is easy to trip over: the cluster name must not contain certain characters, the hyphen among them. A single word is the safest choice, hence clsdb. The remark is stricter than Oracle's documentation of today, which allows hyphens in a cluster name and forbids underscores.

Starting the installer

The installer is a graphical program, so an X server has to run on the PC the installation is driven from. The X libraries on the nodes were installed in the preparation part. The installation media were unpacked in /data/install/grid. The installer was started as grid, after switching from root.

bash
$ su - grid
$ cd /data/install/grid
$ ./runInstaller

The notes do not name the node for this first run. The later run from a response file is on oradb01, and oradb01 is the first node everywhere else, so the installer ran there and reached oradb02 over SSH.

mermaid
sequenceDiagram
  participant i as Installer
  participant a as oradb01
  participant b as oradb02
  i->>a: runInstaller as grid
  i->>b: SSH connectivity, Setup and Test
  i->>i: Answers and prerequisite checks
  i->>a: Install
  i->>b: Install
  Note over i,b: Installer asks for the root scripts on both nodes
  a->>a: orainstRoot.sh as root
  b->>b: orainstRoot.sh as root
  a->>a: root.sh, OLR, ASM, disk group OCR, voting file
  Note over a: must finish before the second node starts
  b->>a: root.sh, CSS finds the active daemon on oradb01
  b->>b: restart and join the cluster
  a->>b: crsctl check cluster -all

The answers that matter

The whole answer set, screen by screen, is a Config document: Grid Infrastructure installer answers. Most screens are a matter of choosing the obvious (skip software updates, install and configure for a cluster, advanced installation, English). The ones below decide how the cluster looks.

ScreenAnswerWhy
Grid Plug and Playcluster clsdb, SCAN oradb-scan, port 1521The SCAN is the one name clients use; it resolves to three addresses in DNS, see Network and DNS plan
Grid Plug and PlayConfigure GNS not checkedAs I understand it, the Grid Naming Service would need a delegated subdomain and DHCP; all names here are static records on dns1
Cluster Node Informationoradb01.example.net with oradb01-vip.example.net, oradb02.example.net with oradb02-vip.example.netEach node gets a virtual address that Clusterware moves to the other node when its home node fails
SSH Connectivityuser grid, home not shared, reuse the existing keysThe keys were generated in the preparation part; the installer only distributes them, "Setup" and then "Test"
Storage OptionOracle ASMOCR and voting file live in a disk group, not on a cluster file system
Create ASM Disk GroupOCR, external redundancy, AU size 1 MB, disk ORCL:ASMDISK1One small LUN for the cluster's own files; redundancy is left to the storage
Failure Isolationno IPMIThe nodes are virtual machines, there is no baseboard controller to fence with
Operating System GroupsOSASM asmadmin, OSDBA for ASM asmdba, OSOPER for ASM oinstallThe groups created in the preparation part
Installation Locationbase /data/u01/app/grid, software /data/u01/app/grid11204As in the header block
Create Inventory/data/u01/app/oraInventoryShared by Grid Infrastructure and the database software installed later

The disk name ORCL:ASMDISK1 is how ASM sees a disk labelled by ASMLib; it is the LUN ORADB-CRS from Storage and ASM disks. External redundancy means ASM keeps one copy of everything and one voting file. With a single LUN there was nothing else to choose.

The SSH screen asks for the operating system password of grid. The installer needs it once, to copy the public keys to the other node; after that it works with the keys.

noteThe password of grid and the ASM password appear in the answer set as the placeholders <GRID_OS_PASSWORD> and <ASM_PASSWORD>.

Which interface does what

The installer lists every interface it finds on the node with its subnet and asks for a role.

InterfaceSubnetRole given
eth010.30.30.0Do Not Use
eth110.30.10.0Public
eth2noneDo Not Use
eth310.30.40.0Do Not Use
eth310.30.20.0Private
eth410.30.50.0Do Not Use

Two rows are not what a textbook installation shows. eth2 has no subnet, and eth3 is listed twice: once with the management subnet, which Clusterware must leave alone, and once with the interconnect subnet 10.30.20.0, which is the alias eth3:INT. This table is the state after the interconnect had been moved off eth2. Why it was moved is the second half of the next part. Backup and iSCSI networks are none of Clusterware's business and are marked "Do Not Use".

Only "Public" and "Private" matter to the cluster: the node addresses, the VIPs and the SCAN addresses are in the public subnet, and the cluster heartbeat and the traffic between the database instances go over the private one.

Prerequisite checks

The notes contradict themselves on this screen. The heading says, in brackets, that these checks stayed in the state Failed; the line under it says every prerequisite is met. Both were probably true at different times, because the installer was run more than once. One of the kept outputs supports the heading: the root.sh output of the first attempt contains the line "User ignored Prerequisites during installation", and the outputs of the later successful runs do not. Which checks failed the first time is not recorded.

The root scripts

Towards the end the installer stops and asks for two scripts to be run as root on both nodes. The first one, orainstRoot.sh, only sets the permissions of the inventory. It was run as root on oradb01 and then on oradb02, with the same output on both.

bash
$ /data/u01/app/oraInventory/orainstRoot.sh
output 6 lines
Changing permissions of /data/u01/app/oraInventory.
Adding read,write permissions for group.
Removing read,write,execute permissions for world.

Changing groupname of /data/u01/app/oraInventory to oinstall.
The execution of the script is complete.

The second one, /data/u01/app/grid11204/root.sh, is where the cluster is created, and it has an order that must be kept. I wrote the rule into the notes: root.sh has to be run on the first node first and has to finish successfully; only then can it be run on the second node. With four nodes it could run on nodes 2 and 3 at the same time, but the first and the last node each have to run alone and finish. That is the rule of the 11.2 installation guide: run root.sh on the first node and wait for it to finish; with four or more nodes it can run concurrently on all nodes but the first and the last.

On the first attempt root.sh failed on both nodes with "USM driver install actions failed". That output is kept as a Config document, root.sh output without the patch, and the fix is in the next part. What follows are the runs after the fix.

root.sh on the first node

The whole output is a Config document: root.sh output on the first node. The script works through these stages on oradb01.

StageDecisive line in the output
Local registry of the nodeOLR initialization - successful, followed by the wallets and certificates of the Grid Plug and Play profile
Start at bootAdding Clusterware entries to inittab
Lower stackora.mdnsd, ora.gpnpd, ora.cssdmonitor, ora.gipcd, ora.cssd and ora.diskmon started one after another
ASMASM created and started successfully.
Disk groupDisk Group OCR created successfully.
Cluster registrySuccessfully accumulated necessary OCR keys.
Voting fileSuccessfully replaced voting disk group with +OCR.
ResultConfigure Oracle Grid Infrastructure for a Cluster ... succeeded

The first node has nobody to join. It starts the Cluster Synchronization Services alone, creates the ASM instance and the disk group OCR on ORCL:ASMDISK1, writes the Oracle Cluster Registry into it and places the voting file there. The script then lists what it has.

output 4 lines
##  STATE    File Universal Id                File Name Disk group
--  -----    -----------------                --------- ---------
 1. ONLINE   0123456789abcdef0123456789abcdef (ORCL:ASMDISK1) [OCR]
Located 1 voting disk(s).

The file identifier printed here is a made-up value of the right shape, not the real one.

root.sh on the second node

The output on oradb02 is a third of the length: root.sh output on the second node. The node initializes its own local registry and its inittab entries, and then two lines say everything.

output 3 lines
CRS-4402: The CSS daemon was started in exclusive mode but found an active CSS daemon on node oradb01, number 1, and is terminating
An active cluster was found during exclusive startup, restarting to join the cluster
Configure Oracle Grid Infrastructure for a Cluster ... succeeded

The second node starts its CSS daemon the same way the first node did, in exclusive mode, as if it were about to create a cluster. It then finds the daemon of oradb01, gives up the exclusive start and restarts as a member. No ASM is created and no disk group: both exist already. Joining is the step that did not work while the interconnect was on a network that passed neither multicast nor broadcast.

Verification

All checks were run as root on oradb01 with crsctl from the Grid home. The commands with their complete output are a Config document: Cluster status checks.

bash
$ /data/u01/app/grid11204/bin/crsctl check cluster -all
output 11 lines
**************************************************************
oradb01:
CRS-4537: Cluster Ready Services is online
CRS-4529: Cluster Synchronization Services is online
CRS-4533: Event Manager is online
**************************************************************
oradb02:
CRS-4537: Cluster Ready Services is online
CRS-4529: Cluster Synchronization Services is online
CRS-4533: Event Manager is online
**************************************************************

Three daemons online on two nodes is the answer to "is there a cluster". What the cluster runs is shown by crsctl status resource -t. Local resources exist once per node, cluster resources once in the cluster.

ResourceKindWhat it isState
ora.LISTENER.lsnrlocalThe node listeneronline on both nodes
ora.OCR.dglocalThe disk group OCR, mountedonline on both nodes
ora.asmlocalThe ASM instanceonline, "Started", on both nodes
ora.gsdlocalGlobal Services Daemon, needed only by Oracle 9i RAC databasesoffline on both nodes, as intended
ora.net1.networklocalThe public network the VIPs depend ononline on both nodes
ora.onslocalOracle Notification Serviceonline on both nodes
ora.registry.acfslocalRegistry of ACFS file systemsonline on both nodes
ora.LISTENER_SCAN1.lsnr to ora.LISTENER_SCAN3.lsnrclusterThe three SCAN listenersone on oradb02, two on oradb01
ora.scan1.vip to ora.scan3.vipclusterThe three SCAN addresseseach beside its listener
ora.oradb01.vip, ora.oradb02.vipclusterThe node VIPseach on its own node
ora.cvuclusterPeriodic checks of the Cluster Verification Utilityonline on oradb01
ora.oc4jclusterA Java container used by Clusterware's management featuresonline on oradb01

The two lines with OFFLINE under ora.gsd look like a fault and are not one: the resource has the target OFFLINE, it is switched off on purpose in this release. The three SCAN listeners are spread as evenly as two nodes allow. ora.registry.acfs being online is, in my reading, a quiet confirmation of the patch from the next part, because the ACFS drivers are what the first root.sh could not install; the notes do not draw that connection. No database resource is in the list yet; it appears in Database creation and user environments.

The last three checks ask the daemons one by one.

bash
$ /data/u01/app/grid11204/bin/crsctl check ctss
$ /data/u01/app/grid11204/bin/crsctl check css
$ /data/u01/app/grid11204/bin/crsctl check evm
output 3 lines
CRS-4700: The Cluster Time Synchronization Service is in Observer mode.
CRS-4529: Cluster Synchronization Services is online
CRS-4533: Event Manager is online

Observer mode is the wanted answer. The Cluster Time Synchronization Service keeps the clocks of the nodes together by itself only when it finds no other time service. Both nodes have NTP configured against dns1, with the -x option that slews the clock instead of stepping it, as set up in Operating system preparation. CTSS sees that configuration, steps back and only watches; the 11.2 guide says it in one sentence: "If NTP is found configured, then the Cluster Time Synchronization Service is started in observer mode".

Reading it today

Checked against the Grid Infrastructure guides for 19c and 26ai.

As builtToday
./runInstaller from unpacked media, software copied into a new homeSince 12.2 the Grid home is an image that is extracted into place, and the wizard is gridSetup.sh
Root scripts run by hand on each nodeThe installer can run them itself, given root or sudo credentials
One 1 GB disk for the disk group OCRBelow every current minimum: 2 GB with external redundancy in 26ai
Interconnect as an alias on the management interfaceStill against the recommendation: each interface should have one role, and the private network should be physically separate
CTSS in observer modeCTSS is deprecated in 19c and desupported in 26ai; the operating system's time service is the only one
ora.gsd, ora.oc4jGSD no longer appears in the 19c and 26ai guides; the oc4j noun is deprecated and the feature it served, Quality of Service Management, is desupported in 26ai
RSA and DSA keys generated for gridThe installer's automatic SSH setup uses RSA keys; the DSA key was never needed, and OpenSSH 10 has removed DSA

The larger change is outside the installer: Release 11.2 left Extended Support in December 2020, and RAC is not available in Standard Edition 2 from 19c on.

What the notes do not hold

The notes keep the answers, the scripts and the checks. They do not keep the response file grid.rsp that the installer can save and that was used for the reinstallation described in the next part, nor the installer's log, nor any timing. The node listener's port is not stated either; only the SCAN port 1521 is. With the cluster up and one disk group in it, the next step on the working path is the database software: Database software, listener and disk groups.

← solutionz