LINUXOR.SK ... open source notes ...

Oracle RAC 07 - Grid patches and the multicast problem

category: solutionz · date: 2016-12-31 · updated: 2026-10-02 · author: LALA

Oracle RAC Solution · Previous: Grid Infrastructure installation · Next: Database software, listener and disk groups

The installation in the previous part reads as one pass. It was not. root.sh failed twice, for two unrelated reasons: first on both nodes, because Grid Infrastructure 11.2.0.4 as shipped could not install its storage drivers on SUSE Linux Enterprise Server 11 SP3, and then on the second node, because the network under the interconnect did not carry what the cluster needs to find its members. The first problem had a patch. The second one looked like it had a patch too, and that cost the most time. This part keeps both, with the dead end, because the dead end is the useful half.

What was tried, and what it led to

mermaid
flowchart TB
  a["Grid 11.2.0.4 installed, root.sh run"] --> b["USM driver install actions failed, on both nodes"]
  b --> c["OPatch replaced, patch 6880880"]
  c --> d["Patch 17475946 applied with opatch napply -local"]
  d --> e["root.sh on oradb01 succeeds"]
  e --> f["root.sh on oradb02 fails, node does not join"]
  f --> g["Cumulative patches 20760982 and 20996923 applied"]
  g --> h["No change, marked in the notes as not this"]
  f -. "when experimenting" .-> i["Deconfigure with rootcrs.pl, reinstall from a response file"]
  f --> j["Cause: Hyper-V networks pass no multicast and no broadcast"]
  j -. "how the move was done is not recorded" .-> k["eth2 unused, interconnect is eth3:INT"]
  k --> l["root.sh on oradb02 joins the cluster"]

The first failure: USM driver install actions failed

After the installer had copied the software, root.sh was run as root on oradb01. It went through the generic part, initialized the local registry, created the wallets, added the Clusterware entries to inittab, and stopped.

output 3 lines
Adding Clusterware entries to inittab
USM driver install actions failed
/data/u01/app/grid11204/perl/bin/perl -I/data/u01/app/grid11204/perl/lib -I/data/u01/app/grid11204/crs/install /data/u01/app/grid11204/crs/install/rootcrs.pl execution failed

The same message came on oradb02. The whole output is a Config document: root.sh output without the patch.

USM is the part of Grid Infrastructure that brings kernel drivers with it, for the ASM cluster file system (ACFS) and its volume manager (ADVM); root.sh installs them. Kernel drivers are built for particular kernels, and the nodes ran SP3 of SLES 11 with kernel 3.0.76-0.11-default. The notes give no analysis of the failure. They give the name of the problem as I found it described at the time, "suse11sp3 execute root.sh treatment failure bug", a reference to a blog post on dreamsanqin.blog.51cto.com, and the number of the patch that solves it. Nothing in this cluster uses ACFS, but root.sh does not ask: without the drivers it does not continue to the step that starts the cluster stack.

The fix: a newer OPatch, then patch 17475946

Two archives were needed in /data/install.

ArchiveWhat it is
p17475946_112040_Linux-x86-64.zipThe patch for Grid Infrastructure 11.2.0.4 on SLES 11 SP3
p6880880_112000_Linux-x86-64.zipA newer OPatch, the tool that applies patches

The patch could not be applied with the OPatch that the 11.2.0.4 media put into the Grid home; the tool had to be updated first. OPatch is not installed, it is replaced: the old directory is moved aside and the new one is copied into the home. I followed a short description of that at jamescoyle.net. All of this ran as root: both archives were unpacked in /data/install, the unpacked patch directory 17475946 got mode 777 recursively, the OPatch directory of the Grid home was renamed to OPatch-old, and the new OPatch was copied into the home and given to grid:oinstall with mode 755. The commands as the notes have them are in the Config document Grid patch commands.

The mode 777 on the unpacked patch is the blunt way to make sure that grid can read what root unpacked.

Then comes the step that is easy to miss. OPatch runs as the owner of the software, grid, and it has to write into the Grid home. The failed root.sh had already been at work in that home, and as I understand it one of the things root.sh does in a cluster installation is to hand parts of the home to root; the notes show the chown commands and give no reason for them. So before the patch, the home directory itself and, recursively, bin and lib were given to grid with chown, as root (the notes start the step with su -). After su - grid the patch was applied as grid.

bash
$ cd /data/u01/app/grid11204/OPatch/
$ ./opatch napply -oh /data/u01/app/grid11204 -local /data/install/17475946

napply applies the patch found in the given directory, -oh names the home, and -local restricts the work to the node where the command runs. With -local nothing is propagated, so the whole procedure belongs on each node. The notes show the commands once and without a node name; the first failure was on both nodes, and the patch had to be in both homes. The complete command set is a Config document: Grid patch commands.

After the patch both root scripts were run again, orainstRoot.sh and then root.sh, first on oradb01. This time the script went past the inittab line, started the stack, created ASM and the disk group and ended with "Configure Oracle Grid Infrastructure for a Cluster ... succeeded": root.sh output on the first node. In that second run the script no longer copies dbhome, oraenv and coraenv to /usr/local/bin but reports that their contents have not changed, which is how the two outputs can be told apart at a glance.

Two things the notes get wrong about this patch

The patch number is written two ways. The step titles and the captions of the kept outputs say 1747594, seven digits. The archive and the directory, which are what was typed and what worked, say 17475946. I take the file names as correct and the titles as a typing error repeated by copying, but the notes do not settle it.

The step that applies the patch is titled "apply patch 1747594 and patch sets 20760982 and 20996923". It shows one opatch command, for the first number only. The other two appear in a later section that is fenced off as a dead end. The title is a leftover from the time when I believed all three belonged together.

The second failure: the second node does not join

With the first node up, root.sh on oradb02 should do what the previous part describes: start CSS in exclusive mode, find the active daemon on oradb01, restart and join. It did not. The notes say that root.sh on the second node did not work and that the node could not be put into the cluster at all.

That is all they say about the symptom. No output of this failure was kept: the three outputs at the end of the notes are the first failure, the success on the first node and the success on the second node. No log excerpt, no error number.

The dead end: cumulative patches

Having just fixed one root.sh failure with a patch, I looked for another patch. Two cumulative patches for 11.2.0.4 were applied, 20760982 and 20996923, in the belief that they would make root.sh run through. The procedure is the one from above: the same three chown commands as root and opatch napply as grid for /data/install/20760982, then two more directories of the home, crs and jlib, given to grid and opatch napply for /data/install/20996923. The lines are in the Config document Grid patch commands, marked there as the dead end.

Then the root scripts again, on oradb01 and on oradb02. It changed nothing. In the notes this section has no number, only an [X], and it stands between two banners that say, in capitals, "not this". I left it in as a warning to myself. The notes do not say whether the two patches were rolled back or stayed in the home that went into operation.

The real cause: the network

The finding is in the title of that same crossed-out section, and nowhere else: root.sh on the second node, and the joining of the second node as such, failed because Hyper-V networks bound to the "Cloud Root Network" support neither multicast nor broadcast, which the Oracle cluster uses for the communication between its nodes.

That fits how the product works, and it was in the documentation all along. The 11.2 installation guide says that with 11.2.0.2 and later releases multicasting is required on the private interconnect, on the subnet ranges 224.0.0.0/24 and 230.0.1.0/24, and its troubleshooting chapter has an entry with nearly my symptom as its title: "root.sh or rootupgrade.sh Script Fails on the Second Node Due to Multicast Issues". The entry names a test tool, mcasttest.pl. A network that silently drops multicast gives exactly this picture: the first node comes up, because it needs nobody, and the second never finds it, even where ping and SSH between the nodes work. A unicast test proves nothing here.

How I got from the symptom to the cause is not in the notes. There is no test command, no trace, no name of the tool; only the conclusion. I will not reconstruct it.

The outcome is visible in the Network and DNS plan. Each node has five interfaces, and the one first meant for the interconnect is not in use.

InterfaceRole in the plan
eth2unused "because of the Hyper-V switches": the switches bound to the Cloud Root Network pass neither multicast nor broadcast
eth3management, 10.30.40.x
eth3:INTOracle RAC interconnect, 10.30.20.x, an alias on the management interface

So the interconnect ended as a second address on the management interface. The notes do not say why that network behaved differently from the one behind eth2; they only show that the cluster formed on it. Three other places show the same end state. The installer's interface table in the answer set lists eth3 twice, once as "Do Not Use" with the management subnet and once as "Private" with 10.30.20.0, and eth2 with no subnet. The Oracle block of sysctl.conf has a comment that eth3 (eth3:INT) is used as the interconnect, and under it net.ipv4.conf.eth3.rp_filter = 0 and net.ipv4.icmp_echo_ignore_broadcasts = 0. And the output of root.sh on the second node ends with the join: root.sh output on the second node.

The move itself is not described. The notes hold no interface configuration file for eth3:INT, no Hyper-V setting, and no kept sequence of commands that reconfigured the cluster for the final network. An interconnect that shares an interface with management is a compromise, not a design; the notes record that it was made, not what it cost.

Deconfigure and reconfigure

The notes have a chapter for taking the configuration off again and putting it back, under the warning that it applies only if you are experimenting with the cluster stack. They do not say when it was used. The remark in step [4.3] about making the change of the network subnet suggests that it belonged to the move of the interconnect; that is my reading, not a statement of the notes. The command set is a Config document: Grid deconfigure and reconfigure commands. I took the procedure from a page on ewan.cc.

StepWhereWhat
Deconfigureoradb02, then oradb01, as rootrootcrs.pl -deconfig -force -verbose; on the last node also -keepdg -lastnode
Put the GPnP profiles asideboth nodesthe two profile.xml files under gpnp are renamed to profile.xml.0
Keep the parametersboth nodescrsconfig_params is copied to crsconfig_params.0
Reinstalloradb01, as gridrunInstaller -responseFile /data/install/grid.rsp, then "Next" through the screens

The order of the first step is the reverse of the installation: the node that joined last is taken out first, and the first node goes last. The two commands, as root.

bash
$ /data/u01/app/grid11204/crs/install/rootcrs.pl -deconfig -force -verbose
$ /data/u01/app/grid11204/crs/install/rootcrs.pl -deconfig -force -verbose -keepdg -lastnode

The first line is for oradb02, the second for oradb01. As I understand the two options, which the notes do not explain, -lastnode tells the script that no cluster remains after this node, so the cluster registry and the voting file go as well; -keepdg keeps the disk group, so that the reinstallation finds OCR on ORCL:ASMDISK1 again.

The Grid Plug and Play profile is, as I understand it, the small XML file in which a node remembers the cluster it belongs to, including its networks. A node that still has the old profile starts with the old idea of the cluster. Hence the renaming, in /data/u01/app/grid11204/gpnp, of profiles/peer/profile.xml and of the copy under the node's own name.

crsconfig_params is the parameter file that root.sh reads; every root.sh output names it in the line "Using configuration parameter file". The note beside the backup says to keep the parameters used during the installation and, where needed, to make the change of the network subnet. That is the only hint in the notes about how a changed interconnect gets into a reconfigured cluster, and it is a hint, not a record: no edited line of the file is shown.

The last step runs the installer again with the answers taken from a response file, with nothing to do in the graphical interface but press "Next". The response file grid.rsp itself is not in the notes.

What the notes establish and what they do not

QuestionIn the notes
Why did the first root.sh fail?The error line and the patch that fixed it; no analysis
On which nodes was the patch applied?Not stated; -local and the failure on both nodes imply both
What did the second failure look like?Not kept
Did the cumulative patches help?No; the section is marked "not this"
What was the cause?Hyper-V networks without multicast and broadcast, stated in a section title
How was it diagnosed?Not recorded
What was changed?Only the result: eth2 unused, interconnect on eth3:INT
With which commands?Not recorded; the deconfigure chapter is the nearest thing

Reading it today

Checked against the Grid Infrastructure guides for 19c and 26ai.

As builtToday
Base software installed, root.sh fails, one-off patch applied afterwardsPatches are applied during setup: gridSetup.sh -applyRU with -applyOneOffs, which would have avoided the failed first root.sh
opatch napply -local as grid after chown by handFor a configured Grid home the documented tool is opatchauto, run as root; plain opatch as the Grid user is documented for the case that the stack does not start
rootcrs.pl -deconfig -forceThe script is rootcrs.sh; -lastnode still completes the deconfiguration including OCR and voting files. -keepdg does not occur in the current guides
GPnP profiles moved aside, reinstallation from a response fileNot a documented procedure, then or now. The documented path is rootcrs.sh -deconfig -force on the failed node, correct the cause, run root.sh again
Multicast on the interconnectStill required: 224.0.0.0/24 and optionally 230.0.1.0/24. Broadcast must work as well
net.ipv4.icmp_echo_ignore_broadcasts = 0Not an Oracle requirement in the 11.2, 19c or 26ai guides. It makes a host answer broadcast pings; it does not make a network carry broadcast

For Hyper-V network virtualisation Microsoft documents that broadcast and subnet multicast are implemented by unicast replication and that IGMP and multicast routing across subnets are not supported. The term "Cloud Root Network" from my notes has no source there; the notes call it the so-called Cloud Root Network and do not say whose term it is.

What I would do differently

The notes support one lesson, and it is about order. The second failure was met with the tool that had fixed the first one, and two cumulative patches went into a home before I asked whether the two nodes could hear each other in the way the cluster talks. Multicast between the interconnect addresses is a property of the platform that can be tested before the installer is ever started, with the mcasttest.pl tool that Oracle's troubleshooting chapter names, and on a virtual platform whose switches I do not control it is the first thing to test. I would also write down the diagnosis, not only the verdict: the title of a crossed-out section is a poor place for the most expensive finding of the whole build.

← solutionz