LINUXOR.SK ... open source notes ...

Linux Storage 10 - RDAC problems and LUN removal

category: solutionz · date: 2010-09-30 · updated: 2026-10-03 · author: LALA

Linux Storage Solution · Previous: RDAC installation and operation

Two problems close the RDAC notes, and both come from the same fact: mppsrv01 was zoned to two arrays, the main DS5300-SITEA and its mirror DS5300-SITEB. The first problem is that the driver dutifully presented both, so the host had two copies of every disk. The second is the general question the first leaves open, how to take a LUN out of a running Linux system without a reboot, which the notes answer with sysfs and end with a small accident that worked anyway. Everything ran as root on mppsrv01; the full listings and command sets are two Config documents, Removing the standby array from RDAC and Hot removal of a LUN. The Article ends with the links my notes kept in 2010.

Problem 1: the mirror is a second array

After the mppBusRescan of the previous Article the host saw two arrays, and lsvdev mapped the main array to /dev/sda and /dev/sdb and the standby array to /dev/sdc and /dev/sdd. The notes state the problem in one sentence: the server sees two disk arrays, with the main one mirrored to the standby one, which is a problem mainly if LVM is in use, where devlabel cannot be used to identify the device being attached. My reading of why, which the notes do not spell out: a mirror carries the same content, so the LVM labels and UUIDs on sdc and sdd are those of sda and sdb, and any tool that identifies a disk by what is written on it, devlabel included, finds two candidates for each volume. The wanted state was a host that knows only the main array; the standby array is there for a failover that would be done on the array side, which the notes do not describe.

mermaid
flowchart LR
  subgraph bad["Faulty state"]
    b1["mppUtil -a, ID 0 DS5300-SITEA"]
    b2["mppUtil -a, ID 1 DS5300-SITEB"]
    b3["devicemapping, 0 and 1"]
    b4["aabbcc01.wwn and aabbcc02.wwn"]
    b5["sda, sdb, sdc, sdd"]
  end
  subgraph good["Wanted state"]
    g1["mppUtil -a, ID 0 DS5300-SITEA"]
    g3["devicemapping, 0 only"]
    g4["aabbcc01.wwn only"]
    g5["sda, sdb"]
  end
  bad -- "chkconfig mpp off, mppUpdate -d, rm .wwn, mppUpdate, init 6" --> good

The faulty state is documented three ways in the notes, and the wanted state the same three ways: mppUtil -a with two arrays against one, cat ./devicemapping with two lines against one, and ls -l in /var/mpp with two .wwn files against one. The two sets differ only in the standby array's lines; the "wanted" outputs even carry the same time stamp, GMT 09/08/2010 10:50:35, as the faulty ones, which suggests they were written by editing the faulty listings rather than captured after the fix; the notes do not say. The disk names in the diagram come from lsvdev, not from these listings.

The fix is six commands, in order:

bash
$ chkconfig mpp off
$ chkconfig --list | grep mpp
$ mppUpdate -d DS5300-SITEB
$ rm ./600a0b800012345700000000aabbcc02.wwn
$ mppUpdate
$ init 6
noteThe array WWN in the file name is a made-up value of the right shape.

First the attempts to discover new arrays at boot were turned off, and checked. The notes' own list of commands explains why: /etc/init.d/mpp discovers devices at boot, and with more than one array it is worth turning it off so that a restart causes no unwanted trouble; otherwise the next boot would find the mirror again. Then the array was taken out of the persistent mapping with mppUpdate -d DS5300-SITEB, which confirmed "The module DS5300-SITEB was removed from the MPP persistence mapping list (/var/mpp/devicemapping).", counted the two QLogic ports and rebuilt the initrd. That command does not touch the array's .wwn file, so the file was removed by hand, in /var/mpp, and mppUpdate was run once more to rebuild the initrd without it; as I understand it the .wwn files are packed into the initrd, which is why the second rebuild was needed. The machine was restarted with init 6, and after the boot mppUtil -a and cat ./devicemapping were to show one array and one line. Their output is not in the notes; the edited "wanted state" listings stand in for it.

What this fix does not do is remove the eight physical devices of the standby array from the SCSI layer while the system runs, and it does not say whether the zoning on the SAN was changed afterwards. That is where the second problem starts.

Problem 2: hot removal of a LUN

The notes collect the ways to tell the SCSI mid-layer to drop a device, all of them a write into sysfs. By disk name, for the two virtual disks of the standby array:

bash
$ echo "1" > /sys/block/sdc/device/delete
$ echo "1" > /sys/block/sdd/device/delete

By SCSI address, where h is the HBA number, c the channel on the HBA, t the SCSI target ID and l the LUN: echo 1 > /sys/class/scsi_device/h:c:t:l/device/delete. Which disk is which, with its address and identifiers, comes from ls -l /dev/disk/by-*; the notes keep the command and no output. Then the notes quote a note in English about taking a path offline with echo offline > /sys/block/sda/device/state, so that any further I/O on that path fails immediately while "Device-mapper-multipath will continue to use the remaining paths to the device". It is a sentence about DM-Multipath, which this host did not run, and its sda is an example name; on mppsrv01 sda was LUN 1 of the main array and nothing to go offline. I keep it because the notes kept it.

The notes also keep a small script titled "who uses what", which walks /sys/block and prints name, vendor and model for every device that has a vendor attribute, so that the VirtualDisk devices of the MPP adapter can be told from the local disks. Its echo line reads echo “${DSK} ${VENDOR} ${MODEL} with a typographic opening quote and no closing quote; as I understand bash, the typographic quote is an ordinary character and the output starts with a stray “, but I did not run it. It is kept exactly as written in the Config document and should be retyped before use.

The case actually done is the removal of all paths to the standby array. First the driver was asked which they are:

bash
$ mppUtil -g 1 | grep PathId
output 4 lines
PathId: 77000002 (hostId: 0, channelId: 0, targetId: 2)
PathId: 77010002 (hostId: 1, channelId: 0, targetId: 2)
PathId: 77000003 (hostId: 0, channelId: 0, targetId: 3)
PathId: 77010003 (hostId: 1, channelId: 0, targetId: 3)

Array ID 1 is DS5300-SITEB; its controllers are targets 2 and 3 on both HBAs, which is what mppBusRescan had found. With two LUNs on each path that is eight SCSI devices, and all eight were deleted, in the order of the path list, with lines of this shape:

bash
$ echo 1>/sys/bus/scsi/devices/0\:0\:2\:1/delete
$ echo 1>/sys/bus/scsi/devices/0\:0\:2\:2/delete

There is no space before the >. The shell reads 1> as a redirection of file descriptor 1, so echo runs with no argument and writes a bare newline into the delete attribute; echo 1 > was meant. As I understand it the kernel removes the device on any write to that attribute, so the eight lines did what was wanted, but by accident, and I keep them as typed in the Config document with that warning. The escaped colons (0\:0\:2\:1) are not needed either, and harmless. What the notes do not hold: whether the two sdc/sdd deletions and the eight path deletions were both done or are alternatives, a listing before and after, and whether mppBusRescan -d, the driver's own way to remove unmapped or disconnected devices, was tried.

References the notes kept

The notes end with three lists of links, which I give here as kept in 2010, with their titles translated where they were not English. The author did not re-check them when the notes were written up, and most of them have moved or gone since; what a check on 2026-10-03 found is after the list.

The forum thread about hot removal of unmapped LUNs and the Red Hat online storage guide are the two that the second problem leans on; the DM-Multipath quotation almost certainly came from the latter, though the notes do not say so. Two things the research for this write-up found when it checked the list on 2026-10-03. The Redbook behind the "Using RDAC multipath driver" link, SG24-6700, is "Configuration and Tuning GPFS for Digital Media Environments", not a storage book; the DS4000 and DS5000 book with the RDAC and AVT guidance is SG24-6363, "IBM Midrange System Storage Implementation and Best Practices Guide", still downloadable from IBM. And several of the links no longer resolve or land on a vendor's front page: the LSI RDAC pages redirect to a Broadcom "offline" page, the Citrix article redirects to the support home, the Sun link is broken while Oracle still hosts the same installation guide, and the Redbook link itself returns 404.

What I would do differently

← solutionz