LINUXOR.SK ... open source notes ...

NetApp 05 - Storage design

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

NetApp Solution · Previous: Network design · Next: SVMs and NFS

A site has 272 disks in twelve shelves, half of them in each datacenter. This article is about who owns which disk, how the disks become mirrored aggregates, and what volumes sit on them. The layout is lopsided on purpose: one cluster owns almost everything, the other lives on a few disks of the shortest shelf. The numbers are those of site 1 (DC1) in the as-built report of 2018, with the changes of the final design noted.

Shelves and disks

DatacenterShelf IDsDisksSeen by cluster A asSeen by cluster B as
A10, 1124 + 242.10, 2.114.10, 4.11
A20, 2124 + 243.20, 3.218.20, 8.21
A30, 3124 + 161.30, 1.315.30, 5.31
B50, 5124 + 246.50, 6.511.50, 1.51
B60, 6124 + 245.60, 5.612.60, 2.61
B70, 7124 + 164.70, 4.713.70, 3.71

The disks are 10K SAS drives that ONTAP reports as 1.2TB_SAS_10k, with 1.09 TB usable each. A disk is named stack, shelf, bay: 2.10.13 is bay 13 of shelf 10. The shelf ID is set on the shelf and is the same for everybody. The stack number in front of it is assigned by each cluster for itself, which is why the same shelf has two names in the table. It did not stay the same over time either: the table is from September 2017, and a later version of the layout spreadsheet has 7.10, 5.30, 4.50 and 1.70. In the working notes the short shelves appear under four names, 1.31, 5.31, 4.71 and 3.71, each with the remark "16 disks".

Both clusters see all twelve shelves, because the shelves hang on the two fabrics and not on a controller. See Physical design and cabling.

Pools and ownership

A mirrored aggregate needs disks in two places. ONTAP keeps them apart with pools. For every node, pool 0 holds the disks in its own datacenter and pool 1 the disks in the other one. An aggregate has one plex in each pool, and SyncMirror writes to both plexes synchronously.

The design spreadsheet gives every shelf a pool, seen from the cluster that owns it:

ShelvesOwnerPool for the owner
10, 11, 20, 21, 30Nodes of cluster A0, local
50, 51, 60, 61, 70Nodes of cluster A1, remote
71, 11 of its 16 disksNodes of cluster B0, local
31, 11 of its 16 disksNodes of cluster B1, remote
31, the other 5 disksNode 1 of cluster A0, local
71, the other 5 disksNode 1 of cluster A1, remote

The spreadsheet writes the last four rows as pool 1/0 for shelf 31 and 0/1 for shelf 71: a shelf with two owners that look at it from opposite sides.

The as-built report counts it up per node:

NodeDisks ownedIn aggregatesSpare
DC1-A-ANAS0011305872
DC1-A-ANAS00212014106
DC1-B-ANAS00116142
DC1-B-ANAS002642

Cluster A owns 250 of the 272 disks of the site, cluster B owns 22. The ten full shelves are split evenly between the two nodes of cluster A, 60 disks each in each pool.

This is not what the system looked like after its first setup. A listing from September 2017 shows the two short shelves owned completely by cluster B, with a data aggregate on each of its nodes. In the final layout node 2 of cluster B has no data aggregate, and five disks of each short shelf have gone to node 1 of cluster A. The commands of that rearrangement were not kept in the notes. The notes file on disk assignment holds two links to the NetApp documentation and one sysconfig -A, followed by material that belongs to the directory integration; the file on physical storage holds the four shelf names above. What can be shown is the state before and the state after.

The commands that were used to look at ownership are in the notes, under a heading that translates as "hmm, questions". On cluster A:

bash
$ storage shelf show
$ storage disk option show
$ storage disk show -ownership
$ storage disk show -ownership -container-name DC1_A_ANAS001_root
$ storage disk show -container-type spare
$ storage disk show -container-type aggregate
$ storage aggregate show -disk

Aggregates

AggregateOwnerRAIDDisksPer plexSizeHolds
DC1_A_ANAS001_rootDC1-A-ANAS001RAID442953.80 GBvol0 of the node
DC1_A_ANAS001_data1DC1-A-ANAS001RAID-DP542721.42 TBKVM and backup SVMs
DC1_A_ANAS002_rootDC1-A-ANAS002RAID442953.80 GBvol0 of the node
DC1_A_ANAS002_data1DC1-A-ANAS002RAID-DP1052.79 TBBalabit SVM
DC1_B_ANAS001_rootDC1-B-ANAS001RAID442953.80 GBvol0 of the node
DC1_B_ANAS001_data1DC1-B-ANAS001RAID-DP1052.79 TBRoot volume of DC1-S-VCVSM004, metadata
DC1_B_ANAS002_rootDC1-B-ANAS002RAID442953.80 GBvol0 of the node

Every aggregate is mirrored, normal. The root aggregates have the high-availability policy cfo, the data aggregates sfo. The maximum RAID group size is 16 for the data aggregates and 8 for the root aggregates.

A root aggregate is two plexes of two whole disks each, one data and one parity. The notes contain a reading list on advanced disk partitioning, which would have put the root on slices of shared disks, and nothing else on the subject. It was never an option: NetApp's documentation lists root-data partitioning as not supported in a fabric-attached MetroCluster. The as-built state shows whole disks: sixteen disks of the site hold nothing but four copies of vol0 and their mirrors, and vol0 fills each root aggregate to 95 per cent.

The large aggregate has 27 disks per plex. With a group size of 16 that is one full RAID group and one of 11, and after two parity disks per group 23 data disks, which matches the 21.42 TB. The final design lists the aggregate with 64 disks, 32 per plex and so two full groups; site 2 has the same aggregate with 38.

Why node 1 holds almost everything

The Source material states the layout and not the reasoning. What the layout itself shows:

The layout on the shelves

The design has a spreadsheet with one cell per disk, coloured by aggregate. Counted per shelf:

mermaid
flowchart LR
  subgraph fa["Datacenter A: plex 0 of cluster A, mirror plexes of cluster B"]
    s10["Shelf 10: data1 of A1 x3, root of A1 x1, data1 of A2 x2, root of A2 x1, spare x17"]
    s11["Shelf 11: data1 of A1 x7, root of A1 x1, data1 of A2 x2, root of A2 x1, spare x13"]
    s20["Shelf 20: data1 of A1 x6, data1 of A2 x1, spare x17"]
    s21["Shelf 21: data1 of A1 x5, spare x19"]
    s30["Shelf 30: data1 of A1 x4, spare x20"]
    s31["Shelf 31: data1 of B1 x5, root of B1 x2, root of B2 x2, spare of B x2, data1 of A1 x2, spare of A x3"]
  end
  subgraph fb["Datacenter B: mirror plexes of cluster A, plex 0 of cluster B"]
    s50["Shelf 50: data1 of A1 x5, root of A1 x1, data1 of A2 x1, root of A2 x2, spare x15"]
    s51["Shelf 51: data1 of A1 x3, root of A1 x1, data1 of A2 x3, spare x17"]
    s60["Shelf 60: data1 of A1 x4, data1 of A2 x1, spare x19"]
    s61["Shelf 61: data1 of A1 x6, spare x18"]
    s70["Shelf 70: data1 of A1 x5, spare x19"]
    s71["Shelf 71: data1 of B1 x5, root of B1 x2, root of B2 x2, spare of B x2, data1 of A1 x4, spare of A x1"]
  end
  fa == "SyncMirror over the two fabrics" === fb

A1 and A2 are the nodes of cluster A, B1 and B2 those of cluster B. The plex names in the spreadsheet follow the same split:

AggregatePlex in datacenter APlex in datacenter B
DC1_A_ANAS001_data1, DC1_A_ANAS002_data1plex0plex1
DC1_A_ANAS001_root, DC1_A_ANAS002_rootplex0plex4
DC1_B_ANAS001_data1plex1plex0
DC1_B_ANAS001_root, DC1_B_ANAS002_rootplex4plex0

Two things in the picture are not tidy, and both have a history.

The disks of an aggregate are not next to each other. The root aggregates and the data aggregate of node 2 sit in single bays spread over shelves 10, 11 and 20, and their mirrors over 50, 51 and 60. The listing of September 2017 has the same bays under the names the first setup gave them, root_DC1_A_ANAS001 and so on. They were renamed and left where they were.

The two plexes of the large aggregate are not symmetrical. In datacenter A it has 25 disks in the full shelves and 2 in shelf 31; in datacenter B, 23 and 4. Those are among the disks taken over from cluster B: several of the bays are ones the removed data aggregate of DC1-B-ANAS002 had used.

Volumes

Cluster A in the as-built report:

VolumeSVMAggregateSizeSnapshot policy
vol0Node DC1-A-ANAS001DC1_A_ANAS001_root902.54 GBNone
vol0Node DC1-A-ANAS002DC1_A_ANAS002_root902.54 GBNone
MDV_CRS_<ID>_AClusterDC1_A_ANAS001_data110 GBnone
MDV_CRS_<ID>_BClusterDC1_A_ANAS002_data110 GBnone
DC1_S_VCVSM001_rootDC1-S-VCVSM001DC1_A_ANAS002_data11 GBdefault
DC1_S_VCVSM001_dataDC1-S-VCVSM001DC1_A_ANAS002_data1100 GBdefault
DC1_S_VCVSM002_rootDC1-S-VCVSM002DC1_A_ANAS001_data11 GBdefault
DC1_S_VCVSM002_dataDC1-S-VCVSM002DC1_A_ANAS001_data110.10 TBnone
DC1_S_VCVSM003_rootDC1-S-VCVSM003DC1_A_ANAS001_data11 GBdefault
DC1_S_VCVSM003_dataDC1-S-VCVSM003DC1_A_ANAS001_data1100 GBdefault

Cluster B has the two vol0, the root volume DC1_S_VCVSM004_root of 1 GB, and its own pair of MDV_CRS volumes, both on DC1_B_ANAS001_data1 because there is no second data aggregate to put the other one on.

The MDV_CRS volumes are not created by an administrator. MetroCluster makes them when it is configured and keeps in them the metadata of its configuration replication, the mechanism by which the partner cluster learns about SVMs, interfaces, exports and policies it must be able to bring up after a switchover. The design document describes them as "root volume for cluster SVM, partition 1 and 2", which is not what they are but is why they appear in its volume table.

The final design adds the volumes of 2019 to DC1_A_ANAS001_data1: root and data volumes for DC1-S-VCVSM005 and DC1-S-VCVSM006, and DC1_S_VCVSM002_data2, described as a temporary data volume "to be removed after agreement with the virtualization team". Their sizes are not in the Source material.

The settings that are the same for every volume:

SettingValue
Space guaranteevolume, so thick provisioned
Deduplication and compressionNot enabled on any volume; the efficiency policies default and inline-only exist and are unused
AutosizeDisabled; the first thing to try on a full volume is set to volume_grow
Snapshot autodeleteDisabled
EncryptionNetApp Volume Encryption on the data volumes; see Encryption and certificates

The snapshot policy default keeps six hourly, two daily and two weekly copies. The cluster also has default-1weekly, which keeps one weekly copy, and none. The KVM datastore is on none: the images of running virtual machines were not snapshotted on the storage. On the partner cluster the policies of a mirrored SVM appear with the suffix -DR.

The same state, as listings in the shape of the ONTAP commands, is a Config document: ONTAP: disk ownership, aggregates and volumes as built. How the volumes are exported is in SVMs and NFS.

Site 2

Site 2 (DC2) has the same shelves with the same IDs, the same four small aggregates, and the same split of the short shelves. The differences in the layout spreadsheet of site 2:

ItemSite 1Site 2
Large aggregate54 disks as built, 64 in the final designDC2_A_ANAS001_data1, 38 disks
Mirror plex namesplex1, plex4plex1, plex6, plex8
Disk states in the legendIn an aggregate or spareAlso "unassigned"
Stack numbers of the shelvesAs in the first tableDifferent: 1.10, 7.20, 6.30, 2.50, 8.60, 3.70
MDV_CRS volumes of cluster AOne on each data aggregateBoth on DC2_A_ANAS001_data1

The higher plex numbers suggest that mirror plexes at site 2 were removed and created again more than once; the notes do not say when or why. The spreadsheet of site 2 also still labels its rows with the datacenters of site 1, and the aggregate table of site 2 in the design contains one row with a site 1 name. No as-built report of site 2 is in the Source material, so the counts of site 2 are those of the design and not of the running system.

Lessons

← solutionz