Namespaces 13 - CGROUP namespace
Linux Namespaces Learning · Previous: USER namespace · Next: TIME namespace
CGROUP (control group) namespaces in Linux
1 Introduction
Control groups (cgroups) are the kernel mechanism that puts processes into a tree of groups and applies resource limits and accounting to each group: memory, CPU, number of processes and so on. The CGROUP namespace does not limit anything. It only changes the view: which directory of that tree a process sees as the root. Two things are virtualised, the paths shown in /proc/PID/cgroup and the root of a cgroup filesystem mounted from inside the namespace. The cgroup a process is in at the moment the namespace is created becomes the root of its view.
The examples were run as an ordinary user on a 6.12 kernel with the unified cgroup v2 hierarchy. The older cgroup v1, with one hierarchy per controller, is not covered. As in the previous articles, --user --map-root-user is there only so that no root is needed.
2 Working with CGROUP namespaces
2.1 Your cgroup in the default namespace
Check that the host uses cgroup v2 and print the cgroup of the shell. On cgroup v2 it is a single line that starts with 0::, followed by the path from the root of the hierarchy. Here it is a scope inside the user's systemd instance.
$ stat -fc %T /sys/fs/cgroup $ cgroup2fs $ cat /proc/self/cgroup $ 0::/user.slice/user-1000.slice/user@1000.service/app.slice/terminal.scope $ readlink /proc/self/ns/cgroup $ cgroup:[4026531835]
2.2 Inside a new CGROUP namespace
In a new CGROUP namespace the same file reads 0::/: the cgroup the process was in is now the root of its view. Paths of processes outside that subtree are shown relative to it, with .. components; this is how the cgroup of the host's PID 1 looks from inside.
$ unshare --user --map-root-user --cgroup sh -c "readlink /proc/self/ns/cgroup; cat /proc/self/cgroup; cat /proc/1/cgroup" $ cgroup:[4026534269] $ 0::/ $ 0::/../../../../../init.scope
2.3 What the host sees for the same process
In terminal 1, leave a process running in a new CGROUP namespace.
$ unshare --user --map-root-user --cgroup sh -c 'exec sleep 633'
In terminal 2, in the default namespace, the process is in the same cgroup as before. Nothing was moved; only its own view changed. lsns -t cgroup shows the namespace, and nsenter -C enters it.
$ cat /proc/1230184/cgroup $ 0::/user.slice/user-1000.slice/user@1000.service/app.slice/terminal.scope $ lsns -t cgroup $ NS TYPE NPROCS PID USER COMMAND $ 4026531835 cgroup 30 4085 user catatonit -P $ 4026534270 cgroup 1 1230184 user sleep 633 $ nsenter -t 1230184 -U -C --preserve-credentials cat /proc/self/cgroup $ 0::/
2.4 Mounting cgroup2 inside
A cgroup filesystem mounted from inside the namespace has the namespace's root cgroup as its top directory. Mounting needs a mount namespace as well. On the host, /sys/fs/cgroup starts at the real root, with all controllers:
$ ls /sys/fs/cgroup | head -5 $ cgroup.controllers $ cgroup.max.depth $ cgroup.max.descendants $ cgroup.procs $ cgroup.stat $ cat /sys/fs/cgroup/cgroup.controllers $ cpuset cpu io memory hugetlb pids rdma misc
Mounted inside, the top directory is the scope from step 2.1: it has the files of an ordinary, non-root cgroup (cgroup.type, memory.max), and only the controllers that were handed down to it.
$ unshare --user --map-root-user --cgroup --mount sh -c "mount -t cgroup2 none /mnt && ls /mnt; cat /mnt/cgroup.controllers; cat /mnt/cgroup.type" $ cgroup.controllers $ cgroup.events $ cgroup.freeze $ cgroup.kill $ ... $ memory.max $ memory.min $ ... $ pids.max $ pids.peak $ memory pids $ domain
Without the CGROUP namespace the same mount is refused for an unprivileged user namespace, because it would expose the whole tree:
$ unshare --user --map-root-user --mount sh -c "mount -t cgroup2 none /mnt" $ mount: /mnt: permission denied. $ dmesg(1) may have more information after failed mount system call.
2.5 Where the limits come from
Limits are set by writing to the controller files of a cgroup (memory.max, cpu.max, pids.max), and that is done from outside, by whoever owns that part of the tree: systemd, or a container engine. As an ordinary user you can ask your own systemd instance for a scope with a memory limit; it creates a new cgroup and writes the limit.
$ systemd-run --user --scope -q -p MemoryMax=100M sh -c "cat /proc/self/cgroup; cat /sys/fs/cgroup\$(cut -d: -f3 /proc/self/cgroup)/memory.max" $ 0::/user.slice/user-1000.slice/user@1000.service/app.slice/run-p1230225-i1230525.scope $ 104857600
Now use both together, a cgroup with a limit and a CGROUP namespace. Inside, the process sees itself at /, and the root of its mounted cgroup2 carries the 100 MB limit that was set outside.
$ systemd-run --user --scope -q -p MemoryMax=100M unshare --user --map-root-user --cgroup --mount sh -c "cat /proc/self/cgroup; mount -t cgroup2 none /mnt; cat /mnt/memory.max" $ 0::/ $ 104857600
The shell you started from has no such limit:
$ cat /sys/fs/cgroup$(cut -d: -f3 /proc/self/cgroup)/memory.max $ max
2.6 Why containers use it
A container gets a cgroup of its own for its limits, and a CGROUP namespace rooted there. The process inside does not learn the host's cgroup path, which would give away how the host is organised and would have to be reproduced when a container is moved to another machine. With the cgroup filesystem mounted from inside, it also cannot reach the cgroup directories above its own, so it cannot touch the limits that confine it. And if the engine hands the subtree over (delegation), the container can create child cgroups in it and run its own systemd. A rootless podman container on this machine (a local Debian based image, shown as IMAGE) sees:
$ podman run --rm IMAGE cat /proc/self/cgroup $ 0::/
3 Check yourself
Answer these before you look at the answers:
- You start a program with
unshare --cgroup. Does it now have less memory available? No. A CGROUP namespace changes only what the program sees as the root of the cgroup tree. Limits come from the controller files of the cgroup it is in. - Inside the namespace
/proc/self/cgroupshows0::/. What does/proc/PID/cgroupshow for the same process on the host? The full path, unchanged. The process was not moved to another cgroup. - Inside the namespace,
/proc/1/cgroupof the host's init shows0::/../../../../../init.scope. Why the dots? That cgroup lies outside the subtree the namespace is rooted at, so its path is given relative to the namespace's root. - Who writes
memory.maxfor a container? The container engine or systemd, from outside, into the container's cgroup. The CGROUP namespace then keeps the container's view, and its mounted cgroup filesystem, inside that subtree.
Sources: