Namespaces 01 - Introduction
Linux Namespaces Learning · Next: MNT namespace
The Slovak original of this document: Namespaces 01 - Úvod (slovensky).
An introduction to Linux namespaces
1 The basic concepts of Linux namespaces
Namespaces in Linux are used to separate (isolate) running programs. The separation lets a program have a different view of the system, or of the system resources, than the other programs. Namespaces can be used for various purposes, but they are mainly used in the implementation of container technologies. The first namespaces (MNT namespace) were added to the kernel already in 2002. The newest addition (2016) is the namespace for control groups (CGROUP namespace). The following 7 namespaces are currently implemented in the Linux kernel.
+-----------+------------------+---------+-------+-----------------------+-------------------------------------+ | Namespace | Clone Flag | Kernel | Year | Required capabilities | Isolates | | | | | | | | +-----------+------------------+---------+-------+-----------------------+-------------------------------------+ | MNT | CLONE_NEWNS | 2.4.19 | 2002 | CAP_SYS_ADMIN | Mount points | | UTS | CLONE_NEWUTS | 2.6.19 | 2006 | CAP_SYS_ADMIN | Hostname and NIS domain name | | PID | CLONE_NEWPID | 2.6.24 | 2008 | CAP_SYS_ADMIN | Process IDs | | NET | CLONE_NEWNET | 2.6.29 | 2009 | CAP_SYS_ADMIN | Network devices, stacks, ports, ... | | IPC | CLONE_NEWIPC | 2.6.30 | 2009 | CAP_SYS_ADMIN | System V IPC, POSIX message queues | | USER | CLONE_NEWUSER | 3.8 | 2013 | Not required | User and group IDs | | CGROUP | CLONE_NEWCGROUP | 4.6 | 2016 | CAP_SYS_ADMIN | Cgroup root directory | +-----------+------------------+---------+-------+-----------------------+-------------------------------------+
The namespaces that have been considered but not yet implemented are: \- Security namespace \- Security keys namespace \- Device namespace \- Time namespace
REF: 2006 - Multiple Instances of the Global Linux Namespaces
1.1 What each namespace is and what it is for
- MNT namespace (mount points and filesystems)
Isolates the mount points a process, or a group of processes, can see, so programs in different MNT namespaces can have different views of the filesystem hierarchy. It is much like the isolation "chroot()" gives, except that an MNT namespace should be the safer and more flexible choice for the purpose.
- UTS namespace (UNIX time-sharing system)
Isolates two system identifiers: the node name (host name) and the domain name, both returned by the "uname()" system call.
- PID namespace (process identifier)
Isolates process identifiers. Programs in different PID namespaces can hold the same PID.
- NET namespace (network stack)
Isolates the network resources. Each network namespace has its own devices, addresses, routing tables, port numbers and its own "/proc/net" directory.
- IPC namespace (interprocess communication)
Isolates IPC resources: System V IPC objects and POSIX message queues. Each IPC namespace has its own set of System V IPC identifiers and its own set of POSIX message queues.
- USER namespace (users and groups)
Isolates user (UID) and group (GID) identifiers. It allows the UID and GID identifiers to be different inside the namespace and outside it. So one set of UIDs and GIDs can be used inside the namespace and another set is used outside it.
- CGROUP namespace (cgroup root directory)
Isolates the control groups a program belongs to. Each CGROUP namespace has its own cgroup root directory.
1.2 The namespace API
The namespace API is three system calls:
- clone()
Creates a new process in a new namespace. It is the more general form of the "fork()" system call, with the "flags" argument deciding what it does.
- "unshare()"
Creates a new namespace without creating a new process, moving the existing program into it (the exception is a PID namespace: the caller stays in its own, and only its children enter the new one). Its purpose is to let a program control what it shares without having to fork.
- "setns()"
Changes the namespace a program belongs to. In other words, this system call lets a program be dissociated from one instance of a given type of namespace (for example an MNT namespace) and re-associated with another instance of the same type, that is a "move" of the program from MNT namespace X to MNT namespace Y.
2 Working with Linux namespaces - the default namespaces
Check which namespaces the Linux INIT process (init or systemd) belongs to. These are the default namespaces. They are used by every process that has no specific namespaces assigned to it; see the SSHD server example below.
Note: the listings show six namespaces. The kernel of the test system did not have the CGROUP namespace yet; it was added in kernel 4.6.
# ls -lh /proc/1/ns/ ---------------------------------------------------------------------------------------------------------------- lrwxrwxrwx. 1 root root 0 Nov 8 09:54 ipc -> ipc:[4026531839] lrwxrwxrwx. 1 root root 0 Nov 8 09:54 mnt -> mnt:[4026531840] lrwxrwxrwx. 1 root root 0 Nov 8 09:54 net -> net:[4026531956] lrwxrwxrwx. 1 root root 0 Nov 8 09:54 pid -> pid:[4026531836] lrwxrwxrwx. 1 root root 0 Nov 8 09:54 user -> user:[4026531837] lrwxrwxrwx. 1 root root 0 Nov 8 09:54 uts -> uts:[4026531838]
Check which namespaces the Linux SSHD process ($SSHD_PID) belongs to. Find the PID of sshd and list the namespaces it belongs to.
# SSHD_PID=`ps xa | grep -m 1 /usr/sbin/sshd | awk '{print $1}'`; ls -lh /proc/$SSHD_PID/ns/
----------------------------------------------------------------------------------------------------------------
lrwxrwxrwx. 1 root root 0 Nov 8 10:04 ipc -> ipc:[4026531839]
lrwxrwxrwx. 1 root root 0 Nov 8 10:04 mnt -> mnt:[4026531840]
lrwxrwxrwx. 1 root root 0 Nov 8 10:04 net -> net:[4026531956]
lrwxrwxrwx. 1 root root 0 Nov 8 10:04 pid -> pid:[4026531836]
lrwxrwxrwx. 1 root root 0 Nov 8 10:04 user -> user:[4026531837]
lrwxrwxrwx. 1 root root 0 Nov 8 10:04 uts -> uts:[4026531838]3 Introducing the "namespaces-info.sh" script
While reading about namespaces I wrote a small tool for the BASH shell called "namespaces-info.sh". It prints the basic facts about the namespaces on a system. The script itself is on GitHub -> github - bash-namespaces-info
The syntax of "namespaces-info.sh" is:
-d List the system/default (parent) namespaces -a List every namespace of every process (PIDs) -n List the non-system/non-default namespaces of every process (PIDs) -p PID List the namespaces of one process (PID) -v Print the version of the script
Running "namespaces-info.sh" with "-d".
# /root/namespaces-info.sh -d --------------- + -------------------------- Linux Namespace | System default namespaces --------------- + -------------------------- IPC namespace | ipc:[4026531839] MNT namespace | mnt:[4026531840] NET namespace | net:[4026531956] PID namespace | pid:[4026531836] USER namespace | user:[4026531837] UTS namespace | uts:[4026531838] --------------- + --------------------------
Running "namespaces-info.sh" with "-a".
# /root/namespaces-info.sh -a ---------- + ---------- + -------------------- + -------- + ---------------------------------------- PID | PPID | NAMESPACE | DEFAULT | COMMAND ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 1 | 0 | ipc:[4026531839] | YES | /usr/lib/systemd/systemd --switched-root 1 | 0 | mnt:[4026531840] | YES | /usr/lib/systemd/systemd --switched-root 1 | 0 | net:[4026531956] | YES | /usr/lib/systemd/systemd --switched-root 1 | 0 | pid:[4026531836] | YES | /usr/lib/systemd/systemd --switched-root 1 | 0 | user:[4026531837] | YES | /usr/lib/systemd/systemd --switched-root 1 | 0 | uts:[4026531838] | YES | /usr/lib/systemd/systemd --switched-root ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 2 | 0 | ipc:[4026531839] | YES | [kthreadd] 2 | 0 | mnt:[4026531840] | YES | [kthreadd] 2 | 0 | net:[4026531956] | YES | [kthreadd] 2 | 0 | pid:[4026531836] | YES | [kthreadd] 2 | 0 | user:[4026531837] | YES | [kthreadd] 2 | 0 | uts:[4026531838] | YES | [kthreadd] ---------- + ---------- + -------------------- + -------- + ----------------------------------------
Running "namespaces-info.sh" with "-n".
# /root/namespaces-info.sh -n ---------- + ---------- + -------------------- + -------- + ---------------------------------------- PID | PPID | NAMESPACE | DEFAULT | COMMAND ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 18 | 2 | mnt:[4026531856] | NO | [kdevtmpfs] ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 582 | 1 | mnt:[4026532423] | NO | /usr/lib/systemd/systemd-udevd ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 718 | 1 | mnt:[4026532450] | NO | /usr/bin/vmtoolsd ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 750 | 1 | mnt:[4026532451] | NO | /usr/sbin/NetworkManager --no-daemon ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 9870 | 750 | mnt:[4026532451] | NO | /sbin/dhclient -d -q -sf /usr/libexec/nm ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 72966 | 9926 | mnt:[4026532578] | NO | bash 72966 | 9926 | net:[4026532452] | NO | bash ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 72975 | 9945 | mnt:[4026532579] | NO | bash 72975 | 9945 | net:[4026532510] | NO | bash ---------- + ---------- + -------------------- + -------- + ----------------------------------------
Running "namespaces-info.sh" with "-p".
# /root/namespaces-info.sh -p 72975 ---------- + ---------- + -------------------- + -------- + ---------------------------------------- PID | PPID | NAMESPACE | DEFAULT | COMMAND ---------- + ---------- + -------------------- + -------- + ---------------------------------------- 72975 | 9945 | ipc:[4026531839] | YES | bash 72975 | 9945 | mnt:[4026532579] | NO | bash 72975 | 9945 | net:[4026532510] | NO | bash 72975 | 9945 | pid:[4026531836] | YES | bash 72975 | 9945 | user:[4026531837] | YES | bash 72975 | 9945 | uts:[4026531838] | YES | bash ---------- + ---------- + -------------------- + -------- + ----------------------------------------
Current practice (checked 2026-10)
- Eight namespaces, not seven: the article lists the time namespace as "considered but not yet implemented". It exists since Linux 5.6 (
CLONE_NEWTIME) and gives a process its own offsets for the monotonic and boot-time clocks./proc/PID/ns/therefore shows more entries than in the listings above:cgroup,time,time_for_childrenandpid_for_childrenwere added. - Root is not needed: the examples run as root because the table says
CAP_SYS_ADMIN. That capability is only required inside the user namespace that owns the new namespace, and a user namespace can be created by any user. So the usual way today is to create a user namespace first and the others inside it, withunshare --user --map-root-userplus the flags for the other namespaces. - What "root" in a user namespace means: the process gets a full capability set, but only over resources that belong to its own namespaces. A file owned by the real UID 0 stays out of reach, because the kernel checks the UID as seen from the host.
- UID maps:
--map-root-usermaps exactly one ID (/proc/self/uid_mapshows0 1000 1); every other ID appears as 65534 (nobody). A whole range of IDs comes from/etc/subuidand/etc/subgidand is written by the setuid helpersnewuidmapandnewgidmap. Rootless podman works this way;podman unshare cat /proc/self/uid_mapshows the result. - Listing namespaces: the
namespaces-info.shscript and thels -l /proc/PID/ns/loops are replaced bylsnsfrom util-linux:lsnsfor all,lsns -p PIDfor one process,lsns -t netfor one type. Run it as root to see the namespaces of all users. - Entering a namespace from the shell: the article describes
setns()only as a system call. Its command-line form isnsenter, for examplensenter -t PID -n -uornsenter -t PID -afor all namespaces of the target process. - Cgroup namespace on cgroup v2: on a host with the unified hierarchy (
stat -fc %T /sys/fs/cgroupprintscgroup2fs)/proc/self/cgroupis a single line0::/path. Insideunshare --user --map-root-user --cgroupit reads0::/. - Namespaces alone are not a sandbox: a process in new namespaces still inherits the whole environment (tokens in variables) and all open file descriptors, and a network namespace is all or nothing. Container engines and sandbox tools add the other kernel pieces: a reduced capability set, a seccomp system call filter and, more recently, Landlock. Containers are built from the same namespaces described here.
bwrap(bubblewrap) is a small unprivileged tool that assembles them, but its upstream says clearly that it is a tool for building a sandbox, not a ready-made security policy.
As an ordinary user, five namespaces at once (the output shows UID 0, a process tree that starts at PID 1 and only a loopback interface):
$ unshare --user --map-root-user --uts --net --pid --fork --mount-proc sh -c 'id -u; ps ax; ip -br link'
Listing instead of the script:
$ lsns $ lsns -p $$ $ lsns -t net
Sources: