LINUXOR.SK ... open source notes ...

Namespaces 15 - From namespaces to a sandbox

category: learnz/namespaces · date: 2026-10-04 · author: LALA · theme: github

Linux Namespaces Learning · Previous: TIME namespace

From namespaces to a sandbox

Introduction

A namespace changes what a process can see: other process IDs, other mounts, other network devices, other user IDs. It does not by itself decide what the process is allowed to do with the things it still sees. A container or a sandbox is what you get when you combine namespaces with a few more kernel mechanisms and with a deliberate choice of what goes in. This article follows the line of the LinuxDays 2026 talk by Michal Vyskočil on modern process isolation (see Sources): first the other building blocks, then a small sandbox built with bwrap, then the things that still get into it. Everything was run as an ordinary user on a 6.12 kernel with bubblewrap 0.10, passt/pasta and podman 5.8.

The other building blocks

Capabilities split the power of root into separate permissions, such as CAP_NET_ADMIN for network configuration or CAP_SYS_CHROOT for chroot(). A process can hold some and lack the rest, so "root" in a container is usually a root with a short list. CAP_SYS_ADMIN covers so many operations that holding it is close to being root again.

Seccomp filters system calls. A process installs a filter, and from then on the kernel checks every system call of that process and its children against it and can refuse the call or kill the process. A filter cannot be removed once installed.

Landlock lets an unprivileged process restrict itself: it declares which parts of the filesystem (and, on newer kernels, which TCP ports) it still needs, and the kernel denies everything else to it and to its children. It needs no root and no system-wide policy.

The state of the first two is visible in /proc/PID/status. An ordinary shell holds no capabilities, has the full bounding set available, and runs without a seccomp filter:

bash
$ grep -E "^(CapEff|CapBnd|Seccomp|NoNewPrivs):" /proc/self/status
$ CapEff:	0000000000000000
$ CapBnd:	000001ffffffffff
$ NoNewPrivs:	0
$ Seccomp:	0

User-mode networking with pasta

A new network namespace is empty: a loopback interface that is down, and no way out.

bash
$ unshare --user --net ip -br addr
$ lo               DOWN

The classic way out is a veth pair and a bridge on the host, which needs root there. pasta does it without: it is one ordinary process that creates the user and network namespace, puts a tap interface into it, and translates between that interface and normal sockets which it opens on the host. There is no bridge, no veth and no firewall rule. By default the namespace gets a copy of the host's interface name, address and routes.

bash
$ pasta --quiet --config-net -- ip -4 -br addr
$ lo               UNKNOWN        127.0.0.1/8
$ eth0             UNKNOWN        192.0.2.10/24
$ pasta --quiet --config-net -- ip -4 route
$ default via 192.0.2.1 dev eth0 proto dhcp metric 600
$ 192.0.2.0/24 dev eth0 proto kernel scope link metric 600

The eth0 inside is a tap device. pasta gives the namespace connectivity that behaves like the host and filters nothing: the namespace can reach whatever the host can reach, and, as shown below, the host's own loopback services as well.

A minimal sandbox with bwrap

bwrap (bubblewrap) is a small unprivileged tool that puts the pieces together. It creates the namespaces, mounts an empty tmpfs as the new root, binds into it only what you name, and then switches the root so that the old tree is no longer reachable. The following command gives the program a read-only /usr and /etc, a fresh /proc, /dev and /tmp, and nothing else. It is kept in a variable because the rest of the article reuses it.

bash
$ BW='bwrap --ro-bind /usr /usr --ro-bind /etc /etc --symlink usr/bin /bin --symlink usr/lib64 /lib64 --proc /proc --dev /dev --tmpfs /tmp --unshare-all --die-with-parent'
$ $BW sh -c 'ls /; id -u; ps ax; ip -br link; ls /home; hostname'
$ bin
$ dev
$ etc
$ lib64
$ proc
$ tmp
$ usr
$ 1000
$     PID TTY      STAT   TIME COMMAND
$       1 ?        S      0:00 bwrap --ro-bind /usr /usr --ro-bind /etc /etc ...
$       2 ?        S      0:00 sh -c ls /; id -u; ps ax; ip -br link; ls /home; hostname
$       5 ?        R      0:00 ps ax
$ lo               UNKNOWN        00:00:00:00:00:00 <LOOPBACK,UP,LOWER_UP>
$ ls: cannot access '/home': No such file or directory
$ host

--unshare-all asks for new user, IPC, PID, network, UTS and cgroup namespaces; the mount namespace is always new. The process tree starts at PID 1, the only network interface is the loopback, and /home does not exist. Two things are still open: the root is a writable tmpfs, and no seccomp filter is installed, because nobody asked for one.

bash
$ $BW sh -c 'findmnt -n -o TARGET,FSTYPE,OPTIONS / ; touch /x && ls /x; grep -E "^(CapEff|Seccomp):" /proc/self/status'
$ / tmpfs rw,nosuid,nodev,relatime,uid=1000,gid=1000,inode64
$ /x
$ CapEff:	0000000000000000
$ Seccomp:	0

What still gets in

The whole environment

Namespaces isolate kernel objects. The environment is not one of them; it is copied from parent to child like in any other program start. A token kept in a variable is passed into the sandbox with it.

bash
$ export DEMO_TOKEN=demo-value-not-real
$ env | wc -l
$ 92
$ $BW sh -c 'env | wc -l'
$ 91
$ $BW sh -c 'echo $DEMO_TOKEN'
$ demo-value-not-real

Restricting files does not cover this. With bwrap you have to clear the environment and pass in what the program needs by name:

bash
$ $BW --clearenv --setenv PATH /usr/bin sh -c 'env; echo $DEMO_TOKEN'
$ PWD=/
$ SHLVL=1
$ PATH=/usr/bin
$ _=/usr/bin/env
$ 

Open file descriptors

A file that is not bound into the sandbox cannot be opened there by its path. A file that was already open when the sandbox started is a different matter: the descriptor is inherited, and it keeps working although the path does not exist inside.

bash
$ $BW sh -c 'cat /home/user/ns-demo/credentials'
$ cat: /home/user/ns-demo/credentials: No such file or directory
$ exec 9< /home/user/ns-demo/credentials
$ $BW sh -c 'cat /proc/self/fd/9'
$ demo_key = DEMO-ONLY-NOT-A-REAL-KEY
$ $BW sh -c 'readlink /proc/self/fd/9'
$ /home/user/ns-demo/credentials

After closing the descriptor in the calling shell, it is gone inside too:

bash
$ exec 9<&-
$ $BW sh -c 'cat /proc/self/fd/9'
$ cat: /proc/self/fd/9: No such file or directory

A program that starts a sandbox has to close everything it does not mean to hand over.

The network is all or nothing

For the next examples a web server listens on the loopback address of the host only, the usual way of running a service that is "just local" and therefore has no authentication. It was started in a second terminal, in a directory that holds one file named secret:

bash
$ python3 -m http.server 8000 --bind 127.0.0.1

From the host itself the page is there:

bash
$ curl -s 127.0.0.1:8000/secret
$ DEMO-ONLY-LOOPBACK-PAGE

With its own network namespace the sandbox reaches nothing at all, which is safe and often useless (curl ends with exit code 7, "could not connect", and prints nothing):

bash
$ $BW curl -s --max-time 2 127.0.0.1:8000/secret

The only other choice a namespace offers is to share the network of the host, and then the sandbox reaches everything the host does, the local-only service included:

bash
$ $BW --share-net curl -s --max-time 2 127.0.0.1:8000/secret
$ DEMO-ONLY-LOOPBACK-PAGE

A network namespace knows no rule like "this one host and port, nothing else". That needs something in front of it: a firewall in the namespace, a filtering proxy, or Landlock rules for TCP ports.

Host loopback through pasta

Giving the sandbox a network of its own through pasta looks like the middle way. By default it is not, because pasta forwards the ports that listen on the host into the namespace, loopback included:

bash
$ pasta --quiet --config-net -- curl -s --max-time 2 127.0.0.1:8000/secret
$ DEMO-ONLY-LOOPBACK-PAGE
$ pasta --quiet --config-net -- $BW --share-net curl -s --max-time 2 127.0.0.1:8000/secret
$ DEMO-ONLY-LOOPBACK-PAGE

--map-host-loopback none does not change this on its own:

bash
$ pasta --quiet --config-net --map-host-loopback none -- curl -s --max-time 2 127.0.0.1:8000/secret
$ DEMO-ONLY-LOOPBACK-PAGE

Turning off the automatic TCP forwarding with -T none closes 127.0.0.1 (the first command prints nothing). The service is then still reachable under the address of the default gateway, which pasta maps to the host:

bash
$ pasta --quiet --config-net -T none -- curl -s --max-time 2 127.0.0.1:8000/secret
$ pasta --quiet --config-net -T none -- curl -s --max-time 2 192.0.2.1:8000/secret
$ DEMO-ONLY-LOOPBACK-PAGE

Only both options together close it on this machine (the command prints nothing):

bash
$ pasta --quiet --config-net -T none --map-host-loopback none -- curl -s --max-time 2 192.0.2.1:8000/secret

pasta is built to make a container behave like the host. A sandbox wants the opposite, so every default has to be checked, and the result tested the way it was tested here. UDP has the same pair of defaults (-U).

Containers use the same mechanisms

A container is made of exactly these parts. Rootless podman on this machine reports that it runs without root, uses pasta for the network, and sits on cgroup v2:

bash
$ podman info --format '{{.Host.Security.Rootless}} {{.Host.NetworkBackend}} {{.Host.RootlessNetworkCmd}} {{.Host.CgroupsVersion}} {{.Host.OCIRuntime.Name}}'
$ true netavark pasta v2 crun

Inside a container (here a local Debian based image, shown as IMAGE) you find what the previous sections built by hand: PID 1, root through a UID map with the subordinate range, a reduced capability set, a seccomp filter (mode 2), the tap interface from pasta. The environment of the calling shell is not passed in.

bash
$ podman run --rm IMAGE sh -c 'echo PID=$$; id -u; cat /proc/self/uid_map; grep -E "^(CapEff|Seccomp):" /proc/self/status; ls /sys/class/net; echo token=${DEMO_TOKEN:-unset}'
$ PID=1
$ 0
$          0       1000          1
$          1     524288      65536
$ CapEff:	00000000800405fb
$ Seccomp:	2
$ lo
$ eth0
$ token=unset

The capability mask, decoded, is the short list mentioned at the beginning. CAP_SYS_ADMIN and CAP_NET_ADMIN are not in it:

bash
$ capsh --decode=00000000800405fb
$ 0x00000000800405fb=cap_chown,cap_dac_override,cap_fowner,cap_fsetid,cap_kill,cap_setgid,cap_setuid,cap_setpcap,cap_net_bind_service,cap_sys_chroot,cap_setfcap

The seccomp filter comes from a profile that denies by default and lists the allowed system calls:

bash
$ grep -m1 defaultAction /usr/share/containers/seccomp.json
$ 	"defaultAction": "SCMP_ACT_ERRNO",

Podman also does not run pasta with pasta's own defaults. The process list shows the options it adds, and they are the ones found by hand in the previous section: no automatic port forwarding in either direction and no gateway mapping. With the loopback web server still running, a default container could not reach it on 127.0.0.1 or on the gateway address.

bash
$ pgrep -a pasta
$ 1215534 /usr/bin/pasta --config-net --dns-forward 169.254.1.1 -t none -u none -T none -U none --no-map-gw --quiet --netns ... --map-guest-addr 169.254.1.2

A container and a sandbox use the same kernel mechanisms and the same helper programs. They differ in the options somebody chose. When you build a sandbox yourself, you choose them, and each one has to be tested.

Check yourself

Answer these before you look at the answers:

Sources:

← learnz/namespaces