Expand description
The seccomp profile a window’s container runs under on docker, and how it reaches the engine.
Podman needs none of this. Rootless podman already lets a process inside a container open an unprivileged user namespace of its own, so a browser-engine application built its own sandbox under podman’s default profile untouched: its renderer processes came up in their own user, PID and network namespaces, each with a filter of its own on top of the container’s.
Docker’s default profile allows unshare and setns only to CAP_SYS_ADMIN and masks the
namespace flags out of clone. Under it the application cannot build the sandbox at all: it
prints that moving to a new namespace was not permitted and exits with 133. The two ways out
are giving the container no filter at all, which loses every other syscall rule with it, or
giving it the default profile plus those three calls. QCode carries the second, in
assets/seccomp/desktop.json, and hands it to docker by path. What it costs is stated in that
file’s own comment: a process in the container can make a user namespace, so the host kernel’s
unprivileged-user-namespace surface is open to it — which is what podman gives by default
anyway.
--no-sandbox is not an option here, and the record’s flags are tested for its absence. It
would let a page or an extension that takes over a renderer reach everything in the container
with the person’s own rights: the workspace’s files, the profile’s home volume, the login.
The file is carried in the binary and written out when it is first needed, named after a digest of its own content, so a QCode that has been updated never hands the engine an older file left in the temporary folder.
Constants§
- PROFILE
- The profile, as it is carried: docker’s own default with
clone,setnsandunshareallowed.
Functions§
- file
- Writes the profile where the engine can read it and answers its path.