]> git.ipfire.org Git - thirdparty/kernel/linux.git/commitdiff
binfmt_misc: don't let an 'F' entry pin its own instance
authorChristian Brauner <brauner@kernel.org>
Tue, 28 Jul 2026 12:26:32 +0000 (14:26 +0200)
committerChristian Brauner <brauner@kernel.org>
Tue, 28 Jul 2026 13:42:32 +0000 (15:42 +0200)
An entry registered with 'F' opens its interpreter at registration time
and holds that file until the entry is freed. Any entry nobody removes
by hand only gets closed once the binfmt_misc superblock is shut down.
If the interpreter lives on a mount that keeps that superblock alive the
two pin each other:

    binfmt_misc sb -> inode -> entry -> interp_file -> vfsmount -> binfmt_misc sb

TL;DR the file is never closed. Once the mount namespace is gone there
is nothing left to unregister through either.

There are two ways to trigger this bug:

- Point the interpreter at the instance itself. Its files are regular
  files owned by the mounter and both bm_get_inode() and
  simple_fill_super() leave i_op at empty_iops. So notify_change() falls
  back to simple_setattr() and chmod +x works. We never set SB_I_NOEXEC
  and so open_exec() accepts it.

- Use the instance as an overlayfs lower layer. The overlay superblock
  holds a clone_private_mount() of every layer until it is destroyed and
  that clone is in no namespace. So umount_tree() never reaches it.

That's a DoS. And it isn't only the superblock that leaks. It pins the
user namespace it was mounted in, so every iteration permanently eats
one of the caller's user namespace charges.

So let's just do the sane thing. SB_I_NOEXEC makes open_exec() fail on
the instance's own files and s_stack_depth makes overlayfs reject the
layer before it ever takes a clone. That also covers the ecryptfs and
fuse passthrough variants. What 'F' promises is unchanged.

The stable tag is narrower than the Fixes tags on purpose. Before
sandboxed mounts this needed global root against the single instance
everyone shares, and the change doesn't apply to those trees anyway.

Note that SB_I_NODEV is implicitly raised for userns mounts but raise it
explicitly here as well.

Link: https://patch.msgid.link/20260728-work-binfmt_misc-selfpin-v1-1-74df5daeca5b@kernel.org
Fixes: 948b701a607f ("binfmt_misc: add persistent opened binary handler for containers")
Fixes: 21ca59b365c0 ("binfmt_misc: enable sandboxed mounts")
Cc: stable@vger.kernel.org # v6.7+
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
fs/binfmt_misc.c

index 5de615ca7a75f1dcfd620f29462feb905f0e12d9..47aeb2b68d3e5eb534552546cf1f50749f21b3e1 100644 (file)
@@ -937,6 +937,10 @@ static int bm_fill_super(struct super_block *sb, struct fs_context *fc)
        if (WARN_ON(user_ns != current_user_ns()))
                return -EINVAL;
 
+       /* Never exec off this instance and never let anything stack on it. */
+       sb->s_iflags |= SB_I_NOEXEC | SB_I_NODEV;
+       sb->s_stack_depth = FILESYSTEM_MAX_STACK_DEPTH;
+
        /*
         * Lazily allocate a new binfmt_misc instance for this namespace, i.e.
         * do it here during the first mount of binfmt_misc. We don't need to