From: Pavel Tikhomirov <ptikhomirov@virtuozzo.com>
To: Mirian Shilakadze <mirian.shilakadze@virtuozzo.com>,
khorenko@virtuozzo.com
Cc: devel@openvz.org
Subject: Re: [Devel] [PATCH vz10 v2 0/2] fs/kernfs, ve: stop a container lookup unmounting host filesystems
Date: Wed, 26 Aug 2026 13:12:36 +0200 [thread overview]
Message-ID: <af71d8d8-c6e7-4477-87db-a2b08bd27625@virtuozzo.com> (raw)
In-Reply-To: <20260826110415.41119-1-mirian.shilakadze@virtuozzo.com>
Reviewed-by: Pavel Tikhomirov <ptikhomirov@virtuozzo.com>
On 8/26/26 13:04, Mirian Shilakadze wrote:
> Starting a container whose configuration carries a bindmount whose source is
> a mount with its own superblock unmounts the host's bpffs and tracefs.
> libvzctl needs bpffs for the cgroup v2 device controller, so once it is gone
> no container on the node can be managed. Every later vzctl command on any
> container, including ones that were already running and were never involved,
> prints "Unable to find mount point for bpf" twice and then reports a stale
> status. Recovery is a manual mount or a reboot. This is VSTOR-142552.
>
> The container start is not what does it. Any task whose VE is a container's,
> resolving a host path under /sys, unmounts what it finds there. setns() on a
> container's ve namespace, staying in the host mount namespace, is enough, and
> one stat() of /sys/fs/bpf both hides the entry from the caller and destroys
> the host's mount.
>
> kernfs_dop_revalidate() ends with a per VE visibility check and answers it
> with the same "return 0" that the staleness checks above it use. Those checks
> are properties of the kernfs node and hold for every observer: the node was
> deactivated, moved, renamed, or retagged. Visibility is a property of the
> calling task's VE, so one host dentry answers "valid" to a ve0 task and
> "stale" to a task inside a container. The VFS reads 0 as a global fact and
> calls d_invalidate(), which hands every mountpoint under that dentry to
> __detach_mounts(), whose mountpoint hash is not scoped to a mount namespace
> and whose m_list holds every mount attached at that dentry in any of them. A
> per VE answer therefore destroys a global object.
>
> Patch 1 reports the name as missing from that check, except on a kernfs
> instance the VE created, where the dentry is dropped as before. Everywhere
> else, the host's sysfs above all, the caller that cannot see the entry is
> told the name is missing, which is what the check is for, and the dentry
> stays valid for everyone else. What a container is told does not change
> either way: the errno for a hidden entry is ENOENT, because today it
> arrives after d_invalidate() and a fresh lookup that ends in a negative
> dentry. kernfs_iop_lookup() has always answered this same condition with a
> plain "not found".
>
> The one exception is a create attempt on a hidden name, and only on an
> instance the VE did not mount, where it now fails with ENOENT rather than
> the EACCES it fails with today. On the VE's own instance nothing changes, a
> create on a hidden name still fails with EACCES. Both fail either way.
>
> Patch 2 adds the regression test to the existing ve_perms selftest. It mounts
> a tmpfs on the entry the fixture already keeps host only, in its own mount
> namespace so the machine running it cannot lose a mount it needs, and
> requires that mount to still be there after a VE has looked the entry up.
>
> Introduced by 3dd8c2499df6 ("ve/kernfs: hide forbidden entries in container")
> in 2021 and reachable ever since. It went unreported because nothing in the
> management stack held a mount under /sys that anyone would miss, until
> libvzctl commit f946fae ("cgroup: switch from cgrou-v1 device controller to
> eBPF program") made it depend on bpffs.
>
> Testing
> =======
>
> Tested on a VHI 8.0.0 node with the same script, the same container and the
> same bindmount on both kernels.
>
> On stock 6.12.0-211.30.1.14.4.vz10 the start fails with rc=255 and "Cancel
> init execution", bpffs and tracefs are both gone afterwards, vzctl status on
> that container and on an unrelated one prints "Unable to find mount point for
> bpf" twice each, and vzctl exec stops working. Losing tracefs also took the
> kprobes the test itself was using.
>
> With patch 1 on 6.12.0-211.39.1.16.9.vz10 the same start returns rc=0, bpffs
> and tracefs are untouched, both status calls are clean, and the bindmount is
> present inside the container and read only as requested. The same holds on a
> debug build with KASAN and lockdep and on the shipping configuration.
>
> Under load, 48 processes inside a container's VE entered with setns(),
> alongside 48 in ve0, resolved /sys/fs/bpf and a tmpfs mounted on a hidden
> sysfs directory, 384000 hidden lookups in total. Every VE process saw ENOENT
> on every lookup and every ve0 process saw the entry on every lookup, with no
> mixed results. A kprobe on d_invalidate() named only the test's own cgroup
> dentries and the /proc pid directories of reaped children, never the hidden
> entries, and __detach_mounts() was never called. gcov on fs/kernfs/dir.c,
> fs/namei.c, fs/dcache.c and fs/namespace.c agrees: the new return ran 384000
> times, the staleness paths in kernfs_dop_revalidate() never ran, and
> __detach_mounts() was never entered.
>
> Granting a path to a VE through ve.sysfs_permissions still makes it visible
> and revoking it hides it again, and the host mount now survives the revoke,
> which it did not before.
>
> ve_perms_test passes 16 of 16 and ve_ns_owner_test 2 of 2, together with the
> filesystems, mount, mount_setattr, move_mount_set_group, nsfs and proc
> selftests. Patch 2 fails on the unpatched kernel with "the VE lookup
> unmounted /sys/power" and passes with patch 1 applied.
>
> The condition added in v2 was checked on both sides. During a container
> start it never fires: of the 13 lookups that answered 0, every one returned
> before reaching the visibility check, from the negative dentry branch or
> from !kernfs_active(), which are the device mapper and uevent nodes churning
> as the disk is set up. The lookups that do reach the check answer ENOENT, 8
> of them, and __detach_mounts() is not called at all. It fires where it is
> meant to: a container looking up a hidden entry in its own sysfs instance
> gets the dentry dropped and the mount on it detached, while the same lookup
> against the host's sysfs answers ENOENT and leaves the mount alone.
>
> Two of the six ->d_revalidate call sites, __lookup_slow() and lookup_open(),
> were not reached at runtime. This tree carries lookup_fast_for_open(), so even
> an O_CREAT open resolves the last component through lookup_fast(), which
> leaves those two reachable only through a dcache race. Both gate
> d_invalidate() on exactly 0, as do the sites that were exercised,
> ovl_revalidate_real() and ecryptfs_d_revalidate().
>
> v2:
> - patch 1: keep the old invalidate on a kernfs instance the VE created,
> and only report the name as missing on any other instance (Pavel)
> - patch 2: detect the mount with openat2(RESOLVE_NO_XDEV) rather than
> reading /proc/self/mountinfo (Pavel)
> - dropped Pavel's Reviewed-by from v1, both patches changed
>
> Mirian Shilakadze (2):
> fs/kernfs, ve: hide entries from a VE without invalidating the dentry
> selftests/ve: check that hiding an entry does not unmount it
>
> fs/kernfs/dir.c | 17 ++++++-
> tools/testing/selftests/ve/ve_perms_test.c | 52 ++++++++++++++++++++++
> tools/testing/selftests/ve/ve_selftest.h | 40 ++++++++++++++++-
> 3 files changed, 105 insertions(+), 4 deletions(-)
>
--
Best regards, Pavel Tikhomirov
Senior Software Developer, Virtuozzo.
_______________________________________________
Devel mailing list
Devel@openvz.org
https://lists.openvz.org/mailman/listinfo/devel
prev parent reply other threads:[~2026-08-26 11:14 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-26 11:04 Mirian Shilakadze
2026-08-26 11:04 ` [Devel] [PATCH vz10 v2 1/2] fs/kernfs, ve: hide entries from a VE without invalidating the dentry Mirian Shilakadze
2026-08-26 16:15 ` [Devel] [PATCH RHEL10 COMMIT] " Konstantin Khorenko
2026-08-26 16:34 ` [Devel] [PATCH vz10 v2 1/2] " Konstantin Khorenko
2026-08-26 11:04 ` [Devel] [PATCH vz10 v2 2/2] selftests/ve: check that hiding an entry does not unmount it Mirian Shilakadze
2026-08-26 16:15 ` [Devel] [PATCH RHEL10 COMMIT] " Konstantin Khorenko
2026-08-26 11:12 ` Pavel Tikhomirov [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=af71d8d8-c6e7-4477-87db-a2b08bd27625@virtuozzo.com \
--to=ptikhomirov@virtuozzo.com \
--cc=devel@openvz.org \
--cc=khorenko@virtuozzo.com \
--cc=mirian.shilakadze@virtuozzo.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.