From mboxrd@z Thu Jan 1 00:00:00 1970 From: Mirian Shilakadze Date: Tue, 25 Aug 2026 17:29:48 +0400 Subject: [Devel] [PATCH vz10 1/2] fs/kernfs, ve: hide entries from a VE without invalidating the dentry In-Reply-To: <20260825133013.704924-1-mirian.shilakadze@virtuozzo.com> References: <20260825133013.704924-1-mirian.shilakadze@virtuozzo.com> Message-ID: <20260825133013.704924-2-mirian.shilakadze@virtuozzo.com> List-Id: kernfs_dop_revalidate() ends with a per VE visibility check and answers it with the same "return 0" that the staleness checks above it use. Those checks are properties of the kernfs node and hold for every observer: the node was deactivated, moved, renamed, or retagged. Visibility is a property of the calling task's VE, so one host dentry answers "valid" to a ve0 task and "stale" to a task inside a Container. The VFS reads 0 as a global fact and calls d_invalidate(), which walks the subtree and hands every mountpoint it finds to __detach_mounts(). The mountpoint hash is not scoped to a mount namespace, and m_list holds every mount attached at that dentry in any of them, so a Container's lookup unmounts the host's mounts. One lookup of /sys/fs/bpf from a task that only did setns() into a Container's ve namespace, staying in the host mount namespace, both hides the entry from the caller and destroys the host's bpffs. A Container start reaches the same path on its own: libvzctl stats every mount point in the namespace to collect the mount flags of a bindmount source, and does it after CLONE_NEWVE and before pivot_root, so the host loses bpffs and tracefs on the way. libvzctl needs bpffs for the cgroup v2 device controller, so no Container on the node can be managed afterwards, and the damage outlives the failed start. Return -ENOENT instead. The caller that cannot see the entry is told the name is missing, which is what the check is for, and the dentry stays valid for everyone else. No caller of ->d_revalidate() reaches d_invalidate() with a negative return: lookup_dcache(), lookup_fast(), __lookup_slow() and lookup_open() in fs/namei.c all gate it on exactly 0, ovl_revalidate_real() gates it the same way, and ecryptfs_d_revalidate() hands the value back without invalidating anything itself. kernfs_iop_lookup() already answers this same condition with a plain "not found". Feature: kernfs: per-CT entries visibility and permissions configuration https://virtuozzo.atlassian.net/browse/VSTOR-142552 Fixes: 3dd8c2499df6 ("ve/kernfs: hide forbidden entries in container") Signed-off-by: Mirian Shilakadze --- fs/kernfs/dir.c | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/fs/kernfs/dir.c b/fs/kernfs/dir.c index be680eb98ed4..7947c49ed1a6 100644 --- a/fs/kernfs/dir.c +++ b/fs/kernfs/dir.c @@ -1199,8 +1199,17 @@ static int kernfs_dop_revalidate(struct dentry *dentry, unsigned int flags) kernfs_info(dentry->d_sb)->ns != kn->ns) goto out_bad; - if (!kernfs_d_visible(kn, kernfs_info(dentry->d_sb))) - goto out_bad; + if (!kernfs_d_visible(kn, kernfs_info(dentry->d_sb))) { + /* + * The node is fine, it is only outside this VE's view. + * Returning 0 would tell the VFS that the dentry is stale, and + * it answers that with d_invalidate(), which detaches every + * mount on that dentry in every mount namespace. Report the + * name as missing to this caller instead. + */ + up_read(&root->kernfs_rwsem); + return -ENOENT; + } up_read(&root->kernfs_rwsem); return 1; -- 2.43.0