From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail.openvz.org (unknown [69.168.225.77]) by lore.virtuozzo.com (Postfix) with ESMTPS id DB78680024 for ; Wed, 26 Aug 2026 15:40:01 +0000 (UTC) Received: from mail.openvz.org (localhost [127.0.0.1]) by mail.openvz.org (8.14.4/8.14.4) with ESMTP id 67QFcl5t010150; Wed, 26 Aug 2026 18:38:48 +0300 DKIM-Filter: OpenDKIM Filter v2.11.0 mail.openvz.org 67QFcl5t010150 Authentication-Results: mail.openvz.org; dkim=fail reason="signature verification failed" (2048-bit key) header.d=virtuozzo.com header.i=@virtuozzo.com header.b="yB42WQlP" Received: from mail-ed1-f71.google.com (mail-ed1-f71.google.com [209.85.208.71]) by mail.openvz.org (8.14.4/8.14.4) with ESMTP id 67QFcjEj010142 (version=TLSv1/SSLv3 cipher=AES128-GCM-SHA256 bits=128 verify=FAIL) for ; Wed, 26 Aug 2026 18:38:45 +0300 DKIM-Filter: OpenDKIM Filter v2.11.0 mail.openvz.org 67QFcjEj010142 Received: by mail-ed1-f71.google.com with SMTP id 4fb4d7f45d1cf-6a5d57a8d51so1237386a12.1 for ; Wed, 26 Aug 2026 08:38:45 -0700 (PDT) X-Gm-Message-State: AFuF++lhPLMXi1LcXREkF2r97mVd7DFC7LQnMr3gYTF0cm5Qf9+1/wW4 nZKa+z/evkt6RtIEi50hCAqEfALNi6lwty+zOsbkMhh9GuozM031tIevh6iufDDzE+U/lxeKt31 0RWE2jTu5YjdH07KMf0EYFJwT9WvUEyUFrFM2E+lJj18f5pG4azTSFw== X-Gm-Gg: AR+sD11WQE0J9O4cwzE6KIOOzowJ6ZsCcDDQCtrEBW9vthhJvD5lJM6ITeqhxLtl7kz 0bDzhz+jWaoF4q0PsxX68iycub738LO8cXwhsmSEj0LVHLldJHGpQTAyFnoyHMwcMhQRhoftIr4 RyTfA65IK0jAJRp2Dp6tt5oyryoNP88EBqqJEK5eyWj+XcXdJ3Bv4o8+X0hHusgSUlM/znGrYDe zil/2UYXqWxKJZo5dthvGK2zmBOv5rbyUVXGT863nrrVtlPzeHGUydjfHnL5L+CRikpnpSiyrrb BSlVbKI+S2YA8aq8rCod7z4IB44PP+/3r9au1tw5gBtgXEJUh/02wXFEfhWjbpFUDPPR2/xtvzp 1+DLrkclz9bV41UrS X-Received: by 2002:a05:6402:2711:b0:6a1:284f:7891 with SMTP id 4fb4d7f45d1cf-6a5df5d92e6mr9866486a12.7.1787758724883; Wed, 26 Aug 2026 08:38:44 -0700 (PDT) X-Received: by 2002:a05:6402:2711:b0:6a1:284f:7891 with SMTP id 4fb4d7f45d1cf-6a5df5d92e6mr9866414a12.7.1787758724360; Wed, 26 Aug 2026 08:38:44 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1787758724; cv=none; d=google.com; s=arc-20260327; b=H3vGIlikNI1k+rlYS9DJ6tIzjVgbzYzl10rSFUEur8i2DgSGMz1GDwN8V++Dg4GLmw u80jyN4EE9Yfh8nPuSLIItt18ExYzZfEO2Ajgdr6KaV/j0DfxWQYV/jXnqL1nmHHqDxg 5PHA7ZOuTSzSoJYY3zA8vMbysHDTyqhZKmqO3ojKcg0cyUjnLtt3pyDd0Z5m2Bb/QWp5 N4z0lj6L0VBfzS74l6rxYr6/wibDqGW6IqXVUC2bAo5fPkYv7z9PnIX5RkZt/6SFNBjg UreN9srzp3YILlmdCU9NxCqHEV39LuMm+HkVFm5aa13BUldP2XgYNi+GFfTFgtLmM45X v1Iw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=subject:in-reply-to:cc:to:from:message-id:date:dkim-signature; bh=oUswbPMWBTQV7VhLuN3xiwaW5ulMhCB/tmnlAlVM9FI=; fh=6WaLqqjLrnoBYT6o6L3rXzHBtCCDnrtj0IcE19DsjNk=; b=WKsEj82UP2zCJ/3Tyg3ZrchGma4PWhqe7ZGqkX+YVtfI9i3SlykgsHrZ53+dl7o8Pr r6exzvZUue6/Ss3jIYqQdm1+URFxRrxq2JlJU/MX+EaAk1pijPFLqMc95DeV2kfLvm37 DLfONsEPSUSh4Y0a9YEiiolhwZd9PciCRvbCj7jn7YQXXfj2vMc1Bw9FpHP7AQzmb22R LN1xVfoLT3XSIme4ZCUG9adPfy6YLSjMo7/oJocvogf16AUbrVm2OFarLLSVSF0XonyV q9yoZ6OvlKBDK1wb6ePEFbkafo+s8C0qmVYFhrtxJktFVPQsteOACccwVFuC8YxDPvL8 NBpg==; dara=google.com ARC-Authentication-Results: i=1; mx.google.com; dkim=pass header.i=@virtuozzo.com header.s=relay header.b=yB42WQlP; spf=pass (google.com: domain of khorenko@virtuozzo.com designates 130.117.225.111 as permitted sender) smtp.mailfrom=khorenko@virtuozzo.com; dmarc=pass (p=QUARANTINE sp=QUARANTINE dis=NONE) header.from=virtuozzo.com Received: from relay.virtuozzo.com (relay.virtuozzo.com. [130.117.225.111]) by mx.google.com with ESMTPS id 4fb4d7f45d1cf-6a5deaa1011si4389794a12.183.2026.08.26.08.38.44 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 08:38:44 -0700 (PDT) Received-SPF: pass (google.com: domain of khorenko@virtuozzo.com designates 130.117.225.111 as permitted sender) client-ip=130.117.225.111; Authentication-Results: mx.google.com; dkim=pass header.i=@virtuozzo.com header.s=relay header.b=yB42WQlP; spf=pass (google.com: domain of khorenko@virtuozzo.com designates 130.117.225.111 as permitted sender) smtp.mailfrom=khorenko@virtuozzo.com; dmarc=pass (p=QUARANTINE sp=QUARANTINE dis=NONE) header.from=virtuozzo.com DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=virtuozzo.com; s=relay; h=Subject:From:Message-Id:Date:Content-Type: MIME-Version; bh=oUswbPMWBTQV7VhLuN3xiwaW5ulMhCB/tmnlAlVM9FI=; b=yB42WQlPLRW7 7qvfC/YSbasab/pRc7zSWvNPTSukLMrfiBwoyr7rEh0DCWP7YOmoZDUlcbUVogtEEWjNUHuZPVxoL j/9vAoRztXIDPRWYHDUACzKe6zpvKvHsgXeciw8eBXe8jgAVrQ8/eIi5CyQqdAUs/+9MHljxV2KhW D/OpbevW2tjOi5iT9cpysVciEzQ720OxLDL4MP4uh/nUoK0CxOAoxz8hhsD0nfEyFF5NnplbJM19M EWJ94Dd8nXqcSeTsCnyTlukZmKByB4rHordGY+MmBD5IO4cJPOPjx+1iEMsNMOpdDolmIKd5isXFH fUCXXpwHqX4ADDnZC1ipmg==; Received: from ch-demo-asa.virtuozzo.com ([130.117.225.8] helo=f0.sw.ru) by relay.virtuozzo.com with esmtps (TLS1.3) tls TLS_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzFft-004iV2-17; Wed, 26 Aug 2026 17:38:43 +0200 Received: from f0.sw.ru (localhost [127.0.0.1]) by f0.sw.ru (8.18.1/8.18.1/Debian-2) with ESMTP id 67QFchJq907545; Wed, 26 Aug 2026 17:38:43 +0200 Received: (from kostja@localhost) by f0.sw.ru (8.18.1/8.18.1/Submit) id 67QFchDO907544; Wed, 26 Aug 2026 17:38:43 +0200 Date: Wed, 26 Aug 2026 17:38:43 +0200 Message-Id: <202608261538.67QFchDO907544@f0.sw.ru> X-Authentication-Warning: f0.sw.ru: kostja set sender to khorenko@virtuozzo.com using -f From: Konstantin Khorenko To: Mirian Shilakadze In-Reply-to: <4fdbc2ee64728e97c0e44cfb07049428166406f6.1786950779.git.mirian.shilakadze@virtuozzo.com> X-OZ-Fwd: true Cc: OpenVZ devel Subject: Re: [Devel] [PATCH RHEL10 COMMIT] ve/fs: transfer mount ownership when a mount enters another VE X-BeenThere: devel@openvz.org X-Mailman-Version: 2.1.12 Precedence: list List-Id: OpenVZ development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , MIME-Version: 1.0 Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit Sender: devel-bounces@openvz.org Errors-To: devel-bounces@openvz.org The commit is pushed to "branch-rh10-6.12.0-211.39.1.16.x.vz10-ovz" and will appear at git@bitbucket.org:openvz/vzkernel.git after rh10-6.12.0-211.39.1.16.10.vz10 ------> commit fce5f612701953f243851cc79415c99e2f995ce6 Author: Mirian Shilakadze Date: Mon Aug 17 11:16:40 2026 +0400 ve/fs: transfer mount ownership when a mount enters another VE ve_owner is assigned once in ve_mount_nr_inc() from alloc_vfsmnt() and nothing updates it afterwards, so a mount the host creates and moves into a container keeps ve_owner == ve0 while it lives in the container's mount namespace. ve_check_trusted_file() reads that field for filesystems with no s_bdev and lets ve0 execute from any mount it considers host owned, so a host tmpfs bindmounted into a container is trusted even though the container can write to it. A tmpfs the container creates itself is refused, the same binary planted by the same container on a mount the host lent it runs: vzctl set 971 --bindmount_add /root/tex_tmpfs:/mnt/bm_tmpfs --save vzctl exec 971 'cp /bin/echo /mnt/bm_tmpfs/planted && chmod 755 /mnt/bm_tmpfs/planted' nsenter -t $INITPID -m /mnt/bm_tmpfs/planted HOST-TMPFS-LENT-TO-CT HOST-TMPFS-LENT-TO-CT Fix it where the mount changes hands. Every path that puts an existing mount into another VE's namespace goes through commit_tree(), both the detached tree the container moves in with move_mount() and the copies the propagation loop in attach_recursive_mnt() commits into a foreign namespace, so adopt the namespace's owner there. Ownership describes where the mount is, so it follows the namespace and the direction of the move is not special cased. That covers every mount that moves. Mounts created with the wrong owner to begin with, which copy_mnt_ns() and open_detached_copy() both do for a ve0 task working inside a container, never pass through commit_tree() and are fixed by the next patch. The same field drives per-VE mount accounting, which was wrong in the same direction: a moved in mount was charged to ve0 rather than to the container holding it. A ve0 process that execs or mmaps from a mount this reowns is now refused, exec with -EACCES and mmap with -EBADF, plus a SIGSEGV for the first few attempts. That is the point of the change, but it is visible to host tooling that reaches into a container's mounts to run something. The transfer does not consult sysctl_ve_mount_nr, commit_tree() cannot fail. A container can therefore be pushed above its mount limit by mounts the host gives it, and while over it the container's own mounts are refused until the count drops. The default limit is 4096 so this takes an unusual number of lent mounts, but it is the host's action that spends the container's budget. The owner change is done under the vfsmount lock, which serializes it against is_sb_ve_accessible() walking sb->s_mounts. ve_check_trusted_file() reads ve_owner without that lock, so both the store and the load are marked. It only compares the pointer against ve0 and the field is never NULL in between, so the reader sees either the old or the new owner. The reference on the old VE is dropped only after the new one is taken. Fixes: d65efacf542b ("trusted/ve/fs/exec: Don't allow a privileged user to execute untrusted files") Fixes: fc7157b84c32 ("trusted/ve/mmap: Protect from unsecure library load from CT image") https://virtuozzo.atlassian.net/browse/VSTOR-141322 Feature: ve: ve generic structures Signed-off-by: Mirian Shilakadze Reviewed-by: Vasileios Almpanis Reviewed-by: Konstantin Khorenko --- fs/mount.h | 2 +- fs/namespace.c | 27 +++++++++++++++++++++++++++ kernel/ve/ve.c | 6 +++++- 3 files changed, 33 insertions(+), 2 deletions(-) diff --git a/fs/mount.h b/fs/mount.h index 5cf06431d5868..1d41fd15265e6 100644 --- a/fs/mount.h +++ b/fs/mount.h @@ -71,7 +71,7 @@ struct mount { }; struct list_head mnt_umounting; /* list entry for umount propagation */ #ifdef CONFIG_VE - struct ve_struct *ve_owner; /* VE in which this mount was created */ + struct ve_struct *ve_owner; /* VE whose mount namespace holds it */ #endif /* CONFIG_VE */ #ifdef CONFIG_FSNOTIFY struct fsnotify_mark_connector __rcu *mnt_fsnotify_marks; diff --git a/fs/namespace.c b/fs/namespace.c index 7e27537dcdaf9..f8319a2b33df8 100644 --- a/fs/namespace.c +++ b/fs/namespace.c @@ -320,6 +320,7 @@ int mnt_get_count(struct mount *mnt) static inline int ve_mount_allowed(void); static inline void ve_mount_nr_inc(struct mount *mnt, struct ve_struct *ve); static inline void ve_mount_nr_dec(struct mount *mnt); +static inline void ve_mount_reown(struct mount *mnt, struct ve_struct *ve); static struct mount *alloc_vfsmnt(const char *name, struct ve_struct *owner_ve) { @@ -1182,6 +1183,7 @@ static void commit_tree(struct mount *mnt) m = list_first_entry(&head, typeof(*m), mnt_list); list_del(&m->mnt_list); + ve_mount_reown(m, n->ve_owner); mnt_add_to_ns(n, m); } n->nr_mounts += n->pending_mounts; @@ -3370,6 +3372,30 @@ static inline void ve_mount_nr_dec(struct mount *mnt) mnt->ve_owner = NULL; } +/* + * A mount that enters the mount namespace of another VE changes hands, so + * that per-VE mount accounting and the trusted-exec check see it as owned by + * the VE it now lives in. Ownership describes where the mount is, so it + * simply follows the namespace and nothing here special cases which way the + * mount travelled. + * + * vfsmount lock must be held for write, it serializes the owner change + * against is_sb_ve_accessible(). The trusted-exec path reads ve_owner + * locklessly but only compares the pointer, which is never NULL here. + */ +static inline void ve_mount_reown(struct mount *mnt, struct ve_struct *ve) +{ + struct ve_struct *old = mnt->ve_owner; + + if (old == ve) + return; + + atomic_inc(&ve->mnt_nr); + WRITE_ONCE(mnt->ve_owner, get_ve(ve)); + atomic_dec(&old->mnt_nr); + put_ve(old); +} + bool is_sb_ve_accessible(struct ve_struct *ve, struct super_block *sb) { struct mount *mnt; @@ -3392,6 +3418,7 @@ bool is_sb_ve_accessible(struct ve_struct *ve, struct super_block *sb) static inline int ve_mount_allowed(void) { return 1; } static inline void ve_mount_nr_inc(struct mount *mnt, struct ve_struct *ve) { } static inline void ve_mount_nr_dec(struct mount *mnt) { } +static inline void ve_mount_reown(struct mount *mnt, struct ve_struct *ve) { } #endif /* CONFIG_VE */ /* diff --git a/kernel/ve/ve.c b/kernel/ve/ve.c index 750a1b2882a7d..b2788b80e7655 100644 --- a/kernel/ve/ve.c +++ b/kernel/ve/ve.c @@ -1727,9 +1727,13 @@ static bool ve_check_trusted_file(struct file *file) /* * bdev can be NULL if the file is on tmpfs, for example. * If this is a host's tmpfs - execution is allowed. + * + * The read is unlocked and pairs with the store in + * ve_mount_reown(): a task holding a descriptor on the mount + * can be executing from it while the mount changes hands. */ file_on_host_mount = ve_is_super( - real_mount(file->f_path.mnt)->ve_owner); + READ_ONCE(real_mount(file->f_path.mnt)->ve_owner)); if (file_on_host_mount) return true; } _______________________________________________ Devel mailing list Devel@openvz.org https://lists.openvz.org/mailman/listinfo/devel