From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail.openvz.org (unknown [69.168.225.77]) by lore.virtuozzo.com (Postfix) with ESMTPS id 31C5D80024 for ; Wed, 26 Aug 2026 13:10:40 +0000 (UTC) Received: from mail.openvz.org (localhost [127.0.0.1]) by mail.openvz.org (8.14.4/8.14.4) with ESMTP id 67QD9Mr0008344; Wed, 26 Aug 2026 16:09:23 +0300 DKIM-Filter: OpenDKIM Filter v2.11.0 mail.openvz.org 67QD9Mr0008344 Authentication-Results: mail.openvz.org; dkim=fail reason="signature verification failed" (2048-bit key) header.d=virtuozzo.com header.i=@virtuozzo.com header.b="AvJCoOm2" Received: from mail-ej1-f72.google.com (mail-ej1-f72.google.com [209.85.218.72]) by mail.openvz.org (8.14.4/8.14.4) with ESMTP id 67QD9ECa008318 (version=TLSv1/SSLv3 cipher=AES128-GCM-SHA256 bits=128 verify=FAIL) for ; Wed, 26 Aug 2026 16:09:14 +0300 DKIM-Filter: OpenDKIM Filter v2.11.0 mail.openvz.org 67QD9ECa008318 Received: by mail-ej1-f72.google.com with SMTP id a640c23a62f3a-c15ddb61f13so73417966b.0 for ; Wed, 26 Aug 2026 06:09:14 -0700 (PDT) X-Gm-Message-State: AFuF++lMctRarcLt/vZXCu9ozvUXj1l5VJzgIonfx6UP9iCNYHKeOvaY eyozDHX7ErRuPODCd+ZKR7ogg+a2U6n4g2rCEPLBaW5NXiimDWuLuXbSo4tLkYuEBLtGZHkH2vK w5e2X4doDQZzHKQ0Rsn5N3tDrMewqHiiztz0bimDUoXD4+rk/tiQWuQ== X-Gm-Gg: AR+sD13WC6DTeCHxL3JFwFkuVuwE449MfAYCgjVHzHtx+wMOXyqO9o5bFO4eDUqW5Ke ps7aQF7uJ7v4dCdIysaqnLCXuyy3yzeL52JmL1riKPkfnsCjYUjRIY2etViu2iERAey7/TCi8Ph uTSXIaBndcFnmOCXee15rPiPISwrp5mLyHfK5Y/240q2n5dSPe7DkhHTEs6UKNyFcU7rglEz4Mq NLdUYylg4MkUiGxNKE+YFMnBN/IWFHiqkyEbIgSrw+2RXaUhCg+JwMCLA48fc1DNaEV+vulOGLr 3M8T8wbefEeiaiCjh2vsUAoY8bgYjrLGQyOVnyw1XsResWR/Qgn0cH9co0WBRUFaPsqGd6DirGl ixbELWtVBDdZBw12C X-Received: by 2002:a17:907:3e9b:b0:c16:8799:fcb4 with SMTP id a640c23a62f3a-c250bc0fe60mr897971066b.19.1787749754209; Wed, 26 Aug 2026 06:09:14 -0700 (PDT) X-Received: by 2002:a17:907:3e9b:b0:c16:8799:fcb4 with SMTP id a640c23a62f3a-c250bc0fe60mr897960366b.19.1787749753547; Wed, 26 Aug 2026 06:09:13 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1787749753; cv=none; d=google.com; s=arc-20260327; b=T6007HELLWQ7q0HM3ogtOhY5WMIj9XiJI1Jij/MvzEuSrECtm++LE6ZFZrOkw2/uzl fTq9DxXJuUGhFgUDfKM0IlVz/iS8Yu1aW49zsDEgCw8OYdduCHW5v0hfWHJTZAqCpwt+ k+2HAVnoXA3vwoOfyq5J0IIN1IW8Cu+xnaX95Xccm/LHJCUF3FKOS9phm6+w+j5BN2tm wd1f72suzC4UP2MFrru6FuMtjgDBHuEa16OjHm/BEcKqWeXVKkKl8CDPClKqg0az34lM xhzPUFFmhIW1q0szrLLHwZ60WxEqPcNu/MFMYYeJijvwn6iXnq8F3KFbz68E6rW6uWRR uVaQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=subject:in-reply-to:cc:to:from:message-id:date:dkim-signature; bh=7HkiR+19SlXTBlNf5esbdQBM3rmd1HTKmDIB+gKVRGk=; fh=UoUx/ScICNSLS7LR9gyUDsZM8Zuq0DvPxxrx16Z82mE=; b=onqxtlDGiBEVKeQ23L+n5WfpOVK7hnlNrfa8GcToAjeBCjE4KZmzlsMHCv291Ql1rO +2I0WIvo3B/jS8fkgKV+AmFlTCryNHpYfaD4/WCWZFZim3iTT1GvirGiVclcil4uQvKY DnKNlPzxjzaSvuFctELgdIEYPxYzJizMIK3PYdsf9+bqvIi8ziZYhwfrovuPYyQmiBtO Ypk6JcUVfTV/7IbLIS9fEPB1q0MYfR/uy0x66pgxHFtkAc1GlJQLemPSpPJ2lwVfVgFF xgqlAscKO6tS9VJK5HBmJ1tdmlw+aoUneIBGqGa+qdNmPqp8whk0AubTp9zsjucaJt5h QTSg==; dara=google.com ARC-Authentication-Results: i=1; mx.google.com; dkim=pass header.i=@virtuozzo.com header.s=relay header.b=AvJCoOm2; spf=pass (google.com: domain of khorenko@virtuozzo.com designates 130.117.225.111 as permitted sender) smtp.mailfrom=khorenko@virtuozzo.com; dmarc=pass (p=QUARANTINE sp=QUARANTINE dis=NONE) header.from=virtuozzo.com Received: from relay.virtuozzo.com (relay.virtuozzo.com. [130.117.225.111]) by mx.google.com with ESMTPS id a640c23a62f3a-c250a721a65si372102866b.83.2026.08.26.06.09.13 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Aug 2026 06:09:13 -0700 (PDT) Received-SPF: pass (google.com: domain of khorenko@virtuozzo.com designates 130.117.225.111 as permitted sender) client-ip=130.117.225.111; Authentication-Results: mx.google.com; dkim=pass header.i=@virtuozzo.com header.s=relay header.b=AvJCoOm2; spf=pass (google.com: domain of khorenko@virtuozzo.com designates 130.117.225.111 as permitted sender) smtp.mailfrom=khorenko@virtuozzo.com; dmarc=pass (p=QUARANTINE sp=QUARANTINE dis=NONE) header.from=virtuozzo.com DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=virtuozzo.com; s=relay; h=Subject:From:Message-Id:Date:Content-Type: MIME-Version; bh=7HkiR+19SlXTBlNf5esbdQBM3rmd1HTKmDIB+gKVRGk=; b=AvJCoOm2q0GQ VDW7CJk//g59cI635Z5aDt8OrKCvD2uB77ZVBoxswxKcw/QObyinRGJ3gfHLp3u4mgtJhqhatKSA+ 6TFJbyXu8kYL1/sCwuHpbwLaQiHbohIg9lNG4LuPMnYEjS84Vkw3oJnb2zjYof/UULrUnSDoJv7i7 0AJxe+jg53q1BFbXvDY0JFrYEPAUrZo0WyblzPl9AEDxOY5fzFQFWna545BIk88YNXfLzEWz7JqVl yFOj1FYNDdOzMS7XzmTgL9FgN0kSEoB6c4tsiypNAH4HMxjY5FeynUTwPk7ReEo5e212aPp7BD+2I hMZ+zGympFMg35tFCoOBuA==; Received: from ch-demo-asa.virtuozzo.com ([130.117.225.8] helo=f0.sw.ru) by relay.virtuozzo.com with esmtps (TLS1.3) tls TLS_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzDLC-001rmX-25; Wed, 26 Aug 2026 15:09:12 +0200 Received: from f0.sw.ru (localhost [127.0.0.1]) by f0.sw.ru (8.18.1/8.18.1/Debian-2) with ESMTP id 67QD9C0e890837; Wed, 26 Aug 2026 15:09:12 +0200 Received: (from kostja@localhost) by f0.sw.ru (8.18.1/8.18.1/Submit) id 67QD9CBZ890835; Wed, 26 Aug 2026 15:09:12 +0200 Date: Wed, 26 Aug 2026 15:09:12 +0200 Message-Id: <202608261309.67QD9CBZ890835@f0.sw.ru> X-Authentication-Warning: f0.sw.ru: kostja set sender to khorenko@virtuozzo.com using -f From: Konstantin Khorenko To: Vasileios Almpanis In-Reply-to: <20260825-connectors-v3-2-7b26773876a0@virtuozzo.com> X-OZ-Fwd: true Cc: OpenVZ devel Subject: Re: [Devel] [PATCH RHEL10 COMMIT] connector: free the per-VE connector state after an RCU grace period X-BeenThere: devel@openvz.org X-Mailman-Version: 2.1.12 Precedence: list List-Id: OpenVZ development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , MIME-Version: 1.0 Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit Sender: devel-bounces@openvz.org Errors-To: devel-bounces@openvz.org The commit is pushed to "branch-rh10-6.12.0-211.39.1.16.x.vz10-ovz" and will appear at git@bitbucket.org:openvz/vzkernel.git after rh10-6.12.0-211.39.1.16.10.vz10 ------> commit 578a77d6c8afd2a6aeb5c25a015fb2d0f22cf198 Author: Vasileios Almpanis Date: Tue Aug 25 16:19:52 2026 +0000 connector: free the per-VE connector state after an RCU grace period cn_fini_ve() tears down everything the proc event delivery path uses: the percpu local_event, the callback device and the kernel netlink socket. It only clears ve->cn at the very end. A reader that fetched ve->cn right before the teardown dereferences freed memory afterwards: the only guard on the delivery path is the ve->cn check in proc_event_num_listeners() and nothing keeps the state alive once the check has passed. Today the window is not reachable: the per-VE delivery path is only entered for a task alive in this VE, a live task keeps the VE pid namespace busy, so zap_pid_ns_processes() -> ve_exit_ns() -> cn_fini_ve() cannot run in parallel. A subsequent patch will make proc_exit_connector() deliver the exit event with a VE reference pinned before exit_notify(), i.e. possibly after the task was reaped and stopped pinning the pid namespace, then the teardown can run in parallel with the delivery. Clear ve->cn and wait for an RCU grace period before freeing anything reachable from it, so that the delivery path can safely use the state it observed within a single RCU read-side critical section (done by the next patch). The clearing is done in cn_proc_fini_ve(): it has to happen before the first thing the delivery path uses (local_event) is freed, and everything else is freed later in cn_fini_ve(). The kernel netlink socket outlives the ve->cn clearing: it is released only later in cn_fini_ve(), after this grace period. A message sent from the host into the dying VE netns (e.g. via nsenter) can therefore still reach cn_call_callback() with ve->cn already NULL, so get_cdev() now returns NULL there. Look the callback device up and walk the queue under rcu_read_lock() and bail out on NULL, so the receive path is covered by the same grace period as the delivery path. cn_init_ve() publishes ve->cn early with rcu_assign_pointer(), before the connector is fully set up, so its error path may already have an RCU reader looking at the state. Wait for a grace period there too before freeing it, symmetric to cn_proc_fini_ve(). https://virtuozzo.atlassian.net/browse/VSTOR-140421 Feature: ve: ve generic structures Signed-off-by: Vasileios Almpanis Reviewed-by: Konstantin Khorenko --- drivers/connector/cn_proc.c | 11 +++++++++++ drivers/connector/connector.c | 25 +++++++++++++++++++++++-- 2 files changed, 34 insertions(+), 2 deletions(-) diff --git a/drivers/connector/cn_proc.c b/drivers/connector/cn_proc.c index d4ce1697dd0bb..6095c7def7cea 100644 --- a/drivers/connector/cn_proc.c +++ b/drivers/connector/cn_proc.c @@ -535,5 +535,16 @@ void cn_proc_fini_ve(struct ve_struct *ve) ve_is_super(ve)); cn_del_callback_ve(ve, &cn_proc_event_id); + + /* + * Hide the connector state from the proc event delivery path, + * which dereferences ve->cn under rcu_read_lock(), and wait for + * the readers to finish before anything reachable from it is + * freed: the percpu local_event here, the callback device, the + * netlink socket and the state itself in cn_fini_ve(). + */ + RCU_INIT_POINTER(ve->cn, NULL); + synchronize_rcu(); + free_percpu(cn->local_event); } diff --git a/drivers/connector/connector.c b/drivers/connector/connector.c index bf4a83f3f8703..05c281bb321be 100644 --- a/drivers/connector/connector.c +++ b/drivers/connector/connector.c @@ -151,7 +151,7 @@ static int cn_call_callback(struct sk_buff *skb) { struct nlmsghdr *nlh; struct cn_callback_entry *i, *cbq = NULL; - struct cn_dev *dev = get_cdev(sock_net(skb->sk)->owner_ve); + struct cn_dev *dev; struct cn_msg *msg = nlmsg_data(nlmsg_hdr(skb)); struct netlink_skb_parms *nsp = &NETLINK_CB(skb); int err = -ENODEV; @@ -161,6 +161,21 @@ static int cn_call_callback(struct sk_buff *skb) if (nlh->nlmsg_len < NLMSG_HDRLEN + sizeof(struct cn_msg) + msg->len) return -EINVAL; + /* + * On VE stop cn_proc_fini_ve() clears ve->cn and waits for a grace + * period before the callback device is freed, but the kernel socket + * is released only afterwards, so a message sent from the host into + * the dying VE netns can still get here. Look the device up and walk + * the callback queue under rcu_read_lock(); the callback itself may + * sleep and is kept alive by the refcount taken here. + */ + rcu_read_lock(); + dev = get_cdev(sock_net(skb->sk)->owner_ve); + if (!dev) { + rcu_read_unlock(); + return -ENODEV; + } + spin_lock_bh(&dev->cbdev->queue_lock); list_for_each_entry(i, &dev->cbdev->queue_list, callback_entry) { if (cn_cb_equal(&i->id.id, &msg->id)) { @@ -170,6 +185,7 @@ static int cn_call_callback(struct sk_buff *skb) } } spin_unlock_bh(&dev->cbdev->queue_lock); + rcu_read_unlock(); if (cbq != NULL) { cbq->callback(msg, nsp); @@ -362,7 +378,12 @@ static int cn_init_ve(void *data) netlink_release: netlink_kernel_release(dev->nls); free_cn: + /* + * The pointer was published: wait for the readers before freeing, + * same as cn_proc_fini_ve() does. + */ RCU_INIT_POINTER(ve->cn, NULL); + synchronize_rcu(); kfree(cn); goto net_unlock; } @@ -394,7 +415,7 @@ static void cn_fini_ve(void *data) cn_queue_free_dev(dev->cbdev); netlink_kernel_release(dev->nls); - RCU_INIT_POINTER(ve->cn, NULL); + /* ve->cn was cleared by cn_proc_fini_ve() before the grace period */ kfree(cn); } _______________________________________________ Devel mailing list Devel@openvz.org https://lists.openvz.org/mailman/listinfo/devel