From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from AS8PR04CU009.outbound.protection.outlook.com (mail-westeuropeazon11021098.outbound.protection.outlook.com [52.101.70.98]) by lore.virtuozzo.com (Postfix) with ESMTPS id 151F380275 for ; Thu, 3 Sep 2026 12:35:05 +0000 (UTC) ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=O0Tx8t6rJ/tT8EW/C1HfeN3Qmy18iqgNun7DSgMj463T0HmpVdrbY1NZWEYqUAGbDISI3CKXTaidAN3iKCGSzHP6F3SHmgEFf6g1wGymNW5i+CRIi6W64bncrLQS4jNTJ7p2v7OkLU4yvlk7DFZwAMecAUsAqR4SQDYOIFLNr8igEgsq+26PL59COMxjFVMAfsMlWBX9neDMneuGDE+FDLWAyT6y9BA9Zvd3A9CtJR7Y5ijYgvlBetms17lWETaCF40ByfmVI7krFzeyV0twQsx3cEFN21b42PfLOTOEZhhOYVLHz+wWLgPDWnvpWhG7tUNQ+9CC/jtMqRjoElhUqQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=Q5sG3Zb16/Bornp3oPL9dTcTbvJkK9hN94fWOQomCQ0=; b=HFjtAdaztoIckh9N1LiEb/pFTfzk4YNY2yB1DYcBOAZGFanCYgOr+ClKNMA4eOSU1m3pnPnma8m+BmmppSw/keicCxZ5V3zjM5OOnVZ+g0W7AoVOLokTvFFMktbYcQCIkIiT9Wrq6bJP2WW6hWCUlnS8nztxMpzqdq7V6rqyEQkAnUnoB+rfD8sl/9TPJybzdU8txh1XJX30M/CLKAYyKuoQ+pT8Rkiah/tULOCeHV7OmrEQ4v88wf788yV+3y7VUg79dyMMR8CFDHPR5GNWlTWYgi9EqxJJEBKUwYcnKLwzjytKVT+Koo9wjIW6wB2uyVYILjcaERjAXQ9aDyF3LQ== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 130.117.225.111) smtp.rcpttodomain=virtuozzo.com smtp.mailfrom=virtuozzo.com; dmarc=pass (p=quarantine sp=quarantine pct=100) action=none header.from=virtuozzo.com; dkim=pass (signature was verified) header.d=virtuozzo.com; arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=virtuozzo.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=Q5sG3Zb16/Bornp3oPL9dTcTbvJkK9hN94fWOQomCQ0=; b=CbGwZClKLhYsqcA6WbIBp4mYOX8OEiZxgnefKGL9rgXacAfpKxCleExxsDArJwKH5K4geJteVIJ4hcBhJ4F0BoQU6PWbZtdxRUKPiI9IcL5FzJrtu33IPzKbtgx1FCQStFfxJ6/jHI8l6Xp3SoMuFtfO2Y2ALUN3I8VTz/t2GslEqXUQy1sLrjdlUC2z/TnVD+szb9ee01Ue2piRC2m/J6+iw5fmqfaATwk1cu0SKb1nR1ZsDhTT2+9n/i35TVcKFiIbQVqqOP4e6NtsCcQb7oJOsYis/+v2Og6JZi5AkPnSSsLhxH/WTvoYDksokCZ/uwjdxoyIpTagoxqY/Qx8Lg== Received: from DUZPR01CA0034.eurprd01.prod.exchangelabs.com (2603:10a6:10:468::10) by PAVPR08MB8992.eurprd08.prod.outlook.com (2603:10a6:102:325::10) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.382.10; Thu, 3 Sep 2026 12:34:58 +0000 Received: from DU2PEPF00028CFC.eurprd03.prod.outlook.com (2603:10a6:10:468:cafe::98) by DUZPR01CA0034.outlook.office365.com (2603:10a6:10:468::10) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.382.11 via Frontend Transport; Thu, 3 Sep 2026 12:34:58 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 130.117.225.111) smtp.mailfrom=virtuozzo.com; dkim=pass (signature was verified) header.d=virtuozzo.com;dmarc=pass action=none header.from=virtuozzo.com; Received-SPF: Pass (protection.outlook.com: domain of virtuozzo.com designates 130.117.225.111 as permitted sender) receiver=protection.outlook.com; client-ip=130.117.225.111; helo=relay.virtuozzo.com; pr=C Received: from relay.virtuozzo.com (130.117.225.111) by DU2PEPF00028CFC.mail.protection.outlook.com (10.167.242.180) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.360.3 via Frontend Transport; Thu, 3 Sep 2026 12:34:58 +0000 DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=virtuozzo.com; s=relay; h=MIME-Version:Message-ID:Date:Subject:From: Content-Type; bh=Q5sG3Zb16/Bornp3oPL9dTcTbvJkK9hN94fWOQomCQ0=; b=JeTV7BZSypkB 9eWmGpeS7ygyMCU+4K4LFpoOJbX5oKDOLwWcigkdn7OI1hFk8+8+IMe/5XbNa3C/Q/qeFX8AJRb8X HswYLTdQ9lddfBd5Kn++JccfzVmJx3hKdYsNU8KEAzohr+F+yf9cJDaw4HzCe1rWJdxXul4GcMtRu yCh3Ps+psdrDMA74IwJMv9xsAVQBRLCL0rS660d+y5vcXi7EoVuaTxib1PQmSIFbSJMzJ8ZiEmaVh cO/JzJ84PAYm8oRA0nsdnIbvpgkn9Q9GEOGYfnqv753B6puw+asNLLGHq5plBXptlukYWWghr4qIk OyoF07nLhheN9B+Sm/vVtQ==; Received: from [130.117.225.5] (helo=vz9-demens-1.aci.vzint.dev) by relay.virtuozzo.com with esmtp (Exim 4.96) (envelope-from ) id 1x26cF-00AcxS-1p; Thu, 03 Sep 2026 14:34:56 +0200 From: Andrey Zhadchenko To: svt-core@virtuozzo.com Cc: den@openvz.org, andrey.drobyshev@virtuozzo.com Subject: [QEMU HCI-8.0 PATCH 5/5] vhost-blk: filter uevents in the kernel Date: Thu, 3 Sep 2026 15:32:04 +0300 Message-ID: <20260903123204.24035-6-andrey.zhadchenko@virtuozzo.com> X-Mailer: git-send-email 2.43.5 In-Reply-To: <20260903123204.24035-1-andrey.zhadchenko@virtuozzo.com> References: <20260903123204.24035-1-andrey.zhadchenko@virtuozzo.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: DU2PEPF00028CFC:EE_|PAVPR08MB8992:EE_ X-MS-Office365-Filtering-Correlation-Id: c374569b-a8f0-427d-40e3-08df09b7c447 X-LD-Processed: 0bc7f26d-0264-416e-a6fc-8352af79c58f,ExtAddr List-Id: svt-core@virtuozzo.com X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|36860700016|82310400026|376014|1800799024|23010399003|10067099003|56012099006|3023799007|6133799003|5113699003|18002099003|13003099007|22082099003; X-Microsoft-Antispam-Message-Info: 1OrLBGGNRz4sj0t4C8NIqJSS+34oNA16uOmtvIdSXYWyRthR+vBir5ePTLL3JdGZ+XEou/r2F0lHzjLV0cVBUNzov3tfuvKPf3TAWpcXyrnv/rXmgePE+QP+Svh/W5ouAdvhJQWbEQ5qWhnWEccn2CvYyayfzbl/FlYmRBj2o5BkjsgYh7aJUPwtiKwTe+s4KuektC8FhCCF2gvGswYgeA6R0P6CvaiZB66+qZf9eSy2zBrOyd0BxqwK5LB3pjToLLNGYBreKR+ZhRySsBmgG1YV2MrDOB/s2yg27wi0uI/A2hl/kTkA9jJybBhzcLhLtTpYgtAiICFefuHxrLdILExJ33tk9nowYu30dUeMbTvlqWmI/XgfgBnkaPmcBmJ/PNRu5xR6EDlJmRLjBmNJ+hD3GJmZAxpfIh8hOkNU7MlF52RiGxOJnw0eCuUyjDOKAaUFBmx7aAFrStGwMQ0/Pair7PJ+jMLB9bP6NG0Q2pdNOAXwiIlHXu0MtNcK3igTWJWTcl5DdIRDjUS/COciNmlyyb+BCeKc0WPZT7UPhYVOvw1P1kGUnnSl+ZPSQW2WoXAVOn7FJAKLzvjWkia1iM2imAlKQlz0AclcLWmMoH7O2xOmN3cWr+dWgC4XToLZdWQBWdHMWCck0Kqr+gR4au31iZiTuIkQWfnra6hMvukTFC3qHLc6/llRB/eD17bSZ+/7Dl626tMoa0rAYy2p4Q== X-Forefront-Antispam-Report: CIP:130.117.225.111;CTRY:CH;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:relay.virtuozzo.com;PTR:relay.virtuozzo.com;CAT:NONE;SFS:(13230040)(36860700016)(82310400026)(376014)(1800799024)(23010399003)(10067099003)(56012099006)(3023799007)(6133799003)(5113699003)(18002099003)(13003099007)(22082099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: jhvLomFJ9vearxqW9EL1oTA2grSr7syqWA+4uH7ua6yTwIZ4R+AjxqyJpQcMv/bkASLoMXgG/pMqWzvNBEAYyG7QYq2/L19qpgPT4P4LqtGIqRLRQ5ja9ux7okv1XCqdOBx/AdYcM9vp9K3h4eWk+BFc9dqqE1XqIAyl/Z1YF/cqltH01m1T9AT6+/3H+RbvkrCRse7AOlP1XCf0pkIJz7CLAq93/RLmXYS+RlXaXaXnyXbf2J/MvxWk4fb6YPZjEh3op92vn/SgKHF3URwEH6rCFGBoVydMF0oMMO2B/AGkB0QOsG36+MvoAkNvUmafn4OviUSaZWurCSyLfJO0L81yNnDCrDEeW6vuy5WQJSg32WjWOCnCw70V9KiOfxQ6rkfNWX7A+Pw+Mg1P+lmN/tO+bczekczZa7agomLFQj2uZkVJZ1QRtyfjdeovyQp1 X-Auto-Response-Suppress: DR, OOF, AutoReply X-OriginatorOrg: virtuozzo.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 03 Sep 2026 12:34:58.6525 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: c374569b-a8f0-427d-40e3-08df09b7c447 X-MS-Exchange-CrossTenant-Id: 0bc7f26d-0264-416e-a6fc-8352af79c58f X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=0bc7f26d-0264-416e-a6fc-8352af79c58f;Ip=[130.117.225.111];Helo=[relay.virtuozzo.com] X-MS-Exchange-CrossTenant-AuthSource: DU2PEPF00028CFC.eurprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: PAVPR08MB8992 The uevent socket receives every uevent broadcast on the host: NETLINK_KOBJECT_UEVENT group 1 has no kernel-side subscription by subsystem or device. On a dense node every add/remove/change event of every device (mass container starts creating dm and loop devices, SCSI rescans, udevadm trigger) wakes up the main loop of every QEMU with a vhost-blk device just to parse and discard the message, and a burst can overflow the socket receive buffer. Attach a classic BPF socket filter which passes only messages starting with "change@" and containing a "RES"-prefixed property within the first 512 bytes. Messages longer than the scan window are passed to userspace instead of being dropped, so the filter can have false positives but never false negatives: vhost_blk_uevent_read() remains the authoritative parser. Also enlarge the receive buffer to 1M: with the filter attached even a large backlog consists of relevant events only. Matching the full "RESIZE=1" property or MAJOR=/MINOR= of the watched devices kernel-side was considered and rejected: the kernel converts classic BPF to eBPF on attach and the converted program must fit in BPF_MAXINSNS, which allows only ~1900 classic instructions of this shape (three per scanned offset). Classic BPF also cannot loop, so device numbers (variable-length decimal strings at variable offsets) would need unrolled matching code regenerated and re-attached on every device plug/unplug. Resize events are rare; the coarse kernel filter drops all of the heavy traffic and userspace keeps doing the exact matching. https://virtuozzo.atlassian.net/browse/VSTOR-143437 Signed-off-by: Andrey Zhadchenko --- hw/block/vhost-blk.c | 123 +++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 123 insertions(+) diff --git a/hw/block/vhost-blk.c b/hw/block/vhost-blk.c index f8eca4d58a..80b0e06b4c 100644 --- a/hw/block/vhost-blk.c +++ b/hw/block/vhost-blk.c @@ -28,6 +28,7 @@ #include #include #include +#include #include "system/runstate.h" static int vhost_blk_uevent_fd = -1; @@ -347,6 +348,126 @@ static void vhost_blk_uevent_read(void *opaque) } } +/* + * The kernel broadcasts uevents of every device on the host to + * NETLINK_KOBJECT_UEVENT group 1 and provides no subscription by subsystem + * or device. Without a filter each uevent (device hotplug, SCSI rescan, + * udevadm trigger, ...) wakes up the main loop of every QEMU with a + * vhost-blk device just to parse and discard the message. + * + * Attach a classic BPF socket filter passing only what we are interested + * in: messages which start with "change@" and contain a property beginning + * with "RES" ("\0RES" match at every offset) within the first + * VHOST_BLK_UEVENT_SCAN_LEN bytes. Messages longer than the scan window + * are passed to userspace instead of being dropped. The filter can have + * false positives but never false negatives: vhost_blk_uevent_read() + * remains the authoritative parser. + * + * Matching the full "\0RESIZE=1\0" property would be nicer, but the kernel + * converts classic BPF to eBPF on attach and every packet load expands to + * several eBPF instructions; the converted program must fit in + * BPF_MAXINSNS (4096) instructions, which allows roughly 1900 classic + * instructions of this shape. Three instructions per scanned offset + * (load, match, accept-jump) fit with a good margin, seven do not. + * + * Filtering by MAJOR=/MINOR= of the watched devices is done in userspace + * only. Classic BPF cannot loop, so matching these variable-length decimal + * strings at variable offsets would require regenerating and re-attaching + * unrolled matching code on every device plug/unplug, and the instruction + * budget above does not allow anything close to that. Resize events are + * rare, all of the heavy traffic is already dropped by the "change@" and + * "\0RES" matches. + */ +#define VHOST_BLK_UEVENT_SCAN_LEN 512 +#define VHOST_BLK_UEVENT_FILTER_HEAD 10 +#define VHOST_BLK_UEVENT_FILTER_BLOCK 3 +#define VHOST_BLK_UEVENT_FILTER_INSNS (VHOST_BLK_UEVENT_FILTER_HEAD + \ + VHOST_BLK_UEVENT_FILTER_BLOCK * \ + VHOST_BLK_UEVENT_SCAN_LEN + 2) + +static void vhost_blk_uevent_apply_filter(int fd) +{ + g_autofree struct sock_filter *insns = + g_new0(struct sock_filter, VHOST_BLK_UEVENT_FILTER_INSNS); + const uint32_t accept = VHOST_BLK_UEVENT_FILTER_INSNS - 1; + struct sock_fprog prog = { + .len = VHOST_BLK_UEVENT_FILTER_INSNS, + .filter = insns, + }; + int rcvbuf = 1024 * 1024; + uint32_t pc = 0; + uint32_t i; + + QEMU_BUILD_BUG_ON(VHOST_BLK_UEVENT_FILTER_INSNS > BPF_MAXINSNS); + + /* Drop everything which does not start with "change@" */ + insns[pc++] = (struct sock_filter) + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, 0); + insns[pc++] = (struct sock_filter) + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, 0x6368616e /* "chan" */, 0, 4); + insns[pc++] = (struct sock_filter) + BPF_STMT(BPF_LD | BPF_H | BPF_ABS, 4); + insns[pc++] = (struct sock_filter) + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, 0x6765 /* "ge" */, 0, 2); + insns[pc++] = (struct sock_filter) + BPF_STMT(BPF_LD | BPF_B | BPF_ABS, 6); + insns[pc++] = (struct sock_filter) + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, '@', 1, 0); + insns[pc++] = (struct sock_filter) + BPF_STMT(BPF_RET | BPF_K, 0); + + /* + * A message longer than the scan window cannot be scanned completely: + * pass it to userspace instead of risking a lost resize event. + */ + insns[pc++] = (struct sock_filter) + BPF_STMT(BPF_LD | BPF_W | BPF_LEN, 0); + insns[pc++] = (struct sock_filter) + BPF_JUMP(BPF_JMP | BPF_JGT | BPF_K, VHOST_BLK_UEVENT_SCAN_LEN, 0, 1); + insns[pc] = (struct sock_filter) + BPF_STMT(BPF_JMP | BPF_JA, accept - pc - 1); + pc++; + + /* + * Scan for "\0RES" at every offset. A load beyond the end of the + * message terminates the program with a drop verdict, which is + * correct: had the message contained the pattern, it would have been + * matched at an earlier, in-bounds offset. + */ + for (i = 0; i < VHOST_BLK_UEVENT_SCAN_LEN; i++) { + insns[pc++] = (struct sock_filter) + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, i); + insns[pc++] = (struct sock_filter) + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, 0x00524553 /* "\0RES" */, + 0, 1); + insns[pc] = (struct sock_filter) + BPF_STMT(BPF_JMP | BPF_JA, accept - pc - 1); + pc++; + } + + /* Drop */ + insns[pc++] = (struct sock_filter)BPF_STMT(BPF_RET | BPF_K, 0); + /* Accept */ + insns[pc++] = (struct sock_filter)BPF_STMT(BPF_RET | BPF_K, 0xffffffff); + assert(pc == VHOST_BLK_UEVENT_FILTER_INSNS); + + if (setsockopt(fd, SOL_SOCKET, SO_ATTACH_FILTER, &prog, sizeof(prog))) { + warn_report("vhost-blk: unable to attach uevent filter: %s", + strerror(errno)); + } + + /* + * Make the socket resilient to main loop stalls. With the filter + * attached even a large backlog consists of relevant events only. + */ + if (setsockopt(fd, SOL_SOCKET, SO_RCVBUFFORCE, + &rcvbuf, sizeof(rcvbuf)) && + setsockopt(fd, SOL_SOCKET, SO_RCVBUF, &rcvbuf, sizeof(rcvbuf))) { + warn_report("vhost-blk: unable to enlarge uevent socket buffer: %s", + strerror(errno)); + } +} + static bool vhost_blk_uevent_init(Error **errp) { struct sockaddr_nl address = { @@ -367,6 +488,8 @@ static bool vhost_blk_uevent_init(Error **errp) return false; } + vhost_blk_uevent_apply_filter(vhost_blk_uevent_fd); + if (bind(vhost_blk_uevent_fd, (struct sockaddr *)&address, sizeof(address)) < 0) { error_setg_errno(errp, errno, -- 2.43.5