OpenVZ / Virtuozzo kernel development (devel@openvz.org)
 help / color / mirror / Atom feed
From: Vasileios Almpanis <vasileios.almpanis@virtuozzo.com>
Subject: Re: [Devel] [PATCH VZ10 6/6] drivers/md/dm-qcow2: seek unallocated L2 entries in one pass within md
Date: Wed, 12 Aug 2026 14:42:23 +0200	[thread overview]
Message-ID: <a9f40e26-03c8-4fee-96cb-48b24e04073b@virtuozzo.com> (raw)
In-Reply-To: <20260810123000.19834-7-andrey.zhadchenko@virtuozzo.com>


On 8/10/26 2:30 PM, Andrey Zhadchenko wrote:
> When a cluster is unallocated but its L2 table exists, the seek code
> walks it one cluster at a time: SEEK_DATA re-parses metadata for
> every cluster of the range, and with a backing image each cluster
> additionally costs a qio and completion allocation for the lower
> delta descent.
>
> If the unmapped range reaches the cluster border, prolong it over the
> following unallocated L2 entries: they are cached in the very md page
> the current entry was parsed from, so this is just scanning for
> non-zero words. Then the lower delta is scanned in one pass over the
> whole range, and SEEK_DATA without a backing image skips it in one go.
>
> Entries with any subcluster or "reads as zeroes" bits set stop the
> scan and spin parse_metadata() machinery.
>
> On a 16G image (128K clusters, extended L2) with all L2 tables
> allocated but no clusters mapped over a backing image with data at
> 15G, warm-cache SEEK_DATA improves from 44.8 ms to 0.4 ms.
>
> https://virtuozzo.atlassian.net/browse/VSTOR-139407
> Feature: dm-qcow2: block device over QCOW2 files driver
> Signed-off-by: Andrey Zhadchenko <andrey.zhadchenko@virtuozzo.com>
> ---
>   drivers/md/dm-qcow2-map.c | 44 +++++++++++++++++++++++++++++++++++++--
>   1 file changed, 42 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/md/dm-qcow2-map.c b/drivers/md/dm-qcow2-map.c
> index 6a73c32689eca..3bda108db725f 100644
> --- a/drivers/md/dm-qcow2-map.c
> +++ b/drivers/md/dm-qcow2-map.c
> @@ -4441,6 +4441,34 @@ static inline void seek_qio_next_clu(struct qio *qio, struct qcow2_map *map)
>   	seek_qio_set_sector(qio, bi_sector);
>   }
>   
> +/*
> + * Manually scan L2 metadata pages until we hit something interesting.
> + * We do not take md_pages_lock here because we have no promises about
> + * concurrent writes anyway.
> + */
> +static loff_t seek_extend_unmapped_end(struct qio *qio, struct qcow2_map *map,
> +				       loff_t end)
> +{
> +	struct qcow2 *qcow2 = qio->qcow2;
> +	loff_t lim = SEEK_QIO_DATA(qio)->lim;
> +	u32 step = 1 + qcow2->ext_l2;
> +	u32 index = map->l2.index_in_page + step;
> +	u64 *entries;
> +
> +	entries = kmap_local_page(map->l2.md->page);
> +	for (; index + step <= PAGE_SIZE / sizeof(u64) && end < lim;
> +	     index += step) {
> +		/* Zero test does not need be64_to_cpu() */
> +		if (READ_ONCE(entries[index]) ||
> +		    (qcow2->ext_l2 && READ_ONCE(entries[index + 1])))
> +			break;
> +		end += qcow2->clu_size;
> +	}
> +	kunmap_local(entries);
> +
> +	return min_t(loff_t, end, lim);
> +}
> +
>   static struct qio *advance_and_spawn_lower_seek_qio(struct qio *old_qio, loff_t end)
>   {
>   	struct qcow2 *lower = old_qio->qcow2->lower;
> @@ -4503,12 +4531,15 @@ static int qcow2_llseek_hole_qio(struct qio *qio, int whence, loff_t *result)
>   				struct qio *new_qio;
>   				loff_t end;
>   
> -				if (!(map.level & L2_LEVEL))
> +				if (!(map.level & L2_LEVEL)) {
>   					end = min_t(loff_t,
>   						    to_bytes(get_next_l2(qio)),
>   						    SEEK_QIO_DATA(qio)->lim);
> -				else
> +				} else {
>   					end = to_bytes(qio->bi_iter.bi_sector) + size;
> +					if (!CLU_OFF(qio->qcow2, end))
> +						end = seek_extend_unmapped_end(qio, &map, end);
> +				}
>   
>   				new_qio = advance_and_spawn_lower_seek_qio(qio, end);
>   				if (!new_qio) {
> @@ -4552,6 +4583,15 @@ static int qcow2_llseek_hole_qio(struct qio *qio, int whence, loff_t *result)
>   				qio->bi_iter.bi_sector += to_sector(size);
>   				qio->bi_iter.bi_size -= size;
>   				goto calc_subclu;
> +			} else if (arg.unmapped && (map.level & L2_LEVEL)) {
> +				loff_t end = to_bytes(qio->bi_iter.bi_sector) + size;
> +
> +				/* Skip the following unallocated entries in one go */
> +				if (!CLU_OFF(qio->qcow2, end)) {
> +					seek_qio_set_sector(qio,

shouldn't you wrap the result of seek_extend_unmapped_end in to_sector? 
It returns bytes and seek_qio_set_sector takes sector_t and we would 
immediately do to_bytes on something that is already bytes.

> +						seek_extend_unmapped_end(qio, &map, end));
> +					continue;
> +				}
>   			}
>   		}
>   

-- 
Best regards, Vasileios Almpanis
Software Developer, Virtuozzo.


  reply	other threads:[~2026-08-12 12:42 UTC|newest]

Thread overview: 16+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-10 12:29 [Devel] [PATCH VZ10 0/6] dm-qcow: lseek improvements Andrey Zhadchenko
2026-08-10 12:29 ` [Devel] [PATCH VZ10 1/6] drivers/md/dm-qcow2: fix lower delta seek window clamp Andrey Zhadchenko
2026-08-10 12:29 ` [Devel] [PATCH VZ10 2/6] drivers/md/dm-qcow2: keep seek parse window within limit Andrey Zhadchenko
2026-08-14  9:55   ` Pavel Tikhomirov
2026-08-14 10:26     ` Andrey Zhadchenko
2026-08-10 12:29 ` [Devel] [PATCH VZ10 3/6] drivers/md/dm-qcow2: fix seek constant comparison Andrey Zhadchenko
2026-08-10 12:29 ` [Devel] [PATCH VZ10 4/6] drivers/md/dm-qcow2: fix seek skip to the next L2 table Andrey Zhadchenko
2026-08-14 10:25   ` Pavel Tikhomirov
2026-08-14 10:33     ` Andrey Zhadchenko
2026-08-10 12:29 ` [Devel] [PATCH VZ10 5/6] drivers/md/dm-qcow2: scan lower delta in one pass over empty L1 entry Andrey Zhadchenko
2026-08-14 11:09   ` Pavel Tikhomirov
2026-08-14 12:00     ` Andrey Zhadchenko
2026-08-10 12:30 ` [Devel] [PATCH VZ10 6/6] drivers/md/dm-qcow2: seek unallocated L2 entries in one pass within md Andrey Zhadchenko
2026-08-12 12:42   ` Vasileios Almpanis [this message]
2026-08-12 12:50     ` Andrey Zhadchenko
2026-08-14 11:55 ` [Devel] [PATCH VZ10 0/6] dm-qcow: lseek improvements Pavel Tikhomirov

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=a9f40e26-03c8-4fee-96cb-48b24e04073b@virtuozzo.com \
    --to=vasileios.almpanis@virtuozzo.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox