From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3D5093002B9; Sun, 2 Aug 2026 19:50:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700252; cv=none; b=DLorpmApGbCGIr+LzKveDbGlaaB1UtGM6vD4obDHf7MM1CkHLCbeQfOyFxWNVzraR0gA81tnBiKoVsZkBX4o2zR5w1mGocoigWTwGrzMLBStI41zf7bPTX91f1K/jRTCYcxao449fmIFtWn7bmIyMTEJRF8RLvIHyy2laoMdrno= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700252; c=relaxed/simple; bh=Iu/k9tdgkPqDHBtJe0dzZZYxYB7j870w1vAFOE8DOKg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=tpt1L4N3nrUseEdcMsC/X0P1lMGY7Zc0Tud/8J3ALrXE39XVbR2NSbCx7qOBtshd5U7hC9BhWytB0pAINUIKPa7F8soRDyLGXFRBdFhD27Dh5qAnLMVX3fYghTTAH3sxrFlGaOxBMb+TJZkO4aBeZWp+GouOq0DCGVRFawTMhfk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=oyzA89nE; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="oyzA89nE" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 73A8E1F00A3A; Sun, 2 Aug 2026 19:50:47 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700250; bh=LwFUXhNbqVsFRuYQdRZ6MNKemfzqMBb3k7WjbIe3rkE=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=oyzA89nEnYCyaiZEI02iEes6e7vnSq3q0oyDhJL2Ucjj2ha/0RqnSmcsdQJ0UBivV wI2aRPH0x3EyGcSTzJvz/6UICultg2wOhz3oD1HlkaZVL6+MrkFRxn7jGUuN1Hb7N/ 5PcCiF22piyfz8EADSlNvX6/jXXUXcItbfzRN69+pdr2kK01bZXbtn3EIbaCvJ4CqV hWRgPmaxefuuoBKZV4fwXla+L2juVUkeG0Iq+NGAZvStklhoQRRIQdP3npRw1aYjTP WDrJ9jOeH57vMw8qTmUYwSb2W8BxOp6M5DwhDZn1kmlpfJxEFQf58iSJtTqtiIk7gE F9q4znwxggjuw== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 01/29] md/md-llbitmap: clear flush state after daemon flush Date: Mon, 3 Aug 2026 03:50:10 +0800 Message-ID: <20260802195038.164272-2-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai llbitmap_flush() sets LLPageFlush on each bitmap page before it queues the daemon worker. The flag tells md_llbitmap_daemon_fn() to ignore the normal barrier_idle expiry check and clean the page immediately. The daemon only tested LLPageFlush. Once a page had been flushed explicitly, the flag stayed set, so later dirty bits on that page also bypassed barrier_idle and were cleaned the next time the daemon ran. That can make a new write look clean much earlier than the configured idle window. Consume LLPageFlush in md_llbitmap_daemon_fn() with test_and_clear_bit() and use the returned value for the current expiry check. The explicit flush sti= ll forces the current daemon pass, while later writes on the same page wait for barrier_idle again. This can be reproduced through normal sysfs operations: 1. Create a small RAID1 with --bitmap=3Dlockless and --assume-clean. 2. Set llbitmap/daemon_sleep=3D1 and llbitmap/barrier_idle=3D10. 3. Toggle md/array_state from active to readonly and back to active to ca= ll llbitmap_flush() without destroying the in-memory bitmap. 4. Write one sector and read llbitmap/bits immediately, after 2 seconds, and after 12 seconds. On the bad kernel the dirty bit is already clean after 2 seconds. With this change it remains dirty until the barrier_idle window expires. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 2a2b38c663c3..71e9a21b98b2 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -1066,14 +1066,14 @@ static void md_llbitmap_daemon_fn(struct work_struc= t *work) =20 for (idx =3D 0; idx < llbitmap->nr_pages; idx++) { struct llbitmap_page_ctl *pctl =3D llbitmap->pctl[idx]; + bool flush =3D test_and_clear_bit(LLPageFlush, &pctl->flags); =20 if (idx > 0) { start =3D end + 1; end =3D min(end + PAGE_SIZE, llbitmap->chunks - 1); } =20 - if (!test_bit(LLPageFlush, &pctl->flags) && - time_before(jiffies, pctl->expire)) { + if (!flush && time_before(jiffies, pctl->expire)) { restart =3D true; continue; } --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9850E30FC1D; Sun, 2 Aug 2026 19:50:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700254; cv=none; b=ceVV3wL4GIbIgcdn8uDM0eVYAQXJBErd2np60G4DfINydcKCsf1XbH+yFBG2z6QOw4p+Nc3iBirZvHySt7dYGpfkA752bJuTHyCTg5kWYnpc7/PPBsb3diVJgJFC6wKVYbIt4okWOG9aTNjfReLTJm/07eSSj08DP67R+pDutCw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700254; c=relaxed/simple; bh=SrShhO8INsztxN7vjuAEUFnP+ZrFXn1nQ1h3HBereC4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=oEMeOspD9Y+QcdgxPZ1rP+oBEi0sIMtrZvhg914vYjXwHZM4S11U+giLAw74ZMACvfXwY2gZDTBctSR9YtB/ff99WiSCsvzUCu/WhscJ0AIe/1FvvLqz1P5UqFgiO4nChDGS0Kt7sevyqHY23drz0WctRn9wk6DC3iLx7N3RJvM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=L1bU1/PH; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="L1bU1/PH" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 5CB2E1F00A3D; Sun, 2 Aug 2026 19:50:51 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700253; bh=iko4rhiQxzapxJGjwNL1TsJTPSuLholuXSDe7HLrqz4=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=L1bU1/PHGdQ0bwVsc3z969iyZPj7mF62/vw26z/nAO8yf8/mTONNREJmS4mplnSnY EnwZEXmi1FMd8x0wrBFbScOiblWIyj3eq2rOvE0m2/uiAMU+8VuYG3A/RKmnjj/SDt GAaGdcZnVG1nIFDMMd9XLdOP+BocNB/6XhCdOqbJdxWeiBSgstNwzQDcxvpqxB98pD olqiZu0KCQy02nuYc2lNNFk+h8XE0Fz9pRXUaDp+ckrnVrTp7/1XZiaVF61+4j5F/M JiSaa+aPturAiqUU3j3ckhRrN+m10HSrbh5lxSuLyuQZab0of7Jp9mOJpuUL/JgCUf ea8445NtI2gpg== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 02/29] md/md-llbitmap: use GFP_NOIO for cache allocations Date: Mon, 3 Aug 2026 03:50:11 +0800 Message-ID: <20260802195038.164272-3-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai llbitmap allocates its in-memory page cache and page-control structures from paths that can already be holding MD reconfiguration or bitmap state locks. For example, component_size_store() takes mddev_lock(), update_size() calls the personality resize method, and llbitmap_resize() can grow the page cache through llbitmap_prepare_resize(). Using GFP_KERNEL in those paths allows direct reclaim to enter filesystem or block I/O while MD resize state is locked. That can recurse back into the same array and wait on state that cannot make progress until the resize path finishes. Use GFP_NOIO for the llbitmap object, cached bitmap pages, page controls, page-control arrays, and percpu_ref initialization. Leave the explicit metadata zeroout path unchanged because it is intentional bitmap I/O rather than reclaim-driven allocation. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 71e9a21b98b2..3cd8373bc9b2 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -521,7 +521,7 @@ static struct page *llbitmap_read_page(struct llbitmap = *llbitmap, int idx) if (page) return page; =20 - page =3D alloc_page(GFP_KERNEL | __GFP_ZERO); + page =3D alloc_page(GFP_NOIO | __GFP_ZERO); if (!page) return ERR_PTR(-ENOMEM); =20 @@ -616,12 +616,12 @@ static int llbitmap_cache_pages(struct llbitmap *llbi= tmap) int i; =20 llbitmap->pctl =3D kmalloc_array(nr_pages, sizeof(void *), - GFP_KERNEL | __GFP_ZERO); + GFP_NOIO | __GFP_ZERO); if (!llbitmap->pctl) return -ENOMEM; =20 size =3D round_up(size, cache_line_size()); - pctl =3D kmalloc_array(nr_pages, size, GFP_KERNEL | __GFP_ZERO); + pctl =3D kmalloc_array(nr_pages, size, GFP_NOIO | __GFP_ZERO); if (!pctl) { kfree(llbitmap->pctl); return -ENOMEM; @@ -640,7 +640,7 @@ static int llbitmap_cache_pages(struct llbitmap *llbitm= ap) } =20 if (percpu_ref_init(&pctl->active, active_release, - PERCPU_REF_ALLOW_REINIT, GFP_KERNEL)) { + PERCPU_REF_ALLOW_REINIT, GFP_NOIO)) { __free_page(page); llbitmap_free_pages(llbitmap); return -ENOMEM; @@ -1110,7 +1110,7 @@ static int llbitmap_create(struct mddev *mddev) if (ret) return ret; =20 - llbitmap =3D kzalloc_obj(*llbitmap); + llbitmap =3D kzalloc_obj(*llbitmap, GFP_NOIO); if (!llbitmap) return -ENOMEM; =20 --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0AEB42F532F; Sun, 2 Aug 2026 19:50:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700258; cv=none; b=COnwvlh8EmaNZqzjF0IpDMMd7wTI9znsE8UikDYEbBViRvKS++GWcVdnMs6vdZ6d4PoI2Ex989+r2sdflmwaWIrMa6WPYQPSkowx/Zj7s6duj8N2JluQmG5VWPdMgyo6W0pBhiqeOqKH0DgBgVvQoOwoJmUB4KHaH5Rg16rv83Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700258; c=relaxed/simple; bh=zlnyAV+EC9LbxYPZuEb2Cdh64EM8UHC3LuKd9rH5S5U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WGWNiCMxQowMuhCESoIt3rOVaZO9lxduVGfPj8CgsFjkTsmaQjSSPuQDl9yM2z4Uf3PyMkUcYvV2nG9VWgspxwtBp0dT14QtwQhr4v0XD6mfKWPNuaVQB3ms4U2qCj7QrtyJoZEuvZhiwmc3qGZN1tjwySFGIdB/o5vQ6D7Tdmo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=OBb+39g3; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="OBb+39g3" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 0E05B1F000E9; Sun, 2 Aug 2026 19:50:53 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700256; bh=hB1LtmiyPLD6jOR2+Q4BxyiLCBViTo8Wy4k3Jlf8ABw=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=OBb+39g3+hQJ3WNofr3tOfehoNH3SwiHwSpqkO1iifbrGw9bQ8kkBcoq/86qeGUNh hKI0VlwowfAjgrB3C2Rfbr0nqm7HdUSrd3mJj64l1GTLTURdTIsOQAm+HiOBfvj7y8 z+i/3luG3bs3FxMrR5FKES2OQl9lJAEKgHsiU52yLAlUolDnTz5dANVx/G+ks8VHrY duvoFzFxMdVPVFM3+MkFXcbgRf7h6spv4QD4DkwTDCRa8FD+CGW9eqFalW+i/aw231 VfU2qxxZPfAyNmoNEUTWkLQ0+l8IQS4B6bSYOUXLoMFaj0kjv+i05PQgGyZ0dreMN3 nBszwZC/75HsA== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 03/29] md/md-llbitmap: only end fully synced chunks Date: Mon, 3 Aug 2026 03:50:12 +0800 Message-ID: <20260802195038.164272-4-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai llbitmap_cond_end_sync() is called with the sync thread's current sector. That value is an exclusive progress boundary: sectors below it have completed, but the llbitmap chunk containing it can still be in progress. The old code converted that sector directly to the last bit passed to BitmapActionEndsync. If resync had only advanced part-way into a large llbitmap chunk, the in-progress chunk was marked synced and flushed before the rest of the chunk was repaired. A later bitmap-assisted RAID1 resync could then skip the remainder of that chunk and leave stale mirror data behind. This can be reproduced without editing bitmap metadata by creating a large RAID1 with a lockless bitmap so llbitmap naturally selects a 524288-sector chunk (with the default 128 KiB bitmap area, an array just over 16 TiB is enough), making one mirror stale through the normal degraded write/re-add path, and throttling resync so the daemon checkpoint runs while resync is still inside the first chunk. On the bad kernel, bit 0 is ended early and a stale sector later in the same chunk is skipped. With this fix, bit 0 remains Syncing until resync reaches the next chunk boundary. Round the exclusive progress sector down to the nearest llbitmap chunk boundary and end only chunks strictly below that boundary. Also honor the force argument so callers that need an immediate checkpoint are not suppressed by daemon_sleep. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 3cd8373bc9b2..948bf64c5ad2 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -1450,22 +1450,27 @@ static void llbitmap_cond_end_sync(struct mddev *md= dev, sector_t sector, bool force) { struct llbitmap *llbitmap =3D mddev->bitmap; + sector_t complete; =20 if (sector =3D=3D 0) { llbitmap->last_end_sync =3D jiffies; return; } =20 - if (time_before(jiffies, llbitmap->last_end_sync + - HZ * mddev->bitmap_info.daemon_sleep)) + if (!force && time_before(jiffies, llbitmap->last_end_sync + + HZ * mddev->bitmap_info.daemon_sleep)) return; =20 wait_event(mddev->recovery_wait, !atomic_read(&mddev->recovery_active)); =20 mddev->curr_resync_completed =3D sector; set_bit(MD_SB_CHANGE_CLEAN, &mddev->sb_flags); - llbitmap_state_machine(llbitmap, 0, sector >> llbitmap->chunkshift, - BitmapActionEndsync); + + complete =3D round_down(sector, llbitmap->chunksize); + if (complete) + llbitmap_state_machine(llbitmap, 0, + (complete >> llbitmap->chunkshift) - 1, + BitmapActionEndsync); __llbitmap_flush(mddev); =20 llbitmap->last_end_sync =3D jiffies; --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9A60D3148A7; Sun, 2 Aug 2026 19:51:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700261; cv=none; b=sCjDG9LMp1aW6fDOb0vQS0L1OArunZJBzbLdakBQKPxgtlyXTKmJt0ye1i98Zpj2fe4CJ71fTaLBB7vtuNomaF89Axp2v+dzzf2g70iE7gFmfbUYoxBkdkqzYgDNab4oKgkDCNeVaOI4laIrWJd84JvD4gE9vQ9IfOAlBTJ22Dg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700261; c=relaxed/simple; bh=zE8MuNNSsd02gI5QB1ckg1GcfIAg2lMqfGsE4HFnzWg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BQ1/1/Eowov7k7qTxSIQZgOfWz0Kyy/N8V3lDgZPd3QuGB/31aULrXAZe70Kg+GritnV2mSdZXiWmn5spHDxs65qEinUt0N6hTjnAwLNGXnZqTJHQIiEPP/5SGQ5AqU2SN3KstLhrQ8/eAcIhQwQDRji7O3SvozujPRi9geIwIc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=hLE2ZYj8; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="hLE2ZYj8" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 25C1A1F00A3A; Sun, 2 Aug 2026 19:50:56 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700260; bh=K/DY461iIKDMYubR0PpYjGrDymcuRKKSv7V6DaMlsKU=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=hLE2ZYj8vCXxbEuuWvYb8za7w6xm4CoIgaWmUGLxlByfQzbgN8Va6oOMFO5tftUvw lxMqM56r6fwLgTtYZHZu9773tqcsjCPb33USBE/JN+tzjhYbOUCInDLzTE7vCkK3Xy zw9AKr+CuiFruEBE525nz1lewf4xg2+LLXdF64+KiWQz9X8vOgCNYsIqmsUl5xCijQ h6gCPmVUdJHTM1GbXCv0PVyI/90Sbhrnsv7Xhs/kfeSr8b21eCrXWo/fanigfMPL+M jlKMadMR+1/tGun+6C2/JO/vGAyIUuG2b3hkO0smbE115RRP9dzkco9bD3ccMQ8dvs OPpihtwvYZUKg== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 04/29] md/raid5: reject zero-sector reshape chunks Date: Mon, 3 Aug 2026 03:50:13 +0800 Message-ID: <20260802195038.164272-5-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Sashiko reported that RAID5 can accept a reshape chunk size that becomes zero sectors. chunk_size_store() stores the sysfs byte value as n >> 9, so writing a value below 512 bytes sets mddev->new_chunk_sectors to zero. RAID5 then accepted that pending reshape geometry and raid5_start_reshape() installed it into conf->chunk_sectors, letting reshape code divide by zero. Reject zero-sector chunks both in check_reshape(), where normal sysfs requests are validated, and in raid5_start_reshape(), so assembly/resume paths also cannot install zero chunk geometry. Test script: in QEMU, create a plain three-disk RAID5 array with 64K chunks, write/read back a small pattern, write 1 to /sys/block/md0/md/chunk_size, add a fourth disk, and run mdadm --grow --raid-devices=3D4 --backup-file=3D... . The script scans dmesg for divide error/Oops/KASAN signatures. Bad kernel, eb29914412c3: echo 1 > /sys/block/md0/md/chunk_size mdadm --grow /dev/md0 --raid-devices=3D4 --backup-file=3D/root/md0-grow.b= ak Oops: divide error: 0000 [#1] SMP KASAN NOPTI RIP: raid5_get_active_stripe+0x863/0xc10 Call Trace: raid5_sync_request md_do_sync md_thread Kernel panic - not syncing: Fatal exception Fixed kernel: echo 1 > /sys/block/md0/md/chunk_size bash: echo: write error: Invalid argument chunk_write_rc=3D1 grow_rc=3Dskipped RESULT: REJECTED_ZERO_CHUNK_NO_OOPS Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid5.c | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c index e2c5a7072aca..d128d238e1da 100644 --- a/drivers/md/raid5.c +++ b/drivers/md/raid5.c @@ -8548,6 +8548,8 @@ static int check_reshape(struct mddev *mddev) return 0; /* nothing to do */ if (has_failed(conf)) return -EINVAL; + if (!mddev->new_chunk_sectors) + return -EINVAL; if (mddev->delta_disks < 0 && mddev->reshape_position =3D=3D MaxSector) { /* We might be able to shrink, but the devices must * be made bigger first. @@ -8591,6 +8593,9 @@ static int raid5_start_reshape(struct mddev *mddev) if (test_bit(MD_RECOVERY_RUNNING, &mddev->recovery)) return -EBUSY; =20 + if (!mddev->new_chunk_sectors) + return -EINVAL; + if (!check_stripe_cache(mddev)) return -ENOSPC; =20 --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 16DBD3002B9; Sun, 2 Aug 2026 19:51:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700265; cv=none; b=cu7cWP2+H8/PnmufchWm7EBg34ET1Zywyl9fE45OxlmmX0uky4HjjXEj7PuEBT9PRbG2CcH5Y0ozTVlFtPC2VUgnr4z+t7+rKY8dTqhHOEsQc5lt9dZkSqBIQMAseMuBwsv9LxHsVkpX3hL7hQQdelWqRTDwDCEyy6P5Cf7MrM8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700265; c=relaxed/simple; bh=jdfEljajARyuhq8YCaqB22dJovLR3agb9y+AMN33aP8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ui2FAUjdealDlYK6ZzEOShMBcg/GL2EpQR25FUonJbOPHBe/JnBWCkQ3FSpy8wMjY5T2WVeY+CipCCwfGyeS2Lc8SxUx41zfTVE72buZ8VXU1q9vWm3myk71Mf/ayMTL+Pw8xypAH6CAb3LUyi0QJGkwHt2cV2LqxJxsiCWP0Bc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=mLT/YXxD; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="mLT/YXxD" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 9FD081F000E9; Sun, 2 Aug 2026 19:51:00 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700263; bh=wuuaKZBEn/vQSwJHTa4GZ8Ic7lEKoex3pkscJSjGRSk=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=mLT/YXxDI1Wx2QSZccmTiahM9cWu1lwFqe3RfCa8UQrzv7XDCv6nMJ2f51r4URRaR ccTNW3f66BHDZlYiNrOhGR1I2rl+nnlYmVSAUAOjdn3A0QAMpb2Roxz15ws0N2YGCm AIs5KXUA0nq99AdpUU7dqI3zLLazGTFbUSbhAcCdWvbblZLs3y1MfeB+T+8ynu+aaR W8KUnn+vf5jnecIQHOqXZcmthAFYdQeXmzFRVi1PSbemyQe5ER1Y3AsBkJUZAKoTfP 7kT/Vh1FaBt5c4RABEanonLkjmSzK9CpXkz520LmfbUl++zURTcpX+5zt9CU3cQfeS kerxIRZaaR02g== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 05/29] md/raid5: round bitmap stripes with sector division Date: Mon, 3 Aug 2026 03:50:14 +0800 Message-ID: <20260802195038.164272-6-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai raid5_bitmap_sector_map() aligns the array range to full RAID5 stripe widths before converting it to component sectors. That width is chunk_sectors multiplied by the number of data disks, and it is not always a power of two. Reproduce with a 4-disk RAID5, 1024-sector chunks, and three data disks. The full-stripe width is 3072 sectors. For a one-sector write at array sector 3072, correct rounding gives array range [3072, 6144), which maps to component range [1024, 2048). The old round_down()/round_up() logic instead gives [1024, 4096), which maps to [0, 1024). Use sector_div() based arithmetic so the rounded range is aligned to the actual RAID5 stripe width. The deterministic mapper test now reports the fixed component range as [1024, 2048), while the old mask-based range was [0, 1024). Fixes: 9c89f604476c ("md/raid5: implement pers->bitmap_sector()") Reported-by: Mykola Marzhan Link: https://lore.kernel.org/all/20260726185916.2223460-1-mykola@meshstor.= io/ Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid5.c | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c index d128d238e1da..2cc2546a29ae 100644 --- a/drivers/md/raid5.c +++ b/drivers/md/raid5.c @@ -6029,8 +6029,11 @@ static void raid5_bitmap_sector(struct mddev *mddev,= sector_t *offset, =20 sectors_per_chunk =3D conf->chunk_sectors * (conf->raid_disks - conf->max_degraded); - start =3D round_down(start, sectors_per_chunk); - end =3D round_up(end, sectors_per_chunk); + sector_div(start, sectors_per_chunk); + start *=3D sectors_per_chunk; + if (sector_div(end, sectors_per_chunk)) + end++; + end *=3D sectors_per_chunk; =20 start =3D raid5_compute_sector(conf, start, 0, &dd_idx, NULL); end =3D raid5_compute_sector(conf, end, 0, &dd_idx, NULL); @@ -6048,8 +6051,10 @@ static void raid5_bitmap_sector(struct mddev *mddev,= sector_t *offset, =20 sectors_per_chunk =3D conf->prev_chunk_sectors * (conf->previous_raid_disks - conf->max_degraded); - prev_start =3D round_down(prev_start, sectors_per_chunk); - prev_end =3D round_down(prev_end, sectors_per_chunk); + sector_div(prev_start, sectors_per_chunk); + prev_start *=3D sectors_per_chunk; + sector_div(prev_end, sectors_per_chunk); + prev_end *=3D sectors_per_chunk; =20 prev_start =3D raid5_compute_sector(conf, prev_start, 1, &dd_idx, NULL); prev_end =3D raid5_compute_sector(conf, prev_end, 1, &dd_idx, NULL); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AD89D327C00; Sun, 2 Aug 2026 19:51:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700268; cv=none; b=IQ/bQh3sMmyqKrwbTmEha4QMxZimD1xD06jr+XvximzokNYckFurMsdj9DToinbmt1YBzs3mlH4zHDm22vW/9SoaLRAUvapp7mYOMPIV3WyRXpTSzLA5cnRSLShYbpRbUiqKzmaBHDD6T/1Ikf52KYoFI1XGNMB7Dudy7HdUsyg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700268; c=relaxed/simple; bh=rwOzmpg3vR0LnvYLs1uZZ+L8hN8RCJEJ3vJ4wEd9gOo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=O2rcevluc5hm3jxd7py9Tz5k3xyAwkwaKz1vpmB6deNBCzOWzm9HDYZKP+yYUi+Yg4dNNOt31FoxxXUJu0eFW5dyMgy1yprnJoeOSHsFEofQNN3d+1rwHqf4P/CK5SG3hUqAVIN80TXGfpEXxFgICVRL0S/zXvPiZlAVyhRKIJA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=dmXUgvWp; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="dmXUgvWp" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 09A8E1F00A3A; Sun, 2 Aug 2026 19:51:04 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700267; bh=wMFvuuW59mwlS5gcr6gtDyHiuNgmW16AAKd9HbpA/jA=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=dmXUgvWp+rvJS+T42wN6meep5EDS71Cw7rZGKGNTrKIIuXCTIrbC2lUJU1jANQ2us 5sWAV7q6yX4p4abVvzkOQEFqVeo6bx8yZ4dKNDv+7G1cs6MmrJV++1i1+PO4fqT5bw LPjbtc5CYGyxywfb3v7dfO3RCvNKAA/A2IdGMDVta/AtRseJ4GDUPzQGpa5qykBAe/ Gp2QYvWxzpF/oyudRiP+/hGO3Lr5uHvQTtZTxLNoxU7OSH00gbVVpBI7i/t644pDgE 5SDFRt1Ju453wK9/VJdY6odIpczmZwOuAf9aXMdBVbJ+EqExZpeC2ym+tXb6IDIeWi 5Rp1tzom/f7Ng== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 06/29] md: wait for behind writes before destroying bitmap Date: Mon, 3 Aug 2026 03:50:15 +0800 Message-ID: <20260802195038.164272-7-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai __md_stop() destroyed the bitmap before calling mddev_detach(). That made mddev_detach() skip bitmap_ops->wait_behind_writes(), because the bitmap was already disconnected from mddev. This was still safe for the legacy bitmap because bitmap_destroy() waits for behind writes itself. llbitmap keeps that wait in its ->wait_behind_writes() operation instead, while ->destroy() tears down the llbitmap storage. With the old ordering, RAID1 behind-write completions could still run after llbitmap storage had been freed. Call mddev_detach() before md_bitmap_destroy() so the common detach path can wait for behind writes while the bitmap is still alive. Only destroy the bitmap after those users are gone. Fixes: 5ab829f1971d ("md/md-llbitmap: introduce new lockless bitmap") Signed-off-by: Yu Kuai Tested-by: Mykola Marzhan --- drivers/md/md.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/md/md.c b/drivers/md/md.c index 51b620edbef7..b61040315aef 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -7085,8 +7085,8 @@ static void __md_stop(struct mddev *mddev) { struct md_personality *pers =3D mddev->pers; =20 - md_bitmap_destroy(mddev); mddev_detach(mddev); + md_bitmap_destroy(mddev); spin_lock(&mddev->lock); mddev->pers =3D NULL; spin_unlock(&mddev->lock); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B60F62C3268; Sun, 2 Aug 2026 19:51:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700273; cv=none; b=OIUDZGF6hmHKtQka7HkYnQ90YDmnmIbkuHUFK9xuy4ZESvJH1dH2jEzl+R2sYOcn44FeFSCFIY2D/J+oDOc/78CTPpxOyeCBsDhJceRl7qGvGpwG/2elX8aICP+2zyM0c91dmzbc4B1jvGTdEH//fQyG8EL8j5usSz5PQp27rM8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700273; c=relaxed/simple; bh=EG1hSCNgt7Vx9kszVsAxRC0ryiRK4TR1WqZWlcEEexY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=iQPM3yd0XWFz+TFjMogPAlM+4x6/FkDMeJo+mTASyFp2gzqF0A2EquBAalBOyusvI9sYtYQlW2h23DLBrMWELW4aqbgc/6z4V2AbctouJutOrWehAKXEbazdyIBPsdCgraQQwhc1h5mtsIlRd7lvDvsLOtmycCWEJipk3DEanFg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Wcbr6A2s; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Wcbr6A2s" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 652441F000E9; Sun, 2 Aug 2026 19:51:08 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700272; bh=p+dw3vVAIj4YAPqd7MgSRpBgA99rCskitLXGzXREk6s=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=Wcbr6A2sse/VYuF0Kd5yi1m2/3qPhYmSLaE4lzO9Ik5g6SLhlbJcIzEVfr/wbVVQY h+bubM4YEGkYrmDp/spuA2ynG2JwH4ID5CdJ/iROpe8RHYHh4558XjKSWGLYexMC+Z 0oQTEvz7AR0cuGKucAQ1l7LaiUN6A1F9/28vRo68N2fy8npfK3WNh2EtAot0Da2fpk 4Ne49Ifo4akk1xfFd2/uToArE9pL+rKCd4JpcoMmmuTRq1NqtnpdCMbDutVrpEfk0L fPdUhQhIndD6C0xLb31kYj5LFtIWMmLTJZahN0jr0jKOYFPhFHrA7U5BHaLdGe6nrM sggNGR7RG9sbA== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 07/29] md: avoid stale clone I/O accounting timestamps Date: Mon, 3 Aug 2026 03:50:16 +0800 Message-ID: <20260802195038.164272-8-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai md_clone_bio() always allocates the clone from mddev->io_clone_set, even when queue I/O stats are disabled. In that case it does not call bio_start_io_acct(), but it also left md_io_clone->start_time untouched. The clone private data comes from a mempool and can contain data from a previous user. md_end_clone_io() checks start_time to decide whether it needs to call bio_end_io_acct(), so a stale non-zero value can make the completion path end accounting that was never started for this bio. Set start_time to 0 in the no-stats branch. This keeps the end path tied to whether bio_start_io_acct() actually ran. Fixes: c687297b8845 ("md: also clone new io if io accounting is disabled") Signed-off-by: Yu Kuai Tested-by: Mykola Marzhan --- drivers/md/md.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/drivers/md/md.c b/drivers/md/md.c index b61040315aef..58fb5453a819 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -9448,6 +9448,8 @@ static void md_clone_bio(struct mddev *mddev, struct = bio **bio) md_io_clone->mddev =3D mddev; if (blk_queue_io_stat(bdev->bd_disk->queue)) md_io_clone->start_time =3D bio_start_io_acct(*bio); + else + md_io_clone->start_time =3D 0; =20 if (bio_data_dir(*bio) =3D=3D WRITE && md_bitmap_enabled(mddev, false)) { md_io_clone->offset =3D (*bio)->bi_iter.bi_sector; --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 56E1B32ED4E; Sun, 2 Aug 2026 19:51:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700276; cv=none; b=R5trnle5MoejFvw5ACDrW1j1e/Yn1bvt5dDKkKT9WyQxyAz4BBuOnG2VSAyLrisV4JH0fI4gfzB+aCciDhcmlSpggVK2eXO2eFOCZQUp2IeDDfjB8tGhdE6eYH55lB1zLTL46oxHBUPmSPQEFWTsOt+nHzq3mE/EUYRIBgHa2ws= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700276; c=relaxed/simple; bh=zHjud01RmE5/01zjpU7ubwPhSXmrMv6FIQHfJg4x2QA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BOC2ASqQ2YQPKNNaB22jCKXHHa+r77mN3wMklJqqDiGDCd4LHxJIezrZch3Xh6XxuuaiTr/GxGwgDWAbjUR8drZULom2tA2Ncph/Ialr9Iz3PoCrdJzn/+DYOwtTc3JyEXVUe9WbCAWYk3iMfuhLfc0rUdau2/QdGsVWt0la4gA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=bj1dDfZZ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="bj1dDfZZ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DA9FF1F00A3A; Sun, 2 Aug 2026 19:51:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700275; bh=0Y2NONpX13C3hOlDmvNXngKpBN72tSpW1BA7NqTzK00=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=bj1dDfZZ20ZRxkFebDQXa9DKphUFt2/gnAuXoijZgoYMfqAiGkYptqjgfWeQl3B0Z 47uZ8EsaZHSMYvIqpwrmYsUPuVFpJsWd+MCU7kSDWUKYTegI0Zi5SOQ/K3hJMJqRvs EdJMe2KknlxbtMnn3k7ZBFg4vFINvNUM0owdXUNzIZQmsj8GhZ0ejLDH2o7rBYQR72 fmP+D2vfIM6DCvJ/IYhvnVgDs4Vq0bSlcB9Fsih45a0RXtvICtr1e93FQCEEX3Vtbs 5EvAGrovf3meJyj7fqC1BUFvOj5cbyq8f1BGjPSYSUUFTzF/XRDzre5giLkGyFxbj1 pK5rUFo+ZMiOw== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 08/29] md/md-llbitmap: prevent create failure bitmap UAF Date: Mon, 3 Aug 2026 03:50:17 +0800 Message-ID: <20260802195038.164272-9-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai llbitmap_create() publishes mddev->bitmap before reading the bitmap superblock. This is needed because llbitmap_read_sb() can initialize a new bitmap and flush it through helpers that use mddev->bitmap. If llbitmap_read_sb() fails, the old cleanup dropped bitmap_info.mutex and freed llbitmap before clearing mddev->bitmap. Readers such as /proc/mdstat rely on bitmap_info.mutex to keep the bitmap pointer stable while collecting bitmap stats, so they could observe the stale pointer after the failed create path released the mutex. Clear mddev->bitmap while still holding bitmap_info.mutex, then free the failed llbitmap after dropping the mutex. This makes mutex-protected readers see either a live bitmap or no bitmap. Fixes: 5ab829f1971d ("md/md-llbitmap: introduce new lockless bitmap") Signed-off-by: Yu Kuai Tested-by: Mykola Marzhan --- drivers/md/md-llbitmap.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 948bf64c5ad2..af80a630bd21 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -1126,10 +1126,11 @@ static int llbitmap_create(struct mddev *mddev) mutex_lock(&mddev->bitmap_info.mutex); mddev->bitmap =3D llbitmap; ret =3D llbitmap_read_sb(llbitmap); + if (ret) + mddev->bitmap =3D NULL; mutex_unlock(&mddev->bitmap_info.mutex); if (ret) { kfree(llbitmap); - mddev->bitmap =3D NULL; } =20 return ret; --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 737B92C3268; Sun, 2 Aug 2026 19:51:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700279; cv=none; b=onZuMR24UGZ3jfEg0y1ZeU25J1SjIdOHFQkasH9w6SSYobe0nuaqNG9UxCtcw64hYthnNY3fpkA8FhWCeExkuXLoI2tHa8gX7y1QwOHK5LRKJ5eYx56gs2oSyhDernqV5clDXKCQ3UbmWCMJicU/It3oZrT3en6T1zBrkSpZG+c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700279; c=relaxed/simple; bh=j+Bx6hcAl8Jh46CtaKGJbg+bQLl23CqsGe+1aHLbspk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=HJ3H+YGdp6UanLfE0IUu/aHck3d79pGWA6DqG2Ds8muXsTZ3iuFNWKS/ZnJ2wwZQKWCeE7UKQNyPqeVV/Ukrj9ZMeNHpiNwTGAVczyyqK1AAUd9pIIHKyk2bJHvxwBYkIHvezbwVsIjTelGvgalopUWsMm97vxn1Q0CVu1VUDE4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=MwhxCDsp; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="MwhxCDsp" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 8B6E71F000E9; Sun, 2 Aug 2026 19:51:15 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700278; bh=b88ojlcMxi5aXfQhkor5RNl8RmZyAxl1XkqT0+eJGYg=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=MwhxCDspJQOE4jfCZmWFZgObuEi5amSdtFK/7xMOvxgndgzlT4uRWvT3C3O6ZjDoe yj8F34fqQj1r7X1SPVMSgzbc7t+MXO7ZhBY5aTcJQ56HHJKzUoH4qGnl5XOq2wELEL kLg4bjgbJUjIOPA+bcteQR+t+C4a4FDei0cuY0tJnkXKphwFvUDIa8KH70ufIkvbQE NdvvKELJ7goUx80xRX9vsuf8G1wN63DTSvkLy1LXdNsIW1s+f0uA/LzH6DwJLOrUnd uV+j2Bk6+kpLxw1+Ma1y2ub2yKZzcgac7GK4vFrV/fnvi8CrRWU5Gn3EXa+7feG59z E3u9Z0lottz8w== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 09/29] md/md-llbitmap: stop daemon timer rearm on destroy Date: Mon, 3 Aug 2026 03:50:18 +0800 Message-ID: <20260802195038.164272-10-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai llbitmap_destroy() deletes pending_timer before flushing md_llbitmap_io_wq. However, daemon_work can still be queued or running after the timer has been deleted, and the daemon path can arm pending_timer again when it finds dirty chunks that are not ready to flush yet. If that happens during teardown, pending_timer can remain armed after llbitmap is freed and later dereference freed memory. Add a BITMAP_SHUTDOWN bit to llbitmap->flags, set it before deleting the timer, and make the timer and daemon paths stop queueing or rearming work once teardown starts. Cancel daemon_work before flushing the shared workqueue so no already queued daemon instance can race with the free. BITMAP_SHUTDOWN is a runtime-only state. Mask it out when reading and updating the llbitmap superblock so the shutdown state is never loaded from disk or persisted to disk. Fixes: 5ab829f1971d ("md/md-llbitmap: introduce new lockless bitmap") Signed-off-by: Yu Kuai Tested-by: Mykola Marzhan --- drivers/md/md-bitmap.h | 1 + drivers/md/md-llbitmap.c | 15 ++++++++++++--- 2 files changed, 13 insertions(+), 3 deletions(-) diff --git a/drivers/md/md-bitmap.h b/drivers/md/md-bitmap.h index 214f623c7e79..890276d9c66e 100644 --- a/drivers/md/md-bitmap.h +++ b/drivers/md/md-bitmap.h @@ -29,6 +29,7 @@ enum bitmap_state { BITMAP_FIRST_USE =3D 3, /* llbitmap is just created */ BITMAP_CLEAN =3D 4, /* llbitmap is created with assume_clean */ BITMAP_DAEMON_BUSY =3D 5, /* llbitmap daemon is not finished after daemon= _sleep */ + BITMAP_SHUTDOWN =3D 6, /* llbitmap is being destroyed */ BITMAP_HOSTENDIAN =3D15, }; =20 diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index af80a630bd21..3ec5b5985d48 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -789,6 +789,7 @@ static enum llbitmap_state llbitmap_state_machine(struc= t llbitmap *llbitmap, if (state =3D=3D BitNeedSync || state =3D=3D BitNeedSyncUnwritten) need_resync =3D !mddev->degraded; else if (state =3D=3D BitDirty && + !test_bit(BITMAP_SHUTDOWN, &llbitmap->flags) && !timer_pending(&llbitmap->pending_timer)) mod_timer(&llbitmap->pending_timer, jiffies + mddev->bitmap_info.daemon_sleep * HZ); @@ -981,7 +982,7 @@ static int llbitmap_read_sb(struct llbitmap *llbitmap) else mddev->bitmap_info.space =3D mddev->bitmap_info.default_space; } - llbitmap->flags =3D le32_to_cpu(sb->state); + llbitmap->flags =3D le32_to_cpu(sb->state) & ~BIT(BITMAP_SHUTDOWN); if (test_and_clear_bit(BITMAP_FIRST_USE, &llbitmap->flags)) { ret =3D llbitmap_init(llbitmap); goto out_put_page; @@ -1037,6 +1038,9 @@ static void llbitmap_pending_timer_fn(struct timer_li= st *pending_timer) struct llbitmap *llbitmap =3D container_of(pending_timer, struct llbitmap, pending_timer); =20 + if (test_bit(BITMAP_SHUTDOWN, &llbitmap->flags)) + return; + if (work_busy(&llbitmap->daemon_work)) { pr_warn("md/llbitmap: %s daemon_work not finished in %lu seconds\n", mdname(llbitmap->mddev), @@ -1057,6 +1061,9 @@ static void md_llbitmap_daemon_fn(struct work_struct = *work) bool restart; int idx; =20 + if (test_bit(BITMAP_SHUTDOWN, &llbitmap->flags)) + return; + if (llbitmap->mddev->degraded) return; retry: @@ -1096,7 +1103,7 @@ static void md_llbitmap_daemon_fn(struct work_struct = *work) goto retry; =20 /* If some page is dirty but not expired, setup timer again */ - if (restart) + if (restart && !test_bit(BITMAP_SHUTDOWN, &llbitmap->flags)) mod_timer(&llbitmap->pending_timer, jiffies + llbitmap->mddev->bitmap_info.daemon_sleep * HZ); } @@ -1179,7 +1186,9 @@ static void llbitmap_destroy(struct mddev *mddev) =20 mutex_lock(&mddev->bitmap_info.mutex); =20 + set_bit(BITMAP_SHUTDOWN, &llbitmap->flags); timer_delete_sync(&llbitmap->pending_timer); + cancel_work_sync(&llbitmap->daemon_work); flush_workqueue(md_llbitmap_io_wq); flush_workqueue(md_llbitmap_unplug_wq); =20 @@ -1523,7 +1532,7 @@ static void llbitmap_update_sb(void *data) =20 sb =3D kmap_local_page(sb_page); sb->events =3D cpu_to_le64(mddev->events); - sb->state =3D cpu_to_le32(llbitmap->flags); + sb->state =3D cpu_to_le32(llbitmap->flags & ~BIT(BITMAP_SHUTDOWN)); sb->chunksize =3D cpu_to_le32(llbitmap->chunksize); sb->sync_size =3D cpu_to_le64(mddev->resync_max_sectors); sb->events_cleared =3D cpu_to_le64(llbitmap->events_cleared); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AFC2D3002B9; Sun, 2 Aug 2026 19:51:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700282; cv=none; b=FBvNprsLlRu3dl9viHu97tWp9jXAsHHOb3Q+VfElBCUDgbl2RHKVXD8CSXGjsTjVqogsoesOIqXFmbztz3PROBWA0YIG6QkshyUL/ON/jVZDA38hRXlRiPtT+mySOPZJD9VNWDWjM6tULysEmV2Fs+6RXE6lv6IXVmbv2sJGOMg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700282; c=relaxed/simple; bh=HyHyVAG3vF6ASrmdXcmDTCd3e/ee9FqThAwpyq5efgc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=HGrn6IsvFVEECCcMhPuAStQ5MBvgLYib41H9KTAI+hxA8PLP+iJ5OwHU0prEZKJMd2vhXGlntwLemhshcVhus2HHtuig4XpbH8xmB1HUk5eg83wewC7oXWfgcViHuVCbzbwSxEHAgZsN23W4vL/YD3P0cLmDRKWmKJvXYf9DOHI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=R1RvNJg2; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="R1RvNJg2" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A35B71F00A3A; Sun, 2 Aug 2026 19:51:18 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700281; bh=PNiC5Z7vUv7Ch8OzLEpTUN8kFtKdUAI0tmDK3WigLE0=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=R1RvNJg2aN7BszuhBWdGXlsNdiVDfqGHsSUT8yTbTvj3NHOWs1SKmWqbzETgJeT1l HtI0V3eH16qjSu4vFw+5CiEmwBjYCWRHUa5Ds9vKTxF06bRWIUQiq+jMI8/IL2SCG/ go5/zyLGAgvwHHaLjpmXKSg8EEbogWiMeq2TIdF7iZwcIilzYNd+rpDFWe9gkFPlPc 5GbBIqRWtVTstN7RpdeoUpFyg0VZKoePwu31n4iGv0y4XYysA/oBZ/yq7Id0fyX4vq LvBQxHGXp1fEUpoTyhCwKtSc4Fq+mlOuF9rI6ovsdWSo/yuSITiQ/4PQ+6vb6eyZ+c 6u5muwi13q8RA== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 10/29] md: skip bitmap accounting for empty write ranges Date: Mon, 3 Aug 2026 03:50:19 +0800 Message-ID: <20260802195038.164272-11-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai mkfs.ext4 can submit zero-sector flush/FUA bios. These bios are WRITE bios for md_write_start() purposes, but they do not cover any data sector and must not dirty bitmap bits. md bitmap accounting currently passes such bios to bitmap start_write(). For llbitmap this reaches llbitmap_start_write() with sectors =3D=3D 0, which underflows the end chunk calculation. Personality bitmap mapping can also turn a non-empty bio into an empty bitmap range when the requested sectors are outside the active bitmap geometry. Treat both cases as not started, so the completion path will not call end_write() for an empty range. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md.c | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-) diff --git a/drivers/md/md.c b/drivers/md/md.c index 58fb5453a819..f88952371b9b 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -9399,6 +9399,8 @@ static void md_bitmap_start(struct mddev *mddev, mddev->pers->bitmap_sector(mddev, &md_io_clone->offset, &md_io_clone->sectors); =20 + if (!md_io_clone->sectors) + return; fn(mddev, md_io_clone->offset, md_io_clone->sectors); } =20 @@ -9419,7 +9421,8 @@ static void md_end_clone_io(struct bio *bio) struct mddev *mddev =3D md_io_clone->mddev; struct completion *reshape_completion =3D bio->bi_private; =20 - if (bio_data_dir(orig_bio) =3D=3D WRITE && md_bitmap_enabled(mddev, false= )) + if (bio_data_dir(orig_bio) =3D=3D WRITE && md_io_clone->sectors && + md_bitmap_enabled(mddev, false)) md_bitmap_end(mddev, md_io_clone); =20 if (bio->bi_status && !orig_bio->bi_status) @@ -9446,12 +9449,14 @@ static void md_clone_bio(struct mddev *mddev, struc= t bio **bio) md_io_clone =3D container_of(clone, struct md_io_clone, bio_clone); md_io_clone->orig_bio =3D *bio; md_io_clone->mddev =3D mddev; + md_io_clone->sectors =3D 0; if (blk_queue_io_stat(bdev->bd_disk->queue)) md_io_clone->start_time =3D bio_start_io_acct(*bio); else md_io_clone->start_time =3D 0; =20 - if (bio_data_dir(*bio) =3D=3D WRITE && md_bitmap_enabled(mddev, false)) { + if (bio_data_dir(*bio) =3D=3D WRITE && bio_sectors(*bio) && + md_bitmap_enabled(mddev, false)) { md_io_clone->offset =3D (*bio)->bi_iter.bi_sector; md_io_clone->sectors =3D bio_sectors(*bio); md_io_clone->rw =3D op_stat_group(bio_op(*bio)); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 792A32DA756; Sun, 2 Aug 2026 19:51:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700287; cv=none; b=IPqqwk84LpHUWUiIyeo9+no8+vcNso9PwFfIQu/zAtaYGwpRP7b4kvCW8yNL+DytrJ7qZfxd/sjMK6D15QBDRuOQFxhfU8jrwAWzuhL3701SEtf4/nio4UFN9y8YneUKXp6iZQzajEuKZ9yz8luHV8qc2+P0PWMRNVtzCSs7rmg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700287; c=relaxed/simple; bh=29Ip17HN0CEdcOr3hVd1G1H1+LsOmFwnkXLsbJdK4Q4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=jqzWYfcldqErFf+zLRaUhM0lUks32QRb08579hM0D+x2SjX0wnsGDqA/zwlXlsDKHFOtUNG4FpR4Q5+dWtL9EHZqjiDAhza0JT/NNpAPG9NG7Ulrk0PgDjHJ86WtYNJ3LH6304sBMP5W9w86mKzbTntPo0Yf1/eBEhZe8iQD6a8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=nA1VLz34; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="nA1VLz34" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 8D0B51F000E9; Sun, 2 Aug 2026 19:51:21 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700286; bh=vKVPXf+g/byhY7EZrFpGCyHm1NZyHDDuQiKiy08ni8A=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=nA1VLz34tITq2ZUFqkxM9M2XfazZAYpVGo/4z17QXtgQd0U7XUfcdKdCfSwlU8LEk LLWSjKqKzdo65EDoYInAkTI3JPvFyfHbizuj7hJfB4qo9ePacZ/H/y/cyOQpDZ4dmf hY3mMXs6KpcFCQbTfDvedP1MkYu3ojmcOTTYBB1izyLZ2zhW3RXv4h8LiyH04ismwx GOVhSBAClqm+f9aH91aYDij3tNY5Ra+8sEH9Q8MsCEfcSnByMjEYF96fU+fuLRC9R1 h3zEacQWSYYN8qprlt3bb5Aca5ASIsAmwxUk+bpNGeqTp985ALKJVXvlGUrWY4xdqy iGtG5vWByk6kQ== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 11/29] md: add helper to split bios at reshape offset Date: Mon, 3 Aug 2026 03:50:20 +0800 Message-ID: <20260802195038.164272-12-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Add mddev_bio_split_at_reshape_offset() so personalities can share reshape-offset bio splitting instead of open-coding the same boundary handling in multiple places. The helper first applies the optional max_sectors limit. If reshape is running and the bio crosses reshape_position, it further limits the front bio to the current reshape boundary so callers can account and submit one side of the reshape at a time. Snapshot reshape_position with READ_ONCE(). RAID5 and RAID10 update this field as reshape progresses, while the I/O path only needs one consistent decision point for the current bio. Using an explicit single load avoids a plain lockless access and prevents the compiler from refetching a different boundary while deciding whether and where to split. When a split is needed, bio_submit_split_bioset() submits the remainder and returns the front bio. Callers must therefore continue processing the returned bio, not the original pointer. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md.c | 39 +++++++++++++++++++++++++++++++++++++++ drivers/md/md.h | 4 ++++ 2 files changed, 43 insertions(+) diff --git a/drivers/md/md.c b/drivers/md/md.c index f88952371b9b..f0eecdfff1cc 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -9388,6 +9388,45 @@ void md_submit_discard_bio(struct mddev *mddev, stru= ct md_rdev *rdev, } EXPORT_SYMBOL_GPL(md_submit_discard_bio); =20 +struct bio *mddev_bio_split_at_reshape_offset(struct mddev *mddev, + struct bio *bio, + unsigned int *max_sectors, + struct bio_set *bs) +{ + sector_t boundary; + sector_t start; + sector_t end; + unsigned int split_sectors; + + split_sectors =3D bio_sectors(bio); + if (max_sectors && *max_sectors && *max_sectors < split_sectors) + split_sectors =3D *max_sectors; + + if (!test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery)) + goto split; + + boundary =3D READ_ONCE(mddev->reshape_position); + start =3D bio->bi_iter.bi_sector; + end =3D bio_end_sector(bio); + if (start >=3D boundary || end <=3D boundary) + goto split; + + if (boundary - start < split_sectors) + split_sectors =3D boundary - start; + +split: + if (max_sectors) + *max_sectors =3D split_sectors; + if (split_sectors < bio_sectors(bio)) { + bio =3D bio_submit_split_bioset(bio, split_sectors, bs); + if (bio) + bio->bi_opf |=3D REQ_NOMERGE; + } + + return bio; +} +EXPORT_SYMBOL_GPL(mddev_bio_split_at_reshape_offset); + static void md_bitmap_start(struct mddev *mddev, struct md_io_clone *md_io_clone) { diff --git a/drivers/md/md.h b/drivers/md/md.h index bb2eb5f39914..8146a6f50a7d 100644 --- a/drivers/md/md.h +++ b/drivers/md/md.h @@ -920,6 +920,10 @@ extern void md_error(struct mddev *mddev, struct md_rd= ev *rdev); extern void md_finish_reshape(struct mddev *mddev); void md_submit_discard_bio(struct mddev *mddev, struct md_rdev *rdev, struct bio *bio, sector_t start, sector_t size); +struct bio *mddev_bio_split_at_reshape_offset(struct mddev *mddev, + struct bio *bio, + unsigned int *max_sectors, + struct bio_set *bs); void md_account_bio(struct mddev *mddev, struct bio **bio); =20 extern bool __must_check md_flush_request(struct mddev *mddev, struct bio = *bio); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B154F2C3268; Sun, 2 Aug 2026 19:51:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700291; cv=none; b=iTQyRaNWwfyVDNC8ZSQX7mXIMTjKrgtoBHcS+EI0/z+hSdHBCodI1/v2FSpQEX8tO0++T6ieg+4iERRzTgr7O2P0Wb1iOEd36XMWgd76TzWXaMRmF+Ncw+xJndYqn5A+xeORaRP4mHxejMVSnyLa5jipFA2osEXMurcrptWNtMo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700291; c=relaxed/simple; bh=1p2nJTQOKn8idEgBa6FWCSV1o/It0iDsnvAKzXSbZkI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ujPeAUOth5aC2bpB1sBjXfqL9qdB2n6CS936ZVVAbqcZ544cYBDv0rsFmWw1+ixMbqblOcn4gPoERcfBF1gZwSZrfaxx4g0Ju/4gqcHJOBtijIPCz8ytkq5xPDllHEa8aY5GTlzcuzOHEC0sv+5l9R7Yb2eyENt4j2sR0TKDypo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=bGIWH899; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="bGIWH899" Received: by smtp.kernel.org (Postfix) with ESMTPSA id BD0F71F000E9; Sun, 2 Aug 2026 19:51:27 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700290; bh=qZ4vXAPCYvJ89FeEjdHZT7m4iU7j8fXxTea+mwBzEFE=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=bGIWH899bok3l3MsVUWXGgXEf+jHhhajZTeDT5MO/V4Wop0Vq9BgnY392Vw2mHBvp Wjse7NyC5g+W60ObXvnwFiRvpr/3cbZys5IaFlfJzSWvP08vwIwIsI8G8/QeKLhDPL VlLD81bpgK9+sQH2XQVFGvxAxDP8k79OFGuPj4/TbCcUON8KnzOrXieoYWtsBumB8u a50Q/87oMB4Y/azQctCkWArzA/OFnjf/QY91YbPybah3dJL7lo0hd0s/If08b7ChV6 0R68u3vAhkp72PRxkd133UtA5HVpJwM7kZIGqlnn+a7bmX7dGGEs0Tuh6EK81Xj7FO Q2HnGD2dDaGZQ== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 12/29] md: add exact bitmap mapping and reshape hooks Date: Mon, 3 Aug 2026 03:50:21 +0800 Message-ID: <20260802195038.164272-13-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Add bitmap mapping and reshape hooks needed by llbitmap reshape support without teaching md core to account a single bio against multiple bitmap ranges. This also adds the old/new bitmap geometry helpers used by personalities to describe reshape mapping to llbitmap. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-bitmap.c | 8 ++++++++ drivers/md/md-bitmap.h | 8 ++++++++ drivers/md/md-llbitmap.c | 8 ++++++++ drivers/md/md.c | 11 ++++++++--- drivers/md/md.h | 4 ++++ 5 files changed, 36 insertions(+), 3 deletions(-) diff --git a/drivers/md/md-bitmap.c b/drivers/md/md-bitmap.c index 7e4fbca93ccb..b7f0d4acce04 100644 --- a/drivers/md/md-bitmap.c +++ b/drivers/md/md-bitmap.c @@ -1730,6 +1730,13 @@ static void bitmap_start_write(struct mddev *mddev, = sector_t offset, } } =20 +static void bitmap_prepare_range(struct mddev *mddev, sector_t *offset, + unsigned long *sectors) +{ + if (mddev->pers->bitmap_sector) + mddev->pers->bitmap_sector(mddev, offset, sectors); +} + static void bitmap_end_write(struct mddev *mddev, sector_t offset, unsigned long sectors) { @@ -3081,6 +3088,7 @@ static struct bitmap_operations bitmap_ops =3D { .flush =3D bitmap_flush, .write_all =3D bitmap_write_all, .dirty_bits =3D bitmap_dirty_bits, + .prepare_range =3D bitmap_prepare_range, .unplug =3D bitmap_unplug, .daemon_work =3D bitmap_daemon_work, =20 diff --git a/drivers/md/md-bitmap.h b/drivers/md/md-bitmap.h index 890276d9c66e..6478cf9d8816 100644 --- a/drivers/md/md-bitmap.h +++ b/drivers/md/md-bitmap.h @@ -94,6 +94,14 @@ struct bitmap_operations { void (*write_all)(struct mddev *mddev); void (*dirty_bits)(struct mddev *mddev, unsigned long s, unsigned long e); + /* Prepare a range for this bitmap implementation. */ + void (*prepare_range)(struct mddev *mddev, + sector_t *offset, + unsigned long *sectors); + void (*reshape_finish)(struct mddev *mddev); + int (*reshape_can_start)(struct mddev *mddev); + void (*reshape_mark)(struct mddev *mddev, sector_t old_pos, + sector_t new_pos); void (*unplug)(struct mddev *mddev, bool sync); void (*daemon_work)(struct mddev *mddev); =20 diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 3ec5b5985d48..4583bbc37c2e 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -1198,6 +1198,13 @@ static void llbitmap_destroy(struct mddev *mddev) mutex_unlock(&mddev->bitmap_info.mutex); } =20 +static void llbitmap_prepare_range(struct mddev *mddev, sector_t *offset, + unsigned long *sectors) +{ + if (mddev->pers->bitmap_sector) + mddev->pers->bitmap_sector(mddev, offset, sectors); +} + static void llbitmap_start_write(struct mddev *mddev, sector_t offset, unsigned long sectors) { @@ -1789,6 +1796,7 @@ static struct bitmap_operations llbitmap_ops =3D { .update_sb =3D llbitmap_update_sb, .get_stats =3D llbitmap_get_stats, .dirty_bits =3D llbitmap_dirty_bits, + .prepare_range =3D llbitmap_prepare_range, .write_all =3D llbitmap_write_all, =20 .groups =3D md_llbitmap_groups, diff --git a/drivers/md/md.c b/drivers/md/md.c index f0eecdfff1cc..538ba7bab060 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -9427,6 +9427,12 @@ struct bio *mddev_bio_split_at_reshape_offset(struct= mddev *mddev, } EXPORT_SYMBOL_GPL(mddev_bio_split_at_reshape_offset); =20 +static void md_bitmap_prepare_range(struct mddev *mddev, sector_t *offset, + unsigned long *sectors) +{ + mddev->bitmap_ops->prepare_range(mddev, offset, sectors); +} + static void md_bitmap_start(struct mddev *mddev, struct md_io_clone *md_io_clone) { @@ -9434,9 +9440,8 @@ static void md_bitmap_start(struct mddev *mddev, mddev->bitmap_ops->start_discard : mddev->bitmap_ops->start_write; =20 - if (mddev->pers->bitmap_sector) - mddev->pers->bitmap_sector(mddev, &md_io_clone->offset, - &md_io_clone->sectors); + md_bitmap_prepare_range(mddev, &md_io_clone->offset, + &md_io_clone->sectors); =20 if (!md_io_clone->sectors) return; diff --git a/drivers/md/md.h b/drivers/md/md.h index 8146a6f50a7d..b6d2e8929a0f 100644 --- a/drivers/md/md.h +++ b/drivers/md/md.h @@ -797,6 +797,10 @@ struct md_personality /* convert io ranges from array to bitmap */ void (*bitmap_sector)(struct mddev *mddev, sector_t *offset, unsigned long *sectors); + void (*bitmap_sector_map)(struct mddev *mddev, sector_t *offset, + unsigned long *sectors, bool previous); + sector_t (*bitmap_sync_size)(struct mddev *mddev, bool previous); + sector_t (*bitmap_array_sectors)(struct mddev *mddev, bool previous); }; =20 struct md_sysfs_entry { --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BABAA33438F; Sun, 2 Aug 2026 19:51:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700295; cv=none; b=i/nbWjyb/6/I347yBMfD5191iqZsJzpGZRTzkDQHzrAs6YF3SF1QANfeV7RjFH2qO9gn2+GOg3nYs3bOISSxb/OPtmLHY44Vw2fTCrH5X+hYFFoIzh5K6aDz7DDEdTjczBTaPdJksk/hsVwO1mOJZdvcsn+bkb4et4meULoV0BU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700295; c=relaxed/simple; bh=1fgKUBcxD4P8+lhYLHQSlrFUNbopgHPT3RBPVc1bQtk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=B/8Diz3kBfFqTkdYQiHu/gK2PN9nWRWnDcDNkEAXH9IfM4o+RZq/K6GemWLyBZ4mS5cy6woj301ibK5Gh95n+BwVetRu+9HCQN2YhkxQUxCFYnUVQm205IX+cDT9zm2e4N97ENVKdkXEX7FYsXjHNB7dos2XrigoQhZM/SRPjG8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=B1DS+0Xh; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="B1DS+0Xh" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 49E5A1F00A3A; Sun, 2 Aug 2026 19:51:31 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700293; bh=PmCAlxRFFGz9k8RIIZAA8jwmjrvbF5WiJfnuAFxFCHk=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=B1DS+0Xha3xJwURwEzMbTfxSecJ191iXlDK+v3Ty4U7HRAx0UcKXnvC+tjYtNgamN Q+gJDga/yeYTfoQmfL74Q7kzzl/8VO5iZ8waqBX7aEkq6jg/fki0Ikz6EITsI6xJpC 6rqDYY9IYfFDp+H/aNaZsYgh7ujxNhxWjagommyscHV47bNBi/IZeuJbXJLjpBbJgE 5u4Vx+TWa8dvnG1aCWXkf/AvVu43wG0SI73zF00l8qpHbR/Tq6z23TG7x0WQdJuK3V Mla/flHudBhP/5YrQKYMO1oyLiBcjkB1KQrrUMsTUu4iZQaM9wp4UqnKj7SpVPmICC b7h6Xj2Rv0dag== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 13/29] md/md-llbitmap: track bitmap sync_size explicitly Date: Mon, 3 Aug 2026 03:50:22 +0800 Message-ID: <20260802195038.164272-14-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Track llbitmap's own sync_size instead of always using mddev->resync_max_sectors directly. This is the minimal bookkeeping needed before llbitmap can track old and new reshape geometry independently. Reviewed-by: Su Yue Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 19 +++++++++++++++++-- 1 file changed, 17 insertions(+), 2 deletions(-) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 4583bbc37c2e..0813cebfbdeb 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -287,6 +287,8 @@ struct llbitmap { unsigned long chunksize; /* total number of chunks */ unsigned long chunks; + /* total number of sectors tracked by current bitmap geometry */ + sector_t sync_size; unsigned long last_end_sync; /* * time in seconds that dirty bits will be cleared if the page is not @@ -919,6 +921,7 @@ static int llbitmap_init(struct llbitmap *llbitmap) llbitmap->chunkshift =3D ffz(~chunksize); llbitmap->chunksize =3D chunksize; llbitmap->chunks =3D chunks; + llbitmap->sync_size =3D blocks; mddev->bitmap_info.daemon_sleep =3D DEFAULT_DAEMON_SLEEP; =20 ret =3D llbitmap_cache_pages(llbitmap); @@ -939,6 +942,7 @@ static int llbitmap_read_sb(struct llbitmap *llbitmap) unsigned long daemon_sleep; unsigned long chunksize; unsigned long events; + sector_t sync_size; struct page *sb_page; bitmap_super_t *sb; int ret =3D -EINVAL; @@ -988,6 +992,14 @@ static int llbitmap_read_sb(struct llbitmap *llbitmap) goto out_put_page; } =20 + sync_size =3D le64_to_cpu(sb->sync_size); + if (!sync_size) + sync_size =3D mddev->resync_max_sectors; + if (sync_size > mddev->resync_max_sectors) { + pr_err("md/llbitmap: %s: sync_size %llu exceeds array sync size %llu", + mdname(mddev), sync_size, mddev->resync_max_sectors); + goto out_put_page; + } chunksize =3D le32_to_cpu(sb->chunksize); if (!is_power_of_2(chunksize)) { pr_err("md/llbitmap: %s: chunksize not a power of 2", @@ -1023,8 +1035,9 @@ static int llbitmap_read_sb(struct llbitmap *llbitmap) =20 llbitmap->barrier_idle =3D DEFAULT_BARRIER_IDLE; llbitmap->chunksize =3D chunksize; - llbitmap->chunks =3D DIV_ROUND_UP_SECTOR_T(mddev->resync_max_sectors, chu= nksize); + llbitmap->chunks =3D DIV_ROUND_UP_SECTOR_T(sync_size, chunksize); llbitmap->chunkshift =3D ffz(~chunksize); + llbitmap->sync_size =3D sync_size; ret =3D llbitmap_cache_pages(llbitmap); =20 out_put_page: @@ -1161,6 +1174,7 @@ static int llbitmap_resize(struct mddev *mddev, secto= r_t blocks, int chunksize) llbitmap->chunkshift =3D ffz(~chunksize); llbitmap->chunksize =3D chunksize; llbitmap->chunks =3D chunks; + llbitmap->sync_size =3D blocks; =20 return 0; } @@ -1541,7 +1555,7 @@ static void llbitmap_update_sb(void *data) sb->events =3D cpu_to_le64(mddev->events); sb->state =3D cpu_to_le32(llbitmap->flags & ~BIT(BITMAP_SHUTDOWN)); sb->chunksize =3D cpu_to_le32(llbitmap->chunksize); - sb->sync_size =3D cpu_to_le64(mddev->resync_max_sectors); + sb->sync_size =3D cpu_to_le64(llbitmap->sync_size); sb->events_cleared =3D cpu_to_le64(llbitmap->events_cleared); sb->sectors_reserved =3D cpu_to_le32(mddev->bitmap_info.space); sb->daemon_sleep =3D cpu_to_le32(mddev->bitmap_info.daemon_sleep); @@ -1559,6 +1573,7 @@ static int llbitmap_get_stats(void *data, struct md_b= itmap_stats *stats) stats->missing_pages =3D 0; stats->pages =3D llbitmap->nr_pages; stats->file_pages =3D llbitmap->nr_pages; + stats->sync_size =3D llbitmap->sync_size; =20 stats->behind_writes =3D atomic_read(&llbitmap->behind_writes); stats->behind_wait =3D wq_has_sleeper(&llbitmap->behind_wait); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CF02A2F1FEA; Sun, 2 Aug 2026 19:51:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700300; cv=none; b=bOb2DXAFg7EOfnMnsJd+VkrErYqPhaHbge5PxE18lb4GMlFaWhduRyInzgS2CrFveD6HiVKSX9VzSOxKJNKf8XOy77waa599j/c7/U6Q6NYwjX/BrAoNGuATmjVCD6TxGLghP9e1E3rZpGImP5wz4+Rmghif5AwogVRtFhyv0jU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700300; c=relaxed/simple; bh=c1vYLYUv45quP/gFZnFqwcW3id/Y6lrfHIFoIdy8k2w=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=IdEXBb9o/NQ0Uvhc4XGDVm3A+p2dURTsLIdnBe5nzAbNtIhlgfOUzJUawF34fKidCR4r5lXMmEtXJs7kgaeNkkNfmCodRzLavm8z5sjLsoSlR2relOzgRzOYIqfBssC07/0b448nOhavwioFB4M1ytUDsKfrQMH2+8OPfW8Jn1o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=R2cHPt0z; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="R2cHPt0z" Received: by smtp.kernel.org (Postfix) with ESMTPSA id ECAD51F000E9; Sun, 2 Aug 2026 19:51:33 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700298; bh=Pyxb5ixBGyVEK4ViIPlMA4hkma7e+3/Jmkws8oBWmEA=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=R2cHPt0zUJV9osj3JZbYbOGNyj19EtXWojRomGOBELQmQAxadA4oFy9jKPeG0khWJ bD2I2hcrjK+ny3fH4FpcIpjog0blyOSiwkJ/U9QoCYSfl9td1DFKb1pV3+Fajzmvkt mqlWBSyvtS0Thsro9ORDyOhZEXF49NtkeYeaYdsIbvQs+ODncegUXTXyrJKPW1EExA TzfYteJVrzCfWhRXxlk25N2PSyWCbKe2AoKJ9IhqKjXejp3DrAd2QeT2MsdRq66NXd l2UdaEtGSxcld/Bz1DjHoktj6Vq4/u5bvs0bn07QbfaVdNAko5IlZMii/iXXQSK4IR ZGHSPVcf3xyhw== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 14/29] md/md-llbitmap: allocate page controls independently Date: Mon, 3 Aug 2026 03:50:23 +0800 Message-ID: <20260802195038.164272-15-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Allocate one llbitmap page-control object at a time and free each object through the same model. Let llbitmap_read_page() return a zeroed page without reading disk when the page index is beyond the current bitmap size, so page-control allocation no longer needs a separate read_existing flag. This keeps the llbitmap page-control lifetime self-consistent and prepares the page-cache code for later in-place growth. Reviewed-by: Su Yue Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 99 +++++++++++++++++++++++++--------------- 1 file changed, 62 insertions(+), 37 deletions(-) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 0813cebfbdeb..300dd8b93b01 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -512,13 +512,19 @@ static void llbitmap_write(struct llbitmap *llbitmap,= enum llbitmap_state state, llbitmap_set_page_dirty(llbitmap, idx, bit, false); } =20 +static unsigned int llbitmap_used_pages(struct llbitmap *llbitmap, + unsigned long chunks) +{ + return DIV_ROUND_UP(chunks + BITMAP_DATA_OFFSET, PAGE_SIZE); +} + static struct page *llbitmap_read_page(struct llbitmap *llbitmap, int idx) { struct mddev *mddev =3D llbitmap->mddev; struct page *page =3D NULL; struct md_rdev *rdev; =20 - if (llbitmap->pctl && llbitmap->pctl[idx]) + if (llbitmap->pctl && idx < llbitmap->nr_pages && llbitmap->pctl[idx]) page =3D llbitmap->pctl[idx]->page; if (page) return page; @@ -526,6 +532,8 @@ static struct page *llbitmap_read_page(struct llbitmap = *llbitmap, int idx) page =3D alloc_page(GFP_NOIO | __GFP_ZERO); if (!page) return ERR_PTR(-ENOMEM); + if (idx >=3D llbitmap_used_pages(llbitmap, llbitmap->chunks)) + return page; =20 rdev_for_each(rdev, mddev) { sector_t sector; @@ -596,61 +604,78 @@ static void llbitmap_free_pages(struct llbitmap *llbi= tmap) for (i =3D 0; i < llbitmap->nr_pages; i++) { struct llbitmap_page_ctl *pctl =3D llbitmap->pctl[i]; =20 - if (!pctl || !pctl->page) - break; - - __free_page(pctl->page); + if (!pctl) + continue; + if (pctl->page) + __free_page(pctl->page); percpu_ref_exit(&pctl->active); + kfree(pctl); } =20 - kfree(llbitmap->pctl[0]); kfree(llbitmap->pctl); llbitmap->pctl =3D NULL; } =20 -static int llbitmap_cache_pages(struct llbitmap *llbitmap) +static struct llbitmap_page_ctl * +llbitmap_alloc_page_ctl(struct llbitmap *llbitmap, int idx) { struct llbitmap_page_ctl *pctl; - unsigned int nr_pages =3D DIV_ROUND_UP(llbitmap->chunks + - BITMAP_DATA_OFFSET, PAGE_SIZE); + struct page *page; unsigned int size =3D struct_size(pctl, dirty, BITS_TO_LONGS( llbitmap->blocks_per_page)); - int i; - - llbitmap->pctl =3D kmalloc_array(nr_pages, sizeof(void *), - GFP_NOIO | __GFP_ZERO); - if (!llbitmap->pctl) - return -ENOMEM; =20 size =3D round_up(size, cache_line_size()); - pctl =3D kmalloc_array(nr_pages, size, GFP_NOIO | __GFP_ZERO); - if (!pctl) { - kfree(llbitmap->pctl); - return -ENOMEM; + pctl =3D kzalloc(size, GFP_NOIO); + if (!pctl) + return ERR_PTR(-ENOMEM); + + page =3D llbitmap_read_page(llbitmap, idx); + + if (IS_ERR(page)) { + kfree(pctl); + return ERR_CAST(page); } =20 - llbitmap->nr_pages =3D nr_pages; + if (percpu_ref_init(&pctl->active, active_release, + PERCPU_REF_ALLOW_REINIT, GFP_NOIO)) { + __free_page(page); + kfree(pctl); + return ERR_PTR(-ENOMEM); + } =20 - for (i =3D 0; i < nr_pages; i++, pctl =3D (void *)pctl + size) { - struct page *page =3D llbitmap_read_page(llbitmap, i); + pctl->page =3D page; + pctl->state =3D page_address(page); + init_waitqueue_head(&pctl->wait); + return pctl; +} =20 - llbitmap->pctl[i] =3D pctl; +static unsigned int llbitmap_reserved_pages(struct llbitmap *llbitmap) +{ + return DIV_ROUND_UP(llbitmap->mddev->bitmap_info.space << SECTOR_SHIFT, + PAGE_SIZE); +} =20 - if (IS_ERR(page)) { - llbitmap_free_pages(llbitmap); - return PTR_ERR(page); - } +static int llbitmap_alloc_pages(struct llbitmap *llbitmap) +{ + unsigned int used_pages =3D llbitmap_used_pages(llbitmap, llbitmap->chunk= s); + unsigned int nr_pages =3D max(used_pages, llbitmap_reserved_pages(llbitma= p)); + int i; + + llbitmap->pctl =3D kcalloc(nr_pages, sizeof(*llbitmap->pctl), GFP_NOIO); + if (!llbitmap->pctl) + return -ENOMEM; =20 - if (percpu_ref_init(&pctl->active, active_release, - PERCPU_REF_ALLOW_REINIT, GFP_NOIO)) { - __free_page(page); + llbitmap->nr_pages =3D nr_pages; + + for (i =3D 0; i < nr_pages; i++) { + llbitmap->pctl[i] =3D llbitmap_alloc_page_ctl(llbitmap, i); + if (IS_ERR(llbitmap->pctl[i])) { + int ret =3D PTR_ERR(llbitmap->pctl[i]); + + llbitmap->pctl[i] =3D NULL; llbitmap_free_pages(llbitmap); - return -ENOMEM; + return ret; } - - pctl->page =3D page; - pctl->state =3D page_address(page); - init_waitqueue_head(&pctl->wait); } =20 return 0; @@ -924,7 +949,7 @@ static int llbitmap_init(struct llbitmap *llbitmap) llbitmap->sync_size =3D blocks; mddev->bitmap_info.daemon_sleep =3D DEFAULT_DAEMON_SLEEP; =20 - ret =3D llbitmap_cache_pages(llbitmap); + ret =3D llbitmap_alloc_pages(llbitmap); if (ret) return ret; =20 @@ -1038,7 +1063,7 @@ static int llbitmap_read_sb(struct llbitmap *llbitmap) llbitmap->chunks =3D DIV_ROUND_UP_SECTOR_T(sync_size, chunksize); llbitmap->chunkshift =3D ffz(~chunksize); llbitmap->sync_size =3D sync_size; - ret =3D llbitmap_cache_pages(llbitmap); + ret =3D llbitmap_alloc_pages(llbitmap); =20 out_put_page: __free_page(sb_page); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A8296334C39; Sun, 2 Aug 2026 19:51:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700303; cv=none; b=r85ZZvTXWyizESaD9OJYiqdcaoaZ8eZuvRQ6ECvAR8ZO30y4CMYsoa7Up7LZ1OU66g/BznaS0ekTDcNqYDGAi0dZDVgbChA74/VC10YsPo34kRgM9jXigbnaCJhDQPhFKp/KFH1Grhhxz5jbUPXSCEwJTp7PsogbsJpZMpvAskg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700303; c=relaxed/simple; bh=9uFQ93/ZlvSilLTbzxrSFw0v98WHtEz0LAOVHAV0qy4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=V/7J/FqXzGx7HnowINuVtG3Tit15tns2RKjQrfCv+qCGTxnRbAs6vlAiXAGjGFTTQ9+TsdmtRiOdXULFVHvrpaX0qu4LvWMbxrjDhuou05OXYauIMZtrE57W9haGX+wBFRB/Ah8EPXFyQK0rIV1h6keaO8nnFdAEL6hqqM2mrJI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=R0Jmp48V; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="R0Jmp48V" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 35A6D1F00A3A; Sun, 2 Aug 2026 19:51:38 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700302; bh=7aaLxNVYe3z6sz70VG4QA59RJHliOPnir9WoCzVL+oE=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=R0Jmp48VxENNPALD3Ss6VyQANQb+qql0X+Z7V7+Wfo0a3r6mWEQBvgBbp9c8YvuJW kZn9F0WtPsvRJScMpEAh/q8pifdd9szHpR2aA/ZFHaApPmJWutx+Lf3A4M83IAZ+1Z jbzuwM/IgyDzt44zyh7RuXz7bwDxltysOej/+SZwFHZtB+vE6DtZjPM0dnB+mL1ZE1 60stK8ftvTUOe4cvcAuet3gBzZVYQQPFZ3O4w0KK97kkLYfA+eCP1EXPqhxkFRETGJ uWmHCIB6LfTxoBQa75ZIzMX4JRrcNeuE3VNS9wdEVIRMOJmpbOb0mE5mySTFCLH8m6 BwiGrpeT5LM3A== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 15/29] md/md-llbitmap: grow the page cache in place for reshape Date: Mon, 3 Aug 2026 03:50:24 +0800 Message-ID: <20260802195038.164272-16-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Use the page-control helpers to grow llbitmap's cached pages in place for resize and later reshape preparation, instead of rebuilding the whole cache. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 145 +++++++++++++++++++++++++++++++++++---- 1 file changed, 133 insertions(+), 12 deletions(-) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 300dd8b93b01..ddeea2098987 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -416,6 +416,19 @@ static char state_machine[BitStateCount][BitmapActionC= ount] =3D { }; =20 static void __llbitmap_flush(struct mddev *mddev); +static void llbitmap_flush(struct mddev *mddev); +static void llbitmap_update_sb(void *data); + +static void llbitmap_calculate_chunks(struct mddev *mddev, sector_t blocks, + unsigned long *chunksize, + unsigned long *chunks) +{ + *chunks =3D DIV_ROUND_UP_SECTOR_T(blocks, *chunksize); + while (*chunks > mddev->bitmap_info.space << SECTOR_SHIFT) { + *chunksize =3D *chunksize << 1; + *chunks =3D DIV_ROUND_UP_SECTOR_T(blocks, *chunksize); + } +} =20 static enum llbitmap_state llbitmap_read(struct llbitmap *llbitmap, loff_t= pos) { @@ -655,6 +668,48 @@ static unsigned int llbitmap_reserved_pages(struct llb= itmap *llbitmap) PAGE_SIZE); } =20 +static int llbitmap_expand_pages(struct llbitmap *llbitmap, + unsigned long chunks) +{ + struct llbitmap_page_ctl **pctl; + unsigned int old_nr_pages =3D llbitmap->nr_pages; + unsigned int nr_pages =3D llbitmap_used_pages(llbitmap, chunks); + unsigned int i; + int ret; + + if (nr_pages <=3D old_nr_pages) + return 0; + + pctl =3D kcalloc(nr_pages, sizeof(*pctl), GFP_NOIO); + if (!pctl) + return -ENOMEM; + + if (llbitmap->pctl) + memcpy(pctl, llbitmap->pctl, + array_size(old_nr_pages, sizeof(*pctl))); + + for (i =3D old_nr_pages; i < nr_pages; i++) { + pctl[i] =3D llbitmap_alloc_page_ctl(llbitmap, i); + if (IS_ERR(pctl[i])) + goto err_alloc_ptr; + } + + kfree(llbitmap->pctl); + llbitmap->pctl =3D pctl; + llbitmap->nr_pages =3D nr_pages; + return 0; + +err_alloc_ptr: + ret =3D PTR_ERR(pctl[i]); + while (i-- > old_nr_pages) { + __free_page(pctl[i]->page); + percpu_ref_exit(&pctl[i]->active); + kfree(pctl[i]); + } + kfree(pctl); + return ret; +} + static int llbitmap_alloc_pages(struct llbitmap *llbitmap) { unsigned int used_pages =3D llbitmap_used_pages(llbitmap, llbitmap->chunk= s); @@ -730,6 +785,34 @@ static bool llbitmap_zero_all_disks(struct llbitmap *l= lbitmap) return true; } =20 +static void llbitmap_mark_range(struct llbitmap *llbitmap, + unsigned long start, + unsigned long end, + enum llbitmap_state state) +{ + while (start <=3D end) { + llbitmap_write(llbitmap, state, start); + start++; + } +} + +static int llbitmap_prepare_resize(struct llbitmap *llbitmap, + unsigned long old_chunks, + unsigned long new_chunks, + unsigned long cache_chunks) +{ + int ret; + + llbitmap_flush(llbitmap->mddev); + ret =3D llbitmap_expand_pages(llbitmap, cache_chunks); + if (ret) + return ret; + if (new_chunks > old_chunks) + llbitmap_mark_range(llbitmap, old_chunks, new_chunks - 1, + BitUnwritten); + return 0; +} + static void llbitmap_init_state(struct llbitmap *llbitmap) { struct mddev *mddev =3D llbitmap->mddev; @@ -1032,10 +1115,10 @@ static int llbitmap_read_sb(struct llbitmap *llbitm= ap) goto out_put_page; } =20 - if (chunksize < DIV_ROUND_UP_SECTOR_T(mddev->resync_max_sectors, + if (chunksize < DIV_ROUND_UP_SECTOR_T(sync_size, mddev->bitmap_info.space << SECTOR_SHIFT)) { pr_err("md/llbitmap: %s: chunksize too small %lu < %llu / %lu", - mdname(mddev), chunksize, mddev->resync_max_sectors, + mdname(mddev), chunksize, sync_size, mddev->bitmap_info.space); goto out_put_page; } @@ -1184,24 +1267,62 @@ static int llbitmap_create(struct mddev *mddev) static int llbitmap_resize(struct mddev *mddev, sector_t blocks, int chunk= size) { struct llbitmap *llbitmap =3D mddev->bitmap; + sector_t old_blocks =3D llbitmap->sync_size; + unsigned long old_chunks =3D llbitmap->chunks; unsigned long chunks; + unsigned long cache_chunks; + int ret =3D 0; + unsigned long bitmap_chunksize; + bool reshape; + bool quiesced =3D false; =20 if (chunksize =3D=3D 0) chunksize =3D llbitmap->chunksize; =20 - /* If there is enough space, leave the chunksize unchanged. */ - chunks =3D DIV_ROUND_UP_SECTOR_T(blocks, chunksize); - while (chunks > mddev->bitmap_info.space << SECTOR_SHIFT) { - chunksize =3D chunksize << 1; - chunks =3D DIV_ROUND_UP_SECTOR_T(blocks, chunksize); - } + bitmap_chunksize =3D chunksize; + llbitmap_calculate_chunks(mddev, blocks, &bitmap_chunksize, &chunks); =20 - llbitmap->chunkshift =3D ffz(~chunksize); - llbitmap->chunksize =3D chunksize; - llbitmap->chunks =3D chunks; - llbitmap->sync_size =3D blocks; + reshape =3D mddev->delta_disks || mddev->new_level !=3D mddev->level || + mddev->new_layout !=3D mddev->layout || + mddev->new_chunk_sectors !=3D mddev->chunk_sectors; + if (!reshape && bitmap_chunksize !=3D llbitmap->chunksize) + return -EOPNOTSUPP; + if (blocks =3D=3D old_blocks && chunks =3D=3D llbitmap->chunks) + return 0; =20 + if (mddev->pers->quiesce) { + mddev->pers->quiesce(mddev, 1); + quiesced =3D true; + } + + mutex_lock(&mddev->bitmap_info.mutex); + cache_chunks =3D reshape ? max(old_chunks, chunks) : chunks; + ret =3D llbitmap_prepare_resize(llbitmap, old_chunks, chunks, cache_chunk= s); + if (ret) + goto out; + + if (reshape) { + llbitmap->chunks =3D max(old_chunks, chunks); + } else { + if (blocks < old_blocks && chunks < old_chunks) + llbitmap_mark_range(llbitmap, chunks, old_chunks - 1, + BitUnwritten); + mddev->bitmap_info.chunksize =3D bitmap_chunksize; + llbitmap->chunks =3D chunks; + llbitmap->sync_size =3D blocks; + llbitmap_update_sb(llbitmap); + } + __llbitmap_flush(mddev); + mutex_unlock(&mddev->bitmap_info.mutex); + if (quiesced) + mddev->pers->quiesce(mddev, 0); return 0; + +out: + mutex_unlock(&mddev->bitmap_info.mutex); + if (quiesced) + mddev->pers->quiesce(mddev, 0); + return ret; } =20 static int llbitmap_load(struct mddev *mddev) --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0AC40331EAB; Sun, 2 Aug 2026 19:51:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700307; cv=none; b=sg3edrTWSV39F1B/J69NzLxhwrl74WMZBTwU6kWZuT+HrYzlQEwbfYFf+/mSbsKKx6p99yu12Bv7oP83U1DSnoc4vbHKJoi1F/+e7hPghePOhH6rPz6jDQk7x4W76funzXOP50Qv6NTzZNwL4zzbzXLkTh+BjeIkAris1O4YDDA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700307; c=relaxed/simple; bh=lAlHtBTsEC8W2Q2572d5X6A3WNwSXAYJi0r/+zXgnDg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=efg6EWs582yhJQBnU9tWt+rNkv75da6L16J5pgWDOeutVgD/Du/hPF5k33xuKMRKuE6YcRIrEDknT7xgslwX9WCCsMPvhno5zEZcBcaZTr7hxurh8RsMb3vTUMCsCmaScNYLR8JhTH0fhbHKLzCYz0DxLIzV2H+KMg4cCuZsdcw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=niamU0uA; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="niamU0uA" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 21C901F000E9; Sun, 2 Aug 2026 19:51:42 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700305; bh=imcEUxtNLe0XJGPCb7oqWIt8GkXPc+zCyHeqJYskS/c=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=niamU0uAwt0eH2npySELIucX1ueU1mQx7R/yxSSotgIZYWvbIO408VWaddOUXesBc WQqdv212zhl+sU+HCSezt7OXzsJmqf4m3MziYDmWLuZ6ipqngOmoIiWqfKvSA2I6gO gFlFoi3E0qavcPIAZ7alrje4Yx+3ozzWAAhGexW6L9FwGeefQvkKW1d23NJJVJjh2L mNINyaoesJHJ/EI697Bqmk0bv9raPGpNGNySjBoZ1lSFELuqDEGYKAKDZVJgHuRf8Z SOdTIBgZKIX4/6ayna/s4bf9E5TOgowpU4eaWIP+bGFHhMU5N6caRu2fTmREgHslnK xOODsjrZJauGg== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 16/29] md/md-llbitmap: track target reshape geometry fields Date: Mon, 3 Aug 2026 03:50:25 +0800 Message-ID: <20260802195038.164272-17-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Track llbitmap bookkeeping for the target reshape geometry while keeping a single live bitmap instance. Add the reshape geometry fields, refresh helper, and update the load and resize paths to keep the target geometry in sync. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 42 ++++++++++++++++++++++++++++++++++++++++ 1 file changed, 42 insertions(+) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index ddeea2098987..71a490390142 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -289,6 +289,9 @@ struct llbitmap { unsigned long chunks; /* total number of sectors tracked by current bitmap geometry */ sector_t sync_size; + unsigned long reshape_chunksize; + unsigned long reshape_chunks; + sector_t reshape_sync_size; unsigned long last_end_sync; /* * time in seconds that dirty bits will be cleared if the page is not @@ -430,6 +433,39 @@ static void llbitmap_calculate_chunks(struct mddev *md= dev, sector_t blocks, } } =20 +static bool llbitmap_reshaping(struct llbitmap *llbitmap) +{ + return llbitmap->mddev->reshape_position !=3D MaxSector; +} + +static sector_t llbitmap_personality_sync_size(struct llbitmap *llbitmap, + bool previous) +{ + struct mddev *mddev =3D llbitmap->mddev; + + if (!llbitmap_reshaping(llbitmap) || !mddev->private || !mddev->pers || + !mddev->pers->bitmap_sync_size) + return llbitmap->sync_size; + return mddev->pers->bitmap_sync_size(mddev, previous); +} + +static void llbitmap_refresh_reshape(struct llbitmap *llbitmap) +{ + unsigned long old_chunks =3D DIV_ROUND_UP_SECTOR_T(llbitmap->sync_size, + llbitmap->chunksize); + sector_t blocks =3D llbitmap_personality_sync_size(llbitmap, false); + unsigned long chunksize =3D llbitmap->chunksize; + unsigned long chunks =3D DIV_ROUND_UP_SECTOR_T(blocks, chunksize); + + llbitmap->reshape_sync_size =3D blocks; + llbitmap->reshape_chunksize =3D chunksize; + llbitmap->reshape_chunks =3D chunks; + llbitmap_calculate_chunks(llbitmap->mddev, blocks, + &llbitmap->reshape_chunksize, + &llbitmap->reshape_chunks); + llbitmap->chunks =3D max(old_chunks, llbitmap->reshape_chunks); +} + static enum llbitmap_state llbitmap_read(struct llbitmap *llbitmap, loff_t= pos) { unsigned int idx; @@ -1030,6 +1066,7 @@ static int llbitmap_init(struct llbitmap *llbitmap) llbitmap->chunksize =3D chunksize; llbitmap->chunks =3D chunks; llbitmap->sync_size =3D blocks; + llbitmap_refresh_reshape(llbitmap); mddev->bitmap_info.daemon_sleep =3D DEFAULT_DAEMON_SLEEP; =20 ret =3D llbitmap_alloc_pages(llbitmap); @@ -1146,6 +1183,7 @@ static int llbitmap_read_sb(struct llbitmap *llbitmap) llbitmap->chunks =3D DIV_ROUND_UP_SECTOR_T(sync_size, chunksize); llbitmap->chunkshift =3D ffz(~chunksize); llbitmap->sync_size =3D sync_size; + llbitmap_refresh_reshape(llbitmap); ret =3D llbitmap_alloc_pages(llbitmap); =20 out_put_page: @@ -1302,6 +1340,9 @@ static int llbitmap_resize(struct mddev *mddev, secto= r_t blocks, int chunksize) goto out; =20 if (reshape) { + llbitmap->reshape_sync_size =3D blocks; + llbitmap->reshape_chunksize =3D bitmap_chunksize; + llbitmap->reshape_chunks =3D chunks; llbitmap->chunks =3D max(old_chunks, chunks); } else { if (blocks < old_blocks && chunks < old_chunks) @@ -1310,6 +1351,7 @@ static int llbitmap_resize(struct mddev *mddev, secto= r_t blocks, int chunksize) mddev->bitmap_info.chunksize =3D bitmap_chunksize; llbitmap->chunks =3D chunks; llbitmap->sync_size =3D blocks; + llbitmap_refresh_reshape(llbitmap); llbitmap_update_sb(llbitmap); } __llbitmap_flush(mddev); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DD1683368A2; Sun, 2 Aug 2026 19:51:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700310; cv=none; b=Xge+/NJu8OdXT4R8D8X9Jt5W8qcDG2ni/zVZDQNjpMvU9pxl2v9hMiJlkw9uRt7vafEu1DUuLQ11oCtEK4oYsKCU6tN/9/cMzO2iJj+KUP+Ut53Ze10gC1ACWjt1QKlBSRX5OQR8pjCnSzEi50i9vgH1l243jLvhMVTqJa+1xO8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700310; c=relaxed/simple; bh=DJ6EU+fzv8iOEUsBftq9+B6ZNyeKCzhFbRnYctpHLt8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gnKAGMN8vVcYvTgPGH/M5FxXDsFXNogiN3UUKb0i9kYLpSiobTgaOYJAsFXnSGWtB3K4oE50roxDlQSoDc47DBsl6fRQheiY5j4vdpNxq5yxJk2TikF8HcViQO1ruHse2/0+zNKinJy1ss7DiYXwc6XKW7VYzy3l6r1/I4gby2Q= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=BoD+aKrt; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="BoD+aKrt" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3B4A51F00A3A; Sun, 2 Aug 2026 19:51:46 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700308; bh=kbUdOuogwlmBTpohYQNpHXGAq9yvsml8HoNAsDi+sfU=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=BoD+aKrtv5sfcEF7Iy/ZKZr/BLsv+T7G/ocgVBqTsBi/MKTScQr4m4qGiovpF28u9 gOSaD+EPtZRB4kAAcK0mXI47FLpcAgZ+CwC5VwiE+at1YRBYHlc5F89NGq2Gzsc7Xd BSrrt4c5Y3Vyo6z0qVJ1XcrD9gv3V94OZ7Q/2dsgo+RTFnteSQ7C/ne39wnDq9OCAj QjLUzXGFIIDpF64H7LDNRcxnztfn+ji/WCHXZvsb96IyWh+K+Z/Ud8skh0bVNOyc1t BPf4Nu9l1rjj3Ttmukw6UZC3XdOGYH4ufbqaBSHmb55bx1/XITj89rDL45eAUCm9gN 1BhkBbT+RIwvg== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 17/29] md/md-llbitmap: finish reshape geometry Date: Mon, 3 Aug 2026 03:50:26 +0800 Message-ID: <20260802195038.164272-18-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Commit the staged llbitmap geometry when reshape finishes. When assembling a stopped reshape, md_run() creates the bitmap before publishing mddev->pers. llbitmap_read_sb() can therefore only initialize the reshape fields from the old on-disk sync size. Refresh the staged reshape geometry again from llbitmap_load(), after mddev->pers is available, and expand the in-memory page controls before replaying bitmap state. Reproduce on the old kernel by creating a RAID10 llbitmap with four active disks and two spares, growing it to six disks, then stopping and assembling while reshape is still running. The llbitmap chunk count was 32704 before grow, 49056 during reshape, then rolled back to 32704 after reassemble. The fixed kernel kept the target geometry across the same stop/reassemble flow: 65440 chunks before grow, 98160 during reshape, and 98160 after reassemble. Reported-by: Mykola Marzhan Link: https://lore.kernel.org/all/20260726185916.2223460-1-mykola@meshstor.= io/ Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 34 ++++++++++++++++++++++++++++++++++ 1 file changed, 34 insertions(+) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 71a490390142..a7c229db3058 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -1371,11 +1371,20 @@ static int llbitmap_load(struct mddev *mddev) { enum llbitmap_action action =3D BitmapActionReload; struct llbitmap *llbitmap =3D mddev->bitmap; + int ret; =20 if (test_and_clear_bit(BITMAP_STALE, &llbitmap->flags)) action =3D BitmapActionStale; =20 + mutex_lock(&mddev->bitmap_info.mutex); + llbitmap_refresh_reshape(llbitmap); + ret =3D llbitmap_expand_pages(llbitmap, llbitmap->chunks); + if (ret) { + mutex_unlock(&mddev->bitmap_info.mutex); + return ret; + } llbitmap_state_machine(llbitmap, 0, llbitmap->chunks - 1, action); + mutex_unlock(&mddev->bitmap_info.mutex); return 0; } =20 @@ -1709,6 +1718,30 @@ static void llbitmap_dirty_bits(struct mddev *mddev,= unsigned long s, llbitmap_state_machine(mddev->bitmap, s, e, BitmapActionStartwrite); } =20 +static void llbitmap_reshape_finish(struct mddev *mddev) +{ + struct llbitmap *llbitmap =3D mddev->bitmap; + + if (mddev->pers->quiesce) + mddev->pers->quiesce(mddev, 1); + + mutex_lock(&mddev->bitmap_info.mutex); + llbitmap_flush(mddev); + + llbitmap->chunksize =3D llbitmap->reshape_chunksize; + llbitmap->chunkshift =3D ffz(~llbitmap->chunksize); + llbitmap->chunks =3D llbitmap->reshape_chunks; + llbitmap->sync_size =3D llbitmap->reshape_sync_size; + llbitmap_refresh_reshape(llbitmap); + mddev->bitmap_info.chunksize =3D llbitmap->chunksize; + llbitmap_update_sb(llbitmap); + __llbitmap_flush(mddev); + mutex_unlock(&mddev->bitmap_info.mutex); + + if (mddev->pers->quiesce) + mddev->pers->quiesce(mddev, 0); +} + static void llbitmap_write_sb(struct llbitmap *llbitmap) { int nr_blocks =3D DIV_ROUND_UP(BITMAP_DATA_OFFSET, llbitmap->io_size); @@ -2000,6 +2033,7 @@ static struct bitmap_operations llbitmap_ops =3D { .get_stats =3D llbitmap_get_stats, .dirty_bits =3D llbitmap_dirty_bits, .prepare_range =3D llbitmap_prepare_range, + .reshape_finish =3D llbitmap_reshape_finish, .write_all =3D llbitmap_write_all, =20 .groups =3D md_llbitmap_groups, --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D741330FC1D; Sun, 2 Aug 2026 19:51:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700314; cv=none; b=I9qSgIwozwxms/6SzhigrNPDToQTv6nbv2hZekhFsiTabbJ4IVxOPqP0coan4IbkyUkE941merP2pWdKFZKI+3ceJLOmbpcaFNRWU36rpZZ7ED0pHePv26iuJCTcLca0vbH5J/ague4d2dVNv5wq7hF7/5/1hQPoBDrcq9O4/KU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700314; c=relaxed/simple; bh=b2d86Chp3gAVBmArBUqqdOV49pAbqhBxNkr3qbqdk6U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=CwaChmsmLRVTX2yFMDLsU6E2XoBjuXqRYQrg2nivPLrVPeRM6mHGOwjvbQDuVu+sYuDZ6dwHqmJE1CSgHmV3M2YC84GElM1DgIsch/D0xcRPld6m5EhVyrdahXBpXnDolPepKc8mRzou7cy81Qa5E5D8U2z6bjDhNAI/jLFdk0w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=aWr3PSYd; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="aWr3PSYd" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 5342F1F000E9; Sun, 2 Aug 2026 19:51:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700312; bh=eUGBX7KepT7gu4vu5Z67JYUXykTW239FwnZo+emIELs=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=aWr3PSYddZ1Td5JXfPBay4cvEHmtYDLBWz1J+RF2aryhsdYfwE1/qSs2dqzLaU+pB JeUzTzGc0hDMx7F3wsxerUBS2sGsshaBzUvGzK7CqMUbkEu7jFff1GiZ1rIPnbp5XT twwsPc8ZgkqPPReHiQHOmeF63LNxjLcGuqE0MPuZNjlTL+iYEHuw5oBzGfuAAl5+fZ i+hH3HnP/e33gdtjHJaDqb+hayccDfQW+lEbZDTFLwo/ysPL9LAwAAztcyp2uFgE8n L148ivK3MTBx2wympI6GoLGXse9iTOLu0VeFf0ZnpjN4DdjYOA4VIgmJ8crwaaIOW0 NYxAmMoRjZkUg== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 18/29] md/md-llbitmap: refuse reshape while llbitmap still needs sync Date: Mon, 3 Aug 2026 03:50:27 +0800 Message-ID: <20260802195038.164272-19-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Reject reshape when llbitmap still contains NeedSync or Syncing bits. This keeps reshape from starting until the current llbitmap state has been reconciled. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 24 ++++++++++++++++++++++++ 1 file changed, 24 insertions(+) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index a7c229db3058..f8a1b0f79be6 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -1718,6 +1718,29 @@ static void llbitmap_dirty_bits(struct mddev *mddev,= unsigned long s, llbitmap_state_machine(mddev->bitmap, s, e, BitmapActionStartwrite); } =20 +static int llbitmap_reshape_can_start(struct mddev *mddev) +{ + struct llbitmap *llbitmap =3D mddev->bitmap; + unsigned long chunk; + int ret =3D 0; + + if (!llbitmap) + return 0; + + mutex_lock(&mddev->bitmap_info.mutex); + for (chunk =3D 0; chunk < llbitmap->chunks; chunk++) { + enum llbitmap_state state =3D llbitmap_read(llbitmap, chunk); + + if (state =3D=3D BitNeedSync || state =3D=3D BitSyncing) { + ret =3D -EBUSY; + break; + } + } + mutex_unlock(&mddev->bitmap_info.mutex); + + return ret; +} + static void llbitmap_reshape_finish(struct mddev *mddev) { struct llbitmap *llbitmap =3D mddev->bitmap; @@ -2034,6 +2057,7 @@ static struct bitmap_operations llbitmap_ops =3D { .dirty_bits =3D llbitmap_dirty_bits, .prepare_range =3D llbitmap_prepare_range, .reshape_finish =3D llbitmap_reshape_finish, + .reshape_can_start =3D llbitmap_reshape_can_start, .write_all =3D llbitmap_write_all, =20 .groups =3D md_llbitmap_groups, --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 21C6D3290AF; Sun, 2 Aug 2026 19:51:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700317; cv=none; b=kZUqftCk6y4UGWKhgHb0Wra25C6WDz+6XXvQyYZj44iP3/ZsZB7VgOmU17f9o0glq8CAZ5jje46PHVEcqbB6W56B8pLCXbJ0dvaWDIzWtCGeXwVS/zCMafULeC5zAU9puUyZa3C3nhz01PBCP7tkCIGuorBVFCfIZha5D+STIXk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700317; c=relaxed/simple; bh=Tf9LxfogAbPa80Jyfg8Xmdxqe+DN3MAF86OrVEFa9ug=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Xt9kwDVcovA2O5UtSMUTXfRRvkLC9RTB8Ka08DB7FWnSVGs+W7D6CfQ6lfRoDSrNnJ9xYafl1Wz+OlhQNxR1oakdE6j3m6umhKtVHMWEkAPPMp7DTycxVEKqQyJWIUVS6QgUG6pmxqSUazWA3rzKEsQ1frzXDqvlIWb25A9nEDw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=CXJoLruQ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="CXJoLruQ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3E2611F00A3A; Sun, 2 Aug 2026 19:51:53 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700316; bh=mlZf7iFI7gH/bQtLhl4ns59GIR8oG8NSlg7Wfv+hq6I=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=CXJoLruQz1Q64ytaEKypnvsX4rDtR3pN//taDTt8FC4GqnLZuUrs3SzAEL0HL0dKv +HpaS99zBdp11KeKKUTd7A9csruiys84ccw2iPQiZbuW8cSWfmw+JtMNMBKMxRTnqC G0eV1ugqMMMTuV0HiOEZ2aGixeAzxwHai6IoTIgF4dBlmAC2vex+mHwaD7aaUJrguc qZUePm/3YcTjVRlO0wvofZQN1wlGXry1WG4h75KOsDE5LNWvqgprskNoMBIJYAkaIP U8Nsg18rqtV1J2ObR10UzWf7lwwIETO8CqHNgfhNYu5AqDLPQTtS0t6NuSLbe75Z2h ySERG9TaaDSpA== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 19/29] md/md-llbitmap: add reshape range mapping helpers Date: Mon, 3 Aug 2026 03:50:28 +0800 Message-ID: <20260802195038.164272-20-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Teach llbitmap to choose old versus new geometry during reshape and to encode exact bitmap ranges for the active geometry. This is the mapping groundwork for checkpoint remapping. Range preparation now distinguishes writes from discards. Normal writes must cover every touched bitmap chunk, while discards may only mark fully covered chunks unwritten. Without this distinction, a discard that starts or ends inside a chunk can make live data look unwritten after the range has been mapped and floored. Reproduce that with a RAID1 llbitmap using 128-sector chunks. A discard starting halfway into chunk 8 with a 128-sector length changed clean bits from 16352 to 16350 and unwritten bits from 0 to 2, even though no chunk was fully discarded. With discard-specific range encoding, both counts stay unchanged for the same test. Range preparation also clamps the pre-map range in the same coordinate space as the incoming IO. RAID5 receives array-sector offsets but tracks llbitmap sync size in component sectors, so steady-state RAID5 must use bitmap_array_sectors() before mapping and keep the existing sync-size clamp after mapping. Reproduce that with a 4-disk RAID5 llbitmap created --assume-clean. A write below dev_sectors changed dirty bits from 0 to 512, but a write at seek=3D2094080 left the count at 512. With the array-sector pre-map limit, writing at seek=3Dcomponent_size + 65536 increased dirty bits from 512 to 1024. Reported-by: Mykola Marzhan Link: https://lore.kernel.org/all/20260726185916.2223460-1-mykola@meshstor.= io/ Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-bitmap.c | 2 +- drivers/md/md-bitmap.h | 3 +- drivers/md/md-llbitmap.c | 137 +++++++++++++++++++++++++++++++++++---- drivers/md/md.c | 11 ++-- 4 files changed, 134 insertions(+), 19 deletions(-) diff --git a/drivers/md/md-bitmap.c b/drivers/md/md-bitmap.c index b7f0d4acce04..b8325cb09a37 100644 --- a/drivers/md/md-bitmap.c +++ b/drivers/md/md-bitmap.c @@ -1731,7 +1731,7 @@ static void bitmap_start_write(struct mddev *mddev, s= ector_t offset, } =20 static void bitmap_prepare_range(struct mddev *mddev, sector_t *offset, - unsigned long *sectors) + unsigned long *sectors, bool discard) { if (mddev->pers->bitmap_sector) mddev->pers->bitmap_sector(mddev, offset, sectors); diff --git a/drivers/md/md-bitmap.h b/drivers/md/md-bitmap.h index 6478cf9d8816..b69c78174f02 100644 --- a/drivers/md/md-bitmap.h +++ b/drivers/md/md-bitmap.h @@ -97,7 +97,8 @@ struct bitmap_operations { /* Prepare a range for this bitmap implementation. */ void (*prepare_range)(struct mddev *mddev, sector_t *offset, - unsigned long *sectors); + unsigned long *sectors, + bool discard); void (*reshape_finish)(struct mddev *mddev); int (*reshape_can_start)(struct mddev *mddev); void (*reshape_mark)(struct mddev *mddev, sector_t old_pos, diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index f8a1b0f79be6..fa16a4224c45 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -9,6 +9,7 @@ #include #include #include +#include #include #include =20 @@ -433,22 +434,28 @@ static void llbitmap_calculate_chunks(struct mddev *m= ddev, sector_t blocks, } } =20 -static bool llbitmap_reshaping(struct llbitmap *llbitmap) -{ - return llbitmap->mddev->reshape_position !=3D MaxSector; -} - static sector_t llbitmap_personality_sync_size(struct llbitmap *llbitmap, bool previous) { struct mddev *mddev =3D llbitmap->mddev; =20 - if (!llbitmap_reshaping(llbitmap) || !mddev->private || !mddev->pers || + if (READ_ONCE(mddev->reshape_position) =3D=3D MaxSector || + !mddev->private || !mddev->pers || !mddev->pers->bitmap_sync_size) return llbitmap->sync_size; return mddev->pers->bitmap_sync_size(mddev, previous); } =20 +static sector_t llbitmap_logical_size(struct llbitmap *llbitmap, bool prev= ious) +{ + struct mddev *mddev =3D llbitmap->mddev; + + if (!mddev->private || !mddev->pers || + !mddev->pers->bitmap_array_sectors) + return llbitmap_personality_sync_size(llbitmap, previous); + return mddev->pers->bitmap_array_sectors(mddev, previous); +} + static void llbitmap_refresh_reshape(struct llbitmap *llbitmap) { unsigned long old_chunks =3D DIV_ROUND_UP_SECTOR_T(llbitmap->sync_size, @@ -466,6 +473,80 @@ static void llbitmap_refresh_reshape(struct llbitmap *= llbitmap) llbitmap->chunks =3D max(old_chunks, llbitmap->reshape_chunks); } =20 +static void llbitmap_map_layout(struct llbitmap *llbitmap, sector_t *offse= t, + unsigned long *sectors, bool previous) +{ + sector_t limit =3D llbitmap_logical_size(llbitmap, previous); + sector_t start =3D *offset; + sector_t end =3D start + *sectors; + + if (start >=3D limit) { + *sectors =3D 0; + return; + } + if (end > limit) + end =3D limit; + + *offset =3D start; + *sectors =3D end - start; + if (!*sectors) + return; + + if (llbitmap->mddev->pers->bitmap_sector_map) + llbitmap->mddev->pers->bitmap_sector_map(llbitmap->mddev, offset, + sectors, previous); + else if (!previous && llbitmap->mddev->pers->bitmap_sector) + llbitmap->mddev->pers->bitmap_sector(llbitmap->mddev, offset, + sectors); +} + +static void llbitmap_encode_range(struct llbitmap *llbitmap, sector_t *off= set, + unsigned long *sectors, bool previous) +{ + unsigned long chunksize =3D previous ? llbitmap->chunksize : + llbitmap->reshape_chunksize; + u64 start; + u64 end; + + if (!*sectors) { + *offset =3D 0; + return; + } + + start =3D div64_u64(*offset, chunksize); + end =3D div64_u64(*offset + *sectors - 1, chunksize); + *offset =3D (sector_t)start << llbitmap->chunkshift; + *sectors =3D (end - start + 1) << llbitmap->chunkshift; +} + +static void llbitmap_encode_discard_range(struct llbitmap *llbitmap, + sector_t *offset, + unsigned long *sectors, + bool previous) +{ + unsigned long chunksize =3D previous ? llbitmap->chunksize : + llbitmap->reshape_chunksize; + sector_t end =3D *offset + *sectors; + u64 start; + u64 last; + + if (!*sectors) { + *offset =3D 0; + return; + } + + start =3D DIV_ROUND_UP_SECTOR_T(*offset, chunksize); + last =3D div64_u64(end, chunksize); + if (start >=3D last) { + *offset =3D 0; + *sectors =3D 0; + return; + } + + *offset =3D (sector_t)start << llbitmap->chunkshift; + *sectors =3D (last - start) << llbitmap->chunkshift; +} + static enum llbitmap_state llbitmap_read(struct llbitmap *llbitmap, loff_t= pos) { unsigned int idx; @@ -1409,11 +1490,35 @@ static void llbitmap_destroy(struct mddev *mddev) mutex_unlock(&mddev->bitmap_info.mutex); } =20 +static bool llbitmap_map_previous(struct llbitmap *llbitmap, sector_t offs= et, + unsigned long sectors) +{ + struct mddev *mddev =3D llbitmap->mddev; + sector_t boundary =3D READ_ONCE(mddev->reshape_position); + + if (boundary =3D=3D MaxSector) + return false; + + WARN_ON_ONCE(sectors && offset < boundary && offset + sectors > boundary); + + return mddev->reshape_backwards ? offset < boundary : offset >=3D boundar= y; +} + static void llbitmap_prepare_range(struct mddev *mddev, sector_t *offset, - unsigned long *sectors) + unsigned long *sectors, bool discard) { - if (mddev->pers->bitmap_sector) - mddev->pers->bitmap_sector(mddev, offset, sectors); + struct llbitmap *llbitmap =3D mddev->bitmap; + bool previous; + + if (!llbitmap) + return; + + previous =3D llbitmap_map_previous(llbitmap, *offset, *sectors); + llbitmap_map_layout(llbitmap, offset, sectors, previous); + if (discard) + llbitmap_encode_discard_range(llbitmap, offset, sectors, previous); + else + llbitmap_encode_range(llbitmap, offset, sectors, previous); } =20 static void llbitmap_start_write(struct mddev *mddev, sector_t offset, @@ -1582,7 +1687,11 @@ static bool llbitmap_blocks_synced(struct mddev *mdd= ev, sector_t offset) { struct llbitmap *llbitmap =3D mddev->bitmap; unsigned long p =3D offset >> llbitmap->chunkshift; - enum llbitmap_state c =3D llbitmap_read(llbitmap, p); + enum llbitmap_state c; + + if (p >=3D llbitmap->chunks) + return false; + c =3D llbitmap_read(llbitmap, p); =20 return c =3D=3D BitClean || c =3D=3D BitDirty || c =3D=3D BitCleanUnwritt= en; } @@ -1592,7 +1701,11 @@ static sector_t llbitmap_skip_sync_blocks(struct mdd= ev *mddev, sector_t offset) struct llbitmap *llbitmap =3D mddev->bitmap; unsigned long p =3D offset >> llbitmap->chunkshift; int blocks =3D llbitmap->chunksize - (offset & (llbitmap->chunksize - 1)); - enum llbitmap_state c =3D llbitmap_read(llbitmap, p); + enum llbitmap_state c; + + if (p >=3D llbitmap->chunks) + return 0; + c =3D llbitmap_read(llbitmap, p); =20 /* always skip unwritten blocks */ if (c =3D=3D BitUnwritten) @@ -1637,6 +1750,8 @@ static bool llbitmap_start_sync(struct mddev *mddev, = sector_t offset, * if md_do_sync() loop more times. */ *blocks =3D llbitmap->chunksize - (offset & (llbitmap->chunksize - 1)); + if (p >=3D llbitmap->chunks) + return false; state =3D llbitmap_state_machine(llbitmap, p, p, BitmapActionStartsync); return state =3D=3D BitSyncing || state =3D=3D BitSyncingUnwritten; } diff --git a/drivers/md/md.c b/drivers/md/md.c index 538ba7bab060..e88381beb209 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -9428,21 +9428,20 @@ struct bio *mddev_bio_split_at_reshape_offset(struc= t mddev *mddev, EXPORT_SYMBOL_GPL(mddev_bio_split_at_reshape_offset); =20 static void md_bitmap_prepare_range(struct mddev *mddev, sector_t *offset, - unsigned long *sectors) + unsigned long *sectors, bool discard) { - mddev->bitmap_ops->prepare_range(mddev, offset, sectors); + mddev->bitmap_ops->prepare_range(mddev, offset, sectors, discard); } =20 static void md_bitmap_start(struct mddev *mddev, struct md_io_clone *md_io_clone) { - md_bitmap_fn *fn =3D unlikely(md_io_clone->rw =3D=3D STAT_DISCARD) ? - mddev->bitmap_ops->start_discard : + bool discard =3D md_io_clone->rw =3D=3D STAT_DISCARD; + md_bitmap_fn *fn =3D discard ? mddev->bitmap_ops->start_discard : mddev->bitmap_ops->start_write; =20 md_bitmap_prepare_range(mddev, &md_io_clone->offset, - &md_io_clone->sectors); - + &md_io_clone->sectors, discard); if (!md_io_clone->sectors) return; fn(mddev, md_io_clone->offset, md_io_clone->sectors); --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0429F334692; Sun, 2 Aug 2026 19:51:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700321; cv=none; b=PfYL5GhHqTz6WJ9wM7F2b5QAisg6814jrqzj1TQtLDUFl/NtaD8IrtaPO2AFrSJsiZq1cvhQSZgvKFdvB/2mpfZFl4ccXy34Mu8wm0rZrBNfi/Pbs2nGbg1DOHS7qsD3YyUCzFzDE4vu06GWx+84aFvLquVdfVv/4kNtoIf5tHk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700321; c=relaxed/simple; bh=kt9sa2AjJSIaZe0my3Fo6pjgkgfQlLdOSQMLLiEnGOs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Gw72a3UA4pFe4DhIlMKaAbeOiWOop0Xi0qvJRMfObCsziwiudnEVINak3OEhhsrKd0sKy65PvzTukJST90fKQaKhupSU7QQms1j4iJ7Pe1bys4mmlzvhOXbVBAEGNL/3kWlBHFAkZs4ZpspBzou9YnLmm7WhIqvXEXffI79kpb8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=X/bQkBsl; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="X/bQkBsl" Received: by smtp.kernel.org (Postfix) with ESMTPSA id F3BBF1F000E9; Sun, 2 Aug 2026 19:51:56 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700319; bh=5MBbMM4FsgaPEuiAk8SX43polaPfRJ6SfrrnLrmeXzE=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=X/bQkBsl2JWRj6NMP5/THMwSUwElNap3fHizsWGyMZcB+punZjBgrcA/2cUFj+H55 CA7OZ1UL7BwulcmtMoytHzKGU12dELCcrltPdM0IZZnFM1GO02gwj0icFmmA3x74yT CQCSInkZzyqymRQ/0m9cjxviDa9NiEU48OWwojSTrGvHSV9zI/xKIKI66dFFfe8uAv p1KnPfwbikX/iEqxUuger5Zh79SNIHcZvqE6O+qjdlUYYRGfWzu+rJAq/ETcwU3P02 u6+dniyynNt6OijfbkxgmpVkRZB8bySf8t+VbFoAALi8EfJ2HerxStrVrnriHVuwWY Z8u0QvMtfIsPg== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 20/29] md/md-llbitmap: don't skip reshape ranges from bitmap state Date: Mon, 3 Aug 2026 03:50:29 +0800 Message-ID: <20260802195038.164272-21-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Reshape progress is tracked by array metadata rather than llbitmap. Do not let llbitmap skip_sync_blocks() suppress reshape ranges based on stale bitmap state before the corresponding checkpoint is persisted. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index fa16a4224c45..a20e55fdf82b 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -1707,6 +1707,14 @@ static sector_t llbitmap_skip_sync_blocks(struct mdd= ev *mddev, sector_t offset) return 0; c =3D llbitmap_read(llbitmap, p); =20 + /* + * Reshape progress is tracked by array metadata rather than llbitmap. + * Skipping reshape ranges from stale bitmap state can lose data after a + * restart before the corresponding bits are checkpointed to disk. + */ + if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery)) + return 0; + /* always skip unwritten blocks */ if (c =3D=3D BitUnwritten) return blocks; --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 22844301493; Sun, 2 Aug 2026 19:52:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700325; cv=none; b=rqaZEDhe+u4Y3DKnwMrm6EMWUSJMfWUiIynmBw9kDy2k0wccTOctSTgnciwOZpdsCLesDqYQ+C+dBD7GhgzLQ3jFnGZKXFk57mNc74PZ9yS/sNnEqZSh0XowhNsziDpg5PEteX4zEjVIOnajFlh836BGFnqwANT1tcKXZc4xtQg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700325; c=relaxed/simple; bh=2r3nJFQqp61u7xrm4tG9zf7L0VoSb1TmaUnmejqzt5g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=RLfakgDpMeYKm0hygXWCxeubaOb/vkr02PcJIRtvU+dMSYm7/zp69qSAYRpMpmCj222xI8HXL3T097szUyJF4MT4jQuqh4x7PBap+zrqBBq9GWuYtM6iblPkjSKU0ZGpIm0mnp/e+Kv58+r1g85ATQO2bumEu8RT18pFht16ZCw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=SLkeUvzr; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="SLkeUvzr" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 42DBB1F00A3A; Sun, 2 Aug 2026 19:51:59 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700323; bh=MVKxX6QpcQmdufbJ1KcquOMcnU3H13Kg9e7iQIi+aMw=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=SLkeUvzr776Qdv/+HoJS+jC5MN1NmMB1Uf6SrYMC1Ei9BeekZpGT8mvNkwVliFpGy oN3lB2DhiC83qSsKTrjzCwbzSK3p9Zll1+0k4P9YiJPf9DpBCtXaN3Ie2Rtuq6CXdp 9N3vqkptpe0jWHi2uQzqgwnseQXkn/ZHwFLotbzvuSeH2ItzJwlcPgiX65eXnSqW4y 9kRZ95ABqpvyR6Buux/yzgpfczygbwFuCf6TVO9a4UmwmyhR5EIGjUceoZ43KHisrM hL3Akd7T4HA+orlAHv56ZhADm/kNoWQzYsurSeyXlsBdfQUvjqPkYFql4ZAN58Ru8m 337MhgUIPeKaA== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 21/29] md/md-llbitmap: remap checkpointed bits as reshape progresses Date: Mon, 3 Aug 2026 03:50:30 +0800 Message-ID: <20260802195038.164272-22-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Merge checkpointed old llbitmap state forward as reshape_position advances and record the checkpoint remap through reshape_mark(). Normal write accounting can run while the reshape thread checkpoints a new reshape position. llbitmap_reshape_mark() reads old state bytes, merges them into destination bits, and writes the result back. If llbitmap_start_write() or llbitmap_start_discard() updates the same state bytes at the same time, the two read/modify/write paths can overwrite each other and lose the state from one side. Serialize only this state-byte race with a rwlock. Normal I/O takes the read side around llbitmap_state_machine(), after page active references are rais= ed, so concurrent normal I/O updates still run in parallel. Reshape checkpointi= ng takes the write side only while merging the checkpointed range, avoiding pa= ge suspension and avoiding a sleeping mutex in the I/O accounting path. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 204 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 204 insertions(+) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index a20e55fdf82b..5d95627ff983 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -302,6 +302,11 @@ struct llbitmap { /* fires on first BitDirty state */ struct timer_list pending_timer; struct work_struct daemon_work; + /* + * Serialize reshape checkpoint remapping against normal I/O bitmap + * updates without blocking concurrent I/O updates on each other. + */ + rwlock_t reshape_lock; =20 unsigned long flags; __u64 events_cleared; @@ -498,6 +503,14 @@ static void llbitmap_map_layout(struct llbitmap *llbit= map, sector_t *offset, else if (!previous && llbitmap->mddev->pers->bitmap_sector) llbitmap->mddev->pers->bitmap_sector(llbitmap->mddev, offset, sectors); + + limit =3D llbitmap_personality_sync_size(llbitmap, previous); + start =3D *offset; + end =3D start + *sectors; + if (start >=3D limit) + *sectors =3D 0; + else if (end > limit) + *sectors =3D limit - start; } =20 static void llbitmap_encode_range(struct llbitmap *llbitmap, sector_t *off= set, @@ -930,6 +943,33 @@ static int llbitmap_prepare_resize(struct llbitmap *ll= bitmap, return 0; } =20 +static enum llbitmap_state +llbitmap_rmerge_state(struct llbitmap *llbitmap, + enum llbitmap_state dst, + enum llbitmap_state src) +{ + bool level_456 =3D raid_is_456(llbitmap->mddev); + + if (dst =3D=3D BitNeedSync || dst =3D=3D BitSyncing || + src =3D=3D BitNeedSync || src =3D=3D BitSyncing) + return BitNeedSync; + + if (dst =3D=3D BitDirty || src =3D=3D BitDirty) + return BitDirty; + + /* + * Reshape generates valid target parity/data for both already-written + * and not-yet-written regions in the checkpointed range, so a mix of + * clean and unwritten still results in a clean destination bit. + */ + if (level_456 && ((dst =3D=3D BitClean && src =3D=3D BitUnwritten) || + (src =3D=3D BitClean && dst =3D=3D BitUnwritten))) + return BitClean; + if (dst =3D=3D BitClean || src =3D=3D BitClean) + return BitClean; + return BitUnwritten; +} + static void llbitmap_init_state(struct llbitmap *llbitmap) { struct mddev *mddev =3D llbitmap->mddev; @@ -1306,6 +1346,7 @@ static void md_llbitmap_daemon_fn(struct work_struct = *work) =20 if (llbitmap->mddev->degraded) return; + retry: start =3D 0; end =3D min(llbitmap->chunks, PAGE_SIZE - BITMAP_DATA_OFFSET) - 1; @@ -1367,6 +1408,7 @@ static int llbitmap_create(struct mddev *mddev) =20 timer_setup(&llbitmap->pending_timer, llbitmap_pending_timer_fn, 0); INIT_WORK(&llbitmap->daemon_work, md_llbitmap_daemon_fn); + rwlock_init(&llbitmap->reshape_lock); atomic_set(&llbitmap->behind_writes, 0); init_waitqueue_head(&llbitmap->behind_wait); =20 @@ -1535,7 +1577,9 @@ static void llbitmap_start_write(struct mddev *mddev,= sector_t offset, page_start++; } =20 + read_lock(&llbitmap->reshape_lock); llbitmap_state_machine(llbitmap, start, end, BitmapActionStartwrite); + read_unlock(&llbitmap->reshape_lock); } =20 static void llbitmap_end_write(struct mddev *mddev, sector_t offset, @@ -1567,7 +1611,9 @@ static void llbitmap_start_discard(struct mddev *mdde= v, sector_t offset, page_start++; } =20 + read_lock(&llbitmap->reshape_lock); llbitmap_state_machine(llbitmap, start, end, BitmapActionDiscard); + read_unlock(&llbitmap->reshape_lock); } =20 static void llbitmap_end_discard(struct mddev *mddev, sector_t offset, @@ -1864,6 +1910,136 @@ static int llbitmap_reshape_can_start(struct mddev = *mddev) return ret; } =20 +struct llbitmap_reshape_range { + sector_t offset; + unsigned long sectors; + sector_t start; + sector_t end; +}; + +static enum llbitmap_state +llbitmap_reshape_init_dst(struct llbitmap *llbitmap, unsigned long dst, + const struct llbitmap_reshape_range *new) +{ + u64 bit_start =3D (u64)dst * llbitmap->reshape_chunksize; + u64 bit_end =3D bit_start + llbitmap->reshape_chunksize; + + if (!llbitmap->mddev->reshape_backwards) + return bit_start < new->offset ? llbitmap_read(llbitmap, dst) : + BitUnwritten; + return bit_end > new->end ? llbitmap_read(llbitmap, dst) : BitUnwritten; +} + +static void llbitmap_reshape_dst_range(struct llbitmap *llbitmap, + unsigned long dst, + const struct llbitmap_reshape_range *new, + struct llbitmap_reshape_range *dst_range) +{ + sector_t dst_bit_start =3D (sector_t)dst * llbitmap->reshape_chunksize; + + dst_range->start =3D max(dst_bit_start, new->offset); + dst_range->end =3D min(dst_bit_start + llbitmap->reshape_chunksize, + new->end); + dst_range->offset =3D dst_range->start; + dst_range->sectors =3D dst_range->end - dst_range->start; +} + +static void llbitmap_reshape_map_range(struct llbitmap *llbitmap, + sector_t lo, sector_t hi, + bool previous, + struct llbitmap_reshape_range *range) +{ + range->offset =3D lo; + range->sectors =3D hi - lo; + llbitmap_map_layout(llbitmap, &range->offset, &range->sectors, previous); + range->start =3D range->offset; + range->end =3D range->offset + range->sectors; +} + +static bool llbitmap_reshape_src_range(const struct llbitmap_reshape_range= *old, + const struct llbitmap_reshape_range *new, + const struct llbitmap_reshape_range *dst, + struct llbitmap_reshape_range *src) +{ + if (!old->sectors) + return false; + + src->start =3D old->offset + + mul_u64_u64_div_u64(dst->start - new->offset, + old->sectors, new->sectors); + src->end =3D old->offset + + mul_u64_u64_div_u64_roundup(dst->end - new->offset, + old->sectors, new->sectors); + if (src->end > old->end) + src->end =3D old->end; + src->offset =3D src->start; + src->sectors =3D src->end - src->start; + + return src->sectors; +} + +static enum llbitmap_state llbitmap_rmerge_src(struct llbitmap *llbitmap, + enum llbitmap_state state, + const struct llbitmap_reshape_range *src) +{ + unsigned long bit =3D div64_u64(src->start, llbitmap->chunksize); + unsigned long end =3D div64_u64(src->end - 1, llbitmap->chunksize); + + while (bit <=3D end) { + enum llbitmap_state src_state =3D llbitmap_read(llbitmap, bit); + + state =3D llbitmap_rmerge_state(llbitmap, state, src_state); + bit++; + } + + return state; +} + +static void llbitmap_reshape_merge(struct llbitmap *llbitmap, + const struct llbitmap_reshape_range *old, + const struct llbitmap_reshape_range *new) +{ + unsigned long dst_start; + unsigned long dst_end; + unsigned long dst; + bool backwards =3D false; + + if (!new->sectors) + return; + + dst_start =3D div64_u64(new->offset, llbitmap->reshape_chunksize); + dst_end =3D div64_u64(new->end - 1, llbitmap->reshape_chunksize); + if (old->sectors) { + unsigned long src_start =3D div64_u64(old->offset, + llbitmap->chunksize); + unsigned long src_end =3D div64_u64(old->end - 1, + llbitmap->chunksize); + + backwards =3D src_start < dst_start && src_end >=3D dst_start; + } + + dst =3D backwards ? dst_end : dst_start; + while (true) { + struct llbitmap_reshape_range dst_range; + struct llbitmap_reshape_range src; + enum llbitmap_state state; + + llbitmap_reshape_dst_range(llbitmap, dst, new, &dst_range); + state =3D llbitmap_reshape_init_dst(llbitmap, dst, new); + if (llbitmap_reshape_src_range(old, new, &dst_range, &src)) + state =3D llbitmap_rmerge_src(llbitmap, state, &src); + else + state =3D llbitmap_rmerge_state(llbitmap, state, BitUnwritten); + llbitmap_write(llbitmap, state, dst); + if (dst =3D=3D (backwards ? dst_start : dst_end)) + break; + if (backwards) + dst--; + else + dst++; + } +} + static void llbitmap_reshape_finish(struct mddev *mddev) { struct llbitmap *llbitmap =3D mddev->bitmap; @@ -1888,6 +2064,33 @@ static void llbitmap_reshape_finish(struct mddev *md= dev) mddev->pers->quiesce(mddev, 0); } =20 +static void llbitmap_reshape_mark(struct mddev *mddev, sector_t old_pos, + sector_t new_pos) +{ + struct llbitmap *llbitmap =3D mddev->bitmap; + sector_t lo; + sector_t hi; + struct llbitmap_reshape_range old; + struct llbitmap_reshape_range new; + + if (!llbitmap || old_pos =3D=3D new_pos) + return; + + lo =3D min(old_pos, new_pos); + hi =3D max(old_pos, new_pos); + if (!hi) + return; + + llbitmap_reshape_map_range(llbitmap, lo, hi, true, &old); + llbitmap_reshape_map_range(llbitmap, lo, hi, false, &new); + if (!new.sectors) + return; + + write_lock(&llbitmap->reshape_lock); + llbitmap_reshape_merge(llbitmap, &old, &new); + write_unlock(&llbitmap->reshape_lock); +} + static void llbitmap_write_sb(struct llbitmap *llbitmap) { int nr_blocks =3D DIV_ROUND_UP(BITMAP_DATA_OFFSET, llbitmap->io_size); @@ -2181,6 +2384,7 @@ static struct bitmap_operations llbitmap_ops =3D { .prepare_range =3D llbitmap_prepare_range, .reshape_finish =3D llbitmap_reshape_finish, .reshape_can_start =3D llbitmap_reshape_can_start, + .reshape_mark =3D llbitmap_reshape_mark, .write_all =3D llbitmap_write_all, =20 .groups =3D md_llbitmap_groups, --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D126B320393; Sun, 2 Aug 2026 19:52:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700327; cv=none; b=VdeONTZcx7D0xgvMZ4Bfky6Z4H5K9Sv1zF+NBoM1UzvLI4zWiczakaH2VQ3RUCyl0RFmhsKz4v3QEoPG1Z1nkKZp5Jskvj6vI/Q6Pi4+uPh3cIr7VflQPwX1Jsl5B94YqPKeD9JicwIDzMfAmsHqJ8cos6lGIQhFFkvA2OmBiJ4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700327; c=relaxed/simple; bh=jG4tqnmF+fhDnanNJAu01qZZRqA7JEfETRe29Ejqdck=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kAs+amfKBSLWFTCRl1mq1wyoKwbiT2tXUsO89D4TYDtRy8nU5YqoQg1HcUTsZBRHHh0SYA1PjMHp8pH+B4NaP8ZRBhKz5F1sQHcKoZNjepJkCyyq9SFvtd4fsz7aOHbSlZs/XQHUTvzPXn5GHmu3vPB/amOvLaJHBxwx7Hvw3j4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Bukuf0lw; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Bukuf0lw" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2F80C1F000E9; Sun, 2 Aug 2026 19:52:03 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700326; bh=rqH4Xu/E74LiTszdv89UTWG+kpu8CAVJ+jzGMnMkj4U=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=Bukuf0lwPtf2Yto9s0JRutyGeam5kbdHsCg/npj/sLjuCSDsG03g16VAGPS88khqn P+lfQNz4U1gcGwU/uwPNttxizNVSWviEuuYQe4qXuL5fs7l5h6AKqZuSTcqfJG/OU/ UbVsL/3jAmAmDtPc5yAbze+TAnPVRtEb9En0jE7ZuFzLAVePr1V8RHapDCRjNLuGvT Impir9WTiM9+uwWnlZ9z2bZzq8QAfg1OyA17pRbBQ5r4xxc8za8QP8mIAjXil0jTmo ydH2p8/d47HWBr2BogGYOVHXptOU3ntuQr2jxLZEKbc8I6JT3lKxE6+xYDi6+F/DbM vj/kvnUv2kaVw== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 22/29] md/md-llbitmap: clamp state-machine walks to tracked bits Date: Mon, 3 Aug 2026 03:50:31 +0800 Message-ID: <20260802195038.164272-23-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai llbitmap_state_machine() can be called with an end bit beyond llbitmap->chunks. In particular, llbitmap_cond_end_sync() passes sector >> chunkshift, and sector can reach the tracked boundary exactly. Clamp the state-machine range to llbitmap->chunks so it cannot walk past the tracked bitmap. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/md-llbitmap.c | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c index 5d95627ff983..2830ce05d941 100644 --- a/drivers/md/md-llbitmap.c +++ b/drivers/md/md-llbitmap.c @@ -1012,7 +1012,10 @@ static enum llbitmap_state llbitmap_state_machine(st= ruct llbitmap *llbitmap, llbitmap_init_state(llbitmap); return BitNone; } - + if (start >=3D llbitmap->chunks) + return BitNone; + if (end >=3D llbitmap->chunks) + end =3D llbitmap->chunks - 1; while (start <=3D end) { enum llbitmap_state c =3D llbitmap_read(llbitmap, start); =20 --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C430E3002B9; Sun, 2 Aug 2026 19:52:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700331; cv=none; b=h1d3+6odtg98hmauGNb/Ji5Xa1e9wkh5a8EGcPW0BJv/5DotyA/AWr5vVv/GGXGlMulxZlO9vp9/FKF7rB5Tq8ybCQJOGTy8TtUxaaClBXZpLSwqVD8lYSbC/TBBBL3L7GZJb/v+hkGk+QURWcwiumnBqr/oZS0YYhvC+64ECq0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700331; c=relaxed/simple; bh=lMhVaco/SDj4s/Zyd25s68j2bJCfLMQMZnOj6XaWJiA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=vB8vT/dt5wRgGaTYru42s1IaLnCJIS/OytZeNmv8+NqNC2sxxiJSIQrnbklBBaVPCnK1pH7MOq7d0zksw6ns83BOhjR3QZ/435KnzQDBlHJl/xRU33+ef/8/rZBKYNIgh0h74FvyCFsK1e2u985/x9xngzaRRXcZ3OCLCA0lGbI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=cXO92nPH; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="cXO92nPH" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2165E1F000E9; Sun, 2 Aug 2026 19:52:07 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700330; bh=pVx/726z0oTAQs76I7NOHZZBH/dzptQlPlB6PmCSfVQ=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=cXO92nPHhawLH2NnRny0sy4J+/rOcXNDSbGsJ+wpkjU8tbiCag0+yFvHWDXM1zAeb MVdO1WtBesDT4nIW9hP6ppC4FPtZmbmS1KOvkM02nhKlYpKd7QOcH8d5NZRdcQVlIW H4+qYFtir1z04pUO29+hlmmlvFISHdXB16NkAUEwhignKt/oQ2i2skqwE4HB5hOvxp xkxUUqFXGk9BpcIvz9jSmFFM2vQrYQqzQhtAhFRnJEAP4tEtuAlUxcLtSTyU8h8nYl OQSwSYvFgPVe/tYTJ/8A/3AbgTG4Lle/xNFmhihQpM/B8YWxMPp4kasKzlA14gTqOK UZexARfkOyBlA== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 23/29] md/raid10: reject llbitmap reshape when md chunk shrinks Date: Mon, 3 Aug 2026 03:50:32 +0800 Message-ID: <20260802195038.164272-24-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai llbitmap reshape keeps one live bitmap and cannot safely make an existing bitmap bit cover a smaller data range. The llbitmap chunksize itself will not shrink when mddev->chunk_sectors stays the same or grows. However, shrinking mddev->chunk_sectors can shrink the effective data range covered by each bit for the RAID10 reshape geometry. Reject that reshape while llbitmap is active. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid10.c | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c index ed3c6fbe65f7..1c3393467667 100644 --- a/drivers/md/raid10.c +++ b/drivers/md/raid10.c @@ -4245,6 +4245,10 @@ static int raid10_check_reshape(struct mddev *mddev) =20 if (conf->geo.far_copies !=3D 1 && !conf->geo.far_offset) return -EINVAL; + if (mddev->bitmap_id =3D=3D ID_LLBITMAP && + mddev->new_chunk_sectors && + mddev->new_chunk_sectors < mddev->chunk_sectors) + return -EOPNOTSUPP; =20 if (setup_geo(&geo, mddev, geo_start) !=3D conf->copies) /* mustn't change number of copies */ --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 48EDF33F368; Sun, 2 Aug 2026 19:52:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700335; cv=none; b=XEzEZclD5JTnmY8caj0QOLJTontA3h2CpC8gcuvTLTkPynZm88/ki2ew0/RegpPUrHUqeR9MbUj7zMI+woLaMOiKuaslUAh6YZ8ncrzwQB0yUamjviFuQ+O1lxvezIggY3NIpV9J3cvPCusfZOudbTJo7qFIdfvhJUqZBItRJkQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700335; c=relaxed/simple; bh=EuyIktEHVmA5SmDcvWqhEPRJ9r0PZUtvrnevxd2wDNo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=IRPpiCNQcARyQnNis+kWP5r4p1prIT2dKkpFujV52GODLIVFpGIfMOcXMzRkbcxUXRJGpwbyQde1ErQZj4fH6EJqfDhUEnfdtmh2FCDzxjR0BOIugR3ul0jkVNLmHuk9aJ2hl7IYquKWIgpQPzvneMeTMY7mqNaeR9QQCqKZ6Rw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Skc7f0jE; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Skc7f0jE" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A1A061F00A3A; Sun, 2 Aug 2026 19:52:11 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700333; bh=MfCLfvgIrt08NkdV2kT4xWFejRZpsGF464/3VXQmSn8=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=Skc7f0jEaxTWlz0zSbTySgQhGe1SDj/n9CVuiLIHjX3vYKYpEkGO1lIgjMqVSJAwl ehEPkVS8KirMQAoqVA5yErpX7WQnHWkvcBrUepueykexv6fG5RpafyRzZiYbjs9EpN 89lce34qyrHiw0oiq/fDLxLEec1nxAFA6QV43lUNhbAmATxGf8Jhajvbwm0+c/7+aw GEk9Ex99SJBvptS87EznRefFQOR5lzcjwcNXixrMwytgmkTaCtlE4LjSp6OzmYCXik 8ZlACHVkNvj0IRrXGBwY0EaPy1BEEVqKZtLZO6EoTJvLow/3xajVavydHMbWhOdDQB krFYsBVnC1htg== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 24/29] md/raid10: wire llbitmap reshape lifecycle Date: Mon, 3 Aug 2026 03:50:33 +0800 Message-ID: <20260802195038.164272-25-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Prepare llbitmap before RAID10 starts growing, checkpoint the bitmap before advancing reshape_position, finish the llbitmap geometry update when reshape completes, and export the old and new tracked sizes. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid10.c | 39 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 39 insertions(+) diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c index 1c3393467667..bac9edd28c97 100644 --- a/drivers/md/raid10.c +++ b/drivers/md/raid10.c @@ -4356,6 +4356,12 @@ static int raid10_start_reshape(struct mddev *mddev) =20 if (test_bit(MD_RECOVERY_RUNNING, &mddev->recovery)) return -EBUSY; + if (md_bitmap_enabled(mddev, false) && + mddev->bitmap_ops->reshape_can_start) { + ret =3D mddev->bitmap_ops->reshape_can_start(mddev); + if (ret) + return ret; + } =20 if (setup_geo(&new, mddev, geo_start) !=3D conf->copies) return -EINVAL; @@ -4679,6 +4685,13 @@ static sector_t reshape_request(struct mddev *mddev,= sector_t sector_nr, time_after(jiffies, conf->reshape_checkpoint + 10*HZ)) { /* Need to update reshape_position in metadata */ wait_barrier(conf); + if (md_bitmap_enabled(mddev, false) && + mddev->bitmap_ops->reshape_mark && + conf->reshape_safe !=3D conf->reshape_progress) { + mddev->bitmap_ops->reshape_mark(mddev, conf->reshape_safe, + conf->reshape_progress); + mddev->bitmap_ops->unplug(mddev, true); + } mddev->reshape_position =3D conf->reshape_progress; if (mddev->reshape_backwards) mddev->curr_resync_completed =3D raid10_size(mddev, 0, 0) @@ -4877,9 +4890,19 @@ static void reshape_request_write(struct mddev *mdde= v, struct r10bio *r10_bio) =20 static void end_reshape(struct r10conf *conf) { + struct mddev *mddev =3D conf->mddev; + if (test_bit(MD_RECOVERY_INTR, &conf->mddev->recovery)) return; =20 + if (md_bitmap_enabled(mddev, false) && + mddev->bitmap_ops->reshape_mark && + conf->reshape_safe !=3D conf->reshape_progress) { + mddev->bitmap_ops->reshape_mark(mddev, conf->reshape_safe, + conf->reshape_progress); + mddev->bitmap_ops->unplug(mddev, true); + } + spin_lock_irq(&conf->device_lock); conf->prev =3D conf->geo; md_finish_reshape(conf->mddev); @@ -5011,10 +5034,15 @@ static void end_reshape_request(struct r10bio *r10_= bio) static void raid10_finish_reshape(struct mddev *mddev) { struct r10conf *conf =3D mddev->private; + bool llbitmap =3D mddev->bitmap_id =3D=3D ID_LLBITMAP && + md_bitmap_enabled(mddev, false); =20 if (test_bit(MD_RECOVERY_INTR, &mddev->recovery)) return; =20 + if (llbitmap && mddev->bitmap_ops->reshape_finish) + mddev->bitmap_ops->reshape_finish(mddev); + if (mddev->delta_disks > 0) { if (mddev->resync_offset > mddev->resync_max_sectors) { mddev->resync_offset =3D mddev->resync_max_sectors; @@ -5041,6 +5069,15 @@ static void raid10_finish_reshape(struct mddev *mdde= v) mddev->reshape_backwards =3D 0; } =20 +static sector_t raid10_bitmap_sync_size(struct mddev *mddev, bool previous) +{ + struct r10conf *conf =3D mddev->private; + + if (previous) + return raid10_size(mddev, 0, 0); + return raid10_size(mddev, 0, conf->geo.raid_disks); +} + static struct md_personality raid10_personality =3D { .head =3D { @@ -5067,6 +5104,8 @@ static struct md_personality raid10_personality =3D .start_reshape =3D raid10_start_reshape, .finish_reshape =3D raid10_finish_reshape, .update_reshape_pos =3D raid10_update_reshape_pos, + .bitmap_sync_size =3D raid10_bitmap_sync_size, + .bitmap_array_sectors =3D raid10_bitmap_sync_size, }; =20 static int __init raid10_init(void) --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3B01533F597; Sun, 2 Aug 2026 19:52:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700338; cv=none; b=eVIKIgPXjA2AWf4EXhvh62jczbL4GFi/VEZT+xJEZYGQeQYfJhZn22wHmfnWFtcVZN+8ILKfcfVi43YW7OUGO1ZKI+4O1atuTWsy/hGBN/k0fj8/c8TcSjT2cnU9wyU8HRZQ6CPg5yhEWIqesMuPHQqQUunI8SfO/MCfkyZfNUc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700338; c=relaxed/simple; bh=I0WCHnswMmVNxkmOkSkooY3CurCC16eDuCPucYJWViw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=R2vXuF0yUYtjFvEAI36JWFoymiSjkYvCZImFxBtTRXdQoTPzfD7mUrNS1+mFbapUlS2+BXhxrIKqgG9yZXdzeJlMIFwErD96FIEzQSEl+/2gZy7a1RnisnWfoejs3bxIOwen3cLe8I0BOOuhTwigNbTmliWKJM/qLXvIYTJvaE0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=N+qW6xsv; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="N+qW6xsv" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 499941F000E9; Sun, 2 Aug 2026 19:52:14 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700336; bh=ryAAtN55XBalJyhi3qRADAZhr6I0H9xs6ZEKweeOw94=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=N+qW6xsvmF0Y/da+OwptuVbTwnomN7DnZqa+1iN4Of3+TeuRvNhpNlQe7r3V61qIt LvpCKj3cpBVlr3FCWh8RW76KAnuyiyLya2Epy5baaCEZAiICFolpRTL1g0aPsFx3d2 5E8bvPTOuwYp07ASSebDRQa8I9bSwCCqKsgl+YtEcBEh/FqhGVgPtabr5Gax8MXBJg fihzgXLfo+CoGzQmf/AdpeBW+DQeOQQDFvFmYhbqv2KI29k5gcJsN8BnkAlS6zo3c5 3Re59O8Sx8CYRlcelWa/NAYBZ3n7e5mpz09+6mx4DPHN4iXj3AyLMvtHyqFkKL0hYJ 08X/O2bVcGy1w== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 25/29] md/raid10: split reshape bios before bitmap accounting Date: Mon, 3 Aug 2026 03:50:34 +0800 Message-ID: <20260802195038.164272-26-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Use the shared mddev_bio_split_at_reshape_offset() helper so RAID10 submits only one-side bios to llbitmap during reshape. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid10.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/drivers/md/raid10.c b/drivers/md/raid10.c index bac9edd28c97..562a325a7195 100644 --- a/drivers/md/raid10.c +++ b/drivers/md/raid10.c @@ -1848,6 +1848,7 @@ static bool raid10_make_request(struct mddev *mddev, = struct bio *bio) { struct r10conf *conf =3D mddev->private; sector_t chunk_mask =3D (conf->geo.chunk_mask & conf->prev.chunk_mask); + const int rw =3D bio_data_dir(bio); int chunk_sects =3D chunk_mask + 1; int sectors =3D bio_sectors(bio); =20 @@ -1873,6 +1874,15 @@ static bool raid10_make_request(struct mddev *mddev,= struct bio *bio) sectors =3D chunk_sects - (bio->bi_iter.bi_sector & (chunk_sects - 1)); + + bio =3D mddev_bio_split_at_reshape_offset(mddev, bio, §ors, + &conf->bio_split); + if (!bio) { + if (rw =3D=3D WRITE) + md_write_end(mddev); + return true; + } + if (!__make_request(mddev, bio, sectors)) md_write_end(mddev); =20 --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 75FBC2DA756; Sun, 2 Aug 2026 19:52:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700344; cv=none; b=YjEgbGRdjNy2DPbuZR8xBkgT1kxDhP3DLT2eW4nhFyDCbUwMkJcfffmAom+AiZWld/b27OmuR+7HFFuQdDMG3NDM+tUM05AM9nUWYNIzXSX2RmzHLN0993yZiGBw2BFtYIgkwymvasZVbgn6JK95FcRWbLX2DZZW8xZaoRlKKxA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700344; c=relaxed/simple; bh=zuBUKN/0By9bwTDiXX3dIKJ0U17ZZbdQqzAgTaYWYCE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BJ1e2ZipyK6T2MqmERKX/FrnZ4a98DuZf4eBiM2FGAxK9vU2+/TfZJrJiEZICFJWa/CCC54oOrwKEERzMta1NvKm5ISUNUt6S/tt1danACxR5dd/bAaBsC9uQzqZuCI0no+WwyQKJGv1B/xpuqkA0P1Ui3luTiXtlBwtKMEQ6eA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=aq7aNS2R; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="aq7aNS2R" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 651F11F00A3A; Sun, 2 Aug 2026 19:52:17 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700343; bh=noao0UvL+Onfed2Lrp0Whb9V04ME13xSoFCoKYK5utY=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=aq7aNS2RFGfzHNKMd2yZ/p1tDjP9Dwb98WFh7UgT0uEX82zLl9N/yMfno+WNdzquv 3pt1Yn5OoZ62y56iANmw0L3l602jcFYFWBwk4YhpcintRA97Cl8kGVn42i8PFfxt2E +CV8Se5OBNV/TYJHM8crxoDkhTX9SnIJmLDVgpJb4VChjIKNayA4SwZTpYLZnS+eeb s/L5FYUiT38lk/RW/eIMbd9m1G8NDk2Y9SPPMGEPu5nPj/+tpJ21cewhjLI2eAVlWW no3nFZU0l44LKoZHxjV6lFxrs8unFv1XhW1vb1cOUse7l6aeEcRKP85CIehtu779s3 tvlr5unwZcriA== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 26/29] md/raid5: add exact old and new llbitmap mapping helpers Date: Mon, 3 Aug 2026 03:50:35 +0800 Message-ID: <20260802195038.164272-27-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Teach RAID5 to export exact old and new llbitmap mappings and the corresponding sync and array sizes for reshape-aware bitmap users. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid5.c | 73 +++++++++++++++++++++++++++++++++------------- 1 file changed, 53 insertions(+), 20 deletions(-) diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c index 2cc2546a29ae..88bf5a9ce573 100644 --- a/drivers/md/raid5.c +++ b/drivers/md/raid5.c @@ -6015,28 +6015,46 @@ static enum reshape_loc get_reshape_loc(struct mdde= v *mddev, return LOC_BEHIND_RESHAPE; } =20 -static void raid5_bitmap_sector(struct mddev *mddev, sector_t *offset, - unsigned long *sectors) +static void raid5_bitmap_sector_map(struct mddev *mddev, sector_t *offset, + unsigned long *sectors, + bool previous) { struct r5conf *conf =3D mddev->private; sector_t start =3D *offset; sector_t end =3D start + *sectors; - sector_t prev_start =3D start; - sector_t prev_end =3D end; int sectors_per_chunk; - enum reshape_loc loc; int dd_idx; =20 - sectors_per_chunk =3D conf->chunk_sectors * - (conf->raid_disks - conf->max_degraded); + if (previous) + sectors_per_chunk =3D conf->prev_chunk_sectors * + (conf->previous_raid_disks - conf->max_degraded); + else + sectors_per_chunk =3D conf->chunk_sectors * + (conf->raid_disks - conf->max_degraded); sector_div(start, sectors_per_chunk); start *=3D sectors_per_chunk; if (sector_div(end, sectors_per_chunk)) end++; end *=3D sectors_per_chunk; =20 - start =3D raid5_compute_sector(conf, start, 0, &dd_idx, NULL); - end =3D raid5_compute_sector(conf, end, 0, &dd_idx, NULL); + start =3D raid5_compute_sector(conf, start, previous, &dd_idx, NULL); + end =3D raid5_compute_sector(conf, end, previous, &dd_idx, NULL); + *offset =3D start; + *sectors =3D end - start; +} + +static void raid5_bitmap_sector(struct mddev *mddev, sector_t *offset, + unsigned long *sectors) +{ + struct r5conf *conf =3D mddev->private; + sector_t start =3D *offset; + sector_t end =3D start + *sectors; + sector_t prev_start =3D start; + unsigned long prev_sectors =3D end - start; + enum reshape_loc loc; + + raid5_bitmap_sector_map(mddev, &start, sectors, false); + end =3D start + *sectors; =20 /* * For LOC_INSIDE_RESHAPE, this IO will wait for reshape to make @@ -6045,19 +6063,10 @@ static void raid5_bitmap_sector(struct mddev *mddev= , sector_t *offset, loc =3D get_reshape_loc(mddev, conf, prev_start); if (likely(loc !=3D LOC_AHEAD_OF_RESHAPE)) { *offset =3D start; - *sectors =3D end - start; return; } =20 - sectors_per_chunk =3D conf->prev_chunk_sectors * - (conf->previous_raid_disks - conf->max_degraded); - sector_div(prev_start, sectors_per_chunk); - prev_start *=3D sectors_per_chunk; - sector_div(prev_end, sectors_per_chunk); - prev_end *=3D sectors_per_chunk; - - prev_start =3D raid5_compute_sector(conf, prev_start, 1, &dd_idx, NULL); - prev_end =3D raid5_compute_sector(conf, prev_end, 1, &dd_idx, NULL); + raid5_bitmap_sector_map(mddev, &prev_start, &prev_sectors, true); =20 /* * for LOC_AHEAD_OF_RESHAPE, reshape can make progress before this IO @@ -6065,7 +6074,7 @@ static void raid5_bitmap_sector(struct mddev *mddev, = sector_t *offset, * we set bits for both. */ *offset =3D min(start, prev_start); - *sectors =3D max(end, prev_end) - *offset; + *sectors =3D max(end, prev_start + prev_sectors) - *offset; } =20 static enum stripe_result make_stripe_request(struct mddev *mddev, @@ -9131,6 +9140,21 @@ static void raid5_prepare_suspend(struct mddev *mdde= v) wake_up(&conf->wait_for_reshape); } =20 +static sector_t raid5_bitmap_sync_size(struct mddev *mddev, bool previous) +{ + return mddev->dev_sectors; +} + +static sector_t raid5_bitmap_array_sectors(struct mddev *mddev, bool previ= ous) +{ + struct r5conf *conf =3D mddev->private; + + if (previous) + return raid5_size(mddev, mddev->dev_sectors, + conf->previous_raid_disks); + return raid5_size(mddev, mddev->dev_sectors, conf->raid_disks); +} + static struct md_personality raid6_personality =3D { .head =3D { @@ -9160,6 +9184,9 @@ static struct md_personality raid6_personality =3D .change_consistency_policy =3D raid5_change_consistency_policy, .prepare_suspend =3D raid5_prepare_suspend, .bitmap_sector =3D raid5_bitmap_sector, + .bitmap_sector_map =3D raid5_bitmap_sector_map, + .bitmap_sync_size =3D raid5_bitmap_sync_size, + .bitmap_array_sectors =3D raid5_bitmap_array_sectors, }; static struct md_personality raid5_personality =3D { @@ -9190,6 +9217,9 @@ static struct md_personality raid5_personality =3D .change_consistency_policy =3D raid5_change_consistency_policy, .prepare_suspend =3D raid5_prepare_suspend, .bitmap_sector =3D raid5_bitmap_sector, + .bitmap_sector_map =3D raid5_bitmap_sector_map, + .bitmap_sync_size =3D raid5_bitmap_sync_size, + .bitmap_array_sectors =3D raid5_bitmap_array_sectors, }; =20 static struct md_personality raid4_personality =3D @@ -9221,6 +9251,9 @@ static struct md_personality raid4_personality =3D .change_consistency_policy =3D raid5_change_consistency_policy, .prepare_suspend =3D raid5_prepare_suspend, .bitmap_sector =3D raid5_bitmap_sector, + .bitmap_sector_map =3D raid5_bitmap_sector_map, + .bitmap_sync_size =3D raid5_bitmap_sync_size, + .bitmap_array_sectors =3D raid5_bitmap_array_sectors, }; =20 static int __init raid5_init(void) --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 60540339363; Sun, 2 Aug 2026 19:52:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700348; cv=none; b=WTakxhwl+z7UmPp/ObJbpfHbpZcdVho8JOkel8PeeEa70d8rcq7kX58PXNo57KrOb99gIPJ/bVasKtGcM3R5/IpZ6yex7iPgz14e+QYUb8ll3JioOB4lp2VO0nQOxSHEKsiwQR8pZ5Ql3ARLSkWSJ8ndBMOtMr8a0mCulzugp4g= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700348; c=relaxed/simple; bh=Rqgvn/E7BhWafURAISqtwi+YnqhKX7fLHXfIH5DzC3w=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=p6gYXrHatMqEOwteB0bXHtSz1goUC31QnYd0r+3N1VWXI1E7F0nd4+lS5ScGvFL1fCMq4YJP/R65xXOPLkDIf2hNRvM3ssip/vrbJ8Vaqc3UB7NuRPPC/Rkvy1UxkB4ZbiDPqwgAyMs7OfoniRweJ0o9VGn3yU/Sl8PG7SVMvIE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=JbL08JSx; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="JbL08JSx" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DFBC01F000E9; Sun, 2 Aug 2026 19:52:23 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700347; bh=4PXxOgVQSzPnexE6Ra77DyXCOwfRDzIjcDTA1bHcSRE=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=JbL08JSxH8dIlxNWSV7R9Iac/PYrDzFTUZGfv+Q6SVe4KO1bJuZ+dtoVnWG9XDlsG P28ZxkogpCkSy76bmU5dc23kGH/TrwP/Xi5y4HdSwMfkACfK+FNr77Aodw8nFOTLNp EihXvqwcgzrhR3Sn7gOAEvtZRcgJ60L1h9ZnvLxc1AHb5iWhewhQoVBKlDWCdLkChy xSFtD066tmEiaaY9Lp0iPXCindcCoT9Utn4Yre4x6hgQN0u7AtdeZyLgwLwirEpEUl R7g9vOGOBoonbe9rWtJOd9l6jPoYIn15/4MIoNCSAmmF3jL+YQ/oGM/894Ke+mgms4 zutsy6vJa/p5A== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 27/29] md/raid5: reject llbitmap reshape when md chunk shrinks Date: Mon, 3 Aug 2026 03:50:36 +0800 Message-ID: <20260802195038.164272-28-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai llbitmap reshape keeps one live bitmap and cannot safely make an existing bitmap bit cover a smaller data range. The llbitmap chunksize itself will not shrink when mddev->chunk_sectors stays the same or grows. However, shrinking mddev->chunk_sectors shrinks sectors_per_chunk used by raid5_bitmap_sector_map(). That can shrink the effective data range covered by each bit across the old and new RAID5 geometry. Reject that reshape while llbitmap is active. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid5.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c index 88bf5a9ce573..67d56c92c8a4 100644 --- a/drivers/md/raid5.c +++ b/drivers/md/raid5.c @@ -8580,6 +8580,9 @@ static int check_reshape(struct mddev *mddev) if (!check_stripe_cache(mddev)) return -ENOSPC; =20 + if (mddev->bitmap_id =3D=3D ID_LLBITMAP && + mddev->new_chunk_sectors < mddev->chunk_sectors) + return -EOPNOTSUPP; if (mddev->new_chunk_sectors > mddev->chunk_sectors || mddev->delta_disks > 0) if (resize_chunks(conf, --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 78CE133C518; Sun, 2 Aug 2026 19:52:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700351; cv=none; b=UHkXWVEdK4Vhbkl1dh7OHlPehw/8H8L8aiOvy3T4iFgrr3Xo5XKMx3CiU9NWacMKJN4ZQUynPo/vN20X2+7X7saqMdYDwi+lQikWGffOSEgxLFn3LgOvZzOuydlwxetVdLHdYnKMryc+27Tv+/RDRwyQCoohCKWKdgujxPdpLeM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700351; c=relaxed/simple; bh=ewheLMNF6mpFLUG/mnveS/nXV6Nb/QOHQrhC0TwSGhk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=S/rWrzqZPNa+BHLcBEg48I5UqMZi2RGd3s4w1LpDr4P0WTpibQb/xb0aRclJgFMVeEyJStTbBpc72qglEWG2dKnQ6VB/LllgpTm7/Ur8f4O+wI/9N/qeOlIE62N0Vcwijwno9+WSJFZyJORuRlnLrmxgVp7SxkFt6nZ2jrNd1rE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=lPp+ntOm; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="lPp+ntOm" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3BD251F00A3A; Sun, 2 Aug 2026 19:52:28 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700350; bh=A1vLhIXj5KEQVFVbQzvp3M2SPBeKYnjOlBcnZmJBjw4=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=lPp+ntOmYWyx/41SsbiIgzWFSnlL/0YQD5ISCSfkwq+H8YjqojWQBZxG5tnRXWutV eXcSZPru8De9blyliwztyFsUnhzwW2h9xXEhF8Dh9WNA43S3m1NdPl4B3US5/JDUKq cUv9vzTTvTd194p3otDC4Ev05mrbFbPtPPCnc8LfkschMvWZDhmaSij4mO5hW/1v6a xlz6buBOhXW5NsSFTjmMEGOvNmht8e9JvVArDaolLubgUoKML91+ib4SuZo8uePfYQ 92Bp7DATpIUI4lmHLszbdMa3fRlVmAWLZJjkCp+T37+pHmMU+mRpcQRFoZ/vnFnpVC 3dCuALh7+n4RQ== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 28/29] md/raid5: wire llbitmap reshape lifecycle Date: Mon, 3 Aug 2026 03:50:37 +0800 Message-ID: <20260802195038.164272-29-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai Prepare llbitmap before RAID5 reshape starts, checkpoint the bitmap before advancing reshape_position, and finish the llbitmap geometry update when reshape completes. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid5.c | 37 +++++++++++++++++++++++++++++++++++++ 1 file changed, 37 insertions(+) diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c index 67d56c92c8a4..5176de5b5956 100644 --- a/drivers/md/raid5.c +++ b/drivers/md/raid5.c @@ -6497,6 +6497,13 @@ static sector_t reshape_request(struct mddev *mddev,= sector_t sector_nr, int *sk || test_bit(MD_RECOVERY_INTR, &mddev->recovery)); if (atomic_read(&conf->reshape_stripes) !=3D 0) return 0; + if (md_bitmap_enabled(mddev, false) && + mddev->bitmap_ops->reshape_mark && + conf->reshape_safe !=3D conf->reshape_progress) { + mddev->bitmap_ops->reshape_mark(mddev, conf->reshape_safe, + conf->reshape_progress); + mddev->bitmap_ops->unplug(mddev, true); + } mddev->reshape_position =3D conf->reshape_progress; mddev->curr_resync_completed =3D sector_nr; if (!mddev->reshape_backwards) @@ -6606,6 +6613,13 @@ static sector_t reshape_request(struct mddev *mddev,= sector_t sector_nr, int *sk || test_bit(MD_RECOVERY_INTR, &mddev->recovery)); if (atomic_read(&conf->reshape_stripes) !=3D 0) goto ret; + if (md_bitmap_enabled(mddev, false) && + mddev->bitmap_ops->reshape_mark && + conf->reshape_safe !=3D conf->reshape_progress) { + mddev->bitmap_ops->reshape_mark(mddev, conf->reshape_safe, + conf->reshape_progress); + mddev->bitmap_ops->unplug(mddev, true); + } mddev->reshape_position =3D conf->reshape_progress; mddev->curr_resync_completed =3D sector_nr; if (!mddev->reshape_backwards) @@ -8648,6 +8662,12 @@ static int raid5_start_reshape(struct mddev *mddev) mdname(mddev)); return -EINVAL; } + if (md_bitmap_enabled(mddev, false) && + mddev->bitmap_id =3D=3D ID_LLBITMAP) { + i =3D mddev->bitmap_ops->resize(mddev, mddev->dev_sectors, 0); + if (i) + return i; + } =20 atomic_set(&conf->reshape_stripes, 0); spin_lock_irq(&conf->device_lock); @@ -8732,10 +8752,19 @@ static int raid5_start_reshape(struct mddev *mddev) */ static void end_reshape(struct r5conf *conf) { + struct mddev *mddev =3D conf->mddev; =20 if (!test_bit(MD_RECOVERY_INTR, &conf->mddev->recovery)) { struct md_rdev *rdev; =20 + if (md_bitmap_enabled(mddev, false) && + mddev->bitmap_ops->reshape_mark && + conf->reshape_safe !=3D conf->reshape_progress) { + mddev->bitmap_ops->reshape_mark(mddev, conf->reshape_safe, + conf->reshape_progress); + mddev->bitmap_ops->unplug(mddev, true); + } + spin_lock_irq(&conf->device_lock); conf->previous_raid_disks =3D conf->raid_disks; md_finish_reshape(conf->mddev); @@ -8762,8 +8791,16 @@ static void raid5_finish_reshape(struct mddev *mddev) { struct r5conf *conf =3D mddev->private; struct md_rdev *rdev; + bool llbitmap =3D mddev->bitmap_id =3D=3D ID_LLBITMAP && + md_bitmap_enabled(mddev, false); =20 if (!test_bit(MD_RECOVERY_INTR, &mddev->recovery)) { + if (llbitmap && mddev->bitmap_ops->reshape_finish) + mddev->bitmap_ops->reshape_finish(mddev); + if (llbitmap) { + mddev->resync_offset =3D 0; + mddev->resync_max_sectors =3D mddev->dev_sectors; + } =20 if (mddev->delta_disks <=3D 0) { int d; --=20 2.51.0 From nobody Fri Oct 2 10:08:01 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3EDCD33C518; Sun, 2 Aug 2026 19:52:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700355; cv=none; b=GYODV9BB8Re3DT4A8GbLTxRX4OxjAvS3AdRtOsMIgp7nFF7UFfJlH5ppx9u5WwD+vuu5fKle3aS7MSd0dQQfwJ1SHCdpW61fyAv9Z6O3ujIMoyz7trOSuXxf1H3ywugk3i0nQ4fsQwItOUSI0ng6o02OYbbEmV2MUFlGbKlYvgc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785700355; c=relaxed/simple; bh=FAjN71/gjsFz604tkKGo3e60EXOL2SxkEFrQ2DSw0WE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rzd1KIxzRRYeO/2xwOKYW2Ikwiis+MOL1edx1+jPME5Z5bb9lVXdbAEa3ypYVBk/7Lcg/xBanJnkJPiBRUJq4PF+HBsBeeSaq1CqaquZhgurk30TWrjZSCAVgnDGpmT7rx0Yo7nwYRXar5lhpArv2HiS7qkFNbBoP1acjTtN5r8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=SUXqtyMn; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="SUXqtyMn" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 533761F000E9; Sun, 2 Aug 2026 19:52:31 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785700353; bh=U5AZPitFOQ2suw5rmOmr7esnHtQb3NLo8m9K2W7ESK8=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=SUXqtyMnGfB5khENlJ1M798OecfyG8cM0fF/Qz0+d//EtvZGMA61Lavi/LzP7nOYr VbTg82DUmqQN2PF0l6ggvQ7Jt/O94P7y5rpCpmipfvAFopNxtxI28w8n/fqKbyJiUE 0TKI/YL//sFrcJdPd7W7YFwATDPzrnm9L8mbCcJp57SUxUUnTRs/9c/iqMnzDzEzGw b+HK7uWLqKQ02mZSm7ofRzk/Uit5IVZM12V7YLszGuzNnC7OFHMUbvmR+mOOh+XvRv X60Fyi9VYsrsY4KeW9RPPFAHeIKKQreTK12BNXX9YKOaqglFZDWa0ALFYL2Ux8YPSi 2Qwv43/243fCQ== From: Yu Kuai To: Song Liu , Li Nan , Xiao Ni Cc: Yu Kuai , linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org, Mykola Marzhan , Su Yue Subject: [PATCH v5 29/29] md/raid5: split reshape bios before bitmap accounting Date: Mon, 3 Aug 2026 03:50:38 +0800 Message-ID: <20260802195038.164272-30-yukuai@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260802195038.164272-1-yukuai@kernel.org> References: <20260802195038.164272-1-yukuai@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Yu Kuai RAID5 maps array sectors through different geometries before and after the reshape position. During llbitmap reshape, md core cannot account one bio against both geometries as a single bitmap range, because the old and new bitmap mappings can cover different chunks. Split bios that cross reshape_position before md_account_bio(), so the bitmap only sees ranges that belong to one side of the reshape boundary. mddev_bio_split_at_reshape_offset() uses bio_submit_split_bioset(), which submits the remainder immediately and returns the front split bio. If that front bio later has to wait for reshape, md_handle_request() must not retry the original bio pointer, because after the split that pointer is the already-submitted remainder. Track whether the split happened, clear the temporary BLK_STS_RESOURCE status after the internal clone completion, and resubmit the front bio directly after the reshape wait. Keep the old return-false retry path for unsplit bios, where md_handle_request() still owns the same bio. Tested-by: Mykola Marzhan Signed-off-by: Yu Kuai --- drivers/md/raid5.c | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c index 5176de5b5956..b91545ce090d 100644 --- a/drivers/md/raid5.c +++ b/drivers/md/raid5.c @@ -6221,9 +6221,11 @@ static bool raid5_make_request(struct mddev *mddev, = struct bio * bi) struct r5conf *conf =3D mddev->private; const int rw =3D bio_data_dir(bi); struct stripe_request_ctx *ctx; + struct bio *front_bio; sector_t logical_sector; enum stripe_result res; int s, stripe_cnt; + bool split =3D false; bool on_wq; =20 if (unlikely(bi->bi_opf & REQ_PREFLUSH)) { @@ -6257,6 +6259,18 @@ static bool raid5_make_request(struct mddev *mddev, = struct bio * bi) return true; } =20 + front_bio =3D bi; + bi =3D mddev_bio_split_at_reshape_offset(mddev, bi, NULL, + &conf->bio_split); + if (!bi) { + if (rw =3D=3D WRITE) + md_write_end(mddev); + return true; + } + if (bi !=3D front_bio) + split =3D true; + front_bio =3D bi; + logical_sector =3D bi->bi_iter.bi_sector & ~((sector_t)RAID5_STRIPE_SECTO= RS(conf)-1); bi->bi_next =3D NULL; =20 @@ -6348,6 +6362,11 @@ static bool raid5_make_request(struct mddev *mddev, = struct bio * bi) bio_endio(bi); =20 wait_for_completion(&done); + front_bio->bi_status =3D BLK_STS_OK; + if (split) { + submit_bio_noacct(front_bio); + return true; + } return false; } =20 --=20 2.51.0