From nobody Tue Sep 29 02:04:18 2026 Received: from mail-wm1-f44.google.com (mail-wm1-f44.google.com [209.85.128.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C5373415F2C for ; Thu, 13 Aug 2026 11:17:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786619843; cv=none; b=FbV2i1yaX8ePvLMEv29GM2/5tOAsoYBRkT6qWfk2VsILLGFrClD34j/EoZQwv1bH7a9Cw1cACnbD8GRiF/a/9UWU5QKrDhsoiaJCDyUyaAP+YXmw6OQXXW7EX69KIBjlt4v7xxcfGB3TGAeFQ6rdeSTsoRdfpM10gNpqIil0WZI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786619843; c=relaxed/simple; bh=pncr+3AAUskME8Gk3iiyUJk6taBi6rFyRMAmMpqFscQ=; h=Date:From:To:Cc:Subject:Message-ID:MIME-Version:Content-Type: Content-Disposition; b=XAfA/hcJCxrTw+LmA4a6RVZKygihsGFf2pA9hnzjcr+qxpWNMeYwtgxa8t5m0LbFUhY2aVTycDOJUxE1N0W/emlQBsHFoXENm+7r/UsaWlfGP/m/JPM9Op7WW7c5sVE55GFaRv6UfdQs+/p+EA/siZF6Vq+NjenA8svsjRGC1Bg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=JlBrFtuv; arc=none smtp.client-ip=209.85.128.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="JlBrFtuv" Received: by mail-wm1-f44.google.com with SMTP id 5b1f17b1804b1-49557167508so16398535e9.1 for ; Thu, 13 Aug 2026 04:17:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786619834; x=1787224634; darn=vger.kernel.org; h=content-disposition:content-type:mime-version:message-id:subject:cc :to:from:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=xJfcjlACAPt4pLrdOvLgFHFNryF7r2+1t+FJ5noYi5Y=; b=JlBrFtuvETreQMUsleoiKlKKOu0yPVaPsbKFYBSydLUWsWBZepV1O1gdgF/hBnxEQ9 DKch7VPXfKiMjuebH1d/Eiz9ousDC9wySy8ucCIBsX5Kp1DqKKi46Rat4blMCJmVfg/+ 7WpS7QJieEk78bu6FbdD8IHCOoDUHSCd+NMvEfpoGHE2pC/gpIwc5mZV2ZoKbzyiY2uE 5MAksGxEQQVwQ1XOwcq8n1xv2QIOdfa6Tl6raLS+0UOQqp461LVd0w/yz98fW2uoOVUP poB+1bd1hO8lGryy86eFkAhH8+QIdeKW84V2vFlPySvm0Ci6GSasT2VzTXOfd02hsQE8 SjtQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786619834; x=1787224634; h=content-disposition:content-type:mime-version:message-id:subject:cc :to:from:date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=xJfcjlACAPt4pLrdOvLgFHFNryF7r2+1t+FJ5noYi5Y=; b=jVuo8NeI1UYGUbAvNprm6nMAyfcex4d2jMjPdl+DefrXmPkA39/pu8ZK1etc7WiAef tbTz6PtOGT4ehKPsOf7dY8XLWxdDD9FGaKSWc1VxO4fSg7AhthesupPTD+6pnHZlGvCK 2qSW73dL5+BxI1G8rdF0R7ntTbvAbQ90gpdxe9z+u0k2TbUUzVU5VLmKvjaZCAt7ahFj xx7sXRdCZzOsLlJoGEMmLfYdcoog4nbgnunp4znnuQbhF13T7jeRN4wNdxH+kedLPTyh SjXoCO/+HOGOWJoJ6Z8x8UlHcaGHztFNmn6WENErWHEPwKA53+VYtJbDwnjP6PvqHujG RL5Q== X-Forwarded-Encrypted: i=1; AHgh+RqMfUR6irz1lozMzY/6/y6wViJUCA8pbkeE6ZMlhbv7BBmUDVGP/r7pNeVVSosblGR3FNGxvYGEDBcEGv8=@vger.kernel.org X-Gm-Message-State: AOJu0YyjaCyIBxfkvQeykxOO8hCo/XIAXSnNpbNM4qegwwxMS6T2S0Pg xfATxti3MF4evGFw2Ra2zYvnbz71PdLvtCBZL4higLpQCt7z9REI+QDV X-Gm-Gg: AR+sD11thWzuRH+y40E2SIn0KmSCgrgQA+8hpsjMTbhEeBOEk7e1TvfFRGCK1UAixBP +OG6cwiVgqhR/tzPNf9PJC9aKGwPqgtsq3EBdQuqjph0Gau0XdF59a5CQW0GKjCcv7RbzhfvhKF Y3YOH0Yu74K5U9ilgeoXkK5X/++krp0AUTmJtspz0W55cchm2wV292NpV2wGp/RPo82I2uq5vKr eaAOkraTLwTDMFb0TV++99L+BPZqMF0uJ4C8xQMCytrB9/NhwEjMcC/yR4u1mA0NGpFG03cS1k4 +8DqJWQ344Bebjhfme980gwpWQJNsGdw9epb1aDQd6zEuiLHpiXi2dD/IF/1Px5Ggn5nfuZOB9m nP/22vjxUfFi0Zrp+jsWtrGPYu0UJ8q4ylYsdCmipIZ0kDz1YBeE2V5/84V19uMBdFujKSlnBQb XpI6oeeK/EJYXK1BnQom6LNMCVJm0CBNYktPAGC+5Y5WULpHn3vUjXqb8kHcDUfezVnlgx99KeT Qqj/Lvg3KnCTLA= X-Received: by 2002:a05:600c:3511:b0:495:4e1d:82df with SMTP id 5b1f17b1804b1-499821f000cmr52259835e9.10.1786619833877; Thu, 13 Aug 2026 04:17:13 -0700 (PDT) Received: from localhost ([37.30.50.141]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49981dd6825sm31796215e9.1.2026.08.13.04.17.12 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 13 Aug 2026 04:17:13 -0700 (PDT) Date: Thu, 13 Aug 2026 13:17:12 +0200 From: Bartosz Chronowski To: linux-btrfs@vger.kernel.org Cc: Qu Wenruo , Chris Mason , David Sterba , "Yan, Zheng" , linux-kernel@vger.kernel.org, syzbot+021d10c4d4edc87daa03@syzkaller.appspotmail.com Subject: [PATCH RFC v2] btrfs: keep mixed block group writable for relocation setup commit Message-ID: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Disposition: inline Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Relocating a nearly full mixed block group can abort the filesystem transaction with -ENOSPC and trigger a warning in cleanup_transaction(). Making the mixed target read-only can lead to a condition where metadata COW cannot use its free space. In particular, btrfs_relocate_block_group() marks the mixed target read-only before prepare_to_relocate() commits the setup transaction. find_free_extent() then skips all free extents in the target. Commit-time COW still needs new tree blocks, so the transaction can fail with -ENOSPC when no suitable extent remains in another block group. Committing before the read-only transition does not fix the bug. Another workload can reserve space or start transaction N+1 between the commit and btrfs_inc_block_group_ro(). Keep a non-zoned, non-remap mixed target writable until its relocation setup transaction finishes. Fence data, tree-log and NOCOW admission while ordinary metadata COW remains allowed. Drain operations that crossed the fence before committing the setup transaction with reloc_ctl unpublished. Implement the boundary at the source files that own each state: - block-group.c owns the setup fence and final read-only transition, treats the fence as read-only for NOCOW and swap-extents admission, and makes other read-only holders wait for setup completion; - extent-tree.c rejects data and tree-log allocation into the fenced target, allows ordinary metadata COW, and keeps block group reservations only for data allocations until ordered extent registration; - inode.c treats the fenced target as read-only during NOCOW checks; - relocation.c drains each pass, binds setup to the running transaction and owns the read-only and reloc_ctl lifecycle; - transaction.c completes setup after switching commit roots and before transaction N+1 can start; - disk-io.c cancels a pending setup when its transaction is cleaned up. At the transaction tail, either mark the target read-only and publish reloc_ctl, or return the read-only transition error to relocation while the transaction completes and the target stays writable. Apply this boundary to every non-remap mixed relocation pass. Keep the existing paths unchanged for zoned, remap-tree and non-mixed block groups. Fixes: 3fd0a5585eb9 ("Btrfs: Metadata ENOSPC handling for balance") Reported-by: syzbot+021d10c4d4edc87daa03@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=3D021d10c4d4edc87daa03 Link: https://lore.kernel.org/r/9d9d207e-ad2a-4af6-9d0b-9a2bfc61d442@suse.c= om Assisted-by: Codex:gpt-5.6-sol syzkaller Signed-off-by: Bartosz Chronowski --- Changes in v2: - Drop the pre-commit-only approach because it leaves an admission window before the block group becomes read-only. - Keep the mixed target writable for setup metadata COW while fencing data, tree-log and NOCOW admission. - Bind setup to the exact transaction and publish the read-only state and reloc_ctl before transaction N+1 can start. - Apply the same boundary to every non-remap relocation pass and handle abort cleanup explicitly. Tested: - Focused and full x86_64 builds. - The syzbot C reproducer completed 16 independent runs without a crash. v1: https://lore.kernel.org/r/a06b5077-baa5-473f-9c65-bf72ac651b14@mail.ker= nel.org fs/btrfs/block-group.c | 128 ++++++++++++++++--- fs/btrfs/block-group.h | 8 +- fs/btrfs/disk-io.c | 1 + fs/btrfs/extent-tree.c | 19 ++- fs/btrfs/extent-tree.h | 1 + fs/btrfs/inode.c | 4 +- fs/btrfs/relocation.c | 280 ++++++++++++++++++++++++++++++++++++----- fs/btrfs/relocation.h | 4 + fs/btrfs/transaction.c | 4 + fs/btrfs/transaction.h | 3 + 10 files changed, 395 insertions(+), 57 deletions(-) diff --git a/fs/btrfs/block-group.c b/fs/btrfs/block-group.c index 8def7abb728f..332fc2721e01 100644 --- a/fs/btrfs/block-group.c +++ b/fs/btrfs/block-group.c @@ -21,6 +21,7 @@ #include "fs.h" #include "accessors.h" #include "extent-tree.h" +#include "relocation.h" =20 static struct kmem_cache *block_group_cache; static struct kmem_cache *free_space_ctl_cache; @@ -363,7 +364,8 @@ struct btrfs_block_group *btrfs_inc_nocow_writers(struc= t btrfs_fs_info *fs_info, return NULL; =20 spin_lock(&bg->lock); - if (bg->ro) + if (bg->ro || test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &bg->runtime_flags)) can_nocow =3D false; else atomic_inc(&bg->nocow_writers); @@ -419,7 +421,8 @@ void btrfs_wait_block_group_reservations(struct btrfs_b= lock_group *bg) { struct btrfs_space_info *space_info =3D bg->space_info; =20 - ASSERT(bg->ro); + ASSERT(bg->ro || test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &bg->runtime_flags)); =20 if (!(bg->flags & BTRFS_BLOCK_GROUP_DATA)) return; @@ -1434,7 +1437,8 @@ struct btrfs_trans_handle *btrfs_start_trans_remove_b= lock_group( * data in this block group. That check should be done by relocation routi= ne, * not this function. */ -static int inc_block_group_ro(struct btrfs_block_group *cache, bool force) +static int __inc_block_group_ro(struct btrfs_block_group *cache, bool forc= e, + bool reloc_setup) { struct btrfs_space_info *sinfo =3D cache->space_info; u64 num_bytes; @@ -1442,6 +1446,11 @@ static int inc_block_group_ro(struct btrfs_block_gro= up *cache, bool force) =20 spin_lock(&sinfo->lock); spin_lock(&cache->lock); + if (!reloc_setup && test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &cache->runtime_flags)) { + ret =3D -EAGAIN; + goto out; + } =20 if (cache->swap_extents) { ret =3D -ETXTBSY; @@ -1504,6 +1513,54 @@ static int inc_block_group_ro(struct btrfs_block_gro= up *cache, bool force) return ret; } =20 +static int inc_block_group_ro(struct btrfs_block_group *cache, bool force) +{ + return __inc_block_group_ro(cache, force, false); +} + +int btrfs_bg_reloc_setup_start(struct btrfs_block_group *cache, bool drop_= ro) +{ + struct btrfs_fs_info *fs_info =3D cache->fs_info; + struct btrfs_space_info *sinfo =3D cache->space_info; + int ret =3D 0; + + ASSERT(!btrfs_is_zoned(fs_info)); + + mutex_lock(&fs_info->ro_block_group_mutex); + spin_lock(&sinfo->lock); + spin_lock(&cache->lock); + if (test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, &cache->runtime_flags) || + cache->ro !=3D (drop_ro ? 1 : 0)) { + ret =3D -EAGAIN; + goto out; + } + + set_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, &cache->runtime_flags); + if (drop_ro) { + cache->ro =3D 0; + sinfo->bytes_readonly -=3D btrfs_block_group_available_space(cache); + list_del_init(&cache->ro_list); + } +out: + spin_unlock(&cache->lock); + spin_unlock(&sinfo->lock); + mutex_unlock(&fs_info->ro_block_group_mutex); + return ret; +} + +int btrfs_bg_reloc_setup_finish(struct btrfs_block_group *cache) +{ + ASSERT(test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, &cache->runtime_flags)); + return __inc_block_group_ro(cache, false, true); +} + +void btrfs_bg_reloc_setup_abort(struct btrfs_block_group *cache) +{ + ASSERT(test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, &cache->runtime_flags)); + clear_and_wake_up_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &cache->runtime_flags); +} + static bool clean_pinned_extents(struct btrfs_trans_handle *trans, const struct btrfs_block_group *bg) { @@ -1945,6 +2002,7 @@ static int btrfs_reclaim_block_group(struct btrfs_blo= ck_group *bg, int *reclaime u64 reserved; u64 old_total; int ret =3D 0; + bool marked_ro =3D false; =20 /* Don't race with allocators so take the groups_sem */ down_write(&space_info->groups_sem); @@ -2018,15 +2076,19 @@ static int btrfs_reclaim_block_group(struct btrfs_b= lock_group *bg, int *reclaime return 0; } =20 - ret =3D inc_block_group_ro(bg, false); + if (!btrfs_relocation_uses_fenced_setup(bg)) { + ret =3D inc_block_group_ro(bg, false); + if (!ret) + marked_ro =3D true; + } up_write(&space_info->groups_sem); if (ret < 0) return ret; =20 /* * The amount of bytes reclaimed corresponds to the sum of the - * "used" and "reserved" counters. We have set the block group - * to RO above, which prevents reservations from happening but + * "used" and "reserved" counters. Relocation prevents new data + * reservations before it drains existing reservations, but * we may have existing reservations for which allocation has * not yet been done - btrfs_update_block_group() was not yet * called, which is where we will transfer a reserved extent's @@ -2048,7 +2110,8 @@ static int btrfs_reclaim_block_group(struct btrfs_blo= ck_group *bg, int *reclaime trace_btrfs_reclaim_block_group(bg); ret =3D btrfs_relocate_chunk(fs_info, bg->start, false); if (ret) { - btrfs_dec_block_group_ro(bg); + if (marked_ro) + btrfs_dec_block_group_ro(bg); btrfs_err(fs_info, "error relocating chunk %llu", bg->start); used =3D 0; @@ -3131,7 +3194,7 @@ int btrfs_inc_block_group_ro(struct btrfs_block_group= *cache, struct btrfs_root *root =3D btrfs_block_group_root(fs_info); u64 alloc_flags; int ret; - bool dirty_bg_running; + bool retry; =20 if (unlikely(!root)) { btrfs_err(fs_info, "missing block group root"); @@ -3145,9 +3208,18 @@ int btrfs_inc_block_group_ro(struct btrfs_block_grou= p *cache, * Thus here we skip all chunk allocations. */ if (sb_rdonly(fs_info->sb)) { - mutex_lock(&fs_info->ro_block_group_mutex); - ret =3D inc_block_group_ro(cache, false); - mutex_unlock(&fs_info->ro_block_group_mutex); + do { + mutex_lock(&fs_info->ro_block_group_mutex); + retry =3D test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &cache->runtime_flags); + if (!retry) + ret =3D inc_block_group_ro(cache, false); + mutex_unlock(&fs_info->ro_block_group_mutex); + if (retry) + ret =3D wait_on_bit(&cache->runtime_flags, + BLOCK_GROUP_FLAG_RELOC_SETUP, + TASK_INTERRUPTIBLE); + } while (retry && !ret); return ret; } =20 @@ -3156,7 +3228,7 @@ int btrfs_inc_block_group_ro(struct btrfs_block_group= *cache, if (IS_ERR(trans)) return PTR_ERR(trans); =20 - dirty_bg_running =3D false; + retry =3D false; =20 /* * We're not allowed to set block groups readonly after the dirty @@ -3164,7 +3236,19 @@ int btrfs_inc_block_group_ro(struct btrfs_block_grou= p *cache, * back off and let this transaction commit. */ mutex_lock(&fs_info->ro_block_group_mutex); - if (test_bit(BTRFS_TRANS_DIRTY_BG_RUN, &trans->transaction->flags)) { + if (test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &cache->runtime_flags)) { + mutex_unlock(&fs_info->ro_block_group_mutex); + btrfs_end_transaction(trans); + + ret =3D wait_on_bit(&cache->runtime_flags, + BLOCK_GROUP_FLAG_RELOC_SETUP, + TASK_INTERRUPTIBLE); + if (ret) + return ret; + retry =3D true; + } else if (test_bit(BTRFS_TRANS_DIRTY_BG_RUN, + &trans->transaction->flags)) { u64 transid =3D trans->transid; =20 mutex_unlock(&fs_info->ro_block_group_mutex); @@ -3173,9 +3257,9 @@ int btrfs_inc_block_group_ro(struct btrfs_block_group= *cache, ret =3D btrfs_wait_for_commit(fs_info, transid); if (ret) return ret; - dirty_bg_running =3D true; + retry =3D true; } - } while (dirty_bg_running); + } while (retry); =20 if (do_chunk_alloc) { /* @@ -3411,7 +3495,9 @@ static void cache_save_setup(struct btrfs_block_group= *block_group, } retries++; =20 - if (block_group->ro) + if (block_group->ro || + test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &block_group->runtime_flags)) goto out_free; =20 ret =3D create_free_space_inode(trans, block_group, path); @@ -3981,6 +4067,7 @@ int btrfs_update_block_group(struct btrfs_trans_handl= e *trans, * @num_bytes except for the compress path. * @num_bytes: The number of bytes in question * @delalloc: The blocks are allocated for the delalloc write + * @allow_reloc_setup: Allow ordinary metadata into a relocation setup tar= get. * * This is called by the allocator when it reserves space. If this is a * reservation and the block group has become read only we cannot make the @@ -3988,7 +4075,8 @@ int btrfs_update_block_group(struct btrfs_trans_handl= e *trans, */ int btrfs_add_reserved_bytes(struct btrfs_block_group *cache, u64 ram_bytes, u64 num_bytes, bool delalloc, - bool force_wrong_size_class) + bool force_wrong_size_class, + bool allow_reloc_setup) { struct btrfs_space_info *space_info =3D cache->space_info; enum btrfs_block_group_size_class size_class; @@ -3996,7 +4084,9 @@ int btrfs_add_reserved_bytes(struct btrfs_block_group= *cache, =20 spin_lock(&space_info->lock); spin_lock(&cache->lock); - if (cache->ro) { + if (cache->ro || + (!allow_reloc_setup && + test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, &cache->runtime_flags))) { ret =3D -EAGAIN; goto out_error; } @@ -4832,7 +4922,7 @@ bool btrfs_inc_block_group_swap_extents(struct btrfs_= block_group *bg) bool ret =3D true; =20 spin_lock(&bg->lock); - if (bg->ro) + if (bg->ro || test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, &bg->runtime_flags)) ret =3D false; else bg->swap_extents++; diff --git a/fs/btrfs/block-group.h b/fs/btrfs/block-group.h index 790c2d467af5..d2b1dd01b45e 100644 --- a/fs/btrfs/block-group.h +++ b/fs/btrfs/block-group.h @@ -95,6 +95,8 @@ enum btrfs_block_group_flags { BLOCK_GROUP_FLAG_NEW, BLOCK_GROUP_FLAG_FULLY_REMAPPED, BLOCK_GROUP_FLAG_STRIPE_REMOVAL_PENDING, + /* Block data, tree-log and NOCOW admission during relocation setup. */ + BLOCK_GROUP_FLAG_RELOC_SETUP, }; =20 enum btrfs_caching_type { @@ -364,6 +366,9 @@ void btrfs_create_pending_block_groups(struct btrfs_tra= ns_handle *trans); int btrfs_inc_block_group_ro(struct btrfs_block_group *cache, bool do_chunk_alloc); void btrfs_dec_block_group_ro(struct btrfs_block_group *cache); +int btrfs_bg_reloc_setup_start(struct btrfs_block_group *cache, bool drop_= ro); +int btrfs_bg_reloc_setup_finish(struct btrfs_block_group *cache); +void btrfs_bg_reloc_setup_abort(struct btrfs_block_group *cache); int btrfs_start_dirty_block_groups(struct btrfs_trans_handle *trans); int btrfs_write_dirty_block_groups(struct btrfs_trans_handle *trans); int btrfs_setup_space_cache(struct btrfs_trans_handle *trans); @@ -371,7 +376,8 @@ int btrfs_update_block_group(struct btrfs_trans_handle = *trans, u64 bytenr, u64 num_bytes, bool alloc); int btrfs_add_reserved_bytes(struct btrfs_block_group *cache, u64 ram_bytes, u64 num_bytes, bool delalloc, - bool force_wrong_size_class); + bool force_wrong_size_class, + bool allow_reloc_setup); void btrfs_free_reserved_bytes(struct btrfs_block_group *cache, u64 num_by= tes, bool is_delalloc); int btrfs_chunk_alloc(struct btrfs_trans_handle *trans, diff --git a/fs/btrfs/disk-io.c b/fs/btrfs/disk-io.c index 2f1666d9544e..eab2fc5bf8b9 100644 --- a/fs/btrfs/disk-io.c +++ b/fs/btrfs/disk-io.c @@ -4936,6 +4936,7 @@ void btrfs_cleanup_one_transaction(struct btrfs_trans= action *cur_trans) } =20 btrfs_destroy_delayed_refs(cur_trans); + btrfs_abort_relocation_setup(cur_trans, cur_trans->aborted); =20 cur_trans->state =3D TRANS_STATE_COMMIT_START; wake_up(&fs_info->transaction_blocked_wait); diff --git a/fs/btrfs/extent-tree.c b/fs/btrfs/extent-tree.c index 624d76e0ca01..962af1840781 100644 --- a/fs/btrfs/extent-tree.c +++ b/fs/btrfs/extent-tree.c @@ -4639,6 +4639,9 @@ static noinline int find_free_extent(struct btrfs_roo= t *root, down_read(&space_info->groups_sem); if (list_empty(&block_group->list) || block_group->ro || + (test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &block_group->runtime_flags) && + (ffe_ctl->is_data || ffe_ctl->for_treelog)) || (block_group->flags & BTRFS_BLOCK_GROUP_REMAPPED)) { /* * someone is removing this block group, @@ -4674,7 +4677,10 @@ static noinline int find_free_extent(struct btrfs_ro= ot *root, ffe_ctl->hinted =3D false; /* If the block group is read-only, we can skip it entirely. */ if (unlikely(block_group->ro || - (block_group->flags & BTRFS_BLOCK_GROUP_REMAPPED))) { + (test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &block_group->runtime_flags) && + (ffe_ctl->is_data || ffe_ctl->for_treelog)) || + (block_group->flags & BTRFS_BLOCK_GROUP_REMAPPED))) { if (ffe_ctl->for_treelog) btrfs_clear_treelog_bg(block_group); if (ffe_ctl->for_data_reloc) @@ -4776,14 +4782,16 @@ static noinline int find_free_extent(struct btrfs_r= oot *root, ret =3D btrfs_add_reserved_bytes(block_group, ffe_ctl->ram_bytes, ffe_ctl->num_bytes, ffe_ctl->delalloc, - ffe_ctl->loop >=3D LOOP_WRONG_SIZE_CLASS); + ffe_ctl->loop >=3D LOOP_WRONG_SIZE_CLASS, + !ffe_ctl->is_data && !ffe_ctl->for_treelog); if (ret =3D=3D -EAGAIN) { btrfs_add_free_space_unused(block_group, ffe_ctl->found_offset, ffe_ctl->num_bytes); goto loop; } - btrfs_inc_block_group_reservations(block_group); + if (ffe_ctl->is_data) + btrfs_inc_block_group_reservations(block_group); =20 /* we are all good, lets return */ ins->objectid =3D ffe_ctl->search_start; @@ -4897,14 +4905,13 @@ int btrfs_reserve_extent(struct btrfs_root *root, u= 64 ram_bytes, ffe_ctl.empty_size =3D empty_size; ffe_ctl.flags =3D flags; ffe_ctl.delalloc =3D delalloc; + ffe_ctl.is_data =3D is_data; ffe_ctl.hint_byte =3D hint_byte; ffe_ctl.for_treelog =3D for_treelog; ffe_ctl.for_data_reloc =3D for_data_reloc; =20 ret =3D find_free_extent(root, ins, &ffe_ctl); - if (!ret && !is_data) { - btrfs_dec_block_group_reservations(fs_info, ins->objectid); - } else if (ret =3D=3D -ENOSPC) { + if (ret =3D=3D -ENOSPC) { if (!final_tried && ins->offset) { num_bytes =3D min(num_bytes >> 1, ins->offset); num_bytes =3D round_down(num_bytes, diff --git a/fs/btrfs/extent-tree.h b/fs/btrfs/extent-tree.h index ff330d4896d6..74ba10a46951 100644 --- a/fs/btrfs/extent-tree.h +++ b/fs/btrfs/extent-tree.h @@ -40,6 +40,7 @@ struct find_free_extent_ctl { bool use_cluster; =20 bool delalloc; + bool is_data; bool have_caching_bg; bool orig_have_caching_bg; =20 diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c index 2534cd9284d5..28c5902636d3 100644 --- a/fs/btrfs/inode.c +++ b/fs/btrfs/inode.c @@ -7398,7 +7398,9 @@ static bool btrfs_extent_readonly(struct btrfs_fs_inf= o *fs_info, u64 bytenr) bool readonly =3D false; =20 block_group =3D btrfs_lookup_block_group(fs_info, bytenr); - if (!block_group || block_group->ro) + if (!block_group || block_group->ro || + test_bit(BLOCK_GROUP_FLAG_RELOC_SETUP, + &block_group->runtime_flags)) readonly =3D true; if (block_group) btrfs_put_block_group(block_group); diff --git a/fs/btrfs/relocation.c b/fs/btrfs/relocation.c index fc5c14b5adad..92059ebc8345 100644 --- a/fs/btrfs/relocation.c +++ b/fs/btrfs/relocation.c @@ -173,11 +173,16 @@ struct reloc_control { =20 u64 search_start; u64 extents_found; + int setup_result; =20 enum reloc_stage stage; bool create_reloc_tree; bool merge_reloc_tree; bool found_file_extent; + bool fenced_setup; + bool setup_pending; + bool block_group_ro; + bool reloc_ctl_set; =20 refcount_t refs; }; @@ -3507,14 +3512,24 @@ int find_next_extent(struct reloc_control *rc, stru= ct btrfs_path *path, return ret; } =20 -static void set_reloc_control(struct reloc_control *rc) +static void __set_reloc_control(struct reloc_control *rc) { struct btrfs_fs_info *fs_info =3D rc->extent_root->fs_info; =20 - mutex_lock(&fs_info->reloc_mutex); + lockdep_assert_held(&fs_info->reloc_mutex); spin_lock(&fs_info->reloc_ctl_lock); + ASSERT(!fs_info->reloc_ctl || fs_info->reloc_ctl =3D=3D rc); fs_info->reloc_ctl =3D rc; + rc->reloc_ctl_set =3D true; spin_unlock(&fs_info->reloc_ctl_lock); +} + +static void set_reloc_control(struct reloc_control *rc) +{ + struct btrfs_fs_info *fs_info =3D rc->extent_root->fs_info; + + mutex_lock(&fs_info->reloc_mutex); + __set_reloc_control(rc); mutex_unlock(&fs_info->reloc_mutex); } =20 @@ -3524,18 +3539,137 @@ static void unset_reloc_control(struct reloc_contr= ol *rc) =20 mutex_lock(&fs_info->reloc_mutex); spin_lock(&fs_info->reloc_ctl_lock); - fs_info->reloc_ctl =3D NULL; + if (rc->reloc_ctl_set) { + ASSERT(fs_info->reloc_ctl =3D=3D rc); + fs_info->reloc_ctl =3D NULL; + rc->reloc_ctl_set =3D false; + } else { + ASSERT(fs_info->reloc_ctl !=3D rc); + } spin_unlock(&fs_info->reloc_ctl_lock); mutex_unlock(&fs_info->reloc_mutex); } =20 +static void complete_relocation_setup(struct reloc_control *rc, int result) +{ + ASSERT(rc->setup_pending); + WRITE_ONCE(rc->setup_result, result); + WRITE_ONCE(rc->setup_pending, false); + btrfs_bg_reloc_setup_abort(rc->block_group); + put_reloc_control(rc); +} + +void btrfs_finish_relocation_setup(struct btrfs_transaction *trans) +{ + struct btrfs_fs_info *fs_info =3D trans->fs_info; + struct reloc_control *rc; + int ret; + + lockdep_assert_held(&fs_info->reloc_mutex); + + spin_lock(&fs_info->trans_lock); + rc =3D trans->reloc_setup; + trans->reloc_setup =3D NULL; + spin_unlock(&fs_info->trans_lock); + if (!rc) + return; + + ret =3D btrfs_bg_reloc_setup_finish(rc->block_group); + if (!ret) { + WRITE_ONCE(rc->block_group_ro, true); + __set_reloc_control(rc); + } + complete_relocation_setup(rc, ret); +} + +void btrfs_abort_relocation_setup(struct btrfs_transaction *trans, int err= or) +{ + struct btrfs_fs_info *fs_info =3D trans->fs_info; + struct reloc_control *rc; + + spin_lock(&fs_info->trans_lock); + rc =3D trans->reloc_setup; + trans->reloc_setup =3D NULL; + spin_unlock(&fs_info->trans_lock); + if (!rc) + return; + + complete_relocation_setup(rc, error ?: -EIO); +} + +static int bind_relocation_setup(struct btrfs_trans_handle *trans, + struct reloc_control *rc, + struct btrfs_transaction **transaction) +{ + struct btrfs_fs_info *fs_info =3D trans->fs_info; + struct btrfs_transaction *cur_trans =3D trans->transaction; + int ret =3D 0; + + mutex_lock(&fs_info->ro_block_group_mutex); + spin_lock(&fs_info->trans_lock); + if (TRANS_ABORTED(cur_trans)) { + ret =3D cur_trans->aborted; + } else if (cur_trans !=3D fs_info->running_transaction || + cur_trans->state !=3D TRANS_STATE_RUNNING || + test_bit(BTRFS_TRANS_DIRTY_BG_RUN, &cur_trans->flags)) { + ret =3D -EAGAIN; + } else if (cur_trans->reloc_setup) { + ret =3D -EBUSY; + } else { + ASSERT(rc->setup_pending); + WRITE_ONCE(rc->setup_result, -EINPROGRESS); + refcount_inc(&rc->refs); + cur_trans->reloc_setup =3D rc; + refcount_inc(&cur_trans->use_count); + *transaction =3D cur_trans; + } + spin_unlock(&fs_info->trans_lock); + mutex_unlock(&fs_info->ro_block_group_mutex); + + return ret; +} + +static int reconcile_relocation_setup(struct btrfs_transaction *trans, + struct reloc_control *rc, + int commit_ret) +{ + struct btrfs_fs_info *fs_info =3D trans->fs_info; + bool cancel =3D false; + bool wait =3D false; + int setup_ret; + + spin_lock(&fs_info->trans_lock); + if (trans->reloc_setup =3D=3D rc && + trans->state < TRANS_STATE_COMMIT_PREP) { + trans->reloc_setup =3D NULL; + cancel =3D true; + } else if (READ_ONCE(rc->setup_result) =3D=3D -EINPROGRESS) { + wait =3D true; + } + spin_unlock(&fs_info->trans_lock); + + if (cancel) + complete_relocation_setup(rc, commit_ret ?: -EIO); + else if (wait) + wait_event(trans->commit_wait, + READ_ONCE(trans->state) >=3D TRANS_STATE_COMPLETED); + + setup_ret =3D READ_ONCE(rc->setup_result); + ASSERT(setup_ret !=3D -EINPROGRESS); + btrfs_put_transaction(trans); + + return commit_ret ?: setup_ret; +} + static noinline_for_stack int prepare_to_relocate(struct reloc_control *rc) { + struct btrfs_fs_info *fs_info =3D rc->extent_root->fs_info; struct btrfs_trans_handle *trans; + struct btrfs_transaction *transaction =3D NULL; int ret; =20 - rc->block_rsv =3D btrfs_alloc_block_rsv(rc->extent_root->fs_info, + rc->block_rsv =3D btrfs_alloc_block_rsv(fs_info, BTRFS_BLOCK_RSV_TEMP); if (!rc->block_rsv) return -ENOMEM; @@ -3546,32 +3680,93 @@ int prepare_to_relocate(struct reloc_control *rc) rc->nodes_relocated =3D 0; rc->merging_rsv_size =3D 0; rc->reserved_bytes =3D 0; - rc->block_rsv->size =3D rc->extent_root->fs_info->nodesize * - RELOCATION_RESERVED_NODES; - ret =3D btrfs_block_rsv_refill(rc->extent_root->fs_info, + rc->block_rsv->size =3D fs_info->nodesize * RELOCATION_RESERVED_NODES; + + if (!rc->fenced_setup) { + ret =3D btrfs_block_rsv_refill(fs_info, + rc->block_rsv, rc->block_rsv->size, + BTRFS_RESERVE_FLUSH_ALL); + if (ret) + return ret; + + rc->create_reloc_tree =3D true; + set_reloc_control(rc); + + trans =3D btrfs_join_transaction(rc->extent_root); + if (IS_ERR(trans)) { + unset_reloc_control(rc); + /* + * The extent tree is not a ref-cow tree and has no reloc + * root to clean up. Callers free the block reserve. + */ + return PTR_ERR(trans); + } + + ret =3D btrfs_commit_transaction(trans); + if (ret) + unset_reloc_control(rc); + return ret; + } + + if (!rc->setup_pending) { + ret =3D btrfs_bg_reloc_setup_start(rc->block_group, + rc->block_group_ro); + if (ret) + return ret; + WRITE_ONCE(rc->setup_pending, true); + WRITE_ONCE(rc->block_group_ro, false); + } else { + ASSERT(!rc->block_group_ro); + } + + btrfs_wait_block_group_reservations(rc->block_group); + btrfs_wait_nocow_writers(rc->block_group); + btrfs_wait_ordered_roots(fs_info, U64_MAX, rc->block_group); + + ret =3D btrfs_block_rsv_refill(fs_info, rc->block_rsv, rc->block_rsv->size, BTRFS_RESERVE_FLUSH_ALL); if (ret) - return ret; + goto abort_setup; =20 + /* The transaction tail publishes reloc_ctl with the new commit roots. */ rc->create_reloc_tree =3D true; - set_reloc_control(rc); + for (;;) { + u64 transid; =20 - trans =3D btrfs_join_transaction(rc->extent_root); - if (IS_ERR(trans)) { - unset_reloc_control(rc); - /* - * extent tree is not a ref_cow tree and has no reloc_root to - * cleanup. And callers are responsible to free the above - * block rsv. - */ - return PTR_ERR(trans); + trans =3D btrfs_join_transaction(rc->extent_root); + if (IS_ERR(trans)) { + ret =3D PTR_ERR(trans); + goto abort_setup; + } + transid =3D trans->transid; + + ret =3D bind_relocation_setup(trans, rc, &transaction); + if (ret =3D=3D -EAGAIN) { + btrfs_end_transaction(trans); + ret =3D btrfs_wait_for_commit(fs_info, transid); + if (ret) + goto abort_setup; + continue; + } + if (ret) { + btrfs_end_transaction(trans); + goto abort_setup; + } + break; } =20 ret =3D btrfs_commit_transaction(trans); - if (ret) + ret =3D reconcile_relocation_setup(transaction, rc, ret); + if (ret && rc->reloc_ctl_set) unset_reloc_control(rc); + return ret; =20 +abort_setup: + ASSERT(rc->setup_pending); + WRITE_ONCE(rc->setup_result, ret); + WRITE_ONCE(rc->setup_pending, false); + btrfs_bg_reloc_setup_abort(rc->block_group); return ret; } =20 @@ -3937,6 +4132,14 @@ static const char *stage_to_string(enum reloc_stage = stage) return "unknown"; } =20 +bool btrfs_relocation_uses_fenced_setup(const struct btrfs_block_group *bg) +{ + const u64 mixed =3D BTRFS_BLOCK_GROUP_DATA | BTRFS_BLOCK_GROUP_METADATA; + + return (bg->flags & mixed) =3D=3D mixed && !btrfs_is_zoned(bg->fs_info) && + !should_relocate_using_remap_tree(bg); +} + static int add_remap_tree_entries(struct btrfs_trans_handle *trans, struct= btrfs_path *path, struct btrfs_key *entries, unsigned int num_entries) { @@ -5404,7 +5607,6 @@ int btrfs_relocate_block_group(struct btrfs_fs_info *= fs_info, u64 group_start, struct inode *inode; struct btrfs_path *path =3D NULL; int ret; - bool bg_is_ro =3D false; =20 if (unlikely(!extent_root)) { btrfs_err(fs_info, @@ -5455,15 +5657,24 @@ int btrfs_relocate_block_group(struct btrfs_fs_info= *fs_info, u64 group_start, rc->extent_root =3D extent_root; /* Block group ref now owned by rc, put_reloc_control() will drop it. */ rc->block_group =3D bg; + rc->fenced_setup =3D btrfs_relocation_uses_fenced_setup(bg); =20 ret =3D reloc_chunk_start(fs_info); if (ret < 0) goto out_put_rc; =20 - ret =3D btrfs_inc_block_group_ro(rc->block_group, true); - if (ret) - goto out; - bg_is_ro =3D true; + if (rc->fenced_setup) { + /* Keep non-metadata writers out until the setup tail marks RO. */ + ret =3D btrfs_bg_reloc_setup_start(rc->block_group, false); + if (ret) + goto out; + WRITE_ONCE(rc->setup_pending, true); + } else { + ret =3D btrfs_inc_block_group_ro(rc->block_group, true); + if (ret) + goto out; + rc->block_group_ro =3D true; + } =20 path =3D btrfs_alloc_path(); if (!path) { @@ -5494,12 +5705,14 @@ int btrfs_relocate_block_group(struct btrfs_fs_info= *fs_info, u64 group_start, if (verbose) describe_relocation(rc->block_group); =20 - btrfs_wait_block_group_reservations(rc->block_group); - btrfs_wait_nocow_writers(rc->block_group); - btrfs_wait_ordered_roots(fs_info, U64_MAX, rc->block_group); + if (!rc->fenced_setup) { + btrfs_wait_block_group_reservations(rc->block_group); + btrfs_wait_nocow_writers(rc->block_group); + btrfs_wait_ordered_roots(fs_info, U64_MAX, rc->block_group); =20 - ret =3D btrfs_zone_finish(rc->block_group); - WARN_ON(ret && ret !=3D -EAGAIN); + ret =3D btrfs_zone_finish(rc->block_group); + WARN_ON(ret && ret !=3D -EAGAIN); + } =20 if (should_relocate_using_remap_tree(bg)) { if (bg->remap_bytes !=3D 0) { @@ -5521,8 +5734,15 @@ int btrfs_relocate_block_group(struct btrfs_fs_info = *fs_info, u64 group_start, } =20 out: - if (ret && bg_is_ro) + if (rc->setup_pending) { + ASSERT(ret); + WRITE_ONCE(rc->setup_pending, false); + btrfs_bg_reloc_setup_abort(rc->block_group); + } + if (ret && rc->block_group_ro) { btrfs_dec_block_group_ro(rc->block_group); + rc->block_group_ro =3D false; + } if (!btrfs_fs_incompat(fs_info, REMAP_TREE)) iput(rc->data_inode); btrfs_free_path(path); diff --git a/fs/btrfs/relocation.h b/fs/btrfs/relocation.h index bb7a86e7dbe3..210d0bbd7d48 100644 --- a/fs/btrfs/relocation.h +++ b/fs/btrfs/relocation.h @@ -11,6 +11,7 @@ struct btrfs_root; struct btrfs_trans_handle; struct btrfs_ordered_extent; struct btrfs_pending_snapshot; +struct btrfs_transaction; =20 static inline bool should_relocate_using_remap_tree(const struct btrfs_blo= ck_group *bg) { @@ -25,6 +26,9 @@ static inline bool should_relocate_using_remap_tree(const= struct btrfs_block_gro =20 int btrfs_relocate_block_group(struct btrfs_fs_info *fs_info, u64 group_st= art, bool verbose); +bool btrfs_relocation_uses_fenced_setup(const struct btrfs_block_group *bg= ); +void btrfs_finish_relocation_setup(struct btrfs_transaction *trans); +void btrfs_abort_relocation_setup(struct btrfs_transaction *trans, int err= or); int btrfs_init_reloc_root(struct btrfs_trans_handle *trans, struct btrfs_r= oot *root); int btrfs_update_reloc_root(struct btrfs_trans_handle *trans, struct btrfs_root *root); diff --git a/fs/btrfs/transaction.c b/fs/btrfs/transaction.c index 8f9419728100..97556bdfdead 100644 --- a/fs/btrfs/transaction.c +++ b/fs/btrfs/transaction.c @@ -173,6 +173,7 @@ void btrfs_put_transaction(struct btrfs_transaction *tr= ansaction) btrfs_put_block_group(cache); } WARN_ON(!list_empty(&transaction->dev_update_list)); + WARN_ON(transaction->reloc_setup); kfree(transaction); } } @@ -379,6 +380,7 @@ static noinline int join_transaction(struct btrfs_fs_in= fo *fs_info, INIT_LIST_HEAD(&cur_trans->dev_update_list); INIT_LIST_HEAD(&cur_trans->switch_commits); INIT_LIST_HEAD(&cur_trans->dirty_bgs); + cur_trans->reloc_setup =3D NULL; INIT_LIST_HEAD(&cur_trans->io_bgs); INIT_LIST_HEAD(&cur_trans->dropped_roots); mutex_init(&cur_trans->cache_write_mutex); @@ -2552,6 +2554,8 @@ int btrfs_commit_transaction(struct btrfs_trans_handl= e *trans) clear_bit(BTRFS_FS_LOG2_ERR, &fs_info->flags); =20 btrfs_trans_release_chunk_metadata(trans); + /* Resolve the relocation setup before transaction N+1 can start. */ + btrfs_finish_relocation_setup(cur_trans); =20 /* * Before changing the transaction state to TRANS_STATE_UNBLOCKED and diff --git a/fs/btrfs/transaction.h b/fs/btrfs/transaction.h index 5e4b1106fd90..bbf3c2b78ce1 100644 --- a/fs/btrfs/transaction.h +++ b/fs/btrfs/transaction.h @@ -23,6 +23,7 @@ struct btrfs_fs_info; struct btrfs_root_item; struct btrfs_root; struct btrfs_path; +struct reloc_control; =20 /* * Signal that a direct IO write is in progress, to avoid deadlock for sync @@ -77,6 +78,8 @@ struct btrfs_transaction { struct list_head dev_update_list; struct list_head switch_commits; struct list_head dirty_bgs; + /* Protected by fs_info->trans_lock. */ + struct reloc_control *reloc_setup; =20 /* * There is no explicit lock which protects io_bgs, rather its base-commit: f5bbbfec59b4e2fb7520a91de3df8a6174325d6a --=20 2.43.0