[PATCH v2 0/2] btrfs: delay mounted sysfs attributes until mount is ready

Jiacheng Xu posted 2 patches 1 month, 1 week ago
fs/btrfs/disk-io.c | 18 ++++++++++++++++-
fs/btrfs/sysfs.c   | 50 ++++++++++++++++++++++++++++++----------------
fs/btrfs/sysfs.h   |  2 ++
3 files changed, 52 insertions(+), 18 deletions(-)
[PATCH v2 0/2] btrfs: delay mounted sysfs attributes until mount is ready
Posted by Jiacheng Xu 1 month, 1 week ago
Here is a potential fix following Wenruo's idea.

btrfs_sysfs_add_mounted() currently publishes the writable label and
feature attributes before the transaction kthread is created. A concurrent
sysfs write can therefore dereference a NULL transaction_kthread in
wake_up_process().

This series follows the suggested lifecycle: create only the required
subdirectories during early mount, publish the fsid attributes after mount
initialization, and remove them before the kthreads are stopped. The
feature attributes are included because their store callback has the same
transaction_kthread dependency as the label callback.

Patch 1 factors the fsid attribute handling into dedicated helpers. Patch 2
moves their publication and removal to the safe mount and unmount stages.
On unmount the cleaner is parked before attribute removal so it cannot
recreate the feature group through sysfs_update_group(). Both patches are
required for stable backports.

The resulting fs/btrfs/sysfs.o and fs/btrfs/disk-io.o were build-tested.

Changes in v2:
- Delay creation of both the root and feature attributes until mount setup
  is complete.
- Remove those attributes while their kthread dependencies are still
  valid.
- Split helper extraction from the lifecycle fix for stable backports.

Jiacheng Xu (2):
  btrfs: sysfs: factor out mounted fsid attribute helpers
  btrfs: delay mounted fsid attributes until the fs is ready

 fs/btrfs/disk-io.c | 18 ++++++++++++++++-
 fs/btrfs/sysfs.c   | 50 ++++++++++++++++++++++++++++++----------------
 fs/btrfs/sysfs.h   |  2 ++
 3 files changed, 52 insertions(+), 18 deletions(-)

base-commit: 0f23d56f17fdfc7db69d51f64c8b91bbab947aa9
-- 
2.25.1


> -----原始邮件-----
> 发件人: "Jiacheng Xu" <stitch@zju.edu.cn>
> 发送时间:2026-08-20 20:27:23 (星期四)
> 收件人: "Chris Mason" <clm@fb.com>
> 抄送: "David Sterba" <dsterba@suse.com>, linux-btrfs@vger.kernel.org, linux-kernel@vger.kernel.org
> 主题: [PATCH] btrfs: drain sysfs callbacks before stopping transaction kthread
> 
> btrfs_label_store() and btrfs_feature_attr_store() wake up the
> transaction kthread through fs_info->transaction_kthread.
> 
> During filesystem teardown, close_ctree() stops the transaction kthread
> before removing the mounted filesystem's sysfs attributes. A concurrent
> sysfs write can therefore enter one of these callbacks after the kthread
> has been stopped and pass an invalid task pointer to wake_up_process().
> 
> This results in a concurrent null-pointer dereference in
> try_to_wake_up(). The scheduler is not the root cause; the invalid
> transaction kthread pointer is used by a Btrfs sysfs callback during
> teardown.
> 
> Split mounted sysfs cleanup into two stages. Remove attributes which may
> have store callbacks before stopping the transaction kthread. The
> remaining sysfs kobjects are removed at the original teardown point,
> after the kthread has been stopped.
> 
> Apply the same ordering to the open_ctree() failure path when the
> transaction kthread has already been created.
> 
> Tested-by: Jiacheng Xu <stitch@zju.edu.cn>
> Signed-off-by: Jiacheng Xu <stitch@zju.edu.cn>
> ---
> fs/btrfs/disk-io.c | 16 ++++++++++++++--
> fs/btrfs/sysfs.c   | 26 +++++++++++++++++++++-----
> fs/btrfs/sysfs.h   |  3 +++
> 3 files changed, 38 insertions(+), 7 deletions(-)
> 
> diff --git a/fs/btrfs/disk-io.c b/fs/btrfs/disk-io.c
> index 2f1666d9544e..4f5bcc576dc6 100644
> --- a/fs/btrfs/disk-io.c
> +++ b/fs/btrfs/disk-io.c
> @@ -3363,6 +3363,7 @@ int __cold open_ctree(struct super_block *sb, struct btrfs_fs_devices *fs_device
>       struct btrfs_root *tree_root;
>       struct btrfs_root *chunk_root;
>       struct btrfs_root *remap_root;
> +     bool sysfs_attrs_removed = false;
>       int ret;
>       int level;
> 
> @@ -3780,6 +3781,9 @@ int __cold open_ctree(struct super_block *sb, struct btrfs_fs_devices *fs_device
> fail_qgroup:
>       btrfs_free_qgroup_config(fs_info);
> fail_trans_kthread:
> +     btrfs_sysfs_remove_mounted_attrs(fs_info);
> +     sysfs_attrs_removed = true;
> +
>       kthread_stop(fs_info->transaction_kthread);
>       btrfs_cleanup_transaction(fs_info);
>       btrfs_free_fs_roots(fs_info);
> @@ -3793,7 +3797,9 @@ int __cold open_ctree(struct super_block *sb, struct btrfs_fs_devices *fs_device
>       filemap_write_and_wait(fs_info->btree_inode->i_mapping);
> 
> fail_sysfs:
> -     btrfs_sysfs_remove_mounted(fs_info);
> +     if (!sysfs_attrs_removed)
> +             btrfs_sysfs_remove_mounted_attrs(fs_info);
> +     btrfs_sysfs_remove_mounted_kobjects(fs_info);
> 
> fail_fsdev_sysfs:
>       btrfs_sysfs_remove_fsid(fs_info->fs_devices);
> @@ -4318,6 +4324,9 @@ void __cold close_ctree(struct btrfs_fs_info *fs_info)
> 
>       set_bit(BTRFS_FS_CLOSING_START, &fs_info->flags);
> 
> +     /* Drain sysfs callbacks before stopping the transaction kthread. */
> +     btrfs_sysfs_remove_mounted_attrs(fs_info);
> +
>       /*
>       * If we had UNFINISHED_DROPS we could still be processing them, so
>       * clear that bit and wake up relocation so it can stop.
> @@ -4538,7 +4547,7 @@ void __cold close_ctree(struct btrfs_fs_info *fs_info)
>                         percpu_counter_sum(&fs_info->ordered_bytes));
> 
> -     btrfs_sysfs_remove_mounted(fs_info);
> +     btrfs_sysfs_remove_mounted_kobjects(fs_info);
>       btrfs_sysfs_remove_fsid(fs_info->fs_devices);
> 
>       btrfs_put_block_group_cache(fs_info);
> 
> diff --git a/fs/btrfs/sysfs.c b/fs/btrfs/sysfs.c
> index 0d14570c8bc2..d90d76a152e9 100644
> --- a/fs/btrfs/sysfs.c
> +++ b/fs/btrfs/sysfs.c
> @@ -1707,11 +1707,23 @@ static void btrfs_sysfs_remove_fs_devices(struct btrfs_fs_devices *fs_devices)
>       }
> }
> 
> -void btrfs_sysfs_remove_mounted(struct btrfs_fs_info *fs_info)
> +/*
> + * Remove attributes which may have store callbacks. kernfs waits for active
> + * callbacks during removal, so this must be done before stopping any kthread
> + * which can be woken up by those callbacks.
> + */
> +void btrfs_sysfs_remove_mounted_attrs(struct btrfs_fs_info *fs_info)
> {
>       struct kobject *fsid_kobj = &fs_info->fs_devices->fsid_kobj;
> 
> -     sysfs_remove_link(fsid_kobj, "bdi");
> +     addrm_unknown_feature_attrs(fs_info, false);
> +     sysfs_remove_group(fsid_kobj, &btrfs_feature_attr_group);
> +     sysfs_remove_files(fsid_kobj, btrfs_attrs);
> +}
> +
> +static void btrfs_sysfs_remove_mounted_dirs(struct btrfs_fs_info *fs_info)
> +{
> +     sysfs_remove_link(&fs_info->fs_devices->fsid_kobj, "bdi");
> 
>       if (fs_info->space_info_kobj) {
>             sysfs_remove_files(fs_info->space_info_kobj, allocation_attrs);
> @@ -1730,9 +1742,18 @@ void btrfs_sysfs_remove_mounted(struct btrfs_fs_info *fs_info)
>             kobject_put(fs_info->debug_kobj);
>       }
> #endif
> -     addrm_unknown_feature_attrs(fs_info, false);
> -     sysfs_remove_group(fsid_kobj, &btrfs_feature_attr_group);
> -     sysfs_remove_files(fsid_kobj, btrfs_attrs);
> +}
> +
> +void btrfs_sysfs_remove_mounted_kobjects(struct btrfs_fs_info *fs_info)
> +{
> +     btrfs_sysfs_remove_mounted_dirs(fs_info);
> +     btrfs_sysfs_remove_fs_devices(fs_info->fs_devices);
> +}
> +
> +void btrfs_sysfs_remove_mounted(struct btrfs_fs_info *fs_info)
> +{
> +     btrfs_sysfs_remove_mounted_dirs(fs_info);
> +     btrfs_sysfs_remove_mounted_attrs(fs_info);
>       btrfs_sysfs_remove_fs_devices(fs_info->fs_devices);
> }
> 
> diff --git a/fs/btrfs/sysfs.h b/fs/btrfs/sysfs.h
> index 05498e5346c3..0d008fc8f1b8 100644
> --- a/fs/btrfs/sysfs.h
> +++ b/fs/btrfs/sysfs.h
> @@ -35,6 +35,9 @@ void btrfs_kobject_uevent(struct block_device *bdev, enum kobject_action action)
> int __init btrfs_init_sysfs(void);
> void __cold btrfs_exit_sysfs(void);
> int btrfs_sysfs_add_mounted(struct btrfs_fs_info *fs_info);
> +void btrfs_sysfs_remove_mounted_attrs(struct btrfs_fs_info *fs_info);
> +void btrfs_sysfs_remove_mounted_kobjects(struct btrfs_fs_info *fs_info);
> void btrfs_sysfs_remove_mounted(struct btrfs_fs_info *fs_info);
> void btrfs_sysfs_add_block_group_type(struct btrfs_block_group *cache);
> int btrfs_sysfs_add_space_info_type(struct btrfs_space_info *space_info);
Re: [PATCH v2 0/2] btrfs: delay mounted sysfs attributes until mount is ready
Posted by Qu Wenruo 1 month, 1 week ago

在 2026/8/22 13:11, Jiacheng Xu 写道:
> Here is a potential fix following Wenruo's idea.
> 
> btrfs_sysfs_add_mounted() currently publishes the writable label and
> feature attributes before the transaction kthread is created. A concurrent
> sysfs write can therefore dereference a NULL transaction_kthread in
> wake_up_process().
> 
> This series follows the suggested lifecycle: create only the required
> subdirectories during early mount, publish the fsid attributes after mount
> initialization, and remove them before the kthreads are stopped. The
> feature attributes are included because their store callback has the same
> transaction_kthread dependency as the label callback.
> 
> Patch 1 factors the fsid attribute handling into dedicated helpers. Patch 2
> moves their publication and removal to the safe mount and unmount stages.
> On unmount the cleaner is parked before attribute removal so it cannot
> recreate the feature group through sysfs_update_group(). Both patches are
> required for stable backports.
> 
> The resulting fs/btrfs/sysfs.o and fs/btrfs/disk-io.o were build-tested.
> 
> Changes in v2:
> - Delay creation of both the root and feature attributes until mount setup
>    is complete.

You don't need to bother feature attributes for now, there is already a 
patch addressing it by completely removing the write support for feature 
attributes:

https://lore.kernel.org/linux-btrfs/8a598d76555b5944d34bb08fa8dbeea28fc05db9.1787307129.git.wqu@suse.com/

Considering it's only extended_iref, removing it should be much simpler.
Until that is determined, you only need to bother the label one.


Furthermore, among all the attr files in the fsid directory, there is 
only label that is writable, it would make more sense to split 
btrfs_attrs into two parts, one for those read-only members, and one for 
the only writebale label one.

Otherwise the series looks much better.


> - Remove those attributes while their kthread dependencies are still
>    valid.
> - Split helper extraction from the lifecycle fix for stable backports.
> 
> Jiacheng Xu (2):
>    btrfs: sysfs: factor out mounted fsid attribute helpers
>    btrfs: delay mounted fsid attributes until the fs is ready
> 
>   fs/btrfs/disk-io.c | 18 ++++++++++++++++-
>   fs/btrfs/sysfs.c   | 50 ++++++++++++++++++++++++++++++----------------
>   fs/btrfs/sysfs.h   |  2 ++
>   3 files changed, 52 insertions(+), 18 deletions(-)
> 
> base-commit: 0f23d56f17fdfc7db69d51f64c8b91bbab947aa9

Re: Re: [PATCH v2 0/2] btrfs: delay mounted sysfs attributes until mount is ready
Posted by Jiacheng Xu 1 month, 1 week ago
Great! Please let me know if the patch is finally merged.

Thanks,
Jiacheng

> -----原始邮件-----
> 发件人: "Qu Wenruo" <wqu@suse.com>
> 发送时间:2026-08-22 12:57:14 (星期六)
> 收件人: "Jiacheng Xu" <stitch@zju.edu.cn>
> 抄送: linux-btrfs@vger.kernel.org, linux-kernel@vger.kernel.org
> 主题: Re: [PATCH v2 0/2] btrfs: delay mounted sysfs attributes until mount is ready
> 
> 
> 
> 在 2026/8/22 13:11, Jiacheng Xu 写道:
> > Here is a potential fix following Wenruo's idea.
> > 
> > btrfs_sysfs_add_mounted() currently publishes the writable label and
> > feature attributes before the transaction kthread is created. A concurrent
> > sysfs write can therefore dereference a NULL transaction_kthread in
> > wake_up_process().
> > 
> > This series follows the suggested lifecycle: create only the required
> > subdirectories during early mount, publish the fsid attributes after mount
> > initialization, and remove them before the kthreads are stopped. The
> > feature attributes are included because their store callback has the same
> > transaction_kthread dependency as the label callback.
> > 
> > Patch 1 factors the fsid attribute handling into dedicated helpers. Patch 2
> > moves their publication and removal to the safe mount and unmount stages.
> > On unmount the cleaner is parked before attribute removal so it cannot
> > recreate the feature group through sysfs_update_group(). Both patches are
> > required for stable backports.
> > 
> > The resulting fs/btrfs/sysfs.o and fs/btrfs/disk-io.o were build-tested.
> > 
> > Changes in v2:
> > - Delay creation of both the root and feature attributes until mount setup
> >    is complete.
> 
> You don't need to bother feature attributes for now, there is already a 
> patch addressing it by completely removing the write support for feature 
> attributes:
> 
> https://lore.kernel.org/linux-btrfs/8a598d76555b5944d34bb08fa8dbeea28fc05db9.1787307129.git.wqu@suse.com/
> 
> Considering it's only extended_iref, removing it should be much simpler.
> Until that is determined, you only need to bother the label one.
> 
> 
> Furthermore, among all the attr files in the fsid directory, there is 
> only label that is writable, it would make more sense to split 
> btrfs_attrs into two parts, one for those read-only members, and one for 
> the only writebale label one.
> 
> Otherwise the series looks much better.
> 
> 
> > - Remove those attributes while their kthread dependencies are still
> >    valid.
> > - Split helper extraction from the lifecycle fix for stable backports.
> > 
> > Jiacheng Xu (2):
> >    btrfs: sysfs: factor out mounted fsid attribute helpers
> >    btrfs: delay mounted fsid attributes until the fs is ready
> > 
> >   fs/btrfs/disk-io.c | 18 ++++++++++++++++-
> >   fs/btrfs/sysfs.c   | 50 ++++++++++++++++++++++++++++++----------------
> >   fs/btrfs/sysfs.h   |  2 ++
> >   3 files changed, 52 insertions(+), 18 deletions(-)
> > 
> > base-commit: 0f23d56f17fdfc7db69d51f64c8b91bbab947aa9
Re: Re: [PATCH v2 0/2] btrfs: delay mounted sysfs attributes until mount is ready
Posted by David Sterba 2 weeks, 6 days ago
On Sat, Aug 22, 2026 at 01:39:31PM +0800, Jiacheng Xu wrote:
> Great! Please let me know if the patch is finally merged.

The patch is now in our for-next branch, but will be merged to Linus'
tree in the next development cycle. Until then it will be in the
linux-next tree.