From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2EBE73DF000; Fri, 14 Aug 2026 09:39:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700400; cv=none; b=uzq8E1+FnmQiMOUIg+gW9XF8RxBzkmu2ajaJ8urXwQmNyO69rWfMcEcZCRn3LByjIfym7x7flKrEcwtxoie4+9K5XeHFO7gle8l40bdEmuogaBe3BguUtueinuiaD9fQq8qT4xbQE4Grr4lU4hFgAdgGUX0HMdQc4L38jNtK4yM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700400; c=relaxed/simple; bh=5aSAcl7xXxzU+2jbngERElG/1SSuffg/Yb9RSxc6sPE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PywM+t92ojKW4oxpEC1jnpxPJFfskj5nKiwfkk7mrmcpfwUm/pAj2JA9Pmv0rSlak2JhFH2nlYPCffQiROJdKo5upcj8/N2lexYx3hRI+fDpt4NEdl4M2IZKBKOg3lFnJgGv5OibakIxfPWZyeim+73aqxDpN8vKEzILnfBwSyM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyQ47wbzKHMP7; Fri, 14 Aug 2026 17:39:26 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 50B1E40570; Fri, 14 Aug 2026 17:39:47 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S5; Fri, 14 Aug 2026 17:39:46 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 01/32] ext4: simplify size updating in ext4_setattr() Date: Fri, 14 Aug 2026 17:33:00 +0800 Message-ID: <20260814093331.1703882-2-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S5 X-Coremail-Antispam: 1UD129KBjvJXoW7Aw4rtry7KFWUtr4UAF4xJFb_yoW8Kr1xpF y3Kw1vkw18WF1q9rn2gF1UZa48ta1093yUXFWUCw4IqFyDC3ZaqF17t3y3WFWrtrWkWw4Y qF4kGrs5Aw1UGrJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmmb4IE77IF4wAFF20E14v26rWj6s0DM7CY07I20VC2zVCF04k2 6cxKx2IYs7xG6rWj6s0DM7CIcVAFz4kK6r1j6r18M28IrcIa0xkI8VA2jI8067AKxVWUGw A2048vs2IY020Ec7CjxVAFwI0_Gr0_Xr1l8cAvFVAK0II2c7xJM28CjxkF64kEwVA0rcxS w2x7M28EF7xvwVC0I7IYx2IY67AKxVW5JVW7JwA2z4x0Y4vE2Ix0cI8IcVCY1x0267AKxV WxJVW8Jr1l84ACjcxK6I8E87Iv67AKxVWxJr0_GcWl84ACjcxK6I8E87Iv6xkF7I0E14v2 6rxl6s0DM2AIxVAIcxkEcVAq07x20xvEncxIr21l5I8CrVACY4xI64kE6c02F40Ex7xfMc Ij6xIIjxv20xvE14v26r1q6rW5McIj6I8E87Iv67AKxVWUJVW8JwAm72CE4IkC6x0Yz7v_ Jr0_Gr1lF7xvr2IYc2Ij64vIr41lF7I21c0EjII2zVCS5cI20VAGYxC7M4IIrI8v6xkF7I 0E8cxan2IY04v7MxkF7I0En4kS14v26r4a6rW5MxAIw28IcxkI7VAKI48JMxC20s026xCa FVCjc4AY6r1j6r4UMI8I3I0E5I8CrVAFwI0_Jr0_Jr4lx2IqxVCjr7xvwVAFwI0_JrI_Jr Wlx4CE17CEb7AF67AKxVW8ZVWrXwCIc40Y0x0EwIxGrwCI42IY6xIIjxv20xvE14v26ryj 6F1UMIIF0xvE2Ix0cI8IcVCY1x0267AKxVWxJVW8Jr1lIxAIcVCF04k26cxKx2IYs7xG6r 1j6r1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr0_Gr1U YxBIdaVFxhVjvjDU0xZFpf9x0piCD7fUUUUU= X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi The logic for updating the file size in ext4_setattr() is currently somewhat messy. By directly entering the error-handling path after failing to add an orphan inode, the unnecessary recovery process involving old_disksize and the file size can be avoided. Signed-off-by: Zhang Yi Reviewed-by: Jan Kara Reviewed-by: Ojaswin Mujoo --- fs/ext4/inode.c | 22 +++++++++------------- 1 file changed, 9 insertions(+), 13 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index bd4b778df9eb..14eab46f4750 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -6080,7 +6080,6 @@ int ext4_setattr(struct mnt_idmap *idmap, struct dent= ry *dentry, if (attr->ia_valid & ATTR_SIZE) { handle_t *handle; loff_t oldsize =3D inode->i_size; - loff_t old_disksize; int shrink =3D (attr->ia_size < inode->i_size); =20 if (!(ext4_test_inode_flag(inode, EXT4_INODE_EXTENTS))) { @@ -6164,6 +6163,8 @@ int ext4_setattr(struct mnt_idmap *idmap, struct dent= ry *dentry, if (ext4_handle_valid(handle) && shrink) { error =3D ext4_orphan_add(handle, inode); orphan =3D 1; + if (error) + goto out_handle; } =20 if (shrink) @@ -6179,23 +6180,18 @@ int ext4_setattr(struct mnt_idmap *idmap, struct de= ntry *dentry, (attr->ia_size > 0 ? attr->ia_size - 1 : 0) >> inode->i_sb->s_blocksize_bits); =20 - down_write(&EXT4_I(inode)->i_data_sem); - old_disksize =3D EXT4_I(inode)->i_disksize; - EXT4_I(inode)->i_disksize =3D attr->ia_size; - /* * We have to update i_size under i_data_sem together * with i_disksize to avoid races with writeback code - * running ext4_wb_update_i_disksize(). + * updating disksize in mpage_map_and_submit_extent(). */ - if (!error) - i_size_write(inode, attr->ia_size); - else - EXT4_I(inode)->i_disksize =3D old_disksize; + down_write(&EXT4_I(inode)->i_data_sem); + i_size_write(inode, attr->ia_size); + EXT4_I(inode)->i_disksize =3D attr->ia_size; up_write(&EXT4_I(inode)->i_data_sem); - rc =3D ext4_mark_inode_dirty(handle, inode); - if (!error) - error =3D rc; + + error =3D ext4_mark_inode_dirty(handle, inode); +out_handle: ext4_journal_stop(handle); if (error) goto out_mmap_sem; --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 312D1372EF0; Fri, 14 Aug 2026 09:39:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700395; cv=none; b=vFoDtOD1Wb0WT/h/YXzfAveUJXjncxdc9dGeNyy4sGwM9SmsC4ApUQapJys/QECMak4KMzUN3d+9SlurOYPSYyIpGSXDKYr9VF1Ze0/czjUiAmi+gIAOt7alH8TwU5+qBeQyL10qfM92PgaOfPk49D723zHB33cb7ONVrAWVB4U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700395; c=relaxed/simple; bh=+EKhVdqv3p8r1A06W61NHGrKxX2hg/mDn643J0UXrHA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=LAQ37OfH+NZexnnqwszFWDQ0i/JOD36r0XFFBcnrlMEvAwoJw/apf9tNNa4ECdsJeeIM2T4Ok6Fttm4sieYQhherHlMJ6CbE7+M4FYI9n08hCfLeurBKn2bqV0hOkoSCXdKMUP4/v4BNBFdbn3YOn/GphE1jhFubQ3kDgPN0+cE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyg1PGnzYQvPK; Fri, 14 Aug 2026 17:39:39 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 688F84056D; Fri, 14 Aug 2026 17:39:47 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S6; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 02/32] ext4: factor out ext4_truncate_[up|down]() Date: Fri, 14 Aug 2026 17:33:01 +0800 Message-ID: <20260814093331.1703882-3-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S6 X-Coremail-Antispam: 1UD129KBjvJXoW3WF48ur1xWFykXrW7ZF4ktFb_yoWxZw4kpF W2ka4Fkw18uFyDWF4Igr4UZF4fta18K3yUGFy2krs2v3Wqyw1ftF1xt3yFgFWUtrWDWw4Y qF4Dtrs3Gw4kJ3DanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmF14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_Jryl82xGYIkIc2 x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2z4x0 Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F4UJw A2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE3s1l e2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2Ix0cI 8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8JwAC jcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2ka0x kIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Yz7v_ Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zVAF1V AY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Xr0_Ar1lIxAI cVC0I7IYx2IY6xkF7I0E14v26F4j6r4UJwCI42IY6xAIw20EY4v20xvaj40_Jr0_JF4lIx AIcVC2z280aVAFwI0_Gr0_Cr1lIxAIcVC2z280aVCY1x0267AKxVW8JVW8JrUvcSsGvfC2 KfnxnUUI43ZEXa7sRiBMKPUUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Refactor ext4_setattr() by introducing two helper functions, ext4_truncate_up() and ext4_truncate_down(), to handle size changes. The current ATTR_SIZE processing consolidates checks for both shrinking and non-shrinking cases, leading to cluttered code. Separating the truncation paths improves readability. Signed-off-by: Zhang Yi Reviewed-by: Ojaswin Mujoo Reviewed-by: Jan Kara --- fs/ext4/inode.c | 199 +++++++++++++++++++++++++++--------------------- 1 file changed, 112 insertions(+), 87 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 14eab46f4750..8654006a57ef 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -5982,6 +5982,112 @@ static void ext4_wait_for_tail_page_commit(struct i= node *inode) } } =20 +/* + * Set i_size and i_disksize to 'newsize'. + * + * Both i_rwsem and i_data_sem are required here to avoid races between + * generic append writeback and concurrent truncate that also modify + * i_size and i_disksize. + */ +static inline void ext4_set_inode_size(struct inode *inode, loff_t newsize) +{ + WARN_ON_ONCE(S_ISREG(inode->i_mode) && !inode_is_locked(inode)); + + down_write(&EXT4_I(inode)->i_data_sem); + i_size_write(inode, newsize); + EXT4_I(inode)->i_disksize =3D newsize; + up_write(&EXT4_I(inode)->i_data_sem); +} + +static int ext4_truncate_up(struct inode *inode, loff_t oldsize, loff_t ne= wsize) +{ + ext4_lblk_t old_lblk, new_lblk; + handle_t *handle; + int ret; + + if (!IS_ALIGNED(oldsize | newsize, i_blocksize(inode))) { + ret =3D ext4_inode_attach_jinode(inode); + if (ret) + return ret; + } + + inode_set_mtime_to_ts(inode, inode_set_ctime_current(inode)); + if (!IS_ALIGNED(oldsize, i_blocksize(inode))) { + ret =3D ext4_block_zero_eof(inode, oldsize, LLONG_MAX); + if (ret) + return ret; + } + + handle =3D ext4_journal_start(inode, EXT4_HT_INODE, 3); + if (IS_ERR(handle)) + return PTR_ERR(handle); + + old_lblk =3D oldsize > 0 ? (oldsize - 1) >> inode->i_blkbits : 0; + new_lblk =3D newsize > 0 ? (newsize - 1) >> inode->i_blkbits : 0; + ext4_fc_track_range(handle, inode, old_lblk, new_lblk); + + ext4_set_inode_size(inode, newsize); + + ret =3D ext4_mark_inode_dirty(handle, inode); + ext4_journal_stop(handle); + if (ret) + return ret; + /* + * isize extend must be called outside an active handle due to + * the lock ordering of transaction start and folio lock in the + * iomap buffered I/O path (folio lock -> transaction start). + */ + pagecache_isize_extended(inode, oldsize, newsize); + return 0; +} + +static int ext4_truncate_down(struct inode *inode, loff_t oldsize, + loff_t newsize, int *orphan) +{ + ext4_lblk_t start_lblk; + handle_t *handle; + int ret; + + /* Do not change i_size. */ + if (newsize =3D=3D oldsize) + goto truncate; + + /* Shrink. */ + handle =3D ext4_journal_start(inode, EXT4_HT_INODE, 3); + if (IS_ERR(handle)) + return PTR_ERR(handle); + + if (ext4_handle_valid(handle)) { + ret =3D ext4_orphan_add(handle, inode); + *orphan =3D 1; + if (ret) { + ext4_journal_stop(handle); + return ret; + } + } + + start_lblk =3D newsize > 0 ? (newsize - 1) >> inode->i_blkbits : 0; + ext4_fc_track_range(handle, inode, start_lblk, EXT_MAX_BLOCKS - 1); + + ext4_set_inode_size(inode, newsize); + + ret =3D ext4_mark_inode_dirty(handle, inode); + ext4_journal_stop(handle); + if (ret) + return ret; + + if (ext4_should_journal_data(inode)) + ext4_wait_for_tail_page_commit(inode); +truncate: + /* + * Truncate pagecache after we've waited for commit in data=3Djournal + * mode to make pages freeable. Call ext4_truncate() even if + * i_size didn't change to truncate possible preallocated blocks. + */ + truncate_pagecache(inode, newsize); + return ext4_truncate(inode); +} + /* * ext4_setattr() * @@ -6078,7 +6184,6 @@ int ext4_setattr(struct mnt_idmap *idmap, struct dent= ry *dentry, } =20 if (attr->ia_valid & ATTR_SIZE) { - handle_t *handle; loff_t oldsize =3D inode->i_size; int shrink =3D (attr->ia_size < inode->i_size); =20 @@ -6130,94 +6235,14 @@ int ext4_setattr(struct mnt_idmap *idmap, struct de= ntry *dentry, goto err_out; } =20 - if (attr->ia_size !=3D inode->i_size) { - /* attach jbd2 jinode for EOF folio tail zeroing */ - if (attr->ia_size & (inode->i_sb->s_blocksize - 1) || - oldsize & (inode->i_sb->s_blocksize - 1)) { - error =3D ext4_inode_attach_jinode(inode); - if (error) - goto out_mmap_sem; - } - - /* - * Update c/mtime and tail zero the EOF folio on - * truncate up. ext4_truncate() handles the shrink case - * below. - */ - if (!shrink) { - inode_set_mtime_to_ts(inode, - inode_set_ctime_current(inode)); - if (oldsize & (inode->i_sb->s_blocksize - 1)) { - error =3D ext4_block_zero_eof(inode, - oldsize, LLONG_MAX); - if (error) - goto out_mmap_sem; - } - } - - handle =3D ext4_journal_start(inode, EXT4_HT_INODE, 3); - if (IS_ERR(handle)) { - error =3D PTR_ERR(handle); - goto out_mmap_sem; - } - if (ext4_handle_valid(handle) && shrink) { - error =3D ext4_orphan_add(handle, inode); - orphan =3D 1; - if (error) - goto out_handle; - } - - if (shrink) - ext4_fc_track_range(handle, inode, - (attr->ia_size > 0 ? attr->ia_size - 1 : 0) >> - inode->i_sb->s_blocksize_bits, - EXT_MAX_BLOCKS - 1); - else - ext4_fc_track_range( - handle, inode, - (oldsize > 0 ? oldsize - 1 : oldsize) >> - inode->i_sb->s_blocksize_bits, - (attr->ia_size > 0 ? attr->ia_size - 1 : 0) >> - inode->i_sb->s_blocksize_bits); - - /* - * We have to update i_size under i_data_sem together - * with i_disksize to avoid races with writeback code - * updating disksize in mpage_map_and_submit_extent(). - */ - down_write(&EXT4_I(inode)->i_data_sem); - i_size_write(inode, attr->ia_size); - EXT4_I(inode)->i_disksize =3D attr->ia_size; - up_write(&EXT4_I(inode)->i_data_sem); - - error =3D ext4_mark_inode_dirty(handle, inode); -out_handle: - ext4_journal_stop(handle); - if (error) - goto out_mmap_sem; - if (!shrink) { - pagecache_isize_extended(inode, oldsize, - inode->i_size); - } else if (ext4_should_journal_data(inode)) { - ext4_wait_for_tail_page_commit(inode); - } + if (attr->ia_size > oldsize) + error =3D ext4_truncate_up(inode, oldsize, attr->ia_size); + else { + /* Shrink or do not change i_size. */ + error =3D ext4_truncate_down(inode, oldsize, + attr->ia_size, &orphan); } =20 - /* - * Truncate pagecache after we've waited for commit - * in data=3Djournal mode to make pages freeable. - */ - truncate_pagecache(inode, inode->i_size); - /* - * Call ext4_truncate() even if i_size didn't change to - * truncate possible preallocated blocks. - */ - if (attr->ia_size <=3D oldsize) { - rc =3D ext4_truncate(inode); - if (rc) - error =3D rc; - } -out_mmap_sem: filemap_invalidate_unlock(inode->i_mapping); } =20 --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9236738D69D; Fri, 14 Aug 2026 09:39:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700403; cv=none; b=slrYANhBiS5gTQGyUgxsrvpZ6v3H3riIs88VkhxaMwZdxZW1+YbOkuXmK89jQoQ+e4sCSOE/1LV4i4JzypNKwPRVRgA53CLz8LFRqnzMQFrI+lt9PMogIZ+JSLlyPNNlEz8UQHp2RLFHYw0ivMADQshwu8B2e5IUWTZtyebhXuE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700403; c=relaxed/simple; bh=kAtOeW3AotyzWlGvjxc0dXev0Sckxwg/8ubM9GtQsa0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=FBkH+Y3eRT5ChQ6Ne40VOGIN1JzjCIXdrFXb8g/Qoczsj4mmajYkXKgfzidbvv/9rwkOMbzhROjrq1wWakNO7zdLzDOIF8PfxXgiIpfr60ypoG8nczAeJ3/twUQY/3RshsWKFEiHCXSsub/AmLoX5aqijDB0Cm2fGWuvioGkeK4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyQ5kVPzKHNKH; Fri, 14 Aug 2026 17:39:26 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 8630A40561; Fri, 14 Aug 2026 17:39:47 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S7; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 03/32] ext4: simplify error handling in ext4_setattr() Date: Fri, 14 Aug 2026 17:33:02 +0800 Message-ID: <20260814093331.1703882-4-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S7 X-Coremail-Antispam: 1UD129KBjvJXoWxZr13JF1kCFW3Zr1UCF1ftFb_yoW5Wr4UpF 13G3Wqkr48XF1UWF48tFy7Zw4Fga1Iq3yUXFyag3Z2kFn5Jr1qgFy2qayFgFy5GrWkGw1a qF4UKry3CFyUA3DanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmI14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JrWl82xGYIkIc2 x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2z4x0 Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F4UJw A2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE3s1l e2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2Ix0cI 8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8JwAC jcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2ka0x kIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Yz7v_ Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zVAF1V AY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Xr0_Ar1lIxAI cVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r1xMI IF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIYCTnI WIevJa73UjIFyTuYvjTRMUDJDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Refactor the error handling in ext4_setattr() for better clarity: - Return directly on ext4_break_layouts() failure. - Propagate ext4_truncate() errors using the existing error variable and jump to the common 'err_out' label. - Propagate posix_acl_chmod() errors also through the error variable, as it theoretically does not return a non-fatal error. With these changes, every error path either returns immediately or jumps to err_out. Consequently, the "if (!error)" condition guarding setattr_copy() and mark_inode_dirty() becomes unreachable for error cases. Remove this redundant check and the unused rc variable can be removed as well. Signed-off-by: Zhang Yi Reviewed-by: Ojaswin Mujoo Reviewed-by: Jan Kara --- fs/ext4/inode.c | 32 +++++++++++++++----------------- 1 file changed, 15 insertions(+), 17 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 8654006a57ef..76bf0e944ebe 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -6116,7 +6116,7 @@ int ext4_setattr(struct mnt_idmap *idmap, struct dent= ry *dentry, struct iattr *attr) { struct inode *inode =3D d_inode(dentry); - int error, rc =3D 0; + int error; int orphan =3D 0; const unsigned int ia_valid =3D attr->ia_valid; bool inc_ivers =3D true; @@ -6229,10 +6229,10 @@ int ext4_setattr(struct mnt_idmap *idmap, struct de= ntry *dentry, =20 filemap_invalidate_lock(inode->i_mapping); =20 - rc =3D ext4_break_layouts(inode); - if (rc) { + error =3D ext4_break_layouts(inode); + if (error) { filemap_invalidate_unlock(inode->i_mapping); - goto err_out; + return error; } =20 if (attr->ia_size > oldsize) @@ -6244,15 +6244,19 @@ int ext4_setattr(struct mnt_idmap *idmap, struct de= ntry *dentry, } =20 filemap_invalidate_unlock(inode->i_mapping); + if (error) + goto err_out; } =20 - if (!error) { - if (inc_ivers) - inode_inc_iversion(inode); - setattr_copy(idmap, inode, attr); - mark_inode_dirty(inode); - } + if (inc_ivers) + inode_inc_iversion(inode); + setattr_copy(idmap, inode, attr); + mark_inode_dirty(inode); =20 + if (ia_valid & ATTR_MODE) + error =3D posix_acl_chmod(idmap, dentry, inode->i_mode); + +err_out: /* * If the call to ext4_truncate failed to get a transaction handle at * all, we need to clean up the in-core orphan list manually. @@ -6260,14 +6264,8 @@ int ext4_setattr(struct mnt_idmap *idmap, struct den= try *dentry, if (orphan && inode->i_nlink) ext4_orphan_del(NULL, inode); =20 - if (!error && (ia_valid & ATTR_MODE)) - rc =3D posix_acl_chmod(idmap, dentry, inode->i_mode); - -err_out: - if (error) + if (error) ext4_std_error(inode->i_sb, error); - if (!error) - error =3D rc; return error; } =20 --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 31435377EC6; Fri, 14 Aug 2026 09:39:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700394; cv=none; b=G+NB5XVFVUjDQ5NdC6H8xK5dIWJ9wQUH/FhZrZE8C47PfhU2hUhhH/5C1TkZR7lR06sH9/9PXboWBIO4EhEmp9Ysowk+KcM4/FtD/SqGFvkl2jLtB94Z2AJOUcglbvE9Vp/fSDWchxopBUrpwhm/EMvU+JNxG7UDYt7C0ZtT1WU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700394; c=relaxed/simple; bh=fHnvvL2PcXYW2npa1XrNItDS5pTvHwExN8HGy/aZcDo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=UWrgURJy2v+zpkpqugx2keecXEAZsukxwkhTtTXY1p6N7m8chSD72NSlBrnsMynZnXXG/IPSd2O+oGFP1Fcz25TvqhEw61gj3U6fB48dF5fAwwvOcus47ggA/wzGC9TuZCq/lQSbs0jdAWhFofUENXL4DrZk0BNykTQqX0johnM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.177]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyg2whMzYQvQl; Fri, 14 Aug 2026 17:39:39 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 9770240590; Fri, 14 Aug 2026 17:39:47 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S8; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 04/32] ext4: skip ordered I/O wait when zeroing beyond i_disksize block Date: Fri, 14 Aug 2026 17:33:03 +0800 Message-ID: <20260814093331.1703882-5-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S8 X-Coremail-Antispam: 1UD129KBjvJXoW7CFy5Jr4fCw1fWw47uw1xXwb_yoW8Cr1xpa y3Gr18Cr4kG3s09w1vq3WIg34Ykan5Ga1rGFZrJr4q9FW3uw1v9F4xt34jvFW2yrs3Ga10 qF45GrW3Z34DA3DanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Xr0_Ar1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi ext4_block_zero_eof() zeros the tail of a partial block beyond EOF. After zeroing, it waits for ordered I/O completion to prevent stale data exposure from concurrent post-EOF mmap writes during folio writeback. However, if the zeroed range lies entirely beyond the block containing i_disksize, no stale data can be exposed because the zeroed region is beyond existing on-disk data. The zeroed pages will be written out before i_disksize is later extended past i_size, so the ordered I/O wait is unnecessary. Add a condition to skip it. Suggested-by: Ojaswin Mujoo Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 76bf0e944ebe..7601fe3618b1 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -4241,9 +4241,16 @@ int ext4_block_zero_eof(struct inode *inode, loff_t = from, loff_t end) * truncating up or performing an append write, because there might be * exposing stale on-disk data which may caused by concurrent post-EOF * mmap write during folio writeback. + * + * Ordered I/O is required only when zeroing the tail of a block that + * overlaps with i_disksize. If the zeroed range falls outside that + * block, the zeroed data lies beyond the existing on-disk data. It + * will be written out before i_disksize is later extended past + * i_size, so no stale data can be exposed. */ if (ext4_should_order_data(inode) && - did_zero && zero_written && !IS_DAX(inode)) { + did_zero && zero_written && !IS_DAX(inode) && + from < round_up(READ_ONCE(EXT4_I(inode)->i_disksize), blocksize)) { handle_t *handle; =20 handle =3D ext4_journal_start(inode, EXT4_HT_MISC, 1); --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4E45B377EC6; Fri, 14 Aug 2026 09:39:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700401; cv=none; b=MN83Nz347G1ROSkCdTqkKrLHVVEYxksxvV9nb87kc0xsZ3oIo+6Yq0nKzdZMlghbmpZIXj7zqawXa+LxUnT/rsaUHEk7IRmGqYlHR+7MkzTCt4fBe4vHtVfVggqwFGCPdCjqhPk9la0WfdrxglSHDu8rKtcqFyGCMORPAG01DCQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700401; c=relaxed/simple; bh=Ze8OLBqaSJHPqqxEAibPW8+QCySn1F2/YqZV0G4UaJo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=E6FAlfVjZUQHNRcAoNaEIJoq+NmaOomg+MxHe/apc3PKu0ebmxqvUsaodAMNIGn3+qHNmTzWO9cZX9CwQYkN7HiSkLkxnv10t46SCB5pdW2puz7xsFCc5QVwt/lm7jA7RXUlZgcj1EcKH2Qc6TVslMKJbQAhGTjw0yZCSElil1I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.177]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyR06qszKHNMR; Fri, 14 Aug 2026 17:39:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id AFB9840593; Fri, 14 Aug 2026 17:39:47 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S9; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 05/32] ext4: set EXT4_MAP_NEW flag for delayed allocated blocks Date: Fri, 14 Aug 2026 17:33:04 +0800 Message-ID: <20260814093331.1703882-6-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S9 X-Coremail-Antispam: 1UD129KBjvdXoW7Gr4xCr4xGrWUZF48XF1rCrg_yoWDZFgE9a yxZr48Zr1rJr1Ska98AF13Zr1vkw4xGr1xua45try3Xa48ArZ8G3ZxAFy3Za4DuF4rurZ8 AryrXrW2kFWIqjkaLaAFLSUrUUUUbb8apTn2vfkv8UJUUUU8Yxn0WfASr-VFAUDa7-sFnT 9fnUUIcSsGvfJTRUUUb98FF20E14v26rWj6s0DM7CY07I20VC2zVCF04k26cxKx2IYs7xG 6rWj6s0DM7CIcVAFz4kK6r1j6r18M28IrcIa0xkI8VA2jI8067AKxVWUAVCq3wA2048vs2 IY020Ec7CjxVAFwI0_Xr0E3s1l8cAvFVAK0II2c7xJM28CjxkF64kEwVA0rcxSw2x7M28E F7xvwVC0I7IYx2IY67AKxVW7JVWDJwA2z4x0Y4vE2Ix0cI8IcVCY1x0267AKxVW8Jr0_Cr 1UM28EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq 3wAS0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7 IYx2IY67AKxVWUtVWrXwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4U M4x0Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2 kIc2xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkE bVWUJVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67 AF67kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVW5JVW7JwCI 42IY6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F4UJwCI42IY6xAIw20EY4v20xvaj40_Jr0_JF 4lIxAIcVC2z280aVAFwI0_Gr0_Cr1lIxAIcVC2z280aVCY1x0267AKxVW8Jr0_Cr1UYxBI daVFxhVjvjDU0xZFpf9x0pRPCzZUUUUU= X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Set EXT4_MAP_NEW in ext4_da_map_blocks() to properly indicate that a new delayed allocation block has been inserted, allowing callers to distinguish newly created delayed extents from existing ones. Reported-by: Ojaswin Mujoo Link: https://lore.kernel.org/linux-ext4/cc05c17d-163e-4251-b2c9-aa3a6f9555= d7@huaweicloud.com/ Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 7601fe3618b1..9dbece14ae56 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -1990,7 +1990,7 @@ static int ext4_da_map_blocks(struct inode *inode, st= ruct ext4_map_blocks *map) } } =20 - map->m_flags |=3D EXT4_MAP_DELAYED; + map->m_flags |=3D EXT4_MAP_DELAYED | EXT4_MAP_NEW; retval =3D ext4_insert_delayed_blocks(inode, map->m_lblk, map->m_len); if (!retval) map->m_seq =3D READ_ONCE(EXT4_I(inode)->i_es_seq); --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2ECBA3E0C73; Fri, 14 Aug 2026 09:39:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700400; cv=none; b=Nj7U0k3h9LSfyaDfMx4HGfE2rEyDCSXwbqFu44Jw6D8UcmTmi4aDDElBRYGxWzi4RQa2jZ82cLdBbl9Dr7RCEeTaw4Q1EhmYFTNbaNtEk51jKllpjwbDocK6PhZbgzvtFsK+rz2dWMTJ/yed7fEydmSocYXQXUeb50il/10MME4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700400; c=relaxed/simple; bh=r5OWhPbD/O7OI1582H0ED157Jfnu0DM1I3zMlrkol8s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=spnxRJ7GsKBXOeYp8LjFFeuG7Bhk/CWomb8OLPs1W1Dk1S9yrcJFzeoOOdxyYRAgFgaGL9uGruFh3GVO2rlKPT4byMqBNMMfSz9Eu8UU/3mu6P0S4B0t7icxTMZF36v3DNvubRtgHtgo0ibOi6jOoS87BPJOSmzWksVpDvkfP0A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyR0VMRzKHNMZ; Fri, 14 Aug 2026 17:39:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id C76F64056D; Fri, 14 Aug 2026 17:39:47 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S10; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 06/32] ext4: recheck extent status tree before block allocation Date: Fri, 14 Aug 2026 17:33:05 +0800 Message-ID: <20260814093331.1703882-7-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S10 X-Coremail-Antispam: 1UD129KBjvJXoW7KrW8GFy3Xry8Cw4xJw17trb_yoW8urW3pr Zak34rGr1vgw1I9rZ7CF18ZF1Ska18XrW7JFZFqr1jvFyUWFyftFyYy3WSyFy5tws3tr4Y qFWrKryUuw4UArJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Xr0_Ar1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi After acquiring i_data_sem in write mode, recheck that the mapping found via the extent status tree or disk query has not changed. A racing truncate may have trimmed the extent between the earlier lookup and the write lock acquisition, since writeback does not hold i_rwsem or the folio locks covering the full extent. This could cause ext4_map_create_blocks() to allocate blocks beyond the truncated range, potentially leading to quota leaks in the upcomming iomap buffered writeback path since the iomap writeback infrastructure caches extents beyond the folio range. Therefore, if we find a valid extent and the sequence number has changed, retry the entire lookup to obtain the correct trimmed mapping. Suggested-by: Jan Kara Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 9dbece14ae56..548a3968c5a7 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -734,6 +734,7 @@ int ext4_map_blocks(handle_t *handle, struct inode *ino= de, else ext4_check_map_extents_env(inode); =20 +create_retry: /* Lookup extent status tree firstly */ if (ext4_es_lookup_extent(inode, map->m_lblk, NULL, &es, &map->m_seq)) { if (ext4_es_is_written(&es) || ext4_es_is_unwritten(&es)) { @@ -820,6 +821,19 @@ int ext4_map_blocks(handle_t *handle, struct inode *in= ode, * with create =3D=3D 1 flag. */ down_write(&EXT4_I(inode)->i_data_sem); + + /* + * Check the validity of the mapping found via the extent status + * tree or the disk query. A racing truncate may have changed the + * extent, since writeback does not hold i_rwsem or the folio locks + * covering the full extent. + */ + if (map->m_seq !=3D READ_ONCE(EXT4_I(inode)->i_es_seq)) { + up_write(&EXT4_I(inode)->i_data_sem); + map->m_flags =3D 0; + map->m_len =3D orig_mlen; + goto create_retry; + } retval =3D ext4_map_create_blocks(handle, inode, map, flags); up_write((&EXT4_I(inode)->i_data_sem)); =20 --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EBD6D3CD8A9; Fri, 14 Aug 2026 09:39:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700399; cv=none; b=GRxY6BEVNj08cmCfrO4sQVhOw520plodpV2r9jQs9rAptzTfRu2ozDKqv+QEd9vTDPB3qlOQiT8NVqP2tYMwsmHmkSdsocn5dqGA+c3yDz68i+KXm095XjJ7fVdjLrsntE/v4Owvxspir0R3Y6PWrpKXgH5M0t37SjkDAQ3K8Rg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700399; c=relaxed/simple; bh=APr+NYktTjIAEK7m2DM/Vmzfna01JCkrnCltT5AwyGo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pTkF3937CoSl8lFZuG89IrklCsrqdRuAqACQExTZiaXilwmoqQZhhBS2gZVmvA6Dp9+if2SPwgPwnbrq7xJtNtvdeR7ESd2ouRkLbzvwKTgqe5cTKyhvCoOU/LCK51kl7U5784xlOv5HY6bk9zA0bYWOQBu5FtFCn9qpBg4tRmA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.198]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyR1D0szKHNMp; Fri, 14 Aug 2026 17:39:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id DCE5740ADD; Fri, 14 Aug 2026 17:39:47 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S11; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 07/32] ext4: fix orig_mlen initialization in ext4_map_blocks() Date: Fri, 14 Aug 2026 17:33:06 +0800 Message-ID: <20260814093331.1703882-8-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S11 X-Coremail-Antispam: 1UD129KBjvdXoWrZF4xur1UAF1fAF1rtr13urg_yoWkGFcEqa y2vr48Gw4rArnakr4kZr4fJFyvkFy8Wr18CrW7Xry8XFn5ta95Ww1vyF9rAr4DW3yYvrZ8 ZFy8J34xKFW7ZjkaLaAFLSUrUUUUbb8apTn2vfkv8UJUUUU8Yxn0WfASr-VFAUDa7-sFnT 9fnUUIcSsGvfJTRUUUb98FF20E14v26rWj6s0DM7CY07I20VC2zVCF04k26cxKx2IYs7xG 6rWj6s0DM7CIcVAFz4kK6r1j6r18M28IrcIa0xkI8VA2jI8067AKxVWUAVCq3wA2048vs2 IY020Ec7CjxVAFwI0_Xr0E3s1l8cAvFVAK0II2c7xJM28CjxkF64kEwVA0rcxSw2x7M28E F7xvwVC0I7IYx2IY67AKxVW7JVWDJwA2z4x0Y4vE2Ix0cI8IcVCY1x0267AKxVW8Jr0_Cr 1UM28EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq 3wAS0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7 IYx2IY67AKxVWUtVWrXwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4U M4x0Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2 kIc2xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkE bVWUJVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67 AF67kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVW7JVWDJwCI 42IY6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F4UJwCI42IY6xAIw20EY4v20xvaj40_Jr0_JF 4lIxAIcVC2z280aVAFwI0_Gr0_Cr1lIxAIcVC2z280aVCY1x0267AKxVW8Jr0_Cr1UYxBI daVFxhVjvjDU0xZFpf9x0pRPCzZUUUUU= X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Save orig_mlen after clamping map->m_len to INT_MAX. Otherwise, the unclamped value may be passed below, bypassing its overflow protection. Fixes: 5bb12b1837c0 ("ext4: Add support for EXT4_GET_BLOCKS_QUERY_LEAF_BLOC= KS") Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 548a3968c5a7..bb4f1079d989 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -703,7 +703,7 @@ int ext4_map_blocks(handle_t *handle, struct inode *ino= de, struct extent_status es; int retval; int ret =3D 0; - unsigned int orig_mlen =3D map->m_len; + unsigned int orig_mlen; #ifdef ES_AGGRESSIVE_TEST struct ext4_map_blocks orig_map; =20 @@ -719,6 +719,7 @@ int ext4_map_blocks(handle_t *handle, struct inode *ino= de, */ if (unlikely(map->m_len > INT_MAX)) map->m_len =3D INT_MAX; + orig_mlen =3D map->m_len; =20 /* We can handle the block number less than EXT_MAX_BLOCKS */ if (unlikely(map->m_lblk >=3D EXT_MAX_BLOCKS)) --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DBA683749EE; Fri, 14 Aug 2026 09:40:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700407; cv=none; b=jkGdb38gOsopi2CGmuUKX+9coNy9JD+f2VdtEtbQC3sAoG/Zx79eK8dS/wW8GBsbJgrEBCgKYEL/SByAIrg0UpBCTNPiRFuvT54cEOtHHkLvdWmkupXvivCfwZ5w6f+3+TdKJs1gX5NAS3OmofXP/aGeRSCM/ISpcY/ucTKRdtg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700407; c=relaxed/simple; bh=6mte78s0bke1ru9B3/dIfd4NVCM2mtnGuNF6c83r3v4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eIJcQNY+dlgHHuCFZsNddT08Nf/mBiK8de33vDIMwiZZ4z9Znk099d/IXhwsojbW+AG1srmrcp0mQK60nPoZgbjG0hs6XOMQhQeHcLQ/92mVa4/7qKQNbgwdhQIlLi8Jxax2Y1WfuE+0BqSh2ynbp5aDKac5OrFaZh2LUdnuxg4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.198]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyR25YKzKHNN0; Fri, 14 Aug 2026 17:39:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 0921B40AE8; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S12; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 08/32] ext4: allow ext4_map_blocks() to start its own transaction handle Date: Fri, 14 Aug 2026 17:33:07 +0800 Message-ID: <20260814093331.1703882-9-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S12 X-Coremail-Antispam: 1UD129KBjvJXoWxGryUJFWfAryDJF43ur17Jrb_yoW5XrW5pr WakryrCr1UWF9a9F4Ska1UZF1aka48KrWUuFWfGryrC34a9rnagF1UK3WYyFWrtrWfWa1j qF45tryUCa1jk3DanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Make ext4_map_blocks() start its own transaction handle when the caller does not provide one. The handle is started after the lookup path confirms that allocation is actually needed, and is stopped at the unified out_handle exit path. This avoids unnecessarily starting a handle for pure mapping queries. This prepares for the buffered iomap writeback conversion, which improves performance for fragile overwrite cases. Suggested-by: Jan Kara Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 34 ++++++++++++++++++++++++++-------- 1 file changed, 26 insertions(+), 8 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index bb4f1079d989..c9904c274347 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -703,6 +703,7 @@ int ext4_map_blocks(handle_t *handle, struct inode *ino= de, struct extent_status es; int retval; int ret =3D 0; + bool internal_handle =3D false; unsigned int orig_mlen; #ifdef ES_AGGRESSIVE_TEST struct ext4_map_blocks orig_map; @@ -790,8 +791,10 @@ int ext4_map_blocks(handle_t *handle, struct inode *in= ode, found: if (retval > 0 && map->m_flags & EXT4_MAP_MAPPED) { ret =3D check_block_validity(inode, map); - if (ret !=3D 0) - return ret; + if (ret !=3D 0) { + retval =3D ret; + goto out_handle; + } } =20 /* If it is only a block(s) look up */ @@ -811,8 +814,15 @@ int ext4_map_blocks(handle_t *handle, struct inode *in= ode, * ext4_ext_map_blocks() */ if (!(flags & EXT4_GET_BLOCKS_CONVERT_UNWRITTEN)) - return retval; + goto out_handle; =20 + if (!handle) { + handle =3D ext4_journal_start(inode, EXT4_HT_MAP_BLOCKS, + ext4_chunk_trans_blocks(inode, orig_mlen)); + if (IS_ERR(handle)) + return PTR_ERR(handle); + internal_handle =3D true; + } =20 ext4_fc_track_inode(handle, inode); /* @@ -841,12 +851,14 @@ int ext4_map_blocks(handle_t *handle, struct inode *i= node, if (retval < 0) ext_debug(inode, "failed with err %d\n", retval); if (retval <=3D 0) - return retval; + goto out_handle; =20 if (map->m_flags & EXT4_MAP_MAPPED) { ret =3D check_block_validity(inode, map); - if (ret !=3D 0) - return ret; + if (ret !=3D 0) { + retval =3D ret; + goto out_handle; + } =20 /* * Inodes with freshly allocated blocks where contents will be @@ -867,12 +879,18 @@ int ext4_map_blocks(handle_t *handle, struct inode *i= node, else ret =3D ext4_jbd2_inode_add_write(handle, inode, start_byte, length); - if (ret) - return ret; + if (ret) { + retval =3D ret; + goto out_handle; + } } } ext4_fc_track_range(handle, inode, map->m_lblk, map->m_lblk + map->m_len - 1); + +out_handle: + if (internal_handle) + ext4_journal_stop(handle); return retval; } =20 --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7B5BC37A83F; Fri, 14 Aug 2026 09:39:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700396; cv=none; b=LBPvk3ZIIHWCCuBsZRLclN1oN9qSgwfR4y0ypnypPtiUzP+kiNyRYJkyPsR2Wh+2ke8TZd9pWGPuk35PziHOSCWMRrpWdLE6Ydv+6I+LMRphgrNTikVOrB4Fuse7GcV49SYeAMcZny7onrAajIpj1vgeq8RlK+3bzqJGaeWJHTg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700396; c=relaxed/simple; bh=H1oWZGvtPgJom00NhUthqXLncw8xGcA4wK1lsyNMcgg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=GOOqWoXStJsT53bgqVquUeUJT8jtUZQZoaUje0+9ef04yyWGT6QXVbi/6YZAzZJ4RgU575O73xZNwoGhoRMwXDwG9ei4J57hQ0C4PBBEP/HK219/u79/vj0kLX/mvqYRhSx5IJsPoz+FrmpZzMzOGpnKyhQpyzID7mxUyRwbu/A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyg64TKzYQvPN; Fri, 14 Aug 2026 17:39:39 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 17A7C40571; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S13; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 09/32] ext4: avoid unnecessary transaction in ext4_map_blocks() for unwritten extents Date: Fri, 14 Aug 2026 17:33:08 +0800 Message-ID: <20260814093331.1703882-10-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S13 X-Coremail-Antispam: 1UD129KBjvJXoW7CFW8Jw1fuFWkJFyUJFWUArb_yoW8Cw1fpa sIkr1rCF40g343WaySyF4jgFWakw1xKFWDuF48KryUZa43Kr1SgF10qF1rGFWUKrWfAay5 XFWjkw18C3Z5CrDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi When ext4_map_blocks() finds an unwritten extent in the extent cache and the caller is willing to accept unwritten extents without conversion, there is no need to start a journal transaction since no metadata update is required. This avoids unnecessary transaction overhead in the upcoming iomap writeback path when overwriting already-allocated unwritten extents. One thing to be careful about, as the comment in ext4_map_blocks() states, if the flags contain EXT4_GET_BLOCKS_CREATE, the function will mark @map as mapped. Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 19 ++++++++++++++----- 1 file changed, 14 insertions(+), 5 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index c9904c274347..5dcc3f7b2ffd 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -807,14 +807,23 @@ int ext4_map_blocks(handle_t *handle, struct inode *i= node, * Note that if blocks have been preallocated * ext4_ext_map_blocks() returns with buffer head unmapped */ - if (retval > 0 && map->m_flags & EXT4_MAP_MAPPED) + if (retval > 0) { /* - * If we need to convert extent to unwritten - * we continue and do the actual work in - * ext4_ext_map_blocks() + * If we need to convert written extent to unwritten or + * convert unwritten extent to written, continue and do + * the actual work in ext4_ext_map_blocks(). */ - if (!(flags & EXT4_GET_BLOCKS_CONVERT_UNWRITTEN)) + if (map->m_flags & EXT4_MAP_MAPPED && + !(flags & EXT4_GET_BLOCKS_CONVERT_UNWRITTEN)) goto out_handle; + if (map->m_flags & EXT4_MAP_UNWRITTEN && + (flags & EXT4_GET_BLOCKS_UNWRIT_EXT) && + !(flags & EXT4_GET_BLOCKS_CONVERT)) { + /* Contains EXT4_GET_BLOCKS_CREATE - mark mapped. */ + map->m_flags |=3D EXT4_MAP_MAPPED; + goto out_handle; + } + } =20 if (!handle) { handle =3D ext4_journal_start(inode, EXT4_HT_MAP_BLOCKS, --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 500EC436BCE; Fri, 14 Aug 2026 09:40:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700407; cv=none; b=qp3g3KVftB5jFIB0/JoptuQi9hLLY0wzgRmGr/Db21DxtAL7bTyk1Gbs2/E06vqwZPIh9utR+ojJUa+u1qS8ACBFSlY0INt3GMVkdx6JL/9ioLGqeHXe+w31pnxcUnqWpC4RRxLGK7hMO4dnXqQt+MRRKdPv88NTEWYXvORuw1U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700407; c=relaxed/simple; bh=ZUs5OSB0xwNvtzbSBvm0XwZqA//Lf5s0xp6XF97fK50=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pkTbYfojJVFUTuQ8AKWeUgTqV6JQ3/ZUjGXA/BFw1SpVO/Px6ZGhp1xyyoKOVBL60vfo+ZIzkfMSaPvHxH9/0F3KUa8F2eiV3eu2xjU09SKMxEa+DdeAoQHmFnF2KWjbaM6bkby/fjsKEi32tqtWW+sTw2mr3mqj8oi3TsGTDHw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyR35r3zKHNM0; Fri, 14 Aug 2026 17:39:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 2B6964056B; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S14; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 10/32] ext4: skip block allocation for holes in the data submission path Date: Fri, 14 Aug 2026 17:33:09 +0800 Message-ID: <20260814093331.1703882-11-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S14 X-Coremail-Antispam: 1UD129KBjvJXoWxAFyxXr18XF17GF1DJFW5KFg_yoW5Jw43pr 9IkryrGr4qgw1a9a97Ca1UXF1Yk3WxGF47CFWrJrWjgry3JF1SqFWUK3WYyFWrKrW8XFWS qF4F934rAr9Yy3DanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi When ext4_map_blocks() is called from the data submission path and I/O end extent conversion path (EXT4_GET_BLOCKS_IO_SUBMIT), it should not allocate blocks if the lookup returns a hole. The writeback path can legitimately encounter dirty ranges that map to holes. For example, when a folio straddles i_size and the tail beyond i_size is dirtied via a mmap write. Allocating blocks for such ranges is wrong because there is no data to write back, the dirty bits should simply be discarded without submitting I/O. This mirrors the existing buffer_head writeback path, where mpage_add_bh_to_extent() skips unmapped buffers and ext4_bio_write_folio() clears their dirty bits. In the ioend extent conversion path, holes are also not expected because we should wait for folio writeback before punching hole. If one is encountered, it likely indicates a failure in the concurrency protection. In this case, to avoid losing data beyond the hole, do not stop conversion, continue on the remaining ranges. This prepares for the buffered iomap writeback conversion. Signed-off-by: Zhang Yi --- fs/ext4/extents.c | 8 ++++++-- fs/ext4/inode.c | 7 +++++++ 2 files changed, 13 insertions(+), 2 deletions(-) diff --git a/fs/ext4/extents.c b/fs/ext4/extents.c index 76038b6c3655..0d62d9312284 100644 --- a/fs/ext4/extents.c +++ b/fs/ext4/extents.c @@ -5167,11 +5167,15 @@ int ext4_convert_unwritten_extents(handle_t *handle= , struct inode *inode, EXT4_GET_BLOCKS_IO_CONVERT_EXT | EXT4_EX_NOCACHE); if (ret <=3D 0) { + /* + * If the ret is zero, an unexpected hole may cause + * conversion to fail. To avoid data loss during I/O + * end conversion, skip the hole and continue + * converting subsequent blocks. + */ ext4_warning(inode->i_sb, "inode #%llu: block %u: len %u: ext4_map_blocks returned %d", inode->i_ino, map.m_lblk, map.m_len, ret); - if (unlikely(ret =3D=3D 0)) - ret =3D -EINVAL; } else { conv_blocks +=3D map.m_len; } diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 5dcc3f7b2ffd..d8c3e5e13b8a 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -823,6 +823,13 @@ int ext4_map_blocks(handle_t *handle, struct inode *in= ode, map->m_flags |=3D EXT4_MAP_MAPPED; goto out_handle; } + } else if (retval =3D=3D 0) { + /* + * Do not allocate blocks for holes in the context of + * data submission path. + */ + if (!map->m_flags && (flags & EXT4_GET_BLOCKS_IO_SUBMIT)) + goto out_handle; } =20 if (!handle) { --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3125836F914; Fri, 14 Aug 2026 09:39:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700395; cv=none; b=g/qDqeqKTAwxZiijoRv8/+8lRe6jRX1kG4qBnww9ZyEM76bU+1WP0A238njnrJHH3YN9gZVKy2YyXka5fF7CAHcL3knlVmOXelP+B1v5sntQrMZBiS1irlZlGMYzmUUowEfAxNQ+4tnq1kfBvCXzoPwIH4b9iHAPtnorsSs+J3w= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700395; c=relaxed/simple; bh=eJmCEzvJstTWcaZTQCEIjieC/4K6tro5SKJzL2Mf1Xk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=fPYCfSlsDZtmFFDhCl7PX5mrhffjR/C2JgDMxOCZvC6jOo4NoA9zpaUboU/s0tEW5aYmrDb6PbjrKGw91WupLHG7VCMITEJo+YzDoVGqZkZ05HCKJzdyFr80APtnWgQNONzpGzxJVHN3iGzN2Khq6via1pnDgx3AVC7t1L0+EeI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.198]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyh0XPNzYQvT9; Fri, 14 Aug 2026 17:39:40 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 4BEED40BFC; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S15; Fri, 14 Aug 2026 17:39:47 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 11/32] ext4: add iomap address space operations for buffered I/O Date: Fri, 14 Aug 2026 17:33:10 +0800 Message-ID: <20260814093331.1703882-12-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S15 X-Coremail-Antispam: 1UD129KBjvJXoWxXF1DuryDWryrtF4rGryfZwb_yoWrJw1fpF 98Kas8GF18XF9F9a1Sqa9rZrWYya4fGw4jgFW3W3ZI9F15GrW2gFW0k3WYyFy5t3ykJr12 qF4j9ry7WF17ArDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Introduce initial support for iomap in the buffered I/O path for regular files on ext4. - Add a new inode state flag EXT4_STATE_BUFFERED_IOMAP to indicate the inode uses iomap instead of buffer_head for buffered I/O - Add helper ext4_inode_buffered_iomap() to check the flag - Add new address space operations ext4_iomap_aops with callbacks that will use generic iomap implementations - Add ext4_iomap_aops to ext4_set_aops() when the flag is set The following callbacks(read_folio(), readahead(), writepages()) are provided as placeholders and will be implemented in later patches. Signed-off-by: Zhang Yi Reviewed-by: Jan Kara Reviewed-by: Ojaswin Mujoo --- fs/ext4/ext4.h | 7 +++++++ fs/ext4/inode.c | 32 ++++++++++++++++++++++++++++++++ 2 files changed, 39 insertions(+) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 724a27e8be61..24ec205da2d7 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -2049,6 +2049,7 @@ enum { EXT4_STATE_FC_FLUSHING_DATA, /* Fast commit flushing data */ EXT4_STATE_ORPHAN_FILE, /* Inode orphaned in orphan file */ EXT4_STATE_FC_REQUEUE, /* Inode modified during fast commit */ + EXT4_STATE_BUFFERED_IOMAP, /* Inode use iomap for buffered IO */ }; =20 #define EXT4_INODE_BIT_FNS(name, field, offset) \ @@ -2148,6 +2149,12 @@ static inline struct mapping_metadata_bhs *ext4_i_me= tadata_bhs( return READ_ONCE(EXT4_I(inode)->i_metadata_bhs); } =20 +/* Whether the inode pass through the iomap infrastructure for buffered I/= O */ +static inline bool ext4_inode_buffered_iomap(struct inode *inode) +{ + return ext4_test_inode_state(inode, EXT4_STATE_BUFFERED_IOMAP); +} + /* * Codes for operating systems */ diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index d8c3e5e13b8a..dda78cf1f68d 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -3956,6 +3956,22 @@ const struct iomap_ops ext4_iomap_report_ops =3D { .iomap_next =3D ext4_iomap_next_report, }; =20 +static int ext4_iomap_read_folio(struct file *file, struct folio *folio) +{ + return 0; +} + +static void ext4_iomap_readahead(struct readahead_control *rac) +{ + +} + +static int ext4_iomap_writepages(struct address_space *mapping, + struct writeback_control *wbc) +{ + return 0; +} + /* * For data=3Djournal mode, folio should be marked dirty only when it was * writeably mapped. When that happens, it was already attached to the @@ -4042,6 +4058,20 @@ static const struct address_space_operations ext4_da= _aops =3D { .swap_activate =3D ext4_iomap_swap_activate, }; =20 +static const struct address_space_operations ext4_iomap_aops =3D { + .read_folio =3D ext4_iomap_read_folio, + .readahead =3D ext4_iomap_readahead, + .writepages =3D ext4_iomap_writepages, + .dirty_folio =3D iomap_dirty_folio, + .bmap =3D ext4_bmap, + .invalidate_folio =3D iomap_invalidate_folio, + .release_folio =3D iomap_release_folio, + .migrate_folio =3D filemap_migrate_folio, + .is_partially_uptodate =3D iomap_is_partially_uptodate, + .error_remove_folio =3D generic_error_remove_folio, + .swap_activate =3D ext4_iomap_swap_activate, +}; + static const struct address_space_operations ext4_dax_aops =3D { .writepages =3D ext4_dax_writepages, .dirty_folio =3D noop_dirty_folio, @@ -4063,6 +4093,8 @@ void ext4_set_aops(struct inode *inode) } if (IS_DAX(inode)) inode->i_mapping->a_ops =3D &ext4_dax_aops; + else if (ext4_inode_buffered_iomap(inode)) + inode->i_mapping->a_ops =3D &ext4_iomap_aops; else if (test_opt(inode->i_sb, DELALLOC)) inode->i_mapping->a_ops =3D &ext4_da_aops; else --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E971843B3D5; Fri, 14 Aug 2026 09:40:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700408; cv=none; b=V9Aew4VPQrW7qYI/nBKpCbIsQNWbKNeAnHN7P+s3kDJli7tjz/SK7J5DOZKSkag+ifalLQFBXEbFnFuTejeEjUmRTQ0KzlTY8KYuH72dBhD6L914L5PTe0V7ideLhEl5H8CNnr4c1ViQ7htZPKGfW4T9byNuB9vrgDX7n22OV74= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700408; c=relaxed/simple; bh=FZtNPEg5aF+ddVQ22+g1fVIG2upqmDvAOr2Tqs5h1pM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=alqaYoLAwHORRvUG8xWF/Z2+MtpvFtb8LqBRg9FW6ZNf8lsWxtJ3rV8VEoV4nxX1dldktXOj9ya9GekBTyZbT9p4iHxPM44g99P1sr3UYfslIpoX514rcLlWbFl+HgXMV0MTO0ZcbS+3rZiMAE8tXmCn6YPl37FLolBRMZyoQAI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.198]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyR4mqbzKHNNF; Fri, 14 Aug 2026 17:39:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 6590B40BFF; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S16; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 12/32] ext4: implement buffered read path using iomap Date: Fri, 14 Aug 2026 17:33:11 +0800 Message-ID: <20260814093331.1703882-13-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S16 X-Coremail-Antispam: 1UD129KBjvJXoWxGrW3tr4DXry8JFWkKryUZFb_yoW5Ww1DpF 90kFy5Gr4UWrnF9F4SqFZrAr1Yka1xJa1UWrWfGwnxWF90krWagayUGF1YvF45t3y7AF18 XF4Ykry8Wa1UArDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Implement the iomap read path for ext4 by introducing a new ext4_iomap_buffered_read_ops instance. This provides the read_folio() and readahead() callbacks for ext4_iomap_aops. The implementation introduces: - ext4_iomap_map_blocks(): Helper function to query extent mappings for a given read range using ext4_map_blocks() and convert the mapping information to iomap type - ext4_iomap_buffered_read_begin(): The iomap_begin callbacks that maps blocks, validates filesystem state, and populates the iomap. It returns -ERANGE for inline data which is not yet supported. Signed-off-by: Zhang Yi Reviewed-by: Jan Kara Reviewed-by: Ojaswin Mujoo --- fs/ext4/inode.c | 48 +++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 47 insertions(+), 1 deletion(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index dda78cf1f68d..376cb9783835 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -3956,14 +3956,60 @@ const struct iomap_ops ext4_iomap_report_ops =3D { .iomap_next =3D ext4_iomap_next_report, }; =20 +static int ext4_iomap_map_blocks(struct inode *inode, loff_t offset, + loff_t length, struct ext4_map_blocks *map) +{ + u8 blkbits =3D inode->i_blkbits; + + if ((offset >> blkbits) > EXT4_MAX_LOGICAL_BLOCK) + return -EINVAL; + + /* Calculate the first and last logical blocks respectively. */ + map->m_lblk =3D offset >> blkbits; + map->m_len =3D min_t(loff_t, (offset + length - 1) >> blkbits, + EXT4_MAX_LOGICAL_BLOCK) - map->m_lblk + 1; + + return ext4_map_blocks(NULL, inode, map, 0); +} + +static int ext4_iomap_buffered_read_begin(struct inode *inode, loff_t offs= et, + loff_t length, unsigned int flags, struct iomap *iomap, + struct iomap *srcmap) +{ + struct ext4_map_blocks map; + int ret; + + if (unlikely(ext4_forced_shutdown(inode->i_sb))) + return -EIO; + + /* Inline data support is not yet available. */ + if (WARN_ON_ONCE(ext4_has_inline_data(inode))) + return -ERANGE; + + ret =3D ext4_iomap_map_blocks(inode, offset, length, &map); + if (ret < 0) + return ret; + + ext4_set_iomap(inode, iomap, &map, offset, length, flags); + return 0; +} + +static DEFINE_IOMAP_ITER_NEXT(ext4_iomap_buffered_read_next, + ext4_iomap_buffered_read_begin); + +const struct iomap_ops ext4_iomap_buffered_read_ops =3D { + .iomap_next =3D ext4_iomap_buffered_read_next, +}; + static int ext4_iomap_read_folio(struct file *file, struct folio *folio) { + iomap_bio_read_folio(folio, &ext4_iomap_buffered_read_ops); return 0; } =20 static void ext4_iomap_readahead(struct readahead_control *rac) { - + iomap_bio_readahead(rac, &ext4_iomap_buffered_read_ops); } =20 static int ext4_iomap_writepages(struct address_space *mapping, --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D237243F4DB; Fri, 14 Aug 2026 09:40:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700409; cv=none; b=dffOse/Wqiq22n08Zc8nITVPEN2r0Dm0MMWawYJEUEXE1YOLz2qiIbEpSRNi5w3J0/dTWdbPR+jea124aqgxapMWECTkT7MfdbRPvC1OwcX5WwXd8edi1FkhmHzEbDX0HrTptDkykSCiKL1iAl0JYyUT+/hOYniOW0m5f0MPMQ0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700409; c=relaxed/simple; bh=LRvGFX15gSlEMgzikCfEppkdkc3qZZAtD8cFysYgEmw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=iDYyUASJSrJ/o8N8jI+FrPeHq+w0E5QkUcicnCFJLGobsjS8nPiQwWSvn/zkP7tR4mwhN+Pp6+lLwqnkzSKAQWz7+zd8RFqE+KvE4NbWDybccMlLsVekaUlrsRZ9trYJ0+FlMl6uErBekaCRNezFVcG5U29Z6OeFPxjHaNouCSk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.198]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyR5Tx5zKHNL5; Fri, 14 Aug 2026 17:39:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 7F8CA40B08; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S17; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 13/32] ext4: pass out extent seq counter when mapping da blocks Date: Fri, 14 Aug 2026 17:33:12 +0800 Message-ID: <20260814093331.1703882-14-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S17 X-Coremail-Antispam: 1UD129KBjvJXoW7Zw4UKFy7uF1xXFyUZr15twb_yoW8uFWDp3 9Ykr15Gw1xZw1v9ayxXF1xZFy5Kay5JrW7GFWfXw1Ygas8WFySgF1jkF12yFykKr4xXr1F vF40kry8Ca4SyFDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi The iomap buffered write path does not hold the folio lock between mapping the inode extent and copying data. Therefore, it can race with writeback that modifies the extent type (e.g., from unwritten to written). This can lead to data corruption on partial writes, as iomap_block_needs_zeroing() may return a false positive based on a stale extent. The iomap infrastructure uses the sequence counter stored in the inode to detect such stale mappings. Commit 07c440e8da8f ("ext4: pass out extent seq counter when mapping blocks") added the m_seq field to ext4_map_blocks to pass out extent sequence numbers, but it missed two callsites within ext4_da_map_blocks(). These callsites are on the delayed allocation path, which is needed in the iomap buffered write path. Pass out the sequence counter to ensure stale mappings can be detected. Signed-off-by: Zhang Yi Reviewed-by: Jan Kara Reviewed-by: Ojaswin Mujoo --- fs/ext4/inode.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 376cb9783835..9098d9a5fc05 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -1970,7 +1970,7 @@ static int ext4_da_map_blocks(struct inode *inode, st= ruct ext4_map_blocks *map) ext4_check_map_extents_env(inode); =20 /* Lookup extent status tree firstly */ - if (ext4_es_lookup_extent(inode, map->m_lblk, NULL, &es, NULL)) { + if (ext4_es_lookup_extent(inode, map->m_lblk, NULL, &es, &map->m_seq)) { map->m_len =3D min_t(unsigned int, map->m_len, es.es_len - (map->m_lblk - es.es_lblk)); =20 @@ -2023,7 +2023,7 @@ static int ext4_da_map_blocks(struct inode *inode, st= ruct ext4_map_blocks *map) * is held in write mode, before inserting a new da entry in * the extent status tree. */ - if (ext4_es_lookup_extent(inode, map->m_lblk, NULL, &es, NULL)) { + if (ext4_es_lookup_extent(inode, map->m_lblk, NULL, &es, &map->m_seq)) { map->m_len =3D min_t(unsigned int, map->m_len, es.es_len - (map->m_lblk - es.es_lblk)); =20 --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 56AC943CE64; Fri, 14 Aug 2026 09:40:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700410; cv=none; b=TpxkzbDw4cXQAbseITlhnxnI2tPUbPZ7B2jzYbY34jgPWF7jccDUp0p0gtDYma5SLUmbV51ZR/j+ZlYYIW5hCNiiR+7NWiCoHyEKsvMWXlvq1W1RBJYbDW4a8vZVvn9FJM6upOe8+/tij7HMNqKhJ0e/4Qjdk2XBkHrz1Cxs1pk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700410; c=relaxed/simple; bh=Z2RtaHCbeerDdnWmNNCCKcU8AQItDLB0+3mXIUNdf0s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=EzVfceDNywYK1a3flIph36mLvYZmKVuVTd0uMeKAJ+1qd/I4s8YSb5WerySIPqZhd9hX1cCb8sZIkn4bCogzcyq5dvJ2r90CZZLm6AOwq7et5OTDG+V8ddD14/B4yIaLC9co+BsmGCoDtNG6yN7atKtxXqVTUGQA5bstHLf4mq4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyR671zzKHNNK; Fri, 14 Aug 2026 17:39:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 95E8440561; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S18; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 14/32] ext4: do not use data=ordered mode for inodes using buffered iomap path Date: Fri, 14 Aug 2026 17:33:13 +0800 Message-ID: <20260814093331.1703882-15-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S18 X-Coremail-Antispam: 1UD129KBjvJXoWxWFWktr1UWr1DCr1kJFWkJFb_yoW7Gryrpr W5K3s8JrZYva47ur1kuFW0qr40y3yUJr47Gry2gFsIgay5J3WIgFyrKa4SkFy5trsxGa4I qr48Ar97Wa1qyrJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ From: Zhang Yi The data=3Dordered mode introduces two fundamental conflicts with the iomap buffered write path, leading to potential deadlocks. 1) Lock ordering conflict In the iomap writeback path, each folio is processed sequentially: the folio lock is acquired first, followed by starting a transaction to create block mappings. In data=3Dordered mode, writeback triggered by the journal commit process may attempt to acquire a folio lock that is already held by iomap background writeback process. Meanwhile, iomap, under that same folio lock, may start a new transaction to map other blocks on this folio and wait for the currently committing transaction to finish, resulting in a deadlock. Trans N commit background writeback(via iomap) journal_submit_data_buffers() ext4_journal_submit_inode_data_buffers() iomap_writepages() iomap_writepages() folio_lock() folio_lock() -- wait iomap_writeback_folio() iomap_writeback_range() ext4_journal_start() start new transaction -- wait for trans N commit, DEADLOCK ext4_map_blocks() Currently, in the buffer_head writeback path, this is handled by starting the transaction before taking any folio locks for writeback. 2) Partial folio submission not supported When block size < folio size, a folio may contain both mapped and unmapped blocks. In data=3Dordered mode, a deadlock can occur if the journal waits (pure JI_WAIT_DATA) for such a folio to be written back while background writeback has already started on it (with the writeback flag set). The problem is that mapping the remaining delalloc blocks can deadlock because the writeback flag is not cleared until the entire folio is processed and committed. T0: Assume we have a folio contains four blocks, from front to back, they are A, B, C, D. The block B and C are holes, and the last block D is written in delalloc mode (the block is not allocated yet). T1: The background writeback process starts to write back data, set writeback flag on the folio, allocates block D, and adds it to transaction N's order list of jbd2 in pure JI_WAIT_DATA mode. T2: This folio completes the writeback and clears the writeback flag. T3: Before transaction N commit, we buffered write block A to C. T4: Transaction N commit and folio writeback are running concurrently. Trans N commit background writeback(via iomap) iomap_writeback_folio() folio_start_writeback() -- set writeback flag jbd2_journal_finish_inode_data_buffers() __filemap_fdatawait_range() -- wait writeback flag to clear iomap_writeback_range() ext4_journal_start() start new transaction -- wait for trans N commit, DEADLOCK ext4_map_block() (B, C) Currently, in the buffer_head writeback path, this is handled by: 1. Partial folio submission =E2=80=94 already-allocated buffers can be submitted first. The writeback flag is cleared after I/O completes, preventing block allocation while the writeback flag is set. 2. Allocation order =E2=80=94 the transaction is started first, then blo= cks are allocated, the writeback flag is set, and finally the allocated buffers submission begins. To support data=3Dordered mode, the iomap core would need two invasive changes: - Acquire the transaction handle before locking any folio for writeback. - Support partial folio submission. Both changes are complicated and risk performance regressions. Therefore, we must avoid using data=3Dordered mode when converting to the iomap path. Currently, data=3Dordered mode is used in three scenarios: - Append write - Post-EOF partial block truncate-up followed by append write - Online defragmentation We can address the first two without data=3Dordered mode: - For append write: always allocate unwritten blocks (i.e. always enable dioread_nolock), preserving the behavior of current extent-type inodes. - For post-EOF truncate-up + append write: postpone updating i_disksize until after the zeroed partial block has been written back. Online defragmentation does not yet support iomap; this can be resolved separately in the future. Signed-off-by: Zhang Yi Reviewed-by: Jan Kara --- fs/ext4/ext4_jbd2.h | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/fs/ext4/ext4_jbd2.h b/fs/ext4/ext4_jbd2.h index 2fbf48b3dfe2..be54e93bde0b 100644 --- a/fs/ext4/ext4_jbd2.h +++ b/fs/ext4/ext4_jbd2.h @@ -379,7 +379,12 @@ static inline int ext4_should_journal_data(struct inod= e *inode) =20 static inline int ext4_should_order_data(struct inode *inode) { - return ext4_inode_journal_mode(inode) & EXT4_INODE_ORDERED_DATA_MODE; + /* + * inodes using the iomap buffered I/O path do not use the + * data=3Dordered mode. + */ + return !ext4_inode_buffered_iomap(inode) && + (ext4_inode_journal_mode(inode) & EXT4_INODE_ORDERED_DATA_MODE); } =20 static inline int ext4_should_writeback_data(struct inode *inode) --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 102A73939A2; Fri, 14 Aug 2026 09:39:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700402; cv=none; b=iHcNYyP3M/3PVbQ3kx3BMh7ioXeTUVd8buI3zXhqPAeImeCUMVqe6yPTyXcGdMdxScfr7HZLyi6R6u3KvVKOKXlnzhe5ucxx4SRG8/mYmqQR8iZ3O4Dwlh0jgnwznOaM9+lnM8SKN5N4U940nKBMfhJaRd3ub+L2OksKPksRWkE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700402; c=relaxed/simple; bh=o6gSSrI/vUZpmeqK+zKp3THDtb2cLUjyn4BoQ3x0588=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=YWgo9E2kA7nsA+CvWjcromKhWkVwlR4kYgAm+zB0504V1AwR+gEDCIeON7bU6949K3Uqsydjc9VDyu+FHg9h2fRzdIIsliP5wRTO0FLfS3ofdGeKCuSz8RLZqi+eCzGgPfVp66fe7VvhO4zTA4gHY+9mhMEAhQYGy6i2f9BXBWc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyh3GX6zYQvN0; Fri, 14 Aug 2026 17:39:40 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id A8C374056D; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S19; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 15/32] ext4: implement buffered write path using iomap Date: Fri, 14 Aug 2026 17:33:14 +0800 Message-ID: <20260814093331.1703882-16-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S19 X-Coremail-Antispam: 1UD129KBjvJXoWfGFyDXFyxZr45Gry7ZFy8uFg_yoWkWr4rpF 98Kry5GFsFqr97ur4fKF4DZr1Fk3WxtrW7urW3Wrn8XF9FyrWIqF40gFyayF15trWxCr40 vF4Y9ry8Wr47CrDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26r4j6F4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Introduce two new iomap_ops instances for ext4 buffered writes: - ext4_iomap_buffered_da_write_ops: for delayed allocation mode, using ext4_da_map_blocks() to map delalloc extents. - ext4_iomap_buffered_write_ops: for non-delayed allocation mode, using ext4_map_blocks() to directly allocate blocks. Also add ext4_iomap_valid() for the iomap infrastructure to check extent validity. Key changes and considerations: - Unwritten extents for new blocks (dioread_nolock always on) Since data=3Dordered mode is not used to prevent stale data exposure in the non-delayed allocation path, new blocks are always allocated as unwritten extents. - Short write and write failure handling a. Delalloc path: On short write or failure, the stale delalloc range must be dropped and its space reservation released. Otherwise, a clean folio may cover leftover delalloc extents, causing inaccurate space reservation accounting. b. Non-delalloc path: No cleanup of allocated blocks is needed on short write. - Lock ordering reversal The folio lock and transaction start ordering is reversed compared to the buffer_head buffered write path. To handle this, the journal handle must be stopped in iomap_begin() callbacks. The lock ordering documentation in super.c has been updated accordingly. Signed-off-by: Zhang Yi Reviewed-by: Ojaswin Mujoo --- fs/ext4/ext4.h | 4 ++ fs/ext4/file.c | 20 +++++++- fs/ext4/inode.c | 130 ++++++++++++++++++++++++++++++++++++++++++++++-- fs/ext4/super.c | 10 ++-- 4 files changed, 156 insertions(+), 8 deletions(-) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 24ec205da2d7..98295ef7069a 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -3165,6 +3165,7 @@ int ext4_walk_page_buffers(handle_t *handle, int do_journal_get_write_access(handle_t *handle, struct inode *inode, struct buffer_head *bh); void ext4_set_inode_mapping_order(struct inode *inode); +int ext4_nonda_switch(struct super_block *sb); #define FALL_BACK_TO_NONDELALLOC 1 #define EXT4_WRITE_DATA_INLINE 2 =20 @@ -4047,6 +4048,9 @@ static inline void ext4_clear_io_unwritten_flag(ext4_= io_end_t *io_end) =20 extern const struct iomap_ops ext4_iomap_ops; extern const struct iomap_ops ext4_iomap_report_ops; +extern const struct iomap_ops ext4_iomap_buffered_write_ops; +extern const struct iomap_ops ext4_iomap_buffered_da_write_ops; +extern const struct iomap_write_ops ext4_iomap_write_ops; =20 int ext4_iomap_begin(struct inode *inode, loff_t offset, loff_t length, unsigned flags, struct iomap *iomap, struct iomap *srcmap); diff --git a/fs/ext4/file.c b/fs/ext4/file.c index 374b4bc25bd5..50d3c92709c8 100644 --- a/fs/ext4/file.c +++ b/fs/ext4/file.c @@ -330,6 +330,21 @@ static ssize_t ext4_write_checks(struct kiocb *iocb, s= truct iov_iter *from) return count; } =20 +static ssize_t ext4_iomap_buffered_write(struct kiocb *iocb, + struct iov_iter *from) +{ + struct inode *inode =3D file_inode(iocb->ki_filp); + const struct iomap_ops *iomap_ops; + + if (test_opt(inode->i_sb, DELALLOC) && !ext4_nonda_switch(inode->i_sb)) + iomap_ops =3D &ext4_iomap_buffered_da_write_ops; + else + iomap_ops =3D &ext4_iomap_buffered_write_ops; + + return iomap_file_buffered_write(iocb, from, iomap_ops, + &ext4_iomap_write_ops, NULL); +} + static ssize_t ext4_buffered_write_iter(struct kiocb *iocb, struct iov_iter *from) { @@ -351,7 +366,10 @@ static ssize_t ext4_buffered_write_iter(struct kiocb *= iocb, if (ret <=3D 0) goto out; =20 - ret =3D generic_perform_write(iocb, from); + if (ext4_inode_buffered_iomap(inode)) + ret =3D ext4_iomap_buffered_write(iocb, from); + else + ret =3D generic_perform_write(iocb, from); =20 out: inode_unlock(inode); diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 9098d9a5fc05..d831d1911a6f 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -3136,7 +3136,7 @@ static int ext4_dax_writepages(struct address_space *= mapping, return ret; } =20 -static int ext4_nonda_switch(struct super_block *sb) +int ext4_nonda_switch(struct super_block *sb) { s64 free_clusters, dirty_clusters; struct ext4_sb_info *sbi =3D EXT4_SB(sb); @@ -3508,6 +3508,15 @@ static bool ext4_inode_datasync_dirty(struct inode *= inode) return inode_state_read_once(inode) & I_DIRTY_DATASYNC; } =20 +static bool ext4_iomap_valid(struct inode *inode, const struct iomap *ioma= p) +{ + return iomap->validity_cookie =3D=3D READ_ONCE(EXT4_I(inode)->i_es_seq); +} + +const struct iomap_write_ops ext4_iomap_write_ops =3D { + .iomap_valid =3D ext4_iomap_valid, +}; + static void ext4_set_iomap(struct inode *inode, struct iomap *iomap, struct ext4_map_blocks *map, loff_t offset, loff_t length, unsigned int flags) @@ -3542,6 +3551,8 @@ static void ext4_set_iomap(struct inode *inode, struc= t iomap *iomap, !ext4_test_inode_flag(inode, EXT4_INODE_EXTENTS)) iomap->flags |=3D IOMAP_F_MERGED; =20 + iomap->validity_cookie =3D map->m_seq; + /* * Flags passed to ext4_map_blocks() for direct I/O writes can result * in m_flags having both EXT4_MAP_MAPPED and EXT4_MAP_UNWRITTEN bits @@ -3957,7 +3968,8 @@ const struct iomap_ops ext4_iomap_report_ops =3D { }; =20 static int ext4_iomap_map_blocks(struct inode *inode, loff_t offset, - loff_t length, struct ext4_map_blocks *map) + loff_t length, struct ext4_map_blocks *map, + int flags) { u8 blkbits =3D inode->i_blkbits; =20 @@ -3969,7 +3981,10 @@ static int ext4_iomap_map_blocks(struct inode *inode= , loff_t offset, map->m_len =3D min_t(loff_t, (offset + length - 1) >> blkbits, EXT4_MAX_LOGICAL_BLOCK) - map->m_lblk + 1; =20 - return ext4_map_blocks(NULL, inode, map, 0); + if (flags & EXT4_GET_BLOCKS_DELALLOC_RESERVE) + return ext4_da_map_blocks(inode, map); + + return ext4_map_blocks(NULL, inode, map, flags); } =20 static int ext4_iomap_buffered_read_begin(struct inode *inode, loff_t offs= et, @@ -3986,7 +4001,7 @@ static int ext4_iomap_buffered_read_begin(struct inod= e *inode, loff_t offset, if (WARN_ON_ONCE(ext4_has_inline_data(inode))) return -ERANGE; =20 - ret =3D ext4_iomap_map_blocks(inode, offset, length, &map); + ret =3D ext4_iomap_map_blocks(inode, offset, length, &map, 0); if (ret < 0) return ret; =20 @@ -3994,6 +4009,113 @@ static int ext4_iomap_buffered_read_begin(struct in= ode *inode, loff_t offset, return 0; } =20 +static int ext4_iomap_buffered_do_write_begin(struct inode *inode, + loff_t offset, loff_t length, unsigned int flags, + struct iomap *iomap, struct iomap *srcmap, bool delalloc) +{ + int ret, retries =3D 0; + struct ext4_map_blocks map; + int map_flags; + + ret =3D ext4_emergency_state(inode->i_sb); + if (unlikely(ret)) + return ret; + + /* Inline data and non-extent are not supported. */ + if (WARN_ON_ONCE(ext4_has_inline_data(inode))) + return -ERANGE; + if (WARN_ON_ONCE(!ext4_test_inode_flag(inode, EXT4_INODE_EXTENTS))) + return -EINVAL; + if (WARN_ON_ONCE(!(flags & IOMAP_WRITE))) + return -EINVAL; + + map_flags =3D delalloc ? EXT4_GET_BLOCKS_DELALLOC_RESERVE : + EXT4_GET_BLOCKS_CREATE_UNWRIT_EXT; +retry: + ret =3D ext4_iomap_map_blocks(inode, offset, length, &map, map_flags); + if (ret =3D=3D -ENOSPC && ext4_should_retry_alloc(inode->i_sb, &retries)) + goto retry; + if (ret < 0) + return ret; + + ext4_set_iomap(inode, iomap, &map, offset, length, flags); + return 0; +} + +static int ext4_iomap_buffered_write_begin(struct inode *inode, + loff_t offset, loff_t length, unsigned int flags, + struct iomap *iomap, struct iomap *srcmap) +{ + return ext4_iomap_buffered_do_write_begin(inode, offset, length, flags, + iomap, srcmap, false); +} + +static int ext4_iomap_buffered_da_write_begin(struct inode *inode, + loff_t offset, loff_t length, unsigned int flags, + struct iomap *iomap, struct iomap *srcmap) +{ + return ext4_iomap_buffered_do_write_begin(inode, offset, length, flags, + iomap, srcmap, true); +} + +/* + * On write failure, drop the stale delayed allocation range and release + * its reserved space for both start and end blocks. Otherwise, we may + * leave a range of delayed extents covered by a clean folio, which can + * result in inaccurate space reservation accounting. + */ +static void ext4_iomap_punch_delalloc(struct inode *inode, loff_t offset, + loff_t length, struct iomap *iomap) +{ + down_write(&EXT4_I(inode)->i_data_sem); + ext4_es_remove_extent(inode, offset >> inode->i_blkbits, + DIV_ROUND_UP_ULL(length, EXT4_BLOCK_SIZE(inode->i_sb))); + up_write(&EXT4_I(inode)->i_data_sem); +} + +static int ext4_iomap_buffered_da_write_end(struct inode *inode, loff_t of= fset, + loff_t length, ssize_t written, + unsigned int flags, + struct iomap *iomap) +{ + loff_t start_byte, end_byte; + + /* If we didn't reserve the blocks, we're not allowed to punch them. */ + if (iomap->type !=3D IOMAP_DELALLOC || !(iomap->flags & IOMAP_F_NEW)) + return 0; + + /* Nothing to do if we've written the entire delalloc extent */ + start_byte =3D iomap_last_written_block(inode, offset, written); + end_byte =3D round_up(offset + length, i_blocksize(inode)); + if (start_byte >=3D end_byte) + return 0; + + filemap_invalidate_lock(inode->i_mapping); + iomap_write_delalloc_release(inode, start_byte, end_byte, flags, + iomap, ext4_iomap_punch_delalloc); + filemap_invalidate_unlock(inode->i_mapping); + return 0; +} + +/* + * Since we always allocate unwritten extents, there is no need for + * iomap_end to clean up allocated blocks on a short write. + */ +static DEFINE_IOMAP_ITER_NEXT(ext4_iomap_buffered_write_next, + ext4_iomap_buffered_write_begin); + +const struct iomap_ops ext4_iomap_buffered_write_ops =3D { + .iomap_next =3D ext4_iomap_buffered_write_next, +}; + +static DEFINE_IOMAP_ITER_NEXT_END(ext4_iomap_buffered_da_write_next, + ext4_iomap_buffered_da_write_begin, + ext4_iomap_buffered_da_write_end); + +const struct iomap_ops ext4_iomap_buffered_da_write_ops =3D { + .iomap_next =3D ext4_iomap_buffered_da_write_next, +}; + static DEFINE_IOMAP_ITER_NEXT(ext4_iomap_buffered_read_next, ext4_iomap_buffered_read_begin); =20 diff --git a/fs/ext4/super.c b/fs/ext4/super.c index bca0dc87d0b7..30150094f2a5 100644 --- a/fs/ext4/super.c +++ b/fs/ext4/super.c @@ -104,9 +104,13 @@ static const struct fs_parameter_spec ext4_param_specs= []; * -> page lock -> i_data_sem (rw) * * buffered write path: - * sb_start_write -> i_mutex -> mmap_lock - * sb_start_write -> i_mutex -> transaction start -> page lock -> - * i_data_sem (rw) + * sb_start_write -> i_rwsem (w) -> mmap_lock + * - buffer_head path: + * sb_start_write -> i_rwsem (w) -> transaction start -> folio lock -> + * i_data_sem (rw) + * - iomap path: + * sb_start_write -> i_rwsem (w) -> transaction start -> i_data_sem (rw) + * sb_start_write -> i_rwsem (w) -> folio lock (not under an active hand= le) * * truncate: * sb_start_write -> i_mutex -> invalidate_lock (w) -> i_mmap_rwsem (w) -> --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 879B6435510; Fri, 14 Aug 2026 09:40:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700408; cv=none; b=EZkG2Me65Iojj0s027WzltnKkJrtLyqkyF8UxIWwsFtjA3Bh3rDYvvJ9TxNNlE3NRYgmk0RUAxidgJFNpiuAe8PhXvtvL1HtVA+5+M2Gep9RzpTWl4sCq/viViLMnwode8bUy92EjZZwarfLmVkYEDHTMlY7SpBPUKiphuOLKDc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700408; c=relaxed/simple; bh=0knHU9HZeRjVrxQFt8k3u902l9ts7OwbHwT9wLhw/qA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=bO+8N8VUvOoz+sVkcXMnKnAcx/IRxZn5kZscRuAB4Xf++t1UgaklWp/P/gg8P7IWAdJ7wgjAHE8IyeKd5jFg/uIC5Msbjf/jDIrpcRJH+tNfspjFwmFmsJiGt3OLXDaS+u81uJSLXr8XjEV7f2G0FrJ/UWa1QgzMx8bGzwgHjco= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyS0bqZzKHNNj; Fri, 14 Aug 2026 17:39:28 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id CAB424056D; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S20; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 16/32] ext4: implement writeback path using iomap Date: Fri, 14 Aug 2026 17:33:15 +0800 Message-ID: <20260814093331.1703882-17-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S20 X-Coremail-Antispam: 1UD129KBjvAXoWfGFy3KFyDXFWrWFWDGrW5Awb_yoW8GF17uo W3ta15Xr48GryYyayF9r1IyryUuan7Gw48Jr45Zrs2va47JF1Y93yfG3y3Wa47Xw4FkFWf AryxJa1rGr4xJF1rn29KB7ZKAUJUUUU5529EdanIXcx71UUUUU7v73VFW2AGmfu7bjvjm3 AaLaJ3UjIYCTnIWjp_UUUO37AC8VAFwI0_Wr0E3s1l1xkIjI8I6I8E6xAIw20EY4v20xva j40_Wr0E3s1l1IIY67AEw4v_Jr0_Jr4l82xGYIkIc2x26280x7IE14v26r126s0DM28Irc Ia0xkI8VCY1x0267AKxVW5JVCq3wA2ocxC64kIII0Yj41l84x0c7CEw4AK67xGY2AK021l 84ACjcxK6xIIjxv20xvE14v26F1j6w1UM28EF7xvwVC0I7IYx2IY6xkF7I0E14v26r4UJV WxJr1l84ACjcxK6I8E87Iv67AKxVWxJr0_GcWl84ACjcxK6I8E87Iv6xkF7I0E14v26rxl 6s0DM2AIxVAIcxkEcVAq07x20xvEncxIr21l5I8CrVACY4xI64kE6c02F40Ex7xfMcIj6x IIjxv20xvE14v26r1q6rW5McIj6I8E87Iv67AKxVWUJVW8JwAm72CE4IkC6x0Yz7v_Jr0_ Gr1lF7xvr2IYc2Ij64vIr41lF7I21c0EjII2zVCS5cI20VAGYxC7M4IIrI8v6xkF7I0E8c xan2IY04v7MxkF7I0En4kS14v26r4a6rW5MxAIw28IcxkI7VAKI48JMxC20s026xCaFVCj c4AY6r1j6r4UMI8I3I0E5I8CrVAFwI0_Jr0_Jr4lx2IqxVCjr7xvwVAFwI0_JrI_JrWlx4 CE17CEb7AF67AKxVW8ZVWrXwCIc40Y0x0EwIxGrwCI42IY6xIIjxv20xvE14v26F1j6w1U MIIF0xvE2Ix0cI8IcVCY1x0267AKxVW8Jr0_Cr1UMIIF0xvE42xK8VAvwI8IcIk0rVWUJV WUCwCI42IY6I8E87Iv67AKxVW8JVWxJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Add the iomap writeback path for ext4 buffered I/O. This introduces: - ext4_iomap_writepages(): the main writeback entry point. - ext4_writeback_ops: a new iomap_writeback_ops instance to handle block mapping and I/O submission. - A new end I/O work handler for converting unwritten extents, updating file size, and handling DATA_ERR_ABORT after I/O completion. Core implementation details: - ->writeback_range() callback Calls ext4_iomap_map_writeback_range() to map and allocate blocks using the enhanced ext4_map_blocks(). ext4_map_blocks() now starts its own transaction internally when it needs to allocate blocks. For performance, when a block range is not yet allocated, it allocates based on the writeback length and delalloc extent length, rather than allocating for a single folio at a time. The folio is then added to an iomap_ioend instance. - ->writeback_submit() callback Registers ext4_iomap_end_bio() as the end bio callback. This callback schedules a worker to handle: - Unwritten extent conversion. - i_disksize update after data is written back. - Journal abort on writeback I/O failure. Key changes and considerations: - Append write and unwritten extents Since data=3Dordered mode is not used to prevent stale data exposure during append writebacks, new blocks are always allocated as unwritten extents (i.e. always enable dioread_nolock), and i_disksize update is postponed until I/O completion. Additionally, the deadlock that the reserve handle was expected to resolve does not occur anymore. Therefore, the end I/O worker can start a normal journal handle instead of a reserve handle when converting unwritten extents. - Lock ordering The ->writeback_range() callback runs under the folio lock, requiring the journal handle to be started under that same lock. This reverses the order compared to the buffer_head writeback path. The lock ordering documentation in super.c has been updated accordingly. - Don't cache writes The iomap infrastructure sets the BIO_COMPLETE_IN_TASK flag when submitting I/O, so the ioend will be processed in task context. However, if a private defer worker is to be started, this flag must be cleared explicitly to avoid double deferral. In the future, all private defer work should be moved to the generic bio complete in task framework. Signed-off-by: Zhang Yi --- fs/ext4/ext4.h | 8 ++- fs/ext4/inode.c | 147 +++++++++++++++++++++++++++++++++++++++++++++- fs/ext4/page-io.c | 122 ++++++++++++++++++++++++++++++++++++++ fs/ext4/super.c | 5 +- 4 files changed, 278 insertions(+), 4 deletions(-) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 98295ef7069a..03fa90d2986f 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -1208,8 +1208,10 @@ struct ext4_inode_info { /* Lock protecting lists below */ spinlock_t i_completed_io_lock; /* - * Completed IOs that need unwritten extents handling and have - * transaction reserved + * Completed IOs that need unwritten extents handling and have a + * transaction reserved for the buffer_head writeback path, and + * also used by the iomap writeback path to queue ioends needing + * unwritten extents conversion, i_disksize update, etc. */ struct list_head i_rsv_conversion_list; struct work_struct i_rsv_conversion_work; @@ -3991,6 +3993,8 @@ void ext4_bio_write_folio(struct ext4_io_submit *io, = struct folio *page, size_t len); extern struct ext4_io_end_vec *ext4_alloc_io_end_vec(ext4_io_end_t *io_end= ); extern struct ext4_io_end_vec *ext4_last_io_end_vec(ext4_io_end_t *io_end); +extern void ext4_iomap_end_io(struct work_struct *work); +extern void ext4_iomap_end_bio(struct bio *bio); =20 /* mmp.c */ extern int ext4_multi_mount_protect(struct super_block *, ext4_fsblk_t); diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index d831d1911a6f..0b3e54e12b78 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -44,6 +44,7 @@ #include =20 #include "ext4_jbd2.h" +#include "ext4_extents.h" #include "xattr.h" #include "acl.h" #include "truncate.h" @@ -4134,10 +4135,154 @@ static void ext4_iomap_readahead(struct readahead_= control *rac) iomap_bio_readahead(rac, &ext4_iomap_buffered_read_ops); } =20 + +static int ext4_iomap_map_writeback_range(struct iomap_writepage_ctx *wpc, + loff_t offset, unsigned int dirty_len) +{ + struct inode *inode =3D wpc->inode; + struct super_block *sb =3D inode->i_sb; + struct journal_s *journal =3D EXT4_SB(sb)->s_journal; + struct ext4_map_blocks map; + unsigned int blkbits =3D inode->i_blkbits; + unsigned int index =3D offset >> blkbits; + unsigned int blk_end, blk_len; + int ret; + + ret =3D ext4_emergency_state(sb); + if (unlikely(ret)) + return ret; + + /* Check validity of the cached writeback mapping. */ + if (offset >=3D wpc->iomap.offset && + offset < wpc->iomap.offset + wpc->iomap.length && + ext4_iomap_valid(inode, &wpc->iomap)) + return 0; + + blk_len =3D dirty_len >> blkbits; + blk_end =3D min_t(unsigned int, (wpc->wbc->range_end >> blkbits), + (UINT_MAX - 1)); + if (blk_end > index + blk_len) + blk_len =3D blk_end - index + 1; + +retry: + map.m_lblk =3D index; + map.m_len =3D min_t(unsigned int, MAX_WRITEPAGES_EXTENT_LEN, blk_len); + ret =3D ext4_map_blocks(NULL, inode, &map, + EXT4_GET_BLOCKS_CREATE_UNWRIT_EXT | + EXT4_GET_BLOCKS_METADATA_NOFAIL | + EXT4_GET_BLOCKS_IO_SUBMIT | + EXT4_EX_NOCACHE); + if (ret < 0) { + if (ext4_emergency_state(sb)) + return ret; + + /* + * Retry transient ENOSPC errors, if + * ext4_count_free_blocks() is non-zero, a commit + * should free up blocks. + */ + if (ret =3D=3D -ENOSPC && journal && ext4_count_free_clusters(sb)) { + jbd2_journal_force_commit_nested(journal); + goto retry; + } + + ext4_msg(sb, KERN_CRIT, + "Delayed block allocation failed for inode %llu at logical offset %llu= with max blocks %u with error %d", + inode->i_ino, (unsigned long long)map.m_lblk, + (unsigned int)map.m_len, -ret); + ext4_msg(sb, KERN_CRIT, + "This should not happen!! Data will be lost\n"); + if (ret =3D=3D -ENOSPC) + ext4_print_free_blocks(inode); + return ret; + } + + ext4_set_iomap(inode, &wpc->iomap, &map, offset, dirty_len, 0); + return 0; +} + +static void ext4_iomap_discard_folio(struct folio *folio, loff_t pos) +{ + struct inode *inode =3D folio->mapping->host; + loff_t length =3D folio_pos(folio) + folio_size(folio) - pos; + + ext4_iomap_punch_delalloc(inode, pos, length, NULL); +} + +static ssize_t ext4_iomap_writeback_range(struct iomap_writepage_ctx *wpc, + struct folio *folio, u64 offset, + unsigned int len, u64 end_pos) +{ + ssize_t ret; + + ret =3D ext4_iomap_map_writeback_range(wpc, offset, len); + if (!ret) + ret =3D iomap_add_to_ioend(wpc, folio, offset, end_pos, len); + if (ret < 0) + ext4_iomap_discard_folio(folio, offset); + return ret; +} + +static int ext4_iomap_writeback_submit(struct iomap_writepage_ctx *wpc, + int error) +{ + struct iomap_ioend *ioend =3D wpc->wb_ctx; + struct ext4_inode_info *ei =3D EXT4_I(ioend->io_inode); + + /* + * After I/O completion, a worker needs to be scheduled when: + * 1) Unwritten extents require conversion. + * 2) The file size needs to be extended. + * 3) The journal needs to be aborted due to an I/O error. + */ + if ((ioend->io_flags & IOMAP_IOEND_UNWRITTEN) || + (ioend->io_offset + ioend->io_size > READ_ONCE(ei->i_disksize)) || + test_opt(ioend->io_inode->i_sb, DATA_ERR_ABORT)) + ioend->io_bio.bi_end_io =3D ext4_iomap_end_bio; + + /* + * ext4_iomap_end_bio() always defers endio processing, disable + * generic BIO in task to avoid double deferral since we will use + * a private defer endio handler in process context. + * + * TODO: Switch all defer handlers to the generic bio complete + * in task framework. + */ + if (ioend->io_bio.bi_end_io) + bio_clear_flag(&ioend->io_bio, BIO_COMPLETE_IN_TASK); + + return iomap_ioend_writeback_submit(wpc, error); +} + +static const struct iomap_writeback_ops ext4_writeback_ops =3D { + .writeback_range =3D ext4_iomap_writeback_range, + .writeback_submit =3D ext4_iomap_writeback_submit, +}; + static int ext4_iomap_writepages(struct address_space *mapping, struct writeback_control *wbc) { - return 0; + struct inode *inode =3D mapping->host; + struct super_block *sb =3D inode->i_sb; + long nr =3D wbc->nr_to_write; + int alloc_ctx, ret; + struct iomap_writepage_ctx wpc =3D { + .inode =3D inode, + .wbc =3D wbc, + .ops =3D &ext4_writeback_ops, + }; + + ret =3D ext4_emergency_state(sb); + if (unlikely(ret)) + return ret; + + alloc_ctx =3D ext4_writepages_down_read(sb); + trace_ext4_writepages(inode, wbc); + ret =3D iomap_writepages(&wpc); + trace_ext4_writepages_result(inode, wbc, ret, nr - wbc->nr_to_write); + ext4_writepages_up_read(sb, alloc_ctx); + + return ret; } =20 /* diff --git a/fs/ext4/page-io.c b/fs/ext4/page-io.c index 0236b6b9785a..2888e0057561 100644 --- a/fs/ext4/page-io.c +++ b/fs/ext4/page-io.c @@ -22,6 +22,7 @@ #include #include #include +#include #include #include #include @@ -547,3 +548,124 @@ void ext4_bio_write_folio(struct ext4_io_submit *io, = struct folio *folio, io_submit_add_bh(io, inode, folio, bh); } while ((bh =3D bh->b_this_page) !=3D head); } + +static int ext4_iomap_wb_update_disksize(handle_t *handle, struct inode *i= node, + loff_t end) +{ + loff_t new_disksize =3D end; + struct ext4_inode_info *ei =3D EXT4_I(inode); + int ret; + + /* + * Races with truncate are avoided by checking i_size under + * i_data_sem. + */ + down_write(&ei->i_data_sem); + new_disksize =3D min(new_disksize, i_size_read(inode)); + if (new_disksize > ei->i_disksize) + ei->i_disksize =3D new_disksize; + up_write(&ei->i_data_sem); + ret =3D ext4_mark_inode_dirty(handle, inode); + if (ret) + EXT4_ERROR_INODE_ERR(inode, -ret, "Failed to mark inode dirty"); + + return ret; +} + +static void ext4_iomap_finish_ioend(struct iomap_ioend *ioend) +{ + struct inode *inode =3D ioend->io_inode; + struct super_block *sb =3D inode->i_sb; + loff_t pos =3D ioend->io_offset; + size_t size =3D ioend->io_size; + loff_t end =3D pos + size; + handle_t *handle; + int credits; + int ret, err; + + ret =3D blk_status_to_errno(ioend->io_bio.bi_status); + if (unlikely(ret)) { + if (test_opt(sb, DATA_ERR_ABORT) && !ext4_emergency_state(sb)) + jbd2_journal_abort(EXT4_SB(sb)->s_journal, ret); + goto out; + } + + if (!(ioend->io_flags & IOMAP_IOEND_UNWRITTEN) && + end <=3D READ_ONCE(EXT4_I(inode)->i_disksize)) + goto out; + + /* + * We may need to convert one extent, update the i_disksize and + * dirty the inode. + */ + credits =3D ext4_chunk_trans_blocks(inode, + EXT4_MAX_BLOCKS(size, pos, inode->i_blkbits)); + handle =3D ext4_journal_start(inode, EXT4_HT_EXT_CONVERT, credits); + if (IS_ERR(handle)) { + ret =3D PTR_ERR(handle); + goto out_err; + } + + /* Update on-disk size after I/O is completed. */ + if (end > READ_ONCE(EXT4_I(inode)->i_disksize)) { + ret =3D ext4_iomap_wb_update_disksize(handle, inode, end); + if (ret) + goto out_journal; + } + + if (ioend->io_flags & IOMAP_IOEND_UNWRITTEN) + ret =3D ext4_convert_unwritten_extents(handle, inode, pos, + size, NULL); + +out_journal: + err =3D ext4_journal_stop(handle); + if (!ret) + ret =3D err; +out_err: + if (ret < 0 && !ext4_emergency_state(sb)) { + ext4_msg(sb, KERN_EMERG, + "failed to convert unwritten extents to written extents or update inod= e size -- potential data loss! (inode %llu, error %d)", + inode->i_ino, ret); + } +out: + iomap_finish_ioends(ioend, ret); +} + +/* + * Work on buffered iomap completed IO, to convert unwritten extents to + * mapped extents + */ +void ext4_iomap_end_io(struct work_struct *work) +{ + struct ext4_inode_info *ei =3D container_of(work, struct ext4_inode_info, + i_rsv_conversion_work); + struct iomap_ioend *ioend; + struct list_head ioend_list; + unsigned long flags; + + spin_lock_irqsave(&ei->i_completed_io_lock, flags); + list_replace_init(&ei->i_rsv_conversion_list, &ioend_list); + spin_unlock_irqrestore(&ei->i_completed_io_lock, flags); + + iomap_sort_ioends(&ioend_list); + while (!list_empty(&ioend_list)) { + ioend =3D list_entry(ioend_list.next, struct iomap_ioend, io_list); + list_del_init(&ioend->io_list); + iomap_ioend_try_merge(ioend, &ioend_list); + ext4_iomap_finish_ioend(ioend); + } +} + +void ext4_iomap_end_bio(struct bio *bio) +{ + struct iomap_ioend *ioend =3D iomap_ioend_from_bio(bio); + struct ext4_inode_info *ei =3D EXT4_I(ioend->io_inode); + unsigned long flags; + + spin_lock_irqsave(&ei->i_completed_io_lock, flags); + if (list_empty(&ei->i_rsv_conversion_list)) + queue_work(EXT4_SB(ioend->io_inode->i_sb)->rsv_conversion_wq, + &ei->i_rsv_conversion_work); + list_add_tail(&ioend->io_list, &ei->i_rsv_conversion_list); + spin_unlock_irqrestore(&ei->i_completed_io_lock, flags); +} diff --git a/fs/ext4/super.c b/fs/ext4/super.c index 30150094f2a5..6d2d323604f9 100644 --- a/fs/ext4/super.c +++ b/fs/ext4/super.c @@ -123,7 +123,10 @@ static const struct fs_parameter_spec ext4_param_specs= []; * sb_start_write -> i_mutex -> transaction start -> i_data_sem (rw) * * writepages: - * transaction start -> page lock(s) -> i_data_sem (rw) + * - buffer_head path: + * transaction start -> folio lock(s) -> i_data_sem (rw) + * - iomap path: + * folio lock -> transaction start -> i_data_sem (rw) */ =20 static const struct fs_context_operations ext4_context_ops =3D { --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9977642885E; Fri, 14 Aug 2026 09:40:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700409; cv=none; b=rFZkgbIFj6FfR7hdkwoVcewMU8JWAVfKK5ff7zmvydSJfwGV02a+h03Xf+Fp5M/qQFF3qVNrraeX99eyGSNoGyH1SVE4lMkG+Usakl8yZucwHu+lA/o3vGKsVNiwOtJHIznirSvi7EL4CQAkdwxhSHRLSt50JtDU88++jDRht/k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700409; c=relaxed/simple; bh=0hrCPBl4XycUSx+v4yYTiOWPbIJZxhHgDfqQbOgrXD0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Ly9SGN7s8PvRsU69nT0mSlbEp0jAfdULVt0jla7nXHa6vTTGU6S3Lme5fu2YNk1Q+/Buz7Td8OASOSEVTn0nFl4JHzH516miOIpxThlYpmTdrr1rx4v9G3wBWUfaPWf+HcX1uONXue64PdtlO0j3Bjw/ft9zswBEPD1pldPVKmc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyS169JzKHNNV; Fri, 14 Aug 2026 17:39:28 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id DDF1840561; Fri, 14 Aug 2026 17:39:48 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S21; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 17/32] ext4: implement mmap path using iomap Date: Fri, 14 Aug 2026 17:33:16 +0800 Message-ID: <20260814093331.1703882-18-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S21 X-Coremail-Antispam: 1UD129KBjvJXoWxWFWfJr17JryxtF43Kw1DGFg_yoWrJw4xpr 95K34rGrsIqwnF9rs7WFsxZr1rKayxKrW7WrW3WrnxZa42y340qa18KF1YvF1rJ3yfCr42 qF4Ykr18Wa47CrDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Introduce ext4_iomap_page_mkwrite() to implement the mmap iomap path for ext4. The heavy lifting is delegated to iomap_page_mkwrite(), which only requires ext4_iomap_buffered_write_ops and ext4_iomap_buffered_da_write_ops to allocate and map blocks. Note that the lock ordering between folio lock and transaction start in this path is reversed compared to the buffer_head buffered write path. The lock ordering documentation in super.c has been updated accordingly. Signed-off-by: Zhang Yi Reviewed-by: Ojaswin Mujoo Reviewed-by: Jan Kara --- fs/ext4/inode.c | 32 +++++++++++++++++++++++++++++++- fs/ext4/super.c | 8 ++++++-- 2 files changed, 37 insertions(+), 3 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 0b3e54e12b78..a05445625895 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -4027,7 +4027,7 @@ static int ext4_iomap_buffered_do_write_begin(struct = inode *inode, return -ERANGE; if (WARN_ON_ONCE(!ext4_test_inode_flag(inode, EXT4_INODE_EXTENTS))) return -EINVAL; - if (WARN_ON_ONCE(!(flags & IOMAP_WRITE))) + if (WARN_ON_ONCE(!(flags & (IOMAP_WRITE | IOMAP_FAULT)))) return -EINVAL; =20 map_flags =3D delalloc ? EXT4_GET_BLOCKS_DELALLOC_RESERVE : @@ -4085,6 +4085,14 @@ static int ext4_iomap_buffered_da_write_end(struct i= node *inode, loff_t offset, if (iomap->type !=3D IOMAP_DELALLOC || !(iomap->flags & IOMAP_F_NEW)) return 0; =20 + /* + * iomap_page_mkwrite() will never fail in a way that requires delalloc + * extents that it allocated to be revoked. Hence never try to release + * them here. + */ + if (flags & IOMAP_FAULT) + return 0; + /* Nothing to do if we've written the entire delalloc extent */ start_byte =3D iomap_last_written_block(inode, offset, written); end_byte =3D round_up(offset + length, i_blocksize(inode)); @@ -7296,6 +7304,23 @@ static int ext4_block_page_mkwrite(struct inode *ino= de, struct folio *folio, return ret; } =20 +static vm_fault_t ext4_iomap_page_mkwrite(struct vm_fault *vmf) +{ + struct inode *inode =3D file_inode(vmf->vma->vm_file); + const struct iomap_ops *iomap_ops; + + /* + * ext4_nonda_switch() could writeback this folio, so have to + * call it before lock folio. + */ + if (test_opt(inode->i_sb, DELALLOC) && !ext4_nonda_switch(inode->i_sb)) + iomap_ops =3D &ext4_iomap_buffered_da_write_ops; + else + iomap_ops =3D &ext4_iomap_buffered_write_ops; + + return iomap_page_mkwrite(vmf, iomap_ops, NULL); +} + vm_fault_t ext4_page_mkwrite(struct vm_fault *vmf) { struct vm_area_struct *vma =3D vmf->vma; @@ -7318,6 +7343,11 @@ vm_fault_t ext4_page_mkwrite(struct vm_fault *vmf) =20 filemap_invalidate_lock_shared(mapping); =20 + if (ext4_inode_buffered_iomap(inode)) { + ret =3D ext4_iomap_page_mkwrite(vmf); + goto out; + } + err =3D ext4_convert_inline_data(inode); if (err) goto out_ret; diff --git a/fs/ext4/super.c b/fs/ext4/super.c index 6d2d323604f9..1c2395aa1d53 100644 --- a/fs/ext4/super.c +++ b/fs/ext4/super.c @@ -100,8 +100,12 @@ static const struct fs_parameter_spec ext4_param_specs= []; * Lock ordering * * page fault path: - * mmap_lock -> sb_start_pagefault -> invalidate_lock (r) -> transaction s= tart - * -> page lock -> i_data_sem (rw) + * - buffer_head path: + * mmap_lock -> sb_start_pagefault -> invalidate_lock (r) -> + * transaction start -> folio lock -> i_data_sem (rw) + * - iomap path: + * mmap_lock -> sb_start_pagefault -> invalidate_lock (r) -> + * folio lock -> transaction start -> i_data_sem (rw) * * buffered write path: * sb_start_write -> i_rwsem (w) -> mmap_lock --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EA7E74195A7; Fri, 14 Aug 2026 09:39:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700404; cv=none; b=s8SfomuJSb9bVgFEGNhm6OOuXKtf3Qj52pfBlNYP9bSQFCZrGK4bz+CMVVskY7i/xG7QJqK37WLGJFX3vkxFGCBSNxULS8CZcvUK/UDXD6gKZJUWlZ2t4SsorbT1KTB8Sru5Ug88PE9yL/x16srVA4k3lXTZ7A3+5Ze5SaOlljQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700404; c=relaxed/simple; bh=rUnFLhM5VwcrvZHLo7VUt9NhobUQbGgmasMsOUkkes0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=H595ZdFq38riMjTs6ZX3cIVE1hBvMadABxmhLLFUSHnvlEAcAUvk1oZwTMYu+5rX1T/ZLEXr4/ltGRESlPFAeTcV9+XAEayLQ/EXFbluzU7r5cmtLrZfsiW26zBRGhotYS6aqQ+xsQ6GEcAMadfLkn7o6g/QLHRIA7gJAYRKOCU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=none smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyh5vvDzYQvTZ; Fri, 14 Aug 2026 17:39:40 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 0DC574056D; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S22; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 18/32] ext4: implement partial block zero range path using iomap Date: Fri, 14 Aug 2026 17:33:17 +0800 Message-ID: <20260814093331.1703882-19-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S22 X-Coremail-Antispam: 1UD129KBjvJXoW3GF47CF4DWw4kKF18ArW5Wrg_yoW3Gr1kpF ZxKry5GrsrX3sF9w4fJFnrXr1Ykw1ftFWUWry3GrnYvas8ZrWxKF1UGFWFvFyUJ3y7Gr12 qF4jy34xKF1UA3DanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Introduce a new iomap_ops instance, ext4_iomap_zero_ops, along with ext4_iomap_block_zero_range() to implement block zeroing via the iomap infrastructure for ext4. ext4_iomap_block_zero_range() calls iomap_zero_range() with ext4_iomap_zero_begin() as the callback. The callback locates the range and populates the iomap mapping. If the range is mapped, iomap_zero_iter() in the iomap core zeros the partial block directly. If the range is an unwritten extent within EOF, the callback collects a dirty folio batch via iomap_fill_dirty_folios() so that iomap_zero_iter() can zero those folios directly, bypassing a separate slow flush operation that would otherwise be needed to convert the unwritten extent. Note that ext4_iomap_zero_begin() can race with concurrent writeback: after it queries an unwritten extent, writeback may convert it to written and complete on the folio before iomap_fill_dirty_folios() scans the range. The empty batch then makes iomap_zero_iter() skip zeroing, leaving stale on-disk data. zero_range writeback ---------------------- ---------------------- ext4_block_zero_range() iomap_zero_range() iomap_iter() ext4_iomap_zero_begin() ext4_iomap_map_blocks() -> extent is UNWRITTEN ext4_convert_unwritten_extents_endio() -> extent is converted to WRITTEN -> folio is clean iomap_fill_dirty_folios() filemap_get_folios_dirty() -> folio is clean, not added to batch ext4_set_iomap() -> IOMAP_UNWRITTEN iomap_zero_iter() __iomap_get_folio() -> NULL (empty batch) iomap_iter_advance_full() <-- zeroing skipped [later read returns stale on-disk data] <-- CORRUPTION Therefore, we retry the extent lookup when iomap_fill_dirty_folios() adds nothing and the i_es_seq cookie captured at the first lookup has advanced, indicating the race actually occurred. Other important constraints: Zeroing out under an active journal handle can cause deadlock, as the lock/handle ordering is inconsistent with the iomap writeback path. Therefore, ext4_iomap_block_zero_range() must not be called under an active handle. In addition, for post-EOF zeroing, the caller cannot rely on data=3Dordered mode to persist the zeroed data before i_disksize is updated. Subsequent patches will address this by deferring i_disksize update to i_size until after the zeroed data has been written back. Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 110 ++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 110 insertions(+) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index a05445625895..d4b4153077bb 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -4106,6 +4106,67 @@ static int ext4_iomap_buffered_da_write_end(struct i= node *inode, loff_t offset, return 0; } =20 +static int ext4_iomap_zero_begin(struct inode *inode, + loff_t offset, loff_t length, unsigned int flags, + struct iomap *iomap, struct iomap *srcmap) +{ + struct iomap_iter *iter =3D container_of(iomap, struct iomap_iter, iomap); + struct ext4_map_blocks map; + u8 blkbits =3D inode->i_blkbits; + unsigned int iomap_flags; + int ret; + + ret =3D ext4_emergency_state(inode->i_sb); + if (unlikely(ret)) + return ret; + + if (WARN_ON_ONCE(!(flags & IOMAP_ZERO))) + return -EINVAL; + +again: + ret =3D ext4_iomap_map_blocks(inode, offset, length, &map, 0); + if (ret < 0) + return ret; + + /* + * Look up dirty folios for unwritten mappings within EOF. Providing + * this bypasses the flush iomap uses to trigger extent conversion + * when unwritten mappings have dirty pagecache in need of zeroing. + */ + iomap_flags =3D 0; + if (map.m_flags & EXT4_MAP_UNWRITTEN) { + loff_t start =3D ((loff_t)map.m_lblk) << blkbits; + loff_t end =3D ((loff_t)map.m_lblk + map.m_len) << blkbits; + unsigned int count; + + count =3D iomap_fill_dirty_folios(iter, &start, end, + &iomap_flags); + if ((start >> blkbits) < map.m_lblk + map.m_len) + map.m_len =3D (start >> blkbits) - map.m_lblk; + + /* + * This can be raced by a concurrent writeback that cleans + * the folio and converts the unwritten extent to written. + * Recheck the mapping after a folio lock round in + * iomap_fill_dirty_folios(). + */ + if (count =3D=3D 0 && + map.m_seq !=3D READ_ONCE(EXT4_I(inode)->i_es_seq)) + goto again; + } + + ext4_set_iomap(inode, iomap, &map, offset, length, flags); + iomap->flags |=3D iomap_flags; + + return 0; +} + +static DEFINE_IOMAP_ITER_NEXT(ext4_iomap_zero_next, ext4_iomap_zero_begin); + +static const struct iomap_ops ext4_iomap_zero_ops =3D { + .iomap_next =3D ext4_iomap_zero_next, +}; + /* * Since we always allocate unwritten extents, there is no need for * iomap_end to clean up allocated blocks on a short write. @@ -4569,6 +4630,48 @@ static int ext4_block_journalled_zero_range(struct i= node *inode, loff_t from, return err; } =20 +static int ext4_block_iomap_zero_range(struct inode *inode, loff_t from, + loff_t length, bool *did_zero, + bool *zero_written) +{ + int ret; + + /* + * Zeroing out under an active handle can cause deadlock since + * the order of acquiring the folio lock and starting a handle is + * inconsistent with the iomap writeback procedure. + */ + if (WARN_ON_ONCE(ext4_handle_valid(journal_current_handle()))) + return -EINVAL; + + /* The zeroing scope should not extend across a block. */ + if (WARN_ON_ONCE((from >> inode->i_blkbits) !=3D + ((from + length - 1) >> inode->i_blkbits))) + return -EINVAL; + + if (!(EXT4_SB(inode->i_sb)->s_mount_state & EXT4_ORPHAN_FS) && + !(inode_state_read_once(inode) & (I_NEW | I_FREEING))) + WARN_ON_ONCE(!inode_is_locked(inode) && + !rwsem_is_locked(&inode->i_mapping->invalidate_lock)); + + ret =3D iomap_zero_range(inode, from, length, did_zero, + &ext4_iomap_zero_ops, &ext4_iomap_write_ops, + NULL); + if (ret) + return ret; + + /* + * TODO: The iomap does not distinguish between different types + * of zeroing operations. So we always set zero_written whenever + * zeroing is performed, which may cause unnecessary folio + * flushing when zeroing occurs on delayed-allocated blocks. + */ + if (did_zero && zero_written) + *zero_written =3D *did_zero; + + return 0; +} + /* * Zeros out a mapping of length 'length' starting from file offset * 'from'. The range to be zero'd must be contained with in one block. @@ -4595,6 +4698,9 @@ static int ext4_block_zero_range(struct inode *inode, } else if (ext4_should_journal_data(inode)) { return ext4_block_journalled_zero_range(inode, from, length, did_zero); + } else if (ext4_inode_buffered_iomap(inode)) { + return ext4_block_iomap_zero_range(inode, from, length, + did_zero, zero_written); } return ext4_block_do_zero_range(inode, from, length, did_zero, zero_written); @@ -4649,6 +4755,10 @@ int ext4_block_zero_eof(struct inode *inode, loff_t = from, loff_t end) * block, the zeroed data lies beyond the existing on-disk data. It * will be written out before i_disksize is later extended past * i_size, so no stale data can be exposed. + * + * TODO: In the iomap path, handle this by tracking the ordered range + * and updating i_disksize to i_size after the zeroed data has been + * written back. */ if (ext4_should_order_data(inode) && did_zero && zero_written && !IS_DAX(inode) && --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4681D400DE8; Fri, 14 Aug 2026 09:39:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700402; cv=none; b=L33V2itXMpsRZCra2gRS15Jn5xx29ny4ev8H5Bnk+RNQWrvp57vbKIYzR9Mm6uLM/CJyMGUWJPQowkU+SpN+x5SkouZX0OYJzs5ztjNdSWqRd8VkEhDy4cl13OFfNrmoLj+P8LRzAd21qB+xObLO+UkoG6YJVz9EXzwvmelVaeA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700402; c=relaxed/simple; bh=J7hB7ZVA/xufntpOwc1H0+mMTzrq/O1ftTvA6QzmsA8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=LQfqdwV5tAndqJUXRV9IsyTBvIT+/Gig/nfjF6SHNsxfxXDEuAa95+Jx2urwp5XAL6tPi2dvoBlFvRYYppeMUD2/T2OW/SdTeSeXfNzW5cTjn63tD1JrfP5LJcJeU8wVof7f1I6VG/Ilsmzz5gA7fAGcDzmSSioMNL1HrbeE9v0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.177]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyh6yFGzYQvY4; Fri, 14 Aug 2026 17:39:40 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 2F02A40592; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S23; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 19/32] ext4: tolerate unexpected holes in ext4_convert_unwritten_extents() Date: Fri, 14 Aug 2026 17:33:18 +0800 Message-ID: <20260814093331.1703882-20-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S23 X-Coremail-Antispam: 1UD129KBjvJXoWxWw4fuFykAF18Kr17tw1UZFb_yoW5Cr4fpF ZIkr15Kr1jga9I9rsrtFWxXr1Fk3WxGF47ZrWfGF13XFyDZr42g3WUKa4FyFy5tFWxJrya qrW8Jry8XF15ZaDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Because the iomap infrastructure does not always create an ifs to manage sub-folio state when folio size is larger than blocksize, invalidating a partial dirty folio during punch hole may fail to clear the dirty state of the affected range. As a result, writeback of that folio may observe a hole. At writeback submit time, ext4_map_blocks() already handles this case and will not allocate blocks. However, when punch hole races with writeback, the following scenario can cause I/O completion to encounter a hole. punch hole writeback ---------- --------- ext4_punch_hole() ext4_truncate_page_cache_block_range() iomap_invalidate_folio() [partial folio] iomap_clear_range_dirty() -- no ifs, sub-block dirty bits NOT cleared ext4_iomap_writepages() iomap_writepages() ext4_iomap_writeback_submit() ext4_iomap_map_writeback_range() ext4_map_blocks(IO_SUBMIT) -> extent exists, not a hole submit_io() -> bio in flight down_write(&i_data_sem) ext4_es_remove_extent() ext4_ext_remove_space() -> extent removed, hole inserted up_write(&i_data_sem) [bio completes] ext4_iomap_finish_ioend() ext4_convert_unwritten_extents() ext4_map_blocks(IO_CONVERT_EXT) -> returns 0 (hole found) Therefore, in ext4_convert_unwritten_extents() we need to be careful about the case where ext4_map_blocks() returns 0. Instead of triggering a warning, we should ignore the hole and continue with the subsequent conversion. Link: https://lore.kernel.org/linux-ext4/a638a8fb-c184-4069-ae33-379ec12cd5= 14@huaweicloud.com/ Signed-off-by: Zhang Yi --- fs/ext4/extents.c | 20 +++++++++++--------- 1 file changed, 11 insertions(+), 9 deletions(-) diff --git a/fs/ext4/extents.c b/fs/ext4/extents.c index 0d62d9312284..5a06259a9b5d 100644 --- a/fs/ext4/extents.c +++ b/fs/ext4/extents.c @@ -5166,19 +5166,21 @@ int ext4_convert_unwritten_extents(handle_t *handle= , struct inode *inode, ret =3D ext4_map_blocks(handle, inode, &map, EXT4_GET_BLOCKS_IO_CONVERT_EXT | EXT4_EX_NOCACHE); - if (ret <=3D 0) { - /* - * If the ret is zero, an unexpected hole may cause - * conversion to fail. To avoid data loss during I/O - * end conversion, skip the hole and continue - * converting subsequent blocks. - */ + /* + * A return value of zero means an unexpected hole was found. + * This can happen when writeback races with a concurrent + * punch hole in the iomap path. Because iomap may not create + * ifs for folios larger than block size, the dirty bit can + * be set again after punching. If writeback happens between + * partial folio invalidation and extent removal, a hole is + * observed at I/O completion. + */ + if (ret < 0) ext4_warning(inode->i_sb, "inode #%llu: block %u: len %u: ext4_map_blocks returned %d", inode->i_ino, map.m_lblk, map.m_len, ret); - } else { + else if (ret > 0) conv_blocks +=3D map.m_len; - } =20 ret2 =3D ext4_mark_inode_dirty(handle, inode); if (credits) { --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5CDBB43CED7; Fri, 14 Aug 2026 09:40:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700411; cv=none; b=WETajSWPqzR3WyESaECUhqDcSf1I6DQ3roXUWtJqFZ5uI9DJ+sbm644qE3k3pzUSaoVmC2Mh0rVn+FRySoChUNPJubEabq6nE5fechRIuXGpbOYRChRqr8p7jlIE6MsoaasmylCrYMjte1vbWvGTYhRYaGttnALsFgyNJ2uAqdM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700411; c=relaxed/simple; bh=zDov8zxYYt9fFGhkHUrlOlPvimkURQvwpLQYfpMmxlQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=QeYQkRezpS9OWMUEAGzvfSyBeei/347irv53ejEpXReOsUh+lwIulzlzqewKUVWVtehiUyLZU7muA/jLIxqOjF0+JO7jzzkerP5ErHXaOpQDbZ5ZjoKwYlC5R1xx/dx0XrJXBS3uc2zS4RCfJtLZKBWJHp7czDRC7kKF93Wid7A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=none smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.198]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyS3vljzKHNNk; Fri, 14 Aug 2026 17:39:28 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 3F99B40C40; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S24; Fri, 14 Aug 2026 17:39:48 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 20/32] ext4: add block mapping tracepoints for iomap buffered I/O path Date: Fri, 14 Aug 2026 17:33:19 +0800 Message-ID: <20260814093331.1703882-21-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S24 X-Coremail-Antispam: 1UD129KBjvJXoWxurW7ZF1rKr18Xr47GFy8Krg_yoWrGFy7pa 4vyFy5KFs3XrsF9w4fWrW3Xr1Fva1xKr4UGry3Wry5ZFWxtr12gF4UGFyjyFy5Jw4jkryf WF4Ykry8G3WUurDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Add tracepoints for iomap buffered read, write, partial block zeroing, and writeback operations to help debug the iomap buffered I/O path. Signed-off-by: Zhang Yi Reviewed-by: Ojaswin Mujoo Reviewed-by: Jan Kara --- fs/ext4/inode.c | 6 +++++ include/trace/events/ext4.h | 45 +++++++++++++++++++++++++++++++++++++ 2 files changed, 51 insertions(+) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index d4b4153077bb..a2f060ae6cbc 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -4006,6 +4006,8 @@ static int ext4_iomap_buffered_read_begin(struct inod= e *inode, loff_t offset, if (ret < 0) return ret; =20 + trace_ext4_iomap_buffered_read_begin(inode, &map, offset, length, + flags); ext4_set_iomap(inode, iomap, &map, offset, length, flags); return 0; } @@ -4039,6 +4041,8 @@ static int ext4_iomap_buffered_do_write_begin(struct = inode *inode, if (ret < 0) return ret; =20 + trace_ext4_iomap_buffered_write_begin(inode, &map, offset, length, + flags); ext4_set_iomap(inode, iomap, &map, offset, length, flags); return 0; } @@ -4155,6 +4159,7 @@ static int ext4_iomap_zero_begin(struct inode *inode, goto again; } =20 + trace_ext4_iomap_zero_begin(inode, &map, offset, length, flags); ext4_set_iomap(inode, iomap, &map, offset, length, flags); iomap->flags |=3D iomap_flags; =20 @@ -4266,6 +4271,7 @@ static int ext4_iomap_map_writeback_range(struct ioma= p_writepage_ctx *wpc, return ret; } =20 + trace_ext4_iomap_map_writeback_range(inode, &map, offset, dirty_len, 0); ext4_set_iomap(inode, &wpc->iomap, &map, offset, dirty_len, 0); return 0; } diff --git a/include/trace/events/ext4.h b/include/trace/events/ext4.h index 7028a28316fa..69596a216dcb 100644 --- a/include/trace/events/ext4.h +++ b/include/trace/events/ext4.h @@ -3157,6 +3157,51 @@ TRACE_EVENT(ext4_move_extent_exit, __entry->ret) ); =20 +DECLARE_EVENT_CLASS(ext4_set_iomap_class, + TP_PROTO(struct inode *inode, struct ext4_map_blocks *map, + loff_t offset, loff_t length, unsigned int flags), + TP_ARGS(inode, map, offset, length, flags), + TP_STRUCT__entry( + __field(dev_t, dev) + __field(u64, ino) + __field(ext4_lblk_t, m_lblk) + __field(unsigned int, m_len) + __field(unsigned int, m_flags) + __field(u64, m_seq) + __field(loff_t, offset) + __field(loff_t, length) + __field(unsigned int, iomap_flags) + ), + TP_fast_assign( + __entry->dev =3D inode->i_sb->s_dev; + __entry->ino =3D inode->i_ino; + __entry->m_lblk =3D map->m_lblk; + __entry->m_len =3D map->m_len; + __entry->m_flags =3D map->m_flags; + __entry->m_seq =3D map->m_seq; + __entry->offset =3D offset; + __entry->length =3D length; + __entry->iomap_flags =3D flags; + + ), + TP_printk("dev %d:%d ino %llu m_lblk %u m_len %u m_flags %s m_seq %llu or= ig_off 0x%llx orig_len 0x%llx iomap_flags 0x%x", + MAJOR(__entry->dev), MINOR(__entry->dev), + __entry->ino, __entry->m_lblk, __entry->m_len, + show_mflags(__entry->m_flags), __entry->m_seq, + __entry->offset, __entry->length, __entry->iomap_flags) +) + +#define DEFINE_SET_IOMAP_EVENT(name) \ +DEFINE_EVENT(ext4_set_iomap_class, name, \ + TP_PROTO(struct inode *inode, struct ext4_map_blocks *map, \ + loff_t offset, loff_t length, unsigned int flags), \ + TP_ARGS(inode, map, offset, length, flags)) + +DEFINE_SET_IOMAP_EVENT(ext4_iomap_buffered_read_begin); +DEFINE_SET_IOMAP_EVENT(ext4_iomap_buffered_write_begin); +DEFINE_SET_IOMAP_EVENT(ext4_iomap_map_writeback_range); +DEFINE_SET_IOMAP_EVENT(ext4_iomap_zero_begin); + #endif /* _TRACE_EXT4_H */ =20 /* This part must be outside protection */ --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 136444570CA; Fri, 14 Aug 2026 09:40:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700414; cv=none; b=M37KOw39uQU6BtDYu0xfGBZeO1KFc+Pa7MaU4Jb8Ws0Ab6xVcUW3VxUsMz95TO9m3BiFo7WwieBvZ+j1GPSR/nq2lDEp6dLiMjZoDS4DjpTT1twEStoJrQ8pKiRW/6YAVTFwv/gYQuyjUEFHDQkgx0metC6j2BQyY07iGO8e5q8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700414; c=relaxed/simple; bh=Kl4yPHS0u6VXq5XBNdH4JzA0QDTF9XxhGxEkTkJHKXk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=A/3/4ABahmvEoR3FOjS+/E9BZvw+VJ4lfqsGmCroipAveju47qAudvnKQ+6/Kk67VrnvaHVhvA9oU1eTMTDo6zLuPiY3gO8GWF/O9cWnaezF4p2DxGHiFF8uzY3LM09SjRYAoMUFFgqLMQuXj60QkDrEJavmCxZS7pPpbCtHmNM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=none smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyS47PzzKHNNw; Fri, 14 Aug 2026 17:39:28 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 4ED2F4056D; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S25; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 21/32] ext4: disable online defrag when inode using iomap buffered I/O path Date: Fri, 14 Aug 2026 17:33:20 +0800 Message-ID: <20260814093331.1703882-22-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S25 X-Coremail-Antispam: 1UD129KBjvJXoW7CFykWw4UJF4fZrWfAr47Arb_yoW8Gr4xp3 Zakw1rGrW8Xa429a9YqF12qw4jga1xGrW2gFWSgr47GFWqyF9Ygr1UKa15Aa45trWUJ34F qF12kryUWw1UA3DanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Online defragmentation does not currently support inodes using the iomap buffered I/O path. The existing implementation relies on buffer_head for sub-folio block management and data=3Dordered mode for data consistency, both of which are incompatible with the iomap path. Signed-off-by: Zhang Yi Reviewed-by: Ojaswin Mujoo Reviewed-by: Jan Kara --- fs/ext4/move_extent.c | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/fs/ext4/move_extent.c b/fs/ext4/move_extent.c index 3329b7ad5dbd..948ef3f44df5 100644 --- a/fs/ext4/move_extent.c +++ b/fs/ext4/move_extent.c @@ -476,6 +476,17 @@ static int mext_check_validity(struct inode *orig_inod= e, return -EOPNOTSUPP; } =20 + /* + * TODO: support online defrag for inodes that use the buffered + * I/O iomap path. + */ + if (ext4_inode_buffered_iomap(orig_inode) || + ext4_inode_buffered_iomap(donor_inode)) { + ext4_msg(sb, KERN_ERR, + "Online defrag not supported for inode with iomap buffered IO path"); + return -EOPNOTSUPP; + } + if (donor_inode->i_mode & (S_ISUID|S_ISGID)) { ext4_debug("ext4 move extent: suid or sgid is set to donor file [ino:ori= g %llu, donor %llu]\n", orig_inode->i_ino, donor_inode->i_ino); --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 88AB94582E2; Fri, 14 Aug 2026 09:40:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700411; cv=none; b=JLgz4V9kR/KDJkKSvpvChlRqwAfgAiECnzerCCH1kLm5XqLsqGihCRdJVduQQ7Ty869NrjSQAeXdmqTLY7PZCx0KxItRc2rP1MrpuN5v7yEjTFT6JUSOUG1Zgz3+vYKDCSh6iEPAbH4wzyRWnKhws/ZdK9ynGQhiPekkelvFq/Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700411; c=relaxed/simple; bh=Td+VEBD6mkH7oIvQ/y/Y0SM9vXaD/MjwygvWVgYJWEU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=tCoJus7bE24qoEBgo/sAiPmkzNTOJtDXNUBacvdwO0giMoPO5sVYepxNEIGvKovgGkjxGbwaNA7wzapk7zRWk8ZQc9QCLgW3D/twHNdNBKxAGQX3JCOjnzXHbkXG8WtbdUSToO3nsoHPOpFP81NI3sNPpp+HuSGhscePAGQSD4U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=none smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyS4tkQzKHNNl; Fri, 14 Aug 2026 17:39:28 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 6923A4056B; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S26; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 22/32] ext4: add EXT4_STATE_DISKSIZE_GROW_PENDING state bit and helpers Date: Fri, 14 Aug 2026 17:33:21 +0800 Message-ID: <20260814093331.1703882-23-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S26 X-Coremail-Antispam: 1UD129KBjvJXoWxJF4kCr4fJr48ur1Uuw4fAFb_yoWrKr4Dpr srKry5Gr15Xr9F9w4SqFy7Zr1Yka1rJw48GFy3Gr4qqFW5WrWxKFn2yFy3ua40yrs5Aw42 qFs8KrZrCw1UCrJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Inodes using the iomap buffered I/O path do not use data=3Dordered mode, so the zeroed EOF block that straddles i_disksize needs explicit tracking to ensure it is written back before i_disksize is advanced. Add the EXT4_STATE_DISKSIZE_GROW_PENDING inode state bit and three helpers: ext4_iomap_clear_disksize_pending() to atomically clear the bit and wake waiters, ext4_iomap_wait_disksize_pending() to block until the bit is cleared, and ext4_iomap_get_disksize_pending_range() to compute the pending range from i_disksize. These will be used by subsequent patches to serialize i_disksize updates with the writeback of the zeroed EOF block. Suggested-by: Jan Kara Signed-off-by: Zhang Yi --- fs/ext4/ext4.h | 7 ++++++ fs/ext4/inode.c | 64 +++++++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 71 insertions(+) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 03fa90d2986f..1c3d736fb700 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -2052,6 +2052,9 @@ enum { EXT4_STATE_ORPHAN_FILE, /* Inode orphaned in orphan file */ EXT4_STATE_FC_REQUEUE, /* Inode modified during fast commit */ EXT4_STATE_BUFFERED_IOMAP, /* Inode use iomap for buffered IO */ + EXT4_STATE_DISKSIZE_GROW_PENDING, + /* Has zeroed EOF block straddles + * i_disksize awaiting writeback */ }; =20 #define EXT4_INODE_BIT_FNS(name, field, offset) \ @@ -3219,6 +3222,10 @@ extern int ext4_chunk_trans_blocks(struct inode *, i= nt nrblocks); extern int ext4_chunk_trans_extent(struct inode *inode, int nrblocks); extern int ext4_meta_trans_blocks(struct inode *inode, int lblocks, int pextents, int alloc_extents); +void ext4_iomap_clear_disksize_pending(struct inode *inode); +void ext4_iomap_wait_disksize_pending(struct inode *inode); +unsigned int ext4_iomap_get_disksize_pending_range(struct inode *inode, + loff_t *start); extern int ext4_block_zero_eof(struct inode *inode, loff_t from, loff_t en= d); =20 #define EXT4_PARTIAL_ZERO_START 0x1 diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index a2f060ae6cbc..e4a4396eaf87 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -126,6 +126,70 @@ void ext4_inode_csum_set(struct inode *inode, struct e= xt4_inode *raw, raw->i_checksum_hi =3D cpu_to_le16(csum >> 16); } =20 +/* + * Clear the disksize-grow-pending state and wake up all waiters. + * Called when the pending zeroed EOF block which straddles i_disksize + * has completed writeback or its folio is discarded. + */ +void ext4_iomap_clear_disksize_pending(struct inode *inode) +{ + ext4_clear_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING); + /* + * Make sure clearing of EXT4_STATE_DISKSIZE_GROW_PENDING is + * visible before we send the wakeup. Pairs with the implicit + * barrier in prepare_to_wait() inside wait_on_bit() in + * ext4_iomap_wait_disksize_pending(). + */ + smp_mb(); + wake_up_bit(ext4_inode_state_wait_word(inode), + ext4_inode_state_wait_bit(EXT4_STATE_DISKSIZE_GROW_PENDING)); +} + +/* + * Wait for the disksize-grow-pending zeroed EOF block which straddles + * i_disksize to be written back or cleared. + */ +void ext4_iomap_wait_disksize_pending(struct inode *inode) +{ + wait_on_bit(ext4_inode_state_wait_word(inode), + ext4_inode_state_wait_bit(EXT4_STATE_DISKSIZE_GROW_PENDING), + TASK_UNINTERRUPTIBLE); +} + +/* + * Get the range of the disksize-grow-pending zeroed EOF block range + * which straddles i_disksize if the EXT4_STATE_DISKSIZE_GROW_PENDING + * bit is set. + * + * Return the pending range, or zero if the BIT has already been cleared. + */ +unsigned int ext4_iomap_get_disksize_pending_range(struct inode *inode, + loff_t *start) +{ + unsigned int blocksize =3D i_blocksize(inode); + loff_t disksize; + + if (!ext4_test_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING)) + return 0; + + /* + * The pending bit should be set only when i_disksize is not + * block-size aligned. While set, i_disksize must not be advanced, + * and the bit must be cleared when i_disksize is shrunk. + */ + down_read(&EXT4_I(inode)->i_data_sem); + disksize =3D READ_ONCE(EXT4_I(inode)->i_disksize); + if (!ext4_test_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING) || + WARN_ON_ONCE(IS_ALIGNED(disksize, blocksize))) { + up_read(&EXT4_I(inode)->i_data_sem); + return 0; + } + + up_read(&EXT4_I(inode)->i_data_sem); + *start =3D disksize; + return blocksize - (disksize & (blocksize - 1)); +} + static inline int ext4_begin_ordered_truncate(struct inode *inode, loff_t new_size) { --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8E79539D6E8; Fri, 14 Aug 2026 09:39:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700404; cv=none; b=aZMK4I8E/0Dan2M/mySiu4/MYhtO7GtEZp0F3/rLbxaYK9aYj497A2N3VO1p9i1U2eLVx68ZULx1mo52bO0T9uyGtum1YGsBkjmAH28m408al7DmfYK6XifC2xiUZxehdy+8AfO4AdFf0JCJ5MLQD6jNbct5QbcuSbQG810FveQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700404; c=relaxed/simple; bh=TKn05nvXB9c1TgJn/Y1KlLuJ3nq7M2VOuAIJ18kvNtU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=V/e/T57GTXQ/D4qcRmVoC0wUTjAHzNT8aYSYmtnI3Hnj6BTbRV0WqDofuJE6+Fj/fzOSn//2u26ILRFlJm8ArwFpsaV80CsgedR4ctOztXCx5D2YyWz4hQcJRaHFFBV3az2YMyksm/1VRpsCKakWT34uFX2j2vr+lbMXYODLoOs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.177]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyj2L1yzYQvYY; Fri, 14 Aug 2026 17:39:41 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 7C93540594; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S27; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 23/32] ext4: submit and wait for pending disksize-grow I/O on writeback Date: Fri, 14 Aug 2026 17:33:22 +0800 Message-ID: <20260814093331.1703882-24-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S27 X-Coremail-Antispam: 1UD129KBjvJXoWxKryUZFy8KF17JrykKr1xKrg_yoWxKrWxp3 s8Kry5Gr4DXr9F9rsaqFW7Zr1Yka1rtr4UJry3Wa98Zry5ury7KFW0gF1a9F1vk393J39F qF4vkrW8Cw17ArJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi When the current writeback pass begins beyond the disksize-grow-pending zeroed EOF block, the ioend worker would otherwise have to wait for the pending EOF block to complete before it can advance i_disksize. Otherwise the old EOF block could be exposed as stale data once i_disksize advances past it. Therefore, introduce the ioend mechanism for the pending range, tag ioends that cover the pending zeroed EOF block which straddles i_disksize with EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO in ext4_iomap_writeback_submit(), and clear the bit and wake up all waiters in ext4_iomap_end_bio() when such an ioend completes. In order to avoid the ioend that passes the pending range waiting for a long time, proactively submit the pending range first in ext4_iomap_writepages() so it completes in parallel with the rest of the writeback. Signed-off-by: Zhang Yi --- fs/ext4/ext4.h | 6 ++++++ fs/ext4/inode.c | 49 ++++++++++++++++++++++++++++++++++++++++++++++- fs/ext4/page-io.c | 40 ++++++++++++++++++++++++++++++++++++++ 3 files changed, 94 insertions(+), 1 deletion(-) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 1c3d736fb700..089dbd39c5c2 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -3986,6 +3986,12 @@ extern int ext4_move_extents(struct file *o_filp, st= ruct file *d_filp, __u64 len, __u64 *moved_len); =20 /* page-io.c */ +/* + * The I/O range covers the zeroed EOF block that straddles i_disksize + * and will advance it upon completion. + */ +#define EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO 1UL + extern int __init ext4_init_pageio(void); extern void ext4_exit_pageio(void); extern ext4_io_end_t *ext4_init_io_end(struct inode *inode, gfp_t flags); diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index e4a4396eaf87..a0707310b464 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -4366,7 +4366,10 @@ static int ext4_iomap_writeback_submit(struct iomap_= writepage_ctx *wpc, int error) { struct iomap_ioend *ioend =3D wpc->wb_ctx; - struct ext4_inode_info *ei =3D EXT4_I(ioend->io_inode); + struct inode *inode =3D ioend->io_inode; + struct ext4_inode_info *ei =3D EXT4_I(inode); + unsigned int blocksize =3D i_blocksize(inode); + loff_t pstart, plen; =20 /* * After I/O completion, a worker needs to be scheduled when: @@ -4379,6 +4382,21 @@ static int ext4_iomap_writeback_submit(struct iomap_= writepage_ctx *wpc, test_opt(ioend->io_inode->i_sb, DATA_ERR_ABORT)) ioend->io_bio.bi_end_io =3D ext4_iomap_end_bio; =20 + /* + * Mark the I/O as DISKSIZE_GROW_IO by setting io_private to + * EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO if it covers the pending range. + * Such I/O will allow or trigger i_disksize advancement in the + * ioend worker. + */ + plen =3D ext4_iomap_get_disksize_pending_range(inode, &pstart); + if (plen && + round_down(ioend->io_offset, blocksize) <=3D pstart && + round_up(ioend->io_offset + ioend->io_size, blocksize) >=3D + pstart + plen) { + ioend->io_bio.bi_end_io =3D ext4_iomap_end_bio; + ioend->io_private =3D (void *)EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO; + } + /* * ext4_iomap_end_bio() always defers endio processing, disable * generic BIO in task to avoid double deferral since we will use @@ -4398,6 +4416,29 @@ static const struct iomap_writeback_ops ext4_writeba= ck_ops =3D { .writeback_submit =3D ext4_iomap_writeback_submit, }; =20 +/* + * If the current writeback range begins after the pending zeroed EOF + * block range which straddles i_disksize, issue a separate writeback to + * flush it first, so as to avoid prolonged waiting. + */ +static void ext4_iomap_wb_submit_zeroed_eof(struct inode *inode, + struct writeback_control *wbc) +{ + struct address_space *mapping =3D inode->i_mapping; + loff_t pstart, plen, range_start; + + if (wbc->range_cyclic) + range_start =3D (loff_t)mapping->writeback_index << PAGE_SHIFT; + else + range_start =3D wbc->range_start; + + plen =3D ext4_iomap_get_disksize_pending_range(inode, &pstart); + if (!plen || range_start < pstart + plen) + return; + + filemap_fdatawrite_range(mapping, pstart, pstart + plen - 1); +} + static int ext4_iomap_writepages(struct address_space *mapping, struct writeback_control *wbc) { @@ -4415,6 +4456,12 @@ static int ext4_iomap_writepages(struct address_spac= e *mapping, if (unlikely(ret)) return ret; =20 + /* + * Submit the pending zeroed EOF block range if the entire + * writeback range lies beyond it. + */ + ext4_iomap_wb_submit_zeroed_eof(inode, wbc); + alloc_ctx =3D ext4_writepages_down_read(sb); trace_ext4_writepages(inode, wbc); ret =3D iomap_writepages(&wpc); diff --git a/fs/ext4/page-io.c b/fs/ext4/page-io.c index 2888e0057561..955ff88045db 100644 --- a/fs/ext4/page-io.c +++ b/fs/ext4/page-io.c @@ -549,6 +549,34 @@ void ext4_bio_write_folio(struct ext4_io_submit *io, s= truct folio *folio, } while ((bh =3D bh->b_this_page) !=3D head); } =20 +/* + * If the current writeback range starts beyond the zeroed EOF pending + * range that straddles i_disksize, wait for the zeroed data from + * ext4_block_zero_eof() to be written out first. Otherwise, extending + * i_disksize may expose stale data in the old EOF block. + */ +static void ext4_iomap_wb_disksize_pending_wait(struct inode *inode, + loff_t pos, size_t size) +{ + loff_t disksize =3D READ_ONCE(EXT4_I(inode)->i_disksize); + loff_t pstart, plen; + + /* + * Overwrite I/Os and I/Os covering the EOF block do not need to + * wait: the former do not advance i_disksize past the pending + * boundary, and the latter are the pending I/O itself (cleared in + * the bio completion path). + */ + if (pos < round_up(disksize, i_blocksize(inode))) + return; + + plen =3D ext4_iomap_get_disksize_pending_range(inode, &pstart); + if (!plen || pos < pstart + plen) + return; + + ext4_iomap_wait_disksize_pending(inode); +} + static int ext4_iomap_wb_update_disksize(handle_t *handle, struct inode *i= node, loff_t end) { @@ -594,6 +622,9 @@ static void ext4_iomap_finish_ioend(struct iomap_ioend = *ioend) end <=3D READ_ONCE(EXT4_I(inode)->i_disksize)) goto out; =20 + /* Wait for disksize-pending zeroed data to be written out. */ + ext4_iomap_wb_disksize_pending_wait(inode, pos, size); + /* * We may need to convert one extent, update the i_disksize and * dirty the inode. @@ -660,8 +691,17 @@ void ext4_iomap_end_bio(struct bio *bio) { struct iomap_ioend *ioend =3D iomap_ioend_from_bio(bio); struct ext4_inode_info *ei =3D EXT4_I(ioend->io_inode); + unsigned long io_mode =3D (unsigned long)ioend->io_private; unsigned long flags; =20 + /* + * This is a disksize-pending I/O: clear the disksize-pending + * state set in ext4_block_zero_eof() and wake up all waiters + * that will update the inode i_disksize. + */ + if (io_mode =3D=3D EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO) + ext4_iomap_clear_disksize_pending(ioend->io_inode); + spin_lock_irqsave(&ei->i_completed_io_lock, flags); if (list_empty(&ei->i_rsv_conversion_list)) queue_work(EXT4_SB(ioend->io_inode->i_sb)->rsv_conversion_wq, --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EF78E39021F; Fri, 14 Aug 2026 09:39:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700401; cv=none; b=SB/6o7LUjFhQt1aBk3qUHKSym8IZuNi4wM2fWps/Hx0bIPEnj3P/Uglp7pOJssmkQpq3ghofQtaaJHQrivzrRDGk4IhSRVnYhibkVlQC4BHNlAoxO5XreAXiaHfas2VJr3GrCCAFecKb731MAEjsE/IJmPuStc8cW1ymtfngHU8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700401; c=relaxed/simple; bh=XgqzvxNmOhF16QmT7ZHdWzE6+Dml6j+CULKFMnfRcQ0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=QHMn+IeHoXjzQVnd9hj5mQXa7K8SVP+HfuN72mlw/rKE51SnPCpGRt/MpW4LT9GHOULYwK2ZzwKcZM8+MNKbLqz6omGMjelxkcF2ihvBdOMIKaWbR2yPLGZsW4sdzCLcKMZCL0CP8r9g7xk1gisPhM/JK5IfYqmrBVICXtwSZ9U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.177]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyj2YcFzYQvZ7; Fri, 14 Aug 2026 17:39:41 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 8B8714058C; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S28; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 24/32] ext4: advance i_disksize to i_size upon disksize-grow I/O completion Date: Fri, 14 Aug 2026 17:33:23 +0800 Message-ID: <20260814093331.1703882-25-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S28 X-Coremail-Antispam: 1UD129KBjvJXoWxZF1kKF1UWF4kJw15Aw43Jrb_yoWrurWxpr W5Krn8Gr4kXr1a9rsaqry0vr1Skw4rA3y8JFW7GrWjqFyYkrsagFWxKFyfWFy0yrs3Zw4q qa1Dtr4UWw1kAr7anT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi When the disksize-grow-pending zeroed EOF I/O completes, i_disksize has to be advanced. Advancing it only to the end of that specific I/O would discard any later i_disksize updates from concurrent fallocate or similar operations, causing filesystem inconsistency. Scanning dirty or writeback folios beyond the current position to compute a safe advance target is expensive and racy with concurrent fallocate, so instead advance i_disksize directly to i_size. This may expose zeroed data (not stale data) after crash recovery when dirty data in the range is not yet on disk, but only for unaligned append writes, which is deemed acceptable. To support this, teach ext4_iomap_wb_update_disksize() to take an is_disksize_grow flag and advance i_disksize to i_size when set, and have ext4_iomap_finish_ioend() pass the flag based on the EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO tag of the completing ioend. Suggested-by: Jan Kara Signed-off-by: Zhang Yi --- fs/ext4/page-io.c | 44 +++++++++++++++++++++++++++++++++++++------- 1 file changed, 37 insertions(+), 7 deletions(-) diff --git a/fs/ext4/page-io.c b/fs/ext4/page-io.c index 955ff88045db..4f1176b9332f 100644 --- a/fs/ext4/page-io.c +++ b/fs/ext4/page-io.c @@ -578,9 +578,9 @@ static void ext4_iomap_wb_disksize_pending_wait(struct = inode *inode, } =20 static int ext4_iomap_wb_update_disksize(handle_t *handle, struct inode *i= node, - loff_t end) + loff_t end, bool is_disksize_grow) { - loff_t new_disksize =3D end; + loff_t new_disksize, i_size; struct ext4_inode_info *ei =3D EXT4_I(inode); int ret; =20 @@ -589,9 +589,36 @@ static int ext4_iomap_wb_update_disksize(handle_t *han= dle, struct inode *inode, * i_data_sem. */ down_write(&ei->i_data_sem); - new_disksize =3D min(new_disksize, i_size_read(inode)); + i_size =3D i_size_read(inode); + + /* + * EXT4_STATE_DISKSIZE_GROW_PENDING is cleared when the pending + * I/O completes. However, another thread may have re-set the bit + * between that point and here, meaning i_disksize has already + * been advanced and a new EOF zeroing has been initiated. In that + * case, do not advance i_disksize to i_size; leave it to the + * next pending grow ioend. + */ + if (is_disksize_grow && + ext4_test_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING)) + is_disksize_grow =3D false; + + /* + * Update i_disksize to i_size when EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO + * completes. This is safe because we never directly allocate written + * blocks during buffered writes. + * + * This ensures that i_disksize is correctly advanced during + * truncate-up or append fallocate on a block-unaligned file, + * preventing it from remaining stale. The tradeoff is that zeroed + * data may be exposed after crash recovery if dirty data in this + * range is not yet on disk, but stale data will never be exposed. + * This is because the extent is only converted to written state + * after the data has been persisted. + */ + new_disksize =3D is_disksize_grow ? i_size : min(end, i_size); if (new_disksize > ei->i_disksize) - ei->i_disksize =3D new_disksize; + WRITE_ONCE(ei->i_disksize, new_disksize); up_write(&ei->i_data_sem); ret =3D ext4_mark_inode_dirty(handle, inode); if (ret) @@ -607,6 +634,8 @@ static void ext4_iomap_finish_ioend(struct iomap_ioend = *ioend) loff_t pos =3D ioend->io_offset; size_t size =3D ioend->io_size; loff_t end =3D pos + size; + unsigned long io_mode =3D (unsigned long)ioend->io_private; + bool is_disksize_grow =3D (io_mode =3D=3D EXT4_IOMAP_IOEND_DISKSIZE_GROW_= IO); handle_t *handle; int credits; int ret, err; @@ -619,7 +648,7 @@ static void ext4_iomap_finish_ioend(struct iomap_ioend = *ioend) } =20 if (!(ioend->io_flags & IOMAP_IOEND_UNWRITTEN) && - end <=3D READ_ONCE(EXT4_I(inode)->i_disksize)) + end <=3D READ_ONCE(EXT4_I(inode)->i_disksize) && !is_disksize_grow) goto out; =20 /* Wait for disksize-pending zeroed data to be written out. */ @@ -638,8 +667,9 @@ static void ext4_iomap_finish_ioend(struct iomap_ioend = *ioend) } =20 /* Update on-disk size after I/O is completed. */ - if (end > READ_ONCE(EXT4_I(inode)->i_disksize)) { - ret =3D ext4_iomap_wb_update_disksize(handle, inode, end); + if (end > READ_ONCE(EXT4_I(inode)->i_disksize) || is_disksize_grow) { + ret =3D ext4_iomap_wb_update_disksize(handle, inode, end, + is_disksize_grow); if (ret) goto out_journal; } --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 13A8D463B96; Fri, 14 Aug 2026 09:40:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700416; cv=none; b=JUBg7QHCNQWwFLV3VfttBn7MNth7tM2sXj0WSjoB6r/Fz+7usp7OR0xlqQmQ29fo2iXXjJVz1/HfrMroD8VUhstquxyG7OVykIX6DHE4PbxFa7vzf1k5OGWRv7l2DJt8ZPKTQj20OhcifUFWx/o37GNEUdkerPWkyoDAeRLSq1M= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700416; c=relaxed/simple; bh=tJ3SKti+kV8IXZXsE1bBAwJicHqq7/5Pc1jCj3Mq3nE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rcquNKJoTgYcFfTcaG6fcdLw2G4oCJ4k9QNcWEX5vrmakDjRYt5mmu1DBnE6Togwkn/bY7DENW/9V0XzR72t6R638Rfa2v1p4+iTxamTZRHf9/QtBg1bqpF0quGSyEKk5uWjka9kbnUa9WKm+mWTaybXCJKAaNO37U0MVuxqOgM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyS6Q5ZzKHNPP; Fri, 14 Aug 2026 17:39:28 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 9E2364056D; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S29; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 25/32] ext4: defer i_disksize update while DISKSIZE_GROW_PENDING is set Date: Fri, 14 Aug 2026 17:33:24 +0800 Message-ID: <20260814093331.1703882-26-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S29 X-Coremail-Antispam: 1UD129KBjvJXoWxtr4rAF43XryfJFWxJFWkZwb_yoW7uFW3pr nxKF95Jr1vq3sF9ws2qryIvr4Fya18Jw47JFy29F4qvFy5Aw4IqF1xtry7GFW8trZ5Ja1j qFZ5Krs5Cw18CrJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmv14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F 4UJwA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE 3s1le2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2I x0cI8IcVAFwI0_Jw0_WrylYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8 JwACjcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2 ka0xkIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Y z7v_Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zV AF1VAY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Ar0_tr1l IxAIcVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r 1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26r4UJVWxJrUv cSsGvfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Operations like append allocate, zero range, and truncate update i_disksize directly. If the new i_disksize exceeds the original value while the zeroed EOF block is still awaiting writeback, metadata may be persisted before the zeroed data, exposing stale data on crash. Defer i_disksize updates while EXT4_STATE_DISKSIZE_GROW_PENDING is set; the ioend worker for the pending block will advance i_disksize to i_size once the zeroed data is written back. The tradeoff is that i_disksize may lag i_size transiently, but this is observable only to callers that read i_disksize directly. Introduce __ext4_set_i_disksize() to centralize the bit check for callers already holding i_data_sem (ext4_ext_truncate and ext4_set_inode_size), and refactor ext4_update_inode_size() to take i_data_sem itself and check the bit atomically with i_size_write(), so the ioend worker observes the latest i_size under the same lock. Suggested-by: Jan Kara Signed-off-by: Zhang Yi --- fs/ext4/ext4.h | 47 ++++++++++++++++++++++++++++++++++++++++++----- fs/ext4/extents.c | 2 +- fs/ext4/inode.c | 8 +++++--- 3 files changed, 48 insertions(+), 9 deletions(-) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 089dbd39c5c2..b26ce3183bac 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -3605,30 +3605,67 @@ do { \ #define EXT4_FREECLUSTERS_WATERMARK 0 #endif =20 -/* Update i_disksize. Requires i_rwsem to avoid races with truncate */ +/* + * Update i_disksize. Requires i_rwsem to avoid races with truncate. + * + * In the iomap buffered I/O path, the EXT4_STATE_DISKSIZE_GROW_PENDING + * inode state bit indicates that the zeroed EOF partial block which + * straddles i_disksize is still waiting writeback. In that case, + * i_disksize will be updated after the pending zeroed data has been + * written out. + */ static inline void ext4_update_i_disksize(struct inode *inode, loff_t news= ize) { WARN_ON_ONCE(S_ISREG(inode->i_mode) && !inode_is_locked(inode)); down_write(&EXT4_I(inode)->i_data_sem); - if (newsize > EXT4_I(inode)->i_disksize) + if (newsize > EXT4_I(inode)->i_disksize && + !ext4_test_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING)) WRITE_ONCE(EXT4_I(inode)->i_disksize, newsize); up_write(&EXT4_I(inode)->i_data_sem); } =20 -/* Update i_size, i_disksize. Requires i_rwsem to avoid races with truncat= e */ +static inline void __ext4_set_i_disksize(struct inode *inode, loff_t newsi= ze) +{ + WARN_ON_ONCE(!rwsem_is_locked(&EXT4_I(inode)->i_data_sem)); + + if (newsize < EXT4_I(inode)->i_disksize || + !ext4_test_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING)) + WRITE_ONCE(EXT4_I(inode)->i_disksize, newsize); +} + +/* + * Update i_size and i_disksize to @newsize. Requires i_rwsem to avoid + * races with truncate. + * + * In the iomap buffered I/O path, i_disksize is updated only if no zeroed + * pending block straddles i_disksize (EXT4_STATE_DISKSIZE_GROW_PENDING + * clear), otherwise the ioend worker for the pending block will advance + * i_disksize once the pending block is written back. Both updates happen + * under i_data_sem so that the writeback ioend worker can always see the + * latest i_size under the same semaphore. + * + * Returns 0 if nothing changed, 1 if i_size was raised, 2 if i_disksize + * was raised, or 3 if both were. + */ static inline int ext4_update_inode_size(struct inode *inode, loff_t newsi= ze) { int changed =3D 0; =20 + if (newsize <=3D inode->i_size && newsize <=3D EXT4_I(inode)->i_disksize) + return 0; + + down_write(&EXT4_I(inode)->i_data_sem); if (newsize > inode->i_size) { i_size_write(inode, newsize); changed =3D 1; } - if (newsize > EXT4_I(inode)->i_disksize) { - ext4_update_i_disksize(inode, newsize); + if (newsize > EXT4_I(inode)->i_disksize && + !ext4_test_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING)) { + WRITE_ONCE(EXT4_I(inode)->i_disksize, newsize); changed |=3D 2; } + up_write(&EXT4_I(inode)->i_data_sem); return changed; } =20 diff --git a/fs/ext4/extents.c b/fs/ext4/extents.c index 5a06259a9b5d..fc5aa2dbefcf 100644 --- a/fs/ext4/extents.c +++ b/fs/ext4/extents.c @@ -4561,7 +4561,7 @@ int ext4_ext_truncate(handle_t *handle, struct inode = *inode) */ =20 /* we have to know where to truncate from in crash case */ - EXT4_I(inode)->i_disksize =3D inode->i_size; + __ext4_set_i_disksize(inode, inode->i_size); err =3D ext4_mark_inode_dirty(handle, inode); if (err) return err; diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index a0707310b464..0fdc31b21be5 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -6622,8 +6622,10 @@ static void ext4_wait_for_tail_page_commit(struct in= ode *inode) * Set i_size and i_disksize to 'newsize'. * * Both i_rwsem and i_data_sem are required here to avoid races between - * generic append writeback and concurrent truncate that also modify - * i_size and i_disksize. + * generic append writeback (or zeroed pending I/O writeback) and + * concurrent operations (e.g., fallocate, truncate) that also modify + * i_size and i_disksize. This also ensures that the writeback ioend worker + * observes the latest i_size under the same lock protection. */ static inline void ext4_set_inode_size(struct inode *inode, loff_t newsize) { @@ -6631,7 +6633,7 @@ static inline void ext4_set_inode_size(struct inode *= inode, loff_t newsize) =20 down_write(&EXT4_I(inode)->i_data_sem); i_size_write(inode, newsize); - EXT4_I(inode)->i_disksize =3D newsize; + __ext4_set_i_disksize(inode, newsize); up_write(&EXT4_I(inode)->i_data_sem); } =20 --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6ED18415F33; Fri, 14 Aug 2026 09:39:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700404; cv=none; b=U+vcg12eIWYiwG1H8H76xwziNdI/Nb5hbv0sWl145O53f3Ck03P2WRg4/QcmwdZSteI6xBCaY7zzJbHLWxw+mDQSZIb6V3Yu2mozpHaavsQXqWwgYNHOO6nKOFOZCJqYPfHq62W4xkL1d+A0X2t7W7rLc/KYZKUvJMmXHLgWFPA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700404; c=relaxed/simple; bh=3Ady3wt/uC5/25alJKtbTUQ/GfPseQtUw5Ix/YPfleI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Ztd0rkXzqVAIDbvHlQQpu3HZtiAYAoVXOUz9V/1jiDrhhmoCHC9R5zYfibMIyoPycL+e3rVgMfV9YeSO0O+0dhKbnd/WEwtx978/sTvfWQU7fRat2axlqhuGq/XnOaGLB6/xWRgoExoyRkySo0Bn2txDN7RF+dxH5qRHMa7mep0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyj3tRCzYQvZ6; Fri, 14 Aug 2026 17:39:41 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id BE54C4056D; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S30; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 26/32] ext4: submit and wait for disksize-grow I/O in fallocate paths Date: Fri, 14 Aug 2026 17:33:25 +0800 Message-ID: <20260814093331.1703882-27-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S30 X-Coremail-Antispam: 1UD129KBjvJXoW3GryDur18CF15Kr4Dur17ZFb_yoWxZF1Dp3 y5Kr15Gr48Wr97ur4SqF4UXr1Yya1xKr48WrZ7ur12vFy5C34xKF1YyFyY9Fy8JrZ5Cw4j vF4qg3y5Cw17A3DanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmq14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Cr1j6r xdM28EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq 3wAS0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7 IYx2IY67AKxVWUtVWrXwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4U M4x0Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2 kIc2xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkE bVWUJVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67 AF67kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVW7JVWDJwCI 42IY6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F4UJwCI42IY6xAIw20EY4v20xvaj40_Jr0_JF 4lIxAIcVC2z280aVAFwI0_Cr0_Gr1UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIY CTnIWIevJa73UjIFyTuYvjTRNF4EDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Collapse range and insert range update i_disksize directly under i_data_sem. If the operation runs while the zeroed EOF block is still awaiting writeback, i_disksize could advance past the zeroed boundary before the zeroed data is persisted, exposing stale data on crash. Deferring i_disksize updates like fallocate and zero_range is not an option here because the shift would move written extents beyond the current i_disksize. So flush and wait for the pending zeroed EOF block before these operations advance i_disksize. Since these operations already perform writeback, the extra flush does not add significant overhead. In addition, for ext4_update_disksize_before_punch(), if the punch discards the pending block, the zeroed data will never be written back before advancing i_disksize, so it is also necessary to sync the pending EOF range there. Finally, for the SYNC variants of zero_range and fallocate, this also guarantees the i_disksize update is persisted on the synchronous return. Signed-off-by: Zhang Yi --- fs/ext4/ext4.h | 2 ++ fs/ext4/extents.c | 53 ++++++++++++++++++++++++++++++++++++++++------- fs/ext4/inode.c | 34 ++++++++++++++++++++++++++++++ 3 files changed, 82 insertions(+), 7 deletions(-) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index b26ce3183bac..504dce9fdc6b 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -3226,6 +3226,8 @@ void ext4_iomap_clear_disksize_pending(struct inode *= inode); void ext4_iomap_wait_disksize_pending(struct inode *inode); unsigned int ext4_iomap_get_disksize_pending_range(struct inode *inode, loff_t *start); +extern int ext4_iomap_sync_zeroed_eof(struct inode *inode, + loff_t offset, loff_t end); extern int ext4_block_zero_eof(struct inode *inode, loff_t from, loff_t en= d); =20 #define EXT4_PARTIAL_ZERO_START 0x1 diff --git a/fs/ext4/extents.c b/fs/ext4/extents.c index fc5aa2dbefcf..dda6d50e96e3 100644 --- a/fs/ext4/extents.c +++ b/fs/ext4/extents.c @@ -4876,6 +4876,16 @@ static long ext4_zero_range(struct file *file, loff_= t offset, return ret; } =20 + /* + * In SYNC mode, sync the pending zeroed EOF block to ensure the + * i_disksize update is persisted. + */ + if (((file->f_flags & O_SYNC) || IS_SYNC(inode)) && new_size) { + ret =3D ext4_iomap_sync_zeroed_eof(inode, 0, LLONG_MAX); + if (ret) + return ret; + } + handle =3D ext4_journal_start(inode, EXT4_HT_MISC, 1); if (IS_ERR(handle)) { ret =3D PTR_ERR(handle); @@ -4928,10 +4938,20 @@ static long ext4_do_fallocate(struct file *file, lo= ff_t offset, if (ret) goto out; =20 - if (((file->f_flags & O_SYNC) || IS_SYNC(inode)) && - EXT4_SB(inode->i_sb)->s_journal) { - ret =3D ext4_fc_commit(EXT4_SB(inode->i_sb)->s_journal, - EXT4_I(inode)->i_sync_tid); + if ((file->f_flags & O_SYNC) || IS_SYNC(inode)) { + /* + * Sync the pending zeroed EOF block to ensure the + * i_disksize update is persisted. + */ + if (new_size) { + ret =3D ext4_iomap_sync_zeroed_eof(inode, 0, LLONG_MAX); + if (ret) + goto out; + } + if (EXT4_SB(inode->i_sb)->s_journal) { + ret =3D ext4_fc_commit(EXT4_SB(inode->i_sb)->s_journal, + EXT4_I(inode)->i_sync_tid); + } } out: trace_ext4_fallocate_exit(inode, offset, @@ -5668,6 +5688,14 @@ static int ext4_collapse_range(struct file *file, lo= ff_t offset, loff_t len) if (end >=3D inode->i_size) return -EINVAL; =20 + /* + * Persist the pending zeroed EOF block to ensure i_disksize + * can be safely updated thereafter. + */ + ret =3D ext4_iomap_sync_zeroed_eof(inode, 0, LLONG_MAX); + if (ret) + return ret; + /* * Write tail of the last page before removed range and data that * will be shifted since they will get removed from the page cache @@ -5715,9 +5743,11 @@ static int ext4_collapse_range(struct file *file, lo= ff_t offset, loff_t len) goto out_handle; } =20 + WARN_ON_ONCE(ext4_test_inode_state(inode, + EXT4_STATE_DISKSIZE_GROW_PENDING)); new_size =3D inode->i_size - len; i_size_write(inode, new_size); - EXT4_I(inode)->i_disksize =3D new_size; + __ext4_set_i_disksize(inode, new_size); =20 up_write(&EXT4_I(inode)->i_data_sem); ret =3D ext4_mark_inode_dirty(handle, inode); @@ -5770,6 +5800,14 @@ static int ext4_insert_range(struct file *file, loff= _t offset, loff_t len) if (len > inode->i_sb->s_maxbytes - inode->i_size) return -EFBIG; =20 + /* + * Persist the pending zeroed EOF block to ensure i_disksize + * can be safely updated thereafter. + */ + ret =3D ext4_iomap_sync_zeroed_eof(inode, 0, LLONG_MAX); + if (ret) + return ret; + /* * Write out all dirty pages. Need to round down to align start offset * to page size boundary for page size > block size. @@ -5789,8 +5827,9 @@ static int ext4_insert_range(struct file *file, loff_= t offset, loff_t len) ext4_fc_mark_ineligible(sb, EXT4_FC_REASON_FALLOC_RANGE, handle); =20 /* Expand file to avoid data loss if there is error while shifting */ - inode->i_size +=3D len; - EXT4_I(inode)->i_disksize +=3D len; + WARN_ON_ONCE(ext4_test_inode_state(inode, + EXT4_STATE_DISKSIZE_GROW_PENDING)); + ext4_update_inode_size(inode, inode->i_size + len); ret =3D ext4_mark_inode_dirty(handle, inode); if (ret) goto out_handle; diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 0fdc31b21be5..056937e27859 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -4823,6 +4823,32 @@ static int ext4_block_zero_range(struct inode *inode, zero_written); } =20 +/* + * Submit and wait for the pending zeroed EOF block range to complete + * if the given range [@offset, @end) fully covers it. Must be called + * outside the context of an active journal handle and hold the i_rwsem. + */ +int ext4_iomap_sync_zeroed_eof(struct inode *inode, loff_t offset, loff_t = end) +{ + loff_t pstart, plen; + int ret; + + if (!ext4_inode_buffered_iomap(inode)) + return 0; + + plen =3D ext4_iomap_get_disksize_pending_range(inode, &pstart); + if (!plen || offset > pstart || end < pstart + plen) + return 0; + + ret =3D filemap_fdatawrite_range(inode->i_mapping, pstart, + pstart + plen - 1); + if (ret) + return ret; + + ext4_iomap_wait_disksize_pending(inode); + return 0; +} + /* * Zero out a mapping from file offset 'from' up to the end of the block * which corresponds to 'from' or to the given 'end' inside this block. @@ -4988,6 +5014,14 @@ int ext4_update_disksize_before_punch(struct inode *= inode, loff_t offset, if (offset > size) return 0; =20 + /* + * We are going to punch the pending zeroed EOF block, persist + * it to ensure i_disksize can be safely updated thereafter. + */ + ret =3D ext4_iomap_sync_zeroed_eof(inode, offset, offset + len); + if (ret) + return ret; + if (offset + len < size) size =3D offset + len; if (EXT4_I(inode)->i_disksize >=3D size) --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 89E3640B108; Fri, 14 Aug 2026 09:39:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700401; cv=none; b=SDwoRG6j4f0XwlIeIsSoHnCjMGpq8zM6cfY0L3ZBKXqKz+4UqaX36cUzGZNsOba9upb2Mf0is0+QOhcrmbooeny0+s8lkMVipod2NpW2xBS2FJUplK2WAoba/EU9GMvOs9Eug6OVJoKuPPCJ6M1C5/mxT9IQw7Cx/hB1Q7sHeek= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700401; c=relaxed/simple; bh=kl0qMxkFE6pnnqyk3HrpVS4AVBkVnvacgrC7uNqsDxk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=e+lkNiLUgI4WjFW3GlR+pNqWOwKQniE80gAgwr8o9sWUliXTyI3pznrwmCjH2f5hjouoKXc6sz8rq1smnXy09OcZO+wpfiZtCrbkVc+xQDILBi+ZVQHLx21ek6JWFEP1h7wGFxJWxzmJvWcdfZkk4L3ZqxTkDeqcDIPNVDDeLC8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyj4hYFzYQvYR; Fri, 14 Aug 2026 17:39:41 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id D98B34056F; Fri, 14 Aug 2026 17:39:49 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S31; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 27/32] ext4: clear DISKSIZE_GROW_PENDING on truncate or error Date: Fri, 14 Aug 2026 17:33:26 +0800 Message-ID: <20260814093331.1703882-28-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S31 X-Coremail-Antispam: 1UD129KBjvJXoWxZFyDur13ZFW8tF13Kw13XFb_yoW7Jw4kpr 98KFn8Ga48Wr9F9r4Sqr1UZr1Fga18ta1UJFZ7uF4kXF15A34IgF18ta43ZF4jkrZ3Jw4a qFW0kr4DWw48GrJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUma14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Cr1j6r xdM28EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq 3wAS0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7 IYx2IY67AKxVWUtVWrXwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4U M4x0Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2 kIc2xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkE bVWUJVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67 AF67kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVW7JVWDJwCI 42IY6xIIjxv20xvEc7CjxVAFwI0_Cr1j6rxdMIIF0xvE42xK8VAvwI8IcIk0rVWUJVWUCw CI42IY6I8E87Iv67AKxVWxJVW8Jr1lIxAIcVC2z280aVCY1x0267AKxVWxJr0_GcJvcSsG vfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ From: Zhang Yi The disksize-grow-pending state is set when a zeroed EOF block is queued for writeback and cleared by the ioend completion path once writeback finishes. However, the zeroed block may be discarded before writeback completes =E2=80=94 through folio discard, truncate, or unlink and inode eviction. Additionally, if the filesystem enters an emergency state, the block will no longer be written back. In any of these cases, leaving the bit set would block subsequent writeback indefinitely. Therefore, we must clear it on all paths that invalidate the pending block before writeback completes: - ext4_iomap_discard_folio() on folio discard. - ext4_evict_inode() when an unlinked inode is destroyed. - ext4_iomap_writepages() when the filesystem is in emergency state. In ext4_truncate_down(), truncating past the pending zeroed EOF block also invalidates the pending disksize update, so the bit must be cleared there as well. Finally, add a WARN_ON in ext4_destroy_inode() to catch any inode destroyed with the bit still set. Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 33 +++++++++++++++++++++++++++++++-- fs/ext4/super.c | 19 +++++++++++++------ 2 files changed, 44 insertions(+), 8 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 056937e27859..a1dfb70127ca 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -273,6 +273,8 @@ void ext4_evict_inode(struct inode *inode) =20 if (ext4_should_order_data(inode)) ext4_begin_ordered_truncate(inode, 0); + if (ext4_inode_buffered_iomap(inode)) + ext4_iomap_clear_disksize_pending(inode); truncate_inode_pages_final(&inode->i_data); =20 /* @@ -4344,8 +4346,17 @@ static void ext4_iomap_discard_folio(struct folio *f= olio, loff_t pos) { struct inode *inode =3D folio->mapping->host; loff_t length =3D folio_pos(folio) + folio_size(folio) - pos; + loff_t pstart, plen; =20 ext4_iomap_punch_delalloc(inode, pos, length, NULL); + + /* + * Clear the disksize-grow-pending state if the zeroed EOF block + * fails to write back and is discarded. + */ + plen =3D ext4_iomap_get_disksize_pending_range(inode, &pstart); + if (plen && pos <=3D pstart && folio_next_pos(folio) >=3D pstart + plen) + ext4_iomap_clear_disksize_pending(inode); } =20 static ssize_t ext4_iomap_writeback_range(struct iomap_writepage_ctx *wpc, @@ -4453,8 +4464,15 @@ static int ext4_iomap_writepages(struct address_spac= e *mapping, }; =20 ret =3D ext4_emergency_state(sb); - if (unlikely(ret)) + if (unlikely(ret)) { + /* + * The filesystem is in an emergency state and no further + * writeback will occur. Clear the disksize-grow-pending + * state to avoid complaints when the inode is destroyed. + */ + ext4_iomap_clear_disksize_pending(inode); return ret; + } =20 /* * Submit the pending zeroed EOF block range if the entire @@ -6741,7 +6759,18 @@ static int ext4_truncate_down(struct inode *inode, l= off_t oldsize, start_lblk =3D newsize > 0 ? (newsize - 1) >> inode->i_blkbits : 0; ext4_fc_track_range(handle, inode, start_lblk, EXT_MAX_BLOCKS - 1); =20 - ext4_set_inode_size(inode, newsize); + down_write(&EXT4_I(inode)->i_data_sem); + /* + * Truncate the zeroed EOF block invalidates the pending disksize + * update, so clear the disksize-grow-pending state. + */ + if (ext4_test_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING) && + (newsize <=3D EXT4_I(inode)->i_disksize)) + ext4_iomap_clear_disksize_pending(inode); + + i_size_write(inode, newsize); + __ext4_set_i_disksize(inode, newsize); + up_write(&EXT4_I(inode)->i_data_sem); =20 ret =3D ext4_mark_inode_dirty(handle, inode); ext4_journal_stop(handle); diff --git a/fs/ext4/super.c b/fs/ext4/super.c index 1c2395aa1d53..86ed5228dbe9 100644 --- a/fs/ext4/super.c +++ b/fs/ext4/super.c @@ -1491,12 +1491,19 @@ static void ext4_destroy_inode(struct inode *inode) dump_stack(); } =20 - if (!(EXT4_SB(inode->i_sb)->s_mount_state & EXT4_ERROR_FS) && - WARN_ON_ONCE(EXT4_I(inode)->i_reserved_data_blocks)) - ext4_msg(inode->i_sb, KERN_ERR, - "Inode %llu (%p): i_reserved_data_blocks (%u) not cleared!", - inode->i_ino, EXT4_I(inode), - EXT4_I(inode)->i_reserved_data_blocks); + if (!(EXT4_SB(inode->i_sb)->s_mount_state & EXT4_ERROR_FS)) { + if (WARN_ON_ONCE(EXT4_I(inode)->i_reserved_data_blocks)) + ext4_msg(inode->i_sb, KERN_ERR, + "Inode %llu (%p): i_reserved_data_blocks (%u) not cleared!", + inode->i_ino, EXT4_I(inode), + EXT4_I(inode)->i_reserved_data_blocks); + + if (WARN_ON_ONCE(ext4_test_inode_state(inode, + EXT4_STATE_DISKSIZE_GROW_PENDING))) + ext4_msg(inode->i_sb, KERN_ERR, + "Inode %llu (%p): EXT4_STATE_DISKSIZE_GROW_PENDING not cleared!", + inode->i_ino, EXT4_I(inode)); + } } =20 static void ext4_shutdown(struct super_block *sb) --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AAFBD37881E; Fri, 14 Aug 2026 09:40:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700417; cv=none; b=BGTrH2satQ7y5UTZ/esQumqOfVMegFyCRuy+i0Ce1cGLdtlPXBKbzzEFCDsnjQhSozoClVEYGe+7cyZdc/ubGtgpI488IeZ4vnsPZY15LRQyS4BkRdar7m00JFNQ1xgOufbQ+AOYx0esWtFmhKApbDm18v850RiaA6qdYliXwo4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700417; c=relaxed/simple; bh=XRaW+WhG0IzW+aLqKfTtlz/OjUhcYQxz45uXY19KKeI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WZ40eR7BoRQAFqM9OakID6LkZQCp7Cs2yaltMWmQ3AWtnhGsRpmPHTvDYL1Rfc3Qg4o+19P0+MbAeKrJ0tgwp+jjMMBqpmqRdUkcOEhflZ1+sQoovM/GWoRNkC1SDebQHkJC9A8C8PhqMuM80OhYs40QQA/OMeCPm2tazxSi9Mc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyT2F6bzKHNPb; Fri, 14 Aug 2026 17:39:29 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 0E1D24056D; Fri, 14 Aug 2026 17:39:50 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S32; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 28/32] ext4: set DISKSIZE_GROW_PENDING after zeroing unaligned EOF block Date: Fri, 14 Aug 2026 17:33:27 +0800 Message-ID: <20260814093331.1703882-29-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S32 X-Coremail-Antispam: 1UD129KBjvJXoW3ArWxXF4UurW8Cw17Xw48tFb_yoW7WF4kp3 y5Kwn5Cr4DKr9F9w4Sq3WxXr1Y9ayrJayUGFZ7Wr42va45WF1IgFyxt348uFyUJrZ3Ga12 qF45GFWDu3WjyrJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUma14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Cr1j6r xdM28EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq 3wAS0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7 IYx2IY67AKxVWUtVWrXwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4U M4x0Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2 kIc2xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkE bVWUJVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67 AF67kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVW7JVWDJwCI 42IY6xIIjxv20xvEc7CjxVAFwI0_Cr1j6rxdMIIF0xvE42xK8VAvwI8IcIk0rVWUJVWUCw CI42IY6I8E87Iv67AKxVWxJVW8Jr1lIxAIcVC2z280aVCY1x0267AKxVWxJr0_GcJvcSsG vfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi In the iomap buffered I/O path, data=3Dordered mode is not used, so the zeroed EOF block has no implicit ordering with later i_disksize updates. Without the pending state being set, i_disksize can be advanced past the zeroed block before writeback completes, exposing stale data after a crash. Previous patches added the consumer side of the disksize-grow-pending mechanism: the state bit, clear and wait helpers, and ioend tagging. Now add ext4_iomap_mark_disksize_pending() and call it from ext4_block_zero_eof() after zeroing the tail of the block that straddles i_disksize. The helper locks the folio, waits for any in-flight writeback on it to complete, then sets EXT4_STATE_DISKSIZE_GROW_PENDING only if the folio is still dirty. Waiting for writeback prevents folio_test_dirty() from returning false mid-writeback, which would cause us to skip the pending state while zeroed data is still in flight. The dirty check then avoids setting the bit when the data has already been written back. Suggested-by: Jan Kara Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 85 ++++++++++++++++++++++++++++++++++++++++++------- 1 file changed, 73 insertions(+), 12 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index a1dfb70127ca..2ec69e8abe54 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -4841,6 +4841,68 @@ static int ext4_block_zero_range(struct inode *inode, zero_written); } =20 +/* + * Inodes using the iomap buffered I/O path do not use data=3Dordered mode. + * Therefore, we mark the inode as disksize-grow-pending after zeroing the + * EOF block. The zeroed block will be submitted before any subsequent + * data. + * + * In the I/O completion path, ext4_iomap_wb_disksize_pending_wait() will + * wait for I/O completion before advancing i_disksize if the write + * extends beyond the zeroed boundary. + * + * When zeroed I/O is in progress, operations that extend i_disksize are + * handled as follows: + * + * - Truncate up, append fallocate and zero_range: + * Defer the update. The file size will be updated to i_size by the + * end_io handler once the ongoing pending I/O completes. + * + * - Insert range and collapse range operations: + * Wait synchronously for the relevant I/O to complete before updating + * i_disksize. + */ +static int ext4_iomap_mark_disksize_pending(struct inode *inode, loff_t fr= om) +{ + struct folio *folio; + + folio =3D filemap_lock_folio(inode->i_mapping, from >> PAGE_SHIFT); + if (IS_ERR(folio)) + /* Already in writeback and cleared? */ + return PTR_ERR(folio) =3D=3D -ENOENT ? 0 : PTR_ERR(folio); + + /* + * Ensure that in-flight writeback, possibly started after + * iomap_zero_range() unlocked the folio, has completed. Without + * this wait folio_test_dirty() below may miss the zeroed data + * (writeback clears PG_dirty), causing us to skip the + * disksize-grow-pending tracking and potentially expose stale + * on-disk data. + */ + folio_wait_writeback(folio); + WARN_ON_ONCE(folio_test_writeback(folio)); + + /* + * Mark the inode as disksize-grow-pending. The zeroed block will + * be written out by the generic writepages cycle or any other + * syncing operation. + * + * Multiple overlapping unaligned EOF writes should not happen, + * because we only mark the pending state after zeroing the on-disk + * EOF block, and i_disksize can only be updated after the previous + * zeroed pending block has been written back or the dirty folio + * has been discared. + */ + if (likely(folio_test_dirty(folio) && + !ext4_test_inode_state(inode, + EXT4_STATE_DISKSIZE_GROW_PENDING))) + ext4_set_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING); + + folio_unlock(folio); + folio_put(folio); + return 0; +} + /* * Submit and wait for the pending zeroed EOF block range to complete * if the given range [@offset, @end) fully covers it. Must be called @@ -4916,22 +4978,21 @@ int ext4_block_zero_eof(struct inode *inode, loff_t= from, loff_t end) * block, the zeroed data lies beyond the existing on-disk data. It * will be written out before i_disksize is later extended past * i_size, so no stale data can be exposed. - * - * TODO: In the iomap path, handle this by tracking the ordered range - * and updating i_disksize to i_size after the zeroed data has been - * written back. */ - if (ext4_should_order_data(inode) && - did_zero && zero_written && !IS_DAX(inode) && + if (did_zero && zero_written && !IS_DAX(inode) && from < round_up(READ_ONCE(EXT4_I(inode)->i_disksize), blocksize)) { - handle_t *handle; + if (ext4_should_order_data(inode)) { + handle_t *handle; =20 - handle =3D ext4_journal_start(inode, EXT4_HT_MISC, 1); - if (IS_ERR(handle)) - return PTR_ERR(handle); + handle =3D ext4_journal_start(inode, EXT4_HT_MISC, 1); + if (IS_ERR(handle)) + return PTR_ERR(handle); =20 - err =3D ext4_jbd2_inode_add_write(handle, inode, from, length); - ext4_journal_stop(handle); + err =3D ext4_jbd2_inode_add_write(handle, inode, from, + length); + ext4_journal_stop(handle); + } else if (ext4_inode_buffered_iomap(inode)) + err =3D ext4_iomap_mark_disksize_pending(inode, from); if (err) return err; } --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A1882463B93; Fri, 14 Aug 2026 09:40:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700417; cv=none; b=J2QSANT004OaUGKcSyOLqk0ciKAN4Aj1Z93qtgs9Xj770idzVwUdCjRq2gUa8RGTCVnHBXuD+awFR1ZAESV5jyBxeRuVhqU61BGAzYoP+CzEQz8XvmnEjHWAgioAa7i8QWKnBmLxaLhxcahFvQrxBlkEU3RFb2NKDW6Q5SE2Sdg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700417; c=relaxed/simple; bh=1N/B8/PkyJjboJ1cO/VS6WU52V+hOGWHQ3Gb2mv8qwg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pUoZ2m9eOH35bijMWiq3xRKctzLFlZY6Qg2bUmV7kSiCKFcyNckWn/oRjrOn67163T3ViuHrQnxISv3NcDUVtwy5S1CR54v+WDi5ekmvQkTy3NihEp2vvhHhenc8f+QNDgRtnOVplMGX65BuOALIOwGerdxCJBYFyvQFzJH8+Uc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.177]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyT3KdpzKHNMD; Fri, 14 Aug 2026 17:39:29 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 285964058F; Fri, 14 Aug 2026 17:39:50 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S33; Fri, 14 Aug 2026 17:39:49 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 29/32] ext4: add tracepoints for DISKSIZE_GROW_PENDING set, clear, and wait Date: Fri, 14 Aug 2026 17:33:28 +0800 Message-ID: <20260814093331.1703882-30-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S33 X-Coremail-Antispam: 1UD129KBjvJXoWxXw1UKF1kGr1ftF4fJw15twb_yoW5uw1Dpr 1DKrn8Gwn7Jr9I9w4Sqry7Zr1YvF4rGr4UGryfCr12qFWrCryxKF4IqF9IkFy8Aw4kA3y2 gFn0k3ykC3WUCrJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUma14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Cr1j6r xdM28EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq 3wAS0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7 IYx2IY67AKxVWUtVWrXwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4U M4x0Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2 kIc2xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkE bVWUJVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67 AF67kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVW7JVWDJwCI 42IY6xIIjxv20xvEc7CjxVAFwI0_Cr1j6rxdMIIF0xvE42xK8VAvwI8IcIk0rVWUJVWUCw CI42IY6I8E87Iv67AKxVWxJVW8Jr1lIxAIcVC2z280aVCY1x0267AKxVWxJr0_GcJvcSsG vfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Add trace events ext4_iomap_mark_disksize_pending(), ext4_iomap_clear_disksize_pending(), and ext4_iomap_wait_disksize_pending() to track disksize-grow-pending state changes and waiting. Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 6 +++++- include/trace/events/ext4.h | 35 +++++++++++++++++++++++++++++++++++ 2 files changed, 40 insertions(+), 1 deletion(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 2ec69e8abe54..bca7d5c33919 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -133,6 +133,7 @@ void ext4_inode_csum_set(struct inode *inode, struct ex= t4_inode *raw, */ void ext4_iomap_clear_disksize_pending(struct inode *inode) { + trace_ext4_iomap_clear_disksize_pending(inode); ext4_clear_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING); /* * Make sure clearing of EXT4_STATE_DISKSIZE_GROW_PENDING is @@ -151,6 +152,7 @@ void ext4_iomap_clear_disksize_pending(struct inode *in= ode) */ void ext4_iomap_wait_disksize_pending(struct inode *inode) { + trace_ext4_iomap_wait_disksize_pending(inode); wait_on_bit(ext4_inode_state_wait_word(inode), ext4_inode_state_wait_bit(EXT4_STATE_DISKSIZE_GROW_PENDING), TASK_UNINTERRUPTIBLE); @@ -4895,8 +4897,10 @@ static int ext4_iomap_mark_disksize_pending(struct i= node *inode, loff_t from) */ if (likely(folio_test_dirty(folio) && !ext4_test_inode_state(inode, - EXT4_STATE_DISKSIZE_GROW_PENDING))) + EXT4_STATE_DISKSIZE_GROW_PENDING))) { + trace_ext4_iomap_mark_disksize_pending(inode); ext4_set_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING); + } =20 folio_unlock(folio); folio_put(folio); diff --git a/include/trace/events/ext4.h b/include/trace/events/ext4.h index 69596a216dcb..4539ef8e5f86 100644 --- a/include/trace/events/ext4.h +++ b/include/trace/events/ext4.h @@ -3202,6 +3202,41 @@ DEFINE_SET_IOMAP_EVENT(ext4_iomap_buffered_write_beg= in); DEFINE_SET_IOMAP_EVENT(ext4_iomap_map_writeback_range); DEFINE_SET_IOMAP_EVENT(ext4_iomap_zero_begin); =20 +DECLARE_EVENT_CLASS(ext4_iomap_disksize_pending, + TP_PROTO(struct inode *inode), + TP_ARGS(inode), + TP_STRUCT__entry( + __field(dev_t, dev) + __field(u64, ino) + __field(loff_t, i_size) + __field(loff_t, i_disksize) + ), + TP_fast_assign( + __entry->dev =3D inode->i_sb->s_dev; + __entry->ino =3D inode->i_ino; + __entry->i_size =3D i_size_read(inode); + __entry->i_disksize =3D READ_ONCE(EXT4_I(inode)->i_disksize); + ), + TP_printk("dev %d:%d ino %llu i_size %lld i_disksize %lld", + MAJOR(__entry->dev), MINOR(__entry->dev), + __entry->ino, __entry->i_size, __entry->i_disksize) +); + +DEFINE_EVENT(ext4_iomap_disksize_pending, ext4_iomap_mark_disksize_pending, + TP_PROTO(struct inode *inode), + TP_ARGS(inode) +); + +DEFINE_EVENT(ext4_iomap_disksize_pending, ext4_iomap_clear_disksize_pendin= g, + TP_PROTO(struct inode *inode), + TP_ARGS(inode) +); + +DEFINE_EVENT(ext4_iomap_disksize_pending, ext4_iomap_wait_disksize_pending, + TP_PROTO(struct inode *inode), + TP_ARGS(inode) +); + #endif /* _TRACE_EXT4_H */ =20 /* This part must be outside protection */ --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A090F3A0EA2; Fri, 14 Aug 2026 09:40:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700406; cv=none; b=o/2PGXj1mBW+6SxjiiFjynfs7x15v0cpGrp0iThSP+xGvkVUtaD3emTX1iV8SSBnY+vzFFuRaLgBTyud6/fCkQpkYmmmWidw4bXxbpGj1a6PzApQA3k4bnQ7ycpcvcCxtQG3Zw4TGc0rfpW1++/CXwBRm/vn2LMXC4eE/YM0IYY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786700406; c=relaxed/simple; bh=paAgJoIYZCFpB17m47tThq8gFbcfpHuo4I11SE4QkPs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=obfe05kD6+rQVxja71hjX55qDuocErRNE9iBlIkkfAP0Cmy3ypnH/b/LBbzEM8Uk8QRk9g2yfEliJjbSrU2kZ4TJ18p6yWf+1VkyMLExdIIjrL+Ge9l3J4kPQljKnAWm681RLZdiXn9u1a10Kt22jDoFHyvgSJheRnmDokVBdXU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLxyk1bPwzYQvYY; Fri, 14 Aug 2026 17:39:42 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 707FD4056B; Fri, 14 Aug 2026 17:39:50 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBHg0JO4n5qk4oXCQ--.21396S34; Fri, 14 Aug 2026 17:39:50 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 30/32] ext4: add tracepoints for EOF block zeroing and disksize-grow I/O Date: Fri, 14 Aug 2026 17:33:29 +0800 Message-ID: <20260814093331.1703882-31-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBHg0JO4n5qk4oXCQ--.21396S34 X-Coremail-Antispam: 1UD129KBjvJXoW3Gw4rZF4xXFWxWFWxZr1xuFg_yoW3ZF17pw 1DGr98Gw1kJr1Y9r4IqryIqr4F9ayruF48Jr9xuFWjv348ZFZ2qF4xKF98uFy0yrsFywsI gF1vkrZ5G3W8XrJanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUma14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JF0E3s1l82xGYI kIc2x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2 z4x0Y4vE2Ix0cI8IcVAFwI0_Ar0_tr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Cr1j6r xdM28EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq 3wAS0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7 IYx2IY67AKxVWUtVWrXwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4U M4x0Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2 kIc2xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkE bVWUJVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67 AF67kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVW7JVWDJwCI 42IY6xIIjxv20xvEc7CjxVAFwI0_Cr1j6rxdMIIF0xvE42xK8VAvwI8IcIk0rVWUJVWUCw CI42IY6I8E87Iv67AKxVWxJVW8Jr1lIxAIcVC2z280aVCY1x0267AKxVWxJr0_GcJvcSsG vfC2KfnxnUUI43ZEXa7sRiyxR3UUUUU== X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Add tracepoints to track the disksize-grow-pending lifecycle in the writeback path and the block-zero-EOF entry point: - ext4_iomap_wb_disksize_pending_submit: ioend is marked as DISKSIZE_GROW_IO in ext4_iomap_writeback_submit. - ext4_iomap_wb_disksize_pending_complete: ioend of type DISKSIZE_GROW_IO completes in ext4_iomap_end_bio. - ext4_iomap_wb_disksize_pending_wait: ioend worker waits for the pending zeroed EOF block to complete. - ext4_iomap_wb_update_disksize: i_disksize is advanced in ext4_iomap_wb_update_disksize, including the new value and whether the update is a disksize-grow completion. - ext4_block_zero_eof: ext4_block_zero_eof is called with the range and zeroing outcome, capturing the producer-side entry point. Together with the previous mark/clear/wait tracepoints, these cover the full lifetime of the disksize-grow-pending state. Signed-off-by: Zhang Yi --- fs/ext4/inode.c | 3 + fs/ext4/page-io.c | 13 ++++- include/trace/events/ext4.h | 107 ++++++++++++++++++++++++++++++++++++ 3 files changed, 121 insertions(+), 2 deletions(-) diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index bca7d5c33919..ee15366422a1 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -4406,6 +4406,8 @@ static int ext4_iomap_writeback_submit(struct iomap_w= ritepage_ctx *wpc, round_down(ioend->io_offset, blocksize) <=3D pstart && round_up(ioend->io_offset + ioend->io_size, blocksize) >=3D pstart + plen) { + trace_ext4_iomap_wb_disksize_pending_submit(inode, + ioend->io_offset, ioend->io_size); ioend->io_bio.bi_end_io =3D ext4_iomap_end_bio; ioend->io_private =3D (void *)EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO; } @@ -5001,6 +5003,7 @@ int ext4_block_zero_eof(struct inode *inode, loff_t f= rom, loff_t end) return err; } =20 + trace_ext4_block_zero_eof(inode, from, length, did_zero, zero_written); return 0; } =20 diff --git a/fs/ext4/page-io.c b/fs/ext4/page-io.c index 4f1176b9332f..4464eb03c972 100644 --- a/fs/ext4/page-io.c +++ b/fs/ext4/page-io.c @@ -31,6 +31,8 @@ #include "xattr.h" #include "acl.h" =20 +#include + static struct kmem_cache *io_end_cachep; static struct kmem_cache *io_end_vec_cachep; =20 @@ -574,6 +576,7 @@ static void ext4_iomap_wb_disksize_pending_wait(struct = inode *inode, if (!plen || pos < pstart + plen) return; =20 + trace_ext4_iomap_wb_disksize_pending_wait(inode, pos, size); ext4_iomap_wait_disksize_pending(inode); } =20 @@ -617,8 +620,11 @@ static int ext4_iomap_wb_update_disksize(handle_t *han= dle, struct inode *inode, * after the data has been persisted. */ new_disksize =3D is_disksize_grow ? i_size : min(end, i_size); - if (new_disksize > ei->i_disksize) + if (new_disksize > ei->i_disksize) { + trace_ext4_iomap_wb_update_disksize(inode, end, i_size, + ei->i_disksize, new_disksize, is_disksize_grow); WRITE_ONCE(ei->i_disksize, new_disksize); + } up_write(&ei->i_data_sem); ret =3D ext4_mark_inode_dirty(handle, inode); if (ret) @@ -729,8 +735,11 @@ void ext4_iomap_end_bio(struct bio *bio) * state set in ext4_block_zero_eof() and wake up all waiters * that will update the inode i_disksize. */ - if (io_mode =3D=3D EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO) + if (io_mode =3D=3D EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO) { + trace_ext4_iomap_wb_disksize_pending_complete(ioend->io_inode, + ioend->io_offset, ioend->io_size); ext4_iomap_clear_disksize_pending(ioend->io_inode); + } =20 spin_lock_irqsave(&ei->i_completed_io_lock, flags); if (list_empty(&ei->i_rsv_conversion_list)) diff --git a/include/trace/events/ext4.h b/include/trace/events/ext4.h index 4539ef8e5f86..c9259c2a3e36 100644 --- a/include/trace/events/ext4.h +++ b/include/trace/events/ext4.h @@ -3237,6 +3237,113 @@ DEFINE_EVENT(ext4_iomap_disksize_pending, ext4_ioma= p_wait_disksize_pending, TP_ARGS(inode) ); =20 +/* disksize pending I/O tracepoints for iomap Buffered I/O path */ +DECLARE_EVENT_CLASS(ext4_iomap_wb_disksize_pending, + TP_PROTO(struct inode *inode, loff_t io_offset, size_t io_size), + TP_ARGS(inode, io_offset, io_size), + TP_STRUCT__entry( + __field(dev_t, dev) + __field(u64, ino) + __field(loff_t, io_offset) + __field(size_t, io_size) + __field(loff_t, i_size) + __field(loff_t, i_disksize) + ), + TP_fast_assign( + __entry->dev =3D inode->i_sb->s_dev; + __entry->ino =3D inode->i_ino; + __entry->io_offset =3D io_offset; + __entry->io_size =3D io_size; + __entry->i_size =3D i_size_read(inode); + __entry->i_disksize =3D READ_ONCE(EXT4_I(inode)->i_disksize); + ), + TP_printk("dev %d:%d ino %llu io_offset %lld io_size %zu i_size %lld i_di= sksize %lld", + MAJOR(__entry->dev), MINOR(__entry->dev), + __entry->ino, __entry->io_offset, __entry->io_size, + __entry->i_size, __entry->i_disksize) +); + +DEFINE_EVENT(ext4_iomap_wb_disksize_pending, + ext4_iomap_wb_disksize_pending_submit, + TP_PROTO(struct inode *inode, loff_t io_offset, size_t io_size), + TP_ARGS(inode, io_offset, io_size) +); + +DEFINE_EVENT(ext4_iomap_wb_disksize_pending, + ext4_iomap_wb_disksize_pending_complete, + TP_PROTO(struct inode *inode, loff_t io_offset, size_t io_size), + TP_ARGS(inode, io_offset, io_size) +); + +DEFINE_EVENT(ext4_iomap_wb_disksize_pending, + ext4_iomap_wb_disksize_pending_wait, + TP_PROTO(struct inode *inode, loff_t io_offset, size_t io_size), + TP_ARGS(inode, io_offset, io_size) +); + +/* i_disksize update tracepoint */ +TRACE_EVENT(ext4_iomap_wb_update_disksize, + TP_PROTO(struct inode *inode, loff_t end, loff_t i_size, + loff_t i_disksize, loff_t new_disksize, + bool is_disksize_grow), + TP_ARGS(inode, end, i_size, i_disksize, new_disksize, + is_disksize_grow), + TP_STRUCT__entry( + __field(dev_t, dev) + __field(u64, ino) + __field(loff_t, end) + __field(loff_t, i_size) + __field(loff_t, i_disksize) + __field(loff_t, new_disksize) + __field(bool, is_disksize_grow) + ), + TP_fast_assign( + __entry->dev =3D inode->i_sb->s_dev; + __entry->ino =3D inode->i_ino; + __entry->end =3D end; + __entry->i_size =3D i_size; + __entry->i_disksize =3D i_disksize; + __entry->new_disksize =3D new_disksize; + __entry->is_disksize_grow =3D is_disksize_grow; + ), + TP_printk("dev %d:%d ino %llu end %lld i_size %lld i_disksize %lld new_di= sksize %lld is_disksize_grow %d", + MAJOR(__entry->dev), MINOR(__entry->dev), + __entry->ino, __entry->end, __entry->i_size, + __entry->i_disksize, __entry->new_disksize, + __entry->is_disksize_grow) +); + +/* Block zero EOF tracepoint */ +TRACE_EVENT(ext4_block_zero_eof, + TP_PROTO(struct inode *inode, loff_t from, loff_t length, + bool did_zero, bool zero_written), + TP_ARGS(inode, from, length, did_zero, zero_written), + TP_STRUCT__entry( + __field(dev_t, dev) + __field(u64, ino) + __field(loff_t, from) + __field(loff_t, length) + __field(loff_t, i_size) + __field(loff_t, i_disksize) + __field(bool, did_zero) + __field(bool, zero_written) + ), + TP_fast_assign( + __entry->dev =3D inode->i_sb->s_dev; + __entry->ino =3D inode->i_ino; + __entry->from =3D from; + __entry->length =3D length; + __entry->i_size =3D inode->i_size; + __entry->i_disksize =3D READ_ONCE(EXT4_I(inode)->i_disksize); + __entry->did_zero =3D did_zero; + __entry->zero_written =3D zero_written; + ), + TP_printk("dev %d:%d ino %llu zero EOF from %lld length %lld i_size %lld = i_disksize %lld did_zero %d zero_written %d", + MAJOR(__entry->dev), MINOR(__entry->dev), __entry->ino, + __entry->from, __entry->length, __entry->i_size, + __entry->i_disksize, __entry->did_zero, __entry->zero_written) +); + #endif /* _TRACE_EXT4_H */ =20 /* This part must be outside protection */ --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 946C739F18F; Fri, 14 Aug 2026 09:52:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786701172; cv=none; b=jOsY6Z6ASZZWAHvwcFy1FxRA/QmlluJXF6mWuzQ4/PIPsgI7biWgtPirlAzx03W9qvWkonUojorbU8i/ADLB/d7SHV4MC3RudSyqUVY2lp/Jt9Dx8pR00ymfwAQxuPxSmP+XUPwNS35vZLaek3h0w7IRELBuf3AxJbdaII+bTl8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786701172; c=relaxed/simple; bh=zl0YNsjvNtZsYQQeWC3ej6cBLjiM2wbGHfmW9HT+nL0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sNuwfjY2WdNlZGc8alAgriPI5pJm7H72sdKYO7Jv3CptEbpWjm8EoGAGW20OY0T/c7lOKaZDDRnPrSZ2GIRAoy4eTXrzN2EYsMZGzS/5UlRCj238WUNtByvGNXpHc+UBWoovNZAjjoQm7t9LN0zuJ+R1rUOoDcaHIaayHpa3l9M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.177]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLyFR75jNzYQvWH; Fri, 14 Aug 2026 17:52:27 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 37A224058C; Fri, 14 Aug 2026 17:52:36 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBXR0ZK5X5qbKAYCQ--.57632S4; Fri, 14 Aug 2026 17:52:35 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 31/32] ext4: partially enable iomap for the buffered I/O path of regular files Date: Fri, 14 Aug 2026 17:46:15 +0800 Message-ID: <20260814094616.1710143-1-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBXR0ZK5X5qbKAYCQ--.57632S4 X-Coremail-Antispam: 1UD129KBjvJXoW3tw1DZF1UKF45GF1UAr43Jrb_yoWDXw1Dpr ZxK34rGr1DX34v9ws7tr4DXr1Yv3WxK3yUGrZ3ur4kAa98Jw1IqFyjyF1YvF15JrZ3Ww12 qF48tw1Uuw1qkrDanT9S1TB71UUUUUJqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUB014x267AKxVW5JVWrJwAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2ocxC64kIII0Yj41l84x0c7CEw4AK67xGY2AK02 1l84ACjcxK6xIIjxv20xvE14v26ryj6F1UM28EF7xvwVC0I7IYx2IY6xkF7I0E14v26F4U JVW0owA2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_Gc CE3s1ln4kS14v26r1Y6r17M2AIxVAIcxkEcVAq07x20xvEncxIr21l5I8CrVACY4xI64kE 6c02F40Ex7xfMcIj6xIIjxv20xvE14v26r1Y6r17McIj6I8E87Iv67AKxVW8JVWxJwAm72 CE4IkC6x0Yz7v_Jr0_Gr1lF7xvr2IYc2Ij64vIr41lF7I21c0EjII2zVCS5cI20VAGYxC7 M4IIrI8v6xkF7I0E8cxan2IY04v7MxkF7I0En4kS14v26r4a6rW5MxAIw28IcxkI7VAKI4 8JMxC20s026xCaFVCjc4AY6r1j6r4UMI8I3I0E5I8CrVAFwI0_Jr0_Jr4lx2IqxVCjr7xv wVAFwI0_JrI_JrWlx4CE17CEb7AF67AKxVW8ZVWrXwCIc40Y0x0EwIxGrwCI42IY6xIIjx v20xvE14v26ryj6F1UMIIF0xvE2Ix0cI8IcVCY1x0267AKxVWxJr0_GcWlIxAIcVCF04k2 6cxKx2IYs7xG6r1j6r1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I 0E14v26F4UJVW0obIYCTnIWIevJa73UjIFyTuYvjTRXku4DUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Introduce ext4_enable_buffered_iomap() to determine whether a regular file inode should use the iomap buffered I/O path. We now support the default filesystem features, mount options, and the bigalloc feature. However, inline data, fsverity, fscrypt, indirect inode type, and data=3Djournal mode are not fully supported. The decision is made at inode initialization time in __ext4_new_inode() and __ext4_iget() by setting the EXT4_STATE_BUFFERED_IOMAP state flag. If any of these unsupported features are met, the inode silently falls back to the traditional buffer_head path. Switching the buffered I/O path on an active inode is not supported, with the exception of changing a per-inode journal flag. For features like encryption, verity, and inline data that can be dynamically enabled at the superblock level, checking the global feature flag avoids the complexity of toggling the path on individual inodes. Additionally: - Extend ext4_inode_journal_mode() to force ordered mode for inodes using the iomap path under a data=3Djournal mount. For the global data journal mode (EXT4_MOUNT_JOURNAL_DATA), dynamic enablement is deferred until the next inode re-initialization. For the per-inode data journal mode (EXT4_INODE_JOURNAL_DATA), dynamic changes take effect immediately, as it is safe to switch address_space operations and drop all page cache under i_rwsem and filemap_invalidate_lock. - Add WARN_ON_ONCE() guards in _ext4_get_block() and ext4_do_writepages() to catch inodes using the iomap path from accidentally entering the legacy buffer_head writeback path. - Reject extent-to-indirect migration via ext4_ind_migrate() for inodes on the iomap path. Signed-off-by: Zhang Yi --- fs/ext4/ext4.h | 1 + fs/ext4/ext4_jbd2.c | 8 +++- fs/ext4/ialloc.c | 1 + fs/ext4/inode.c | 99 ++++++++++++++++++++++++++++++++++++++++++++- fs/ext4/migrate.c | 2 + 5 files changed, 107 insertions(+), 4 deletions(-) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 504dce9fdc6b..405e256a1802 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -3170,6 +3170,7 @@ int ext4_walk_page_buffers(handle_t *handle, int do_journal_get_write_access(handle_t *handle, struct inode *inode, struct buffer_head *bh); void ext4_set_inode_mapping_order(struct inode *inode); +void ext4_enable_buffered_iomap(struct inode *inode); int ext4_nonda_switch(struct super_block *sb); #define FALL_BACK_TO_NONDELALLOC 1 #define EXT4_WRITE_DATA_INLINE 2 diff --git a/fs/ext4/ext4_jbd2.c b/fs/ext4/ext4_jbd2.c index 53ddedb52a6f..a4664ddecdcd 100644 --- a/fs/ext4/ext4_jbd2.c +++ b/fs/ext4/ext4_jbd2.c @@ -17,8 +17,12 @@ int ext4_inode_journal_mode(struct inode *inode) test_opt(inode->i_sb, DATA_FLAGS) =3D=3D EXT4_MOUNT_JOURNAL_DATA || (ext4_test_inode_flag(inode, EXT4_INODE_JOURNAL_DATA) && !test_opt(inode->i_sb, DELALLOC))) { - /* We do not support data journalling for encrypted data */ - if (S_ISREG(inode->i_mode) && IS_ENCRYPTED(inode)) + /* + * We do not support data journalling for encrypted data + * and buffered IOMAP path. + */ + if (S_ISREG(inode->i_mode) && + (IS_ENCRYPTED(inode) || ext4_inode_buffered_iomap(inode))) return EXT4_INODE_ORDERED_DATA_MODE; /* ordered */ return EXT4_INODE_JOURNAL_DATA_MODE; /* journal data */ } diff --git a/fs/ext4/ialloc.c b/fs/ext4/ialloc.c index a5831fc536db..f97a2f4904eb 100644 --- a/fs/ext4/ialloc.c +++ b/fs/ext4/ialloc.c @@ -1346,6 +1346,7 @@ struct inode *__ext4_new_inode(struct mnt_idmap *idma= p, } } =20 + ext4_enable_buffered_iomap(inode); ext4_set_inode_mapping_order(inode); =20 ext4_update_inode_fsync_trans(handle, inode, 1); diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index ee15366422a1..c9ee78fba4d0 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -1034,6 +1034,9 @@ static int _ext4_get_block(struct inode *inode, secto= r_t iblock, =20 if (ext4_has_inline_data(inode)) return -ERANGE; + /* inode using the iomap buffered I/O path should not go here. */ + if (WARN_ON_ONCE(ext4_inode_buffered_iomap(inode))) + return -EINVAL; =20 map.m_lblk =3D iblock; map.m_len =3D bh->b_size >> inode->i_blkbits; @@ -2904,6 +2907,12 @@ static int ext4_do_writepages(struct mpage_da_data *= mpd) if (!mapping->nrpages || !mapping_tagged(mapping, PAGECACHE_TAG_DIRTY)) goto out_writepages; =20 + /* inode using the iomap buffered I/O path should not go here. */ + if (WARN_ON_ONCE(ext4_inode_buffered_iomap(inode))) { + ret =3D -EINVAL; + goto out_writepages; + } + /* * If the filesystem has aborted, it is read-only, so return * right away instead of dumping stack traces later on that @@ -4042,6 +4051,9 @@ static int ext4_iomap_map_blocks(struct inode *inode,= loff_t offset, { u8 blkbits =3D inode->i_blkbits; =20 + /* inode using the buffer_head buffered I/O path should not go here. */ + if (WARN_ON_ONCE(!ext4_inode_buffered_iomap(inode))) + return -EINVAL; if ((offset >> blkbits) > EXT4_MAX_LOGICAL_BLOCK) return -EINVAL; =20 @@ -4467,6 +4479,10 @@ static int ext4_iomap_writepages(struct address_spac= e *mapping, .ops =3D &ext4_writeback_ops, }; =20 + /* inode using the buffer_head buffered I/O path should not go here. */ + if (WARN_ON_ONCE(!ext4_inode_buffered_iomap(inode))) + return -EINVAL; + ret =3D ext4_emergency_state(sb); if (unlikely(ret)) { /* @@ -6037,6 +6053,81 @@ static int check_igot_inode(struct inode *inode, ext= 4_iget_flags flags, return -EFSCORRUPTED; } =20 +/* + * Determine whether an inode should use the iomap buffered I/O path. + * EXT4_STATE_BUFFERED_IOMAP is generally set at inode initialization + * time. Online switching of the buffered I/O path on an active inode is + * NOT supported, with the exception of changing a per-inode journal + * flag. + * + * For features like inline data, fsverity, and encryption that can be + * dynamically enabled or disabled, we check the superblock-level + * feature flags. If any of these is globally enabled, no inode is + * allowed into the iomap buffered I/O path. This avoids the complexity + * of dynamic toggling. + * + * For the global data journal mode (EXT4_MOUNT_JOURNAL_DATA), dynamic + * change through remount is deferred. It will only become available + * after the inode is re-initialized (i.e., after the last reference + * drops and the inode is re-read from disk with the journal flag + * cleared). + * + * For the per-inode data journal mode (EXT4_INODE_JOURNAL_DATA), + * dynamic changes take effect immediately. This is safe because + * address_space operations can be switched and all page cache can be + * dropped under i_rwsem and filemap_invalidate_lock. + * + * For extent-to-indirect block migration (via EXT4_IOC_SETFLAGS + * clearing EXT4_EXTENTS_FL), this operation is directly rejected for + * inodes using the iomap path. + */ +void ext4_enable_buffered_iomap(struct inode *inode) +{ + struct super_block *sb =3D inode->i_sb; + + if (!S_ISREG(inode->i_mode)) + return; + if (ext4_test_inode_flag(inode, EXT4_INODE_EA_INODE)) + return; + + /* Unsupported Features */ + if (ext4_has_feature_inline_data(sb)) + return; + if (ext4_has_feature_verity(sb)) + return; + if (ext4_has_feature_encrypt(sb)) + return; + if (test_opt(sb, DATA_FLAGS) =3D=3D EXT4_MOUNT_JOURNAL_DATA || + ext4_test_inode_flag(inode, EXT4_INODE_JOURNAL_DATA)) + return; + if (!(ext4_test_inode_flag(inode, EXT4_INODE_EXTENTS))) + return; + + ext4_set_inode_state(inode, EXT4_STATE_BUFFERED_IOMAP); + + /* + * Install the iomap end_io handler on the shared conversion + * work. This is safe at inode initialization and during the + * buffered I/O path changes where we flush all pending + * writebacks and drop page cache under i_rwsem and + * filemap_invalidate_lock. + */ + INIT_WORK(&EXT4_I(inode)->i_rsv_conversion_work, ext4_iomap_end_io); +} + +static void ext4_disable_buffered_iomap(struct inode *inode) +{ + ext4_clear_inode_state(inode, EXT4_STATE_BUFFERED_IOMAP); + + /* + * Reinstall the buffer_head end_io handler on the shared + * conversion work. This is safe during the buffered I/O path + * changes where we flush all pending writebacks and drop page + * cache under i_rwsem and filemap_invalidate_lock. + */ + INIT_WORK(&EXT4_I(inode)->i_rsv_conversion_work, ext4_end_io_rsv_work); +} + void ext4_set_inode_mapping_order(struct inode *inode) { struct super_block *sb =3D inode->i_sb; @@ -6351,6 +6442,8 @@ struct inode *__ext4_iget(struct super_block *sb, uns= igned long ino, if (ret) goto bad_inode; =20 + ext4_enable_buffered_iomap(inode); + if (S_ISREG(inode->i_mode)) { inode->i_op =3D &ext4_file_inode_operations; inode->i_fop =3D &ext4_file_operations; @@ -7574,9 +7667,10 @@ int ext4_change_inode_journal_flag(struct inode *ino= de, int val) * the inode's in-core data-journaling state flag now. */ =20 - if (val) + if (val) { ext4_set_inode_flag(inode, EXT4_INODE_JOURNAL_DATA); - else { + ext4_disable_buffered_iomap(inode); + } else { err =3D jbd2_journal_flush(journal, 0); if (err < 0) { jbd2_journal_unlock_updates(journal); @@ -7585,6 +7679,7 @@ int ext4_change_inode_journal_flag(struct inode *inod= e, int val) return err; } ext4_clear_inode_flag(inode, EXT4_INODE_JOURNAL_DATA); + ext4_enable_buffered_iomap(inode); } ext4_set_aops(inode); ext4_set_inode_mapping_order(inode); diff --git a/fs/ext4/migrate.c b/fs/ext4/migrate.c index 5d60ef10fe11..09931d3ba2c6 100644 --- a/fs/ext4/migrate.c +++ b/fs/ext4/migrate.c @@ -621,6 +621,8 @@ int ext4_ind_migrate(struct inode *inode) =20 if (ext4_has_feature_bigalloc(inode->i_sb)) return -EOPNOTSUPP; + if (ext4_inode_buffered_iomap(inode)) + return -EOPNOTSUPP; =20 /* * In order to get correct extent info, force all delayed allocation --=20 2.52.0 From nobody Tue Sep 29 00:36:44 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 438AC3C0A13; Fri, 14 Aug 2026 09:52:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786701171; cv=none; b=TvOI3KagaejfnoSM5KeWU5kkGYxxKJY1Fpng9VJuDYoGkSiUe5JirkWVk7TxsaLds4m5EXnOpS7Dnf89gs59INn/4QsKRBguJrWdN6tzPoJW4TMnwp4twrf0ZulKCN3F8OtZLO+0Fwv/PuYwZ2IjWXE855bTqWFlYPTxa9ZEMZg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786701171; c=relaxed/simple; bh=CSqV7drS2LhKrxjqgy2PweyXzpaQs16hCrIFJeznxEQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=HQyKUUuclWJG4sZ11QNL2zFPkJEvPSrAGqAoc1gGfhiAYEiy3y4M1TJ9pnYgYY7rt0Vq8TQEOP+9UjapRxd6CD14+TStwTlSBFnd0FVLzEPJd1Px1qPfrFu9qB/0JY6XwgFKlE+Qx7OovTq8tUH30NatwrKYDyeCDvcgstE5qgU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.177]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hLyFS0YxgzYQvWv; Fri, 14 Aug 2026 17:52:28 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.252]) by mail.maildlp.com (Postfix) with ESMTP id 4B8E14058D; Fri, 14 Aug 2026 17:52:36 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP3 (Coremail) with UTF8SMTPSA id _Ch0CgBXR0ZK5X5qbKAYCQ--.57632S5; Fri, 14 Aug 2026 17:52:36 +0800 (CST) From: Zhang Yi To: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org Cc: linux-kernel@vger.kernel.org, tytso@mit.edu, adilger.kernel@dilger.ca, libaokun@linux.alibaba.com, jack@suse.cz, ojaswin@linux.ibm.com, ritesh.list@gmail.com, djwong@kernel.org, hch@infradead.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, chengzhihao1@huawei.com, yangerkun@huawei.com, yukuai@fnnas.com Subject: [PATCH -next v5 32/32] ext4: introduce a mount option for iomap buffered I/O path Date: Fri, 14 Aug 2026 17:46:16 +0800 Message-ID: <20260814094616.1710143-2-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> References: <20260814093331.1703882-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: _Ch0CgBXR0ZK5X5qbKAYCQ--.57632S5 X-Coremail-Antispam: 1UD129KBjvJXoWxWr47Kr15Gr1xZry3ZFy7KFg_yoWrGr1xpr W5KFyrKr1kXryF9w48uF4kXr1Yy3Zaka1UCrZYgr47Ga9rAryIqFyfKF13AFWagrW8X340 qF1rWw17Wa13CrDanT9S1TB71UUUUUDqnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmlb4IE77IF4wAFF20E14v26rWj6s0DM7CY07I20VC2zVCF04k2 6cxKx2IYs7xG6rWj6s0DM7CIcVAFz4kK6r1j6r18M28IrcIa0xkI8VA2jI8067AKxVWUGw A2048vs2IY020Ec7CjxVAFwI0_Gr0_Xr1l8cAvFVAK0II2c7xJM28CjxkF64kEwVA0rcxS w2x7M28EF7xvwVC0I7IYx2IY67AKxVW5JVW7JwA2z4x0Y4vE2Ix0cI8IcVCY1x0267AKxV WxJr0_GcWl84ACjcxK6I8E87Iv67AKxVWxJr0_GcWl84ACjcxK6I8E87Iv6xkF7I0E14v2 6rxl6s0DM2AIxVAIcxkEcVAq07x20xvEncxIr21l5I8CrVACY4xI64kE6c02F40Ex7xfMc Ij6xIIjxv20xvE14v26r1Y6r17McIj6I8E87Iv67AKxVW8JVWxJwAm72CE4IkC6x0Yz7v_ Jr0_Gr1lF7xvr2IYc2Ij64vIr41lF7I21c0EjII2zVCS5cI20VAGYxC7M4IIrI8v6xkF7I 0E8cxan2IY04v7MxkF7I0En4kS14v26r4a6rW5MxAIw28IcxkI7VAKI48JMxC20s026xCa FVCjc4AY6r1j6r4UMI8I3I0E5I8CrVAFwI0_Jr0_Jr4lx2IqxVCjr7xvwVAFwI0_JrI_Jr Wlx4CE17CEb7AF67AKxVW8ZVWrXwCIc40Y0x0EwIxGrwCI42IY6xIIjxv20xvE14v26ryj 6F1UMIIF0xvE2Ix0cI8IcVCY1x0267AKxVWxJr0_GcWlIxAIcVCF04k26cxKx2IYs7xG6r 1j6r1xMIIF0xvEx4A2jsIE14v26F4j6r4UJwCI42IY6I8E87Iv6xkF7I0E14v26F4UJVW0 obIYCTnIWIevJa73UjIFyTuYvjTRvZXoUUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi Since the iomap buffered I/O path does not yet support all existing ext4 features, it cannot be enabled by default. Introduce the 'buffered_iomap' and 'nobuffered_iomap' mount options to explicitly enable or disable the iomap buffered I/O path for regular files. Toggling this option via remount is allowed. The change of I/O path will not take effect immediately. It will be deferred. The new setting will only take effect after the inode is re-initialized (i.e., after the last reference is dropped and the inode is re-read from disk). Signed-off-by: Zhang Yi --- fs/ext4/ext4.h | 1 + fs/ext4/inode.c | 6 ++++++ fs/ext4/super.c | 7 +++++++ 3 files changed, 14 insertions(+) diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 405e256a1802..8cfe77b45e21 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -1321,6 +1321,7 @@ struct ext4_inode_info { * scanning in mballoc */ #define EXT4_MOUNT2_ABORT 0x00000100 /* Abort filesystem */ +#define EXT4_MOUNT2_BUFFERED_IOMAP 0x00000200 /* Use iomap for buffered I/= O */ =20 #define clear_opt(sb, opt) EXT4_SB(sb)->s_mount_opt &=3D \ ~EXT4_MOUNT_##opt diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index c9ee78fba4d0..c7466ecc44c2 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -6080,11 +6080,17 @@ static int check_igot_inode(struct inode *inode, ex= t4_iget_flags flags, * For extent-to-indirect block migration (via EXT4_IOC_SETFLAGS * clearing EXT4_EXTENTS_FL), this operation is directly rejected for * inodes using the iomap path. + * + * When remounting to toggle the buffered_iomap mount option, the change + * of I/O path is deferred as well, it will be available after the inode + * is re-initialized. */ void ext4_enable_buffered_iomap(struct inode *inode) { struct super_block *sb =3D inode->i_sb; =20 + if (!test_opt2(sb, BUFFERED_IOMAP)) + return; if (!S_ISREG(inode->i_mode)) return; if (ext4_test_inode_flag(inode, EXT4_INODE_EA_INODE)) diff --git a/fs/ext4/super.c b/fs/ext4/super.c index 86ed5228dbe9..0855f9801df4 100644 --- a/fs/ext4/super.c +++ b/fs/ext4/super.c @@ -1746,6 +1746,7 @@ enum { Opt_discard, Opt_nodiscard, Opt_init_itable, Opt_noinit_itable, Opt_max_dir_size_kb, Opt_nojournal_checksum, Opt_nombcache, Opt_no_prefetch_block_bitmaps, Opt_mb_optimize_scan, + Opt_buffered_iomap, Opt_nobuffered_iomap, Opt_errors, Opt_data, Opt_data_err, Opt_jqfmt, Opt_dax_type, #ifdef CONFIG_EXT4_DEBUG Opt_fc_debug_max_replay, Opt_fc_debug_force @@ -1884,6 +1885,8 @@ static const struct fs_parameter_spec ext4_param_spec= s[] =3D { fsparam_flag ("no_prefetch_block_bitmaps", Opt_no_prefetch_block_bitmaps), fsparam_s32 ("mb_optimize_scan", Opt_mb_optimize_scan), + fsparam_flag ("buffered_iomap", Opt_buffered_iomap), + fsparam_flag ("nobuffered_iomap", Opt_nobuffered_iomap), fsparam_string ("check", Opt_removed), /* mount option from ext2/3 */ fsparam_flag ("nocheck", Opt_removed), /* mount option from ext2/3 */ fsparam_flag ("reservation", Opt_removed), /* mount option from ext2/3 */ @@ -1977,6 +1980,10 @@ static const struct mount_opts { {Opt_nombcache, EXT4_MOUNT_NO_MBCACHE, MOPT_SET}, {Opt_no_prefetch_block_bitmaps, EXT4_MOUNT_NO_PREFETCH_BLOCK_BITMAPS, MOPT_SET}, + {Opt_buffered_iomap, EXT4_MOUNT2_BUFFERED_IOMAP, + MOPT_SET | MOPT_2 | MOPT_EXT4_ONLY}, + {Opt_nobuffered_iomap, EXT4_MOUNT2_BUFFERED_IOMAP, + MOPT_CLEAR | MOPT_2 | MOPT_EXT4_ONLY}, #ifdef CONFIG_EXT4_DEBUG {Opt_fc_debug_force, EXT4_MOUNT2_JOURNAL_FAST_COMMIT, MOPT_SET | MOPT_2 | MOPT_EXT4_ONLY}, --=20 2.52.0