From nobody Fri Sep 25 06:04:07 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D85CD4A8A3D; Wed, 16 Sep 2026 09:32:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789551198; cv=none; b=oTb2dBZbCAE/B/x8shq6yhJJREpyJ1RySAvsZAvzzRkYcL+2C+mzohwLJSAGBnun/Av8/CpxgA35kBKnTve8JTmJwOxTR+8Qy3TcuKxhJEw9iZmhsqnXLwQ4NP4onulYVAQC5pZ11m7MFLx7IhbTlf2qKR9VO0fjRhUgwJgsSSQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789551198; c=relaxed/simple; bh=T7YYHppyuAuUvkLx3+mSDaSFFlYvvU7NkTEotrRDTM0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aLPmsOyuFN0suZ95kPSLPlk8H7qcF/arWjEVxKLjoH/Mn+wP3LQECZtA7wDznIeErndPP3ltn7vO5eKtvpYeg+dSmPDGdS7jk9NGycxBpedY43+cGlV+YazY4wb6dLuBziuS4Nfyf/bNFXNPguOdsVl/Q1qs1Pj0sFxUep+QNb4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hlDFD4WdLzKHMcS; Wed, 16 Sep 2026 17:32:32 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.128]) by mail.maildlp.com (Postfix) with ESMTP id C90CD4056D; Wed, 16 Sep 2026 17:32:46 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP4 (Coremail) with UTF8SMTPSA id gCh0CgCHVikvYqpq8i0GAg--.65015S5; Wed, 16 Sep 2026 17:32:46 +0800 (CST) From: Zhang Yi To: linux-mm@kvack.org Cc: linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-ext4@vger.kernel.org, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, hughd@google.com, baolin.wang@linux.alibaba.com, willy@infradead.org, jack@suse.cz, ziy@nvidia.com, bfoster@redhat.com, joannelkoong@gmail.com, djwong@kernel.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, yangerkun@huawei.com, chengzhihao1@huawei.com, wangkefeng.wang@huawei.com, yukuai@fnnas.com Subject: [PATCH v3 1/3] mm/truncate: fix data loss when splitting straddling large folios fails Date: Wed, 16 Sep 2026 17:24:48 +0800 Message-ID: <20260916092450.654408-2-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260916092450.654408-1-yi.zhang@huaweicloud.com> References: <20260916092450.654408-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: gCh0CgCHVikvYqpq8i0GAg--.65015S5 X-Coremail-Antispam: 1UD129KBjvJXoW3ur45WrWUWr48XryrCry8Xwb_yoWDtr48pa yUK3sxKrZ8Ww4jkrnrXa1UZw4Yyas3XFWUAFyxGwnxAan0qwnrKF1Ut3W8KFW3J3s7A34F qF1jyay7WF1UJFJanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmY14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_Jr4l82xGYIkIc2 x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2z4x0 Y4vE2Ix0cI8IcVAFwI0_Jr0_JF4l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Cr0_Gr1UM2 8EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq3wAS 0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7IYx2 IY67AKxVWUJVWUGwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4UM4x0 Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2kIc2 xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkEbVWU JVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67AF67 kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVWUJVWUCwCI42IY 6xIIjxv20xvEc7CjxVAFwI0_Cr0_Gr1UMIIF0xvE42xK8VAvwI8IcIk0rVWUJVWUCwCI42 IY6I8E87Iv67AKxVWUJVW8JwCI42IY6I8E87Iv6xkF7I0E14v26r4j6r4UJbIYCTnIWIev Ja73UjIFyTuYvjTRMfOzDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi truncate_inode_partial_folio() splits a large folio so that the caller's truncate loop can drop the in-range sub-folios while keeping the out-of-range tail. The first split at the punch start edge is non-uniform, which leaves the sub-folio at the truncation end edge as large as possible, this means it may still straddle the range, holding both zeroed in-range and valid out-of-range data. The function then attempts a second split at offset + length to isolate that tail. If the second split fails the straddling sub-folio stays merged. The function returned true unconditionally on all exit paths of the success block, telling the caller it was fully handled. The caller kept its default end and the truncate loop truncated every sub-folio below it, including the merged straddler, discarding the valid out-of-range tail. For example, a 4-page order-2 folio punched from offset 0 to the middle of the last page: truncate_inode_pages_range() truncate_inode_partial_folio() # same_folio =3D=3D true 1st split at page0 -> [p0, p1, p2-3] # non-uniform, success folio2 =3D p2-3 # straddles: p2 zeroed, p3 tail valid 2nd split of folio2 fails / cannot lock return true # BUG: caller keeps default end end =3D 3 loop truncates p0, p1, p2-3 # p3's valid tail is lost This became reachable after commit 7460b470a131 ("mm/truncate: use folio_split() in truncate operation") replaced the atomic split_folio() with folio_split(), whose non-uniform split can partially split a folio and leave the end edge merged. It has gone unnoticed because a dirty large folio normally carries the filesystem's private data, for example buffer_head, so filemap_release_folio() -> iomap_release_folio() returns false on a dirty folio and folio_split() aborts with -EBUSY before any split, leaving the straddler safely unsplit. The bug is only reachable on paths that produce dirty large folios without filesystem private data, and it was caught on the upcoming ext4 iomap buffered I/O path when no ifs is attached. Rework the contract so the caller is told the page range to discard: - Add pgoff_t *pstart and *pend out-parameters that receive the page range fully covered by [lstart, lend] after any split (or none), i.e. the pages wholly within the range and safe to discard. - Adjust the ordering of the validate check when splitting folio2. folio2->index is only reliable after the reference count and lock have been successfully acquired, since it may have been split concurrently, or freed and recycled to an unrelated mapping. On any failure to obtain a reliable end position, fall back to folio->index, which is safe but leaves the sub-folios split off at the offset edge in the page cache. - Rename the byte-range parameters start/end to lstart/lend to better express their semantics. Callers in truncate_inode_pages_range() and shmem_undo_range() pass &pstart for the folio at the start edge and &pend for the folio at the end edge, so the truncate loop drops exactly the fully covered pages and never touches a straddling folio that still holds valid out-of-range data. Suggested-by: Brian Foster Link: https://lore.kernel.org/linux-fsdevel/anH-WKA1coW6wtfG@bfoster/ Fixes: 7460b470a131 ("mm/truncate: use folio_split() in truncate operation") Signed-off-by: Zhang Yi --- mm/internal.h | 4 +-- mm/shmem.c | 13 +++----- mm/truncate.c | 89 ++++++++++++++++++++++++++++++++++++--------------- 3 files changed, 71 insertions(+), 35 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index 68db5abd0a4c..6e6ad3187378 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -627,8 +627,8 @@ unsigned find_lock_entries(struct address_space *mappin= g, pgoff_t *start, unsigned find_get_entries(struct address_space *mapping, pgoff_t *start, pgoff_t end, struct folio_batch *fbatch, pgoff_t *indices); int truncate_inode_folio(struct address_space *mapping, struct folio *foli= o); -bool truncate_inode_partial_folio(struct folio *folio, loff_t start, - loff_t end); +bool truncate_inode_partial_folio(struct folio *folio, loff_t lstart, + loff_t lend, pgoff_t *pstart, pgoff_t *pend); long mapping_evict_folio(struct address_space *mapping, struct folio *foli= o); unsigned long mapping_try_invalidate(struct address_space *mapping, pgoff_t start, pgoff_t end, unsigned long *nr_failed); diff --git a/mm/shmem.c b/mm/shmem.c index 89a1495e55f7..712dc3effe02 100644 --- a/mm/shmem.c +++ b/mm/shmem.c @@ -1175,11 +1175,8 @@ static void shmem_undo_range(struct inode *inode, lo= ff_t lstart, uoff_t lend, if (folio) { same_folio =3D lend < folio_next_pos(folio); folio_mark_dirty(folio); - if (!truncate_inode_partial_folio(folio, lstart, lend)) { - start =3D folio_next_index(folio); - if (same_folio) - end =3D folio->index; - } + truncate_inode_partial_folio(folio, lstart, lend, &start, + same_folio ? &end : NULL); folio_unlock(folio); folio_put(folio); folio =3D NULL; @@ -1189,8 +1186,7 @@ static void shmem_undo_range(struct inode *inode, lof= f_t lstart, uoff_t lend, folio =3D shmem_get_partial_folio(inode, lend >> PAGE_SHIFT); if (folio) { folio_mark_dirty(folio); - if (!truncate_inode_partial_folio(folio, lstart, lend)) - end =3D folio->index; + truncate_inode_partial_folio(folio, lstart, lend, NULL, &end); folio_unlock(folio); folio_put(folio); } @@ -1258,7 +1254,8 @@ static void shmem_undo_range(struct inode *inode, lof= f_t lstart, uoff_t lend, =20 if (!folio_test_large(folio)) { truncate_inode_folio(mapping, folio); - } else if (truncate_inode_partial_folio(folio, lstart, lend)) { + } else if (truncate_inode_partial_folio(folio, + lstart, lend, NULL, NULL)) { /* * If we split a page, reset the loop so * that we pick up the new sub pages. diff --git a/mm/truncate.c b/mm/truncate.c index b58ba940be47..bec6d881d022 100644 --- a/mm/truncate.c +++ b/mm/truncate.c @@ -206,15 +206,21 @@ static int folio_split_or_unmap(struct folio *folio, = struct page *split_at, /* * Handle partial folios. The folio may be entirely within the * range if a split has raced with us. If not, we zero the part of the - * folio that's within the [start, end] range, and then split the folio if + * folio that's within the [lstart, lend] range, and then split the folio = if * it's large. split_page_range() will discard pages which now lie beyond * i_size, and we rely on the caller to discard pages which lie within a * newly created hole. * + * When @pstart and/or @pend are non-NULL they receive the indexes of the + * page range fully covered by [lstart, lend] after any split (or none), + * i.e. the range of pages wholly within [lstart, lend] and so safe to + * discard. + * * Returns false if splitting failed so the caller can avoid * discarding the entire folio which is stubbornly unsplit. */ -bool truncate_inode_partial_folio(struct folio *folio, loff_t start, loff_= t end) +bool truncate_inode_partial_folio(struct folio *folio, loff_t lstart, + loff_t lend, pgoff_t *pstart, pgoff_t *pend) { loff_t pos =3D folio_pos(folio); size_t size =3D folio_size(folio); @@ -222,14 +228,20 @@ bool truncate_inode_partial_folio(struct folio *folio= , loff_t start, loff_t end) struct page *split_at, *split_at2; unsigned int min_order; =20 - if (pos < start) - offset =3D start - pos; + if (pos < lstart) + offset =3D lstart - pos; else offset =3D 0; - if (pos + size <=3D (u64)end) + if (pos + size <=3D (u64)lend) length =3D size - offset; else - length =3D end + 1 - pos - offset; + length =3D lend + 1 - pos - offset; + + if (pstart) + *pstart =3D offset ? folio_next_index(folio) : folio->index; + if (pend) + *pend =3D (pos + size > (u64)lend) ? folio->index : + folio_next_index(folio); =20 folio_wait_writeback(folio); if (length =3D=3D size) { @@ -259,32 +271,62 @@ bool truncate_inode_partial_folio(struct folio *folio= , loff_t start, loff_t end) * for shmem truncate */ struct folio *folio2; + pgoff_t end, aligned_end =3D (pos + offset + length) >> + PAGE_SHIFT; =20 - if (offset + length =3D=3D size) - goto no_split; + if (pstart) + *pstart =3D round_up(pos + offset, PAGE_SIZE) >> + PAGE_SHIFT; + + if (offset + length =3D=3D size) { + end =3D aligned_end; + goto out; + } =20 split_at2 =3D folio_page(folio, PAGE_ALIGN_DOWN(offset + length) / PAGE_SIZE); folio2 =3D page_folio(split_at2); =20 + /* + * folio2 may become stale due to a concurrent split or + * freeing, so validate it before and after taking its lock. + * If it fails, we can't get an accurate end position and fall + * back to folio->index, which may leave sub-folios split off + * at the offset edge in the page cache this round. + */ + end =3D folio->index; if (!folio_try_get(folio2)) - goto no_split; - - if (!folio_test_large(folio2)) goto out; =20 + if (folio2->mapping !=3D folio->mapping || + !folio_test_large(folio2)) + goto out_put; + if (!folio_trylock(folio2)) - goto out; + goto out_put; =20 - /* make sure folio2 is large and does not change its mapping */ - if (folio_test_large(folio2) && - folio2->mapping =3D=3D folio->mapping) - folio_split_or_unmap(folio2, split_at2, min_order); + if (page_folio(split_at2) !=3D folio2) { + folio_unlock(folio2); + goto out_put; + } + if (!folio_test_large(folio2)) { + end =3D aligned_end; + folio_unlock(folio2); + goto out_put; + } + + /* Split failed: back off to the head of the straddler */ + if (folio_split_or_unmap(folio2, split_at2, min_order)) + end =3D folio2->index; + else + end =3D aligned_end; =20 folio_unlock(folio2); -out: +out_put: folio_put(folio2); -no_split: +out: + if (pend) + *pend =3D end; return true; } if (folio_test_dirty(folio)) @@ -413,11 +455,8 @@ void truncate_inode_pages_range(struct address_space *= mapping, folio =3D __filemap_get_folio(mapping, lstart >> PAGE_SHIFT, FGP_LOCK, 0); if (!IS_ERR(folio)) { same_folio =3D lend < folio_next_pos(folio); - if (!truncate_inode_partial_folio(folio, lstart, lend)) { - start =3D folio_next_index(folio); - if (same_folio) - end =3D folio->index; - } + truncate_inode_partial_folio(folio, lstart, lend, &start, + same_folio ? &end : NULL); folio_unlock(folio); folio_put(folio); folio =3D NULL; @@ -427,8 +466,8 @@ void truncate_inode_pages_range(struct address_space *m= apping, folio =3D __filemap_get_folio(mapping, lend >> PAGE_SHIFT, FGP_LOCK, 0); if (!IS_ERR(folio)) { - if (!truncate_inode_partial_folio(folio, lstart, lend)) - end =3D folio->index; + truncate_inode_partial_folio(folio, lstart, lend, + NULL, &end); folio_unlock(folio); folio_put(folio); } --=20 2.52.0 From nobody Fri Sep 25 06:04:07 2026 Received: from dggsgout12.his.huawei.com (dggsgout12.his.huawei.com [45.249.212.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0D8B04AC14E; Wed, 16 Sep 2026 09:32:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.56 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789551180; cv=none; b=l5gnX9GZUNaMJJJyY11xJSFo4D5fziY3PNfJOd/TXJRek0eQ53NK28nm12Qt3Vdsrz0QPp3aL5HjC6u7NQbuX85NaM371ZYrEoXrV8eWKwFPFXfbTRp3SqcVARe0zMrT/4jHcj27EruJ9ufF9zcg3J4hqHpf26FHA5y/ZpdvJpY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789551180; c=relaxed/simple; bh=NtPjiCyCqEW56gksSfAtrRaeqQJ7ypvr+plInLxt+3k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=GUb3+FXXmdT2HCURINGEx8/K6IwwFUOxDPS2HTFkcr3pwrQPPj+HhxSOR5MbSTDrJ5Kr0fBda/TI5GbvLMSv6ryfNOrVValjS3SNuQOXSSPhap6Tdqv/QC/9UauxdnHUh+FcApytS6LctGhlxPtmY/ELKxVo/0LPkUE4b1+zeKs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.170]) by dggsgout12.his.huawei.com (SkyGuard) with ESMTPS id 4hlDFD58mlzKHMcf; Wed, 16 Sep 2026 17:32:32 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.128]) by mail.maildlp.com (Postfix) with ESMTP id DEF8540570; Wed, 16 Sep 2026 17:32:46 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP4 (Coremail) with UTF8SMTPSA id gCh0CgCHVikvYqpq8i0GAg--.65015S6; Wed, 16 Sep 2026 17:32:46 +0800 (CST) From: Zhang Yi To: linux-mm@kvack.org Cc: linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-ext4@vger.kernel.org, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, hughd@google.com, baolin.wang@linux.alibaba.com, willy@infradead.org, jack@suse.cz, ziy@nvidia.com, bfoster@redhat.com, joannelkoong@gmail.com, djwong@kernel.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, yangerkun@huawei.com, chengzhihao1@huawei.com, wangkefeng.wang@huawei.com, yukuai@fnnas.com Subject: [PATCH v3 2/3] mm/truncate: align truncation boundaries to mapping minimum folio order Date: Wed, 16 Sep 2026 17:24:49 +0800 Message-ID: <20260916092450.654408-3-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260916092450.654408-1-yi.zhang@huaweicloud.com> References: <20260916092450.654408-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: gCh0CgCHVikvYqpq8i0GAg--.65015S6 X-Coremail-Antispam: 1UD129KBjvJXoWxWr15ZFyfXryUWryUJFykGrg_yoWrXr45pF W2g3ZxArWUXr12kr4Dua1DZr45X39xX3WUAFWxGF93C3Z0q3WqkFyjg3W8Zw48GryxAryr ZF1jyasrWF1DAFJanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmY14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_Jryl82xGYIkIc2 x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2z4x0 Y4vE2Ix0cI8IcVAFwI0_Jr0_JF4l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Cr0_Gr1UM2 8EF7xvwVC2z280aVAFwI0_Cr1j6rxdM28EF7xvwVC2z280aVCY1x0267AKxVW0oVCq3wAS 0I0E0xvYzxvE52x082IY62kv0487Mc02F40EFcxC0VAKzVAqx4xG6I80ewAv7VC0I7IYx2 IY67AKxVWUJVWUGwAv7VC2z280aVAFwI0_Jr0_Gr1lOx8S6xCaFVCjc4AY6r1j6r4UM4x0 Y48IcxkI7VAKI48JM4x0x7Aq67IIx4CEVc8vx2IErcIFxwACI402YVCY1x02628vn2kIc2 xKxwCY1x0262kKe7AKxVW8ZVWrXwCF04k20xvY0x0EwIxGrwCFx2IqxVCFs4IE7xkEbVWU JVW8JwC20s026c02F40E14v26r1j6r18MI8I3I0E7480Y4vE14v26r106r1rMI8E67AF67 kF1VAFwI0_GFv_WrylIxkGc2Ij64vIr41lIxAIcVC0I7IYx2IY67AKxVWUJVWUCwCI42IY 6xIIjxv20xvEc7CjxVAFwI0_Cr0_Gr1UMIIF0xvE42xK8VAvwI8IcIk0rVWUJVWUCwCI42 IY6I8E87Iv67AKxVWUJVW8JwCI42IY6I8E87Iv6xkF7I0E14v26r4j6r4UJbIYCTnIWIev Ja73UjIFyTuYvjTRNiSHDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi When the mapping has a non-zero minimum folio order (min_order), folio_split() stops at min_order instead of order 0, so the sub-folio containing a split point stays aligned to 1 << min_order rather than to a single page. The original boundaries were based on page granularity, so either boundary could land inside the min_order chunk at its edge, and the truncation loop would drop that whole chunk, valid out-of-range tail included. For example, a 64K (order-4) folio with min_order =3D 2 (16K) punched from offset 0 to 36K: split @p0 -> [p0-p3, p4-p7, p8-p15] # non-uniform, min_order folio2 =3D p8-p15 # straddles: p8 in range, p9-p15 tail valid 2nd split of folio2 -> [p8-p11, p12-p15] # success end(old) =3D p9 # BUG: p9 inside [p8-p11] loop truncates ... p8-p11 # p9-p11's valid tail is lost It has gone unnoticed so far for two reasons. A non-zero min_order is only used by filesystems with a block or sector size larger than the page size, and those either always write back the affected range before punching a hole or truncating, or they carry filesystem private data on dirty folios (e.g. buffer_head), which makes filemap_release_folio() fail and folio_split() abort with -EBUSY, so the folio is never split and the old start/end boundaries remain valid. The bug only becomes reachable on paths that truncate dirty large folios without prior writeback and without filesystem private data, such as the upcoming ext4 iomap buffered I/O path. Align both start (rounded up) and end (rounded down) to the mapping minimum folio order so they always fall on a folio boundary. Reported-by: Joanne Koong Link: https://lore.kernel.org/linux-mm/CAJnrk1bQYUe6+1ryyJur5EEnZYrC+_5AYsy= =3DOWzVRgD4202y1g@mail.gmail.com/ Fixes: e220917fa5077 ("mm: split a folio in minimum folio order chunks") Signed-off-by: Zhang Yi --- mm/truncate.c | 14 ++++++++------ 1 file changed, 8 insertions(+), 6 deletions(-) diff --git a/mm/truncate.c b/mm/truncate.c index bec6d881d022..5ab7a40b1e25 100644 --- a/mm/truncate.c +++ b/mm/truncate.c @@ -213,8 +213,8 @@ static int folio_split_or_unmap(struct folio *folio, st= ruct page *split_at, * * When @pstart and/or @pend are non-NULL they receive the indexes of the * page range fully covered by [lstart, lend] after any split (or none), - * i.e. the range of pages wholly within [lstart, lend] and so safe to - * discard. + * aligned inwards to min_order, i.e. the range of folios wholly within + * [lstart, lend] and so safe to discard. * * Returns false if splitting failed so the caller can avoid * discarding the entire folio which is stubbornly unsplit. @@ -226,6 +226,7 @@ bool truncate_inode_partial_folio(struct folio *folio, = loff_t lstart, size_t size =3D folio_size(folio); unsigned int offset, length; struct page *split_at, *split_at2; + unsigned long min_nrbytes; unsigned int min_order; =20 if (pos < lstart) @@ -263,6 +264,7 @@ bool truncate_inode_partial_folio(struct folio *folio, = loff_t lstart, return true; =20 min_order =3D mapping_min_folio_order(folio->mapping); + min_nrbytes =3D mapping_min_folio_nrbytes(folio->mapping); split_at =3D folio_page(folio, PAGE_ALIGN_DOWN(offset) / PAGE_SIZE); if (!folio_split_or_unmap(folio, split_at, min_order)) { /* @@ -271,12 +273,12 @@ bool truncate_inode_partial_folio(struct folio *folio= , loff_t lstart, * for shmem truncate */ struct folio *folio2; - pgoff_t end, aligned_end =3D (pos + offset + length) >> - PAGE_SHIFT; + pgoff_t end, aligned_end =3D round_down(pos + offset + length, + min_nrbytes) >> PAGE_SHIFT; =20 if (pstart) - *pstart =3D round_up(pos + offset, PAGE_SIZE) >> - PAGE_SHIFT; + *pstart =3D round_up(pos + offset, + min_nrbytes) >> PAGE_SHIFT; =20 if (offset + length =3D=3D size) { end =3D aligned_end; --=20 2.52.0 From nobody Fri Sep 25 06:04:07 2026 Received: from dggsgout11.his.huawei.com (dggsgout11.his.huawei.com [45.249.212.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DB2084AEBD4; Wed, 16 Sep 2026 09:32:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.51 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789551186; cv=none; b=kOGAvcXWpv+mhn+7+sL0yMet45g2t44O/TvRNFJNnjmNvgffwhVtle/5NPXUj+lDGCn4JOqM6mfl5WLN9wMCyx/5I7rOSLAR0n1OIIpW049rU42mNljfnrpOhVlBP2cb0kDX+D2G51iJ0+4KSa+8wmsW3sIY477Pgr5ypj+gW/w= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789551186; c=relaxed/simple; bh=jtrV9wUz7keq+svAMMiwMuzlfCzYTh7F673W4idWNKY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=DCVDdfuE5VP+2gNo8cFh36RfxxCluc2bfpul+hQN+JrM40WKhaYEqeNuFMeZGuJ/OLxruMLP2k4IjNdhQpF3lNkKcQYFLTYZaGwDx/u8KJ4Ch4DZX0sqC9/vnH9yVI1IMaQXgbQIampblAsOiM9FxRCzqGWPeXStyO1wH5cPFEI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com; spf=pass smtp.mailfrom=huaweicloud.com; arc=none smtp.client-ip=45.249.212.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=huaweicloud.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huaweicloud.com Received: from mail.maildlp.com (unknown [172.19.163.198]) by dggsgout11.his.huawei.com (SkyGuard) with ESMTPS id 4hlDFQ2KkVzYQv1G; Wed, 16 Sep 2026 17:32:42 +0800 (CST) Received: from mail02.huawei.com (unknown [10.116.40.128]) by mail.maildlp.com (Postfix) with ESMTP id 0AD3540577; Wed, 16 Sep 2026 17:32:47 +0800 (CST) Received: from huaweicloud.com (unknown [10.50.85.155]) by APP4 (Coremail) with UTF8SMTPSA id gCh0CgCHVikvYqpq8i0GAg--.65015S7; Wed, 16 Sep 2026 17:32:46 +0800 (CST) From: Zhang Yi To: linux-mm@kvack.org Cc: linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-ext4@vger.kernel.org, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, hughd@google.com, baolin.wang@linux.alibaba.com, willy@infradead.org, jack@suse.cz, ziy@nvidia.com, bfoster@redhat.com, joannelkoong@gmail.com, djwong@kernel.org, yi.zhang@huawei.com, yi.zhang@huaweicloud.com, yizhang089@gmail.com, yangerkun@huawei.com, chengzhihao1@huawei.com, wangkefeng.wang@huawei.com, yukuai@fnnas.com Subject: [PATCH v3 3/3] mm/truncate: clarify return value of truncate_inode_partial_folio() Date: Wed, 16 Sep 2026 17:24:50 +0800 Message-ID: <20260916092450.654408-4-yi.zhang@huaweicloud.com> X-Mailer: git-send-email 2.52.0 In-Reply-To: <20260916092450.654408-1-yi.zhang@huaweicloud.com> References: <20260916092450.654408-1-yi.zhang@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: gCh0CgCHVikvYqpq8i0GAg--.65015S7 X-Coremail-Antispam: 1UD129KBjvJXoW7ZFW3CF4xuryfAry3XF15CFg_yoW8Kr43pa y7G3sxJ395Wa1Ikw1xuF48Zw4YyFZagrWUAFWxG3s7CFn8Zw4UKr1Ut3WUtw4fJr1kZ340 qF1jyFW3W3WUJF7anT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDU0xBIdaVrnRJUUUmI14x267AKxVWrJVCq3wAFc2x0x2IEx4CE42xK8VAvwI8IcIk0 rVWrJVCq3wAFIxvE14AKwVWUJVWUGwA2048vs2IY020E87I2jVAFwI0_JrWl82xGYIkIc2 x26xkF7I0E14v26ryj6s0DM28lY4IEw2IIxxk0rwA2F7IY1VAKz4vEj48ve4kI8wA2z4x0 Y4vE2Ix0cI8IcVAFwI0_JFI_Gr1l84ACjcxK6xIIjxv20xvEc7CjxVAFwI0_Gr1j6F4UJw A2z4x0Y4vEx4A2jsIE14v26F4UJVW0owA2z4x0Y4vEx4A2jsIEc7CjxVAFwI0_GcCE3s1l e2I262IYc4CY6c8Ij28IcVAaY2xG8wAqx4xG64xvF2IEw4CE5I8CrVC2j2WlYx0E2Ix0cI 8IcVAFwI0_Jr0_Jr4lYx0Ex4A2jsIE14v26r1j6r4UMcvjeVCFs4IE7xkEbVWUJVW8JwAC jcxG0xvY0x0EwIxGrwACjI8F5VA0II8E6IAqYI8I648v4I1lFIxGxcIEc7CjxVA2Y2ka0x kIwI1lc7CjxVAaw2AFwI0_GFv_Wryl42xK82IYc2Ij64vIr41l4I8I3I0E4IkC6x0Yz7v_ Jr0_Gr1lx2IqxVAqx4xG67AKxVWUJVWUGwC20s026x8GjcxK67AKxVWUGVWUWwC2zVAF1V AY17CE14v26r4a6rW5MIIYrxkI7VAKI48JMIIF0xvE2Ix0cI8IcVAFwI0_Jr0_JF4lIxAI cVC0I7IYx2IY6xkF7I0E14v26r4UJVWxJr1lIxAIcVCF04k26cxKx2IYs7xG6r1j6r1xMI IF0xvEx4A2jsIE14v26r1j6r4UMIIF0xvEx4A2jsIEc7CjxVAFwI0_Gr1j6F4UJbIYCTnI WIevJa73UjIFyTuYvjTRM6wCDUUUU X-CM-SenderInfo: d1lo6xhdqjqx5xdzvxpfor3voofrz/ Content-Type: text/plain; charset="utf-8" From: Zhang Yi With the earlier rework the callers no longer rely on the return value of truncate_inode_partial_folio() to decide whether to adjust the truncation range. The pstart/pend out-parameters carry that information instead. The callers now only use the return value as a flag indicating whether the loop should be reset to pick up newly split sub-folios on the shmem path. Return true if at least one split succeeded, and false otherwise. This clarifies the existing confusing return value semantics. Signed-off-by: Zhang Yi --- mm/truncate.c | 14 ++++++-------- 1 file changed, 6 insertions(+), 8 deletions(-) diff --git a/mm/truncate.c b/mm/truncate.c index 5ab7a40b1e25..23c90f00b530 100644 --- a/mm/truncate.c +++ b/mm/truncate.c @@ -216,8 +216,7 @@ static int folio_split_or_unmap(struct folio *folio, st= ruct page *split_at, * aligned inwards to min_order, i.e. the range of folios wholly within * [lstart, lend] and so safe to discard. * - * Returns false if splitting failed so the caller can avoid - * discarding the entire folio which is stubbornly unsplit. + * Return %true if at least one split succeeded, %false otherwise. */ bool truncate_inode_partial_folio(struct folio *folio, loff_t lstart, loff_t lend, pgoff_t *pstart, pgoff_t *pend) @@ -247,7 +246,7 @@ bool truncate_inode_partial_folio(struct folio *folio, = loff_t lstart, folio_wait_writeback(folio); if (length =3D=3D size) { truncate_inode_folio(folio->mapping, folio); - return true; + return false; } =20 /* @@ -261,7 +260,7 @@ bool truncate_inode_partial_folio(struct folio *folio, = loff_t lstart, if (folio_needs_release(folio)) folio_invalidate(folio, offset, length); if (!folio_test_large(folio)) - return true; + return false; =20 min_order =3D mapping_min_folio_order(folio->mapping); min_nrbytes =3D mapping_min_folio_nrbytes(folio->mapping); @@ -331,10 +330,9 @@ bool truncate_inode_partial_folio(struct folio *folio,= loff_t lstart, *pend =3D end; return true; } - if (folio_test_dirty(folio)) - return false; - truncate_inode_folio(folio->mapping, folio); - return true; + if (!folio_test_dirty(folio)) + truncate_inode_folio(folio->mapping, folio); + return false; } =20 /* --=20 2.52.0