From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2446D37E5FD; Fri, 17 Jul 2026 18:22:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312542; cv=none; b=OJTFdn0uIg6GmUTgs+o9p5I4VRmbmNt3rTP9S1P107dgOS32q91bMlk3rXIwjEeBl0ll58ZrT/JtYnLjKcB2xyeKZUlAqJwibIlrrWGSZac+4U8lldlyKuA+HYbXlAjoabo55cp6HJwuwyM+6kHxweF6RXepHEzbBDG8xBf6+1A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312542; c=relaxed/simple; bh=YXEfVWh2K58M+fDJD08jE97qBHPU4vg8rEAVLna7Ku4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Eqju9M+AbdTTdiBn7b3Hdbi+mbQXQOG59Bab6lL2NlqaPurhT7U74uFbJ8TD5am2RE18lHq4X1s6SxegVyCpP69awl9TZqKU7NsqO2IbvG5c0kDyTbfmnxmYg5zLAkGxkDMW5EAG6BSgD85dccJCSHICuo9rVQ66LjipRvB3pvw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=c8eIgB7Q; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="c8eIgB7Q" Received: by smtp.kernel.org (Postfix) with ESMTPSA id AB6751F00A3A; Fri, 17 Jul 2026 18:22:10 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312539; bh=TBERY1CFqX0D7t50CAbelyZ+L1fjNG+q7nyi8SSJxkg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=c8eIgB7QeydVoinrdlpcbGKnB55ieFOn62QyloUFSD/xDwVImtEpD1Uc6MjxCaPAZ DZO3S9lVGYz8Zsd2OyvF2CCi/7g6ZvOhyIQmpetfb39EXpqZl69WbWiXJqF2wNhlYD 9BUIAH5kVtQrz7GZhGbhH21OAjE2/rphIeArw9PRPf4hJ/ZzqA1Rc6vbdRxdgvPFUw /O9JtubscUgSdUVRM58U1xPiRDDye0Vdkyb2FU/J9x+Ce3X+bq+9Uif/ENzHO3iwj7 YDJEwmvXl2M94/Lux5lXAXJieBsCZ8/fCQcHsp6kkGQRuWSwAw34W9BeHqXThCf2Ae 6lYMHXjLnQoAA== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:39 +0100 Subject: [PATCH 01/15] mm/vma: introduce VMA virtual page offset field and add helpers Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-1-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=7578; i=ljs@kernel.org; h=from:subject:message-id; bh=YXEfVWh2K58M+fDJD08jE97qBHPU4vg8rEAVLna7Ku4=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivZqG/FtFZhhy1e24Vpp0o5/93M/T/xeLvvOriQrg kvBev/xjlIWBjEuBlkxRZbnX8T3B4mEzeu84O8GM4eVCWQIAxenAEwkZjEjw4zN7Id9mz9d9RQz rXA98rNj95pNe5juJTP05vU1Lvv0mIfhf0jhxhlLu725b5zePG3itCiX2e/M0p43P2wU3ezTZfh agRUA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This patch establishes fields within the vm_area_struct type to store the virtual page offset of VMAs. The virtual page offset of a VMA is equal to vma->vm_start >> PAGE_SHIFT if they are unfaulted or were not remapped, otherwise it is equal to this value at the point of first fault. Currently, anonymous folios belonging to CoW'd MAP_PRIVATE-mapped file-backed VMAs are tracked by their file offset. By adding virtual offset as a property of VMAs, we can now track them by their virtual page offset instead. By tracking this, we provide the means by which to eliminate this inconsistency, and more importantly lay the foundations for future work for the scalable CoW anonymous rmap rework. This patch simply adds the fields and some simple helpers. Subsequent patches will update mm code to make use of these fields correctly. The fields chosen are packed in the VMA such that, for 64-bit kernel builds, no additional space is taken up. The first field is present on cacheline 0 containing key VMA fields, and the second on cacheline 3, which contains file-backed reverse mapping fields. Given the relative time spent accessing reverse mapping fields as well as updating them, there shouldn't be any performance impact here from false sharing. Update the VMA userland tests to account for this change. No callsites are updated yet, so no functional change intended. Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/mm.h | 59 +++++++++++++++++++++++++++++++++++++= ++++ include/linux/mm_types.h | 4 +++ mm/vma.h | 14 ++++++++++ mm/vma_init.c | 1 + tools/testing/vma/include/dup.h | 26 ++++++++++++++++++ 5 files changed, 104 insertions(+) diff --git a/include/linux/mm.h b/include/linux/mm.h index 87feaa5a2b78..59b98cc60402 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -4393,6 +4393,65 @@ static inline pgoff_t vma_last_pgoff(const struct vm= _area_struct *vma) return vma_end_pgoff(vma) - 1; } =20 +/** + * vma_start_virt_pgoff() - Get the virtual page offset of the start of @v= ma + * @vma: The VMA whose virtual page offset is required. + * + * If unfaulted, then this is vma->vm_start >> PAGE_SHIFT, if faulted then= the + * virtual page offset at the time of first fault. + * + * If the VMA is anonymous, this returns the same value as vma_start_pgoff= (). + * + * This value is used for tracking MAP_PRIVATE file-backed mappings by the= ir + * virtual page offset. + * + * Returns: The virtual page offset of the start of @vma. + */ +static inline pgoff_t vma_start_virt_pgoff(const struct vm_area_struct *vm= a) +{ + pgoff_t pgoff =3D 0; + +#ifdef CONFIG_64BIT + pgoff +=3D vma->__vm_virt_pgoff_hi; + pgoff <<=3D 32; +#endif + pgoff +=3D vma->__vm_virt_pgoff_lo; + return pgoff; +} + +/** + * vma_end_virt_pgoff() - Get the virtual page offset of the exclusive end= of + * @vma. + * @vma: The VMA whose end virtual page offset is required. + * + * This returns the virtual exclusive end page offset of @vma, which is us= eful + * for expressing page offset ranges. + * + * See the description of vma_start_virt_pgoff() for a description of VMA + * virtual page offsets. + * + * Returns: The exclusive end virtual page offset of @vma. + */ +static inline pgoff_t vma_end_virt_pgoff(const struct vm_area_struct *vma) +{ + return vma_start_virt_pgoff(vma) + vma_pages(vma); +} + +/** + * vma_last_virt_pgoff() - Get the virtual page offset of the last page in + * @vma. + * @vma: The VMA whose last virtual page offset is required. + * + * See the description of vma_start_virt_pgoff() for a description of VMA + * virtual page offsets. + * + * Returns: The last virtual page offset of @vma. + */ +static inline pgoff_t vma_last_virt_pgoff(const struct vm_area_struct *vma) +{ + return vma_end_virt_pgoff(vma) - 1; +} + static inline unsigned long vma_desc_size(const struct vm_area_desc *desc) { return desc->end - desc->start; diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 939b5ea8c9e0..2710628059b1 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -967,6 +967,7 @@ struct vm_area_struct { */ unsigned int vm_lock_seq; #endif + unsigned int __vm_virt_pgoff_lo; /* Low 32-bits of virtual pgoff. */ /* * A file's MAP_PRIVATE vma can be in both i_mmap tree and anon_vma * list, after a COW of one of the file pages. A MAP_SHARED vma @@ -1041,6 +1042,9 @@ struct vm_area_struct { #ifdef CONFIG_DEBUG_LOCK_ALLOC struct lockdep_map vmlock_dep_map; #endif +#endif +#ifdef CONFIG_64BIT + unsigned int __vm_virt_pgoff_hi; /* High 32-bits of virtual pgoff. */ #endif /* * For areas with an address space and backing store, diff --git a/mm/vma.h b/mm/vma.h index 0bc7d521e976..04c5c1c546f6 100644 --- a/mm/vma.h +++ b/mm/vma.h @@ -283,6 +283,20 @@ static inline void vma_set_pgoff(struct vm_area_struct= *vma, pgoff_t pgoff) vma->vm_pgoff =3D pgoff; } =20 +static inline void __vma_set_virt_pgoff(struct vm_area_struct *vma, pgoff_= t pgoff) +{ +#ifdef CONFIG_64BIT + vma->__vm_virt_pgoff_hi =3D pgoff >> 32; +#endif + vma->__vm_virt_pgoff_lo =3D pgoff & GENMASK(31, 0); +} + +static inline void vma_set_virt_pgoff(struct vm_area_struct *vma, pgoff_t = pgoff) +{ + vma_assert_can_modify(vma); + __vma_set_virt_pgoff(vma, pgoff); +} + static inline void vma_add_pgoff(struct vm_area_struct *vma, pgoff_t delta) { vma_assert_can_modify(vma); diff --git a/mm/vma_init.c b/mm/vma_init.c index 715feee283f0..710b18849a36 100644 --- a/mm/vma_init.c +++ b/mm/vma_init.c @@ -51,6 +51,7 @@ static void vm_area_init_from(const struct vm_area_struct= *src, dest->vm_end =3D src->vm_end; dest->anon_vma =3D src->anon_vma; dest->vm_pgoff =3D vma_start_pgoff(src); + __vma_set_virt_pgoff(dest, vma_start_virt_pgoff(src)); dest->vm_file =3D src->vm_file; dest->vm_private_data =3D src->vm_private_data; vm_flags_init(dest, src->vm_flags); diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index c52b23773cd2..00d416d79ec1 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -577,6 +577,7 @@ struct vm_area_struct { */ unsigned int vm_lock_seq; #endif + unsigned int __vm_virt_pgoff_lo; =20 /* * A file's MAP_PRIVATE vma can be in both i_mmap tree and anon_vma @@ -612,6 +613,9 @@ struct vm_area_struct { #ifdef CONFIG_PER_VMA_LOCK /* Unstable RCU readers are allowed to read this. */ refcount_t vm_refcnt; +#endif +#ifdef CONFIG_64BIT + unsigned int __vm_virt_pgoff_hi; #endif /* * For areas with an address space and backing store, @@ -1320,6 +1324,28 @@ static inline pgoff_t vma_end_pgoff(const struct vm_= area_struct *vma) return vma_start_pgoff(vma) + vma_pages(vma); } =20 +static inline pgoff_t vma_start_virt_pgoff(const struct vm_area_struct *vm= a) +{ + pgoff_t pgoff =3D 0; + +#ifdef CONFIG_64BIT + pgoff +=3D vma->__vm_virt_pgoff_hi; + pgoff <<=3D 32; +#endif + pgoff +=3D vma->__vm_virt_pgoff_lo; + return pgoff; +} + +static inline pgoff_t vma_end_virt_pgoff(const struct vm_area_struct *vma) +{ + return vma_start_virt_pgoff(vma) + vma_pages(vma); +} + +static inline pgoff_t vma_last_virt_pgoff(const struct vm_area_struct *vma) +{ + return vma_end_virt_pgoff(vma) - 1; +} + static inline int vfs_mmap_prepare(struct file *file, struct vm_area_desc = *desc) { return file->f_op->mmap_prepare(desc); --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C5D8F37F732; Fri, 17 Jul 2026 18:22:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312555; cv=none; b=MJ/FaZaXpiyuGfLQUmBfbuF1kHVXUXOb4hoViczXVy4TjqMcXLiL7DB5zrYX8Jhv+2nw5wFIGkZMHFU+wJ5GWE3Itcivgy56jsQmQt7vudEAdlMxdW8leLvmNL2OmiAQy6fYHLW/Dmp1Pn4WfQjqkhWa94aZ0g3rX8q/zllaeSw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312555; c=relaxed/simple; bh=uH5Jc7cUxcyEjswUfjNSH8zXPOA1SoCYNdQzvopkVXQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Db9zzJXlMUwrmgOVIbm0IFHFrtqJJj3K7E32XBP2sDspxPP7aK0TSzOIcJLuuB0GUkNzSImKkSLdSsg3KBNFpUHp1GFmIeS7kdR+1QVedLVmCjYOOXpTu3lD325f4PiIFcQ//PNfzjln75Noys0b2FJ8pPY8iU682BT35IyGsr4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Iv0O7c/c; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Iv0O7c/c" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1402D1F000E9; Fri, 17 Jul 2026 18:22:19 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312548; bh=z7hUpZUpGSNT5pTheBg1aLusw4z3n2IV4NIpFOCmsM0=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=Iv0O7c/cyThXK+BwGq+t0bqFQIEcSP9JnjGHjL88OjKYam1KO93erpmi6Qqv2i6wj KLVl/KIvIIIXULRFqK6km63WOb1NSMAy19EAtsTqTqQrUOLO+xK1tS67UP5df987od lnvHCnhS+Yd2CkvHYczJb/wUnu6KJN5K+3KY4+jBydAZI5rw8gGhV3u/A5win8Ci/c yNfHr/Pt2aH9GyAchR6StrnHTINBJq534AmnPBZJWBrMMZEyU+oR4g/ORKm3Q2fONq tVgAzQghEMFtZ9J7vKCT9wA3sIYhpRzZRxKFuniOEOb/FVH4yJUGlKtC8hBmwDuxDl +nk6Wlknpcilw== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:40 +0100 Subject: [PATCH 02/15] mm: introduce linear_virt_page_index() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-2-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4596; i=ljs@kernel.org; h=from:subject:message-id; bh=uH5Jc7cUxcyEjswUfjNSH8zXPOA1SoCYNdQzvopkVXQ=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivZ+PFhuNm29vdiBi4qBnraTuVSfLl5j+9X4S3Bpm nx5ce2DjlIWBjEuBlkxRZbnX8T3B4mEzeu84O8GM4eVCWQIAxenAEzk3zOGf1ZWJUt/GxTc3f00 5+1L3apjr0y6BTiitCcslJgRJXbryhlGhnnSWxflRaelzcn/cNduHUe7yu2IhJOzJ+ZcfKjm8tp VnQsA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This function provides the virtual equivalent of linear_page_index(), instead offsetting based on the virtual page offset of the VMA. It is valid only for anonymous or MAP_PRIVATE file-backed mappings. It must not be called for shared file-backed mappings. For pure anon VMAs, this will be equal to linear_page_index(). We implement the algorithm in __linear_virt_page_index(), which is provided for internal mm code that might be interacting with shared VMAs. In linear_virt_page_index() we assert that both of these invariants are true. Note that MAP_PRIVATE-/dev/zero mappings will satisfy vma_is_anonymous() but not fulfill this invariant, so when asserting this we check vma->vm_file to account for this. We do not update callsites yet, so no functional change intended. Also const-ify vma_is_anonymous() to make it compatible with the const-ified linear_virt_page_index(). VMA userland tests are updated accordingly. Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/mm.h | 2 +- include/linux/pagemap.h | 41 +++++++++++++++++++++++++++++++++++++= ++++ tools/testing/vma/include/dup.h | 22 ++++++++++++++++++++++ 3 files changed, 64 insertions(+), 1 deletion(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index 59b98cc60402..b6503b5f0010 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -1556,7 +1556,7 @@ static inline void vma_desc_set_anonymous(struct vm_a= rea_desc *desc) desc->vm_ops =3D NULL; } =20 -static inline bool vma_is_anonymous(struct vm_area_struct *vma) +static inline bool vma_is_anonymous(const struct vm_area_struct *vma) { return !vma->vm_ops; } diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h index c6fc783aaee5..81da91e103b3 100644 --- a/include/linux/pagemap.h +++ b/include/linux/pagemap.h @@ -1101,6 +1101,47 @@ static inline pgoff_t linear_page_index(const struct= vm_area_struct *vma, return pgoff; } =20 +static inline pgoff_t __linear_virt_page_index(const struct vm_area_struct= *vma, + const unsigned long address) +{ + pgoff_t pgoff; + + pgoff =3D linear_page_delta(vma, address); + pgoff +=3D vma_start_virt_pgoff(vma); + return pgoff; +} + +/** + * linear_virt_page_index() - Determine the absolute virtual page offset of + * @address within @vma. + * @vma: An anonymous or MAP_PRIVATE file-backed VMA in which @address res= ides. + * @address: The address whose absolute page offset is required. + * + * This returns the virtual page offset of @address, which is the page off= set + * the address possessed at the time the VMA was first faulted. + * + * For anonymous mappings, this returns the same value as linear_page_inde= x(). + * + * For MAP_PRIVATE file-backed mappings, this returns the virtual page off= set of + * @address, which is the page offset the address possessed at the time th= e VMA + * was first faulted. + * + * It is not valid to call this function for shared file-backed mappings. + * + * Returns: The absolute virtual page offset of @address within @vma. + */ +static inline pgoff_t linear_virt_page_index(const struct vm_area_struct *= vma, + const unsigned long address) +{ + const pgoff_t pgoff =3D __linear_virt_page_index(vma, address); + + VM_WARN_ON_ONCE(vma_test(vma, VMA_SHARED_BIT)); + if (!vma->vm_file) /* Is anonymous except MAP_PRIVATE-/dev/zero */ + VM_WARN_ON_ONCE(pgoff !=3D linear_page_index(vma, address)); + + return pgoff; +} + struct wait_page_key { struct folio *folio; int bit_nr; diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index 00d416d79ec1..51605ade06fb 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -1610,3 +1610,25 @@ static inline pgprot_t vma_get_page_prot(const struc= t vm_area_struct *vma) { return vma_flags_to_page_prot(vma->flags); } + +static inline pgoff_t __linear_virt_page_index(const struct vm_area_struct= *vma, + const unsigned long address) +{ + pgoff_t pgoff; + + pgoff =3D linear_page_delta(vma, address); + pgoff +=3D vma_start_virt_pgoff(vma); + return pgoff; +} + +static inline pgoff_t linear_virt_page_index(const struct vm_area_struct *= vma, + const unsigned long address) +{ + const pgoff_t pgoff =3D __linear_virt_page_index(vma, address); + + VM_WARN_ON_ONCE(vma_test(vma, VMA_SHARED_BIT)); + if (!vma->vm_file) /* Is anonymous except MAP_PRIVATE-/dev/zero */ + VM_WARN_ON_ONCE(pgoff !=3D linear_page_index(vma, address)); + + return pgoff; +} --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 578E737B40C; Fri, 17 Jul 2026 18:22:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312563; cv=none; b=ebIajEBdbeUHtw6WBFBKoz06FGnNW91bz6VoqabAwFsny/TonHZZ8Wp/tZBA8hxxw3O1WtDCVkT0v0b7JYt8UN8jhZXXIvYLwuEHAK201uC+13xMtbt5pPISpqcXhWVzkW2TQ3dAbZA3wg1Gmm7BFF+TU5NXpgELDzEudpzvCww= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312563; c=relaxed/simple; bh=RyscjNpggtQJ4miTDtxwbhTzbm6OnS6a6lUgnowGnY8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=YcHXye/p9AWvYTzbyTbsZIH7wjNE27oCjtY6lgYW9CLLKft65LRKFBfY6ca2yXgRb5CS3kDchAXDwmKwa6cQDpe/K0/6Ui9x9NXs86bPphiZFSO8Fv0tKNcXtpdeZgwcIXMa8B0ZJQChwne9XaEgAFZqqQlmp0Mc5h4tZLN8kRk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=JcVCkXYb; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="JcVCkXYb" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 6D2291F00A3A; Fri, 17 Jul 2026 18:22:29 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312558; bh=mgnvDCamwFs796LwtMmRe0oIoMrcoFkkE3Z0cH2KMyM=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=JcVCkXYbqHwnz/lv+52TG6IeFP5mcYERwkcsuQUadn1YjT+dBFptC5HqOqi3OM+zD iwyOeiSdGMXoozab7Cy0MAodevehkDpNww7QWxCdO3tK8ONU/ZCABlw/wz9qpXm7M8 30l1oXnFmtHGeT0l9BpADeMgWLz+jsDYtN7MJ3bes9HMF2rRHjlw11sQdeuDEePd/l MTbz1rSAzbZ6sejQsfm9xCviBOSMCC6mwoOPMNI+rpflGY//L+fqKjvhbX2dG0hWN3 TG6XjEatCuDuXysY9EllWP0C95dgnbBEYFi6uw1rGmf7Rk/2onDTW0a5L1aNLz0tHA 9vhC9PMU8wesw== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:41 +0100 Subject: [PATCH 03/15] mm: abstract vma_address() and introduce vma_anon_address() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-3-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3665; i=ljs@kernel.org; h=from:subject:message-id; bh=RyscjNpggtQJ4miTDtxwbhTzbm6OnS6a6lUgnowGnY8=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivZuKWncvuNBOpfV8iuztPcVGTK/zP8pEr5g8pp7l ztvXmEN6ShlYRDjYpAVU2R5/kV8f5BI2LzOC/5uMHNYmUCGMHBxCsBElh1nZHj22efli9O3TSY/ Dr/bsahig2BfcZum2+97DQ8/bxCNCd3EyHA4KO7o1ikMG+7V/jS5e07vhVPXslO1n/ST/YPNW6/ M9mACAA== X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Introduce __vma_address() which abstracts the VMA start page offset field as pgoff_start, then update vma_address() to use it. Then introduce vma_anon_address() which does the equivalent of vma_address(), only using the virtual page offset of the VMA rather than the file-backed one. Also add an assert to ensure that the function is not called for mappings which are file-backed but not MAP_PRIVATE to ensure it is only used in the correct places. This will be necessary for determining the address of a folio's index within a VMA when the folio belongs to a MAP_PRIVATE file-backed VMA but has been CoW'd, and thus is anonymous, once the anonymous VMA page offset field is used for the reverse mapping. No callers are updated, so no functional change intended. Signed-off-by: Lorenzo Stoakes (ARM) --- mm/internal.h | 49 +++++++++++++++++++++++++++++++++++++------------ 1 file changed, 37 insertions(+), 12 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index f26423de4ca2..04acfeffa066 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1005,19 +1005,9 @@ void mlock_drain_remote(int cpu); =20 extern pmd_t maybe_pmd_mkwrite(pmd_t pmd, struct vm_area_struct *vma); =20 -/** - * vma_address - Find the virtual address a page range is mapped at - * @vma: The vma which maps this object. - * @pgoff: The page offset within its object. - * @nr_pages: The number of pages to consider. - * - * If any page in this range is mapped by this VMA, return the first addre= ss - * where any of these pages appear. Otherwise, return -EFAULT. - */ -static inline unsigned long vma_address(const struct vm_area_struct *vma, - pgoff_t pgoff, unsigned long nr_pages) +static inline unsigned long __vma_address(const struct vm_area_struct *vma, + pgoff_t pgoff, pgoff_t pgoff_start, unsigned long nr_pages) { - const pgoff_t pgoff_start =3D vma_start_pgoff(vma); unsigned long address; =20 if (pgoff >=3D pgoff_start) { @@ -1035,6 +1025,41 @@ static inline unsigned long vma_address(const struct= vm_area_struct *vma, return address; } =20 +/** + * vma_address - Find the virtual address a page range is mapped at. + * @vma: The vma which maps this object. + * @pgoff: The page offset within its object. + * @nr_pages: The number of pages to consider. + * + * If any page in this range is mapped by this VMA, return the first addre= ss + * where any of these pages appear. Otherwise, return -EFAULT. + */ +static inline unsigned long vma_address(const struct vm_area_struct *vma, + pgoff_t pgoff, unsigned long nr_pages) +{ + return __vma_address(vma, pgoff, vma_start_pgoff(vma), nr_pages); +} + +/** + * vma_anon_address - Find the address an anonymous folio with index @pgof= f_virt + * is mapped at. + * @vma: The vma which maps this object. + * @pgoff_virt: The virtual page index belonging to the folio. + * @nr_pages: The number of pages to consider. + * + * This is only valid for anonymous or MAP_PRIVATE-mapped file-backed VMAs. + * + * Returns: If any page in this range is mapped by this VMA, return the fi= rst address + * where any of these pages appear. Otherwise, return -EFAULT. + */ +static inline unsigned long vma_anon_address(const struct vm_area_struct *= vma, + pgoff_t pgoff_virt, unsigned long nr_pages) +{ + VM_WARN_ON_ONCE(!vma_is_anonymous(vma) && vma_test(vma, VMA_SHARED_BIT)); + + return __vma_address(vma, pgoff_virt, vma_start_virt_pgoff(vma), nr_pages= ); +} + /* * Then at what user virtual address will none of the range be found in vm= a? * Assumes that vma_address() already returned a good starting address. --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4307637F015; Fri, 17 Jul 2026 18:22:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312570; cv=none; b=f/PpPAd3mLaBwBaH1zO9TrL9nqf8X3aKdG0vfVQCAaM6IGa4IeZlss0XUlXPtAf/MAkW1dDDOlYOzFTVr2ZP2++ZNreyEm11BBp58vnmwyMARYy7RXuSDDy/nmS80AHN+OzWtRud+3vvaMq+9dgLj8ld4yaANDPzR5TzQB4WE94= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312570; c=relaxed/simple; bh=T5Gj/piAc/UlxcCiA9iiJMiRS1jgA5HHtjYX3YWMzOQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=PWqyYHEdze9XN5UVvpv8pJv+69KCUkz0YlTmal0ujzUzP/SQkq/4uJTWMCjVekFaM3OchTdtmGlWCjI8KAAGy1qOjNHqp/QWuCZzwR53GdWkDr8dI0RqKrZXM91BrHJU+mkt2CohpBfdEaFvOjN3uCkf494LV6qwPXd7N6kLWl8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=cQb0NqBC; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="cQb0NqBC" Received: by smtp.kernel.org (Postfix) with ESMTPSA id C5D8C1F000E9; Fri, 17 Jul 2026 18:22:38 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312567; bh=ZUSSGoClEIjzcszURqrS6Zg8fTue7uAoFYyLToW70HM=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=cQb0NqBCHQGHC1RFzL4Kq1JuN+fgnh3yxUX3z9TzKhvfgzLhXPkosagslKQGpb6ik jBM648jP4hMwYdweCkXVfjisIz2NgT6YNfxYnF4kza07w4Lf0PZ+TUeW5Sk8g8Pd0u 2bW2+mOyJN1QFVP22fB7jI8AGzEQ2pCS8rTpVg+2udEuBzyC5Iy6+DkEx7iKPAMCZe QWv5IHzNgQN6mfE5qeCXeoBpGYK5ujKOAoTCPH8wIpGSRtgU0N3EpRd2jqdVxpFxmb EP/b8mAoXtCira7rFVJ0wp0C7OSakTTEj2eGUOiOtWY/R20C+KLO8wlSfu9VRRe3Y1 RTHy+SMU2k59Q== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:42 +0100 Subject: [PATCH 04/15] mm: update print_bad_page_map() to show virtual page index Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-4-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=1716; i=ljs@kernel.org; h=from:subject:message-id; bh=T5Gj/piAc/UlxcCiA9iiJMiRS1jgA5HHtjYX3YWMzOQ=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivZaFhx+LXA8WfhVkt0Hq1l1+dZVuTueldzvOJmpO i9p03mxjlIWBjEuBlkxRZbnX8T3B4mEzeu84O8GM4eVCWQIAxenAEzEIJqR4cS79waejYfa5U85 d6TaKbms1DtksyqP60vuPEkPlWlefgz/vc4vVDhzVSJViO1zmTKn5o3tHPV3ln7+U9xzdp9T09c yJgA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This is potentially useful debugging information and matches the existing page offset provided. Use the raw __linear_virt_page_index() function so as to always output this value regardless of whether the mapping is file-backed or not. Signed-off-by: Lorenzo Stoakes (ARM) --- mm/memory.c | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/mm/memory.c b/mm/memory.c index d5e87624f692..56b244552f13 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -631,13 +631,14 @@ static void print_bad_page_map(struct vm_area_struct = *vma, { struct address_space *mapping; char entry_str[PTVAL_STR_MAX]; - pgoff_t index; + pgoff_t index, virt_index; =20 if (is_bad_page_map_ratelimited()) return; =20 mapping =3D vma->vm_file ? vma->vm_file->f_mapping : NULL; index =3D linear_page_index(vma, addr); + virt_index =3D __linear_virt_page_index(vma, addr); =20 ptval_bytes_to_hex_str(entry_str, sizeof(entry_str), entry, entry_size); pr_alert("BUG: Bad page map in process %s %s:%s", current->comm, @@ -645,8 +646,9 @@ static void print_bad_page_map(struct vm_area_struct *v= ma, __print_bad_page_map_pgtable(vma->vm_mm, addr); if (page) dump_page(page, "bad page map"); - pr_alert("addr:%px vm_flags:%08lx anon_vma:%px mapping:%px index:%lx\n", - (void *)addr, vma->vm_flags, vma->anon_vma, mapping, index); + pr_alert("addr:%px vm_flags:%08lx anon_vma:%px mapping:%px index:%lx virt= _index:%lx\n", + (void *)addr, vma->vm_flags, vma->anon_vma, mapping, index, + virt_index); pr_alert("file:%pD fault:%ps mmap:%ps mmap_prepare: %ps read_folio:%ps\n", vma->vm_file, vma->vm_ops ? vma->vm_ops->fault : NULL, --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 50EC237E2E7; Fri, 17 Jul 2026 18:22:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312581; cv=none; b=uokBAUVXPM6iV3a8ufqFhY15bLR7OLQhFay4Lc4dexRe7OZ0a4FWUQFaVdTp23yXg04BNhUALBNqgnoUQ/9ZElrX2/UrlS7z1eD0rn3ktIEay4yOwzaqqqZ08cEz42ziXm4mwGgG74JvKhcpGPung1mE3qH3hX0z+WoZygt9Nw0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312581; c=relaxed/simple; bh=wV4Xs9C/oteickrPVjPK0/HJqNh07r+G5RYKSasgEQI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=lMxh62FWcGKJLEqevzDmccUjWWvOrhTFUmBqiMB3PjRLmMMgWB8qJl5V8k1ei9d1suRRHZGLLChO/hDvPe36Zzs7XHGQV5fp8aV94r0e/0q1xWR9vGmOcRBbxan4lcs1syHMSmTJzjq1vFwCNMIUShPjoE+rjnbfxZmh1qGif0I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=UPaEdbOj; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="UPaEdbOj" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2A5BC1F00A3A; Fri, 17 Jul 2026 18:22:47 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312577; bh=hWH16RMl+8FexBLFh6ScXidqQvRnE14iA6MdxcccmsM=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=UPaEdbOjipnQeZVXns57A4+tKFRvb0G3mGTTDnuRx8lhCbgfamRLCaxC5Dn2lAje8 t4SPcjV+dETPer94Y/Iqfz6M9ip447qkNetc5M5FLbc/UqzIN7wTGRhWh9YS9uYEhV iAkK/p7lNbrlApeUMXnGASP1b+9EKfyLwayGXeRy++ru0UKZFRufVP1DaKO/UaNQUZ +oF/1WnnOI40UWX+brIZq526ZOF9ei1YtlekU/PQw6yde4fOEoMj7e9XgpUEQ7Knl1 fFMidcddxj3u9T9UB7/q8wS6OKOp7Gle12Dio9ylJqhlBvG+qGM9P0eEYscfOO3OGM nXM9zkp9iCdsA== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:43 +0100 Subject: [PATCH 05/15] mm: introduce and use vma_filebacked_address() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-5-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4889; i=ljs@kernel.org; h=from:subject:message-id; bh=wV4Xs9C/oteickrPVjPK0/HJqNh07r+G5RYKSasgEQI=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivbOXH1cVvHc5SudNg+3C4v63ja3lnsSvz/x5zTx5 PvF/xardJSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAik4QZ/qnbfL10XjZE0/W8 z/mtio1177e80F91QO3R70Prtu+WidvLyLBarztkVusfUbGuLh1z95Y9EQ/9VKU9f0lv9XHYE2b yjQcA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 In cases where we know that the VMA is file-backed, use vma_filebacked_address() rather than vma_address(). This lays the foundation for using the virtual page offset for anonymous VMAs via vma_anon_address(). Also add an assert to ensure that the VMA whose address is required is not anonymous. No functional change intended. Signed-off-by: Lorenzo Stoakes (ARM) --- mm/internal.h | 18 ++++++++++++++++++ mm/memory-failure.c | 4 ++-- mm/page_vma_mapped.c | 6 +++++- mm/rmap.c | 10 ++++++---- 4 files changed, 31 insertions(+), 7 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index 04acfeffa066..73f90aaa979d 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1025,6 +1025,24 @@ static inline unsigned long __vma_address(const stru= ct vm_area_struct *vma, return address; } =20 +/** + * vma_filebacked_address - Find the virtual address a file-backed page ra= nge is + * mapped at. + * @vma: The vma which maps this object. + * @pgoff: The page offset within its object. + * @nr_pages: The number of pages to consider. + * + * Returns: If any page in this range is mapped by this VMA, return the fi= rst + * address where any of these pages appear. Otherwise, return -EFAULT. + */ +static inline unsigned long vma_filebacked_address(const struct vm_area_st= ruct *vma, + pgoff_t pgoff, unsigned long nr_pages) +{ + VM_WARN_ON_ONCE(vma_is_anonymous(vma)); + + return __vma_address(vma, pgoff, vma_start_pgoff(vma), nr_pages); +} + /** * vma_address - Find the virtual address a page range is mapped at. * @vma: The vma which maps this object. diff --git a/mm/memory-failure.c b/mm/memory-failure.c index aaf14608b30e..a8b03e2920ba 100644 --- a/mm/memory-failure.c +++ b/mm/memory-failure.c @@ -620,7 +620,7 @@ static void add_to_kill_fsdax(struct task_struct *tsk, = const struct page *p, struct vm_area_struct *vma, struct list_head *to_kill, pgoff_t pgoff) { - unsigned long addr =3D vma_address(vma, pgoff, 1); + unsigned long addr =3D vma_filebacked_address(vma, pgoff, 1); __add_to_kill(tsk, p, vma, to_kill, addr); } =20 @@ -2265,7 +2265,7 @@ static void add_to_kill_pgoff(struct task_struct *tsk, } =20 /* Check for pgoff not backed by struct page */ - tk->addr =3D vma_address(vma, pgoff, 1); + tk->addr =3D vma_filebacked_address(vma, pgoff, 1); tk->size_shift =3D PAGE_SHIFT; =20 if (tk->addr =3D=3D -EFAULT) diff --git a/mm/page_vma_mapped.c b/mm/page_vma_mapped.c index d7670ba4147b..081e483cc7bf 100644 --- a/mm/page_vma_mapped.c +++ b/mm/page_vma_mapped.c @@ -356,6 +356,7 @@ unsigned long page_mapped_in_vma(const struct page *pag= e, struct vm_area_struct *vma) { const struct folio *folio =3D page_folio(page); + const pgoff_t pgoff =3D page_pgoff(folio, page); struct page_vma_mapped_walk pvmw =3D { .pfn =3D page_to_pfn(page), .nr_pages =3D 1, @@ -363,7 +364,10 @@ unsigned long page_mapped_in_vma(const struct page *pa= ge, .flags =3D PVMW_SYNC, }; =20 - pvmw.address =3D vma_address(vma, page_pgoff(folio, page), 1); + if (folio_test_anon(folio)) + pvmw.address =3D vma_address(vma, pgoff, 1); + else + pvmw.address =3D vma_filebacked_address(vma, pgoff, 1); if (pvmw.address =3D=3D -EFAULT) goto out; if (!page_vma_mapped_walk(&pvmw)) diff --git a/mm/rmap.c b/mm/rmap.c index ad820fe86f7d..5798427d007f 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -865,14 +865,15 @@ unsigned long page_address_in_vma(const struct folio = *folio, if (!vma->anon_vma || !anon_vma || vma->anon_vma->root !=3D anon_vma->root) return -EFAULT; + /* KSM folios don't reach here because of the !anon_vma check */ + return vma_address(vma, page_pgoff(folio, page), 1); } else if (!vma->vm_file) { return -EFAULT; } else if (vma->vm_file->f_mapping !=3D folio->mapping) { return -EFAULT; } =20 - /* KSM folios don't reach here because of the !anon_vma check */ - return vma_address(vma, page_pgoff(folio, page), 1); + return vma_filebacked_address(vma, page_pgoff(folio, page), 1); } =20 /* @@ -1321,7 +1322,7 @@ int pfn_mkclean_range(unsigned long pfn, unsigned lon= g nr_pages, pgoff_t pgoff, if (invalid_mkclean_vma(vma, NULL)) return 0; =20 - pvmw.address =3D vma_address(vma, pgoff, nr_pages); + pvmw.address =3D vma_filebacked_address(vma, pgoff, nr_pages); VM_BUG_ON_VMA(pvmw.address =3D=3D -EFAULT, vma); =20 return page_vma_mkclean_one(&pvmw); @@ -3100,7 +3101,8 @@ static void __rmap_walk_file(struct folio *folio, str= uct address_space *mapping, } lookup: mapping_rmap_tree_foreach(vma, mapping, pgoff_start, pgoff_end) { - unsigned long address =3D vma_address(vma, pgoff_start, nr_pages); + unsigned long address =3D vma_filebacked_address(vma, pgoff_start, + nr_pages); =20 VM_BUG_ON_VMA(address =3D=3D -EFAULT, vma); cond_resched(); --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BB15937F015; Fri, 17 Jul 2026 18:23:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312592; cv=none; b=koUicug8wwp3gvNgtzf8cmWfbWK9l8sUwFP29yS8MCXuQVfEss6r5Hdmh1tkxDgXPwRMLaBpGcZoJeh49cfjfgtlDu7T1pOYSj863m7AKr0gXElrNm0btTxxB1rQscx0Tm5N380WhWnJOJNBImdiZ6Oie5eiFn98isHZ0V60/xI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312592; c=relaxed/simple; bh=bAL004VzAX7A4Tw7tZ5pMBFAkxdI+imb7ffhtNnNsm0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=QQ/tMqHZyIwS6+AR6fv/CUm6L4e/ACeHo8c4KC81cyIJg50kJXBwOS5sY57YH+0DclsDYPXCzVjl/dtopENCgqskzHk8sKI+XER+dsNSVN1KX5i6PzMzmkXXsDU6QQpVbmbil6yHAwdkh7rx48/1IfjDDsOhBXzqTc03bK6Nbr0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=XishULLS; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="XishULLS" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 82EE01F000E9; Fri, 17 Jul 2026 18:22:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312586; bh=XMjFEdTH6k3hj5Oy3hKgeo/PLhtrcN2I3TWsmNiCNVQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=XishULLSAFO4zVL2jpcFQSof4hhC7vwzkRRH+mIWJLtV/jseY8OXcWPZcrC2dooc4 g4YwcQA3LFRZImEDmkbN8ak2IX/ZodpL2i9gyBsrk+yHqekF6w6wpjgcqjyRewrnNP Us5t27Fk9m/UJJJEibXWMsNe3xO9a18ud5/oSrrW7WrLQG4706XJNS7dDpsee3+5PH iptmCoqdgVNo18hmwlNgri4B37m+HqSRetlBZTrPJJgaUQJEKZ0NY1y/AlC1Il/Fwr msCjFcgiuDxdtxGe7WqC5ezmD3c76lKgMJlIqLscZ8wHfEsH/F/OdcGSlDRGbD378T LbhpcDLruRfcg== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:44 +0100 Subject: [PATCH 06/15] mm: propagate VMA virtual page offset on map, remap, split + merge Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-6-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=21281; i=ljs@kernel.org; h=from:subject:message-id; bh=bAL004VzAX7A4Tw7tZ5pMBFAkxdI+imb7ffhtNnNsm0=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivYu/fu198umGkPfPOawjj3uZqpvdpfmqRyI/f7u7 KcvcUJhHaUsDGJcDLJiiizPv4jvDxIJm9d5wd8NZg4rE8gQBi5OAZiIfjYjw6ODz+a9fC6y4OrJ Zz3dt8pZJFreJAmp96Vf1pJKnPpM24KRYeYBKYftT7Xiv3tIuzA2h+sXlPxV/vFC+ePXkNC+CLN wDgA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 We must correctly update VMA virtual page offset state on all VMA operations that would result in it changing, with special attention given to remapping. We cover most cases by simply updating vma_set_range() to do so (with a new virtual page offset parameter), but also notably must update the merging and mapping logic to propagate this parameter correctly. The remap logic remains the same - we may update the virtual page offset if the VMA is unfaulted, but now this applies to MAP_PRIVATE file-backed mappings too, so we update the code to reflect this. Note that we use __linear_virt_page_index() upon remap as the VMA may be shared, in order that we update the field consistently regardless of VMA type. Also while we're here, replace a VM_BUG_ON_VMA() with a VM_WARN_ON_ONCE_VMA(). We also introduce vma_anon_pgoff_addr(), vma_start_anon_pgoff(), and vma_end_anon_pgoff() which differ from their virtual page offset equivalents in that a shared file-backed mapping returns its file page offset otherwise the virtual page offset is used. This means we don't predicate merges for shared file-backed mappings on virtual page offset. We simply ensure state is correctly propagated here, so no functional changes are intended. Finally, we update insert_vm_struct() to correctly set the virtual page offset on insertion of a VMA. Also update VMA userland tests to reflect this change. Signed-off-by: Lorenzo Stoakes (ARM) --- mm/mremap.c | 6 +- mm/vma.c | 56 ++++++++++++----- mm/vma.h | 127 ++++++++++++++++++++++++++++++-----= ---- mm/vma_exec.c | 2 +- tools/testing/vma/shared.c | 3 +- tools/testing/vma/tests/merge.c | 4 +- tools/testing/vma/tests/vma.c | 4 +- tools/testing/vma/vma_internal.h | 1 + 8 files changed, 152 insertions(+), 51 deletions(-) diff --git a/mm/mremap.c b/mm/mremap.c index b64aa1f6e07e..f07fc4e3ef2e 100644 --- a/mm/mremap.c +++ b/mm/mremap.c @@ -1265,7 +1265,9 @@ static void unmap_source_vma(struct vma_remap_struct = *vrm) static int copy_vma_and_data(struct vma_remap_struct *vrm, struct vm_area_struct **new_vma_ptr) { - const unsigned long new_pgoff =3D linear_page_index(vrm->vma, vrm->addr); + const pgoff_t new_pgoff =3D linear_page_index(vrm->vma, vrm->addr); + const pgoff_t new_virt_pgoff =3D + __linear_virt_page_index(vrm->vma, vrm->addr); struct vm_area_struct *vma =3D vrm->vma; struct vm_area_struct *new_vma; unsigned long moved_len; @@ -1273,7 +1275,7 @@ static int copy_vma_and_data(struct vma_remap_struct = *vrm, PAGETABLE_MOVE(pmc, NULL, NULL, vrm->addr, vrm->new_addr, vrm->old_len); =20 new_vma =3D copy_vma(&vma, vrm->new_addr, vrm->new_len, new_pgoff, - &pmc.need_rmap_locks); + new_virt_pgoff, &pmc.need_rmap_locks); if (!new_vma) { vrm_uncharge(vrm); *new_vma_ptr =3D NULL; diff --git a/mm/vma.c b/mm/vma.c index b5bc3eec961c..e45d43493446 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -18,6 +18,7 @@ struct mmap_state { unsigned long addr; unsigned long end; pgoff_t pgoff; + pgoff_t virt_pgoff; unsigned long pglen; union { vm_flags_t vm_flags; @@ -46,13 +47,22 @@ struct mmap_state { bool file_doesnt_need_get :1; }; =20 -#define MMAP_STATE(name, mm_, vmi_, addr_, len_, pgoff_, vma_flags_, file_= ) \ +static inline pgoff_t map_anon_pgoff(const struct mmap_state *map) +{ + if (vma_flags_test(&map->vma_flags, VMA_SHARED_BIT)) + return map->pgoff; + + return map->virt_pgoff; +} + +#define MMAP_STATE(name, mm_, vmi_, addr_, len_, pgoff_, virt_pgoff_, vma_= flags_, file_) \ struct mmap_state name =3D { \ .mm =3D mm_, \ .vmi =3D vmi_, \ .addr =3D addr_, \ .end =3D (addr_) + (len_), \ .pgoff =3D pgoff_, \ + .virt_pgoff =3D virt_pgoff_, \ .pglen =3D PHYS_PFN(len_), \ .vma_flags =3D vma_flags_, \ .file =3D file_, \ @@ -67,6 +77,7 @@ struct mmap_state { .end =3D (map_)->end, \ .vma_flags =3D (map_)->vma_flags, \ .pgoff =3D (map_)->pgoff, \ + .anon_pgoff =3D map_anon_pgoff(map_), \ .file =3D (map_)->file, \ .prev =3D (map_)->prev, \ .middle =3D vma_, \ @@ -82,10 +93,11 @@ static void __vma_set_range(struct vm_area_struct *vma,= unsigned long start, } =20 static void vma_set_range(struct vm_area_struct *vma, unsigned long start, - unsigned long end, pgoff_t pgoff) + unsigned long end, pgoff_t pgoff, pgoff_t virt_pgoff) { __vma_set_range(vma, start, end); vma_set_pgoff(vma, pgoff); + vma_set_virt_pgoff(vma, virt_pgoff); } =20 /* Was this VMA ever forked from a parent, i.e. maybe contains CoW mapping= s? */ @@ -812,7 +824,8 @@ static int commit_merge(struct vma_merge_struct *vmg) */ vma_adjust_trans_huge(vma, vmg->start, vmg->end, vmg->__adjust_middle_start ? vmg->middle : NULL); - vma_set_range(vma, vmg->start, vmg->end, vmg_start_pgoff(vmg)); + vma_set_range(vma, vmg->start, vmg->end, vmg_start_pgoff(vmg), + vmg_start_anon_pgoff(vmg)); vmg_adjust_set_range(vmg); vma_iter_store_overwrite(vmg->vmi, vmg->target); =20 @@ -982,6 +995,7 @@ static __must_check struct vm_area_struct *vma_merge_ex= isting_range( vmg->start =3D prev->vm_start; vmg->end =3D next->vm_end; vmg->pgoff =3D vma_start_pgoff(prev); + vmg->anon_pgoff =3D vma_start_anon_pgoff(prev); =20 /* * We already ensured anon_vma compatibility above, so now it's @@ -1000,6 +1014,7 @@ static __must_check struct vm_area_struct *vma_merge_= existing_range( */ vmg->start =3D prev->vm_start; vmg->pgoff =3D vma_start_pgoff(prev); + vmg->anon_pgoff =3D vma_start_anon_pgoff(prev); =20 if (!vmg->__remove_middle) vmg->__adjust_middle_start =3D true; @@ -1022,12 +1037,14 @@ static __must_check struct vm_area_struct *vma_merg= e_existing_range( if (vmg->__remove_middle) { vmg->end =3D next->vm_end; vmg->pgoff =3D vma_start_pgoff(next) - pglen; + vmg->anon_pgoff =3D vma_start_anon_pgoff(next) - pglen; } else { /* We shrink middle and expand next. */ vmg->__adjust_next_start =3D true; vmg->start =3D middle->vm_start; vmg->end =3D start; vmg->pgoff =3D vma_start_pgoff(middle); + vmg->anon_pgoff =3D vma_start_anon_pgoff(middle); } =20 err =3D dup_anon_vma(next, middle, &anon_dup); @@ -1137,6 +1154,7 @@ struct vm_area_struct *vma_merge_new_range(struct vma= _merge_struct *vmg) vmg->start =3D prev->vm_start; vmg->target =3D prev; vmg->pgoff =3D vma_start_pgoff(prev); + vmg->anon_pgoff =3D vma_start_anon_pgoff(prev); =20 /* * If this merge would result in removal of the next VMA but we @@ -1908,9 +1926,10 @@ static int vma_link(struct mm_struct *mm, struct vm_= area_struct *vma) */ struct vm_area_struct *copy_vma(struct vm_area_struct **vmap, unsigned long addr, unsigned long len, pgoff_t pgoff, - bool *need_rmap_locks) + pgoff_t virt_pgoff, bool *need_rmap_locks) { struct vm_area_struct *vma =3D *vmap; + const bool is_shared =3D vma_test(vma, VMA_SHARED_BIT); unsigned long vma_start =3D vma->vm_start; struct mm_struct *mm =3D vma->vm_mm; struct vm_area_struct *new_vma; @@ -1919,11 +1938,14 @@ struct vm_area_struct *copy_vma(struct vm_area_stru= ct **vmap, VMG_VMA_STATE(vmg, &vmi, NULL, vma, addr, addr + len); =20 /* - * If anonymous vma has not yet been faulted, update new pgoff - * to match new location, to increase its chance of merging. + * If a vma has not yet been faulted, update its virtual pgoff to match + * the new location to increase its chance of merging. */ - if (unlikely(vma_is_anonymous(vma) && !vma->anon_vma)) { - pgoff =3D addr >> PAGE_SHIFT; + if (!vma->anon_vma && !is_shared) { + virt_pgoff =3D addr >> PAGE_SHIFT; + + if (vma_is_anonymous(vma)) + pgoff =3D virt_pgoff; faulted_in_anon_vma =3D false; } =20 @@ -1940,6 +1962,7 @@ struct vm_area_struct *copy_vma(struct vm_area_struct= **vmap, return NULL; /* should never get here */ =20 vmg.pgoff =3D pgoff; + vmg.anon_pgoff =3D is_shared ? pgoff : virt_pgoff; vmg.next =3D vma_iter_next_rewind(&vmi, NULL); new_vma =3D vma_merge_copied_range(&vmg); =20 @@ -1961,7 +1984,7 @@ struct vm_area_struct *copy_vma(struct vm_area_struct= **vmap, * safe. It is only safe to keep the vm_pgoff * linear if there are no pages mapped yet. */ - VM_BUG_ON_VMA(faulted_in_anon_vma, new_vma); + VM_WARN_ON_ONCE_VMA(faulted_in_anon_vma, new_vma); *vmap =3D vma =3D new_vma; } *need_rmap_locks =3D @@ -1970,7 +1993,7 @@ struct vm_area_struct *copy_vma(struct vm_area_struct= **vmap, new_vma =3D vm_area_dup(vma); if (!new_vma) goto out; - vma_set_range(new_vma, addr, addr + len, pgoff); + vma_set_range(new_vma, addr, addr + len, pgoff, virt_pgoff); if (vma_dup_policy(vma, new_vma)) goto out_free_vma; if (anon_vma_clone(new_vma, vma, VMA_OP_REMAP)) @@ -2612,7 +2635,7 @@ static int __mmap_new_vma(struct mmap_state *map, str= uct vm_area_struct **vmap, if (is_anon) vma_set_anonymous(vma); =20 - vma_set_range(vma, map->addr, map->end, map->pgoff); + vma_set_range(vma, map->addr, map->end, map->pgoff, map->virt_pgoff); vma->flags =3D map->vma_flags; vma->vm_page_prot =3D map->page_prot; =20 @@ -2801,7 +2824,8 @@ static unsigned long __mmap_region(struct file *file,= unsigned long addr, struct vm_area_struct *vma =3D NULL; bool have_mmap_prepare =3D file && file->f_op->mmap_prepare; VMA_ITERATOR(vmi, mm, addr); - MMAP_STATE(map, mm, &vmi, addr, len, pgoff, vma_flags, file); + const pgoff_t virt_pgoff =3D addr >> PAGE_SHIFT; + MMAP_STATE(map, mm, &vmi, addr, len, pgoff, virt_pgoff, vma_flags, file); struct vm_area_desc desc =3D { .mm =3D mm, .file =3D file, @@ -2946,6 +2970,7 @@ int do_brk_flags(struct vma_iterator *vmi, struct vm_= area_struct *vma, unsigned long addr, unsigned long len, vma_flags_t vma_flags) { struct mm_struct *mm =3D current->mm; + const pgoff_t pgoff =3D addr >> PAGE_SHIFT; =20 /* * Check against address space limits by the changed size @@ -2970,7 +2995,7 @@ int do_brk_flags(struct vma_iterator *vmi, struct vm_= area_struct *vma, * occur after forking, so the expand will only happen on new VMAs. */ if (vma && vma->vm_end =3D=3D addr) { - VMG_STATE(vmg, mm, vmi, addr, addr + len, vma_flags, PHYS_PFN(addr)); + VMG_STATE(vmg, mm, vmi, addr, addr + len, vma_flags, pgoff, pgoff); =20 vmg.prev =3D vma; /* vmi is positioned at prev, which this mode expects. */ @@ -2990,7 +3015,7 @@ int do_brk_flags(struct vma_iterator *vmi, struct vm_= area_struct *vma, goto unacct_fail; =20 vma_set_anonymous(vma); - vma_set_range(vma, addr, addr + len, addr >> PAGE_SHIFT); + vma_set_range(vma, addr, addr + len, pgoff, pgoff); vma->flags =3D vma_flags; vma->vm_page_prot =3D vm_get_page_prot(vma_flags_to_legacy(vma_flags)); vma_start_write(vma); @@ -3382,6 +3407,7 @@ int insert_vm_struct(struct mm_struct *mm, struct vm_= area_struct *vma) WARN_ON_ONCE(vma->anon_vma); vma_set_pgoff(vma, vma->vm_start >> PAGE_SHIFT); } + vma_set_virt_pgoff(vma, vma->vm_start >> PAGE_SHIFT); =20 if (vma_link(mm, vma)) { if (vma_test(vma, VMA_ACCOUNT_BIT)) @@ -3437,7 +3463,7 @@ struct vm_area_struct *__install_special_mapping( =20 vma->vm_ops =3D ops; vma->vm_private_data =3D priv; - vma_set_range(vma, addr, addr + len, 0); + vma_set_range(vma, addr, addr + len, 0, addr >> PAGE_SHIFT); =20 ret =3D insert_vm_struct(mm, vma); if (ret) diff --git a/mm/vma.h b/mm/vma.h index 04c5c1c546f6..c1d65060071f 100644 --- a/mm/vma.h +++ b/mm/vma.h @@ -104,6 +104,7 @@ struct vma_merge_struct { unsigned long start; unsigned long end; pgoff_t pgoff; + pgoff_t anon_pgoff; =20 union { /* Temporary while VMA flags are being converted. */ @@ -237,11 +238,6 @@ static inline bool vmg_nomem(struct vma_merge_struct *= vmg) return vmg->state =3D=3D VMA_MERGE_ERROR_NOMEM; } =20 -static inline pgoff_t vmg_start_pgoff(const struct vma_merge_struct *vmg) -{ - return vmg->pgoff; -} - static inline pgoff_t vmg_pages(const struct vma_merge_struct *vmg) { const unsigned long size =3D vmg->end - vmg->start; @@ -249,6 +245,11 @@ static inline pgoff_t vmg_pages(const struct vma_merge= _struct *vmg) return size >> PAGE_SHIFT; } =20 +static inline pgoff_t vmg_start_pgoff(const struct vma_merge_struct *vmg) +{ + return vmg->pgoff; +} + static inline pgoff_t vmg_end_pgoff(const struct vma_merge_struct *vmg) { return vmg_start_pgoff(vmg) + vmg_pages(vmg); @@ -283,6 +284,16 @@ static inline void vma_set_pgoff(struct vm_area_struct= *vma, pgoff_t pgoff) vma->vm_pgoff =3D pgoff; } =20 +static inline pgoff_t vmg_start_anon_pgoff(const struct vma_merge_struct *= vmg) +{ + return vmg->anon_pgoff; +} + +static inline pgoff_t vmg_end_anon_pgoff(const struct vma_merge_struct *vm= g) +{ + return vmg_start_anon_pgoff(vmg) + vmg_pages(vmg); +} + static inline void __vma_set_virt_pgoff(struct vm_area_struct *vma, pgoff_= t pgoff) { #ifdef CONFIG_64BIT @@ -301,44 +312,102 @@ static inline void vma_add_pgoff(struct vm_area_stru= ct *vma, pgoff_t delta) { vma_assert_can_modify(vma); vma_set_pgoff(vma, vma_start_pgoff(vma) + delta); + vma_set_virt_pgoff(vma, vma_start_virt_pgoff(vma) + delta); } =20 static inline void vma_sub_pgoff(struct vm_area_struct *vma, pgoff_t delta) { vma_assert_can_modify(vma); vma_set_pgoff(vma, vma_start_pgoff(vma) - delta); + vma_set_virt_pgoff(vma, vma_start_virt_pgoff(vma) - delta); +} + +/** + * vma_anon_pgoff_addr() - Calculates the absolute anonymous page offset of + * @address. + * @vma: The VMA whose anonymous page offset is required. + * @address: The address whose absolute page offset is required. + * + * If the VMA is a shared file-backed mapping, then the file-based page of= fset + * is returned. + * + * Otherwise, the virtual page offset is returned. + * + * This means that shared file-backed mappings are correctly merged based = on + * their file page offset compatibility. + * + * Returns: The absolute anonymous page offset of @address within @vma. + */ +static inline pgoff_t vma_anon_pgoff_addr(const struct vm_area_struct *vma, + unsigned long address) +{ + if (vma_test(vma, VMA_SHARED_BIT)) + return linear_page_index(vma, address); + + return linear_virt_page_index(vma, address); +} + +/** + * vma_start_anon_pgoff() - Calculates the absolute anonymous page offset = used + * for purposes of merge compatibility. + * @vma: The VMA whose anonymous page offset is required. + * + * See vma_anon_pgoff_addr(). + * + * Returns: The absolute anonymous page offset of @vma for purposes of mer= ging. + */ +static inline pgoff_t vma_start_anon_pgoff(const struct vm_area_struct *vm= a) +{ + return vma_anon_pgoff_addr(vma, vma->vm_start); } =20 -#define VMG_STATE(name, mm_, vmi_, start_, end_, vma_flags_, pgoff_) \ +/** + * vma_end_anon_pgoff() - Calculates the absolute exclusive end anonymous = page + * offset used for purposes of merge compatibility. + * @vma: The VMA whosse anonymous end page offset is required. + * + * See vma_start_anon_pgoff(). + * + * Returns: The absolute exclusive end anonymous page offset of @vma for + * purposes of merging. + */ +static inline pgoff_t vma_end_anon_pgoff(const struct vm_area_struct *vma) +{ + return vma_start_anon_pgoff(vma) + vma_pages(vma); +} + +#define VMG_STATE(name, mm_, vmi_, start_, end_, vma_flags_, pgoff_, anon_= pgoff_) \ + struct vma_merge_struct name =3D { \ + .mm =3D mm_, \ + .vmi =3D vmi_, \ + .start =3D start_, \ + .end =3D end_, \ + .vma_flags =3D vma_flags_, \ + .pgoff =3D pgoff_, \ + .anon_pgoff =3D anon_pgoff_, \ + .state =3D VMA_MERGE_START, \ + } + +#define VMG_VMA_STATE(name, vmi_, prev_, vma_, start_, end_) \ struct vma_merge_struct name =3D { \ - .mm =3D mm_, \ + .mm =3D vma_->vm_mm, \ .vmi =3D vmi_, \ + .prev =3D prev_, \ + .middle =3D vma_, \ + .next =3D NULL, \ .start =3D start_, \ .end =3D end_, \ - .vma_flags =3D vma_flags_, \ - .pgoff =3D pgoff_, \ + .vm_flags =3D vma_->vm_flags, \ + .pgoff =3D linear_page_index(vma_, start_), \ + .anon_pgoff =3D vma_anon_pgoff_addr(vma_, start_), \ + .file =3D vma_->vm_file, \ + .anon_vma =3D vma_->anon_vma, \ + .policy =3D vma_policy(vma_), \ + .uffd_ctx =3D vma_->vm_userfaultfd_ctx, \ + .anon_name =3D anon_vma_name(vma_), \ .state =3D VMA_MERGE_START, \ } =20 -#define VMG_VMA_STATE(name, vmi_, prev_, vma_, start_, end_) \ - struct vma_merge_struct name =3D { \ - .mm =3D vma_->vm_mm, \ - .vmi =3D vmi_, \ - .prev =3D prev_, \ - .middle =3D vma_, \ - .next =3D NULL, \ - .start =3D start_, \ - .end =3D end_, \ - .vm_flags =3D vma_->vm_flags, \ - .pgoff =3D linear_page_index(vma_, start_), \ - .file =3D vma_->vm_file, \ - .anon_vma =3D vma_->anon_vma, \ - .policy =3D vma_policy(vma_), \ - .uffd_ctx =3D vma_->vm_userfaultfd_ctx, \ - .anon_name =3D anon_vma_name(vma_), \ - .state =3D VMA_MERGE_START, \ - } - #ifdef CONFIG_DEBUG_VM_MAPLE_TREE void validate_mm(struct mm_struct *mm); #else @@ -520,7 +589,7 @@ void unlink_file_vma_batch_add(struct unlink_vma_file_b= atch *vb, =20 struct vm_area_struct *copy_vma(struct vm_area_struct **vmap, unsigned long addr, unsigned long len, pgoff_t pgoff, - bool *need_rmap_locks); + pgoff_t anon_pgoff, bool *need_rmap_locks); =20 struct anon_vma *find_mergeable_anon_vma(struct vm_area_struct *vma); =20 diff --git a/mm/vma_exec.c b/mm/vma_exec.c index 7af1260689b9..586c52155942 100644 --- a/mm/vma_exec.c +++ b/mm/vma_exec.c @@ -41,7 +41,7 @@ int relocate_vma_down(struct vm_area_struct *vma, unsigne= d long shift) unsigned long new_end =3D old_end - shift; VMA_ITERATOR(vmi, mm, new_start); VMG_STATE(vmg, mm, &vmi, new_start, old_end, EMPTY_VMA_FLAGS, - vma_start_pgoff(vma)); + vma_start_pgoff(vma), vma_start_anon_pgoff(vma)); struct vm_area_struct *next; struct mmu_gather tlb; PAGETABLE_MOVE(pmc, vma, vma, old_start, new_start, length); diff --git a/tools/testing/vma/shared.c b/tools/testing/vma/shared.c index bea9ea6db02a..f410bb6f858e 100644 --- a/tools/testing/vma/shared.c +++ b/tools/testing/vma/shared.c @@ -23,7 +23,8 @@ struct vm_area_struct *alloc_vma(struct mm_struct *mm, =20 vma->vm_start =3D start; vma->vm_end =3D end; - vma->vm_pgoff =3D pgoff; + vma_set_pgoff(vma, pgoff); + vma_set_virt_pgoff(vma, start >> PAGE_SHIFT); vma->flags =3D vma_flags; vma_assert_detached(vma); =20 diff --git a/tools/testing/vma/tests/merge.c b/tools/testing/vma/tests/merg= e.c index e357accc8499..48418b82b01d 100644 --- a/tools/testing/vma/tests/merge.c +++ b/tools/testing/vma/tests/merge.c @@ -45,6 +45,7 @@ void vmg_set_range(struct vma_merge_struct *vmg, unsigned= long start, vmg->start =3D start; vmg->end =3D end; vmg->pgoff =3D pgoff; + vmg->anon_pgoff =3D start >> PAGE_SHIFT; vmg->vma_flags =3D vma_flags; =20 vmg->just_expand =3D false; @@ -108,6 +109,7 @@ static bool test_simple_merge(void) .end =3D 0x2000, .vma_flags =3D vma_flags, .pgoff =3D 1, + .anon_pgoff =3D 1, }; =20 ASSERT_FALSE(attach_vma(&mm, vma_left)); @@ -1431,7 +1433,7 @@ static bool test_expand_only_mode(void) struct mm_struct mm =3D {}; VMA_ITERATOR(vmi, &mm, 0); struct vm_area_struct *vma_prev, *vma; - VMG_STATE(vmg, &mm, &vmi, 0x5000, 0x9000, vma_flags, 5); + VMG_STATE(vmg, &mm, &vmi, 0x5000, 0x9000, vma_flags, 5, 5); =20 /* * Place a VMA prior to the one we're expanding so we assert that we do diff --git a/tools/testing/vma/tests/vma.c b/tools/testing/vma/tests/vma.c index 754a2da06321..7ca5289e0f95 100644 --- a/tools/testing/vma/tests/vma.c +++ b/tools/testing/vma/tests/vma.c @@ -38,7 +38,7 @@ static bool test_copy_vma(void) /* Move backwards and do not merge. */ =20 vma =3D alloc_and_link_vma(&mm, 0x3000, 0x5000, 3, vma_flags); - vma_new =3D copy_vma(&vma, 0, 0x2000, 0, &need_locks); + vma_new =3D copy_vma(&vma, 0, 0x2000, 0, 3, &need_locks); ASSERT_NE(vma_new, vma); ASSERT_EQ(vma_new->vm_start, 0); ASSERT_EQ(vma_new->vm_end, 0x2000); @@ -51,7 +51,7 @@ static bool test_copy_vma(void) =20 vma =3D alloc_and_link_vma(&mm, 0, 0x2000, 0, vma_flags); vma_next =3D alloc_and_link_vma(&mm, 0x6000, 0x8000, 6, vma_flags); - vma_new =3D copy_vma(&vma, 0x4000, 0x2000, 4, &need_locks); + vma_new =3D copy_vma(&vma, 0x4000, 0x2000, 4, 4, &need_locks); vma_assert_attached(vma_new); =20 ASSERT_EQ(vma_new, vma_next); diff --git a/tools/testing/vma/vma_internal.h b/tools/testing/vma/vma_inter= nal.h index 4f6c5666ac07..8a48b231aa7a 100644 --- a/tools/testing/vma/vma_internal.h +++ b/tools/testing/vma/vma_internal.h @@ -53,6 +53,7 @@ typedef __bitwise unsigned int vm_fault_t; =20 #define VM_WARN_ON(_expr) (WARN_ON(_expr)) #define VM_WARN_ON_ONCE(_expr) (WARN_ON_ONCE(_expr)) +#define VM_WARN_ON_ONCE_VMA(_expr, _vma) (WARN_ON_ONCE(_expr)) #define VM_WARN_ON_VMG(_expr, _vmg) (WARN_ON(_expr)) #define VM_BUG_ON(_expr) (BUG_ON(_expr)) #define VM_BUG_ON_VMA(_expr, _vma) (BUG_ON(_expr)) --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 01D8037F732; Fri, 17 Jul 2026 18:23:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312600; cv=none; b=dR27kW6I4qwHdgBZleXQoPyMrQV4i/Zw9lO+frkrbsTgWHbylMvchRW7kYLfflp25A1XjBmsrnKxfy4XAk0iY0Q622l1SGDTypGGzUNmhOkWey0dXQPST9udb6z3dLJq0PzMYxOdwUNxd7UDdLhtlWPw5sAl9XywSUdVeN3lAoo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312600; c=relaxed/simple; bh=M2DuEa+QMY+cZQChTlOi3eUq8pPFDxokrv+qbMEdb+c=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=EMPvGB/T/zOuDYy68AWeXcZWZ+f+pNogC8BYLWQnTcdV8Zn8PNLqUQSyGpqPugmepniS1yPL1OJ4mAo6Tbyh/MMY6e/7pbniztghqj2NEjaDo+fAa6DKJdKm4Q75Cnin7jekru/o2Yy+IX0a5TmudPeKPhJGCf43aOlPTpS8lQA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=dU5t7A9x; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="dU5t7A9x" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DC5601F00A3A; Fri, 17 Jul 2026 18:23:06 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312595; bh=K9J65l4/3tAXQjoApUsRo6wpq7Iiti7b+Q54PqfKy2U=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=dU5t7A9xxE1rqZ1PA5/5f+2Wu21cdQEJUSBHMauC5kMnYT1kDL/cyheogHKuFwzqj nvvku9uzuFSlcxKvF6cMFaQKmuQ+XEQbHDw76+48dB8c6jtLQ6mzA9yFY2jYf9S8Qk RMT42K1WAT+Rd9rPdoqum4UcK1jGG/mognui1GF0gBoawsjVNZlFz+ahHQKymI6PH/ 43xsPxmL5ftrK1PmXmi/TQzpS23MDE5C/fzN+mP39M2jWR0qKWuJ7gqIPvmeB9/JNy i6O/EuYQ7aNMaKns/e8vPC/7Yh18CEFXXoq/Ekqu/sxlQUiqneWIVkyE292QIejW/q dah186Oj1sn/w== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:45 +0100 Subject: [PATCH 07/15] mm/rmap: track whether the page VMA mapped walk is anonymous Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-7-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3223; i=ljs@kernel.org; h=from:subject:message-id; bh=M2DuEa+QMY+cZQChTlOi3eUq8pPFDxokrv+qbMEdb+c=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivblzmr2WeH22/fBFcFjy2Ou3tW+uCCw3bjgQOHjH 0eXX1S06yhlYRDjYpAVU2R5/kV8f5BI2LzOC/5uMHNYmUCGMHBxCsBE/MwZ/opfdj+rxXiu75RM 47b+P7fu5Oz4s/ceV0vY+g871Itau4IZ/sfvNxJQ6JsxdS6DQqDQ/FbW4832Wlw/JYqdVn11E7q txAIA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Update the page_vma_mapped_walk structure to track whether the walk is over an anonymous folio or not. This is necessary in order to determine the correct VMA page offset (virtual or not) in vma_address_end() ready for a subsequent change which adjusts which page offset to use depending on this parameter, and update the comment slightly. No functional change intended. Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/rmap.h | 2 ++ mm/internal.h | 13 ++++++++----- mm/rmap.c | 3 +++ 3 files changed, 13 insertions(+), 5 deletions(-) diff --git a/include/linux/rmap.h b/include/linux/rmap.h index 8dc0871e5f00..a48ae9575bd2 100644 --- a/include/linux/rmap.h +++ b/include/linux/rmap.h @@ -871,6 +871,7 @@ struct page_vma_mapped_walk { pte_t *pte; spinlock_t *ptl; unsigned int flags; + bool is_anon_walk; }; =20 #define DEFINE_FOLIO_VMA_WALK(name, _folio, _vma, _address, _flags) \ @@ -881,6 +882,7 @@ struct page_vma_mapped_walk { .vma =3D _vma, \ .address =3D _address, \ .flags =3D _flags, \ + .is_anon_walk =3D folio_test_anon(_folio), \ } =20 static inline void page_vma_mapped_walk_done(struct page_vma_mapped_walk *= pvmw) diff --git a/mm/internal.h b/mm/internal.h index 73f90aaa979d..7ed8bde7890d 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1079,22 +1079,25 @@ static inline unsigned long vma_anon_address(const = struct vm_area_struct *vma, } =20 /* - * Then at what user virtual address will none of the range be found in vm= a? + * At what user virtual address will none of the range be found in vma? * Assumes that vma_address() already returned a good starting address. */ static inline unsigned long vma_address_end(struct page_vma_mapped_walk *p= vmw) { - struct vm_area_struct *vma =3D pvmw->vma; - pgoff_t pgoff; + const struct vm_area_struct *vma =3D pvmw->vma; + const pgoff_t pgoff =3D pvmw->pgoff; + pgoff_t pgoff_vma_start; unsigned long address; + pgoff_t pgoff_end; =20 /* Common case, plus ->pgoff is invalid for KSM */ if (pvmw->nr_pages =3D=3D 1) return pvmw->address + PAGE_SIZE; =20 - pgoff =3D pvmw->pgoff + pvmw->nr_pages; + pgoff_vma_start =3D vma_start_pgoff(vma); + pgoff_end =3D pgoff + pvmw->nr_pages; address =3D vma->vm_start + - ((pgoff - vma_start_pgoff(vma)) << PAGE_SHIFT); + ((pgoff_end - pgoff_vma_start) << PAGE_SHIFT); /* Check for address beyond vma (or wrapped through 0?) */ if (address < vma->vm_start || address > vma->vm_end) address =3D vma->vm_end; diff --git a/mm/rmap.c b/mm/rmap.c index 5798427d007f..cfd4195ff1ae 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -1240,8 +1240,10 @@ static bool mapping_wrprotect_range_one(struct folio= *folio, .vma =3D vma, .address =3D address, .flags =3D PVMW_SYNC, + .is_anon_walk =3D false, }; =20 + VM_WARN_ON_ONCE(folio_test_anon(folio)); state->cleaned +=3D page_vma_mkclean_one(&pvmw); =20 return true; @@ -1317,6 +1319,7 @@ int pfn_mkclean_range(unsigned long pfn, unsigned lon= g nr_pages, pgoff_t pgoff, .pgoff =3D pgoff, .vma =3D vma, .flags =3D PVMW_SYNC, + .is_anon_walk =3D false, }; =20 if (invalid_mkclean_vma(vma, NULL)) --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 38D7B349CD7; Fri, 17 Jul 2026 18:23:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312608; cv=none; b=IMuLlF2t4ND0Mmf9F22GHu5rwUYRf9Md6lHJdVNDUD7MolOmmb57PDVR2cfNR8m6SVqe8O7Yii06q6lzjjUaZNije/e96ivhxn7nWBaELysKHkjRfzqfulsXz9tw2fOB8VePXMsd0Qcs+T/kFT5oRoP75D6jrF2n1ZpxBteE8W4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312608; c=relaxed/simple; bh=6pAegtlPfBQ3G3CGe2dDWkuXSlF6+Udh4ZexgB7rXRw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=tQe+EDSzWWt/y4UnWKamD9l9h7geQHIpeKKdSyDvtXPugOoHKLAXwq0vn+R5G9U+fvfDvhMAbXOv1AZN240tD5S0x0vvrdwKl+AkKkTb5qHUjejnM4eESHOhYI2uK+k8ywvfnCX0XUgfbUxO3qpuHu56NToE6vVX6J7R4ppEfiQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=iZ50G7B9; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="iZ50G7B9" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 42F6B1F00A3D; Fri, 17 Jul 2026 18:23:16 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312605; bh=LAYYiM2rpTmaPPo2VlPYNoqH0gFENrcwglvYiZo/Bls=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=iZ50G7B9X1eEJ+gjLEVKfTpi5LG2qZTmjqad3Qg4Kly7DE9FC9veVfrNbHX7mY8UC 3koAjIihABNFqIt/r2bxQ1SLL4TO1GVk1j9dn39GBDtf+quhPEqqWZ77guX1Ls2nsT YozctpZgHRXUVZZHfCLBJi9p/6ZGu26Nmto8hBvm4AZhcRnNIwd5TAUU34CjzDuqn/ 7W0HPy5aPvO7h6PCL3F8CK27lvT4Jgq6JB7PkiVHEF+jaK9xFpygzAsLbdo3xIV1Bf dEbDUlcZxYXUdY5GdfQJhuqaYzEKWmrR2d2ZOwg+kcLFljrJ91dXHOC0flCwocph5j 3pY8V+ezEsFnQ== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:46 +0100 Subject: [PATCH 08/15] mm: introduce and use linear_folio_page_index() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-8-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4022; i=ljs@kernel.org; h=from:subject:message-id; bh=6pAegtlPfBQ3G3CGe2dDWkuXSlF6+Udh4ZexgB7rXRw=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivbNurLzn7vj8qY5HOe75/ZoLd3/KXmjaJgIx53Zt 5XS/hxN6ShlYRDjYpAVU2R5/kV8f5BI2LzOC/5uMHNYmUCGMHBxCsBE3ucx/LN8F3ziz3yR9Zu/ Cv2pSqxmUuMPWvw8SOXG2ja2kMJNzPsY/sfZJlVsiDxTHHpo4e85L9z2u+wVlD3ztN6p95/s/gt XSpgB X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This function is, for now, a placeholder; it will be used in future to determine whether to use the virtual page index or not, based on whether the folio is anonymous or not. Currently it simply wraps linear_page_index(), so this does not change behaviour. We update callers that will, once the change is introduced to track anonymous folios by virtual page offset if MAP_PRIVATE file-backed, need to determine which index to use based on folio type. No functional change intended. Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/pagemap.h | 18 ++++++++++++++++++ mm/huge_memory.c | 3 ++- mm/migrate.c | 6 ++++-- mm/userfaultfd.c | 6 ++++-- 4 files changed, 28 insertions(+), 5 deletions(-) diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h index 81da91e103b3..dab355eda0ef 100644 --- a/include/linux/pagemap.h +++ b/include/linux/pagemap.h @@ -1142,6 +1142,24 @@ static inline pgoff_t linear_virt_page_index(const s= truct vm_area_struct *vma, return pgoff; } =20 +/** + * linear_folio_page_index() - Determine the absolute page offset of + * @address within @vma from @folio. + * @folio: The folio whose linear page index is sought. + * @vma: The VMA in which @address resides. + * @address: The address whose absolute page offset is required. + * + * For compatibility, currently identical to linear_page_index(). + * + * Returns: The absolute page offset of @address within @vma. + */ +static inline pgoff_t linear_folio_page_index(const struct folio *folio, + const struct vm_area_struct *vma, + const unsigned long address) +{ + return linear_page_index(vma, address); +} + struct wait_page_key { struct folio *folio; int bit_nr; diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 9b1f3b24f7e0..abc65d608c23 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2887,7 +2887,8 @@ int move_pages_huge_pmd(struct mm_struct *mm, pmd_t *= dst_pmd, pmd_t *src_pmd, pm } =20 folio_move_anon_rmap(src_folio, dst_vma); - src_folio->index =3D linear_page_index(dst_vma, dst_addr); + src_folio->index =3D linear_folio_page_index(src_folio, dst_vma, + dst_addr); =20 _dst_pmd =3D folio_mk_pmd(src_folio, dst_vma->vm_page_prot); /* Follow mremap() behavior and treat the entry dirty after the move */ diff --git a/mm/migrate.c b/mm/migrate.c index 222c8c15f782..37fe7a9b3fac 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -363,8 +363,10 @@ static bool remove_migration_pte(struct folio *folio, unsigned long idx =3D 0; =20 /* pgoff is invalid for ksm pages, but they are never large */ - if (folio_test_large(folio) && !folio_test_hugetlb(folio)) - idx =3D linear_page_index(vma, pvmw.address) - pvmw.pgoff; + if (folio_test_large(folio) && !folio_test_hugetlb(folio)) { + idx +=3D linear_folio_page_index(folio, vma, pvmw.address); + idx -=3D pvmw.pgoff; + } new =3D folio_page(folio, idx); =20 #ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index 8fd24c8b428e..258b03182a78 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -1352,7 +1352,8 @@ static long move_present_ptes(struct mm_struct *mm, } =20 folio_move_anon_rmap(src_folio, dst_vma); - src_folio->index =3D linear_page_index(dst_vma, dst_addr); + src_folio->index =3D linear_folio_page_index(src_folio, dst_vma, + dst_addr); =20 orig_dst_pte =3D folio_mk_pte(src_folio, dst_vma->vm_page_prot); /* Set soft dirty bit so userspace can notice the pte was moved */ @@ -1428,7 +1429,8 @@ static int move_swap_pte(struct mm_struct *mm, struct= vm_area_struct *dst_vma, */ if (src_folio) { folio_move_anon_rmap(src_folio, dst_vma); - src_folio->index =3D linear_page_index(dst_vma, dst_addr); + src_folio->index =3D linear_folio_page_index(src_folio, dst_vma, + dst_addr); } else { /* * Check if the swap entry is cached after acquiring the src_pte --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C9CEB2233A; Fri, 17 Jul 2026 18:23:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312616; cv=none; b=NS6/qb2H4G8uqw17IjHd/ofNnr32IXQp91+N+pJHsX6+vnBzfd/SpUrG2zcfDJwqx/IcxFyXWDApec1LwIFHqNqzuRjHgQW6IQJgonPBxkiZTZnOplWML4ZvEI+PZmC8NtHIXfjuJ6jHPSkL6eAD2jtIjvSzdbpTwny+dkD6Ay0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312616; c=relaxed/simple; bh=trcZDCZGb1+fFeGFvg3327sbE4yWDvlmUSh3MOsJRQI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=cUXbPxvmzQLRNlJFxLnBAwL+BKaHaQgTdZBzRkV0sanVxmEErMLsm5J00dwBxRasHq5DlHFHb7IcZoUt28vMQZzo8jvnRhBo5B0wL2FF4yL2FoTRrrtLgCP2xqGH7oFlpLErd761Qxvq0MqB1XUOsqW4GQt7FcUxji/raSgRwZ4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=dmE5lYer; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="dmE5lYer" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 9C2201F000E9; Fri, 17 Jul 2026 18:23:25 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312614; bh=THcSTCW+eH3M326ttZSOcln1sIFibD7ABxlX6LgmtOs=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=dmE5lYerIgZlq5rUsfess6SgtAzDQjiz+vmhBsj9u0sgXzson+piHCvZjmwwkCIry rpM6SwHqZiqpHrN9sIx83cI60ZK1oPD6oX/pzpFtMc4W4yNFtjEV1HG3ai3WJQZbw/ lsL2av1Yu6aUVFzs6xrCndBa0FTTIulWTrXCq6GjbeOkl6OTreNYCjFndDGxhsZXGD CPwcvnY+ezL2LQh+u1QclXrR6QF2NkIn0bnquRAUxmW6DS4ukP2HnLO/fXqMO/BzKL oOfxBZCo+iQeS3kdRVLPbSV7Ail+2aHQN1d2pxplUOCBmUJPEUnKmp1z0O5yQmy9Vy OQgE1nxOhqMXg== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:47 +0100 Subject: [PATCH 09/15] mm/rmap: use virt pgoff for MAP_PRIVATE file-backed anon folios Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-9-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=11839; i=ljs@kernel.org; h=from:subject:message-id; bh=trcZDCZGb1+fFeGFvg3327sbE4yWDvlmUSh3MOsJRQI=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivZ57q9ZtbpYadn2wNNMB1ZMz+gWNhc8INM87cKPl q2ncrw3dpSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAiwocY/lmJpbyIu7O3NePD 9kOTl6ekpsWse60rU5g/bV++GWNX2DaGfyZR10XCerSXrp0T+0I9QNix8K6rX6Ko6rr9amsEN0/ y4gEA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Currently, anonymous folios belonging to CoW'd MAP_PRIVATE file-backed mappings are indexed by their page offset within the file in which they were originally mapped. This differs from anonymous folios belonging to pure anon mappings which are indexed by their virtual page offset (the address at which they'd belong in the VMA when first faulted). This change fixes this inconsistency, always indexing anonymous folios by their virtual page offset regardless of the VMA to which they belong. We have laid the foundations for making this change to the point where we need only 'switch it on', and this patch switches it on by: * Using linear_virt_page_index() in __folio_set_anon() to assign the folio's index to the anonymous linear index rather than the file-backed one. * Otherwise using linear_virt_page_index() in all instances where anonymous folios are being referenced or manipulated. * Replacing vma_address() with vma_filebacked_address() or vma_anon_address() as appropriate. * Updating the rmap lock logic in copy_vma() to also account for virtual page offsets. * Updating the merging logic to check that virtual page offsets are aligned as well as filebacked ones for anonymous or MAP_PRIVATE file-backed VMAs. * Updating linear_folio_page_index() to invoke linear_virt_page_index() if the folio is anonymous. * Correcting folio_within_range() to use virtual page offset for anonymous folios. This will have no impact on merging of anonymous VMAs or shared file-backed VMAs, whose page offset and anonymous page offset will be identical. However, MAP_PRIVATE file-backed mappings must now be aligned on virtual page offset as well. In most instances this should have no impact on merging of file-backed mappings, which are usually not merged all that often, let alone MAP_PRIVATE mapped ones, and rarely remapped and faulted before being moved back in place (the case in which a merge may now fail). This change lays the foundations for future scalable CoW work which needs to track at least some remaps. This change means that most remap tracking can be avoided, and in nearly all cases the anonymous page offset can be used to quickly find the VMA in an mm. Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/pagemap.h | 9 ++++++++- mm/internal.h | 28 +++++++++------------------- mm/interval_tree.c | 4 ++-- mm/ksm.c | 7 ++++--- mm/page_vma_mapped.c | 2 +- mm/rmap.c | 12 ++++++------ mm/vma.c | 14 ++++++++++++-- 7 files changed, 42 insertions(+), 34 deletions(-) diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h index dab355eda0ef..90130f28e7a8 100644 --- a/include/linux/pagemap.h +++ b/include/linux/pagemap.h @@ -1149,7 +1149,11 @@ static inline pgoff_t linear_virt_page_index(const s= truct vm_area_struct *vma, * @vma: The VMA in which @address resides. * @address: The address whose absolute page offset is required. * - * For compatibility, currently identical to linear_page_index(). + * Determines whether to obtain the virtual linear page index based on whe= ther + * @folio is anonymous or not. + * + * See the descriptions of linear_virt_page_index() and linear_page_index(= ) for + * details of each. * * Returns: The absolute page offset of @address within @vma. */ @@ -1157,6 +1161,9 @@ static inline pgoff_t linear_folio_page_index(const s= truct folio *folio, const struct vm_area_struct *vma, const unsigned long address) { + if (folio_test_anon(folio)) + return linear_virt_page_index(vma, address); + return linear_page_index(vma, address); } =20 diff --git a/mm/internal.h b/mm/internal.h index 7ed8bde7890d..22f63cb85bd3 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -933,7 +933,8 @@ folio_within_range(struct folio *folio, struct vm_area_= struct *vma, return false; =20 pgoff_folio =3D folio_pgoff(folio); - pgoff_vma_start =3D vma_start_pgoff(vma); + pgoff_vma_start =3D folio_test_anon(folio) ? + vma_start_virt_pgoff(vma) : vma_start_pgoff(vma); =20 if (start < vma->vm_start) start =3D vma->vm_start; @@ -1044,23 +1045,8 @@ static inline unsigned long vma_filebacked_address(c= onst struct vm_area_struct * } =20 /** - * vma_address - Find the virtual address a page range is mapped at. - * @vma: The vma which maps this object. - * @pgoff: The page offset within its object. - * @nr_pages: The number of pages to consider. - * - * If any page in this range is mapped by this VMA, return the first addre= ss - * where any of these pages appear. Otherwise, return -EFAULT. - */ -static inline unsigned long vma_address(const struct vm_area_struct *vma, - pgoff_t pgoff, unsigned long nr_pages) -{ - return __vma_address(vma, pgoff, vma_start_pgoff(vma), nr_pages); -} - -/** - * vma_anon_address - Find the address an anonymous folio with index @pgof= f_virt - * is mapped at. + * vma_anon_address - Find the virtual address an anonymous page range is = mapped + * at. * @vma: The vma which maps this object. * @pgoff_virt: The virtual page index belonging to the folio. * @nr_pages: The number of pages to consider. @@ -1094,7 +1080,11 @@ static inline unsigned long vma_address_end(struct p= age_vma_mapped_walk *pvmw) if (pvmw->nr_pages =3D=3D 1) return pvmw->address + PAGE_SIZE; =20 - pgoff_vma_start =3D vma_start_pgoff(vma); + if (pvmw->is_anon_walk) + pgoff_vma_start =3D vma_start_virt_pgoff(vma); + else + pgoff_vma_start =3D vma_start_pgoff(vma); + pgoff_end =3D pgoff + pvmw->nr_pages; address =3D vma->vm_start + ((pgoff_end - pgoff_vma_start) << PAGE_SHIFT); diff --git a/mm/interval_tree.c b/mm/interval_tree.c index 3ae9e106d3af..26b8437e3b1b 100644 --- a/mm/interval_tree.c +++ b/mm/interval_tree.c @@ -83,12 +83,12 @@ mapping_rmap_tree_iter_next(struct vm_area_struct *vma, =20 static pgoff_t avc_start_pgoff(struct anon_vma_chain *avc) { - return vma_start_pgoff(avc->vma); + return vma_start_virt_pgoff(avc->vma); } =20 static pgoff_t avc_last_pgoff(struct anon_vma_chain *avc) { - return vma_last_pgoff(avc->vma); + return vma_last_virt_pgoff(avc->vma); } =20 INTERVAL_TREE_DEFINE(struct anon_vma_chain, rb, pgoff_t, rb_subtree_last, diff --git a/mm/ksm.c b/mm/ksm.c index 47006f494fcb..b9b32f7bfb69 100644 --- a/mm/ksm.c +++ b/mm/ksm.c @@ -1625,7 +1625,8 @@ static int try_to_merge_with_ksm_page(struct ksm_rmap= _item *rmap_item, * stable_tree, break_cow() will clean it up. */ rmap_item->anon_vma =3D vma->anon_vma; - rmap_item->linear_page_index =3D linear_page_index(vma, rmap_item->addres= s); + /* The VMA is always anon/MAP_PRIVATE-file backed so use anon index. */ + rmap_item->linear_page_index =3D linear_virt_page_index(vma, rmap_item->a= ddress); get_anon_vma(vma->anon_vma); out: mmap_read_unlock(mm); @@ -3152,7 +3153,7 @@ struct folio *ksm_might_need_to_copy(struct folio *fo= lio, return folio; /* no need to copy it */ } else if (!anon_vma) { return folio; /* no need to copy it */ - } else if (folio->index =3D=3D linear_page_index(vma, addr) && + } else if (folio->index =3D=3D linear_virt_page_index(vma, addr) && anon_vma->root =3D=3D vma->anon_vma->root) { return folio; /* still no need to copy it */ } @@ -3222,7 +3223,7 @@ void rmap_walk_ksm(struct folio *folio, struct rmap_w= alk_control *rwc) /* * Currently, KSM folios are always small folios, so it's * sufficient to search for a single page. We can simply use - * the linear_page_index of the original de-duplicate + * the linear_virt_page_index of the original de-duplicate * anonymous page that we remembered in the rmap_item while * de-duplicating. Note that mremap() always de-duplicates KSM * folios: so if there was mremap() in our parent or our child, diff --git a/mm/page_vma_mapped.c b/mm/page_vma_mapped.c index 081e483cc7bf..4e964545e5e8 100644 --- a/mm/page_vma_mapped.c +++ b/mm/page_vma_mapped.c @@ -365,7 +365,7 @@ unsigned long page_mapped_in_vma(const struct page *pag= e, }; =20 if (folio_test_anon(folio)) - pvmw.address =3D vma_address(vma, pgoff, 1); + pvmw.address =3D vma_anon_address(vma, pgoff, 1); else pvmw.address =3D vma_filebacked_address(vma, pgoff, 1); if (pvmw.address =3D=3D -EFAULT) diff --git a/mm/rmap.c b/mm/rmap.c index cfd4195ff1ae..6854baf4b877 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -866,7 +866,7 @@ unsigned long page_address_in_vma(const struct folio *f= olio, vma->anon_vma->root !=3D anon_vma->root) return -EFAULT; /* KSM folios don't reach here because of the !anon_vma check */ - return vma_address(vma, page_pgoff(folio, page), 1); + return vma_anon_address(vma, page_pgoff(folio, page), 1); } else if (!vma->vm_file) { return -EFAULT; } else if (vma->vm_file->f_mapping !=3D folio->mapping) { @@ -1486,7 +1486,7 @@ static void __folio_set_anon(struct folio *folio, str= uct vm_area_struct *vma, */ anon_vma =3D (void *) anon_vma + FOLIO_MAPPING_ANON; WRITE_ONCE(folio->mapping, (struct address_space *) anon_vma); - folio->index =3D linear_page_index(vma, address); + folio->index =3D linear_virt_page_index(vma, address); } =20 /** @@ -1513,8 +1513,8 @@ static void __page_check_anon_rmap(const struct folio= *folio, */ VM_BUG_ON_FOLIO(folio_anon_vma(folio)->root !=3D vma->anon_vma->root, folio); - VM_BUG_ON_PAGE(page_pgoff(folio, page) !=3D linear_page_index(vma, addres= s), - page); + VM_BUG_ON_PAGE(page_pgoff(folio, page) !=3D + linear_virt_page_index(vma, address), page); } =20 static __always_inline void __folio_add_anon_rmap(struct folio *folio, @@ -3041,10 +3041,10 @@ static void rmap_walk_anon(struct folio *folio, pgoff_end =3D pgoff_start + folio_nr_pages(folio) - 1; anon_rmap_tree_foreach(avc, anon_vma, pgoff_start, pgoff_end) { struct vm_area_struct *vma =3D avc->vma; - unsigned long address =3D vma_address(vma, pgoff_start, + const unsigned long address =3D vma_anon_address(vma, pgoff_start, folio_nr_pages(folio)); =20 - VM_BUG_ON_VMA(address =3D=3D -EFAULT, vma); + VM_WARN_ON_ONCE_VMA(address =3D=3D -EFAULT, vma); cond_resched(); =20 if (rwc->invalid_vma && rwc->invalid_vma(vma, rwc->arg)) diff --git a/mm/vma.c b/mm/vma.c index e45d43493446..6dc095d03382 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -233,6 +233,8 @@ static bool can_vma_merge_before(struct vma_merge_struc= t *vmg) return false; if (vmg_end_pgoff(vmg) !=3D vma_start_pgoff(vmg->next)) return false; + if (vmg_end_anon_pgoff(vmg) !=3D vma_start_anon_pgoff(vmg->next)) + return false; return true; } =20 @@ -253,6 +255,8 @@ static bool can_vma_merge_after(struct vma_merge_struct= *vmg) return false; if (vma_end_pgoff(vmg->prev) !=3D vmg_start_pgoff(vmg)) return false; + if (vma_end_anon_pgoff(vmg->prev) !=3D vmg_start_anon_pgoff(vmg)) + return false; return true; } =20 @@ -1988,7 +1992,8 @@ struct vm_area_struct *copy_vma(struct vm_area_struct= **vmap, *vmap =3D vma =3D new_vma; } *need_rmap_locks =3D - (vma_start_pgoff(new_vma) <=3D vma_start_pgoff(vma)); + (vma_start_pgoff(new_vma) <=3D vma_start_pgoff(vma)) || + (vma_start_anon_pgoff(new_vma) <=3D vma_start_anon_pgoff(vma)); } else { new_vma =3D vm_area_dup(vma); if (!new_vma) @@ -2059,7 +2064,12 @@ static int anon_vma_compatible(struct vm_area_struct= *a, struct vm_area_struct * if (!vma_flags_empty(&diff)) return false; /* Page offset must align. */ - return vma_end_pgoff(a) =3D=3D vma_start_pgoff(b); + if (vma_end_pgoff(a) !=3D vma_start_pgoff(b)) + return false; + /* Anon page offset must align. */ + if (vma_end_anon_pgoff(a) !=3D vma_start_anon_pgoff(b)) + return false; + return true; } =20 /* --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 30B822F8EB4; Fri, 17 Jul 2026 18:23:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312625; cv=none; b=Qgjhoax3RHvsvkUPNF8WGQ4BDL9xLeCEEUjeL5fSbYE3XbRtkxnSCcn8C5XIJX7Yyv6yGRmsWRVUbwXLmaVse7hqfdbYu8lo/bFIJq00OZiR8Zz1ZGHnit8nNSRQB9Idj0DgcaF2EPxgtB3ufQHYii2T8FdS0rbRWkGv5w4H7E4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312625; c=relaxed/simple; bh=2613GakduscTW0HQyjciqfb4Fv/xmKQjLs/wrRcxexM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=bNbGFmQUIMJnt+NdQhZcZRG4v8dnQ3Gkbo4xkjNoOVEKHkg3M4Ku5RM5pKRQWE0g8O+0kaGdfPDrz4syYVFyq1/JVeJmE6m50jXa3HGWAvaEpToQur2DDXgE83J56hCQ2Ltzruy27hexNfA8l5iflgF6HoeuIv17o72jfVDdgss= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=i9kZrGVy; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="i9kZrGVy" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 006091F00A3A; Fri, 17 Jul 2026 18:23:34 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312623; bh=e/iL8+q9nchQYoecoHRozmG0upgg4wUWQpXFNukM+ng=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=i9kZrGVy7LfrF2gGQSO3BKdoe0YIuNz4INwdJaJvFm61yNcpoGBzPrlU00k5ugyrh l6lAXaeqTdjovB0IXT+eL7XxI2vtdg194NykBlvIGaIUs1Fhi7rXGg9vknAhEW8xAx xOl+qXalWG2nob3yFTyUsM/E6MXEU5JunIdeVjYan46xBrf06Zvc8Kb9Uo9l8MPRE4 czOaIRBvoN5mYCi1tUtmstx2reEnop856osJafp3GYSxB8JJHoiCQu0thIo/eXF9Tv +0CweEZGGTwCKAlDBULOs7yelMUgrnXW2ID8MfKVRuDknL26r8RG4zJJ5ndQ84gWbf R7Nzq2ZG3fk3A== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:48 +0100 Subject: [PATCH 10/15] tools/testing/vma: expand VMA merge tests to assert virt pgoff Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-10-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=6335; i=ljs@kernel.org; h=from:subject:message-id; bh=2613GakduscTW0HQyjciqfb4Fv/xmKQjLs/wrRcxexM=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivZ53Ywy26t1/84nP/stiqvOhX69Kqejw3Orf7dW6 vEDbjdmd5SyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAiK9cxMrSs+PZMUNFZ9aLE z2XN/a73ai3Sb3kbNUfzyh5cf+YYwx2G/4ln2f0f+Dy6k8m8ctOns0khFSWbWG0vCSolhgifN96 Rzw8A X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Now we have introduced the VMA virtual page offset attribute and update it when VMAs are manipulated, update VMA merge tests to assert that the virtual page offset is as expected. Signed-off-by: Lorenzo Stoakes (ARM) --- tools/testing/vma/tests/merge.c | 18 ++++++++++++++++++ 1 file changed, 18 insertions(+) diff --git a/tools/testing/vma/tests/merge.c b/tools/testing/vma/tests/merg= e.c index 48418b82b01d..a01107ab8e21 100644 --- a/tools/testing/vma/tests/merge.c +++ b/tools/testing/vma/tests/merge.c @@ -121,6 +121,7 @@ static bool test_simple_merge(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x3000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); ASSERT_FLAGS_SAME_MASK(&vma->flags, vma_flags); =20 detach_free_vma(vma); @@ -153,6 +154,7 @@ static bool test_simple_modify(void) ASSERT_EQ(vma->vm_start, 0x1000); ASSERT_EQ(vma->vm_end, 0x2000); ASSERT_EQ(vma_start_pgoff(vma), 1); + ASSERT_EQ(vma_start_virt_pgoff(vma), 1); =20 /* * Now walk through the three split VMAs and make sure they are as @@ -165,6 +167,7 @@ static bool test_simple_modify(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x1000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); =20 detach_free_vma(vma); vma_iter_clear(&vmi); @@ -174,6 +177,7 @@ static bool test_simple_modify(void) ASSERT_EQ(vma->vm_start, 0x1000); ASSERT_EQ(vma->vm_end, 0x2000); ASSERT_EQ(vma_start_pgoff(vma), 1); + ASSERT_EQ(vma_start_virt_pgoff(vma), 1); =20 detach_free_vma(vma); vma_iter_clear(&vmi); @@ -183,6 +187,7 @@ static bool test_simple_modify(void) ASSERT_EQ(vma->vm_start, 0x2000); ASSERT_EQ(vma->vm_end, 0x3000); ASSERT_EQ(vma_start_pgoff(vma), 2); + ASSERT_EQ(vma_start_virt_pgoff(vma), 2); =20 detach_free_vma(vma); mtree_destroy(&mm.mm_mt); @@ -212,6 +217,7 @@ static bool test_simple_expand(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x3000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); =20 detach_free_vma(vma); mtree_destroy(&mm.mm_mt); @@ -234,6 +240,7 @@ static bool test_simple_shrink(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x1000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); =20 detach_free_vma(vma); mtree_destroy(&mm.mm_mt); @@ -346,6 +353,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x5000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 3); @@ -367,6 +375,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0x6000); ASSERT_EQ(vma->vm_end, 0x9000); ASSERT_EQ(vma_start_pgoff(vma), 6); + ASSERT_EQ(vma_start_virt_pgoff(vma), 6); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 3); @@ -387,6 +396,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x9000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 2); @@ -407,6 +417,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0xa000); ASSERT_EQ(vma->vm_end, 0xc000); ASSERT_EQ(vma_start_pgoff(vma), 0xa); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0xa); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 2); @@ -426,6 +437,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0xc000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 1); @@ -446,6 +458,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0xc000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); =20 detach_free_vma(vma); @@ -808,6 +821,7 @@ static bool test_vma_merge_new_with_close(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x5000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); ASSERT_EQ(vma->vm_ops, &vm_ops); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 2); @@ -868,6 +882,7 @@ static bool __test_merge_existing(bool prev_is_sticky, = bool middle_is_sticky, bo ASSERT_EQ(vma->vm_start, 0x2000); ASSERT_EQ(vma->vm_end, 0x3000); ASSERT_EQ(vma_start_pgoff(vma), 2); + ASSERT_EQ(vma_start_virt_pgoff(vma), 2); ASSERT_TRUE(vma_write_started(vma)); ASSERT_TRUE(vma_write_started(vma_next)); ASSERT_EQ(mm.map_count, 2); @@ -934,6 +949,7 @@ static bool __test_merge_existing(bool prev_is_sticky, = bool middle_is_sticky, bo ASSERT_EQ(vma->vm_start, 0x6000); ASSERT_EQ(vma->vm_end, 0x7000); ASSERT_EQ(vma_start_pgoff(vma), 6); + ASSERT_EQ(vma_start_virt_pgoff(vma), 6); ASSERT_TRUE(vma_write_started(vma_prev)); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 2); @@ -1419,6 +1435,7 @@ static bool test_merge_extend(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x4000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 1); =20 @@ -1459,6 +1476,7 @@ static bool test_expand_only_mode(void) ASSERT_EQ(vma->vm_start, 0x3000); ASSERT_EQ(vma->vm_end, 0x9000); ASSERT_EQ(vma_start_pgoff(vma), 3); + ASSERT_EQ(vma_start_virt_pgoff(vma), 3); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(vma_iter_addr(&vmi), 0x3000); vma_assert_attached(vma); --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 80FAF27587D; Fri, 17 Jul 2026 18:23:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312634; cv=none; b=UkzMylY06wYy5v7NVg8g1Q0yLV6971XgwFQOQZk9yF3Tg6QgtpKudgwFwsaSlDdZi1gMHoLD8Qp+EGFtEJw3eZuln41ytEqQg8tOuAOJy27QN2CiRV/RpVjujRCokgG38t6zhExBYLkrGMP8JDVvBQr8yigo1Lc+HfEBdQZXFnY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312634; c=relaxed/simple; bh=01jUabKK64NHmCAfzlhC9tplycbpQo1xdwDkiXe4HZs=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=uOyq6Ds5C6CU5irUGA6MlaxgLlATzHPedMay8uTxT2vbRXdA7m+vSHgn1xnlyVQEBBd+HTG8FDpjE6WfTA+BtzpoBnPVdChuDkXFvb/uU7MkvfkcAoWxF5W3R5tBzX3M22N8g5Kn7YesLtmshnGCQDkXkfqZap8QicTt4BJKiNU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=oHDmxbPk; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="oHDmxbPk" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 58F601F000E9; Fri, 17 Jul 2026 18:23:44 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312633; bh=TVlIRx355t0P5ERcL5Eqs6eD8BqACD3+tbQgzhdx6t4=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=oHDmxbPkeZ2daJD26fnM+oSVULS5ZBTBm6laTClU9DxY3FQ2nY+UbQ0IESJ40tz25 AUwaH7OU9xiiEdGVIcHG4CHOwi5w6XK9DnbdVU+QKhtDUlHye7LQ6HLCTuoRzsHT7S J72/Tv1VCcCLpctWACUDNPH7BZoY4UKdrQrKSaW8miBkU8G+UWtYrFO4WugRv4k4pW GOl7REesFJUyAosHDnPQkgXY86dAWEgYQw7mXZdjxnBdsqntB4uQ0haOBciZ4x9cd3 sNZwF7UdnmPM0r0btx+i+obWK9DgjevDwoxiANm8YWp3biTuzvIqoAggy6ZJNmQiwF olIYKjnWXCeBw== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:49 +0100 Subject: [PATCH 11/15] tools/testing/selftests/mm: test virtual page offset merge behaviour Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-11-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2961; i=ljs@kernel.org; h=from:subject:message-id; bh=01jUabKK64NHmCAfzlhC9tplycbpQo1xdwDkiXe4HZs=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivZZRr3pjj9ycdukFK0HjybKqS6x5PtSNmeSwuIzz ZqxK9o6O0pZGMS4GGTFFFmefxHfHyQSNq/zgr8bzBxWJpAhDFycAjCRle6MDPudjzufvXAlRnZ9 xC++l0s1te6E7NvhMsum3C+U91K+9FtGhiuGt76qxysqLpgml1vBlPflsCireUL6n/vHT72o2z/ PmRcA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 While maintaining virtual page offsets for VMAs has no impact for most merge cases, it does impact MAP_PRIVATE-mapped file-backed mappings which happen to have matching page offset but not matching virtual page offset. Assert this behaviour by attempting to map an unfaulted MAP_PRIVATE-memfd region with a faulted one with compatible file page offsets but incompatible virtual page offsets. Signed-off-by: Lorenzo Stoakes (ARM) --- tools/testing/selftests/mm/merge.c | 57 ++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 57 insertions(+) diff --git a/tools/testing/selftests/mm/merge.c b/tools/testing/selftests/m= m/merge.c index 519e5ac02db7..c093e27d0aea 100644 --- a/tools/testing/selftests/mm/merge.c +++ b/tools/testing/selftests/mm/merge.c @@ -1305,6 +1305,63 @@ TEST_F(merge, merge_vmas_with_mseal) ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 2 * page_size); } =20 +TEST_F(merge, virt_and_page_offset_mismatch_memfd) +{ + struct procmap_fd *procmap =3D &self->procmap; + unsigned int page_size =3D self->page_size; + char *carveout =3D self->carveout; + char *ptr, *ptr2; + int fd; + + /* Create a 10 page memfd descriptor. */ + fd =3D memfd_create("virt_page_offset_test", MFD_CLOEXEC); + ASSERT_NE(fd, -1); + ASSERT_EQ(ftruncate(fd, 10 * page_size), 0); + + /* Map a region using the memfd at page offset 0. */ + ptr =3D mmap(carveout, 5 * page_size, PROT_READ | PROT_WRITE, + MAP_FIXED | MAP_PRIVATE, fd, 0); + ASSERT_NE(ptr, MAP_FAILED); + + /* + * Map another separately, and fault in, at page offset 5: + * + * |-----------| |---------| + * | unfaulted | | faulted | + * |-----------| |---------| + */ + ptr2 =3D mmap(&carveout[10 * page_size], 5 * page_size, + PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE, + fd, 5 * page_size); + ASSERT_NE(ptr2, MAP_FAILED); + ptr2[0] =3D 'x'; + + /* + * Now move it in place: + * + * |----------| + * | | + * v | + * |-----------| |---------| + * | unfaulted | | faulted | + * |-----------| |---------| + * + * Because virtual page offset of the faulted region is now + * &carveout[10 * page_size], despite the two regions being mergeable + * due to file page offset, they are NOT mergeable due to virtual page + * offset. + */ + ptr2 =3D sys_mremap(ptr2, 5 * page_size, 5 * page_size, + MREMAP_MAYMOVE | MREMAP_FIXED, + &carveout[5 * page_size]); + ASSERT_NE(ptr2, MAP_FAILED); + + /* Assert that they did not merge. */ + ASSERT_TRUE(find_vma_procmap(procmap, ptr)); + ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr); + ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 5 * page_size); +} + TEST_F(merge_with_fork, mremap_faulted_to_unfaulted_prev) { struct procmap_fd *procmap =3D &self->procmap; --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E73EC3370E4; Fri, 17 Jul 2026 18:24:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312644; cv=none; b=BiBejW+MHZUdM8o6M1/PE6WojPxRaFS82U91r0kcUfwN4iEUF54AzWk1+Ezz+cVcxt4zPaDJcjoxyRAMOwjZSSCVxDqEBvwgkzSuxGMp0RFoL4vmq6tFjCuU8cYMAvjePONE0KRuKcbqKXNr/nA7kNF7NtDHKzAxLsz9EIMLnDA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312644; c=relaxed/simple; bh=efwTnIHzkKX0RplIWXulhNQpbVQxVROQFPLdI6z77uY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=e5djMWAt2dhFA6UMaR6OYm3puszpK434N9v+PYxFXv6OlKIeEvAhTRAQBNliuo8nIJFzgpS46vs4BgMeU7AtDUU6wOEwjALsRW+BCVjyAL3ubbSmZQfI+qRUt81X4FPUmRDnV+BKkGOL4WeinsCgL5AFFDQDJJURabo+bJOK9ug= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=odkRKlJj; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="odkRKlJj" Received: by smtp.kernel.org (Postfix) with ESMTPSA id B31EB1F00A3D; Fri, 17 Jul 2026 18:23:53 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312642; bh=AMXMZPPGRWd9Xw8i5HYWkiF0Bq2ZcnqCXrjlidS4hik=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=odkRKlJjGCbmUafVPgYl86ptaM4bULyKnomTxwVv6fOeiTnjQV945B301hF7qZsHd ckdjpBgNXQYY6Zq2c2G1UQIo5+TffcppfpHiAErGoz3l8soQ9NVDcxD5s3ugi/8Un7 42nBwEuAkoE6M2UkQyAsJRq0AyNHhzE052DIRJqEB+o+7cMfdkqHL+rLU9bW7lQN/0 RgcLWReYD8LNeeYwCQ7MbjNg8KM21a7NkxUa/6N0s655halOIB9hsQhzJOEMpet7xn y4qEnEEPiSedlrJZc79XLDTfrkBqyUbWQArVD1TIw6ZNQqkZZ2jp187OcABa2Kzf0O kebbKXvGg6tew== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:50 +0100 Subject: [PATCH 12/15] mm/vma: only permit MAP_PRIVATE /dev/zero to be mapped anonymous Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-12-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=7326; i=ljs@kernel.org; h=from:subject:message-id; bh=efwTnIHzkKX0RplIWXulhNQpbVQxVROQFPLdI6z77uY=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivbZxV65PsvQ7eQ/v50PLQRT1V9m1xm82xq40yxZx WqDjOPdjlIWBjEuBlkxRZbnX8T3B4mEzeu84O8GM4eVCWQIAxenAEzkkS4jw7l94f/FtS4vthdm 8bXWfX/nbPKsv1NVtauTY1+9dt9bJcnIcHjnB2cFsQ05Lu7H380JP864Tm2Z7c53p2e9eO9sZ77 oNDsA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 In order to use mmap_prepare() with MAP_PRIVATE mappings of /dev/zero without the success_hook hack we explicitly permitted mmap_prepare handlers to set NULL vm_ops. However this is dangerous and we really only want to allow this for MAP_PRIVATE-mapped /dev/zero. Make it possible to explicitly identify /dev/zero by setting a global DEVZERO_MINOR device minor number then explicitly check for this in mmap code for a MAP_PRIVATE mapping and only set the VMA anonymous if we have positively identified it. Then remove all ability for mmap_prepare or mmap hooks to set a VMA anonymous and update mmap_zero_prepare() to leave it to the core mmap code to mark the VMA anonymous. Note that this disallows nested MAP_PRIVATE-mappings of /dev/zero regions. Doing this would be broken in any case. We therefore do not need to update the mmap_prepare() compatibility layer to reflect these changes, as the mmap hook check suffices to disallow this behaviour. Also update the VMA userland tests to reflect the change. Signed-off-by: Lorenzo Stoakes (ARM) --- drivers/char/mem.c | 8 ++------ include/linux/mm.h | 3 +++ mm/internal.h | 4 ++++ mm/vma.c | 33 +++++++++++++++++++++++++++------ mm/vma_internal.h | 1 + tools/testing/vma/include/dup.h | 28 ++++++++++++++++++++++++++++ 6 files changed, 65 insertions(+), 12 deletions(-) diff --git a/drivers/char/mem.c b/drivers/char/mem.c index 63253d1de5d7..dcfd896b733d 100644 --- a/drivers/char/mem.c +++ b/drivers/char/mem.c @@ -506,11 +506,7 @@ static int mmap_zero_prepare(struct vm_area_desc *desc) if (vma_desc_test(desc, VMA_SHARED_BIT)) return shmem_zero_setup_desc(desc); =20 - /* - * This is a highly unique situation where we mark a MAP_PRIVATE mapping - * of /dev/zero anonymous, despite it not being. - */ - vma_desc_set_anonymous(desc); + /* MAP_PRIVATE semantics are taken care for us by core mm. */ return 0; } =20 @@ -698,7 +694,7 @@ static const struct memdev { #ifdef CONFIG_DEVPORT [4] =3D { "port", &port_fops, 0, 0 }, #endif - [5] =3D { "zero", &zero_fops, FMODE_NOWAIT, 0666 }, + [DEVZERO_MINOR] =3D { "zero", &zero_fops, FMODE_NOWAIT, 0666 }, [7] =3D { "full", &full_fops, 0, 0666 }, [8] =3D { "random", &random_fops, FMODE_NOWAIT, 0666 }, [9] =3D { "urandom", &urandom_fops, FMODE_NOWAIT, 0666 }, diff --git a/include/linux/mm.h b/include/linux/mm.h index b6503b5f0010..7614afe99c96 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -740,6 +740,9 @@ static inline bool fault_flag_allow_retry_first(enum fa= ult_flag flags) { FAULT_FLAG_INTERRUPTIBLE, "INTERRUPTIBLE" }, \ { FAULT_FLAG_VMA_LOCK, "VMA_LOCK" } =20 +/* /dev/zero minor device number. Special due to MAP_PRIVATE semantics. */ +#define DEVZERO_MINOR 5 + /* * vm_fault is filled by the pagefault handler and passed to the vma's * ->fault function. The vma's ->fault is responsible for returning a bitm= ask diff --git a/mm/internal.h b/mm/internal.h index 22f63cb85bd3..41561fdeb56d 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -240,6 +240,10 @@ static inline int mmap_file(struct file *file, struct = vm_area_struct *vma) { int err =3D vfs_mmap(file, vma); =20 + /* Hooks cannot mark themselves anonymous. */ + if (WARN_ON_ONCE(vma_is_anonymous(vma))) + err =3D -EINVAL; + if (likely(!err)) return 0; =20 diff --git a/mm/vma.c b/mm/vma.c index 6dc095d03382..6ccd097b8813 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -2612,6 +2612,27 @@ static int __mmap_new_file_vma(struct mmap_state *ma= p, return 0; } =20 +static bool map_is_dev_zero(const struct mmap_state *map) +{ + const struct file *file =3D map->file; + const struct inode *inode =3D file_inode(file); + + return imajor(inode) =3D=3D MEM_MAJOR && iminor(inode) =3D=3D DEVZERO_MIN= OR; +} + +static bool map_is_private(const struct mmap_state *map) +{ + return !vma_flags_test(&map->vma_flags, VMA_SHARED_BIT); +} + +static bool map_is_anon(const struct mmap_state *map) +{ + if (!map_is_private(map)) + return false; + + return !map->file || map_is_dev_zero(map); +} + /* * __mmap_new_vma() - Allocate a new VMA for the region, as merging was not * possible. @@ -2625,8 +2646,7 @@ static int __mmap_new_file_vma(struct mmap_state *map, static int __mmap_new_vma(struct mmap_state *map, struct vm_area_struct **= vmap, struct mmap_action *action) { - const bool is_anon =3D !map->file && - !vma_flags_test(&map->vma_flags, VMA_SHARED_BIT); + const bool is_anon =3D map_is_anon(map); struct vma_iterator *vmi =3D map->vmi; int error =3D 0; struct vm_area_struct *vma; @@ -2768,6 +2788,10 @@ static int call_mmap_prepare(struct mmap_state *map, if (err) return err; =20 + /* Hooks cannot mark themselves anonymous. */ + if (!desc->vm_ops) + return -EINVAL; + err =3D call_action_prepare(map, desc); if (err) return err; @@ -2790,10 +2814,7 @@ static int call_mmap_prepare(struct mmap_state *map, static void set_vma_user_defined_fields(struct vm_area_struct *vma, struct mmap_state *map) { - if (map->vm_ops) - vma->vm_ops =3D map->vm_ops; - else /* Only /dev/zero should do this. */ - vma_set_anonymous(vma); + vma->vm_ops =3D map->vm_ops; vma->vm_private_data =3D map->vm_private_data; } =20 diff --git a/mm/vma_internal.h b/mm/vma_internal.h index 4d300e7bbaf4..385c0ab13777 100644 --- a/mm/vma_internal.h +++ b/mm/vma_internal.h @@ -23,6 +23,7 @@ #include #include #include +#include #include #include #include diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index 51605ade06fb..93049241b297 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -15,6 +15,15 @@ struct task_struct *get_current(void); #define MMF_HAS_MDWE 28 #define current get_current() =20 +#define MINORBITS 20 +#define MINORMASK ((1U << MINORBITS) - 1) + +#define MAJOR(dev) ((unsigned int) ((dev) >> MINORBITS)) +#define MINOR(dev) ((unsigned int) ((dev) & MINORMASK)) + +#define MEM_MAJOR 1 +#define DEVZERO_MINOR 5 + /* * Define the task command name length as enum, then it can be visible to * BPF programs. @@ -45,6 +54,9 @@ struct address_space { unsigned long flags; atomic_t i_mmap_writable; }; +struct inode { + dev_t i_rdev; +}; struct file_operations { int (*mmap)(struct file *, struct vm_area_struct *); int (*mmap_prepare)(struct vm_area_desc *); @@ -52,6 +64,7 @@ struct file_operations { struct file { struct address_space *f_mapping; const struct file_operations *f_op; + struct inode *f_inode; }; struct anon_vma_chain { struct anon_vma *anon_vma; @@ -1632,3 +1645,18 @@ static inline pgoff_t linear_virt_page_index(const s= truct vm_area_struct *vma, =20 return pgoff; } + +static inline struct inode *file_inode(const struct file *f) +{ + return f->f_inode; +} + +static inline unsigned iminor(const struct inode *inode) +{ + return MINOR(inode->i_rdev); +} + +static inline unsigned imajor(const struct inode *inode) +{ + return MAJOR(inode->i_rdev); +} --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 14EB83370E4; Fri, 17 Jul 2026 18:24:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312653; cv=none; b=V20d4kTS8XC/cBLil3gcgnLYAcwVfve/hkfSEV1UxEnl8+Uvu1ZCPWbDooNvPNFUXASGmven+Gd1TpDX+B2N2ZQ2b7+PnjKpkVVFa1657VFGc94sbd0po/RDg5UxTCp4COz5ZPz93viEsXeWDFJcWmRIs43OjJt3tXXl8wfaESk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312653; c=relaxed/simple; bh=C4IaEx/DcAW96RFPFZ/bWWEFrGWe3/+6S/w8PUQQWRk=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=EB5fj5AwhIAGSy5FuxVePM6TRwO0pwYgAaxcBFed0g4JnIRKx0TLqdPtE1rJ19noq0eFoDf/U1OU8iB7YKGJZmltOPWcLA5jof1C+3/JIOeuMKHNZul1vgWnS//oIyo+tskqtoXsoI2UsILTILbxzBgfva8/GnKjxaMwLR5ai7s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Nok5y9V7; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Nok5y9V7" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1C0311F000E9; Fri, 17 Jul 2026 18:24:02 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312652; bh=KkiX41oRUtbyELsD72Aha5EOEJUotYzdMaBBIICRLMU=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=Nok5y9V7uz97J68uVLNxkH9lWj0x2uAj7uK5VE0cKXSiIYTR/bogQ4ikpvYEEJnXz NrVinhIg3yTbgiWfyStlcb8GWIt6LRFzjNX0iRmVzJSZpyO02oJLyo8hHnMTP/UFKs +91eaKvRooR1qs76/DtIuE+9w1whqb86KouOz754CZmJt/64FB8mFMBrXfxx1qES/w nplHl1qI86bWKKMhcsdLFVSGVMaWJZYrsXllxvcsfJofyLaQNGmDiiSUCAs1ZzqPBI uMHsMSsJUtBE2mv18kPAIUeSscuCfo7hLx+rZb7BFYEY11dvgtciWupXWzDPg7Hd2F DBCsiaKDltn9Q== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:51 +0100 Subject: [PATCH 13/15] mm/vma: make MAP_PRIVATE-mapped /dev/zero mappings truly anonymous Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-13-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4776; i=ljs@kernel.org; h=from:subject:message-id; bh=C4IaEx/DcAW96RFPFZ/bWWEFrGWe3/+6S/w8PUQQWRk=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivZ9VN962Us85H6JBdP+UpEo98o/NdE83AdOdbfWm kSFr7fpKGVhEONikBVTZHn+RXx/kEjYvM4L/m4wc1iZQIYwcHEKwEQOTmb4X/Dc+NG7De5Mf6KD Kk5G/ORgOGrvUKZ0/vNhs/xgzbBZcxj+u66a+ml1oLxcv+X5TKF/JyNjpmR+L18cFrtzb/2lVXZ bGQA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 When mapping /dev/zero with MAP_PRIVATE, one ends up with strange VMAs originating from Linux's distant past. These have vma->vm_file set but NULL vma->vm_ops, meaning they satisfy vma_is_anonymous() but otherwise resemble a file-backed VMA. The introduction of virtual page offsets and their subsequent use as indexes for MAP_PRIVATE-file-backed mappings mean the rmap does the right thing with these but we are left with inconsistencies. The vma_start_pgoff(vma) =3D=3D vma_start_virt_pgoff(vma) invariant is true= for all other anonymous VMAs, but not these. These VMAs are also observable as files in /proc//[s]maps but otherwise behave like anonymous mappings. Therefore let's make these VMAs actually anonymous at mapping time which will activate the anonymous code path for mappings. This means we no longer have to account for this discrepancy anywhere and no longer have to think about these at all. A previous commit gave us map_is_dev_zero() to positively identify these mappings, so we expressly only do so for these alone. The impact of this change should be low as likely very few are relying upon this in any case, and in using them are asking for anonymous memory, so no longer seeing these as file mappings in smaps should have no meaningful impact. Update assert_sane_pgoff() and the comment for vma_start_pgoff() to reflect the change. We make this change in call_mmap_prepare() alone as /dev/zero has been converted to an mmap_prepare hook and we do not permit nested MAP_PRIVATE mapping of /dev/zero. We also remove the now defunct vma_desc_set_anonymous(). Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/mm.h | 10 ++-------- mm/vma.c | 21 +++++++++++++++++---- mm/vma.h | 3 --- 3 files changed, 19 insertions(+), 15 deletions(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index 7614afe99c96..7fabe6c66b4b 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -1554,11 +1554,6 @@ static inline void vma_set_anonymous(struct vm_area_= struct *vma) vma->vm_ops =3D NULL; } =20 -static inline void vma_desc_set_anonymous(struct vm_area_desc *desc) -{ - desc->vm_ops =3D NULL; -} - static inline bool vma_is_anonymous(const struct vm_area_struct *vma) { return !vma->vm_ops; @@ -4352,9 +4347,8 @@ static inline unsigned long vma_pages(const struct vm= _area_struct *vma) * If @vma is a MAP_PRIVATE file-backed mapping, then this returns the * page offset within the file. * - * Edge cases: nommu does not abide by these, MAP_PRIVATE-/dev/zero satisf= ies - * vma_is_anonymous() but has file-backed page offset, and MAP_PRIVATE-pfn= map - * regions have their page offset set to the first PFN in the range. + * Edge cases: nommu does not abide by these and CoW MAP_PRIVATE-pfnmap re= gions + * have their page offset set to the first PFN in the range. * * Returns: The page offset of the start of @vma. */ diff --git a/mm/vma.c b/mm/vma.c index 6ccd097b8813..f8001d5c23aa 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -2620,6 +2620,14 @@ static bool map_is_dev_zero(const struct mmap_state = *map) return imajor(inode) =3D=3D MEM_MAJOR && iminor(inode) =3D=3D DEVZERO_MIN= OR; } =20 +static void map_set_anon(struct mmap_state *map) +{ + map->file =3D NULL; + map->file_doesnt_need_get =3D false; + map->pgoff =3D map->addr >> PAGE_SHIFT; + map->vm_ops =3D NULL; +} + static bool map_is_private(const struct mmap_state *map) { return !vma_flags_test(&map->vma_flags, VMA_SHARED_BIT); @@ -2627,10 +2635,7 @@ static bool map_is_private(const struct mmap_state *= map) =20 static bool map_is_anon(const struct mmap_state *map) { - if (!map_is_private(map)) - return false; - - return !map->file || map_is_dev_zero(map); + return map_is_private(map) && !map->file; } =20 /* @@ -2808,6 +2813,14 @@ static int call_mmap_prepare(struct mmap_state *map, map->vm_ops =3D desc->vm_ops; map->vm_private_data =3D desc->private_data; =20 + /* + * MAP_PRIVATE-/dev/zero mappings are an ancient way of getting + * anonymous mappings. Rather than allowing these mappings to be odd + * outliers, simply make them truly anonymous. + */ + if (map_is_private(map) && map_is_dev_zero(map)) + map_set_anon(map); + return 0; } =20 diff --git a/mm/vma.h b/mm/vma.h index c1d65060071f..8caee3b4ba26 100644 --- a/mm/vma.h +++ b/mm/vma.h @@ -267,9 +267,6 @@ static inline void assert_sane_pgoff(struct vm_area_str= uct *vma, pgoff_t pgoff) */ if (!vma_is_anonymous(vma)) return; - /* MAP_PRIVATE-/dev/zero is anon, non-NULL vm_file, but has file pgoff. */ - if (vma->vm_file) - return; /* If faulted in, could have been remapped. */ if (vma->anon_vma) return; --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 889C02F8EB4; Fri, 17 Jul 2026 18:24:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312662; cv=none; b=KdDRIiUctoQkmWSoHJZLI/4EmJv855FEadJEWwh+IWO8R++C8689kS7/jpSENKYtqKuX7B13sOqtaT2Q5dM+HatXcoE9FtBDx9D6IJmznga1jYi++5GPw9OUa6Q3p/Y7UOhmtnTl8Kp9TlEJYMh19LbNGSR0onJq86+07HAjTxc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312662; c=relaxed/simple; bh=2PKWXL01QnSXKidEY/TWtRsvAkg/KT5p1XRORzualh4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=MEzVL8HQmCNKHA5uXnw/EtgJenn4HhfZCCFK+PtmEwQ2orV4ZZR6n2i1aiHALO25iZMj0nOnoHpoRr/ZmKXtdRFfPg+0639Lall/gbNTSLJJMYsq0osGO+bu+NZRs50rf0lFq0tgQiGx/cNX6SG+0NVAJwggg1EY/IKA3VwngyM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=FJMX7hxD; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="FJMX7hxD" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 76D501F00A3A; Fri, 17 Jul 2026 18:24:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312661; bh=HlBBcS3yoy86Hd8Tk6gkvICDLvh7sVRSb6A5C2ruF/k=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=FJMX7hxDC041A9zZr9hhXc4BMh6S0gtjj56+ThInE05IAekRkJBz1uxfru/bZbJL+ UT5kerNVwDA7zM63We4ymno/jVhzefiBRhPCgsvS6otAaNI/xEYowngkTmylTPVsY3 ybLLX8thyqHcCh3+kJ+7xCKR2dSGJKQmSoVmvk/iCVbPQ7L619ffw+2yn4N4ffWtAD lq+lBaIaqDSOHIqdmYbEDWt3qT5aRZQ+fqezWb0azl2zNcEwpJYm6v1NQuP41RNWz/ b0ARS/VR+cVKjktU62J5KKwpDcyDYcDXxH78WJEHctdiHDCLJuTR4ArPI82rU6sN5S ibsabnGCRIoJA== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:52 +0100 Subject: [PATCH 14/15] tools/testing/vma: add test to assert MAP_PRIVATE-/dev/zero is anon Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-14-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2619; i=ljs@kernel.org; h=from:subject:message-id; bh=2PKWXL01QnSXKidEY/TWtRsvAkg/KT5p1XRORzualh4=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivaF7IjUzTZd3XfC3aN2vYhS7cr2B8qdptqW3r23v fnf+vV1lLIwiHExyIopsjz/Ir4/SCRsXucFfzeYOaxMIEMYuDgFYCLG3Az/LPXZ3P5eXiq39NLa Cot3zRKZGkZTX14vNb+iUfb4D4dWESPD2gB2S1tZzeD9S6ykq0/mnY6dEjT1hNRX3Z2Sotv5lJ5 yAwA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Now we've made MAP_PRIVATE-mapped /dev/zero mappings truly anonymous, add a VMA userland test to assert that this is the case and everything is as we would expect for an anonymous mapping. Signed-off-by: Lorenzo Stoakes (ARM) --- tools/testing/vma/include/dup.h | 1 + tools/testing/vma/tests/mmap.c | 49 +++++++++++++++++++++++++++++++++++++= ++++ 2 files changed, 50 insertions(+) diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index 93049241b297..b4f0f7c24fcb 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -20,6 +20,7 @@ struct task_struct *get_current(void); =20 #define MAJOR(dev) ((unsigned int) ((dev) >> MINORBITS)) #define MINOR(dev) ((unsigned int) ((dev) & MINORMASK)) +#define MKDEV(ma, mi) (((ma) << MINORBITS) | (mi)) =20 #define MEM_MAJOR 1 #define DEVZERO_MINOR 5 diff --git a/tools/testing/vma/tests/mmap.c b/tools/testing/vma/tests/mmap.c index c85bc000d1cb..0c8a3446b905 100644 --- a/tools/testing/vma/tests/mmap.c +++ b/tools/testing/vma/tests/mmap.c @@ -45,7 +45,56 @@ static bool test_mmap_region_basic(void) return true; } =20 +static int dummy_mmap_prepare(struct vm_area_desc *desc) +{ + return 0; +} + +static bool test_pure_anon_dev_zero(void) +{ + const vma_flags_t vma_flags =3D mk_vma_flags(VMA_READ_BIT, VMA_WRITE_BIT, + VMA_MAYREAD_BIT, VMA_MAYWRITE_BIT); + const struct file_operations f_op =3D { + .mmap_prepare =3D dummy_mmap_prepare, + }; + struct inode inode =3D { + .i_rdev =3D MKDEV(MEM_MAJOR, DEVZERO_MINOR), + }; + struct file file =3D { + .f_inode =3D &inode, + .f_op =3D &f_op, + }; + struct mm_struct mm =3D {}; + struct vm_area_struct *vma; + unsigned long addr; + VMA_ITERATOR(vmi, &mm, 0); + + current->mm =3D &mm; + + /* + * Map a MAP_PRIVATE-/dev/zero mapping at address 0x300000 with a page + * offset of 0x10, which we expect to be reset to the virtual page + * offset. + */ + addr =3D __mmap_region(&file, 0x300000, 0x3000, vma_flags, 0x10, NULL); + ASSERT_EQ(addr, 0x300000); + + /* Assert that it truly is an anonymous mapping. */ + vma =3D vma_lookup(&mm, addr); + ASSERT_NE(vma, NULL); + ASSERT_TRUE(vma_is_anonymous(vma)); + ASSERT_EQ(vma->vm_file, NULL); + ASSERT_EQ(vma->vm_private_data, NULL); + /* Expect virtual page offsets. */ + ASSERT_EQ(vma->vm_pgoff, 0x300); + ASSERT_EQ(vma_start_virt_pgoff(vma), 0x300); + + cleanup_mm(&mm, &vmi); + return true; +} + static void run_mmap_tests(int *num_tests, int *num_fail) { TEST(mmap_region_basic); + TEST(pure_anon_dev_zero); } --=20 2.55.0 From nobody Sat Jul 25 04:59:09 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1EBD12D77F7; Fri, 17 Jul 2026 18:24:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312672; cv=none; b=daLi6ynj32giW6vZuIq8s8OZIx3xJzhbvs+z7UDgkmMaAlhkpTsUPGITKRrPYP9V8Icd77eg1/1KPZoFOm8y+sLTsbWl9/P1UYIfX83tSPw3f/Csk5ZQoI+H9NucEf7E2BNYM87qq2EcW7P1xbQ7ZegsrAQnkSfblAWLqHSbKrg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784312672; c=relaxed/simple; bh=yk0gma+yTZdel+ljE629bRHp848gvZkgeJregd069ZY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=FyZc/AfvtJ4LLnpzH6saJ1ZwWhSe/uQ9o87hdPdaCa0gYXE2b95iqNl7WxBQrQ/VVq91n91zSZ24XzTDFNrjOQtmBYLapAwgbSSErN+dZ86UjcwH4NHFTFH4HyFTqNQAdocKqj5irO7QFpFOu9AmWhEilU0owSps5NcmzUPQ1jQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=omXE4CsQ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="omXE4CsQ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id CFCB71F000E9; Fri, 17 Jul 2026 18:24:21 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784312670; bh=Q+hgooHeYYaYv/oiqa66MnVBmMzHrm3fz5OuZIxN9Pg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=omXE4CsQ2Z64k5LeffWwtnN50HMH1/GHkEgGbyKyF+mXWBnl8u4Ei1dERW0FAekOB zQLtjUZRKIg4sQD1ChlPZUv/HdqV3J9HP4cBFR9GIUVhiUfG50Ij9EbHeruFLhsyOT 2LeI0NkUowD1qy9ds0ErFPOqZFEcCPHtNLQ2A12lkujjycPHc+BhEV3ecB3iZIb4rI TMCdXJBdgeyKiCbtMA7pIdPANJrInzPE8TG0Bh03d4R6HLx3I1VDk4BILnb+jg89A5 QacXczPre4f19oIX8EhVmBOnv4OL0FXLjd+A+LAIRiOYDXQdLLtEEkAVfAB9/1zmGL 9JeWLb0BtKVZA== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 19:19:53 +0100 Subject: [PATCH 15/15] tools/testing/selftests/mm: add MAP_PRIVATE-/dev/zero merge tests Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-b4-scalable-cow-virt-pgoff-v1-15-cf24910ef094@kernel.org> References: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> In-Reply-To: <20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4278; i=ljs@kernel.org; h=from:subject:message-id; bh=yk0gma+yTZdel+ljE629bRHp848gvZkgeJregd069ZY=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKiivYVNv2si77OGnzswM/3ItzKPx4dfOBkKvBj1cI75 ucurbO83lHKwiDGxSArpsjy/Iv4/iCRsHmdF/zdYOawMoEMYeDiFICJ8FsxMjzZUBtmE/Eifvml ncobjFkv3K4I6Ww9zvHr6NL/sqf/1j1kZGjbHRG2r664eP6GNL57BXHqtjtMuvLiW19+mdY/80C cJhMA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Assert that MAP_PRIVATE-mapped /dev/zero mappings behave like they are anonymous. We test both unfaulted and faulted/unfaulted merges - each with the regions having page offset of 0, which would not merge if the mappings were treated as if they were file-backed. With the recent change that makes them behave as pure anonymous mappings, the merges should succeed as their page offsets are equal to their virtual page offsets. Signed-off-by: Lorenzo Stoakes (ARM) --- tools/testing/selftests/mm/merge.c | 104 +++++++++++++++++++++++++++++++++= ++++ 1 file changed, 104 insertions(+) diff --git a/tools/testing/selftests/mm/merge.c b/tools/testing/selftests/m= m/merge.c index c093e27d0aea..edfbd39ae58e 100644 --- a/tools/testing/selftests/mm/merge.c +++ b/tools/testing/selftests/mm/merge.c @@ -1362,6 +1362,110 @@ TEST_F(merge, virt_and_page_offset_mismatch_memfd) ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 5 * page_size); } =20 +TEST_F(merge, merge_map_private_dev_zero_unfaulted) +{ + struct procmap_fd *procmap =3D &self->procmap; + unsigned int page_size =3D self->page_size; + char *carveout =3D self->carveout; + char *ptr, *ptr2; + int fd_zero; + + if (access("/dev/zero", F_OK)) + SKIP(return, "No /dev/zero."); + fd_zero =3D open("/dev/zero", O_RDWR); + ASSERT_NE(fd_zero, -1); + + /* + * Map two MAP_PRIVATE-/dev/zero VMAs next to one another with offset 0 + * each. + * + * With these being made truly anonymous upon mapping, they will + * merge. If they were file-backed VMAs the page offsets would prevent + * merge: + * + * |-----||------| |-------------| + * | ptr || ptr2 | -> | ptr | + * |-----||------| |-------------| + */ + ptr =3D mmap(carveout, 5 * page_size, PROT_READ | PROT_WRITE, + MAP_FIXED | MAP_PRIVATE, fd_zero, 0); + if (ptr =3D=3D MAP_FAILED) { + close(fd_zero); + ASSERT_TRUE(false); + } + ptr2 =3D mmap(&carveout[5 * page_size], 5 * page_size, + PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE, fd_zero, 0); + if (ptr2 =3D=3D MAP_FAILED) { + close(fd_zero); + ASSERT_TRUE(false); + } + close(fd_zero); + + /* Assert that they merged. */ + ASSERT_TRUE(find_vma_procmap(procmap, ptr)); + ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr); + ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 10 * page_size); +} + +TEST_F(merge, merge_map_private_dev_zero_faulted_unfaulted) +{ + struct procmap_fd *procmap =3D &self->procmap; + unsigned int page_size =3D self->page_size; + char *carveout =3D self->carveout; + char *ptr, *ptr2; + int fd_zero; + + if (access("/dev/zero", F_OK)) + SKIP(return, "No /dev/zero."); + fd_zero =3D open("/dev/zero", O_RDWR); + ASSERT_NE(fd_zero, -1); + + /* + * Map a MAP_PRIVATE mapping of /dev/zero with page offset 0, then fault + * it in: + * + * |-------------------------------| + * | faulted | + * |-------------------------------| + */ + ptr =3D mmap(carveout, 15 * page_size, PROT_READ | PROT_WRITE, + MAP_FIXED | MAP_PRIVATE, fd_zero, 0); + if (ptr =3D=3D MAP_FAILED) { + close(fd_zero); + ASSERT_TRUE(false); + } + memset(ptr, 'x', 15 * page_size); + + /* + * Unmap the middle: + * + * |---------| |---------| + * | faulted | | faulted | + * |---------| |---------| + */ + ASSERT_EQ(munmap(&ptr[5 * page_size], 5 * page_size), 0); + + /* + * Map in a new unfaulted mapping in the middle with page offset 0 - + * this should merge and would not if it were treated as a file rather + * than pure anon: + * + * |---------|-----------|---------| + * | faulted | unfaulted | faulted | + * |---------|-----------|---------| + */ + ptr2 =3D mmap(&carveout[5 * page_size], 5 * page_size, + PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE, + fd_zero, 0); + close(fd_zero); + ASSERT_NE(ptr2, MAP_FAILED); + + /* Assert that they merged. */ + ASSERT_TRUE(find_vma_procmap(procmap, ptr)); + ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr); + ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 15 * page_size); +} + TEST_F(merge_with_fork, mremap_faulted_to_unfaulted_prev) { struct procmap_fd *procmap =3D &self->procmap; --=20 2.55.0