From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4AF74388E68; Thu, 13 Aug 2026 17:33:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642405; cv=none; b=GlV42g7SbF17Kw18jfE5kFXK0+NeNPtX/XUxP68W9rOxMvFCtsgKtdhGp4SpBdO0KPcFRZfduup6hiEgAR5Pjxzpzrm1AFsiVJpUb8B+Z4mBCZNInC+7WJF5X0cDY/2+Tr9NUWZK7b2MD54lHKVswo6eePEXoG7ZDMM4prXNH3k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642405; c=relaxed/simple; bh=k4nsxhVksl7Id9wH8PF9j9dwGi3ny6iiiLSmQsKoHuc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=si9NOZZ9gO9GVCCnxM7DkF+//If50hXTyYpeyHrS6zicyTd2GTfuN94LcHqmyRbIyg1f1Rtbw7ZXqUO+POyf9d50Po5RSJeVNQVXzp/D2jgT7TpQ2PDEIEivWpgJqMUSsSVe7Miw1W6SZIl9l5jprPdfmP+9Pg4YTvSx4Fvh/Ag= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=edE8IdYl; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="edE8IdYl" Received: by smtp.kernel.org (Postfix) with ESMTPSA id E39091F00A3A; Thu, 13 Aug 2026 17:33:03 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642403; bh=DbE5DHkTNtU6STxh6PeRdePiOC2bhJBQkXlvls/SNVM=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=edE8IdYlIINlenX9lgeJHUProrbMovmH4UE9do8ZRCxNT+52JZtnG6J2Ec8WlhYzD mnMRt30TRAMo9hbN95yoUl3r5ptkkjszHnUjIhsTqtojYCS8LU7H67SdBhjPmAA0sg HkIWDikanTZsQ20dEo918r70Jk/xF4Lx4M42CXZ19ty99RoYBeFc+uGbUHuBd0kLzG IS6dykK3gesLDRnSdFEtApvN9u+GwqGLjS3xDU/AdlqajrIYGCKAykmzpATw+EJfFk 2gZXLhLCdnt8tcpxccFRnqIHR4EO/wc8YrDZnnGVb+yayLMza76oRlmPvCMt7ZpWUH fthlzsDNOCBAQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:18 +0100 Subject: [PATCH v5 01/16] mm/vma: introduce VMA anon page offset field and add helpers Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-1-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=8006; i=ljs@kernel.org; h=from:subject:message-id; bh=k4nsxhVksl7Id9wH8PF9j9dwGi3ny6iiiLSmQsKoHuc=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6+e0XPxYsziC6zu6zNPXOK8/03SJX4jx3bW5mDtO wIrmm/FdZSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAipfsYGT4rrbO/rfyIRUP2 6b0ldx7fsr+9/fzrCsvZFtvCVnflHqpm+GfXf/Blx12VfXeLdUzf5fv+nOh323A7/4Xfv27aWeU +NuMAAA== X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Establish fields in vm_area_struct to store the anonymous page offset of VMAs. Initially, the anonymous page offset of a VMA is vma->vm_start >> PAGE_SHIFT. When a VMA is remapped to new_address its anonymous page offset is either updated to new_address >> PAGE_SHIFT if unfaulted or, if faulted, remains equal to the anonymous page offset it had when first faulted. Currently, anonymous folios belonging to CoW'd MAP_PRIVATE-mapped file-backed VMAs are tracked by their file offsets. By adding anonymous offset as a property of VMAs, we can now track them by their anonymous page offset instead. By tracking this, we provide the means by which to eliminate this inconsistency, and more importantly lay the foundations for future work for the scalable CoW anonymous rmap rework. This patch simply adds the fields and some simple helpers. Subsequent patches will update mm code to make use of these fields correctly. The fields chosen are packed in the VMA such that, for 64-bit kernel builds, no additional space is taken up. The first field is present on cacheline 0 containing key VMA fields, and the second on cacheline 3, which contains file-backed reverse mapping fields. Given the relative time spent accessing reverse mapping fields as well as updating them, there shouldn't be any performance impact here from false sharing. Update the VMA userland tests to account for this change. No callsites are updated yet, so no functional change intended. Acked-by: David Hildenbrand (Arm) Reviewed-by: Gregory Price (Meta) Reviewed-by: Xu Xin Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/mm.h | 59 +++++++++++++++++++++++++++++++++++++= ++++ include/linux/mm_types.h | 12 +++++++++ mm/vma.h | 14 ++++++++++ mm/vma_init.c | 1 + tools/testing/vma/include/dup.h | 26 ++++++++++++++++++ 5 files changed, 112 insertions(+) diff --git a/include/linux/mm.h b/include/linux/mm.h index 87feaa5a2b78..df78847f5f07 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -4393,6 +4393,65 @@ static inline pgoff_t vma_last_pgoff(const struct vm= _area_struct *vma) return vma_end_pgoff(vma) - 1; } =20 +/** + * vma_start_anon_pgoff() - Get the anonymous page offset of the start of = @vma + * @vma: The VMA whose anonymous page offset is required. + * + * If unfaulted, then this is vma->vm_start >> PAGE_SHIFT, if faulted then= the + * anonymous page offset at the time of first fault. + * + * If the VMA is anonymous, this returns the same value as vma_start_pgoff= (). + * + * This value is used for tracking MAP_PRIVATE file-backed mappings by the= ir + * anonymous page offset. + * + * Returns: The anonymous page offset of the start of @vma. + */ +static inline pgoff_t vma_start_anon_pgoff(const struct vm_area_struct *vm= a) +{ + pgoff_t pgoff =3D 0; + +#ifdef CONFIG_64BIT + pgoff +=3D vma->__vm_anon_pgoff_hi; + pgoff <<=3D 32; +#endif + pgoff +=3D vma->__vm_anon_pgoff_lo; + return pgoff; +} + +/** + * vma_end_anon_pgoff() - Get the anonymous page offset of the exclusive e= nd of + * @vma. + * @vma: The VMA whose end anonymous page offset is required. + * + * This returns the anonymous exclusive end page offset of @vma, which is = useful + * for expressing page offset ranges. + * + * See the description of vma_start_anon_pgoff() for a description of VMA + * anonymous page offsets. + * + * Returns: The exclusive end anonymous page offset of @vma. + */ +static inline pgoff_t vma_end_anon_pgoff(const struct vm_area_struct *vma) +{ + return vma_start_anon_pgoff(vma) + vma_pages(vma); +} + +/** + * vma_last_anon_pgoff() - Get the anonymous page offset of the last page = in + * @vma. + * @vma: The VMA whose last anonymous page offset is required. + * + * See the description of vma_start_anon_pgoff() for a description of VMA + * anonymous page offsets. + * + * Returns: The last anonymous page offset of @vma. + */ +static inline pgoff_t vma_last_anon_pgoff(const struct vm_area_struct *vma) +{ + return vma_end_anon_pgoff(vma) - 1; +} + static inline unsigned long vma_desc_size(const struct vm_area_desc *desc) { return desc->end - desc->start; diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 939b5ea8c9e0..ebf0d912be7d 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -967,6 +967,11 @@ struct vm_area_struct { */ unsigned int vm_lock_seq; #endif + /* + * Low 32-bits of anonymous page offset. + * See vma_start_anon_pgoff() comment for details. + */ + unsigned int __vm_anon_pgoff_lo; /* * A file's MAP_PRIVATE vma can be in both i_mmap tree and anon_vma * list, after a COW of one of the file pages. A MAP_SHARED vma @@ -1041,6 +1046,13 @@ struct vm_area_struct { #ifdef CONFIG_DEBUG_LOCK_ALLOC struct lockdep_map vmlock_dep_map; #endif +#endif +#ifdef CONFIG_64BIT + /* + * High 32-bits of anonymous page offset. + * See vma_start_anon_pgoff() comment for details. + */ + unsigned int __vm_anon_pgoff_hi; #endif /* * For areas with an address space and backing store, diff --git a/mm/vma.h b/mm/vma.h index 0bc7d521e976..54ed7c744e3b 100644 --- a/mm/vma.h +++ b/mm/vma.h @@ -283,6 +283,20 @@ static inline void vma_set_pgoff(struct vm_area_struct= *vma, pgoff_t pgoff) vma->vm_pgoff =3D pgoff; } =20 +static inline void __vma_set_anon_pgoff(struct vm_area_struct *vma, pgoff_= t pgoff) +{ +#ifdef CONFIG_64BIT + vma->__vm_anon_pgoff_hi =3D pgoff >> 32; +#endif + vma->__vm_anon_pgoff_lo =3D pgoff & GENMASK(31, 0); +} + +static inline void vma_set_anon_pgoff(struct vm_area_struct *vma, pgoff_t = pgoff) +{ + vma_assert_can_modify(vma); + __vma_set_anon_pgoff(vma, pgoff); +} + static inline void vma_add_pgoff(struct vm_area_struct *vma, pgoff_t delta) { vma_assert_can_modify(vma); diff --git a/mm/vma_init.c b/mm/vma_init.c index 715feee283f0..baa7e82f47e3 100644 --- a/mm/vma_init.c +++ b/mm/vma_init.c @@ -51,6 +51,7 @@ static void vm_area_init_from(const struct vm_area_struct= *src, dest->vm_end =3D src->vm_end; dest->anon_vma =3D src->anon_vma; dest->vm_pgoff =3D vma_start_pgoff(src); + __vma_set_anon_pgoff(dest, vma_start_anon_pgoff(src)); dest->vm_file =3D src->vm_file; dest->vm_private_data =3D src->vm_private_data; vm_flags_init(dest, src->vm_flags); diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index c52b23773cd2..26d9210b1e8b 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -577,6 +577,7 @@ struct vm_area_struct { */ unsigned int vm_lock_seq; #endif + unsigned int __vm_anon_pgoff_lo; =20 /* * A file's MAP_PRIVATE vma can be in both i_mmap tree and anon_vma @@ -612,6 +613,9 @@ struct vm_area_struct { #ifdef CONFIG_PER_VMA_LOCK /* Unstable RCU readers are allowed to read this. */ refcount_t vm_refcnt; +#endif +#ifdef CONFIG_64BIT + unsigned int __vm_anon_pgoff_hi; #endif /* * For areas with an address space and backing store, @@ -1320,6 +1324,28 @@ static inline pgoff_t vma_end_pgoff(const struct vm_= area_struct *vma) return vma_start_pgoff(vma) + vma_pages(vma); } =20 +static inline pgoff_t vma_start_anon_pgoff(const struct vm_area_struct *vm= a) +{ + pgoff_t pgoff =3D 0; + +#ifdef CONFIG_64BIT + pgoff +=3D vma->__vm_anon_pgoff_hi; + pgoff <<=3D 32; +#endif + pgoff +=3D vma->__vm_anon_pgoff_lo; + return pgoff; +} + +static inline pgoff_t vma_end_anon_pgoff(const struct vm_area_struct *vma) +{ + return vma_start_anon_pgoff(vma) + vma_pages(vma); +} + +static inline pgoff_t vma_last_anon_pgoff(const struct vm_area_struct *vma) +{ + return vma_end_anon_pgoff(vma) - 1; +} + static inline int vfs_mmap_prepare(struct file *file, struct vm_area_desc = *desc) { return file->f_op->mmap_prepare(desc); --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6EF6338BF8D; Thu, 13 Aug 2026 17:33:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642426; cv=none; b=o4csWi0XxWuTGXkU2VsO3iMgjDFAFprBuU1mi1vXbuJisjh8OmangaKHQs8lKG/jbQK3Cpgh+aofVf7q86iHyjfS/VP2tcPeAHh9PaPFCSLuEp9OMKUny7DAs5SutUSMTCSGz82zx6V3VZUFW+c/n7ebnxM/+cBVcLuX9mhC16o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642426; c=relaxed/simple; bh=MS3khbE1o4A4JymXRIsA9QW+UBIPnd1IKSMFhya8LNs=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=VmJwZGjwyUHH3mKqpjb1bR3DJIthfFBzPrMdd9/bbI8hHqte1a9OjW1zwI+COSuhjPNZDuBJQ4e+R75+YSXIz9l/gtsCZVH+noAjgKkjoQqFKcQPNCBq9eytkSClffoXgYRT0LagtUZWqatb+OJjB/OdyZsIVh0tvS7msRW79pU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=D4ZqlEcK; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="D4ZqlEcK" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3A5AF1F000E9; Thu, 13 Aug 2026 17:33:24 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642424; bh=2vLcy4JbU53/8GMqn+VKiBNeIud51sx6m46KO4U/098=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=D4ZqlEcKlZJQh8znwKyelQly/cq3mREQ6jdSS3Cl8Fv6PBQcprPs/j0ryQfCx5+by E5wYgl1qk7gd04YIbK8EVujLGwUYS5TcXjDghHzb05Qs3Vt6wU5M+k/kehxXzi8DbM XiyH1MhY/v8Lhnq3FI63hqbwjetpuEzUTk1h8Ql/L3yf8K9iHobCpImytnM/UZtihy qzLtXey/Szz0un7vTSceDCupfV//WorLS3Ar9adQ5bfCQcEXQtXFTkuwRXNwxYtEb4 3v2bSv+3iKv29T7yPa+umDA6sNt6wJqY015X0gWpSConcz1dq/unRiTP98fkPNrVdj ic9ZjekM7DA5Q== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:19 +0100 Subject: [PATCH v5 02/16] mm: provide vma_[flags_]is_cow_mapping() and remove is_cow_mapping() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-2-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=18667; i=ljs@kernel.org; h=from:subject:message-id; bh=MS3khbE1o4A4JymXRIsA9QW+UBIPnd1IKSMFhya8LNs=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/69O0L3iP6fBaXbBj4lTj3vF5Btffemi5l+ydz3P8 seGGs/ndpSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAi28QYGU7ofw7nCP27aNpO EdZsvbR/l408+c2kjJfUO8r8C6qN3M3IcKa5X9UkkKtR7vfisqPv9z/cYPYn6Oqxezob7fSvLjt rxgIA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 All remaining callers of is_cow_mapping() are invoking it in the form of is_cow_mapping(vma->vm_flags) or an indirected version of this. Therefore, provide a helper - vma_is_cow_mapping() to directly test the VMA. Additionally provide a new helper vma_flags_is_cow_mapping() which performs the check using the new vma_flags_t type, and share this logic between vma_is_cow_mapping() and vma_desc_is_cow_mapping(). With these changes, no callers of is_cow_mapping() remain, so remove it. Also update the userland VMA tests to reflect the change. No functional change intended. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- arch/s390/mm/gmap_helpers.c | 2 +- drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c | 4 +- drivers/gpu/drm/drm_gem_shmem_helper.c | 2 +- drivers/gpu/drm/panthor/panthor_gem.c | 2 +- drivers/gpu/drm/ttm/ttm_bo_vm.c | 2 +- drivers/gpu/drm/xe/xe_device.c | 2 +- fs/proc/task_mmu.c | 2 +- include/linux/mm.h | 71 +++++++++++++++++++++++++++++= +--- kernel/events/uprobes.c | 2 +- mm/gup.c | 2 +- mm/huge_memory.c | 8 ++-- mm/hugetlb.c | 2 +- mm/internal.h | 2 +- mm/memory.c | 25 ++++++------ mm/mempolicy.c | 2 +- tools/testing/vma/include/dup.h | 11 +++++ 16 files changed, 105 insertions(+), 36 deletions(-) diff --git a/arch/s390/mm/gmap_helpers.c b/arch/s390/mm/gmap_helpers.c index 4bf7c9012feb..cd5fded159c0 100644 --- a/arch/s390/mm/gmap_helpers.c +++ b/arch/s390/mm/gmap_helpers.c @@ -200,7 +200,7 @@ static int find_zeropage_pte_entry(pte_t *pte, unsigned= long addr, * currently only works in COW mappings, which is also where * mm_forbids_zeropage() is checked. */ - if (!is_cow_mapping(walk->vma->vm_flags)) + if (!vma_is_cow_mapping(walk->vma)) return -EFAULT; =20 *found_addr =3D addr; diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c b/drivers/gpu/drm/amd/= amdgpu/amdgpu_gem.c index 6a0699746fbc..0c7309080a7a 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c @@ -377,9 +377,9 @@ static int amdgpu_gem_object_mmap(struct drm_gem_object= *obj, struct vm_area_str /* Workaround for Thunk bug creating PROT_NONE,MAP_PRIVATE mappings * for debugger access to invisible VRAM. Should have used MAP_SHARED * instead. Clearing VM_MAYWRITE prevents the mapping from ever - * becoming writable and makes is_cow_mapping(vm_flags) false. + * becoming writable and makes vma_is_cow_mapping(vma) false. */ - if (is_cow_mapping(vma->vm_flags) && + if (vma_is_cow_mapping(vma) && !(vma->vm_flags & VM_ACCESS_FLAGS)) vm_flags_clear(vma, VM_MAYWRITE); =20 diff --git a/drivers/gpu/drm/drm_gem_shmem_helper.c b/drivers/gpu/drm/drm_g= em_shmem_helper.c index 06d019d51d3e..177d0e0b9334 100644 --- a/drivers/gpu/drm/drm_gem_shmem_helper.c +++ b/drivers/gpu/drm/drm_gem_shmem_helper.c @@ -753,7 +753,7 @@ int drm_gem_shmem_mmap(struct drm_gem_shmem_object *shm= em, struct vm_area_struct return ret; } =20 - if (is_cow_mapping(vma->vm_flags)) + if (vma_is_cow_mapping(vma)) return -EINVAL; =20 dma_resv_lock(shmem->base.resv, NULL); diff --git a/drivers/gpu/drm/panthor/panthor_gem.c b/drivers/gpu/drm/pantho= r/panthor_gem.c index 770556353968..d2eec46f7abe 100644 --- a/drivers/gpu/drm/panthor/panthor_gem.c +++ b/drivers/gpu/drm/panthor/panthor_gem.c @@ -761,7 +761,7 @@ static int panthor_gem_mmap(struct drm_gem_object *obj,= struct vm_area_struct *v return ret; } =20 - if (is_cow_mapping(vma->vm_flags)) + if (vma_is_cow_mapping(vma)) return -EINVAL; =20 if (!refcount_inc_not_zero(&bo->cmap.mmap_count)) { diff --git a/drivers/gpu/drm/ttm/ttm_bo_vm.c b/drivers/gpu/drm/ttm/ttm_bo_v= m.c index 88babf435ac2..872bf444b1f0 100644 --- a/drivers/gpu/drm/ttm/ttm_bo_vm.c +++ b/drivers/gpu/drm/ttm/ttm_bo_vm.c @@ -489,7 +489,7 @@ static const struct vm_operations_struct ttm_bo_vm_ops = =3D { int ttm_bo_mmap_obj(struct vm_area_struct *vma, struct ttm_buffer_object *= bo) { /* Enforce no COW since would have really strange behavior with it. */ - if (is_cow_mapping(vma->vm_flags)) + if (vma_is_cow_mapping(vma)) return -EINVAL; =20 drm_gem_object_get(&bo->base); diff --git a/drivers/gpu/drm/xe/xe_device.c b/drivers/gpu/drm/xe/xe_device.c index 838797cc65d7..ab16937ebbe0 100644 --- a/drivers/gpu/drm/xe/xe_device.c +++ b/drivers/gpu/drm/xe/xe_device.c @@ -330,7 +330,7 @@ static int xe_pci_barrier_mmap(struct file *filp, if (vma->vm_end - vma->vm_start > SZ_4K) return -EINVAL; =20 - if (is_cow_mapping(vma->vm_flags)) + if (vma_is_cow_mapping(vma)) return -EINVAL; =20 if (vma->vm_flags & (VM_READ | VM_EXEC)) diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c index 817e3e0f9194..5c54aebe2118 100644 --- a/fs/proc/task_mmu.c +++ b/fs/proc/task_mmu.c @@ -1693,7 +1693,7 @@ static inline bool pte_is_pinned(struct vm_area_struc= t *vma, unsigned long addr, =20 if (!pte_write(pte)) return false; - if (!is_cow_mapping(vma->vm_flags)) + if (!vma_is_cow_mapping(vma)) return false; if (likely(!mm_flags_test(MMF_HAS_PINNED, vma->vm_mm))) return false; diff --git a/include/linux/mm.h b/include/linux/mm.h index df78847f5f07..199d0a64b14a 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -2271,17 +2271,76 @@ void unpin_user_pages(struct page **pages, unsigned= long npages); void unpin_user_folio(struct folio *folio, unsigned long npages); void unpin_folios(struct folio **folios, unsigned long nfolios); =20 -static inline bool is_cow_mapping(vm_flags_t flags) +/** + * vma_flags_is_cow_mapping() - Do these VMA flags imply a CoW mapping? + * @flags: The VMA flags to check. + * + * Mappings which could be CoW'd (subject to Copy-On-Write faults) are + * described as CoW mappings. + * + * All mappings backed by anonymous folios (all anonymous mappings and most + * MAP_PRIVATE-file backed ranges) are CoW mappings. + * + * All other mappings (including all MAP_SHARED mappings) are non-CoW. + * + * The criteria are !VMA_SHARED_BIT, VMA_MAYWRITE_BIT. + * + * VMA_MAYWRITE_BIT is checked instead of VMA_WRITE_BIT to account for both + * future mprotect() calls which can render a read-only mapping writable, = and + * GUP with FOLL_FORCE (e.g. ptrace) which can CoW a read-only mapping. + * + * - No anonymous mapping can ever clear VMA_MAYWRITE_BIT. + * + * - Writes to anonymous mappings do not immediately result in CoW faults = but + * may do so after the process is forked or if a read is followed by a + * write. + * + * - Writes to MAP_PRIVATE file-backed mappings result in CoW faults and m= ay + * do so again after fork. + * + * - MAP_SHARED mappings of a file opened read-only are transformed into + * VMA_MAYSHARE_BIT, !VMA_SHARED_BIT, !VMA_MAYWRITE_BIT mappings, so rem= ain + * non-CoW. + * + * - Drivers may clear VMA_MAYWRITE_BIT but do so at mmap() time and cannot + * mark themselves anonymous. Having cleared this flag it is not valid f= or + * them to leave the VMA_WRITE_BIT flag set. + * + * As a consequence, the anonymous reverse mapping only tracks CoW mapping= s. + * + * Returns: true if the flags indicate a CoW mapping, otherwise false. + */ +static inline bool vma_flags_is_cow_mapping(const vma_flags_t *flags) +{ + return vma_flags_test(flags, VMA_MAYWRITE_BIT) && + !vma_flags_test(flags, VMA_SHARED_BIT); +} + +/** + * vma_is_cow_mapping() - Is this VMA a CoW mapping? + * @desc: The VMA to check. + * + * See vma_flags_is_cow_mapping() for details. + * + * Returns: true if the VMA is a CoW mapping, otherwise false. + */ +static inline bool vma_is_cow_mapping(const struct vm_area_struct *vma) { - return (flags & (VM_SHARED | VM_MAYWRITE)) =3D=3D VM_MAYWRITE; + return vma_flags_is_cow_mapping(&vma->flags); } =20 +/** + * vma_desc_is_cow_mapping() - Is this VMA descriptor a CoW mapping? + * @desc: The VMA descriptor to check. + * + * See vma_flags_is_cow_mapping() for details. + * + * Returns: true if the VMA descriptor describes a CoW mapping, otherwise + * false. + */ static inline bool vma_desc_is_cow_mapping(struct vm_area_desc *desc) { - const vma_flags_t *flags =3D &desc->vma_flags; - - return vma_flags_test(flags, VMA_MAYWRITE_BIT) && - !vma_flags_test(flags, VMA_SHARED_BIT); + return vma_flags_is_cow_mapping(&desc->vma_flags); } =20 #ifndef CONFIG_MMU diff --git a/kernel/events/uprobes.c b/kernel/events/uprobes.c index ae2f3b9f8d50..eb0d11092fb3 100644 --- a/kernel/events/uprobes.c +++ b/kernel/events/uprobes.c @@ -513,7 +513,7 @@ int uprobe_write(struct arch_uprobe *auprobe, struct vm= _area_struct *vma, =20 uprobe =3D container_of(auprobe, struct uprobe, arch); =20 - if (WARN_ON_ONCE(!is_cow_mapping(vma->vm_flags))) + if (WARN_ON_ONCE(!vma_is_cow_mapping(vma))) return -EINVAL; =20 /* diff --git a/mm/gup.c b/mm/gup.c index 1d32e9a3dc79..b534ef58c46a 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -1236,7 +1236,7 @@ static int check_vma_flags(struct vm_area_struct *vma= , unsigned long gup_flags) * Anon pages in shared mappings are surprising: now * just reject it. */ - if (!is_cow_mapping(vm_flags)) + if (!vma_is_cow_mapping(vma)) return -EFAULT; } } else if (!(vm_flags & VM_READ)) { diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 9b1f3b24f7e0..6b0cabd45b2d 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -1638,7 +1638,7 @@ vm_fault_t vmf_insert_pfn_pmd(struct vm_fault *vmf, u= nsigned long pfn, BUG_ON(!(vma->vm_flags & (VM_PFNMAP|VM_MIXEDMAP))); BUG_ON((vma->vm_flags & (VM_PFNMAP|VM_MIXEDMAP)) =3D=3D (VM_PFNMAP|VM_MIXEDMAP)); - BUG_ON((vma->vm_flags & VM_PFNMAP) && is_cow_mapping(vma->vm_flags)); + BUG_ON((vma->vm_flags & VM_PFNMAP) && vma_is_cow_mapping(vma)); =20 pfnmap_setup_cachemode_pfn(pfn, &pgprot); =20 @@ -1746,7 +1746,7 @@ vm_fault_t vmf_insert_pfn_pud(struct vm_fault *vmf, u= nsigned long pfn, BUG_ON(!(vma->vm_flags & (VM_PFNMAP|VM_MIXEDMAP))); BUG_ON((vma->vm_flags & (VM_PFNMAP|VM_MIXEDMAP)) =3D=3D (VM_PFNMAP|VM_MIXEDMAP)); - BUG_ON((vma->vm_flags & VM_PFNMAP) && is_cow_mapping(vma->vm_flags)); + BUG_ON((vma->vm_flags & VM_PFNMAP) && vma_is_cow_mapping(vma)); =20 pfnmap_setup_cachemode_pfn(pfn, &pgprot); =20 @@ -1888,7 +1888,7 @@ int copy_huge_pmd(struct mm_struct *dst_mm, struct mm= _struct *src_mm, * applied special bit, or we made the PRIVATE mapping be * able to wrongly write to the backend MMIO. */ - VM_WARN_ON_ONCE(is_cow_mapping(src_vma->vm_flags) && pmd_write(pmd)); + VM_WARN_ON_ONCE(vma_is_cow_mapping(src_vma) && pmd_write(pmd)); goto set_pmd; } =20 @@ -2009,7 +2009,7 @@ int copy_huge_pud(struct mm_struct *dst_mm, struct mm= _struct *src_mm, * TODO: once we support anonymous pages, use * folio_try_dup_anon_rmap_*() and split if duplicating fails. */ - if (is_cow_mapping(vma->vm_flags) && pud_write(pud)) { + if (vma_is_cow_mapping(vma) && pud_write(pud)) { pudp_set_wrprotect(src_mm, addr, src_pud); pud =3D pud_wrprotect(pud); } diff --git a/mm/hugetlb.c b/mm/hugetlb.c index d86a27c8819c..560e85ca1d1c 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -4905,7 +4905,7 @@ int copy_hugetlb_page_range(struct mm_struct *dst, st= ruct mm_struct *src, pte_t *src_pte, *dst_pte, entry; struct folio *pte_folio; unsigned long addr; - bool cow =3D is_cow_mapping(src_vma->vm_flags); + bool cow =3D vma_is_cow_mapping(src_vma); struct hstate *h =3D hstate_vma(src_vma); unsigned long sz =3D huge_page_size(h); unsigned long npages =3D pages_per_huge_page(h); diff --git a/mm/internal.h b/mm/internal.h index f26423de4ca2..a75a1344e9ba 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1341,7 +1341,7 @@ static inline bool gup_must_unshare(struct vm_area_st= ruct *vma, * ... because we only care about writable private ("COW") * mappings where we have to break COW early. */ - return is_cow_mapping(vma->vm_flags); + return vma_is_cow_mapping(vma); } =20 /* Paired with a memory barrier in folio_try_share_anon_rmap_*(). */ diff --git a/mm/memory.c b/mm/memory.c index d5e87624f692..e1349e18f046 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -780,7 +780,7 @@ static inline struct page *__vm_normal_page(struct vm_a= rea_struct *vma, /* Only CoW'ed anon folios are "normal". */ if (pfn =3D=3D index) return NULL; - if (!is_cow_mapping(vma->vm_flags)) + if (!vma_is_cow_mapping(vma)) return NULL; } } @@ -1002,7 +1002,6 @@ copy_nonpresent_pte(struct mm_struct *dst_mm, struct = mm_struct *src_mm, pte_t *dst_pte, pte_t *src_pte, struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma, unsigned long addr, int *rss) { - vm_flags_t vm_flags =3D dst_vma->vm_flags; pte_t orig_pte =3D ptep_get(src_pte); softleaf_t entry =3D softleaf_from_pte(orig_pte); pte_t pte =3D orig_pte; @@ -1026,7 +1025,7 @@ copy_nonpresent_pte(struct mm_struct *dst_mm, struct = mm_struct *src_mm, rss[mm_counter(folio)]++; =20 if (!softleaf_is_migration_read(entry) && - is_cow_mapping(vm_flags)) { + vma_is_cow_mapping(dst_vma)) { /* * COW mappings require pages in both parent and child * to be set to read. A previously exclusive entry is @@ -1067,7 +1066,7 @@ copy_nonpresent_pte(struct mm_struct *dst_mm, struct = mm_struct *src_mm, * save and restore device driver state). */ if (softleaf_is_device_private_write(entry) && - is_cow_mapping(vm_flags)) { + vma_is_cow_mapping(dst_vma)) { entry =3D make_readable_device_private_entry( swp_offset(entry)); pte =3D swp_entry_to_pte(entry); @@ -1082,7 +1081,7 @@ copy_nonpresent_pte(struct mm_struct *dst_mm, struct = mm_struct *src_mm, * exclusive entries currently only support private writable * (ie. COW) mappings. */ - VM_BUG_ON(!is_cow_mapping(src_vma->vm_flags)); + VM_BUG_ON(!vma_is_cow_mapping(src_vma)); if (try_restore_exclusive_pte(src_vma, addr, src_pte, orig_pte)) return -EBUSY; return -ENOENT; @@ -1181,7 +1180,7 @@ static __always_inline void __copy_present_ptes(struc= t vm_area_struct *dst_vma, } =20 /* If it's a COW mapping, write protect it both processes. */ - if (is_cow_mapping(src_vma->vm_flags) && writable) { + if (vma_is_cow_mapping(src_vma) && writable) { wrprotect_ptes(src_mm, addr, src_pte, nr); pte =3D pte_wrprotect(pte); } @@ -1602,9 +1601,9 @@ copy_page_range(struct vm_area_struct *dst_vma, struc= t vm_area_struct *src_vma) * We need to invalidate the secondary MMU mappings only when * there could be a permission downgrade on the ptes of the * parent mm. And a permission downgrade will only happen if - * is_cow_mapping() returns true. + * vma_is_cow_mapping() returns true. */ - is_cow =3D is_cow_mapping(src_vma->vm_flags); + is_cow =3D vma_is_cow_mapping(src_vma); =20 if (is_cow) { mmu_notifier_range_init(&range, MMU_NOTIFY_PROTECTION_PAGE, @@ -2388,7 +2387,7 @@ static bool vm_mixed_zeropage_allowed(struct vm_area_= struct *vma) if (mm_forbids_zeropage(vma->vm_mm)) return false; /* zeropages in COW mappings are common and unproblematic. */ - if (is_cow_mapping(vma->vm_flags)) + if (vma_is_cow_mapping(vma)) return true; /* Mappings that do not allow for writable PTEs are unproblematic. */ if (!(vma->vm_flags & (VM_WRITE | VM_MAYWRITE))) @@ -2839,7 +2838,7 @@ vm_fault_t vmf_insert_pfn_prot(struct vm_area_struct = *vma, unsigned long addr, BUG_ON(!(vma->vm_flags & (VM_PFNMAP|VM_MIXEDMAP))); BUG_ON((vma->vm_flags & (VM_PFNMAP|VM_MIXEDMAP)) =3D=3D (VM_PFNMAP|VM_MIXEDMAP)); - BUG_ON((vma->vm_flags & VM_PFNMAP) && is_cow_mapping(vma->vm_flags)); + BUG_ON((vma->vm_flags & VM_PFNMAP) && vma_is_cow_mapping(vma)); BUG_ON((vma->vm_flags & VM_MIXEDMAP) && pfn_valid(pfn)); =20 if (addr < vma->vm_start || addr >=3D vma->vm_end) @@ -3251,7 +3250,7 @@ static int remap_pfn_range_prepare_vma(struct vm_area= _struct *vma, unsigned long size) { const unsigned long end =3D addr + PAGE_ALIGN(size); - const bool is_cow =3D is_cow_mapping(vma->vm_flags); + const bool is_cow =3D vma_is_cow_mapping(vma); int err; =20 err =3D get_remap_pgoff(is_cow, addr, end, vma->vm_start, vma->vm_end, @@ -6751,7 +6750,7 @@ static vm_fault_t sanitize_fault_flags(struct vm_area= _struct *vma, * FAULT_FLAG_UNSHARE only applies to COW mappings. Let's * just treat it like an ordinary read-fault otherwise. */ - if (!is_cow_mapping(vma->vm_flags)) + if (!vma_is_cow_mapping(vma)) *flags &=3D ~FAULT_FLAG_UNSHARE; } else if (*flags & FAULT_FLAG_WRITE) { /* Write faults on read-only mappings are impossible ... */ @@ -6759,7 +6758,7 @@ static vm_fault_t sanitize_fault_flags(struct vm_area= _struct *vma, return VM_FAULT_SIGSEGV; /* ... and FOLL_FORCE only applies to COW mappings. */ if (WARN_ON_ONCE(!(vma->vm_flags & VM_WRITE) && - !is_cow_mapping(vma->vm_flags))) + !vma_is_cow_mapping(vma))) return VM_FAULT_SIGSEGV; } #ifdef CONFIG_PER_VMA_LOCK diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 5720f7f54d94..3498a5651d50 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -844,7 +844,7 @@ bool folio_can_map_prot_numa(struct folio *folio, struc= t vm_area_struct *vma, return false; =20 /* Also skip shared copy-on-write folios */ - if (is_cow_mapping(vma->vm_flags) && folio_maybe_mapped_shared(folio)) + if (vma_is_cow_mapping(vma) && folio_maybe_mapped_shared(folio)) return false; =20 /* Folios are pinned and can't be migrated */ diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index 26d9210b1e8b..40ad83936b28 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -1162,6 +1162,17 @@ static inline bool vma_is_shared_maywrite(struct vm_= area_struct *vma) return is_shared_maywrite(&vma->flags); } =20 +static inline bool vma_flags_is_cow_mapping(const vma_flags_t *flags) +{ + return vma_flags_test(flags, VMA_MAYWRITE_BIT) && + !vma_flags_test(flags, VMA_SHARED_BIT); +} + +static inline bool vma_is_cow_mapping(const struct vm_area_struct *vma) +{ + return vma_flags_is_cow_mapping(&vma->flags); +} + static inline struct vm_area_struct *vma_next(struct vma_iterator *vmi) { /* --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A14C238D3E2; Thu, 13 Aug 2026 17:34:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642446; cv=none; b=EeVdKCyiOst6ScavRtQvPUWbQ1dCfN1FMiC/SM4HXPLF0k7QZ65W+dLYp1klBPZFm5qoMCIkG62tvFCIZ8RvwflSOdRbu+/Y771B0TcS0ooOhSI7zYRBEePKTc57lSdpms1bvW3bq6X9DLLxTajxL2U3kPckEp7I6/AramrO/m4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642446; c=relaxed/simple; bh=y6jzJtnWVZLf67Htx6t0c40EgHiFxaw6UOaxbttS95s=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=OSrN0wqAiqiChEI1f4qb4DIAO2IjBDcY58cuHRhJ5o97jMMcu8gmEFo0AuEXI+VLAR0D7bJIWvxZs0zKxLSwRVa8IAu44kCuhJ68K8NG6W8WjguPQyFhZylMNd1kG1vGPln0yaQ0CEZK6vFyBlOSGFcL2rbrOYiycF2YMZF8un4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=SgK/s4ib; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="SgK/s4ib" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 85A501F00A3A; Thu, 13 Aug 2026 17:33:44 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642444; bh=guTiHCsFNVG67Ttnk4SCPHHfd57LUZl91cPGNyrXtXg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=SgK/s4ibLKzjYPT7hLeXQwQh7ZM1+cr4gLJ9JOd4pcTUc0m/OXFDYZE0rK5Wk4+yp 31h2S/hTDIh4N0B2EzJrSYvW2y/4U2LshW8p6Xt5g676OTsfNMaztzJfV6FLu+G3bS m37cOGQtMxnuKYufRhLmRuyCtSdjcrkRomYAzdDEw69NDcwf2PE+8+ytoaUPUxY6tE f8NUWZoVFWcRfebSyF9BSNUobYZsem567XB2Z+WmKIBvr4fFh+KU436lNvtI69CxQ/ AMoADn5gyY67bkkypYFrf91VTnOTA9qr/0skHJ1nb5StGhDVc1IeTXhdJkpyW04wML ebAgynBxTPE7A== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:20 +0100 Subject: [PATCH v5 03/16] mm: introduce linear_anon_page_index() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-3-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=5197; i=ljs@kernel.org; h=from:subject:message-id; bh=y6jzJtnWVZLf67Htx6t0c40EgHiFxaw6UOaxbttS95s=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/69e1nbnifDj9bYaryaZ9Hi+mGUdXCh4T3y/veOmx pr85JIHHaUsDGJcDLJiiizPv4jvDxIJm9d5wd8NZg4rE8gQBi5OAZhI4TpGhk/+4c8YPnNsnRed dcW7TvuYttHbo3slig+L65cKrbtZVsDIcPri3uJLk38cuHLfttB3t03fk7l2i3+1MhdwhgtX7/1 mzwwA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This function provides the anonymous equivalent of linear_page_index(), instead offsetting based on the anonymous page offset of the VMA. It is valid only for anonymous or MAP_PRIVATE file-backed mappings, in other words CoW mappings. For pure anon VMAs, this will be equal to linear_page_index(). Assert that both of these invariants are true in linear_anon_page_index() and implement the algorithm in __linear_anon_page_index(). Note that MAP_PRIVATE-/dev/zero mappings will satisfy vma_is_anonymous() but not fulfill this invariant, so when asserting this we check vma->vm_file to account for this. We do not update callsites yet, so no functional change intended. Also const-ify vma_is_anonymous() to make it compatible with the const-ified linear_anon_page_index(). While we're here, update linear_page_index() to be more succinct. VMA userland tests are also updated accordingly. Reviewed-by: Gregory Price (Meta) Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/mm.h | 2 +- include/linux/pagemap.h | 40 +++++++++++++++++++++++++++++++++++++= --- tools/testing/vma/include/dup.h | 25 ++++++++++++++++++++++++- 3 files changed, 62 insertions(+), 5 deletions(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index 199d0a64b14a..0829e0d3b2d1 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -1556,7 +1556,7 @@ static inline void vma_desc_set_anonymous(struct vm_a= rea_desc *desc) desc->vm_ops =3D NULL; } =20 -static inline bool vma_is_anonymous(struct vm_area_struct *vma) +static inline bool vma_is_anonymous(const struct vm_area_struct *vma) { return !vma->vm_ops; } diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h index c6fc783aaee5..0adfa6605653 100644 --- a/include/linux/pagemap.h +++ b/include/linux/pagemap.h @@ -1094,10 +1094,44 @@ static inline pgoff_t linear_page_delta(const struc= t vm_area_struct *vma, static inline pgoff_t linear_page_index(const struct vm_area_struct *vma, const unsigned long address) { - pgoff_t pgoff; + return linear_page_delta(vma, address) + vma_start_pgoff(vma); +} + +static inline pgoff_t __linear_anon_page_index(const struct vm_area_struct= *vma, + const unsigned long address) +{ + return linear_page_delta(vma, address) + vma_start_anon_pgoff(vma); +} + +/** + * linear_anon_page_index() - Determine the absolute anonymous page offset= of + * @address within @vma. + * @vma: An anonymous or MAP_PRIVATE file-backed VMA in which @address res= ides. + * @address: The address whose absolute page offset is required. + * + * This returns the anonymous page offset of @address, which is the page o= ffset + * the address possessed at the time the VMA was first faulted. + * + * For anonymous mappings, this returns the same value as linear_page_inde= x(). + * + * For MAP_PRIVATE file-backed mappings, this returns the anonymous page o= ffset + * of @address, which is the page offset the address possessed at the time= the + * VMA was first faulted. + * + * It is not valid to call this function for shared file-backed mappings. + * + * Returns: The absolute anonymous page offset of @address within @vma. + */ +static inline pgoff_t linear_anon_page_index(const struct vm_area_struct *= vma, + const unsigned long address) +{ + const pgoff_t pgoff =3D __linear_anon_page_index(vma, address); + + VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma)); + /* Account for MAP_PRIVATE-/dev/zero which is only semi-anonymous. */ + if (vma_is_anonymous(vma) && !vma->vm_file) + VM_WARN_ON_ONCE(pgoff !=3D linear_page_index(vma, address)); =20 - pgoff =3D linear_page_delta(vma, address); - pgoff +=3D vma_start_pgoff(vma); return pgoff; } =20 diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index 40ad83936b28..4c58487b764e 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -1428,7 +1428,7 @@ static inline void vma_iter_set(struct vma_iterator *= vmi, unsigned long addr) mas_set(&vmi->mas, addr); } =20 -static inline bool vma_is_anonymous(struct vm_area_struct *vma) +static inline bool vma_is_anonymous(const struct vm_area_struct *vma) { return !vma->vm_ops; } @@ -1621,3 +1621,26 @@ static inline pgprot_t vma_get_page_prot(const struc= t vm_area_struct *vma) { return vma_flags_to_page_prot(vma->flags); } + +static inline pgoff_t __linear_anon_page_index(const struct vm_area_struct= *vma, + const unsigned long address) +{ + pgoff_t pgoff; + + pgoff =3D linear_page_delta(vma, address); + pgoff +=3D vma_start_anon_pgoff(vma); + return pgoff; +} + +static inline pgoff_t linear_anon_page_index(const struct vm_area_struct *= vma, + const unsigned long address) +{ + const pgoff_t pgoff =3D __linear_anon_page_index(vma, address); + + VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma)); + /* Account for MAP_PRIVATE-/dev/zero which is only semi-anonymous. */ + if (vma_is_anonymous(vma) && !vma->vm_file) + VM_WARN_ON_ONCE(pgoff !=3D linear_page_index(vma, address)); + + return pgoff; +} --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0F9EC384233; Thu, 13 Aug 2026 17:34:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642466; cv=none; b=funve0948BD7gkG7GFG6elzhkrACZ1IF8AinajZjbdKD9dRfIGYdCy8LCavJkPDPRlJ0o8YketO3DKwORiP3qaLSvgmb3KvI+ckkINnwl2DYYgPLjME+lleIsVOWvDIf+JpzHdNLz7X2k5T2F73Mt15XSk2fRumNgd8zaeVW0O4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642466; c=relaxed/simple; bh=TcLlQnmra6maBGHXmsGw4bE6RvgdQMZRWpLIfm3Ze4w=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=OXMmZGQl9WW2utEkN7V5mKClVqBzf1HeeW6fIFvJsWIIhbIigCq6o+0pSTqIfX4E/0PBflDVy/rtF8Ut5LXljQJo82/1UTu4P9wAl2TqHctaUNol74QvqZ4caimckR4app+5xgzJKusHJHcbp7OQ5ccunAmS5qPkSOi8/XcQUEs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=YugjBca5; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="YugjBca5" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DC8B21F000E9; Thu, 13 Aug 2026 17:34:04 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642464; bh=GlrT7r3tcep46r6Il05ij5JS2SxeagPngYBWHI1YDNw=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=YugjBca5WND+n3QaqWYkEL7gPewL8hAiL78fU5MbiGSIG6TMa9V7N1Od8KlaB1Zm0 EANGb2/0A0U2107blNX9VfRctXD15pC7sEti0k/dEIk4q3cg/HdHZeU2CHMchhYRJM 7C7Yjv17+AeCZvJvWcSCmGT5AO5lEZpFqT8zQkSb6rmu/qb7PglulSmb7q+y1sd+r1 X6SGHj1oJm7HLJULbGdHvPcRFKZR4cxfRBtHEwS8mBCQEYMK3isfql/LsTOtFirYte vqbd6POdg05Jn8DcGcoQJcCf/SW0NMBpmw0SqMWB+8vzrZLSOvjw36Akp2jBAQYqJf ufZdjAYY63PwA== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:21 +0100 Subject: [PATCH v5 04/16] mm: abstract vma_address() and introduce vma_anon_address() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-4-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3692; i=ljs@kernel.org; h=from:subject:message-id; bh=TcLlQnmra6maBGHXmsGw4bE6RvgdQMZRWpLIfm3Ze4w=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6/Wbt46vb61rk04ZfEGqcDu4L8/E0LvvuMzL/xVq 3/hyX/NjlIWBjEuBlkxRZbnX8T3B4mEzeu84O8GM4eVCWQIAxenAEyEM5GRYUqZfbsf1/frl++u L6lOUQ8prt7EaHKgzWrNmVxm5XjGJwz/HVSi7nZ7qYt8ljy/X7hmVt0Ug/0513/13qkWT1h332o xMwA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Introduce __vma_address() which abstracts the VMA start page offset field as pgoff_start, then update vma_address() to use it. Then introduce vma_anon_address() which does the equivalent of vma_address(), only using the anonymous page offset of the VMA rather than the file-backed one. Also add an assert to ensure that the function is not called for mappings which are file-backed but not MAP_PRIVATE to ensure it is only used in the correct places. This will be necessary for determining the address of a folio's index within a VMA when the folio belongs to a MAP_PRIVATE file-backed VMA but has been CoW'd, and thus is anonymous, once the anonymous VMA page offset field is used for the reverse mapping. No callers are updated, so no functional change intended. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- mm/internal.h | 49 +++++++++++++++++++++++++++++++++++++------------ 1 file changed, 37 insertions(+), 12 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index a75a1344e9ba..10b6662be8ed 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1005,19 +1005,9 @@ void mlock_drain_remote(int cpu); =20 extern pmd_t maybe_pmd_mkwrite(pmd_t pmd, struct vm_area_struct *vma); =20 -/** - * vma_address - Find the virtual address a page range is mapped at - * @vma: The vma which maps this object. - * @pgoff: The page offset within its object. - * @nr_pages: The number of pages to consider. - * - * If any page in this range is mapped by this VMA, return the first addre= ss - * where any of these pages appear. Otherwise, return -EFAULT. - */ -static inline unsigned long vma_address(const struct vm_area_struct *vma, - pgoff_t pgoff, unsigned long nr_pages) +static inline unsigned long __vma_address(const struct vm_area_struct *vma, + pgoff_t pgoff, pgoff_t pgoff_start, unsigned long nr_pages) { - const pgoff_t pgoff_start =3D vma_start_pgoff(vma); unsigned long address; =20 if (pgoff >=3D pgoff_start) { @@ -1035,6 +1025,41 @@ static inline unsigned long vma_address(const struct= vm_area_struct *vma, return address; } =20 +/** + * vma_address - Find the virtual address a page range is mapped at. + * @vma: The vma which maps this object. + * @pgoff: The page offset within its object. + * @nr_pages: The number of pages to consider. + * + * If any page in this range is mapped by this VMA, return the first addre= ss + * where any of these pages appear. Otherwise, return -EFAULT. + */ +static inline unsigned long vma_address(const struct vm_area_struct *vma, + pgoff_t pgoff, unsigned long nr_pages) +{ + return __vma_address(vma, pgoff, vma_start_pgoff(vma), nr_pages); +} + +/** + * vma_anon_address - Find the address an anonymous folio with index @pgof= f_anon + * is mapped at. + * @vma: The vma which maps this object. + * @pgoff_anon: The anonymous page index belonging to the folio. + * @nr_pages: The number of pages to consider. + * + * This is only valid for anonymous or MAP_PRIVATE-mapped file-backed VMAs. + * + * Returns: If any page in this range is mapped by this VMA, return the fi= rst + * address where any of these pages appear. Otherwise, return -EFAULT. + */ +static inline unsigned long vma_anon_address(const struct vm_area_struct *= vma, + pgoff_t pgoff_anon, unsigned long nr_pages) +{ + VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma)); + + return __vma_address(vma, pgoff_anon, vma_start_anon_pgoff(vma), nr_pages= ); +} + /* * Then at what user virtual address will none of the range be found in vm= a? * Assumes that vma_address() already returned a good starting address. --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2633B38C427; Thu, 13 Aug 2026 17:34:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642486; cv=none; b=KmzY6mDRQg1HlUGNoxCSzcop7nWEf5QVWwOYlBUDmEmX0JfSa8Wxl8OZmBJE36UM0JikdMfKOpfm/Cp7tEKfmAqDrS1CoiKXSFdHuOeqEr7e48pHzXgylRIRDSy9D/Q0QlTffmY2bbItUC3F79UTcD38lnnSr1rkrsfzXlUKr7c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642486; c=relaxed/simple; bh=n0kfl3sX7UAIKHIjCWNuf6tCoy4kVBTkNam6LpgmDSU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=CDzmbEdxufdZTov0HbjPMwHFautOLCJnheE3f7/SUnkyndk8D3CaYdtf2PALcdzLNpXS1Csm8cirdoWSuWThCvUw/Q7qEq6+w7ZZ+ZN2nA+wvYuP40yrJtC07BuRzazg3REjIdpOvR79G+BwDVAOXHijJaRZx1Lgr67yQsRJCJE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=UJHrHFR6; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="UJHrHFR6" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 36FC91F00A3D; Thu, 13 Aug 2026 17:34:25 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642485; bh=7Nmbnq0r/tdkn7LzRpvbEkYwmsG/p8Yg0FdX7/Lw0rs=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=UJHrHFR6u9p/SM5kX+G8cV/Zu8+jiCGjo+Jl8zwrOiYWKlhFQHjrMyRDRNsZdldmv qJ0aMYRTHVmuUDG4AHUa3Bb3WaxwtrGHmulqnc8mlxF9KrA0vAEeVIomi2htKiZrgb YZh0DG6jlBGdxPV3Wl2SHTPL3OqhwEuftXqsDqCpS9zoJ2i4s6yIUE03wCgiNiUFs0 jN3yqFoE+yZmmjdF8c/mz5tD1/heVisks35JsytpQPAcWSDW4QSWP8FZ7NsCznwGdk ZVBLVAm+mEJLe3w7sFC0/d2PWfgR+77ptjHLg6PgEVZGVdeGk14zwD6i3v978rzzou wJHNd1ySkH6JQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:22 +0100 Subject: [PATCH v5 05/16] mm: update print_bad_page_map() to show anon index if appropriate Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-5-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2529; i=ljs@kernel.org; h=from:subject:message-id; bh=n0kfl3sX7UAIKHIjCWNuf6tCoy4kVBTkNam6LpgmDSU=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/69+Xib95eGvG37S4v815IoWPK0tbjuZt/u6LHdR6 a6nEybHdZSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAi84sYGa65SWk+CRW/1Sku LxRzNzi47kdjggp768q9Qnr9+XkalYwMh2PfPbxWn2G4NyWQ6VL1h9zHN831Z6yZwneohWMye2g uIwA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 If the VMA is a CoW mapping page offset may differ from anon page offset, indicating different positions in the relevant rmap trees. Update print_bad_page_map() to reflect that - if the mapping is non-CoW or the indexes match, then output only one index as before, otherwise output both with (file) or (anon) suffixes to reflect which is which. It's not possible to give only one index as there is no folio available to perform folio_test_anon() upon (the page table entry is bad so this is unavailable). This is potentially useful debugging information and matches the existing page offset provided. Use the raw __linear_anon_page_index() function so as to always output this value regardless of whether the mapping is file-backed or not and to avoid asserts that shouldn't apply here. Acked-by: David Hildenbrand (Arm) Reviewed-by: Gregory Price (Meta) Signed-off-by: Lorenzo Stoakes (ARM) --- mm/memory.c | 13 ++++++++++--- 1 file changed, 10 insertions(+), 3 deletions(-) diff --git a/mm/memory.c b/mm/memory.c index e1349e18f046..3bd3616b27d2 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -631,13 +631,14 @@ static void print_bad_page_map(struct vm_area_struct = *vma, { struct address_space *mapping; char entry_str[PTVAL_STR_MAX]; - pgoff_t index; + pgoff_t index, anon_index; =20 if (is_bad_page_map_ratelimited()) return; =20 mapping =3D vma->vm_file ? vma->vm_file->f_mapping : NULL; index =3D linear_page_index(vma, addr); + anon_index =3D __linear_anon_page_index(vma, addr); =20 ptval_bytes_to_hex_str(entry_str, sizeof(entry_str), entry, entry_size); pr_alert("BUG: Bad page map in process %s %s:%s", current->comm, @@ -645,8 +646,14 @@ static void print_bad_page_map(struct vm_area_struct *= vma, __print_bad_page_map_pgtable(vma->vm_mm, addr); if (page) dump_page(page, "bad page map"); - pr_alert("addr:%px vm_flags:%08lx anon_vma:%px mapping:%px index:%lx\n", - (void *)addr, vma->vm_flags, vma->anon_vma, mapping, index); + pr_alert("addr:%px vm_flags:%08lx anon_vma:%px mapping:%px", + (void *)addr, vma->vm_flags, vma->anon_vma, mapping); + if (!vma_is_cow_mapping(vma) || index =3D=3D anon_index) { + pr_cont(" index:%lx\n", index); + } else { + pr_cont(" index:%lx (file) %lx (anon)\n", index, anon_index); + } + pr_alert("file:%pD fault:%ps mmap:%ps mmap_prepare: %ps read_folio:%ps\n", vma->vm_file, vma->vm_ops ? vma->vm_ops->fault : NULL, --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0682F3911A1; Thu, 13 Aug 2026 17:35:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642507; cv=none; b=bFA37SPYchKzeUI+Iu9GIf2l/OTzhxXYbCRR0VeMcoHKPcRb90FBuGiSxqepfsILOMNIWJrK+XlXURCWA38jmR11Eial9AFJWr5rKk5tF/9siVtQSGg7i/BVRZyC7oLFmVo9Z8s06kRLVVlRoqqYDRQB+wOpMCWpLr7gvbRxkGc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642507; c=relaxed/simple; bh=p77ev8vuIe2NzptTm2ieSg/F7FGTlLmZaKMePPxwgJc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=gX4z6YtUiIbxQhI2+3hhBZTkO87rZGTNTnsg6ednbfS9R8WIcoa6Iz0QeXrQM/STpHzYx17RdPTYgMnKKTOyS7CcwVRDCeXO0Gd06VUe4nQ6sqeVoEwnPKpgUB/KnUPchQyXbH4Kql0qwgKd7uaVSoPdEzLODAg2R9THJb7+5ck= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=OtSqY/H9; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="OtSqY/H9" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 8889F1F000E9; Thu, 13 Aug 2026 17:34:45 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642505; bh=jHeRoJJr2O8/TaKftY2r1DUlL3KME6EykG7g/IHvkjk=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=OtSqY/H9sqJEAHUIMd6G0CbKd8RO2TgYWZL9dRMCbGMnxyxtrTZ04YIhDmmPnZyIL dnCdS7meG7AfAt7cLXD3GxZEn2WAtW+YfBj+PQxek5U7GYVd1tFylkXiMKk9h41H7z uCLFDAsS3nGgEjWs13IqMvGiqlx5syYnHtUwxZKa3t16U0wX6QtP4TfnBu/Y/6O7zy HU4BlORDw5SdeJ67S+cNDuC9Ls6tRrs9gcd78uXLdG1DP83NeG02vLznkJkgdka12a GMhsfDn3hS4W0UXSMb/G8QLz6iVo5jISGfrVbGsRKzp/Gc+W5CS/92+ObTJ1J2xv/d SjUsZsi6nh6aQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:23 +0100 Subject: [PATCH v5 06/16] mm: introduce and use vma_filebacked_address() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-6-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org, syzbot@syzkaller.appspotmail.com X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4971; i=ljs@kernel.org; h=from:subject:message-id; bh=p77ev8vuIe2NzptTm2ieSg/F7FGTlLmZaKMePPxwgJc=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6/eIlF58FHJ3ilGt/2fyeoEGh1/4X/wOcPUA13zX 982fyUr31HKwiDGxSArpsjy/Iv4/iCRsHmdF/zdYOawMoEMYeDiFICJmN5jZLgr3nP/0lvxFM1u Lk+hijXPjXq0/brua4Y/XOLMuiBKrpvhf1HgrzWhx+6Z7u+w6692iE7o/v/gyv0fBbm1TcJ9iis +MwAA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 In cases where we know that the VMA is file-backed, use vma_filebacked_address() rather than vma_address(). This lays the foundation for using the anonymous page offset via vma_anon_address(). Also add an assert to ensure that the VMA whose address is required is not anonymous. No functional change intended. Tested-by: syzbot@syzkaller.appspotmail.com Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- mm/internal.h | 18 ++++++++++++++++++ mm/memory-failure.c | 4 ++-- mm/page_vma_mapped.c | 6 +++++- mm/rmap.c | 10 ++++++---- 4 files changed, 31 insertions(+), 7 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index 10b6662be8ed..03145b8d0d56 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1025,6 +1025,24 @@ static inline unsigned long __vma_address(const stru= ct vm_area_struct *vma, return address; } =20 +/** + * vma_filebacked_address - Find the virtual address a file-backed page ra= nge is + * mapped at. + * @vma: The vma which maps this object. + * @pgoff: The page offset within its object. + * @nr_pages: The number of pages to consider. + * + * Returns: If any page in this range is mapped by this VMA, return the fi= rst + * address where any of these pages appear. Otherwise, return -EFAULT. + */ +static inline unsigned long vma_filebacked_address(const struct vm_area_st= ruct *vma, + pgoff_t pgoff, unsigned long nr_pages) +{ + VM_WARN_ON_ONCE(vma_is_anonymous(vma)); + + return __vma_address(vma, pgoff, vma_start_pgoff(vma), nr_pages); +} + /** * vma_address - Find the virtual address a page range is mapped at. * @vma: The vma which maps this object. diff --git a/mm/memory-failure.c b/mm/memory-failure.c index aaf14608b30e..a8b03e2920ba 100644 --- a/mm/memory-failure.c +++ b/mm/memory-failure.c @@ -620,7 +620,7 @@ static void add_to_kill_fsdax(struct task_struct *tsk, = const struct page *p, struct vm_area_struct *vma, struct list_head *to_kill, pgoff_t pgoff) { - unsigned long addr =3D vma_address(vma, pgoff, 1); + unsigned long addr =3D vma_filebacked_address(vma, pgoff, 1); __add_to_kill(tsk, p, vma, to_kill, addr); } =20 @@ -2265,7 +2265,7 @@ static void add_to_kill_pgoff(struct task_struct *tsk, } =20 /* Check for pgoff not backed by struct page */ - tk->addr =3D vma_address(vma, pgoff, 1); + tk->addr =3D vma_filebacked_address(vma, pgoff, 1); tk->size_shift =3D PAGE_SHIFT; =20 if (tk->addr =3D=3D -EFAULT) diff --git a/mm/page_vma_mapped.c b/mm/page_vma_mapped.c index d7670ba4147b..081e483cc7bf 100644 --- a/mm/page_vma_mapped.c +++ b/mm/page_vma_mapped.c @@ -356,6 +356,7 @@ unsigned long page_mapped_in_vma(const struct page *pag= e, struct vm_area_struct *vma) { const struct folio *folio =3D page_folio(page); + const pgoff_t pgoff =3D page_pgoff(folio, page); struct page_vma_mapped_walk pvmw =3D { .pfn =3D page_to_pfn(page), .nr_pages =3D 1, @@ -363,7 +364,10 @@ unsigned long page_mapped_in_vma(const struct page *pa= ge, .flags =3D PVMW_SYNC, }; =20 - pvmw.address =3D vma_address(vma, page_pgoff(folio, page), 1); + if (folio_test_anon(folio)) + pvmw.address =3D vma_address(vma, pgoff, 1); + else + pvmw.address =3D vma_filebacked_address(vma, pgoff, 1); if (pvmw.address =3D=3D -EFAULT) goto out; if (!page_vma_mapped_walk(&pvmw)) diff --git a/mm/rmap.c b/mm/rmap.c index ad820fe86f7d..5798427d007f 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -865,14 +865,15 @@ unsigned long page_address_in_vma(const struct folio = *folio, if (!vma->anon_vma || !anon_vma || vma->anon_vma->root !=3D anon_vma->root) return -EFAULT; + /* KSM folios don't reach here because of the !anon_vma check */ + return vma_address(vma, page_pgoff(folio, page), 1); } else if (!vma->vm_file) { return -EFAULT; } else if (vma->vm_file->f_mapping !=3D folio->mapping) { return -EFAULT; } =20 - /* KSM folios don't reach here because of the !anon_vma check */ - return vma_address(vma, page_pgoff(folio, page), 1); + return vma_filebacked_address(vma, page_pgoff(folio, page), 1); } =20 /* @@ -1321,7 +1322,7 @@ int pfn_mkclean_range(unsigned long pfn, unsigned lon= g nr_pages, pgoff_t pgoff, if (invalid_mkclean_vma(vma, NULL)) return 0; =20 - pvmw.address =3D vma_address(vma, pgoff, nr_pages); + pvmw.address =3D vma_filebacked_address(vma, pgoff, nr_pages); VM_BUG_ON_VMA(pvmw.address =3D=3D -EFAULT, vma); =20 return page_vma_mkclean_one(&pvmw); @@ -3100,7 +3101,8 @@ static void __rmap_walk_file(struct folio *folio, str= uct address_space *mapping, } lookup: mapping_rmap_tree_foreach(vma, mapping, pgoff_start, pgoff_end) { - unsigned long address =3D vma_address(vma, pgoff_start, nr_pages); + unsigned long address =3D vma_filebacked_address(vma, pgoff_start, + nr_pages); =20 VM_BUG_ON_VMA(address =3D=3D -EFAULT, vma); cond_resched(); --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 55A2A2DF6F4; Thu, 13 Aug 2026 17:35:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642527; cv=none; b=LCSRFL2eM87BguNfLSCwk5/YqRVqp29wwF8cx6K+mrSLd2wneM0PsWg9PrDgMomgWYwsjy8wDDk326uqPEG2827j0tQaEziJ4xRUXC3eybztAyKY3PuiIC+JC0B3TVR7zSdEpYDwCZbfIzLtbQZ6E5Mi6NCx7Yk4hOnHcmPU3vY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642527; c=relaxed/simple; bh=iRuXFKdiG3IkLvRnILJFDmJ//0Yv8O86ToLE8ObFySw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Lfv2kniupPgIcvw6AV4CBt1Witv/xBXoyWPG4HzJWHufmEocA6JrfdMyrTOFJsT1xRN59ujcLd91sNemhc6Kd+4DoEZ+mR8TKfL4M1RvgG9cYECFV6d2ao0/++U9tmBWJ9ojnACfVxuc/GwX5Xj6huNaD2uQ0uMl0lXvRGkTxlQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=XNf5Wqev; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="XNf5Wqev" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1B3501F00A3A; Thu, 13 Aug 2026 17:35:05 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642525; bh=X99YrB4uFf+BkWFFnJLytE8bkb5ZotMsaFUlyvfCIZQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=XNf5WqevkY16a6WJ2U35Q3OHl/tyzB1cz5rTvdXpzbP6zLFiN95i55598rx8OAokr pAt7jjRlHuk9RD+4vZpkMJzcqKLxezgQf77U/kqBSynG0MeCmG1Tvxp9+bvQB+HGzc hpPffbwvYrFa6DyItZpYfJuijUl+Q4SGIE0zvG85wwJq6ybwi1lt09dr3WEblMKiSE gKii5/3FH+zzeUXlpmZtjcBB2kh3UuEN2q/H9PWmVTA2n8k5eafEdV4AzRd6xtju+y aFBZY6bfsuFCQen98b16YCyG8l+tx77HOZ1rqjB/IAIskliMSgfLdMybbLDj5LSRb0 mVoJD2PYoWxDA== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:24 +0100 Subject: [PATCH v5 07/16] mm/vma: fix self-merge check in copy_vma() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-7-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=5550; i=ljs@kernel.org; h=from:subject:message-id; bh=iRuXFKdiG3IkLvRnILJFDmJ//0Yv8O86ToLE8ObFySw=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6/+/Ep125oQKTu+H/vzcu/edg7UWT2J586dghdq8 m0CyYxyHaUsDGJcDLJiiizPv4jvDxIJm9d5wd8NZg4rE8gQBi5OAZhIfR8jwzbrbNGa5mgTTpnY dZzW39vq2y7dS5rWIvdz5lqhFJetXIwMm++YLbCKOdf9y/WU3Yb7O1+d1TDb+iCo9cidqKWPy6a xMgIA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 The existing logic is very confusing so improve things. Firstly rename the confusing faulted_in_anon_vma variable to can_self_merge and update this when the page offset is updated. What is being checked for is a 'self-merge' - that is between the VMA being remapped and its prior VMA (remember that this is copy_vma() - if a non-MREMAP_DONTUNMAP remap the original VMA is only removed afterwards). This can happen if the VMA is moved immediately adjacent to itself, either before or after it: |----------------|----------------| | | | v | v |...............||---------------||...............| | new || old || new | |...............||---------------||---------------| In these cases the old VMA is simply expanded to cover the new range. It is also possible for the move to both self-merge and merge with a prior VMA if it is placed between a preceding VMA and its old self: |---------------| | | v | |---------------||...............||---------------| | prev || new || old | |---------------||...............||---------------| In this case, the old VMA is removed and 'prev' is expanded and replaces it. Since copy_vma_and_data() which calls copy_vma() intends to reference the old VMA after the merge, it must have this pointer updated. This kind of self-merge is not possible with a succeeding merge, as the merge always prefers to expand the preceding VMA if possible. copy_vma() accounts for this by explicitly checking to see if a self-merge occurred and updating the vmap pointer if so. However it incorrect did so even for a subsequent merge (this is simply a noop so it had no impact). So change this to only check for the case which matters - a backwards merge - and rearrange the parameters to make it clearer we're doing that - i.e. check new_vma->vm_start < old_vma_start (having already renamed vma_start to old_vma_start to make it clear this is the previous VMA). Also update the existing wall-of-text comment to be a lot clearer. While we're here, replace the VM_BUG_ON_VMA() with a VM_WARN_ON_ONCE_VMA() and update the VMA userland tests accordingly. No functional change intended. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- mm/vma.c | 35 ++++++++++++++++------------------- tools/testing/vma/vma_internal.h | 1 + 2 files changed, 17 insertions(+), 19 deletions(-) diff --git a/mm/vma.c b/mm/vma.c index b5bc3eec961c..18c8f2546765 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -1911,10 +1911,10 @@ struct vm_area_struct *copy_vma(struct vm_area_stru= ct **vmap, bool *need_rmap_locks) { struct vm_area_struct *vma =3D *vmap; - unsigned long vma_start =3D vma->vm_start; + unsigned long old_vma_start =3D vma->vm_start; struct mm_struct *mm =3D vma->vm_mm; struct vm_area_struct *new_vma; - bool faulted_in_anon_vma =3D true; + bool can_self_merge =3D false; VMA_ITERATOR(vmi, mm, addr); VMG_VMA_STATE(vmg, &vmi, NULL, vma, addr, addr + len); =20 @@ -1924,7 +1924,7 @@ struct vm_area_struct *copy_vma(struct vm_area_struct= **vmap, */ if (unlikely(vma_is_anonymous(vma) && !vma->anon_vma)) { pgoff =3D addr >> PAGE_SHIFT; - faulted_in_anon_vma =3D false; + can_self_merge =3D true; } =20 /* @@ -1944,24 +1944,21 @@ struct vm_area_struct *copy_vma(struct vm_area_stru= ct **vmap, new_vma =3D vma_merge_copied_range(&vmg); =20 if (new_vma) { - /* - * Source vma may have been merged into new_vma - */ - if (unlikely(vma_start >=3D new_vma->vm_start && - vma_start < new_vma->vm_end)) { + /* Self-merged and VMA replaced. */ + if (unlikely(new_vma->vm_start < old_vma_start && + new_vma->vm_end > old_vma_start)) { /* - * The only way we can get a vma_merge with - * self during an mremap is if the vma hasn't - * been faulted in yet and we were allowed to - * reset the dst vma->vm_pgoff to the - * destination address of the mremap to allow - * the merge to happen. mremap must change the - * vm_pgoff linearity between src and dst vmas - * (in turn preventing a vma_merge) to be - * safe. It is only safe to keep the vm_pgoff - * linear if there are no pages mapped yet. + * The only way a VMA can both self-merge and be + * replaced is if the remap places the new VMA + * immediately prior to its old self ('next') and + * immediately after another VMA ('prev') causing the + * next to be removed and prev to be expanded to cover + * the entire range. + * + * This should only be possible if the page offset was + * updated, i.e. the VMA is unfaulted. */ - VM_BUG_ON_VMA(faulted_in_anon_vma, new_vma); + VM_WARN_ON_ONCE_VMA(!can_self_merge, new_vma); *vmap =3D vma =3D new_vma; } *need_rmap_locks =3D diff --git a/tools/testing/vma/vma_internal.h b/tools/testing/vma/vma_inter= nal.h index 4f6c5666ac07..8a48b231aa7a 100644 --- a/tools/testing/vma/vma_internal.h +++ b/tools/testing/vma/vma_internal.h @@ -53,6 +53,7 @@ typedef __bitwise unsigned int vm_fault_t; =20 #define VM_WARN_ON(_expr) (WARN_ON(_expr)) #define VM_WARN_ON_ONCE(_expr) (WARN_ON_ONCE(_expr)) +#define VM_WARN_ON_ONCE_VMA(_expr, _vma) (WARN_ON_ONCE(_expr)) #define VM_WARN_ON_VMG(_expr, _vmg) (WARN_ON(_expr)) #define VM_BUG_ON(_expr) (BUG_ON(_expr)) #define VM_BUG_ON_VMA(_expr, _vma) (BUG_ON(_expr)) --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A750A38BF8D; Thu, 13 Aug 2026 17:35:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642548; cv=none; b=JS4+yktaG4cVEVsi4/UDmMgngUwLKN2DbSHrsab9Sk7UOIoW2WD+I0LcbNUHSCcJt5BGZVfiMD0+b7C44DK3A2XsRXGpOX/xBYTjMuWFrMvbrOIP9+Deel+BD/VO7T3fVF+y+cK408O5BtXpl+yFcPnba31sYOup+NZSp1U+U8g= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642548; c=relaxed/simple; bh=0z311QDVDbRkLzxpR827T28yvRS4ZLJp7flLC+XA2kc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=KyH8iu0goSNNqXWN6/utGOT3JoXpZfNn0LHE3Ku062/SYn7uJ8yhymE+7GQK8AOhnBfdzKkJIHlkqhsUi/WWoA0IjA5ZicJSrxZ5mgXfstXvqdZKxNhoYk+UmHtrxOM+WdYsd8vKBw4SDMT5aZ0jh3vW1/XkABfG4BQ8dkF0mFs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=PkmLlMyr; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="PkmLlMyr" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 6B4661F000E9; Thu, 13 Aug 2026 17:35:26 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642546; bh=e7idCKxvuxc1VD3/Y+D8HZ1eKufEBlsCtRsBSzrHM2Y=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=PkmLlMyrsX2m5fanA2c5rpxa/EifX3xwSAX8xxMpwcWxhpca2+mz8b35ptW4+WUVm z5Jv6LMlKfeILZji0OtaDQxFCCuYwYZKqz/fII4SzWAm8mnDZXWbA+khF9g338T4J5 iOjn02HhY5DO7GFnYnzbPjLi98OSMn6LpTiXqX9llMaiiS12b12txoGGobH8Gmhl8J bgk0Umr5fyJRavOes1pEb0n1WbIsDmOsjkxuME1WI6oPzyciES452Tf3RI3jr3YGo9 ivCA7S0VWHHo22BAdkUUVcxX8wBXPSxScIu20TJaK4Zg4SXT7cznU5Tr4mK24yPlsf esvmH+OknqRBQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:25 +0100 Subject: [PATCH v5 08/16] tools/testing/vma: add tests for copy_vma() self-merge Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-8-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2596; i=ljs@kernel.org; h=from:subject:message-id; bh=0z311QDVDbRkLzxpR827T28yvRS4ZLJp7flLC+XA2kc=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/69Z1VT5fCvv9WieeoYZE4xYRd0e5rmd/HaXyaCjr b07ME22o5SFQYyLQVZMkeX5F/H9QSJh8zov+LvBzGFlAhnCwMUpABNR6mJk+HCb/2br05jeXX67 rD18dnZHCr6eJ7+M+Uf/ovddiaw/chgZjjKwZfxl+8Z1YvPXU3t1Djz67nk4X3rWBO7UMzaPco1 U+QE= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Assert that a VMA can be moved backwards, forwards and between a preceding VMA and its old self. In the cases in which the VMA merges only with itself expect that to be achieved by expanding its old self, so assert that these function correctly. However in the case of a merge between a preceding VMA and itself the original VMA is removed, so assert that the preceding VMA replaces the one passed in as vmap and the merge is as expected. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- tools/testing/vma/tests/vma.c | 46 +++++++++++++++++++++++++++++++++++++++= +++- 1 file changed, 45 insertions(+), 1 deletion(-) diff --git a/tools/testing/vma/tests/vma.c b/tools/testing/vma/tests/vma.c index 754a2da06321..0d40d7ba2181 100644 --- a/tools/testing/vma/tests/vma.c +++ b/tools/testing/vma/tests/vma.c @@ -33,7 +33,51 @@ static bool test_copy_vma(void) struct mm_struct mm =3D {}; bool need_locks =3D false; VMA_ITERATOR(vmi, &mm, 0); - struct vm_area_struct *vma, *vma_new, *vma_next; + struct vm_area_struct *vma, *vma_prev, *vma_new, *vma_next, *vma_orig; + + /* Move forwards, adjacent to old self - self-merge. */ + + vma =3D alloc_and_link_vma(&mm, 0x1000, 0x2000, 1, vma_flags); + vma_set_anonymous(vma); + vma_orig =3D vma; + vma_new =3D copy_vma(&vma, 0x2000, 0x1000, 1, &need_locks); + ASSERT_EQ(vma_new, vma_orig); + ASSERT_EQ(vma, vma_orig); + ASSERT_EQ(vma_new->vm_start, 0x1000); + ASSERT_EQ(vma_new->vm_end, 0x3000); + + cleanup_mm(&mm, &vmi); + + /* Move backwards, adjacent to old self - self-merge. */ + + vma =3D alloc_and_link_vma(&mm, 0x2000, 0x3000, 2, vma_flags); + vma_set_anonymous(vma); + vma_orig =3D vma; + vma_new =3D copy_vma(&vma, 0x1000, 0x1000, 2, &need_locks); + ASSERT_EQ(vma_new, vma_orig); + ASSERT_EQ(vma, vma_orig); + ASSERT_EQ(vma_new->vm_start, 0x1000); + ASSERT_EQ(vma_new->vm_end, 0x3000); + + cleanup_mm(&mm, &vmi); + + /* + * Move backwards between prior VMA and old self - self-merge and vma + * updated to a new VMA. + */ + + vma_prev =3D alloc_and_link_vma(&mm, 0x1000, 0x2000, 1, vma_flags); + vma_set_anonymous(vma_prev); + vma =3D alloc_and_link_vma(&mm, 0x3000, 0x4000, 3, vma_flags); + vma_set_anonymous(vma); + vma_orig =3D vma; + vma_new =3D copy_vma(&vma, 0x2000, 0x1000, 3, &need_locks); + ASSERT_NE(vma_new, vma_orig); + ASSERT_EQ(vma_new, vma); + ASSERT_EQ(vma_new->vm_start, 0x1000); + ASSERT_EQ(vma_new->vm_end, 0x4000); + + cleanup_mm(&mm, &vmi); =20 /* Move backwards and do not merge. */ =20 --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 64EF33911C0; Thu, 13 Aug 2026 17:36:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642570; cv=none; b=gFHvwsfSzDruFirHTZSpidn3vpxTcZt0aQrjJXAGvnMGqz6OlMis0y4p9WxdqdARL1RXrHLdxC2/VgBlaS/R1Jrk3j3DnSujD4e/5kWVJAc4LLpQNKdIRkOMebfrpBbYsFUurLclDtZyKUuCJkLgi8Ie7CUJv541HA0sDjvJZo0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642570; c=relaxed/simple; bh=AyIYrhzUHJF8dUestntVqmDUqdrez55Qr20v7Ceu/jA=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=ofEjZdxKlo8m9RQOocruJ8pZtGRo0rmjx5d/tfL14f8w8348+NpSZMw7YDCjHsh5C/5jFggvPHr9qTi0lUXhc86l6NGp8wiay/6iAJ4ZuRNtF4AWLW34xJGqjhC2gphaEX25Pp5iIrB3G/U2dfacHz8IDh7zDH74HP1OYqKdvmY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=J6BxMzOm; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="J6BxMzOm" Received: by smtp.kernel.org (Postfix) with ESMTPSA id BA9F91F00A3E; Thu, 13 Aug 2026 17:35:46 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642566; bh=oao9Xk4uxKSjDN12QXQzhenCGYwnNUJY+Ax5MPFwIgA=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=J6BxMzOmmUwaiYkcH1/6YHF2jmW/nz/poIZw0iyFg/zjXbXj01l7U026vA27TJn2S /nMclsWo2ZkqPm3kFLI7IbEGlrZ1d/QCETmq00xq+XTYHsVqnN27MBqpNVOu8EFR/L hBcW2GqWVZ5eNngNPTB6CJlejWsJCKtNqxuU5QyIdDMpTnV3aytV7e8bYsEH4lIy3X LprI3HbfTI9Nf4kG22g6nWlRjVIeoC64X8zI1e43ksK7TEsDTcIcxkqF6u/mWN9oUN tte+a0ltLD1I7E6s1P8w22yyU3W3n45zieWn3bls/mXxpYiMKh/GClzSdQa/qNNRcR zKGC35Xl4kZgQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:26 +0100 Subject: [PATCH v5 09/16] mm: propagate VMA anonymous page offset on map, remap, split + merge Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-9-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=20248; i=ljs@kernel.org; h=from:subject:message-id; bh=AyIYrhzUHJF8dUestntVqmDUqdrez55Qr20v7Ceu/jA=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/69JtWg/HXFu8udTPDqppT2qJsGaTzSLw49tUtduF zp0fNXXjlIWBjEuBlkxRZbnX8T3B4mEzeu84O8GM4eVCWQIAxenAEyEvY3hr9CJCpFNU9hvGKTk ujLO1ftmKRrw4cJk0a781SfXbX476xEjw28VCZe/mZy++dckzn22Ldhe/+Owkb2ldO8Sva5T31P f8wEA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 We must correctly update VMA anonymous page offset state on all VMA operations that would result in it changing, with special attention given to remapping. We cover most cases by simply updating vma_set_range() to do so (with a new anonymous page offset parameter), but also notably must update the merging and mapping logic to propagate this parameter correctly. The remap logic remains the same - we may update the anonymous page offset if the VMA is unfaulted, but now this applies to MAP_PRIVATE file-backed mappings too, so we update the code to reflect this. Note that we use __linear_anon_page_index() upon remap as the VMA may be shared, in order that we update the field consistently regardless of VMA type. Similarly, pass through anon page offset to the merge logic, updating the vma_merge_struct struct to propagate it, and also use __linear_anon_page_index() to obtain the anonymous page index so it can be safely used for both shared and MAP_PRIVATE file-backed mappings. In copy_vma(), the anonymous page offset is updated regardless of whether the mapping is a CoW mapping or not. This is both to keep the anonymous page offset consistent even for non-CoW mappings (it is set so should at least remain correct) and makes the logic cleaner. A self-merge however remains permitted only for mappings which can have a populated vma->anon_vma and do not require alignment on a separate file offset - that is pure anonymous VMAs, so only set can_self_merge if vma_is_anonymous(). Finally, we update insert_vm_struct() to correctly set the anonymous page offset on insertion of a VMA. We simply ensure state is correctly propagated here, so no functional changes are intended. Also update VMA userland tests to reflect this change. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- mm/mremap.c | 6 ++-- mm/vma.c | 52 ++++++++++++++++++---------- mm/vma.h | 75 ++++++++++++++++++++++++-------------= ---- mm/vma_exec.c | 2 +- tools/testing/vma/shared.c | 3 +- tools/testing/vma/tests/merge.c | 4 ++- tools/testing/vma/tests/vma.c | 10 +++--- 7 files changed, 95 insertions(+), 57 deletions(-) diff --git a/mm/mremap.c b/mm/mremap.c index b64aa1f6e07e..9ea1707eafa5 100644 --- a/mm/mremap.c +++ b/mm/mremap.c @@ -1265,7 +1265,9 @@ static void unmap_source_vma(struct vma_remap_struct = *vrm) static int copy_vma_and_data(struct vma_remap_struct *vrm, struct vm_area_struct **new_vma_ptr) { - const unsigned long new_pgoff =3D linear_page_index(vrm->vma, vrm->addr); + const pgoff_t new_pgoff =3D linear_page_index(vrm->vma, vrm->addr); + const pgoff_t new_anon_pgoff =3D + __linear_anon_page_index(vrm->vma, vrm->addr); struct vm_area_struct *vma =3D vrm->vma; struct vm_area_struct *new_vma; unsigned long moved_len; @@ -1273,7 +1275,7 @@ static int copy_vma_and_data(struct vma_remap_struct = *vrm, PAGETABLE_MOVE(pmc, NULL, NULL, vrm->addr, vrm->new_addr, vrm->old_len); =20 new_vma =3D copy_vma(&vma, vrm->new_addr, vrm->new_len, new_pgoff, - &pmc.need_rmap_locks); + new_anon_pgoff, &pmc.need_rmap_locks); if (!new_vma) { vrm_uncharge(vrm); *new_vma_ptr =3D NULL; diff --git a/mm/vma.c b/mm/vma.c index 18c8f2546765..b55015952f0d 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -18,6 +18,7 @@ struct mmap_state { unsigned long addr; unsigned long end; pgoff_t pgoff; + pgoff_t anon_pgoff; unsigned long pglen; union { vm_flags_t vm_flags; @@ -46,13 +47,14 @@ struct mmap_state { bool file_doesnt_need_get :1; }; =20 -#define MMAP_STATE(name, mm_, vmi_, addr_, len_, pgoff_, vma_flags_, file_= ) \ +#define MMAP_STATE(name, mm_, vmi_, addr_, len_, pgoff_, anon_pgoff_, vma_= flags_, file_) \ struct mmap_state name =3D { \ .mm =3D mm_, \ .vmi =3D vmi_, \ .addr =3D addr_, \ .end =3D (addr_) + (len_), \ .pgoff =3D pgoff_, \ + .anon_pgoff =3D anon_pgoff_, \ .pglen =3D PHYS_PFN(len_), \ .vma_flags =3D vma_flags_, \ .file =3D file_, \ @@ -67,6 +69,7 @@ struct mmap_state { .end =3D (map_)->end, \ .vma_flags =3D (map_)->vma_flags, \ .pgoff =3D (map_)->pgoff, \ + .anon_pgoff =3D (map_)->anon_pgoff, \ .file =3D (map_)->file, \ .prev =3D (map_)->prev, \ .middle =3D vma_, \ @@ -82,10 +85,11 @@ static void __vma_set_range(struct vm_area_struct *vma,= unsigned long start, } =20 static void vma_set_range(struct vm_area_struct *vma, unsigned long start, - unsigned long end, pgoff_t pgoff) + unsigned long end, pgoff_t pgoff, pgoff_t anon_pgoff) { __vma_set_range(vma, start, end); vma_set_pgoff(vma, pgoff); + vma_set_anon_pgoff(vma, anon_pgoff); } =20 /* Was this VMA ever forked from a parent, i.e. maybe contains CoW mapping= s? */ @@ -812,7 +816,8 @@ static int commit_merge(struct vma_merge_struct *vmg) */ vma_adjust_trans_huge(vma, vmg->start, vmg->end, vmg->__adjust_middle_start ? vmg->middle : NULL); - vma_set_range(vma, vmg->start, vmg->end, vmg_start_pgoff(vmg)); + vma_set_range(vma, vmg->start, vmg->end, vmg_start_pgoff(vmg), + vmg_start_anon_pgoff(vmg)); vmg_adjust_set_range(vmg); vma_iter_store_overwrite(vmg->vmi, vmg->target); =20 @@ -982,6 +987,7 @@ static __must_check struct vm_area_struct *vma_merge_ex= isting_range( vmg->start =3D prev->vm_start; vmg->end =3D next->vm_end; vmg->pgoff =3D vma_start_pgoff(prev); + vmg->anon_pgoff =3D vma_start_anon_pgoff(prev); =20 /* * We already ensured anon_vma compatibility above, so now it's @@ -1000,6 +1006,7 @@ static __must_check struct vm_area_struct *vma_merge_= existing_range( */ vmg->start =3D prev->vm_start; vmg->pgoff =3D vma_start_pgoff(prev); + vmg->anon_pgoff =3D vma_start_anon_pgoff(prev); =20 if (!vmg->__remove_middle) vmg->__adjust_middle_start =3D true; @@ -1022,12 +1029,14 @@ static __must_check struct vm_area_struct *vma_merg= e_existing_range( if (vmg->__remove_middle) { vmg->end =3D next->vm_end; vmg->pgoff =3D vma_start_pgoff(next) - pglen; + vmg->anon_pgoff =3D vma_start_anon_pgoff(next) - pglen; } else { /* We shrink middle and expand next. */ vmg->__adjust_next_start =3D true; vmg->start =3D middle->vm_start; vmg->end =3D start; vmg->pgoff =3D vma_start_pgoff(middle); + vmg->anon_pgoff =3D vma_start_anon_pgoff(middle); } =20 err =3D dup_anon_vma(next, middle, &anon_dup); @@ -1137,6 +1146,7 @@ struct vm_area_struct *vma_merge_new_range(struct vma= _merge_struct *vmg) vmg->start =3D prev->vm_start; vmg->target =3D prev; vmg->pgoff =3D vma_start_pgoff(prev); + vmg->anon_pgoff =3D vma_start_anon_pgoff(prev); =20 /* * If this merge would result in removal of the next VMA but we @@ -1908,7 +1918,7 @@ static int vma_link(struct mm_struct *mm, struct vm_a= rea_struct *vma) */ struct vm_area_struct *copy_vma(struct vm_area_struct **vmap, unsigned long addr, unsigned long len, pgoff_t pgoff, - bool *need_rmap_locks) + pgoff_t anon_pgoff, bool *need_rmap_locks) { struct vm_area_struct *vma =3D *vmap; unsigned long old_vma_start =3D vma->vm_start; @@ -1919,12 +1929,16 @@ struct vm_area_struct *copy_vma(struct vm_area_stru= ct **vmap, VMG_VMA_STATE(vmg, &vmi, NULL, vma, addr, addr + len); =20 /* - * If anonymous vma has not yet been faulted, update new pgoff - * to match new location, to increase its chance of merging. + * If a vma has not yet been faulted, update its anonymous pgoff to + * match the new location to increase its chance of merging. */ - if (unlikely(vma_is_anonymous(vma) && !vma->anon_vma)) { - pgoff =3D addr >> PAGE_SHIFT; - can_self_merge =3D true; + if (!vma->anon_vma) { + anon_pgoff =3D addr >> PAGE_SHIFT; + + if (vma_is_anonymous(vma)) { + pgoff =3D anon_pgoff; + can_self_merge =3D true; + } } =20 /* @@ -1940,6 +1954,7 @@ struct vm_area_struct *copy_vma(struct vm_area_struct= **vmap, return NULL; /* should never get here */ =20 vmg.pgoff =3D pgoff; + vmg.anon_pgoff =3D anon_pgoff; vmg.next =3D vma_iter_next_rewind(&vmi, NULL); new_vma =3D vma_merge_copied_range(&vmg); =20 @@ -1955,8 +1970,8 @@ struct vm_area_struct *copy_vma(struct vm_area_struct= **vmap, * next to be removed and prev to be expanded to cover * the entire range. * - * This should only be possible if the page offset was - * updated, i.e. the VMA is unfaulted. + * This should only be possible if the anonymous page + * offset was updated, i.e. the VMA is unfaulted. */ VM_WARN_ON_ONCE_VMA(!can_self_merge, new_vma); *vmap =3D vma =3D new_vma; @@ -1967,7 +1982,7 @@ struct vm_area_struct *copy_vma(struct vm_area_struct= **vmap, new_vma =3D vm_area_dup(vma); if (!new_vma) goto out; - vma_set_range(new_vma, addr, addr + len, pgoff); + vma_set_range(new_vma, addr, addr + len, pgoff, anon_pgoff); if (vma_dup_policy(vma, new_vma)) goto out_free_vma; if (anon_vma_clone(new_vma, vma, VMA_OP_REMAP)) @@ -2609,7 +2624,7 @@ static int __mmap_new_vma(struct mmap_state *map, str= uct vm_area_struct **vmap, if (is_anon) vma_set_anonymous(vma); =20 - vma_set_range(vma, map->addr, map->end, map->pgoff); + vma_set_range(vma, map->addr, map->end, map->pgoff, map->anon_pgoff); vma->flags =3D map->vma_flags; vma->vm_page_prot =3D map->page_prot; =20 @@ -2798,7 +2813,8 @@ static unsigned long __mmap_region(struct file *file,= unsigned long addr, struct vm_area_struct *vma =3D NULL; bool have_mmap_prepare =3D file && file->f_op->mmap_prepare; VMA_ITERATOR(vmi, mm, addr); - MMAP_STATE(map, mm, &vmi, addr, len, pgoff, vma_flags, file); + const pgoff_t anon_pgoff =3D addr >> PAGE_SHIFT; + MMAP_STATE(map, mm, &vmi, addr, len, pgoff, anon_pgoff, vma_flags, file); struct vm_area_desc desc =3D { .mm =3D mm, .file =3D file, @@ -2943,6 +2959,7 @@ int do_brk_flags(struct vma_iterator *vmi, struct vm_= area_struct *vma, unsigned long addr, unsigned long len, vma_flags_t vma_flags) { struct mm_struct *mm =3D current->mm; + const pgoff_t pgoff =3D addr >> PAGE_SHIFT; =20 /* * Check against address space limits by the changed size @@ -2967,7 +2984,7 @@ int do_brk_flags(struct vma_iterator *vmi, struct vm_= area_struct *vma, * occur after forking, so the expand will only happen on new VMAs. */ if (vma && vma->vm_end =3D=3D addr) { - VMG_STATE(vmg, mm, vmi, addr, addr + len, vma_flags, PHYS_PFN(addr)); + VMG_STATE(vmg, mm, vmi, addr, addr + len, vma_flags, pgoff, pgoff); =20 vmg.prev =3D vma; /* vmi is positioned at prev, which this mode expects. */ @@ -2987,7 +3004,7 @@ int do_brk_flags(struct vma_iterator *vmi, struct vm_= area_struct *vma, goto unacct_fail; =20 vma_set_anonymous(vma); - vma_set_range(vma, addr, addr + len, addr >> PAGE_SHIFT); + vma_set_range(vma, addr, addr + len, pgoff, pgoff); vma->flags =3D vma_flags; vma->vm_page_prot =3D vm_get_page_prot(vma_flags_to_legacy(vma_flags)); vma_start_write(vma); @@ -3379,6 +3396,7 @@ int insert_vm_struct(struct mm_struct *mm, struct vm_= area_struct *vma) WARN_ON_ONCE(vma->anon_vma); vma_set_pgoff(vma, vma->vm_start >> PAGE_SHIFT); } + vma_set_anon_pgoff(vma, vma->vm_start >> PAGE_SHIFT); =20 if (vma_link(mm, vma)) { if (vma_test(vma, VMA_ACCOUNT_BIT)) @@ -3434,7 +3452,7 @@ struct vm_area_struct *__install_special_mapping( =20 vma->vm_ops =3D ops; vma->vm_private_data =3D priv; - vma_set_range(vma, addr, addr + len, 0); + vma_set_range(vma, addr, addr + len, 0, addr >> PAGE_SHIFT); =20 ret =3D insert_vm_struct(mm, vma); if (ret) diff --git a/mm/vma.h b/mm/vma.h index 54ed7c744e3b..024fabe63560 100644 --- a/mm/vma.h +++ b/mm/vma.h @@ -104,6 +104,7 @@ struct vma_merge_struct { unsigned long start; unsigned long end; pgoff_t pgoff; + pgoff_t anon_pgoff; =20 union { /* Temporary while VMA flags are being converted. */ @@ -237,11 +238,6 @@ static inline bool vmg_nomem(struct vma_merge_struct *= vmg) return vmg->state =3D=3D VMA_MERGE_ERROR_NOMEM; } =20 -static inline pgoff_t vmg_start_pgoff(const struct vma_merge_struct *vmg) -{ - return vmg->pgoff; -} - static inline pgoff_t vmg_pages(const struct vma_merge_struct *vmg) { const unsigned long size =3D vmg->end - vmg->start; @@ -249,6 +245,11 @@ static inline pgoff_t vmg_pages(const struct vma_merge= _struct *vmg) return size >> PAGE_SHIFT; } =20 +static inline pgoff_t vmg_start_pgoff(const struct vma_merge_struct *vmg) +{ + return vmg->pgoff; +} + static inline pgoff_t vmg_end_pgoff(const struct vma_merge_struct *vmg) { return vmg_start_pgoff(vmg) + vmg_pages(vmg); @@ -283,6 +284,16 @@ static inline void vma_set_pgoff(struct vm_area_struct= *vma, pgoff_t pgoff) vma->vm_pgoff =3D pgoff; } =20 +static inline pgoff_t vmg_start_anon_pgoff(const struct vma_merge_struct *= vmg) +{ + return vmg->anon_pgoff; +} + +static inline pgoff_t vmg_end_anon_pgoff(const struct vma_merge_struct *vm= g) +{ + return vmg_start_anon_pgoff(vmg) + vmg_pages(vmg); +} + static inline void __vma_set_anon_pgoff(struct vm_area_struct *vma, pgoff_= t pgoff) { #ifdef CONFIG_64BIT @@ -301,44 +312,48 @@ static inline void vma_add_pgoff(struct vm_area_struc= t *vma, pgoff_t delta) { vma_assert_can_modify(vma); vma_set_pgoff(vma, vma_start_pgoff(vma) + delta); + vma_set_anon_pgoff(vma, vma_start_anon_pgoff(vma) + delta); } =20 static inline void vma_sub_pgoff(struct vm_area_struct *vma, pgoff_t delta) { vma_assert_can_modify(vma); vma_set_pgoff(vma, vma_start_pgoff(vma) - delta); -} + vma_set_anon_pgoff(vma, vma_start_anon_pgoff(vma) - delta); +} + +#define VMG_STATE(name, mm_, vmi_, start_, end_, vma_flags_, pgoff_, anon_= pgoff_) \ + struct vma_merge_struct name =3D { \ + .mm =3D mm_, \ + .vmi =3D vmi_, \ + .start =3D start_, \ + .end =3D end_, \ + .vma_flags =3D vma_flags_, \ + .pgoff =3D pgoff_, \ + .anon_pgoff =3D anon_pgoff_, \ + .state =3D VMA_MERGE_START, \ + } =20 -#define VMG_STATE(name, mm_, vmi_, start_, end_, vma_flags_, pgoff_) \ +#define VMG_VMA_STATE(name, vmi_, prev_, vma_, start_, end_) \ struct vma_merge_struct name =3D { \ - .mm =3D mm_, \ + .mm =3D vma_->vm_mm, \ .vmi =3D vmi_, \ + .prev =3D prev_, \ + .middle =3D vma_, \ + .next =3D NULL, \ .start =3D start_, \ .end =3D end_, \ - .vma_flags =3D vma_flags_, \ - .pgoff =3D pgoff_, \ + .vm_flags =3D vma_->vm_flags, \ + .pgoff =3D linear_page_index(vma_, start_), \ + .anon_pgoff =3D __linear_anon_page_index(vma_, start_), \ + .file =3D vma_->vm_file, \ + .anon_vma =3D vma_->anon_vma, \ + .policy =3D vma_policy(vma_), \ + .uffd_ctx =3D vma_->vm_userfaultfd_ctx, \ + .anon_name =3D anon_vma_name(vma_), \ .state =3D VMA_MERGE_START, \ } =20 -#define VMG_VMA_STATE(name, vmi_, prev_, vma_, start_, end_) \ - struct vma_merge_struct name =3D { \ - .mm =3D vma_->vm_mm, \ - .vmi =3D vmi_, \ - .prev =3D prev_, \ - .middle =3D vma_, \ - .next =3D NULL, \ - .start =3D start_, \ - .end =3D end_, \ - .vm_flags =3D vma_->vm_flags, \ - .pgoff =3D linear_page_index(vma_, start_), \ - .file =3D vma_->vm_file, \ - .anon_vma =3D vma_->anon_vma, \ - .policy =3D vma_policy(vma_), \ - .uffd_ctx =3D vma_->vm_userfaultfd_ctx, \ - .anon_name =3D anon_vma_name(vma_), \ - .state =3D VMA_MERGE_START, \ - } - #ifdef CONFIG_DEBUG_VM_MAPLE_TREE void validate_mm(struct mm_struct *mm); #else @@ -520,7 +535,7 @@ void unlink_file_vma_batch_add(struct unlink_vma_file_b= atch *vb, =20 struct vm_area_struct *copy_vma(struct vm_area_struct **vmap, unsigned long addr, unsigned long len, pgoff_t pgoff, - bool *need_rmap_locks); + pgoff_t anon_pgoff, bool *need_rmap_locks); =20 struct anon_vma *find_mergeable_anon_vma(struct vm_area_struct *vma); =20 diff --git a/mm/vma_exec.c b/mm/vma_exec.c index 7af1260689b9..586c52155942 100644 --- a/mm/vma_exec.c +++ b/mm/vma_exec.c @@ -41,7 +41,7 @@ int relocate_vma_down(struct vm_area_struct *vma, unsigne= d long shift) unsigned long new_end =3D old_end - shift; VMA_ITERATOR(vmi, mm, new_start); VMG_STATE(vmg, mm, &vmi, new_start, old_end, EMPTY_VMA_FLAGS, - vma_start_pgoff(vma)); + vma_start_pgoff(vma), vma_start_anon_pgoff(vma)); struct vm_area_struct *next; struct mmu_gather tlb; PAGETABLE_MOVE(pmc, vma, vma, old_start, new_start, length); diff --git a/tools/testing/vma/shared.c b/tools/testing/vma/shared.c index bea9ea6db02a..4a39c9d50489 100644 --- a/tools/testing/vma/shared.c +++ b/tools/testing/vma/shared.c @@ -23,7 +23,8 @@ struct vm_area_struct *alloc_vma(struct mm_struct *mm, =20 vma->vm_start =3D start; vma->vm_end =3D end; - vma->vm_pgoff =3D pgoff; + vma_set_pgoff(vma, pgoff); + vma_set_anon_pgoff(vma, start >> PAGE_SHIFT); vma->flags =3D vma_flags; vma_assert_detached(vma); =20 diff --git a/tools/testing/vma/tests/merge.c b/tools/testing/vma/tests/merg= e.c index e357accc8499..48418b82b01d 100644 --- a/tools/testing/vma/tests/merge.c +++ b/tools/testing/vma/tests/merge.c @@ -45,6 +45,7 @@ void vmg_set_range(struct vma_merge_struct *vmg, unsigned= long start, vmg->start =3D start; vmg->end =3D end; vmg->pgoff =3D pgoff; + vmg->anon_pgoff =3D start >> PAGE_SHIFT; vmg->vma_flags =3D vma_flags; =20 vmg->just_expand =3D false; @@ -108,6 +109,7 @@ static bool test_simple_merge(void) .end =3D 0x2000, .vma_flags =3D vma_flags, .pgoff =3D 1, + .anon_pgoff =3D 1, }; =20 ASSERT_FALSE(attach_vma(&mm, vma_left)); @@ -1431,7 +1433,7 @@ static bool test_expand_only_mode(void) struct mm_struct mm =3D {}; VMA_ITERATOR(vmi, &mm, 0); struct vm_area_struct *vma_prev, *vma; - VMG_STATE(vmg, &mm, &vmi, 0x5000, 0x9000, vma_flags, 5); + VMG_STATE(vmg, &mm, &vmi, 0x5000, 0x9000, vma_flags, 5, 5); =20 /* * Place a VMA prior to the one we're expanding so we assert that we do diff --git a/tools/testing/vma/tests/vma.c b/tools/testing/vma/tests/vma.c index 0d40d7ba2181..c8ef7b8cd46b 100644 --- a/tools/testing/vma/tests/vma.c +++ b/tools/testing/vma/tests/vma.c @@ -40,7 +40,7 @@ static bool test_copy_vma(void) vma =3D alloc_and_link_vma(&mm, 0x1000, 0x2000, 1, vma_flags); vma_set_anonymous(vma); vma_orig =3D vma; - vma_new =3D copy_vma(&vma, 0x2000, 0x1000, 1, &need_locks); + vma_new =3D copy_vma(&vma, 0x2000, 0x1000, 1, 1, &need_locks); ASSERT_EQ(vma_new, vma_orig); ASSERT_EQ(vma, vma_orig); ASSERT_EQ(vma_new->vm_start, 0x1000); @@ -53,7 +53,7 @@ static bool test_copy_vma(void) vma =3D alloc_and_link_vma(&mm, 0x2000, 0x3000, 2, vma_flags); vma_set_anonymous(vma); vma_orig =3D vma; - vma_new =3D copy_vma(&vma, 0x1000, 0x1000, 2, &need_locks); + vma_new =3D copy_vma(&vma, 0x1000, 0x1000, 2, 2, &need_locks); ASSERT_EQ(vma_new, vma_orig); ASSERT_EQ(vma, vma_orig); ASSERT_EQ(vma_new->vm_start, 0x1000); @@ -71,7 +71,7 @@ static bool test_copy_vma(void) vma =3D alloc_and_link_vma(&mm, 0x3000, 0x4000, 3, vma_flags); vma_set_anonymous(vma); vma_orig =3D vma; - vma_new =3D copy_vma(&vma, 0x2000, 0x1000, 3, &need_locks); + vma_new =3D copy_vma(&vma, 0x2000, 0x1000, 3, 3, &need_locks); ASSERT_NE(vma_new, vma_orig); ASSERT_EQ(vma_new, vma); ASSERT_EQ(vma_new->vm_start, 0x1000); @@ -82,7 +82,7 @@ static bool test_copy_vma(void) /* Move backwards and do not merge. */ =20 vma =3D alloc_and_link_vma(&mm, 0x3000, 0x5000, 3, vma_flags); - vma_new =3D copy_vma(&vma, 0, 0x2000, 0, &need_locks); + vma_new =3D copy_vma(&vma, 0, 0x2000, 0, 3, &need_locks); ASSERT_NE(vma_new, vma); ASSERT_EQ(vma_new->vm_start, 0); ASSERT_EQ(vma_new->vm_end, 0x2000); @@ -95,7 +95,7 @@ static bool test_copy_vma(void) =20 vma =3D alloc_and_link_vma(&mm, 0, 0x2000, 0, vma_flags); vma_next =3D alloc_and_link_vma(&mm, 0x6000, 0x8000, 6, vma_flags); - vma_new =3D copy_vma(&vma, 0x4000, 0x2000, 4, &need_locks); + vma_new =3D copy_vma(&vma, 0x4000, 0x2000, 4, 4, &need_locks); vma_assert_attached(vma_new); =20 ASSERT_EQ(vma_new, vma_next); --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 090DF2DF6F4; Thu, 13 Aug 2026 17:36:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642588; cv=none; b=Y/3a1ePw1FS1wco7EKVjBudqSY6UMHygBm9KWtyB8b9Az/5+xhAavslwITWkF0kKsP5xG5Q7VBNtMxSAw3XNXhajxjI6xUy2Pg8SazyiFsUGj9d2z0u/fSEY0T7kIcO1RptSNkdnMfwtFZ4gY7suB2ML40xnW26GJyK4MsNfexc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642588; c=relaxed/simple; bh=795PEltcwnxXCtrgqOcEpUiQ44/tIXKOdzD9y1SPcjQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=tq1310RMMSDuOVoT2SrrEcVDxUKys1yR5/3WagSVLkAIgFqDw+LhJwf6depsSmi9JuysShLnbZGqcst2e+42O5fuz3IV63Aju+14UBRuboxSSL0OU7eBHYb7k9WAHd8Cu/Ve4xSNxTHSCH+/FZSDL4z3kzGBfRQL+P43ZM3WyJw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=QZ6OwOIV; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="QZ6OwOIV" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 171001F000E9; Thu, 13 Aug 2026 17:36:06 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642586; bh=p0ckbLWugOOUGCY6PMxy8vsH8dtr8rQdR8d8zrRRsVc=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=QZ6OwOIV6rnPZ7gDW//HB57G0jreHQ3d0pRsOiPKbI/u0+vpItGyz2BKfdF3roLos jLuBLuws6qd9Xepg6sAzESfGaaDZh8sIkysFXu9qEOXl8VIGq0pX08wNsSl1tXnzgM RzW0q2FCKrXi3FF7ErrAUvRLoSVu78UIUvBDqgzuQKHW5hKBmaZmqZH4D4Oq6OxFFo LAGKNBQevXYf6YqhwSznZ63wKUSLpV6AFeVyjkoK9HbQ8IklLOCoL4JOQcBfoBhmfd JnyG5EGLVTook5lrh3MF/eEEeaYKOwqwnVGKsDq5le5qtJkrr8yr5z9SbTDvvL50wm W1NraZ6QVjotw== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:27 +0100 Subject: [PATCH v5 10/16] mm/rmap: track whether the page VMA mapped pgoff is anonymous Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-10-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2340; i=ljs@kernel.org; h=from:subject:message-id; bh=795PEltcwnxXCtrgqOcEpUiQ44/tIXKOdzD9y1SPcjQ=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6+5yFEtVfRauCfxzfpoC4mgIKvnczfxaAtwtuW16 +1QW/62o5SFQYyLQVZMkeX5F/H9QSJh8zov+LvBzGFlAhnCwMUpABPR+M3I8OTS7YWmL358cbo8 d1nro1b2P8+XTl0ft0VdYHOJS9fTlD6GPxwpba8V47wMT5lzv5eqvF4rN/Pnz2CxZOcM95Tlehn lnAA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Update the page_vma_mapped_walk structure to track whether the pgoff being tracked is an anonymous pgoff or not and update the comments to reflect this. This is necessary in order to determine the correct VMA page offset in vma_address_end() when pvmw->nr_pages > 1. Also document that pvmw->pgoff is meaningless for pvmw->nr_pages =3D=3D 1 a= nd for KSM. Do not set this field where pgoff is not specified. This is laying the groundwork for eventually using anonymous page offsets as the index for all anonymous folios. No functional change intended. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/rmap.h | 4 +++- mm/rmap.c | 2 ++ 2 files changed, 5 insertions(+), 1 deletion(-) diff --git a/include/linux/rmap.h b/include/linux/rmap.h index 8dc0871e5f00..0574537a355c 100644 --- a/include/linux/rmap.h +++ b/include/linux/rmap.h @@ -864,13 +864,14 @@ struct page *make_device_exclusive(struct mm_struct *= mm, unsigned long addr, struct page_vma_mapped_walk { unsigned long pfn; unsigned long nr_pages; - pgoff_t pgoff; + pgoff_t pgoff; /* Only meaningful if nr_pages > 1 and not a KSM walk */ struct vm_area_struct *vma; unsigned long address; pmd_t *pmd; pte_t *pte; spinlock_t *ptl; unsigned int flags; + bool pgoff_is_anon : 1; }; =20 #define DEFINE_FOLIO_VMA_WALK(name, _folio, _vma, _address, _flags) \ @@ -881,6 +882,7 @@ struct page_vma_mapped_walk { .vma =3D _vma, \ .address =3D _address, \ .flags =3D _flags, \ + .pgoff_is_anon =3D folio_test_anon(_folio), \ } =20 static inline void page_vma_mapped_walk_done(struct page_vma_mapped_walk *= pvmw) diff --git a/mm/rmap.c b/mm/rmap.c index 5798427d007f..ab3f879454de 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -1240,6 +1240,7 @@ static bool mapping_wrprotect_range_one(struct folio = *folio, .vma =3D vma, .address =3D address, .flags =3D PVMW_SYNC, + .pgoff_is_anon =3D false, }; =20 state->cleaned +=3D page_vma_mkclean_one(&pvmw); @@ -1317,6 +1318,7 @@ int pfn_mkclean_range(unsigned long pfn, unsigned lon= g nr_pages, pgoff_t pgoff, .pgoff =3D pgoff, .vma =3D vma, .flags =3D PVMW_SYNC, + .pgoff_is_anon =3D false, }; =20 if (invalid_mkclean_vma(vma, NULL)) --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9048C395ACD; Thu, 13 Aug 2026 17:36:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642608; cv=none; b=LQhMl8GjbhqGKJkzIG11RfG9Zhxs39iyJnWUMBfnF3QxOA6aV0DAilug7mrw4YA3WiBiqjwy2GW35NERi7JzVoT9XccWWVc5WY3Y7G3jx3piQB20NWErzLRl/7yyE13U0svMIFxsABi+5HI6vtwMZeWQTafk8iNfA9pgkkXXRbY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642608; c=relaxed/simple; bh=rEAXpM+Q48loiQvQcclMM0xWbuZorGJAZULvM46znVw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=A4lDzsDDFYx6riJ9k80fSfUYK4gnq9UMt6Bhz1DjX/j4PMj9dPcqZHeXdm6Gam3EKN2QSa+qqZtQjJ83wog9Dm08z1Dimn6J3M7UV83Xkv7UpJ3xA0tvpva9qZaTrk7NS9ZcpgJqpn5ecauXVs8BizAXbrwE4IkmLLG5F93e8ho= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Rp9Z/BjE; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Rp9Z/BjE" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 631C31F00A3A; Thu, 13 Aug 2026 17:36:27 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642607; bh=5EXc5iOEg9H0d0e1HfYbF7zEjYZ4oBYHKpBUsUiBnw0=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=Rp9Z/BjEPxpYUoyxaFds6p1oBZAw1SzuHFYiH/cSF6KJkjwec8ieKG8PtE6ebH0Oh XbFmJ+JOri7Q0P5RfkiRIkg7AlWnDHfYrpxRm1e8iXTFNx4Nsul77z22qVOlnUkQ8a uPGrHtUBs6LHOtSIC+tCPR1DG11boYqayR16QQnlpYjEnpvl725Lg28dz7gZnVluGl BfV1JcBTm32fY9GvnTvAN9vQUPJRs9g7qfW/StEW0gC4aLs9Dg+IB3GrdoCpvzyOJp 31rLMMpzLNA+nX43lEvPi/d0gs6sRAs2Gxsb4Necef7josS+e4lxADOFVZ5w8/XQ7t jj+g1IHiSK5iA== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:28 +0100 Subject: [PATCH v5 11/16] mm: clean up vma_address_end() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-11-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2040; i=ljs@kernel.org; h=from:subject:message-id; bh=rEAXpM+Q48loiQvQcclMM0xWbuZorGJAZULvM46znVw=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6/RMJi5nSPlvT53x+qFS++Jt00McDb5oxCeU6vV1 Vx+fFNIRykLgxgXg6yYIsvzL+L7g0TC5nVe8HeDmcPKBDKEgYtTACayuorhn/YPuR0iaxKVZ/XX 1h39v2nu6qlREnOOcTZsuvVzcdRxIR+Gv1JinWoLribkBHpV5ngUvzrE2Z0eo+d27syP3CUHJOb x8wEA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 vma_address_end() is a confusing function with a lot of moving parts so clean it up prior to extending it for anon page indexed mappings. Const-ify some variables and establish pgoff_vma_start and pgoff_end variables to clearly identify the page offset for the start of the VMA and the end of the page offset range specified by the page walk. This simplifies the function significantly and lays the groundwork for a future change to update this function to account for anonymously page indexed folios. No functional change intended. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- mm/internal.h | 12 +++++++----- 1 file changed, 7 insertions(+), 5 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index 03145b8d0d56..98b067472ceb 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1079,22 +1079,24 @@ static inline unsigned long vma_anon_address(const = struct vm_area_struct *vma, } =20 /* - * Then at what user virtual address will none of the range be found in vm= a? + * At what user virtual address will none of the range be found in vma? * Assumes that vma_address() already returned a good starting address. */ static inline unsigned long vma_address_end(struct page_vma_mapped_walk *p= vmw) { - struct vm_area_struct *vma =3D pvmw->vma; - pgoff_t pgoff; + const pgoff_t pgoff_end =3D pvmw->pgoff + pvmw->nr_pages; + const struct vm_area_struct *vma =3D pvmw->vma; + pgoff_t pgoff_vma_start; unsigned long address; =20 /* Common case, plus ->pgoff is invalid for KSM */ if (pvmw->nr_pages =3D=3D 1) return pvmw->address + PAGE_SIZE; =20 - pgoff =3D pvmw->pgoff + pvmw->nr_pages; + pgoff_vma_start =3D vma_start_pgoff(vma); + address =3D vma->vm_start + - ((pgoff - vma_start_pgoff(vma)) << PAGE_SHIFT); + ((pgoff_end - pgoff_vma_start) << PAGE_SHIFT); /* Check for address beyond vma (or wrapped through 0?) */ if (address < vma->vm_start || address > vma->vm_end) address =3D vma->vm_end; --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0F68038758C; Thu, 13 Aug 2026 17:37:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642629; cv=none; b=j/nKCyeBSxlj1dphqT7an8AnTM/g4dkOzmZZo57WFgHZko6CYXJpBCvw+KH5Yajh8K6erEfaDgNa4C+yvK4qQ9km1JTEZtc4/JCrrpdH+OvNPj02VyGeEI2s+O9mhajd2k5q8MvPNvC0DDkD/QnNL+zcsICVNV7hbAT5hj2iFq4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642629; c=relaxed/simple; bh=q0oMVyXXBC+siqfel45AxXFQKXhgvOdgu44yz/skNq4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Y1ifBLvyqY9luVpKTKCcBP5wHf1lF4ozGvdKc7/YuPG/6hC+aakaRQT/hVm3GhYJoTlns3viUlJE2qaHYlsEjjRldj3zDZ61RRwihIzkZQivWiFbLERkMKhyOajkJwh1LFwGpgRXUDVGG5liD3nywnO96M1mCd1sApXwW1WkCo8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ixA7p747; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ixA7p747" Received: by smtp.kernel.org (Postfix) with ESMTPSA id B4A411F00A3E; Thu, 13 Aug 2026 17:36:47 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642627; bh=vaw3s6T/GSWhGIBxH9c44M6soIaJSA08keLAt3jySb4=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=ixA7p747hccwgImJ59EcIdFEdSb9rc3LOw5kVbJpKoEo+0MrDfKmMlBaTs4ef7wRp Y09zqNUCXL/Q1nXEnswtP39scxNB9QBcvSpkUWXqiJes3uks/0pES6sNqGYVl9zfIH q1UdbZqJfKD2AjguBvWW/DtZR+cJY2l/Oa3RXrVhsaKxiZXggZapMoIY6l1WbpapBK soW+xVWeTSY/FLuyemB/fTPfHP/6sgSHr5/1/K6gAGYMJccN3jC4vFXgCNgBT8mAwc TAhftfuvaFfkEVyxyj0At3cGYA9N/gt8Q0gNTrFfMACpD2FzWDj1uimQe82B7+ml91 7mrNMhuTrg+7w== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:29 +0100 Subject: [PATCH v5 12/16] mm/huge_memory: update remove_migration_pmd() to accept a folio Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-12-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3818; i=ljs@kernel.org; h=from:subject:message-id; bh=q0oMVyXXBC+siqfel45AxXFQKXhgvOdgu44yz/skNq4=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6/5F1G55rl8yiLzKSurYtcednm68v2b5LfO17iun E6r5tWz6ShlYRDjYpAVU2R5/kV8f5BI2LzOC/5uMHNYmUCGMHBxCsBESq4x/E959/lG4ZRWo3vN /Px/yn09V2vV3/rQ2563THRvA/eG/t8Mf8VmWS7u+1xZK/v/296rxqzNa68cNXc6UXH325RilYd 3s7kB X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This function does not need to accept a page and requiring it to is unnecessary and misleading. make_[writable, readable]_device_private_entry() must be passed a PMD-aligned PFN as they immediately used to obtain a softleaf PMD entry and the same argument applies to folio_add_[anon, file]_rmap_pmd(). While we are here, update a VM_BUG_ON() to a VM_WARN_ON_ONCE(). No functional change intended. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/swapops.h | 6 +++--- mm/huge_memory.c | 16 +++++++--------- mm/migrate.c | 2 +- 3 files changed, 11 insertions(+), 13 deletions(-) diff --git a/include/linux/swapops.h b/include/linux/swapops.h index c956bc445ee0..1f3ff3b93e16 100644 --- a/include/linux/swapops.h +++ b/include/linux/swapops.h @@ -325,8 +325,8 @@ struct page_vma_mapped_walk; extern int set_pmd_migration_entry(struct page_vma_mapped_walk *pvmw, struct page *page); =20 -extern void remove_migration_pmd(struct page_vma_mapped_walk *pvmw, - struct page *new); +void remove_migration_pmd(struct page_vma_mapped_walk *pvmw, + struct folio *folio); =20 extern void pmd_migration_entry_wait(struct mm_struct *mm, pmd_t *pmd); =20 @@ -346,7 +346,7 @@ static inline int set_pmd_migration_entry(struct page_v= ma_mapped_walk *pvmw, } =20 static inline void remove_migration_pmd(struct page_vma_mapped_walk *pvmw, - struct page *new) + struct folio *folio) { BUILD_BUG(); } diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 6b0cabd45b2d..47c9c6e32eba 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -5020,9 +5020,8 @@ int set_pmd_migration_entry(struct page_vma_mapped_wa= lk *pvmw, return 0; } =20 -void remove_migration_pmd(struct page_vma_mapped_walk *pvmw, struct page *= new) +void remove_migration_pmd(struct page_vma_mapped_walk *pvmw, struct folio = *folio) { - struct folio *folio =3D page_folio(new); struct vm_area_struct *vma =3D pvmw->vma; struct mm_struct *mm =3D vma->vm_mm; unsigned long address =3D pvmw->address; @@ -5058,11 +5057,9 @@ void remove_migration_pmd(struct page_vma_mapped_wal= k *pvmw, struct page *new) swp_entry_t entry; =20 if (pmd_write(pmde)) - entry =3D make_writable_device_private_entry( - page_to_pfn(new)); + entry =3D make_writable_device_private_entry(folio_pfn(folio)); else - entry =3D make_readable_device_private_entry( - page_to_pfn(new)); + entry =3D make_readable_device_private_entry(folio_pfn(folio)); pmde =3D softleaf_to_pmd(entry); =20 if (pmd_swp_soft_dirty(*pvmw->pmd)) @@ -5077,11 +5074,12 @@ void remove_migration_pmd(struct page_vma_mapped_wa= lk *pvmw, struct page *new) if (!softleaf_is_migration_read(entry)) rmap_flags |=3D RMAP_EXCLUSIVE; =20 - folio_add_anon_rmap_pmd(folio, new, vma, haddr, rmap_flags); + folio_add_anon_rmap_pmd(folio, &folio->page, vma, haddr, rmap_flags); } else { - folio_add_file_rmap_pmd(folio, new, vma); + folio_add_file_rmap_pmd(folio, &folio->page, vma); } - VM_BUG_ON(pmd_write(pmde) && folio_test_anon(folio) && !PageAnonExclusive= (new)); + VM_WARN_ON_ONCE(pmd_write(pmde) && folio_test_anon(folio) && + !PageAnonExclusive(&folio->page)); set_pmd_at(mm, haddr, pvmw->pmd, pmde); =20 /* No need to invalidate - it was non-present before */ diff --git a/mm/migrate.c b/mm/migrate.c index 222c8c15f782..d08eff028483 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -372,7 +372,7 @@ static bool remove_migration_pte(struct folio *folio, if (!pvmw.pte) { VM_BUG_ON_FOLIO(folio_test_hugetlb(folio) || !folio_test_pmd_mappable(folio), folio); - remove_migration_pmd(&pvmw, new); + remove_migration_pmd(&pvmw, folio); continue; } #endif --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3F89B391822; Thu, 13 Aug 2026 17:37:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642649; cv=none; b=HwDTkMIVagbIRm92w/5YtUcFa5UKrxZh6pW3rBEptu8GTF+iQwM3tRTwJ4UBGIadWxFnHKO4CwNXrocWdXXKV2QlkG4H3lKw98++yspelieonZY5WUa32eKK1mtEIk+V2d5TucLbNk5+5jDBohdwCA0Lc8LPMYlbh4a1/Xg2MlI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642649; c=relaxed/simple; bh=4CWJyPC33HnfZF8tYHNty46xzNZUZPa7yJy7Pu2iMAw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=a2AdAKzt/DQBT8jMXzExMbXxAoXFUQABUgknbqokwFB8IOiinprwKMIcANTOicyKSHI7Xlxdqz7SNTvzrC08EmlXaicDOTJe5NaksU3l4VtSPI0FZHEtr0NVDJZTtFxstmvZ69R4jvQStebB9cM5fnycsb6txgi/IbZnlKIZ4Fc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=g87o/8Lw; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="g87o/8Lw" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1700F1F000E9; Thu, 13 Aug 2026 17:37:07 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642647; bh=mJPTehiXiSdt1BNIK6xJnSzDQS1j/lxnN+kNI1EurjI=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=g87o/8LwPgowJqtU2bHG3Mufu+nohJlc3U6lun9oXAPmaBCwkixuTZDN7rXbdKakS mXhifbZ1+rJrS/PCWRJsyf5c8BbMzX5pKEosCqVYUJYGMUEECrhrE+HsbfIOgBaKda 8MN2Hz2yijd7TM54oIuAh0fVtzMy5Tg2nV9eUqsGOegNPbXNuCITbdnv32FuFF7IuX 5Sbi/VYekz3di2IadYIzKl1LPsRv52IP2u4e/5/SRHMxZ4WS+lLSeL6YudyBGHQknB U1iKaNCFQ95sTMI7Orqnpl+RFqOw+kW0Y0CDpbh0KewYLf91Zrz7qF+EK0CUGDKjHg ZlOMTADzLm4Zg== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:30 +0100 Subject: [PATCH v5 13/16] mm/migrate: calculate large folio page index using PFN Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-13-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2257; i=ljs@kernel.org; h=from:subject:message-id; bh=4CWJyPC33HnfZF8tYHNty46xzNZUZPa7yJy7Pu2iMAw=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/69Rd/JoXfQxfUbOWnaZwqzqy/9jQmbvev5hzQFh0 QVWb4wEOkpZGMS4GGTFFFmefxHfHyQSNq/zgr8bzBxWJpAhDFycAjCRFccZfrMU3rB5/yarPODj V133utms9+Zaec2wdTwoXaIvX7lhMgvD/6j1bxLuvr0pu2920YVPrKcffZ8n/Vv10ffg3wWL+S9 uY2YEAA== X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Rather than having to figure out the page index to use using linear_page_index(), calculate it using PFN. This is a more natural fit as the linear page index is immaterial to determining the folio page index. Derive the page index from the offset between migration entry PFN and folio PFN - pvmw.pfn (set via DEFINE_FOLIO_VMA_WALK() which uses folio_pfn() to obtain it). Additionally remove a not so useful comment and clean the code layout up. No functional change intended. Suggested-by: David Hildenbrand (Arm) Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- mm/migrate.c | 19 +++++++++---------- 1 file changed, 9 insertions(+), 10 deletions(-) diff --git a/mm/migrate.c b/mm/migrate.c index d08eff028483..50d0b547bca3 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -356,16 +356,11 @@ static bool remove_migration_pte(struct folio *folio, =20 while (page_vma_mapped_walk(&pvmw)) { rmap_t rmap_flags =3D RMAP_NONE; - pte_t old_pte; - pte_t pte; + unsigned long idx =3D 0; softleaf_t entry; struct page *new; - unsigned long idx =3D 0; - - /* pgoff is invalid for ksm pages, but they are never large */ - if (folio_test_large(folio) && !folio_test_hugetlb(folio)) - idx =3D linear_page_index(vma, pvmw.address) - pvmw.pgoff; - new =3D folio_page(folio, idx); + pte_t old_pte; + pte_t pte; =20 #ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES /* PMD-mapped THP migration entry */ @@ -381,14 +376,18 @@ static bool remove_migration_pte(struct folio *folio, pvmw.pte); else old_pte =3D ptep_get(pvmw.pte); + + entry =3D softleaf_from_pte(old_pte); + if (folio_test_large(folio) && !folio_test_hugetlb(folio)) + idx =3D softleaf_to_pfn(entry) - pvmw.pfn; + if (rmap_walk_arg->map_unused_to_zeropage && try_to_map_unused_to_zeropage(&pvmw, folio, old_pte, idx)) continue; =20 folio_get(folio); + new =3D folio_page(folio, idx); pte =3D mk_pte(new, READ_ONCE(vma->vm_page_prot)); - - entry =3D softleaf_from_pte(old_pte); if (!softleaf_is_migration_young(entry)) pte =3D pte_mkold(pte); if (folio_test_dirty(folio) && softleaf_is_migration_dirty(entry)) --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 58C9D3955EA; Thu, 13 Aug 2026 17:37:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642670; cv=none; b=p60jAg1ljzCJEAA31Oyo36HqV2k9+Dm191TeR1v2kXM7bDuw9KlqTJQuc0cpwmtDD6N08RWMR0/R0avFXQp6eKPQLkzX+mIKvyWJW+QVCaSMDBnSTDTG33yOZjPHkT6SevPrxhGZxt+zmz3Wh0DskZr5zJn+SVTo/6bS6BoTgcc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642670; c=relaxed/simple; bh=21X8tD3ymqApiXmdDvrPodzuRSm0jreU8242LYwalmY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=tu9MA85a8Q/lxT1T5rKC7WhPGhcTNMJajzmMMrzUs+HYk9ZxIPMlBsPQoTO5bpiKDcbLq5OWl5jrTZl2f84jv2Z6J0c2T307XrLeE8ZTYrWI+nDdzYtiGD3nix5fv4mRGWp8vZs5U2TdosChJpFCEUKphgofMSYI7d1+CS0sMYw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=EjDtfCZw; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="EjDtfCZw" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 6983A1F00A3A; Thu, 13 Aug 2026 17:37:28 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642668; bh=k7V3KbJWuqhI1wrXEM1lc4hRJ9fsgJc7Eafz5Ed8bps=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=EjDtfCZwSgC5n/poUhpAkUiu22SB+fza++HAtn/5Tnu+tDvuPXsEU2S4fdHF0gHbq dyCUERBZ4xjE0SOp5SN0GKrNzUpVJhagbUkMiTLZLohzCfA4D8QK9DxzRWPD8QUsUJ oxtF6c06T61iML2LjWvQZgD/tf6ZyDGmQfBNpreQHZpackHxoWN9zXiRxnDMLv+yOF TMc12Xwi2HKHOMI82+SuOvcbCyKKJTgpOcNPZ/ur2fHPYJkzh294+nzl8epupHWFZu emR3UOI9M1eSf8M6V42YUbUmJ16JKUeHmt9lIfxBhIwtoLI1XxvwZ6VXELg8in9k8I FDW5X9LgkCv1g== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:31 +0100 Subject: [PATCH v5 14/16] mm/rmap: use anon pgoff to track MAP_PRIVATE file-backed anon folios Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-14-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=14135; i=ljs@kernel.org; h=from:subject:message-id; bh=21X8tD3ymqApiXmdDvrPodzuRSm0jreU8242LYwalmY=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6/ZWzo/1POEjZHw5SLNL3+PRXSLJZ+6p2Z1pURjj sSHrOWyHaUsDGJcDLJiiizPv4jvDxIJm9d5wd8NZg4rE8gQBi5OAZgIiwjDf9/ljS7xdr6qjx8v OzHlnFsud/S5at3D2kzT/u4/f9JYQoCRYU1x/cL2vAlOHB8+5P44M+Hlwh1yZ7bNi+WrX12WePJ oJBMA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Currently anonymous folios belonging to CoW'd MAP_PRIVATE file-backed mappings are indexed by their page offset within the file in which they were originally mapped. This differs from anonymous folios belonging to pure anon mappings which are indexed by their anonymous page offset (the address at which they'd belong in the VMA when first faulted). This change fixes this inconsistency, always indexing anonymous folios by their anonymous page offset regardless of the VMA to which they belong. The foundations have been laid such that we need only switch this functionality on such by: * Using linear_anon_page_index() in __folio_set_anon() to assign the folio's index to the anonymous linear index rather than the file-backed one. * Otherwise using linear_anon_page_index() in all instances where anonymous folios are being referenced or manipulated. * Replacing vma_address() with vma_filebacked_address() or vma_anon_address() as appropriate. * Updating the merging logic to check that anonymous page offsets are aligned as well as filebacked ones for MAP_PRIVATE file-backed VMAs, introducing needs_adjacent_anon_pgoff() to figure out when this is required. * Updating linear_folio_page_index() to invoke linear_anon_page_index() if the folio is anonymous. * Updating vma_address_end() to use the VMA's anonymous page offset when pvmw->pgoff is anonymous. * Correcting folio_within_range() to use anonymous page offset for anonymous folios. This will have no impact on merging of anonymous VMAs, whose page offset and anonymous page offset are identical, nor will it impact shared file-backed VMAs, which will continue to be merged based on the file-backed page offset. However, MAP_PRIVATE file-backed mappings must now be aligned on anonymous page offset as well. In most instances this should have no impact on merging of file-backed mappings, which are usually not merged all that often, let alone MAP_PRIVATE mapped ones, and rarely remapped and faulted before being moved back in place (the case in which a merge may now fail). One subtle impact of this change is in NUMA interleaving - since commit 88c91dc58582 ("mempolicy: migration attempt to match interleave nodes"), migration heuristically tries to maintain interleaving behaviour matching the policy using folio indices. When doing migration of CoW'd MAP_PRIVATE-file backed ranges, the 'base' upon which the interleaving behaviour is performed will vary for these ranges. However the commit notes that ranges spanning multiple VMAs will already cause varying bases, and that this is an acceptable approximation. It is very unlikely real world use-cases will be impacted by this (MAP_PRIVATE file-backed mappings are already an edge case), and all that will happen is that such ranges will cause interleaving to be rotated over the CoW'd range, with little to no impact. This commit lays the foundations for future scalable CoW work which needs to track some remaps, meaning that most remap tracking can be avoided, and in nearly all cases the anonymous page offset will be able to be used to quickly find the VMA in an mm. Note that the need_rmap_locks check doesn't need to be updated, as any remapping will offset both the anonymous and file-backed page offset, so it suffices to check only one. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- mm/huge_memory.c | 2 +- mm/internal.h | 27 ++++++++------------------- mm/interval_tree.c | 4 ++-- mm/ksm.c | 6 +++--- mm/page_vma_mapped.c | 2 +- mm/rmap.c | 12 ++++++------ mm/userfaultfd.c | 4 ++-- mm/vma.c | 32 +++++++++++++++++++++++++++++++- 8 files changed, 54 insertions(+), 35 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 47c9c6e32eba..e46b3a4750ee 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2887,7 +2887,7 @@ int move_pages_huge_pmd(struct mm_struct *mm, pmd_t *= dst_pmd, pmd_t *src_pmd, pm } =20 folio_move_anon_rmap(src_folio, dst_vma); - src_folio->index =3D linear_page_index(dst_vma, dst_addr); + src_folio->index =3D linear_anon_page_index(dst_vma, dst_addr); =20 _dst_pmd =3D folio_mk_pmd(src_folio, dst_vma->vm_page_prot); /* Follow mremap() behavior and treat the entry dirty after the move */ diff --git a/mm/internal.h b/mm/internal.h index 98b067472ceb..5892b7453b54 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -933,7 +933,8 @@ folio_within_range(struct folio *folio, struct vm_area_= struct *vma, return false; =20 pgoff_folio =3D folio_pgoff(folio); - pgoff_vma_start =3D vma_start_pgoff(vma); + pgoff_vma_start =3D folio_test_anon(folio) ? + vma_start_anon_pgoff(vma) : vma_start_pgoff(vma); =20 if (start < vma->vm_start) start =3D vma->vm_start; @@ -1044,23 +1045,8 @@ static inline unsigned long vma_filebacked_address(c= onst struct vm_area_struct * } =20 /** - * vma_address - Find the virtual address a page range is mapped at. - * @vma: The vma which maps this object. - * @pgoff: The page offset within its object. - * @nr_pages: The number of pages to consider. - * - * If any page in this range is mapped by this VMA, return the first addre= ss - * where any of these pages appear. Otherwise, return -EFAULT. - */ -static inline unsigned long vma_address(const struct vm_area_struct *vma, - pgoff_t pgoff, unsigned long nr_pages) -{ - return __vma_address(vma, pgoff, vma_start_pgoff(vma), nr_pages); -} - -/** - * vma_anon_address - Find the address an anonymous folio with index @pgof= f_anon - * is mapped at. + * vma_anon_address - Find the virtual address an anonymous page range is = mapped + * at. * @vma: The vma which maps this object. * @pgoff_anon: The anonymous page index belonging to the folio. * @nr_pages: The number of pages to consider. @@ -1093,7 +1079,10 @@ static inline unsigned long vma_address_end(struct p= age_vma_mapped_walk *pvmw) if (pvmw->nr_pages =3D=3D 1) return pvmw->address + PAGE_SIZE; =20 - pgoff_vma_start =3D vma_start_pgoff(vma); + if (pvmw->pgoff_is_anon) + pgoff_vma_start =3D vma_start_anon_pgoff(vma); + else + pgoff_vma_start =3D vma_start_pgoff(vma); =20 address =3D vma->vm_start + ((pgoff_end - pgoff_vma_start) << PAGE_SHIFT); diff --git a/mm/interval_tree.c b/mm/interval_tree.c index 3ae9e106d3af..7bbbf15cfbf0 100644 --- a/mm/interval_tree.c +++ b/mm/interval_tree.c @@ -83,12 +83,12 @@ mapping_rmap_tree_iter_next(struct vm_area_struct *vma, =20 static pgoff_t avc_start_pgoff(struct anon_vma_chain *avc) { - return vma_start_pgoff(avc->vma); + return vma_start_anon_pgoff(avc->vma); } =20 static pgoff_t avc_last_pgoff(struct anon_vma_chain *avc) { - return vma_last_pgoff(avc->vma); + return vma_last_anon_pgoff(avc->vma); } =20 INTERVAL_TREE_DEFINE(struct anon_vma_chain, rb, pgoff_t, rb_subtree_last, diff --git a/mm/ksm.c b/mm/ksm.c index 47006f494fcb..99738a213243 100644 --- a/mm/ksm.c +++ b/mm/ksm.c @@ -1625,7 +1625,7 @@ static int try_to_merge_with_ksm_page(struct ksm_rmap= _item *rmap_item, * stable_tree, break_cow() will clean it up. */ rmap_item->anon_vma =3D vma->anon_vma; - rmap_item->linear_page_index =3D linear_page_index(vma, rmap_item->addres= s); + rmap_item->linear_page_index =3D linear_anon_page_index(vma, rmap_item->a= ddress); get_anon_vma(vma->anon_vma); out: mmap_read_unlock(mm); @@ -3152,7 +3152,7 @@ struct folio *ksm_might_need_to_copy(struct folio *fo= lio, return folio; /* no need to copy it */ } else if (!anon_vma) { return folio; /* no need to copy it */ - } else if (folio->index =3D=3D linear_page_index(vma, addr) && + } else if (folio->index =3D=3D linear_anon_page_index(vma, addr) && anon_vma->root =3D=3D vma->anon_vma->root) { return folio; /* still no need to copy it */ } @@ -3222,7 +3222,7 @@ void rmap_walk_ksm(struct folio *folio, struct rmap_w= alk_control *rwc) /* * Currently, KSM folios are always small folios, so it's * sufficient to search for a single page. We can simply use - * the linear_page_index of the original de-duplicate + * the linear_anon_page_index of the original de-duplicate * anonymous page that we remembered in the rmap_item while * de-duplicating. Note that mremap() always de-duplicates KSM * folios: so if there was mremap() in our parent or our child, diff --git a/mm/page_vma_mapped.c b/mm/page_vma_mapped.c index 081e483cc7bf..4e964545e5e8 100644 --- a/mm/page_vma_mapped.c +++ b/mm/page_vma_mapped.c @@ -365,7 +365,7 @@ unsigned long page_mapped_in_vma(const struct page *pag= e, }; =20 if (folio_test_anon(folio)) - pvmw.address =3D vma_address(vma, pgoff, 1); + pvmw.address =3D vma_anon_address(vma, pgoff, 1); else pvmw.address =3D vma_filebacked_address(vma, pgoff, 1); if (pvmw.address =3D=3D -EFAULT) diff --git a/mm/rmap.c b/mm/rmap.c index ab3f879454de..2a9ae25ae72d 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -866,7 +866,7 @@ unsigned long page_address_in_vma(const struct folio *f= olio, vma->anon_vma->root !=3D anon_vma->root) return -EFAULT; /* KSM folios don't reach here because of the !anon_vma check */ - return vma_address(vma, page_pgoff(folio, page), 1); + return vma_anon_address(vma, page_pgoff(folio, page), 1); } else if (!vma->vm_file) { return -EFAULT; } else if (vma->vm_file->f_mapping !=3D folio->mapping) { @@ -1485,7 +1485,7 @@ static void __folio_set_anon(struct folio *folio, str= uct vm_area_struct *vma, */ anon_vma =3D (void *) anon_vma + FOLIO_MAPPING_ANON; WRITE_ONCE(folio->mapping, (struct address_space *) anon_vma); - folio->index =3D linear_page_index(vma, address); + folio->index =3D linear_anon_page_index(vma, address); } =20 /** @@ -1512,8 +1512,8 @@ static void __page_check_anon_rmap(const struct folio= *folio, */ VM_BUG_ON_FOLIO(folio_anon_vma(folio)->root !=3D vma->anon_vma->root, folio); - VM_BUG_ON_PAGE(page_pgoff(folio, page) !=3D linear_page_index(vma, addres= s), - page); + VM_BUG_ON_PAGE(page_pgoff(folio, page) !=3D + linear_anon_page_index(vma, address), page); } =20 static __always_inline void __folio_add_anon_rmap(struct folio *folio, @@ -3040,10 +3040,10 @@ static void rmap_walk_anon(struct folio *folio, pgoff_end =3D pgoff_start + folio_nr_pages(folio) - 1; anon_rmap_tree_foreach(avc, anon_vma, pgoff_start, pgoff_end) { struct vm_area_struct *vma =3D avc->vma; - unsigned long address =3D vma_address(vma, pgoff_start, + const unsigned long address =3D vma_anon_address(vma, pgoff_start, folio_nr_pages(folio)); =20 - VM_BUG_ON_VMA(address =3D=3D -EFAULT, vma); + VM_WARN_ON_ONCE_VMA(address =3D=3D -EFAULT, vma); cond_resched(); =20 if (rwc->invalid_vma && rwc->invalid_vma(vma, rwc->arg)) diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index 8fd24c8b428e..24a4d92ffa3c 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -1352,7 +1352,7 @@ static long move_present_ptes(struct mm_struct *mm, } =20 folio_move_anon_rmap(src_folio, dst_vma); - src_folio->index =3D linear_page_index(dst_vma, dst_addr); + src_folio->index =3D linear_anon_page_index(dst_vma, dst_addr); =20 orig_dst_pte =3D folio_mk_pte(src_folio, dst_vma->vm_page_prot); /* Set soft dirty bit so userspace can notice the pte was moved */ @@ -1428,7 +1428,7 @@ static int move_swap_pte(struct mm_struct *mm, struct= vm_area_struct *dst_vma, */ if (src_folio) { folio_move_anon_rmap(src_folio, dst_vma); - src_folio->index =3D linear_page_index(dst_vma, dst_addr); + src_folio->index =3D linear_anon_page_index(dst_vma, dst_addr); } else { /* * Check if the swap entry is cached after acquiring the src_pte diff --git a/mm/vma.c b/mm/vma.c index b55015952f0d..117ca94fc907 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -204,6 +204,25 @@ static void init_multi_vma_prep(struct vma_prepare *vp, vp->skip_vma_uprobe =3D true; } =20 +/* + * Does this merge require that adjacent VMAs must have adjacent anonymous= page + * offsets in addition to having adjacent vma->vm_pgoff? + * + * This is only required for MAP_PRIVATE-file backed mappings as the page = offset + * for pure anonymous VMAs is equal to the anonymous page offset. + * + * Read-only shared mappings (with VMA_SHARED_BIT cleared) are always unfa= ulted + * so automatically have correct anonymous page offset (as it is always up= dated + * on remap). + * + * 'Special' mappings in the sense of VDSO, VVAR etc. have !file but would= in + * any case not be candidates for merge nor be mergeable. + */ +static bool needs_adjacent_anon_pgoff(const struct vma_merge_struct *vmg) +{ + return vmg->file && vma_flags_is_cow_mapping(&vmg->vma_flags); +} + /* * Return true if we can merge this (vma_flags,anon_vma,file,vm_pgoff) * in front of (at a lower virtual address and file offset than) the vma. @@ -225,6 +244,9 @@ static bool can_vma_merge_before(struct vma_merge_struc= t *vmg) return false; if (vmg_end_pgoff(vmg) !=3D vma_start_pgoff(vmg->next)) return false; + if (needs_adjacent_anon_pgoff(vmg) && + vmg_end_anon_pgoff(vmg) !=3D vma_start_anon_pgoff(vmg->next)) + return false; return true; } =20 @@ -245,6 +267,9 @@ static bool can_vma_merge_after(struct vma_merge_struct= *vmg) return false; if (vma_end_pgoff(vmg->prev) !=3D vmg_start_pgoff(vmg)) return false; + if (needs_adjacent_anon_pgoff(vmg) && + vma_end_anon_pgoff(vmg->prev) !=3D vmg_start_anon_pgoff(vmg)) + return false; return true; } =20 @@ -2048,7 +2073,12 @@ static int anon_vma_compatible(struct vm_area_struct= *a, struct vm_area_struct * if (!vma_flags_empty(&diff)) return false; /* Page offset must align. */ - return vma_end_pgoff(a) =3D=3D vma_start_pgoff(b); + if (vma_end_pgoff(a) !=3D vma_start_pgoff(b)) + return false; + /* Only reached from anon path, so either MAP_PRIVATE file or anon. */ + if (vma_end_anon_pgoff(a) !=3D vma_start_anon_pgoff(b)) + return false; + return true; } =20 /* --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0E890391822; Thu, 13 Aug 2026 17:38:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642690; cv=none; b=TatLAIpp+Z+Etcl502tTTeIHCZyOe7h1gjWxwe3jPyO0O0z0FUOwrYFJo7WUjMqUJRRTQUyBo8I/PEKkhDHOrK64Agg7YL3jxXr1o9ymRXibJgA/6ms8mgj8Kjod7mxWMiLS7xkjaWgHQKq/UmgQLlupY2lrS8SpS93pNTLapAI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642690; c=relaxed/simple; bh=HHjostKOppt8OOF7vuuqCIn8DUKnlUowdatCvDzIb/8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=QYX30wRbFm6tLI8jMnXe3J8yyO32O3Juyt7L/klFUHyLhvmBq+DDToU5sRiotkAUweqmtCLz7ksM91TQo32jgakD+4MIn+5H4FBFxjp5uErkbJG2cSWgvyS/nBgwFdV4lYAD2O+FqINAjr6C0RuaswN5JtaX3XqYZ5dozN3Jioc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=HgWtg/iw; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="HgWtg/iw" Received: by smtp.kernel.org (Postfix) with ESMTPSA id BB74E1F000E9; Thu, 13 Aug 2026 17:37:48 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642688; bh=ecVSw+SFDAYeoiVDEKAfjYEP3TnSRWs7gN9D15vLI6g=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=HgWtg/iwsgFGuYlOl4ZcPwgqvhzFf9rOpPcS5i3CrS3IVlakxeVhKQM9yCdSaI2oA lwUxxFk+4oWVSSzb45Nv/7rXUHdxSzEScO3TKmT+ibLp1H1Px0UmPVP2pkixCsVSQ+ JM1UNEYjcH743rNGDXJfYcJ9YvbSl6fyC/dOmD9hOMcfjvdBbPWLdBcBjDqxX9zg/I FhnabeQbsFJWmUISWiKjF+INQb3ScKnqrv8mdv4QYqLmCtOKwcEe7nys81C8TLuOZ6 CXmeGk284l/pSIcRw+rY1Bj3qZ5zCxCK+ntxhmEBVZgeDKEztxG+woWpcaFxWywHWu Au91bfHax7Dhw== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:32 +0100 Subject: [PATCH v5 15/16] tools/testing/vma: expand VMA merge tests to assert anon pgoff Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-15-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=10115; i=ljs@kernel.org; h=from:subject:message-id; bh=HHjostKOppt8OOF7vuuqCIn8DUKnlUowdatCvDzIb/8=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6+pyGpg3bRYqlPzq8Kl7U/11ObaKjgJlqVxzP63P NX26OH4jlIWBjEuBlkxRZbnX8T3B4mEzeu84O8GM4eVCWQIAxenAEzkyVGG/1H3vwa8O7Bi6Ut5 M6Uvkr0bIk/5PbJjy5gkdvqMSXX+txkM/4OOtdY9edjce8us8NhSE3MtFt3XKa875jg7bOvLqzt nzQcA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Now we have introduced the VMA anonymous page offset attribute and update it when VMAs are manipulated, update VMA merge tests to assert that the anonymous page offset is as expected. Also update instances where we could use vma_start_pgoff() to do so. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- tools/testing/vma/tests/merge.c | 45 ++++++++++++++++++++++++++++++++-----= ---- 1 file changed, 36 insertions(+), 9 deletions(-) diff --git a/tools/testing/vma/tests/merge.c b/tools/testing/vma/tests/merg= e.c index 48418b82b01d..acaab282939c 100644 --- a/tools/testing/vma/tests/merge.c +++ b/tools/testing/vma/tests/merge.c @@ -121,6 +121,7 @@ static bool test_simple_merge(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x3000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); ASSERT_FLAGS_SAME_MASK(&vma->flags, vma_flags); =20 detach_free_vma(vma); @@ -153,6 +154,7 @@ static bool test_simple_modify(void) ASSERT_EQ(vma->vm_start, 0x1000); ASSERT_EQ(vma->vm_end, 0x2000); ASSERT_EQ(vma_start_pgoff(vma), 1); + ASSERT_EQ(vma_start_anon_pgoff(vma), 1); =20 /* * Now walk through the three split VMAs and make sure they are as @@ -165,6 +167,7 @@ static bool test_simple_modify(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x1000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); =20 detach_free_vma(vma); vma_iter_clear(&vmi); @@ -174,6 +177,7 @@ static bool test_simple_modify(void) ASSERT_EQ(vma->vm_start, 0x1000); ASSERT_EQ(vma->vm_end, 0x2000); ASSERT_EQ(vma_start_pgoff(vma), 1); + ASSERT_EQ(vma_start_anon_pgoff(vma), 1); =20 detach_free_vma(vma); vma_iter_clear(&vmi); @@ -183,6 +187,7 @@ static bool test_simple_modify(void) ASSERT_EQ(vma->vm_start, 0x2000); ASSERT_EQ(vma->vm_end, 0x3000); ASSERT_EQ(vma_start_pgoff(vma), 2); + ASSERT_EQ(vma_start_anon_pgoff(vma), 2); =20 detach_free_vma(vma); mtree_destroy(&mm.mm_mt); @@ -212,6 +217,7 @@ static bool test_simple_expand(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x3000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); =20 detach_free_vma(vma); mtree_destroy(&mm.mm_mt); @@ -234,6 +240,7 @@ static bool test_simple_shrink(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x1000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); =20 detach_free_vma(vma); mtree_destroy(&mm.mm_mt); @@ -346,6 +353,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x5000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 3); @@ -367,6 +375,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0x6000); ASSERT_EQ(vma->vm_end, 0x9000); ASSERT_EQ(vma_start_pgoff(vma), 6); + ASSERT_EQ(vma_start_anon_pgoff(vma), 6); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 3); @@ -387,6 +396,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x9000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 2); @@ -407,6 +417,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0xa000); ASSERT_EQ(vma->vm_end, 0xc000); ASSERT_EQ(vma_start_pgoff(vma), 0xa); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0xa); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 2); @@ -426,6 +437,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0xc000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 1); @@ -446,6 +458,7 @@ static bool __test_merge_new(bool is_sticky, bool a_is_= sticky, bool b_is_sticky, ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0xc000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); ASSERT_EQ(vma->anon_vma, &dummy_anon_vma); =20 detach_free_vma(vma); @@ -642,7 +655,8 @@ static bool test_vma_merge_with_close(void) ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_prev->vm_start, 0); ASSERT_EQ(vma_prev->vm_end, 0x5000); - ASSERT_EQ(vma_prev->vm_pgoff, 0); + ASSERT_EQ(vma_start_pgoff(vma_prev), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma_prev), 0); =20 ASSERT_EQ(cleanup_mm(&mm, &vmi), 2); =20 @@ -753,7 +767,8 @@ static bool test_vma_merge_with_close(void) ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_prev->vm_start, 0); ASSERT_EQ(vma_prev->vm_end, 0x5000); - ASSERT_EQ(vma_prev->vm_pgoff, 0); + ASSERT_EQ(vma_start_pgoff(vma_prev), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma_prev), 0); =20 ASSERT_EQ(cleanup_mm(&mm, &vmi), 2); =20 @@ -808,6 +823,7 @@ static bool test_vma_merge_new_with_close(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x5000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); ASSERT_EQ(vma->vm_ops, &vm_ops); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 2); @@ -863,11 +879,13 @@ static bool __test_merge_existing(bool prev_is_sticky= , bool middle_is_sticky, bo ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_next->vm_start, 0x3000); ASSERT_EQ(vma_next->vm_end, 0x9000); - ASSERT_EQ(vma_next->vm_pgoff, 3); + ASSERT_EQ(vma_start_pgoff(vma_next), 3); + ASSERT_EQ(vma_start_anon_pgoff(vma_next), 3); ASSERT_EQ(vma_next->anon_vma, &dummy_anon_vma); ASSERT_EQ(vma->vm_start, 0x2000); ASSERT_EQ(vma->vm_end, 0x3000); ASSERT_EQ(vma_start_pgoff(vma), 2); + ASSERT_EQ(vma_start_anon_pgoff(vma), 2); ASSERT_TRUE(vma_write_started(vma)); ASSERT_TRUE(vma_write_started(vma_next)); ASSERT_EQ(mm.map_count, 2); @@ -897,7 +915,8 @@ static bool __test_merge_existing(bool prev_is_sticky, = bool middle_is_sticky, bo ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_next->vm_start, 0x2000); ASSERT_EQ(vma_next->vm_end, 0x9000); - ASSERT_EQ(vma_next->vm_pgoff, 2); + ASSERT_EQ(vma_start_pgoff(vma_next), 2); + ASSERT_EQ(vma_start_anon_pgoff(vma_next), 2); ASSERT_EQ(vma_next->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma_next)); ASSERT_EQ(mm.map_count, 1); @@ -929,11 +948,13 @@ static bool __test_merge_existing(bool prev_is_sticky= , bool middle_is_sticky, bo ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_prev->vm_start, 0); ASSERT_EQ(vma_prev->vm_end, 0x6000); - ASSERT_EQ(vma_prev->vm_pgoff, 0); + ASSERT_EQ(vma_start_pgoff(vma_prev), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma_prev), 0); ASSERT_EQ(vma_prev->anon_vma, &dummy_anon_vma); ASSERT_EQ(vma->vm_start, 0x6000); ASSERT_EQ(vma->vm_end, 0x7000); ASSERT_EQ(vma_start_pgoff(vma), 6); + ASSERT_EQ(vma_start_anon_pgoff(vma), 6); ASSERT_TRUE(vma_write_started(vma_prev)); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 2); @@ -964,7 +985,8 @@ static bool __test_merge_existing(bool prev_is_sticky, = bool middle_is_sticky, bo ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_prev->vm_start, 0); ASSERT_EQ(vma_prev->vm_end, 0x7000); - ASSERT_EQ(vma_prev->vm_pgoff, 0); + ASSERT_EQ(vma_start_pgoff(vma_prev), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma_prev), 0); ASSERT_EQ(vma_prev->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma_prev)); ASSERT_EQ(mm.map_count, 1); @@ -996,7 +1018,8 @@ static bool __test_merge_existing(bool prev_is_sticky,= bool middle_is_sticky, bo ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_prev->vm_start, 0); ASSERT_EQ(vma_prev->vm_end, 0x9000); - ASSERT_EQ(vma_prev->vm_pgoff, 0); + ASSERT_EQ(vma_start_pgoff(vma_prev), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma_prev), 0); ASSERT_EQ(vma_prev->anon_vma, &dummy_anon_vma); ASSERT_TRUE(vma_write_started(vma_prev)); ASSERT_EQ(mm.map_count, 1); @@ -1126,7 +1149,8 @@ static bool test_anon_vma_non_mergeable(void) ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_prev->vm_start, 0); ASSERT_EQ(vma_prev->vm_end, 0x7000); - ASSERT_EQ(vma_prev->vm_pgoff, 0); + ASSERT_EQ(vma_start_pgoff(vma_prev), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma_prev), 0); ASSERT_TRUE(vma_write_started(vma_prev)); ASSERT_FALSE(vma_write_started(vma_next)); =20 @@ -1157,7 +1181,8 @@ static bool test_anon_vma_non_mergeable(void) ASSERT_EQ(vmg.state, VMA_MERGE_SUCCESS); ASSERT_EQ(vma_prev->vm_start, 0); ASSERT_EQ(vma_prev->vm_end, 0x7000); - ASSERT_EQ(vma_prev->vm_pgoff, 0); + ASSERT_EQ(vma_start_pgoff(vma_prev), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma_prev), 0); ASSERT_TRUE(vma_write_started(vma_prev)); ASSERT_FALSE(vma_write_started(vma_next)); =20 @@ -1419,6 +1444,7 @@ static bool test_merge_extend(void) ASSERT_EQ(vma->vm_start, 0); ASSERT_EQ(vma->vm_end, 0x4000); ASSERT_EQ(vma_start_pgoff(vma), 0); + ASSERT_EQ(vma_start_anon_pgoff(vma), 0); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(mm.map_count, 1); =20 @@ -1459,6 +1485,7 @@ static bool test_expand_only_mode(void) ASSERT_EQ(vma->vm_start, 0x3000); ASSERT_EQ(vma->vm_end, 0x9000); ASSERT_EQ(vma_start_pgoff(vma), 3); + ASSERT_EQ(vma_start_anon_pgoff(vma), 3); ASSERT_TRUE(vma_write_started(vma)); ASSERT_EQ(vma_iter_addr(&vmi), 0x3000); vma_assert_attached(vma); --=20 2.55.0 From nobody Tue Sep 29 02:01:27 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4E42639989B; Thu, 13 Aug 2026 17:38:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642710; cv=none; b=gNWTNZLxeZrfqqZkYqfxJiWKSq7dv4pR1vA3yKEkS+Z8Qv7Pdzs76M7mGGTqlKT/OC1fC/9f/rjg8TYIgS63ER12hM4Bb4Nr32xirqutXgKupjCWyv44uWDZ8GcwgZ1CgXREKYA48Qwgq0zgN9WvFgzQobBzKtE/QJf5XR/yTfA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786642710; c=relaxed/simple; bh=mYrfVGgyI30/pbLGvSPI/0MARQQYwT9E9OZ2ujBLXv8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=oBzcl9ORh/yC6tCvw1ecCAohsx/RHUkOg7cTrSpKcqqb2MrESHk/mG0ur2folk4wEoHc+tkeAhHCHQVImE7Rrmgg/qBN115ZCMUTnzoxIzXqCv1GYTSUU7djoECZLL98q+gqvgndYChmVoSEN14rXKGSUaKPQzJAO8JDPsDFxZg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=flX92SuR; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="flX92SuR" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1BA651F00A3A; Thu, 13 Aug 2026 17:38:08 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786642708; bh=hMo1I5EICsnpWK2NyeWJQT5c5+lum9qWOStjoHG+C7w=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=flX92SuR9DTMANg5HOpHnyNYcgOcMymFbvVD3clXUl0rT1qSy6jmD+PyCw3AOEdZ0 AF94M4rwAMN3kBcOk8UPefw8GBepiPJCWLemuuu7ghdebqrDPLyCqT0+i3qXb3g5Cr MZmnse0QjJA/thOdojIs56h10ZQvCMfm/hOG7lfytW4p/iANjZ4ed1ZYj5cH7aAk/d ZxUVtkGz7fw7aaehO7Rai1OqFhiDrzI+tAKXH5xfTUFcYFO8vNgB+Evua53OFQ+zdK 5p67yd44WTqiBbNyd5q3qQDlb7mQt+MATFFu27DJ1Gw51kr3Rm4BJ6WNGb2VK511p4 4uFoGeeLj72QQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 13 Aug 2026 18:32:33 +0100 Subject: [PATCH v5 16/16] tools/testing/selftests/mm: test anonymous page offset merge behaviour Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-b4-scalable-cow-virt-pgoff-v5-16-c21581c0c3c8@kernel.org> References: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> In-Reply-To: <20260813-b4-scalable-cow-virt-pgoff-v5-0-c21581c0c3c8@kernel.org> To: Andrew Morton , David Hildenbrand , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jann Horn , Pedro Falcato , "Matthew Wilcox (Oracle)" , Jan Kara , Miaohe Lin , Naoya Horiguchi , Rik van Riel , Harry Yoo , Lance Yang , Kees Cook , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Usama Arif , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Peter Xu , Xu Xin , Chengming Zhou , Arnd Bergmann , Greg Kroah-Hartman , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , Sven Schnelle , Alex Deucher , =?utf-8?q?Christian_K=C3=B6nig?= , David Airlie , Simona Vetter , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , Boris Brezillon , Steven Price , Liviu Dudau , Huang Rui , Matthew Auld , =?utf-8?q?Thomas_Hellstr=C3=B6m?= , Rodrigo Vivi , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark , Jason Gunthorpe , John Hubbard , Muchun Song , Oscar Salvador , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park Cc: "Lorenzo Stoakes (ARM)" , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kselftest@vger.kernel.org, kvm@vger.kernel.org, linux-s390@vger.kernel.org, amd-gfx@lists.freedesktop.org, dri-devel@lists.freedesktop.org, intel-xe@lists.freedesktop.org, linux-perf-users@vger.kernel.org, linux-trace-kernel@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3038; i=ljs@kernel.org; h=from:subject:message-id; bh=mYrfVGgyI30/pbLGvSPI/0MARQQYwT9E9OZ2ujBLXv8=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLJq/6/5OCPtm43YsWzDl1uy+WO2RH0/Pfn8bK7pemKNN 6NSxeaodpSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAimx8xMmy8djfnZS3D8vdn XsRrCJmdsXyxpuzQ0he3LmvP3q419fgJhv8RjKlaMtPnTY5dtPdhqX9izFflLRwv/hg1/9v/Sn/ /3Fo2AA== X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 While maintaining anonymous page offsets for VMAs has no impact for most merge cases, it does impact MAP_PRIVATE-mapped file-backed mappings which happen to have matching page offset but not matching anonymous page offset. Assert this behaviour by attempting to map an unfaulted MAP_PRIVATE-memfd region with a faulted one with compatible file page offsets but incompatible anonymous page offsets. Acked-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- tools/testing/selftests/mm/merge.c | 57 ++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 57 insertions(+) diff --git a/tools/testing/selftests/mm/merge.c b/tools/testing/selftests/m= m/merge.c index 519e5ac02db7..52b8727b6628 100644 --- a/tools/testing/selftests/mm/merge.c +++ b/tools/testing/selftests/mm/merge.c @@ -1305,6 +1305,63 @@ TEST_F(merge, merge_vmas_with_mseal) ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 2 * page_size); } =20 +TEST_F(merge, anon_and_page_offset_mismatch_memfd) +{ + struct procmap_fd *procmap =3D &self->procmap; + unsigned int page_size =3D self->page_size; + char *carveout =3D self->carveout; + char *ptr, *ptr2; + int fd; + + /* Create a 10 page memfd descriptor. */ + fd =3D memfd_create("anon_page_offset_test", MFD_CLOEXEC); + ASSERT_NE(fd, -1); + ASSERT_EQ(ftruncate(fd, 10 * page_size), 0); + + /* Map a region using the memfd at page offset 0. */ + ptr =3D mmap(carveout, 5 * page_size, PROT_READ | PROT_WRITE, + MAP_FIXED | MAP_PRIVATE, fd, 0); + ASSERT_NE(ptr, MAP_FAILED); + + /* + * Map another separately and trigger a CoW fault at page offset 5: + * + * |-----------| |---------| + * | unfaulted | | faulted | + * |-----------| |---------| + */ + ptr2 =3D mmap(&carveout[10 * page_size], 5 * page_size, + PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE, + fd, 5 * page_size); + ASSERT_NE(ptr2, MAP_FAILED); + ptr2[0] =3D 'x'; + + /* + * Now move it in place: + * + * |----------| + * | | + * v | + * |-----------| |---------| + * | unfaulted | | faulted | + * |-----------| |---------| + * + * Because the anonymous page offset of the faulted region is now + * &carveout[10 * page_size], despite the two regions being mergeable + * due to file page offset, they are NOT mergeable due to anonymous + * page offset. + */ + ptr2 =3D sys_mremap(ptr2, 5 * page_size, 5 * page_size, + MREMAP_MAYMOVE | MREMAP_FIXED, + &carveout[5 * page_size]); + ASSERT_NE(ptr2, MAP_FAILED); + + /* Assert that they did not merge. */ + ASSERT_TRUE(find_vma_procmap(procmap, ptr)); + ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr); + ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 5 * page_size); +} + TEST_F(merge_with_fork, mremap_faulted_to_unfaulted_prev) { struct procmap_fd *procmap =3D &self->procmap; --=20 2.55.0