[PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff

Lorenzo Stoakes (ARM) posted 16 patches 1 month, 2 weeks ago
arch/s390/mm/gmap_helpers.c             |   2 +-
drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c |   4 +-
drivers/gpu/drm/drm_gem_shmem_helper.c  |   2 +-
drivers/gpu/drm/panthor/panthor_gem.c   |   2 +-
drivers/gpu/drm/ttm/ttm_bo_vm.c         |   2 +-
drivers/gpu/drm/xe/xe_device.c          |   2 +-
fs/proc/task_mmu.c                      |   2 +-
include/linux/mm.h                      | 132 ++++++++++++++++++++++++++++++--
include/linux/mm_types.h                |  12 +++
include/linux/pagemap.h                 |  40 +++++++++-
include/linux/rmap.h                    |   4 +-
include/linux/swapops.h                 |   6 +-
kernel/events/uprobes.c                 |   2 +-
mm/gup.c                                |   2 +-
mm/huge_memory.c                        |  26 +++----
mm/hugetlb.c                            |   2 +-
mm/internal.h                           |  72 ++++++++++++-----
mm/interval_tree.c                      |   4 +-
mm/ksm.c                                |   6 +-
mm/memory-failure.c                     |   4 +-
mm/memory.c                             |  38 +++++----
mm/mempolicy.c                          |   2 +-
mm/migrate.c                            |  21 +++--
mm/mremap.c                             |   6 +-
mm/page_vma_mapped.c                    |   6 +-
mm/rmap.c                               |  22 +++---
mm/userfaultfd.c                        |   4 +-
mm/vma.c                                | 113 +++++++++++++++++++--------
mm/vma.h                                |  89 +++++++++++++--------
mm/vma_exec.c                           |   2 +-
mm/vma_init.c                           |   1 +
tools/testing/selftests/mm/merge.c      |  57 ++++++++++++++
tools/testing/vma/include/dup.h         |  62 ++++++++++++++-
tools/testing/vma/shared.c              |   3 +-
tools/testing/vma/tests/merge.c         |  49 +++++++++---
tools/testing/vma/tests/vma.c           |  50 +++++++++++-
tools/testing/vma/vma_internal.h        |   1 +
37 files changed, 667 insertions(+), 187 deletions(-)
[PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff
Posted by Lorenzo Stoakes (ARM) 1 month, 2 weeks ago
In memory management we've managed to manufacture a great deal of confusion
around the concept of anonymous memory. We have:

1. 'Pure anon' memory - anonymous VMAs whose folios are anonymous and
   swap-backed (thus for reclaim purposes, treated as anonymous). These are
   simple enough.

2. shmem - file-backed VMAs, file-backed folios (from rmap perspective) so
   present in the page cache and mapped by an address_space object, but
   whose folios are also swap-backed (thus treated as anonymous for reclaim
   purposes).

3. MAP_PRIVATE-mapped /dev/zero - a strange beast whose VMAs have
   vma->vm_file set, but which clears vma->vm_ops to satisfy
   vma_is_anonymous(), resulting in VMAs that were mmap()'d referencing a
   file, but are in every other sense anonymous, including the folios.

4. Other MAP_PRIVATE-file backed mappings - These possess file-backed VMAs
   and have file-backed folios until CoW'd, at which point those CoW'd
   folios are anonymous.

This series fixes issue 3.

In order for us to traverse VMAs using the reverse mapping, we require two
fields - folio->mapping and folio->index. The first tells the rmap code
where to look for VMAs, and the second tells it at which offset the folio
starts within the referenced object.

For anonymous folios, folio->mapping points at an anon_vma object. For
file-backed folios, it points at an address_space. And:

* For file-backed folios folio->index is simply the page offset of the start
  of the folio within the file.

* For anonymous folios belonging to pure anon mappings, folio->index is
  equal to the anonymous page offset of the folio.

* For anonymous folios belonging to file-backed mappings (i.e. CoW'd folios
  of a MAP_PRIVATE file-backed mapping), folio->index is equal to the file
  page offset.

This series establishes a new anonymous page offset property of VMAs to
allow us to map anonymous folios at their anonymous page offset, consistent
with pure anon.

The purpose of doing so is to lay the foundations for the scalable CoW
work. This is necessary because scalable CoW looks in the maple tree for
the VMA located at folio->index << PAGE_SHIFT, before falling back to
looking up tracked remaps if necessary.

The MAP_PRIVATE file-backed case means that folio indices will very often
conflict with one another and this remap tracking becomes substantially
more contended, and of course the fast path can never be used.

This also makes it possible, in future, to unshare anonymously mapped
folios with deep fork hierarchies on remap, eliminating the need for remap
tracking in the vast majority of cases.

Similar to page offset of pure anonymous VMAs, we update the anonymous page
offset of unfaulted file-backed VMAs on remap, but do not once
CoW'd (i.e. vma->anon_vma is non-NULL).

Overall, there is little impact on mergeability, which remains exactly the same
for pure anonymous and shared file-backed mappings, with the only impact being
on MAP_PRIVATE-mapped file-backed mappings, which must now match on anonymous
page offset as well as file page offset to be merged.

To fail to merge like this would require CoW'ing the mapping, then finding
another VMA with identical file and compatible page offset to remap next
to.

This is therefore very much an edge case that should have very little
impact (and which scalable CoW may very well address in any case).

Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
---
v5:
* Accumulated tags (thanks everybody!)
* Removed the final 4 patches to be handled later as there are nuances with
  the /dev/zero stuff we need to figure out, as discussed with David.
* Updated the cover letter to reflect this.
* Added comments to vma_flags_is_cow_mapping(),
  vma_[desc_]is_cow_mapping() as per Suren.
* Correct typo as per Suren.
* Reworded test comment in patch 16 from 'fault in' to 'trigger a CoW
  fault' as per David.
* Fix -> 75 char limit in patch 12's commit msg.

v4:
- Updated tags (thanks everyone!)
- Adjusted some prose as per David.
- Adjusted whitespace to 2 tabs in 2/15 as per David.
- Improved linear_page_index(), __linear_anon_page_index() to be more
  succinct in 2/15 as per David and updated commit message to reflect it.
- Added commit to provide vma_[flags_]is_cow_mapping() - nearly all callers
  are calling is_cow_mapping() as is_cow_mapping(vma->vm_flags) so just
  provide a helper to do this for them.
- In the same commit update all is_cow_mapping() callers and remove the
  now-unused function.
- Change linear_anon_page_index() to assert on !CoW mapping using
  vma_is_cow_mapping() in 2/15 as per David.
- Updated 4/15 to output index differently depending on whether the VMA is
  a CoW mapping or not or whether the file index differs from the anon
  index as per Gregory.
- Updated 6/15 to change the 'update page offset' logic in copy_vma() to
  not be gated on CoW or VMA_SHARED_BIT as discussed with David.
- Updated 6/15 to improve the 'faulted in anon vma' assert stuff. It was
  very unclear so rename the variable sensibly and update the comment.
- Added a separate commit to fix the mess that is the faulted_in_anon_vma
  and the VM_WARN_ON_ONCE_VMA() assert in copy_vma() - make it actually
  only update the vmap in cases where the VMA was replaced (backwards
  remap), update the checks to reflect this in a way that's actually
  understandable and improve the comment.
- Updated needs_adjacent_anon_gpoff() in 9/15 to use
  vma_flags_is_cow_mapping() as per David. Updated the comment to reflect
  it.
- Broke out changes to vma_address_end() into a separate patch as per David.
- Eliminated pgoff in vma_address_end() as it adds confusion - just use
  pgoff_end, which was what pgoff used to be (confusingly).
- Placed more variables in vma_address_end() as constants at the start of
  the function.
- Dropped KSM comment in 9/15 as per David.
- Dropped linear_folio_page_index() patch altogether as per
  David.
- Realised all off uffd is only anon so use linear_anon_page_index()
  throughout there and also the huge memory case for the same reason.
- Added a patch to make remove_migration_pmd() accept a folio instead of a
  page.
- Added a patch to calculate large folio index using PFN as suggested by
  David, eliminating the need for linear page index lookup at all.
- Added a patch to add self merge VMA userland tests.
https://patch.msgid.link/20260806-b4-scalable-cow-virt-pgoff-v4-0-ab318a350404@kernel.org

v3:
- Renamed virtual page offset to anonymous page offset across the series as
  per David and adjusted prose to reflect it.
- As part of the rename, eliminated  vma_anon_pgoff_addr() and the
  conflicting vma_[start, end]_anon_pgoff() functions by passing anon pgoff
  to the merge logic and having that figure out which page offset to use,
  significantly simplifying things.
- Introduced needs_adjacent_anon_pgoff() to be really clear about what the
  merge logic is doing.
- Passing through the anon pgoff fixes the issue Sashiko raised with shared
  file-backed mappings having incorrect anon pgoff (inconsequential but a
  wrinkle nonetheless).
- Dropped unnecessary needs_rmap_locks change in 9/15 - this checks to see
  if rmap locks need to be taken due to the range being moved
  backwards. Since both file-backed and anon page offset are updated on
  such a move, it suffices to check only the former. Updated commit message
  to reflect it.
- Sashiko complained about a couple missing .is_anon_walk entries in
  pvmw's - page_mapped_in_vma() and migrate_vma_collect_huge_pmd() -
  however neither impact anything - the field is only meaningful for
  vma_address_end() if nr_pages > 1 and neither site is impacted. Moreover,
  neither site sets pgoff either. Update the commit message to explain
  this.
- Renamed is_anon_walk to pgoff_is_anon and document that pvmw->pgoff is
  only meaningful if pvmw->nr_pages > 1.
- Highlighted the user-visible changes to /dev/zero being true-anon in the
  relevant commit message as per Yang.
- Fixed a bug in the proc-self-map-files-00[1,2].c procfs selftests - the
  code was trying to MAP_PRIVATE 'an arbitrary file' then asserts that
  file-backed procfs entries exist, but happened to choose
  /dev/zero. Updated to the guaranteed-available /proc/self/exe, as
  reported by Mark.
- Fixed an issue with drivers that intentionally mark vma->vm_ops as NULL
  (using the legacy ->mmap callback). If they do this set dummy ops, which
  is what they meant. The mmap_prepare case is fine as nobody does this
  there and this will be fixed when all drivers are finally converted to
  mmap_prepare. As reported by Sashiko.
- Added some missed VMA selftest vma_start_[anon]_pgoff() conversions in
  the merge test commit as per Sashiko.
- Various small prose/comment fixups.
https://patch.msgid.link/20260729-b4-scalable-cow-virt-pgoff-v3-0-e8ecfefea812@kernel.org

v2:
- Removed incorrect assert on always-NULL folio from 7/15, as per syzbot.
- Updated 2/15 so linear_virt_page_index() checks for vma_is_anonymous() as
  well to be cautious about 'special' (VDSO, VVAR, etc.) VMAs accidentally
  being asserted when CONFIG_DEBUG_VM is set, as per Sashiko.
- Updated commit message of 9/15 to mention the subtle change in NUMA
  interleaving behaviour, as per Sashiko.
- Updated 10/15 to assert virtual page offset for adjacent VMAs for various
  VMA userland tests, as per Sashiko.
- Updated 13/15 to remove the !vma->vm_file check altogether after
  MAP_PRIVATE-/dev/zero is made pure anon in vma_start_virt_pgoff().
- Updated 12/15 to check that the /dev/zero device is a character device
  since it turns out that block and character devices have their own
  separate major/minor device number namespaces... :) as per Sashiko.
- Updated 12/15 to fix a bisection hazard where vma->vm_ops would be
  overwritten by vma_dummy_vm_ops for MAP_PRIVATE-/dev/zero, as per
  Sashiko.
- Updated 12/15 to fix another bisection hazard (...!) due to
  ordering of vma_set_anonymous(). Removed in 13/15.
- Added comments to __vm_virt_pgoff[lo, hi] fields referencing
  vma_start_virt_pgoff()'s comment to be clearer what these are as per
  Xu Xin in 1/15.
- Various typo fixes + cleanups in prose.
- Updated the cover letter to point out that the MAP_PRIVATE-/dev/zero
  issue is addressed in this series too.
https://patch.msgid.link/20260720-b4-scalable-cow-virt-pgoff-v2-0-2d549757a76f@kernel.org

v1:
- Rebased onto mm-new.
- Dependent series heavily reviewed and looks highly likely to land, so
  un-RFC.
- Added explicit check for /dev/zero and removed ability for arbitrary
  mmap/mmap_prepare hooks to make themselves anonymous.
- Made MAP_PRIVATE-/dev/zero mappings truly anonymous.
- Added MAP_PRIVATE file-backed mapping merge test to selftests.
- Added a MAP_PRIVATE-/dev/zero VMA userland test to assert that the VMA
  really is made anonymous.
- Added MAP_PRIVATE-/dev/zero merge tests to selftests.
- Fixed missed virtual page index site in try_to_merge_with_ksm_page().
- Updated folio_within_range() to use virtual page offset for anon
  folio. This had no impact as it is only called for large folios at the
  moment (and MAP_PRIVATE-file backed mappings can't currently be backed by
  a large folio) but making the change now protects us for the future.
- Fixed typos etc.
https://patch.msgid.link/20260717-b4-scalable-cow-virt-pgoff-v1-0-cf24910ef094@kernel.org

RFC:
https://patch.msgid.link/cover.1782745153.git.ljs@kernel.org

To: Andrew Morton <akpm@linux-foundation.org>
To: David Hildenbrand <david@kernel.org>
To: "Liam R. Howlett" <liam@infradead.org>
To: Vlastimil Babka <vbabka@kernel.org>
To: Mike Rapoport <rppt@kernel.org>
To: Suren Baghdasaryan <surenb@google.com>
To: Michal Hocko <mhocko@suse.com>
To: Jann Horn <jannh@google.com>
To: Pedro Falcato <pfalcato@suse.de>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
To: Miaohe Lin <linmiaohe@huawei.com>
To: Naoya Horiguchi <nao.horiguchi@gmail.com>
To: Rik van Riel <riel@surriel.com>
To: Harry Yoo <harry@kernel.org>
To: Lance Yang <lance.yang@linux.dev>
To: Kees Cook <kees@kernel.org>
To: Zi Yan <ziy@nvidia.com>
To: Baolin Wang <baolin.wang@linux.alibaba.com>
To: Nico Pache <npache@redhat.com>
To: Ryan Roberts <ryan.roberts@arm.com>
To: Dev Jain <dev.jain@arm.com>
To: Barry Song <baohua@kernel.org>
To: Usama Arif <usama.arif@linux.dev>
To: Matthew Brost <matthew.brost@intel.com>
To: Joshua Hahn <joshua.hahnjy@gmail.com>
To: Rakie Kim <rakie.kim@sk.com>
To: Byungchul Park <byungchul@sk.com>
To: Gregory Price <gourry@gourry.net>
To: Ying Huang <ying.huang@linux.alibaba.com>
To: Alistair Popple <apopple@nvidia.com>
To: Peter Xu <peterx@redhat.com>
To: Xu Xin <xu.xin16@zte.com.cn>
To: Chengming Zhou <chengming.zhou@linux.dev>
To: Arnd Bergmann <arnd@arndb.de>
To: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
To: Christian Borntraeger <borntraeger@linux.ibm.com>
To: Janosch Frank <frankja@linux.ibm.com>
To: Claudio Imbrenda <imbrenda@linux.ibm.com>
To: Alexander Gordeev <agordeev@linux.ibm.com>
To: Gerald Schaefer <gerald.schaefer@linux.ibm.com>
To: Heiko Carstens <hca@linux.ibm.com>
To: Vasily Gorbik <gor@linux.ibm.com>
To: Sven Schnelle <svens@linux.ibm.com>
To: Alex Deucher <alexander.deucher@amd.com>
To: Christian König <christian.koenig@amd.com>
To: David Airlie <airlied@gmail.com>
To: Simona Vetter <simona@ffwll.ch>
To: Maarten Lankhorst <maarten.lankhorst@linux.intel.com>
To: Maxime Ripard <mripard@kernel.org>
To: Thomas Zimmermann <tzimmermann@suse.de>
To: Boris Brezillon <boris.brezillon@collabora.com>
To: Steven Price <steven.price@arm.com>
To: Liviu Dudau <liviu.dudau@arm.com>
To: Huang Rui <ray.huang@amd.com>
To: Matthew Auld <matthew.auld@intel.com>
To: Thomas Hellström <thomas.hellstrom@linux.intel.com>
To: Rodrigo Vivi <rodrigo.vivi@intel.com>
To: Masami Hiramatsu <mhiramat@kernel.org>
To: Oleg Nesterov <oleg@redhat.com>
To: Peter Zijlstra <peterz@infradead.org>
To: Ingo Molnar <mingo@redhat.com>
To: Arnaldo Carvalho de Melo <acme@kernel.org>
To: Namhyung Kim <namhyung@kernel.org>
To: Mark Rutland <mark.rutland@arm.com>
To: Alexander Shishkin <alexander.shishkin@linux.intel.com>
To: Jiri Olsa <jolsa@kernel.org>
To: Ian Rogers <irogers@google.com>
To: Adrian Hunter <adrian.hunter@intel.com>
To: James Clark <james.clark@linaro.org>
To: Jason Gunthorpe <jgg@ziepe.ca>
To: John Hubbard <jhubbard@nvidia.com>
To: Muchun Song <muchun.song@linux.dev>
To: Oscar Salvador <osalvador@suse.de>
To: Chris Li <chrisl@kernel.org>
To: Kairui Song <kasong@tencent.com>
To: Kemeng Shi <shikemeng@huaweicloud.com>
To: Nhat Pham <nphamcs@gmail.com>
To: Baoquan He <baoquan.he@linux.dev>
To: Youngjun Park <youngjun.park@lge.com>
Cc: ljs@kernel.org
Cc: linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-kselftest@vger.kernel.org
Cc: kvm@vger.kernel.org
Cc: linux-s390@vger.kernel.org
Cc: amd-gfx@lists.freedesktop.org
Cc: dri-devel@lists.freedesktop.org
Cc: intel-xe@lists.freedesktop.org
Cc: linux-perf-users@vger.kernel.org
Cc: linux-trace-kernel@vger.kernel.org

---
Lorenzo Stoakes (ARM) (16):
      mm/vma: introduce VMA anon page offset field and add helpers
      mm: provide vma_[flags_]is_cow_mapping() and remove is_cow_mapping()
      mm: introduce linear_anon_page_index()
      mm: abstract vma_address() and introduce vma_anon_address()
      mm: update print_bad_page_map() to show anon index if appropriate
      mm: introduce and use vma_filebacked_address()
      mm/vma: fix self-merge check in copy_vma()
      tools/testing/vma: add tests for copy_vma() self-merge
      mm: propagate VMA anonymous page offset on map, remap, split + merge
      mm/rmap: track whether the page VMA mapped pgoff is anonymous
      mm: clean up vma_address_end()
      mm/huge_memory: update remove_migration_pmd() to accept a folio
      mm/migrate: calculate large folio page index using PFN
      mm/rmap: use anon pgoff to track MAP_PRIVATE file-backed anon folios
      tools/testing/vma: expand VMA merge tests to assert anon pgoff
      tools/testing/selftests/mm: test anonymous page offset merge behaviour

 arch/s390/mm/gmap_helpers.c             |   2 +-
 drivers/gpu/drm/amd/amdgpu/amdgpu_gem.c |   4 +-
 drivers/gpu/drm/drm_gem_shmem_helper.c  |   2 +-
 drivers/gpu/drm/panthor/panthor_gem.c   |   2 +-
 drivers/gpu/drm/ttm/ttm_bo_vm.c         |   2 +-
 drivers/gpu/drm/xe/xe_device.c          |   2 +-
 fs/proc/task_mmu.c                      |   2 +-
 include/linux/mm.h                      | 132 ++++++++++++++++++++++++++++++--
 include/linux/mm_types.h                |  12 +++
 include/linux/pagemap.h                 |  40 +++++++++-
 include/linux/rmap.h                    |   4 +-
 include/linux/swapops.h                 |   6 +-
 kernel/events/uprobes.c                 |   2 +-
 mm/gup.c                                |   2 +-
 mm/huge_memory.c                        |  26 +++----
 mm/hugetlb.c                            |   2 +-
 mm/internal.h                           |  72 ++++++++++++-----
 mm/interval_tree.c                      |   4 +-
 mm/ksm.c                                |   6 +-
 mm/memory-failure.c                     |   4 +-
 mm/memory.c                             |  38 +++++----
 mm/mempolicy.c                          |   2 +-
 mm/migrate.c                            |  21 +++--
 mm/mremap.c                             |   6 +-
 mm/page_vma_mapped.c                    |   6 +-
 mm/rmap.c                               |  22 +++---
 mm/userfaultfd.c                        |   4 +-
 mm/vma.c                                | 113 +++++++++++++++++++--------
 mm/vma.h                                |  89 +++++++++++++--------
 mm/vma_exec.c                           |   2 +-
 mm/vma_init.c                           |   1 +
 tools/testing/selftests/mm/merge.c      |  57 ++++++++++++++
 tools/testing/vma/include/dup.h         |  62 ++++++++++++++-
 tools/testing/vma/shared.c              |   3 +-
 tools/testing/vma/tests/merge.c         |  49 +++++++++---
 tools/testing/vma/tests/vma.c           |  50 +++++++++++-
 tools/testing/vma/vma_internal.h        |   1 +
 37 files changed, 667 insertions(+), 187 deletions(-)
---
base-commit: 5093dba1014c1d7f7e247fd118f0fa8f22136046
change-id: 20260711-b4-scalable-cow-virt-pgoff-a0cc0eb14bc6

Cheers,
-- 
Lorenzo Stoakes (ARM) <ljs@kernel.org>

Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff
Posted by Andrew Morton 1 month, 2 weeks ago
On Thu, 13 Aug 2026 18:32:17 +0100 "Lorenzo Stoakes (ARM)" <ljs@kernel.org> wrote:

> In memory management we've managed to manufacture a great deal of confusion
> around the concept of anonymous memory. We have:
> 
> 1. 'Pure anon' memory - anonymous VMAs whose folios are anonymous and
>    swap-backed (thus for reclaim purposes, treated as anonymous). These are
>    simple enough.
> 
> 2. shmem - file-backed VMAs, file-backed folios (from rmap perspective) so
>    present in the page cache and mapped by an address_space object, but
>    whose folios are also swap-backed (thus treated as anonymous for reclaim
>    purposes).
> 
> 3. MAP_PRIVATE-mapped /dev/zero - a strange beast whose VMAs have
>    vma->vm_file set, but which clears vma->vm_ops to satisfy
>    vma_is_anonymous(), resulting in VMAs that were mmap()'d referencing a
>    file, but are in every other sense anonymous, including the folios.
> 
> 4. Other MAP_PRIVATE-file backed mappings - These possess file-backed VMAs
>    and have file-backed folios until CoW'd, at which point those CoW'd
>    folios are anonymous.
> 
> This series fixes issue 3.

Thanks.  I've updated mm.git's mm-unstable branch to this version. 
Looking good for the second week of the upcoming merge window.


You'll be mortified to hear that Sashiko wasn't able to find anything
to which to apply this.

Sashiko can be guided with a base-commit: tag but I'm not sure how to
tell it what tree/branch to try, or even if that's necessary.  Perhaps
someone can figure this out sometime.

maybe

hp2:/usr/src/linux-next> git log --oneline | grep "mm/vma: introduce VMA anon page offset field and add helpers"
249646a587dc mm/vma: introduce VMA anon page offset field and add helpers

base-commit: 249646a587dc^

But that requires that Sashiko be able to poke around in linux-next
from previous days.

> v5:
> * Accumulated tags (thanks everybody!)
> * Removed the final 4 patches to be handled later as there are nuances with
>   the /dev/zero stuff we need to figure out, as discussed with David.
> * Updated the cover letter to reflect this.
> * Added comments to vma_flags_is_cow_mapping(),
>   vma_[desc_]is_cow_mapping() as per Suren.
> * Correct typo as per Suren.
> * Reworded test comment in patch 16 from 'fault in' to 'trigger a CoW
>   fault' as per David.
> * Fix -> 75 char limit in patch 12's commit msg.

Here's how v5 altered mm.git.  It's rather substantial, but mainly
selftests:


 drivers/char/mem.c                                     |    8 
 include/linux/mm.h                                     |   18 -
 include/linux/pagemap.h                                |    3 
 mm/internal.h                                          |   17 -
 mm/vma.c                                               |   52 ----
 mm/vma.h                                               |    3 
 mm/vma_internal.h                                      |    1 
 tools/testing/selftests/mm/merge.c                     |  106 ----------
 tools/testing/selftests/proc/proc-self-map-files-001.c |    2 
 tools/testing/selftests/proc/proc-self-map-files-002.c |    2 
 tools/testing/vma/include/dup.h                        |   40 ---
 tools/testing/vma/tests/mmap.c                         |   50 ----
 12 files changed, 40 insertions(+), 262 deletions(-)

--- a/drivers/char/mem.c~b
+++ a/drivers/char/mem.c
@@ -506,7 +506,11 @@ static int mmap_zero_prepare(struct vm_a
 	if (vma_desc_test(desc, VMA_SHARED_BIT))
 		return shmem_zero_setup_desc(desc);
 
-	/* MAP_PRIVATE semantics are taken care for us by core mm. */
+	/*
+	 * This is a highly unique situation where we mark a MAP_PRIVATE mapping
+	 * of /dev/zero anonymous, despite it not being.
+	 */
+	vma_desc_set_anonymous(desc);
 	return 0;
 }
 
@@ -694,7 +698,7 @@ static const struct memdev {
 #ifdef CONFIG_DEVPORT
 	[4] = { "port", &port_fops, 0, 0 },
 #endif
-	[DEVZERO_MINOR] = { "zero", &zero_fops, FMODE_NOWAIT, 0666 },
+	[5] = { "zero", &zero_fops, FMODE_NOWAIT, 0666 },
 	[7] = { "full", &full_fops, 0, 0666 },
 	[8] = { "random", &random_fops, FMODE_NOWAIT, 0666 },
 	[9] = { "urandom", &urandom_fops, FMODE_NOWAIT, 0666 },
--- a/include/linux/mm.h~b
+++ a/include/linux/mm.h
@@ -740,9 +740,6 @@ static inline bool fault_flag_allow_retr
 	{ FAULT_FLAG_INTERRUPTIBLE,	"INTERRUPTIBLE" }, \
 	{ FAULT_FLAG_VMA_LOCK,		"VMA_LOCK" }
 
-/* /dev/zero minor device number. Special due to MAP_PRIVATE semantics. */
-#define DEVZERO_MINOR	5
-
 /*
  * vm_fault is filled by the pagefault handler and passed to the vma's
  * ->fault function. The vma's ->fault is responsible for returning a bitmask
@@ -1554,6 +1551,11 @@ static inline void vma_set_anonymous(str
 	vma->vm_ops = NULL;
 }
 
+static inline void vma_desc_set_anonymous(struct vm_area_desc *desc)
+{
+	desc->vm_ops = NULL;
+}
+
 static inline bool vma_is_anonymous(const struct vm_area_struct *vma)
 {
 	return !vma->vm_ops;
@@ -2279,8 +2281,7 @@ void unpin_folios(struct folio **folios,
  * All mappings backed by anonymous folios (all anonymous mappings and most
  * MAP_PRIVATE-file backed ranges) are CoW mappings.
  *
- * All other mappings (including all writable MAP_SHARED mappings) are
- * non-CoW.
+ * All other mappings (including all MAP_SHARED mappings) are non-CoW.
  *
  * The criteria are !VMA_SHARED_BIT, VMA_MAYWRITE_BIT.
  *
@@ -2317,7 +2318,7 @@ static inline bool vma_flags_is_cow_mapp
 
 /**
  * vma_is_cow_mapping() - Is this VMA a CoW mapping?
- * @vma: The VMA to check.
+ * @desc: The VMA to check.
  *
  * See vma_flags_is_cow_mapping() for details.
  *
@@ -4407,8 +4408,9 @@ static inline unsigned long vma_pages(co
  * If @vma is a MAP_PRIVATE file-backed mapping, then this returns the
  * page offset within the file.
  *
- * Edge cases: nommu does not abide by these and CoW MAP_PRIVATE-pfnmap regions
- * have their page offset set to the first PFN in the range.
+ * Edge cases: nommu does not abide by these, MAP_PRIVATE-/dev/zero satisfies
+ * vma_is_anonymous() but has file-backed page offset, and MAP_PRIVATE-pfnmap
+ * regions have their page offset set to the first PFN in the range.
  *
  * Returns: The page offset of the start of @vma.
  */
--- a/include/linux/pagemap.h~b
+++ a/include/linux/pagemap.h
@@ -1128,7 +1128,8 @@ static inline pgoff_t linear_anon_page_i
 	const pgoff_t pgoff = __linear_anon_page_index(vma, address);
 
 	VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma));
-	if (vma_is_anonymous(vma))
+	/* Account for MAP_PRIVATE-/dev/zero which is only semi-anonymous. */
+	if (vma_is_anonymous(vma) && !vma->vm_file)
 		VM_WARN_ON_ONCE(pgoff != linear_page_index(vma, address));
 
 	return pgoff;
--- a/mm/internal.h~b
+++ a/mm/internal.h
@@ -240,18 +240,15 @@ static inline int mmap_file(struct file
 {
 	int err = vfs_mmap(file, vma);
 
+	if (likely(!err))
+		return 0;
+
 	/*
-	 * Either we tried to call the file hook for mmap() and an error arose
-	 * or a driver set vma->vm_ops = NULL intending there to be no VMA
-	 * operations.
-	 *
-	 * In the former case the VMA is in an inconsistent state and we mustn't
-	 * invoke any further hooks on it, in the latter case the hook actually
-	 * wanted no further hooks to be invoked, so fix both by setting dummy
-	 * VMA ops.
+	 * OK, we tried to call the file hook for mmap(), but an error
+	 * arose. The mapping is in an inconsistent state and we must not invoke
+	 * any further hooks on it.
 	 */
-	if (unlikely(err || !vma->vm_ops))
-		vma->vm_ops = &vma_dummy_vm_ops;
+	vma->vm_ops = &vma_dummy_vm_ops;
 
 	return err;
 }
--- a/mm/vma.c~b
+++ a/mm/vma.c
@@ -2621,36 +2621,6 @@ static int __mmap_new_file_vma(struct mm
 	return 0;
 }
 
-static bool map_is_dev_zero(const struct mmap_state *map)
-{
-	const struct file *file = map->file;
-	struct inode *inode;
-
-	if (!file)
-		return false;
-	inode = file_inode(file);
-	if (!S_ISCHR(inode->i_mode))
-		return false;
-	return imajor(inode) == MEM_MAJOR && iminor(inode) == DEVZERO_MINOR;
-}
-
-static void map_set_anon(struct mmap_state *map)
-{
-	map->file = NULL;
-	map->vm_ops = NULL;
-	map->pgoff = map->addr >> PAGE_SHIFT;
-}
-
-static bool map_is_private(const struct mmap_state *map)
-{
-	return !vma_flags_test(&map->vma_flags, VMA_SHARED_BIT);
-}
-
-static bool map_is_anon(const struct mmap_state *map)
-{
-	return map_is_private(map) && !map->file;
-}
-
 /*
  * __mmap_new_vma() - Allocate a new VMA for the region, as merging was not
  * possible.
@@ -2664,7 +2634,8 @@ static bool map_is_anon(const struct mma
 static int __mmap_new_vma(struct mmap_state *map, struct vm_area_struct **vmap,
 	struct mmap_action *action)
 {
-	const bool is_anon = map_is_anon(map);
+	const bool is_anon = !map->file &&
+		!vma_flags_test(&map->vma_flags, VMA_SHARED_BIT);
 	struct vma_iterator *vmi = map->vmi;
 	int error = 0;
 	struct vm_area_struct *vma;
@@ -2806,10 +2777,6 @@ static int call_mmap_prepare(struct mmap
 	if (err)
 		return err;
 
-	/* Hooks cannot mark themselves anonymous. */
-	if (!desc->vm_ops)
-		return -EINVAL;
-
 	err = call_action_prepare(map, desc);
 	if (err)
 		return err;
@@ -2826,21 +2793,16 @@ static int call_mmap_prepare(struct mmap
 	map->vm_ops = desc->vm_ops;
 	map->vm_private_data = desc->private_data;
 
-	/*
-	 * MAP_PRIVATE-/dev/zero mappings are an ancient way of getting
-	 * anonymous mappings. Rather than allowing these mappings to be odd
-	 * outliers, simply make them truly anonymous.
-	 */
-	if (map_is_private(map) && map_is_dev_zero(map))
-		map_set_anon(map);
-
 	return 0;
 }
 
 static void set_vma_user_defined_fields(struct vm_area_struct *vma,
 		struct mmap_state *map)
 {
-	vma->vm_ops = map->vm_ops;
+	if (map->vm_ops)
+		vma->vm_ops = map->vm_ops;
+	else	/* Only /dev/zero should do this. */
+		vma_set_anonymous(vma);
 	vma->vm_private_data = map->vm_private_data;
 }
 
@@ -2920,7 +2882,7 @@ static unsigned long __mmap_region(struc
 		allocated_new = true;
 	}
 
-	if (have_mmap_prepare && !map_is_anon(&map))
+	if (have_mmap_prepare)
 		set_vma_user_defined_fields(vma, &map);
 
 	__mmap_complete(&map, vma);
--- a/mm/vma.h~b
+++ a/mm/vma.h
@@ -267,6 +267,9 @@ static inline void assert_sane_pgoff(str
 	 */
 	if (!vma_is_anonymous(vma))
 		return;
+	/* MAP_PRIVATE-/dev/zero is anon, non-NULL vm_file, but has file pgoff. */
+	if (vma->vm_file)
+		return;
 	/* If faulted in, could have been remapped. */
 	if (vma->anon_vma)
 		return;
--- a/mm/vma_internal.h~b
+++ a/mm/vma_internal.h
@@ -23,7 +23,6 @@
 #include <linux/ksm.h>
 #include <linux/khugepaged.h>
 #include <linux/list.h>
-#include <linux/major.h>
 #include <linux/maple_tree.h>
 #include <linux/mempolicy.h>
 #include <linux/mm.h>
--- a/tools/testing/selftests/mm/merge.c~b
+++ a/tools/testing/selftests/mm/merge.c
@@ -1324,7 +1324,7 @@ TEST_F(merge, anon_and_page_offset_misma
 	ASSERT_NE(ptr, MAP_FAILED);
 
 	/*
-	 * Map another separately and trigger a CoW fault, at page offset 5:
+	 * Map another separately and trigger a CoW fault at page offset 5:
 	 *
 	 * |-----------|           |---------|
 	 * | unfaulted |           | faulted |
@@ -1362,110 +1362,6 @@ TEST_F(merge, anon_and_page_offset_misma
 	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 5 * page_size);
 }
 
-TEST_F(merge, merge_map_private_dev_zero_unfaulted)
-{
-	struct procmap_fd *procmap = &self->procmap;
-	unsigned int page_size = self->page_size;
-	char *carveout = self->carveout;
-	char *ptr, *ptr2;
-	int fd_zero;
-
-	if (access("/dev/zero", F_OK))
-		SKIP(return, "No /dev/zero.");
-	fd_zero = open("/dev/zero", O_RDWR);
-	ASSERT_NE(fd_zero, -1);
-
-	/*
-	 * Map two MAP_PRIVATE-/dev/zero VMAs next to one another with offset 0
-	 * each.
-	 *
-	 * With these being made truly anonymous upon mapping, they will
-	 * merge. If they were file-backed VMAs the page offsets would prevent
-	 * merge:
-	 *
-	 * |-----||------|    |-------------|
-	 * | ptr || ptr2 | -> |     ptr     |
-	 * |-----||------|    |-------------|
-	 */
-	ptr = mmap(carveout, 5 * page_size, PROT_READ | PROT_WRITE,
-		   MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
-	if (ptr == MAP_FAILED) {
-		close(fd_zero);
-		ASSERT_TRUE(false);
-	}
-	ptr2 = mmap(&carveout[5 * page_size], 5 * page_size,
-		   PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
-	if (ptr2 == MAP_FAILED) {
-		close(fd_zero);
-		ASSERT_TRUE(false);
-	}
-	close(fd_zero);
-
-	/* Assert that they merged. */
-	ASSERT_TRUE(find_vma_procmap(procmap, ptr));
-	ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr);
-	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 10 * page_size);
-}
-
-TEST_F(merge, merge_map_private_dev_zero_faulted_unfaulted)
-{
-	struct procmap_fd *procmap = &self->procmap;
-	unsigned int page_size = self->page_size;
-	char *carveout = self->carveout;
-	char *ptr, *ptr2;
-	int fd_zero;
-
-	if (access("/dev/zero", F_OK))
-		SKIP(return, "No /dev/zero.");
-	fd_zero = open("/dev/zero", O_RDWR);
-	ASSERT_NE(fd_zero, -1);
-
-	/*
-	 * Map a MAP_PRIVATE mapping of /dev/zero with page offset 0, then fault
-	 * it in:
-	 *
-	 * |-------------------------------|
-	 * |           faulted             |
-	 * |-------------------------------|
-	 */
-	ptr = mmap(carveout, 15 * page_size, PROT_READ | PROT_WRITE,
-		   MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
-	if (ptr == MAP_FAILED) {
-		close(fd_zero);
-		ASSERT_TRUE(false);
-	}
-	memset(ptr, 'x', 15 * page_size);
-
-	/*
-	 * Unmap the middle:
-	 *
-	 * |---------|           |---------|
-	 * | faulted |           | faulted |
-	 * |---------|           |---------|
-	 */
-	ASSERT_EQ(munmap(&ptr[5 * page_size], 5 * page_size), 0);
-
-	/*
-	 * Map in a new unfaulted mapping in the middle with page offset 0 -
-	 * this should merge and would not if it were treated as a file rather
-	 * than pure anon:
-	 *
-	 * |---------|-----------|---------|
-	 * | faulted | unfaulted | faulted |
-	 * |---------|-----------|---------|
-	 */
-	ptr2 = mmap(&carveout[5 * page_size], 5 * page_size,
-		    PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE,
-		    fd_zero, 0);
-	close(fd_zero);
-	ASSERT_NE(ptr2, MAP_FAILED);
-
-	/* Assert that they merged. */
-	ASSERT_TRUE(find_vma_procmap(procmap, ptr));
-	ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr);
-	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 15 * page_size);
-}
-
 TEST_F(merge_with_fork, mremap_faulted_to_unfaulted_prev)
 {
 	struct procmap_fd *procmap = &self->procmap;
--- a/tools/testing/selftests/proc/proc-self-map-files-001.c~b
+++ a/tools/testing/selftests/proc/proc-self-map-files-001.c
@@ -51,7 +51,7 @@ int main(void)
 	int fd;
 	unsigned long a, b;
 
-	fd = open("/proc/self/exe", O_RDONLY);
+	fd = open("/dev/zero", O_RDONLY);
 	if (fd == -1)
 		return 1;
 
--- a/tools/testing/selftests/proc/proc-self-map-files-002.c~b
+++ a/tools/testing/selftests/proc/proc-self-map-files-002.c
@@ -57,7 +57,7 @@ int main(void)
 	int fd;
 	unsigned long a, b;
 
-	fd = open("/proc/self/exe", O_RDONLY);
+	fd = open("/dev/zero", O_RDONLY);
 	if (fd == -1)
 		return 1;
 
--- a/tools/testing/vma/include/dup.h~b
+++ a/tools/testing/vma/include/dup.h
@@ -15,21 +15,6 @@ struct task_struct *get_current(void);
 #define MMF_HAS_MDWE	28
 #define current get_current()
 
-#define MINORBITS	20
-#define MINORMASK	((1U << MINORBITS) - 1)
-
-#define MAJOR(dev)	((unsigned int) ((dev) >> MINORBITS))
-#define MINOR(dev)	((unsigned int) ((dev) & MINORMASK))
-#define MKDEV(ma, mi)	(((ma) << MINORBITS) | (mi))
-
-#define S_IFMT  00170000
-#define S_IFCHR  0020000
-
-#define S_ISCHR(m)	(((m) & S_IFMT) == S_IFCHR)
-
-#define MEM_MAJOR		1
-#define DEVZERO_MINOR	5
-
 /*
  * Define the task command name length as enum, then it can be visible to
  * BPF programs.
@@ -38,8 +23,6 @@ enum {
 	TASK_COMM_LEN = 16,
 };
 
-typedef unsigned short		umode_t;
-
 /* PARTIALLY implemented types. */
 struct mm_struct {
 	struct maple_tree mm_mt;
@@ -62,10 +45,6 @@ struct address_space {
 	unsigned long		flags;
 	atomic_t		i_mmap_writable;
 };
-struct inode {
-	umode_t			i_mode;
-	dev_t			i_rdev;
-};
 struct file_operations {
 	int (*mmap)(struct file *, struct vm_area_struct *);
 	int (*mmap_prepare)(struct vm_area_desc *);
@@ -73,7 +52,6 @@ struct file_operations {
 struct file {
 	struct address_space	*f_mapping;
 	const struct file_operations	*f_op;
-	struct inode			*f_inode;
 };
 struct anon_vma_chain {
 	struct anon_vma *anon_vma;
@@ -1660,23 +1638,9 @@ static inline pgoff_t linear_anon_page_i
 	const pgoff_t pgoff = __linear_anon_page_index(vma, address);
 
 	VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma));
-	if (vma_is_anonymous(vma))
+	/* Account for MAP_PRIVATE-/dev/zero which is only semi-anonymous. */
+	if (vma_is_anonymous(vma) && !vma->vm_file)
 		VM_WARN_ON_ONCE(pgoff != linear_page_index(vma, address));
 
 	return pgoff;
 }
-
-static inline struct inode *file_inode(const struct file *f)
-{
-	return f->f_inode;
-}
-
-static inline unsigned iminor(const struct inode *inode)
-{
-	return MINOR(inode->i_rdev);
-}
-
-static inline unsigned imajor(const struct inode *inode)
-{
-	return MAJOR(inode->i_rdev);
-}
--- a/tools/testing/vma/tests/mmap.c~b
+++ a/tools/testing/vma/tests/mmap.c
@@ -45,57 +45,7 @@ static bool test_mmap_region_basic(void)
 	return true;
 }
 
-static int dummy_mmap_prepare(struct vm_area_desc *desc)
-{
-	return 0;
-}
-
-static bool test_pure_anon_dev_zero(void)
-{
-	const vma_flags_t vma_flags = mk_vma_flags(VMA_READ_BIT, VMA_WRITE_BIT,
-			VMA_MAYREAD_BIT, VMA_MAYWRITE_BIT);
-	const struct file_operations f_op = {
-		.mmap_prepare = dummy_mmap_prepare,
-	};
-	struct inode inode = {
-		.i_mode = S_IFCHR,
-		.i_rdev = MKDEV(MEM_MAJOR, DEVZERO_MINOR),
-	};
-	struct file file = {
-		.f_inode = &inode,
-		.f_op = &f_op,
-	};
-	struct mm_struct mm = {};
-	struct vm_area_struct *vma;
-	unsigned long addr;
-	VMA_ITERATOR(vmi, &mm, 0);
-
-	current->mm = &mm;
-
-	/*
-	 * Map a MAP_PRIVATE-/dev/zero mapping at address 0x300000 with a page
-	 * offset of 0x10, which we expect to be reset to the anonymous page
-	 * offset.
-	 */
-	addr = __mmap_region(&file, 0x300000, 0x3000, vma_flags, 0x10, NULL);
-	ASSERT_EQ(addr, 0x300000);
-
-	/* Assert that it truly is an anonymous mapping. */
-	vma = vma_lookup(&mm, addr);
-	ASSERT_NE(vma, NULL);
-	ASSERT_TRUE(vma_is_anonymous(vma));
-	ASSERT_EQ(vma->vm_file, NULL);
-	ASSERT_EQ(vma->vm_private_data, NULL);
-	/* Expect anonymous page offsets. */
-	ASSERT_EQ(vma->vm_pgoff, 0x300);
-	ASSERT_EQ(vma_start_anon_pgoff(vma), 0x300);
-
-	cleanup_mm(&mm, &vmi);
-	return true;
-}
-
 static void run_mmap_tests(int *num_tests, int *num_fail)
 {
 	TEST(mmap_region_basic);
-	TEST(pure_anon_dev_zero);
 }
_
Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff
Posted by Lorenzo Stoakes (ARM) 1 month, 2 weeks ago
On Thu, Aug 13, 2026 at 11:53:46AM -0700, Andrew Morton wrote:
> You'll be mortified to hear that Sashiko wasn't able to find anything
> to which to apply this.

:))

Well, when it's right it's useful, when it's wrong or suggesting unrelated
what-nots it's less useful :>)

I do locally put things through claude + Chris Mason's prompts a lot, I
don't always invoke local sashiko as it's very slow and token-heavy or has
been so far, but am planning to do that more also in future.

>
> Sashiko can be guided with a base-commit: tag but I'm not sure how to
> tell it what tree/branch to try, or even if that's necessary.  Perhaps
> someone can figure this out sometime.

b4 gives a base commit, but I think because the trees are rebased it ends
up being the incorrect one.

Not sure what the solution is!

>
> maybe
>
> hp2:/usr/src/linux-next> git log --oneline | grep "mm/vma: introduce VMA anon page offset field and add helpers"
> 249646a587dc mm/vma: introduce VMA anon page offset field and add helpers
>
> base-commit: 249646a587dc^
>
> But that requires that Sashiko be able to poke around in linux-next
> from previous days.
>
> > v5:
> > * Accumulated tags (thanks everybody!)
> > * Removed the final 4 patches to be handled later as there are nuances with
> >   the /dev/zero stuff we need to figure out, as discussed with David.
> > * Updated the cover letter to reflect this.
> > * Added comments to vma_flags_is_cow_mapping(),
> >   vma_[desc_]is_cow_mapping() as per Suren.
> > * Correct typo as per Suren.
> > * Reworded test comment in patch 16 from 'fault in' to 'trigger a CoW
> >   fault' as per David.
> > * Fix -> 75 char limit in patch 12's commit msg.
>
> Here's how v5 altered mm.git.  It's rather substantial, but mainly
> selftests:

Thanks for the diff, always useful!

The noise it's mostly because of dropping the final 4 commits, and as you
say mostly test stuff that will be sent with whichever approach we decide
on for MAP_PRIVATE-/dev/zero in the next cycle.

The actual changes elsewhere are rather trivial otherwise.

What remains, targeting 2nd week of the merge window, is heavily tested +
fully reviewed, so all is still very sane :)

>
>
>  drivers/char/mem.c                                     |    8
>  include/linux/mm.h                                     |   18 -
>  include/linux/pagemap.h                                |    3
>  mm/internal.h                                          |   17 -
>  mm/vma.c                                               |   52 ----
>  mm/vma.h                                               |    3
>  mm/vma_internal.h                                      |    1
>  tools/testing/selftests/mm/merge.c                     |  106 ----------
>  tools/testing/selftests/proc/proc-self-map-files-001.c |    2
>  tools/testing/selftests/proc/proc-self-map-files-002.c |    2
>  tools/testing/vma/include/dup.h                        |   40 ---
>  tools/testing/vma/tests/mmap.c                         |   50 ----
>  12 files changed, 40 insertions(+), 262 deletions(-)
>
> --- a/drivers/char/mem.c~b
> +++ a/drivers/char/mem.c
> @@ -506,7 +506,11 @@ static int mmap_zero_prepare(struct vm_a
>  	if (vma_desc_test(desc, VMA_SHARED_BIT))
>  		return shmem_zero_setup_desc(desc);
>
> -	/* MAP_PRIVATE semantics are taken care for us by core mm. */
> +	/*
> +	 * This is a highly unique situation where we mark a MAP_PRIVATE mapping
> +	 * of /dev/zero anonymous, despite it not being.
> +	 */
> +	vma_desc_set_anonymous(desc);
>  	return 0;
>  }
>
> @@ -694,7 +698,7 @@ static const struct memdev {
>  #ifdef CONFIG_DEVPORT
>  	[4] = { "port", &port_fops, 0, 0 },
>  #endif
> -	[DEVZERO_MINOR] = { "zero", &zero_fops, FMODE_NOWAIT, 0666 },
> +	[5] = { "zero", &zero_fops, FMODE_NOWAIT, 0666 },
>  	[7] = { "full", &full_fops, 0, 0666 },
>  	[8] = { "random", &random_fops, FMODE_NOWAIT, 0666 },
>  	[9] = { "urandom", &urandom_fops, FMODE_NOWAIT, 0666 },
> --- a/include/linux/mm.h~b
> +++ a/include/linux/mm.h
> @@ -740,9 +740,6 @@ static inline bool fault_flag_allow_retr
>  	{ FAULT_FLAG_INTERRUPTIBLE,	"INTERRUPTIBLE" }, \
>  	{ FAULT_FLAG_VMA_LOCK,		"VMA_LOCK" }
>
> -/* /dev/zero minor device number. Special due to MAP_PRIVATE semantics. */
> -#define DEVZERO_MINOR	5
> -
>  /*
>   * vm_fault is filled by the pagefault handler and passed to the vma's
>   * ->fault function. The vma's ->fault is responsible for returning a bitmask
> @@ -1554,6 +1551,11 @@ static inline void vma_set_anonymous(str
>  	vma->vm_ops = NULL;
>  }
>
> +static inline void vma_desc_set_anonymous(struct vm_area_desc *desc)
> +{
> +	desc->vm_ops = NULL;
> +}
> +
>  static inline bool vma_is_anonymous(const struct vm_area_struct *vma)
>  {
>  	return !vma->vm_ops;
> @@ -2279,8 +2281,7 @@ void unpin_folios(struct folio **folios,
>   * All mappings backed by anonymous folios (all anonymous mappings and most
>   * MAP_PRIVATE-file backed ranges) are CoW mappings.
>   *
> - * All other mappings (including all writable MAP_SHARED mappings) are
> - * non-CoW.
> + * All other mappings (including all MAP_SHARED mappings) are non-CoW.
>   *
>   * The criteria are !VMA_SHARED_BIT, VMA_MAYWRITE_BIT.
>   *
> @@ -2317,7 +2318,7 @@ static inline bool vma_flags_is_cow_mapp
>
>  /**
>   * vma_is_cow_mapping() - Is this VMA a CoW mapping?
> - * @vma: The VMA to check.
> + * @desc: The VMA to check.
>   *
>   * See vma_flags_is_cow_mapping() for details.
>   *
> @@ -4407,8 +4408,9 @@ static inline unsigned long vma_pages(co
>   * If @vma is a MAP_PRIVATE file-backed mapping, then this returns the
>   * page offset within the file.
>   *
> - * Edge cases: nommu does not abide by these and CoW MAP_PRIVATE-pfnmap regions
> - * have their page offset set to the first PFN in the range.
> + * Edge cases: nommu does not abide by these, MAP_PRIVATE-/dev/zero satisfies
> + * vma_is_anonymous() but has file-backed page offset, and MAP_PRIVATE-pfnmap
> + * regions have their page offset set to the first PFN in the range.
>   *
>   * Returns: The page offset of the start of @vma.
>   */
> --- a/include/linux/pagemap.h~b
> +++ a/include/linux/pagemap.h
> @@ -1128,7 +1128,8 @@ static inline pgoff_t linear_anon_page_i
>  	const pgoff_t pgoff = __linear_anon_page_index(vma, address);
>
>  	VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma));
> -	if (vma_is_anonymous(vma))
> +	/* Account for MAP_PRIVATE-/dev/zero which is only semi-anonymous. */
> +	if (vma_is_anonymous(vma) && !vma->vm_file)
>  		VM_WARN_ON_ONCE(pgoff != linear_page_index(vma, address));
>
>  	return pgoff;
> --- a/mm/internal.h~b
> +++ a/mm/internal.h
> @@ -240,18 +240,15 @@ static inline int mmap_file(struct file
>  {
>  	int err = vfs_mmap(file, vma);
>
> +	if (likely(!err))
> +		return 0;
> +
>  	/*
> -	 * Either we tried to call the file hook for mmap() and an error arose
> -	 * or a driver set vma->vm_ops = NULL intending there to be no VMA
> -	 * operations.
> -	 *
> -	 * In the former case the VMA is in an inconsistent state and we mustn't
> -	 * invoke any further hooks on it, in the latter case the hook actually
> -	 * wanted no further hooks to be invoked, so fix both by setting dummy
> -	 * VMA ops.
> +	 * OK, we tried to call the file hook for mmap(), but an error
> +	 * arose. The mapping is in an inconsistent state and we must not invoke
> +	 * any further hooks on it.
>  	 */
> -	if (unlikely(err || !vma->vm_ops))
> -		vma->vm_ops = &vma_dummy_vm_ops;
> +	vma->vm_ops = &vma_dummy_vm_ops;
>
>  	return err;
>  }
> --- a/mm/vma.c~b
> +++ a/mm/vma.c
> @@ -2621,36 +2621,6 @@ static int __mmap_new_file_vma(struct mm
>  	return 0;
>  }
>
> -static bool map_is_dev_zero(const struct mmap_state *map)
> -{
> -	const struct file *file = map->file;
> -	struct inode *inode;
> -
> -	if (!file)
> -		return false;
> -	inode = file_inode(file);
> -	if (!S_ISCHR(inode->i_mode))
> -		return false;
> -	return imajor(inode) == MEM_MAJOR && iminor(inode) == DEVZERO_MINOR;
> -}
> -
> -static void map_set_anon(struct mmap_state *map)
> -{
> -	map->file = NULL;
> -	map->vm_ops = NULL;
> -	map->pgoff = map->addr >> PAGE_SHIFT;
> -}
> -
> -static bool map_is_private(const struct mmap_state *map)
> -{
> -	return !vma_flags_test(&map->vma_flags, VMA_SHARED_BIT);
> -}
> -
> -static bool map_is_anon(const struct mmap_state *map)
> -{
> -	return map_is_private(map) && !map->file;
> -}
> -
>  /*
>   * __mmap_new_vma() - Allocate a new VMA for the region, as merging was not
>   * possible.
> @@ -2664,7 +2634,8 @@ static bool map_is_anon(const struct mma
>  static int __mmap_new_vma(struct mmap_state *map, struct vm_area_struct **vmap,
>  	struct mmap_action *action)
>  {
> -	const bool is_anon = map_is_anon(map);
> +	const bool is_anon = !map->file &&
> +		!vma_flags_test(&map->vma_flags, VMA_SHARED_BIT);
>  	struct vma_iterator *vmi = map->vmi;
>  	int error = 0;
>  	struct vm_area_struct *vma;
> @@ -2806,10 +2777,6 @@ static int call_mmap_prepare(struct mmap
>  	if (err)
>  		return err;
>
> -	/* Hooks cannot mark themselves anonymous. */
> -	if (!desc->vm_ops)
> -		return -EINVAL;
> -
>  	err = call_action_prepare(map, desc);
>  	if (err)
>  		return err;
> @@ -2826,21 +2793,16 @@ static int call_mmap_prepare(struct mmap
>  	map->vm_ops = desc->vm_ops;
>  	map->vm_private_data = desc->private_data;
>
> -	/*
> -	 * MAP_PRIVATE-/dev/zero mappings are an ancient way of getting
> -	 * anonymous mappings. Rather than allowing these mappings to be odd
> -	 * outliers, simply make them truly anonymous.
> -	 */
> -	if (map_is_private(map) && map_is_dev_zero(map))
> -		map_set_anon(map);
> -
>  	return 0;
>  }
>
>  static void set_vma_user_defined_fields(struct vm_area_struct *vma,
>  		struct mmap_state *map)
>  {
> -	vma->vm_ops = map->vm_ops;
> +	if (map->vm_ops)
> +		vma->vm_ops = map->vm_ops;
> +	else	/* Only /dev/zero should do this. */
> +		vma_set_anonymous(vma);
>  	vma->vm_private_data = map->vm_private_data;
>  }
>
> @@ -2920,7 +2882,7 @@ static unsigned long __mmap_region(struc
>  		allocated_new = true;
>  	}
>
> -	if (have_mmap_prepare && !map_is_anon(&map))
> +	if (have_mmap_prepare)
>  		set_vma_user_defined_fields(vma, &map);
>
>  	__mmap_complete(&map, vma);
> --- a/mm/vma.h~b
> +++ a/mm/vma.h
> @@ -267,6 +267,9 @@ static inline void assert_sane_pgoff(str
>  	 */
>  	if (!vma_is_anonymous(vma))
>  		return;
> +	/* MAP_PRIVATE-/dev/zero is anon, non-NULL vm_file, but has file pgoff. */
> +	if (vma->vm_file)
> +		return;
>  	/* If faulted in, could have been remapped. */
>  	if (vma->anon_vma)
>  		return;
> --- a/mm/vma_internal.h~b
> +++ a/mm/vma_internal.h
> @@ -23,7 +23,6 @@
>  #include <linux/ksm.h>
>  #include <linux/khugepaged.h>
>  #include <linux/list.h>
> -#include <linux/major.h>
>  #include <linux/maple_tree.h>
>  #include <linux/mempolicy.h>
>  #include <linux/mm.h>
> --- a/tools/testing/selftests/mm/merge.c~b
> +++ a/tools/testing/selftests/mm/merge.c
> @@ -1324,7 +1324,7 @@ TEST_F(merge, anon_and_page_offset_misma
>  	ASSERT_NE(ptr, MAP_FAILED);
>
>  	/*
> -	 * Map another separately and trigger a CoW fault, at page offset 5:
> +	 * Map another separately and trigger a CoW fault at page offset 5:
>  	 *
>  	 * |-----------|           |---------|
>  	 * | unfaulted |           | faulted |
> @@ -1362,110 +1362,6 @@ TEST_F(merge, anon_and_page_offset_misma
>  	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 5 * page_size);
>  }
>
> -TEST_F(merge, merge_map_private_dev_zero_unfaulted)
> -{
> -	struct procmap_fd *procmap = &self->procmap;
> -	unsigned int page_size = self->page_size;
> -	char *carveout = self->carveout;
> -	char *ptr, *ptr2;
> -	int fd_zero;
> -
> -	if (access("/dev/zero", F_OK))
> -		SKIP(return, "No /dev/zero.");
> -	fd_zero = open("/dev/zero", O_RDWR);
> -	ASSERT_NE(fd_zero, -1);
> -
> -	/*
> -	 * Map two MAP_PRIVATE-/dev/zero VMAs next to one another with offset 0
> -	 * each.
> -	 *
> -	 * With these being made truly anonymous upon mapping, they will
> -	 * merge. If they were file-backed VMAs the page offsets would prevent
> -	 * merge:
> -	 *
> -	 * |-----||------|    |-------------|
> -	 * | ptr || ptr2 | -> |     ptr     |
> -	 * |-----||------|    |-------------|
> -	 */
> -	ptr = mmap(carveout, 5 * page_size, PROT_READ | PROT_WRITE,
> -		   MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
> -	if (ptr == MAP_FAILED) {
> -		close(fd_zero);
> -		ASSERT_TRUE(false);
> -	}
> -	ptr2 = mmap(&carveout[5 * page_size], 5 * page_size,
> -		   PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
> -	if (ptr2 == MAP_FAILED) {
> -		close(fd_zero);
> -		ASSERT_TRUE(false);
> -	}
> -	close(fd_zero);
> -
> -	/* Assert that they merged. */
> -	ASSERT_TRUE(find_vma_procmap(procmap, ptr));
> -	ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr);
> -	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 10 * page_size);
> -}
> -
> -TEST_F(merge, merge_map_private_dev_zero_faulted_unfaulted)
> -{
> -	struct procmap_fd *procmap = &self->procmap;
> -	unsigned int page_size = self->page_size;
> -	char *carveout = self->carveout;
> -	char *ptr, *ptr2;
> -	int fd_zero;
> -
> -	if (access("/dev/zero", F_OK))
> -		SKIP(return, "No /dev/zero.");
> -	fd_zero = open("/dev/zero", O_RDWR);
> -	ASSERT_NE(fd_zero, -1);
> -
> -	/*
> -	 * Map a MAP_PRIVATE mapping of /dev/zero with page offset 0, then fault
> -	 * it in:
> -	 *
> -	 * |-------------------------------|
> -	 * |           faulted             |
> -	 * |-------------------------------|
> -	 */
> -	ptr = mmap(carveout, 15 * page_size, PROT_READ | PROT_WRITE,
> -		   MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
> -	if (ptr == MAP_FAILED) {
> -		close(fd_zero);
> -		ASSERT_TRUE(false);
> -	}
> -	memset(ptr, 'x', 15 * page_size);
> -
> -	/*
> -	 * Unmap the middle:
> -	 *
> -	 * |---------|           |---------|
> -	 * | faulted |           | faulted |
> -	 * |---------|           |---------|
> -	 */
> -	ASSERT_EQ(munmap(&ptr[5 * page_size], 5 * page_size), 0);
> -
> -	/*
> -	 * Map in a new unfaulted mapping in the middle with page offset 0 -
> -	 * this should merge and would not if it were treated as a file rather
> -	 * than pure anon:
> -	 *
> -	 * |---------|-----------|---------|
> -	 * | faulted | unfaulted | faulted |
> -	 * |---------|-----------|---------|
> -	 */
> -	ptr2 = mmap(&carveout[5 * page_size], 5 * page_size,
> -		    PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE,
> -		    fd_zero, 0);
> -	close(fd_zero);
> -	ASSERT_NE(ptr2, MAP_FAILED);
> -
> -	/* Assert that they merged. */
> -	ASSERT_TRUE(find_vma_procmap(procmap, ptr));
> -	ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr);
> -	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 15 * page_size);
> -}
> -
>  TEST_F(merge_with_fork, mremap_faulted_to_unfaulted_prev)
>  {
>  	struct procmap_fd *procmap = &self->procmap;
> --- a/tools/testing/selftests/proc/proc-self-map-files-001.c~b
> +++ a/tools/testing/selftests/proc/proc-self-map-files-001.c
> @@ -51,7 +51,7 @@ int main(void)
>  	int fd;
>  	unsigned long a, b;
>
> -	fd = open("/proc/self/exe", O_RDONLY);
> +	fd = open("/dev/zero", O_RDONLY);
>  	if (fd == -1)
>  		return 1;
>
> --- a/tools/testing/selftests/proc/proc-self-map-files-002.c~b
> +++ a/tools/testing/selftests/proc/proc-self-map-files-002.c
> @@ -57,7 +57,7 @@ int main(void)
>  	int fd;
>  	unsigned long a, b;
>
> -	fd = open("/proc/self/exe", O_RDONLY);
> +	fd = open("/dev/zero", O_RDONLY);
>  	if (fd == -1)
>  		return 1;
>
> --- a/tools/testing/vma/include/dup.h~b
> +++ a/tools/testing/vma/include/dup.h
> @@ -15,21 +15,6 @@ struct task_struct *get_current(void);
>  #define MMF_HAS_MDWE	28
>  #define current get_current()
>
> -#define MINORBITS	20
> -#define MINORMASK	((1U << MINORBITS) - 1)
> -
> -#define MAJOR(dev)	((unsigned int) ((dev) >> MINORBITS))
> -#define MINOR(dev)	((unsigned int) ((dev) & MINORMASK))
> -#define MKDEV(ma, mi)	(((ma) << MINORBITS) | (mi))
> -
> -#define S_IFMT  00170000
> -#define S_IFCHR  0020000
> -
> -#define S_ISCHR(m)	(((m) & S_IFMT) == S_IFCHR)
> -
> -#define MEM_MAJOR		1
> -#define DEVZERO_MINOR	5
> -
>  /*
>   * Define the task command name length as enum, then it can be visible to
>   * BPF programs.
> @@ -38,8 +23,6 @@ enum {
>  	TASK_COMM_LEN = 16,
>  };
>
> -typedef unsigned short		umode_t;
> -
>  /* PARTIALLY implemented types. */
>  struct mm_struct {
>  	struct maple_tree mm_mt;
> @@ -62,10 +45,6 @@ struct address_space {
>  	unsigned long		flags;
>  	atomic_t		i_mmap_writable;
>  };
> -struct inode {
> -	umode_t			i_mode;
> -	dev_t			i_rdev;
> -};
>  struct file_operations {
>  	int (*mmap)(struct file *, struct vm_area_struct *);
>  	int (*mmap_prepare)(struct vm_area_desc *);
> @@ -73,7 +52,6 @@ struct file_operations {
>  struct file {
>  	struct address_space	*f_mapping;
>  	const struct file_operations	*f_op;
> -	struct inode			*f_inode;
>  };
>  struct anon_vma_chain {
>  	struct anon_vma *anon_vma;
> @@ -1660,23 +1638,9 @@ static inline pgoff_t linear_anon_page_i
>  	const pgoff_t pgoff = __linear_anon_page_index(vma, address);
>
>  	VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma));
> -	if (vma_is_anonymous(vma))
> +	/* Account for MAP_PRIVATE-/dev/zero which is only semi-anonymous. */
> +	if (vma_is_anonymous(vma) && !vma->vm_file)
>  		VM_WARN_ON_ONCE(pgoff != linear_page_index(vma, address));
>
>  	return pgoff;
>  }
> -
> -static inline struct inode *file_inode(const struct file *f)
> -{
> -	return f->f_inode;
> -}
> -
> -static inline unsigned iminor(const struct inode *inode)
> -{
> -	return MINOR(inode->i_rdev);
> -}
> -
> -static inline unsigned imajor(const struct inode *inode)
> -{
> -	return MAJOR(inode->i_rdev);
> -}
> --- a/tools/testing/vma/tests/mmap.c~b
> +++ a/tools/testing/vma/tests/mmap.c
> @@ -45,57 +45,7 @@ static bool test_mmap_region_basic(void)
>  	return true;
>  }
>
> -static int dummy_mmap_prepare(struct vm_area_desc *desc)
> -{
> -	return 0;
> -}
> -
> -static bool test_pure_anon_dev_zero(void)
> -{
> -	const vma_flags_t vma_flags = mk_vma_flags(VMA_READ_BIT, VMA_WRITE_BIT,
> -			VMA_MAYREAD_BIT, VMA_MAYWRITE_BIT);
> -	const struct file_operations f_op = {
> -		.mmap_prepare = dummy_mmap_prepare,
> -	};
> -	struct inode inode = {
> -		.i_mode = S_IFCHR,
> -		.i_rdev = MKDEV(MEM_MAJOR, DEVZERO_MINOR),
> -	};
> -	struct file file = {
> -		.f_inode = &inode,
> -		.f_op = &f_op,
> -	};
> -	struct mm_struct mm = {};
> -	struct vm_area_struct *vma;
> -	unsigned long addr;
> -	VMA_ITERATOR(vmi, &mm, 0);
> -
> -	current->mm = &mm;
> -
> -	/*
> -	 * Map a MAP_PRIVATE-/dev/zero mapping at address 0x300000 with a page
> -	 * offset of 0x10, which we expect to be reset to the anonymous page
> -	 * offset.
> -	 */
> -	addr = __mmap_region(&file, 0x300000, 0x3000, vma_flags, 0x10, NULL);
> -	ASSERT_EQ(addr, 0x300000);
> -
> -	/* Assert that it truly is an anonymous mapping. */
> -	vma = vma_lookup(&mm, addr);
> -	ASSERT_NE(vma, NULL);
> -	ASSERT_TRUE(vma_is_anonymous(vma));
> -	ASSERT_EQ(vma->vm_file, NULL);
> -	ASSERT_EQ(vma->vm_private_data, NULL);
> -	/* Expect anonymous page offsets. */
> -	ASSERT_EQ(vma->vm_pgoff, 0x300);
> -	ASSERT_EQ(vma_start_anon_pgoff(vma), 0x300);
> -
> -	cleanup_mm(&mm, &vmi);
> -	return true;
> -}
> -
>  static void run_mmap_tests(int *num_tests, int *num_fail)
>  {
>  	TEST(mmap_region_basic);
> -	TEST(pure_anon_dev_zero);
>  }
> _
>

--
Cheers, Lorenzo
Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff
Posted by Mike Rapoport 1 month, 2 weeks ago
On Fri, Aug 14, 2026 at 10:01:19AM +0100, Lorenzo Stoakes (ARM) wrote:
> On Thu, Aug 13, 2026 at 11:53:46AM -0700, Andrew Morton wrote:
> > You'll be mortified to hear that Sashiko wasn't able to find anything
> > to which to apply this.
> 
> :))
> 
> Well, when it's right it's useful, when it's wrong or suggesting unrelated
> what-nots it's less useful :>)
> 
> I do locally put things through claude + Chris Mason's prompts a lot, I
> don't always invoke local sashiko as it's very slow and token-heavy or has
> been so far, but am planning to do that more also in future.
> 
> >
> > Sashiko can be guided with a base-commit: tag but I'm not sure how to
> > tell it what tree/branch to try, or even if that's necessary.  Perhaps
> > someone can figure this out sometime.
> 
> b4 gives a base commit, but I think because the trees are rebased it ends
> up being the incorrect one.

It's not only that the tree is rebased, but also that mm-unstable carries
the previous version of the patches.

> Not sure what the solution is!
> 
> --
> Cheers, Lorenzo

-- 
Sincerely yours,
Mike.
Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff
Posted by Lorenzo Stoakes (ARM) 1 month, 2 weeks ago
On Fri, Aug 14, 2026 at 12:18:50PM +0300, Mike Rapoport wrote:
> On Fri, Aug 14, 2026 at 10:01:19AM +0100, Lorenzo Stoakes (ARM) wrote:
> > On Thu, Aug 13, 2026 at 11:53:46AM -0700, Andrew Morton wrote:
> > > You'll be mortified to hear that Sashiko wasn't able to find anything
> > > to which to apply this.
> >
> > :))
> >
> > Well, when it's right it's useful, when it's wrong or suggesting unrelated
> > what-nots it's less useful :>)
> >
> > I do locally put things through claude + Chris Mason's prompts a lot, I
> > don't always invoke local sashiko as it's very slow and token-heavy or has
> > been so far, but am planning to do that more also in future.
> >
> > >
> > > Sashiko can be guided with a base-commit: tag but I'm not sure how to
> > > tell it what tree/branch to try, or even if that's necessary.  Perhaps
> > > someone can figure this out sometime.
> >
> > b4 gives a base commit, but I think because the trees are rebased it ends
> > up being the incorrect one.
>
> It's not only that the tree is rebased, but also that mm-unstable carries
> the previous version of the patches.

Yeah that's part of the issue, I can't rebase on mm-unstable as a result
obviously.

Is hard to know what is sensible to rebase on - if I rebase on base of my old
version of the series in mm-unstable that'll probably be a bogus commit at some
point due to other rebasing.

>
> > Not sure what the solution is!
> >
> > --
> > Cheers, Lorenzo
>
> --
> Sincerely yours,
> Mike.

--
Cheers, Lorenzo
Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff
Posted by Mike Rapoport 1 month, 2 weeks ago
On Fri, Aug 14, 2026 at 10:23:33AM +0100, Lorenzo Stoakes (ARM) wrote:
> On Fri, Aug 14, 2026 at 12:18:50PM +0300, Mike Rapoport wrote:
> > On Fri, Aug 14, 2026 at 10:01:19AM +0100, Lorenzo Stoakes (ARM) wrote:
> > > On Thu, Aug 13, 2026 at 11:53:46AM -0700, Andrew Morton wrote:
> > > > You'll be mortified to hear that Sashiko wasn't able to find anything
> > > > to which to apply this.
> > >
> > > :))
> > >
> > > Well, when it's right it's useful, when it's wrong or suggesting unrelated
> > > what-nots it's less useful :>)
> > >
> > > I do locally put things through claude + Chris Mason's prompts a lot, I
> > > don't always invoke local sashiko as it's very slow and token-heavy or has
> > > been so far, but am planning to do that more also in future.
> > >
> > > >
> > > > Sashiko can be guided with a base-commit: tag but I'm not sure how to
> > > > tell it what tree/branch to try, or even if that's necessary.  Perhaps
> > > > someone can figure this out sometime.
> > >
> > > b4 gives a base commit, but I think because the trees are rebased it ends
> > > up being the incorrect one.
> >
> > It's not only that the tree is rebased, but also that mm-unstable carries
> > the previous version of the patches.
> 
> Yeah that's part of the issue, I can't rebase on mm-unstable as a result
> obviously.
> 
> Is hard to know what is sensible to rebase on - if I rebase on base of my old
> version of the series in mm-unstable that'll probably be a bogus commit at some
> point due to other rebasing.

If the work does not depend on other commits in mm-unstable, it's possible
to base on -rcX. But often it's not the case.

> > > Not sure what the solution is!
> > >
> > > --
> > > Cheers, Lorenzo
> >
> > --
> > Sincerely yours,
> > Mike.
> 
> --
> Cheers, Lorenzo

-- 
Sincerely yours,
Mike.
Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff
Posted by Matthew Brost 1 month, 2 weeks ago
On Fri, Aug 14, 2026 at 10:01:19AM +0100, Lorenzo Stoakes (ARM) wrote:
> On Thu, Aug 13, 2026 at 11:53:46AM -0700, Andrew Morton wrote:
> > You'll be mortified to hear that Sashiko wasn't able to find anything
> > to which to apply this.
> 
> :))
> 
> Well, when it's right it's useful, when it's wrong or suggesting unrelated
> what-nots it's less useful :>)
> 

Questioning your assumptions is useful, even when they turn out to be wrong.
Show more lines

> I do locally put things through claude + Chris Mason's prompts a lot, I
> don't always invoke local sashiko as it's very slow and token-heavy or has
> been so far, but am planning to do that more also in future.
> 

Yes, it's kind of odd that Sashiko burns more tokens than a full day of
breakfast, lunch, and dinner service. Running Sashiko is a bottleneck in
my workflow, so I'll defer to others on this list.

> >
> > Sashiko can be guided with a base-commit: tag but I'm not sure how to
> > tell it what tree/branch to try, or even if that's necessary.  Perhaps
> > someone can figure this out sometime.
> 
> b4 gives a base commit, but I think because the trees are rebased it ends
> up being the incorrect one.
> 
> Not sure what the solution is!
> 

We have seen this on the Xe list (our list is based on drm-tip),
typically with cross-subsystem patches. Some cross-subsystem patches
apply and run correctly, while others do not but public CI flows run
based on drm-tip. I do not have a bisect or a clear understanding of
what works and what doesn't, but I think it would be very useful if the
community could better understand the root cause.

Matt

> >
> > maybe
> >
> > hp2:/usr/src/linux-next> git log --oneline | grep "mm/vma: introduce VMA anon page offset field and add helpers"
> > 249646a587dc mm/vma: introduce VMA anon page offset field and add helpers
> >
> > base-commit: 249646a587dc^
> >
> > But that requires that Sashiko be able to poke around in linux-next
> > from previous days.
> >
> > > v5:
> > > * Accumulated tags (thanks everybody!)
> > > * Removed the final 4 patches to be handled later as there are nuances with
> > >   the /dev/zero stuff we need to figure out, as discussed with David.
> > > * Updated the cover letter to reflect this.
> > > * Added comments to vma_flags_is_cow_mapping(),
> > >   vma_[desc_]is_cow_mapping() as per Suren.
> > > * Correct typo as per Suren.
> > > * Reworded test comment in patch 16 from 'fault in' to 'trigger a CoW
> > >   fault' as per David.
> > > * Fix -> 75 char limit in patch 12's commit msg.
> >
> > Here's how v5 altered mm.git.  It's rather substantial, but mainly
> > selftests:
> 
> Thanks for the diff, always useful!
> 
> The noise it's mostly because of dropping the final 4 commits, and as you
> say mostly test stuff that will be sent with whichever approach we decide
> on for MAP_PRIVATE-/dev/zero in the next cycle.
> 
> The actual changes elsewhere are rather trivial otherwise.
> 
> What remains, targeting 2nd week of the merge window, is heavily tested +
> fully reviewed, so all is still very sane :)
> 
> >
> >
> >  drivers/char/mem.c                                     |    8
> >  include/linux/mm.h                                     |   18 -
> >  include/linux/pagemap.h                                |    3
> >  mm/internal.h                                          |   17 -
> >  mm/vma.c                                               |   52 ----
> >  mm/vma.h                                               |    3
> >  mm/vma_internal.h                                      |    1
> >  tools/testing/selftests/mm/merge.c                     |  106 ----------
> >  tools/testing/selftests/proc/proc-self-map-files-001.c |    2
> >  tools/testing/selftests/proc/proc-self-map-files-002.c |    2
> >  tools/testing/vma/include/dup.h                        |   40 ---
> >  tools/testing/vma/tests/mmap.c                         |   50 ----
> >  12 files changed, 40 insertions(+), 262 deletions(-)
> >
> > --- a/drivers/char/mem.c~b
> > +++ a/drivers/char/mem.c
> > @@ -506,7 +506,11 @@ static int mmap_zero_prepare(struct vm_a
> >  	if (vma_desc_test(desc, VMA_SHARED_BIT))
> >  		return shmem_zero_setup_desc(desc);
> >
> > -	/* MAP_PRIVATE semantics are taken care for us by core mm. */
> > +	/*
> > +	 * This is a highly unique situation where we mark a MAP_PRIVATE mapping
> > +	 * of /dev/zero anonymous, despite it not being.
> > +	 */
> > +	vma_desc_set_anonymous(desc);
> >  	return 0;
> >  }
> >
> > @@ -694,7 +698,7 @@ static const struct memdev {
> >  #ifdef CONFIG_DEVPORT
> >  	[4] = { "port", &port_fops, 0, 0 },
> >  #endif
> > -	[DEVZERO_MINOR] = { "zero", &zero_fops, FMODE_NOWAIT, 0666 },
> > +	[5] = { "zero", &zero_fops, FMODE_NOWAIT, 0666 },
> >  	[7] = { "full", &full_fops, 0, 0666 },
> >  	[8] = { "random", &random_fops, FMODE_NOWAIT, 0666 },
> >  	[9] = { "urandom", &urandom_fops, FMODE_NOWAIT, 0666 },
> > --- a/include/linux/mm.h~b
> > +++ a/include/linux/mm.h
> > @@ -740,9 +740,6 @@ static inline bool fault_flag_allow_retr
> >  	{ FAULT_FLAG_INTERRUPTIBLE,	"INTERRUPTIBLE" }, \
> >  	{ FAULT_FLAG_VMA_LOCK,		"VMA_LOCK" }
> >
> > -/* /dev/zero minor device number. Special due to MAP_PRIVATE semantics. */
> > -#define DEVZERO_MINOR	5
> > -
> >  /*
> >   * vm_fault is filled by the pagefault handler and passed to the vma's
> >   * ->fault function. The vma's ->fault is responsible for returning a bitmask
> > @@ -1554,6 +1551,11 @@ static inline void vma_set_anonymous(str
> >  	vma->vm_ops = NULL;
> >  }
> >
> > +static inline void vma_desc_set_anonymous(struct vm_area_desc *desc)
> > +{
> > +	desc->vm_ops = NULL;
> > +}
> > +
> >  static inline bool vma_is_anonymous(const struct vm_area_struct *vma)
> >  {
> >  	return !vma->vm_ops;
> > @@ -2279,8 +2281,7 @@ void unpin_folios(struct folio **folios,
> >   * All mappings backed by anonymous folios (all anonymous mappings and most
> >   * MAP_PRIVATE-file backed ranges) are CoW mappings.
> >   *
> > - * All other mappings (including all writable MAP_SHARED mappings) are
> > - * non-CoW.
> > + * All other mappings (including all MAP_SHARED mappings) are non-CoW.
> >   *
> >   * The criteria are !VMA_SHARED_BIT, VMA_MAYWRITE_BIT.
> >   *
> > @@ -2317,7 +2318,7 @@ static inline bool vma_flags_is_cow_mapp
> >
> >  /**
> >   * vma_is_cow_mapping() - Is this VMA a CoW mapping?
> > - * @vma: The VMA to check.
> > + * @desc: The VMA to check.
> >   *
> >   * See vma_flags_is_cow_mapping() for details.
> >   *
> > @@ -4407,8 +4408,9 @@ static inline unsigned long vma_pages(co
> >   * If @vma is a MAP_PRIVATE file-backed mapping, then this returns the
> >   * page offset within the file.
> >   *
> > - * Edge cases: nommu does not abide by these and CoW MAP_PRIVATE-pfnmap regions
> > - * have their page offset set to the first PFN in the range.
> > + * Edge cases: nommu does not abide by these, MAP_PRIVATE-/dev/zero satisfies
> > + * vma_is_anonymous() but has file-backed page offset, and MAP_PRIVATE-pfnmap
> > + * regions have their page offset set to the first PFN in the range.
> >   *
> >   * Returns: The page offset of the start of @vma.
> >   */
> > --- a/include/linux/pagemap.h~b
> > +++ a/include/linux/pagemap.h
> > @@ -1128,7 +1128,8 @@ static inline pgoff_t linear_anon_page_i
> >  	const pgoff_t pgoff = __linear_anon_page_index(vma, address);
> >
> >  	VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma));
> > -	if (vma_is_anonymous(vma))
> > +	/* Account for MAP_PRIVATE-/dev/zero which is only semi-anonymous. */
> > +	if (vma_is_anonymous(vma) && !vma->vm_file)
> >  		VM_WARN_ON_ONCE(pgoff != linear_page_index(vma, address));
> >
> >  	return pgoff;
> > --- a/mm/internal.h~b
> > +++ a/mm/internal.h
> > @@ -240,18 +240,15 @@ static inline int mmap_file(struct file
> >  {
> >  	int err = vfs_mmap(file, vma);
> >
> > +	if (likely(!err))
> > +		return 0;
> > +
> >  	/*
> > -	 * Either we tried to call the file hook for mmap() and an error arose
> > -	 * or a driver set vma->vm_ops = NULL intending there to be no VMA
> > -	 * operations.
> > -	 *
> > -	 * In the former case the VMA is in an inconsistent state and we mustn't
> > -	 * invoke any further hooks on it, in the latter case the hook actually
> > -	 * wanted no further hooks to be invoked, so fix both by setting dummy
> > -	 * VMA ops.
> > +	 * OK, we tried to call the file hook for mmap(), but an error
> > +	 * arose. The mapping is in an inconsistent state and we must not invoke
> > +	 * any further hooks on it.
> >  	 */
> > -	if (unlikely(err || !vma->vm_ops))
> > -		vma->vm_ops = &vma_dummy_vm_ops;
> > +	vma->vm_ops = &vma_dummy_vm_ops;
> >
> >  	return err;
> >  }
> > --- a/mm/vma.c~b
> > +++ a/mm/vma.c
> > @@ -2621,36 +2621,6 @@ static int __mmap_new_file_vma(struct mm
> >  	return 0;
> >  }
> >
> > -static bool map_is_dev_zero(const struct mmap_state *map)
> > -{
> > -	const struct file *file = map->file;
> > -	struct inode *inode;
> > -
> > -	if (!file)
> > -		return false;
> > -	inode = file_inode(file);
> > -	if (!S_ISCHR(inode->i_mode))
> > -		return false;
> > -	return imajor(inode) == MEM_MAJOR && iminor(inode) == DEVZERO_MINOR;
> > -}
> > -
> > -static void map_set_anon(struct mmap_state *map)
> > -{
> > -	map->file = NULL;
> > -	map->vm_ops = NULL;
> > -	map->pgoff = map->addr >> PAGE_SHIFT;
> > -}
> > -
> > -static bool map_is_private(const struct mmap_state *map)
> > -{
> > -	return !vma_flags_test(&map->vma_flags, VMA_SHARED_BIT);
> > -}
> > -
> > -static bool map_is_anon(const struct mmap_state *map)
> > -{
> > -	return map_is_private(map) && !map->file;
> > -}
> > -
> >  /*
> >   * __mmap_new_vma() - Allocate a new VMA for the region, as merging was not
> >   * possible.
> > @@ -2664,7 +2634,8 @@ static bool map_is_anon(const struct mma
> >  static int __mmap_new_vma(struct mmap_state *map, struct vm_area_struct **vmap,
> >  	struct mmap_action *action)
> >  {
> > -	const bool is_anon = map_is_anon(map);
> > +	const bool is_anon = !map->file &&
> > +		!vma_flags_test(&map->vma_flags, VMA_SHARED_BIT);
> >  	struct vma_iterator *vmi = map->vmi;
> >  	int error = 0;
> >  	struct vm_area_struct *vma;
> > @@ -2806,10 +2777,6 @@ static int call_mmap_prepare(struct mmap
> >  	if (err)
> >  		return err;
> >
> > -	/* Hooks cannot mark themselves anonymous. */
> > -	if (!desc->vm_ops)
> > -		return -EINVAL;
> > -
> >  	err = call_action_prepare(map, desc);
> >  	if (err)
> >  		return err;
> > @@ -2826,21 +2793,16 @@ static int call_mmap_prepare(struct mmap
> >  	map->vm_ops = desc->vm_ops;
> >  	map->vm_private_data = desc->private_data;
> >
> > -	/*
> > -	 * MAP_PRIVATE-/dev/zero mappings are an ancient way of getting
> > -	 * anonymous mappings. Rather than allowing these mappings to be odd
> > -	 * outliers, simply make them truly anonymous.
> > -	 */
> > -	if (map_is_private(map) && map_is_dev_zero(map))
> > -		map_set_anon(map);
> > -
> >  	return 0;
> >  }
> >
> >  static void set_vma_user_defined_fields(struct vm_area_struct *vma,
> >  		struct mmap_state *map)
> >  {
> > -	vma->vm_ops = map->vm_ops;
> > +	if (map->vm_ops)
> > +		vma->vm_ops = map->vm_ops;
> > +	else	/* Only /dev/zero should do this. */
> > +		vma_set_anonymous(vma);
> >  	vma->vm_private_data = map->vm_private_data;
> >  }
> >
> > @@ -2920,7 +2882,7 @@ static unsigned long __mmap_region(struc
> >  		allocated_new = true;
> >  	}
> >
> > -	if (have_mmap_prepare && !map_is_anon(&map))
> > +	if (have_mmap_prepare)
> >  		set_vma_user_defined_fields(vma, &map);
> >
> >  	__mmap_complete(&map, vma);
> > --- a/mm/vma.h~b
> > +++ a/mm/vma.h
> > @@ -267,6 +267,9 @@ static inline void assert_sane_pgoff(str
> >  	 */
> >  	if (!vma_is_anonymous(vma))
> >  		return;
> > +	/* MAP_PRIVATE-/dev/zero is anon, non-NULL vm_file, but has file pgoff. */
> > +	if (vma->vm_file)
> > +		return;
> >  	/* If faulted in, could have been remapped. */
> >  	if (vma->anon_vma)
> >  		return;
> > --- a/mm/vma_internal.h~b
> > +++ a/mm/vma_internal.h
> > @@ -23,7 +23,6 @@
> >  #include <linux/ksm.h>
> >  #include <linux/khugepaged.h>
> >  #include <linux/list.h>
> > -#include <linux/major.h>
> >  #include <linux/maple_tree.h>
> >  #include <linux/mempolicy.h>
> >  #include <linux/mm.h>
> > --- a/tools/testing/selftests/mm/merge.c~b
> > +++ a/tools/testing/selftests/mm/merge.c
> > @@ -1324,7 +1324,7 @@ TEST_F(merge, anon_and_page_offset_misma
> >  	ASSERT_NE(ptr, MAP_FAILED);
> >
> >  	/*
> > -	 * Map another separately and trigger a CoW fault, at page offset 5:
> > +	 * Map another separately and trigger a CoW fault at page offset 5:
> >  	 *
> >  	 * |-----------|           |---------|
> >  	 * | unfaulted |           | faulted |
> > @@ -1362,110 +1362,6 @@ TEST_F(merge, anon_and_page_offset_misma
> >  	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 5 * page_size);
> >  }
> >
> > -TEST_F(merge, merge_map_private_dev_zero_unfaulted)
> > -{
> > -	struct procmap_fd *procmap = &self->procmap;
> > -	unsigned int page_size = self->page_size;
> > -	char *carveout = self->carveout;
> > -	char *ptr, *ptr2;
> > -	int fd_zero;
> > -
> > -	if (access("/dev/zero", F_OK))
> > -		SKIP(return, "No /dev/zero.");
> > -	fd_zero = open("/dev/zero", O_RDWR);
> > -	ASSERT_NE(fd_zero, -1);
> > -
> > -	/*
> > -	 * Map two MAP_PRIVATE-/dev/zero VMAs next to one another with offset 0
> > -	 * each.
> > -	 *
> > -	 * With these being made truly anonymous upon mapping, they will
> > -	 * merge. If they were file-backed VMAs the page offsets would prevent
> > -	 * merge:
> > -	 *
> > -	 * |-----||------|    |-------------|
> > -	 * | ptr || ptr2 | -> |     ptr     |
> > -	 * |-----||------|    |-------------|
> > -	 */
> > -	ptr = mmap(carveout, 5 * page_size, PROT_READ | PROT_WRITE,
> > -		   MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
> > -	if (ptr == MAP_FAILED) {
> > -		close(fd_zero);
> > -		ASSERT_TRUE(false);
> > -	}
> > -	ptr2 = mmap(&carveout[5 * page_size], 5 * page_size,
> > -		   PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
> > -	if (ptr2 == MAP_FAILED) {
> > -		close(fd_zero);
> > -		ASSERT_TRUE(false);
> > -	}
> > -	close(fd_zero);
> > -
> > -	/* Assert that they merged. */
> > -	ASSERT_TRUE(find_vma_procmap(procmap, ptr));
> > -	ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr);
> > -	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 10 * page_size);
> > -}
> > -
> > -TEST_F(merge, merge_map_private_dev_zero_faulted_unfaulted)
> > -{
> > -	struct procmap_fd *procmap = &self->procmap;
> > -	unsigned int page_size = self->page_size;
> > -	char *carveout = self->carveout;
> > -	char *ptr, *ptr2;
> > -	int fd_zero;
> > -
> > -	if (access("/dev/zero", F_OK))
> > -		SKIP(return, "No /dev/zero.");
> > -	fd_zero = open("/dev/zero", O_RDWR);
> > -	ASSERT_NE(fd_zero, -1);
> > -
> > -	/*
> > -	 * Map a MAP_PRIVATE mapping of /dev/zero with page offset 0, then fault
> > -	 * it in:
> > -	 *
> > -	 * |-------------------------------|
> > -	 * |           faulted             |
> > -	 * |-------------------------------|
> > -	 */
> > -	ptr = mmap(carveout, 15 * page_size, PROT_READ | PROT_WRITE,
> > -		   MAP_FIXED | MAP_PRIVATE, fd_zero, 0);
> > -	if (ptr == MAP_FAILED) {
> > -		close(fd_zero);
> > -		ASSERT_TRUE(false);
> > -	}
> > -	memset(ptr, 'x', 15 * page_size);
> > -
> > -	/*
> > -	 * Unmap the middle:
> > -	 *
> > -	 * |---------|           |---------|
> > -	 * | faulted |           | faulted |
> > -	 * |---------|           |---------|
> > -	 */
> > -	ASSERT_EQ(munmap(&ptr[5 * page_size], 5 * page_size), 0);
> > -
> > -	/*
> > -	 * Map in a new unfaulted mapping in the middle with page offset 0 -
> > -	 * this should merge and would not if it were treated as a file rather
> > -	 * than pure anon:
> > -	 *
> > -	 * |---------|-----------|---------|
> > -	 * | faulted | unfaulted | faulted |
> > -	 * |---------|-----------|---------|
> > -	 */
> > -	ptr2 = mmap(&carveout[5 * page_size], 5 * page_size,
> > -		    PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE,
> > -		    fd_zero, 0);
> > -	close(fd_zero);
> > -	ASSERT_NE(ptr2, MAP_FAILED);
> > -
> > -	/* Assert that they merged. */
> > -	ASSERT_TRUE(find_vma_procmap(procmap, ptr));
> > -	ASSERT_EQ(procmap->query.vma_start, (unsigned long)ptr);
> > -	ASSERT_EQ(procmap->query.vma_end, (unsigned long)ptr + 15 * page_size);
> > -}
> > -
> >  TEST_F(merge_with_fork, mremap_faulted_to_unfaulted_prev)
> >  {
> >  	struct procmap_fd *procmap = &self->procmap;
> > --- a/tools/testing/selftests/proc/proc-self-map-files-001.c~b
> > +++ a/tools/testing/selftests/proc/proc-self-map-files-001.c
> > @@ -51,7 +51,7 @@ int main(void)
> >  	int fd;
> >  	unsigned long a, b;
> >
> > -	fd = open("/proc/self/exe", O_RDONLY);
> > +	fd = open("/dev/zero", O_RDONLY);
> >  	if (fd == -1)
> >  		return 1;
> >
> > --- a/tools/testing/selftests/proc/proc-self-map-files-002.c~b
> > +++ a/tools/testing/selftests/proc/proc-self-map-files-002.c
> > @@ -57,7 +57,7 @@ int main(void)
> >  	int fd;
> >  	unsigned long a, b;
> >
> > -	fd = open("/proc/self/exe", O_RDONLY);
> > +	fd = open("/dev/zero", O_RDONLY);
> >  	if (fd == -1)
> >  		return 1;
> >
> > --- a/tools/testing/vma/include/dup.h~b
> > +++ a/tools/testing/vma/include/dup.h
> > @@ -15,21 +15,6 @@ struct task_struct *get_current(void);
> >  #define MMF_HAS_MDWE	28
> >  #define current get_current()
> >
> > -#define MINORBITS	20
> > -#define MINORMASK	((1U << MINORBITS) - 1)
> > -
> > -#define MAJOR(dev)	((unsigned int) ((dev) >> MINORBITS))
> > -#define MINOR(dev)	((unsigned int) ((dev) & MINORMASK))
> > -#define MKDEV(ma, mi)	(((ma) << MINORBITS) | (mi))
> > -
> > -#define S_IFMT  00170000
> > -#define S_IFCHR  0020000
> > -
> > -#define S_ISCHR(m)	(((m) & S_IFMT) == S_IFCHR)
> > -
> > -#define MEM_MAJOR		1
> > -#define DEVZERO_MINOR	5
> > -
> >  /*
> >   * Define the task command name length as enum, then it can be visible to
> >   * BPF programs.
> > @@ -38,8 +23,6 @@ enum {
> >  	TASK_COMM_LEN = 16,
> >  };
> >
> > -typedef unsigned short		umode_t;
> > -
> >  /* PARTIALLY implemented types. */
> >  struct mm_struct {
> >  	struct maple_tree mm_mt;
> > @@ -62,10 +45,6 @@ struct address_space {
> >  	unsigned long		flags;
> >  	atomic_t		i_mmap_writable;
> >  };
> > -struct inode {
> > -	umode_t			i_mode;
> > -	dev_t			i_rdev;
> > -};
> >  struct file_operations {
> >  	int (*mmap)(struct file *, struct vm_area_struct *);
> >  	int (*mmap_prepare)(struct vm_area_desc *);
> > @@ -73,7 +52,6 @@ struct file_operations {
> >  struct file {
> >  	struct address_space	*f_mapping;
> >  	const struct file_operations	*f_op;
> > -	struct inode			*f_inode;
> >  };
> >  struct anon_vma_chain {
> >  	struct anon_vma *anon_vma;
> > @@ -1660,23 +1638,9 @@ static inline pgoff_t linear_anon_page_i
> >  	const pgoff_t pgoff = __linear_anon_page_index(vma, address);
> >
> >  	VM_WARN_ON_ONCE(!vma_is_cow_mapping(vma));
> > -	if (vma_is_anonymous(vma))
> > +	/* Account for MAP_PRIVATE-/dev/zero which is only semi-anonymous. */
> > +	if (vma_is_anonymous(vma) && !vma->vm_file)
> >  		VM_WARN_ON_ONCE(pgoff != linear_page_index(vma, address));
> >
> >  	return pgoff;
> >  }
> > -
> > -static inline struct inode *file_inode(const struct file *f)
> > -{
> > -	return f->f_inode;
> > -}
> > -
> > -static inline unsigned iminor(const struct inode *inode)
> > -{
> > -	return MINOR(inode->i_rdev);
> > -}
> > -
> > -static inline unsigned imajor(const struct inode *inode)
> > -{
> > -	return MAJOR(inode->i_rdev);
> > -}
> > --- a/tools/testing/vma/tests/mmap.c~b
> > +++ a/tools/testing/vma/tests/mmap.c
> > @@ -45,57 +45,7 @@ static bool test_mmap_region_basic(void)
> >  	return true;
> >  }
> >
> > -static int dummy_mmap_prepare(struct vm_area_desc *desc)
> > -{
> > -	return 0;
> > -}
> > -
> > -static bool test_pure_anon_dev_zero(void)
> > -{
> > -	const vma_flags_t vma_flags = mk_vma_flags(VMA_READ_BIT, VMA_WRITE_BIT,
> > -			VMA_MAYREAD_BIT, VMA_MAYWRITE_BIT);
> > -	const struct file_operations f_op = {
> > -		.mmap_prepare = dummy_mmap_prepare,
> > -	};
> > -	struct inode inode = {
> > -		.i_mode = S_IFCHR,
> > -		.i_rdev = MKDEV(MEM_MAJOR, DEVZERO_MINOR),
> > -	};
> > -	struct file file = {
> > -		.f_inode = &inode,
> > -		.f_op = &f_op,
> > -	};
> > -	struct mm_struct mm = {};
> > -	struct vm_area_struct *vma;
> > -	unsigned long addr;
> > -	VMA_ITERATOR(vmi, &mm, 0);
> > -
> > -	current->mm = &mm;
> > -
> > -	/*
> > -	 * Map a MAP_PRIVATE-/dev/zero mapping at address 0x300000 with a page
> > -	 * offset of 0x10, which we expect to be reset to the anonymous page
> > -	 * offset.
> > -	 */
> > -	addr = __mmap_region(&file, 0x300000, 0x3000, vma_flags, 0x10, NULL);
> > -	ASSERT_EQ(addr, 0x300000);
> > -
> > -	/* Assert that it truly is an anonymous mapping. */
> > -	vma = vma_lookup(&mm, addr);
> > -	ASSERT_NE(vma, NULL);
> > -	ASSERT_TRUE(vma_is_anonymous(vma));
> > -	ASSERT_EQ(vma->vm_file, NULL);
> > -	ASSERT_EQ(vma->vm_private_data, NULL);
> > -	/* Expect anonymous page offsets. */
> > -	ASSERT_EQ(vma->vm_pgoff, 0x300);
> > -	ASSERT_EQ(vma_start_anon_pgoff(vma), 0x300);
> > -
> > -	cleanup_mm(&mm, &vmi);
> > -	return true;
> > -}
> > -
> >  static void run_mmap_tests(int *num_tests, int *num_fail)
> >  {
> >  	TEST(mmap_region_basic);
> > -	TEST(pure_anon_dev_zero);
> >  }
> > _
> >
> 
> --
> Cheers, Lorenzo
fixing sashiko failure to apply (was Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff)
Posted by Lorenzo Stoakes (ARM) 1 month, 2 weeks ago
[ trim cc list ]

+cc Roman

On Fri, Aug 14, 2026 at 02:13:51AM -0700, Matthew Brost wrote:
> On Fri, Aug 14, 2026 at 10:01:19AM +0100, Lorenzo Stoakes (ARM) wrote:
> > On Thu, Aug 13, 2026 at 11:53:46AM -0700, Andrew Morton wrote:
> > > You'll be mortified to hear that Sashiko wasn't able to find anything
> > > to which to apply this.
> >
> > :))
> >
> > Well, when it's right it's useful, when it's wrong or suggesting unrelated
> > what-nots it's less useful :>)
> >
>
> Questioning your assumptions is useful, even when they turn out to be wrong.
> Show more lines
>
> > I do locally put things through claude + Chris Mason's prompts a lot, I
> > don't always invoke local sashiko as it's very slow and token-heavy or has
> > been so far, but am planning to do that more also in future.
> >
>
> Yes, it's kind of odd that Sashiko burns more tokens than a full day of
> breakfast, lunch, and dinner service. Running Sashiko is a bottleneck in
> my workflow, so I'll defer to others on this list.

Yup, not sure if there are recommended configs for something saner :)

Maybe Roman has some advice on that?

>
> > >
> > > Sashiko can be guided with a base-commit: tag but I'm not sure how to
> > > tell it what tree/branch to try, or even if that's necessary.  Perhaps
> > > someone can figure this out sometime.
> >
> > b4 gives a base commit, but I think because the trees are rebased it ends
> > up being the incorrect one.
> >
> > Not sure what the solution is!
> >
>
> We have seen this on the Xe list (our list is based on drm-tip),
> typically with cross-subsystem patches. Some cross-subsystem patches
> apply and run correctly, while others do not but public CI flows run
> based on drm-tip. I do not have a bisect or a clear understanding of
> what works and what doesn't, but I think it would be very useful if the
> community could better understand the root cause.

As Mike said, mm-unstable/mm-new is heavily rebased and also carries the old
version of the series before the new one is applied, so it's super unclear what
the base commit should be there.

But in general, I wonder if it's possible that we could tell sashiko
after-the-fact what base commit to look at once the series is in, or re-trigger
it somehow once it's in-tree?

Roman - any suggestions on what we could do to help sashiko find things?

(Once mm-next is in place everything with change again, but can address that
then :)

>
> Matt
>

--
Cheers, Lorenzo
Re: fixing sashiko failure to apply (was Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff)
Posted by Roman Gushchin 1 month, 2 weeks ago

> On Aug 14, 2026, at 11:30 AM, Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote:
> 
> [ trim cc list ]
> 
> +cc Roman
> 
>> On Fri, Aug 14, 2026 at 02:13:51AM -0700, Matthew Brost wrote:
>>> On Fri, Aug 14, 2026 at 10:01:19AM +0100, Lorenzo Stoakes (ARM) wrote:
>>> On Thu, Aug 13, 2026 at 11:53:46AM -0700, Andrew Morton wrote:
>>>> You'll be mortified to hear that Sashiko wasn't able to find anything
>>>> to which to apply this.
>>> 
>>> :))
>>> 
>>> Well, when it's right it's useful, when it's wrong or suggesting unrelated
>>> what-nots it's less useful :>)
>>> 
>> 
>> Questioning your assumptions is useful, even when they turn out to be wrong.
>> Show more lines
>> 
>>> I do locally put things through claude + Chris Mason's prompts a lot, I
>>> don't always invoke local sashiko as it's very slow and token-heavy or has
>>> been so far, but am planning to do that more also in future.
>>> 
>> 
>> Yes, it's kind of odd that Sashiko burns more tokens than a full day of
>> breakfast, lunch, and dinner service. Running Sashiko is a bottleneck in
>> my workflow, so I'll defer to others on this list.
> 
> Yup, not sure if there are recommended configs for something saner :)
> 
> Maybe Roman has some advice on that?

Sorry, no magic way to save tokens without hurting the quality. But I am curious what are your numbers?
Can be model-dependent too. In prod on average it burns 3-4M tokens per patch with Gemini 3.1 Pro, 
but maybe mm patches are more complex than average, Idk.

One option is to run only some discovery stages (—stages), but this unlikely will save you that much.
I’d say use a cheaper and faster model for the development, but it has it’s downsides too.

If you have an example of a patch(set) which is particularly token-hungry, I can take a look.
> 
>> 
>>>> 
>>>> Sashiko can be guided with a base-commit: tag but I'm not sure how to
>>>> tell it what tree/branch to try, or even if that's necessary.  Perhaps
>>>> someone can figure this out sometime.
>>> 
>>> b4 gives a base commit, but I think because the trees are rebased it ends
>>> up being the incorrect one.
>>> 
>>> Not sure what the solution is!
>>> 
>> 
>> We have seen this on the Xe list (our list is based on drm-tip),
>> typically with cross-subsystem patches. Some cross-subsystem patches
>> apply and run correctly, while others do not but public CI flows run
>> based on drm-tip. I do not have a bisect or a clear understanding of
>> what works and what doesn't, but I think it would be very useful if the
>> community could better understand the root cause.
> 
> As Mike said, mm-unstable/mm-new is heavily rebased and also carries the old
> version of the series before the new one is applied, so it's super unclear what
> the base commit should be there.
> 
> But in general, I wonder if it's possible that we could tell sashiko
> after-the-fact what base commit to look at once the series is in, or re-trigger
> it somehow once it's in-tree?
> 
> Roman - any suggestions on what we could do to help sashiko find things?
> 
> (Once mm-next is in place everything with change again, but can address that
> then :)

I can implement any reasonable logic here, the problem is that my understanding is
the current mm process is a bit vague here. Which likely will be also an issue for the mm ci.
I’ll merge a support for b4-like dependencies specification soon.

Re re-starting with manual selection it’s on my todo list, but maybe a bit lfurther away, as it requires
an authorization, etc.

Thanks
Re: fixing sashiko failure to apply (was Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff)
Posted by Lorenzo Stoakes (ARM) 1 month, 2 weeks ago
On Fri, Aug 14, 2026 at 03:53:36PM +0200, Roman Gushchin wrote:
> >> On Fri, Aug 14, 2026 at 02:13:51AM -0700, Matthew Brost wrote:
> >>> On Fri, Aug 14, 2026 at 10:01:19AM +0100, Lorenzo Stoakes (ARM) wrote:
> >>> On Thu, Aug 13, 2026 at 11:53:46AM -0700, Andrew Morton wrote:
> >>>> You'll be mortified to hear that Sashiko wasn't able to find anything
> >>>> to which to apply this.
> >>>
> >>> :))
> >>>
> >>> Well, when it's right it's useful, when it's wrong or suggesting unrelated
> >>> what-nots it's less useful :>)
> >>>
> >>
> >> Questioning your assumptions is useful, even when they turn out to be wrong.
> >> Show more lines
> >>
> >>> I do locally put things through claude + Chris Mason's prompts a lot, I
> >>> don't always invoke local sashiko as it's very slow and token-heavy or has
> >>> been so far, but am planning to do that more also in future.
> >>>
> >>
> >> Yes, it's kind of odd that Sashiko burns more tokens than a full day of
> >> breakfast, lunch, and dinner service. Running Sashiko is a bottleneck in
> >> my workflow, so I'll defer to others on this list.
> >
> > Yup, not sure if there are recommended configs for something saner :)
> >
> > Maybe Roman has some advice on that?
>
> Sorry, no magic way to save tokens without hurting the quality. But I am curious what are your numbers?
> Can be model-dependent too. In prod on average it burns 3-4M tokens per patch with Gemini 3.1 Pro,
> but maybe mm patches are more complex than average, Idk.
>
> One option is to run only some discovery stages (—stages), but this unlikely will save you that much.
> I’d say use a cheaper and faster model for the development, but it has it’s downsides too.
>
> If you have an example of a patch(set) which is particularly token-hungry, I can take a look.

Oh well damn, no 3-4M per patch sounds about right actually. I guess that just
is what it is then!

> >
> >>
> >>>>
> >>>> Sashiko can be guided with a base-commit: tag but I'm not sure how to
> >>>> tell it what tree/branch to try, or even if that's necessary.  Perhaps
> >>>> someone can figure this out sometime.
> >>>
> >>> b4 gives a base commit, but I think because the trees are rebased it ends
> >>> up being the incorrect one.
> >>>
> >>> Not sure what the solution is!
> >>>
> >>
> >> We have seen this on the Xe list (our list is based on drm-tip),
> >> typically with cross-subsystem patches. Some cross-subsystem patches
> >> apply and run correctly, while others do not but public CI flows run
> >> based on drm-tip. I do not have a bisect or a clear understanding of
> >> what works and what doesn't, but I think it would be very useful if the
> >> community could better understand the root cause.
> >
> > As Mike said, mm-unstable/mm-new is heavily rebased and also carries the old
> > version of the series before the new one is applied, so it's super unclear what
> > the base commit should be there.
> >
> > But in general, I wonder if it's possible that we could tell sashiko
> > after-the-fact what base commit to look at once the series is in, or re-trigger
> > it somehow once it's in-tree?
> >
> > Roman - any suggestions on what we could do to help sashiko find things?
> >
> > (Once mm-next is in place everything with change again, but can address that
> > then :)
>
> I can implement any reasonable logic here, the problem is that my understanding is
> the current mm process is a bit vague here. Which likely will be also an issue for the mm ci.
> I’ll merge a support for b4-like dependencies specification soon.

Yeah I suspect things might be tricky with mm given the rebases honestly.

>
> Re re-starting with manual selection it’s on my todo list, but maybe a bit lfurther away, as it requires
> an authorization, etc.

Yeah that's the fly in the ointment I guess for many things like giving instant
feedback on accuracy, well you want to make sure the person giving it is who you
think they are :)

Probably an email -> author with magic link or something but thinking through
how to avoid abuse/spam/rate limiting everything etc. is surely all a pain :)

Good to hear it's the TODO list though!

>
> Thanks

--
Cheers, Lorenzo
Re: fixing sashiko failure to apply (was Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff)
Posted by Zi Yan 1 month, 2 weeks ago
<snip>
>> >>
>> >>>>
>> >>>> Sashiko can be guided with a base-commit: tag but I'm not sure how to
>> >>>> tell it what tree/branch to try, or even if that's necessary.  Perhaps
>> >>>> someone can figure this out sometime.
>> >>>
>> >>> b4 gives a base commit, but I think because the trees are rebased it ends
>> >>> up being the incorrect one.
>> >>>
>> >>> Not sure what the solution is!
>> >>>
>> >>
>> >> We have seen this on the Xe list (our list is based on drm-tip),
>> >> typically with cross-subsystem patches. Some cross-subsystem patches
>> >> apply and run correctly, while others do not but public CI flows run
>> >> based on drm-tip. I do not have a bisect or a clear understanding of
>> >> what works and what doesn't, but I think it would be very useful if the
>> >> community could better understand the root cause.
>> >
>> > As Mike said, mm-unstable/mm-new is heavily rebased and also carries the old
>> > version of the series before the new one is applied, so it's super unclear what
>> > the base commit should be there.
>> >
>> > But in general, I wonder if it's possible that we could tell sashiko
>> > after-the-fact what base commit to look at once the series is in, or re-trigger
>> > it somehow once it's in-tree?
>> >
>> > Roman - any suggestions on what we could do to help sashiko find things?
>> >
>> > (Once mm-next is in place everything with change again, but can address that
>> > then :)
>>
>> I can implement any reasonable logic here, the problem is that my understanding is
>> the current mm process is a bit vague here. Which likely will be also an issue for the mm ci.
>> I’ll merge a support for b4-like dependencies specification soon.
>
> Yeah I suspect things might be tricky with mm given the rebases honestly.
>
>>
>> Re re-starting with manual selection it’s on my todo list, but maybe a bit lfurther away, as it requires
>> an authorization, etc.
>
> Yeah that's the fly in the ointment I guess for many things like giving instant
> feedback on accuracy, well you want to make sure the person giving it is who you
> think they are :)
>
> Probably an email -> author with magic link or something but thinking through
> how to avoid abuse/spam/rate limiting everything etc. is surely all a pain :)

If b4 is used, the cover letter comes with a fixed change-id. That can
be a good id for new patch series replacement. For patches do not use
b4, authors can add their own change-id. Without change-id, if the new
version comes with a link to the prior version, that can be helpful too.
If none is present, author email + patch title might be the last resort.
But I think we need to encourage authors to provide some id for their
patches to make replacement easier instead of spending too much effort
on identifying different patch versions.

-- 
Best Regards,
Yan, Zi
Re: fixing sashiko failure to apply (was Re: [PATCH v5 00/16] mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff)
Posted by Lorenzo Stoakes (ARM) 1 month, 1 week ago
On Fri, Aug 14, 2026 at 10:19:10PM -0400, Zi Yan wrote:
> <snip>
> >> >>
> >> >>>>
> >> >>>> Sashiko can be guided with a base-commit: tag but I'm not sure how to
> >> >>>> tell it what tree/branch to try, or even if that's necessary.  Perhaps
> >> >>>> someone can figure this out sometime.
> >> >>>
> >> >>> b4 gives a base commit, but I think because the trees are rebased it ends
> >> >>> up being the incorrect one.
> >> >>>
> >> >>> Not sure what the solution is!
> >> >>>
> >> >>
> >> >> We have seen this on the Xe list (our list is based on drm-tip),
> >> >> typically with cross-subsystem patches. Some cross-subsystem patches
> >> >> apply and run correctly, while others do not but public CI flows run
> >> >> based on drm-tip. I do not have a bisect or a clear understanding of
> >> >> what works and what doesn't, but I think it would be very useful if the
> >> >> community could better understand the root cause.
> >> >
> >> > As Mike said, mm-unstable/mm-new is heavily rebased and also carries the old
> >> > version of the series before the new one is applied, so it's super unclear what
> >> > the base commit should be there.
> >> >
> >> > But in general, I wonder if it's possible that we could tell sashiko
> >> > after-the-fact what base commit to look at once the series is in, or re-trigger
> >> > it somehow once it's in-tree?
> >> >
> >> > Roman - any suggestions on what we could do to help sashiko find things?
> >> >
> >> > (Once mm-next is in place everything with change again, but can address that
> >> > then :)
> >>
> >> I can implement any reasonable logic here, the problem is that my understanding is
> >> the current mm process is a bit vague here. Which likely will be also an issue for the mm ci.
> >> I’ll merge a support for b4-like dependencies specification soon.
> >
> > Yeah I suspect things might be tricky with mm given the rebases honestly.
> >
> >>
> >> Re re-starting with manual selection it’s on my todo list, but maybe a bit lfurther away, as it requires
> >> an authorization, etc.
> >
> > Yeah that's the fly in the ointment I guess for many things like giving instant
> > feedback on accuracy, well you want to make sure the person giving it is who you
> > think they are :)
> >
> > Probably an email -> author with magic link or something but thinking through
> > how to avoid abuse/spam/rate limiting everything etc. is surely all a pain :)
>
> If b4 is used, the cover letter comes with a fixed change-id. That can
> be a good id for new patch series replacement. For patches do not use
> b4, authors can add their own change-id. Without change-id, if the new
> version comes with a link to the prior version, that can be helpful too.
> If none is present, author email + patch title might be the last resort.
> But I think we need to encourage authors to provide some id for their
> patches to make replacement easier instead of spending too much effort
> on identifying different patch versions.

Agree, but with mm constantly rebasing these hashes become potentially less
useful :)

>
> --
> Best Regards,
> Yan, Zi
>

--
Cheers, Lorenzo