[PATCH v4 0/7] read proc/pid/smaps_rollup under per-vma lock

Suren Baghdasaryan posted 7 patches 1 week, 6 days ago
There is a newer version of this series
fs/proc/task_mmu.c                            | 339 ++++++++----------
tools/testing/selftests/proc/proc-maps-race.c | 186 +++++++++-
2 files changed, 334 insertions(+), 191 deletions(-)
[PATCH v4 0/7] read proc/pid/smaps_rollup under per-vma lock
Posted by Suren Baghdasaryan 1 week, 6 days ago
proc/pid/smaps_rollup can be read using the combination of RCU and
VMA read locks, similar to proc/pid/{maps|smaps|numa_maps}. RCU is
required to safely traverse the VMA tree and VMA lock stabilizes the
VMA being processed and the pagetable walk.
Note that we have to keep the logic to drop mmap_lock on contention
because even when using per-VMA locks we might have to fall back to
holding the mmap_lock.

The first 5 patches are cleanups making later change simpler. The main
change is in patch 6. Patch 7 extends existing proc-maps-race tearing
test to verify smaps_rollup content.

Changes since v3 [1]:
Patch 2:
- Added Acked-by, per David Hildenbrand
Patch 3:
- Moved conversion of (start != 0) to is_partial into the next patch,
  per Lorenzo Stoakes
Patch 4:
- Moved conversion of (start != 0) to is_partial into the next patch,
  per Lorenzo Stoakes
- Converted multiple lines of arguments in modified functions to
  two-tab indents, per David Hildenbrand
- Introduced smap_gather_stats_range() for partial VMA walks,
  per Lorenzo Stoakes
Patch 5:
- Added Reviewed-by and Acked-by, per David Hildenbrand and
  Lorenzo Stoakes
- Added a comment explaining the check for SENTINEL_VMA_GATE in m_start,
  per Lorenzo Stoakes
Patch 6:
- Added Reviewed-by, per Lorenzo Stoakes
Patch 7:
- Converted multiple lines of arguments in modified functions to
  two-tab indents
- Added Acked-by, per Lorenzo Stoakes

[1] https://lore.kernel.org/all/20260910234737.1340642-1-surenb@google.com/

Patchset applies over mm-new branch.

Suren Baghdasaryan (7):
  proc/task_mmu: remove unnecessary helpers
  proc/task_mmu: remove unnecessary inlines in function definitions
  proc/task_mmu: clarify shmem mapping walk conditions in
    smap_gather_stats()
  proc/task_mmu: remove special-casing of smap_gather_stats() start
    parameter
  proc/task_mmu: change proc_get_vma() to stop returning gate VMA at the
    end
  proc/task_mmu: read proc/pid/smaps_rollup under per-vma lock
  selftests/proc: add /proc/pid/smaps_rollup tearing tests

 fs/proc/task_mmu.c                            | 339 ++++++++----------
 tools/testing/selftests/proc/proc-maps-race.c | 186 +++++++++-
 2 files changed, 334 insertions(+), 191 deletions(-)


base-commit: 1f78f28a2945f9e856b4ea8428ac41a6788e4b67
-- 
2.55.0.1007.g17ff1f9808-goog
Re: [PATCH v4 0/7] read proc/pid/smaps_rollup under per-vma lock
Posted by Andrew Morton 1 week, 6 days ago
On Fri, 11 Sep 2026 12:41:38 -0700 Suren Baghdasaryan <surenb@google.com> wrote:

> proc/pid/smaps_rollup can be read using the combination of RCU and
> VMA read locks, similar to proc/pid/{maps|smaps|numa_maps}. RCU is
> required to safely traverse the VMA tree and VMA lock stabilizes the
> VMA being processed and the pagetable walk.
> Note that we have to keep the logic to drop mmap_lock on contention
> because even when using per-VMA locks we might have to fall back to
> holding the mmap_lock.

Nice.  Queued, thanks.

The speedups described in [6/7] are significant, although I don't know
how representative Paul's tests are.  Probably not very.

So I don't know how much improvement our users will be seeing from
these changes?

And I don't know how much usage smaps_rollup gets in the real world?


As I've no doubt you know, Sashiko is saying things.  About these
patches and about the current code.  Its pre-existing
clear_soft_dirty_pmd() driveby find looks significant.

	https://sashiko.dev/#/patchset/20260911194145.1781926-1-surenb@google.com
Re: [PATCH v4 0/7] read proc/pid/smaps_rollup under per-vma lock
Posted by Suren Baghdasaryan 1 week, 4 days ago
On Sat, Sep 12, 2026 at 12:24 AM Andrew Morton
<akpm@linux-foundation.org> wrote:
>
> On Fri, 11 Sep 2026 12:41:38 -0700 Suren Baghdasaryan <surenb@google.com> wrote:
>
> > proc/pid/smaps_rollup can be read using the combination of RCU and
> > VMA read locks, similar to proc/pid/{maps|smaps|numa_maps}. RCU is
> > required to safely traverse the VMA tree and VMA lock stabilizes the
> > VMA being processed and the pagetable walk.
> > Note that we have to keep the logic to drop mmap_lock on contention
> > because even when using per-VMA locks we might have to fall back to
> > holding the mmap_lock.
>
> Nice.  Queued, thanks.
>
> The speedups described in [6/7] are significant, although I don't know
> how representative Paul's tests are.  Probably not very.
>
> So I don't know how much improvement our users will be seeing from
> these changes?
>
> And I don't know how much usage smaps_rollup gets in the real world?
>
>
> As I've no doubt you know, Sashiko is saying things.  About these
> patches and about the current code.  Its pre-existing
> clear_soft_dirty_pmd() driveby find looks significant.
>
>         https://sashiko.dev/#/patchset/20260911194145.1781926-1-surenb@google.com

Thanks! I'll investigate the preexisting ones and will try to fix the
ones that look real.
Re: [PATCH v4 0/7] read proc/pid/smaps_rollup under per-vma lock
Posted by Paul E. McKenney 1 week, 5 days ago
On Sat, Sep 12, 2026 at 12:24:36AM -0700, Andrew Morton wrote:
> On Fri, 11 Sep 2026 12:41:38 -0700 Suren Baghdasaryan <surenb@google.com> wrote:
> 
> > proc/pid/smaps_rollup can be read using the combination of RCU and
> > VMA read locks, similar to proc/pid/{maps|smaps|numa_maps}. RCU is
> > required to safely traverse the VMA tree and VMA lock stabilizes the
> > VMA being processed and the pagetable walk.
> > Note that we have to keep the logic to drop mmap_lock on contention
> > because even when using per-VMA locks we might have to fall back to
> > holding the mmap_lock.
> 
> Nice.  Queued, thanks.
> 
> The speedups described in [6/7] are significant, although I don't know
> how representative Paul's tests are.  Probably not very.
> 
> So I don't know how much improvement our users will be seeing from
> these changes?
> 
> And I don't know how much usage smaps_rollup gets in the real world?

It is used heavily in some applications for monitoring.  This should allow
us to more tightly fence our monitoring software, leaving more CPU for
the application.

So thank you all!!!

							Thanx, Paul

> As I've no doubt you know, Sashiko is saying things.  About these
> patches and about the current code.  Its pre-existing
> clear_soft_dirty_pmd() driveby find looks significant.
> 
> 	https://sashiko.dev/#/patchset/20260911194145.1781926-1-surenb@google.com