[PATCH v3 0/5] mm: Unconditional per-VMA locks and cleanups

Suren Baghdasaryan posted 5 patches 2 months ago
There is a newer version of this series
arch/arm/Kconfig                       |  1 -
arch/arm64/Kconfig                     |  1 -
arch/loongarch/Kconfig                 |  1 -
arch/powerpc/platforms/powernv/Kconfig |  1 -
arch/powerpc/platforms/pseries/Kconfig |  1 -
arch/riscv/Kconfig                     |  1 -
arch/s390/Kconfig                      |  1 -
arch/x86/Kconfig                       |  2 -
drivers/android/binder/page_range.rs   | 19 +-----
drivers/android/binder_alloc.c         | 46 ++++++-------
fs/proc/internal.h                     |  2 -
fs/proc/task_mmu.c                     | 93 --------------------------
include/linux/mm.h                     | 12 ----
include/linux/mm_types.h               |  8 +--
include/linux/mmap_lock.h              | 65 +++---------------
kernel/bpf/stackmap.c                  | 16 +----
kernel/bpf/task_iter.c                 |  5 --
kernel/fork.c                          |  2 -
mm/Kconfig                             | 13 ----
mm/Kconfig.debug                       |  1 -
mm/debug.c                             |  4 --
mm/init-mm.c                           |  2 -
mm/memory.c                            |  2 -
mm/mmap_lock.c                         | 53 ++++++++-------
mm/pagewalk.c                          |  2 -
mm/rmap.c                              |  2 -
mm/userfaultfd.c                       | 61 ++---------------
net/ipv4/tcp.c                         | 31 +++------
rust/kernel/mm.rs                      | 38 +++++++----
tools/testing/vma/include/dup.h        |  5 +-
tools/testing/vma/vma_internal.h       |  1 -
31 files changed, 106 insertions(+), 386 deletions(-)
[PATCH v3 0/5] mm: Unconditional per-VMA locks and cleanups
Posted by Suren Baghdasaryan 2 months ago
v2 version of this patchset [1] was written by Dave Hansen and per his
request, I'm taking over this series.

tl;dr: Make per-VMA locks available in all configs. Simplify some
of the per-VMA lock users now that they can rely on them being
always available.

Binder and networking folks: Your code is the target of the cleanups.
I'm cc'ing you now on v2 because there's emerging consensus on the mm
side that the approach here is sane. I'm not quite sure how this pile
would get merged, but ack/review tags would be appreciated if this
looks good to you.

Longer version:

When working on some x86 shadow stack code, it was a real pain to
avoid causing recursive locking problems with mmap_lock. One way
to avoid those was to avoid mmap_lock and use per-VMA locks instead.
They are great, but they are not available in all configs which
makes them unusable in generic code, or if you want to completely
avoid mmap_lock.

Make per-VMA locks available in all configs. Right now, they are
only available on select architectures when SMP and MMU are enabled.
But all of the primitives that per-VMA locks are built on (RCU, maple
trees, refcounts) work just fine without SMP or MMU.

The only real downside is that making VMAs a wee bit bigger on !MMU
and !SMP builds.

The upside is much cleaner code, lower complexity and less #ifdeffery.

Clean up a binder VMA locking site now that it can rely on per-VMA
locks.

Building on top of universally-available per-VMA locks, introduce a
new helper. Since the new API does not require callers to have a
fallback to mmap_lock, it's much easier to use. Callers can
potentially replace this very common kernel idiom:

	mmap_read_lock(mm);
	vma = vma_lookup()
	// fiddle with vma
	mmap_read_unlock(mm);

with:

	vma = vma_start_read_unlocked(mm, address);
	// fiddle with vma
	vma_end_read(vma);

Which avoids mmap_lock entirely in the fast path.

Use that new API for another binder site and one in the TCP code.

Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: "Liam R. Howlett" <Liam.Howlett@oracle.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: linux-mm@kvack.org
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Arve Hjønnevåg <arve@android.com>
Cc: Todd Kjos <tkjos@android.com>
Cc: Christian Brauner <christian@brauner.io>
Cc: Carlos Llamas <cmllamas@google.com>
Cc: Alice Ryhl <aliceryhl@google.com>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: David Ahern <dsahern@kernel.org>
Cc: netdev@vger.kernel.org

Changes from v2 [1]:
Patch 1
- Removed new CONFIG_PER_VMA_LOCK usage in kernel/bpf/stackmap.c
- Modified lock_vma_under_rcu Rust implementation, per Alice Ryhl
Patch 2
- Added binder_alloc_is_mapped() check to allow page freeing if the vma is
unmapped, per Alice Ryhl and Carlos Llamas
- Removed earlier Reviewed-by's and Acked-by's
Patch 3
- Modified comments for vma_start_read_locked_nested(),
vma_start_read_locked() and uffd_lock_vma(), per Vlastimil Babka
- Updated vma_start_read_unlocked() to call vma_lookup(), per Dave Hansen
- Added a check for VMA to be valid before calling
vma_start_read_locked(), per Vlastimil Babka
- Replaced guard with mmap_read_lock/mmap_read_unlock
- Modified the comment to indicate mmap_locking is temporary,
per Vlastimil Babka
Patch 4
- Removed earlier Reviewed-by's and Acked-by's
- Added Rust version of vma_start_read_unlocked() and used it in Rust
version of the binder driver, per Alice Ryhl
Patch 5
- Replaced lock_vma_under_rcu_wait() with vma_start_read_unlocked() in the
changelog, per Vlastimil Babka

Applies cleanly over mm-unstable

[1] https://lore.kernel.org/all/20260610230409.A44D29FA@davehans-spike.ostc.intel.com/

Dave Hansen (5):
  mm: Make per-VMA locks available universally
  binder: Make shrinker rely solely on per-VMA lock
  mm: Add RCU-based VMA lookup helper that waits for writers
  binder: Remove mmap_lock fallback
  tcp: Remove mmap_lock fallback path

 arch/arm/Kconfig                       |  1 -
 arch/arm64/Kconfig                     |  1 -
 arch/loongarch/Kconfig                 |  1 -
 arch/powerpc/platforms/powernv/Kconfig |  1 -
 arch/powerpc/platforms/pseries/Kconfig |  1 -
 arch/riscv/Kconfig                     |  1 -
 arch/s390/Kconfig                      |  1 -
 arch/x86/Kconfig                       |  2 -
 drivers/android/binder/page_range.rs   | 19 +-----
 drivers/android/binder_alloc.c         | 46 ++++++-------
 fs/proc/internal.h                     |  2 -
 fs/proc/task_mmu.c                     | 93 --------------------------
 include/linux/mm.h                     | 12 ----
 include/linux/mm_types.h               |  8 +--
 include/linux/mmap_lock.h              | 65 +++---------------
 kernel/bpf/stackmap.c                  | 16 +----
 kernel/bpf/task_iter.c                 |  5 --
 kernel/fork.c                          |  2 -
 mm/Kconfig                             | 13 ----
 mm/Kconfig.debug                       |  1 -
 mm/debug.c                             |  4 --
 mm/init-mm.c                           |  2 -
 mm/memory.c                            |  2 -
 mm/mmap_lock.c                         | 53 ++++++++-------
 mm/pagewalk.c                          |  2 -
 mm/rmap.c                              |  2 -
 mm/userfaultfd.c                       | 61 ++---------------
 net/ipv4/tcp.c                         | 31 +++------
 rust/kernel/mm.rs                      | 38 +++++++----
 tools/testing/vma/include/dup.h        |  5 +-
 tools/testing/vma/vma_internal.h       |  1 -
 31 files changed, 106 insertions(+), 386 deletions(-)


base-commit: 94f9b3980dd446b56acf1dfed649e9b32a9f3813
-- 
2.55.0.508.g3f0d502094-goog
Re: [PATCH v3 0/5] mm: Unconditional per-VMA locks and cleanups
Posted by Barry Song 2 months ago
On Mon, Aug 3, 2026 at 5:58 AM Suren Baghdasaryan <surenb@google.com> wrote:
>
> v2 version of this patchset [1] was written by Dave Hansen and per his
> request, I'm taking over this series.
>
> tl;dr: Make per-VMA locks available in all configs. Simplify some
> of the per-VMA lock users now that they can rely on them being
> always available.
>
> Binder and networking folks: Your code is the target of the cleanups.
> I'm cc'ing you now on v2 because there's emerging consensus on the mm
> side that the approach here is sane. I'm not quite sure how this pile
> would get merged, but ack/review tags would be appreciated if this
> looks good to you.
>
> Longer version:
>
> When working on some x86 shadow stack code, it was a real pain to
> avoid causing recursive locking problems with mmap_lock. One way
> to avoid those was to avoid mmap_lock and use per-VMA locks instead.
> They are great, but they are not available in all configs which
> makes them unusable in generic code, or if you want to completely
> avoid mmap_lock.
>
> Make per-VMA locks available in all configs. Right now, they are
> only available on select architectures when SMP and MMU are enabled.
> But all of the primitives that per-VMA locks are built on (RCU, maple
> trees, refcounts) work just fine without SMP or MMU.
>
> The only real downside is that making VMAs a wee bit bigger on !MMU
> and !SMP builds.
>
> The upside is much cleaner code, lower complexity and less #ifdeffery.
>
> Clean up a binder VMA locking site now that it can rely on per-VMA
> locks.
>
> Building on top of universally-available per-VMA locks, introduce a
> new helper. Since the new API does not require callers to have a
> fallback to mmap_lock, it's much easier to use. Callers can
> potentially replace this very common kernel idiom:
>
>         mmap_read_lock(mm);
>         vma = vma_lookup()
>         // fiddle with vma
>         mmap_read_unlock(mm);
>
> with:
>
>         vma = vma_start_read_unlocked(mm, address);
>         // fiddle with vma
>         vma_end_read(vma);
>
> Which avoids mmap_lock entirely in the fast path.
>
> Use that new API for another binder site and one in the TCP code.

Nice, Suren and Dave.

I wonder if we could use the same approach in the page fault
path. Instead of falling back to mmap_lock when
lock_vma_under_rcu() fails the first time, could we wait for the
writer to finish and then retry acquiring the VMA lock?
For example:

diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c
index 85e23388f9bb..684f38cc4e74 100644
--- a/arch/arm64/mm/fault.c
+++ b/arch/arm64/mm/fault.c
@@ -677,7 +677,7 @@ static int __kprobes do_page_fault(unsigned long
far, unsigned long esr,
        if (!(mm_flags & FAULT_FLAG_USER))
                goto lock_mmap;

-       vma = lock_vma_under_rcu(mm, addr);
+       vma = vma_start_read_unlocked(mm, addr);
        if (!vma)
                goto lock_mmap;

diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
index 45b99c3b1442..a3a4c4741e30 100644
--- a/arch/x86/mm/fault.c
+++ b/arch/x86/mm/fault.c
@@ -1331,7 +1331,7 @@ void do_user_addr_fault(struct pt_regs *regs,
        if (!(flags & FAULT_FLAG_USER))
                goto lock_mmap;

-       vma = lock_vma_under_rcu(mm, address);
+       vma = vma_start_read_unlocked(mm, address);
        if (!vma)
                goto lock_mmap;

Best Regards
Barry
Re: [PATCH v3 0/5] mm: Unconditional per-VMA locks and cleanups
Posted by Suren Baghdasaryan 1 month, 4 weeks ago
On Sun, Aug 2, 2026 at 7:11 PM Barry Song <baohua@kernel.org> wrote:
>
> On Mon, Aug 3, 2026 at 5:58 AM Suren Baghdasaryan <surenb@google.com> wrote:
> >
> > v2 version of this patchset [1] was written by Dave Hansen and per his
> > request, I'm taking over this series.
> >
> > tl;dr: Make per-VMA locks available in all configs. Simplify some
> > of the per-VMA lock users now that they can rely on them being
> > always available.
> >
> > Binder and networking folks: Your code is the target of the cleanups.
> > I'm cc'ing you now on v2 because there's emerging consensus on the mm
> > side that the approach here is sane. I'm not quite sure how this pile
> > would get merged, but ack/review tags would be appreciated if this
> > looks good to you.
> >
> > Longer version:
> >
> > When working on some x86 shadow stack code, it was a real pain to
> > avoid causing recursive locking problems with mmap_lock. One way
> > to avoid those was to avoid mmap_lock and use per-VMA locks instead.
> > They are great, but they are not available in all configs which
> > makes them unusable in generic code, or if you want to completely
> > avoid mmap_lock.
> >
> > Make per-VMA locks available in all configs. Right now, they are
> > only available on select architectures when SMP and MMU are enabled.
> > But all of the primitives that per-VMA locks are built on (RCU, maple
> > trees, refcounts) work just fine without SMP or MMU.
> >
> > The only real downside is that making VMAs a wee bit bigger on !MMU
> > and !SMP builds.
> >
> > The upside is much cleaner code, lower complexity and less #ifdeffery.
> >
> > Clean up a binder VMA locking site now that it can rely on per-VMA
> > locks.
> >
> > Building on top of universally-available per-VMA locks, introduce a
> > new helper. Since the new API does not require callers to have a
> > fallback to mmap_lock, it's much easier to use. Callers can
> > potentially replace this very common kernel idiom:
> >
> >         mmap_read_lock(mm);
> >         vma = vma_lookup()
> >         // fiddle with vma
> >         mmap_read_unlock(mm);
> >
> > with:
> >
> >         vma = vma_start_read_unlocked(mm, address);
> >         // fiddle with vma
> >         vma_end_read(vma);
> >
> > Which avoids mmap_lock entirely in the fast path.
> >
> > Use that new API for another binder site and one in the TCP code.
>
> Nice, Suren and Dave.
>
> I wonder if we could use the same approach in the page fault
> path. Instead of falling back to mmap_lock when
> lock_vma_under_rcu() fails the first time, could we wait for the
> writer to finish and then retry acquiring the VMA lock?

Yeah, we might be able to do that. Matthew is working on moving common
page-fault handling code into a single arch-independent place. Your
suggested change would be simpler if done after Matthew's refactoring.

> For example:
>
> diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c
> index 85e23388f9bb..684f38cc4e74 100644
> --- a/arch/arm64/mm/fault.c
> +++ b/arch/arm64/mm/fault.c
> @@ -677,7 +677,7 @@ static int __kprobes do_page_fault(unsigned long
> far, unsigned long esr,
>         if (!(mm_flags & FAULT_FLAG_USER))
>                 goto lock_mmap;
>
> -       vma = lock_vma_under_rcu(mm, addr);
> +       vma = vma_start_read_unlocked(mm, addr);
>         if (!vma)
>                 goto lock_mmap;
>
> diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
> index 45b99c3b1442..a3a4c4741e30 100644
> --- a/arch/x86/mm/fault.c
> +++ b/arch/x86/mm/fault.c
> @@ -1331,7 +1331,7 @@ void do_user_addr_fault(struct pt_regs *regs,
>         if (!(flags & FAULT_FLAG_USER))
>                 goto lock_mmap;
>
> -       vma = lock_vma_under_rcu(mm, address);
> +       vma = vma_start_read_unlocked(mm, address);
>         if (!vma)
>                 goto lock_mmap;
>
> Best Regards
> Barry
Re: [PATCH v3 0/5] mm: Unconditional per-VMA locks and cleanups
Posted by Lorenzo Stoakes (ARM) 1 month, 4 weeks ago
On Mon, Aug 03, 2026 at 10:51:25AM -0700, Suren Baghdasaryan wrote:
> On Sun, Aug 2, 2026 at 7:11 PM Barry Song <baohua@kernel.org> wrote:
> >
> > On Mon, Aug 3, 2026 at 5:58 AM Suren Baghdasaryan <surenb@google.com> wrote:
> > >
> > > v2 version of this patchset [1] was written by Dave Hansen and per his
> > > request, I'm taking over this series.
> > >
> > > tl;dr: Make per-VMA locks available in all configs. Simplify some
> > > of the per-VMA lock users now that they can rely on them being
> > > always available.
> > >
> > > Binder and networking folks: Your code is the target of the cleanups.
> > > I'm cc'ing you now on v2 because there's emerging consensus on the mm
> > > side that the approach here is sane. I'm not quite sure how this pile
> > > would get merged, but ack/review tags would be appreciated if this
> > > looks good to you.
> > >
> > > Longer version:
> > >
> > > When working on some x86 shadow stack code, it was a real pain to
> > > avoid causing recursive locking problems with mmap_lock. One way
> > > to avoid those was to avoid mmap_lock and use per-VMA locks instead.
> > > They are great, but they are not available in all configs which
> > > makes them unusable in generic code, or if you want to completely
> > > avoid mmap_lock.
> > >
> > > Make per-VMA locks available in all configs. Right now, they are
> > > only available on select architectures when SMP and MMU are enabled.
> > > But all of the primitives that per-VMA locks are built on (RCU, maple
> > > trees, refcounts) work just fine without SMP or MMU.
> > >
> > > The only real downside is that making VMAs a wee bit bigger on !MMU
> > > and !SMP builds.
> > >
> > > The upside is much cleaner code, lower complexity and less #ifdeffery.
> > >
> > > Clean up a binder VMA locking site now that it can rely on per-VMA
> > > locks.
> > >
> > > Building on top of universally-available per-VMA locks, introduce a
> > > new helper. Since the new API does not require callers to have a
> > > fallback to mmap_lock, it's much easier to use. Callers can
> > > potentially replace this very common kernel idiom:
> > >
> > >         mmap_read_lock(mm);
> > >         vma = vma_lookup()
> > >         // fiddle with vma
> > >         mmap_read_unlock(mm);
> > >
> > > with:
> > >
> > >         vma = vma_start_read_unlocked(mm, address);
> > >         // fiddle with vma
> > >         vma_end_read(vma);
> > >
> > > Which avoids mmap_lock entirely in the fast path.
> > >
> > > Use that new API for another binder site and one in the TCP code.
> >
> > Nice, Suren and Dave.

And Lorenzo ;)

> >
> > I wonder if we could use the same approach in the page fault
> > path. Instead of falling back to mmap_lock when
> > lock_vma_under_rcu() fails the first time, could we wait for the
> > writer to finish and then retry acquiring the VMA lock?
>
> Yeah, we might be able to do that. Matthew is working on moving common

I think it could really help clean things up actually.

Though obviously you still need to account for the fault retry stuff that still
needs to fall back to mmap lock, so sadly we DO need a fallback path (ugh).

It'd be good to remove some of the duplication.

> page-fault handling code into a single arch-independent place. Your
> suggested change would be simpler if done after Matthew's refactoring.

Yeah might be worth waiting for that :)

>
> > For example:
> >
> > diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c
> > index 85e23388f9bb..684f38cc4e74 100644
> > --- a/arch/arm64/mm/fault.c
> > +++ b/arch/arm64/mm/fault.c
> > @@ -677,7 +677,7 @@ static int __kprobes do_page_fault(unsigned long
> > far, unsigned long esr,
> >         if (!(mm_flags & FAULT_FLAG_USER))
> >                 goto lock_mmap;
> >
> > -       vma = lock_vma_under_rcu(mm, addr);
> > +       vma = vma_start_read_unlocked(mm, addr);
> >         if (!vma)
> >                 goto lock_mmap;
> >
> > diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c
> > index 45b99c3b1442..a3a4c4741e30 100644
> > --- a/arch/x86/mm/fault.c
> > +++ b/arch/x86/mm/fault.c
> > @@ -1331,7 +1331,7 @@ void do_user_addr_fault(struct pt_regs *regs,
> >         if (!(flags & FAULT_FLAG_USER))
> >                 goto lock_mmap;
> >
> > -       vma = lock_vma_under_rcu(mm, address);
> > +       vma = vma_start_read_unlocked(mm, address);
> >         if (!vma)
> >                 goto lock_mmap;
> >
> > Best Regards
> > Barry

--
Cheers, Lorenzo