From nobody Sat Sep 26 13:47:00 2026 Received: from mail-pg1-f198.google.com (mail-pg1-f198.google.com [209.85.215.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B684038B142 for ; Mon, 31 Aug 2026 20:31:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208267; cv=none; b=L4ZLQxfg7MbjsHh/EAPO4jJQ+QS4Ss35TrRpAuYFd8ms9kEci/m52rxU9JU4/yJOj5x6q9ysKeTBkNTSmgjpLVXrvEUFTTOzK6qLTOjZT5oc5Mv4tEDYqnaqqQWcTEtBcZSt6J1XgpMPJwQR776K4quAJEsqr7kE4YjyuyBU+IU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208267; c=relaxed/simple; bh=Ouv2uOgETbMRJXR8YkrtDs0dcskv5u5kBsaCdK4/3xk=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=oPUkZmXLSLICJJhiGeoboconrmhvh+CiB5rz6qTiGQ4jRY1Urn7KBKRf3VFNyIyHr6MUxj9zWEpIfpLur/gmtYvq+uBvok86h91IHNbYGEaZFxv3W4+rw3VGKhiibn02vaUE9BcBSR/S/W5HQmEjsf8KF8KvIpk8vgFCCFnMqjo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Rf6c+ur5; arc=none smtp.client-ip=209.85.215.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Rf6c+ur5" Received: by mail-pg1-f198.google.com with SMTP id 41be03b00d2f7-cbedb8673ceso174849a12.0 for ; Mon, 31 Aug 2026 13:31:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788208264; x=1788813064; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=MD955j7svwcecXFODkmvNY3vK8UjCpMQPU48awBPMVA=; b=Rf6c+ur5FaUAgN12pDIBmtVUvYNgjRN3m2uhfl2CNBzB3jE3aO5/BMuIiXF0mmXzSZ wDSNyRF7AUNkhUXnwDWeoHyOMWCgNs7vFH1I/8u2MEFaq0Fy1nzM249n8doHvRHKn/F9 wyd/75f3LHTaN/brxcTnfxpugFTNMvWZYvWEllr54o6MyLA5JBDkIfJsoxEDR9EmhANp EIHt5NaNhLVMIWD4D7sdtvIePxyH0qXXznrItKOhsigNhT+KoY9prpnx9IIt9SIAbQhf +pNsxJpCRyZ2velNEgnF2RGRfEF8uEEAd69w9THp22aJukFFpCyiQDCg9pJElm6+obGq QJgA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788208264; x=1788813064; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=MD955j7svwcecXFODkmvNY3vK8UjCpMQPU48awBPMVA=; b=VcsP+9FJGwSxrQhdCr0QWFGIcPrka/vir0Pqaqs2g9F9hUqXw8q8n4DqB/SMSkJoP0 IQ4LpZ69N2aqBC8JPCxtS/2gmvqzlmfvc7k0d3Ht0/lRqtpiztG+jSwWc1eJbPdQO3nt ruPsdx+fNuqbOHLFIzSxTsFSPzNbw8Zvbs+G07jj87bPBIHcx59fUN78JzUxEogAYQ37 XhOIwCDWAqUzVbf+NFbIZ1UTL9sNHqqmEvDOcUfB3rO15xLVMb1xTc5cIZBOUf4UUqbw AAKZlwXnaZEa42T8wYkhkxEgfiMGpOFVQCr3RTqvFgVkSGBYFmRmwFd06O2AMKuZ09DH nuZQ== X-Forwarded-Encrypted: i=1; AHgh+RrtUOMAFDOsYISBDfdCt5ec/n0oFyGkVZG+p7n9brjKSBX3SmngsrlAPrRx5rObxpBChyLUaB4KmOUorNg=@vger.kernel.org X-Gm-Message-State: AFuF++n3NVt89VILyXuvWQWdvspJL1DqyCBu2vTEYEGHMcn+tCKFunhW +Dns+e/mlzz2MRrh5q24sNL2hxLoYqfREqgmAZZ8OFDkNVkr0GQlknoEyl8tTEU6UPkXwy90ADU U+2qJnA== X-Received: from dlbvt12.prod.google.com ([2002:a05:7022:3f8c:b0:142:d756:9ce9]) (user=surenb job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:6a29:b0:3d1:be60:281b with SMTP id adf61e73a8af0-3d268ef9eb5mr42600453637.15.1788208263607; Mon, 31 Aug 2026 13:31:03 -0700 (PDT) Date: Mon, 31 Aug 2026 13:30:52 -0700 In-Reply-To: <20260831203056.838265-1-surenb@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260831203056.838265-1-surenb@google.com> X-Mailer: git-send-email 2.55.0.966.g6673acef38-goog Message-ID: <20260831203056.838265-2-surenb@google.com> Subject: [PATCH v7 1/5] mm: Make per-VMA locks available universally From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: dave.hansen@linux.intel.com, Liam.Howlett@oracle.com, ljs@kernel.org, david@redhat.com, willy@infradead.org, shakeel.butt@linux.dev, vbabka@kernel.org, jannh@google.com, aliceryhl@google.com, arve@android.com, cmllamas@google.com, christian@brauner.io, tkjos@android.com, dsahern@kernel.org, davem@davemloft.net, gregkh@linuxfoundation.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, netdev@vger.kernel.org, surenb@google.com Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Dave Hansen The per-VMA locks have been around for several years. They've had some bugs worked out of them and have seen quite wide use. However, they are still only available when architectures explicitly enable them. Remove the conditional compilation around the per-VMA locks, making them available on all architectures and configs. The approach up to now seemed to be to add ARCH_SUPPORTS_PER_VMA_LOCK when the architecture started using per-VMA locks in the fault handler. But, contrary to the naming, the Kconfig option does not really indicate whether the architecture supports per-VMA locks or not. It is more of a marker for whether the architecture is likely to benefit from per-VMA locks. To me, the most important thing side-effect of universal availability is letting per-VMA locks be used in SMP=3Dn configs. This lets us use per-VMA locking in all x86 code without fallbacks. Overall, this just generally makes the kernel simpler. Just look at the diffstat. It also opens the door to users that want to use the per-VMA locks in common code. Doing *that* brings additional simplifications. The downside of this is adding some fields to vm_area_struct and mm_struct. There are likely ways to optimize this, especially for things like SMP=3Dn configs. For now, do the simplest thing: use the same implementation everywhere. =3D=3D Considerations for NOMMU config =3D=3D NOMMU systems do not write-lock VMAs, therefore read-locking a VMA would always succeed unless VMA is detached. Therefore for NOMMU config we make vma_mark_attached() a NOOP, which keeps VMAs always in detached state. This causes VMA read-locking to always fail and the caller falls back to locking mmap_lock. The following functions will have a different implementation in NOMMU config: - vma_mark_attached(), vma_mark_detached() are made NOOPs, keeping VMAs always in a detached state and preventing assertions and refcount underflows; - vma_start_write(), vma_start_write_killable() are made NOOPs to avoid warnings in __vma_start_write() due to VMAs being detached. These functions are not used in NOMMU code but __vma_start_write() is an exported function, therefore might be used by drivers. - vma_assert_attached() is made NOOP because it's reachable from NOMMU code via split_vma()->vma_iter_store_new()->vma_iter_store_overwrite(); - vma_assert_write_locked() is asserting vma->vm_mm is write-locked, as was done before this change; - vma_assert_locked() is asserting vma->vm_mm is locked, as was done before this change; The following functions work for both MMU and NOMMU configs: - vma_lock_init() performs the same initialization as for MMU config; - mm_lock_seqcount_init(), mm_lock_seqcount_begin(), mm_lock_seqcount_end() are called from mmap_write_{lock|unlock} and update mm_lock_seq correctly. - mmap_lock_speculate_try_begin(), mmap_lock_speculate_retry() work as is because mm_lock_seq is updated correctly; - vma_start_read(), vma_start_read_locked() will always fail because VMAs are always detached; - vma_end_read() will never be called because vma_start_read() never succeeds; - vma_is_attached() always return false because VMAs are always detached; - vma_assert_detached() will never trigger because VMAs are never attached; - vma_start_read_locked() always return false because VMAs are always detached; - lock_vma_under_rcu() will be safe as the attempted read lock will bail; Changes in the following files are not affecting NOMMU config: task_mmu.c - not compiled when CONFIG_MMU=3Dn; pagewalk.c - not compiled when CONFIG_MMU=3Dn; userfaultfd.c - not compiled when CONFIG_MMU=3Dn (CONFIG_USERFAULTFD depends on CONFIG_MMU); The following changes in the BPF code are made to keep NOMMU config working like before: stack_map_lock_vma() - keeps mmap_lock in NOMMU config; bpf_iter_task_vma_new() - bails out in NOMMU config; Signed-off-by: Dave Hansen Cc: Suren Baghdasaryan Cc: Andrew Morton Cc: Liam R. Howlett Cc: Lorenzo Stoakes Cc: Vlastimil Babka Cc: Shakeel Butt Cc: linux-mm@kvack.org Cc: Greg Kroah-Hartman Cc: Arve Hj=C3=B8nnev=C3=A5g Cc: Todd Kjos Cc: Christian Brauner Cc: Carlos Llamas Cc: Alice Ryhl Cc: David S. Miller Cc: David Ahern Cc: netdev@vger.kernel.org Signed-off-by: Suren Baghdasaryan Reviewed-by: Lorenzo Stoakes (ARM) Acked-by: Vlastimil Babka (SUSE) Acked-by: David Hildenbrand (Arm) --- arch/arm/Kconfig | 1 - arch/arm64/Kconfig | 1 - arch/loongarch/Kconfig | 1 - arch/powerpc/platforms/powernv/Kconfig | 1 - arch/powerpc/platforms/pseries/Kconfig | 1 - arch/riscv/Kconfig | 1 - arch/s390/Kconfig | 1 - arch/x86/Kconfig | 2 - fs/proc/internal.h | 2 - fs/proc/task_mmu.c | 93 -------------------------- include/linux/mm.h | 12 ---- include/linux/mm_types.h | 8 +-- include/linux/mmap_lock.h | 75 +++++++-------------- kernel/bpf/stackmap.c | 17 ++--- kernel/bpf/task_iter.c | 2 +- kernel/fork.c | 2 - mm/Kconfig | 12 ---- mm/Kconfig.debug | 1 - mm/debug.c | 4 -- mm/init-mm.c | 2 - mm/memory.c | 2 - mm/mmap_lock.c | 26 +------ mm/pagewalk.c | 2 - mm/rmap.c | 2 - mm/userfaultfd.c | 55 --------------- rust/kernel/mm.rs | 32 +++------ tools/testing/vma/include/dup.h | 5 +- tools/testing/vma/vma_internal.h | 1 - 28 files changed, 48 insertions(+), 316 deletions(-) diff --git a/arch/arm/Kconfig b/arch/arm/Kconfig index f10379dfe7a3..3075e5c60a25 100644 --- a/arch/arm/Kconfig +++ b/arch/arm/Kconfig @@ -42,7 +42,6 @@ config ARM select ARCH_SUPPORTS_ATOMIC_RMW select ARCH_SUPPORTS_CFI select ARCH_SUPPORTS_HUGETLBFS if ARM_LPAE - select ARCH_SUPPORTS_PER_VMA_LOCK select ARCH_SUPPORTS_RT select ARCH_USE_BUILTIN_BSWAP select ARCH_USE_CMPXCHG_LOCKREF diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig index b5a51b0ef944..2bbeded33da0 100644 --- a/arch/arm64/Kconfig +++ b/arch/arm64/Kconfig @@ -81,7 +81,6 @@ config ARM64 select ARCH_HAS_PTE_PROTNONE select ARCH_SUPPORTS_NUMA_BALANCING select ARCH_SUPPORTS_PAGE_TABLE_CHECK - select ARCH_SUPPORTS_PER_VMA_LOCK select ARCH_SUPPORTS_HUGE_PFNMAP if TRANSPARENT_HUGEPAGE select ARCH_SUPPORTS_RT select ARCH_SUPPORTS_SCHED_SMT diff --git a/arch/loongarch/Kconfig b/arch/loongarch/Kconfig index a21f51e5815e..9c5def706222 100644 --- a/arch/loongarch/Kconfig +++ b/arch/loongarch/Kconfig @@ -69,7 +69,6 @@ config LOONGARCH select ARCH_SUPPORTS_MSEAL_SYSTEM_MAPPINGS select ARCH_HAS_PTE_PROTNONE if 64BIT select ARCH_SUPPORTS_NUMA_BALANCING if NUMA - select ARCH_SUPPORTS_PER_VMA_LOCK select ARCH_SUPPORTS_RT select ARCH_SUPPORTS_SCHED_SMT if SMP select ARCH_SUPPORTS_SCHED_MC if SMP diff --git a/arch/powerpc/platforms/powernv/Kconfig b/arch/powerpc/platform= s/powernv/Kconfig index b5ad7c173ef0..dd8f6060fb7a 100644 --- a/arch/powerpc/platforms/powernv/Kconfig +++ b/arch/powerpc/platforms/powernv/Kconfig @@ -17,7 +17,6 @@ config PPC_POWERNV select PPC_DOORBELL select MMU_NOTIFIER select FORCE_SMP - select ARCH_SUPPORTS_PER_VMA_LOCK select PPC_RADIX_BROADCAST_TLBIE if PPC_RADIX_MMU default y =20 diff --git a/arch/powerpc/platforms/pseries/Kconfig b/arch/powerpc/platform= s/pseries/Kconfig index 74910ce3a541..7d125e288f6e 100644 --- a/arch/powerpc/platforms/pseries/Kconfig +++ b/arch/powerpc/platforms/pseries/Kconfig @@ -23,7 +23,6 @@ config PPC_PSERIES select HOTPLUG_CPU select FORCE_SMP select SWIOTLB - select ARCH_SUPPORTS_PER_VMA_LOCK select PPC_RADIX_BROADCAST_TLBIE if PPC_RADIX_MMU default y =20 diff --git a/arch/riscv/Kconfig b/arch/riscv/Kconfig index f8e26c4bed2b..505eed4af932 100644 --- a/arch/riscv/Kconfig +++ b/arch/riscv/Kconfig @@ -72,7 +72,6 @@ config RISCV select ARCH_SUPPORTS_LTO_CLANG_THIN select ARCH_SUPPORTS_MSEAL_SYSTEM_MAPPINGS if 64BIT && MMU select ARCH_SUPPORTS_PAGE_TABLE_CHECK if MMU - select ARCH_SUPPORTS_PER_VMA_LOCK if MMU select ARCH_HAS_PTE_PROTNONE if MMU select ARCH_SUPPORTS_RT select ARCH_SUPPORTS_SHADOW_CALL_STACK if HAVE_SHADOW_CALL_STACK diff --git a/arch/s390/Kconfig b/arch/s390/Kconfig index 4b51bc6e8948..b88b85042136 100644 --- a/arch/s390/Kconfig +++ b/arch/s390/Kconfig @@ -156,7 +156,6 @@ config S390 select ARCH_HAS_PTE_PROTNONE select ARCH_SUPPORTS_NUMA_BALANCING select ARCH_SUPPORTS_PAGE_TABLE_CHECK - select ARCH_SUPPORTS_PER_VMA_LOCK select ARCH_USES_CFI_GENERIC_LLVM_PASS if CC_IS_CLANG select ARCH_USE_BUILTIN_BSWAP select ARCH_USE_CMPXCHG_LOCKREF diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig index 15fd9ec5ecac..ac92b3fd70c6 100644 --- a/arch/x86/Kconfig +++ b/arch/x86/Kconfig @@ -27,7 +27,6 @@ config X86_64 select ARCH_HAS_GIGANTIC_PAGE select ARCH_SUPPORTS_MSEAL_SYSTEM_MAPPINGS select ARCH_SUPPORTS_INT128 if CC_HAS_INT128 - select ARCH_SUPPORTS_PER_VMA_LOCK select ARCH_SUPPORTS_HUGE_PFNMAP if TRANSPARENT_HUGEPAGE select HAVE_ARCH_SOFT_DIRTY select MODULES_USE_ELF_RELA @@ -1849,7 +1848,6 @@ config X86_USER_SHADOW_STACK bool "X86 userspace shadow stack" depends on AS_WRUSS depends on X86_64 - depends on PER_VMA_LOCK select ARCH_USES_HIGH_VMA_FLAGS select ARCH_HAS_USER_SHADOW_STACK select X86_CET diff --git a/fs/proc/internal.h b/fs/proc/internal.h index 04bd6c9e65a7..623bb43ede55 100644 --- a/fs/proc/internal.h +++ b/fs/proc/internal.h @@ -385,10 +385,8 @@ struct mem_size_stats; =20 struct proc_maps_locking_ctx { struct mm_struct *mm; -#ifdef CONFIG_PER_VMA_LOCK bool mmap_locked; struct vm_area_struct *locked_vma; -#endif }; =20 struct proc_maps_private { diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c index 5c54aebe2118..e671b4fd8ded 100644 --- a/fs/proc/task_mmu.c +++ b/fs/proc/task_mmu.c @@ -130,8 +130,6 @@ static void release_task_mempolicy(struct proc_maps_pri= vate *priv) } #endif =20 -#ifdef CONFIG_PER_VMA_LOCK - static inline int lock_ctx_mm(struct proc_maps_locking_ctx *lock_ctx) { int ret =3D mmap_read_lock_killable(lock_ctx->mm); @@ -233,46 +231,6 @@ static inline void reacquire_rcu(struct proc_maps_priv= ate *priv) vma_iter_set(&priv->iter, priv->lock_ctx.locked_vma->vm_end); } =20 -#else /* CONFIG_PER_VMA_LOCK */ - -static inline int lock_ctx_mm(struct proc_maps_locking_ctx *lock_ctx) -{ - return mmap_read_lock_killable(lock_ctx->mm); -} - -static inline void unlock_ctx_mm(struct proc_maps_locking_ctx *lock_ctx) -{ - mmap_read_unlock(lock_ctx->mm); -} - -static inline bool lock_vma_range(struct seq_file *m, - struct proc_maps_locking_ctx *lock_ctx) -{ - return lock_ctx_mm(lock_ctx) =3D=3D 0; -} - -static inline void unlock_vma_range(struct proc_maps_locking_ctx *lock_ctx) -{ - unlock_ctx_mm(lock_ctx); -} - -static struct vm_area_struct *get_next_vma(struct proc_maps_private *priv, - loff_t last_pos) -{ - return vma_next(&priv->iter); -} - -static inline bool fallback_to_mmap_lock(struct proc_maps_private *priv, - loff_t pos) -{ - return false; -} - -static inline void drop_rcu(struct proc_maps_private *priv) {} -static inline void reacquire_rcu(struct proc_maps_private *priv) {} - -#endif /* CONFIG_PER_VMA_LOCK */ - static struct vm_area_struct *proc_get_vma(struct seq_file *m, loff_t *ppo= s) { struct proc_maps_private *priv =3D m->private; @@ -560,8 +518,6 @@ static int pid_maps_open(struct inode *inode, struct fi= le *file) PROCMAP_QUERY_VMA_FLAGS \ ) =20 -#ifdef CONFIG_PER_VMA_LOCK - static int query_vma_setup(struct proc_maps_locking_ctx *lock_ctx) { reset_lock_ctx(lock_ctx); @@ -612,26 +568,6 @@ static struct vm_area_struct *query_vma_find_by_addr(s= truct proc_maps_locking_ct return vma; } =20 -#else /* CONFIG_PER_VMA_LOCK */ - -static int query_vma_setup(struct proc_maps_locking_ctx *lock_ctx) -{ - return mmap_read_lock_killable(lock_ctx->mm); -} - -static void query_vma_teardown(struct proc_maps_locking_ctx *lock_ctx) -{ - mmap_read_unlock(lock_ctx->mm); -} - -static struct vm_area_struct *query_vma_find_by_addr(struct proc_maps_lock= ing_ctx *lock_ctx, - unsigned long addr) -{ - return find_vma(lock_ctx->mm, addr); -} - -#endif /* CONFIG_PER_VMA_LOCK */ - static struct vm_area_struct *query_matching_vma(struct proc_maps_locking_= ctx *lock_ctx, unsigned long addr, u32 flags) { @@ -1314,8 +1250,6 @@ static const struct mm_walk_ops smaps_shmem_walk_ops = =3D { .walk_lock =3D PGWALK_RDLOCK, }; =20 -#ifdef CONFIG_PER_VMA_LOCK - static const struct mm_walk_ops smaps_walk_vma_lock_ops =3D { .pmd_entry =3D smaps_pte_range, .hugetlb_entry =3D smaps_hugetlb_range, @@ -1345,22 +1279,6 @@ get_smaps_shmem_walk_ops(struct proc_maps_private *p= riv) return &smaps_shmem_walk_vma_lock_ops; } =20 -#else /* CONFIG_PER_VMA_LOCK */ - -static inline const struct mm_walk_ops * -get_smaps_walk_ops(struct proc_maps_private *priv) -{ - return &smaps_walk_ops; -} - -static inline const struct mm_walk_ops * -get_smaps_shmem_walk_ops(struct proc_maps_private *priv) -{ - return &smaps_shmem_walk_ops; -} - -#endif /* CONFIG_PER_VMA_LOCK */ - /* * Gather mem stats from @vma with the indicated beginning * address @start, and keep them in @mss. @@ -3497,7 +3415,6 @@ static const struct mm_walk_ops show_numa_ops =3D { .walk_lock =3D PGWALK_RDLOCK, }; =20 -#ifdef CONFIG_PER_VMA_LOCK static const struct mm_walk_ops show_numa_vma_lock_ops =3D { .hugetlb_entry =3D gather_hugetlb_stats, .pmd_entry =3D gather_pte_stats, @@ -3512,16 +3429,6 @@ get_show_numa_ops(struct proc_maps_private *priv) return &show_numa_vma_lock_ops; } =20 -#else /* CONFIG_PER_VMA_LOCK */ - -static inline const struct mm_walk_ops * -get_show_numa_ops(struct proc_maps_private *priv) -{ - return &show_numa_ops; -} - -#endif /* CONFIG_PER_VMA_LOCK */ - /* * Display pages allocated per node and memory policy via /proc. */ diff --git a/include/linux/mm.h b/include/linux/mm.h index dd09c438fa23..7aaa25abc241 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -928,7 +928,6 @@ static inline void vma_numab_state_free(struct vm_area_= struct *vma) {} * These must be here rather than mmap_lock.h as dependent on vm_fault typ= e, * declared in this header. */ -#ifdef CONFIG_PER_VMA_LOCK static inline void release_fault_lock(struct vm_fault *vmf) { if (vmf->flags & FAULT_FLAG_VMA_LOCK) @@ -944,17 +943,6 @@ static inline void assert_fault_locked(const struct vm= _fault *vmf) else mmap_assert_locked(vmf->vma->vm_mm); } -#else -static inline void release_fault_lock(struct vm_fault *vmf) -{ - mmap_read_unlock(vmf->vma->vm_mm); -} - -static inline void assert_fault_locked(const struct vm_fault *vmf) -{ - mmap_assert_locked(vmf->vma->vm_mm); -} -#endif /* CONFIG_PER_VMA_LOCK */ =20 static inline bool mm_flags_test(int flag, const struct mm_struct *mm) { diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 6d815f6440c9..5413bd10fff2 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -950,7 +950,6 @@ struct vm_area_struct { vma_flags_t flags; }; =20 -#ifdef CONFIG_PER_VMA_LOCK /* * Can only be written (using WRITE_ONCE()) while holding both: * - mmap_lock (in write mode) @@ -966,7 +965,7 @@ struct vm_area_struct { * slowpath. */ unsigned int vm_lock_seq; -#endif + /* * Low 32-bits of anonymous page offset. * See vma_start_anon_pgoff() comment for details. @@ -1003,7 +1002,6 @@ struct vm_area_struct { #ifdef CONFIG_NUMA_BALANCING struct vma_numab_state *numab_state; /* NUMA Balancing state */ #endif -#ifdef CONFIG_PER_VMA_LOCK /* * Used to keep track of firstly, whether the VMA is attached, secondly, * if attached, how many read locks are taken, and thirdly, if the @@ -1046,7 +1044,6 @@ struct vm_area_struct { #ifdef CONFIG_DEBUG_LOCK_ALLOC struct lockdep_map vmlock_dep_map; #endif -#endif #ifdef CONFIG_64BIT /* * High 32-bits of anonymous page offset. @@ -1254,7 +1251,6 @@ struct mm_struct { * init_mm.mmlist, and are protected * by mmlist_lock */ -#ifdef CONFIG_PER_VMA_LOCK struct rcuwait vma_writer_wait; /* * This field has lock-like semantics, meaning it is sometimes @@ -1274,7 +1270,7 @@ struct mm_struct { * mmap_lock. */ seqcount_t mm_lock_seq; -#endif + struct futex_mm_data futex; =20 unsigned long hiwater_rss; /* High-watermark of RSS usage */ diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index bec0eab6ef03..27c00bac29f9 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -76,8 +76,6 @@ static inline void mmap_assert_write_locked(const struct = mm_struct *mm) rwsem_assert_held_write(&mm->mmap_lock); } =20 -#ifdef CONFIG_PER_VMA_LOCK - #ifdef CONFIG_LOCKDEP #define __vma_lockdep_map(vma) (&vma->vmlock_dep_map) #else @@ -297,6 +295,9 @@ int __vma_start_write(struct vm_area_struct *vma, int s= tate); */ static inline void vma_start_write(struct vm_area_struct *vma) { + if (!IS_ENABLED(CONFIG_MMU)) + return; + if (__is_vma_write_locked(vma)) return; =20 @@ -319,6 +320,9 @@ static inline void vma_start_write(struct vm_area_struc= t *vma) static inline __must_check int vma_start_write_killable(struct vm_area_struct *vma) { + if (!IS_ENABLED(CONFIG_MMU)) + return 0; + if (__is_vma_write_locked(vma)) return 0; =20 @@ -331,6 +335,11 @@ int vma_start_write_killable(struct vm_area_struct *vm= a) */ static inline void vma_assert_write_locked(struct vm_area_struct *vma) { + if (!IS_ENABLED(CONFIG_MMU)) { + mmap_assert_write_locked(vma->vm_mm); + return; + } + VM_WARN_ON_ONCE_VMA(!__is_vma_write_locked(vma), vma); } =20 @@ -343,6 +352,11 @@ static inline void vma_assert_locked(struct vm_area_st= ruct *vma) { unsigned int refcnt; =20 + if (!IS_ENABLED(CONFIG_MMU)) { + mmap_assert_locked(vma->vm_mm); + return; + } + if (IS_ENABLED(CONFIG_LOCKDEP)) { if (!lock_is_held(__vma_lockdep_map(vma))) vma_assert_write_locked(vma); @@ -432,6 +446,9 @@ static inline bool vma_is_attached(struct vm_area_struc= t *vma) */ static inline void vma_assert_attached(struct vm_area_struct *vma) { + if (!IS_ENABLED(CONFIG_MMU)) + return; + WARN_ON_ONCE(!vma_is_attached(vma)); } =20 @@ -442,6 +459,9 @@ static inline void vma_assert_detached(struct vm_area_s= truct *vma) =20 static inline void vma_mark_attached(struct vm_area_struct *vma) { + if (!IS_ENABLED(CONFIG_MMU)) + return; + vma_assert_write_locked(vma); vma_assert_detached(vma); refcount_set_release(&vma->vm_refcnt, 1); @@ -451,6 +471,9 @@ void __vma_exclude_readers_for_detach(struct vm_area_st= ruct *vma); =20 static inline void vma_mark_detached(struct vm_area_struct *vma) { + if (!IS_ENABLED(CONFIG_MMU)) + return; + vma_assert_write_locked(vma); vma_assert_attached(vma); =20 @@ -484,54 +507,6 @@ struct vm_area_struct *lock_next_vma(struct mm_struct = *mm, struct vma_iterator *iter, unsigned long address); =20 -#else /* CONFIG_PER_VMA_LOCK */ - -static inline void mm_lock_seqcount_init(struct mm_struct *mm) {} -static inline void mm_lock_seqcount_begin(struct mm_struct *mm) {} -static inline void mm_lock_seqcount_end(struct mm_struct *mm) {} - -static inline bool mmap_lock_speculate_try_begin(struct mm_struct *mm, uns= igned int *seq) -{ - return false; -} - -static inline bool mmap_lock_speculate_retry(struct mm_struct *mm, unsigne= d int seq) -{ - return true; -} -static inline void vma_lock_init(struct vm_area_struct *vma, bool reset_re= fcnt) {} -static inline void vma_end_read(struct vm_area_struct *vma) {} -static inline void vma_start_write(struct vm_area_struct *vma) {} -static inline __must_check -int vma_start_write_killable(struct vm_area_struct *vma) { return 0; } -static inline void vma_assert_write_locked(struct vm_area_struct *vma) - { mmap_assert_write_locked(vma->vm_mm); } -static inline bool vma_is_attached(struct vm_area_struct *vma) - { return true; } -static inline void vma_assert_attached(struct vm_area_struct *vma) {} -static inline void vma_assert_detached(struct vm_area_struct *vma) {} -static inline void vma_mark_attached(struct vm_area_struct *vma) {} -static inline void vma_mark_detached(struct vm_area_struct *vma) {} - -static inline struct vm_area_struct *lock_vma_under_rcu(struct mm_struct *= mm, - unsigned long address) -{ - return NULL; -} - -static inline void vma_assert_locked(struct vm_area_struct *vma) -{ - mmap_assert_locked(vma->vm_mm); -} - -static inline void vma_assert_stabilised(struct vm_area_struct *vma) -{ - /* If no VMA locks, then either mmap lock suffices to stabilise. */ - mmap_assert_locked(vma->vm_mm); -} - -#endif /* CONFIG_PER_VMA_LOCK */ - static inline void vma_assert_can_modify(struct vm_area_struct *vma) { if (vma_is_attached(vma)) diff --git a/kernel/bpf/stackmap.c b/kernel/bpf/stackmap.c index a839041e0d00..8fb70c2a9c8e 100644 --- a/kernel/bpf/stackmap.c +++ b/kernel/bpf/stackmap.c @@ -272,13 +272,10 @@ struct stack_map_vma_lock { /* * Acquire a stable read-side reference on the VMA covering @ip. * - * With CONFIG_PER_VMA_LOCK=3Dy this returns a VMA with its per-VMA read - * lock held and mmap_lock dropped, so the caller may sleep. - * - * With CONFIG_PER_VMA_LOCK=3Dn it returns a VMA with mmap_lock still - * held; the caller must snapshot any fields it needs and pin vm_file - * with get_file() before stack_map_unlock_vma() drops mmap_lock, as - * the VMA may be split, merged, or freed after that. + * On NOMMU configurations, returns with the mmap_lock held. If the MMU + * is enabled, the per-VMA lock will be held instead. The lock + * should be released with stack_map_unlock_vma() which will release the + * appropriate lock. Once the lock is released, the VMA may be freed. * * Returns NULL on failure, in which case no lock is held. */ @@ -288,7 +285,6 @@ stack_map_lock_vma(struct stack_map_vma_lock *lock, uns= igned long ip) struct mm_struct *mm =3D lock->mm; struct vm_area_struct *vma; =20 - /* noop under !CONFIG_PER_VMA_LOCK */ vma =3D lock_vma_under_rcu(mm, ip); if (vma) { lock->vma =3D vma; @@ -308,21 +304,20 @@ stack_map_lock_vma(struct stack_map_vma_lock *lock, u= nsigned long ip) return NULL; } =20 -#ifdef CONFIG_PER_VMA_LOCK +#ifdef CONFIG_MMU if (!vma_start_read_locked(vma)) { mmap_read_unlock(mm); return NULL; } mmap_read_unlock(mm); #endif - lock->vma =3D vma; return vma; } =20 static void stack_map_unlock_vma(struct stack_map_vma_lock *lock) { -#ifdef CONFIG_PER_VMA_LOCK +#ifdef CONFIG_MMU vma_end_read(lock->vma); #else mmap_read_unlock(lock->mm); diff --git a/kernel/bpf/task_iter.c b/kernel/bpf/task_iter.c index 13e1aabe6f88..c65ba1dcd866 100644 --- a/kernel/bpf/task_iter.c +++ b/kernel/bpf/task_iter.c @@ -869,7 +869,7 @@ __bpf_kfunc int bpf_iter_task_vma_new(struct bpf_iter_t= ask_vma *it, BUILD_BUG_ON(sizeof(struct bpf_iter_task_vma_kern) !=3D sizeof(struct bpf= _iter_task_vma)); BUILD_BUG_ON(__alignof__(struct bpf_iter_task_vma_kern) !=3D __alignof__(= struct bpf_iter_task_vma)); =20 - if (!IS_ENABLED(CONFIG_PER_VMA_LOCK)) { + if (!IS_ENABLED(CONFIG_MMU)) { kit->data =3D NULL; return -EOPNOTSUPP; } diff --git a/kernel/fork.c b/kernel/fork.c index 416758c8a3d4..22283bf849e1 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -1083,9 +1083,7 @@ static void mmap_init_lock(struct mm_struct *mm) { init_rwsem(&mm->mmap_lock); mm_lock_seqcount_init(mm); -#ifdef CONFIG_PER_VMA_LOCK rcuwait_init(&mm->vma_writer_wait); -#endif } =20 static struct mm_struct *mm_init(struct mm_struct *mm, struct task_struct = *p) diff --git a/mm/Kconfig b/mm/Kconfig index 604c58199acb..ece5d37b4eb7 100644 --- a/mm/Kconfig +++ b/mm/Kconfig @@ -1430,18 +1430,6 @@ config LRU_GEN_WALKS_MMU depends on LRU_GEN && ARCH_HAS_HW_PTE_YOUNG # } =20 -config ARCH_SUPPORTS_PER_VMA_LOCK - def_bool n - -config PER_VMA_LOCK - def_bool y - depends on ARCH_SUPPORTS_PER_VMA_LOCK && MMU && SMP - help - Allow per-vma locking during page fault handling. - - This feature allows locking each virtual memory area separately when - handling page faults instead of taking mmap_lock. - config LOCK_MM_AND_FIND_VMA bool depends on !STACK_GROWSUP diff --git a/mm/Kconfig.debug b/mm/Kconfig.debug index 15dca19dd07d..9eaa25d1cf23 100644 --- a/mm/Kconfig.debug +++ b/mm/Kconfig.debug @@ -310,7 +310,6 @@ config DEBUG_KMEMLEAK_VERBOSE =20 config PER_VMA_LOCK_STATS bool "Statistics for per-vma locks" - depends on PER_VMA_LOCK help Say Y here to enable success, retry and failure counters of page faults handled under protection of per-vma locks. When enabled, the diff --git a/mm/debug.c b/mm/debug.c index 9a0297b3988d..655e6bcc0e8d 100644 --- a/mm/debug.c +++ b/mm/debug.c @@ -157,17 +157,13 @@ void dump_vma(const struct vm_area_struct *vma) pr_emerg("vma %px start %px end %px mm %px\n" "prot %lx anon_vma %px vm_ops %px\n" "pgoff %lx file %px private_data %px\n" -#ifdef CONFIG_PER_VMA_LOCK "refcnt %x\n" -#endif "flags: %#lx(%pGv)\n", vma, (void *)vma->vm_start, (void *)vma->vm_end, vma->vm_mm, (unsigned long)pgprot_val(vma->vm_page_prot), vma->anon_vma, vma->vm_ops, vma_start_pgoff(vma), vma->vm_file, vma->vm_private_data, -#ifdef CONFIG_PER_VMA_LOCK refcount_read(&vma->vm_refcnt), -#endif vma->vm_flags, &vma->vm_flags); } EXPORT_SYMBOL(dump_vma); diff --git a/mm/init-mm.c b/mm/init-mm.c index 3e792aad7626..a1bb2c2d0284 100644 --- a/mm/init-mm.c +++ b/mm/init-mm.c @@ -39,10 +39,8 @@ struct mm_struct init_mm =3D { .page_table_lock =3D __SPIN_LOCK_UNLOCKED(init_mm.page_table_lock), .arg_lock =3D __SPIN_LOCK_UNLOCKED(init_mm.arg_lock), .mmlist =3D LIST_HEAD_INIT(init_mm.mmlist), -#ifdef CONFIG_PER_VMA_LOCK .vma_writer_wait =3D __RCUWAIT_INITIALIZER(init_mm.vma_writer_wait), .mm_lock_seq =3D SEQCNT_ZERO(init_mm.mm_lock_seq), -#endif #ifdef CONFIG_SCHED_MM_CID .mm_cid.lock =3D __RAW_SPIN_LOCK_UNLOCKED(init_mm.mm_cid.lock), #endif diff --git a/mm/memory.c b/mm/memory.c index 8b0c2c735d3d..7bd660d48ff7 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -6817,7 +6817,6 @@ static vm_fault_t sanitize_fault_flags(struct vm_area= _struct *vma, !vma_is_cow_mapping(vma))) return VM_FAULT_SIGSEGV; } -#ifdef CONFIG_PER_VMA_LOCK /* * Per-VMA locks can't be used with FAULT_FLAG_RETRY_NOWAIT because of * the assumption that lock is dropped on VM_FAULT_RETRY. @@ -6826,7 +6825,6 @@ static vm_fault_t sanitize_fault_flags(struct vm_area= _struct *vma, (FAULT_FLAG_VMA_LOCK | FAULT_FLAG_RETRY_NOWAIT)) =3D=3D (FAULT_FLAG_VMA_LOCK | FAULT_FLAG_RETRY_NOWAIT))) return VM_FAULT_SIGSEGV; -#endif =20 return 0; } diff --git a/mm/mmap_lock.c b/mm/mmap_lock.c index 898c2ef1e958..272f9ac762b9 100644 --- a/mm/mmap_lock.c +++ b/mm/mmap_lock.c @@ -43,9 +43,6 @@ void __mmap_lock_do_trace_released(struct mm_struct *mm, = bool write) EXPORT_SYMBOL(__mmap_lock_do_trace_released); #endif /* CONFIG_TRACING */ =20 -#ifdef CONFIG_MMU -#ifdef CONFIG_PER_VMA_LOCK - /* State shared across __vma_[start, end]_exclude_readers. */ struct vma_exclude_readers_state { /* Input parameters. */ @@ -299,6 +296,8 @@ struct vm_area_struct *lock_vma_under_rcu(struct mm_str= uct *mm, MA_STATE(mas, &mm->mm_mt, address, address); struct vm_area_struct *vma; =20 + if (!IS_ENABLED(CONFIG_MMU)) + return NULL; retry: rcu_read_lock(); vma =3D mas_walk(&mas); @@ -431,7 +430,6 @@ struct vm_area_struct *lock_next_vma(struct mm_struct *= mm, =20 return vma; } -#endif /* CONFIG_PER_VMA_LOCK */ =20 #ifdef CONFIG_LOCK_MM_AND_FIND_VMA #include @@ -548,23 +546,3 @@ struct vm_area_struct *lock_mm_and_find_vma(struct mm_= struct *mm, return NULL; } #endif /* CONFIG_LOCK_MM_AND_FIND_VMA */ - -#else /* CONFIG_MMU */ - -/* - * At least xtensa ends up having protection faults even with no - * MMU.. No stack expansion, at least. - */ -struct vm_area_struct *lock_mm_and_find_vma(struct mm_struct *mm, - unsigned long addr, struct pt_regs *regs) -{ - struct vm_area_struct *vma; - - mmap_read_lock(mm); - vma =3D vma_lookup(mm, addr); - if (!vma) - mmap_read_unlock(mm); - return vma; -} - -#endif /* CONFIG_MMU */ diff --git a/mm/pagewalk.c b/mm/pagewalk.c index cc07fcf50e87..7411702a37f5 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -444,7 +444,6 @@ static inline void process_mm_walk_lock(struct mm_struc= t *mm, static inline void process_vma_walk_lock(struct vm_area_struct *vma, enum page_walk_lock walk_lock) { -#ifdef CONFIG_PER_VMA_LOCK switch (walk_lock) { case PGWALK_WRLOCK: vma_start_write(vma); @@ -459,7 +458,6 @@ static inline void process_vma_walk_lock(struct vm_area= _struct *vma, /* PGWALK_RDLOCK is handled by process_mm_walk_lock */ break; } -#endif } =20 /* diff --git a/mm/rmap.c b/mm/rmap.c index d1819fd69938..fed0362e0bd0 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -260,11 +260,9 @@ static void check_anon_vma_clone(struct vm_area_struct= *dst, /* For the anon_vma to be compatible, it can only be singular. */ VM_WARN_ON_ONCE(operation =3D=3D VMA_OP_MERGE_UNFAULTED && !list_is_singular(&src->anon_vma_chain)); -#ifdef CONFIG_PER_VMA_LOCK /* Only merging an unfaulted VMA leaves the destination attached. */ VM_WARN_ON_ONCE(operation !=3D VMA_OP_MERGE_UNFAULTED && vma_is_attached(dst)); -#endif } =20 static void maybe_reuse_anon_vma(struct vm_area_struct *dst, diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index 74f04c323c50..d157e15fc77c 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -122,7 +122,6 @@ struct vm_area_struct *find_vma_and_prepare_anon(struct= mm_struct *mm, return vma; } =20 -#ifdef CONFIG_PER_VMA_LOCK /* * uffd_lock_vma() - Lookup and lock vma corresponding to @address. * @mm: mm to search vma in. @@ -182,34 +181,6 @@ static void uffd_mfill_unlock(struct vm_area_struct *v= ma) vma_end_read(vma); } =20 -#else - -static struct vm_area_struct *uffd_mfill_lock(struct mm_struct *dst_mm, - unsigned long dst_start, - unsigned long len) -{ - struct vm_area_struct *dst_vma; - - mmap_read_lock(dst_mm); - dst_vma =3D find_vma_and_prepare_anon(dst_mm, dst_start); - if (IS_ERR(dst_vma)) - goto out_unlock; - - if (validate_dst_vma(dst_vma, dst_start + len)) - return dst_vma; - - dst_vma =3D ERR_PTR(-ENOENT); -out_unlock: - mmap_read_unlock(dst_mm); - return dst_vma; -} - -static void uffd_mfill_unlock(struct vm_area_struct *vma) -{ - mmap_read_unlock(vma->vm_mm); -} -#endif - static void mfill_put_vma(struct mfill_state *state) { if (!state->vma) @@ -1850,7 +1821,6 @@ int find_vmas_mm_locked(struct mm_struct *mm, return 0; } =20 -#ifdef CONFIG_PER_VMA_LOCK static int uffd_move_lock(struct mm_struct *mm, unsigned long dst_start, unsigned long src_start, @@ -1925,31 +1895,6 @@ static void uffd_move_unlock(struct vm_area_struct *= dst_vma, vma_end_read(dst_vma); } =20 -#else - -static int uffd_move_lock(struct mm_struct *mm, - unsigned long dst_start, - unsigned long src_start, - struct vm_area_struct **dst_vmap, - struct vm_area_struct **src_vmap) -{ - int err; - - mmap_read_lock(mm); - err =3D find_vmas_mm_locked(mm, dst_start, src_start, dst_vmap, src_vmap); - if (err) - mmap_read_unlock(mm); - return err; -} - -static void uffd_move_unlock(struct vm_area_struct *dst_vma, - struct vm_area_struct *src_vma) -{ - mmap_assert_locked(src_vma->vm_mm); - mmap_read_unlock(dst_vma->vm_mm); -} -#endif - /** * move_pages - move arbitrary anonymous pages of an existing vma * @ctx: pointer to the userfaultfd context diff --git a/rust/kernel/mm.rs b/rust/kernel/mm.rs index 4764d7b68f2a..f4fa54616085 100644 --- a/rust/kernel/mm.rs +++ b/rust/kernel/mm.rs @@ -170,30 +170,20 @@ pub unsafe fn from_raw<'a>(ptr: *const bindings::mm_s= truct) -> &'a MmWithUser { /// /// This is an optimistic trylock operation, so it may fail if there i= s contention. In that /// case, you should fall back to taking the mmap read lock. - /// - /// When per-vma locks are disabled, this always returns `None`. #[inline] pub fn lock_vma_under_rcu(&self, vma_addr: usize) -> Option> { - #[cfg(CONFIG_PER_VMA_LOCK)] - { - // SAFETY: Calling `bindings::lock_vma_under_rcu` is always ok= ay given an mm where - // `mm_users` is non-zero. - let vma =3D unsafe { bindings::lock_vma_under_rcu(self.as_raw(= ), vma_addr) }; - if !vma.is_null() { - return Some(VmaReadGuard { - // SAFETY: If `lock_vma_under_rcu` returns a non-null = ptr, then it points at a - // valid vma. The vma is stable for as long as the vma= read lock is held. - vma: unsafe { VmaRef::from_raw(vma) }, - _nts: NotThreadSafe, - }); - } + // SAFETY: Calling `bindings::lock_vma_under_rcu` is always okay g= iven an mm where + // `mm_users` is non-zero. + let vma =3D unsafe { bindings::lock_vma_under_rcu(self.as_raw(), v= ma_addr) }; + if vma.is_null() { + return None; } - - // Silence warnings about unused variables. - #[cfg(not(CONFIG_PER_VMA_LOCK))] - let _ =3D vma_addr; - - None + Some(VmaReadGuard { + // SAFETY: If `lock_vma_under_rcu` returns a non-null ptr, the= n it points at a + // valid vma. The vma is stable for as long as the vma read lo= ck is held. + vma: unsafe { VmaRef::from_raw(vma) }, + _nts: NotThreadSafe, + }) } =20 /// Lock the mmap read lock. diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index 4c58487b764e..57046d8ac81d 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -560,7 +560,6 @@ struct vm_area_struct { vma_flags_t flags; }; =20 -#ifdef CONFIG_PER_VMA_LOCK /* * Can only be written (using WRITE_ONCE()) while holding both: * - mmap_lock (in write mode) @@ -576,7 +575,7 @@ struct vm_area_struct { * slowpath. */ unsigned int vm_lock_seq; -#endif + unsigned int __vm_anon_pgoff_lo; =20 /* @@ -610,10 +609,8 @@ struct vm_area_struct { #ifdef CONFIG_NUMA_BALANCING struct vma_numab_state *numab_state; /* NUMA Balancing state */ #endif -#ifdef CONFIG_PER_VMA_LOCK /* Unstable RCU readers are allowed to read this. */ refcount_t vm_refcnt; -#endif #ifdef CONFIG_64BIT unsigned int __vm_anon_pgoff_hi; #endif diff --git a/tools/testing/vma/vma_internal.h b/tools/testing/vma/vma_inter= nal.h index 8a48b231aa7a..54d5c3360aa2 100644 --- a/tools/testing/vma/vma_internal.h +++ b/tools/testing/vma/vma_internal.h @@ -15,7 +15,6 @@ #include =20 #define CONFIG_MMU 1 -#define CONFIG_PER_VMA_LOCK 1 =20 #ifdef __CONCAT #undef __CONCAT --=20 2.55.0.966.g6673acef38-goog From nobody Sat Sep 26 13:47:00 2026 Received: from mail-pj1-f69.google.com (mail-pj1-f69.google.com [209.85.216.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 53B1A38E5C8 for ; Mon, 31 Aug 2026 20:31:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208268; cv=none; b=pw3uDwmMsuPkQnzw+9/OAchRT8RcbrdmITansiNpPEWVInIpa5xn/u3Vw382NZQdlt5IEhXuUgmgcyXGDHsLzGRFQIQg05tRLYnCToWV2jqCfPgKAmzh6r/3d6K41yFxS302wDoqIyWWcmA5EQ2tWInFualTFTvMnnS8U9ib4hw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208268; c=relaxed/simple; bh=nMRPhuGFkFLmIP849hxdJsTzN087jM+7bAJ4YB45i7Y=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=ua/EDxyZVH5GbZYiy980WEJVksxsG32DEeGdyS3YhppgrTlUFv4/2PRA5+9DzKuDjjFcAeGGDZH/RbY16tmBtPPMnc4175tCltkFtZV+1+mwdFwFXJXxhADBU2+5X0ksFQo+2y/EXvL6qQSgn8UsNAudbd+d1n4Y1lRrkEI3jfI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=M7yg/B2y; arc=none smtp.client-ip=209.85.216.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="M7yg/B2y" Received: by mail-pj1-f69.google.com with SMTP id 98e67ed59e1d1-398ba5404d2so3068671a91.1 for ; Mon, 31 Aug 2026 13:31:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788208266; x=1788813066; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=PJ7fafUlBPjfr1KwszEslsPjbtrF6UKljnaxbzPgF8k=; b=M7yg/B2yTxv7rPh/oNGqzlsa8/ULj7Fmkn4epB6Z8NqIZxHLoKPcduTRkM5l3ePql2 PEewJYzPAMF7rF6S2fZ9LgDpbHImsTDTUGI6KTJabXHsm1sZRYF5ztxXaH7tStwEmv+k aDzLAn56A8mnIGmLXlqq8Ubtt2mexqLAdq8AI6WvHEWUcjnUxgRPmSg4Nbqlc+Dh5T9y EVWYGFglE7n9MIBgclERMsAYufB127vdThAz/vl4WPwJMV1xmsX6IIVuhpNA/IraYHO/ pFyAw1mZWX3cQZXbwC+58pTpiySA4D81GFTLc6tSgeXOr0DWUCE/j7OdfVobZbvKS4CZ SCqg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788208266; x=1788813066; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=PJ7fafUlBPjfr1KwszEslsPjbtrF6UKljnaxbzPgF8k=; b=jIuh9CwPoeTbs1cc5LliIoK/qJNp/suwQ8OEPtb9BC9YB4OQMDIMaKt0an7Seh+5Pa pPKyxwC0FOeKK/+lEtdLIZJACaZc9kPLZmdlWElUK7DX0JopIfJ7ZlgwsRTgVfnNswen Fd3MOflZxUA4eFuybQlQ+BrfrDV95/YFHm9Zc768nPFA1e0PVmKMTIiE8YaRLfixkAAz OwdlFtC/qTl2CJAF/BrWri/WPV7riEv0T9fftcKJmlk71hOrQIMJyqTAH6yCGV5aWIhy ExxQB0WF0NVcO8H2gW4/6RWD1wjNkGW6fFzUyyJFpwd1avUW4S+iZdC0MJrf77mJvNxz dOig== X-Forwarded-Encrypted: i=1; AKwUvBxIyx2jBFkBqsXxScijznbGZyNZQ9RzCgRKlUfQTFZYftFpUlrwJG20Hd12L9oTZkbUg8YYuLgXnWHi5qc=@vger.kernel.org X-Gm-Message-State: AFuF++lXUQH+AVSswuasoMpJsjxP95fenaLwR0Rm8ZLX1h7ftMKmJWE7 CxPkgi6DM11V1ZuGOBWVRFObMoMUwSSBH0x7K5YxPSrPXF++yAHwsmFQ1tnSq3tDlzutE14XREM VNxdglQ== X-Received: from dlboy1.prod.google.com ([2002:a05:7022:1281:b0:141:5392:d9c1]) (user=surenb job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90a:d446:b0:398:9be6:f998 with SMTP id 98e67ed59e1d1-39907e15804mr4310486a91.23.1788208266284; Mon, 31 Aug 2026 13:31:06 -0700 (PDT) Date: Mon, 31 Aug 2026 13:30:53 -0700 In-Reply-To: <20260831203056.838265-1-surenb@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260831203056.838265-1-surenb@google.com> X-Mailer: git-send-email 2.55.0.966.g6673acef38-goog Message-ID: <20260831203056.838265-3-surenb@google.com> Subject: [PATCH v7 2/5] binder: Make shrinker rely solely on per-VMA lock From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: dave.hansen@linux.intel.com, Liam.Howlett@oracle.com, ljs@kernel.org, david@redhat.com, willy@infradead.org, shakeel.butt@linux.dev, vbabka@kernel.org, jannh@google.com, aliceryhl@google.com, arve@android.com, cmllamas@google.com, christian@brauner.io, tkjos@android.com, dsahern@kernel.org, davem@davemloft.net, gregkh@linuxfoundation.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, netdev@vger.kernel.org, surenb@google.com Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Dave Hansen tl;dr: lock_vma_under_rcu() is already a trylock. No need to do both it and mmap_read_trylock(). Long Version: =3D=3D Background =3D=3D Historically, binder used an mmap_read_trylock() in its shrinker code. This ensures that reclaim is not blocked on an mmap_lock. Commit 95bc2d4a9020 ("binder: use per-vma lock in page reclaiming") added support for the per-VMA lock, but left mmap_read_trylock() as a fallback. This was presumably because the per-VMA locking can fail for several reasons and most (all?) lock_vma_under_rcu() callers have a fallback to mmap_read_trylock(). =3D=3D Problem =3D=3D The fallback is not worth the complexity here. lock_vma_under_rcu() is essentially already a non-blocking trylock. The main reason it fails is also the reason mmap_read_trylock() fails: something is holding mmap_write_lock(). The only remedy for a collision with mmap_write_lock() is to wait, which this code can not do. So the "fallback" after lock_vma_under_rcu() failure is not really a fallback: it is really likely to just be retrying in vain. That retry in an of itself isn't horrible. But it adds complexity. =3D=3D Solution =3D=3D Now that per-VMA locks are universally available, lock_vma_under_rcu() will not persistently fail. Rely on it alone and simplify the code. The removal of the fallback does not affect NOMMU case because binder driver depends on CONFIG_MMU. While at it we also make the handling of the cases where the original binder VMA is gone consistent. There are two cases to consider when Binder VMA is gone: 1. there is no VMA at that location anymore. 2. there is now another unrelated VMA at that location. Before this change we handle case 1 by having the shrinker proceed to free the page, and just skip the zap_vma_range() call. And we handle case 2 by having the shrinker return LRU_SKIP. While either behavior is acceptable, we need to handle them in a consistent way. Handle both cases by freeing the page without touching the VMA (skipping the zap_vma_range()). Full disclosure: I originally tried to do this with lock_vma_under_rcu_wait(), but it did not fit well with the mmap_lock trylock semantics. Claude caught this in a review and suggested the approach in this path. It seemed sane to me. So, Suggesed-by: Claude, I guess. Signed-off-by: Dave Hansen Cc: Andrew Morton Cc: Liam R. Howlett Cc: Vlastimil Babka Cc: Shakeel Butt Cc: linux-mm@kvack.org Cc: Greg Kroah-Hartman Cc: Arve Hj=C3=B8nnev=C3=A5g Cc: Todd Kjos Cc: Christian Brauner Cc: Carlos Llamas Cc: Alice Ryhl Cc: David S. Miller Cc: David Ahern Cc: netdev@vger.kernel.org Signed-off-by: Suren Baghdasaryan Reviewed-by: Alice Ryhl Acked-by: Lorenzo Stoakes (ARM) Acked-by: Carlos Llamas --- drivers/android/binder_alloc.c | 46 ++++++++++++++++------------------ 1 file changed, 21 insertions(+), 25 deletions(-) diff --git a/drivers/android/binder_alloc.c b/drivers/android/binder_alloc.c index e4488ad86a65..fcb744088e77 100644 --- a/drivers/android/binder_alloc.c +++ b/drivers/android/binder_alloc.c @@ -1142,7 +1142,6 @@ enum lru_status binder_alloc_free_page(struct list_he= ad *item, struct vm_area_struct *vma; struct page *page_to_free; unsigned long page_addr; - int mm_locked =3D 0; size_t index; =20 if (!mmget_not_zero(mm)) @@ -1151,27 +1150,25 @@ enum lru_status binder_alloc_free_page(struct list_= head *item, index =3D mdata->page_index; page_addr =3D alloc->vm_start + index * PAGE_SIZE; =20 - /* attempt per-vma lock first */ + /* + * Attempt per-vma lock. This is essentially a + * "trylock". It can fail even if the VMA exists + * for 'page_addr'. + */ vma =3D lock_vma_under_rcu(mm, page_addr); if (!vma) { - /* fall back to mmap_lock */ - if (!mmap_read_trylock(mm)) - goto err_mmap_read_lock_failed; - mm_locked =3D 1; - vma =3D vma_lookup(mm, page_addr); + /* + * If the vma exists, we can't continue because we cannot + * remove the page from the vma. However, if the vma was + * unmapped, it's okay to continue. + */ + if (binder_alloc_is_mapped(alloc)) + goto err_vma_lock_failed; } =20 if (!mutex_trylock(&alloc->mutex)) goto err_get_alloc_mutex_failed; =20 - /* - * Since a binder_alloc can only be mapped once, we ensure - * the vma corresponds to this mapping by checking whether - * the binder_alloc is still mapped. - */ - if (vma && !binder_alloc_is_mapped(alloc)) - goto err_invalid_vma; - trace_binder_unmap_kernel_start(alloc, index); =20 page_to_free =3D alloc->pages[index]; @@ -1182,7 +1179,12 @@ enum lru_status binder_alloc_free_page(struct list_h= ead *item, list_lru_isolate(lru, item); spin_unlock(&lru->lock); =20 - if (vma) { + /* + * Since a binder_alloc can only be mapped once, we ensure + * the vma corresponds to this mapping by checking whether + * the binder_alloc is still mapped. + */ + if (vma && binder_alloc_is_mapped(alloc)) { trace_binder_unmap_user_start(alloc, index); =20 zap_vma_range(vma, page_addr, PAGE_SIZE); @@ -1191,23 +1193,17 @@ enum lru_status binder_alloc_free_page(struct list_= head *item, } =20 mutex_unlock(&alloc->mutex); - if (mm_locked) - mmap_read_unlock(mm); - else + if (vma) vma_end_read(vma); mmput_async(mm); binder_free_page(page_to_free); =20 return LRU_REMOVED_RETRY; =20 -err_invalid_vma: - mutex_unlock(&alloc->mutex); err_get_alloc_mutex_failed: - if (mm_locked) - mmap_read_unlock(mm); - else + if (vma) vma_end_read(vma); -err_mmap_read_lock_failed: +err_vma_lock_failed: mmput_async(mm); err_mmget: return LRU_SKIP; --=20 2.55.0.966.g6673acef38-goog From nobody Sat Sep 26 13:47:00 2026 Received: from mail-pj1-f71.google.com (mail-pj1-f71.google.com [209.85.216.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F2BC83911C9 for ; Mon, 31 Aug 2026 20:31:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.71 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208271; cv=none; b=jHAr1gXPbcRrPUzkBG53NxtPQWfblqcKNfhIvvL3wkp1bB62PJSWyIB6Db9jFvpExZjCiCFISgEISizzCNdHQakGT32yJX8Q3D0afL5cRJ2lZiBY+fzUG2Gseh2zU4lLDdS1MK56mfV6Z6exGMJ5TEZcYHEwZ0+1svFYzBT7+20= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208271; c=relaxed/simple; bh=0yZiSioIL8z81HUTM4QdI25XMPcYXqTlcPZ2v8++geA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=FMODb4HXmgjXAPwIqy5lZ0MKTsm5sZSluPH2eJsGDvU0pt/2PGeEFml3G4ORNH8+QOw7LmR4KqAnkxSE4o/Jjg4rviNTqSroKOJnwTO/gDUMXVf5TyP8/sH66Z9UQD2IdV0Fipx1sGVB5hirSnuSt+SlasTfgXB0Sw5xNHrft8s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=if6LhE4h; arc=none smtp.client-ip=209.85.216.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="if6LhE4h" Received: by mail-pj1-f71.google.com with SMTP id 98e67ed59e1d1-398ba5404d2so3068736a91.1 for ; Mon, 31 Aug 2026 13:31:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788208269; x=1788813069; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=2d63lXSBmx7NRIUq5ZJpOYybhygf0jA0W5buHn0/eWM=; b=if6LhE4hEY5TUG9h4NJHIj/PXTGXgQcjOooFA7otSYa7YchEDHN9ioKHLd1mBPsEew uXTU5o8cihMu/DvcG7l7aoO2FKbShGKQzvkTEe+2/0D6i35cF/zki0ZgMkQwiOX/N2GH AT51yjzSG1MqreVULMBz3bPYCihKPfWj6hosWSGx/LqcgTtErLJQ9hdskrWh9MC9P105 cPkR5Eex2Np+oGNFiKy5tanTogmVzVYXJ/401SaPt+Hx5Bokoebzgs3ooKxfkfctv3i1 iPbASbYtW8Mj4mo/w8U3waHYiMnQP5haSysIcvm/Dx4JnF6I2XRAcMEMd3Q7NFC7ghuN 2Yzw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788208269; x=1788813069; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=2d63lXSBmx7NRIUq5ZJpOYybhygf0jA0W5buHn0/eWM=; b=OfYDdmOFa0bbPhn1oybQcWVqhj9vW4g6KmYXNi2KxfX0FeKQVXUVEjI88ark7GlYcc /UpgiblNajRbhIT9RED/1Xueyf7d5dhdAd+OM0r4dQbPXJhAG9+Xu+zmwVrqozK9ZI2W zJcwVYjPQIOmumyP4mxlAGZz5Z1XBeVtpxa4JToYOpbfOW9n7790h0aVnIoKnNRNxfoY r7Pa8fCViGN1hEHw3TeLclzphCMYVsxXrN07QeKLHnv+JuNkrpW3JMv8LlEzw3aFogVJ MElUf3yI3JmunMyeB1v62T3hkgNQqXQflMj7DFluV4fhCBnWUrssVAJzijbaXS4U0V1C gAXg== X-Forwarded-Encrypted: i=1; AKwUvBydwSURPmVkr03yqBfQJfWa96VD4Zi+w2ArDnbkFUTL5PuaaOpkiBOQ67wDhvja072D10XW4TBuCmZDqag=@vger.kernel.org X-Gm-Message-State: AFuF++nIpKP+jhUklC45G/0UqzRNZUany+f1T3wmYZxHt1/HjQpA25CV AGHY4T9qTJJRbMQu/V8eXtZG2hC5vglUu+i9DzbcaUQsYY/rXMVuL62sEiI95YhS81QEqSYI5Tc SpKC0hg== X-Received: from dlbeg32.prod.google.com ([2002:a05:7022:fa0:b0:141:4d1b:41b8]) (user=surenb job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:17ca:b0:395:4de5:1054 with SMTP id 98e67ed59e1d1-39907d65648mr3496744a91.16.1788208268999; Mon, 31 Aug 2026 13:31:08 -0700 (PDT) Date: Mon, 31 Aug 2026 13:30:54 -0700 In-Reply-To: <20260831203056.838265-1-surenb@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260831203056.838265-1-surenb@google.com> X-Mailer: git-send-email 2.55.0.966.g6673acef38-goog Message-ID: <20260831203056.838265-4-surenb@google.com> Subject: [PATCH v7 3/5] mm: Add RCU-based VMA lookup helper that waits for writers From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: dave.hansen@linux.intel.com, Liam.Howlett@oracle.com, ljs@kernel.org, david@redhat.com, willy@infradead.org, shakeel.butt@linux.dev, vbabka@kernel.org, jannh@google.com, aliceryhl@google.com, arve@android.com, cmllamas@google.com, christian@brauner.io, tkjos@android.com, dsahern@kernel.org, davem@davemloft.net, gregkh@linuxfoundation.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, netdev@vger.kernel.org, surenb@google.com Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Dave Hansen There are basically two parallel ways to look up a VMA: the traditional way, which is protected by mmap_read_lock, and the RCU-based per-VMA lock way which is based on RCU and refcounts. However, per-VMA locks will fail if the lock is help by a writer and therefore never waits. In a number of places we need to wait for the lock and it's done by falling back to mmap_read_lock, locking the VMA and releasing the mmap_lock once VMA is locked. Add vma_start_read_unlocked() - a variant of the RCU-based lookup that waits for writers. This is basically the same as the existing RCU-based lookup, but on a failure to lock it temporarily takes mmap_lock for read and waits for writers to finish before locking the VMA, dropping the mmap_read_lock and returning the locked VMA. This has some advantages: 1. Callers do not need to have a fallback path for when they collide with writers. 2. Its fast path does not require taking mmap_lock for read. Basically, when applied correctly, this approach results in faster *and* simpler code. While at it, fix the comments for vma_start_read_locked(), vma_start_read_locked_nested(), and uffd_lock_vma(). Suggested-by: Lorenzo Stoakes (ARM) Signed-off-by: Dave Hansen Cc: Suren Baghdasaryan Cc: Andrew Morton Cc: Liam R. Howlett Cc: Lorenzo Stoakes Cc: Vlastimil Babka Cc: Shakeel Butt Cc: linux-mm@kvack.org Cc: Greg Kroah-Hartman Cc: Arve Hj=C3=B8nnev=C3=A5g Cc: Todd Kjos Cc: Christian Brauner Cc: Carlos Llamas Cc: Alice Ryhl Cc: David S. Miller Cc: David Ahern Cc: netdev@vger.kernel.org Reviewed-by: Lorenzo Stoakes (ARM) Acked-by: Vlastimil Babka (SUSE) Signed-off-by: Suren Baghdasaryan --- include/linux/mmap_lock.h | 19 +++++++++++++++---- mm/mmap_lock.c | 35 +++++++++++++++++++++++++++++++++++ mm/userfaultfd.c | 6 ++++-- 3 files changed, 54 insertions(+), 6 deletions(-) diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 27c00bac29f9..00eae65b74bd 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -228,10 +228,14 @@ static inline void vma_refcount_put(struct vm_area_st= ruct *vma) } =20 /* - * Use only while holding mmap read lock which guarantees that locking wil= l not - * fail (nobody can concurrently write-lock the vma). vma_start_read() sho= uld + * Use only while holding mmap read lock which guarantees that vma lock is= not + * contended (nobody can concurrently write-lock the vma). vma_start_read(= ) should * not be used in such cases because it might fail due to mm_lock_seq over= flow. * This functionality is used to obtain vma read lock and drop the mmap re= ad lock. + * + * VMA can't be detached while we are holding mmap lock, therefore in prac= tice this + * function can fail only when there are so many readers that vm_refcnt ov= erflows. + * The failure case is very unlikely and is already annotated as such inte= rnally. */ static inline bool vma_start_read_locked_nested(struct vm_area_struct *vma= , int subclass) { @@ -247,16 +251,23 @@ static inline bool vma_start_read_locked_nested(struc= t vm_area_struct *vma, int } =20 /* - * Use only while holding mmap read lock which guarantees that locking wil= l not - * fail (nobody can concurrently write-lock the vma). vma_start_read() sho= uld + * Use only while holding mmap read lock which guarantees that vma lock is= not + * contended (nobody can concurrently write-lock the vma). vma_start_read(= ) should * not be used in such cases because it might fail due to mm_lock_seq over= flow. * This functionality is used to obtain vma read lock and drop the mmap re= ad lock. + * + * VMA can't be detached while we are holding mmap lock, therefore in prac= tice this + * function can fail only when there are so many readers that vm_refcnt ov= erflows. + * The failure case is very unlikely and is already annotated as such inte= rnally. */ static inline bool vma_start_read_locked(struct vm_area_struct *vma) { return vma_start_read_locked_nested(vma, 0); } =20 +struct vm_area_struct *vma_start_read_unlocked(struct mm_struct *mm, + unsigned long address); + static inline void vma_end_read(struct vm_area_struct *vma) { vma_refcount_put(vma); diff --git a/mm/mmap_lock.c b/mm/mmap_lock.c index 272f9ac762b9..2f94ee0fdee2 100644 --- a/mm/mmap_lock.c +++ b/mm/mmap_lock.c @@ -340,6 +340,41 @@ struct vm_area_struct *lock_vma_under_rcu(struct mm_st= ruct *mm, return NULL; } =20 +/** + * vma_start_read_unlocked() - Find the VMA covering 'address' and read-lo= ck it. + * @mm: the mm_struct of the address space to search + * @address: address that the vma should contain + * + * The fast path does not take mmap_lock. Waits for writers to finish if t= he + * VMA is being modified by taking mmap_lock. + * Use when mmap_lock is not held, otherwise use vma_start_read_locked(). + * Nothing prevents VMAs being unmapped/mapped before or after the VMA is + * looked up, if a stronger guarantee is required, take an mmap_lock. + * + * Return: If a VMA exists which spans @address, return that VMA, read-loc= ked. + * If no VMA is mapped there or, very unlikely, a reference count overflow + * occurred, return NULL. + */ +struct vm_area_struct *vma_start_read_unlocked(struct mm_struct *mm, + unsigned long address) +{ + struct vm_area_struct *vma; + + /* Fast path: return stable VMA covering 'address': */ + vma =3D lock_vma_under_rcu(mm, address); + if (vma) + return vma; + + /* Slow path: preclude VMA writers by temporarily getting mmap read lock.= */ + mmap_read_lock(mm); + vma =3D vma_lookup(mm, address); + if (vma && !vma_start_read_locked(vma)) + vma =3D NULL; + mmap_read_unlock(mm); + + return vma; +} + static struct vm_area_struct *lock_next_vma_under_mmap_lock(struct mm_stru= ct *mm, struct vma_iterator *vmi, unsigned long from_addr) diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index d157e15fc77c..6ee933a3803c 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -129,8 +129,10 @@ struct vm_area_struct *find_vma_and_prepare_anon(struc= t mm_struct *mm, * * Should be called without holding mmap_lock. * - * Return: A locked vma containing @address, -ENOENT if no vma is found, or - * -ENOMEM if anon_vma couldn't be allocated. + * Return: A locked vma containing @address, -ENOENT if no vma is found, + * -ENOMEM if anon_vma couldn't be allocated, or -EAGAIN if vma refcount + * overflow happened due to high number of readers and the caller should + * retry later. */ static struct vm_area_struct *uffd_lock_vma(struct mm_struct *mm, unsigned long address) --=20 2.55.0.966.g6673acef38-goog From nobody Sat Sep 26 13:47:00 2026 Received: from mail-pj1-f71.google.com (mail-pj1-f71.google.com [209.85.216.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AEF6B3932D8 for ; Mon, 31 Aug 2026 20:31:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.71 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208276; cv=none; b=WE6sgiCcp/v8ediEgrKTHwYckMCddAwoN1ICTNPYtsExbptOjAobsfVy59ev/bxUD6mJ/gfjGbpiNO99bLrEe5mepoJI/aXAKOHdXoKLyRY+9lxpLpXtjwV3LOKeYTCE8erqhJizsciAjV23N0qv7GSa7YpkqQdnqr0rlwG98Vg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208276; c=relaxed/simple; bh=E5omcHuGcH/5sMqY70uPWGBn/Pv7a09NS/jGJbvxUDA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=NVXPl1FmZqZa3R66TFjoQKzZYM3R28TwXaEBMcGr6hr0Wc69w4VVvD2rISZL6C8M9H1CUKjr3qxPD0xxblw4dM3pCZWny+oF2lXPUE/4/QPjfGp04AV+qwtoGXKDCSho8GeJhPZX1Td9+b6J1kyOh+qkTdOPEVbphVoatQqnZfM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=UwjO+kRP; arc=none smtp.client-ip=209.85.216.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="UwjO+kRP" Received: by mail-pj1-f71.google.com with SMTP id 98e67ed59e1d1-38e7ff7b375so229536a91.1 for ; Mon, 31 Aug 2026 13:31:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788208272; x=1788813072; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=fMb1BIWAt7k71K1WS7NJDkr5bMmZbsfs3e75iDz3R4U=; b=UwjO+kRPdRYCN7XOF0SKWzCdwZYzLGaLUA83mFXoGZnSam+2QZ+MsWDfjCBo1gkpAo SyQerbaqC0HXCD9H2GvE56Z4tfLD0Z5C9p655IDZYQbiM1U6dH/fr6+ARJ4lQTBCX4Ms aACs42ZrAFCCGFELTv5+8OaqVR22S3KigQyE4H1lpL2qIH0yiMZP0e1p1HuGKBh+NokL tsglwUqoIzkzkMW/fE2CoEKGcjfw7MqvEumtSU3naNJx1i5WCP5bbXSCgJLXFkqh7hDD OKrta5LeTC0s+YDVSpCmb2ycN/7dyWzBDblWgOwtNG19/CmaWP+ciTAJSkbwy78dwr+B P3KQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788208272; x=1788813072; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=fMb1BIWAt7k71K1WS7NJDkr5bMmZbsfs3e75iDz3R4U=; b=RlpoabwWb/9uCpyqqhBOV4koaWX26qNafy5VTOd8G1ZBdho2yyBBHG+pjGE4JE4d9R hLEWN0Jd7NCJuEGgzIqR+n9mvEMpBgqjzZP3VewlYztnGjWecD6qPEp/qoiQlNp9Edx3 hmHuo8QYF1D0mGVNe5aN9n2VthnTs1t23PmE3zufqji4Wk7kbRCZJhjmqzHZN7moMhKk wzgONWK5DKDS53Mt0Cf5COHRoBvdUEdl9k/4+s/irjCxnXivIDyE9/vtvFb/wpFYesgp JC/v0gx4DbaYyu6VniPxYBG+/tQdWSmtD6ftEO2cp8I9pNV8TW0r6gnUDL9SiMTPgu4k xweA== X-Forwarded-Encrypted: i=1; AKwUvBxtauCzhAweP9prGHzIeF7dziF+xbMKxsACiMq/Z3mL5/0ErBNvONRe0BlqCPSY1ljlOa/TUY/ZymCUTJE=@vger.kernel.org X-Gm-Message-State: AFuF++miTfnsXs9VyGBdmHp0EXB9k54Eu+mzBBvurX5LYGp3Lv2xxAb/ nLlcJjTiWFxbJxwvGOYOvzNgGpx3cXiELfjqz54iE1fPUlS9gmzMVvqP+M34CQpM/K589D29UCa 1Zkscgg== X-Received: from dlak6.prod.google.com ([2002:a05:701b:2906:b0:141:5987:7f07]) (user=surenb job=prod-delivery.src-stubby-dispatcher) by 2002:a17:90b:3c8a:b0:38e:524:8797 with SMTP id 98e67ed59e1d1-396d10076e9mr43087877a91.13.1788208271612; Mon, 31 Aug 2026 13:31:11 -0700 (PDT) Date: Mon, 31 Aug 2026 13:30:55 -0700 In-Reply-To: <20260831203056.838265-1-surenb@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260831203056.838265-1-surenb@google.com> X-Mailer: git-send-email 2.55.0.966.g6673acef38-goog Message-ID: <20260831203056.838265-5-surenb@google.com> Subject: [PATCH v7 4/5] binder: Remove mmap_lock fallback From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: dave.hansen@linux.intel.com, Liam.Howlett@oracle.com, ljs@kernel.org, david@redhat.com, willy@infradead.org, shakeel.butt@linux.dev, vbabka@kernel.org, jannh@google.com, aliceryhl@google.com, arve@android.com, cmllamas@google.com, christian@brauner.io, tkjos@android.com, dsahern@kernel.org, davem@davemloft.net, gregkh@linuxfoundation.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, netdev@vger.kernel.org, surenb@google.com Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Dave Hansen Previously, the per-VMA locking could fail in the face of writers which necessitate a fallback to mmap_lock. The new vma_start_read_unlocked() will wait for writers instead of failing. Use the new helper. Wait for writers. Remove the fallback to mmap_lock. Signed-off-by: Dave Hansen Reviewed-by: Alice Ryhl Acked-by: Lorenzo Stoakes (ARM) Cc: Andrew Morton Cc: Liam R. Howlett Cc: Vlastimil Babka Cc: Shakeel Butt Cc: linux-mm@kvack.org Cc: Greg Kroah-Hartman Cc: Arve Hj=C3=B8nnev=C3=A5g Cc: Todd Kjos Cc: Christian Brauner Cc: Carlos Llamas Cc: Alice Ryhl Cc: David S. Miller Cc: David Ahern Cc: netdev@vger.kernel.org Signed-off-by: Suren Baghdasaryan --- drivers/android/binder/page_range.rs | 19 +++---------------- drivers/android/binder_alloc.c | 17 +++++------------ rust/kernel/mm.rs | 28 ++++++++++++++++++++++++++++ 3 files changed, 36 insertions(+), 28 deletions(-) diff --git a/drivers/android/binder/page_range.rs b/drivers/android/binder/= page_range.rs index 52ffbf3504e7..71febd3d5b07 100644 --- a/drivers/android/binder/page_range.rs +++ b/drivers/android/binder/page_range.rs @@ -439,22 +439,9 @@ unsafe fn use_page_slow(&self, i: usize) -> Result<()>= { // workqueue. let mm =3D MmWithUser::into_mmput_async(self.mm.mmget_not_zero().o= k_or(ESRCH)?); { - let vma_read; - let mmap_read; - let vma =3D if let Some(ret) =3D mm.lock_vma_under_rcu(vma_add= r) { - vma_read =3D ret; - check_vma(&vma_read, self) - } else { - mmap_read =3D mm.mmap_read_lock(); - mmap_read - .vma_lookup(vma_addr) - .and_then(|vma| check_vma(vma, self)) - }; - - match vma { - Some(vma) =3D> vma.vm_insert_page(user_page_addr, &new_pag= e)?, - None =3D> return Err(ESRCH), - } + let vma_read_guard =3D mm.vma_start_read_unlocked(vma_addr).ok= _or(ESRCH)?; + let vma =3D check_vma(&vma_read_guard, self).ok_or(ESRCH)?; + vma.vm_insert_page(user_page_addr, &new_page)?; } =20 let inner =3D self.lock.lock(); diff --git a/drivers/android/binder_alloc.c b/drivers/android/binder_alloc.c index fcb744088e77..d6eae0aa7085 100644 --- a/drivers/android/binder_alloc.c +++ b/drivers/android/binder_alloc.c @@ -259,21 +259,14 @@ static int binder_page_insert(struct binder_alloc *al= loc, struct vm_area_struct *vma; int ret =3D -ESRCH; =20 - /* attempt per-vma lock first */ - vma =3D lock_vma_under_rcu(mm, addr); - if (vma) { - if (binder_alloc_is_mapped(alloc)) - ret =3D vm_insert_page(vma, addr, page); - vma_end_read(vma); + vma =3D vma_start_read_unlocked(mm, addr); + if (!vma) return ret; - } =20 - /* fall back to mmap_lock */ - mmap_read_lock(mm); - vma =3D vma_lookup(mm, addr); - if (vma && binder_alloc_is_mapped(alloc)) + if (binder_alloc_is_mapped(alloc)) ret =3D vm_insert_page(vma, addr, page); - mmap_read_unlock(mm); + + vma_end_read(vma); =20 return ret; } diff --git a/rust/kernel/mm.rs b/rust/kernel/mm.rs index f4fa54616085..58bc1793fdaf 100644 --- a/rust/kernel/mm.rs +++ b/rust/kernel/mm.rs @@ -186,6 +186,34 @@ pub fn lock_vma_under_rcu(&self, vma_addr: usize) -> O= ption> { }) } =20 + /// Find the VMA covering 'address' and read-lock it. + /// + /// The fast path does not take mmap_lock. Waits for writers to finish= if the + /// VMA is being modified by taking mmap_lock. + /// Use when mmap_lock is not held, otherwise use vma_start_read_locke= d(). + /// Nothing prevents VMAs being unmapped/mapped before or after the VM= A is + /// looked up, if a stronger guarantee is required, take an mmap_lock. + /// + /// Return: If a VMA exists which spans @address, return that VMA, rea= d-locked. + /// If no VMA is mapped there or, very unlikely, a reference count ove= rflow + /// occurred, return NULL. + #[inline] + pub fn vma_start_read_unlocked(&self, vma_addr: usize) -> Option> { + // SAFETY: We may invoke `vma_start_read_unlocked` because we know= this `mm` has non-zero + // `mm_users`. + let vma =3D unsafe { bindings::vma_start_read_unlocked(self.as_raw= (), vma_addr) }; + if vma.is_null() { + return None; + } + // INVARIANT: We just acquired the VMA read lock. + Some(VmaReadGuard { + // SAFETY: If `vma_start_read_unlocked` returns a non-null ptr= , then it points at a + // valid vma. The vma is stable for as long as the vma read lo= ck is held. + vma: unsafe { VmaRef::from_raw(vma) }, + _nts: NotThreadSafe, + }) + } + /// Lock the mmap read lock. #[inline] pub fn mmap_read_lock(&self) -> MmapReadGuard<'_> { --=20 2.55.0.966.g6673acef38-goog From nobody Sat Sep 26 13:47:00 2026 Received: from mail-pl1-f200.google.com (mail-pl1-f200.google.com [209.85.214.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9779839524B for ; Mon, 31 Aug 2026 20:31:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.200 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208277; cv=none; b=YPYmY1tPA9d77Vdka8ZIqsSoqjPD6XZ+LVILAQEviOk+vmi5d/3AmXe+p8DkCWYhtlfBVgzft0doHHXjowQ0epoX75IlxREgpgZqahq5KfHfAmrRP+qBrGrFQzuld86KJHaPHXA36DN6Zkuy1GXIKH8DcfE5zcjhAGqmS1VNcWg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788208277; c=relaxed/simple; bh=H5Lc6BdEbXCjpa9lGiNJgyEyPhusZ9dehbsNMdHLR0w=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=ofkUIeuDNdgapDGOtsrBan6+F+PX/MHybvq93EeLz9ot+YmZ/JHueHo0lXGlnncM80ZenmPV3LQtSDQL/nWoRSGj7wrf7YAfuaeS4zFvCkj/MkjBOuOCkw/qOegzZhj2kaTR2HtLIF8ycJPKj0UJdr3Ium+O5yA6hxDXRJv9kRQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=ZBLLhBHv; arc=none smtp.client-ip=209.85.214.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--surenb.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="ZBLLhBHv" Received: by mail-pl1-f200.google.com with SMTP id d9443c01a7336-2d54187d8b0so1476615ad.0 for ; Mon, 31 Aug 2026 13:31:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1788208274; x=1788813074; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=ELv/N+6IjeKMBTz6azi//X4BSCo29f+bI+vir4cQf2g=; b=ZBLLhBHvOCgw7kUF5dNd2jmtO+8THjspA1cTInUjnr87Ed2Ovju96LfdVPFxSbAS14 KRKMo9vlSmYswJ26SLiM7Et/xnvVkyUZXmabcktlbxKkYQBDk1rOdpdC9HBocMYRb4SG tCqa6YhN7PoqsFaFPK9ojkQNLmQh4MicQCb8UN3ZutCoceHEXNwB/Pv0P2cCRm9CLGW2 T/ebZLjCYNJykqeKx+vbXKyWQnapnEbVZI23xa3xEVOZ4jY3eKVuAbkgXl0wi1ZbNjAJ aCfCEmMuitYlWHk8kIvPo8vazc88Od7+jJ/Jq6bmZfz5I0oZfF9ocsBhuis4nU/ZVMhF ehow== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788208274; x=1788813074; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=ELv/N+6IjeKMBTz6azi//X4BSCo29f+bI+vir4cQf2g=; b=cHWJZfprxGzxYM4nkumpB5qS0MixG0zjSXrOcsPPBF/dhUK60TfMJMbo81aEBVlRx+ F057XrpDk+5b/XTlnvnWtv3b0U7AFrxexgDPR2Gwu038vvueVmxjKx4gLhfzlBaB4IZ+ ++F8AMVzImlAlUdIVJI5a0WABGAv23+iS7cKAW4Sfct7TD1eCira3S8M4ZD5WVd8iLDl Xllu5zvqKcAvV3RWEcnICHcM1WiFCVPK4B9Qt7GiFtplZcjrKhDpGROm6F8FDkzzRTb4 maHLjTk1kQ5FfsOsTIumQDYk6MsnlPUh8AZ/P4zrHuLSrsoOri7rh7RGTFgo31qjNbca xkpw== X-Forwarded-Encrypted: i=1; AHgh+RrnbvQ9GyT0vU6E03GqKctST9s36aODjocRvIxox/TKOZUEN7OjAGItvsNRoxApkSppNNT7+QrGRTD+nTk=@vger.kernel.org X-Gm-Message-State: AFuF++ld3u2+OM8FJ5iYs5/YQTu68sxhpPU94ZgfrzIdVMnA9mfxyNjA F5jLCrrmVGUgW3jhM5/BA5tMMspMpxdevbpNcrxdCyziXVMJgMK3L0kRjfCjllhcralSa3Wad8k b1tsZYQ== X-Received: from dlbrl15.prod.google.com ([2002:a05:7022:f50f:b0:13e:5fe8:61ef]) (user=surenb job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:a93:b0:3d1:e510:9052 with SMTP id adf61e73a8af0-3d265e3f792mr48387570637.7.1788208274250; Mon, 31 Aug 2026 13:31:14 -0700 (PDT) Date: Mon, 31 Aug 2026 13:30:56 -0700 In-Reply-To: <20260831203056.838265-1-surenb@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260831203056.838265-1-surenb@google.com> X-Mailer: git-send-email 2.55.0.966.g6673acef38-goog Message-ID: <20260831203056.838265-6-surenb@google.com> Subject: [PATCH v7 5/5] tcp: Remove mmap_lock fallback path From: Suren Baghdasaryan To: akpm@linux-foundation.org Cc: dave.hansen@linux.intel.com, Liam.Howlett@oracle.com, ljs@kernel.org, david@redhat.com, willy@infradead.org, shakeel.butt@linux.dev, vbabka@kernel.org, jannh@google.com, aliceryhl@google.com, arve@android.com, cmllamas@google.com, christian@brauner.io, tkjos@android.com, dsahern@kernel.org, davem@davemloft.net, gregkh@linuxfoundation.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, netdev@vger.kernel.org, surenb@google.com, syzbot@syzkaller.appspotmail.com Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Dave Hansen Previously, the per-VMA locking could fail in the face of writers which necessitates a fallback to mmap_lock. The new vma_start_read_unlocked() will wait for writers instead of failing. Use the new helper. Wait for writers. Remove the fallback to mmap_lock. The fallback removal does not affect NOMMU case because TCP_ZEROCOPY is gated on CONFIG_MMU. This really is a nice cleanup. It removes the need to pass the lock state back and forth to find_tcp_vma(). Signed-off-by: Dave Hansen Acked-by: Lorenzo Stoakes Acked-by: Vlastimil Babka (SUSE) Tested-by: syzbot@syzkaller.appspotmail.com Cc: Andrew Morton Cc: Liam R. Howlett Cc: Vlastimil Babka Cc: Shakeel Butt Cc: linux-mm@kvack.org Cc: Greg Kroah-Hartman Cc: Arve Hj=C3=B8nnev=C3=A5g Cc: Todd Kjos Cc: Christian Brauner Cc: Carlos Llamas Cc: Alice Ryhl Cc: David S. Miller Cc: David Ahern Cc: netdev@vger.kernel.org Signed-off-by: Suren Baghdasaryan --- net/ipv4/tcp.c | 31 +++++++++---------------------- 1 file changed, 9 insertions(+), 22 deletions(-) diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c index 455441f1b694..62defe70f3ce 100644 --- a/net/ipv4/tcp.c +++ b/net/ipv4/tcp.c @@ -2167,27 +2167,18 @@ static void tcp_zc_finalize_rx_tstamp(struct sock *= sk, } =20 static struct vm_area_struct *find_tcp_vma(struct mm_struct *mm, - unsigned long address, - bool *mmap_locked) + unsigned long address) { - struct vm_area_struct *vma =3D lock_vma_under_rcu(mm, address); + struct vm_area_struct *vma =3D vma_start_read_unlocked(mm, address); =20 - if (vma) { - if (vma->vm_ops !=3D &tcp_vm_ops) { - vma_end_read(vma); - return NULL; - } - *mmap_locked =3D false; - return vma; - } + if (!vma) + return NULL; =20 - mmap_read_lock(mm); - vma =3D vma_lookup(mm, address); - if (!vma || vma->vm_ops !=3D &tcp_vm_ops) { - mmap_read_unlock(mm); + if (vma->vm_ops !=3D &tcp_vm_ops) { + vma_end_read(vma); return NULL; } - *mmap_locked =3D true; + return vma; } =20 @@ -2208,7 +2199,6 @@ static int tcp_zerocopy_receive(struct sock *sk, u32 seq =3D tp->copied_seq; u32 total_bytes_to_map; int inq =3D tcp_inq(sk); - bool mmap_locked; int ret; =20 zc->copybuf_len =3D 0; @@ -2233,7 +2223,7 @@ static int tcp_zerocopy_receive(struct sock *sk, return 0; } =20 - vma =3D find_tcp_vma(current->mm, address, &mmap_locked); + vma =3D find_tcp_vma(current->mm, address); if (!vma) return -EINVAL; =20 @@ -2315,10 +2305,7 @@ static int tcp_zerocopy_receive(struct sock *sk, zc, total_bytes_to_map); } out: - if (mmap_locked) - mmap_read_unlock(current->mm); - else - vma_end_read(vma); + vma_end_read(vma); /* Try to copy straggler data. */ if (!ret) copylen =3D tcp_zc_handle_leftover(zc, sk, skb, &seq, copybuf_len, tss); --=20 2.55.0.966.g6673acef38-goog