From nobody Tue Sep 29 00:36:44 2026 Received: from mail-wm2-f9.google.com (mail-wm2-f9.google.com [74.125.225.137]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8E46F2222C5 for ; Fri, 14 Aug 2026 08:55:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.137 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786697718; cv=none; b=Ivep1j4wccDyNbZNBXQIIyFCXkJt0dnoSo2fV9NYQhRT1bOewamnN9750a89fso6/uLGO59HQRm4/l4qCMSHdPy+8icPoxwCh9iZKd2zb/6R+gtkSXviVJ9Vv0fWJrOsxTRubQ69HlevJ3+4UjWVI7iVzhox04QUcgWeF5huVO4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786697718; c=relaxed/simple; bh=Y752F3HpLuE35vHFFRya1+CBBqX+7hjYThuLCnAzcug=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=h7GNImqlRombTejNUvtB8EB2WhVpE7T9MGRxj5UT7pFaFqc5uH6emnzC0p+SLLzc4wBF2Z8ZKk1lPtEOr+zuOUVWCIdMcz6ilPN3Yq+VTAKfj0HNqDz7Pi4+7u7jSDKYvzguqbEbdUQ7K/SClnhLlapOHTEbOIero9cj1QK8SkQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=kJIBjgDG; arc=none smtp.client-ip=74.125.225.137 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="kJIBjgDG" Received: by mail-wm2-f9.google.com with SMTP id 5b1f17b1804b1-49554457f71so515225e9.1 for ; Fri, 14 Aug 2026 01:55:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786697715; x=1787302515; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=M7Pbg3qHy9dUvw/2GBMFVLtK9eJ8QsOBHb0QGZCzF6E=; b=kJIBjgDGG6fbD+PgNbfPXevPrBggi79OpZyiWaT7F1cIscCIZdkk4HbVK6OR6SdvbV 0XfH0GEQmXi8pS9RldKte12YtOhEe6UbiYycXnUAERm8dcPH2pUNEM7nXMbS4Kweq+Do 2SVh9Gi+0sircJ7yXtQTJ9FdLeMqViHzaHYP7XUu6+Qbx9qxCxgu929ub7fY1RmCmdgS nKHyasQuNTIvTGTMvwqDKCAB2tRNMUdzYsAquXj7YfEJkyV9vUefAgLqxMphQUEZxMMR 5vwtNo6dFZeKRonTk4qG6AHcMCWgF4ap2nQMV3AvA/07Ku1+XcPClDXcVubgqNVlKlVL Ks+Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786697715; x=1787302515; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=M7Pbg3qHy9dUvw/2GBMFVLtK9eJ8QsOBHb0QGZCzF6E=; b=kcKbD5DCMkfrs6tbbXV3pTTebpdFdRss+Uc8GYwwygE7unJn8J8EXBOyO+U7AT7MDJ 9rMiSRSGPPDXd6rO2zCZPZuv+AF8bwaJ0m7DJWc9E+qettkLbRIwf6WoKpBd4f4NngXL oMGVEi1syaR9fwgZV6yEBViAfm6xWaCS3JXqDiZt4gxgOtZ0etUWVXVjcunTX1v/F6uZ WkvNTH9ay+kwt2zkLRLJMjzoC3QLdTJbBDHtrOlodILLxsxLjj8ZYkfL9Nm7NdAESUjI NIfb4S9KGlbrAWLezdbFc3gUw4Zyj2iW001ArCi5DsjRCFgpS7eTljzqM4EV8Key8Si3 pcKQ== X-Forwarded-Encrypted: i=1; AHgh+RrLZMta8ZPS9JHX4e+r8DtMSdQDbECgJ/1s80FrNPUDOHifqMOjsy7uBEjFtgr1jlm2cZyUwHb+O6QNtPA=@vger.kernel.org X-Gm-Message-State: AOJu0YyW8f5Bq+d39rFGy47Wu+OweyOfauqtp3Dz7Ya88Hu6GJizuLwU /mCOmgwlc5M7dk2fjV8V7I8PEbsr7qBQIMfcsoFUDRr2KC5kH34Jgs6K X-Gm-Gg: AR+sD10tcu3gbLoVCKa/HyzLQEiOWrAUnFOjOEebexCn31fpm2HSqB1910N0RXiN8TG nGifqvVTY0OKs402FAe139t7ZOTd/jv3gtvj8AvhvrdBGD/H02OjxZX4x8IAtmJKWYtnWf3araz fOO+UpbEY+2auo2MBqIBbfQN2rE5pjYmgbVIZsHd/8zlf80CO1/zXv37/u8/eDoc9+z6Ptr68he 1yGFpjUNx4va8BVRIuPO8wt9Og00SOsfjYkJzCQH0pR/R4HiyQRA9Ft031ZzoL9FZ3BB6MyEOUI 2enauC4kajp8CoQw4ATeEyDCbLYydJN1W/69mM+hds5kKkJoRtFO1E5V8KGayPocxrXB6X5ke+M Jhb1RNhH47SGKqhoeAkVj1p7/hCjjJDt+YcxfXvo9Zh/R5c4CdImWAzGNAfopbhtRE83jdCbprh U0jz6rAX485Zv+hXVaW+6uq6hdCc+f6xu/VwdM5Q1vmz38pxUlMkyhC9U4nlhR3ReVgw5t5wofa b2N7tersyhdUE8= X-Received: by 2002:a05:600c:4e14:b0:495:3bc6:d381 with SMTP id 5b1f17b1804b1-49987a4f03bmr50711495e9.2.1786697714519; Fri, 14 Aug 2026 01:55:14 -0700 (PDT) Received: from zhr-ThinkStation-K.mioffice.cn ([2408:8607:1b00:8:10f3:285:31cf:47aa]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4998777d428sm28760795e9.0.2026.08.14.01.55.08 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 14 Aug 2026 01:55:12 -0700 (PDT) From: Hongru Zhang X-Google-Original-From: Hongru Zhang To: akpm@linux-foundation.org, linux-mm@kvack.org Cc: david@kernel.org, liam@infradead.org, linux-kernel@vger.kernel.org, ljs@kernel.org, baohua@kernel.org, mhocko@suse.com, rppt@kernel.org, shakeel.butt@linux.dev, surenb@google.com, vbabka@kernel.org, willy@infradead.org, zhanghongru@xiaomi.com Subject: [RFC PATCH v5] mm: retry page faults once under the per-VMA lock Date: Fri, 14 Aug 2026 16:53:00 +0800 Message-ID: <20260814085300.399107-1-zhanghongru@xiaomi.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Hongru Zhang The per-VMA lock fault path falls back to mmap_lock on VM_FAULT_RETRY. When mmap_lock is write-contended, the fallback can reduce page-fault throughput. Add a single retry under the per-VMA lock in the arch fault handler instead of adding a new VM_FAULT_* flag. This avoids overloading vm_fault_t with retry policy and leaves the lower-level fault handlers unchanged. It lets faults that can make progress on an immediate retry stay on the per-VMA lock path, avoiding waits on mmap_lock when it is write-contended and thereby improving page-fault throughput. Some faults may retry unnecessarily, for example in the __vmf_anon_prepare() path or device-private fault handling, but these cases are expected to be infrequent and only add one cheap per-VMA lock attempt. If the second attempt still returns VM_FAULT_RETRY, the fault continues through the existing mmap_lock path. Based on the stress model from Kunwu Chan and Wang Lian in RFC v2, we adapted a benchmark [1] to a 20-core Intel i7-12700 desktop by reducing the thread count and adjusting the memcg limits. The benchmark uses concurrent page faults under memcg pressure with parallel munmap to amplify mmap_lock read-write contention. Filemap Throughput (higher is better): +---------+------------+---------------------+ | Threads | Vanilla | Patched | +---------+------------+---------------------+ | 40 | 1069.34 /s | 1400.13 /s (+30.9%) | +---------+------------+---------------------+ | 60 | 1038.12 /s | 1683.37 /s (+62.2%) | +---------+------------+---------------------+ | 80 | 1042.62 /s | 1767.83 /s (+69.6%) | +---------+------------+---------------------+ mmap_lock contention count (lower is better): +---------+-----------+---------+-----------+ | Threads | Vanilla | Patched | Reduction | +---------+-----------+---------+-----------+ | 40 | 3,187,336 | 52,086 | -98.4% | +---------+-----------+---------+-----------+ | 60 | 4,385,154 | 65,079 | -98.5% | +---------+-----------+---------+-----------+ | 80 | 5,337,890 | 69,708 | -98.7% | +---------+-----------+---------+-----------+ These results show that retrying once under the per-VMA lock keeps more file-backed faults on the fast path, improving throughput and reducing mmap_lock contention. Using benchmark [2], we tested this on a 20-core Intel i7-12700 desktop with a 2GB swapfile. The benchmark uses one pressure thread under memcg limits to keep a 128MB non-zero anonymous mapping under swap pressure, 12 reader threads to fault it back in, and optional mmap writer threads to amplify mmap_lock read-write contention. Each test ran for 60 seconds and reported completed reader rounds per second under swap pressure. Swap Throughput (higher is better): +--------------+-------------+---------------------------+ | mmap writers | Vanilla | Patched | +--------------+-------------+---------------------------+ | 0 | 17303.09 /s | 17899.48 /s (+3.4%) | +--------------+-------------+---------------------------+ | 4 | 12596.23 /s | 16095.20 /s (+27.8%) | +--------------+-------------+---------------------------+ | 8 | 0.58 /s | 15420.57 /s (+2658619.0%) | +--------------+-------------+---------------------------+ With increasing mmap_lock write pressure, Vanilla degrades sharply and drops to near zero at eight writers. Patched kernel holds up much better. [1] https://gist.github.com/zhr250/c36c2c54d9351df37e12fd072d4926ef [2] https://gist.github.com/zhr250/218ffe693f842346b56434483127422c Signed-off-by: Hongru Zhang Suggested-by: Barry Song Suggested-by: Suren Baghdasaryan Suggested-by: Lorenzo Stoakes (ARM) --- Changes since RFC v4: - Drop `VM_FAULT_MAY_USE_VMA_LOCK` and always retry once under the per-VMA lock, based on feedback from Lorenzo and Barry. Thanks! Changes since RFC v3: - Keep VM_FAULT_RETRY unchanged and add VM_FAULT_MAY_USE_VMA_LOCK as an adv= isory bit - Bound VMA-lock retries with FAULT_FLAG_TRIED - Opt in filemap_fault() and do_swap_page() to VM_FAULT_MAY_USE_VMA_LOCK - Rebased on mm-unstable Changes since RFC v2: - Redesigned as a single blacklist-based patch (v2 was 5 per-path patches) - Added retry_vma loop to all architectures (not just x86) - Rebased on mm-unstable Changes since RFC v1: - collect tags from Pedro, Kunwu and Lian, thanks! - handle case (2), for uptodate folios, don't retry PF Link to RFC v4: https://lore.kernel.org/lkml/20260804095135.45897-1-zhanghongru@xiaomi.com/ Link to RFC v3: https://lore.kernel.org/all/20260626075019.1833065-1-zhanghongru@xiaomi.com/ Link to RFC v2: https://lore.kernel.org/all/20260430040427.4672-1-baohua@kernel.org/ Link to RFC v1: https://lore.kernel.org/all/20251127011438.6918-1-21cnbao@gmail.com/ arch/arm/mm/fault.c | 8 ++++++++ arch/arm64/mm/fault.c | 8 ++++++++ arch/loongarch/mm/fault.c | 8 ++++++++ arch/powerpc/mm/fault.c | 7 +++++++ arch/riscv/mm/fault.c | 8 ++++++++ arch/s390/mm/fault.c | 6 ++++++ arch/x86/mm/fault.c | 8 ++++++++ 7 files changed, 53 insertions(+) diff --git a/arch/arm/mm/fault.c b/arch/arm/mm/fault.c index e62cc4be5adf..319fd89c4263 100644 --- a/arch/arm/mm/fault.c +++ b/arch/arm/mm/fault.c @@ -340,6 +340,7 @@ do_page_fault(unsigned long addr, unsigned int fsr, str= uct pt_regs *regs) vm_fault_t fault; unsigned int flags =3D FAULT_FLAG_DEFAULT; vm_flags_t vm_flags =3D VM_ACCESS_FLAGS; + bool vma_lock_retried =3D false; =20 if (kprobe_page_fault(regs, fsr)) return 0; @@ -391,6 +392,7 @@ do_page_fault(unsigned long addr, unsigned int fsr, str= uct pt_regs *regs) if (!(flags & FAULT_FLAG_USER)) goto lock_mmap; =20 +lock_vma: vma =3D lock_vma_under_rcu(mm, addr); if (!vma) goto lock_mmap; @@ -420,6 +422,12 @@ do_page_fault(unsigned long addr, unsigned int fsr, st= ruct pt_regs *regs) goto no_context; return 0; } + + if (!vma_lock_retried) { + vma_lock_retried =3D true; + goto lock_vma; + } + lock_mmap: =20 retry: diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c index 0b52557652be..b8633863380e 100644 --- a/arch/arm64/mm/fault.c +++ b/arch/arm64/mm/fault.c @@ -610,6 +610,7 @@ static int __kprobes do_page_fault(unsigned long far, u= nsigned long esr, struct vm_area_struct *vma; int si_code; int pkey =3D -1; + bool vma_lock_retried =3D false; =20 if (kprobe_page_fault(regs, esr)) return 0; @@ -678,6 +679,7 @@ static int __kprobes do_page_fault(unsigned long far, u= nsigned long esr, if (!(mm_flags & FAULT_FLAG_USER)) goto lock_mmap; =20 +lock_vma: vma =3D lock_vma_under_rcu(mm, addr); if (!vma) goto lock_mmap; @@ -724,6 +726,12 @@ static int __kprobes do_page_fault(unsigned long far, = unsigned long esr, goto no_context; return 0; } + + if (!vma_lock_retried) { + vma_lock_retried =3D true; + goto lock_vma; + } + lock_mmap: =20 retry: diff --git a/arch/loongarch/mm/fault.c b/arch/loongarch/mm/fault.c index 2c93d33356e5..ef6ea847b1e0 100644 --- a/arch/loongarch/mm/fault.c +++ b/arch/loongarch/mm/fault.c @@ -181,6 +181,7 @@ static void __kprobes __do_page_fault(struct pt_regs *r= egs, struct mm_struct *mm =3D tsk->mm; struct vm_area_struct *vma =3D NULL; vm_fault_t fault; + bool vma_lock_retried =3D false; =20 if (kprobe_page_fault(regs, current->thread.trap_nr)) return; @@ -219,6 +220,7 @@ static void __kprobes __do_page_fault(struct pt_regs *r= egs, if (!(flags & FAULT_FLAG_USER)) goto lock_mmap; =20 +lock_vma: vma =3D lock_vma_under_rcu(mm, address); if (!vma) goto lock_mmap; @@ -265,6 +267,12 @@ static void __kprobes __do_page_fault(struct pt_regs *= regs, no_context(regs, write, address); return; } + + if (!vma_lock_retried) { + vma_lock_retried =3D true; + goto lock_vma; + } + lock_mmap: =20 retry: diff --git a/arch/powerpc/mm/fault.c b/arch/powerpc/mm/fault.c index 806c74e0d5ab..06018b6d7086 100644 --- a/arch/powerpc/mm/fault.c +++ b/arch/powerpc/mm/fault.c @@ -422,6 +422,7 @@ static int ___do_page_fault(struct pt_regs *regs, unsig= ned long address, int is_write =3D page_fault_is_write(error_code); vm_fault_t fault, major =3D 0; bool kprobe_fault =3D kprobe_page_fault(regs, 11); + bool vma_lock_retried =3D false; =20 if (unlikely(debugger_fault_handler(regs) || kprobe_fault)) return 0; @@ -487,6 +488,7 @@ static int ___do_page_fault(struct pt_regs *regs, unsig= ned long address, if (!(flags & FAULT_FLAG_USER)) goto lock_mmap; =20 +lock_vma: vma =3D lock_vma_under_rcu(mm, address); if (!vma) goto lock_mmap; @@ -517,6 +519,11 @@ static int ___do_page_fault(struct pt_regs *regs, unsi= gned long address, if (fault_signal_pending(fault, regs)) return user_mode(regs) ? 0 : SIGBUS; =20 + if (!vma_lock_retried) { + vma_lock_retried =3D true; + goto lock_vma; + } + lock_mmap: =20 /* When running in the kernel we expect faults to occur only to diff --git a/arch/riscv/mm/fault.c b/arch/riscv/mm/fault.c index 04ed6f8acae4..ff861793dba9 100644 --- a/arch/riscv/mm/fault.c +++ b/arch/riscv/mm/fault.c @@ -284,6 +284,7 @@ void handle_page_fault(struct pt_regs *regs) unsigned int flags =3D FAULT_FLAG_DEFAULT; int code =3D SEGV_MAPERR; vm_fault_t fault; + bool vma_lock_retried =3D false; =20 cause =3D regs->cause; addr =3D regs->badaddr; @@ -347,6 +348,7 @@ void handle_page_fault(struct pt_regs *regs) if (!(flags & FAULT_FLAG_USER)) goto lock_mmap; =20 +lock_vma: vma =3D lock_vma_under_rcu(mm, addr); if (!vma) goto lock_mmap; @@ -376,6 +378,12 @@ void handle_page_fault(struct pt_regs *regs) no_context(regs, addr); return; } + + if (!vma_lock_retried) { + vma_lock_retried =3D true; + goto lock_vma; + } + lock_mmap: =20 retry: diff --git a/arch/s390/mm/fault.c b/arch/s390/mm/fault.c index 028aeb9c48d6..faae0e91b90b 100644 --- a/arch/s390/mm/fault.c +++ b/arch/s390/mm/fault.c @@ -271,6 +271,7 @@ static void do_exception(struct pt_regs *regs, int acce= ss) unsigned int flags; vm_fault_t fault; bool is_write; + bool vma_lock_retried =3D false; =20 /* * The instruction that caused the program check has @@ -294,6 +295,7 @@ static void do_exception(struct pt_regs *regs, int acce= ss) flags |=3D FAULT_FLAG_WRITE; if (!(flags & FAULT_FLAG_USER)) goto lock_mmap; +lock_vma: vma =3D lock_vma_under_rcu(mm, address); if (!vma) goto lock_mmap; @@ -318,6 +320,10 @@ static void do_exception(struct pt_regs *regs, int acc= ess) handle_fault_error_nolock(regs, 0); return; } + if (!vma_lock_retried) { + vma_lock_retried =3D true; + goto lock_vma; + } lock_mmap: retry: vma =3D lock_mm_and_find_vma(mm, address, regs); diff --git a/arch/x86/mm/fault.c b/arch/x86/mm/fault.c index 45b99c3b1442..c3ab30d32a15 100644 --- a/arch/x86/mm/fault.c +++ b/arch/x86/mm/fault.c @@ -1222,6 +1222,7 @@ void do_user_addr_fault(struct pt_regs *regs, struct mm_struct *mm; vm_fault_t fault; unsigned int flags =3D FAULT_FLAG_DEFAULT; + bool vma_lock_retried =3D false; =20 tsk =3D current; mm =3D tsk->mm; @@ -1331,6 +1332,7 @@ void do_user_addr_fault(struct pt_regs *regs, if (!(flags & FAULT_FLAG_USER)) goto lock_mmap; =20 +lock_vma: vma =3D lock_vma_under_rcu(mm, address); if (!vma) goto lock_mmap; @@ -1360,6 +1362,12 @@ void do_user_addr_fault(struct pt_regs *regs, ARCH_DEFAULT_PKEY); return; } + + if (!vma_lock_retried) { + vma_lock_retried =3D true; + goto lock_vma; + } + lock_mmap: =20 retry: base-commit: e737cebb8de0d38e8f64584a8bbfbcf9176c7537 --=20 2.43.0