From nobody Fri Sep 25 15:15:43 2026 Received: from mail-qk1-f181.google.com (mail-qk1-f181.google.com [209.85.222.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ED75632C8B for ; Fri, 11 Sep 2026 00:18:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.181 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789085917; cv=none; b=AZTLyLYBiSSnTlM6xa9DttB4aAqPXeuTSFJpqpMoQ90WurIxwvrhLKUCxSNFL/H0Qgti1Rlxc77FbmeIUTmABvlk2vTPWxaOdD4t0tbqa7Ary/cyaGmAfGt5wSBkfUrf/l/yJpUHn5qkD+OLtjwhs3SA+H2aVD+YQOqV7iFITbk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789085917; c=relaxed/simple; bh=Jljp3ztDhFP6O+PAfvk83+8GLPXNCVYLh8IZWY/KNiw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dtHpwTsqkq6n+Ba76xOavZXrJ0hLfqbWbiNvV78j+pA2muecr7fcIcL1kJ0g18L67iw5bNxjGFx1bHBTSab+UaQ1NrHI/7SJwvwon88y10d6SdOPXKIXW/SzYmuMyW7Y2GQvvdqyz7ms9+WOrnmY5ijQpVID++X55QqFfqXy1vM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=s2N61fim; arc=none smtp.client-ip=209.85.222.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="s2N61fim" Received: by mail-qk1-f181.google.com with SMTP id af79cd13be357-9309d4ea213so34380985a.1 for ; Thu, 10 Sep 2026 17:18:35 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1789085915; x=1789690715; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=XlfoGyQgrNSRVIxWxP7qSy7UsF5M3AH+3KUxqb60W44=; b=s2N61fim4h9Hyo1KG80u1cKEqeKti0m5D1U0+FTkwcCoc7rOkZzPPx6Pty7Jjojqkf xmzAC1oIaD/AvxQJQI5isg6B2kBJ7OhwYUiDyzjq/NtY5z8LghxbWKFg2lZUQsF+3dHX noFPRbxcoFRvL5jmnafdtrWxWG/eK/JYj7bSWnLWfixKDAotQyoK0PebTZxqnI9Ll/fv /8wTWgTHzPv7zyDfhFcZWXoxHYsnAwCyR3wYgUsq95SIVJYKb7EzOBJ0wWAV4CszNCi4 rIjJZwMbZToZdUjNRm8Tzi60Ws2RkTB899x5fSbHd4KE+iBiG/loyAsu24eSV/hNwvjB byQA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789085915; x=1789690715; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=XlfoGyQgrNSRVIxWxP7qSy7UsF5M3AH+3KUxqb60W44=; b=DsNs3ugazU4lLVRi907NjzwLKG7M0FUuF98r9fWaLvhU9SZsCPUNEquZzm/SaeSNO7 wqsGxf4bLyo7P0GZqBBLWrIjfp6SCefZmayBmSgG8OBlhvYjCX0q7gI9GGHnBdNYUlf1 PTXe/gCd3RFlbjnEI52s4D/olv3rrtpoXvTrjBHh7bGP8UKkCsCZvTiSZHbWlxNn4/7v MZjm8OCFoZHDEHksqEgMoHTmuCHVeKx9kwedhAmCdi/jouEltHH2asDxOaCe8MLy70LY oY9SdMmT3USBRPFQfiiqXhJrFf5KcB510A2H1lFVysKdRDHQPE2hbWxyHiohZUcAhrNc xpGA== X-Gm-Message-State: AFuF++kQERwp3JtV5z7RqdOvA/ewnBNUdm20925WvpPhSb4eYaU9O9hG nypdt2P+S0i2negjmR4taQkfWm0f4cDJAxOHWHc0wzb8uwzEobzEntHL/N4XgO6n2tc= X-Gm-Gg: AYBFou1TZX2HSqT9u3X7kngsaW/WRpMh5iZ2S5+OKMU+A+rwP3S8FnkZC8W0+eM8UGP nhrEiGRzLOWZP/qnNZPJIYHZWWYl7i/bqnfN9E74Bx9WsQUWTeF0JxV5OLmHAvIVQGpM15nQ7U6 71abl3lfvpRMPT885bQEwFR6wbxbdWMV4VQH78mzEparo359ZnZvvl6ILdyx2k80jFfSxah6b5O ryOGm7xNRoaJ0tNsJvtfwhZg6Kk6pa7H1boivXAUDyJ+sJxZ3ujjg0BuZ0OX7fDK+aI0ZL9g8dp m1WzEAtgB2/MBinswmWdJPCzIuRSOis4uHKcfrrtKvvl5h0Q6BOC0RS0PDK/IMOLssCsBcDeRRB g4CgzrGB20lzAhOaBEuIkuUTfGQU3oGI0ShGnGf84+sk0H76zaO6mDY32BtQ00odKCGMpNd1/QZ 7kfcb/SuyV/gLCmRsnMstkGPO7bFpPZa8iKz5w8aw+FrLJ4rIR1y9nG6Dde93ccbL+Xb+y2+0qS 1RZcbnzGx0fRhDb1STAqV4BAf3suwd8H2IMORKEXFTRhAx6KP2pb8MbHuQibA== X-Received: by 2002:a05:620a:2989:b0:939:7f2f:b17e with SMTP id af79cd13be357-939e9efcaf8mr205827185a.30.1789085914637; Thu, 10 Sep 2026 17:18:34 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-939e80e4c36sm108896485a.40.2026.09.10.17.18.33 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 10 Sep 2026 17:18:34 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, osalvador@suse.de, hannes@cmpxchg.org, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v2 1/4] mm: support promotion-only NUMA hinting scans Date: Thu, 10 Sep 2026 20:18:23 -0400 Message-ID: <20260911001826.2109390-2-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911001826.2109390-1-gourry@gourry.net> References: <20260911001826.2109390-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" folio_can_map_prot_numa() derives the eligible memory from the global balancing mode. Combined mode needs the scanner to distinguish placement scans from promotion-only scans. Add MM_CP_PROT_NUMA_PROMO_ONLY and let change_prot_numa() callers request a promotion-only protection walk. Carry the choice with each walk so the PTE and PMD paths use the same decision. Initially derive the value from the normal balancing mode, preserving the existing behavior for the following fixes. Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory ti= ering system") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- include/linux/mm.h | 4 +++- kernel/sched/fair.c | 7 ++++++- mm/huge_memory.c | 3 ++- mm/internal.h | 5 +++-- mm/mempolicy.c | 22 +++++++++++++--------- mm/mprotect.c | 4 +++- 6 files changed, 30 insertions(+), 15 deletions(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index 969594074fd2..2d1b59a27629 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -3408,6 +3408,8 @@ int get_cmdline(struct task_struct *task, char *buffe= r, int buflen); #define MM_CP_UFFD_RWP_RESOLVE (1UL << 5) /* resolve rwp */ #define MM_CP_UFFD_RWP_ALL (MM_CP_UFFD_RWP | \ MM_CP_UFFD_RWP_RESOLVE) +/* Whether a MM_CP_PROT_NUMA change is for promotion only */ +#define MM_CP_PROT_NUMA_PROMO_ONLY (1UL << 6) =20 bool can_change_pte_writable(struct vm_area_struct *vma, unsigned long add= r, pte_t pte); @@ -4736,7 +4738,7 @@ void vma_set_file(struct vm_area_struct *vma, struct = file *file); =20 #ifdef CONFIG_NUMA_BALANCING unsigned long change_prot_numa(struct vm_area_struct *vma, - unsigned long start, unsigned long end); + unsigned long start, unsigned long end, bool promo_only); #endif =20 struct vm_area_struct *find_extend_vma_locked(struct mm_struct *, diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 8dff37059faf..81359b414947 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4129,8 +4129,10 @@ static void task_numa_work(struct callback_head *wor= k) unsigned long nr_pte_updates =3D 0; long pages, virtpages; struct vma_iterator vmi; + unsigned int numab_mode =3D READ_ONCE(sysctl_numa_balancing_mode); bool vma_pids_skipped; bool vma_pids_forced =3D false; + bool promo_only; =20 WARN_ON_ONCE(p !=3D container_of(work, struct task_struct, numa_work)); =20 @@ -4304,11 +4306,14 @@ static void task_numa_work(struct callback_head *wo= rk) continue; } =20 + promo_only =3D !(numab_mode & NUMA_BALANCING_NORMAL); + do { start =3D max(start, vma->vm_start); end =3D ALIGN(start + (pages << PAGE_SHIFT), HPAGE_SIZE); end =3D min(end, vma->vm_end); - nr_pte_updates =3D change_prot_numa(vma, start, end); + nr_pte_updates =3D change_prot_numa(vma, start, end, + promo_only); =20 /* * Try to scan sysctl_numa_balancing_size worth of diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 30b7c63b0e35..2f9ada1bbcfc 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2784,7 +2784,8 @@ int change_huge_pmd(struct mmu_gather *tlb, struct vm= _area_struct *vma, goto unlock; =20 if (!folio_can_map_prot_numa(pmd_folio(*pmd), vma, - vma_is_single_threaded_private(vma))) + vma_is_single_threaded_private(vma), + cp_flags & MM_CP_PROT_NUMA_PROMO_ONLY)) goto unlock; } /* diff --git a/mm/internal.h b/mm/internal.h index 0dca33db068f..f3e093168f84 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1237,11 +1237,12 @@ static inline bool vma_is_single_threaded_private(s= truct vm_area_struct *vma) =20 #ifdef CONFIG_NUMA_BALANCING bool folio_can_map_prot_numa(struct folio *folio, struct vm_area_struct *v= ma, - bool is_private_single_threaded); + bool is_private_single_threaded, bool promo_only); =20 #else static inline bool folio_can_map_prot_numa(struct folio *folio, - struct vm_area_struct *vma, bool is_private_single_threaded) + struct vm_area_struct *vma, bool is_private_single_threaded, + bool promo_only) { return false; } diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 2ad0a5f18280..a082ccfa09ec 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -846,6 +846,7 @@ static int queue_folios_hugetlb(pte_t *pte, unsigned lo= ng hmask, * @folio: The folio whose mapping considered for being made NUMA hintable * @vma: The VMA that the folio belongs to. * @is_private_single_threaded: Is this a single-threaded private VMA or n= ot + * @promo_only: Whether this scan should only collect promotion candidates * * This function checks to see if the folio actually indicates that * we need to make the mapping one which causes a NUMA hinting fault, @@ -855,7 +856,7 @@ static int queue_folios_hugetlb(pte_t *pte, unsigned lo= ng hmask, * Return: True if the mapping of the folio needs to be changed, false oth= erwise. */ bool folio_can_map_prot_numa(struct folio *folio, struct vm_area_struct *v= ma, - bool is_private_single_threaded) + bool is_private_single_threaded, bool promo_only) { int nid; =20 @@ -886,12 +887,8 @@ bool folio_can_map_prot_numa(struct folio *folio, stru= ct vm_area_struct *vma, if (is_private_single_threaded && (nid =3D=3D numa_node_id())) return false; =20 - /* - * Skip scanning top tier node if normal numa - * balancing is disabled - */ - if (!(sysctl_numa_balancing_mode & NUMA_BALANCING_NORMAL) && - node_is_toptier(nid)) + /* Promotion-only scans do not collect NUMA placement samples */ + if (promo_only && node_is_toptier(nid)) return false; =20 if (folio_use_access_time(folio)) @@ -905,19 +902,26 @@ bool folio_can_map_prot_numa(struct folio *folio, str= uct vm_area_struct *vma, * These are later cleared by a NUMA hinting fault. Depending on these * faults, pages may be migrated for better NUMA placement. * + * With @promo_only only folios eligible for promotion are made + * hint-faultable, without also sampling for task placement. + * * This is assuming that NUMA faults are handled using PROT_NONE. If * an architecture makes a different choice, it will need further * changes to the core. */ unsigned long change_prot_numa(struct vm_area_struct *vma, - unsigned long addr, unsigned long end) + unsigned long addr, unsigned long end, bool promo_only) { + unsigned long cp_flags =3D MM_CP_PROT_NUMA; struct mmu_gather tlb; long nr_updated; =20 + if (promo_only) + cp_flags |=3D MM_CP_PROT_NUMA_PROMO_ONLY; + tlb_gather_mmu(&tlb, vma->vm_mm); =20 - nr_updated =3D change_protection(&tlb, vma, addr, end, MM_CP_PROT_NUMA); + nr_updated =3D change_protection(&tlb, vma, addr, end, cp_flags); if (nr_updated > 0) { count_vm_numa_events(NUMA_PTE_UPDATES, nr_updated); count_memcg_events_mm(vma->vm_mm, NUMA_PTE_UPDATES, nr_updated); diff --git a/mm/mprotect.c b/mm/mprotect.c index 2888ee638d87..4e66e4c7665b 100644 --- a/mm/mprotect.c +++ b/mm/mprotect.c @@ -337,6 +337,7 @@ static long change_pte_range(struct mmu_gather *tlb, long pages =3D 0; bool is_private_single_threaded; bool prot_numa =3D cp_flags & MM_CP_PROT_NUMA; + bool numa_promo_only =3D cp_flags & MM_CP_PROT_NUMA_PROMO_ONLY; bool uffd_rwp =3D cp_flags & MM_CP_UFFD_RWP; bool uffd_wp =3D cp_flags & MM_CP_UFFD_WP; int nr_ptes; @@ -384,7 +385,8 @@ static long change_pte_range(struct mmu_gather *tlb, */ if (prot_numa && !folio_can_map_prot_numa(folio, vma, - is_private_single_threaded)) { + is_private_single_threaded, + numa_promo_only)) { =20 /* determine batch to skip */ nr_ptes =3D mprotect_folio_pte_batch(folio, --=20 2.55.0 From nobody Fri Sep 25 15:15:43 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A7D901A6809 for ; Fri, 11 Sep 2026 00:18:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789085919; cv=none; b=Pl7lWcEBaCrEX6GIQGWO9cVJ9CwuEfPSlWWyR0UNu1M+J0pbg8hPb7NScOuNsDilu3GFzFZZFXCHgFDRgQaCkgvOU50Uc7l8vM5yDeX64IfObxwk+dSDYay+XmDyH28FTgq6E9zZznica/ODK5yFLhrm6kUQLhnciqhVDCLTI7c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789085919; c=relaxed/simple; bh=tJeL+TtL8fcQ3ySRpIEsJeq0/SmJ3J8dhXSj75dUeBQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Z6Eq6xmDd6Rmp00rAYA8gjjA5LpyM2dDZKxYxs6Gw5qLHwrai1WUj2b2R4QDszhrx8FGEuI8vqMZcqeoLsKWgZJowKAOraV1r17gj+TB5pVw13QBKz9SQQ9chdJRAeXU/BIUe0WdIT23qQxU7FjxocBzvfBaFyEm2kGGNWGiE/o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=XXe2jlVm; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="XXe2jlVm" Received: by mail-qk2-f12.google.com with SMTP id af79cd13be357-93910cadeb4so40201085a.1 for ; Thu, 10 Sep 2026 17:18:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1789085916; x=1789690716; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=LhZIogGFAKBWcn9TqaxnEfdXlwBjbgcMrXypWpyrcPc=; b=XXe2jlVmMcsR0HgFbWr2jUgdv0ECoWc0g6YBEu5SzUZPWVdW2gq3GbroxvD7PYi1g9 3V0Y7L34PO11TNXPE2Pz8ov/dw72LudE2YD3x3ln3I7uN1chHCjzPsZAEHjvT+evKXh2 9EhpcwOhIVQ7kn5xQKTDBe0ESua9ZugAJ07ojvoYfOhEyi+FdZ1A4eWDWFuCHbMcG1e2 uLT0roITZsOoaLkqneQlGVRiFCXO98+5exGSR1bfbE2ikHXEimjmw2VRo+oZ3K8Cr2dG jkMN8o+DLaTvo45gNj84oN1/z5MonsT0Pv5RWQOEJEbd0OE5ns413zH6bZiJt3jO+O6W DNUA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789085916; x=1789690716; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=LhZIogGFAKBWcn9TqaxnEfdXlwBjbgcMrXypWpyrcPc=; b=SGoTikQuTfTEuQto2dpmxNuA8VGvDEXZUtA+ilqSbNG9n1c8shj55Cn0RVt/QQAyKS jPPDV86X1lT0XnkYMzNKpJpQNFASMnJTBO0Cu5p7Qz7GxioUuOFlpBkEKIduA/Gz830Y bjm5628nEAJEK3nMfKYSGlDWyu4tRcrdwrRtuNUgoK/JkgW7f62ih3BwSZ+ByVQ9VCce w3njle7zKVEzG3l8xeTa9oETWMi4ePIhfj60yZTsRzIS+HmsnmBxBolczbvtReJGpOqU +R60DiePo2j4EnkZ3o/T65Y5hpKhti6guI9M2cWAt2OeLGXzRRPy2dUoTXcEuGZAhJOJ tPUg== X-Gm-Message-State: AFuF++nKuR2iHyeKl4qaJX4SkmXpr+c7tNbprsvM/JzbidodNx3T1bhk 5XAvcSx8Pr7oKwgqeDekckIswqPxZnQUr3dfviv1kIrenDk47tDHfxe0AFp6Vcw23RM= X-Gm-Gg: AYBFou0XXRM3bZEeqV6heQrb3AXTETzaMxmI6kqjQe+eomVziLo8onpxP3Tx8PyNWRB 5odZMzF6WqSWQ/meHH7eCRBUE0UNR3lvs0k0bZ34VR3a+R0hIf4F5iO/qyETvUQKK/N40iHbsGk 8uObnm0di51BkQXZ8PZ9CIV6b/I1cIOIpZynTdwLYeAczaycA9Jc/J9NNgjKFvaChbhL15zCM6a CSP66Dm6IGi7VNlmj3r0qNxIu80TJhcHFPPQDgim74dCMXzDD5iViYCuTYOe2x+o1YSJr45oV9N vURtItv/5slY+znlrrqHrpYqAf6r0bOOb8eVJkMfMlES2Y/ISc5+81+D12k6dkaLE2fuMUc8wr+ voTxI6hu2aWqhV3ecd6gFvoSV/R5ov/zRqmbwfbA+a3Shngp8KojLvdKMwvh8CWRYW8k/J0p/N0 pAeMfZs8llZOZwDTlazzIwtOe7+sr19fHbRlzUoLvDtajV+DtJ0kivQe00vBZOo2/gM3CFKYmnd j9f3fidiIkrXwtMVDnKo8ca8cnX5yjNzeclTMD8RSmnTUgdlelTdGp8+C6jmnd2liivCVI= X-Received: by 2002:a05:620a:4109:b0:938:fd60:4f9 with SMTP id af79cd13be357-939ea16a07dmr234110685a.27.1789085916408; Thu, 10 Sep 2026 17:18:36 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-939e80e4c36sm108896485a.40.2026.09.10.17.18.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 10 Sep 2026 17:18:36 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, osalvador@suse.de, hannes@cmpxchg.org, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v2 2/4] mm: allow shared folios to be promoted to a fast tier Date: Thu, 10 Sep 2026 20:18:24 -0400 Message-ID: <20260911001826.2109390-3-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911001826.2109390-1-gourry@gourry.net> References: <20260911001826.2109390-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" NUMA balancing rejects shared copy-on-write folios and executable file folios mapped by multiple processes to avoid placement bouncing. These checks also block promotion from slow memory. Allow such folios to participate when moving from a slow tier to a fast tier. Keep the existing restrictions for ordinary placement. Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory ti= ering system") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- mm/mempolicy.c | 8 ++++++-- mm/migrate.c | 6 ++++-- 2 files changed, 10 insertions(+), 4 deletions(-) diff --git a/mm/mempolicy.c b/mm/mempolicy.c index a082ccfa09ec..19b599bc2dd1 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -863,8 +863,12 @@ bool folio_can_map_prot_numa(struct folio *folio, stru= ct vm_area_struct *vma, if (!folio || folio_is_zone_device(folio) || folio_test_ksm(folio)) return false; =20 - /* Also skip shared copy-on-write folios */ - if (vma_is_cow_mapping(vma) && folio_maybe_mapped_shared(folio)) + /* + * Shared copy-on-write folios are poor NUMA placement candidates, but + * a hot folio on a slow tier still needs a hint fault for promotion. + */ + if (vma_is_cow_mapping(vma) && folio_maybe_mapped_shared(folio) && + !folio_use_access_time(folio)) return false; =20 /* Folios are pinned and can't be migrated */ diff --git a/mm/migrate.c b/mm/migrate.c index a369d0c95c38..afd9c97d2389 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -2697,12 +2697,14 @@ int migrate_misplaced_folio_prepare(struct folio *f= olio, /* * Do not migrate file folios that are mapped in multiple * processes with execute permissions as they are probably - * shared libraries. + * shared libraries, unless this is a promotion from a slow tier. * * See folio_maybe_mapped_shared() on possible imprecision * when we cannot easily detect if a folio is shared. */ - if ((vma->vm_flags & VM_EXEC) && folio_maybe_mapped_shared(folio)) + if ((vma->vm_flags & VM_EXEC) && + folio_maybe_mapped_shared(folio) && + (!folio_use_access_time(folio) || !node_is_toptier(node))) return -EACCES; =20 /* --=20 2.55.0 From nobody Fri Sep 25 15:15:43 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 76160275AEB for ; Fri, 11 Sep 2026 00:18:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789085921; cv=none; b=uevijEwaPlHihAKzvJNXKdFqKlHwBUo/4HKk2r3EKlRsvxPtFpX1eRImRdKwaIRE6zvkkgIs2u36daQbj/tR7aaJDYQvLX24b7nRXdQ0fRqX8Q8r1BJiVA4tAIf3B8N4hoEeq5eJ1Ybjt00gsvggileyN/kMlQhTMA39Ex6si8c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789085921; c=relaxed/simple; bh=MbCDYm5ROgW+FhjLtzZnAOPPKxmb4d+M7CPyKUkZdHw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=o+K9bBt7d9R2pvzOereyFIsnosLa/gtfAXOcRhSWIZKJE2HfhPR2fPuzLkyq2lzWZ7hCIQBTF7bd/OZkJkMuBHpBbnbzRJ+PIpwYCqh0fQXfFSU+vXuGz6daH8eXgog8zp8bWlRCHHgxE0OW9CZpPCquOtWrHPlw+THPtAeIw0k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=eKzHeCjB; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="eKzHeCjB" Received: by mail-qk2-f12.google.com with SMTP id af79cd13be357-93910ca5aa6so43986185a.0 for ; Thu, 10 Sep 2026 17:18:39 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1789085918; x=1789690718; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PdeGUk78rOtPCIG/0kxB5vZ5w0EIZpg3hPrMBzGs7bI=; b=eKzHeCjBxsA1oJHX/yRG79/lLg6Sjgoog4OZohyMsVneXKA7/sF+q5fztO42LviodM FHQw4taURmCmDSQPL/6kZSjMO1QQ8dz8dRz67d9kFcuGHSNseAJCn1qh+0w0ACLtGKQV iXy8UIXsL8PPnF5AVn9BdGbIbAZ6+GtWeFf7tTd/X/dafkw/xMqofNxm9nYrtjqc5fYR W1YC42wT4fBLz/JA2T1lpj5QttccY/8ymVrUGKhWSc/wdKT4Z2w/L1dPka0JO2hBcqeA 9QlyVM9il7EwEN2iKnVf2AQW2QJeNwmS+/S40/VkcbFjgObAnjuWhZRBpPm4C/ynTqHQ UJIw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789085918; x=1789690718; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=PdeGUk78rOtPCIG/0kxB5vZ5w0EIZpg3hPrMBzGs7bI=; b=KkD62z2Vtr6lwPYtC+UOgCAnsjV8C/zrUjHed1zSb7JYsY1OjjduVBMt3MIp9A01GP wx00UwDOV9BgtEiYSMpOlxcJVtr/HvNUkx6HzF3z8eLA1Ki5zxeF4C1/lVdRGi1HBlkY W2KfqA121EZxVSwATZ2sQVhsmbFPOS2DOSQUfVgLx1WfwwKjJTksVLwCOdrku1Zx6FcP cfyIlj2ZkX3R98rgI4wT6M+TbzbgHN8FhLTDWgZTSMyPNuE8kyCOcDbkLx0nCV5P3U5r FDfeM7wQKg3DbUpa3kBlc1JNJCmmM+YkGNT0bpNMBRDy8WDSiwhe5XtG35Cddl+YY6Ci NjkA== X-Gm-Message-State: AFuF++m7Snr6fTeexApVQ3bQfD+g6+h8q2PaitocCGvtNUF2aepko50i 9B0tDZgJBv6Vx/E5nk8Byt48NUae1XMmBxMOqazEIpEvThlYwshWtsHfBf8y2S/0ws0= X-Gm-Gg: AYBFou2MLT37iBSWrsXQ2PJpjFSCcFbJwuZ8olRVaLVVcWJMa5/kXXCHczLk9hV/UHl jdDYWgCbhQX9lg5VZEsjCwxdr9Y3CfelPC+C7+zt4709Y6Dg1NNWVegmItfQCMC7RBT4LwLI0SO 8HCWgzXeBUNOXQfBbZha7gG9+Hvs9cFUl2p7ZWjFfSXNmRsXJThJkfm3+6ngaTz5rl3jRmeR0Bu Dg1MfgU+tJ4tUuXApEEfaUIWO3KnH1rQ6YMLieaRuz4CQ4MRiMxP7lrRKiEoqbcY0plz3pdQDqI r0nB47hlMcIb3lVKOcdW7vwE7G4QPDmiH3CfdHd2sM7X9tz3JQbLQLlTapIGMcGiJSVALwvjmFO Uiv5YVgIsEeXCJP9EqeDi6Ok2b8jdfXxFOR4BJoHGkJWmJgrt4/aeBSv5xQmUlBiiTbOdYDej1u eu/THPoS5qoTvGZF1Hl4c17ZAbFZZC0X3tCRAiOv2qLdF4Fz9e0eVIe88iXhpd++ic5v7fsARSK Wj8xsloksqrqptWCrshNW2sO0lbDizbIZ/Od2dKRNVCUBzd+X9xyHCz7hyS X-Received: by 2002:a05:620a:472b:b0:939:d440:e530 with SMTP id af79cd13be357-939ea0e3717mr224092885a.9.1789085918183; Thu, 10 Sep 2026 17:18:38 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-939e80e4c36sm108896485a.40.2026.09.10.17.18.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 10 Sep 2026 17:18:37 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, osalvador@suse.de, hannes@cmpxchg.org, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v2 3/4] sched/numa: scan read-only file mappings in tiering mode Date: Thu, 10 Sep 2026 20:18:25 -0400 Message-ID: <20260911001826.2109390-4-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911001826.2109390-1-gourry@gourry.net> References: <20260911001826.2109390-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" Commit 4591ce4f2d22 ("sched/numa: Do not trap hinting faults for shared libraries") excludes read-only file mappings from NUMA hinting to prevent placement bouncing. This also hides hot file folios on slow memory from the tiering code. Scan these mappings when memory tiering is enabled, but make their scans promotion-only. Ordinary NUMA placement retains the existing restriction. On a host with 768 GB of DRAM and 256 GB of CXL memory running a roughly 430 GB database service, 169 MB of its 185 MB main binary accumulated on CXL before this change. Afterwards its tier residency tracked runtime load. Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory ti= ering system") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- kernel/sched/fair.c | 23 +++++++++++++++++------ 1 file changed, 17 insertions(+), 6 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 81359b414947..e636e8de53f1 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4061,6 +4061,16 @@ void task_numa_fault(int last_cpupid, int mem_node, = int pages, int flags) p->numa_faults_locality[local] +=3D pages; } =20 +/* + * Read-only file-backed mappings are expected to be cache replicated betw= een + * accessor nodes, so they are not worth sampling for placement. They can + * still strand on the slow tier like anything else. + */ +static bool vma_is_ro_file(struct vm_area_struct *vma) +{ + return vma->vm_file && (vma->vm_flags & (VM_READ | VM_WRITE)) =3D=3D VM_R= EAD; +} + static void reset_ptenuma_scan(struct task_struct *p) { /* @@ -4220,13 +4230,13 @@ static void task_numa_work(struct callback_head *wo= rk) } =20 /* - * Shared library pages mapped by multiple processes are not - * migrated as it is expected they are cache replicated. Avoid - * hinting faults in read-only file-backed mappings or the vDSO - * as migrating the pages will be of marginal benefit. + * Read-only file-backed folios are poor NUMA placement + * candidates, but slow-tier folios still need to be scanned for + * promotion. */ if (!vma->vm_mm || - (vma->vm_file && (vma->vm_flags & (VM_READ|VM_WRITE)) =3D=3D (VM_REA= D))) { + (vma_is_ro_file(vma) && + !(numab_mode & NUMA_BALANCING_MEMORY_TIERING))) { trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_SHARED_RO); continue; } @@ -4306,7 +4316,8 @@ static void task_numa_work(struct callback_head *work) continue; } =20 - promo_only =3D !(numab_mode & NUMA_BALANCING_NORMAL); + promo_only =3D !(numab_mode & NUMA_BALANCING_NORMAL) || + vma_is_ro_file(vma); =20 do { start =3D max(start, vma->vm_start); --=20 2.55.0 From nobody Fri Sep 25 15:15:43 2026 Received: from mail-qk1-f182.google.com (mail-qk1-f182.google.com [209.85.222.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C6BA02836F for ; Fri, 11 Sep 2026 00:18:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.182 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789085924; cv=none; b=WtBj/OkOZnzrbxOTcfbpwYHV/KC3XyZ1OnNbZtd5IyWbDFYFZ6gCsz3cY2eCSWj5YYLbPA1BPoyPQLCrFf7FGHjoxjWsMub3lt1lPXdPhJzmH5x9TNEn/jJE9AZT7H8nfoaesiLLFECMBmLdGGkVeTdqOb5xpBaDNV3zUeAnFIM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789085924; c=relaxed/simple; bh=Egofb6FnPwJOzI2OGYLWyIi9uKbwLgCrzOqXLr2nxM0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=cTsOEXPFVr7PQi9I9/0PLSb1ACEMVsFDv865AGtjXRXhxDEv6cLFo03aYJgfhELkJxrbVCQUhlz48awxmSdNLn5MqXKBOIux5GrnVtP0tpud0Bu1Y+wMBx6OOq0uA7y9U0r16nGGqGDKjBP7Z1aq1D40jAI2tqz5nLmJd4k68pY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=WhWxV6o4; arc=none smtp.client-ip=209.85.222.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="WhWxV6o4" Received: by mail-qk1-f182.google.com with SMTP id af79cd13be357-939949d9c33so31023085a.0 for ; Thu, 10 Sep 2026 17:18:41 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1789085920; x=1789690720; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Sdb+na/h6ecDzLaWx78ptxSH3vM79ZXkyzQFAMwq7RI=; b=WhWxV6o4tedF4d1afpB6KR3vD1WhgVxoDVIdBhjDHlo9I1XaoAei587YIp82o1QGwC Z6LasDuv5mubK4t2ZnAuEbaqPvYWEVxwe+B38U1WbaW1CC+r2TldG4nwqfhzgL2LShuu +X65qViLESm3vimf5xzoPDKcga3QqlE4m96ldI44HJLA5+wCwvTMX/B0gTuTpetqtir+ ztpAt58mAf0pK3X349znXMN4Gkc2ybgHv6fmD13dMe+f1flblJHJOPksxQc0P8mftoNJ JhQDAOgL5ZFnxf1Y3SLPPqaZRzqzY19mHyHGZ/BMZjAB7pCsm6sZ6Lvo91AJKphZtfB4 O5Sg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789085920; x=1789690720; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Sdb+na/h6ecDzLaWx78ptxSH3vM79ZXkyzQFAMwq7RI=; b=lo7fjSxcPL9jyEDhSlY++JSAk7dbiqkrYbrjWh4xmqhp6XainnswlUWvMnAQoKu53J +wAYjKfRLBwetBqMQ0TlR/3xIKfD5itN9uBzPIzX5dZsqf9560C0dDxn7qP6xH9UmGSG XTrddezh7IgIswM8nuoCIt+pVuN63gpjMmBVG9BDuzKMfdaymYgTLHuCBnVBSJfhuXMA OyNiVhfq3rRxpRFJ1yaj4caeD33Xa8dakBaCL4pSfL2HwXXWaAnhyZ9fCNo5gbVhw1k6 FfAvu/7LCYYJ3yXCBiPJR33FaWzL1a2X+7UePXrCWMEkMVfz9f271aDx06kn8NofD5t6 JdVg== X-Gm-Message-State: AFuF++k4D8NoiQQwaGwCQJW9fxmSyyLN5pMSIpaw+2eMZAbTARlMI9gJ sFXStQ2Wr60wzQwoHWCGkH9XFLHY8bas9D5C1+blIyo05SIShkWyo2aptWJW7t1Op+g= X-Gm-Gg: AYBFou0lvG6Ol6ePXk6dXdCPcqXLdBsAiru69/e1y6OeGqBxTFOdPafn7N0MVHcnxtB ZhSGKayVs7IdbXj8fqwApXetZt/j9cX2LsEYSp+AY6rSiopiXcr6NvMpgL9gL97VexzlljRePpi XVHY0CikDaVTc5SxKEsOG/egCDoCqiDWGqdgjuWNNK5cMXjqE6WpYdc/71kgF7dLKGM2lSANa8I XOC/zOxSHiSX/0ybvDIaAL2/+3e/oU90SluUz9fzfejMf23FOwQkwzmvJzbKklQ+TpDGcvkze5K URAbrttKiIp/pRc4BsC+IMyYK7+0Ek/MjZEa0ze+kG+a2qW5lXdhwtVEdS3EWVivHnKQXdL7dYo CTYxG2jyTZ1PIQULZ+OXIOPZ1IqIA9/8klJ91GVHEAHL2uCBr+gflSyCnt7VakIzoVbxhxBeBpe mAggUCPL2ghYp10B8PcGRYzHZMvFy9idyd7YUsuRmY6XUpw6jgefMVEkOBtuSmu1i/BSTl2lLTC D7/lBkbY0hegYXCaRBhBYBqO5IVZ4R9R7Z7+LMvJly6LmiidjKI6H57zDly X-Received: by 2002:a05:620a:270c:b0:939:a3a0:9a70 with SMTP id af79cd13be357-939ea0a39e2mr191963485a.18.1789085920053; Thu, 10 Sep 2026 17:18:40 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-939e80e4c36sm108896485a.40.2026.09.10.17.18.38 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 10 Sep 2026 17:18:39 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, osalvador@suse.de, hannes@cmpxchg.org, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v2 4/4] sched/numa: do not let VMA PID activity gate promotion Date: Thu, 10 Sep 2026 20:18:26 -0400 Message-ID: <20260911001826.2109390-5-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911001826.2109390-1-gourry@gourry.net> References: <20260911001826.2109390-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" Commit fc137c0ddab2 ("sched/numa: enhance vma scanning logic") skips VMAs without recent PID activity. Since only NUMA hint faults record that activity, the filter can suppress the fault needed to promote hot slow-tier memory. In tiering mode, scan PID-inactive VMAs using promotion-only scans. Reevaluate this choice whenever the scanner visits a VMA and pass it with each protection walk. Record socket-placement scans in prev_placement_scan_seq. Promotion-only scans still update prev_scan_seq, but no longer postpone the starvation fallback for placement scans. On a host with 768 GB of DRAM and 256 GB of CXL memory running two roughly 430 GB database workloads, a large shmem VMA occupied each scan while 2,537 other VMAs covering 84 GB were skipped as inactive. A hot 20 GB hash table remained entirely on CXL before this change and was split evenly between DRAM and CXL afterwards. Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- include/linux/mm_types.h | 7 ++++++ kernel/sched/fair.c | 48 ++++++++++++++++++++++++++-------------- 2 files changed, 39 insertions(+), 16 deletions(-) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 5413bd10fff2..9f042d6ad465 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -803,6 +803,13 @@ struct vma_numab_state { * A VMA is not eligible for scanning if prev_scan_seq =3D=3D numa_scan_s= eq */ int prev_scan_seq; + + /* + * MM scan sequence ID when the VMA was last scanned for placement. + * The starvation horizon in vma_is_accessed() counts against this, so + * promotion-only scans cannot postpone placement indefinitely. + */ + int prev_placement_scan_seq; }; =20 #ifdef __HAVE_PFNMAP_TRACKING diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index e636e8de53f1..6d1da13a2ef5 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4085,6 +4085,10 @@ static void reset_ptenuma_scan(struct task_struct *p) p->mm->numa_scan_offset =3D 0; } =20 +/* + * Decide whether this VMA should be sampled for NUMA placement. In addit= ion + * to recent accesses, periodically allow a scan to avoid starvation. + */ static bool vma_is_accessed(struct mm_struct *mm, struct vm_area_struct *v= ma) { unsigned long pids; @@ -4101,22 +4105,13 @@ static bool vma_is_accessed(struct mm_struct *mm, s= truct vm_area_struct *vma) if (test_bit(hash_32(current->pid, ilog2(BITS_PER_LONG)), &pids)) return true; =20 - /* - * Complete a scan that has already started regardless of PID access, or - * some VMAs may never be scanned in multi-threaded applications: - */ - if (mm->numa_scan_offset > vma->vm_start) { - trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_IGNORE_PID); - return true; - } - /* * This vma has not been accessed for a while, and if the number * the threads in the same process is low, which means no other * threads can help scan this vma, force a vma scan. */ if (READ_ONCE(mm->numa_scan_seq) > - (vma->numab_state->prev_scan_seq + get_nr_threads(current))) + (vma->numab_state->prev_placement_scan_seq + get_nr_threads(current))) return true; =20 return false; @@ -4142,7 +4137,7 @@ static void task_numa_work(struct callback_head *work) unsigned int numab_mode =3D READ_ONCE(sysctl_numa_balancing_mode); bool vma_pids_skipped; bool vma_pids_forced =3D false; - bool promo_only; + bool accessed, scan_started, promo_only; =20 WARN_ON_ONCE(p !=3D container_of(work, struct task_struct, numa_work)); =20 @@ -4278,6 +4273,7 @@ static void task_numa_work(struct callback_head *work) * first scan: */ vma->numab_state->prev_scan_seq =3D mm->numa_scan_seq - 1; + vma->numab_state->prev_placement_scan_seq =3D mm->numa_scan_seq - 1; } =20 /* @@ -4307,17 +4303,30 @@ static void task_numa_work(struct callback_head *wo= rk) } =20 /* - * Do not scan the VMA if task has not accessed it, unless no other - * VMA candidate exists. + * The PID filter must not gate promotion. Scan PID-inactive + * VMAs in tiering mode using promotion-only scans. */ - if (!vma_pids_forced && !vma_is_accessed(mm, vma)) { + accessed =3D vma_is_accessed(mm, vma); + scan_started =3D mm->numa_scan_offset > vma->vm_start; + + if (!vma_pids_forced && !accessed && !scan_started && + !(numab_mode & NUMA_BALANCING_MEMORY_TIERING)) { vma_pids_skipped =3D true; trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_PID_INACTIVE); continue; } + if (!vma_pids_forced && !accessed && + !(numab_mode & NUMA_BALANCING_MEMORY_TIERING) && + scan_started) + trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_IGNORE_PID); =20 + /* + * In combined mode, only VMAs the task uses need placement + * samples. Without tiering, every VMA reaching here does. + */ promo_only =3D !(numab_mode & NUMA_BALANCING_NORMAL) || - vma_is_ro_file(vma); + ((numab_mode & NUMA_BALANCING_MEMORY_TIERING) && + !accessed) || vma_is_ro_file(vma); =20 do { start =3D max(start, vma->vm_start); @@ -4345,8 +4354,15 @@ static void task_numa_work(struct callback_head *wor= k) cond_resched(); } while (end !=3D vma->vm_end); =20 - /* VMA scan is complete, do not scan until next sequence. */ + /* + * VMA scan is complete, do not scan until next sequence. A + * promotion-only scan reached the end of the VMA but did not + * sample placement, so it does not count towards the starvation + * horizon in vma_is_accessed(). + */ vma->numab_state->prev_scan_seq =3D mm->numa_scan_seq; + if (!promo_only) + vma->numab_state->prev_placement_scan_seq =3D mm->numa_scan_seq; =20 /* * Only force scan within one VMA at a time, to limit the --=20 2.55.0