From nobody Thu Sep 24 15:11:02 2026 Received: from mail-qk2-f42.google.com (mail-qk2-f42.google.com [74.125.230.234]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C1EF548F000 for ; Tue, 22 Sep 2026 18:29:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.234 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101777; cv=none; b=noh2/hmKbAuWPKfyeL9k+/scBusqH5BuhnULI8xA3HbBO7SXoiQMa5crrwQHtkSM4h4erjca2oj5oSNxajuL3n0ihusL8aEbMCCOdUtFEnhp78mGJx9khPrydr82lakpQ0yKVhsZAY/rx3IoxN/78+xLUnpgx8flupEkIsoRuPk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101777; c=relaxed/simple; bh=DUuMJtRLPKdLJIOxwHMx3Bh+ZGPd3Aw74L7rounBegw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=FoKKWrt9uP2K2vKKLIxVKp1mvZZJ5BIVoig2V/YctNS5XR6miHhOyXHeVI6xHPhXsenXC+kWRNdVO4ky0styspF7YT5ekfZCb97wyNSoTu0vNewKbUntUL912pnpvGfmevhODQYjR4cqa+2SrzO5CvPWmjDiSDgBhd+YHVsaJuU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=TRtFFBM/; arc=none smtp.client-ip=74.125.230.234 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="TRtFFBM/" Received: by mail-qk2-f42.google.com with SMTP id af79cd13be357-93bd580489dso20476385a.1 for ; Tue, 22 Sep 2026 11:29:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1790101773; x=1790706573; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=I23ituCzWdUZNH9RRsiqmg7x+KgPhziEZ7cFA0c9iYo=; b=TRtFFBM/Vq/UQ4KkmD0Sl9Z+RtjQzb22ocMyp2NqeBz9I5ORzci9b6CwrCFmVnzQDf +ctDnr0s3vbAQXChnMzgcEaFD3BztNbZEp/8IG+0XCcqNrMOPs6LpThb9Jg6H+ca57Ck O5H78VYZZWkOKx10mbU+xVc0h+a7sMZiCZS2vMrZjwFUNNOAzOy0dAv/2BeNh7NdmoEl 58k7fS2E2xU9sWtLismvXd6gWMcYb1j6E+ieaH4bUb9X3L3rk7ZOvwUDFgvKHXXxUQXW 51EOp0tZ6RUPyjM8mj+Mo7vn5ryiFiTvhx1odveFEFEG2ha1jEbdjXbKe0uaJY5Xcom5 VPaQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790101773; x=1790706573; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=I23ituCzWdUZNH9RRsiqmg7x+KgPhziEZ7cFA0c9iYo=; b=FczBiYbYAVUAIuKm5tkmq4ut0XWl+FtE+HLOaDFufYOtYqv/yrfHLYfO+kjuad9wfg zxFl0pTYetYIHkWH07XvwkpQXMLoH4Ftmd2Z+ecztMi7xS4ZAJdb/ZZ6ePq1kVWKQLs8 DwZp5FOEVrfcxKFa9lsd1OiJdk/LMxdeVAu3LN4aKo1Ijb0ivUNYOjo4MXKj3wJndO+A DnHZW3iYJStvkqQpSpbJ+CbcIKWgDRfxr43uijqyVJcv5A4YDfRHoKLhybtJlxxX9cVd sbUoa3WauPekL93iJaXr4BVSR7kxsAovFVZIw6Nc0tjCzFPqyYbmayLeEskTiD3ztWql qzKg== X-Gm-Message-State: AFuF++lcnz4MOHmvJQO31MgB3gASnp9DHnZF4eUNmWi9P/w2I3bGUIiK 58WFzleHEyCbnANDhPYez8MqsmYS6tKnP+CPAKnTjGx4O/or3EUm/rTOu7Ke/hwBGWk= X-Gm-Gg: AYBFou2WNDjYQN4aZfLhGPljFBmtI2t1HFx/syzb9sTDAI3LErDTs1KES6ZcrDFRC1h Z7zraGBM6cM/7TzO1l8ovYyY1gzFpmDTEv/fEtivPNRRXusACktKUVbgZ7/GtBkhvryzjC0GFl/ VWKgVIbgfXjsaj451d9QyM3DCaEQunti5LC0yI00F+2AaKV5D87oIDWlLzmXqga+rlK4adiJ+Su jydGiTA+LIseWoXkOIvJLpACRGf7ukQVLx2mcFxSCiwfp76TDKkgl0uL23rJQs5HTuJHRb+Y3CW RSEFkBaFvRLKJBEDXSCQcoERxPgAD1TQVQ9LbsZaE66YzVfmijFMUYGOaeQwyMp+nhSezKSVQPV JmLzf3nk35kABeu3hzDEJY/5DYFEBe/jmYYU5zyzRW1QdatLyqCXWpecgViCJ0wRSdTmF89mWUr /adzGHMkyo5scyFbJSt63cOVXZcyFEDKe62+sp6SRH6R7fJtRczJAOhxJh385XVQou3Nl3q/Ejp Iuv0KRTePVI9eEYW35RUg26yh25AQKPMoEVHE/H/De/gMgGJ/gUtbX/kNhp X-Received: by 2002:a05:620a:31a4:b0:93a:2e8:5afa with SMTP id af79cd13be357-93c2514141amr37871485a.30.1790101773423; Tue, 22 Sep 2026 11:29:33 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c24889ac9sm37130085a.27.2026.09.22.11.29.32 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 11:29:33 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v3 1/7] mm: support promotion-only NUMA hinting scans Date: Tue, 22 Sep 2026 14:29:22 -0400 Message-ID: <20260922182928.2199090-2-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260922182928.2199090-1-gourry@gourry.net> References: <20260922182928.2199090-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" folio_can_map_prot_numa() derives folio eligibility from the global balancing mode. The mode (normal, tiering, or combined) describes which balancing mechanisms are enabled (placement vs promotion). The global setting itself cannot describe the intent of an individual protection walk because normal-mode tempers its scanning activity based on a number of heuristics (read-only VMAs, activeness of VMA, etc). In tiering or combined mode, applying these normal-mode optimizations to an entire VMA is incorrect and breaks tiering. - Opting VMAs out of scanning because they became inactive obviously breaks tiering - because the intent is to identify when a VMA becomes active. Applying it in tiering modes causes - Opting Read-only file VMAs out of scanning is just incorrect for tiering modes because its intent is to prevent bouncing between sockets (east-west), while tiering controls tier migration (north-south). The result is hot read-only files can overload lower tier bandwidth. Add MM_CP_PROT_NUMA_PROMO_ONLY and let task_numa_work() select the type of walk. Build the change-protection flags there and pass them unchanged through change_prot_numa() so its PTE/PMD paths use the same decision. Have folio_can_map_prot_numa() derive the single-threaded private state at the point of use instead of adding another precomputed boolean to the protection-walk interface. This keeps the interface focused on scan intent and avoids plumbing VMA-derived state beside cp_flags. The tradeoff for the cleaner interface is an atomic mm_users read for each folio, rather than once per PTE range. Check the promotion-only top-tier exclusion bit first as a mild optimization. The scan-intent interface is required by the following memory-tiering fixes and must accompany them when backported. Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory ti= ering system") Cc: stable@vger.kernel.org Suggested-by: David Hildenbrand Link: https://lore.kernel.org/r/4d2853c9-edf4-4685-b186-7214ed6abd84@kernel= .org Assisted-by: LLM Signed-off-by: Gregory Price (Meta) Acked-by: David Hildenbrand (Arm) --- include/linux/mm.h | 6 ++++-- kernel/sched/fair.c | 7 ++++++- mm/huge_memory.c | 3 +-- mm/internal.h | 4 ++-- mm/mempolicy.c | 27 ++++++++++++--------------- mm/mprotect.c | 7 +------ 6 files changed, 26 insertions(+), 28 deletions(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index e3d29f87567f8..beab621e6f88d 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -3608,6 +3608,8 @@ int get_cmdline(struct task_struct *task, char *buffe= r, int buflen); #define MM_CP_UFFD_RWP_RESOLVE (1UL << 5) /* resolve rwp */ #define MM_CP_UFFD_RWP_ALL (MM_CP_UFFD_RWP | \ MM_CP_UFFD_RWP_RESOLVE) +/* Whether a MM_CP_PROT_NUMA change is for promotion only */ +#define MM_CP_PROT_NUMA_PROMO_ONLY (1UL << 6) =20 bool can_change_pte_writable(struct vm_area_struct *vma, unsigned long add= r, pte_t pte); @@ -4980,8 +4982,8 @@ static inline void vma_set_page_prot(struct vm_area_s= truct *vma) void vma_set_file(struct vm_area_struct *vma, struct file *file); =20 #ifdef CONFIG_NUMA_BALANCING -unsigned long change_prot_numa(struct vm_area_struct *vma, - unsigned long start, unsigned long end); +unsigned long change_prot_numa(struct vm_area_struct *vma, unsigned long s= tart, + unsigned long end, unsigned long cp_flags); #endif =20 struct vm_area_struct *find_extend_vma_locked(struct mm_struct *, diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index ae6c1a606eb5d..9f544e3df9490 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4124,6 +4124,7 @@ static void task_numa_work(struct callback_head *work) struct mm_struct *mm =3D p->mm; u64 runtime =3D p->se.sum_exec_runtime; struct vm_area_struct *vma; + unsigned long cp_flags =3D MM_CP_PROT_NUMA; unsigned long start, end; unsigned long nr_pte_updates =3D 0; long pages, virtpages; @@ -4131,6 +4132,9 @@ static void task_numa_work(struct callback_head *work) bool vma_pids_skipped; bool vma_pids_forced =3D false; =20 + if (!(READ_ONCE(sysctl_numa_balancing_mode) & NUMA_BALANCING_NORMAL)) + cp_flags |=3D MM_CP_PROT_NUMA_PROMO_ONLY; + WARN_ON_ONCE(p !=3D container_of(work, struct task_struct, numa_work)); =20 work->next =3D work; @@ -4307,7 +4311,8 @@ static void task_numa_work(struct callback_head *work) start =3D max(start, vma->vm_start); end =3D ALIGN(start + (pages << PAGE_SHIFT), HPAGE_SIZE); end =3D min(end, vma->vm_end); - nr_pte_updates =3D change_prot_numa(vma, start, end); + nr_pte_updates =3D change_prot_numa(vma, start, end, + cp_flags); =20 /* * Try to scan sysctl_numa_balancing_size worth of diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 493b29cabee59..da9cf87903a55 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2769,8 +2769,7 @@ int change_huge_pmd(struct mmu_gather *tlb, struct vm= _area_struct *vma, if (is_huge_zero_pmd(*pmd)) goto unlock; =20 - if (!folio_can_map_prot_numa(pmd_folio(*pmd), vma, - vma_is_single_threaded_private(vma))) + if (!folio_can_map_prot_numa(pmd_folio(*pmd), vma, cp_flags)) goto unlock; } /* diff --git a/mm/internal.h b/mm/internal.h index 0434dfcfc36f1..cbeaf1fc38cfc 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1248,11 +1248,11 @@ static inline bool vma_is_single_threaded_private(s= truct vm_area_struct *vma) =20 #ifdef CONFIG_NUMA_BALANCING bool folio_can_map_prot_numa(struct folio *folio, struct vm_area_struct *v= ma, - bool is_private_single_threaded); + unsigned long cp_flags); =20 #else static inline bool folio_can_map_prot_numa(struct folio *folio, - struct vm_area_struct *vma, bool is_private_single_threaded) + struct vm_area_struct *vma, unsigned long cp_flags) { return false; } diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 4c0b8ff1a7e66..1427e1b6213b6 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -847,7 +847,7 @@ static int queue_folios_hugetlb(pte_t *pte, unsigned lo= ng hmask, * folio_can_map_prot_numa() - check whether the folio can map prot numa * @folio: The folio whose mapping considered for being made NUMA hintable * @vma: The VMA that the folio belongs to. - * @is_private_single_threaded: Is this a single-threaded private VMA or n= ot + * @cp_flags: Flags describing the protection change * * This function checks to see if the folio actually indicates that * we need to make the mapping one which causes a NUMA hinting fault, @@ -857,7 +857,7 @@ static int queue_folios_hugetlb(pte_t *pte, unsigned lo= ng hmask, * Return: True if the mapping of the folio needs to be changed, false oth= erwise. */ bool folio_can_map_prot_numa(struct folio *folio, struct vm_area_struct *v= ma, - bool is_private_single_threaded) + unsigned long cp_flags) { int nid; =20 @@ -880,22 +880,17 @@ bool folio_can_map_prot_numa(struct folio *folio, str= uct vm_area_struct *vma, if (folio_is_file_lru(folio) && folio_test_dirty(folio)) return false; =20 - /* - * Don't mess with PTEs if folio is already on the node - * a single-threaded process is running on. - */ nid =3D folio_nid(folio); - if (is_private_single_threaded && (nid =3D=3D numa_node_id())) + /* Promotion-only scans do not mark top-tier folios. */ + if ((cp_flags & MM_CP_PROT_NUMA_PROMO_ONLY) && node_is_toptier(nid)) return false; =20 /* - * Skip scanning top tier node if normal numa - * balancing is disabled + * Don't mess with PTEs if folio is already on the node + * a single-threaded process is running on. */ - if (!(sysctl_numa_balancing_mode & NUMA_BALANCING_NORMAL) && - node_is_toptier(nid)) + if (vma_is_single_threaded_private(vma) && nid =3D=3D numa_node_id()) return false; - if (folio_use_access_time(folio)) folio_xchg_access_time(folio, jiffies_to_msecs(jiffies)); =20 @@ -907,19 +902,21 @@ bool folio_can_map_prot_numa(struct folio *folio, str= uct vm_area_struct *vma, * These are later cleared by a NUMA hinting fault. Depending on these * faults, pages may be migrated for better NUMA placement. * + * @cp_flags carries the NUMA scan policy through the protection walk. + * * This is assuming that NUMA faults are handled using PROT_NONE. If * an architecture makes a different choice, it will need further * changes to the core. */ -unsigned long change_prot_numa(struct vm_area_struct *vma, - unsigned long addr, unsigned long end) +unsigned long change_prot_numa(struct vm_area_struct *vma, unsigned long a= ddr, + unsigned long end, unsigned long cp_flags) { struct mmu_gather tlb; long nr_updated; =20 tlb_gather_mmu(&tlb, vma->vm_mm); =20 - nr_updated =3D change_protection(&tlb, vma, addr, end, MM_CP_PROT_NUMA); + nr_updated =3D change_protection(&tlb, vma, addr, end, cp_flags); if (nr_updated > 0) { count_vm_numa_events(NUMA_PTE_UPDATES, nr_updated); count_memcg_events_mm(vma->vm_mm, NUMA_PTE_UPDATES, nr_updated); diff --git a/mm/mprotect.c b/mm/mprotect.c index e59c69cb5a238..7a247ed4cd55f 100644 --- a/mm/mprotect.c +++ b/mm/mprotect.c @@ -335,7 +335,6 @@ static long change_pte_range(struct mmu_gather *tlb, pte_t *pte, oldpte; spinlock_t *ptl; long pages =3D 0; - bool is_private_single_threaded; bool prot_numa =3D cp_flags & MM_CP_PROT_NUMA; bool uffd_rwp =3D cp_flags & MM_CP_UFFD_RWP; bool uffd_wp =3D cp_flags & MM_CP_UFFD_WP; @@ -346,9 +345,6 @@ static long change_pte_range(struct mmu_gather *tlb, if (!pte) return -EAGAIN; =20 - if (prot_numa) - is_private_single_threaded =3D vma_is_single_threaded_private(vma); - flush_tlb_batched_pending(vma->vm_mm); lazy_mmu_mode_enable(); do { @@ -383,8 +379,7 @@ static long change_pte_range(struct mmu_gather *tlb, * must set protnone regardless of NUMA placement. */ if (prot_numa && - !folio_can_map_prot_numa(folio, vma, - is_private_single_threaded)) { + !folio_can_map_prot_numa(folio, vma, cp_flags)) { =20 /* determine batch to skip */ nr_ptes =3D mprotect_folio_pte_batch(folio, --=20 2.53.0-Meta From nobody Thu Sep 24 15:11:02 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6DB80493650 for ; Tue, 22 Sep 2026 18:29:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101778; cv=none; b=rMGD1X9JiOTuEeWW2tS1aNpVbjDW/BoO6iFFbH7eAqHZIi8O35D4sLN7UitLn8du6oRPrrQOM78xtYHVSt6C7PuFormR4hKWNXTov41ljuYnuHZRLhI1qOZ7QKJuBaW258rVycU3KEIUGfA8f2eNLefh0AFWpfg5rqdo4iAghCI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101778; c=relaxed/simple; bh=3VZh/Vv8oefReRSPIMQMONopSVybIFo51Rb6XMBQI3E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=cQSM1UMRCsqrPwd1QVy+Ionul6vuwJF++dUCROVEbeoStHgcR9m9YiKvT/7oSgH7oklXUFwSWn6QcGpHAPgfW3fpHbrkIDj0QOGzbkEejKhZTg5QGD0U/klrmgqMhHh47yKPb9/IyBhdrSIgDD2f04V/d8IBqCb0ExD2ELEHFHw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=PGUlStLP; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="PGUlStLP" Received: by mail-qk2-f13.google.com with SMTP id af79cd13be357-93910ad2273so20944085a.0 for ; Tue, 22 Sep 2026 11:29:36 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1790101775; x=1790706575; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=NRwVfr0lfTUMkgCOhTSEinsHZrcpdOMoOOSVu7rzRt8=; b=PGUlStLPB5+U5nyzh48gC+bmjptn8iNWrTjSpZsu1v2m43jtZNI4LPT+b1E0Mqn5kL 6xhRI++xdTTehL2mu/6Ln2SwDgRIQArqNB5hP7vJnWwm1AwdT34kUI72NC8TdMisEjhE TsCLxxzP7ED/xr4RE6VJ8Ri3ksaEdSiZCPLELt0NMj0vnqeKs/3L7lybMarPtPhSeNay Wp6bk5IY08xoPEeMwQcifY52Ax/i1fme9P/SlsJjhQfae24yNgaRsOGqt9xm/ipVd1lK fQFFIsM7/G/R5zoQDh3fy5Rs4BcaJvI2qSkf+VY9c/a13uyTUJ9MoUzTBY15vBxrhYNF G6PQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790101775; x=1790706575; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=NRwVfr0lfTUMkgCOhTSEinsHZrcpdOMoOOSVu7rzRt8=; b=EFpoYlaEi2X1bp3FuQSLBwdud66ObW7ieCl0bAy8ogU5O3SpwwiZiYsLi8VUdSw80Q EKM0n2BkhPMAuRxi6WheOJujEMWGWGcO3zz9BhZWVDMOwNOSUSZhcCSQCJHAtXO2psJU GD+qR6jCQyy5edOwaHlWa84GKMcCx/5hqcYvdY3bqpl8lEBwo5YZhLyTaXszu6SMj8QE qIaPv+AqRPmwoxDW4DYh4TojTm48DwONW8R2YAq5veMtDfTIXG67w4R1fV6yhS21dhOO 0+ipLviHv9XNdQTndRO7Jnq5ZGPuu3oLXA7NiH8CUZ9nsL8PkXRdPjhFIDYLERK0tytm 3r7Q== X-Gm-Message-State: AFuF++ny0jXRBklAsWpr3NF/2+W2bNVxXpY5/42tP4LdDx14aF0jZ0ek YfOqCgrV7EeqjAHGIPI1xZ9Fmrvt9sFr+vF35AjPjoXXTRHrCCEjVowcklXMJDWrueY= X-Gm-Gg: AYBFou0KJRCH2XMka/NszEuOa5rpJfEIKXLgznJlzPVGl5zwVoI4tlSVj2Ws870j6SG INH5aI40KJQ+mow122Gx2TY+UuKj/0zDo9KVi4WxQVCFJniH/+a8Zh1XQFtzZSXCX3tSI/NbFfR LmnnwNkBrRsmOqBfEps7qNAtRURBcXcWwgXRjqk30+Zci5Yfsg+dHtt02uoZdPZ3QkWgOz6YA7K cYhH5tnZL2fQSoywzlTy/7kZoqdtmP3vgZTb1rgg7JPkBmU2CpIKMVLu1aBlc3+8/SnJyoXtra7 NMKZw5nij5CEMxOwaHfuKBUO0Q/1Cz+CkCVC6WhnjbLssALohzh/DiW1qXYo4lCw3J0fc77ZWsT Wm8GDTWrdLg6s6HBsh5zZQM+lf439VQFDitFXP0n+qQA1Erczp88nKicnNY+EHQLeDyDxemnOUX pVyOEiwYohkZGt0Dx2O88FWAaigMTX3YN22fiNeP++xxrST9ebU5VRitxg/aaMi2KgtadlaHFXf S3XaccmqMtxSoZCA4D8Y2VW7fVUZAZv4wXQsUML0Z6l/+gIoP7Vaqzw0PTtZN+PPJkr/GY= X-Received: by 2002:a05:620a:8006:b0:93b:c209:b6c0 with SMTP id af79cd13be357-93c250ad128mr41734385a.5.1790101775234; Tue, 22 Sep 2026 11:29:35 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c24889ac9sm37130085a.27.2026.09.22.11.29.34 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 11:29:34 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v3 2/7] mm: allow shared folios to be promoted to a fast tier Date: Tue, 22 Sep 2026 14:29:23 -0400 Message-ID: <20260922182928.2199090-3-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260922182928.2199090-1-gourry@gourry.net> References: <20260922182928.2199090-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" NUMA balancing rejects shared copy-on-write folios and executable file folios mapped by multiple processes to avoid east-west migration bouncing. These checks break promotion from slow memory. Allow such folios to participate when the folio being migrated is a low-tier folio and the destination is top-tier (south->north). This allows promotion, but prevents east-west or north->south migrations (north->south is handled by reclaim demotion). Keep the existing restrictions for ordinary placement and for migrations that are not slow-to-top-tier promotions. Rename folio_use_access_time() to folio_in_lowtier() so the source-tier condition is more obvious (timing is the mechanism, not the condition). The helper retains its existing behavior. Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory ti= ering system") Cc: stable@vger.kernel.org Suggested-by: Zi Yan Link: https://lore.kernel.org/r/DLHS4KFPQ86I.1J4LN3352RI71@nvidia.com Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- include/linux/mm.h | 5 +++-- kernel/sched/fair.c | 2 +- mm/memory-tiers.c | 9 ++++++--- mm/memory.c | 2 +- mm/mempolicy.c | 12 +++++++++--- mm/migrate.c | 14 ++++++++++---- 6 files changed, 30 insertions(+), 14 deletions(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index beab621e6f88d..225ba26c9d503 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -2671,7 +2671,7 @@ static inline void vma_set_access_pid_bit(struct vm_a= rea_struct *vma) } } =20 -bool folio_use_access_time(struct folio *folio); +bool folio_in_lowtier(struct folio *folio); #else /* !CONFIG_NUMA_BALANCING */ static inline int folio_xchg_last_cpupid(struct folio *folio, int cpupid) { @@ -2725,7 +2725,8 @@ static inline bool cpupid_match_pid(struct task_struc= t *task, int cpupid) static inline void vma_set_access_pid_bit(struct vm_area_struct *vma) { } -static inline bool folio_use_access_time(struct folio *folio) + +static inline bool folio_in_lowtier(struct folio *folio) { return false; } diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 9f544e3df9490..dc78d24ed8bc0 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -2734,7 +2734,7 @@ bool should_numa_migrate_memory(struct task_struct *p= , struct folio *folio, * The pages in slow memory node should be migrated according * to hot/cold instead of private/shared. */ - if (folio_use_access_time(folio)) { + if (folio_in_lowtier(folio)) { struct pglist_data *pgdat; unsigned long rate_limit; unsigned int latency, th, def_th; diff --git a/mm/memory-tiers.c b/mm/memory-tiers.c index 25e121851b586..082ca74d51ae9 100644 --- a/mm/memory-tiers.c +++ b/mm/memory-tiers.c @@ -53,16 +53,19 @@ static const struct bus_type memory_tier_subsys =3D { =20 #ifdef CONFIG_NUMA_BALANCING /** - * folio_use_access_time - check if a folio reuses cpupid for page access = time + * folio_in_lowtier - check if a folio is in a tiering-managed lower tier * @folio: folio to check * * folio's _last_cpupid field is repurposed by memory tiering. In memory * tiering mode, cpupid of slow memory folio (not toptier memory) is used = to * record page access time. * - * Return: the folio _last_cpupid is used to record page access time + * If memory tiering is disabled, then lowtier has no appreciable meaning, + * so we return false (the folio should not be migrated on this distinctio= n). + * + * Return: true if memory tiering can promote the folio. */ -bool folio_use_access_time(struct folio *folio) +bool folio_in_lowtier(struct folio *folio) { return (sysctl_numa_balancing_mode & NUMA_BALANCING_MEMORY_TIERING) && !node_is_toptier(folio_nid(folio)); diff --git a/mm/memory.c b/mm/memory.c index 79fa57a381ce0..f05c469e32a2a 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -6240,7 +6240,7 @@ int numa_migrate_check(struct folio *folio, struct vm= _fault *vmf, * For memory tiering mode, cpupid of slow memory page is used * to record page access time. So use default value. */ - if (folio_use_access_time(folio)) + if (folio_in_lowtier(folio)) *last_cpupid =3D (-1 & LAST_CPUPID_MASK); else *last_cpupid =3D folio_last_cpupid(folio); diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 1427e1b6213b6..3c3ee28f0142f 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -864,8 +864,13 @@ bool folio_can_map_prot_numa(struct folio *folio, stru= ct vm_area_struct *vma, if (!folio || folio_is_zone_device(folio) || folio_test_ksm(folio)) return false; =20 - /* Also skip shared copy-on-write folios */ - if (vma_is_cow_mapping(vma) && folio_maybe_mapped_shared(folio)) + /* + * Shared copy-on-write folios are poor east-west placement candidates. + * When tiering is enabled, folio_in_lowtier() identifies a promotable + * folio on a low tier, which needs a hint fault for promotion. + */ + if (vma_is_cow_mapping(vma) && folio_maybe_mapped_shared(folio) && + !folio_in_lowtier(folio)) return false; =20 /* Folios are pinned and can't be migrated */ @@ -891,7 +896,8 @@ bool folio_can_map_prot_numa(struct folio *folio, struc= t vm_area_struct *vma, */ if (vma_is_single_threaded_private(vma) && nid =3D=3D numa_node_id()) return false; - if (folio_use_access_time(folio)) + + if (folio_in_lowtier(folio)) folio_xchg_access_time(folio, jiffies_to_msecs(jiffies)); =20 return true; diff --git a/mm/migrate.c b/mm/migrate.c index 7bdcdb57652f8..6a08690cfe219 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -2694,14 +2694,20 @@ int migrate_misplaced_folio_prepare(struct folio *f= olio, =20 if (folio_is_file_lru(folio)) { /* - * Do not migrate file folios that are mapped in multiple - * processes with execute permissions as they are probably - * shared libraries. + * Limit east-east migration of file folios mapped in + * multiple processes with execute permissions as they + * are probably shared libraries (limits bouncing). + * + * If this is a low-tier folio, only migrate if the target + * node is toptier (this allows south->north migration while + * disallowing east-west migration between slow tiers). * * See folio_maybe_mapped_shared() on possible imprecision * when we cannot easily detect if a folio is shared. */ - if ((vma->vm_flags & VM_EXEC) && folio_maybe_mapped_shared(folio)) + if ((vma->vm_flags & VM_EXEC) && + folio_maybe_mapped_shared(folio) && + (!folio_in_lowtier(folio) || !node_is_toptier(node))) return -EACCES; =20 /* --=20 2.53.0-Meta From nobody Thu Sep 24 15:11:02 2026 Received: from mail-qk1-f180.google.com (mail-qk1-f180.google.com [209.85.222.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8AB25490BF7 for ; Tue, 22 Sep 2026 18:29:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.180 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101780; cv=none; b=qT1L6Axuj3P0dt8ft4Ql7w+aY15EYwWjSoXdfL4DIUcowrOdogEhPhDDdwctKjJXrN17C2Nl+m+n5NyVUGTIczakQOOt/KculBFTiuySXxiQmaP25O42S7tX0kar7lOKzLEU7QbCDbS6ZOBLjbguicEWsV0S++FS8iFD0r7Yggk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101780; c=relaxed/simple; bh=GheSuOituCPdrEAFObtrvIwVgE1xHij2yZfuS4AU0v0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=js+nJwpYbWH4ysM3p2CxYa77W32hxREOUcknaRw42yfjUIgN3dXuHkm8ABeZyk6p3owJB4eC10i7DPpMDcEsMGgQsKeboO5HtcZscI6XBD7zNNpO67aXlbUPA6G8jFO6E0Le3DlLPBKKzse8ktWWLOfJt/MtxM751jV7rTcKUgo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=N6HOwn19; arc=none smtp.client-ip=209.85.222.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="N6HOwn19" Received: by mail-qk1-f180.google.com with SMTP id af79cd13be357-93a0fb2f9efso5946285a.1 for ; Tue, 22 Sep 2026 11:29:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1790101777; x=1790706577; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=1HZmb4EJkk3UUTgAceH+rUT2N1g6qzeO8Xgy2w+ERPg=; b=N6HOwn19eqW+60+jCqY99oqqiOLfRC4xlpuboPicOa607w6hq2Y29o+IqoPLTe2rbY hMUnzfAJQR6SKk4lp1DbHQQPHubwD7vTAys3x9E9Fafitp9HOM8O4ktQE1w9SjV5CLFE vaWX+Zmmf0HlPshyr09HgIBO4kI09LCBbFjd/TWKSl4l6MOAsKKJfFB8oNEyoq/eEx/F OE6cMK6mped1jkQ9e4xlk5pzSjWAX6mSsMlkp5H7OYpKuomOAFWJ+d8qCMnlaelTIuls JJS9d0vUqAFhyzzHzOIcBMbDkMqgQH2fIR7lHy7auhny62qpmwEis0o4KCDKAXzuDmee Lylg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790101777; x=1790706577; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=1HZmb4EJkk3UUTgAceH+rUT2N1g6qzeO8Xgy2w+ERPg=; b=LswmnIHf6OM/lJUxOd5E9lKkP8S2OW7iemln6D3/aIxpfhG1zFDNfPN8nFFHORn9Kc X7PB3slCYNaKEQpkpNAAo8jlTc11n52zA4x5a66MyTvgWvTlOquNrBv0JF4Ud/JSFDfh ywt6Kz3yzTCJMXakWCmhTCC5bLUxLhZKb+gKkrs0Jc4mqUqVMLIoGtHRMtojQbagneGb ydLm5MYHFR1vxunPzMb+zb3uMYYRWSykeolsCErtvt41K4DRSswOrqOq96zEQzywMWn6 VepJp6UHtR2KiQEKkQOtDi0yvtnUJA0eZSOkpilBvVq6k4fTf6Da7y8971TwqRvSedSy 9OUA== X-Gm-Message-State: AFuF++mDyJcDjUoRN67k7l4gRXkjNXgX1I8zrTOVPORq5R4FR/eDfrsW Bk7gDptQtjAsicCh1dwP0Ays6LsP7i71ySIiG48YBlInLdtBWvPOlmnwOXjGDXaguNGuvWDw2rq B4ZWIfCE= X-Gm-Gg: AYBFou22+yqYC7W7lN8oOL/UDadw/5tQlvRF3iPr0/bA9GsUdATOefQqGCyKy1sj0Nr hHXzwGebmBkw8wJaotZNgrMcOZjlPAb+Ux0FfUCHOwy8bhndesmn/FYW+U0STHksZQ826jpVk3H O2XPLlerrjRI/Ug2A9cxs1ahu5JOLRvuaimEtf4xCdhqNvrN6jPVWUVetep9SgHR4rJNCWMGCQF lXYvLa5B3HEUDctvpxmL9ZBNW1tEAhlBjC2ujPLHRTTIdxDjrk1Iy+6TSotel1+tTO2bajMs2BD QCju1GRcZahZ8jf1ZXJQxysvjzs/C2ERjofP+SUres9AeBE1ryeRsttwFtrlZiaoMnZWQyq1bik GgRbYihxyCTSrjjRV9ZeDuLUmftVefQkWRgTTyxSa+6+kPa10z1aPaRaFvl2lXHWUvvU1MAIpQ2 l6HyTfa1u5cwWLeDtGH6fUIYDDH+1OYqFI0WrrKJw63cpFbhW6FetFCXF7UnYFwgGbNpLaR9LO0 HHg14uXrmzMlpiJONat2+5sEp5uSJKuUxbYnijRa3uep9js8h5GTrSsN46j X-Received: by 2002:a05:620a:370f:b0:939:a9d4:7023 with SMTP id af79cd13be357-93c17e32171mr566886685a.10.1790101777045; Tue, 22 Sep 2026 11:29:37 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c24889ac9sm37130085a.27.2026.09.22.11.29.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 11:29:36 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v3 3/7] sched/numa: scan read-only file mappings in tiering mode Date: Tue, 22 Sep 2026 14:29:24 -0400 Message-ID: <20260922182928.2199090-4-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260922182928.2199090-1-gourry@gourry.net> References: <20260922182928.2199090-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" Commit 4591ce4f2d22 ("sched/numa: Do not trap hinting faults for shared libraries") excludes file-backed read-only VMAs from NUMA hint faulting to prevent east-west placement bouncing. This filter hides hot file folios on slow memory from promotion. Scan those VMAs when tiering is enabled, but make their scans promotion only to retains the existing restriction. Keep the historical VMA predicate unchanged for backport-ability. Read the balancing mode once per task_numa_work() invocation and build the protection flags for each VMA from that snapshot. This keeps the PTE and PMD paths on the same policy for an entire protection walk (which may be split across multiple scanning periods). On a host with 768 GB of DRAM and 256 GB of CXL memory running two database services using ~430GB each, 169 MB of their shared 185 MB main binary accumulated on CXL before permanently stuck there. With the change, the binary tier residency tracks its runtime hotness. Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory ti= ering system") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- kernel/sched/fair.c | 31 +++++++++++++++++++++---------- 1 file changed, 21 insertions(+), 10 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index dc78d24ed8bc0..8a4687f67d82f 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4119,21 +4119,21 @@ static bool vma_is_accessed(struct mm_struct *mm, s= truct vm_area_struct *vma) */ static void task_numa_work(struct callback_head *work) { + const unsigned int numab_mode =3D READ_ONCE(sysctl_numa_balancing_mode); + const bool tiering =3D numab_mode & NUMA_BALANCING_MEMORY_TIERING; unsigned long migrate, next_scan, now =3D jiffies; struct task_struct *p =3D current; struct mm_struct *mm =3D p->mm; u64 runtime =3D p->se.sum_exec_runtime; struct vm_area_struct *vma; - unsigned long cp_flags =3D MM_CP_PROT_NUMA; + unsigned long cp_flags; unsigned long start, end; unsigned long nr_pte_updates =3D 0; long pages, virtpages; struct vma_iterator vmi; bool vma_pids_skipped; bool vma_pids_forced =3D false; - - if (!(READ_ONCE(sysctl_numa_balancing_mode) & NUMA_BALANCING_NORMAL)) - cp_flags |=3D MM_CP_PROT_NUMA_PROMO_ONLY; + bool placement_scan; =20 WARN_ON_ONCE(p !=3D container_of(work, struct task_struct, numa_work)); =20 @@ -4221,13 +4221,19 @@ static void task_numa_work(struct callback_head *wo= rk) } =20 /* - * Shared library pages mapped by multiple processes are not - * migrated as it is expected they are cache replicated. Avoid - * hinting faults in read-only file-backed mappings or the vDSO - * as migrating the pages will be of marginal benefit. + * Shared library pages mapped by multiple processes are limited + * to south->north migrations as it is expected they are cache + * replicated. The benefit of east-west migration in this case + * is at best marginal and may be harmful due to TLB/cache + * invalidation. + * + * Allow promotion as a cold page incurring many cache-misses + * under cache pressure can drive considerable bandwidth. */ - if (!vma->vm_mm || - (vma->vm_file && (vma->vm_flags & (VM_READ|VM_WRITE)) =3D=3D (VM_REA= D))) { + placement_scan =3D !(vma->vm_file && + (vma->vm_flags & (VM_READ | VM_WRITE)) =3D=3D VM_READ); + + if (!vma->vm_mm || (!placement_scan && !tiering)) { trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_SHARED_RO); continue; } @@ -4307,6 +4313,11 @@ static void task_numa_work(struct callback_head *wor= k) continue; } =20 + placement_scan &=3D numab_mode & NUMA_BALANCING_NORMAL; + cp_flags =3D MM_CP_PROT_NUMA; + if (!placement_scan) + cp_flags |=3D MM_CP_PROT_NUMA_PROMO_ONLY; + do { start =3D max(start, vma->vm_start); end =3D ALIGN(start + (pages << PAGE_SHIFT), HPAGE_SIZE); --=20 2.53.0-Meta From nobody Thu Sep 24 15:11:02 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E3699492E52 for ; Tue, 22 Sep 2026 18:29:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101783; cv=none; b=EsIKAZ8TJ5LR9mXA6vpfeKAhrskBF/bL93c/y2dxKxHDQWduo2HaXZEiGgeffFn2lHNM2mngNVn1lu2nBYPnaH6WBBMsFnBu7uolj2TW4+Lls4e+UK2zwCe5XD2O6hnODT9fqcXY0El9YT4jciI12dmQZ0HESKwm1pfFowu2nDQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101783; c=relaxed/simple; bh=7YChvBHzQ8DGOUO9tClkDuQmXSsjmELLb2ZSb68vG4c=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lg34q+nPei+Rldj0f+USBTSfFix16y/LYVZi7QQn8jLKllyMY1iISBuMeRFaC5Vzao9+9IuWxbkIpngnZ9Q1hTucOofzkfUyhdUQfkQ9zqa+DNRP8Fbon86zZX1kttg2aieaLY3ZXgmpWKW/cD0V/npvVLFBYn8zovWZqfnBRK8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=UTjSOxVp; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="UTjSOxVp" Received: by mail-qk2-f13.google.com with SMTP id af79cd13be357-93910a0cb7cso15774485a.0 for ; Tue, 22 Sep 2026 11:29:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1790101779; x=1790706579; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=nfgiKa+YRyMtK2RiU6jcCWMC566+XGPfa3bnqM4rs00=; b=UTjSOxVpdI8mvzzZDoYYaq9lRgOlymMuD5SWMdZOqm7xqmxK1ogfYNnelvFk/HlwEg mv7FGX+JvL0Q7NyJZ9AJsuQofX6Hs4HKRCrPn3Cg4p86CQ1wZS0pVzF2SAG78Gkw5/RF /Ocg5tuEUhHQp1anP6MqSYX5QTVM45EenfqzRlGV5yeKsIIofaDrTSDlqDFJikbMbNxX ixg2X45zr12ecZqwf8sxXWRakga38xDnZM7orcMhHKGkpgvOSlwh/3PnWMosdHjPNWIE PRw/0Bb1sLSirI7vXmZzjx1j2ddhMRNE4BwtWemjfHFjhof0Xyc+3KNfhc5saKyDcscs ug5A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790101779; x=1790706579; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=nfgiKa+YRyMtK2RiU6jcCWMC566+XGPfa3bnqM4rs00=; b=yD9qyCZsOScXovlYqOfGpYn1FBVFbwoP8TwTJ0w6nDAaVIgahpW43IAV/OGJHX8CZc aMClkk9Hkm6/72e20JeJk1jNjIW+WTN9djBbmfXTGi/UQVDSVhHAHZ5EIeoRhgdOeJIF xp7IcslzrBF5TqBLyZxhtiq8izaRU6ZZs0qLvOIHB6ZF/v+Jr+2VLHFMC7HbbkvLufij JyodmhJ5aofE2Gi+Fnz7uwFN1CkoUXmdpFmTrfu9P6bAqIGlw9lGfXnIUr9sctbhRwyE 1lqehbp5e3QLX/HFk/Bri/l8o0h62DG4jmddi0DS1C/MbVtECk4R5BzA0QETY1jukGNg Sevg== X-Gm-Message-State: AFuF++n4SVTwUe7Dz19TW1D2X0B6Xdxz4vFuv2bOe/63O9p9WwvuI0x7 NMAKm81Wdo6qDZXZyGRJe5vNr3US2sia73w8vOGKNhj4nAp8wwG/SbOduxijYHa6Kxk= X-Gm-Gg: AYBFou0AcCwUH9/JUGlUcIfsxE+2i2FgbrsUqjBGzO4bAZdD0O5VriOVAf1r+KZNz5P skj1sKNpP0n9GSDF2G8cW5uTArbGscY+Beqywd1X5JoxLe76QVn2/30HfD1oJWoHvYd4tqGXImz 3NVc9VWSFWXgTuw7+WNX4Qvgo2BwpZWIoW+cU4pBLwuEpA9UV+6rW1khLZO0B8A6HLX//lN7z6l JhADEvbEsVxDCKhNwPf9YDveOkxtHfg7pY908ls15miWyuL/6KsDHmcIgXYGRD7jKtCzDqTtE0M l2U0NA+FnKkaaIrwRN1Ht0hHaj2K4xbiwyHoFHYBftN+YxVObdJCx1noSmp/nqSWygV9s4688BE y1i5Ov/aDbkmpLypSvd9PjHrN1l5ROY1MNd5miDCGyrV0bdVc1gXniN7DH7raLPfFylBjZPBvR1 V4PJUOLXRLF1ZaoIB1SfQVkKVj5ejmkUnIu/UhPy4XbLPF1px5O5YQPAUcVGnJVCZV7VEQrL17p evqGv0vd2iXhR9PAnzmg8mlyxtWs5dvK0H+XvXZOXluz18RmO6C/u3U0MK6 X-Received: by 2002:a05:620a:27d5:b0:93b:d7a2:83d5 with SMTP id af79cd13be357-93c252aabb2mr34297885a.68.1790101778939; Tue, 22 Sep 2026 11:29:38 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c24889ac9sm37130085a.27.2026.09.22.11.29.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 11:29:38 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v3 4/7] sched/numa: separate VMA placement from scan continuation Date: Tue, 22 Sep 2026 14:29:25 -0400 Message-ID: <20260922182928.2199090-5-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260922182928.2199090-1-gourry@gourry.net> References: <20260922182928.2199090-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" vma_is_accessed() returns true both when a VMA needs placement (east-west) sampling and when a VMA is already mid-scan across multiple scan windows. Callers cannot distinguish placement eligibility from scan progress. Rename the helper to vma_needs_placement_scan() and handle continuation state in task_numa_work(). Cache the scan eligibility decision while a VMA is scanned in chunks so a resumed scan retains the existing policy. This separation allows PID-inactive VMAs to be scanned for promotion without treating scan continuation as evidence of placement eligibility. It is a prerequisite for the following fix and must accompany it when backported. Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- include/linux/mm_types.h | 3 +++ kernel/sched/fair.c | 42 ++++++++++++++++++++++++---------------- 2 files changed, 28 insertions(+), 17 deletions(-) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 77c90b1994131..fd35db969bc94 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -803,6 +803,9 @@ struct vma_numab_state { * A VMA is not eligible for scanning if prev_scan_seq =3D=3D numa_scan_s= eq */ int prev_scan_seq; + + /* Preserve placement-scan eligibility during an in-progress scan. */ + bool placement_scan; }; =20 #ifdef __HAVE_PFNMAP_TRACKING diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 8a4687f67d82f..a2849e72c4e26 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4074,7 +4074,8 @@ static void reset_ptenuma_scan(struct task_struct *p) p->mm->numa_scan_offset =3D 0; } =20 -static bool vma_is_accessed(struct mm_struct *mm, struct vm_area_struct *v= ma) +static bool vma_needs_placement_scan(struct mm_struct *mm, + struct vm_area_struct *vma) { unsigned long pids; /* @@ -4090,15 +4091,6 @@ static bool vma_is_accessed(struct mm_struct *mm, st= ruct vm_area_struct *vma) if (test_bit(hash_32(current->pid, ilog2(BITS_PER_LONG)), &pids)) return true; =20 - /* - * Complete a scan that has already started regardless of PID access, or - * some VMAs may never be scanned in multi-threaded applications: - */ - if (mm->numa_scan_offset > vma->vm_start) { - trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_IGNORE_PID); - return true; - } - /* * This vma has not been accessed for a while, and if the number * the threads in the same process is low, which means no other @@ -4133,7 +4125,8 @@ static void task_numa_work(struct callback_head *work) struct vma_iterator vmi; bool vma_pids_skipped; bool vma_pids_forced =3D false; - bool placement_scan; + bool pid_scan_allowed, placement_due; + bool placement_scan, scan_started; =20 WARN_ON_ONCE(p !=3D container_of(work, struct task_struct, numa_work)); =20 @@ -4304,16 +4297,30 @@ static void task_numa_work(struct callback_head *wo= rk) } =20 /* - * Do not scan the VMA if task has not accessed it, unless no other - * VMA candidate exists. + * Do not scan the VMA if a task has not accessed it, unless no other + * VMA candidate exists. If a scan is already in-progress, finish it, + * but track continuation separately from starting a new one. */ - if (!vma_pids_forced && !vma_is_accessed(mm, vma)) { - vma_pids_skipped =3D true; - trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_PID_INACTIVE); - continue; + placement_due =3D vma_needs_placement_scan(mm, vma); + scan_started =3D mm->numa_scan_offset > vma->vm_start; + pid_scan_allowed =3D vma_pids_forced || placement_due; + + if (!pid_scan_allowed) { + if (scan_started) { + trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_IGNORE_PID); + } else { + vma_pids_skipped =3D true; + trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_PID_INACTIVE); + continue; + } } =20 + /* Keep scan policy stable while processing a VMA in chunks.*/ placement_scan &=3D numab_mode & NUMA_BALANCING_NORMAL; + if (scan_started) + placement_scan &=3D vma->numab_state->placement_scan; + + vma->numab_state->placement_scan =3D placement_scan; cp_flags =3D MM_CP_PROT_NUMA; if (!placement_scan) cp_flags |=3D MM_CP_PROT_NUMA_PROMO_ONLY; @@ -4346,6 +4353,7 @@ static void task_numa_work(struct callback_head *work) =20 /* VMA scan is complete, do not scan until next sequence. */ vma->numab_state->prev_scan_seq =3D mm->numa_scan_seq; + vma->numab_state->placement_scan =3D false; =20 /* * Only force scan within one VMA at a time, to limit the --=20 2.53.0-Meta From nobody Thu Sep 24 15:11:02 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 21552493D2D for ; Tue, 22 Sep 2026 18:29:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101784; cv=none; b=CcCnYODSoS/dFIORnH50I3lvl4H7e7Rsepij4GpnU2WKxobo0WPkVd4f5pJkEigHj/c1HWemomv3gF6yZsrwVgCZzNXUu+5ZCdyHVSotvEjIf2sK2bM1bqrROrpnShg375iO+lVp1D/h1ey7eVlCEW86HdK0dZ9yjE+ABzCnNM4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101784; c=relaxed/simple; bh=k4NugTQESVNugUzaMMuu3OYQl8Um++AA2GKnIZL6+Sk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NqAzwl3irvLTyykXpjEPMayem5b/Hq2rUlnt6wJLRCenHRStLq0mlEZ0ORVHeP/Z59SU6B5BxFX1kh3g94i90Jk6YmjwrI8VM371Wmt2Vu9D6jnM9p6iBzfUptxBiVh8EydnfydlCJkXMytpYQBKDu7S/vi7SLdD0ZV+xlnetzY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=iIAb+tlh; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="iIAb+tlh" Received: by mail-qk2-f13.google.com with SMTP id af79cd13be357-93a0299c787so15808485a.1 for ; Tue, 22 Sep 2026 11:29:41 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1790101781; x=1790706581; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=axWZscC+v0nVIusaiMTv07a6IEIvhtM/ylITkcHKQpY=; b=iIAb+tlh2BOcjhJcOrzMjhd4U8PGcMXr/VmEIkt0r/qX5dYnOCBBBz+L3fbUg8b9ko aEl/h9+uN5s/pgSDqerKUIYiUsGJCMwkAqi8EBWWYd91ImqART76LVXn3CIJekS/eRhS eXR66VE+rQeXqgUobSYlX4psLSDyOPa5x8GeHpyZnYT96liwR+cB6MhFd08gqdWVPsUD tEcDVkqF9kvEqUzt7NoGueZ8ZmeEkQccVnLFGGP/AG8ullsb6DV4q8eleFXScAYg7f0m PvWRw1lGqwXWYQGefToMK30D5anbpG2YeLJ97UTK5jB+EZ9FQ0wAXXyuUVmWNB3wHA6F t1MQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790101781; x=1790706581; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=axWZscC+v0nVIusaiMTv07a6IEIvhtM/ylITkcHKQpY=; b=kCCvX142rVDnWwHsIe5FfZLBLEx0qr7//e267DVs+FAKUZysQilR91yZ4eZFEs1rdw bXvpyCTZJ4Vb5ck5gy21r+xOT7kQ+nqWDmEjs4z+raKalHsieJCHrPFrTncxEy0hhgpa LxUBNYmNBly4e1O8Cb4toyjgrgYuRPGM/nVb3odp8A8MU6psVE0X0dwQJx+IrhBcdsfl bBXJuG+KK708mbGMdc3JHWiklLMvGk3IyzyvZqclmQrFyWU99frym2WIccUwU38dWRCf KMvwW63Hmm0ZIua4Yz1Vz79i3qR8zaQ9XQaMQGsXhlXn9E5gFKPj9zk3wP74jC5XEFOC 31Ag== X-Gm-Message-State: AFuF++kCGwyjg2SKTZQtbDyzU95Qadc1gmihy7ZjG6QtuaEnJEc/xhOG gA/dQjrMsO4RinZhxyOsB8FjL0nD9FeAIYyiMOqBV73CxPzYc60pyqPIaGKyZ1qM5OA= X-Gm-Gg: AYBFou16Ua6QiVBQWz0BkH0zArrrC08W425drmdieIDJG+ao+KziLJixtagCEnfUvdu GXrjIVXOzOD3IkDIhqwNCp5M65sMkBH7iR4VudD9yWzRBP8Wbd0TJb7IWkESVXMhkKhB9XxYL0i zqwVzvJ7mbGQh7Kl0V2SzOvhcFFwdG4kmJYgYgVDweI4bqLDfVe600Lf9w4HFH02ORfkVw5kEBD xXtmfFCKjJvtMW5sViiL/xxQkOrDuscCiMJSd+TUgVSYU4Ve8NTUXDc0eJG/RYqEvsga3gA15r/ RqxELIYpWcQzD+9ndsJX7CDa/xgQDodKQ+2nYWxSUR35nyFCxmUceG/YNcgAe6dYIKtAyzUcjOz M1d7F3nDpu2/BsrlIAvaE+CVaH0k1HQd8Lw4AWBue+1M+dmmlMOye3DYGfi3QvAECIuVQjxHBSN fRmGT9lrWb8uv1PLhzO3rS3J3qOdlTjE/gSj91+x+ulANVRyzNAODEwGnWG5aNLBodZUk6KCJnx mJq7Hf2HDyf44ebNiv1RPE5VbSle8nlBv7Nd/KfdiBusHpy6cLaTVCNoHuXsmhWfBahpno= X-Received: by 2002:a05:620a:1d02:b0:93b:d7a3:45de with SMTP id af79cd13be357-93c252538f6mr36699785a.71.1790101780908; Tue, 22 Sep 2026 11:29:40 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c24889ac9sm37130085a.27.2026.09.22.11.29.39 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 11:29:40 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com, stable@vger.kernel.org Subject: [PATCH v3 5/7] sched/numa: scan PID-inactive VMAs for promotion Date: Tue, 22 Sep 2026 14:29:26 -0400 Message-ID: <20260922182928.2199090-6-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260922182928.2199090-1-gourry@gourry.net> References: <20260922182928.2199090-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" Commit fc137c0ddab2 ("sched/numa: enhance vma scanning logic") skips VMAs without recent PID activity. Since only NUMA hint faults record that activity, the filter can suppress the fault needed to promote hot slow-tier memory. Let memory tiering bypass the PID scan filter. In combined mode, use the placement decision to restrict top-tier sampling to VMAs that need it, while inactive VMAs still receive promotion-only scans. Track the last completed placement scan separately from scans of any kind. Promotion-only scans still update prev_scan_seq, but do not advance the placement-starvation horizon. On a host with 768 GB of DRAM and 256 GB of CXL memory, one large shmem VMA consumed most scanning activity, while 2,537 other VMAs covering 84 GB were skipped as inactive. One stand-out result: a hot 20 GB hash table ended up trapped entirely on CXL and drove CXL bandwidth utilization beyond sustainable levels - resulting in a large regression. With this series, the hot hash table ends up split evenly between DRAM and CXL, tier residency tracked runtime load, and CXL bandwidth utilization drops from 45GB/s (maxed) to 5-10GB/s, while DRAM bandwidth utilization increases from ~200GB/s to 250GB/s+, resulting in major throughput improvements for the database workload. Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic") Cc: stable@vger.kernel.org Assisted-by: OpenAI:gpt-5 Signed-off-by: Gregory Price (Meta) --- include/linux/mm_types.h | 7 +++++++ kernel/sched/fair.c | 26 +++++++++++++++++++++----- 2 files changed, 28 insertions(+), 5 deletions(-) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index fd35db969bc94..dcca3ead9db59 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -804,6 +804,13 @@ struct vma_numab_state { */ int prev_scan_seq; =20 + /* + * MM scan sequence ID when the VMA was last scanned for placement. + * The starvation horizon in vma_needs_placement_scan() counts against + * this, so promotion-only scans cannot postpone placement indefinitely. + */ + int prev_placement_scan_seq; + /* Preserve placement-scan eligibility during an in-progress scan. */ bool placement_scan; }; diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index a2849e72c4e26..412c72084a63d 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4097,7 +4097,7 @@ static bool vma_needs_placement_scan(struct mm_struct= *mm, * threads can help scan this vma, force a vma scan. */ if (READ_ONCE(mm->numa_scan_seq) > - (vma->numab_state->prev_scan_seq + get_nr_threads(current))) + (vma->numab_state->prev_placement_scan_seq + get_nr_threads(current))) return true; =20 return false; @@ -4267,7 +4267,8 @@ static void task_numa_work(struct callback_head *work) * to prevent VMAs being skipped prematurely on the * first scan: */ - vma->numab_state->prev_scan_seq =3D mm->numa_scan_seq - 1; + vma->numab_state->prev_scan_seq =3D mm->numa_scan_seq - 1; + vma->numab_state->prev_placement_scan_seq =3D mm->numa_scan_seq - 1; } =20 /* @@ -4300,10 +4301,13 @@ static void task_numa_work(struct callback_head *wo= rk) * Do not scan the VMA if a task has not accessed it, unless no other * VMA candidate exists. If a scan is already in-progress, finish it, * but track continuation separately from starting a new one. + * + * The PID filter must not gate promotion. Allow PID-inactive VMAs + * to proceed when memory tiering is enabled. */ placement_due =3D vma_needs_placement_scan(mm, vma); scan_started =3D mm->numa_scan_offset > vma->vm_start; - pid_scan_allowed =3D vma_pids_forced || placement_due; + pid_scan_allowed =3D tiering || vma_pids_forced || placement_due; =20 if (!pid_scan_allowed) { if (scan_started) { @@ -4315,10 +4319,16 @@ static void task_numa_work(struct callback_head *wo= rk) } } =20 - /* Keep scan policy stable while processing a VMA in chunks.*/ + /* + * Keep scan policy stable while processing a VMA in chunks. + * A fault in one chunk can make a VMA placement-eligible. Keep a + * promotion-only decision sticky for the rest of a partial scan. + */ placement_scan &=3D numab_mode & NUMA_BALANCING_NORMAL; if (scan_started) placement_scan &=3D vma->numab_state->placement_scan; + else if (tiering) + placement_scan &=3D placement_due; =20 vma->numab_state->placement_scan =3D placement_scan; cp_flags =3D MM_CP_PROT_NUMA; @@ -4351,8 +4361,14 @@ static void task_numa_work(struct callback_head *wor= k) cond_resched(); } while (end !=3D vma->vm_end); =20 - /* VMA scan is complete, do not scan until next sequence. */ + /* + * VMA scan is complete, do not scan until next sequence. + * A promotion-only scan did not cover top-tier folios, so it + * does not count towards the placement-scan starvation check. + */ vma->numab_state->prev_scan_seq =3D mm->numa_scan_seq; + if (placement_scan) + vma->numab_state->prev_placement_scan_seq =3D mm->numa_scan_seq; vma->numab_state->placement_scan =3D false; =20 /* --=20 2.53.0-Meta From nobody Thu Sep 24 15:11:02 2026 Received: from mail-qk2-f43.google.com (mail-qk2-f43.google.com [74.125.230.235]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0D9DF49480F for ; Tue, 22 Sep 2026 18:29:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.235 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101786; cv=none; b=BAp4E/sx5LHZrn9f4CaQElBUy3b7zvR+eSBW9JWjDlyR1GvOhaQ4fWKedkhhuNYAgNb68RswuBXjHiPZ6LAuJZE8G4A5KfHPDtTk7VpcQM61/8/YDH7/+8zJoJo9gQGv918Ann7vcithG6e+qclvwnJF1Ue8Gbg/8ZZS1mq4YbQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101786; c=relaxed/simple; bh=h5CQimn+pnGcNICzqnW9+AOKntw/eMMcYRRlL+m7m9E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ZU73kHJXklIs0WMAe5ioUS87E+btfRrdYJStEGmufiVedXaNChOar5x8OumuE0G1iPlkVu14o/vW250dybRsTfySK8d9IAJD9s5NFxdxpmxFbqkMSXSknywdhSFJ0n1H4032HqgTbipKAPctOvFP53yWI7eRwtHC0WCWWd3L8/o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=SBD7C8Gi; arc=none smtp.client-ip=74.125.230.235 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="SBD7C8Gi" Received: by mail-qk2-f43.google.com with SMTP id af79cd13be357-93910cadeb0so14589185a.3 for ; Tue, 22 Sep 2026 11:29:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1790101783; x=1790706583; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=qVt6tYIl/AHtUN/63hgHvGCs0fPV42UvfQxmNYHhn/o=; b=SBD7C8Gia8h4vFBY3RUXVtwy2dNdNLs+cNdaecaR++yTtXzW+UNFho7nAXf3WNT1NE vVO6wA2VDa7yhRky3t8FVAKOwb1VFyig2M3BQFZ0G5YCOZTxyhcapJB7mjpY6go9fzwz 33wIzxfo5GVo43m7stLuA0Jow3J6rs4Prb7GaUqB+U07Mkim3lYYuxhswGiKVVXv5wU2 FbV0aOQtf1IprAi9rRfroZtGfubF+H6Ppmwtc3MQZb/emfBpsEP4BuThcUehWm1HmIMB J5QXBMb1sIOYAUkUqhYSyVu+WHqj8HoB/zTQf3LRczYy7taRAB8smdsWPHxNNp8CvrZf 1c1w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790101783; x=1790706583; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=qVt6tYIl/AHtUN/63hgHvGCs0fPV42UvfQxmNYHhn/o=; b=UUPQmVXaMaa7xZE3BuugXXGhoQJW22YP0Gzf5SR0aWW0k224fC+xKP7DaUEGFMYNRl SU6o9ouF923cDY82XhKhA4eQ57k1Ku8VqK47Hky86JmBbfK0sTTq7VdDkX4wUIsg/Zzl XEg17AlN0UyI5U2V5PjmZ08R6f6KfnT+apJg2VurB7XGfIlAH9lPfmDBsDZpdPj78r4B oieIJG96Himo9MRh9lFLwzyk5KEDkz9jDt6xVDW7cMoYjloVngQZ6fI6I0Yr2DSTNunU aA9FEpGiAWIdmmmJ4EppInr4zjtSti+5NApVPn6Ix4Dbk54NKs/2NYsTKCy5KwzhDnnq GZFg== X-Gm-Message-State: AFuF++l3XoH/uNhd8v+S9Gf4w93zZs90r/PMI1e/NB8QNoB448cUdH7j M8/YqLVnXrEDE5URwvDSXpcgJQ4CVFs+flQ56tHPYeyRnJDWg5RGTPQg0g95L2rqCJA= X-Gm-Gg: AYBFou14fFNXyFmb6goIhMSzSpMi0iId4WV+k+ieJsI7/rI+m66z5IC7GZlcAg6US4e 9joEB7zHhoGa9pGfRpe6JfD+1E9x3ntk1aVIUCaYGEiUqZC8gF/FdF3Z1+r14bEltNAkULQimay /3flmn4lty9inFnpXuoN+O6WtBitGH0SYzWpPHLJJgbU/3lSKOALKFNnP7pSj5M/S4UwxutEXqe 6/V/ksRpINUQL+0zLGGyyq2GEC6u68Ll85dYJi5eEB7MNGuS3N6wVYXUjU3HH5Mhe2gbiLl2bKq xW1Zn9XfBTPJzXYsBMdaP57rl9VZgjy70Ik+6phl7rsx4nRClSzZ5hSAKQxfE/L12sMr3Wvj+Y7 0y+sxDZ83mwXhILCAjZIsyEKkGM8w2BDvUoqHjMY4H0NCXsJFljlZknfn6BFeMZLm8+zdHXtbbF Xkb+8p0dmBjS+5SlMxEh4wY9+fTbR2MfygJD/GJ0dBYi0BUtEVouPHORUNQ4w0BGmDaVQBX8w9b vcp3whr6amtRFuPjlfUKGuOPErWOCrdXN656ceogFdOU40duBqdsNmNwPFZDw== X-Received: by 2002:a05:620a:19a3:b0:930:9585:e08e with SMTP id af79cd13be357-93c25080ba2mr41386185a.10.1790101782763; Tue, 22 Sep 2026 11:29:42 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c24889ac9sm37130085a.27.2026.09.22.11.29.41 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 11:29:42 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com Subject: [PATCH v3 6/7] mm: use BIT() for change_protection() flags Date: Tue, 22 Sep 2026 14:29:27 -0400 Message-ID: <20260922182928.2199090-7-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260922182928.2199090-1-gourry@gourry.net> References: <20260922182928.2199090-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" The MM_CP_* flags are individual bits expressed as open-coded shifts. Use BIT() consistently for the complete block. No functional change intended. Suggested-by: David Hildenbrand Link: https://lore.kernel.org/r/9b5e1573-6fd8-48a5-a274-a6e7ff83d2fe@kernel= .org Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- include/linux/mm.h | 14 +++++++------- 1 file changed, 7 insertions(+), 7 deletions(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index 225ba26c9d503..c2ffba42019c0 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -3596,21 +3596,21 @@ int get_cmdline(struct task_struct *task, char *buf= fer, int buflen); * because something (e.g., COW, uffd-wp) blocks that from happening for a= ll * PTEs automatically in a writable mapping. */ -#define MM_CP_TRY_CHANGE_WRITABLE (1UL << 0) +#define MM_CP_TRY_CHANGE_WRITABLE BIT(0) /* Whether this protection change is for NUMA hints */ -#define MM_CP_PROT_NUMA (1UL << 1) +#define MM_CP_PROT_NUMA BIT(1) /* Whether this change is for write protecting */ -#define MM_CP_UFFD_WP (1UL << 2) /* do wp */ -#define MM_CP_UFFD_WP_RESOLVE (1UL << 3) /* Resolve wp */ +#define MM_CP_UFFD_WP BIT(2) /* do wp */ +#define MM_CP_UFFD_WP_RESOLVE BIT(3) /* Resolve wp */ #define MM_CP_UFFD_WP_ALL (MM_CP_UFFD_WP | \ MM_CP_UFFD_WP_RESOLVE) /* Whether this change is for uffd RWP */ -#define MM_CP_UFFD_RWP (1UL << 4) /* do rwp */ -#define MM_CP_UFFD_RWP_RESOLVE (1UL << 5) /* resolve rwp */ +#define MM_CP_UFFD_RWP BIT(4) /* do rwp */ +#define MM_CP_UFFD_RWP_RESOLVE BIT(5) /* resolve rwp */ #define MM_CP_UFFD_RWP_ALL (MM_CP_UFFD_RWP | \ MM_CP_UFFD_RWP_RESOLVE) /* Whether a MM_CP_PROT_NUMA change is for promotion only */ -#define MM_CP_PROT_NUMA_PROMO_ONLY (1UL << 6) +#define MM_CP_PROT_NUMA_PROMO_ONLY BIT(6) =20 bool can_change_pte_writable(struct vm_area_struct *vma, unsigned long add= r, pte_t pte); --=20 2.53.0-Meta From nobody Thu Sep 24 15:11:02 2026 Received: from mail-qk2-f41.google.com (mail-qk2-f41.google.com [74.125.230.233]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B61049B44C for ; Tue, 22 Sep 2026 18:29:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.233 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101787; cv=none; b=hPv9+Jqdh1zSZkwVYKAi/LgA/wwb65mO/Mu8tjBPWDjxdRbT7UYY+6wmnRDw+UZW3n9MZK8XZvsazpQeqqR+vl356Gz2GgOBGRgSHA7j/vE3956fibPf/4F/Ufp5uC4XYDntzBO6bo5mRn2dMudA7/HCBQ8fN2nCbCN/Dz2C5Qw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790101787; c=relaxed/simple; bh=n6JqVCExDT3lN80dM6YcCQVpCL3hT3SrhDfa6lpEbp4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=s/177sd1CG/J8FYO8/VrD4m4wKzXYvj6zfE2guglQGVP8oq0pV0HupA6ZckpkPgobvzgSgkDgBjg+/zdwdSi3jERkyjA5/MHfSUWk0gLcPNBPhrOlT6+3Z7CvStwdox2WAYArZpIdvf+bqJc8itgeihPd7NGLdCLbpDfA0jYWls= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=jQQCzHcj; arc=none smtp.client-ip=74.125.230.233 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="jQQCzHcj" Received: by mail-qk2-f41.google.com with SMTP id af79cd13be357-93bfa7b2093so22299785a.1 for ; Tue, 22 Sep 2026 11:29:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1790101785; x=1790706585; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=IEBH2G+jGtGbGG4yhB6a6/Qs6IdlZYnf6Lv/AMN6/lM=; b=jQQCzHcjFH+hl2ldSJxKE1x/ZZwI/1emLrqVH98PFsYRXP+F1EpQ1/LQWdoxT65kdH s4KB+ytslUFA4ToNRCObgM/o28BaZpOPbUKgW4YG871Zv1bdX5g4SqwMLCb9JU3Oc1Ks PMLjd4lP1Mi+7DThyEJcmOrVEnB4IlfnjZ3W3VZCHETV3VN4KpscoIzcOGwsFo/8AaKv GNe7EB3P9EPLx55jFK3+BSnGDF7+1JJnuzkNqAn9xt+sNU1tQei57ZdQ3D2+n6GPpwa4 O/EpCvjc+0i/ZRnK/Edtb8BaAaDc71xj7qi2ktvp0MHN6Xk3Wu+u056quj2SJSDLzabD j2zA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790101785; x=1790706585; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=IEBH2G+jGtGbGG4yhB6a6/Qs6IdlZYnf6Lv/AMN6/lM=; b=ArGZfmee7jkO+VrokPfm9wgHMTCG8WtjoJAbmbwwRWX5rbVwJJN6kEOqQP1HKrmfDK MkZeZlpsxDpTu0flQR9amD52kUW4DLwY7PNNhqVi7EAF1+YG0pskdp0oif9TucR7Or1c d+YDxni9mIgcz4Y0Cv7khV/vMyL0vVDwVRurlglZsyHjsprmZj8O9fBFuPYZMrHxaCn4 LR87YhUiVRpCtyPJ5vUXyrolSJkgmN7jrks5bW0r+vSFjVS8nDvZyh+L+RPUfKclp7t5 TMHuUGH6+F+87qWvSvBwxw9edFmX32gTRYYUDDqbuGyz30+NKEGo1sY4vso5lOsi6W6j +qvA== X-Gm-Message-State: AFuF++kmSm98PM8348ukHCa+c8P0ozSGeqARA4p4+lueVFaue/m+3iU+ twLbgt0FfnGMpnCN0XnoVxhyY6Ir5viKIzFsZXZVeMo9uwNrYja694dzlPEHepISlI8= X-Gm-Gg: AYBFou2DB0rB9PVYyagkRIPSZYKnsuQbJZzEkx5cufL0SChMy7gaww2MGb8nz5e45i0 Q4Yf4+GQKjGINsQd2+xW9yCE4+m3z50zyWNwT3vIjl/LHvwWdTLPHKrf3KgMOYLoqusQ0yJyhgm TQ1FsXLWStnwS4on9IFbYvUh7m6+5xxwwnUH9FMEW9pBIqLRjYCmOs3V5Fp5Mk0LdpmxiiO3nyn REqCof86Gg8hZw8CHWfgjPFrkhVjuzvte2/BSqx8d7Zykd75aVnIbYflc1MHhZMCKWOaKKDWvLM m5sTA7mwasYbj0saA8Q4NMRUAKpNnahsOY3H0XoezY0hYEE+GGkXxzGBDxHUnb5s4NcOJQt3Kra eS1stFPbY7uBUQ3tCBzMgPJBkncRkwOoZ1+YkxoxCkw22pwWs3dBwJU15RPgmpfG2dW17Gar3Kp p3Eb4VGgqfgi9FOERYu38p+ciWD3CTxUSX3215eRJcNo6k1gNlgSCVZtPd7E8anutEa+IkBMfKQ o6TgTvSlDS2oxDf88FNfNhPSEgra9VW9G8gwpSPYzlw7qXDbUqa7gYA3fQ8 X-Received: by 2002:a05:620a:29c8:b0:939:2c5c:7977 with SMTP id af79cd13be357-93c250a3567mr41425285a.20.1790101784800; Tue, 22 Sep 2026 11:29:44 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c24889ac9sm37130085a.27.2026.09.22.11.29.43 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 22 Sep 2026 11:29:44 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, baolin.wang@linux.alibaba.com, nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com, baohua@kernel.org, lance.yang@linux.dev, usama.arif@linux.dev, kas@kernel.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, jannh@google.com, pfalcato@suse.de, hannes@cmpxchg.org, shy828301@gmail.com, raghavendra.kt@amd.com Subject: [PATCH v3 7/7] mm: use VMA flag helpers in NUMA balancing Date: Tue, 22 Sep 2026 14:29:28 -0400 Message-ID: <20260922182928.2199090-8-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260922182928.2199090-1-gourry@gourry.net> References: <20260922182928.2199090-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Gregory Price (Meta)" Use vma_test() instead of accessing vm_flags directly in the NUMA balancing code touched by the preceding fixes. Keep the existing flag predicates unchanged. No functional change intended. Suggested-by: Lorenzo Stoakes Link: https://lore.kernel.org/r/aq1FDhy00epXXtgd@gremlin Assisted-by: LLM Signed-off-by: Gregory Price (Meta) --- kernel/sched/fair.c | 3 ++- mm/internal.h | 2 +- mm/migrate.c | 2 +- 3 files changed, 4 insertions(+), 3 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 412c72084a63d..0d48d595740bf 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4224,7 +4224,8 @@ static void task_numa_work(struct callback_head *work) * under cache pressure can drive considerable bandwidth. */ placement_scan =3D !(vma->vm_file && - (vma->vm_flags & (VM_READ | VM_WRITE)) =3D=3D VM_READ); + vma_test(vma, VMA_READ_BIT) && + !vma_test(vma, VMA_WRITE_BIT)); =20 if (!vma->vm_mm || (!placement_scan && !tiering)) { trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_SHARED_RO); diff --git a/mm/internal.h b/mm/internal.h index cbeaf1fc38cfc..560406c36a757 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1240,7 +1240,7 @@ size_t splice_folio_into_pipe(struct pipe_inode_info = *pipe, =20 static inline bool vma_is_single_threaded_private(struct vm_area_struct *v= ma) { - if (vma->vm_flags & VM_SHARED) + if (vma_test(vma, VMA_SHARED_BIT)) return false; =20 return atomic_read(&vma->vm_mm->mm_users) =3D=3D 1; diff --git a/mm/migrate.c b/mm/migrate.c index 6a08690cfe219..8e1e5ff1cb580 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -2705,7 +2705,7 @@ int migrate_misplaced_folio_prepare(struct folio *fol= io, * See folio_maybe_mapped_shared() on possible imprecision * when we cannot easily detect if a folio is shared. */ - if ((vma->vm_flags & VM_EXEC) && + if (vma_test(vma, VMA_EXEC_BIT) && folio_maybe_mapped_shared(folio) && (!folio_in_lowtier(folio) || !node_is_toptier(node))) return -EACCES; --=20 2.53.0-Meta