From nobody Sat Sep 26 03:57:33 2026 Received: from mail-qk1-f176.google.com (mail-qk1-f176.google.com [209.85.222.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A18B151A759 for ; Fri, 4 Sep 2026 18:20:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.176 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788546021; cv=none; b=PfEji1OjUV2cRz7S0wlFJzKX/QSxFc/Q/lL5VSbOwx2MV6CLiW0jW61f1hgenvFOB98mHccIctuB0xTE6GkwI7OegJhdbNHaZROhhNabCzF6ZrYvW37oJtvt6oNDJGGlxQn45dXClFVhGG0vfviErESFZUMJcwZDDephpQAxlKs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788546021; c=relaxed/simple; bh=DJOU1m53y2E4/TLdlgmhdCS/oFkXY5aFxDtOM6g4YSQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=N0KByqDeck9JJxdS5jMLeFHIlquEIHzQtT2XGOBH7CjynlX+ANdo1Krkxo7IQq7LVxp1omzvZFEXt1Zky7ycvJgOyJRMNZ95ssKoPPdT5UiNI9TxCl+meHIQk4KVeb11wOrf8WaSf/TPkcICTh6iC+2dTDV6qfsPvlzQZ2TNn+8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=FS0EUrKx; arc=none smtp.client-ip=209.85.222.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="FS0EUrKx" Received: by mail-qk1-f176.google.com with SMTP id af79cd13be357-9397e2994dcso113502985a.2 for ; Fri, 04 Sep 2026 11:20:19 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1788546018; x=1789150818; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=p8Z4JBi3sjE2KlaCLrWKRkykHYLdYdycFQqiI4ckxMI=; b=FS0EUrKx3TIwUxcSqoKuByIJtaSUPXh1iqRG1vqhbp9T2md1RsJyE3O2MCOFof4lYK Sfinz6B3x3YrRZAN6ae+tK4IqfurMQ2k7HLHtN+UT6Q+cTx/JCCfivaxGFp22Yr42IMz UQlrPfs+FWQB9p3UKpC7Ol1LW5fqGCyc6U7sOawmbPaMJfRb5rfZccgMXvfFWV/FioWR 0QNDNAMzHJ2ZfoOZ8Uj8RTOUri3HoitZd6l5pksfOYBpo65AU+jSFMz7caODvDVrGcxW pqXC5XotwPJswvqFf1ycRPXGe8z2QwI9VFwRZqeLGFzBVfjqQhRsib4LFUXqA59ayRP0 +xUg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788546018; x=1789150818; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=p8Z4JBi3sjE2KlaCLrWKRkykHYLdYdycFQqiI4ckxMI=; b=ijwQbKDiiUKJVwR4gwVxhDtAF4ilvWZY9im27eD6+4Nb4TbwQ3F9WuWjzpp9ITQhYV kcKSgP/0T2gFyGZAwSsScyu++/vQK4VlaYkLTRD95xCBREJoNx84cQokomOc5OAoSuxp rytbH3OI2Cqwjv2pjODN2YgMSp0T8P72TfOa8q5bopiInoZjAwe81aKvGV0IJumf+P0v 4JgWssNIv6szcqFGQucxmP76KRNYguwO44gl9/lg3TzdMgjVQE10wqeO5I1HichEC6c2 CmtIRMuayQtrJDPPsAS2q71vJIR1xg6UIm6BGMrSS2N7FblcXlesuG67X5+YVeUqO2ya M1sA== X-Gm-Message-State: AFuF++ldDfwAAk2QuFsNMNGQ3Hvo7/3p1IWhSUWxwUJ4aFSrkIeESy5i aQatU6WSd3IwowZ+2fK0iS2BkcuKr1VrjrfeYV7OAyb+C+3c795Wc3SxmAJxhhz7yuc= X-Gm-Gg: AYBFou2KAbTow+dNSsj3k6pf41DtxlSqbO4tVgODSwuCLxwtQM8/Nj0C1J96bymRPL9 ybg1QURpf5aeQVKgRFyboe6ylET4/If2EPhL8EEiCh+htb/0S08w3d9JUvdrdZlcXfJNBO9hTsy EpZN42M0ZCyn3oWpA5cCe/9sj+8qgGRX87NNMwuf0OcNykRzy03jM0/Po4QGYAt+RBbXF8nAMdG 0g9Rgkujm/LAQlcb0j4BXOiSOpNBTod4VbFIdX3oT7600RZfTF8eXpS07YxBttoUij8PbQjOIMw kfLZFnjSg5T7pQxhw7gXB0KAxVYE7jLXLfAAFyCF2h9MU7jE8vFa2IkiHodihu/PBFsVF2NoWl3 /tPH59iqlJorxiOqOgCQArgvbVh9jgx1VtzlYE+FS0f694ubeLDRTxdhJ7RezWxQoBiQZFB4gIS xVVeSKI8UUsVDu6PTY5HmLHmpCCQ9Xn8AVJLus5aVCth3VwmwZHdaITjn0MEWZmCS6c2In7tGSH mn9iIBH9zpJqiur5pg9ZjZtdfSI01OVeX3LtnYnF9ZNfL5VbQ== X-Received: by 2002:a05:620a:2715:b0:939:8a03:f7e3 with SMTP id af79cd13be357-9398a03ff25mr362976385a.6.1788546018205; Fri, 04 Sep 2026 11:20:18 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93982db6a95sm207374585a.9.2026.09.04.11.20.17 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 04 Sep 2026 11:20:17 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, raghavendra.kt@amd.com, shy828301@gmail.com, baolin.wang@linux.alibaba.com, osalvador@suse.de Subject: [PATCH 1/2] sched/numa: do not let the per-VMA PID filter gate promotion Date: Fri, 4 Sep 2026 14:20:05 -0400 Message-ID: <20260904182006.1562449-2-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260904182006.1562449-1-gourry@gourry.net> References: <20260904182006.1562449-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" task_numa_work() will not scan a VMA that vma_is_accessed() rejects: if (!vma_pids_forced && !vma_is_accessed(mm, vma)) { vma_pids_skipped =3D true; trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_PID_INACTIVE); continue; } vma_is_accessed() wants the scanning thread's hash bit in vma->numab_state->pids_active[]. Those bits have one setter, vma_set_access_pid_bit(), called on a NUMA hint fault. The PROT_NONE scan is the only source of hint faults, so a VMA the task has not hint-faulted looks untouched - preventing new scans from causing faults for potentially hours or days. The existing escape mechanisms do not solve this problem: numa_scan_offset - rescues only the VMA holding the scan cursor get_nr_threads() - the horizon from commit f22cde4371f3 ("sched/numa: Fix the vma scan starving issue") costs nr_threads scan sequences, which can be incredibly large. vma_pids_forced retry - fires only at the end of the VMA list, and then scans exactly one VMA Observed effect in the following environment: - 768G DRAM + 256G CXL Host - two ~430GB highly-threaded database workload - a single large (300GB+) shmem/tmpfs VMA in each workload - numa_balancing =3D 2 (default tuning) In this environment, the large shmem VMAs held the cursor for 100% of a scan, while 84G across 2537 other VMAs were skipped due to "inactivity". Among them, a 20G hash table ended up 100% on CXL. After this patch it was consistently split 50:50 between DRAM and CXL. Scan the rejected VMAs, looking only at folios on non-toptier nodes. Promotion candidates then come from the whole address space, while toptier folios are only marked where vma_is_accessed() already allowed it, so top-tier marking is unchanged in every mode. A restricted scan reaches the end of the VMA but may not completely scan it, so it leaves prev_scan_seq alone. Setting it there would permanently disarm the get_nr_threads() horizon above. Before this change: DRAM Bandwidth utilization dropped from 200GB/s to 150GB/s CXL Bandwidth utilization increased from 5GB/s to 45GB/s (capped) Request Latency increases from 800us to 5ms tracked with CXL bandwidth After this change: DRAM Bandwidth maintains between 200GB/s and 250GB/s CXL Bandwidth maintains at ~7-10GB/s Request Latency hovers between 800us-2ms tracked with request load. commit d230991493b5 ("mm: mempolicy: fix automatic numa balancing for shmem= ") surfaced this issue when it fixed shmem numa balance scanning. Fixes: fc137c0ddab2 ("sched/numa: enhance vma scanning logic") Signed-off-by: Gregory Price (Meta) Assisted-by: Claude:claude-opus-5 --- include/linux/mm_types.h | 7 +++++++ kernel/sched/fair.c | 25 ++++++++++++++++++++----- mm/mempolicy.c | 9 +++++---- 3 files changed, 32 insertions(+), 9 deletions(-) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 5413bd10fff2..d986fcab9c8f 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -803,6 +803,13 @@ struct vma_numab_state { * A VMA is not eligible for scanning if prev_scan_seq =3D=3D numa_scan_s= eq */ int prev_scan_seq; + + /* + * Set by the scanner for the duration of a scan of this VMA when only + * folios on non-toptier nodes should be made NUMA hintable, because + * the scanning task has not accessed the VMA recently. + */ + bool slow_only; }; =20 #ifdef __HAVE_PFNMAP_TRACKING diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 6d881e530f89..990af0e45d69 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4297,11 +4297,20 @@ static void task_numa_work(struct callback_head *wo= rk) /* * Do not scan the VMA if task has not accessed it, unless no other * VMA candidate exists. + * + * Force-scan only slow-tier folios when in tiering mode, as a large + * VMA can cause others to become "permanently unaccessed" if a scan + * cycle is consumed entirely by the large VMA. */ + vma->numab_state->slow_only =3D false; if (!vma_pids_forced && !vma_is_accessed(mm, vma)) { - vma_pids_skipped =3D true; - trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_PID_INACTIVE); - continue; + if (!(sysctl_numa_balancing_mode & + NUMA_BALANCING_MEMORY_TIERING)) { + vma_pids_skipped =3D true; + trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_PID_INACTIVE); + continue; + } + vma->numab_state->slow_only =3D true; } =20 do { @@ -4329,8 +4338,14 @@ static void task_numa_work(struct callback_head *wor= k) cond_resched(); } while (end !=3D vma->vm_end); =20 - /* VMA scan is complete, do not scan until next sequence. */ - vma->numab_state->prev_scan_seq =3D mm->numa_scan_seq; + /* + * VMA scan is complete, do not scan until next sequence. + * A restricted scan reached the end of the VMA but may not + * have completely scanned it - it must not disarm the sequence + * counting horizon in vma_is_accessed() for this VMA. + */ + if (!vma->numab_state->slow_only) + vma->numab_state->prev_scan_seq =3D mm->numa_scan_seq; =20 /* * Only force scan within one VMA at a time, to limit the diff --git a/mm/mempolicy.c b/mm/mempolicy.c index ce10ce437464..e3fe08c4f109 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -887,11 +887,12 @@ bool folio_can_map_prot_numa(struct folio *folio, str= uct vm_area_struct *vma, return false; =20 /* - * Skip scanning top tier node if normal numa - * balancing is disabled + * Skip scanning top tier node if normal numa balancing is disabled, + * or if the scanner only wants promotion candidates out of this VMA. */ - if (!(sysctl_numa_balancing_mode & NUMA_BALANCING_NORMAL) && - node_is_toptier(nid)) + if (node_is_toptier(nid) && + (!(sysctl_numa_balancing_mode & NUMA_BALANCING_NORMAL) || + (vma->numab_state && vma->numab_state->slow_only))) return false; =20 if (folio_use_access_time(folio)) --=20 2.55.0 From nobody Sat Sep 26 03:57:33 2026 Received: from mail-qk1-f170.google.com (mail-qk1-f170.google.com [209.85.222.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6EC9251D500 for ; Fri, 4 Sep 2026 18:20:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788546023; cv=none; b=jRiijq/XMmDblEyuNhUnAgbxxW2gjJAmuNeL3fwEJj7iNahenGer8hBdx29j/sPjHqlAtUOq2sba3rGf62AWarLSyJpAkYb36n7BhstkT4FCNB6dJ/70fbd/2heUaNpGdOwLsYiX87nfZpwg74MvXlFE3OkUgGVYqS/NcLYau3Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788546023; c=relaxed/simple; bh=sRMAupkd+/z8tPz3CPJ9c9jvDKH4MQCfdZ1Wr8kL6Bo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PzbK21aQ+5uKRrY4cLg4K3kyYy81wP8mZ2eKFEKieLRVm03XpykaMSoJjNArSqdikITfxvbH72zvOeBJwfT83Y3uW3s1qrkaNJ29gcRCLauRR+MfPU5spLlYrlmL93BL2c3AYz3RULaX5tGE3uYjoS2gafXCSuHDrCxXMAyXKBc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=pKoF4lDb; arc=none smtp.client-ip=209.85.222.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="pKoF4lDb" Received: by mail-qk1-f170.google.com with SMTP id af79cd13be357-92ed19f4d60so134423285a.0 for ; Fri, 04 Sep 2026 11:20:21 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1788546020; x=1789150820; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=GSRixT7jynn3gCypVDydAS92P1PrLhM0zXSgAeyy22g=; b=pKoF4lDbTnokPdaBIm1CyUlldv+Ql80dF8EZqy55hQYHUuXlyan1FzO9rAlwxBXnqV aRYDKyrC08/LsQZgAkd4+zYUGMGsb2u0Sq3ue306afzbJ2iwXxkOPDpLyOXz6ZTY+8N4 UT9UnPd1jwT22onjHMf2rkrxrnHJG3dRAGeq4F9qzEPKIhirEL4HewpriAOnDnhsRMkq AT+QUd6m5lnUBycNrWcQ+Dk8ds8Uy+RPz/yZ/fNsHzqjJQl87ZV+MyAJyNeZ5ggkVede XCarp4zVfl4no06t6w+ArvrhLaWcIAB4HbjDyZZZakqLmwhGp/8SJ7afrObCWQlIP/MU w5KA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788546020; x=1789150820; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=GSRixT7jynn3gCypVDydAS92P1PrLhM0zXSgAeyy22g=; b=PGk376OZFOBVMD9BX6hRZQ2vZJE2EATKCzWKeVze8WxEB1dHGaEMTAExzP34L1+dsh bwZCs771kHWiz56szHXkRBImRHSUE8NXqCaRtQ45w+tckLqv/1X/PEovClHCahIAV/9x 8StAoOGqoAsLa3U3sQKJzOCKrznYM7or6eqQKbadkCSA5Kjh8Je9zqMOpY+bJsbe3Vu3 d36VfOd04lzB7Wql4h3ihGHG0bJ6yKfeS4tin3SQ58Jq3bO9xFfdZyBkRP03hbV+gYae f1kitqaIcm+ct6IN6BC0SnDa7o5w0bE+4DBh+Hb+f3jGAleO6xv3Cp54Pj/AnQ4EShA7 bVWA== X-Gm-Message-State: AFuF++mhAYGz8p/zaCMjP3yTFPqwi061DKf72r+2oIanCMn3mYMZqIqJ biFE24SW4qA1bPmJKUSaRHyK/isWEJPcHpRuGHfo7qRVy0LcKNHbxJ0BVesRGI9Y49c= X-Gm-Gg: AYBFou13/oM/UrUuz5edUBwkCGbzla4Mn7MeyNaRLDfMUg8CIVU9LshNhvTnQ89aTpR O0L0EpsS/ne/69oasMUn2xP3RnjPGSkTIS4iB3IFOdsD78fCEWa6B/TJtjvPT8ADRguHdwxsLlw cnkByCr4PDLEr+kVMxhHeGFgJx5ZZS4P6EyYpd9LK6u2ZppB2sKou+vxpoXrPYuEh8M9b97bDB9 4CUosRoRiLcblyPHqGhBPWU961hcMBrW9RhouSmNk6JrqZRfK0ayzV1UXbWRktgSqMRtOKKd60H TPIjm2BDCE3HXZJrcRJ8bBUebgkMOalT4FYMMlAdUtfse2RKfN6whSRCA+eaROPIKfj0yp1LI4w kxu+CwIvr/LK2+Fg/Dqdm1mHRHV+KDnqx/uK3CUIeoOs+W8pMR3qivt8QLA3EDZveiwNo/eWZ7+ cZ89md5zoOZFVDUmij50Ztt+46DbDhRwzRjcI3xJGMjENEet8vFXIUmUY3X00rIywtM6ZLoeeQ/ GqhGRrbAH7vdPuFo1LV9bBW2iCmjdfmOiw9XhvRqwDxgOYvrE7T77cHcBaaYWU= X-Received: by 2002:a05:620a:63c8:b0:939:8d2:325d with SMTP id af79cd13be357-939817398b7mr507496985a.21.1788546020031; Fri, 04 Sep 2026 11:20:20 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93982db6a95sm207374585a.9.2026.09.04.11.20.19 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 04 Sep 2026 11:20:19 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com, akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, raghavendra.kt@amd.com, shy828301@gmail.com, baolin.wang@linux.alibaba.com, osalvador@suse.de Subject: [PATCH 2/2] sched/numa: scan read-only file mappings in tiering mode Date: Fri, 4 Sep 2026 14:20:06 -0400 Message-ID: <20260904182006.1562449-3-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260904182006.1562449-1-gourry@gourry.net> References: <20260904182006.1562449-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" task_numa_work() refuses to scan any read-only file-backed mapping: if (vma->vm_file && (vma->vm_flags & (VM_READ|VM_WRITE)) =3D=3D (VM_READ)) { trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_SHARED_RO); continue; } commit 4591ce4f2d22 ("sched/numa: Do not trap hinting faults for shared libraries") added this because such pages are expected to be cache replicated and bounce between accessor nodes, so: If we are never going to migrate the page, it is overhead for no gain Memory tiering broke that premise. A slow tier folio is promoted on hint fault based on hotness, so these can migrate now - they just never get the chance, because the VMA is never scanned. Observed on a 768G DRAM + 256G CXL host running a ~430G database service: 91% of its main binary (169M of 185M) accumulated on CXL. After this patch the binary's tier residency tracked runtime load. Scan these VMAs with the slow_only restriction. Only non-toptier folios are made hint-faultable, so the scan never marks a toptier folio in a read-only file mapping in any mode. A page already on the fast tier costs nothing, and a page on the slow tier costs one hint fault, which is what buys the promotion. That fault is not free of placement effect: a successful promotion reports its toptier destination node to task_numa_fault(), so it lands in numa_faults[] like any other promotion. Only the set of VMAs that can produce one changes. Fixes: c574bbe91703 ("NUMA balancing: optimize page placement for memory ti= ering system") Signed-off-by: Gregory Price (Meta) Assisted-by: Claude:claude-opus-5 --- kernel/sched/fair.c | 18 ++++++++++++++++-- 1 file changed, 16 insertions(+), 2 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 990af0e45d69..5c5c80b17649 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -4075,6 +4075,11 @@ static void reset_ptenuma_scan(struct task_struct *p) p->mm->numa_scan_offset =3D 0; } =20 +static bool vma_is_ro_file(struct vm_area_struct *vma) +{ + return vma->vm_file && (vma->vm_flags & (VM_READ | VM_WRITE)) =3D=3D VM_R= EAD; +} + static bool vma_is_accessed(struct mm_struct *mm, struct vm_area_struct *v= ma) { unsigned long pids; @@ -4211,6 +4216,8 @@ static void task_numa_work(struct callback_head *work) } =20 for (; vma; vma =3D vma_next(&vmi)) { + bool ro_file; + if (!vma_migratable(vma) || !vma_policy_mof(vma) || is_vm_hugetlb_page(vma) || (vma->vm_flags & VM_MIXEDMAP)) { trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_UNSUITABLE); @@ -4222,9 +4229,16 @@ static void task_numa_work(struct callback_head *wor= k) * migrated as it is expected they are cache replicated. Avoid * hinting faults in read-only file-backed mappings or the vDSO * as migrating the pages will be of marginal benefit. + * + * Tiering still wants them. A read-only mapping strands just + * as much on the slow tier as any other, so scan it for slow + * tier folios alone. Nothing on the top tier is marked here, + * in any mode. */ + ro_file =3D vma_is_ro_file(vma); if (!vma->vm_mm || - (vma->vm_file && (vma->vm_flags & (VM_READ|VM_WRITE)) =3D=3D (VM_REA= D))) { + (ro_file && !(sysctl_numa_balancing_mode & + NUMA_BALANCING_MEMORY_TIERING))) { trace_sched_skip_vma_numa(mm, vma, NUMAB_SKIP_SHARED_RO); continue; } @@ -4302,7 +4316,7 @@ static void task_numa_work(struct callback_head *work) * VMA can cause others to become "permanently unaccessed" if a scan * cycle is consumed entirely by the large VMA. */ - vma->numab_state->slow_only =3D false; + vma->numab_state->slow_only =3D ro_file; if (!vma_pids_forced && !vma_is_accessed(mm, vma)) { if (!(sysctl_numa_balancing_mode & NUMA_BALANCING_MEMORY_TIERING)) { --=20 2.55.0