From nobody Sat Sep 26 13:08:16 2026 Received: from mta1.migadu.com (out-5.mta1.migadu.com [95.215.58.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DC1993546D0 for ; Tue, 1 Sep 2026 09:13:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788253982; cv=none; b=ct/ihXBJErIC4VFdLGOdMWp0ZbUwUrtH8yqC56x1y5PQVXBckfsPp0SpIFrvmqJ850p26C4qDPJ5/pbH7K+mYQrLVNksHP5cRIIbv7YfPyuJ+ZhYXSIUv+vBnK96hCw16zm/+3u/x9+UiKJSsfqPtRdQDDA+uKQCr3QFBDc6pus= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788253982; c=relaxed/simple; bh=kaupwl2CerM+wcHSCXXaFwLigAE3E8c74QApBbEU49A=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=DICJ4GduYC8qx/IOvsLORWF7yAj3azpE+HHi2jioHSVSSUUD5QWHRaVEdRbl8Jf6TvgSoFKRpy/uAlRKHTqkXfyID8j2FBabMgt8OoSxdfhYtb0RJ8Qrla0ncYec7WFpZL+u88/r1EpyVK8th5j1speLZHM0JnL9TphxVRThm5Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=irU2sJO9; arc=none smtp.client-ip=95.215.58.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="irU2sJO9" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=kaupwl2CerM+wcHSCXXaFwLigAE3E8c74QApBbEU49A=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1788253978; v=1; x=1788858778; b=irU2sJO9q4WUJoILGuRTchvv90VDGfFzrX7swFprOZS1WtztcdnKaSoZVz/BNamTFcR8oQdd Sb6IEq9iJkws2UFE5KgXWKdaDwItzXwWCn1zAAoWIjiy9hxR82tLy8OmGyvn35qdm8N8/DGQCib YeyCuuGdmZX9TgfQ60U62DZc= X-Envelope-To: linux-kernel@vger.kernel.org Received: by smtp.migadu.com with ESMTPS id b5c480e00ff4c702; Tue, 01 Sep 2026 09:12:58 +0000 X-Mizu-Trace-ID: b5c480e00ff4c702 X-Migadu-Flow: FLOW_OUT From: "Hui Zhu" To: Andrew Morton , Kairui Song , Qi Zheng , Shakeel Butt , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Johannes Weiner , David Hildenbrand , Michal Hocko , Lorenzo Stoakes , Baolin Wang , linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: Hui Zhu Subject: [PATCH mm-unstable v5 1/2] mm/vmscan: fix missing NR_ISOLATED counter update in MGLRU reclaim path Date: Tue, 1 Sep 2026 17:12:18 +0800 Message-ID: <65cf410fc9e281bc1de00e938522e5140eb49499.1788252038.git.zhuhui@kylinos.cn> X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Hui Zhu MGLRU evict_folios() isolates folios from the LRU without updating the NR_ISOLATED_ANON/FILE counters, unlike the legacy shrink_inactive_list() path. This causes compaction's too_many_isolated() check to under-count isolated pages when MGLRU reclaim is active. Add NR_ISOLATED counter updates in evict_folios(): increment after isolate_folios() and decrement after all retry passes complete, using the existing nr_isolated which holds the original isolated count. Signed-off-by: Hui Zhu Reviewed-by: Baolin Wang Reviewed-by: Barry Song --- mm/vmscan.c | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/mm/vmscan.c b/mm/vmscan.c index c1404a59523d..98226bb021f3 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -4892,6 +4892,9 @@ static int evict_folios(unsigned long nr_to_scan, str= uct lruvec *lruvec, scanned =3D isolate_folios(nr_to_scan, lruvec, sc, swappiness, &list, &isolated, &type, &type_scanned); nr_isolated =3D isolated; + if (nr_isolated) + __mod_node_page_state(pgdat, NR_ISOLATED_ANON + type, + nr_isolated); =20 /* Scanning may have emptied the oldest gen, flush it */ if (scanned) @@ -4954,6 +4957,8 @@ static int evict_folios(unsigned long nr_to_scan, str= uct lruvec *lruvec, goto retry; } =20 + mod_node_page_state(pgdat, NR_ISOLATED_ANON + type, -nr_isolated); + if (nr_isolated > total_reclaimed) mod_lruvec_state(lruvec, PGROTATE_ANON + type, nr_isolated - total_reclaimed); --=20 2.53.0 From nobody Sat Sep 26 13:08:16 2026 Received: from mta0.migadu.com (out-35.mta0.migadu.com [91.218.175.35]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 069A84779B9 for ; Tue, 1 Sep 2026 09:13:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.35 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788253990; cv=none; b=qRIYbcNOrnE9WuIe/zzMZotNwHnSFIJF2018anuZkm2AOtPBWXAmPTOXssYL5bTB78p0MjNYC7KdWBOsaO6heOZfpKgJM5F7B6VL6eiKt+3x0fAwUNr1frjJdlSGRN8Bg0fxoQFxj0D/Tkjsmu5/x1RrC7ZxsRmPSaE7kV0vIsM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788253990; c=relaxed/simple; bh=8uJOowBp8VCu1y67aq7L73a1vzkStQA+SCBPipldYWU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aLjv4aiQH/3hyLGKiNu7sVhxRLum7tEceYypRcMz6nINducwW/D48gllFSFioNyaUPNRISffobCO1dJVs3MknZJzEXn2MdIXqKdxH6o0ModmWHpSLwIYqrvcvs6d56Js+y79Om8qKqqROtl7sM/+cWjHJfLSHel1J1B1vhKzIek= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=QZh2AOfo; arc=none smtp.client-ip=91.218.175.35 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="QZh2AOfo" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=8uJOowBp8VCu1y67aq7L73a1vzkStQA+SCBPipldYWU=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1788253985; v=1; x=1788858785; b=QZh2AOfokA5HVuMaAADAD9aDdWjbY9f/zIPKSkmnIf6igP+lDhmSA8pC6D8kLrGHjTX+egz1 tGkBPtb9L5ZXXceHwA4sd8i9NQUKfmQPy29r5ioNBl8ba2BlJBCTLzxndyq5Y9VJM2xjotDMEgU DGZ7oYwCsQ4/6EvWxRb9jP5c= X-Envelope-To: linux-kernel@vger.kernel.org Received: by smtp.migadu.com with ESMTPS id 41a52598f51d4228; Tue, 01 Sep 2026 09:13:05 +0000 X-Mizu-Trace-ID: 41a52598f51d4228 X-Migadu-Flow: FLOW_OUT From: "Hui Zhu" To: Andrew Morton , Kairui Song , Qi Zheng , Shakeel Butt , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Johannes Weiner , David Hildenbrand , Michal Hocko , Lorenzo Stoakes , Baolin Wang , linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: Hui Zhu Subject: [PATCH mm-unstable v5 2/2] mm/vmscan: throttle MGLRU eviction when isolated folios pile up Date: Tue, 1 Sep 2026 17:12:19 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Hui Zhu The legacy path throttles direct reclaim via too_many_isolated() when isolated folios pile up, but MGLRU's evict_folios() isolates folios without this check, which can lead to unnecessary swapping, thrashing and OOM. Add lru_gen_too_many_isolated() and throttle_evictable_types(), called from evict_folios() before the lruvec lock is taken since throttling sleeps. The check is per lruvec rather than node-wide: reclaimers contend per lruvec, and a node-wide check would let isolation in one memcg throttle reclaim in another. The threshold is the evictable size of the type divided by MIN_NR_GENS, approximating one generation. Since isolate_folios() picks the type from refault feedback and may fall back to the other one, throttle_evictable_types() collects the types that are not over-isolated and have evictable folios, and sleeps only when all of them are over-isolated. The mask is passed to isolate_folios() to restrict its choice and fallback, so an over-isolated type never blocks the other, and isolation never lands on a throttled or empty type. Like the legacy path, kswapd and !writeback_throttling_sane() reclaim are exempt, the threshold is raised for GFP_NOIO/GFP_NOFS callers, and a dying task fakes progress to exit reclaim quickly. Tested with threads write-touching a 48MB anonymous working set in a 16MB memcg backed by swap delayed 50ms via dm-delay. With 192 threads the task is OOM-killed without this patch and survives with it. With 64 threads both kernels survive the 60s run, and the patched kernel completes ~22% more touches (301B vs 246B). Signed-off-by: Hui Zhu --- include/linux/mmzone.h | 2 + mm/vmscan.c | 170 +++++++++++++++++++++++++++++++++++++---- 2 files changed, 158 insertions(+), 14 deletions(-) diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h index 94f9c3ff5416..819b6c059116 100644 --- a/include/linux/mmzone.h +++ b/include/linux/mmzone.h @@ -580,6 +580,8 @@ struct lru_gen_folio { /* can be modified without holding the LRU lock */ atomic_long_t evicted[NR_HIST_GENS][ANON_AND_FILE][MAX_NR_TIERS]; atomic_long_t refaulted[NR_HIST_GENS][ANON_AND_FILE][MAX_NR_TIERS]; + /* the number of folios isolated from this lruvec and being reclaimed */ + atomic_long_t nr_isolated[ANON_AND_FILE]; /* whether the multi-gen LRU is enabled */ bool enabled; /* the memcg generation this lru_gen_folio belongs to */ diff --git a/mm/vmscan.c b/mm/vmscan.c index 98226bb021f3..b5b38d296004 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -4151,25 +4151,34 @@ static void set_initial_priority(struct pglist_data= *pgdat, struct scan_control sc->priority =3D clamp(priority, DEF_PRIORITY / 2, DEF_PRIORITY); } =20 -static unsigned long lruvec_evictable_size(struct lruvec *lruvec, int swap= piness) +static unsigned long lruvec_type_evictable_size(struct lruvec *lruvec, int= type) { - int gen, type, zone; + int gen, zone; unsigned long seq, total =3D 0; struct lru_gen_folio *lrugen =3D &lruvec->lrugen; DEFINE_MAX_SEQ(lruvec); DEFINE_MIN_SEQ(lruvec); =20 - for_each_evictable_type(type, swappiness) { - for (seq =3D min_seq[type]; seq <=3D max_seq; seq++) { - gen =3D lru_gen_from_seq(seq); - for (zone =3D 0; zone < MAX_NR_ZONES; zone++) - total +=3D max(READ_ONCE(lrugen->nr_pages[gen][type][zone]), 0L); - } + for (seq =3D min_seq[type]; seq <=3D max_seq; seq++) { + gen =3D lru_gen_from_seq(seq); + for (zone =3D 0; zone < MAX_NR_ZONES; zone++) + total +=3D max(READ_ONCE(lrugen->nr_pages[gen][type][zone]), 0L); } =20 return total; } =20 +static unsigned long lruvec_evictable_size(struct lruvec *lruvec, int swap= piness) +{ + unsigned long total =3D 0; + int type; + + for_each_evictable_type(type, swappiness) + total +=3D lruvec_type_evictable_size(lruvec, type); + + return total; +} + static bool lruvec_is_sizable(struct lruvec *lruvec, struct scan_control *= sc) { unsigned long total; @@ -4833,15 +4842,128 @@ static int get_type_to_scan(struct lruvec *lruvec,= int swappiness) return positive_ctrl_err(&sp, &pv); } =20 +/* + * The MGLRU counterpart to too_many_isolated(). + * + * too_many_isolated() compares node-wide isolated counts against + * node-wide inactive sizes. It predates per-memcg LRU lists, which + * moved isolation to per-lruvec granularity. MGLRU reclaimers contend + * per lruvec, so this check is per lruvec as well, and isolation in + * one memcg does not throttle reclaim in another. + * + * The legacy threshold, the size of the inactive list, is not usable + * here either: in MGLRU the inactive counters are only a compatibility + * shim - they jump around when swap runs out or a memcg hits its swap + * limit, and proactive aging can inflate them. Compare against the + * total number of evictable pages of the type divided by MIN_NR_GENS, + * approximating the size of one generation - the same unit the + * original MGLRU aging heuristics used. + */ +static bool lru_gen_too_many_isolated(struct lruvec *lruvec, int type, + struct scan_control *sc) +{ + unsigned long isolated, evictable; + + if (current_is_kswapd()) + return false; + + if (!writeback_throttling_sane(sc)) + return false; + + isolated =3D atomic_long_read(&lruvec->lrugen.nr_isolated[type]); + evictable =3D lruvec_type_evictable_size(lruvec, type); + + /* + * GFP_NOIO/GFP_NOFS callers are allowed to isolate more pages, so + * they won't be blocked by normal direct-reclaimers, forming a + * circular deadlock. + */ + if (gfp_has_io_fs(sc->gfp_mask)) + evictable >>=3D 3; + + return isolated > evictable / MIN_NR_GENS; +} + +/* + * Unlike the legacy path, where the LRU list to isolate from is known + * before isolation, isolate_folios() picks the type from the refault + * feedback and may fall back to the other one. Therefore, instead of + * throttling on a single type, collect the evictable types that do not + * have too many isolated folios, and only sleep when all of them do. + * + * Returns the mask of the types isolate_folios() may isolate from, or + * 0 if reclaim should stop. Also sets @fatal to tell the caller that + * the task received a fatal signal while waiting. + */ +static unsigned int throttle_evictable_types(struct lruvec *lruvec, + int swappiness, + struct scan_control *sc, + bool *fatal) +{ + unsigned int allowed; + bool stalled =3D false; + struct pglist_data *pgdat =3D lruvec_pgdat(lruvec); + int i; + + *fatal =3D false; + + for (;;) { + allowed =3D 0; + for_each_evictable_type(i, swappiness) { + /* + * A type with no evictable folios has nothing to + * isolate. Don't allow it, otherwise reclaim can be + * redirected to it and spin making no progress, e.g. + * when the only type with folios is over-isolated. + */ + if (!lruvec_type_evictable_size(lruvec, i)) + continue; + + if (!lru_gen_too_many_isolated(lruvec, i, sc)) + allowed |=3D BIT(i); + } + + if (allowed) { + /* Wake up reclaimers waiting on the isolation to go down. */ + wake_throttle_isolated(pgdat); + return allowed; + } + + /* + * All evictable types are over-isolated. Like the legacy + * path, wait once for concurrent reclaimers to put their + * isolated folios back; give up if that makes no progress. + */ + if (stalled) + return 0; + + stalled =3D true; + reclaim_throttle(pgdat, VMSCAN_THROTTLE_ISOLATED); + + /* We are about to die and free our memory. Return now. */ + if (fatal_signal_pending(current)) { + *fatal =3D true; + return 0; + } + } +} + static int isolate_folios(unsigned long nr_to_scan, struct lruvec *lruvec, struct scan_control *sc, int swappiness, - struct list_head *list, int *isolated, - int *isolate_type, int *isolate_scanned) + unsigned int allowed, struct list_head *list, + int *isolated, int *isolate_type, int *isolate_scanned) { int i; int total_scanned =3D 0; int type =3D get_type_to_scan(lruvec, swappiness); =20 + /* + * The preferred type may have been excluded by + * throttle_evictable_types(); start from the other one. + */ + if (!(allowed & BIT(type))) + type =3D !type; + for_each_evictable_type(i, swappiness) { int scanned; int tier =3D get_tier_idx(lruvec, type); @@ -4858,9 +4980,10 @@ static int isolate_folios(unsigned long nr_to_scan, = struct lruvec *lruvec, /* * If scanned > 0 and isolated =3D=3D 0, avoid falling back to the * other type, as this type remains sufficient. Falling back - * too readily can disrupt the positive_ctrl_err() bias. + * too readily can disrupt the positive_ctrl_err() bias. Only + * fall back to a type that is not throttled. */ - if (!scanned) + if (!scanned && (allowed & BIT(!type))) type =3D !type; } =20 @@ -4883,18 +5006,36 @@ static int evict_folios(unsigned long nr_to_scan, s= truct lruvec *lruvec, bool skip_retry =3D false; struct mem_cgroup *memcg =3D lruvec_memcg(lruvec); struct pglist_data *pgdat =3D lruvec_pgdat(lruvec); + unsigned int allowed; + bool fatal; + + allowed =3D throttle_evictable_types(lruvec, swappiness, sc, &fatal); + if (!allowed) { + /* + * We are about to die and free our memory. Like the legacy + * path, pretend some pages were reclaimed so reclaim + * unwinds quickly instead of looping back into the + * throttle. + */ + if (fatal) + sc->nr_reclaimed +=3D SWAP_CLUSTER_MAX; + + return 0; + } =20 lruvec_lock_irq(lruvec); =20 /* In case folio deletion left empty old gens, flush them */ try_to_inc_min_seq(lruvec, swappiness); =20 - scanned =3D isolate_folios(nr_to_scan, lruvec, sc, swappiness, + scanned =3D isolate_folios(nr_to_scan, lruvec, sc, swappiness, allowed, &list, &isolated, &type, &type_scanned); nr_isolated =3D isolated; - if (nr_isolated) + if (nr_isolated) { __mod_node_page_state(pgdat, NR_ISOLATED_ANON + type, nr_isolated); + atomic_long_add(nr_isolated, &lruvec->lrugen.nr_isolated[type]); + } =20 /* Scanning may have emptied the oldest gen, flush it */ if (scanned) @@ -4957,6 +5098,7 @@ static int evict_folios(unsigned long nr_to_scan, str= uct lruvec *lruvec, goto retry; } =20 + atomic_long_sub(nr_isolated, &lruvec->lrugen.nr_isolated[type]); mod_node_page_state(pgdat, NR_ISOLATED_ANON + type, -nr_isolated); =20 if (nr_isolated > total_reclaimed) --=20 2.53.0