From nobody Sat Jul 25 02:10:53 2026 Received: from out-172.mta0.migadu.com (out-172.mta0.migadu.com [91.218.175.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AFF673FA5C8 for ; Mon, 20 Jul 2026 16:42:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784565749; cv=none; b=a5El3fabhZeuwLvRBMHPZQApVOsKgT38GzTKmdDXAOaYyyGlvqh2m1YFzwZQbF79ATju4CTuXd1UHbeUxmuk7xpi41HLEu1qphnz9Zx6fEKs+ccn7QWlLFVhwZTNjvmTpUb5D39mjh/TlOvmfJjzc4st09RJP+PUM47gN+11BHU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784565749; c=relaxed/simple; bh=pUBlWuRQMiTLJTLYkrKc+8utTMOhy7TEVGGqWyB4sSk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=oGQvTqp8vT36r4rndEi9wtrZbTNOsPICis0nJxiUZvOtvHAvjbbCpkWtZUCxO9DDtj8tUmbAzdoL+LOBwbSpfOtoHQ3uP94slaNZz8qKeYjlQZ5MMnTihJ7iwgP3w4cf9gj6u7lO3U1nvjx7haIxJ4Zb48SKrwpNR3TI51x9D50= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=bSf73LT6; arc=none smtp.client-ip=91.218.175.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="bSf73LT6" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1784565743; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=EqMnQVIkWoxb+RrEs+jMVyYBu0LSbT8WfW7w+P6K7uQ=; b=bSf73LT6zmv6iHRzh1nyG6HQ5XmmTwZZ07TNO0fiFD3dsGWdYHcv3gwPKWOmvfUK0/+1eJ 3meZXRvGIYrzUoY7rsap1WGZOkw9i86k/1/MJkiRzvub+OTV74RSAJfuouRFSJPCf3BpGf Q316LzrD3OIedsEgeruFH6ijtO5euzc= From: Usama Arif To: Andrew Morton , david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, chrisl@kernel.org, nphamcs@gmail.com, baoquan.he@linux.dev, youngjun.park@lge.com, hannes@cmpxchg.org, roman.gushchin@linux.dev, muchun.song@linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, rientjes@google.com, kernel-team@meta.com Cc: Usama Arif Subject: [PATCH v4 1/2] mm/vmstat, mm/memcontrol: add _monotonic vmstat readers Date: Mon, 20 Jul 2026 09:41:22 -0700 Message-ID: <20260720164207.450685-2-usama.arif@linux.dev> In-Reply-To: <20260720164207.450685-1-usama.arif@linux.dev> References: <20260720164207.450685-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" lruvec_page_state(), node_page_state(), and global_node_page_state() all clamp negative reads to zero on CONFIG_SMP so that a transient per-CPU delta skew presents as zero pages rather than as a garbage unsigned value. This is the right behaviour for non-monotonic page-count readers. It is however incorrect for callers that snapshot a monotonically- incremented event counter and compute a delta from two samples. Once the underlying signed long wraps past LONG_MAX, the clamped read drops to zero while the previously-recorded snapshot still holds the pre-wrap value; the unsigned subtraction then underflows into a ~2^31 spurious delta for 32-bit architecture and corrupts the caller's accumulator. Add non-clamping siblings that return the underlying state value cast to unsigned long: global_node_page_state_monotonic() node_page_state_monotonic() lruvec_page_state_monotonic() With both samples read via the _monotonic variant, unsigned modular subtraction stays correct across a signed-long wraparound as long as the true growth between two samples fits in unsigned long (< 2^32 on 32-bit, < 2^64 on 64-bit); the 32-bit bound is the practically-reachable one that motivates this helper. The variants are only safe for monotonically-incremented counters. Non-monotonic page-count readers must keep using the existing clamped helpers so transient negative reads still present as zero. This is a prerequisite for the following patch which replaces the producer-side anon_cost/file_cost accumulators with a read-side accumulator in prepare_scan_control() that samples monotonic per-LRU vmstat counters (PGROTATE_*, PGRECLAIM_PAGEOUT_*, WORKINGSET_RESTORE_*) via lruvec_page_state_monotonic() and folds the unsigned modular delta into a per-lruvec cost_accum[]. Acked-by: Johannes Weiner Signed-off-by: Usama Arif Acked-by: Shakeel Butt --- include/linux/memcontrol.h | 8 ++++++++ include/linux/vmstat.h | 16 ++++++++++++++++ mm/memcontrol.c | 36 ++++++++++++++++++++++++++++++++++++ mm/vmstat.c | 11 +++++++++++ 4 files changed, 71 insertions(+) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index e1f46a0016fc..b40bc4f6fe4a 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -931,6 +931,8 @@ unsigned long memcg_page_state_output(struct mem_cgroup= *memcg, int item); bool memcg_stat_item_valid(int idx); bool memcg_vm_event_item_valid(enum vm_event_item idx); unsigned long lruvec_page_state(struct lruvec *lruvec, enum node_stat_item= idx); +unsigned long lruvec_page_state_monotonic(struct lruvec *lruvec, + enum node_stat_item idx); unsigned long lruvec_page_state_local(struct lruvec *lruvec, enum node_stat_item idx); =20 @@ -1378,6 +1380,12 @@ static inline unsigned long lruvec_page_state(struct= lruvec *lruvec, return node_page_state(lruvec_pgdat(lruvec), idx); } =20 +static inline unsigned long lruvec_page_state_monotonic(struct lruvec *lru= vec, + enum node_stat_item idx) +{ + return node_page_state_monotonic(lruvec_pgdat(lruvec), idx); +} + static inline unsigned long lruvec_page_state_local(struct lruvec *lruvec, enum node_stat_item idx) { diff --git a/include/linux/vmstat.h b/include/linux/vmstat.h index 3c9c266cf782..fb8c76289e02 100644 --- a/include/linux/vmstat.h +++ b/include/linux/vmstat.h @@ -194,6 +194,19 @@ unsigned long global_node_page_state_pages(enum node_s= tat_item item) return x; } =20 +/* + * Non-clamping variant of global_node_page_state() intended for callers t= hat + * snapshot a monotonically-incremented counter and subtract two samples. + * Returns the raw wrapping value so that unsigned modular subtraction sta= ys + * correct across a signed-long overflow (a real hazard on 32-bit) that the + * clamp in global_node_page_state() would otherwise turn into a huge spur= ious + * delta. Do NOT use for non-monotonic page-count reads. + */ +static inline unsigned long global_node_page_state_monotonic(enum node_sta= t_item item) +{ + return (unsigned long)atomic_long_read(&vm_node_stat[item]); +} + static inline unsigned long global_node_page_state(enum node_stat_item ite= m) { VM_WARN_ON_ONCE(vmstat_item_in_bytes(item)); @@ -259,11 +272,14 @@ extern unsigned long node_page_state(struct pglist_da= ta *pgdat, enum node_stat_item item); extern unsigned long node_page_state_pages(struct pglist_data *pgdat, enum node_stat_item item); +extern unsigned long node_page_state_monotonic(struct pglist_data *pgdat, + enum node_stat_item item); extern void fold_vm_numa_events(void); #else #define sum_zone_node_page_state(node, item) global_zone_page_state(item) #define node_page_state(node, item) global_node_page_state(item) #define node_page_state_pages(node, item) global_node_page_state_pages(ite= m) +#define node_page_state_monotonic(node, item) global_node_page_state_monot= onic(item) static inline void fold_vm_numa_events(void) { } diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 56cd4af08232..b4a357c5f7e0 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -502,6 +502,42 @@ unsigned long lruvec_page_state(struct lruvec *lruvec,= enum node_stat_item idx) return x; } =20 +/** + * lruvec_page_state_monotonic - non-clamping lruvec stat read for delta s= ampling + * @lruvec: the LRU vector to read from + * @idx: the node_stat_item to read + * + * Returns the raw state[idx] value cast to unsigned long, skipping the + * clamp-negative-to-zero step in lruvec_page_state(). Intended for callers + * that snapshot a monotonically-incremented counter and subtract two + * samples: unsigned modular arithmetic then yields the correct delta acro= ss + * a signed-long wraparound (a real hazard on 32-bit) that the clamp would + * otherwise turn into a huge spurious delta. + * + * Do NOT use for non-monotonic page-count reads where a transient negative + * reading from per-CPU delta skew must present as zero. + * + * XXX: This helper (and its node/global peers) exists because we place + * monotonically-incremented event counters (PGROTATE_*, PGRECLAIM_PAGEOUT= _*) + * into enum node_stat_item. + */ +unsigned long lruvec_page_state_monotonic(struct lruvec *lruvec, + enum node_stat_item idx) +{ + struct mem_cgroup_per_node *pn; + int i; + + if (mem_cgroup_disabled()) + return node_page_state_monotonic(lruvec_pgdat(lruvec), idx); + + i =3D memcg_stats_index(idx); + if (WARN_ONCE(BAD_STAT_IDX(i), "%s: missing stat item %d\n", __func__, id= x)) + return 0; + + pn =3D container_of(lruvec, struct mem_cgroup_per_node, lruvec); + return (unsigned long)READ_ONCE(pn->lruvec_stats->state[i]); +} + unsigned long lruvec_page_state_local(struct lruvec *lruvec, enum node_stat_item idx) { diff --git a/mm/vmstat.c b/mm/vmstat.c index f534972f517d..c4364f0eb08a 100644 --- a/mm/vmstat.c +++ b/mm/vmstat.c @@ -1024,6 +1024,17 @@ unsigned long node_page_state(struct pglist_data *pg= dat, =20 return node_page_state_pages(pgdat, item); } + +/* + * Non-clamping variant of node_page_state() intended for callers that + * snapshot a monotonically-incremented counter and subtract two samples. + * See global_node_page_state_monotonic() for the rationale. + */ +unsigned long node_page_state_monotonic(struct pglist_data *pgdat, + enum node_stat_item item) +{ + return (unsigned long)atomic_long_read(&pgdat->vm_stat[item]); +} #endif =20 /* --=20 2.53.0-Meta From nobody Sat Jul 25 02:10:53 2026 Received: from out-173.mta1.migadu.com (out-173.mta1.migadu.com [95.215.58.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0F3E23DA5AE for ; Mon, 20 Jul 2026 16:42:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.173 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784565760; cv=none; b=frnsrkZ/1C3NlY1YMQbgwQpYS8oJOHjfmbYXBxSvjwIp4Xc2HyfhEHerymUqtxjPHKHVuvovWoIdMej5cfeVFO2RLgyBt8rOFz3tjdH7AqVTFVmPBrpjCYyoJ/DYlyDZKpQGF68kQyzidVRpSN3zMKNaXpTh/xu8zB9CCIAd/kc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784565760; c=relaxed/simple; bh=XkkueEqDJj2EMJaLrH+isYHS2/2C3fvYvQvkJYjIz2g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sl2zJ2Llcx0NiJvJos/GpyWG6mPGw7KGf4ZOuIi6M6gSC0+zR0N8wABA9VAFY8Nm6njnXYSXLCy1iV7TSniRb0U5WLRZglNY08TA9/0tgdN7Vp76tQGXgSyu31qusXQa+TYdLTMRc8xeXglu6e0bpHUTemloFDUKIksJu2RafKo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=cN6G++7W; arc=none smtp.client-ip=95.215.58.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="cN6G++7W" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1784565755; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=bPozRyj5PIXs2wgYkG/A+cTHgkjXSiK5aMeh3QMvNHY=; b=cN6G++7W/9+OLuxSiAhn4yElmeDmmioK3+PeDVk4/IZUcu12BjMaSlmwGwdepyQrIB6Qts 6osZU3K6VDjYTtYClFVU5i+SwBw1O4pfZaBzUFKccKC/taN0Md79mEPi29n2Sb+6pLkhLw BJB37IWptOXBA+BmV8Yg4wt+d8JuV+A= From: Usama Arif To: Andrew Morton , david@kernel.org, ljs@kernel.org, liam@infradead.org, vbabka@kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, kasong@tencent.com, qi.zheng@linux.dev, shakeel.butt@linux.dev, axelrasmussen@google.com, yuanchu@google.com, weixugc@google.com, chrisl@kernel.org, nphamcs@gmail.com, baoquan.he@linux.dev, youngjun.park@lge.com, hannes@cmpxchg.org, roman.gushchin@linux.dev, muchun.song@linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, rientjes@google.com, kernel-team@meta.com Cc: Usama Arif Subject: [PATCH v4 2/2] mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance cost Date: Mon, 20 Jul 2026 09:41:23 -0700 Message-ID: <20260720164207.450685-3-usama.arif@linux.dev> In-Reply-To: <20260720164207.450685-1-usama.arif@linux.dev> References: <20260720164207.450685-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" The anon/file scan balance in get_scan_count() is driven by two scalars in struct lruvec, anon_cost and file_cost, accumulated by every reclaim producer under lruvec->lru_lock. The acquisition sites for cost work specifically are: - shrink_inactive_list() re-takes lru_lock at function exit purely to call lru_note_cost_unlock_irq() with (nr_pageout, nr_scanned - nr_reclaimed). One acquisition per inactive shrink. - shrink_active_list() does the same with (0, nr_rotated). One acquisition per active shrink. - workingset_refault() takes the lock via folio_lruvec_lock_irq() purely to record the refault cost. One acquisition per refault. - prepare_scan_control() takes lru_lock just to snapshot the two scalars into sc->{anon,file}_cost. - lru_note_cost_unlock_irq() itself walks parent_lruvec and re-acquires lru_lock on each ancestor to propagate the update, adding O(memcg-depth) acquisitions per producer call. This hurts because lru_lock is already a heavy contention point on memory-heavy workloads: every isolate_lru_folios(), move_folios_to_lru() and folio_add_lru() takes it. The cost work itself is trivial (two scalar bumps and one comparison), but it contends with and causes contention for actual LRU manipulation. The parent_lruvec() walk also multiplies cost-update overhead by memcg hierarchy depth. Replace the producer-side accumulators with a read-side accumulator fed from per-LRU vmstat counters. The old producer formula was: cost =3D nr_io * SWAP_CLUSTER_MAX + nr_rotated Reuse NR_VMSCAN_WRITE for reclaim-driven anon pageout submissions. It is already bumped by writeout() for the same successful outcome that fed reclaim_stat.nr_pageout. Reclaim does not submit filesystem folios from this path, so there is no file pageout term. Charge NR_VMSCAN_WRITE via lruvec_stat_mod_folio() and include it in memcg_node_stat_items so it can be sampled per lruvec and aggregated through the memcg hierarchy. Add explicit PGROTATE_{ANON,FILE} node_stat counters for the remaining producer-local input. They are bumped from shrink_inactive_list() by nr_scanned - nr_reclaimed and from shrink_active_list() by nr_rotated. WORKINGSET_RESTORE_{ANON,FILE} already captures the refault IO that lru_note_cost_refault() used to bill. Add a per-side struct lru_cost { count, last_rotated, last_io } to struct lruvec. In prepare_scan_control() the two monotonic inputs are sampled separately - rotated from PGROTATE_ANON/FILE, io from WORKINGSET_RESTORE_BASE + f plus (for anon) NR_VMSCAN_WRITE - and the raw per-side deltas are computed against cost->last_rotated and cost->last_io before the SWAP_CLUSTER_MAX IO weighting is applied. Extracting the deltas from the individual counters (rather than from a pre-weighted sum) keeps the unsigned modular subtraction bounded by the true per-counter growth, so a signed-long wraparound of any underlying vmstat still yields the correct delta on 32-bit. The weighted delta is folded into cost->count. Since one vmstat delta can cover many producer events between reclaim passes, halve cost->count on both sides until their sum is back within the lrusize/4 bound instead of halving only once. Moving accumulation and decay to the reclaim side also improves the cost model across reclaim gaps. With producer-side decay, events that happen while reclaim is idle still age each other before reclaim ever samples the costs. If a workload refaults a large anon set and then a smaller file set before reclaim runs again, the later file activity can age the earlier anon activity out of the cost model. The new scheme observes the whole between-reclaim delta and decays anon and file proportionally, so the scan-balance history better represents what happened since the last reclaim pass. A dedicated per-lruvec spinlock, cost_lock, serialises the delta extraction, the cost->count update and the halving loop against concurrent reclaimers in the same memcg+node. Hierarchy aggregation is now implicit in the vmstat accounting. The producer-side parent_lruvec() walk and lru_reparent_memcg() cost splice existed only because anon_cost/file_cost were private lruvec fields. With the cost expressed as lruvec vmstats, rstat propagates the underlying counters through the memcg hierarchy and prepare_scan_control() consumes the same ratelimited rstat view as the surrounding reclaim heuristics. NR_VMSCAN_WRITE is accounted at writeout(), so reclaim_stat.nr_pageout is no longer needed and is removed. memcg-v1's memory.stat anon_cost/file_cost is now sourced from cost[].count instead of the removed lruvec anon_cost/file_cost fields. The reported values only refresh when prepare_scan_control() runs and are bounded at ~lrusize/4 by the halving loop; the scan-balance signal they express is unchanged. Under pure MGLRU the scan-balance signal itself is not consumed (both prepare_scan_control() and get_scan_count() are short-circuited on the MGLRU paths, and MGLRU's own type/tier selection comes from read_ctrl_pos() on lrugen->{avg_refaulted,avg_total,refaulted,evicted}, not from anon_cost/file_cost). NR_VMSCAN_WRITE naturally covers writeout from either reclaim implementation, and PGROTATE_{ANON,FILE} are bumped from evict_folios() so per-memcg observability of rotation-driven reclaim work stays consistent across both implementations. Signed-off-by: Usama Arif Acked-by: Johannes Weiner --- include/linux/mmzone.h | 15 ++++++- include/linux/swap.h | 3 -- include/linux/vmstat.h | 1 - mm/memcontrol-v1.c | 4 +- mm/memcontrol.c | 5 ++- mm/mmzone.c | 1 + mm/swap.c | 69 ------------------------------- mm/vmscan.c | 93 +++++++++++++++++++++++++++++++++++------- mm/vmstat.c | 2 + mm/workingset.c | 5 --- 10 files changed, 101 insertions(+), 97 deletions(-) diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h index ca2712187147..85303c5867c8 100644 --- a/include/linux/mmzone.h +++ b/include/linux/mmzone.h @@ -323,6 +323,8 @@ enum node_stat_item { PGSCAN_PROACTIVE, PGSCAN_ANON, PGSCAN_FILE, + PGROTATE_ANON, + PGROTATE_FILE, PGREFILL, #ifdef CONFIG_HUGETLB_PAGE NR_HUGETLB, @@ -755,6 +757,12 @@ void lru_gen_reparent_memcg(struct mem_cgroup *memcg, = struct mem_cgroup *parent, =20 #endif /* CONFIG_LRU_GEN */ =20 +struct lru_cost { + unsigned long count; + unsigned long last_rotated; + unsigned long last_io; +}; + struct lruvec { struct list_head lists[NR_LRU_LISTS]; /* per lruvec lru_lock for memcg */ @@ -763,9 +771,12 @@ struct lruvec { * These track the cost of reclaiming one LRU - file or anon - * over the other. As the observed cost of reclaiming one LRU * increases, the reclaim scan balance tips toward the other. + * Updated and decayed at prepare_scan_control() time; cost_lock + * serialises that update. */ - unsigned long anon_cost; - unsigned long file_cost; + struct lru_cost cost[ANON_AND_FILE]; + /* Protects cost[]. */ + spinlock_t cost_lock; /* Non-resident age, driven by LRU movement */ atomic_long_t nonresident_age; /* Refaults at the time of last reclaim cycle */ diff --git a/include/linux/swap.h b/include/linux/swap.h index 6d72778e6cc3..d35a4761ebd7 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -309,9 +309,6 @@ extern unsigned long totalreserve_pages; =20 =20 /* linux/mm/swap.c */ -void lru_note_cost_unlock_irq(struct lruvec *lruvec, bool file, - unsigned int nr_io, unsigned int nr_rotated); -void lru_note_cost_refault(struct folio *); void folio_add_lru(struct folio *); void folio_add_lru_vma(struct folio *, struct vm_area_struct *); void mark_page_accessed(struct page *); diff --git a/include/linux/vmstat.h b/include/linux/vmstat.h index fb8c76289e02..5b31d8e7ae40 100644 --- a/include/linux/vmstat.h +++ b/include/linux/vmstat.h @@ -20,7 +20,6 @@ struct reclaim_stat { unsigned nr_congested; unsigned nr_writeback; unsigned nr_immediate; - unsigned nr_pageout; unsigned nr_activate[ANON_AND_FILE]; unsigned nr_ref_keep; unsigned nr_unmap_fail; diff --git a/mm/memcontrol-v1.c b/mm/memcontrol-v1.c index 765069211567..091bc9ffee44 100644 --- a/mm/memcontrol-v1.c +++ b/mm/memcontrol-v1.c @@ -1988,8 +1988,8 @@ void memcg1_stat_format(struct mem_cgroup *memcg, str= uct seq_buf *s) for_each_online_pgdat(pgdat) { mz =3D memcg->nodeinfo[pgdat->node_id]; =20 - anon_cost +=3D mz->lruvec.anon_cost; - file_cost +=3D mz->lruvec.file_cost; + anon_cost +=3D mz->lruvec.cost[WORKINGSET_ANON].count; + file_cost +=3D mz->lruvec.cost[WORKINGSET_FILE].count; } seq_buf_printf(s, "anon_cost %lu\n", anon_cost); seq_buf_printf(s, "file_cost %lu\n", file_cost); diff --git a/mm/memcontrol.c b/mm/memcontrol.c index b4a357c5f7e0..8693aad26ca2 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -393,6 +393,7 @@ static const unsigned int memcg_node_stat_items[] =3D { NR_SHMEM_THPS, NR_FILE_THPS, NR_ANON_THPS, + NR_VMSCAN_WRITE, NR_VMALLOC, NR_KERNEL_STACK_KB, NR_PAGETABLE, @@ -419,6 +420,8 @@ static const unsigned int memcg_node_stat_items[] =3D { PGSCAN_PROACTIVE, PGSCAN_ANON, PGSCAN_FILE, + PGROTATE_ANON, + PGROTATE_FILE, PGREFILL, #ifdef CONFIG_HUGETLB_PAGE NR_HUGETLB, @@ -518,7 +521,7 @@ unsigned long lruvec_page_state(struct lruvec *lruvec, = enum node_stat_item idx) * reading from per-CPU delta skew must present as zero. * * XXX: This helper (and its node/global peers) exists because we place - * monotonically-incremented event counters (PGROTATE_*, PGRECLAIM_PAGEOUT= _*) + * monotonically-incremented event counters (NR_VMSCAN_WRITE and PGROTATE_= *) * into enum node_stat_item. */ unsigned long lruvec_page_state_monotonic(struct lruvec *lruvec, diff --git a/mm/mmzone.c b/mm/mmzone.c index 0c8f181d9d50..17139db4d291 100644 --- a/mm/mmzone.c +++ b/mm/mmzone.c @@ -78,6 +78,7 @@ void lruvec_init(struct lruvec *lruvec) =20 memset(lruvec, 0, sizeof(struct lruvec)); spin_lock_init(&lruvec->lru_lock); + spin_lock_init(&lruvec->cost_lock); zswap_lruvec_state_init(lruvec); =20 for_each_lru(lru) diff --git a/mm/swap.c b/mm/swap.c index 588f50d8f1a8..74b281778cbc 100644 --- a/mm/swap.c +++ b/mm/swap.c @@ -272,73 +272,6 @@ void folio_rotate_reclaimable(struct folio *folio) folio_batch_add_and_move(folio, lru_move_tail); } =20 -void lru_note_cost_unlock_irq(struct lruvec *lruvec, bool file, - unsigned int nr_io, unsigned int nr_rotated) - __releases(lruvec->lru_lock) - __releases(rcu) -{ - unsigned long cost; - - /* - * Reflect the relative cost of incurring IO and spending CPU - * time on rotations. This doesn't attempt to make a precise - * comparison, it just says: if reloads are about comparable - * between the LRU lists, or rotations are overwhelmingly - * different between them, adjust scan balance for CPU work. - */ - cost =3D nr_io * SWAP_CLUSTER_MAX + nr_rotated; - if (!cost) { - spin_unlock_irq(&lruvec->lru_lock); - rcu_read_unlock(); - return; - } - - for (;;) { - unsigned long lrusize; - - /* Record cost event */ - if (file) - lruvec->file_cost +=3D cost; - else - lruvec->anon_cost +=3D cost; - - /* - * Decay previous events - * - * Because workloads change over time (and to avoid - * overflow) we keep these statistics as a floating - * average, which ends up weighing recent refaults - * more than old ones. - */ - lrusize =3D lruvec_page_state(lruvec, NR_INACTIVE_ANON) + - lruvec_page_state(lruvec, NR_ACTIVE_ANON) + - lruvec_page_state(lruvec, NR_INACTIVE_FILE) + - lruvec_page_state(lruvec, NR_ACTIVE_FILE); - - if (lruvec->file_cost + lruvec->anon_cost > lrusize / 4) { - lruvec->file_cost /=3D 2; - lruvec->anon_cost /=3D 2; - } - - spin_unlock_irq(&lruvec->lru_lock); - lruvec =3D parent_lruvec(lruvec); - if (!lruvec) { - rcu_read_unlock(); - break; - } - spin_lock_irq(&lruvec->lru_lock); - } -} - -void lru_note_cost_refault(struct folio *folio) -{ - struct lruvec *lruvec; - - lruvec =3D folio_lruvec_lock_irq(folio); - lru_note_cost_unlock_irq(lruvec, folio_is_file_lru(folio), - folio_nr_pages(folio), 0); -} - static void lru_activate(struct lruvec *lruvec, struct folio *folio) { long nr_pages =3D folio_nr_pages(folio); @@ -1164,8 +1097,6 @@ void lru_reparent_memcg(struct mem_cgroup *memcg, str= uct mem_cgroup *parent, int =20 child_lruvec =3D mem_cgroup_lruvec(memcg, NODE_DATA(nid)); parent_lruvec =3D mem_cgroup_lruvec(parent, NODE_DATA(nid)); - parent_lruvec->anon_cost +=3D child_lruvec->anon_cost; - parent_lruvec->file_cost +=3D child_lruvec->file_cost; =20 for_each_lru(lru) lruvec_reparent_lru(child_lruvec, parent_lruvec, lru, nid); diff --git a/mm/vmscan.c b/mm/vmscan.c index e8a90911bf88..0f6334005610 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -641,7 +641,7 @@ static pageout_t writeout(struct folio *folio, struct a= ddress_space *mapping, folio_clear_reclaim(folio); =20 trace_mm_vmscan_write_folio(folio); - node_stat_add_folio(folio, NR_VMSCAN_WRITE); + lruvec_stat_mod_folio(folio, NR_VMSCAN_WRITE, folio_nr_pages(folio)); return PAGE_SUCCESS; } =20 @@ -1418,8 +1418,6 @@ static unsigned int shrink_folio_list(struct list_hea= d *folio_list, sc->nr_scanned -=3D (nr_pages - 1); nr_pages =3D 1; } - stat->nr_pageout +=3D nr_pages; - if (folio_test_writeback(folio)) goto keep; if (folio_test_dirty(folio)) @@ -2043,10 +2041,10 @@ static unsigned long shrink_inactive_list(unsigned = long nr_to_scan, item =3D PGSTEAL_KSWAPD + reclaimer_offset(sc); mod_lruvec_state(lruvec, item, nr_reclaimed); mod_lruvec_state(lruvec, PGSTEAL_ANON + file, nr_reclaimed); + if (nr_scanned > nr_reclaimed) + mod_lruvec_state(lruvec, PGROTATE_ANON + file, + nr_scanned - nr_reclaimed); =20 - lruvec_lock_irq(lruvec); - lru_note_cost_unlock_irq(lruvec, file, stat.nr_pageout, - nr_scanned - nr_reclaimed); handle_reclaim_writeback(nr_taken, pgdat, sc, &stat); trace_mm_vmscan_lru_shrink_inactive(pgdat->node_id, nr_scanned, nr_reclaimed, &stat, sc->priority, file); @@ -2152,9 +2150,9 @@ static void shrink_active_list(unsigned long nr_to_sc= an, count_vm_events(PGDEACTIVATE, nr_deactivate); count_memcg_events(lruvec_memcg(lruvec), PGDEACTIVATE, nr_deactivate); mod_node_page_state(pgdat, NR_ISOLATED_ANON + file, -nr_taken); + if (nr_rotated) + mod_lruvec_state(lruvec, PGROTATE_ANON + file, nr_rotated); =20 - lruvec_lock_irq(lruvec); - lru_note_cost_unlock_irq(lruvec, file, 0, nr_rotated); trace_mm_vmscan_lru_shrink_active(pgdat->node_id, nr_taken, nr_activate, nr_deactivate, nr_rotated, sc->priority, file); } @@ -2287,8 +2285,10 @@ enum scan_balance { =20 static void prepare_scan_control(pg_data_t *pgdat, struct scan_control *sc) { - unsigned long file; + struct lru_cost *anon_cost, *file_cost; struct lruvec *target_lruvec; + unsigned long lrusize; + unsigned long file; =20 if (lru_gen_enabled() && !lru_gen_switching()) return; @@ -2304,11 +2304,69 @@ static void prepare_scan_control(pg_data_t *pgdat, = struct scan_control *sc) =20 /* * Determine the scan balance between anon and file LRUs. + * + * The cost model is based on rotations, refaults and + * reclaim-driven writes (anon only) on each side. + * + * These event counters are monotonic, so each reclaim cycle + * the delta since the last scan is extracted and incorporated + * into a decaying average. This ensures currency, as workloads + * change over time, and avoids overflow in the calculations. + * + * Use lruvec_page_state_monotonic() so unsigned subtraction + * yields the correct delta across a signed-long wraparound of + * the underlying counter (a real hazard on 32-bit that the + * clamp in lruvec_page_state() would otherwise turn into a huge + * spurious delta). */ - spin_lock_irq(&target_lruvec->lru_lock); - sc->anon_cost =3D target_lruvec->anon_cost; - sc->file_cost =3D target_lruvec->file_cost; - spin_unlock_irq(&target_lruvec->lru_lock); + spin_lock(&target_lruvec->cost_lock); + + for (int f =3D 0; f <=3D 1; f++) { + struct lru_cost *cost =3D &target_lruvec->cost[f]; + unsigned long rotated, io, nr_rotated, nr_io; + + rotated =3D lruvec_page_state_monotonic(target_lruvec, + PGROTATE_ANON + f); + io =3D lruvec_page_state_monotonic(target_lruvec, + WORKINGSET_RESTORE_BASE + f); + if (f =3D=3D WORKINGSET_ANON) + io +=3D lruvec_page_state_monotonic(target_lruvec, + NR_VMSCAN_WRITE); + + nr_rotated =3D rotated - cost->last_rotated; + nr_io =3D io - cost->last_io; + + /* + * Reflect the relative cost of incurring IO and spending + * CPU time on rotations. This doesn't attempt to make a + * precise comparison, it just says: if reloads are about + * comparable between the LRU lists, or rotations are + * overwhelmingly different between them, adjust scan + * balance for CPU work. + */ + cost->count +=3D nr_io * SWAP_CLUSTER_MAX + nr_rotated; + + cost->last_rotated =3D rotated; + cost->last_io =3D io; + } + + anon_cost =3D &target_lruvec->cost[WORKINGSET_ANON]; + file_cost =3D &target_lruvec->cost[WORKINGSET_FILE]; + + lrusize =3D lruvec_page_state(target_lruvec, NR_INACTIVE_ANON) + + lruvec_page_state(target_lruvec, NR_ACTIVE_ANON) + + lruvec_page_state(target_lruvec, NR_INACTIVE_FILE) + + lruvec_page_state(target_lruvec, NR_ACTIVE_FILE); + + while (anon_cost->count + file_cost->count > lrusize / 4) { + anon_cost->count /=3D 2; + file_cost->count /=3D 2; + } + + sc->anon_cost =3D anon_cost->count; + sc->file_cost =3D file_cost->count; + + spin_unlock(&target_lruvec->cost_lock); =20 /* * Target desirable inactive:active list ratios for the anon @@ -4815,7 +4873,8 @@ static int evict_folios(unsigned long nr_to_scan, str= uct lruvec *lruvec, struct reclaim_stat stat; struct lru_gen_mm_walk *walk; int scanned, reclaimed; - int isolated =3D 0, type, type_scanned; + int isolated =3D 0, nr_isolated =3D 0, type, type_scanned; + unsigned long total_reclaimed =3D 0; bool skip_retry =3D false; struct mem_cgroup *memcg =3D lruvec_memcg(lruvec); struct pglist_data *pgdat =3D lruvec_pgdat(lruvec); @@ -4827,6 +4886,7 @@ static int evict_folios(unsigned long nr_to_scan, str= uct lruvec *lruvec, =20 scanned =3D isolate_folios(nr_to_scan, lruvec, sc, swappiness, &list, &isolated, &type, &type_scanned); + nr_isolated =3D isolated; =20 /* Scanning may have emptied the oldest gen, flush it */ if (scanned) @@ -4839,6 +4899,7 @@ static int evict_folios(unsigned long nr_to_scan, str= uct lruvec *lruvec, retry: reclaimed =3D shrink_folio_list(&list, pgdat, sc, &stat, false, memcg); sc->nr_reclaimed +=3D reclaimed; + total_reclaimed +=3D reclaimed; /* Retry pass is only meant for clean folios without new isolation */ if (isolated) handle_reclaim_writeback(isolated, pgdat, sc, &stat); @@ -4892,6 +4953,10 @@ static int evict_folios(unsigned long nr_to_scan, st= ruct lruvec *lruvec, goto retry; } =20 + if (nr_isolated > total_reclaimed) + mod_lruvec_state(lruvec, PGROTATE_ANON + type, + nr_isolated - total_reclaimed); + return scanned; } =20 diff --git a/mm/vmstat.c b/mm/vmstat.c index c4364f0eb08a..87d4a6781367 100644 --- a/mm/vmstat.c +++ b/mm/vmstat.c @@ -1300,6 +1300,8 @@ const char * const vmstat_text[] =3D { [I(PGSCAN_PROACTIVE)] =3D "pgscan_proactive", [I(PGSCAN_ANON)] =3D "pgscan_anon", [I(PGSCAN_FILE)] =3D "pgscan_file", + [I(PGROTATE_ANON)] =3D "pgrotate_anon", + [I(PGROTATE_FILE)] =3D "pgrotate_file", [I(PGREFILL)] =3D "pgrefill", #ifdef CONFIG_HUGETLB_PAGE [I(NR_HUGETLB)] =3D "nr_hugetlb", diff --git a/mm/workingset.c b/mm/workingset.c index f351798e723a..7ac2b88c80ae 100644 --- a/mm/workingset.c +++ b/mm/workingset.c @@ -584,11 +584,6 @@ void workingset_refault(struct folio *folio, void *sha= dow) /* Folio was active prior to eviction */ if (workingset) { folio_set_workingset(folio); - /* - * XXX: Move to folio_add_lru() when it supports new vs - * putback - */ - lru_note_cost_refault(folio); mod_lruvec_state(lruvec, WORKINGSET_RESTORE_BASE + file, nr); } out: --=20 2.53.0-Meta