From nobody Sat Jul 25 03:47:43 2026 Received: from sender-pp-o93.zoho.in (sender-pp-o93.zoho.in [103.117.158.93]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4BC3B26059D for ; Sun, 19 Jul 2026 18:56:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=pass smtp.client-ip=103.117.158.93 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487390; cv=pass; b=kgAzKW0AZjlOR3kCvQDCRAO4FVl4ZNbMsPkl1B5MDazOY+7cxWW8+G9ydDezNXUI0epm1FNlEB70dJBMr6YkvmFqCwUvo99eDwghodlwjmOiBJJce7CND+g743C+xOkjRW7GWZaBfGQs8h9kV3cQ5+hFycs/Re/0a6DXqLFN1Nc= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487390; c=relaxed/simple; bh=DWqpV7HV35H5sSmmuvc8SHajZmC2s/uaGRbQuy3Z+NA=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=a2cdFBOU3v0GYL3a4X4wv/Y9kO6syYtPjM+9HTyf/Qt+S+255jQI8i/LddZX4hNPQS1I6Td5CHUJ41yhzuELltVSvhKOBdhmBwEobfiuPZ2BIhqyM7W5HXsHGbhAU/YDBt9vUakpxFUmsuNkga26qnmeIAPXL0K6AsG6wPekLCs= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in; spf=pass smtp.mailfrom=zohomail.in; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b=KxhcUFcb; arc=pass smtp.client-ip=103.117.158.93 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zohomail.in Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b="KxhcUFcb" ARC-Seal: i=1; a=rsa-sha256; t=1784487280; cv=none; d=zohomail.in; s=zohoarc; b=XEiIyIz5mE0Bp26X0wV6jXA/VaVEDP0Vsd8y+wtWApSpa0QnNb1QflaNTLWFhUnGhhhtxzfJCuzo0fiHM9keOIa6/HZdzmITOPROxQGgPC+KqksvGbqDzNUa3EBFq+euvnUukms+YSB2J5UPUhW16rYMvfndG0EowDoBNKl6VhE= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.in; s=zohoarc; t=1784487280; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:MIME-Version:Message-ID:Subject:Subject:To:To:Message-Id:Reply-To; bh=2WFeeZ7KFA8P2LraO+LyQSTzhNZOCYdC5G9o/v637Mc=; b=N/saLD1un8BAWR2AyHmL00UBCNQGgS9/mzsoMKFPls2nqPvcQd+TE66HlrKt8RGiL89qVR9M5f6W4l12uUOC2nvBrukr1+OWVMMtyIRcFb7CXatipTkE8sSZF+7gU4erUjorKMdKAeU6W7rF8Y7xqAnl9zXd6GUjmF2aCGtfdpU= ARC-Authentication-Results: i=1; mx.zohomail.in; dkim=pass header.i=zohomail.in; spf=pass smtp.mailfrom=adi.sharma@zohomail.in; dmarc=pass header.from= DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; t=1784487280; s=zoho; d=zohomail.in; i=adi.sharma@zohomail.in; h=From:From:To:To:Cc:Cc:Subject:Subject:Date:Date:Message-Id:Message-Id:In-Reply-To:MIME-Version:Content-Transfer-Encoding:Reply-To; bh=2WFeeZ7KFA8P2LraO+LyQSTzhNZOCYdC5G9o/v637Mc=; b=KxhcUFcbC6fZlSEYws8W0Ligew3B89Dy34756H9Y+FnvLbM3HTHSHcq86uuSw8oE n0Ztxv6GW1uqGFGw83uoBxi03Di1KvQYhCIeLO3m80kIr5Uz1H/Sh6dVKWRaGyqsAVD lbsjcFg5UpRN2pbF/Is6Fg7wpUjH+nwu2Tyjw80g= Received: by mx.zoho.in with SMTPS id 1784487278994775.8847779555358; Mon, 20 Jul 2026 00:24:38 +0530 (IST) From: Aditya Sharma To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko Cc: David Rientjes , Shakeel Butt , Jonathan Corbet , Shuah Khan , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Kees Cook , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-mm@kvack.org, linux-doc@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org, imbrenda@linux.ibm.com, Aditya Sharma Subject: [RFC PATCH 1/7] mm: add CONFIG_ASYNC_MM_TEARDOWN scaffolding for off-CPU exit teardown Date: Mon, 20 Jul 2026 00:24:03 +0530 Message-Id: <20260719185409.409685-2-adi.sharma@zohomail.in> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260719185409.409685-1-adi.sharma@zohomail.in> References: <20260719185409.409685-1-adi.sharma@zohomail.in> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ZohoMailClient: External Content-Type: text/plain; charset="utf-8" Address-space teardown on process exit runs synchronously in the dying task's context: exit_mm() -> mmput() -> __mmput() -> exit_mmap() walks page tables, updates rmap, frees the RSS and drops file refs, all on the exiting CPU. Add the scaffolding to later move __mmput() off-CPU for large exiting address spaces, with no functional change in this patch: - new config ASYNC_MM_TEARDOWN (bool, depends on MMU, default n) - struct mm_struct::async_reap_node (llist_node), CONFIG-gated, with an explicit include (it was only ever available transitively via spinlock.h) - mmput_exit() declaration, with a static inline fallback that is a plain mmput When the config is disabled mmput_exit() compiles to mmput() with zero delta, so nothing changes until the kthread and dispatch land. Signed-off-by: Aditya Sharma --- include/linux/mm_types.h | 4 ++++ include/linux/sched/mm.h | 6 ++++++ mm/Kconfig | 11 +++++++++++ 3 files changed, 21 insertions(+) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 939b5ea8c..fffa285ef 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -7,6 +7,7 @@ #include #include #include +#include #include #include #include @@ -1367,6 +1368,9 @@ struct mm_struct { atomic_long_t hugetlb_usage; #endif struct work_struct async_put_work; +#ifdef CONFIG_ASYNC_MM_TEARDOWN + struct llist_node async_reap_node; +#endif /* CONFIG_ASYNC_MM_TEARDOWN */ =20 #ifdef CONFIG_IOMMU_MM_DATA struct iommu_mm_data *iommu_mm; diff --git a/include/linux/sched/mm.h b/include/linux/sched/mm.h index 10be8a54b..20aeb1d7e 100644 --- a/include/linux/sched/mm.h +++ b/include/linux/sched/mm.h @@ -147,6 +147,12 @@ extern void mmput(struct mm_struct *); void mmput_async(struct mm_struct *); #endif =20 +#ifdef CONFIG_ASYNC_MM_TEARDOWN +void mmput_exit(struct mm_struct *); +#else +static inline void mmput_exit(struct mm_struct *mm) { mmput(mm); } +#endif + /* Grab a reference to a task's mm, if it is not already going away */ extern struct mm_struct *get_task_mm(struct task_struct *task); /* diff --git a/mm/Kconfig b/mm/Kconfig index c52ab6afc..296b16ad4 100644 --- a/mm/Kconfig +++ b/mm/Kconfig @@ -1500,6 +1500,17 @@ config LAZY_MMU_MODE_KUNIT_TEST =20 If unsure, say N. =20 +config ASYNC_MM_TEARDOWN + bool "Async off-CPU address space teardown on exit" + depends on MMU + default n + help + Defer the address space teardown (exit_mmap()) of large exiting process= es + to a kernel thread instead of running it on the exiting CPU. The feature + is off by default. + + If unsure, say N. + source "mm/damon/Kconfig" =20 endmenu --=20 2.34.1 From nobody Sat Jul 25 03:47:43 2026 Received: from sender-pp-o91.zoho.in (sender-pp-o91.zoho.in [103.117.158.91]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8DF4C20C00C for ; Sun, 19 Jul 2026 18:56:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=pass smtp.client-ip=103.117.158.91 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487390; cv=pass; b=BlBhAqyA5/L0r0LjRRVy1m/ba+MWOhCdkZ15mz/MxuwVBPX4e3V6UIaV/P1gIRuPGYSRl7+KU0q7RB2r/Au7Dzl9/vW+PXPH2jxCyn5Q/Gw3DK2JqGf3472uJE+CsuzVZPg7B2xarQIbky79BUW7l5bYGSBolwAJk642yWXAnyA= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487390; c=relaxed/simple; bh=qoHgkdUKWRzS9NepWwg9eepnHtaReUoqGg549GR2i48=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=SJCt9MM5iLBJQNhfn0doNJ/xuCC5kOZNNSIfIkK/xQucswWdn3d3YgPIWb4ZUEm+LV1iEc+n/yjOlEMZhj4YzJGVQq8jQNakInm41i1RdAsWigCE6hRSURePe/ELUvXyo/b6fJM33q3nQO+Vs0vAlW2PAdQ6zTjEDOheSUfHnxM= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in; spf=pass smtp.mailfrom=zohomail.in; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b=wu2EFQEi; arc=pass smtp.client-ip=103.117.158.91 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zohomail.in Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b="wu2EFQEi" ARC-Seal: i=1; a=rsa-sha256; t=1784487282; cv=none; d=zohomail.in; s=zohoarc; b=dLhc3+CEGqis5T4rqStjmYcL0cArClyv4uSj5dmJskdIDaS2dkoOkbI5gnFiAaPqppOKOGqybHO6sWjHZchibtxWosF6Siu60p3ZpRkEwteIe9BSQsiJyxpZ/5v9WKPW9VYmdgw0Y62QB/wlXHoOW6YxALeAcU0hIL4sMZbd4JM= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.in; s=zohoarc; t=1784487282; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:MIME-Version:Message-ID:Subject:Subject:To:To:Message-Id:Reply-To; bh=xzuHmVh+q9OHnLUm7h9D4LXtKeaYnS+MienvvkAwdTQ=; b=WvfxIsfbFLFFfwp8Ho0Bg96cN63/PAHN/OsqkROcPTKeEL3ujIsYv/gt9Z1PPA5OnyPUbKHdYWLVK5cukKG9kY7HSOLLm1k1dRFe8WRtGMiFpab+Evceij4BJ6ZhWDCHswk3N28rgqKON34CIKlZdsoXVyg0lUOxOa+v8niJUo8= ARC-Authentication-Results: i=1; mx.zohomail.in; dkim=pass header.i=zohomail.in; spf=pass smtp.mailfrom=adi.sharma@zohomail.in; dmarc=pass header.from= DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; t=1784487282; s=zoho; d=zohomail.in; i=adi.sharma@zohomail.in; h=From:From:To:To:Cc:Cc:Subject:Subject:Date:Date:Message-Id:Message-Id:In-Reply-To:MIME-Version:Content-Transfer-Encoding:Reply-To; bh=xzuHmVh+q9OHnLUm7h9D4LXtKeaYnS+MienvvkAwdTQ=; b=wu2EFQEihaI2ze8F242df+yReXM4jf25ryHY4KKKanZgjoR/bZpD3mWFA8kfKM0f xR3+l3h/x55psY5nD93Hy8+QyEfe7mO4ZGfhk6o+TWuGgYFXpn8G0krpgYdiq7qy00E UTwVTpV6p2zbO/LCM5UiqzVh8xk6KepwdUX3yG3k= Received: by mx.zoho.in with SMTPS id 1784487280644201.2764833788417; Mon, 20 Jul 2026 00:24:40 +0530 (IST) From: Aditya Sharma To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko Cc: David Rientjes , Shakeel Butt , Jonathan Corbet , Shuah Khan , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Kees Cook , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-mm@kvack.org, linux-doc@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org, imbrenda@linux.ibm.com, Aditya Sharma Subject: [RFC PATCH 2/7] mm: dispatch __mmput() to an mm_reaper kthread via llist Date: Mon, 20 Jul 2026 00:24:04 +0530 Message-Id: <20260719185409.409685-3-adi.sharma@zohomail.in> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260719185409.409685-1-adi.sharma@zohomail.in> References: <20260719185409.409685-1-adi.sharma@zohomail.in> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ZohoMailClient: External Content-Type: text/plain; charset="utf-8" Introduce a single dedicated, freezable, sleepable kthread (mm_reaper) modeled on oom_reaper, and the machinery to route the deferrable unit of teardown -- __mmput() -- to it through a lock-free llist plus waitqueue. mmput_exit() does llist_add() and wakes the kthread, which drains with llist_del_all() and runs __mmput() per entry with cond_resched(). The drain loop also calls try_to_freeze() after each entry, so a freeze request is honoured mid-batch rather than only at the wait point. Nothing calls mmput_exit() yet: exit_mm() is not routed to it, and there is no eligibility check or gate, so it would queue every final-drop mm unconditionally if invoked. This changes no behavior -- the function is built but unreferenced, the kthread receives no work, and __mmput() runs inline on the exiting CPU exactly as today. The exit_mm() hook and the eligibility/gating policy follow in later patches. llist (rather than oom_reaper's spinlock+list) avoids a global lock on the exit path under many-core parallel exits. llist_for_each_entry_safe() caches the next pointer before __mmput(), so iteration is safe even when the final mmdrop() inside __mmput() frees mm. Lifetime: the mm_count reference that mm_users>0 stands for is released only by that final mmdrop() inside __mmput(). Deferring __mmput() keeps the reference held, so the mm_struct and its embedded async_reap_node stay alive until the kthread runs. No extra mmgrab() is needed, for the same reason mmput_async() needs none. The kthread is started at subsys_initcall and set to nice 19 to minimize competition with other EEVDF tasks. Its affinity is expressed as a preference via kthread_affine_preferred() over the HK_TYPE_KTHREAD housekeeping mask, so it keeps off isolcpus=3D and cpuset-isolated CPUs without the hard binding kthread_bind_mask() would impose -- PF_NO_SETAFFINITY would otherwise leave admins unable to re-affine it. Signed-off-by: Aditya Sharma --- kernel/fork.c | 54 +++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 54 insertions(+) diff --git a/kernel/fork.c b/kernel/fork.c index f0e2e131a..884e4e767 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -112,6 +112,7 @@ #include #include #include +#include =20 #include #include @@ -3407,3 +3408,56 @@ static int __init init_fork_sysctl(void) } =20 subsys_initcall(init_fork_sysctl); + +#ifdef CONFIG_ASYNC_MM_TEARDOWN +static LLIST_HEAD(mm_reaper_list); +static DECLARE_WAIT_QUEUE_HEAD(mm_reaper_wait); + +static void async_mm_teardown_queue(struct mm_struct *mm) +{ + if (llist_add(&mm->async_reap_node, &mm_reaper_list)) + wake_up(&mm_reaper_wait); +} + +static int mm_reaper(void *unused) +{ + set_freezable(); + while (true) { + struct llist_node *batch; + struct mm_struct *mm, *n; + + wait_event_freezable(mm_reaper_wait, !llist_empty(&mm_reaper_list)); + batch =3D llist_del_all(&mm_reaper_list); + llist_for_each_entry_safe(mm, n, batch, async_reap_node) { + __mmput(mm); /* may free mm via mmdrop */ + cond_resched(); + try_to_freeze(); + } + } + return 0; +} + +void mmput_exit(struct mm_struct *mm) +{ + might_sleep(); + if (!atomic_dec_and_test(&mm->mm_users)) + return; + async_mm_teardown_queue(mm); +} + +static int __init mm_reaper_init(void) +{ + struct task_struct *th; + + th =3D kthread_create(mm_reaper, NULL, "mm_reaper"); + if (IS_ERR(th)) { + pr_err("mm_reaper: failed to start kthread: %ld\n", PTR_ERR(th)); + return PTR_ERR(th); + } + set_user_nice(th, 19); /* minimize competition with other fair class task= s */ + kthread_affine_preferred(th, housekeeping_cpumask(HK_TYPE_KTHREAD)); + wake_up_process(th); + return 0; +} +subsys_initcall(mm_reaper_init); +#endif --=20 2.34.1 From nobody Sat Jul 25 03:47:43 2026 Received: from sender-pp-o91.zoho.in (sender-pp-o91.zoho.in [103.117.158.91]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2F4B62EEE9B for ; Sun, 19 Jul 2026 18:56:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=pass smtp.client-ip=103.117.158.91 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487394; cv=pass; b=gKy7zeckMa+lmYxsmrmhAFnIPGyODRVTXzKsfBxHjzdEX60PBLo/J+MWavdOxscj9vnpiLpEQFTdFawgBR1324lzyMJQWifpmlGx2ntvZGC68vKyAI7TbQpunq2nrebhK3SndJgbI9O268swUeaMh2eiiC9cghiErAVJshA+PUw= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487394; c=relaxed/simple; bh=ktiquAK6YZSdvni2NVu6IG0204AB7h8sJ8n1Cxt4zzw=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=HtN18t/+2RjQq99dzd6J8Z8ye0W4VT/fqEbJw1/oyA6EBDYqwteSbiCWg2EhZ9DXqR6NmmX2xC4Burkn3xjIwuLakii/cHcThLK1W7lOOzuOjW45GrScd24y1G7tB8yrEEBBqZRsxrldtdviCFjQLZ9FvVYE+RvelC8BWig57l4= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in; spf=pass smtp.mailfrom=zohomail.in; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b=K1//Kb3N; arc=pass smtp.client-ip=103.117.158.91 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zohomail.in Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b="K1//Kb3N" ARC-Seal: i=1; a=rsa-sha256; t=1784487283; cv=none; d=zohomail.in; s=zohoarc; b=YAoNz96vR2exNu7ryJumatDIvUlceHraaIe2VjWA1qn5AOm0SSW1L6zOEjjn+5x1ziMoy2mNV8xodHKvJP/vXRUsGuPxJ5sf774OYJIPWgGI+1GYyCLELWqVgmO6XdBVlr2LtWOTLIw+qyl/P0IQstdhCyBHj+R0uoYKJobwJvg= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.in; s=zohoarc; t=1784487283; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:MIME-Version:Message-ID:Subject:Subject:To:To:Message-Id:Reply-To; bh=dvpCD42RtrpG9qRaXZM07VjGfK7+/4ydMfeXZD1P/yA=; b=KmgtdkokWN+PeUA55Bn4lNzR81kn0pTxoAzDD+ItjstForqxmVIIFA/JS+BbtZTk7XbwcDfacr4FE+id4q/7Dv1fpsvOPu/Oc1dk7luKpzbCdtN3SN+HALogke5tVKKGBg/2yDxW9HVHdlyl27R5FX48QJEgJ/JzqVfVxrBC0ms= ARC-Authentication-Results: i=1; mx.zohomail.in; dkim=pass header.i=zohomail.in; spf=pass smtp.mailfrom=adi.sharma@zohomail.in; dmarc=pass header.from= DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; t=1784487283; s=zoho; d=zohomail.in; i=adi.sharma@zohomail.in; h=From:From:To:To:Cc:Cc:Subject:Subject:Date:Date:Message-Id:Message-Id:In-Reply-To:MIME-Version:Content-Transfer-Encoding:Reply-To; bh=dvpCD42RtrpG9qRaXZM07VjGfK7+/4ydMfeXZD1P/yA=; b=K1//Kb3NX+WsfCNSoYvhbrab8aFMLVW4+UqIamzfwUvmbRjd2asQh4LC/IZ6O4Su HEgpLiuqlGhoSiOaraF4MebSm+LXFDr1CQzO2aQDtG8X0Hbx7pPMXx/Jj8UvoKWzLY8 monoVJxYxvEDkIWKcAZrcAfG4Dm0tDpXQJ46795w= Received: by mx.zoho.in with SMTPS id 1784487282148903.7584683853768; Mon, 20 Jul 2026 00:24:42 +0530 (IST) From: Aditya Sharma To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko Cc: David Rientjes , Shakeel Butt , Jonathan Corbet , Shuah Khan , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Kees Cook , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-mm@kvack.org, linux-doc@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org, imbrenda@linux.ibm.com, Aditya Sharma Subject: [RFC PATCH 3/7] mm/oom_kill: mark an OOM victim's mm with MMF_OOM_TARGETED Date: Mon, 20 Jul 2026 00:24:05 +0530 Message-Id: <20260719185409.409685-4-adi.sharma@zohomail.in> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260719185409.409685-1-adi.sharma@zohomail.in> References: <20260719185409.409685-1-adi.sharma@zohomail.in> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ZohoMailClient: External Content-Type: text/plain; charset="utf-8" A later part of this series defers the address-space teardown of large exiting processes to a kernel thread (mm_reaper) instead of running it inline on the exiting CPU. That deferral must never be applied to an OOM victim: its memory has to be released promptly to relieve the pressure that triggered the kill, and queuing it behind the low-priority, housekeeping-CPU-confined mm_reaper would work directly against the oom_reaper. The teardown path therefore needs a cheap "is this mm an OOM victim?" test. The eventual consumer runs from mmput_exit() on the exit path, after exit_mm() has already cleared current->mm, so the only handle it has is the mm pointer itself. Existing OOM state does not answer the question well from an mm alone: - signal->oom_mm and TIF_MEMDIE are per-task / per-signal and are set only on the single task passed to mark_oom_victim(). A sibling thread, or a CLONE_VM-sharing process in another thread group (which __oom_kill_process() SIGKILLs but never marks), can be the task that drops the last mm_users reference and runs the teardown. A per-task check misses those. - MMF_OOM_SKIP is mm-scoped but does not mean "OOM victim". exit_mmap() sets it on every exiting mm once the memory is freed, and dup_mmap() sets it on the fork-failure cleanup path, so it produces false positives on ordinary exits. It is also not reliably early: on a real victim it is normally set only once exit_mmap() or oom_reap_task_mm() has already released the memory -- the exception being the mm-pinned-by-init case in __oom_kill_process(), which sets it at kill time with RSS fully intact. Wrong in both directions, so it cannot serve as the predicate this path needs. Add MMF_OOM_TARGETED, an mm-scoped flag, and set it in two places: - in mark_oom_victim(), covering every caller uniformly (this is the only marking for the two task_will_free_mem() fast paths in oom_kill_process() and out_of_memory()); - early in __oom_kill_process(), before any SIGKILL for the kill event is dispatched, to victim's own thread group or to other thread groups sharing the mm. Observing this flag is not load-bearing for correctness. Should it ever be missed, the consequence is bounded: the victim's mm is queued to mm_reaper instead of being torn down inline, i.e. a latency and priority-inversion regression against the oom_reaper -- not corruption, use-after-free or a leak. The mm is still fully torn down and the oom_reaper still reaps the victim independently. The visibility argument below is therefore belt-and-suspenders. Ordering and coverage: this flag is reliably visible to whichever thread ends up dropping the last reference in mmput_exit(), because marking and queuing for async teardown can never overlap in time. Every mark site runs while some task's ->mm still points at this mm -- either tsk is current, marking itself in program order (the task_will_free_mem() fast paths), or tsk is task_lock()ed by the caller (oom_kill_process()'s fast path, and __oom_kill_process() via find_lock_task_mm()) -- while mmput_exit(), the only path that queues for teardown, is reached only after mm_users hits 0, which requires that same task's own exit_mm() (or exec_mmap(), the other path that sheds ->mm) to have already cleared ->mm under that same task_lock(). Program order, or task_lock() release/acquire, therefore orders the flag store before that task's own eventual mmput_exit(). This is stated for userspace tasks: any userspace task's ->mm holds an mm_users reference, whereas kthread_use_mm() borrows an mm under mmgrab() alone -- kthreads are not signal-killable and are never OOM victims, so the borrowed-mm case does not arise here. If a different CLONE_VM sharer in another thread group performs the actual final decrement instead -- a task mark_oom_victim() never marks directly -- the ordering still carries, though not from the flag store itself. mm_flags_set() is set_bit(), a non-value-returning RMW, so it is unordered and contributes nothing on its own. The ordering comes from the marking task's own atomic_dec_and_test(), a value-returning RMW and therefore fully ordered (Documentation/atomic_t.txt: an smp_mb() before and after). Decrements are totally ordered in mm_users' modification order with the zeroing one last, so the general barriers on both sides make the flag store visible to whichever thread performs it. The redundant set in mark_oom_victim() on the __oom_kill_process() path is harmless and keeps the invariant "mark_oom_victim() implies MMF_OOM_TARGETED" independent of future code motion. The flag is intentionally sticky: it is never cleared. That is fine because a marked mm is dying: __oom_kill_process() SIGKILLs every process sharing the victim's mm -- only global init (whose mm also gets MMF_OOM_SKIP set on the spot, keeping it off the async path anyway) and kthreads are exempt from the sweep, and a kthread's kthread_unuse_mm() drop is a plain mmput(), never mmput_exit(). So no task that could reach the async path survives owning a marked mm. Should some future path leave a survivor, the consequence is only conservative: that mm keeps tearing down synchronously. The flag is structurally outside legacy-mask inheritance: mm_init() clears the whole flags bitmap and re-seeds only word 0 from MMF_INIT_LEGACY_MASK, so flags above bit 31 are never copied on fork(). NUM_MM_FLAG_BITS is a fixed 64 on all architectures, so bit 32 exists on 32-bit builds as well. This commit only defines and sets the flag; nothing reads it yet. The write is gated by IS_ENABLED(CONFIG_ASYNC_MM_TEARDOWN) and compiles out entirely on kernels that will not consume it. The reader (async_mm_teardown_eligible()) is added in the next patch, and the deferred teardown is not routed onto the exit path until later in the series, so this change is inert on its own. Note: a bit named MMF_OOM_VICTIM existed until 2022 and was removed as dead state once exit_mmap()'s special-case reap-before-teardown was replaced by mmap_lock/MMF_OOM_SKIP-based concurrency between exit_mmap() and oom_reap_task_mm(). MMF_OOM_TARGETED is a deliberately distinct name for a new consumer -- deferred exit-path teardown -- and does not revive the old exit_mmap() behaviour. Signed-off-by: Aditya Sharma --- include/linux/mm_types.h | 10 +++++++ mm/oom_kill.c | 61 ++++++++++++++++++++++++++++++++++++++++ 2 files changed, 71 insertions(+) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index fffa285ef..3240c029f 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -1989,6 +1989,16 @@ enum { #define MMF_TOPDOWN 31 /* mm searches top down by default */ #define MMF_TOPDOWN_MASK BIT(MMF_TOPDOWN) =20 +/* + * mm is the target of an in-flight OOM kill. Only written under + * CONFIG_ASYNC_MM_TEARDOWN today. Sticky: never cleared -- the mm is + * dying (__oom_kill_process() kills every process sharing it), and a + * hypothetical surviving sharer would merely keep tearing down + * synchronously. Above bit 31, so it is outside MMF_INIT_LEGACY_MASK's + * domain entirely and is never inherited on fork. + */ +#define MMF_OOM_TARGETED 32 + #define MMF_INIT_LEGACY_MASK (MMF_DUMP_FILTER_MASK |\ MMF_DISABLE_THP_MASK | MMF_HAS_MDWE_MASK |\ MMF_VM_MERGE_ANY_MASK | MMF_TOPDOWN_MASK) diff --git a/mm/oom_kill.c b/mm/oom_kill.c index 5f372f6e2..b5490d1b1 100644 --- a/mm/oom_kill.c +++ b/mm/oom_kill.c @@ -754,6 +754,40 @@ static void mark_oom_victim(struct task_struct *tsk) struct mm_struct *mm =3D tsk->mm; =20 WARN_ON(oom_killer_disabled); + + /* + * Mark the mm itself, not just this task/signal, as an OOM target, + * so the deferred-teardown path can test it from the mm alone. Set + * ahead of the TIF_MEMDIE test-and-set below so that every call + * marks the mm, including a repeat mark on a task that is already + * TIF_MEMDIE. + * + * This is reliably visible to whichever thread ends up dropping the + * last reference in mmput_exit(): marking and queuing for async + * teardown can never overlap in time. Every caller reaches this with + * tsk->mm still set -- either tsk is current, marking itself in + * program order, or tsk is task_lock()ed by the caller (the + * task_will_free_mem() fast path in oom_kill_process()) -- while + * mmput_exit(), the only path that queues for teardown, is reached + * only after mm_users hits 0, which requires tsk's own exit_mm() (or + * exec_mmap(), the other path that sheds ->mm) to have already + * cleared ->mm under that same task_lock(). Program order or + * task_lock() release/acquire therefore orders this store before + * tsk's own eventual mmput_exit(). + * + * If a different CLONE_VM sharer performs the actual final decrement + * instead, the ordering does not come from this store: + * mm_flags_set() is set_bit(), a non-value-returning RMW, so it is + * unordered and contributes nothing on its own. It comes from the + * marking task's own atomic_dec_and_test(), a value-returning RMW + * and therefore fully ordered (an smp_mb() before and after). + * Decrements are totally ordered in mm_users' modification order + * with the zeroing one last, so the general barriers on both sides + * make this store visible to whichever thread performs it. + */ + if (IS_ENABLED(CONFIG_ASYNC_MM_TEARDOWN)) + mm_flags_set(MMF_OOM_TARGETED, mm); + /* OOM killer might race with memcg OOM */ if (test_and_set_tsk_thread_flag(tsk, TIF_MEMDIE)) return; @@ -931,6 +965,33 @@ static void __oom_kill_process(struct task_struct *vic= tim, const char *message) mm =3D victim->mm; mmgrab(mm); =20 + /* + * Set this before any SIGKILL for this kill event goes out below, + * while task_lock(victim) (held since find_lock_task_mm() above) is + * still held. This is reliably visible to whichever thread ends up + * running mmput_exit() and dropping the last reference to this mm -- + * victim itself, a sibling, or a CLONE_VM sharer in another thread + * group (which mark_oom_victim() never marks, but which still + * observes this store on the shared mm). Marking and queuing for + * async teardown can never overlap: mmput_exit() is reached only + * after mm_users hits 0, and that requires victim's own exit_mm() + * (or exec_mmap(), the other path that sheds ->mm) to have cleared + * ->mm under this same task_lock() first, which release/acquire + * orders after the store above. + * + * If some other sharer performs the actual final decrement instead, + * the ordering does not come from this store: mm_flags_set() is + * set_bit(), a non-value-returning RMW, so it is unordered and + * contributes nothing on its own. It comes from the marking task's + * own atomic_dec_and_test(), a value-returning RMW and therefore + * fully ordered (an smp_mb() before and after). Decrements are + * totally ordered in mm_users' modification order with the zeroing + * one last, so the general barriers on both sides make this store + * visible to whichever thread performs it. + */ + if (IS_ENABLED(CONFIG_ASYNC_MM_TEARDOWN)) + mm_flags_set(MMF_OOM_TARGETED, mm); + /* Raise event before sending signal: task reaper must see this */ count_vm_event(OOM_KILL); memcg_memory_event_mm(mm, MEMCG_OOM_KILL); --=20 2.34.1 From nobody Sat Jul 25 03:47:43 2026 Received: from sender-pp-o91.zoho.in (sender-pp-o91.zoho.in [103.117.158.91]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 97B8A2E6CCD for ; Sun, 19 Jul 2026 18:56:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=pass smtp.client-ip=103.117.158.91 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487394; cv=pass; b=ANv0FaWWsZbqKOdubXe8W+N9FOZLX/2i8O3ikn+sDyTYJ+ZNOiCQ/yw7gBcYgc/UrN27uy1ebkUdaorWucWP48Uk5nRTGjpMmZh2go5prj98u1hAyJzeReCQTkgU0hUf1dd9GZcGu+Nhovj7Aaxb4bmUv5Bj8lpckRqt5f0mySA= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487394; c=relaxed/simple; bh=vx3OPMUUoe68Q/4cGI8bVpPqsREryBia9lB9xO54sCU=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=JpEWrexFMo4dzTSl008k7bjrPQTmlm4vXQdCl4aF22U9YtdlEXjYmihG9MSVrbDTIq4ClPp91VCyzNxKqQWUbVRDyIorpsSt5U5lliQPdUiZU+i3+Dzp8L2C7M9RBqV/YwwpGs7r3x/JJGmZmDXHTvxIecmfNs7Nd/2Zzj8d0lM= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in; spf=pass smtp.mailfrom=zohomail.in; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b=ti12zkcG; arc=pass smtp.client-ip=103.117.158.91 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zohomail.in Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b="ti12zkcG" ARC-Seal: i=1; a=rsa-sha256; t=1784487285; cv=none; d=zohomail.in; s=zohoarc; b=QGfoxSY0NHJRFFFk4MJ9D1OB/xRzPDpysy3mZGzokE97/lbhUbLF77gCFtqzOuZSBiXQteNIsxbM3iIM5hrSSi0Yp0ymhI9sv4scTaXI8I1vvW0kv0BZkRwg4ZAxVWVMa7AOccAPmy6z8T1FRCWBPBkHYOkbbRpxdd8LFO0Lvo4= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.in; s=zohoarc; t=1784487285; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:MIME-Version:Message-ID:Subject:Subject:To:To:Message-Id:Reply-To; bh=FgZdTLPY7PuLb4yQ+exzFKyeCVLnL9WHdNDZ74r1sCI=; b=A6yGeJKs8Gq7QoiYlYdtWwmum4OJEXYscvDJLkgU6br3+fb6XS2Tkkwc4aJWn+7pUvWuleZp+1nXKqn1ufZfHfncrHmToUEqJam4vyKAKxMIFR57LVa/ObVcYQVHvrzrHFzIa7JkfqiPRWfPR+5LeK7NafClE8e5NxM7TB14thE= ARC-Authentication-Results: i=1; mx.zohomail.in; dkim=pass header.i=zohomail.in; spf=pass smtp.mailfrom=adi.sharma@zohomail.in; dmarc=pass header.from= DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; t=1784487285; s=zoho; d=zohomail.in; i=adi.sharma@zohomail.in; h=From:From:To:To:Cc:Cc:Subject:Subject:Date:Date:Message-Id:Message-Id:In-Reply-To:MIME-Version:Content-Transfer-Encoding:Reply-To; bh=FgZdTLPY7PuLb4yQ+exzFKyeCVLnL9WHdNDZ74r1sCI=; b=ti12zkcGnBfLTBw9fHFNt9PB2IhRRcjLzCwdqRQEYBhvAlUGAO71teK3pE48YoGT qIfK5gTDP/uPH/+O9mqXHPD0yPt7ZL4KfWQ/syFSmtMUfHc5ZEWp8W5hHTE/DU07xAl ybvePEbyoPFa6kwl3iFBOoTaVc6PbMNrHZccOmy0= Received: by mx.zoho.in with SMTPS id 1784487283605355.28475880973383; Mon, 20 Jul 2026 00:24:43 +0530 (IST) From: Aditya Sharma To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko Cc: David Rientjes , Shakeel Butt , Jonathan Corbet , Shuah Khan , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Kees Cook , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-mm@kvack.org, linux-doc@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org, imbrenda@linux.ibm.com, Aditya Sharma Subject: [RFC PATCH 4/7] mm: gate async teardown on RSS with pending-pages backpressure Date: Mon, 20 Jul 2026 00:24:06 +0530 Message-Id: <20260719185409.409685-5-adi.sharma@zohomail.in> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260719185409.409685-1-adi.sharma@zohomail.in> References: <20260719185409.409685-1-adi.sharma@zohomail.in> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ZohoMailClient: External Content-Type: text/plain; charset="utf-8" Deferring every exit would churn small processes for no benefit and let unreaped memory grow without bound. Split the defer decision into a pure predicate and a budget reservation: - async_mm_teardown_eligible() (side-effect free): defer only when RSS >=3D async_mm_teardown_thresh_pages and both MMF_OOM_SKIP and MMF_OOM_TARGETED are clear: * MMF_OOM_SKIP -- the mm is hidden from the OOM killer and reaper. Usually that means it has already been reaped or run through exit_mmap() and has no RSS left worth deferring; in the mm-pinned-by-init case (__oom_kill_process()) the flag is set at kill time with RSS intact, and keeping such an mm synchronous is equally what we want; * MMF_OOM_TARGETED -- the mm is the target of an in-flight OOM kill (added in the preceding patch). Such a victim is torn down synchronously so its memory is released promptly rather than queued behind the low-priority, housekeeping-confined mm_reaper, which would work against the oom_reaper during the exact pressure that triggered the kill. The flag is set while some task still holds a live reference to the mm, and this check is reached only after mm_users has dropped to 0 -- which requires that task's own exit_mm() (or exec_mmap(), the other path that sheds ->mm) to have already cleared ->mm under that same task_lock() -- so marking and checking can never overlap in time. Whichever thread performs the final decrement is therefore guaranteed to observe the flag already set, per the preceding patch. - async_mm_teardown_reserve(): charge RSS against a pending-pages budget with atomic_long_add_return(); if the result exceeds async_mm_teardown_max_pending_pages it subtracts back and returns false, so the caller falls back to synchronous __mmput(). mmput_exit() reads RSS once and checks eligibility before reservation, so the budget is only touched for processes that pass the cheap checks. On success it stores the charged amount in mm->async_reap_node's companion field async_reap_rss before enqueueing, and the reaper releases exactly that stored charge after __mmput(). The charge cannot be re-derived at reap time: RSS keeps shrinking after mm_users reaches zero -- rmap-based reclaim can still swap out anon pages and truncation can still unmap file pages, neither of which requires mm_users -- and get_mm_rss() is an approximate per-CPU counter read besides. Two independent reads at enqueue and reap would systematically under-release, drifting the pending counter upward until the cap permanently forces synchronous fallback. An eligible mm would otherwise have its entire __mmput() -- including exit_aio() -- deferred onto the single mm_reaper kthread. exit_aio() can block indefinitely waiting for in-flight AIO to drain, so mmput_exit() calls it eagerly, on the exiting task's own context, before reservation: a stuck AIO context then only stalls this task, same as a synchronous mmput() would, instead of blocking mm_reaper and stranding every other mm already queued behind it. exit_aio() is idempotent -- it is a no-op once mm->ioctx_table has been torn down -- so __mmput() calling it again unconditionally is harmless, whether that second call lands on the reaper or -- when the reservation fails -- back on the exiting task itself. Because reservation is add-then-check rather than read-then-add, committed pending never stays above the cap: an exit that would exceed it tears down synchronously instead. Under concurrent exits a failing reservation can transiently inflate the counter between its add and roll-back, which may cause another exiter to fall back to sync conservatively -- it can over-reject but never over-admit. The cap is a coarse safety bound on unreaped memory, not a reclaim mechanism: pages queued here are invisible to reclaim until the reaper frees them. Integrating the queue with reclaim (unmap-first draining, a shrinker) is future work; see the cover letter. The RSS threshold defaults to a fixed 64MB at init; the pending cap defaults to totalram_pages() / 4 as a placeholder pending tuning. Both are static module variables here and not yet user-tunable. The static key that gates the whole path, and sysctl registration, follow in the next patch. This patch is dormant: nothing calls mmput_exit() yet (exit_mm() is routed to it later in the series), so no teardown is deferred. Signed-off-by: Aditya Sharma --- include/linux/mm_types.h | 1 + kernel/fork.c | 79 +++++++++++++++++++++++++++++++++++++++- 2 files changed, 79 insertions(+), 1 deletion(-) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 3240c029f..ee6c0ac84 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -1370,6 +1370,7 @@ struct mm_struct { struct work_struct async_put_work; #ifdef CONFIG_ASYNC_MM_TEARDOWN struct llist_node async_reap_node; + unsigned long async_reap_rss; #endif /* CONFIG_ASYNC_MM_TEARDOWN */ =20 #ifdef CONFIG_IOMMU_MM_DATA diff --git a/kernel/fork.c b/kernel/fork.c index 884e4e767..d45418e66 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -113,6 +113,7 @@ #include #include #include +#include =20 #include #include @@ -3412,6 +3413,48 @@ subsys_initcall(init_fork_sysctl); #ifdef CONFIG_ASYNC_MM_TEARDOWN static LLIST_HEAD(mm_reaper_list); static DECLARE_WAIT_QUEUE_HEAD(mm_reaper_wait); +static atomic_long_t mm_reaper_pending_pages; +static unsigned long sysctl_async_mm_teardown_thresh_pages; +static unsigned long sysctl_async_mm_teardown_max_pending_pages; + +static bool async_mm_teardown_eligible(struct mm_struct *mm, unsigned long= rss) +{ + if (rss < READ_ONCE(sysctl_async_mm_teardown_thresh_pages)) + return false; + if (mm_flags_test(MMF_OOM_SKIP, mm)) /* reaped, or hidden from the reaper= */ + return false; + /* + * MMF_OOM_TARGETED is set by the OOM killer while some task still + * holds a live reference to this mm (see mark_oom_victim() and + * __oom_kill_process()), and this eligibility check is reached only + * from mmput_exit(), after mm_users has dropped to 0 -- so marking + * and this check can never overlap in time. Whichever thread + * performs the final decrement is therefore guaranteed to observe + * the flag already set, regardless of whether that thread was + * itself ever passed to mark_oom_victim() (a CLONE_VM-sharing + * process in another thread group never is, but still observes + * this flag on the shared mm). + */ + if (mm_flags_test(MMF_OOM_TARGETED, mm)) + return false; + return true; +} + +/* + * Charge rss against the pending-teardown budget. Returns true if it fits + * under the cap, in which case the caller enqueues and the reaper releases + * the same rss after __mmput(). On overshoot nothing stays charged and t= he + * caller must tear down synchronously. + */ +static bool async_mm_teardown_reserve(unsigned long rss) +{ + if (atomic_long_add_return(rss, &mm_reaper_pending_pages) > + READ_ONCE(sysctl_async_mm_teardown_max_pending_pages)) { + atomic_long_sub(rss, &mm_reaper_pending_pages); + return false; + } + return true; +} =20 static void async_mm_teardown_queue(struct mm_struct *mm) { @@ -3429,7 +3472,10 @@ static int mm_reaper(void *unused) wait_event_freezable(mm_reaper_wait, !llist_empty(&mm_reaper_list)); batch =3D llist_del_all(&mm_reaper_list); llist_for_each_entry_safe(mm, n, batch, async_reap_node) { + unsigned long pages =3D mm->async_reap_rss; + __mmput(mm); /* may free mm via mmdrop */ + atomic_long_sub(pages, &mm_reaper_pending_pages); cond_resched(); try_to_freeze(); } @@ -3439,10 +3485,38 @@ static int mm_reaper(void *unused) =20 void mmput_exit(struct mm_struct *mm) { + unsigned long rss; + might_sleep(); + /* + * Fully ordered RMW: combined with exit_mm() (or exec_mmap(), the + * other path that sheds ->mm) clearing ->mm under task_lock() before + * this call, it guarantees MMF_OOM_TARGETED set at any mark site is + * visible here -- see mark_oom_victim(). Any new caller of + * mmput_exit() outside exit_mm() must re-verify that argument. + */ if (!atomic_dec_and_test(&mm->mm_users)) return; - async_mm_teardown_queue(mm); + + rss =3D get_mm_rss(mm); + + if (async_mm_teardown_eligible(mm, rss)) { + /* + * exit_aio() can block indefinitely. Run it here so a stuck + * AIO only hangs this task (same as how it would happen in + * case of synchronous mmput()) instead of stranding every + * teardown queued behind this mm + */ + exit_aio(mm); + + if (async_mm_teardown_reserve(rss)) { + mm->async_reap_rss =3D rss; + async_mm_teardown_queue(mm); + return; + } + } + + __mmput(mm); } =20 static int __init mm_reaper_init(void) @@ -3456,7 +3530,10 @@ static int __init mm_reaper_init(void) } set_user_nice(th, 19); /* minimize competition with other fair class task= s */ kthread_affine_preferred(th, housekeeping_cpumask(HK_TYPE_KTHREAD)); + sysctl_async_mm_teardown_thresh_pages =3D SZ_64M >> PAGE_SHIFT; + sysctl_async_mm_teardown_max_pending_pages =3D totalram_pages() / 4; /* = TODO: placeholder */ wake_up_process(th); + return 0; } subsys_initcall(mm_reaper_init); --=20 2.34.1 From nobody Sat Jul 25 03:47:43 2026 Received: from sender-pp-o91.zoho.in (sender-pp-o91.zoho.in [103.117.158.91]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 07C55282F29 for ; Sun, 19 Jul 2026 18:56:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=pass smtp.client-ip=103.117.158.91 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487394; cv=pass; b=oFlb5JxggPdXFRYuARGR7US/0Wg11AJVwXEQ1lRy7F4jipBc/lpO5FXe8z1DW7aC9mRObSBRXb2asXAZ8GQUSH5a75agtz/c7nXZQvUnRuKpbJG2/9hRAOrsz8e5ICAT/0v+KQImUBdwhdACjTKLuZ+UkUEekmsUsjG2AfvKKyk= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487394; c=relaxed/simple; bh=NXR8Pc4xTaselcNDKYjw6CmxtOIzVpM9iH+zR8NlLpQ=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=p2agJwockJXRDb0Qs9sQo3q028QHDeVrPxwlVNi0mPbHuC8AzmzxDHVpXLbKF6uiS0uE4bgdK80lrhaEA0X5hDvHfbkmL4X3zl8STf1LilAidZCeq/F1YPPsb1tFYrYDLHeH/cGDMG24s6X7luPJ3nCyNWyQXeaa/EBs+YQkhlw= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in; spf=pass smtp.mailfrom=zohomail.in; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b=C05pp2fR; arc=pass smtp.client-ip=103.117.158.91 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zohomail.in Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b="C05pp2fR" ARC-Seal: i=1; a=rsa-sha256; t=1784487286; cv=none; d=zohomail.in; s=zohoarc; b=eoZkNbwBdZ3sAgobW/+MD+mXYVUfuXZoh0PW0dkgph8ontDVOymnyy1uza09p8k77a70nNrVW96Qn+v20s7gIEdyp2cENVyM1hAe5gVkW1GJ8Wkf7+xzOjnVW756s8T8BDOt102znL4MIw1E64+K/vwgxKTx6+pp2Dg5zGYvwAU= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.in; s=zohoarc; t=1784487286; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:MIME-Version:Message-ID:Subject:Subject:To:To:Message-Id:Reply-To; bh=qa4YjEBQvZIx5KKQpuyATR3R1yqVaUF8NDQIptYbS9w=; b=RumEku92gkeZHIU445cPHdCm6F2NbcdDI8sOZqYysLMAXLb1ia2xJVM1sWdPbMOZCavKkoFq4gHjFPycY/l2woYI0GaTPtS8798o7u69arMuhea2EZK4BqlA6OGC0kkVfg7mmd07XdZxF2OvqnuPD71ytNVE+Whco2F8m61JEoc= ARC-Authentication-Results: i=1; mx.zohomail.in; dkim=pass header.i=zohomail.in; spf=pass smtp.mailfrom=adi.sharma@zohomail.in; dmarc=pass header.from= DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; t=1784487286; s=zoho; d=zohomail.in; i=adi.sharma@zohomail.in; h=From:From:To:To:Cc:Cc:Subject:Subject:Date:Date:Message-Id:Message-Id:In-Reply-To:MIME-Version:Content-Transfer-Encoding:Reply-To; bh=qa4YjEBQvZIx5KKQpuyATR3R1yqVaUF8NDQIptYbS9w=; b=C05pp2fRMkzVr8EPumtScLZSIAxiwWdUw5nGAZ6j0AL9hOYugDThqSgbodVFyuVJ A1Y2esHcmBZrPnoeJjnr1gPKNpdfV00uui6p8qfuO2jb3SBK4h7a0P2GXYiStBTMM2e p5ydWfmf3m05+2byNLHwxHVe1msEDc+gV6JSpuXQ= Received: by mx.zoho.in with SMTPS id 1784487285082857.0950681993331; Mon, 20 Jul 2026 00:24:45 +0530 (IST) From: Aditya Sharma To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko Cc: David Rientjes , Shakeel Butt , Jonathan Corbet , Shuah Khan , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Kees Cook , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-mm@kvack.org, linux-doc@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org, imbrenda@linux.ibm.com, Aditya Sharma Subject: [RFC PATCH 5/7] mm: add runtime toggle and sysctls for async teardown Date: Mon, 20 Jul 2026 00:24:07 +0530 Message-Id: <20260719185409.409685-6-adi.sharma@zohomail.in> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260719185409.409685-1-adi.sharma@zohomail.in> References: <20260719185409.409685-1-adi.sharma@zohomail.in> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ZohoMailClient: External Content-Type: text/plain; charset="utf-8" Add the userspace interface that arms and tunes off-CPU exit teardown, and the static key that gates it. Introduce async_mm_teardown_key (static key, default false) and gate the deferral path in mmput_exit() on it via static_branch_unlikely(); when off, mmput_exit() runs __mmput() inline. The get_mm_rss() read moves inside the static_branch_unlikely() block, alongside eligibility and exit_aio(), so the default (key off) path costs nothing beyond the branch itself -- no RSS scan for every exiting process on kernels or boots that never enable the feature. Under CONFIG_SYSCTL, register three knobs under vm. with register_sysctl_init(): - async_mm_teardown: write 1/0 to enable/disable. A custom handler flips the static key, serialized by a mutex so the key state always matches the last value written. - async_mm_teardown_thresh_pages: RSS threshold (default ~64MB). - async_mm_teardown_max_pending_pages: backpressure cap. The threshold defaults are set in the initcall outside any CONFIG_SYSCTL guard, since eligible()/reserve() read them whenever the key is on; only the toggle variable, its lock, the handler and the table depend on CONFIG_SYSCTL. With CONFIG_SYSCTL=3Dn there is no path to enable the key, so the feature stays off and the thresholds are inert. This patch is still dormant: nothing calls mmput_exit() yet, so flipping async_mm_teardown on has no observable effect. exit_mm() is routed through mmput_exit() later in the series; only then does enabling the key defer teardown. The toggle and gate are introduced here so that, once routing lands, the feature defaults off and an explicit opt-in is required. Signed-off-by: Aditya Sharma --- Documentation/admin-guide/sysctl/vm.rst | 38 +++++++++++ kernel/fork.c | 90 ++++++++++++++++++++----- mm/Kconfig | 3 +- 3 files changed, 114 insertions(+), 17 deletions(-) diff --git a/Documentation/admin-guide/sysctl/vm.rst b/Documentation/admin-= guide/sysctl/vm.rst index 5b318d17a..ccda6c75f 100644 --- a/Documentation/admin-guide/sysctl/vm.rst +++ b/Documentation/admin-guide/sysctl/vm.rst @@ -22,6 +22,9 @@ the writeout of dirty data to disk. Currently, these files are in /proc/sys/vm: =20 - admin_reserve_kbytes +- async_mm_teardown +- async_mm_teardown_max_pending_pages +- async_mm_teardown_thresh_pages - compact_memory - compaction_proactiveness - compact_unevictable_allowed @@ -108,6 +111,41 @@ On x86_64 this is about 128MB. Changing this takes effect whenever an application requests memory. =20 =20 +async_mm_teardown +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Available only when CONFIG_ASYNC_MM_TEARDOWN is set. Controls whether the +address-space teardown (exit_mmap()) of large exiting processes is deferred +to the mm_reaper kernel thread instead of running on the exiting CPU. + +Writing 1 enables deferral, 0 disables it. Disabling makes mmput_exit() te= ar +down inline again; it does not drain mms already queued. Defaults to 0. + +Deferred teardown consumes CPU in the mm_reaper kernel thread, which runs = in +the root cgroup: the teardown work of a task in a CPU-limited cgroup is not +charged to that cgroup, similar to kswapd and the oom_reaper. Consider this +before enabling the feature on systems that rely on strict per-cgroup CPU +accounting. + + +async_mm_teardown_max_pending_pages +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Available only when CONFIG_ASYNC_MM_TEARDOWN is set. Backpressure cap, in +pages, on the total RSS of mms queued for the mm_reaper but not yet torn +down. An exit that would push the total over this cap tears down +synchronously instead of queuing. Defaults to totalram_pages() / 4. + + +async_mm_teardown_thresh_pages +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D + +Available only when CONFIG_ASYNC_MM_TEARDOWN is set. Minimum RSS, in pages, +for an exiting mm to be eligible for deferral to the mm_reaper. Exiting mms +below this threshold are always torn down synchronously. Defaults to 64MB +worth of pages. + + compact_memory =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D =20 diff --git a/kernel/fork.c b/kernel/fork.c index d45418e66..0976770db 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -3414,8 +3414,13 @@ subsys_initcall(init_fork_sysctl); static LLIST_HEAD(mm_reaper_list); static DECLARE_WAIT_QUEUE_HEAD(mm_reaper_wait); static atomic_long_t mm_reaper_pending_pages; +static DEFINE_STATIC_KEY_FALSE(async_mm_teardown_key); static unsigned long sysctl_async_mm_teardown_thresh_pages; static unsigned long sysctl_async_mm_teardown_max_pending_pages; +#ifdef CONFIG_SYSCTL +static u8 sysctl_async_mm_teardown_enabled; +static DEFINE_MUTEX(async_mm_teardown_lock); +#endif =20 static bool async_mm_teardown_eligible(struct mm_struct *mm, unsigned long= rss) { @@ -3485,8 +3490,6 @@ static int mm_reaper(void *unused) =20 void mmput_exit(struct mm_struct *mm) { - unsigned long rss; - might_sleep(); /* * Fully ordered RMW: combined with exit_mm() (or exec_mmap(), the @@ -3498,27 +3501,80 @@ void mmput_exit(struct mm_struct *mm) if (!atomic_dec_and_test(&mm->mm_users)) return; =20 - rss =3D get_mm_rss(mm); + if (static_branch_unlikely(&async_mm_teardown_key)) { + unsigned long rss =3D get_mm_rss(mm); =20 - if (async_mm_teardown_eligible(mm, rss)) { - /* - * exit_aio() can block indefinitely. Run it here so a stuck - * AIO only hangs this task (same as how it would happen in - * case of synchronous mmput()) instead of stranding every - * teardown queued behind this mm - */ - exit_aio(mm); + if (async_mm_teardown_eligible(mm, rss)) { + /* + * exit_aio() can block indefinitely. Run it here so a stuck + * AIO only hangs this task (same as how it would happen in + * case of synchronous mmput()) instead of stranding every + * teardown queued behind this mm + */ + exit_aio(mm); =20 - if (async_mm_teardown_reserve(rss)) { - mm->async_reap_rss =3D rss; - async_mm_teardown_queue(mm); - return; + if (async_mm_teardown_reserve(rss)) { + mm->async_reap_rss =3D rss; + async_mm_teardown_queue(mm); + return; + } } } =20 __mmput(mm); } =20 +#ifdef CONFIG_SYSCTL +static int async_mm_teardown_enabled_handler(const struct ctl_table *table, + int write, void *buffer, + size_t *lenp, loff_t *ppos) +{ + int ret; + + /* + * Serialize concurrent writers so the static key state always matches + * the last value written to the variable. + */ + guard(mutex)(&async_mm_teardown_lock); + + ret =3D proc_dou8vec_minmax(table, write, buffer, lenp, ppos); + if (ret || !write) + return ret; + + if (sysctl_async_mm_teardown_enabled) + static_branch_enable(&async_mm_teardown_key); + else + static_branch_disable(&async_mm_teardown_key); + return 0; +} + +static const struct ctl_table async_mm_teardown_table[] =3D { + { + .procname =3D "async_mm_teardown", + .data =3D &sysctl_async_mm_teardown_enabled, + .maxlen =3D sizeof(u8), + .mode =3D 0644, + .proc_handler =3D async_mm_teardown_enabled_handler, + .extra1 =3D SYSCTL_ZERO, + .extra2 =3D SYSCTL_ONE, + }, + { + .procname =3D "async_mm_teardown_thresh_pages", + .data =3D &sysctl_async_mm_teardown_thresh_pages, + .maxlen =3D sizeof(unsigned long), + .mode =3D 0644, + .proc_handler =3D proc_doulongvec_minmax, + }, + { + .procname =3D "async_mm_teardown_max_pending_pages", + .data =3D &sysctl_async_mm_teardown_max_pending_pages, + .maxlen =3D sizeof(unsigned long), + .mode =3D 0644, + .proc_handler =3D proc_doulongvec_minmax, + }, +}; +#endif /* CONFIG_SYSCTL */ + static int __init mm_reaper_init(void) { struct task_struct *th; @@ -3533,7 +3589,9 @@ static int __init mm_reaper_init(void) sysctl_async_mm_teardown_thresh_pages =3D SZ_64M >> PAGE_SHIFT; sysctl_async_mm_teardown_max_pending_pages =3D totalram_pages() / 4; /* = TODO: placeholder */ wake_up_process(th); - +#ifdef CONFIG_SYSCTL + register_sysctl_init("vm", async_mm_teardown_table); +#endif return 0; } subsys_initcall(mm_reaper_init); diff --git a/mm/Kconfig b/mm/Kconfig index 296b16ad4..a9d925406 100644 --- a/mm/Kconfig +++ b/mm/Kconfig @@ -1507,7 +1507,8 @@ config ASYNC_MM_TEARDOWN help Defer the address space teardown (exit_mmap()) of large exiting process= es to a kernel thread instead of running it on the exiting CPU. The feature - is off by default. + is off by default and is enabled at runtime via the vm.async_mm_teardown + sysctl. =20 If unsure, say N. =20 --=20 2.34.1 From nobody Sat Jul 25 03:47:43 2026 Received: from sender-pp-o91.zoho.in (sender-pp-o91.zoho.in [103.117.158.91]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A01EE2571B8 for ; Sun, 19 Jul 2026 18:56:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=pass smtp.client-ip=103.117.158.91 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487390; cv=pass; b=fm0uzd5onlfmxgxOjzJal0osSKwy1kS2l+M3Nhr1RTVmPJ1RUjT6Uw1L+aYwm0J9lBeEqtvuVpFGfX/DRa9MlyzLXmuZaId3Rm3X27OwsmbfDCkeHVn/4I/FJD/OrqcuhKpjrlfKmltR3/U4+GEEXICN2fljI5ZDSl7dB4OEt5U= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487390; c=relaxed/simple; bh=r5G2wi50yD8h9o+0dBa0rO1l4YvK8s1tJBUzz2j94yk=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=pfUKO1Ww/tx824LUijOLGqQzZWDa9FWhyyltVxxr/R41syifvIL8ikTOHw+DoGlb8KQNzoAYAsQOvGuftBkuySdtBYaHIKoXRWnyBLAdUYZLBZfH5niKKWiSgDbAlrwTjhnBzpdQi/Qsr1TlSRmSHN1Pk8uccO4vzXuT5kbuHWk= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in; spf=pass smtp.mailfrom=zohomail.in; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b=YnW8pALY; arc=pass smtp.client-ip=103.117.158.91 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zohomail.in Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b="YnW8pALY" ARC-Seal: i=1; a=rsa-sha256; t=1784487288; cv=none; d=zohomail.in; s=zohoarc; b=Hg3KfhqrSbZ77reLafznqJhwV1LJSzO5mmcoXRp+JmqH1LlItiJruE8VW/W/1aDyRAoqh32F1e48HM1XC6ohw4TLFQjiziV78DYDC+bnMar0lNCcJK2p0MPrTBWUOPNklKb3SyVaCTNCvqhS7AxZdMjJFes5otKZ64i9Ee2gl2w= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.in; s=zohoarc; t=1784487288; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:MIME-Version:Message-ID:Subject:Subject:To:To:Message-Id:Reply-To; bh=oj9fDzX1wMrErAiwiqiXxnM4cGXgD5lIVRx7urULVtQ=; b=G/beN0jv6Q8FIOuAqbFFROOF1nkiv8xTb2iZishK3qidN0iSg/pHWKgLGCi2cSPRRcx/PoGu+D/Hc+EUabM3CYQA6LTskvx/cNIZAN2bvTTbmphlPzxx2NKKkjHZqVFiScJ0IqHidhneoVhmN3C2H+bfSgTH8Gm7dcMh+8nUiwo= ARC-Authentication-Results: i=1; mx.zohomail.in; dkim=pass header.i=zohomail.in; spf=pass smtp.mailfrom=adi.sharma@zohomail.in; dmarc=pass header.from= DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; t=1784487288; s=zoho; d=zohomail.in; i=adi.sharma@zohomail.in; h=From:From:To:To:Cc:Cc:Subject:Subject:Date:Date:Message-Id:Message-Id:In-Reply-To:MIME-Version:Content-Transfer-Encoding:Reply-To; bh=oj9fDzX1wMrErAiwiqiXxnM4cGXgD5lIVRx7urULVtQ=; b=YnW8pALYKS+7BrhWfZSMS4owuUZYNoMgUCQiPgBceYMAA9mcp7oH2vkTqn7Azxuq HATY+/+WkSG7WqveYQ/CJRAqclFelyYI+bNA/Ud7IurKaWNtDInCNL26LMPBf8xJEIN hKMVys2fVq4gSF0M638PiWT2OJ7rM3PdtgMuUWV8= Received: by mx.zoho.in with SMTPS id 1784487286510813.0873877162367; Mon, 20 Jul 2026 00:24:46 +0530 (IST) From: Aditya Sharma To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko Cc: David Rientjes , Shakeel Butt , Jonathan Corbet , Shuah Khan , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Kees Cook , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-mm@kvack.org, linux-doc@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org, imbrenda@linux.ibm.com, Aditya Sharma Subject: [RFC PATCH 6/7] mm: add tracepoints and vmstat counters for async teardown Date: Mon, 20 Jul 2026 00:24:08 +0530 Message-Id: <20260719185409.409685-7-adi.sharma@zohomail.in> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260719185409.409685-1-adi.sharma@zohomail.in> References: <20260719185409.409685-1-adi.sharma@zohomail.in> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ZohoMailClient: External Content-Type: text/plain; charset="utf-8" Add observability for the deferred teardown path: - vmstat (CONFIG_ASYNC_MM_TEARDOWN, in /proc/vmstat): async_mm_teardown_queued - deferred to the mm_reaper kthread async_mm_teardown_sync - torn down inline on the exit path: feature off, mm below the RSS threshold, MMF_OOM_SKIP or MMF_OOM_TARGETED set, or rejected by backpressure. Counts only teardowns reaching mmput_exit(); final drops via mmput()/mmput_async() elsewhere in the kernel are inline but not counted here. async_mm_teardown_rejected - subset of the above that was eligible but lost to async_mm_teardown_reserve() (counted for both this and sync, so sync stays the total for the exit path) - tracepoints in a new include/trace/events/mm_reaper.h: mm_async_teardown_queue at enqueue, and mm_async_teardown_reap when the reaper picks an mm up, just before __mmput(). queue carries the mm, its RSS, and the exiting task's pid/comm, captured while current is still that task, plus the node of the exiting task. reap carries the mm, the charged RSS (mm->async_reap_rss, the amount mmput_exit() reserved and what the pending-pages accounting releases), a fresh live RSS read taken just before __mmput(), and the node of the reaper. The gap between charged and live measures how much reclaim ate out of the queue before the reaper got to it. reap is emitted before __mmput() because the final mmdrop() inside __mmput() may free the mm. Signed-off-by: Aditya Sharma --- include/linux/vm_event_item.h | 5 +++ include/trace/events/mm_reaper.h | 73 ++++++++++++++++++++++++++++++++ kernel/fork.c | 14 +++++- mm/vmstat.c | 5 +++ 4 files changed, 95 insertions(+), 2 deletions(-) create mode 100644 include/trace/events/mm_reaper.h diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h index 2628ccda0..de4dc20b7 100644 --- a/include/linux/vm_event_item.h +++ b/include/linux/vm_event_item.h @@ -179,6 +179,11 @@ enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT, NRSWPIN, NRSWPOUT, #endif /* CONFIG_SWAP */ +#ifdef CONFIG_ASYNC_MM_TEARDOWN + ASYNC_MM_TEARDOWN_QUEUED, + ASYNC_MM_TEARDOWN_SYNC, + ASYNC_MM_TEARDOWN_REJECTED, +#endif NR_VM_EVENT_ITEMS }; =20 diff --git a/include/trace/events/mm_reaper.h b/include/trace/events/mm_rea= per.h new file mode 100644 index 000000000..315bced7c --- /dev/null +++ b/include/trace/events/mm_reaper.h @@ -0,0 +1,73 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#undef TRACE_SYSTEM +#define TRACE_SYSTEM mm_reaper + +#if !defined(_TRACE_MM_REAPER_H) || defined(TRACE_HEADER_MULTI_READ) +#define _TRACE_MM_REAPER_H + +#include +#include +#include + +TRACE_EVENT(mm_async_teardown_queue, + + TP_PROTO(struct mm_struct *mm, unsigned long rss), + + TP_ARGS(mm, rss), + + TP_STRUCT__entry( + __field(struct mm_struct *, mm) + __field(int, pid) + __array(char, comm, TASK_COMM_LEN) + __field(unsigned long, rss) + __field(int, node) + ), + + TP_fast_assign( + __entry->mm =3D mm; + __entry->pid =3D current->pid; + memcpy(__entry->comm, current->comm, TASK_COMM_LEN); + __entry->rss =3D rss; + __entry->node =3D numa_node_id(); + ), + + TP_printk("mm=3D%p pid=3D%d comm=3D%s rss=3D%lukB node=3D%d", + __entry->mm, + __entry->pid, + __entry->comm, + __entry->rss << (PAGE_SHIFT - 10), + __entry->node + ) +); + +TRACE_EVENT(mm_async_teardown_reap, + + TP_PROTO(struct mm_struct *mm, unsigned long charged_rss, unsigned long l= ive_rss), + + TP_ARGS(mm, charged_rss, live_rss), + + TP_STRUCT__entry( + __field(struct mm_struct *, mm) + __field(unsigned long, charged_rss) + __field(unsigned long, live_rss) + __field(int, node) + ), + + TP_fast_assign( + __entry->mm =3D mm; + __entry->charged_rss =3D charged_rss; + __entry->live_rss =3D live_rss; + __entry->node =3D numa_node_id(); + ), + + TP_printk("mm=3D%p charged_rss=3D%lukB live_rss=3D%lukB node=3D%d", + __entry->mm, + __entry->charged_rss << (PAGE_SHIFT - 10), + __entry->live_rss << (PAGE_SHIFT - 10), + __entry->node + ) +); + +#endif /* _TRACE_MM_REAPER_H */ + +#include diff --git a/kernel/fork.c b/kernel/fork.c index 0976770db..4bdffa188 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -126,6 +126,9 @@ =20 #define CREATE_TRACE_POINTS #include +#ifdef CONFIG_ASYNC_MM_TEARDOWN +#include +#endif =20 #include =20 @@ -3461,8 +3464,10 @@ static bool async_mm_teardown_reserve(unsigned long = rss) return true; } =20 -static void async_mm_teardown_queue(struct mm_struct *mm) +static void async_mm_teardown_queue(struct mm_struct *mm, unsigned long rs= s) { + count_vm_event(ASYNC_MM_TEARDOWN_QUEUED); + trace_mm_async_teardown_queue(mm, rss); if (llist_add(&mm->async_reap_node, &mm_reaper_list)) wake_up(&mm_reaper_wait); } @@ -3478,6 +3483,9 @@ static int mm_reaper(void *unused) batch =3D llist_del_all(&mm_reaper_list); llist_for_each_entry_safe(mm, n, batch, async_reap_node) { unsigned long pages =3D mm->async_reap_rss; + unsigned long live =3D get_mm_rss(mm); + + trace_mm_async_teardown_reap(mm, pages, live); =20 __mmput(mm); /* may free mm via mmdrop */ atomic_long_sub(pages, &mm_reaper_pending_pages); @@ -3515,12 +3523,14 @@ void mmput_exit(struct mm_struct *mm) =20 if (async_mm_teardown_reserve(rss)) { mm->async_reap_rss =3D rss; - async_mm_teardown_queue(mm); + async_mm_teardown_queue(mm, rss); return; } + count_vm_event(ASYNC_MM_TEARDOWN_REJECTED); } } =20 + count_vm_event(ASYNC_MM_TEARDOWN_SYNC); __mmput(mm); } =20 diff --git a/mm/vmstat.c b/mm/vmstat.c index 4e26e5fd6..e3ff30027 100644 --- a/mm/vmstat.c +++ b/mm/vmstat.c @@ -1494,6 +1494,11 @@ const char * const vmstat_text[] =3D { [I(NRSWPIN)] =3D "nrswpin", [I(NRSWPOUT)] =3D "nrswpout", #endif /* CONFIG_SWAP */ +#ifdef CONFIG_ASYNC_MM_TEARDOWN + [I(ASYNC_MM_TEARDOWN_QUEUED)] =3D "async_mm_teardown_queued", + [I(ASYNC_MM_TEARDOWN_SYNC)] =3D "async_mm_teardown_sync", + [I(ASYNC_MM_TEARDOWN_REJECTED)] =3D "async_mm_teardown_rejected", +#endif #undef I #endif /* CONFIG_VM_EVENT_COUNTERS */ }; --=20 2.34.1 From nobody Sat Jul 25 03:47:43 2026 Received: from sender-pp-o91.zoho.in (sender-pp-o91.zoho.in [103.117.158.91]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EA1F625B08B for ; Sun, 19 Jul 2026 18:56:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=pass smtp.client-ip=103.117.158.91 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487390; cv=pass; b=onXd62RGrE3TjhNV8LH2V4NFQ4vkgKqrE0qoXSSeIiCa6/nm88rTX5gXUaiJzTaIM2+Qsch2NoREyVE1Hm8z+lSmxRYe86UMTjiwKCHm4905x1NdIYtO1byUjU3nzkYkTnuQqh9SuXYefwS3mn7v/qd0WcnuK49Sfcd8l2jp7v0= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784487390; c=relaxed/simple; bh=N+21rkpaZHsE/JH8ThNMpu23yT81kMKbJpy7npclSVA=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=CH1sgSsxVsWHwxWGEH4j2b/yypvTUQsN6w8WCcvRatjvMRnkfNe4SxODUDZW0pRzAscukHAnmA0Rz5yB7kGUUjzf0AM8xVYDWYC0BLJy9GSHQVu8JcOR7bPA3H3/k3OKO42CKObAZoHIwcc6kdvvGzLtHEHYkxKzme3xLQyGBsM= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in; spf=pass smtp.mailfrom=zohomail.in; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b=oSDCqezK; arc=pass smtp.client-ip=103.117.158.91 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=zohomail.in Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=zohomail.in Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=zohomail.in header.i=adi.sharma@zohomail.in header.b="oSDCqezK" ARC-Seal: i=1; a=rsa-sha256; t=1784487289; cv=none; d=zohomail.in; s=zohoarc; b=McJqNokg8RJLtID0rDpHo9dZHGTcCImBNyjNUYGUNKx7Hy6zqpz3HtC81y5rc0sfXObKBS4Y+e3lTVp572kZ92nL9WRYPvz//dOHQDIeg+gxOl+WN+PHsSnsz52S5WD4bEC9Z53LbdghM0M9Ze7kQ74LRQEBinvoaEWzn3KnLUU= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.in; s=zohoarc; t=1784487289; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:MIME-Version:Message-ID:Subject:Subject:To:To:Message-Id:Reply-To; bh=XAzpNKpcNCtyrAHnItiN1RWU6bwJlMRMpPXgZD5jnVo=; b=Cj+k8yjMZlLPaVGvfAwb5N7TWBlR2uu0H1Y0tQi/gckdCPxFbiix3Cp0dsDTOk46SEZrU85QatitPy3wLhs10Y95LVTBNhmKcpn50pPxVFVfZJZJSJtjy7hcCFYBaZdsTEQt2QM4cCTszi4dKGTXQw5qoLu+mtPUxiEs+BoIrF4= ARC-Authentication-Results: i=1; mx.zohomail.in; dkim=pass header.i=zohomail.in; spf=pass smtp.mailfrom=adi.sharma@zohomail.in; dmarc=pass header.from= DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; t=1784487289; s=zoho; d=zohomail.in; i=adi.sharma@zohomail.in; h=From:From:To:To:Cc:Cc:Subject:Subject:Date:Date:Message-Id:Message-Id:In-Reply-To:MIME-Version:Content-Transfer-Encoding:Reply-To; bh=XAzpNKpcNCtyrAHnItiN1RWU6bwJlMRMpPXgZD5jnVo=; b=oSDCqezKKsi/qJYD/MBrcfiWX378g9IN6DcHLRiXs7+RN7K7Vq+gHxkoR4s4DpBV L2JKMZvU1ciBmMO3VARq/OG04R6t+GbSS0nZ1/VSTzyqv/lZKKlM6vvNI9KpBKFEJV8 3guQw0Ax2RZ/tTah8w+9X8Z8wsJJ+ouXl2Xy2unc= Received: by mx.zoho.in with SMTPS id 178448728809969.40393444362564; Mon, 20 Jul 2026 00:24:48 +0530 (IST) From: Aditya Sharma To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko Cc: David Rientjes , Shakeel Butt , Jonathan Corbet , Shuah Khan , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Kees Cook , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , linux-mm@kvack.org, linux-doc@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-kernel@vger.kernel.org, imbrenda@linux.ibm.com, Aditya Sharma Subject: [RFC PATCH 7/7] exit: route exit_mm()'s final mmput() through mmput_exit() Date: Mon, 20 Jul 2026 00:24:09 +0530 Message-Id: <20260719185409.409685-8-adi.sharma@zohomail.in> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260719185409.409685-1-adi.sharma@zohomail.in> References: <20260719185409.409685-1-adi.sharma@zohomail.in> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ZohoMailClient: External Content-Type: text/plain; charset="utf-8" This makes the series operative. So far, nothing called mmput_exit(). Switch exit_mm() from mmput() to mmput_exit(), so the exiting task's final reference drop can defer the address-space teardown to mm_reaper instead of running exit_mmap() on the exiting CPU. Only the exit path is routed. Every other mmput() caller (get_task_mm() users like ptrace and /proc, kthread_unuse_mm(), etc.) still tears down inline, and if one of those ends up holding the actual last reference, behavior is unchanged -- the deferral applies only when the exiting task's own drop is the final one. The deferral triggers only when all of the following hold, and falls back to inline __mmput() otherwise: - vm.async_mm_teardown=3D1 (static key, default off) - RSS >=3D vm.async_mm_teardown_thresh_pages (default 64MB) - the mm is not an OOM target (MMF_OOM_TARGETED) and has not already been reaped (MMF_OOM_SKIP) - charging the RSS to the pending budget stays under vm.async_mm_teardown_max_pending_pages Results, bare metal, 32 CPUs / 32GB RAM (31 GiB usable), single NUMA node, performance governor, medians of 10 reps, THP off unless noted, vm.async_mm_teardown_max_pending_pages raised to 3/4 of RAM for the async runs (the RAM/4 default would reject a single 16GB mm on this box; see the cover letter's open question on the cap default): reap latency (SIGKILL -> waitpid() returns), async off -> on: 1GB 18.01 ms -> 0.14 ms 4GB 73.71 ms -> 0.14 ms 16GB 282.16 ms -> 0.15 ms 16GB 27.68 ms -> 0.13 ms (THP on) voluntary exit (_exit() -> reap) matches within noise: 16GB 271.27 ms -> 0.12 ms redis-server with a 16GB populated dataset (real heap, VmRSS ~16.2GB): 292.97 ms -> 0.24 ms 32 simultaneous 0.5GB exits, tail until all reaped: 157.42 ms -> 0.32 ms below-threshold 30MB process with the feature enabled: 0.56 ms -> 0.60 ms (sync path, within noise) A CONFIG_ASYNC_MM_TEARDOWN=3Dn build of the same tree reproduces the async-off column within noise (18.01 vs 18.01 ms at 1GB, 282.14 vs 282.16 ms at 16GB), so the config itself costs nothing when dark. The work is moved, not eliminated: with async on, the 16GB region is fully back in the buddy allocator ~240 ms after the kill (vs ~283 ms when torn down inline), and the reaper spends ~1.18x the inline CPU on the same teardowns (nice-19 kthread, cache-cold on another CPU). The cover letter carries the full matrix: freeing-latency checks, backpressure fallback, OOM-killer interaction (an OOM victim is never queued, tracepoint-verified, including the CLONE_VM-sharer case), tmpfs inode eviction in reaper context, sysctl toggling under load, and a freezer cycle with a queued backlog. Signed-off-by: Aditya Sharma --- kernel/exit.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/kernel/exit.c b/kernel/exit.c index 1056422bc..6f97d204f 100644 --- a/kernel/exit.c +++ b/kernel/exit.c @@ -607,7 +607,7 @@ static void exit_mm(void) task_unlock(current); mmap_read_unlock(mm); mm_update_next_owner(mm); - mmput(mm); + mmput_exit(mm); if (test_thread_flag(TIF_MEMDIE)) exit_oom_victim(); } --=20 2.34.1