From nobody Mon Sep 28 14:48:02 2026 Received: from forwardcorp1b.mail.yandex.net (forwardcorp1b.mail.yandex.net [178.154.239.136]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 10EAC470110 for ; Thu, 20 Aug 2026 13:37:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=178.154.239.136 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787233054; cv=none; b=tIxHpjwTnFvdfevMbkVSOjYAxvNTsA2c0vTcESN/f+ZJwnw/sWLdjumL1WRd4wZTsbhHEvVyIlfb9eO8mrXqN0j/bZ15pnqEqeiH4b3HO12QVjLhDyX8XTq17aQVnEFcZfCeZOJ5pHtQn/04rfQjJHhFlw8YP4k9e7yqFvabW2E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787233054; c=relaxed/simple; bh=tZWD+5pyobzwbGyfAt98Ke5uWbwBPyxzgnfcvqx0s3I=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=ZMjK6iJ84Rk870ZrEj42IGFbP61O/wVrFUw0+tCC+Htis03+R1Amufk7d1KWAhNgKxoj2O7GekimLjNdk24W8LfXVNmZr0OQBAoakRQVKRZEbCzqeXlCrigoxIMyL37aeMt6D2+CSo0xFApP8oXZ83/5iDsNlJ0IRyGcTPBWkj8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=yandex-team.ru; spf=pass smtp.mailfrom=yandex-team.ru; dkim=pass (1024-bit key) header.d=yandex-team.ru header.i=@yandex-team.ru header.b=sQ7/vEfs; arc=none smtp.client-ip=178.154.239.136 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=yandex-team.ru Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=yandex-team.ru Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=yandex-team.ru header.i=@yandex-team.ru header.b="sQ7/vEfs" Received: from mail-nwsmtp-smtp-corp-main-80.iva.yp-c.yandex.net (mail-nwsmtp-smtp-corp-main-80.iva.yp-c.yandex.net [IPv6:2a02:6b8:c0c:118b:0:640:b49:0]) by forwardcorp1b.mail.yandex.net (postfix) with ESMTPS id 12ECA80848; Thu, 20 Aug 2026 16:37:27 +0300 (MSK) Received: from i101646577.yandex-team.ru (unknown [2a02:6bf:8080:a66::1:2b]) by mail-nwsmtp-smtp-corp-main-80.iva.yp-c.yandex.net (smtpcorp) with ESMTPSA id NbXRRG1eQ0U0-VPv2Jkxg; Thu, 20 Aug 2026 16:37:26 +0300 X-Yandex-Fwd: 1 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=yandex-team.ru; s=default; t=1787233046; bh=M33AQZULeMLglpMel8DACoCzUVp1sCxykav1O7jduSw=; h=Message-ID:Date:Cc:Subject:To:From; b=sQ7/vEfs9B82gh0TzyYSDISaVxSKQ/DEV1E057jYCdbEiPNLBqowBxdmkqPKHwcgx aj+vpDUD1cjDcDZ7xptrkoiXWZ9o+k3COKVEVJCPd4BghOzNeIINfnW+RSlhOrItGi fDR3nvicxW4oQMTVNtRIp7TJ0g+1vcyRGkJ9QllI= Authentication-Results: mail-nwsmtp-smtp-corp-main-80.iva.yp-c.yandex.net; dkim=pass header.i=@yandex-team.ru From: Daniil Tatianin To: Andrew Morton , linux-mm@kvack.org Cc: Daniil Tatianin , Vlastimil Babka , Suren Baghdasaryan , Michal Hocko , Brendan Jackman , Johannes Weiner , Zi Yan , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Mike Rapoport , linux-kernel@vger.kernel.org Subject: [PATCH] mm/vmstat: add per-order allocation slow path statistics Date: Thu, 20 Aug 2026 16:36:58 +0300 Message-ID: <20260820133659.712111-1-d-tatianin@yandex-team.ru> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Production incidents caused by bursts of high-order allocations all entering direct compaction are currently hard to attribute from /proc/vmstat: pgalloc_* has no order breakdown, and compact_stall does not say which order stalled. Tracepoints can recover this on a single machine, but they are impractical as an always-on fleet-wide monitoring source, which is what is needed to correlate latency regressions with allocation behavior after the fact. Add per-order event counters to /proc/vmstat, covering only the allocation slow path, so the page allocator fast path is not touched at all: - pgalloc_slowpath_orderN: entries into __alloc_pages_slowpath(), counted once per allocation, before the restart loop - pgalloc_fail_orderN: allocations that returned NULL to the caller (including a successful allocation freed by memcg charge failure) - compact_stall_orderN / compact_success_orderN: per-order split of the existing direct compaction counters, order 0 is omitted since direct compaction is never entered for it All new counters are purely additive: the existing keys are untouched and compact_stall =3D=3D sum of compact_stall_orderN. alloc_pages_nolock() is deliberately not counted: it is opportunistic, never enters the slow path, and its NULL returns are expected rather than failures. Counter names are generated for any MAX_PAGE_ORDER the arch Kconfig ranges allow (10..13), a static_assert catches larger values. A per-order split of PGALLOC itself was proposed in 2017 but stalled over fast path overhead concerns, restricting the counters to the slow path avoids that overhead entirely while still capturing the allocations that cause latency. Link: https://lore.kernel.org/all/1499346271-15653-1-git-send-email-guro@fb= .com/ Signed-off-by: Daniil Tatianin --- include/linux/vm_event_item.h | 19 +++++++++++++ mm/page_alloc.c | 7 +++++ mm/vmstat.c | 52 +++++++++++++++++++++++++++++++++++ 3 files changed, 78 insertions(+) diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h index 03fe95f5a020..724561b7e800 100644 --- a/include/linux/vm_event_item.h +++ b/include/linux/vm_event_item.h @@ -175,6 +175,25 @@ enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT, KSTACK_REST, #endif #endif /* CONFIG_DEBUG_STACK_USAGE */ + /* + * Per-order allocation statistics: each *_FIRST..*_LAST range + * is indexed by allocation order. + */ + PGALLOC_SLOWPATH_ORDER_FIRST, + PGALLOC_SLOWPATH_ORDER_LAST =3D + PGALLOC_SLOWPATH_ORDER_FIRST + MAX_PAGE_ORDER, + PGALLOC_FAIL_ORDER_FIRST, + PGALLOC_FAIL_ORDER_LAST =3D + PGALLOC_FAIL_ORDER_FIRST + MAX_PAGE_ORDER, +#ifdef CONFIG_COMPACTION + /* Direct compaction is never entered for order 0 */ + COMPACTSTALL_ORDER_FIRST, + COMPACTSTALL_ORDER_LAST =3D + COMPACTSTALL_ORDER_FIRST + MAX_PAGE_ORDER - 1, + COMPACTSUCCESS_ORDER_FIRST, + COMPACTSUCCESS_ORDER_LAST =3D + COMPACTSUCCESS_ORDER_FIRST + MAX_PAGE_ORDER - 1, +#endif NR_VM_EVENT_ITEMS }; =20 diff --git a/mm/page_alloc.c b/mm/page_alloc.c index ee902a468c2f..2a2b14f3b516 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -4169,6 +4169,7 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned= int order, * count a compaction stall */ count_vm_event(COMPACTSTALL); + count_vm_event(COMPACTSTALL_ORDER_FIRST + order - 1); =20 /* Prep a captured page if available */ if (page) @@ -4184,6 +4185,7 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned= int order, zone->compact_blockskip_flush =3D false; compaction_defer_reset(zone, order, true); count_vm_event(COMPACTSUCCESS); + count_vm_event(COMPACTSUCCESS_ORDER_FIRST + order - 1); return page; } =20 @@ -4757,6 +4759,8 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int o= rder, WARN_ON_ONCE(current->flags & PF_MEMALLOC); } =20 + count_vm_event(PGALLOC_SLOWPATH_ORDER_FIRST + order); + restart: compaction_retries =3D 0; no_progress_loops =3D 0; @@ -5323,6 +5327,9 @@ struct page *__alloc_frozen_pages_noprof(gfp_t gfp, u= nsigned int order, page =3D NULL; } =20 + if (unlikely(!page)) + count_vm_event(PGALLOC_FAIL_ORDER_FIRST + order); + trace_mm_page_alloc(page, order, alloc_gfp, ac.migratetype); kmsan_alloc_page(page, order, alloc_gfp); =20 diff --git a/mm/vmstat.c b/mm/vmstat.c index f534972f517d..08596edc53ca 100644 --- a/mm/vmstat.c +++ b/mm/vmstat.c @@ -1184,6 +1184,50 @@ int fragmentation_index(struct zone *zone, unsigned = int order) [xx##_MOVABLE] =3D yy "_movable", \ TEXT_FOR_DEVICE(xx, yy) =20 +#if MAX_PAGE_ORDER >=3D 11 +#define TEXT_FOR_ORDER_11(xx, yy) [I((xx) + 11)] =3D yy "11", +#else +#define TEXT_FOR_ORDER_11(xx, yy) +#endif + +#if MAX_PAGE_ORDER >=3D 12 +#define TEXT_FOR_ORDER_12(xx, yy) [I((xx) + 12)] =3D yy "12", +#else +#define TEXT_FOR_ORDER_12(xx, yy) +#endif + +#if MAX_PAGE_ORDER >=3D 13 +#define TEXT_FOR_ORDER_13(xx, yy) [I((xx) + 13)] =3D yy "13", +#else +#define TEXT_FOR_ORDER_13(xx, yy) +#endif + +static_assert(MAX_PAGE_ORDER <=3D 13, + "extend TEXT_FOR_ORDER_* for this MAX_PAGE_ORDER"); + +/* + * (xx) + n must resolve to the vm_event_item for order n, so ranges that + * start at order 1 pass their *_ORDER_FIRST item minus one. + */ +#define TEXTS_FOR_NONZERO_ORDERS(xx, yy) \ + [I((xx) + 1)] =3D yy "1", \ + [I((xx) + 2)] =3D yy "2", \ + [I((xx) + 3)] =3D yy "3", \ + [I((xx) + 4)] =3D yy "4", \ + [I((xx) + 5)] =3D yy "5", \ + [I((xx) + 6)] =3D yy "6", \ + [I((xx) + 7)] =3D yy "7", \ + [I((xx) + 8)] =3D yy "8", \ + [I((xx) + 9)] =3D yy "9", \ + [I((xx) + 10)] =3D yy "10", \ + TEXT_FOR_ORDER_11(xx, yy) \ + TEXT_FOR_ORDER_12(xx, yy) \ + TEXT_FOR_ORDER_13(xx, yy) + +#define TEXTS_FOR_ORDERS(xx, yy) \ + [I(xx)] =3D yy "0", \ + TEXTS_FOR_NONZERO_ORDERS(xx, yy) + const char * const vmstat_text[] =3D { /* enum zone_stat_item counters */ #define I(x) (x) @@ -1488,6 +1532,14 @@ const char * const vmstat_text[] =3D { #if THREAD_SIZE > 65536 [I(KSTACK_REST)] =3D "kstack_rest", #endif +#endif + TEXTS_FOR_ORDERS(PGALLOC_SLOWPATH_ORDER_FIRST, "pgalloc_slowpath_order") + TEXTS_FOR_ORDERS(PGALLOC_FAIL_ORDER_FIRST, "pgalloc_fail_order") +#ifdef CONFIG_COMPACTION + TEXTS_FOR_NONZERO_ORDERS(COMPACTSTALL_ORDER_FIRST - 1, + "compact_stall_order") + TEXTS_FOR_NONZERO_ORDERS(COMPACTSUCCESS_ORDER_FIRST - 1, + "compact_success_order") #endif #undef I #endif /* CONFIG_VM_EVENT_COUNTERS */