From nobody Fri Jul 24 22:54:56 2026 Received: from mail-qt1-f178.google.com (mail-qt1-f178.google.com [209.85.160.178]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BB83D3624AB for ; Wed, 22 Jul 2026 15:00:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.178 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784732415; cv=none; b=SivVT4ynUQprHEPAuDgOMG41cf+3CpnuwS/SX5VLp1bXIEtR5nSyxteMz73ieSL3DUDkhQW/GDX7BgB0N/NhcAGqcj7t3uMZE1ga4pNcYfXyLT4q5ud/OrzD8TWnqfrhnyf4MQ6TG75jfgpFFSCY7Vl61QxsWDFTY84tJ3cBnYU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784732415; c=relaxed/simple; bh=EoGwkVjZ83g4P4XGelFN68Dd+/8XYcpHXKQhq0ckNng=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=LtkDvw/CBgI+LacH3Stc1UaqEoWCkD/jyp8QAGSPeIuQqWNexey7/noa9cy81HRPsDNZh9nb69gO5ePzDnt1MMhnKkUipe0x9KdCdR+PUeVPr9AIydPfqzNxA8u+NGiXEY5Hv04ywlnmPJi7FO/UGI6nrd5RklmOGDZrfQn4A/E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=o9Df/i1f; arc=none smtp.client-ip=209.85.160.178 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="o9Df/i1f" Received: by mail-qt1-f178.google.com with SMTP id d75a77b69052e-51a868b6962so145174371cf.2 for ; Wed, 22 Jul 2026 08:00:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1784732412; x=1785337212; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=HEuJeVpO7/an9TIgcv/7XS3Xju/554aT8fkPf9kxlBg=; b=o9Df/i1fRfAP4PDXDf6UfU+xhakIH7Yh5iBEKt+g3QROZVusN1yCn+ct45gMVD7ddi 4M3El3WpTiONA5DXBDVOp325KYbOOIqQl8RE7tFtNzkbYLZwWwmjAcSKxS9mpzPrWcLS ojMxHoB2td1Am0NlyXNzejWcJxD4MPby2GYPfsABesqE6eKfy4OTkvSvE5DwpZZhrWbm NJ7q1RiOzdbN2VEIC0wWLadJLpgmBojjNCO3ud7dLEMdMw3nTAqAqGReABwNSYtaXBXK tVjsHbHYa2L9h0kkySGtouW2pqs2ygKRAM6mn4JUaY8ulw90txcMQvDGDWCH6vuGN4In vS3A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784732412; x=1785337212; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=HEuJeVpO7/an9TIgcv/7XS3Xju/554aT8fkPf9kxlBg=; b=dEGI5uRVPo5419smASdJgKzUja0ZWNZQHNG/Xqn/ZebM4HGfu3Ecl77T96UDyFfe6D A5B6kr2eVi9iwWJhAeSfZFNXbKW0NLZjtb+o0zkjTwKH3e1Op2sO1DVkWY/+AE7K9Fac Ev2iIc4q+kvPj/6I6RNocy7wP8i6wN8GqUq19fwXgR8102z1RJS9DaygBJU1HujoggSX 7CQfTHKrvX2NxXEBeZH1+qf//Dyc3ccj5BF+UGN259YMGJKsPBtIvgcQBwK9jSezV+3u 9QfKynpDaP30BzP4+gTADMUOXdr90NLN/7VEmJClu9k7sI8xv3N145Hwp3TLbLgpOw6t pHng== X-Forwarded-Encrypted: i=1; AHgh+RqlfF24zVl5fotMASvxWk988wXCOEMZv8e5eyqZ3MQPvyy0hdrxuFI7t3KOhzWdPM5OJYKgoQZgx3pziIA=@vger.kernel.org X-Gm-Message-State: AOJu0YwxTFTV8hCNjKPQ/2vjudvKTgXTV4JVnoDCVTBEHfzMaJAoQdq+ pVVdqYOqJFzjZcJ9M5U8ojZEhahqY/XMz3Zf66bLrYLKkH44fpfpa0qhd6ILLNUwCqg= X-Gm-Gg: AR+sD10tmw1lDIZDXr4Qbxa+KYsX7uW41gotTCLmKpLUfj8yMfXO5dw5hItGvrgTagV Bdd9IZAM5mn3bCTQdBIGs/TcMKr2Yu763bIMakX1kw2mS3w89R73aCbjXB0kp5JDg2/lfyPNGvW ESc0uwl1w1Cb3VSMqm7lwBRwwZCJ+mXvf1t5MwpLInKpDRv+hb/xcuRtJKAtIyizE+XLHaywd8S mH3/lLVIIRRtR6oqg2mRB/07j5xDE0ZDhf9H+pwRmAolIAdBmjhSRl4Dz7P4dJWZ6p8JVupS+ad V8G8XKku5Nwp+UULrjkSzp2WrBDNZYEp5VqeOLmgRQYU/ykaWAg/GqPKZOHFN+cIIR5Dc0oT8M5 kZSDLeqiy/Yq8oUyEf3uPPrK60irpXD0pBIgrcXYCQ/g2mQfTGcgXreUWfIKQGXDCvh6gRJ6rIw AOVaCIM4fqAvU= X-Received: by 2002:a05:622a:1186:b0:517:6804:3732 with SMTP id d75a77b69052e-5213e77cf30mr202908201cf.55.1784732412231; Wed, 22 Jul 2026 08:00:12 -0700 (PDT) Received: from localhost ([2603:7001:f100:500:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-907ba8d6180sm22721736d6.16.2026.07.22.08.00.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 08:00:11 -0700 (PDT) From: Johannes Weiner To: Andrew Morton , Vlastimil Babka Cc: Suren Baghdasaryan , Michal Hocko , Brendan Jackman , Zi Yan , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Mike Rapoport , Shakeel Butt , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [PATCH v2 1/4] mm: page_alloc: __GFP_FS lockdep annotation for direct compaction Date: Wed, 22 Jul 2026 10:56:44 -0400 Message-ID: <20260722150006.3848560-2-hannes@cmpxchg.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260722150006.3848560-1-hannes@cmpxchg.org> References: <20260722150006.3848560-1-hannes@cmpxchg.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" A subsequent patch will have some order-0 allocations participate in compaction under defrag_mode, to stave off extfrag events. Since this is a sprawling expansion of entry points, and compaction can enter filesystem paths, add lockdep annotations that catches __GFP_FS passing errors. Direct reclaim has had this annotation for a while, and since reclaim and compaction are usually used in conjunction, this is unlikely to unearth old bugs. It's more about future proofing and peace of mind. Reviewed-by: Vlastimil Babka (SUSE) Acked-by: Shakeel Butt Signed-off-by: Johannes Weiner --- mm/page_alloc.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/mm/page_alloc.c b/mm/page_alloc.c index ee902a468c2f..cb422505c6ef 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -4152,12 +4152,14 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsign= ed int order, =20 psi_memstall_enter(&pflags); delayacct_compact_start(); + fs_reclaim_acquire(gfp_mask); noreclaim_flag =3D memalloc_noreclaim_save(); =20 *compact_result =3D try_to_compact_pages(gfp_mask, order, alloc_flags, ac, prio, &page); =20 memalloc_noreclaim_restore(noreclaim_flag); + fs_reclaim_release(gfp_mask); psi_memstall_leave(&pflags); delayacct_compact_end(); =20 --=20 2.55.0 From nobody Fri Jul 24 22:54:56 2026 Received: from mail-qk1-f177.google.com (mail-qk1-f177.google.com [209.85.222.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AED963749F6 for ; Wed, 22 Jul 2026 15:00:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.177 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784732417; cv=none; b=D+R+UnZL9ULkDi9YrgKf1IlJiajNCkoGxofO83qbWjp7L7U18bf1DnTvjNno3ii99SXL7H820YODEiKZdUz+USnvmV+NKMV+6qJ3RnL71YWTj4AaN5OGpfrUE4OFBydpjAC6kS0MG7Lw4ruBtSlrkrqRurBWfa4DftbSeAwTzFY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784732417; c=relaxed/simple; bh=INHUvS4Xk54MpBqgU7R471x+YkLibsTfRlnI/xBZiMk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gnlLE0b+QEV2fB8WgDiS9SPU6eYg7YkwU2y/gj627Xku6tCBTumxFC7Bk7ysK7HgBaM/0yIMOmb7VFJp3/aHJlwnLbbNNUJs8uU+PC7iyaAxp7sMHHavTPUP0lOL7vvwlXuPENPNL2RjvbDSZfSsV/ILY0crcajLLEm8G/9M9oM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=Tyu+LVCE; arc=none smtp.client-ip=209.85.222.177 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="Tyu+LVCE" Received: by mail-qk1-f177.google.com with SMTP id af79cd13be357-92e67555e24so672520885a.3 for ; Wed, 22 Jul 2026 08:00:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1784732414; x=1785337214; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=WzWEJqU7j0ZiAwHq29jiPVthqAQlNA5+hbQ4+opl0xY=; b=Tyu+LVCEldgJ4p+k6aXx2lli1Sps4iSLwArxt8IsHvn300w9XRxEA8YwdmKR7jh1Mi PGJsYvlxa9pHhOlDtn7JIsNiHzXMBiboZw5kOH6xMMebhglpf9LWZY/IqbuQKEZ0WtVI SYXzNknxggsP7oOh9s6afaRTraj2UpnWhUiTMuM+X3BSXxAm3j2BbBsmdlb+SYcFc+hW F6qN6kdmOu2J675/m+lrKXPaeL+b1fX8xa0rsEhK2FzwjiDESNGz74dMPHqBNb89ISpG c0Ehy83UQ5nxSz3rO0SBysN6xGNGudCEVNvz9uq1YgY2z1lg0ucT5YLeeWTvgzXABJHR VLgw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784732414; x=1785337214; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=WzWEJqU7j0ZiAwHq29jiPVthqAQlNA5+hbQ4+opl0xY=; b=MG31E4nhUz2IXQJVjfYNo6r3pHbowcV7g3TFHhaHmEWqmW4mi9x7c7CS7L+HCIxpJ/ G+CfWbPSOwR2MvZbfeZ1TG7xm4SgtKe/lVTcH/xwPo1rsDunnMoYdm6WAILlRUa58BdD mn6NdpFBPD2i8zeI4mbN73/WkvvB2cvxb+Trtxf6h1fQSNcg4+AeIovEEyYHxfavUzBg ZZFveAZJru4owcIXXhacMW6PGNiA1elZNNfwmh0A3Ds+QRpNn2YBoRsE+cSAdfcaqNrw Ed+LHm8VyAsBl47Zf7tXV+32aTv3sRW/BnguKScAaOW1Dy8ETb3k6dJYiyXoowhNXnxG NokA== X-Forwarded-Encrypted: i=1; AHgh+RqzxgqirGeBA8H6VnDuLWsP8Ijezy6Ifa/FlxiiUVE++W1QYGkdBhTogZ5Y72c+ApAX1yHsaZY4S43Vbss=@vger.kernel.org X-Gm-Message-State: AOJu0YzMeIQ4zR7fTJ3qdJqDps4NZ/YV2bV6l9yxXGXGvrnO7qZonlxW Tg/eIhJV2XBxu5R4cK2DTiCmGN7m5wyBrx+s5I0psUmD9ZfwYoc0U9d4a6W7Fp3E84U= X-Gm-Gg: AR+sD12ciodkqc+dB9hcxLyQVySMzhyxhiF27pj8alPCAD50SIS3lhAZeaQcFXpo6nr GrqLQZHqhDnPpXqQizFJAdQ1zurbO8mh05VYptmHzGYa7DpwqCcGGIG7CWE4pYFFi+T+Bb41L61 gKrcKZxcqcA8WnwoCMLepWyAQLER6JHA5xWxdQTSjF2vop2/2CdKVEPXSUv/dozLupq/ELZzrnl xFJ7CQSyr8Tw1QhkjgwS74Z0Rwh8LNBo0QOxJMeRGRv2Vavr7b4gmkC3dI9noU7AbEkfg6XSVrP kox/yPnhnZLUS5M4ly1JTAnFjhM+epPoUVUi2JJxyVXT+rucaJg3Vy7pxp3QQpeSRV2VKi3xkwH W9keAmyfPFPXx8ovjeyHTjCkGnSjiPpQIuOdtl56v1gsnCt5G47gMcLBVr44iZm0tTVrK/mFByn Zauc5X81aDA3U= X-Received: by 2002:a05:620a:4008:b0:92d:5504:72dd with SMTP id af79cd13be357-930b3ed9d45mr2222220285a.10.1784732414019; Wed, 22 Jul 2026 08:00:14 -0700 (PDT) Received: from localhost ([2603:7001:f100:500:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-907ba672806sm23149956d6.0.2026.07.22.08.00.13 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 08:00:13 -0700 (PDT) From: Johannes Weiner To: Andrew Morton , Vlastimil Babka Cc: Suren Baghdasaryan , Michal Hocko , Brendan Jackman , Zi Yan , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Mike Rapoport , Shakeel Butt , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [PATCH v2 2/4] mm: compaction: support non-movable compaction for pageblock requests Date: Wed, 22 Jul 2026 10:56:45 -0400 Message-ID: <20260722150006.3848560-3-hannes@cmpxchg.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260722150006.3848560-1-hannes@cmpxchg.org> References: <20260722150006.3848560-1-hannes@cmpxchg.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" While trying to fix a reclaim storm in defrag_mode, I noticed that non-movable direct compaction is extremely inefficient. When searching for space to evacuate, compaction only allows blocks of the same type as the incoming request. This is to prevent migratetype pollution, where a small non-movable request frees space in a movable block and provokes the allocator to fall back and pollute it. This protection is reasonable on one hand, but the downside is that it makes non-movable direct compaction nearly useless: if we get the type annotations right, by definition there aren't any movable pages inside the non-movable blocks it is allowed to scan. With defrag_mode, the goal is the production of whole blocks, which are essentially type neutral: __rmqueue_claim() will convert them wholesale on alloc. This makes type mixing and pollution a non-issue. Fix the pollution gates to take the requested order into account, and allow whole-block requests to scan blocks of other types. The only exception is CMA blocks. That type is sticky and these blocks cannot be claimed to other types. Continue to be strict with them, and allow only explicit ALLOC_CMA requests and kcompactd to evacuate them. Reviewed-by: Vlastimil Babka (SUSE) Signed-off-by: Johannes Weiner Reviewed-by: Gregory Price --- mm/compaction.c | 46 +++++++++++++++++++++++++++++++++++++++------- 1 file changed, 39 insertions(+), 7 deletions(-) diff --git a/mm/compaction.c b/mm/compaction.c index f08765ade014..225862a00380 100644 --- a/mm/compaction.c +++ b/mm/compaction.c @@ -1381,12 +1381,44 @@ static bool suitable_migration_source(struct compac= t_control *cc, if (pageblock_skip_persistent(page)) return false; =20 - if ((cc->mode !=3D MIGRATE_ASYNC) || !cc->direct_compaction) + /* + * Background compaction produces blocks for the zone at + * large, with no particular allocation context. Allow all + * block types, including CMA. + */ + if (!cc->direct_compaction) return true; =20 block_mt =3D get_pageblock_migratetype(page); =20 - if (cc->migratetype =3D=3D MIGRATE_MOVABLE) + /* + * CMA pages can only be taken by ALLOC_CMA requests. For anybody + * else, vacating a CMA block consumes free pages the caller + * could have used, and produces free pages it cannot. + */ + if (is_migrate_cma(block_mt) && !(cc->alloc_flags & ALLOC_CMA)) + return false; + + /* + * Per default, scans are restricted to blocks compatible with + * the request, to prevent cross-contamination. Once + * compaction priority escalates to synchronous scans, though, + * scan all blocks to try to make forward progress. For + * movable request, this likely helps little: there shouldn't + * be many migratable pages inside non-movable blocks besides + * allocator fallbacks. For non-movable requests, this helps a + * lot, as they can finally scan movable blocks. + */ + if (cc->mode !=3D MIGRATE_ASYNC) + return true; + + /* + * Prevent migratetype =3D=3D MIGRATE_MOVABLE || cc->order >=3D pageblock_or= der) return is_migrate_movable(block_mt); else return block_mt =3D=3D cc->migratetype; @@ -1974,12 +2006,12 @@ static unsigned long fast_find_migrateblock(struct = compact_control *cc) return pfn; =20 /* - * Only allow kcompactd and direct requests for movable pages to - * quickly clear out a MOVABLE pageblock for allocation. This - * reduces the risk that a large movable pageblock is freed for - * an unmovable/reclaimable small allocation. + * Prevent direct_compaction && cc->migratetype !=3D MIGRATE_MOVABLE) + if (cc->direct_compaction && cc->migratetype !=3D MIGRATE_MOVABLE && + cc->order < pageblock_order) return pfn; =20 /* --=20 2.55.0 From nobody Fri Jul 24 22:54:56 2026 Received: from mail-qv1-f46.google.com (mail-qv1-f46.google.com [209.85.219.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0990737702E for ; Wed, 22 Jul 2026 15:00:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.46 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784732420; cv=none; b=J8x9Qr3exnMi23MiOufcQE4NnUm03z0A6lzUI6CobiolrdmVrXm1Pwjon/oZNt5cE0+xityaqsq8rKcBAqceuFH7fzSEwJmUrWW/yvT7VubFyJwfS5LYQHH8EkTv8l08RapnNr12eg/8bnNHc0sGd1mWdy3N6U3B9ytAIudD1Uw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784732420; c=relaxed/simple; bh=OGdOpi7gb6g7qGXkw7ZMzS314GNbfqv/MOKmMqeaXwc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=NtC44QTZVHLclcRCf2ygDsnZ4Y4z008skZHP2jt0Uqxk2a8NeMSxUo/blM/TSQLazZj1q2hkrvUJ41a+ZGngbG4uLs+hNGvZtRYPZoAsAu9Q967uhmJa5GtwGhhQQ2txfyF96Xf9AKIXghRSp4ZYYUWkw1DPGN5vu4yvj1nCFC8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=BWMdfCoq; arc=none smtp.client-ip=209.85.219.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="BWMdfCoq" Received: by mail-qv1-f46.google.com with SMTP id 6a1803df08f44-8efbafa1bacso104245116d6.1 for ; Wed, 22 Jul 2026 08:00:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1784732417; x=1785337217; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=9w4/w+DgNeKgvmYqz2X0YQytmiACV0MWUeK4q1M7V1c=; b=BWMdfCoq4Z6uUgHmHQoNui7Mnp1xqHnQDeRKWzrSoTjHCQ+UQ5429qCMdj/dWEQLE5 LnK51KbF+VwEzzog8vtW5e/3/ijIZLF0BSEq+qgtt/HEGYmQPTB1e3o9xs3422fFOIwZ stpeC7+hNNl+09XtueSKOUW/czAgVCbQFNtxfzjG2aH8blumbT1uKcwDintRmSd7jI17 j+iDDAXLoz6gqQY+f9jOjeW0R5VbmtTh12Yp68Qm5lf2g4Ob537BEYvDlFdjWWzRA+OP MWBJizEnORj9/RhbjJXLMgy4ckEoa8+bFDNWc+8NMQk84kDad+6mShL6DZF0NHfa8YSD gUhA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784732417; x=1785337217; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=9w4/w+DgNeKgvmYqz2X0YQytmiACV0MWUeK4q1M7V1c=; b=nanPVfzCguWM/JxIO7QOCLXevEpPux2fCd5pH2RDDOHAf/J/YhMdrK+cR7PGucNfnS OGForsIeOD8LGdM2TIShS6lNz7HGMEqqCAVzk10Hvnzy07bB8rXfPBouR7vH08S8Tjq9 hnAm4D+8fp9dDAaCftf6eqJVrguhzgMCWv97NudlGLQcofN+yE4CFLNtigPYgbA4yjRo rH7Ga+RYV19FM6jJBpAViCbfPNrHdgM0ZqArEG/yWpapNkdz7SK5tBvjJsX2b0ZNfAa9 4dB+VHQMFF8fwN6lZ92Qqj5nADDm1qEBDBKR8lHil8klWZSVc9l6wxmBUUSuHxQ5ZQpB pnMA== X-Forwarded-Encrypted: i=1; AHgh+RpuGaACIn+UwMbUppiiTfWlhCLWz6fVLw3d5fVAanmckYju16m0l/abqYy+rQGLqDCAp9MtFifbuFcHIz4=@vger.kernel.org X-Gm-Message-State: AOJu0YzxDNAyAS8MPb4vt+5VflS9aLRuvk607xVWUwVgIyFUtQM6TdX6 8l9QrOSdudg1nHEKmEUXXWqEbrNQJaNd6d4PGH350f4uCd/o4Hei9vlQ3GMelf/RvlU= X-Gm-Gg: AR+sD12NgyriEQlikEYMzHge/jlqlXO6TJJG2QfFsdVrQc74n3tk7HdH5Q+oJKMn05f aHViTZxKV15s4nJPmt3y4F0QridywDFfDFtvQAwU5DnJrxb8oFhm1JK+IfWNLbY5kdS1V4tms4N IQC2jlfvPI9KpR39lyTkV9EFASMk34mMNK2hkZNcssW9e/lfCG5CTuxbrRap+FSx4zNvOJxI9Er 1Tg+Us2odX8ZZgPTXPJHDXiPKisrIjmQtfRaPyl881M+o4ezDg2//BBZkwaySIVfMkSPObybg7C y/MKx66bLL9FM6J+s1BkrgiRtM541dJqU3eQLQB6H55BpevH2fDvbXMYREFTUbt1jwGO7lzUnpL w7dFNhCHFl1MkdAxUywU4rDLoDq1YUgfR6wHRra/hpRnLr/kgxwdu4KyNcgEULvrQudcAlTjZzK ljrUIx4rYJOho= X-Received: by 2002:a05:6214:3d0d:b0:8ea:184f:c15a with SMTP id 6a1803df08f44-9077836d21amr240711036d6.17.1784732415754; Wed, 22 Jul 2026 08:00:15 -0700 (PDT) Received: from localhost ([2603:7001:f100:500:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-907baa0ef54sm22285326d6.39.2026.07.22.08.00.14 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 08:00:15 -0700 (PDT) From: Johannes Weiner To: Andrew Morton , Vlastimil Babka Cc: Suren Baghdasaryan , Michal Hocko , Brendan Jackman , Zi Yan , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Mike Rapoport , Shakeel Butt , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: [PATCH v2 3/4] mm: page_alloc: move capture_control to the page allocator Date: Wed, 22 Jul 2026 10:56:46 -0400 Message-ID: <20260722150006.3848560-4-hannes@cmpxchg.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260722150006.3848560-1-hannes@cmpxchg.org> References: <20260722150006.3848560-1-hannes@cmpxchg.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Vlastimil Babka (SUSE)" The compaction capturing code assumes the allocation request order and compaction target order are the same. That won't be true once defrag_mode promotes sub-block allocations to pageblock-order compaction: compaction targets the larger order, while capture should remain at the original allocation order. Move the capture_control to the page allocator and give it its own copies of what the page freeing path matches against - zone, migratetype and the allocation order - rather than reaching into compaction's live compact_control. __alloc_pages_direct_compact() fills in migratetype and order, and installs and hides current->capture_control around the whole compaction call; try_to_compact_pages() aims capc->zone at each zone while it is being compacted. compact_zone_order() no longer deals with capture at all. Pass the capture_control through try_to_compact_pages() / compact_zone_order() in place of the bare struct page **. No functional change. Signed-off-by: Vlastimil Babka (SUSE) Co-developed-by: Johannes Weiner Signed-off-by: Johannes Weiner Reviewed-by: Gregory Price --- include/linux/compaction.h | 3 ++- mm/compaction.c | 50 +++++++++++--------------------------- mm/internal.h | 4 ++- mm/page_alloc.c | 45 ++++++++++++++++++++++++++++------ 4 files changed, 57 insertions(+), 45 deletions(-) diff --git a/include/linux/compaction.h b/include/linux/compaction.h index f29ef0653546..66a2f70e9e01 100644 --- a/include/linux/compaction.h +++ b/include/linux/compaction.h @@ -58,6 +58,7 @@ enum compact_result { }; =20 struct alloc_context; /* in mm/internal.h */ +struct capture_control; /* in mm/internal.h */ =20 /* * Number of free order-0 pages that should be available above given water= mark @@ -92,7 +93,7 @@ extern int fragmentation_index(struct zone *zone, unsigne= d int order); extern enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order, unsigned int alloc_flags, const struct alloc_context *ac, enum compact_priority prio, - struct page **page); + struct capture_control *capc); extern void reset_isolation_suitable(pg_data_t *pgdat); extern bool compaction_suitable(struct zone *zone, int order, unsigned long watermark, int highest_zoneidx); diff --git a/mm/compaction.c b/mm/compaction.c index 225862a00380..977e30b3a2bd 100644 --- a/mm/compaction.c +++ b/mm/compaction.c @@ -2802,9 +2802,8 @@ compact_zone(struct compact_control *cc, struct captu= re_control *capc) static enum compact_result compact_zone_order(struct zone *zone, int order, gfp_t gfp_mask, enum compact_priority prio, unsigned int alloc_flags, int highest_zoneidx, - struct page **capture) + struct capture_control *capc) { - enum compact_result ret; struct compact_control cc =3D { .order =3D order, .search_order =3D order, @@ -2819,38 +2818,8 @@ static enum compact_result compact_zone_order(struct= zone *zone, int order, .ignore_skip_hint =3D (prio =3D=3D MIN_COMPACT_PRIORITY), .ignore_block_suitable =3D (prio =3D=3D MIN_COMPACT_PRIORITY) }; - struct capture_control capc =3D { - .cc =3D &cc, - .page =3D NULL, - }; - - /* - * Make sure the structs are really initialized before we expose the - * capture control, in case we are interrupted and the interrupt handler - * frees a page. - */ - barrier(); - WRITE_ONCE(current->capture_control, &capc); =20 - ret =3D compact_zone(&cc, &capc); - - /* - * Make sure we hide capture control first before we read the captured - * page pointer, otherwise an interrupt could free and capture a page - * and we would leak it. - */ - WRITE_ONCE(current->capture_control, NULL); - *capture =3D READ_ONCE(capc.page); - /* - * Technically, it is also possible that compaction is skipped but - * the page is still captured out of luck(IRQ came and freed the page). - * Returning COMPACT_SUCCESS in such cases helps in properly accounting - * the COMPACT[STALL|FAIL] when compaction is skipped. - */ - if (*capture) - ret =3D COMPACT_SUCCESS; - - return ret; + return compact_zone(&cc, capc); } =20 /** @@ -2860,13 +2829,13 @@ static enum compact_result compact_zone_order(struc= t zone *zone, int order, * @alloc_flags: The allocation flags of the current allocation * @ac: The context of current allocation * @prio: Determines how hard direct compaction should try to succeed - * @capture: Pointer to free page created by compaction will be stored here + * @capc: Free page capture bypassing the freelist * * This is the main entry point for direct page compaction. */ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int orde= r, unsigned int alloc_flags, const struct alloc_context *ac, - enum compact_priority prio, struct page **capture) + enum compact_priority prio, struct capture_control *capc) { struct zoneref *z; struct zone *zone; @@ -2893,8 +2862,17 @@ enum compact_result try_to_compact_pages(gfp_t gfp_m= ask, unsigned int order, continue; } =20 + WRITE_ONCE(capc->zone, zone); + status =3D compact_zone_order(zone, order, gfp_mask, prio, - alloc_flags, ac->highest_zoneidx, capture); + alloc_flags, ac->highest_zoneidx, capc); + + WRITE_ONCE(capc->zone, NULL); + + /* Stop if a page has been captured */ + if (READ_ONCE(capc->page)) + status =3D COMPACT_SUCCESS; + rc =3D max(status, rc); =20 /* The allocation should succeed, stop compacting */ diff --git a/mm/internal.h b/mm/internal.h index 181e79f1d6a2..6d3402001b93 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1059,7 +1059,9 @@ struct compact_control { * immediately when one is created during the free path. */ struct capture_control { - struct compact_control *cc; + struct zone *zone; + int migratetype; + int order; struct page *page; }; =20 diff --git a/mm/page_alloc.c b/mm/page_alloc.c index cb422505c6ef..4b9686fedb71 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -721,14 +721,14 @@ static inline struct capture_control *task_capc(struc= t zone *zone) return unlikely(capc) && !(current->flags & PF_KTHREAD) && !capc->page && - capc->cc->zone =3D=3D zone ? capc : NULL; + capc->zone =3D=3D zone ? capc : NULL; } =20 static inline bool compaction_capture(struct capture_control *capc, struct page *page, int order, int migratetype) { - if (!capc || order !=3D capc->cc->order) + if (!capc || order !=3D capc->order) return false; =20 /* Do not accidentally pollute CMA or isolated regions*/ @@ -744,12 +744,12 @@ compaction_capture(struct capture_control *capc, stru= ct page *page, * have trouble finding a high-order free page. */ if (order < pageblock_order && migratetype =3D=3D MIGRATE_MOVABLE && - capc->cc->migratetype !=3D MIGRATE_MOVABLE) + capc->migratetype !=3D MIGRATE_MOVABLE) return false; =20 - if (migratetype !=3D capc->cc->migratetype) - trace_mm_page_alloc_extfrag(page, capc->cc->order, order, - capc->cc->migratetype, migratetype); + if (migratetype !=3D capc->migratetype) + trace_mm_page_alloc_extfrag(page, capc->order, order, + capc->migratetype, migratetype); =20 capc->page =3D page; return true; @@ -4146,6 +4146,12 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigne= d int order, struct page *page =3D NULL; unsigned long pflags; unsigned int noreclaim_flag; + struct capture_control capc =3D { + .zone =3D NULL, + .migratetype =3D ac->migratetype, + .order =3D order, + .page =3D NULL, + }; =20 if (!order) return NULL; @@ -4155,8 +4161,33 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigne= d int order, fs_reclaim_acquire(gfp_mask); noreclaim_flag =3D memalloc_noreclaim_save(); =20 + /* + * Make sure the structs are really initialized before we expose the + * capture control, in case we are interrupted and the interrupt handler + * frees a page. + */ + barrier(); + WRITE_ONCE(current->capture_control, &capc); + *compact_result =3D try_to_compact_pages(gfp_mask, order, alloc_flags, ac, - prio, &page); + prio, &capc); + + /* + * Make sure we hide capture control first before we read the captured + * page pointer, otherwise an interrupt could free and capture a page + * and we would leak it. + */ + WRITE_ONCE(current->capture_control, NULL); + page =3D READ_ONCE(capc.page); + + /* + * Technically, it is also possible that compaction is skipped but + * the page is still captured out of luck(IRQ came and freed the page). + * Returning COMPACT_SUCCESS in such cases helps in properly accounting + * the COMPACT[STALL|FAIL] when compaction is skipped. + */ + if (page) + *compact_result =3D COMPACT_SUCCESS; =20 memalloc_noreclaim_restore(noreclaim_flag); fs_reclaim_release(gfp_mask); --=20 2.55.0 From nobody Fri Jul 24 22:54:56 2026 Received: from mail-qv1-f48.google.com (mail-qv1-f48.google.com [209.85.219.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4C52B37DE84 for ; Wed, 22 Jul 2026 15:00:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.48 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784732424; cv=none; b=rYBPK7C4ZpAmNr+odKBXMHNoC+IrFQFbv7tbtsS11PHaRIGX65QExbsquHULk+WJ4kiKfcQ+GaCu9is0RP7Q+3Exan868nLDaMfOaEdYWHul+iwp78aafUCP8LPOE7abEsFyOnKOJr0uu5Hsg64aNE45XQnaTKbGpaKhd6zuMSI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784732424; c=relaxed/simple; bh=QKg5YGY/yCLwZN3vmT5c5UttyqkYjSANMRTryN3kfVI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=LCbtyeArtw65RoBjxHHBTqy/JoIDFiww0sqUpIn1T7Wolao05aXCsGIwsECM0i5PY2XB+NFJ+ea6bEdAAmCbLou7snb1Ub1HzWJqUQQ1JSqnlsWrCL497V6IsKm4pKLv7kJ3zK4uifIVFkhhk2kQi7vb3bPJLuqa/RTX6vFuUCo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=YquYS/A8; arc=none smtp.client-ip=209.85.219.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="YquYS/A8" Received: by mail-qv1-f48.google.com with SMTP id 6a1803df08f44-8ff20870ac7so102871106d6.1 for ; Wed, 22 Jul 2026 08:00:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1784732422; x=1785337222; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ssUPrZXhmKEtcaomL/NSzb6q5Qkj8exc1H0XDUvt4yE=; b=YquYS/A81i4NStxAz2sixHxEV0vy/5FSbqv9AzkNe9wO41B/Q14lKBPu5zR6toHgVl dcLMhkU/wYun0sQKnBOSM3sSiz7G1NBONgJZDRdueIaC43Qir3PZj6ZEdk6lDW26IETh WltWVga59seLsifIWTqK6+yHhMBVRqaZbGjEhWaQmIjh4Z/sMTjlr3wdCcirq4JOPlkW A+4lye9IOCWnVYAs/4hfhEZcOd6E9QnmF5K4ddc7Unr2wfA8ezMpg6LU+sh1pFVEoFlG oE4/HiAFPIb+InIxP9GKbEr3fKNaJxemnET9GMMJdHDGE8Ht4aflCjPF/OEbL5zcAGCB E0Nw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784732422; x=1785337222; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=ssUPrZXhmKEtcaomL/NSzb6q5Qkj8exc1H0XDUvt4yE=; b=VLBYw9zu/ckT4tGJVaPtAhDb37R5dZsVsLjnQvYOYrVNXROGBoKsGy1gm5L+pzShG+ nvc9jO+RnxyNaQUsjBkX6MfAx5s5N6iCLE0d6RYvxWeTxK0zcz2s7ZhURs/cvTy/OAA9 OirNk/uHucMBkBzOClBIv+cAJ5G4JkfyypDq0eBIQlIVgAKRmLzbcka4FM0ubrfGcqNC wu7+MfKIlh3Ef4vRoiFwzDRfTam/Liz8BOi2O31t0VERNsqYshuvP7lHOhVF7naDhpbC GMXSLsq7ik+4H8Et9hFyNGoS9ShJkI1VDRlXlkJmAS7IeR1+KdPNMxY4J92RRb/vYe15 tY3g== X-Forwarded-Encrypted: i=1; AHgh+RpczDZ4xdRlTYxUjMQcx32N7FRGZ9z5N/ifua1KttaRDaoZZofs7iTied8o/5LpFdkET1Zgqepg+EkB88c=@vger.kernel.org X-Gm-Message-State: AOJu0Yy4XFSBjGKLKLwlr2R3EqixwXu6iGkxJqBTWUkCXqW4y5rg9upQ 3KrpNJ3bKH6ZPO8MAQ5CgnRpQpFL+pCzmB/xrvd1F5T6jXZTDaOEmC/tJWq2v461zZQ= X-Gm-Gg: AR+sD12V7PD92NM+hmcfH4JR8GNO06uupykqrcHU8gekbMk7O7cWK/PFr66OLQF6c78 3hUP+KpZ/ygE8hQDs7f9pK6UoVy5C4WnB8TwRi7z4TXeaRNX3PdfdWPpDKi348SjjGNsnxFGyEG Kl81WhjK6rzin2pXC5PujLFRS4UhGsyzNRQbmhsRrJxRMRQw857D94TAIQO/j9J1zxY4N7jZJz3 iISuzyQwNkC3mcdI2YnLUhvUPekVvGQ7UCIyfpyEA/GEpYNnE+2Nsy/MXcQClhP2tUI1yzI5fi+ dfd4rtIZ45HAOY5XDD5NjWageGdkGaZhWmpIJQQJUjyaI0KlFSI0ZCU1QIy4ojUdhP+ZcySCbcv IMDGaNj8Y7n5h+dxMKxr2f6kN7oShVhpRd9OydM+Xhhfdk50FLCVcTs5MBGEFOn4YGZJkEtsQ8f Y2 X-Received: by 2002:a05:6214:3912:b0:8f1:323a:eede with SMTP id 6a1803df08f44-9077835ab03mr252766896d6.14.1784732418083; Wed, 22 Jul 2026 08:00:18 -0700 (PDT) Received: from localhost ([2603:7001:f100:500:365a:60ff:fe62:ff29]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-907baa0a146sm22359506d6.33.2026.07.22.08.00.17 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 08:00:17 -0700 (PDT) From: Johannes Weiner To: Andrew Morton , Vlastimil Babka Cc: Suren Baghdasaryan , Michal Hocko , Brendan Jackman , Zi Yan , David Hildenbrand , Lorenzo Stoakes , "Liam R . Howlett" , Mike Rapoport , Shakeel Butt , linux-mm@kvack.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: [PATCH v2 4/4] mm: page_alloc: fix non-movable reclaim storm in defrag_mode Date: Wed, 22 Jul 2026 10:56:47 -0400 Message-ID: <20260722150006.3848560-5-hannes@cmpxchg.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260722150006.3848560-1-hannes@cmpxchg.org> References: <20260722150006.3848560-1-hannes@cmpxchg.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" As we deployed defrag_mode into Meta production, pressure spikes and excessive swapping were observed on some workloads. Tracing confirmed that this is unmovable/reclaimable requests spinning in the allocator and direct reclaim, causing excessive amounts of swap. The initial plan for defrag_mode was to rely on kswapd/kcompactd to produce blocks, and if those are overwhelmed under high pressure, let the allocator fall back (__rmqueue_steal()) after its retry loops. However, that retrying results in more reclaim on some of these workloads than we'd hoped, sometimes excessively so, spurred on by the !costly order conditions in should_reclaim_retry(). The storms are dependent on the request type. Reclaim will inevitably make room in existing movable blocks, since that's where the LRU pages live. So if movable requests retry on reclaim, they make progress. When non-movable requests spin in reclaim that isn't productive. They cannot use the individually freed pages, and the process is unlikely to accidentally free whole blocks to meet the ALLOC_NOFRAGMENT bar. They spin and overreclaim excessively, which tanks performance and triggers userspace guards like swap exhaustion or pressure based OOM. To fix this, send non-movable requests, regardless of order, into pageblock reclaim/compaction. This way, they help move things along to meet the ALLOC_NOFRAGMENT bar. After this patch, the reclaim storms and excess OOM rates are no longer observed in production. The longer-term plan is still to have all requests, including the movable ones, help make blocks to spread the cost of defragmenting more evenly and fairly; combined with proper watermarking to reduce allocation latencies in the common case. However, doing this naively unearths scaling and concurrency limitations in compaction that need to be addressed first. Promoting just non-movables for now is the minimally viable bug fix for the above issue. Fixes: e3aa7df331bc ("mm: page_alloc: defrag_mode") Cc: Signed-off-by: Johannes Weiner Reviewed-by: Vlastimil Babka (SUSE) --- mm/internal.h | 6 ++++++ mm/page_alloc.c | 31 ++++++++++++++++++++++++++----- 2 files changed, 32 insertions(+), 5 deletions(-) diff --git a/mm/internal.h b/mm/internal.h index 6d3402001b93..5acba3470659 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1061,6 +1061,12 @@ struct compact_control { struct capture_control { struct zone *zone; int migratetype; + /* + * Allocation request order. May differ from the compaction + * order: defrag_mode promotes sub-block allocations to + * pageblock-order compaction; capture still matches at the + * original allocation order so prep_new_page() is consistent. + */ int order; struct page *page; }; diff --git a/mm/page_alloc.c b/mm/page_alloc.c index 4b9686fedb71..f92055827ae9 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -4152,8 +4152,24 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigne= d int order, .order =3D order, .page =3D NULL, }; + int compact_order =3D order; =20 - if (!order) + /* + * If fallbacks are not permitted (defrag_mode), we either + * need to reclaim space in a block of matching type, or clear + * out an entire block to allow __rmqueue_claim() to convert. + * + * Reclaim by itself is primarily freeing space in movable + * blocks, since that's where the LRU pages live. So this + * works for movable requests, but not for others. + * + * For those, promote the order to help make blocks, instead + * of spinning in reclaim alone unproductively. + */ + if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype !=3D MIGRATE_MOVA= BLE) + compact_order =3D max(order, pageblock_order); + + if (!compact_order) return NULL; =20 psi_memstall_enter(&pflags); @@ -4169,8 +4185,8 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned= int order, barrier(); WRITE_ONCE(current->capture_control, &capc); =20 - *compact_result =3D try_to_compact_pages(gfp_mask, order, alloc_flags, ac, - prio, &capc); + *compact_result =3D try_to_compact_pages(gfp_mask, compact_order, + alloc_flags, ac, prio, &capc); =20 /* * Make sure we hide capture control first before we read the captured @@ -4215,7 +4231,7 @@ __alloc_pages_direct_compact(gfp_t gfp_mask, unsigned= int order, struct zone *zone =3D page_zone(page); =20 zone->compact_blockskip_flush =3D false; - compaction_defer_reset(zone, order, true); + compaction_defer_reset(zone, compact_order, true); count_vm_event(COMPACTSUCCESS); return page; } @@ -4455,9 +4471,14 @@ __alloc_pages_direct_reclaim(gfp_t gfp_mask, unsigne= d int order, struct page *page =3D NULL; unsigned long pflags; bool drained =3D false; + int reclaim_order =3D order; + + /* Match the slowpath compaction promotion in __alloc_pages_direct_compac= t */ + if ((alloc_flags & ALLOC_NOFRAGMENT) && ac->migratetype !=3D MIGRATE_MOVA= BLE) + reclaim_order =3D max(order, pageblock_order); =20 psi_memstall_enter(&pflags); - *did_some_progress =3D __perform_reclaim(gfp_mask, order, ac); + *did_some_progress =3D __perform_reclaim(gfp_mask, reclaim_order, ac); if (unlikely(!(*did_some_progress))) goto out; =20 --=20 2.55.0