From nobody Tue Sep 29 07:39:37 2026 Received: from mail-pf1-f199.google.com (mail-pf1-f199.google.com [209.85.210.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9977737CD5A for ; Mon, 10 Aug 2026 23:08:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786403328; cv=none; b=AM04D7ARYzYCh3Q0xyy4yYqT2ZRhn2Kq++ODBXZ8DE9vSstVg6plQJx1XEpQv1gJepgCPtHI1Pp/ws92QfBZ41aRQzwn/XDSShI6Gr2Sp+tERQ0WrrUdRpUnu/zoCgge7SouweFIYXEjoY2FxShzO/6cPGKxAnGbq/2qcJbrdnA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786403328; c=relaxed/simple; bh=OF56ZlCi2CSMBEVryreRjRAoKwZjdUoi8HiqG2WhZEs=; h=Date:Mime-Version:Message-ID:Subject:From:To:Cc:Content-Type; b=M1BXjEwJQWB11qcvzizGF1r+lPMoQSQh6CY3CfZ0UuoqLVpPdp/Hb77u1arGBFGQNXuZcsmAylAPVCIwAY1zNhK5xeGlfY/THIGpzRU+UuYAJ6AkXKNxYWFWRZ7Q5RhSuhydmo9RElXkd8XMP6ketwGkLNgxZbeGzLDswVGqBfg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--souravpanda.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=n2xGZcWx; arc=none smtp.client-ip=209.85.210.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--souravpanda.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="n2xGZcWx" Received: by mail-pf1-f199.google.com with SMTP id d2e1a72fcca58-84e4ef486c7so314126b3a.1 for ; Mon, 10 Aug 2026 16:08:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786403326; x=1787008126; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:mime-version:date:from :to:cc:subject:date:message-id:reply-to:content-type; bh=ItB67UhN5xHbGpjuS6XYU+rFxP6+nTJipcuCpIdU/ZU=; b=n2xGZcWxQ/kh4HU5QJFOg6kPMXWvtr/zhxFgxWDy1tyy8qdGNILQDb0wzHvlkY4MpJ cNmKHd490Clv/+VeMUmbSTwsVQD0dB6crXV2u0W2b9PUF1JNGDWYYC0NYIA2v1bHLOlH L1Z32vuwN/iUaAYJ6RkJGWsZYDdmoXKwfydx/HUeBbEJQW27GlXzvYa6SSBryL8ima9o 59VK/oIDqa3Eovy3EMRQeRmttXB5vqhRhKxDUQzzCv1gobAxaYghOCCiZ+xWYeKjueSr daPHQPNdjpgnTAHQsDyZZPV+xyygNMQcAQbYqv3GkCwVbsOFpNj9KSFGaOf2LqRwXW3v dZ+g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786403326; x=1787008126; h=content-type:cc:to:from:subject:message-id:mime-version:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=ItB67UhN5xHbGpjuS6XYU+rFxP6+nTJipcuCpIdU/ZU=; b=ZLSHO9JFikc1mdp5j3+1uIcLYX1ceeD0SKxboPUj2WH3uCgWWkYxlwvnVNnpwYK5vW 4CF26DlehatcStEFBMPfoqJYsUSp68Vv9F+QFpJIGZ0Ys8/pI1LoYPPkygyxIGMa0QFi uOZAAOUtp7niTWG6RYZzZsQP8LG/u6K/2ustQzucyGDkH4IdCd5z0JKNAgvl7y3VPaLr 3MfMtJW7ggZjrdXhtmc1TRJcegTeAF0lIExvHQD7uin4Z58KFbfXkIl2HEL7TgMct3pF jNAuleoUZSuATFjXelVsWEujumICa1ryOUOAhfHv9BY3WQQH3vkb3d24NFVbOf6PxKnq Dlsw== X-Forwarded-Encrypted: i=1; AHgh+Rpo1Ci9Qw5d2762cE4r3aQZVawHXPoXk+WdW1laIvk1lagC76TVggzL8QfYg3tp0sC2UsWbqTDBIVlvcjs=@vger.kernel.org X-Gm-Message-State: AOJu0YyySUx+wH3REAmTatGbx+6Plg7VFXlA8QdyKIkSxhaFkf9n7VOv MzfX/Ownz4KyBhs/OjZ1JGlKe6/pI+7PqpevMWapJNqcwK0QwhJzzYYknhZg/UOtFZ/vOIU5rn9 T6r69BX9T7JK7we47Js4iWSGJHg== X-Received: from pgng7.prod.google.com ([2002:a63:3747:0:b0:cbe:afd1:613e]) (user=souravpanda job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:1485:b0:847:9dd7:2683 with SMTP id d2e1a72fcca58-84fa1851e75mr1737003b3a.9.1786403325561; Mon, 10 Aug 2026 16:08:45 -0700 (PDT) Date: Mon, 10 Aug 2026 23:08:44 +0000 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 X-Mailer: git-send-email 2.55.0.679.g6767b8d81c-goog Message-ID: <20260810230844.3778931-1-souravpanda@google.com> Subject: [PATCH v6] mm/hugetlb_cma: Fix null nodemask dereference in hugetlb_cma_alloc_frozen_folio From: Sourav Panda To: muchun.song@linux.dev, osalvador@suse.de, akpm@linux-foundation.org Cc: usama.arif@linux.dev, shakeel.butt@linux.dev, wangkefeng.wang@huawei.com, anshuman.khandual@arm.com, david@kernel.org, surenb@google.com, fvdl@google.com, gthelen@google.com, hannes@cmpxchg.org, riel@surriel.com, sj@kernel.org, vbabka@suse.cz, mhocko@suse.com, bjackman@google.com, zi.yan@sent.com, souravpanda@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" alloc_buddy_hugetlb_folio_with_mpol() can pass a NULL nodemask to alloc_fresh_hugetlb_folio() as a fallback to allocate from all nodes. If order is gigantic, alloc_fresh_hugetlb_folio() propagates the NULL nodemask down to hugetlb_cma_alloc_frozen_folio() via alloc_gigantic_frozen_folio(). Additionally, hugetlb_cma_alloc_frozen_folio() previously attempted allocation on hugetlb_cma[nid] without verifying if nid is included in the caller's nodemask. Adding a node_isset(nid, *nodemask) check ensures the initial preferred node allocation honors the memory policy / nodemask. However, hugetlb_cma_alloc_frozen_folio() dereferences the nodemask in node_isset(nid, *nodemask) and for_each_node_mask(node, *nodemask), leading to a null pointer dereference kernel panic when nodemask is NULL. Fix this by checking if nodemask is NULL in hugetlb_cma_alloc_frozen_folio() and defaulting it to cpuset_current_mems_allowed. Enclose the allocation attempts within the cpuset seqcount retry loop so that if the cpuset changes concurrently during allocation, the attempts are retried using the updated nodemask. This ensures that the initial node check and fallback loop safely honor the task's cpuset without violating cpuset constraints or causing NULL pointer dereferences or unexpected allocation failures. From a userspace perspective, this bug allows an unprivileged user to crash the kernel (trigger a panic) by requesting a gigantic hugepage allocation with MPOL_PREFERRED_MANY on a system where CMA is only configured on a subset of NUMA nodes. This can be reproduced by booting a VM with two NUMA nodes, restricting CMA to Node 1 (e.g., hugetlb_cma=3D1:1G default_hugepagesz=3D1G hugepagesz=3D1G hugepages=3D0), and running a program that allocates a 1GB hugepage area without reserving, restricts allocation to Node 0 using mbind() with MPOL_PREFERRED_MANY, and triggers a page fault: void *ptr =3D mmap(NULL, 1UL << 30, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB | MAP_HUGE_1GB | MAP_NORESERVE, -1, 0); unsigned long nodemask =3D 1; /* Node 0 */ mbind(ptr, 1UL << 30, MPOL_PREFERRED_MANY, &nodemask, sizeof(nodemask) * 8, 0); memset(ptr, 0, 1UL << 30); /* Trigger fault */ This results in a NULL pointer dereference: BUG: kernel NULL pointer dereference, address: 0000000000000000 #PF: supervisor read access in kernel mode #PF: error_code(0x0000) - not-present page Oops: Oops: 0000 [#1] SMP NOPTI RIP: 0010:hugetlb_cma_alloc_frozen_folio+0x75/0x120 Call Trace: only_alloc_fresh_hugetlb_folio.isra.0+0x2c/0x160 alloc_surplus_hugetlb_folio+0x6d/0x100 alloc_hugetlb_folio+0x3c5/0x660 hugetlb_no_page+0x3d9/0x650 Fixes: eb02f14c4a2b ("mm/hugetlb: allow overcommitting gigantic hugepages") Cc: stable@vger.kernel.org Signed-off-by: Sourav Panda --- Changes in v6: - Enclosed the CMA allocation attempts within the cpuset seqcount retry loo= p, retrying allocation upon cpuset mems_allowed updates to prevent unexpected allocation failures as suggested by Muchun Song. - v5: https://lore.kernel.org/linux-mm/20260809043250.2917406-1-souravpanda= @google.com/ - v4: https://lore.kernel.org/linux-mm/20260726072935.3513996-1-souravpanda= @google.com/ - v3: https://lore.kernel.org/linux-mm/20260705175119.440599-1-souravpanda@= google.com/ - v2: https://lore.kernel.org/linux-mm/20260704174930.2885785-1-souravpanda= @google.com/ - v1: https://lore.kernel.org/linux-mm/20260702215713.627941-1-souravpanda@= google.com/ mm/hugetlb_cma.c | 23 ++++++++++++++++++----- 1 file changed, 18 insertions(+), 5 deletions(-) diff --git a/mm/hugetlb_cma.c b/mm/hugetlb_cma.c index 39344d6c78d8..9debf033d4fd 100644 --- a/mm/hugetlb_cma.c +++ b/mm/hugetlb_cma.c @@ -3,6 +3,7 @@ #include #include #include +#include #include =20 #include @@ -30,15 +31,27 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order,= gfp_t gfp_mask, int node; struct folio *folio; struct page *page =3D NULL; + const nodemask_t *nmask; + nodemask_t local_node_mask; + unsigned int cpuset_mems_cookie; =20 if (!hugetlb_cma_size) return NULL; =20 - if (hugetlb_cma[nid]) +retry_cpuset: + if (!nodemask) { + cpuset_mems_cookie =3D read_mems_allowed_begin(); + local_node_mask =3D cpuset_current_mems_allowed; + nmask =3D &local_node_mask; + } else { + nmask =3D nodemask; + } + + if (hugetlb_cma[nid] && node_isset(nid, *nmask)) page =3D cma_alloc_frozen_compound(hugetlb_cma[nid], order); =20 if (!page && !(gfp_mask & __GFP_THISNODE)) { - for_each_node_mask(node, *nodemask) { + for_each_node_mask(node, *nmask) { if (node =3D=3D nid || !hugetlb_cma[node]) continue; =20 @@ -48,8 +61,12 @@ struct folio *hugetlb_cma_alloc_frozen_folio(int order, = gfp_t gfp_mask, } } =20 - if (!page) + if (!page) { + if (!nodemask && + unlikely(read_mems_allowed_retry(cpuset_mems_cookie))) + goto retry_cpuset; return NULL; + } =20 folio =3D page_folio(page); folio_set_hugetlb_cma(folio); --=20 2.55.0