From nobody Mon Feb 9 19:30:54 2026 Received: from mail-pj1-f74.google.com (mail-pj1-f74.google.com [209.85.216.74]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 952771DE4CD for ; Mon, 27 Jan 2025 23:22:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.74 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1738020159; cv=none; b=vGNAHqP9D7nPX229aWiqGbBaBCXePVRgxyT37HQA7wDYj261YNSNAt72cN9HjM7eakaIx0uH9z0ZVRBhLZku+aVol8SW5Bzz8+sBbHkOHBYSxL5ceMZL3rnEa3IGUmHCfnsGA2xgcEN8UJSNmQgg4nkCxfhuydV4sYKRigXvBxU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1738020159; c=relaxed/simple; bh=xPMmll6gntzOHQsFpuklKUBoh66rniQVtsXPPHWHupU=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=RTp0amh82zdvuP/GD0J2zfZ0UcYW1l9IG44666lVV+66IY86myv+YrNFJ1R2q7RLzR1ym3Rbl6Kl2hhJxEL+nIWy+ghst1635Gjx2LYX0eDJrY/nyxid7r/DkVCNtxLL+T1rrviZQ/+iZE2cdyx7dS3r5xfQQed6uZb0Vru7mus= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--fvdl.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=oUwqogea; arc=none smtp.client-ip=209.85.216.74 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--fvdl.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="oUwqogea" Received: by mail-pj1-f74.google.com with SMTP id 98e67ed59e1d1-2ef6ef86607so11126168a91.0 for ; Mon, 27 Jan 2025 15:22:37 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1738020157; x=1738624957; darn=vger.kernel.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=sEk7wWJBoXYCWMM5+k1sYt6L/V24bQrmW4i63Y+H+6w=; b=oUwqogeaJs+I+qsnWlYfwxY14GzT3y9sJYwU3kOfSBk5EAl9c014NYIkXa0FQ4aAnS zWG2kU/tsASOThcdGc0/UCWOWlCqGSloUnjSb8vnbxr1qZxr8s6p67lTbo8+WX3Z5MQ8 6QYEc74jqAgFIjsMjCtXn46nG1NVRpbgP5NM6s/Qk16Aa0v72hRw8yGOmQTV3aLm3ay0 V4qGat2iQjX/kaelxMfqV8LWUzsdPhSOsbquDqzjbXdwxxtmsVpizOuFooBUPvXEr6K7 5Ov1IBb35+uRtLQ58msMbVhiI+B8pET+Gl/aUFI0XfaztBq9JM6J4MSZD+lHPPfeyQ88 JBHA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1738020157; x=1738624957; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=sEk7wWJBoXYCWMM5+k1sYt6L/V24bQrmW4i63Y+H+6w=; b=TCtRswVZfpF9+4BciICI6mf56EsQ0jqPVyJbyxawmps7n+tk0p0eW/i61JCeB/xD3H gpdUXrMVBQ+5nks0keLrvVX+9/uokP6v3ajGNDuEIVAs2dP2P0si/iOXJpZtnmfIWTG4 rfl4hQ2yrna0dwXkk0YU5mUvoXW2EDL2nN6HU22J/5vS+DYWXT+GDoca/dSUA81DfxU2 pOl5uEmaJh0QigQ+AItm3OClyhzaUojc0f37R+1k7+kAxoWugQuKZmjMNyxs91xz6GIh WTomy+98hOjFQAd2wCeAMSGxBxbZCZRE9F03tCEb9I4BlA93uihfhpylgQU3ZEd4zknd n2JQ== X-Forwarded-Encrypted: i=1; AJvYcCVrnau4IjxH0HaWnG4fhaSus/RsHQzBwOjFYyqB4albnGfzE6EdfoaEz5aT+yBbZU2w19WmTPhC90BMbT0=@vger.kernel.org X-Gm-Message-State: AOJu0YxBlS6j2gywm8TkU0H6pAwxuAeV1PfrILNKt6sehJ7a+/IzFjF5 TGvCAG1WZi6cm7vFl3Zz4bxqBYfq5MyoxcUuKXHSJ+IKu8jSqGLTVS4aVTrKjbC2QxHITg== X-Google-Smtp-Source: AGHT+IHfBibcmt8vtXZ+mym7NU2SB0DPkxQSWz4/Xq+ZhKrlvbhzuCjqjkBbya8StVC/rWEyywcCYFQJ X-Received: from pfblh4.prod.google.com ([2002:a05:6a00:7104:b0:72f:bfd9:14c5]) (user=fvdl job=prod-delivery.src-stubby-dispatcher) by 2002:aa7:88d2:0:b0:729:1c0f:b94e with SMTP id d2e1a72fcca58-72fc0917f6bmr1506666b3a.6.1738020156872; Mon, 27 Jan 2025 15:22:36 -0800 (PST) Date: Mon, 27 Jan 2025 23:21:48 +0000 In-Reply-To: <20250127232207.3888640-1-fvdl@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20250127232207.3888640-1-fvdl@google.com> X-Mailer: git-send-email 2.48.1.262.g85cc9f2d1e-goog Message-ID: <20250127232207.3888640-9-fvdl@google.com> Subject: [PATCH 08/27] mm/hugetlb: convert cmdline parameters from setup to early From: Frank van der Linden To: akpm@linux-foundation.org, muchun.song@linux.dev, linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: yuzhao@google.com, usama.arif@bytedance.com, joao.m.martins@oracle.com, roman.gushchin@linux.dev, Frank van der Linden Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Convert the cmdline parameters (hugepagesz, hugepages, default_hugepagesz and hugetlb_free_vmemmap) to early parameters. Since parse_early_param might run before MMU setups on some platforms (powerpc), validation of huge page sizes as specified in command line parameters would fail. So instead, for the hstate-related values, just record the them and parse them on demand, from hugetlb_bootmem_alloc. The allocation of hugetlb bootmem pages is now done in hugetlb_bootmem_alloc, which is called explicitly at the start of mm_core_init(). core_initcall would be too late, as that happens with memblock already torn down. This change will allow earlier allocation and initialization of bootmem hugetlb pages later on. No functional change intended. Signed-off-by: Frank van der Linden --- include/linux/hugetlb.h | 6 ++ mm/hugetlb.c | 133 +++++++++++++++++++++++++++++++--------- mm/hugetlb_vmemmap.c | 6 +- mm/mm_init.c | 3 + 4 files changed, 119 insertions(+), 29 deletions(-) diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h index ec8c0ccc8f95..9cd7c9dacb88 100644 --- a/include/linux/hugetlb.h +++ b/include/linux/hugetlb.h @@ -174,6 +174,8 @@ struct address_space *hugetlb_folio_mapping_lock_write(= struct folio *folio); extern int sysctl_hugetlb_shm_group; extern struct list_head huge_boot_pages[MAX_NUMNODES]; =20 +void hugetlb_bootmem_alloc(void); + /* arch callbacks */ =20 #ifndef CONFIG_HIGHPTE @@ -1250,6 +1252,10 @@ static inline bool hugetlbfs_pagecache_present( { return false; } + +static inline void hugetlb_bootmem_alloc(void) +{ +} #endif /* CONFIG_HUGETLB_PAGE */ =20 static inline spinlock_t *huge_pte_lock(struct hstate *h, diff --git a/mm/hugetlb.c b/mm/hugetlb.c index a67339ca65b4..a95ab44d5545 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -40,6 +40,7 @@ #include #include #include +#include =20 #include #include @@ -62,6 +63,24 @@ static unsigned long hugetlb_cma_size __initdata; =20 __initdata struct list_head huge_boot_pages[MAX_NUMNODES]; =20 +/* + * Due to ordering constraints across the init code for various + * architectures, hugetlb hstate cmdline parameters can't simply + * be early_param. early_param might call the setup function + * before valid hugetlb page sizes are determined, leading to + * incorrect rejection of valid hugepagesz=3D options. + * + * So, record the parameters early and consume them whenever the + * init code is ready for them, by calling hugetlb_parse_params(). + */ + +/* one (hugepagesz=3D,hugepages=3D) pair per hstate, one default_hugepages= z */ +#define HUGE_MAX_CMDLINE_ARGS (2 * HUGE_MAX_HSTATE + 1) +struct hugetlb_cmdline { + char *val; + int (*setup)(char *val); +}; + /* for command line parsing */ static struct hstate * __initdata parsed_hstate; static unsigned long __initdata default_hstate_max_huge_pages; @@ -69,6 +88,20 @@ static bool __initdata parsed_valid_hugepagesz =3D true; static bool __initdata parsed_default_hugepagesz; static unsigned int default_hugepages_in_node[MAX_NUMNODES] __initdata; =20 +static char hstate_cmdline_buf[COMMAND_LINE_SIZE] __initdata; +static int hstate_cmdline_index __initdata; +static struct hugetlb_cmdline hugetlb_params[HUGE_MAX_CMDLINE_ARGS] __init= data; +static int hugetlb_param_index __initdata; +static __init int hugetlb_add_param(char *s, int (*setup)(char *val)); +static __init void hugetlb_parse_params(void); + +#define hugetlb_early_param(str, func) \ +static __init int func##args(char *s) \ +{ \ + return hugetlb_add_param(s, func); \ +} \ +early_param(str, func##args) + /* * Protects updates to hugepage_freelists, hugepage_activelist, nr_huge_pa= ges, * free_huge_pages, and surplus_huge_pages. @@ -3488,6 +3521,8 @@ static void __init hugetlb_hstate_alloc_pages(struct = hstate *h) =20 for (i =3D 0; i < MAX_NUMNODES; i++) INIT_LIST_HEAD(&huge_boot_pages[i]); + h->next_nid_to_alloc =3D first_online_node; + h->next_nid_to_free =3D first_online_node; initialized =3D true; } =20 @@ -4550,8 +4585,6 @@ void __init hugetlb_add_hstate(unsigned int order) for (i =3D 0; i < MAX_NUMNODES; ++i) INIT_LIST_HEAD(&h->hugepage_freelists[i]); INIT_LIST_HEAD(&h->hugepage_activelist); - h->next_nid_to_alloc =3D first_online_node; - h->next_nid_to_free =3D first_online_node; snprintf(h->name, HSTATE_NAME_LEN, "hugepages-%lukB", huge_page_size(h)/SZ_1K); =20 @@ -4576,6 +4609,42 @@ static void __init hugepages_clear_pages_in_node(voi= d) } } =20 +static __init int hugetlb_add_param(char *s, int (*setup)(char *)) +{ + size_t len; + char *p; + + if (hugetlb_param_index >=3D HUGE_MAX_CMDLINE_ARGS) + return -EINVAL; + + len =3D strlen(s) + 1; + if (len + hstate_cmdline_index > sizeof(hstate_cmdline_buf)) + return -EINVAL; + + p =3D &hstate_cmdline_buf[hstate_cmdline_index]; + memcpy(p, s, len); + hstate_cmdline_index +=3D len; + + hugetlb_params[hugetlb_param_index].val =3D p; + hugetlb_params[hugetlb_param_index].setup =3D setup; + + hugetlb_param_index++; + + return 0; +} + +static __init void hugetlb_parse_params(void) +{ + int i; + struct hugetlb_cmdline *hcp; + + for (i =3D 0; i < hugetlb_param_index; i++) { + hcp =3D &hugetlb_params[i]; + + hcp->setup(hcp->val); + } +} + /* * hugepages command line processing * hugepages normally follows a valid hugepagsz or default_hugepagsz @@ -4595,7 +4664,7 @@ static int __init hugepages_setup(char *s) if (!parsed_valid_hugepagesz) { pr_warn("HugeTLB: hugepages=3D%s does not follow a valid hugepagesz, ign= oring\n", s); parsed_valid_hugepagesz =3D true; - return 1; + return -EINVAL; } =20 /* @@ -4649,24 +4718,16 @@ static int __init hugepages_setup(char *s) } } =20 - /* - * Global state is always initialized later in hugetlb_init. - * But we need to allocate gigantic hstates here early to still - * use the bootmem allocator. - */ - if (hugetlb_max_hstate && hstate_is_gigantic(parsed_hstate)) - hugetlb_hstate_alloc_pages(parsed_hstate); - last_mhp =3D mhp; =20 - return 1; + return 0; =20 invalid: pr_warn("HugeTLB: Invalid hugepages parameter %s\n", p); hugepages_clear_pages_in_node(); - return 1; + return -EINVAL; } -__setup("hugepages=3D", hugepages_setup); +hugetlb_early_param("hugepages", hugepages_setup); =20 /* * hugepagesz command line processing @@ -4685,7 +4746,7 @@ static int __init hugepagesz_setup(char *s) =20 if (!arch_hugetlb_valid_size(size)) { pr_err("HugeTLB: unsupported hugepagesz=3D%s\n", s); - return 1; + return -EINVAL; } =20 h =3D size_to_hstate(size); @@ -4700,7 +4761,7 @@ static int __init hugepagesz_setup(char *s) if (!parsed_default_hugepagesz || h !=3D &default_hstate || default_hstate.max_huge_pages) { pr_warn("HugeTLB: hugepagesz=3D%s specified twice, ignoring\n", s); - return 1; + return -EINVAL; } =20 /* @@ -4710,14 +4771,14 @@ static int __init hugepagesz_setup(char *s) */ parsed_hstate =3D h; parsed_valid_hugepagesz =3D true; - return 1; + return 0; } =20 hugetlb_add_hstate(ilog2(size) - PAGE_SHIFT); parsed_valid_hugepagesz =3D true; - return 1; + return 0; } -__setup("hugepagesz=3D", hugepagesz_setup); +hugetlb_early_param("hugepagesz", hugepagesz_setup); =20 /* * default_hugepagesz command line input @@ -4731,14 +4792,14 @@ static int __init default_hugepagesz_setup(char *s) parsed_valid_hugepagesz =3D false; if (parsed_default_hugepagesz) { pr_err("HugeTLB: default_hugepagesz previously specified, ignoring %s\n"= , s); - return 1; + return -EINVAL; } =20 size =3D (unsigned long)memparse(s, NULL); =20 if (!arch_hugetlb_valid_size(size)) { pr_err("HugeTLB: unsupported default_hugepagesz=3D%s\n", s); - return 1; + return -EINVAL; } =20 hugetlb_add_hstate(ilog2(size) - PAGE_SHIFT); @@ -4755,17 +4816,33 @@ static int __init default_hugepagesz_setup(char *s) */ if (default_hstate_max_huge_pages) { default_hstate.max_huge_pages =3D default_hstate_max_huge_pages; - for_each_online_node(i) - default_hstate.max_huge_pages_node[i] =3D - default_hugepages_in_node[i]; - if (hstate_is_gigantic(&default_hstate)) - hugetlb_hstate_alloc_pages(&default_hstate); + /* + * Since this is an early parameter, we can't check + * NUMA node state yet, so loop through MAX_NUMNODES. + */ + for (i =3D 0; i < MAX_NUMNODES; i++) { + if (default_hugepages_in_node[i] !=3D 0) + default_hstate.max_huge_pages_node[i] =3D + default_hugepages_in_node[i]; + } default_hstate_max_huge_pages =3D 0; } =20 - return 1; + return 0; +} +hugetlb_early_param("default_hugepagesz", default_hugepagesz_setup); + +void __init hugetlb_bootmem_alloc(void) +{ + struct hstate *h; + + hugetlb_parse_params(); + + for_each_hstate(h) { + if (hstate_is_gigantic(h)) + hugetlb_hstate_alloc_pages(h); + } } -__setup("default_hugepagesz=3D", default_hugepagesz_setup); =20 static unsigned int allowed_mems_nr(struct hstate *h) { diff --git a/mm/hugetlb_vmemmap.c b/mm/hugetlb_vmemmap.c index 57b7f591eee8..326cdf94192e 100644 --- a/mm/hugetlb_vmemmap.c +++ b/mm/hugetlb_vmemmap.c @@ -444,7 +444,11 @@ DEFINE_STATIC_KEY_FALSE(hugetlb_optimize_vmemmap_key); EXPORT_SYMBOL(hugetlb_optimize_vmemmap_key); =20 static bool vmemmap_optimize_enabled =3D IS_ENABLED(CONFIG_HUGETLB_PAGE_OP= TIMIZE_VMEMMAP_DEFAULT_ON); -core_param(hugetlb_free_vmemmap, vmemmap_optimize_enabled, bool, 0); +static int __init hugetlb_vmemmap_optimize_param(char *buf) +{ + return kstrtobool(buf, &vmemmap_optimize_enabled); +} +early_param("hugetlb_free_vmemmap", hugetlb_vmemmap_optimize_param); =20 static int __hugetlb_vmemmap_restore_folio(const struct hstate *h, struct folio *folio, unsigned long flags) diff --git a/mm/mm_init.c b/mm/mm_init.c index 2630cc30147e..d2dee53e95dd 100644 --- a/mm/mm_init.c +++ b/mm/mm_init.c @@ -30,6 +30,7 @@ #include #include #include +#include #include "internal.h" #include "slab.h" #include "shuffle.h" @@ -2641,6 +2642,8 @@ static void __init mem_init_print_info(void) */ void __init mm_core_init(void) { + hugetlb_bootmem_alloc(); + /* Initializations relying on SMP setup */ BUILD_BUG_ON(MAX_ZONELISTS > 2); build_all_zonelists(NULL); --=20 2.48.1.262.g85cc9f2d1e-goog