From nobody Tue Sep 29 04:43:03 2026 Received: from stravinsky.debian.org (stravinsky.debian.org [82.195.75.108]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6A3DE434E50 for ; Wed, 12 Aug 2026 11:32:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=82.195.75.108 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786534338; cv=none; b=uKJ78QxYqwc42jx3D4EnkeoflDbnRDC5DCMhOsxxB/wrNL7crS70tLNHvAT7fug5S8bz2a/3q89EubzkZqSWZZpP87O1f4W2UF4phjKmPoMmCV4dE/2wfaOzB9l3TqYlFFxBsDUHhByVexvFs5A2kxrABK/qZG3W7wuIk2ZPG6s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786534338; c=relaxed/simple; bh=FmvGEkh2rRrjTsnxxpplsZ7+GnyJWiiOZt1q+BoClXI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=PO8aDXOAuGZ3DmKHArcmNx9Tx07HTYce7auk4PkdhZXs8gKvjgKhne9Zu9C7vIy3flb9L95pAO1lNqvEGsC9DmNTQt2JgFY3hPj5VCXoxHP2+QUecF1bSmPJzdfBJIiwZvTAFLvWKh/19wCSFkEpFaywlY7T+0oLiom8LfP1ZTo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org; spf=pass smtp.mailfrom=debian.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b=Bgkzq1hs; arc=none smtp.client-ip=82.195.75.108 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=debian.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b="Bgkzq1hs" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=debian.org; s=smtpauto.stravinsky; h=X-Debian-User:Cc:To:In-Reply-To:References: Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date: From:Reply-To:Content-ID:Content-Description; bh=96LAzBjfGAx9nfZL8SX73XYEA1A5G1phryUkEZQ35/w=; b=Bgkzq1hsvSlzCUSnGj2E/Lk8tY 1dMyAR9AzJbQC4LpK1CNVMN3IMg2XiBBdGGpHkZxiyzZHzvIp3lUvfGekHCX5qso78dVIIz/RZzCU hVvG+mObEJqzW7JMq1qj9/03SAiuonCVJ7bOj/yPU94AS13q/qPSjmQLbpm0nHqqRNDSL9GTfJTki 7cfn0GrGjen4rSZ9S7hRf+NKHoJNEtW8bP6rjSgwUqTSuMZ4Tvjn4LUfe7bBvffBhNfU4VO18ZqA4 jZM4srYTplz+S6iuZ0U3OAwuI7bklGXi8w3NpWly3ODPPkf2aua3U5d1w/Dfozz2IlQpJhZVdTTkj YSjXuH3g==; Received: from authenticated-user by stravinsky.debian.org with esmtpsa (TLS1.3:ECDHE_X25519__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim 4.96) (envelope-from ) id 1wu7Bo-004JJH-16; Wed, 12 Aug 2026 11:32:08 +0000 From: Breno Leitao Date: Wed, 12 Aug 2026 04:31:51 -0700 Subject: [PATCH v6 1/2] kexec_file: stop the top-down search before it underflows Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260812-kexec_posioned-v6-1-e477887086f0@debian.org> References: <20260812-kexec_posioned-v6-0-e477887086f0@debian.org> In-Reply-To: <20260812-kexec_posioned-v6-0-e477887086f0@debian.org> To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Baoquan He , Pasha Tatashin , Pratyush Yadav , Miaohe Lin , Naoya Horiguchi Cc: Breno Leitao , linux-mm@kvack.org, linux-kernel@vger.kernel.org, kexec@lists.infradead.org, rmikey@meta.com, riel@surriel.com, kernel-team@meta.com X-Mailer: b4 0.16-dev-f8e9d X-Developer-Signature: v=1; a=openpgp-sha256; l=1669; i=leitao@debian.org; h=from:subject:message-id; bh=FmvGEkh2rRrjTsnxxpplsZ7+GnyJWiiOZt1q+BoClXI=; b=owEBbQKS/ZANAwAIATWjk5/8eHdtAcsmYgBqfFmsyEno7+ZM8jNDTttS6/P8qtEqZPqoheexT 8D0MQRonQuJAjMEAAEIAB0WIQSshTmm6PRnAspKQ5s1o5Of/Hh3bQUCanxZrAAKCRA1o5Of/Hh3 bXSJD/9jfrIbSU3oHTQv0L6n64XBO1LZMWRmaCcv4mpdb64PYKBuJ1vpA5i8Tqe5xgTSKXXA23c CuwqNJzcw1l3YXFoQsX3nttO24qrXYmd1hXK/jlkiBlgRxi+K3ue/Xmcz9108cNsRdieQmwVc6A h7veNRxUviUn1rvqU4UrdmgY5FZC5phM5rnABwFMz/hOju74BJ2YBer7UB23/KC81X5lPomJxQu DjRMiHU5OgXRDphVIdC1ehkv6d+D4L+kv2WGaJDqJ6qA6QsFScNowu+jY+WIbq+7nO2VY0RooHp IUm06Yv2t0eN2jtT05X0mve/wbxwzrHerIWupYY15hnhalruDg4iPLEv0UER95+Z28F/c7LpfXe 8a6LbuH6JVQSr/lfxxoNSoqSatUe0YEGHWNecjZS85BdNG1pqw3mmN5N/1z9hxNC2KpPXtQXwSe nkAhZCovOV5Lnma3VIV4uFACGpv1O7NjkHsPX/xjEl/saVEWhnZW+Kk4/n577us+SJbLXQW+0Lm N1imHnRQQYBGccSPA8iyjrYOZhhJdUxEbGuISLllYItLB+g5UzJdx8wvHro047EPIzHVqffnTe8 gofvKGUZtDyNj0/gzahda+W8EGVLlfGHHnyUE1bA3xQDQIiiTfONMjN28AA3hd0JGadH48d7q1+ sSI7+Vo85i3O9sw== X-Developer-Key: i=leitao@debian.org; a=openpgp; fpr=AC8539A6E8F46702CA4A439B35A3939FFC78776D X-Debian-User: leitao locate_mem_hole_top_down() walks candidates downwards by subtracting PAGE_SIZE whenever the window conflicts with an existing segment or with an architecture exclude range. Nothing stops that subtraction at zero, so a search that reaches the bottom of the address space wraps temp_start around and continues. The walk starts inside the range being scanned and only moves down, so bail out once a candidate ends up above end. That covers every step in the loop rather than the subtractions alone, and it matches locate_mem_hole_bottom_up(), which already bounds its candidate on both sides. This is a better check than subtracting with check_sub_overflow(), given that we would have 3 subtractions in this block, and this single fix would take care of them all (instead of three check_sub_overflow()). Fixes: cb1052581e2b ("kexec: implementation of new syscall kexec_file_load") Signed-off-by: Breno Leitao Reviewed-by: Bradley Morgan --- kernel/kexec_file.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/kernel/kexec_file.c b/kernel/kexec_file.c index 59fb9d71e9d86..01a64d98fbcd7 100644 --- a/kernel/kexec_file.c +++ b/kernel/kexec_file.c @@ -484,7 +484,9 @@ static int locate_mem_hole_top_down(unsigned long start= , unsigned long end, /* align down start */ temp_start =3D ALIGN_DOWN(temp_start, kbuf->buf_align); =20 - if (temp_start < start || temp_start < kbuf->buf_min) + /* A candidate above the range means the walk wrapped around */ + if (temp_start < start || temp_start < kbuf->buf_min || + temp_start > end) return 0; =20 temp_end =3D temp_start + kbuf->memsz - 1; --=20 2.53.0-Meta From nobody Tue Sep 29 04:43:03 2026 Received: from stravinsky.debian.org (stravinsky.debian.org [82.195.75.108]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D9460435AA1 for ; Wed, 12 Aug 2026 11:32:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=82.195.75.108 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786534346; cv=none; b=MiBRWDb99Vku72zqWWknKFJNteWxHRiApXm+Kn7ODAk2GZFnB8WVkcaTtl9jIK8WrweNtNkSEx0uG21KnDNzC+lDfeWLkyDt8M1ID+cyFh9zrA1c8Z93DLZ9loJzGVk2cnHz7WGVhqcNqDfBx3fOEB8dFFo12EcllmdlH9czwJ8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786534346; c=relaxed/simple; bh=DngnSfPX8XBe6I/yjyE8XagV4YyZPUZlPnCTWnEo5k4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=qv5+/N6z8NXzBRdfUk8SfpYSmhvrvGRSDO7O6oE9TWObP5ws97+r/vOE2QOXLa1urH7BE1+w23lVRZplSChjQYGWlGR78CwIQEPRy4OzjNxh/oFzBzpSY/XVt45MsjIuLMtpowv7lh0/bUPuXu/zSvwzIerMHeG/Xqbn65BzGMc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org; spf=pass smtp.mailfrom=debian.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b=lMBLiFDX; arc=none smtp.client-ip=82.195.75.108 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=debian.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b="lMBLiFDX" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=debian.org; s=smtpauto.stravinsky; h=X-Debian-User:Cc:To:In-Reply-To:References: Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date: From:Reply-To:Content-ID:Content-Description; bh=eEO9P1mvess5y0n0KUnCYAeP/+IifPhRXEhfxRdjhi4=; b=lMBLiFDXe7xrvPugP2EHgSR/nH L6WHW8wD4ND+1rUNkQxYbTMBSHt1nDX8MJINlBdl1vpnF4+Trm6iXThbx/JbiNs51St77PFsP5b1P qFxuqyste7/BkihRa9OGHUaMM5LtxY9z0DIeeeXiBG0e673pZr4woHJYRrFP0uawOhd1V4Euojm9+ znpnH9BeGXEtJLzX8BS9zUAwbro6R/SImQ6IGvRZPZ0w9VOpfq+/aUxFKNjLmm5T7KCAOGW4TyzbU w1m2yF2+H7oLYG0P9JN0vOA2dROstVc2+phcoYg65fULWiTk5EBkChKE3BevZm+w3p7ujzUfnWS8W v2Vgw5Yg==; Received: from authenticated-user by stravinsky.debian.org with esmtpsa (TLS1.3:ECDHE_X25519__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim 4.96) (envelope-from ) id 1wu7Bt-004JJP-1k; Wed, 12 Aug 2026 11:32:13 +0000 From: Breno Leitao Date: Wed, 12 Aug 2026 04:31:52 -0700 Subject: [PATCH v6 2/2] kexec: keep the next kernel off hardware-poisoned pages Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260812-kexec_posioned-v6-2-e477887086f0@debian.org> References: <20260812-kexec_posioned-v6-0-e477887086f0@debian.org> In-Reply-To: <20260812-kexec_posioned-v6-0-e477887086f0@debian.org> To: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Baoquan He , Pasha Tatashin , Pratyush Yadav , Miaohe Lin , Naoya Horiguchi Cc: Breno Leitao , linux-mm@kvack.org, linux-kernel@vger.kernel.org, kexec@lists.infradead.org, rmikey@meta.com, riel@surriel.com, kernel-team@meta.com, Kiryl Shutsemau , "Kiryl Shutsemau (Meta)" , Bradley Morgan X-Mailer: b4 0.16-dev-f8e9d X-Developer-Signature: v=1; a=openpgp-sha256; l=6504; i=leitao@debian.org; h=from:subject:message-id; bh=DngnSfPX8XBe6I/yjyE8XagV4YyZPUZlPnCTWnEo5k4=; b=owEBbQKS/ZANAwAIATWjk5/8eHdtAcsmYgBqfFmsufCoDT5cFM1bcdNKZCitkrQyQY/M2Zqyk xSH76lNFveJAjMEAAEIAB0WIQSshTmm6PRnAspKQ5s1o5Of/Hh3bQUCanxZrAAKCRA1o5Of/Hh3 bdGFD/92nUsa8eZs63kMuOvyNiWyDpQeZzq//w2XgiP+7o6YoTuBSKV85sb0nS3tF+MRwuRswiK kb+9r1A32Qkkg1UGU2wnVMKiGJ9+wUejnaLVVYCpsT6oiPRC8Ew5vrbPmFemzhRD4HoiERfXqoy i2tF1+0PEsrL4bCpZNgpmYbgZXB5ZMJnGjymNBohy2bwbVCvusbh48Gi599QMoI+QU+/tHo/rdA UFMzRqfLcdMeEpV90Vyln3FTcsfm7ten97cC0jCQrzrBtbXHc1jHELYcSuDEs1nkTaFZe8AQGSs keyopJCpKyX8dO415r7FmadedOtaaoCC/5HsjBV8Zb1CcP+dKD8FDirBfa7G3F1dPz/oiUZfhLv hthfFOhT4/qi1XVIaPFURSqibxmfqUMGNQbcDBT36O0nQr6KlIr5hoF1LNOSeRnluQRjRmCMsbA Wje6t8igbkRm/I9JN1sPhOP5EiDZmNlvzqb6gVE1vuQH/9c+k+ga3+vo2CvrVhqRhONgLd9MzdB pn4DDzlqht+TGzDyX9bMuZmC6dwDEdvjno0gx041/whLGeoNS3VuxK53s1KA7u2PFfx5M/IoPzS XYdiG1bTrs80KIKtacvHg/Rb1SxCSDLzs1VpzLqzhWgDuIPqY2AHmh4Dl5I+EEAVSPZvyq+7q72 r5jTQEUCzjaONKQ== X-Developer-Key: i=leitao@debian.org; a=openpgp; fpr=AC8539A6E8F46702CA4A439B35A3939FFC78776D X-Debian-User: leitao Memory failures (such as unrecoverable ECCs errors) are getting more and more common. The kernel knows how to handle it while running, marking it as poisoned (and SIGBUS user tasks). Poisoned memory is removed from the buddy allocator, but, not from other places. A current problem is that kexec will load new kernel on top of a bad/poisoned memory, which is undesirable. If the next kernel's image, initrd or purgatory lands on a poisoned frame, the relocation copy puts it on memory that is known bad. The error happens on the first read from a bad page, and that is what we want to avoid. Skip hardware-poisoned frames that were detected by the memory failure subsystem earlier when placing kexec segments. To do so, add a helper that reports the first or the last poisoned page in a range: memory is walked top-down by locate_mem_hole_top_down() and bottom-up by locate_mem_hole_bottom_up(), so each direction needs a different answer to stay clear of the poison. kexec_load() gets its destinations from userspace and cannot move them, so there sanity_check_segment_list() just rejects a segment that happens to have a poisoned page. is_page_hwpoison() also covers hugetlb, so a poisoned hugetlb folio is skipped as a whole. Suggested-by: Kiryl Shutsemau Signed-off-by: Breno Leitao Reviewed-by: Kiryl Shutsemau (Meta) Reviewed-by: Pratyush Yadav Reviewed-by: Bradley Morgan Reviewed-by: Miaohe Lin --- include/linux/mm.h | 14 ++++++++++++++ kernel/kexec_core.c | 10 ++++++++++ kernel/kexec_file.c | 16 ++++++++++++++++ mm/memory-failure.c | 40 ++++++++++++++++++++++++++++++++++++++++ 4 files changed, 80 insertions(+) diff --git a/include/linux/mm.h b/include/linux/mm.h index 7fabe6c66b4b7..41b923901b193 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -5192,6 +5192,8 @@ extern const struct attribute_group memory_failure_at= tr_group; extern void memory_failure_queue(unsigned long pfn, int flags); void num_poisoned_pages_inc(unsigned long pfn); void num_poisoned_pages_sub(unsigned long pfn, long i); +phys_addr_t range_first_hwpoison(phys_addr_t start, unsigned long size); +phys_addr_t range_last_hwpoison(phys_addr_t start, unsigned long size); #else static inline void memory_failure_queue(unsigned long pfn, int flags) { @@ -5204,6 +5206,18 @@ static inline void num_poisoned_pages_inc(unsigned l= ong pfn) static inline void num_poisoned_pages_sub(unsigned long pfn, long i) { } + +static inline phys_addr_t range_first_hwpoison(phys_addr_t start, + unsigned long size) +{ + return PHYS_ADDR_MAX; +} + +static inline phys_addr_t range_last_hwpoison(phys_addr_t start, + unsigned long size) +{ + return PHYS_ADDR_MAX; +} #endif =20 #if defined(CONFIG_MEMORY_FAILURE) && defined(CONFIG_MEMORY_HOTPLUG) diff --git a/kernel/kexec_core.c b/kernel/kexec_core.c index dc770b9a6d053..7ee8c9f078f6b 100644 --- a/kernel/kexec_core.c +++ b/kernel/kexec_core.c @@ -212,6 +212,16 @@ int sanity_check_segment_list(struct kimage *image) } #endif =20 + /* + * Reject destinations that land on hardware-poisoned memory: the + * relocation copy would machine-check on the bad frame. + */ + for (i =3D 0; i < nr_segments; i++) { + if (range_first_hwpoison(image->segment[i].mem, + image->segment[i].memsz) !=3D PHYS_ADDR_MAX) + return -EHWPOISON; + } + /* * The destination addresses are searched from system RAM rather than * being allocated from the buddy allocator, so they are not guaranteed diff --git a/kernel/kexec_file.c b/kernel/kexec_file.c index 01a64d98fbcd7..8b0fc8a5d3c36 100644 --- a/kernel/kexec_file.c +++ b/kernel/kexec_file.c @@ -475,6 +475,7 @@ static int locate_mem_hole_top_down(unsigned long start= , unsigned long end, { struct kimage *image =3D kbuf->image; unsigned long temp_start, temp_end; + phys_addr_t poison; =20 temp_end =3D min(end, kbuf->buf_max); temp_start =3D temp_end - kbuf->memsz + 1; @@ -506,6 +507,13 @@ static int locate_mem_hole_top_down(unsigned long star= t, unsigned long end, continue; } =20 + poison =3D range_first_hwpoison(temp_start, kbuf->memsz); + if (poison !=3D PHYS_ADDR_MAX) { + /* we hit a poisoned page */ + temp_start =3D poison - kbuf->memsz; + continue; + } + /* We found a suitable memory range */ break; } while (1); @@ -522,6 +530,7 @@ static int locate_mem_hole_bottom_up(unsigned long star= t, unsigned long end, { struct kimage *image =3D kbuf->image; unsigned long temp_start, temp_end; + phys_addr_t poison; =20 temp_start =3D max(start, kbuf->buf_min); =20 @@ -548,6 +557,13 @@ static int locate_mem_hole_bottom_up(unsigned long sta= rt, unsigned long end, continue; } =20 + poison =3D range_last_hwpoison(temp_start, kbuf->memsz); + if (poison !=3D PHYS_ADDR_MAX) { + /* we hit a poisoned page */ + temp_start =3D poison + PAGE_SIZE; + continue; + } + /* We found a suitable memory range */ break; } while (1); diff --git a/mm/memory-failure.c b/mm/memory-failure.c index a8b03e2920ba8..a2ca8df501cae 100644 --- a/mm/memory-failure.c +++ b/mm/memory-failure.c @@ -96,6 +96,46 @@ void num_poisoned_pages_sub(unsigned long pfn, long i) memblk_nr_poison_sub(pfn, i); } =20 +/* + * Return the first or the last hardware-poisoned online page in [start, + * start + size), or PHYS_ADDR_MAX if the range is clean. + */ +static phys_addr_t range_hwpoison(phys_addr_t start, unsigned long size, + bool first) +{ + phys_addr_t poison =3D PHYS_ADDR_MAX; + unsigned long pfn, end_pfn; + + if (!size || !atomic_long_read(&num_poisoned_pages)) + return poison; + + end_pfn =3D PHYS_PFN(start + size - 1); + for (pfn =3D PHYS_PFN(start); pfn <=3D end_pfn; pfn++) { + struct page *page =3D pfn_to_online_page(pfn); + + if (page && is_page_hwpoison(page)) { + if (first) + return PFN_PHYS(pfn); + + poison =3D PFN_PHYS(pfn); + } + + cond_resched(); + } + + return poison; +} + +phys_addr_t range_first_hwpoison(phys_addr_t start, unsigned long size) +{ + return range_hwpoison(start, size, true); +} + +phys_addr_t range_last_hwpoison(phys_addr_t start, unsigned long size) +{ + return range_hwpoison(start, size, false); +} + /** * MF_ATTR_RO - Create sysfs entry for each memory failure statistics. * @_name: name of the file in the per NUMA sysfs directory. --=20 2.53.0-Meta