include/linux/kexec_handover.h | 10 +++ kernel/liveupdate/kexec_handover.c | 127 ++++++++++++++++++++++++----- 2 files changed, 115 insertions(+), 22 deletions(-)
Introduction ============ This series is required for the ongoing effort to preserve DMA allocations across KHO [1]. It addresses a fundamental mismatch between the current KHO restoration logic and the physical reality of high-order buddy allocations. The Problem =========== The current KHO restore implementation treats all multi-page blocks as split pages during restoration. Specifically, kho_restore_pages() initializes every 4KB sub-page with a refcount of 1. However, many kernel subsystems, most notably the DMA allocator (via dma_alloc_coherent), frequently return high-order non-compound pages. In this state, only the head page carries a refcount of 1, while all tail pages have a refcount of 0. Consequently, when these contiguous blocks are restored by KHO in the new kernel, the forced reference count of 1 on tail pages causes some trouble with the buddy allocator. Downstream of the eventual free path, __free_pages_prepare() [2] ends up calling page_expected_state() [3] when is_check_pages_enabled() returns true (triggered when CONFIG_DEBUG_VM is enabled or debug_pagealloc=on). This detects the unexpected non-zero reference counts on tail pages [4] and incorrectly taints the kernel while leaking the physical pages in question. Proposed Solution ================= Following feedback on the v1 RFC, this series moves away from auto type detection and instead introduces explicit preserve / restore APIs for high-order pages. Callers now explicitly preserve these high-order blocks as a single unit by using kho_preserve_page() and kho_restore_page(). These functions apply a refcount of 1 to the head page while leaving tail pages at 0. The existing APIs (kho_preserve_pages / kho_restore_pages) remain as is for ranges of independent 4KB pages, continuing to use the split refcount. The internal initialization logic is refactored to provide a helper: kho_init_high_order_page(), which is shared between folios and high-order page restore APIs. We also consolidate the common metadata validation, state clearing, and managed page accounting into __kho_restore_page() to avoid duplication. [v4] - Consolidated adjust_managed_page_count() within __kho_restore_page() [v3] - Renamed "unsplit" terminology to "high-order". - Consolidated the common restoration code (magic checks, private clearing etc.) into the internal __kho_restore_page() helper. [v2] - https://lore.kernel.org/all/20260713204935.3069000-1-praan@google.com/ - Dropped automatic type detection via higher bits in Radix key. - Introduced explicit kho_preserve_page and kho_restore_page helpers. - Refactored internal init logic to share code between folios & high-order pages. [v1] https://lore.kernel.org/all/20260703020832.1731864-1-praan@google.com/ Thanks, Praan [1] https://lore.kernel.org/all/20260708234854.4044652-1-skhawaja@google.com/ [2] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1370 [3] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1027 [4] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1034 Pranjal Shrivastava (2): kho: Introduce a helper to init high order pages kho: Introduce preserve/restore APIs for high-order pages include/linux/kexec_handover.h | 10 +++ kernel/liveupdate/kexec_handover.c | 127 ++++++++++++++++++++++++----- 2 files changed, 115 insertions(+), 22 deletions(-) base-commit: 8ba098e6b6ff0db8edf28528d1552be261af30d4 -- 2.55.0.508.g3f0d502094-goog
Hi Pranjal, On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote: > Introduction > ============ > This series is required for the ongoing effort to preserve DMA allocations > across KHO [1]. It addresses a fundamental mismatch between the current KHO > restoration logic and the physical reality of high-order buddy allocations. I skimmed through the patches, they look fine to me before the in-depth review :) But we are really close to the merge window, so we'll anyway need to reiterate after v7.3-rc1. > The Problem > =========== > The current KHO restore implementation treats all multi-page blocks as > split pages during restoration. Specifically, kho_restore_pages() > initializes every 4KB sub-page with a refcount of 1. > > However, many kernel subsystems, most notably the DMA allocator (via > dma_alloc_coherent), frequently return high-order non-compound pages. > In this state, only the head page carries a refcount of 1, while > all tail pages have a refcount of 0. This hints that these patches could be a part of the DMA preservation series, unless you expect other users of the new API. Generally, we don't merge new APIs without the users and if DMA preservation is the only user, it's better to fold these two patches there. -- Sincerely yours, Mike.
On Tue, Aug 11 2026, Mike Rapoport wrote: > Hi Pranjal, > > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote: >> Introduction >> ============ >> This series is required for the ongoing effort to preserve DMA allocations >> across KHO [1]. It addresses a fundamental mismatch between the current KHO >> restoration logic and the physical reality of high-order buddy allocations. > > I skimmed through the patches, they look fine to me before the in-depth > review :) > > But we are really close to the merge window, so we'll anyway need to > reiterate after v7.3-rc1. > >> The Problem >> =========== >> The current KHO restore implementation treats all multi-page blocks as >> split pages during restoration. Specifically, kho_restore_pages() >> initializes every 4KB sub-page with a refcount of 1. >> >> However, many kernel subsystems, most notably the DMA allocator (via >> dma_alloc_coherent), frequently return high-order non-compound pages. >> In this state, only the head page carries a refcount of 1, while >> all tail pages have a refcount of 0. > > This hints that these patches could be a part of the DMA preservation > series, unless you expect other users of the new API. > > Generally, we don't merge new APIs without the users and if DMA > preservation is the only user, it's better to fold these two patches there. Makes sense I think. We can review the patches here, but then they can go in with the DMA series. -- Regards, Pratyush Yadav
On Wed, Aug 12, 2026 at 12:32:25PM +0200, Pratyush Yadav wrote: > On Tue, Aug 11 2026, Mike Rapoport wrote: > > > Hi Pranjal, > > > > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote: > >> Introduction > >> ============ > >> This series is required for the ongoing effort to preserve DMA allocations > >> across KHO [1]. It addresses a fundamental mismatch between the current KHO > >> restoration logic and the physical reality of high-order buddy allocations. > > > > I skimmed through the patches, they look fine to me before the in-depth > > review :) > > > > But we are really close to the merge window, so we'll anyway need to > > reiterate after v7.3-rc1. > > > >> The Problem > >> =========== > >> The current KHO restore implementation treats all multi-page blocks as > >> split pages during restoration. Specifically, kho_restore_pages() > >> initializes every 4KB sub-page with a refcount of 1. > >> > >> However, many kernel subsystems, most notably the DMA allocator (via > >> dma_alloc_coherent), frequently return high-order non-compound pages. > >> In this state, only the head page carries a refcount of 1, while > >> all tail pages have a refcount of 0. > > > > This hints that these patches could be a part of the DMA preservation > > series, unless you expect other users of the new API. > > > > Generally, we don't merge new APIs without the users and if DMA > > preservation is the only user, it's better to fold these two patches there. > > Makes sense I think. We can review the patches here, but then they can > go in with the DMA series. > So.. IIUC, I'll post a v5 here as a standalone series till we get consensus, and finally the reviewed patches can be folded with the DMA series? Thanks, Praan
On Wed, Aug 12 2026, Pranjal Shrivastava wrote: > On Wed, Aug 12, 2026 at 12:32:25PM +0200, Pratyush Yadav wrote: >> On Tue, Aug 11 2026, Mike Rapoport wrote: >> >> > Hi Pranjal, >> > >> > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote: >> >> Introduction >> >> ============ >> >> This series is required for the ongoing effort to preserve DMA allocations >> >> across KHO [1]. It addresses a fundamental mismatch between the current KHO >> >> restoration logic and the physical reality of high-order buddy allocations. >> > >> > I skimmed through the patches, they look fine to me before the in-depth >> > review :) >> > >> > But we are really close to the merge window, so we'll anyway need to >> > reiterate after v7.3-rc1. >> > >> >> The Problem >> >> =========== >> >> The current KHO restore implementation treats all multi-page blocks as >> >> split pages during restoration. Specifically, kho_restore_pages() >> >> initializes every 4KB sub-page with a refcount of 1. >> >> >> >> However, many kernel subsystems, most notably the DMA allocator (via >> >> dma_alloc_coherent), frequently return high-order non-compound pages. >> >> In this state, only the head page carries a refcount of 1, while >> >> all tail pages have a refcount of 0. >> > >> > This hints that these patches could be a part of the DMA preservation >> > series, unless you expect other users of the new API. >> > >> > Generally, we don't merge new APIs without the users and if DMA >> > preservation is the only user, it's better to fold these two patches there. >> >> Makes sense I think. We can review the patches here, but then they can >> go in with the DMA series. >> > > So.. IIUC, I'll post a v5 here as a standalone series till we get > consensus, and finally the reviewed patches can be folded with the DMA > series? That sounds good to me at least. -- Regards, Pratyush Yadav
On Wed, Aug 12, 2026 at 03:46:20PM +0200, Pratyush Yadav wrote: > On Wed, Aug 12 2026, Pranjal Shrivastava wrote: > > > On Wed, Aug 12, 2026 at 12:32:25PM +0200, Pratyush Yadav wrote: > >> On Tue, Aug 11 2026, Mike Rapoport wrote: > >> > >> > Hi Pranjal, > >> > > >> > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote: > >> >> Introduction > >> >> ============ > >> >> This series is required for the ongoing effort to preserve DMA allocations > >> >> across KHO [1]. It addresses a fundamental mismatch between the current KHO > >> >> restoration logic and the physical reality of high-order buddy allocations. > >> > > >> > I skimmed through the patches, they look fine to me before the in-depth > >> > review :) > >> > > >> > But we are really close to the merge window, so we'll anyway need to > >> > reiterate after v7.3-rc1. > >> > > >> >> The Problem > >> >> =========== > >> >> The current KHO restore implementation treats all multi-page blocks as > >> >> split pages during restoration. Specifically, kho_restore_pages() > >> >> initializes every 4KB sub-page with a refcount of 1. > >> >> > >> >> However, many kernel subsystems, most notably the DMA allocator (via > >> >> dma_alloc_coherent), frequently return high-order non-compound pages. > >> >> In this state, only the head page carries a refcount of 1, while > >> >> all tail pages have a refcount of 0. > >> > > >> > This hints that these patches could be a part of the DMA preservation > >> > series, unless you expect other users of the new API. > >> > > >> > Generally, we don't merge new APIs without the users and if DMA > >> > preservation is the only user, it's better to fold these two patches there. > >> > >> Makes sense I think. We can review the patches here, but then they can > >> go in with the DMA series. > >> > > > > So.. IIUC, I'll post a v5 here as a standalone series till we get > > consensus, and finally the reviewed patches can be folded with the DMA > > series? > > That sounds good to me at least. Yes, although when the patches will be a part of the DMA series we might notices something else :) > -- > Regards, > Pratyush Yadav -- Sincerely yours, Mike.
On Thu, Aug 13, 2026 at 02:39:46PM +0300, Mike Rapoport wrote: > On Wed, Aug 12, 2026 at 03:46:20PM +0200, Pratyush Yadav wrote: > > On Wed, Aug 12 2026, Pranjal Shrivastava wrote: > > > > > On Wed, Aug 12, 2026 at 12:32:25PM +0200, Pratyush Yadav wrote: > > >> On Tue, Aug 11 2026, Mike Rapoport wrote: > > >> > > >> > Hi Pranjal, > > >> > > > >> > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote: > > >> >> Introduction > > >> >> ============ > > >> >> This series is required for the ongoing effort to preserve DMA allocations > > >> >> across KHO [1]. It addresses a fundamental mismatch between the current KHO > > >> >> restoration logic and the physical reality of high-order buddy allocations. > > >> > > > >> > I skimmed through the patches, they look fine to me before the in-depth > > >> > review :) > > >> > > > >> > But we are really close to the merge window, so we'll anyway need to > > >> > reiterate after v7.3-rc1. > > >> > > > >> >> The Problem > > >> >> =========== > > >> >> The current KHO restore implementation treats all multi-page blocks as > > >> >> split pages during restoration. Specifically, kho_restore_pages() > > >> >> initializes every 4KB sub-page with a refcount of 1. > > >> >> > > >> >> However, many kernel subsystems, most notably the DMA allocator (via > > >> >> dma_alloc_coherent), frequently return high-order non-compound pages. > > >> >> In this state, only the head page carries a refcount of 1, while > > >> >> all tail pages have a refcount of 0. > > >> > > > >> > This hints that these patches could be a part of the DMA preservation > > >> > series, unless you expect other users of the new API. > > >> > > > >> > Generally, we don't merge new APIs without the users and if DMA > > >> > preservation is the only user, it's better to fold these two patches there. > > >> > > >> Makes sense I think. We can review the patches here, but then they can > > >> go in with the DMA series. > > >> > > > > > > So.. IIUC, I'll post a v5 here as a standalone series till we get > > > consensus, and finally the reviewed patches can be folded with the DMA > > > series? > > > > That sounds good to me at least. > > Yes, although when the patches will be a part of the DMA series we might > notices something else :) > Ohh yes, definitely! Just asking to know what to follow for the next version :) > > -- > > Regards, > > Pratyush Yadav > > -- > Sincerely yours, > Mike. Thanks, Praan
© 2016 - 2026 Red Hat, Inc.