Introduction
============
This series is required for the ongoing effort to preserve DMA allocations
across KHO [1]. It addresses a fundamental mismatch between the current KHO
restoration logic and the physical reality of high-order buddy allocations.
The Problem
===========
The current KHO restore implementation treats all multi-page blocks as
split pages during restoration. Specifically, kho_restore_pages()
initializes every 4KB sub-page with a refcount of 1.
However, many kernel subsystems, most notably the DMA allocator (via
dma_alloc_coherent), frequently return high-order non-compound pages.
In this state, only the head page carries a refcount of 1, while
all tail pages have a refcount of 0.
Consequently, when these contiguous blocks are restored by KHO in
the new kernel, the forced reference count of 1 on tail pages causes some
trouble with the buddy allocator. Downstream of the eventual free path,
__free_pages_prepare() [2] ends up calling page_expected_state() [3]
when is_check_pages_enabled() returns true (triggered when CONFIG_DEBUG_VM
is enabled or debug_pagealloc=on).
This detects the unexpected non-zero reference counts on tail pages [4] and
incorrectly taints the kernel while leaking the physical pages in question.
Proposed Solution
=================
Following feedback on the v1 RFC, this series moves away from auto type
detection and instead introduces explicit preserve / restore APIs for
high-order pages.
Callers now explicitly preserve these high-order blocks as a single unit by using
kho_preserve_page() and kho_restore_page(). These functions apply a refcount
of 1 to the head page while leaving tail pages at 0.
The existing APIs (kho_preserve_pages / kho_restore_pages) remain as is
for ranges of independent 4KB pages, continuing to use the split refcount.
The internal initialization logic is refactored to provide a helper:
kho_init_high_order_page(), which is shared between folios and high-order page
restore APIs. We also consolidate the common metadata validation and state
clearing into __kho_restore_page() to avoid duplication.
[v3]
- Renamed "unsplit" terminology to "high-order".
- Consolidated the common restoration code (magic checks, private
clearing etc.) into the internal __kho_restore_page() helper.
[v2]
- https://lore.kernel.org/all/20260713204935.3069000-1-praan@google.com/
- Dropped automatic type detection via higher bits in Radix key.
- Introduced explicit kho_preserve_page and kho_restore_page helpers.
- Refactored internal init logic to share code between folios & high-order pages.
[v1] https://lore.kernel.org/all/20260703020832.1731864-1-praan@google.com/
Thanks,
Praan
[1] https://lore.kernel.org/all/20260708234854.4044652-1-skhawaja@google.com/
[2] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1370
[3] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1027
[4] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1034
Pranjal Shrivastava (2):
kho: Introduce a helper to init high order pages
kho: Introduce preserve/restore APIs for high-order pages
include/linux/kexec_handover.h | 10 +++
kernel/liveupdate/kexec_handover.c | 130 ++++++++++++++++++++++++-----
2 files changed, 117 insertions(+), 23 deletions(-)
--
2.55.0.229.g6434b31f56-goog