[PATCH v3 0/2] kho: support preserving high-order non-compound pages

Pranjal Shrivastava posted 2 patches 1 day ago
include/linux/kexec_handover.h     |  10 +++
kernel/liveupdate/kexec_handover.c | 130 ++++++++++++++++++++++++-----
2 files changed, 117 insertions(+), 23 deletions(-)
[PATCH v3 0/2] kho: support preserving high-order non-compound pages
Posted by Pranjal Shrivastava 1 day ago
Introduction
============
This series is required for the ongoing effort to preserve DMA allocations
across KHO [1]. It addresses a fundamental mismatch between the current KHO
restoration logic and the physical reality of high-order buddy allocations.

The Problem
===========
The current KHO restore implementation treats all multi-page blocks as
split pages during restoration. Specifically, kho_restore_pages()
initializes every 4KB sub-page with a refcount of 1.

However, many kernel subsystems, most notably the DMA allocator (via
dma_alloc_coherent), frequently return high-order non-compound pages.
In this state, only the head page carries a refcount of 1, while
all tail pages have a refcount of 0.

Consequently, when these contiguous blocks are restored by KHO in
the new kernel, the forced reference count of 1 on tail pages causes some 
trouble with the buddy allocator. Downstream of the eventual free path,
__free_pages_prepare() [2] ends up calling page_expected_state() [3]
when is_check_pages_enabled() returns true (triggered when CONFIG_DEBUG_VM
is enabled or debug_pagealloc=on).

This detects the unexpected non-zero reference counts on tail pages [4] and
incorrectly taints the kernel while leaking the physical pages in question.

Proposed Solution
=================
Following feedback on the v1 RFC, this series moves away from auto type 
detection and instead introduces explicit preserve / restore APIs for 
high-order pages.

Callers now explicitly preserve these high-order blocks as a single unit by using
kho_preserve_page() and kho_restore_page(). These functions apply a refcount
of 1 to the head page while leaving tail pages at 0.

The existing APIs (kho_preserve_pages / kho_restore_pages) remain as is 
for ranges of independent 4KB pages, continuing to use the split refcount.

The internal initialization logic is refactored to provide a helper:
kho_init_high_order_page(), which is shared between folios and high-order page 
restore APIs. We also consolidate the common metadata validation and state
clearing into __kho_restore_page() to avoid duplication.

[v3]
 - Renamed "unsplit" terminology to "high-order".
 - Consolidated the common restoration code (magic checks, private 
   clearing etc.) into the internal __kho_restore_page() helper.

[v2]
 - https://lore.kernel.org/all/20260713204935.3069000-1-praan@google.com/
 - Dropped automatic type detection via higher bits in Radix key.
 - Introduced explicit kho_preserve_page and kho_restore_page helpers.
 - Refactored internal init logic to share code between folios & high-order pages.

[v1] https://lore.kernel.org/all/20260703020832.1731864-1-praan@google.com/

Thanks,
Praan

[1] https://lore.kernel.org/all/20260708234854.4044652-1-skhawaja@google.com/
[2] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1370
[3] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1027
[4] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1034


Pranjal Shrivastava (2):
  kho: Introduce a helper to init high order pages
  kho: Introduce preserve/restore APIs for high-order pages

 include/linux/kexec_handover.h     |  10 +++
 kernel/liveupdate/kexec_handover.c | 130 ++++++++++++++++++++++++-----
 2 files changed, 117 insertions(+), 23 deletions(-)

-- 
2.55.0.229.g6434b31f56-goog