[PATCH v4 0/2] kho: support preserving high-order non-compound pages

Pranjal Shrivastava posted 2 patches 1 month, 4 weeks ago
include/linux/kexec_handover.h     |  10 +++
kernel/liveupdate/kexec_handover.c | 127 ++++++++++++++++++++++++-----
2 files changed, 115 insertions(+), 22 deletions(-)
[PATCH v4 0/2] kho: support preserving high-order non-compound pages
Posted by Pranjal Shrivastava 1 month, 4 weeks ago
Introduction
============
This series is required for the ongoing effort to preserve DMA allocations
across KHO [1]. It addresses a fundamental mismatch between the current KHO
restoration logic and the physical reality of high-order buddy allocations.

The Problem
===========
The current KHO restore implementation treats all multi-page blocks as
split pages during restoration. Specifically, kho_restore_pages()
initializes every 4KB sub-page with a refcount of 1.

However, many kernel subsystems, most notably the DMA allocator (via
dma_alloc_coherent), frequently return high-order non-compound pages.
In this state, only the head page carries a refcount of 1, while
all tail pages have a refcount of 0.

Consequently, when these contiguous blocks are restored by KHO in
the new kernel, the forced reference count of 1 on tail pages causes some 
trouble with the buddy allocator. Downstream of the eventual free path,
__free_pages_prepare() [2] ends up calling page_expected_state() [3]
when is_check_pages_enabled() returns true (triggered when CONFIG_DEBUG_VM
is enabled or debug_pagealloc=on).

This detects the unexpected non-zero reference counts on tail pages [4] and
incorrectly taints the kernel while leaking the physical pages in question.

Proposed Solution
=================
Following feedback on the v1 RFC, this series moves away from auto type 
detection and instead introduces explicit preserve / restore APIs for 
high-order pages.

Callers now explicitly preserve these high-order blocks as a single unit by using
kho_preserve_page() and kho_restore_page(). These functions apply a refcount
of 1 to the head page while leaving tail pages at 0.

The existing APIs (kho_preserve_pages / kho_restore_pages) remain as is 
for ranges of independent 4KB pages, continuing to use the split refcount.

The internal initialization logic is refactored to provide a helper:
kho_init_high_order_page(), which is shared between folios and high-order page 
restore APIs. We also consolidate the common metadata validation, state
clearing, and managed page accounting into __kho_restore_page() to avoid duplication.

[v4]
 - Consolidated adjust_managed_page_count() within __kho_restore_page()

[v3]
 - Renamed "unsplit" terminology to "high-order".
 - Consolidated the common restoration code (magic checks, private 
   clearing etc.) into the internal __kho_restore_page() helper.

[v2]
 - https://lore.kernel.org/all/20260713204935.3069000-1-praan@google.com/
 - Dropped automatic type detection via higher bits in Radix key.
 - Introduced explicit kho_preserve_page and kho_restore_page helpers.
 - Refactored internal init logic to share code between folios & high-order pages.

[v1] https://lore.kernel.org/all/20260703020832.1731864-1-praan@google.com/

Thanks,
Praan

[1] https://lore.kernel.org/all/20260708234854.4044652-1-skhawaja@google.com/
[2] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1370
[3] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1027
[4] https://elixir.bootlin.com/linux/v7.1.1/source/mm/page_alloc.c#L1034

Pranjal Shrivastava (2):
  kho: Introduce a helper to init high order pages
  kho: Introduce preserve/restore APIs for high-order pages

 include/linux/kexec_handover.h     |  10 +++
 kernel/liveupdate/kexec_handover.c | 127 ++++++++++++++++++++++++-----
 2 files changed, 115 insertions(+), 22 deletions(-)

base-commit: 8ba098e6b6ff0db8edf28528d1552be261af30d4
-- 
2.55.0.508.g3f0d502094-goog
Re: [PATCH v4 0/2] kho: support preserving high-order non-compound pages
Posted by Mike Rapoport 1 month, 3 weeks ago
Hi Pranjal,

On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote:
> Introduction
> ============
> This series is required for the ongoing effort to preserve DMA allocations
> across KHO [1]. It addresses a fundamental mismatch between the current KHO
> restoration logic and the physical reality of high-order buddy allocations.

I skimmed through the patches, they look fine to me before the in-depth
review :)

But we are really close to the merge window, so we'll anyway need to
reiterate after v7.3-rc1.
 
> The Problem
> ===========
> The current KHO restore implementation treats all multi-page blocks as
> split pages during restoration. Specifically, kho_restore_pages()
> initializes every 4KB sub-page with a refcount of 1.
> 
> However, many kernel subsystems, most notably the DMA allocator (via
> dma_alloc_coherent), frequently return high-order non-compound pages.
> In this state, only the head page carries a refcount of 1, while
> all tail pages have a refcount of 0.

This hints that these patches could be a part of the DMA preservation
series, unless you expect other users of the new API.

Generally, we don't merge new APIs without the users and if DMA
preservation is the only user, it's better to fold these two patches there.

-- 
Sincerely yours,
Mike.
Re: [PATCH v4 0/2] kho: support preserving high-order non-compound pages
Posted by Pratyush Yadav 1 month, 2 weeks ago
On Tue, Aug 11 2026, Mike Rapoport wrote:

> Hi Pranjal,
>
> On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote:
>> Introduction
>> ============
>> This series is required for the ongoing effort to preserve DMA allocations
>> across KHO [1]. It addresses a fundamental mismatch between the current KHO
>> restoration logic and the physical reality of high-order buddy allocations.
>
> I skimmed through the patches, they look fine to me before the in-depth
> review :)
>
> But we are really close to the merge window, so we'll anyway need to
> reiterate after v7.3-rc1.
>  
>> The Problem
>> ===========
>> The current KHO restore implementation treats all multi-page blocks as
>> split pages during restoration. Specifically, kho_restore_pages()
>> initializes every 4KB sub-page with a refcount of 1.
>> 
>> However, many kernel subsystems, most notably the DMA allocator (via
>> dma_alloc_coherent), frequently return high-order non-compound pages.
>> In this state, only the head page carries a refcount of 1, while
>> all tail pages have a refcount of 0.
>
> This hints that these patches could be a part of the DMA preservation
> series, unless you expect other users of the new API.
>
> Generally, we don't merge new APIs without the users and if DMA
> preservation is the only user, it's better to fold these two patches there.

Makes sense I think. We can review the patches here, but then they can
go in with the DMA series.

-- 
Regards,
Pratyush Yadav
Re: [PATCH v4 0/2] kho: support preserving high-order non-compound pages
Posted by Pranjal Shrivastava 1 month, 2 weeks ago
On Wed, Aug 12, 2026 at 12:32:25PM +0200, Pratyush Yadav wrote:
> On Tue, Aug 11 2026, Mike Rapoport wrote:
> 
> > Hi Pranjal,
> >
> > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote:
> >> Introduction
> >> ============
> >> This series is required for the ongoing effort to preserve DMA allocations
> >> across KHO [1]. It addresses a fundamental mismatch between the current KHO
> >> restoration logic and the physical reality of high-order buddy allocations.
> >
> > I skimmed through the patches, they look fine to me before the in-depth
> > review :)
> >
> > But we are really close to the merge window, so we'll anyway need to
> > reiterate after v7.3-rc1.
> >  
> >> The Problem
> >> ===========
> >> The current KHO restore implementation treats all multi-page blocks as
> >> split pages during restoration. Specifically, kho_restore_pages()
> >> initializes every 4KB sub-page with a refcount of 1.
> >> 
> >> However, many kernel subsystems, most notably the DMA allocator (via
> >> dma_alloc_coherent), frequently return high-order non-compound pages.
> >> In this state, only the head page carries a refcount of 1, while
> >> all tail pages have a refcount of 0.
> >
> > This hints that these patches could be a part of the DMA preservation
> > series, unless you expect other users of the new API.
> >
> > Generally, we don't merge new APIs without the users and if DMA
> > preservation is the only user, it's better to fold these two patches there.
> 
> Makes sense I think. We can review the patches here, but then they can
> go in with the DMA series.
> 

So.. IIUC, I'll post a v5 here as a standalone series till we get
consensus, and finally the reviewed patches can be folded with the DMA
series?

Thanks,
Praan
Re: [PATCH v4 0/2] kho: support preserving high-order non-compound pages
Posted by Pratyush Yadav 1 month, 2 weeks ago
On Wed, Aug 12 2026, Pranjal Shrivastava wrote:

> On Wed, Aug 12, 2026 at 12:32:25PM +0200, Pratyush Yadav wrote:
>> On Tue, Aug 11 2026, Mike Rapoport wrote:
>> 
>> > Hi Pranjal,
>> >
>> > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote:
>> >> Introduction
>> >> ============
>> >> This series is required for the ongoing effort to preserve DMA allocations
>> >> across KHO [1]. It addresses a fundamental mismatch between the current KHO
>> >> restoration logic and the physical reality of high-order buddy allocations.
>> >
>> > I skimmed through the patches, they look fine to me before the in-depth
>> > review :)
>> >
>> > But we are really close to the merge window, so we'll anyway need to
>> > reiterate after v7.3-rc1.
>> >  
>> >> The Problem
>> >> ===========
>> >> The current KHO restore implementation treats all multi-page blocks as
>> >> split pages during restoration. Specifically, kho_restore_pages()
>> >> initializes every 4KB sub-page with a refcount of 1.
>> >> 
>> >> However, many kernel subsystems, most notably the DMA allocator (via
>> >> dma_alloc_coherent), frequently return high-order non-compound pages.
>> >> In this state, only the head page carries a refcount of 1, while
>> >> all tail pages have a refcount of 0.
>> >
>> > This hints that these patches could be a part of the DMA preservation
>> > series, unless you expect other users of the new API.
>> >
>> > Generally, we don't merge new APIs without the users and if DMA
>> > preservation is the only user, it's better to fold these two patches there.
>> 
>> Makes sense I think. We can review the patches here, but then they can
>> go in with the DMA series.
>> 
>
> So.. IIUC, I'll post a v5 here as a standalone series till we get
> consensus, and finally the reviewed patches can be folded with the DMA
> series?

That sounds good to me at least.

-- 
Regards,
Pratyush Yadav
Re: [PATCH v4 0/2] kho: support preserving high-order non-compound pages
Posted by Mike Rapoport 1 month, 2 weeks ago
On Wed, Aug 12, 2026 at 03:46:20PM +0200, Pratyush Yadav wrote:
> On Wed, Aug 12 2026, Pranjal Shrivastava wrote:
> 
> > On Wed, Aug 12, 2026 at 12:32:25PM +0200, Pratyush Yadav wrote:
> >> On Tue, Aug 11 2026, Mike Rapoport wrote:
> >> 
> >> > Hi Pranjal,
> >> >
> >> > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote:
> >> >> Introduction
> >> >> ============
> >> >> This series is required for the ongoing effort to preserve DMA allocations
> >> >> across KHO [1]. It addresses a fundamental mismatch between the current KHO
> >> >> restoration logic and the physical reality of high-order buddy allocations.
> >> >
> >> > I skimmed through the patches, they look fine to me before the in-depth
> >> > review :)
> >> >
> >> > But we are really close to the merge window, so we'll anyway need to
> >> > reiterate after v7.3-rc1.
> >> >  
> >> >> The Problem
> >> >> ===========
> >> >> The current KHO restore implementation treats all multi-page blocks as
> >> >> split pages during restoration. Specifically, kho_restore_pages()
> >> >> initializes every 4KB sub-page with a refcount of 1.
> >> >> 
> >> >> However, many kernel subsystems, most notably the DMA allocator (via
> >> >> dma_alloc_coherent), frequently return high-order non-compound pages.
> >> >> In this state, only the head page carries a refcount of 1, while
> >> >> all tail pages have a refcount of 0.
> >> >
> >> > This hints that these patches could be a part of the DMA preservation
> >> > series, unless you expect other users of the new API.
> >> >
> >> > Generally, we don't merge new APIs without the users and if DMA
> >> > preservation is the only user, it's better to fold these two patches there.
> >> 
> >> Makes sense I think. We can review the patches here, but then they can
> >> go in with the DMA series.
> >> 
> >
> > So.. IIUC, I'll post a v5 here as a standalone series till we get
> > consensus, and finally the reviewed patches can be folded with the DMA
> > series?
> 
> That sounds good to me at least.

Yes, although when the patches will be a part of the DMA series we might
notices something else :)
 
> -- 
> Regards,
> Pratyush Yadav

-- 
Sincerely yours,
Mike.
Re: [PATCH v4 0/2] kho: support preserving high-order non-compound pages
Posted by Pranjal Shrivastava 1 month, 2 weeks ago
On Thu, Aug 13, 2026 at 02:39:46PM +0300, Mike Rapoport wrote:
> On Wed, Aug 12, 2026 at 03:46:20PM +0200, Pratyush Yadav wrote:
> > On Wed, Aug 12 2026, Pranjal Shrivastava wrote:
> > 
> > > On Wed, Aug 12, 2026 at 12:32:25PM +0200, Pratyush Yadav wrote:
> > >> On Tue, Aug 11 2026, Mike Rapoport wrote:
> > >> 
> > >> > Hi Pranjal,
> > >> >
> > >> > On Mon, Aug 03, 2026 at 11:39:41AM +0000, Pranjal Shrivastava wrote:
> > >> >> Introduction
> > >> >> ============
> > >> >> This series is required for the ongoing effort to preserve DMA allocations
> > >> >> across KHO [1]. It addresses a fundamental mismatch between the current KHO
> > >> >> restoration logic and the physical reality of high-order buddy allocations.
> > >> >
> > >> > I skimmed through the patches, they look fine to me before the in-depth
> > >> > review :)
> > >> >
> > >> > But we are really close to the merge window, so we'll anyway need to
> > >> > reiterate after v7.3-rc1.
> > >> >  
> > >> >> The Problem
> > >> >> ===========
> > >> >> The current KHO restore implementation treats all multi-page blocks as
> > >> >> split pages during restoration. Specifically, kho_restore_pages()
> > >> >> initializes every 4KB sub-page with a refcount of 1.
> > >> >> 
> > >> >> However, many kernel subsystems, most notably the DMA allocator (via
> > >> >> dma_alloc_coherent), frequently return high-order non-compound pages.
> > >> >> In this state, only the head page carries a refcount of 1, while
> > >> >> all tail pages have a refcount of 0.
> > >> >
> > >> > This hints that these patches could be a part of the DMA preservation
> > >> > series, unless you expect other users of the new API.
> > >> >
> > >> > Generally, we don't merge new APIs without the users and if DMA
> > >> > preservation is the only user, it's better to fold these two patches there.
> > >> 
> > >> Makes sense I think. We can review the patches here, but then they can
> > >> go in with the DMA series.
> > >> 
> > >
> > > So.. IIUC, I'll post a v5 here as a standalone series till we get
> > > consensus, and finally the reviewed patches can be folded with the DMA
> > > series?
> > 
> > That sounds good to me at least.
> 
> Yes, although when the patches will be a part of the DMA series we might
> notices something else :)
>  

Ohh yes, definitely! Just asking to know what to follow for the next version :)

> > -- 
> > Regards,
> > Pratyush Yadav
> 
> -- 
> Sincerely yours,
> Mike.

Thanks,
Praan