include/trace/events/huge_memory.h | 18 +++++++++--------- mm/khugepaged.c | 23 ++++++++++++++++------- 2 files changed, 25 insertions(+), 16 deletions(-)
From: Vernon Yang <yanglincheng@kylinos.cn> The khugepaged tracepoints take a folio pointer and call folio_pfn(), but by then the folio may no longer be valid: freed after folio_put(), folio_unlock() or pte_unmap_unlock(), or not a folio at all but an xarray-encoded swap entry. On classic SPARSEMEM, dereferencing it oopses khugepaged as soon as the trace event is enabled; on other memory models it merely prints a bogus pfn. Pass the pfn to the tracepoints directly, captured while the folio is still pinned, closing the use-after-free windows in mm_khugepaged_scan_file(), mm_khugepaged_scan_pmd() and mm_khugepaged_collapse_file(). V1 -> V2: - Instead of passing the folio, just pass the pfn directly. - Using the folio_pfn() before dropping the reference or the page table lock. V1 : https://lore.kernel.org/linux-mm/20260811133655.267739-1-vernon2gm@gmail.com/ Vernon Yang (3): mm: khugepaged: fix swap entry value to folio_pfn() mm: khugepaged: fix folio is used after pte_unmap_unlock() mm: khugepaged: fix folio is used after folio_put/unlock() include/trace/events/huge_memory.h | 18 +++++++++--------- mm/khugepaged.c | 23 ++++++++++++++++------- 2 files changed, 25 insertions(+), 16 deletions(-) base-commit: 075b74841bd0065a3bda3440873c747938e69b68 -- 2.53.0
On Sat, Aug 15, 2026 at 01:19:21PM +0800, Vernon Yang wrote: >From: Vernon Yang <yanglincheng@kylinos.cn> > >The khugepaged tracepoints take a folio pointer and call folio_pfn(), >but by then the folio may no longer be valid: freed after folio_put(), >folio_unlock() or pte_unmap_unlock(), or not a folio at all but an >xarray-encoded swap entry. On classic SPARSEMEM, dereferencing it oopses >khugepaged as soon as the trace event is enabled; on other memory models >it merely prints a bogus pfn. > >Pass the pfn to the tracepoints directly, captured while the folio is >still pinned, closing the use-after-free windows in >mm_khugepaged_scan_file(), mm_khugepaged_scan_pmd() and >mm_khugepaged_collapse_file(). Well spotted! Gave the series a run on x86_64 (KVM), all good (only classic SPARSEMEM untested) :) Tested-by: Lance Yang <lance.yang@linux.dev>
+Cc Baolin
On Sun, Aug 16, 2026 at 01:44:44AM +0800, Lance Yang wrote:
>
>On Sat, Aug 15, 2026 at 01:19:21PM +0800, Vernon Yang wrote:
>>From: Vernon Yang <yanglincheng@kylinos.cn>
>>
>>The khugepaged tracepoints take a folio pointer and call folio_pfn(),
>>but by then the folio may no longer be valid: freed after folio_put(),
>>folio_unlock() or pte_unmap_unlock(), or not a folio at all but an
>>xarray-encoded swap entry. On classic SPARSEMEM, dereferencing it oopses
>>khugepaged as soon as the trace event is enabled; on other memory models
>>it merely prints a bogus pfn.
>>
>>Pass the pfn to the tracepoints directly, captured while the folio is
>>still pinned, closing the use-after-free windows in
>>mm_khugepaged_scan_file(), mm_khugepaged_scan_pmd() and
>>mm_khugepaged_collapse_file().
>
>Well spotted!
>
>Gave the series a run on x86_64 (KVM), all good (only classic SPARSEMEM
>untested) :)
Hmm ... stumbled over something else while testing this ...
With tmpfs mounted huge=advise, one MADV_HUGEPAGE isn't enough to get
an unregistered mm onto khugepaged's list. Do it twice, and khugepaged
starts scanning right away.
The pending flags make it into khugepaged just fine:
int hugepage_madvise(struct vm_area_struct *vma,
vm_flags_t *vm_flags, int advice)
{
switch (advice) {
case MADV_HUGEPAGE:
*vm_flags &= ~VM_NOHUGEPAGE;
*vm_flags |= VM_HUGEPAGE;
...
khugepaged_enter_vma(vma, *vm_flags);
break;
...
}
return 0;
}
and survive the common eligibility check:
unsigned long __thp_vma_allowable_orders(struct vm_area_struct *vma,
vm_flags_t vm_flags,
enum tva_type type,
unsigned long orders)
{
...
/*
* Enabled via shmem mount options or sysfs settings.
* Must be done before hugepage flags check since shmem has its
* own flags.
*/
if (!in_pf && shmem_file(vma->vm_file))
return orders & shmem_allowable_huge_orders(file_inode(vma->vm_file),
vma, vma_start_pgoff(vma), 0,
forced_collapse);
...
}
But then the shmem helper reads them back from the VMA:
unsigned long shmem_allowable_huge_orders(struct inode *inode,
struct vm_area_struct *vma, pgoff_t index,
loff_t write_end, bool shmem_huge_force)
{
...
vm_flags_t vm_flags = vma ? vma->vm_flags : 0;
...
}
At that point vma->vm_flags still has the old value, so huge=advise
quietly gives us no allowable order.
First madvise still succeeds, of course. The second one works because it
finds VM_HUGEPAGE already installed by the first call.
Looked at history too ... we've been here before. 2cf1338454a8 ("mm: fix
khugepaged with shmem_enabled=advise") fixed this exact ordering bug and
tagged cd89fb065099 as the culprit. Then 6beeab870e70 was meant to be
"No functional changes", but moving shmem_huge_global_enabled() into
shmem_allowable_huge_orders() seems to have wired the stale vma->vm_flags
read back in.
So AFAICT, this regressed in v6.12 with 6beeab870e70.
@Baolin, does that ring a bell? Any reason shmem_allowable_huge_orders()
can't just take the pending vm_flags as well?
Cheers, Lance
On Sun, 16 Aug 2026 02:16:32 +0800 Lance Yang <lance.yang@linux.dev> wrote: > >>Pass the pfn to the tracepoints directly, captured while the folio is > >>still pinned, closing the use-after-free windows in > >>mm_khugepaged_scan_file(), mm_khugepaged_scan_pmd() and > >>mm_khugepaged_collapse_file(). > > > >Well spotted! > > > >Gave the series a run on x86_64 (KVM), all good (only classic SPARSEMEM > >untested) :) > > Hmm ... stumbled over something else while testing this ... Sashiko might have found a couple of pre-existing things. https://sashiko.dev/#/patchset/20260815051924.194810-1-vernon2gm@gmail.com Vernon, you may decide that the second one is on-topic for this patchset?
On Mon, Aug 17, 2026 at 03:30:49PM -0700, Andrew Morton wrote: > On Sun, 16 Aug 2026 02:16:32 +0800 Lance Yang <lance.yang@linux.dev> wrote: > > > >>Pass the pfn to the tracepoints directly, captured while the folio is > > >>still pinned, closing the use-after-free windows in > > >>mm_khugepaged_scan_file(), mm_khugepaged_scan_pmd() and > > >>mm_khugepaged_collapse_file(). > > > > > >Well spotted! > > > > > >Gave the series a run on x86_64 (KVM), all good (only classic SPARSEMEM > > >untested) :) > > > > Hmm ... stumbled over something else while testing this ... > > Sashiko might have found a couple of pre-existing things. > https://sashiko.dev/#/patchset/20260815051924.194810-1-vernon2gm@gmail.com > > Vernon, you may decide that the second one is on-topic for this patchset? Hi, Andrew, the second one already is on-topic for this patchset, detail look for the PATCH#3 please. -- Cheers, Vernon
Hi Lance,
On 8/16/26 2:16 AM, Lance Yang wrote:
> +Cc Baolin
>
> On Sun, Aug 16, 2026 at 01:44:44AM +0800, Lance Yang wrote:
>>
>> On Sat, Aug 15, 2026 at 01:19:21PM +0800, Vernon Yang wrote:
>>> From: Vernon Yang <yanglincheng@kylinos.cn>
>>>
>>> The khugepaged tracepoints take a folio pointer and call folio_pfn(),
>>> but by then the folio may no longer be valid: freed after folio_put(),
>>> folio_unlock() or pte_unmap_unlock(), or not a folio at all but an
>>> xarray-encoded swap entry. On classic SPARSEMEM, dereferencing it oopses
>>> khugepaged as soon as the trace event is enabled; on other memory models
>>> it merely prints a bogus pfn.
>>>
>>> Pass the pfn to the tracepoints directly, captured while the folio is
>>> still pinned, closing the use-after-free windows in
>>> mm_khugepaged_scan_file(), mm_khugepaged_scan_pmd() and
>>> mm_khugepaged_collapse_file().
>>
>> Well spotted!
>>
>> Gave the series a run on x86_64 (KVM), all good (only classic SPARSEMEM
>> untested) :)
>
> Hmm ... stumbled over something else while testing this ...
>
> With tmpfs mounted huge=advise, one MADV_HUGEPAGE isn't enough to get
> an unregistered mm onto khugepaged's list. Do it twice, and khugepaged
> starts scanning right away.
>
> The pending flags make it into khugepaged just fine:
>
> int hugepage_madvise(struct vm_area_struct *vma,
> vm_flags_t *vm_flags, int advice)
> {
> switch (advice) {
> case MADV_HUGEPAGE:
> *vm_flags &= ~VM_NOHUGEPAGE;
> *vm_flags |= VM_HUGEPAGE;
> ...
> khugepaged_enter_vma(vma, *vm_flags);
> break;
> ...
> }
>
> return 0;
> }
>
> and survive the common eligibility check:
>
> unsigned long __thp_vma_allowable_orders(struct vm_area_struct *vma,
> vm_flags_t vm_flags,
> enum tva_type type,
> unsigned long orders)
> {
> ...
> /*
> * Enabled via shmem mount options or sysfs settings.
> * Must be done before hugepage flags check since shmem has its
> * own flags.
> */
> if (!in_pf && shmem_file(vma->vm_file))
> return orders & shmem_allowable_huge_orders(file_inode(vma->vm_file),
> vma, vma_start_pgoff(vma), 0,
> forced_collapse);
> ...
> }
>
> But then the shmem helper reads them back from the VMA:
>
> unsigned long shmem_allowable_huge_orders(struct inode *inode,
> struct vm_area_struct *vma, pgoff_t index,
> loff_t write_end, bool shmem_huge_force)
> {
> ...
> vm_flags_t vm_flags = vma ? vma->vm_flags : 0;
> ...
> }
>
> At that point vma->vm_flags still has the old value, so huge=advise
> quietly gives us no allowable order.
>
> First madvise still succeeds, of course. The second one works because it
> finds VM_HUGEPAGE already installed by the first call.
>
> Looked at history too ... we've been here before. 2cf1338454a8 ("mm: fix
> khugepaged with shmem_enabled=advise") fixed this exact ordering bug and
> tagged cd89fb065099 as the culprit. Then 6beeab870e70 was meant to be
> "No functional changes", but moving shmem_huge_global_enabled() into
> shmem_allowable_huge_orders() seems to have wired the stale vma->vm_flags
> read back in.
>
> So AFAICT, this regressed in v6.12 with 6beeab870e70.
>
> @Baolin, does that ring a bell? Any reason shmem_allowable_huge_orders()
> can't just take the pending vm_flags as well?
Good catch. Sorry for my mistake. Would you like to send a fix?
Otherwise, I can fix it. Thanks for your report and analysis.
On 2026/8/17 10:25, Baolin Wang wrote: > Hi Lance, > > On 8/16/26 2:16 AM, Lance Yang wrote: >> +Cc Baolin >> >> On Sun, Aug 16, 2026 at 01:44:44AM +0800, Lance Yang wrote: >>> >>> On Sat, Aug 15, 2026 at 01:19:21PM +0800, Vernon Yang wrote: [...] > Otherwise, I can fix it. Thanks for your report and analysis. Yeah, please go ahead :) Thanks, Baolin!
On 8/17/26 10:53 AM, Lance Yang wrote: > > > On 2026/8/17 10:25, Baolin Wang wrote: >> Hi Lance, >> >> On 8/16/26 2:16 AM, Lance Yang wrote: >>> +Cc Baolin >>> >>> On Sun, Aug 16, 2026 at 01:44:44AM +0800, Lance Yang wrote: >>>> >>>> On Sat, Aug 15, 2026 at 01:19:21PM +0800, Vernon Yang wrote: > [...] >> Otherwise, I can fix it. Thanks for your report and analysis. > > Yeah, please go ahead :) Thanks, Baolin! FYI: I've posted the fix and verified it resolves the issue from my testing. https://lore.kernel.org/all/ed34ca03ae7d65e89467fb87bc961f5497049c00.1786948410.git.baolin.wang@linux.alibaba.com/
© 2016 - 2026 Red Hat, Inc.