tools/testing/selftests/mm/pagemap_ioctl.c | 18 +- tools/testing/selftests/mm/soft-dirty.c | 6 +- tools/testing/selftests/mm/split_huge_page_test.c | 27 ++- tools/testing/selftests/mm/vm_util.c | 212 ++++++++++++++++------ tools/testing/selftests/mm/vm_util.h | 3 + 5 files changed, 192 insertions(+), 74 deletions(-)
split_huge_page_test can fail for the following reasons:
1. During the test, khugepaged may collapse previously split pages again,
causing intermittent failures.
2. Since glibc commit 321e1fc73f (“malloc: Enable 2MB THP by default on AArch64”),
glibc may call madvise(MADV_HUGEPAGE) for sufficiently large allocations
made by memalign(). The underlying VMA may start at a different address
from the aligned address returned by memalign(). Moreover, a subsequent
madvise(MADV_HUGEPAGE) call does not split the VMA because it already
has the same advice.
This causes the test to fail because the check_huge_xxx() helpers
incorrectly require the address returned by memalign() to match the
VMA start address reported in /proc/self/smaps.
Address these issues by applying MADV_NOHUGEPAGE after faulting in the
huge page, preventing khugepaged from collapsing it again, and instead of
relying on /proc/self/smaps, use /proc/self/pagemap and
/proc/kpageflags to detect huge-page mappings and large folios:
1. If hpage_size == pmd_pagesize, check PAGE_IS_HUGE instead of
using check_large_folios(), since only the mapping type matters.
This identifies PMD-mapped huge pages.
2. Otherwise, use check_large_folios() to detect large folios. This
covers mTHP cases.
3. Check the folio flags according to the type of huge page.
Also, current usage of memalign() would result memory area may
unexpectedly merge with an adjacent VMA, causing tests
that inspect it through /proc/self/smaps to fail.
Eliminate this potential source of test flakiness by introducing
alloc_isolated_mem(), which places guard regions before and after the
allocated area to prevent unexpected VMA merging.
This patch based on mm-unstable
---
Changes in v5:
- rename madv_nohuge() to disable_khugepaged().
- Link to v4: https://lore.kernel.org/r/20260902-fix_split-v4-0-85f03905f7b1@arm.com
Changes in v4:
- reparse commit message.
- introduce alloc_isolated_mem().
- make madv_nohuge() helper according to suggestion.
- Link to v3: https://lore.kernel.org/r/20260828-fix_split-v3-0-374022586a4b@arm.com
Changes in v3:
- change message in case of failure of madvise() with MADV_NOHUGEPAGE.
- fix collapse_single_mthp (mthp_khugepaged:anon) test case failure.
- Link to v2: https://lore.kernel.org/all/20260826-fix_split-v2-0-71153c7f579a@arm.com/
Changes in v2:
- rebase to mm-unstable.
- add message in case of failure of madvise() with MADV_NOHUGEPAGE.
- fix wrong setup expected_huge in check_huge_shmem().
- Link to v1: https://lore.kernel.org/r/20260820-fix_split-v1-0-ab430c58c7cf@arm.com
---
Yeoreum Yun (3):
kselftest: mm: prevent random failure of huge page split for khugepaged
kselftest: mm: replace usage of /proc/self/smaps for check_huge_xxx() helper
kselftest: mm: introduce alloc_isolated_mem()
tools/testing/selftests/mm/pagemap_ioctl.c | 18 +-
tools/testing/selftests/mm/soft-dirty.c | 6 +-
tools/testing/selftests/mm/split_huge_page_test.c | 27 ++-
tools/testing/selftests/mm/vm_util.c | 212 ++++++++++++++++------
tools/testing/selftests/mm/vm_util.h | 3 +
5 files changed, 192 insertions(+), 74 deletions(-)
---
base-commit: d118502628f8b673be9023db8bdf878f64a7ed45
change-id: 20260820-fix_split-f44939ec44b8
Best regards,
--
Sincerely,
Yeoreum Yun
On 9/7/26 10:19, Yeoreum Yun wrote: > split_huge_page_test can fail for the following reasons: > > 1. During the test, khugepaged may collapse previously split pages again, > causing intermittent failures. > > 2. Since glibc commit 321e1fc73f (“malloc: Enable 2MB THP by default on AArch64”), > glibc may call madvise(MADV_HUGEPAGE) for sufficiently large allocations > made by memalign(). The underlying VMA may start at a different address > from the aligned address returned by memalign(). Moreover, a subsequent > madvise(MADV_HUGEPAGE) call does not split the VMA because it already > has the same advice. > > This causes the test to fail because the check_huge_xxx() helpers > incorrectly require the address returned by memalign() to match the > VMA start address reported in /proc/self/smaps. > > Address these issues by applying MADV_NOHUGEPAGE after faulting in the > huge page, preventing khugepaged from collapsing it again, and instead of > relying on /proc/self/smaps, use /proc/self/pagemap and > /proc/kpageflags to detect huge-page mappings and large folios: > > 1. If hpage_size == pmd_pagesize, check PAGE_IS_HUGE instead of > using check_large_folios(), since only the mapping type matters. > This identifies PMD-mapped huge pages. > 2. Otherwise, use check_large_folios() to detect large folios. This > covers mTHP cases. > 3. Check the folio flags according to the type of huge page. > > Also, current usage of memalign() would result memory area may > unexpectedly merge with an adjacent VMA, causing tests > that inspect it through /proc/self/smaps to fail. I'm not particularly happy about this. Relying on VMA merging details rather hints that we shouldn't be using smaps to query some stats/properties. Which exact things are test querying through /proc/self/smaps? Could we convert the code to just query that stuff through different interfaces? -- Cheers, David
> On 9/7/26 10:19, Yeoreum Yun wrote: > > split_huge_page_test can fail for the following reasons: > > > > 1. During the test, khugepaged may collapse previously split pages again, > > causing intermittent failures. > > > > 2. Since glibc commit 321e1fc73f (“malloc: Enable 2MB THP by default on AArch64”), > > glibc may call madvise(MADV_HUGEPAGE) for sufficiently large allocations > > made by memalign(). The underlying VMA may start at a different address > > from the aligned address returned by memalign(). Moreover, a subsequent > > madvise(MADV_HUGEPAGE) call does not split the VMA because it already > > has the same advice. > > > > This causes the test to fail because the check_huge_xxx() helpers > > incorrectly require the address returned by memalign() to match the > > VMA start address reported in /proc/self/smaps. > > > > Address these issues by applying MADV_NOHUGEPAGE after faulting in the > > huge page, preventing khugepaged from collapsing it again, and instead of > > relying on /proc/self/smaps, use /proc/self/pagemap and > > /proc/kpageflags to detect huge-page mappings and large folios: > > > > 1. If hpage_size == pmd_pagesize, check PAGE_IS_HUGE instead of > > using check_large_folios(), since only the mapping type matters. > > This identifies PMD-mapped huge pages. > > 2. Otherwise, use check_large_folios() to detect large folios. This > > covers mTHP cases. > > 3. Check the folio flags according to the type of huge page. > > > > Also, current usage of memalign() would result memory area may > > unexpectedly merge with an adjacent VMA, causing tests > > that inspect it through /proc/self/smaps to fail. > I'm not particularly happy about this. > > Relying on VMA merging details rather hints that we shouldn't be using smaps to > query some stats/properties. > > Which exact things are test querying through /proc/self/smaps? Could we convert > the code to just query that stuff through different interfaces? Well, users currently for using /proc/self/smaps are for check vm-flags: - guard-regions where using check_vmflags_guard() - pfnmap test where uses check_vmflag_pfnmap() AFAIK, there is no interface to get vm-flags except /proc/self/smaps, we should add new interface but I'm not sure whether it's useful except for test purpose. Since for a testing, the most chance for VMA merge is when using anon private mapping and this would be enough with allocate_isolated_mem(). -- Sincerely, Yeoreum Yun
On 9/10/26 12:31, Yeoreum Yun wrote: >> On 9/7/26 10:19, Yeoreum Yun wrote: >>> split_huge_page_test can fail for the following reasons: >>> >>> 1. During the test, khugepaged may collapse previously split pages again, >>> causing intermittent failures. >>> >>> 2. Since glibc commit 321e1fc73f (“malloc: Enable 2MB THP by default on AArch64”), >>> glibc may call madvise(MADV_HUGEPAGE) for sufficiently large allocations >>> made by memalign(). The underlying VMA may start at a different address >>> from the aligned address returned by memalign(). Moreover, a subsequent >>> madvise(MADV_HUGEPAGE) call does not split the VMA because it already >>> has the same advice. >>> >>> This causes the test to fail because the check_huge_xxx() helpers >>> incorrectly require the address returned by memalign() to match the >>> VMA start address reported in /proc/self/smaps. >>> >>> Address these issues by applying MADV_NOHUGEPAGE after faulting in the >>> huge page, preventing khugepaged from collapsing it again, and instead of >>> relying on /proc/self/smaps, use /proc/self/pagemap and >>> /proc/kpageflags to detect huge-page mappings and large folios: >>> >>> 1. If hpage_size == pmd_pagesize, check PAGE_IS_HUGE instead of >>> using check_large_folios(), since only the mapping type matters. >>> This identifies PMD-mapped huge pages. >>> 2. Otherwise, use check_large_folios() to detect large folios. This >>> covers mTHP cases. >>> 3. Check the folio flags according to the type of huge page. >>> >>> Also, current usage of memalign() would result memory area may >>> unexpectedly merge with an adjacent VMA, causing tests >>> that inspect it through /proc/self/smaps to fail. >> I'm not particularly happy about this. >> >> Relying on VMA merging details rather hints that we shouldn't be using smaps to >> query some stats/properties. >> >> Which exact things are test querying through /proc/self/smaps? Could we convert >> the code to just query that stuff through different interfaces? > > Well, users currently for using /proc/self/smaps are for check vm-flags: > - guard-regions where using check_vmflags_guard() > - pfnmap test where uses check_vmflag_pfnmap() Most vm-flags should not be an issue when it comes to merging. The only exception are vmflags that do not prevent VMA merging. So it's VM_SOFTDIRTY and VM_MAYBE_GUARD. And I agree that for guard-regions.c we likely have to care such that we don't merge by accident with other VMAs (guard regions). But that's independent of memalign. ptr = mmap_(self, variant, NULL, 10 * page_size, PROT_READ | PROT_WRITE, 0, 0); ASSERT_FALSE(check_vmflag_guard(ptr)); could be problematic on its own (unlikely but possible). For other flags, you really only have to find the smaps area that covers the given address and look at the vm-flags. Or am I missing something important? (merging vnas with VM_PFNMAP is impossible right now IIRC) -- Cheers, David
On Thu, Sep 10, 2026 at 12:45:39PM +0200, David Hildenbrand (Arm) wrote: > On 9/10/26 12:31, Yeoreum Yun wrote: > >> On 9/7/26 10:19, Yeoreum Yun wrote: > >>> split_huge_page_test can fail for the following reasons: > >>> > >>> 1. During the test, khugepaged may collapse previously split pages again, > >>> causing intermittent failures. > >>> > >>> 2. Since glibc commit 321e1fc73f (“malloc: Enable 2MB THP by default on AArch64”), > >>> glibc may call madvise(MADV_HUGEPAGE) for sufficiently large allocations > >>> made by memalign(). The underlying VMA may start at a different address > >>> from the aligned address returned by memalign(). Moreover, a subsequent > >>> madvise(MADV_HUGEPAGE) call does not split the VMA because it already > >>> has the same advice. > >>> > >>> This causes the test to fail because the check_huge_xxx() helpers > >>> incorrectly require the address returned by memalign() to match the > >>> VMA start address reported in /proc/self/smaps. > >>> > >>> Address these issues by applying MADV_NOHUGEPAGE after faulting in the > >>> huge page, preventing khugepaged from collapsing it again, and instead of > >>> relying on /proc/self/smaps, use /proc/self/pagemap and > >>> /proc/kpageflags to detect huge-page mappings and large folios: > >>> > >>> 1. If hpage_size == pmd_pagesize, check PAGE_IS_HUGE instead of > >>> using check_large_folios(), since only the mapping type matters. > >>> This identifies PMD-mapped huge pages. > >>> 2. Otherwise, use check_large_folios() to detect large folios. This > >>> covers mTHP cases. > >>> 3. Check the folio flags according to the type of huge page. > >>> > >>> Also, current usage of memalign() would result memory area may > >>> unexpectedly merge with an adjacent VMA, causing tests > >>> that inspect it through /proc/self/smaps to fail. > >> I'm not particularly happy about this. > >> > >> Relying on VMA merging details rather hints that we shouldn't be using smaps to > >> query some stats/properties. > >> > >> Which exact things are test querying through /proc/self/smaps? Could we convert > >> the code to just query that stuff through different interfaces? > > > > Well, users currently for using /proc/self/smaps are for check vm-flags: > > - guard-regions where using check_vmflags_guard() > > - pfnmap test where uses check_vmflag_pfnmap() > > Most vm-flags should not be an issue when it comes to merging. The only > exception are vmflags that do not prevent VMA merging. > > So it's VM_SOFTDIRTY and VM_MAYBE_GUARD. And I agree that for guard-regions.c we > likely have to care such that we don't merge by accident with other VMAs (guard > regions). > > But that's independent of memalign. > > ptr = mmap_(self, variant, NULL, 10 * page_size, PROT_READ | PROT_WRITE, 0, 0); > ASSERT_FALSE(check_vmflag_guard(ptr)); > > could be problematic on its own (unlikely but possible). Yes. That's why I'm think it would be good to use alloc_isolated_mem() in case of ANON mapping for this case. > > For other flags, you really only have to find the smaps area that covers the > given address and look at the vm-flags. > > Or am I missing something important? > > (merging vnas with VM_PFNMAP is impossible right now IIRC) No. what I want to say including the patch #3 is for the above case where you point out -- ASSERT_FALSE(check_vmflag_guard(ptr)). Since we don't have any interface to check vm_flags execpt smap and for memory mmaped with anon would have a chance to merge, We need something to replace memalign() with preventing unexpected VMA merge. (But I forgot to change the those case in guard test case). -- Sincerely, Yeoreum Yun
On 9/10/26 13:16, Yeoreum Yun wrote: > On Thu, Sep 10, 2026 at 12:45:39PM +0200, David Hildenbrand (Arm) wrote: >> On 9/10/26 12:31, Yeoreum Yun wrote: >>> >>> Well, users currently for using /proc/self/smaps are for check vm-flags: >>> - guard-regions where using check_vmflags_guard() >>> - pfnmap test where uses check_vmflag_pfnmap() >> >> Most vm-flags should not be an issue when it comes to merging. The only >> exception are vmflags that do not prevent VMA merging. >> >> So it's VM_SOFTDIRTY and VM_MAYBE_GUARD. And I agree that for guard-regions.c we >> likely have to care such that we don't merge by accident with other VMAs (guard >> regions). >> >> But that's independent of memalign. >> >> ptr = mmap_(self, variant, NULL, 10 * page_size, PROT_READ | PROT_WRITE, 0, 0); >> ASSERT_FALSE(check_vmflag_guard(ptr)); >> >> could be problematic on its own (unlikely but possible). > > Yes. That's why I'm think it would be good to use alloc_isolated_mem() > in case of ANON mapping for this case. It's really only guest-region code that needs this. > >> >> For other flags, you really only have to find the smaps area that covers the >> given address and look at the vm-flags. >> >> Or am I missing something important? >> >> (merging vnas with VM_PFNMAP is impossible right now IIRC) > > No. what I want to say including the patch #3 is for the above case > where you point out -- ASSERT_FALSE(check_vmflag_guard(ptr)). > > Since we don't have any interface to check vm_flags execpt smap > and for memory mmaped with anon would have a chance to merge, > We need something to replace memalign() with preventing unexpected VMA > merge. Only for the cases that actually really needs this, which is in my understanding guard-regions. And I repeat, this is not a memalign() problem. -- Cheers, David
On Thu, Sep 10, 2026 at 01:23:15PM +0200, David Hildenbrand (Arm) wrote: > On 9/10/26 13:16, Yeoreum Yun wrote: > > On Thu, Sep 10, 2026 at 12:45:39PM +0200, David Hildenbrand (Arm) wrote: > >> On 9/10/26 12:31, Yeoreum Yun wrote: > >>> > >>> Well, users currently for using /proc/self/smaps are for check vm-flags: > >>> - guard-regions where using check_vmflags_guard() > >>> - pfnmap test where uses check_vmflag_pfnmap() > >> > >> Most vm-flags should not be an issue when it comes to merging. The only > >> exception are vmflags that do not prevent VMA merging. > >> > >> So it's VM_SOFTDIRTY and VM_MAYBE_GUARD. And I agree that for guard-regions.c we > >> likely have to care such that we don't merge by accident with other VMAs (guard > >> regions). > >> > >> But that's independent of memalign. > >> > >> ptr = mmap_(self, variant, NULL, 10 * page_size, PROT_READ | PROT_WRITE, 0, 0); > >> ASSERT_FALSE(check_vmflag_guard(ptr)); > >> > >> could be problematic on its own (unlikely but possible). > > > > Yes. That's why I'm think it would be good to use alloc_isolated_mem() > > in case of ANON mapping for this case. > > It's really only guest-region code that needs this. > > > > >> > >> For other flags, you really only have to find the smaps area that covers the > >> given address and look at the vm-flags. > >> > >> Or am I missing something important? > >> > >> (merging vnas with VM_PFNMAP is impossible right now IIRC) > > > > No. what I want to say including the patch #3 is for the above case > > where you point out -- ASSERT_FALSE(check_vmflag_guard(ptr)). > > > > Since we don't have any interface to check vm_flags execpt smap > > and for memory mmaped with anon would have a chance to merge, > > We need something to replace memalign() with preventing unexpected VMA > > merge. > Only for the cases that actually really needs this, which is in my understanding > guard-regions. > > And I repeat, this is not a memalign() problem. Yes. I'm not claim memalign() is problem but want to prevent unwanted VMA merge. That's all. And as I mentioned in another reply, if we get rid of ASSERT_FALSE(check_vmflag_guard(ptr)), this patch would be droppable. -- Sincerely, Yeoreum Yun
© 2016 - 2026 Red Hat, Inc.