drivers/gpu/drm/drm_gpuvm.c | 11 +- include/drm/drm_gpuvm.h | 6 + include/linux/bpf_verifier.h | 3 + include/linux/folio_pool.h | 279 +++++++++++++++++++++++++++++++++++++ include/net/netfilter/nf_tables.h | 3 + kernel/bpf/syscall.c | 17 ++- kernel/bpf/verifier.c | 16 ++- kernel/locking/lockdep.c | 278 +++++++++++++++++++++++++++++------- kernel/locking/lockdep_internals.h | 2 + kernel/locking/lockdep_proc.c | 16 ++- lib/Makefile | 2 +- lib/folio_pool.c | 230 ++++++++++++++++++++++++++++++ net/netfilter/nf_tables_api.c | 51 +++++-- 13 files changed, 837 insertions(+), 77 deletions(-)
Introduce two light-weight bump allocators backed directly by compound
folio pages from the buddy allocator (skipping SLUB metadata overhead):
1. struct folio_scratchpad - Variable-sized, alignment-aware bump allocator
for bursty, append-mostly lifecycles with bulk teardown.
2. struct folio_pool - Fixed uniform-slot bump allocator for homogeneous
descriptors.
These serve 2 kinds of use-cases:
1: Ephemeral Batch/transaction Lifecycles (Netfilter, BPF, DRM GPUVM)
- Transactional subsystems allocate bursts of hundreds or thousands of
homogeneous or heterogeneous objects, only to tear them all down
simultaneously at batch completion or error abort.
- Under SLUB: Each descriptor incurs freelist traversal, lock
contention, and O(N) kfree() loops on teardown.
- Under Folio-Scratchpad: Allocations are straight-line pointer bumps,
bulk teardown is O(1) folio_put(), and consecutive Netlink
transactions reuse warm L1/L2 cachelines without buddy lock churn
via folio_scratchpad_reset().
2: Dynamic Long-Lived Graph Scaling (Lockdep)
- Core subsystems historically avoided SLUB by statically
preallocating massive compile-time arrays in .bss
(e.g. list_entries[32768], 1.31 MB) to prevent SLUB allocator
recursion deadlocks.
- Fixed arrays exhaust their limit and permanently disable validation
(BUG: MAX_LOCKDEP_ENTRIES too low!).
- Under Folio-Pool: Lockdep uses a 4,096-entry __initdata bootstrap
buffer for early boot prior to buddy initialization, then seamlessly
migrates and compacts all dependency edges into 64 KB direct-map
folios at late_initcall.
- Permanent static .bss allocation for dependency edges is reduced
from 1.31 MB to 0 KB (100% reclamation at free_initmem()).
Core Primitives in include/linux/folio_pool.h & lib/folio_pool.c:
-----------------------------------------------------------------
0. struct folio_scratchpad (Variable-Sized Elements):
- Dynamic alignment-aware bump pointer with zero padding waste.
- API: folio_scratchpad_init(), folio_scratchpad_alloc(),
folio_scratchpad_reset(), folio_scratchpad_free(),
folio_scratchpad_stats().
- Typed macros: folio_scratchpad_alloc_obj(),
folio_scratchpad_alloc_bytes(), DEFINE_FREE(folio_scratchpad, ...).
1. struct folio_pool (Fixed Uniform Elements):
- Thin wrapper embedding struct folio_scratchpad for homogeneous objects.
- API: folio_pool_init(), folio_pool_alloc(), folio_pool_free(),
folio_pool_stats().
- Typed macros: folio_pool_alloc_obj(), folio_pool_alloc_type().
2. Encapsulated Slab Discrimination & Reallocation:
- is_folio_pool_ptr(ptr): Direct-map folios are identified via
virt_to_folio(ptr) && !folio_test_slab(f), avoiding custom bitflags.
- folio_scratchpad_free_elem(ptr): Safe no-op for folio-backed objects;
delegates to kvfree() for SLUB/vmalloc fallback objects.
- folio_scratchpad_realloc(ptr, old_size, new_size, gfp): Delegates to
krealloc() for SLUB objects; allocates a fresh SLUB buffer and copies
payload for folio-backed descriptors.
3. Subsystem Autonomy & Runtime Static-Key Control:
- DEFINE_FOLIO_POOL_STATIC_KEY_PARAM() allows maintainers of each
subsystem (Netfilter, Lockdep, DRM, BPF) to choose their own default
enablement policy (DECLARE_STATIC_KEY_TRUE vs FALSE) and independent
module/boot parameters (e.g. lockdep.folio_pool, nf_tables.trans_scratchpad).
- Enables isolated, zero-overhead A/B benchmarking and production triage
without cross-subsystem coupling or kernel recompilation.
Active Proofs & Performance Profiling:
--------------------------------------
0. Virtualized Guest Netfilter Benchmark (KVM / virtme-ng ftrace):
Controlled A/B benchmark across 10,025 Netlink transaction descriptors
and 5 asynchronous commit/abort teardown passes:
Function: nft_trans_alloc (10,025 descriptor allocations)
-----------------------------------------------------------------------------
Mode Total Time Avg / Alloc Delta
-----------------------------------------------------------------------------
folio_scratchpad Enabled (Bump) 9,158.02 us 0.914 us -19.4%
SLUB Fallback (kzalloc) 11,366.53 us 1.134 us Baseline
-----------------------------------------------------------------------------
Net allocation latency reduction: -2,208.52 us (-2.21 ms)
Function: nf_tables_trans_destroy_work (5 asynchronous teardown passes)
-----------------------------------------------------------------------------
Mode Total Time Avg / Pass Delta
-----------------------------------------------------------------------------
folio_scratchpad Enabled (Bulk) 36,065.65 us 7.21 ms -8.6%
SLUB Fallback (O(N) kfree) 39,449.65 us 7.89 ms Baseline
-----------------------------------------------------------------------------
Net teardown latency reduction: -3,384.00 us (-3.38 ms)
Combined In-Kernel Netlink Transaction Lifecycle Savings: -5.59 ms (-13.3%)
1. Bare-Metal Host Netfilter Profile (AMD Zen 3 Hardware PMU & ftrace):
- Asynchronous Teardown: 53.77 ms vs 56.03 ms (-2.26 ms / -4.0%).
- Instruction Retirement: 281.29M -> 282.46M (-1,174,062 instructions).
- Branch Instructions: 62.94M -> 63.20M (-265,961 branches).
2. Bare-Metal Host Lockdep Telemetry (Physical Host frodo, AMD Zen SMP):
- Pre-Buddy Watermark: 928/4096 bootstrap entries consumed during boot.
- Graph Compaction: Migrated and compacted bootstrap dependency edges
into folio_pool at late_initcall; bootstrap array reclaimed by
free_initmem() (0 KB permanent .bss cost).
- Live procfs Telemetry (/proc/lockdep_stats):
lock-classes: 996 [max: 8192]
direct dependencies: 2324 [dynamic: 7 x 64 kB, tail: 17 kB/64 kB]
indirect dependencies: 24612
all direct dependencies: 81884
zapped classes: 2
zapped lock chains: 173
- Active Invariant Validation: Live module unloads exercised modernized
zap_class() direct list unlinking across dynamic folios with zero
global bitmap scans and zero assertions.
What's Unchanged:
-----------------
- All consumer object structures, alignment rules, and lifecycle boundaries
remain strictly identical.
- Memory safety invariants are preserved: individual deallocation is a safe
no-op for arena-backed objects while SLUB objects continue to use kfree().
Subsystem Adoptions in this Series:
-----------------------------------
0. lib/folio_pool: Core direct-map large-folio pool and scratchpad allocators.
1. netfilter/nf_tables: Pack netlink transaction descriptors into per-net
trans_scratchpad with bulk reclaim & warm chunk reset.
2. bpf/verifier: Route verifier stack state node allocations to folio_pool.
3. drm/gpuvm: Route gpuva_op allocations to folio_scratchpad.
4. bpf/syscall: Route generic_map_update_batch key/value allocations to
folio_scratchpad.
5. locking/lockdep: Fallback to folio_pool in alloc_list_entry when static
pool is full and expose dynamic folio telemetry in /proc/lockdep_stats.
6. locking/lockdep: Traverse adjacency lists directly in zap_class().
7. locking/lockdep: Shrink static list_entries array to early bootstrap buffer.
8. locking/lockdep: Migrate and compact boot-time dependency graph from __initdata.
Patches in this series:
-----------------------
[PATCH 1/9] lib/folio_pool: Introduce Direct-Map Large Folio Pool & Scratchpad bump allocators
[PATCH 2/9] netfilter/nf_tables: Add folio_scratchpad collector to struct nftables_pernet
[PATCH 3/9] bpf/verifier: Route verifier stack state node allocations to folio_pool
[PATCH 4/9] drm/gpuvm: Route gpuva_op allocations to folio_scratchpad
[PATCH 5/9] bpf/syscall: Route generic_map_update_batch key/value allocations to folio_scratchpad
[PATCH 6/9] locking/lockdep: Fallback to folio_pool in alloc_list_entry when static pool is full
[PATCH 7/9] locking/lockdep: Traverse adjacency lists directly in zap_class()
[PATCH 8/9] locking/lockdep: Shrink static list_entries array to early bootstrap buffer
[PATCH 9/9] locking/lockdep: Migrate and compact boot-time dependency graph from __initdata
Signed-off-by: Jim Cromie <jim.cromie@gmail.com>
---
Jim Cromie (9):
lib/folio_pool: Introduce Direct-Map Large Folio Pool & Scratchpad bump allocators
netfilter/nf_tables: Add folio_scratchpad collector to struct nftables_pernet
bpf/verifier: Route verifier stack state node allocations to folio_pool
drm/gpuvm: Route gpuva_op allocations to folio_scratchpad
bpf/syscall: Route generic_map_update_batch key/value allocations to folio_scratchpad
locking/lockdep: Fallback to folio_pool in alloc_list_entry when static pool is full
locking/lockdep: Traverse adjacency lists directly in zap_class()
locking/lockdep: Shrink static list_entries array to early bootstrap buffer
locking/lockdep: Migrate and compact boot-time dependency graph from __initdata
drivers/gpu/drm/drm_gpuvm.c | 11 +-
include/drm/drm_gpuvm.h | 6 +
include/linux/bpf_verifier.h | 3 +
include/linux/folio_pool.h | 279 +++++++++++++++++++++++++++++++++++++
include/net/netfilter/nf_tables.h | 3 +
kernel/bpf/syscall.c | 17 ++-
kernel/bpf/verifier.c | 16 ++-
kernel/locking/lockdep.c | 278 +++++++++++++++++++++++++++++-------
kernel/locking/lockdep_internals.h | 2 +
kernel/locking/lockdep_proc.c | 16 ++-
lib/Makefile | 2 +-
lib/folio_pool.c | 230 ++++++++++++++++++++++++++++++
net/netfilter/nf_tables_api.c | 51 +++++--
13 files changed, 837 insertions(+), 77 deletions(-)
---
base-commit: 1a7ac48cd3e62f281ce62f8ed89dc2cc8866eaa4
change-id: 20260815-folio-pool-v1-cb75c247692c
Best regards,
--
Jim Cromie <jim.cromie@gmail.com>
On Mon, Aug 17, 2026 at 11:22:14AM -0600, Jim Cromie via B4 Relay wrote: > Introduce two light-weight bump allocators backed directly by compound > folio pages from the buddy allocator (skipping SLUB metadata overhead): > > 1. struct folio_scratchpad - Variable-sized, alignment-aware bump allocator > for bursty, append-mostly lifecycles with bulk teardown. > 2. struct folio_pool - Fixed uniform-slot bump allocator for homogeneous > descriptors. This is an inappropriate use of folios. Folios are for memory that's in the page cache and anonymous user memory, not for general kernel use. They would need their own memdesc type. But Byungchul is already doing that work. So you should just reuse what he's done. > These serve 2 kinds of use-cases: > > 1: Ephemeral Batch/transaction Lifecycles (Netfilter, BPF, DRM GPUVM) > > - Transactional subsystems allocate bursts of hundreds or thousands of > homogeneous or heterogeneous objects, only to tear them all down > simultaneously at batch completion or error abort. > > - Under SLUB: Each descriptor incurs freelist traversal, lock > contention, and O(N) kfree() loops on teardown. > > - Under Folio-Scratchpad: Allocations are straight-line pointer bumps, > bulk teardown is O(1) folio_put(), and consecutive Netlink > transactions reuse warm L1/L2 cachelines without buddy lock churn > via folio_scratchpad_reset(). > > 2: Dynamic Long-Lived Graph Scaling (Lockdep) > > - Core subsystems historically avoided SLUB by statically > preallocating massive compile-time arrays in .bss > (e.g. list_entries[32768], 1.31 MB) to prevent SLUB allocator > recursion deadlocks. > > - Fixed arrays exhaust their limit and permanently disable validation > (BUG: MAX_LOCKDEP_ENTRIES too low!). > > - Under Folio-Pool: Lockdep uses a 4,096-entry __initdata bootstrap > buffer for early boot prior to buddy initialization, then seamlessly > migrates and compacts all dependency edges into 64 KB direct-map > folios at late_initcall. > > - Permanent static .bss allocation for dependency edges is reduced > from 1.31 MB to 0 KB (100% reclamation at free_initmem()). > > Core Primitives in include/linux/folio_pool.h & lib/folio_pool.c: > ----------------------------------------------------------------- > 0. struct folio_scratchpad (Variable-Sized Elements): > - Dynamic alignment-aware bump pointer with zero padding waste. > - API: folio_scratchpad_init(), folio_scratchpad_alloc(), > folio_scratchpad_reset(), folio_scratchpad_free(), > folio_scratchpad_stats(). > - Typed macros: folio_scratchpad_alloc_obj(), > folio_scratchpad_alloc_bytes(), DEFINE_FREE(folio_scratchpad, ...). > > 1. struct folio_pool (Fixed Uniform Elements): > - Thin wrapper embedding struct folio_scratchpad for homogeneous objects. > - API: folio_pool_init(), folio_pool_alloc(), folio_pool_free(), > folio_pool_stats(). > - Typed macros: folio_pool_alloc_obj(), folio_pool_alloc_type(). > > 2. Encapsulated Slab Discrimination & Reallocation: > - is_folio_pool_ptr(ptr): Direct-map folios are identified via > virt_to_folio(ptr) && !folio_test_slab(f), avoiding custom bitflags. > - folio_scratchpad_free_elem(ptr): Safe no-op for folio-backed objects; > delegates to kvfree() for SLUB/vmalloc fallback objects. > - folio_scratchpad_realloc(ptr, old_size, new_size, gfp): Delegates to > krealloc() for SLUB objects; allocates a fresh SLUB buffer and copies > payload for folio-backed descriptors. > > 3. Subsystem Autonomy & Runtime Static-Key Control: > - DEFINE_FOLIO_POOL_STATIC_KEY_PARAM() allows maintainers of each > subsystem (Netfilter, Lockdep, DRM, BPF) to choose their own default > enablement policy (DECLARE_STATIC_KEY_TRUE vs FALSE) and independent > module/boot parameters (e.g. lockdep.folio_pool, nf_tables.trans_scratchpad). > - Enables isolated, zero-overhead A/B benchmarking and production triage > without cross-subsystem coupling or kernel recompilation. > > Active Proofs & Performance Profiling: > -------------------------------------- > 0. Virtualized Guest Netfilter Benchmark (KVM / virtme-ng ftrace): > Controlled A/B benchmark across 10,025 Netlink transaction descriptors > and 5 asynchronous commit/abort teardown passes: > > Function: nft_trans_alloc (10,025 descriptor allocations) > ----------------------------------------------------------------------------- > Mode Total Time Avg / Alloc Delta > ----------------------------------------------------------------------------- > folio_scratchpad Enabled (Bump) 9,158.02 us 0.914 us -19.4% > SLUB Fallback (kzalloc) 11,366.53 us 1.134 us Baseline > ----------------------------------------------------------------------------- > Net allocation latency reduction: -2,208.52 us (-2.21 ms) > > Function: nf_tables_trans_destroy_work (5 asynchronous teardown passes) > ----------------------------------------------------------------------------- > Mode Total Time Avg / Pass Delta > ----------------------------------------------------------------------------- > folio_scratchpad Enabled (Bulk) 36,065.65 us 7.21 ms -8.6% > SLUB Fallback (O(N) kfree) 39,449.65 us 7.89 ms Baseline > ----------------------------------------------------------------------------- > Net teardown latency reduction: -3,384.00 us (-3.38 ms) > > Combined In-Kernel Netlink Transaction Lifecycle Savings: -5.59 ms (-13.3%) > > 1. Bare-Metal Host Netfilter Profile (AMD Zen 3 Hardware PMU & ftrace): > - Asynchronous Teardown: 53.77 ms vs 56.03 ms (-2.26 ms / -4.0%). > - Instruction Retirement: 281.29M -> 282.46M (-1,174,062 instructions). > - Branch Instructions: 62.94M -> 63.20M (-265,961 branches). > > 2. Bare-Metal Host Lockdep Telemetry (Physical Host frodo, AMD Zen SMP): > - Pre-Buddy Watermark: 928/4096 bootstrap entries consumed during boot. > - Graph Compaction: Migrated and compacted bootstrap dependency edges > into folio_pool at late_initcall; bootstrap array reclaimed by > free_initmem() (0 KB permanent .bss cost). > - Live procfs Telemetry (/proc/lockdep_stats): > lock-classes: 996 [max: 8192] > direct dependencies: 2324 [dynamic: 7 x 64 kB, tail: 17 kB/64 kB] > indirect dependencies: 24612 > all direct dependencies: 81884 > zapped classes: 2 > zapped lock chains: 173 > - Active Invariant Validation: Live module unloads exercised modernized > zap_class() direct list unlinking across dynamic folios with zero > global bitmap scans and zero assertions. > > What's Unchanged: > ----------------- > - All consumer object structures, alignment rules, and lifecycle boundaries > remain strictly identical. > - Memory safety invariants are preserved: individual deallocation is a safe > no-op for arena-backed objects while SLUB objects continue to use kfree(). > > Subsystem Adoptions in this Series: > ----------------------------------- > 0. lib/folio_pool: Core direct-map large-folio pool and scratchpad allocators. > 1. netfilter/nf_tables: Pack netlink transaction descriptors into per-net > trans_scratchpad with bulk reclaim & warm chunk reset. > 2. bpf/verifier: Route verifier stack state node allocations to folio_pool. > 3. drm/gpuvm: Route gpuva_op allocations to folio_scratchpad. > 4. bpf/syscall: Route generic_map_update_batch key/value allocations to > folio_scratchpad. > 5. locking/lockdep: Fallback to folio_pool in alloc_list_entry when static > pool is full and expose dynamic folio telemetry in /proc/lockdep_stats. > 6. locking/lockdep: Traverse adjacency lists directly in zap_class(). > 7. locking/lockdep: Shrink static list_entries array to early bootstrap buffer. > 8. locking/lockdep: Migrate and compact boot-time dependency graph from __initdata. > > Patches in this series: > ----------------------- > [PATCH 1/9] lib/folio_pool: Introduce Direct-Map Large Folio Pool & Scratchpad bump allocators > [PATCH 2/9] netfilter/nf_tables: Add folio_scratchpad collector to struct nftables_pernet > [PATCH 3/9] bpf/verifier: Route verifier stack state node allocations to folio_pool > [PATCH 4/9] drm/gpuvm: Route gpuva_op allocations to folio_scratchpad > [PATCH 5/9] bpf/syscall: Route generic_map_update_batch key/value allocations to folio_scratchpad > [PATCH 6/9] locking/lockdep: Fallback to folio_pool in alloc_list_entry when static pool is full > [PATCH 7/9] locking/lockdep: Traverse adjacency lists directly in zap_class() > [PATCH 8/9] locking/lockdep: Shrink static list_entries array to early bootstrap buffer > [PATCH 9/9] locking/lockdep: Migrate and compact boot-time dependency graph from __initdata > > Signed-off-by: Jim Cromie <jim.cromie@gmail.com> > --- > Jim Cromie (9): > lib/folio_pool: Introduce Direct-Map Large Folio Pool & Scratchpad bump allocators > netfilter/nf_tables: Add folio_scratchpad collector to struct nftables_pernet > bpf/verifier: Route verifier stack state node allocations to folio_pool > drm/gpuvm: Route gpuva_op allocations to folio_scratchpad > bpf/syscall: Route generic_map_update_batch key/value allocations to folio_scratchpad > locking/lockdep: Fallback to folio_pool in alloc_list_entry when static pool is full > locking/lockdep: Traverse adjacency lists directly in zap_class() > locking/lockdep: Shrink static list_entries array to early bootstrap buffer > locking/lockdep: Migrate and compact boot-time dependency graph from __initdata > > drivers/gpu/drm/drm_gpuvm.c | 11 +- > include/drm/drm_gpuvm.h | 6 + > include/linux/bpf_verifier.h | 3 + > include/linux/folio_pool.h | 279 +++++++++++++++++++++++++++++++++++++ > include/net/netfilter/nf_tables.h | 3 + > kernel/bpf/syscall.c | 17 ++- > kernel/bpf/verifier.c | 16 ++- > kernel/locking/lockdep.c | 278 +++++++++++++++++++++++++++++------- > kernel/locking/lockdep_internals.h | 2 + > kernel/locking/lockdep_proc.c | 16 ++- > lib/Makefile | 2 +- > lib/folio_pool.c | 230 ++++++++++++++++++++++++++++++ > net/netfilter/nf_tables_api.c | 51 +++++-- > 13 files changed, 837 insertions(+), 77 deletions(-) > --- > base-commit: 1a7ac48cd3e62f281ce62f8ed89dc2cc8866eaa4 > change-id: 20260815-folio-pool-v1-cb75c247692c > > Best regards, > -- > Jim Cromie <jim.cromie@gmail.com> > > >
On 8/17/26 19:22, Jim Cromie via B4 Relay wrote: > Introduce two light-weight bump allocators backed directly by compound > folio pages from the buddy allocator (skipping SLUB metadata overhead): We are running into a bit of a naming problem here: folios are really supposed to describe things that we will map to user space, whereby we have to track mapping/index/mapcount etc. That's really just anon + pagecache memory. Right now that naming difference doesn't matter, because each compound page really is a folio. But once we start decoupling folios from the rest (memdescs), we really only expect anon+pagecache to resemble "folios". There will be other compound page types (similar to slab pages), and things like page_folio() will fail for them (like slab pages). And we are actually working on making page_folio() fail earlier for anything that didn't go through page_rmappable_folio() as part of the (folio) allocation. So, does what you do here fit the definition of folios, or is it really just some kind of compound pages? -- Cheers, David
© 2016 - 2026 Red Hat, Inc.