[RFC PATCH 0/2] mm/slub: reduce list_lock contention with slab parking

Hao Li posted 2 patches 1 month ago
Only 1 patches received!
mm/slub.c | 326 +++++++++++++++++++++++++++++++++++++++++++++---------
1 file changed, 274 insertions(+), 52 deletions(-)
[RFC PATCH 0/2] mm/slub: reduce list_lock contention with slab parking
Posted by Hao Li 1 month ago
This patch series might sound a bit wild, but the initial numbers don't
look too bad so far. I would really appreciate any feedback and
discussion :)

On a will-it-scale mmap1 run with 192 processes, list_lock is the top
contention point: __slab_free() and __refill_objects_node() together
spend 44% of cycles in native_queued_spin_lock_slowpath. The free
slowpath takes the lock mainly to add slabs that became non-full to the
partial list.

By adding extra instrumentation to __slab_free(), I collect the following
data for the maple_node cache (in counts):

  partial->partial    843017414
  full->partial       550719384
  partial->empty      17459564
  full->empty         2

We can see that full -> partial transitions account for a significant
proportion, and optimizing them can help reduce lock contention to some
extent.

This series introduces the parking mechanism to address this issue.
When the trylock fails during a full -> partial/empty transition,
__slab_free() parks the slab on a per-node llist instead of waiting. The
paths that consume the partial list (sheaf refill, alloc slowpath,
shrink, cache destruction) unpark the slabs after taking the lock, and a
delayed work covers the case where none of them runs.

Patch 1 cleans up the case handling in __slab_free(), no functional
change. Patch 2 introduces the parking mechanism.

Tested with will-it-scale mmap1 (192 processes).

Summary data
------------

throughput             29237910 -> 35585663  (+21.7%)
alloc_slab,free_slab   -52%
cmpxchg_double_fail    -85%

perf data without this patchset:
- 44.22% [kernel] [k] native_queued_spin_lock_slowpath
   43.48% native_queued_spin_lock_slowpath
    - _raw_spin_lock_irqsave
       - 23.80% __refill_objects_node
       - 19.12% __slab_free

perf data with this patchset:
- 30.39% [kernel] [k] native_queued_spin_lock_slowpath
   29.82% native_queued_spin_lock_slowpath
    - _raw_spin_lock_irqsave
       - 29.06% __refill_objects_node

Additionally, the number of partial slabs and the number of objects show
no noticeable change before and after applying this patchset, indicating
that this change has a negligible impact on slab fragmentation.

Detailed data
-------------

metric                              before             after             delta      change
==========================================================================================
alloc_fastpath                     155,417           168,534            13,117      +8.44%
alloc_slab                      55,679,646        26,702,510       -28,977,136     -52.04%
alloc_slowpath                           0                 0                 0      +0.00%
barn_get                             2,715             2,771                56      +2.06%
barn_get_fail                            2                 0                -2    -100.00%
barn_put                             2,715             2,770                55      +2.03%
barn_put_fail                1,370,488,457     1,668,257,974       297,769,517     +21.73%
cmpxchg_double_fail              3,827,935           549,427        -3,278,508     -85.65%
free_add_partial             1,684,340,964     2,161,921,625       477,580,661     +28.35%
free_fastpath                       31,766            32,979             1,213      +3.82%
free_rcu_sheaf              43,855,689,417    53,384,314,553     9,528,625,136     +21.73%
free_rcu_sheaf_fail                      0                 0                 0      +0.00%
free_remove_partial             55,678,459        26,685,274       -28,993,185     -52.07%
free_slab                       55,678,459        26,701,099       -28,977,360     -52.04%
free_slowpath                  107,771,739        63,197,344       -44,574,395     -41.36%
min_partial                              5                 5                 0      +0.00%
object_size                            256               256                 0      +0.00%
objects                             14,504            14,336              -168      -1.16%
objects_partial                     14,504            14,208              -296      -2.04%
objs_per_slab                           64                64                 0      +0.00%
park_slab                                -     2,147,963,530                 -      absent
partial                              1,398             1,367               -31      -2.22%
sheaf_alloc                    743,660,288     1,284,618,991       540,958,703     +72.74%
sheaf_capacity                          32                32                 0      +0.00%
sheaf_flush                 43,855,651,858    53,384,274,983     9,528,623,125     +21.73%
sheaf_free                     743,660,280     1,284,618,967       540,958,687     +72.74%
sheaf_prefill_fast          17,585,335,850    21,378,958,833     3,793,622,983     +21.57%
sheaf_prefill_oversize                   0                 0                 0      +0.00%
sheaf_prefill_slow                   2,060             2,051                -9      -0.44%
sheaf_refill                43,963,424,505    53,447,473,167     9,484,048,662     +21.57%
sheaf_return_fast           17,585,336,335    21,378,959,334     3,793,622,999     +21.57%
sheaf_return_slow                    1,402             1,277              -125      -8.92%
slabs                                1,398             1,369               -29      -2.07%
total_objects                       89,472            87,616            -1,856      -2.07%
unpark_event                             -       315,397,838                 -      absent
unpark_slab                              -     2,147,963,530                 -      absent

derived                                                  before             after      change
=============================================================================================
PARK_SLAB / FREE_ADD_PARTIAL                                  -            99.35%      absent
UNPARK_SLAB / UNPARK_EVENT                                    -              6.81      absent
PARK_SLAB - UNPARK_SLAB                                       -                 0      absent
page allocator churn (alloc_slab + free_slab)       111,358,105        53,403,609     -52.04%

Note that some metrics have very small absolute values (such as
alloc_fastpath, partial, slabs, and total_objects) and are subject to
noise. Across multiple test runs, their rate of change fluctuates
between positive and negative, which supports the hypothesis that this
is measurement noise and demonstrates that this patch has no noticeable
impact on these metrics.

For metrics with large absolute values, their trends are distinct. The
data indicates that the primary benefit of this approach is
significantly relieved pressure on the buddy system, with page allocator
churn reduced by 52%. Additionally, free_slowpath decreases by 41%, and
cmpxchg_double_fail decreases by 85%.

The PARK_SLAB / FREE_ADD_PARTIAL ratio reaches 99.35%, which indicates
that the vast majority of partial slabs are added back to the partial
list via the parking mechanism, reflecting that the lock stayed
saturated and nearly all additions avoided waiting for the lock. The
ratio of UNPARK_SLAB / UNPARK_EVENT shows that each unpark event
processes roughly 6 slabs. PARK_SLAB - UNPARK_SLAB being 0 confirms
that no parked slabs are left stranded.

I also observe increases in both sheaf_alloc and sheaf_free, which
could currently be attributed to faster allocation and free paths
resulting from the overall performance improvement. However, I'm not
sure about this, which is part of why this is posted as an RFC.

Based on slab/for-next.

Hao Li (2):
  mm/slub: make the case handling in __slab_free() easier to follow
  mm/slub: introduce slab parking to reduce list_lock contention

 mm/slub.c | 326 +++++++++++++++++++++++++++++++++++++++++++++---------
 1 file changed, 274 insertions(+), 52 deletions(-)

base-commit: e7f630142df2afccce90555e4972e60008222311
-- 
2.55.0
Re: [RFC PATCH 0/2] mm/slub: reduce list_lock contention with slab parking
Posted by Vlastimil Babka (SUSE) 2 weeks, 6 days ago
On 8/24/26 14:19, Hao Li wrote:
> This patch series might sound a bit wild, but the initial numbers don't
> look too bad so far. I would really appreciate any feedback and
> discussion :)
> 
> On a will-it-scale mmap1 run with 192 processes, list_lock is the top
> contention point: __slab_free() and __refill_objects_node() together
> spend 44% of cycles in native_queued_spin_lock_slowpath. The free
> slowpath takes the lock mainly to add slabs that became non-full to the
> partial list.
> 
> By adding extra instrumentation to __slab_free(), I collect the following
> data for the maple_node cache (in counts):
> 
>   partial->partial    843017414
>   full->partial       550719384
>   partial->empty      17459564
>   full->empty         2
> 
> We can see that full -> partial transitions account for a significant
> proportion, and optimizing them can help reduce lock contention to some
> extent.
> 
> This series introduces the parking mechanism to address this issue.
> When the trylock fails during a full -> partial/empty transition,
> __slab_free() parks the slab on a per-node llist instead of waiting. The
> paths that consume the partial list (sheaf refill, alloc slowpath,
> shrink, cache destruction) unpark the slabs after taking the lock, and a
> delayed work covers the case where none of them runs.
> 
> Patch 1 cleans up the case handling in __slab_free(), no functional
> change. Patch 2 introduces the parking mechanism.
> 
> Tested with will-it-scale mmap1 (192 processes).
> 
> Summary data
> ------------
> 
> throughput             29237910 -> 35585663  (+21.7%)
> alloc_slab,free_slab   -52%
> cmpxchg_double_fail    -85%
> 
> perf data without this patchset:
> - 44.22% [kernel] [k] native_queued_spin_lock_slowpath
>    43.48% native_queued_spin_lock_slowpath
>     - _raw_spin_lock_irqsave
>        - 23.80% __refill_objects_node
>        - 19.12% __slab_free
> 
> perf data with this patchset:
> - 30.39% [kernel] [k] native_queued_spin_lock_slowpath
>    29.82% native_queued_spin_lock_slowpath
>     - _raw_spin_lock_irqsave
>        - 29.06% __refill_objects_node
> 
> Additionally, the number of partial slabs and the number of objects show
> no noticeable change before and after applying this patchset, indicating
> that this change has a negligible impact on slab fragmentation.
> 
> Detailed data
> -------------
> 
> metric                              before             after             delta      change
> ==========================================================================================
> alloc_fastpath                     155,417           168,534            13,117      +8.44%
> alloc_slab                      55,679,646        26,702,510       -28,977,136     -52.04%

It's interesting that this is reduced so much. Is it because parked slabs
cause more slabs to stay around for reuse, despite they are unparked
immediately when trying to allocate/refill? That seems odd?

> alloc_slowpath                           0                 0                 0      +0.00%
> barn_get                             2,715             2,771                56      +2.06%
> barn_get_fail                            2                 0                -2    -100.00%
> barn_put                             2,715             2,770                55      +2.03%
> barn_put_fail                1,370,488,457     1,668,257,974       297,769,517     +21.73%
> cmpxchg_double_fail              3,827,935           549,427        -3,278,508     -85.65%
> free_add_partial             1,684,340,964     2,161,921,625       477,580,661     +28.35%
> free_fastpath                       31,766            32,979             1,213      +3.82%
> free_rcu_sheaf              43,855,689,417    53,384,314,553     9,528,625,136     +21.73%

This metric (and others with similar numbers) should not be affected by the
change. Does it mean the benchmark has a fixed time to run, but manages to
do more work in that time thanks to the increased throughput?

> free_rcu_sheaf_fail                      0                 0                 0      +0.00%
> free_remove_partial             55,678,459        26,685,274       -28,993,185     -52.07%
> free_slab                       55,678,459        26,701,099       -28,977,360     -52.04%
> free_slowpath                  107,771,739        63,197,344       -44,574,395     -41.36%
> min_partial                              5                 5                 0      +0.00%
> object_size                            256               256                 0      +0.00%
> objects                             14,504            14,336              -168      -1.16%
> objects_partial                     14,504            14,208              -296      -2.04%
> objs_per_slab                           64                64                 0      +0.00%
> park_slab                                -     2,147,963,530                 -      absent
> partial                              1,398             1,367               -31      -2.22%
> sheaf_alloc                    743,660,288     1,284,618,991       540,958,703     +72.74%
> sheaf_capacity                          32                32                 0      +0.00%
> sheaf_flush                 43,855,651,858    53,384,274,983     9,528,623,125     +21.73%
> sheaf_free                     743,660,280     1,284,618,967       540,958,687     +72.74%
> sheaf_prefill_fast          17,585,335,850    21,378,958,833     3,793,622,983     +21.57%
> sheaf_prefill_oversize                   0                 0                 0      +0.00%
> sheaf_prefill_slow                   2,060             2,051                -9      -0.44%
> sheaf_refill                43,963,424,505    53,447,473,167     9,484,048,662     +21.57%
> sheaf_return_fast           17,585,336,335    21,378,959,334     3,793,622,999     +21.57%
> sheaf_return_slow                    1,402             1,277              -125      -8.92%
> slabs                                1,398             1,369               -29      -2.07%
> total_objects                       89,472            87,616            -1,856      -2.07%
> unpark_event                             -       315,397,838                 -      absent
> unpark_slab                              -     2,147,963,530                 -      absent
> 
> derived                                                  before             after      change
> =============================================================================================
> PARK_SLAB / FREE_ADD_PARTIAL                                  -            99.35%      absent
> UNPARK_SLAB / UNPARK_EVENT                                    -              6.81      absent
> PARK_SLAB - UNPARK_SLAB                                       -                 0      absent
> page allocator churn (alloc_slab + free_slab)       111,358,105        53,403,609     -52.04%
> 
> Note that some metrics have very small absolute values (such as
> alloc_fastpath, partial, slabs, and total_objects) and are subject to
> noise. Across multiple test runs, their rate of change fluctuates
> between positive and negative, which supports the hypothesis that this
> is measurement noise and demonstrates that this patch has no noticeable
> impact on these metrics.
> 
> For metrics with large absolute values, their trends are distinct. The
> data indicates that the primary benefit of this approach is
> significantly relieved pressure on the buddy system, with page allocator
> churn reduced by 52%. Additionally, free_slowpath decreases by 41%, and
> cmpxchg_double_fail decreases by 85%.
> 
> The PARK_SLAB / FREE_ADD_PARTIAL ratio reaches 99.35%, which indicates
> that the vast majority of partial slabs are added back to the partial
> list via the parking mechanism, reflecting that the lock stayed
> saturated and nearly all additions avoided waiting for the lock. The
> ratio of UNPARK_SLAB / UNPARK_EVENT shows that each unpark event
> processes roughly 6 slabs. PARK_SLAB - UNPARK_SLAB being 0 confirms
> that no parked slabs are left stranded.
> 
> I also observe increases in both sheaf_alloc and sheaf_free, which
> could currently be attributed to faster allocation and free paths
> resulting from the overall performance improvement. However, I'm not
> sure about this, which is part of why this is posted as an RFC.
> 
> Based on slab/for-next.
> 
> Hao Li (2):
>   mm/slub: make the case handling in __slab_free() easier to follow
>   mm/slub: introduce slab parking to reduce list_lock contention
> 
>  mm/slub.c | 326 +++++++++++++++++++++++++++++++++++++++++++++---------
>  1 file changed, 274 insertions(+), 52 deletions(-)
> 
> base-commit: e7f630142df2afccce90555e4972e60008222311
Re: [RFC PATCH 0/2] mm/slub: reduce list_lock contention with slab parking
Posted by Hao Li 2 weeks, 2 days ago
On Mon, Sep 07, 2026 at 03:44:08PM +0200, Vlastimil Babka (SUSE) wrote:
> On 8/24/26 14:19, Hao Li wrote:
> > This patch series might sound a bit wild, but the initial numbers don't
> > look too bad so far. I would really appreciate any feedback and
> > discussion :)
> > 
> > On a will-it-scale mmap1 run with 192 processes, list_lock is the top
> > contention point: __slab_free() and __refill_objects_node() together
> > spend 44% of cycles in native_queued_spin_lock_slowpath. The free
> > slowpath takes the lock mainly to add slabs that became non-full to the
> > partial list.
> > 
> > By adding extra instrumentation to __slab_free(), I collect the following
> > data for the maple_node cache (in counts):
> > 
> >   partial->partial    843017414
> >   full->partial       550719384
> >   partial->empty      17459564
> >   full->empty         2
> > 
> > We can see that full -> partial transitions account for a significant
> > proportion, and optimizing them can help reduce lock contention to some
> > extent.
> > 
> > This series introduces the parking mechanism to address this issue.
> > When the trylock fails during a full -> partial/empty transition,
> > __slab_free() parks the slab on a per-node llist instead of waiting. The
> > paths that consume the partial list (sheaf refill, alloc slowpath,
> > shrink, cache destruction) unpark the slabs after taking the lock, and a
> > delayed work covers the case where none of them runs.
> > 
> > Patch 1 cleans up the case handling in __slab_free(), no functional
> > change. Patch 2 introduces the parking mechanism.
> > 
> > Tested with will-it-scale mmap1 (192 processes).
> > 
> > Summary data
> > ------------
> > 
> > throughput             29237910 -> 35585663  (+21.7%)
> > alloc_slab,free_slab   -52%
> > cmpxchg_double_fail    -85%
> > 
> > perf data without this patchset:
> > - 44.22% [kernel] [k] native_queued_spin_lock_slowpath
> >    43.48% native_queued_spin_lock_slowpath
> >     - _raw_spin_lock_irqsave
> >        - 23.80% __refill_objects_node
> >        - 19.12% __slab_free
> > 
> > perf data with this patchset:
> > - 30.39% [kernel] [k] native_queued_spin_lock_slowpath
> >    29.82% native_queued_spin_lock_slowpath
> >     - _raw_spin_lock_irqsave
> >        - 29.06% __refill_objects_node
> > 
> > Additionally, the number of partial slabs and the number of objects show
> > no noticeable change before and after applying this patchset, indicating
> > that this change has a negligible impact on slab fragmentation.
> > 
> > Detailed data
> > -------------
> > 
> > metric                              before             after             delta      change
> > ==========================================================================================
> > alloc_fastpath                     155,417           168,534            13,117      +8.44%
> > alloc_slab                      55,679,646        26,702,510       -28,977,136     -52.04%
> 
> It's interesting that this is reduced so much. Is it because parked slabs
> cause more slabs to stay around for reuse, despite they are unparked
> immediately when trying to allocate/refill? That seems odd?

Yeah, it does look counter-intuitive at first glance. My understanding is that
the parking mechanism helps by bypassing the partial list spinlock, allowing
partial slabs to be parked onto the llist locklessly at any point. Because of
this, the refill path gets to see and reuse more partial slabs instead of
having to allocate fresh ones so frequently.

> 
> > alloc_slowpath                           0                 0                 0      +0.00%
> > barn_get                             2,715             2,771                56      +2.06%
> > barn_get_fail                            2                 0                -2    -100.00%
> > barn_put                             2,715             2,770                55      +2.03%
> > barn_put_fail                1,370,488,457     1,668,257,974       297,769,517     +21.73%
> > cmpxchg_double_fail              3,827,935           549,427        -3,278,508     -85.65%
> > free_add_partial             1,684,340,964     2,161,921,625       477,580,661     +28.35%
> > free_fastpath                       31,766            32,979             1,213      +3.82%
> > free_rcu_sheaf              43,855,689,417    53,384,314,553     9,528,625,136     +21.73%
> 
> This metric (and others with similar numbers) should not be affected by the
> change. Does it mean the benchmark has a fixed time to run, but manages to
> do more work in that time thanks to the increased throughput?

Exactly! In fact, the increase in free_rcu_sheaf matches the throughput gain
almost 1:1. A ~21% increase in throughput naturally translates to ~21% more object
allocations and frees.

> 
> > free_rcu_sheaf_fail                      0                 0                 0      +0.00%
> > free_remove_partial             55,678,459        26,685,274       -28,993,185     -52.07%
> > free_slab                       55,678,459        26,701,099       -28,977,360     -52.04%
> > free_slowpath                  107,771,739        63,197,344       -44,574,395     -41.36%
> > min_partial                              5                 5                 0      +0.00%
> > object_size                            256               256                 0      +0.00%
> > objects                             14,504            14,336              -168      -1.16%
> > objects_partial                     14,504            14,208              -296      -2.04%
> > objs_per_slab                           64                64                 0      +0.00%
> > park_slab                                -     2,147,963,530                 -      absent
> > partial                              1,398             1,367               -31      -2.22%
> > sheaf_alloc                    743,660,288     1,284,618,991       540,958,703     +72.74%
> > sheaf_capacity                          32                32                 0      +0.00%
> > sheaf_flush                 43,855,651,858    53,384,274,983     9,528,623,125     +21.73%
> > sheaf_free                     743,660,280     1,284,618,967       540,958,687     +72.74%
> > sheaf_prefill_fast          17,585,335,850    21,378,958,833     3,793,622,983     +21.57%
> > sheaf_prefill_oversize                   0                 0                 0      +0.00%
> > sheaf_prefill_slow                   2,060             2,051                -9      -0.44%
> > sheaf_refill                43,963,424,505    53,447,473,167     9,484,048,662     +21.57%
> > sheaf_return_fast           17,585,336,335    21,378,959,334     3,793,622,999     +21.57%
> > sheaf_return_slow                    1,402             1,277              -125      -8.92%
> > slabs                                1,398             1,369               -29      -2.07%
> > total_objects                       89,472            87,616            -1,856      -2.07%
> > unpark_event                             -       315,397,838                 -      absent
> > unpark_slab                              -     2,147,963,530                 -      absent
> > 
> > derived                                                  before             after      change
> > =============================================================================================
> > PARK_SLAB / FREE_ADD_PARTIAL                                  -            99.35%      absent
> > UNPARK_SLAB / UNPARK_EVENT                                    -              6.81      absent
> > PARK_SLAB - UNPARK_SLAB                                       -                 0      absent
> > page allocator churn (alloc_slab + free_slab)       111,358,105        53,403,609     -52.04%
> > 
> > Note that some metrics have very small absolute values (such as
> > alloc_fastpath, partial, slabs, and total_objects) and are subject to
> > noise. Across multiple test runs, their rate of change fluctuates
> > between positive and negative, which supports the hypothesis that this
> > is measurement noise and demonstrates that this patch has no noticeable
> > impact on these metrics.
> > 
> > For metrics with large absolute values, their trends are distinct. The
> > data indicates that the primary benefit of this approach is
> > significantly relieved pressure on the buddy system, with page allocator
> > churn reduced by 52%. Additionally, free_slowpath decreases by 41%, and
> > cmpxchg_double_fail decreases by 85%.
> > 
> > The PARK_SLAB / FREE_ADD_PARTIAL ratio reaches 99.35%, which indicates
> > that the vast majority of partial slabs are added back to the partial
> > list via the parking mechanism, reflecting that the lock stayed
> > saturated and nearly all additions avoided waiting for the lock. The
> > ratio of UNPARK_SLAB / UNPARK_EVENT shows that each unpark event
> > processes roughly 6 slabs. PARK_SLAB - UNPARK_SLAB being 0 confirms
> > that no parked slabs are left stranded.
> > 
> > I also observe increases in both sheaf_alloc and sheaf_free, which
> > could currently be attributed to faster allocation and free paths
> > resulting from the overall performance improvement. However, I'm not
> > sure about this, which is part of why this is posted as an RFC.
> > 
> > Based on slab/for-next.
> > 
> > Hao Li (2):
> >   mm/slub: make the case handling in __slab_free() easier to follow
> >   mm/slub: introduce slab parking to reduce list_lock contention
> > 
> >  mm/slub.c | 326 +++++++++++++++++++++++++++++++++++++++++++++---------
> >  1 file changed, 274 insertions(+), 52 deletions(-)
> > 
> > base-commit: e7f630142df2afccce90555e4972e60008222311
> 

-- 
Thanks,
Hao
Re: [RFC PATCH 0/2] mm/slub: reduce list_lock contention with slab parking
Posted by Vlastimil Babka (SUSE) 1 week, 6 days ago
On 9/11/26 13:24, Hao Li wrote:
> On Mon, Sep 07, 2026 at 03:44:08PM +0200, Vlastimil Babka (SUSE) wrote:
>> On 8/24/26 14:19, Hao Li wrote:
>> > 
>> > Detailed data
>> > -------------
>> > 
>> > metric                              before             after             delta      change
>> > ==========================================================================================
>> > alloc_fastpath                     155,417           168,534            13,117      +8.44%
>> > alloc_slab                      55,679,646        26,702,510       -28,977,136     -52.04%
>> 
>> It's interesting that this is reduced so much. Is it because parked slabs
>> cause more slabs to stay around for reuse, despite they are unparked
>> immediately when trying to allocate/refill? That seems odd?
> 
> Yeah, it does look counter-intuitive at first glance. My understanding is that
> the parking mechanism helps by bypassing the partial list spinlock, allowing
> partial slabs to be parked onto the llist locklessly at any point. Because of
> this, the refill path gets to see and reuse more partial slabs instead of
> having to allocate fresh ones so frequently.

Hmm that's an interesting effect, surprisingly large if true. Because the
slabs we'd be parking are transitioning from full, there would be most often
just a single free object to reuse?

>> 
>> > alloc_slowpath                           0                 0                 0      +0.00%
>> > barn_get                             2,715             2,771                56      +2.06%
>> > barn_get_fail                            2                 0                -2    -100.00%
>> > barn_put                             2,715             2,770                55      +2.03%
>> > barn_put_fail                1,370,488,457     1,668,257,974       297,769,517     +21.73%
>> > cmpxchg_double_fail              3,827,935           549,427        -3,278,508     -85.65%
>> > free_add_partial             1,684,340,964     2,161,921,625       477,580,661     +28.35%
>> > free_fastpath                       31,766            32,979             1,213      +3.82%
>> > free_rcu_sheaf              43,855,689,417    53,384,314,553     9,528,625,136     +21.73%
>> 
>> This metric (and others with similar numbers) should not be affected by the
>> change. Does it mean the benchmark has a fixed time to run, but manages to
>> do more work in that time thanks to the increased throughput?
> 
> Exactly! In fact, the increase in free_rcu_sheaf matches the throughput gain
> almost 1:1. A ~21% increase in throughput naturally translates to ~21% more object
> allocations and frees.

Ack!
Re: [RFC PATCH 0/2] mm/slub: reduce list_lock contention with slab parking
Posted by Hao Li 1 week, 4 days ago
On Tue, Sep 15, 2026 at 09:10:22AM +0200, Vlastimil Babka (SUSE) wrote:
> On 9/11/26 13:24, Hao Li wrote:
> > On Mon, Sep 07, 2026 at 03:44:08PM +0200, Vlastimil Babka (SUSE) wrote:
> >> On 8/24/26 14:19, Hao Li wrote:
> >> > 
> >> > Detailed data
> >> > -------------
> >> > 
> >> > metric                              before             after             delta      change
> >> > ==========================================================================================
> >> > alloc_fastpath                     155,417           168,534            13,117      +8.44%
> >> > alloc_slab                      55,679,646        26,702,510       -28,977,136     -52.04%
> >> 
> >> It's interesting that this is reduced so much. Is it because parked slabs
> >> cause more slabs to stay around for reuse, despite they are unparked
> >> immediately when trying to allocate/refill? That seems odd?
> > 
> > Yeah, it does look counter-intuitive at first glance. My understanding is that
> > the parking mechanism helps by bypassing the partial list spinlock, allowing
> > partial slabs to be parked onto the llist locklessly at any point. Because of
> > this, the refill path gets to see and reuse more partial slabs instead of
> > having to allocate fresh ones so frequently.
> 
> Hmm that's an interesting effect, surprisingly large if true. Because the
> slabs we'd be parking are transitioning from full, there would be most often
> just a single free object to reuse?

That makes sense. That said, it might be more than a single object, since
build_detached_freelist() tries to free objects from the same slab together.

Even though we "park" the parking series for now, I was still curious to
understand the underlying behavior. So I add a counter in
get_partial_node_bulk() to see how many partial slabs it can see.
(counting freed objects directly is tough, so the partial slab count gives us a
rough picture)

Interestingly, each get_partial_node_bulk() on the patched kernel saw about
3-4x more partial slabs on average compared to baseline. This suggest that the
parking mechanism can supply more allocatable objects even under lock
contention, since it can avoid waiting for list lock.

For example, in the baseline kernel, if the lock is held by an allocator, a
freer has to wait on it. Once it's released, the lock might just get grabbed by
yet another allocator, leaving the freer stuck waiting. As a result, the
partial list ends up starved of free objects.


-- 
Thanks,
Hao
Re: [RFC PATCH 0/2] mm/slub: reduce list_lock contention with slab parking
Posted by Pedro Falcato 1 month ago
On Mon, Aug 24, 2026 at 08:19:50PM +0800, Hao Li wrote:
> This patch series might sound a bit wild, but the initial numbers don't
> look too bad so far. I would really appreciate any feedback and
> discussion :)

The numbers look great, I have to say ;)

> 
> On a will-it-scale mmap1 run with 192 processes, list_lock is the top
> contention point: __slab_free() and __refill_objects_node() together
> spend 44% of cycles in native_queued_spin_lock_slowpath. The free
> slowpath takes the lock mainly to add slabs that became non-full to the
> partial list.

Doesn't this mean you're hitting the slow path way too many times? I think
that's the actual issue, no?

> 
> By adding extra instrumentation to __slab_free(), I collect the following
> data for the maple_node cache (in counts):
> 
>   partial->partial    843017414
>   full->partial       550719384
>   partial->empty      17459564
>   full->empty         2
> 
> We can see that full -> partial transitions account for a significant
> proportion, and optimizing them can help reduce lock contention to some
> extent.
> 
> This series introduces the parking mechanism to address this issue.
> When the trylock fails during a full -> partial/empty transition,
> __slab_free() parks the slab on a per-node llist instead of waiting. The
> paths that consume the partial list (sheaf refill, alloc slowpath,
> shrink, cache destruction) unpark the slabs after taking the lock, and a
> delayed work covers the case where none of them runs.
> 
> Patch 1 cleans up the case handling in __slab_free(), no functional
> change. Patch 2 introduces the parking mechanism.

This looks like fundamentally the wrong fix, when we want to find out why
1) you're hitting the alloc slowpath so hard
2) you're hitting the free slowpath so hard

and ideally find some way to tune it properly. Maybe if sheaf size is scaled
up/down in some way (by default at least), according to amount of RAM or amount of
CPUs.

-- 
Pedro
Re: [RFC PATCH 0/2] mm/slub: reduce list_lock contention with slab parking
Posted by Hao Li 4 weeks ago
On Thu, Aug 27, 2026 at 05:24:32PM +0100, Pedro Falcato wrote:
> On Mon, Aug 24, 2026 at 08:19:50PM +0800, Hao Li wrote:
> > This patch series might sound a bit wild, but the initial numbers don't
> > look too bad so far. I would really appreciate any feedback and
> > discussion :)
> 
> The numbers look great, I have to say ;)

Thanks and sorry for the delayed reply!

> 
> > 
> > On a will-it-scale mmap1 run with 192 processes, list_lock is the top
> > contention point: __slab_free() and __refill_objects_node() together
> > spend 44% of cycles in native_queued_spin_lock_slowpath. The free
> > slowpath takes the lock mainly to add slabs that became non-full to the
> > partial list.
> 
> Doesn't this mean you're hitting the slow path way too many times? I think
> that's the actual issue, no?

Partly, let me separate the two sides.

> 
> > 
> > By adding extra instrumentation to __slab_free(), I collect the following
> > data for the maple_node cache (in counts):
> > 
> >   partial->partial    843017414
> >   full->partial       550719384
> >   partial->empty      17459564
> >   full->empty         2
> > 
> > We can see that full -> partial transitions account for a significant
> > proportion, and optimizing them can help reduce lock contention to some
> > extent.
> > 
> > This series introduces the parking mechanism to address this issue.
> > When the trylock fails during a full -> partial/empty transition,
> > __slab_free() parks the slab on a per-node llist instead of waiting. The
> > paths that consume the partial list (sheaf refill, alloc slowpath,
> > shrink, cache destruction) unpark the slabs after taking the lock, and a
> > delayed work covers the case where none of them runs.
> > 
> > Patch 1 cleans up the case handling in __slab_free(), no functional
> > change. Patch 2 introduces the parking mechanism.
> 
> This looks like fundamentally the wrong fix, when we want to find out why
> 1) you're hitting the alloc slowpath so hard

Regarding the alloc side, it mostly goes through sheaf refilling, which isn't
quite the actual slowpath in the usual sense. The full slowpath would be
allocating a single object directly from the partial list, and from what we've
observed in testing, that path is rarely ever hit.

> 2) you're hitting the free slowpath so hard

Yes, the free side does indeed take the slowpath, which is precisely the
problem this series aims to address.

> 
> and ideally find some way to tune it properly. Maybe if sheaf size is scaled
> up/down in some way (by default at least), according to amount of RAM or amount of
> CPUs.

I think the idea makes sense, though it might only offer partial relief. Since
much of the overhead seems to come from free-side lock contention, a lockless
approach could be a more fundamental fix.

Thanks for the discussion!

-- 
Thanks,
Hao