[PATCH v2 00/12] Convert barrier pairs to acquire/release for better performance

Jinjie Ruan posted 12 patches 3 weeks, 4 days ago
There is a newer version of this series
fs/aio.c                  | 10 ++++------
fs/ext4/balloc.c          |  2 +-
fs/ext4/ext4.h            | 10 +++-------
fs/ext4/mballoc.c         |  6 ++----
fs/ext4/resize.c          | 19 +++++++++++--------
fs/file.c                 | 10 ++++------
fs/mnt_idmapping.c        |  5 ++---
fs/pidfs.c                |  6 ++----
fs/super.c                |  6 ++----
kernel/user_namespace.c   | 24 +++++++++++++-----------
lib/vsprintf.c            | 11 ++++-------
net/8021q/vlan.c          |  6 ++----
net/8021q/vlan.h          | 12 +++++-------
net/core/sock_reuseport.c | 20 ++++++++------------
net/sched/act_gact.c      | 12 ++++--------
15 files changed, 67 insertions(+), 92 deletions(-)
[PATCH v2 00/12] Convert barrier pairs to acquire/release for better performance
Posted by Jinjie Ruan 3 weeks, 4 days ago
Hi,

This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
smp_store_release()/smp_load_acquire() across various subsystems.

Background
==========

Many architectures support load acquire and store release instructions
which can replace explicit memory barriers and save cycles. As noted
in the ARM architecture reference [1]:

  "Weaker ordering requirements that are imposed by Load-Acquire and
   Store-Release instructions allow for micro-architectural
   optimizations, which could reduce some of the performance impacts
   that are otherwise imposed by an explicit memory barrier.

   If the ordering requirement is satisfied using either a Load-Acquire
   or Store-Release, then it would be preferable to use these
   instructions instead of a DMB."

On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
barriers. Replacing the read barrier with smp_load_acquire() reduces
this to 8 cycles on an Ampere Altra.

We also observed significant barrier overhead while profiling Unxibench
syscall test on arm64: a single getuid() call is ~8ns slower than on
a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
the use of LDAR, eliminating the measurable overhead.

This motivated a broader search for existing barrier pairs that can
be converted to the lighter acquire/release semantics.

Changes
=======

Each patch in this series targets a specific barrier pair where the
publish/subscribe pattern is already present:

- Writers populate data, then publish a flag/count/pointer via
  smp_store_release()

- Readers load the flag/count/pointer via smp_load_acquire(), then
  consume the data

This preserves the existing memory ordering guarantees while allowing
architectures with native acquire/release instructions (e.g. arm64's
STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
On architectures without native support, the generated code is
generally no worse than the explicit barrier pair.

The conversions are mechanical and no functional change is intended.

Testing (arm64 Kunpeng HIP09 server)
================

1. UNIXBENCH syscall
	Baseline: 715.27
	Patched:  718.83
	Improvement: +0.50%

2. fs/aio (fio + null_blk, 4 jobs):
	Baseline: 1441k IOPS, 86.46us
	Patched:  1452k IOPS, 85.80us
	Improvement: ~0.8%

3. soreuseport (wrk, 8 servers):
	Baseline: 162.6k req/s, 452.5us
	Patched:  164.2k req/s, 449.4us
	Improvement: ~1.0%

Both improvements are consistent across runs and align with the
expected savings from replacing DMB with LDAR/STLR on arm64.

[1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
[2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba

Changes in v2:
- Fix pre-existing issue for ext4 and 8021q [3].
- Fix missing copy_mnt_idmap() udapte [3].
- Drop nacked isotp patch.
- Add test data.
- Add Reviewed-by and update fs patch as Jan suggested.

[3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com

Jinjie Ruan (12):
  user_namespace: Use acquire/release for nr_extents synchronization
  lib/vsprintf: Use acquire/release for ptr_key publication
  fs: aio: Use acquire/release for ring->tail publication
  fs: Use acquire/release for fdtable resize synchronization
  pidfs: Use test_bit_acquire() for attr flag tests
  super: Use acquire for SB_BORN check in super_cache_count()
  ext4: Fix out-of-bounds read in ext4_get_group_info()
  ext4: Convert group-count barrier protocol to acquire/release
  soreuseport: publish num_socks with acquire/release
  net: sched: act_gact: use acquire/release for tcfg_ptype
  8021q: Fix data race when publishing vlan net_device pointers
  8021q: publish vlan_devices_arrays entries with acquire/release

 fs/aio.c                  | 10 ++++------
 fs/ext4/balloc.c          |  2 +-
 fs/ext4/ext4.h            | 10 +++-------
 fs/ext4/mballoc.c         |  6 ++----
 fs/ext4/resize.c          | 19 +++++++++++--------
 fs/file.c                 | 10 ++++------
 fs/mnt_idmapping.c        |  5 ++---
 fs/pidfs.c                |  6 ++----
 fs/super.c                |  6 ++----
 kernel/user_namespace.c   | 24 +++++++++++++-----------
 lib/vsprintf.c            | 11 ++++-------
 net/8021q/vlan.c          |  6 ++----
 net/8021q/vlan.h          | 12 +++++-------
 net/core/sock_reuseport.c | 20 ++++++++------------
 net/sched/act_gact.c      | 12 ++++--------
 15 files changed, 67 insertions(+), 92 deletions(-)

-- 
2.34.1
Re: [PATCH v2 00/12] Convert barrier pairs to acquire/release for better performance
Posted by Kuniyuki Iwashima 3 weeks, 4 days ago
On Mon, Aug 31, 2026 at 7:42 PM Jinjie Ruan <ruanjinjie@huawei.com> wrote:
>
> Hi,
>
> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
> smp_store_release()/smp_load_acquire() across various subsystems.
>
> Background
> ==========
>
> Many architectures support load acquire and store release instructions
> which can replace explicit memory barriers and save cycles. As noted
> in the ARM architecture reference [1]:
>
>   "Weaker ordering requirements that are imposed by Load-Acquire and
>    Store-Release instructions allow for micro-architectural
>    optimizations, which could reduce some of the performance impacts
>    that are otherwise imposed by an explicit memory barrier.
>
>    If the ordering requirement is satisfied using either a Load-Acquire
>    or Store-Release, then it would be preferable to use these
>    instructions instead of a DMB."
>
> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
> barriers. Replacing the read barrier with smp_load_acquire() reduces
> this to 8 cycles on an Ampere Altra.
>
> We also observed significant barrier overhead while profiling Unxibench
> syscall test on arm64: a single getuid() call is ~8ns slower than on
> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
> the use of LDAR, eliminating the measurable overhead.
>
> This motivated a broader search for existing barrier pairs that can
> be converted to the lighter acquire/release semantics.
>
> Changes
> =======
>
> Each patch in this series targets a specific barrier pair where the
> publish/subscribe pattern is already present:
>
> - Writers populate data, then publish a flag/count/pointer via
>   smp_store_release()
>
> - Readers load the flag/count/pointer via smp_load_acquire(), then
>   consume the data
>
> This preserves the existing memory ordering guarantees while allowing
> architectures with native acquire/release instructions (e.g. arm64's
> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
> On architectures without native support, the generated code is
> generally no worse than the explicit barrier pair.
>
> The conversions are mechanical and no functional change is intended.
>
> Testing (arm64 Kunpeng HIP09 server)
> ================
>
> 1. UNIXBENCH syscall
>         Baseline: 715.27
>         Patched:  718.83
>         Improvement: +0.50%
>
> 2. fs/aio (fio + null_blk, 4 jobs):
>         Baseline: 1441k IOPS, 86.46us
>         Patched:  1452k IOPS, 85.80us
>         Improvement: ~0.8%
>
> 3. soreuseport (wrk, 8 servers):
>         Baseline: 162.6k req/s, 452.5us
>         Patched:  164.2k req/s, 449.4us
>         Improvement: ~1.0%
>
> Both improvements are consistent across runs and align with the
> expected savings from replacing DMB with LDAR/STLR on arm64.
>
> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
>
> Changes in v2:
> - Fix pre-existing issue for ext4 and 8021q [3].
> - Fix missing copy_mnt_idmap() udapte [3].
> - Drop nacked isotp patch.
> - Add test data.
> - Add Reviewed-by and update fs patch as Jan suggested.
>
> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
>
> Jinjie Ruan (12):
>   user_namespace: Use acquire/release for nr_extents synchronization
>   lib/vsprintf: Use acquire/release for ptr_key publication
>   fs: aio: Use acquire/release for ring->tail publication
>   fs: Use acquire/release for fdtable resize synchronization
>   pidfs: Use test_bit_acquire() for attr flag tests
>   super: Use acquire for SB_BORN check in super_cache_count()
>   ext4: Fix out-of-bounds read in ext4_get_group_info()
>   ext4: Convert group-count barrier protocol to acquire/release
>   soreuseport: publish num_socks with acquire/release
>   net: sched: act_gact: use acquire/release for tcfg_ptype
>   8021q: Fix data race when publishing vlan net_device pointers
>   8021q: publish vlan_devices_arrays entries with acquire/release

Please post networking patches separately with the target tree specified:

  Subject: [PATCH vX net-next] soreuseport: ...

8021q changes can be posted a series.
Re: [PATCH v2 00/12] Convert barrier pairs to acquire/release for better performance
Posted by Jinjie Ruan 3 weeks, 4 days ago

在 2026/9/1 11:06, Kuniyuki Iwashima 写道:
> On Mon, Aug 31, 2026 at 7:42 PM Jinjie Ruan <ruanjinjie@huawei.com> wrote:
>>
>> Hi,
>>
>> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
>> smp_store_release()/smp_load_acquire() across various subsystems.
>>
>> Background
>> ==========
>>
>> Many architectures support load acquire and store release instructions
>> which can replace explicit memory barriers and save cycles. As noted
>> in the ARM architecture reference [1]:
>>
>>   "Weaker ordering requirements that are imposed by Load-Acquire and
>>    Store-Release instructions allow for micro-architectural
>>    optimizations, which could reduce some of the performance impacts
>>    that are otherwise imposed by an explicit memory barrier.
>>
>>    If the ordering requirement is satisfied using either a Load-Acquire
>>    or Store-Release, then it would be preferable to use these
>>    instructions instead of a DMB."
>>
>> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
>> barriers. Replacing the read barrier with smp_load_acquire() reduces
>> this to 8 cycles on an Ampere Altra.
>>
>> We also observed significant barrier overhead while profiling Unxibench
>> syscall test on arm64: a single getuid() call is ~8ns slower than on
>> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
>> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
>> the use of LDAR, eliminating the measurable overhead.
>>
>> This motivated a broader search for existing barrier pairs that can
>> be converted to the lighter acquire/release semantics.
>>
>> Changes
>> =======
>>
>> Each patch in this series targets a specific barrier pair where the
>> publish/subscribe pattern is already present:
>>
>> - Writers populate data, then publish a flag/count/pointer via
>>   smp_store_release()
>>
>> - Readers load the flag/count/pointer via smp_load_acquire(), then
>>   consume the data
>>
>> This preserves the existing memory ordering guarantees while allowing
>> architectures with native acquire/release instructions (e.g. arm64's
>> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
>> On architectures without native support, the generated code is
>> generally no worse than the explicit barrier pair.
>>
>> The conversions are mechanical and no functional change is intended.
>>
>> Testing (arm64 Kunpeng HIP09 server)
>> ================
>>
>> 1. UNIXBENCH syscall
>>         Baseline: 715.27
>>         Patched:  718.83
>>         Improvement: +0.50%
>>
>> 2. fs/aio (fio + null_blk, 4 jobs):
>>         Baseline: 1441k IOPS, 86.46us
>>         Patched:  1452k IOPS, 85.80us
>>         Improvement: ~0.8%
>>
>> 3. soreuseport (wrk, 8 servers):
>>         Baseline: 162.6k req/s, 452.5us
>>         Patched:  164.2k req/s, 449.4us
>>         Improvement: ~1.0%
>>
>> Both improvements are consistent across runs and align with the
>> expected savings from replacing DMB with LDAR/STLR on arm64.
>>
>> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
>> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
>>
>> Changes in v2:
>> - Fix pre-existing issue for ext4 and 8021q [3].
>> - Fix missing copy_mnt_idmap() udapte [3].
>> - Drop nacked isotp patch.
>> - Add test data.
>> - Add Reviewed-by and update fs patch as Jan suggested.
>>
>> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
>>
>> Jinjie Ruan (12):
>>   user_namespace: Use acquire/release for nr_extents synchronization
>>   lib/vsprintf: Use acquire/release for ptr_key publication
>>   fs: aio: Use acquire/release for ring->tail publication
>>   fs: Use acquire/release for fdtable resize synchronization
>>   pidfs: Use test_bit_acquire() for attr flag tests
>>   super: Use acquire for SB_BORN check in super_cache_count()
>>   ext4: Fix out-of-bounds read in ext4_get_group_info()
>>   ext4: Convert group-count barrier protocol to acquire/release
>>   soreuseport: publish num_socks with acquire/release
>>   net: sched: act_gact: use acquire/release for tcfg_ptype
>>   8021q: Fix data race when publishing vlan net_device pointers
>>   8021q: publish vlan_devices_arrays entries with acquire/release
> 
> Please post networking patches separately with the target tree specified:
> 
>   Subject: [PATCH vX net-next] soreuseport: ...
> 
> 8021q changes can be posted a series.

Thanks for the review. I will split the series as suggested — the
networking patches will be posted separately.