[PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance

Jinjie Ruan posted 8 patches 3 weeks, 3 days ago
fs/aio.c                | 10 ++++------
fs/ext4/balloc.c        |  2 +-
fs/ext4/ext4.h          | 10 +++-------
fs/ext4/mballoc.c       |  6 ++----
fs/ext4/resize.c        | 19 +++++++++++--------
fs/file.c               | 10 ++++------
fs/mnt_idmapping.c      |  5 ++---
fs/pidfs.c              |  6 ++----
fs/super.c              |  6 ++----
kernel/user_namespace.c | 24 +++++++++++++-----------
lib/vsprintf.c          | 11 ++++-------
11 files changed, 48 insertions(+), 61 deletions(-)
[PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Posted by Jinjie Ruan 3 weeks, 3 days ago
Hi,

This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
smp_store_release()/smp_load_acquire() across various subsystems.

Background
==========

Many architectures support load acquire and store release instructions
which can replace explicit memory barriers and save cycles. As noted
in the ARM architecture reference [1]:

  "Weaker ordering requirements that are imposed by Load-Acquire and
   Store-Release instructions allow for micro-architectural
   optimizations, which could reduce some of the performance impacts
   that are otherwise imposed by an explicit memory barrier.

   If the ordering requirement is satisfied using either a Load-Acquire
   or Store-Release, then it would be preferable to use these
   instructions instead of a DMB."

On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
barriers. Replacing the read barrier with smp_load_acquire() reduces
this to 8 cycles on an Ampere Altra.

We also observed significant barrier overhead while profiling Unxibench
syscall test on arm64: a single getuid() call is ~8ns slower than on
a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
the use of LDAR, eliminating the measurable overhead.

This motivated a broader search for existing barrier pairs that can
be converted to the lighter acquire/release semantics.

Changes
=======

Each patch in this series targets a specific barrier pair where the
publish/subscribe pattern is already present:

- Writers populate data, then publish a flag/count/pointer via
  smp_store_release()

- Readers load the flag/count/pointer via smp_load_acquire(), then
  consume the data

This preserves the existing memory ordering guarantees while allowing
architectures with native acquire/release instructions (e.g. arm64's
STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
On architectures without native support, the generated code is
generally no worse than the explicit barrier pair.

The conversions are mechanical and no functional change is intended.

Testing (Kunpeng HIP09 arm64 server)
====================================

1. UNIXBENCH syscall
	Baseline: 715.27
	Patched:  718.83
	Improvement: +0.50%

2. fs/aio (fio + null_blk, 4 jobs):
	Baseline: 1441k IOPS, 86.46us
	Patched:  1452k IOPS, 85.80us
	Improvement: ~0.8%

Both improvements are consistent across runs and align with the
expected savings from replacing DMB with LDAR/STLR on arm64.

[1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
[2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba

Changes in v3:
- Add Reviewed-by.
- Split out network patch set as Kuniyuki suggested.
- Link to v2: https://lore.kernel.org/all/20260901024234.135119-1-ruanjinjie@huawei.com/

Changes in v2:
- Fix pre-existing issue for ext4 and 8021q [3].
- Fix missing copy_mnt_idmap() udapte [3].
- Drop nacked isotp patch.
- Add test data.
- Add Reviewed-by and update fs patch as Jan suggested.

[3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com

Jinjie Ruan (8):
  user_namespace: Use acquire/release for nr_extents synchronization
  lib/vsprintf: Use acquire/release for ptr_key publication
  fs: aio: Use acquire/release for ring->tail publication
  fs: Use acquire/release for fdtable resize synchronization
  pidfs: Use test_bit_acquire() for attr flag tests
  super: Use acquire for SB_BORN check in super_cache_count()
  ext4: Fix out-of-bounds read in ext4_get_group_info()
  ext4: Convert group-count barrier protocol to acquire/release

 fs/aio.c                | 10 ++++------
 fs/ext4/balloc.c        |  2 +-
 fs/ext4/ext4.h          | 10 +++-------
 fs/ext4/mballoc.c       |  6 ++----
 fs/ext4/resize.c        | 19 +++++++++++--------
 fs/file.c               | 10 ++++------
 fs/mnt_idmapping.c      |  5 ++---
 fs/pidfs.c              |  6 ++----
 fs/super.c              |  6 ++----
 kernel/user_namespace.c | 24 +++++++++++++-----------
 lib/vsprintf.c          | 11 ++++-------
 11 files changed, 48 insertions(+), 61 deletions(-)

-- 
2.34.1
Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Posted by Theodore Tso 3 weeks, 2 days ago
On Wed, Sep 02, 2026 at 03:47:57PM -0500, Jinjie Ruan wrote:
> 
> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
> smp_store_release()/smp_load_acquire() across various subsystems.

I would like to land the ext4 related patches through the ext4 tree,
so please detach them from this patch series.  That's because I want
to be able to monitor the regression testing before sending a pull
request to Linus.

More generally, I think we should be looking at this kind of barrier
conversion on a per site basis.  If there is a bug that isn't notice,
it can result in a very hard to debug problem.  And if the benefit of
speeding up a particular code path isn't particularly that great,
maybe the cost vs benefit vs risk calculation might cause us to decide
to take a pass on the change.

Cheers,

					- Ted
Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Posted by Jinjie Ruan 3 weeks, 1 day ago

在 2026/9/2 22:19, Theodore Tso 写道:
> On Wed, Sep 02, 2026 at 03:47:57PM -0500, Jinjie Ruan wrote:
>>
>> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
>> smp_store_release()/smp_load_acquire() across various subsystems.
> 
> I would like to land the ext4 related patches through the ext4 tree,
> so please detach them from this patch series.  That's because I want
> to be able to monitor the regression testing before sending a pull
> request to Linus.

Hi Ted,

Thanks for your feedback. I will split out the ext4-related patches.

Thanks,
Jinjie

> 
> More generally, I think we should be looking at this kind of barrier
> conversion on a per site basis.  If there is a bug that isn't notice,
> it can result in a very hard to debug problem.  And if the benefit of
> speeding up a particular code path isn't particularly that great,
> maybe the cost vs benefit vs risk calculation might cause us to decide
> to take a pass on the change.
> 
> Cheers,
> 
> 					- Ted

Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Posted by Jan Kara 3 weeks, 2 days ago
On Wed 02-09-26 10:19:46, Theodore Tso wrote:
> On Wed, Sep 02, 2026 at 03:47:57PM -0500, Jinjie Ruan wrote:
> > 
> > This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
> > smp_store_release()/smp_load_acquire() across various subsystems.
> 
> I would like to land the ext4 related patches through the ext4 tree,
> so please detach them from this patch series.  That's because I want
> to be able to monitor the regression testing before sending a pull
> request to Linus.
> 
> More generally, I think we should be looking at this kind of barrier
> conversion on a per site basis.  If there is a bug that isn't notice,
> it can result in a very hard to debug problem.  And if the benefit of
> speeding up a particular code path isn't particularly that great,
> maybe the cost vs benefit vs risk calculation might cause us to decide
> to take a pass on the change.

FWIW I believe each patch in this series is standalone so you can just pick
the ext4 ones yourself. Regarding the benefit, I actually find load_acquire
/ store_release more self-documenting, what exactly they are protecting, so
generally I'm in favor of such conversions if they logically make sense
(i.e., it's really the "prepare some data + store to publish" pattern) -
which was the case in the patches I've reviewed in this series.

								Honza
-- 
Jan Kara <jack@suse.com>
SUSE Labs, CR
Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Posted by Petr Mladek 3 weeks ago
Adding Risc-V list into Cc.

On Wed 2026-09-02 15:47:57, Jinjie Ruan wrote:
> Hi,
> 
> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
> smp_store_release()/smp_load_acquire() across various subsystems.
> 
> Background
> ==========
> 
> Many architectures support load acquire and store release instructions
> which can replace explicit memory barriers and save cycles. As noted
> in the ARM architecture reference [1]:
> 
>   "Weaker ordering requirements that are imposed by Load-Acquire and
>    Store-Release instructions allow for micro-architectural
>    optimizations, which could reduce some of the performance impacts
>    that are otherwise imposed by an explicit memory barrier.
> 
>    If the ordering requirement is satisfied using either a Load-Acquire
>    or Store-Release, then it would be preferable to use these
>    instructions instead of a DMB."
> 
> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
> barriers. Replacing the read barrier with smp_load_acquire() reduces
> this to 8 cycles on an Ampere Altra.

I wonder if this is true on all other architectures:

  + It seems that Arm gets the gain because the instruction
    does both load/store + barrier. It helps even when
    the barrier is full.

  + Some other architectures need two instructions. One for the
    load/store and the other for the barrier. But the barrier
    is weaker, it synchronizes just reads or just writes.

For example, I see the following in riscv/include/asm/barrier.h:

<paste riscv/include/asm/barrier.h>
#define smp_mb()	RISCV_FENCE(rw, rw)
#define smp_rmb()	RISCV_FENCE(r, r)
#define smp_wmb()	RISCV_FENCE(w, w)

#define smp_store_release(p, v)						\
do {									\
	RISCV_FENCE(rw, w);						\
	WRITE_ONCE(*p, v);						\
} while (0)

#define smp_load_acquire(p)						\
({									\
	typeof(*p) ___p1 = READ_ONCE(*p);				\
	RISCV_FENCE(r, rw);						\
	___p1;								\
})
</paste riscv/include/asm/barrier.h>

I wonder whether:

  + RISCV_FENCE(r, r) is faster than RISCV_FENCE(r, rw)
  + RISCV_FENCE(w, w) is faster than RISCV_FENCE(rw, w)

so it might cause performance regression there...

Best Regards,
Petr

> We also observed significant barrier overhead while profiling Unxibench
> syscall test on arm64: a single getuid() call is ~8ns slower than on
> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
> the use of LDAR, eliminating the measurable overhead.
> 
> This motivated a broader search for existing barrier pairs that can
> be converted to the lighter acquire/release semantics.
> 
> Changes
> =======
> 
> Each patch in this series targets a specific barrier pair where the
> publish/subscribe pattern is already present:
> 
> - Writers populate data, then publish a flag/count/pointer via
>   smp_store_release()
> 
> - Readers load the flag/count/pointer via smp_load_acquire(), then
>   consume the data
> 
> This preserves the existing memory ordering guarantees while allowing
> architectures with native acquire/release instructions (e.g. arm64's
> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
> On architectures without native support, the generated code is
> generally no worse than the explicit barrier pair.
> 
> The conversions are mechanical and no functional change is intended.
> 
> Testing (Kunpeng HIP09 arm64 server)
> ====================================
> 
> 1. UNIXBENCH syscall
> 	Baseline: 715.27
> 	Patched:  718.83
> 	Improvement: +0.50%
> 
> 2. fs/aio (fio + null_blk, 4 jobs):
> 	Baseline: 1441k IOPS, 86.46us
> 	Patched:  1452k IOPS, 85.80us
> 	Improvement: ~0.8%
> 
> Both improvements are consistent across runs and align with the
> expected savings from replacing DMB with LDAR/STLR on arm64.
> 
> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
> 
> Changes in v3:
> - Add Reviewed-by.
> - Split out network patch set as Kuniyuki suggested.
> - Link to v2: https://lore.kernel.org/all/20260901024234.135119-1-ruanjinjie@huawei.com/
> 
> Changes in v2:
> - Fix pre-existing issue for ext4 and 8021q [3].
> - Fix missing copy_mnt_idmap() udapte [3].
> - Drop nacked isotp patch.
> - Add test data.
> - Add Reviewed-by and update fs patch as Jan suggested.
> 
> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
> 
> Jinjie Ruan (8):
>   user_namespace: Use acquire/release for nr_extents synchronization
>   lib/vsprintf: Use acquire/release for ptr_key publication
>   fs: aio: Use acquire/release for ring->tail publication
>   fs: Use acquire/release for fdtable resize synchronization
>   pidfs: Use test_bit_acquire() for attr flag tests
>   super: Use acquire for SB_BORN check in super_cache_count()
>   ext4: Fix out-of-bounds read in ext4_get_group_info()
>   ext4: Convert group-count barrier protocol to acquire/release
> 
>  fs/aio.c                | 10 ++++------
>  fs/ext4/balloc.c        |  2 +-
>  fs/ext4/ext4.h          | 10 +++-------
>  fs/ext4/mballoc.c       |  6 ++----
>  fs/ext4/resize.c        | 19 +++++++++++--------
>  fs/file.c               | 10 ++++------
>  fs/mnt_idmapping.c      |  5 ++---
>  fs/pidfs.c              |  6 ++----
>  fs/super.c              |  6 ++----
>  kernel/user_namespace.c | 24 +++++++++++++-----------
>  lib/vsprintf.c          | 11 ++++-------
>  11 files changed, 48 insertions(+), 61 deletions(-)
> 
> -- 
> 2.34.1
Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Posted by Jinjie Ruan 2 weeks, 4 days ago

在 2026/9/4 20:28, Petr Mladek 写道:
> Adding Risc-V list into Cc.
> 
> On Wed 2026-09-02 15:47:57, Jinjie Ruan wrote:
>> Hi,
>>
>> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
>> smp_store_release()/smp_load_acquire() across various subsystems.
>>
>> Background
>> ==========
>>
>> Many architectures support load acquire and store release instructions
>> which can replace explicit memory barriers and save cycles. As noted
>> in the ARM architecture reference [1]:
>>
>>   "Weaker ordering requirements that are imposed by Load-Acquire and
>>    Store-Release instructions allow for micro-architectural
>>    optimizations, which could reduce some of the performance impacts
>>    that are otherwise imposed by an explicit memory barrier.
>>
>>    If the ordering requirement is satisfied using either a Load-Acquire
>>    or Store-Release, then it would be preferable to use these
>>    instructions instead of a DMB."
>>
>> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
>> barriers. Replacing the read barrier with smp_load_acquire() reduces
>> this to 8 cycles on an Ampere Altra.
> 
> I wonder if this is true on all other architectures:
> 
>   + It seems that Arm gets the gain because the instruction
>     does both load/store + barrier. It helps even when
>     the barrier is full.
> 
>   + Some other architectures need two instructions. One for the
>     load/store and the other for the barrier. But the barrier
>     is weaker, it synchronizes just reads or just writes.
> 
> For example, I see the following in riscv/include/asm/barrier.h:
> 
> <paste riscv/include/asm/barrier.h>
> #define smp_mb()	RISCV_FENCE(rw, rw)
> #define smp_rmb()	RISCV_FENCE(r, r)
> #define smp_wmb()	RISCV_FENCE(w, w)
> 
> #define smp_store_release(p, v)						\
> do {									\
> 	RISCV_FENCE(rw, w);						\
> 	WRITE_ONCE(*p, v);						\
> } while (0)
> 
> #define smp_load_acquire(p)						\
> ({									\
> 	typeof(*p) ___p1 = READ_ONCE(*p);				\
> 	RISCV_FENCE(r, rw);						\
> 	___p1;								\
> })
> </paste riscv/include/asm/barrier.h>
> 
> I wonder whether:
> 
>   + RISCV_FENCE(r, r) is faster than RISCV_FENCE(r, rw)
>   + RISCV_FENCE(w, w) is faster than RISCV_FENCE(rw, w)
> 
> so it might cause performance regression there...

Hi Petr,

Thanks for the detailed analysis. You're right that on some
architectures the acquire/release variants use a slightly heavier
fence than the plain smp_wmb()/smp_rmb() pair. The full picture by
architecture:

  arm64:    Improvement   (DMB ISHST/ISHLD + STR/LDR → STLR/LDAR)
  x86:      Neutral       (both are compiler barriers)
  s390:     Neutral       (both are compiler barriers)
  ppc64:    Neutral       (both use lwsync)
  loongarch: Slightly heavier fence (DBAR(o_w_w) → DBAR(orw_w),
                           DBAR(or_r_) → DBAR(or_rw))
  riscv:    Slightly heavier fence (fence w,w → fence rw,w,
                           fence r,r → fence r,rw)

So on Loongarch and Riscv, there may be a slight performance regression.

Regards,
Jinjie

> 
> Best Regards,
> Petr
> 
>> We also observed significant barrier overhead while profiling Unxibench
>> syscall test on arm64: a single getuid() call is ~8ns slower than on
>> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
>> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
>> the use of LDAR, eliminating the measurable overhead.
>>
>> This motivated a broader search for existing barrier pairs that can
>> be converted to the lighter acquire/release semantics.
>>
>> Changes
>> =======
>>
>> Each patch in this series targets a specific barrier pair where the
>> publish/subscribe pattern is already present:
>>
>> - Writers populate data, then publish a flag/count/pointer via
>>   smp_store_release()
>>
>> - Readers load the flag/count/pointer via smp_load_acquire(), then
>>   consume the data
>>
>> This preserves the existing memory ordering guarantees while allowing
>> architectures with native acquire/release instructions (e.g. arm64's
>> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
>> On architectures without native support, the generated code is
>> generally no worse than the explicit barrier pair.
>>
>> The conversions are mechanical and no functional change is intended.
>>
>> Testing (Kunpeng HIP09 arm64 server)
>> ====================================
>>
>> 1. UNIXBENCH syscall
>> 	Baseline: 715.27
>> 	Patched:  718.83
>> 	Improvement: +0.50%
>>
>> 2. fs/aio (fio + null_blk, 4 jobs):
>> 	Baseline: 1441k IOPS, 86.46us
>> 	Patched:  1452k IOPS, 85.80us
>> 	Improvement: ~0.8%
>>
>> Both improvements are consistent across runs and align with the
>> expected savings from replacing DMB with LDAR/STLR on arm64.
>>
>> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
>> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
>>
>> Changes in v3:
>> - Add Reviewed-by.
>> - Split out network patch set as Kuniyuki suggested.
>> - Link to v2: https://lore.kernel.org/all/20260901024234.135119-1-ruanjinjie@huawei.com/
>>
>> Changes in v2:
>> - Fix pre-existing issue for ext4 and 8021q [3].
>> - Fix missing copy_mnt_idmap() udapte [3].
>> - Drop nacked isotp patch.
>> - Add test data.
>> - Add Reviewed-by and update fs patch as Jan suggested.
>>
>> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
>>
>> Jinjie Ruan (8):
>>   user_namespace: Use acquire/release for nr_extents synchronization
>>   lib/vsprintf: Use acquire/release for ptr_key publication
>>   fs: aio: Use acquire/release for ring->tail publication
>>   fs: Use acquire/release for fdtable resize synchronization
>>   pidfs: Use test_bit_acquire() for attr flag tests
>>   super: Use acquire for SB_BORN check in super_cache_count()
>>   ext4: Fix out-of-bounds read in ext4_get_group_info()
>>   ext4: Convert group-count barrier protocol to acquire/release
>>
>>  fs/aio.c                | 10 ++++------
>>  fs/ext4/balloc.c        |  2 +-
>>  fs/ext4/ext4.h          | 10 +++-------
>>  fs/ext4/mballoc.c       |  6 ++----
>>  fs/ext4/resize.c        | 19 +++++++++++--------
>>  fs/file.c               | 10 ++++------
>>  fs/mnt_idmapping.c      |  5 ++---
>>  fs/pidfs.c              |  6 ++----
>>  fs/super.c              |  6 ++----
>>  kernel/user_namespace.c | 24 +++++++++++++-----------
>>  lib/vsprintf.c          | 11 ++++-------
>>  11 files changed, 48 insertions(+), 61 deletions(-)
>>
>> -- 
>> 2.34.1

Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Posted by David Laight 2 weeks, 4 days ago
On Mon, 7 Sep 2026 19:29:06 +0800
Jinjie Ruan <ruanjinjie@huawei.com> wrote:

> 在 2026/9/4 20:28, Petr Mladek 写道:
> > Adding Risc-V list into Cc.
> > 
> > On Wed 2026-09-02 15:47:57, Jinjie Ruan wrote:  
> >> Hi,
> >>
> >> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
> >> smp_store_release()/smp_load_acquire() across various subsystems.
> >>
> >> Background
> >> ==========
> >>
> >> Many architectures support load acquire and store release instructions
> >> which can replace explicit memory barriers and save cycles. As noted
> >> in the ARM architecture reference [1]:
> >>
> >>   "Weaker ordering requirements that are imposed by Load-Acquire and
> >>    Store-Release instructions allow for micro-architectural
> >>    optimizations, which could reduce some of the performance impacts
> >>    that are otherwise imposed by an explicit memory barrier.
> >>
> >>    If the ordering requirement is satisfied using either a Load-Acquire
> >>    or Store-Release, then it would be preferable to use these
> >>    instructions instead of a DMB."
> >>
> >> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
> >> barriers. Replacing the read barrier with smp_load_acquire() reduces
> >> this to 8 cycles on an Ampere Altra.  
> > 
> > I wonder if this is true on all other architectures:
> > 
> >   + It seems that Arm gets the gain because the instruction
> >     does both load/store + barrier. It helps even when
> >     the barrier is full.
> > 
> >   + Some other architectures need two instructions. One for the
> >     load/store and the other for the barrier. But the barrier
> >     is weaker, it synchronizes just reads or just writes.
> > 
> > For example, I see the following in riscv/include/asm/barrier.h:
> > 
> > <paste riscv/include/asm/barrier.h>
> > #define smp_mb()	RISCV_FENCE(rw, rw)
> > #define smp_rmb()	RISCV_FENCE(r, r)
> > #define smp_wmb()	RISCV_FENCE(w, w)
> > 
> > #define smp_store_release(p, v)						\
> > do {									\
> > 	RISCV_FENCE(rw, w);						\
> > 	WRITE_ONCE(*p, v);						\
> > } while (0)
> > 
> > #define smp_load_acquire(p)						\
> > ({									\
> > 	typeof(*p) ___p1 = READ_ONCE(*p);				\
> > 	RISCV_FENCE(r, rw);						\
> > 	___p1;								\
> > })
> > </paste riscv/include/asm/barrier.h>
> > 
> > I wonder whether:
> > 
> >   + RISCV_FENCE(r, r) is faster than RISCV_FENCE(r, rw)
> >   + RISCV_FENCE(w, w) is faster than RISCV_FENCE(rw, w)
> > 
> > so it might cause performance regression there...  
> 
> Hi Petr,
> 
> Thanks for the detailed analysis. You're right that on some
> architectures the acquire/release variants use a slightly heavier
> fence than the plain smp_wmb()/smp_rmb() pair. The full picture by
> architecture:
> 
>   arm64:    Improvement   (DMB ISHST/ISHLD + STR/LDR → STLR/LDAR)
>   x86:      Neutral       (both are compiler barriers)
>   s390:     Neutral       (both are compiler barriers)
>   ppc64:    Neutral       (both use lwsync)
>   loongarch: Slightly heavier fence (DBAR(o_w_w) → DBAR(orw_w),
>                            DBAR(or_r_) → DBAR(or_rw))
>   riscv:    Slightly heavier fence (fence w,w → fence rw,w,
>                            fence r,r → fence r,rw)
> 
> So on Loongarch and Riscv, there may be a slight performance regression.

Is that a bug in the riscv definitions?
Nothing in the commit messages seems to indicate why the stronger barriers
are used.
The original commit 8d235b17 was done to avoid the rw,rw barrier in the
generic code (which might since have been relaxed).

David

> 
> Regards,
> Jinjie
> 
> > 
> > Best Regards,
> > Petr
> >   
> >> We also observed significant barrier overhead while profiling Unxibench
> >> syscall test on arm64: a single getuid() call is ~8ns slower than on
> >> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
> >> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
> >> the use of LDAR, eliminating the measurable overhead.
> >>
> >> This motivated a broader search for existing barrier pairs that can
> >> be converted to the lighter acquire/release semantics.
> >>
> >> Changes
> >> =======
> >>
> >> Each patch in this series targets a specific barrier pair where the
> >> publish/subscribe pattern is already present:
> >>
> >> - Writers populate data, then publish a flag/count/pointer via
> >>   smp_store_release()
> >>
> >> - Readers load the flag/count/pointer via smp_load_acquire(), then
> >>   consume the data
> >>
> >> This preserves the existing memory ordering guarantees while allowing
> >> architectures with native acquire/release instructions (e.g. arm64's
> >> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
> >> On architectures without native support, the generated code is
> >> generally no worse than the explicit barrier pair.
> >>
> >> The conversions are mechanical and no functional change is intended.
> >>
> >> Testing (Kunpeng HIP09 arm64 server)
> >> ====================================
> >>
> >> 1. UNIXBENCH syscall
> >> 	Baseline: 715.27
> >> 	Patched:  718.83
> >> 	Improvement: +0.50%
> >>
> >> 2. fs/aio (fio + null_blk, 4 jobs):
> >> 	Baseline: 1441k IOPS, 86.46us
> >> 	Patched:  1452k IOPS, 85.80us
> >> 	Improvement: ~0.8%
> >>
> >> Both improvements are consistent across runs and align with the
> >> expected savings from replacing DMB with LDAR/STLR on arm64.
> >>
> >> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
> >> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
> >>
> >> Changes in v3:
> >> - Add Reviewed-by.
> >> - Split out network patch set as Kuniyuki suggested.
> >> - Link to v2: https://lore.kernel.org/all/20260901024234.135119-1-ruanjinjie@huawei.com/
> >>
> >> Changes in v2:
> >> - Fix pre-existing issue for ext4 and 8021q [3].
> >> - Fix missing copy_mnt_idmap() udapte [3].
> >> - Drop nacked isotp patch.
> >> - Add test data.
> >> - Add Reviewed-by and update fs patch as Jan suggested.
> >>
> >> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
> >>
> >> Jinjie Ruan (8):
> >>   user_namespace: Use acquire/release for nr_extents synchronization
> >>   lib/vsprintf: Use acquire/release for ptr_key publication
> >>   fs: aio: Use acquire/release for ring->tail publication
> >>   fs: Use acquire/release for fdtable resize synchronization
> >>   pidfs: Use test_bit_acquire() for attr flag tests
> >>   super: Use acquire for SB_BORN check in super_cache_count()
> >>   ext4: Fix out-of-bounds read in ext4_get_group_info()
> >>   ext4: Convert group-count barrier protocol to acquire/release
> >>
> >>  fs/aio.c                | 10 ++++------
> >>  fs/ext4/balloc.c        |  2 +-
> >>  fs/ext4/ext4.h          | 10 +++-------
> >>  fs/ext4/mballoc.c       |  6 ++----
> >>  fs/ext4/resize.c        | 19 +++++++++++--------
> >>  fs/file.c               | 10 ++++------
> >>  fs/mnt_idmapping.c      |  5 ++---
> >>  fs/pidfs.c              |  6 ++----
> >>  fs/super.c              |  6 ++----
> >>  kernel/user_namespace.c | 24 +++++++++++++-----------
> >>  lib/vsprintf.c          | 11 ++++-------
> >>  11 files changed, 48 insertions(+), 61 deletions(-)
> >>
> >> -- 
> >> 2.34.1  
> 
> 
Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Posted by Jinjie Ruan 2 weeks, 3 days ago

在 2026/9/7 20:59, David Laight 写道:
> On Mon, 7 Sep 2026 19:29:06 +0800
> Jinjie Ruan <ruanjinjie@huawei.com> wrote:
> 
>> 在 2026/9/4 20:28, Petr Mladek 写道:
>>> Adding Risc-V list into Cc.
>>>
>>> On Wed 2026-09-02 15:47:57, Jinjie Ruan wrote:  
>>>> Hi,
>>>>
>>>> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
>>>> smp_store_release()/smp_load_acquire() across various subsystems.
>>>>
>>>> Background
>>>> ==========
>>>>
>>>> Many architectures support load acquire and store release instructions
>>>> which can replace explicit memory barriers and save cycles. As noted
>>>> in the ARM architecture reference [1]:
>>>>
>>>>   "Weaker ordering requirements that are imposed by Load-Acquire and
>>>>    Store-Release instructions allow for micro-architectural
>>>>    optimizations, which could reduce some of the performance impacts
>>>>    that are otherwise imposed by an explicit memory barrier.
>>>>
>>>>    If the ordering requirement is satisfied using either a Load-Acquire
>>>>    or Store-Release, then it would be preferable to use these
>>>>    instructions instead of a DMB."
>>>>
>>>> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
>>>> barriers. Replacing the read barrier with smp_load_acquire() reduces
>>>> this to 8 cycles on an Ampere Altra.  
>>>
>>> I wonder if this is true on all other architectures:
>>>
>>>   + It seems that Arm gets the gain because the instruction
>>>     does both load/store + barrier. It helps even when
>>>     the barrier is full.
>>>
>>>   + Some other architectures need two instructions. One for the
>>>     load/store and the other for the barrier. But the barrier
>>>     is weaker, it synchronizes just reads or just writes.
>>>
>>> For example, I see the following in riscv/include/asm/barrier.h:
>>>
>>> <paste riscv/include/asm/barrier.h>
>>> #define smp_mb()	RISCV_FENCE(rw, rw)
>>> #define smp_rmb()	RISCV_FENCE(r, r)
>>> #define smp_wmb()	RISCV_FENCE(w, w)
>>>
>>> #define smp_store_release(p, v)						\
>>> do {									\
>>> 	RISCV_FENCE(rw, w);						\
>>> 	WRITE_ONCE(*p, v);						\
>>> } while (0)
>>>
>>> #define smp_load_acquire(p)						\
>>> ({									\
>>> 	typeof(*p) ___p1 = READ_ONCE(*p);				\
>>> 	RISCV_FENCE(r, rw);						\
>>> 	___p1;								\
>>> })
>>> </paste riscv/include/asm/barrier.h>
>>>
>>> I wonder whether:
>>>
>>>   + RISCV_FENCE(r, r) is faster than RISCV_FENCE(r, rw)
>>>   + RISCV_FENCE(w, w) is faster than RISCV_FENCE(rw, w)
>>>
>>> so it might cause performance regression there...  
>>
>> Hi Petr,
>>
>> Thanks for the detailed analysis. You're right that on some
>> architectures the acquire/release variants use a slightly heavier
>> fence than the plain smp_wmb()/smp_rmb() pair. The full picture by
>> architecture:
>>
>>   arm64:    Improvement   (DMB ISHST/ISHLD + STR/LDR → STLR/LDAR)
>>   x86:      Neutral       (both are compiler barriers)
>>   s390:     Neutral       (both are compiler barriers)
>>   ppc64:    Neutral       (both use lwsync)
>>   loongarch: Slightly heavier fence (DBAR(o_w_w) → DBAR(orw_w),
>>                            DBAR(or_r_) → DBAR(or_rw))
>>   riscv:    Slightly heavier fence (fence w,w → fence rw,w,
>>                            fence r,r → fence r,rw)
>>
>> So on Loongarch and Riscv, there may be a slight performance regression.
> 
> Is that a bug in the riscv definitions?
> Nothing in the commit messages seems to indicate why the stronger barriers
> are used.
> The original commit 8d235b17 was done to avoid the rw,rw barrier in the
> generic code (which might since have been relaxed).

Hi David,

I think the riscv definition is correct, which are consistent with the
descriptions found in the ARM64 architecture manual [2]. Per the RISC-V
RVWMO specification [1], the semantics of these fences are:

- FENCE RW, W (for Store-Release):

"Release orderings work exactly the same as acquire orderings, just in
the opposite direction. Release semantics require all loads and stores
preceding the release operation in program order to also precede the
release operation in the global memory order. This ensures, for example,
that memory accesses in a critical section appear before the
lock-releasing store in the global memory order. Just as for acquire
semantics, release semantics can be enforced using release annotations
or with a FENCE RW,W operation."

- FENCE R, RW (for Load-Acquire):

"An acquire operation, as would be used at the start of a critical
section, requires all memory operations following the acquire in program
order to also follow the acquire in the global memory order. This
ensures, for example, that all loads and stores inside the critical
section are up to date with respect to the synchronization variable
being used to protect it. Acquire ordering can be enforced in one of two
ways: with an acquire annotation, which enforces ordering with respect
to just the synchronization variable itself, or with a FENCE R,RW, which
enforces ordering with respect to all previous loads."

Regardless, I believe that replacing explicit with acquire/release
annotations would make the current code semantics clearer.

[1]: https://docs.riscv.org/reference/isa/v20260120/unpriv/mm-eplan.html
[2]:
https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions

Best regards,
Jinjie

> 
> David
> 
>>
>> Regards,
>> Jinjie
>>
>>>
>>> Best Regards,
>>> Petr
>>>   
>>>> We also observed significant barrier overhead while profiling Unxibench
>>>> syscall test on arm64: a single getuid() call is ~8ns slower than on
>>>> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
>>>> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
>>>> the use of LDAR, eliminating the measurable overhead.
>>>>
>>>> This motivated a broader search for existing barrier pairs that can
>>>> be converted to the lighter acquire/release semantics.
>>>>
>>>> Changes
>>>> =======
>>>>
>>>> Each patch in this series targets a specific barrier pair where the
>>>> publish/subscribe pattern is already present:
>>>>
>>>> - Writers populate data, then publish a flag/count/pointer via
>>>>   smp_store_release()
>>>>
>>>> - Readers load the flag/count/pointer via smp_load_acquire(), then
>>>>   consume the data
>>>>
>>>> This preserves the existing memory ordering guarantees while allowing
>>>> architectures with native acquire/release instructions (e.g. arm64's
>>>> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
>>>> On architectures without native support, the generated code is
>>>> generally no worse than the explicit barrier pair.
>>>>
>>>> The conversions are mechanical and no functional change is intended.
>>>>
>>>> Testing (Kunpeng HIP09 arm64 server)
>>>> ====================================
>>>>
>>>> 1. UNIXBENCH syscall
>>>> 	Baseline: 715.27
>>>> 	Patched:  718.83
>>>> 	Improvement: +0.50%
>>>>
>>>> 2. fs/aio (fio + null_blk, 4 jobs):
>>>> 	Baseline: 1441k IOPS, 86.46us
>>>> 	Patched:  1452k IOPS, 85.80us
>>>> 	Improvement: ~0.8%
>>>>
>>>> Both improvements are consistent across runs and align with the
>>>> expected savings from replacing DMB with LDAR/STLR on arm64.
>>>>
>>>> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
>>>> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
>>>>
>>>> Changes in v3:
>>>> - Add Reviewed-by.
>>>> - Split out network patch set as Kuniyuki suggested.
>>>> - Link to v2: https://lore.kernel.org/all/20260901024234.135119-1-ruanjinjie@huawei.com/
>>>>
>>>> Changes in v2:
>>>> - Fix pre-existing issue for ext4 and 8021q [3].
>>>> - Fix missing copy_mnt_idmap() udapte [3].
>>>> - Drop nacked isotp patch.
>>>> - Add test data.
>>>> - Add Reviewed-by and update fs patch as Jan suggested.
>>>>
>>>> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
>>>>
>>>> Jinjie Ruan (8):
>>>>   user_namespace: Use acquire/release for nr_extents synchronization
>>>>   lib/vsprintf: Use acquire/release for ptr_key publication
>>>>   fs: aio: Use acquire/release for ring->tail publication
>>>>   fs: Use acquire/release for fdtable resize synchronization
>>>>   pidfs: Use test_bit_acquire() for attr flag tests
>>>>   super: Use acquire for SB_BORN check in super_cache_count()
>>>>   ext4: Fix out-of-bounds read in ext4_get_group_info()
>>>>   ext4: Convert group-count barrier protocol to acquire/release
>>>>
>>>>  fs/aio.c                | 10 ++++------
>>>>  fs/ext4/balloc.c        |  2 +-
>>>>  fs/ext4/ext4.h          | 10 +++-------
>>>>  fs/ext4/mballoc.c       |  6 ++----
>>>>  fs/ext4/resize.c        | 19 +++++++++++--------
>>>>  fs/file.c               | 10 ++++------
>>>>  fs/mnt_idmapping.c      |  5 ++---
>>>>  fs/pidfs.c              |  6 ++----
>>>>  fs/super.c              |  6 ++----
>>>>  kernel/user_namespace.c | 24 +++++++++++++-----------
>>>>  lib/vsprintf.c          | 11 ++++-------
>>>>  11 files changed, 48 insertions(+), 61 deletions(-)
>>>>
>>>> -- 
>>>> 2.34.1  
>>
>>
>