[PATCH 0/6] x86: add missing vzeroupper instructions

Eric Biggers posted 6 patches 1 month, 2 weeks ago
arch/x86/crypto/aria-aesni-avx2-asm_64.S  |  6 ++++++
arch/x86/crypto/aria-gfni-avx512-asm_64.S |  3 +++
lib/raid/raid6/x86/avx2.c                 |  6 ++++++
lib/raid/raid6/x86/avx512.c               |  6 ++++++
lib/raid/raid6/x86/recov_avx2.c           |  2 ++
lib/raid/raid6/x86/recov_avx512.c         |  2 ++
lib/raid/xor/x86/xor-avx.c                |  1 +
net/netfilter/nft_set_pipapo_avx2.c       | 17 ++++++++---------
8 files changed, 34 insertions(+), 9 deletions(-)
[PATCH 0/6] x86: add missing vzeroupper instructions
Posted by Eric Biggers 1 month, 2 weeks ago
Assembly code using YMM or ZMM registers is supposed to end with the
vzeroupper instruction in order to avoid degrading the performance of
any later SSE code.  Since this only affects performance and not
correctness, it is sometimes overlooked.  Most kernel code does it
correctly, but a few cases were missed.  This series fixes them.

It should be easiest to take the full series through the x86 tree.

Eric Biggers (6):
  xor: add missing vzeroupper to AVX code
  raid6: add missing vzeroupper to AVX2 code
  raid6: add missing vzeroupper to AVX-512 code
  crypto: x86/aria - add missing vzeroupper in AVX2 code
  crypto: x86/aria - add missing vzeroupper in AVX-512 code
  netfilter: nft_set_pipapo_avx2: add missing vzeroupper

 arch/x86/crypto/aria-aesni-avx2-asm_64.S  |  6 ++++++
 arch/x86/crypto/aria-gfni-avx512-asm_64.S |  3 +++
 lib/raid/raid6/x86/avx2.c                 |  6 ++++++
 lib/raid/raid6/x86/avx512.c               |  6 ++++++
 lib/raid/raid6/x86/recov_avx2.c           |  2 ++
 lib/raid/raid6/x86/recov_avx512.c         |  2 ++
 lib/raid/xor/x86/xor-avx.c                |  1 +
 net/netfilter/nft_set_pipapo_avx2.c       | 17 ++++++++---------
 8 files changed, 34 insertions(+), 9 deletions(-)


base-commit: db2ddb87143519e20a95aa36c60b36107b736a58
-- 
2.55.0
Re: [PATCH 0/6] x86: add missing vzeroupper instructions
Posted by David Laight 1 month, 1 week ago
On Sat, 15 Aug 2026 13:57:44 -0700
Eric Biggers <ebiggers@kernel.org> wrote:

> Assembly code using YMM or ZMM registers is supposed to end with the
> vzeroupper instruction in order to avoid degrading the performance of
> any later SSE code.  Since this only affects performance and not
> correctness, it is sometimes overlooked.  Most kernel code does it
> correctly, but a few cases were missed.  This series fixes them.
> 
> It should be easiest to take the full series through the x86 tree.

Would it be better to an an unconditional vzeroupper in kernel_fpu_end()?
It could go in kernel_fpu_start() but that might have a bigger effect
on latency.
(I assume there is one in kernel_fpu_start() if it actually saves
the user registers?)

Looking the latency/uops seems reasonably on everything 'recent' except
zen-1 and knights-landing.
Although the microcode patch for zen-2 might make that a lot worse.

	David  


> 
> Eric Biggers (6):
>   xor: add missing vzeroupper to AVX code
>   raid6: add missing vzeroupper to AVX2 code
>   raid6: add missing vzeroupper to AVX-512 code
>   crypto: x86/aria - add missing vzeroupper in AVX2 code
>   crypto: x86/aria - add missing vzeroupper in AVX-512 code
>   netfilter: nft_set_pipapo_avx2: add missing vzeroupper
> 
>  arch/x86/crypto/aria-aesni-avx2-asm_64.S  |  6 ++++++
>  arch/x86/crypto/aria-gfni-avx512-asm_64.S |  3 +++
>  lib/raid/raid6/x86/avx2.c                 |  6 ++++++
>  lib/raid/raid6/x86/avx512.c               |  6 ++++++
>  lib/raid/raid6/x86/recov_avx2.c           |  2 ++
>  lib/raid/raid6/x86/recov_avx512.c         |  2 ++
>  lib/raid/xor/x86/xor-avx.c                |  1 +
>  net/netfilter/nft_set_pipapo_avx2.c       | 17 ++++++++---------
>  8 files changed, 34 insertions(+), 9 deletions(-)
> 
> 
> base-commit: db2ddb87143519e20a95aa36c60b36107b736a58
Re: [PATCH 0/6] x86: add missing vzeroupper instructions
Posted by Eric Biggers 1 month, 1 week ago
On Sun, Aug 16, 2026 at 04:14:50PM +0100, David Laight wrote:
> On Sat, 15 Aug 2026 13:57:44 -0700
> Eric Biggers <ebiggers@kernel.org> wrote:
> 
> > Assembly code using YMM or ZMM registers is supposed to end with the
> > vzeroupper instruction in order to avoid degrading the performance of
> > any later SSE code.  Since this only affects performance and not
> > correctness, it is sometimes overlooked.  Most kernel code does it
> > correctly, but a few cases were missed.  This series fixes them.
> > 
> > It should be easiest to take the full series through the x86 tree.
> 
> Would it be better to an an unconditional vzeroupper in kernel_fpu_end()?
> It could go in kernel_fpu_start() but that might have a bigger effect
> on latency.
> (I assume there is one in kernel_fpu_start() if it actually saves
> the user registers?)
> 
> Looking the latency/uops seems reasonably on everything 'recent' except
> zen-1 and knights-landing.
> Although the microcode patch for zen-2 might make that a lot worse.
> 
> 	David  

I would like to do that, but there are some issues:

- In some cases, within a single kernel-mode FPU section the kernel
  executes AVX instructions, then SSE instructions afterwards.
  It typically occurs when large input lengths are optimized specially
  with AVX and then shorter lengths fall back to SSE code.
  chacha_dosimd() in lib/crypto/x86/chacha.h is an example of this.

  In these cases the internal vzeroupper is definitely needed.

- Unnecessary overhead if only SSE instructions are used, which is still
  frequent since we don't provide both SSE and AVX versions of the same
  code when the AVX doesn't provide a notable performance benefit.
  lib/crypto/x86/sha256-ni-asm.S is an example of this.

- Further divergence from the userspace ABI, which can be annoying when
  importing assembly code to or from userspace projects, or developing
  or testing the assembly files in userspace.

As for kernel_fpu_begin(), no, it doesn't do vzeroupper.

I do think that some years down the line, we'll drop the use of SSE in
the kernel entirely.  At that point, vzeroupper in kernel_fpu_end()
would make sense.

But until then, I think we should stick with the existing, standard
convention of having the vzeroupper at the end of the assembly routines.

- Eric
Re: [PATCH 0/6] x86: add missing vzeroupper instructions
Posted by Christoph Hellwig 1 month, 1 week ago
On Sun, Aug 16, 2026 at 10:31:59AM -0700, Eric Biggers wrote:
> As for kernel_fpu_begin(), no, it doesn't do vzeroupper.
> 
> I do think that some years down the line, we'll drop the use of SSE in
> the kernel entirely.  At that point, vzeroupper in kernel_fpu_end()
> would make sense.

Or add kernel_avx_{begin,end} wrappers that include the vzeroupper
in kernel_avx_end.  That would be a lot easier to use than the manual
vzeroupper in every modern user of in-kernel AVX.
Re: [PATCH 0/6] x86: add missing vzeroupper instructions
Posted by David Laight 1 month, 1 week ago
On Mon, 17 Aug 2026 02:15:19 -0700
Christoph Hellwig <hch@infradead.org> wrote:

> On Sun, Aug 16, 2026 at 10:31:59AM -0700, Eric Biggers wrote:
> > As for kernel_fpu_begin(), no, it doesn't do vzeroupper.
> > 
> > I do think that some years down the line, we'll drop the use of SSE in
> > the kernel entirely.  At that point, vzeroupper in kernel_fpu_end()
> > would make sense.  
> 
> Or add kernel_avx_{begin,end} wrappers that include the vzeroupper
> in kernel_avx_end.  That would be a lot easier to use than the manual
> vzeroupper in every modern user of in-kernel AVX.
> 

You might want one in the start as well.
I have a theory that the avx512 logic was added as a completely separate block.
This meant it could be included in cpu for testing but disabled in any
released to customers.
(Or maybe the it is the original avx logic that used latches not in the
normal register file.)
A side effect is that different latches are used for the low bits of the
registers - so when you change to/from avx512 the register contents have to
be transferred between the blocks - adding latency.
So if the wrong registers are live for the code you are going to execute
the data has to be transferred across.

There are also other effects as well.
I found this link: https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html
It is a few years old now (2020) but probably still relevant.
A quick summary is that the first 256 or 512 bit instruction starts a 9us
window where the cpu runs at 1/4 speed, for 512 bit that is followed by 11us
where nothing happens at all. 

	David
Re: [PATCH 0/6] x86: add missing vzeroupper instructions
Posted by Eric Biggers 1 month, 1 week ago
On Mon, Aug 17, 2026 at 11:55:23AM +0100, David Laight wrote:
> On Mon, 17 Aug 2026 02:15:19 -0700
> Christoph Hellwig <hch@infradead.org> wrote:
> 
> > On Sun, Aug 16, 2026 at 10:31:59AM -0700, Eric Biggers wrote:
> > > As for kernel_fpu_begin(), no, it doesn't do vzeroupper.
> > > 
> > > I do think that some years down the line, we'll drop the use of SSE in
> > > the kernel entirely.  At that point, vzeroupper in kernel_fpu_end()
> > > would make sense.  
> > 
> > Or add kernel_avx_{begin,end} wrappers that include the vzeroupper
> > in kernel_avx_end.  That would be a lot easier to use than the manual
> > vzeroupper in every modern user of in-kernel AVX.
> > 
> 
> You might want one in the start as well.
> I have a theory that the avx512 logic was added as a completely separate block.
> This meant it could be included in cpu for testing but disabled in any
> released to customers.
> (Or maybe the it is the original avx logic that used latches not in the
> normal register file.)
> A side effect is that different latches are used for the low bits of the
> registers - so when you change to/from avx512 the register contents have to
> be transferred between the blocks - adding latency.
> So if the wrong registers are live for the code you are going to execute
> the data has to be transferred across.
> 
> There are also other effects as well.
> I found this link: https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html
> It is a few years old now (2020) but probably still relevant.
> A quick summary is that the first 256 or 512 bit instruction starts a 9us
> window where the cpu runs at 1/4 speed, for 512 bit that is followed by 11us
> where nothing happens at all. 

The linked article is about Skylake, which is an older Intel CPU that
has a bad AVX-512 implementation with overly-eager downclocking.  Later
Intel CPUs improved the implementation.  And of course, AMD just
implemented it properly from the start without the downclocking issues.
Information about AMD Zen 5's AVX-512 implementation can be found here:
https://www.numberworld.org/blogs/2024_8_7_zen5_avx512_teardown/

Most of the AVX-512 optimized code in the kernel already requires
!X86_FEATURE_PREFER_YMM, excluding Skylake as well as Ice Lake.

That being said, if I recall correctly, even with Intel's improved
implementation on Sapphire Rapids and Emerald Rapids, Intel does still
have some start-up latency for accessing ZMM registers.  AMD doesn't.  I
don't believe vzeroupper helps, unfortunately.

I've considered setting X86_FEATURE_PREFER_YMM on all Intel CPUs, but
then even workloads that would benefit from ZMM registers wouldn't use
them.  And I suspect the Intel folks wouldn't agree with that either.

- Eric