From nobody Tue Sep 29 12:02:21 2026 Received: from galois.linutronix.de (Galois.linutronix.de [193.142.43.55]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 94C8438B7D2; Fri, 7 Aug 2026 16:00:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=193.142.43.55 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786118412; cv=none; b=tjgs9OKQ51AH406IMpFTOf7sE1ak/t2lO/TLMdUXjt0dSwRJ/bajms4WGY/blXIdLETUqgAVNe2EfoMxrawHCcVfO+KoMISLZq60zUidrEWqeTnbUpwMYD5iFCiUQeRySuaqe+dgMPppYTWVO9SDuPDHQaumpRv6COTDiu9GwZg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786118412; c=relaxed/simple; bh=YoDrJHBKQ4Qky9lSA6fAE35TCtrmQwxCn8k2LPjOfT4=; h=Date:From:To:Subject:Cc:In-Reply-To:References:MIME-Version: Message-ID:Content-Type; b=hmCuSTnZGmBskMYvSNCLDgogXStcw7v+ZFkHWlu6oN7w57ELBtVRvLx6j1ZUHKITQs/9DPoKNfLrz9911ob9oPfI2L6Sdks7QyPxTTNySkG622+5lBVsa6L4FXvzMrJ0wbbbvdNw742jPFwybAyAZTiyG3p4+//9V+LY3wuyFDw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de; spf=pass smtp.mailfrom=linutronix.de; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=BCRjX4gg; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=zjzdeXu1; arc=none smtp.client-ip=193.142.43.55 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linutronix.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="BCRjX4gg"; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="zjzdeXu1" Date: Fri, 07 Aug 2026 16:00:07 -0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020; t=1786118409; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=1wKVYGoJpuVxgQ4KuwrGmGCnV4YciylTPFPgsBygMZY=; b=BCRjX4ggTwpJyki/MUA59viwx7F84uz4yrq7pbOPtQW9a86038yyQQXdq8rSqb5WKUtoJu tF+fj0p0IEwZllosPR+Qac7EjUE5ZyAfH1uzbQKXeY0+wEUg6wudd6AZL4psmhAaCgqfKb Gi+7mXeZl46iCJYeNYxc/sbguHdGj5eWCrQc4O2ljwbsL+M5ULAKzANcr6BlhMQu+o81ip ulOSDi2SWCvoyS3ImJ4N2LlzzC//b2i1Et70HR871hmah4zfBd34qCweUCShIzb1nyZH2/ xTjUgInrNCck3OUuVbIX4vj+lV9jOHl9aSR9EwAZBFLTa9ld0EFE0MTsdU2gag== DKIM-Signature: v=1; a=ed25519-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020e; t=1786118409; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=1wKVYGoJpuVxgQ4KuwrGmGCnV4YciylTPFPgsBygMZY=; b=zjzdeXu1C5hBF5wYxZsKAf+vVJ4yNeu8GQOT47+20j5aiyYADCKo1yG7/LN5yuS6d1afZs ECQ2XDp72iRCJnDg== From: "tip-bot2 for Peter Zijlstra" Sender: tip-bot2@linutronix.de Reply-to: linux-kernel@vger.kernel.org To: linux-tip-commits@vger.kernel.org Subject: [tip: locking/core] x86/paravirt: Use static_call() for the paravirt spinlock ops Cc: "Peter Zijlstra (Intel)" , Dmitry Ilvokhin , Juergen Gross , x86@kernel.org, linux-kernel@vger.kernel.org In-Reply-To: <9a32ae399eb804a02a31af04dcabe7e7ee4f3fdf.1785778551.git.d@ilvokhin.com> References: <9a32ae399eb804a02a31af04dcabe7e7ee4f3fdf.1785778551.git.d@ilvokhin.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Message-ID: <178611840704.708.8737883032124119716.tip-bot2@tip-bot2> Robot-ID: Robot-Unsubscribe: Contact to get blacklisted from these emails Precedence: bulk Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable The following commit has been merged into the locking/core branch of tip: Commit-ID: fc6ad5eadcd1694d2f1ee2717c8024efd8152e20 Gitweb: https://git.kernel.org/tip/fc6ad5eadcd1694d2f1ee2717c8024efd= 8152e20 Author: Peter Zijlstra AuthorDate: Tue, 04 Aug 2026 07:15:41=20 Committer: Peter Zijlstra CommitterDate: Fri, 07 Aug 2026 17:58:09 +02:00 x86/paravirt: Use static_call() for the paravirt spinlock ops queued_spin_lock_slowpath() and queued_spin_unlock() are dispatched through pv_ops_lock via the paravirt-ops ALTERNATIVE machinery, which picks the target (native inline store / hypervisor call) once at boot and cannot change at runtime. Convert both to static_call(). The site becomes a direct call patched in place (one byte smaller), and on native the unlock still collapses to the inline "movb $0, (%rdi)" store, so the fast path is unchanged. Unlike the ALTERNATIVE mechanism, a static_call() target can also be updated at runtime via static_call_update(). This is a prerequisite for the contended_release tracepoint, which has to swap in a traced unlock while the system is running. [ ilvokhin: commit message; fix PARAVIRT_SPINLOCKS=3Dn build; teach __static_call_validate() about the inline unlock insn; make the slowpath site module-safe: static_call_mod() + EXPORT_STATIC_CALL_TRAMP(); pass @lock to the callee-save unlock, fixing a boot hang under CALL_DEPTH_TRACKING. Boot tested native + KVM PV guest. ] Signed-off-by: Peter Zijlstra (Intel) Co-developed-by: Dmitry Ilvokhin Signed-off-by: Dmitry Ilvokhin Signed-off-by: Peter Zijlstra (Intel) Acked-by: Juergen Gross Link: https://lore.kernel.org/all/20260603120811.GW3493090@noisy.programmin= g.kicks-ass.net/ Link: https://patch.msgid.link/9a32ae399eb804a02a31af04dcabe7e7ee4f3fdf.178= 5778551.git.d@ilvokhin.com --- arch/x86/hyperv/hv_spinlock.c | 4 +-- arch/x86/include/asm/cpufeatures.h | 2 +- arch/x86/include/asm/paravirt-spinlock.h | 19 ++++++++++------ arch/x86/kernel/kvm.c | 5 +--- arch/x86/kernel/paravirt-spinlocks.c | 12 +++++----- arch/x86/kernel/static_call.c | 27 +++++++++++++++++++++++- arch/x86/xen/spinlock.c | 5 +--- tools/arch/x86/include/asm/cpufeatures.h | 2 +- 8 files changed, 53 insertions(+), 23 deletions(-) diff --git a/arch/x86/hyperv/hv_spinlock.c b/arch/x86/hyperv/hv_spinlock.c index 210b494..6b4bdea 100644 --- a/arch/x86/hyperv/hv_spinlock.c +++ b/arch/x86/hyperv/hv_spinlock.c @@ -78,8 +78,8 @@ void __init hv_init_spinlocks(void) pr_info("PV spinlocks enabled\n"); =20 __pv_init_lock_hash(); - pv_ops_lock.queued_spin_lock_slowpath =3D __pv_queued_spin_lock_slowpath; - pv_ops_lock.queued_spin_unlock =3D PV_CALLEE_SAVE(__pv_queued_spin_unlock= ); + static_call_update(queued_spin_lock_slowpath, __pv_queued_spin_lock_slowp= ath); + static_call_update(queued_spin_unlock, __raw_callee_save___pv_queued_spin= _unlock); pv_ops_lock.wait =3D hv_qlock_wait; pv_ops_lock.kick =3D hv_qlock_kick; pv_ops_lock.vcpu_is_preempted =3D PV_CALLEE_SAVE(hv_vcpu_is_preempted); diff --git a/arch/x86/include/asm/cpufeatures.h b/arch/x86/include/asm/cpuf= eatures.h index 1b4a48b..5522415 100644 --- a/arch/x86/include/asm/cpufeatures.h +++ b/arch/x86/include/asm/cpufeatures.h @@ -225,7 +225,7 @@ #define X86_FEATURE_EPT_AD ( 8*32+17) /* "ept_ad" Intel Extended Page Tab= le access-dirty bit */ #define X86_FEATURE_VMCALL ( 8*32+18) /* Hypervisor supports the VMCALL i= nstruction */ #define X86_FEATURE_VMW_VMMCALL ( 8*32+19) /* VMware prefers VMMCALL hype= rcall instruction */ -#define X86_FEATURE_PVUNLOCK ( 8*32+20) /* PV unlock function */ +// free: was #define X86_FEATURE_PVUNLOCK ( 8*32+20) /* PV u= nlock function */ #define X86_FEATURE_VCPUPREEMPT ( 8*32+21) /* PV vcpu_is_preempted functi= on */ #define X86_FEATURE_TDX_GUEST ( 8*32+22) /* "tdx_guest" Intel Trust Domai= n Extensions Guest */ =20 diff --git a/arch/x86/include/asm/paravirt-spinlock.h b/arch/x86/include/as= m/paravirt-spinlock.h index 7beffcb..ff73583 100644 --- a/arch/x86/include/asm/paravirt-spinlock.h +++ b/arch/x86/include/asm/paravirt-spinlock.h @@ -3,6 +3,7 @@ #define _ASM_X86_PARAVIRT_SPINLOCK_H =20 #include +#include =20 #ifdef CONFIG_SMP #include @@ -11,9 +12,6 @@ struct qspinlock; =20 struct pv_lock_ops { - void (*queued_spin_lock_slowpath)(struct qspinlock *lock, u32 val); - struct paravirt_callee_save queued_spin_unlock; - void (*wait)(u8 *ptr, u8 val); void (*kick)(int cpu); =20 @@ -26,20 +24,27 @@ extern struct pv_lock_ops pv_ops_lock; extern void native_queued_spin_lock_slowpath(struct qspinlock *lock, u32 v= al); extern void __pv_init_lock_hash(void); extern void __pv_queued_spin_lock_slowpath(struct qspinlock *lock, u32 val= ); +extern void __raw_callee_save___native_queued_spin_unlock(struct qspinlock= *lock); extern void __raw_callee_save___pv_queued_spin_unlock(struct qspinlock *lo= ck); extern bool nopvspin; =20 +DECLARE_STATIC_CALL(queued_spin_lock_slowpath, native_queued_spin_lock_slo= wpath); +DECLARE_STATIC_CALL(queued_spin_unlock, __raw_callee_save___native_queued_= spin_unlock); + static __always_inline void pv_queued_spin_lock_slowpath(struct qspinlock = *lock, u32 val) { - PVOP_VCALL2(pv_ops_lock, queued_spin_lock_slowpath, lock, val); + static_call_mod(queued_spin_lock_slowpath)(lock, val); } =20 static __always_inline void pv_queued_spin_unlock(struct qspinlock *lock) { - PVOP_ALT_VCALLEE1(pv_ops_lock, queued_spin_unlock, lock, - "movb $0, (%%" _ASM_ARG1 ")", - ALT_NOT(X86_FEATURE_PVUNLOCK)); + PVOP_CALL_ARGS; + __STATIC_CALL_MOD_ADDRESSABLE(queued_spin_unlock); + asm volatile ("call " STATIC_CALL_TRAMP_STR(queued_spin_unlock) + : PVOP_VCALLEE_CLOBBERS, ASM_CALL_CONSTRAINT + : PVOP_CALL_ARG1(lock) + : "memory", "cc"); } =20 static __always_inline bool pv_vcpu_is_preempted(long cpu) diff --git a/arch/x86/kernel/kvm.c b/arch/x86/kernel/kvm.c index dcef84d..253c159 100644 --- a/arch/x86/kernel/kvm.c +++ b/arch/x86/kernel/kvm.c @@ -1136,9 +1136,8 @@ void __init kvm_spinlock_init(void) pr_info("PV spinlocks enabled\n"); =20 __pv_init_lock_hash(); - pv_ops_lock.queued_spin_lock_slowpath =3D __pv_queued_spin_lock_slowpath; - pv_ops_lock.queued_spin_unlock =3D - PV_CALLEE_SAVE(__pv_queued_spin_unlock); + static_call_update(queued_spin_lock_slowpath, __pv_queued_spin_lock_slowp= ath); + static_call_update(queued_spin_unlock, __raw_callee_save___pv_queued_spin= _unlock); pv_ops_lock.wait =3D kvm_wait; pv_ops_lock.kick =3D kvm_kick_cpu; =20 diff --git a/arch/x86/kernel/paravirt-spinlocks.c b/arch/x86/kernel/paravir= t-spinlocks.c index 9545244..ddc19dc 100644 --- a/arch/x86/kernel/paravirt-spinlocks.c +++ b/arch/x86/kernel/paravirt-spinlocks.c @@ -25,9 +25,14 @@ __visible void __native_queued_spin_unlock(struct qspinl= ock *lock) } PV_CALLEE_SAVE_REGS_THUNK(__native_queued_spin_unlock); =20 +DEFINE_STATIC_CALL(queued_spin_lock_slowpath, native_queued_spin_lock_slow= path); +EXPORT_STATIC_CALL_TRAMP(queued_spin_lock_slowpath); +DEFINE_STATIC_CALL(queued_spin_unlock, __raw_callee_save___native_queued_s= pin_unlock); +EXPORT_STATIC_CALL_TRAMP(queued_spin_unlock); + bool pv_is_native_spin_unlock(void) { - return pv_ops_lock.queued_spin_unlock.func =3D=3D + return static_call_query(queued_spin_unlock) =3D=3D __raw_callee_save___native_queued_spin_unlock; } =20 @@ -45,16 +50,11 @@ bool pv_is_native_vcpu_is_preempted(void) =20 void __init paravirt_set_cap(void) { - if (!pv_is_native_spin_unlock()) - setup_force_cpu_cap(X86_FEATURE_PVUNLOCK); - if (!pv_is_native_vcpu_is_preempted()) setup_force_cpu_cap(X86_FEATURE_VCPUPREEMPT); } =20 struct pv_lock_ops pv_ops_lock =3D { - .queued_spin_lock_slowpath =3D native_queued_spin_lock_slowpath, - .queued_spin_unlock =3D PV_CALLEE_SAVE(__native_queued_spin_unlock), .wait =3D paravirt_nop, .kick =3D paravirt_nop, .vcpu_is_preempted =3D PV_CALLEE_SAVE(__native_vcpu_is_preempted), diff --git a/arch/x86/kernel/static_call.c b/arch/x86/kernel/static_call.c index 61592e4..bab9406 100644 --- a/arch/x86/kernel/static_call.c +++ b/arch/x86/kernel/static_call.c @@ -4,6 +4,12 @@ #include #include =20 +/* Declared locally to avoid pulling asm/paravirt-spinlock.h header. */ +#ifdef CONFIG_PARAVIRT_SPINLOCKS +struct qspinlock; +void __raw_callee_save___native_queued_spin_unlock(struct qspinlock *lock); +#endif + enum insn_type { CALL =3D 0, /* site call */ NOP =3D 1, /* site cond-call */ @@ -31,6 +37,17 @@ static const u8 retinsn[] =3D { RET_INSN_OPCODE, 0xcc, 0= xcc, 0xcc, 0xcc }; */ static const u8 warninsn[] =3D { 0x67, 0x48, 0x0f, 0xb9, 0x3a }; =20 +#ifdef CONFIG_PARAVIRT_SPINLOCKS +/* + * ds ds movb $0, (_ASM_ARG1) + */ +#ifdef CONFIG_64BIT +static const u8 unlockinsn[] =3D { 0x3e, 0x3e, 0xc6, 0x07, 0x00 }; +#else +static const u8 unlockinsn[] =3D { 0x3e, 0x3e, 0xc6, 0x00, 0x00 }; +#endif +#endif + static u8 __is_Jcc(u8 *insn) /* Jcc.d32 */ { u8 ret =3D 0; @@ -78,6 +95,12 @@ static void __ref __static_call_transform(void *insn, en= um insn_type type, emulate =3D code; code =3D &warninsn; } +#ifdef CONFIG_PARAVIRT_SPINLOCKS + if (func =3D=3D &__raw_callee_save___native_queued_spin_unlock) { + emulate =3D code; + code =3D &unlockinsn; + } +#endif break; =20 case NOP: @@ -139,6 +162,10 @@ static void __static_call_validate(u8 *insn, bool tail= , bool tramp) !memcmp(insn, xor5rax, 5) || !memcmp(insn, warninsn, 5)) return; +#ifdef CONFIG_PARAVIRT_SPINLOCKS + if (!memcmp(insn, unlockinsn, 5)) + return; +#endif } =20 /* diff --git a/arch/x86/xen/spinlock.c b/arch/x86/xen/spinlock.c index 83ac24e..f718e53 100644 --- a/arch/x86/xen/spinlock.c +++ b/arch/x86/xen/spinlock.c @@ -134,9 +134,8 @@ void __init xen_init_spinlocks(void) printk(KERN_DEBUG "xen: PV spinlocks enabled\n"); =20 __pv_init_lock_hash(); - pv_ops_lock.queued_spin_lock_slowpath =3D __pv_queued_spin_lock_slowpath; - pv_ops_lock.queued_spin_unlock =3D - PV_CALLEE_SAVE(__pv_queued_spin_unlock); + static_call_update(queued_spin_lock_slowpath, __pv_queued_spin_lock_slowp= ath); + static_call_update(queued_spin_unlock, __raw_callee_save___pv_queued_spin= _unlock); pv_ops_lock.wait =3D xen_qlock_wait; pv_ops_lock.kick =3D xen_qlock_kick; pv_ops_lock.vcpu_is_preempted =3D PV_CALLEE_SAVE(xen_vcpu_stolen); diff --git a/tools/arch/x86/include/asm/cpufeatures.h b/tools/arch/x86/incl= ude/asm/cpufeatures.h index 86d17b1..425452f 100644 --- a/tools/arch/x86/include/asm/cpufeatures.h +++ b/tools/arch/x86/include/asm/cpufeatures.h @@ -225,7 +225,7 @@ #define X86_FEATURE_EPT_AD ( 8*32+17) /* "ept_ad" Intel Extended Page Tab= le access-dirty bit */ #define X86_FEATURE_VMCALL ( 8*32+18) /* Hypervisor supports the VMCALL i= nstruction */ #define X86_FEATURE_VMW_VMMCALL ( 8*32+19) /* VMware prefers VMMCALL hype= rcall instruction */ -#define X86_FEATURE_PVUNLOCK ( 8*32+20) /* PV unlock function */ +// free: was #define X86_FEATURE_PVUNLOCK ( 8*32+20) /* PV u= nlock function */ #define X86_FEATURE_VCPUPREEMPT ( 8*32+21) /* PV vcpu_is_preempted functi= on */ #define X86_FEATURE_TDX_GUEST ( 8*32+22) /* "tdx_guest" Intel Trust Domai= n Extensions Guest */ =20