From nobody Sat Jul 25 02:12:28 2026 Received: from galois.linutronix.de (Galois.linutronix.de [193.142.43.55]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A1C271A9FAF; Mon, 20 Jul 2026 18:43:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=193.142.43.55 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784573026; cv=none; b=TpTuKtxCKQiOV9X31zKyR2N07CEK3Bl09bdlnwVGh6pIj1PAwo8ufpTOxTQJM3z1lnmrzupDfWA44JBgMF7RmUdrYnzvEoIsDsucUl7V6QkBPnHhXdzHTDCRKMra6dkAdWKkTwu6OL02ZdWeH9ekXWve3vb1l83AwKXJXFan2eA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784573026; c=relaxed/simple; bh=3MaqStR0uk+DXRLjc+AH/+Bx2T0zEHtDZz1yeA+yKV4=; h=Date:From:To:Subject:Cc:In-Reply-To:References:MIME-Version: Message-ID:Content-Type; b=R8Ku6fJ0fruFOH8u602D/rx6RcmdtB0gtiGtgQajilGVky1XGu1jj02LAcccB2qsVJJDW6Y2PcjWoAI/gejqxlU7OAar/FejAumkcMjtiH/3QQUdVQngn3GiGB39Xju1JdQPLd6FvnZVkEJlWrRDNFnNdn46FLTXnaxCNDMfdyI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de; spf=pass smtp.mailfrom=linutronix.de; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=Au5OEXH1; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=Xg2UM3jC; arc=none smtp.client-ip=193.142.43.55 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linutronix.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="Au5OEXH1"; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="Xg2UM3jC" Date: Mon, 20 Jul 2026 18:43:35 -0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020; t=1784573017; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=6EquUAiD18xeulPnKIAYeNCfxnlUk05aRn0CjsRExqQ=; b=Au5OEXH1KW+eCXfTvNAC+osdxwwMdSjeyC+5CVCh/aw01+xK2j0VpntpPVLNYjVrPM0lh4 +/bhhikOPoWcaDwTXvhqG9u5n8kfMTOSAIOaqvL7Wi7A0ktVQZNfdgOzlgi3aG46Lfpr7Q wXKxvY6bsLve49qmOScMSCynGQHQa9EbWa518prKTKmxXqWkzSvBK/hBqHkhQhOfp7xYC1 XQ5f7ffZ1BDcaATrLirwVb4LoNfF2wVmyd8Bp6Yo5yexQWo9ZgkLUsPCVRTF1hPGtIX9hy 7ZJUfVBJcLNAvUskC8CFJQoJ9K70uM4DA4KT6yp88NXNEpbr6lzLeXvFfuHIaQ== DKIM-Signature: v=1; a=ed25519-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020e; t=1784573017; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=6EquUAiD18xeulPnKIAYeNCfxnlUk05aRn0CjsRExqQ=; b=Xg2UM3jCiuZZQbzCtRn1QP9zBWQ3qx9d5dWgoK/2lAjaC9DmNSSzdk6LNFLvZtqOuXl71j otabgO5SOj8KASDg== From: "tip-bot2 for Thomas Gleixner" Sender: tip-bot2@linutronix.de Reply-to: linux-kernel@vger.kernel.org To: linux-tip-commits@vger.kernel.org Subject: [tip: core/entry] entry, treewide: Make syscall_enter_from_user_mode[_work]() indicate syscall execution Cc: msuchanek@suse.de, Thomas Gleixner , x86@kernel.org, linux-kernel@vger.kernel.org In-Reply-To: <20260712141346.772209074@kernel.org> References: <20260712141346.772209074@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Message-ID: <178457301568.2943223.1730079655708381522.tip-bot2@tip-bot2> Robot-ID: Robot-Unsubscribe: Contact to get blacklisted from these emails Precedence: bulk Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable The following commit has been merged into the core/entry branch of tip: Commit-ID: 05c033db7e9ad3c34f6968ec568cb6ee01051c57 Gitweb: https://git.kernel.org/tip/05c033db7e9ad3c34f6968ec568cb6ee0= 1051c57 Author: Thomas Gleixner AuthorDate: Sun, 12 Jul 2026 23:25:32 +02:00 Committer: Thomas Gleixner CommitterDate: Mon, 20 Jul 2026 20:38:40 +02:00 entry, treewide: Make syscall_enter_from_user_mode[_work]() indicate syscal= l execution The return values of syscall_enter_from_user_mode[_work]() are non-intuitive. Both functions return the syscall number which should be invoked by the architecture specific syscall entry code. The returned number can be: - the unmodified syscall number which was handed in by the caller - a modified syscall number (ptrace, seccomp, trace/probe/bpf) That has an additional twist. If the return value is -1L then the caller is not allowed to modify the return value as that indicates that the modifying entity requests to abort the syscall and set the return value already. That can obviously not be differentiated from a syscall which handed in -1 as syscall number. The most trivial way to deal with that is: set_return_value(regs, -ENOSYS); nr =3D syscall_enter_from_user_mode(regs, nr); if (valid(nr)) handle_syscall(regs, nr); That's what LOONGARCH, RISCV, and X86 do. But PowerPC and S390 do not preset the return value, so when user space hands in -1 and there is nothing setting the return value in the entry work code, then the syscall is skipped but the return value is whatever random data has been in the return value register. Change the return values of syscall_enter_from_user_mode[_work]() to boolean and return false, when either ptrace or seccomp request to skip the syscall. If they return true, update the syscall number as it might have been changed. That results in slightly different behaviour of the architectures versus tracing. If the syscall tracepoint has probe/BPF attached, those might set the syscall number to -1 and also set the return value. PowerPC and S390 will then overwrite that value with -ENOSYS. The other architectures will just ignore it like any other invalid syscall and use the modified one. Originally-by: Michal Such=C3=A1nek Signed-off-by: Thomas Gleixner Tested-by: Michal Such=C3=A1nek Link: https://patch.msgid.link/20260712141346.772209074@kernel.org --- Documentation/core-api/entry.rst | 45 ++++++++++++++++++++++++------- arch/loongarch/kernel/syscall.c | 14 +++++----- arch/powerpc/kernel/syscall.c | 3 +- arch/riscv/kernel/traps.c | 11 +++----- arch/s390/kernel/syscall.c | 7 +++-- arch/x86/entry/syscall_32.c | 25 ++++++++--------- arch/x86/entry/syscall_64.c | 12 ++++---- include/linux/entry-common.h | 32 +++++++++++----------- 8 files changed, 88 insertions(+), 61 deletions(-) diff --git a/Documentation/core-api/entry.rst b/Documentation/core-api/entr= y.rst index 44e9608..79fdaed 100644 --- a/Documentation/core-api/entry.rst +++ b/Documentation/core-api/entry.rst @@ -58,26 +58,51 @@ state transitions must run with interrupts disabled. Syscalls -------- =20 -Syscall-entry code starts in assembly code and calls out into low-level C = code -after establishing low-level architecture-specific state and stack frames.= This -low-level C code must not be instrumented. A typical syscall handling func= tion -invoked from low-level assembly code looks like this: +Syscall-entry code starts in assembly code and calls out into low-level C +code after establishing low-level architecture-specific state and stack +frames. This low-level C code must not be instrumented. The recommended +syscall handling function invoked from low-level assembly code looks like +this: =20 .. code-block:: c =20 - noinstr void syscall(struct pt_regs *regs, int nr) + noinstr void syscall(struct pt_regs *regs, long nr) { arch_syscall_enter(regs); - nr =3D syscall_enter_from_user_mode_randomize_stack(regs, nr); + result_reg(regs) =3D -ENOSYS; + if (syscall_enter_from_user_mode_randomize_stack(regs, &nr)) { + instrumentation_begin(); + if (valid(nr) + result_reg(regs) =3D invoke_syscall(regs, nr); + instrumentation_end(); + } + syscall_exit_to_user_mode(regs); + } =20 - instrumentation_begin(); - if (!invoke_syscall(regs, nr) && nr !=3D -1) - result_reg(regs) =3D __sys_ni_syscall(regs); - instrumentation_end(); +This is the most resilent variant as it has always a guaranteed valid +return code. The alternative variant is: =20 +.. code-block:: c + + noinstr void syscall(struct pt_regs *regs, long nr) + { + arch_syscall_enter(regs); + if (syscall_enter_from_user_mode_randomize_stack(regs, &nr)) { + instrumentation_begin(); + if (valid(nr) + result_reg(regs) =3D invoke_syscall(regs, nr); + else + result_reg(regs) =3D -ENOSYS; + instrumentation_end(); + } syscall_exit_to_user_mode(regs); } =20 +That works for most situations except when a probe/BPF attached to the +syscall tracepoint sets an invalid syscall number e.g. -1 and also modifies +the result register. So this variant will obviously overwrite the modified +result with -ENOSYS. + syscall_enter_from_user_mode_randomize_stack() first invokes enter_from_user_mode_randomize_stack() which establishes state in the following order: diff --git a/arch/loongarch/kernel/syscall.c b/arch/loongarch/kernel/syscal= l.c index 142a9c2..62264e4 100644 --- a/arch/loongarch/kernel/syscall.c +++ b/arch/loongarch/kernel/syscall.c @@ -57,8 +57,8 @@ typedef long (*sys_call_fn)(unsigned long, unsigned long, =20 void noinstr __no_stack_protector do_syscall(struct pt_regs *regs) { - unsigned long nr; sys_call_fn syscall_fn; + unsigned long nr; =20 nr =3D regs->regs[11]; /* Set for syscall restarting */ @@ -69,12 +69,12 @@ void noinstr __no_stack_protector do_syscall(struct pt_= regs *regs) regs->orig_a0 =3D regs->regs[4]; regs->regs[4] =3D -ENOSYS; =20 - nr =3D syscall_enter_from_user_mode_randomize_stack(regs, nr); - - if (nr < NR_syscalls) { - syscall_fn =3D sys_call_table[array_index_nospec(nr, NR_syscalls)]; - regs->regs[4] =3D syscall_fn(regs->orig_a0, regs->regs[5], regs->regs[6], - regs->regs[7], regs->regs[8], regs->regs[9]); + if (likely(syscall_enter_from_user_mode_randomize_stack(regs, &nr))) { + if (nr < NR_syscalls) { + syscall_fn =3D sys_call_table[array_index_nospec(nr, NR_syscalls)]; + regs->regs[4] =3D syscall_fn(regs->orig_a0, regs->regs[5], regs->regs[6= ], + regs->regs[7], regs->regs[8], regs->regs[9]); + } } =20 syscall_exit_to_user_mode(regs); diff --git a/arch/powerpc/kernel/syscall.c b/arch/powerpc/kernel/syscall.c index b4d0159..1440dca 100644 --- a/arch/powerpc/kernel/syscall.c +++ b/arch/powerpc/kernel/syscall.c @@ -18,7 +18,8 @@ notrace long system_call_exception(struct pt_regs *regs, = unsigned long r0) long ret; syscall_fn f; =20 - r0 =3D syscall_enter_from_user_mode_randomize_stack(regs, r0); + if (unlikely(!syscall_enter_from_user_mode_randomize_stack(regs, &r0))) + return syscall_get_error(current, regs); =20 if (unlikely(r0 >=3D NR_syscalls)) { if (unlikely(trap_is_unsupported_scv(regs))) { diff --git a/arch/riscv/kernel/traps.c b/arch/riscv/kernel/traps.c index bb18011..2f57fd4 100644 --- a/arch/riscv/kernel/traps.c +++ b/arch/riscv/kernel/traps.c @@ -332,13 +332,12 @@ void do_trap_ecall_u(struct pt_regs *regs) =20 riscv_v_vstate_discard(regs); =20 - syscall =3D syscall_enter_from_user_mode_randomize_stack(regs, syscall); - - if (syscall >=3D 0 && syscall < NR_syscalls) { - syscall =3D array_index_nospec(syscall, NR_syscalls); - syscall_handler(regs, syscall); + if (likely(syscall_enter_from_user_mode_randomize_stack(regs, &syscall))= ) { + if (syscall >=3D 0 && syscall < NR_syscalls) { + syscall =3D array_index_nospec(syscall, NR_syscalls); + syscall_handler(regs, syscall); + } } - syscall_exit_to_user_mode(regs); } else { irqentry_state_t state =3D irqentry_nmi_enter(regs); diff --git a/arch/s390/kernel/syscall.c b/arch/s390/kernel/syscall.c index 73abda2..4ac80e2 100644 --- a/arch/s390/kernel/syscall.c +++ b/arch/s390/kernel/syscall.c @@ -96,6 +96,7 @@ SYSCALL_DEFINE0(ni_syscall) void noinstr __do_syscall(struct pt_regs *regs, int per_trap) { unsigned long nr; + bool permit; =20 enter_from_user_mode_randomize_stack(regs); =20 @@ -121,7 +122,9 @@ void noinstr __do_syscall(struct pt_regs *regs, int per= _trap) regs->psw.addr =3D current->restart_block.arch_data; current->restart_block.arch_data =3D 1; } - nr =3D syscall_enter_from_user_mode_work(regs, nr); + + permit =3D syscall_enter_from_user_mode_work(regs, &nr); + /* * In the s390 ptrace ABI, both the syscall number and the return value * use gpr2. However, userspace puts the syscall number either in the @@ -129,7 +132,7 @@ void noinstr __do_syscall(struct pt_regs *regs, int per= _trap) * work, the ptrace code sets PIF_SYSCALL_RET_SET, which is checked here * and if set, the syscall will be skipped. */ - if (unlikely(test_and_clear_pt_regs_flag(regs, PIF_SYSCALL_RET_SET))) + if (unlikely(test_and_clear_pt_regs_flag(regs, PIF_SYSCALL_RET_SET) || !p= ermit)) goto out; regs->gprs[2] =3D -ENOSYS; if (likely(nr < NR_syscalls)) { diff --git a/arch/x86/entry/syscall_32.c b/arch/x86/entry/syscall_32.c index 92369d8..91123a9 100644 --- a/arch/x86/entry/syscall_32.c +++ b/arch/x86/entry/syscall_32.c @@ -161,8 +161,9 @@ __visible noinstr void do_int80_emulation(struct pt_reg= s *regs) nr =3D syscall_32_enter(regs); =20 local_irq_enable(); - nr =3D syscall_enter_from_user_mode_work(regs, nr); - do_syscall_32_irqs_on(regs, nr); + + if (likely(syscall_enter_from_user_mode_work(regs, &nr))) + do_syscall_32_irqs_on(regs, nr); =20 instrumentation_end(); syscall_exit_to_user_mode(regs); @@ -223,8 +224,8 @@ DEFINE_FREDENTRY_RAW(int80_emulation) nr =3D syscall_32_enter(regs); =20 local_irq_enable(); - nr =3D syscall_enter_from_user_mode_work(regs, nr); - do_syscall_32_irqs_on(regs, nr); + if (likely(syscall_enter_from_user_mode_work(regs, &nr))) + do_syscall_32_irqs_on(regs, nr); =20 instrumentation_end(); syscall_exit_to_user_mode(regs); @@ -243,13 +244,13 @@ __visible noinstr void do_int80_syscall_32(struct pt_= regs *regs) * orig_ax, the int return value truncates it. This matches * the semantics of syscall_get_nr(). */ - nr =3D syscall_enter_from_user_mode_randomize_stack(regs, nr); - - instrumentation_begin(); + if (likely(syscall_enter_from_user_mode_randomize_stack(regs, &nr))) { + instrumentation_begin(); =20 - do_syscall_32_irqs_on(regs, nr); + do_syscall_32_irqs_on(regs, nr); =20 - instrumentation_end(); + instrumentation_end(); + } syscall_exit_to_user_mode(regs); } #endif /* !CONFIG_IA32_EMULATION */ @@ -286,10 +287,8 @@ static noinstr bool __do_fast_syscall_32(struct pt_reg= s *regs) return false; } =20 - nr =3D syscall_enter_from_user_mode_work(regs, nr); - - /* Now this is just like a normal syscall. */ - do_syscall_32_irqs_on(regs, nr); + if (likely(syscall_enter_from_user_mode_work(regs, &nr))) + do_syscall_32_irqs_on(regs, nr); =20 instrumentation_end(); syscall_exit_to_user_mode(regs); diff --git a/arch/x86/entry/syscall_64.c b/arch/x86/entry/syscall_64.c index 9a2762e..966d3d1 100644 --- a/arch/x86/entry/syscall_64.c +++ b/arch/x86/entry/syscall_64.c @@ -78,14 +78,14 @@ static __always_inline void do_syscall_x32(struct pt_re= gs *regs, unsigned long n /* Returns true to return using SYSRET, or false to use IRET */ __visible noinstr bool do_syscall_64(struct pt_regs *regs, long nr) { - nr =3D syscall_enter_from_user_mode_randomize_stack(regs, nr); + if (likely(syscall_enter_from_user_mode_randomize_stack(regs, &nr))) { + instrumentation_begin(); =20 - instrumentation_begin(); + if (!do_syscall_x64(regs, nr)) + do_syscall_x32(regs, nr); =20 - if (!do_syscall_x64(regs, nr)) - do_syscall_x32(regs, nr); - - instrumentation_end(); + instrumentation_end(); + } syscall_exit_to_user_mode(regs); =20 /* diff --git a/include/linux/entry-common.h b/include/linux/entry-common.h index df221d9..6574b71 100644 --- a/include/linux/entry-common.h +++ b/include/linux/entry-common.h @@ -114,16 +114,15 @@ static __always_inline long syscall_trace_enter(struc= t pt_regs *regs, unsigned l * @regs: Pointer to currents pt_regs * @syscall: The syscall number * - * Invoked from architecture specific syscall entry code with interrupts - * enabled after invoking enter_from_user_mode(), enabling interrupts and - * extra architecture specific work. + * Invoked from architecture specific syscall entry code with interrupts e= nabled + * after invoking enter_from_user_mode(), enabling interrupts and extra + * architecture specific work with the syscall return value preset to -ENO= SYS. * - * Returns: The original or a modified syscall number + * Returns: True if the syscall should be invoked, False otherwise. * - * If the returned syscall number is -1 then the syscall should be - * skipped. In this case the caller may invoke syscall_set_error() or - * syscall_set_return_value() first. If neither of those are called and -1 - * is returned, then the syscall will fail with ENOSYS. + * If the return value is false, the caller must skip the syscall and leav= e the + * syscall return value unmodified as it might have been set by one of the= entry + * work functions. * * It handles the following work items: * @@ -131,19 +130,20 @@ static __always_inline long syscall_trace_enter(struc= t pt_regs *regs, unsigned l * ptrace_report_syscall_permit_entry(), __seccomp_permit_syscall(), t= race_sys_enter() * 2) Invocation of audit_syscall_entry() */ -static __always_inline long syscall_enter_from_user_mode_work(struct pt_re= gs *regs, long syscall) +static __always_inline bool syscall_enter_from_user_mode_work(struct pt_re= gs *regs, long *syscall) { unsigned long work =3D READ_ONCE(current_thread_info()->syscall_work); =20 - if (work & SYSCALL_WORK_ENTER) { - if (!syscall_trace_enter(regs, work, syscall)) - return -1L; + if (!(work & SYSCALL_WORK_ENTER)) + return true; =20 - /* Reread the syscall number as it might have been modified */ - syscall =3D syscall_get_nr(current, regs); - } + if (unlikely(!syscall_trace_enter(regs, work, *syscall))) + return false; =20 - return syscall; + /* Reread the syscall number as it might have been modified */ + *syscall =3D syscall_get_nr(current, regs); + + return true; } =20 /**