From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oo2-f12.google.com (mail-oo2-f12.google.com [74.125.231.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 61EFB471437 for ; Fri, 11 Sep 2026 15:41:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141324; cv=none; b=ugMPY5FxtkpXgQObyHg4SaOmX9C4yHgyi1mgmkYs+dcQP6gcRWKkLrTkH9a1C63t1V2Z7GOlhT5IcSdOakMGDxQPniLtoku+JdTwlWgYW7+osbuJK9ECQsGVrKydjMkKcsTcbZCj6HOAyXofp6N+70Anyo+SLPz1VeNmCUS86YQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141324; c=relaxed/simple; bh=r/qsE4zGgo7MhImrb2TlshD5UTHqgg274D7HoUGeQkY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=JgIDl423/TDpsApZT65hTiVDU7O9fHYdtkfdAFTgsM0XTSdlZeiqyTJGkslqVecx7iz6kVMAJ8kKTUtgWKJ3wt0QWTHNMo9AOv4XZlVCAP1ZmeKmZCdnp2HIHjUyi9BDG3HQxzi5uR9D5OakH4TL0xUifwYzh/jWKpI5rdBgdvk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=0r1hjtZ6; arc=none smtp.client-ip=74.125.231.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="0r1hjtZ6" Received: by mail-oo2-f12.google.com with SMTP id 006d021491bc7-6b1ae6f5fc5so608078eaf.0 for ; Fri, 11 Sep 2026 08:41:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141317; x=1789746117; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=LQGIaUxtlzAP3sjiLv3UoTOfP6jJ6w5b9DVleKCMx9s=; b=0r1hjtZ6LSHw1RdShsmODshx0ksFvZNxUF/S/zxsGU5PnKiin3G9XIlhpZqd5gSDZu ShnTZ/wUoRsKU0CtPxIWtuj1ShiG1ZV84OPjyXpS2H91Sxir75P+cRQSdgS7bpD7X3Sg lJzF8wgdT5JEjCgIMJ+NGnbobBY6HvsiFqsiBbkDJUmtBr9Q90g2CX/sLcf9L5MoEOtn yvTwnwx3zed0dFGWWT0U7u5uRwOJtonLFiNKvxK4JEVtSv0TX0bRtu5UBEC6Md/Dh3HT swBqbw+1uEa5WNZo0VJ5L4dra31JKBJG/tEk1zJiYQXpZKODHdkdGGivCRbgenkOm4VU a3zQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141317; x=1789746117; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=LQGIaUxtlzAP3sjiLv3UoTOfP6jJ6w5b9DVleKCMx9s=; b=HKsqj55HmESPFFgsfNOEmhmtX3wyShKeUQxllfSeSI6i5acwDiQgM4s7O3+Dd8DY6t v2WwmgSCgUys79ClGoo+kqiXL+3nHY6RdqucbVkIoV3quzCqGJmqCFSh70qoI+pNc0Rr CITS0HA9i1JP2EvM36nvYVuauAT68mpWjhvopIU9Qy7Mt9ZX80g5WIO0d3L6I4FNoS8T 1Q+hFLuzxlDJZwGhE6q6DJZ7EiUTBPeYx3zvZZ3U/E+BMFC9s+e0zDEcabsc7VvQbRbq LfM2g0WxYfPYhNCtBWMLkdJJWqGf5kJj8LKkNUCkhMqpNJeVvW5yjPMrt7IHfj7SSSE3 4Yig== X-Forwarded-Encrypted: i=1; AKwUvBzj+QUACSs5EOVRm1Vuv6sS2AMwocrEuKqYFiTUQYi52sHzI2xEi+40QExnN2x7bUhZq4rGuwitdwRC7GM=@vger.kernel.org X-Gm-Message-State: AFuF++nf02LLq6k6NiwAwSlQF7MpTpZmCpt5iXP6RKAtcDA7RxQsFEGm 4kQsOIbkPdDtW9fXpjNoEnnlp8vekF0EApZ338EsWB7HZrO6zKQtS3e570Kw9lhU94wC6nfKK7D 5E6/cV0U= X-Gm-Gg: AYBFou0XMsmjPCT47WpynODtDiW4bhvr1MLy2Z/YVe//yWxuRh11H2mBl4Lsef+w5T9 MSBmgHD/88/kPIjoshud2ltaQI9s1IugEqBxIS/2BQbvbZTRehr/BYxvPsEmMrBqWD1ZqaG0lHf oach0YPDh5cSYhKP7Atj3RwjEoHbSDjxiEhlMFWWcLLXCAd/sEZdk7qk1m9aUy2pN6o4N/YC1mW 6scMY9UwuH9LJl21vbmR8tm8s3gQ+0tlEWYXph8RgUE1XzgUN50CbtgRoyeImDEBY2fqQU97irF cgIvQvfDG7WIItmNfA/6MsJTt3wVfVp0HL3Znb3FCB8RFWbHxQ36dUUIDKXKRHoRo5IzCQM4jJJ cDQmWYSYUHkvxMPd/ZJXOjS1bLDwPytUvGJxCQ0CN9jofSGT1lL/1b33DD3yoA3rHYZ3peg2Q7O oAv6vLDhek7a4c8Yy0ON77uYMVsXwgttShLgURgqobIcApyv7jWHS31lbcpDX8sh4x8xRYkCyld LI8mvmAV59jtLSRK8ZXWfN0ZvJah/5c1YABtyzS0Og= X-Received: by 2002:a05:6820:1f07:b0:6b1:b867:1a1 with SMTP id 006d021491bc7-6c0bb2f3cadmr6386168eaf.12.1789141316858; Fri, 11 Sep 2026 08:41:56 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.41.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:41:56 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 01/15] kernel: add thread identity handoff Date: Fri, 11 Sep 2026 09:40:51 -0600 Message-ID: <20260911154148.644489-2-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add infrastructure for moving the user visible identity of a thread that is about to block in the kernel to another thread in the same process, which then returns to userspace on its behalf. The blocked thread keeps its task_struct and kernel stack, only what userspace observes moves: tid, signal state, rseq and robust list, credentials, scheduling attributes, mempolicy, comm and the user register state. Thread group leadership follows the tid. thread_handoff_allowed() vets the source and thread_handoff_compatible() the destination. thread_handoff_prepare() runs on the source right before it blocks, thread_handoff_finish() on the destination moves the identity over. Architectures provide the arch_thread_handoff_*() hooks and select ARCH_HAS_THREAD_HANDOFF. Signed-off-by: Jens Axboe --- arch/Kconfig | 7 + include/linux/thread_handoff.h | 72 +++++ init/Kconfig | 11 + kernel/Makefile | 1 + kernel/sched/core.c | 33 +++ kernel/thread_handoff.c | 484 +++++++++++++++++++++++++++++++++ 6 files changed, 608 insertions(+) create mode 100644 include/linux/thread_handoff.h create mode 100644 kernel/thread_handoff.c diff --git a/arch/Kconfig b/arch/Kconfig index 45c657772362..5ce7e1713f59 100644 --- a/arch/Kconfig +++ b/arch/Kconfig @@ -604,6 +604,13 @@ config ARCH_HAVE_EXTRA_ELF_NOTES config ARCH_HAS_NMI_SAFE_THIS_CPU_OPS bool =20 +config ARCH_HAS_THREAD_HANDOFF + bool + help + Architecture provides the arch_thread_handoff_*() hooks for moving + the user visible register state of a thread that is blocked in the + kernel to another thread of the same process. + config HAVE_ALIGNED_STRUCT_PAGE bool help diff --git a/include/linux/thread_handoff.h b/include/linux/thread_handoff.h new file mode 100644 index 000000000000..e1c17b833e7c --- /dev/null +++ b/include/linux/thread_handoff.h @@ -0,0 +1,72 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef _LINUX_THREAD_HANDOFF_H +#define _LINUX_THREAD_HANDOFF_H + +#include + +/* per-thread accounting that follows the identity */ +struct thread_handoff_stats { + u64 utime; + u64 stime; + u64 gtime; + u64 exec_runtime; + unsigned long min_flt; + unsigned long maj_flt; + unsigned long nvcsw; + unsigned long nivcsw; + struct task_io_accounting ioac; +}; + +/* + * Move the user visible identity of a thread that is about to block in the + * kernel (tid, signals, registers, rseq, creds, sched attributes, ...) to + * another thread in the same group, which then returns to userspace on its + * behalf. The blocked thread keeps its task_struct and in-kernel state. + */ +#ifdef CONFIG_THREAD_HANDOFF +bool thread_handoff_allowed(struct task_struct *tsk); +bool thread_handoff_compatible(struct task_struct *src, + struct task_struct *dst); +bool thread_handoff_prepare(struct task_struct *tsk); +void thread_handoff_stats_take(struct thread_handoff_stats *st); +int thread_handoff_finish(struct task_struct *src, + struct thread_handoff_stats *st); + +u64 sched_exec_runtime_take(struct task_struct *p); +void sched_exec_runtime_add(struct task_struct *p, u64 ns); + +/* + * Arch hooks. _prepare() syncs the live user register state on the source + * before it blocks, _finish() copies it over and loads what the return to + * userspace won't. @leader tells it that thread group leadership moved. + */ +bool arch_thread_handoff_allowed(struct task_struct *tsk); +bool arch_thread_handoff_compatible(struct task_struct *src, + struct task_struct *dst); +bool arch_thread_handoff_prepare(void); +int arch_thread_handoff_finish(struct task_struct *src, bool leader); +#else +static inline bool thread_handoff_allowed(struct task_struct *tsk) +{ + return false; +} +static inline bool thread_handoff_compatible(struct task_struct *src, + struct task_struct *dst) +{ + return false; +} +static inline bool thread_handoff_prepare(struct task_struct *tsk) +{ + return false; +} +static inline void thread_handoff_stats_take(struct thread_handoff_stats *= st) +{ +} +static inline int thread_handoff_finish(struct task_struct *src, + struct thread_handoff_stats *st) +{ + return -EOPNOTSUPP; +} +#endif + +#endif diff --git a/init/Kconfig b/init/Kconfig index 8583d9f06c52..b376f802a77a 100644 --- a/init/Kconfig +++ b/init/Kconfig @@ -1957,9 +1957,20 @@ config AIO by some high performance threaded applications. Disabling this option saves about 7k. =20 +config THREAD_HANDOFF + bool + depends on ARCH_HAS_THREAD_HANDOFF + help + Support for handing the user visible identity of a thread that is + about to block in the kernel to another thread of the same process, + which then returns to userspace on its behalf. Used by io_uring to + issue requests inline and only offload them to a worker thread if + they actually block. + config IO_URING bool "Enable IO uring support" if EXPERT select IO_WQ + select THREAD_HANDOFF if ARCH_HAS_THREAD_HANDOFF default y help This option enables support for the io_uring interface, enabling diff --git a/kernel/Makefile b/kernel/Makefile index 1e1a31673577..5843d4866db7 100644 --- a/kernel/Makefile +++ b/kernel/Makefile @@ -14,6 +14,7 @@ obj-y =3D fork.o exec_domain.o exec_state.o panic.o \ =20 obj-$(CONFIG_MULTIUSER) +=3D groups.o obj-$(CONFIG_VHOST_TASK) +=3D vhost_task.o +obj-$(CONFIG_THREAD_HANDOFF) +=3D thread_handoff.o =20 ifdef CONFIG_FUNCTION_TRACER # Do not trace internal ftrace files diff --git a/kernel/sched/core.c b/kernel/sched/core.c index b998ef6b87af..eeb55367c0c6 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -96,6 +96,7 @@ =20 #include "../workqueue_internal.h" #include "../../io_uring/io-wq.h" +#include #include "../smpboot.h" #include "../locking/mutex.h" =20 @@ -5684,6 +5685,38 @@ static inline void prefetch_curr_exec_start(struct t= ask_struct *p) * In case the task is currently running, return the runtime plus current's * pending runtime that have not been accounted yet. */ +#ifdef CONFIG_THREAD_HANDOFF +/* keep the prev_sum_exec_runtime delta intact, slice accounting uses it */ +u64 sched_exec_runtime_take(struct task_struct *p) +{ + struct rq_flags rf; + struct rq *rq; + u64 ns; + + rq =3D task_rq_lock(p, &rf); + if (task_current_donor(rq, p) && task_on_rq_queued(p)) { + update_rq_clock(rq); + p->sched_class->update_curr(rq); + } + ns =3D p->se.sum_exec_runtime; + p->se.sum_exec_runtime =3D 0; + p->se.prev_sum_exec_runtime -=3D ns; + task_rq_unlock(rq, p, &rf); + return ns; +} + +void sched_exec_runtime_add(struct task_struct *p, u64 ns) +{ + struct rq_flags rf; + struct rq *rq; + + rq =3D task_rq_lock(p, &rf); + p->se.sum_exec_runtime +=3D ns; + p->se.prev_sum_exec_runtime +=3D ns; + task_rq_unlock(rq, p, &rf); +} +#endif + unsigned long long task_sched_runtime(struct task_struct *p) { struct rq_flags rf; diff --git a/kernel/thread_handoff.c b/kernel/thread_handoff.c new file mode 100644 index 000000000000..1901eb85bae8 --- /dev/null +++ b/kernel/thread_handoff.c @@ -0,0 +1,484 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Thread identity handoff, see include/linux/thread_handoff.h + * + * Copyright (C) 2026 Jens Axboe + */ +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +/* prctl state that moves. Not PF_MEMALLOC_NOIO, kernel code sets that too= */ +#define THREAD_HANDOFF_PF_FLAGS (PF_MCE_PROCESS | PF_MCE_EARLY) + +/* can the identity of @tsk (current) be handed off, errs on the safe side= */ +bool thread_handoff_allowed(struct task_struct *tsk) +{ + WARN_ON_ONCE(tsk !=3D current); + + if (tsk->flags & (PF_EXITING | PF_KTHREAD | PF_IO_WORKER)) + return false; + if (tsk->ptrace) + return false; + /* tracees point back at the tracer task */ + if (!list_empty(&tsk->ptraced)) + return false; + if (tsk->signal->flags & SIGNAL_GROUP_EXIT) + return false; +#ifdef CONFIG_PERF_EVENTS + /* per-task perf contexts are bound to the task_struct */ + if (tsk->perf_event_ctxp) + return false; +#endif +#ifdef CONFIG_FUTEX + /* PI futex ownership is tied to the task_struct */ + if (!list_empty(&tsk->futex.pi_state_list)) + return false; +#endif +#ifdef CONFIG_GENERIC_ENTRY + if (test_syscall_work(SYSCALL_USER_DISPATCH)) + return false; +#endif + /* only the fair class moves, see thread_handoff_sched() */ + if (rt_or_dl_task(tsk)) + return false; +#ifdef CONFIG_SCHED_CORE + /* the core scheduling cookie is bound to the task */ + if (tsk->core_cookie) + return false; +#endif +#ifdef CONFIG_KCOV + /* coverage collection is per-thread */ + if (tsk->kcov) + return false; +#endif +#ifdef CONFIG_POSIX_TIMERS + /* armed per-thread CPU timers would sample the destination's clock */ + if (tsk->posix_cputimers.timers_active) + return false; +#endif + /* the syscall never exits, its audit record would never be emitted */ + if (!audit_dummy_context()) + return false; +#ifdef CONFIG_UPROBES + /* pending uretprobes, the return address bookkeeping is in our utask */ + if (tsk->utask && tsk->utask->return_instances) + return false; +#endif + /* a vfork() parent waits on this task_struct, not on the identity */ + if (tsk->vfork_done) + return false; + + return arch_thread_handoff_allowed(tsk); +} + +/* can @dst take over the identity of @src, called from sched_submit_work(= ) */ +bool thread_handoff_compatible(struct task_struct *src, struct task_struct= *dst) +{ + if (dst->flags & PF_EXITING) + return false; + /* the identity would land under a tracer that never attached to it */ + if (dst->ptrace) + return false; + if (!same_thread_group(src, dst)) + return false; +#ifdef CONFIG_GENERIC_ENTRY + /* inherited from the thread that forked the destination */ + if (test_task_syscall_work(dst, SYSCALL_USER_DISPATCH)) + return false; +#endif + if (dst->mm !=3D src->mm || dst->files !=3D src->files || + dst->fs !=3D src->fs || dst->nsproxy !=3D src->nsproxy) + return false; + /* the identity must not gain no_new_privs */ + if (task_no_new_privs(dst) && !task_no_new_privs(src)) + return false; +#ifdef CONFIG_SECCOMP + /* seccomp filters are per-thread, the identity must not escape them */ + if (dst->seccomp.mode !=3D src->seccomp.mode || + dst->seccomp.filter !=3D src->seccomp.filter) + return false; +#endif +#ifdef CONFIG_SYSVIPC + /* SEM_UNDO adjustments are accounted per undo list */ + if (dst->sysvsem.undo_list !=3D src->sysvsem.undo_list) + return false; +#endif + /* thread_handoff_creds() doesn't switch namespaces */ + scoped_guard(rcu) { + if (__task_cred(dst)->user_ns !=3D __task_cred(src)->user_ns) + return false; + } + return arch_thread_handoff_compatible(src, dst); +} + +/* runs on the source right before it blocks, syncs its live user state */ +bool thread_handoff_prepare(struct task_struct *tsk) +{ + WARN_ON_ONCE(tsk !=3D current); + /* re-check, may have changed since thread_handoff_allowed() */ + if (tsk->ptrace) + return false; +#ifdef CONFIG_PERF_EVENTS + if (tsk->perf_event_ctxp) + return false; +#endif +#ifdef CONFIG_FUTEX + if (!list_empty(&tsk->futex.pi_state_list)) + return false; +#endif + return arch_thread_handoff_prepare(); +} + +/* + * Take the source's per-thread accounting for the destination, so the tid= 's + * counters stay monotonic. Runs as current, the tick writes these from ir= q. + */ +void thread_handoff_stats_take(struct thread_handoff_stats *st) +{ + struct task_struct *p =3D current; + + st->exec_runtime =3D sched_exec_runtime_take(p); + + local_irq_disable(); + st->utime =3D p->utime; + st->stime =3D p->stime; + st->gtime =3D p->gtime; + p->utime =3D p->stime =3D p->gtime =3D 0; +#ifndef CONFIG_VIRT_CPU_ACCOUNTING_NATIVE + raw_spin_lock(&p->prev_cputime.lock); + p->prev_cputime.utime =3D p->prev_cputime.stime =3D 0; + raw_spin_unlock(&p->prev_cputime.lock); +#endif + local_irq_enable(); + + st->min_flt =3D p->min_flt; + st->maj_flt =3D p->maj_flt; + st->nvcsw =3D p->nvcsw; + st->nivcsw =3D p->nivcsw; + p->min_flt =3D p->maj_flt =3D p->nvcsw =3D p->nivcsw =3D 0; + st->ioac =3D p->ioac; + memset(&p->ioac, 0, sizeof(p->ioac)); +} + +static void thread_handoff_stats_add(struct task_struct *p, + struct thread_handoff_stats *st) +{ + sched_exec_runtime_add(p, st->exec_runtime); + + local_irq_disable(); + p->utime +=3D st->utime; + p->stime +=3D st->stime; + p->gtime +=3D st->gtime; + local_irq_enable(); + + p->min_flt +=3D st->min_flt; + p->maj_flt +=3D st->maj_flt; + p->nvcsw +=3D st->nvcsw; + p->nivcsw +=3D st->nivcsw; + task_io_accounting_add(&p->ioac, &st->ioac); +} + +/* signal state follows the identity, the source gets a worker's mask */ +static void thread_handoff_signals(struct task_struct *dst, + struct task_struct *src) + __must_hold(&dst->sighand->siglock) +{ + dst->blocked =3D src->blocked; + dst->real_blocked =3D src->real_blocked; + dst->saved_sigmask =3D src->saved_sigmask; + dst->sas_ss_sp =3D src->sas_ss_sp; + dst->sas_ss_size =3D src->sas_ss_size; + dst->sas_ss_flags =3D src->sas_ss_flags; + dst->restart_block =3D src->restart_block; + + siginitsetinv(&src->blocked, sigmask(SIGKILL) | sigmask(SIGSTOP)); + sigemptyset(&src->real_blocked); + sas_ss_reset(src); + + list_splice_tail_init(&src->pending.list, &dst->pending.list); + sigorsets(&dst->pending.signal, &dst->pending.signal, + &src->pending.signal); + sigemptyset(&src->pending.signal); +} + +/* the tgid is the leader's tid, so leadership follows. Like de_thread() */ +static void thread_handoff_leader(struct task_struct *dst, + struct task_struct *src) + __must_hold(&tasklist_lock) +{ + struct list_head *prev =3D dst->thread_node.prev; + struct task_struct *t; + + /* + * The leader must be first on ->thread_head, swap the two list + * positions. RCU readers may see a thread twice, never miss one. + */ + if (prev =3D=3D &src->thread_node) + prev =3D &dst->thread_node; + list_del_rcu(&dst->thread_node); + list_replace_rcu(&src->thread_node, &dst->thread_node); + list_add_rcu(&src->thread_node, prev); + + transfer_pid(src, dst, PIDTYPE_TGID); + transfer_pid(src, dst, PIDTYPE_PGID); + transfer_pid(src, dst, PIDTYPE_SID); + + list_replace_rcu(&src->tasks, &dst->tasks); + list_replace_init(&src->sibling, &dst->sibling); + + for_each_thread(dst, t) + t->group_leader =3D dst; + + dst->exit_signal =3D src->exit_signal; + src->exit_signal =3D -1; +} + +/* children are parented to the forking thread, move them along */ +static void thread_handoff_children(struct task_struct *dst, + struct task_struct *src) + __must_hold(&tasklist_lock) +{ + struct task_struct *p; + + list_for_each_entry(p, &src->children, sibling) { + RCU_INIT_POINTER(p->real_parent, dst); + if (rcu_access_pointer(p->parent) =3D=3D src) + RCU_INIT_POINTER(p->parent, dst); + } + list_splice_init(&src->children, &dst->children); + dst->self_exec_id =3D src->self_exec_id; +} + +static bool thread_handoff_sched_same(struct task_struct *dst, + struct task_struct *src) +{ + if (dst->policy !=3D src->policy || task_nice(dst) !=3D task_nice(src)) + return false; + if (dst->sched_reset_on_fork !=3D src->sched_reset_on_fork) + return false; + if (dst->se.custom_slice !=3D src->se.custom_slice || + (src->se.custom_slice && dst->se.slice !=3D src->se.slice)) + return false; +#ifdef CONFIG_UCLAMP_TASK + for (int i =3D 0; i < UCLAMP_CNT; i++) { + if (dst->uclamp_req[i].user_defined !=3D src->uclamp_req[i].user_defined) + return false; + if (src->uclamp_req[i].user_defined && + dst->uclamp_req[i].value !=3D src->uclamp_req[i].value) + return false; + } +#endif + return true; +} + +/* sched attributes follow the identity, sched_setattr() only if needed */ +static void thread_handoff_sched(struct task_struct *dst, + struct task_struct *src) +{ + struct sched_attr attr =3D { + .sched_policy =3D src->policy, + .sched_nice =3D task_nice(src), + }; + + if (thread_handoff_sched_same(dst, src)) + return; + if (src->sched_reset_on_fork) + attr.sched_flags |=3D SCHED_FLAG_RESET_ON_FORK; + if (src->se.custom_slice) + attr.sched_runtime =3D src->se.slice; +#ifdef CONFIG_UCLAMP_TASK + if (src->uclamp_req[UCLAMP_MIN].user_defined) { + attr.sched_flags |=3D SCHED_FLAG_UTIL_CLAMP_MIN; + attr.sched_util_min =3D src->uclamp_req[UCLAMP_MIN].value; + } + if (src->uclamp_req[UCLAMP_MAX].user_defined) { + attr.sched_flags |=3D SCHED_FLAG_UTIL_CLAMP_MAX; + attr.sched_util_max =3D src->uclamp_req[UCLAMP_MAX].value; + } +#endif + WARN_ON_ONCE(sched_setattr_nocheck(dst, &attr)); +} + +static void thread_handoff_mempolicy(struct task_struct *dst, + struct task_struct *src) +{ +#ifdef CONFIG_NUMA + struct mempolicy *pol =3D mpol_dup(src->mempolicy); + + if (IS_ERR(pol)) + return; + task_lock(dst); + swap(dst->mempolicy, pol); + task_unlock(dst); + mpol_put(pol); +#endif +} + +#ifdef CONFIG_BLOCK +/* ionice'd threads keep their IO priority, the source's ioc may be in use= */ +static void thread_handoff_ioprio(struct task_struct *dst, + struct task_struct *src) +{ + struct io_context *ioc =3D src->io_context; + + /* workers share the ioc of the thread that forked them, detach */ + if (dst->io_context) + exit_io_context(dst); + if (ioc && ioprio_valid(ioc->ioprio)) + WARN_ON_ONCE(set_task_ioprio(dst, ioc->ioprio)); +} +#else +static void thread_handoff_ioprio(struct task_struct *dst, + struct task_struct *src) +{ +} +#endif + +#ifdef CONFIG_CGROUPS +/* threaded cgroup placement follows the identity, the source stays put */ +static void thread_handoff_cgroup(struct task_struct *dst, + struct task_struct *src) +{ + if (rcu_access_pointer(src->cgroups) =3D=3D rcu_access_pointer(dst->cgrou= ps)) + return; + WARN_ON_ONCE(cgroup_attach_task_all(src, dst)); +} +#else +static void thread_handoff_cgroup(struct task_struct *dst, + struct task_struct *src) +{ +} +#endif + +/* + * Adopt the source's creds. Not commit_creds(), the process isn't changing + * credentials, an existing identity is just moving between two of its tas= ks. + */ +static void thread_handoff_creds(struct task_struct *dst, + struct task_struct *src) +{ + /* neither side changes its own creds while a handoff is in flight */ + const struct cred *old =3D rcu_dereference_protected(dst->real_cred, true= ); + const struct cred *new =3D rcu_dereference_protected(src->real_cred, true= ); + + WARN_ON_ONCE(rcu_dereference_protected(dst->cred, true) !=3D old); + if (new =3D=3D old) + return; + + get_cred_many(new, 2); + if (new->user !=3D old->user) + inc_rlimit_ucounts(new->ucounts, UCOUNT_RLIMIT_NPROC, 1); + rcu_assign_pointer(dst->real_cred, new); + rcu_assign_pointer(dst->cred, new); + if (new->user !=3D old->user) + dec_rlimit_ucounts(old->ucounts, UCOUNT_RLIMIT_NPROC, 1); + put_cred_many(old, 2); +} + +/* the user requested affinity follows, the effective mask derives from it= */ +static void thread_handoff_affinity(struct task_struct *dst, + struct task_struct *src) +{ + /* dup_user_cpus_ptr() wants no user mask, nothing can race us here */ + release_user_cpus_ptr(dst); + dup_user_cpus_ptr(dst, src, NUMA_NO_NODE); + set_cpus_allowed_ptr(dst, src->cpus_ptr); +} + +/* runs on the destination, the source never looks at the moved state agai= n */ +int thread_handoff_finish(struct task_struct *src, + struct thread_handoff_stats *st) +{ + struct task_struct *dst =3D current; + struct sighand_struct *sighand; + char comm[TASK_COMM_LEN]; + bool leader; + + /* the same for both, neither can be mid exec */ + sighand =3D rcu_dereference_protected(dst->sighand, true); + WARN_ON_ONCE(sighand !=3D rcu_dereference_protected(src->sighand, true)); + + /* tid, leadership, and signal state swap in one go */ + cgroup_threadgroup_change_begin(dst); + write_lock_irq(&tasklist_lock); + spin_lock(&sighand->siglock); + leader =3D thread_group_leader(src); + exchange_tids(dst, src); + if (leader) + thread_handoff_leader(dst, src); + thread_handoff_children(dst, src); + thread_handoff_signals(dst, src); + spin_unlock(&sighand->siglock); + write_unlock_irq(&tasklist_lock); + cgroup_threadgroup_change_end(dst); + recalc_sigpending(); + +#ifdef CONFIG_FUTEX + dst->futex.robust_list =3D src->futex.robust_list; + src->futex.robust_list =3D NULL; +#ifdef CONFIG_COMPAT + dst->futex.compat_robust_list =3D src->futex.compat_robust_list; + src->futex.compat_robust_list =3D NULL; +#endif +#endif + dst->clear_child_tid =3D src->clear_child_tid; + src->clear_child_tid =3D NULL; + +#ifdef CONFIG_RSEQ + scoped_guard(irqsave) { + dst->rseq =3D src->rseq; + memset(&src->rseq, 0, sizeof(src->rseq)); + src->rseq.ids.cpu_id =3D RSEQ_CPU_ID_UNINITIALIZED; + } + rseq_force_update(); +#endif + + thread_handoff_creds(dst, src); +#ifdef CONFIG_AUDIT + dst->loginuid =3D src->loginuid; + dst->sessionid =3D src->sessionid; +#endif + + dst->personality =3D src->personality; + dst->pdeath_signal =3D src->pdeath_signal; + dst->timer_slack_ns =3D src->timer_slack_ns; + dst->default_timer_slack_ns =3D src->default_timer_slack_ns; + dst->start_time =3D src->start_time; + dst->start_boottime =3D src->start_boottime; + if (task_no_new_privs(src)) + task_set_no_new_privs(dst); + /* prctl driven per-task flags, the source keeps them for its work */ + dst->flags =3D (dst->flags & ~THREAD_HANDOFF_PF_FLAGS) | + (src->flags & THREAD_HANDOFF_PF_FLAGS); + + thread_handoff_mempolicy(dst, src); + thread_handoff_cgroup(dst, src); + thread_handoff_ioprio(dst, src); + thread_handoff_stats_add(dst, st); + + thread_handoff_sched(dst, src); + thread_handoff_affinity(dst, src); + + get_task_comm(comm, src); + set_task_comm(dst, comm); + + return arch_thread_handoff_finish(src, leader); +} --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oi1-f174.google.com (mail-oi1-f174.google.com [209.85.167.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4DD8E443C20 for ; Fri, 11 Sep 2026 15:41:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.174 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141322; cv=none; b=Y1vcNfgAZSpF8o6kRjJjjXAlkNlOO0TUm8MxaIF5LJ5iMbHvkuav8NnuZLqsiXoEIktXKBCzhm1DfnYluG5oWLw7fr9O2B0J6f13b/A0w8i3de+32WMMrWGTPW/K0DgNLDSKzYryVtAXWNmHpsbIlyCGqdr1lP/KPDIXnho7Ubg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141322; c=relaxed/simple; bh=pHbCp/BMag/nnpNzMxAMuj5uUap5oH5CLoIai3Pt0L8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=GiBPULGW6cqK2FVwLcPgaZ1KdyX1i5EvL2j9vPAxULqbGBeT76r2cMNuzOrnVcBx2Jl9jYzjS/ag8x+NhRQ9Jj2GeygS0o/W64v40m0cu6gNxTsLK2K+pONTM4SGAqVMBVx1GVHmeLur/KBgsqLngSY0Ug33hSsNJ5QrSJiAUIY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=uG44yyeS; arc=none smtp.client-ip=209.85.167.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="uG44yyeS" Received: by mail-oi1-f174.google.com with SMTP id 5614622812f47-4c14d66d922so549426b6e.0 for ; Fri, 11 Sep 2026 08:41:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141318; x=1789746118; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=/RcGks6YdaJ2UkXyVy3FKEoUGQQvU1qrezFrY/Mkmuw=; b=uG44yyeS2p/eswpm9iO/SEecQ25QuCzgS3aUtbAGyWkopEBCZjHUWByJwLnUJV3agY 70iXi8ihypvBrjWzdi4OwlyGl0k5pTDcC6c1Qwo+dh3/up9TfE/W4hQS6RoW8yItkLUw 11UpwLSVx+iQqOT/wUCPy/BjnFKAtaYlujGfHOzxPHq1MenRifRGL45J5oFz7m5FDILM NI7gGt2oNoj2d0sfjbaLfKmI9SDumoScj6L6mxGSnx0SfaV2jxz4vo/DFnB9Dtane6BL rgpezV4fhaVCRM/R6/bCLyKyP3zoB+vPF6J0QaTfPdO6IMBtUBbwdeVGyAaxk2PSjRlJ FnnA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141318; x=1789746118; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=/RcGks6YdaJ2UkXyVy3FKEoUGQQvU1qrezFrY/Mkmuw=; b=Lb5VTNKpz3NRNdGNDZQCbhcNT31tuI0pYmhxcDhU5YJpQHSW8++/+lqS3aYeUBuRRB qqRCSG6V68B85Fq9hfs9txBhaYyXvOn/uTOhekk9gv3IiZ19IHIxnokzfwrO/SGMx4It Vm6NcUkSIl9j0dd8g87vZciKT7xjNhcX8IDm8v1z8OULtyLnpjQ1v8YbL0oiFRtfIFAg SeyL6Cmizz+JFd86aXh7nV2j7ZOFC3KnrAqyRyF1aWrEqxLS3RusVWsTQaYIqQeN6JVm gZkiBaexHcuSHJW/SXok+BEZkfpNwW/RbfUtvG1o4PCyxd+3mo49ySza0KOqBOWdo1C6 /rkg== X-Forwarded-Encrypted: i=1; AKwUvBxYcwn7pF7Ym/qirGYN5LJjlH1/6yZYV6HCr6qp8NoyREHT/9BMHONwVpiIIO/2E86pfPgj6dHBerhlPg0=@vger.kernel.org X-Gm-Message-State: AFuF++n/mdwO6cS1btnQq3FgsmYUKNwQGs3a8iXiyWMNmW0HtTAszNgz litMt4IEOpOeqpBGa41aQYZifaBCr/Y7021/S6ucby3qYKKUL0ZtcG0pmL1HUK8S05Q= X-Gm-Gg: AYBFou2DvFe54Ya/YYoVKmzZtZzTNaGavR6iI3+taoGIh+JmSwxnvIWJXBWFEbiYTWo vtG9/6Vx1lkn9qcDpAX8GtjPRFQlRaqLqJtqO7TukzHGRaDKjwdCd2spuuvgOy5UVukSR1H+050 B9+7APjc6Yp2ijqr1Hd2Uwe5dAWeO53LujEty+mXBVTC+BT2D/fXqnHC5hTbR9fQ+SCfr98nME1 eghvGcUkvY82EXWrXrSWH3vQ20M/r60S2p+/zaC+5+DP6tT55v9PuUEqDdTqteoUngcTJ1ZxBT4 6sOuBaFNwJU+bSP+8M2sd8Mf0bOvzfUCN5zwfe+qMVrYrri9Kxm4b0r0VqMLFTySFqFK/teetNA 9eT33Duuj+dEiMGX9/g3RhMRyOdoHHNkvM8+hRALESpJBpV4aB0JbFz2BuzDwCidnwMCiSGtNz3 xElJzIbq55aOpw0GJla9gEiZpzfd0evdva+GzAb79k2r3+rdiZb1q93U28/Y/+3JQDMHrWhxasY JLtJT7q9+FEZ0NtBLrUoflAuHASCnvlaEnx9v1svoQ= X-Received: by 2002:a05:6820:c446:20b0:69d:a4f0:4191 with SMTP id 006d021491bc7-6c0b9a57494mr2410199eaf.12.1789141317982; Fri, 11 Sep 2026 08:41:57 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.41.57 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:41:57 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 02/15] sched: call into io_uring when a PF_IO_HANDOFF task blocks Date: Fri, 11 Sep 2026 09:40:52 -0600 Message-ID: <20260911154148.644489-3-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add PF_IO_HANDOFF, set by io_uring on a task for the duration of an inline request issue that may block, and have sched_submit_work() call io_uring_task_sleeping() when such a task blocks. Placeholder for now. Signed-off-by: Jens Axboe --- include/linux/io_uring.h | 5 +++++ include/linux/sched.h | 2 +- kernel/fork.c | 3 ++- kernel/sched/core.c | 3 +++ 4 files changed, 11 insertions(+), 2 deletions(-) diff --git a/include/linux/io_uring.h b/include/linux/io_uring.h index d1aa4edfc2a5..969de22c3d0f 100644 --- a/include/linux/io_uring.h +++ b/include/linux/io_uring.h @@ -60,4 +60,9 @@ static inline int io_uring_fork(struct task_struct *tsk) } #endif =20 +/* called from sched_submit_work() when a PF_IO_HANDOFF task blocks */ +static inline void io_uring_task_sleeping(struct task_struct *tsk) +{ +} + #endif diff --git a/include/linux/sched.h b/include/linux/sched.h index 8b3d47a325cc..310310865029 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -1816,7 +1816,7 @@ extern struct pid *cad_pid; * I am cleaning dirty pages from some other bdi. */ #define PF_KTHREAD 0x00200000 /* I am a kernel thread */ #define PF_RANDOMIZE 0x00400000 /* Randomize virtual address space */ -#define PF__HOLE__00800000 0x00800000 +#define PF_IO_HANDOFF 0x00800000 /* io_uring: hand identity off if the ta= sk blocks */ #define PF__HOLE__01000000 0x01000000 #define PF__HOLE__02000000 0x02000000 #define PF_NO_SETAFFINITY 0x04000000 /* Userland is not allowed to meddle = with cpus_mask */ diff --git a/kernel/fork.c b/kernel/fork.c index 416758c8a3d4..510c8a9aa870 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -2190,7 +2190,8 @@ __latent_entropy struct task_struct *copy_process( goto bad_fork_cleanup_count; =20 delayacct_tsk_init(p); /* Must remain after dup_task_struct() */ - p->flags &=3D ~(PF_SUPERPRIV | PF_WQ_WORKER | PF_IDLE | PF_NO_SETAFFINITY= ); + p->flags &=3D ~(PF_SUPERPRIV | PF_WQ_WORKER | PF_IDLE | + PF_NO_SETAFFINITY | PF_IO_HANDOFF); p->flags |=3D PF_FORKNOEXEC; INIT_LIST_HEAD(&p->children); INIT_LIST_HEAD(&p->sibling); diff --git a/kernel/sched/core.c b/kernel/sched/core.c index eeb55367c0c6..bba1c3b26b7e 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -96,6 +96,7 @@ =20 #include "../workqueue_internal.h" #include "../../io_uring/io-wq.h" +#include #include #include "../smpboot.h" #include "../locking/mutex.h" @@ -7352,6 +7353,8 @@ static inline void sched_submit_work(struct task_stru= ct *tsk) wq_worker_sleeping(tsk); else if (task_flags & PF_IO_WORKER) io_wq_worker_sleeping(tsk); + else if (task_flags & PF_IO_HANDOFF) + io_uring_task_sleeping(tsk); =20 /* * spinlock and rwlock must not flush block requests. This will --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oa2-f12.google.com (mail-oa2-f12.google.com [74.125.231.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E6FF1476694 for ; Fri, 11 Sep 2026 15:42:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.76 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141323; cv=none; b=rJrWRgaLMjy7X4rboHLTk6jHw3UyLYzzKqAErKLe0lXyaW2TyAn7OVWPm6C8hV1Or31K1tEk7YZr6uT8Z2PdRpGpURA3QBOPr57XLqHsIc7xhPlyA8tlAENO5W3dDjyw9trZD4oZq3pEIkyJzAoukzbgyzgDIwKuPr+ErdEOk2A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141323; c=relaxed/simple; bh=imsgKgY1OaIa8G+gQkSklejkPfoEILr53ZwC+FKcriY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gTo6YV7Stg+cxNBumRjzw+AKJzqeaLKQFcwEWQuoi91ahWMcBMliN144mN59KY5cyrLiihS2iVKQKqI36VJY8aLbymBJzaoDMDJXgS1NSxVdAivI004MrAryTmDu0jGzxdGqNIeBhMG4RsREC35v/2d34pXVTf9GH3EDoDGZ32Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=OYxWRbk/; arc=none smtp.client-ip=74.125.231.76 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="OYxWRbk/" Received: by mail-oa2-f12.google.com with SMTP id 586e51a60fabf-46accbdfc20so28352fac.0 for ; Fri, 11 Sep 2026 08:42:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141319; x=1789746119; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=BaSkBg1gwRLv4QW0pPStlZCwsjMD0iYU2zavQE68qlw=; b=OYxWRbk/7Me8tLmO4K+1EYGEernm3x4ncDAfwqgJ5wsXQk3Sr4R1uR8S8nniNf3SAk iTuTZuHzZHx+ZthER4+tbAt3Gfj0vYWv7G2JPI+MjSs2s4Nrsmr47dQnJbRy+O7+oy1I s4wKHI+g+TxQITPnCZGOOeNOdoZqh9jHLt9Hg0AWGCwF5kTEvlbGw8eHz2ckvQtE2bvz Kp3epyPWkB9bs+jzlKDQrQL2rpPG2grTZDnj/0QmO4iB8+Gk5+gHix2wJRWQ/ig440VP mZzQQA3uRfs+W628IUcKvGT1lT/OSAECIP90RMYGMltanqUA2ozTQRagBAl4Ab4unG7m f/Tw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141319; x=1789746119; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=BaSkBg1gwRLv4QW0pPStlZCwsjMD0iYU2zavQE68qlw=; b=TEVBo/PYL8n7exAdAvzMI3HOnegmH/cXyJNxLhv999AH3dWgzFZhNNHdIbHSMgUcq+ t6KGh/wwmpA8B7bSp1b0Q+NvN4+w33/9fY2qJnIPuzUqE4cvTY4rg9rOfv6LUmnDTvNG UnKJLBnDf+ght3pXET7dMN2ep8wBjK+rrX0AxLGxYDLx1uxghDSZI0401OAzVCYWp090 UcY6U9FXopfglaEyiVfq3PL4cG20pI2wZ0Xnmxtb4aXKmPAaYez2AFsuomDs9B6FtvP7 r3HMZwCZIOXzoxhWiF82sYubhJ+kzr9XTHrjDl9wZr+TsKJ++8+YmfJIcKpoPuR3kcu5 yk7A== X-Forwarded-Encrypted: i=1; AKwUvBzos0URN+BzlwIrMx/CFxqsYp0+74g+AGnatvjH9Tb0cEVX5lX6HGQ4wP5QcFPXkdLv5XrvTYDZ3PlX9Is=@vger.kernel.org X-Gm-Message-State: AFuF++n79S/IPv2eU1u4bbi8KJtUe1y7XyjCKMYclGoLJi8brN8ZP+PC NLnBzV/osMzn2xb2Xp5qB9Z+OF7m9YS4N9snfe1tPjtZf2UY4ydT6wtBIG4su6WCANE= X-Gm-Gg: AYBFou3IjaAJUopR/xnnFf2Mm0s/xT8bCzMgqqSNVwSfcwN2udxfGr7HThsozJ/dISy l2fzicWmSjIetFfgWYwr3GJyFFuU6DBr+IFwNHjqesOcnUGjNyQyMEP+jNTU8xGSuploa7IMPmL HxDqjqxlt3Sw8gRXL+dVsuBrfC7W1+T2G0hrzphcg5xNIYcxALCANi/kn644FP30Cq8BEwacSTH kHgOv4NCttPuZuTyKaJmld/xXlJND8VGcRIri64Gf/fXIqZF9WDpEmngYMzdTaKDRkcsXeeQqpF OzhWZMemjSGzaOMUZRDjc/Fh6xrJZiSD30E54KUnSS5awrnXjpFGPQkm/vdJ1i0Pc4Q4rK54gO8 l8wBcg05JhB3f2SYXoPb15YDPm+DfRgDkqn56VZoy92vMotrp4knxnu5wAfHICOT3Eu1NuycGzR Je4Y9+2lm0ypq/Ii/sraZLvksK+KMPUet+tzStwOw86rCCxE13bowbC9nxQq4YGAMIJFtzKc5A3 ie7q3n07AlFV3SpIcc3RlB7pl3mkU3+LglrszaIwxoU X-Received: by 2002:a05:6820:571c:20b0:6b6:4f90:a62a with SMTP id 006d021491bc7-6c0ba8280bbmr4922795eaf.11.1789141319190; Fri, 11 Sep 2026 08:41:59 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.41.58 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:41:58 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 03/15] arm64: implement thread identity handoff Date: Fri, 11 Sep 2026 09:40:53 -0600 Message-ID: <20260911154148.644489-4-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Implement the arch_thread_handoff_*() hooks and select ARCH_HAS_THREAD_HANDOFF. The source is inside a syscall, so only the FPSIMD view of the vector registers needs to move. Tasks with SME streaming mode or ZA enabled, compat tasks, tasks with a guarded control stack and tasks trapping counter-timer access are refused. Signed-off-by: Jens Axboe --- arch/arm64/Kconfig | 1 + arch/arm64/kernel/process.c | 109 ++++++++++++++++++++++++++++++++++++ 2 files changed, 110 insertions(+) diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig index b5a51b0ef944..a919570c80cd 100644 --- a/arch/arm64/Kconfig +++ b/arch/arm64/Kconfig @@ -46,6 +46,7 @@ config ARM64 select ARCH_HAS_PTE_SPECIAL select ARCH_HAS_HW_PTE_YOUNG select ARCH_HAS_SETUP_DMA_OPS + select ARCH_HAS_THREAD_HANDOFF select ARCH_HAS_SET_DIRECT_MAP select ARCH_HAS_SET_MEMORY select ARCH_HAS_FORCE_DMA_UNENCRYPTED diff --git a/arch/arm64/kernel/process.c b/arch/arm64/kernel/process.c index 581f80e9b9b7..e669a2717106 100644 --- a/arch/arm64/kernel/process.c +++ b/arch/arm64/kernel/process.c @@ -11,6 +11,7 @@ #include #include #include +#include #include #include #include @@ -1003,3 +1004,111 @@ int set_tsc_mode(unsigned int val) =20 return do_set_tsc_mode(val); } + +#ifdef CONFIG_THREAD_HANDOFF +/* + * Thread identity handoff. The source is inside a syscall, so only the FP= SIMD + * view of the vector registers needs to move. SME state is refused. + */ +static bool thread_handoff_task_ok(struct task_struct *tsk) +{ + if (is_compat_thread(task_thread_info(tsk))) + return false; + /* the GCS is per-thread and would have to move along */ + if (task_gcs_el0_enabled(tsk)) + return false; + /* counter-timer trapping doesn't move, see update_cntkctl_el1() */ + if (test_tsk_thread_flag(tsk, TIF_TSC_SIGSEGV)) + return false; + return true; +} + +bool arch_thread_handoff_allowed(struct task_struct *tsk) +{ + return thread_handoff_task_ok(tsk); +} + +bool arch_thread_handoff_compatible(struct task_struct *src, + struct task_struct *dst) +{ + return thread_handoff_task_ok(dst); +} + +/* sync the live user state, fpsimd_syscall_enter() guarantees FPSIMD form= at */ +bool arch_thread_handoff_prepare(void) +{ + struct thread_struct *thread =3D ¤t->thread; + + fpsimd_preserve_current_state(); + tls_preserve_current_state(); + if (system_supports_poe()) + thread->por_el0 =3D read_sysreg_s(SYS_POR_EL0); + + if (WARN_ON_ONCE(thread->fp_type !=3D FP_STATE_FPSIMD)) + return false; + /* ZA and streaming mode state doesn't move, for now */ + if (thread_sm_enabled(thread) || thread_za_enabled(thread)) + return false; + return true; +} + +/* copy the user register state over, load what the return to user won't */ +int arch_thread_handoff_finish(struct task_struct *src, bool leader) +{ + struct task_struct *dst =3D current; + struct pt_regs *regs =3D task_pt_regs(dst); + struct frame_record_meta stackframe =3D regs->stackframe; + + /* the syscall frame, this is what the return to userspace restores */ + *regs =3D *task_pt_regs(src); + regs->stackframe =3D stackframe; + + /* + * Marks our FP state foreign so nothing saves over what's copied in + * below, and frees the SVE/SME buffers, the vector lengths may differ. + */ + fpsimd_flush_thread(); + + preempt_disable(); + + dst->thread.uw =3D src->thread.uw; + dst->thread.fp_type =3D FP_STATE_FPSIMD; + memcpy(dst->thread.vl, src->thread.vl, sizeof(dst->thread.vl)); + memcpy(dst->thread.vl_onexec, src->thread.vl_onexec, + sizeof(dst->thread.vl_onexec)); + dst->thread.svcr =3D src->thread.svcr; + dst->thread.tpidr2_el0 =3D src->thread.tpidr2_el0; + dst->thread.por_el0 =3D src->thread.por_el0; + dst->thread.sctlr_user =3D src->thread.sctlr_user; +#ifdef CONFIG_ARM64_MTE + dst->thread.mte_ctrl =3D src->thread.mte_ctrl; +#endif +#ifdef CONFIG_ARM64_PTR_AUTH + dst->thread.keys_user =3D src->thread.keys_user; +#endif + update_tsk_thread_flag(dst, TIF_SVE_VL_INHERIT, + test_tsk_thread_flag(src, TIF_SVE_VL_INHERIT)); + update_tsk_thread_flag(dst, TIF_SME_VL_INHERIT, + test_tsk_thread_flag(src, TIF_SME_VL_INHERIT)); + update_tsk_thread_flag(dst, TIF_TAGGED_ADDR, + test_tsk_thread_flag(src, TIF_TAGGED_ADDR)); + update_tsk_thread_flag(dst, TIF_SSBD, + test_tsk_thread_flag(src, TIF_SSBD)); + if (test_and_clear_tsk_thread_flag(src, TIF_MTE_ASYNC_FAULT)) + set_tsk_thread_flag(dst, TIF_MTE_ASYNC_FAULT); + + /* load what __switch_to() would have, FPSIMD gets restored on exit */ + write_sysreg(dst->thread.uw.tp_value, tpidr_el0); + if (system_supports_tpidr2()) + write_sysreg_s(dst->thread.tpidr2_el0, SYS_TPIDR2_EL0); + if (system_supports_poe()) + write_sysreg_s(dst->thread.por_el0, SYS_POR_EL0); + contextidr_thread_switch(dst); + ptrauth_thread_switch_user(dst); + mte_thread_switch(dst); + update_sctlr_el1(dst->thread.sctlr_user); + + preempt_enable(); + return 0; +} +#endif --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oi1-f179.google.com (mail-oi1-f179.google.com [209.85.167.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DFBFC472068 for ; Fri, 11 Sep 2026 15:42:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141325; cv=none; b=r/aqkj7AiU0s3cEyiEAHqbmXdvaRC/t6/Tt9Cvs+Jh9TJOYi8LLdcmMM8V/WNa8DFIxkZxqNavAJuJW405RVutDLzBnwIMJoNK65J05eB1rij0KEQCLRygJuRpWYgAl98yF+2BJzsDINU+7PEPSD04Y8rv0mTzs6zB0dvytlxcw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141325; c=relaxed/simple; bh=WoQD+GEr/rVa5dRIa/E9lO9G5gP37D3TEnDGbuP6Hvs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=bySeiTx3Lsvy4WdMlnmqF/H5YF+E09cRCQT3QCsX+TwNCZWaJ5gc/zg3vLDH6PHEwhKHz+fkVgQ+oFwvwpT+XHMKOOXkpmcBTa0Puy19Fpv5IsRKXEmMfI1dvnb7J1PIFqoqinlwXCmcn7AKlMrp/S8idpmZwqmcPjXJh+CWy5k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=UG8smZmZ; arc=none smtp.client-ip=209.85.167.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="UG8smZmZ" Received: by mail-oi1-f179.google.com with SMTP id 5614622812f47-4a4cb36ae00so886872b6e.0 for ; Fri, 11 Sep 2026 08:42:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141320; x=1789746120; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Vd+hpYguRm0/LXSau2U/e2KAtn148hROAqJ+zClJj0g=; b=UG8smZmZ1hZ/C+Tjnt3eyuLFAG0SkuIw265DY6JXitZdILCc7DGsVbPFGjw51VWSde xdPxOJscTS85eAroDCSHHkVhikuCCBObLPefq4DPvEAfVuZaRo9Lt7eL2SK/JGGewW1C Vkyj2F8aYk4mOQ99v0UYA3QNg4A0Pcys+AWgI75CLlh6PWznJTcf0j52RAZpb3JsRcoy q/CQtRlJFuYM7YliAYtshooKifCLbqPSOAzRUpG952l//XNNffj7/jCxI+vIJ1P4NwTS wzMsPqcZZirbI6z3UFif9fc8bx0XQnFIYjsY1vLmC9SEnIZJpcfo5JXYYH53QHJSNzWl Brhw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141320; x=1789746120; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Vd+hpYguRm0/LXSau2U/e2KAtn148hROAqJ+zClJj0g=; b=pgiLCEwigVAT2nPjkPCOwba3xFhM5x/y8GOgWWXWIZh8afbdjkqVgdRyJoaS8lm3Mw 0ruOuqL3YXpjs5kZxmCJRF5FWYVXie1dtmvCVyBdO6MdJZ+qce7IBZ62Ie5Tw5ek0kZ9 fSP97S+xbOnb8Bvz7kM8VkLaGA3qdec6Ado1I0jlA7Wc7Ybdi1VsajLulY8Y8Px8B6+S KBU6Rli8/rxUT5BksOiqEtaYeKhqkHkDcZ5iHiA+uz4PBSXE622AJ/cwUAu5ENT6SEPZ DGrpZ0B2iKRMnhngD8qFbS5fDgomE2lDX+/wYNkyl2ounr7mPqBTaKVc50avFt+7huOu mhiA== X-Forwarded-Encrypted: i=1; AKwUvBxnGJAY/X3pr6HAzU6nLZ678rTMUgddNhOXKSi+02lYNNBqZgCGvK9izezaMZm7GkfswyKBxOVT3PFQukk=@vger.kernel.org X-Gm-Message-State: AFuF++n5NCqdKJIcUOwKZibp9wst9Ad25gxFO0AszNklOqppMibNYb1d zTjwsG7yDKdBTxBI12gDJGGoDBacHPqRQLVyBSkzOezy/nmlKV7Qcl6I5OPVkb2AYdE= X-Gm-Gg: AYBFou1k1HhguswTqJgbne+EKDIqHzJCeDzFUzxLvPewQpWynaYWtXkDHXmqZXYhxag gTQXNqVRPMDDV48tTtm5UToEzRcKM+vjesMqUA7kKi2+vWwfLR+QUdQ/FtjUeVfcnrftw0FnK8B Vix6IoIQwzwpKc7xHyy13x2OrJifjx+iNHYKlnf4sbHV+MMFmBw3bXE0JV13fbDtahyn9tzu54P VNYntOKcheYJIIb9zknFsULYZImTG+DGOI4Bs6lQs3VKGFPZWaWrfaY03NcYEBq5qGZSAGwoplR u1nKSbq+o7RMtA+3u7wH9vuPkf/iiSObEir3cTiE7eV6fsNNPn30QmRfHr2ETkf/jkDFoByFWrX MZ0gTLf3o9exn/pmjYEl5gSeQWT5/vI3DEQjR10dzD8zLpGv5NZ0cnVBm2u7PUs1mkRxmunzyiW qe4ANBvW6IekMw9x9TWraF3adpHVBYFV/OpQIciRksOdzhLnZTyascJN6ccsn3AAa7tDx5+v43f AabP9Q02LlKJeRgeZekLsebwwKX7gGbLm1ydLo3TgJv X-Received: by 2002:a05:6820:491a:b0:6b7:46fc:1d3 with SMTP id 006d021491bc7-6c0bd587be9mr2976582eaf.50.1789141320381; Fri, 11 Sep 2026 08:42:00 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.41.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:41:59 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 04/15] x86: implement thread identity handoff Date: Fri, 11 Sep 2026 09:40:54 -0600 Message-ID: <20260911154148.644489-5-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Implement the arch_thread_handoff_*() hooks for 64-bit x86 and select ARCH_HAS_THREAD_HANDOFF. Prepare saves FS/GS, PKRU and the FPU state. Finish copies the syscall pt_regs, fault info and FPU image over, and loads what __switch_to() would have. Refused are 32-bit tasks, I/O bitmaps and emulated iopl, per-thread speculation and CPUID/TSC controls, user shadow stacks and non-default sized fpstates. ret_from_fork() now returns what the thread function returns in regs->ax rather than 0. A kernel thread returning from kernel_execve() returns 0 anyway, an io-wq worker that got handed a user identity returns the result of the syscall it took over. Signed-off-by: Jens Axboe --- arch/x86/Kconfig | 1 + arch/x86/kernel/process.c | 10 +-- arch/x86/kernel/process_64.c | 139 +++++++++++++++++++++++++++++++++++ 3 files changed, 145 insertions(+), 5 deletions(-) diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig index 15fd9ec5ecac..4f53859d0228 100644 --- a/arch/x86/Kconfig +++ b/arch/x86/Kconfig @@ -109,6 +109,7 @@ config X86 select ARCH_HAS_STRICT_MODULE_RWX select ARCH_HAS_SYNC_CORE_BEFORE_USERMODE select ARCH_HAS_SYSCALL_WRAPPER + select ARCH_HAS_THREAD_HANDOFF if X86_64 select ARCH_HAS_UBSAN select ARCH_HAS_DEBUG_WX select ARCH_HAS_ZONE_DMA_SET if EXPERT diff --git a/arch/x86/kernel/process.c b/arch/x86/kernel/process.c index 346c438ac880..ed52af862392 100644 --- a/arch/x86/kernel/process.c +++ b/arch/x86/kernel/process.c @@ -155,13 +155,13 @@ __visible void ret_from_fork(struct task_struct *prev= , struct pt_regs *regs, =20 /* Is this a kernel thread? */ if (unlikely(fn)) { - fn(fn_arg); + long ret =3D fn(fn_arg); + /* - * A kernel thread is allowed to return here after successfully - * calling kernel_execve(). Exit to userspace to complete the - * execve() syscall. + * A kernel thread returning from kernel_execve(), or an io-wq + * worker returning the result of a syscall it took over. */ - regs->ax =3D 0; + regs->ax =3D ret; } =20 syscall_exit_to_user_mode(regs); diff --git a/arch/x86/kernel/process_64.c b/arch/x86/kernel/process_64.c index 2bce7b3f97ed..0f07e9a6bb02 100644 --- a/arch/x86/kernel/process_64.c +++ b/arch/x86/kernel/process_64.c @@ -41,10 +41,12 @@ #include #include #include +#include =20 #include #include #include +#include #include #include #include @@ -980,3 +982,140 @@ long do_arch_prctl_64(struct task_struct *task, int o= ption, unsigned long arg2) =20 return ret; } + +#ifdef CONFIG_THREAD_HANDOFF +/* Thread identity handoff, see include/linux/thread_handoff.h */ + +/* prctl driven per-thread controls that __switch_to_xtra() applies */ +#define THREAD_HANDOFF_TIF_MATCH \ + (_TIF_SSBD | _TIF_SPEC_IB | _TIF_NOCPUID | _TIF_NOTSC) + +/* state bound to the task that neither side may have */ +static bool thread_handoff_task_ok(struct task_struct *tsk) +{ + /* I/O permissions, the bitmap hangs off the task */ + if (test_tsk_thread_flag(tsk, TIF_IO_BITMAP) || tsk->thread.iopl_emul) + return false; +#ifdef CONFIG_X86_USER_SHADOW_STACK + /* the shadow stack is per-thread and would have to move along */ + if (tsk->thread.features & ARCH_SHSTK_SHSTK) + return false; +#endif + /* only the default sized FPU state gets copied over, no AMX */ + if (x86_task_fpu(tsk)->fpstate->is_valloc) + return false; + return true; +} + +bool arch_thread_handoff_allowed(struct task_struct *tsk) +{ + /* 64-bit tasks only */ + if (test_tsk_thread_flag(tsk, TIF_ADDR32)) + return false; + return thread_handoff_task_ok(tsk); +} + +bool arch_thread_handoff_compatible(struct task_struct *src, + struct task_struct *dst) +{ + /* these don't move, must match. They're usually applied process wide */ + if ((read_task_thread_flags(src) ^ read_task_thread_flags(dst)) & + THREAD_HANDOFF_TIF_MATCH) + return false; + return thread_handoff_task_ok(dst); +} + +/* + * Sync the live user register state. TIF_NEED_FPU_LOAD makes the in-memory + * FPU image final, later context switches won't write it again. + */ +bool arch_thread_handoff_prepare(void) +{ + current_save_fsgs(); + /* thread.pkru is only valid when scheduled out, make it so */ + if (cpu_feature_enabled(X86_FEATURE_OSPKE)) + current->thread.pkru =3D read_pkru(); + fpregs_lock(); + if (!test_thread_flag(TIF_NEED_FPU_LOAD)) { + save_fpregs_to_fpstate(x86_task_fpu(current)); + set_thread_flag(TIF_NEED_FPU_LOAD); + } + fpregs_unlock(); + return true; +} + +/* copy the user register state over, load what __switch_to() would have */ +int arch_thread_handoff_finish(struct task_struct *src, bool leader) +{ + struct task_struct *dst =3D current; + struct thread_struct *t =3D &dst->thread, *s =3D &src->thread; + struct fpu *dst_fpu =3D x86_task_fpu(dst), *src_fpu =3D x86_task_fpu(src); + struct thread_struct prev; + + /* the syscall frame, this is what the return to userspace restores */ + *task_pt_regs(dst) =3D *task_pt_regs(src); + + /* fault info, in case a signal for it is pending */ + t->cr2 =3D s->cr2; + t->trap_nr =3D s->trap_nr; + t->error_code =3D s->error_code; + + /* + * Dynamic xstate permissions are a property of the process but live + * in the group leader's struct fpu, see xstate_get_group_perm(). + */ + if (leader && fpu_state_size_dynamic()) { + struct sighand_struct *sighand; + + sighand =3D rcu_dereference_protected(dst->sighand, true); + spin_lock_irq(&sighand->siglock); + dst_fpu->perm =3D src_fpu->perm; + dst_fpu->guest_perm =3D src_fpu->guest_perm; + spin_unlock_irq(&sighand->siglock); + } + + /* both sides have the default sized fpstate, reload on the way out */ + fpregs_lock(); + memcpy(&dst_fpu->fpstate->regs, &src_fpu->fpstate->regs, + src_fpu->fpstate->size); + dst_fpu->last_cpu =3D -1; + set_thread_flag(TIF_NEED_FPU_LOAD); + fpregs_unlock(); + + preempt_disable(); + + memcpy(t->tls_array, s->tls_array, sizeof(t->tls_array)); + load_TLS(t, smp_processor_id()); + + savesegment(es, t->es); + if (unlikely(t->es | s->es)) + loadsegment(es, s->es); + t->es =3D s->es; + savesegment(ds, t->ds); + if (unlikely(t->ds | s->ds)) + loadsegment(ds, s->ds); + t->ds =3D s->ds; + + /* FS/GS, the legacy load path needs to know what the CPU holds now */ + local_irq_disable(); + save_fsgs(dst); + prev.fsindex =3D t->fsindex; + prev.fsbase =3D t->fsbase; + prev.gsindex =3D t->gsindex; + prev.gsbase =3D t->gsbase; + t->fsindex =3D s->fsindex; + t->fsbase =3D s->fsbase; + t->gsindex =3D s->gsindex; + t->gsbase =3D s->gsbase; + x86_fsgsbase_load(&prev, t); + local_irq_enable(); + + if (cpu_feature_enabled(X86_FEATURE_OSPKE)) { + t->pkru =3D s->pkru; + write_pkru(t->pkru); + } + + preempt_enable(); + return 0; +} +#endif --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oo2-f43.google.com (mail-oo2-f43.google.com [74.125.231.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 67D93485954 for ; Fri, 11 Sep 2026 15:42:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141326; cv=none; b=jnXlA8N3cgDMrF0maGLNGslOnfhtnIgKBDq6gWVrkIOzcMpcMfbPRg9PyXjfYrUw45nxO/If2hRLQWrutPlG2wZS0S3fcEp8kH2B7aV9Krg4Q33zIAd7ytKpCQQfTwZR2CvDx9s/o5uYNcivKbntoMYPMTH4/4Xmg7/C5K+7vLU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141326; c=relaxed/simple; bh=KJRnZ7USBDtOxaI9oLCCZhPnMHJSo9U/cVcNXiZ7Jrk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=l8YuXZj43/QA/KV8rM5kw4kG+iLy8QlY9GW1M1qycJoyj4I6GVbKe4BEqgv/2qB1yTHupvhBDlB2Li5FJ247Uh0W0pNPMOnbBjg8lIKqIU/hnA3ZdxMtYTjW2Fg7393OnrB5OWuJbtrnV8wPbDI194UZWQIrsZhPfNudNqhzYWM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=XCPu4FNb; arc=none smtp.client-ip=74.125.231.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="XCPu4FNb" Received: by mail-oo2-f43.google.com with SMTP id 46e09a7af769-7f4f0cfb33cso13757a34.0 for ; Fri, 11 Sep 2026 08:42:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141322; x=1789746122; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=xN5FMpx5BomrnwwRlZAaFF7UH1EwB4frjE8eycmnS4k=; b=XCPu4FNbY6WeYi1MpCfIjFhynJfGO6JxH/YNqvjIrRbksojK1QaTuf9RKdrOTdCgTn GymY3mqZFcj18u+h6+cgX36X3NaRkfX0NLQxRUMc/16bUbKaSm/vYWOYeF2K8tgqoKqn A2iP96F31aekg35+EE6m8PUKMktOD3W6lrIAbj0Tgi2FloGaWqcQej18oKFSJlpUHoX3 n2sHhDT0d8D62XwjkoxN1FZDDxfK6+CgHUwrprRnhpIoRwgU1RRvw2fRYVUCwscE+EMu 1vIhkM0VWZFZPMdEH4ltwPR/BV2sobr1kwxSpFdlIKLcYbUrSshQZYt2UjcFMHHWMlew 2AuA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141322; x=1789746122; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=xN5FMpx5BomrnwwRlZAaFF7UH1EwB4frjE8eycmnS4k=; b=V86nQFG3V4cky4ZrTiVHqvg7kzUmYyvy0EfKOdurTnyipqW8kf7D8RXBBAtnXw7Q1/ ekXH5T2tCmM2eyV9WGPfyV9PYTyZ5js6I/hcrzzxf97I/JeKbU4KXtrcUp+WfDoR6jNt mp2EvEyg2xseMne0hQXrlTWoQOli14emwBCnlFc+v6v64JgW80bJw2/PLxcxS7Ol7m6b wCR4wm9aS/iXlNsU/CG/KH7phPTnQUbJ7D/b0CnsrJf+AtDAARLYpGKRNnoFcnChoQUQ Bb7o5+LJTpyBfJDI6tYBFm/tYfmK08XGP2yQyNbBav3tP9lAM1WhvnTfQMeBiYPE7C8B 1jmw== X-Forwarded-Encrypted: i=1; AKwUvBy6WuzfBWNMVknnlxM5diIq7AGjwQTKxIck0t86rDSz16QllyNtemTs53be3STuOqGfOFgnLbVn6XAITB0=@vger.kernel.org X-Gm-Message-State: AFuF++lbjiCcQsYOJva4tsoK4ZpuQ8lx1arOgIHJHTgQO15KcxUzXEMC AP5AbzUNhE1VBsuFetH8IbevZykNmzqV+PSRyzUjsnv3+GIX4c3LyTIQTYzdh8n3y2E= X-Gm-Gg: AYBFou1OQpuiTYNvJrkRhLCQjGINHuJ6R1LW4pRs9uZJyNwXUJ3R04nDrBnjhi3RyAo jCo1HQsNOkd0ibXEjPD9uMh1LTSNchl6vyEsiucrVvxlmBDXbtzcGIjlMAjH757dhmvJwYbV3qU 9ROfwUrGE/kJUdvGS8PBwzNDwzEGkv70K6Sn/eRNYh+rKL4V7d0fXhL4PnqCL4QmXrIwqV41s7L vo23VHAfvySKCXHbQW6d5DK6HE28m53QokeBZPzD2AHImyQv4ew691uuDMzCLQNC3ZKXdeM8kv4 SnVSiGpM8xmhodRmnTRQykFx6ndHR1c4Kv6YvggTE8M315MV3INs2cPygzl7Eze7n7IxqIqQBAS dNEXrFqRIzqLAyEZwU2ncqHDWHtsbYTKQezFuon84Qvu5cyKyO5dLwzl6NW+Y1xor+dNRT3MnDM jIm8BQ3BIepVGxbK+q0ZIcx+KY+tfcy6tHfr1FKnA6BsT7I79QtX1+oEX0caNxU27e23ktGO5HK 2t9txWn17HW/bYD/GdNMeYCP7MT8IzCNFrJZ5Fy4Lw= X-Received: by 2002:a05:6820:208b:b0:6b3:50e5:c625 with SMTP id 006d021491bc7-6c0bde62624mr6480780eaf.32.1789141322172; Fri, 11 Sep 2026 08:42:02 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:00 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 05/15] io_uring/kbuf: use io_ring_submit_unlock() helper Date: Fri, 11 Sep 2026 09:40:55 -0600 Message-ID: <20260911154148.644489-6-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" io_buffers_select() open-codes how it unlocks the ring based on the issue_flags, use the generic helper instead. Signed-off-by: Jens Axboe --- io_uring/kbuf.c | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/io_uring/kbuf.c b/io_uring/kbuf.c index 7c309173dd19..8f04e6cc773b 100644 --- a/io_uring/kbuf.c +++ b/io_uring/kbuf.c @@ -371,10 +371,9 @@ int io_buffers_select(struct io_kiocb *req, struct buf= _sel_arg *arg, ret =3D io_provided_buffers_select(req, &arg->out_len, sel->buf_list, ar= g->iovs); } out_unlock: - if (issue_flags & IO_URING_F_UNLOCKED) { + if (issue_flags & IO_URING_F_UNLOCKED) sel->buf_list =3D NULL; - mutex_unlock(&ctx->uring_lock); - } + io_ring_submit_unlock(ctx, issue_flags); return ret; } =20 --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-ot1-f44.google.com (mail-ot1-f44.google.com [209.85.210.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D6BA9489877 for ; Fri, 11 Sep 2026 15:42:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141327; cv=none; b=W0MN6Ocsfr57xFKjY/ToS/jX0XUwwkVFaBd+xReAP5n6L8wZpDgAT1UrKeyAU5j2P+xVQrvLPFts1O7dJjm3A/22d3mRrKV2A+3NZERu6iGmdatEkF7IXkcJ0akRpVnl+Zmoq8Oodr/Mm3nq8ogZLGQiWZjndbVphAN4eb2DHGs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141327; c=relaxed/simple; bh=0NwaKUmf44Ot+C4Vhin0K2ohE4gDm+QbMEeRlNNgATA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eUUlLjQzz46oicempmRyBoVMzeMo8iAXPWj4BM1AEKjbgE0pQJcVHS7bYgTkp3IbxCuMtgpyKqkZgZGHO4z5R9Tq9R+KUZlvM1NPsM54IFN/uyjaQWFUgtz56QnLDvGQmoAi6ITQ/ww6wiwhn0zhugVb6ImCbJdVSijdPxnKRLM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=g3tLiZG9; arc=none smtp.client-ip=209.85.210.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="g3tLiZG9" Received: by mail-ot1-f44.google.com with SMTP id 46e09a7af769-8049e44d0c3so560464a34.3 for ; Fri, 11 Sep 2026 08:42:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141324; x=1789746124; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=26NN2voef8YBl9mHMwb0pP++w/F7eMXxYYY8spi7/jc=; b=g3tLiZG9SIvwMgDaxYDCMXEg0m2r3sJqQ575QFWcccwf/+OP9S1ZkyTTDcBZ/GEeqc 22T1hHUq7pVox3RSQf8p+jzX+ypj7E+ZgbYIlpQ/UvWonFGKC6qN6ydshnJTBrwK/DQ4 IZcytgIsQHRVwafwNR7ULW8gKzl2NdfXAXjW5oFFaNWWUBenUu+2+Di/uEIQUhYtBJdH AsIACc1hFpYvv3bBqGhVuPl/jW+T86XefpKVdlgM01if0B3YcLaiIPzTCZTjhORT6aCW Tbs3BqGhwkfvp6iIeVdghIRzK670dVENUyFY0gfdmQq8X5a7cUcNLX1zHbonW1n69zrQ FORQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141324; x=1789746124; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=26NN2voef8YBl9mHMwb0pP++w/F7eMXxYYY8spi7/jc=; b=Z9weFb1DjSiLgiUNpnH57bSgRCcCSS3WEP0fc7Z5KKa7R+en1MHFBNNGQamGZWOgKd PtmgnnSoaRRkV1HnKVZV4UaTGQsuJp+moI2mt6DsKpN9OzImzu0DtBjwsXoMTDWcrnUk DDlCu93XtpOeDJRaz2zdT868dlWuNeuuIZYtH0X4AiBdY3BodrJYyPbSMZpenJ7QSJ2c eOBgXJEC0MmdDksfrInKuKXcuFUs8jWgIU7k1F1+NCrZKclCbHwXikz5ccJgfCrk5bmc LqLydvO5OllYrPzObZrx7F+v77N/jJQz3zxa9sP4nJ8eRECNUX/DyUEzXD4gh5r9KBYO N3mw== X-Forwarded-Encrypted: i=1; AKwUvBwCO9crn4ENh0v3xGV6WicfaVoNQi3PQG/kYyGmg2DFvbLdVaoN07R2YBlYVWTZskozotcbnyCsUfGkU1E=@vger.kernel.org X-Gm-Message-State: AFuF++kozLDxtZK0V3ZxtSKqfzBzQ4p5vLHYpQ7PmiQYxtV4tHWw2oEk QeoFH+jE5vn+zW0VoAv0xty3WPDrRH9bysvXjzbJ/oaE9DQ2L3+TMVwvMpx4kaeIwWY= X-Gm-Gg: AYBFou27fqqYvp0GFICDIMfNNijbFEMFSIC16tK2jBVhAASpdmAlhel6QGr6sPd7xyW qmCPz/x460nfSq7QGZ2emToWWy7vkNBdUm8SjBIZwBFU1msZ6ZODdDN8SU2BEQ4eQ5iiUAxl8wV uTVL6KJz6iOLmSUyiLoQAzBoi1phWu1xi2ux761zsU2jAwdxdRZlcUHqlyia3W5yoXNrpz7J4T/ uEbLUvmE3Yf30bNO0YWOwioqK+zZoC3UGwjFtvVOWQJm3Up4JB/JOHHmjgPn77cwCOxIOg5XSEa nLD26A53+nawnYsm4t8QCo2HdkwCywwLKGt84uU+WUrgKLsrksOFGTes9ZmIIaLavneqLlHJIjy DxdG5I+yEITyXSX95aUrZIn8T7sMO7TsgL3Ri1SSpF4AU+h9ER+FBHN5qM9bP1RZKWTZ6WJoKki MeZZEIRyHpmQtc+5vRhlWiWmc1+P1jUBngmE+QkCuTwW93NJPDl7lrCofGvU4V/jPv++MNRPbg9 ClBCLe9FU1HfV4xDGWgzejqw9rr6Gi4+X7L6Be8GTTk X-Received: by 2002:a05:6820:2903:b0:6ba:3635:7ffd with SMTP id 006d021491bc7-6c0bde5bcd5mr3300243eaf.64.1789141323624; Fri, 11 Sep 2026 08:42:03 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.02 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:02 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 06/15] io_uring: keep the tctx nodes on a list Date: Fri, 11 Sep 2026 09:40:56 -0600 Message-ID: <20260911154148.644489-7-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" tctx->xa is indexed by the ring pointer, which makes it a deep and sparse xarray. Iterating it with xa_for_each() is expensive, 12 usec for a single entry in testing. Keep the nodes on a list as well and walk that instead, lookups by ring stay in the xarray. Signed-off-by: Jens Axboe --- include/linux/io_uring_types.h | 2 ++ io_uring/tctx.c | 10 ++++++---- io_uring/tctx.h | 2 ++ 3 files changed, 10 insertions(+), 4 deletions(-) diff --git a/include/linux/io_uring_types.h b/include/linux/io_uring_types.h index 39629ee77b91..4af3d579ead6 100644 --- a/include/linux/io_uring_types.h +++ b/include/linux/io_uring_types.h @@ -153,6 +153,8 @@ struct io_uring_task { struct file *registered_rings[IO_RINGFD_REG_MAX]; =20 struct xarray xa; + /* the nodes in ->xa, for walking without the xarray lookup cost */ + struct list_head node_list; struct wait_queue_head wait; atomic_t in_cancel; atomic_t inflight_tracked; diff --git a/io_uring/tctx.c b/io_uring/tctx.c index 466b7300e208..737dfad4a976 100644 --- a/io_uring/tctx.c +++ b/io_uring/tctx.c @@ -105,6 +105,7 @@ __cold struct io_uring_task *io_uring_alloc_task_contex= t(struct task_struct *tas =20 tctx->task =3D task; xa_init(&tctx->xa); + INIT_LIST_HEAD(&tctx->node_list); init_waitqueue_head(&tctx->wait); atomic_set(&tctx->in_cancel, 0); atomic_set(&tctx->inflight_tracked, 0); @@ -135,6 +136,7 @@ static int io_tctx_install_node(struct io_ring_ctx *ctx, kfree(node); return ret; } + list_add(&node->tctx_link, &tctx->node_list); =20 mutex_lock(&ctx->tctx_lock); list_add(&node->ctx_node, &ctx->tctx_list); @@ -227,6 +229,7 @@ __cold void io_uring_del_tctx_node(unsigned long index) =20 WARN_ON_ONCE(current !=3D node->task); WARN_ON_ONCE(list_empty(&node->ctx_node)); + list_del(&node->tctx_link); =20 mutex_lock(&node->ctx->tctx_lock); list_del(&node->ctx_node); @@ -243,11 +246,10 @@ __cold void io_uring_del_tctx_node(unsigned long inde= x) __cold void io_uring_clean_tctx(struct io_uring_task *tctx) { struct io_wq *wq =3D tctx->io_wq; - struct io_tctx_node *node; - unsigned long index; + struct io_tctx_node *node, *tmp; =20 - xa_for_each(&tctx->xa, index, node) { - io_uring_del_tctx_node(index); + list_for_each_entry_safe(node, tmp, &tctx->node_list, tctx_link) { + io_uring_del_tctx_node((unsigned long)node->ctx); cond_resched(); } if (wq) { diff --git a/io_uring/tctx.h b/io_uring/tctx.h index 76ad1ad4594e..9a139816ce4f 100644 --- a/io_uring/tctx.h +++ b/io_uring/tctx.h @@ -2,6 +2,8 @@ =20 struct io_tctx_node { struct list_head ctx_node; + /* on tctx->node_list, only ever touched by the owning task */ + struct list_head tctx_link; struct task_struct *task; struct io_ring_ctx *ctx; }; --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oa2-f12.google.com (mail-oa2-f12.google.com [74.125.231.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 65EFC48C3EF for ; Fri, 11 Sep 2026 15:42:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.76 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141329; cv=none; b=PSG2hPOBkvg8i6P4kWM/T9wWkGMAdh8A2JegQTbVv+pxEAaBgv4j67FXUaTYM1G5YuwAv4nDBIPuy37D+L4yU7gfbVI7PX5jl4ennrvkCUBBY6pQ6Wxd+yYtMTw2gxe46+5mRFnDLRYLiyy2f0UlxxI/jwuIBsELuKw57zNFDog= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141329; c=relaxed/simple; bh=rrWhPwLXUMHPMn9y2nXl5b0RAfyzWV0Gp95H7S6fUlo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=OTWUMsobaTX+GsfbDUR2jQjGCwadmuP3W2GNE2Ix3MqEwYQD0JQ23gDVDc86PEpsfmaZ0SASrvK2ChOj8d1MmdYZo0c6YriMd4BEI897xUwNZdKf0c6wXjKUET9TPU09PGqLSm9+eQz1wcNX4UHnXUoiN8aOEWrVWZpTnimra4g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=AUg+VxdR; arc=none smtp.client-ip=74.125.231.76 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="AUg+VxdR" Received: by mail-oa2-f12.google.com with SMTP id 586e51a60fabf-466cc9b4b63so120344fac.1 for ; Fri, 11 Sep 2026 08:42:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141325; x=1789746125; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rIQRpDFpvkA6v5lh/5UPtLNtecnKXIQ2k9kxGqhZIEM=; b=AUg+VxdRZ5Hk5lqCCpEItfNbMJEzk2c033f4UtKsH0LxJoQgZ3drmpiKGBYTuIDvUY y+HI+lhlKRqjF63pFHC4xbD10fUb2JmsGC36TYjCy25eQuF8i5UFM4S1FzB7ARR9wVKf bgqoU/bF6ouoEX9gNCoc7rrMxHGCdRmnQJAnFtpZwq5wqx7Z7B7xnyn977q3rPLWl9cM 70QDGB0M4AK9RXaDTT7AY4lmc0kN3r0Nj7JACMitn7ZkJE53RGbmme6e5wISoBBxIxLx o+fu/VoOzdcQG+aFbyW1dZ3AHWwu+OsfF7dRkQICCLvCSD2guOwSyq5hf20Dr1G/twk1 zptw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141325; x=1789746125; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=rIQRpDFpvkA6v5lh/5UPtLNtecnKXIQ2k9kxGqhZIEM=; b=kB+soIzo+mgFl7J7wm6OHxe3tHfcB83CXlTtNKAsAhA4xRRxUQAreBn2Op2zEj5GYR /QTxL1hIJj70FEUuJvWg9pHFhmx/sAcBQDy21IsUikhVba1UPu7T4IiAFa2aJrtxZVr4 aOc3mZashEia/FVC1HVcnv0doFbxjptbKO5WUe+6x23BAcu+xY6TWXR/5DKtK8fKiHN5 Zdf/gWv0ljwDiR8SUv+rMLhLKEMbY8QCL56Nr4iS5yJQ/raXeGZSdyVWafYhNMJa8kOf X8mo6RlDlXBUvrvnb0ERHkMtXyT6c8uGOteBIrDj2Lwqm5LfkkVaGiJrreLa4gmTiS81 br/A== X-Forwarded-Encrypted: i=1; AKwUvBxJ8X4qb7hUa0VSzVELMSaIYKExU9e2AW7KN/wTQitOkMt8pnLShXgbkCqIWtCR48CLphdOnSNh1Hv3KhY=@vger.kernel.org X-Gm-Message-State: AFuF++nwUeQNz1y7diX3Olhf/W7pfHttnZpPuPpSQaiioev4ZnvTCwz4 B1IIIWamw6CgSvuObU7BM22I+pQN42+TL+3soIRW3ov9K4VLDJUS2EfOx0E17tZjvcM= X-Gm-Gg: AYBFou0TaDSe5y+/RmVx/d/ntgU30jZKLYLI33n+5/MB7Iz3P9kc9cszaEYKNdS+I2C nNc1qnso++A2vt1Fry/F7AhlWY7Ta29VT9erqj5rq4dVNz4eCjFUuwDbJJ96Zlxx24UTPRm5Izw lh4g/7Cf8ECsyk/qvZCKGtvn785h0HOvKZuFdzbJ79jt3dFeZ2g7RzYEdiH3gan1kIKbo/93QrW ux6lp/XuLcRi9he+p1ZDxKEr7J23A3ffR7KgvrR0Wu8fgCl45UvtUAPIIDD2dE28dxL0X35kCnX 8amTH2jgoBCmc1kWVGQ2Q2irxaQErCfkHSvSaY4zzQk23Mr/Q/kDdExLz37xqErx7D8vreD+9Tx 1uIQi2ww8ZGHeDGJ2uDSI9Gr8pgOA8uyAwwRJZZYVNXrLqnxxPxJ8BIYVAvSBv8Kbfw/gfsrz3O PM0dIfHUwRDzJaBTA5tH7jvplVVP1b1hzY2HCNemo4LcZQL22Cr/lOwBNKzJmIttQ0LTOhgXuul 990DA3EbUx8LiShXXG9LUHx5jfMTc+UoL4aIc/WaB8= X-Received: by 2002:a05:6820:190a:b0:6b5:ec3f:4985 with SMTP id 006d021491bc7-6c0a99f860cmr2369093eaf.32.1789141324900; Fri, 11 Sep 2026 08:42:04 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.03 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:04 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 07/15] io_uring: add uring_lock section depth tracking and blockable opdef flag Date: Fri, 11 Sep 2026 09:40:57 -0600 Message-ID: <20260911154148.644489-8-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Prep patch for issuing requests inline in blocking mode and catching the sleep when it happens, rather than punting to io-wq upfront because an operation may block. If a request blocks inline, uring_lock must be dropped on behalf of the sleeping task, which is only safe outside the sections that rely on the lock being held. Track those with a depth counter in io_ring_submit_lock() and io_ring_submit_unlock(). Add a "blockable" flag to io_issue_def for opcodes whose issue path can cope with blocking inline: read/write, the forced async fs ops, open, close and splice/tee. uring_cmd is excluded for now, drivers may bind state to the submitting task. No functional changes in this patch. Signed-off-by: Jens Axboe --- include/linux/io_uring_types.h | 5 +++++ io_uring/io_uring.h | 3 +++ io_uring/opdef.c | 29 +++++++++++++++++++++++++++++ io_uring/opdef.h | 2 ++ 4 files changed, 39 insertions(+) diff --git a/include/linux/io_uring_types.h b/include/linux/io_uring_types.h index 4af3d579ead6..50a4a0ad222f 100644 --- a/include/linux/io_uring_types.h +++ b/include/linux/io_uring_types.h @@ -352,6 +352,11 @@ struct io_ring_ctx { /* submission data */ struct { struct mutex uring_lock; + /* + * io_ring_submit_lock() nesting depth, non-zero means the + * issue path relies on the lock being held. + */ + unsigned int submit_lock_depth; =20 /* * Ring buffer of indices into array of io_uring_sqe, which is diff --git a/io_uring/io_uring.h b/io_uring/io_uring.h index 896aab1ed026..870bb4dcc415 100644 --- a/io_uring/io_uring.h +++ b/io_uring/io_uring.h @@ -393,6 +393,8 @@ static inline void io_ring_submit_unlock(struct io_ring= _ctx *ctx, unsigned issue_flags) { lockdep_assert_held(&ctx->uring_lock); + lockdep_assert(ctx->submit_lock_depth > 0); + ctx->submit_lock_depth--; if (unlikely(issue_flags & IO_URING_F_UNLOCKED)) mutex_unlock(&ctx->uring_lock); } @@ -409,6 +411,7 @@ static inline void io_ring_submit_lock(struct io_ring_c= tx *ctx, if (unlikely(issue_flags & IO_URING_F_UNLOCKED)) mutex_lock(&ctx->uring_lock); lockdep_assert_held(&ctx->uring_lock); + ctx->submit_lock_depth++; } =20 static inline void io_commit_cqring(struct io_ring_ctx *ctx) diff --git a/io_uring/opdef.c b/io_uring/opdef.c index cf3aa2242cd7..fa07a2b94536 100644 --- a/io_uring/opdef.c +++ b/io_uring/opdef.c @@ -69,6 +69,7 @@ const struct io_issue_def io_issue_defs[] =3D { .iopoll =3D 1, .vectored =3D 1, .async_size =3D sizeof(struct io_async_rw), + .blockable =3D 1, .prep =3D io_prep_readv, .issue =3D io_read, }, @@ -83,12 +84,14 @@ const struct io_issue_def io_issue_defs[] =3D { .iopoll =3D 1, .vectored =3D 1, .async_size =3D sizeof(struct io_async_rw), + .blockable =3D 1, .prep =3D io_prep_writev, .issue =3D io_write, }, [IORING_OP_FSYNC] =3D { .needs_file =3D 1, .audit_skip =3D 1, + .blockable =3D 1, .prep =3D io_fsync_prep, .issue =3D io_fsync, }, @@ -101,6 +104,7 @@ const struct io_issue_def io_issue_defs[] =3D { .ioprio =3D 1, .iopoll =3D 1, .async_size =3D sizeof(struct io_async_rw), + .blockable =3D 1, .prep =3D io_prep_read_fixed, .issue =3D io_read_fixed, }, @@ -114,6 +118,7 @@ const struct io_issue_def io_issue_defs[] =3D { .ioprio =3D 1, .iopoll =3D 1, .async_size =3D sizeof(struct io_async_rw), + .blockable =3D 1, .prep =3D io_prep_write_fixed, .issue =3D io_write_fixed, }, @@ -132,6 +137,7 @@ const struct io_issue_def io_issue_defs[] =3D { [IORING_OP_SYNC_FILE_RANGE] =3D { .needs_file =3D 1, .audit_skip =3D 1, + .blockable =3D 1, .prep =3D io_sfr_prep, .issue =3D io_sync_file_range, }, @@ -215,16 +221,19 @@ const struct io_issue_def io_issue_defs[] =3D { [IORING_OP_FALLOCATE] =3D { .needs_file =3D 1, .hash_reg_file =3D 1, + .blockable =3D 1, .prep =3D io_fallocate_prep, .issue =3D io_fallocate, }, [IORING_OP_OPENAT] =3D { .filter_pdu_size =3D sizeof_field(struct io_uring_bpf_ctx, open), + .blockable =3D 1, .prep =3D io_openat_prep, .issue =3D io_openat, .filter_populate =3D io_openat_bpf_populate, }, [IORING_OP_CLOSE] =3D { + .blockable =3D 1, .prep =3D io_close_prep, .issue =3D io_close, }, @@ -236,6 +245,7 @@ const struct io_issue_def io_issue_defs[] =3D { }, [IORING_OP_STATX] =3D { .audit_skip =3D 1, + .blockable =3D 1, .prep =3D io_statx_prep, .issue =3D io_statx, }, @@ -249,6 +259,7 @@ const struct io_issue_def io_issue_defs[] =3D { .ioprio =3D 1, .iopoll =3D 1, .async_size =3D sizeof(struct io_async_rw), + .blockable =3D 1, .prep =3D io_prep_read, .issue =3D io_read, }, @@ -262,17 +273,20 @@ const struct io_issue_def io_issue_defs[] =3D { .ioprio =3D 1, .iopoll =3D 1, .async_size =3D sizeof(struct io_async_rw), + .blockable =3D 1, .prep =3D io_prep_write, .issue =3D io_write, }, [IORING_OP_FADVISE] =3D { .needs_file =3D 1, .audit_skip =3D 1, + .blockable =3D 1, .prep =3D io_fadvise_prep, .issue =3D io_fadvise, }, [IORING_OP_MADVISE] =3D { .audit_skip =3D 1, + .blockable =3D 1, .prep =3D io_madvise_prep, .issue =3D io_madvise, }, @@ -308,6 +322,7 @@ const struct io_issue_def io_issue_defs[] =3D { }, [IORING_OP_OPENAT2] =3D { .filter_pdu_size =3D sizeof_field(struct io_uring_bpf_ctx, open), + .blockable =3D 1, .prep =3D io_openat2_prep, .issue =3D io_openat2, .filter_populate =3D io_openat_bpf_populate, @@ -327,6 +342,7 @@ const struct io_issue_def io_issue_defs[] =3D { .hash_reg_file =3D 1, .unbound_nonreg_file =3D 1, .audit_skip =3D 1, + .blockable =3D 1, .prep =3D io_splice_prep, .issue =3D io_splice, }, @@ -347,6 +363,7 @@ const struct io_issue_def io_issue_defs[] =3D { .hash_reg_file =3D 1, .unbound_nonreg_file =3D 1, .audit_skip =3D 1, + .blockable =3D 1, .prep =3D io_tee_prep, .issue =3D io_tee, }, @@ -360,22 +377,27 @@ const struct io_issue_def io_issue_defs[] =3D { #endif }, [IORING_OP_RENAMEAT] =3D { + .blockable =3D 1, .prep =3D io_renameat_prep, .issue =3D io_renameat, }, [IORING_OP_UNLINKAT] =3D { + .blockable =3D 1, .prep =3D io_unlinkat_prep, .issue =3D io_unlinkat, }, [IORING_OP_MKDIRAT] =3D { + .blockable =3D 1, .prep =3D io_mkdirat_prep, .issue =3D io_mkdirat, }, [IORING_OP_SYMLINKAT] =3D { + .blockable =3D 1, .prep =3D io_symlinkat_prep, .issue =3D io_symlinkat, }, [IORING_OP_LINKAT] =3D { + .blockable =3D 1, .prep =3D io_linkat_prep, .issue =3D io_linkat, }, @@ -387,19 +409,23 @@ const struct io_issue_def io_issue_defs[] =3D { }, [IORING_OP_FSETXATTR] =3D { .needs_file =3D 1, + .blockable =3D 1, .prep =3D io_fsetxattr_prep, .issue =3D io_fsetxattr, }, [IORING_OP_SETXATTR] =3D { + .blockable =3D 1, .prep =3D io_setxattr_prep, .issue =3D io_setxattr, }, [IORING_OP_FGETXATTR] =3D { .needs_file =3D 1, + .blockable =3D 1, .prep =3D io_fgetxattr_prep, .issue =3D io_fgetxattr, }, [IORING_OP_GETXATTR] =3D { + .blockable =3D 1, .prep =3D io_getxattr_prep, .issue =3D io_getxattr, }, @@ -497,6 +523,7 @@ const struct io_issue_def io_issue_defs[] =3D { [IORING_OP_FTRUNCATE] =3D { .needs_file =3D 1, .hash_reg_file =3D 1, + .blockable =3D 1, .prep =3D io_ftruncate_prep, .issue =3D io_ftruncate, }, @@ -553,6 +580,7 @@ const struct io_issue_def io_issue_defs[] =3D { .iopoll =3D 1, .vectored =3D 1, .async_size =3D sizeof(struct io_async_rw), + .blockable =3D 1, .prep =3D io_prep_readv_fixed, .issue =3D io_read, }, @@ -567,6 +595,7 @@ const struct io_issue_def io_issue_defs[] =3D { .iopoll =3D 1, .vectored =3D 1, .async_size =3D sizeof(struct io_async_rw), + .blockable =3D 1, .prep =3D io_prep_writev_fixed, .issue =3D io_write, }, diff --git a/io_uring/opdef.h b/io_uring/opdef.h index 667f981e63b0..45c2f77cf782 100644 --- a/io_uring/opdef.h +++ b/io_uring/opdef.h @@ -29,6 +29,8 @@ struct io_issue_def { unsigned vectored : 1; /* set to 1 if this opcode uses 128b sqes in a mixed sq */ unsigned is_128 : 1; + /* issue path is safe to run inline in blocking mode */ + unsigned blockable : 1; =20 /* size of async data needed, if any */ unsigned short async_size; --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oo2-f43.google.com (mail-oo2-f43.google.com [74.125.231.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A1DDB48EC8C for ; Fri, 11 Sep 2026 15:42:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141331; cv=none; b=FhG89g1FcwuURzAy9+Uzh57HOBtpKTG27Mh3B0EcdtK6+CYzorKm7c6qYxm6DdJsZZE4+DJ2ZFztBPw0L0WCO4njUYSHJZmJYNz/JjWcTrH4kGkvkspX3iHwQSsHyCs576rHnWdF8cERN5WQmZ+/pYhpooWrBCvaty0iGPNcWWo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141331; c=relaxed/simple; bh=yVh0rVZ6GS3pgp+sdjQXZ5uUZjKAgLGO2BaRo77OYf8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=H8/lHF/vE8cm2TpJjuGDGXtqgJjdQvqTfgc6GKvuXM9hGPP1Oi9y5ZsQgIDvO2Oy9dbUJfdKDEzdI4NED/Bj0jh+5oxISB47eWGeSnNSivvQb7NlfzRzlNeoc2Pm3d0SCu2V9EnVYx23T3kVR/x9mGIbQmxGTYxrQd/Xvhf03+M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=etXEfYo4; arc=none smtp.client-ip=74.125.231.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="etXEfYo4" Received: by mail-oo2-f43.google.com with SMTP id 006d021491bc7-6b1ae6c9b81so170111eaf.1 for ; Fri, 11 Sep 2026 08:42:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141326; x=1789746126; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=1mwNlgJ0Ok5ZgfwwPHL43bLSQUr6PFWje+Owka8lGt4=; b=etXEfYo43ilZ0Limel5MTE5uqsKiY6bHcXxumueRC/YW9U9CxektBF0U17It0a09JP PAPM7sWXepSlegDfCq+i97pcOwiLqhZyR6DkOcQjOOZi+XKj0owWjAdTImbARkugJ5ZK R5hyQwOeql6VaThMSBTBV2c7/AktgWy9nx5oeurr0FezoWPiLP936OYM/pNLf9JQGrtX w+bpQuDK2HR/dlQxUH11zuT4N2XxpwMi+wduJ/bEgYbYEoM76EaGIB8yz4j265YeoGTr haPAmJo79uzkeQjCB8IzeZgjA2hu8EY8kP46igWqTBJgpM0jWd9522Bbbub7/5b73qPq gIhA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141326; x=1789746126; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=1mwNlgJ0Ok5ZgfwwPHL43bLSQUr6PFWje+Owka8lGt4=; b=UYUjj+fcBDJ8V7QHZXVSxOY6kXpB7CaoL/tCo/qZUJElvwvFRXWQA4FQglgGfgq695 KNUxlqOeMHlsmYVjwM7QJUCiE34LnY1uU5ZMUUJHjsELlCxDUbqkl7E7Gb3+Xhg89/ky Dd1rEa5zVGLCrsZK/ovnL4DaBW1AyYJpvCTQn0th1N7h2aSqCpA7WuA1tJoXifzTc6E5 rLXmmzuhaXAAumOQxuQxsmJAa2j7ahpIFXATVA77AWGDWHGanMUA4vmH1TWhrLAnw0wC 2c9OMYFLwhfWX5hMOBq5CMtstIuN3HqxCrQ53RndS7CZe3rSaTT+T98uxUJu3bYZOS9H m+HQ== X-Forwarded-Encrypted: i=1; AKwUvByNR1aO0cFEW7q1WwvggSDoTdDazdNR2gJsREmfBjdpVMUfDbr3CPyPpxcmHZij5aqkfOxEzTLIJXMDumg=@vger.kernel.org X-Gm-Message-State: AFuF++n6BoZqwzGLV2vuC/4jVt418iqKLWOZKo46ozxiDx/gO/RNrnyu fbDlc8e0q372LqvXqhYlM7z5Dr/uQrEasnevCU7xZ2CzTDtMx0pPhD/oltE6jqVUmVg= X-Gm-Gg: AYBFou1QuJV7eqmkbBjSuAiVStVcQdys70fnX4eMoLozrYh0j2+dEQkWjB2kqaR6Hsx cZ8Qm7phbUwh4ofP5Pyg2n8s2PE47d7YDCBrsTAdaBdxNmsh1ld7HNGr5UiInf27p2ldgAACNFe IJPG9+8/uHTPXpxXU7mN3kHwMSeZBGVl0TBs07SI08Qnztf1Ea4IAwfXGYHVFxuGk/UhdGv/WAf L0eiM0xjBElCOqp4yXNlr4JzePOChFdCEu9fKYDqcdD/QOKcN31QIq6/EAFJ30CxffLSnkQPUtT WoE1lawZhnOM/2JE5jWlvsOJ4k4HwaZkB4XK0sEwjpiCzRwEsbGmqQBHh1Fzy9ZTU3DeClqZWX3 QrxMhhksfaXpfevJvDBaRCht7Qw4Lihynol5rc3pMbWtSyWR3ubT754J1wRYgQgPBcwAaAHcpsK Z5nnSZr2uXjXRc2ljHrcweNWKU/Wnc1S1RgNZIZUg9xhOc2MQ3iedLsjT+ixBkJ/Vbjat6sa8YI XI6AFQXq2XdB1yJFl1aqKh6nA+1l7p4B9wYtE9V0A8= X-Received: by 2002:a05:6820:c0d9:20b0:6b5:ec3f:497d with SMTP id 006d021491bc7-6c0a8a27f9dmr1999129eaf.24.1789141326313; Fri, 11 Sep 2026 08:42:06 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.05 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:05 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 08/15] io_uring: split io_uring_enter() and io_submit_sqes() into helpers Date: Fri, 11 Sep 2026 09:40:58 -0600 Message-ID: <20260911154148.644489-9-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Move the tail of io_submit_sqes() into io_submit_sqes_end(), and split the IORING_ENTER_GETEVENTS handling out of io_uring_enter() into io_uring_enter_finish() and io_uring_getevents(). A later patch needs to resume io_uring_enter() from a different task than the one that started the syscall. No functional changes in this patch. Signed-off-by: Jens Axboe --- io_uring/io_uring.c | 179 ++++++++++++++++++++++++++++---------------- 1 file changed, 116 insertions(+), 63 deletions(-) diff --git a/io_uring/io_uring.c b/io_uring/io_uring.c index 61053421d809..b09221e239c0 100644 --- a/io_uring/io_uring.c +++ b/io_uring/io_uring.c @@ -2010,12 +2010,31 @@ static bool io_get_sqe(struct io_ring_ctx *ctx, con= st struct io_uring_sqe **sqe) return true; } =20 +static int io_submit_sqes_end(struct io_ring_ctx *ctx, unsigned int entrie= s, + unsigned int left) + __must_hold(&ctx->uring_lock) +{ + int ret =3D entries; + + if (unlikely(left)) { + ret -=3D left; + /* try again if it submitted nothing and can't allocate a req */ + if (!ret && io_req_cache_empty(ctx)) + ret =3D -EAGAIN; + current->io_uring->cached_refs +=3D left; + } + + io_submit_state_end(ctx); + /* Commit SQ ring head once we've consumed and submitted all SQEs */ + io_commit_sqring(ctx); + return ret; +} + int io_submit_sqes(struct io_ring_ctx *ctx, unsigned int nr) __must_hold(&ctx->uring_lock) { unsigned int entries; unsigned int left; - int ret; =20 if (ctx->flags & IORING_SETUP_SQ_REWIND) entries =3D ctx->sq_entries; @@ -2026,7 +2045,7 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigned = int nr) if (unlikely(!entries)) return 0; =20 - ret =3D left =3D entries; + left =3D entries; io_get_task_refs(left); io_submit_state_start(&ctx->submit_state, left); =20 @@ -2052,18 +2071,7 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigned= int nr) } } while (--left); =20 - if (unlikely(left)) { - ret -=3D left; - /* try again if it submitted nothing and can't allocate a req */ - if (!ret && io_req_cache_empty(ctx)) - ret =3D -EAGAIN; - current->io_uring->cached_refs +=3D left; - } - - io_submit_state_end(ctx); - /* Commit SQ ring head once we've consumed and submitted all SQEs */ - io_commit_sqring(ctx); - return ret; + return io_submit_sqes_end(ctx, entries, left); } =20 static void io_rings_free(struct io_ring_ctx *ctx) @@ -2583,6 +2591,94 @@ struct file *io_uring_ctx_get_file(unsigned int fd, = bool registered) } =20 =20 +static int io_iopoll_getevents(struct io_ring_ctx *ctx, u32 min_complete, + u32 flags, const void __user *argp, size_t argsz) + __must_hold(&ctx->uring_lock) +{ + int ret; + + ret =3D io_validate_ext_arg(ctx, flags, argp, argsz); + if (likely(!ret)) + return io_iopoll_check(ctx, min_complete); + return ret; +} + +static int io_wait_getevents(struct io_ring_ctx *ctx, u32 min_complete, + u32 flags, const void __user *argp, size_t argsz) +{ + struct ext_arg ext_arg =3D { .argsz =3D argsz }; + int ret; + + ret =3D io_get_ext_arg(ctx, flags, argp, &ext_arg); + if (likely(!ret)) + return io_cqring_wait(ctx, min_complete, flags, &ext_arg); + return ret; +} + +static int io_getevents_ret(struct io_ring_ctx *ctx, int ret, int ret2) +{ + if (ret) + return ret; + /* + * EBADR indicates that one or more CQE were dropped. Once the user has + * been informed we can clear the bit as they are obviously ok with + * those drops. + */ + if (unlikely(ret2 =3D=3D -EBADR)) + clear_bit(IO_CHECK_CQ_DROPPED_BIT, &ctx->check_cq); + return ret2; +} + +static int io_uring_getevents(struct io_ring_ctx *ctx, int ret, + u32 min_complete, u32 flags, + const void __user *argp, size_t argsz) +{ + int ret2; + + if (ctx->int_flags & IO_RING_F_SYSCALL_IOPOLL) { + /* + * We disallow the app entering submit/complete with polling, + * but we still need to lock the ring to prevent racing with + * polled issue that got punted to a workqueue. + */ + mutex_lock(&ctx->uring_lock); + ret2 =3D io_iopoll_getevents(ctx, min_complete, flags, argp, + argsz); + mutex_unlock(&ctx->uring_lock); + } else { + ret2 =3D io_wait_getevents(ctx, min_complete, flags, argp, argsz); + } + return io_getevents_ret(ctx, ret, ret2); +} + +/* Finish an io_uring_enter() call that submitted and holds the uring_lock= */ +static int io_uring_enter_finish(struct io_ring_ctx *ctx, int ret, u32 min= _complete, + u32 flags, const void __user *argp, size_t argsz) +{ + int ret2; + + if (!(flags & IORING_ENTER_GETEVENTS)) { + mutex_unlock(&ctx->uring_lock); + return ret; + } + + if (ctx->int_flags & IO_RING_F_SYSCALL_IOPOLL) { + ret2 =3D io_iopoll_getevents(ctx, min_complete, flags, argp, + argsz); + mutex_unlock(&ctx->uring_lock); + } else { + /* + * Ignore errors, we'll soon call io_cqring_wait() and it + * should handle ownership problems if any. + */ + if (ctx->flags & IORING_SETUP_DEFER_TASKRUN) + (void)io_run_local_work_locked(ctx, min_complete); + mutex_unlock(&ctx->uring_lock); + ret2 =3D io_wait_getevents(ctx, min_complete, flags, argp, argsz); + } + return io_getevents_ret(ctx, ret, ret2); +} + SYSCALL_DEFINE6(io_uring_enter, unsigned int, fd, u32, to_submit, u32, min_complete, u32, flags, const void __user *, argp, size_t, argsz) @@ -2639,57 +2735,14 @@ SYSCALL_DEFINE6(io_uring_enter, unsigned int, fd, u= 32, to_submit, mutex_unlock(&ctx->uring_lock); goto out; } - if (flags & IORING_ENTER_GETEVENTS) { - if (ctx->int_flags & IO_RING_F_SYSCALL_IOPOLL) - goto iopoll_locked; - /* - * Ignore errors, we'll soon call io_cqring_wait() and - * it should handle ownership problems if any. - */ - if (ctx->flags & IORING_SETUP_DEFER_TASKRUN) - (void)io_run_local_work_locked(ctx, min_complete); - } - mutex_unlock(&ctx->uring_lock); + ret =3D io_uring_enter_finish(ctx, ret, min_complete, flags, + argp, argsz); + goto out; } =20 - if (flags & IORING_ENTER_GETEVENTS) { - int ret2; - - if (ctx->int_flags & IO_RING_F_SYSCALL_IOPOLL) { - /* - * We disallow the app entering submit/complete with - * polling, but we still need to lock the ring to - * prevent racing with polled issue that got punted to - * a workqueue. - */ - mutex_lock(&ctx->uring_lock); -iopoll_locked: - ret2 =3D io_validate_ext_arg(ctx, flags, argp, argsz); - if (likely(!ret2)) - ret2 =3D io_iopoll_check(ctx, min_complete); - mutex_unlock(&ctx->uring_lock); - } else { - struct ext_arg ext_arg =3D { .argsz =3D argsz }; - - ret2 =3D io_get_ext_arg(ctx, flags, argp, &ext_arg); - if (likely(!ret2)) - ret2 =3D io_cqring_wait(ctx, min_complete, flags, - &ext_arg); - } - - if (!ret) { - ret =3D ret2; - - /* - * EBADR indicates that one or more CQE were dropped. - * Once the user has been informed we can clear the bit - * as they are obviously ok with those drops. - */ - if (unlikely(ret2 =3D=3D -EBADR)) - clear_bit(IO_CHECK_CQ_DROPPED_BIT, - &ctx->check_cq); - } - } + if (flags & IORING_ENTER_GETEVENTS) + ret =3D io_uring_getevents(ctx, ret, min_complete, flags, argp, + argsz); out: if (!(flags & IORING_ENTER_REGISTERED_RING)) fput(file); --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-ot1-f52.google.com (mail-ot1-f52.google.com [209.85.210.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D32D84921B7 for ; Fri, 11 Sep 2026 15:42:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141331; cv=none; b=BCwKxa4+i31bIRG0CLGxCfEgCDB4wIqMM6wm3QwWvqICsHdUn5kauwI+rvUhbHxpkZtTnccYM+erIZlzr46m+Sodbq1qGXZF8h/fmB+wXXzBA/G51YgS6tCTXe0dlxwestehPJiCZre1kPjzl7NY289wEavrvgYdBvS/1XOKF9Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141331; c=relaxed/simple; bh=3cOzorvOxRrM2357Kv03cIHM3KvIHiOeSsYnysFbQ18=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=d8ReoOheiwu4IdO/yub2RJg6qwnuLE2m3NasAi6a13EFTvBeKVVjo1aS6t4evG12fAxyBstR8QbdzSwD/Y97Tt+ZkDa4MuCuP4TawE+3qqJ9rUPSnhyFjN+y/SBIHGYfRh4TYpdVlFfb9SAqBClA9bf9jTBX5XfNlj8UqqnqfIM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=f2eYMXof; arc=none smtp.client-ip=209.85.210.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="f2eYMXof" Received: by mail-ot1-f52.google.com with SMTP id 46e09a7af769-7f4f53975e6so1215281a34.3 for ; Fri, 11 Sep 2026 08:42:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141327; x=1789746127; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=0vHIYEHJqme0gI4jLTtintEZJNWn/ID2vfmcQE7LyaU=; b=f2eYMXoftvJtz0vfLTfqNO1EESRZFXkbbiyBBu5RKsswHkgIIXH/RmfzSjNDdbEkRp F/02xB2agpO5uabIwtYoH7iypBMTMbCS3v/C2q2cchAw7OiP3cKxzaoOtDfBWarrtzUM DX8lUpuIzsyZPARK7GwEdUWtISe1C5P7V6w2Vo1oWuwa/dZGMRL5srddUd0r7BZ/tgt2 3ipw5X8HJC93VlgMnOKBUY0V6/g4oaAGcEqYrEZqC4baaqE/9vOy8MczJzCTSb8P0AGz NqadnBu5dYqC+ZeWT1AqCuorqo44nunkhaMaL/Z9E6+rbTLUpPOXuSMygEjsdwL1k7lL eWsg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141327; x=1789746127; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=0vHIYEHJqme0gI4jLTtintEZJNWn/ID2vfmcQE7LyaU=; b=WgVzfiDRZUvEBB/ubw5uxYOL/NJ9Cev2RojQ1qNbokE33GBba3OsjCobGutyDPX5u5 gvjtwy5nL+jm6hjks4iSHRYbGdTePdHURSwQj/XD+YZqUCWtYHcACFdslIAJw/Sy8S5z Mulv9WB1i5755NwSr2cP4rXU2LZ/TRVZB+9MPNc5UJndEg073BijTep5nCgd1iHL1fXJ NEBlHDIjMYXln6EMGSY31P4vTxwKa/Fel/8De/o2l974F+B669Bwno/fQSHC0eDLTZwi sQed/R2ylMgb/2v6OWI/p7Gvm1pKoESBcitKDnu09kqotacLpbYr/Yqip5y8lzIh6lAY kihw== X-Forwarded-Encrypted: i=1; AKwUvBwp6U4WJtlMmW8ks9eoKwJo7WMBI4lSYPA2Lu78/ZZhXTGyrxqd4N08DrvFxJBYPlqhJ6eV0csoFTbLvjg=@vger.kernel.org X-Gm-Message-State: AFuF++nXxWKzk3BjsAW6pB2myfEMdkSE7TC2dbu9sf0q0xp5XPg0ZVoq fYuy3K2eXZm8tq4OjBwDDmzSczpQFnj7ixsRwvbaRW4B+avfkTxVkd5NI3aE3BsiW7o= X-Gm-Gg: AYBFou3WZZzTXbgV+z7gdBqJ52ngQ44D/MvJT9roB7jl9tyvCBnHqbjstHJX7okK0Wf NCgt0W1g8OXL0Qwi5cZ1mX7WPqY62vNA3RiTSwjTX2t+UhKgG2E311mM0zuK/TmBG02jn5EXfo5 DiaZMR10nB9BTEl3nV7Y3Jmp9x89Y21dFk3X+3v8F+Byzh02RTyIUQJBTSisomzwDe2tDJEx2OC s+ZrPoC9482I4rMG9ycqHAftJ3ocPpdKKWJgmrCZqZdPADbaeoxlx2gGTiv+fO/PD4pIyuw7Nu/ HYKr2LM2NsZnGLWj75x9H6Yxi2ukuylHYC/LRg64tynvqKJOLNG6sZdidCtCNNKCBbVWNoQA8AY mQydzE0/MWjj0NTT0/zabVKs+XSfyaGf9+OU805feYjV9FBCsDk7tr7ppQ6Th8h3ogZRUgwdwk6 Rr+1E/KJcvqZa2TgRfjr2TwK0xxQ4Ws+vbvBhyQ1XgFsr8jk3agN2t0FcuTuizDtNOEuDKVNBkJ 41NpNfzWIRZawFM1UadVp8uJ/CZ6eXpi+Uyq6Xl7emt X-Received: by 2002:a05:6820:c0d4:10b0:6be:d716:524d with SMTP id 006d021491bc7-6c0bcea90bcmr2738418eaf.64.1789141327530; Fri, 11 Sep 2026 08:42:07 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:06 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 09/15] io_uring: keep the submission plug on the io_submit_sqes() stack Date: Fri, 11 Sep 2026 09:40:59 -0600 Message-ID: <20260911154148.644489-10-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The submission plug is embedded in the ring's submit state, which only works while a single task runs the whole batch. With a submitter identity handoff another task finishes the batch, and the block layer caches current->plug across sleeps. Put the plug on the stack of io_submit_sqes() instead, it stays with the task that started it. No functional changes in this patch. Signed-off-by: Jens Axboe --- include/linux/io_uring_types.h | 3 ++- io_uring/io_uring.c | 10 +++++++--- 2 files changed, 9 insertions(+), 4 deletions(-) diff --git a/include/linux/io_uring_types.h b/include/linux/io_uring_types.h index 50a4a0ad222f..0b0d73688b8c 100644 --- a/include/linux/io_uring_types.h +++ b/include/linux/io_uring_types.h @@ -298,7 +298,8 @@ struct io_submit_state { bool need_plug; bool cq_flush; unsigned short submit_nr; - struct blk_plug plug; + /* the submitting task's plug, lives on its stack */ + struct blk_plug *plug; }; =20 struct io_alloc_cache { diff --git a/io_uring/io_uring.c b/io_uring/io_uring.c index b09221e239c0..100ade1eee3e 100644 --- a/io_uring/io_uring.c +++ b/io_uring/io_uring.c @@ -1810,7 +1810,7 @@ static int io_init_req(struct io_ring_ctx *ctx, struc= t io_kiocb *req, if (state->need_plug && def->plug) { state->plug_started =3D true; state->need_plug =3D false; - blk_start_plug_nr_ios(&state->plug, state->submit_nr); + blk_start_plug_nr_ios(state->plug, state->submit_nr); } } =20 @@ -1938,13 +1938,14 @@ static void io_submit_state_end(struct io_ring_ctx = *ctx) /* flush only after queuing links as they can generate completions */ io_submit_flush_completions(ctx); if (state->plug_started) - blk_finish_plug(&state->plug); + blk_finish_plug(state->plug); } =20 /* * Start submission side cache. */ static void io_submit_state_start(struct io_submit_state *state, + struct blk_plug *plug, unsigned int max_ios) { state->plug_started =3D false; @@ -1952,6 +1953,8 @@ static void io_submit_state_start(struct io_submit_st= ate *state, state->submit_nr =3D max_ios; /* set only head, no need to init link_last in advance */ state->link.head =3D NULL; + /* on the submitter's stack, the block layer caches current->plug */ + state->plug =3D plug; } =20 static void io_commit_sqring(struct io_ring_ctx *ctx) @@ -2035,6 +2038,7 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigned = int nr) { unsigned int entries; unsigned int left; + struct blk_plug plug; =20 if (ctx->flags & IORING_SETUP_SQ_REWIND) entries =3D ctx->sq_entries; @@ -2047,7 +2051,7 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigned = int nr) =20 left =3D entries; io_get_task_refs(left); - io_submit_state_start(&ctx->submit_state, left); + io_submit_state_start(&ctx->submit_state, &plug, left); =20 do { const struct io_uring_sqe *sqe; --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oa1-f43.google.com (mail-oa1-f43.google.com [209.85.160.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5B737495021 for ; Fri, 11 Sep 2026 15:42:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.43 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141333; cv=none; b=Rc88gK6esrbq/paAmBVtIdAwcquwafRRoEnagfPl4uQhq1gpBwhmZGJSpr70Hj/+se02ZUEGTUvBOgOKkCUAxjDuXhMHmPgc21VXN3aDKOeS2i/BDA7bwQWw2xda8KKOzURZSX+cD7sgEgBzYvnmSjPnINcPmWEUGyQR9qYIn7Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141333; c=relaxed/simple; bh=JITu/mXsVAuxIfuBVo9KN86BDNBFb2HVXwMUtq89zcA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=OYDjXWARqEdEt+U3StgIr6crSOGZ2ASLTVZ3veoawjaa+MfHps/f9grB74rQwvFGIb542gO3FrdFdAUL7xWD6gV/GBBtzG1Zu/DipLbi5TD0cB/b5dxNswaYA/ccTXX/yTy8h44t1tNX2abov2jQzasmALszZXN7c0cVsUQ82Qg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=P6DxBvFt; arc=none smtp.client-ip=209.85.160.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="P6DxBvFt" Received: by mail-oa1-f43.google.com with SMTP id 586e51a60fabf-455ca262ccbso915867fac.1 for ; Fri, 11 Sep 2026 08:42:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141329; x=1789746129; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=RAOeNgcoqL33KiDySFvl0CYTM4qzSmTk0iyHPZgyOXc=; b=P6DxBvFtv2TUFVYPhJyv0S4KCgQ1NX4oHW6db56ScSAu2rnwW01rxg8DU3fbMQLY/r 7EFefOHCQuR7IxZbA8fcvvHpReJe5cyEvhR9z0AVHPZnoC/mhdE7DppcXNC5EvknW2gq 1kJJgoYYDE7VvkNz+78txws9//N2wGc7RcZZwTM1UzrjCVFldGP4a0Y5KwykMo+HpnS4 oqYloEIGCH0XU+/M0Qx7xSmoslVDer4kKyJIg+h/A9MNGHr82hcIMkPj5n3gPeI6aCHN OT9l6iPOVr2ar1EV7pyvZiYCbP5Z7SPP/XXXFl4L4kZZ8j38QfDeIIKfo/C9EiWa7nOw fe7A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141329; x=1789746129; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=RAOeNgcoqL33KiDySFvl0CYTM4qzSmTk0iyHPZgyOXc=; b=QfKK+fjJc8JHBmZ0RckWRdp8f8vCccJ2wA+jxHqk/I62VwBheyoulrATDt4LhZNR5z 401k4FYZJlUSsggo04/A4OqodHMR9x1QroRB06MAIbSaSn54N4ayNfrLkxewc5SJs1h3 /6VDO5f/YCKKuXMGHgARSXjw2ZG+mvckSVoV1HrM0GLvSygXPiREtRFo+hy5Mj1V0+O4 xjeGegtZ+vvAzMCKG6AvuUzMa6obqwOpudpEff0ajWEHp9gpBctyvqunTX/2vt3y9K/b FBbTzroHtVHrr51h76sJIGttfT865TnEW/G3N77+Q7u2KwFIiWdqRoa3SGvMll0LgCt6 U5wA== X-Forwarded-Encrypted: i=1; AKwUvBxfZHvvpx+5mjx0XgUBKNCmA+j5BAEekbeYSDuoT7+qus+ns6DftqaCSvW7XMeoqPyepjZhc88oGfQaHcw=@vger.kernel.org X-Gm-Message-State: AFuF++mSKkX7wyqobJaKGpeLkB/Bz7n4udP6htDATwKss7JaCiROliLM SF0PMqyVG5uti3UyYMhHqZ8AYuHDxeTQ3N7f67vZuelKT9Y/zr5Id/8MZd5WVSdZdN25XN1wF0n jx1xP7zo= X-Gm-Gg: AYBFou0xOcTnKXCJC2xV2ekybJLXaFGuPDDGVV577vibpp84lMqW5gjJNF4dy4sNy9j ocN4b2sPxLdA3X5p9pHH5dhMxXO0uIbyGlLXOGKIDENxU+w1CvuN/g7rxBwTHPur676X9g0d5Xz heODt/DYXsKpa1YwW1PxqHc5fiq+NT6rsSqsGFsJWvCZTfVN0rbkLrdklXQJNkuW+mby5LNAsnL S5fWSwzldW127KTl97X+lPs1hM3LxPEmuCSZ0ivWoMf0YwIzcBc/4srJu4ctbNQO+5nIsLpwpLX ZCOmiketKh9w16o5Aw93PrSvJe7jpAS82nXPUQZ22UnQ4QZx22QCX4i19M5i0o1/eda+p4wR710 zjinMMfB+6LJ8YjdU9dHZ5yEO3YuP8lsRElTw32AwDh5ZJRmYJNLJARW/+GlcOCdCwRLkNtlpge LLUdFilJZi+LundfEjY0XKZek1oCO7WN3kdzK+s+KYwhN2j68bV74OWPNpNPBj9jrh49qpXusdk bZfqwKBtDoiwDY1nb5GOAThFwVd3RMJuGQ+I2sSJhg= X-Received: by 2002:a05:6820:208d:b0:6c0:14e9:1850 with SMTP id 006d021491bc7-6c0bae0b281mr3497013eaf.21.1789141328879; Fri, 11 Sep 2026 08:42:08 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.07 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:08 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 10/15] io-wq: support handing a task identity to an idle worker Date: Fri, 11 Sep 2026 09:41:00 -0600 Message-ID: <20260911154148.644489-11-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add the io-wq side of handing a blocked submitter's identity to an idle worker. io_wq_handoff_claim() picks an idle worker that can take over the identity of the current task, and arranges for it to run a caller supplied function when it wakes instead of continuing as a worker. Only a worker inside the idle sleep of the worker loop is claimable, being on the free list isn't enough as it may be sleeping elsewhere. This is tracked in acct->nr_iosleep. Redundant workers exit through the normal idle timeout. Signed-off-by: Jens Axboe --- io_uring/io-wq.c | 263 ++++++++++++++++++++++++++++++++++++++++++++--- io_uring/io-wq.h | 14 +++ kernel/fork.c | 3 +- 3 files changed, 267 insertions(+), 13 deletions(-) diff --git a/io_uring/io-wq.c b/io_uring/io-wq.c index 2ca223e47d41..3d4eb4992d5b 100644 --- a/io_uring/io-wq.c +++ b/io_uring/io-wq.c @@ -20,6 +20,7 @@ #include #include #include +#include =20 #include "io-wq.h" #include "slist.h" @@ -32,6 +33,15 @@ enum { IO_WORKER_F_UP =3D 0, /* up and active */ IO_WORKER_F_RUNNING =3D 1, /* account as running */ IO_WORKER_F_FREE =3D 2, /* worker on free list */ + IO_WORKER_F_IDLE_SLEEP =3D 3, /* in the idle sleep of the worker loop */ +}; + +/* worker->handoff handshake */ +enum { + IO_WORKER_HANDOFF_NONE =3D 0, + IO_WORKER_HANDOFF_PROMOTE, /* claimed, set under ->workers_lock */ + IO_WORKER_HANDOFF_DONE, /* claimer took over the worker */ + IO_WORKER_HANDOFF_FINISHED, /* identity moved, demoted released */ }; =20 enum { @@ -64,6 +74,10 @@ struct io_worker { struct callback_head create_work; int init_retries; =20 + atomic_t handoff; + io_wq_handoff_fn *handoff_fn; + struct task_struct *handoff_task; + union { struct rcu_head rcu; struct delayed_work work; @@ -88,6 +102,9 @@ struct io_wq_acct { unsigned max_workers; atomic_t nr_running; =20 + /* workers in the idle sleep, claimable for a handoff */ + atomic_t nr_iosleep; + /** * The list of free workers. Protected by #workers_lock * (write) and RCU (read). @@ -151,6 +168,8 @@ static bool io_acct_cancel_pending_work(struct io_wq *w= q, struct io_wq_acct *acct, struct io_cb_cancel_data *match); static void create_worker_cb(struct callback_head *cb); +static void create_worker_cont(struct callback_head *cb); +static bool io_task_work_match(struct callback_head *cb, void *data); static void io_wq_cancel_tw_create(struct io_wq *wq); =20 static inline unsigned int __io_get_work_hash(unsigned int work_flags) @@ -233,7 +252,7 @@ static bool io_task_worker_match(struct callback_head *= cb, void *data) return worker =3D=3D data; } =20 -static void io_worker_exit(struct io_worker *worker) +static void __noreturn io_worker_exit(struct io_worker *worker) { struct io_wq *wq =3D worker->wq; struct io_wq_acct *acct =3D io_wq_get_acct(worker); @@ -411,7 +430,7 @@ static bool io_queue_worker_create(struct io_worker *wo= rker, =20 atomic_inc(&wq->worker_refs); init_task_work(&worker->create_work, func); - if (!task_work_add(wq->task, &worker->create_work, TWA_SIGNAL)) { + if (!io_wq_task_work_add(wq->task, &worker->create_work, TWA_SIGNAL)) { /* * EXIT may have been set after checking it above, check after * adding the task_work and remove any creation item if it is @@ -687,19 +706,48 @@ static void io_worker_handle_work(struct io_wq_acct *= acct, } while (1); } =20 -static int io_wq_worker(void *data) +/* + * Leave the idle sleep, check if we got claimed while in it. Serialized w= ith + * the claimer by ->workers_lock, so a claim can't be missed or raced. + */ +static bool io_wq_worker_idle_done(struct io_wq_acct *acct, + struct io_worker *worker) { - struct io_worker *worker =3D data; - struct io_wq_acct *acct =3D io_wq_get_acct(worker); - struct io_wq *wq =3D worker->wq; - bool exit_mask =3D false, last_timeout =3D false; - char buf[TASK_COMM_LEN] =3D {}; + bool promoted; =20 - set_mask_bits(&worker->flags, 0, - BIT(IO_WORKER_F_UP) | BIT(IO_WORKER_F_RUNNING)); + atomic_dec(&acct->nr_iosleep); + raw_spin_lock(&acct->workers_lock); + clear_bit(IO_WORKER_F_IDLE_SLEEP, &worker->flags); + /* the claimer may have committed already, DONE rather than PROMOTE */ + promoted =3D atomic_read(&worker->handoff) !=3D IO_WORKER_HANDOFF_NONE; + raw_spin_unlock(&acct->workers_lock); =20 - snprintf(buf, sizeof(buf), "iou-wrk-%d", wq->task->pid); - set_task_comm(current, buf); + WARN_ON_ONCE(promoted && worker->handoff_task !=3D current); + return promoted; +} + +/* claimed, wait for the claimer to finish taking over the worker struct */ +static io_wq_handoff_fn *io_wq_worker_promoted(struct io_worker *worker) +{ + io_wq_handoff_fn *fn =3D worker->handoff_fn; + + /* stop the scheduler from treating us as a worker while we wait */ + current->flags &=3D ~(PF_IO_WORKER | PF_USER_WORKER); + + wait_var_event(&worker->handoff, + atomic_read_acquire(&worker->handoff) =3D=3D IO_WORKER_HANDOFF_DONE); + + WARN_ON_ONCE(current->worker_private); + atomic_set(&worker->handoff, IO_WORKER_HANDOFF_NONE); + return fn; +} + +/* the worker loop, only returns if claimed for a handoff */ +static io_wq_handoff_fn *io_wq_worker_run(struct io_worker *worker) +{ + struct io_wq_acct *acct =3D io_wq_get_acct(worker); + bool exit_mask =3D false, last_timeout =3D false; + struct io_wq *wq =3D worker->wq; =20 while (!test_bit(IO_WQ_BIT_EXIT, &wq->state)) { long ret; @@ -724,6 +772,11 @@ static int io_wq_worker(void *data) if ((last_timeout && (exit_mask || acct->nr_workers > 1)) || test_bit(IO_WQ_BIT_EXIT_ON_IDLE, &wq->state)) { acct->nr_workers--; + /* a handoff can't claim an exiting worker */ + if (test_bit(IO_WORKER_F_FREE, &worker->flags)) { + clear_bit(IO_WORKER_F_FREE, &worker->flags); + hlist_nulls_del_rcu(&worker->nulls_node); + } raw_spin_unlock(&acct->workers_lock); __set_current_state(TASK_RUNNING); break; @@ -733,7 +786,12 @@ static int io_wq_worker(void *data) raw_spin_unlock(&acct->workers_lock); if (io_run_task_work()) continue; + /* claimable only in this sleep, the free list isn't enough */ + set_bit(IO_WORKER_F_IDLE_SLEEP, &worker->flags); + atomic_inc(&acct->nr_iosleep); ret =3D schedule_timeout(WORKER_IDLE_TIMEOUT); + if (unlikely(io_wq_worker_idle_done(acct, worker))) + return io_wq_worker_promoted(worker); if (signal_pending(current)) { struct ksignal ksig; =20 @@ -752,9 +810,190 @@ static int io_wq_worker(void *data) io_worker_handle_work(acct, worker); =20 io_worker_exit(worker); +} + +static int io_wq_worker(void *data) +{ + struct io_worker *worker =3D data; + struct io_wq *wq =3D worker->wq; + io_wq_handoff_fn *fn; + char buf[TASK_COMM_LEN] =3D {}; + + set_mask_bits(&worker->flags, 0, + BIT(IO_WORKER_F_UP) | BIT(IO_WORKER_F_RUNNING)); + + snprintf(buf, sizeof(buf), "iou-wrk-%d", wq->task->pid); + set_task_comm(current, buf); + + /* + * Only returns if we got handed an identity. -EIOCBQUEUED means we got + * demoted again while running it, back to the worker loop. + */ + for (;;) { + long ret; + + fn =3D io_wq_worker_run(worker); + ret =3D fn(); + /* what we return is what userspace gets on some archs */ + if (ret !=3D -EIOCBQUEUED) + return ret; + worker =3D current->worker_private; + } +} + +/* find and claim an idle sleeping worker, see io_wq_worker_idle_done() */ +static struct io_worker *io_wq_acct_handoff_claim(struct io_wq *wq, + struct io_wq_acct *acct, + io_wq_handoff_fn *fn) +{ + struct io_worker *worker, *found =3D NULL; + struct hlist_nulls_node *n; + + raw_spin_lock(&acct->workers_lock); + if (test_bit(IO_WQ_BIT_EXIT, &wq->state)) + goto out_unlock; + hlist_nulls_for_each_entry(worker, n, &acct->free_list, nulls_node) { + /* only claimable inside the idle sleep of the worker loop */ + if (!test_bit(IO_WORKER_F_IDLE_SLEEP, &worker->flags)) + continue; + if (!thread_handoff_compatible(current, worker->task)) + continue; + clear_bit(IO_WORKER_F_FREE, &worker->flags); + hlist_nulls_del_init_rcu(&worker->nulls_node); + worker->handoff_fn =3D fn; + worker->handoff_task =3D worker->task; + atomic_set_release(&worker->handoff, IO_WORKER_HANDOFF_PROMOTE); + found =3D worker; + break; + } +out_unlock: + raw_spin_unlock(&acct->workers_lock); + if (found) + wake_up_process(found->handoff_task); + return found; +} + +/* + * task_work_add() that doesn't notify a task inside a blocking inline iss= ue, + * it'd interrupt the sleep. io_handoff_end() picks pending work up instea= d. + */ +int io_wq_task_work_add(struct task_struct *task, struct callback_head *cb, + enum task_work_notify_mode notify) +{ + int ret; + + if (notify !=3D TWA_SIGNAL && notify !=3D TWA_SIGNAL_NO_IPI) + return task_work_add(task, cb, notify); + + ret =3D task_work_add(task, cb, TWA_NONE); + if (ret) + return ret; + if (READ_ONCE(task->flags) & PF_IO_HANDOFF) + return 0; + if (notify =3D=3D TWA_SIGNAL) + set_notify_signal(task); + else + __set_notify_signal(task); return 0; } =20 +/* claim an idle worker to hand our identity to, pairs with _commit() */ +struct task_struct *io_wq_handoff_claim(struct io_wq *wq, bool bound, + io_wq_handoff_fn *fn) +{ + struct io_worker *worker; + + worker =3D io_wq_acct_handoff_claim(wq, io_get_acct(wq, bound), fn); + if (!worker) + worker =3D io_wq_acct_handoff_claim(wq, io_get_acct(wq, !bound), fn); + if (worker) + return worker->task; + return NULL; +} + +/* we take over the worker @dst was, @dst goes on to run the handoff fn */ +void io_wq_handoff_commit(struct task_struct *dst) +{ + struct io_worker *worker =3D dst->worker_private; + struct task_struct *src =3D current; + struct io_wq *wq =3D worker->wq; + struct callback_head *cb; + + WARN_ON_ONCE(src->worker_private); + WARN_ON_ONCE(worker->task !=3D dst); + WARN_ON_ONCE(wq->task !=3D src); + + /* the sched hooks around @dst's wakeup cope with NULL worker_private */ + WRITE_ONCE(dst->worker_private, NULL); + src->worker_private =3D worker; + WRITE_ONCE(worker->task, src); + src->flags |=3D PF_IO_WORKER | PF_USER_WORKER; + + /* the user task owns the wq */ + get_task_struct(dst); + WRITE_ONCE(wq->task, dst); + + /* move pending worker creations along, we may block for a while */ + while ((cb =3D task_work_cancel_match(src, io_task_work_match, wq))) { + struct io_worker *w =3D container_of(cb, struct io_worker, + create_work); + + if (!task_work_add(dst, cb, TWA_SIGNAL)) + continue; + io_worker_cancel_cb(w); + if (cb->func =3D=3D create_worker_cont) + kfree(w); + } + put_task_struct(src); + + atomic_set_release(&worker->handoff, IO_WORKER_HANDOFF_DONE); + wake_up_var(&worker->handoff); +} + +/* worker loop entry for a demoted task, only returns on another handoff */ +io_wq_handoff_fn *io_wq_handoff_worker(void) +{ + struct io_worker *worker =3D current->worker_private; + char buf[TASK_COMM_LEN] =3D {}; + + WARN_ON_ONCE(!io_wq_current_is_worker()); + + /* the promoted task reads our state until it's done migrating it */ + wait_var_event(&worker->handoff, + atomic_read_acquire(&worker->handoff) =3D=3D IO_WORKER_HANDOFF_FINISHED); + atomic_set(&worker->handoff, IO_WORKER_HANDOFF_NONE); + + snprintf(buf, sizeof(buf), "iou-wrk-%d", worker->wq->task->pid); + set_task_comm(current, buf); + set_cpus_allowed_ptr(current, worker->wq->cpu_mask); + + return io_wq_worker_run(worker); +} + +/* release the demoted task @tsk to run as the worker it now is */ +void io_wq_handoff_finished(struct task_struct *tsk) +{ + struct io_worker *worker =3D tsk->worker_private; + + atomic_set_release(&worker->handoff, IO_WORKER_HANDOFF_FINISHED); + wake_up_var(&worker->handoff); +} + +/* idle sleepers to keep around as handoff targets */ +#define IO_WQ_HANDOFF_SPARES 2 + +/* true if a worker is claimable, @topup forks one if below the target */ +bool io_wq_handoff_spare(struct io_wq *wq, bool bound, bool topup) +{ + struct io_wq_acct *acct =3D io_get_acct(wq, bound); + unsigned int idle; + + idle =3D atomic_read(&acct->nr_iosleep); + if (topup && idle < IO_WQ_HANDOFF_SPARES) + io_wq_create_worker(wq, acct); + return idle > 0; +} + /* * Called when a worker is scheduled in. Mark us as currently running. */ diff --git a/io_uring/io-wq.h b/io_uring/io-wq.h index 42f00a47a9c9..98357b665e54 100644 --- a/io_uring/io-wq.h +++ b/io_uring/io-wq.h @@ -4,6 +4,7 @@ =20 #include #include +#include =20 struct io_wq; =20 @@ -47,6 +48,19 @@ void io_wq_set_exit_on_idle(struct io_wq *wq, bool enabl= e); void io_wq_enqueue(struct io_wq *wq, struct io_wq_work *work); void io_wq_hash_work(struct io_wq_work *work, void *val); =20 +typedef long (io_wq_handoff_fn)(void); + +/* claim an idle worker, it runs @fn instead of the worker loop when woken= */ +struct task_struct *io_wq_handoff_claim(struct io_wq *wq, bool bound, + io_wq_handoff_fn *fn); + +void io_wq_handoff_commit(struct task_struct *dst); +io_wq_handoff_fn *io_wq_handoff_worker(void); +bool io_wq_handoff_spare(struct io_wq *wq, bool bound, bool topup); +void io_wq_handoff_finished(struct task_struct *tsk); +int io_wq_task_work_add(struct task_struct *task, struct callback_head *cb, + enum task_work_notify_mode notify); + int io_wq_cpu_affinity(struct io_uring_task *tctx, cpumask_var_t mask); int io_wq_max_workers(struct io_wq *wq, int *new_count); bool io_wq_worker_stopped(void); diff --git a/kernel/fork.c b/kernel/fork.c index 510c8a9aa870..f31af9c5bae7 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -2688,11 +2688,12 @@ struct task_struct * __init fork_idle(int cpu) * creating io_uring workers. It returns a created task, or an error point= er. * The returned task is inactive, and the caller must fire it up through * wake_up_new_task(p). All signals are blocked in the created task. + * CLONE_SYSVSEM as a worker may take over a user thread's identity. */ struct task_struct *create_io_thread(int (*fn)(void *), void *arg, int nod= e) { unsigned long flags =3D CLONE_FS|CLONE_FILES|CLONE_SIGHAND|CLONE_THREAD| - CLONE_IO|CLONE_VM|CLONE_UNTRACED; + CLONE_IO|CLONE_VM|CLONE_UNTRACED|CLONE_SYSVSEM; struct kernel_clone_args args =3D { .flags =3D flags, .fn =3D fn, --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oo1-f53.google.com (mail-oo1-f53.google.com [209.85.161.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3B5B049690C for ; Fri, 11 Sep 2026 15:42:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.53 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141337; cv=none; b=lJymKm8aWudE1WYQtU++SGWdTM5MlVUepzqOJWubFHjv/uMK33yDj3cNNgVs3aACmJjUn6rBzQNT+lhYSTBh2k00zy4jYmSv6YiS/c1DYQmQJ40LlqGVkeJRg5zmYjEPN5NcY4c5Ttkr0uUG702T8Dol3Z4kECIZ1icjGDTokoE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141337; c=relaxed/simple; bh=xWUtciFBYWJsNLzGvaU/kzObElfZHwwMSlHjuUQOHF4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dUYMCJuPGEOgtZVTqSLXaiRNyUnwIB+rYRxtBSTP4LbNYC9PvsXwieCj5ul6/13ijiLsNpwDuU+2YLq4AiLU/u2rP1TTADYNf+Bz/dGvfqHgbogSfXVhJNChYrGz7SJX+AImNgjHnlPp44f2zcO1P8Ye4TpaKEmQWgF5C7+DDso= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=F3r5Y288; arc=none smtp.client-ip=209.85.161.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="F3r5Y288" Received: by mail-oo1-f53.google.com with SMTP id 006d021491bc7-6c16080f8bcso379123eaf.1 for ; Fri, 11 Sep 2026 08:42:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141331; x=1789746131; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=RKc5Jg/Dnuhi3Awy6UWcgXwLQtmBLJceWKQ5shu2ooU=; b=F3r5Y288EPyoU8E9/5Xh1sMm73uSoNPFe/DHV+iFSJrt61Mnz18n1JE50izwVHRMD/ WzC6ZJT9PDVnskpyWddmjaGh1a0CvCLSIyVpHd7j6grhmM6WU7WRbKyxSYImhZ7uBdk+ gqb8G5ZNUYxvilbUp5BuhGw0TnYXwHf6/evn38rMIFFe7FN+M0Mopl1grt/ZOcKvBTRd lYQdghSubSaAYfy16Mp9tveZ/f2GQGWi7Y+lY9fju0pXLVGKXKEyRa/JwEG5h9alS2Xf F3nZZtWI7WYr3jnLJdJgmfOHSksOsNtIINygwWaU6grxOyKCR7kHLv9BoIFoKogk7fZX 7t8w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141331; x=1789746131; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=RKc5Jg/Dnuhi3Awy6UWcgXwLQtmBLJceWKQ5shu2ooU=; b=S0eARwWZwP443/E45BfW5iVZB7rgaaWcrKCUL4CXbvJPEcUD4vn/B+KirW3MRQdLvQ x8FeLgvyET4OuziLRktQYnF0R1JELZLBQm/VqtsfsgrEZISWAq9eVFORjVPr+jya/sMs ee8eTYkxT1phk7bp+97LZ4Hu28mpyY357auVqcyQqwWRZAlDvWXFWGM7WRHjtJZEUHod UxhjQfWbYnY65Lol4Jus3fhiq2/E88h7GIaQLp/WP8X8nP0Fq9ajBcG4RoXedt/ntii7 VRF5ximeSblwtMuP3HNq6jlnZeJCo/7ETg7IMvUL47H5dv8PVXyMns74cDR4xMt91fci gHgQ== X-Forwarded-Encrypted: i=1; AKwUvBxE9v2mNJATIp+KO5pkmJZmDJgblKfMDBhMgFx5r735w7G0yoRKPxAsVSHgH0D1cWnKukGk+XYosIePIsg=@vger.kernel.org X-Gm-Message-State: AFuF++lnI7FGxKjuzQKwysceJzqyGLw2Rkb4F8N+DhQutoHS6e4c2yBL Se0K2hhseGCpsoL7QSwHhMmrFr7qMxBUU/DbEZ3sWaipBbAH2ZwhJnZGjVgMSvIBJ54= X-Gm-Gg: AYBFou1Nfo9dmhuT+9jl0G4GYe8yg0cwhpZepNGMm/mtgRSURhg1APTdbj3o8MtPgyv zdZUo2ztFRHYUEb8UzIH+hspNisxFBKxKDp9AOsWmoChqEiuaxHzU7Ibl8NIHgHXUYkaYKGmRQZ 0LTuGd05DeIGyFmH8WcmrEs3WzCeY8GhdpozydDXwq0ZzXENZK/ZTTxvexNNjUVoJ48keaz/k2I /Nfk5YoPfQrDRCgOHIR+6P4YvdNBiRj+Ws6kO/PCHGVH3/EZf4BFRNp5vSfU9RFzJZ7V2dw9Yt5 6bJ2Ka9qx130pCMJZN9mjrAd9bxbutr2H3YfqPFXQ36MUp1vW/ouv9LbHs/LVq5mw1RxQD/Po6Z 1OkERPSgngnKbtr5Hhy/aST73ko2ixCWN9oupxYEOXO1zKOfj/XQ/iHzucUQTZuhcsYQZQmdVdG SzbTD7nggggaK4F2+dLVeXXwZJ8BFp+GJw4621m4ewreF+h+ft2wV8jWekr2t+H9DQOFZOHiBio +0eh3DTG5CeLxPRcp4HCrGsZc1y X-Received: by 2002:a05:6820:3414:10b0:6b3:4b80:3100 with SMTP id 006d021491bc7-6c0ba92040dmr2407722eaf.23.1789141330439; Fri, 11 Sep 2026 08:42:10 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:09 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 11/15] io_uring: enable handing submitter identity to an io-wq worker Date: Fri, 11 Sep 2026 09:41:01 -0600 Message-ID: <20260911154148.644489-12-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" If a blockable request issued inline from io_uring_enter() blocks, hand the submitter's identity to an idle io-wq worker. The worker finishes the io_uring_enter() call and returns to userspace as the submitter, while the original task finishes the request and lives on as the worker. Requests are still issued with IO_URING_F_NONBLOCK, but any sleep they hit anyway (page fault, lock, allocation stall) no longer blocks the submitter. Signed-off-by: Jens Axboe --- include/linux/io_uring.h | 4 + include/linux/io_uring_types.h | 34 ++++ io_uring/Makefile | 1 + io_uring/handoff.c | 326 +++++++++++++++++++++++++++++++++ io_uring/handoff.h | 102 +++++++++++ io_uring/io_uring.c | 84 +++++++-- io_uring/io_uring.h | 36 +++- io_uring/msg_ring.c | 2 +- io_uring/rw.c | 2 +- io_uring/tctx.c | 6 +- io_uring/tw.c | 14 +- io_uring/uring_cmd.c | 5 +- 12 files changed, 596 insertions(+), 20 deletions(-) create mode 100644 io_uring/handoff.c create mode 100644 io_uring/handoff.h diff --git a/include/linux/io_uring.h b/include/linux/io_uring.h index 969de22c3d0f..4137276ba8f1 100644 --- a/include/linux/io_uring.h +++ b/include/linux/io_uring.h @@ -61,8 +61,12 @@ static inline int io_uring_fork(struct task_struct *tsk) #endif =20 /* called from sched_submit_work() when a PF_IO_HANDOFF task blocks */ +#if defined(CONFIG_IO_URING) && defined(CONFIG_THREAD_HANDOFF) +void io_uring_task_sleeping(struct task_struct *tsk); +#else static inline void io_uring_task_sleeping(struct task_struct *tsk) { } +#endif =20 #endif diff --git a/include/linux/io_uring_types.h b/include/linux/io_uring_types.h index 0b0d73688b8c..81bc4810fcab 100644 --- a/include/linux/io_uring_types.h +++ b/include/linux/io_uring_types.h @@ -4,9 +4,11 @@ #include #include #include +#include #include #include #include +#include #include =20 struct iou_loop_params; @@ -139,12 +141,42 @@ struct io_br_sel { */ #define IO_RINGFD_REG_MAX 16 =20 +/* handoff state of a submitter that blocked inline, see io_uring/handoff.= c */ +struct io_handoff { + /* the request being issued, while a handoff is possible */ + struct io_kiocb *req; + /* its ring and io-wq pool, @req is the demoted task's after that */ + struct io_ring_ctx *ctx; + bool bound; + /* the submitter's signal mask while blocking issues run without */ + sigset_t sigmask; + bool sigsaved; + /* the task the identity came from, and the task refs it held */ + struct task_struct *src; + unsigned int src_refs; + struct thread_handoff_stats stats; + /* io_uring_enter() arguments, to resume the syscall */ + struct file *file; + u32 to_submit; + /* SQEs consumed so far by this syscall, across handoffs */ + u32 consumed; + u32 min_complete; + u32 flags; + const void __user *argp; + size_t argsz; +}; + struct io_uring_task { /* submission side */ int cached_refs; const struct io_ring_ctx *last; struct task_struct *task; + /* serializes ->task changes against off-task reference puts */ + raw_spinlock_t task_ref_lock; struct io_wq *io_wq; +#ifdef CONFIG_THREAD_HANDOFF + struct io_handoff handoff; +#endif /* * Consumer cursor for ->task_list. Only popped by the task itself, * or by ->fallback_work once the task can no longer run task_work. @@ -298,6 +330,8 @@ struct io_submit_state { bool need_plug; bool cq_flush; unsigned short submit_nr; + /* cached SQ head at the start of the batch */ + unsigned int sq_head; /* the submitting task's plug, lives on its stack */ struct blk_plug *plug; }; diff --git a/io_uring/Makefile b/io_uring/Makefile index c54e328d1410..9fd99118d26a 100644 --- a/io_uring/Makefile +++ b/io_uring/Makefile @@ -18,6 +18,7 @@ obj-$(CONFIG_IO_URING) +=3D io_uring.o opdef.o kbuf.o rs= rc.o notif.o \ =20 obj-$(CONFIG_IO_URING_ZCRX) +=3D zcrx.o obj-$(CONFIG_IO_WQ) +=3D io-wq.o +obj-$(CONFIG_THREAD_HANDOFF) +=3D handoff.o obj-$(CONFIG_FUTEX) +=3D futex.o obj-$(CONFIG_EPOLL) +=3D epoll.o obj-$(CONFIG_NET_RX_BUSY_POLL) +=3D napi.o diff --git a/io_uring/handoff.c b/io_uring/handoff.c new file mode 100644 index 000000000000..25e9e06812a7 --- /dev/null +++ b/io_uring/handoff.c @@ -0,0 +1,326 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Inline issue handoff: if an inline issue blocks, the submitter's identi= ty + * moves to an idle io-wq worker which returns to userspace as the submitt= er, + * while the submitter finishes the request as the worker. + * + * Copyright (C) 2026 Jens Axboe + */ +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "io_uring.h" +#include "io-wq.h" +#include "opdef.h" +#include "tctx.h" +#include "handoff.h" + +int sysctl_io_uring_handoff __read_mostly =3D 1; + +static long io_handoff_resume(void); + +/* fork a spare worker upfront, so the first blockable issue has a target = */ +void io_handoff_prime(struct io_uring_task *tctx, struct io_ring_ctx *ctx) +{ + if (!sysctl_io_uring_handoff || !tctx->io_wq) + return; + /* handoffs are never done for these, see __io_handoff_begin() */ + if (ctx->flags & (IORING_SETUP_IOPOLL | IORING_SETUP_SQPOLL | + IORING_SETUP_SQ_REWIND)) + return; + io_wq_handoff_spare(tctx->io_wq, true, true); +} + +/* + * Blocking inline issues run with an io-wq worker's signal mask, a request + * shouldn't fail with -EINTR because the submitter has a timer. + */ +static void io_handoff_block_signals(struct io_handoff *ho) +{ + sigset_t mask; + + if (ho->sigsaved) + return; + ho->sigsaved =3D true; + ho->sigmask =3D current->blocked; + siginitsetinv(&mask, sigmask(SIGKILL) | sigmask(SIGSTOP)); + set_current_blocked(&mask); +} + +void __io_handoff_restore_signals(struct io_handoff *ho) +{ + ho->sigsaved =3D false; + set_current_blocked(&ho->sigmask); +} + +/* + * Arm a handoff for the inline issue of @req. Until io_handoff_end() the = issue + * runs like on io-wq, neither normal signals nor task_work interrupt it. + */ +bool __io_handoff_begin(struct io_kiocb *req) +{ + struct io_ring_ctx *ctx =3D req->ctx; + struct io_uring_task *tctx =3D current->io_uring; + struct io_handoff *ho =3D &tctx->handoff; + + if (!sysctl_io_uring_handoff) + return false; + /* nonblocking semantics were asked for, -EAGAIN is the answer */ + if (req->flags & REQ_F_NOWAIT) + return false; + /* IOPOLL/SQPOLL issue differently, SQ_REWIND can't resume mid-batch */ + if (ctx->flags & (IORING_SETUP_IOPOLL | IORING_SETUP_SQPOLL | + IORING_SETUP_SQ_REWIND)) + return false; + /* pollable files keep the nonblocking issue + poll retry path */ + if (io_file_can_poll(req)) + return false; + if (!tctx->io_wq) + return false; + if (!thread_handoff_allowed(current)) + return false; + /* the SQ head is published while we may still be running */ + if (io_req_sqe_copy(req, IO_URING_F_INLINE)) + return false; + /* have a worker ready to take over */ + if (!io_wq_handoff_spare(tctx->io_wq, !io_req_unbound(req), false)) + return false; + /* would interrupt the issue right away, and can't be handled here */ + if (signal_pending(current)) + return false; + + ho->req =3D req; + io_handoff_block_signals(ho); + current->flags |=3D PF_IO_HANDOFF; + return true; +} + +/* Returns true if the identity got handed off during the issue */ +bool io_handoff_end(void) +{ + bool handed_off =3D current->flags & PF_IO_WORKER; + + current->flags &=3D ~PF_IO_HANDOFF; + /* pairs with io_wq_task_work_add(), notify for work queued meanwhile */ + smp_mb(); + if (task_work_pending(current)) + set_notify_signal(current); + + /* once handed off the tctx isn't ours anymore */ + if (likely(!handed_off)) + current->io_uring->handoff.req =3D NULL; + return handed_off; +} + +/* close our part of the batch and drop uring_lock for the promoted task */ +static void io_handoff_release_ring(struct io_ring_ctx *ctx, + struct io_handoff *ho) + __releases(&ctx->uring_lock) +{ + lockdep_assert_held(&ctx->uring_lock); + + ho->consumed +=3D io_submit_sqes_abandon(ctx); + mutex_unlock(&ctx->uring_lock); +} + +/* move the outstanding tctx task refs from @src to @dst, see io_put_task(= ) */ +static void io_handoff_task_refs(struct io_uring_task *tctx, + struct task_struct *src, + struct task_struct *dst) +{ + unsigned int nr; + + raw_spin_lock(&tctx->task_ref_lock); + nr =3D percpu_counter_sum(&tctx->inflight); + refcount_add(nr, &dst->usage); + WRITE_ONCE(tctx->task, dst); + raw_spin_unlock(&tctx->task_ref_lock); + + /* dropped by the promoted task once it's done taking over */ + tctx->handoff.src_refs =3D nr; +} + +/* move tctx task_work queued on @task along to the tctx's new task */ +void io_handoff_tw_moved(struct io_uring_task *tctx, struct task_struct *t= ask) +{ + struct task_struct *cur; + + while (task_work_cancel(task, &tctx->task_work)) { + cur =3D READ_ONCE(tctx->task); + if (WARN_ON_ONCE(task_work_add(cur, &tctx->task_work, TWA_SIGNAL))) + break; + if (READ_ONCE(tctx->task) =3D=3D cur) + break; + task =3D cur; + } +} + +/* Move the io_uring task state from @src to @dst */ +static void io_handoff_move_tctx(struct io_uring_task *tctx, + struct task_struct *src, + struct task_struct *dst) +{ + struct io_tctx_node *node; + + io_handoff_task_refs(tctx, src, dst); + + dst->io_uring =3D tctx; + src->io_uring =3D NULL; + dst->io_uring_restrict =3D src->io_uring_restrict; + src->io_uring_restrict =3D NULL; + + list_for_each_entry(node, &tctx->node_list, tctx_link) { + struct io_ring_ctx *ctx =3D node->ctx; + + node->task =3D dst; + if (READ_ONCE(ctx->submitter_task) =3D=3D src) { + get_task_struct(dst); + WRITE_ONCE(ctx->submitter_task, dst); + put_task_struct(src); + } + } + + io_handoff_tw_moved(tctx, src); +} + +/* + * Called from sched_submit_work() when a task blocks inside an inline iss= ue, + * hand our identity to an idle worker. Can't block, uring_lock is held. + */ +void io_uring_task_sleeping(struct task_struct *tsk) +{ + struct io_uring_task *tctx =3D tsk->io_uring; + struct io_handoff *ho =3D &tctx->handoff; + struct io_kiocb *req =3D ho->req; + struct io_ring_ctx *ctx =3D req->ctx; + struct task_struct *dst; + bool bound; + + WARN_ON_ONCE(tsk !=3D current); + + /* the issue path is touching state that needs the ring lock held */ + if (ctx->submit_lock_depth) + return; + if (!thread_handoff_prepare(tsk)) + return; + + /* don't let the woken worker preempt us before we've committed */ + preempt_disable(); + bound =3D !io_req_unbound(req); + dst =3D io_wq_handoff_claim(tctx->io_wq, bound, io_handoff_resume); + if (!dst) { + preempt_enable(); + return; + } + + /* committed, @req is ours as the worker from here on */ + ho->src =3D tsk; + ho->ctx =3D ctx; + ho->bound =3D bound; + thread_handoff_stats_take(&ho->stats); + + io_handoff_release_ring(ctx, ho); + io_handoff_move_tctx(tctx, tsk, dst); + io_wq_handoff_commit(dst); + + /* do what sched_submit_work() would have done for an io-wq worker */ + io_wq_worker_sleeping(tsk); + preempt_enable(); +} + +/* the issue of @req blocked and we're a worker now, finish it like io-wq = */ +int io_handoff_complete(struct io_kiocb *req, int ret) +{ + WARN_ON_ONCE(!io_wq_current_is_worker()); + + if (ret =3D=3D IOU_COMPLETE) { + req->io_task_work.func =3D io_req_task_complete; + io_req_task_work_add(req); + } else if (ret =3D=3D IOU_ISSUE_SKIP_COMPLETE) { + /* completes on its own */ + } else if ((ret =3D=3D -EAGAIN && !(req->flags & REQ_F_NOWAIT)) || + io_issue_wants_restart(ret)) { + /* wants a blocking retry, or got interrupted, io-wq does that */ + io_queue_iowq(req); + } else { + io_req_task_queue_fail(req, ret); + } + + return -EIOCBQUEUED; +} + +/* runs on the promoted task, finishes io_uring_enter() for the submitter = */ +static long io_handoff_resume(void) +{ + struct io_uring_task *tctx =3D current->io_uring; + struct io_handoff *ho =3D &tctx->handoff; + struct task_struct *src =3D ho->src; + struct io_ring_ctx *ctx =3D ho->ctx; + bool bound =3D ho->bound; + long ret; + + if (WARN_ON_ONCE(thread_handoff_finish(src, &ho->stats))) + force_sig(SIGKILL); + io_wq_handoff_finished(src); + put_task_struct_many(src, ho->src_refs); + ho->src_refs =3D 0; + ho->src =3D NULL; + ho->req =3D NULL; + ho->ctx =3D NULL; + + /* flush what the blocked batch left behind, then submit the rest */ + io_run_task_work(); + mutex_lock(&ctx->uring_lock); + io_submit_flush_completions(ctx); + if (ho->consumed < ho->to_submit) { + ret =3D io_submit_sqes(ctx, ho->to_submit - ho->consumed); + if (ret =3D=3D -EIOCBQUEUED) + return ret; + if (ret > 0) + ho->consumed +=3D ret; + } + /* the identity came with the mask the blocking issue ran under */ + io_handoff_submit_end(); + ret =3D ho->consumed; + if (ret !=3D ho->to_submit) { + mutex_unlock(&ctx->uring_lock); + } else { + ret =3D io_uring_enter_finish(ctx, ret, ho->min_complete, + ho->flags, ho->argp, ho->argsz); + } + if (!(ho->flags & IORING_ENTER_REGISTERED_RING)) + fput(ho->file); + + /* we took a worker, top the spare pool back up now the work is done */ + io_wq_handoff_spare(tctx->io_wq, bound, true); + + syscall_set_return_value(current, task_pt_regs(current), + ret < 0 ? ret : 0, ret); + return ret; +} + +/* run the worker loop after a demotion, returns once handed an identity */ +long io_uring_handoff_worker(void) +{ + io_wq_handoff_fn *fn; + long ret; + + do { + fn =3D io_wq_handoff_worker(); + ret =3D fn(); + } while (ret =3D=3D -EIOCBQUEUED); + + return ret; +} diff --git a/io_uring/handoff.h b/io_uring/handoff.h new file mode 100644 index 000000000000..833c6314d3b7 --- /dev/null +++ b/io_uring/handoff.h @@ -0,0 +1,102 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef IOU_HANDOFF_H +#define IOU_HANDOFF_H + +#include +#include "opdef.h" +#include "tw.h" + +/* a blocking issue got interrupted, retry on io-wq rather than restart */ +static inline bool io_issue_wants_restart(int ret) +{ + return ret =3D=3D -ERESTARTSYS || ret =3D=3D -ERESTARTNOINTR || + ret =3D=3D -ERESTARTNOHAND || ret =3D=3D -ERESTART_RESTARTBLOCK; +} + +#ifdef CONFIG_THREAD_HANDOFF +extern int sysctl_io_uring_handoff; + +bool __io_handoff_begin(struct io_kiocb *req); +void io_handoff_prime(struct io_uring_task *tctx, struct io_ring_ctx *ctx); +bool io_handoff_end(void); +void __io_handoff_restore_signals(struct io_handoff *ho); +int io_handoff_complete(struct io_kiocb *req, int ret); +long io_uring_handoff_worker(void); +void io_handoff_tw_moved(struct io_uring_task *tctx, struct task_struct *t= ask); + +/* + * Stash the io_uring_enter() arguments so a promoted task can resume it, = and + * run pending task_work so it doesn't interrupt a blocking issue later. + */ +static inline void io_handoff_enter(struct file *file, u32 to_submit, + u32 min_complete, u32 flags, + const void __user *argp, size_t argsz) +{ + struct io_handoff *ho =3D ¤t->io_uring->handoff; + + io_run_task_work(); + ho->file =3D file; + ho->to_submit =3D to_submit; + ho->consumed =3D 0; + ho->min_complete =3D min_complete; + ho->flags =3D flags; + ho->argp =3D argp; + ho->argsz =3D argsz; +} + +/* a submit call is done issuing, restore the signal mask if we changed it= */ +static inline void io_handoff_submit_end(void) +{ + struct io_handoff *ho =3D ¤t->io_uring->handoff; + + if (unlikely(ho->sigsaved)) + __io_handoff_restore_signals(ho); +} + +/* if true, @req gets a blocking inline issue. Pair with io_handoff_end() = */ +static inline bool io_handoff_begin(struct io_kiocb *req, + const struct io_issue_def *def, + unsigned int issue_flags) +{ + if (!(issue_flags & IO_URING_F_INLINE) || !def->blockable) + return false; + return __io_handoff_begin(req); +} +#else +static inline void io_handoff_enter(struct file *file, u32 to_submit, + u32 min_complete, u32 flags, + const void __user *argp, size_t argsz) +{ +} +static inline bool io_handoff_begin(struct io_kiocb *req, + const struct io_issue_def *def, + unsigned int issue_flags) +{ + return false; +} +static inline void io_handoff_submit_end(void) +{ +} +static inline bool io_handoff_end(void) +{ + return false; +} +static inline int io_handoff_complete(struct io_kiocb *req, int ret) +{ + return -EFAULT; +} +static inline long io_uring_handoff_worker(void) +{ + return -EFAULT; +} +static inline void io_handoff_tw_moved(struct io_uring_task *tctx, + struct task_struct *task) +{ +} +static inline void io_handoff_prime(struct io_uring_task *tctx, + struct io_ring_ctx *ctx) +{ +} +#endif + +#endif diff --git a/io_uring/io_uring.c b/io_uring/io_uring.c index 100ade1eee3e..289e9ddc8c24 100644 --- a/io_uring/io_uring.c +++ b/io_uring/io_uring.c @@ -98,6 +98,7 @@ #include "wait.h" #include "bpf_filter.h" #include "loop.h" +#include "handoff.h" =20 #define SQE_COMMON_FLAGS (IOSQE_FIXED_FILE | IOSQE_IO_LINK | \ IOSQE_IO_HARDLINK | IOSQE_ASYNC) @@ -119,7 +120,7 @@ /* requests with any of those set should undergo io_disarm_next() */ #define IO_DISARM_MASK (REQ_F_ARM_LTIMEOUT | REQ_F_LINK_TIMEOUT | REQ_F_FA= IL) =20 -static void io_queue_sqe(struct io_kiocb *req, unsigned int extra_flags); +static int io_queue_sqe(struct io_kiocb *req, unsigned int extra_flags); static void __io_req_caches_free(struct io_ring_ctx *ctx); =20 static __read_mostly DEFINE_STATIC_KEY_DEFERRED_FALSE(io_key_has_sqarray, = HZ); @@ -132,6 +133,17 @@ static int __read_mostly sysctl_io_uring_group =3D -1; =20 #ifdef CONFIG_SYSCTL static const struct ctl_table kernel_io_uring_disabled_table[] =3D { +#ifdef CONFIG_THREAD_HANDOFF + { + .procname =3D "io_uring_handoff", + .data =3D &sysctl_io_uring_handoff, + .maxlen =3D sizeof(sysctl_io_uring_handoff), + .mode =3D 0644, + .proc_handler =3D proc_dointvec_minmax, + .extra1 =3D SYSCTL_ZERO, + .extra2 =3D SYSCTL_ONE, + }, +#endif { .procname =3D "io_uring_disabled", .data =3D &sysctl_io_uring_disabled, @@ -384,9 +396,8 @@ static void io_prep_async_work(struct io_kiocb *req) should_hash =3D false; if (should_hash || (req->flags & REQ_F_IOPOLL)) io_wq_hash_work(&req->work, file_inode(req->file)); - } else if (!req->file || !S_ISBLK(file_inode(req->file)->i_mode)) { - if (def->unbound_nonreg_file) - atomic_or(IO_WQ_WORK_UNBOUND, &req->work.flags); + } else if (io_req_unbound(req)) { + atomic_or(IO_WQ_WORK_UNBOUND, &req->work.flags); } } =20 @@ -595,10 +606,16 @@ static inline void io_put_task(struct io_kiocb *req) if (likely(tctx->task =3D=3D current)) { tctx->cached_refs++; } else { + struct task_struct *task; + + /* ->task can change under us, see io_handoff_task_refs() */ + raw_spin_lock(&tctx->task_ref_lock); percpu_counter_sub(&tctx->inflight, 1); + task =3D tctx->task; + raw_spin_unlock(&tctx->task_ref_lock); if (unlikely(atomic_read(&tctx->in_cancel))) wake_up(&tctx->wait); - put_task_struct(tctx->task); + put_task_struct(task); } } =20 @@ -1401,12 +1418,16 @@ static inline int __io_issue_sqe(struct io_kiocb *r= eq, static int io_issue_sqe(struct io_kiocb *req, unsigned int issue_flags) { const struct io_issue_def *def =3D &io_issue_defs[req->opcode]; + bool handoff; int ret; =20 if (unlikely(!io_assign_file(req, def, issue_flags))) return -EBADF; =20 + handoff =3D io_handoff_begin(req, def, issue_flags); ret =3D __io_issue_sqe(req, issue_flags, def); + if (handoff && unlikely(io_handoff_end())) + return io_handoff_complete(req, ret); =20 if (ret =3D=3D IOU_COMPLETE) { if (issue_flags & IO_URING_F_COMPLETE_DEFER) @@ -1590,7 +1611,7 @@ struct file *io_file_get_normal(struct io_kiocb *req,= int fd) return file; } =20 -static int io_req_sqe_copy(struct io_kiocb *req, unsigned int issue_flags) +int io_req_sqe_copy(struct io_kiocb *req, unsigned int issue_flags) { const struct io_cold_def *def =3D &io_cold_defs[req->opcode]; =20 @@ -1630,7 +1651,8 @@ static void io_queue_async(struct io_kiocb *req, unsi= gned int issue_flags, int r } } =20 -static inline void io_queue_sqe(struct io_kiocb *req, unsigned int extra_f= lags) +/* returns -EIOCBQUEUED if the identity got handed off, we're a worker now= */ +static inline int io_queue_sqe(struct io_kiocb *req, unsigned int extra_fl= ags) __must_hold(&req->ctx->uring_lock) { unsigned int issue_flags =3D IO_URING_F_NONBLOCK | @@ -1638,6 +1660,8 @@ static inline void io_queue_sqe(struct io_kiocb *req,= unsigned int extra_flags) int ret; =20 ret =3D io_issue_sqe(req, issue_flags); + if (unlikely(ret =3D=3D -EIOCBQUEUED)) + return ret; =20 /* * We async punt it if the file wasn't marked NOWAIT, or if the file @@ -1645,6 +1669,7 @@ static inline void io_queue_sqe(struct io_kiocb *req,= unsigned int extra_flags) */ if (unlikely(ret)) io_queue_async(req, issue_flags, ret); + return 0; } =20 static void io_queue_sqe_fallback(struct io_kiocb *req) @@ -1922,8 +1947,7 @@ static inline int io_submit_sqe(struct io_ring_ctx *c= tx, struct io_kiocb *req, return 0; } =20 - io_queue_sqe(req, IO_URING_F_INLINE); - return 0; + return io_queue_sqe(req, IO_URING_F_INLINE); } =20 /* @@ -2033,6 +2057,25 @@ static int io_submit_sqes_end(struct io_ring_ctx *ct= x, unsigned int entries, return ret; } =20 +#ifdef CONFIG_THREAD_HANDOFF +/* + * The submitter blocked mid-batch, return the task refs for what it won't + * submit and publish the SQ head. Returns the number of entries consumed. + */ +unsigned int io_submit_sqes_abandon(struct io_ring_ctx *ctx) + __must_hold(&ctx->uring_lock) +{ + struct io_submit_state *state =3D &ctx->submit_state; + unsigned int consumed =3D ctx->cached_sq_head - state->sq_head; + + WARN_ON_ONCE(state->link.head); + current->io_uring->cached_refs +=3D state->submit_nr - consumed; + state->plug_started =3D false; + io_commit_sqring(ctx); + return consumed; +} +#endif + int io_submit_sqes(struct io_ring_ctx *ctx, unsigned int nr) __must_hold(&ctx->uring_lock) { @@ -2052,10 +2095,12 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigne= d int nr) left =3D entries; io_get_task_refs(left); io_submit_state_start(&ctx->submit_state, &plug, left); + ctx->submit_state.sq_head =3D ctx->cached_sq_head; =20 do { const struct io_uring_sqe *sqe; struct io_kiocb *req; + int ret; =20 if (unlikely(!io_alloc_req(ctx, &req))) break; @@ -2064,17 +2109,24 @@ int io_submit_sqes(struct io_ring_ctx *ctx, unsigne= d int nr) break; } =20 + ret =3D io_submit_sqe(ctx, req, sqe, &left); + /* handed off, the promoted task finishes the batch */ + if (unlikely(ret =3D=3D -EIOCBQUEUED)) { + if (current->plug =3D=3D &plug) + blk_finish_plug(&plug); + return ret; + } /* * Continue submitting even for sqe failure if the * ring was setup with IORING_SETUP_SUBMIT_ALL */ - if (unlikely(io_submit_sqe(ctx, req, sqe, &left)) && - !(ctx->flags & IORING_SETUP_SUBMIT_ALL)) { + if (unlikely(ret) && !(ctx->flags & IORING_SETUP_SUBMIT_ALL)) { left--; break; } } while (--left); =20 + io_handoff_submit_end(); return io_submit_sqes_end(ctx, entries, left); } =20 @@ -2656,7 +2708,7 @@ static int io_uring_getevents(struct io_ring_ctx *ctx= , int ret, } =20 /* Finish an io_uring_enter() call that submitted and holds the uring_lock= */ -static int io_uring_enter_finish(struct io_ring_ctx *ctx, int ret, u32 min= _complete, +int io_uring_enter_finish(struct io_ring_ctx *ctx, int ret, u32 min_comple= te, u32 flags, const void __user *argp, size_t argsz) { int ret2; @@ -2733,8 +2785,16 @@ SYSCALL_DEFINE6(io_uring_enter, unsigned int, fd, u3= 2, to_submit, if (unlikely(ret)) goto out; =20 + io_handoff_enter(file, to_submit, min_complete, flags, argp, + argsz); mutex_lock(&ctx->uring_lock); ret =3D io_submit_sqes(ctx, to_submit); + /* handed off, the promoted task finishes the syscall */ + if (unlikely(ret =3D=3D -EIOCBQUEUED)) { + /* uring_lock was dropped for the promoted task */ + __release(&ctx->uring_lock); + return io_uring_handoff_worker(); + } if (ret !=3D to_submit) { mutex_unlock(&ctx->uring_lock); goto out; diff --git a/io_uring/io_uring.h b/io_uring/io_uring.h index 870bb4dcc415..79db0a8b9cc8 100644 --- a/io_uring/io_uring.h +++ b/io_uring/io_uring.h @@ -211,6 +211,10 @@ void io_free_req(struct io_kiocb *req); void io_queue_next(struct io_kiocb *req); void io_task_refs_refill(struct io_uring_task *tctx); bool __io_alloc_req_refill(struct io_ring_ctx *ctx); +int io_req_sqe_copy(struct io_kiocb *req, unsigned int issue_flags); +unsigned int io_submit_sqes_abandon(struct io_ring_ctx *ctx); +int io_uring_enter_finish(struct io_ring_ctx *ctx, int ret, u32 min_comple= te, + u32 flags, const void __user *argp, size_t argsz); =20 void io_activate_pollwq(struct io_ring_ctx *ctx); void io_restriction_clone(struct io_restriction *dst, struct io_restrictio= n *src); @@ -389,13 +393,41 @@ static inline void io_put_file(struct io_kiocb *req) fput(req->file); } =20 +/* a handed off issue keeps its issue_flags but no longer holds uring_lock= */ +static inline bool io_issue_handed_off(unsigned int issue_flags) +{ + return !(issue_flags & IO_URING_F_UNLOCKED) && + (current->flags & (PF_IO_HANDOFF | PF_IO_WORKER)) =3D=3D + (PF_IO_HANDOFF | PF_IO_WORKER); +} + +/* which io-wq pool a request belongs in, bound unless a non-reg file op */ +static inline bool io_req_unbound(struct io_kiocb *req) +{ + if (!io_issue_defs[req->opcode].unbound_nonreg_file) + return false; + if (req->file) { + umode_t mode =3D file_inode(req->file)->i_mode; + + if (S_ISREG(mode) || S_ISBLK(mode)) + return false; + } + return true; +} + +static inline bool io_issue_needs_lock(unsigned int issue_flags) +{ + return (issue_flags & IO_URING_F_UNLOCKED) || + io_issue_handed_off(issue_flags); +} + static inline void io_ring_submit_unlock(struct io_ring_ctx *ctx, unsigned issue_flags) { lockdep_assert_held(&ctx->uring_lock); lockdep_assert(ctx->submit_lock_depth > 0); ctx->submit_lock_depth--; - if (unlikely(issue_flags & IO_URING_F_UNLOCKED)) + if (unlikely(io_issue_needs_lock(issue_flags))) mutex_unlock(&ctx->uring_lock); } =20 @@ -408,7 +440,7 @@ static inline void io_ring_submit_lock(struct io_ring_c= tx *ctx, * The only exception is when we've detached the request and issue it * from an async worker thread, grab the lock for that case. */ - if (unlikely(issue_flags & IO_URING_F_UNLOCKED)) + if (unlikely(io_issue_needs_lock(issue_flags))) mutex_lock(&ctx->uring_lock); lockdep_assert_held(&ctx->uring_lock); ctx->submit_lock_depth++; diff --git a/io_uring/msg_ring.c b/io_uring/msg_ring.c index 3067c9343991..04f2ffe4381d 100644 --- a/io_uring/msg_ring.c +++ b/io_uring/msg_ring.c @@ -245,7 +245,7 @@ static int io_msg_fd_remote(struct io_kiocb *req) struct task_struct *task =3D ctx->submitter_task; =20 init_task_work(&msg->tw, io_msg_tw_fd_complete); - if (task_work_add(task, &msg->tw, TWA_SIGNAL)) + if (io_wq_task_work_add(task, &msg->tw, TWA_SIGNAL)) return -EOWNERDEAD; =20 return IOU_ISSUE_SKIP_COMPLETE; diff --git a/io_uring/rw.c b/io_uring/rw.c index 95106dd1d7eb..23b078fe2991 100644 --- a/io_uring/rw.c +++ b/io_uring/rw.c @@ -134,7 +134,7 @@ static bool io_rw_recycle(struct io_kiocb *req, unsigne= d int issue_flags) { struct io_async_rw *rw =3D req->async_data; =20 - if (unlikely(issue_flags & IO_URING_F_UNLOCKED)) + if (unlikely(io_issue_needs_lock(issue_flags))) return false; =20 io_alloc_cache_vec_kasan(&rw->vec); diff --git a/io_uring/tctx.c b/io_uring/tctx.c index 737dfad4a976..f86aca489d15 100644 --- a/io_uring/tctx.c +++ b/io_uring/tctx.c @@ -10,6 +10,7 @@ #include =20 #include "io_uring.h" +#include "handoff.h" #include "tctx.h" #include "bpf_filter.h" =20 @@ -104,6 +105,7 @@ __cold struct io_uring_task *io_uring_alloc_task_contex= t(struct task_struct *tas } =20 tctx->task =3D task; + raw_spin_lock_init(&tctx->task_ref_lock); xa_init(&tctx->xa); INIT_LIST_HEAD(&tctx->node_list); init_waitqueue_head(&tctx->wait); @@ -175,8 +177,10 @@ int __io_uring_add_tctx_node(struct io_ring_ctx *ctx) * been marked for idle-exit when the task temporarily had no active * io_uring instances. */ - if (tctx->io_wq) + if (tctx->io_wq) { io_wq_set_exit_on_idle(tctx->io_wq, false); + io_handoff_prime(tctx, ctx); + } =20 if (new_tctx) current->io_uring =3D tctx; diff --git a/io_uring/tw.c b/io_uring/tw.c index f573bcc3af6a..f26faeb5d21c 100644 --- a/io_uring/tw.c +++ b/io_uring/tw.c @@ -9,6 +9,7 @@ #include =20 #include "io_uring.h" +#include "handoff.h" #include "tctx.h" #include "poll.h" #include "rw.h" @@ -130,6 +131,10 @@ void tctx_task_work(struct callback_head *cb) unsigned int count =3D 0; =20 tctx =3D container_of(cb, struct io_uring_task, task_work); + /* the tctx may have moved while queued, run it there if so */ + if (unlikely(READ_ONCE(tctx->task) !=3D current) && + !task_work_add(tctx->task, cb, TWA_SIGNAL)) + return; tctx_task_work_run(tctx, UINT_MAX, &count); } =20 @@ -209,6 +214,7 @@ void io_req_normal_work_add(struct io_kiocb *req) { struct io_uring_task *tctx =3D req->tctx; struct io_ring_ctx *ctx =3D req->ctx; + struct task_struct *task; =20 /* tw run already pending, nothing else to do */ if (!mpscq_push(&tctx->task_list, &req->io_task_work.node)) @@ -227,8 +233,14 @@ void io_req_normal_work_add(struct io_kiocb *req) return; } =20 - if (likely(!task_work_add(tctx->task, &tctx->task_work, ctx->notify_metho= d))) + task =3D READ_ONCE(tctx->task); + if (likely(!io_wq_task_work_add(task, &tctx->task_work, + ctx->notify_method))) { + /* the tctx moved while we were adding, move the work along */ + if (unlikely(READ_ONCE(tctx->task) !=3D task)) + io_handoff_tw_moved(tctx, task); return; + } =20 io_fallback_tw(tctx); } diff --git a/io_uring/uring_cmd.c b/io_uring/uring_cmd.c index 726a659f38c3..ba36a555db98 100644 --- a/io_uring/uring_cmd.c +++ b/io_uring/uring_cmd.c @@ -28,7 +28,7 @@ static void io_req_uring_cleanup(struct io_kiocb *req, un= signed int issue_flags) struct io_uring_cmd *ioucmd =3D io_kiocb_to_cmd(req, struct io_uring_cmd); struct io_async_cmd *ac =3D req->async_data; =20 - if (issue_flags & IO_URING_F_UNLOCKED) + if (io_issue_needs_lock(issue_flags)) return; =20 io_alloc_cache_vec_kasan(&ac->vec); @@ -172,7 +172,8 @@ void __io_uring_cmd_done(struct io_uring_cmd *ioucmd, s= 32 ret, u64 res2, if (req->flags & REQ_F_IOPOLL) { /* order with io_do_iopoll() checking ->iopoll_completed */ smp_store_release(&req->iopoll_completed, 1); - } else if (issue_flags & IO_URING_F_COMPLETE_DEFER) { + } else if ((issue_flags & IO_URING_F_COMPLETE_DEFER) && + !io_issue_handed_off(issue_flags)) { if (WARN_ON_ONCE(issue_flags & IO_URING_F_UNLOCKED)) return; io_req_complete_defer(req); --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oo1-f47.google.com (mail-oo1-f47.google.com [209.85.161.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1955148F02C for ; Fri, 11 Sep 2026 15:42:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.161.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141337; cv=none; b=CaB0oOByEvlWQCpEUOKSueG+NeiEiAwkDPaibLi2VIO7FaCgS7muNlG4js6rXrgrzx04HzXFLfePWKC+dzwT+aTMtI2ZYJrcC/LZPHXwx86KXdgXHw8SKzQ61fycG0jUd/Uq8er2cqjyRHfO3OCH4fJ6kDkhuiwLMD2fp2qKO9Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141337; c=relaxed/simple; bh=hynIoVRf6KTWjB7dioo8ii+57RiEIaRTzKGrkDPGesE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=fo0NEI+ddWggRwa4LBtH+PJclZ3ALPR9PH6/nC+VVbSRMJ3b17PXVXj5tm2lDmE4jiXm2H2yik3yRvRxYx79Q1XMVY106kmY2KwvEw1T2va3K4YAB0Q9+7lEcDp+7iBa8TQBxXnaXUKCr76cN5ZvOJYA2z6R3522yWpl2kQGuGo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=wvVVJEHF; arc=none smtp.client-ip=209.85.161.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="wvVVJEHF" Received: by mail-oo1-f47.google.com with SMTP id 006d021491bc7-6b1b1cb72d0so598195eaf.1 for ; Fri, 11 Sep 2026 08:42:13 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141332; x=1789746132; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=mIwUeFtb2rgpZypFvq5Uvr0k1u6Xe8moD+czEvfbQMY=; b=wvVVJEHFlWUMgEz5jWgI9Lb5EYa6DMnqHBa29HUbnu8XLbKgN3Oaagk2ENdS25sz0+ dOyxfXnVCKzr84kCWlzRqaKypxZMvjGznQSlVh5I+3QVUwEXlePfmGmZPl8dA0e+SaVU I9tZzxeuEDeu/fYudq8PLgu9utVr5EGl1myBm45yq6Fkw0oO9AAVStCDEsCa+9x0z5Sg tden1iTWW9XATM6dxID0LM0fhjkNtZplRAkg9RiWBQhRY9ZenWU52rDIHymtPaSQrj96 /aU8C73xrNKIkRbvxMLxtvGKn6fE/v4TWXSjrYjubNtqD6GYjc3RqanOj3hGsXfCU+HI KN9g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141332; x=1789746132; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=mIwUeFtb2rgpZypFvq5Uvr0k1u6Xe8moD+czEvfbQMY=; b=O9pfoD+sXOQFG5Wxc4VIBQrHuiJI28fNfPd3uZg/Ko2HBGiqvNs5DZFDJkCFFI1Ad+ gQjwuSui+7vvoAenHMUP0tRNJUM2Ipu3SqXlz2n8E/EfARvi7xe0P+j004RVO7RWxuKJ ry28KVIC8DgCtMtKSaMRnyyDMUzifO6jQDs4b4uz65jHJa3UmNNreaINt6H0H6oGpvzi /FCcK6FEHnbLQlB5vyJ+QvxMLP24N3TfTFm8CTarZn578UGNqFz8x1uHUzEEge5aKCoR yblCt6SSZYlRg5bW1pckB82P6bE5Js/ou6z/aprr7pxOZaQlOKvkZT2TUhkx1oD+N3W/ nSvQ== X-Forwarded-Encrypted: i=1; AKwUvBx9uCM6kvYl84k1V4RwyBEfdopOP/7y1B1HuGf0mC4EtwzZWPBBH4XmHmo4OrVA4OrLTOFcvpdqLV5JDMU=@vger.kernel.org X-Gm-Message-State: AFuF++ldCmsuWs9XO6EELnED5ANmeGoIY1IbWEOEhZ3hsW72nsycf8xc BMtY1j6+NwICZonXxjvDgAeo512GMB/p5H5hHY7oPfbDJZxLDiVaP8NLo7dhYZLoxj8= X-Gm-Gg: AYBFou2sBRyGhIgNgBUzsEg/WYlvdybwJ3pWBCSk1/Eo2Fx0aR1fd9QsEi2E9f7JJBZ iUCJ8h7uj96Lk5jad6aX4gz47LP65+tVb+/htC4zpP0vvnTq/xogIf58+8rkDcomTJhMwf/HIAp GulJgLNFIS9YoSKlwkoDzVNiybDrEU5elPGBrfjo4ZaMo6FrYe31yLoiNuWfD7sPGlSug0dnym6 FAc/1koveaBPTm25aOf2QQ6sjWWSdA70UcqoyrswyTuDeTJ+/MklthNMX9m6BmHXMItyACFeDO5 cZb9R04QqbvkhCq++wvw2kDu0Ha+QpM3XIVF3HTgG3gsdUAhsElF1bt5/Jpq0EZdlaWPSga6tMA RoV+MDQ61x8DFQqf/kuhRGccs4hSpwkBYb3Skr9SBunw/3YddH2Qf8V1uXItvbJmnQzDhjabiI0 wVj/M2iPidVH8eFkw5PEyc3TsPbiTeJGo1IKcAw3G5/tQkDqe856FPJAV4rLJWgtPbyUF7bmYVQ ZUKV6mj8EitDhTJOMEjHKrs+9gM X-Received: by 2002:a05:6820:55d4:10b0:6b1:3534:a5b8 with SMTP id 006d021491bc7-6c0ba03ff05mr2428688eaf.15.1789141331976; Fri, 11 Sep 2026 08:42:11 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.10 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:10 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 12/15] io_uring: defer the identity migration to the end of the submission Date: Fri, 11 Sep 2026 09:41:02 -0600 Message-ID: <20260911154148.644489-13-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" A handoff currently migrates the full identity (tid, signals, cgroup, sched attributes, register state) before the promoted worker resumes the SQ ring. With N blockable SQEs in one io_uring_enter(), that puts a complete migration between each of them, where the old behaviour was N cheap io-wq punts. None of that is needed to run kernel code on the submitter's behalf, only its creds and io_uring context are. Have the promoted task adopt the creds and continue the submission right away. If it blocks and hands off again, it just goes back to being a worker. Only the task that ends the submission migrates the identity, once, from the original submitter which is parked in io_wq_handoff_worker() until then. A task demoted while running the handoff function now also goes through io_wq_handoff_worker() rather than straight into the worker loop, so it doesn't exit or rename itself while its state is still being read. Signed-off-by: Jens Axboe --- include/linux/io_uring_types.h | 5 +-- include/linux/thread_handoff.h | 4 +++ io_uring/handoff.c | 64 +++++++++++++++++++++++----------- io_uring/handoff.h | 7 ++-- io_uring/io-wq.c | 33 +++++++++++------- io_uring/io-wq.h | 3 +- kernel/thread_handoff.c | 6 ++++ 7 files changed, 83 insertions(+), 39 deletions(-) diff --git a/include/linux/io_uring_types.h b/include/linux/io_uring_types.h index 81bc4810fcab..6c8fe7232aa2 100644 --- a/include/linux/io_uring_types.h +++ b/include/linux/io_uring_types.h @@ -151,9 +151,10 @@ struct io_handoff { /* the submitter's signal mask while blocking issues run without */ sigset_t sigmask; bool sigsaved; - /* the task the identity came from, and the task refs it held */ + /* identity source, and the task this hop took the worker from */ struct task_struct *src; - unsigned int src_refs; + struct task_struct *prev; + unsigned int prev_refs; struct thread_handoff_stats stats; /* io_uring_enter() arguments, to resume the syscall */ struct file *file; diff --git a/include/linux/thread_handoff.h b/include/linux/thread_handoff.h index e1c17b833e7c..203ab6ef73e0 100644 --- a/include/linux/thread_handoff.h +++ b/include/linux/thread_handoff.h @@ -29,6 +29,7 @@ bool thread_handoff_compatible(struct task_struct *src, struct task_struct *dst); bool thread_handoff_prepare(struct task_struct *tsk); void thread_handoff_stats_take(struct thread_handoff_stats *st); +void thread_handoff_adopt_creds(struct task_struct *src); int thread_handoff_finish(struct task_struct *src, struct thread_handoff_stats *st); =20 @@ -62,6 +63,9 @@ static inline bool thread_handoff_prepare(struct task_str= uct *tsk) static inline void thread_handoff_stats_take(struct thread_handoff_stats *= st) { } +static inline void thread_handoff_adopt_creds(struct task_struct *src) +{ +} static inline int thread_handoff_finish(struct task_struct *src, struct thread_handoff_stats *st) { diff --git a/io_uring/handoff.c b/io_uring/handoff.c index 25e9e06812a7..ddc3c4d6a4f3 100644 --- a/io_uring/handoff.c +++ b/io_uring/handoff.c @@ -89,7 +89,8 @@ bool __io_handoff_begin(struct io_kiocb *req) return false; if (!tctx->io_wq) return false; - if (!thread_handoff_allowed(current)) + /* an intermediate task's own user state doesn't matter, it stays */ + if (!tctx->handoff.src && !thread_handoff_allowed(current)) return false; /* the SQ head is published while we may still be running */ if (io_req_sqe_copy(req, IO_URING_F_INLINE)) @@ -98,12 +99,15 @@ bool __io_handoff_begin(struct io_kiocb *req) if (!io_wq_handoff_spare(tctx->io_wq, !io_req_unbound(req), false)) return false; /* would interrupt the issue right away, and can't be handled here */ - if (signal_pending(current)) + if (task_sigpending(current)) return false; =20 ho->req =3D req; io_handoff_block_signals(ho); current->flags |=3D PF_IO_HANDOFF; + /* already queued task_work gets picked up by io_handoff_end() too */ + if (test_thread_flag(TIF_NOTIFY_SIGNAL)) + clear_notify_signal(); return true; } =20 @@ -148,8 +152,8 @@ static void io_handoff_task_refs(struct io_uring_task *= tctx, WRITE_ONCE(tctx->task, dst); raw_spin_unlock(&tctx->task_ref_lock); =20 - /* dropped by the promoted task once it's done taking over */ - tctx->handoff.src_refs =3D nr; + /* dropped by the promoted task */ + tctx->handoff.prev_refs =3D nr; } =20 /* move tctx task_work queued on @task along to the tctx's new task */ @@ -205,6 +209,8 @@ void io_uring_task_sleeping(struct task_struct *tsk) struct io_handoff *ho =3D &tctx->handoff; struct io_kiocb *req =3D ho->req; struct io_ring_ctx *ctx =3D req->ctx; + /* the identity being handed around, ours unless we're intermediate */ + struct task_struct *src =3D ho->src ?: tsk; struct task_struct *dst; bool bound; =20 @@ -213,23 +219,26 @@ void io_uring_task_sleeping(struct task_struct *tsk) /* the issue path is touching state that needs the ring lock held */ if (ctx->submit_lock_depth) return; - if (!thread_handoff_prepare(tsk)) + if (src =3D=3D tsk && !thread_handoff_prepare(tsk)) return; =20 /* don't let the woken worker preempt us before we've committed */ preempt_disable(); bound =3D !io_req_unbound(req); - dst =3D io_wq_handoff_claim(tctx->io_wq, bound, io_handoff_resume); + dst =3D io_wq_handoff_claim(tctx->io_wq, bound, io_handoff_resume, src); if (!dst) { preempt_enable(); return; } =20 /* committed, @req is ours as the worker from here on */ - ho->src =3D tsk; + ho->src =3D src; + ho->prev =3D tsk; ho->ctx =3D ctx; ho->bound =3D bound; - thread_handoff_stats_take(&ho->stats); + /* our accounting follows the identity, an intermediate's doesn't */ + if (src =3D=3D tsk) + thread_handoff_stats_take(&ho->stats); =20 io_handoff_release_ring(ctx, ho); io_handoff_move_tctx(tctx, tsk, dst); @@ -261,24 +270,29 @@ int io_handoff_complete(struct io_kiocb *req, int ret) return -EIOCBQUEUED; } =20 -/* runs on the promoted task, finishes io_uring_enter() for the submitter = */ +/* + * Runs on the promoted task, finishes io_uring_enter() for the submitter.= Only + * takes its identity if it gets through the submission without handing of= f. + */ static long io_handoff_resume(void) { struct io_uring_task *tctx =3D current->io_uring; struct io_handoff *ho =3D &tctx->handoff; - struct task_struct *src =3D ho->src; + struct task_struct *src =3D ho->src, *prev =3D ho->prev; struct io_ring_ctx *ctx =3D ho->ctx; bool bound =3D ho->bound; long ret; =20 - if (WARN_ON_ONCE(thread_handoff_finish(src, &ho->stats))) - force_sig(SIGKILL); - io_wq_handoff_finished(src); - put_task_struct_many(src, ho->src_refs); - ho->src_refs =3D 0; - ho->src =3D NULL; + /* enough of the identity to issue requests on its behalf */ + thread_handoff_adopt_creds(src); + put_task_struct_many(prev, ho->prev_refs); + ho->prev_refs =3D 0; + ho->prev =3D NULL; ho->req =3D NULL; ho->ctx =3D NULL; + /* an intermediate task has nothing we still need, let it work */ + if (prev !=3D src) + io_wq_handoff_finished(prev); =20 /* flush what the blocked batch left behind, then submit the rest */ io_run_task_work(); @@ -291,12 +305,20 @@ static long io_handoff_resume(void) if (ret > 0) ho->consumed +=3D ret; } - /* the identity came with the mask the blocking issue ran under */ - io_handoff_submit_end(); + + mutex_unlock(&ctx->uring_lock); + + /* submission done, become the submitter and return to userspace */ + if (WARN_ON_ONCE(thread_handoff_finish(src, &ho->stats))) + force_sig(SIGKILL); + if (ho->sigsaved) + __io_handoff_restore_signals(ho); + io_wq_handoff_finished(src); + ho->src =3D NULL; + ret =3D ho->consumed; - if (ret !=3D ho->to_submit) { - mutex_unlock(&ctx->uring_lock); - } else { + if (ret =3D=3D ho->to_submit && (ho->flags & IORING_ENTER_GETEVENTS)) { + mutex_lock(&ctx->uring_lock); ret =3D io_uring_enter_finish(ctx, ret, ho->min_complete, ho->flags, ho->argp, ho->argsz); } diff --git a/io_uring/handoff.h b/io_uring/handoff.h index 833c6314d3b7..b8ded4916606 100644 --- a/io_uring/handoff.h +++ b/io_uring/handoff.h @@ -44,12 +44,15 @@ static inline void io_handoff_enter(struct file *file, = u32 to_submit, ho->argsz =3D argsz; } =20 -/* a submit call is done issuing, restore the signal mask if we changed it= */ +/* + * Done issuing, restore the signal mask if we changed it. Not with a hand= off + * in flight, io_handoff_resume() does that once it has the identity. + */ static inline void io_handoff_submit_end(void) { struct io_handoff *ho =3D ¤t->io_uring->handoff; =20 - if (unlikely(ho->sigsaved)) + if (unlikely(ho->sigsaved) && !ho->src) __io_handoff_restore_signals(ho); } =20 diff --git a/io_uring/io-wq.c b/io_uring/io-wq.c index 3d4eb4992d5b..d29e5a80eddd 100644 --- a/io_uring/io-wq.c +++ b/io_uring/io-wq.c @@ -829,22 +829,22 @@ static int io_wq_worker(void *data) * Only returns if we got handed an identity. -EIOCBQUEUED means we got * demoted again while running it, back to the worker loop. */ + fn =3D io_wq_worker_run(worker); for (;;) { - long ret; + long ret =3D fn(); =20 - fn =3D io_wq_worker_run(worker); - ret =3D fn(); /* what we return is what userspace gets on some archs */ if (ret !=3D -EIOCBQUEUED) return ret; - worker =3D current->worker_private; + fn =3D io_wq_handoff_worker(); } } =20 /* find and claim an idle sleeping worker, see io_wq_worker_idle_done() */ static struct io_worker *io_wq_acct_handoff_claim(struct io_wq *wq, struct io_wq_acct *acct, - io_wq_handoff_fn *fn) + io_wq_handoff_fn *fn, + struct task_struct *src) { struct io_worker *worker, *found =3D NULL; struct hlist_nulls_node *n; @@ -856,7 +856,7 @@ static struct io_worker *io_wq_acct_handoff_claim(struc= t io_wq *wq, /* only claimable inside the idle sleep of the worker loop */ if (!test_bit(IO_WORKER_F_IDLE_SLEEP, &worker->flags)) continue; - if (!thread_handoff_compatible(current, worker->task)) + if (!thread_handoff_compatible(src, worker->task)) continue; clear_bit(IO_WORKER_F_FREE, &worker->flags); hlist_nulls_del_init_rcu(&worker->nulls_node); @@ -897,15 +897,17 @@ int io_wq_task_work_add(struct task_struct *task, str= uct callback_head *cb, return 0; } =20 -/* claim an idle worker to hand our identity to, pairs with _commit() */ +/* claim an idle worker to hand @src's identity to, pairs with _commit() */ struct task_struct *io_wq_handoff_claim(struct io_wq *wq, bool bound, - io_wq_handoff_fn *fn) + io_wq_handoff_fn *fn, + struct task_struct *src) { struct io_worker *worker; =20 - worker =3D io_wq_acct_handoff_claim(wq, io_get_acct(wq, bound), fn); + worker =3D io_wq_acct_handoff_claim(wq, io_get_acct(wq, bound), fn, src); if (!worker) - worker =3D io_wq_acct_handoff_claim(wq, io_get_acct(wq, !bound), fn); + worker =3D io_wq_acct_handoff_claim(wq, io_get_acct(wq, !bound), + fn, src); if (worker) return worker->task; return NULL; @@ -958,9 +960,14 @@ io_wq_handoff_fn *io_wq_handoff_worker(void) =20 WARN_ON_ONCE(!io_wq_current_is_worker()); =20 - /* the promoted task reads our state until it's done migrating it */ - wait_var_event(&worker->handoff, - atomic_read_acquire(&worker->handoff) =3D=3D IO_WORKER_HANDOFF_FINISHED); + /* + * Wait until nobody needs our state anymore, which may be a while if + * an identity is still parked on us. Hence TASK_IDLE. + */ + ___wait_var_event(&worker->handoff, + atomic_read_acquire(&worker->handoff) =3D=3D + IO_WORKER_HANDOFF_FINISHED, + TASK_IDLE, 0, 0, schedule()); atomic_set(&worker->handoff, IO_WORKER_HANDOFF_NONE); =20 snprintf(buf, sizeof(buf), "iou-wrk-%d", worker->wq->task->pid); diff --git a/io_uring/io-wq.h b/io_uring/io-wq.h index 98357b665e54..df451838828d 100644 --- a/io_uring/io-wq.h +++ b/io_uring/io-wq.h @@ -52,7 +52,8 @@ typedef long (io_wq_handoff_fn)(void); =20 /* claim an idle worker, it runs @fn instead of the worker loop when woken= */ struct task_struct *io_wq_handoff_claim(struct io_wq *wq, bool bound, - io_wq_handoff_fn *fn); + io_wq_handoff_fn *fn, + struct task_struct *src); =20 void io_wq_handoff_commit(struct task_struct *dst); io_wq_handoff_fn *io_wq_handoff_worker(void); diff --git a/kernel/thread_handoff.c b/kernel/thread_handoff.c index 1901eb85bae8..822fdd9a0e7f 100644 --- a/kernel/thread_handoff.c +++ b/kernel/thread_handoff.c @@ -393,6 +393,12 @@ static void thread_handoff_creds(struct task_struct *d= st, put_cred_many(old, 2); } =20 +/* the part of thread_handoff_finish() needed to run kernel code for @src = */ +void thread_handoff_adopt_creds(struct task_struct *src) +{ + thread_handoff_creds(current, src); +} + /* the user requested affinity follows, the effective mask derives from it= */ static void thread_handoff_affinity(struct task_struct *dst, struct task_struct *src) --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oa2-f12.google.com (mail-oa2-f12.google.com [74.125.231.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 179B9496D43 for ; Fri, 11 Sep 2026 15:42:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.76 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141338; cv=none; b=TTJimV6rtnrqi5g0z82xw4MMqYz8sQHzuDWeRmC0DAZWXa73t1nBVpSHAn3WrerD/SZpHeQZdFkvpmG88WqBGk2CPMmJ4xMnsRDWnjKWAkyDJsR1nxMgg4Gn68l2jH5DTL27nA6f+EezcYQIJnVf6/u9uNzs6B9Ut8OZCpbHIDo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141338; c=relaxed/simple; bh=dtEkrDn9TuULG8Rmh0vGe3duGIN/MxnGeLhOxWrexcY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uHiZKElP4APMLk3Xy+6bbz9M6P2K0cd1tXQgQElYkbp2wEIKCjGwEoOkFTTBaDlP4djW44S/2d3/iNXOXGk+duO0sPJR8XG6hLF/5MONFllzEfyVY+WKhCresVZHyYwKf3/Pums9/BRBydUP2lQrWqNCSbST0uzI0rvdeA6gD+w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=NQ/tanlM; arc=none smtp.client-ip=74.125.231.76 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="NQ/tanlM" Received: by mail-oa2-f12.google.com with SMTP id 586e51a60fabf-470417e5b52so805285fac.3 for ; Fri, 11 Sep 2026 08:42:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141333; x=1789746133; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=TNGI4wdJO9MfPQNZHymsObp/iaVWwbU1kTAYKqVGN8A=; b=NQ/tanlMP1DsJz9UozBHS0Jq2qRVnowwWpEXTwqyeBINoXGma4LGYFKL7wJwj0485t AKalzR8B2fquE4dyx2CxKYc+wEWr17kLPxvwyiik+JHKMxAdfJkv0TD7q1c1ptQGVIsd qG+l2SL7iFpO6nWHLldEc3KgLWkOhqNXh17r4+wfCKDWJZimUJSXLOCbj0la4gZj0hT5 qZZ9Osbe5twOl/rRHZ9WFc1ycCBLKEaoUoxERVfdues+oXY18UzUPhQEvufC8qMU004I NsWz8Zs9qdQu8qRHucULs37AsMHSC9UrhKhZkmobbiOgxozUExknZYw9yU9ipnaFNyr/ vdfA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141333; x=1789746133; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=TNGI4wdJO9MfPQNZHymsObp/iaVWwbU1kTAYKqVGN8A=; b=YkoSapx4D/ky8K5nFuRRhsCc/W3ivFFa/jTdNRvL9860iJM9S0O/c/gqNR6q4mGC+V 4h/YNSI5s6AsKXhR9Z3ZDGbcZ7oYSuW48GnKWUgO70CDQK5xhMt8tcvADQc82/d15Puq evMCa65zRA40GQKkTO01fklm0iAZdxaypE7PRuYeu1N7x4VpPliWeHfZ0R29uBzu55rN J6SLMY/NQJ4l8hFePKaQGznANKyrAGarl/xI3gnTp39sFVa58ZzmmAq2oCss3iDcQp58 ZXZn4ulpno5GHHrYgrbYb7JI7q7eRNnONo8YVmePtoEaNeF7vQppFsWqxWEn2pndTq+u 8KEA== X-Forwarded-Encrypted: i=1; AKwUvBwvY9uOoOfOPTuycsXLOYql1RmLnNNval5U7sXExUO1XT+rGh2qf7YEfHsRbfS5PIUc7DUDcR8RaEfuDF8=@vger.kernel.org X-Gm-Message-State: AFuF++kU2l3h/tTKzqJDWg65K32OO7Yt9SUJ9lbEVzu7LFNRJolW3EnI P061+p8JNI167TkpKfGFegtm7fqdx0Po4sBs2zzVRnO+Xa3IqgQAjyOi2aGwPCgfXU4= X-Gm-Gg: AYBFou2onGjvGQLnZIkfGaLytovKyH/QjfR5f2tIpPaarGqYjxuw10LMnj2x6Hd84fT HZYJRVckaFD9BsRg/uQUd9UlI85R9VgrY/O1VfafzP3oyovgRtl7FaTX5LX7Pb5A5zAdDabZ6pI DtW7sv+hp7sb35qo85w4hNRREuo/zVkHskuMT3n0yF/L/UcA0eMJfuUqcnAhGOrhvIVbQKH7ehm apusHZLPXxy1SGZA80iznWirZeO5lz/8u5xfmeDvbEBXkeTRWLJSWsP0UCNkBv7wmb1d+yrEV/t aO/i8v6+Jul0iT9nr/klqoNDr+Bz+qzooI67XLD3BWz7mAIavqyaFInCGsU23onuW0SimagXXM6 pRx75pmbsdczIaswGnb1QsSUM99J/WzGzRry0YSoKmJK3Zd9GJtWIWV68sjjYdseScFtIgmy1JW XxcPAWqpFh7N62WGrJEHMBao6eivl7Q6DdFZzAloiY2s+vOTO25scwy2ETfnTVFm7to/sLHZaeb bRLP9cudwJQ1/CmlTuZ2GuCe6Bfg+GVSOsiXuMuoSA= X-Received: by 2002:a05:6820:a283:20b0:6ac:8e23:3078 with SMTP id 006d021491bc7-6c0b9a5f10cmr2546524eaf.5.1789141333306; Fri, 11 Sep 2026 08:42:13 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.12 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:12 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 13/15] io_uring: issue blockable requests inline in blocking mode Date: Fri, 11 Sep 2026 09:41:03 -0600 Message-ID: <20260911154148.644489-14-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" With a handoff available, there's no point in issuing a blockable opcode nonblocking first. Clear IO_URING_F_NONBLOCK if io_handoff_begin() succeeds, and issue REQ_F_FORCE_ASYNC requests inline rather than punting them to io-wq upfront. Requests with a working nonblocking issue path, like reads and writes on FMODE_NOWAIT files, behave as before. The handoff is for requests that otherwise would have required an io-wq punt upfront, most of which never block. SQEs marked IOSQE_ASYNC keep their explicit io-wq offload. Signed-off-by: Jens Axboe --- include/linux/io_uring_types.h | 6 +++ io_uring/handoff.c | 75 ++++++++++++++++++++++------------ io_uring/handoff.h | 12 +++--- io_uring/io_uring.c | 48 +++++++++++++++++++--- io_uring/io_uring.h | 7 ++++ io_uring/splice.c | 6 +++ 6 files changed, 116 insertions(+), 38 deletions(-) diff --git a/include/linux/io_uring_types.h b/include/linux/io_uring_types.h index 6c8fe7232aa2..37c56ad37e05 100644 --- a/include/linux/io_uring_types.h +++ b/include/linux/io_uring_types.h @@ -651,6 +651,8 @@ enum { REQ_F_IMPORT_BUFFER_BIT, REQ_F_SQE_COPIED_BIT, REQ_F_IOPOLL_BIT, + REQ_F_ASYNC_USER_BIT, + REQ_F_HANDOFF_BIT, =20 /* not a real bit, just to check we're not overflowing the space */ __REQ_F_LAST_BIT, @@ -746,6 +748,10 @@ enum { REQ_F_SQE_COPIED =3D IO_REQ_FLAG(REQ_F_SQE_COPIED_BIT), /* request must be iopolled to completion (set in ->issue()) */ REQ_F_IOPOLL =3D IO_REQ_FLAG(REQ_F_IOPOLL_BIT), + /* IOSQE_ASYNC was set on the SQE, not just by prep */ + REQ_F_ASYNC_USER =3D IO_REQ_FLAG(REQ_F_ASYNC_USER_BIT), + /* vetted at submit for an inline blocking issue with a handoff */ + REQ_F_HANDOFF =3D IO_REQ_FLAG(REQ_F_HANDOFF_BIT), }; =20 struct io_tw_req { diff --git a/io_uring/handoff.c b/io_uring/handoff.c index ddc3c4d6a4f3..9c9bb7ba99f0 100644 --- a/io_uring/handoff.c +++ b/io_uring/handoff.c @@ -31,12 +31,58 @@ int sysctl_io_uring_handoff __read_mostly =3D 1; =20 static long io_handoff_resume(void); =20 +/* + * Can @req be issued inline in blocking mode with a handoff ready. Everyt= hing + * but the spare worker check is static, REQ_F_HANDOFF caches that part. + */ +bool io_handoff_possible(struct io_kiocb *req) +{ + const struct io_issue_def *def =3D &io_issue_defs[req->opcode]; + struct io_ring_ctx *ctx =3D req->ctx; + struct io_uring_task *tctx =3D current->io_uring; + + if (!sysctl_io_uring_handoff) + return false; + if (req->flags & REQ_F_HANDOFF) + goto check_spare; + if (!def->blockable) + return false; + /* nonblocking semantics were asked for, -EAGAIN is the answer */ + if (req->flags & REQ_F_NOWAIT) + return false; + /* IOPOLL/SQPOLL issue differently, SQ_REWIND can't resume mid-batch */ + if (ctx->flags & (IORING_SETUP_IOPOLL | IORING_SETUP_SQPOLL | + IORING_SETUP_SQ_REWIND)) + return false; + /* pollable files keep the nonblocking issue + poll retry path */ + if (io_file_can_poll(req)) + return false; + /* FMODE_NOWAIT files have a working nonblocking path, keep using it */ + if ((def->pollin || def->pollout) && req->file && + (req->file->f_mode & FMODE_NOWAIT)) + return false; + if (!tctx->io_wq) + return false; + /* an intermediate task's own user state doesn't matter, it stays */ + if (!tctx->handoff.src && !thread_handoff_allowed(current)) + return false; + /* the SQ head is published while we may still be running */ + if (io_req_sqe_copy(req, IO_URING_F_INLINE)) + return false; + req->flags |=3D REQ_F_HANDOFF; +check_spare: + /* have a worker ready to take over */ + if (!io_wq_handoff_spare(tctx->io_wq, !io_req_unbound(req), false)) + return false; + return true; +} + /* fork a spare worker upfront, so the first blockable issue has a target = */ void io_handoff_prime(struct io_uring_task *tctx, struct io_ring_ctx *ctx) { if (!sysctl_io_uring_handoff || !tctx->io_wq) return; - /* handoffs are never done for these, see __io_handoff_begin() */ + /* handoffs are never done for these, see io_handoff_possible() */ if (ctx->flags & (IORING_SETUP_IOPOLL | IORING_SETUP_SQPOLL | IORING_SETUP_SQ_REWIND)) return; @@ -71,32 +117,9 @@ void __io_handoff_restore_signals(struct io_handoff *ho) */ bool __io_handoff_begin(struct io_kiocb *req) { - struct io_ring_ctx *ctx =3D req->ctx; - struct io_uring_task *tctx =3D current->io_uring; - struct io_handoff *ho =3D &tctx->handoff; + struct io_handoff *ho =3D ¤t->io_uring->handoff; =20 - if (!sysctl_io_uring_handoff) - return false; - /* nonblocking semantics were asked for, -EAGAIN is the answer */ - if (req->flags & REQ_F_NOWAIT) - return false; - /* IOPOLL/SQPOLL issue differently, SQ_REWIND can't resume mid-batch */ - if (ctx->flags & (IORING_SETUP_IOPOLL | IORING_SETUP_SQPOLL | - IORING_SETUP_SQ_REWIND)) - return false; - /* pollable files keep the nonblocking issue + poll retry path */ - if (io_file_can_poll(req)) - return false; - if (!tctx->io_wq) - return false; - /* an intermediate task's own user state doesn't matter, it stays */ - if (!tctx->handoff.src && !thread_handoff_allowed(current)) - return false; - /* the SQ head is published while we may still be running */ - if (io_req_sqe_copy(req, IO_URING_F_INLINE)) - return false; - /* have a worker ready to take over */ - if (!io_wq_handoff_spare(tctx->io_wq, !io_req_unbound(req), false)) + if (!io_handoff_possible(req)) return false; /* would interrupt the issue right away, and can't be handled here */ if (task_sigpending(current)) diff --git a/io_uring/handoff.h b/io_uring/handoff.h index b8ded4916606..f315f2ae8d80 100644 --- a/io_uring/handoff.h +++ b/io_uring/handoff.h @@ -6,16 +6,10 @@ #include "opdef.h" #include "tw.h" =20 -/* a blocking issue got interrupted, retry on io-wq rather than restart */ -static inline bool io_issue_wants_restart(int ret) -{ - return ret =3D=3D -ERESTARTSYS || ret =3D=3D -ERESTARTNOINTR || - ret =3D=3D -ERESTARTNOHAND || ret =3D=3D -ERESTART_RESTARTBLOCK; -} - #ifdef CONFIG_THREAD_HANDOFF extern int sysctl_io_uring_handoff; =20 +bool io_handoff_possible(struct io_kiocb *req); bool __io_handoff_begin(struct io_kiocb *req); void io_handoff_prime(struct io_uring_task *tctx, struct io_ring_ctx *ctx); bool io_handoff_end(void); @@ -77,6 +71,10 @@ static inline bool io_handoff_begin(struct io_kiocb *req, { return false; } +static inline bool io_handoff_possible(struct io_kiocb *req) +{ + return false; +} static inline void io_handoff_submit_end(void) { } diff --git a/io_uring/io_uring.c b/io_uring/io_uring.c index 289e9ddc8c24..7c2aa0cfacc8 100644 --- a/io_uring/io_uring.c +++ b/io_uring/io_uring.c @@ -1424,10 +1424,25 @@ static int io_issue_sqe(struct io_kiocb *req, unsig= ned int issue_flags) if (unlikely(!io_assign_file(req, def, issue_flags))) return -EBADF; =20 + /* + * No point in a nonblocking attempt with a handoff armed. A force-async + * request can't do nonblocking at all, punt if no handoff is possible. + */ handoff =3D io_handoff_begin(req, def, issue_flags); + if (handoff) { + issue_flags &=3D ~IO_URING_F_NONBLOCK; + } else if ((issue_flags & IO_URING_F_INLINE) && + (req->flags & REQ_F_FORCE_ASYNC)) { + return -EAGAIN; + } ret =3D __io_issue_sqe(req, issue_flags, def); - if (handoff && unlikely(io_handoff_end())) - return io_handoff_complete(req, ret); + if (handoff) { + if (unlikely(io_handoff_end())) + return io_handoff_complete(req, ret); + /* interrupted regardless (fatal signal, stop), io-wq retries */ + if (unlikely(io_issue_wants_restart(ret))) + return -EAGAIN; + } =20 if (ret =3D=3D IOU_COMPLETE) { if (issue_flags & IO_URING_F_COMPLETE_DEFER) @@ -1756,6 +1771,8 @@ static int io_init_req(struct io_ring_ctx *ctx, struc= t io_kiocb *req, /* same numerical values with corresponding REQ_F_*, safe to copy */ sqe_flags =3D READ_ONCE(sqe->flags); req->flags =3D (__force io_req_flags_t) sqe_flags; + if (sqe_flags & IOSQE_ASYNC) + req->flags |=3D REQ_F_ASYNC_USER; req->cqe.user_data =3D READ_ONCE(sqe->user_data); req->file =3D NULL; req->tctx =3D current->io_uring; @@ -1895,6 +1912,25 @@ static __cold int io_submit_fail_init(const struct i= o_uring_sqe *sqe, return 0; } =20 +/* a blockable force-async request issued inline beats an io-wq punt */ +static bool io_req_force_async(struct io_kiocb *req) +{ + if (req->flags & REQ_F_FAIL) + return true; + if (!(req->flags & REQ_F_FORCE_ASYNC)) + return false; + /* userspace asked for it, keep the explicit offload */ + if (req->flags & REQ_F_ASYNC_USER) + return true; + if (req->ctx->int_flags & IO_RING_F_DRAIN_ACTIVE) + return true; + /* the file decides on pollability, resolve it now if fixed */ + if (!io_assign_file(req, &io_issue_defs[req->opcode], + IO_URING_F_INLINE)) + return true; + return !io_handoff_possible(req); +} + static inline int io_submit_sqe(struct io_ring_ctx *ctx, struct io_kiocb *= req, const struct io_uring_sqe *sqe, unsigned int *left) __must_hold(&ctx->uring_lock) @@ -1932,7 +1968,7 @@ static inline int io_submit_sqe(struct io_ring_ctx *c= tx, struct io_kiocb *req, /* last request of the link, flush it */ req =3D link->head; link->head =3D NULL; - if (req->flags & (REQ_F_FORCE_ASYNC | REQ_F_FAIL)) + if (io_req_force_async(req)) goto fallback; =20 } else if (unlikely(req->flags & (IO_REQ_LINK_FLAGS | @@ -1940,11 +1976,13 @@ static inline int io_submit_sqe(struct io_ring_ctx = *ctx, struct io_kiocb *req, if (req->flags & IO_REQ_LINK_FLAGS) { link->head =3D req; link->last =3D req; - } else { + return 0; + } + if (io_req_force_async(req)) { fallback: io_queue_sqe_fallback(req); + return 0; } - return 0; } =20 return io_queue_sqe(req, IO_URING_F_INLINE); diff --git a/io_uring/io_uring.h b/io_uring/io_uring.h index 79db0a8b9cc8..1f536617b584 100644 --- a/io_uring/io_uring.h +++ b/io_uring/io_uring.h @@ -421,6 +421,13 @@ static inline bool io_issue_needs_lock(unsigned int is= sue_flags) io_issue_handed_off(issue_flags); } =20 +/* a blocking issue got interrupted, retry on io-wq rather than restart */ +static inline bool io_issue_wants_restart(int ret) +{ + return ret =3D=3D -ERESTARTSYS || ret =3D=3D -ERESTARTNOINTR || + ret =3D=3D -ERESTARTNOHAND || ret =3D=3D -ERESTART_RESTARTBLOCK; +} + static inline void io_ring_submit_unlock(struct io_ring_ctx *ctx, unsigned issue_flags) { diff --git a/io_uring/splice.c b/io_uring/splice.c index e81ebbb91925..7d464deb0972 100644 --- a/io_uring/splice.c +++ b/io_uring/splice.c @@ -100,6 +100,9 @@ int io_tee(struct io_kiocb *req, unsigned int issue_fla= gs) =20 if (!(sp->flags & SPLICE_F_FD_IN_FIXED)) fput(in); + /* interrupted before making progress, have the core retry it */ + if (io_issue_wants_restart(ret)) + return -EAGAIN; done: if (ret !=3D sp->len) req_set_fail(req); @@ -141,6 +144,9 @@ int io_splice(struct io_kiocb *req, unsigned int issue_= flags) =20 if (!(sp->flags & SPLICE_F_FD_IN_FIXED)) fput(in); + /* interrupted before making progress, have the core retry it */ + if (io_issue_wants_restart(ret)) + return -EAGAIN; done: if (ret !=3D sp->len) req_set_fail(req); --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oa2-f12.google.com (mail-oa2-f12.google.com [74.125.231.76]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5C71D496D53 for ; Fri, 11 Sep 2026 15:42:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.76 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141339; cv=none; b=PUhsYvzL+XXbvcccuw/NzwWNU1bybsUWZdW+Mi61h+iYz/l8iIG2hrDw5eBB+8/Z20mexySuRdCJf5t/EUMGJnq8mmWiMQuEaydnJUyEXLHfMQ2aC6PeDvlkgyazs6qW7NsfU1NrEsLqL/2bbUT5bX+I/dAzIaYERA+jmrWd9ak= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141339; c=relaxed/simple; bh=0IcDUuFgsZRW2mHnNDY7QMSvPwoA7pARp/KcOd2nG6Y=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eiiNVtIhoa1PNHDzgUVI9Xhk9P2Edul+bNtvkgKtcsL+0oyf5BXuUCtQS3UFXTF4ZRLb+8rjaIgYXkKX+oRR6JajKsG21svtKkX2XMMeY+qAvTojwSI3LParq6PCRAqzBA0syR7SnYigbs+UlpqJi4MXAZduB5G/HeOXr+m9UwA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=KH7mZYIW; arc=none smtp.client-ip=74.125.231.76 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="KH7mZYIW" Received: by mail-oa2-f12.google.com with SMTP id 586e51a60fabf-47b5043f191so618418fac.3 for ; Fri, 11 Sep 2026 08:42:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141334; x=1789746134; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=a9YT6ic9x6FyZ148rG7PDLceVb4/eE9JqVTzoSzdUWU=; b=KH7mZYIW+Nd8JN84rRI5On75xjBWNm/NkJDjBUpY67qDqAxL1AxQWDYw2aBMyUPXVg +gMRXEGzQ6yuriNPMSfi8ZYvlQ7oRKDSbAngrjaOIWxKgQJyExCX5EBAae/lc4TD+QpT 4xofkkSi6Wpy/k4pzOkyTjYO295G0BfsDopa9b9eYJwxCPn16+K+XRltCTYBgrc8cIbb /wN7hbGWks1yzQmTNtKIxI2Je++T2qdsg4p4feo7Y6zpia4M15j4qbaUnJ0Z/MTGczmd itHISjpA7153hN4jlwqgMQPVDD3MmxfBjIPXEzRWudIPsXAv471tnKRjp5GNTnHu47FX S79A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141334; x=1789746134; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=a9YT6ic9x6FyZ148rG7PDLceVb4/eE9JqVTzoSzdUWU=; b=O/BkGmmFwlnFpzYgD38BYNUJ0j7jygrEgSiG7t0z72mnvA5ZK6PpnEqfjbfpYjl3m8 JJx8/D1TEE83M+VDDBDVZs2IvsQgyNkC1UOxKzKAyjBTccy2KPoxEmEhdWkACJPQSkXW Nz9B01N3OEZyyHIat8XU4dx5brnq1dYYcjoMo4rUHu0aNBZDZUY4hjU9o63tOnkOeSV7 2GMWLjSHWwzKAWICjJwc3PfMssyu7IM6U/VFqmVx3a/LGFaeWAYh8h0tvIxWkAy4TjOk CXmtdiiiPZMlFJ0sTZpX31TmA9BBj07umWdpfK0A8OUMvuvDMOlphrHiiVIb2uwha6WT Fd0Q== X-Forwarded-Encrypted: i=1; AKwUvBw3VdTT9PFEhLfUWRPA/5Z6B60CLOPPQs0KPWTR/bcUWX9XzI5OvqSjTz+EeSZ3McUocTQ/sS3cDr/fj50=@vger.kernel.org X-Gm-Message-State: AFuF++nMQZro3uybm7t3Mp7Sys4fLfjzfkg+lJGJS1a/WY/f0oEKPJ4z 2uEmIouOD5GpIwnA1IDBFL7V68ZHk/1gQiuVJZ0T6vvpFk8k9X6w/jBPzEWxR71Y4A4= X-Gm-Gg: AYBFou1YGmBJAwaHh4XfGeXMF0axOxFLmfATBPKtRA7AlovaqkRochW8j6q0NUHQ8Vf BS6aWjB7vlLOcBNATefp8VSTXUP3FPHxdGm+6EUp82AHJmMwa6Ulz32WkoS8xMr5fqkmFw1+cb7 t7H/e1kMVXoZmn4f1vIQfsGwrgDgqf1nIIIG1Zav4qmGEVhVuBuEcysE5wRsrjLTDNEIi3b+sO5 wqpnqglNvbmFCZamrps7jHRSv24dzPrUqq7lDDVwUpwiD+njRY79vMGAbrJCV+KNXbk5/qgehr0 lekPCnij0shzw/M3ifnJ6SXuWw0w1O8e6AwQo5JktQJbfvHO68eWgxr/DZZkHgmrvi+v0Meamnv E/33K7NIctvVMpbsrbxB+/jV0jbAebK4Z+D8TI863r9raMq1BjJCRSxksWHMwuEU0R//IMj+WL/ UxVrgl9lV6Q7kC3iry2yvUYprXpAOPGy6LIzwFZT+pkhBds+UnTrtSXEY8PaOrrWZ4e1hlKms3K h2kgbtje7jDOLFz5tRQ0fyE+PJ4ucVlCzQinottJnE= X-Received: by 2002:a05:6820:83d2:10b0:6b9:7d9e:5708 with SMTP id 006d021491bc7-6c0b9e5081cmr5077670eaf.3.1789141334570; Fri, 11 Sep 2026 08:42:14 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.13 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:13 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 14/15] io_uring: add tracepoints for the handoff operation Date: Fri, 11 Sep 2026 09:41:04 -0600 Message-ID: <20260911154148.644489-15-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add tracepoints for a handoff, a handoff that didn't happen with the reason, and the promoted task resuming the submission. Signed-off-by: Jens Axboe --- include/trace/events/io_uring.h | 114 ++++++++++++++++++++++++++++++++ io_uring/handoff.c | 42 +++++++++--- 2 files changed, 147 insertions(+), 9 deletions(-) diff --git a/include/trace/events/io_uring.h b/include/trace/events/io_urin= g.h index 34b31a855ea4..6043c5d46dbd 100644 --- a/include/trace/events/io_uring.h +++ b/include/trace/events/io_uring.h @@ -671,6 +671,120 @@ TRACE_EVENT(io_uring_local_work_run, TP_printk("ring %p, count %d, loops %u", __entry->ctx, __entry->count, __= entry->loops) ); =20 +/** + * io_uring_handoff - a blocked submitter hands its identity to a worker + * + * @req: pointer to a submitted request + * @dst: the idle io-wq worker task taking over + */ +TRACE_EVENT(io_uring_handoff, + + TP_PROTO(struct io_kiocb *req, struct task_struct *dst), + + TP_ARGS(req, dst), + + TP_STRUCT__entry ( + __field( void *, ctx ) + __field( void *, req ) + __field( u64, user_data ) + __field( u8, opcode ) + __field( pid_t, src_pid ) + __field( pid_t, dst_pid ) + + __string( op_str, io_uring_get_opcode(req->opcode) ) + ), + + TP_fast_assign( + __entry->ctx =3D req->ctx; + __entry->req =3D req; + __entry->user_data =3D req->cqe.user_data; + __entry->opcode =3D req->opcode; + __entry->src_pid =3D task_pid_nr(current); + __entry->dst_pid =3D task_pid_nr(dst); + + __assign_str(op_str); + ), + + TP_printk("ring %p, request %p, user_data 0x%llx, opcode %s, identity %d = handed to worker %d", + __entry->ctx, __entry->req, __entry->user_data, + __get_str(op_str), __entry->src_pid, __entry->dst_pid) +); + +/** + * io_uring_handoff_fail - a handoff didn't happen for a request + * + * @req: pointer to the request being issued + * @reason: why. "lock", "prepare" and "worker" mean the task blocked in + * place, anything else that it took the io-wq punt path instead. + */ +TRACE_EVENT(io_uring_handoff_fail, + + TP_PROTO(struct io_kiocb *req, const char *reason), + + TP_ARGS(req, reason), + + TP_STRUCT__entry ( + __field( void *, ctx ) + __field( void *, req ) + __field( u64, user_data ) + __field( u8, opcode ) + + __string( op_str, io_uring_get_opcode(req->opcode) ) + __string( reason, reason ) + ), + + TP_fast_assign( + __entry->ctx =3D req->ctx; + __entry->req =3D req; + __entry->user_data =3D req->cqe.user_data; + __entry->opcode =3D req->opcode; + + __assign_str(op_str); + __assign_str(reason); + ), + + TP_printk("ring %p, request %p, user_data 0x%llx, opcode %s, %s", + __entry->ctx, __entry->req, __entry->user_data, + __get_str(op_str), __get_str(reason)) +); + +/** + * io_uring_handoff_resume - a promoted worker continues the submission + * + * @ctx: pointer to a ring context structure + * @req: the request that blocked, owned by the demoted task by now + * @worker: pid the demoted task now runs under + * @consumed: SQEs consumed by earlier handoffs of this syscall + * @to_submit: SQE count the syscall asked for + */ +TRACE_EVENT(io_uring_handoff_resume, + + TP_PROTO(void *ctx, void *req, pid_t worker, unsigned int consumed, + unsigned int to_submit), + + TP_ARGS(ctx, req, worker, consumed, to_submit), + + TP_STRUCT__entry ( + __field( void *, ctx ) + __field( void *, req ) + __field( pid_t, worker ) + __field( unsigned int, consumed ) + __field( unsigned int, to_submit ) + ), + + TP_fast_assign( + __entry->ctx =3D ctx; + __entry->req =3D req; + __entry->worker =3D worker; + __entry->consumed =3D consumed; + __entry->to_submit =3D to_submit; + ), + + TP_printk("ring %p, request %p now on worker %d, consumed %u, to_submit %= u", + __entry->ctx, __entry->req, __entry->worker, + __entry->consumed, __entry->to_submit) +); + #endif /* _TRACE_IO_URING_H */ =20 /* This part must be outside protection */ diff --git a/io_uring/handoff.c b/io_uring/handoff.c index 9c9bb7ba99f0..8aefea5d326e 100644 --- a/io_uring/handoff.c +++ b/io_uring/handoff.c @@ -20,6 +20,7 @@ #include #include #include +#include =20 #include "io_uring.h" #include "io-wq.h" @@ -52,28 +53,40 @@ bool io_handoff_possible(struct io_kiocb *req) return false; /* IOPOLL/SQPOLL issue differently, SQ_REWIND can't resume mid-batch */ if (ctx->flags & (IORING_SETUP_IOPOLL | IORING_SETUP_SQPOLL | - IORING_SETUP_SQ_REWIND)) + IORING_SETUP_SQ_REWIND)) { + trace_io_uring_handoff_fail(req, "ring"); return false; + } /* pollable files keep the nonblocking issue + poll retry path */ - if (io_file_can_poll(req)) + if (io_file_can_poll(req)) { + trace_io_uring_handoff_fail(req, "poll"); return false; + } /* FMODE_NOWAIT files have a working nonblocking path, keep using it */ if ((def->pollin || def->pollout) && req->file && - (req->file->f_mode & FMODE_NOWAIT)) + (req->file->f_mode & FMODE_NOWAIT)) { + trace_io_uring_handoff_fail(req, "nowait-file"); return false; + } if (!tctx->io_wq) return false; /* an intermediate task's own user state doesn't matter, it stays */ - if (!tctx->handoff.src && !thread_handoff_allowed(current)) + if (!tctx->handoff.src && !thread_handoff_allowed(current)) { + trace_io_uring_handoff_fail(req, "task"); return false; + } /* the SQ head is published while we may still be running */ - if (io_req_sqe_copy(req, IO_URING_F_INLINE)) + if (io_req_sqe_copy(req, IO_URING_F_INLINE)) { + trace_io_uring_handoff_fail(req, "sqe"); return false; + } req->flags |=3D REQ_F_HANDOFF; check_spare: /* have a worker ready to take over */ - if (!io_wq_handoff_spare(tctx->io_wq, !io_req_unbound(req), false)) + if (!io_wq_handoff_spare(tctx->io_wq, !io_req_unbound(req), false)) { + trace_io_uring_handoff_fail(req, "spare"); return false; + } return true; } =20 @@ -122,8 +135,10 @@ bool __io_handoff_begin(struct io_kiocb *req) if (!io_handoff_possible(req)) return false; /* would interrupt the issue right away, and can't be handled here */ - if (task_sigpending(current)) + if (task_sigpending(current)) { + trace_io_uring_handoff_fail(req, "signal"); return false; + } =20 ho->req =3D req; io_handoff_block_signals(ho); @@ -240,10 +255,14 @@ void io_uring_task_sleeping(struct task_struct *tsk) WARN_ON_ONCE(tsk !=3D current); =20 /* the issue path is touching state that needs the ring lock held */ - if (ctx->submit_lock_depth) + if (ctx->submit_lock_depth) { + trace_io_uring_handoff_fail(req, "lock"); return; - if (src =3D=3D tsk && !thread_handoff_prepare(tsk)) + } + if (src =3D=3D tsk && !thread_handoff_prepare(tsk)) { + trace_io_uring_handoff_fail(req, "prepare"); return; + } =20 /* don't let the woken worker preempt us before we've committed */ preempt_disable(); @@ -251,6 +270,7 @@ void io_uring_task_sleeping(struct task_struct *tsk) dst =3D io_wq_handoff_claim(tctx->io_wq, bound, io_handoff_resume, src); if (!dst) { preempt_enable(); + trace_io_uring_handoff_fail(req, "worker"); return; } =20 @@ -262,6 +282,7 @@ void io_uring_task_sleeping(struct task_struct *tsk) /* our accounting follows the identity, an intermediate's doesn't */ if (src =3D=3D tsk) thread_handoff_stats_take(&ho->stats); + trace_io_uring_handoff(req, dst); =20 io_handoff_release_ring(ctx, ho); io_handoff_move_tctx(tctx, tsk, dst); @@ -306,6 +327,9 @@ static long io_handoff_resume(void) bool bound =3D ho->bound; long ret; =20 + trace_io_uring_handoff_resume(ctx, ho->req, task_pid_nr(prev), + ho->consumed, ho->to_submit); + /* enough of the identity to issue requests on its behalf */ thread_handoff_adopt_creds(src); put_task_struct_many(prev, ho->prev_refs); --=20 2.55.0 From nobody Fri Sep 25 13:53:56 2026 Received: from mail-oo2-f43.google.com (mail-oo2-f43.google.com [74.125.231.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C7F89496D26 for ; Fri, 11 Sep 2026 15:42:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.231.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141340; cv=none; b=Str21fgi6XzfBD1yWj4bnN7MFwCH7zhrQWhblYPQKkgwzlQZ8pyOYtUUIPy0D6oo08kuVEX6MOU4zYap31RoDYjBogX+B9zIkx5k9c4GmSw7BsewGhQIJMR/TAIGTj1U71AFXU3fNvOib63TC4/3bgofKSqTHIGrX6IBGvaSzPY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789141340; c=relaxed/simple; bh=5A1wu8NEeK9mJoM9/OY+CglS8dpgeZrlo6DgdNg46HM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dTQoYbpVAKq6Dj/qulAbfnOKf/JpHEiENewwVHIzUl6GJT2zBJ4PMuo0cTIyftSEO9ZFOC34AXh6QLiyGUvfuB0s6vykYJYDytNPPJhFSkb/zS7PlJaIH5dYeGm+C21XbrGVSPST5wtA+lgtawGMRW3KbxVWEQe3m8ouj99EK2U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=GdCuJ436; arc=none smtp.client-ip=74.125.231.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="GdCuJ436" Received: by mail-oo2-f43.google.com with SMTP id 006d021491bc7-6b1ae6b200fso713332eaf.1 for ; Fri, 11 Sep 2026 08:42:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1789141336; x=1789746136; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=UfA4gKLRjMNP2H10G64YdooBaWutPGm7Kf7SsyC2UwA=; b=GdCuJ436yBQPGmZmv0GkA1LKAVy0QL/LDG1+alUkOkMSCpTPAhG+dKJPJuNkbCQyge vQk4wFjY8CFZXuXB6VVANodNvSNNRHNnlw4pf26ZOWq6E98n7o5oDEQqU1WNwkjSiQjm olNH8hDPutJiP2MDuKB+Ti6NF1eyZ4DK38nQXW+agzjJaAhc8cZML7+kSv2oRO/PIYAK LnfjFm+vhMu02NiBVZXc6qCw7leT3hsjaIpI8pca0rGSYCHNn3Wb0OWIAfTTY1untzVF uIAypQj7+5jjIXYGptlEBXPBDaBS9hWpffmyvGXJkpLkWIQIG65qVS6zIAg91rKhgwcs olPg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789141336; x=1789746136; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=UfA4gKLRjMNP2H10G64YdooBaWutPGm7Kf7SsyC2UwA=; b=Gibd8w7SAWTh1pU7/uvw2wCPIQ0ieeZH39B1ItfyeQ3hSuRusgR/cSxEq4xzksJbrP o+eG0BKmDbMxE+ozhcByPVhpi8wKoiGgI0C86HDd77MRpZyvFaT2eiNVsI+oYMAo55XX lueZ89G3D9QHTUnzTvG9uGqjIokbAd0maQu0H7gude6UTTxIv5IU9QSpseU4zRw20gDX caosh9j2FcO/Sr0FL0Lmj1LTCViJtpQXSyGZaXYYhnAg6f7YT3UMAF5FmZPKFFpPN4t/ RF3b+xuFxNFegGD9xAwEWp6G5QeU9BnslX5aDEv490z54xNRjduEtBZAmLn4FYPHaHrk NQIA== X-Forwarded-Encrypted: i=1; AKwUvBztVaUkzf8/V+XRqeX8Ar/yhmOO1YG+oJB+3ayOnL/gvhU5MX+bCZ5E/BaEAVtmKHTyplE4Vkqasqs9QeI=@vger.kernel.org X-Gm-Message-State: AFuF++kV6EJulhx2cKHNsAIyrW5XMBeQFU8qCYdEBo2x8IXpRJvki7/X bl6hYWARgs9CFf5VAEQeesn1ZDwq+WJ4PHM6/PC+3oUDqXYTWEwJfq8DvIJSCYl1vRc= X-Gm-Gg: AYBFou2sH/TTMgv3LWwjTQZuAj48bN3zrq+GXYJ5lp8L3ubDRKaY9y/Y2/fKNkv50S8 7QI5HFjETnuExijcg0QUet9EwpeawFNRk/XWOWZ8itQlgAa6eYpgOQztqrGDQZX1gNskkI3+Lwx cL/aFW3Hcq9zVx2waiihAsoOQperrhCB5NyIzBbruprScdvNXykI00p2sRC2bI4khISH7K9RHRy 80LFpusF+8k3lBQY/zkWa+VQnqgKHV1uGc0/4GaL9JilIgBZgsMIjnLCOttlQ61xhsizrOgVWqS vfujuZLSVX0oVJMWb+hmOs/qsg5SpIV/UXY0sziyGTGf5O1/Y3phWuzR2zOHgrwmc3tU5xqf1Cb SS16AC993/IlV2ZrsDwfq1pojGdvvOPjsvS3qFTco6TwUy6FDxdIZ50yt76ljHu/Nhjp1u404Ra lWiTJpMaxZJ93RE9kMi/gSFOefjMKhaZMJWbPFgNJGYAJQomiuEXAn2nWzbnu/ZPy8KVby8Ycyx oueA7A/6L/Ur18w8/wwx3F6IGChQQiJjLmmqZkpajo= X-Received: by 2002:a05:6820:f00b:b0:6bd:cd33:b885 with SMTP id 006d021491bc7-6c0bc4c8a27mr3367535eaf.42.1789141335900; Fri, 11 Sep 2026 08:42:15 -0700 (PDT) Received: from m2max ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 006d021491bc7-6c09690af1dsm2802199eaf.1.2026.09.11.08.42.14 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 11 Sep 2026 08:42:15 -0700 (PDT) From: Jens Axboe To: io-uring@vger.kernel.org Cc: linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, tglx@kernel.org, mingo@redhat.com, peterz@infradead.org, Jens Axboe Subject: [PATCH 15/15] io_uring: issue IOSQE_ASYNC requests inline when a handoff is possible Date: Fri, 11 Sep 2026 09:41:05 -0600 Message-ID: <20260911154148.644489-16-axboe@kernel.dk> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260911154148.644489-1-axboe@kernel.dk> References: <20260911154148.644489-1-axboe@kernel.dk> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" NOTE: undecided if this should be the behavior. Signed-off-by: Jens Axboe --- include/linux/io_uring_types.h | 3 --- io_uring/io_uring.c | 5 ----- 2 files changed, 8 deletions(-) diff --git a/include/linux/io_uring_types.h b/include/linux/io_uring_types.h index 37c56ad37e05..dadbe3655138 100644 --- a/include/linux/io_uring_types.h +++ b/include/linux/io_uring_types.h @@ -651,7 +651,6 @@ enum { REQ_F_IMPORT_BUFFER_BIT, REQ_F_SQE_COPIED_BIT, REQ_F_IOPOLL_BIT, - REQ_F_ASYNC_USER_BIT, REQ_F_HANDOFF_BIT, =20 /* not a real bit, just to check we're not overflowing the space */ @@ -748,8 +747,6 @@ enum { REQ_F_SQE_COPIED =3D IO_REQ_FLAG(REQ_F_SQE_COPIED_BIT), /* request must be iopolled to completion (set in ->issue()) */ REQ_F_IOPOLL =3D IO_REQ_FLAG(REQ_F_IOPOLL_BIT), - /* IOSQE_ASYNC was set on the SQE, not just by prep */ - REQ_F_ASYNC_USER =3D IO_REQ_FLAG(REQ_F_ASYNC_USER_BIT), /* vetted at submit for an inline blocking issue with a handoff */ REQ_F_HANDOFF =3D IO_REQ_FLAG(REQ_F_HANDOFF_BIT), }; diff --git a/io_uring/io_uring.c b/io_uring/io_uring.c index 7c2aa0cfacc8..04e2fbdcb3dd 100644 --- a/io_uring/io_uring.c +++ b/io_uring/io_uring.c @@ -1771,8 +1771,6 @@ static int io_init_req(struct io_ring_ctx *ctx, struc= t io_kiocb *req, /* same numerical values with corresponding REQ_F_*, safe to copy */ sqe_flags =3D READ_ONCE(sqe->flags); req->flags =3D (__force io_req_flags_t) sqe_flags; - if (sqe_flags & IOSQE_ASYNC) - req->flags |=3D REQ_F_ASYNC_USER; req->cqe.user_data =3D READ_ONCE(sqe->user_data); req->file =3D NULL; req->tctx =3D current->io_uring; @@ -1919,9 +1917,6 @@ static bool io_req_force_async(struct io_kiocb *req) return true; if (!(req->flags & REQ_F_FORCE_ASYNC)) return false; - /* userspace asked for it, keep the explicit offload */ - if (req->flags & REQ_F_ASYNC_USER) - return true; if (req->ctx->int_flags & IO_RING_F_DRAIN_ACTIVE) return true; /* the file decides on pollability, resolve it now if fixed */ --=20 2.55.0