From nobody Fri Jul 24 21:53:37 2026 Received: from galois.linutronix.de (Galois.linutronix.de [193.142.43.55]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 84B38439F78; Thu, 23 Jul 2026 10:21:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=193.142.43.55 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784802087; cv=none; b=kA3QQXmG8avlJaZHYowtoN9dS1D44n9mRYWHFP+l68UQpUMplfuhKqm8m7pgItQLnEABJoOuepiixe9mjp6C849Nez/N466QpfGeZF+p8DAIfl70AH38oclGMuyPn5IPVdbKfkNDM3W+bSzLGK/0Vp0cne3j9s0/k2HJso0qc+o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784802087; c=relaxed/simple; bh=duWKyY2Lg9SE2B79EUNuuOWTB7bIQtPAptIVCvMx9Hc=; h=Date:From:To:Subject:Cc:In-Reply-To:References:MIME-Version: Message-ID:Content-Type; b=tDrH/Z+Unh66qkHXu+7Z6LG/fL3VVIKpuCyCpeCbOiEX7E886G44HAePUdT49wsR1ch+ukyxCc9YZHYmuXm9bV4hjfgCZb6Ozm0KxvuX5IAyncPn1FdnPLE33gew2Xi3krBSUc6KE+MhU6rwWb8audByk+I8VIxbC2LzrZHiTl4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de; spf=pass smtp.mailfrom=linutronix.de; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=B90SjylG; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=q0IxWNOS; arc=none smtp.client-ip=193.142.43.55 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linutronix.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="B90SjylG"; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="q0IxWNOS" Date: Thu, 23 Jul 2026 10:21:22 -0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020; t=1784802084; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=7x2N9VSVk3F2LuhIs0/LfIjfQ5Tf+ya2hz2/Inpm3a0=; b=B90SjylGha265EKtpPy4UFZaJpOwKAMNSwv9UP3Lz/489oZiFErzk//6aVTQat7A07RHWN 3CSZN+ZvT2FrZKoSSVgq+Lc44t84UTQ19ClNpIGH9Dy6Da3dsOhygQyldQHz8wz3c/DV74 jBSKrGXVlxhS15tv/Yoh/hAiOESuj9vR2ycoH1aBknsg1iksD+2QofiJiAyyCuH4JNviUc qjYug0yMf+45vNxf30M95/AOtI6CpstB8vdJzM6x168pe8pz0smppCaaf3rizyO3EmZ2oI Zlcob77DVrPFaodQ5T/ZxT1R6J1p1fDMkM3zbZoHFLKbhIXqagznkDyWk6qfnQ== DKIM-Signature: v=1; a=ed25519-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020e; t=1784802084; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=7x2N9VSVk3F2LuhIs0/LfIjfQ5Tf+ya2hz2/Inpm3a0=; b=q0IxWNOSXtx035zJR/n28+U/jWkd4UrsLW8xvEvspz+x8VzUW2YkWkhxhWCWI16iL1Ce9U RVQMmsJ2UbqtG8DA== From: "tip-bot2 for Chuyi Zhou" Sender: tip-bot2@linutronix.de Reply-to: linux-kernel@vger.kernel.org To: linux-tip-commits@vger.kernel.org Subject: [tip: smp/core] smp: Use task-local IPI cpumask in smp_call_function_many_cond() Cc: Chuyi Zhou , Thomas Gleixner , "Paul E. McKenney" , Sebastian Andrzej Siewior , x86@kernel.org, linux-kernel@vger.kernel.org In-Reply-To: <20260709122933.4021501-5-zhouchuyi@bytedance.com> References: <20260709122933.4021501-5-zhouchuyi@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Message-ID: <178480208231.2943223.11680828613174536505.tip-bot2@tip-bot2> Robot-ID: Robot-Unsubscribe: Contact to get blacklisted from these emails Precedence: bulk Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable The following commit has been merged into the smp/core branch of tip: Commit-ID: 9a560af15fd31b3b93235b1584a2509445a17963 Gitweb: https://git.kernel.org/tip/9a560af15fd31b3b93235b1584a250944= 5a17963 Author: Chuyi Zhou AuthorDate: Thu, 09 Jul 2026 20:29:23 +08:00 Committer: Thomas Gleixner CommitterDate: Thu, 16 Jul 2026 09:24:55 +02:00 smp: Use task-local IPI cpumask in smp_call_function_many_cond() smp_call_function_many_cond() uses the per-CPU cfd->cpumask as the list of remote CPUs to wait for. That is safe while the caller remains pinned to the current CPU for the whole operation, because another task cannot run on the same CPU and reuse the per-CPU mask. The synchronous wait is the long-latency part of the operation. To make that wait preemptible, the mask iterated by csd_lock_wait() must remain stable even if the task is preempted or migrates. If the wait used the per-CPU cfd->cpumask after dropping CPU pinning, another task scheduled on the original CPU could enter smp_call_function_many_cond() and overwrite the mask while the first task is still iterating it. Give each task private IPI cpumask storage and use it as the wait mask in smp_call_function_many_cond(). Other cpumask storage choices do not fit this use case: - Per-CPU storage is the state that becomes unsafe once the wait is made preemptible. After the caller drops CPU pinning, another task scheduled on the original CPU can enter smp_call_function_many_cond() and reuse the same per-CPU mask. - Stack storage is not suitable for large NR_CPUS or CONFIG_CPUMASK_OFFSTACK=3Dy configurations. The wait mask needs to scale with cpumask_size(), and putting that storage on the stack is not acceptable on large systems. - Allocating the mask inside smp_call_function_many_cond() would put an allocation and a failure path in the generic IPI path. A sleeping allocation is not suitable because callers have historically only provided a preempt-disabled context, not a sleepable one. GFP_ATOMIC would avoid sleeping, but a failure fallback would make the latency improvement opportunistic instead of guaranteed. The users are not limited to a small, pre-identifiable class of tasks. On x86, ordinary tasks can reach this path through TLB flushes during exit, unmap and reclaim, so allocating the mask only for a known subset of tasks is not straightforward. The memory cost is explicit: one word is added to task_struct. When cpumask_size() fits in that word, the mask is stored inline and no separate allocation is needed. Larger systems allocate cpumask_size() per task; on x86-64 NR_CPUS=3D8192 this is 1 KiB per task. For context, x86 already carries several KiB of per-task architecture and FPU state, depending on the enabled features and configuration. That does not make the extra cpumask free, but it puts the large-NR_CPUS case in perspective. Signed-off-by: Chuyi Zhou Signed-off-by: Thomas Gleixner Tested-by: Paul E. McKenney Reviewed-by: Sebastian Andrzej Siewior Link: https://patch.msgid.link/20260709122933.4021501-5-zhouchuyi@bytedance= .com --- include/linux/sched.h | 12 +++++++- include/linux/smp.h | 12 +++++++- kernel/fork.c | 9 ++++- kernel/smp.c | 71 +++++++++++++++++++++++++++++++++++++----- 4 files changed, 95 insertions(+), 9 deletions(-) diff --git a/include/linux/sched.h b/include/linux/sched.h index 373bcc0..5738c54 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -823,6 +823,17 @@ struct kmap_ctrl { #endif }; =20 +#if defined(CONFIG_SMP) && defined(CONFIG_PREEMPTION) +struct task_ipi_mask { + union { + cpumask_t *ipi_mask_ptr; + unsigned long ipi_mask_val; + }; +}; +#else +struct task_ipi_mask { }; +#endif + struct task_struct { #ifdef CONFIG_THREAD_INFO_IN_TASK /* @@ -1359,6 +1370,7 @@ struct task_struct { struct list_head perf_event_list; struct perf_ctx_data __rcu *perf_ctx_data; #endif + struct task_ipi_mask __private ipi_mask; #ifdef CONFIG_DEBUG_PREEMPT unsigned long preempt_disable_ip; #endif diff --git a/include/linux/smp.h b/include/linux/smp.h index 11e36c7..2dfa739 100644 --- a/include/linux/smp.h +++ b/include/linux/smp.h @@ -238,6 +238,18 @@ static inline int get_boot_cpu_id(void) =20 #endif /* !SMP */ =20 +#if defined(CONFIG_PREEMPTION) && defined(CONFIG_SMP) +int smp_task_ipi_mask_alloc(struct task_struct *task); +void smp_task_ipi_mask_free(struct task_struct *task); +#else +static inline int smp_task_ipi_mask_alloc(struct task_struct *task) +{ + return 0; +} + +static inline void smp_task_ipi_mask_free(struct task_struct *task) { } +#endif + /* * raw_smp_processor_id() - get the current (unstable) CPU id * diff --git a/kernel/fork.c b/kernel/fork.c index 13e38e8..ac3fc49 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -537,6 +537,7 @@ void free_task(struct task_struct *tsk) #endif release_user_cpus_ptr(tsk); scs_release(tsk); + smp_task_ipi_mask_free(tsk); =20 #ifndef CONFIG_THREAD_INFO_IN_TASK /* @@ -935,10 +936,14 @@ static struct task_struct *dup_task_struct(struct tas= k_struct *orig, int node) #endif account_kernel_stack(tsk, 1); =20 - err =3D scs_prepare(tsk, node); + err =3D smp_task_ipi_mask_alloc(tsk); if (err) goto free_stack; =20 + err =3D scs_prepare(tsk, node); + if (err) + goto free_ipi_mask; + #ifdef CONFIG_SECCOMP /* * We must handle setting up seccomp filters once we're under @@ -1011,6 +1016,8 @@ static struct task_struct *dup_task_struct(struct tas= k_struct *orig, int node) #endif return tsk; =20 +free_ipi_mask: + smp_task_ipi_mask_free(tsk); free_stack: exit_task_stack_account(tsk); free_thread_stack(tsk); diff --git a/kernel/smp.c b/kernel/smp.c index 5c05029..19fdee6 100644 --- a/kernel/smp.c +++ b/kernel/smp.c @@ -16,6 +16,7 @@ #include #include #include +#include #include #include #include @@ -806,6 +807,50 @@ int smp_call_function_any(const struct cpumask *mask, } EXPORT_SYMBOL_GPL(smp_call_function_any); =20 +static DEFINE_STATIC_KEY_FALSE(ipi_mask_inlined); + +#ifdef CONFIG_PREEMPTION + +int smp_task_ipi_mask_alloc(struct task_struct *task) +{ + if (static_branch_unlikely(&ipi_mask_inlined)) + return 0; + + ACCESS_PRIVATE(task, ipi_mask).ipi_mask_ptr =3D + kmalloc(cpumask_size(), GFP_KERNEL); + if (!ACCESS_PRIVATE(task, ipi_mask).ipi_mask_ptr) + return -ENOMEM; + + return 0; +} + +void smp_task_ipi_mask_free(struct task_struct *task) +{ + if (static_branch_unlikely(&ipi_mask_inlined)) + return; + + kfree(ACCESS_PRIVATE(task, ipi_mask).ipi_mask_ptr); +} + +static cpumask_t *smp_task_ipi_mask(struct task_struct *cur) +{ + /* + * If cpumask_size() is smaller than or equal to the pointer + * size, it stashes the cpumask in the pointer itself to + * avoid extra memory allocations. + */ + if (static_branch_unlikely(&ipi_mask_inlined)) + return (cpumask_t *)&ACCESS_PRIVATE(cur, ipi_mask).ipi_mask_val; + + return ACCESS_PRIVATE(cur, ipi_mask).ipi_mask_ptr; +} +#else +static cpumask_t *smp_task_ipi_mask(struct task_struct *cur) +{ + return NULL; +} +#endif + /* * Flags to be used as scf_flags argument of smp_call_function_many_cond(). * @@ -821,13 +866,21 @@ static void smp_call_function_many_cond(const struct = cpumask *mask, smp_cond_func_t cond_func) { int cpu, last_cpu, this_cpu =3D smp_processor_id(); - struct call_function_data *cfd; + struct cpumask *cpumask, *task_mask; bool wait =3D scf_flags & SCF_WAIT; - int nr_cpus =3D 0; + struct call_function_data *cfd; bool run_remote =3D false; + int nr_cpus =3D 0; =20 lockdep_assert_preemption_disabled(); =20 + cfd =3D this_cpu_ptr(&cfd_data); + task_mask =3D smp_task_ipi_mask(current); + if (task_mask) + cpumask =3D task_mask; + else + cpumask =3D cfd->cpumask; + /* * Can deadlock when called with interrupts disabled. * We allow cpu's that are not yet online though, as no one else can @@ -848,16 +901,15 @@ static void smp_call_function_many_cond(const struct = cpumask *mask, =20 /* Check if we need remote execution, i.e., any CPU excluding this one. */ if (cpumask_any_and_but(mask, cpu_online_mask, this_cpu) < nr_cpu_ids) { - cfd =3D this_cpu_ptr(&cfd_data); - cpumask_and(cfd->cpumask, mask, cpu_online_mask); - __cpumask_clear_cpu(this_cpu, cfd->cpumask); + cpumask_and(cpumask, mask, cpu_online_mask); + __cpumask_clear_cpu(this_cpu, cpumask); =20 cpumask_clear(cfd->cpumask_ipi); - for_each_cpu(cpu, cfd->cpumask) { + for_each_cpu(cpu, cpumask) { call_single_data_t *csd =3D per_cpu_ptr(cfd->csd, cpu); =20 if (cond_func && !cond_func(cpu, info)) { - __cpumask_clear_cpu(cpu, cfd->cpumask); + __cpumask_clear_cpu(cpu, cpumask); continue; } =20 @@ -908,7 +960,7 @@ static void smp_call_function_many_cond(const struct cp= umask *mask, } =20 if (run_remote && wait) { - for_each_cpu(cpu, cfd->cpumask) { + for_each_cpu(cpu, cpumask) { call_single_data_t *csd; =20 csd =3D per_cpu_ptr(cfd->csd, cpu); @@ -1022,6 +1074,9 @@ EXPORT_SYMBOL(nr_cpu_ids); void __init setup_nr_cpu_ids(void) { set_nr_cpu_ids(find_last_bit(cpumask_bits(cpu_possible_mask), NR_CPUS) + = 1); + + if (IS_ENABLED(CONFIG_PREEMPTION) && cpumask_size() <=3D sizeof(unsigned = long)) + static_branch_enable(&ipi_mask_inlined); } =20 /* Called by boot processor to activate the rest. */