include/linux/sched.h | 31 ++- init/init_task.c | 5 + kernel/fork.c | 5 + kernel/locking/mutex.c | 57 +++- kernel/sched/core.c | 580 ++++++++++++++++++++++++++++++++++++++-- kernel/sched/deadline.c | 15 +- kernel/sched/fair.c | 16 +- kernel/sched/sched.h | 6 + kernel/sched/stats.h | 2 +- 9 files changed, 666 insertions(+), 51 deletions(-)
Hello folks,
I had promised John I would share an insane idea if I got it
working(ish) so this RFC presents an alternate approach to handle
enqueued donor wakeup vs owner's chain-wakeup race via ttwu_runnable()
path and make the activate path as light as possible for tasks that
don't need to do a chain wakeup.
The series essentially breaks down John's large patch at
https://lore.kernel.org/lkml/20260807035232.1881495-9-jstultz@google.com/
into smaller chunks while introducing the new approach and fixing a few
snags encountered along the way (Patch 1 - Patch 4 are fixes that can be
discussed without getting into the rest of the RFC).
Disclaimer: Series in not bisectible in any way at the moment - related
bits are introduced together at and intermediate builds may fail.
Introduction
============
Handling sleeping owner requires proxy-chains to queue on the said owner
and perform a chain activation when the blocked owner wakes up.
Since a blocked donor can be woken up by another concurrent wake event,
the activation path becomes complicated and wakeup has to take an extra
lock (p->blocked_lock) to prevent any modifications to "p->blocked_head"
and miss activating any blocked donors.
The nesting rules for the locks are as follows:
p->pi_lock
__task_rq_lock()
lock->wait_lock
p->blocked_lock
Current chain-wakeup scheme juggles p->pi_lock, and p->blocked_lock to
prevent concurrent wake events from modifying the chain.
This RFC introduces a __task_rq_lock() based scheme that requires just
the __task_rq_lock() (always held) + p->blocked_lock (juggled for each
owner in chain) when queued tasks are detected to minimize the lock
bouncing and introduces a no-extra-locking fast-path for activation.
Trade offs are discussed at the end.
Approach
========
To implement this, we try a bunch of bad ideas:
o Waking up blocked donors linked to sleeping owner via ttwu_runnable()
To prevent chain-wakeup transitioning "p->on_rq" to QUEUED while a
concurrent wakeup is in progress, "p->on_rq" is unionized with a new
state "p->is_linked" as:
union {
struct {
u8 on_rq; /* Task on CPU */
u8 is_linked; /* Task on sleeping owner */
};
u16 needs_rq_sync;
}
Instead of checking READ_ONCE(p->on_rq), to send task down the
ttwu_runnable() route, the new "needs_rq_sync" indicator is used which
can be atomically inspected outside rq_lock.
In ttwu_runnable() a fast-path first tries to remove the blocked donor
from the sleeping owner by just grabbing the "owner->blocked_lock" and
forces a full __task_rq_lock() if the blocked donor's state has changed
while trying to grab the blocked lock.
The chain-wakeup path takes a single __task_rq_lock() to prevent any
race from concurrent wakeups (described later).
o Proactively removing the blocked donor from owner if it wakes up
Proactively remove the blocked donor from sleeping owner in
proxy_enqueue_on_owner() if owner->on_rq is transitioned to != 0. This
ensures owner need not care about any additional donors that come in
after it transitions p->on_rq.
proxy_enqueue_on_owner() holds lock->wait_lock + owner->blocked_lock.
Since owner must grab the wait_lock again to wake up the blocked donor,
it is guaranteed that owner cannot disappear and get_task_struct() /
put_task_struct() dance is avoided.
o Migrating enqueued donors while blocking (SLEEP | MIGRATING)
To ensure a single __task_rq_lock() can handle the entire chain, a
blocked donor is migrated to a deterministic CPU before it is
transitioned to p->on_rq = 0. This ensures the whole chain is owned by a
single CPU.
o No delayed tasks on chain
Delayed tasks pose a challenge that they are still on_rq and require
special handling to ensure the above condition of chain resolving to a
single CPU holds true. For simplicity, all delayed tasks on the chain
are fully blocked before they are queued.
This can be handled differently but I got a bit busy lately so that will
be handled in RFC v2 (if folks think this approach holds water).
o Single shot chain-wakeup (ab)using TASK_ON_RQ_MIGRATING
Tasks on chain are transitioned from being blocked to
TASK_ON_RQ_MIGRATING under __task_rq_lock(). Tasks on ttwu_runnable()
path that are trying to grab __task_rq_lock() is stalled until the
task is woken up on destination rq and the rq_lock() there is dropped.
o Tracking task that can potentially resolve to __mutex_owner()
This indicator is used to selectively set the lock holding
tasks to the slow-path.
Interleaving
============
How are different races handled (individual patches have these
highlighted too, either in comment message, or via code comments):
o Owner wakeup vs donor enqueue
activate_task(owner)
owner->on_rq = MIGRATING proxy_enqueue_on_owner()
guard(raw_spinlock)(&owner->blocked_lock)
list_add(donor, owner->blocked_head)
smp_mb() /* MIGRATING is visible */ smp_mb() /* List addition is visible */
if (list_empty(&owner->blocked_head)) if (!READ_ONCE(owner->on_rq))
return __activate_task(owner) block_task(donor)
... /* Do chain-activation */ __proxy_dequeue_from_owner(donor)
o Chain-wakeup vs donor wakeup
proxy_activate_donor_task() try_to_wake_up(donor)
rq_lock(rq_of(onwner->blocked_cpu)) proxy_try_dequeue_from_owner(p)
guard(raw_spinlock)(&owner->blocked_lock) guard(raw_spinlock)(&owner->blocked_lock)
for_each(donor, owner->blocked_head)
donor->on_rq = MIGRATING;
smp_mb(); if (donor->on_rq)
goto slowpath;
... /* More stuff */
rq_unlock(rq_of(onwner->blocked_cpu)); __task_rq_lock(donor)
if (donor->on_rq == MIGRATING)
cpu_relax();
__proxy_dequeue_from_owner(donor)
__activate_task(donor)
p->on_rq = QUEUED; /* Acquired */
if (p->is_linked)
proxy_dequeue_from_owner(donor)
/* Do wakeup. */
Chain wakeup is covered extensively in Patch 14.
Series breakdown
================
Patch 01-04: Fixes to existing bits (Patch 4 needs more thinking)
Patch 05-07: Donor queuing with some convenience added to skip delayed
handling later on.
Patch 08-13: Putting tasks on ttwu_running() path and swapping
task_cpu() while blocking
Patch 14 : Chain wake-up
Patch 15-16: Optimizations to limit slowpath.
Trade-off
=========
Advantages:
o Only need to juggle blocked_lock(s) under a single rq_lock().
o Re-use ttwu_runnable bit to handle removal of proxy donor.
o No need to do get_task_stuct() / put_task_struct() juggling.
Disadvantages:
o Possible increased rq_lock contention on the chain-wakeup path but
those events are generally rare.
o Lack of delayed task handling since the src_rq need to be locked
separately to migrate it.
For the delayed handling, it is possible to use proxy_migrate_task() to
move the task after blocking. Series does not implement this yet to
limit the number of bad ideas.
How to apply
============
This series is based on John's tree at
https://github.com/johnstultz-work/linux-dev.git proxy-exec-v31-7.2-rc4
at commit 18bea15ce9fa ("sched: Migrate whole chain in
proxy_migrate_task()"). For convenience, a tree is present at:
https://github.com/kudureranganath/linux.git sched/proxy/sleeping_owner_rfc_v1
Sorry in advance to folks whose VPNs hate dealing with Github but I'm
told pulling it locally and inspecting the patches on tree works better.
Testing
=======
Tested with perf bench sched messaging, test-ww_mutex, and a kernel
module that implements a duty-cycle with sleep_ms() + busy-loop in
kernel.
I made sure to build and test !SCHED_PROXY_EXEC version this time around
to save John some trouble :-)
Acknowledgement
===============
Most of the good work is by John and all the bad ideas are by me.
Sorry in advance.
---
K Prateek Nayak (16):
sched/core: Break activation of blocked task into a separate helper
sched/core: Use enqueue/dequeue flags instead of
task_on_rq_migrating()
sched/fair: Use enqueue flags for DO_ATTACH in update_load_avg()
sched/core: Activate blocked donor when no owner is found
sched/core: Do not queue blocked donor on a delayed owner
sched/core: Queue blocked donor onto sleeping owner for chain-wakeup
sched/core: Avoid delaying blocked donors queued on sleeping owner
sched/deadline: Prepare for blocking and proxy activation with
MIGRATING flag
sched/core: Track CPU where the task was blocked on
sched/core: Introduce p->is_linked to track if task is queued on
sleeping owner
sched/core: Prepare to inspect ->is_linked alongside ->on_rq during
wakeup
sched:core: Add MIGRATING flags when blocking and activating linked
donors
sched/core: Use p->is_linked state to unlink from sleeping owner early
sched/core: Introduce chain-wakeup to activate blocked donors
locking/mutex: Track locks owned by a task in a per-task counter
sched/core: Set activation of non lock-holders to fast-path
include/linux/sched.h | 31 ++-
init/init_task.c | 5 +
kernel/fork.c | 5 +
kernel/locking/mutex.c | 57 +++-
kernel/sched/core.c | 580 ++++++++++++++++++++++++++++++++++++++--
kernel/sched/deadline.c | 15 +-
kernel/sched/fair.c | 16 +-
kernel/sched/sched.h | 6 +
kernel/sched/stats.h | 2 +-
9 files changed, 666 insertions(+), 51 deletions(-)
base-commit: 18bea15ce9faaa0c1bd827012e6c60dbf35578c7
--
2.34.1
On Tue, Aug 25, 2026 at 11:29 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote: > > I had promised John I would share an insane idea if I got it > working(ish) so this RFC presents an alternate approach to handle > enqueued donor wakeup vs owner's chain-wakeup race via ttwu_runnable() > path and make the activate path as light as possible for tasks that > don't need to do a chain wakeup. > > The series essentially breaks down John's large patch at > https://lore.kernel.org/lkml/20260807035232.1881495-9-jstultz@google.com/ > into smaller chunks while introducing the new approach and fixing a few > snags encountered along the way (Patch 1 - Patch 4 are fixes that can be > discussed without getting into the rest of the RFC). > > Disclaimer: Series in not bisectible in any way at the moment - related > bits are introduced together at and intermediate builds may fail. > > Introduction > ============ > > Handling sleeping owner requires proxy-chains to queue on the said owner > and perform a chain activation when the blocked owner wakes up. > > Since a blocked donor can be woken up by another concurrent wake event, > the activation path becomes complicated and wakeup has to take an extra > lock (p->blocked_lock) to prevent any modifications to "p->blocked_head" > and miss activating any blocked donors. > ... > Trade-off > ========= > > Advantages: > > o Only need to juggle blocked_lock(s) under a single rq_lock(). > o Re-use ttwu_runnable bit to handle removal of proxy donor. > o No need to do get_task_stuct() / put_task_struct() juggling. > > Disadvantages: > > o Possible increased rq_lock contention on the chain-wakeup path but > those events are generally rare. > o Lack of delayed task handling since the src_rq need to be locked > separately to migrate it. > > For the delayed handling, it is possible to use proxy_migrate_task() to > move the task after blocking. Series does not implement this yet to > limit the number of bad ideas. > Sorry again for being so slow to respond here! The series definitely looks interesting, and it seems to be holding up ok in testing. Though figuring out how to refactor them so that they are also bisectable looks like a real challenge (part of why my sleeping-owner enqueuing is basically one big patch)! That said, I can't say I've truly gotten my head fully around your series yet. I've definitely had way more time with my change, so I'm a little biased in feeling its somewhat more bounded (even though the activate_blocked_waiters() function and the multiple lists of tasks are very complicated). My initial sense of the downsides here with your series are: It adds a lot of new per-task state (blocked_cpu, is_linked/needs_rq_sync, sched_migrated_on_blocking, lock_nesting) to keep track of, and some of rules for the new state have dependencies (sched_migrated_on_blocking is tied to is_linked), and leveraging the ENQUEUE/DEQUEUE_MIGRATING flags feels a little subtle (I have often gotten the ENQUEUE/DEQUEUE flags wrong in my patch series, and unfortunately the general documentation around those flags is lagging a bit, so this adds to it). The locking being simpler is a clear benefit and definitely sound appealing - though I do agree the rq_lock hold time in proxy_activate_blocked_task() seems like it might be an issue. I've got a few more questions on specific patches, so I'll reply there. Again, it definitely is interesting and appears to be more integrated into the scheduler logic, so I'd expect that will give us better results then my maybe more isolated and tacked on activation logic. Have you gotten a sense of what Peter thinks of it? thanks -john
On Tue, Sep 15, 2026 at 10:22 PM John Stultz <jstultz@google.com> wrote: > The series definitely looks interesting, and it seems to be holding up > ok in testing. Though figuring out how to refactor them so that they Well, I did hit the following last night during stress testing: [11176.672951] WARNING: kernel/sched/sched.h:1645 at sched_change_begin+0x1b6/0x2b0, CPU#14: torture_shuffle/433 [11176.676650] CPU: 14 UID: 0 PID: 433 Comm: torture_shuffle Not tainted 7.2.0-rc4-00032-gf14e36dcc837-dirty #138 PREEMPT(full) [11176.680917] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.17.0-debian-1.17.0-1 04/01/2014 [11176.684794] RIP: 0010:sched_change_begin+0x1b6/0x2b0 [11176.686748] Code: c9 1e 48 89 03 80 7b 1c 00 0f 85 ca 00 00 00 80 7b 1d 00 0f 85 91 00 00 00 48 89 d8 5b 5d 41 5c 41 5d 41 5e e9 1b 63 23 01 90 <0f> 0b 90 41 f6 c5 08 0f 85 c7 fe ff ff 48 89 ef 41 83 cd 08 e8 b1 [11176.693913] RSP: 0018:ffffc900010b3e10 EFLAGS: 00010046 [11176.695962] RAX: 0000000000000000 RBX: ffff8881b9198020 RCX: 0000000000000001 [11176.698747] RDX: 0000000000000001 RSI: ffffffff82c8b519 RDI: ffffffff82cad7fe [11176.701554] RBP: ffff8881b91abdc0 R08: 0000000000000001 R09: 00000000d19ffe29 [11176.704561] R10: 000000000000000e R11: ffff88810264cff0 R12: ffff888106694400 [11176.707423] R13: 0000000000000002 R14: 0000000000000000 R15: ffffc90000017d58 [11176.710161] FS: 0000000000000000(0000) GS:ffff88823563e000(0000) knlGS:0000000000000000 [11176.713196] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [11176.715628] CR2: 0000000000000000 CR3: 00000001072a7006 CR4: 0000000000370ef0 [11176.718462] Call Trace: [11176.719460] <TASK> [11176.720319] __set_cpus_allowed_ptr_locked+0x146/0x1f0 [11176.722343] ? __pfx_torture_shuffle+0x10/0x10 [11176.724063] __set_cpus_allowed_ptr+0x64/0xa0 [11176.725913] set_cpus_allowed_ptr+0x3a/0x60 [11176.727513] torture_shuffle+0x12b/0x230 [11176.729014] kthread+0x103/0x130 [11176.730304] ? __pfx_kthread+0x10/0x10 [11176.731758] ret_from_fork+0x27c/0x340 [11176.733220] ? __pfx_kthread+0x10/0x10 [11176.734680] ret_from_fork_asm+0x1a/0x30 [11176.736366] </TASK> Followed shortly after by: [11663.223681] kernel BUG at kernel/sched/rt.c:1020! [11663.226093] Oops: invalid opcode: 0000 [#1] SMP NOPTI [11663.228669] CPU: 46 UID: 0 PID: 0 Comm: swapper/46 Tainted: G W 7.2.0-rc4-00032-gf14e36dcc837-dirty #138 PREEMPT(full) [11663.234796] Tainted: [W]=WARN [11663.236392] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.17.0-debian-1.17.0-1 04/01/2014 [11663.241196] RIP: 0010:dequeue_rt_stack+0x2ad/0x2c0 [11663.243652] Code: 83 a8 fe ff ff 39 c1 0f 8e 47 ff ff ff 90 0f 0b 90 e9 b7 fd ff ff 90 0f 0b 90 e9 e8 fd ff ff f3 48 0f bc d0 e9 80 fe ff ff 90 <0f> 0b 48 8d 3d 1a 79 dc 01 67 48 0f b9 3a e9 e9 fe ff ff 90 90 90 [11663.253981] RSP: 0018:ffffc90000c10dd8 EFLAGS: 00010046 [11663.256630] RAX: 0000000000000000 RBX: 0000000000000000 RCX: ffff88810ad50080 [11663.260225] RDX: ffff888235e3e000 RSI: 0000000000000000 RDI: ffff88810ad501c0 [11663.263862] RBP: ffff8881b99abdc0 R08: 0000000000000000 R09: ffffffff83b9a280 [11663.266876] R10: 000000000000002e R11: 0000000000000000 R12: ffff8881b99abdc0 [11663.269626] R13: ffff88810ad501c0 R14: 0000000000000000 R15: 0000000000000025 [11663.272412] FS: 0000000000000000(0000) GS:ffff888235e3e000(0000) knlGS:0000000000000000 [11663.276347] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [11663.279185] CR2: 00007f03955d60ac CR3: 00000001072a7006 CR4: 0000000000370ef0 [11663.282698] Call Trace: [11663.284061] <IRQ> [11663.285116] enqueue_task_rt+0x77/0x3f0 [11663.287276] enqueue_task+0x1cf/0x330 [11663.289121] proxy_activate_blocked_task+0x498/0x640 [11663.291579] ttwu_do_activate+0x287/0x290 [11663.293582] sched_ttwu_pending+0x10b/0x220 [11663.295726] __flush_smp_call_function_queue+0x1f8/0x680 [11663.298320] __sysvec_call_function_single+0x35/0x180 [11663.300877] sysvec_call_function_single+0x8a/0xb0 [11663.303287] </IRQ> [11663.304376] <TASK> [11663.305467] asm_sysvec_call_function_single+0x1a/0x20 [11663.307994] RIP: 0010:pv_native_safe_halt+0xf/0x20 The only change I made to your tree was passing the rf into proxy_activate_blocked_task, to use rq_(un)lock() instead of raw_spin_rq_(un)lock(). thanks -john
Hello John, On 9/16/2026 11:27 PM, John Stultz wrote: > On Tue, Sep 15, 2026 at 10:22 PM John Stultz <jstultz@google.com> wrote: >> The series definitely looks interesting, and it seems to be holding up >> ok in testing. Though figuring out how to refactor them so that they > > Well, I did hit the following last night during stress testing: > > [11176.672951] WARNING: kernel/sched/sched.h:1645 at > sched_change_begin+0x1b6/0x2b0, CPU#14: torture_shuffle/433 > [11176.676650] CPU: 14 UID: 0 PID: 433 Comm: torture_shuffle Not > tainted 7.2.0-rc4-00032-gf14e36dcc837-dirty #138 PREEMPT(full) > [11176.680917] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), > BIOS 1.17.0-debian-1.17.0-1 04/01/2014 > [11176.684794] RIP: 0010:sched_change_begin+0x1b6/0x2b0 > [11176.686748] Code: c9 1e 48 89 03 80 7b 1c 00 0f 85 ca 00 00 00 80 > 7b 1d 00 0f 85 91 00 00 00 48 89 d8 5b 5d 41 5c 41 5d 41 5e e9 1b 63 > 23 01 90 <0f> 0b 90 41 f6 c5 08 0f 85 c7 fe ff ff 48 89 ef 41 83 cd 08 > e8 b1 > [11176.693913] RSP: 0018:ffffc900010b3e10 EFLAGS: 00010046 > [11176.695962] RAX: 0000000000000000 RBX: ffff8881b9198020 RCX: 0000000000000001 > [11176.698747] RDX: 0000000000000001 RSI: ffffffff82c8b519 RDI: ffffffff82cad7fe > [11176.701554] RBP: ffff8881b91abdc0 R08: 0000000000000001 R09: 00000000d19ffe29 > [11176.704561] R10: 000000000000000e R11: ffff88810264cff0 R12: ffff888106694400 > [11176.707423] R13: 0000000000000002 R14: 0000000000000000 R15: ffffc90000017d58 > [11176.710161] FS: 0000000000000000(0000) GS:ffff88823563e000(0000) > knlGS:0000000000000000 > [11176.713196] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 > [11176.715628] CR2: 0000000000000000 CR3: 00000001072a7006 CR4: 0000000000370ef0 > [11176.718462] Call Trace: > [11176.719460] <TASK> > [11176.720319] __set_cpus_allowed_ptr_locked+0x146/0x1f0 > [11176.722343] ? __pfx_torture_shuffle+0x10/0x10 > [11176.724063] __set_cpus_allowed_ptr+0x64/0xa0 > [11176.725913] set_cpus_allowed_ptr+0x3a/0x60 > [11176.727513] torture_shuffle+0x12b/0x230 > [11176.729014] kthread+0x103/0x130 > [11176.730304] ? __pfx_kthread+0x10/0x10 > [11176.731758] ret_from_fork+0x27c/0x340 > [11176.733220] ? __pfx_kthread+0x10/0x10 > [11176.734680] ret_from_fork_asm+0x1a/0x30 > [11176.736366] </TASK> > > Followed shortly after by: > [11663.223681] kernel BUG at kernel/sched/rt.c:1020! > [11663.226093] Oops: invalid opcode: 0000 [#1] SMP NOPTI > [11663.228669] CPU: 46 UID: 0 PID: 0 Comm: swapper/46 Tainted: G > W 7.2.0-rc4-00032-gf14e36dcc837-dirty #138 PREEMPT(full) > [11663.234796] Tainted: [W]=WARN > [11663.236392] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), > BIOS 1.17.0-debian-1.17.0-1 04/01/2014 > [11663.241196] RIP: 0010:dequeue_rt_stack+0x2ad/0x2c0 > [11663.243652] Code: 83 a8 fe ff ff 39 c1 0f 8e 47 ff ff ff 90 0f 0b > 90 e9 b7 fd ff ff 90 0f 0b 90 e9 e8 fd ff ff f3 48 0f bc d0 e9 80 fe > ff ff 90 <0f> 0b 48 8d 3d 1a 79 dc 01 67 48 0f b9 3a e9 e9 fe ff ff 90 > 90 90 > [11663.253981] RSP: 0018:ffffc90000c10dd8 EFLAGS: 00010046 > [11663.256630] RAX: 0000000000000000 RBX: 0000000000000000 RCX: ffff88810ad50080 > [11663.260225] RDX: ffff888235e3e000 RSI: 0000000000000000 RDI: ffff88810ad501c0 > [11663.263862] RBP: ffff8881b99abdc0 R08: 0000000000000000 R09: ffffffff83b9a280 > [11663.266876] R10: 000000000000002e R11: 0000000000000000 R12: ffff8881b99abdc0 > [11663.269626] R13: ffff88810ad501c0 R14: 0000000000000000 R15: 0000000000000025 > [11663.272412] FS: 0000000000000000(0000) GS:ffff888235e3e000(0000) > knlGS:0000000000000000 > [11663.276347] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 > [11663.279185] CR2: 00007f03955d60ac CR3: 00000001072a7006 CR4: 0000000000370ef0 > [11663.282698] Call Trace: > [11663.284061] <IRQ> > [11663.285116] enqueue_task_rt+0x77/0x3f0 > [11663.287276] enqueue_task+0x1cf/0x330 > [11663.289121] proxy_activate_blocked_task+0x498/0x640 > [11663.291579] ttwu_do_activate+0x287/0x290 > [11663.293582] sched_ttwu_pending+0x10b/0x220 > [11663.295726] __flush_smp_call_function_queue+0x1f8/0x680 > [11663.298320] __sysvec_call_function_single+0x35/0x180 > [11663.300877] sysvec_call_function_single+0x8a/0xb0 > [11663.303287] </IRQ> > [11663.304376] <TASK> > [11663.305467] asm_sysvec_call_function_single+0x1a/0x20 > [11663.307994] RIP: 0010:pv_native_safe_halt+0xf/0x20 > > > The only change I made to your tree was passing the rf into > proxy_activate_blocked_task, to use rq_(un)lock() instead of > raw_spin_rq_(un)lock(). Thank you for testing and the report! Let me go find what I've broken. -- Thanks and Regards, Prateek
Hello John, On 9/16/2026 10:52 AM, John Stultz wrote: > On Tue, Aug 25, 2026 at 11:29 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote: >> >> I had promised John I would share an insane idea if I got it >> working(ish) so this RFC presents an alternate approach to handle >> enqueued donor wakeup vs owner's chain-wakeup race via ttwu_runnable() >> path and make the activate path as light as possible for tasks that >> don't need to do a chain wakeup. >> >> The series essentially breaks down John's large patch at >> https://lore.kernel.org/lkml/20260807035232.1881495-9-jstultz@google.com/ >> into smaller chunks while introducing the new approach and fixing a few >> snags encountered along the way (Patch 1 - Patch 4 are fixes that can be >> discussed without getting into the rest of the RFC). >> >> Disclaimer: Series in not bisectible in any way at the moment - related >> bits are introduced together at and intermediate builds may fail. >> >> Introduction >> ============ >> >> Handling sleeping owner requires proxy-chains to queue on the said owner >> and perform a chain activation when the blocked owner wakes up. >> >> Since a blocked donor can be woken up by another concurrent wake event, >> the activation path becomes complicated and wakeup has to take an extra >> lock (p->blocked_lock) to prevent any modifications to "p->blocked_head" >> and miss activating any blocked donors. >> > ... >> Trade-off >> ========= >> >> Advantages: >> >> o Only need to juggle blocked_lock(s) under a single rq_lock(). >> o Re-use ttwu_runnable bit to handle removal of proxy donor. >> o No need to do get_task_stuct() / put_task_struct() juggling. >> >> Disadvantages: >> >> o Possible increased rq_lock contention on the chain-wakeup path but >> those events are generally rare. >> o Lack of delayed task handling since the src_rq need to be locked >> separately to migrate it. >> >> For the delayed handling, it is possible to use proxy_migrate_task() to >> move the task after blocking. Series does not implement this yet to >> limit the number of bad ideas. >> > > Sorry again for being so slow to respond here! > > The series definitely looks interesting, and it seems to be holding up > ok in testing. Though figuring out how to refactor them so that they > are also bisectable looks like a real challenge (part of why my > sleeping-owner enqueuing is basically one big patch)! > > That said, I can't say I've truly gotten my head fully around your > series yet. I've definitely had way more time with my change, so I'm a > little biased in feeling its somewhat more bounded (even though the > activate_blocked_waiters() function and the multiple lists of tasks > are very complicated). > > My initial sense of the downsides here with your series are: It adds a > lot of new per-task state (blocked_cpu, is_linked/needs_rq_sync, > sched_migrated_on_blocking, lock_nesting) to keep track of, and some > of rules for the new state have dependencies > (sched_migrated_on_blocking is tied to is_linked), and leveraging the > ENQUEUE/DEQUEUE_MIGRATING flags feels a little subtle (I have often > gotten the ENQUEUE/DEQUEUE flags wrong in my patch series, and > unfortunately the general documentation around those flags is lagging > a bit, so this adds to it). > > The locking being simpler is a clear benefit and definitely sound > appealing - though I do agree the rq_lock hold time in > proxy_activate_blocked_task() seems like it might be an issue. Ack! My original thinking was it is only done in the slow-path and is only held for a few ->on_rq and list manipulations so I might be able to get away with that small overhead. > > I've got a few more questions on specific patches, so I'll reply there. Looking forward to those. > Again, it definitely is interesting and appears to be more integrated > into the scheduler logic, so I'd expect that will give us better > results then my maybe more isolated and tacked on activation logic. > > Have you gotten a sense of what Peter thinks of it? I think Peter is slowly getting through the more stable stuff first and may arrive at this at some point but I'll try to get a word in at LPC if he is planning on attending. That said, knowing Peter, I have a hunch he might like your approach better since it handles delayed tasks too (tasks may be eligible at the time of chain-wakeup), does not add stuff into ttwu_runnable(), and does not have the insane (DEQUEUE_SLEEP | DEQUEUE_MIGRATING) behavior which has larger implications for PELT tracking and SCHED_DEADLINE. I sent out the RFC since it makes for an interesting discussion but I have a feeling lot more things need ironing out if w go this way but I can always do it later after the stuff from you and Suleiman lands once I can prove some benefit. -- Thanks and Regards, Prateek
On Tue, Sep 15, 2026 at 11:10 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote: > On 9/16/2026 10:52 AM, John Stultz wrote: > > On Tue, Aug 25, 2026 at 11:29 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote: > > Again, it definitely is interesting and appears to be more integrated > > into the scheduler logic, so I'd expect that will give us better > > results then my maybe more isolated and tacked on activation logic. > > > > Have you gotten a sense of what Peter thinks of it? > > I think Peter is slowly getting through the more stable stuff first > and may arrive at this at some point but I'll try to get a word in > at LPC if he is planning on attending. > > That said, knowing Peter, I have a hunch he might like your approach > better since it handles delayed tasks too (tasks may be eligible at the > time of chain-wakeup), does not add stuff into ttwu_runnable(), and does > not have the insane (DEQUEUE_SLEEP | DEQUEUE_MIGRATING) behavior which > has larger implications for PELT tracking and SCHED_DEADLINE. > > I sent out the RFC since it makes for an interesting discussion but I > have a feeling lot more things need ironing out if w go this way but > I can always do it later after the stuff from you and Suleiman lands > once I can prove some benefit. > Ok. I'd like to try to make more progress moving the sleeping owner enqueuing upstream as soon as possible, so deciding to push forward with my big patch or migrate to focusing on your set would be a good call to make quickly here. Your lock nesting optimization at the end seems like it might still be applicable with my patch, no? Maybe we can pull that in at least? I'm also wondering if trying to refactor the is_linked bits in after my change might make sense, though I suspect that will require the _MIGRATING flag dependencies? thanks -john
Hello John,
On 9/16/2026 11:53 AM, John Stultz wrote:
> On Tue, Sep 15, 2026 at 11:10 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
>> On 9/16/2026 10:52 AM, John Stultz wrote:
>>> On Tue, Aug 25, 2026 at 11:29 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
>>> Again, it definitely is interesting and appears to be more integrated
>>> into the scheduler logic, so I'd expect that will give us better
>>> results then my maybe more isolated and tacked on activation logic.
>>>
>>> Have you gotten a sense of what Peter thinks of it?
>>
>> I think Peter is slowly getting through the more stable stuff first
>> and may arrive at this at some point but I'll try to get a word in
>> at LPC if he is planning on attending.
>>
>> That said, knowing Peter, I have a hunch he might like your approach
>> better since it handles delayed tasks too (tasks may be eligible at the
>> time of chain-wakeup), does not add stuff into ttwu_runnable(), and does
>> not have the insane (DEQUEUE_SLEEP | DEQUEUE_MIGRATING) behavior which
>> has larger implications for PELT tracking and SCHED_DEADLINE.
>>
>> I sent out the RFC since it makes for an interesting discussion but I
>> have a feeling lot more things need ironing out if w go this way but
>> I can always do it later after the stuff from you and Suleiman lands
>> once I can prove some benefit.
>>
>
> Ok. I'd like to try to make more progress moving the sleeping owner
> enqueuing upstream as soon as possible, so deciding to push forward
> with my big patch or migrate to focusing on your set would be a good
> call to make quickly here.
>
> Your lock nesting optimization at the end seems like it might still be
> applicable with my patch, no? Maybe we can pull that in at least?
Lock nesting optimizations only work if we fully block delayed tasks
at the very least.
The only optimization I think which might work readily with your
series is proxy_enqueue_on_owner() bits in Patch 6 with:
activate_task(rq, p, flags)
{
if (sched_proxy_exec()) {
WRITE_ONCE(p->on_rq, TASK_ON_RQ_MIGRATING);
/*
* Once p->on_rq != 0, donors cannot queue on this task.
* pairs with smp_mb() in proxy_enqueue_on_owner() that
* orders list addition against owner's ->on_rq check.
*/
smp_mb();
}
if (list_empty(&p->blocked_head))
__activate_task(rq, p, flags);
guard(raw_spinlock)(p->blcoked_lock);
/* Slow-path */
}
Trades off needing to acquire blocked_lock always with an
ordered store instead. Also depends on the first 3 fixes from this
series to handle p->on_rq = MIGRATING properly on the
ttwu_do_activate() path.
>
> I'm also wondering if trying to refactor the is_linked bits in after
> my change might make sense, though I suspect that will require the
> _MIGRATING flag dependencies?
I think so but I can take a stab at it and send out next version
re-based just before / after proxy futex bits.
--
Thanks and Regards,
Prateek
© 2016 - 2026 Red Hat, Inc.