[RFC PATCH 00/16][PoC] sched/core: Alternate approach to sleeping-owner handling in PROXY_EXEC

K Prateek Nayak posted 16 patches 1 month ago
include/linux/sched.h   |  31 ++-
init/init_task.c        |   5 +
kernel/fork.c           |   5 +
kernel/locking/mutex.c  |  57 +++-
kernel/sched/core.c     | 580 ++++++++++++++++++++++++++++++++++++++--
kernel/sched/deadline.c |  15 +-
kernel/sched/fair.c     |  16 +-
kernel/sched/sched.h    |   6 +
kernel/sched/stats.h    |   2 +-
9 files changed, 666 insertions(+), 51 deletions(-)
[RFC PATCH 00/16][PoC] sched/core: Alternate approach to sleeping-owner handling in PROXY_EXEC
Posted by K Prateek Nayak 1 month ago
Hello folks,

I had promised John I would share an insane idea if I got it
working(ish) so this RFC presents an alternate approach to handle
enqueued donor wakeup vs owner's chain-wakeup race via ttwu_runnable()
path and make the activate path as light as possible for tasks that
don't need to do a chain wakeup.

The series essentially breaks down John's large patch at
https://lore.kernel.org/lkml/20260807035232.1881495-9-jstultz@google.com/
into smaller chunks while introducing the new approach and fixing a few
snags encountered along the way (Patch 1 - Patch 4 are fixes that can be
discussed without getting into the rest of the RFC).

Disclaimer: Series in not bisectible in any way at the moment - related
bits are introduced together at and intermediate builds may fail.

Introduction
============

Handling sleeping owner requires proxy-chains to queue on the said owner
and perform a chain activation when the blocked owner wakes up.

Since a blocked donor can be woken up by another concurrent wake event,
the activation path becomes complicated and wakeup has to take an extra
lock (p->blocked_lock) to prevent any modifications to "p->blocked_head"
and miss activating any blocked donors.

The nesting rules for the locks are as follows:

  p->pi_lock
    __task_rq_lock()
      lock->wait_lock
        p->blocked_lock

Current chain-wakeup scheme juggles p->pi_lock, and p->blocked_lock to
prevent concurrent wake events from modifying the chain.

This RFC introduces a __task_rq_lock() based scheme that requires just
the __task_rq_lock() (always held) + p->blocked_lock (juggled for each
owner in chain) when queued tasks are detected to minimize the lock
bouncing and introduces a no-extra-locking fast-path for activation.

Trade offs are discussed at the end.

Approach
========

To implement this, we try a bunch of bad ideas:

o Waking up blocked donors linked to sleeping owner via ttwu_runnable()

To prevent chain-wakeup transitioning "p->on_rq" to QUEUED while a
concurrent wakeup is in progress, "p->on_rq" is unionized with a new
state "p->is_linked" as:

    union {
        struct {
            u8    on_rq;	/* Task on CPU */
            u8    is_linked;	/* Task on sleeping owner */
        };
        u16       needs_rq_sync;
    }

Instead of checking READ_ONCE(p->on_rq), to send task down the
ttwu_runnable() route, the new "needs_rq_sync" indicator is used which
can be atomically inspected outside rq_lock.

In ttwu_runnable() a fast-path first tries to remove the blocked donor
from the sleeping owner by just grabbing the "owner->blocked_lock" and
forces a full __task_rq_lock() if the blocked donor's state has changed
while trying to grab the blocked lock.

The chain-wakeup path takes a single __task_rq_lock() to prevent any
race from concurrent wakeups (described later).


o Proactively removing the blocked donor from owner if it wakes up

Proactively remove the blocked donor from sleeping owner in
proxy_enqueue_on_owner() if owner->on_rq is transitioned to != 0. This
ensures owner need not care about any additional donors that come in
after it transitions p->on_rq.

proxy_enqueue_on_owner() holds lock->wait_lock + owner->blocked_lock.
Since owner must grab the wait_lock again to wake up the blocked donor,
it is guaranteed that owner cannot disappear and get_task_struct() /
put_task_struct() dance is avoided.


o Migrating enqueued donors while blocking (SLEEP | MIGRATING)

To ensure a single __task_rq_lock() can handle the entire chain, a
blocked donor is migrated to a deterministic CPU  before it is
transitioned to p->on_rq = 0. This ensures the whole chain is owned by a
single CPU.


o No delayed tasks on chain

Delayed tasks pose a challenge that they are still on_rq and require
special handling to ensure the above condition of chain resolving to a
single CPU holds true. For simplicity, all delayed tasks on the chain
are fully blocked before they are queued.

This can be handled differently but I got a bit busy lately so that will
be handled in RFC v2 (if folks think this approach holds water).


o Single shot chain-wakeup (ab)using TASK_ON_RQ_MIGRATING

Tasks on chain are transitioned from being blocked to
TASK_ON_RQ_MIGRATING under __task_rq_lock(). Tasks on ttwu_runnable()
path that are trying to grab __task_rq_lock() is stalled until the
task is woken up on destination rq and the rq_lock() there is dropped.


o Tracking task that can potentially resolve to __mutex_owner()

This indicator is used to selectively set the lock holding
tasks to the slow-path.


Interleaving
============

How are different races handled (individual patches have these
highlighted too, either in comment message, or via code comments):

o Owner wakeup vs donor enqueue

  activate_task(owner)
    owner->on_rq = MIGRATING			proxy_enqueue_on_owner()
						  guard(raw_spinlock)(&owner->blocked_lock)
						  list_add(donor, owner->blocked_head)

    smp_mb() /* MIGRATING is visible */		  smp_mb() /* List addition is visible */

    if (list_empty(&owner->blocked_head))	  if (!READ_ONCE(owner->on_rq))
      return __activate_task(owner)		    block_task(donor)

    ... /* Do chain-activation */		  __proxy_dequeue_from_owner(donor)


o Chain-wakeup vs donor wakeup

  proxy_activate_donor_task()			try_to_wake_up(donor)

    rq_lock(rq_of(onwner->blocked_cpu))		  proxy_try_dequeue_from_owner(p)
      guard(raw_spinlock)(&owner->blocked_lock)	    guard(raw_spinlock)(&owner->blocked_lock)

      for_each(donor, owner->blocked_head)
        donor->on_rq = MIGRATING;
        smp_mb();				    if (donor->on_rq)
						      goto slowpath;
        ... /* More stuff */
    rq_unlock(rq_of(onwner->blocked_cpu));	    __task_rq_lock(donor)
						      if (donor->on_rq == MIGRATING)
						        cpu_relax();

    __proxy_dequeue_from_owner(donor)
    __activate_task(donor)
      p->on_rq = QUEUED;			    /* Acquired */
						    if (p->is_linked)
						      proxy_dequeue_from_owner(donor)

						    /* Do wakeup. */

Chain wakeup is covered extensively in Patch 14.


Series breakdown
================

Patch 01-04: Fixes to existing bits (Patch 4 needs more thinking)
Patch 05-07: Donor queuing with some convenience added to skip delayed
             handling later on.
Patch 08-13: Putting tasks on ttwu_running() path and swapping
             task_cpu() while blocking
Patch 14   : Chain wake-up
Patch 15-16: Optimizations to limit slowpath.


Trade-off
=========

Advantages:

o Only need to juggle blocked_lock(s) under a single rq_lock().
o Re-use ttwu_runnable bit to handle removal of proxy donor.
o No need to do get_task_stuct() / put_task_struct() juggling.

Disadvantages:

o Possible increased rq_lock contention on the chain-wakeup path but
  those events are generally rare.
o Lack of delayed task handling since the src_rq need to be locked
  separately to migrate it.

For the delayed handling, it is possible to use proxy_migrate_task() to
move the task after blocking. Series does not implement this yet to
limit the number of bad ideas.


How to apply
============

This series is based on John's tree at 

  https://github.com/johnstultz-work/linux-dev.git proxy-exec-v31-7.2-rc4

at commit 18bea15ce9fa ("sched: Migrate whole chain in
proxy_migrate_task()"). For convenience, a tree is present at:

  https://github.com/kudureranganath/linux.git  sched/proxy/sleeping_owner_rfc_v1

Sorry in advance to folks whose VPNs hate dealing with Github but I'm
told pulling it locally and inspecting the patches on tree works better.


Testing
=======

Tested with perf bench sched messaging, test-ww_mutex, and a kernel
module that implements a duty-cycle with sleep_ms() + busy-loop in
kernel.

I made sure to build and test !SCHED_PROXY_EXEC version this time around
to save John some trouble :-)


Acknowledgement
===============

Most of the good work is by John and all the bad ideas are by me.
Sorry in advance.

---
K Prateek Nayak (16):
  sched/core: Break activation of blocked task into a separate helper
  sched/core: Use enqueue/dequeue flags instead of
    task_on_rq_migrating()
  sched/fair: Use enqueue flags for DO_ATTACH in update_load_avg()
  sched/core: Activate blocked donor when no owner is found
  sched/core: Do not queue blocked donor on a delayed owner
  sched/core: Queue blocked donor onto sleeping owner for chain-wakeup
  sched/core: Avoid delaying blocked donors queued on sleeping owner
  sched/deadline: Prepare for blocking and proxy activation with
    MIGRATING flag
  sched/core: Track CPU where the task was blocked on
  sched/core: Introduce p->is_linked to track if task is queued on
    sleeping owner
  sched/core: Prepare to inspect ->is_linked alongside ->on_rq during
    wakeup
  sched:core: Add MIGRATING flags when blocking and activating linked
    donors
  sched/core: Use p->is_linked state to unlink from sleeping owner early
  sched/core: Introduce chain-wakeup to activate blocked donors
  locking/mutex: Track locks owned by a task in a per-task counter
  sched/core: Set activation of non lock-holders to fast-path

 include/linux/sched.h   |  31 ++-
 init/init_task.c        |   5 +
 kernel/fork.c           |   5 +
 kernel/locking/mutex.c  |  57 +++-
 kernel/sched/core.c     | 580 ++++++++++++++++++++++++++++++++++++++--
 kernel/sched/deadline.c |  15 +-
 kernel/sched/fair.c     |  16 +-
 kernel/sched/sched.h    |   6 +
 kernel/sched/stats.h    |   2 +-
 9 files changed, 666 insertions(+), 51 deletions(-)


base-commit: 18bea15ce9faaa0c1bd827012e6c60dbf35578c7
-- 
2.34.1
Re: [RFC PATCH 00/16][PoC] sched/core: Alternate approach to sleeping-owner handling in PROXY_EXEC
Posted by John Stultz 1 week, 4 days ago
On Tue, Aug 25, 2026 at 11:29 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
>
> I had promised John I would share an insane idea if I got it
> working(ish) so this RFC presents an alternate approach to handle
> enqueued donor wakeup vs owner's chain-wakeup race via ttwu_runnable()
> path and make the activate path as light as possible for tasks that
> don't need to do a chain wakeup.
>
> The series essentially breaks down John's large patch at
> https://lore.kernel.org/lkml/20260807035232.1881495-9-jstultz@google.com/
> into smaller chunks while introducing the new approach and fixing a few
> snags encountered along the way (Patch 1 - Patch 4 are fixes that can be
> discussed without getting into the rest of the RFC).
>
> Disclaimer: Series in not bisectible in any way at the moment - related
> bits are introduced together at and intermediate builds may fail.
>
> Introduction
> ============
>
> Handling sleeping owner requires proxy-chains to queue on the said owner
> and perform a chain activation when the blocked owner wakes up.
>
> Since a blocked donor can be woken up by another concurrent wake event,
> the activation path becomes complicated and wakeup has to take an extra
> lock (p->blocked_lock) to prevent any modifications to "p->blocked_head"
> and miss activating any blocked donors.
>
...
> Trade-off
> =========
>
> Advantages:
>
> o Only need to juggle blocked_lock(s) under a single rq_lock().
> o Re-use ttwu_runnable bit to handle removal of proxy donor.
> o No need to do get_task_stuct() / put_task_struct() juggling.
>
> Disadvantages:
>
> o Possible increased rq_lock contention on the chain-wakeup path but
>   those events are generally rare.
> o Lack of delayed task handling since the src_rq need to be locked
>   separately to migrate it.
>
> For the delayed handling, it is possible to use proxy_migrate_task() to
> move the task after blocking. Series does not implement this yet to
> limit the number of bad ideas.
>

Sorry again for being so slow to respond here!

The series definitely looks interesting, and it seems to be holding up
ok in testing.  Though figuring out how to refactor them so that they
are also bisectable looks like a real challenge (part of why my
sleeping-owner enqueuing is basically one big patch)!

That said, I can't say I've truly gotten my head fully around your
series yet. I've definitely had way more time with my change, so I'm a
little biased in feeling its somewhat more bounded (even though the
activate_blocked_waiters() function and the multiple lists of tasks
are very complicated).

My initial sense of the downsides here with your series are: It adds a
lot of new per-task state (blocked_cpu, is_linked/needs_rq_sync,
sched_migrated_on_blocking, lock_nesting) to keep track of, and some
of rules for the new state have dependencies
(sched_migrated_on_blocking is tied to is_linked), and leveraging the
ENQUEUE/DEQUEUE_MIGRATING flags feels a little subtle (I have often
gotten the ENQUEUE/DEQUEUE flags wrong in my patch series, and
unfortunately the general documentation around those flags is lagging
a bit, so this adds to it).

The locking being simpler is a clear benefit and definitely sound
appealing - though I do agree the rq_lock hold time in
proxy_activate_blocked_task() seems like it might be an issue.

I've got a few more questions on specific patches, so I'll reply there.

Again, it definitely is interesting and appears to be more integrated
into the scheduler logic, so I'd expect that will give us better
results then my maybe more isolated and tacked on activation logic.

Have you gotten a sense of what Peter thinks of it?

thanks
-john
Re: [RFC PATCH 00/16][PoC] sched/core: Alternate approach to sleeping-owner handling in PROXY_EXEC
Posted by John Stultz 1 week, 4 days ago
On Tue, Sep 15, 2026 at 10:22 PM John Stultz <jstultz@google.com> wrote:
> The series definitely looks interesting, and it seems to be holding up
> ok in testing.  Though figuring out how to refactor them so that they

Well, I did hit the following last night during stress testing:

[11176.672951] WARNING: kernel/sched/sched.h:1645 at
sched_change_begin+0x1b6/0x2b0, CPU#14: torture_shuffle/433
[11176.676650] CPU: 14 UID: 0 PID: 433 Comm: torture_shuffle Not
tainted 7.2.0-rc4-00032-gf14e36dcc837-dirty #138 PREEMPT(full)
[11176.680917] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996),
BIOS 1.17.0-debian-1.17.0-1 04/01/2014
[11176.684794] RIP: 0010:sched_change_begin+0x1b6/0x2b0
[11176.686748] Code: c9 1e 48 89 03 80 7b 1c 00 0f 85 ca 00 00 00 80
7b 1d 00 0f 85 91 00 00 00 48 89 d8 5b 5d 41 5c 41 5d 41 5e e9 1b 63
23 01 90 <0f> 0b 90 41 f6 c5 08 0f 85 c7 fe ff ff 48 89 ef 41 83 cd 08
e8 b1
[11176.693913] RSP: 0018:ffffc900010b3e10 EFLAGS: 00010046
[11176.695962] RAX: 0000000000000000 RBX: ffff8881b9198020 RCX: 0000000000000001
[11176.698747] RDX: 0000000000000001 RSI: ffffffff82c8b519 RDI: ffffffff82cad7fe
[11176.701554] RBP: ffff8881b91abdc0 R08: 0000000000000001 R09: 00000000d19ffe29
[11176.704561] R10: 000000000000000e R11: ffff88810264cff0 R12: ffff888106694400
[11176.707423] R13: 0000000000000002 R14: 0000000000000000 R15: ffffc90000017d58
[11176.710161] FS:  0000000000000000(0000) GS:ffff88823563e000(0000)
knlGS:0000000000000000
[11176.713196] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[11176.715628] CR2: 0000000000000000 CR3: 00000001072a7006 CR4: 0000000000370ef0
[11176.718462] Call Trace:
[11176.719460]  <TASK>
[11176.720319]  __set_cpus_allowed_ptr_locked+0x146/0x1f0
[11176.722343]  ? __pfx_torture_shuffle+0x10/0x10
[11176.724063]  __set_cpus_allowed_ptr+0x64/0xa0
[11176.725913]  set_cpus_allowed_ptr+0x3a/0x60
[11176.727513]  torture_shuffle+0x12b/0x230
[11176.729014]  kthread+0x103/0x130
[11176.730304]  ? __pfx_kthread+0x10/0x10
[11176.731758]  ret_from_fork+0x27c/0x340
[11176.733220]  ? __pfx_kthread+0x10/0x10
[11176.734680]  ret_from_fork_asm+0x1a/0x30
[11176.736366]  </TASK>

Followed shortly after by:
[11663.223681] kernel BUG at kernel/sched/rt.c:1020!
[11663.226093] Oops: invalid opcode: 0000 [#1] SMP NOPTI
[11663.228669] CPU: 46 UID: 0 PID: 0 Comm: swapper/46 Tainted: G
 W           7.2.0-rc4-00032-gf14e36dcc837-dirty #138 PREEMPT(full)
[11663.234796] Tainted: [W]=WARN
[11663.236392] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996),
BIOS 1.17.0-debian-1.17.0-1 04/01/2014
[11663.241196] RIP: 0010:dequeue_rt_stack+0x2ad/0x2c0
[11663.243652] Code: 83 a8 fe ff ff 39 c1 0f 8e 47 ff ff ff 90 0f 0b
90 e9 b7 fd ff ff 90 0f 0b 90 e9 e8 fd ff ff f3 48 0f bc d0 e9 80 fe
ff ff 90 <0f> 0b 48 8d 3d 1a 79 dc 01 67 48 0f b9 3a e9 e9 fe ff ff 90
90 90
[11663.253981] RSP: 0018:ffffc90000c10dd8 EFLAGS: 00010046
[11663.256630] RAX: 0000000000000000 RBX: 0000000000000000 RCX: ffff88810ad50080
[11663.260225] RDX: ffff888235e3e000 RSI: 0000000000000000 RDI: ffff88810ad501c0
[11663.263862] RBP: ffff8881b99abdc0 R08: 0000000000000000 R09: ffffffff83b9a280
[11663.266876] R10: 000000000000002e R11: 0000000000000000 R12: ffff8881b99abdc0
[11663.269626] R13: ffff88810ad501c0 R14: 0000000000000000 R15: 0000000000000025
[11663.272412] FS:  0000000000000000(0000) GS:ffff888235e3e000(0000)
knlGS:0000000000000000
[11663.276347] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[11663.279185] CR2: 00007f03955d60ac CR3: 00000001072a7006 CR4: 0000000000370ef0
[11663.282698] Call Trace:
[11663.284061]  <IRQ>
[11663.285116]  enqueue_task_rt+0x77/0x3f0
[11663.287276]  enqueue_task+0x1cf/0x330
[11663.289121]  proxy_activate_blocked_task+0x498/0x640
[11663.291579]  ttwu_do_activate+0x287/0x290
[11663.293582]  sched_ttwu_pending+0x10b/0x220
[11663.295726]  __flush_smp_call_function_queue+0x1f8/0x680
[11663.298320]  __sysvec_call_function_single+0x35/0x180
[11663.300877]  sysvec_call_function_single+0x8a/0xb0
[11663.303287]  </IRQ>
[11663.304376]  <TASK>
[11663.305467]  asm_sysvec_call_function_single+0x1a/0x20
[11663.307994] RIP: 0010:pv_native_safe_halt+0xf/0x20


The only change I made to your tree was passing the rf into
proxy_activate_blocked_task, to use rq_(un)lock() instead of
raw_spin_rq_(un)lock().

thanks
-john
Re: [RFC PATCH 00/16][PoC] sched/core: Alternate approach to sleeping-owner handling in PROXY_EXEC
Posted by K Prateek Nayak 1 week, 3 days ago
Hello John,

On 9/16/2026 11:27 PM, John Stultz wrote:
> On Tue, Sep 15, 2026 at 10:22 PM John Stultz <jstultz@google.com> wrote:
>> The series definitely looks interesting, and it seems to be holding up
>> ok in testing.  Though figuring out how to refactor them so that they
> 
> Well, I did hit the following last night during stress testing:
> 
> [11176.672951] WARNING: kernel/sched/sched.h:1645 at
> sched_change_begin+0x1b6/0x2b0, CPU#14: torture_shuffle/433
> [11176.676650] CPU: 14 UID: 0 PID: 433 Comm: torture_shuffle Not
> tainted 7.2.0-rc4-00032-gf14e36dcc837-dirty #138 PREEMPT(full)
> [11176.680917] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996),
> BIOS 1.17.0-debian-1.17.0-1 04/01/2014
> [11176.684794] RIP: 0010:sched_change_begin+0x1b6/0x2b0
> [11176.686748] Code: c9 1e 48 89 03 80 7b 1c 00 0f 85 ca 00 00 00 80
> 7b 1d 00 0f 85 91 00 00 00 48 89 d8 5b 5d 41 5c 41 5d 41 5e e9 1b 63
> 23 01 90 <0f> 0b 90 41 f6 c5 08 0f 85 c7 fe ff ff 48 89 ef 41 83 cd 08
> e8 b1
> [11176.693913] RSP: 0018:ffffc900010b3e10 EFLAGS: 00010046
> [11176.695962] RAX: 0000000000000000 RBX: ffff8881b9198020 RCX: 0000000000000001
> [11176.698747] RDX: 0000000000000001 RSI: ffffffff82c8b519 RDI: ffffffff82cad7fe
> [11176.701554] RBP: ffff8881b91abdc0 R08: 0000000000000001 R09: 00000000d19ffe29
> [11176.704561] R10: 000000000000000e R11: ffff88810264cff0 R12: ffff888106694400
> [11176.707423] R13: 0000000000000002 R14: 0000000000000000 R15: ffffc90000017d58
> [11176.710161] FS:  0000000000000000(0000) GS:ffff88823563e000(0000)
> knlGS:0000000000000000
> [11176.713196] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [11176.715628] CR2: 0000000000000000 CR3: 00000001072a7006 CR4: 0000000000370ef0
> [11176.718462] Call Trace:
> [11176.719460]  <TASK>
> [11176.720319]  __set_cpus_allowed_ptr_locked+0x146/0x1f0
> [11176.722343]  ? __pfx_torture_shuffle+0x10/0x10
> [11176.724063]  __set_cpus_allowed_ptr+0x64/0xa0
> [11176.725913]  set_cpus_allowed_ptr+0x3a/0x60
> [11176.727513]  torture_shuffle+0x12b/0x230
> [11176.729014]  kthread+0x103/0x130
> [11176.730304]  ? __pfx_kthread+0x10/0x10
> [11176.731758]  ret_from_fork+0x27c/0x340
> [11176.733220]  ? __pfx_kthread+0x10/0x10
> [11176.734680]  ret_from_fork_asm+0x1a/0x30
> [11176.736366]  </TASK>
> 
> Followed shortly after by:
> [11663.223681] kernel BUG at kernel/sched/rt.c:1020!
> [11663.226093] Oops: invalid opcode: 0000 [#1] SMP NOPTI
> [11663.228669] CPU: 46 UID: 0 PID: 0 Comm: swapper/46 Tainted: G
>  W           7.2.0-rc4-00032-gf14e36dcc837-dirty #138 PREEMPT(full)
> [11663.234796] Tainted: [W]=WARN
> [11663.236392] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996),
> BIOS 1.17.0-debian-1.17.0-1 04/01/2014
> [11663.241196] RIP: 0010:dequeue_rt_stack+0x2ad/0x2c0
> [11663.243652] Code: 83 a8 fe ff ff 39 c1 0f 8e 47 ff ff ff 90 0f 0b
> 90 e9 b7 fd ff ff 90 0f 0b 90 e9 e8 fd ff ff f3 48 0f bc d0 e9 80 fe
> ff ff 90 <0f> 0b 48 8d 3d 1a 79 dc 01 67 48 0f b9 3a e9 e9 fe ff ff 90
> 90 90
> [11663.253981] RSP: 0018:ffffc90000c10dd8 EFLAGS: 00010046
> [11663.256630] RAX: 0000000000000000 RBX: 0000000000000000 RCX: ffff88810ad50080
> [11663.260225] RDX: ffff888235e3e000 RSI: 0000000000000000 RDI: ffff88810ad501c0
> [11663.263862] RBP: ffff8881b99abdc0 R08: 0000000000000000 R09: ffffffff83b9a280
> [11663.266876] R10: 000000000000002e R11: 0000000000000000 R12: ffff8881b99abdc0
> [11663.269626] R13: ffff88810ad501c0 R14: 0000000000000000 R15: 0000000000000025
> [11663.272412] FS:  0000000000000000(0000) GS:ffff888235e3e000(0000)
> knlGS:0000000000000000
> [11663.276347] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [11663.279185] CR2: 00007f03955d60ac CR3: 00000001072a7006 CR4: 0000000000370ef0
> [11663.282698] Call Trace:
> [11663.284061]  <IRQ>
> [11663.285116]  enqueue_task_rt+0x77/0x3f0
> [11663.287276]  enqueue_task+0x1cf/0x330
> [11663.289121]  proxy_activate_blocked_task+0x498/0x640
> [11663.291579]  ttwu_do_activate+0x287/0x290
> [11663.293582]  sched_ttwu_pending+0x10b/0x220
> [11663.295726]  __flush_smp_call_function_queue+0x1f8/0x680
> [11663.298320]  __sysvec_call_function_single+0x35/0x180
> [11663.300877]  sysvec_call_function_single+0x8a/0xb0
> [11663.303287]  </IRQ>
> [11663.304376]  <TASK>
> [11663.305467]  asm_sysvec_call_function_single+0x1a/0x20
> [11663.307994] RIP: 0010:pv_native_safe_halt+0xf/0x20
> 
> 
> The only change I made to your tree was passing the rf into
> proxy_activate_blocked_task, to use rq_(un)lock() instead of
> raw_spin_rq_(un)lock().

Thank you for testing and the report! Let me go find what I've
broken.

-- 
Thanks and Regards,
Prateek

Re: [RFC PATCH 00/16][PoC] sched/core: Alternate approach to sleeping-owner handling in PROXY_EXEC
Posted by K Prateek Nayak 1 week, 4 days ago
Hello John,

On 9/16/2026 10:52 AM, John Stultz wrote:
> On Tue, Aug 25, 2026 at 11:29 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
>>
>> I had promised John I would share an insane idea if I got it
>> working(ish) so this RFC presents an alternate approach to handle
>> enqueued donor wakeup vs owner's chain-wakeup race via ttwu_runnable()
>> path and make the activate path as light as possible for tasks that
>> don't need to do a chain wakeup.
>>
>> The series essentially breaks down John's large patch at
>> https://lore.kernel.org/lkml/20260807035232.1881495-9-jstultz@google.com/
>> into smaller chunks while introducing the new approach and fixing a few
>> snags encountered along the way (Patch 1 - Patch 4 are fixes that can be
>> discussed without getting into the rest of the RFC).
>>
>> Disclaimer: Series in not bisectible in any way at the moment - related
>> bits are introduced together at and intermediate builds may fail.
>>
>> Introduction
>> ============
>>
>> Handling sleeping owner requires proxy-chains to queue on the said owner
>> and perform a chain activation when the blocked owner wakes up.
>>
>> Since a blocked donor can be woken up by another concurrent wake event,
>> the activation path becomes complicated and wakeup has to take an extra
>> lock (p->blocked_lock) to prevent any modifications to "p->blocked_head"
>> and miss activating any blocked donors.
>>
> ...
>> Trade-off
>> =========
>>
>> Advantages:
>>
>> o Only need to juggle blocked_lock(s) under a single rq_lock().
>> o Re-use ttwu_runnable bit to handle removal of proxy donor.
>> o No need to do get_task_stuct() / put_task_struct() juggling.
>>
>> Disadvantages:
>>
>> o Possible increased rq_lock contention on the chain-wakeup path but
>>   those events are generally rare.
>> o Lack of delayed task handling since the src_rq need to be locked
>>   separately to migrate it.
>>
>> For the delayed handling, it is possible to use proxy_migrate_task() to
>> move the task after blocking. Series does not implement this yet to
>> limit the number of bad ideas.
>>
> 
> Sorry again for being so slow to respond here!
> 
> The series definitely looks interesting, and it seems to be holding up
> ok in testing.  Though figuring out how to refactor them so that they
> are also bisectable looks like a real challenge (part of why my
> sleeping-owner enqueuing is basically one big patch)!
> 
> That said, I can't say I've truly gotten my head fully around your
> series yet. I've definitely had way more time with my change, so I'm a
> little biased in feeling its somewhat more bounded (even though the
> activate_blocked_waiters() function and the multiple lists of tasks
> are very complicated).
> 
> My initial sense of the downsides here with your series are: It adds a
> lot of new per-task state (blocked_cpu, is_linked/needs_rq_sync,
> sched_migrated_on_blocking, lock_nesting) to keep track of, and some
> of rules for the new state have dependencies
> (sched_migrated_on_blocking is tied to is_linked), and leveraging the
> ENQUEUE/DEQUEUE_MIGRATING flags feels a little subtle (I have often
> gotten the ENQUEUE/DEQUEUE flags wrong in my patch series, and
> unfortunately the general documentation around those flags is lagging
> a bit, so this adds to it).
> 
> The locking being simpler is a clear benefit and definitely sound
> appealing - though I do agree the rq_lock hold time in
> proxy_activate_blocked_task() seems like it might be an issue.

Ack! My original thinking was it is only done in the slow-path and
is only held for a few ->on_rq and list manipulations so I might be
able to get away with that small overhead.

> 
> I've got a few more questions on specific patches, so I'll reply there.

Looking forward to those.

> Again, it definitely is interesting and appears to be more integrated
> into the scheduler logic, so I'd expect that will give us better
> results then my maybe more isolated and tacked on activation logic.
> 
> Have you gotten a sense of what Peter thinks of it?

I think Peter is slowly getting through the more stable stuff first
and may arrive at this at some point but I'll try to get a word in
at LPC if he is planning on attending.

That said, knowing Peter, I have a hunch he might like your approach
better since it handles delayed tasks too (tasks may be eligible at the
time of chain-wakeup), does not add stuff into ttwu_runnable(), and does
not have the insane (DEQUEUE_SLEEP | DEQUEUE_MIGRATING) behavior which
has larger implications for PELT tracking and SCHED_DEADLINE.

I sent out the RFC since it makes for an interesting discussion but I
have a feeling lot more things need ironing out if w go this way but
I can always do it later after the stuff from you and Suleiman lands
once I can prove some benefit.

-- 
Thanks and Regards,
Prateek

Re: [RFC PATCH 00/16][PoC] sched/core: Alternate approach to sleeping-owner handling in PROXY_EXEC
Posted by John Stultz 1 week, 4 days ago
On Tue, Sep 15, 2026 at 11:10 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
> On 9/16/2026 10:52 AM, John Stultz wrote:
> > On Tue, Aug 25, 2026 at 11:29 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
> > Again, it definitely is interesting and appears to be more integrated
> > into the scheduler logic, so I'd expect that will give us better
> > results then my maybe more isolated and tacked on activation logic.
> >
> > Have you gotten a sense of what Peter thinks of it?
>
> I think Peter is slowly getting through the more stable stuff first
> and may arrive at this at some point but I'll try to get a word in
> at LPC if he is planning on attending.
>
> That said, knowing Peter, I have a hunch he might like your approach
> better since it handles delayed tasks too (tasks may be eligible at the
> time of chain-wakeup), does not add stuff into ttwu_runnable(), and does
> not have the insane (DEQUEUE_SLEEP | DEQUEUE_MIGRATING) behavior which
> has larger implications for PELT tracking and SCHED_DEADLINE.
>
> I sent out the RFC since it makes for an interesting discussion but I
> have a feeling lot more things need ironing out if w go this way but
> I can always do it later after the stuff from you and Suleiman lands
> once I can prove some benefit.
>

Ok. I'd like to try to make more progress moving the sleeping owner
enqueuing upstream as soon as possible, so deciding to push forward
with my big patch or migrate to focusing on your set would be a good
call to make quickly here.

Your lock nesting optimization at the end seems like it might still be
applicable with my patch, no? Maybe we can pull that in at least?

I'm also wondering if trying to refactor the is_linked bits in after
my change might make sense, though I suspect that will require the
_MIGRATING flag dependencies?

thanks
-john
Re: [RFC PATCH 00/16][PoC] sched/core: Alternate approach to sleeping-owner handling in PROXY_EXEC
Posted by K Prateek Nayak 1 week, 4 days ago
Hello John,

On 9/16/2026 11:53 AM, John Stultz wrote:
> On Tue, Sep 15, 2026 at 11:10 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
>> On 9/16/2026 10:52 AM, John Stultz wrote:
>>> On Tue, Aug 25, 2026 at 11:29 PM K Prateek Nayak <kprateek.nayak@amd.com> wrote:
>>> Again, it definitely is interesting and appears to be more integrated
>>> into the scheduler logic, so I'd expect that will give us better
>>> results then my maybe more isolated and tacked on activation logic.
>>>
>>> Have you gotten a sense of what Peter thinks of it?
>>
>> I think Peter is slowly getting through the more stable stuff first
>> and may arrive at this at some point but I'll try to get a word in
>> at LPC if he is planning on attending.
>>
>> That said, knowing Peter, I have a hunch he might like your approach
>> better since it handles delayed tasks too (tasks may be eligible at the
>> time of chain-wakeup), does not add stuff into ttwu_runnable(), and does
>> not have the insane (DEQUEUE_SLEEP | DEQUEUE_MIGRATING) behavior which
>> has larger implications for PELT tracking and SCHED_DEADLINE.
>>
>> I sent out the RFC since it makes for an interesting discussion but I
>> have a feeling lot more things need ironing out if w go this way but
>> I can always do it later after the stuff from you and Suleiman lands
>> once I can prove some benefit.
>>
> 
> Ok. I'd like to try to make more progress moving the sleeping owner
> enqueuing upstream as soon as possible, so deciding to push forward
> with my big patch or migrate to focusing on your set would be a good
> call to make quickly here.
> 
> Your lock nesting optimization at the end seems like it might still be
> applicable with my patch, no? Maybe we can pull that in at least?

Lock nesting optimizations only work if we fully block delayed tasks
at the very least.

The only optimization I think which might work readily with your
series is proxy_enqueue_on_owner() bits in Patch 6 with:

activate_task(rq, p, flags)
{

    if (sched_proxy_exec()) {
        WRITE_ONCE(p->on_rq, TASK_ON_RQ_MIGRATING);

        /*
         * Once p->on_rq != 0, donors cannot queue on this task.
         * pairs with smp_mb() in proxy_enqueue_on_owner() that
         * orders list addition against owner's ->on_rq check.
         */
        smp_mb();
    }

    if (list_empty(&p->blocked_head))
        __activate_task(rq, p, flags);

    guard(raw_spinlock)(p->blcoked_lock);

    /* Slow-path */
}

Trades off needing to acquire blocked_lock always with an
ordered store instead. Also depends on the first 3 fixes from this
series to handle p->on_rq = MIGRATING properly on the
ttwu_do_activate() path.

> 
> I'm also wondering if trying to refactor the is_linked bits in after
> my change might make sense, though I suspect that will require the
> _MIGRATING flag dependencies?

I think so but I can take a stab at it and send out next version
re-based just before / after proxy futex bits.

-- 
Thanks and Regards,
Prateek