[PATCH] sched_ext: don't deliver duplicate ops.cgroup_set_idle() for same value

Tao Cui posted 1 patch 3 weeks, 4 days ago
There is a newer version of this series
kernel/sched/ext/ext.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
[PATCH] sched_ext: don't deliver duplicate ops.cgroup_set_idle() for same value
Posted by Tao Cui 3 weeks, 4 days ago
From: Tao Cui <cuitao@kylinos.cn>

ops.cgroup_set_idle() is documented to be invoked when a cgroup
transitions between idle and non-idle states, and scx_group_set_weight()
already skips value-preserving writes. scx_group_set_idle() delivers
every write unconditionally, so rewriting an already-correct cpu.idle
value feeds the BPF scheduler a transition callback each time, which
toggle- or accounting-based schedulers miscount. Mirror the weight
guard and only deliver on an actual change.

Verified with a probe scheduler printing each callback: rewriting
cpu.idle=1 twice on an already-idle cgroup delivered two callbacks
before and none after.

Fixes: 347ed2d566da ("sched/ext: Implement cgroup_set_idle() callback")
Link: https://lore.kernel.org/r/b53c61a1-4d7d-4232-941f-d48b0563d4ed
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
 kernel/sched/ext/ext.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 8041c87a3562..8b3625107b72 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -4933,7 +4933,8 @@ void scx_group_set_idle(struct task_group *tg, bool idle)
 	percpu_down_read(&scx_cgroup_ops_rwsem);
 	sch = scx_tg_knob_sched(tg);
 
-	if (scx_cgroup_enabled && sch && SCX_HAS_OP(sch, cgroup_set_idle))
+	if (scx_cgroup_enabled && sch && SCX_HAS_OP(sch, cgroup_set_idle) &&
+	    tg->scx.sched_idle != idle)
 		SCX_CALL_OP(sch, cgroup_set_idle, NULL, tg_cgrp(tg), idle);
 
 	/* Update the task group's idle state */
-- 
2.43.0
Re: [PATCH] sched_ext: don't deliver duplicate ops.cgroup_set_idle() for same value
Posted by Andrea Righi 3 weeks, 4 days ago
Hi Tao,

On Tue, Sep 01, 2026 at 11:11:01AM +0800, Tao Cui wrote:
> From: Tao Cui <cuitao@kylinos.cn>
> 
> ops.cgroup_set_idle() is documented to be invoked when a cgroup
> transitions between idle and non-idle states, and scx_group_set_weight()
> already skips value-preserving writes. scx_group_set_idle() delivers
> every write unconditionally, so rewriting an already-correct cpu.idle
> value feeds the BPF scheduler a transition callback each time, which
> toggle- or accounting-based schedulers miscount. Mirror the weight
> guard and only deliver on an actual change.
> 
> Verified with a probe scheduler printing each callback: rewriting
> cpu.idle=1 twice on an already-idle cgroup delivered two callbacks
> before and none after.
> 
> Fixes: 347ed2d566da ("sched/ext: Implement cgroup_set_idle() callback")
> Link: https://lore.kernel.org/r/b53c61a1-4d7d-4232-941f-d48b0563d4ed

This link seems broken, I think the right one is:

Link: https://lore.kernel.org/r/b53c61a1-4d7d-4232-941f-d48b0563d4ed@linux.dev

> Signed-off-by: Tao Cui <cuitao@kylinos.cn>

Other than that looks good to me.

Reviewed-by: Andrea Righi <arighi@nvidia.com>

Thanks,
-Andrea

> ---
>  kernel/sched/ext/ext.c | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
> 
> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
> index 8041c87a3562..8b3625107b72 100644
> --- a/kernel/sched/ext/ext.c
> +++ b/kernel/sched/ext/ext.c
> @@ -4933,7 +4933,8 @@ void scx_group_set_idle(struct task_group *tg, bool idle)
>  	percpu_down_read(&scx_cgroup_ops_rwsem);
>  	sch = scx_tg_knob_sched(tg);
>  
> -	if (scx_cgroup_enabled && sch && SCX_HAS_OP(sch, cgroup_set_idle))
> +	if (scx_cgroup_enabled && sch && SCX_HAS_OP(sch, cgroup_set_idle) &&
> +	    tg->scx.sched_idle != idle)
>  		SCX_CALL_OP(sch, cgroup_set_idle, NULL, tg_cgrp(tg), idle);
>  
>  	/* Update the task group's idle state */
> -- 
> 2.43.0
>
Re: [PATCH] sched_ext: don't deliver duplicate ops.cgroup_set_idle() for same value
Posted by Tao Cui 3 weeks, 4 days ago
Hi Andrea,

在 2026/9/1 15:25, Andrea Righi 写道:
> Hi Tao,
> 
> On Tue, Sep 01, 2026 at 11:11:01AM +0800, Tao Cui wrote:
>> From: Tao Cui <cuitao@kylinos.cn>
>>
>> ops.cgroup_set_idle() is documented to be invoked when a cgroup
>> transitions between idle and non-idle states, and scx_group_set_weight()
>> already skips value-preserving writes. scx_group_set_idle() delivers
>> every write unconditionally, so rewriting an already-correct cpu.idle
>> value feeds the BPF scheduler a transition callback each time, which
>> toggle- or accounting-based schedulers miscount. Mirror the weight
>> guard and only deliver on an actual change.
>>
>> Verified with a probe scheduler printing each callback: rewriting
>> cpu.idle=1 twice on an already-idle cgroup delivered two callbacks
>> before and none after.
>>
>> Fixes: 347ed2d566da ("sched/ext: Implement cgroup_set_idle() callback")
>> Link: https://lore.kernel.org/r/b53c61a1-4d7d-4232-941f-d48b0563d4ed
> 
> This link seems broken, I think the right one is:
> 
> Link: https://lore.kernel.org/r/b53c61a1-4d7d-4232-941f-d48b0563d4ed@linux.dev
> 

Thanks! Somehow my vim seems to have eaten the `@linux.dev` part of the Message-ID. I'll fix the Link tag in the next revision.

Thanks for the review!

Best,
Tao

>> Signed-off-by: Tao Cui <cuitao@kylinos.cn>
> 
> Other than that looks good to me.
> 
> Reviewed-by: Andrea Righi <arighi@nvidia.com>
> 
> Thanks,
> -Andrea
> 
>> ---
>>  kernel/sched/ext/ext.c | 3 ++-
>>  1 file changed, 2 insertions(+), 1 deletion(-)
>>
>> diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
>> index 8041c87a3562..8b3625107b72 100644
>> --- a/kernel/sched/ext/ext.c
>> +++ b/kernel/sched/ext/ext.c
>> @@ -4933,7 +4933,8 @@ void scx_group_set_idle(struct task_group *tg, bool idle)
>>  	percpu_down_read(&scx_cgroup_ops_rwsem);
>>  	sch = scx_tg_knob_sched(tg);
>>  
>> -	if (scx_cgroup_enabled && sch && SCX_HAS_OP(sch, cgroup_set_idle))
>> +	if (scx_cgroup_enabled && sch && SCX_HAS_OP(sch, cgroup_set_idle) &&
>> +	    tg->scx.sched_idle != idle)
>>  		SCX_CALL_OP(sch, cgroup_set_idle, NULL, tg_cgrp(tg), idle);
>>  
>>  	/* Update the task group's idle state */
>> -- 
>> 2.43.0
>>