[PATCH v2] sched_ext: don't deliver duplicate ops.cgroup_set_idle() for same value

Tao Cui posted 1 patch 3 weeks, 3 days ago
kernel/sched/ext/ext.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
[PATCH v2] sched_ext: don't deliver duplicate ops.cgroup_set_idle() for same value
Posted by Tao Cui 3 weeks, 3 days ago
From: Tao Cui <cuitao@kylinos.cn>

ops.cgroup_set_idle() is documented to be invoked when a cgroup
transitions between idle and non-idle states, and scx_group_set_weight()
already skips value-preserving writes. scx_group_set_idle() delivers
every write unconditionally, so rewriting an already-correct cpu.idle
value feeds the BPF scheduler a transition callback each time, which
toggle- or accounting-based schedulers miscount. Mirror the weight
guard and only deliver on an actual change.

Verified with a probe scheduler printing each callback: rewriting
cpu.idle=1 twice on an already-idle cgroup delivered two callbacks
before and none after.

Fixes: 347ed2d566da ("sched/ext: Implement cgroup_set_idle() callback")
Link: https://lore.kernel.org/r/b53c61a1-4d7d-4232-941f-d48b0563d4ed@linux.dev
Signed-off-by: Tao Cui <cuitao@kylinos.cn>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
---
v1 -> v2: Fix the Link: msgid (missing @linux.dev, Andrea).

 kernel/sched/ext/ext.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c
index 8041c87a3562..8b3625107b72 100644
--- a/kernel/sched/ext/ext.c
+++ b/kernel/sched/ext/ext.c
@@ -4933,7 +4933,8 @@ void scx_group_set_idle(struct task_group *tg, bool idle)
 	percpu_down_read(&scx_cgroup_ops_rwsem);
 	sch = scx_tg_knob_sched(tg);
 
-	if (scx_cgroup_enabled && sch && SCX_HAS_OP(sch, cgroup_set_idle))
+	if (scx_cgroup_enabled && sch && SCX_HAS_OP(sch, cgroup_set_idle) &&
+	    tg->scx.sched_idle != idle)
 		SCX_CALL_OP(sch, cgroup_set_idle, NULL, tg_cgrp(tg), idle);
 
 	/* Update the task group's idle state */
-- 
2.43.0
Re: [PATCH v2] sched_ext: don't deliver duplicate ops.cgroup_set_idle() for same value
Posted by Tejun Heo 3 weeks, 2 days ago
> ops.cgroup_set_idle() is documented to be invoked when a cgroup
> transitions between idle and non-idle states, and scx_group_set_weight()
> already skips value-preserving writes. scx_group_set_idle() delivers
> every write unconditionally, so rewriting an already-correct cpu.idle
> value feeds the BPF scheduler a transition callback each time, which
> toggle- or accounting-based schedulers miscount. Mirror the weight
> guard and only deliver on an actual change.

Applied to sched_ext/for-7.3-fixes with the subject capitalized and the
new comparison changed to tg->scx.idle, which is the field's name there:

-	    tg->scx.sched_idle != idle)
+	    tg->scx.idle != idle)

The rename to tg->scx.sched_idle is on for-7.4 and the for-next merge
switches the comparison back.

Thanks.

--
tejun