[PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff

Shrikanth Hegde posted 12 patches 1 month, 2 weeks ago
There is a newer version of this series
.../ABI/testing/sysfs-devices-system-cpu      |  14 +
Documentation/driver-api/index.rst            |   1 +
Documentation/driver-api/steal-governor.rst   | 137 +++++++++
Documentation/scheduler/index.rst             |   1 +
Documentation/scheduler/sched-paravirt.rst    |  67 ++++
MAINTAINERS                                   |   9 +
arch/s390/kernel/hiperdispatch.c              |   8 +-
drivers/base/cpu.c                            |  12 +
drivers/virt/Kconfig                          |  18 ++
drivers/virt/Makefile                         |   1 +
drivers/virt/steal_governor.c                 | 286 ++++++++++++++++++
fs/proc/uptime.c                              |   6 +-
include/linux/cpumask.h                       |  24 ++
include/linux/kernel_stat.h                   |  11 +
include/linux/sched.h                         |   1 +
kernel/Kconfig.preempt                        |   4 +
kernel/cpu.c                                  |   6 +
kernel/sched/core.c                           | 112 ++++++-
kernel/sched/debug.c                          |   1 +
kernel/sched/fair.c                           |   8 +-
kernel/sched/sched.h                          |   8 +
21 files changed, 717 insertions(+), 18 deletions(-)
create mode 100644 Documentation/driver-api/steal-governor.rst
create mode 100644 Documentation/scheduler/sched-paravirt.rst
create mode 100644 drivers/virt/steal_governor.c
[PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Shrikanth Hegde 1 month, 2 weeks ago
If you have already read v8,v9 cover-letter then see only revision
changes. everything else is pretty much same. :) 

This patch series represents the result of multiple iterations, 
redesigns and community feedback. What started as an arch-specific RFC
has evolved into a scheduler mechanism paired with a virtualization
driver.

Special thanks to Yury Norov for the rigorous reviews that greatly 
improved the series and to everyone who have provided their review
comments so far. Really appreciated! _/\_

I have put a detailed context around problem statement, design, best
practises and performance numbers below. This cover-letter is a good
starting point for anyone looking into this solution without the pain of
browsing through all the previous patches/videos.

Apologies in advance if any review comments are missed or missed any
implementation for the new driver. If so would be purely
accidental, not in any way intentional.

Background and Problem Statement
================================

As hardware scales, the density of physical CPUs (pCPUs) per server is
increasing across many architectures. On these massive systems, deploying
a single bare-metal OS for general workloads becomes increasingly difficult
to manage if not impossible. The natural shift is to deploy
Virtual Machines(VMs). For example, on IBM PowerPC architecture customers
frequently deploy Shared Processor LPARs (SPLPARs) to maximize hardware ROI.

Typical enterprise workloads are combination of bursty and long running;
their average CPU utilization is low, but they require high core counts
during peak transactions. To accommodate this, customers often
use CPU overcommit strategies i.e. configuring VMs with a large number
of virtual CPUs (vCPUs) while backing them with a smaller, shared pool
of physical CPUs (pCPUs). This achieves a high server consolidation
and excellent cost efficiency.

However, when multiple such VMs have high utilization simultaneously,
the shared pCPU pool becomes contended. The hypervisor is forced to preempt
one vCPU to run another to maintain fairness. It maybe schedule vCPU of same
VM or different VM.  If a vCPU is preempted while holding a lock or
irq disabled section, overall forward progress collapses. There are some
mitigation strategies such as yielding the vCPU to lock-holder, but they
don't cover all the cases. In addition there are hidden costs such as cache,
tlb misses, cost of vCPU preemption, host scheduling overheads etc.

Under heavy contention, the most effective mitigation strategy is for
the guests/VMs to voluntarily fold its workload onto a smaller subset
of its vCPUs. By demanding fewer pCPUs, the VMs reduce overall host 
contention, which decreases vCPU preemption and improves total throughput
for the system. 

Limitations of Existing Approaches
==================================

CPU Hotplug, Isolated cpusets, cpuset: 
- This is a heavy and administrative operation that requires topology rebuild.
  Crucially, it breaks userspace CPU affinities. 

Explicit task affinity:
- Very difficult to manage for the users, if not impossible.

We need a fast, co-operative backoff mechanism inside the kernel that can
dynamically react to contention without violating user/task affinity
contracts. Since reacting to the contention is agnostic to the user
it cannot violate user affinity contracts.

When there is high contention, fold the workload and use limited vCPUs
and when there is no contention, use all the vCPUs again. This natural
expansion/contraction gives the best possible performance to the users
based on the underlying contention.

Proposed Architecture
=====================

Current design is built on basis that contention is effectively
quantified by steal time as seen in guest kernel.
Steal time is already a well established construct today in
para-virtualization world.  All major archs support this feature.
It is indication of the contention of physical CPU. It scales according
to the amount of contention. Today it is used by administrative users
for changing the VM configurations. During high contention the steal
time shows up in each guest based on its configuration. The proposed
solution works well when all VMs honor the hint and work in co-operative
manner. Note there is still no inter-guest communication to achieve this
co-operation. Read the section on best practises on how to get the
best out of this solution.

This series introduces a dynamic vCPU backoff mechanism.
It is separated into a core scheduler mechanism and a loadable
virtualization policy module.

Layer A: The Scheduler Mechanism (preferred CPUs)
=================================================

Series introduces a new CPU state called preferred. It indicates that
vCPU can be safely used and using that vCPU won't increase contention
for underlying physical CPUs. This state info is made available via
cpu_preferred_mask, which is strictly maintained as a subset of
cpu_active_mask.

The scheduler uses this mask as a hint to fold workloads onto preferred
CPUs using a few mechanisms.

1. Wakeup: is_cpu_allowed() checks if CPU is preferred. If not calls
   select_fallback_rq, which selects a preferred CPU if tasks's affinity
   permits.

2. The Tick (Push): During sched_tick(), if the current CPU is non-preferred,
   the scheduler actively pushes the running task onto a preferred CPU
   using a stopper thread. 

3. Load Balance: sched_balance_rq restricts its domain span to
   cpu_preferred_mask, preventing tasks from being pulled toward
   non-preferred CPUs.

Design Constraint: The scheduler strictly respects user affinities.
If a task is pinned exclusively to non-preferred CPUs, it will remain there.
The kernel will not break user/task affinity contracts.

Layer B: The Policy Engine (virt/steal_governor)
================================================

The core scheduler should not dictate virtualization policy.
Therefore, the policy is isolated into a new driver: steal_governor.
(Can be selected by CONFIG_STEAL_GOVERNOR)
This module latches onto that concept that contention is quantified by
steal time. It periodically samples the steal time values across the
system and depending on high/low steal values, takes appropriate action.

When it sees high steal times, i.e. steal time exceeds high_threshold
(default 5%), driver reduces the preferred CPUs by 1 core. 
When it sees Low Steal Times, i.e.  steal time drops below low_threshold
(default 2%), driver increases the preferred CPUs by 1 core.

This creates a dynamic, self-maintained stepwise loop. The guest automatically
shrinks its pCPU footprint when the host is saturated, and expands it when
the noise clears while requiring zero cross-VM communication.

Policy Design Constraints:
- Ensure at least one core is kept as preferred.
- Ensure preferred is always subset of active.

Best Practises
==============
1. Ensure all the VM run kernel which has the patches.

2. Keep CONFIG_STEAL_GOVERNOR=m. Build it as module, but don't load it by
   default. When the administrative user enables it in one VM, he/she
   will likely enable it in all VMs. Also module parameters can
   only be changed at module load. Having it as module also allows one
   to disable it to remove additional overhead it brings.

3. Keep the interval_ms=500 to 5000. I.e. between 500ms to 5 second.
   Though parameters allows slightly higher range. 

4. Fine tune low and high threshold depending on your platform for best
   results. Even where is no contention, very small steal values
   might show up. So it might be better to keep low threshold higher
   than 0.

Baseline and Revision History
==============================

tip/sched/core at commit:
'f2c2ba7219e5 ("sched/topology: Restore SD_PREFER_SIBLING in domains with asymmetric capacity")'

For a detailed talk on the problem and discussion on this issue, one can also
refer to the OSPM26 talk[1]. 

[1]: https://youtu.be/adxUKFPlOp0
[2]: https://www.ibm.com/support/pages/ibm-power-virtualization-best-practices-guide
[3]: https://www.ibm.com/docs/en/linux-on-systems?topic=bad-daytrader

v9->v10:
- Introduce kcpustat_field_total helper. (Yury Norov)
- Always do the design checks. This helps to avoid placing design
  constraints in core hotplug code. 
- Remove cpu_preferred check in idle balancing. This helps to naturally
  take care update of nohz.next_balance.
- find_new_ilb changes are deferred as it isn't applicable for most
  common use cases.
- Move scheduler documentation to sched-paravirt.rst. (Yury Norov)
- Add details of limitation of default values in documentation. (Yury Norov)
- Remove task_can_sched_on_preferred out of sched.h (Mete Durlu)
- Updated suggested-by tags for few patches. (I know i should have
  done it earlier, sorry about that)
- Minor polish of all changelogs.

v8->v9:
- Move to simpler layout. Everything in drivers/virt/steal_governor.c
  (Yury Norov)
- Move design checks into a helper function (Yury Norov)
- Comments update (Yury Norov)
- Requeue work without further checks when steal ratio is within the 
  low/high threshold window. (Yury Norov)
- Renamed sg_core_ctx to sg_ctx.
- refactoring like kcpustat_field_total for CPUTIME_STEAL will be
  picked up post the series.

v7->v8:
- Rename to STEAL_GOVERNOR from STEAL_MONITOR.
- Remove additional defaults.c and move it to core.c (Yury Norov)
- Remove SM_DIR gating for direction control. (Yury Norov)
- Enforce design constraint and restore the state if not met (Yury
  Norov)
- Drop nohz_full tick enable patch.
- Move Kconfig patch as the last patch for enablement. (Yury Norov)
- Use disable_delayed_work_sync to avoid race condition during
  module unload. (Yury Norov)
- Add same kconfig dependency and fail to compile the driver (Yury Norov)
- Make low < high comparison during module init instead as they
  are dependent parameters (Sashiko)
- Update sysfs file helper section (Yury Norov)
- Make preferred sysfs file available only with CONFIG_PREFERRED_CPU=y
  (Yury Norov)
- A few documentation and comments fixes. (Randy Dunlap)
- Fix possible race in sched_push_current_non_preferred_cpu (Yury Norov)
- Move is_migration_disabled check just before actual migration.
- Make 100ms as minimal interval_ms from 10ms.
- Make helper functions static and remove from header file as there
  are no other callers.
- Collapse helper functions and periodic work into one patch.

Short summary on previous versions:
v6->v7:
- Consolidate new driver code to 4-5 patches.
- defer the arch specific interface.
- Use possible CPUs instead of active for steal value calculations.
- Simplify is_cpu_allowed.
- Make module parameters fixed at module load
- Define CONFIG_STEAL_MONITOR and Make it select CONFIG_PREFERRED_CPU

v5->v6:
- Drop the optimization of caching the preferred state
  in select_fallback_rq
- Drop wakeup patch

v4->v5:
- Move the computation of steal time and decide on preferred CPU state
  to a driver. i.e new driver called STEAL_MONITOR

v3->v4:
- Make preferred subset of active instead of online. 
- Dropped RT patch and Defer sched_ext. Support only FAIR class.

v2->v3:
- Introduce a new config CONFIG_PREFERRED_CPU

v1->v2:
- A new name - Preferred CPUs and cpu_preferred_mask
- Arch independent code. Everything happens in scheduler.
- Steal time computation is gated with sched feature STEAL_MONITOR

RFC v3-> RFC v4:
- Introduced computation of steal time in arch/powerpc.

RFC PATCH v1:
- push task mechanism.
- No steal time computation. Manual sysfs hint for preferred CPUs 

v1: https://lore.kernel.org/all/236f4925-dd3c-41ef-be04-47708c9ce129@linux.ibm.com/
v2: https://lore.kernel.org/all/20260407191950.643549-1-sshegde@linux.ibm.com/#t
v3: https://lore.kernel.org/all/20260514152204.481115-1-sshegde@linux.ibm.com/#r
v4: https://lore.kernel.org/all/20260617174139.155540-1-sshegde@linux.ibm.com/#t
v5: https://lore.kernel.org/all/20260625124648.802832-1-sshegde@linux.ibm.com/
v6: https://lore.kernel.org/all/20260701141654.500125-1-sshegde@linux.ibm.com/#t
v7: https://lore.kernel.org/all/20260709215648.1246821-1-sshegde@linux.ibm.com/
v8: https://lore.kernel.org/all/20260720172250.2257582-1-sshegde@linux.ibm.com/
v9: https://lore.kernel.org/all/20260724140732.2683314-1-sshegde@linux.ibm.com/
Even earlier version:
https://lore.kernel.org/all/236f4925-dd3c-41ef-be04-47708c9ce129@linux.ibm.com/ 

========================================
Performance Numbers (powerpc, x86, s390)
========================================

PowerPC:
===================
VM1: 60VP/30EC and VM2: 30VP/20EC
Shared physical CPU pool size: 50 Cores. Each core is SMT8.
(VP - Virtual Core, EC - Entitles Core) -  PowerVM terminologies of SPLPAR[2]

Default parameter values: 1000ms, 200 low threshold, 500 high threshold
Both the VMs are running the same workload. Total throughput/time of VM1+VM2
is being mentioned in all cases.

Hackbench
              baseline    steal_governor        steal_governor
                             disabled               enabled
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

10 groups        5.20   |    5.40 (-3.85%)  |     4.65 (+10.58%)
20 groups       11.39   |   12.01 (-5.44%)  |     7.09 (+37.75%)
40 groups       20.32   |   19.80 (+2.56%)  |    11.31 (+44.34%)
10 groups(-p)    2.37   |    2.26 (+4.64%)  |     2.06 (+13.08%)
20 groups(-p)    3.34   |    3.28 (+1.80%)  |     3.20 (+4.19%)
40 groups(-p)    4.46   |    4.83 (-8.30%)  |     4.26 (+4.48%)
Remarks: Net improvement with steal_governor specially high load points.

schbench ( -L -n 0 -r 30 -s 0)
              baseline    steal_governor          steal_governor
                             disabled                 enabled
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
-m 1 -t 128     2475162 |    2621246 (+5.90%)  |      2527299 (+2.11%)
-m 1 -t 256     1467350 |    1470032 (+0.18%)  |      1492372 (+1.71%)
-m 1 -t 512     1408813 |    1454687 (+3.26%)  |      1437605 (+2.04%)
Remarks: Effectively means no-improvements or regressions

kernbench	baseline    steal_governor     steal_governor
(elapsed time)	               disabled            enabled
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
-j nr_cpus	231      |      235 (-1.7%) |    199 (+14%)
Remarks: Net improvement in elapsed time.

Daytrader - A real life work which is a proxy for trading based
on db2[3]
              baseline      steal_governor   steal_governor
                              disabled          enabled
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Load@30%	1x	|	0.96x	|	 1.53x			
Load@60%	1x	|	0.94x	|	 1.41x
Remarks: Good improvement seen at different load points.

When there is no steal time (such as dedicated LPAR, or only VM2
is running) throughput was same with steal_governor enabled/disabled
which indicates minimal overhead of steal_governor. 

I have run v10 also on a smaller powerpc LPAR system and it shows
good improvements.

=======================================================================

Data from x86,s390 KVM which Ilya Leoshkevich carried out during OSPM26
time. *This was based on v2*. Idea is still the name, numbers are
expected to be better in v10 as some of the overhead has been removed.
Note: Other variations of the benchmark shows no observable
difference.

x86:
====
cascade-lake: 32 threads = 16 cores
Benchmark      #VMs    #CPUs/VM  ΔRPS     (%std)
===============================================
hackbench         8          16  90.73% ± 9.97%
hackbench         4          24  52.67% ± 7.43%
hackbench         4          16  37.96% ± 11.19%
hackbench         4          32  37.82% ± 4.38%
hackbench        12           8  36.90% ± 4.74%
hackbench         8           8  35.30% ± 3.61%
pgbench          16           4  31.77% ± 2.44%
hackbench         2          24  25.85% ± 8.63%
hackbench        16           8  24.87% ± 3.46%
pgbench          16           8  21.83% ± 2.20%
pgbench          12           8  21.35% ± 2.15%
pgbench           8           8  18.46% ± 1.01%
hackbench         2          32  15.56% ± 4.53%
pgbench          12           4  14.28% ± 2.04%
hackbench        16           4  14.07% ± 2.90%
hackbench        12           4  9.60% ± 3.49%
[...]
pgbench           4           8  -1.16% ± 3.60%
hackbench         4           4  -1.80% ± 9.55%
sysbench         12           4  -2.19% ± 0.78%
pgbench           4          24  -2.43% ± 4.38%
pgbench           4          32  -3.21% ± 0.79%
sysbench         16           4  -3.22% ± 1.09%

S390:
=====
z16: 16 threads = 8 cores (SMT-2)
Benchmark      #VMs    #CPUs/VM  ΔRPS    (std%)
===============================================
pgbench           2           8  73.50% ± 35.91%
pgbench          16           4  61.30% ± 4.09%
hackbench        16           4  54.11% ± 4.38%
hackbench        12           4  36.34% ± 4.63%
pgbench          12           4  34.83% ± 2.57%
hackbench         8           4  29.75% ± 5.86%
hackbench         8           8  25.98% ± 5.09%
pgbench           2           4  23.31% ± 33.44%
pgbench           2          16  19.95% ± 17.12%
hackbench         4           8  19.43% ± 9.33%
pgbench           8           4  19.32% ± 4.50%
[...]
schbench          8           8  -0.79% ± 0.33%
sysbench          8           8  -0.81% ± 0.39%
hackbench         4          16  -1.11% ± 5.82%
sysbench          8           4  -1.62% ± 0.49%
sysbench         16           4  -2.70% ± 0.58%
schbench         16           4  -2.73% ± 0.91%
sysbench         12           4  -2.91% ± 0.61%
hackbench         2          24  -4.99% ± 3.31%

Summary:
- Many improvement across archs specially with real life workloads.
- No major regressions observed.
- Overhead of steal_governor looks minimal when there is no steal time.
- Overhead when STEAL_GOVERNOR=n is negligible.

Testing and Validation
======================

Apart from performance, To ensure the robustness of the preferred
CPU masking and push mechanisms, the following scenarios were tested:
- CPU Hotplug: bringing CPUs up/down change the preferred mask
  accordingly under no-contention and contention.
- Housekeeping cores: Verified with different combinations of
  nohz_full=<beginning, middle, end set of CPUs> to ensure that
  policy engine restricts to first housekeeping core in extreme cases.
- User Affinity: Confirmed that tasks explicitly pinned to non-preferred
  CPUs via taskset remain on their assigned CPUs.
- Affine Move: Confirmed the affinity move using "taskset -cp" happens
  on all combinations of non-preferred, non-preferred under contention.
- Affinity and hotplug: It works as expected. I.e affinity gets
  reset if all the CPUs of p->cpus_ptr go offline even if they are
  non-preferred CPUs.
- Extreme load and running threads: for example 4800 stress-ng threads
  on 480 CPU system and it still packs to preferred CPUs.

Known Limitations & Future Work
===============================

To keep this initial implementation clean and minimal, a few optimizations
have been deferred:

- Push all tasks on rq: Currently, the stopper thread only pushes the current
  running task off a non-preferred CPU. Future optimizations may look into
  migrating all queued tasks on that runqueue.

- Sched Classes: This feature currently only works for the FAIR
  class. Real-time (RT) and sched_ext classes are deferred for now,
  as there is no need for it.

- Arch specific hints and framework for it as been deferred to
  the future.

- NUMA Splicing: The steal_governor currently removes last active core
  based on CPU number. It does not yet do complex NUMA-aware splicing,
  expecting that CPUs are spread out uniformly across nodes in
  most cases.

Shrikanth Hegde (12):
  sched/cputime: Add kcpustat_field_total helper
  sched/docs: Document cpu_preferred_mask and Preferred CPU concept
  cpumask: Introduce cpu_preferred_mask
  sysfs: Add preferred CPU file
  sched/core: Try to use a preferred CPU in is_cpu_allowed
  sched/fair: Load balance only among preferred CPUs
  sched/core: Push current task from non preferred CPU
  sched/debug: Add migration stats due to non preferred CPUs
  virt: Introduce steal governor driver
  virt/steal_governor: Add control knobs for handling steal values
  virt/steal_governor: Implement steal_governor policy loop
  virt/steal_governor: Enable the driver

 .../ABI/testing/sysfs-devices-system-cpu      |  14 +
 Documentation/driver-api/index.rst            |   1 +
 Documentation/driver-api/steal-governor.rst   | 137 +++++++++
 Documentation/scheduler/index.rst             |   1 +
 Documentation/scheduler/sched-paravirt.rst    |  67 ++++
 MAINTAINERS                                   |   9 +
 arch/s390/kernel/hiperdispatch.c              |   8 +-
 drivers/base/cpu.c                            |  12 +
 drivers/virt/Kconfig                          |  18 ++
 drivers/virt/Makefile                         |   1 +
 drivers/virt/steal_governor.c                 | 286 ++++++++++++++++++
 fs/proc/uptime.c                              |   6 +-
 include/linux/cpumask.h                       |  24 ++
 include/linux/kernel_stat.h                   |  11 +
 include/linux/sched.h                         |   1 +
 kernel/Kconfig.preempt                        |   4 +
 kernel/cpu.c                                  |   6 +
 kernel/sched/core.c                           | 112 ++++++-
 kernel/sched/debug.c                          |   1 +
 kernel/sched/fair.c                           |   8 +-
 kernel/sched/sched.h                          |   8 +
 21 files changed, 717 insertions(+), 18 deletions(-)
 create mode 100644 Documentation/driver-api/steal-governor.rst
 create mode 100644 Documentation/scheduler/sched-paravirt.rst
 create mode 100644 drivers/virt/steal_governor.c

-- 
2.47.3

Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Shrikanth Hegde 1 month, 1 week ago
Hi.

In addition to what's currently planned for v11 which was posted here,
https://lore.kernel.org/all/895a058a-475e-42ca-a7a3-2c854598eea4@linux.ibm.com/

I was going through sashiko's comments at:
https://sashiko.dev/#/patchset/20260812054033.95658-1-sshegde%40linux.ibm.com
This has revealed some gaps. Thanks to some really nice insights too.
Report quality improving day by day!


Vincent, Dietmar, please check the 32-bit task issue fix on ARM64.


On 8/12/26 11:10 AM, Shrikanth Hegde wrote:
> If you have already read v8,v9 cover-letter then see only revision
> changes. everything else is pretty much same. :)
> 

> v9->v10:
> - Introduce kcpustat_field_total helper. (Yury Norov)
> - Always do the design checks. This helps to avoid placing design
>    constraints in core hotplug code.
> - Remove cpu_preferred check in idle balancing. This helps to naturally
>    take care update of nohz.next_balance.
> - find_new_ilb changes are deferred as it isn't applicable for most
>    common use cases.
> - Move scheduler documentation to sched-paravirt.rst. (Yury Norov)
> - Add details of limitation of default values in documentation. (Yury Norov)
> - Remove task_can_sched_on_preferred out of sched.h (Mete Durlu)
> - Updated suggested-by tags for few patches. (I know i should have
>    done it earlier, sorry about that)
> - Minor polish of all changelogs.
> 



++++ Patch [1]: [PATCH v10 01/12] sched/cputime: Add kcpustat_field_total helper ++++
Issue:
======
>>  	int cpus, cpu;
>Does this code leave the local variable cpu unused?
>Since the for_each_cpu loop was replaced with a call to kcpustat_field_total,
>cpu does not appear to be referenced anymore in hd_calculate_steal_percentage.

Fix:
====
I did miss to notice it. Will fix it in v11.

index e5c7c818c178..c21496f0a141 100644
--- a/arch/s390/kernel/hiperdispatch.c
+++ b/arch/s390/kernel/hiperdispatch.c
@@ -207,7 +207,7 @@ static unsigned long hd_calculate_steal_percentage(void)
  {
  	unsigned long time_delta, steal_delta, steal, percentage;
  	static ktime_t prev;
-	int cpus, cpu;
+	int cpus;
  	ktime_t now;
  

++++  Patch [5]: [PATCH v10 05/12] sched/core: Try to use a preferred CPU in is_cpu_allowed ++++

Issue1: Possible crash on 32-bit tasks on ARM64.
=======
>> +static inline bool task_can_sched_on_preferred(int cpu, struct task_struct *p)
>> +{
>> +	if (cpu_preferred(cpu))
>> +		return false;
>> +
>> +	/* Only FAIR tasks honor preferred CPU state */
>> +	if (unlikely(p->sched_class != &fair_sched_class))
>> +		return false;
>> +
>> +	return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
>> +}
>Does this intersection check need to account for the architectural CPU mask?
>On asymmetric systems, 32-bit tasks are architecturally restricted by
>task_cpu_possible_mask(). If a 32-bit task's mask intersects with
>64-bit-only preferred CPUs, this function might return true, causing
>is_cpu_allowed() to falsely return false for valid 32-bit non-preferred CPUs.
>Since 64-bit CPUs are rightfully rejected by task_allowed_on_cpu(), all CPUs
>end up rejected. Could this regression cause the select_fallback_rq() loop
>to exhaust all options and hit the BUG() case for 32-bit tasks?

Fix:
====
I wasn;t aware of this case, thanks to sashiko for bring it up.
Yes, it could potentially cause a BUG in select_fallback_rq.

Do a simple check if mask differ from possible mask which indicates we are on 32-bit task on 64 bit
kernel. Do the below. I think that should solve it.

  static inline bool task_can_sched_on_preferred(int cpu, struct task_struct *p)
  {
+	const struct cpumask *valid_mask;
+	int i;
[...]
+	valid_mask = task_cpu_possible_mask(p);
+	if (likely(valid_mask == cpu_possible_mask))
+		return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
+
+	/* 32-bit task */
+	for_each_cpu_and(i, p->cpus_ptr, cpu_preferred_mask) {
+		if (cpumask_test_cpu(i, valid_mask))
+			return true;
+	}


Issue2:
=======
>How does this impact the migration stopper thread during sched_setaffinity?
>When sched_setaffinity updates p->cpus_ptr, it schedules a stopper thread
>to migrate the task. The destination CPU is selected without knowledge of the
>new preference logic in __set_cpus_allowed_ptr_locked():
>    dest_cpu = cpumask_any_and_distribute(cpu_valid_mask, ctx->new_mask);
>If this randomly chosen destination is a non-preferred CPU, __migrate_task()
>will observe is_cpu_allowed() returning false:
>kernel/sched/core.c:__migrate_task() {
>    ...
>    if (!is_cpu_allowed(p, dest_cpu))
>        return rq;
>    ...
>}
>Will this regression cause __migrate_task() to silently abort, leaving the
>task running outside its new explicitly allowed cpumask?

Fix:
====
This is a possible race with script which tries to change affinity at the same window
where preferred CPU mask has changed, but sched_tick hasn't run yet and new affinity
mask is spanning across both preferred and non-preferred CPUs and dest_cpu chosen by
cpumask_any_and_distribute is a non-preferred CPU.
So it is a very narrow case, but a possible one. This results in a
task which never gives up CPU, might run outside of its affinity.

Fix it by checking new affinity allows it run where it is currently running.
If not, don't make decision based on preferred CPU state.
cpus_ptr is updated before __migrate_task.

@@ -2503,7 +2506,20 @@ static inline bool task_can_sched_on_preferred(int cpu, struct task_struct *p)
  	if (unlikely(p->sched_class != &fair_sched_class))
  		return false;
  
+	if (unlikely(!cpumask_test_cpu(task_cpu(p), p->cpus_ptr)))
+		return false;


++++ Patch [7]: [PATCH v10 07/12] sched/core: Push current task from non preferred CPU ++++

Issue1:
>> +#ifdef CONFIG_PREFERRED_CPU
>> +static DEFINE_PER_CPU(struct cpu_stop_work, npc_push_task_work);
>> +
>> +static int sched_non_preferred_cpu_push_stop(void *arg)
>> +{
>> +	struct task_struct *p = arg;
>> +	struct rq *rq = this_rq();
>> +	struct rq_flags rf;
>> +	int cpu;
>> +
>> +	if (cpu_preferred(rq->cpu)) {
>> +		scoped_guard(rq_lock, rq)
>Does acquiring rq_lock here without disabling local interrupts create a
>deadlock window?
>The stopper thread executes this callback with preemption disabled but local
>interrupts enabled. If a timer interrupt fires while this lock is held, the
>interrupt handler will call sched_tick() which attempts to acquire the local
>rq->lock via rq_lock(rq, &rf), leading to an immediate self-deadlock.
>Should this early exit path use rq_lock_irq or rq_lock_irqsave instead?

Fix:
===
Yes a possible window exists. Disable irq when taking rq lock.

diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index b2c011824a7d..51cf0b40a7dd 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -11335,7 +11335,7 @@ static int sched_non_preferred_cpu_push_stop(void *arg)
         int cpu;
  
         if (cpu_preferred(rq->cpu)) {
-               scoped_guard(rq_lock, rq)
+               scoped_guard(rq_lock_irqsave, rq)
                         rq->push_task_work_done = false;
                 put_task_struct(p);
                 return 0;


Issue2:
======
>> +void sched_push_current_non_preferred_cpu(struct rq *rq)
>> +{
>> +	struct task_struct *push_task = rq->curr;
>> +
>> +	scoped_guard(rq_lock, rq) {
>> +		/* Push the task if its explicit affinity allows */
>> +		if (!task_can_sched_on_preferred(rq->cpu, push_task))
>> +			return;
>> +
>> +		/* There is already a stopper thread. Don't race with it. */
>> +		if (rq->push_task_work_done)
>> +			return;
>> +
>> +		rq->push_task_work_done = true;
>> +	}
>Could missing a check for is_migration_disabled(push_task) here cause a
>severe preemption storm?
>Since the stopper thread unconditionally queues without verifying if the task's
>migration is temporarily disabled, it will preempt the task. When the stopper
>observes !is_migration_disabled(p) is false, it resets rq->push_task_work_done
>and exits without migrating the task.
>On the very next timer tick, sched_tick() will observe the task is still on a
>non-preferred CPU and queue the stopper thread again. This continuous queueing
>of the highest-priority stopper thread every tick while migration is disabled
>defeats the purpose of the migration-disabled section.

Fix:
====
Do the same check before queuing the stopper too.
i.e. same check that happens in stopper.

@@ -11378,6 +11394,9 @@ void sched_push_current_non_preferred_cpu(struct rq *rq)
  		if (rq->push_task_work_done)
  			return;
  
+		if (is_migration_disabled(push_task))
+			return;
+
  		rq->push_task_work_done = true;



============================================================

Other comments which are worth noting, but are not a concern.

- time of use, time of check issue in select_fallback_rq w.r.t to preferred
   mask change. As explained in earlier changeset, this cannot happen since
   select_fallback_rq does two loop. First of nodemask, and then cpus_ptr.
   Lets due to concurrent mask change, first one fails, then by second loop, mask
   will be stable, and cannot race again. Mask updates by 100ms at least.

- Overloading of preferred CPUs. That is expected by design.

- Could the __read_mostly annotation on __cpu_preferred_mask cause cache line
   bouncing and false sharing? Kept as __read_mostly as majority of the time is
   isn't changing.

- Ping-pong doesn't happen since load balance doesn't push tasks onto preferred
   CPUs.

- Does triggering select_fallback_rq() on the hot wakeup path introduce a
   lock contention bottleneck? - yes but not too much, but adding more
   checks there, add more overhead in generic case.
   So it is optimization that is avoided at the moment.

- A non-preferred CPU isn't expected to pull any load and there is no load balancing
   among non-preferred CPUs as said in the changelog.
Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Dietmar Eggemann 1 month ago
Hi Shrikanth,

On 17.08.26 09:39, Shrikanth Hegde wrote:
> Hi.
> 
> In addition to what's currently planned for v11 which was posted here,
> https://lore.kernel.org/all/895a058a-475e-42ca-
> a7a3-2c854598eea4@linux.ibm.com/
> 
> I was going through sashiko's comments at:
> https://sashiko.dev/#/patchset/20260812054033.95658-1-
> sshegde%40linux.ibm.com
> This has revealed some gaps. Thanks to some really nice insights too.
> Report quality improving day by day!
> 
> 
> Vincent, Dietmar, please check the 32-bit task issue fix on ARM64.

See below.

> On 8/12/26 11:10 AM, Shrikanth Hegde wrote:

[...]

> Issue1: Possible crash on 32-bit tasks on ARM64.
> =======
>>> +static inline bool task_can_sched_on_preferred(int cpu, struct
>>> task_struct *p)
>>> +{
>>> +    if (cpu_preferred(cpu))
>>> +        return false;
>>> +
>>> +    /* Only FAIR tasks honor preferred CPU state */
>>> +    if (unlikely(p->sched_class != &fair_sched_class))
>>> +        return false;
>>> +
>>> +    return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
>>> +}
>> Does this intersection check need to account for the architectural CPU
>> mask?
>> On asymmetric systems, 32-bit tasks are architecturally restricted by
>> task_cpu_possible_mask(). If a 32-bit task's mask intersects with
>> 64-bit-only preferred CPUs, this function might return true, causing
>> is_cpu_allowed() to falsely return false for valid 32-bit non-
>> preferred CPUs.
>> Since 64-bit CPUs are rightfully rejected by task_allowed_on_cpu(),
>> all CPUs
>> end up rejected. Could this regression cause the select_fallback_rq()
>> loop
>> to exhaust all options and hit the BUG() case for 32-bit tasks?
> 
> Fix:
> ====
> I wasn;t aware of this case, thanks to sashiko for bring it up.
> Yes, it could potentially cause a BUG in select_fallback_rq.
> 
> Do a simple check if mask differ from possible mask which indicates we
> are on 32-bit task on 64 bit
> kernel. Do the below. I think that should solve it.
> 
>  static inline bool task_can_sched_on_preferred(int cpu, struct
> task_struct *p)
>  {
> +    const struct cpumask *valid_mask;
> +    int i;
> [...]
> +    valid_mask = task_cpu_possible_mask(p);
> +    if (likely(valid_mask == cpu_possible_mask))
> +        return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
> +
> +    /* 32-bit task */
> +    for_each_cpu_and(i, p->cpus_ptr, cpu_preferred_mask) {
> +        if (cpumask_test_cpu(i, valid_mask))
> +            return true;
> +    }

I assume the question is whether task_can_sched_on_preferred() would
have to be changed:

- return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
+ return cpumask_first_and_and(p->cpus_ptr, cpu_preferred_mask,
+                               task_cpu_possible_mask(p)) < nr_cpu_ids;

so that cpu_preferred_mask can play together nicely with the 'asymmetric
AArch32 EL0 (executing 32-bit Arm userspace under an AArch64 kernel)
support' feature on some mobile Arm64 Socs.

IMHO, this is not necessary since for those tasks p->cpus_ptr is always
a subset of task_cpu_possible_mask(p). 'p->cpus_ptr ∩
cpu_preferred_mask' already cannot contain an architecturally impossible
CPU for those 32-bit Arm userspace tasks.

[...]

Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Shrikanth Hegde 1 month ago

On 8/27/26 4:20 PM, Dietmar Eggemann wrote:
> Hi Shrikanth,
> 

Hi Dietmar. Thanks for checking it.

> On 17.08.26 09:39, Shrikanth Hegde wrote:
>> Hi.
>>
>> In addition to what's currently planned for v11 which was posted here,
>> https://lore.kernel.org/all/895a058a-475e-42ca-
>> a7a3-2c854598eea4@linux.ibm.com/
>>
>> I was going through sashiko's comments at:
>> https://sashiko.dev/#/patchset/20260812054033.95658-1-
>> sshegde%40linux.ibm.com
>> This has revealed some gaps. Thanks to some really nice insights too.
>> Report quality improving day by day!
>>
>>
>> Vincent, Dietmar, please check the 32-bit task issue fix on ARM64.
> 
> See below.
> 
>> On 8/12/26 11:10 AM, Shrikanth Hegde wrote:
> 
> [...]
> 
>> Issue1: Possible crash on 32-bit tasks on ARM64.
>> =======
>>>> +static inline bool task_can_sched_on_preferred(int cpu, struct
>>>> task_struct *p)
>>>> +{
>>>> +    if (cpu_preferred(cpu))
>>>> +        return false;
>>>> +
>>>> +    /* Only FAIR tasks honor preferred CPU state */
>>>> +    if (unlikely(p->sched_class != &fair_sched_class))
>>>> +        return false;
>>>> +
>>>> +    return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
>>>> +}
>>> Does this intersection check need to account for the architectural CPU
>>> mask?
>>> On asymmetric systems, 32-bit tasks are architecturally restricted by
>>> task_cpu_possible_mask(). If a 32-bit task's mask intersects with
>>> 64-bit-only preferred CPUs, this function might return true, causing
>>> is_cpu_allowed() to falsely return false for valid 32-bit non-
>>> preferred CPUs.
>>> Since 64-bit CPUs are rightfully rejected by task_allowed_on_cpu(),
>>> all CPUs
>>> end up rejected. Could this regression cause the select_fallback_rq()
>>> loop
>>> to exhaust all options and hit the BUG() case for 32-bit tasks?
>>
>> Fix:
>> ====
>> I wasn;t aware of this case, thanks to sashiko for bring it up.
>> Yes, it could potentially cause a BUG in select_fallback_rq.
>>
>> Do a simple check if mask differ from possible mask which indicates we
>> are on 32-bit task on 64 bit
>> kernel. Do the below. I think that should solve it.
>>
>>   static inline bool task_can_sched_on_preferred(int cpu, struct
>> task_struct *p)
>>   {
>> +    const struct cpumask *valid_mask;
>> +    int i;
>> [...]
>> +    valid_mask = task_cpu_possible_mask(p);
>> +    if (likely(valid_mask == cpu_possible_mask))
>> +        return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
>> +
>> +    /* 32-bit task */
>> +    for_each_cpu_and(i, p->cpus_ptr, cpu_preferred_mask) {
>> +        if (cpumask_test_cpu(i, valid_mask))
>> +            return true;
>> +    }
> 
> I assume the question is whether task_can_sched_on_preferred() would
> have to be changed:
> 
> - return cpumask_intersects(p->cpus_ptr, cpu_preferred_mask);
> + return cpumask_first_and_and(p->cpus_ptr, cpu_preferred_mask,
> +                               task_cpu_possible_mask(p)) < nr_cpu_ids;
>> so that cpu_preferred_mask can play together nicely with the 'asymmetric
> AArch32 EL0 (executing 32-bit Arm userspace under an AArch64 kernel)
> support' feature on some mobile Arm64 Socs.
> 
> IMHO, this is not necessary since for those tasks p->cpus_ptr is always
> a subset of task_cpu_possible_mask(p). 'p->cpus_ptr ∩
> cpu_preferred_mask' already cannot contain an architecturally impossible
> CPU for those 32-bit Arm userspace tasks.
> 
> [...]
> 

Based on the report, I thought there maybe cases where p->cpus_ptr may contain.
Specifically after looking at:

static inline bool task_allowed_on_cpu(struct task_struct *p, int cpu)
{
         /* When not in the task's cpumask, no point in looking further. */
         if (!cpumask_test_cpu(cpu, p->cpus_ptr))
                 return false;

         /* Can @cpu run a user thread? */
         if (!(p->flags & PF_KTHREAD) && !task_cpu_possible(cpu, p))
                 return false;

         return true;
}

If p->cpus_ptr is cannot contain an architecturally impossible, then check for
task_cpu_possible again is necessary? I thought there may be cases.
So i kept the defensive check not to fall into BUG later on.

If you think cpumask_intersects(p->cpus_ptr, cpu_preferred_mask) is sufficient and cover
all cases of 32 bit tasks, then i can drop that v11 change specific to 32-bit tasks.

What do you suggest?
Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Yury Norov 1 month, 1 week ago
> ========================================
> Performance Numbers (powerpc, x86, s390)
> ========================================
> 
> PowerPC:
> ===================
> VM1: 60VP/30EC and VM2: 30VP/20EC
> Shared physical CPU pool size: 50 Cores. Each core is SMT8.
> (VP - Virtual Core, EC - Entitles Core) -  PowerVM terminologies of SPLPAR[2]
> 
> Default parameter values: 1000ms, 200 low threshold, 500 high threshold
> Both the VMs are running the same workload. Total throughput/time of VM1+VM2
> is being mentioned in all cases.
> 
> Hackbench
>               baseline    steal_governor        steal_governor
>                              disabled               enabled
> ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
> 
> 10 groups        5.20   |    5.40 (-3.85%)  |     4.65 (+10.58%)
> 20 groups       11.39   |   12.01 (-5.44%)  |     7.09 (+37.75%)
> 40 groups       20.32   |   19.80 (+2.56%)  |    11.31 (+44.34%)
> 10 groups(-p)    2.37   |    2.26 (+4.64%)  |     2.06 (+13.08%)
> 20 groups(-p)    3.34   |    3.28 (+1.80%)  |     3.20 (+4.19%)
> 40 groups(-p)    4.46   |    4.83 (-8.30%)  |     4.26 (+4.48%)
> Remarks: Net improvement with steal_governor specially high load points.
> 
> schbench ( -L -n 0 -r 30 -s 0)
>               baseline    steal_governor          steal_governor
>                              disabled                 enabled
> ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
> -m 1 -t 128     2475162 |    2621246 (+5.90%)  |      2527299 (+2.11%)
> -m 1 -t 256     1467350 |    1470032 (+0.18%)  |      1492372 (+1.71%)
> -m 1 -t 512     1408813 |    1454687 (+3.26%)  |      1437605 (+2.04%)
> Remarks: Effectively means no-improvements or regressions
> 
> kernbench	baseline    steal_governor     steal_governor
> (elapsed time)	               disabled            enabled
> ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
> -j nr_cpus	231      |      235 (-1.7%) |    199 (+14%)
> Remarks: Net improvement in elapsed time.
> 
> Daytrader - A real life work which is a proxy for trading based
> on db2[3]
>               baseline      steal_governor   steal_governor
>                               disabled          enabled
> ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
> Load@30%	1x	|	0.96x	|	 1.53x			
> Load@60%	1x	|	0.94x	|	 1.41x
> Remarks: Good improvement seen at different load points.
> 
> When there is no steal time (such as dedicated LPAR, or only VM2
> is running) throughput was same with steal_governor enabled/disabled
> which indicates minimal overhead of steal_governor. 
> 
> I have run v10 also on a smaller powerpc LPAR system and it shows
> good improvements.
> 
> =======================================================================
> 
> Data from x86,s390 KVM which Ilya Leoshkevich carried out during OSPM26
> time. *This was based on v2*. Idea is still the name, numbers are
> expected to be better in v10 as some of the overhead has been removed.
> Note: Other variations of the benchmark shows no observable
> difference.
> 
> x86:
> ====
> cascade-lake: 32 threads = 16 cores
> Benchmark      #VMs    #CPUs/VM  ΔRPS     (%std)
> ===============================================
> hackbench         8          16  90.73% ± 9.97%
> hackbench         4          24  52.67% ± 7.43%
> hackbench         4          16  37.96% ± 11.19%
> hackbench         4          32  37.82% ± 4.38%
> hackbench        12           8  36.90% ± 4.74%
> hackbench         8           8  35.30% ± 3.61%
> pgbench          16           4  31.77% ± 2.44%
> hackbench         2          24  25.85% ± 8.63%
> hackbench        16           8  24.87% ± 3.46%
> pgbench          16           8  21.83% ± 2.20%
> pgbench          12           8  21.35% ± 2.15%
> pgbench           8           8  18.46% ± 1.01%
> hackbench         2          32  15.56% ± 4.53%
> pgbench          12           4  14.28% ± 2.04%
> hackbench        16           4  14.07% ± 2.90%
> hackbench        12           4  9.60% ± 3.49%
> [...]
> pgbench           4           8  -1.16% ± 3.60%
> hackbench         4           4  -1.80% ± 9.55%
> sysbench         12           4  -2.19% ± 0.78%
> pgbench           4          24  -2.43% ± 4.38%
> pgbench           4          32  -3.21% ± 0.79%
> sysbench         16           4  -3.22% ± 1.09%
> 
> S390:
> =====
> z16: 16 threads = 8 cores (SMT-2)
> Benchmark      #VMs    #CPUs/VM  ΔRPS    (std%)
> ===============================================
> pgbench           2           8  73.50% ± 35.91%
> pgbench          16           4  61.30% ± 4.09%
> hackbench        16           4  54.11% ± 4.38%
> hackbench        12           4  36.34% ± 4.63%
> pgbench          12           4  34.83% ± 2.57%
> hackbench         8           4  29.75% ± 5.86%
> hackbench         8           8  25.98% ± 5.09%
> pgbench           2           4  23.31% ± 33.44%
> pgbench           2          16  19.95% ± 17.12%
> hackbench         4           8  19.43% ± 9.33%
> pgbench           8           4  19.32% ± 4.50%
> [...]
> schbench          8           8  -0.79% ± 0.33%
> sysbench          8           8  -0.81% ± 0.39%
> hackbench         4          16  -1.11% ± 5.82%
> sysbench          8           4  -1.62% ± 0.49%
> sysbench         16           4  -2.70% ± 0.58%
> schbench         16           4  -2.73% ± 0.91%
> sysbench         12           4  -2.91% ± 0.61%
> hackbench         2          24  -4.99% ± 3.31%
> 
> Summary:
> - Many improvement across archs specially with real life workloads.
> - No major regressions observed.
> - Overhead of steal_governor looks minimal when there is no steal time.
> - Overhead when STEAL_GOVERNOR=n is negligible.

OK, I gave it some testing on my laptop.

The results are pretty consistent: the steal ratio is converged to a
number withing the threshold, but the overall performance is 3-5% worse
comparing to baseline. I tried 2-5% and 1.5-15% boundaries.

It's 4 VMs, each running 8 vCPUs on 8 pCPU machine, the payload is
running for 2 minutes.

  Steal governor off:
  VM         THROUGHPUT       AVG STEAL%
  -------- ------------ ----------------
  0               13802            74.59
  1               11929            76.35
  2               13470            74.61
  3               10259            78.67
  -------- ------------ ----------------
  TOTAL           49460                -
  
  Steal governor on:
  VM         THROUGHPUT       AVG STEAL%
  -------- ------------ ----------------
  0               12921             5.72
  1               12762             6.11
  2               10160             5.99
  3               12098             6.00
  -------- ------------ ----------------
  TOTAL           47941                -


The test is attached below. The results are quite differ from the numbers
above, so maybe I misconfigured something? I didn't use hackbench or
similar benchmarks, just a basic math.

Shrikanth, can you please check my test and results? Is there something 
that I have missed?

I think this series should include some testing. The scripts below look
bulky and they depend on virtme, but they allow to build the proper
kernel and run tests with a single command.

Thanks,
Yury

From 27e548d65bbe0ec01308d3329017619e165d0488 Mon Sep 17 00:00:00 2001
From: Yury Norov <ynorov@nvidia.com>
Date: Wed, 19 Aug 2026 16:31:33 -0400
Subject: [PATCH] steal governor: add testing harness based on vng

Signed-off-by: Yury Norov <ynorov@nvidia.com>
---
 .../virtme-steal-payload.sh                   | 220 ++++++++++++
 .../steal_governor_test/virtme-steal-time.sh  | 321 ++++++++++++++++++
 2 files changed, 541 insertions(+)
 create mode 100755 drivers/virt/steal_governor_test/virtme-steal-payload.sh
 create mode 100755 drivers/virt/steal_governor_test/virtme-steal-time.sh

diff --git a/drivers/virt/steal_governor_test/virtme-steal-payload.sh b/drivers/virt/steal_governor_test/virtme-steal-payload.sh
new file mode 100755
index 000000000000..02a245a900b8
--- /dev/null
+++ b/drivers/virt/steal_governor_test/virtme-steal-payload.sh
@@ -0,0 +1,220 @@
+#!/bin/sh
+# Guest-side measurable workload for the steal governor test.
+
+set -eu
+
+case $# in
+2|3|5) ;;
+*)
+	echo "usage: ${0##*/} VM_ID DURATION_SECONDS [DRIVER [LOW_THRESHOLD HIGH_THRESHOLD]]" >&2
+	exit 2
+	;;
+esac
+
+vm_id=$1
+duration=$2
+driver=${3-}
+low_threshold=${4-}
+high_threshold=${5-}
+
+verify_parameter()
+{
+	parameter=$1
+	expected=$2
+	parameter_file=/sys/module/$driver_sysfs/parameters/$parameter
+
+	if [ ! -r "$parameter_file" ]; then
+		echo "error: driver $driver has no readable $parameter parameter" >&2
+		exit 1
+	fi
+	actual=$(cat "$parameter_file")
+	if [ "$actual" != "$expected" ]; then
+		echo "error: driver $driver $parameter is $actual, expected $expected" >&2
+		exit 1
+	fi
+}
+
+case $vm_id in
+*[!0-9]*|'')
+	echo "error: VM_ID must be a non-negative integer" >&2
+	exit 2
+	;;
+esac
+case $duration in
+*[!0-9]*|'')
+	echo "error: DURATION_SECONDS must be a positive integer" >&2
+	exit 2
+	;;
+esac
+if [ "$duration" -eq 0 ]; then
+	echo "error: DURATION_SECONDS must be greater than zero" >&2
+	exit 2
+fi
+if [ -n "$driver" ]; then
+	case $driver in
+	*[!A-Za-z0-9_.-]*)
+		echo "error: invalid driver module name: $driver" >&2
+		exit 2
+		;;
+	esac
+	if [ -n "$low_threshold" ]; then
+		case $low_threshold:$high_threshold in
+		*[!0-9:]*|:*|*:)
+			echo "error: thresholds must be non-negative integers in percent * 100" >&2
+			exit 2
+			;;
+		esac
+		if [ "$low_threshold" -ge "$high_threshold" ] ||
+		   [ "$high_threshold" -ge 10000 ]; then
+			echo "error: thresholds must satisfy 0 <= LOW < HIGH < 10000" >&2
+			exit 2
+		fi
+	fi
+	if ! command -v modprobe >/dev/null; then
+		echo "error: modprobe is required to load driver $driver" >&2
+		exit 1
+	fi
+	driver_sysfs=$(printf '%s\n' "$driver" | tr '-' '_')
+	if [ ! -d "/sys/module/$driver_sysfs" ]; then
+		if [ -n "$low_threshold" ]; then
+			modprobe "$driver" \
+				low_threshold="$low_threshold" \
+				high_threshold="$high_threshold"
+		else
+			modprobe "$driver"
+		fi
+	fi
+	if [ ! -d "/sys/module/$driver_sysfs" ]; then
+		echo "error: driver $driver has no /sys/module/$driver_sysfs entry after modprobe" >&2
+		exit 1
+	fi
+	if [ -n "$low_threshold" ]; then
+		verify_parameter low_threshold "$low_threshold"
+		verify_parameter high_threshold "$high_threshold"
+	fi
+fi
+
+tmpdir=$(mktemp -d "${TMPDIR:-/tmp}/virtme-steal-payload.XXXXXXXX")
+before=$tmpdir/before
+after=$tmpdir/after
+pids=
+
+stop_workers()
+{
+	for pid in $pids; do
+		kill "$pid" 2>/dev/null || :
+	done
+	for pid in $pids; do
+		wait "$pid" 2>/dev/null || :
+	done
+	pids=
+}
+
+cleanup()
+{
+	stop_workers
+	rm -rf -- "$tmpdir"
+}
+trap cleanup EXIT
+trap 'exit 130' INT
+trap 'exit 143' TERM
+
+awk '$1 ~ /^cpu[0-9]+$/ { print $1, $9 }' /proc/stat >"$before"
+start_time=$(awk '{ print $1 }' /proc/uptime)
+deadline=$(awk -v start="$start_time" -v seconds="$duration" \
+	'BEGIN { printf "%.2f", start + seconds }')
+workers=$(awk '$1 ~ /^cpu[0-9]+$/ { n++ } END { print n }' /proc/stat)
+
+i=0
+while [ "$i" -lt "$workers" ]; do
+	result=$tmpdir/work-$i
+	awk -v deadline="$deadline" -v result="$result" '
+		function uptime(    line, fields) {
+			getline line < "/proc/uptime"
+			close("/proc/uptime")
+			split(line, fields)
+			return fields[1]
+		}
+		BEGIN {
+			batch = 1000
+			units = 0
+			value = 1
+			while (uptime() < deadline) {
+				for (iteration = 0; iteration < batch; iteration++)
+					value = (value * 1103515245 + 12345) % 2147483647
+				units++
+			}
+			printf "%.0f %.0f\n", units, units * batch > result
+		}
+	' </dev/null &
+	pids="$pids $!"
+	i=$((i + 1))
+done
+
+worker_failed=0
+for pid in $pids; do
+	wait "$pid" || worker_failed=1
+done
+pids=
+if [ "$worker_failed" -ne 0 ]; then
+	echo "error: one or more workload processes failed" >&2
+	exit 1
+fi
+
+awk '$1 ~ /^cpu[0-9]+$/ { print $1, $9 }' /proc/stat >"$after"
+end_time=$(awk '{ print $1 }' /proc/uptime)
+elapsed=$(awk -v start="$start_time" -v end="$end_time" \
+	'BEGIN { printf "%.2f", end - start }')
+clk_tck=$(getconf CLK_TCK 2>/dev/null || echo 100)
+
+echo VIRTME_STEAL_REPORT_BEGIN
+driver_status=none
+[ -n "$driver" ] && driver_status=$driver:loaded
+[ -n "$low_threshold" ] && \
+	driver_status="$driver_status,thresholds=$low_threshold..$high_threshold"
+echo "VM $vm_id  kernel=$(uname -r)  vCPUs=$workers  sample=${elapsed}s  driver=$driver_status"
+echo "Workload (unpinned workers)"
+printf '%-10s %12s %16s %16s\n' \
+	WORKER WORK_UNITS ITERATIONS ITERATIONS/s
+printf '%-10s %12s %16s %16s\n' \
+	---------- ------------ ---------------- ----------------
+total_units=0
+total_iterations=0
+i=0
+while [ "$i" -lt "$workers" ]; do
+	read -r units iterations <"$tmpdir/work-$i"
+	rate=$(awk -v iterations="$iterations" -v elapsed="$elapsed" \
+		'BEGIN { printf "%.0f", iterations / elapsed }')
+	printf '%-10s %12s %16s %16s\n' \
+		"worker$i" "$units" "$iterations" "$rate"
+	total_units=$((total_units + units))
+	total_iterations=$((total_iterations + iterations))
+	i=$((i + 1))
+done
+total_rate=$(awk -v iterations="$total_iterations" -v elapsed="$elapsed" \
+	'BEGIN { printf "%.0f", iterations / elapsed }')
+throughput=$(awk -v units="$total_units" -v elapsed="$elapsed" \
+	'BEGIN { printf "%.0f", units / elapsed }')
+printf '%-10s %12s %16s %16s\n' \
+	TOTAL "$total_units" "$total_iterations" "$total_rate"
+
+echo "Steal time by guest CPU"
+printf '%-8s %12s %12s %10s\n' CPU STEAL_TICKS STEAL_s STEAL_%
+printf '%-8s %12s %12s %10s\n' -------- ------------ ------------ ----------
+average_file=$tmpdir/average-steal
+awk -v hz="$clk_tck" -v elapsed="$elapsed" -v cpus="$workers" \
+	-v average_file="$average_file" '
+	NR == FNR { before[$1] = $2; next }
+	{
+		delta = $2 - before[$1]
+		total += delta
+		printf "%-8s %12d %12.2f %10.2f\n", $1, delta,
+		       delta / hz, 100 * delta / hz / elapsed
+	}
+	END {
+		printf "%.2f\n", 100 * total / hz / elapsed / cpus > average_file
+	}
+' "$before" "$after"
+average_steal=$(cat "$average_file")
+echo "VIRTME_STEAL_SUMMARY $vm_id $throughput $average_steal"
+echo VIRTME_STEAL_REPORT_END
diff --git a/drivers/virt/steal_governor_test/virtme-steal-time.sh b/drivers/virt/steal_governor_test/virtme-steal-time.sh
new file mode 100755
index 000000000000..a41d362c7888
--- /dev/null
+++ b/drivers/virt/steal_governor_test/virtme-steal-time.sh
@@ -0,0 +1,321 @@
+#!/usr/bin/env bash
+# Build a minimal KVM guest kernel and measure per-vCPU steal time under load.
+
+set -euo pipefail
+
+script_dir=$(CDPATH= cd -- "$(dirname -- "$0")" && pwd)
+kernel_dir=$(CDPATH= cd -- "$script_dir/../../.." && pwd)
+vng=${VNG:-vng}
+build_dir=${BUILD_DIR:-"$kernel_dir/.virtme-steal"}
+payload=${PAYLOAD:-"$script_dir/virtme-steal-payload.sh"}
+vm_count=4
+vcpu_count=$(nproc)
+duration=20
+memory=512M
+driver=
+low_threshold=
+high_threshold=
+verbose=0
+declare -a config_items=()
+
+usage()
+{
+	cat <<EOF
+Usage: ${0##*/} [options] [O=DIR]
+
+Build a virtme-ng minimal kernel with paravirtual steal-time accounting,
+start multiple CPU-bound VMs, and report the steal time of every guest CPU.
+
+Options:
+  -n VMS       number of VMs to run (default: $vm_count)
+  -p VCPUS     vCPUs per VM (default: $vcpu_count)
+  -d SECONDS   workload duration (default: $duration)
+  -m MEMORY    memory per VM (default: $memory)
+  --configitem CONFIG[=VALUE]
+                enable or set a kernel config option (repeatable)
+  --driver MODULE
+                load and verify this module before running the payload
+  -lo VALUE     low steal threshold in percent * 100 (for example, 150 = 1.5%)
+  -hi VALUE     high steal threshold in percent * 100 (for example, 550 = 5.5%)
+  -O DIR       kernel build directory (default: $build_dir)
+  -v, --verbose show per-worker and per-CPU result tables
+  -h            show this help
+
+Environment equivalents: VNG, BUILD_DIR, and PAYLOAD. A custom payload is
+called inside each VM as: PAYLOAD VM_ID DURATION_SECONDS
+[DRIVER [LOW_THRESHOLD HIGH_THRESHOLD]].
+
+Local virtme-ng checkout example:
+  VNG=../virtme-ng/vng ${0##*/}
+
+The positional O=DIR form is equivalent to -O DIR, for example:
+  ${0##*/} -n 4 O=../build-linux-virtme-steal
+
+Config examples:
+  ${0##*/} --configitem CONFIG_SCHEDSTATS --configitem CONFIG_HZ_1000=y
+
+Threshold example:
+  ${0##*/} --driver steal_governor -lo 150 -hi 550
+EOF
+}
+
+while (($#)); do
+	case $1 in
+	-n|-p|-d|-m|-O|-lo|-hi|--configitem|--driver)
+		if (($# < 2)); then
+			echo "error: $1 requires an argument" >&2
+			exit 2
+		fi
+		case $1 in
+		-n) vm_count=$2 ;;
+		-p) vcpu_count=$2 ;;
+		-d) duration=$2 ;;
+		-m) memory=$2 ;;
+		-O) build_dir=$2 ;;
+		-lo) low_threshold=$2 ;;
+		-hi) high_threshold=$2 ;;
+		--configitem) config_items+=("$2") ;;
+		--driver) driver=$2 ;;
+		esac
+		shift 2
+		;;
+	--configitem=?*)
+		config_items+=("${1#*=}")
+		shift
+		;;
+	--driver=?*)
+		driver=${1#*=}
+		shift
+		;;
+	-v|--verbose)
+		verbose=1
+		shift
+		;;
+	-h|--help)
+		usage
+		exit 0
+		;;
+	O=?*)
+		build_dir=${1#O=}
+		shift
+		;;
+	O=|--configitem=|--driver=)
+		echo "error: ${1%%=*}= requires a non-empty argument" >&2
+		exit 2
+		;;
+	*)
+		echo "error: unexpected argument: $1" >&2
+		usage >&2
+		exit 2
+		;;
+	esac
+done
+
+require_positive_integer()
+{
+	local name=$1 value=$2
+
+	if [[ ! $value =~ ^[1-9][0-9]*$ ]]; then
+		echo "error: $name must be a positive integer (got '$value')" >&2
+		exit 2
+	fi
+}
+
+require_positive_integer "VM count" "$vm_count"
+require_positive_integer "vCPU count" "$vcpu_count"
+require_positive_integer "duration" "$duration"
+
+if [[ -n $low_threshold || -n $high_threshold ]]; then
+	if [[ -z $low_threshold || -z $high_threshold ]]; then
+		echo "error: -lo and -hi must be specified together" >&2
+		exit 2
+	fi
+	if [[ ! $low_threshold =~ ^[0-9]+$ || ! $high_threshold =~ ^[0-9]+$ ]]; then
+		echo "error: -lo and -hi must be non-negative integers in percent * 100" >&2
+		exit 2
+	fi
+	low_threshold=$((10#$low_threshold))
+	high_threshold=$((10#$high_threshold))
+	if ((low_threshold >= high_threshold)); then
+		echo "error: -lo must be less than -hi" >&2
+		exit 2
+	fi
+	if ((high_threshold >= 10000)); then
+		echo "error: -hi must be less than 10000 (100%)" >&2
+		exit 2
+	fi
+	if [[ -z $driver ]]; then
+		echo "error: -lo and -hi require --driver MODULE" >&2
+		exit 2
+	fi
+fi
+
+if [[ -n $driver && ! $driver =~ ^[A-Za-z0-9_.-]+$ ]]; then
+	echo "error: invalid driver module name: $driver" >&2
+	exit 2
+fi
+
+for index in "${!config_items[@]}"; do
+	if [[ ! ${config_items[index]} =~ ^CONFIG_[A-Z0-9_]+(=.*)?$ ]]; then
+		echo "error: invalid kernel config item: ${config_items[index]}" >&2
+		exit 2
+	fi
+	if [[ ${config_items[index]} != *=* ]]; then
+		config_items[index]="${config_items[index]}=y"
+	fi
+done
+
+if [[ ! -x $payload ]]; then
+	echo "error: guest payload is not executable: $payload" >&2
+	exit 1
+fi
+if [[ ! -r /dev/kvm || ! -w /dev/kvm ]]; then
+	echo "error: /dev/kvm is not accessible; KVM is required for steal-time accounting" >&2
+	exit 1
+fi
+
+case $build_dir in
+/*) ;;
+*) build_dir=$PWD/$build_dir ;;
+esac
+
+mkdir -p "$build_dir"
+cd "$kernel_dir"
+
+echo "==> Configuring paravirtual kernel in $build_dir"
+if [[ ! -f $build_dir/.config ]]; then
+	config_args=()
+	for config_item in "${config_items[@]}"; do
+		config_args+=(--configitem "$config_item")
+	done
+	"$vng" --kconfig "${config_args[@]}" \
+		--configitem CONFIG_HYPERVISOR_GUEST=y \
+		--configitem CONFIG_PARAVIRT=y \
+		--configitem CONFIG_KVM_GUEST=y \
+		--configitem CONFIG_PARAVIRT_TIME_ACCOUNTING=y \
+		-- "O=$build_dir"
+else
+	# Preserve an existing local minimal config and add the options needed by
+	# this test. olddefconfig resolves their dependencies for the current tree.
+	config_args=(--file "$build_dir/.config")
+	for config_item in "${config_items[@]}"; do
+		config_name=${config_item%%=*}
+		config_value=${config_item#*=}
+		config_args+=(--set-val "$config_name" "$config_value")
+	done
+	"$kernel_dir/scripts/config" "${config_args[@]}" \
+		-e HYPERVISOR_GUEST \
+		-e PARAVIRT \
+		-e KVM_GUEST \
+		-e PARAVIRT_TIME_ACCOUNTING
+	make -s O="$build_dir" olddefconfig
+fi
+
+echo "==> Building kernel"
+declare -a module_args=()
+if [[ -z $driver ]]; then
+	module_args+=(--skip-modules)
+fi
+"$vng" --build "${module_args[@]}" -- "O=$build_dir"
+
+for option in HYPERVISOR_GUEST PARAVIRT KVM_GUEST PARAVIRT_TIME_ACCOUNTING; do
+	if ! grep -qx "CONFIG_${option}=y" "$build_dir/.config"; then
+		echo "error: CONFIG_${option}=y is required but is absent after configuration" >&2
+		exit 1
+	fi
+done
+
+tmpdir=$(mktemp -d "${TMPDIR:-/tmp}/virtme-steal.XXXXXXXX")
+declare -a vm_pids=()
+
+cleanup()
+{
+	local pid
+
+	for pid in "${vm_pids[@]}"; do
+		kill "$pid" 2>/dev/null || true
+	done
+	wait 2>/dev/null || true
+	rm -rf -- "$tmpdir"
+}
+trap cleanup EXIT
+trap 'exit 130' INT
+trap 'exit 143' TERM
+
+echo "==> Starting $vm_count VMs ($vcpu_count vCPUs each)"
+for ((vm = 0; vm < vm_count; vm++)); do
+	printf -v guest_script '%q %q %q' "$payload" "$vm" "$duration"
+	if [[ -n $driver ]]; then
+		printf -v driver_arg ' %q' "$driver"
+		guest_script+=$driver_arg
+		if [[ -n $low_threshold ]]; then
+			printf -v threshold_args ' %q %q' \
+				"$low_threshold" "$high_threshold"
+			guest_script+=$threshold_args
+		fi
+	fi
+
+	log=$tmpdir/vm-$vm.log
+	"$vng" \
+		"${module_args[@]}" \
+		--name "steal-vm-$vm" \
+		--cpus "$vcpu_count" \
+		--memory "$memory" \
+		--exec "$guest_script" \
+		-- "O=$build_dir" >"$log" 2>&1 &
+	vm_pids+=("$!")
+done
+
+status=0
+total_throughput=0
+if (( ! verbose )); then
+	printf '%-8s %12s %16s\n' VM THROUGHPUT 'AVG STEAL%'
+	printf '%-8s %12s %16s\n' -------- ------------ ----------------
+fi
+for ((vm = 0; vm < vm_count; vm++)); do
+	if ! wait "${vm_pids[vm]}"; then
+		echo "error: VM $vm failed; its complete log follows" >&2
+		status=1
+		echo "--- VM $vm log ---"
+		cat "$tmpdir/vm-$vm.log"
+		continue
+	fi
+	if ! grep -qx VIRTME_STEAL_REPORT_BEGIN "$tmpdir/vm-$vm.log" ||
+	   ! grep -qx VIRTME_STEAL_REPORT_END "$tmpdir/vm-$vm.log"; then
+		echo "error: VM $vm exited without a steal-time report; its complete log follows" >&2
+		cat "$tmpdir/vm-$vm.log"
+		status=1
+		continue
+	fi
+
+	summary=$(awk '/^VIRTME_STEAL_SUMMARY / { print $2, $3, $4 }' \
+		"$tmpdir/vm-$vm.log")
+	if [[ -z $summary ]]; then
+		echo "error: VM $vm exited without a summary; its complete log follows" >&2
+		cat "$tmpdir/vm-$vm.log"
+		status=1
+		continue
+	fi
+
+	if (( verbose )); then
+		echo "--- VM $vm results ---"
+		awk '
+			/^VIRTME_STEAL_REPORT_BEGIN$/ { report = 1; next }
+			/^VIRTME_STEAL_REPORT_END$/ { report = 0 }
+			/^VIRTME_STEAL_SUMMARY / { next }
+			report
+		' "$tmpdir/vm-$vm.log"
+	else
+		read -r summary_vm summary_throughput summary_steal <<<"$summary"
+		printf '%-8s %12s %16s\n' \
+			"$summary_vm" "$summary_throughput" "$summary_steal"
+		total_throughput=$((total_throughput + summary_throughput))
+	fi
+done
+
+if (( ! verbose )); then
+	printf '%-8s %12s %16s\n' -------- ------------ ----------------
+	printf '%-8s %12s %16s\n' TOTAL "$total_throughput" -
+fi
+
+exit "$status"
-- 
2.53.0

Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Shrikanth Hegde 1 month, 1 week ago
Hi Yury.

On 8/22/26 3:57 AM, Yury Norov wrote:
>> ========================================
>> Performance Numbers (powerpc, x86, s390)
>> ========================================
>>
>> PowerPC:
>> ===================
>> VM1: 60VP/30EC and VM2: 30VP/20EC
>> Shared physical CPU pool size: 50 Cores. Each core is SMT8.
>> (VP - Virtual Core, EC - Entitles Core) -  PowerVM terminologies of SPLPAR[2]
>>
>> Default parameter values: 1000ms, 200 low threshold, 500 high threshold
>> Both the VMs are running the same workload. Total throughput/time of VM1+VM2
>> is being mentioned in all cases.
>>
>> Hackbench
>>                baseline    steal_governor        steal_governor
>>                               disabled               enabled
>> ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
>>
>> 10 groups        5.20   |    5.40 (-3.85%)  |     4.65 (+10.58%)
>> 20 groups       11.39   |   12.01 (-5.44%)  |     7.09 (+37.75%)
>> 40 groups       20.32   |   19.80 (+2.56%)  |    11.31 (+44.34%)
>> 10 groups(-p)    2.37   |    2.26 (+4.64%)  |     2.06 (+13.08%)
>> 20 groups(-p)    3.34   |    3.28 (+1.80%)  |     3.20 (+4.19%)
>> 40 groups(-p)    4.46   |    4.83 (-8.30%)  |     4.26 (+4.48%)
>> Remarks: Net improvement with steal_governor specially high load points.
>>
>> schbench ( -L -n 0 -r 30 -s 0)
>>                baseline    steal_governor          steal_governor
>>                               disabled                 enabled
>> ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
>> -m 1 -t 128     2475162 |    2621246 (+5.90%)  |      2527299 (+2.11%)
>> -m 1 -t 256     1467350 |    1470032 (+0.18%)  |      1492372 (+1.71%)
>> -m 1 -t 512     1408813 |    1454687 (+3.26%)  |      1437605 (+2.04%)
>> Remarks: Effectively means no-improvements or regressions
>>
>> kernbench	baseline    steal_governor     steal_governor
>> (elapsed time)	               disabled            enabled
>> ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
>> -j nr_cpus	231      |      235 (-1.7%) |    199 (+14%)
>> Remarks: Net improvement in elapsed time.
>>
>> Daytrader - A real life work which is a proxy for trading based
>> on db2[3]
>>                baseline      steal_governor   steal_governor
>>                                disabled          enabled
>> ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
>> Load@30%	1x	|	0.96x	|	 1.53x			
>> Load@60%	1x	|	0.94x	|	 1.41x
>> Remarks: Good improvement seen at different load points.
>>
>> When there is no steal time (such as dedicated LPAR, or only VM2
>> is running) throughput was same with steal_governor enabled/disabled
>> which indicates minimal overhead of steal_governor.
>>
>> I have run v10 also on a smaller powerpc LPAR system and it shows
>> good improvements.
>>
>> =======================================================================
>>
>> Data from x86,s390 KVM which Ilya Leoshkevich carried out during OSPM26
>> time. *This was based on v2*. Idea is still the name, numbers are
>> expected to be better in v10 as some of the overhead has been removed.
>> Note: Other variations of the benchmark shows no observable
>> difference.
>>
>> x86:
>> ====
>> cascade-lake: 32 threads = 16 cores
>> Benchmark      #VMs    #CPUs/VM  ΔRPS     (%std)
>> ===============================================
>> hackbench         8          16  90.73% ± 9.97%
>> hackbench         4          24  52.67% ± 7.43%
>> hackbench         4          16  37.96% ± 11.19%
>> hackbench         4          32  37.82% ± 4.38%
>> hackbench        12           8  36.90% ± 4.74%
>> hackbench         8           8  35.30% ± 3.61%
>> pgbench          16           4  31.77% ± 2.44%
>> hackbench         2          24  25.85% ± 8.63%
>> hackbench        16           8  24.87% ± 3.46%
>> pgbench          16           8  21.83% ± 2.20%
>> pgbench          12           8  21.35% ± 2.15%
>> pgbench           8           8  18.46% ± 1.01%
>> hackbench         2          32  15.56% ± 4.53%
>> pgbench          12           4  14.28% ± 2.04%
>> hackbench        16           4  14.07% ± 2.90%
>> hackbench        12           4  9.60% ± 3.49%
>> [...]
>> pgbench           4           8  -1.16% ± 3.60%
>> hackbench         4           4  -1.80% ± 9.55%
>> sysbench         12           4  -2.19% ± 0.78%
>> pgbench           4          24  -2.43% ± 4.38%
>> pgbench           4          32  -3.21% ± 0.79%
>> sysbench         16           4  -3.22% ± 1.09%
>>
>> S390:
>> =====
>> z16: 16 threads = 8 cores (SMT-2)
>> Benchmark      #VMs    #CPUs/VM  ΔRPS    (std%)
>> ===============================================
>> pgbench           2           8  73.50% ± 35.91%
>> pgbench          16           4  61.30% ± 4.09%
>> hackbench        16           4  54.11% ± 4.38%
>> hackbench        12           4  36.34% ± 4.63%
>> pgbench          12           4  34.83% ± 2.57%
>> hackbench         8           4  29.75% ± 5.86%
>> hackbench         8           8  25.98% ± 5.09%
>> pgbench           2           4  23.31% ± 33.44%
>> pgbench           2          16  19.95% ± 17.12%
>> hackbench         4           8  19.43% ± 9.33%
>> pgbench           8           4  19.32% ± 4.50%
>> [...]
>> schbench          8           8  -0.79% ± 0.33%
>> sysbench          8           8  -0.81% ± 0.39%
>> hackbench         4          16  -1.11% ± 5.82%
>> sysbench          8           4  -1.62% ± 0.49%
>> sysbench         16           4  -2.70% ± 0.58%
>> schbench         16           4  -2.73% ± 0.91%
>> sysbench         12           4  -2.91% ± 0.61%
>> hackbench         2          24  -4.99% ± 3.31%
>>
>> Summary:
>> - Many improvement across archs specially with real life workloads.
>> - No major regressions observed.
>> - Overhead of steal_governor looks minimal when there is no steal time.
>> - Overhead when STEAL_GOVERNOR=n is negligible.
> 
> OK, I gave it some testing on my laptop.
> 

Hi Yury, thanks for trying.

> The results are pretty consistent: the steal ratio is converged to a
> number withing the threshold, but the overall performance is 3-5% worse
> comparing to baseline. I tried 2-5% and 1.5-15% boundaries.
> 
> It's 4 VMs, each running 8 vCPUs on 8 pCPU machine, the payload is
> running for 2 minutes.
> 
>    Steal governor off:
>    VM         THROUGHPUT       AVG STEAL%
>    -------- ------------ ----------------
>    0               13802            74.59
>    1               11929            76.35
>    2               13470            74.61
>    3               10259            78.67
>    -------- ------------ ----------------
>    TOTAL           49460                -
>    
>    Steal governor on:
>    VM         THROUGHPUT       AVG STEAL%
>    -------- ------------ ----------------
>    0               12921             5.72
>    1               12762             6.11
>    2               10160             5.99
>    3               12098             6.00
>    -------- ------------ ----------------
>    TOTAL           47941                -
> 
> 
> The test is attached below. The results are quite differ from the numbers
> above, so maybe I misconfigured something? I didn't use hackbench or
> similar benchmarks, just a basic math.
> 

If cputime is all that matters to workload, i.e. if workload is just math, or
without any locks or critical sections you may not see gains.
I have seen similar observation with stress-ng --cpu --metrics IIRC. Not much different.

The reason being, previously it was vCPU but thet vCPU is preempted 75% time.
But now, it is on vCPU but it is sharing that vCPU with 3 more tasks. Effectively
it is preempted 75% time still. Plus context switch overhead will show up.
That likley accounts for your 3-5% regression.

Can you give a try with hackbench, or pgbench, etc if possible?

> Shrikanth, can you please check my test and results? Is there something
> that I have missed?
> 
> I think this series should include some testing. The scripts below look
> bulky and they depend on virtme, but they allow to build the proper
> kernel and run tests with a single command.
> 

Ok, I will check the scripts you have attached. If they are just math, they may not be
the right benchmark to see gains for the reason explained above.


> Thanks,
> Yury
> 
>  From 27e548d65bbe0ec01308d3329017619e165d0488 Mon Sep 17 00:00:00 2001
> From: Yury Norov <ynorov@nvidia.com>
> Date: Wed, 19 Aug 2026 16:31:33 -0400
> Subject: [PATCH] steal governor: add testing harness based on vng
> 
> Signed-off-by: Yury Norov <ynorov@nvidia.com>
> ---
>   .../virtme-steal-payload.sh                   | 220 ++++++++++++
>   .../steal_governor_test/virtme-steal-time.sh  | 321 ++++++++++++++++++
>   2 files changed, 541 insertions(+)
>   create mode 100755 drivers/virt/steal_governor_test/virtme-steal-payload.sh
>   create mode 100755 drivers/virt/steal_governor_test/virtme-steal-time.sh
> 
> diff --git a/drivers/virt/steal_governor_test/virtme-steal-payload.sh b/drivers/virt/steal_governor_test/virtme-steal-payload.sh
> new file mode 100755
> index 000000000000..02a245a900b8
> --- /dev/null
> +++ b/drivers/virt/steal_governor_test/virtme-steal-payload.sh
> @@ -0,0 +1,220 @@
> +#!/bin/sh
> +# Guest-side measurable workload for the steal governor test.
> +
> +set -eu
> +
> +case $# in
> +2|3|5) ;;
> +*)
> +	echo "usage: ${0##*/} VM_ID DURATION_SECONDS [DRIVER [LOW_THRESHOLD HIGH_THRESHOLD]]" >&2
> +	exit 2
> +	;;
> +esac
> +
> +vm_id=$1
> +duration=$2
> +driver=${3-}
> +low_threshold=${4-}
> +high_threshold=${5-}
> +
> +verify_parameter()
> +{
> +	parameter=$1
> +	expected=$2
> +	parameter_file=/sys/module/$driver_sysfs/parameters/$parameter
> +
> +	if [ ! -r "$parameter_file" ]; then
> +		echo "error: driver $driver has no readable $parameter parameter" >&2
> +		exit 1
> +	fi
> +	actual=$(cat "$parameter_file")
> +	if [ "$actual" != "$expected" ]; then
> +		echo "error: driver $driver $parameter is $actual, expected $expected" >&2
> +		exit 1
> +	fi
> +}
> +
> +case $vm_id in
> +*[!0-9]*|'')
> +	echo "error: VM_ID must be a non-negative integer" >&2
> +	exit 2
> +	;;
> +esac
> +case $duration in
> +*[!0-9]*|'')
> +	echo "error: DURATION_SECONDS must be a positive integer" >&2
> +	exit 2
> +	;;
> +esac
> +if [ "$duration" -eq 0 ]; then
> +	echo "error: DURATION_SECONDS must be greater than zero" >&2
> +	exit 2
> +fi
> +if [ -n "$driver" ]; then
> +	case $driver in
> +	*[!A-Za-z0-9_.-]*)
> +		echo "error: invalid driver module name: $driver" >&2
> +		exit 2
> +		;;
> +	esac
> +	if [ -n "$low_threshold" ]; then
> +		case $low_threshold:$high_threshold in
> +		*[!0-9:]*|:*|*:)
> +			echo "error: thresholds must be non-negative integers in percent * 100" >&2
> +			exit 2
> +			;;
> +		esac
> +		if [ "$low_threshold" -ge "$high_threshold" ] ||
> +		   [ "$high_threshold" -ge 10000 ]; then
> +			echo "error: thresholds must satisfy 0 <= LOW < HIGH < 10000" >&2
> +			exit 2
> +		fi
> +	fi
> +	if ! command -v modprobe >/dev/null; then
> +		echo "error: modprobe is required to load driver $driver" >&2
> +		exit 1
> +	fi
> +	driver_sysfs=$(printf '%s\n' "$driver" | tr '-' '_')
> +	if [ ! -d "/sys/module/$driver_sysfs" ]; then
> +		if [ -n "$low_threshold" ]; then
> +			modprobe "$driver" \
> +				low_threshold="$low_threshold" \
> +				high_threshold="$high_threshold"
> +		else
> +			modprobe "$driver"
> +		fi
> +	fi
> +	if [ ! -d "/sys/module/$driver_sysfs" ]; then
> +		echo "error: driver $driver has no /sys/module/$driver_sysfs entry after modprobe" >&2
> +		exit 1
> +	fi
> +	if [ -n "$low_threshold" ]; then
> +		verify_parameter low_threshold "$low_threshold"
> +		verify_parameter high_threshold "$high_threshold"
> +	fi
> +fi
> +
> +tmpdir=$(mktemp -d "${TMPDIR:-/tmp}/virtme-steal-payload.XXXXXXXX")
> +before=$tmpdir/before
> +after=$tmpdir/after
> +pids=
> +
> +stop_workers()
> +{
> +	for pid in $pids; do
> +		kill "$pid" 2>/dev/null || :
> +	done
> +	for pid in $pids; do
> +		wait "$pid" 2>/dev/null || :
> +	done
> +	pids=
> +}
> +
> +cleanup()
> +{
> +	stop_workers
> +	rm -rf -- "$tmpdir"
> +}
> +trap cleanup EXIT
> +trap 'exit 130' INT
> +trap 'exit 143' TERM
> +
> +awk '$1 ~ /^cpu[0-9]+$/ { print $1, $9 }' /proc/stat >"$before"
> +start_time=$(awk '{ print $1 }' /proc/uptime)
> +deadline=$(awk -v start="$start_time" -v seconds="$duration" \
> +	'BEGIN { printf "%.2f", start + seconds }')
> +workers=$(awk '$1 ~ /^cpu[0-9]+$/ { n++ } END { print n }' /proc/stat)
> +
> +i=0
> +while [ "$i" -lt "$workers" ]; do
> +	result=$tmpdir/work-$i
> +	awk -v deadline="$deadline" -v result="$result" '
> +		function uptime(    line, fields) {
> +			getline line < "/proc/uptime"
> +			close("/proc/uptime")
> +			split(line, fields)
> +			return fields[1]
> +		}
> +		BEGIN {
> +			batch = 1000
> +			units = 0
> +			value = 1
> +			while (uptime() < deadline) {
> +				for (iteration = 0; iteration < batch; iteration++)
> +					value = (value * 1103515245 + 12345) % 2147483647
> +				units++
> +			}
> +			printf "%.0f %.0f\n", units, units * batch > result
> +		}
> +	' </dev/null &
> +	pids="$pids $!"
> +	i=$((i + 1))
> +done
> +
> +worker_failed=0
> +for pid in $pids; do
> +	wait "$pid" || worker_failed=1
> +done
> +pids=
> +if [ "$worker_failed" -ne 0 ]; then
> +	echo "error: one or more workload processes failed" >&2
> +	exit 1
> +fi
> +
> +awk '$1 ~ /^cpu[0-9]+$/ { print $1, $9 }' /proc/stat >"$after"
> +end_time=$(awk '{ print $1 }' /proc/uptime)
> +elapsed=$(awk -v start="$start_time" -v end="$end_time" \
> +	'BEGIN { printf "%.2f", end - start }')
> +clk_tck=$(getconf CLK_TCK 2>/dev/null || echo 100)
> +
> +echo VIRTME_STEAL_REPORT_BEGIN
> +driver_status=none
> +[ -n "$driver" ] && driver_status=$driver:loaded
> +[ -n "$low_threshold" ] && \
> +	driver_status="$driver_status,thresholds=$low_threshold..$high_threshold"
> +echo "VM $vm_id  kernel=$(uname -r)  vCPUs=$workers  sample=${elapsed}s  driver=$driver_status"
> +echo "Workload (unpinned workers)"
> +printf '%-10s %12s %16s %16s\n' \
> +	WORKER WORK_UNITS ITERATIONS ITERATIONS/s
> +printf '%-10s %12s %16s %16s\n' \
> +	---------- ------------ ---------------- ----------------
> +total_units=0
> +total_iterations=0
> +i=0
> +while [ "$i" -lt "$workers" ]; do
> +	read -r units iterations <"$tmpdir/work-$i"
> +	rate=$(awk -v iterations="$iterations" -v elapsed="$elapsed" \
> +		'BEGIN { printf "%.0f", iterations / elapsed }')
> +	printf '%-10s %12s %16s %16s\n' \
> +		"worker$i" "$units" "$iterations" "$rate"
> +	total_units=$((total_units + units))
> +	total_iterations=$((total_iterations + iterations))
> +	i=$((i + 1))
> +done
> +total_rate=$(awk -v iterations="$total_iterations" -v elapsed="$elapsed" \
> +	'BEGIN { printf "%.0f", iterations / elapsed }')
> +throughput=$(awk -v units="$total_units" -v elapsed="$elapsed" \
> +	'BEGIN { printf "%.0f", units / elapsed }')
> +printf '%-10s %12s %16s %16s\n' \
> +	TOTAL "$total_units" "$total_iterations" "$total_rate"
> +
> +echo "Steal time by guest CPU"
> +printf '%-8s %12s %12s %10s\n' CPU STEAL_TICKS STEAL_s STEAL_%
> +printf '%-8s %12s %12s %10s\n' -------- ------------ ------------ ----------
> +average_file=$tmpdir/average-steal
> +awk -v hz="$clk_tck" -v elapsed="$elapsed" -v cpus="$workers" \
> +	-v average_file="$average_file" '
> +	NR == FNR { before[$1] = $2; next }
> +	{
> +		delta = $2 - before[$1]
> +		total += delta
> +		printf "%-8s %12d %12.2f %10.2f\n", $1, delta,
> +		       delta / hz, 100 * delta / hz / elapsed
> +	}
> +	END {
> +		printf "%.2f\n", 100 * total / hz / elapsed / cpus > average_file
> +	}
> +' "$before" "$after"
> +average_steal=$(cat "$average_file")
> +echo "VIRTME_STEAL_SUMMARY $vm_id $throughput $average_steal"
> +echo VIRTME_STEAL_REPORT_END
> diff --git a/drivers/virt/steal_governor_test/virtme-steal-time.sh b/drivers/virt/steal_governor_test/virtme-steal-time.sh
> new file mode 100755
> index 000000000000..a41d362c7888
> --- /dev/null
> +++ b/drivers/virt/steal_governor_test/virtme-steal-time.sh
> @@ -0,0 +1,321 @@
> +#!/usr/bin/env bash
> +# Build a minimal KVM guest kernel and measure per-vCPU steal time under load.
> +
> +set -euo pipefail
> +
> +script_dir=$(CDPATH= cd -- "$(dirname -- "$0")" && pwd)
> +kernel_dir=$(CDPATH= cd -- "$script_dir/../../.." && pwd)
> +vng=${VNG:-vng}
> +build_dir=${BUILD_DIR:-"$kernel_dir/.virtme-steal"}
> +payload=${PAYLOAD:-"$script_dir/virtme-steal-payload.sh"}
> +vm_count=4
> +vcpu_count=$(nproc)
> +duration=20
> +memory=512M
> +driver=
> +low_threshold=
> +high_threshold=
> +verbose=0
> +declare -a config_items=()
> +
> +usage()
> +{
> +	cat <<EOF
> +Usage: ${0##*/} [options] [O=DIR]
> +
> +Build a virtme-ng minimal kernel with paravirtual steal-time accounting,
> +start multiple CPU-bound VMs, and report the steal time of every guest CPU.
> +
> +Options:
> +  -n VMS       number of VMs to run (default: $vm_count)
> +  -p VCPUS     vCPUs per VM (default: $vcpu_count)
> +  -d SECONDS   workload duration (default: $duration)
> +  -m MEMORY    memory per VM (default: $memory)
> +  --configitem CONFIG[=VALUE]
> +                enable or set a kernel config option (repeatable)
> +  --driver MODULE
> +                load and verify this module before running the payload
> +  -lo VALUE     low steal threshold in percent * 100 (for example, 150 = 1.5%)
> +  -hi VALUE     high steal threshold in percent * 100 (for example, 550 = 5.5%)
> +  -O DIR       kernel build directory (default: $build_dir)
> +  -v, --verbose show per-worker and per-CPU result tables
> +  -h            show this help
> +
> +Environment equivalents: VNG, BUILD_DIR, and PAYLOAD. A custom payload is
> +called inside each VM as: PAYLOAD VM_ID DURATION_SECONDS
> +[DRIVER [LOW_THRESHOLD HIGH_THRESHOLD]].
> +
> +Local virtme-ng checkout example:
> +  VNG=../virtme-ng/vng ${0##*/}
> +
> +The positional O=DIR form is equivalent to -O DIR, for example:
> +  ${0##*/} -n 4 O=../build-linux-virtme-steal
> +
> +Config examples:
> +  ${0##*/} --configitem CONFIG_SCHEDSTATS --configitem CONFIG_HZ_1000=y
> +
> +Threshold example:
> +  ${0##*/} --driver steal_governor -lo 150 -hi 550
> +EOF
> +}
> +
> +while (($#)); do
> +	case $1 in
> +	-n|-p|-d|-m|-O|-lo|-hi|--configitem|--driver)
> +		if (($# < 2)); then
> +			echo "error: $1 requires an argument" >&2
> +			exit 2
> +		fi
> +		case $1 in
> +		-n) vm_count=$2 ;;
> +		-p) vcpu_count=$2 ;;
> +		-d) duration=$2 ;;
> +		-m) memory=$2 ;;
> +		-O) build_dir=$2 ;;
> +		-lo) low_threshold=$2 ;;
> +		-hi) high_threshold=$2 ;;
> +		--configitem) config_items+=("$2") ;;
> +		--driver) driver=$2 ;;
> +		esac
> +		shift 2
> +		;;
> +	--configitem=?*)
> +		config_items+=("${1#*=}")
> +		shift
> +		;;
> +	--driver=?*)
> +		driver=${1#*=}
> +		shift
> +		;;
> +	-v|--verbose)
> +		verbose=1
> +		shift
> +		;;
> +	-h|--help)
> +		usage
> +		exit 0
> +		;;
> +	O=?*)
> +		build_dir=${1#O=}
> +		shift
> +		;;
> +	O=|--configitem=|--driver=)
> +		echo "error: ${1%%=*}= requires a non-empty argument" >&2
> +		exit 2
> +		;;
> +	*)
> +		echo "error: unexpected argument: $1" >&2
> +		usage >&2
> +		exit 2
> +		;;
> +	esac
> +done
> +
> +require_positive_integer()
> +{
> +	local name=$1 value=$2
> +
> +	if [[ ! $value =~ ^[1-9][0-9]*$ ]]; then
> +		echo "error: $name must be a positive integer (got '$value')" >&2
> +		exit 2
> +	fi
> +}
> +
> +require_positive_integer "VM count" "$vm_count"
> +require_positive_integer "vCPU count" "$vcpu_count"
> +require_positive_integer "duration" "$duration"
> +
> +if [[ -n $low_threshold || -n $high_threshold ]]; then
> +	if [[ -z $low_threshold || -z $high_threshold ]]; then
> +		echo "error: -lo and -hi must be specified together" >&2
> +		exit 2
> +	fi
> +	if [[ ! $low_threshold =~ ^[0-9]+$ || ! $high_threshold =~ ^[0-9]+$ ]]; then
> +		echo "error: -lo and -hi must be non-negative integers in percent * 100" >&2
> +		exit 2
> +	fi
> +	low_threshold=$((10#$low_threshold))
> +	high_threshold=$((10#$high_threshold))
> +	if ((low_threshold >= high_threshold)); then
> +		echo "error: -lo must be less than -hi" >&2
> +		exit 2
> +	fi
> +	if ((high_threshold >= 10000)); then
> +		echo "error: -hi must be less than 10000 (100%)" >&2
> +		exit 2
> +	fi
> +	if [[ -z $driver ]]; then
> +		echo "error: -lo and -hi require --driver MODULE" >&2
> +		exit 2
> +	fi
> +fi
> +
> +if [[ -n $driver && ! $driver =~ ^[A-Za-z0-9_.-]+$ ]]; then
> +	echo "error: invalid driver module name: $driver" >&2
> +	exit 2
> +fi
> +
> +for index in "${!config_items[@]}"; do
> +	if [[ ! ${config_items[index]} =~ ^CONFIG_[A-Z0-9_]+(=.*)?$ ]]; then
> +		echo "error: invalid kernel config item: ${config_items[index]}" >&2
> +		exit 2
> +	fi
> +	if [[ ${config_items[index]} != *=* ]]; then
> +		config_items[index]="${config_items[index]}=y"
> +	fi
> +done
> +
> +if [[ ! -x $payload ]]; then
> +	echo "error: guest payload is not executable: $payload" >&2
> +	exit 1
> +fi
> +if [[ ! -r /dev/kvm || ! -w /dev/kvm ]]; then
> +	echo "error: /dev/kvm is not accessible; KVM is required for steal-time accounting" >&2
> +	exit 1
> +fi
> +
> +case $build_dir in
> +/*) ;;
> +*) build_dir=$PWD/$build_dir ;;
> +esac
> +
> +mkdir -p "$build_dir"
> +cd "$kernel_dir"
> +
> +echo "==> Configuring paravirtual kernel in $build_dir"
> +if [[ ! -f $build_dir/.config ]]; then
> +	config_args=()
> +	for config_item in "${config_items[@]}"; do
> +		config_args+=(--configitem "$config_item")
> +	done
> +	"$vng" --kconfig "${config_args[@]}" \
> +		--configitem CONFIG_HYPERVISOR_GUEST=y \
> +		--configitem CONFIG_PARAVIRT=y \
> +		--configitem CONFIG_KVM_GUEST=y \
> +		--configitem CONFIG_PARAVIRT_TIME_ACCOUNTING=y \
> +		-- "O=$build_dir"
> +else
> +	# Preserve an existing local minimal config and add the options needed by
> +	# this test. olddefconfig resolves their dependencies for the current tree.
> +	config_args=(--file "$build_dir/.config")
> +	for config_item in "${config_items[@]}"; do
> +		config_name=${config_item%%=*}
> +		config_value=${config_item#*=}
> +		config_args+=(--set-val "$config_name" "$config_value")
> +	done
> +	"$kernel_dir/scripts/config" "${config_args[@]}" \
> +		-e HYPERVISOR_GUEST \
> +		-e PARAVIRT \
> +		-e KVM_GUEST \
> +		-e PARAVIRT_TIME_ACCOUNTING
> +	make -s O="$build_dir" olddefconfig
> +fi
> +
> +echo "==> Building kernel"
> +declare -a module_args=()
> +if [[ -z $driver ]]; then
> +	module_args+=(--skip-modules)
> +fi
> +"$vng" --build "${module_args[@]}" -- "O=$build_dir"
> +
> +for option in HYPERVISOR_GUEST PARAVIRT KVM_GUEST PARAVIRT_TIME_ACCOUNTING; do
> +	if ! grep -qx "CONFIG_${option}=y" "$build_dir/.config"; then
> +		echo "error: CONFIG_${option}=y is required but is absent after configuration" >&2
> +		exit 1
> +	fi
> +done
> +
> +tmpdir=$(mktemp -d "${TMPDIR:-/tmp}/virtme-steal.XXXXXXXX")
> +declare -a vm_pids=()
> +
> +cleanup()
> +{
> +	local pid
> +
> +	for pid in "${vm_pids[@]}"; do
> +		kill "$pid" 2>/dev/null || true
> +	done
> +	wait 2>/dev/null || true
> +	rm -rf -- "$tmpdir"
> +}
> +trap cleanup EXIT
> +trap 'exit 130' INT
> +trap 'exit 143' TERM
> +
> +echo "==> Starting $vm_count VMs ($vcpu_count vCPUs each)"
> +for ((vm = 0; vm < vm_count; vm++)); do
> +	printf -v guest_script '%q %q %q' "$payload" "$vm" "$duration"
> +	if [[ -n $driver ]]; then
> +		printf -v driver_arg ' %q' "$driver"
> +		guest_script+=$driver_arg
> +		if [[ -n $low_threshold ]]; then
> +			printf -v threshold_args ' %q %q' \
> +				"$low_threshold" "$high_threshold"
> +			guest_script+=$threshold_args
> +		fi
> +	fi
> +
> +	log=$tmpdir/vm-$vm.log
> +	"$vng" \
> +		"${module_args[@]}" \
> +		--name "steal-vm-$vm" \
> +		--cpus "$vcpu_count" \
> +		--memory "$memory" \
> +		--exec "$guest_script" \
> +		-- "O=$build_dir" >"$log" 2>&1 &
> +	vm_pids+=("$!")
> +done
> +
> +status=0
> +total_throughput=0
> +if (( ! verbose )); then
> +	printf '%-8s %12s %16s\n' VM THROUGHPUT 'AVG STEAL%'
> +	printf '%-8s %12s %16s\n' -------- ------------ ----------------
> +fi
> +for ((vm = 0; vm < vm_count; vm++)); do
> +	if ! wait "${vm_pids[vm]}"; then
> +		echo "error: VM $vm failed; its complete log follows" >&2
> +		status=1
> +		echo "--- VM $vm log ---"
> +		cat "$tmpdir/vm-$vm.log"
> +		continue
> +	fi
> +	if ! grep -qx VIRTME_STEAL_REPORT_BEGIN "$tmpdir/vm-$vm.log" ||
> +	   ! grep -qx VIRTME_STEAL_REPORT_END "$tmpdir/vm-$vm.log"; then
> +		echo "error: VM $vm exited without a steal-time report; its complete log follows" >&2
> +		cat "$tmpdir/vm-$vm.log"
> +		status=1
> +		continue
> +	fi
> +
> +	summary=$(awk '/^VIRTME_STEAL_SUMMARY / { print $2, $3, $4 }' \
> +		"$tmpdir/vm-$vm.log")
> +	if [[ -z $summary ]]; then
> +		echo "error: VM $vm exited without a summary; its complete log follows" >&2
> +		cat "$tmpdir/vm-$vm.log"
> +		status=1
> +		continue
> +	fi
> +
> +	if (( verbose )); then
> +		echo "--- VM $vm results ---"
> +		awk '
> +			/^VIRTME_STEAL_REPORT_BEGIN$/ { report = 1; next }
> +			/^VIRTME_STEAL_REPORT_END$/ { report = 0 }
> +			/^VIRTME_STEAL_SUMMARY / { next }
> +			report
> +		' "$tmpdir/vm-$vm.log"
> +	else
> +		read -r summary_vm summary_throughput summary_steal <<<"$summary"
> +		printf '%-8s %12s %16s\n' \
> +			"$summary_vm" "$summary_throughput" "$summary_steal"
> +		total_throughput=$((total_throughput + summary_throughput))
> +	fi
> +done
> +
> +if (( ! verbose )); then
> +	printf '%-8s %12s %16s\n' -------- ------------ ----------------
> +	printf '%-8s %12s %16s\n' TOTAL "$total_throughput" -
> +fi
> +
> +exit "$status"

[PATCH] Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Ionut Nechita (Sunlight Linux) 1 month, 2 weeks ago
On Wed, Aug 12, 2026 at 11:10:21AM +0530, Shrikanth Hegde wrote:
> This patch series represents the result of multiple iterations,
> redesigns and community feedback.

Nice work, and thanks for the very readable cover letter.

I looked at this from the KVM and Xen guest angle rather than from
PowerVM, since that is what I run.  Three observations below.  All of
them are from code inspection only -- I have not measured any of this,
so please treat the numbers as arithmetic rather than as results.

Code references are against next-20260812, which already carries your
base commit f2c2ba7219e5, so the series applies there directly.


1) Default thresholds are unreachable on common QEMU command lines
==================================================================

get_system_cpus() returns num_possible_cpus(), and steal_governor_loop()
divides by it:

	delta_ns = max_t(u64, div_u64(delta_ns * get_system_cpus(), 10000), 1);
	steal_ratio = div64_u64(delta_steal, delta_ns);

On x86 the possible map is sized by topology_init_possible_cpus()
(arch/x86/kernel/cpu/topology.c) from assigned + disabled CPUs, where
"disabled" is incremented by topo_register_apic() for every APIC that is
registered but not present.

QEMU emits exactly such MADT entries for the range [smp, maxcpus), and
acpi_is_processor_usable() in arch/x86/kernel/acpi/boot.c documents that
this is deliberate:

	/*
	 * QEMU expects legacy "Enabled=0" LAPIC entries to be counted as
	 * usable in order to support CPU hotplug in guests.
	 */

So those vCPUs are registered as usable, land in nr_disabled_cpus, and
end up in the possible map.  A guest started with

	-smp 4,maxcpus=32

has num_possible_cpus() == 32 while only 4 vCPUs ever run.  The steal
ratio is then diluted 8x, and the default thresholds of 5% / 2% become
40% / 16% of the steal that is actually observable.  The governor never
leaves the "do nothing" window, and the only symptom is that nothing
happens.

Xen PV guests are affected in the same way, and often more strongly,
because the possible map there tends to be sized for vCPU hotplug.

I understand from the v9 changelog why possible CPUs were chosen -- it
keeps the accumulated steal monotonic across hotplug, which is a real
property worth having.  The documentation does describe the effect and
gives a worked example for recomputing the thresholds by hand.  But for
a mechanism whose whole premise is that every VM on the host opts in
with the same policy, requiring each operator to first derive their own
thresholds seems likely to translate into low real-world adoption.

Some options, roughly in increasing order of intrusiveness:

  - emit a pr_info() (or pr_warn()) at module init when
    num_possible_cpus() significantly exceeds num_online_cpus(), naming
    the ratio and the effective thresholds.  Cheap, and turns a silent
    no-op into something diagnosable.

  - scale the thresholds by num_possible_cpus() / num_online_cpus() at
    init, so the documented defaults keep their intended meaning.

  - keep summing steal over the possible mask, as today, but divide by
    the online count, and handle the hotplug discontinuity by resetting
    the sg_ctx.steal baseline from a hotplug notifier.

I do not have a strong preference among these, and the first one alone
would already be a large improvement.


2) Nothing stops the driver from folding Xen dom0
=================================================

dom0 accounts steal time like any other domain -- xen_time_setup_guest()
in arch/x86/xen/time.c wires up pv_steal_clock unconditionally, with no
feature negotiation and no privileged-domain exemption.

So loading steal_governor in dom0 makes it shrink its own preferred mask
under contention.  That is precisely when the blkback and netback
threads serving every other guest need CPU, and the driver has no notion
that this domain is different from the ones it is trying to be polite
towards.  The effect would be host-wide, not confined to the domain that
loaded the module.

Given that Kconfig carries "default m", the module is built on any
distro kernel with PARAVIRT=y, which includes dom0 kernels.  It is one
modprobe away from being loaded there, quite plausibly by someone who
read the documentation's advice to enable it uniformly across all VMs.

A xen_initial_domain() check that refuses to load, or at minimum a loud
warning, seems worth having.  More generally it may be worth stating in
the documentation that the driver is meant for guests only -- the same
argument applies to a KVM host that is itself running nested guests.

Juergen and the virtualization list are already on Cc -- I would value
their view on this one in particular.


3) Core granularity degenerates to single vCPUs on KVM and Xen
==============================================================

decrease_preferred_cpus() and increase_preferred_cpus() step by
topology_sibling_cpumask(), which matches PowerVM, where the hypervisor
schedules whole cores.

On Xen PV that mask is always the CPU itself, by design.  The header
comment of arch/x86/xen/smp_pv.c is explicit about both the behaviour
and the reason for it:

	/*
	 * Because virtual CPUs can be scheduled onto any real CPU, there's
	 * no useful topology information for the kernel to make use of.  As
	 * a result, all CPUs are treated as if they're single-core and
	 * single-threaded.
	 */

A plain "-smp N" under QEMU gives one thread per core and one core per
socket, with the same result.  So on both hypervisors the step size is
one vCPU, and convergence to a folded state takes proportionally longer
than the PowerVM numbers would suggest.

I do not think this is a correctness problem -- per-vCPU is arguably the
right granularity there, for exactly the reason the Xen comment gives.
The case that does not work as intended is the opposite one: a guest
given a synthetic topology such as

	-smp 16,sockets=1,cores=8,threads=2

will fold what it believes is a core, but those two vCPU threads need
not be co-located on any host core, so nothing in particular is freed.

So mainly a documentation request: a sentence in
Documentation/driver-api/steal-governor.rst noting that the core-level
step assumes the guest topology reflects host scheduling granularity,
and that on KVM and Xen it commonly does not.  It would also explain to
readers why convergence looks slower there than in your PowerVM numbers.


--

I would be happy to run this on x86 KVM and on Xen PV guests if that is
useful -- the series has no Xen coverage that I can see, and points 1
and 3 above would both show up there first.

Thanks,
Ionut
Re: [PATCH] Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Yury Norov 1 month, 2 weeks ago
Hi Ionut,

Thanks for the details. It was indeed an informative reading. Let me
add a couple notes below.

Thanks,
Yury

On Wed, Aug 12, 2026 at 10:45:56PM +0300, Ionut Nechita (Sunlight Linux) wrote:
> On Wed, Aug 12, 2026 at 11:10:21AM +0530, Shrikanth Hegde wrote:
> > This patch series represents the result of multiple iterations,
> > redesigns and community feedback.
> 
> Nice work, and thanks for the very readable cover letter.
> 
> I looked at this from the KVM and Xen guest angle rather than from
> PowerVM, since that is what I run.  Three observations below.  All of
> them are from code inspection only -- I have not measured any of this,
> so please treat the numbers as arithmetic rather than as results.
> 
> Code references are against next-20260812, which already carries your
> base commit f2c2ba7219e5, so the series applies there directly.
> 
> 
> 1) Default thresholds are unreachable on common QEMU command lines
> ==================================================================
> 
> get_system_cpus() returns num_possible_cpus(), and steal_governor_loop()
> divides by it:
> 
> 	delta_ns = max_t(u64, div_u64(delta_ns * get_system_cpus(), 10000), 1);
> 	steal_ratio = div64_u64(delta_steal, delta_ns);
> 
> On x86 the possible map is sized by topology_init_possible_cpus()
> (arch/x86/kernel/cpu/topology.c) from assigned + disabled CPUs, where
> "disabled" is incremented by topo_register_apic() for every APIC that is
> registered but not present.
> 
> QEMU emits exactly such MADT entries for the range [smp, maxcpus), and
> acpi_is_processor_usable() in arch/x86/kernel/acpi/boot.c documents that
> this is deliberate:
> 
> 	/*
> 	 * QEMU expects legacy "Enabled=0" LAPIC entries to be counted as
> 	 * usable in order to support CPU hotplug in guests.
> 	 */
> 
> So those vCPUs are registered as usable, land in nr_disabled_cpus, and
> end up in the possible map.  A guest started with
> 
> 	-smp 4,maxcpus=32
> 
> has num_possible_cpus() == 32 while only 4 vCPUs ever run.  The steal
> ratio is then diluted 8x, and the default thresholds of 5% / 2% become
> 40% / 16% of the steal that is actually observable.  The governor never
> leaves the "do nothing" window, and the only symptom is that nothing
> happens.
> 
> Xen PV guests are affected in the same way, and often more strongly,
> because the possible map there tends to be sized for vCPU hotplug.
> 
> I understand from the v9 changelog why possible CPUs were chosen -- it
> keeps the accumulated steal monotonic across hotplug, which is a real
> property worth having.  The documentation does describe the effect and
> gives a worked example for recomputing the thresholds by hand.  But for
> a mechanism whose whole premise is that every VM on the host opts in
> with the same policy, requiring each operator to first derive their own
> thresholds seems likely to translate into low real-world adoption.
> 
> Some options, roughly in increasing order of intrusiveness:
> 
>   - emit a pr_info() (or pr_warn()) at module init when
>     num_possible_cpus() significantly exceeds num_online_cpus(), naming
>     the ratio and the effective thresholds.  Cheap, and turns a silent
>     no-op into something diagnosable.
> 
>   - scale the thresholds by num_possible_cpus() / num_online_cpus() at
>     init, so the documented defaults keep their intended meaning.
> 
>   - keep summing steal over the possible mask, as today, but divide by
>     the online count, and handle the hotplug discontinuity by resetting
>     the sg_ctx.steal baseline from a hotplug notifier.
> 
> I do not have a strong preference among these, and the first one alone
> would already be a large improvement.

I asked the same question on v9 iteration, but the problem is still
there. I don't think that VM with a big number of offlined CPUs is
something that needs to bring admin's attention with pr_info(). The
solution should be found withing the driver.

I thought about saving the steal time for each CPU in an array, so on
the next iteration we count steal time from the same subset of CPUs,
regardless of hot plugging. 

Your approach with num_possible_cpus() / num_online_cpus() scale looks
nicer, if it works. Definitely need a testing for both. Resetting the
steal baseline on each hotplug is another option, but to me the least
attractive.

This problem must be resolved this way or another before we merge the
driver.

> 2) Nothing stops the driver from folding Xen dom0
> =================================================
> 
> dom0 accounts steal time like any other domain -- xen_time_setup_guest()
> in arch/x86/xen/time.c wires up pv_steal_clock unconditionally, with no
> feature negotiation and no privileged-domain exemption.
> 
> So loading steal_governor in dom0 makes it shrink its own preferred mask
> under contention.  That is precisely when the blkback and netback
> threads serving every other guest need CPU, and the driver has no notion
> that this domain is different from the ones it is trying to be polite
> towards.  The effect would be host-wide, not confined to the domain that
> loaded the module.
> 
> Given that Kconfig carries "default m", the module is built on any
> distro kernel with PARAVIRT=y, which includes dom0 kernels.  It is one
> modprobe away from being loaded there, quite plausibly by someone who
> read the documentation's advice to enable it uniformly across all VMs.
> 
> A xen_initial_domain() check that refuses to load, or at minimum a loud
> warning, seems worth having.  More generally it may be worth stating in
> the documentation that the driver is meant for guests only -- the same
> argument applies to a KVM host that is itself running nested guests.
> 
> Juergen and the virtualization list are already on Cc -- I would value
> their view on this one in particular.
> 
> 
> 3) Core granularity degenerates to single vCPUs on KVM and Xen
> ==============================================================
> 
> decrease_preferred_cpus() and increase_preferred_cpus() step by
> topology_sibling_cpumask(), which matches PowerVM, where the hypervisor
> schedules whole cores.
> 
> On Xen PV that mask is always the CPU itself, by design.  The header
> comment of arch/x86/xen/smp_pv.c is explicit about both the behaviour
> and the reason for it:
> 
> 	/*
> 	 * Because virtual CPUs can be scheduled onto any real CPU, there's
> 	 * no useful topology information for the kernel to make use of.  As
> 	 * a result, all CPUs are treated as if they're single-core and
> 	 * single-threaded.
> 	 */
> 
> A plain "-smp N" under QEMU gives one thread per core and one core per
> socket, with the same result.  So on both hypervisors the step size is
> one vCPU, and convergence to a folded state takes proportionally longer
> than the PowerVM numbers would suggest.
> 
> I do not think this is a correctness problem -- per-vCPU is arguably the
> right granularity there, for exactly the reason the Xen comment gives.
> The case that does not work as intended is the opposite one: a guest
> given a synthetic topology such as
> 
> 	-smp 16,sockets=1,cores=8,threads=2
> 
> will fold what it believes is a core, but those two vCPU threads need
> not be co-located on any host core, so nothing in particular is freed.

On 2 and 3.

If the governor will become a practical tool for admins, I believe
there will be the more corner cases, the more adoption the driver
gets. 

That's why I suggested the modular architecture for the governor. Now
we have 2 options:

1. Keep the driver as a single module, and add those checks for DOM 0
   and per-core vs per-thread behavior.
2. Rename this driver to powervm_steal_governor, and let KVM and XEN
   engineers implement their own drivers.

The options are not mutually exclusive. The earlier versions of the
driver had a 'core' part and hooks for different arches.

At this point, I believe, keeping the driver simple is the most
important goal for making it upstreamable. When it's merged, it
will be much easier to add more functionality and tuning for
different VMs and architectures.
 
> So mainly a documentation request: a sentence in
> Documentation/driver-api/steal-governor.rst noting that the core-level
> step assumes the guest topology reflects host scheduling granularity,
> and that on KVM and Xen it commonly does not.  It would also explain to
> readers why convergence looks slower there than in your PowerVM numbers.

I agree, this should be documented.

> --
> 
> I would be happy to run this on x86 KVM and on Xen PV guests if that is
> useful -- the series has no Xen coverage that I can see, and points 1
> and 3 above would both show up there first.

I didn't test the patches myself yet, mostly because configuring VMs
is a separate topic. If you share a script that configures an
environment with paravirtual VMs that run some payload that overcommit
the vCPUs, it would be a valuable addition to the series.

Andrea Righi with his virtme is already in CC. It's a very friendly
and highly configurable wrapper around qemu that allows to compile a
and run kernel with a couple commands. Andrea, can the virtme prepare
such an environment? For testing we need:

 - several paravirtual VMs, so that the number of vCPUs exceeds the
   number of CPUs, and
 - some payload running on them to overcommit the vCPUs, triggering
   the steal time counter.

Thanks,
Yury
Re: [PATCH] Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Shrikanth Hegde 1 month, 2 weeks ago
Hi Ionut, Yury

On 8/13/26 5:43 AM, Yury Norov wrote:
> Hi Ionut,
> 
> Thanks for the details. It was indeed an informative reading. Let me
> add a couple notes below.
> 
> Thanks,
> Yury
> 
> On Wed, Aug 12, 2026 at 10:45:56PM +0300, Ionut Nechita (Sunlight Linux) wrote:
>> On Wed, Aug 12, 2026 at 11:10:21AM +0530, Shrikanth Hegde wrote:
>>> This patch series represents the result of multiple iterations,
>>> redesigns and community feedback.
>>
>> Nice work, and thanks for the very readable cover letter.

Thanks for reading and going through the series. Really appreciated.

>>
>> I looked at this from the KVM and Xen guest angle rather than from
>> PowerVM, since that is what I run.  Three observations below.  All of
>> them are from code inspection only -- I have not measured any of this,
>> so please treat the numbers as arithmetic rather than as results.
>>
>> Code references are against next-20260812, which already carries your
>> base commit f2c2ba7219e5, so the series applies there directly.
>>
>>
>> 1) Default thresholds are unreachable on common QEMU command lines
>> ==================================================================
>>
>> get_system_cpus() returns num_possible_cpus(), and steal_governor_loop()
>> divides by it:
>>
>> 	delta_ns = max_t(u64, div_u64(delta_ns * get_system_cpus(), 10000), 1);
>> 	steal_ratio = div64_u64(delta_steal, delta_ns);
>>
>> On x86 the possible map is sized by topology_init_possible_cpus()
>> (arch/x86/kernel/cpu/topology.c) from assigned + disabled CPUs, where
>> "disabled" is incremented by topo_register_apic() for every APIC that is
>> registered but not present.
>>
>> QEMU emits exactly such MADT entries for the range [smp, maxcpus), and
>> acpi_is_processor_usable() in arch/x86/kernel/acpi/boot.c documents that
>> this is deliberate:
>>
>> 	/*
>> 	 * QEMU expects legacy "Enabled=0" LAPIC entries to be counted as
>> 	 * usable in order to support CPU hotplug in guests.
>> 	 */
>>
>> So those vCPUs are registered as usable, land in nr_disabled_cpus, and
>> end up in the possible map.  A guest started with
>>
>> 	-smp 4,maxcpus=32
>>
>> has num_possible_cpus() == 32 while only 4 vCPUs ever run.  The steal
>> ratio is then diluted 8x, and the default thresholds of 5% / 2% become
>> 40% / 16% of the steal that is actually observable.  The governor never
>> leaves the "do nothing" window, and the only symptom is that nothing
>> happens.
>>
>> Xen PV guests are affected in the same way, and often more strongly,
>> because the possible map there tends to be sized for vCPU hotplug.
>>
>> I understand from the v9 changelog why possible CPUs were chosen -- it
>> keeps the accumulated steal monotonic across hotplug, which is a real
>> property worth having.  The documentation does describe the effect and
>> gives a worked example for recomputing the thresholds by hand.  But for
>> a mechanism whose whole premise is that every VM on the host opts in
>> with the same policy, requiring each operator to first derive their own
>> thresholds seems likely to translate into low real-world adoption.
>>
>> Some options, roughly in increasing order of intrusiveness:
>>
>>    - emit a pr_info() (or pr_warn()) at module init when
>>      num_possible_cpus() significantly exceeds num_online_cpus(), naming
>>      the ratio and the effective thresholds.  Cheap, and turns a silent
>>      no-op into something diagnosable.
>>
>>    - scale the thresholds by num_possible_cpus() / num_online_cpus() at
>>      init, so the documented defaults keep their intended meaning.
>>
>>    - keep summing steal over the possible mask, as today, but divide by
>>      the online count, and handle the hotplug discontinuity by resetting
>>      the sg_ctx.steal baseline from a hotplug notifier.
>>
>> I do not have a strong preference among these, and the first one alone
>> would already be a large improvement.
> 
> I asked the same question on v9 iteration, but the problem is still
> there. I don't think that VM with a big number of offlined CPUs is
> something that needs to bring admin's attention with pr_info(). The
> solution should be found withing the driver.
> 
> I thought about saving the steal time for each CPU in an array, so on
> the next iteration we count steal time from the same subset of CPUs,
> regardless of hot plugging.
> 
> Your approach with num_possible_cpus() / num_online_cpus() scale looks
> nicer, if it works. Definitely need a testing for both. Resetting the
> steal baseline on each hotplug is another option, but to me the least
> attractive.
> 
> This problem must be resolved this way or another before we merge the
> driver.
> 

Looking back at based on what you guys suggested,
i am thinking if we do number of active CPUs, it might work.

The reason being, only active/online CPUs will contribute to delta increase.
Since we sum across possible CPUs, spikes due to hotplug will not happen.

See the diff at the end about returning active CPUs.

>> 2) Nothing stops the driver from folding Xen dom0
>> =================================================
>>
>> dom0 accounts steal time like any other domain -- xen_time_setup_guest()
>> in arch/x86/xen/time.c wires up pv_steal_clock unconditionally, with no
>> feature negotiation and no privileged-domain exemption.
>>
>> So loading steal_governor in dom0 makes it shrink its own preferred mask
>> under contention.  That is precisely when the blkback and netback
>> threads serving every other guest need CPU, and the driver has no notion
>> that this domain is different from the ones it is trying to be polite
>> towards.  The effect would be host-wide, not confined to the domain that
>> loaded the module.
>>
>> Given that Kconfig carries "default m", the module is built on any
>> distro kernel with PARAVIRT=y, which includes dom0 kernels.  It is one
>> modprobe away from being loaded there, quite plausibly by someone who
>> read the documentation's advice to enable it uniformly across all VMs.
>>
>> A xen_initial_domain() check that refuses to load, or at minimum a loud
>> warning, seems worth having.  More generally it may be worth stating in
>> the documentation that the driver is meant for guests only -- the same
>> argument applies to a KVM host that is itself running nested guests.
>>
>> Juergen and the virtualization list are already on Cc -- I would value
>> their view on this one in particular.
>>

I prefer this could be done by arch specific hooks. For the initial series,
this has been deferred.

I had __weak symbols way of allowing arch specific hooks. But it isn't very common
practise as yury suggested. So for now, it has been deferred.

For now, i have added check for not loading on dom0.

>>
>> 3) Core granularity degenerates to single vCPUs on KVM and Xen
>> ==============================================================
>>
>> decrease_preferred_cpus() and increase_preferred_cpus() step by
>> topology_sibling_cpumask(), which matches PowerVM, where the hypervisor
>> schedules whole cores.
>>
>> On Xen PV that mask is always the CPU itself, by design.  The header
>> comment of arch/x86/xen/smp_pv.c is explicit about both the behaviour
>> and the reason for it:
>>
>> 	/*
>> 	 * Because virtual CPUs can be scheduled onto any real CPU, there's
>> 	 * no useful topology information for the kernel to make use of.  As
>> 	 * a result, all CPUs are treated as if they're single-core and
>> 	 * single-threaded.
>> 	 */
>>
>> A plain "-smp N" under QEMU gives one thread per core and one core per
>> socket, with the same result.  So on both hypervisors the step size is
>> one vCPU, and convergence to a folded state takes proportionally longer
>> than the PowerVM numbers would suggest.
>>
>> I do not think this is a correctness problem -- per-vCPU is arguably the
>> right granularity there, for exactly the reason the Xen comment gives.
>> The case that does not work as intended is the opposite one: a guest
>> given a synthetic topology such as
>>
>> 	-smp 16,sockets=1,cores=8,threads=2
>>
>> will fold what it believes is a core, but those two vCPU threads need
>> not be co-located on any host core, so nothing in particular is freed.
> 

vCPU effectively becomes a running thread/task in host. So even with KVM,
doing core level means for example SMT2, it means two less tasks in host
compared to 1. Even though it can run anywhere, it should converge faster.

I don't see why using core will hurt. Maybe I am missing to see it.

> On 2 and 3.
> 
> If the governor will become a practical tool for admins, I believe
> there will be the more corner cases, the more adoption the driver
> gets.

Yes.

> 
> That's why I suggested the modular architecture for the governor. Now
> we have 2 options:
> 
> 1. Keep the driver as a single module, and add those checks for DOM 0

You mean add below check in this driver?

I have added it and once we have arch spefific implementation, we can remove it.

#ifdef CONFIG_XEN
	if (xen_initial_domain()) {
		pr_err("Cannot load in Xen Dom0 (Host OS). Driver is for guests only.\n");
		return -ENODEV;
	}
#endif

See the diff in the end.

>     and per-core vs per-thread behavior.

Unless we see using core level hurts, i don't see a reason why we should
switch to per CPU. If one has numbers for using per cpu method is better for KVM/XEN,
then we can bring in that optimization IMHO.


> 2. Rename this driver to powervm_steal_governor, and let KVM and XEN
>     engineers implement their own drivers.


I prefer to keep it as generic, since it would work fine for XEN too.
I am okay to even add XEN specific checks so it stays as one generic module.
This would help avoid many archs doing same thing more or less with few addition
debug checks. Renaming it to powervm_steal_governor would likely
fragment the solution.

> 
> The options are not mutually exclusive. The earlier versions of the
> driver had a 'core' part and hooks for different arches.
> 
> At this point, I believe, keeping the driver simple is the most
> important goal for making it upstreamable. When it's merged, it
> will be much easier to add more functionality and tuning for
> different VMs and architectures.

That's what I would prefer :)

>   
>> So mainly a documentation request: a sentence in
>> Documentation/driver-api/steal-governor.rst noting that the core-level
>> step assumes the guest topology reflects host scheduling granularity,
>> and that on KVM and Xen it commonly does not.  It would also explain to
>> readers why convergence looks slower there than in your PowerVM numbers.
> 
> I agree, this should be documented.

Ok. See the diff it the end.

> 
>> --
>>
>> I would be happy to run this on x86 KVM and on Xen PV guests if that is
>> useful -- the series has no Xen coverage that I can see, and points 1
>> and 3 above would both show up there first.
> 

If you can try the series it would help.
take the below patch and apply it on top.


> I didn't test the patches myself yet, mostly because configuring VMs
> is a separate topic. If you share a script that configures an
> environment with paravirtual VMs that run some payload that overcommit
> the vCPUs, it would be a valuable addition to the series.
> 
> Andrea Righi with his virtme is already in CC. It's a very friendly
> and highly configurable wrapper around qemu that allows to compile a
> and run kernel with a couple commands. Andrea, can the virtme prepare
> such an environment? For testing we need:
> 
>   - several paravirtual VMs, so that the number of vCPUs exceeds the
>     number of CPUs, and
>   - some payload running on them to overcommit the vCPUs, triggering
>     the steal time counter.
> 
> Thanks,
> Yury

I think below should address all the three concerns.
Let me know if i missed anything.

Only compiled tested for x86 CONFIG_XEN.

---

diff --git a/Documentation/driver-api/steal-governor.rst b/Documentation/driver-api/steal-governor.rst
index 672eeccabfe8..62add9f59032 100644
--- a/Documentation/driver-api/steal-governor.rst
+++ b/Documentation/driver-api/steal-governor.rst
@@ -102,22 +102,17 @@ reduce the preferred CPUs by 1 core.
  Limitations of default values
  -----------------------------
  
-Because of the vast diversity in VM configurations (e.g., highly populated
-vs. sparsely populated CPU masks, few offlined CPUs etc), the default
-thresholds may not be optimal for all systems. Users may need to tune these
-parameters based on the system under test to achieve the best results.
-
-For example:
-Possible CPUs = 128 and Active CPUs = 8
-Steal on online CPUs = 50%
-steal ratio: (50% * 8 + 0% * 120) / 128 = 3.125%
-This would fall in between and default values won't work.
-In this example, if one wants effective 2% and 5% limits, then set,
-low_threshold  = (2% * 8 + 0% * 120) / 128 = 0.1250% = 12
-high_threshold = (5% * 8 + 0% * 120) / 128 = 0.3125% = 31
+Because of the vast diversity in VM configurations, the default
+thresholds may not be optimal for all systems.
  
  Using possible CPUs helps to handle spikes during CPU hotplug as the steal
-time across possible CPUs is a monotonically increasing value.
+time across possible CPUs is a monotonically increasing value. Effective delta
+for steal time can come via only active CPUs. Current logic is effective for even
+sparsely populated systems. But, some corner cases can still linger.
+
+Driver reduces/increases preferred CPUs by core-level. This could provide faster
+convergence for hypervisors such as powerVM. But on KVM and Xen convergence
+could be slower depending on the configuration.
  
  Reasons for CONFIG_STEAL_GOVERNOR=m
  ===================================
@@ -132,6 +127,6 @@ It is recommended to build CONFIG_STEAL_GOVERNOR=m due to below reasons:
  
  2. This works well when all VMs work in co-operative manner. When an
     administrative user enables it in one VM, he/she will likely enable
-   it all VMs.
+   it all VMs. But Note that, it should be enabled it only on actual guests.
  
  3. User can tweak the module parameters by reloading the module.
diff --git a/drivers/virt/steal_governor.c b/drivers/virt/steal_governor.c
index ee92152e64cc..a25eb50a4c7b 100644
--- a/drivers/virt/steal_governor.c
+++ b/drivers/virt/steal_governor.c
@@ -28,6 +28,10 @@
  #include <linux/types.h>
  #include <linux/workqueue.h>
  
+#ifdef CONFIG_XEN
+#include <xen/xen.h>
+#endif
+
  #if !IS_ENABLED(CONFIG_PREFERRED_CPU)
  #error "Steal Governor requires CONFIG_PREFERRED_CPU"
  #endif
@@ -122,7 +126,7 @@ static u64 get_system_steal_time(void)
  /* Return number of CPUs to consider for steal ratio. */
  static unsigned int get_system_cpus(void)
  {
-       return num_possible_cpus();
+       return num_active_cpus();
  }
  
  /*
@@ -254,6 +258,13 @@ static void steal_governor_loop(struct work_struct *work)
  
  static int __init steal_governor_init(void)
  {
+
+#ifdef CONFIG_XEN
+       if (xen_initial_domain()) {
+               pr_err("Cannot load in Xen Dom0 (Host OS). Driver is for guests only.\n");
+               return -ENODEV;
+       }
+#endif
         if (sg_ctx.low_threshold >= sg_ctx.high_threshold) {
                 pr_err("low_threshold (%u) must be less than high_threshold (%u)\n",
                        sg_ctx.low_threshold, sg_ctx.high_threshold);
Re: [PATCH] Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Mete Durlu 1 month, 2 weeks ago
Hi all,

All of the points raised by Ionut are about governor lacking
certain logic for a specific arch. Considering what Shrikanth
mentions under "Future Work";

"""
Known Limitations & Future Work
===============================
...
- Arch specific hints and framework for it as been deferred to
   the future.
...
"""

I think all of these points can be addressed better if
Xen used said framework and implemented their own governor
module. That way we wouldn't see an overinflated single
steal_governor but instead nicely separated arch/platform
specific ones, that are tailored best for their needs.
The current implementation could be the fallback option
if platform does not implement their own and would also
serve as an example.

I really think we should follow the example of cpuidle
drivers and how their framework brings together so many
platforms under a single roof.

Thanks!
-Mete
Re: [PATCH] Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Shrikanth Hegde 1 month, 2 weeks ago
Hi Mete, Thanks for going through the patches/discussions.

On 8/13/26 12:20 PM, Mete Durlu wrote:
> Hi all,
> 
> All of the points raised by Ionut are about governor lacking
> certain logic for a specific arch. Considering what Shrikanth
> mentions under "Future Work";
> 
> """
> Known Limitations & Future Work
> ===============================
> ...
> - Arch specific hints and framework for it as been deferred to
>    the future.
> ...
> """
> 
> I think all of these points can be addressed better if
> Xen used said framework and implemented their own governor
> module. That way we wouldn't see an overinflated single
> steal_governor but instead nicely separated arch/platform
> specific ones, that are tailored best for their needs.
> The current implementation could be the fallback option
> if platform does not implement their own and would also
> serve as an example.

For now, I prefer adding a defensive check for dom0
and keep the driver simple.

> 
> I really think we should follow the example of cpuidle
> drivers and how their framework brings together so many
> platforms under a single roof.
> 

But I remember fredric saying it isn't ideal either and he had
planned to clean it up. (Fredric, correct me if i remember it wrong)

> Thanks!
> -Mete

I prefer we defer the arch specific hooks for now, until there is a need 
for one. If you guys insist it should be done, then i can start looking 
at cpuidle framework. But it will be a bigger rework.

But if the patch given in other thread is good enough for XEN, then we 
could keep it simple one file for the time being.

If the driver eventually outgrows a single file, we can work on a 
modular framework post-merge. But for now, let's keep it simple and get 
the simple version upstream.

Thoughts?
Re: [PATCH] Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Mete Durlu 1 month, 2 weeks ago
On 13/08/2026 13:12, Shrikanth Hegde wrote:
> Hi Mete, Thanks for going through the patches/discussions.
> 
> On 8/13/26 12:20 PM, Mete Durlu wrote:

...

>>
>> I think all of these points can be addressed better if
>> Xen used said framework and implemented their own governor
>> module. That way we wouldn't see an overinflated single
>> steal_governor but instead nicely separated arch/platform
>> specific ones, that are tailored best for their needs.
>> The current implementation could be the fallback option
>> if platform does not implement their own and would also
>> serve as an example.
> 
> For now, I prefer adding a defensive check for dom0
> and keep the driver simple.

Fair enough.

...
> I prefer we defer the arch specific hooks for now, until there is a need 
> for one. If you guys insist it should be done, then i can start looking 
> at cpuidle framework. But it will be a bigger rework.

s390 plans to adopt and start using the preferred CPU approach along
with the governor. The concern is that there are some enhancements
planned which would not really fit into the current governor.

Later on, s390 will probably introduce its own governor module and
for that I was hoping that there would be a framework similar to
cpuidle drivers.

What I mean essentially is a common infrastructure to initialize
the basics required for the preferred CPUs management and maybe
the update loop mechanism. That should ideally leave just the
decision making part to the individual arch/platform to implement.
I imagine the whole thing being much more simpler than cpuidle
drivers as it had a lot more moving parts involved.

...

> If the driver eventually outgrows a single file, we can work on a 
> modular framework post-merge. But for now, let's keep it simple and get 
> the simple version upstream.

I understand the concern and I think it's the right approach to keep
it simple initially.

Thank you!
Re: [PATCH] Re: [PATCH v10 00/12] sched, steal_governor: Introduce preferred CPUs and steal-driven vCPU backoff
Posted by Shrikanth Hegde 1 month, 2 weeks ago
Hi Mete,

On 8/14/26 2:52 PM, Mete Durlu wrote:
> On 13/08/2026 13:12, Shrikanth Hegde wrote:
>> Hi Mete, Thanks for going through the patches/discussions.
>>
>> On 8/13/26 12:20 PM, Mete Durlu wrote:
> 
> ...
> 
>>>
>>> I think all of these points can be addressed better if
>>> Xen used said framework and implemented their own governor
>>> module. That way we wouldn't see an overinflated single
>>> steal_governor but instead nicely separated arch/platform
>>> specific ones, that are tailored best for their needs.
>>> The current implementation could be the fallback option
>>> if platform does not implement their own and would also
>>> serve as an example.
>>
>> For now, I prefer adding a defensive check for dom0
>> and keep the driver simple.
> 
> Fair enough.
> 
> ...
>> I prefer we defer the arch specific hooks for now, until there is a 
>> need for one. If you guys insist it should be done, then i can start 
>> looking at cpuidle framework. But it will be a bigger rework.
> 
> s390 plans to adopt and start using the preferred CPU approach along

That's nice. I am happy to hear that it will come in soon.

> with the governor. The concern is that there are some enhancements
> planned which would not really fit into the current governor.
> 

> Later on, s390 will probably introduce its own governor module and
> for that I was hoping that there would be a framework similar to
> cpuidle drivers.
> 


I was skimming through cpuidle logic. The problem with steal_governor is that,
it is not a built in module. For loadable modules, core_initcall and
device_initcall will evaluate to the same thing,
which makes a direct cpuidle-like approach tricky.

However, as you mentioned, we just need a common infrastructure for the
init/exit/methods leaving the decision-making to the arch.
I've thought of a few ways we can cleanly pull this off post-merge:

1. ifdefs and ops function pointers:
   We define a struct of function pointers and a bit ifdefs for init etc.
   The core initializes ops to default or s390 depending on the config.

2. __weak Functions:
   We define the main routines as __weak and let s390 simply override them.
   This is the lowest boilerplate, though its usage is not preferred.

3. Multiple Modules (The cpuidle module approach):
   We split it into steal_governor_core.ko and steal_governor_s390.ko.
   The core exports a steal_governor_register_driver() symbol, and the
   s390 module registers its specific ops when loaded.
   However, this would still need elements of the first approach to
   handle Dom0-like cases natively.

Depending on users and adoption of the feature by different archs,
we can go about it which is more appropriate.


> What I mean essentially is a common infrastructure to initialize
> the basics required for the preferred CPUs management and maybe
> the update loop mechanism. That should ideally leave just the
> decision making part to the individual arch/platform to implement.
> I imagine the whole thing being much more simpler than cpuidle
> drivers as it had a lot more moving parts involved.
> 
> ...
> 
>> If the driver eventually outgrows a single file, we can work on a 
>> modular framework post-merge. But for now, let's keep it simple and 
>> get the simple version upstream.
> 
> I understand the concern and I think it's the right approach to keep
> it simple initially.

Thanks. Yes. Lets keep it simple for now.

Once this series is merged, We can work out the
framework for s390-specific enhancements.

> 
> Thank you!
> 

PS: I will wait for few days to hear from Ionut/Yury on the patch addressing
all the comments. If i don't hear anything back, i will post v11 next week.