[PATCH v2 0/3] block: skip the blkcg walk in blk_cgroup_congested() when nothing is throttled

Usama Arif posted 3 patches 1 month, 2 weeks ago
block/blk-cgroup.c         | 15 ++++++++++++++-
block/blk-cgroup.h         | 25 +++++++++++++++++++++----
block/blk-iocost.c         |  7 +++++++
block/blk-iolatency.c      |  9 +++++++++
include/linux/blk-cgroup.h | 23 ++++++++++++++++++++++-
5 files changed, 73 insertions(+), 6 deletions(-)
[PATCH v2 0/3] block: skip the blkcg walk in blk_cgroup_congested() when nothing is throttled
Posted by Usama Arif 1 month, 2 weeks ago
blk_cgroup_congested() walks the current task's blkcg ancestor chain on every
readahead decision and, once swap is in use, on every anonymous and shmem
folio allocation.  The answer is almost always "no", but finding that out
costs two loads per level on two cold cache lines, plus an out-of-line
kthread_blkcg() and an RCU read-side pair.  On a fleet profile of hosts
running containers with 5-10 level hierarchies it costs about as much as all
of mutex_lock(), 99.4% of it under __folio_throttle_swaprate().

Patch 3 gates the walk on a global count of blkcgs with a non-zero
congestion_count, so the common case is a load and a predicted branch.

That only works if the count is correctly maintained, currently two teardown
paths can leave a blkcg permanently marked congested.  Today that only hurts
tasks in the affected cgroup, but it hurts them for the life of the cgroup -
readahead cut to a single page, async readahead skipped, and a throttle
scheduled on every anonymous folio allocation.  With a global gate it would
cost every other task on the machine the walk as well.  Patches 1 and 2 fix
those two paths and stand on their own as bugfixes; patch 3 depends on them.

v1 -> v2 (Tejun):
- Rename blkcg_congested_blkcgs to blkcg_nr_congested to make it clear
  that the global tracks a count rather than a boolean.
- Warn if blkcg_css_free() finds a residual congestion_count, while still
  dropping its contribution so it cannot disable the fast path permanently.
- Use atomic_dec_and_test() for the congestion_count 1 -> 0 transition.

Usama Arif (3):
  blk-iolatency: clear delay state when freeing policy data
  blk-iocost: clear delay state when freeing policy data
  block: skip blkcg walk in blk_cgroup_congested() when nothing
    throttled

 block/blk-cgroup.c         | 15 ++++++++++++++-
 block/blk-cgroup.h         | 25 +++++++++++++++++++++----
 block/blk-iocost.c         |  7 +++++++
 block/blk-iolatency.c      |  9 +++++++++
 include/linux/blk-cgroup.h | 23 ++++++++++++++++++++++-
 5 files changed, 73 insertions(+), 6 deletions(-)

-- 
2.53.0-Meta
Re: [PATCH v2 0/3] block: skip the blkcg walk in blk_cgroup_congested() when nothing is throttled
Posted by Jens Axboe 1 month, 1 week ago
On Fri, 14 Aug 2026 09:56:36 -0700, Usama Arif wrote:
> blk_cgroup_congested() walks the current task's blkcg ancestor chain on every
> readahead decision and, once swap is in use, on every anonymous and shmem
> folio allocation.  The answer is almost always "no", but finding that out
> costs two loads per level on two cold cache lines, plus an out-of-line
> kthread_blkcg() and an RCU read-side pair.  On a fleet profile of hosts
> running containers with 5-10 level hierarchies it costs about as much as all
> of mutex_lock(), 99.4% of it under __folio_throttle_swaprate().
> 
> [...]

Applied, thanks!

[1/3] blk-iolatency: clear delay state when freeing policy data
      commit: 8935bf22c0a0db517a7f72f7097300e05dd852f5
[2/3] blk-iocost: clear delay state when freeing policy data
      commit: 97cb95d2148835ae86ff916b145aef332d40439d
[3/3] block: skip blkcg walk in blk_cgroup_congested() when nothing throttled
      commit: 4febfe7d98948bf6693f5c6a0a7e198e8fb4e584

Best regards,
-- 
Jens Axboe
Re: [PATCH v2 0/3] block: skip the blkcg walk in blk_cgroup_congested() when nothing is throttled
Posted by Tejun Heo 1 month, 2 weeks ago
On Fri, Aug 14, 2026 at 09:56:36AM -0700, Usama Arif wrote:
> blk_cgroup_congested() walks the current task's blkcg ancestor chain on every
> readahead decision and, once swap is in use, on every anonymous and shmem
> folio allocation.  The answer is almost always "no", but finding that out
> costs two loads per level on two cold cache lines, plus an out-of-line
> kthread_blkcg() and an RCU read-side pair.  On a fleet profile of hosts
> running containers with 5-10 level hierarchies it costs about as much as all
> of mutex_lock(), 99.4% of it under __folio_throttle_swaprate().
> 
> Patch 3 gates the walk on a global count of blkcgs with a non-zero
> congestion_count, so the common case is a load and a predicted branch.
> 
> That only works if the count is correctly maintained, currently two teardown
> paths can leave a blkcg permanently marked congested.  Today that only hurts
> tasks in the affected cgroup, but it hurts them for the life of the cgroup -
> readahead cut to a single page, async readahead skipped, and a throttle
> scheduled on every anonymous folio allocation.  With a global gate it would
> cost every other task on the machine the walk as well.  Patches 1 and 2 fix
> those two paths and stand on their own as bugfixes; patch 3 depends on them.

For the series,

Acked-by: Tejun Heo <tj@kernel.org>

Thanks.

-- 
tejun