[PATCH 0/4] kernfs: remove kernfs_rwsem from dentry revalidation

Shakeel Butt posted 4 patches 1 month, 1 week ago
fs/kernfs/dir.c             | 76 +++++++++++++++----------------------
fs/kernfs/kernfs-internal.h |  9 ++---
2 files changed, 35 insertions(+), 50 deletions(-)
[PATCH 0/4] kernfs: remove kernfs_rwsem from dentry revalidation
Posted by Shakeel Butt 1 month, 1 week ago
At Meta, we are seeing important system daemons that poll cgroupfs and
sysfs geth stuck in kernfs_dop_revalidate() for minutes. The two that
hurt most are the ones we can least afford to lose: oomd, which decides
what to kill when a machine runs out of memory, and below[1], which
records the telemetry used to understand what happened afterwards.

kernfs_dop_revalidate() takes kernfs_rwsem for read once per path component
of every walk into a kernfs mount. Linux rwsems do not permit reader lock
stealing once a writer is queued, so one writer -- a cgroup created or
destroyed, a device renamed -- parks the entire incoming reader stream in
uninterruptible sleep. Daemons polling cgroup files in a loop are exactly
the workload that turns this into a convoy, and cgroup churn is exactly
what a busy machine does.

Nothing the callback reads needs the semaphore. kn->active is an atomic_t
already tested lock-free elsewhere, kn->__parent and kn->name are RCU
pointers, kn->ns can be compared rather than dereferenced, and
parent->dir.rev is a plain counter.

  1/4 uses the parent inode and name the VFS already passes to
      ->d_revalidate() rather than recovering them from mutable dentry
      fields, comparing the name by explicit length
  2/4 annotates the directory revision counter for lockless access
  3/4 compares namespace tags by pointer
  4/4 removes kernfs_rwsem from the callback

LOOKUP_RCU still returns -ECHILD. kernfs_iop_permission() forces every walk
out of RCU-walk before children are revalidated, so lifting it here would
have no effect until that path is fixed; left to a separate series.

Readers walking cgroupfs and sysfs while another thread churns cgroups,
renames netdevs and adds/removes devices, on an 8-CPU VM:

                    kernfs_rwsem read   contentions   revalidate among
                        acquisitions                  top call sites
  before                  48,593,360     1,744,517    #1 and #2
  after                    1,280,280       429,846    absent

Reader path-walk throughput improved 35-53% over the same workload.

Tested against an unpatched control of the same tree, built and booted with
KASAN, KCSAN (default and STRICT), PROVE_LOCKING, PROVE_RCU,
DEBUG_ATOMIC_SLEEP and LOCK_STAT. The deactivated, renamed and
namespace-moved reject paths and negative-dentry invalidation all behave as
before. KCSAN_STRICT over 180s reports no data race involving
kernfs_dop_revalidate() or any field it reads, and there are no KASAN,
lockdep or might-sleep reports across millions of concurrent path walks.
The only kernfs KCSAN reports are in kernfs_refresh_inode(), present
identically on the control and addressed separately.

Link: https://github.com/facebookincubator/below [1]

Shakeel Butt (4):
  kernfs: Use VFS lookup context in d_revalidate()
  kernfs: Prepare directory revisions for lockless reads
  kernfs: Avoid namespace dereference in d_revalidate()
  kernfs: Remove kernfs_rwsem from dentry revalidation

 fs/kernfs/dir.c             | 76 +++++++++++++++----------------------
 fs/kernfs/kernfs-internal.h |  9 ++---
 2 files changed, 35 insertions(+), 50 deletions(-)


base-commit: 7079a12d7506b07fb53b54a664bfad5fa9b16d70
-- 
2.53.0-Meta
Re: [PATCH 0/4] kernfs: remove kernfs_rwsem from dentry revalidation
Posted by Christian Brauner 1 month ago
On Thu, 20 Aug 2026 22:05:03 -0700, Shakeel Butt wrote:
> kernfs: remove kernfs_rwsem from dentry revalidation
> 
> At Meta, we are seeing important system daemons that poll cgroupfs and
> sysfs geth stuck in kernfs_dop_revalidate() for minutes. The two that
> hurt most are the ones we can least afford to lose: oomd, which decides
> what to kill when a machine runs out of memory, and below[1], which
> records the telemetry used to understand what happened afterwards.
> 
> [...]

Applied to the vfs-7.4.kernfs branch of the vfs/vfs.git tree.
Patches in the vfs-7.4.kernfs branch should appear in linux-next soon.

Please report any outstanding bugs that were missed during review in a
new review to the original patch series allowing us to drop it.

It's encouraged to provide Acked-bys and Reviewed-bys even though the
patch has now been applied. If possible patch trailers will be updated.

Note that commit hashes shown below are subject to change due to rebase,
trailer updates or similar. If in doubt, please check the listed branch.

tree:   https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
branch: vfs-7.4.kernfs

[1/4] kernfs: Use VFS lookup context in d_revalidate()
      https://git.kernel.org/vfs/vfs/c/b3c97f5cb084
[2/4] kernfs: Prepare directory revisions for lockless reads
      https://git.kernel.org/vfs/vfs/c/272e0997df44
[3/4] kernfs: Avoid namespace dereference in d_revalidate()
      https://git.kernel.org/vfs/vfs/c/e1c629a8dbc7
[4/4] kernfs: Remove kernfs_rwsem from dentry revalidation
      https://git.kernel.org/vfs/vfs/c/8c6ad578e2f0
Re: [PATCH 0/4] kernfs: remove kernfs_rwsem from dentry revalidation
Posted by Ian Kent 1 month ago
On 25/8/26 22:58, Christian Brauner wrote:
> On Thu, 20 Aug 2026 22:05:03 -0700, Shakeel Butt wrote:
>> kernfs: remove kernfs_rwsem from dentry revalidation
>>
>> At Meta, we are seeing important system daemons that poll cgroupfs and
>> sysfs geth stuck in kernfs_dop_revalidate() for minutes. The two that
>> hurt most are the ones we can least afford to lose: oomd, which decides
>> what to kill when a machine runs out of memory, and below[1], which
>> records the telemetry used to understand what happened afterwards.
>>
>> [...]
> Applied to the vfs-7.4.kernfs branch of the vfs/vfs.git tree.
> Patches in the vfs-7.4.kernfs branch should appear in linux-next soon.
>
> Please report any outstanding bugs that were missed during review in a
> new review to the original patch series allowing us to drop it.
>
> It's encouraged to provide Acked-bys and Reviewed-bys even though the
> patch has now been applied. If possible patch trailers will be updated.


I have a report against kernfs where I cannot find a reason for significant

contention, primarily on ->revaliate(), in either of two vmcores, exactly

the case described in this series.


I reviewed the series and it looks good to me.

The descriptions justifying the changes also look sound.


While I think renames and removals should be ok based on the descriptions

justifying the change (my biggest concern with removing the rwsem) I'll

continue to ponder its implications for a while.


Reviewed-by: Ian Kent <raven@themaw.net>


>
> Note that commit hashes shown below are subject to change due to rebase,
> trailer updates or similar. If in doubt, please check the listed branch.
>
> tree:   https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
> branch: vfs-7.4.kernfs
>
> [1/4] kernfs: Use VFS lookup context in d_revalidate()
>        https://git.kernel.org/vfs/vfs/c/b3c97f5cb084
> [2/4] kernfs: Prepare directory revisions for lockless reads
>        https://git.kernel.org/vfs/vfs/c/272e0997df44
> [3/4] kernfs: Avoid namespace dereference in d_revalidate()
>        https://git.kernel.org/vfs/vfs/c/e1c629a8dbc7
> [4/4] kernfs: Remove kernfs_rwsem from dentry revalidation
>        https://git.kernel.org/vfs/vfs/c/8c6ad578e2f0
>
>