fs/pidfs.c | 43 +++++++++++++++++++++++++++++++------------ 1 file changed, 31 insertions(+), 12 deletions(-)
From: Chen Linxuan <me@black-desk.cn>
The PIDFD_GET_*_NAMESPACE ioctls in pidfd_ioctl() perform a filesystem
credentials ptrace access check before handing out a namespace file
descriptor. The accompanying comment states that the code "mirrors nsfs
behavior", but, unlike the corresponding procfs paths, it does so without
holding the target task's exec_update_lock.
proc_ns_get_link() and proc_ns_readlink() both take exec_update_lock for
reading around the ptrace check and the namespace lookup, so that the
credentials used for the access decision match those of the task when its
namespace is read. Without it, a caller can pass the check against the
target's old credentials and then read the namespace after the target has
execve()'d a setuid binary and committed new credentials -- accessing
namespace information it should have been denied.
Hold exec_update_lock for reading around the ptrace check and the
namespace lookup so that pidfd truly mirrors nsfs behavior, as the comment
already claims. open_namespace() itself runs outside the lock: once a
namespace reference is obtained it carries its own refcount and is opened
with the caller's own credentials, so a concurrent execve() on the target
can no longer affect the outcome.
Fixes: 5b08bd408534 ("pidfs: allow retrieval of namespace file descriptors")
Cc: stable@vger.kernel.org
Assisted-by: Codex:glm-5.2
Signed-off-by: Chen Linxuan <me@black-desk.cn>
---
fs/pidfs.c | 43 +++++++++++++++++++++++++++++++------------
1 file changed, 31 insertions(+), 12 deletions(-)
diff --git a/fs/pidfs.c b/fs/pidfs.c
index b57ecc96e967..a6a643f15d08 100644
--- a/fs/pidfs.c
+++ b/fs/pidfs.c
@@ -532,6 +532,7 @@ static long pidfd_ioctl(struct file *file, unsigned int cmd, unsigned long arg)
struct task_struct *task __free(put_task) = NULL;
struct nsproxy *nsp __free(put_nsproxy) = NULL;
struct ns_common *ns_common = NULL;
+ int error;
if (!pidfs_ioctl_valid(cmd))
return -ENOIOCTLCMD;
@@ -555,20 +556,33 @@ static long pidfd_ioctl(struct file *file, unsigned int cmd, unsigned long arg)
if (arg)
return -EINVAL;
+ /*
+ * We're trying to open a file descriptor to the namespace so perform a
+ * filesystem cred ptrace check. Hold @task's exec_update_lock for the
+ * duration of the ptrace check and the namespace lookup so that the
+ * credentials used for the access decision match those of @task at the
+ * time its namespace is read, preventing a concurrent execve() from
+ * swapping the task's credentials in between the check and the use. We
+ * mirror nsfs behavior.
+ */
+ error = down_read_killable(&task->signal->exec_update_lock);
+ if (error)
+ return error;
+
+ if (!ptrace_may_access(task, PTRACE_MODE_READ_FSCREDS)) {
+ error = -EACCES;
+ goto out_unlock;
+ }
+
scoped_guard(task_lock, task) {
nsp = task->nsproxy;
if (nsp)
get_nsproxy(nsp);
}
- if (!nsp)
- return -ESRCH; /* just pretend it didn't exist */
-
- /*
- * We're trying to open a file descriptor to the namespace so perform a
- * filesystem cred ptrace check. Also, we mirror nsfs behavior.
- */
- if (!ptrace_may_access(task, PTRACE_MODE_READ_FSCREDS))
- return -EACCES;
+ if (!nsp) {
+ error = -ESRCH; /* just pretend it didn't exist */
+ goto out_unlock;
+ }
switch (cmd) {
/* Namespaces that hang of nsproxy. */
@@ -650,11 +664,16 @@ static long pidfd_ioctl(struct file *file, unsigned int cmd, unsigned long arg)
#endif
break;
default:
- return -ENOIOCTLCMD;
+ error = -ENOIOCTLCMD;
}
- if (!ns_common)
- return -EOPNOTSUPP;
+ if (!error && !ns_common)
+ error = -EOPNOTSUPP;
+
+out_unlock:
+ up_read(&task->signal->exec_update_lock);
+ if (error)
+ return error;
/* open_namespace() unconditionally consumes the reference */
return open_namespace(ns_common);
---
base-commit: 8ba098e6b6ff0db8edf28528d1552be261af30d4
change-id: 20260731-pidfd-exec-update-lock-de8e990154c6
Best regards,
--
Chen Linxuan <me@black-desk.cn>
On Fri, 31 Jul 2026 22:50:08 +0800, Chen Linxuan wrote:
> pidfd: hold exec_update_lock around namespace ioctl
Applied to the vfs-7.3.misc branch of the vfs/vfs.git tree.
Patches in the vfs-7.3.misc branch should appear in linux-next soon.
Please report any outstanding bugs that were missed during review in a
new review to the original patch series allowing us to drop it.
It's encouraged to provide Acked-bys and Reviewed-bys even though the
patch has now been applied. If possible patch trailers will be updated.
Note that commit hashes shown below are subject to change due to rebase,
trailer updates or similar. If in doubt, please check the listed branch.
tree: https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
branch: vfs-7.3.misc
[1/1] pidfd: hold exec_update_lock around namespace ioctl
https://git.kernel.org/vfs/vfs/c/a852a3066d3f
On Fri, Jul 31, 2026 at 4:50 PM Chen Linxuan via B4 Relay <devnull+me.black-desk.cn@kernel.org> wrote: > The PIDFD_GET_*_NAMESPACE ioctls in pidfd_ioctl() perform a filesystem > credentials ptrace access check before handing out a namespace file > descriptor. The accompanying comment states that the code "mirrors nsfs > behavior", but, unlike the corresponding procfs paths, it does so without > holding the target task's exec_update_lock. > > proc_ns_get_link() and proc_ns_readlink() both take exec_update_lock for > reading around the ptrace check and the namespace lookup, so that the > credentials used for the access decision match those of the task when its > namespace is read. Without it, a caller can pass the check against the > target's old credentials and then read the namespace after the target has > execve()'d a setuid binary and committed new credentials -- accessing > namespace information it should have been denied. > > Hold exec_update_lock for reading around the ptrace check and the > namespace lookup so that pidfd truly mirrors nsfs behavior, as the comment > already claims. open_namespace() itself runs outside the lock: once a > namespace reference is obtained it carries its own refcount and is opened > with the caller's own credentials, so a concurrent execve() on the target > can no longer affect the outcome. I think this makes sense. Given that the rest of this function is written with scope-based cleanup (https://docs.kernel.org/core-api/cleanup.html), I wonder if this patch would look cleaner if it also used scope-based cleanup... that documentation also says: "Lastly, given that the benefit of cleanup helpers is removal of “goto”, and that the “goto” statement can jump between scopes, the expectation is that usage of “goto” and cleanup helpers is never mixed in the same function." But we probably shouldn't be holding the exec_update_lock across the open_namespace() call... so I guess this will require either refactoring this function into two, or indenting most of the function body, or some explicit "drop this guard" operation. @Christian, do you have an opinion on this?
On Mon, Aug 03, 2026 at 03:44:21PM +0200, Jann Horn wrote: > On Fri, Jul 31, 2026 at 4:50 PM Chen Linxuan via B4 Relay > <devnull+me.black-desk.cn@kernel.org> wrote: > > The PIDFD_GET_*_NAMESPACE ioctls in pidfd_ioctl() perform a filesystem > > credentials ptrace access check before handing out a namespace file > > descriptor. The accompanying comment states that the code "mirrors nsfs > > behavior", but, unlike the corresponding procfs paths, it does so without > > holding the target task's exec_update_lock. > > > > proc_ns_get_link() and proc_ns_readlink() both take exec_update_lock for > > reading around the ptrace check and the namespace lookup, so that the > > credentials used for the access decision match those of the task when its > > namespace is read. Without it, a caller can pass the check against the > > target's old credentials and then read the namespace after the target has > > execve()'d a setuid binary and committed new credentials -- accessing > > namespace information it should have been denied. > > > > Hold exec_update_lock for reading around the ptrace check and the > > namespace lookup so that pidfd truly mirrors nsfs behavior, as the comment > > already claims. open_namespace() itself runs outside the lock: once a > > namespace reference is obtained it carries its own refcount and is opened > > with the caller's own credentials, so a concurrent execve() on the target > > can no longer affect the outcome. > > I think this makes sense. > > Given that the rest of this function is written with scope-based > cleanup (https://docs.kernel.org/core-api/cleanup.html), I wonder if > this patch would look cleaner if it also used scope-based cleanup... > that documentation also says: > > "Lastly, given that the benefit of cleanup helpers is removal of > “goto”, and that the “goto” statement can jump between scopes, the > expectation is that usage of “goto” and cleanup helpers is never mixed > in the same function." > > But we probably shouldn't be holding the exec_update_lock across the > open_namespace() call... so I guess this will require either > refactoring this function into two, or indenting most of the function > body, or some explicit "drop this guard" operation. > > @Christian, do you have an opinion on this? No strong opinion tbh. If I don't like the next version I can also just massage it when applying.
On Tue, Aug 11, 2026 at 4:54 PM Christian Brauner <brauner@kernel.org> wrote: > > On Mon, Aug 03, 2026 at 03:44:21PM +0200, Jann Horn wrote: > > On Fri, Jul 31, 2026 at 4:50 PM Chen Linxuan via B4 Relay > > <devnull+me.black-desk.cn@kernel.org> wrote: > > > The PIDFD_GET_*_NAMESPACE ioctls in pidfd_ioctl() perform a filesystem > > > credentials ptrace access check before handing out a namespace file > > > descriptor. The accompanying comment states that the code "mirrors nsfs > > > behavior", but, unlike the corresponding procfs paths, it does so without > > > holding the target task's exec_update_lock. > > > > > > proc_ns_get_link() and proc_ns_readlink() both take exec_update_lock for > > > reading around the ptrace check and the namespace lookup, so that the > > > credentials used for the access decision match those of the task when its > > > namespace is read. Without it, a caller can pass the check against the > > > target's old credentials and then read the namespace after the target has > > > execve()'d a setuid binary and committed new credentials -- accessing > > > namespace information it should have been denied. > > > > > > Hold exec_update_lock for reading around the ptrace check and the > > > namespace lookup so that pidfd truly mirrors nsfs behavior, as the comment > > > already claims. open_namespace() itself runs outside the lock: once a > > > namespace reference is obtained it carries its own refcount and is opened > > > with the caller's own credentials, so a concurrent execve() on the target > > > can no longer affect the outcome. > > > > I think this makes sense. > > > > Given that the rest of this function is written with scope-based > > cleanup (https://docs.kernel.org/core-api/cleanup.html), I wonder if > > this patch would look cleaner if it also used scope-based cleanup... > > that documentation also says: > > > > "Lastly, given that the benefit of cleanup helpers is removal of > > “goto”, and that the “goto” statement can jump between scopes, the > > expectation is that usage of “goto” and cleanup helpers is never mixed > > in the same function." > > > > But we probably shouldn't be holding the exec_update_lock across the > > open_namespace() call... so I guess this will require either > > refactoring this function into two, or indenting most of the function > > body, or some explicit "drop this guard" operation. > > > > @Christian, do you have an opinion on this? > > No strong opinion tbh. If I don't like the next version I can also just > massage it when applying. I did consider using scope-based cleanup, but there's a gap in the available guard infrastructure: rwsem read locks have conditional guard variants for _try and _intr, but not for _kill (down_read_killable). I'd prefer to keep the current manual lock/unlock + goto style for now rather than switch to interruptible acquisition or add a new guard definition just for this.
On Fri, Jul 31, 2026 at 10:50 PM Chen Linxuan via B4 Relay <devnull+me.black-desk.cn@kernel.org> wrote: > > From: Chen Linxuan <me@black-desk.cn> > > The PIDFD_GET_*_NAMESPACE ioctls in pidfd_ioctl() perform a filesystem > credentials ptrace access check before handing out a namespace file > descriptor. The accompanying comment states that the code "mirrors nsfs > behavior", but, unlike the corresponding procfs paths, it does so without > holding the target task's exec_update_lock. After sending v1 I realised that the exec_update_lock in the procfs namespace paths was just added by commit 6650527444da to fix CVE-2026-64371. So the pidfd PIDFD_GET_*_NAMESPACE ioctls, which mirror those procfs paths, should be fixed the same way. CC Jann.
© 2016 - 2026 Red Hat, Inc.