From nobody Sat Jul 25 05:59:38 2026 Received: from mail-pj1-f42.google.com (mail-pj1-f42.google.com [209.85.216.42]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1DC0141684E for ; Fri, 24 Jul 2026 22:02:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.42 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930543; cv=none; b=XFyc+XyPs+aDSCdMzpqXVAplwCqCtvWV8zT75e4Rr+VQ4ELBMYXjushpbx93CpXqa+W1RowZxoFnhUoA/zt17aRxQTjbtJK3WEs5P2eN15Jy6eSMPinp7QUpdtoaURnooxlhrvjHdcFa0NmxVAiwOU3+Dx1kuAmdKl63ghTJ3Vw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930543; c=relaxed/simple; bh=Tw1zHtn6CRvQptldaz10B1Em/cClkzWrEZZkFPoeiB8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Drz4/8ycaIpkzR3x0YsdLud48Fso0BMXkGyo2RhfduSmxOonXMZUJHQtHFqJR0dNp+wncAajDEM+7TURsPdCDW1l81FIK63g2LquVsux9F2mYRmGDF+LgYQRwk4x4FpE03rVkPTn/OH7JyNY9av41mQ8wT2A1CwXwz05qiMOhWo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=K56jzAIp; arc=none smtp.client-ip=209.85.216.42 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="K56jzAIp" Received: by mail-pj1-f42.google.com with SMTP id 98e67ed59e1d1-381c51fde6bso981317a91.2 for ; Fri, 24 Jul 2026 15:02:18 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784930538; x=1785535338; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=8IX3CybTbQbjQjEh5WM2j4ukbfHE6wlKdSHlOsft0Ic=; b=K56jzAIp6li8Dg63fJpdDKMF83RnyTPRoGGod/ECFyeBklvdPWsw5wLcXCZvAlPhs7 PxiIwhFtG8PiDoQIPQRDeRGJ4mXHy/Ywrk9geFvohdyw8uH66Ft5BGi67UuB+58V/1b2 PkA/dcO4UCyTT5iaM9C5A2MVfSSzVbvFQcngZhoVHlShKG82RwUbTpVoGhBWLF1DJ+Sh Os4c73BOzrTW9T3cg8no8YNS5JVd/Pl8WHXX1BNeQZFJ3Eym+GEe5Gi/wUR1ZJ93z1+c 6QZzN6tEFfvj5gTEX017iVsFRtvnOz5CF+zDj6XEAGx2D/v1tPsln8bcW/hnOcqAQTCo N/xw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784930538; x=1785535338; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=8IX3CybTbQbjQjEh5WM2j4ukbfHE6wlKdSHlOsft0Ic=; b=FLUl54j9fzNw2ILjT6C4dXS5WQY9tWDj3Cka1BeZZDK8RzUryPG3hL/clQhol5ggBf 8HaiILwJZ6lIxo+wITcBZ16tFq+vlFsl0UgBlGYbvuO4YGZEoCF5j6UD0gj382UbTNFP oFrustx+AHOy2+RXTd9PjxK9OTlaGcaSFeKY9qFcdhsHVwqjs+SST8mvcYBUj0RtpMaM 4oaUPQwUyCas0FvpO3zGg6mbtYCnM9XSXK8FCc5sc4r3hQGPnIQIyAhQXBQ3lREpXTrx braTxsc1hz+S+g5GC3cCNoGqj67tfOQxcg3nFPlOU2tsjNiaRveIhpqOra6bsnG0skXw OsOQ== X-Forwarded-Encrypted: i=1; AHgh+RosgpOUuzJ06KRTJ2qvDKl8ZdoS6bjFQVrMyvOhhaxy4U1v3+ntmYeT0nwAxXA1eWq40XkQfgHqFAsiS9A=@vger.kernel.org X-Gm-Message-State: AOJu0Yz2JxRAkqw/WbJ5ZbgnymLQKjBwoVQpMnBePOR8cW2aYy57/BZw 5q8DbjkHihD36L4HYl0dukWgYSdhFtr8Lbf3AFB91HB2+qXdmG3snEEL X-Gm-Gg: AR+sD13jQ+gIbzHhRAE2Flesq+9h3kXZFohM419QRKC/A8P7w2sJFgzW86juG0LqAF5 urWX3WnnsApuBImiM4xTFtxozvxH5IXwO02+4AI4kpzEh8msTMhVb006J6namT9EBG66fZm6yuv s3fCt394zgR5qRjsXLmlWWUM2RTCW7YvXfmlG7/6xnAndz2EF+Uh8fjeWKtkW+8w36trHCepE3r PKKE7nKq/ZRlp0x14cFtKj+xTVHAHy+VrD12QCJyWQs4tm+60KLEbaVNXgFNdBplx537gkhvFhN tPOPqkhe2IFT7ugtGHOFB+atKNsJhgSiL+04tYrGfRDwfdOCS+W6cxveuNpQVYcpwmhtIhLVQnH IVT7uPz1vGV8WYs1Pb6zlY+WUw/p1TbcF5YUYQH8toHbhp+GjE9igTNihbuR85r8LZKrwPsklAG vdVM8NcGYtwTSr5KajZV3MWLC9xhuwYUVEy6Rkf8hVn4uMwQtGHetldAag9iZ5 X-Received: by 2002:a17:90b:2dc5:b0:38d:e0c4:c955 with SMTP id 98e67ed59e1d1-38f2950e5b0mr219022a91.15.1784930537606; Fri, 24 Jul 2026 15:02:17 -0700 (PDT) Received: from pop-os.tail4adac7.ts.net ([2601:647:6802:dbc0:546a:e1b0:9574:3a29]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc5a67f3sm2962631eec.29.2026.07.24.15.02.15 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 15:02:16 -0700 (PDT) From: Cong Wang To: Andy Lutomirski Cc: Kees Cook , linux-kernel@vger.kernel.org, Will Drewry , Christian Brauner , Andrew Morton , linux-mm@kvack.org, Cong Wang Subject: [PATCH v7 1/8] mm: pass the target mm parameter through get_unmapped_area family Date: Fri, 24 Jul 2026 15:01:40 -0700 Message-ID: <20260724220147.214396-2-xiyou.wangcong@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260724220147.214396-1-xiyou.wangcong@gmail.com> References: <20260724220147.214396-1-xiyou.wangcong@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Cong Wang Every placement callback in the mmap path implicitly resolves against current->mm, so a placement cannot be computed in another task's address space. vm_mmap_remote() in the next patch needs this: without it a remote install cannot run the file's own ->get_unmapped_area() and file-specific placement rules are bypassed or rejected at MAP_FIXED validation. Make the mm an explicit first parameter of the get_unmapped_area family across the entire tree. It is difficult and not elegant to avoid such a big change. The get_unmapped_area() convenience inline keeps its old signature and passes current->mm, so ordinary callers do not change. So no functional change. Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Cong Wang --- Documentation/filesystems/locking.rst | 2 +- Documentation/filesystems/vfs.rst | 2 +- arch/alpha/kernel/osf_sys.c | 21 +++--- arch/arc/mm/mmap.c | 7 +- arch/arm/mm/mmap.c | 19 +++-- arch/csky/abiv1/mmap.c | 7 +- arch/loongarch/mm/mmap.c | 28 ++++---- arch/mips/mm/mmap.c | 28 ++++---- arch/parisc/kernel/sys_parisc.c | 28 ++++---- arch/powerpc/include/asm/book3s/64/slice.h | 3 +- arch/powerpc/mm/book3s64/slice.c | 28 ++++---- arch/s390/mm/mmap.c | 25 ++++--- arch/sh/mm/mmap.c | 21 +++--- arch/sparc/include/asm/pgtable_64.h | 6 +- arch/sparc/kernel/sys_sparc_32.c | 7 +- arch/sparc/kernel/sys_sparc_64.c | 32 +++++---- arch/x86/include/asm/elf.h | 2 +- arch/x86/kernel/cpu/sgx/driver.c | 13 ++-- arch/x86/kernel/sys_x86_64.c | 32 +++++---- arch/x86/kernel/uprobes.c | 4 +- arch/x86/mm/mmap.c | 4 +- arch/xtensa/kernel/syscall.c | 8 +-- drivers/char/mem.c | 15 ++-- drivers/dax/device.c | 11 +-- drivers/gpu/drm/drm_gem.c | 9 ++- drivers/gpu/drm/drm_gem_dma_helper.c | 4 +- drivers/media/v4l2-core/v4l2-dev.c | 3 +- drivers/mtd/mtdchar.c | 3 +- drivers/video/fbdev/core/fb_chrdev.c | 3 +- fs/cramfs/inode.c | 3 +- fs/hugetlbfs/inode.c | 9 +-- fs/proc/inode.c | 15 ++-- fs/ramfs/file-mmu.c | 5 +- fs/ramfs/file-nommu.c | 6 +- fs/romfs/mmap-nommu.c | 3 +- include/drm/drm_gem.h | 3 +- include/drm/drm_gem_dma_helper.h | 3 +- include/linux/bpf.h | 3 +- include/linux/fs.h | 5 +- include/linux/huge_mm.h | 18 ++--- include/linux/hugetlb.h | 6 +- include/linux/mm.h | 12 ++-- include/linux/proc_fs.h | 5 +- include/linux/sched/mm.h | 37 +++++----- include/linux/shmem_fs.h | 5 +- io_uring/memmap.c | 8 ++- io_uring/memmap.h | 3 +- ipc/shm.c | 8 +-- kernel/bpf/arena.c | 5 +- kernel/bpf/syscall.c | 8 ++- mm/huge_memory.c | 27 +++---- mm/mmap.c | 83 ++++++++++++---------- mm/nommu.c | 5 +- mm/shmem.c | 13 ++-- mm/vma.c | 16 +++-- mm/vma.h | 6 +- sound/core/pcm_native.c | 3 +- 57 files changed, 391 insertions(+), 307 deletions(-) diff --git a/Documentation/filesystems/locking.rst b/Documentation/filesyst= ems/locking.rst index 08d01bc62c31..b5defd74e302 100644 --- a/Documentation/filesystems/locking.rst +++ b/Documentation/filesystems/locking.rst @@ -471,7 +471,7 @@ prototypes:: int (*fsync) (struct file *, loff_t start, loff_t end, int datasync); int (*fasync) (int, struct file *, int); int (*lock) (struct file *, int, struct file_lock *); - unsigned long (*get_unmapped_area)(struct file *, unsigned long, + unsigned long (*get_unmapped_area)(struct mm_struct *, struct file *, uns= igned long, unsigned long, unsigned long, unsigned long); int (*check_flags)(int); int (*flock) (struct file *, int, struct file_lock *); diff --git a/Documentation/filesystems/vfs.rst b/Documentation/filesystems/= vfs.rst index 7c753148af88..4f95a272eab8 100644 --- a/Documentation/filesystems/vfs.rst +++ b/Documentation/filesystems/vfs.rst @@ -1023,7 +1023,7 @@ This describes how the VFS can manipulate an open fil= e. As of kernel int (*fsync) (struct file *, loff_t, loff_t, int datasync); int (*fasync) (int, struct file *, int); int (*lock) (struct file *, int, struct file_lock *); - unsigned long (*get_unmapped_area)(struct file *, unsigned long, unsigne= d long, unsigned long, unsigned long); + unsigned long (*get_unmapped_area)(struct mm_struct *, struct file *, un= signed long, unsigned long, unsigned long, unsigned long); int (*check_flags)(int); int (*flock) (struct file *, int, struct file_lock *); ssize_t (*splice_write)(struct pipe_inode_info *, struct file *, loff_t = *, size_t, unsigned int); diff --git a/arch/alpha/kernel/osf_sys.c b/arch/alpha/kernel/osf_sys.c index 7b6543d2cca3..80ae6d03e266 100644 --- a/arch/alpha/kernel/osf_sys.c +++ b/arch/alpha/kernel/osf_sys.c @@ -1201,21 +1201,22 @@ SYSCALL_DEFINE1(old_adjtimex, struct timex32 __user= *, txc_p) /* Get an address range which is currently unmapped. */ =20 static unsigned long -arch_get_unmapped_area_1(unsigned long addr, unsigned long len, - unsigned long limit) +arch_get_unmapped_area_1(struct mm_struct *mm, unsigned long addr, + unsigned long len, unsigned long limit) { struct vm_unmapped_area_info info =3D {}; =20 info.length =3D len; info.low_limit =3D addr; info.high_limit =3D limit; - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 unsigned long -arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +arch_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { unsigned long limit =3D TASK_SIZE; =20 @@ -1236,19 +1237,19 @@ arch_get_unmapped_area(struct file *filp, unsigned = long addr, this feature should be incorporated into all ports? */ =20 if (addr) { - addr =3D arch_get_unmapped_area_1 (PAGE_ALIGN(addr), len, limit); + addr =3D arch_get_unmapped_area_1(mm, PAGE_ALIGN(addr), len, limit); if (addr !=3D (unsigned long) -ENOMEM) return addr; } =20 /* Next, try allocating at TASK_UNMAPPED_BASE. */ - addr =3D arch_get_unmapped_area_1 (PAGE_ALIGN(TASK_UNMAPPED_BASE), - len, limit); + addr =3D arch_get_unmapped_area_1(mm, PAGE_ALIGN(TASK_UNMAPPED_BASE), + len, limit); if (addr !=3D (unsigned long) -ENOMEM) return addr; =20 /* Finally, try allocating in low memory. */ - addr =3D arch_get_unmapped_area_1 (PAGE_SIZE, len, limit); + addr =3D arch_get_unmapped_area_1(mm, PAGE_SIZE, len, limit); =20 return addr; } diff --git a/arch/arc/mm/mmap.c b/arch/arc/mm/mmap.c index 2185afe8d59f..d33d5443b616 100644 --- a/arch/arc/mm/mmap.c +++ b/arch/arc/mm/mmap.c @@ -22,11 +22,10 @@ * SHMLBA bytes. */ unsigned long -arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, +arch_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; struct vm_unmapped_area_info info =3D {}; =20 @@ -56,7 +55,7 @@ arch_get_unmapped_area(struct file *filp, unsigned long a= ddr, info.low_limit =3D mm->mmap_base; info.high_limit =3D TASK_SIZE; info.align_offset =3D pgoff << PAGE_SHIFT; - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 static const pgprot_t protection_map[16] =3D { diff --git a/arch/arm/mm/mmap.c b/arch/arm/mm/mmap.c index 3dbb383c26d5..5d1169539ce3 100644 --- a/arch/arm/mm/mmap.c +++ b/arch/arm/mm/mmap.c @@ -27,11 +27,10 @@ * in the VIVT case, we optimise out the alignment rules. */ unsigned long -arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, +arch_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; int do_align =3D 0; int aliasing =3D cache_is_vipt_aliasing(); @@ -74,16 +73,16 @@ arch_get_unmapped_area(struct file *filp, unsigned long= addr, info.high_limit =3D TASK_SIZE; info.align_mask =3D do_align ? (PAGE_MASK & (SHMLBA - 1)) : 0; info.align_offset =3D pgoff << PAGE_SHIFT; - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 unsigned long -arch_get_unmapped_area_topdown(struct file *filp, const unsigned long addr= 0, - const unsigned long len, const unsigned long pgoff, - const unsigned long flags, vm_flags_t vm_flags) +arch_get_unmapped_area_topdown(struct mm_struct *mm, struct file *filp, + const unsigned long addr0, const unsigned long len, + const unsigned long pgoff, const unsigned long flags, + vm_flags_t vm_flags) { struct vm_area_struct *vma; - struct mm_struct *mm =3D current->mm; unsigned long addr =3D addr0; int do_align =3D 0; int aliasing =3D cache_is_vipt_aliasing(); @@ -125,7 +124,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const= unsigned long addr0, info.high_limit =3D mm->mmap_base; info.align_mask =3D do_align ? (PAGE_MASK & (SHMLBA - 1)) : 0; info.align_offset =3D pgoff << PAGE_SHIFT; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); =20 /* * A failed mmap() very likely causes application failure, @@ -138,7 +137,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const= unsigned long addr0, info.flags =3D 0; info.low_limit =3D mm->mmap_base; info.high_limit =3D TASK_SIZE; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); } =20 return addr; diff --git a/arch/csky/abiv1/mmap.c b/arch/csky/abiv1/mmap.c index 1047865e82a9..62eaca7a0d50 100644 --- a/arch/csky/abiv1/mmap.c +++ b/arch/csky/abiv1/mmap.c @@ -22,11 +22,10 @@ * We unconditionally provide this function for all cases. */ unsigned long -arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, +arch_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; int do_align =3D 0; struct vm_unmapped_area_info info =3D { @@ -68,5 +67,5 @@ arch_get_unmapped_area(struct file *filp, unsigned long a= ddr, } =20 info.align_mask =3D do_align ? (PAGE_MASK & (SHMLBA - 1)) : 0; - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } diff --git a/arch/loongarch/mm/mmap.c b/arch/loongarch/mm/mmap.c index 1df9e99582cc..600943121362 100644 --- a/arch/loongarch/mm/mmap.c +++ b/arch/loongarch/mm/mmap.c @@ -18,11 +18,11 @@ =20 enum mmap_allocation_direction {UP, DOWN}; =20 -static unsigned long arch_get_unmapped_area_common(struct file *filp, - unsigned long addr0, unsigned long len, unsigned long pgoff, - unsigned long flags, enum mmap_allocation_direction dir) +static unsigned long arch_get_unmapped_area_common(struct mm_struct *mm, + struct file *filp, unsigned long addr0, unsigned long len, + unsigned long pgoff, unsigned long flags, + enum mmap_allocation_direction dir) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; unsigned long addr =3D addr0; int do_color_align; @@ -74,7 +74,7 @@ static unsigned long arch_get_unmapped_area_common(struct= file *filp, info.flags =3D VM_UNMAPPED_AREA_TOPDOWN; info.low_limit =3D PAGE_SIZE; info.high_limit =3D mm->mmap_base; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); =20 if (!(addr & ~PAGE_MASK)) return addr; @@ -89,14 +89,14 @@ static unsigned long arch_get_unmapped_area_common(stru= ct file *filp, =20 info.low_limit =3D mm->mmap_base; info.high_limit =3D TASK_SIZE; - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 -unsigned long arch_get_unmapped_area(struct file *filp, unsigned long addr= 0, - unsigned long len, unsigned long pgoff, unsigned long flags, - vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area(struct mm_struct *mm, struct file *fi= lp, + unsigned long addr0, unsigned long len, unsigned long pgoff, + unsigned long flags, vm_flags_t vm_flags) { - return arch_get_unmapped_area_common(filp, + return arch_get_unmapped_area_common(mm, filp, addr0, len, pgoff, flags, UP); } =20 @@ -104,11 +104,11 @@ unsigned long arch_get_unmapped_area(struct file *fil= p, unsigned long addr0, * There is no need to export this but sched.h declares the function as * extern so making it static here results in an error. */ -unsigned long arch_get_unmapped_area_topdown(struct file *filp, - unsigned long addr0, unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area_topdown(struct mm_struct *mm, + struct file *filp, unsigned long addr0, unsigned long len, + unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { - return arch_get_unmapped_area_common(filp, + return arch_get_unmapped_area_common(mm, filp, addr0, len, pgoff, flags, DOWN); } =20 diff --git a/arch/mips/mm/mmap.c b/arch/mips/mm/mmap.c index 5d2a1225785b..8905bba30aa0 100644 --- a/arch/mips/mm/mmap.c +++ b/arch/mips/mm/mmap.c @@ -26,11 +26,11 @@ EXPORT_SYMBOL(shm_align_mask); =20 enum mmap_allocation_direction {UP, DOWN}; =20 -static unsigned long arch_get_unmapped_area_common(struct file *filp, - unsigned long addr0, unsigned long len, unsigned long pgoff, - unsigned long flags, enum mmap_allocation_direction dir) +static unsigned long arch_get_unmapped_area_common(struct mm_struct *mm, + struct file *filp, unsigned long addr0, unsigned long len, + unsigned long pgoff, unsigned long flags, + enum mmap_allocation_direction dir) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; unsigned long addr =3D addr0; int do_color_align; @@ -79,7 +79,7 @@ static unsigned long arch_get_unmapped_area_common(struct= file *filp, info.flags =3D VM_UNMAPPED_AREA_TOPDOWN; info.low_limit =3D PAGE_SIZE; info.high_limit =3D mm->mmap_base; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); =20 if (!(addr & ~PAGE_MASK)) return addr; @@ -94,14 +94,14 @@ static unsigned long arch_get_unmapped_area_common(stru= ct file *filp, =20 info.low_limit =3D mm->mmap_base; info.high_limit =3D TASK_SIZE; - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 -unsigned long arch_get_unmapped_area(struct file *filp, unsigned long addr= 0, - unsigned long len, unsigned long pgoff, unsigned long flags, - vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area(struct mm_struct *mm, struct file *fi= lp, + unsigned long addr0, unsigned long len, unsigned long pgoff, + unsigned long flags, vm_flags_t vm_flags) { - return arch_get_unmapped_area_common(filp, + return arch_get_unmapped_area_common(mm, filp, addr0, len, pgoff, flags, UP); } =20 @@ -109,11 +109,11 @@ unsigned long arch_get_unmapped_area(struct file *fil= p, unsigned long addr0, * There is no need to export this but sched.h declares the function as * extern so making it static here results in an error. */ -unsigned long arch_get_unmapped_area_topdown(struct file *filp, - unsigned long addr0, unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area_topdown(struct mm_struct *mm, + struct file *filp, unsigned long addr0, unsigned long len, + unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { - return arch_get_unmapped_area_common(filp, + return arch_get_unmapped_area_common(mm, filp, addr0, len, pgoff, flags, DOWN); } =20 diff --git a/arch/parisc/kernel/sys_parisc.c b/arch/parisc/kernel/sys_paris= c.c index 1a676a9bf80c..b9bf734cf678 100644 --- a/arch/parisc/kernel/sys_parisc.c +++ b/arch/parisc/kernel/sys_parisc.c @@ -100,11 +100,11 @@ unsigned long mmap_upper_limit(const struct rlimit *r= lim_stack) =20 enum mmap_allocation_direction {UP, DOWN}; =20 -static unsigned long arch_get_unmapped_area_common(struct file *filp, - unsigned long addr, unsigned long len, unsigned long pgoff, - unsigned long flags, enum mmap_allocation_direction dir) +static unsigned long arch_get_unmapped_area_common(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + enum mmap_allocation_direction dir) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma, *prev; unsigned long filp_pgoff; int do_color_align; @@ -152,7 +152,7 @@ static unsigned long arch_get_unmapped_area_common(stru= ct file *filp, info.flags =3D VM_UNMAPPED_AREA_TOPDOWN; info.low_limit =3D PAGE_SIZE; info.high_limit =3D mm->mmap_base; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); if (!(addr & ~PAGE_MASK)) return addr; VM_BUG_ON(addr !=3D -ENOMEM); @@ -167,22 +167,22 @@ static unsigned long arch_get_unmapped_area_common(st= ruct file *filp, =20 info.low_limit =3D mm->mmap_base; info.high_limit =3D mmap_upper_limit(NULL); - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 -unsigned long arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, unsigned long flags, - vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area(struct mm_struct *mm, struct file *fi= lp, + unsigned long addr, unsigned long len, unsigned long pgoff, + unsigned long flags, vm_flags_t vm_flags) { - return arch_get_unmapped_area_common(filp, + return arch_get_unmapped_area_common(mm, filp, addr, len, pgoff, flags, UP); } =20 -unsigned long arch_get_unmapped_area_topdown(struct file *filp, - unsigned long addr, unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area_topdown(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { - return arch_get_unmapped_area_common(filp, + return arch_get_unmapped_area_common(mm, filp, addr, len, pgoff, flags, DOWN); } =20 diff --git a/arch/powerpc/include/asm/book3s/64/slice.h b/arch/powerpc/incl= ude/asm/book3s/64/slice.h index 6e2f7a74cd75..d95d9e292443 100644 --- a/arch/powerpc/include/asm/book3s/64/slice.h +++ b/arch/powerpc/include/asm/book3s/64/slice.h @@ -25,7 +25,8 @@ =20 struct mm_struct; =20 -unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long le= n, +unsigned long slice_get_unmapped_area(struct mm_struct *mm, + unsigned long addr, unsigned long len, unsigned long flags, unsigned int psize, int topdown); =20 diff --git a/arch/powerpc/mm/book3s64/slice.c b/arch/powerpc/mm/book3s64/sl= ice.c index 28bec5bc7879..ab8b24f6287d 100644 --- a/arch/powerpc/mm/book3s64/slice.c +++ b/arch/powerpc/mm/book3s64/slice.c @@ -311,7 +311,7 @@ static unsigned long slice_find_area_bottomup(struct mm= _struct *mm, } info.high_limit =3D addr; =20 - found =3D vm_unmapped_area(&info); + found =3D vm_unmapped_area(mm, &info); if (!(found & ~PAGE_MASK)) return found; } @@ -362,7 +362,7 @@ static unsigned long slice_find_area_topdown(struct mm_= struct *mm, } info.low_limit =3D addr; =20 - found =3D vm_unmapped_area(&info); + found =3D vm_unmapped_area(mm, &info); if (!(found & ~PAGE_MASK)) return found; } @@ -422,7 +422,8 @@ static inline void slice_andnot_mask(struct slice_mask = *dst, #define MMU_PAGE_BASE MMU_PAGE_4K #endif =20 -unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long le= n, +unsigned long slice_get_unmapped_area(struct mm_struct *mm, + unsigned long addr, unsigned long len, unsigned long flags, unsigned int psize, int topdown) { @@ -433,7 +434,6 @@ unsigned long slice_get_unmapped_area(unsigned long add= r, unsigned long len, int fixed =3D (flags & MAP_FIXED); int pshift =3D max_t(int, mmu_psize_defs[psize].shift, PAGE_SHIFT); unsigned long page_size =3D 1UL << pshift; - struct mm_struct *mm =3D current->mm; unsigned long newaddr; unsigned long high_limit; =20 @@ -649,7 +649,8 @@ static int file_to_psize(struct file *file) } #endif =20 -unsigned long arch_get_unmapped_area(struct file *filp, +unsigned long arch_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, @@ -659,17 +660,19 @@ unsigned long arch_get_unmapped_area(struct file *fil= p, unsigned int psize; =20 if (radix_enabled()) - return generic_get_unmapped_area(filp, addr, len, pgoff, flags, vm_flags= ); + return generic_get_unmapped_area(mm, filp, addr, len, pgoff, + flags, vm_flags); =20 if (filp && is_file_hugepages(filp)) psize =3D file_to_psize(filp); else - psize =3D mm_ctx_user_psize(¤t->mm->context); + psize =3D mm_ctx_user_psize(&mm->context); =20 - return slice_get_unmapped_area(addr, len, flags, psize, 0); + return slice_get_unmapped_area(mm, addr, len, flags, psize, 0); } =20 -unsigned long arch_get_unmapped_area_topdown(struct file *filp, +unsigned long arch_get_unmapped_area_topdown(struct mm_struct *mm, + struct file *filp, const unsigned long addr0, const unsigned long len, const unsigned long pgoff, @@ -679,14 +682,15 @@ unsigned long arch_get_unmapped_area_topdown(struct f= ile *filp, unsigned int psize; =20 if (radix_enabled()) - return generic_get_unmapped_area_topdown(filp, addr0, len, pgoff, flags,= vm_flags); + return generic_get_unmapped_area_topdown(mm, filp, addr0, len, + pgoff, flags, vm_flags); =20 if (filp && is_file_hugepages(filp)) psize =3D file_to_psize(filp); else - psize =3D mm_ctx_user_psize(¤t->mm->context); + psize =3D mm_ctx_user_psize(&mm->context); =20 - return slice_get_unmapped_area(addr0, len, flags, psize, 1); + return slice_get_unmapped_area(mm, addr0, len, flags, psize, 1); } =20 unsigned int notrace get_slice_psize(struct mm_struct *mm, unsigned long a= ddr) diff --git a/arch/s390/mm/mmap.c b/arch/s390/mm/mmap.c index ef7bfc87758c..92332e9f5229 100644 --- a/arch/s390/mm/mmap.c +++ b/arch/s390/mm/mmap.c @@ -75,11 +75,11 @@ static unsigned long get_align_mask(struct file *filp, = unsigned long flags) return 0; } =20 -unsigned long arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area(struct mm_struct *mm, struct file *fi= lp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; struct vm_unmapped_area_info info =3D {}; =20 @@ -103,7 +103,7 @@ unsigned long arch_get_unmapped_area(struct file *filp,= unsigned long addr, info.align_mask =3D get_align_mask(filp, flags); if (!(filp && is_file_hugepages(filp))) info.align_offset =3D pgoff << PAGE_SHIFT; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); if (offset_in_page(addr)) return addr; =20 @@ -111,12 +111,15 @@ unsigned long arch_get_unmapped_area(struct file *fil= p, unsigned long addr, return check_asce_limit(mm, addr, len); } =20 -unsigned long arch_get_unmapped_area_topdown(struct file *filp, unsigned l= ong addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area_topdown(struct mm_struct *mm, + struct file *filp, + unsigned long addr, + unsigned long len, + unsigned long pgoff, + unsigned long flags, + vm_flags_t vm_flags) { struct vm_area_struct *vma; - struct mm_struct *mm =3D current->mm; struct vm_unmapped_area_info info =3D {}; =20 /* requested length too big for entire address space */ @@ -142,7 +145,7 @@ unsigned long arch_get_unmapped_area_topdown(struct fil= e *filp, unsigned long ad info.align_mask =3D get_align_mask(filp, flags); if (!(filp && is_file_hugepages(filp))) info.align_offset =3D pgoff << PAGE_SHIFT; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); =20 /* * A failed mmap() very likely causes application failure, @@ -155,7 +158,7 @@ unsigned long arch_get_unmapped_area_topdown(struct fil= e *filp, unsigned long ad info.flags =3D 0; info.low_limit =3D TASK_UNMAPPED_BASE; info.high_limit =3D TASK_SIZE; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); if (offset_in_page(addr)) return addr; } diff --git a/arch/sh/mm/mmap.c b/arch/sh/mm/mmap.c index c442734d9b0c..d7a790521539 100644 --- a/arch/sh/mm/mmap.c +++ b/arch/sh/mm/mmap.c @@ -51,11 +51,10 @@ static inline unsigned long COLOUR_ALIGN(unsigned long = addr, return base + off; } =20 -unsigned long arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, unsigned long flags, - vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area(struct mm_struct *mm, struct file *fi= lp, + unsigned long addr, unsigned long len, unsigned long pgoff, + unsigned long flags, vm_flags_t vm_flags) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; int do_colour_align; struct vm_unmapped_area_info info =3D {}; @@ -94,16 +93,16 @@ unsigned long arch_get_unmapped_area(struct file *filp,= unsigned long addr, info.high_limit =3D TASK_SIZE; info.align_mask =3D do_colour_align ? (PAGE_MASK & shm_align_mask) : 0; info.align_offset =3D pgoff << PAGE_SHIFT; - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 unsigned long -arch_get_unmapped_area_topdown(struct file *filp, const unsigned long addr= 0, - const unsigned long len, const unsigned long pgoff, - const unsigned long flags, vm_flags_t vm_flags) +arch_get_unmapped_area_topdown(struct mm_struct *mm, struct file *filp, + const unsigned long addr0, const unsigned long len, + const unsigned long pgoff, const unsigned long flags, + vm_flags_t vm_flags) { struct vm_area_struct *vma; - struct mm_struct *mm =3D current->mm; unsigned long addr =3D addr0; int do_colour_align; struct vm_unmapped_area_info info =3D {}; @@ -144,7 +143,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const= unsigned long addr0, info.high_limit =3D mm->mmap_base; info.align_mask =3D do_colour_align ? (PAGE_MASK & shm_align_mask) : 0; info.align_offset =3D pgoff << PAGE_SHIFT; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); =20 /* * A failed mmap() very likely causes application failure, @@ -157,7 +156,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const= unsigned long addr0, info.flags =3D 0; info.low_limit =3D TASK_UNMAPPED_BASE; info.high_limit =3D TASK_SIZE; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); } =20 return addr; diff --git a/arch/sparc/include/asm/pgtable_64.h b/arch/sparc/include/asm/p= gtable_64.h index 74ede706fb32..b8bed39cddc9 100644 --- a/arch/sparc/include/asm/pgtable_64.h +++ b/arch/sparc/include/asm/pgtable_64.h @@ -1141,9 +1141,9 @@ static inline bool pte_access_permitted(pte_t pte, bo= ol write) /* We provide a special get_unmapped_area for framebuffer mmaps to try and= use * the largest alignment possible such that larget PTEs can be used. */ -unsigned long get_fb_unmapped_area(struct file *filp, unsigned long, - unsigned long, unsigned long, - unsigned long); +unsigned long get_fb_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags); #define HAVE_ARCH_FB_UNMAPPED_AREA =20 void sun4v_register_fault_status(void); diff --git a/arch/sparc/kernel/sys_sparc_32.c b/arch/sparc/kernel/sys_sparc= _32.c index fb31bc0c5b48..930ca7f8a7dc 100644 --- a/arch/sparc/kernel/sys_sparc_32.c +++ b/arch/sparc/kernel/sys_sparc_32.c @@ -40,7 +40,10 @@ SYSCALL_DEFINE0(getpagesize) return PAGE_SIZE; /* Possibly older binaries want 8192 on sun4's? */ } =20 -unsigned long arch_get_unmapped_area(struct file *filp, unsigned long addr= , unsigned long len, unsigned long pgoff, unsigned long flags, vm_flags_t v= m_flags) +unsigned long arch_get_unmapped_area(struct mm_struct *mm, struct file *fi= lp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { struct vm_unmapped_area_info info =3D {}; bool file_hugepage =3D false; @@ -74,7 +77,7 @@ unsigned long arch_get_unmapped_area(struct file *filp, u= nsigned long addr, unsi } else { info.align_mask =3D huge_page_mask_align(filp); } - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 /* diff --git a/arch/sparc/kernel/sys_sparc_64.c b/arch/sparc/kernel/sys_sparc= _64.c index ecefcffcf7b1..e9f5e1b37ee9 100644 --- a/arch/sparc/kernel/sys_sparc_64.c +++ b/arch/sparc/kernel/sys_sparc_64.c @@ -98,9 +98,11 @@ static unsigned long get_align_mask(struct file *filp, u= nsigned long flags) return 0; } =20 -unsigned long arch_get_unmapped_area(struct file *filp, unsigned long addr= , unsigned long len, unsigned long pgoff, unsigned long flags, vm_flags_t v= m_flags) +unsigned long arch_get_unmapped_area(struct mm_struct *mm, struct file *fi= lp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct * vma; unsigned long task_size =3D TASK_SIZE; int do_color_align; @@ -147,25 +149,25 @@ unsigned long arch_get_unmapped_area(struct file *fil= p, unsigned long addr, unsi info.align_mask =3D get_align_mask(filp, flags); if (!file_hugepage) info.align_offset =3D pgoff << PAGE_SHIFT; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); =20 if ((addr & ~PAGE_MASK) && task_size > VA_EXCLUDE_END) { VM_BUG_ON(addr !=3D -ENOMEM); info.low_limit =3D VA_EXCLUDE_END; info.high_limit =3D task_size; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); } =20 return addr; } =20 unsigned long -arch_get_unmapped_area_topdown(struct file *filp, const unsigned long addr= 0, - const unsigned long len, const unsigned long pgoff, - const unsigned long flags, vm_flags_t vm_flags) +arch_get_unmapped_area_topdown(struct mm_struct *mm, struct file *filp, + const unsigned long addr0, const unsigned long len, + const unsigned long pgoff, const unsigned long flags, + vm_flags_t vm_flags) { struct vm_area_struct *vma; - struct mm_struct *mm =3D current->mm; unsigned long task_size =3D STACK_TOP32; unsigned long addr =3D addr0; int do_color_align; @@ -215,7 +217,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const= unsigned long addr0, info.align_mask =3D get_align_mask(filp, flags); if (!file_hugepage) info.align_offset =3D pgoff << PAGE_SHIFT; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); =20 /* * A failed mmap() very likely causes application failure, @@ -228,20 +230,22 @@ arch_get_unmapped_area_topdown(struct file *filp, con= st unsigned long addr0, info.flags =3D 0; info.low_limit =3D TASK_UNMAPPED_BASE; info.high_limit =3D STACK_TOP32; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); } =20 return addr; } =20 /* Try to align mapping such that we align it as much as possible. */ -unsigned long get_fb_unmapped_area(struct file *filp, unsigned long orig_a= ddr, unsigned long len, unsigned long pgoff, unsigned long flags) +unsigned long get_fb_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long orig_addr, unsigned long len, + unsigned long pgoff, unsigned long flags) { unsigned long align_goal, addr =3D -ENOMEM; =20 if (flags & MAP_FIXED) { /* Ok, don't mess with it. */ - return mm_get_unmapped_area(NULL, orig_addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, NULL, orig_addr, len, pgoff, flags); } flags &=3D ~MAP_SHARED; =20 @@ -254,7 +258,7 @@ unsigned long get_fb_unmapped_area(struct file *filp, u= nsigned long orig_addr, u align_goal =3D (64UL * 1024); =20 do { - addr =3D mm_get_unmapped_area(NULL, orig_addr, + addr =3D mm_get_unmapped_area(mm, NULL, orig_addr, len + (align_goal - PAGE_SIZE), pgoff, flags); if (!(addr & ~PAGE_MASK)) { addr =3D (addr + (align_goal - 1UL)) & ~(align_goal - 1UL); @@ -273,7 +277,7 @@ unsigned long get_fb_unmapped_area(struct file *filp, u= nsigned long orig_addr, u * be obtained. */ if (addr & ~PAGE_MASK) - addr =3D mm_get_unmapped_area(NULL, orig_addr, len, pgoff, flags); + addr =3D mm_get_unmapped_area(mm, NULL, orig_addr, len, pgoff, flags); =20 return addr; } diff --git a/arch/x86/include/asm/elf.h b/arch/x86/include/asm/elf.h index 0de9df759c99..12c66609889e 100644 --- a/arch/x86/include/asm/elf.h +++ b/arch/x86/include/asm/elf.h @@ -307,7 +307,7 @@ static inline int mmap_is_ia32(void) =20 extern unsigned long task_size_32bit(void); extern unsigned long task_size_64bit(int full_addr_space); -extern unsigned long get_mmap_base(int is_legacy); +extern unsigned long get_mmap_base(struct mm_struct *mm, int is_legacy); extern bool mmap_address_hint_valid(unsigned long addr, unsigned long len); extern unsigned long get_sigframe_size(void); =20 diff --git a/arch/x86/kernel/cpu/sgx/driver.c b/arch/x86/kernel/cpu/sgx/dri= ver.c index 9268289cd9f9..f81652086909 100644 --- a/arch/x86/kernel/cpu/sgx/driver.c +++ b/arch/x86/kernel/cpu/sgx/driver.c @@ -121,11 +121,12 @@ static int sgx_mmap(struct file *file, struct vm_area= _struct *vma) return 0; } =20 -static unsigned long sgx_get_unmapped_area(struct file *file, - unsigned long addr, - unsigned long len, - unsigned long pgoff, - unsigned long flags) +static unsigned long sgx_get_unmapped_area(struct mm_struct *mm, + struct file *file, + unsigned long addr, + unsigned long len, + unsigned long pgoff, + unsigned long flags) { if ((flags & MAP_TYPE) =3D=3D MAP_PRIVATE) return -EINVAL; @@ -133,7 +134,7 @@ static unsigned long sgx_get_unmapped_area(struct file = *file, if (flags & MAP_FIXED) return addr; =20 - return mm_get_unmapped_area(file, addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, file, addr, len, pgoff, flags); } =20 #ifdef CONFIG_COMPAT diff --git a/arch/x86/kernel/sys_x86_64.c b/arch/x86/kernel/sys_x86_64.c index 776ae6fa7f2d..d9c4353913a4 100644 --- a/arch/x86/kernel/sys_x86_64.c +++ b/arch/x86/kernel/sys_x86_64.c @@ -89,8 +89,9 @@ SYSCALL_DEFINE6(mmap, unsigned long, addr, unsigned long,= len, return ksys_mmap_pgoff(addr, len, prot, flags, fd, off >> PAGE_SHIFT); } =20 -static void find_start_end(unsigned long addr, unsigned long flags, - unsigned long *begin, unsigned long *end) +static void find_start_end(struct mm_struct *mm, unsigned long addr, + unsigned long flags, unsigned long *begin, + unsigned long *end) { if (!in_32bit_syscall() && (flags & MAP_32BIT)) { /* This is usually used needed to map code in small @@ -108,7 +109,7 @@ static void find_start_end(unsigned long addr, unsigned= long flags, return; } =20 - *begin =3D get_mmap_base(1); + *begin =3D get_mmap_base(mm, 1); if (in_32bit_syscall()) *end =3D task_size_32bit(); else @@ -124,10 +125,11 @@ static inline unsigned long stack_guard_placement(vm_= flags_t vm_flags) } =20 unsigned long -arch_get_unmapped_area(struct file *filp, unsigned long addr, unsigned lon= g len, - unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) +arch_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma; struct vm_unmapped_area_info info =3D {}; unsigned long begin, end; @@ -135,7 +137,7 @@ arch_get_unmapped_area(struct file *filp, unsigned long= addr, unsigned long len, if (flags & MAP_FIXED) return addr; =20 - find_start_end(addr, flags, &begin, &end); + find_start_end(mm, addr, flags, &begin, &end); =20 if (len > end) return -ENOMEM; @@ -160,16 +162,16 @@ arch_get_unmapped_area(struct file *filp, unsigned lo= ng addr, unsigned long len, info.align_offset +=3D get_align_bits(); } =20 - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 unsigned long -arch_get_unmapped_area_topdown(struct file *filp, unsigned long addr0, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +arch_get_unmapped_area_topdown(struct mm_struct *mm, struct file *filp, + unsigned long addr0, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { struct vm_area_struct *vma; - struct mm_struct *mm =3D current->mm; unsigned long addr =3D addr0; struct vm_unmapped_area_info info =3D {}; =20 @@ -204,7 +206,7 @@ arch_get_unmapped_area_topdown(struct file *filp, unsig= ned long addr0, else info.low_limit =3D PAGE_SIZE; =20 - info.high_limit =3D get_mmap_base(0); + info.high_limit =3D get_mmap_base(mm, 0); if (!(filp && is_file_hugepages(filp))) { info.start_gap =3D stack_guard_placement(vm_flags); info.align_offset =3D pgoff << PAGE_SHIFT; @@ -224,7 +226,7 @@ arch_get_unmapped_area_topdown(struct file *filp, unsig= ned long addr0, info.align_mask =3D get_align_mask(filp); info.align_offset +=3D get_align_bits(); } - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); if (!(addr & ~PAGE_MASK)) return addr; VM_BUG_ON(addr !=3D -ENOMEM); @@ -236,5 +238,5 @@ arch_get_unmapped_area_topdown(struct file *filp, unsig= ned long addr0, * can happen with large stack limits and large mmap() * allocations. */ - return arch_get_unmapped_area(filp, addr0, len, pgoff, flags, 0); + return arch_get_unmapped_area(mm, filp, addr0, len, pgoff, flags, 0); } diff --git a/arch/x86/kernel/uprobes.c b/arch/x86/kernel/uprobes.c index 3af979fb41d3..417153d02c54 100644 --- a/arch/x86/kernel/uprobes.c +++ b/arch/x86/kernel/uprobes.c @@ -661,13 +661,13 @@ static unsigned long find_nearest_trampoline(unsigned= long vaddr) /* Search up from the caller address. */ info.low_limit =3D call_end; info.high_limit =3D min(high_limit, TASK_SIZE); - high_tramp =3D vm_unmapped_area(&info); + high_tramp =3D vm_unmapped_area(current->mm, &info); =20 /* Search down from the caller address. */ info.low_limit =3D max(low_limit, PAGE_SIZE); info.high_limit =3D call_end; info.flags =3D VM_UNMAPPED_AREA_TOPDOWN; - low_tramp =3D vm_unmapped_area(&info); + low_tramp =3D vm_unmapped_area(current->mm, &info); =20 if (IS_ERR_VALUE(high_tramp) && IS_ERR_VALUE(low_tramp)) return -ENOMEM; diff --git a/arch/x86/mm/mmap.c b/arch/x86/mm/mmap.c index 82f3a987f7cf..75c864f3c0d8 100644 --- a/arch/x86/mm/mmap.c +++ b/arch/x86/mm/mmap.c @@ -143,10 +143,8 @@ void arch_pick_mmap_layout(struct mm_struct *mm, const= struct rlimit *rlim_stack #endif } =20 -unsigned long get_mmap_base(int is_legacy) +unsigned long get_mmap_base(struct mm_struct *mm, int is_legacy) { - struct mm_struct *mm =3D current->mm; - #ifdef CONFIG_HAVE_ARCH_COMPAT_MMAP_BASES if (in_32bit_syscall()) { return is_legacy ? mm->mmap_compat_legacy_base diff --git a/arch/xtensa/kernel/syscall.c b/arch/xtensa/kernel/syscall.c index dc54f854c2f5..51bfe06746d7 100644 --- a/arch/xtensa/kernel/syscall.c +++ b/arch/xtensa/kernel/syscall.c @@ -54,9 +54,9 @@ asmlinkage long xtensa_fadvise64_64(int fd, int advice, } =20 #ifdef CONFIG_MMU -unsigned long arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, unsigned long flags, - vm_flags_t vm_flags) +unsigned long arch_get_unmapped_area(struct mm_struct *mm, struct file *fi= lp, + unsigned long addr, unsigned long len, unsigned long pgoff, + unsigned long flags, vm_flags_t vm_flags) { struct vm_area_struct *vmm; struct vma_iterator vmi; @@ -81,7 +81,7 @@ unsigned long arch_get_unmapped_area(struct file *filp, u= nsigned long addr, else addr =3D PAGE_ALIGN(addr); =20 - vma_iter_init(&vmi, current->mm, addr); + vma_iter_init(&vmi, mm, addr); for_each_vma(vmi, vmm) { /* At this point: (addr < vmm->vm_end). */ if (addr + len <=3D vm_start_gap(vmm)) diff --git a/drivers/char/mem.c b/drivers/char/mem.c index 63253d1de5d7..204bf406173d 100644 --- a/drivers/char/mem.c +++ b/drivers/char/mem.c @@ -280,7 +280,8 @@ static pgprot_t phys_mem_access_prot(struct file *file,= unsigned long pfn, #endif =20 #ifndef CONFIG_MMU -static unsigned long get_unmapped_area_mem(struct file *file, +static unsigned long get_unmapped_area_mem(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, @@ -515,14 +516,16 @@ static int mmap_zero_prepare(struct vm_area_desc *des= c) } =20 #ifndef CONFIG_MMU -static unsigned long get_unmapped_area_zero(struct file *file, +static unsigned long get_unmapped_area_zero(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { return -ENOSYS; } #else -static unsigned long get_unmapped_area_zero(struct file *file, +static unsigned long get_unmapped_area_zero(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { @@ -534,7 +537,7 @@ static unsigned long get_unmapped_area_zero(struct file= *file, * get_unmapped_area(), so as not to confuse shmem with our * handle on "/dev/zero". */ - return shmem_get_unmapped_area(NULL, addr, len, pgoff, flags); + return shmem_get_unmapped_area(mm, NULL, addr, len, pgoff, flags); } =20 /* @@ -543,9 +546,9 @@ static unsigned long get_unmapped_area_zero(struct file= *file, * fall back to system page size mappings. */ #ifdef CONFIG_TRANSPARENT_HUGEPAGE - return thp_get_unmapped_area(file, addr, len, pgoff, flags); + return thp_get_unmapped_area(mm, file, addr, len, pgoff, flags); #else - return mm_get_unmapped_area(file, addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, file, addr, len, pgoff, flags); #endif } #endif /* CONFIG_MMU */ diff --git a/drivers/dax/device.c b/drivers/dax/device.c index d0c9b4e03b47..80df897b6535 100644 --- a/drivers/dax/device.c +++ b/drivers/dax/device.c @@ -295,9 +295,9 @@ static int dax_mmap_prepare(struct vm_area_desc *desc) } =20 /* return an unmapped area aligned to the dax region specified alignment */ -static unsigned long dax_get_unmapped_area(struct file *filp, - unsigned long addr, unsigned long len, unsigned long pgoff, - unsigned long flags) +static unsigned long dax_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags) { unsigned long off, off_end, off_align, len_align, addr_align, align; struct dev_dax *dev_dax =3D filp ? filp->private_data : NULL; @@ -317,13 +317,14 @@ static unsigned long dax_get_unmapped_area(struct fil= e *filp, if ((off + len_align) < off) goto out; =20 - addr_align =3D mm_get_unmapped_area(filp, addr, len_align, pgoff, flags); + addr_align =3D mm_get_unmapped_area(mm, filp, addr, len_align, pgoff, + flags); if (!IS_ERR_VALUE(addr_align)) { addr_align +=3D (off - addr_align) & (align - 1); return addr_align; } out: - return mm_get_unmapped_area(filp, addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, filp, addr, len, pgoff, flags); } =20 static const struct address_space_operations dev_dax_aops =3D { diff --git a/drivers/gpu/drm/drm_gem.c b/drivers/gpu/drm/drm_gem.c index e3ed684ddcf2..f3d4f71e61d4 100644 --- a/drivers/gpu/drm/drm_gem.c +++ b/drivers/gpu/drm/drm_gem.c @@ -1316,6 +1316,7 @@ drm_gem_object_lookup_at_offset(struct file *filp, un= signed long start, #ifdef CONFIG_MMU /** * drm_gem_get_unmapped_area - get memory mapping region routine for GEM o= bjects + * @mm: mm_struct the mapping is placed in * @filp: DRM file pointer * @uaddr: User address hint * @len: Mapping length @@ -1336,7 +1337,8 @@ drm_gem_object_lookup_at_offset(struct file *filp, un= signed long start, * If a GEM object is not available at the given offset or if the caller i= s not * granted access to it, fall back to mm_get_unmapped_area(). */ -unsigned long drm_gem_get_unmapped_area(struct file *filp, unsigned long u= addr, +unsigned long drm_gem_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long uaddr, unsigned long len, unsigned long pgoff, unsigned long flags) { @@ -1348,9 +1350,10 @@ unsigned long drm_gem_get_unmapped_area(struct file = *filp, unsigned long uaddr, obj =3D NULL; =20 if (!obj || !obj->filp || !obj->filp->f_op->get_unmapped_area) - ret =3D mm_get_unmapped_area(filp, uaddr, len, 0, flags); + ret =3D mm_get_unmapped_area(mm, filp, uaddr, len, 0, flags); else - ret =3D obj->filp->f_op->get_unmapped_area(obj->filp, uaddr, len, 0, fla= gs); + ret =3D obj->filp->f_op->get_unmapped_area(mm, obj->filp, uaddr, + len, 0, flags); =20 drm_gem_object_put(obj); =20 diff --git a/drivers/gpu/drm/drm_gem_dma_helper.c b/drivers/gpu/drm/drm_gem= _dma_helper.c index 1c00a71ab3c9..59f517a4f283 100644 --- a/drivers/gpu/drm/drm_gem_dma_helper.c +++ b/drivers/gpu/drm/drm_gem_dma_helper.c @@ -330,6 +330,7 @@ EXPORT_SYMBOL_GPL(drm_gem_dma_vm_ops); #ifndef CONFIG_MMU /** * drm_gem_dma_get_unmapped_area - propose address for mapping in noMMU ca= ses + * @mm: mm_struct the mapping is placed in * @filp: file object * @addr: memory address * @len: buffer size @@ -344,7 +345,8 @@ EXPORT_SYMBOL_GPL(drm_gem_dma_vm_ops); * Returns: * mapping address on success or a negative error code on failure. */ -unsigned long drm_gem_dma_get_unmapped_area(struct file *filp, +unsigned long drm_gem_dma_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, diff --git a/drivers/media/v4l2-core/v4l2-dev.c b/drivers/media/v4l2-core/v= 4l2-dev.c index 5516b2bbb08f..1cf7c36fc945 100644 --- a/drivers/media/v4l2-core/v4l2-dev.c +++ b/drivers/media/v4l2-core/v4l2-dev.c @@ -373,7 +373,8 @@ static long v4l2_ioctl(struct file *filp, unsigned int = cmd, unsigned long arg) #ifdef CONFIG_MMU #define v4l2_get_unmapped_area NULL #else -static unsigned long v4l2_get_unmapped_area(struct file *filp, +static unsigned long v4l2_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { diff --git a/drivers/mtd/mtdchar.c b/drivers/mtd/mtdchar.c index bf01e6ac7293..e208a8de8716 100644 --- a/drivers/mtd/mtdchar.c +++ b/drivers/mtd/mtdchar.c @@ -1340,7 +1340,8 @@ static long mtdchar_compat_ioctl(struct file *file, u= nsigned int cmd, * mappings) */ #ifndef CONFIG_MMU -static unsigned long mtdchar_get_unmapped_area(struct file *file, +static unsigned long mtdchar_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, diff --git a/drivers/video/fbdev/core/fb_chrdev.c b/drivers/video/fbdev/cor= e/fb_chrdev.c index ba1d0bc214c5..cf8590b64b54 100644 --- a/drivers/video/fbdev/core/fb_chrdev.c +++ b/drivers/video/fbdev/core/fb_chrdev.c @@ -383,7 +383,8 @@ __releases(&info->lock) } =20 #if defined(CONFIG_FB_PROVIDE_GET_FB_UNMAPPED_AREA) && !defined(CONFIG_MMU) -static unsigned long get_fb_unmapped_area(struct file *filp, +static unsigned long get_fb_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { diff --git a/fs/cramfs/inode.c b/fs/cramfs/inode.c index 4edbfccd0bbe..d3aee6705b8d 100644 --- a/fs/cramfs/inode.c +++ b/fs/cramfs/inode.c @@ -449,7 +449,8 @@ static int cramfs_physmem_mmap(struct file *file, struc= t vm_area_struct *vma) return is_nommu_shared_mapping(vma->vm_flags) ? 0 : -ENOSYS; } =20 -static unsigned long cramfs_physmem_get_unmapped_area(struct file *file, +static unsigned long cramfs_physmem_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c index 216e1a0dd0b2..f5d3c182cc7c 100644 --- a/fs/hugetlbfs/inode.c +++ b/fs/hugetlbfs/inode.c @@ -170,9 +170,9 @@ static int hugetlbfs_file_mmap(struct file *file, struc= t vm_area_struct *vma) */ =20 unsigned long -hugetlb_get_unmapped_area(struct file *file, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags) +hugetlb_get_unmapped_area(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags) { unsigned long addr0 =3D 0; struct hstate *h =3D hstate_file(file); @@ -184,7 +184,8 @@ hugetlb_get_unmapped_area(struct file *file, unsigned l= ong addr, if (addr) addr0 =3D ALIGN(addr, huge_page_size(h)); =20 - return mm_get_unmapped_area_vmflags(file, addr0, len, pgoff, flags, 0); + return mm_get_unmapped_area_vmflags(mm, file, addr0, len, pgoff, + flags, 0); } =20 /* diff --git a/fs/proc/inode.c b/fs/proc/inode.c index b7634f975d98..3e542fe3118d 100644 --- a/fs/proc/inode.c +++ b/fs/proc/inode.c @@ -435,22 +435,25 @@ static int proc_reg_mmap(struct file *file, struct vm= _area_struct *vma) } =20 static unsigned long -pde_get_unmapped_area(struct proc_dir_entry *pde, struct file *file, unsig= ned long orig_addr, +pde_get_unmapped_area(struct proc_dir_entry *pde, struct mm_struct *mm, + struct file *file, unsigned long orig_addr, unsigned long len, unsigned long pgoff, unsigned long flags) { if (pde->proc_ops->proc_get_unmapped_area) - return pde->proc_ops->proc_get_unmapped_area(file, orig_addr, len, pgoff= , flags); + return pde->proc_ops->proc_get_unmapped_area(mm, file, + orig_addr, len, pgoff, flags); =20 #ifdef CONFIG_MMU - return mm_get_unmapped_area(file, orig_addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, file, orig_addr, len, pgoff, flags); #endif =20 return orig_addr; } =20 static unsigned long -proc_reg_get_unmapped_area(struct file *file, unsigned long orig_addr, +proc_reg_get_unmapped_area(struct mm_struct *mm, struct file *file, + unsigned long orig_addr, unsigned long len, unsigned long pgoff, unsigned long flags) { @@ -458,9 +461,9 @@ proc_reg_get_unmapped_area(struct file *file, unsigned = long orig_addr, unsigned long rv =3D -EIO; =20 if (pde_is_permanent(pde)) { - return pde_get_unmapped_area(pde, file, orig_addr, len, pgoff, flags); + return pde_get_unmapped_area(pde, mm, file, orig_addr, len, pgoff, flags= ); } else if (use_pde(pde)) { - rv =3D pde_get_unmapped_area(pde, file, orig_addr, len, pgoff, flags); + rv =3D pde_get_unmapped_area(pde, mm, file, orig_addr, len, pgoff, flags= ); unuse_pde(pde); } return rv; diff --git a/fs/ramfs/file-mmu.c b/fs/ramfs/file-mmu.c index c3ed1c5117b2..e070ca06499b 100644 --- a/fs/ramfs/file-mmu.c +++ b/fs/ramfs/file-mmu.c @@ -31,11 +31,12 @@ =20 #include "internal.h" =20 -static unsigned long ramfs_mmu_get_unmapped_area(struct file *file, +static unsigned long ramfs_mmu_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { - return mm_get_unmapped_area(file, addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, file, addr, len, pgoff, flags); } =20 const struct file_operations ramfs_file_operations =3D { diff --git a/fs/ramfs/file-nommu.c b/fs/ramfs/file-nommu.c index 2f79bcb89d2e..2e20fc9c0d37 100644 --- a/fs/ramfs/file-nommu.c +++ b/fs/ramfs/file-nommu.c @@ -23,7 +23,8 @@ #include "internal.h" =20 static int ramfs_nommu_setattr(struct mnt_idmap *, struct dentry *, struct= iattr *); -static unsigned long ramfs_nommu_get_unmapped_area(struct file *file, +static unsigned long ramfs_nommu_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, @@ -199,7 +200,8 @@ static int ramfs_nommu_setattr(struct mnt_idmap *idmap, * - the pages to be mapped must exist * - the pages be physically contiguous in sequence */ -static unsigned long ramfs_nommu_get_unmapped_area(struct file *file, +static unsigned long ramfs_nommu_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { diff --git a/fs/romfs/mmap-nommu.c b/fs/romfs/mmap-nommu.c index 7c3a1a7fecee..18c1e8545b0d 100644 --- a/fs/romfs/mmap-nommu.c +++ b/fs/romfs/mmap-nommu.c @@ -15,7 +15,8 @@ * mappings) * - attempts to map through to the underlying MTD device */ -static unsigned long romfs_get_unmapped_area(struct file *file, +static unsigned long romfs_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, diff --git a/include/drm/drm_gem.h b/include/drm/drm_gem.h index 8a704f6a65c1..a99d525f0369 100644 --- a/include/drm/drm_gem.h +++ b/include/drm/drm_gem.h @@ -536,7 +536,8 @@ int drm_gem_mmap_obj(struct drm_gem_object *obj, unsign= ed long obj_size, int drm_gem_mmap(struct file *filp, struct vm_area_struct *vma); =20 #ifdef CONFIG_MMU -unsigned long drm_gem_get_unmapped_area(struct file *filp, unsigned long u= addr, +unsigned long drm_gem_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long uaddr, unsigned long len, unsigned long pgoff, unsigned long flags); #else diff --git a/include/drm/drm_gem_dma_helper.h b/include/drm/drm_gem_dma_hel= per.h index f2678e7ecb98..127f52f192a9 100644 --- a/include/drm/drm_gem_dma_helper.h +++ b/include/drm/drm_gem_dma_helper.h @@ -232,7 +232,8 @@ drm_gem_dma_prime_import_sg_table_vmap(struct drm_devic= e *drm, */ =20 #ifndef CONFIG_MMU -unsigned long drm_gem_dma_get_unmapped_area(struct file *filp, +unsigned long drm_gem_dma_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, diff --git a/include/linux/bpf.h b/include/linux/bpf.h index 7719f6528445..d72a5ba49654 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -147,7 +147,8 @@ struct bpf_map_ops { int (*map_mmap)(struct bpf_map *map, struct vm_area_struct *vma); __poll_t (*map_poll)(struct bpf_map *map, struct file *filp, struct poll_table_struct *pts); - unsigned long (*map_get_unmapped_area)(struct file *filep, unsigned long = addr, + unsigned long (*map_get_unmapped_area)(struct mm_struct *mm, + struct file *filep, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags); =20 diff --git a/include/linux/fs.h b/include/linux/fs.h index 50ce731a2b78..58e00b320efd 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -1939,7 +1939,10 @@ struct file_operations { int (*fsync) (struct file *, loff_t, loff_t, int datasync); int (*fasync) (int, struct file *, int); int (*lock) (struct file *, int, struct file_lock *); - unsigned long (*get_unmapped_area)(struct file *, unsigned long, unsigned= long, unsigned long, unsigned long); + unsigned long (*get_unmapped_area)(struct mm_struct *mm, + struct file *file, unsigned long addr, + unsigned long len, unsigned long pgoff, + unsigned long flags); int (*check_flags)(int); int (*flock) (struct file *, int, struct file_lock *); ssize_t (*splice_write)(struct pipe_inode_info *, struct file *, loff_t *= , size_t, unsigned int); diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h index ad20f7f8c179..5760710c3c5d 100644 --- a/include/linux/huge_mm.h +++ b/include/linux/huge_mm.h @@ -388,11 +388,12 @@ static inline bool thp_disabled_by_hw(void) return transparent_hugepage_flags & (1 << TRANSPARENT_HUGEPAGE_UNSUPPORTE= D); } =20 -unsigned long thp_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, unsigned long flags); -unsigned long thp_get_unmapped_area_vmflags(struct file *filp, unsigned lo= ng addr, - unsigned long len, unsigned long pgoff, unsigned long flags, - vm_flags_t vm_flags); +unsigned long thp_get_unmapped_area(struct mm_struct *mm, struct file *fil= p, + unsigned long addr, unsigned long len, unsigned long pgoff, + unsigned long flags); +unsigned long thp_get_unmapped_area_vmflags(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags); =20 enum split_type { SPLIT_TYPE_UNIFORM, @@ -614,9 +615,10 @@ static inline unsigned long thp_vma_allowable_orders(s= truct vm_area_struct *vma, #define thp_get_unmapped_area NULL =20 static inline unsigned long -thp_get_unmapped_area_vmflags(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +thp_get_unmapped_area_vmflags(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { return 0; } diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h index 2abaf99321e9..b2743931f927 100644 --- a/include/linux/hugetlb.h +++ b/include/linux/hugetlb.h @@ -543,9 +543,9 @@ static inline struct hstate *hstate_inode(struct inode = *i) #endif /* !CONFIG_HUGETLBFS */ =20 unsigned long -hugetlb_get_unmapped_area(struct file *file, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags); +hugetlb_get_unmapped_area(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags); =20 /* * huegtlb page specific state flags. These flags are located in page.pri= vate diff --git a/include/linux/mm.h b/include/linux/mm.h index 485df9c2dbdd..560bc369a54c 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -4138,14 +4138,17 @@ unsigned long randomize_stack_top(unsigned long sta= ck_top); unsigned long randomize_page(unsigned long start, unsigned long range); =20 unsigned long -__get_unmapped_area(struct file *file, unsigned long addr, unsigned long l= en, - unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags); +__get_unmapped_area(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags); =20 static inline unsigned long get_unmapped_area(struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { - return __get_unmapped_area(file, addr, len, pgoff, flags, 0); + return __get_unmapped_area(current->mm, file, addr, len, pgoff, + flags, 0); } =20 extern unsigned long do_mmap(struct file *file, unsigned long addr, @@ -4194,7 +4197,8 @@ struct vm_unmapped_area_info { unsigned long start_gap; }; =20 -extern unsigned long vm_unmapped_area(struct vm_unmapped_area_info *info); +extern unsigned long vm_unmapped_area(struct mm_struct *mm, + struct vm_unmapped_area_info *info); =20 /* truncate.c */ void truncate_inode_pages(struct address_space *mapping, loff_t lstart); diff --git a/include/linux/proc_fs.h b/include/linux/proc_fs.h index 47d7deaeed8f..e7998a303519 100644 --- a/include/linux/proc_fs.h +++ b/include/linux/proc_fs.h @@ -47,7 +47,10 @@ struct proc_ops { long (*proc_compat_ioctl)(struct file *, unsigned int, unsigned long); #endif int (*proc_mmap)(struct file *, struct vm_area_struct *); - unsigned long (*proc_get_unmapped_area)(struct file *, unsigned long, uns= igned long, unsigned long, unsigned long); + unsigned long (*proc_get_unmapped_area)(struct mm_struct *mm, + struct file *file, unsigned long addr, + unsigned long len, unsigned long pgoff, + unsigned long flags); } __randomize_layout; =20 /* definitions for hide_pid field */ diff --git a/include/linux/sched/mm.h b/include/linux/sched/mm.h index 95d0040df584..adc11009fedf 100644 --- a/include/linux/sched/mm.h +++ b/include/linux/sched/mm.h @@ -181,19 +181,22 @@ extern void arch_pick_mmap_layout(struct mm_struct *m= m, const struct rlimit *rlim_stack); =20 unsigned long -arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags); +arch_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags); unsigned long -arch_get_unmapped_area_topdown(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t); +arch_get_unmapped_area_topdown(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t); =20 -unsigned long mm_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags); +unsigned long mm_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags); =20 -unsigned long mm_get_unmapped_area_vmflags(struct file *filp, +unsigned long mm_get_unmapped_area_vmflags(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, @@ -201,13 +204,15 @@ unsigned long mm_get_unmapped_area_vmflags(struct fil= e *filp, vm_flags_t vm_flags); =20 unsigned long -generic_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags); +generic_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags); unsigned long -generic_get_unmapped_area_topdown(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags); +generic_get_unmapped_area_topdown(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags); #else static inline void arch_pick_mmap_layout(struct mm_struct *mm, const struct rlimit *rlim_stack) {} diff --git a/include/linux/shmem_fs.h b/include/linux/shmem_fs.h index e729b9b0e38d..3912836b375f 100644 --- a/include/linux/shmem_fs.h +++ b/include/linux/shmem_fs.h @@ -109,8 +109,9 @@ extern struct file *shmem_file_setup_with_mnt(struct vf= smount *mnt, const char *name, loff_t size, vma_flags_t flags); int shmem_zero_setup(struct vm_area_struct *vma); int shmem_zero_setup_desc(struct vm_area_desc *desc); -extern unsigned long shmem_get_unmapped_area(struct file *, unsigned long = addr, - unsigned long len, unsigned long pgoff, unsigned long flags); +extern unsigned long shmem_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags); extern int shmem_lock(struct file *file, int lock, struct ucounts *ucounts= ); #ifdef CONFIG_SHMEM bool shmem_mapping(const struct address_space *mapping); diff --git a/io_uring/memmap.c b/io_uring/memmap.c index 23e8a85111bc..255fd80fd2f2 100644 --- a/io_uring/memmap.c +++ b/io_uring/memmap.c @@ -318,7 +318,8 @@ __cold int io_uring_mmap(struct file *file, struct vm_a= rea_struct *vma) return io_region_mmap(ctx, region, vma, page_limit); } =20 -unsigned long io_uring_get_unmapped_area(struct file *filp, unsigned long = addr, +unsigned long io_uring_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { @@ -361,7 +362,7 @@ unsigned long io_uring_get_unmapped_area(struct file *f= ilp, unsigned long addr, #else addr =3D 0UL; #endif - return mm_get_unmapped_area(filp, addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, filp, addr, len, pgoff, flags); } =20 #else /* !CONFIG_MMU */ @@ -420,7 +421,8 @@ unsigned int io_uring_nommu_mmap_capabilities(struct fi= le *file) return NOMMU_MAP_DIRECT | NOMMU_MAP_READ | NOMMU_MAP_WRITE; } =20 -unsigned long io_uring_get_unmapped_area(struct file *file, unsigned long = addr, +unsigned long io_uring_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { diff --git a/io_uring/memmap.h b/io_uring/memmap.h index f4cfbb6b9a1f..40104ed26dbe 100644 --- a/io_uring/memmap.h +++ b/io_uring/memmap.h @@ -12,7 +12,8 @@ struct page **io_pin_pages(unsigned long uaddr, unsigned = long len, int *npages); #ifndef CONFIG_MMU unsigned int io_uring_nommu_mmap_capabilities(struct file *file); #endif -unsigned long io_uring_get_unmapped_area(struct file *file, unsigned long = addr, +unsigned long io_uring_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags); int io_uring_mmap(struct file *file, struct vm_area_struct *vma); diff --git a/ipc/shm.c b/ipc/shm.c index b3e8a58e177d..c49e1461ff9c 100644 --- a/ipc/shm.c +++ b/ipc/shm.c @@ -649,13 +649,13 @@ static long shm_fallocate(struct file *file, int mode= , loff_t offset, return sfd->file->f_op->fallocate(file, mode, offset, len); } =20 -static unsigned long shm_get_unmapped_area(struct file *file, - unsigned long addr, unsigned long len, unsigned long pgoff, - unsigned long flags) +static unsigned long shm_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags) { struct shm_file_data *sfd =3D shm_file_data(file); =20 - return sfd->file->f_op->get_unmapped_area(sfd->file, addr, len, + return sfd->file->f_op->get_unmapped_area(mm, sfd->file, addr, len, pgoff, flags); } =20 diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c index 80b7b8a69446..839ebaae301d 100644 --- a/kernel/bpf/arena.c +++ b/kernel/bpf/arena.c @@ -543,7 +543,8 @@ static const struct vm_operations_struct arena_vm_ops = =3D { .fault =3D arena_vm_fault, }; =20 -static unsigned long arena_get_unmapped_area(struct file *filp, unsigned l= ong addr, +static unsigned long arena_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { @@ -566,7 +567,7 @@ static unsigned long arena_get_unmapped_area(struct fil= e *filp, unsigned long ad return -EINVAL; } =20 - ret =3D mm_get_unmapped_area(filp, addr, len * 2, 0, flags); + ret =3D mm_get_unmapped_area(mm, filp, addr, len * 2, 0, flags); if (IS_ERR_VALUE(ret)) return ret; if ((ret >> 32) =3D=3D ((ret + len - 1) >> 32)) diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c index 6db306d23b47..5092fc8707c5 100644 --- a/kernel/bpf/syscall.c +++ b/kernel/bpf/syscall.c @@ -1150,16 +1150,18 @@ static __poll_t bpf_map_poll(struct file *filp, str= uct poll_table_struct *pts) return EPOLLERR; } =20 -static unsigned long bpf_get_unmapped_area(struct file *filp, unsigned lon= g addr, +static unsigned long bpf_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { struct bpf_map *map =3D filp->private_data; =20 if (map->ops->map_get_unmapped_area) - return map->ops->map_get_unmapped_area(filp, addr, len, pgoff, flags); + return map->ops->map_get_unmapped_area(mm, filp, addr, len, + pgoff, flags); #ifdef CONFIG_MMU - return mm_get_unmapped_area(filp, addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, filp, addr, len, pgoff, flags); #else return addr; #endif diff --git a/mm/huge_memory.c b/mm/huge_memory.c index b5d1e9d4463d..d75c5bd5f94d 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -1193,8 +1193,8 @@ static inline bool is_transparent_hugepage(const stru= ct folio *folio) folio_test_large_rmappable(folio); } =20 -static unsigned long __thp_get_unmapped_area(struct file *filp, - unsigned long addr, unsigned long len, +static unsigned long __thp_get_unmapped_area(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, loff_t off, unsigned long flags, unsigned long size, vm_flags_t vm_flags) { @@ -1212,7 +1212,7 @@ static unsigned long __thp_get_unmapped_area(struct f= ile *filp, if (len_pad < len || (off + len_pad) < off) return 0; =20 - ret =3D mm_get_unmapped_area_vmflags(filp, addr, len_pad, + ret =3D mm_get_unmapped_area_vmflags(mm, filp, addr, len_pad, off >> PAGE_SHIFT, flags, vm_flags); =20 /* @@ -1231,32 +1231,35 @@ static unsigned long __thp_get_unmapped_area(struct= file *filp, =20 off_sub =3D (off - ret) & (size - 1); =20 - if (mm_flags_test(MMF_TOPDOWN, current->mm) && !off_sub) + if (mm_flags_test(MMF_TOPDOWN, mm) && !off_sub) return ret + size; =20 ret +=3D off_sub; return ret; } =20 -unsigned long thp_get_unmapped_area_vmflags(struct file *filp, unsigned lo= ng addr, - unsigned long len, unsigned long pgoff, unsigned long flags, - vm_flags_t vm_flags) +unsigned long thp_get_unmapped_area_vmflags(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { unsigned long ret; loff_t off =3D (loff_t)pgoff << PAGE_SHIFT; =20 - ret =3D __thp_get_unmapped_area(filp, addr, len, off, flags, PMD_SIZE, vm= _flags); + ret =3D __thp_get_unmapped_area(mm, filp, addr, len, off, flags, + PMD_SIZE, vm_flags); if (ret) return ret; =20 - return mm_get_unmapped_area_vmflags(filp, addr, len, pgoff, flags, + return mm_get_unmapped_area_vmflags(mm, filp, addr, len, pgoff, flags, vm_flags); } =20 -unsigned long thp_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, unsigned long flags) +unsigned long thp_get_unmapped_area(struct mm_struct *mm, struct file *fil= p, + unsigned long addr, unsigned long len, unsigned long pgoff, + unsigned long flags) { - return thp_get_unmapped_area_vmflags(filp, addr, len, pgoff, flags, 0); + return thp_get_unmapped_area_vmflags(mm, filp, addr, len, pgoff, + flags, 0); } EXPORT_SYMBOL_GPL(thp_get_unmapped_area); =20 diff --git a/mm/mmap.c b/mm/mmap.c index 2311ae7c2ff4..54915ac478ab 100644 --- a/mm/mmap.c +++ b/mm/mmap.c @@ -405,7 +405,7 @@ unsigned long do_mmap(struct file *file, unsigned long = addr, /* Obtain the address to map to. we verify (or select) it and ensure * that it represents a valid section of the address space. */ - addr =3D __get_unmapped_area(file, addr, len, pgoff, flags, vm_flags); + addr =3D __get_unmapped_area(mm, file, addr, len, pgoff, flags, vm_flags); if (IS_ERR_VALUE(addr)) return addr; =20 @@ -662,14 +662,15 @@ static inline unsigned long stack_guard_placement(vm_= flags_t vm_flags) * - is at least the desired size. * - satisfies (begin_addr & align_mask) =3D=3D (align_offset & align_mask) */ -unsigned long vm_unmapped_area(struct vm_unmapped_area_info *info) +unsigned long vm_unmapped_area(struct mm_struct *mm, + struct vm_unmapped_area_info *info) { unsigned long addr; =20 if (info->flags & VM_UNMAPPED_AREA_TOPDOWN) - addr =3D unmapped_area_topdown(info); + addr =3D unmapped_area_topdown(mm, info); else - addr =3D unmapped_area(info); + addr =3D unmapped_area(mm, info); =20 trace_vm_unmapped_area(addr, info); return addr; @@ -687,11 +688,11 @@ unsigned long vm_unmapped_area(struct vm_unmapped_are= a_info *info) * This function "knows" that -ENOMEM has the bits set. */ unsigned long -generic_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +generic_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma, *prev; struct vm_unmapped_area_info info =3D {}; const unsigned long mmap_end =3D arch_get_mmap_end(addr, len, flags); @@ -717,16 +718,17 @@ generic_get_unmapped_area(struct file *filp, unsigned= long addr, info.start_gap =3D stack_guard_placement(vm_flags); if (filp && is_file_hugepages(filp)) info.align_mask =3D huge_page_mask_align(filp); - return vm_unmapped_area(&info); + return vm_unmapped_area(mm, &info); } =20 #ifndef HAVE_ARCH_UNMAPPED_AREA unsigned long -arch_get_unmapped_area(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +arch_get_unmapped_area(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { - return generic_get_unmapped_area(filp, addr, len, pgoff, flags, + return generic_get_unmapped_area(mm, filp, addr, len, pgoff, flags, vm_flags); } #endif @@ -736,12 +738,12 @@ arch_get_unmapped_area(struct file *filp, unsigned lo= ng addr, * stack's low limit (the base): */ unsigned long -generic_get_unmapped_area_topdown(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +generic_get_unmapped_area_topdown(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { struct vm_area_struct *vma, *prev; - struct mm_struct *mm =3D current->mm; struct vm_unmapped_area_info info =3D {}; const unsigned long mmap_end =3D arch_get_mmap_end(addr, len, flags); =20 @@ -769,7 +771,7 @@ generic_get_unmapped_area_topdown(struct file *filp, un= signed long addr, info.start_gap =3D stack_guard_placement(vm_flags); if (filp && is_file_hugepages(filp)) info.align_mask =3D huge_page_mask_align(filp); - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); =20 /* * A failed mmap() very likely causes application failure, @@ -782,7 +784,7 @@ generic_get_unmapped_area_topdown(struct file *filp, un= signed long addr, info.flags =3D 0; info.low_limit =3D TASK_UNMAPPED_BASE; info.high_limit =3D mmap_end; - addr =3D vm_unmapped_area(&info); + addr =3D vm_unmapped_area(mm, &info); } =20 return addr; @@ -790,31 +792,36 @@ generic_get_unmapped_area_topdown(struct file *filp, = unsigned long addr, =20 #ifndef HAVE_ARCH_UNMAPPED_AREA_TOPDOWN unsigned long -arch_get_unmapped_area_topdown(struct file *filp, unsigned long addr, - unsigned long len, unsigned long pgoff, - unsigned long flags, vm_flags_t vm_flags) +arch_get_unmapped_area_topdown(struct mm_struct *mm, struct file *filp, + unsigned long addr, unsigned long len, + unsigned long pgoff, unsigned long flags, + vm_flags_t vm_flags) { - return generic_get_unmapped_area_topdown(filp, addr, len, pgoff, flags, - vm_flags); + return generic_get_unmapped_area_topdown(mm, filp, addr, len, pgoff, + flags, vm_flags); } #endif =20 -unsigned long mm_get_unmapped_area_vmflags(struct file *filp, unsigned lon= g addr, +unsigned long mm_get_unmapped_area_vmflags(struct mm_struct *mm, + struct file *filp, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { - if (mm_flags_test(MMF_TOPDOWN, current->mm)) - return arch_get_unmapped_area_topdown(filp, addr, len, pgoff, - flags, vm_flags); - return arch_get_unmapped_area(filp, addr, len, pgoff, flags, vm_flags); + if (mm_flags_test(MMF_TOPDOWN, mm)) + return arch_get_unmapped_area_topdown(mm, filp, addr, len, + pgoff, flags, vm_flags); + return arch_get_unmapped_area(mm, filp, addr, len, pgoff, flags, + vm_flags); } =20 unsigned long -__get_unmapped_area(struct file *file, unsigned long addr, unsigned long l= en, +__get_unmapped_area(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags, vm_flags_t vm_flags) { - unsigned long (*get_area)(struct file *, unsigned long, - unsigned long, unsigned long, unsigned long) + unsigned long (*get_area)(struct mm_struct *, struct file *, + unsigned long, unsigned long, + unsigned long, unsigned long) =3D NULL; =20 unsigned long error =3D arch_mmap_check(addr, len, flags); @@ -841,15 +848,15 @@ __get_unmapped_area(struct file *file, unsigned long = addr, unsigned long len, pgoff =3D 0; =20 if (get_area) { - addr =3D get_area(file, addr, len, pgoff, flags); + addr =3D get_area(mm, file, addr, len, pgoff, flags); } else if (IS_ENABLED(CONFIG_TRANSPARENT_HUGEPAGE) && !file && !addr /* no hint */ && IS_ALIGNED(len, PMD_SIZE)) { /* Ensures that larger anonymous mappings are THP aligned. */ - addr =3D thp_get_unmapped_area_vmflags(file, addr, len, + addr =3D thp_get_unmapped_area_vmflags(mm, file, addr, len, pgoff, flags, vm_flags); } else { - addr =3D mm_get_unmapped_area_vmflags(file, addr, len, + addr =3D mm_get_unmapped_area_vmflags(mm, file, addr, len, pgoff, flags, vm_flags); } if (IS_ERR_VALUE(addr)) @@ -865,10 +872,12 @@ __get_unmapped_area(struct file *file, unsigned long = addr, unsigned long len, } =20 unsigned long -mm_get_unmapped_area(struct file *file, unsigned long addr, unsigned long = len, +mm_get_unmapped_area(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { - return mm_get_unmapped_area_vmflags(file, addr, len, pgoff, flags, 0); + return mm_get_unmapped_area_vmflags(mm, file, addr, len, pgoff, + flags, 0); } EXPORT_SYMBOL(mm_get_unmapped_area); =20 diff --git a/mm/nommu.c b/mm/nommu.c index ed3934bc2de4..d6393bd2844e 100644 --- a/mm/nommu.c +++ b/mm/nommu.c @@ -1145,8 +1145,9 @@ unsigned long do_mmap(struct file *file, * tell us the location of a shared mapping */ if (capabilities & NOMMU_MAP_DIRECT) { - addr =3D file->f_op->get_unmapped_area(file, addr, len, - pgoff, flags); + addr =3D file->f_op->get_unmapped_area(current->mm, file, + addr, len, pgoff, + flags); if (IS_ERR_VALUE(addr)) { ret =3D addr; if (ret !=3D -ENOSYS) diff --git a/mm/shmem.c b/mm/shmem.c index b51f83c970bb..1349e00bc7fd 100644 --- a/mm/shmem.c +++ b/mm/shmem.c @@ -2711,7 +2711,8 @@ static vm_fault_t shmem_fault(struct vm_fault *vmf) return ret; } =20 -unsigned long shmem_get_unmapped_area(struct file *file, +unsigned long shmem_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long uaddr, unsigned long len, unsigned long pgoff, unsigned long flags) { @@ -2725,7 +2726,7 @@ unsigned long shmem_get_unmapped_area(struct file *fi= le, if (len > TASK_SIZE) return -ENOMEM; =20 - addr =3D mm_get_unmapped_area(file, uaddr, len, pgoff, flags); + addr =3D mm_get_unmapped_area(mm, file, uaddr, len, pgoff, flags); =20 if (!IS_ENABLED(CONFIG_TRANSPARENT_HUGEPAGE)) return addr; @@ -2803,7 +2804,8 @@ unsigned long shmem_get_unmapped_area(struct file *fi= le, if (inflated_len < len) return addr; =20 - inflated_addr =3D mm_get_unmapped_area(NULL, uaddr, inflated_len, 0, flag= s); + inflated_addr =3D mm_get_unmapped_area(mm, NULL, uaddr, inflated_len, 0, + flags); if (IS_ERR_VALUE(inflated_addr)) return addr; if (inflated_addr & ~PAGE_MASK) @@ -5730,11 +5732,12 @@ void shmem_unlock_mapping(struct address_space *map= ping) } =20 #ifdef CONFIG_MMU -unsigned long shmem_get_unmapped_area(struct file *file, +unsigned long shmem_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, unsigned long flags) { - return mm_get_unmapped_area(file, addr, len, pgoff, flags); + return mm_get_unmapped_area(mm, file, addr, len, pgoff, flags); } #endif =20 diff --git a/mm/vma.c b/mm/vma.c index 9eea2850818a..9a6cb0338ed8 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -2957,20 +2957,22 @@ int do_brk_flags(struct vma_iterator *vmi, struct v= m_area_struct *vma, =20 /** * unmapped_area() - Find an area between the low_limit and the high_limit= with - * the correct alignment and offset, all from @info. Note: current->mm is = used + * the correct alignment and offset, all from @info. Note: @mm is used * for the search. * + * @mm: The mm_struct to search. * @info: The unmapped area information including the range [low_limit - * high_limit), the alignment offset and mask. * * Return: A memory address or -ENOMEM. */ -unsigned long unmapped_area(struct vm_unmapped_area_info *info) +unsigned long unmapped_area(struct mm_struct *mm, + struct vm_unmapped_area_info *info) { unsigned long length, gap; unsigned long low_limit, high_limit; struct vm_area_struct *tmp; - VMA_ITERATOR(vmi, current->mm, 0); + VMA_ITERATOR(vmi, mm, 0); =20 /* Adjust search length to account for worst case alignment overhead */ length =3D info->length + info->align_mask + info->start_gap; @@ -3016,19 +3018,21 @@ unsigned long unmapped_area(struct vm_unmapped_area= _info *info) /** * unmapped_area_topdown() - Find an area between the low_limit and the * high_limit with the correct alignment and offset at the highest availab= le - * address, all from @info. Note: current->mm is used for the search. + * address, all from @info. Note: @mm is used for the search. * + * @mm: The mm_struct to search. * @info: The unmapped area information including the range [low_limit - * high_limit), the alignment offset and mask. * * Return: A memory address or -ENOMEM. */ -unsigned long unmapped_area_topdown(struct vm_unmapped_area_info *info) +unsigned long unmapped_area_topdown(struct mm_struct *mm, + struct vm_unmapped_area_info *info) { unsigned long length, gap, gap_end; unsigned long low_limit, high_limit; struct vm_area_struct *tmp; - VMA_ITERATOR(vmi, current->mm, 0); + VMA_ITERATOR(vmi, mm, 0); =20 /* Adjust search length to account for worst case alignment overhead */ length =3D info->length + info->align_mask + info->start_gap; diff --git a/mm/vma.h b/mm/vma.h index 8e4b61a7304c..d0a45718deb9 100644 --- a/mm/vma.h +++ b/mm/vma.h @@ -467,8 +467,10 @@ int do_brk_flags(struct vma_iterator *vmi, struct vm_a= rea_struct *brkvma, unsigned long addr, unsigned long request, vma_flags_t vma_flags); =20 -unsigned long unmapped_area(struct vm_unmapped_area_info *info); -unsigned long unmapped_area_topdown(struct vm_unmapped_area_info *info); +unsigned long unmapped_area(struct mm_struct *mm, + struct vm_unmapped_area_info *info); +unsigned long unmapped_area_topdown(struct mm_struct *mm, + struct vm_unmapped_area_info *info); =20 static inline bool vma_wants_manual_pte_write_upgrade(struct vm_area_struc= t *vma) { diff --git a/sound/core/pcm_native.c b/sound/core/pcm_native.c index 7dc0060617f1..f01465a70ddd 100644 --- a/sound/core/pcm_native.c +++ b/sound/core/pcm_native.c @@ -4202,7 +4202,8 @@ static int snd_pcm_hw_params_old_user(struct snd_pcm_= substream *substream, #endif /* CONFIG_SND_SUPPORT_OLD_API */ =20 #ifndef CONFIG_MMU -static unsigned long snd_pcm_get_unmapped_area(struct file *file, +static unsigned long snd_pcm_get_unmapped_area(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long pgoff, --=20 2.43.0 From nobody Sat Jul 25 05:59:38 2026 Received: from mail-pf1-f179.google.com (mail-pf1-f179.google.com [209.85.210.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B61BC40BCAB for ; Fri, 24 Jul 2026 22:02:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930542; cv=none; b=YCpOMb72eQefk01MmA+iBWa+/yIyQq+ry/A3OlMfuyOgZ34WihuofA9/BU9cItmqNxthrtohMIUHtEF4nRNvlFcVvV/HdV6nz23e2UnzBrFYqgGsyt03h4ZbKgQFCe+LFhb+H97qe5o/c0isReiSGm+tAsLrXvv+MaLVLM1iA+A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930542; c=relaxed/simple; bh=RxZJ3yGFkquK1zXEEtJeWjfnPiAvm4PtLP2tNh0/X34=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=f8tx9Paf7lZW6dusDFTGj9rpUpF1M8WePJYtZsxn1VGWkHcO0Mi/RkMQQVvzFrQe/heSHs8mxAWClQ0PoDN3guRg+gB8/uUCj4EbE2Xs5g7J15MdYQrdUsyQ1J2l/tNfngL5CDNj8t3AqnhMYT43d4rzaaFv8+dHOQIS9syDQto= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=ojsAqlG/; arc=none smtp.client-ip=209.85.210.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="ojsAqlG/" Received: by mail-pf1-f179.google.com with SMTP id d2e1a72fcca58-84862b0d5aeso992637b3a.2 for ; Fri, 24 Jul 2026 15:02:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784930540; x=1785535340; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=SJ13WqemEDPKEjZObDcoideCrDvXpg5YU7m60Etb0C8=; b=ojsAqlG/QX2llA7HhtLaP9HbqcRormih/Q2+NMc/zhKDmNYSraNCdGnfvx79i2m9Me 3H+U1rNLGbRpXt7as2V9rLgMJR1uFGJAoE7je1ZmHWzh6a4YjUMUvkik7fMaeRxakQL/ EFwH0QtPoprfY4C0TXpO0VGHWiRPr64UQsWXkwkhGQqwFV6Ohrd8iYyhY/dpr2MQm0S4 m37nAA/B1CFdMAvakkJjAkopwdmOkh9lXBc740OHUBP1gFM+9iV3LWPyZtN7sA0109na 30g5sqW9JesRvFFLfZPyoOB97rcQ5aprO3FRUhlFN8XeiX6S54kK3OwRyDqhNHL9PBAw 1yAg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784930540; x=1785535340; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=SJ13WqemEDPKEjZObDcoideCrDvXpg5YU7m60Etb0C8=; b=BcPc51f6ZkrHiuEdroC42ltULYTscouf4hZ3w2QGejLhslOjQzLC3QhGWDkJABZ+O9 KeSuKvBch5TvlQt1XbfHpztUUFC2VjEpiIU1B837EkCkbWU3b1pf8PjJGRtda4HFeysn xt9T65G7aHn6i9ewMPljCBko7lAx3Kj81tLnMchtqP4YbZA0kSRx09swWPOOmdIv+7HW XIM0ClTbVkmq9dlcJEN9P6ygyRqjwd58B+UG41QTTWFwHwIX7+HJDUGtMpSavrhDQDqV Lgxb3wtAxNkmvq+q68Cvh+Cyx+E8Irno12lx6EbERv8nQMrNiS1A2nfdjzm6meURDHH9 v/xw== X-Forwarded-Encrypted: i=1; AHgh+RqpxRp3lwlZaqFGPtS1ka0ysPr7T/VvOMBG0H+HZzQKXHO7MUnqz9R3xGVhCsnpQKGZaHHfb46HjV7fMjk=@vger.kernel.org X-Gm-Message-State: AOJu0YyG7wxUw6SjWl2p3xd4VFPScpxAY4ckxDM4tNpqA/EVfYTmWXf8 PH4ZzWIaJZyMldTvWXMOtgMu0/wgbb5o3EwSYNGS3lF22EdjRHKQ5FhJ X-Gm-Gg: AR+sD12UEH6uosC7NaeEduS4cBscJrVhX7BI1It52eyEryiVC1SreaDnUijMBIh07ZQ 8vLPPykpUEdOPN3hzzE4TBy6FzHBVyfBmaD12GYKFeahv1+sIfBT0BJkUFv+TnFAJB0MGk6lGUt 7p+U2U+U8gDFjbukXgi0Sc8I9d1yvzuvGL4DYjN4HcJDCd/dSVtaD49Oya60Ls6oC4TXQMnBSZD Pqw/7/NH5AZyLqptlJ3b/9Q24nJ/y+o298p0ikzFhqoq/1OLbgGQT62xzoYe8IPNJorCAjqclC+ r7x7aFINwtFzMneQLXQ65zrkmjjQKZbS1zZl0B1AdZzhA20Nf6ludqu7L75CeT2Z6QPEejNumQD BQOAT80w6y3ACMzVuwnt8xBkOP4Ky+ZbHW+qaWbp4IwATft4aONs5HTeXfF9LfbA2QbPOfTJ2T9 cKP7aslS4PT5c162FWizrW6R1Ax1QGq5GCq4WfLlDO5gRN/1pWt3uTpR7OeWYS X-Received: by 2002:a05:6a20:a128:b0:3c3:b57b:6459 with SMTP id adf61e73a8af0-3c67d9b4062mr167283637.7.1784930539791; Fri, 24 Jul 2026 15:02:19 -0700 (PDT) Received: from pop-os.tail4adac7.ts.net ([2601:647:6802:dbc0:546a:e1b0:9574:3a29]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc5a67f3sm2962631eec.29.2026.07.24.15.02.17 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 15:02:18 -0700 (PDT) From: Cong Wang To: Andy Lutomirski Cc: Kees Cook , linux-kernel@vger.kernel.org, Will Drewry , Christian Brauner , Andrew Morton , linux-mm@kvack.org, Cong Wang Subject: [PATCH v7 2/8] mm: add __do_mmap() and vm_mmap_remote()/vm_munmap_remote() Date: Fri, 24 Jul 2026 15:01:41 -0700 Message-ID: <20260724220147.214396-3-xiyou.wangcong@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260724220147.214396-1-xiyou.wangcong@gmail.com> References: <20260724220147.214396-1-xiyou.wangcong@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Cong Wang Add __do_mmap(), a variant of do_mmap() that installs the mapping into a caller-supplied mm rather than current->mm; do_mmap() becomes a wrapper passing current->mm and keeps the current-task READ_IMPLIES_EXEC handling. mmap_region()/__mmap_region() gain an mm argument so the target mm reaches VMA insertion. On top of that, add vm_mmap_remote() and vm_munmap_remote() in mm/util.c: entry points that install / remove a mapping in a caller-specified mm without target-side cooperation. The intended user is the seccomp unotify syscall redirect. Kernel-chosen placement resolves through __get_unmapped_area() against the target mm, so (via the previous patch) the file's own ->get_unmapped_area() runs and file-specific placement is honored. The result is bounds-checked against the target's mm->task_size, and the munmap machinery is likewise threaded with the target mm so counter bookkeeping and range checks apply to it rather than current->mm. vm_munmap_remote() delivers the target's UFFD_EVENT_UNMAP. Both are MMU-only and guarded with CONFIG_MMU, with -EOPNOTSUPP stubs so NOMMU keeps linking. LSM/fsnotify hooks run against current, the supervisor, paralleling pidfd_getfd(); cross-task authorization is left to the caller. Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Cong Wang --- include/linux/mm.h | 19 ++++++ mm/internal.h | 5 ++ mm/mmap.c | 51 ++++++++------ mm/mprotect.c | 2 +- mm/nommu.c | 22 ++++-- mm/util.c | 117 ++++++++++++++++++++++++++++++++ mm/vma.c | 46 +++++++------ mm/vma.h | 12 ++-- tools/testing/vma/include/dup.h | 1 + tools/testing/vma/tests/mmap.c | 8 +-- 10 files changed, 228 insertions(+), 55 deletions(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index 560bc369a54c..fa1d96bc379d 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -4155,6 +4155,25 @@ extern unsigned long do_mmap(struct file *file, unsi= gned long addr, unsigned long len, unsigned long prot, unsigned long flags, vm_flags_t vm_flags, unsigned long pgoff, unsigned long *populate, struct list_head *uf); +#ifdef CONFIG_MMU +unsigned long vm_mmap_remote(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, unsigned long prot, + unsigned long flags, unsigned long pgoff, vm_flags_t vm_flags); +int vm_munmap_remote(struct mm_struct *mm, unsigned long start, size_t len= ); +#else +static inline unsigned long vm_mmap_remote(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, + unsigned long prot, unsigned long flags, unsigned long pgoff, + vm_flags_t vm_flags) +{ + return -EOPNOTSUPP; +} +static inline int vm_munmap_remote(struct mm_struct *mm, unsigned long sta= rt, + size_t len) +{ + return -EOPNOTSUPP; +} +#endif extern int do_vmi_munmap(struct vma_iterator *vmi, struct mm_struct *mm, unsigned long start, size_t len, struct list_head *uf, bool unlock); diff --git a/mm/internal.h b/mm/internal.h index 181e79f1d6a2..6a510a2fd59e 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -1436,6 +1436,11 @@ extern unsigned long __must_check vm_mmap_pgoff(str= uct file *, unsigned long, unsigned long, unsigned long, unsigned long, unsigned long); =20 +unsigned long __do_mmap(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, unsigned long prot, + unsigned long flags, vm_flags_t vm_flags, unsigned long pgoff, + unsigned long *populate, struct list_head *uf); + extern void set_pageblock_order(void); unsigned long reclaim_pages(struct list_head *folio_list); unsigned int reclaim_clean_pages_from_list(struct zone *zone, diff --git a/mm/mmap.c b/mm/mmap.c index 54915ac478ab..23cf45ed22e1 100644 --- a/mm/mmap.c +++ b/mm/mmap.c @@ -277,8 +277,8 @@ static inline bool file_mmap_ok(struct file *file, stru= ct inode *inode, } =20 /** - * do_mmap() - Perform a userland memory mapping into the current process - * address space of length @len with protection bits @prot, mmap flags @fl= ags + * __do_mmap() - Perform a userland memory mapping into @mm's address space + * of length @len with protection bits @prot, mmap flags @flags * (from which VMA flags will be inferred), and any additional VMA flags to * apply @vm_flags. If this is a file-backed mapping then the file is spec= ified * in @file and page offset into the file via @pgoff. @@ -307,8 +307,11 @@ static inline bool file_mmap_ok(struct file *file, str= uct inode *inode, * start of a VMA, rather only the start of a valid mapped range of length * @len bytes, rounded down to the nearest page size. * - * The caller must write-lock current->mm->mmap_lock. + * The caller must write-lock @mm->mmap_lock. do_mmap() is the common + * wrapper that targets current->mm. * + * @mm: The mm_struct to install the mapping into. The caller must hold a + * reference and write-lock its mmap_lock. * @file: An optional struct file pointer describing the file which is to = be * mapped, if a file-backed mapping. * @addr: If non-zero, hints at (or if @flags has MAP_FIXED set, specifies= ) the @@ -333,13 +336,12 @@ static inline bool file_mmap_ok(struct file *file, st= ruct inode *inode, * Returns: Either an error, or the address at which the requested mapping= has * been performed. */ -unsigned long do_mmap(struct file *file, unsigned long addr, - unsigned long len, unsigned long prot, - unsigned long flags, vm_flags_t vm_flags, - unsigned long pgoff, unsigned long *populate, - struct list_head *uf) +unsigned long __do_mmap(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, + unsigned long prot, unsigned long flags, + vm_flags_t vm_flags, unsigned long pgoff, + unsigned long *populate, struct list_head *uf) { - struct mm_struct *mm =3D current->mm; int pkey =3D 0; =20 *populate =3D 0; @@ -349,16 +351,6 @@ unsigned long do_mmap(struct file *file, unsigned long= addr, if (!len) return -EINVAL; =20 - /* - * Does the application expect PROT_READ to imply PROT_EXEC? - * - * (the exception is when the underlying filesystem is noexec - * mounted, in which case we don't add PROT_EXEC.) - */ - if ((prot & PROT_READ) && (current->personality & READ_IMPLIES_EXEC)) - if (!(file && path_noexec(&file->f_path))) - prot |=3D PROT_EXEC; - /* force arch specific MAP_FIXED handling in get_unmapped_area */ if (flags & MAP_FIXED_NOREPLACE) flags |=3D MAP_FIXED; @@ -557,7 +549,7 @@ unsigned long do_mmap(struct file *file, unsigned long = addr, vm_flags |=3D VM_NORESERVE; } =20 - addr =3D mmap_region(file, addr, len, vm_flags, pgoff, uf); + addr =3D mmap_region(mm, file, addr, len, vm_flags, pgoff, uf); if (!IS_ERR_VALUE(addr) && ((vm_flags & VM_LOCKED) || (flags & (MAP_POPULATE | MAP_NONBLOCK)) =3D=3D MAP_POPULATE)) @@ -565,6 +557,25 @@ unsigned long do_mmap(struct file *file, unsigned long= addr, return addr; } =20 +unsigned long do_mmap(struct file *file, unsigned long addr, unsigned long= len, + unsigned long prot, unsigned long flags, + vm_flags_t vm_flags, unsigned long pgoff, + unsigned long *populate, struct list_head *uf) +{ + /* + * Does the application expect PROT_READ to imply PROT_EXEC? + * + * (the exception is when the underlying filesystem is noexec + * mounted, in which case we don't add PROT_EXEC.) + */ + if ((prot & PROT_READ) && (current->personality & READ_IMPLIES_EXEC)) + if (!(file && path_noexec(&file->f_path))) + prot |=3D PROT_EXEC; + + return __do_mmap(current->mm, file, addr, len, prot, flags, vm_flags, + pgoff, populate, uf); +} + unsigned long ksys_mmap_pgoff(unsigned long addr, unsigned long len, unsigned long prot, unsigned long flags, unsigned long fd, unsigned long pgoff) diff --git a/mm/mprotect.c b/mm/mprotect.c index 9cbf932b028c..02e90577db25 100644 --- a/mm/mprotect.c +++ b/mm/mprotect.c @@ -939,7 +939,7 @@ static int do_mprotect_pkey(unsigned long start, size_t= len, break; } =20 - if (map_deny_write_exec(&vma->flags, &new_vma_flags)) { + if (map_deny_write_exec(vma->vm_mm, &vma->flags, &new_vma_flags)) { error =3D -EACCES; break; } diff --git a/mm/nommu.c b/mm/nommu.c index d6393bd2844e..15390ef5b784 100644 --- a/mm/nommu.c +++ b/mm/nommu.c @@ -1009,7 +1009,8 @@ static int do_mmap_private(struct vm_area_struct *vma, /* * handle mapping creation for uClinux */ -unsigned long do_mmap(struct file *file, +unsigned long __do_mmap(struct mm_struct *mm, + struct file *file, unsigned long addr, unsigned long len, unsigned long prot, @@ -1024,7 +1025,7 @@ unsigned long do_mmap(struct file *file, struct rb_node *rb; unsigned long capabilities, result; int ret; - VMA_ITERATOR(vmi, current->mm, 0); + VMA_ITERATOR(vmi, mm, 0); =20 *populate =3D 0; =20 @@ -1049,7 +1050,7 @@ unsigned long do_mmap(struct file *file, if (!region) goto error_getting_region; =20 - vma =3D vm_area_alloc(current->mm); + vma =3D vm_area_alloc(mm); if (!vma) goto error_getting_vma; =20 @@ -1191,7 +1192,7 @@ unsigned long do_mmap(struct file *file, /* okay... we have a mapping; now we have to register it */ result =3D vma->vm_start; =20 - current->mm->total_vm +=3D len >> PAGE_SHIFT; + mm->total_vm +=3D len >> PAGE_SHIFT; =20 share: BUG_ON(!vma->vm_region); @@ -1199,8 +1200,8 @@ unsigned long do_mmap(struct file *file, if (vma_iter_prealloc(&vmi, vma)) goto error_just_free; =20 - setup_vma_to_mm(vma, current->mm); - current->mm->map_count++; + setup_vma_to_mm(vma, mm); + mm->map_count++; /* add the VMA to the tree */ vma_iter_store_new(&vmi, vma); =20 @@ -1247,6 +1248,15 @@ unsigned long do_mmap(struct file *file, return -ENOMEM; } =20 +unsigned long do_mmap(struct file *file, unsigned long addr, unsigned long= len, + unsigned long prot, unsigned long flags, + vm_flags_t vm_flags, unsigned long pgoff, + unsigned long *populate, struct list_head *uf) +{ + return __do_mmap(current->mm, file, addr, len, prot, flags, vm_flags, pgo= ff, + populate, uf); +} + unsigned long ksys_mmap_pgoff(unsigned long addr, unsigned long len, unsigned long prot, unsigned long flags, unsigned long fd, unsigned long pgoff) diff --git a/mm/util.c b/mm/util.c index af2c2103f0d9..a7d499012795 100644 --- a/mm/util.c +++ b/mm/util.c @@ -588,6 +588,123 @@ unsigned long vm_mmap_pgoff(struct file *file, unsign= ed long addr, return ret; } =20 +#ifdef CONFIG_MMU + +#define VM_MMAP_REMOTE_FLAGS (MAP_FIXED | MAP_FIXED_NOREPLACE) + +/** + * vm_mmap_remote - install a file (or anonymous) mapping into @mm, without + * target-side cooperation. + * @mm: Target mm; caller holds a reference (e.g. get_task_mm()). + * @file: Backing file, or NULL for an anonymous mapping. + * @addr: Page-aligned placement. Zero asks the kernel to choose a free ar= ea + * in @mm and return it. Non-zero requires MAP_FIXED or + * MAP_FIXED_NOREPLACE in @flags. + * @len: Length in bytes. + * @prot: PROT_* protection bits. + * @flags: MAP_* flags. + * @pgoff: Page offset into @file. + * @vm_flags: Extra VMA flags to OR in (e.g. VM_SEALED), or 0. + * + * The general form of the remote install primitive: a supervisor places a + * mapping into a remote task's address space. LSM/fsnotify hooks run agai= nst + * %current (the installer), paralleling pidfd_getfd()'s cross-task instal= l; + * cross-task authorization is the caller's responsibility. Only MAP_SHARE= D is + * supported, and @flags outside VM_MMAP_REMOTE_FLAGS are refused with + * -EOPNOTSUPP. + * + * Kernel-chosen placement (@addr =3D=3D 0) resolves through + * __get_unmapped_area() against @mm, so the file's own + * ->get_unmapped_area() runs and file-specific placement (hugetlb page + * size, THP PMD alignment, SHMLBA cache coloring, device dax alignment) + * is honored exactly as for a local mmap(). + * + * Returns the mapped address on success, or a negative errno. + */ +unsigned long vm_mmap_remote(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, unsigned long prot, + unsigned long flags, unsigned long pgoff, vm_flags_t vm_flags) +{ + loff_t off =3D (loff_t)pgoff << PAGE_SHIFT; + unsigned long aligned_len =3D PAGE_ALIGN(len); + unsigned long ret; + unsigned long populate; + LIST_HEAD(uf); + + if (WARN_ON_ONCE(!mm)) + return -EINVAL; + + if (addr && !(flags & (MAP_FIXED | MAP_FIXED_NOREPLACE))) + return -EINVAL; + + if ((flags & MAP_TYPE) !=3D MAP_SHARED) + return -EOPNOTSUPP; + + if (flags & ~(MAP_TYPE | VM_MMAP_REMOTE_FLAGS)) + return -EOPNOTSUPP; + + ret =3D security_mmap_file(file, prot, flags); + if (!ret) + ret =3D fsnotify_mmap_perm(file, prot, off, len); + if (ret) + return ret; + + if (mmap_write_lock_killable(mm)) + return -EINTR; + + /* + * Resolve kernel-chosen placement up front and pin it with + * MAP_FIXED, so the mm->task_size check below runs before the + * mapping is installed. The file's own ->get_unmapped_area() + * performs the search against @mm. + */ + if (!addr) { + addr =3D __get_unmapped_area(mm, file, 0, aligned_len, pgoff, + flags, vm_flags); + if (IS_ERR_VALUE(addr)) { + ret =3D addr; + goto unlock; + } + flags |=3D MAP_FIXED; + } + + if (aligned_len > mm->task_size || addr > mm->task_size - aligned_len) { + ret =3D -ENOMEM; + goto unlock; + } + + ret =3D __do_mmap(mm, file, addr, len, prot, flags, vm_flags, + pgoff, &populate, &uf); +unlock: + mmap_write_unlock(mm); + userfaultfd_unmap_complete(mm, &uf); + return ret; +} + +/** + * vm_munmap_remote - unmap a range from @mm, mirroring vm_munmap() but + * targeting an explicit mm. + * @mm: Target mm; caller holds a reference (e.g. get_task_mm()). + * @start: Start address of the range to unmap. + * @len: Length in bytes. + * + * Returns 0 on success, or a negative errno. + */ +int vm_munmap_remote(struct mm_struct *mm, unsigned long start, size_t len) +{ + VMA_ITERATOR(vmi, mm, start); + LIST_HEAD(uf); + int ret; + + if (mmap_write_lock_killable(mm)) + return -EINTR; + ret =3D do_vmi_munmap(&vmi, mm, start, len, &uf, false); + mmap_write_unlock(mm); + userfaultfd_unmap_complete(mm, &uf); + return ret; +} +#endif /* CONFIG_MMU */ + /* * Perform a userland memory mapping into the current process address spac= e. See * the comment for do_mmap() for more details on this operation in general. diff --git a/mm/vma.c b/mm/vma.c index 9a6cb0338ed8..788c2a0f35ec 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -1333,7 +1333,7 @@ static void vms_complete_munmap_vmas(struct vma_munma= p_struct *vms, struct vm_area_struct *vma; struct mm_struct *mm; =20 - mm =3D current->mm; + mm =3D vms->mm; mm->map_count -=3D vms->vma_count; mm->locked_vm -=3D vms->locked_vm; if (vms->unlock) @@ -1544,10 +1544,11 @@ static int vms_gather_munmap_vmas(struct vma_munmap= _struct *vms, * @unlock: Unlock after the operation. Only unlocked on success */ static void init_vma_munmap(struct vma_munmap_struct *vms, - struct vma_iterator *vmi, struct vm_area_struct *vma, - unsigned long start, unsigned long end, struct list_head *uf, - bool unlock) + struct mm_struct *mm, struct vma_iterator *vmi, + struct vm_area_struct *vma, unsigned long start, + unsigned long end, struct list_head *uf, bool unlock) { + vms->mm =3D mm; vms->vmi =3D vmi; vms->vma =3D vma; if (vma) { @@ -1591,7 +1592,7 @@ int do_vmi_align_munmap(struct vma_iterator *vmi, str= uct vm_area_struct *vma, struct vma_munmap_struct vms; int error; =20 - init_vma_munmap(&vms, vmi, vma, start, end, uf, unlock); + init_vma_munmap(&vms, mm, vmi, vma, start, end, uf, unlock); error =3D vms_gather_munmap_vmas(&vms, &mas_detach); if (error) goto gather_failed; @@ -1634,7 +1635,8 @@ int do_vmi_munmap(struct vma_iterator *vmi, struct mm= _struct *mm, unsigned long end; struct vm_area_struct *vma; =20 - if ((offset_in_page(start)) || start > TASK_SIZE || len > TASK_SIZE-start) + if ((offset_in_page(start)) || start > mm->task_size || + len > mm->task_size - start) return -EINVAL; =20 end =3D start + PAGE_ALIGN(len); @@ -2426,7 +2428,7 @@ static int __mmap_setup(struct mmap_state *map, struc= t vm_area_desc *desc, =20 /* Find the first overlapping VMA and initialise unmap state. */ vms->vma =3D vma_find(vmi, map->end); - init_vma_munmap(vms, vmi, vms->vma, map->addr, map->end, uf, + init_vma_munmap(vms, map->mm, vmi, vms->vma, map->addr, map->end, uf, /* unlock =3D */ false); =20 /* OK, we have overlapping VMAs - prepare to unmap them. */ @@ -2731,11 +2733,10 @@ static bool can_set_ksm_flags_early(struct mmap_sta= te *map) return false; } =20 -static unsigned long __mmap_region(struct file *file, unsigned long addr, - unsigned long len, vma_flags_t vma_flags, +static unsigned long __mmap_region(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, vma_flags_t vma_flags, unsigned long pgoff, struct list_head *uf) { - struct mm_struct *mm =3D current->mm; struct vm_area_struct *vma =3D NULL; bool have_mmap_prepare =3D file && file->f_op->mmap_prepare; VMA_ITERATOR(vmi, mm, addr); @@ -2809,14 +2810,16 @@ static unsigned long __mmap_region(struct file *fil= e, unsigned long addr, =20 /** * mmap_region() - Actually perform the userland mapping of a VMA into - * current->mm with known, aligned and overflow-checked @addr and @len, and + * @mm with known, aligned and overflow-checked @addr and @len, and * correctly determined VMA flags @vm_flags and page offset @pgoff. * * This is an internal memory management function, and should not be used * directly. * - * The caller must write-lock current->mm->mmap_lock. + * The caller must write-lock @mm->mmap_lock. * + * @mm: The mm_struct to install the mapping into. The caller must hold a + * reference and write-lock its mmap_lock. * @file: If a file-backed mapping, a pointer to the struct file describin= g the * file to be mapped, otherwise NULL. * @addr: The page-aligned address at which to perform the mapping. @@ -2830,18 +2833,19 @@ static unsigned long __mmap_region(struct file *fil= e, unsigned long addr, * Returns: Either an error, or the address at which the requested mapping= has * been performed. */ -unsigned long mmap_region(struct file *file, unsigned long addr, - unsigned long len, vm_flags_t vm_flags, - unsigned long pgoff, struct list_head *uf) +unsigned long mmap_region(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, + vm_flags_t vm_flags, unsigned long pgoff, + struct list_head *uf) { unsigned long ret; bool writable_file_mapping =3D false; const vma_flags_t vma_flags =3D legacy_to_vma_flags(vm_flags); =20 - mmap_assert_write_locked(current->mm); + mmap_assert_write_locked(mm); =20 /* Check to see if MDWE is applicable. */ - if (map_deny_write_exec(&vma_flags, &vma_flags)) + if (map_deny_write_exec(mm, &vma_flags, &vma_flags)) return -EACCES; =20 /* Allow architectures to sanity-check the vm_flags. */ @@ -2857,13 +2861,13 @@ unsigned long mmap_region(struct file *file, unsign= ed long addr, writable_file_mapping =3D true; } =20 - ret =3D __mmap_region(file, addr, len, vma_flags, pgoff, uf); + ret =3D __mmap_region(mm, file, addr, len, vma_flags, pgoff, uf); =20 /* Clear our write mapping regardless of error. */ if (writable_file_mapping) mapping_unmap_writable(file->f_mapping); =20 - validate_mm(current->mm); + validate_mm(mm); return ret; } =20 @@ -2983,6 +2987,8 @@ unsigned long unmapped_area(struct mm_struct *mm, if (low_limit < mmap_min_addr) low_limit =3D mmap_min_addr; high_limit =3D info->high_limit; + if (mm !=3D current->mm) + high_limit =3D min(high_limit, mm->task_size); retry: if (vma_iter_area_lowest(&vmi, low_limit, high_limit, length)) return -ENOMEM; @@ -3043,6 +3049,8 @@ unsigned long unmapped_area_topdown(struct mm_struct = *mm, if (low_limit < mmap_min_addr) low_limit =3D mmap_min_addr; high_limit =3D info->high_limit; + if (mm !=3D current->mm) + high_limit =3D min(high_limit, mm->task_size); retry: if (vma_iter_area_highest(&vmi, low_limit, high_limit, length)) return -ENOMEM; diff --git a/mm/vma.h b/mm/vma.h index d0a45718deb9..0c0b9d9c5248 100644 --- a/mm/vma.h +++ b/mm/vma.h @@ -32,6 +32,7 @@ struct unlink_vma_file_batch { * vma munmap operation */ struct vma_munmap_struct { + struct mm_struct *mm; /* The mm being operated on */ struct vma_iterator *vmi; struct vm_area_struct *vma; /* The first vma to munmap */ struct vm_area_struct *prev; /* vma before the munmap area */ @@ -459,9 +460,9 @@ bool vma_wants_writenotify(struct vm_area_struct *vma, = pgprot_t vm_page_prot); int mm_take_all_locks(struct mm_struct *mm); void mm_drop_all_locks(struct mm_struct *mm); =20 -unsigned long mmap_region(struct file *file, unsigned long addr, - unsigned long len, vm_flags_t vm_flags, unsigned long pgoff, - struct list_head *uf); +unsigned long mmap_region(struct mm_struct *mm, struct file *file, + unsigned long addr, unsigned long len, vm_flags_t vm_flags, + unsigned long pgoff, struct list_head *uf); =20 int do_brk_flags(struct vma_iterator *vmi, struct vm_area_struct *brkvma, unsigned long addr, unsigned long request, @@ -732,11 +733,12 @@ int relocate_vma_down(struct vm_area_struct *vma, uns= igned long shift); * * Return: false if proposed change is OK, true if not ok and should be de= nied. */ -static inline bool map_deny_write_exec(const vma_flags_t *old, +static inline bool map_deny_write_exec(struct mm_struct *mm, + const vma_flags_t *old, const vma_flags_t *new) { /* If MDWE is disabled, we have nothing to deny. */ - if (!mm_flags_test(MMF_HAS_MDWE, current->mm)) + if (!mm_flags_test(MMF_HAS_MDWE, mm)) return false; =20 /* If the new VMA is not executable, we have nothing to deny. */ diff --git a/tools/testing/vma/include/dup.h b/tools/testing/vma/include/du= p.h index bf26b3f48d3a..6786e43a5a6b 100644 --- a/tools/testing/vma/include/dup.h +++ b/tools/testing/vma/include/dup.h @@ -32,6 +32,7 @@ struct mm_struct { unsigned long data_vm; /* VM_WRITE & ~VM_SHARED & ~VM_STACK */ unsigned long exec_vm; /* VM_EXEC & ~VM_WRITE & ~VM_STACK */ unsigned long stack_vm; /* VM_STACK */ + unsigned long task_size; /* size of task vm space */ =20 union { vm_flags_t def_flags; diff --git a/tools/testing/vma/tests/mmap.c b/tools/testing/vma/tests/mmap.c index c85bc000d1cb..46c9bf12d79f 100644 --- a/tools/testing/vma/tests/mmap.c +++ b/tools/testing/vma/tests/mmap.c @@ -12,19 +12,19 @@ static bool test_mmap_region_basic(void) current->mm =3D &mm; =20 /* Map at 0x300000, length 0x3000. */ - addr =3D __mmap_region(NULL, 0x300000, 0x3000, vma_flags, 0x300, NULL); + addr =3D __mmap_region(&mm, NULL, 0x300000, 0x3000, vma_flags, 0x300, NUL= L); ASSERT_EQ(addr, 0x300000); =20 /* Map at 0x250000, length 0x3000. */ - addr =3D __mmap_region(NULL, 0x250000, 0x3000, vma_flags, 0x250, NULL); + addr =3D __mmap_region(&mm, NULL, 0x250000, 0x3000, vma_flags, 0x250, NUL= L); ASSERT_EQ(addr, 0x250000); =20 /* Map at 0x303000, merging to 0x300000 of length 0x6000. */ - addr =3D __mmap_region(NULL, 0x303000, 0x3000, vma_flags, 0x303, NULL); + addr =3D __mmap_region(&mm, NULL, 0x303000, 0x3000, vma_flags, 0x303, NUL= L); ASSERT_EQ(addr, 0x303000); =20 /* Map at 0x24d000, merging to 0x250000 of length 0x6000. */ - addr =3D __mmap_region(NULL, 0x24d000, 0x3000, vma_flags, 0x24d, NULL); + addr =3D __mmap_region(&mm, NULL, 0x24d000, 0x3000, vma_flags, 0x24d, NUL= L); ASSERT_EQ(addr, 0x24d000); =20 ASSERT_EQ(mm.map_count, 2); --=20 2.43.0 From nobody Sat Jul 25 05:59:38 2026 Received: from mail-pl1-f176.google.com (mail-pl1-f176.google.com [209.85.214.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 52C0E41685B for ; Fri, 24 Jul 2026 22:02:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.176 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930543; cv=none; b=cnqKyY5TmxOnk7EJVjWeYH5aWnmYxYGNWMFN5kaw8mJNNBwWykZBrAOrNcnSc9Uy1ca7Ug5zwXN75WZSxxQIShdF9qB2fkQJKGwR+/oG/N/SgY2Mr7fg3NFjhu4poOFDqOcbXGH1afMd/25hvEu6Jsk6CS/heprz98tYS8xkBOg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930543; c=relaxed/simple; bh=oRGrlChxBMsxVyxpsQ43J9Y1ZaxNsqhcVFf/I6c2bm4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=o260ULv7sTxBsQLbfPjZL3IQe/OpeGAFD32fX1+dHTcBtCidR7/I9rEn3OHzmkt2c/NBVQym7Ex+GGmh7KaWTucsJHE7gCau1gD1y95o4zIey/1E9eHK24ovbS7S4+Q7NHng/K0X5G+Blx1AjOCRmTOtsWEFLdFxItGqpycO28Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=I+WYZcIl; arc=none smtp.client-ip=209.85.214.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="I+WYZcIl" Received: by mail-pl1-f176.google.com with SMTP id d9443c01a7336-2cf27856f9cso12841645ad.2 for ; Fri, 24 Jul 2026 15:02:22 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784930542; x=1785535342; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=dNz1nVhQciAPDe5/sp7R44bDI4wMfSaLt+M5kzutlc4=; b=I+WYZcIlU6cjFeBCxgGFlm/WVizV45lhRpsO7BxANf+wiGXZ6WgmDiXyvwNXSJqt1f r0vubLa4Eb7zmFCSqSYlm+sq9HBIAeeQs6sEPbpliuATGzismKZC9LMTvyhq00YjoN4F lrzGj+vlNXRhRA7xdiAjxOBH2ayJXxgoTz2Tyo/ikiGyejtmdGrm6rKAY2aJ/s+2u+42 z5vzIFqb2xVz+7eoIKIA/iz+v297u0Qz6K0tV9HmLphb2WUrR/7Xn+O4ehuexqaeaYOp m4a1WYtxWK8WREBhW2qmdZaCR9+OIi0pdSWN6AQ1ypk/UWxCOqxp/NaCs9IPhpFL3FBD +FQQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784930542; x=1785535342; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=dNz1nVhQciAPDe5/sp7R44bDI4wMfSaLt+M5kzutlc4=; b=i095A6an9DZM3XXT1rCN/qMa42T8DCFRMLNLFpxvJ4I8rz1AAY9Jlgnm2HtdK6madO X3qKiGpOVGX8VIo9zUaYPUThwGVBPJ5ep7rPxrD92F89olXHLOpjx0ChcM9VI2hTL1SG oQbAzzZgs1cEhdGMdWrv083PYbQ7VzvTnnhzS4i+S3dw8btOSQdPVRg2XGzj+yoWXGqu Nxdt/NLSVSKM/Rb7nLsbpoyhDuG13a8yccZlVhIKN0yuA6rw56FZrt2v5yeRtUl3CON5 YdLdEeZ8UmLC4eKtI9HcsPGPCzaMN8/VMtzdc4XNgOMU0LzLGSrYXX0Roz398C+hxO6B AQhA== X-Forwarded-Encrypted: i=1; AHgh+Rq9teVHSApjANSEndq3RkGmuIRBKMCWca3sgMtmKVqlJfd7yn85dFfdU5hWIbzTX+5WbGTrzvMkPDL1rWs=@vger.kernel.org X-Gm-Message-State: AOJu0YwBJ/+pQWLXlU7aJHA8imjHHbzqdYFFVp0C/qd67T/VApRWu7+U QnJExxpnh4Cv2Lbrxi1pY4HrRhtqeBF59qrJc8HSqPCP5fUaE1wQ20rK X-Gm-Gg: AR+sD11edkGBmDdyFd37vbfD97aGJFyYW+NyOXnd7j17SV94pUVBfUeNaJ6QFdNr16m oC65PcEgxIBR+eN2z53FkMzeUa59xEXIuuNlea9Qc5Wd97BHNyF3qV347DBz/4dYubVPPL/JjHg S8uWyM4bvbngH1kC8SfBsA/+ZJlez8j7jWLTGzI/H+gVGabKDTFos1Jd6thQsjkv8zvcrb/4jhj uerz5TeyadDd8eweIcaJZqo6vEMZxHsBSrIt4ZaX/Pc9KG5FaN5q7AAXzebXs8OM1Vdt1OqsvMi dvwlSVAqALYrSLnf3BF9fZ29d2sulBBegkHQflu3OCOej9loTVHzcnfD178nNdmX/6fdqXDKose bK8xq6FYbLUMIevqY71bqVdTMj80C6Tr7F7J3douKcr81vbyRBgcltTWwFxriCibLemcufdVHDq wGuf68KdybzvasBSMmq7sVmTuu9tUGNZuJbyVLUcyRNa8JAHNUJnN+6xGEY/bn X-Received: by 2002:a17:90b:1e46:b0:387:e0bb:57f4 with SMTP id 98e67ed59e1d1-38f2978d427mr191838a91.37.1784930541375; Fri, 24 Jul 2026 15:02:21 -0700 (PDT) Received: from pop-os.tail4adac7.ts.net ([2601:647:6802:dbc0:546a:e1b0:9574:3a29]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc5a67f3sm2962631eec.29.2026.07.24.15.02.19 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 15:02:20 -0700 (PDT) From: Cong Wang To: Andy Lutomirski Cc: Kees Cook , linux-kernel@vger.kernel.org, Will Drewry , Christian Brauner , Andrew Morton , linux-mm@kvack.org, Cong Wang Subject: [PATCH v7 3/8] seccomp: introduce SECCOMP_IOCTL_NOTIF_PIN_INSTALL Date: Fri, 24 Jul 2026 15:01:42 -0700 Message-ID: <20260724220147.214396-4-xiyou.wangcong@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260724220147.214396-1-xiyou.wangcong@gmail.com> References: <20260724220147.214396-1-xiyou.wangcong@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Cong Wang SECCOMP_IOCTL_NOTIF_PIN_INSTALL maps a supervisor-owned @memfd at @target_addr in the trapped task's mm via vm_mmap_remote(), PROT_READ, MAP_SHARED, MAP_FIXED_NOREPLACE and VM_SEALED. Because the mapping is sealed, neither the target nor a CLONE_VM peer can munmap, mremap, mprotect or MAP_FIXED-stomp it; its contents are immutable from the target's side while the supervisor retains write access through its own mapping of the same memfd. The install needs no target-side cooperation, which is what makes the feature usable for fork+execve sandbox wrappers (Sandlock, Firejail, Bubblewrap-style) that have no trusted post-exec window to install their own mappings. The pin is just a sealed VMA owned by the target's mm: it persists until the task execve()s or exits (a sealed VMA cannot be unmapped piecemeal), and the kernel keeps no per-pin bookkeeping. A supervisor reuses one region across many redirects. Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Cong Wang --- include/linux/seccomp.h | 5 ++ include/uapi/linux/seccomp.h | 34 ++++++++++ kernel/seccomp.c | 126 +++++++++++++++++++++++++++++++++++ 3 files changed, 165 insertions(+) diff --git a/include/linux/seccomp.h b/include/linux/seccomp.h index 9b959972bf4a..a91d1fc8a2b8 100644 --- a/include/linux/seccomp.h +++ b/include/linux/seccomp.h @@ -16,6 +16,11 @@ #define SECCOMP_NOTIFY_ADDFD_SIZE_VER0 24 #define SECCOMP_NOTIFY_ADDFD_SIZE_LATEST SECCOMP_NOTIFY_ADDFD_SIZE_VER0 =20 +/* sizeof() the first published struct seccomp_notif_pin_install */ +#define SECCOMP_NOTIFY_PIN_INSTALL_SIZE_VER0 32 /* up to @size */ +#define SECCOMP_NOTIFY_PIN_INSTALL_SIZE_VER1 40 /* adds @offset */ +#define SECCOMP_NOTIFY_PIN_INSTALL_SIZE_LATEST SECCOMP_NOTIFY_PIN_INSTALL_= SIZE_VER1 + #ifdef CONFIG_SECCOMP =20 #include diff --git a/include/uapi/linux/seccomp.h b/include/uapi/linux/seccomp.h index dbfc9b37fcae..d3249294788b 100644 --- a/include/uapi/linux/seccomp.h +++ b/include/uapi/linux/seccomp.h @@ -137,6 +137,37 @@ struct seccomp_notif_addfd { __u32 newfd_flags; }; =20 +/** + * struct seccomp_notif_pin_install - have the kernel install a sealed + * MAP_SHARED mapping of @memfd into the trapped task's mm at @target_addr. + * + * The supervisor owns @memfd and the kernel installs the mapping without + * target-side cooperation. It is read-only and VM_SEALED, so the target a= nd + * any CLONE_VM peer cannot munmap, mremap, mprotect or MAP_FIXED-stomp it. + * @memfd must be write-sealed (F_SEAL_WRITE or F_SEAL_FUTURE_WRITE, -EINV= AL + * otherwise) so its bytes cannot be rewritten through any other reference= to + * the same memfd. + * + * @id: The ID of an active seccomp notification on this listener, + * identifying the trapped task whose mm receives the pin. + * @flags: Reserved, must be 0. + * @memfd: Supervisor-side fd for the backing memfd. Must be write-sealed. + * @target_addr: Page-aligned address in the trapped task's mm to install = at. + * If non-zero it is MAP_FIXED (no existing mapping may over= lap + * [@target_addr, @target_addr + @size)); if zero the kernel + * picks a free area. The actual address is written back her= e. + * @size: Size of the pin in bytes. Must be page-aligned. + * @offset: Page-aligned byte offset into @memfd to map from. + */ +struct seccomp_notif_pin_install { + __u64 id; + __u32 flags; + __u32 memfd; + __u64 target_addr; + __u64 size; + __u64 offset; +}; + #define SECCOMP_IOC_MAGIC '!' #define SECCOMP_IO(nr) _IO(SECCOMP_IOC_MAGIC, nr) #define SECCOMP_IOR(nr, type) _IOR(SECCOMP_IOC_MAGIC, nr, type) @@ -154,4 +185,7 @@ struct seccomp_notif_addfd { =20 #define SECCOMP_IOCTL_NOTIF_SET_FLAGS SECCOMP_IOW(4, __u64) =20 +#define SECCOMP_IOCTL_NOTIF_PIN_INSTALL SECCOMP_IOWR(5, \ + struct seccomp_notif_pin_install) + #endif /* _UAPI_LINUX_SECCOMP_H */ diff --git a/kernel/seccomp.c b/kernel/seccomp.c index 066909393c38..e894af0e7c78 100644 --- a/kernel/seccomp.c +++ b/kernel/seccomp.c @@ -37,12 +37,18 @@ #ifdef CONFIG_SECCOMP_FILTER #include #include +#include #include #include #include #include #include #include +#include +#include +#include +#include +#include =20 /* * When SECCOMP_IOCTL_NOTIF_ID_VALID was first introduced, it had the @@ -1823,6 +1829,123 @@ static long seccomp_notify_addfd(struct seccomp_fil= ter *filter, return ret; } =20 +static unsigned long seccomp_install_pin(struct mm_struct *mm, + struct file *memfd_file, + unsigned long target_addr, size_t size, + unsigned long offset) +{ + unsigned long ret; + + if (!VM_SEALED) + return -EOPNOTSUPP; + + /* + * Install a sealed, read-only mapping. A fixed request (@target_addr + * !=3D 0) is MAP_FIXED_NOREPLACE: an existing mapping yields -EEXIST + * rather than being silently clobbered. A request of 0 lets the kernel + * pick a free area in the target mm. + */ + ret =3D vm_mmap_remote(mm, memfd_file, target_addr, size, PROT_READ, + MAP_SHARED | MAP_FIXED_NOREPLACE, + offset >> PAGE_SHIFT, VM_SEALED); + if (IS_ERR_VALUE(ret)) + return ret; + if (target_addr && ret !=3D target_addr) + return -ENOMEM; + return ret; +} + +static long seccomp_notify_pin_install(struct seccomp_filter *filter, + struct seccomp_notif_pin_install __user *upin, + unsigned int size) +{ + struct seccomp_notif_pin_install pin; + struct seccomp_knotif *knotif; + struct task_struct *target; + struct file *memfd_file; + struct mm_struct *mm; + unsigned long addr, as_limit, npages; + int seals; + long ret; + + BUILD_BUG_ON(sizeof(pin) < SECCOMP_NOTIFY_PIN_INSTALL_SIZE_VER0); + BUILD_BUG_ON(sizeof(pin) !=3D SECCOMP_NOTIFY_PIN_INSTALL_SIZE_LATEST); + + if (size < SECCOMP_NOTIFY_PIN_INSTALL_SIZE_VER0 || size >=3D PAGE_SIZE) + return -EINVAL; + + ret =3D copy_struct_from_user(&pin, sizeof(pin), upin, size); + if (ret) + return ret; + + if (pin.flags) + return -EINVAL; + if (!pin.size || !IS_ALIGNED(pin.target_addr, PAGE_SIZE) || + !IS_ALIGNED(pin.size, PAGE_SIZE) || !IS_ALIGNED(pin.offset, PAGE_SIZE= )) + return -EINVAL; + if (pin.target_addr + pin.size < pin.target_addr) + return -EINVAL; + if (pin.offset + pin.size < pin.offset) + return -EINVAL; + + memfd_file =3D fget(pin.memfd); + if (!memfd_file) + return -EBADF; + + seals =3D memfd_get_seals(memfd_file); + if (seals < 0 || !(seals & (F_SEAL_WRITE | F_SEAL_FUTURE_WRITE))) { + ret =3D -EINVAL; + goto out_fput; + } + + ret =3D mutex_lock_interruptible(&filter->notify_lock); + if (ret < 0) + goto out_fput; + + knotif =3D find_notification(filter, pin.id); + if (!knotif) { + ret =3D -ENOENT; + goto out_unlock; + } + if (knotif->state !=3D SECCOMP_NOTIFY_SENT) { + ret =3D -EINPROGRESS; + goto out_unlock; + } + + target =3D knotif->task; + mm =3D get_task_mm(target); + as_limit =3D task_rlimit(target, RLIMIT_AS) >> PAGE_SHIFT; + mutex_unlock(&filter->notify_lock); + if (!mm) { + ret =3D -ESRCH; + goto out_fput; + } + + npages =3D pin.size >> PAGE_SHIFT; + if (npages > as_limit || READ_ONCE(mm->total_vm) > as_limit - npages) { + mmput(mm); + ret =3D -ENOMEM; + goto out_fput; + } + + addr =3D seccomp_install_pin(mm, memfd_file, pin.target_addr, pin.size, + pin.offset); + mmput(mm); + if (IS_ERR_VALUE(addr)) + ret =3D addr; + else if (put_user(addr, &upin->target_addr)) + ret =3D -EFAULT; + else + ret =3D 0; + goto out_fput; + +out_unlock: + mutex_unlock(&filter->notify_lock); +out_fput: + fput(memfd_file); + return ret; +} + static long seccomp_notify_ioctl(struct file *file, unsigned int cmd, unsigned long arg) { @@ -1847,6 +1970,9 @@ static long seccomp_notify_ioctl(struct file *file, u= nsigned int cmd, switch (EA_IOCTL(cmd)) { case EA_IOCTL(SECCOMP_IOCTL_NOTIF_ADDFD): return seccomp_notify_addfd(filter, buf, _IOC_SIZE(cmd)); + case EA_IOCTL(SECCOMP_IOCTL_NOTIF_PIN_INSTALL): + return seccomp_notify_pin_install(filter, buf, + _IOC_SIZE(cmd)); default: return -EINVAL; } --=20 2.43.0 From nobody Sat Jul 25 05:59:38 2026 Received: from mail-pg1-f171.google.com (mail-pg1-f171.google.com [209.85.215.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 887DB3FA5E7 for ; Fri, 24 Jul 2026 22:02:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930546; cv=none; b=WeYPdn0+mlRsfRcj3OysxF9rl+30a6D3I7tiZlXnYsiNjsh3LZedKwejzcBaAGobMnTdSFsoJIlhLwE8SfGxYCDEPq6BMvcrnEB2H1HoytlqLQlSZOp8rYDTNIZhTMfvGb20sNYmdXCx1SvWCg18yEb7olsXzlgAi6VCDpGKWi8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930546; c=relaxed/simple; bh=l8LrloIzOK32CMuzElY05harW0usRO0LsWxE3MqYBd0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=B/fH5Xaub8Atw1KcaHDo8INdHrR9wXDH/cT+xGG/q42o8uGjnnZXtKwJlsLhERZAZVcgVqn911rDBifIf2dq7H+DDCrnVNnvojdVR3u3nbEbQtLouwqGpoDLVHSW71RDo3kPa4vHJAogRYaJuAfGBA1HSkeKSfIPox9GfEoodBI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=mctNviMZ; arc=none smtp.client-ip=209.85.215.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="mctNviMZ" Received: by mail-pg1-f171.google.com with SMTP id 41be03b00d2f7-ca7bea5e5b3so707483a12.1 for ; Fri, 24 Jul 2026 15:02:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784930544; x=1785535344; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=6q5nQLizQGV2XlL8vTWKU8w6FEVGfMIRL5kdTFKa2t4=; b=mctNviMZ/WDakgjh1+XmbvqO0bqbQlL5USYYyrrA7e9zkOshlESagjYAAOFY0+fjHY /gkaAjaPQlJHcZZuqOltAw8i9RcN+xTlD8/O8tJfkumLDemUEwK8Ggfl7kn/g3PWFozT RaEFDZK9clKLBoze0wvqHjy7FVhQhon4o1ojEILdNE0ZRylyH1nP//UqHLBKkEB5olaC OL7vNQFhXanoo2cZTP5VCv0mdEPxnqnZJG04Dc6DKIYX8wlZL6IUEwWU3pUDOeYEro67 zoR+39zRmlUTKanhM/1V0LUMJNRVLlrbY/h+SCdgj7pZFjuscu001Sl7Du/httvgOXhr +2VQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784930544; x=1785535344; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=6q5nQLizQGV2XlL8vTWKU8w6FEVGfMIRL5kdTFKa2t4=; b=bvfM8j+PGDxnmZ0jgqrDE48rJvyPJp4TcFbc6MITqE6Z8GxdwW3f1/Elrt9stk0PHN xTlllpKS60XHgvqGaO+5sz+Zso3FjCfRvQrx16/RxVZgQfLWjRjFXLuwmoHeWYjOI6wx uy0J+LrqGMwnHQYRjWYbMEoOqUsgfto5Kk2lrHvh3sUDCGnvMH0wdImKTnGC/DMC5xaw ERcnwJLIpwdPNn9O+/I2OfI0Pv6qM6lI0aJJpo1fuivT9LVOvRb3UC6cBrY+4xd6hD9h s2f69OiDZfm89Tw6EaIltrTfMDvnYaGIu1p4qNxsTPywjqpwELg3Y86u7bpOndCOa0sn V3wQ== X-Forwarded-Encrypted: i=1; AHgh+RrCOfA1KfrCsDxYfkYD4alN1h6eVRyEl4qXMNwr7TTFKHcSr94tltT3Sjpe9m0YNkebH/xO6NGCvfOLAok=@vger.kernel.org X-Gm-Message-State: AOJu0YxKxTFK6VrOmpKfKlR9lAljZGngkdwOVm5KrBuqGQgFUv8LCBK5 fMaEfIGxAHFNN66qt0FJisBGlZnzBbXytV4G35lJBMIr5Zo9BFuVzhNw X-Gm-Gg: AR+sD1271McuNmzxFcWPrqtWZedKQXfBRIFVri3NtVNzYCK2zDjEhKNVHDFPB6ARtbo OpZPxLiKTvNW+VZhzIQd+nDbw7TJweuQHByPnEESbON4JiVZljldbe5egIJAIr6odoYIujabraw Eej3cc+OgyUQkqv6x1SBNHBjNAAXrKqYk12XmG2HuUBvBVoTPMMOtjjN/rq/GCyWb2EA/D5fIsr /ak4zCtpBNP+7lwuVI3qSFhOAttobBg9vUzhuwK6aP10Yos1IZVbVmBjKGLCHV/LEPdvuZ+a/Wa nQRoEu2n4OGdOxMQ1fCuzXP0YxLg/2sbM6swaYJJD5HhAP+ia89Ht525kCRWu8QQhEXt/gBvgBd KjHbMefFOaXqkNPRvNHv+zFNnzjOwBbQqT1qJ7yjhViXkGNZG6MUP+rWMPff6I6stxgSmn2DF5w sJiAbWLB3eWrBpkalvXef/jjr+f7ZTp3y7eAdc8GSw5pXWyQa0GrVfVpoD6hULzNtf3pwK1w8= X-Received: by 2002:a05:6a21:4a8c:b0:3c3:83e6:c6e2 with SMTP id adf61e73a8af0-3c67e0e101fmr109624637.50.1784930543994; Fri, 24 Jul 2026 15:02:23 -0700 (PDT) Received: from pop-os.tail4adac7.ts.net ([2601:647:6802:dbc0:546a:e1b0:9574:3a29]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc5a67f3sm2962631eec.29.2026.07.24.15.02.21 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 15:02:22 -0700 (PDT) From: Cong Wang To: Andy Lutomirski Cc: Kees Cook , linux-kernel@vger.kernel.org, Will Drewry , Christian Brauner , Andrew Morton , linux-mm@kvack.org, Cong Wang Subject: [PATCH v7 4/8] seccomp: add __NR_seccomp_* aliases for rt_sigreturn and clone/fork Date: Fri, 24 Jul 2026 15:01:43 -0700 Message-ID: <20260724220147.214396-5-xiyou.wangcong@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260724220147.214396-1-xiyou.wangcong@gmail.com> References: <20260724220147.214396-1-xiyou.wangcong@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Cong Wang The existing __NR_seccomp_* aliases name only the strict-mode syscalls (read/write/exit/sigreturn). SEND_REDIRECT must also recognise rt_sigreturn and the clone/fork task-creation family, native and compat, so it can refuse to redirect them. Gate these behind a new SECCOMP_ARCH_REDIRECT opt-in, mirroring how SECCOMP_ARCH_NATIVE gates the bitmap cache: an arch declares it only once it supplies a complete, verified set of these numbers, and where it is undefined the feature is compiled out. This avoids a generic fallback silently handing an arch a wrong number that would drop a syscall from the deny list. x86_64 opts in and supplies the ia32 compat numbers. Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Cong Wang --- arch/x86/include/asm/seccomp.h | 15 ++++++++++++++- include/asm-generic/seccomp.h | 30 ++++++++++++++++++++++++++++++ 2 files changed, 44 insertions(+), 1 deletion(-) diff --git a/arch/x86/include/asm/seccomp.h b/arch/x86/include/asm/seccomp.h index 42bcd42d70d1..a911fce504ac 100644 --- a/arch/x86/include/asm/seccomp.h +++ b/arch/x86/include/asm/seccomp.h @@ -6,6 +6,7 @@ =20 #ifdef CONFIG_X86_32 #define __NR_seccomp_sigreturn __NR_sigreturn +#define __NR_seccomp_rt_sigreturn __NR_rt_sigreturn #endif =20 #ifdef CONFIG_COMPAT @@ -14,12 +15,18 @@ #define __NR_seccomp_write_32 __NR_ia32_write #define __NR_seccomp_exit_32 __NR_ia32_exit #define __NR_seccomp_sigreturn_32 __NR_ia32_sigreturn +#define __NR_seccomp_rt_sigreturn_32 __NR_ia32_rt_sigreturn +#define __NR_seccomp_clone_32 __NR_ia32_clone +#define __NR_seccomp_clone3_32 __NR_ia32_clone3 +#define __NR_seccomp_fork_32 __NR_ia32_fork +#define __NR_seccomp_vfork_32 __NR_ia32_vfork #endif =20 #ifdef CONFIG_X86_64 # define SECCOMP_ARCH_NATIVE AUDIT_ARCH_X86_64 # define SECCOMP_ARCH_NATIVE_NR NR_syscalls # define SECCOMP_ARCH_NATIVE_NAME "x86_64" +# define SECCOMP_ARCH_REDIRECT 1 # ifdef CONFIG_COMPAT # define SECCOMP_ARCH_COMPAT AUDIT_ARCH_I386 # define SECCOMP_ARCH_COMPAT_NR IA32_NR_syscalls @@ -28,8 +35,14 @@ /* * x32 will have __X32_SYSCALL_BIT set in syscall number. We don't support * caching them and they are treated as out of range syscalls, which will - * always pass through the BPF filter. + * always pass through the BPF filter. It shares AUDIT_ARCH_X86_64 with the + * native ABI, so refuse to redirect it: the generic denylist keys off pla= in + * syscall numbers and cannot name x32's sigreturn/clone. */ +static inline bool arch_seccomp_redirect_deny(const struct seccomp_data *s= d) +{ + return sd->nr & __X32_SYSCALL_BIT; +} #else /* !CONFIG_X86_64 */ # define SECCOMP_ARCH_NATIVE AUDIT_ARCH_I386 # define SECCOMP_ARCH_NATIVE_NR NR_syscalls diff --git a/include/asm-generic/seccomp.h b/include/asm-generic/seccomp.h index 6b6f42bc58f9..42b1b9b79ddf 100644 --- a/include/asm-generic/seccomp.h +++ b/include/asm-generic/seccomp.h @@ -26,6 +26,36 @@ #define __NR_seccomp_sigreturn __NR_rt_sigreturn #endif =20 +#ifdef SECCOMP_ARCH_REDIRECT +#ifndef __NR_seccomp_rt_sigreturn +#define __NR_seccomp_rt_sigreturn __NR_seccomp_sigreturn +#endif +#ifndef __NR_seccomp_clone +#define __NR_seccomp_clone __NR_clone +#endif +#ifndef __NR_seccomp_clone3 +#ifdef __NR_clone3 +#define __NR_seccomp_clone3 __NR_clone3 +#else +#define __NR_seccomp_clone3 (-1) +#endif +#endif +#ifndef __NR_seccomp_fork +#ifdef __NR_fork +#define __NR_seccomp_fork __NR_fork +#else +#define __NR_seccomp_fork (-1) +#endif +#endif +#ifndef __NR_seccomp_vfork +#ifdef __NR_vfork +#define __NR_seccomp_vfork __NR_vfork +#else +#define __NR_seccomp_vfork (-1) +#endif +#endif +#endif /* SECCOMP_ARCH_REDIRECT */ + #ifdef CONFIG_COMPAT #ifndef get_compat_mode1_syscalls static inline const int *get_compat_mode1_syscalls(void) --=20 2.43.0 From nobody Sat Jul 25 05:59:38 2026 Received: from mail-pj1-f45.google.com (mail-pj1-f45.google.com [209.85.216.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3EF7141D12D for ; Fri, 24 Jul 2026 22:02:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.45 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930550; cv=none; b=Qjp8s0XQJvWBK9Igkfp/1eMFLQxR5rBagDxOmhVlAl/uWz70Xy4GkFsC61e5qL/fNZlVnc9KyZ8ef7BTRgftAf8K8qUpxpf/bH/KrwZWTx7hg4tS8qzcJXU4ch9rgy7sDWUuyqxsazGDlHKcfB3XeSxST27qAWbiN8oF4vWlIpE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930550; c=relaxed/simple; bh=MXNQjtf+TsyeRxmRCDaUn98Q0K/87hjwilGQeNwqNN0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=svyxi1qVZxAmcD4PucfRM1FKy+uqYJ1NaRYMK7M5Fxf5rQy1UeVdI+yK80yeLQtmn+eHaQpQyIPNwqgvIy5WfsrkTqktjQLZKF7uFWI8qrjMYi8gJJyQWH1p768IVFTcBgaeTUOLc/MsP8YcygJyKZly29gJ8qENTFuil0MfDlY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=iFEcZ4JG; arc=none smtp.client-ip=209.85.216.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="iFEcZ4JG" Received: by mail-pj1-f45.google.com with SMTP id 98e67ed59e1d1-38dc69c74b8so884232a91.0 for ; Fri, 24 Jul 2026 15:02:28 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784930547; x=1785535347; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ZL7AIlq0MhT1K4P23mjl1bIfDpk6kn//PjdUqpevit0=; b=iFEcZ4JGJzi1gjfMcAQgd0ER6T94Q1N7WprvaaM9X6VDmN9aw3Dr/lx9WK9wf6qPNc 9852uE6gHixOmusncyMVwlrJ+gf42AEx2x06t/E4R/NeKrlbYetdKoEkqMfJbtr4e3UG PsAyvnu6KPM8Csk2VeSBpkZVTnh+4PmWQK2o9mG/ANnZW83qayfW+aESqcBUmfB051Pg Qgo2X6rDTCgurTot1+MkfEDG8fvk1IhmH+9EezI2JWgh8fsMIvi4BAvHuHS03XTANfwn tFWwjH4IW7sDFyiCLe6kJVIPvI870/hc7SeJGsV1EQTqRifrssg4qFiCirmvoRc5Xigs nJKA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784930547; x=1785535347; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=ZL7AIlq0MhT1K4P23mjl1bIfDpk6kn//PjdUqpevit0=; b=Nq53czkZt8gNr4hG14WOPKfsh2JzFsFLYoRVHm9FowVMWioiGm7fiCxzmmEhZRYytd DQgo1neVBM1axxEYQvgtVLdqNJeaq9jixnz2psPsayfp5Fq/oZDJytzmPKEZ5Og10WFz C88BW8pl7FASATDzH8A2KQTjggk5DFyRQc7wVSCWiSEVzgCUzNZnSpNt3ZVw7FyqW9SG ZBGk41ABcqTJvQuf5lefn1fcZKTjVS42ykzB1lzeVjCCeuOHpWk6tvxXRktoIRMcqEpm 19cZrq9PgG9OdWq6xPZb8C1qkV2WObwmngO/i212pYWJcDomSJATpALq9Ns8VUoTws5d bLRA== X-Forwarded-Encrypted: i=1; AHgh+RodDxge5AZYuJtzScsBFaLJB5aNh3pXAD7F1uaCrYJvb1DciQiThMJt4ykzQw+pDoODXNIreUxLF6mFYfs=@vger.kernel.org X-Gm-Message-State: AOJu0YzvRq8CEj2DGLEucygK0D3xtqPClDRjnDzbcz4QDCtt9bt/PXt7 TKsDINKBfgEsSSd5T/D/zCNC6PUgGgzEEgqc1nkPgm/T8VuYUfOdiddV X-Gm-Gg: AR+sD119QF6It7mwX3dGQ1NAFG1LShE1fUqW/lm/h1tOmVyAv1/RcYDbOuVFAJEA06j /4QBGwudTXlgVmVoXjwycD28e8AQKcjXODnJGAfhYV6122ggm6yXwuEBuKtaPWCgtjtlVrJgvzs segQ7yx+ezL2O6miRXfE7dl45YDAesnDnpaAnnt8HPX2BwUTb5tHbnOHgKU+Q6cbfok19lQ4MVI uE+2MSRPYaGZliM7B9C1Iw9KB8H/mDoEpdHZw0ZU9NFVPZWEbFe5zTMTRrUIDTu8E3YIdP/a6O2 Atazov1YWjvnGOf0vOBRt4yQ8j39C15fKkknhA39rwAoPAu+1PEwxAd5ts3BRNWZUou8yvp9uNV kJuH/sXYFKWMlRfp6hKAbTItAr2Ot5cGunVTVvsMY4E+76nFupXNH7oZPzPbm6Y7XUyxure7bpw pHTEQl9Bf9UgnIsNzDbi4vAOe0vek8LMOAVqlL/vy5GFpykRvrYh0CmsV4d9RG X-Received: by 2002:a17:90b:4b86:b0:38e:a555:8fc3 with SMTP id 98e67ed59e1d1-38f2965654bmr216020a91.41.1784930547299; Fri, 24 Jul 2026 15:02:27 -0700 (PDT) Received: from pop-os.tail4adac7.ts.net ([2601:647:6802:dbc0:546a:e1b0:9574:3a29]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc5a67f3sm2962631eec.29.2026.07.24.15.02.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 15:02:25 -0700 (PDT) From: Cong Wang To: Andy Lutomirski Cc: Kees Cook , linux-kernel@vger.kernel.org, Will Drewry , Christian Brauner , Andrew Morton , linux-mm@kvack.org, Cong Wang Subject: [PATCH v7 5/8] seccomp: add kernel-installed pinned-memfd redirect Date: Fri, 24 Jul 2026 15:01:44 -0700 Message-ID: <20260724220147.214396-6-xiyou.wangcong@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260724220147.214396-1-xiyou.wangcong@gmail.com> References: <20260724220147.214396-1-xiyou.wangcong@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Cong Wang Add SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, which resumes a trapped syscall (like SECCOMP_USER_NOTIF_FLAG_CONTINUE) with selected argument registers rewritten to point into a pin installed by SECCOMP_IOCTL_NOTIF_PIN_INSTALL. This closes the user-notification TOCTOU for fork+execve sandboxes: the kernel acts on an immutable, supervisor-controlled sealed mapping instead of memory a CLONE_VM peer can rewrite after the check. The feature is gated behind SECCOMP_FILTER_FLAG_REDIRECT, declared at listener creation (it requires SECCOMP_FILTER_FLAG_NEW_LISTENER) and required by both ioctls. At most one redirect-capable filter may exist in a chain (-EBUSY otherwise), so a redirect has a single, unambiguous register fixup. The supervisor supplies an args_mask, a ptr_mask and replacement values. Each pointer substitution is validated by seccomp_pin_check(): the access [args[i], args[i] + ptr_len[i]) must lie in a single VM_SEALED, read-only, MAP_SHARED VMA still backed by the named memfd. The kernel keeps no bookkeeping; after execve or exit the VMA is gone and validation returns -EFAULT. Original arg registers are saved and restored at user-mode return (via a TWA_RESUME task_work, so a restartable syscall is not turned into a livelock), preserving the caller-saved arg-register ABI. The restore is skipped after a successful execve. rt_sigreturn is refused (-EOPNOTSUPP): it restores the whole register frame and takes no arguments to substitute. The whole redirect path is now gated by SECCOMP_ARCH_REDIRECT, the opt-in arch declares once its deny list syscall numbers are complete and verified. Only x86_64 opts in so far. Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Cong Wang --- include/linux/seccomp.h | 7 +- include/uapi/linux/seccomp.h | 55 +++++- kernel/seccomp.c | 320 +++++++++++++++++++++++++++++++++++ 3 files changed, 380 insertions(+), 2 deletions(-) diff --git a/include/linux/seccomp.h b/include/linux/seccomp.h index a91d1fc8a2b8..5d53f8fce508 100644 --- a/include/linux/seccomp.h +++ b/include/linux/seccomp.h @@ -10,7 +10,8 @@ SECCOMP_FILTER_FLAG_SPEC_ALLOW | \ SECCOMP_FILTER_FLAG_NEW_LISTENER | \ SECCOMP_FILTER_FLAG_TSYNC_ESRCH | \ - SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV) + SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV | \ + SECCOMP_FILTER_FLAG_REDIRECT) =20 /* sizeof() the first published struct seccomp_notif_addfd */ #define SECCOMP_NOTIFY_ADDFD_SIZE_VER0 24 @@ -21,6 +22,10 @@ #define SECCOMP_NOTIFY_PIN_INSTALL_SIZE_VER1 40 /* adds @offset */ #define SECCOMP_NOTIFY_PIN_INSTALL_SIZE_LATEST SECCOMP_NOTIFY_PIN_INSTALL_= SIZE_VER1 =20 +/* sizeof() the first published struct seccomp_notif_resp_redirect */ +#define SECCOMP_NOTIFY_RESP_REDIRECT_SIZE_VER0 120 +#define SECCOMP_NOTIFY_RESP_REDIRECT_SIZE_LATEST SECCOMP_NOTIFY_RESP_REDIR= ECT_SIZE_VER0 + #ifdef CONFIG_SECCOMP =20 #include diff --git a/include/uapi/linux/seccomp.h b/include/uapi/linux/seccomp.h index d3249294788b..bb5875f72556 100644 --- a/include/uapi/linux/seccomp.h +++ b/include/uapi/linux/seccomp.h @@ -25,6 +25,12 @@ #define SECCOMP_FILTER_FLAG_TSYNC_ESRCH (1UL << 4) /* Received notifications wait in killable state (only respond to fatal si= gnals) */ #define SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (1UL << 5) +/* + * Declares that this listener's notifier may issue + * SECCOMP_IOCTL_NOTIF_PIN_INSTALL / SECCOMP_IOCTL_NOTIF_SEND_REDIRECT. At= most + * one such filter may exist in a task's filter chain. Requires NEW_LISTEN= ER. + */ +#define SECCOMP_FILTER_FLAG_REDIRECT (1UL << 6) =20 /* * All BPF programs must return a 32-bit value. @@ -139,7 +145,9 @@ struct seccomp_notif_addfd { =20 /** * struct seccomp_notif_pin_install - have the kernel install a sealed - * MAP_SHARED mapping of @memfd into the trapped task's mm at @target_addr. + * MAP_SHARED mapping of @memfd into the trapped task's mm at @target_addr, + * which SECCOMP_IOCTL_NOTIF_SEND_REDIRECT can then use as a target for + * substituted pointer arguments. * * The supervisor owns @memfd and the kernel installs the mapping without * target-side cooperation. It is read-only and VM_SEALED, so the target a= nd @@ -168,6 +176,45 @@ struct seccomp_notif_pin_install { __u64 offset; }; =20 +#define SECCOMP_REDIRECT_ARGS 6 + +/** + * struct seccomp_notif_resp_redirect - resume the trapped syscall with + * substituted arg-register values, optionally pointing into an installed + * pinned-memfd region. + * + * Like SECCOMP_USER_NOTIF_FLAG_CONTINUE the syscall runs, but the kernel + * first rewrites the arg registers in @args_mask. Pointer substitutions + * (@ptr_mask) are validated against the trapped task's live mapping of + * @memfd, so a target that has exited or execve()d simply fails validatio= n. + * Original registers are restored at syscall exit, skipped after a succes= sful + * execve whose fresh register file must not be clobbered. + * + * @id: The ID of the seccomp notification this response consumes. + * @flags: SECCOMP_REDIRECT_FLAG_*. CONTINUE must be set. + * @args_mask: Bit i set means args[i] replaces arg register i before the + * syscall runs. + * @ptr_mask: Subset of @args_mask. Bit i set means args[i] is a pointer w= hose + * access [args[i], args[i] + ptr_len[i]) must lie inside a sin= gle + * VM_SEALED, read-only mapping of @memfd. Scalars (in @args_ma= sk + * but not @ptr_mask) are written verbatim. + * @memfd: Supervisor-side fd for the backing memfd. Consulted only when + * @ptr_mask is non-zero. + * @args: Replacement values for the arg registers. + * @ptr_len: For each bit set in @ptr_mask, the byte length of the access = at + * args[i]; must be non-zero and args[i] + ptr_len[i] must not + * overflow. Must be 0 where @ptr_mask bit i is clear. + */ +struct seccomp_notif_resp_redirect { + __u64 id; + __u32 flags; + __u32 args_mask; + __u32 ptr_mask; + __u32 memfd; + __u64 args[SECCOMP_REDIRECT_ARGS]; + __u64 ptr_len[SECCOMP_REDIRECT_ARGS]; +}; + #define SECCOMP_IOC_MAGIC '!' #define SECCOMP_IO(nr) _IO(SECCOMP_IOC_MAGIC, nr) #define SECCOMP_IOR(nr, type) _IOR(SECCOMP_IOC_MAGIC, nr, type) @@ -188,4 +235,10 @@ struct seccomp_notif_pin_install { #define SECCOMP_IOCTL_NOTIF_PIN_INSTALL SECCOMP_IOWR(5, \ struct seccomp_notif_pin_install) =20 +#define SECCOMP_IOCTL_NOTIF_SEND_REDIRECT SECCOMP_IOW(6, \ + struct seccomp_notif_resp_redirect) + +/* Valid flags for struct seccomp_notif_resp_redirect. */ +#define SECCOMP_REDIRECT_FLAG_CONTINUE (1UL << 0) + #endif /* _UAPI_LINUX_SECCOMP_H */ diff --git a/kernel/seccomp.c b/kernel/seccomp.c index e894af0e7c78..f23dbee12ab7 100644 --- a/kernel/seccomp.c +++ b/kernel/seccomp.c @@ -211,6 +211,9 @@ static inline void seccomp_cache_prepare(struct seccomp= _filter *sfilter) * @log: true if all actions except for SECCOMP_RET_ALLOW should be logged * @wait_killable_recv: Put notifying process in killable state once the * notification is received by the userspace listener. + * @redirect_capable: true if installed with SECCOMP_FILTER_FLAG_REDIRECT,= so + * this filter's notifier may issue SEND_REDIRECT. At most + * one filter in a stack may be redirect-capable. * @prev: points to a previously installed, or inherited, filter * @prog: the BPF program to evaluate * @notif: the struct that holds all notification related information @@ -232,6 +235,7 @@ struct seccomp_filter { refcount_t users; bool log; bool wait_killable_recv; + bool redirect_capable; struct action_cache cache; struct seccomp_filter *prev; struct bpf_prog *prog; @@ -952,6 +956,13 @@ static long seccomp_attach_filter(unsigned int flags, } } =20 + if (flags & SECCOMP_FILTER_FLAG_REDIRECT) { + for (walker =3D current->seccomp.filter; walker; + walker =3D walker->prev) + if (walker->redirect_capable) + return -EBUSY; + } + /* Set log flag, if present. */ if (flags & SECCOMP_FILTER_FLAG_LOG) filter->log =3D true; @@ -960,6 +971,10 @@ static long seccomp_attach_filter(unsigned int flags, if (flags & SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV) filter->wait_killable_recv =3D true; =20 + /* Set redirect-capable flag, if present. */ + if (flags & SECCOMP_FILTER_FLAG_REDIRECT) + filter->redirect_capable =3D true; + /* * If there is an existing filter, make it the prev and don't drop its * task reference. @@ -1946,6 +1961,299 @@ static long seccomp_notify_pin_install(struct secco= mp_filter *filter, return ret; } =20 +#ifdef SECCOMP_ARCH_REDIRECT +static bool seccomp_pin_access_ok(struct mm_struct *mm, struct file *memfd= _file, + u64 ptr, u64 len) +{ + struct vm_area_struct *vma; + u64 end =3D ptr + len; + + if (!len || end < ptr) + return false; + + vma =3D vma_lookup(mm, ptr); + if (!vma || end > vma->vm_end) + return false; + /* + * The access must lie in a single sealed, read-only, MAP_SHARED, + * memfd-backed VMA. VM_SHARED is required so the bytes the kernel reads + * are the memfd's own pages: a MAP_PRIVATE mapping would resolve to + * anonymous COW copies the target could have written before sealing it + * read-only, defeating the guarantee. + */ + if (!(vma->vm_flags & VM_SEALED) || !(vma->vm_flags & VM_SHARED) || + (vma->vm_flags & VM_WRITE)) + return false; + if (!vma->vm_file) + return false; + return file_inode(vma->vm_file) =3D=3D file_inode(memfd_file); +} + +static bool seccomp_pin_check(struct task_struct *target, + struct file *memfd_file, u32 ptr_mask, + const u64 *args, const u64 *ptr_len) +{ + struct mm_struct *mm; + bool ok =3D true; + int i; + + mm =3D get_task_mm(target); + if (!mm) + return false; + + mmap_read_lock(mm); + for (i =3D 0; i < SECCOMP_REDIRECT_ARGS; i++) { + if (!(ptr_mask & (1U << i))) + continue; + if (!seccomp_pin_access_ok(mm, memfd_file, args[i], + ptr_len[i])) { + ok =3D false; + break; + } + } + mmap_read_unlock(mm); + mmput(mm); + return ok; +} + +struct seccomp_redirect_restore { + struct callback_head twork; + unsigned long orig_args[SECCOMP_REDIRECT_ARGS]; + u32 args_mask; /* bit i: arg i was substituted, restore it */ + u64 self_exec_id; /* snapshot to detect an intervening execve */ +}; + +static void seccomp_redirect_restore_cb(struct callback_head *cb) +{ + struct seccomp_redirect_restore *r =3D + container_of(cb, struct seccomp_redirect_restore, twork); + unsigned long args[SECCOMP_REDIRECT_ARGS]; + long ret, err; + int i; + + if (READ_ONCE(current->self_exec_id) !=3D r->self_exec_id) { + kfree(r); + return; + } + + err =3D syscall_get_error(current, current_pt_regs()); + ret =3D syscall_get_return_value(current, current_pt_regs()); + + syscall_get_arguments(current, current_pt_regs(), args); + for (i =3D 0; i < SECCOMP_REDIRECT_ARGS; i++) + if (r->args_mask & (1U << i)) + args[i] =3D r->orig_args[i]; + syscall_set_arguments(current, current_pt_regs(), args); + + syscall_set_return_value(current, current_pt_regs(), err, ret); + + kfree(r); +} + +/* + * sigreturn/rt_sigreturn restore the entire register frame from the user + * signal stack; the SEND_REDIRECT register-restore (run from task_work at + * user-mode return) would corrupt that frame, and the syscall takes no + * arguments to substitute anyway. Refuse to redirect any of them, includi= ng + * the compat variants (legacy sigreturn and rt_sigreturn are distinct + * numbers). x32 rt_sigreturn carries __X32_SYSCALL_BIT and is not matched + * here; the deprecated x32 ABI is out of scope for redirect. + */ +static bool seccomp_redirect_is_sigreturn(const struct seccomp_data *sd) +{ +#ifdef SECCOMP_ARCH_COMPAT + if (sd->arch =3D=3D SECCOMP_ARCH_COMPAT) + return sd->nr =3D=3D __NR_seccomp_sigreturn_32 || + sd->nr =3D=3D __NR_seccomp_rt_sigreturn_32; +#endif + return sd->nr =3D=3D __NR_seccomp_sigreturn || + sd->nr =3D=3D __NR_seccomp_rt_sigreturn; +} + +/* + * clone/fork-family syscalls copy the trapped task's register frame into = the + * new child, which gets no restore task_work of its own. A redirected task + * creation would leave the child in user space with the substituted (pinn= ed) + * values still in its caller-saved arg registers, breaking the ABI the re= store + * preserves for the parent. There is no use case for redirecting task cre= ation, + * so refuse it. Syscalls absent on an arch resolve to -1 and never match. + */ +static bool seccomp_redirect_is_task_create(const struct seccomp_data *sd) +{ +#ifdef SECCOMP_ARCH_COMPAT + if (sd->arch =3D=3D SECCOMP_ARCH_COMPAT) + return sd->nr =3D=3D __NR_seccomp_clone_32 || + sd->nr =3D=3D __NR_seccomp_clone3_32 || + sd->nr =3D=3D __NR_seccomp_fork_32 || + sd->nr =3D=3D __NR_seccomp_vfork_32; +#endif + return sd->nr =3D=3D __NR_seccomp_clone || + sd->nr =3D=3D __NR_seccomp_clone3 || + sd->nr =3D=3D __NR_seccomp_fork || + sd->nr =3D=3D __NR_seccomp_vfork; +} + +static bool seccomp_redirect_deny(const struct seccomp_data *sd) +{ + return arch_seccomp_redirect_deny(sd) || + seccomp_redirect_is_sigreturn(sd) || + seccomp_redirect_is_task_create(sd); +} + +static long seccomp_notify_send_redirect(struct seccomp_filter *filter, + struct seccomp_notif_resp_redirect __user *uresp, + unsigned int size) +{ + unsigned long args[SECCOMP_REDIRECT_ARGS]; + struct seccomp_redirect_restore *restore; + struct seccomp_notif_resp_redirect resp; + struct file *memfd_file =3D NULL; + struct seccomp_knotif *knotif; + struct pt_regs *target_regs; + long ret; + int i; + + BUILD_BUG_ON(sizeof(resp) < SECCOMP_NOTIFY_RESP_REDIRECT_SIZE_VER0); + BUILD_BUG_ON(sizeof(resp) !=3D SECCOMP_NOTIFY_RESP_REDIRECT_SIZE_LATEST); + + if (!filter->redirect_capable) + return -EPERM; + + if (size < SECCOMP_NOTIFY_RESP_REDIRECT_SIZE_VER0 || size >=3D PAGE_SIZE) + return -EINVAL; + + ret =3D copy_struct_from_user(&resp, sizeof(resp), uresp, size); + if (ret) + return ret; + + if (!(resp.flags & SECCOMP_REDIRECT_FLAG_CONTINUE)) + return -EINVAL; + if (resp.flags & ~SECCOMP_REDIRECT_FLAG_CONTINUE) + return -EINVAL; + if (resp.args_mask & ~((1U << SECCOMP_REDIRECT_ARGS) - 1)) + return -EINVAL; + if (resp.ptr_mask & ~resp.args_mask) + return -EINVAL; + if (!resp.args_mask) + return -EINVAL; + for (i =3D 0; i < SECCOMP_REDIRECT_ARGS; i++) { + if (resp.ptr_mask & (1U << i)) { + if (!resp.ptr_len[i]) + return -EINVAL; + } else if (resp.ptr_len[i]) { + return -EINVAL; + } + } + if (resp.ptr_mask) { + memfd_file =3D fget(resp.memfd); + if (!memfd_file) + return -EBADF; + } + + restore =3D kzalloc_obj(*restore, GFP_KERNEL_ACCOUNT); + if (!restore) { + ret =3D -ENOMEM; + goto out_free; + } + init_task_work(&restore->twork, seccomp_redirect_restore_cb); + + ret =3D mutex_lock_interruptible(&filter->notify_lock); + if (ret < 0) + goto out_free; + + knotif =3D find_notification(filter, resp.id); + if (!knotif) { + ret =3D -ENOENT; + goto out_unlock_free; + } + if (knotif->state !=3D SECCOMP_NOTIFY_SENT) { + ret =3D -EINPROGRESS; + goto out_unlock_free; + } + + if (seccomp_redirect_deny(knotif->data)) { + ret =3D -EOPNOTSUPP; + goto out_unlock_free; + } + +#ifdef SECCOMP_ARCH_COMPAT + if (knotif->data->arch =3D=3D SECCOMP_ARCH_COMPAT) { + for (i =3D 0; i < SECCOMP_REDIRECT_ARGS; i++) + if (resp.ptr_mask & (1U << i)) + resp.args[i] =3D (u32)resp.args[i]; + } +#endif + + if (resp.ptr_mask && + !seccomp_pin_check(knotif->task, memfd_file, resp.ptr_mask, + resp.args, resp.ptr_len)) { + ret =3D -EFAULT; + goto out_unlock_free; + } + + target_regs =3D task_pt_regs(knotif->task); + syscall_get_arguments(knotif->task, target_regs, args); + + for (i =3D 0; i < SECCOMP_REDIRECT_ARGS; i++) + restore->orig_args[i] =3D args[i]; + restore->args_mask =3D resp.args_mask; + restore->self_exec_id =3D READ_ONCE(knotif->task->self_exec_id); + + for (i =3D 0; i < SECCOMP_REDIRECT_ARGS; i++) + if (resp.args_mask & (1U << i)) + args[i] =3D resp.args[i]; + syscall_set_arguments(knotif->task, target_regs, args); + + /* + * Use TWA_RESUME, not TWA_SIGNAL. TWA_SIGNAL sets TIF_NOTIFY_SIGNAL, + * which makes signal_pending() true for the entire redirected syscall + * An interruptible syscall would then bail out with -ERESTARTSYS before + * doing any work, restart, re-trap and get redirected again, it is a + * livelock. TWA_RESUME does not feed signal_pending(), and the restore + * still runs before signal delivery: get_signal() runs task_work_run() + * before it dequeues a signal, so the original args are back in pt_regs + * before handle_signal() builds the sigframe or the -ERESTART* path + * rewinds for restart. + */ + ret =3D task_work_add(knotif->task, &restore->twork, TWA_RESUME); + if (ret) { + for (i =3D 0; i < SECCOMP_REDIRECT_ARGS; i++) + args[i] =3D restore->orig_args[i]; + syscall_set_arguments(knotif->task, target_regs, args); + goto out_unlock_free; + } + + knotif->state =3D SECCOMP_NOTIFY_REPLIED; + knotif->error =3D 0; + knotif->val =3D 0; + knotif->flags =3D SECCOMP_USER_NOTIF_FLAG_CONTINUE; + if (filter->notif->flags & SECCOMP_USER_NOTIF_FD_SYNC_WAKE_UP) + complete_on_current_cpu(&knotif->ready); + else + complete(&knotif->ready); + + mutex_unlock(&filter->notify_lock); + if (memfd_file) + fput(memfd_file); + return 0; + +out_unlock_free: + mutex_unlock(&filter->notify_lock); +out_free: + if (memfd_file) + fput(memfd_file); + kfree(restore); + return ret; +} +#else /* !SECCOMP_ARCH_REDIRECT */ +static long seccomp_notify_send_redirect(struct seccomp_filter *filter, + struct seccomp_notif_resp_redirect __user *uresp, + unsigned int size) +{ + return -EOPNOTSUPP; +} +#endif /* SECCOMP_ARCH_REDIRECT */ + static long seccomp_notify_ioctl(struct file *file, unsigned int cmd, unsigned long arg) { @@ -1973,6 +2281,9 @@ static long seccomp_notify_ioctl(struct file *file, u= nsigned int cmd, case EA_IOCTL(SECCOMP_IOCTL_NOTIF_PIN_INSTALL): return seccomp_notify_pin_install(filter, buf, _IOC_SIZE(cmd)); + case EA_IOCTL(SECCOMP_IOCTL_NOTIF_SEND_REDIRECT): + return seccomp_notify_send_redirect(filter, buf, + _IOC_SIZE(cmd)); default: return -EINVAL; } @@ -2112,6 +2423,15 @@ static long seccomp_set_mode_filter(unsigned int fla= gs, ((flags & SECCOMP_FILTER_FLAG_NEW_LISTENER) =3D=3D 0)) return -EINVAL; =20 +#ifndef SECCOMP_ARCH_REDIRECT + /* Redirect denylist inputs are not verified on this arch. */ + if (flags & SECCOMP_FILTER_FLAG_REDIRECT) + return -EINVAL; +#endif + if ((flags & SECCOMP_FILTER_FLAG_REDIRECT) && + ((flags & SECCOMP_FILTER_FLAG_NEW_LISTENER) =3D=3D 0)) + return -EINVAL; + /* Prepare the new filter before holding any locks. */ prepared =3D seccomp_prepare_user_filter(filter); if (IS_ERR(prepared)) --=20 2.43.0 From nobody Sat Jul 25 05:59:38 2026 Received: from mail-pf1-f171.google.com (mail-pf1-f171.google.com [209.85.210.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5966641BA9B for ; Fri, 24 Jul 2026 22:02:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930551; cv=none; b=bkvKitMSI5d1mtOzGQkSngmSmzy5aM2SlbLprGUUEevNASH5ql0+ihC123IdhN/PlT7LRftmJqm8ZzPI3vd2Vnb6nwzy1FI9stNy0J2cNlH3D6B1gbcpGvb847lsxNaJ7UMG/hgX6EOhVK0FK4BdX63p44VV99rWCDWwakTgjIk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930551; c=relaxed/simple; bh=b9+clYgP6mL9bW6u2CvbC18jZah0tWLSqjNjkcpfOTw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XR7VySLZL+/5aSk0B0KLgHLD/vrrhmqnJSMqy2NU3ogqLzV3Eud3bpj0lxCU+Aqwav0NrGOpdGiQmXRkhSZMZ5+HGc4QkEJgZIfK7HI9WBucLINVPLXNv/b1LMn17mst8U86EfommVnrae20NFuzRyWZs4wBrqg0Skn/lAASYTc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=qWsxaAVq; arc=none smtp.client-ip=209.85.210.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="qWsxaAVq" Received: by mail-pf1-f171.google.com with SMTP id d2e1a72fcca58-845c92bc464so778343b3a.2 for ; Fri, 24 Jul 2026 15:02:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784930550; x=1785535350; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PAzNTf9Ao1Jd+i4HQ1y+PwV6f9J/K6SmB5ui1Xz+21c=; b=qWsxaAVqtl+bmvXjs9/deQojKDEOV5XRxWkm8kBVnH5Js2P+rHcrdeMlbQ5RzFNwgs 5o1PqiJN8nUxmk2DehWbagYR2ua7oPlsGXneAINu8yLKnfVwJ8Lb+SScSQlEwSU2TgY2 OzT4XrYBsei2Uy9C+sTtB/P8kQz58fIKD70XVCZKMqiwQ4dv0eL20rN8G6nCvPejPF3j rr7KsiSYnM49u+PSAoqehxp+dkOxse7K9i0BPIfvUFBEMJ8Tv7+aPKl6SanXQ818R7f9 42B38reVq9kyGuHgvQ3u3A7dzpqPQCxxZ1yYfG/GT3sw0jDtBYVTya8SXB1GnZ/qlkll VKkA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784930550; x=1785535350; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=PAzNTf9Ao1Jd+i4HQ1y+PwV6f9J/K6SmB5ui1Xz+21c=; b=rDdrXOgCSZFjwu1wcnnifUjaxMrQVJDHQZU4VK2QrLdSm6vdm6JsaMDPMqFjvmne8F NvlSp9eeSUaHkQNWTONlTkrEQ6KsQmZwoDdd/g0FCpBKx5Pype0cbXsrcpvismJn1VeA nYymn5RrHj0RlIRdltA33JhrTgo6eEIZyZTTtsF2CoqCiKR9rhkM7g/3M4hGKgEMklPJ 2ZdWOlWddAr4xbN+AWKoSn8OXuwPvkNUAZf9H3q42RvOdH1J8JlznMG3uhfFnWkpJvsS XjdH49KSMRd230so+cjY7wxEZ7pOFQgZRb4+Ksb/4QiPnxzddxCPO6G41kxjbMRS55dw mS1w== X-Forwarded-Encrypted: i=1; AHgh+RqOMn7baxOa0lXSxkuZvMwHikZp8T+nZGj2DYBkuPrR/6rpxgiR1QI/kkXcRcMWtthdS0LIZHXibOwjVY0=@vger.kernel.org X-Gm-Message-State: AOJu0Yy6Fr5TWzauEzRz4u/yWucEUHzrkexdSh2XW/0Vi0gMTkFijj1g jJx7S6RCNhfHLo+ufRrCgyTsRc+Ry7T2rATi71zmwji/aq3kMX3w6KCR X-Gm-Gg: AR+sD13M6Ic1P3Pgd6BLey2mDFvl1l5Ij5kNoE5bkUJQWK5kumKo9iYpmAzUsJgMZ4P drey60CHBAjbvtP+jEAYcLCFhDUXTHIr+IaXF/I5lIxQcqQaetBVwwAVU2Pksmc2EqZt+uMBwvj EkU1oeOaYEUaElGFQmEX+kqGwF8NMIodt+f7I00EGFQGfrRCAMNji4z7Ass07O+2vRQiLz4xCHZ DFtcNqX8/cTo+694SK5BcI/7yjC7NvC71sGLW0DB2knQPUbUMEek/KljsebzpRhbeYMlKJf8rGZ LZoE0qIMpSMhDVior9Lknc/1LEnIW7p7Lqmap/R7/TKcAIf/cyFMPGLoSzZccHyJ2TScwpdRFLL WjyOGNztYHuRMeKwyJPmd/DyyMU9xHterGtfHAshb9WT0mhI0x6EMzzmsA2yKTlm7mXbTX0i/V7 u1wbcdHVKF2JBt9TmvqY+Z+iV8scUtOqBMaLpqQtrPiI2lO5u7RnQ65JVFU9IN X-Received: by 2002:a05:6a21:495:b0:39f:59c8:f302 with SMTP id adf61e73a8af0-3c67df10d07mr125091637.37.1784930549615; Fri, 24 Jul 2026 15:02:29 -0700 (PDT) Received: from pop-os.tail4adac7.ts.net ([2601:647:6802:dbc0:546a:e1b0:9574:3a29]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc5a67f3sm2962631eec.29.2026.07.24.15.02.27 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 15:02:28 -0700 (PDT) From: Cong Wang To: Andy Lutomirski Cc: Kees Cook , linux-kernel@vger.kernel.org, Will Drewry , Christian Brauner , Andrew Morton , linux-mm@kvack.org, Cong Wang Subject: [PATCH v7 6/8] seccomp: re-validate a redirected syscall against outer filters Date: Fri, 24 Jul 2026 15:01:45 -0700 Message-ID: <20260724220147.214396-7-xiyou.wangcong@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260724220147.214396-1-xiyou.wangcong@gmail.com> References: <20260724220147.214396-1-xiyou.wangcong@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Cong Wang Stacked filters compose by taking the most restrictive verdict over one evaluation of a single seccomp_data, assuming the syscall they voted on is the syscall that runs. SECCOMP_IOCTL_NOTIF_SEND_REDIRECT breaks that: the supervisor rewrites the argument registers and the syscall resumes without the stack being re-consulted, so an inner, container-installed filter can redirect a syscall into a form an outer filter would have blocked. Close the hole with seccomp_redirect_revalidate(): after a redirect it walks from the notifier outward, judging the substituted syscall one filter at a time; the innermost filter that does not allow it decides. ALLOW and LOG fall through; ERRNO, TRAP and KILL are terminal; USER_NOTIF consults the outer supervisor, whose plain FLAG_CONTINUE keeps the walk going; TRACE fails closed with -ENOSYS, since a tracer rewrite cannot be soundly re-composed mid-walk. The walk is strictly outward, so the notifier is never reconsulted and no re-notify loop exists, so a deep chain cannot exhaust the kernel stack. A redirect never changes the syscall number, only the argument registers differ. Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Cong Wang --- kernel/seccomp.c | 141 ++++++++++++++++++++++++++++++++++++++++++----- 1 file changed, 126 insertions(+), 15 deletions(-) diff --git a/kernel/seccomp.c b/kernel/seccomp.c index f23dbee12ab7..2475aff55ec2 100644 --- a/kernel/seccomp.c +++ b/kernel/seccomp.c @@ -93,6 +93,13 @@ struct seccomp_knotif { long val; u32 flags; =20 + /* + * Set by SEND_REDIRECT: the reply rewrote the syscall's registers, + * so on resume the syscall must be re-evaluated against the filters + * outer to the one that notified (see __seccomp_filter()). + */ + bool redirect; + /* * Signals when this has changed states, such as the listener * dying, a new seccomp addfd message, or changing to REPLIED @@ -1183,10 +1190,12 @@ static bool should_sleep_killable(struct seccomp_fi= lter *match, =20 static int seccomp_do_user_notification(int this_syscall, struct seccomp_filter *match, - const struct seccomp_data *sd) + const struct seccomp_data *sd, + bool *redirected) { int err; u32 flags =3D 0; + bool redirect =3D false; long ret =3D 0; struct seccomp_knotif n =3D {}; struct seccomp_kaddfd *addfd, *tmp; @@ -1243,6 +1252,7 @@ static int seccomp_do_user_notification(int this_sysc= all, ret =3D n.val; err =3D n.error; flags =3D n.flags; + redirect =3D n.redirect; =20 interrupted: /* If there were any pending addfd calls, clear them out */ @@ -1269,19 +1279,120 @@ static int seccomp_do_user_notification(int this_s= yscall, mutex_unlock(&match->notify_lock); =20 /* Userspace requests to continue the syscall. */ - if (flags & SECCOMP_USER_NOTIF_FLAG_CONTINUE) + if (flags & SECCOMP_USER_NOTIF_FLAG_CONTINUE) { + *redirected =3D redirect; return 0; + } =20 syscall_set_return_value(current, current_pt_regs(), err, ret); return -1; } =20 +static void seccomp_kill_task(int this_syscall, u32 action, int data) +{ + current->seccomp.mode =3D SECCOMP_MODE_DEAD; + seccomp_log(this_syscall, SIGSYS, action, true); + /* Dump core only if this is the last remaining thread. */ + if (action !=3D SECCOMP_RET_KILL_THREAD || + (atomic_read(¤t->signal->live) =3D=3D 1)) { + /* Show the original registers in the dump. */ + syscall_rollback(current, current_pt_regs()); + /* Trigger a coredump with SIGSYS */ + force_sig_seccomp(this_syscall, data, true); + } else { + do_exit(SIGSYS); + } +} + +static int seccomp_redirect_revalidate(struct seccomp_filter *notifier) +{ + struct seccomp_filter *f; + struct seccomp_data sd; + bool redirected =3D false; + int this_syscall; + u32 action; + int data; + + populate_seccomp_data(&sd); + this_syscall =3D sd.nr; + + for (f =3D notifier->prev; f; f =3D f->prev) { + u32 cur_ret =3D bpf_prog_run_pin_on_cpu(f->prog, &sd); + + data =3D cur_ret & SECCOMP_RET_DATA; + action =3D cur_ret & SECCOMP_RET_ACTION_FULL; + + switch (action) { + case SECCOMP_RET_ALLOW: + continue; + + case SECCOMP_RET_LOG: + seccomp_log(this_syscall, 0, action, true); + continue; + + case SECCOMP_RET_ERRNO: + /* Set low-order bits as an errno, capped at MAX_ERRNO. */ + if (data > MAX_ERRNO) + data =3D MAX_ERRNO; + syscall_set_return_value(current, current_pt_regs(), + -data, 0); + goto skip; + + case SECCOMP_RET_TRAP: + /* Show the handler the original registers. */ + syscall_rollback(current, current_pt_regs()); + /* Let the filter pass back 16 bits of data. */ + force_sig_seccomp(this_syscall, data, false); + goto skip; + + case SECCOMP_RET_USER_NOTIF: + /* + * The outer supervisor judges the substituted call: + * an error reply skips it, a plain FLAG_CONTINUE + * keeps the walk going, or this reply would slip the + * call past a stricter filter further out. It cannot + * redirect again: at most one redirect-capable + * listener exists in a chain, and the walk starts + * outside it. + */ + if (seccomp_do_user_notification(this_syscall, f, &sd, + &redirected)) + goto skip; + continue; + + case SECCOMP_RET_TRACE: + /* + * A tracer may rewrite the syscall, and there is no + * defensible way to restart composition mid-walk. + * Fail closed exactly like TRACE with no tracer + * attached: skip with -ENOSYS. + */ + syscall_set_return_value(current, current_pt_regs(), + -ENOSYS, 0); + goto skip; + + case SECCOMP_RET_KILL_THREAD: + case SECCOMP_RET_KILL_PROCESS: + default: + seccomp_kill_task(this_syscall, action, data); + return -1; + } + } + + return 0; + +skip: + seccomp_log(this_syscall, 0, action, f->log); + return -1; +} + static int __seccomp_filter(int this_syscall, const bool recheck_after_tra= ce) { u32 filter_ret, action; struct seccomp_data sd; struct seccomp_filter *match =3D NULL; + bool redirected =3D false; int data; =20 /* @@ -1356,9 +1467,19 @@ static int __seccomp_filter(int this_syscall, const = bool recheck_after_trace) return 0; =20 case SECCOMP_RET_USER_NOTIF: - if (seccomp_do_user_notification(this_syscall, match, &sd)) + if (seccomp_do_user_notification(this_syscall, match, &sd, + &redirected)) goto skip; =20 + /* + * A redirect rewrote the argument registers; every filter + * outer to the notifier must judge the substituted syscall + * before it runs. A redirect from the outermost filter has + * no outer filter left to judge it. + */ + if (redirected && match->prev) + return seccomp_redirect_revalidate(match); + return 0; =20 case SECCOMP_RET_LOG: @@ -1376,18 +1497,7 @@ static int __seccomp_filter(int this_syscall, const = bool recheck_after_trace) case SECCOMP_RET_KILL_THREAD: case SECCOMP_RET_KILL_PROCESS: default: - current->seccomp.mode =3D SECCOMP_MODE_DEAD; - seccomp_log(this_syscall, SIGSYS, action, true); - /* Dump core only if this is the last remaining thread. */ - if (action !=3D SECCOMP_RET_KILL_THREAD || - (atomic_read(¤t->signal->live) =3D=3D 1)) { - /* Show the original registers in the dump. */ - syscall_rollback(current, current_pt_regs()); - /* Trigger a coredump with SIGSYS */ - force_sig_seccomp(this_syscall, data, true); - } else { - do_exit(SIGSYS); - } + seccomp_kill_task(this_syscall, action, data); return -1; /* skip the syscall go directly to signal handling */ } =20 @@ -2223,6 +2333,7 @@ static long seccomp_notify_send_redirect(struct secco= mp_filter *filter, goto out_unlock_free; } =20 + knotif->redirect =3D true; knotif->state =3D SECCOMP_NOTIFY_REPLIED; knotif->error =3D 0; knotif->val =3D 0; --=20 2.43.0 From nobody Sat Jul 25 05:59:38 2026 Received: from mail-pg1-f175.google.com (mail-pg1-f175.google.com [209.85.215.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7FE2541A931 for ; Fri, 24 Jul 2026 22:02:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.175 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930554; cv=none; b=o+Yov7vFRkZ/l4UCXrd9LzmRoKv2O36ld1Uue7Skzox5tZi+XBoYWkLHVMQhYNHVSXpmVK5i7+UaFlnBIjolj1958cw8qHN/59orG4aJZ1u1clFy9pB/hj8L6ABGsU1fkvcBipR3XYsyFf+s+Oue0Yc8IX6JukMgbTs5lTcXHEs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930554; c=relaxed/simple; bh=5HCIi5ySWkKQMRa9Wr9spZFtfcF1MXpOzBTg6XYfGww=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=INyNckSahPGWT9ul3dn/L7fAjA52OeqIBNoEJiMfTJP8Cj1E8x1rfCaUHnkM7JglvoRhi4j/oUw5TCis+dBpKw4vTdtUVzH1XxfhWJV6whRjJi9rOiJu+EElsPkSbCDTV05mx83B31r/uISUyLMcoBHdPmJmkoxZisMebrrUTaQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=rPMGYTcA; arc=none smtp.client-ip=209.85.215.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="rPMGYTcA" Received: by mail-pg1-f175.google.com with SMTP id 41be03b00d2f7-c9d1fff21edso590602a12.1 for ; Fri, 24 Jul 2026 15:02:32 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784930552; x=1785535352; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=e1GNNEUTrXfcasiBfiMqOJUDgRwUpOC+gV6LgyALCys=; b=rPMGYTcAVhKyfnbu0fBbhxfdTbEu5GGQ6TsLw6S2glwO3Ta2GxwnyE/cCUa51sd2Nd i9rZsTc41Ycz/3fRnUG8G+ZZoBtfFFw2SKMLG+INSaThLgXx/trQnDVjsT9+kwvYZSU+ GrxuxVMf0iVNmLUN204CX6DQF1ZVThS97XbMVSl8DGANzA6s1TdGsRYHTlPeHEshzPmr rL21szI+JHZkgRwXhQFVxMhbK8qmFxAWg5uHYO0gfGudRaXs8JD5N1QQKYqrOgpvmUu5 4nVC81bzWQLVrcSE/7kCYOyORrhl6Ca2hNt7JSyg9KiYdyerR8Ozi+WvV5kATV+9lnX4 m9pw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784930552; x=1785535352; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=e1GNNEUTrXfcasiBfiMqOJUDgRwUpOC+gV6LgyALCys=; b=jMVeOKT38H4CYAQS/MbXlocdPVgtMku2pchFAh18huiP2oUPyiOqWC7e/YfH8M2xuU FBRE9S1rgMGQzBK5YFq9etwfdnQAMgGfRmBchELXXmu3+iQ0zo1ZJK8qQkcs9CSIQb+b 6tGtneUjZk3MSHrnzRDOmsbSWI1moCHcgHXLUw7u+/z+ePMhVVFxENNpTYlaTcAaZ+/3 4JNdGodhWR3SkFPzMRjP1DG2cRL8u8S4uueRZFHmEp0GfoHPOxTP6rAy1bLKIrPgfI64 DjcUMXyzcBV0EvzrqJ7q4Df5gUGy4DWIhiiRoM//jm2tpiWB38mWWT8kngTgnMwmldLj sH9A== X-Forwarded-Encrypted: i=1; AHgh+RqW7FoZX81SmLgX8LAIJcQIk3r0rDetKag0PnvUHYVWNN0MswBqrMKk5YMRoo7JMVI+trVE7O+7pnXnWg8=@vger.kernel.org X-Gm-Message-State: AOJu0YzUbJw8alD8/h8uafRrbKz6up+ISu+mqWEY3HPYBuZr2vl450Qw XlnQz4u1/b7p1ruM4SnXb8HoUSjIBX32FRak8+oMELuulyBPt1O5KPjr X-Gm-Gg: AR+sD10fmtTSYyW23C0QlTap3Iy0ygh4+tNYKKgG6TxgyN0ua7+Kq6kTYpuGUBTWTOX jlaCUbwAUsjogJVGhJWHqDheF9DVbpncuI863cgXqOAkiRDXlaLQxhJdnJ4O/Lj+P7EeQ9RH+gl M4IZe3pgYediHT83fp1k22MNoq4+VgYuiI7y5GhXXwA2AE6f2wJX9r6k2nEt/ouiJVwyQOUjWNo aCj7aDv5S+v7cvsh6DX0CGmK9Qe0VRMrxViW4wxWU7dkR/iwtn+3Ig7RayNQdHG/Iyyu7ty8HOl Vc30a/BfxiyJKtBRelOyzq9Pwxjl0VCxywTFaxITD4nnBTgxEFzKVQDt1SxHB67NcIC9/7qIExM 5bYjJH6SFeWYGxOO1SvhFwE3tdWzDr/p1nFyqlxIklbsebJ4CYUUm3NiBrKczX8lQm6Q0G/ozfy OI2X4jHexOAG8kHIHuxhB4AMBndnwGIOQ93HDbX5o2un/6P2KCbPqCTGg6EzwA X-Received: by 2002:a05:6a21:6e96:b0:3c0:9c1a:8941 with SMTP id adf61e73a8af0-3c67e17d26dmr101841637.73.1784930551744; Fri, 24 Jul 2026 15:02:31 -0700 (PDT) Received: from pop-os.tail4adac7.ts.net ([2601:647:6802:dbc0:546a:e1b0:9574:3a29]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc5a67f3sm2962631eec.29.2026.07.24.15.02.29 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 15:02:30 -0700 (PDT) From: Cong Wang To: Andy Lutomirski Cc: Kees Cook , linux-kernel@vger.kernel.org, Will Drewry , Christian Brauner , Andrew Morton , linux-mm@kvack.org, Cong Wang Subject: [PATCH v7 7/8] docs/seccomp: document pinned-memfd redirect ioctls Date: Fri, 24 Jul 2026 15:01:46 -0700 Message-ID: <20260724220147.214396-8-xiyou.wangcong@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260724220147.214396-1-xiyou.wangcong@gmail.com> References: <20260724220147.214396-1-xiyou.wangcong@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Cong Wang Document SECCOMP_IOCTL_NOTIF_PIN_INSTALL and SECCOMP_IOCTL_NOTIF_SEND_REDIRECT in the userspace API guide: the SECCOMP_FILTER_FLAG_REDIRECT opt-in and the single-redirector restriction, the two response structures, and how the pair closes the user-notification TOCTOU for non-cooperative fork+execve sandboxes. Also spell out the scope the implementation deliberately enforces or relies on: read-only input pointers only, same-syscall-number only (rt_sigreturn and the clone/fork family are refused), the per-interruption re-notification of restartable syscalls and the restart-block behaviour, and the ptrace syscall-stop semantics. Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Cong Wang --- .../userspace-api/seccomp_filter.rst | 109 ++++++++++++++++++ 1 file changed, 109 insertions(+) diff --git a/Documentation/userspace-api/seccomp_filter.rst b/Documentation= /userspace-api/seccomp_filter.rst index cff0fa7f3175..be2c338224ac 100644 --- a/Documentation/userspace-api/seccomp_filter.rst +++ b/Documentation/userspace-api/seccomp_filter.rst @@ -289,6 +289,115 @@ above in this document: all arguments being read from= the tracee's memory should be read into the tracer's memory before any policy decisions are ma= de. This allows for an atomic decision on syscall arguments. =20 +Non-cooperative pinned-memfd redirect +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The TOCTOU described above means ``SECCOMP_USER_NOTIF_FLAG_CONTINUE`` cann= ot +enforce a policy on pointer arguments: after the supervisor inspects the +target's memory and lets the syscall continue, the target (or a thread sha= ring +its address space) can rewrite that memory before the kernel reads it. The +cooperative workaround, the target ``mmap()`` + ``mseal()``-ing a shared +buffer, is unavailable in the fork+execve sandbox model, where the supervi= sor +confines a binary it did not write. + +Two ioctls let the supervisor close this race without target cooperation. = The +redirect step (below) requires a listener created with +``SECCOMP_FILTER_FLAG_REDIRECT`` (in addition to +``SECCOMP_FILTER_FLAG_NEW_LISTENER``). Because it rewrites another task's +registers, at most one such listener may exist in a task's filter chain; a +second fails with ``-EBUSY``: + +.. code-block:: c + + fd =3D seccomp(SECCOMP_SET_MODE_FILTER, + SECCOMP_FILTER_FLAG_NEW_LISTENER | SECCOMP_FILTER_FLAG_RE= DIRECT, + &prog); + +``ioctl(SECCOMP_IOCTL_NOTIF_PIN_INSTALL)`` installs a sealed mapping of a +supervisor-owned ``memfd`` directly into the trapped task's address space: + +.. code-block:: c + + struct seccomp_notif_pin_install { + __u64 id; + __u32 flags; /* reserved, must be 0 */ + __u32 memfd; + __u64 target_addr; + __u64 size; + __u64 offset; /* page-aligned offset into memfd */ + }; + +``id`` names an active notification (the trapped task to install into). +``target_addr``, ``size`` and ``offset`` are page-aligned; ``offset`` sele= cts +where in ``memfd`` the mapping starts, so one memfd can back several pins.= If +``target_addr`` is ``0`` the kernel picks a free address and writes it bac= k; +otherwise an existing mapping there yields ``-EEXIST``. The pin is read-on= ly +and sealed, the target and its threads cannot unmap, move, reprotect or +overwrite it, and lasts until the target calls ``execve()`` or exits. + +``memfd`` must be write-sealed (``F_SEAL_WRITE`` or ``F_SEAL_FUTURE_WRITE`= `) +or the ioctl returns ``-EINVAL``; otherwise the target could rewrite the p= in's +bytes through a separate writable handle to the same memfd. +``F_SEAL_FUTURE_WRITE`` still lets the supervisor update the contents thro= ugh +its own mapping made before the seal. + +``ioctl(SECCOMP_IOCTL_NOTIF_SEND_REDIRECT)`` then resumes the trapped sysc= all +like ``SECCOMP_USER_NOTIF_FLAG_CONTINUE``, but with selected argument +registers replaced: + +.. code-block:: c + + struct seccomp_notif_resp_redirect { + __u64 id; + __u32 flags; /* SECCOMP_REDIRECT_FLAG_CONTINUE must be set */ + __u32 args_mask; /* which arg registers to replace */ + __u32 ptr_mask; /* which of those are pointers into a pin */ + __u32 memfd; /* the pin's backing memfd */ + __u64 args[6]; /* replacement values */ + __u64 ptr_len[6]; /* validated access length for each pointer arg = */ + }; + +Each bit in ``ptr_mask`` (a subset of ``args_mask``) marks ``args[i]`` as a +pointer; the access ``[args[i], args[i] + ptr_len[i])`` must lie within a +single read-only pin of ``memfd`` in the target, or the ioctl returns +``-EFAULT``. ``ptr_len[i]`` must be non-zero for those bits and ``0`` +otherwise. Bits in ``args_mask`` but not ``ptr_mask`` are scalar replaceme= nts +written verbatim, e.g. to set the length register that goes with a redirec= ted +pointer. The original registers are restored at syscall exit, so the +substitution is invisible to the target and the TOCTOU is closed. + +Scope and limitations +--------------------- + +The redirect mechanism is deliberately narrow and is *not* a general sysca= ll +rewriting facility: + +- **Read-only input pointers only.** A pin is read-only, so only an argume= nt + the syscall *reads* (a pathname, a ``sockaddr``) may be redirected into = it. + Aiming an output or in/out argument at a pin makes the syscall fail with + ``-EFAULT`` when it writes back. + +- **Same syscall only.** A redirect replaces arguments, never the syscall + number. ``rt_sigreturn()`` (and its compat variant) cannot be redirected= and + return ``-EOPNOTSUPP``. + +- **Signals and restarts.** The redirected syscall really runs, so it can = be + interrupted and restarted. On a restart the original arguments are resto= red + and the syscall re-traps, so the supervisor is notified again and must a= nswer + consistently. Syscalls the kernel restarts without re-trapping (e.g. + ``nanosleep()``, ``futex(FUTEX_WAIT)``) keep the substituted arguments -- + safe for read-only inputs, but a reason not to redirect arguments of sys= calls + that block or wait. + +- **clone()/fork().** The task-creation family (``clone``, ``clone3``, + ``fork``, ``vfork``) cannot be redirected and returns ``-EOPNOTSUPP``. A= new + child inherits the substituted argument registers but not the restore, s= o it + would return to user space with the pinned values still in its registers. + +- **ptrace.** A tracer sees the substituted arguments at the syscall-exit = stop; + they are restored before the task resumes, so a ``PTRACE_SETREGS`` of a + substituted register at that stop is overwritten. + Sysctls =3D=3D=3D=3D=3D=3D=3D =20 --=20 2.43.0 From nobody Sat Jul 25 05:59:38 2026 Received: from mail-pg1-f177.google.com (mail-pg1-f177.google.com [209.85.215.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 222794137B2 for ; Fri, 24 Jul 2026 22:02:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.177 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930560; cv=none; b=PBYmdGl9qcwLFvMVDrFf2rb8sS6iI0M046eRmE23oovgtDBtxwR7OG5xDJJcfI17qgneGbhwVDDyY/9yTx2GqfJxRrjUhnIU7TI9Xl5vdVkB+JaaVUSgt7OMtOdpFoA8VCg4urzKqUAgeLz3AM6KaIMVo4020rYA4be1+UcHBMk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784930560; c=relaxed/simple; bh=B1Ug9Pjo34gbmCbzCERBbGlgJdYErqsiuLKTI9GQL70=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=kapRXqSGspGvZUNf+DR4HWpDdxWAC/L3odwKfLONROfris279tLTBLdd4NWPBPfjDmEXJkxV1sM2faNXyAc8oFOJz4etxvBWOwLengSPnsGdokL2e98R/Ebeq1T9uWxdftRFTkK79+lxt3gqUPTJPv28RiKtffBEJTF5qK9L3ZA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=L9ZcAnfV; arc=none smtp.client-ip=209.85.215.177 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="L9ZcAnfV" Received: by mail-pg1-f177.google.com with SMTP id 41be03b00d2f7-ca7bea5e5b3so707564a12.1 for ; Fri, 24 Jul 2026 15:02:35 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784930554; x=1785535354; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=FKyxT8v03tfemxpC4Cf8boPeiHU+d+LC0S9/PzgfnXI=; b=L9ZcAnfVZxAtgCjTH7iwsnSu8xB5VrdY3UEM6mcirm25T8HfBl8pYF8yLzFQgeOtKp aFlY6i8A02+b71INM1Fm3MOmHbq9vFpZyQBlCXbcE6+maTTSLjN7HAVtPpJoWMsPEkWF b3GafoHsXBvJOvOo8u6w6GXMAIbvTO9JIocCWNsvFpXK42lpLmul8J/lUKr1JU/ll5RD 2aKRtddUXBSi4swTll85I7Q83LDGm+iQV/JoSN2iec33xjD8M0IZzbazujH0aRnUQ28u 0qutYx8uJ64t42CYUKwnpnM22cCRJp921DRvqmExOSxnK16fhQc27oneYXyKkFcXM8sC SKHw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784930554; x=1785535354; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=FKyxT8v03tfemxpC4Cf8boPeiHU+d+LC0S9/PzgfnXI=; b=mj69n8IamTYVzHeGrBuosjJNWLY654NAjP3A3KPncNq4BNq4D5jS+WPG5kfSt7I9x3 LOJYQKDp7YAUpUfD70mfGKWKYXJBzSZP6QQWS4kWGM5dTtGF1uyu9+sdqh2OP6lf/vyi Pfuos/Y3fCOeoShiU18spqX5akYDMAWnkUAhsxOxYE2+WuwpvgfHGLqZA3aCiDzDJf1f cfh89XOMLPDkE/JrJoBATvNVm6QKpgk54xtFq2FRB/bmS0S3RWeo7QMv8Aqr4mrGHR73 ApYVoTJAtVLMX30AzThkFrkqVtRwF5KX9QK0+OZ9NluhsMq8lIJYJRu3CJZUc6phOvWw udGQ== X-Forwarded-Encrypted: i=1; AHgh+Rqmvv65rUdanJK9MPMeCjhXCOIb12fLpAj1OeR43sI77wQitEBku2oW5pM9IaWSs6valAHTG5aa8E0QqpA=@vger.kernel.org X-Gm-Message-State: AOJu0Yyb8PoN6BRkkvmlGT7Jb8rWt50AAjnsGFyE6mIcig/1LcBDPPSq WdwFMT4PKbGFaFFPMyyAjUVpGS4s1gIvpMyuckk3ZRys+icMmWqi1+8z X-Gm-Gg: AR+sD12TzFGdcZWbFk6Pljt8R6AwnKhJ6aRiRzsRLzUl6squI7WIfjCQ/dZ8lUUPWm9 S9nohbWGLwlUHk1yWq/h9iI6z4oKhroJm4pLIshTbJRVYibRVIpIIgE4CbwoSIQh4pip4K8zLKs L2H6Qe1P7pL7FpIpsmYV+X+4UCEJnDk3Xsu/BFk4xlMORZbW3Bxg8Ab0StVs/F+XlaKGXkYQJDM 4LpKpFwoQVHDy3SJOMgGQBs2XdP4ldQjMZT5pWNtonSwMvGb11TzZVKsj5xeWY5uhvNz/wjHeHu kf/7cDmg0VqE5VxecR4T0u2VAjWGybJyCo4VZMFcATMmJ/LvMoSzZBEudQCLj5Aa7Yg1tWqGQE+ asMe6tSRVXnFnRYo9XozaMO6tuwSBwzsRZ4e/sErxFcQA4NTKBuEOtZcm8sx166/Fxp1SdFxk0R ghNadKColaIbr8dcGcqvzfMsDjN4YVYhEoCQYOOHB3UTFuuj2CX2LEZzKzSWsr X-Received: by 2002:a05:6a20:4323:b0:3c0:9c19:65b1 with SMTP id adf61e73a8af0-3c67e1780bemr109759637.73.1784930553882; Fri, 24 Jul 2026 15:02:33 -0700 (PDT) Received: from pop-os.tail4adac7.ts.net ([2601:647:6802:dbc0:546a:e1b0:9574:3a29]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-314bc5a67f3sm2962631eec.29.2026.07.24.15.02.31 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 15:02:32 -0700 (PDT) From: Cong Wang To: Andy Lutomirski Cc: Kees Cook , linux-kernel@vger.kernel.org, Will Drewry , Christian Brauner , Andrew Morton , linux-mm@kvack.org, Cong Wang Subject: [PATCH v7 8/8] selftests/seccomp: cover non-cooperative pinned-memfd install Date: Fri, 24 Jul 2026 15:01:47 -0700 Message-ID: <20260724220147.214396-9-xiyou.wangcong@gmail.com> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260724220147.214396-1-xiyou.wangcong@gmail.com> References: <20260724220147.214396-1-xiyou.wangcong@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Cong Wang Add 11 tests for SECCOMP_IOCTL_NOTIF_PIN_INSTALL and SECCOMP_IOCTL_NOTIF_SEND_REDIRECT: - pinned_memfd_remote: basic install + redirect, plus the unsealed memfd and out-of-pin rejection paths - pinned_memfd_target_cannot_unmap: the sealed pin survives the target's munmap/mprotect/mremap/MAP_FIXED attacks (all EPERM) - pinned_memfd_execve_scm: SCM_RIGHTS supervisor handoff and re-pin in the fresh post-execve mm - pinned_memfd_churn: one listener serves many short-lived targets, no per-target state - redirect_outer_refilter, redirect_revalidate_chain: a redirect is re-validated by every outer filter, not just the nearest - redirect_outer_trace, redirect_outer_notify: outer TRACE and USER_NOTIF verdicts over a redirected call fail closed with ENOSYS; the tracer must not get a PTRACE_EVENT_SECCOMP for the substituted call, and a listener-less outer USER_NOTIF must still block it - redirect_denied_syscalls: EOPNOTSUPP for rt_sigreturn and the clone/fork family - pinned_memfd_abi, redirect_signal_abi: the original arg register is restored before returning to user mode, and before any signal frame is built Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Cong Wang --- tools/testing/selftests/seccomp/seccomp_bpf.c | 1569 +++++++++++++++++ 1 file changed, 1569 insertions(+) diff --git a/tools/testing/selftests/seccomp/seccomp_bpf.c b/tools/testing/= selftests/seccomp/seccomp_bpf.c index 358b6c65e120..6122dc58f897 100644 --- a/tools/testing/selftests/seccomp/seccomp_bpf.c +++ b/tools/testing/selftests/seccomp/seccomp_bpf.c @@ -217,6 +217,10 @@ struct seccomp_metadata { #define SECCOMP_FILTER_FLAG_NEW_LISTENER (1UL << 3) #endif =20 +#ifndef SECCOMP_FILTER_FLAG_REDIRECT +#define SECCOMP_FILTER_FLAG_REDIRECT (1UL << 6) +#endif + #ifndef SECCOMP_RET_USER_NOTIF #define SECCOMP_RET_USER_NOTIF 0x7fc00000U =20 @@ -295,6 +299,35 @@ struct seccomp_notif_addfd_big { #define PTRACE_EVENTMSG_SYSCALL_EXIT 2 #endif =20 +#ifndef SECCOMP_IOCTL_NOTIF_PIN_INSTALL +struct seccomp_notif_pin_install { + __u64 id; + __u32 flags; + __u32 memfd; + __u64 target_addr; + __u64 size; + __u64 offset; +}; +#define SECCOMP_IOCTL_NOTIF_PIN_INSTALL SECCOMP_IOWR(5, \ + struct seccomp_notif_pin_install) +#endif + +#ifndef SECCOMP_IOCTL_NOTIF_SEND_REDIRECT +#define SECCOMP_REDIRECT_FLAG_CONTINUE (1UL << 0) +#define SECCOMP_REDIRECT_ARGS 6 +struct seccomp_notif_resp_redirect { + __u64 id; + __u32 flags; + __u32 args_mask; + __u32 ptr_mask; + __u32 memfd; + __u64 args[SECCOMP_REDIRECT_ARGS]; + __u64 ptr_len[SECCOMP_REDIRECT_ARGS]; +}; +#define SECCOMP_IOCTL_NOTIF_SEND_REDIRECT SECCOMP_IOW(6, \ + struct seccomp_notif_resp_redirect) +#endif + #ifndef SECCOMP_USER_NOTIF_FLAG_CONTINUE #define SECCOMP_USER_NOTIF_FLAG_CONTINUE 0x00000001 #endif @@ -4368,6 +4401,1542 @@ TEST(user_notification_addfd_rlimit) close(memfd); } =20 +/* + * Create a write-sealed memfd of @size for PIN_INSTALL and map a supervis= or + * writable view, primed with @content. F_SEAL_FUTURE_WRITE keeps this + * pre-seal mapping writable (so the test can still stage content) while + * barring any other writable reference, as PIN_INSTALL requires. Returns + * the memfd. + */ +static int make_pin_memfd(struct __test_metadata *_metadata, const char *n= ame, + size_t size, char **sup_view, const char *content) +{ + int memfd =3D memfd_create(name, MFD_ALLOW_SEALING); + + ASSERT_GE(memfd, 0); + ASSERT_EQ(0, ftruncate(memfd, size)); + ASSERT_EQ(0, fcntl(memfd, F_ADD_SEALS, F_SEAL_SHRINK | F_SEAL_GROW)); + + *sup_view =3D mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_SHARED, + memfd, 0); + ASSERT_NE(MAP_FAILED, *sup_view); + ASSERT_EQ(0, fcntl(memfd, F_ADD_SEALS, F_SEAL_FUTURE_WRITE)); + memcpy(*sup_view, content, strlen(content) + 1); + return memfd; +} + +/* + * Non-cooperative pinned-memfd: the target only traps on openat(); the + * supervisor PIN_INSTALLs a sealed mapping of its memfd into the target m= m and + * SEND_REDIRECTs args[1] into it, so the openat reads the supervisor's pa= th. + * Also covers the unsealed-memfd and out-of-pin rejections. + */ +TEST(user_notification_pinned_memfd_remote) +{ + pid_t pid; + long ret; + int status, listener, memfd, unsealed; + struct seccomp_notif req =3D {}; + struct seccomp_notif_pin_install pin =3D {}; + struct seccomp_notif_pin_install unsealed_pin =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + char *sup_view; + const size_t PIN_SIZE =3D 4096; + const char *safe_path =3D "/dev/null"; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + memfd =3D make_pin_memfd(_metadata, "pinned-remote", PIN_SIZE, + &sup_view, safe_path); + + listener =3D user_notif_syscall(__NR_openat, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(goto cleanup, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + int fd; + + /* + * Target performs no setup. Just trap on openat. Kernel + * (driven by the supervisor) will install the pin in this + * process's mm at a kernel-chosen address behind our back, + * and our openat will be redirected to read from there. + */ + fd =3D syscall(__NR_openat, AT_FDCWD, + "/this/should/never/be/touched", O_RDONLY, 0); + if (fd < 0) + _exit(11); + _exit(0); + } + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_openat); + + pin.id =3D req.id; + pin.memfd =3D memfd; + pin.target_addr =3D 0; + pin.size =3D PIN_SIZE; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_PIN_INSTALL, &pin)) { + if (errno =3D=3D EINVAL) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + SKIP(goto cleanup, + "Kernel does not support pinned-memfd remote install"); + } + TH_LOG("PIN_INSTALL failed: errno=3D%d", errno); + } + + /* The kernel wrote a non-zero, page-aligned address back to us. */ + EXPECT_NE(0, pin.target_addr); + EXPECT_EQ(0, pin.target_addr & (PIN_SIZE - 1)); + + /* Reject: the backing memfd must be write-sealed. */ + unsealed =3D memfd_create("unsealed", MFD_ALLOW_SEALING); + ASSERT_GE(unsealed, 0); + ASSERT_EQ(0, ftruncate(unsealed, PIN_SIZE)); + unsealed_pin.id =3D req.id; + unsealed_pin.memfd =3D unsealed; + unsealed_pin.size =3D PIN_SIZE; + EXPECT_EQ(-1, ioctl(listener, SECCOMP_IOCTL_NOTIF_PIN_INSTALL, + &unsealed_pin)); + EXPECT_EQ(EINVAL, errno); + close(unsealed); + + /* Reject: redirect outside any installed pin. */ + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 1; + redir.ptr_mask =3D 1U << 1; + redir.memfd =3D memfd; + redir.ptr_len[1] =3D strlen(safe_path) + 1; + redir.args[1] =3D pin.target_addr + PIN_SIZE; /* one byte past */ + EXPECT_EQ(-1, ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, + &redir)); + EXPECT_EQ(EFAULT, errno); + + /* Reject: base is inside the pin but the extent runs past its end. */ + redir.args[1] =3D pin.target_addr; + redir.ptr_len[1] =3D PIN_SIZE + 1; + EXPECT_EQ(-1, ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, + &redir)); + EXPECT_EQ(EFAULT, errno); + + /* Happy path: redirect into the kernel-installed pin. */ + redir.args[1] =3D pin.target_addr; + redir.ptr_len[1] =3D strlen(safe_path) + 1; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, + &redir)) { + /* Unblock the trapped child so waitpid() cannot hang. */ + kill(pid, SIGKILL); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + TH_LOG("child exit %d (11=3Dopenat fail)", WEXITSTATUS(status)); + } + +cleanup: + munmap(sup_view, PIN_SIZE); + close(memfd); + close(listener); +} + +/* + * The pin is VM_SEALED: a target that learns its address can read it but = must + * not unmap, move, reprotect or MAP_FIXED-stomp it. That is what keeps a + * redirected pointer aimed at supervisor-controlled bytes for the whole c= all. + */ +TEST(user_notification_pinned_memfd_target_cannot_unmap) +{ + pid_t pid; + long ret; + int status, listener, memfd; + struct seccomp_notif req =3D {}; + struct seccomp_notif_pin_install pin =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + char *sup_view; + int addrpipe[2]; + const size_t PIN_SIZE =3D 4096; + const char *safe_path =3D "/dev/null"; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + ASSERT_EQ(0, pipe(addrpipe)); + + /* + * The parent writes the pin address to the child through addrpipe. + * If the child has already exited (e.g. a failed redirect), that + * write must fail cleanly rather than kill the parent with SIGPIPE. + */ + signal(SIGPIPE, SIG_IGN); + + memfd =3D make_pin_memfd(_metadata, "pin-nounmap", PIN_SIZE, + &sup_view, safe_path); + + listener =3D user_notif_syscall(__NR_openat, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(goto cleanup, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + unsigned long pin_addr; + char *p; + int fd; + + close(addrpipe[1]); + + /* Trap; the supervisor installs+seals the pin and redirects. */ + fd =3D syscall(__NR_openat, AT_FDCWD, + "/this/should/never/be/touched", O_RDONLY, 0); + if (fd < 0) + _exit(11); + + /* Learn where the kernel installed the pin. */ + if (read(addrpipe[0], &pin_addr, sizeof(pin_addr)) !=3D + sizeof(pin_addr)) + _exit(12); + p =3D (char *)pin_addr; + + /* Mapped and readable: it holds the redirected path. */ + if (*p !=3D '/') + _exit(13); + + /* Sealed: unmap must fail and tear nothing down. */ + if (munmap((void *)pin_addr, PIN_SIZE) =3D=3D 0) + _exit(20); + if (errno !=3D EPERM) + _exit(21); + + /* Sealed: cannot reprotect it (even to PROT_NONE). */ + if (mprotect((void *)pin_addr, PIN_SIZE, PROT_NONE) =3D=3D 0) + _exit(22); + if (errno !=3D EPERM) + _exit(23); + + /* Sealed: cannot move or resize it. */ + if (mremap((void *)pin_addr, PIN_SIZE, PIN_SIZE * 2, + MREMAP_MAYMOVE) !=3D MAP_FAILED) + _exit(24); + if (errno !=3D EPERM) + _exit(25); + + /* Sealed: cannot be replaced by a MAP_FIXED mapping. */ + if (mmap((void *)pin_addr, PIN_SIZE, PROT_READ, + MAP_PRIVATE | MAP_ANONYMOUS | MAP_FIXED, -1, 0) !=3D + MAP_FAILED) + _exit(26); + if (errno !=3D EPERM) + _exit(27); + + /* Survived every attack, still mapped and readable. */ + if (*p !=3D '/') + _exit(28); + _exit(0); + } + + close(addrpipe[0]); + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_openat); + + pin.id =3D req.id; + pin.memfd =3D memfd; + pin.target_addr =3D 0; + pin.size =3D PIN_SIZE; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_PIN_INSTALL, &pin)) { + if (errno =3D=3D EINVAL) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + SKIP(goto cleanup, + "Kernel does not support pinned-memfd remote install"); + } + TH_LOG("PIN_INSTALL failed: errno=3D%d", errno); + } + + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 1; + redir.ptr_mask =3D 1U << 1; + redir.memfd =3D memfd; + redir.ptr_len[1] =3D strlen(safe_path) + 1; + redir.args[1] =3D pin.target_addr; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, &redir)) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + goto cleanup; + } + + /* Hand the target the pin address so it can try to attack it. */ + EXPECT_EQ(sizeof(pin.target_addr), + write(addrpipe[1], &pin.target_addr, sizeof(pin.target_addr))); + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + TH_LOG("child exit %d (>=3D20 =3D seal breach)", + WEXITSTATUS(status)); + } + +cleanup: + close(addrpipe[1]); + munmap(sup_view, PIN_SIZE); + close(memfd); + close(listener); +} + +/* + * Helper for the execve test: read up to @max bytes of a NUL-terminated + * string from @pid's mm at @addr into @out. Returns the length read + * (excluding the NUL), or -1 on failure or no NUL. + */ +static ssize_t read_remote_string(pid_t pid, unsigned long addr, + char *out, size_t max) +{ + struct iovec local =3D { .iov_base =3D out, .iov_len =3D max }; + struct iovec remote =3D { .iov_base =3D (void *)addr, .iov_len =3D max }; + ssize_t n; + size_t i; + + n =3D process_vm_readv(pid, &local, 1, &remote, 1, 0); + if (n <=3D 0) + return -1; + for (i =3D 0; i < (size_t)n; i++) + if (out[i] =3D=3D '\0') + return (ssize_t)i; + return -1; +} + +/* + * Send a file descriptor over a connected UNIX socket via SCM_RIGHTS. + * Used by the execve_scm test so the target child can hand its + * SECCOMP_FILTER_FLAG_NEW_LISTENER fd to the supervising parent + * without the parent having to inherit the seccomp filter itself. + */ +static int send_fd(int sock, int fd) +{ + char cbuf[CMSG_SPACE(sizeof(int))] =3D {}; + char data =3D 'x'; + struct iovec iov =3D { .iov_base =3D &data, .iov_len =3D 1 }; + struct msghdr msg =3D { + .msg_iov =3D &iov, .msg_iovlen =3D 1, + .msg_control =3D cbuf, .msg_controllen =3D sizeof(cbuf), + }; + struct cmsghdr *cmsg =3D CMSG_FIRSTHDR(&msg); + + cmsg->cmsg_level =3D SOL_SOCKET; + cmsg->cmsg_type =3D SCM_RIGHTS; + cmsg->cmsg_len =3D CMSG_LEN(sizeof(int)); + memcpy(CMSG_DATA(cmsg), &fd, sizeof(int)); + return sendmsg(sock, &msg, 0) < 0 ? -1 : 0; +} + +static int recv_fd(int sock) +{ + char cbuf[CMSG_SPACE(sizeof(int))] =3D {}; + char data; + struct iovec iov =3D { .iov_base =3D &data, .iov_len =3D 1 }; + struct msghdr msg =3D { + .msg_iov =3D &iov, .msg_iovlen =3D 1, + .msg_control =3D cbuf, .msg_controllen =3D sizeof(cbuf), + }; + struct cmsghdr *cmsg; + int fd; + + if (recvmsg(sock, &msg, 0) < 0) + return -1; + cmsg =3D CMSG_FIRSTHDR(&msg); + if (!cmsg || cmsg->cmsg_level !=3D SOL_SOCKET || + cmsg->cmsg_type !=3D SCM_RIGHTS || + cmsg->cmsg_len !=3D CMSG_LEN(sizeof(int))) + return -1; + memcpy(&fd, CMSG_DATA(cmsg), sizeof(int)); + return fd; +} + +struct addr_range { + unsigned long start, end; +}; + +/* + * Parse /proc//maps looking for the dynamic linker's executable + * mapping (glibc ld-linux-*.so, musl ld-musl-*.so, etc.). The trapped + * task's instruction_pointer falling in this range identifies a + * loader-bootstrap syscall (race-free, kernel-truth) so the supervisor + * can auto-allow it without inspecting argument content via the racy + * process_vm_readv path. + * + * Requires the supervisor not to be subject to the seccomp filter + * itself -- fopen() internally calls openat(). The execve_scm test + * structure (child installs filter, sends listener fd to parent via + * SCM_RIGHTS) satisfies that. + * + * Returns 0 on success with @out populated, -1 if not found. + */ +static int find_loader_text_range(pid_t pid, struct addr_range *out) +{ + char maps_path[64]; + char line[512]; + FILE *f; + int found =3D 0; + + snprintf(maps_path, sizeof(maps_path), "/proc/%d/maps", pid); + f =3D fopen(maps_path, "r"); + if (!f) + return -1; + + while (fgets(line, sizeof(line), f)) { + unsigned long start, end; + char perms[8]; + char *path; + + if (sscanf(line, "%lx-%lx %7s", &start, &end, perms) !=3D 3) + continue; + if (!strchr(perms, 'x')) + continue; + path =3D strchr(line, '/'); + if (!path) + continue; + /* + * Match common dynamic-linker basenames: ld-linux-*.so + * (glibc), ld-musl-*.so (musl), ld-*.so (older glibc). + */ + if (strstr(path, "/ld-") || strstr(path, "/ld.so")) { + out->start =3D start; + out->end =3D end; + found =3D 1; + break; + } + } + fclose(f); + return found ? 0 : -1; +} + +/* + * Pinned-memfd across a real execve. The child installs the filter on its= elf + * and hands the listener to the parent over SCM_RIGHTS, so the parent (not + * filtered) can read /proc//maps for race-free loader detection. The + * supervisor PIN_INSTALLs+SEND_REDIRECTs before execve, then again in the + * fresh post-execve mm (the old pin VMA died with the old mm), proving the + * redirect survives an mm replacement, not just the install side. + */ +TEST(user_notification_pinned_memfd_execve_scm) +{ + pid_t pid; + int status, listener, memfd, sv[2]; + struct seccomp_notif req =3D {}; + struct seccomp_notif_pin_install pin =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + struct seccomp_notif_resp cont_resp =3D {}; + char *sup_view; + const size_t PIN_SIZE =3D 4096; + const char *safe_path =3D "/dev/null"; + const char *bait =3D "/seccomp_pinned_memfd_test_bait_scm"; + bool post_exec_install_ok =3D false; + bool post_exec_redirect_done =3D false; + bool loader_known =3D false; + bool loader_check_attempted =3D false; + struct addr_range loader_range =3D {}; + int phase =3D 0; + int trap_count =3D 0; + const int trap_limit =3D 200; + + if (access("/bin/cat", X_OK) !=3D 0) + SKIP(return, "/bin/cat not present"); + + memfd =3D make_pin_memfd(_metadata, "pin-execve-scm", PIN_SIZE, + &sup_view, safe_path); + + ASSERT_EQ(0, socketpair(AF_UNIX, SOCK_SEQPACKET, 0, sv)); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + struct sock_filter filter[] =3D { + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, + offsetof(struct seccomp_data, nr)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, __NR_openat, + 0, 1), + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_USER_NOTIF), + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_ALLOW), + }; + struct sock_fprog prog =3D { + .len =3D (unsigned short)ARRAY_SIZE(filter), + .filter =3D filter, + }; + int my_listener; + int fd; + + close(sv[0]); + if (prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0)) + _exit(20); + my_listener =3D seccomp(SECCOMP_SET_MODE_FILTER, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT, + &prog); + if (my_listener < 0) + _exit(21); + if (send_fd(sv[1], my_listener) < 0) + _exit(22); + close(my_listener); + close(sv[1]); + + /* Pre-execve trap. */ + fd =3D syscall(__NR_openat, AT_FDCWD, + "/this/should/never/be/touched", O_RDONLY, 0); + if (fd < 0) + _exit(11); + + execl("/bin/cat", "cat", bait, (char *)NULL); + _exit(12); + } + + close(sv[1]); + listener =3D recv_fd(sv[0]); + close(sv[0]); + if (listener < 0) { + /* Child exits 21 when the kernel lacks the REDIRECT flag. */ + if (waitpid(pid, &status, 0) =3D=3D pid && WIFEXITED(status) && + WEXITSTATUS(status) =3D=3D 21) + SKIP(goto cleanup_scm, + "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + } + ASSERT_GE(listener, 0); + + /* + * Parent has the listener fd and does NOT have the seccomp + * filter. fopen(/proc//maps) below works without + * deadlocking on the parent's own openat. + */ + for (;;) { + struct pollfd pfd =3D { .fd =3D listener, .events =3D POLLIN }; + int pret =3D poll(&pfd, 1, 500); + pid_t reaped; + bool ip_in_loader; + + if (pret < 0) + break; + if (pret =3D=3D 0 || !(pfd.revents & POLLIN)) { + reaped =3D waitpid(pid, &status, WNOHANG); + if (reaped =3D=3D pid) + break; + if (pfd.revents & (POLLHUP | POLLERR)) + break; + continue; + } + + memset(&req, 0, sizeof(req)); + if (ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req) < 0) { + TH_LOG("NOTIF_RECV failed: errno=3D%d", errno); + break; + } + if (++trap_count > trap_limit) { + TH_LOG("trap_limit (%d) exceeded", trap_limit); + break; + } + + if (phase =3D=3D 0) { + pin.id =3D req.id; + pin.memfd =3D memfd; + pin.target_addr =3D 0; + pin.size =3D PIN_SIZE; + if (ioctl(listener, + SECCOMP_IOCTL_NOTIF_PIN_INSTALL, + &pin) !=3D 0) { + TH_LOG("pre-exec PIN_INSTALL failed: errno=3D%d", + errno); + if (errno =3D=3D EINVAL) + SKIP(goto cleanup_scm, + "Kernel lacks pinned-memfd remote"); + goto cleanup_scm; + } + + memset(&redir, 0, sizeof(redir)); + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 1; + redir.ptr_mask =3D 1U << 1; + redir.memfd =3D memfd; + redir.ptr_len[1] =3D strlen(safe_path) + 1; + redir.args[1] =3D pin.target_addr; + if (ioctl(listener, + SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, + &redir) !=3D 0) { + TH_LOG("pre-exec SEND_REDIRECT failed: errno=3D%d", + errno); + goto cleanup_scm; + } + phase =3D 1; + continue; + } + + /* + * Post-execve. Lazily resolve the loader range. The + * supervisor's own openat (fopen on /proc//maps) + * doesn't trap because the filter lives on the child, + * not on us. + */ + if (!loader_known && !loader_check_attempted) { + if (find_loader_text_range(req.pid, + &loader_range) =3D=3D 0) + loader_known =3D true; + loader_check_attempted =3D true; + } + + ip_in_loader =3D loader_known && + req.data.instruction_pointer >=3D loader_range.start && + req.data.instruction_pointer < loader_range.end; + + if (ip_in_loader) { + memset(&cont_resp, 0, sizeof(cont_resp)); + cont_resp.id =3D req.id; + cont_resp.flags =3D SECCOMP_USER_NOTIF_FLAG_CONTINUE; + ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND, &cont_resp); + continue; + } + + /* Program code: inspect the path to identify the bait. */ + { + char path[PATH_MAX]; + ssize_t n; + + n =3D read_remote_string(req.pid, req.data.args[1], + path, sizeof(path)); + if (n < 0 || strcmp(path, bait) !=3D 0) { + memset(&cont_resp, 0, sizeof(cont_resp)); + cont_resp.id =3D req.id; + cont_resp.flags =3D + SECCOMP_USER_NOTIF_FLAG_CONTINUE; + ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND, + &cont_resp); + continue; + } + + pin.id =3D req.id; + pin.memfd =3D memfd; + pin.target_addr =3D 0; + pin.size =3D PIN_SIZE; + if (ioctl(listener, + SECCOMP_IOCTL_NOTIF_PIN_INSTALL, + &pin) =3D=3D 0) { + post_exec_install_ok =3D true; + } else { + TH_LOG("post-exec PIN_INSTALL failed: errno=3D%d", + errno); + memset(&cont_resp, 0, sizeof(cont_resp)); + cont_resp.id =3D req.id; + cont_resp.flags =3D + SECCOMP_USER_NOTIF_FLAG_CONTINUE; + ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND, + &cont_resp); + continue; + } + + memset(&redir, 0, sizeof(redir)); + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 1; + redir.ptr_mask =3D 1U << 1; + redir.memfd =3D memfd; + redir.ptr_len[1] =3D strlen(safe_path) + 1; + redir.args[1] =3D pin.target_addr; + if (ioctl(listener, + SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, + &redir) =3D=3D 0) { + post_exec_redirect_done =3D true; + } else { + TH_LOG("post-exec SEND_REDIRECT failed: errno=3D%d", + errno); + memset(&cont_resp, 0, sizeof(cont_resp)); + cont_resp.id =3D req.id; + cont_resp.flags =3D + SECCOMP_USER_NOTIF_FLAG_CONTINUE; + ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND, + &cont_resp); + } + } + } + + if (waitpid(pid, &status, WNOHANG) =3D=3D 0) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + } + EXPECT_EQ(true, loader_known) { + TH_LOG("find_loader_text_range never resolved"); + } + EXPECT_EQ(true, post_exec_install_ok); + EXPECT_EQ(true, post_exec_redirect_done); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + +cleanup_scm: + if (waitpid(pid, &status, WNOHANG) =3D=3D 0) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + } + munmap(sup_view, PIN_SIZE); + close(memfd); + close(listener); +} + +/* + * One listener serves many short-lived targets. PIN_INSTALL keeps no + * per-target state (the sealed VMA is the only record, re-validated at + * SEND_REDIRECT), so every install/redirect must succeed across the loop = with + * nothing accumulating (run under kmemleak/KASAN to confirm). + */ +TEST(user_notification_pinned_memfd_churn) +{ + const size_t PIN_SIZE =3D 4096; + const char *safe_path =3D "/dev/null"; + const int iters =3D 16; + int listener, memfd, i; + char *sup_view; + long ret; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + memfd =3D make_pin_memfd(_metadata, "pinned-reap", PIN_SIZE, + &sup_view, safe_path); + + listener =3D user_notif_syscall(__NR_openat, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(goto cleanup, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + for (i =3D 0; i < iters; i++) { + struct seccomp_notif req =3D {}; + struct seccomp_notif_pin_install pin =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + int status; + pid_t pid; + + pid =3D fork(); + ASSERT_GE(pid, 0); + if (pid =3D=3D 0) { + int fd =3D syscall(__NR_openat, AT_FDCWD, + "/never/touched", O_RDONLY, 0); + _exit(fd < 0 ? 11 : 0); + } + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_openat); + + pin.id =3D req.id; + pin.memfd =3D memfd; + pin.target_addr =3D 0; + pin.size =3D PIN_SIZE; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_PIN_INSTALL, + &pin)) { + if (errno =3D=3D EINVAL) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + SKIP(goto cleanup, + "Kernel lacks pinned-memfd remote install"); + } + TH_LOG("iter %d PIN_INSTALL failed: errno=3D%d", i, errno); + } + + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 1; + redir.ptr_mask =3D 1U << 1; + redir.memfd =3D memfd; + redir.ptr_len[1] =3D strlen(safe_path) + 1; + redir.args[1] =3D pin.target_addr; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, + &redir)) { + kill(pid, SIGKILL); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + TH_LOG("iter %d child exit %d (11=3Dopenat fail)", + i, WEXITSTATUS(status)); + } + /* + * Target is dead now; its pin (this iter's mm, at the + * kernel-chosen address) is stale. The next iteration's + * PIN_INSTALL walk must reap it rather than leak the range + + * mm + memfd reference. + */ + } + +cleanup: + munmap(sup_view, PIN_SIZE); + close(memfd); + close(listener); +} + +#ifdef __NR_socket +/* + * A redirect must not smuggle a syscall past an outer filter. Stack: outer + * blocks socket(AF_INET) with EACCES (else ALLOW), inner notifies on sock= et. + * The child calls socket(AF_UNIX) (outer allows, inner fires); the superv= isor + * SEND_REDIRECTs arg0 to AF_INET. The kernel must re-run the outer filter + * against the rewritten args and block it with EACCES. + */ +TEST(user_notification_redirect_outer_refilter) +{ + struct sock_filter outer_filter[] =3D { + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, + offsetof(struct seccomp_data, nr)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, __NR_socket, 0, 3), + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, syscall_arg(0)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, AF_INET, 0, 1), + BPF_STMT(BPF_RET | BPF_K, + SECCOMP_RET_ERRNO | (EACCES & SECCOMP_RET_DATA)), + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_ALLOW), + }; + struct sock_fprog outer_prog =3D { + .len =3D (unsigned short)ARRAY_SIZE(outer_filter), + .filter =3D outer_filter, + }; + struct seccomp_notif req =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + int status, listener; + pid_t pid; + long ret; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + /* Outer filter first =3D> it becomes the outer/root of the stack. */ + ASSERT_EQ(0, seccomp(SECCOMP_SET_MODE_FILTER, 0, &outer_prog)); + + /* Inner USER_NOTIF filter second (innermost); returns the listener. */ + listener =3D user_notif_syscall(__NR_socket, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(return, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + int fd =3D syscall(__NR_socket, AF_UNIX, SOCK_STREAM, 0); + + if (fd >=3D 0) + _exit(12); + if (errno !=3D EACCES) + _exit(13); + _exit(0); + } + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_socket); + EXPECT_EQ(req.data.args[0], AF_UNIX); + + /* Scalar redirect of arg0 (no pin needed): AF_UNIX -> AF_INET. */ + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 0; + redir.args[0] =3D AF_INET; + ret =3D ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, &redir); + if (ret < 0 && errno =3D=3D EINVAL) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + close(listener); + SKIP(return, "Kernel lacks SECCOMP_IOCTL_NOTIF_SEND_REDIRECT"); + } + EXPECT_EQ(0, ret); + if (ret) + kill(pid, SIGKILL); + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + switch (WEXITSTATUS(status)) { + case 12: + TH_LOG("child exit 12: redirect bypassed the outer filter"); + break; + case 13: + TH_LOG("child exit 13: socket failed with unexpected errno"); + break; + default: + TH_LOG("child exit %d (unexpected)", WEXITSTATUS(status)); + } + } + + close(listener); +} +#endif /* __NR_socket */ + +/* + * SEND_REDIRECT returns -EOPNOTSUPP for syscalls whose register substitut= ion + * is unsafe: sigreturn (frame restore fights the redirect restore) and the + * clone/fork family (the child would inherit substituted regs with no + * restore). The target installs the filter and hands the listener over a + * socketpair, since a filter trapping fork/clone can't be installed before + * the supervisor forks the target. + */ +TEST(user_notification_redirect_denied_syscalls) +{ + static const int denied[] =3D { + __NR_rt_sigreturn, + __NR_clone, + __NR_clone3, +#ifdef __NR_fork + __NR_fork, +#endif +#ifdef __NR_vfork + __NR_vfork, +#endif + }; + unsigned int i; + long ret; + + for (i =3D 0; i < ARRAY_SIZE(denied); i++) { + struct seccomp_notif req =3D {}; + struct seccomp_notif_resp resp =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + int sk[2], listener, status; + pid_t pid; + + if (denied[i] < 0) + continue; + + ASSERT_EQ(0, socketpair(AF_UNIX, SOCK_STREAM, 0, sk)); + + pid =3D fork(); + ASSERT_GE(pid, 0); + if (pid =3D=3D 0) { + close(sk[0]); + if (prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0)) + _exit(1); + listener =3D user_notif_syscall( + denied[i], + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0) + _exit(2); + if (send_fd(sk[1], listener)) + _exit(3); + syscall(denied[i], 0, 0, 0, 0, 0, 0); + _exit(0); + } + + close(sk[1]); + listener =3D recv_fd(sk[0]); + if (listener < 0) { + waitpid(pid, &status, 0); + close(sk[0]); + SKIP(return, "SECCOMP_FILTER_FLAG_REDIRECT unsupported"); + } + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, denied[i]); + + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 0; + redir.args[0] =3D 0; + errno =3D 0; + ret =3D ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, &redir); + EXPECT_EQ(-1, ret); + EXPECT_EQ(EOPNOTSUPP, errno) { + TH_LOG("nr %d: SEND_REDIRECT errno %d, want EOPNOTSUPP", + denied[i], errno); + } + + resp.id =3D req.id; + resp.error =3D -EPERM; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND, &resp)); + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + TH_LOG("nr %d: child exit %d", denied[i], + WEXITSTATUS(status)); + } + close(listener); + close(sk[0]); + } +} + +#ifdef __NR_socket +/* + * Re-validation walks *every* outer filter, not just the nearest. Stack: + * outer blocks socket(AF_INET) with EACCES, middle ALLOWs all, inner noti= fies. + * After the redirect to AF_INET, the walk must pass the permissive middle= and + * still reach the outer EACCES; stopping at the middle would slip it past. + */ +TEST(user_notification_redirect_revalidate_chain) +{ + struct sock_filter outer_filter[] =3D { + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, + offsetof(struct seccomp_data, nr)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, __NR_socket, 0, 3), + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, syscall_arg(0)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, AF_INET, 0, 1), + BPF_STMT(BPF_RET | BPF_K, + SECCOMP_RET_ERRNO | (EACCES & SECCOMP_RET_DATA)), + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_ALLOW), + }; + struct sock_fprog outer_prog =3D { + .len =3D (unsigned short)ARRAY_SIZE(outer_filter), + .filter =3D outer_filter, + }; + struct sock_filter allow_filter[] =3D { + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_ALLOW), + }; + struct sock_fprog allow_prog =3D { + .len =3D (unsigned short)ARRAY_SIZE(allow_filter), + .filter =3D allow_filter, + }; + struct seccomp_notif req =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + int status, listener; + pid_t pid; + long ret; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + /* outer (root) -> middle (permissive) -> inner (notifier). */ + ASSERT_EQ(0, seccomp(SECCOMP_SET_MODE_FILTER, 0, &outer_prog)); + ASSERT_EQ(0, seccomp(SECCOMP_SET_MODE_FILTER, 0, &allow_prog)); + listener =3D user_notif_syscall(__NR_socket, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(return, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + int fd =3D syscall(__NR_socket, AF_UNIX, SOCK_STREAM, 0); + + if (fd >=3D 0) + _exit(12); + if (errno !=3D EACCES) + _exit(13); + _exit(0); + } + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_socket); + EXPECT_EQ(req.data.args[0], AF_UNIX); + + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 0; + redir.args[0] =3D AF_INET; + ret =3D ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, &redir); + if (ret < 0 && errno =3D=3D EINVAL) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + close(listener); + SKIP(return, "Kernel lacks SECCOMP_IOCTL_NOTIF_SEND_REDIRECT"); + } + EXPECT_EQ(0, ret); + if (ret) + kill(pid, SIGKILL); + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + switch (WEXITSTATUS(status)) { + case 12: + TH_LOG("exit 12: walk stopped at middle; outer bypassed"); + break; + case 13: + TH_LOG("exit 13: socket failed with unexpected errno"); + break; + default: + TH_LOG("child exit %d (unexpected)", WEXITSTATUS(status)); + } + } + + close(listener); +} + +/* + * An outer TRACE verdict over a redirected syscall fails closed: a tracer + * rewrite can't be re-composed mid-walk, so the kernel skips the call with + * -ENOSYS (as if no tracer) and fires no PTRACE_EVENT_SECCOMP. Stack: out= er + * TRACEs socket(AF_INET) (else ALLOW), inner notifies. The traced child c= alls + * socket(AF_UNIX); after the redirect to AF_INET the child must see ENOSY= S and + * the tracer a plain exit, not a seccomp event stop. + */ +TEST(user_notification_redirect_outer_trace) +{ + struct sock_filter outer_filter[] =3D { + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, + offsetof(struct seccomp_data, nr)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, __NR_socket, 0, 3), + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, syscall_arg(0)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, AF_INET, 0, 1), + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_TRACE | 0x23), + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_ALLOW), + }; + struct sock_fprog outer_prog =3D { + .len =3D (unsigned short)ARRAY_SIZE(outer_filter), + .filter =3D outer_filter, + }; + struct seccomp_notif req =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + int status, listener; + pid_t pid; + long ret; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + /* Outer filter first =3D> it becomes the outer/root of the stack. */ + ASSERT_EQ(0, seccomp(SECCOMP_SET_MODE_FILTER, 0, &outer_prog)); + + /* Inner USER_NOTIF filter second (innermost); returns the listener. */ + listener =3D user_notif_syscall(__NR_socket, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(return, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + int fd; + + if (ptrace(PTRACE_TRACEME, 0, NULL, NULL)) + _exit(14); + if (raise(SIGSTOP)) + _exit(15); + + fd =3D syscall(__NR_socket, AF_UNIX, SOCK_STREAM, 0); + if (fd >=3D 0) + _exit(12); + if (errno !=3D ENOSYS) + _exit(13); + _exit(0); + } + + /* Tracer handshake: enable PTRACE_EVENT_SECCOMP reporting. */ + ASSERT_EQ(pid, waitpid(pid, &status, 0)); + ASSERT_EQ(true, WIFSTOPPED(status)); + ASSERT_EQ(SIGSTOP, WSTOPSIG(status)); + ASSERT_EQ(0, ptrace(PTRACE_SETOPTIONS, pid, NULL, + PTRACE_O_TRACESECCOMP)); + ASSERT_EQ(0, ptrace(PTRACE_CONT, pid, NULL, 0)); + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_socket); + EXPECT_EQ(req.data.args[0], AF_UNIX); + + /* Scalar redirect of arg0 (no pin needed): AF_UNIX -> AF_INET. */ + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 0; + redir.args[0] =3D AF_INET; + ret =3D ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, &redir); + if (ret < 0 && (errno =3D=3D EINVAL || errno =3D=3D EOPNOTSUPP)) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + close(listener); + SKIP(return, "Kernel lacks SECCOMP_IOCTL_NOTIF_SEND_REDIRECT"); + } + EXPECT_EQ(0, ret); + + /* + * A plain exit, not a ptrace stop: the tracer must never see a + * PTRACE_EVENT_SECCOMP for the substituted call. + */ + EXPECT_EQ(pid, waitpid(pid, &status, 0)); + EXPECT_EQ(true, WIFEXITED(status)) { + if (WIFSTOPPED(status) && + status >> 8 =3D=3D (SIGTRAP | (PTRACE_EVENT_SECCOMP << 8))) + TH_LOG("tracer got PTRACE_EVENT_SECCOMP for the redirected call"); + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + } + EXPECT_EQ(0, WEXITSTATUS(status)) { + switch (WEXITSTATUS(status)) { + case 12: + TH_LOG("child exit 12: redirected socket ran past the outer TRACE filte= r"); + break; + case 13: + TH_LOG("child exit 13: socket failed with unexpected errno (want ENOSYS= )"); + break; + default: + TH_LOG("child exit %d (unexpected)", WEXITSTATUS(status)); + } + } + + close(listener); +} + +/* + * An outer USER_NOTIF verdict over a redirected syscall, where the outer + * filter has no listener (installed without NEW_LISTENER): it resolves li= ke + * any listener-less USER_NOTIF, skipping with -ENOSYS. Stack: outer notif= ies + * on socket(AF_INET) (no listener, else ALLOW), inner notifies (the + * redirector). After the redirect to AF_INET the child must see ENOSYS. + * + * (An outer filter with a live listener is also reachable -- only a second + * redirect-capable filter is rejected with EBUSY, not a second plain list= ener + * -- and drives the walk's notify path; not covered here.) + */ +TEST(user_notification_redirect_outer_notify) +{ + struct sock_filter outer_filter[] =3D { + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, + offsetof(struct seccomp_data, nr)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, __NR_socket, 0, 3), + BPF_STMT(BPF_LD | BPF_W | BPF_ABS, syscall_arg(0)), + BPF_JUMP(BPF_JMP | BPF_JEQ | BPF_K, AF_INET, 0, 1), + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_USER_NOTIF), + BPF_STMT(BPF_RET | BPF_K, SECCOMP_RET_ALLOW), + }; + struct sock_fprog outer_prog =3D { + .len =3D (unsigned short)ARRAY_SIZE(outer_filter), + .filter =3D outer_filter, + }; + struct seccomp_notif req =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + int status, listener; + pid_t pid; + long ret; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + /* Outer filter first =3D> it becomes the outer/root of the stack. */ + ASSERT_EQ(0, seccomp(SECCOMP_SET_MODE_FILTER, 0, &outer_prog)); + + /* Inner USER_NOTIF filter second (innermost); returns the listener. */ + listener =3D user_notif_syscall(__NR_socket, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(return, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + int fd =3D syscall(__NR_socket, AF_UNIX, SOCK_STREAM, 0); + + if (fd >=3D 0) + _exit(12); + if (errno !=3D ENOSYS) + _exit(13); + _exit(0); + } + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_socket); + EXPECT_EQ(req.data.args[0], AF_UNIX); + + /* Scalar redirect of arg0 (no pin needed): AF_UNIX -> AF_INET. */ + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 0; + redir.args[0] =3D AF_INET; + ret =3D ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, &redir); + if (ret < 0 && (errno =3D=3D EINVAL || errno =3D=3D EOPNOTSUPP)) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + close(listener); + SKIP(return, "Kernel lacks SECCOMP_IOCTL_NOTIF_SEND_REDIRECT"); + } + EXPECT_EQ(0, ret); + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + switch (WEXITSTATUS(status)) { + case 12: + TH_LOG("child exit 12: redirect ran past the outer USER_NOTIF filter"); + break; + case 13: + TH_LOG("child exit 13: socket failed with unexpected errno (want ENOSYS= )"); + break; + default: + TH_LOG("child exit %d (unexpected)", WEXITSTATUS(status)); + } + } + + close(listener); +} +#endif /* __NR_socket */ + +#ifdef __x86_64__ +/* + * ABI check: after SEND_REDIRECT the redirected arg register must be rest= ored + * before user mode resumes (via task_work_add(TWA_RESUME) -> + * seccomp_redirect_restore_cb). Raw asm bypasses libc's syscall() wrapper + * (which caller-saves args and would mask a restore bug) and captures RSI + * right after the SYSCALL. The child openat's with RSI =3D sentinel_path;= the + * supervisor redirects RSI into the pin. A correct restore leaves RSI =3D= =3D + * sentinel; a broken one leaves the pin address. + */ +TEST(user_notification_pinned_memfd_abi) +{ + pid_t pid; + long ret; + int status, listener, memfd; + struct seccomp_notif req =3D {}; + struct seccomp_notif_pin_install pin =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + char *sup_view; + const size_t PIN_SIZE =3D 4096; + const char *safe_path =3D "/dev/null"; + /* + * The "sentinel" is a real string the child can also pass as + * the openat path. Its address is captured pre-syscall as RSI; + * post-syscall RSI must equal the same address. + */ + static const char sentinel_path[] =3D "/seccomp_abi_sentinel"; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + memfd =3D make_pin_memfd(_metadata, "pin-abi", PIN_SIZE, + &sup_view, safe_path); + + listener =3D user_notif_syscall(__NR_openat, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(goto cleanup, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + register long r10_val asm("r10") =3D 0; + unsigned long rsi_after; + long fd; + + asm volatile( + "syscall\n\t" + "mov %%rsi, %[after]" + : "=3Da"(fd), [after] "=3D&r"(rsi_after) + : "0"((long)__NR_openat), + "D"((long)AT_FDCWD), + "S"((unsigned long)sentinel_path), + "d"((long)O_RDONLY), + "r"(r10_val) + : "rcx", "r11", "memory" + ); + + if (fd < 0) + _exit(11); + /* + * Load-bearing check: RSI immediately post-SYSCALL must + * still be the sentinel pointer the child passed in. The + * kernel's REDIRECT-then-restore mechanism is the only + * thing that guarantees this; a broken restore would leave + * the pin address in RSI. + */ + if (rsi_after !=3D (unsigned long)sentinel_path) + _exit(12); + _exit(0); + } + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_openat); + EXPECT_EQ(req.data.args[1], (unsigned long)sentinel_path); + + pin.id =3D req.id; + pin.memfd =3D memfd; + pin.target_addr =3D 0; + pin.size =3D PIN_SIZE; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_PIN_INSTALL, &pin)) { + if (errno =3D=3D EINVAL) { + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + SKIP(goto cleanup, + "Kernel lacks pinned-memfd remote install"); + } + } + + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 1; + redir.ptr_mask =3D 1U << 1; + redir.memfd =3D memfd; + redir.ptr_len[1] =3D strlen(safe_path) + 1; + redir.args[1] =3D pin.target_addr; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, + &redir)) { + kill(pid, SIGKILL); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + switch (WEXITSTATUS(status)) { + case 11: + TH_LOG("child exit 11: openat returned -errno"); + break; + case 12: + TH_LOG("child exit 12: ABI violation -- RSI not restored after redirect= "); + break; + default: + TH_LOG("child exit %d (unexpected)", WEXITSTATUS(status)); + } + } + +cleanup: + munmap(sup_view, PIN_SIZE); + close(memfd); + close(listener); +} + +static void redir_sigusr1_handler(int signo) +{ + /* _exit() is async-signal-safe; bail with a distinct code if the + * signal frame was clobbered so the handler sees the wrong signo. + */ + if (signo !=3D SIGUSR1) + _exit(12); +} + +/* + * The redirect's deferred arg-register restore must run before a signal f= rame + * is built. get_signal() runs task_work_run() before dequeuing a signal, = so + * the TWA_RESUME restore lands before handle_signal() sets up the handler + * frame (regs->di =3D signo, ...); a restore running after would clobber = it. The + * child traps on pause() with a sentinel in RDI, the supervisor redirects + * arg0, then sends SIGUSR1; the handler must see signo =3D=3D SIGUSR1. + */ +TEST(user_notification_redirect_signal_abi) +{ + pid_t pid; + long ret; + int status, listener; + struct seccomp_notif req =3D {}; + struct seccomp_notif_resp_redirect redir =3D {}; + /* A recognizable original RDI the broken restore would leak in. */ + const unsigned long RDI_SENTINEL =3D 0x5a5a5a5aUL; + + ret =3D prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0); + ASSERT_EQ(0, ret) { + TH_LOG("Kernel does not support PR_SET_NO_NEW_PRIVS!"); + } + + listener =3D user_notif_syscall(__NR_pause, + SECCOMP_FILTER_FLAG_NEW_LISTENER | + SECCOMP_FILTER_FLAG_REDIRECT); + if (listener < 0 && errno =3D=3D EINVAL) + SKIP(goto cleanup, "Kernel lacks SECCOMP_FILTER_FLAG_REDIRECT"); + ASSERT_GE(listener, 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + struct sigaction sa =3D { + .sa_handler =3D redir_sigusr1_handler, + }; + long rc; + + if (sigaction(SIGUSR1, &sa, NULL)) + _exit(10); + + /* Raw pause() carrying a controlled RDI sentinel. */ + asm volatile( + "syscall" + : "=3Da"(rc) + : "0"((long)__NR_pause), + "D"(RDI_SENTINEL) + : "rcx", "r11", "memory"); + + if (rc !=3D -EINTR) + _exit(11); + _exit(0); + } + + ASSERT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_RECV, &req)); + EXPECT_EQ(req.data.nr, __NR_pause); + EXPECT_EQ(req.data.args[0], RDI_SENTINEL); + + /* Redirect arg0 (non-pointer); this arms the original-RDI restore. */ + redir.id =3D req.id; + redir.flags =3D SECCOMP_REDIRECT_FLAG_CONTINUE; + redir.args_mask =3D 1U << 0; + redir.args[0] =3D 0; + EXPECT_EQ(0, ioctl(listener, SECCOMP_IOCTL_NOTIF_SEND_REDIRECT, + &redir)) { + int einval =3D (errno =3D=3D EINVAL); + + kill(pid, SIGKILL); + waitpid(pid, &status, 0); + if (einval) + SKIP(goto cleanup, + "Kernel lacks SECCOMP_IOCTL_NOTIF_SEND_REDIRECT"); + goto cleanup; + } + + usleep(100000); + EXPECT_EQ(0, kill(pid, SIGUSR1)); + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)) { + switch (WEXITSTATUS(status)) { + case 10: + TH_LOG("child exit 10: sigaction failed"); + break; + case 11: + TH_LOG("child exit 11: pause() did not return -EINTR"); + break; + case 12: + TH_LOG("child exit 12: handler saw wrong signo (frame clobbered)"); + break; + default: + TH_LOG("child exit %d (unexpected)", WEXITSTATUS(status)); + } + } + +cleanup: + close(listener); +} +#endif /* __x86_64__ */ + #ifndef SECCOMP_USER_NOTIF_FD_SYNC_WAKE_UP #define SECCOMP_USER_NOTIF_FD_SYNC_WAKE_UP (1UL << 0) #define SECCOMP_IOCTL_NOTIF_SET_FLAGS SECCOMP_IOW(4, __u64) --=20 2.43.0