From nobody Sat Sep 26 09:19:24 2026 Received: from mail-qk1-f174.google.com (mail-qk1-f174.google.com [209.85.222.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E6E5A3B388C for ; Wed, 2 Sep 2026 19:47:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.174 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378444; cv=none; b=cdApO0NymW/CG6Id6NxBX2gOzZ1C67h6niGWq705MMObPsa3DLRCGWM18AML0mU+EBUmsqK09eioVYSAYSkoIP67miijWaYP8CFGgppOroGs2+1WOBzpg00GU/ss7JR6hMMH3L+DtUoTAcXE7p2/XHaf3x5j6TUxYJBT0A91GKg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378444; c=relaxed/simple; bh=0nDpyHAZVWDLsoBhhy63Ifby8fFRcSPcb4Ws9onqaLg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=VqPZijOpMQLeGeK2BZ90p77EPATxJv1kLG0pCNyXbWOr61vNCnfQE7er5Q9Dk1NBpO3HrnCbpnI4vWB79WGpJEjaY/7kdXAU1MKAZA3oG3+q6oxynRpIMMJxccufDemv8UV8jSLQXa33mHVf6ahecP3SbVxWxxESdZveEeYFqQg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=JNrTBf0Q; arc=none smtp.client-ip=209.85.222.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="JNrTBf0Q" Received: by mail-qk1-f174.google.com with SMTP id af79cd13be357-9387752a4d0so134003985a.2 for ; Wed, 02 Sep 2026 12:47:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1788378425; x=1788983225; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=kwKfYOaniBtFKBvRoVSK69Bg/ErimmVxxQRLm+ngzMY=; b=JNrTBf0QaqiJDpswctkn5+CjUZNAAAW0SellCxPU4lcN/kf+eoAXqdf243NdTrkwyI PWSnrGVu/E9VS2/SIp+hykiDowIqM7WldrqNVoNnmJeYvLo1EgIeE93PDWWqsi+EYqXF hqo11CYhlUS56yzIqBqdPduY4K0m1+Ke5TXYH5Sd7ObQ7x08xUHEGXBwCJbdftHhPeeX 2cXNW+UWJV9rAj5GHfUUZL+Toab7eJ0Rt8GkeHaXYfMZs4l0hyTJnyC+NIYi7wCuQ3bC 7XLEFs2oMNGLmS7qUdQfwPu0AKFM8wHSsKGQG1IAF+ZHeiNkQKO2nS3XjxoceH2Y1bjV 9RXA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788378425; x=1788983225; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=kwKfYOaniBtFKBvRoVSK69Bg/ErimmVxxQRLm+ngzMY=; b=jdfjbcWd4Zl+IX0tJ4eNFw14lgFvyxkAYsxcSABs2CLjj0r2iFunxckpZ3I08Pjz60 uuHFXA0kp6Bt8z0Oo6cp6S9i8d5ECc6TGBuRHtbBq625dwic4P901HPNPMJtxMwxBC3I CzQwy6ZBHFnlhqHd8zFUo/qs+q/6ksik/s2GuhI3e0SYzPOX/KpE+yjcitrw00gMLgVq VR2p8LwGxR7Eg+RAdCWAf0GsQIU+BjlcDg60iJcKmOFOFQaXgcg6b9Tf2dhxWPERHEkL SLJhoSGpEAH3FNjuXnyKhSuqLHTdQfnZrVW7NCuQWJKHx1rcF2iQS9LbjJzkJ1cMwkSw oy+Q== X-Forwarded-Encrypted: i=1; AKwUvBx3R06KEvj5nafwyGN57u9VmNX4Kgwbd7a03L0f/jUvBmIZu0kSAsOt3SWBQlBFu+SU9+IJXnMBeJ2wXKg=@vger.kernel.org X-Gm-Message-State: AFuF++m44117aPQYZtQSvC+UGJN/jdaxgpZlhKS8EKVFSuX2uRR1hYlR GvVB2iCAFopjjgrY2lZ3oEWLAZXjj06+pnHPjTY0T3zMFp3ZJC68HsftFq90FRIjWjI= X-Gm-Gg: AYBFou2GTp/oFcPjaR+3oiBWzbmgingNkbc9mCWgM3THuxA9/n9PkaJ0QBPpYg5ViRT 2YiQ4IP5H8ZujNk1MjX0CkntMKPE79AbSvzasxZbjD+I57gYoimYSUKW2qmLrXofICaUdxSNTNn z2LFskEXZXRhtTzlwe/zRxUooEXYVcGEIg7Z7+uqa+CyOh3HqfRYOoWZGO8V+99Z0KzXUkpJm1N 7jNxG9D/Kcm5/ksFQZMS4OnA8qFMgiBqKyGAsvkYK6X+7gws3VD7bSSR1+g9fIWcjpA0NBEfx+9 VyAjFQdR1CYydOIC+960s9T6Xl87NQgqzIu430DgxqKxZguJsZ0SMte7A99znsBXlokZWVaiwdG jPz/KYxftwvWOaKp+fYENe+Vzc9rJsKty+2tfLG8XOisF+cBrAUHB+PcQP9VfTpZBaJ5W6m7SRM bXDg12i3hMR1JFrZUMeHK2ibc8itcvrQVBlRT0Iug+zl7XYE1dEsXBf3nQFXzT7HrMrHBknHJ4b Ubq9y2LPoIE/4uVWDJmKDMglw5iBkwdYWiuCr+palN3yg7NKQ== X-Received: by 2002:a05:620a:4586:b0:930:9585:e08e with SMTP id af79cd13be357-9396f83aa9cmr61071585a.10.1788378421813; Wed, 02 Sep 2026 12:47:01 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-9395f18801asm299300585a.16.2026.09.02.12.47.01 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 02 Sep 2026 12:47:01 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, pbonzini@redhat.com, seanjc@google.com, akpm@linux-foundation.org, david@kernel.org, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, shuah@kernel.org Subject: [PATCH 1/5] mm/mempolicy: add mempolicy_create() Date: Wed, 2 Sep 2026 15:46:53 -0400 Message-ID: <20260902194657.79075-2-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260902194657.79075-1-gourry@gourry.net> References: <20260902194657.79075-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" do_set_mempolicy() builds a validated policy and installs it into the current running task. An in-kernel user that wants to construct a mempolicy wants the same validation without the install step. Add mempolicy_create(mode, flags, nodes): it allocates the policy and contextualises it against the calling task's cpuset, and returns the mempolicy to the caller (or ERR_PTR on error). do_set_mempolicy() is deliberately left alone rather than reimplemented on top of it. For syscall users, mpol_set_nodemask() and the current task policy swap must run under the task_lock(current) acquisition. Export the symbol to kvm for use in guest_memfd integration. Signed-off-by: Gregory Price --- include/linux/mempolicy.h | 3 +++ mm/mempolicy.c | 40 ++++++++++++++++++++++++++++++++++++++- 2 files changed, 42 insertions(+), 1 deletion(-) diff --git a/include/linux/mempolicy.h b/include/linux/mempolicy.h index 65c732d440d2f..aef018ad92317 100644 --- a/include/linux/mempolicy.h +++ b/include/linux/mempolicy.h @@ -128,6 +128,9 @@ void mpol_free_shared_policy(struct shared_policy *sp); struct mempolicy *mpol_shared_policy_lookup(struct shared_policy *sp, pgoff_t idx); =20 +struct mempolicy *mempolicy_create(unsigned short mode, unsigned short fla= gs, + nodemask_t *nodes); + struct mempolicy *get_task_policy(struct task_struct *p); struct mempolicy *__get_vma_policy(struct vm_area_struct *vma, unsigned long addr, pgoff_t *ilx); diff --git a/mm/mempolicy.c b/mm/mempolicy.c index 2ad0a5f18280a..da133ffe0b1c8 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -1085,7 +1085,45 @@ static int mbind_range(struct vma_iterator *vmi, str= uct vm_area_struct *vma, return vma_replace_policy(vma, new_pol); } =20 -/* Set the process memory policy */ +/** + * mempolicy_create - build a validated, cpuset-contextualised mempolicy + * @mode: MPOL_* mode + * @flags: MPOL_F_* flags + * @nodes: target nodemask, or NULL (interpreted per @mode; see mpol_new()) + * + * Creates a new policy and constrains it to the task's cpuset. + * + * The caller owns the returned reference and frees it with mpol_put(). + * + * Return: the policy (NULL for a default policy), or an ERR_PTR on failur= e. + */ +struct mempolicy *mempolicy_create(unsigned short mode, unsigned short fla= gs, + nodemask_t *nodes) +{ + struct mempolicy *pol; + NODEMASK_SCRATCH(scratch); + int err; + + if (!scratch) + return ERR_PTR(-ENOMEM); + + pol =3D mpol_new(mode, flags, nodes); + if (IS_ERR(pol)) + goto out; + + task_lock(current); + err =3D mpol_set_nodemask(pol, nodes, scratch); + task_unlock(current); + if (err) { + mpol_put(pol); + pol =3D ERR_PTR(err); + } +out: + NODEMASK_SCRATCH_FREE(scratch); + return pol; +} +EXPORT_SYMBOL_FOR_MODULES(mempolicy_create, "kvm"); + static long do_set_mempolicy(unsigned short mode, unsigned short flags, nodemask_t *nodes) { --=20 2.53.0-Meta From nobody Sat Sep 26 09:19:24 2026 Received: from mail-qk1-f180.google.com (mail-qk1-f180.google.com [209.85.222.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 541634AA1CA for ; Wed, 2 Sep 2026 19:47:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.180 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378439; cv=none; b=ZWoT0aYAvHPVH2Z4nqKYQ8bPpY3NrWaVzu7PPScVpg1cn2+tt3od843SEYKrkoej7d6Riu9rDvgicO3btF+ZW64pAaQJPITBgKr6o2kotzSCPumZJSnL8bfyIX6BrDhPiBUEStmelo9LoN4lkgQzAgPMExGKX1btwNsUVdgbsyg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378439; c=relaxed/simple; bh=wOiGSHTvuaGU2hGJD2n9TN+YSN62eWKKXnnL46WXyRo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ZAn10nWrHG7OqAzpyDz+i1ZlLsgrgrl9uHJKCysmTUloVhCbOEvQnGruUne4UzxpeLexoQwjTY2KmGCHlPEdwCn5EOYLnYntCT76sun56HUChzNnw/Yo+FHToOh3GFDGSjC5cgdfvpfESoVEYfc4xkexTseS9csaTplC2TAciSA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=TqzEg1GQ; arc=none smtp.client-ip=209.85.222.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="TqzEg1GQ" Received: by mail-qk1-f180.google.com with SMTP id af79cd13be357-9390dd46b45so109420485a.0 for ; Wed, 02 Sep 2026 12:47:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1788378423; x=1788983223; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=VW5hThLcNiS9B9Y8k40LTfPGZJYf7iWoviGkGwyNijs=; b=TqzEg1GQ0mynWcwNXcKYtfBFIQpz9l41oqKUyll6YjdI/vbPvKurVd3ahcqF5Z+9vC diu/Sc4/gcVQqbwhZKKA/r4yMh+2LIKGsuHPLT0kudAppx86QSYaPVUdEDZV7CimPozC uW0/aBAgRla3TRFKjL13fdzbjOYMSub4btlS7O83eMysyi1D4kPcfCSL9S2qQi6sCZgL 57JHb85z4KSV4+Pkg8cKbZ59Of6yO84dqwC5vZw6CG5C2STQ+51C1jBqYLKq7algduMm gNFo8AkMfCOTy0o5iht6PGXA3YaL9R1B4KFl+Ytl2BlSfzAU2mHJ65Q9S/z6tJIOmncc fpiw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788378423; x=1788983223; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=VW5hThLcNiS9B9Y8k40LTfPGZJYf7iWoviGkGwyNijs=; b=W7Or9Kn80kkCxMAPCyaI3le+uWfTqiPsM0WojLTiflneDWrM8wpJLoGCGHa6FtZQQm Tp2c134IokenTkAX0GuiyUdKBnq08WOK4ywl5koVsy7FPZhHaO/o1xEXcfAbY9qalYNy ch7sEmK8fOnNJe5bNlQuk1P0zv++vH412EHtB9yuh1p1PaDJWg8NAtDcCwI+2uTtxVXU Vhw50PZhiBNuxta/ZjC//weMcFatv8dO5tso5HeV0UMsucyucVxq1Op6uW8ZPW5kWsye mQ0JWqs1BbtfkM8wBTSOUKXcXqX5j1oR3ri3Oncy6g6PMyz7znPGRsAGE/vsP5a0KEuB CL7w== X-Forwarded-Encrypted: i=1; AKwUvBz4EB3gOkLNvYeeE65lWtL1RZArq6uQp+zJJ3dZBhEwhKqvlo4skWu2o5jYAZlJtt5xpLQ3zjRL09kXHxg=@vger.kernel.org X-Gm-Message-State: AFuF++l9u32ZhPaORVgGQjcTVzjsGBvLDaohkO4zfJi0+7hKl0TXwSpH +rfM7ivEtdj9TRW0gnm9vZAvNFhokt/06OBDPr6gK+tVfM8Xu1nS4CjLw9nDc2tdmRo= X-Gm-Gg: AYBFou3dJL2aipSxSPMT4i+JYc1GukP3E3LkopF8aGB2fc0EVGQ5YoJ+ZM/t9h6AwhC HPu3Kc2Pdssq3PzZvc0seDTM67rYw+1gVR7fV/Ea4MuE13WMNtL/xH+S68zKcjpxe5ZiNCKQzKh FNhJvIDBOF4QZMVQ6fgKcC6ohg/0BKBWt+0m5s3tUFOF8a94Qi38TvqoftmnkPXTBSEQw3nv0xl tEqEGH1FxfaLIO77zkiX32VGccvOYIoYiw02g87tlaK+wjRgdU1pkM5HmD0+la7EdHkbNUwa7wF cHsyz6seTjiwVcmbK//UgNK9yfxP34bVKxavbO2nVfnBZn/Emz6TBh2MHNmyvIIA4My7n9hridh jAxOmw2seBfzzVNqbYyXpwmnA1f7AUrnHWaj0w1t/Vw2S8+bPBAjFUfUeooU7a//0J4MORfiWcF Bs/9781TwYMDDN0RZ9uux43hupjEXCt01r7WpznNnef2fXF/OpJlhFQCtasB8kQmggH9dGUc5Ww zqz3YlVtbwCKSKiXloRuiHvcar0QIjr1+c8LRNg7zlx4nky9g== X-Received: by 2002:a05:620a:a40f:20b0:939:191c:cac7 with SMTP id af79cd13be357-93960e19503mr738239685a.13.1788378423217; Wed, 02 Sep 2026 12:47:03 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-9395f18801asm299300585a.16.2026.09.02.12.47.02 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 02 Sep 2026 12:47:02 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, pbonzini@redhat.com, seanjc@google.com, akpm@linux-foundation.org, david@kernel.org, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, shuah@kernel.org, Dave Jiang Subject: [PATCH 2/5] mm/mempolicy: add mpol_set_shared_policy_range() Date: Wed, 2 Sep 2026 15:46:54 -0400 Message-ID: <20260902194657.79075-3-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260902194657.79075-1-gourry@gourry.net> References: <20260902194657.79075-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" mpol_set_shared_policy() installs a policy over a VMA's page-offset range, which is the only programmatic way to populate an inode's shared policy after init. Two limitations make it unusable for binding an entire backing inode from in-kernel code: - It requires a VMA, so it cannot cover unmapped offsets. (e.g. unmapped file folios faulted by pagecache) - mpol_shared_policy_init(), the only no-VMA installer, reconstructs the policy from mpol->w.user_nodemask. That field is only populated for static/relative or mount-string policies. a policy built programmatically (e.g. by mempolicy_create()) leaves it empty and stores w.cpuset_mems_allowed instead and _init mangles it. Both matter to a caller that wants a whole inode bound at creation. guest_memfd is one: it has no VMA at that point and may never gain one for a given offset, and it builds its policy in-kernel rather than from a mount string, so neither existing installer can express it. Add mpol_set_shared_policy_range(), which installs an already-built, fully contextualised policy verbatim over an arbitrary [start, end) page range with no VMA. Reimplement mpol_set_shared_policy() as a thin wrapper that derives the range from the VMA, so both share a single underlying path. Suggested-by: Dave Jiang Co-developed-by: Dave Jiang Signed-off-by: Dave Jiang Signed-off-by: Gregory Price Assisted-by: Claude:claude-opus-4-8 --- include/linux/mempolicy.h | 2 ++ mm/mempolicy.c | 35 +++++++++++++++++++++++++++++------ 2 files changed, 31 insertions(+), 6 deletions(-) diff --git a/include/linux/mempolicy.h b/include/linux/mempolicy.h index aef018ad92317..398318175cec7 100644 --- a/include/linux/mempolicy.h +++ b/include/linux/mempolicy.h @@ -124,6 +124,8 @@ int vma_dup_policy(struct vm_area_struct *src, struct v= m_area_struct *dst); void mpol_shared_policy_init(struct shared_policy *sp, struct mempolicy *m= pol); int mpol_set_shared_policy(struct shared_policy *sp, struct vm_area_struct *vma, struct mempolicy *mpol); +int mpol_set_shared_policy_range(struct shared_policy *sp, pgoff_t start, + pgoff_t end, struct mempolicy *mpol); void mpol_free_shared_policy(struct shared_policy *sp); struct mempolicy *mpol_shared_policy_lookup(struct shared_policy *sp, pgoff_t idx); diff --git a/mm/mempolicy.c b/mm/mempolicy.c index da133ffe0b1c8..ce10ce4374643 100644 --- a/mm/mempolicy.c +++ b/mm/mempolicy.c @@ -3312,24 +3312,47 @@ void mpol_shared_policy_init(struct shared_policy *= sp, struct mempolicy *mpol) } EXPORT_SYMBOL_FOR_MODULES(mpol_shared_policy_init, "kvm"); =20 -int mpol_set_shared_policy(struct shared_policy *sp, - struct vm_area_struct *vma, struct mempolicy *pol) +/** + * mpol_set_shared_policy_range - install @pol over [@start, @end) of @sp + * @sp: the shared policy tree + * @start: first page offset (inclusive) + * @end: last page offset (exclusive) + * @pol: a fully-built, validated policy, or NULL to clear the range + * + * Installs @pol over the given range, replacing any overlapping policy. + * @sp takes its own reference, the caller retains its reference on @pol. + * + * The policy is not reconstructed, so the policy is preserved exactly. + * + * Unlike mpol_set_shared_policy(), no VMA is required, so a range that + * is never mapped into a VMA can be covered, including the whole file. + * + * Return: 0 on success, -ENOMEM on allocation failure. + */ +int mpol_set_shared_policy_range(struct shared_policy *sp, pgoff_t start, + pgoff_t end, struct mempolicy *pol) { - const pgoff_t pgoff =3D vma_start_pgoff(vma); - const pgoff_t pgoff_end =3D vma_end_pgoff(vma); struct sp_node *new =3D NULL; int err; =20 if (pol) { - new =3D sp_alloc(pgoff, pgoff_end, pol); + new =3D sp_alloc(start, end, pol); if (!new) return -ENOMEM; } - err =3D shared_policy_replace(sp, pgoff, pgoff_end, new); + err =3D shared_policy_replace(sp, start, end, new); if (err && new) sp_free(new); return err; } +EXPORT_SYMBOL_FOR_MODULES(mpol_set_shared_policy_range, "kvm"); + +int mpol_set_shared_policy(struct shared_policy *sp, + struct vm_area_struct *vma, struct mempolicy *pol) +{ + return mpol_set_shared_policy_range(sp, vma->vm_pgoff, + vma->vm_pgoff + vma_pages(vma), pol); +} EXPORT_SYMBOL_FOR_MODULES(mpol_set_shared_policy, "kvm"); =20 /* Free a backing policy store on inode delete. */ --=20 2.53.0-Meta From nobody Sat Sep 26 09:19:24 2026 Received: from mail-qk1-f178.google.com (mail-qk1-f178.google.com [209.85.222.178]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DCF943939D2 for ; Wed, 2 Sep 2026 19:47:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.178 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378439; cv=none; b=mop+iXDNeVIoagT2FESOV7qOaaV3ZVXMa+HOpjWacdHOd1BQBf8A6Rt1C0pFamguXEZp9ZjXfylLP2S8nm3i3yjGrhJUrczMXrbSLWu+2oPIqZDGOcIAz1vTSpR5LZaOlDDjBhhyJqauPZVOqpz4tuswgyvMMLj7iRCxrnnJry0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378439; c=relaxed/simple; bh=X0I+aFHOnDW1nXFeqe9LLIh2Xl92y2wIRvotCjhFCQA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uJWBONxjN3V+rwaWJi2blFqZX1GHluIeD8eOHGjDYpZvUliS9zJTamljsDgV9+bkzTlPRKALzENOv0860JCJT/9SVh4Mwf7m6fzjA1TlGjidPEjDi8XjIjoBmC8PM8AeAfR8NA+KZmyZw/9RPGymyeZtjxioxk5NOZDoXT9tu2w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=rTcSqMBt; arc=none smtp.client-ip=209.85.222.178 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="rTcSqMBt" Received: by mail-qk1-f178.google.com with SMTP id af79cd13be357-936e8bd9caaso137189385a.2 for ; Wed, 02 Sep 2026 12:47:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1788378425; x=1788983225; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=CBcqkxQZqwfVczJ4Zyf97FMnYegDKAQ3ju3Ot5Bn64M=; b=rTcSqMBtlRahGp1rVVVaUt/Vpi0CeT61OOyPS99c398F9fqbGdfv2lCtvkhOlLyUCF fWvKm8FzXO4VLEa6cjYLHztENQ7hDfD5lI8KlAMR0A5AtTnwGGHF/9le0tWkdywuKSbw MIQ1XbX3JnZInw7RsvsDYXSHZvY9uLDowG+fKQzQ746sOQwOCDTec0JNh8lOrlqaVEbq Gdq9g6+3yCRnm1rnrUTrSkvZNss5V2bZmgtnuMscmsESGo1neTVTyKralUjDNsnigIXM wYXMdmtZBIq/NpIa39K/2lZGqVfv5XXOJVurvsL/hRV2sUpRyq5BeJm1zd/vYd5XADZ7 uebg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788378425; x=1788983225; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=CBcqkxQZqwfVczJ4Zyf97FMnYegDKAQ3ju3Ot5Bn64M=; b=N429n8mA+JvUXL738Xz6snFRyDgFj/J7INNmVChiIt+x3LhGEBSwCjENCGR5G5aG3b QLtDRinGD6x4X1+sWZwLSy2FCuQ44UcdbjeGm+ZAVMAU+PTiDcknV3JJ2L34N1jLwitS rra5Jti0u9NVpaulTTLmtUSDlkOdtEjBX2T0dQAJrc+7NSxsYrobS3LX3L+NsyuzCcFn MBfY2ej7le1ztWm4N9FBFnxR6tL+2n+UAazh8mF1yd5DU8DOBaUfnW0E2kQoxj16eXr6 DXL+5v9TBBf1I9ondSm7es9ukdIAH++7E/crV5Y4Vqb+ZPXSy2oXP+zXCpu+913LegTr itnw== X-Forwarded-Encrypted: i=1; AKwUvBxsHcFprZ6GfcK21Xwp+tq2bnrY6cJRxNdniAYTy4H+c0NYuU8UcUqe+uxzukVBNmyTMSEHdW4M6J6WhP0=@vger.kernel.org X-Gm-Message-State: AFuF++nnKfT/EvdoOAKbfNduKeZ3T9kfi74uQYjnwhTWEiLG3d0kojGQ dO6j8cta2uRr219gmKKI4v2fPCoTQeL2Z0640jErZXLqNARw9fIGyaiNbcJuBbFjmmw= X-Gm-Gg: AYBFou2wJjqFTwKmfMlIOw6nW0enuoPjDzJI72Wn+mc0ZeJX0eVJHgMzsURfqdk72ET eN018JfQBODDR7HLD1ORxzQo3xPghp0vJx6WkuiOPVvy2n5SVMg87cgOFWrqAoF43O46uAR4nBJ /y3Y27Wweh2Ctas6aOTSjK0+QtJhi7GKUuNsd86sJdOKRHFvElswpuDZbXJlicBAJmBQwO/r+9S 2cMPtnimQotT5fgxC+xdoeEHqDXoAyXyP5Le3T7cGkUyqVL/oT5QrARQbAA34fGyIOMmH1zF8bo jnc6FRyHgG4Bz5aZD4aEewg1E8CxzjTc2fydEm790hpuT8hv+LSgWpbuFelnhPFbkB2d7FZbpI2 4PXL6NI1Eh/nC7arovGrPnnwg0spfGTAHu9pSA+7tGuQdJ8WYu2rJ1Kx+5n9odOiTdL87zO457F 7sn4b93rwu3nvwRJunmtfotMTpiPkCGYQJtHCvUEms+YBKEy2QKHf7Y+4ZyfPOEbCCFMR9sNZmv J+cX9sTPkOMRDVj2iu8eogSBFBvYMf0Prgk46tYO4bEAUMxORcuIB2FpIMO X-Received: by 2002:a05:620a:bd4:b0:939:3a8c:28f4 with SMTP id af79cd13be357-93960de25e9mr874489085a.10.1788378425144; Wed, 02 Sep 2026 12:47:05 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-9395f18801asm299300585a.16.2026.09.02.12.47.04 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 02 Sep 2026 12:47:04 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, pbonzini@redhat.com, seanjc@google.com, akpm@linux-foundation.org, david@kernel.org, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, shuah@kernel.org, Dave Jiang Subject: [PATCH 3/5] KVM: guest_memfd: bind backing memory to a NUMA node at creation Date: Wed, 2 Sep 2026 15:46:55 -0400 Message-ID: <20260902194657.79075-4-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260902194657.79075-1-gourry@gourry.net> References: <20260902194657.79075-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" guest_memfd presently allocates its page-cache folios through a per-inode shared mempolicy (kvm_gmem_get_folio()). Today that policy can only be set after the fact, via mbind() on a host mmap of the fd. This requires the fd to be mmap-able and cannot reach folios that are only ever guest-faulted (no host VMA). Neither holds for a non-mappable (confidential) guest_memfd. Add GUEST_MEMFD_FLAG_BIND_NODE. When set, KVM builds an MPOL_BIND policy for the requested node and installs it over the whole inode's shared policy, so every folio is allocated on the requested node with no userspace mbind(). The flag is advertised through KVM_CAP_GUEST_MEMFD_FLAGS only when CONFIG_NUMA is enabled. Two user-visible behaviors worth noting: - mempolicy_create() constrains the request against the calling task's cpuset. Requesting a node outside the cpuset mems_allowed results in the ioctl failing with -EINVAL. Binding is essentially subject to the same cpuset constraint as mbind(). This behavior is correct - a task cannot grant a guest_memfd access to a node it cannot access itself. - The policy hangs off the inode, and nothing rebinds an inode's shared policy on a later cpuset change. mpol_rebind_task() walks tsk->mempolicy mpol_rebind_mm() walks vma->vm_policy Neither walker reaches a struct shared_policy. The binding is therefore fixed for the life of the fd. That matches shmem, whose inode policy behaves the same way, and is the intent here: the node is a property of the guest's backing memory, not of whoever happens to hold the fd. Suggested-by: Dave Jiang Co-developed-by: Dave Jiang Signed-off-by: Dave Jiang Signed-off-by: Gregory Price Assisted-by: Claude:claude-opus-4-8 --- include/linux/kvm_host.h | 3 +++ include/uapi/linux/kvm.h | 5 ++++- virt/kvm/guest_memfd.c | 44 ++++++++++++++++++++++++++++++++++++++-- 3 files changed, 49 insertions(+), 3 deletions(-) diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 03bfc92864b6e..738e276633c1e 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -739,6 +739,9 @@ static inline u64 kvm_gmem_get_supported_flags(struct k= vm *kvm) if (!kvm || kvm_arch_supports_gmem_init_shared(kvm)) flags |=3D GUEST_MEMFD_FLAG_INIT_SHARED; =20 + if (IS_ENABLED(CONFIG_NUMA)) + flags |=3D GUEST_MEMFD_FLAG_BIND_NODE; + return flags; } #endif diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index ac2d77d149635..8d3ae7e2ead8e 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -1658,11 +1658,14 @@ struct kvm_memory_attributes { #define KVM_CREATE_GUEST_MEMFD _IOWR(KVMIO, 0xd4, struct kvm_create_guest= _memfd) #define GUEST_MEMFD_FLAG_MMAP (1ULL << 0) #define GUEST_MEMFD_FLAG_INIT_SHARED (1ULL << 1) +#define GUEST_MEMFD_FLAG_BIND_NODE (1ULL << 2) =20 struct kvm_create_guest_memfd { __u64 size; __u64 flags; - __u64 reserved[6]; + __u32 node; + __u32 pad; + __u64 reserved[5]; }; =20 #define KVM_PRE_FAULT_MEMORY _IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memor= y) diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 625e62e1a0318..dc9f071dd969b 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -423,6 +423,31 @@ static struct mempolicy *kvm_gmem_get_policy(struct vm= _area_struct *vma, */ return mpol_shared_policy_lookup(&GMEM_I(inode)->policy, pgoff); } + +static int kvm_gmem_bind_node(struct inode *inode, int node) +{ + struct mempolicy *pol; + nodemask_t nodes; + int err; + + if ((unsigned int)node >=3D MAX_NUMNODES) + return -EINVAL; + + init_nodemask_of_node(&nodes, node); + pol =3D mempolicy_create(MPOL_BIND, 0, &nodes); + if (IS_ERR(pol)) + return PTR_ERR(pol); + + err =3D mpol_set_shared_policy_range(&GMEM_I(inode)->policy, 0, + MAX_LFS_FILESIZE >> PAGE_SHIFT, pol); + mpol_put(pol); + return err; +} +#else +static int kvm_gmem_bind_node(struct inode *inode, int node) +{ + return -EINVAL; +} #endif /* CONFIG_NUMA */ =20 static const struct vm_operations_struct kvm_gmem_vm_ops =3D { @@ -520,7 +545,7 @@ bool __weak kvm_arch_supports_gmem_init_shared(struct k= vm *kvm) return true; } =20 -static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags) +static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags, int = node) { static const char *name =3D "[kvm-gmem]"; struct gmem_file *f; @@ -561,6 +586,12 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t s= ize, u64 flags) =20 GMEM_I(inode)->flags =3D flags; =20 + if (flags & GUEST_MEMFD_FLAG_BIND_NODE) { + err =3D kvm_gmem_bind_node(inode, node); + if (err) + goto err_inode; + } + file =3D alloc_file_pseudo(inode, kvm_gmem_mnt, name, O_RDWR, &kvm_gmem_f= ops); if (IS_ERR(file)) { err =3D PTR_ERR(file); @@ -593,6 +624,7 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_= guest_memfd *args) { loff_t size =3D args->size; u64 flags =3D args->flags; + int node =3D NUMA_NO_NODE; =20 if (flags & ~kvm_gmem_get_supported_flags(kvm)) return -EINVAL; @@ -600,7 +632,15 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create= _guest_memfd *args) if (size <=3D 0 || !PAGE_ALIGNED(size)) return -EINVAL; =20 - return __kvm_gmem_create(kvm, size, flags); + if (flags & GUEST_MEMFD_FLAG_BIND_NODE) { + if (args->pad || args->node >=3D MAX_NUMNODES) + return -EINVAL; + node =3D args->node; + } else if (args->node || args->pad) { + return -EINVAL; + } + + return __kvm_gmem_create(kvm, size, flags, node); } =20 int kvm_gmem_bind(struct kvm *kvm, struct kvm_memory_slot *slot, --=20 2.53.0-Meta From nobody Sat Sep 26 09:19:24 2026 Received: from mail-qv1-f45.google.com (mail-qv1-f45.google.com [209.85.219.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C55A52D876A for ; Wed, 2 Sep 2026 19:47:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.45 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378439; cv=none; b=QIRA3qAF0M/aBuJiHvXdPAisBN6a/L1vhRbSLwHIdn00NpoBJa5yPvOjksPBc26BMdkKlHqjJCr2K42cZiQU6P9E2yqQHCYcEt7isiSCLXaVJJRjRN7l8/B23mBdXM02ofZOrmQzqgWSItO8CJ7Av6CVJ4S1y1VNL7Jhf5Jigug= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378439; c=relaxed/simple; bh=3roPg4ZfXvCu9H2QE9Rl5CyrdVkiobJpwL6O2Q76emc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ROHLYm5SbV8p+/kG1JHfQ4xJX2+IleOX/H1xShBckBKfuwdFZyvAJywJNYIC9Mo6JgWIuVjQrzcCL6TQqCkL09lKVNmKvCvl8v1beXOv6CKNdGwPTRb3GQcjqjp9/sfUnB6nRq80f6Czw01FSTyybu0TsiqkomxW9t7haL4m3pU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=RqW2ef64; arc=none smtp.client-ip=209.85.219.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="RqW2ef64" Received: by mail-qv1-f45.google.com with SMTP id 6a1803df08f44-90ce3971a86so14631026d6.0 for ; Wed, 02 Sep 2026 12:47:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1788378427; x=1788983227; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=jK2MmrU5PHerjBBjkrNgI0UI4L0yhiHtWM989h7yfSE=; b=RqW2ef64Dc/9ZRfrHgRDo90wDTDKIu9w78a73wauD9AG7v8UkzkHNEhej8zt2y1ben S34SCE/9Ww4KgV7MmtL9YljeZ/UlvtKXcMqFf8OG7uE3/Xt4Z4jKh1RfZIKHFUET11GS XSWsIijEmmWLnx+Ccl2AYR8XNBqzxYEQz2YkBaFqsjQOTSaQ/SRmCGZn7taedg4FVFgH QoW3oVS0EoawlLIGaZQHooI8lFXwao0pSSxwvbvS9mwA/bPz0JVnJ6iZ2j+bDl3wArJL 9vANN+zk0eTfF0zh+lWRJMnfGqWXPNSlUwA8zGpqcW9ZrX4Y2rfndyJu/yFtJhVPf4I7 SXiA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788378427; x=1788983227; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=jK2MmrU5PHerjBBjkrNgI0UI4L0yhiHtWM989h7yfSE=; b=B9GlkR8AF7NjSja0lHa5P4bgJSIrg+Gao+cFLWeKf3ULdOpkzcxtxze9bBw4pDh9IJ z0q8YkhYn9k+3DBDHbzo8Y0CkqhumgTobY3+3pPuBH7w7aKXoaRWPomC9ssRBx8VZekD XRsbjYry+i+ut9uQ5DgMys1p/HAOhIL2OQMVYCgn/jOXL410b2fs+X5irUd5ZHVeoRrV IZGc/oz9f/CL/5vcmlehwQRdVBtqpi0bquwqzKjuTeoIHHnu5+utKPUiCm/bolOMOmFg d+8kRDuJXP84L1rc87nlaV5zrHHCoyOb9OxadNb0Q+s95xC777zM4sgxF+lP27nFvULK kLQQ== X-Forwarded-Encrypted: i=1; AKwUvBz/2ILLjB2Gc2ZC6qrBPTsbCDOwjjm3xZbXdg8hq/njQaJcz0/S8lm3QQFBts6YVS17hrocGydddlRUmFA=@vger.kernel.org X-Gm-Message-State: AFuF++nlK1KJlag47OcksZ9oSH/SIXZ3DRCH4Vi1lgdI37JMlERe09Ry +GZ64S2vdEutfj0QHYghTmIIfOd7B2t2e4Na0cproFVGIGhCx/SEuhc5Qhp3/u8oBGc= X-Gm-Gg: AYBFou2/GqEUaZ61q+kncrTxVGAtD042ZhbCz1UBcdoM9emn2sVBVPYhaZq/f2BODJB uuVfA/n2pDH1g3KBIZtfUilD+2URCLJSBWY4zRv7dySkzmBi9YxHQoZaEo0kx9KkG73nLgXAL9t hEklyW3ciDEjMc+A2nuwD8alT7/7+lNyAtPGlauVQi4dvGBaPRS7GBzoZrNsAPrILKGZZDS+5Bm FJ2WSYqoezVuvN5MfsmI/30E8/9zeMSJ6nseTcILrBIn4DHGqHLmV+rmITBm10f88vFtRv0MUHx BHAA1MvWOn0ctnjJXEOu7jVrNtXAIhVRjnHy7kfUhQ1eOnlr06Sskzxy80aaOawRMLr+rGss+Er A5WJde+dE/DJO5gF/mxRdxe/geZ0QG3xlFWnX+1ruKNhcrKxgWyQk/O2zjGcvYJeBmhVS+FBt+t qxnmCVJs+PEJBvU5bUnouAEnsGkBA1BeGasZdg9uI122tWdcUWd0LHdMqRRXxpKry3p4U+WyQ5t SUl2A01JxgAXDJOtRA4YFp9xZs2YNkkPt1FayhGwFHA4K7AV9UvJ0euaz8r X-Received: by 2002:a05:620a:3192:b0:936:ea3d:9313 with SMTP id af79cd13be357-9396ef7c45bmr114112385a.20.1788378426620; Wed, 02 Sep 2026 12:47:06 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-9395f18801asm299300585a.16.2026.09.02.12.47.05 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 02 Sep 2026 12:47:06 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, pbonzini@redhat.com, seanjc@google.com, akpm@linux-foundation.org, david@kernel.org, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, shuah@kernel.org Subject: [PATCH 4/5] selftests: KVM: guest_memfd: let the gmem_test() harness bind a node Date: Wed, 2 Sep 2026 15:46:56 -0400 Message-ID: <20260902194657.79075-5-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260902194657.79075-1-gourry@gourry.net> References: <20260902194657.79075-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" To test GUEST_MEMFD_FLAG_BIND_NODE - extend create_guest_memfd() to accept a node. Ignore the node argument if the BIND_NODE flag is not provided. Signed-off-by: Gregory Price Assisted-by: Claude:claude-opus-5 --- .../testing/selftests/kvm/guest_memfd_test.c | 50 ++++++++++++++++--- 1 file changed, 42 insertions(+), 8 deletions(-) diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing= /selftests/kvm/guest_memfd_test.c index 2233d871a38f4..1818e0fea5690 100644 --- a/tools/testing/selftests/kvm/guest_memfd_test.c +++ b/tools/testing/selftests/kvm/guest_memfd_test.c @@ -25,6 +25,31 @@ =20 static size_t page_size; =20 +static int __create_guest_memfd_node(struct kvm_vm *vm, u64 size, u64 flag= s, + u32 node, u32 pad) +{ + struct kvm_create_guest_memfd guest_memfd =3D { + .size =3D size, + .flags =3D flags, + .node =3D node, + .pad =3D pad, + }; + + return __vm_ioctl(vm, KVM_CREATE_GUEST_MEMFD, &guest_memfd); +} + +static int create_guest_memfd(struct kvm_vm *vm, u64 size, u64 flags, u32 = node) +{ + int fd; + + if (!(flags & GUEST_MEMFD_FLAG_BIND_NODE)) + return vm_create_guest_memfd(vm, size, flags); + + fd =3D __create_guest_memfd_node(vm, size, flags, node, 0); + TEST_ASSERT(fd >=3D 0, KVM_IOCTL_ERROR(KVM_CREATE_GUEST_MEMFD, fd)); + return fd; +} + static void test_file_read_write(int fd, size_t total_size) { char buf[64]; @@ -418,26 +443,35 @@ static void test_guest_memfd_flags(struct kvm_vm *vm) } } =20 -#define ____gmem_test(__test, __vm, __flags, __gmem_size, args...) \ -do { \ - int fd =3D vm_create_guest_memfd(__vm, __gmem_size, __flags); \ - \ - test_##__test(args); \ - close(fd); \ +#define ____gmem_test(__test, __vm, __flags, __gmem_size, __node, args...)= \ +do { \ + int fd =3D create_guest_memfd(__vm, __gmem_size, __flags, __node); \ + \ + test_##__test(args); \ + close(fd); \ } while (0) =20 #define __gmem_test(__test, __vm, __flags, __gmem_size) \ - ____gmem_test(__test, __vm, __flags, __gmem_size, fd, __gmem_size) + ____gmem_test(__test, __vm, __flags, __gmem_size, 0, fd, __gmem_size) =20 #define gmem_test(__test, __vm, __flags) \ __gmem_test(__test, __vm, __flags, page_size * 4) =20 #define __gmem_test_vm(__test, __vm, __flags, __gmem_size) \ - ____gmem_test(__test, __vm, __flags, __gmem_size, __vm, fd, __gmem_size) + ____gmem_test(__test, __vm, __flags, __gmem_size, 0, \ + __vm, fd, __gmem_size) =20 #define gmem_test_vm(__test, __vm, __flags) \ __gmem_test_vm(__test, __vm, __flags, page_size * 4) =20 +#define __gmem_test_node(__test, __vm, __flags, __gmem_size, __node) \ + ____gmem_test(__test, __vm, \ + (__flags) | GUEST_MEMFD_FLAG_BIND_NODE, \ + __gmem_size, __node, fd, __gmem_size, __node) + +#define gmem_test_node(__test, __vm, __flags, __node) \ + __gmem_test_node(__test, __vm, __flags, page_size * 4, __node) + static void __test_guest_memfd(struct kvm_vm *vm, u64 flags) { test_create_guest_memfd_multiple(vm); --=20 2.53.0-Meta From nobody Sat Sep 26 09:19:24 2026 Received: from mail-qk1-f182.google.com (mail-qk1-f182.google.com [209.85.222.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4FC933B0AF3 for ; Wed, 2 Sep 2026 19:47:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.182 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378446; cv=none; b=Vs2qTroZIRLTCr2G6afWttf0vugu6zABroB6yvW97DMtXo3biCmnX/f8LU7WP1nybHCH+X+MG0Qd5zpfWuyu1v66oour2xrMcoB4/UIg63i2AMNQZtJxiG4iBLJhjM3ts33zZ5ApfuhVCmXqTgzhVGeyuZ39W5PkkX2n1HrqMbQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788378446; c=relaxed/simple; bh=l/O097XS3NY2FSXC1MdvNZgKvqwkv0yyxmFMASNwv2U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ihjj8ru9zN/FgNh4fRBBr0Um1lasyk7uJiDreqGoPZFJ2YQg0/7Sg2wF04EaM/poFhR1Gt34sMpDhz4uu+KJe/wtSdvsszXEJ6PCfgwIlLYiKUP57JTA7pEVis13wK/DYmFml6gJq4ZWtEWiVdIAWYaIvgKgkC7hcphlPVl8enc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net; spf=pass smtp.mailfrom=gourry.net; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b=Q685jDrV; arc=none smtp.client-ip=209.85.222.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=gourry.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gourry.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gourry.net header.i=@gourry.net header.b="Q685jDrV" Received: by mail-qk1-f182.google.com with SMTP id af79cd13be357-92e50a650a0so148014385a.1 for ; Wed, 02 Sep 2026 12:47:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gourry.net; s=google; t=1788378433; x=1788983233; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Z0iWDp9wFRQQy4gDc5F5inYbcQx9NkSfWa5yO100B34=; b=Q685jDrVXFW/FTii7/GmvC4KwUTsKsJ9jthixJ92NX4rX4X5a4GhnDTQgxrY+RCb1A 4STa6qsKww4weVCTLbYccGP2XfhKNI/WwmThhq385Ik0sC/cLWeTwrpt8rKNYdc9+4UT n3N2pHp7aQ5nFGrdGp3Ex0p+iu1I1LLhsPwsWuefEooUzZ8rZYgbPowvlvTFAjhhSxm8 kZH23OuVbW2Vv9oIriDPeeA/0Hvga6JvQuahod/FTYm8BEbVEKz/AvbA4Rp4iArtjWHJ RnejA+P94LL3xFsK/8YyHGuuNpZbWwPC6bCoJ2LWc63McQph8eJWNW/BOPTyRZvcYDj2 Cuzg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788378433; x=1788983233; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=Z0iWDp9wFRQQy4gDc5F5inYbcQx9NkSfWa5yO100B34=; b=UmOacIBrMPpHnbQ5217x5dgPca9zLvGxsy3HM8shBytaVtymAzPMwrpBWojPqaunwD EktSUCUAGOCH+grU+h80gWJ+rCnfT2J5BdN+WeG2pMYPfuxvKFzFLtB+bF8JyLVvYiji j3gFbYJREWXlUjkmq1JJ1M03Sge12asptdiqZy1uZTeNV/EnWOlf1UWcWL/CbK2tsotA Gu7fME/KoRSoUEb2zuBfMoVraBNnAzHQfvdu84rqzwlsB7HteuMJ8XwDOQpHWVtwRi9Z 4k/G4S8fvNwSnEfho/MLaLQLFCuW5OP+hF9L1vUzQVz/8n9pXBwQEvcUNMLak5snFXxV 5tCQ== X-Forwarded-Encrypted: i=1; AKwUvBwjXMHk7Oha1PI+9oem8UZ0XNPeLCb8IsOzd+ntQ5FmpC4JuKopXJfQ0V8mjiyNHBWV9JyEheN8td7cNhU=@vger.kernel.org X-Gm-Message-State: AFuF++kNvsHC2kapEQUxJ80aBsLOrac3MwmjI3NRak9mIa/PuHFHETnd +MAx/h84qOTc1nWb9fpckq5Z6aBAJ3jI4XWWwIhGshYL/oUkSEGF3U0WWPeYenI3B1c= X-Gm-Gg: AYBFou0zzMEDtbwhKAzMOBJp6esLMPF9ytroSkJmelO58ygIeuaK/5h2efXSv+BylCy rBKMKDruqLCqwhvB9utylAYOjAMxnMBe3mYIMEzSvaqmMd8EgqmaiTMrkGmxnbdSI/eXswCPpiw 3PfxvZ/JeTOBeOJjy7EEyCAGIFury+oqMZdmiWIlUU5/KJd1oxrmepm9OZ27DrN4sEHgcjXHzhQ KEQf/r8b0j70cfrXgRoVcv2b4cXyC7At/NJWx14RdwBsLj6rf/m1uE7/Dmp+2hgYTa5x4axekyf AJQ+pf4YKUADUo0nqlgKIXz2D1Dl5FWyApn+PCka9oV6ecp3fDVNu0ZlDFciSZ17piDVGr7HSgu od7LzmQvqK+Mzj8XCu+5iqXNn99lP3WDThemAWa0EAeCqO8aRHmRo99MjBbydgf2Kd0vgfC42xX Lh5DchVQgW65X6ntphNEW5utF0fnqzb0LIT9um4TReIfmD4E9qxmRuzA+mJFPsPOvPI4tbOvFjn +Qso+O5+SCZoNf76RqpbnZjBy5W85+hWDPRlj03iKLfNrAaiw== X-Received: by 2002:a05:620a:7002:b0:933:aa0:bb83 with SMTP id af79cd13be357-93960f69d21mr852874685a.34.1788378428558; Wed, 02 Sep 2026 12:47:08 -0700 (PDT) Received: from gourry-fedora-PF4VCD3F.lan (pool-173-79-60-52.washdc.fios.verizon.net. [173.79.60.52]) by smtp.gmail.com with ESMTPSA id af79cd13be357-9395f18801asm299300585a.16.2026.09.02.12.47.07 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 02 Sep 2026 12:47:07 -0700 (PDT) From: Gregory Price To: linux-mm@kvack.org Cc: kvm@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, pbonzini@redhat.com, seanjc@google.com, akpm@linux-foundation.org, david@kernel.org, ziy@nvidia.com, matthew.brost@intel.com, joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, shuah@kernel.org Subject: [PATCH 5/5] selftests: KVM: guest_memfd: test GUEST_MEMFD_FLAG_BIND_NODE Date: Wed, 2 Sep 2026 15:46:57 -0400 Message-ID: <20260902194657.79075-6-gourry@gourry.net> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260902194657.79075-1-gourry@gourry.net> References: <20260902194657.79075-1-gourry@gourry.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add tests for an umapped guest_memfd mempolicy configured at creation via GUEST_MEMFD_FLAG_BIND_NODE. Handles: - !CONFIG_NUMA (skips) - single node system - multi-node system (faults onto remote node) - invalid arguments (non-zero pad, bad node, node w/o flag). BIND_NODE is skipped in test_guest_memfd_flags() and testing later because that loop asserts every advertised flag succeeds on its own, but BIND_NODE is the only flag that depends on a second field. Signed-off-by: Gregory Price Assisted-by: Claude:claude-opus-5 --- .../testing/selftests/kvm/guest_memfd_test.c | 84 +++++++++++++++++++ 1 file changed, 84 insertions(+) diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing= /selftests/kvm/guest_memfd_test.c index 1818e0fea5690..b333cb42fab29 100644 --- a/tools/testing/selftests/kvm/guest_memfd_test.c +++ b/tools/testing/selftests/kvm/guest_memfd_test.c @@ -196,6 +196,83 @@ static void test_numa_allocation(int fd, size_t total_= size) kvm_munmap(mem, total_size); } =20 +static bool has_bind_node(struct kvm_vm *vm) +{ + return vm_check_cap(vm, KVM_CAP_GUEST_MEMFD_FLAGS) & + GUEST_MEMFD_FLAG_BIND_NODE; +} + +static void test_bind_node_invalid(struct kvm_vm *vm, u64 flags) +{ + int fd; + + if (!has_bind_node(vm)) + return; + + fd =3D __create_guest_memfd_node(vm, page_size, + flags | GUEST_MEMFD_FLAG_BIND_NODE, 0, 1); + TEST_ASSERT(fd < 0 && errno =3D=3D EINVAL, + "guest_memfd() with non-zero pad should fail with EINVAL"); + + fd =3D __create_guest_memfd_node(vm, page_size, + flags | GUEST_MEMFD_FLAG_BIND_NODE, + 1 << 20, 0); + TEST_ASSERT(fd < 0 && errno =3D=3D EINVAL, + "guest_memfd() with out-of-range node should fail with EINVAL"); + + fd =3D __create_guest_memfd_node(vm, page_size, flags, 1, 0); + TEST_ASSERT(fd < 0 && errno =3D=3D EINVAL, + "guest_memfd() with a node but no BIND_NODE flag should fail with EI= NVAL"); +} + +static void test_bind_node(int fd, size_t total_size, int node) +{ + const unsigned long other_mask =3D 1UL << (node ? 0 : 1); + const unsigned long maxnode =3D BITS_PER_TYPE(other_mask); + bool steer_away =3D is_multi_numa_node_system(); + void *pages[4]; + int status[4]; + char *mem; + int i; + + mem =3D kvm_mmap(total_size, PROT_READ | PROT_WRITE, MAP_SHARED, fd); + for (i =3D 0; i < 4; i++) + pages[i] =3D mem + page_size * i; + + /* + * Bind on a different node if possible order to check whether faulting + * happens as desired. Without a second node use the local node and + * just get coverage of create/mmap/fault paths. + */ + if (steer_away) + kvm_set_mempolicy(MPOL_BIND, &other_mask, maxnode); + + /* Deliberately no mbind() on this mapping. */ + memset(mem, 0xaa, total_size); + + kvm_move_pages(0, 4, pages, NULL, status, 0); + for (i =3D 0; i < 4; i++) + TEST_ASSERT(status[i] =3D=3D node, + "Expected page %d on node %d, got it on node %d", + i, node, status[i]); + + /* Dropped memory should fault back onto the same node */ + kvm_fallocate(fd, FALLOC_FL_PUNCH_HOLE | FALLOC_FL_KEEP_SIZE, 0, + total_size); + memset(mem, 0xaa, total_size); + + kvm_move_pages(0, 4, pages, NULL, status, 0); + for (i =3D 0; i < 4; i++) + TEST_ASSERT(status[i] =3D=3D node, + "Expected page %d back on node %d, got it on node %d", + i, node, status[i]); + + if (steer_away) + kvm_set_mempolicy(MPOL_DEFAULT, NULL, 0); + + kvm_munmap(mem, total_size); +} + static void test_collapse(int fd, u64 flags) { const size_t pmd_size =3D get_trans_hugepagesz(); @@ -429,6 +506,10 @@ static void test_guest_memfd_flags(struct kvm_vm *vm) int fd; =20 for (flag =3D BIT(0); flag; flag <<=3D 1) { + /* BIND_NODE depends on a valid node field, test separately */ + if (flag =3D=3D GUEST_MEMFD_FLAG_BIND_NODE) + continue; + fd =3D __vm_create_guest_memfd(vm, page_size, flag); if (flag & valid_flags) { TEST_ASSERT(fd >=3D 0, @@ -476,6 +557,7 @@ static void __test_guest_memfd(struct kvm_vm *vm, u64 f= lags) { test_create_guest_memfd_multiple(vm); test_create_guest_memfd_invalid_sizes(vm, flags); + test_bind_node_invalid(vm, flags); =20 gmem_test(file_read_write, vm, flags); =20 @@ -486,6 +568,8 @@ static void __test_guest_memfd(struct kvm_vm *vm, u64 f= lags) gmem_test(mmap_supported, vm, flags); gmem_test(fault_overflow, vm, flags); gmem_test(numa_allocation, vm, flags); + if (has_bind_node(vm)) + gmem_test_node(bind_node, vm, flags, 0); __gmem_test(collapse, vm, flags, pmd_size); } else { gmem_test(fault_private, vm, flags); --=20 2.53.0-Meta