From nobody Thu Sep 24 12:10:22 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 22E7A3955CB for ; Tue, 4 Aug 2026 12:05:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845141; cv=none; b=sG4TZUDjd3Mc3am1ROnJBWDspG9bu14ZJV6ZBJdldLStyLPI90yRlbCw4MgpE5ST5PKwI6i05SunJq3hjEayK9Qz3iU1c3Vc/y+VYM0xyQ0EfU5MtuqUkcxKtUevlCyiYY+poizbA+Cd1tT6aDdB6i8yrfeRVNH56N97Wr3DtBA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845141; c=relaxed/simple; bh=2NmEb6/L7QYcfhxZ8+lab83yjq9i9GZownbXwTIqVvs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uX8/WjIWEwVJphKflYHXLnsRINKI2QKDeCpAL4zFEJF21WSvtBnwvvQbw6aUOpEoZ4OqiDypGNY6wLr6EpD9QNUDFrGWyCPhf/RUx3EuhU/Nu6aEHfiGxU4nklCgyi6raQctq0a5fU7vErXxE+WV2XmJEoDpFkDDmkq31XZn04U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=UbW8QkjM; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=VyG6mxr6; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="UbW8QkjM"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="VyG6mxr6" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785845138; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=VnnVy48ydCXauSW+w5ZraHIy4YMvaYkPwdBxROmLWeM=; b=UbW8QkjMHh6y2r3PT7FYazUcub8AkJcFBNTUZlhd+KXxIS31AmUQnGKR8NWe+cAVgqdx3H TpfpBuHm3bhmp08uh31f+9k+WqU8tH6u9g64TPlgXvE7WXti9SiqWC470C6ZigqpL/MQ3W Z6Y9VOHbWDvqo1n2g4DKyc+3HphJmyY= Received: from mail-wm1-f70.google.com (mail-wm1-f70.google.com [209.85.128.70]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-57--Mam4piiPAGzqlKmq-QaVA-1; Tue, 04 Aug 2026 08:05:37 -0400 X-MC-Unique: -Mam4piiPAGzqlKmq-QaVA-1 X-Mimecast-MFC-AGG-ID: -Mam4piiPAGzqlKmq-QaVA_1785845136 Received: by mail-wm1-f70.google.com with SMTP id 5b1f17b1804b1-4955843c6cdso29383475e9.1 for ; Tue, 04 Aug 2026 05:05:36 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1785845136; x=1786449936; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=VnnVy48ydCXauSW+w5ZraHIy4YMvaYkPwdBxROmLWeM=; b=VyG6mxr6G93O8udUC6Utp604TT4xcBQs5xFqo3/6/bApCGY6l9S9sskeOtsh+yynaY NtchMC0CZpuiCZsUCfcGnOHAIRmTY731wrS/z508prdMgDtQabDOoBFMddSN1qDP5S0k hJ6l/KWGOOc4FAWS9NlPnvxmONoF4ppoe7kEQE/onXsM+ibXEmUvmXFY/28rf47l2mvV JNZqrmKJY8aq8znRuHqBLQWKuF6Zk74YLTwajnz5F0RhIVCGlEfeifue/wbJ2HQDlfpL 9w5xTHIX3foRNnDqR9LrvojV0ZwoydsBgZg5zrirSCw7KYk7oO6/2XM7gydhyL5qB9Zw zQOQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785845136; x=1786449936; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=VnnVy48ydCXauSW+w5ZraHIy4YMvaYkPwdBxROmLWeM=; b=eBiqDxv3RFL86K76FZ4jKuxaCIdR/EZ1uiJ+jitzn4ZfcV/6v/WAmEDMpRpl4SgMN6 5EFJh6uGVAGF0Ck/OOv35FhJ8CPUnfExuCCXmky2pU/IQXzwv98vsui9eK4pOQSqJbD0 q4FQY9FC9Z0piy7svWVQiGHY3VbcVSyhqN1W51Lj9FHK1OMgbSth85rw64z/fR1Y/TCv kX+4kXgyogiADpVOFUv3VB4p3bGWSY8o4YLwYoEukqmfNTuMu5eWqXz5mzLv7hzJUatC IGRJzLJU458hjHMvbNlAv5MYOvF65IOyruSfcLaUxXerDoTJWBFKGeTw1XWA0GOhaOv3 +rQA== X-Gm-Message-State: AOJu0YwT2NCF1dkmJerehRtO4S2TKr16k6lXBKWoFLZJ0yKgcGqFANBH b5M/Uqk9hSEbCFfXMlqw2MlNDp0O7CVkSbYhUSbURrEZGTDbcRyBqPeKGrIfOU+dr2Vdumd7xHs 2ZoP7GeQIbDXAJer8RE0m0E2iEDBlfMc6/U2RPxN8rb95cUV3v47QKwynY5Rb3GqN7PzTsx/2tU Mtd3092YWSaBKeHbr7v59J3uWgEoEE9r4fmoWOSNRYcstHp/EotQ== X-Gm-Gg: AR+sD11cRJpaHmmkihZLJbRfbpWMpjVjlL8swFOiqAu2H2ztwyE+0yrxJuZTk9FkEEW wGPvXOIiqFVoKDm9YKRDNO+F2Ojv4e8YXBq2dAm5CVbiF79dQ0MZ0UQf6NBDVL9FLkJ08Ku0Ltf //DJ46sPxZhHIuZu+49LdlAhD4jGjEa6TLt85IlUhPKPLnvfhxpBCBk60idIleHOLNRobERGH7V rs19jfvo5XGllwS3X467uNGk6+YIIsx0oalFeVFJLz/tqaeC5kKgyb6ehxan1z1siAisX9YsqwG lsR2A5OabSKb0QU2RzSV++d6r24STGp+PEFng+fr0rOiuSiLwdNh2APIzSHTXl3lpaiY4H1KGzu yraVtf2lpnIIFcvLl/B3H3A+JY8B+Crvymu5YcoQr/Sy0I7y/GTUAC/7pPlmiOVr+RH0dHSLQGL /p9m4= X-Received: by 2002:a05:600c:3507:b0:496:cb48:5eb8 with SMTP id 5b1f17b1804b1-4980c6747fdmr217895395e9.15.1785845135349; Tue, 04 Aug 2026 05:05:35 -0700 (PDT) X-Received: by 2002:a05:600c:3507:b0:496:cb48:5eb8 with SMTP id 5b1f17b1804b1-4980c6747fdmr217893395e9.15.1785845134667; Tue, 04 Aug 2026 05:05:34 -0700 (PDT) Received: from [192.168.10.48] ([151.95.34.92]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49949f66168sm92311515e9.0.2026.08.04.05.05.33 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 04 Aug 2026 05:05:34 -0700 (PDT) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: Alex Williamson , bcm-kernel-feedback-list@broadcom.com, Boris Brezillon , Christian Koenig , David Hildenbrand , dri-devel@lists.freedesktop.org, Fei Li , Huang Rui , linux-mm@kvack.org, linux-s390@vger.kernel.org, Michal Hocko , Peter Xu , Sergio Lopez , Sean Christopherson , Thomas Zimmermann , stable@vger.kernel.org Subject: [PATCH v2 1/6] mm: export vmf_insert_pfn_prot_mkwrite(), change variants to inline Date: Tue, 4 Aug 2026 14:05:23 +0200 Message-ID: <20260804120529.1730187-2-pbonzini@redhat.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260804120529.1730187-1-pbonzini@redhat.com> References: <20260804120529.1730187-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Right now, users of .pfn_mkwrite() have no way to create a PTE that has gone through maybe_mkwrite(). Because vma_set_page_prot() will have cleared the writable PTE bit, users of fixup_user_fault() will see a read-only PTE and have no clue that the page needs a *second* fault to reach its final status. Handling this in fixup_user_fault() is problematic: the information about the presence of *_mkwrite is only recorded in vma->vm_page_prot, which is an opaque pgprot_t, therefore only follow_pfnmap_start() knows how to retrieve it. There are actually some preexisting functions that suggest how this is supposed to be handled, namely vmf_insert_page_mkwrite() and vmf_insert_pfn_pmd(). Fixing the drivers requires similar variants of vm_insert_pfn(), namely vmf_insert_pfn_mkwrite() for the common case where vma->vm_page_prot is okay, and vmf_insert_pfn_prot_mkwrite() when really all parameters are needed. This makes it possible to fix drivers that use .pfn_mkwrite together with vmf_insert_pfn() and vmf_insert_pfn_prot(). Since vmf_insert_pfn_prot_mkwrite() is the most general variant and all the others are just special cases, turn them into inline functions in the header. Fixes: 28e3918179aa ("drm/gem-shmem: Track folio accessed/dirty status in m= map") Cc: stable@vger.kernel.org Signed-off-by: Paolo Bonzini --- include/linux/mm.h | 81 +++++++++++++++++++++++++++++++++++++++++--- mm/huge_memory.c | 2 +- mm/memory.c | 84 ++++++++++++++++++++-------------------------- 3 files changed, 114 insertions(+), 53 deletions(-) diff --git a/include/linux/mm.h b/include/linux/mm.h index 485df9c2dbdd..01184a4bdd6f 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -4544,16 +4544,89 @@ int vm_map_pages_zero(struct vm_area_struct *vma, s= truct page **pages, unsigned long num); vm_fault_t vmf_insert_page_mkwrite(struct vm_fault *vmf, struct page *page, bool write); -vm_fault_t vmf_insert_pfn(struct vm_area_struct *vma, unsigned long addr, - unsigned long pfn); -vm_fault_t vmf_insert_pfn_prot(struct vm_area_struct *vma, unsigned long a= ddr, - unsigned long pfn, pgprot_t pgprot); +vm_fault_t vmf_insert_pfn_prot_mkwrite(struct vm_area_struct *vma, unsigne= d long addr, + unsigned long pfn, pgprot_t pgprot, bool mkwrite); vm_fault_t vmf_insert_mixed(struct vm_area_struct *vma, unsigned long addr, unsigned long pfn); vm_fault_t vmf_insert_mixed_mkwrite(struct vm_area_struct *vma, unsigned long addr, unsigned long pfn); int vm_iomap_memory(struct vm_area_struct *vma, phys_addr_t start, unsigne= d long len); =20 + +/** + * vmf_insert_pfn_prot - insert single pfn into user vma with specified pg= prot + * @vma: user vma to map to + * @addr: target user address of this page + * @pfn: source kernel pfn + * @pgprot: pgprot flags for the inserted page + * + * This is exactly like vmf_insert_pfn(), except that it allows drivers + * to override pgprot on a per-page basis. For more information, + * see vmf_insert_pfn_prot_mkwrite(). + * + * This only makes sense for IO mappings, and it makes no sense for + * COW mappings. In general, using multiple vmas is preferable; + * vmf_insert_pfn_prot should only be used if using multiple VMAs is + * impractical. + * + * Context: Process context. May allocate using %GFP_KERNEL. + * Return: vm_fault_t value. + */ +static inline vm_fault_t vmf_insert_pfn_prot(struct vm_area_struct *vma, + unsigned long addr, unsigned long pfn, pgprot_t pgprot) +{ + return vmf_insert_pfn_prot_mkwrite(vma, addr, pfn, pgprot, false); +} + +/** + * vmf_insert_pfn_mkwrite - insert single pfn into user vma, possibly writ= able + * @vma: user vma to map to + * @addr: target user address of this page + * @pfn: source kernel pfn + * @write: whether the PTE should be installed writable + * + * Like vmf_insert_pfn(), except that @write allows installing a writable + * PTE even when @vma is under write notification. For more information, + * see vmf_insert_pfn_prot_mkwrite(). + * + * Note that neither .pfn_mkwrite() nor .page_mkwrite() is invoked, so the + * caller must itself do whatever they would have done if @write is true. + * + * Context: Process context. May allocate using %GFP_KERNEL. + * Return: vm_fault_t value. + */ +static inline vm_fault_t vmf_insert_pfn_mkwrite(struct vm_area_struct *vma, + unsigned long addr, unsigned long pfn, bool write) +{ + return vmf_insert_pfn_prot_mkwrite(vma, addr, pfn, vma->vm_page_prot, wri= te); +} + +/** + * vmf_insert_pfn - insert single pfn into user vma + * @vma: user vma to map to + * @addr: target user address of this page + * @pfn: source kernel pfn + * + * Similar to vm_insert_page, this allows drivers to insert individual pag= es + * they've allocated into a user vma. Same comments apply. + * + * This function should only be called from a vm_ops->fault handler, and + * in that case the handler should return the result of this function. + * + * vma cannot be a COW mapping. + * + * As this is called only for pages that do not currently exist, we + * do not need to flush old virtual caches or the TLB. + * + * Context: Process context. May allocate using %GFP_KERNEL. + * Return: vm_fault_t value. + */ +static inline vm_fault_t vmf_insert_pfn(struct vm_area_struct *vma, + unsigned long addr, unsigned long pfn) +{ + return vmf_insert_pfn_mkwrite(vma, addr, pfn, false); +} + static inline vm_fault_t vmf_insert_page(struct vm_area_struct *vma, unsigned long addr, struct page *page) { diff --git a/mm/huge_memory.c b/mm/huge_memory.c index b5d1e9d4463d..2f4dcaa819b7 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -1615,7 +1615,7 @@ static vm_fault_t insert_pmd(struct vm_area_struct *v= ma, unsigned long addr, * @pfn: pfn to insert * @write: whether it's a write fault * - * Insert a pmd size pfn. See vmf_insert_pfn() for additional info. + * Insert a pmd size pfn. See vmf_insert_pfn_mkwrite() for additional info. * * Return: vm_fault_t value. */ diff --git a/mm/memory.c b/mm/memory.c index ff338c2abe92..b5555217b121 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -2719,40 +2719,55 @@ static vm_fault_t insert_pfn(struct vm_area_struct = *vma, unsigned long addr, } =20 /** - * vmf_insert_pfn_prot - insert single pfn into user vma with specified pg= prot + * vmf_insert_pfn_prot_mkwrite - insert single pfn into user vma with spec= ified pgprot * @vma: user vma to map to * @addr: target user address of this page * @pfn: source kernel pfn * @pgprot: pgprot flags for the inserted page + * @mkwrite: whether to make the page writable. * - * This is exactly like vmf_insert_pfn(), except that it allows drivers - * to override pgprot on a per-page basis. + * This is the function underlying all the others in the vmf_insert_pfn() + * family. It is the most flexible, as it allows drivers to override pgpr= ot + * on a per-page basis, as well as to insert the pfn as if it already had + * a write fault. vmf_insert_pfn() is usually sufficient, however. + * + * These functions should only be called from a vm_ops->fault handler, and + * in that case the handler should return the result of these functions. * * This only makes sense for IO mappings, and it makes no sense for - * COW mappings. In general, using multiple vmas is preferable; - * vmf_insert_pfn_prot should only be used if using multiple VMAs is - * impractical. + * COW mappings. * - * pgprot typically only differs from @vma->vm_page_prot when drivers set - * caching- and encryption bits different than those of @vma->vm_page_prot, - * because the caching- or encryption mode may not be known at mmap() time. + * For vmf_insert_pfn_prot_mkwrite() and vmf_insert_pfn_mkwrite(), the + * @mkwrite argument allows installing a writable PTE even when @vma is + * under write notification, i.e. when it has a .pfn_mkwrite() callback. + * In this case, vma_set_page_prot() has cleared the write bit from + * @vma->vm_page_prot. This lets the fault() callback install a writable + * PTE in response to write faults; note that .pfn_mkwrite() is not called, + * and therefore the caller has to do by itself whatever the callback would + * have done. * - * This is ok as long as @vma->vm_page_prot is not used by the core vm + * For vmf_insert_pfn_prot_mkwrite() and vmf_insert_pfn_prot(), + * pgprot can differ from @vma->vm_page_prot. This typically happens only + * for caching and encryption bits, which may not be known at mmap() time; + * it is ok as long as @vma->vm_page_prot is not used by the core vm * to set caching and encryption bits for those vmas (except for COW pages= ). - * This is ensured by core vm only modifying these page table entries using - * functions that don't touch caching- or encryption bits, using pte_modif= y() - * if needed. (See for example mprotect()). + * This is ensured in two ways: * - * Also when new page-table entries are created, this is only done using t= he - * fault() callback, and never using the value of vma->vm_page_prot, - * except for page-table entries that point to anonymous pages as the resu= lt - * of COW. + * - core vm only modifies these page table entries using functions that d= on't + * touch caching- or encryption bits, using pte_modify() if needed. (See + * for example mprotect()). + * + * - when new page-table entries are created, this is only done using the + * fault() callback, and never using the value of vma->vm_page_prot, + * except for page-table entries that point to anonymous pages as the re= sult + * of COW. * * Context: Process context. May allocate using %GFP_KERNEL. * Return: vm_fault_t value. */ -vm_fault_t vmf_insert_pfn_prot(struct vm_area_struct *vma, unsigned long a= ddr, - unsigned long pfn, pgprot_t pgprot) +vm_fault_t vmf_insert_pfn_prot_mkwrite(struct vm_area_struct *vma, + unsigned long addr, unsigned long pfn, pgprot_t pgprot, + bool mkwrite) { /* * Technically, architectures with pte_special can avoid all these @@ -2774,36 +2789,9 @@ vm_fault_t vmf_insert_pfn_prot(struct vm_area_struct= *vma, unsigned long addr, =20 pfnmap_setup_cachemode_pfn(pfn, &pgprot); =20 - return insert_pfn(vma, addr, pfn, pgprot, false); + return insert_pfn(vma, addr, pfn, pgprot, mkwrite); } -EXPORT_SYMBOL(vmf_insert_pfn_prot); - -/** - * vmf_insert_pfn - insert single pfn into user vma - * @vma: user vma to map to - * @addr: target user address of this page - * @pfn: source kernel pfn - * - * Similar to vm_insert_page, this allows drivers to insert individual pag= es - * they've allocated into a user vma. Same comments apply. - * - * This function should only be called from a vm_ops->fault handler, and - * in that case the handler should return the result of this function. - * - * vma cannot be a COW mapping. - * - * As this is called only for pages that do not currently exist, we - * do not need to flush old virtual caches or the TLB. - * - * Context: Process context. May allocate using %GFP_KERNEL. - * Return: vm_fault_t value. - */ -vm_fault_t vmf_insert_pfn(struct vm_area_struct *vma, unsigned long addr, - unsigned long pfn) -{ - return vmf_insert_pfn_prot(vma, addr, pfn, vma->vm_page_prot); -} -EXPORT_SYMBOL(vmf_insert_pfn); +EXPORT_SYMBOL(vmf_insert_pfn_prot_mkwrite); =20 static bool vm_mixed_ok(struct vm_area_struct *vma, unsigned long pfn, bool mkwrite) --=20 2.55.0 From nobody Thu Sep 24 12:10:22 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 26F13394EA7 for ; Tue, 4 Aug 2026 12:05:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845143; cv=none; b=A1wa58xV5Wk12tJ+R1cMOJzWZQuTy/38B9dEjoPms7BL/FKU9erF+f6XccWoInucoE31omaJff9QtI4QNLS69Jf4vc2tWpaK8zqcOxqZsOrkB5+P6kec8OW+tLHdJjjDycMmnC96aU4vujsiNN1p90KYr4hkO/CoIvXZ9z/1taM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845143; c=relaxed/simple; bh=tWhbpLrndV3dvFj6rw3wUEqAvQhUNv7Z7XChRUIb0M0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Pkel2Io61bFG0FWqCvF3VWeOSzebO0ZHGpHx+oefbjDcTq8/DnpbMubuND+bdYSVgqXcLMgWqALuSWNu09m+nWb0bPw8QnYyrYBEiZNFaEy5pnhMn+NSTYd2WPu6bLPTXKMn6qjYiOk88nbgYap8ZxNcI9hHxVYckkmo/1wgnCM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=ZHLPLsk4; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=uDdqbUZv; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="ZHLPLsk4"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="uDdqbUZv" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785845140; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=taQNLA8VzPi8pIAuKLKHCaHnQR8Qtlo0OLpkQO4wnlM=; b=ZHLPLsk4NNkZ1MUg20CMCvFb47FhYoquPXJtd1knMZSb7GQRhdBQ2RFYyEhz3TN9fh6zrg SWvBRUNTHZHZRVz1LLnFHlwebUWD+GpyMtnfOamxDBHuNpyf0G7h5VK6vNOPrGFKUCjNM9 0p1pP3CbfFBSr2B1kmyPZgycKM/+mFw= Received: from mail-wm1-f70.google.com (mail-wm1-f70.google.com [209.85.128.70]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-206-pvo8xRoYMOWYnd1gEXtAwQ-1; Tue, 04 Aug 2026 08:05:39 -0400 X-MC-Unique: pvo8xRoYMOWYnd1gEXtAwQ-1 X-Mimecast-MFC-AGG-ID: pvo8xRoYMOWYnd1gEXtAwQ_1785845138 Received: by mail-wm1-f70.google.com with SMTP id 5b1f17b1804b1-49545071724so32917315e9.0 for ; Tue, 04 Aug 2026 05:05:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1785845138; x=1786449938; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=taQNLA8VzPi8pIAuKLKHCaHnQR8Qtlo0OLpkQO4wnlM=; b=uDdqbUZvYP79RqimY1ZqTEuTXAixIlW/QtJ7a4KlNLDotqiHC1jC+BqaNbdrzoB8iP cU+AHAMYM89wwmVfsuSmTWMfB2/Fx34ztmYWncP0qP1zK32pIaOH9F/xITa1sfSTE/gu ndWjrHWkUfTLCpWYWdkXxS+Nd1po8h50lTtsj9o2CyLeCsSM6FYnnl1uBQEdAg6nPNhI lblviG/InhzcdCosHgo3agQfYY5c/+oKxQPbzW7EO8dIquZ8LX6maNtNVJ9hDC7FY49t 5I0I1tJH9MkIOd8zQNWwNqDi/ouIyyGx551nNd9HdphXflkxVPFuaJ7HjIH+8RH5EoxQ 10Dg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785845138; x=1786449938; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=taQNLA8VzPi8pIAuKLKHCaHnQR8Qtlo0OLpkQO4wnlM=; b=h/dbRotH6mSxQkbb7jj7mCAW8Pn0F3RwLkrotqld496bczXClpMJx/CCvSHtiX8A1l lHm/NY6g/zvNqY9RAlJow6mLmM+53zrvyf4xScb0D6ltN56T8kf+3qHzkSalVkZPldK4 zk3W9eMaek8v4DoEarxo1M36nQonc82k6nxIpzc8AI5y91tDL+AAG3yiIP5OyCgkxdkD Lo6RXVSux6YwQFr8FmOjujuGHfAzUd4JF96AD+NG1OFQN4IlC7eRDIzxxY2k3uoUn5hb K8p/ig0ChVbEwCpEr6jle9X9zU2FCNropU8yJkIJxzCTDMQyDrxIkRW0uES3lcmJZR9a lj4A== X-Gm-Message-State: AOJu0YzN+uXSkzm4yUrDNjlIvTBcWdnFXK+16Kk+qcmAddWjmI+YIatb BMnfia0VqIJ8ugzp80SYPM6K/EJERnSGT/DE8T8JJuaUxvFR7upqbYlDkSAdyvfR1RVoJn9nZWN tfHg6HVbwRljWgZvT+xlOb55kdD1mlpmI+22MUheLtNv4THgK73rYuQrdrQn+K7fs4Y7mEqgWP7 azR7RFGoEpBu7LrW09dcR4LaB0RhekGesEmGFSdNFkBHOf5DmZWw== X-Gm-Gg: AR+sD11gEwRbc7EsXREZLIrRpNPq1jZZPnKUTvvBonbhRokFJpDN+IbTGkiaeKe9cUa pf70oUKKa1ABZ1qKUuyRwRa+cd2RrlKSa5lARdIiCYRxs4dAwU8o/tDDCjqpdAYv31KgmGlB7hV gwR/dJCKU2Onff5l9/TXteS8nNXaKOJn0uVy+taFaya/cVq+ZAv+5AEFdh1qADMtJjuk7tOEtFz GlnZiB2zhwNIm2ZsJ1FOgR/kOfg5yYaFK29FKm/ZY6WG4bXM99XdFHhDWDLDCx2CB/EtxWO/+oN LweuqctR/9G1TQy7OKFUQkEFU5B0UcVDa3lWR8lN1ChTSuAMP3aPdma0HzJ2xlrYBR/aCM8kKUK 6UBilhnaYU9phGQsN7dXEp8s1qAj4dbRYujFFSWnBgigCXWJufmolveO2yek8bzZ4ZPYTfFMRp1 IsLlA= X-Received: by 2002:a05:600c:6298:b0:499:484a:7644 with SMTP id 5b1f17b1804b1-499484a7693mr122467415e9.9.1785845137757; Tue, 04 Aug 2026 05:05:37 -0700 (PDT) X-Received: by 2002:a05:600c:6298:b0:499:484a:7644 with SMTP id 5b1f17b1804b1-499484a7693mr122465915e9.9.1785845136990; Tue, 04 Aug 2026 05:05:36 -0700 (PDT) Received: from [192.168.10.48] ([151.95.34.92]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49807b8d04fsm453070475e9.3.2026.08.04.05.05.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 04 Aug 2026 05:05:36 -0700 (PDT) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: Alex Williamson , bcm-kernel-feedback-list@broadcom.com, Boris Brezillon , Christian Koenig , David Hildenbrand , dri-devel@lists.freedesktop.org, Fei Li , Huang Rui , linux-mm@kvack.org, linux-s390@vger.kernel.org, Michal Hocko , Peter Xu , Sergio Lopez , Sean Christopherson , Thomas Zimmermann , stable@vger.kernel.org Subject: [PATCH v2 2/6] drm/shmem_helper: use vmf_insert_pfn_mkwrite() Date: Tue, 4 Aug 2026 14:05:24 +0200 Message-ID: <20260804120529.1730187-3-pbonzini@redhat.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260804120529.1730187-1-pbonzini@redhat.com> References: <20260804120529.1730187-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" This ensures that KVM or VFIO correctly see a writable PTE when they request one. Otherwise, a guest write to an unpopulated PTE from a mapping backed by a DRM GEM BO triggers a VM exit with EFAULT. The code actually is simpler, because the same logic already applied to the hugepage mapping case using vmf_insert_pfn_pmd(). Reported-by: Sergio Lopez Link: https://lore.kernel.org/kvm/20260729072044.25796-1-slp@redhat.com/ Tested-by: Sergio Lopez Reviewed-by: Boris Brezillon Fixes: 28e3918179aa ("drm/gem-shmem: Track folio accessed/dirty status in m= map") Cc: stable@vger.kernel.org Signed-off-by: Paolo Bonzini --- drivers/gpu/drm/drm_gem_shmem_helper.c | 38 ++++++++++++++------------ 1 file changed, 20 insertions(+), 18 deletions(-) diff --git a/drivers/gpu/drm/drm_gem_shmem_helper.c b/drivers/gpu/drm/drm_g= em_shmem_helper.c index c989459eb215..c81be3e97317 100644 --- a/drivers/gpu/drm/drm_gem_shmem_helper.c +++ b/drivers/gpu/drm/drm_gem_shmem_helper.c @@ -589,11 +589,25 @@ static void drm_gem_shmem_record_mkwrite(struct vm_fa= ult *vmf) folio_mark_dirty(page_folio(shmem->pages[page_offset])); } =20 +/* + * Because the vm_ops have a .pfn_mkwrite() callback, vma_set_page_prot() + * has cleared the write bit from vma->vm_page_prot. vmf_insert_pfn() + * would install a read-only entry even for a write fault, relying on a + * second fault to reach .pfn_mkwrite() and upgrade it, but that second + * fault never happens for fixup_user_fault() callers that directly + * walk the page tables with follow_pfnmap_start(). To ensure that + * they don't see the read-only entry, pass FAULT_FLAG_WRITE info down + * to install a writable entry right away. Because .pfn_mkwrite() is + * not invoked, record the write afterwards. + */ static vm_fault_t try_insert_pfn(struct vm_fault *vmf, unsigned int order, unsigned long pfn) { + bool write =3D vmf->flags & FAULT_FLAG_WRITE; + vm_fault_t ret =3D VM_FAULT_FALLBACK; + if (!order) { - return vmf_insert_pfn(vmf->vma, vmf->address, pfn); + ret =3D vmf_insert_pfn_mkwrite(vmf->vma, vmf->address, pfn, write); #ifdef CONFIG_ARCH_SUPPORTS_PMD_PFNMAP } else if (order =3D=3D PMD_ORDER) { unsigned long paddr =3D pfn << PAGE_SHIFT; @@ -601,27 +615,15 @@ static vm_fault_t try_insert_pfn(struct vm_fault *vmf= , unsigned int order, =20 if (aligned && folio_test_pmd_mappable(page_folio(pfn_to_page(pfn)))) { - vm_fault_t ret; - pfn &=3D PMD_MASK >> PAGE_SHIFT; - - /* Unlike PTEs which are automatically upgraded to - * writeable entries, the PMD upgrades go through - * .huge_fault(). Make sure we pass the "write" info - * along in that case. - * This also means we have to record the write fault - * here, instead of in .pfn_mkwrite(). - */ - ret =3D vmf_insert_pfn_pmd(vmf, pfn, - vmf->flags & FAULT_FLAG_WRITE); - if (ret =3D=3D VM_FAULT_NOPAGE && (vmf->flags & FAULT_FLAG_WRITE)) - drm_gem_shmem_record_mkwrite(vmf); - - return ret; + ret =3D vmf_insert_pfn_pmd(vmf, pfn, write); } #endif } - return VM_FAULT_FALLBACK; + + if (ret =3D=3D VM_FAULT_NOPAGE && write) + drm_gem_shmem_record_mkwrite(vmf); + return ret; } =20 static vm_fault_t drm_gem_shmem_any_fault(struct vm_fault *vmf, unsigned i= nt order) --=20 2.55.0 From nobody Thu Sep 24 12:10:22 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D4D4B3A7F55 for ; Tue, 4 Aug 2026 12:05:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845147; cv=none; b=oeJc7WvDHr2c60XboNJ6GlGioBKBtOnGrlVpxRnwt6qq7lRLVtd8pObF5nOHSoXragMPjfY+QlBSOS7pfjBWXvs75p3hFUF3AN/2PwbQZ+YNJvIc0oI6OMdyQ8JKfNbL+sp3W/AMCNv+xptBODLi8KV/XsuEuEy7KMQWd2AsRmM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845147; c=relaxed/simple; bh=cDZoklWx0/N7wex/E0OmJ7dckrnB/Lv/S0imotkXjpA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XygvmusjPlt8ort29rBm8DontkvXqEDyO0DahH/EM1Fk/xKJDgWwZ2cqdiaqNKmWZVOf4SVah4Srm/OnHTlwt+tDkmPvXln7M5Jp60cspzvYbIWESoh/CpbZ84VVVCqhogCmfaC6R7IhREOcA0PILxWFTdcx809NNDrXVaI3rVw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=JtAD0yik; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=PBlyJqgg; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="JtAD0yik"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="PBlyJqgg" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785845144; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=tjD464vOIu2Voqk/KOGceV068H0b4O8GMn0RbRgTxRU=; b=JtAD0yikt7uIbSB66QLsjCSn9FY/gZqBzemeBQRRqyk7eYeTVDtC3+o7EfTjoGBQOVwtOv VJmUFOuvFLEsE0GK0v+6WtzqLWu/wxM5dBJryGUkUA342Joy35D/ClWZn0QkdG1hIhuGfP C+n4N6sg5R0GaEmS+zGX7G4aw6aSaUQ= Received: from mail-wm1-f70.google.com (mail-wm1-f70.google.com [209.85.128.70]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-84-ot9ZS2IZPVKBErk2RW9YpQ-1; Tue, 04 Aug 2026 08:05:43 -0400 X-MC-Unique: ot9ZS2IZPVKBErk2RW9YpQ-1 X-Mimecast-MFC-AGG-ID: ot9ZS2IZPVKBErk2RW9YpQ_1785845142 Received: by mail-wm1-f70.google.com with SMTP id 5b1f17b1804b1-495529a93f9so33867305e9.3 for ; Tue, 04 Aug 2026 05:05:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1785845142; x=1786449942; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=tjD464vOIu2Voqk/KOGceV068H0b4O8GMn0RbRgTxRU=; b=PBlyJqggW557o63F/kMXtvdzr2cYcTMTTjJqXhJYSUppJ8jSsDSuNBQDc8qYBT3iTU K4KnBO6ItEFRf/kI5UjlrfAuGEQfJsOYxAhhMXxbVkDiciKBVL3ljk3HV+UwvXajOhr4 UQuz2Ddmdw/sZDZRMsjuIeXuU/KXy84ll79fckKqYsiw+Scf5uev7spLgc8nn8Jby9Mf kC/vFboTw2C6Gd2l5LecAQVE4BuR/5RCDHBzPK+Y9tnUWL0bakvh/EX9ftAmZL/ba8RW Q+oHjgYaWQps0V9rkmJrz4pvSSUAGf8d88sp3v9Q7fYQp5FxRbjfLxZDf38IG9P/pVc3 B29Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785845142; x=1786449942; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=tjD464vOIu2Voqk/KOGceV068H0b4O8GMn0RbRgTxRU=; b=OmFXfC+YrZPbTNmcDLH0UL5YOAKZ64jdG8u0e0pyQnq+h05EiktHbB/9LxFW2GaCW6 L10uAi4vG/NRmV5wSViKqgKUKti4dlC1LuDEolXkybzET0RG7gA8ttR8D8jlU1p4voEs 4qNmOCtNCPt0FhoDY8pMDrOD9XjlgpU0DtkEDgJgzBdvF10oTPtavWMCKrwFdxgEhp28 s2bqB7STwMXMsKHp9vZ28c1qmm8mubXooU8/IHWz9tD0ougUWj4XwKFc3x0Y0q8p7OOS hYaMySf2PGC+zUqxuOoTvmBHZw55M3yiTpBgG9ArTwQWyonQyeN+7h+MpxBYFTIX9WCm 6/qQ== X-Gm-Message-State: AOJu0YxE8i+HqgdarqApaciFPSJ/qODgAFxAtcKPcutx71qA++ttxsnl z7/HuPDevUnmpMcRUX+VTYtFXeOzUMeqPrghqRVwXuGBu04Nzy2ar0jWcR5O4ROByCxft7M3irR wNx7UUyobmoC8LIpxLknMlmk5RTtkAHQraPm0VYgOAgLmur3B/pT8fikDFbHbrSEArCQeoxzRRr c6pUWJ9+J7AgbAe93dVxSoHrvhnzG/TSngAnGR22VqXJYcNv04tg== X-Gm-Gg: AR+sD11WTuD8EuLZWOLGaJ3FzMFwPo60pozAtsgJX/LQXivk2mea+NtAMEWBedgG7ZT /AmVbtCMES2ZvQ187wbKRrKaH1wafCEqI9ca2mb9Nha4bzCn/75GLGkHjEZM964kkExQWk/iLAq PXUtVmrCdmQ+EybQs+Rh4bJmSPD/bHheS36SjEW2jnRLiEdj72ylXqP2MrA+PgMk3UAYXiVunW+ vm4a6RQhjCFYWs7wVd9eUU0iK5crzJ6waLiA4tqH+WUpkSIKsZH8dlacOVL3+C008xMcfvLi8+7 nfyZYHsa+T5jzm6smcepwqgtXScBMW/IM5kX8xS6cVyFX4cbi6uxEYXDhDq1J6ubDA8HyVmNydH Tf49Gu9W5OoQ18gd8VsjjUIinLeQ1R8lWHZrT4/l08+ig+9AcjhEN0uIOvuN8Mtwn+Q32oE9i3z X0TSQ= X-Received: by 2002:a05:600c:c8c:b0:495:4589:707 with SMTP id 5b1f17b1804b1-4980c6452f7mr312396685e9.5.1785845142393; Tue, 04 Aug 2026 05:05:42 -0700 (PDT) X-Received: by 2002:a05:600c:c8c:b0:495:4589:707 with SMTP id 5b1f17b1804b1-4980c6452f7mr312394755e9.5.1785845141410; Tue, 04 Aug 2026 05:05:41 -0700 (PDT) Received: from [192.168.10.48] ([151.95.34.92]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49949fc2ff6sm79768615e9.1.2026.08.04.05.05.38 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 04 Aug 2026 05:05:39 -0700 (PDT) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: Alex Williamson , bcm-kernel-feedback-list@broadcom.com, Boris Brezillon , Christian Koenig , David Hildenbrand , dri-devel@lists.freedesktop.org, Fei Li , Huang Rui , linux-mm@kvack.org, linux-s390@vger.kernel.org, Michal Hocko , Peter Xu , Sergio Lopez , Sean Christopherson , Thomas Zimmermann , stable@vger.kernel.org Subject: [PATCH v2 3/6] drm/ttm, drm/vmwgfx: directly create writable PTEs when mkwrite is in use Date: Tue, 4 Aug 2026 14:05:25 +0200 Message-ID: <20260804120529.1730187-4-pbonzini@redhat.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260804120529.1730187-1-pbonzini@redhat.com> References: <20260804120529.1730187-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" This ensures that fixup_user_fault() users see a writable PTE when they request one. The flip side is that vmw_bo_vm_fault() now has to record by hand the write fault, because .pfn_mkwrite() is not invoked. Prefaulting works as before because only the first entry comes out writable, while the following ones still end up executing the .pfn_mkwrite() callback. Cc: stable@vger.kernel.org Signed-off-by: Paolo Bonzini --- drivers/gpu/drm/ttm/ttm_bo_vm.c | 7 ++-- drivers/gpu/drm/vmwgfx/vmwgfx_page_dirty.c | 42 ++++++++++++---------- 2 files changed, 29 insertions(+), 20 deletions(-) diff --git a/drivers/gpu/drm/ttm/ttm_bo_vm.c b/drivers/gpu/drm/ttm/ttm_bo_v= m.c index a80510489c45..3ebde936ce60 100644 --- a/drivers/gpu/drm/ttm/ttm_bo_vm.c +++ b/drivers/gpu/drm/ttm/ttm_bo_vm.c @@ -191,6 +191,7 @@ vm_fault_t ttm_bo_vm_fault_reserved(struct vm_fault *vm= f, unsigned long pfn; struct ttm_tt *ttm =3D NULL; struct page *page; + bool mkwrite; int err; pgoff_t i; vm_fault_t ret =3D VM_FAULT_NOPAGE; @@ -242,6 +243,7 @@ vm_fault_t ttm_bo_vm_fault_reserved(struct vm_fault *vm= f, * Speculatively prefault a number of pages. Only error on * first page. */ + mkwrite =3D !!(vmf->flags & FAULT_FLAG_WRITE); for (i =3D 0; i < num_prefault; ++i) { if (bo->resource->bus.is_iomem) { pfn =3D ttm_bo_io_mem_pfn(bo, page_offset); @@ -263,9 +265,10 @@ vm_fault_t ttm_bo_vm_fault_reserved(struct vm_fault *v= mf, * at arbitrary times while the data is mmap'ed. * See vmf_insert_pfn_prot() for a discussion. */ - ret =3D vmf_insert_pfn_prot(vma, address, pfn, prot); + ret =3D vmf_insert_pfn_prot_mkwrite(vma, address, pfn, prot, mkwrite); =20 - /* Never error on prefaulted PTEs */ + /* Never error on prefaulted PTEs and never map them writable */ + mkwrite =3D false; if (unlikely((ret & VM_FAULT_ERROR))) { if (i =3D=3D 0) return VM_FAULT_NOPAGE; diff --git a/drivers/gpu/drm/vmwgfx/vmwgfx_page_dirty.c b/drivers/gpu/drm/v= mwgfx/vmwgfx_page_dirty.c index 45561bc1c9ef..3099558c0762 100644 --- a/drivers/gpu/drm/vmwgfx/vmwgfx_page_dirty.c +++ b/drivers/gpu/drm/vmwgfx/vmwgfx_page_dirty.c @@ -398,15 +398,33 @@ void vmw_bo_dirty_clear_res(struct vmw_resource *res) dirty->end =3D res_start; } =20 +static vm_fault_t vmw_bo_dirty_mkwrite(struct vm_fault *vmf, struct ttm_bu= ffer_object *bo) +{ + unsigned long page_offset; + struct vmw_bo *vbo =3D to_vmw_bo(&bo->base); + + page_offset =3D vmf->pgoff - drm_vma_node_start(&bo->base.vma_node); + if (unlikely(page_offset >=3D PFN_UP(bo->resource->size))) + return VM_FAULT_SIGBUS; + + if (vbo->dirty && vbo->dirty->method =3D=3D VMW_BO_DIRTY_MKWRITE && + !test_bit(page_offset, &vbo->dirty->bitmap[0])) { + struct vmw_bo_dirty *dirty =3D vbo->dirty; + + __set_bit(page_offset, &dirty->bitmap[0]); + dirty->start =3D min(dirty->start, page_offset); + dirty->end =3D max(dirty->end, page_offset + 1); + } + return 0; +} + vm_fault_t vmw_bo_vm_mkwrite(struct vm_fault *vmf) { struct vm_area_struct *vma =3D vmf->vma; struct ttm_buffer_object *bo =3D (struct ttm_buffer_object *) vma->vm_private_data; vm_fault_t ret; - unsigned long page_offset; unsigned int save_flags; - struct vmw_bo *vbo =3D to_vmw_bo(&bo->base); =20 /* * mkwrite() doesn't handle the VM_FAULT_RETRY return value correctly. @@ -419,22 +437,7 @@ vm_fault_t vmw_bo_vm_mkwrite(struct vm_fault *vmf) if (ret) return ret; =20 - page_offset =3D vmf->pgoff - drm_vma_node_start(&bo->base.vma_node); - if (unlikely(page_offset >=3D PFN_UP(bo->resource->size))) { - ret =3D VM_FAULT_SIGBUS; - goto out_unlock; - } - - if (vbo->dirty && vbo->dirty->method =3D=3D VMW_BO_DIRTY_MKWRITE && - !test_bit(page_offset, &vbo->dirty->bitmap[0])) { - struct vmw_bo_dirty *dirty =3D vbo->dirty; - - __set_bit(page_offset, &dirty->bitmap[0]); - dirty->start =3D min(dirty->start, page_offset); - dirty->end =3D max(dirty->end, page_offset + 1); - } - -out_unlock: + ret =3D vmw_bo_dirty_mkwrite(vmf, bo); dma_resv_unlock(bo->base.resv); return ret; } @@ -484,6 +487,9 @@ vm_fault_t vmw_bo_vm_fault(struct vm_fault *vmf) prot =3D vm_get_page_prot(vma->vm_flags); =20 ret =3D ttm_bo_vm_fault_reserved(vmf, prot, num_prefault); + if (ret =3D=3D VM_FAULT_NOPAGE && (vmf->flags & FAULT_FLAG_WRITE)) + WARN_ON_ONCE(vmw_bo_dirty_mkwrite(vmf, bo)); + if (ret =3D=3D VM_FAULT_RETRY && !(vmf->flags & FAULT_FLAG_RETRY_NOWAIT)) return ret; =20 --=20 2.55.0 From nobody Thu Sep 24 12:10:22 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B148838DC65 for ; Tue, 4 Aug 2026 12:05:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845157; cv=none; b=Tp+VC4XIM8KNAmB/snYSV/gD4WbUP7BYvUxKEyiYD5qh4Atym4DdCuyVmMCBS63jszDOwGThxTZXxq4aGoMBsRtcqz5i6OVNH2hkpHP8IP+0DPTG58z9WMHwc7fgCPSnV9gGpeZ3S0JBPvDqxGKherR+naE8IIKDuWWgvjDqTNw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845157; c=relaxed/simple; bh=ShgJllGtoj3B6lj3SHwP+387ixSZNj2YWvPFDV6BseQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BGXXxnyMGw1Th+wK2sB4PoML/yqLV5U0A0dbA7BVbcpjF0FMl+Pk1e7muZ6O/E72x4p3P9c3zOivfd1gVFvfxUZ4kcBabaXsyw4iaOIbZouM4wS9XAMz7vejX2OOZu1REh4cPW0NRfjfHCEvF+ABdLQgefWR/RQNwoOPS9cRl0U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=Fc+V+8IK; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=cAIhJJ5i; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="Fc+V+8IK"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="cAIhJJ5i" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785845153; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=7VbhR/EeseCGUhv0h0yNN2Py/m+SxPzo7FfzMk4g7pY=; b=Fc+V+8IKcZCcp6czP4p2omH3FWU8+EnyxkruDA0g15z08MySPBOaaq3FTUC6F5ClD3O2rd 36ONEIWQtToahldeprLEYidSpcqA7+d2DHNNtHRjS+MbL1dliXGq4rUY4nt5wzJC2b+X+M gZ2Wzfd5km3LLilFPutcFkhOrzethJI= Received: from mail-wr1-f69.google.com (mail-wr1-f69.google.com [209.85.221.69]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-216-2-M5-ChmMV2V13PC3mVqmg-1; Tue, 04 Aug 2026 08:05:47 -0400 X-MC-Unique: 2-M5-ChmMV2V13PC3mVqmg-1 X-Mimecast-MFC-AGG-ID: 2-M5-ChmMV2V13PC3mVqmg_1785845146 Received: by mail-wr1-f69.google.com with SMTP id ffacd0b85a97d-47f83416551so4290353f8f.3 for ; Tue, 04 Aug 2026 05:05:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1785845146; x=1786449946; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=7VbhR/EeseCGUhv0h0yNN2Py/m+SxPzo7FfzMk4g7pY=; b=cAIhJJ5iypZOUTPkcOfJVA/+T+eNpiURGumey1wvCCiEqCPTbkXpjDSGd2kXAlHDgI it6fL70CYJYt++jBX6Pg7dkNwpXmtH4S8MZJHfIyPz7qSzaoI7AJWObi6LyXq1QAfh/N Nr2H8i0KrzVACddd0yxSRxEE6+k/odK4HgyRlQphWwyA20wzRZgCDegc5FmKqS2c5ye6 RpHmIy+9gHvLWdHzrHXywY3ZWc4sDBUJ7s6R48hgs3rRopd8PtRMx2yA9PQRiWon403Q VhxLqw3teQZewLTMwxlTswIepR9ZxT8+XJNtkWPVglCnSf0NucOUpwbbhFuuxmtfrjBV 3ijw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785845146; x=1786449946; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=7VbhR/EeseCGUhv0h0yNN2Py/m+SxPzo7FfzMk4g7pY=; b=Jf7CabhWFinoq4/NAvQ+QD5LaNNyb2x5n5y1Vz22RQy/yCjIkC/lsu1syabY4t4HsF TCDNCJ8p5GcV6elLWKE9mO5LflP4QKY6IyEokQYLw48i8KcimwRZhgPK/dxST60BsJOZ jp097CxqJpQW9LXzD8TngyEst1KzogQRgX6axhWSNtFi7URAcV+jpsCbDv85jFARVZha s4znEWbq0hO4xeVwvloHGhg6ikCihJwlYEtyvubW5gW/ZvSDzziGnPdH36f765a3CkMU jmmS7ybNLNr3ORQqgZJrSXVJXlM7R2hiry5I2iaEYCr2slxQ5H9lSKg5+H35ToFE58Ra KmUg== X-Gm-Message-State: AOJu0YwXK5Ssw7BqYFhMyRIbvx+mPCXHK8IYBUBAGlqW5GQt1YJNA42r 4z281TL2z1pO+Rkh6kOUUSEHfRTKRjvaltwD/eCWNPna9T6XCCLr7cR0x9GpxWCXnsF7WOjzqu+ QxV6iVIuP0EuRZDICDBZxLguYIfO6TJmu4GGH9YH4E2JV8DZFoaH6Mdgr9WUdzXBeEKnnxiNKkZ erlhmlYJOdVa8V869dsrjKLuEzhTUHMA5wox03uDEuyHs+3R6nNA== X-Gm-Gg: AR+sD11/2+4O+KA241rUaJbV6z386gWuAC2Jb/MA+uSXw+2TyWNXpFFH4Pqn9N23LA4 xHRawlcoPVSi0kjnHmiHB5mjWutoXWyDJkKpnZnNKtJRNIbZ69EZIJhE6Zh5WTcLBik/h224fM/ q3aZmXP54HIgs1yPysKyTZxP29CvqCfTfGbWRtU1dHJ7cjeC6MdpNPVT99lRLkT4a8Wof5VKTZ3 6tOeGUgv1hi50wE9fvudb6qsFi3YtnPf/b0UM1G9lTcYUFtG/zizGCh6vVEfndJvj1qdt4qSh+M oAXiaZlGMCNA22ICR3Sb5vseWBnFQJpR9tBuMf0DE6+ujathTrD/tpfDOFzSHNPN/J80z4OgLYF tUqOM8jXdM8TUHfbpVOTLKga0oJSLyVTbzWTGvp3FZh1pHm5JeyheNFSP2823uPQm5B2VkgPXSs ER2ck= X-Received: by 2002:a05:6000:2989:20b0:47f:7154:9dfc with SMTP id ffacd0b85a97d-47fd72a8e68mr29813808f8f.9.1785845146115; Tue, 04 Aug 2026 05:05:46 -0700 (PDT) X-Received: by 2002:a05:6000:2989:20b0:47f:7154:9dfc with SMTP id ffacd0b85a97d-47fd72a8e68mr29813677f8f.9.1785845145601; Tue, 04 Aug 2026 05:05:45 -0700 (PDT) Received: from [192.168.10.48] ([151.95.34.92]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-47fd4562667sm45367660f8f.24.2026.08.04.05.05.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 04 Aug 2026 05:05:43 -0700 (PDT) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: Alex Williamson , bcm-kernel-feedback-list@broadcom.com, Boris Brezillon , Christian Koenig , David Hildenbrand , dri-devel@lists.freedesktop.org, Fei Li , Huang Rui , linux-mm@kvack.org, linux-s390@vger.kernel.org, Michal Hocko , Peter Xu , Sergio Lopez , Sean Christopherson , Thomas Zimmermann , stable@vger.kernel.org Subject: [PATCH v2 4/6] kvm: apply VM_READ/VM_WRITE checks to all VMA types Date: Tue, 4 Aug 2026 14:05:26 +0200 Message-ID: <20260804120529.1730187-5-pbonzini@redhat.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260804120529.1730187-1-pbonzini@redhat.com> References: <20260804120529.1730187-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The VM_READ and VM_WRITE flags are checked only at the very end of hva_to_pfn(). For both the hva_to_pfn_remapped() case and for regular mappings, this adds unnecessary cases and inconsistent error behavior. For hva_to_pfn_remapped(), the code is relying on fixup_user_fault() to detect this situation. This is fragile because hva_to_pfn_remapped() returns different error codes for a !VM_WRITE VMA depending on whether the PTE happens to be mapped: * if the PTE is present, follow_pfnmap_start() sets args.writable to false and KVM_PFN_ERR_RO_FAULT is returned; * if no PTE is present, fixup_user_fault(FAULT_FLAG_WRITE) returns -EFAULT after checking vma_permits_fault(), and hva_to_pfn() ends up returning KVM_PFN_ERR_FAULT. With this patch KVM_PFN_ERR_RO_FAULT is returned uniformly. Likewise, a PROT_NONE pfnmap VMA would be mapped into the guest if the PTE was pte_present()[1] when the guest attempted to read it; with the patch instead KVM uniformly returns KVM_PFN_ERR_FAULT. Doing the check early avoids these special cases and also sidesteps the issue pointed out at https://sashiko.dev/#/patchset/20260731160514.1101989-1-pbonzini%40redhat.c= om. For regular mappings a PROT_READ VMA, if placed in a writable memslot, would return KVM_PFN_ERR_FAULT instead of KVM_PFN_ERR_RO_FAULT when the guest writes to it. This would cause a -EFAULT exit to userspace, instead of triggering emulation as the VM_IO|VM_PFNMAP arm would do; however it should be considered part of the KVM API because mmu_stress_test relies on it. Still, even with this snag about the returned pfn error code, pull the vm_flags checks in front so that they are done for all VMAs and the above inconsistency goes away for the VM_IO|VM_PFNMAP case. [1] on x86, for example, such a page would have _PAGE_PRESENT clear but _PAGE_PROTNONE set Fixes: 28e3918179aa ("drm/gem-shmem: Track folio accessed/dirty status in m= map") Cc: stable@vger.kernel.org Signed-off-by: Paolo Bonzini --- virt/kvm/kvm_main.c | 34 ++++++++++++++++------------------ 1 file changed, 16 insertions(+), 18 deletions(-) diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 45e784462ec6..576bcb21be3a 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -2925,17 +2925,6 @@ static int hva_to_pfn_slow(struct kvm_follow_pfn *kf= p, kvm_pfn_t *pfn) return npages; } =20 -static bool vma_is_valid(struct vm_area_struct *vma, bool write_fault) -{ - if (unlikely(!(vma->vm_flags & VM_READ))) - return false; - - if (write_fault && (unlikely(!(vma->vm_flags & VM_WRITE)))) - return false; - - return true; -} - static int hva_to_pfn_remapped(struct vm_area_struct *vma, struct kvm_follow_pfn *kfp, kvm_pfn_t *p_pfn) { @@ -3008,20 +2997,29 @@ kvm_pfn_t hva_to_pfn(struct kvm_follow_pfn *kfp) retry: vma =3D vma_lookup(current->mm, kfp->hva); =20 - if (vma =3D=3D NULL) + /* + * GUP failed. It could be an inaccessible mapping, a pfnmap one, + * or the page might be absent. + */ + + if (vma =3D=3D NULL || unlikely(!(vma->vm_flags & VM_READ))) { pfn =3D KVM_PFN_ERR_FAULT; - else if (vma->vm_flags & (VM_IO | VM_PFNMAP)) { + } else if ((kfp->flags & FOLL_WRITE) && unlikely(!(vma->vm_flags & VM_WRI= TE))) { + /* + * Exit to userspace for PROT_READ mappings in a writable + * memslot, as this is part of the API. + */ + pfn =3D vma->vm_flags & (VM_IO | VM_PFNMAP) ? KVM_PFN_ERR_RO_FAULT : + KVM_PFN_ERR_FAULT; + } else if (vma->vm_flags & (VM_IO | VM_PFNMAP)) { r =3D hva_to_pfn_remapped(vma, kfp, &pfn); if (r =3D=3D -EAGAIN) goto retry; if (r < 0) pfn =3D KVM_PFN_ERR_FAULT; } else { - if ((kfp->flags & FOLL_NOWAIT) && - vma_is_valid(vma, kfp->flags & FOLL_WRITE)) - pfn =3D KVM_PFN_ERR_NEEDS_IO; - else - pfn =3D KVM_PFN_ERR_FAULT; + pfn =3D kfp->flags & FOLL_NOWAIT ? KVM_PFN_ERR_NEEDS_IO : + KVM_PFN_ERR_FAULT; } mmap_read_unlock(current->mm); return pfn; --=20 2.55.0 From nobody Thu Sep 24 12:10:22 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E08AD3A16A1 for ; Tue, 4 Aug 2026 12:05:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845156; cv=none; b=jXyuwGzJwAbElPQpP1k5QCns+st5ThzL30s2umWkPCFrBVQcnjDVRzJWB0rxKP80ZU9U27yAqTak+rt0XILzIJ/be4K0j6g/EvV/5BZYnNlEHP8sEgxeZQXvR3cEZh+DKQo2XnI6aRYy2e9Ct+EvuJ8jwwL9TQy4Ysfd+DnNjSU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845156; c=relaxed/simple; bh=pfZ9447zK/9Z8PR/ugAEc+1h5xI66Sw/HNM5zHbRVZw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rx66ueqdM0Gn+6TMV1oKcXeJ3a+LCU0jMkuXgfdpdrfid7UammH3nUwvUHn9hgcMjVA3tjQ0r4zWA9lAsc2inHNsaBx36cUt/ZwMiS9y8/h3IySHWuh9Pstra+5RTqMaGGpN79rLcWpWSblffkOi/D5cpGazNjpvI8DAYaNHnw4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=hziiORRS; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=E6C/yKqM; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="hziiORRS"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="E6C/yKqM" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785845153; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=j37JFINyTYCydy6a1WRFKL3cyTkOscZBfjMnbKqhj70=; b=hziiORRSCX3td/DtoTVXZKbUu11WEU0DDNqJU3uRCS1EwEZzSyBqg+cq9rv1TexrCesJTy /tss3rjxqeHTk44ymO2f+L2jVMdICLYxBUK9ZcrWCKdUjpd//9uVtveUCgjd+MKW38nFxz detb/B9Wy3tbNOewh7efxcTjfWsi4G0= Received: from mail-wr1-f72.google.com (mail-wr1-f72.google.com [209.85.221.72]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-13-quHdnuZsP-uiRW4B8Vw58A-1; Tue, 04 Aug 2026 08:05:50 -0400 X-MC-Unique: quHdnuZsP-uiRW4B8Vw58A-1 X-Mimecast-MFC-AGG-ID: quHdnuZsP-uiRW4B8Vw58A_1785845150 Received: by mail-wr1-f72.google.com with SMTP id ffacd0b85a97d-47f6db77430so3331656f8f.0 for ; Tue, 04 Aug 2026 05:05:50 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1785845149; x=1786449949; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=j37JFINyTYCydy6a1WRFKL3cyTkOscZBfjMnbKqhj70=; b=E6C/yKqMRyFadcJC6fKjGVc9ZurbscECGfuXwXa5veKYxO3wgWpV1hfVxJE3y0jUpV ugyjpClLjCphtzH/JVprAl5b504ncEXNhLkFOqouMcmwnG7a1G4XiUjpGJQKfzLqCT1g F6avX0ZWzs2DfGnzUW3DgRS5Zc3jk/kXPf2ygWIBUJ9cdmjvRdsHUCPA0jcYOX2NI1gQ J7Rrx9t9y9RCOyECw4ICAphTMOPxbaU3qDz6cYKFi8WBHPoLXOO3qU1loSVj4JNRb+5M iQrli2L7qbh9e9uVNAxTlnZH9HnxMlSsEmsM88LBCF6O1bWNMk/Mn6G5j1j/+O5GwkX1 RVkQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785845149; x=1786449949; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=j37JFINyTYCydy6a1WRFKL3cyTkOscZBfjMnbKqhj70=; b=Tc227gpAhGBWoaobmZ7JjP1R8i7nqBORWLniI52hWMB6rWacikX7NP8vRL7PDi6ki0 h3n6Se1ek3y0qzmT9/t8cgy8F0eP/2S9/nypfn7q7VYG6bFohx1DDsT54+uR0dwpQDT2 pVSOHqnsgz4wcn9j2pFMWYHDa7rZOakqcgiHZ8idfxL70IHAWIsT4EYuEZoYUh9wC3uD 2TYWhBiNkGgxvFKC9ir2N72vhiAbIi9iZH1nPRs7tZ92Ep38h/bfvaYGwXZYgELsCfIE 0Go7pHfi/uauH3396IJwrsbNiLhjuPLqDpDag+Bxkh2uDWtfRswfPgqLyxhRFsMH2qqw yX2A== X-Gm-Message-State: AOJu0Yw0IBNIP3v8r/KB/nmLzPid8sGVdyrrPfXgIui+oCONwttew04f 1ocLrJhCQ7Go70VPgQ5andsku0/Mxx3BJCnugZexNHQ7o7LwAc5zC4+cE+cAiG/p5pCzwVUFols GrHs0DZXNVcRRCVN6O6S1DnlwHNbp3YlNCCt+86Nx6ZZ3gwAYYrjwez8TexbZFnaQeo/btlGoud pjZfujRe5kdBDqRrT/ibUfBlKkRBnfQVn44Hpu5oZ266CYYMmvGg== X-Gm-Gg: AR+sD10kIfdxkC0K6GYeJJVE0jX9ekd8FYMY31Zj+0R+elHPyz6O4ES6RvT0v93mGBe f6S729NpP5jl5TQ0zpzRtI+Unegoyz1qDfgA/WWB83IEXG2dyI7s6TlrAcLrywkIbZ5MAklivWi e+u3y69b53pdLx8n5j6P/fSR90ZYJ0Jb/qK1CCEFW8iv0/bhpvPjB196BJdQuhaqQW33ZJEns3U j08ByeDMQMTyeGQw8Pisdcawq10D9UasnEfCKr0VNpv/8KNyOVajg8EXFUxng8sMEPwyUmm2V2V jXcvx174Y6FxVcqtS77FTS/fO2jWFxnO/ZSoV1hZqYRZ5CMq0GwJVrH5lCFgzOK51gZillUIa01 RWKiDGB0aIZLpqv1Vims4VRhHhH2/jJbBMGoax5RtsrIZiycIWsKOB6KabN47X5aIlAUOx8fe0g l/fSc= X-Received: by 2002:adf:eb8f:0:b0:47f:97f6:d39a with SMTP id ffacd0b85a97d-47fd72c7121mr27900835f8f.17.1785845149474; Tue, 04 Aug 2026 05:05:49 -0700 (PDT) X-Received: by 2002:adf:eb8f:0:b0:47f:97f6:d39a with SMTP id ffacd0b85a97d-47fd72c7121mr27900691f8f.17.1785845148854; Tue, 04 Aug 2026 05:05:48 -0700 (PDT) Received: from [192.168.10.48] ([151.95.34.92]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-47fd41d1756sm41086159f8f.4.2026.08.04.05.05.46 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 04 Aug 2026 05:05:47 -0700 (PDT) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: Alex Williamson , bcm-kernel-feedback-list@broadcom.com, Boris Brezillon , Christian Koenig , David Hildenbrand , dri-devel@lists.freedesktop.org, Fei Li , Huang Rui , linux-mm@kvack.org, linux-s390@vger.kernel.org, Michal Hocko , Peter Xu , Sergio Lopez , Sean Christopherson , Thomas Zimmermann , stable@vger.kernel.org Subject: [PATCH v2 5/6] mm: pull writability check to follow_pfnmap_start() Date: Tue, 4 Aug 2026 14:05:27 +0200 Message-ID: <20260804120529.1730187-6-pbonzini@redhat.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260804120529.1730187-1-pbonzini@redhat.com> References: <20260804120529.1730187-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" All callers of follow_pfnmap_start() except s390_pci_mmio_write() are following it, if they are doing a write, with a check that args.writable is true; for s390_pci_mmio_write() that's a bug. Also, most of them return -EFAULT if it is not. Pull the check directly into follow_pfnmap_start() through another input parameter args.write_fault, to eliminate the need to do it in the caller. This also fixes an issue where follow_pfnmap_start() would return 0 for a PFN that is mapped read-only, and the caller would not attempt to call fixup_user_fault() on it; this can happen with vm_ops that set .pfn_mkwrite(), for example. Instead, now the caller (for example hva_to_pfn_remapped()) sees an error, does attempt to fix it, and only returns -EFAULT if the fixup was fruitless. Reported-by: Sergio Lopez Fixes: 28e3918179aa ("drm/gem-shmem: Track folio accessed/dirty status in m= map") Link: https://lore.kernel.org/kvm/CAAiTLFU1ALsDoJoKW3d9bUvv990AozAoX=3DbEHm= fnG54qyBAHFg@mail.gmail.com/ Cc: stable@vger.kernel.org Signed-off-by: Paolo Bonzini --- arch/s390/pci/pci_mmio.c | 2 ++ drivers/vfio/vfio_iommu_type1.c | 17 +++++---- drivers/virt/acrn/mm.c | 10 +----- include/linux/mm.h | 3 ++ mm/memory.c | 62 ++++++++++++++++++++------------- virt/kvm/kvm_main.c | 15 ++++---- 6 files changed, 58 insertions(+), 51 deletions(-) diff --git a/arch/s390/pci/pci_mmio.c b/arch/s390/pci/pci_mmio.c index 51e7a28af899..d9d5b3318cbc 100644 --- a/arch/s390/pci/pci_mmio.c +++ b/arch/s390/pci/pci_mmio.c @@ -180,6 +180,7 @@ SYSCALL_DEFINE3(s390_pci_mmio_write, unsigned long, mmi= o_addr, =20 args.address =3D mmio_addr; args.vma =3D vma; + args.write =3D true; ret =3D follow_pfnmap_start(&args); if (ret) { fixup_user_fault(current->mm, mmio_addr, FAULT_FLAG_WRITE, NULL); @@ -332,6 +333,7 @@ SYSCALL_DEFINE3(s390_pci_mmio_read, unsigned long, mmio= _addr, =20 args.vma =3D vma; args.address =3D mmio_addr; + args.write =3D false; ret =3D follow_pfnmap_start(&args); if (ret) { fixup_user_fault(current->mm, mmio_addr, 0, NULL); diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type= 1.c index c8151ba54de3..e6d3a2311a99 100644 --- a/drivers/vfio/vfio_iommu_type1.c +++ b/drivers/vfio/vfio_iommu_type1.c @@ -541,7 +541,11 @@ static int follow_fault_pfn(struct vm_area_struct *vma= , struct mm_struct *mm, unsigned long vaddr, unsigned long *pfn, unsigned long *addr_mask, bool write_fault) { - struct follow_pfnmap_args args =3D { .vma =3D vma, .address =3D vaddr }; + struct follow_pfnmap_args args =3D { + .vma =3D vma, + .address =3D vaddr, + .write =3D write_fault, + }; int ret; =20 ret =3D follow_pfnmap_start(&args); @@ -563,15 +567,10 @@ static int follow_fault_pfn(struct vm_area_struct *vm= a, struct mm_struct *mm, return ret; } =20 - if (write_fault && !args.writable) { - ret =3D -EFAULT; - } else { - *pfn =3D args.pfn; - *addr_mask =3D args.addr_mask; - } - + *pfn =3D args.pfn; + *addr_mask =3D args.addr_mask; follow_pfnmap_end(&args); - return ret; + return 0; } =20 /* diff --git a/drivers/virt/acrn/mm.c b/drivers/virt/acrn/mm.c index 5bca500a83e0..2f9808399f19 100644 --- a/drivers/virt/acrn/mm.c +++ b/drivers/virt/acrn/mm.c @@ -177,7 +177,6 @@ int acrn_vm_ram_map(struct acrn_vm *vm, struct acrn_vm_= memmap *memmap) vma =3D vma_lookup(current->mm, memmap->vma_base); if (vma && ((vma->vm_flags & VM_PFNMAP) !=3D 0)) { unsigned long start_pfn, cur_pfn; - bool writable; =20 if ((memmap->vma_base + memmap->len) > vma->vm_end) { mmap_read_unlock(current->mm); @@ -188,6 +187,7 @@ int acrn_vm_ram_map(struct acrn_vm *vm, struct acrn_vm_= memmap *memmap) struct follow_pfnmap_args args =3D { .vma =3D vma, .address =3D memmap->vma_base + i * PAGE_SIZE, + .write =3D !!(memmap->attr & ACRN_MEM_ACCESS_WRITE), }; =20 ret =3D follow_pfnmap_start(&args); @@ -197,16 +197,8 @@ int acrn_vm_ram_map(struct acrn_vm *vm, struct acrn_vm= _memmap *memmap) cur_pfn =3D args.pfn; if (i =3D=3D 0) start_pfn =3D cur_pfn; - writable =3D args.writable; follow_pfnmap_end(&args); =20 - /* Disallow write access if the PTE is not writable. */ - if (!writable && - (memmap->attr & ACRN_MEM_ACCESS_WRITE)) { - ret =3D -EFAULT; - break; - } - /* Disallow refcounted pages. */ if (pfn_valid(cur_pfn) && !PageReserved(pfn_to_page(cur_pfn))) { diff --git a/include/linux/mm.h b/include/linux/mm.h index 01184a4bdd6f..1659cb8f42fd 100644 --- a/include/linux/mm.h +++ b/include/linux/mm.h @@ -3136,9 +3136,12 @@ struct follow_pfnmap_args { * Inputs: * @vma: Pointer to @vm_area_struct struct * @address: the virtual address to walk + * @write: if true, fail with -EFAULT unless the mapping is + * writable */ struct vm_area_struct *vma; unsigned long address; + bool write; /** * Internals: * diff --git a/mm/memory.c b/mm/memory.c index b5555217b121..27f5dcc319c8 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -6774,12 +6774,15 @@ int __pmd_alloc(struct mm_struct *mm, pud_t *pud, u= nsigned long address) } #endif /* __PAGETABLE_PMD_FOLDED */ =20 -static inline void pfnmap_args_setup(struct follow_pfnmap_args *args, - spinlock_t *lock, pte_t *ptep, - pgprot_t pgprot, unsigned long pfn_base, - unsigned long addr_mask, bool writable, - bool special) +static inline int pfnmap_args_setup(struct follow_pfnmap_args *args, + spinlock_t *lock, pte_t *ptep, + pgprot_t pgprot, unsigned long pfn_base, + unsigned long addr_mask, bool writable, + bool special) { + if (!writable && args->write) + return -EFAULT; + args->lock =3D lock; args->ptep =3D ptep; args->pfn =3D pfn_base + ((args->address & ~addr_mask) >> PAGE_SHIFT); @@ -6787,6 +6790,7 @@ static inline void pfnmap_args_setup(struct follow_pf= nmap_args *args, args->pgprot =3D pgprot; args->writable =3D writable; args->special =3D special; + return 0; } =20 static inline void pfnmap_lockdep_assert(struct vm_area_struct *vma) @@ -6808,8 +6812,9 @@ static inline void pfnmap_lockdep_assert(struct vm_ar= ea_struct *vma) * @args: Pointer to struct @follow_pfnmap_args * * The caller needs to setup args->vma and args->address to point to the - * virtual address as the target of such lookup. On a successful return, - * the results will be put into other output fields. + * virtual address as the target of such lookup, and optionally set + * args->write to require a writable mapping. On a successful + * return, the results will be put into other output fields. * * After the caller finished using the fields, the caller must invoke * another follow_pfnmap_end() to proper releases the locks and resources @@ -6832,7 +6837,8 @@ static inline void pfnmap_lockdep_assert(struct vm_ar= ea_struct *vma) * * This function must not be used to modify PTE content. * - * Return: zero on success, negative otherwise. + * Return: zero on success, -EFAULT if @args->write was set but the + * mapping is not writable, -EINVAL if there is no mapping at all. */ int follow_pfnmap_start(struct follow_pfnmap_args *args) { @@ -6845,6 +6851,7 @@ int follow_pfnmap_start(struct follow_pfnmap_args *ar= gs) pud_t *pudp, pud; pmd_t *pmdp, pmd; pte_t *ptep, pte; + int r =3D -EINVAL; =20 pfnmap_lockdep_assert(vma); =20 @@ -6878,10 +6885,12 @@ int follow_pfnmap_start(struct follow_pfnmap_args *= args) spin_unlock(lock); goto retry; } - pfnmap_args_setup(args, lock, NULL, pud_pgprot(pud), - pud_pfn(pud), PUD_MASK, pud_write(pud), - pud_special(pud)); - return 0; + r =3D pfnmap_args_setup(args, lock, NULL, pud_pgprot(pud), + pud_pfn(pud), PUD_MASK, pud_write(pud), + pud_special(pud)); + if (r) + spin_unlock(lock); + return r; } =20 pmdp =3D pmd_offset(pudp, address); @@ -6899,10 +6908,12 @@ int follow_pfnmap_start(struct follow_pfnmap_args *= args) spin_unlock(lock); goto retry; } - pfnmap_args_setup(args, lock, NULL, pmd_pgprot(pmd), - pmd_pfn(pmd), PMD_MASK, pmd_write(pmd), - pmd_special(pmd)); - return 0; + r =3D pfnmap_args_setup(args, lock, NULL, pmd_pgprot(pmd), + pmd_pfn(pmd), PMD_MASK, pmd_write(pmd), + pmd_special(pmd)); + if (r) + spin_unlock(lock); + return r; } =20 ptep =3D pte_offset_map_lock(mm, pmdp, address, &lock); @@ -6911,14 +6922,16 @@ int follow_pfnmap_start(struct follow_pfnmap_args *= args) pte =3D ptep_get(ptep); if (!pte_present(pte)) goto unlock; - pfnmap_args_setup(args, lock, ptep, pte_pgprot(pte), - pte_pfn(pte), PAGE_MASK, pte_write(pte), - pte_special(pte)); + r =3D pfnmap_args_setup(args, lock, ptep, pte_pgprot(pte), + pte_pfn(pte), PAGE_MASK, pte_write(pte), + pte_special(pte)); + if (r) + goto unlock; return 0; unlock: pte_unmap_unlock(ptep, lock); out: - return -EINVAL; + return r; } EXPORT_SYMBOL_GPL(follow_pfnmap_start); =20 @@ -6960,7 +6973,11 @@ int generic_access_phys(struct vm_area_struct *vma, = unsigned long addr, int offset =3D offset_in_page(addr); int ret =3D -EINVAL; bool writable; - struct follow_pfnmap_args args =3D { .vma =3D vma, .address =3D addr }; + struct follow_pfnmap_args args =3D { + .vma =3D vma, + .address =3D addr, + .write =3D !!(write & FOLL_WRITE) + }; =20 retry: if (follow_pfnmap_start(&args)) @@ -6970,9 +6987,6 @@ int generic_access_phys(struct vm_area_struct *vma, u= nsigned long addr, writable =3D args.writable; follow_pfnmap_end(&args); =20 - if ((write & FOLL_WRITE) && !writable) - return -EINVAL; - maddr =3D ioremap_prot(phys_addr, PAGE_ALIGN(len + offset), prot); if (!maddr) return -ENOMEM; diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 576bcb21be3a..b7c21a48a45c 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -2928,8 +2928,11 @@ static int hva_to_pfn_slow(struct kvm_follow_pfn *kf= p, kvm_pfn_t *pfn) static int hva_to_pfn_remapped(struct vm_area_struct *vma, struct kvm_follow_pfn *kfp, kvm_pfn_t *p_pfn) { - struct follow_pfnmap_args args =3D { .vma =3D vma, .address =3D kfp->hva = }; - bool write_fault =3D kfp->flags & FOLL_WRITE; + struct follow_pfnmap_args args =3D { + .vma =3D vma, + .address =3D kfp->hva, + .write =3D !!(kfp->flags & FOLL_WRITE), + }; int r; =20 /* @@ -2948,7 +2951,7 @@ static int hva_to_pfn_remapped(struct vm_area_struct = *vma, */ bool unlocked =3D false; r =3D fixup_user_fault(current->mm, kfp->hva, - (write_fault ? FAULT_FLAG_WRITE : 0), + (args.write ? FAULT_FLAG_WRITE : 0), &unlocked); if (unlocked) return -EAGAIN; @@ -2960,13 +2963,7 @@ static int hva_to_pfn_remapped(struct vm_area_struct= *vma, return r; } =20 - if (write_fault && !args.writable) { - *p_pfn =3D KVM_PFN_ERR_RO_FAULT; - goto out; - } - *p_pfn =3D kvm_resolve_pfn(kfp, NULL, &args, args.writable); -out: follow_pfnmap_end(&args); return r; } --=20 2.55.0 From nobody Thu Sep 24 12:10:22 2026 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2410838E5F9 for ; Tue, 4 Aug 2026 12:05:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845159; cv=none; b=iJRSOkVkLtm0r7iBE/zaUTfniR0OVD2SaNQK+rbLst2XJBTLS7SIPku/PyX6W+0yhx5CnYLZi4WYSiUfLoP0B/7ncT6k/TBbzyokeyyx+rqIFh73G0sKReeP6zloepnY65xGnHKSR77/CWcfY5jCIINFPJU1zwVlwgszonDCdlA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785845159; c=relaxed/simple; bh=Zcn8O1erPuSdDnXyXFymHAYFaLae2f2+vOh43hT0kQw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=hwPTRQVW1dTOb4h7CFRr5cfJIDF+K48hj2RWw6lklU6LFXj3wfr0zWkKx8gd9gv/7dxdZ0XlzVe4PshoUAOaLeY7Bz+jXOh+5Xnn54lOLyIYWSuN2wK09L6ckCsSxmWqHrqXEcFJldnNh2iNzfqNs1TehcRmIqElyv/Zi7k2h6w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=Stz6n6MW; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=gMALMmIM; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="Stz6n6MW"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="gMALMmIM" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1785845154; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=6MKpfy5ngorVqbj+qf8v5CmiQT/j6dhAeNWojtcyJWQ=; b=Stz6n6MWhRtr/PC82EV+/NB6DqC0gl6zHlAldhR645+srckGA+UyB94hUsZ/PNnQ0IBCU2 gJTjdTOXXmID/jfXGUj0NAItHbyRWmBgMAatDFfBrUL3x66twsDr/3+HGSvjETG3NxgDKs XQZVDUwWnDSEaoM6bb1FpVooPWcle78= Received: from mail-wm1-f72.google.com (mail-wm1-f72.google.com [209.85.128.72]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-241-Hw0gPYp-OPu_2jEXOAMCpQ-1; Tue, 04 Aug 2026 08:05:53 -0400 X-MC-Unique: Hw0gPYp-OPu_2jEXOAMCpQ-1 X-Mimecast-MFC-AGG-ID: Hw0gPYp-OPu_2jEXOAMCpQ_1785845152 Received: by mail-wm1-f72.google.com with SMTP id 5b1f17b1804b1-496bbcf7d1eso20947645e9.0 for ; Tue, 04 Aug 2026 05:05:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1785845152; x=1786449952; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=6MKpfy5ngorVqbj+qf8v5CmiQT/j6dhAeNWojtcyJWQ=; b=gMALMmIMGhsOWE09Di5u3dFbxkUKvOJmHguRc88lNDsV3F+LtSCWcOj7dt0VMNRM50 kAbR8X9w1Q1cVGJ0m7rMwgm5GvAIbMfEcnPfB7w07Ak1wG/RZTISnxZyNg3iLsg2XamU pB+EqBEXbS8HpDOn8FS+3OjUxYmneC4D0K6eMN6wwxzghh/zlXjqOtGaeK7JdjNeF/+5 xXuo8k2tWrHNXM4uL+ZxnlxcwHwMT6lt7XX2LY9XUP/VZCRbI117hkDLY+AGHsoKcJ5G E0KN+KZf81g0aVO2CvCWVohIuNLWgxw0oT2D+l1DNn0RT+Gc1uK1EMiKru98/xPOUyrl zevg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785845152; x=1786449952; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=6MKpfy5ngorVqbj+qf8v5CmiQT/j6dhAeNWojtcyJWQ=; b=OB7ESO+G7z7bADrb2OR4mLqy1T6u0Ch9kXcFOYj/vGXXEqUX0AITuFCYjXNkVJ/+ib 55PdnIRBAVq3o9z3KhGLFE5VG3c5bfbaTRbvf5kM+BAOgjzjVoKmxRHgX0AI/Gf5lDSi pDqKDkhhwynXEdAJFGHcRsjVkCf+Hv70/ZhpjuXV/wugosjqnOBhuCduPqyxpxt6w6gR afuhXazj8FkEEfaLQxBqy+jEkqKrj7JGl8M6A6x4nmoyh0VoRjO2QFPiQiNrHKa2cazc lzXemTbTNZg40Usrxkpt00AaqSITVLdV9EGImmHjnkxMsy0Y7cIq/q3rBi179Y+pezJG J9NQ== X-Gm-Message-State: AOJu0Yw81rQoPBLqR1718v7q5tl+DYA/rY+vffsf/YKQQMfGvHIUGyW7 K3qHYdNEyeB9wJ/LKDoIGvEwqwp2K4Pq+XfylDY/b0FDFPhhOg688A8fudU7nwfqyVj5ppJB80K tzIdvnqsurSU6lK/K4VJRJ7g3jOUx7HqyoWqudKqgs6G2hkSXZnQq6+VXlgwx3j0OfePGw2bXIT VwHwDdSVee5piU/2QnEvSH9OVMSnKMsHHL8OGxJ3/ZC70309xVUQ== X-Gm-Gg: AR+sD11hYRXLosHz2/ZtWQ0xy71udu6ZFtSC74EjZB+zRk6a6TbeOFb+BkBlYo1N9u1 fpvmtsiyDvkhbDNlUmKHQ9gQJD4+m/PPireFI/UNrhboC6JebLR/CjpCR3vLXR6XAqhMXbMdEBT PgynekUpJx3vDQlA9WFH1HZC+6hOotZJS55fYJAQHQOHtlLqQxjNrOuIz1N83Z3wzq/IgTZDwJx ghin/ZxkNpwYTFyqPQkOEFikUXa/FTpxhegXsZSbDBKCMaKu/ikx+vqMHSBGOe+/kr6qyjR9g6u T2vOcjqnrmsvxU1dgkHRPl+9Wn9JtFvhySFjA6W0KFp3jeaaVrh20iSRHTKuXyK/Ak9WL18fRp5 EbSFbduBp0RPYzENZkI0h8XoVJ1GcTCqikd897AqetcG4wtw7d/GSu6CzE5X8iw/Yd5k5Djl7+U cTqRs= X-Received: by 2002:a05:600c:840f:b0:493:c42e:5be0 with SMTP id 5b1f17b1804b1-4980c60705amr298312455e9.0.1785845151977; Tue, 04 Aug 2026 05:05:51 -0700 (PDT) X-Received: by 2002:a05:600c:840f:b0:493:c42e:5be0 with SMTP id 5b1f17b1804b1-4980c60705amr298311535e9.0.1785845151462; Tue, 04 Aug 2026 05:05:51 -0700 (PDT) Received: from [192.168.10.48] ([151.95.34.92]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49949fd8df5sm88354535e9.9.2026.08.04.05.05.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 04 Aug 2026 05:05:50 -0700 (PDT) From: Paolo Bonzini To: linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: Alex Williamson , bcm-kernel-feedback-list@broadcom.com, Boris Brezillon , Christian Koenig , David Hildenbrand , dri-devel@lists.freedesktop.org, Fei Li , Huang Rui , linux-mm@kvack.org, linux-s390@vger.kernel.org, Michal Hocko , Peter Xu , Sergio Lopez , Sean Christopherson , Thomas Zimmermann Subject: [PATCH v2 6/6] kvm: return -EFAULT for writes to !VM_WRITE IO mappings Date: Tue, 4 Aug 2026 14:05:28 +0200 Message-ID: <20260804120529.1730187-7-pbonzini@redhat.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260804120529.1730187-1-pbonzini@redhat.com> References: <20260804120529.1730187-1-pbonzini@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" KVM's behavior when the guest writes to a non-writable VMA is inconsistent. For regular, page-backed mappings it returns KVM_PFN_ERR_FAULT and thus returns -EFAULT to userspace (which is ABI, and relied upon by tests); for VM_IO/VM_PFNMAP mappings instead it returns KVM_PFN_ERR_RO_FAULT and thus exits to userspace with KVM_EXIT_MMIO. This behavior for VM_{IO,PFNMAP} was added by commit bd2fae8da794 ("KVM: do not assume PTE is writable after follow_pfn"), and even if it has been in place for five years it is unlikely that it is relied upon by userspace, since it is inconsistent with KVM itself. Change hva_to_pfn() to return KVM_PFN_ERR_FAULT for all non-writable VMAs, and restrict KVM_EXIT_MMIO to the case of an explicitly read-only memslot. Suggested-by: Sean Christopherson Signed-off-by: Paolo Bonzini Reviewed-by: Sean Christopherson --- virt/kvm/kvm_main.c | 11 +++-------- 1 file changed, 3 insertions(+), 8 deletions(-) diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index b7c21a48a45c..da5b0bb62590 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -2999,15 +2999,10 @@ kvm_pfn_t hva_to_pfn(struct kvm_follow_pfn *kfp) * or the page might be absent. */ =20 - if (vma =3D=3D NULL || unlikely(!(vma->vm_flags & VM_READ))) { + if (vma =3D=3D NULL || + unlikely(!(vma->vm_flags & VM_READ)) || + ((kfp->flags & FOLL_WRITE) && unlikely(!(vma->vm_flags & VM_WRITE))))= { pfn =3D KVM_PFN_ERR_FAULT; - } else if ((kfp->flags & FOLL_WRITE) && unlikely(!(vma->vm_flags & VM_WRI= TE))) { - /* - * Exit to userspace for PROT_READ mappings in a writable - * memslot, as this is part of the API. - */ - pfn =3D vma->vm_flags & (VM_IO | VM_PFNMAP) ? KVM_PFN_ERR_RO_FAULT : - KVM_PFN_ERR_FAULT; } else if (vma->vm_flags & (VM_IO | VM_PFNMAP)) { r =3D hva_to_pfn_remapped(vma, kfp, &pfn); if (r =3D=3D -EAGAIN) --=20 2.55.0