From nobody Thu Apr 9 19:24:33 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8E6AF30CDBC; Fri, 6 Mar 2026 17:18:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772817537; cv=none; b=jNc8jQc5r2UDHufAg6Tu2JmxRs/HloJBWCB1SN3ktO2pXgGCDHUe+aaZpMXQfimCNVj6RbFtKGKC7M7pO/PzJ4XaPzGjUKjINX/yDnpTXuE6+6Amm+6gy3vn/4wDrraOnn1wPHj7q/GWvnhzkM3eJcayTNbwToKcM7oMyC9MUW0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772817537; c=relaxed/simple; bh=uK6YVn1uE8/KerH/03wILxjQbYg54EfQnPDfF+duTRs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=hy+p2VJcGhU3SNYsYNsakfDBRRfzgrUElr5nCc97Ac7lMPYIm26GR9cNb61YwBdvalJKt3B2VhVAWFVDvScAFk+2CpQgErZPLCID2tjWuIzmY9MW+wmExXT/U6PxzaNIFCLvvmv0vmnU96qaPji53shUn+3QQD2/Zx5CyhVcmnw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=L9/3UiF0; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="L9/3UiF0" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 8E93AC2BCB5; Fri, 6 Mar 2026 17:18:51 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1772817537; bh=uK6YVn1uE8/KerH/03wILxjQbYg54EfQnPDfF+duTRs=; h=From:To:Cc:Subject:Date:In-Reply-To:References:From; b=L9/3UiF0bLNrJW644W0ccpADpS02wFJO0vTsEAx70zQnWCc+CmzLC0z5GNaljf838 9kmqMyDToEUngqQ0VlIu9ohLy+iRzBw5PM10ms0ZQHGffSwdW7ZtQLvPOEb/LG72K5 Tyf2dRIFqQqYxpoPFkaaSXkl1Y3ixUfXKoG2LPnxrMu1X3zUB9mdO75veRF4I0Nxam i3rxEdp3qAx0Y5UmJj9Yx209e0qhdG6Ui7U2tA1sWXPKk8npq4f9C/K9lIebErZiRW mGCHmipp5578CvkR5Qrznyly///YQ/XqjCCFnDB4oUAP9zMjtseOfZO+kh8Dpt3KvP bjE4FGXWmF+vw== From: Mike Rapoport To: Andrew Morton Cc: Andrea Arcangeli , Axel Rasmussen , Baolin Wang , David Hildenbrand , Hugh Dickins , James Houghton , "Liam R. Howlett" , Lorenzo Stoakes , "Matthew Wilcox (Oracle)" , Michal Hocko , Mike Rapoport , Muchun Song , Nikita Kalyazin , Oscar Salvador , Paolo Bonzini , Peter Xu , Sean Christopherson , Shuah Khan , Suren Baghdasaryan , Vlastimil Babka , kvm@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-mm@kvack.org Subject: [PATCH v2 05/15] userfaultfd: retry copying with locks dropped in mfill_atomic_pte_copy() Date: Fri, 6 Mar 2026 19:18:05 +0200 Message-ID: <20260306171815.3160826-6-rppt@kernel.org> X-Mailer: git-send-email 2.51.0 In-Reply-To: <20260306171815.3160826-1-rppt@kernel.org> References: <20260306171815.3160826-1-rppt@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Mike Rapoport (Microsoft)" Implementation of UFFDIO_COPY for anonymous memory might fail to copy data from userspace buffer when the destination VMA is locked (either with mm_lock or with per-VMA lock). In that case, mfill_atomic() releases the locks, retries copying the data with locks dropped and then re-locks the destination VMA and re-establishes PMD. Since this retry-reget dance is only relevant for UFFDIO_COPY and it never happens for other UFFDIO_ operations, make it a part of mfill_atomic_pte_copy() that actually implements UFFDIO_COPY for anonymous memory. As a temporal safety measure to avoid breaking biscection mfill_atomic_pte_copy() makes sure to never return -ENOENT so that the loop in mfill_atomic() won't retry copiyng outside of mmap_lock. This is removed later when shmem implementation will be updated later and the loop in mfill_atomic() will be adjusted. Signed-off-by: Mike Rapoport (Microsoft) --- mm/userfaultfd.c | 78 +++++++++++++++++++++++++++++++++--------------- 1 file changed, 54 insertions(+), 24 deletions(-) diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index baff11e83101..828f252c720c 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -159,6 +159,9 @@ static void uffd_mfill_unlock(struct vm_area_struct *vm= a) =20 static void mfill_put_vma(struct mfill_state *state) { + if (!state->vma) + return; + up_read(&state->ctx->map_changing_lock); uffd_mfill_unlock(state->vma); state->vma =3D NULL; @@ -404,35 +407,63 @@ static int mfill_copy_folio_locked(struct folio *foli= o, unsigned long src_addr) return ret; } =20 +static int mfill_copy_folio_retry(struct mfill_state *state, struct folio = *folio) +{ + unsigned long src_addr =3D state->src_addr; + void *kaddr; + int err; + + /* retry copying with mm_lock dropped */ + mfill_put_vma(state); + + kaddr =3D kmap_local_folio(folio, 0); + err =3D copy_from_user(kaddr, (const void __user *) src_addr, PAGE_SIZE); + kunmap_local(kaddr); + if (unlikely(err)) + return -EFAULT; + + flush_dcache_folio(folio); + + /* reget VMA and PMD, they could change underneath us */ + err =3D mfill_get_vma(state); + if (err) + return err; + + err =3D mfill_get_pmd(state); + if (err) + return err; + + return 0; +} + static int mfill_atomic_pte_copy(struct mfill_state *state) { - struct vm_area_struct *dst_vma =3D state->vma; unsigned long dst_addr =3D state->dst_addr; unsigned long src_addr =3D state->src_addr; uffd_flags_t flags =3D state->flags; - pmd_t *dst_pmd =3D state->pmd; struct folio *folio; int ret; =20 - if (!state->folio) { - ret =3D -ENOMEM; - folio =3D vma_alloc_folio(GFP_HIGHUSER_MOVABLE, 0, dst_vma, - dst_addr); - if (!folio) - goto out; + folio =3D vma_alloc_folio(GFP_HIGHUSER_MOVABLE, 0, state->vma, dst_addr); + if (!folio) + return -ENOMEM; =20 - ret =3D mfill_copy_folio_locked(folio, src_addr); + ret =3D -ENOMEM; + if (mem_cgroup_charge(folio, state->vma->vm_mm, GFP_KERNEL)) + goto out_release; =20 - /* fallback to copy_from_user outside mmap_lock */ - if (unlikely(ret)) { - ret =3D -ENOENT; - state->folio =3D folio; - /* don't free the page */ - goto out; - } - } else { - folio =3D state->folio; - state->folio =3D NULL; + ret =3D mfill_copy_folio_locked(folio, src_addr); + if (unlikely(ret)) { + /* + * Fallback to copy_from_user outside mmap_lock. + * If retry is successful, mfill_copy_folio_locked() returns + * with locks retaken by mfill_get_vma(). + * If there was an error, we must mfill_put_vma() anyway and it + * will take care of unlocking if needed. + */ + ret =3D mfill_copy_folio_retry(state, folio); + if (ret) + goto out_release; } =20 /* @@ -442,17 +473,16 @@ static int mfill_atomic_pte_copy(struct mfill_state *= state) */ __folio_mark_uptodate(folio); =20 - ret =3D -ENOMEM; - if (mem_cgroup_charge(folio, dst_vma->vm_mm, GFP_KERNEL)) - goto out_release; - - ret =3D mfill_atomic_install_pte(dst_pmd, dst_vma, dst_addr, + ret =3D mfill_atomic_install_pte(state->pmd, state->vma, dst_addr, &folio->page, true, flags); if (ret) goto out_release; out: return ret; out_release: + /* Don't return -ENOENT so that our caller won't retry */ + if (ret =3D=3D -ENOENT) + ret =3D -EFAULT; folio_put(folio); goto out; } --=20 2.51.0