From nobody Fri Jul 24 22:17:46 2026 Received: from mail-pf1-f170.google.com (mail-pf1-f170.google.com [209.85.210.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E93B843F4C3 for ; Wed, 22 Jul 2026 21:44:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756691; cv=none; b=CKN87JRjj7mY3L2cQWmOVZTVm8GPNQ4j0U8K/9+zbFmlZ9g1Y/H3rpRKrxeKUdZdmRhUmkxK6qK4CuPlHQVUw8+Zh+T7Pvr8rSL3u9ZkWI42l5Pdp3jbwNn+osdfpHHJ4do0UM3yXefptFWrKjaeUB0VSweP8Te9FNOFQWsNxYw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756691; c=relaxed/simple; bh=e05SOoAaH6rdxtJDrz7A9LUM75TsCH+ORPFdPnvT+/o=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=fOfj6UPC8loLdL5P7tAK+o7TFjQOchURURXBV1/jFd40TnESm58s8l17d42RemMcymsaDOMJkvQqZ7CvIjcpv6cPvXaDvxyamn2Zsm/I3PO+FxZjsIqW4L6FCTLMDIGBY6q9DPJpq5XBPa/eIsSwa00MSRkZ6qdIT/ByxQhXIoE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=a9znE91D; arc=none smtp.client-ip=209.85.210.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="a9znE91D" Received: by mail-pf1-f170.google.com with SMTP id d2e1a72fcca58-84867f07d63so14622118b3a.2 for ; Wed, 22 Jul 2026 14:44:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784756688; x=1785361488; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JmjPr4ajWT+pNjYdNv6NRB3T+nqQJNeM5CFnnbec2W4=; b=a9znE91DJL9XltxYXAd1baUEX69HzErOPO71bvy8Uk/dV5+RuKZHAbfYEdH3rcvvGZ zKkT/EQhG1mRbWpRRBMN8V7rFYVuwc7iQ/iZbKGutVO3F47KU4Q2S2oXzYzUyjkvlDbr drTPSAg4GIx6cH2rxrM55KPRsE+FIOQ5LpQ83NyaWupdg8HI3wqLsiR58Ff7V29Dmy8u lBIGVlJkzXOXwvJFO8TbfGm2grUMgc/2BlQ29h570uEp42JH1OZ4KtaDHg7kkEtYlRN0 LtfVE9vRmodRbzGHRuM6e0WcX4FzHG9T+hTY3/6n/b2ugIDqtVv+DUT33ZUHReFqX/+H ss+A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784756688; x=1785361488; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=JmjPr4ajWT+pNjYdNv6NRB3T+nqQJNeM5CFnnbec2W4=; b=s653s6L4zEUPJDtKUyF7QPMAGfHdhZWuap+sOuZ+bNo3IA4xne+k0jac3ENDj1orcm jG5zagzRUaIXjW4d0pJynnIpYnzsFGClLVjPg0RMdwP+abSqTBda05DJEpscP1YFIBJV p/4vDR1WQPoXznE+6pLCyBtunkGqD0YlluCn8hPNXkiV7yl4BSDgv3pppl6h5m6GU54v wrpQ93dx9K200A9YSZqe1gHMRZQt0GmMjy2ZxxPZmWpJ18vVMM0y+gGNpw19xVyxjzWC Sn2IX9+bFrt8wanCg2MmuPRi9xs+JmtksMSLoUP8VwbWcB1WoIn4l9jshSNGyrDGsfFg urTw== X-Forwarded-Encrypted: i=1; AHgh+RphgcQiSBQDM35FS9REkpf9XuOOF2TBi2uLI6Le72sfC9lKDvpMMyenyERmNyzIgBdza3XF3md82u7zRxA=@vger.kernel.org X-Gm-Message-State: AOJu0YyjZXa8XL3UH0YHTcDT2Fazf5adIk0R/clxSKO7VYJ9DNFjaPKJ Pcxtqg6772IeWl/NJczaRB+ZYoAnmDXAlxXuS1HdjiYwftXfxpEl/tVU X-Gm-Gg: AR+sD13gAR/x4qfvWqtK7idoMBOQDc1mxG+KCgQYfwqPuSBuWEI/lyUpHYgvUcOIxBF lC8Vpk2dA90Q1GU1GX4cdzMKrondyTy4NgK5+9drvO1HvFlP37fMOKRM6juk1Ov9jXSA5OV+gKM rxy+N6RAKSNt/bnJHV6LnjR3WkliESD+i4KjuMIZLgaLlKs/J8tgWA9U2R0DGNuGRk0WltpxOUC 4ThP1Pufl/Qjy9kjv5QhsuGEYlriZNnTdW63gNCbY6eIli63hJfY9i2V0q9XiV683sdUEChHmRz cArJQl5fTVrXWHG7b1xOO4C6boRU9ZGCLsc90K8B/cN1Grza3wUWjWM4Xqz7bFPs56FdMi+vQsB vQUyFu/PymR+k+Rf85/j36v71JaqqakYp6A1kNbcn0g47sCfyMraJmN9js6g7rxux32V0i9qF1f DL+TZnfIA+NHO/UjvFI7AePQsiyy/6t6jEP3VJvR+oJgGhK/CJ X-Received: by 2002:a05:6a00:174c:b0:848:2f7a:2e5e with SMTP id d2e1a72fcca58-84e2bbdcc09mr629731b3a.77.1784756688050; Wed, 22 Jul 2026 14:44:48 -0700 (PDT) Received: from [192.168.0.160] (c-98-225-44-182.hsd1.wa.comcast.net. [98.225.44.182]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84e17237c3fsm1946297b3a.9.2026.07.22.14.44.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 14:44:47 -0700 (PDT) From: Stanislav Kinsburskii Date: Wed, 22 Jul 2026 14:44:23 -0700 Subject: [PATCH v10 1/8] mm/hmm: move page fault handling out of walk callbacks Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hmm-v10-v1-1-606464dd601a@gmail.com> References: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> In-Reply-To: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> To: Jason Gunthorpe , Leon Romanovsky , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Shuah Khan , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Min Ma , Lizhi Hou , Oded Gabbay , skinsburskii@gmail.com Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-hyperv@vger.kernel.org, dri-devel@lists.freedesktop.org, nouveau@lists.freedesktop.org, linux-rdma@vger.kernel.org, Jason Gunthorpe X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784756683; l=9376; i=skinsburskii@gmail.com; s=20260722; h=from:subject:message-id; bh=e05SOoAaH6rdxtJDrz7A9LUM75TsCH+ORPFdPnvT+/o=; b=6UEcui0Mgdw7SEXvw/fnz0FGAQeRe0lav9XGzNUVUHMiVOsV2b7xmxa+kesWfMdeztUFG7U59 1Kn/V2daOYUCUD0magnsyOym2mUZFFsnubIphBn1Z2zSdggOddwVTuP X-Developer-Key: i=skinsburskii@gmail.com; a=ed25519; pk=bDpriHBYgeTdkIDweZDCemxsU93neJBOCn3YLIuJpnE= hmm_range_fault() currently triggers page faults from inside the page-table walk callbacks: hmm_vma_walk_pmd(), hmm_vma_walk_pud(), hmm_vma_walk_hugetlb_entry() and the pte-level helper all call hmm_vma_fault(), which in turn calls handle_mm_fault() while the walker still holds nested locks. The pte spinlock is dropped explicitly by each caller, and the hugetlb path manually drops and retakes hugetlb_vma_lock_read around the fault to dodge a deadlock against the walk framework's unconditional unlock. This layering does not extend cleanly to fault handlers that may release mmap_lock (VM_FAULT_RETRY, VM_FAULT_COMPLETED). If the lock is dropped while walk_page_range() is mid-traversal, the VMA can be freed before the walk framework's matching hugetlb_vma_unlock_read(), turning that unlock into a use-after-free. Split the responsibilities the way get_user_pages() does. Walk callbacks become inspect-only: when they detect a range that needs to be faulted in, they record it in struct hmm_vma_walk and return a private sentinel (HMM_FAULT_PENDING). The outer loop in hmm_range_fault() then drops out of walk_page_range(), invokes a new helper hmm_do_fault() that calls handle_mm_fault() with only mmap_lock held, and restarts the walk so the now-present entries are collected into hmm_pfns. No functional change for existing callers. As a side effect the hugetlb callback no longer needs the hugetlb_vma_{un}lock_read dance, and every fault-path exit from the callbacks now releases the pte spinlock on a single, common path. This refactor is also a precursor for adding an unlockable variant of hmm_range_fault() in a follow-up patch. Reviewed-by: Jason Gunthorpe Signed-off-by: Stanislav Kinsburskii --- mm/hmm.c | 118 ++++++++++++++++++++++++++++++++++++++++-------------------= ---- 1 file changed, 75 insertions(+), 43 deletions(-) diff --git a/mm/hmm.c b/mm/hmm.c index e5c1f4deed24..bc9361a715fa 100644 --- a/mm/hmm.c +++ b/mm/hmm.c @@ -33,8 +33,17 @@ struct hmm_vma_walk { struct hmm_range *range; unsigned long last; + unsigned long end; + unsigned int required_fault; }; =20 +/* + * Internal sentinel returned by walk callbacks when they need a page faul= t. + * The callback stores end/required_fault in hmm_vma_walk; the outer loop + * consumes the sentinel and never propagates it to the caller. + */ +#define HMM_FAULT_PENDING -EAGAIN + enum { HMM_NEED_FAULT =3D 1 << 0, HMM_NEED_WRITE_FAULT =3D 1 << 1, @@ -60,37 +69,25 @@ static int hmm_pfns_fill(unsigned long addr, unsigned l= ong end, } =20 /* - * hmm_vma_fault() - fault in a range lacking valid pmd or pte(s) - * @addr: range virtual start address (inclusive) - * @end: range virtual end address (exclusive) - * @required_fault: HMM_NEED_* flags - * @walk: mm_walk structure - * Return: -EBUSY after page fault, or page fault error + * hmm_record_fault() - record a range that needs to be faulted in * - * This function will be called whenever pmd_none() or pte_none() returns = true, - * or whenever there is no page directory covering the virtual address ran= ge. + * Called by the walk callbacks when they discover that part of the range + * needs a page fault. The callback records what to fault and returns + * HMM_FAULT_PENDING; the outer loop in hmm_range_fault() drops back out of + * walk_page_range() and invokes handle_mm_fault() from a context where no + * page-table or hugetlb_vma_lock is held. */ -static int hmm_vma_fault(unsigned long addr, unsigned long end, - unsigned int required_fault, struct mm_walk *walk) +static int hmm_record_fault(unsigned long addr, unsigned long end, + unsigned int required_fault, + struct mm_walk *walk) { struct hmm_vma_walk *hmm_vma_walk =3D walk->private; - struct vm_area_struct *vma =3D walk->vma; - unsigned int fault_flags =3D FAULT_FLAG_REMOTE; =20 WARN_ON_ONCE(!required_fault); hmm_vma_walk->last =3D addr; - - if (required_fault & HMM_NEED_WRITE_FAULT) { - if (!(vma->vm_flags & VM_WRITE)) - return -EPERM; - fault_flags |=3D FAULT_FLAG_WRITE; - } - - for (; addr < end; addr +=3D PAGE_SIZE) - if (handle_mm_fault(vma, addr, fault_flags, NULL) & - VM_FAULT_ERROR) - return -EFAULT; - return -EBUSY; + hmm_vma_walk->end =3D end; + hmm_vma_walk->required_fault =3D required_fault; + return HMM_FAULT_PENDING; } =20 static unsigned int hmm_pte_need_fault(const struct hmm_vma_walk *hmm_vma_= walk, @@ -174,7 +171,7 @@ static int hmm_vma_walk_hole(unsigned long addr, unsign= ed long end, return hmm_pfns_fill(addr, end, range, HMM_PFN_ERROR); } if (required_fault) - return hmm_vma_fault(addr, end, required_fault, walk); + return hmm_record_fault(addr, end, required_fault, walk); return hmm_pfns_fill(addr, end, range, 0); } =20 @@ -209,7 +206,7 @@ static int hmm_vma_handle_pmd(struct mm_walk *walk, uns= igned long addr, required_fault =3D hmm_range_need_fault(hmm_vma_walk, hmm_pfns, npages, cpu_flags); if (required_fault) - return hmm_vma_fault(addr, end, required_fault, walk); + return hmm_record_fault(addr, end, required_fault, walk); =20 pfn =3D pmd_pfn(pmd) + ((addr & ~PMD_MASK) >> PAGE_SHIFT); for (i =3D 0; addr < end; addr +=3D PAGE_SIZE, i++, pfn++) { @@ -328,7 +325,7 @@ static int hmm_vma_handle_pte(struct mm_walk *walk, uns= igned long addr, fault: pte_unmap(ptep); /* Fault any virtual address we were asked to fault */ - return hmm_vma_fault(addr, end, required_fault, walk); + return hmm_record_fault(addr, end, required_fault, walk); } =20 #ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES @@ -371,7 +368,7 @@ static int hmm_vma_handle_absent_pmd(struct mm_walk *wa= lk, unsigned long start, npages, 0); if (required_fault) { if (softleaf_is_device_private(entry)) - return hmm_vma_fault(addr, end, required_fault, walk); + return hmm_record_fault(addr, end, required_fault, walk); else return -EFAULT; } @@ -517,7 +514,7 @@ static int hmm_vma_walk_pud(pud_t *pudp, unsigned long = start, unsigned long end, npages, cpu_flags); if (required_fault) { spin_unlock(ptl); - return hmm_vma_fault(addr, end, required_fault, walk); + return hmm_record_fault(addr, end, required_fault, walk); } =20 pfn =3D pud_pfn(pud) + ((addr & ~PUD_MASK) >> PAGE_SHIFT); @@ -564,21 +561,8 @@ static int hmm_vma_walk_hugetlb_entry(pte_t *pte, unsi= gned long hmask, required_fault =3D hmm_pte_need_fault(hmm_vma_walk, pfn_req_flags, cpu_flags); if (required_fault) { - int ret; - spin_unlock(ptl); - hugetlb_vma_unlock_read(vma); - /* - * Avoid deadlock: drop the vma lock before calling - * hmm_vma_fault(), which will itself potentially take and - * drop the vma lock. This is also correct from a - * protection point of view, because there is no further - * use here of either pte or ptl after dropping the vma - * lock. - */ - ret =3D hmm_vma_fault(addr, end, required_fault, walk); - hugetlb_vma_lock_read(vma); - return ret; + return hmm_record_fault(addr, end, required_fault, walk); } =20 pfn =3D pte_pfn(entry) + ((start & ~hmask) >> PAGE_SHIFT); @@ -637,6 +621,44 @@ static const struct mm_walk_ops hmm_walk_ops =3D { .walk_lock =3D PGWALK_RDLOCK, }; =20 +/* + * hmm_do_fault - fault in a range recorded by a walk callback + * + * Called from the outer loop in hmm_range_fault() after a callback + * returned HMM_FAULT_PENDING. At this point we hold only mmap_lock; + * the page-table spinlock and any hugetlb_vma_lock acquired by the walk + * framework have already been released by the unwind. + * + * Returns -EBUSY on success (all pages faulted, caller should re-walk). + * Returns a negative errno on failure. + */ +static int hmm_do_fault(struct mm_struct *mm, + struct hmm_vma_walk *hmm_vma_walk) +{ + unsigned long addr =3D hmm_vma_walk->last; + unsigned long end =3D hmm_vma_walk->end; + unsigned int required_fault =3D hmm_vma_walk->required_fault; + unsigned int fault_flags =3D FAULT_FLAG_REMOTE; + struct vm_area_struct *vma; + + vma =3D vma_lookup(mm, addr); + if (!vma) + return -EFAULT; + + if (required_fault & HMM_NEED_WRITE_FAULT) { + if (!(vma->vm_flags & VM_WRITE)) + return -EPERM; + fault_flags |=3D FAULT_FLAG_WRITE; + } + + for (; addr < end; addr +=3D PAGE_SIZE) + if (handle_mm_fault(vma, addr, fault_flags, NULL) & + VM_FAULT_ERROR) + return -EFAULT; + + return -EBUSY; +} + /** * hmm_range_fault - try to fault some address in a virtual address range * @range: argument structure @@ -674,6 +696,16 @@ int hmm_range_fault(struct hmm_range *range) return -EBUSY; ret =3D walk_page_range(mm, hmm_vma_walk.last, range->end, &hmm_walk_ops, &hmm_vma_walk); + /* + * When HMM_FAULT_PENDING is returned a walk callback + * recorded a range that needs handle_mm_fault(); + * hmm_do_fault() runs the fault outside walk_page_range() + * (so no page-table or hugetlb_vma_lock is held) and + * returns -EBUSY so the loop re-walks and picks up the + * now-present entries. + */ + if (ret =3D=3D HMM_FAULT_PENDING) + ret =3D hmm_do_fault(mm, &hmm_vma_walk); /* * When -EBUSY is returned the loop restarts with * hmm_vma_walk.last set to an address that has not been stored --=20 2.43.0 From nobody Fri Jul 24 22:17:46 2026 Received: from mail-pf1-f169.google.com (mail-pf1-f169.google.com [209.85.210.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B6CB84399CB for ; Wed, 22 Jul 2026 21:44:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756693; cv=none; b=P1ZH8cREvdV+RmgB7vQadf2nnef/tgMsuKYqxFiHBi6d3A+IPD36CW3thLLTV33QQSafCYuhRMzbcf/rV7g9vEqwD0CBVQsxTYm8JC2mDh1Xrk2jDlIiZp5CrKq4MwGBYXZkc0J2zCtcFBwKh7zLXtyttKWSPVB0Pr9X72GmuNQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756693; c=relaxed/simple; bh=phNfi9QTZkEVMxoyJ4Ni2C683NZl0Qp7ozSXMuOXsZc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=OxCQ6NVnakJW2QDGEy2mCjgyQT+6yMH7tp4V9/dG7dwgDFztIJXpgckgxNm9dgIm4HOshKujI1vuAaM+kqN8eqQT8NSmtvWsB8FcpUORTlo3tx3MvIfusb6wOpoWfnXWysevhmZ9gC4yPahoHaZYae5lK8SCQ8BRDuXaVLOGRDE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=KAbJUOla; arc=none smtp.client-ip=209.85.210.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="KAbJUOla" Received: by mail-pf1-f169.google.com with SMTP id d2e1a72fcca58-848595b338cso14772518b3a.0 for ; Wed, 22 Jul 2026 14:44:50 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784756690; x=1785361490; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=qdx7YDLo+QMjlb8j72PZyp2xuVmsQJocEMpMp1wNFRQ=; b=KAbJUOla5v1vLbKVFq+UC2UqmP0DOgG2EcUo//avQ5FOytM5HenjA/gc146YA9SciW 6o5nZY/Qm9I3/wWA+JsgAEW/p4lFQZnghFjH17PXpnm5utPlECn8pQyWiz4JEda5uVkV rhOIXx6V+BHoAprx1Q+e0AaD4YjXebxSN4Y1G196wk4jtzHZ9gxWzeFFmpySq1S2Z1Xf Btju9zFKFDvRMUeeau94ws40rzFXwEjrVu5oSWDbmprLBsXAiht5Da3hFSu64gIOma9L 7ZW+nLtUdQBRSluBj52mjuFNzLx36Z6YZnVA3Qc0/oyQfhiUgQY9GQzYtmnmC1yFTDLt C8EQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784756690; x=1785361490; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=qdx7YDLo+QMjlb8j72PZyp2xuVmsQJocEMpMp1wNFRQ=; b=E08J7smuYQWVzUrShyDxHmf1v2088lz9RnV4FzWwhnL5HSWz/kHP1IrozyHpuyMn3W 4g45rJko4qUZ0GswsK/XvgK4i4QFk1J3GKgk9VBNL7bIWeWMF+5qe2pZ4KaRPJBWlLLn Ct+3neCWl8/kcw3WquPZ2m+8eZa5JmQyOqrtz1ZkDWBRiZqzvjCxtlwRHYrYH5mu3e0l hlwSgDsLHBoz+/Qe7/uhAshj6Gj2ZouP76z6oS53pkUI30/6IKKlvePywR2jltaMdTDZ Ru4fWAOQEivwzMeBiEErMp4cCKbXB0ZI7xA3YoX4MNhTTb6PhtD35wtJpUAYAc6ACPwQ K1vA== X-Forwarded-Encrypted: i=1; AHgh+Rq2Y9zi6NmKnJaadbd/G8JmlPWIoyDMnZdlmgAjT5Lbmy51CJhkEXkcUCzD8GcaQIl2GLM1fPrb3KNu1yw=@vger.kernel.org X-Gm-Message-State: AOJu0YxA1bV0D5nSAEFlSh0gw4WgtpVgw3c7zfeTDmvsr6ZG8UCI2IEl jfi9FGK4j/unXcdQVGmdzUEAiVgA804kgjlSyTzvepFgf1hDWaAXpg1j X-Gm-Gg: AR+sD12imWDV3LunsaJNsWC+9F3QGUSjj1huJkZSuMPQ9wDyotLAU24sfE2CAExaFOD mwyzVBBynK9ZMhIJxnVEIjtWV0Nep7CUm25Rc52i6dGyROf9RhE+V+NdWrx39Rf7kCGq3TskF39 PlsWuPsSjDn31B2pLGTjvVI/p6k3sN1hVaxVzpPnymLlA+o2T4JT5yCj9N6Dbw5ADzb2DDTxNLb GM6bSTP72900HfX3OLCTGhUX3m9kO8mdFdNPu2BKsKkKheAmt4cmCoSRUDw0siLakkzcWiaGYHK gUXU7VFIlFeG5VfYIYpKt6DQHORhrBn5yE93on5LMW8jSCOl7yCNawHX9paqempz6e3682dbIgh n/4lPHXdrw03lDRxN803FMwq+IBMMCw8oHjr/akIKPFHhhngjHhIJGoQVnGh0lE56oLwQXWC2Uy QPeiZLV80pI0Ek7YdPlv1bC2TwBBoeMd+nPuVaB1CFVBIJtpQ1 X-Received: by 2002:a05:6a00:3d12:b0:847:98ff:4af5 with SMTP id d2e1a72fcca58-84e2b8c0da6mr711681b3a.26.1784756689863; Wed, 22 Jul 2026 14:44:49 -0700 (PDT) Received: from [192.168.0.160] (c-98-225-44-182.hsd1.wa.comcast.net. [98.225.44.182]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84e17237c3fsm1946297b3a.9.2026.07.22.14.44.48 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 14:44:49 -0700 (PDT) From: Stanislav Kinsburskii Date: Wed, 22 Jul 2026 14:44:24 -0700 Subject: [PATCH v10 2/8] mm/hmm: add hmm_range_fault_unlocked_timeout() for mmap lock-drop support Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hmm-v10-v1-2-606464dd601a@gmail.com> References: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> In-Reply-To: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> To: Jason Gunthorpe , Leon Romanovsky , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Shuah Khan , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Min Ma , Lizhi Hou , Oded Gabbay , skinsburskii@gmail.com Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-hyperv@vger.kernel.org, dri-devel@lists.freedesktop.org, nouveau@lists.freedesktop.org, linux-rdma@vger.kernel.org X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784756683; l=17250; i=skinsburskii@gmail.com; s=20260722; h=from:subject:message-id; bh=phNfi9QTZkEVMxoyJ4Ni2C683NZl0Qp7ozSXMuOXsZc=; b=kdDY7nuq9dEb1iZ4WbkS/DHv4ppZRm5lPwJw6yGAjbvELwTl1EY7EymTqTd5YgPoTrv7RCNjb 88j94dyszYkAJMGpJIEF8Ri3EHQCYEvc2wgoWxeNxsDxqo9XW7LOnwV X-Developer-Key: i=skinsburskii@gmail.com; a=ed25519; pk=bDpriHBYgeTdkIDweZDCemxsU93neJBOCn3YLIuJpnE= hmm_range_fault() requires the caller to hold the mmap read lock for the duration of the call. This is incompatible with mappings whose fault handler may release the mmap lock, notably userfaultfd-managed regions, where handle_mm_fault() can return VM_FAULT_RETRY or VM_FAULT_COMPLETED after dropping the lock. Drivers that need to populate device page tables for such mappings have no way to do so today. Add hmm_range_fault_unlocked_timeout() for callers that do not need to hold mmap_lock across any work outside the HMM fault itself. The helper takes mmap_read_lock_killable() internally, calls the common HMM fault implementation, and releases the lock before returning if it is still held. The timeout is specified in jiffies; passing 0 retries indefinitely, while a non-zero timeout makes the helper return -EBUSY when the retry budget expires. The retry deadline is set before refreshing the notifier sequence and acquiring mmap_lock, so contended mmap_lock acquisition is included in the retry budget. When handle_mm_fault() drops mmap_lock, or when the range is invalidated, hmm_range_fault_unlocked_timeout() refreshes range->notifier_seq and retries the walk internally. If the lock was dropped, the retry deadline is also restarted because a lock-dropping fault handler made progress. Ordinary -EBUSY retries keep the existing deadline, preserving the caller's timeout policy for repeated mmu-notifier invalidations. The caller only needs to perform the usual post-success mmu_interval_read_retry() check while holding its update lock before consuming the pfns. If mmap_lock acquisition is interrupted or a fatal signal is pending during retry handling, -EINTR is returned instead. The common implementation conditionally sets FAULT_FLAG_ALLOW_RETRY and FAULT_FLAG_KILLABLE only for hmm_range_fault_unlocked_timeout(). The existing hmm_range_fault() path still passes no locked state, does not allow handle_mm_fault() to drop mmap_lock, and remains a thin wrapper preserving the existing API contract for current callers. The previous refactor that moved page fault handling out of the page-table walk callbacks is what makes this change small. Faults now run after walk_page_range() has unwound, with only mmap_lock held, so dropping it does not interact with the walker's pte spinlock or hugetlb_vma_lock. Hugetlb regions therefore participate in the unlocked path uniformly with PTE- and PMD-level mappings; no special case is required. Documentation/mm/hmm.rst is updated with a description of the new API and the recommended caller pattern. Signed-off-by: Stanislav Kinsburskii --- Documentation/mm/hmm.rst | 79 ++++++++++++++++------- include/linux/hmm.h | 2 + mm/hmm.c | 162 ++++++++++++++++++++++++++++++++++++++-----= ---- 3 files changed, 191 insertions(+), 52 deletions(-) diff --git a/Documentation/mm/hmm.rst b/Documentation/mm/hmm.rst index 7d61b7a8b65b..e021218ada58 100644 --- a/Documentation/mm/hmm.rst +++ b/Documentation/mm/hmm.rst @@ -156,42 +156,57 @@ During the ops->invalidate() callback the device driv= er must perform the update action to the range (mark range read only, or fully unmap, etc.). T= he device must complete the update before the driver callback returns. =20 -When the device driver wants to populate a range of virtual addresses, it = can -use:: +When the device driver wants to populate a range of virtual addresses, the +normal interface is:: =20 - int hmm_range_fault(struct hmm_range *range); + int hmm_range_fault_unlocked_timeout(struct hmm_range *range, + unsigned long timeout); =20 It will trigger a page fault on missing or read-only entries if write acce= ss is requested (see below). Page faults use the generic mm page fault code path= just -like a CPU page fault. The usage pattern is:: +like a CPU page fault. + +The caller must not hold ``mmap_read_lock`` before the call. +``hmm_range_fault_unlocked_timeout()`` takes the mmap read lock internally= and +allows ``handle_mm_fault()`` to drop it during fault handling. This is req= uired +for VMAs whose fault handlers may release the mmap lock, for example regio= ns +managed by ``userfaultfd``. + +If the mmap lock is dropped or the range is invalidated, the function refr= eshes +``range->notifier_seq`` and restarts the walk internally. ``-EINTR`` is re= turned +if mmap lock acquisition is interrupted or a fatal signal is pending during +retry handling. + +The timeout is specified in jiffies; passing ``0`` means retry indefinitel= y. The +timeout exists to preserve caller policy for repeated mmu-notifier invalid= ation +and is checked between retry attempts. HMM does not interrupt page fault +handling when the timeout expires, but returns ``-EBUSY`` if the retry bud= get is +exhausted before a stable range is obtained. + +The usage pattern is:: =20 int driver_populate_range(...) { struct hmm_range range; + unsigned long timeout; ... =20 + timeout =3D msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); range.notifier =3D &interval_sub; range.start =3D ...; range.end =3D ...; range.hmm_pfns =3D ...; =20 - if (!mmget_not_zero(interval_sub->notifier.mm)) + if (!mmget_not_zero(interval_sub.mm)) return -EFAULT; =20 again: - range.notifier_seq =3D mmu_interval_read_begin(&interval_sub); - mmap_read_lock(mm); - ret =3D hmm_range_fault(&range); - if (ret) { - mmap_read_unlock(mm); - if (ret =3D=3D -EBUSY) - goto again; - return ret; - } - mmap_read_unlock(mm); + ret =3D hmm_range_fault_unlocked_timeout(&range, timeout); + if (ret) + goto out_put; =20 take_lock(driver->update); - if (mmu_interval_read_retry(&ni, range.notifier_seq) { + if (mmu_interval_read_retry(range.notifier, range.notifier_seq)) { release_lock(driver->update); goto again; } @@ -200,13 +215,31 @@ like a CPU page fault. The usage pattern is:: * under the update lock */ =20 release_lock(driver->update); - return 0; + ret =3D 0; + + out_put: + mmput(interval_sub.mm); + return ret; } =20 The driver->update lock is the same lock that the driver takes inside its invalidate() callback. That lock must be held before calling mmu_interval_read_retry() to avoid any race with a concurrent CPU page tab= le -update. +update. The retry check must use the same notifier and sequence number sto= red +in ``range`` by ``hmm_range_fault_unlocked_timeout()``. + +Holding the mmap lock across HMM faults +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Most callers should use ``hmm_range_fault_unlocked_timeout()``. If a driver +really needs to hold the mmap lock across work outside HMM, it can use:: + + int hmm_range_fault(struct hmm_range *range); + +The mmap lock must be held by the caller and will remain held on return. T= his +interface cannot support VMAs whose fault handlers need to drop the mmap l= ock. +New callers should prefer ``hmm_range_fault_unlocked_timeout()`` unless th= ey +have a specific requirement to keep the mmap lock held across the call. =20 Leverage default_flags and pfn_flags_mask =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D @@ -221,8 +254,8 @@ permission, it sets:: range->default_flags =3D HMM_PFN_REQ_FAULT; range->pfn_flags_mask =3D 0; =20 -and calls hmm_range_fault() as described above. This will fill fault all p= ages -in the range with at least read permission. +and calls the HMM range fault helper as described above. This will fault +all pages in the range with at least read permission. =20 Now let's say the driver wants to do the same except for one page in the r= ange for which it wants to have write permission. Now driver set:: @@ -236,9 +269,9 @@ address =3D=3D range->start + (index_of_write << PAGE_S= HIFT) it will fault with write permission i.e., if the CPU pte does not have write permission set t= hen HMM will call handle_mm_fault(). =20 -After hmm_range_fault completes the flag bits are set to the current state= of -the page tables, ie HMM_PFN_VALID | HMM_PFN_WRITE will be set if the page = is -writable. +After the HMM range fault helper completes the flag bits are set to the +current state of the page tables, ie HMM_PFN_VALID | HMM_PFN_WRITE will be +set if the page is writable. =20 =20 Represent and manage device memory from core kernel point of view diff --git a/include/linux/hmm.h b/include/linux/hmm.h index db75ffc949a7..6f04e3932f5b 100644 --- a/include/linux/hmm.h +++ b/include/linux/hmm.h @@ -123,6 +123,8 @@ struct hmm_range { * Please see Documentation/mm/hmm.rst for how to use the range API. */ int hmm_range_fault(struct hmm_range *range); +int hmm_range_fault_unlocked_timeout(struct hmm_range *range, + unsigned long timeout); =20 /* * HMM_RANGE_DEFAULT_TIMEOUT - default timeout (ms) when waiting for a ran= ge diff --git a/mm/hmm.c b/mm/hmm.c index bc9361a715fa..87952c1259c5 100644 --- a/mm/hmm.c +++ b/mm/hmm.c @@ -32,6 +32,7 @@ =20 struct hmm_vma_walk { struct hmm_range *range; + bool *locked; unsigned long last; unsigned long end; unsigned int required_fault; @@ -44,6 +45,14 @@ struct hmm_vma_walk { */ #define HMM_FAULT_PENDING -EAGAIN =20 +/* + * Internal sentinel returned by hmm_do_fault() when handle_mm_fault() + * completes a page fault with the mmap lock dropped. hmm_do_fault() sets + * *locked =3D false; the outer loop consumes the sentinel and never propa= gates + * it to the caller. + */ +#define HMM_FAULT_UNLOCKED -ENOLCK + enum { HMM_NEED_FAULT =3D 1 << 0, HMM_NEED_WRITE_FAULT =3D 1 << 1, @@ -73,9 +82,9 @@ static int hmm_pfns_fill(unsigned long addr, unsigned lon= g end, * * Called by the walk callbacks when they discover that part of the range * needs a page fault. The callback records what to fault and returns - * HMM_FAULT_PENDING; the outer loop in hmm_range_fault() drops back out of - * walk_page_range() and invokes handle_mm_fault() from a context where no - * page-table or hugetlb_vma_lock is held. + * HMM_FAULT_PENDING; the outer loop in hmm_range_fault_locked() drops + * back out of walk_page_range() and invokes handle_mm_fault() from a cont= ext + * where no page-table or hugetlb_vma_lock is held. */ static int hmm_record_fault(unsigned long addr, unsigned long end, unsigned int required_fault, @@ -624,7 +633,7 @@ static const struct mm_walk_ops hmm_walk_ops =3D { /* * hmm_do_fault - fault in a range recorded by a walk callback * - * Called from the outer loop in hmm_range_fault() after a callback + * Called from the outer loop in hmm_range_fault_locked() after a callback * returned HMM_FAULT_PENDING. At this point we hold only mmap_lock; * the page-table spinlock and any hugetlb_vma_lock acquired by the walk * framework have already been released by the unwind. @@ -641,6 +650,9 @@ static int hmm_do_fault(struct mm_struct *mm, unsigned int fault_flags =3D FAULT_FLAG_REMOTE; struct vm_area_struct *vma; =20 + if (hmm_vma_walk->locked) + fault_flags |=3D FAULT_FLAG_ALLOW_RETRY | FAULT_FLAG_KILLABLE; + vma =3D vma_lookup(mm, addr); if (!vma) return -EFAULT; @@ -651,37 +663,34 @@ static int hmm_do_fault(struct mm_struct *mm, fault_flags |=3D FAULT_FLAG_WRITE; } =20 - for (; addr < end; addr +=3D PAGE_SIZE) - if (handle_mm_fault(vma, addr, fault_flags, NULL) & - VM_FAULT_ERROR) - return -EFAULT; + for (; addr < end; addr +=3D PAGE_SIZE) { + vm_fault_t ret; + + ret =3D handle_mm_fault(vma, addr, fault_flags, NULL); + + if (ret & (VM_FAULT_COMPLETED | VM_FAULT_RETRY)) { + *hmm_vma_walk->locked =3D false; + return HMM_FAULT_UNLOCKED; + } + + if (ret & VM_FAULT_ERROR) { + int err =3D vm_fault_to_errno(ret, 0); + + if (WARN_ON(!err)) + err =3D -EINVAL; + + return err; + } + } =20 return -EBUSY; } =20 -/** - * hmm_range_fault - try to fault some address in a virtual address range - * @range: argument structure - * - * Returns 0 on success or one of the following error codes: - * - * -EINVAL: Invalid arguments or mm or virtual address is in an invalid vma - * (e.g., device file vma). - * -ENOMEM: Out of memory. - * -EPERM: Invalid permission (e.g., asking for write and range is read - * only). - * -EBUSY: The range has been invalidated and the caller needs to wait for - * the invalidation to finish. - * -EFAULT: A page was requested to be valid and could not be made val= id - * ie it has no backing VMA or it is illegal to access - * - * This is similar to get_user_pages(), except that it can read the page t= ables - * without mutating them (ie causing faults). - */ -int hmm_range_fault(struct hmm_range *range) +static int hmm_range_fault_locked(struct hmm_range *range, bool *locked) { struct hmm_vma_walk hmm_vma_walk =3D { .range =3D range, + .locked =3D locked, .last =3D range->start, }; struct mm_struct *mm =3D range->notifier->mm; @@ -704,8 +713,14 @@ int hmm_range_fault(struct hmm_range *range) * returns -EBUSY so the loop re-walks and picks up the * now-present entries. */ - if (ret =3D=3D HMM_FAULT_PENDING) + if (ret =3D=3D HMM_FAULT_PENDING) { ret =3D hmm_do_fault(mm, &hmm_vma_walk); + if (ret =3D=3D HMM_FAULT_UNLOCKED) { + if (fatal_signal_pending(current)) + return -EINTR; + return -EBUSY; + } + } /* * When -EBUSY is returned the loop restarts with * hmm_vma_walk.last set to an address that has not been stored @@ -715,8 +730,97 @@ int hmm_range_fault(struct hmm_range *range) } while (ret =3D=3D -EBUSY); return ret; } + +/** + * hmm_range_fault - try to fault some address in a virtual address range + * @range: argument structure + * + * Returns 0 on success or one of the following error codes: + * + * -EINVAL: Invalid arguments or mm or virtual address is in an invalid vma + * (e.g., device file vma). + * -ENOMEM: Out of memory. + * -EPERM: Invalid permission (e.g., asking for write and range is read + * only). + * -EBUSY: The range has been invalidated and the caller needs to wait for + * the invalidation to finish. + * -EFAULT: A page was requested to be valid and could not be made val= id + * ie it has no backing VMA or it is illegal to access + * + * This is similar to get_user_pages(), except that it can read the page t= ables + * without mutating them (ie causing faults). + * + * The mmap lock must be held by the caller and will remain held on return. + * New users should prefer hmm_range_fault_unlocked_timeout() unless they + * specifically need to keep the mmap lock held across the call. This help= er + * cannot support VMAs whose fault handlers need to drop the mmap lock. + */ +int hmm_range_fault(struct hmm_range *range) +{ + return hmm_range_fault_locked(range, NULL); +} EXPORT_SYMBOL(hmm_range_fault); =20 +/** + * hmm_range_fault_unlocked_timeout - fault in a range with a retry timeout + * @range: argument structure + * @timeout: timeout in jiffies for internal -EBUSY retries, or 0 to retry + * indefinitely + * + * The caller must not hold the mmap lock. The function takes the mmap read + * lock internally and allows handle_mm_fault() to drop it during faults. = If + * the mmap lock is dropped or the range is invalidated, the function refr= eshes + * range->notifier_seq and restarts the walk internally. + * + * Passing 0 for @timeout retries indefinitely. A non-zero @timeout is a c= aller + * policy limit for repeated mmu-notifier invalidation retries. HMM does n= ot + * interrupt page fault handling when the timeout expires, but returns -EB= USY + * if the retry budget is exhausted before a stable range is obtained. + * + * Returns 0 on success or one of the error codes documented for + * hmm_range_fault(). -EINTR is returned if mmap_lock acquisition is + * interrupted or a fatal signal is pending during retry handling. + */ +int hmm_range_fault_unlocked_timeout(struct hmm_range *range, + unsigned long timeout) +{ + struct mm_struct *mm =3D range->notifier->mm; + unsigned long deadline =3D 0; + bool locked =3D false; + int ret; + + do { + /* + * If the previous fault dropped mmap_lock, then the fault + * handler made progress. Restart the retry timeout in that + * case, but keep the existing deadline for ordinary -EBUSY + * retries. + */ + if (timeout && !locked) + deadline =3D jiffies + timeout; + + range->notifier_seq =3D + mmu_interval_read_begin(range->notifier); + + ret =3D mmap_read_lock_killable(mm); + if (ret) + return ret; + + if (timeout && time_after(jiffies, deadline)) { + mmap_read_unlock(mm); + return -EBUSY; + } + + locked =3D true; + ret =3D hmm_range_fault_locked(range, &locked); + if (locked) + mmap_read_unlock(mm); + } while (ret =3D=3D -EBUSY); + + return ret; +} +EXPORT_SYMBOL(hmm_range_fault_unlocked_timeout); + /** * hmm_dma_map_alloc - Allocate HMM map structure * @dev: device to allocate structure for --=20 2.43.0 From nobody Fri Jul 24 22:17:46 2026 Received: from mail-pf1-f179.google.com (mail-pf1-f179.google.com [209.85.210.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6688B426ECF for ; Wed, 22 Jul 2026 21:44:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756696; cv=none; b=Eqxk7R7Rp5ApetCmB3SOZJ7wvw3a8dNgTKL9Suvz2GBHETJ1tLJpBbLh+mAtmVIHgrkoLGt8GBOZzjNM8YpIDPCiAuOoiNpbXF7QQMZaOSuNo2ncj4D1Jvu8UPWGRfhMUUmp2D9pQRt/P8feb5a50riQIG5ylExZ0yO9ugt2swU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756696; c=relaxed/simple; bh=AMv7G8JGzmbmB+WwCDHthj1N45IFPqE+bBj6HkmdaG8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=GJljBNH3o1SYwayVLRN0v+tSX4poP+sfzEjhdMneAFhOLHs9Qv/R8yXIhTqaFAOcxO2VkbeDZn6N8MEaBYKNNgpXQJeDjHHazIjMbOmsZqR+j6AWwqqrJEYPS0I+dImUBMdLknTko0qDHZZHvVFYXsYf2ihi8JX4xFiaj5RJ1TE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=JsiCibWu; arc=none smtp.client-ip=209.85.210.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="JsiCibWu" Received: by mail-pf1-f179.google.com with SMTP id d2e1a72fcca58-8487b7b3fc8so13327886b3a.3 for ; Wed, 22 Jul 2026 14:44:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784756692; x=1785361492; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=3cb8/RR6ht/AFgcmZ+ZVlj45R/ECVd/hJBDF8v5cizE=; b=JsiCibWuZ6Mu2dxBfJgq0iBT8KQcyA2JXHjIOcQMS50BQLm3wraD9babzpM9u65Jk+ ouzA1YuMcmfw9ibALTGf4y0ranJpCs+KJTAOPVOenqSlfoQGjOAaK97iuL3wcB4rdxRw Kg5fglp5s6w1eNpkET60p+5m7bgZ5KTVlBNfLSyADlQ6R3qQvpBleNdfWG0h9ck8t4g8 KqMXeNCvccBsSc0+iwMbUxXJ2lmV+dOJfIhr07bDUcWlJsHCjT32B72JFmN1u+qUhy+0 GkgkO3MGXUk23GrNEcoWTOLawAp54SNWlPWJW21icx+T2JAm22iVUo4X1v5idpo0dn3u YlHA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784756692; x=1785361492; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=3cb8/RR6ht/AFgcmZ+ZVlj45R/ECVd/hJBDF8v5cizE=; b=RjWPHukZXVMpx5sOt2VSx83HuRZ/VlZCH0I6CfdQP5n4vkTbE1LfUpD/0a4Ng+NXep 94Lrqzk5n4DM0JW+2KhJgiXMZMea+mNk+sIXL/hlAehwiGKFHjDIrqHQQfCd5m3BmlhU 1+q0tHVMlolSrrj6LJG2T8EAGHh1mD+FWl2mUI1qaa+FrGsYvY4T8hwq+D/AcCb9T44w RxN3sBHTf53dsqY98WnNPpw1MGwPnEA9qoVi38HTH0KCjUasx3fCN8rwlP793NfA62Jq 8Nju5J8Uz2zWiUmY63h9I8VX9gt1JXxbPvthF1rOx3YqAbblpo+6L6Sd19AZF4EPxn27 QJ1Q== X-Forwarded-Encrypted: i=1; AHgh+RpPWJcmaqQXRHFUS7Gyd4krDoTqW1Qkqy/rf9c8v0rMBQ553lrxLKfpMSRY08YmeYYZXDoHnr+CUtmusGM=@vger.kernel.org X-Gm-Message-State: AOJu0YzalKZLoV6jUHwK1ZZIVB9jt/PymWPQahamWSZ0rw7Pn+ec0zHE JtH8de9Ie1AjBcscOqsLqOZCD+Du3CgIj5ZMz1AEvdf+EEq49wZz62tD X-Gm-Gg: AR+sD13CocjGBsT4uaz5zuB2C1U7e3shOnply+qlkRMvQ4wIGLvh5CaBdRQXEXP2Ecm DVh/J5/RNoGVhvjsnY1xQ9elqeDzVuPc84/3/ZJOEKpMHgq3Wdb8nq+E1SycM2MhccigIib65Wm bVDo+hZf8SjnpqAm65BdFmurjiI2i50zZe9RxvTiM+wy216q4ImDTld+FRgMD6frAveAAUu1rsX dUSB8f55rfIvm5u/81vc1uM4H+sWE44V3BfI4PTsFeggBOBDYmVTyMRSb9NiPiAtJ43EeXW+7X8 7X9jaIpZb6EdWKTsVZpvxJQyZ1VPY4xDCNjQfxQvbEueZ4Pyls70pjYUbcch3/s+2GNJ0nknyJC qQI0QbCfSc4egs5oONUTCBzZMXBT6KuJB6DkAP0T66qsuxypHLpU40HRGLTSviEBdEnKh5gL+BH WncBKvzg2CDy/7EUMt86xDH4c8KwU4ob2GeiGG0qT0iLKMbQQDu6WGNwwMbN0= X-Received: by 2002:a05:6a00:39a7:b0:848:7835:bbac with SMTP id d2e1a72fcca58-84e2c230baemr657561b3a.65.1784756691553; Wed, 22 Jul 2026 14:44:51 -0700 (PDT) Received: from [192.168.0.160] (c-98-225-44-182.hsd1.wa.comcast.net. [98.225.44.182]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84e17237c3fsm1946297b3a.9.2026.07.22.14.44.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 14:44:51 -0700 (PDT) From: Stanislav Kinsburskii Date: Wed, 22 Jul 2026 14:44:25 -0700 Subject: [PATCH v10 3/8] selftests/mm: add HMM test for mmap lock-dropping faults Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hmm-v10-v1-3-606464dd601a@gmail.com> References: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> In-Reply-To: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> To: Jason Gunthorpe , Leon Romanovsky , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Shuah Khan , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Min Ma , Lizhi Hou , Oded Gabbay , skinsburskii@gmail.com Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-hyperv@vger.kernel.org, dri-devel@lists.freedesktop.org, nouveau@lists.freedesktop.org, linux-rdma@vger.kernel.org X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784756683; l=9280; i=skinsburskii@gmail.com; s=20260722; h=from:subject:message-id; bh=AMv7G8JGzmbmB+WwCDHthj1N45IFPqE+bBj6HkmdaG8=; b=jxDT00qgIp0eckAnCieFk/yrH4UjraxIJmUAkidlQfO9sY38TOaO5atlFw18CN2G75txgsawI w9ZNh/wRC0pA9t3+i2E84MWvjLWEgZoL8WEkW8uvNLBA7ijJ2Vro8vy X-Developer-Key: i=skinsburskii@gmail.com; a=ed25519; pk=bDpriHBYgeTdkIDweZDCemxsU93neJBOCn3YLIuJpnE= Add test_hmm coverage for the HMM lock-dropping fault path. The test module gets a new HMM_DMIRROR_READ_UNLOCKED ioctl that calls hmm_range_fault_unlocked_timeout() with a timeout of 0, exercising the unbounded retry mode while allowing the mmap lock to be dropped during fault handling. Add a userfaultfd_read selftest that registers an anonymous mapping with UFFDIO_REGISTER_MODE_MISSING, services the faults from a handler thread with UFFDIO_COPY, and verifies that HMM can read back the data supplied by the handler. This exercises the path where handle_mm_fault() drops mmap_lock and hmm_range_fault_unlocked_timeout() restarts the walk internally. Assisted-by: GitHub-Copilot:claude-opus-4.6 Signed-off-by: Stanislav Kinsburskii --- lib/test_hmm.c | 107 ++++++++++++++++++++++- lib/test_hmm_uapi.h | 1 + tools/testing/selftests/mm/hmm-tests.c | 150 +++++++++++++++++++++++++++++= ++++ 3 files changed, 257 insertions(+), 1 deletion(-) diff --git a/lib/test_hmm.c b/lib/test_hmm.c index 45c0cb992218..6205fb313bd0 100644 --- a/lib/test_hmm.c +++ b/lib/test_hmm.c @@ -389,6 +389,67 @@ static int dmirror_range_fault(struct dmirror *dmirror, return ret; } =20 +static int dmirror_range_fault_unlocked(struct dmirror *dmirror, + struct hmm_range *range, + unsigned long timeout) +{ + int ret; + + while (true) { + ret =3D hmm_range_fault_unlocked_timeout(range, timeout); + if (ret) + goto out; + + mutex_lock(&dmirror->mutex); + if (mmu_interval_read_retry(range->notifier, + range->notifier_seq)) { + mutex_unlock(&dmirror->mutex); + continue; + } + break; + } + + ret =3D dmirror_do_fault(dmirror, range); + + mutex_unlock(&dmirror->mutex); +out: + return ret; +} + +static int dmirror_fault_unlocked(struct dmirror *dmirror, + unsigned long start, + unsigned long end, bool write, + unsigned long timeout) +{ + struct mm_struct *mm =3D dmirror->notifier.mm; + unsigned long addr; + unsigned long pfns[32]; + struct hmm_range range =3D { + .notifier =3D &dmirror->notifier, + .hmm_pfns =3D pfns, + .pfn_flags_mask =3D 0, + .default_flags =3D + HMM_PFN_REQ_FAULT | (write ? HMM_PFN_REQ_WRITE : 0), + .dev_private_owner =3D dmirror->mdevice, + }; + int ret =3D 0; + + if (!mmget_not_zero(mm)) + return -EFAULT; + + for (addr =3D start; addr < end; addr =3D range.end) { + range.start =3D addr; + range.end =3D min(addr + (ARRAY_SIZE(pfns) << PAGE_SHIFT), end); + + ret =3D dmirror_range_fault_unlocked(dmirror, &range, timeout); + if (ret) + break; + } + + mmput(mm); + return ret; +} + static int dmirror_fault(struct dmirror *dmirror, unsigned long start, unsigned long end, bool write) { @@ -488,6 +549,48 @@ static int dmirror_read(struct dmirror *dmirror, struc= t hmm_dmirror_cmd *cmd) return ret; } =20 +static int dmirror_read_unlocked(struct dmirror *dmirror, + struct hmm_dmirror_cmd *cmd, + unsigned long timeout) +{ + struct dmirror_bounce bounce; + unsigned long start, end; + unsigned long size =3D cmd->npages << PAGE_SHIFT; + int ret; + + start =3D cmd->addr; + end =3D start + size; + if (end < start) + return -EINVAL; + + ret =3D dmirror_bounce_init(&bounce, start, size); + if (ret) + return ret; + + while (1) { + mutex_lock(&dmirror->mutex); + ret =3D dmirror_do_read(dmirror, start, end, &bounce); + mutex_unlock(&dmirror->mutex); + if (ret !=3D -ENOENT) + break; + + start =3D cmd->addr + (bounce.cpages << PAGE_SHIFT); + ret =3D dmirror_fault_unlocked(dmirror, start, end, false, timeout); + if (ret) + break; + cmd->faults++; + } + + if (ret =3D=3D 0) { + if (copy_to_user(u64_to_user_ptr(cmd->ptr), bounce.ptr, + bounce.size)) + ret =3D -EFAULT; + } + cmd->cpages =3D bounce.cpages; + dmirror_bounce_fini(&bounce); + return ret; +} + static int dmirror_do_write(struct dmirror *dmirror, unsigned long start, unsigned long end, struct dmirror_bounce *bounce) { @@ -1572,7 +1675,9 @@ static long dmirror_fops_unlocked_ioctl(struct file *= filp, dmirror->flags =3D cmd.npages; ret =3D 0; break; - + case HMM_DMIRROR_READ_UNLOCKED: + ret =3D dmirror_read_unlocked(dmirror, &cmd, 0); + break; default: return -EINVAL; } diff --git a/lib/test_hmm_uapi.h b/lib/test_hmm_uapi.h index f94c6d457338..ea9b0ec404fb 100644 --- a/lib/test_hmm_uapi.h +++ b/lib/test_hmm_uapi.h @@ -38,6 +38,7 @@ struct hmm_dmirror_cmd { #define HMM_DMIRROR_CHECK_EXCLUSIVE _IOWR('H', 0x06, struct hmm_dmirror_cm= d) #define HMM_DMIRROR_RELEASE _IOWR('H', 0x07, struct hmm_dmirror_cmd) #define HMM_DMIRROR_FLAGS _IOWR('H', 0x08, struct hmm_dmirror_cmd) +#define HMM_DMIRROR_READ_UNLOCKED _IOWR('H', 0x09, struct hmm_dmirror_cmd) =20 #define HMM_DMIRROR_FLAG_FAIL_ALLOC (1ULL << 0) =20 diff --git a/tools/testing/selftests/mm/hmm-tests.c b/tools/testing/selftes= ts/mm/hmm-tests.c index 6fccbdab02ee..5acb728666f8 100644 --- a/tools/testing/selftests/mm/hmm-tests.c +++ b/tools/testing/selftests/mm/hmm-tests.c @@ -29,6 +29,10 @@ #include #include #include +#include +#include +#include +#include =20 /* * This is a private UAPI to the kernel test module so it isn't exported @@ -2952,4 +2956,150 @@ TEST_F_TIMEOUT(hmm, benchmark_thp_migration, 120) &thp_results, ®ular_results); } } +/* + * Test that HMM can fault in pages backed by userfaultfd using the + * hmm_range_fault_unlocked_timeout() path with no timeout. This exercises + * the lock-drop retry logic in the HMM framework. + */ +struct uffd_thread_args { + int uffd; + int stop_fd; + void *page_buffer; + unsigned long page_size; +}; + +static void *uffd_handler_thread(void *arg) +{ + struct uffd_thread_args *args =3D arg; + struct uffd_msg msg; + struct uffdio_copy copy; + struct pollfd pollfd[2]; + int ret; + + pollfd[0].fd =3D args->uffd; + pollfd[0].events =3D POLLIN; + pollfd[1].fd =3D args->stop_fd; + pollfd[1].events =3D POLLIN; + + while (1) { + ret =3D poll(pollfd, 2, -1); + if (ret <=3D 0) + break; + if (pollfd[1].revents) + break; + if (!(pollfd[0].revents & POLLIN)) + break; + + ret =3D read(args->uffd, &msg, sizeof(msg)); + if (ret !=3D sizeof(msg)) + break; + + if (msg.event !=3D UFFD_EVENT_PAGEFAULT) + break; + + /* Fill the page with a known pattern */ + memset(args->page_buffer, 0xAB, args->page_size); + + copy.dst =3D msg.arg.pagefault.address & ~(args->page_size - 1); + copy.src =3D (unsigned long)args->page_buffer; + copy.len =3D args->page_size; + copy.mode =3D 0; + copy.copy =3D 0; + + ret =3D ioctl(args->uffd, UFFDIO_COPY, ©); + if (ret < 0) + break; + } + + return NULL; +} + +TEST_F(hmm, userfaultfd_read) +{ + struct hmm_buffer *buffer; + struct uffd_thread_args uffd_args; + unsigned long npages; + unsigned long size; + unsigned long i; + unsigned char *ptr; + pthread_t thread; + int uffd; + int stop_fd; + int ret; + struct uffdio_api api; + struct uffdio_register reg; + uint64_t stop =3D 1; + ssize_t nwrite; + + npages =3D 4; + size =3D npages << self->page_shift; + + /* Create userfaultfd */ + uffd =3D syscall(__NR_userfaultfd, O_CLOEXEC | O_NONBLOCK); + if (uffd < 0) + SKIP(return, "userfaultfd not available"); + + api.api =3D UFFD_API; + api.features =3D 0; + ret =3D ioctl(uffd, UFFDIO_API, &api); + ASSERT_EQ(ret, 0); + + buffer =3D malloc(sizeof(*buffer)); + ASSERT_NE(buffer, NULL); + + buffer->fd =3D -1; + buffer->size =3D size; + buffer->mirror =3D malloc(size); + ASSERT_NE(buffer->mirror, NULL); + + /* Create anonymous mapping */ + buffer->ptr =3D mmap(NULL, size, + PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS, + -1, 0); + ASSERT_NE(buffer->ptr, MAP_FAILED); + + /* Register the region with userfaultfd */ + reg.range.start =3D (unsigned long)buffer->ptr; + reg.range.len =3D size; + reg.mode =3D UFFDIO_REGISTER_MODE_MISSING; + ret =3D ioctl(uffd, UFFDIO_REGISTER, ®); + ASSERT_EQ(ret, 0); + + /* Set up the handler thread */ + uffd_args.uffd =3D uffd; + stop_fd =3D eventfd(0, EFD_CLOEXEC); + ASSERT_GE(stop_fd, 0); + uffd_args.stop_fd =3D stop_fd; + uffd_args.page_buffer =3D malloc(self->page_size); + ASSERT_NE(uffd_args.page_buffer, NULL); + uffd_args.page_size =3D self->page_size; + + ret =3D pthread_create(&thread, NULL, uffd_handler_thread, &uffd_args); + ASSERT_EQ(ret, 0); + + /* + * Use the unlocked read path which allows the mmap lock to be + * dropped during the fault, enabling userfaultfd resolution. + */ + ret =3D hmm_dmirror_cmd(self->fd, HMM_DMIRROR_READ_UNLOCKED, + buffer, npages); + ASSERT_EQ(ret, 0); + ASSERT_EQ(buffer->cpages, npages); + + /* Verify the device read the data filled by the uffd handler */ + ptr =3D buffer->mirror; + for (i =3D 0; i < size; ++i) + ASSERT_EQ(ptr[i], (unsigned char)0xAB); + + nwrite =3D write(stop_fd, &stop, sizeof(stop)); + ASSERT_EQ(nwrite, sizeof(stop)); + pthread_join(thread, NULL); + close(stop_fd); + free(uffd_args.page_buffer); + close(uffd); + hmm_buffer_free(buffer); +} + + TEST_HARNESS_MAIN --=20 2.43.0 From nobody Fri Jul 24 22:17:46 2026 Received: from mail-pf1-f171.google.com (mail-pf1-f171.google.com [209.85.210.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3EAB5442B12 for ; Wed, 22 Jul 2026 21:44:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756697; cv=none; b=h1oMcCWdoWRlKv6KUbW0CFNJX4Lqt1IcaxZWbvIq1LspbxMhb1f1PVMR26IQhcDDGNgqwm6vPu56ep0uGgZ5NNxAWcd78Y0MMq0V4iHwUWnRbcddkbibCUUTsi1uvteB06E0n2lgHVBebdwO7p3LotAPVFbGNsAYRHYw9rNPr9c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756697; c=relaxed/simple; bh=ZyAhBAfWZePcXjSsL7QqvsrSYU5bhX/lu/bgMYHTJjg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=ciG5IJrlpC/r+yFtLNwulQm9nOv/tQ4obQAStdWLNvBDohn+gi8MQAQ7UoJ0nFvUJSiTpkhQQcLaxwui1sP8AQyeYM+KENkfnN5k9g3XlOo6i6u7G+hQ792c6YvkJC4QoqIkEgxK5YGnC3BsRM2TtTAZhkTnjQP1HSJNxbA0d2A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=r2b1GNxZ; arc=none smtp.client-ip=209.85.210.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="r2b1GNxZ" Received: by mail-pf1-f171.google.com with SMTP id d2e1a72fcca58-84862b0d5aeso14253751b3a.2 for ; Wed, 22 Jul 2026 14:44:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784756693; x=1785361493; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=3slUHnl834IP76VoxdiRinKrcbFYlUZTVsrXFJZ4JdY=; b=r2b1GNxZG0pMLdmLeYUJ66eog3ws7mntKym23Cgz76GD4C0nsBH3f9m19Hrcw373vY MPWtKIY6/nnSTSrV4hxZ2Qi5iWYBvD8PtV/YL/RdqRvyZntwz2M1BXvYgNcrmp4ugxDS ZBcwsrYtDo7vcMTi9eLs9+q9nXv5LZlEP2NphZlze0yRmKypPyBeSJOVE61JJMgq4sek nE5PMZeAiq/qBmTTWpjLnLhZoLnBoie8tXP2TL5aqjFgygETjzB5LNn28ue2CagrPNK+ 367OCC0BdQLtNLgYOMZksa/UfMZTzrtKUVcYO/YSqHKXz8yJnzhvh4sf7VVa1QEqxfss auaA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784756693; x=1785361493; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=3slUHnl834IP76VoxdiRinKrcbFYlUZTVsrXFJZ4JdY=; b=OOsUCc5kmpdNs1tCgBllyYXIW2wDVq+DJ8P//IgPkkASTiyP/UtuEl7c+oxpjS2eSr oPnRmUVdrZUA2/dHb4QuJXv71RysnxzVWwqH2YOSDGRZfy2SIy9/7I9eYGwiz023vhm4 UgAnkexf/LLto30l3vpHVSfEM9soKoy3nZZu48kfI1M/KqHzm+ggozDrPwMwWAkQ89bg n5IwP8RMH1yvKODjckGAmcH2thAEOlfYXTPgELeWKiFyoMoJLJRaJflSzt2Lu3bUu3G9 oCrLFF59vo+7Fyc3MV981goJq9mwuHTP3WPj+l099R1KJayh0Pg6BM0nyQD4+aP6FIe4 yHnQ== X-Forwarded-Encrypted: i=1; AHgh+RrHZnj/biPVOY5l1IPniYFmUjROhKZyEHqoa82GyhcRxgPTUnyJaoByEKbhpExj/nGcyQZ9Ofba/BUJ0jM=@vger.kernel.org X-Gm-Message-State: AOJu0Yy21F8Fsm0/5schTsfPg0mB5SeOwGJufKuC+Ujque5nLaSX3oqF 5MlkArREDkndeQOvec76Bq9/R3OnnRNZmS0RgieOUeUrdJ9LCt139238 X-Gm-Gg: AR+sD12ORggZqWiRjy3ynNIPh2Su9QBXrbTGZmgZzrXX/EHe93DOD+AaQglLTzvuF7E 7XbzQu7Ut4EgwfnAEvd1W7LD+XYHY9uiAa1C0YmLoFcvZInRRH/rEGVfEM51x0/Empzym2xoyta PKXf3zJ7ZwAjQ8mqMdH0wTJge9t+UX7AV0o75Mn6cXJn+o9F+rRRwTbGcDMnVbFd8/X+y/yWIU9 +M27t+QEVi4FV1oryaqmgWInRiM6+wSZK/cv2uD5B9MDqC2YxT5lsgpsvY/A/HEqR8SL0aQC9cT 304zohfNqfncaCEySTRZPh3oM3E9tSpukngiNooRLn4pZRKW4SLP2ttTAXcxmsRozKOFQpAPNf1 pyT1myOTbrI6oiQ6fbIg442gl7Fk1Te4ZF/kkDbGXWggIChrnkjonCC0Fob6kXtz4RMMB6kQcmk COK68SU6ERXV2zeb4sqItbu73JFrJ5ilYe+EacuWOqI1nYU0+k X-Received: by 2002:a05:6a00:4fd3:b0:848:3fe2:c88b with SMTP id d2e1a72fcca58-84e2b7f59camr674000b3a.6.1784756693352; Wed, 22 Jul 2026 14:44:53 -0700 (PDT) Received: from [192.168.0.160] (c-98-225-44-182.hsd1.wa.comcast.net. [98.225.44.182]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84e17237c3fsm1946297b3a.9.2026.07.22.14.44.51 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 14:44:52 -0700 (PDT) From: Stanislav Kinsburskii Date: Wed, 22 Jul 2026 14:44:26 -0700 Subject: [PATCH v10 4/8] mshv: Use hmm_range_fault_unlocked_timeout() for region faults Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hmm-v10-v1-4-606464dd601a@gmail.com> References: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> In-Reply-To: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> To: Jason Gunthorpe , Leon Romanovsky , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Shuah Khan , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Min Ma , Lizhi Hou , Oded Gabbay , skinsburskii@gmail.com Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-hyperv@vger.kernel.org, dri-devel@lists.freedesktop.org, nouveau@lists.freedesktop.org, linux-rdma@vger.kernel.org, Jason Gunthorpe X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784756683; l=3646; i=skinsburskii@gmail.com; s=20260722; h=from:subject:message-id; bh=ZyAhBAfWZePcXjSsL7QqvsrSYU5bhX/lu/bgMYHTJjg=; b=tdDpNRsCNvmVGCNOZZ9HC3yxGmAH4ax4m27NmzACrsQovmkbRz97eJ93jI+rhX+xJ0J5zVBrq FSNwxKzSGzhDXRe3C8R+8iVP4estzVdVfYmBtDp3huLkA3AiN268Vyv X-Developer-Key: i=skinsburskii@gmail.com; a=ed25519; pk=bDpriHBYgeTdkIDweZDCemxsU93neJBOCn3YLIuJpnE= MSHV currently faults movable memory regions by taking mmap_read_lock() around hmm_range_fault(). That prevents the fault path from handling VMAs whose fault handlers need to drop mmap_lock, such as userfaultfd-backed mappings. Use hmm_range_fault_unlocked_timeout() instead. Passing a timeout of 0 preserves MSHV's existing unbounded retry behavior while letting the HMM helper own mmap_lock acquisition and refresh range->notifier_seq internally before walking the range. After the fault succeeds, MSHV still takes mreg_mutex and checks mmu_interval_read_retry() before installing the pages into the region, so the existing invalidation synchronization is preserved. Fold the small fault-and-lock helper into mshv_region_range_fault(), since the remaining retry path is just the standard "fault, take the driver lock, check the interval notifier sequence" pattern. Reviewed-by: Jason Gunthorpe Signed-off-by: Stanislav Kinsburskii --- drivers/hv/mshv_regions.c | 54 +++++++++----------------------------------= ---- 1 file changed, 10 insertions(+), 44 deletions(-) diff --git a/drivers/hv/mshv_regions.c b/drivers/hv/mshv_regions.c index 6d65e5b42152..dddaade31b5d 100644 --- a/drivers/hv/mshv_regions.c +++ b/drivers/hv/mshv_regions.c @@ -381,46 +381,6 @@ int mshv_region_get(struct mshv_mem_region *region) return kref_get_unless_zero(®ion->mreg_refcount); } =20 -/** - * mshv_region_hmm_fault_and_lock - Handle HMM faults and lock the memory = region - * @region: Pointer to the memory region structure - * @range: Pointer to the HMM range structure - * - * This function performs the following steps: - * 1. Reads the notifier sequence for the HMM range. - * 2. Acquires a read lock on the memory map. - * 3. Handles HMM faults for the specified range. - * 4. Releases the read lock on the memory map. - * 5. If successful, locks the memory region mutex. - * 6. Verifies if the notifier sequence has changed during the operation. - * If it has, releases the mutex and returns -EBUSY to match with - * hmm_range_fault() return code for repeating. - * - * Return: 0 on success, a negative error code otherwise. - */ -static int mshv_region_hmm_fault_and_lock(struct mshv_mem_region *region, - struct hmm_range *range) -{ - int ret; - - range->notifier_seq =3D mmu_interval_read_begin(range->notifier); - mmap_read_lock(region->mreg_mni.mm); - ret =3D hmm_range_fault(range); - mmap_read_unlock(region->mreg_mni.mm); - if (ret) - return ret; - - mutex_lock(®ion->mreg_mutex); - - if (mmu_interval_read_retry(range->notifier, range->notifier_seq)) { - mutex_unlock(®ion->mreg_mutex); - cond_resched(); - return -EBUSY; - } - - return 0; -} - /** * mshv_region_range_fault - Handle memory range faults for a given region. * @region: Pointer to the memory region structure. @@ -452,13 +412,19 @@ static int mshv_region_range_fault(struct mshv_mem_re= gion *region, range.start =3D region->start_uaddr + page_offset * HV_HYP_PAGE_SIZE; range.end =3D range.start + page_count * HV_HYP_PAGE_SIZE; =20 - do { - ret =3D mshv_region_hmm_fault_and_lock(region, &range); - } while (ret =3D=3D -EBUSY); - +again: + ret =3D hmm_range_fault_unlocked_timeout(&range, 0); if (ret) goto out; =20 + mutex_lock(®ion->mreg_mutex); + + if (mmu_interval_read_retry(range.notifier, range.notifier_seq)) { + mutex_unlock(®ion->mreg_mutex); + cond_resched(); + goto again; + } + for (i =3D 0; i < page_count; i++) region->mreg_pages[page_offset + i] =3D hmm_pfn_to_page(pfns[i]); =20 --=20 2.43.0 From nobody Fri Jul 24 22:17:46 2026 Received: from mail-pf1-f179.google.com (mail-pf1-f179.google.com [209.85.210.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1724943D515 for ; Wed, 22 Jul 2026 21:44:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.179 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756698; cv=none; b=ai4PV3hQvCM3OTMGARo7/4J7T+cQCZun9NzS+1J2y4Xaws+bID1iA4Mm3EQQZ18GLoaSUidBQPXuYKLS8SOPWKhZ2I/UMNW9j9ru7C+7Y2yIhE19C3NfXqnoeIoSY0WThXwvbUUS7/s7eOId6w1/pcDXi3CWv7yGZsCvhE/o7kA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756698; c=relaxed/simple; bh=KaVy0LDD2WZ56zIlG4jDVljfxrz2j27JL0AnOSACkpk=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=npuWi56j3qzIKT02eGM4tzvcPmwwZH0VLq2THtNERTxpudqPlBm7j2C+1QW0If/hJv4t8EoWFmHgYMljBIPBH6QmeFgLTHWqBDIh2QBiiGCtje1WrxJNzOMRq1Xgy2zf1WwwzlssMC5RrhBQctyO9VZ6pkK+NU929ftufUXwGAQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=MjXGis/0; arc=none smtp.client-ip=209.85.210.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="MjXGis/0" Received: by mail-pf1-f179.google.com with SMTP id d2e1a72fcca58-8487214ad2bso12675933b3a.1 for ; Wed, 22 Jul 2026 14:44:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784756695; x=1785361495; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=WVIK4CTS2ebWqFuX16f+6TzH9lK4l3CZHr+WHLsMqrs=; b=MjXGis/08C6P1DbvTDS6XJJiOLeKA3eQfTgRH93dbv1sJJCfHSRkL3aF/mNyyMkS+M 2/bzui3U5K4koDZ1W+02OsJlccKvZw20haIwOpjsQL6hK4muM/KJT0lqRDPbyLUU18VR pm1Mgq7UGLlKE0DYF+SgPGFSWxgOKUnaQ85PNqqo5naqIWdQKyv4vHEXcrv87MBWVXKn zra+0ehpmUoykyEljanDVPzp3AfQRYC/c5nVsAoTFdRDqc+tF2ByiA7xh7hs5Dj0b0x1 OEtwLJzVmH1UNoPWyOsX/CF+qI+idQQHhRLt2Svm9kDITmGxLhByKG6WSxAoISesyuHF es8Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784756695; x=1785361495; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=WVIK4CTS2ebWqFuX16f+6TzH9lK4l3CZHr+WHLsMqrs=; b=dxiF8n5XPhsWOctakGR6bdzLODVBhzewH77YRY+jRubXOyb6BtFGRjRV1nxWcBgirY Sz9iWCybyD0nxandp3p5HewDYONgOxHb1lce6tGpdwVb9E6HMKT5a+Y8j/ROaQrZGOyG 4pH6oUG7yggK8b5Y0OQXJdND9ZQx/F2REZIyy3eZspkuORxJNWkyhId5fVw4o2KFpLR2 K1ICBqmX+DJ4AfzGRzoAaF4jELx0rIUT4jlzR0EwAHYjfqM7VE5C19UtdUwVYAfYojXH H/o+4wu6tmBedJsLQuhJA2ZxzmmOfVzE+TP8i1hclfcwYcWixOvU/9qf0F/usbowQ9Ww 960A== X-Forwarded-Encrypted: i=1; AHgh+RoolVIrHtLd8xs9ulpR0ER+iWc918X16J2A3HNgz9PMcsn2v+0rZLDIFbs/rZHTo/JHfr1QBXoCyDG+00E=@vger.kernel.org X-Gm-Message-State: AOJu0YyKHkMx7McS2ch2h4oOXK8xbyNUsT6CUFeVa150rKNN5yEYs3qb CVOkWyf83jxG9FZ4+1ln7l4AQvziR7kZS5f9SOqtyqtKtKz169OXU7ld X-Gm-Gg: AR+sD13VZWeeKcoVU1AvebFkutQH2mdah7+u8me1SMIg/T/J5qVNmieBtzst2MBolPV WmUHUPd9UEyl8lfFeriC4ovTlhdV8iKpGd9O4kT4mIVG6fUU7bqpGiyp25S38Gj1l1m80f338+B CS6TPOs+BPZ+GO7p1PdkKHS5iqZjqUndfFNYOiLFo/rae/Tdw7f6Yt+Be10mQH1hQ3Di4Mzj5+t r+YcjTRyznkbedOyvL1k47scxFG5kpyDR3ugGuEXavAPiZXoxfP28UZ5bJy+3Ogwu5+lkKFhxB+ aE08YJG68GLBlqc56mJGlsk8V7WUcwKcAO+4skMEkjkt5Kbx1i303YXhNl9pxZwaYFgS7StL7ap TxVay12YIxzl1PJPgVcgr98IWs9eEhYw/YzsXSQZA4soR/VQmhk1+2PD/PzedQtnoFop+CfxOik /0z+aurFwKStym1xV+zCKqPu03RDzPwlgdOyJtRw2oSlQlXvNF X-Received: by 2002:a05:6a00:852:b0:848:46ea:bbe7 with SMTP id d2e1a72fcca58-84e2bc6b5d6mr673956b3a.19.1784756695183; Wed, 22 Jul 2026 14:44:55 -0700 (PDT) Received: from [192.168.0.160] (c-98-225-44-182.hsd1.wa.comcast.net. [98.225.44.182]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84e17237c3fsm1946297b3a.9.2026.07.22.14.44.53 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 14:44:54 -0700 (PDT) From: Stanislav Kinsburskii Date: Wed, 22 Jul 2026 14:44:27 -0700 Subject: [PATCH v10 5/8] drm/nouveau: Use hmm_range_fault_unlocked_timeout() for SVM faults Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hmm-v10-v1-5-606464dd601a@gmail.com> References: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> In-Reply-To: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> To: Jason Gunthorpe , Leon Romanovsky , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Shuah Khan , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Min Ma , Lizhi Hou , Oded Gabbay , skinsburskii@gmail.com Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-hyperv@vger.kernel.org, dri-devel@lists.freedesktop.org, nouveau@lists.freedesktop.org, linux-rdma@vger.kernel.org, Jason Gunthorpe X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784756683; l=2163; i=skinsburskii@gmail.com; s=20260722; h=from:subject:message-id; bh=KaVy0LDD2WZ56zIlG4jDVljfxrz2j27JL0AnOSACkpk=; b=uKgRFxKY7gMhEFoK0hHnwQl6N8UsouNcazo70jIAxend3BGPO0koGFjjGoE03g3KTWZd66DI6 5tAsc5/qbDAAjPQ/pjqrtB3km5I7U9rctkdARaqcdzu58Lj+SzHd5bT X-Developer-Key: i=skinsburskii@gmail.com; a=ed25519; pk=bDpriHBYgeTdkIDweZDCemxsU93neJBOCn3YLIuJpnE= nouveau_range_fault() takes mmap_read_lock() only to call hmm_range_fault(). It also keeps a single HMM_RANGE_DEFAULT_TIMEOUT deadline across both HMM -EBUSY retries and post-fault mmu_interval_read_retry() retries. Use hmm_range_fault_unlocked_timeout() instead. The HMM helper now owns the mmap lock and refreshes range->notifier_seq for its internal retries. Nouveau keeps its existing absolute deadline in the outer loop and passes the remaining jiffies to the helper for each fault attempt, so retries caused by mmu_interval_read_retry() do not reset the overall retry budget. Nouveau still validates the interval notifier sequence while holding svmm->mutex before programming the GPU mapping. Reviewed-by: Jason Gunthorpe Signed-off-by: Stanislav Kinsburskii --- drivers/gpu/drm/nouveau/nouveau_svm.c | 20 +++++++++++--------- 1 file changed, 11 insertions(+), 9 deletions(-) diff --git a/drivers/gpu/drm/nouveau/nouveau_svm.c b/drivers/gpu/drm/nouvea= u/nouveau_svm.c index dcc92131488e..58735446d783 100644 --- a/drivers/gpu/drm/nouveau/nouveau_svm.c +++ b/drivers/gpu/drm/nouveau/nouveau_svm.c @@ -678,20 +678,22 @@ static int nouveau_range_fault(struct nouveau_svmm *s= vmm, range.end =3D notifier->notifier.interval_tree.last + 1; =20 while (true) { - if (time_after(jiffies, timeout)) { + long remaining =3D timeout - jiffies; + + /* + * The HMM timeout only bounds retries while HMM is walking and + * faulting the range. This fault is handled by a kernel worker, + * so fatal signals from the faulting process cannot stop an + * endless stream of invalidations here. + */ + if (time_after_eq(jiffies, timeout)) { ret =3D -EBUSY; goto out; } =20 - range.notifier_seq =3D mmu_interval_read_begin(range.notifier); - mmap_read_lock(mm); - ret =3D hmm_range_fault(&range); - mmap_read_unlock(mm); - if (ret) { - if (ret =3D=3D -EBUSY) - continue; + ret =3D hmm_range_fault_unlocked_timeout(&range, remaining); + if (ret) goto out; - } =20 mutex_lock(&svmm->mutex); if (mmu_interval_read_retry(range.notifier, --=20 2.43.0 From nobody Fri Jul 24 22:17:46 2026 Received: from mail-pf1-f176.google.com (mail-pf1-f176.google.com [209.85.210.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 23C8E444719 for ; Wed, 22 Jul 2026 21:44:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.176 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756700; cv=none; b=mCixraRr/DMKivabimfr0rXw2Xc7HUR6tlHFAbe2AmHLsimaHW3LbsHLQ6LA0BDL50UhQ8b0XwzWj1RlgQNcnNdiB+Bt54G0EsoI0uDEs9oaAG2RLjCjugIl+f/gEJ0o5YITO6iIx4BY3kcI4Gy51F7nIKQlHOaEReGoJQ0+kK4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756700; c=relaxed/simple; bh=wCDSwsZ/5jgwWzEqFja4q7w3SxWPNbitK21qCohy6H8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=b2+3trM9mHqcbND2lShjbYNMUBpjylwKi+brCo08L9SLC9hf3I9P0WpiLpSwn9wEuFZJhVY1en+C6izMOsk1zzYbCOhSf/gGsKYMQqr8f3A3H42BpKvKgRxnnuMZndcm8L3HlhWhhkerxX00S6C1WqB/b8dD2T+lB8Q0D653/nM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=PTMOc09F; arc=none smtp.client-ip=209.85.210.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="PTMOc09F" Received: by mail-pf1-f176.google.com with SMTP id d2e1a72fcca58-84e27035206so598733b3a.3 for ; Wed, 22 Jul 2026 14:44:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784756697; x=1785361497; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Hwr2PA6k/yzOHBkhyRcV60oQVyvCsFMHOEQ+QEbH+Ag=; b=PTMOc09FpcMlM1DZNUhPtBUi/K2aOWkLARcKfkanolkLZxB6W1E4j8SxD8pdcvvU3Z 4OUnqhRDqm7SnYk5Wyg0CyjAOu+KMvNSVcH0xAukoVzeUEA5fYXzT/o94NBxxSD6MGj+ IN+s+1LTizRBVmM49hhXv/iwOHDtKTfca33m7E/iLv0BsXJr7if4H293U+9R5KwknEDc qNWwMa8WmksgPgWpWSvXsaG/Wv3WYnO6t2667aKZv/CyHt9b/835czW4PgMGP4VHlIfH fuGhVRvQNL4OgtpnRhLvGsI9w+lVynv7fPfX3GsEp3GJiNoLJqpqbdjC1mTY91lWJ+TV 613A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784756697; x=1785361497; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=Hwr2PA6k/yzOHBkhyRcV60oQVyvCsFMHOEQ+QEbH+Ag=; b=Cc7tqFW7LLLezhFyhbiX0pYyXSgZDUB7qPr1QGXcYey3dI9eYE+se+nuq0MU2xiwKy 0fwPbhIq008Uk52wiL9MuFRtsZdgxzdAlMkoKOyUFj7zHJ4NbFlyTiQLN7m8pmmVdxIR CxaZygS9rSBGCUO4jNfam6cXtdahlumPux04EYZS/zQMCqTIkbhQqmUeqUxwQ7qPidEx S5lmbOHej8ZNYX8o6Bco2INFOA5/x2JMJq9s2BImddEjacSIWuCEcGTdh4c+HnAwIKp7 I4EIRu6Z9cFkig+xNlw7VQjI+UekgGVmOMH2CRKhmk4zO6QlF7s5DGZo7jIHlETZXyqA yC6w== X-Forwarded-Encrypted: i=1; AHgh+RpEJIgEA8oUgr0SUUMYuttNdWKK562fGDZ9a08aKwr2qO5fubodlkIcOu8KAr25ydoaZT5EKWsqMoRIO8U=@vger.kernel.org X-Gm-Message-State: AOJu0Yw3rmRfQ44hZ8KCqQDrFQxxxWjNC6yoTtaBYWDbJCFky5I0VBuI MqwmSWdb/s8W334eVS3EVj3SIPLKQs1Yu8Izlk9NAJntJGQjofazRq9Z X-Gm-Gg: AR+sD12SszoDxPRozKnp8l1R1/HJY/7aa4H1blJbQLcnjcBnjT8QSZzPhSwg1lgAZ8m K1rgrOHKWfQvtp0J7nMwdo9jsL8nUH9GNRdz77db/h02H1pdxcNkkP9utc2JWED+GNP7LhUBiFt dwFTgifxSeLoWXEMf54vz+yjareiPJ3G0X+My1rhZoL4ULvryg7XJUZ7guHEhdDJLJQ9LLPwh/W YMJs5tTON3lHL7s05vZcRJM5E0orLvkCvSxsl0HckkFC57ARbOFmxrBmB0PwAR0rQGYJ3MdgcTG QtMBthIVAxWRpvdCv6i78OsUSjz2VhvVqRCkFBEzGonXsAiq/QcGV3UqXdsDZ3Hf3gR0wxEO2wi QtwSFeHvIrsdQKkMASAx+8ndw/6cssB4ySQYwwpjm838eEV8S5MHi/VYGWJn1qnRukFSpLdHypu RDS3BUt+8hj3HhKTaM0Ye+6HFq8Us7G+QGP+CVjg4jHM8nCmz/ X-Received: by 2002:a05:6a00:1388:b0:848:7b92:8e74 with SMTP id d2e1a72fcca58-84e2c1b9f43mr598680b3a.39.1784756697112; Wed, 22 Jul 2026 14:44:57 -0700 (PDT) Received: from [192.168.0.160] (c-98-225-44-182.hsd1.wa.comcast.net. [98.225.44.182]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84e17237c3fsm1946297b3a.9.2026.07.22.14.44.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 14:44:56 -0700 (PDT) From: Stanislav Kinsburskii Date: Wed, 22 Jul 2026 14:44:28 -0700 Subject: [PATCH v10 6/8] RDMA/umem: Use hmm_range_fault_unlocked_timeout() for ODP faults Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hmm-v10-v1-6-606464dd601a@gmail.com> References: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> In-Reply-To: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> To: Jason Gunthorpe , Leon Romanovsky , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Shuah Khan , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Min Ma , Lizhi Hou , Oded Gabbay , skinsburskii@gmail.com Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-hyperv@vger.kernel.org, dri-devel@lists.freedesktop.org, nouveau@lists.freedesktop.org, linux-rdma@vger.kernel.org, Jason Gunthorpe X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784756683; l=2457; i=skinsburskii@gmail.com; s=20260722; h=from:subject:message-id; bh=wCDSwsZ/5jgwWzEqFja4q7w3SxWPNbitK21qCohy6H8=; b=oTLl/A3KKM6QVfczDQir+MiwWd6xegg7AP0nwym7Jb+EYIV4zgzMMZlGO0Uy3Tv+0wtFJYifg zfmA30BQ0D/AoP9W2yBY0iwSakKcdy31nmnUIbQ0QmfiWeEgduaY6fE X-Developer-Key: i=skinsburskii@gmail.com; a=ed25519; pk=bDpriHBYgeTdkIDweZDCemxsU93neJBOCn3YLIuJpnE= ib_umem_odp_map_dma_and_lock() takes mmap_read_lock() only around hmm_range_fault(), then retries -EBUSY until HMM_RANGE_DEFAULT_TIMEOUT expires. Use hmm_range_fault_unlocked_timeout() instead. The HMM helper now owns the mmap lock and refreshes range->notifier_seq for its internal retries. ODP keeps using HMM_RANGE_DEFAULT_TIMEOUT for each HMM fault attempt, while interval invalidation retries continue to be handled by the existing outer loop. ODP still validates the interval notifier sequence while holding umem_mutex before DMA mapping pages. Reviewed-by: Jason Gunthorpe Signed-off-by: Stanislav Kinsburskii --- drivers/infiniband/core/umem_odp.c | 18 +++++------------- 1 file changed, 5 insertions(+), 13 deletions(-) diff --git a/drivers/infiniband/core/umem_odp.c b/drivers/infiniband/core/u= mem_odp.c index 404fa1cc3254..9cc21cd762d9 100644 --- a/drivers/infiniband/core/umem_odp.c +++ b/drivers/infiniband/core/umem_odp.c @@ -329,7 +329,7 @@ int ib_umem_odp_map_dma_and_lock(struct ib_umem_odp *um= em_odp, u64 user_virt, struct mm_struct *owning_mm =3D umem_odp->umem.owning_mm; int pfn_index, dma_index, ret =3D 0, start_idx; unsigned int page_shift, hmm_order, pfn_start_idx; - unsigned long num_pfns, current_seq; + unsigned long num_pfns; struct hmm_range range =3D {}; unsigned long timeout; =20 @@ -363,26 +363,18 @@ int ib_umem_odp_map_dma_and_lock(struct ib_umem_odp *= umem_odp, u64 user_virt, } =20 range.hmm_pfns =3D &(umem_odp->map.pfn_list[pfn_start_idx]); - timeout =3D jiffies + msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); + timeout =3D msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); =20 retry: - current_seq =3D range.notifier_seq =3D - mmu_interval_read_begin(&umem_odp->notifier); - - mmap_read_lock(owning_mm); - ret =3D hmm_range_fault(&range); - mmap_read_unlock(owning_mm); - if (unlikely(ret)) { - if (ret =3D=3D -EBUSY && !time_after(jiffies, timeout)) - goto retry; + ret =3D hmm_range_fault_unlocked_timeout(&range, timeout); + if (unlikely(ret)) goto out_put_mm; - } =20 start_idx =3D (range.start - ib_umem_start(umem_odp)) >> page_shift; dma_index =3D start_idx; =20 mutex_lock(&umem_odp->umem_mutex); - if (mmu_interval_read_retry(&umem_odp->notifier, current_seq)) { + if (mmu_interval_read_retry(&umem_odp->notifier, range.notifier_seq)) { mutex_unlock(&umem_odp->umem_mutex); goto retry; } --=20 2.43.0 From nobody Fri Jul 24 22:17:46 2026 Received: from mail-pf1-f171.google.com (mail-pf1-f171.google.com [209.85.210.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A3B8244606A for ; Wed, 22 Jul 2026 21:44:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756702; cv=none; b=dy0vCV+hhYIkXbhkDEBTuVFQq7pGG0ssU0WAv69XFmLcl7V7KJKFCVGrHcoFiM+c/mIrhu7s526xTHW1UqM6lJb2Q1FKo//mbxy7p9i2DQeSK53kAXcehBOSk8YGduzuZX/DPojdAPKTTFDwcW60fuH+al94BCyPKAMg0/0a4b0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756702; c=relaxed/simple; bh=Js9Z5fScNTLs9yZDEPq5MMQDWBFdccrrYdkvKMfBR0c=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=XHgh2LMCECcmQIckuwF47gizBvvv/jpPDPGM2LhRxkb9WWx9+U4PX9At6fnF+0SpPKBacgCmvfM7Tl/OjeDE3f8bW4k5OPydCle+8fi4LPZEE1H1HT2En0bkptZJZn+d+OrJafe/F2d3XSXLf/21VFKnI5bXdjeTb9K0otPKUO0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=bqfwTEge; arc=none smtp.client-ip=209.85.210.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="bqfwTEge" Received: by mail-pf1-f171.google.com with SMTP id d2e1a72fcca58-8485ef63b68so12772353b3a.1 for ; Wed, 22 Jul 2026 14:44:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784756699; x=1785361499; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=/3EHZQlaXq9sLmhVgtuP5ODA2wxc36yvs54yemcncOc=; b=bqfwTEgeLLBzB5btm+3q41Gfm+tsHhfYbXFwI58ASYtVh9ExV7vdkWmUMY3W9XDpcQ /7a3wuLg87gEATfBwM7u6oitGBFUa39ChxwuVNJBxRF5Ak1PoUFo+L11pE45bP6ya/bg po9Y/JtM8iRcws9GV/n53TuQp0N9DDeyysQITt8sLvx82dVDk3AHl6ES5s560p1EIxIh SKAkCGq/X6mzwji+sQSViy1DJMv462IYVEYVhRbDr5vZuA5ErWeQIivIDyGOl4Wuq19S AfHVZTJ4snaeY8S9TL+mwHfhE0yJr033IVhHgsDXAp574jP8wgYVYClSB7JWyYBenWW4 01ZA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784756699; x=1785361499; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=/3EHZQlaXq9sLmhVgtuP5ODA2wxc36yvs54yemcncOc=; b=A2wlMYoyDEx4bj6EMwXXCefvYkNLdXRrrxyqChWqzA2dTFVhzGn9zacTZWoBaGw2/S 1du1CU9iQUM0uXVIaZkuulbBgEGnwThoLoYqqmjEjxGT400bWms7d3e2vjjhHszHcw29 yhoP9DxmbZ/0wnbhD5bknmUxkrOl3AX6Ac9zMYmKMxcJU6VK/SJNfWZqvCNFwHd8+GnS 7tO987Rf5LcnDcfZQ7agykoGmhxsx7th3P9suESSIAdck6QtG4XRtUfo2cGh103ojLPm pg66Tx+0P2N1yI/b3p0llwYrePZZ4EwiMq18r1phv69i1jwchrPLBsC9Ufp7IJ8fz267 fwjw== X-Forwarded-Encrypted: i=1; AHgh+Rr/BD6Ewz8r1nuOgA04JOqEzX/orDl7pmM1BR5mmHXC+MtEbiUcxzbvvJHkrokbwzDRROLFnPZuOpGZUdI=@vger.kernel.org X-Gm-Message-State: AOJu0YzYQ4qM+8lTIhipbYpfgqnUJYXaDvWWtZ1vubVkUlSEMrMXavrG Dc+hs92nBIyuj30pjBiRWgqC/sUoLT4drr8bKTTBAPKGFW9IYGR2yq8k X-Gm-Gg: AR+sD10CG31Zl/zcmASOmU6VBq+4MjakvURuLRIZXjwxEnCvYODAsI6Y/KMJy7rIKOI 13xZwH8jWZOXtdwlXmtIgN1wWfW+fLN/WofCCsMblnqnQkNUuKiL8+ZXV0QXdDbfAjotRW8ctKY huUTdbVhIUi9zYa2t6TEmeP5rKsVILzLktI0eDzFN6G8FyBKQGFMV6Sj6gZTyk90oNk8DXVrZpf DoVAH+hN9pE0PZXzoMoZdhTFeRZPo1k/14yEAFlitWuMrCQZ1zxDCEycqKemrwSyheKnzZeg39V VO0CC0f9fPnLjt0GnY7Qml+khVFWnsMlKZVQU/ZEgX+id8A8fv6vckkuyZEUa4FbVSZbXhGxiv7 Jl74nIgvGmFAE/CjAXHh4uzh97bLZwC9hBxdf3eDsCMRpq9WiDAKM4pqtq/DkLGMtQOm55tLyrD mLYGjLcsXSD68uG9Vm1BJFZTi4p4N8pazn3FWBz1H0mq/1zl5g X-Received: by 2002:a05:6a00:2e21:b0:848:5ca1:bae0 with SMTP id d2e1a72fcca58-84e2bb3728fmr677188b3a.42.1784756698735; Wed, 22 Jul 2026 14:44:58 -0700 (PDT) Received: from [192.168.0.160] (c-98-225-44-182.hsd1.wa.comcast.net. [98.225.44.182]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84e17237c3fsm1946297b3a.9.2026.07.22.14.44.57 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 14:44:58 -0700 (PDT) From: Stanislav Kinsburskii Date: Wed, 22 Jul 2026 14:44:29 -0700 Subject: [PATCH v10 7/8] accel/amdxdna: Use hmm_range_fault_unlocked_timeout() for range population Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hmm-v10-v1-7-606464dd601a@gmail.com> References: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> In-Reply-To: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> To: Jason Gunthorpe , Leon Romanovsky , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Shuah Khan , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Min Ma , Lizhi Hou , Oded Gabbay , skinsburskii@gmail.com Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-hyperv@vger.kernel.org, dri-devel@lists.freedesktop.org, nouveau@lists.freedesktop.org, linux-rdma@vger.kernel.org, Jason Gunthorpe X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784756683; l=2650; i=skinsburskii@gmail.com; s=20260722; h=from:subject:message-id; bh=Js9Z5fScNTLs9yZDEPq5MMQDWBFdccrrYdkvKMfBR0c=; b=OvUAdAx0y7f+AZFnN5LSfteCYnBEUqdMQ1drzLwHbqAw6OKXvHLZibJ29BghZSxMaAsL2Igct 7dhM0ZwBa4kCtizE7Rk+AH/sCkM9/0Xx1LPVtAK4e1kYz6EQ4qhx7vD X-Developer-Key: i=skinsburskii@gmail.com; a=ed25519; pk=bDpriHBYgeTdkIDweZDCemxsU93neJBOCn3YLIuJpnE= aie2_populate_range() takes mmap_read_lock() only around hmm_range_fault(). It also open-codes the mmu interval sequence setup before each HMM walk and retries -EBUSY until HMM_RANGE_DEFAULT_TIMEOUT expires. Use hmm_range_fault_unlocked_timeout() instead. The HMM helper now owns the mmap lock and refreshes mapp->range.notifier_seq for its internal retries, so the driver only needs to call the helper and then validate the sequence before marking the mapping populated. Pass HMM_RANGE_DEFAULT_TIMEOUT as the helper retry budget for each HMM population attempt. This scopes the timeout to repeated HMM notifier retries while preserving the existing outer loop that moves between invalid mappings and restarts when the interval is invalidated before the driver updates its mapping state. Keep returning -ETIME when the HMM retry budget expires, matching the driver's existing timeout error convention. Reviewed-by: Jason Gunthorpe Signed-off-by: Stanislav Kinsburskii --- drivers/accel/amdxdna/aie2_ctx.c | 23 ++++------------------- 1 file changed, 4 insertions(+), 19 deletions(-) diff --git a/drivers/accel/amdxdna/aie2_ctx.c b/drivers/accel/amdxdna/aie2_= ctx.c index 54486960cbf5..b5b4ca263002 100644 --- a/drivers/accel/amdxdna/aie2_ctx.c +++ b/drivers/accel/amdxdna/aie2_ctx.c @@ -1034,7 +1034,7 @@ static int aie2_populate_range(struct amdxdna_gem_obj= *abo) bool found; int ret; =20 - timeout =3D jiffies + msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); + timeout =3D msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); again: found =3D false; down_write(&xdna->notifier_lock); @@ -1061,24 +1061,9 @@ static int aie2_populate_range(struct amdxdna_gem_ob= j *abo) return -EFAULT; } =20 - mapp->range.notifier_seq =3D mmu_interval_read_begin(&mapp->notifier); - mmap_read_lock(mm); - ret =3D hmm_range_fault(&mapp->range); - mmap_read_unlock(mm); - if (ret) { - if (time_after(jiffies, timeout)) { - ret =3D -ETIME; - goto put_mm; - } - - if (ret =3D=3D -EBUSY) { - amdxdna_umap_put(mapp); - mmput(mm); - goto again; - } - + ret =3D hmm_range_fault_unlocked_timeout(&mapp->range, timeout); + if (ret) goto put_mm; - } =20 down_write(&xdna->notifier_lock); if (mmu_interval_read_retry(&mapp->notifier, mapp->range.notifier_seq)) { @@ -1096,7 +1081,7 @@ static int aie2_populate_range(struct amdxdna_gem_obj= *abo) put_mm: amdxdna_umap_put(mapp); mmput(mm); - return ret; + return ret =3D=3D -EBUSY ? -ETIME : ret; } =20 int aie2_cmd_submit(struct amdxdna_hwctx *hwctx, struct amdxdna_sched_job = *job, u64 *seq) --=20 2.43.0 From nobody Fri Jul 24 22:17:46 2026 Received: from mail-pf1-f182.google.com (mail-pf1-f182.google.com [209.85.210.182]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0D0A7440A35 for ; Wed, 22 Jul 2026 21:45:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.182 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756703; cv=none; b=eRdhqgWUvGNBQkJCmGTntQE9UuraFdR+cpu2VZI0uJPnTGSOQbDsxj7fHL7Lpnm7phWB8jn5EXy4AYdlbSAJVABmwFzUUTspjx800vtvRpgOghCftgcpyRZAn0ZjX9Y9U20ybJN8JXPzOsoX6CenRJ90SvQADDxndzOAbAwuj4E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784756703; c=relaxed/simple; bh=JxcTBTW6lEzlUH9pe80rZJYyQZGUPa2uP7JaZIpbDO4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=AqWPyKBsECJQgIyJw3jFsfjAEuCwaOaCUnZTwP6fdW2Az73g6Ml+vmRSLzjvfmuUYURRXilutGmHKFOrWSXboQC7grhFYPC4vcYnkKFyF20bmDFijzb1JuAzr/8BTTyzCl3Fy54BVLnDLqw8yKw2oQhUQ0a8qrxFZ9c8HVl0JxA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=QEUHhGML; arc=none smtp.client-ip=209.85.210.182 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="QEUHhGML" Received: by mail-pf1-f182.google.com with SMTP id d2e1a72fcca58-84867f07d63so14622260b3a.2 for ; Wed, 22 Jul 2026 14:45:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784756700; x=1785361500; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=TSbIIq+MW5ujNCU6weGrEYnGd53JSmfWB5BOONHtR2c=; b=QEUHhGMLlERdp/ksj43uAtQfYB+uqlIx+aW6wnIUJLXOblYCFHTacX1v8ZLFDaRmBb ab77uFTIEJJ4vIYBPI33HDnSx59zsUnf2fnv8wBy9NWyQSivjbsekbWa3q2OM3pbk+cr vjK0lSgyPyEUQSjvrR2mi/XHNb3Ym/7Xd9gQrOyvm+sU+srsKQQgweud9ATK2ZVqGnAA Lcvx8JIV5RPb5ShYWnXIH2/Jk024MogiOW80uEDsrhOE9fTwU4lWHi25YbqLkonyqme9 kR1A4zPE2CtfcPM8bXudQiXO5/aEzI6C9a+fil8Tz1qe+SrWjrNIgyqT+P8mafVfVkUS NFag== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784756700; x=1785361500; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=TSbIIq+MW5ujNCU6weGrEYnGd53JSmfWB5BOONHtR2c=; b=YVuDJ95bepuxk1N/oPUQ2hKi42tTr7un4IH24NrfxDl1LddGDAz1xqswPfhc9xZetL /L/+B3CiMyDMqAjyexun+xwDz9y32WSn5hCbW6l5tfetPvSNdPcGNBo1Nv5oYWi4j6q3 VHozpUggMpYxBThVVKWYCDTuEjHcBuEEefpSSlkfUuaacOmaoyqoDH62a66YI6uSb+zx M2j5pS2phhweJ9cHB2lP6p+T9dh63NcQ0sdL4jBdEi2YQd+2hTdNoCHWxLeumxYWPImb OsUxJiD3MCUpWqFJCSSfSLdeSoWbE0UmJRco7lR3A8H1gh3IzrWzF8Hb8vqdn7MqqUFC 0B+A== X-Forwarded-Encrypted: i=1; AHgh+RpbfXFtiI+Rbligm58fNB5u3LarhL1bVRp4ZAajNC6bVQcai6e/g48AzkYhyaeZKdZp+Y2P7Un76qH/M/U=@vger.kernel.org X-Gm-Message-State: AOJu0YwE85kv+WuMORzVJNfSRfI+vXb7LA24HJS211K0ssXwEzSyyToT iQKkIJvpUNHqbQALoZ4GKoCwommpbWIg9990eEeztz1UQpuhwB+bN6mf X-Gm-Gg: AR+sD11NSXQPwXOwJ6a0WW+s0l4ojcaqnL3NWg3cXIph0vDQipkcU5R72aCFUhUzBXj XjNrGN84FfgQ1chJ1HOlkVhbtMHYdUy0INdEE8MUdyujOZh6HX2/F8aAV7nF3d+mVL8rJnyJqNE WbUoGeHdUYcVa3BqKjrAMVpKbxNNoN4nxl/rm/3GvwxqpKOW1mm8CeAWN2845wCTw14Sduea9xm xUrb9Jhu+l1PGYqtfAnnGl1GjgNDhN0i+9rzWmsVcWYRK07Tv3zj1z2weoowKaad8au3wRZ8PWE b6hsIoOHz0gHOkqZyX61mB6FY0zYxiwv+026sGrN7a9LeU4QIIZPUDTS0pwdVGKyqOJBtQoxV22 6hTmUYirPk/DGJKtyFyEt7LTSBQ0BMezzqCY6o1GPYfZdRc7U/cXs6LM9tsUAuAGdVb7675wTa+ mNUTUsdwX6X/MZ/npKbSzKRn3wAx+uzXAAq1lI8mxbs71ChA7+ X-Received: by 2002:a05:6a00:94f7:b0:848:467d:293e with SMTP id d2e1a72fcca58-84e2baddf66mr627847b3a.43.1784756700316; Wed, 22 Jul 2026 14:45:00 -0700 (PDT) Received: from [192.168.0.160] (c-98-225-44-182.hsd1.wa.comcast.net. [98.225.44.182]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-84e17237c3fsm1946297b3a.9.2026.07.22.14.44.58 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 14:44:59 -0700 (PDT) From: Stanislav Kinsburskii Date: Wed, 22 Jul 2026 14:44:30 -0700 Subject: [PATCH v10 8/8] drm/gpusvm: Use hmm_range_fault_unlocked_timeout() for range faults Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hmm-v10-v1-8-606464dd601a@gmail.com> References: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> In-Reply-To: <20260722-hmm-v10-v1-0-606464dd601a@gmail.com> To: Jason Gunthorpe , Leon Romanovsky , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Shuah Khan , "K. Y. Srinivasan" , Haiyang Zhang , Wei Liu , Dexuan Cui , Long Li , Lyude Paul , Danilo Krummrich , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Min Ma , Lizhi Hou , Oded Gabbay , skinsburskii@gmail.com Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-hyperv@vger.kernel.org, dri-devel@lists.freedesktop.org, nouveau@lists.freedesktop.org, linux-rdma@vger.kernel.org, Jason Gunthorpe X-Mailer: b4 0.13.0 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784756683; l=5001; i=skinsburskii@gmail.com; s=20260722; h=from:subject:message-id; bh=JxcTBTW6lEzlUH9pe80rZJYyQZGUPa2uP7JaZIpbDO4=; b=t5pXaIaExV86XNOZX4rodLg/qlkvO7z8UzQMAKDvBSpVBwUgOqPA+1uKWVHHstjr3FsmGAV1u FhiXQ8BfH3FAK6QXFIhXuSWoEAqA3lYTf5A7cJupNOmjpSFDJW/pEX3 X-Developer-Key: i=skinsburskii@gmail.com; a=ed25519; pk=bDpriHBYgeTdkIDweZDCemxsU93neJBOCn3YLIuJpnE= Several GPU SVM paths take mmap_read_lock() only to call hmm_range_fault() and open-code mmu interval sequence setup before each HMM walk. They also retry -EBUSY until HMM_RANGE_DEFAULT_TIMEOUT expires. Use hmm_range_fault_unlocked_timeout() for those faults. The HMM helper now owns mmap_lock acquisition and refreshes range->notifier_seq for its internal retries, while GPU SVM keeps its existing driver-lock validation with mmu_interval_read_retry() after a successful fault. drm_gpusvm_scan_mm() and drm_gpusvm_range_evict() pass HMM_RANGE_DEFAULT_TIMEOUT as the helper retry budget for each HMM fault attempt. drm_gpusvm_get_pages() keeps its existing absolute outer deadline because it can be reached from GPU page-fault workers, where fatal signals from the faulting process cannot stop an endless invalidation retry loop. It passes the remaining time from that deadline to HMM for each fault attempt. Leave drm_gpusvm_check_pages() on hmm_range_fault() because that path is called with the mmap lock already held by its caller. Reviewed-by: Jason Gunthorpe Signed-off-by: Stanislav Kinsburskii --- drivers/gpu/drm/drm_gpusvm.c | 60 ++++++++--------------------------------= ---- 1 file changed, 10 insertions(+), 50 deletions(-) diff --git a/drivers/gpu/drm/drm_gpusvm.c b/drivers/gpu/drm/drm_gpusvm.c index 958cb605aedd..b946f920b7a0 100644 --- a/drivers/gpu/drm/drm_gpusvm.c +++ b/drivers/gpu/drm/drm_gpusvm.c @@ -773,8 +773,7 @@ enum drm_gpusvm_scan_result drm_gpusvm_scan_mm(struct d= rm_gpusvm_range *range, .end =3D end, .dev_private_owner =3D dev_private_owner, }; - unsigned long timeout =3D - jiffies + msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); + unsigned long timeout =3D msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); enum drm_gpusvm_scan_result state =3D DRM_GPUSVM_SCAN_UNPOPULATED, new_st= ate; unsigned long *pfns; unsigned long npages =3D npages_in_range(start, end); @@ -788,22 +787,7 @@ enum drm_gpusvm_scan_result drm_gpusvm_scan_mm(struct = drm_gpusvm_range *range, hmm_range.hmm_pfns =3D pfns; =20 retry: - hmm_range.notifier_seq =3D mmu_interval_read_begin(notifier); - mmap_read_lock(range->gpusvm->mm); - - while (true) { - err =3D hmm_range_fault(&hmm_range); - if (err =3D=3D -EBUSY) { - if (time_after(jiffies, timeout)) - break; - - hmm_range.notifier_seq =3D - mmu_interval_read_begin(notifier); - continue; - } - break; - } - mmap_read_unlock(range->gpusvm->mm); + err =3D hmm_range_fault_unlocked_timeout(&hmm_range, timeout); if (err) goto err_free; =20 @@ -1408,6 +1392,7 @@ int drm_gpusvm_get_pages(struct drm_gpusvm *gpusvm, void *zdd; unsigned long timeout =3D jiffies + msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); + unsigned long remaining; unsigned long i, j; unsigned long npages =3D npages_in_range(pages_start, pages_end); unsigned long num_dma_mapped; @@ -1422,9 +1407,11 @@ int drm_gpusvm_get_pages(struct drm_gpusvm *gpusvm, struct dma_iova_state *state =3D &svm_pages->state; =20 retry: - if (time_after(jiffies, timeout)) + if (time_after_eq(jiffies, timeout)) return -EBUSY; =20 + remaining =3D timeout - jiffies; + hmm_range.notifier_seq =3D mmu_interval_read_begin(notifier); if (drm_gpusvm_pages_valid_unlocked(gpusvm, svm_pages)) goto set_seqno; @@ -1439,21 +1426,7 @@ int drm_gpusvm_get_pages(struct drm_gpusvm *gpusvm, } =20 hmm_range.hmm_pfns =3D pfns; - while (true) { - mmap_read_lock(mm); - err =3D hmm_range_fault(&hmm_range); - mmap_read_unlock(mm); - - if (err =3D=3D -EBUSY) { - if (time_after(jiffies, timeout)) - break; - - hmm_range.notifier_seq =3D - mmu_interval_read_begin(notifier); - continue; - } - break; - } + err =3D hmm_range_fault_unlocked_timeout(&hmm_range, remaining); mmput(mm); if (err) goto err_free; @@ -1720,8 +1693,7 @@ int drm_gpusvm_range_evict(struct drm_gpusvm *gpusvm, .end =3D drm_gpusvm_range_end(range), .dev_private_owner =3D NULL, }; - unsigned long timeout =3D - jiffies + msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); + unsigned long timeout =3D msecs_to_jiffies(HMM_RANGE_DEFAULT_TIMEOUT); unsigned long *pfns; unsigned long npages =3D npages_in_range(drm_gpusvm_range_start(range), drm_gpusvm_range_end(range)); @@ -1736,24 +1708,12 @@ int drm_gpusvm_range_evict(struct drm_gpusvm *gpusv= m, return -ENOMEM; =20 hmm_range.hmm_pfns =3D pfns; - while (!time_after(jiffies, timeout)) { - hmm_range.notifier_seq =3D mmu_interval_read_begin(notifier); - if (time_after(jiffies, timeout)) { - err =3D -ETIME; - break; - } - - mmap_read_lock(mm); - err =3D hmm_range_fault(&hmm_range); - mmap_read_unlock(mm); - if (err !=3D -EBUSY) - break; - } + err =3D hmm_range_fault_unlocked_timeout(&hmm_range, timeout); =20 kvfree(pfns); mmput(mm); =20 - return err; + return err =3D=3D -EBUSY ? -ETIME : err; } EXPORT_SYMBOL_GPL(drm_gpusvm_range_evict); =20 --=20 2.43.0