From nobody Mon Aug 24 07:02:17 2026 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E7D873DCD80; Wed, 5 Aug 2026 19:35:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958546; cv=none; b=fcjtaHvkXYE8BNM/cK8kEXsTnKjCy8b8uDaK62NFs9fX+fqbrb3zcAEiaXaS0iuIJfU6EfudWmXwLjbtUI1f1GZKO9Xa2iPGsbycChdKi4zOCGgjnS580JNPJaLYOObsh82UXrfUfUVINrjQ2Irpgf2aafKG8rDwee13JVCWReQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958546; c=relaxed/simple; bh=0YG7iyYcbF3G1EFKh9ioNjadG3W1OhFnqKUemvuHbQY=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version:Content-Type; b=iQzsGLYLQ5DA2iJ4tFsx48njpZL9nF/vT5axRrdtaZMXO5LrIEyO6Rmso0rkRlfxtKhY1xzo+UIudPy9j6TagjcXXig2wKKY0sxdv+rPdzMQhCyWGsvtbpkyG4fTa3RKCnOI2tQSsHjTsvPeqes7vciYUQ7YHbdJnqDol4o28gI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=WxIEp9Uy; arc=none smtp.client-ip=198.175.65.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="WxIEp9Uy" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1785958544; x=1817494544; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=0YG7iyYcbF3G1EFKh9ioNjadG3W1OhFnqKUemvuHbQY=; b=WxIEp9UyybAK5OHwZ+yljQuyX5fEFKaSMzyYwoI5T6znV+qmzQpf2jGY i2VhcB8pkTfzPuIPoMrrJJXxcIypOoEl7y9EfAEedtJRaUgY74l6ilMHh YXdQ9dRf/HwSIYFUCHRnkDuSAxdUQOAaQ2raNzpc5J+xWXi3m+c3ohNTl tDGsw/iOy2WAwmYksEyjb1ZV4o0Gb1OiNuPUhb3eY3l4e/2E7QT65s0xn pSrDcFJCcllsagTiht5oyaVJvfZ/sx/cPitcnNazQVMUkGGNHRvk85xaL gSlEI/97B1WMnBRITsK31UgYAmz/digX6+lQtn4E50HYkPsSlzYlH0v+I Q==; X-CSE-ConnectionGUID: H+0PxfBzS/G21SqIgdc+VA== X-CSE-MsgGUID: ztfvh61TQKe9UyXQMmf0OA== X-IronPort-AV: E=McAfee;i="6800,10657,11866"; a="109332918" X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="109332918" Received: from fmviesa002.fm.intel.com ([10.60.135.142]) by orvoesa101.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:42 -0700 X-CSE-ConnectionGUID: +IlgXOyXTU+bAoeSzcnfsg== X-CSE-MsgGUID: aH+F8fM8RHmiW/7pNd/qeA== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="285263519" Received: from gsse-cloud1.jf.intel.com ([10.54.39.91]) by fmviesa002-auth.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:42 -0700 From: Matthew Brost To: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: Sashiko , Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R . Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Balbir Singh , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , =?UTF-8?q?Thomas=20Hellstr=C3=B6m?= , Francois Dugast , stable@vger.kernel.org Subject: [PATCH v2 1/5] mm/migrate_device: Clear MIGRATE_PFN_MIGRATE on all sub-folios of a split THP Date: Wed, 5 Aug 2026 12:35:32 -0700 Message-Id: <20260805193536.3756457-2-matthew.brost@intel.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260805193536.3756457-1-matthew.brost@intel.com> References: <20260805193536.3756457-1-matthew.brost@intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable migrate_vma_split_unmapped_folio() propagates MIGRATE_PFN_MIGRATE from the head entry to all HPAGE_PMD_NR entries of src_pfns[]. The two bailouts below it in __migrate_device_pages() only cleared the head, and the "next" label then advances by @nr, so the tails keep the flag and a valid destination without ever going through folio_migrate_mapping(). migrate_vma_finalize() then maps unpopulated destination folios into userspace. Clear the flag across the whole @nr range at both bailouts. Reported-by: Sashiko Fixes: 4265d67e405a ("mm/migrate_device: add THP splitting during migration= ") Cc: Andrew Morton Cc: David Hildenbrand Cc: Lorenzo Stoakes Cc: Zi Yan Cc: Baolin Wang Cc: Liam R. Howlett Cc: Nico Pache Cc: Ryan Roberts Cc: Dev Jain Cc: Barry Song Cc: Lance Yang Cc: Usama Arif Cc: Joshua Hahn Cc: Rakie Kim Cc: Byungchul Park Cc: Gregory Price Cc: Ying Huang Cc: Alistair Popple Cc: Balbir Singh Cc: Maarten Lankhorst Cc: Maxime Ripard Cc: Thomas Zimmermann Cc: David Airlie Cc: Simona Vetter Cc: Thomas Hellstr=C3=B6m Cc: Francois Dugast Cc: dri-devel@lists.freedesktop.org Cc: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org Cc: stable@vger.kernel.org Assisted-by: GitHub_Copilot:claude-opus-5 Signed-off-by: Matthew Brost --- mm/migrate_device.c | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-) diff --git a/mm/migrate_device.c b/mm/migrate_device.c index 908d2d4ec43a..d37a96cc6335 100644 --- a/mm/migrate_device.c +++ b/mm/migrate_device.c @@ -1199,10 +1199,14 @@ static void __migrate_device_pages(unsigned long *s= rc_pfns, * device private or coherent memory. * * Try to get rid of swap cache if possible. + * + * @folio may have been split into @nr folios + * above, so clear all of them. */ if (!folio_test_anon(folio) || !folio_free_swap(folio)) { - src_pfns[i] &=3D ~MIGRATE_PFN_MIGRATE; + for (j =3D 0; j < nr && i + j < npages; j++) + src_pfns[i+j] &=3D ~MIGRATE_PFN_MIGRATE; goto next; } } @@ -1210,7 +1214,8 @@ static void __migrate_device_pages(unsigned long *src= _pfns, /* * Other types of ZONE_DEVICE page are not supported. */ - src_pfns[i] &=3D ~MIGRATE_PFN_MIGRATE; + for (j =3D 0; j < nr && i + j < npages; j++) + src_pfns[i+j] &=3D ~MIGRATE_PFN_MIGRATE; goto next; } =20 --=20 2.34.1 From nobody Mon Aug 24 07:02:17 2026 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 139FA399003; Wed, 5 Aug 2026 19:35:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958547; cv=none; b=Z3i0WbajKzHV1LxYjqAeC6Oz44BAfCiHSTCxaka+CE6u6rwnB4r5ZpLIJMSszLDq8Lw47lfp+E/SZPn3kti5+WUA4Jcdhqsv3r3bSUcL0JAFvcoomhRz9Yy8ih6WEWMJQH73oZLFjORTkSZEUbJP9WPlemt4siY/L4CnL2mI/lM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958547; c=relaxed/simple; bh=Q8T+Js7ezU9vaLDf0JKr8untYueFAoOH+rF/XAfETZI=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version:Content-Type; b=Nd0WYg+DKFxu3kHzVKVwmDgoncYDmQTDu31ga/KhknM19S3Ri0uxvsTPf0R7nPo1u8t7b54SCdAGd1L2WrXvg8BaJug6opMGbpwWCiyHF+wzLLt+YuNU7emunV7QkYaUfDEBKp9PFNRiNqGJ+I7l16zIM1qq5p69n7hQk5RzEVc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=HLDSJwL9; arc=none smtp.client-ip=198.175.65.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="HLDSJwL9" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1785958545; x=1817494545; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=Q8T+Js7ezU9vaLDf0JKr8untYueFAoOH+rF/XAfETZI=; b=HLDSJwL96qSu/QcY1SelnSrjbxGxB5cCJeevJwIjcVJkFIwUAgahrEBd A7QlsSmDSDdSqPi0jmH/XaoDDZXfwvuQ4Kiw6t6eIbldPQxzCdz+6eAy2 bzQJWG7c+xZ1tyaFvVRnLRcqglEfsOPwgjyX+1siiyfIl0tc3Qkvp3Iib QKFWW82D33w2VpGMJLfik9mRvF08DPtqIXkb4jD2sO2e9QP3h6T3XHo2U 28LlsXkIw7ylPHHB2CayEk7Vu+w5Aesa3Z7/IlBsO+2vQw2LF5XJSm9BZ yZ/cIzkvPXNIhgSJM36iSz4Ivug1eO+zMeAyxR88KrudS3ktKrwmmH0S6 Q==; X-CSE-ConnectionGUID: ZXcxo3WXSliFfXYNdyn4zQ== X-CSE-MsgGUID: Wio2e10jTEaLeC+qkpcayw== X-IronPort-AV: E=McAfee;i="6800,10657,11866"; a="109332938" X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="109332938" Received: from fmviesa002.fm.intel.com ([10.60.135.142]) by orvoesa101.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:43 -0700 X-CSE-ConnectionGUID: KHiaN7iMRTau0I+DyND8PQ== X-CSE-MsgGUID: 71gF13OgSAuigqw8aZtv1Q== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="285263524" Received: from gsse-cloud1.jf.intel.com ([10.54.39.91]) by fmviesa002-auth.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:42 -0700 From: Matthew Brost To: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R . Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Balbir Singh , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , =?UTF-8?q?Thomas=20Hellstr=C3=B6m?= , Francois Dugast , stable@vger.kernel.org Subject: [PATCH v2 2/5] mm/migrate_device: Fix THP splitting of a CPU faulted device private folio Date: Wed, 5 Aug 2026 12:35:33 -0700 Message-Id: <20260805193536.3756457-3-matthew.brost@intel.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260805193536.3756457-1-matthew.brost@intel.com> References: <20260805193536.3756457-1-matthew.brost@intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable When a CPU faults on a device private PMD and the device driver can only allocate order-0 destination folios, __migrate_device_pages() has to split the source THP via migrate_vma_split_unmapped_folio(). That path is broken in two independent ways when the fault is what triggered the migration. First, the split never succeeds. At the point folio_split_unmapped() is called the folio carries two references beyond the ones it is entitled to: 1 - taken by do_huge_pmd_device_private() for the duration of the ->migrate_to_ram() callback 2 - taken by migrate_vma_collect_huge_pmd() when the folio was collected (the mapping reference having been dropped by set_pmd_migration_entry()). folio_split_unmapped() requires folio_expected_ref_count(folio) =3D=3D folio_ref_count(folio) - 1, i.e. it tolerates exactly one caller reference. With both of the above held the check sees 2 against an expected 0 and returns -EAGAIN, so the migration is abandoned and the CPU fault makes no progress. The PTE-based split path does not have this problem: migrate_vma_split_folio() is called before any collect reference is taken and explicitly skips folio_get() for the fault folio, so the fault reference is the single caller reference the split expects. Fix it by dropping the fault reference across the split and re-taking it afterwards. do_huge_pmd_device_private() derives the fault page from the PMD entry, so it is always the head page of the folio and always ends up in the head folio of an uniform split to order 0; re-taking the reference on the folio therefore puts it back exactly where do_huge_pmd_device_private() will release it. The folio cannot be freed while the reference is dropped because the collect reference is still held. Second, the folio is split globally but the page tables were demoted only locally: split_huge_pmd_address(migrate->vma, addr, true); ret =3D folio_split_unmapped(folio, 0); migrate_device_unmap() unmaps via try_to_migrate(folio, 0), deliberately without TTU_SPLIT_HUGE_PMD, so every VMA that PMD maps the folio is left holding a PMD sized migration entry. A folio that was PMD mapped in more than one VMA -- after fork(), for example -- therefore keeps huge migration entries in all the other VMAs while only migrate->vma is demoted. folio_split_unmapped() does not notice: the folio is fully unmapped, so it only looks at the refcount and happily splits to order 0. The other VMAs are then left pointing a huge PMD at an order-0 folio, and migrate_vma_finalize() -> remove_migration_ptes() walks into it: page dumped because: VM_BUG_ON_FOLIO(folio_test_hugetlb(folio) || !folio_test_pmd_mappable(folio)) kernel BUG at mm/migrate.c:368! RIP: 0010:remove_migration_pte+0x56a/0x9b0 Call Trace: rmap_walk_anon+0xfc/0x260 remove_migration_ptes+0x79/0xb0 __migrate_device_finalize+0x113/0x290 __drm_pagemap_migrate_to_ram+0x278/0x360 [drm_gpusvm_helper] drm_pagemap_migrate_to_ram+0x5c/0x80 [drm_gpusvm_helper] do_huge_pmd_device_private+0x160/0x280 Without CONFIG_DEBUG_VM the VM_BUG_ON_FOLIO() is compiled out and remove_migration_pmd() installs a huge PMD pointing at an order-0 page instead, along with add_mm_counter(mm, MM_ANONPAGES, HPAGE_PMD_NR). The victim mm then maps 2MB of address space onto a single 4K page, which shows up later as bad rss-counter state, leaked page tables and page allocator freelist corruption in unrelated processes. Note this second problem was latent before the refcount fix above: the split always failed, and the failed attempt left migrate->vma demoted, so the retried fault took the PTE path, where __folio_split() unmaps with TTU_SPLIT_HUGE_PMD and demotes every VMA. Fix it by walking the rmap and demoting every PMD sized migration entry mapping the folio before splitting it. Demote with freeze =3D false: entry creation in __split_huge_pmd_locked() is dispatched on pmd_is_migration_entry(), not on freeze, so a migration PMD becomes PTE sized migration entries either way, and freeze only controls a trailing put_page(). With freeze =3D false there is no refcount change at all, which makes the demotion idempotent across N VMAs. rmap_walk_control.anon_lock is deliberately left unset: folio_lock_anon_vma_read() depends on folio_mapped(), and the folio is already fully unmapped here. This mirrors remove_migration_ptes(). Finally, refuse the split for a folio that is not anonymous. The rmap walk would otherwise reach a file backed VMA, where split_huge_pmd_address() zaps the PMD instead of demoting it. Fixes: 4265d67e405a ("mm/migrate_device: add THP splitting during migration= ") Cc: Andrew Morton Cc: David Hildenbrand Cc: Lorenzo Stoakes Cc: Zi Yan Cc: Baolin Wang Cc: Liam R. Howlett Cc: Nico Pache Cc: Ryan Roberts Cc: Dev Jain Cc: Barry Song Cc: Lance Yang Cc: Usama Arif Cc: Joshua Hahn Cc: Rakie Kim Cc: Byungchul Park Cc: Gregory Price Cc: Ying Huang Cc: Alistair Popple Cc: Balbir Singh Cc: Maarten Lankhorst Cc: Maxime Ripard Cc: Thomas Zimmermann Cc: David Airlie Cc: Simona Vetter Cc: Thomas Hellstr=C3=B6m Cc: Francois Dugast Cc: dri-devel@lists.freedesktop.org Cc: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org Cc: stable@vger.kernel.org Assisted-by: GitHub_Copilot:claude-opus-5 Signed-off-by: Matthew Brost --- mm/migrate_device.c | 98 ++++++++++++++++++++++++++++++++++++++++----- 1 file changed, 89 insertions(+), 9 deletions(-) diff --git a/mm/migrate_device.c b/mm/migrate_device.c index d37a96cc6335..142920a464d8 100644 --- a/mm/migrate_device.c +++ b/mm/migrate_device.c @@ -899,22 +899,104 @@ static int migrate_vma_insert_huge_pmd_page(struct m= igrate_vma *migrate, return 0; } =20 +static bool migrate_vma_split_pmd_one(struct folio *folio, + struct vm_area_struct *vma, + unsigned long addr, void *arg) +{ + DEFINE_FOLIO_VMA_WALK(pvmw, folio, vma, addr, PVMW_SYNC | PVMW_MIGRATION); + + while (page_vma_mapped_walk(&pvmw)) { + if (pvmw.pte) + continue; + + addr =3D pvmw.address; + page_vma_mapped_walk_done(&pvmw); + + /* + * Demote with freeze =3D false: the PMD already holds a + * migration entry, so __split_huge_pmd_locked() creates PTE + * sized migration entries from it and leaves the refcount + * alone. There is at most one PMD mapping @folio per VMA, so + * stop the walk here. + */ + split_huge_pmd_address(vma, addr, false); + break; + } + + return true; +} + +/* + * Demote every PMD sized migration entry that maps @folio to PTE sized on= es. + * + * migrate_device_unmap() unmaps with try_to_migrate(folio, 0), i.e. witho= ut + * TTU_SPLIT_HUGE_PMD, so a folio that was PMD mapped in several VMAs -- a= fter + * fork(), for instance -- ends up with a PMD sized migration entry in eve= ry one + * of them. folio_split_unmapped() below does not care, it only looks at t= he + * refcount, so splitting the folio without demoting all of those first wo= uld + * leave the other VMAs pointing a huge PMD at what is now an order-0 foli= o. + * remove_migration_ptes() trips over that in migrate_vma_finalize(). + */ +static void migrate_vma_split_pmd_mappings(struct folio *folio) +{ + struct rmap_walk_control rwc =3D { + .rmap_one =3D migrate_vma_split_pmd_one, + }; + + /* + * Do not pass .anon_lock: folio_lock_anon_vma_read() requires + * folio_mapped(), and @folio is already fully unmapped here. + */ + rmap_walk(folio, &rwc); +} + static int migrate_vma_split_unmapped_folio(struct migrate_vma *migrate, - unsigned long idx, unsigned long addr, + unsigned long idx, struct folio *folio) { unsigned long i; unsigned long pfn; unsigned long flags; + bool fault_folio; int ret =3D 0; =20 /* - * take a reference, since split_huge_pmd_address() with freeze =3D true - * drops a reference at the end. + * migrate_vma_split_pmd_mappings() walks the rmap, and + * split_huge_pmd_address() zaps rather than demotes a PMD in a VMA that + * is not anonymous. migrate_vma_collect_huge_pmd() does not check the + * VMA type, so a file THP can reach here; the rest of the migrate_vma() + * machinery only supports anonymous memory anyway. */ - folio_get(folio); - split_huge_pmd_address(migrate->vma, addr, true); + if (!folio_test_anon(folio)) + return -EINVAL; + + /* + * A CPU fault on a device private PMD holds an extra reference on the + * folio, taken by do_huge_pmd_device_private(). folio_split_unmapped() + * only tolerates a single caller reference, so the split would always + * fail with -EAGAIN while this fault reference is held. + * + * do_huge_pmd_device_private() derives the fault page from the PMD + * entry, so it is always the head page of @folio, and therefore always + * ends up in the head folio after an uniform split to order 0. Drop + * the reference across the split and re-take it on the head folio + * afterwards, leaving the reference exactly where it is expected to be + * released. + * + * The folio cannot go away while the reference is dropped: the + * reference taken by migrate_vma_collect_huge_pmd() is still held. + */ + fault_folio =3D migrate->fault_page && + page_folio(migrate->fault_page) =3D=3D folio; + + migrate_vma_split_pmd_mappings(folio); + + if (fault_folio) + folio_put(folio); ret =3D folio_split_unmapped(folio, 0); + if (fault_folio) + folio_get(folio); + if (ret) return ret; migrate->src[idx] &=3D ~MIGRATE_PFN_COMPOUND; @@ -935,7 +1017,7 @@ static int migrate_vma_insert_huge_pmd_page(struct mig= rate_vma *migrate, } =20 static int migrate_vma_split_unmapped_folio(struct migrate_vma *migrate, - unsigned long idx, unsigned long addr, + unsigned long idx, struct folio *folio) { return 0; @@ -1103,7 +1185,6 @@ static void __migrate_device_pages(unsigned long *src= _pfns, struct mmu_notifier_range range; unsigned long i, j; bool notified =3D false; - unsigned long addr; =20 for (i =3D 0; i < npages; ) { struct page *newpage =3D migrate_pfn_to_page(dst_pfns[i]); @@ -1177,8 +1258,7 @@ static void __migrate_device_pages(unsigned long *src= _pfns, goto next; } nr =3D 1 << folio_order(folio); - addr =3D migrate->start + i * PAGE_SIZE; - if (migrate_vma_split_unmapped_folio(migrate, i, addr, folio)) { + if (migrate_vma_split_unmapped_folio(migrate, i, folio)) { src_pfns[i] &=3D ~(MIGRATE_PFN_MIGRATE | MIGRATE_PFN_COMPOUND); goto next; --=20 2.34.1 From nobody Mon Aug 24 07:02:17 2026 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 005CC3E49EF; Wed, 5 Aug 2026 19:35:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958547; cv=none; b=gGEV7vIDnNbuZNBQhKFEvXI6k4zG8VK1s6brM7+ltkojwvJVpByC/JxEY6xYYy3v0FC8HhLhKFDrqQ/1S4eHHEdlpDHSBinqg9gKbj49zE6dabLHyK+pKqDTZKEpkt4xS7xKJHUmhVC6bZEhI0k+CIqsZsnTbO8IY9qf67aY1HQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958547; c=relaxed/simple; bh=DwXjrp0YOMKqBIa0CyQY66LJvCKzj79JzvSlROjOOxA=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version:Content-Type; b=skvmkiJWSRA6zXICYts+H43IjwRhgL1Ze4lJzZG7avpipGuvtFblcJYd81SwofIVr89D7QmoerliDuc9MNl6XZXVWn3BrjIL9Jidfnoqj2B8jXsSx51mkfi6+IrA7QypZdN2Bi1y7/g9kPqasYttqfGWVaqjaA3VDdnz9EWr1lA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=Qk/CHIeV; arc=none smtp.client-ip=198.175.65.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="Qk/CHIeV" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1785958546; x=1817494546; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=DwXjrp0YOMKqBIa0CyQY66LJvCKzj79JzvSlROjOOxA=; b=Qk/CHIeVJn38sklPcAFJDamTISxuj/S4tz3BjbcDb4gCH6eTL9DMkene Px95NR9oh0XKdNytqH6qDNQDWkkO4lILpFVuA8RxgwW9rs/oHOZNqj55D Wmr7cDr3KeL+ONFrpUXoeIKKNDIXHk1YVpfY6CuuPHlY83BbArT+fA75g Rbhiq4YJ6twrbF0X02WA2+PUGPrDgIzIQ9jYFXqs+hlfnniqp48o6cEIh Q3ArWMioRNPV8aJMEJ2eCMSwVB4VtwNmfpEm/uYqrsqHiuJUFhSSRQZcB FgSjcQI8GqTgNTqq5q6AAFhPJ9gi7lAQT7WtKexG0s1kHkrW9zYQJyZC4 g==; X-CSE-ConnectionGUID: upvVpFnEQf6as8jJxHrgHg== X-CSE-MsgGUID: 5gQgvINeQ7eFcYOS0oOeKQ== X-IronPort-AV: E=McAfee;i="6800,10657,11866"; a="109332954" X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="109332954" Received: from fmviesa002.fm.intel.com ([10.60.135.142]) by orvoesa101.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:44 -0700 X-CSE-ConnectionGUID: CujDtT0BQmaBotcPTuJkwA== X-CSE-MsgGUID: sX6rfqqERo6FgsdinoenMQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="285263527" Received: from gsse-cloud1.jf.intel.com ([10.54.39.91]) by fmviesa002-auth.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:43 -0700 From: Matthew Brost To: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R . Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Balbir Singh , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , =?UTF-8?q?Thomas=20Hellstr=C3=B6m?= , Francois Dugast , stable@vger.kernel.org Subject: [PATCH v2 3/5] mm/migrate_device: Apply the fault reference to the correct folio Date: Wed, 5 Aug 2026 12:35:34 -0700 Message-Id: <20260805193536.3756457-4-matthew.brost@intel.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260805193536.3756457-1-matthew.brost@intel.com> References: <20260805193536.3756457-1-matthew.brost@intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable __migrate_device_pages() computed extra_cnt once, from the head page of the source folio, and then passed the same value to folio_migrate_mapping() for every one of the @nr sub-folios produced by migrate_vma_split_unmapped_folio(). The extra reference the CPU fault holds only exists on the single folio that ends up containing the fault page. Claiming it for all of them makes folio_migrate_mapping() expect one reference too many on every other sub-folio, so it returns -EAGAIN and MIGRATE_PFN_MIGRATE is cleared for them. Only the fault page would migrate; the remaining HPAGE_PMD_NR - 1 pages would be restored to device memory, and the faulting access would immediately fault again. Compute extra_cnt per sub-folio instead, comparing against the source page for that entry. While at it, use the sub-folio's own mapping rather than the mapping that was read from the pre-split folio. This has been latent so far because the split it depends on could never succeed while the fault reference was held. Fixes: 4265d67e405a ("mm/migrate_device: add THP splitting during migration= ") Cc: Andrew Morton Cc: David Hildenbrand Cc: Lorenzo Stoakes Cc: Zi Yan Cc: Baolin Wang Cc: Liam R. Howlett Cc: Nico Pache Cc: Ryan Roberts Cc: Dev Jain Cc: Barry Song Cc: Lance Yang Cc: Usama Arif Cc: Joshua Hahn Cc: Rakie Kim Cc: Byungchul Park Cc: Gregory Price Cc: Ying Huang Cc: Alistair Popple Cc: Balbir Singh Cc: Maarten Lankhorst Cc: Maxime Ripard Cc: Thomas Zimmermann Cc: David Airlie Cc: Simona Vetter Cc: Thomas Hellstr=C3=B6m Cc: Francois Dugast Cc: dri-devel@lists.freedesktop.org Cc: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org Cc: stable@vger.kernel.org Assisted-by: GitHub_Copilot:claude-opus-5 Signed-off-by: Matthew Brost --- mm/migrate_device.c | 22 +++++++++++++++++----- 1 file changed, 17 insertions(+), 5 deletions(-) diff --git a/mm/migrate_device.c b/mm/migrate_device.c index 142920a464d8..4ee09801efe6 100644 --- a/mm/migrate_device.c +++ b/mm/migrate_device.c @@ -1191,7 +1191,7 @@ static void __migrate_device_pages(unsigned long *src= _pfns, struct page *page =3D migrate_pfn_to_page(src_pfns[i]); struct address_space *mapping; struct folio *newfolio, *folio; - int r, extra_cnt =3D 0; + int r; unsigned long nr =3D 1; =20 if (!newpage) { @@ -1301,13 +1301,25 @@ static void __migrate_device_pages(unsigned long *s= rc_pfns, =20 BUG_ON(folio_test_writeback(folio)); =20 - if (migrate && migrate->fault_page =3D=3D page) - extra_cnt =3D 1; for (j =3D 0; j < nr && i + j < npages; j++) { - folio =3D page_folio(migrate_pfn_to_page(src_pfns[i+j])); + struct page *src_page =3D migrate_pfn_to_page(src_pfns[i+j]); + int extra_cnt =3D 0; + + folio =3D page_folio(src_page); newfolio =3D page_folio(migrate_pfn_to_page(dst_pfns[i+j])); =20 - r =3D folio_migrate_mapping(mapping, newfolio, folio, extra_cnt); + /* + * The CPU fault holds an extra reference on the folio + * containing the fault page. @folio may have been + * split above, so the fault page only accounts for an + * extra reference on the folio it actually ended up + * in, not on every folio of the original THP. + */ + if (migrate && migrate->fault_page =3D=3D src_page) + extra_cnt =3D 1; + + r =3D folio_migrate_mapping(folio_mapping(folio), newfolio, + folio, extra_cnt); if (r) src_pfns[i+j] &=3D ~MIGRATE_PFN_MIGRATE; else --=20 2.34.1 From nobody Mon Aug 24 07:02:17 2026 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 74F0E3F4DFD; Wed, 5 Aug 2026 19:35:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958551; cv=none; b=U+/YUEGsW6WOK/zIJxsTlBcaQUdSy2Icv48E8N5+blmCXGoKWTU3yTGge2CEtWYVqmTNy2dGq1LL85Z7GNDRR9CH85mPVRh9B1OcYPyNU7YEwFMwO2tv44YvBeEzczCSSoXomf4ddfqFhNN463eo891NQpQ9hpR4Jh440XKxS9c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958551; c=relaxed/simple; bh=zZRj+b/NzvOD47gFxglRIPtyT0lDGG2KHr/65GQvaMk=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version:Content-Type; b=JxP4GPXnIj6XxjICet9tbZZU9sy8bhHiA1yDIeynREJIzdkZwBbfyjOy5bXz7OCU8OJe9L2qysqJJw3bxtdNIY+pPvm3m8//plpvE66t+Gz6Hmb4OXtJgSi9yXPPTQVpwSw3+zgNPC3/mmei5luoB8TsfHH516CK80/D6iDlPOk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=NQ871CCP; arc=none smtp.client-ip=198.175.65.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="NQ871CCP" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1785958546; x=1817494546; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=zZRj+b/NzvOD47gFxglRIPtyT0lDGG2KHr/65GQvaMk=; b=NQ871CCPmZq6iFjaAPhog7UBiX91R63iQ9jD3xQLvjUvqc8xz9dHgZxP T0H6/L7EPNAmD3yHe/Mii2a9e9MLVWGE1lYTJyFBNXKWcD306ziDtSgZU jEgQTEBlE+v0gna7kw+paXbxPdqv4XpAHdKaP9UcklRIdJlPfqpoxLwU7 skfgyoUDmaTjCVZTE8JV1pWHnbIj9vCc5eBZQ9R1L4/u0nta0T7uRdHCT I6xALy2pdLAUfsUfsMZr9gBclRhVukCZ+8ZdyfOAqeQQGKQXG6Ykixq3+ eu4vERzJej5FO89NI9vuC0JN1CEbGuxdX6MG+mb6DxTl9WDeb8qS9MDKU Q==; X-CSE-ConnectionGUID: bLJmhXgGROWSTSLolIvJBw== X-CSE-MsgGUID: hlmdfQPZQ0eWDQtgqd3uIw== X-IronPort-AV: E=McAfee;i="6800,10657,11866"; a="109332968" X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="109332968" Received: from fmviesa002.fm.intel.com ([10.60.135.142]) by orvoesa101.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:44 -0700 X-CSE-ConnectionGUID: 9lSYi9ToSZiwsdkFe3a+OQ== X-CSE-MsgGUID: SAeRpc4fQWWc+k0u9sIZ1Q== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="285263531" Received: from gsse-cloud1.jf.intel.com ([10.54.39.91]) by fmviesa002-auth.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:43 -0700 From: Matthew Brost To: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R . Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Balbir Singh , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , =?UTF-8?q?Thomas=20Hellstr=C3=B6m?= , Francois Dugast , stable@vger.kernel.org Subject: [PATCH v2 4/5] drm/pagemap: Fix folio allocation fallback and use-after-put Date: Wed, 5 Aug 2026 12:35:35 -0700 Message-Id: <20260805193536.3756457-5-matthew.brost@intel.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260805193536.3756457-1-matthew.brost@intel.com> References: <20260805193536.3756457-1-matthew.brost@intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable drm_pagemap_migrate_populate_ram_pfn() had two issues when populating RAM PFNs with higher-order folios: 1. The higher-order vma_alloc_folio()/folio_alloc() calls did not pass __GFP_NOWARN, so a THP allocation failure under memory pressure would spam the kernel log, and there was no fallback path despite a TODO comment stating one was needed. Add __GFP_NOWARN to the higher-order allocation and, on failure, fall back to order-0 allocations for the entire range originally covered by the failed higher-order allocation, leaving MIGRATE_PFN_COMPOUND unset for those PFNs. 2. In the free_pages error path, order was computed via folio_order(page_folio(page)) *after* put_page(page) had already dropped the reference, resulting in a use-after-free/put when that was the last reference on the page. Compute order before releasing the page. Introducing the fallback in 1. also requires the source page array handed to ->copy_to_ram() to be built differently. Both callers only populated the entry at the head of each source folio, relying on the copy callback to derive the rest of the folio from the order recorded in the matching drm_pagemap_addr. Once the destination has been demoted to order-0 folios the drm_pagemap_addr entries are per-page, so a source page is needed for every one of them; leaving them NULL makes the copy callback stop after the first page and the remainder of the range is never copied. The source folio is only split later, by migrate_vma_pages() / migrate_device_pages(), so its order cannot be used to detect the demotion - test the destination for MIGRATE_PFN_COMPOUND instead. Factor the array population out into drm_pagemap_migrate_populate_src_pages() and use it from both drm_pagemap_evict_to_ram() and __drm_pagemap_migrate_to_ram(). Fixes: ddeda6136038 ("drm/pagemap: Allocate folios when possible") Cc: Andrew Morton Cc: David Hildenbrand Cc: Lorenzo Stoakes Cc: Zi Yan Cc: Baolin Wang Cc: Liam R. Howlett Cc: Nico Pache Cc: Ryan Roberts Cc: Dev Jain Cc: Barry Song Cc: Lance Yang Cc: Usama Arif Cc: Joshua Hahn Cc: Rakie Kim Cc: Byungchul Park Cc: Gregory Price Cc: Ying Huang Cc: Alistair Popple Cc: Balbir Singh Cc: Maarten Lankhorst Cc: Maxime Ripard Cc: Thomas Zimmermann Cc: David Airlie Cc: Simona Vetter Cc: Thomas Hellstr=C3=B6m Cc: Francois Dugast Cc: dri-devel@lists.freedesktop.org Cc: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org Cc: stable@vger.kernel.org Assisted-by: GitHub_Copilot:claude-opus-5 Signed-off-by: Matthew Brost --- v2: Add THP-mid-PMD invariant (Sashiko) --- drivers/gpu/drm/drm_pagemap.c | 128 +++++++++++++++++++++++++++------- 1 file changed, 103 insertions(+), 25 deletions(-) diff --git a/drivers/gpu/drm/drm_pagemap.c b/drivers/gpu/drm/drm_pagemap.c index 892b325fa99b..05eb7254028f 100644 --- a/drivers/gpu/drm/drm_pagemap.c +++ b/drivers/gpu/drm/drm_pagemap.c @@ -383,6 +383,58 @@ drm_pagemap_migrate_map_system_pages(struct device *de= v, return 0; } =20 +/** + * drm_pagemap_migrate_populate_src_pages() - Populate the source page arr= ay + * @pages: Array of source pages to populate + * @src_mpfn: Source array of migrate PFNs + * @dst_mpfn: Destination array of migrate PFNs + * @npages: Number of pages in the arrays + * + * Populate @pages with the device pages the copy callback is to read from. + * + * Entries are normally only populated at the head of each source folio, w= ith + * the copy callback deriving the rest of the folio from the order recorde= d in + * the corresponding drm_pagemap_addr. That does not work where + * drm_pagemap_migrate_populate_ram_pfn() had to demote a higher-order sou= rce + * folio to order-0 destination folios: the drm_pagemap_addr entries are t= hen + * per-page, and the copy callback needs a source page for each of them. + * Populate every entry for those ranges. + * + * Note that the source folio itself is only split later, by + * migrate_vma_pages() / migrate_device_pages(), so its order cannot be us= ed to + * detect the demotion - the destination has to be inspected instead. + */ +static void drm_pagemap_migrate_populate_src_pages(struct page **pages, + unsigned long *src_mpfn, + unsigned long *dst_mpfn, + unsigned long npages) +{ + unsigned long i; + + for (i =3D 0; i < npages;) { + struct page *page =3D migrate_pfn_to_page(src_mpfn[i]); + unsigned int order =3D 0; + unsigned long j, nr; + + if (!page) { + i++; + continue; + } + + order =3D folio_order(page_folio(page)); + nr =3D NR_PAGES(order); + + if (order && !(dst_mpfn[i] & MIGRATE_PFN_COMPOUND)) { + for (j =3D 0; j < nr && i + j < npages; j++) + pages[i + j] =3D folio_page(page_folio(page), j); + } else { + pages[i] =3D page; + } + + i +=3D nr; + } +} + /** * drm_pagemap_migrate_unmap_pages() - Unmap pages previously mapped for G= PU SVM migration * @dev: The device for which the pages were mapped @@ -875,6 +927,7 @@ static int drm_pagemap_migrate_populate_ram_pfn(struct = vm_area_struct *vas, struct page *page =3D NULL, *src_page; struct folio *folio; unsigned int order =3D 0; + gfp_t gfp =3D GFP_HIGHUSER; =20 if (!(src_mpfn[i] & MIGRATE_PFN_MIGRATE)) goto next; @@ -891,11 +944,51 @@ static int drm_pagemap_migrate_populate_ram_pfn(struc= t vm_area_struct *vas, =20 order =3D folio_order(page_folio(src_page)); =20 - /* TODO: Support fallback to single pages if THP allocation fails */ + /* + * A large source folio is always collected whole, at its head + * page, PMD aligned and flagged MIGRATE_PFN_COMPOUND: anything + * else is split before it reaches us, either by + * migrate_vma_collect_pmd() or, for the eviction path, by + * migrate_device_pfns(). Both the order-0 fallback below and + * drm_pagemap_migrate_populate_src_pages() rely on that, as + * they index the folio from @i. + */ + WARN_ON_ONCE(order && + (src_page !=3D folio_page(page_folio(src_page), 0) || + !(src_mpfn[i] & MIGRATE_PFN_COMPOUND))); + + if (order) + gfp |=3D __GFP_NOWARN; + if (vas) - folio =3D vma_alloc_folio(GFP_HIGHUSER, order, vas, addr); + folio =3D vma_alloc_folio(gfp, order, vas, addr); else - folio =3D folio_alloc(GFP_HIGHUSER, order); + folio =3D folio_alloc(gfp, order); + + if (!folio && order) { + /* + * Higher-order allocation failed, fall back to + * order-0 allocations for the entire range covered + * by the original higher-order allocation, without + * setting MIGRATE_PFN_COMPOUND, until we move past + * that range. + */ + unsigned long nr =3D NR_PAGES(order); + unsigned long j; + + gfp &=3D ~__GFP_NOWARN; + for (j =3D 0; j < nr && i < npages; j++, i++, addr +=3D PAGE_SIZE) { + folio =3D vas ? + vma_alloc_folio(gfp, 0, vas, addr) : + folio_alloc(gfp, 0); + if (!folio) + goto free_pages; + + page =3D folio_page(folio, 0); + mpfn[i] =3D migrate_pfn(page_to_pfn(page)); + } + continue; + } =20 if (!folio) goto free_pages; @@ -940,11 +1033,11 @@ static int drm_pagemap_migrate_populate_ram_pfn(stru= ct vm_area_struct *vas, if (!page) goto next_put; =20 + order =3D folio_order(page_folio(page)); + put_page(page); mpfn[i] =3D 0; =20 - order =3D folio_order(page_folio(page)); - next_put: i +=3D NR_PAGES(order); } @@ -1120,7 +1213,7 @@ int drm_pagemap_evict_to_ram(struct drm_pagemap_devme= m *devmem_allocation) unsigned long *src, *dst; struct drm_pagemap_addr *pagemap_addr; void *buf; - int i, err =3D 0; + int err =3D 0; unsigned int retry_count =3D 2; =20 npages =3D devmem_allocation->size >> PAGE_SHIFT; @@ -1160,15 +1253,7 @@ int drm_pagemap_evict_to_ram(struct drm_pagemap_devm= em *devmem_allocation) if (err) goto err_finalize; =20 - for (i =3D 0; i < npages;) { - unsigned int order =3D 0; - - pages[i] =3D migrate_pfn_to_page(src[i]); - if (pages[i]) - order =3D folio_order(page_folio(pages[i])); - - i +=3D NR_PAGES(order); - } + drm_pagemap_migrate_populate_src_pages(pages, src, dst, npages); =20 err =3D ops->copy_to_ram(pages, pagemap_addr, npages, NULL); if (err) @@ -1235,7 +1320,7 @@ static int __drm_pagemap_migrate_to_ram(struct vm_are= a_struct *vas, struct drm_pagemap_addr *pagemap_addr; unsigned long start, end; void *buf; - int i, err =3D 0; + int err =3D 0; =20 zdd =3D drm_pagemap_page_zone_device_data(page); if (time_before64(get_jiffies_64(), zdd->devmem_allocation->timeslice_exp= iration)) @@ -1290,15 +1375,8 @@ static int __drm_pagemap_migrate_to_ram(struct vm_ar= ea_struct *vas, if (err) goto err_finalize; =20 - for (i =3D 0; i < npages;) { - unsigned int order =3D 0; - - pages[i] =3D migrate_pfn_to_page(migrate.src[i]); - if (pages[i]) - order =3D folio_order(page_folio(pages[i])); - - i +=3D NR_PAGES(order); - } + drm_pagemap_migrate_populate_src_pages(pages, migrate.src, migrate.dst, + npages); =20 err =3D ops->copy_to_ram(pages, pagemap_addr, npages, NULL); if (err) --=20 2.34.1 From nobody Mon Aug 24 07:02:17 2026 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.9]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5BF114248C0 for ; Wed, 5 Aug 2026 19:35:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.9 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958557; cv=none; b=CZZ0K7Ekn5Rwh4IvjrZA7np81pMfA3ND3O0I3aGrtw+taa9xDGP9v9l3nGNQhXhSDqwBBsKDh+3qoMMSjJpi/25l9tJg6IBKwJTfCPhgn0VZI+ZhY7+oyBJfpqd/d4ggVCzuWGp6hYHPYTdNCVYOip0Lo/FcKZWeoNJtC3UyBOA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785958557; c=relaxed/simple; bh=QYVG22Ril2YRFp2fcAJ/FBSXr0ANtT+bhitFTakpD74=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version:Content-Type; b=XBwonBO98aljjmi9znkqMoj4gunPCib0JTbvwpeiB5Vn5EOyOeqLspbc+43uJMj5cF6VyiqNpYrq9/q5Cux6gk/mpOaBE9ruS/I6LhUPjeeWNnHrIJM+CpEj2ZoJ/7jmsndohKluQbvGlAq0EYpW9Suzf9faqqea+hSTWkxCQ9o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=UrNPm7wa; arc=none smtp.client-ip=198.175.65.9 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="UrNPm7wa" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1785958547; x=1817494547; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=QYVG22Ril2YRFp2fcAJ/FBSXr0ANtT+bhitFTakpD74=; b=UrNPm7wa+ReTWzY0x1XjnJhlhjELjNpc2AY8WAGjyDlv0CEpNXvgbZOA 6q7TU4bzhjxHlBRjOHoN3vRVxqvcHN0pv2uDazuP/Dr4iBF9bRezGy0Fq vfI+iddk1unHywr45cWU5jGlXEBH6GnT7E6+QAUBmvndJ21TxvBEbAy6A +vg3deWGUs8R9QX0/UmyZrbVEl6CyAD6dS+jTCXoSV6Rg8FFcqzcZZe5K QYstMHDOW3piwVp6J4AnFZIYQdZ39IwUly+IgZkzo0BpxsOwOwe5C5C7L 9E3lo6U/6o6qXlq7Kpx9jnXHRRPiPFJd6x2aXTMZ4FIdOdd/9VXk89RHb w==; X-CSE-ConnectionGUID: fHhO9HYPSDiBO+Q9GkYcwg== X-CSE-MsgGUID: KUkSSfb+QvabhhC//nbFog== X-IronPort-AV: E=McAfee;i="6800,10657,11866"; a="109332985" X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="109332985" Received: from fmviesa002.fm.intel.com ([10.60.135.142]) by orvoesa101.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:45 -0700 X-CSE-ConnectionGUID: /fGuctSQS9q21fF+FOoZZg== X-CSE-MsgGUID: 9KQEbQ3VSVe6d7wvAWHz2Q== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,207,1779174000"; d="scan'208";a="285263536" Received: from gsse-cloud1.jf.intel.com ([10.54.39.91]) by fmviesa002-auth.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 05 Aug 2026 12:35:44 -0700 From: Matthew Brost To: intel-xe@lists.freedesktop.org, dri-devel@lists.freedesktop.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Cc: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R . Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Balbir Singh , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , =?UTF-8?q?Thomas=20Hellstr=C3=B6m?= , Francois Dugast Subject: [PATCH v2 5/5] drm/pagemap: Add fault injection for higher-order RAM folio allocation Date: Wed, 5 Aug 2026 12:35:36 -0700 Message-Id: <20260805193536.3756457-6-matthew.brost@intel.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260805193536.3756457-1-matthew.brost@intel.com> References: <20260805193536.3756457-1-matthew.brost@intel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Migrating a device-private THP back to system memory has two distinct paths in __migrate_device_pages(): the fast path where both source and destination carry MIGRATE_PFN_COMPOUND, and the fallback path where the destination could only be satisfied with order-0 folios and the source THP therefore has to be split via migrate_vma_split_unmapped_folio(). The fallback path only triggers under genuine memory pressure, which makes it both rare and awkward to reproduce, yet it is the path where the interesting refcounting happens (the CPU fault holds an extra reference on the device folio taken by do_huge_pmd_device_private()). Add a fault_attr, modelled on backup_fault_inject in ttm_pool.c, that forces the higher-order allocation in drm_pagemap_migrate_populate_ram_pfn() to fail so the existing order-0 fallback is taken deterministically. The attribute is exposed at /sys/kernel/debug/drm_pagemap_fault_inject and requires CONFIG_FAULT_INJECTION_DEBUG_FS. With CONFIG_FAULT_INJECTION disabled the helper compiles out to a constant false and the injection has no cost. Cc: Andrew Morton Cc: David Hildenbrand Cc: Lorenzo Stoakes Cc: Zi Yan Cc: Baolin Wang Cc: Liam R. Howlett Cc: Nico Pache Cc: Ryan Roberts Cc: Dev Jain Cc: Barry Song Cc: Lance Yang Cc: Usama Arif Cc: Joshua Hahn Cc: Rakie Kim Cc: Byungchul Park Cc: Gregory Price Cc: Ying Huang Cc: Alistair Popple Cc: Balbir Singh Cc: Maarten Lankhorst Cc: Maxime Ripard Cc: Thomas Zimmermann Cc: David Airlie Cc: Simona Vetter Cc: Thomas Hellstr=C3=B6m Cc: Francois Dugast Cc: dri-devel@lists.freedesktop.org Cc: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org Assisted-by: GitHub_Copilot:claude-opus-5 Signed-off-by: Matthew Brost --- drivers/gpu/drm/drm_pagemap.c | 36 ++++++++++++++++++++++++++++++++++- 1 file changed, 35 insertions(+), 1 deletion(-) diff --git a/drivers/gpu/drm/drm_pagemap.c b/drivers/gpu/drm/drm_pagemap.c index 05eb7254028f..01c639da767c 100644 --- a/drivers/gpu/drm/drm_pagemap.c +++ b/drivers/gpu/drm/drm_pagemap.c @@ -3,6 +3,7 @@ * Copyright =C2=A9 2024-2025 Intel Corporation */ =20 +#include #include #include #include @@ -12,6 +13,27 @@ #include #include =20 +#ifdef CONFIG_FAULT_INJECTION +#include +static DECLARE_FAULT_ATTR(migrate_to_ram_fault_inject); + +/* + * Force a higher-order destination folio allocation to fail in + * drm_pagemap_migrate_populate_ram_pfn(), exercising the order-0 fallback + * (and, in turn, the THP split path in __migrate_device_pages()) without + * having to drive the system into actual memory pressure. + */ +static bool drm_pagemap_fault_inject_folio(void) +{ + return should_fail(&migrate_to_ram_fault_inject, 1); +} +#else +static bool drm_pagemap_fault_inject_folio(void) +{ + return false; +} +#endif + /** * DOC: Overview * @@ -960,7 +982,9 @@ static int drm_pagemap_migrate_populate_ram_pfn(struct = vm_area_struct *vas, if (order) gfp |=3D __GFP_NOWARN; =20 - if (vas) + if (order && drm_pagemap_fault_inject_folio()) + folio =3D NULL; + else if (vas) folio =3D vma_alloc_folio(gfp, order, vas, addr); else folio =3D folio_alloc(gfp, order); @@ -1554,6 +1578,16 @@ void drm_pagemap_destroy(struct drm_pagemap *dpagema= p, bool is_atomic_or_reclaim kfree(dpagemap); } =20 +static int __init drm_pagemap_module_init(void) +{ +#if defined(CONFIG_DEBUG_FS) && defined(CONFIG_FAULT_INJECTION) + fault_create_debugfs_attr("drm_pagemap_fault_inject", NULL, + &migrate_to_ram_fault_inject); +#endif + return 0; +} +module_init(drm_pagemap_module_init); + static void drm_pagemap_exit(void) { flush_work(&drm_pagemap_work); --=20 2.34.1