From nobody Thu Dec 18 07:21:06 2025 Received: from mail-oi1-f171.google.com (mail-oi1-f171.google.com [209.85.167.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 71BF73F9C5 for ; Wed, 1 May 2024 04:27:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1714537637; cv=none; b=uGkNmnxsnjDK1xa+osUcp9h/cJAbucs2z8Hk2s/7ippQxu5jQcMwjh1jv8yjqX4uAsppyP6SjvMhFPrp5aN9FZjxOAV/hEbTjJUP2xOeVPhq9d9EchvzJj/1vWS+fNduCs/FB+7yBz2+8O8wHU5UmudX9W8NT0CP5TwhhkmZivQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1714537637; c=relaxed/simple; bh=gWXuoid7dRKeMQUZwz4YAEAXa4JIxXi9Uf3wRsrmHxE=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=JeUO/W04FmGiv1tsIo14hiquvEKwtKJn622hd+Kw3jqWeGPJB4hffXkgYKvmuJcnUB6YoD3q6HnsMdGXF6xBbZPQA57BpiTRAiXaiAyNyWs88rSLFkQyF+Fz0KCZP2dzySBxnz515cp4LH+PuMjgfjMmUnsvdG8HNIuZ7S6kE6c= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=C6nA8Zz5; arc=none smtp.client-ip=209.85.167.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="C6nA8Zz5" Received: by mail-oi1-f171.google.com with SMTP id 5614622812f47-3c70999ff96so3357601b6e.2 for ; Tue, 30 Apr 2024 21:27:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1714537635; x=1715142435; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=qxwHmPfQRT1njG/0ZnhsAAHKQR4nWB4DSExpUtfTagU=; b=C6nA8Zz599momHHPwCc3iHD/zL+9GZcqMnTPEgIOmhK3IfhCpfHyWpbC/7Pg2z/lpL ymVCayPgM2/nKM6ANzbFfXuj/rLS9ZtXyL7oQZ1D3oMBni4DySZkQ/yHZvFjAhuZi27o T0z1H3syoKFI4cYXlNY8HEH2SnPu+cW0zDJYCkSMHoDLWHmdcMMn0vV3ZuRkpGSQ3Lvy SHRU5E3RBgrKWLb1VxNZr2HS9dYq9jbmS0c+5yn0hBqBq4jDc4e9wrdcgVWWn4aeOieq Cdm3V669FO0kjUQx3fCngT/xW26l5bYB6+1XSvFzYNLabX6TOMxdkv49G/gt9RwxaWw0 2vKQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1714537635; x=1715142435; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=qxwHmPfQRT1njG/0ZnhsAAHKQR4nWB4DSExpUtfTagU=; b=Lf0cy5c8h5NrfMVdd/r82CdZvliPlLV5YdtxOM8FrsOy44f6+WkQG9oy1joLyMc5xp LipDs9JHZy7p0Pv+Pre/RKfU/dUzyZ3a26W0/e/FNjTceeUwtHxvJRdMmQafBSM4Y1pE +XkP8syalgcDEJy38UlnxlVx2KJ/9NY32pSaPr1q92iCoF1E7/g5ErZg+JbznOO2ZCa8 Cp1rxOmBlvJorMRWToaibOJ8QP4ZIwyr80uVCyiQhdBNaIuHRO5pcR6sRAt/dbfkdQut XtVP8Q+cWiB+idzxC+7+W/q/x1NNZtIdcvVUfy00sdBR6cbpsae9h1Q4NdaIj9PdG7aa 1qgQ== X-Forwarded-Encrypted: i=1; AJvYcCWkeKKxEY+VRHws6LX8VD0gXcM/8QSyx5pMfKEg17aXlyjxqo9MAmJ8y+etejsc7kQiiqYypnzpMC9UHX/VqDBY0jTAkAaBYvF4OLyb X-Gm-Message-State: AOJu0YwD8mOCXGNMnn6Ox6ztWke0CMQ4b/LSG9d05vq0hach8Of03/wt hZjh1i+mjRGQlu+Tyjezo7R4lYZFcOWpZYInTVO4oJYCejxOgAtR X-Google-Smtp-Source: AGHT+IEl+/L1ituyaJTnk0BjJi3vRmGfqMR+zcexC0PXTPj0xeHACJfp1bfVqYeei+6Mig3AUi1xRg== X-Received: by 2002:a05:6808:1813:b0:3c9:147c:bd22 with SMTP id bh19-20020a056808181300b003c9147cbd22mr1266333oib.29.1714537635509; Tue, 30 Apr 2024 21:27:15 -0700 (PDT) Received: from LancedeMBP.lan ([112.10.225.242]) by smtp.gmail.com with ESMTPSA id m15-20020a656a0f000000b005dc4806ad7dsm19165970pgu.40.2024.04.30.21.27.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 30 Apr 2024 21:27:15 -0700 (PDT) From: Lance Yang To: akpm@linux-foundation.org Cc: willy@infradead.org, sj@kernel.org, maskray@google.com, ziy@nvidia.com, ryan.roberts@arm.com, david@redhat.com, 21cnbao@gmail.com, mhocko@suse.com, fengwei.yin@intel.com, zokeefe@google.com, shy828301@gmail.com, xiehuan09@gmail.com, libang.li@antgroup.com, wangkefeng.wang@huawei.com, songmuchun@bytedance.com, peterx@redhat.com, minchan@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Lance Yang Subject: [PATCH v4 1/3] mm/rmap: remove duplicated exit code in pagewalk loop Date: Wed, 1 May 2024 12:26:58 +0800 Message-Id: <20240501042700.83974-2-ioworker0@gmail.com> X-Mailer: git-send-email 2.33.1 In-Reply-To: <20240501042700.83974-1-ioworker0@gmail.com> References: <20240501042700.83974-1-ioworker0@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Introduce the labels walk_done and walk_done_err as exit points to eliminate duplicated exit code in the pagewalk loop. Signed-off-by: Lance Yang Tested-by: SeongJae Park --- mm/rmap.c | 40 +++++++++++++++------------------------- 1 file changed, 15 insertions(+), 25 deletions(-) diff --git a/mm/rmap.c b/mm/rmap.c index 7faa60bc3e4d..7e2575d669a9 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -1675,9 +1675,7 @@ static bool try_to_unmap_one(struct folio *folio, str= uct vm_area_struct *vma, /* Restore the mlock which got missed */ if (!folio_test_large(folio)) mlock_vma_folio(folio, vma); - page_vma_mapped_walk_done(&pvmw); - ret =3D false; - break; + goto walk_done_err; } =20 pfn =3D pte_pfn(ptep_get(pvmw.pte)); @@ -1715,11 +1713,8 @@ static bool try_to_unmap_one(struct folio *folio, st= ruct vm_area_struct *vma, */ if (!anon) { VM_BUG_ON(!(flags & TTU_RMAP_LOCKED)); - if (!hugetlb_vma_trylock_write(vma)) { - page_vma_mapped_walk_done(&pvmw); - ret =3D false; - break; - } + if (!hugetlb_vma_trylock_write(vma)) + goto walk_done_err; if (huge_pmd_unshare(mm, vma, address, pvmw.pte)) { hugetlb_vma_unlock_write(vma); flush_tlb_range(vma, @@ -1734,8 +1729,7 @@ static bool try_to_unmap_one(struct folio *folio, str= uct vm_area_struct *vma, * actual page and drop map count * to zero. */ - page_vma_mapped_walk_done(&pvmw); - break; + goto walk_done; } hugetlb_vma_unlock_write(vma); } @@ -1807,9 +1801,7 @@ static bool try_to_unmap_one(struct folio *folio, str= uct vm_area_struct *vma, if (unlikely(folio_test_swapbacked(folio) !=3D folio_test_swapcache(folio))) { WARN_ON_ONCE(1); - ret =3D false; - page_vma_mapped_walk_done(&pvmw); - break; + goto walk_done_err; } =20 /* MADV_FREE page check */ @@ -1848,23 +1840,17 @@ static bool try_to_unmap_one(struct folio *folio, s= truct vm_area_struct *vma, */ set_pte_at(mm, address, pvmw.pte, pteval); folio_set_swapbacked(folio); - ret =3D false; - page_vma_mapped_walk_done(&pvmw); - break; + goto walk_done_err; } =20 if (swap_duplicate(entry) < 0) { set_pte_at(mm, address, pvmw.pte, pteval); - ret =3D false; - page_vma_mapped_walk_done(&pvmw); - break; + goto walk_done_err; } if (arch_unmap_one(mm, vma, address, pteval) < 0) { swap_free(entry); set_pte_at(mm, address, pvmw.pte, pteval); - ret =3D false; - page_vma_mapped_walk_done(&pvmw); - break; + goto walk_done_err; } =20 /* See folio_try_share_anon_rmap(): clear PTE first. */ @@ -1872,9 +1858,7 @@ static bool try_to_unmap_one(struct folio *folio, str= uct vm_area_struct *vma, folio_try_share_anon_rmap_pte(folio, subpage)) { swap_free(entry); set_pte_at(mm, address, pvmw.pte, pteval); - ret =3D false; - page_vma_mapped_walk_done(&pvmw); - break; + goto walk_done_err; } if (list_empty(&mm->mmlist)) { spin_lock(&mmlist_lock); @@ -1914,6 +1898,12 @@ static bool try_to_unmap_one(struct folio *folio, st= ruct vm_area_struct *vma, if (vma->vm_flags & VM_LOCKED) mlock_drain_local(); folio_put(folio); + continue; +walk_done_err: + ret =3D false; +walk_done: + page_vma_mapped_walk_done(&pvmw); + break; } =20 mmu_notifier_invalidate_range_end(&range); --=20 2.33.1 From nobody Thu Dec 18 07:21:06 2025 Received: from mail-pf1-f173.google.com (mail-pf1-f173.google.com [209.85.210.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0DD6B433D5 for ; Wed, 1 May 2024 04:27:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.173 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1714537642; cv=none; b=X4UNz9JCZ+GRRZ0O4IjAp8nG4EZBLJazCQ3YjPl1/dTFOzZDwYURIM8noKEeJSj6FeQzb/FUgZoz36Nne2Lqv96ULmZe21yuT2CQoObr7DqWuLJpazDUwmQS1bHV3co6IW4IaqwJp/O9plWIKTHUOJn8f27OMOwEF3X33RCMylQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1714537642; c=relaxed/simple; bh=uV2/14QUugpYgH1LYX0hwV/yJj37Hyi7gXmkQUiVLdk=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=RjZKIHt1H+SeWY6EuZ1/UNZn7nlH5sKb0MV07ikhU8QZQXFCvVm2WQyHn43XkYLgFj+HWyWPPlXgfCe6SPOqDHjvS1Mpp2W3bcczOvWlY7mwV/05s51vizDKG99x3ZMd/mJAPdXQAA+B3eCOSjuyp3HezzW0LM8lkMDkx6vt3CE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=SOqTeFp/; arc=none smtp.client-ip=209.85.210.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="SOqTeFp/" Received: by mail-pf1-f173.google.com with SMTP id d2e1a72fcca58-6f0aeee172dso366943b3a.1 for ; Tue, 30 Apr 2024 21:27:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1714537640; x=1715142440; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=ROFqZ6lfkIoUlwY4du6p7l8Lb0MkeIYOr3uqszlhzvk=; b=SOqTeFp/DHlvDeGUEzwOznd9SXymQehtJab0fVQJ+LthkK/EGpsdl3QiiIELqeUxg6 PfpNegWgldBzc6Trj/4pAN10dmbtXZoA+HG6wwldafX1eVruybiylFAv4a85eKyu3UEx SVr0UGxIkc0Xp67WxK5kakl0qyRh+iwY5KiSJo6yhh0t77vGkMTvco3QavB3wgG1OZWq Q94bGtFk/qHl2izXNePihycDJndJtBdMnOFIozGCPd1ZRjjzB03nFD91Be3bNKI/46iw LuEUiiZ/MhdOku2v7asnNrN5VtVxdCrwvMXbGHFHUdOzgGoqhMh5uf16I7ReuW2rLd4g T4XQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1714537640; x=1715142440; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=ROFqZ6lfkIoUlwY4du6p7l8Lb0MkeIYOr3uqszlhzvk=; b=Bd6UdYW30Hz6WXh9jryYDP6uV2N0ned/3jxDgH8SI6y3VBrg18/82oBQfV6VT0UkMM lP63EQGsrGiXRrxP2g9tqCA/bHg9zG3aFGE9aOQt7UFWWTze34UtgXZR/WvXmJr1AZI9 bkCXOfOkWAdx7anCiVs0UylI7PRY8aln9bUUDDF722BulzUMjNPgeUrmmb4aQJXAJGCc qTmIZpEhzDUFARg1LcB2cQPjwYprrSGBqiTjaDFfh+WuywZQ8fSdOAkPnr1SNFbqJhjq sT26LlGfA4AQ47e0qmREeQbEhI5qNIxY24W8rWVd8bAFCWL2OJxeWpVJhisWqYMgCGzm 9HZw== X-Forwarded-Encrypted: i=1; AJvYcCVt4g210taKe57HWiQfL+c8tWj97FRMOKp2+KYMOOoWu5Yvxe9MFOIgP/9GdOi6xFyDciKMFHiJajedoUAYXUEVwMUVlwUNqHbMh8S3 X-Gm-Message-State: AOJu0Yx5BxF5AWe5LHHI32x7Ar3+zW5C0Iai2SnT46lIIxRCO3ibjmaw nwOeCDFO1Jf+Yk2RauhDGGJJBIBIolSjBEWETpRYsIoZjEEbrmu9 X-Google-Smtp-Source: AGHT+IF2W0eVQxBQaMDY8AWoFZ/AvARxWOBrEttWhO+PAqpgYqb/8VwaVvuvijythAV5TvB4SnJUzQ== X-Received: by 2002:a05:6a20:431a:b0:1ad:4978:adf4 with SMTP id h26-20020a056a20431a00b001ad4978adf4mr7731975pzk.1.1714537640184; Tue, 30 Apr 2024 21:27:20 -0700 (PDT) Received: from LancedeMBP.lan ([112.10.225.242]) by smtp.gmail.com with ESMTPSA id m15-20020a656a0f000000b005dc4806ad7dsm19165970pgu.40.2024.04.30.21.27.15 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 30 Apr 2024 21:27:19 -0700 (PDT) From: Lance Yang To: akpm@linux-foundation.org Cc: willy@infradead.org, sj@kernel.org, maskray@google.com, ziy@nvidia.com, ryan.roberts@arm.com, david@redhat.com, 21cnbao@gmail.com, mhocko@suse.com, fengwei.yin@intel.com, zokeefe@google.com, shy828301@gmail.com, xiehuan09@gmail.com, libang.li@antgroup.com, wangkefeng.wang@huawei.com, songmuchun@bytedance.com, peterx@redhat.com, minchan@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Lance Yang Subject: [PATCH v4 2/3] mm/rmap: integrate PMD-mapped folio splitting into pagewalk loop Date: Wed, 1 May 2024 12:26:59 +0800 Message-Id: <20240501042700.83974-3-ioworker0@gmail.com> X-Mailer: git-send-email 2.33.1 In-Reply-To: <20240501042700.83974-1-ioworker0@gmail.com> References: <20240501042700.83974-1-ioworker0@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" In preparation for supporting try_to_unmap_one() to unmap PMD-mapped folios, start the pagewalk first, then call split_huge_pmd_address() to split the folio. Suggested-by: David Hildenbrand Signed-off-by: Lance Yang Tested-by: SeongJae Park --- include/linux/huge_mm.h | 20 ++++++++++++++++++++ mm/huge_memory.c | 42 +++++++++++++++++++++-------------------- mm/rmap.c | 24 +++++++++++++++++------ 3 files changed, 60 insertions(+), 26 deletions(-) diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h index c8d3ec116e29..38c4b5537715 100644 --- a/include/linux/huge_mm.h +++ b/include/linux/huge_mm.h @@ -409,6 +409,20 @@ static inline bool thp_migration_supported(void) return IS_ENABLED(CONFIG_ARCH_ENABLE_THP_MIGRATION); } =20 +void split_huge_pmd_locked(struct vm_area_struct *vma, unsigned long addre= ss, + pmd_t *pmd, bool freeze, struct folio *folio); + +static inline void align_huge_pmd_range(struct vm_area_struct *vma, + unsigned long *start, + unsigned long *end) +{ + *start =3D ALIGN(*start, HPAGE_PMD_SIZE); + *end =3D ALIGN_DOWN(*end, HPAGE_PMD_SIZE); + + VM_WARN_ON_ONCE(vma->vm_start > *start); + VM_WARN_ON_ONCE(vma->vm_end < *end); +} + #else /* CONFIG_TRANSPARENT_HUGEPAGE */ =20 static inline bool folio_test_pmd_mappable(struct folio *folio) @@ -471,6 +485,12 @@ static inline void __split_huge_pmd(struct vm_area_str= uct *vma, pmd_t *pmd, unsigned long address, bool freeze, struct folio *folio) {} static inline void split_huge_pmd_address(struct vm_area_struct *vma, unsigned long address, bool freeze, struct folio *folio) {} +static inline void split_huge_pmd_locked(struct vm_area_struct *vma, + unsigned long address, pmd_t *pmd, + bool freeze, struct folio *folio) {} +static inline void align_huge_pmd_range(struct vm_area_struct *vma, + unsigned long *start, + unsigned long *end) {} =20 #define split_huge_pud(__vma, __pmd, __address) \ do { } while (0) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 8261b5669397..145505a1dd05 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2584,6 +2584,27 @@ static void __split_huge_pmd_locked(struct vm_area_s= truct *vma, pmd_t *pmd, pmd_populate(mm, pmd, pgtable); } =20 +void split_huge_pmd_locked(struct vm_area_struct *vma, unsigned long addre= ss, + pmd_t *pmd, bool freeze, struct folio *folio) +{ + VM_WARN_ON_ONCE(folio && !folio_test_pmd_mappable(folio)); + VM_WARN_ON_ONCE(!IS_ALIGNED(address, HPAGE_PMD_SIZE)); + VM_WARN_ON_ONCE(folio && !folio_test_locked(folio)); + VM_BUG_ON(freeze && !folio); + + /* + * When the caller requests to set up a migration entry, we + * require a folio to check the PMD against. Otherwise, there + * is a risk of replacing the wrong folio. + */ + if (pmd_trans_huge(*pmd) || pmd_devmap(*pmd) || + is_pmd_migration_entry(*pmd)) { + if (folio && folio !=3D pmd_folio(*pmd)) + return; + __split_huge_pmd_locked(vma, pmd, address, freeze); + } +} + void __split_huge_pmd(struct vm_area_struct *vma, pmd_t *pmd, unsigned long address, bool freeze, struct folio *folio) { @@ -2595,26 +2616,7 @@ void __split_huge_pmd(struct vm_area_struct *vma, pm= d_t *pmd, (address & HPAGE_PMD_MASK) + HPAGE_PMD_SIZE); mmu_notifier_invalidate_range_start(&range); ptl =3D pmd_lock(vma->vm_mm, pmd); - - /* - * If caller asks to setup a migration entry, we need a folio to check - * pmd against. Otherwise we can end up replacing wrong folio. - */ - VM_BUG_ON(freeze && !folio); - VM_WARN_ON_ONCE(folio && !folio_test_locked(folio)); - - if (pmd_trans_huge(*pmd) || pmd_devmap(*pmd) || - is_pmd_migration_entry(*pmd)) { - /* - * It's safe to call pmd_page when folio is set because it's - * guaranteed that pmd is present. - */ - if (folio && folio !=3D pmd_folio(*pmd)) - goto out; - __split_huge_pmd_locked(vma, pmd, range.start, freeze); - } - -out: + split_huge_pmd_locked(vma, range.start, pmd, freeze, folio); spin_unlock(ptl); mmu_notifier_invalidate_range_end(&range); } diff --git a/mm/rmap.c b/mm/rmap.c index 7e2575d669a9..432601154583 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -1636,9 +1636,6 @@ static bool try_to_unmap_one(struct folio *folio, str= uct vm_area_struct *vma, if (flags & TTU_SYNC) pvmw.flags =3D PVMW_SYNC; =20 - if (flags & TTU_SPLIT_HUGE_PMD) - split_huge_pmd_address(vma, address, false, folio); - /* * For THP, we have to assume the worse case ie pmd for invalidation. * For hugetlb, it could be much worse if we need to do pud @@ -1650,6 +1647,8 @@ static bool try_to_unmap_one(struct folio *folio, str= uct vm_area_struct *vma, range.end =3D vma_address_end(&pvmw); mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, vma->vm_mm, address, range.end); + if (flags & TTU_SPLIT_HUGE_PMD) + align_huge_pmd_range(vma, &range.start, &range.end); if (folio_test_hugetlb(folio)) { /* * If sharing is possible, start and end will be adjusted @@ -1664,9 +1663,6 @@ static bool try_to_unmap_one(struct folio *folio, str= uct vm_area_struct *vma, mmu_notifier_invalidate_range_start(&range); =20 while (page_vma_mapped_walk(&pvmw)) { - /* Unexpected PMD-mapped THP? */ - VM_BUG_ON_FOLIO(!pvmw.pte, folio); - /* * If the folio is in an mlock()d vma, we must not swap it out. */ @@ -1678,6 +1674,22 @@ static bool try_to_unmap_one(struct folio *folio, st= ruct vm_area_struct *vma, goto walk_done_err; } =20 + if (!pvmw.pte && (flags & TTU_SPLIT_HUGE_PMD)) { + /* + * We temporarily have to drop the PTL and start once + * again from that now-PTE-mapped page table. + */ + split_huge_pmd_locked(vma, range.start, pvmw.pmd, false, + folio); + pvmw.pmd =3D NULL; + spin_unlock(pvmw.ptl); + flags &=3D ~TTU_SPLIT_HUGE_PMD; + continue; + } + + /* Unexpected PMD-mapped THP? */ + VM_BUG_ON_FOLIO(!pvmw.pte, folio); + pfn =3D pte_pfn(ptep_get(pvmw.pte)); subpage =3D folio_page(folio, pfn - folio_pfn(folio)); address =3D pvmw.address; --=20 2.33.1 From nobody Thu Dec 18 07:21:06 2025 Received: from mail-pf1-f169.google.com (mail-pf1-f169.google.com [209.85.210.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9E2034437A for ; Wed, 1 May 2024 04:27:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1714537647; cv=none; b=ZGSDP7vZtyAQwcB/snAwEHj25InIbUh5Hf2CH89NZiXfBOW9u7T8IH/KczKLNrwZ81R7exevQYekWDWRAvUDW1UGcXCsJPncyeGMmL2VG3o7oKNx4xSxq21cLSJeT9o96g+KAzVCQwnDuDoFyDTsLjYPjW+cfCV2FvpX87OQj2I= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1714537647; c=relaxed/simple; bh=6sHQDtC40RV0vFu1QW1zzKktIZ+5vhLAKVpXfbHtYwI=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=IJ1MHZ+Jhmod9ADH/VewtQD6zQXyclYbvGiNFe9PHZVWS0/GLeI9jTQeEIyALpL2M0wTkenIRnmmvEjIWp0uXHIUYZzyQrlo6K5RoSkzq9Ag2TkM7Ms014r+vpVeVhZzXeXIKqL4rAF9tjm9mAj6yiRgFquTZiHji+KcqHg6y0A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=TzXtF7jS; arc=none smtp.client-ip=209.85.210.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="TzXtF7jS" Received: by mail-pf1-f169.google.com with SMTP id d2e1a72fcca58-6ed9fc77bbfso5124486b3a.1 for ; Tue, 30 Apr 2024 21:27:25 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1714537645; x=1715142445; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=hOzQapCM9/nhQFCiXcuQFHHgx8BSo7e2vP87qJAgbT0=; b=TzXtF7jSBWuz5DdzLV2xfFUR7FmxjpR6jaRMJhPqVQFvSPT4QlAmvn1Hcz1r7yCcIk NUuzfFCRlSjufrXQN/bETS94gY43dc7F45yP4A7VGNHJY4p3otguqQUDEQnqLyqM2Quu JwLU5nOVz7kuIpPjhLhyaSfcygUbZ1L7khq8KW/1d4LI2sHMcypbeGtG91WcCVn2xn5Q MCzBzC3TwIuPOPVWkBv0LoPvMnsqtXUbJkU/YlJJx9EwD5y3MvLGkjef7kLtJmkxtT3J yS4+aKhSj5iN95eR8bRKCtFdVBJd9oFZ3uHTeG3T84DjvZA5heP1gQqxRbOM6TE4PlWO VgdA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1714537645; x=1715142445; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=hOzQapCM9/nhQFCiXcuQFHHgx8BSo7e2vP87qJAgbT0=; b=H+KIXkzpuFyjCbRsbDLpLNvVdJ8E0InzP73DlTocipG2ZvI8iP0ZBETxejvYYinia6 xDqpCfEKrTAblA+SZ3GK9qFznOW416DVa6TEuqsmF8jdx19MhCAT9hedvy2FNQKHl3hu ocMPuoOAp05UrGKkA6C3q+msLSjYmSnPWA/sD+el7Dqxo4/sxGGPpqQsQwrJVkiEtkkK q0QR2AZBVNuMJuMff+a40JlsdZhLRj3h+eFxLlg7Cn3bFjrkmCbmTg8dxlXkgq2jtG39 MDHlE0BrkN2p0DM3xtF0an/ufqpDjOo1UrXX6kxuFCpFv3jv5fNX7xjd9am7frUaG/sc GH+w== X-Forwarded-Encrypted: i=1; AJvYcCWUJXNuuWUjxSG9w2EA6myMiz5Vg2wOsGN2/EmJFU4ofjXnX5LDoiJgLd7/xnLSVDYPp5WW4Oy4vlpkHFn6Bjsc5xT+kkvutcz3Ii+A X-Gm-Message-State: AOJu0Yy3eDWUE7dSSEIm4jR0pXyPZA/9XA0JTOY7JNCa57Cel4ZlImXt INaNhlx4ATs66Hx7qGaXFkJqaXr8aS/5euSTZVm0u5Hs8ynD7GOk X-Google-Smtp-Source: AGHT+IExH8/mQKcC+a0yuiE7nJbe/y3Zn478CUKsgvKf8sffmUuOcVVFi/q1aIYi5BB/0pju90QCCg== X-Received: by 2002:a05:6a00:2345:b0:6ed:d8d2:503d with SMTP id j5-20020a056a00234500b006edd8d2503dmr2150218pfj.20.1714537644912; Tue, 30 Apr 2024 21:27:24 -0700 (PDT) Received: from LancedeMBP.lan ([112.10.225.242]) by smtp.gmail.com with ESMTPSA id m15-20020a656a0f000000b005dc4806ad7dsm19165970pgu.40.2024.04.30.21.27.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 30 Apr 2024 21:27:24 -0700 (PDT) From: Lance Yang To: akpm@linux-foundation.org Cc: willy@infradead.org, sj@kernel.org, maskray@google.com, ziy@nvidia.com, ryan.roberts@arm.com, david@redhat.com, 21cnbao@gmail.com, mhocko@suse.com, fengwei.yin@intel.com, zokeefe@google.com, shy828301@gmail.com, xiehuan09@gmail.com, libang.li@antgroup.com, wangkefeng.wang@huawei.com, songmuchun@bytedance.com, peterx@redhat.com, minchan@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, Lance Yang Subject: [PATCH v4 3/3] mm/vmscan: avoid split lazyfree THP during shrink_folio_list() Date: Wed, 1 May 2024 12:27:00 +0800 Message-Id: <20240501042700.83974-4-ioworker0@gmail.com> X-Mailer: git-send-email 2.33.1 In-Reply-To: <20240501042700.83974-1-ioworker0@gmail.com> References: <20240501042700.83974-1-ioworker0@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" When the user no longer requires the pages, they would use madvise(MADV_FREE) to mark the pages as lazy free. Subsequently, they typically would not re-write to that memory again. During memory reclaim, if we detect that the large folio and its PMD are both still marked as clean and there are no unexpected references (such as GUP), so we can just discard the memory lazily, improving the efficiency of memory reclamation in this case. On an Intel i5 CPU, reclaim= ing 1GiB of lazyfree THPs using mem_cgroup_force_empty() results in the following runtimes in seconds (shorter is better): Suggested-by: David Hildenbrand Suggested-by: Zi Yan Tested-by: SeongJae Park -------------------------------------------- | Old | New | Change | -------------------------------------------- | 0.683426 | 0.049197 | -92.80% | -------------------------------------------- Suggested-by: Zi Yan Suggested-by: David Hildenbrand Signed-off-by: Lance Yang --- include/linux/huge_mm.h | 9 +++++ mm/huge_memory.c | 73 +++++++++++++++++++++++++++++++++++++++++ mm/rmap.c | 3 ++ 3 files changed, 85 insertions(+) diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h index 38c4b5537715..017cee864080 100644 --- a/include/linux/huge_mm.h +++ b/include/linux/huge_mm.h @@ -411,6 +411,8 @@ static inline bool thp_migration_supported(void) =20 void split_huge_pmd_locked(struct vm_area_struct *vma, unsigned long addre= ss, pmd_t *pmd, bool freeze, struct folio *folio); +bool unmap_huge_pmd_locked(struct vm_area_struct *vma, unsigned long addr, + pmd_t *pmdp, struct folio *folio); =20 static inline void align_huge_pmd_range(struct vm_area_struct *vma, unsigned long *start, @@ -492,6 +494,13 @@ static inline void align_huge_pmd_range(struct vm_area= _struct *vma, unsigned long *start, unsigned long *end) {} =20 +static inline bool unmap_huge_pmd_locked(struct vm_area_struct *vma, + unsigned long addr, pmd_t *pmdp, + struct folio *folio) +{ + return false; +} + #define split_huge_pud(__vma, __pmd, __address) \ do { } while (0) =20 diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 145505a1dd05..90fdef847a88 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2690,6 +2690,79 @@ static void unmap_folio(struct folio *folio) try_to_unmap_flush(); } =20 +static bool __discard_trans_pmd_locked(struct vm_area_struct *vma, + unsigned long addr, pmd_t *pmdp, + struct folio *folio) +{ + struct mm_struct *mm =3D vma->vm_mm; + int ref_count, map_count; + pmd_t orig_pmd =3D *pmdp; + struct mmu_gather tlb; + struct page *page; + + if (pmd_dirty(orig_pmd) || folio_test_dirty(folio)) + return false; + if (unlikely(!pmd_present(orig_pmd) || !pmd_trans_huge(orig_pmd))) + return false; + + page =3D pmd_page(orig_pmd); + if (unlikely(page_folio(page) !=3D folio)) + return false; + + tlb_gather_mmu(&tlb, mm); + orig_pmd =3D pmdp_huge_get_and_clear(mm, addr, pmdp); + tlb_remove_pmd_tlb_entry(&tlb, pmdp, addr); + + /* + * Syncing against concurrent GUP-fast: + * - clear PMD; barrier; read refcount + * - inc refcount; barrier; read PMD + */ + smp_mb(); + + ref_count =3D folio_ref_count(folio); + map_count =3D folio_mapcount(folio); + + /* + * Order reads for folio refcount and dirty flag + * (see comments in __remove_mapping()). + */ + smp_rmb(); + + /* + * If the PMD or folio is redirtied at this point, or if there are + * unexpected references, we will give up to discard this folio + * and remap it. + * + * The only folio refs must be one from isolation plus the rmap(s). + */ + if (ref_count !=3D map_count + 1 || folio_test_dirty(folio) || + pmd_dirty(orig_pmd)) { + set_pmd_at(mm, addr, pmdp, orig_pmd); + return false; + } + + folio_remove_rmap_pmd(folio, page, vma); + zap_deposited_table(mm, pmdp); + add_mm_counter(mm, MM_ANONPAGES, -HPAGE_PMD_NR); + folio_put(folio); + + return true; +} + +bool unmap_huge_pmd_locked(struct vm_area_struct *vma, unsigned long addr, + pmd_t *pmdp, struct folio *folio) +{ + VM_WARN_ON_FOLIO(!folio_test_pmd_mappable(folio), folio); + VM_WARN_ON_FOLIO(!folio_test_locked(folio), folio); + VM_WARN_ON_ONCE(!IS_ALIGNED(addr, HPAGE_PMD_SIZE)); + + if (folio_test_anon(folio) && !folio_test_swapbacked(folio)) + return __discard_trans_pmd_locked(vma, addr, pmdp, folio); + + return false; +} + static void remap_page(struct folio *folio, unsigned long nr) { int i =3D 0; diff --git a/mm/rmap.c b/mm/rmap.c index 432601154583..1d3d30cb752c 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -1675,6 +1675,9 @@ static bool try_to_unmap_one(struct folio *folio, str= uct vm_area_struct *vma, } =20 if (!pvmw.pte && (flags & TTU_SPLIT_HUGE_PMD)) { + if (unmap_huge_pmd_locked(vma, range.start, pvmw.pmd, + folio)) + goto walk_done; /* * We temporarily have to drop the PTL and start once * again from that now-PTE-mapped page table. --=20 2.33.1