From nobody Fri Jul 24 04:48:26 2026 Received: from mail-wm1-f69.google.com (mail-wm1-f69.google.com [209.85.128.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B5A433B9608 for ; Thu, 23 Jul 2026 18:21:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784830920; cv=none; b=UUPw+P5Yoy6WX1KTPDer/QYICpHybu4yzvIc7kxM/hcdrK2N7o0GF7W2NRlS+FtBD1VAM+EZywfr1ySZbZN4Ru9ATq1uAwLJw4Q7nZ1Pp1gscEqQFQ41HcmGdceVF8YUS1DS6wYhcHIOMY0zmTq5x06qJ3iiO/8bo5SB49OLpMY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784830920; c=relaxed/simple; bh=mMr1WmosQxXB2EcGfwJiQaiR9DVyE3hjJJoEXxfwmHY=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=ojwQ8LzLgRWvYh7tDGFbpJoFBCpX8YyF8bwQ0Y8tK0rtQw6Po60LPrbtqUSSo3ZKOxSEgeh57JasOBvC71YxLadFgNR1sy6TCLn0NhRPhQk/KlracK0qWsnvDXpcdiwHx/yhaGl+k0Ccp2yoNgc7WtiGiyKjkGJLAZCcmIrDB1M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--smostafa.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=bZVi5euH; arc=none smtp.client-ip=209.85.128.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--smostafa.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="bZVi5euH" Received: by mail-wm1-f69.google.com with SMTP id 5b1f17b1804b1-4955ce558d8so8929645e9.3 for ; Thu, 23 Jul 2026 11:21:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784830908; x=1785435708; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=9FpdytHdhxDISqer6403yWjEFqnIJW3e2H1iU7+f5M0=; b=bZVi5euHXY9WGrnNFYaeYUO9ZeeXX+KvEDPPjBEZoQ3o2/FPjyd5ngTsPvTL2b9IOO x6r/aurifUU2SdlJu7Dv8zDdFaUg2LzvzTMTv68TeLO0VDP2JPA+otEXnyPO4PmCFnHi EC7gDHaiye9HRue9qiXUFmpydyAy26T+n8Ta2UpAIVV2p4BmkYAMMjyyTdMF1VrWU4eX US4Q9DUdBF1oktfz6B9uNozlNZNC6BZxlMyZIungKDS0BAMCkqZOLw2NUCZ3CVerAu4F hsjLFPP+L13CynGH9aqc/GiSUsdSzyepeFG2O/eiSlC4uqz1uSp9itISlYWzDmFVXGU4 hDVw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784830908; x=1785435708; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=9FpdytHdhxDISqer6403yWjEFqnIJW3e2H1iU7+f5M0=; b=dtNNfvf2xPd4QsC9xv96RAAIvI/SDZYuQvQkeAjgELMlyEJg+xdaogLZDt5zUomYVD HOdimd+7JxbCmY3zIlNwSAWtSxsOxKMpxafVt/vV7mgZe4Bu1CWY+eosxAmvW15yyrQc Bh/CiPZmWxoQtULEzA0fNZgxiLLHe76WH+iaF9zl+iD4jEoiElg8j27mJ/6aav1a2quZ YiN/YfZ1ei+PhGZfOI0LMA7p98x82lVKI4MlZPJNCyWT5+GljzE4G4YQ1+4rhaggliDk 4mM1Wtvl90ncjlCWBQAagveNUNgXpRgLtlTtW1ogZAjRousP9yJdw+rGSfGRmJnEYNZj aTzA== X-Gm-Message-State: AOJu0YzSuSRVpyRZB2vl3/EcZ1prRgpjdns+j8s71X2xovQTU5PTnzYy uNL4/RhkGl4U179acciY5yxSa1Zzj6NBcUPTBB9ZdTrCGX5NBCcS3z9fYewRf4XyMqHbGZ8QiFO UjbKyu/ET5wlUiBTNdt8wOwafaef0BWUNP47biu8HJ3AAzw4pIkTWdW/YvIbbjRYd/rAw4gawzf dS8jvmhXert4W/DrvBl60no+JZlq1Gn1K3+19tTzU5EP/vh/33kfpx6X4= X-Received: from wmbgw6.prod.google.com ([2002:a05:600c:8506:b0:495:4855:c65d]) (user=smostafa job=prod-delivery.src-stubby-dispatcher) by 2002:a05:600c:21c3:b0:493:f6e6:ce2a with SMTP id 5b1f17b1804b1-49573cbe93emr35666835e9.2.1784830908237; Thu, 23 Jul 2026 11:21:48 -0700 (PDT) Date: Thu, 23 Jul 2026 18:21:39 +0000 In-Reply-To: <20260723182140.4025575-1-smostafa@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260723182140.4025575-1-smostafa@google.com> X-Mailer: git-send-email 2.55.0.229.g6434b31f56-goog Message-ID: <20260723182140.4025575-2-smostafa@google.com> Subject: [RFC PATCH v2 1/2] KVM: arm64: Add stage2_clean_old_pte() From: Mostafa Saleh To: linux-kernel@vger.kernel.org, kvmarm@lists.linux.dev, linux-arm-kernel@lists.infradead.org Cc: maz@kernel.org, oupton@kernel.org, seiden@linux.ibm.com, joey.gouly@arm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, vdonnefort@google.com, tabba@google.com, sebastianene@google.com, keirf@google.com, linu.cherian@arm.com, Mostafa Saleh Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" At the moment, the pgtable code rely on BBM in SW which looks like: Break: stage2_try_break_pte() 1) Break PTE and lock it 2) TLBI 3) Put the ref on the old PTE Make: stage2_make_pte() 1) Get a ref on the new PTE 2) Install the new PTE With BBML3, the sequence will look as 1) Get ref on the new PTE 2) Install new PTE 3) TLBI 4) Put the ref on the old PTE Which requires moving step #2 #3 from the break function to the make function, although it is possible to do that for SW BBM also, that means the stage2_try_break_pte() did not fully break the PTE as it is referenced in TLBs, although that works it seems fragile. Instead, move this logic to a new function stage2_clean_old_pte() so that can be called from BBML3. Signed-off-by: Mostafa Saleh --- arch/arm64/kvm/hyp/pgtable.c | 67 ++++++++++++++++++++---------------- 1 file changed, 37 insertions(+), 30 deletions(-) diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c index b74dd5ce1efd..d670da8882a5 100644 --- a/arch/arm64/kvm/hyp/pgtable.c +++ b/arch/arm64/kvm/hyp/pgtable.c @@ -810,39 +810,10 @@ static bool stage2_try_set_pte(const struct kvm_pgtab= le_visit_ctx *ctx, kvm_pte_ return cmpxchg(ctx->ptep, ctx->old, new) =3D=3D ctx->old; } =20 -/** - * stage2_try_break_pte() - Invalidates a pte according to the - * 'break-before-make' requirements of the - * architecture. - * - * @ctx: context of the visited pte. - * @mmu: stage-2 mmu - * - * Returns: true if the pte was successfully broken. - * - * If the removed pte was valid, performs the necessary serialization and = TLB - * invalidation for the old value. For counted ptes, drops the reference c= ount - * on the containing table page. - */ -static bool stage2_try_break_pte(const struct kvm_pgtable_visit_ctx *ctx, +static void stage2_clean_old_pte(const struct kvm_pgtable_visit_ctx *ctx, struct kvm_s2_mmu *mmu) { struct kvm_pgtable_mm_ops *mm_ops =3D ctx->mm_ops; - kvm_pte_t locked_pte; - - if (stage2_pte_is_locked(ctx->old)) { - /* - * Should never occur if this walker has exclusive access to the - * page tables. - */ - WARN_ON(!kvm_pgtable_walk_shared(ctx)); - return false; - } - - locked_pte =3D FIELD_PREP(KVM_INVALID_PTE_TYPE_MASK, - KVM_INVALID_PTE_TYPE_LOCKED); - if (!stage2_try_set_pte(ctx, locked_pte)) - return false; =20 if (!kvm_pgtable_walk_skip_bbm_tlbi(ctx)) { /* @@ -862,6 +833,42 @@ static bool stage2_try_break_pte(const struct kvm_pgta= ble_visit_ctx *ctx, =20 if (stage2_pte_is_counted(ctx->old)) mm_ops->put_page(ctx->ptep); +} + +/** + * stage2_try_break_pte() - Invalidates a pte according to the + * 'break-before-make' requirements of the + * architecture. + * + * @ctx: context of the visited pte. + * @mmu: stage-2 mmu + * + * Returns: true if the pte was successfully broken. + * + * If the removed pte was valid, performs the necessary serialization and = TLB + * invalidation for the old value. For counted ptes, drops the reference c= ount + * on the containing table page. + */ +static bool stage2_try_break_pte(const struct kvm_pgtable_visit_ctx *ctx, + struct kvm_s2_mmu *mmu) +{ + kvm_pte_t locked_pte; + + if (stage2_pte_is_locked(ctx->old)) { + /* + * Should never occur if this walker has exclusive access to the + * page tables. + */ + WARN_ON(!kvm_pgtable_walk_shared(ctx)); + return false; + } + + locked_pte =3D FIELD_PREP(KVM_INVALID_PTE_TYPE_MASK, + KVM_INVALID_PTE_TYPE_LOCKED); + if (!stage2_try_set_pte(ctx, locked_pte)) + return false; + + stage2_clean_old_pte(ctx, mmu); =20 return true; } --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 04:48:26 2026 Received: from mail-wm1-f69.google.com (mail-wm1-f69.google.com [209.85.128.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7372E3B6BF1 for ; Thu, 23 Jul 2026 18:21:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784830918; cv=none; b=hIH8Fs+ntF7/YdOOLPwHtIV6B+tA1n1Hv3Iuf7rcz2/tQzj0D77rEev+uTNMQtgoCF96+XPzxVPZa2e75Yly14AOaVVcAeJYV8ocY+a/yzazJJMYtqmLOYtnWHtDZY4860uWJ3C1cLlZxfbmmbOCJgLqGurG7qAzqn2jHdlg7Y4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784830918; c=relaxed/simple; bh=0tb67XafB5UfHWrQ/N2Zv8/FnZ2jU3xBrAVpocbKjDE=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=uLfx7Qw4sNl5bofIpJ76uyu1irUCSpEa4/yT3ySZAAnlW+W31Zd0dQjHiAOHhFkwWO1LYvEl4iY6YgeAK2ExPGX/Br1TZ/iUVl0fjY9euTbUflywog8L20FpAa41DlNBemRkzohWiM4eHLf+3fJhpaj93woK/tRKYHRhPNx8Yz8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--smostafa.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=nW1Aaqle; arc=none smtp.client-ip=209.85.128.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--smostafa.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="nW1Aaqle" Received: by mail-wm1-f69.google.com with SMTP id 5b1f17b1804b1-4957287363bso6601675e9.0 for ; Thu, 23 Jul 2026 11:21:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784830911; x=1785435711; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=N52o3BVp773wN3Idd00jSA0m0Lpy2g57ImYGMyeejhc=; b=nW1AaqlePuddkWPD2G4pvYqCAvZXjfux6bKNj5CHAcoi/EqyMxK9N4g6e+KtzETT44 xKpXfJkC3D5qDti4rqw3TEujFq2iEFIMAnTDesMKw9fIPVX5L34F++eVcHEbPl3qfhYf zSsrNMheuX/iFpqGyzjNn2asKVTFx/nlhmFKUbUbmzpvnlci1tmeFiSxX+hB2vvcy7M8 tdZMCzTmLt28xSumE49ppcFtiWrZGfgUcI0uR2bAQ2mqcNX3XynzVfnun1467KFLvqbK m4tKOYOA5roratlqXMYKIvV9Dhp9HnK3zvbjMgaAdqUJq+uxcx+MYvvt7ds1rL6nruMC 4iSQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784830911; x=1785435711; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=N52o3BVp773wN3Idd00jSA0m0Lpy2g57ImYGMyeejhc=; b=N13K5CHUJgDVmsCM4OLjxWrgDAhHI82ye7ZPiky7gxlA4qTYAICDBHj9jTksON+uxg 036Jw1uY2Jmhhf5ky5YGus61y3NBcx8qN+M5RuWcrtpgIzdJNdqBv+uDQZkm8zbJSxfs 5VHEbDUPlA85Qj6ZbtyfZ3F/vvx7xgXGo/mtxePILIzI8unwCaVa3vVhcDeYWzy+Bl42 X6pOIyeJ/m73Qz0MD7+9PQxFxqgDIDoGYsY/rsCIWdhC/Ovj9ZcSd9yn5n68O34ob3ur lFPHvH+Q5aaO+46FMkNwjdla7wxS3DLuF+DfI8a2vNqK0K6/CqLQsfwR1kThn9xFZDx7 eKEQ== X-Gm-Message-State: AOJu0YzkICwAHCTJ+l2NZwIbjgZJbHQprl/Ha7QQzDmgstQq/1+GvP1+ ozSLdbynijAB6jKdYm9nYToGSXG3DEYaiAb3JSLOX/kiDxDfJL4E2NWnBXL7uPHfOdaHm8M0px5 9PgH6+CnN/OLVKMVeXAo2UQCNEspkuVLFy7WVrXdqJQySWal8KrcCaAKYTMoqHeAHesIpI4jpjL UiNfbetDuYH2ZqnCiOXqN9A4b4L5dJ2RNr7T/o9+BkzXdTVYIM2tkTqiA= X-Received: from wrhm19.prod.google.com ([2002:a05:6000:1813:b0:472:9520:f359]) (user=smostafa job=prod-delivery.src-stubby-dispatcher) by 2002:a7b:c84c:0:b0:495:3eb2:b763 with SMTP id 5b1f17b1804b1-49573cffa2emr35967485e9.21.1784830911229; Thu, 23 Jul 2026 11:21:51 -0700 (PDT) Date: Thu, 23 Jul 2026 18:21:40 +0000 In-Reply-To: <20260723182140.4025575-1-smostafa@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260723182140.4025575-1-smostafa@google.com> X-Mailer: git-send-email 2.55.0.229.g6434b31f56-goog Message-ID: <20260723182140.4025575-3-smostafa@google.com> Subject: [RFC PATCH v2 2/2] KVM: arm64: Support BBM level 3 From: Mostafa Saleh To: linux-kernel@vger.kernel.org, kvmarm@lists.linux.dev, linux-arm-kernel@lists.infradead.org Cc: maz@kernel.org, oupton@kernel.org, seiden@linux.ibm.com, joey.gouly@arm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, vdonnefort@google.com, tabba@google.com, sebastianene@google.com, keirf@google.com, linu.cherian@arm.com, Mostafa Saleh Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" If the system supports hardware Break-Before-Make (BBM) level 3, use it to replace stage-2 PTEs directly. Otherwise, fall back to the software BBM sequence. For BBML3 the sequence is: 1) Get a reference count on the containing table for the new PTE. 2) Atomically update the PTE with the new valid descriptor. 3) Invalidate the TLB for the old PTE. 4) Drop the reference count holding the old PTE. One interesting case, as BBML3 will update the PTE atomically, it can only know it raced with another core at the point of the cmpxchg failing, unlike the SW implementation which locks the PTE first. And as we must issue CMOs to the new mapped page before the update, that means with BBML3 racing cores will issue redundant CMOs. To avoid this, limit BBML3 support for systems with DIC and FWB, which does not require CMOs. Signed-off-by: Mostafa Saleh --- arch/arm64/kvm/hyp/pgtable.c | 58 +++++++++++++++++++++++++++++++----- 1 file changed, 51 insertions(+), 7 deletions(-) diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c index d670da8882a5..4644b596f020 100644 --- a/arch/arm64/kvm/hyp/pgtable.c +++ b/arch/arm64/kvm/hyp/pgtable.c @@ -835,10 +835,24 @@ static void stage2_clean_old_pte(const struct kvm_pgt= able_visit_ctx *ctx, mm_ops->put_page(ctx->ptep); } =20 +/* + * We assume that KVM will never change the OA of an active translation. + * If the host needs to move the backing PFN, it should do an explicit + * unmap to issue the required TLBI. + */ +static bool stage2_use_bbml3(void) +{ + return system_supports_bbml3() && + cpus_have_final_cap(ARM64_HAS_STAGE2_FWB) && + cpus_have_final_cap(ARM64_HAS_CACHE_DIC); +} + /** * stage2_try_break_pte() - Invalidates a pte according to the * 'break-before-make' requirements of the - * architecture. + * architecture, if BMML3 is supported it + * will be used, meaning that this function + * won't break the PTE. * * @ctx: context of the visited pte. * @mmu: stage-2 mmu @@ -854,6 +868,10 @@ static bool stage2_try_break_pte(const struct kvm_pgta= ble_visit_ctx *ctx, { kvm_pte_t locked_pte; =20 + /* All handled in stage2_make_pte() */ + if (stage2_use_bbml3() && kvm_pte_valid(ctx->old)) + return true; + if (stage2_pte_is_locked(ctx->old)) { /* * Should never occur if this walker has exclusive access to the @@ -873,16 +891,35 @@ static bool stage2_try_break_pte(const struct kvm_pgt= able_visit_ctx *ctx, return true; } =20 -static void stage2_make_pte(const struct kvm_pgtable_visit_ctx *ctx, kvm_p= te_t new) +static bool stage2_make_pte(const struct kvm_pgtable_visit_ctx *ctx, struc= t kvm_s2_mmu *mmu, + kvm_pte_t new) { struct kvm_pgtable_mm_ops *mm_ops =3D ctx->mm_ops; =20 - WARN_ON(!stage2_pte_is_locked(*ctx->ptep)); - if (stage2_pte_is_counted(new)) mm_ops->get_page(ctx->ptep); =20 + if (stage2_use_bbml3() && kvm_pte_valid(ctx->old)) { + /* + * Barrier is required because stage2_try_set_pte() uses + * WRITE_ONCE for non-shared walks, lacking release semantics + * used in the software BBM case. + */ + smp_wmb(); + if (!stage2_try_set_pte(ctx, new)) { + /* Raced with another core. */ + if (stage2_pte_is_counted(new)) + mm_ops->put_page(ctx->ptep); + return false; + } + + stage2_clean_old_pte(ctx, mmu); + return true; + } + + WARN_ON(!stage2_pte_is_locked(*ctx->ptep)); smp_store_release(ctx->ptep, new); + return true; } =20 static bool stage2_unmap_defer_tlb_flush(struct kvm_pgtable *pgt) @@ -1014,7 +1051,8 @@ static int stage2_map_walker_try_leaf(const struct kv= m_pgtable_visit_ctx *ctx, stage2_pte_executable(new)) mm_ops->icache_inval_pou(kvm_pte_follow(new, mm_ops), granule); =20 - stage2_make_pte(ctx, new); + if (!stage2_make_pte(ctx, data->mmu, new)) + return -EAGAIN; =20 return 0; } @@ -1069,7 +1107,10 @@ static int stage2_map_walk_leaf(const struct kvm_pgt= able_visit_ctx *ctx, * will be mapped lazily. */ new =3D kvm_init_table_pte(childp, mm_ops); - stage2_make_pte(ctx, new); + if (!stage2_make_pte(ctx, data->mmu, new)) { + mm_ops->put_page(childp); + return -EAGAIN; + } =20 return 0; } @@ -1560,7 +1601,10 @@ static int stage2_split_walker(const struct kvm_pgta= ble_visit_ctx *ctx, * writes the PTE using smp_store_release(). */ new =3D kvm_init_table_pte(childp, mm_ops); - stage2_make_pte(ctx, new); + if (!stage2_make_pte(ctx, mmu, new)) { + kvm_pgtable_stage2_free_unlinked(mm_ops, childp, level); + return -EAGAIN; + } return 0; } =20 --=20 2.55.0.229.g6434b31f56-goog