From nobody Sat Jul 25 05:22:53 2026 Received: from mail-wr1-f74.google.com (mail-wr1-f74.google.com [209.85.221.74]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6894C3E7BCE for ; Fri, 17 Jul 2026 13:09:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.74 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293749; cv=none; b=XXjtllncmkHcwRWKTKS8mUto1Gd48twALcMYaqMD1m67alxbxupFmYeGtr8ZdZzvh+jhu6fGmiIkUvqFho2ZzPq5Xk4p/ERBH6Zfz+iod2ZEMH/TUooXl4EY7rvd93VuH7bPa4gM1xmMrzF6+KAo8B4ljAce5cVBo6A8RD2fbho= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293749; c=relaxed/simple; bh=8bWbCVTmqjsdXLL+5FJBNFKc14o//Ho5EEc1NgDWZ00=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=Dfkasv1EEGQ29hQFzOAPoEKrfSBsNI1HWFILe3b1fYrwtARbWrdLvMukQKXbCuEp1VA15nXySbz00k9HoHig30x4fVJyPo3ST8RZMwu2L8bze5gxQQboq6zSSjeRdZp8dbCIf+qXboqlNVju8MCLJ2rZLcL5+iZE8LI42aFNQ/o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--smostafa.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=MDb4IhYo; arc=none smtp.client-ip=209.85.221.74 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--smostafa.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="MDb4IhYo" Received: by mail-wr1-f74.google.com with SMTP id ffacd0b85a97d-475e540a0ffso4574252f8f.3 for ; Fri, 17 Jul 2026 06:09:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784293747; x=1784898547; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=VImkqaBVxIups1vL1ulSnRZBEUIJVUK8xJrNhCCDu1A=; b=MDb4IhYoVhzpvm740ORwIPsZ2RySDgar+4qRnNe0JbEiepAYUDLByn0pCOgjLCTC3z kJn4x7vPcDpU1GDbCG362mvx1UI7ZG0sQkGDJLzrxUEXxsf3SBFKNPZGjYa/DZjydyI8 MKgveHL3Ezdz9SApLExt9ffKk+4QnrTu4177DZNJ0KpbKYhnIrdh0FGaZMFP379I0+H8 3qB9qQ4BZ6dxw1I9fdLstepgh5Kz2V6IvdPFNGJIrY16AdoALovrMtHFXwLbvZnHO3eH P1ya0/eeLmsEHKqw/zNcZozlrPYzMWaCJD9KKKNk7u64C+tLFPgTuflzb8ddZFIkUVk8 tCAA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293747; x=1784898547; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=VImkqaBVxIups1vL1ulSnRZBEUIJVUK8xJrNhCCDu1A=; b=apbMgF1thgoCO0EzicGnwBZkpDLawkLvVINcCN6nyGA01SshpMj3iSee92moSur19W uPX07xL3IifCg/qhIs1Bny2UTvjxxK5vmM1SQ3jvMYj2ghWEwAtbL/ZtkT2tfjaySRgL If453bImXOnBK3JNxJqkIGIbuF+9ExOt4lH/oOYQzVfDBqjxEWPnCHporkofabpmh9Bd Z/3c415eoJeEyL0JYzFpO737D+gItciiINe84h0oBBMJK/Nu0BBOEn0L8TF3qVmvLqE+ G6Uzb/g2uQanRygTP1n9WTcSKm41OttLEj27G9L7i2kA2icTnvT2EOftDJ8mx7PqAXN5 Pa/g== X-Gm-Message-State: AOJu0YzZJaS1cvpmQtVJnf9sNgAXLIQ0vEree78dWOhtiggFAe8lh54B W/W2afat0qzSqVf6UiJVtFkJgzy1qVUCnbiW6mbtCJ+Z+ksI7wkAmiAH/Ckk7tHsVhyXaVsV8Eu eTYtlSmjkFCxnOn+LvcqoIPJ8Jtbgvd1uSk/6d1X8QW918hXYIEiVsv4V4NG6ciLEo1Fv7p02Bq Ao+7JnCpFiK9ejH9ftysp0nkqQstYrAG038HPPJ5Pzbyp2WKoVmsT9v7s= X-Received: from wrpk9.prod.google.com ([2002:adf:f5c9:0:b0:460:d09:f6ad]) (user=smostafa job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6000:25f0:b0:476:dd2b:611 with SMTP id ffacd0b85a97d-47f623049e7mr3779169f8f.3.1784293746031; Fri, 17 Jul 2026 06:09:06 -0700 (PDT) Date: Fri, 17 Jul 2026 13:08:59 +0000 In-Reply-To: <20260717130901.2239134-1-smostafa@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260717130901.2239134-1-smostafa@google.com> X-Mailer: git-send-email 2.55.0.229.g6434b31f56-goog Message-ID: <20260717130901.2239134-2-smostafa@google.com> Subject: [RFC PATCH 1/2] KVM: arm64: Add stage2_clean_old_pte() From: Mostafa Saleh To: linux-kernel@vger.kernel.org, kvmarm@lists.linux.dev, linux-arm-kernel@lists.infradead.org Cc: maz@kernel.org, oupton@kernel.org, seiden@linux.ibm.com, joey.gouly@arm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, vdonnefort@google.com, tabba@google.com, Mostafa Saleh Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" At the moment, the pgtable code rely on BBM in SW which looks like: Break: stage2_try_break_pte() 1) Break PTE and lock it 2) TLBI 3) Put the ref on the old PTE Make: stage2_make_pte() 1) Get a ref on the new PTE 2) Install the live PTE With BBML3, the sequence will look as 1) Get ref on the new PTE 2) Install new PTE 3) TLBI 4) Put the ref on the old PTE Which requires moving step #2 #3 from the break function to the make function, although it is possible to do that for SW BBM also, that means the stage2_try_break_pte() did not fully break the PTE as it is referenced in TLBs, although that works it seems fragile. Instead, move this logic to a new function stage2_clean_old_pte() that can be called from BBML3. Signed-off-by: Mostafa Saleh --- arch/arm64/kvm/hyp/pgtable.c | 67 ++++++++++++++++++++---------------- 1 file changed, 37 insertions(+), 30 deletions(-) diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c index 91a7dfad6686..127b7f9541b1 100644 --- a/arch/arm64/kvm/hyp/pgtable.c +++ b/arch/arm64/kvm/hyp/pgtable.c @@ -810,39 +810,10 @@ static bool stage2_try_set_pte(const struct kvm_pgtab= le_visit_ctx *ctx, kvm_pte_ return cmpxchg(ctx->ptep, ctx->old, new) =3D=3D ctx->old; } =20 -/** - * stage2_try_break_pte() - Invalidates a pte according to the - * 'break-before-make' requirements of the - * architecture. - * - * @ctx: context of the visited pte. - * @mmu: stage-2 mmu - * - * Returns: true if the pte was successfully broken. - * - * If the removed pte was valid, performs the necessary serialization and = TLB - * invalidation for the old value. For counted ptes, drops the reference c= ount - * on the containing table page. - */ -static bool stage2_try_break_pte(const struct kvm_pgtable_visit_ctx *ctx, +static void stage2_clean_old_pte(const struct kvm_pgtable_visit_ctx *ctx, struct kvm_s2_mmu *mmu) { struct kvm_pgtable_mm_ops *mm_ops =3D ctx->mm_ops; - kvm_pte_t locked_pte; - - if (stage2_pte_is_locked(ctx->old)) { - /* - * Should never occur if this walker has exclusive access to the - * page tables. - */ - WARN_ON(!kvm_pgtable_walk_shared(ctx)); - return false; - } - - locked_pte =3D FIELD_PREP(KVM_INVALID_PTE_TYPE_MASK, - KVM_INVALID_PTE_TYPE_LOCKED); - if (!stage2_try_set_pte(ctx, locked_pte)) - return false; =20 if (!kvm_pgtable_walk_skip_bbm_tlbi(ctx)) { /* @@ -862,6 +833,42 @@ static bool stage2_try_break_pte(const struct kvm_pgta= ble_visit_ctx *ctx, =20 if (stage2_pte_is_counted(ctx->old)) mm_ops->put_page(ctx->ptep); +} + +/** + * stage2_try_break_pte() - Invalidates a pte according to the + * 'break-before-make' requirements of the + * architecture. + * + * @ctx: context of the visited pte. + * @mmu: stage-2 mmu + * + * Returns: true if the pte was successfully broken. + * + * If the removed pte was valid, performs the necessary serialization and = TLB + * invalidation for the old value. For counted ptes, drops the reference c= ount + * on the containing table page. + */ +static bool stage2_try_break_pte(const struct kvm_pgtable_visit_ctx *ctx, + struct kvm_s2_mmu *mmu) +{ + kvm_pte_t locked_pte; + + if (stage2_pte_is_locked(ctx->old)) { + /* + * Should never occur if this walker has exclusive access to the + * page tables. + */ + WARN_ON(!kvm_pgtable_walk_shared(ctx)); + return false; + } + + locked_pte =3D FIELD_PREP(KVM_INVALID_PTE_TYPE_MASK, + KVM_INVALID_PTE_TYPE_LOCKED); + if (!stage2_try_set_pte(ctx, locked_pte)) + return false; + + stage2_clean_old_pte(ctx, mmu); =20 return true; } --=20 2.55.0.229.g6434b31f56-goog From nobody Sat Jul 25 05:22:53 2026 Received: from mail-wm1-f73.google.com (mail-wm1-f73.google.com [209.85.128.73]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 053AD3FC5D3 for ; Fri, 17 Jul 2026 13:09:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.73 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293751; cv=none; b=kMrQT1CFtwQ3pS5ujueDiJcLul0xtbHkt8CnToaEF72tJTaGywxJAFAMyS7i06YOXckM7QnnS1+fpeNl7IjRvqe5REuwD6n+jR5jS/6zZNFHQu/O+aF+TMJxosReDnVJ6XWmP575UeO5GmqDL2qoXqI0rbaBm0smUJ/zp1wiRMU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784293751; c=relaxed/simple; bh=Mjz8ivQNJtbkbZVgumLby1xtvMOe4isndBrW/gdbi9E=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=TlEQtARkjR+nHy61S0818HgMAPg7vgI9RIT4xr3q2y6AprCYI0WsbnbapDZa7o9h9d+Ztq9QZ9Z9u2vhw4K5oCyIUA9oftGHNoSWl8y1BK25yZDPqfV16I7lVbs5ITJ5VZLX61lOkOzcEqHmIi3ltNiiSccfzu3GF0Z5Igh6q1E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--smostafa.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=QC1Qakon; arc=none smtp.client-ip=209.85.128.73 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--smostafa.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="QC1Qakon" Received: by mail-wm1-f73.google.com with SMTP id 5b1f17b1804b1-495474a5fbcso12389165e9.1 for ; Fri, 17 Jul 2026 06:09:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1784293748; x=1784898548; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=+n0quTBKTLkuwq9QawKwspQRW3gBb6Rua2VTww6C7Lw=; b=QC1QakonfQ1kiZ/+LDAMCdrwjAhLDqNFsU7cQWvn0N2uVi77ugRRqFgHv1u6ZpSaSO hyNX92fKf2R1mvCgBbpUFs8dW7DNM5F3ghIo20Rmbtw3QGgF7UlFcgdaKICilIv+vDgM EbpPg1U5y+EsPf0FVecIEcDyn3xhLPOziaiQKaA8SM00/b/6k1pTuZCpArz2SA/S7kK+ idcRHSq7pypR9G/MgvjErj8h2Z3iU+VuKS7vcKx9fYLoEdR55Su13PCsL5YPzVoQDnCq QiPf2zSR0i9khVAKjboTXlDsC2rdrezTAPHPvNBo9dxnBMYAEEOqdUqBBHP8o7rtP6VV jMzQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784293748; x=1784898548; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=+n0quTBKTLkuwq9QawKwspQRW3gBb6Rua2VTww6C7Lw=; b=HVow2dI0Nnc3CWcPXJ3ujY7PZ7uQCpV5ivSQ+BAF0xQOqNDpfeJmk5GWqFGaGbH93H fvpN7gHu+l7o9ebmPRftQDylqH3ztq11Bk7/3ObThbTefhonQyo79WW3QMs7VYV9fQjH bH4C73TzEd7fa5esLqARJmdIWQdYkllQGtMjDc+pleN+BUYsym56wWhASX+zrtk3MlsZ Df2sSEs5tNaDcANeQAwwPkGEzJcqOJvoepEGrtotJjnarkKYch5dKcWz2K6VhtDV1wSZ +tDRm0nDLFAR83YIFtMLCh9MhiZBa0IVqyESRLapH1OuJRH31LDg9Qe4GC2F655GE+fm +q1g== X-Gm-Message-State: AOJu0YymUFV1aZCoXKtvoHBvydOLG8QznPcx2E8vHR9OJSlMhQzOOBcU Mj4ikBofEcseofDq8VLkwcfCcKFBvb7yIM0Xn4SYne5kl/tPn+mD224yhDz31v4eKgn3E2HFZws 5rJoBShvw7/ZjnfY2GnL0X/AcczDQ+Srv8FUJF9TaQVzhtTNTb2QxJzjrAeQmrqIc7oXr9OTluI +Ete994Q+eUBISvUqD5oQN/Xg0RtoC0HF8EsgSEblneTFC/CzNCE+Xtg8= X-Received: from wrbcg5.prod.google.com ([2002:a5d:5cc5:0:b0:47f:4f99:b306]) (user=smostafa job=prod-delivery.src-stubby-dispatcher) by 2002:a05:600c:154e:b0:495:49eb:1ccc with SMTP id 5b1f17b1804b1-4954a3eff56mr32478345e9.13.1784293747452; Fri, 17 Jul 2026 06:09:07 -0700 (PDT) Date: Fri, 17 Jul 2026 13:09:00 +0000 In-Reply-To: <20260717130901.2239134-1-smostafa@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260717130901.2239134-1-smostafa@google.com> X-Mailer: git-send-email 2.55.0.229.g6434b31f56-goog Message-ID: <20260717130901.2239134-3-smostafa@google.com> Subject: [RFC PATCH 2/2] KVM: arm64: Support BBM level 3 From: Mostafa Saleh To: linux-kernel@vger.kernel.org, kvmarm@lists.linux.dev, linux-arm-kernel@lists.infradead.org Cc: maz@kernel.org, oupton@kernel.org, seiden@linux.ibm.com, joey.gouly@arm.com, suzuki.poulose@arm.com, yuzenghui@huawei.com, catalin.marinas@arm.com, will@kernel.org, vdonnefort@google.com, tabba@google.com, Mostafa Saleh Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" If the system supports hardware Break-Before-Make (BBM) level 3, use it to replace stage-2 PTEs directly instead of falling back to the software break-before-make sequence. 1) Get a reference count on the containing table for the new PTE. 2) Atomically update the PTE with the new valid descriptor. 3) Invalidate the TLB for the old PTE. 4) Drop the reference count holding the old PTE. One interesting case, as BBML3 will update the PTE atomically, it can only know it raced with another core at the point of the cmpxchg failing, unlike the SW implementation which locks the PTE first. And as we must issue CMOs to the new mapped page before the update, that means with BBML3 racing cores will issue redundant CMOs, to improve this: - We only use BBML3 if the old PTE was live. - To reduce the window of the race, an early check is added before the CMO to exit early, but that does not eliminate the race. Signed-off-by: Mostafa Saleh --- arch/arm64/kvm/hyp/pgtable.c | 53 +++++++++++++++++++++++++++++++----- 1 file changed, 46 insertions(+), 7 deletions(-) diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c index 127b7f9541b1..69d52308236f 100644 --- a/arch/arm64/kvm/hyp/pgtable.c +++ b/arch/arm64/kvm/hyp/pgtable.c @@ -838,7 +838,8 @@ static void stage2_clean_old_pte(const struct kvm_pgtab= le_visit_ctx *ctx, /** * stage2_try_break_pte() - Invalidates a pte according to the * 'break-before-make' requirements of the - * architecture. + * architecture, if BMML3 is supported it + * will be used, otherwise fallback to SW. * * @ctx: context of the visited pte. * @mmu: stage-2 mmu @@ -854,6 +855,18 @@ static bool stage2_try_break_pte(const struct kvm_pgta= ble_visit_ctx *ctx, { kvm_pte_t locked_pte; =20 + if (system_supports_bbml3() && kvm_pte_valid(ctx->old)) { + kvm_pte_t curr_pte =3D READ_ONCE(*ctx->ptep); + + /* + * All handled in stage2_make_pte(). However exit early if we already + * lost the race to avoid extra CMOs. + */ + if (curr_pte !=3D ctx->old) + return false; + return true; + } + if (stage2_pte_is_locked(ctx->old)) { /* * Should never occur if this walker has exclusive access to the @@ -873,16 +886,35 @@ static bool stage2_try_break_pte(const struct kvm_pgt= able_visit_ctx *ctx, return true; } =20 -static void stage2_make_pte(const struct kvm_pgtable_visit_ctx *ctx, kvm_p= te_t new) +/* Must be paired with stage2_try_break_pte() */ +static bool stage2_make_pte(const struct kvm_pgtable_visit_ctx *ctx, struc= t kvm_s2_mmu *mmu, + kvm_pte_t new) { struct kvm_pgtable_mm_ops *mm_ops =3D ctx->mm_ops; =20 - WARN_ON(!stage2_pte_is_locked(*ctx->ptep)); - if (stage2_pte_is_counted(new)) mm_ops->get_page(ctx->ptep); =20 + if (system_supports_bbml3() && kvm_pte_valid(ctx->old)) { + /* + * Barrier is required because stage2_try_set_pte() uses + * WRITE_ONCE for non-shared walks, lacking release semantics + * used in the software BBM case. + */ + smp_wmb(); + if (!stage2_try_set_pte(ctx, new)) { + if (stage2_pte_is_counted(new)) + mm_ops->put_page(ctx->ptep); + return false; + } + + stage2_clean_old_pte(ctx, mmu); + return true; + } + + WARN_ON(!stage2_pte_is_locked(*ctx->ptep)); smp_store_release(ctx->ptep, new); + return true; } =20 static bool stage2_unmap_defer_tlb_flush(struct kvm_pgtable *pgt) @@ -1014,7 +1046,8 @@ static int stage2_map_walker_try_leaf(const struct kv= m_pgtable_visit_ctx *ctx, stage2_pte_executable(new)) mm_ops->icache_inval_pou(kvm_pte_follow(new, mm_ops), granule); =20 - stage2_make_pte(ctx, new); + if (!stage2_make_pte(ctx, data->mmu, new)) + return -EAGAIN; =20 return 0; } @@ -1069,7 +1102,10 @@ static int stage2_map_walk_leaf(const struct kvm_pgt= able_visit_ctx *ctx, * will be mapped lazily. */ new =3D kvm_init_table_pte(childp, mm_ops); - stage2_make_pte(ctx, new); + if (!stage2_make_pte(ctx, data->mmu, new)) { + mm_ops->put_page(childp); + return -EAGAIN; + } =20 return 0; } @@ -1557,7 +1593,10 @@ static int stage2_split_walker(const struct kvm_pgta= ble_visit_ctx *ctx, * writes the PTE using smp_store_release(). */ new =3D kvm_init_table_pte(childp, mm_ops); - stage2_make_pte(ctx, new); + if (!stage2_make_pte(ctx, mmu, new)) { + kvm_pgtable_stage2_free_unlinked(mm_ops, childp, level); + return -EAGAIN; + } return 0; } =20 --=20 2.55.0.229.g6434b31f56-goog