From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8E5C52609E3 for ; Mon, 21 Sep 2026 00:48:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951722; cv=none; b=jmaKo9a7reoopMVXkW1zso/I5rBHFuyFLA+sABKh1y11FREDUF0fFwMqB/o0+xkA74g0i/Y04hPSansVLyC1NQ/q2GgvzVY5c6WjebXD2AMfpPflyGP/7hdZCtYPnhUfEhC8HXebIbaguBlQLDU1m46RnZz+hJkXcDBzKAQa/m4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951722; c=relaxed/simple; bh=smolpSPLyRzt3C+D1JHi3AqucD1pUoZl1tQGGaZ45JA=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=XUITyT1dITCJwGnhfDGpsrwb1+FlYY3GvufCvyQI+Lvogtmf12kd5Deo9yfwMKJgu5+cn30GlTauqlA3vkkycWPahi5JVSTkIZZMzXPnSa1oMvCwGp3fJabPmWeafAwf0SSh+0heijYsR0sn+k8fsuz3RAcg1spLN1IKAWw9s34= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=SDAMdGij; arc=none smtp.client-ip=209.85.214.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="SDAMdGij" Received: by mail-pl1-f198.google.com with SMTP id d9443c01a7336-2d55d8cd938so44523035ad.1 for ; Sun, 20 Sep 2026 17:48:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951720; x=1790556520; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=5rIYSaUxr2dmbQ4fWz0MikFLRE1YKdMT0WLnLm0z0eI=; b=SDAMdGijtguYAtnSMOHATZqLSgU6Pc72BPWLuWQVRa0bW1T5dFdDo2iB4ZikDOFm/O 6O1aSdIZjjM4VMNwn2aCNJhHMwQ5ZoJ3eOJfjuXrhvrtgw1FQ3uSUfLF7Tw8oE14T/3x JzYx81cxXbHqd27Q+Lwc4rHrPipxAx2LuRd77RvVbcdJU89MS33s8aGYB2+j8vTVMeK5 H73js2scbKmCoybbipl1ffOjejiMUKYYNx91mS0fI/S5/jxzQB9rOmFjI3fupwaDAWBw id5RSE4Ml0niI0KWLWGzN7pY28g3gnrmKnatG7xbLC7aDw3d5MuOg2kjWI3mJPCwYV36 JzDw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951720; x=1790556520; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=5rIYSaUxr2dmbQ4fWz0MikFLRE1YKdMT0WLnLm0z0eI=; b=lWUMZOFvu780D5kk5pWLveA7cNjAzlVyhYDt1uWKlnoDeER7P4nnTRojV6m5/0aGGO e4jbruQm3US/aVuvO/uomyqEU4v463m6+kd3IraA5P0+ORb2lignUeQExVlqFh9h7Cvh 2I4t5G6fajAOZH4pTEuU9oUeEgdaR00n6rD3J83vLWZZxjmkRB5hQNCieRhbtmE7t2JE g5SP3g55hehf4xIaNHXWze1BjgQzI+ybB+qdJPPMr+fQw2z5NBTmCdkYa7UEdVSgy2oZ ifGuBZrD6mZPcA7hkvrzhHngxm1vdanuiIU1Kgb3O+Lg2gGJm/S1N19SrO1AOSFWAg5S NLKw== X-Forwarded-Encrypted: i=1; AKwUvBwZFLmsbeUoJ7xt/60imFfJix3/y2esfZrlsdCmAUvPVzGwwK3MBLGIo94+G3I9vP9e1Fp1W7jqPboI0Ew=@vger.kernel.org X-Gm-Message-State: AFuF++lLkShHMG3KoM8cCkjw5TsTbDXOxvlzvsUP+E0XLjamWSLIr4bn 5ubGcE4GADQnzh+P2PHUrQDHWK6aeGbk2B+OzlibIuZSIIslrfPYAAddUuWs6EIf+A3f0kVdn+u RU6/v391H3jGTgg== X-Received: from plld8.prod.google.com ([2002:a17:902:7288:b0:2dd:1e7c:bcec]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:3c6b:b0:2dd:c100:9433 with SMTP id d9443c01a7336-2ddc1009c3cmr88814875ad.49.1789951719834; Sun, 20 Sep 2026 17:48:39 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:17 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-2-skhawaja@google.com> Subject: [PATCH v5 01/18] memfd: export memfd_get_seals() From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Hugh Dickins , Baolin Wang , Andrew Morton , linux-mm@kvack.org, "Pratyush Yadav (Google)" , Ankit Soni , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pasha Tatashin , David Matlack , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" memfd_get_seals() is used by iommufd during preservation to make sure that he preserved memfd was sealed when it was mapped into iommufd. Since iommufd can be built as a module, export memfd_get_seals() to avoid linker error. Cc: Hugh Dickins Cc: Baolin Wang Cc: Andrew Morton Cc: linux-mm@kvack.org Acked-by: Pratyush Yadav (Google) Reviewed-by: Ankit Soni Signed-off-by: Samiullah Khawaja --- mm/memfd.c | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/mm/memfd.c b/mm/memfd.c index c708d92533f4..888ccf3778c7 100644 --- a/mm/memfd.c +++ b/mm/memfd.c @@ -6,6 +6,7 @@ * use by hugetlbfs as well as tmpfs. */ =20 +#include #include #include #include @@ -310,12 +311,19 @@ int memfd_add_seals(struct file *file, unsigned int s= eals) return error; } =20 +/** + * memfd_get_seals - Gets current seals on a memfd + * @file: struct file representing the memfd + * + * Returns seals if file is a memfd, otherwise returns -EINVAL + */ int memfd_get_seals(struct file *file) { unsigned int *seals =3D memfd_file_seals_ptr(file); =20 return seals ? *seals : -EINVAL; } +EXPORT_SYMBOL_GPL(memfd_get_seals); =20 long memfd_fcntl(struct file *file, unsigned int cmd, unsigned int arg) { --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f197.google.com (mail-pl1-f197.google.com [209.85.214.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BC07E2BE053 for ; Mon, 21 Sep 2026 00:48:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.197 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951725; cv=none; b=k3FAbTogW1trQz/qgZnfelJX7Ydqk7LwTp2g0jyV0hDyK2kIuKyfmMC+REIIe1DQrr4uKDBbt0m5nVYdbb7SSRiQjd5hzTIU9rb9YHBz6/dY//035MCrlqOJFLw9rCc2NBPh6RK+gAUk4mnzRYJh/1P1cjL510Pl6Y6y8qsO5Mw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951725; c=relaxed/simple; bh=JrBVxQgkB/CBntA4xh0rGxbH3KEFoXKIAsymS8b+Xqw=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=UzySzN0BMmWZVJZDjgofbV1blm2/AiL7CC31Xeu7Cmc+CDwWz9q5SqjH6VQSwiMOs/Eww4xw1YfnBzn4QgvCQawqwiqfVO4/5BIw2pilEPdC0AA9LIDHbE2Y+RV4WtuCIm2ExMjsJ7bk2cycjwOyamvdNvvq8WLgaxb4AFYBEuQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=DSWiDBnm; arc=none smtp.client-ip=209.85.214.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="DSWiDBnm" Received: by mail-pl1-f197.google.com with SMTP id d9443c01a7336-2d55d8cd938so44523615ad.1 for ; Sun, 20 Sep 2026 17:48:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951722; x=1790556522; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=p3MOa0K8GJJXFbzSXjxcRG0git6pUSRUszMCcW9ffPE=; b=DSWiDBnmp00MGK/9u974L69HAcJ+w7noCN5vJ1K6R2iLGP0CrgXwO7kO42z5/STxMc Lo1TYjIen8iOfFO+w9+hQDotncQ8CQBKqqDf6M/kICnTv/tBsMfcClKSBB1Y5McoMR7+ 8AOP6ZxaSPyS2SaB6iiee49iUgvQ9skITYxVlDB7dw5O24E80P3gJOZgAsKfOjqCmfdO JNvgxzLTVjWclL0DBrhB4DzhWRT8cQJgYM74xf3tIRqYzXHVXdHG0XWHLNKOlqqgW0V9 tPQPxGlwcJPbp0UAf/ZkKb4CgzoVMQeJ2/91RPyaulLzIuOGwuxGvFsnLJGZbe4TZiQA 5faA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951722; x=1790556522; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=p3MOa0K8GJJXFbzSXjxcRG0git6pUSRUszMCcW9ffPE=; b=peovqTDPNxDIhIw4IKYmah51RsnThAOO56U91bZLQ7/M6MD9c9Eio+QEibw70ZClob uJpHY9PwIDgJp46S73y8vEw4+pCkwtDty3TDn2aCbKbAqPt73K8es7vbno+BH+VvFBMj VqKoROO7Ar8W/amg64lh/qGJfY6r26Gxu9cbHJ+ivEpJ3lOwNVVu0gzIuqjiws4e/8f8 cSkk8JlHELEvieLHwbl0HgkpdVV9nuZadWcN0er3KukEuDKrt26x/mWNi3UQxNNsIQ75 v9Q3e9cBYyBtPazOcNhcV5z9GoMZ5TSgUKf0zU994PDe8Mh5Jsmnx+Rzrm/BARQl9QVl tl7Q== X-Forwarded-Encrypted: i=1; AKwUvBy8BsrI6CZ9Lo06UYdMgWznT3r2r/4hU5zY+DIbXh4H5mWuM+Rky8frFHTvIsUiC1zzMiKeEbfPDLDC75A=@vger.kernel.org X-Gm-Message-State: AFuF++mR4s6WAcQI5J95QWDFfHUW58BSFUFyeraeIx2xB5z4JyxMWncu 2HMe+i3CLX5CrrJe38iyD3NlLyHhAkg1Jj3GcGPGDz7qp9axQoFdwzprBTQqQ7txUm8GMpFW+aM kmWA0XWPxrTfyAQ== X-Received: from plcr5.prod.google.com ([2002:a17:903:145:b0:2df:4e6f:dbe5]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:ffce:b0:2dd:ad74:ac7d with SMTP id d9443c01a7336-2ddb1ba3a91mr151523875ad.24.1789951721813; Sun, 20 Sep 2026 17:48:41 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:18 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-3-skhawaja@google.com> Subject: [PATCH v5 02/18] iommu: Implement IOMMU Live update FLB callbacks From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add liveupdate FLB for IOMMU state preservation. Use KHO preserve memory alloc/free helper functions to allocate memory for the IOMMU Live update FLB object and the serialization structs for device, domain and iommu. During retrieve, walk through the preserved obj array headers and restore each folio. Also recreate the FLB obj. Signed-off-by: Samiullah Khawaja --- MAINTAINERS | 9 ++ drivers/iommu/Kconfig | 12 ++ drivers/iommu/Makefile | 1 + drivers/iommu/liveupdate.c | 253 +++++++++++++++++++++++++++++++ include/linux/iommu-liveupdate.h | 28 ++++ include/linux/kho/abi/iommu.h | 212 ++++++++++++++++++++++++++ 6 files changed, 515 insertions(+) create mode 100644 drivers/iommu/liveupdate.c create mode 100644 include/linux/iommu-liveupdate.h create mode 100644 include/linux/kho/abi/iommu.h diff --git a/MAINTAINERS b/MAINTAINERS index 2b62a0b8a48c..2114b50412ee 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -13715,6 +13715,15 @@ F: include/linux/iova.h F: include/linux/of_iommu.h F: rust/kernel/iommu/ =20 +IOMMU LIVEUPDATE +M: Samiullah Khawaja +R: Pranjal Shrivastava +L: iommu@lists.linux.dev +S: Maintained +F: drivers/iommu/liveupdate.c +F: include/linux/iommu-liveupdate.h +F: include/linux/kho/abi/iommu.h + IOMMUFD M: Jason Gunthorpe M: Kevin Tian diff --git a/drivers/iommu/Kconfig b/drivers/iommu/Kconfig index 6e07bd69467a..56a96d22cc8b 100644 --- a/drivers/iommu/Kconfig +++ b/drivers/iommu/Kconfig @@ -405,6 +405,18 @@ config IOMMU_DEBUG_PAGEALLOC line to activate the runtime checks. =20 If unsure, say N. + +config IOMMU_LIVEUPDATE + bool "IOMMU live update state preservation support" + depends on LIVEUPDATE && IOMMUFD && 64BIT + help + Enable support for preserving IOMMU state across a kexec live update. + + This allows devices managed by iommufd to maintain their DMA mappings + during kexec base kernel update. + + If unsure, say N. + endif # IOMMU_SUPPORT =20 source "drivers/iommu/generic_pt/Kconfig" diff --git a/drivers/iommu/Makefile b/drivers/iommu/Makefile index 2f05725eaab1..a775fb526e91 100644 --- a/drivers/iommu/Makefile +++ b/drivers/iommu/Makefile @@ -17,6 +17,7 @@ obj-$(CONFIG_IOMMU_IO_PGTABLE_LPAE) +=3D io-pgtable-arm.o obj-$(CONFIG_IOMMU_IO_PGTABLE_LPAE_KUNIT_TEST) +=3D io-pgtable-arm-selftes= ts.o obj-$(CONFIG_IOMMU_IO_PGTABLE_DART) +=3D io-pgtable-dart.o obj-$(CONFIG_IOMMU_IOVA) +=3D iova.o +obj-$(CONFIG_IOMMU_LIVEUPDATE) +=3D liveupdate.o obj-$(CONFIG_OF_IOMMU) +=3D of_iommu.o obj-$(CONFIG_MSM_IOMMU) +=3D msm_iommu.o obj-$(CONFIG_IPMMU_VMSA) +=3D ipmmu-vmsa.o diff --git a/drivers/iommu/liveupdate.c b/drivers/iommu/liveupdate.c new file mode 100644 index 000000000000..b644f4792532 --- /dev/null +++ b/drivers/iommu/liveupdate.c @@ -0,0 +1,253 @@ +// SPDX-License-Identifier: GPL-2.0-only + +/* + * Copyright (C) 2026, Google LLC + * Author: Samiullah Khawaja + */ + +/** + * DOC: IOMMU Live Update + * + * The IOMMU subsystem participates in the Live Update process to preserve= its + * IOMMU domains and IOMMU specific state of preserved devices. + * + * File-Lifecycle-Bound (FLB) Data + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D + * + * Preserved state of IOMMU HW needs to be restored during boot when the I= OMMU is + * initialized and registered. The preserved IOMMU domains also need to be + * restored and associated with the preserved devices early during boot. F= or + * this reason, IOMMU subsystem uses LUO File-Lifecycle-Bound Data to stor= e some + * state globally. + * + * During preservation of IOMMUFD into LUO, some of the state is stored in= the + * IOMMU FLB so it can be restored during boot. Once the FDs are retrieved= from LUO + * the restored state can be reassociated with the relevant IOMMUFDs. + * + * The FLB also contains the state of IOMMU HW that might be shared between + * multiple devices, domains and iommufds. This is restored early during b= oot + * and no re-association of this state is needed later. + * + * As the state is preserved using KHO, the FLB handler callbacks are call= ed by + * LUO only when the KHO is enabled. So there is no need to check whether = KHO is + * enabled in the callback implementation. + * + */ + +#define pr_fmt(fmt) "iommu: liveupdate: " fmt + +#include +#include +#include +#include +#include + +struct iommu_flb_obj { + struct mutex lock; + struct iommu_flb_ser *ser; + + struct iommu_hw_array_ser *curr_iommu_array; + struct iommu_domain_array_ser *curr_domain_array; + struct iommu_device_array_ser *curr_device_array; +}; + +static void *iommu_liveupdate_restore_array(u64 array_phys) +{ + struct iommu_array_hdr_ser *array_hdr; + void *vaddr =3D array_phys ? phys_to_virt(array_phys) : NULL; + + while (array_phys) { + /* + * Failure to restore preserved IOMMU state is considered fatal. + * + * This is because the IOMMU translations for preserved IOMMUs + * were kept enabled in the previous kernel and the preserved + * devices have their IOMMU domains still present. Not being + * able to restore means that the memory mapped into preserved + * domains might be already corrupted by the preserved devices. + * + * There is no way to confirm the integrity of the memory that + * was mapped. BUG_ON is the safest option at this point. + */ + BUG_ON(!kho_restore_folio(array_phys)); + array_hdr =3D phys_to_virt(array_phys); + array_phys =3D array_hdr->next_array_phys; + } + + return vaddr; +} + +static void iommu_liveupdate_unpreserve_free(u64 array_phys) +{ + struct iommu_array_hdr_ser *array_hdr; + + while (array_phys) { + array_hdr =3D phys_to_virt(array_phys); + array_phys =3D array_hdr->next_array_phys; + kho_unpreserve_free(array_hdr); + } +} + +static void iommu_liveupdate_folio_put(u64 array_phys) +{ + struct iommu_array_hdr_ser *array_hdr; + + while (array_phys) { + array_hdr =3D phys_to_virt(array_phys); + array_phys =3D array_hdr->next_array_phys; + folio_put(virt_to_folio(array_hdr)); + } +} + +static void iommu_liveupdate_flb_free(struct iommu_flb_obj *obj) +{ + iommu_liveupdate_unpreserve_free(obj->ser->iommu_domain_array_phys); + iommu_liveupdate_unpreserve_free(obj->ser->device_array_phys); + iommu_liveupdate_unpreserve_free(obj->ser->iommu_array_phys); + + kho_unpreserve_free(obj->ser); + mutex_destroy(&obj->lock); + kfree(obj); +} + +static int iommu_liveupdate_flb_preserve(struct liveupdate_flb_op_args *ar= gp) +{ + struct iommu_flb_obj *obj; + struct iommu_flb_ser *ser; + void *mem; + + /* obj exists only in the current kernel to track preserved state */ + obj =3D kzalloc_obj(*obj, GFP_KERNEL); + if (!obj) + return -ENOMEM; + + mutex_init(&obj->lock); + + /* mem is allocated via KHO and will survive the kexec */ + mem =3D kho_alloc_preserve(sizeof(*ser)); + if (IS_ERR(mem)) + goto err_free_obj; + + ser =3D mem; + obj->ser =3D ser; + ser->version =3D IOMMU_LUO_FLB_VERSION; + + mem =3D kho_alloc_preserve(PAGE_SIZE); + if (IS_ERR(mem)) + goto err_free_ser; + + obj->curr_domain_array =3D mem; + ser->iommu_domain_array_phys =3D virt_to_phys(obj->curr_domain_array); + + mem =3D kho_alloc_preserve(PAGE_SIZE); + if (IS_ERR(mem)) + goto err_free_domains; + + obj->curr_device_array =3D mem; + ser->device_array_phys =3D virt_to_phys(obj->curr_device_array); + + mem =3D kho_alloc_preserve(PAGE_SIZE); + if (IS_ERR(mem)) + goto err_free_devices; + + obj->curr_iommu_array =3D mem; + ser->iommu_array_phys =3D virt_to_phys(obj->curr_iommu_array); + + argp->obj =3D obj; + argp->data =3D virt_to_phys(ser); + return 0; + +err_free_devices: + kho_unpreserve_free(obj->curr_device_array); +err_free_domains: + kho_unpreserve_free(obj->curr_domain_array); +err_free_ser: + kho_unpreserve_free(obj->ser); +err_free_obj: + mutex_destroy(&obj->lock); + kfree(obj); + return PTR_ERR(mem); +} + +static void iommu_liveupdate_flb_unpreserve(struct liveupdate_flb_op_args = *argp) +{ + iommu_liveupdate_flb_free(argp->obj); +} + +static void iommu_liveupdate_flb_finish(struct liveupdate_flb_op_args *arg= p) +{ + struct iommu_flb_obj *obj =3D argp->obj; + + iommu_liveupdate_folio_put(obj->ser->iommu_domain_array_phys); + iommu_liveupdate_folio_put(obj->ser->device_array_phys); + iommu_liveupdate_folio_put(obj->ser->iommu_array_phys); + + folio_put(virt_to_folio(obj->ser)); + mutex_destroy(&obj->lock); + kfree(obj); +} + +static int iommu_liveupdate_flb_retrieve(struct liveupdate_flb_op_args *ar= gp) +{ + struct iommu_flb_obj *obj; + struct iommu_flb_ser *ser; + + obj =3D kzalloc_obj(*obj, GFP_KERNEL); + if (!obj) { + /* + * If retrieve fails, the finish path won't be called as + * can_finish() will fail, preventing the restore. + */ + return -ENOMEM; + } + + /* Data must be present and valid from the previous kernel */ + BUG_ON(!kho_restore_folio(argp->data)); + + mutex_init(&obj->lock); + ser =3D phys_to_virt(argp->data); + obj->ser =3D ser; + + obj->curr_domain_array =3D iommu_liveupdate_restore_array(ser->iommu_doma= in_array_phys); + obj->curr_device_array =3D iommu_liveupdate_restore_array(ser->device_arr= ay_phys); + obj->curr_iommu_array =3D iommu_liveupdate_restore_array(ser->iommu_array= _phys); + argp->obj =3D obj; + return 0; +} + +static struct liveupdate_flb_ops iommu_flb_ops =3D { + .preserve =3D iommu_liveupdate_flb_preserve, + .unpreserve =3D iommu_liveupdate_flb_unpreserve, + .retrieve =3D iommu_liveupdate_flb_retrieve, + .finish =3D iommu_liveupdate_flb_finish, +}; + +static struct liveupdate_flb iommu_flb =3D { + .compatible =3D IOMMU_LUO_FLB_COMPATIBLE, + .ops =3D &iommu_flb_ops, +}; + +/** + * iommu_liveupdate_register_flb() - Register IOMMU Live Update FLB + * @handler: Live Update file handler to associate FLB with + * + * Register a Live Update FLB for IOMMU state preservation with the file + * handler. + * + * Return: 0 on success, or negative error code. + */ +int iommu_liveupdate_register_flb(struct liveupdate_file_handler *handler) +{ + return liveupdate_register_flb(handler, &iommu_flb); +} +EXPORT_SYMBOL(iommu_liveupdate_register_flb); + +/** + * iommu_liveupdate_unregister_flb() - Unregister IOMMU Live Update FLB + * @handler: Live Update file handler to associate FLB with + */ +void iommu_liveupdate_unregister_flb(struct liveupdate_file_handler *handl= er) +{ + liveupdate_unregister_flb(handler, &iommu_flb); +} +EXPORT_SYMBOL(iommu_liveupdate_unregister_flb); diff --git a/include/linux/iommu-liveupdate.h b/include/linux/iommu-liveupd= ate.h new file mode 100644 index 000000000000..4755ab3cd67a --- /dev/null +++ b/include/linux/iommu-liveupdate.h @@ -0,0 +1,28 @@ +/* SPDX-License-Identifier: GPL-2.0 */ + +/* + * Copyright (C) 2026, Google LLC + * Author: Samiullah Khawaja + */ + +#ifndef _LINUX_IOMMU_LIVEUPDATE_H +#define _LINUX_IOMMU_LIVEUPDATE_H + +#include +#include +#include + +#ifdef CONFIG_IOMMU_LIVEUPDATE +int iommu_liveupdate_register_flb(struct liveupdate_file_handler *handler); +void iommu_liveupdate_unregister_flb(struct liveupdate_file_handler *handl= er); +#else +static inline int iommu_liveupdate_register_flb(struct liveupdate_file_han= dler *handler) +{ + return 0; +} + +static inline void iommu_liveupdate_unregister_flb(struct liveupdate_file_= handler *handler) +{ +} +#endif +#endif /* _LINUX_IOMMU_LIVEUPDATE_H */ diff --git a/include/linux/kho/abi/iommu.h b/include/linux/kho/abi/iommu.h new file mode 100644 index 000000000000..aa42085409e5 --- /dev/null +++ b/include/linux/kho/abi/iommu.h @@ -0,0 +1,212 @@ +/* SPDX-License-Identifier: GPL-2.0 */ + +/* + * Copyright (C) 2026, Google LLC + * Author: Samiullah Khawaja + */ + +#ifndef _LINUX_KHO_ABI_IOMMU_H +#define _LINUX_KHO_ABI_IOMMU_H + +#include +#include +#include + +/** + * DOC: IOMMU File-Lifecycle Bound (FLB) Live Update ABI + * + * This header defines the ABI for preserving IOMMU state across kexec usi= ng + * Live Update File-Lifecycle Bound (FLB) data. + * + * This interface is a contract. Any modification to any of the serializat= ion + * structs defined here constitutes a breaking change. Such changes require + * incrementing the version number in IOMMU_LUO_FLB_VERSION. + * + * Memory Layout of Serialization Structures: + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + * + * Each serialized type (IOMMU, Domain, Device) is stored in a linked list= of + * arrays. The first array is allocated initially. When an array is full, = a new + * array is allocated and its physical address is stored in the next_array= _phys + * field of the hdr of the current array. + * + * Top Level (struct iommu_flb_ser): + * +---------------------------+ + * | - iommu_array_phys | + * | - iommu_domain_array_phys | + * | - device_array_phys | + * +---------------------------+ + * + * Each Array contains the serialized objects of the respective type. For + * example see below the representation of struct iommu_domain_array_ser. + * + * +---------------------------+ +---------------------------+ + * | iommu_domain_array_ser |-->| iommu_domain_array_ser |--> NULL + * | - hdr.next_array_phys | | - hdr.next_array_phys | + * | - hdr.nr_objects | | - hdr.nr_objects | + * | | | | + * | objects[]: | | objects[]: | + * | [ iommu_domain_ser ] | | [ iommu_domain_ser ] | + * | [ iommu_domain_ser ] | | [ iommu_domain_ser ] | + * | ... | | ... | + * +---------------------------+ +---------------------------+ + * + * Each object in the array starts with a common header (iommu_hdr_ser). + * For example, the layout of struct iommu_domain_ser is: + * + * +-----------------------------+ + * | iommu_domain_ser | + * | +-------------------------+ | + * | | hdr (iommu_hdr_ser) | | + * | | - ref_count | | + * | | - deleted / incoming | | + * | +-------------------------+ | + * | - top_table_phys | | + * | - top_level | | + * | - restored_domain | | + * +-----------------------------+ + * + * This pattern applies identically to iommu_device_ser and iommu_hw_ser. + */ + +#define IOMMU_LUO_FLB_COMPATIBLE "iommu-liveupdate" +#define IOMMU_LUO_FLB_VERSION 1 + +/** + * enum iommu_type_ser - Type of the IOMMU being preserved + * @IOMMU_INVALID: Invalid type of IOMMU + * + * IOMMU type is stored in the IOMMU HW state to differentiate between var= ious + * IOMMU HWs. + */ +enum iommu_type_ser { + IOMMU_INVALID, +}; + +#define IOMMU_SER_FLAG_DELETED (1 << 0) +#define IOMMU_SER_FLAG_INCOMING (1 << 1) + +/** + * struct iommu_hdr_ser - Common header for all serialized IOMMU objects + * @ref_count: Reference count for the object + * @flags: Bitmask of IOMMU_SER_FLAG_* flags indicating object state + */ +struct iommu_hdr_ser { + u32 ref_count; + u32 flags; +} __packed; + +/** + * struct iommu_domain_ser - Serialized state of an IOMMU domain + * @hdr: Common object header + * @top_table_phys: Physical address of the top-level page table + * @top_level: Level of the top-level page table + * @vasz: Virtual Address Size + * @sign_extend: FEAT_SIGN_EXTEND is enabled for this domain + * @restored_domain: Pointer to the restored domain (valid only after rest= ore) + */ +struct iommu_domain_ser { + struct iommu_hdr_ser hdr; + u64 top_table_phys; + u64 top_level; + u32 vasz; + u8 sign_extend; + u8 padding[3]; + struct iommu_domain *restored_domain; +} __packed; + +/** + * struct iommu_dev_map_ser - Serialized mapping between device, domain, + * and IOMMU instance. + * @attachment_id: ID of the attachment between device and domain. + * @domain_phys: Physical address of the domain + * @iommu_phys: Physical address of the IOMMU + */ +struct iommu_dev_map_ser { + u64 attachment_id; + u64 domain_phys; + u64 iommu_phys; +} __packed; + +/** + * struct iommu_device_ser - Serialized state of a device + * @hdr: Common object header + * @devid: Device ID + * @pci_domain_nr: PCI domain number + * @dma_owner_token: Token to identify the DMA owner of this device + * @domain_iommu_ser: Domain and IOMMU mapping + */ +struct iommu_device_ser { + struct iommu_hdr_ser hdr; + u32 devid; + u32 pci_domain_nr; + u64 dma_owner_token; + struct iommu_dev_map_ser domain_iommu_ser; +} __packed; + +/** + * struct iommu_hw_ser - Serialized state of an IOMMU instance + * @hdr: Common object header + * @token: Unique token for the IOMMU + * @type: IOMMU type serialized state belongs to + */ +struct iommu_hw_ser { + struct iommu_hdr_ser hdr; + u64 token; + u64 type; +} __packed; + +/** + * struct iommu_array_hdr_ser - Header for an array of serialized objects + * @next_array_phys: Physical address of the next array of objects + * @nr_objects: Number of objects in the current array + */ +struct iommu_array_hdr_ser { + u64 next_array_phys; + u64 nr_objects; +} __packed; + +/** + * struct iommu_hw_array_ser - An array containing serialized IOMMU HWs + * @hdr: Array header + * @objects: Array of serialized IOMMU devices + */ +struct iommu_hw_array_ser { + struct iommu_array_hdr_ser hdr; + struct iommu_hw_ser objects[]; +} __packed; + +/** + * struct iommu_domain_array_ser - An array containing serialized domains + * @hdr: Array header + * @objects: Array of serialized domains + */ +struct iommu_domain_array_ser { + struct iommu_array_hdr_ser hdr; + struct iommu_domain_ser objects[]; +} __packed; + +/** + * struct iommu_device_array_ser - An array containing serialized devices + * @hdr: Array header + * @objects: Array of serialized devices + */ +struct iommu_device_array_ser { + struct iommu_array_hdr_ser hdr; + struct iommu_device_ser objects[]; +} __packed; + +/** + * struct iommu_flb_ser - Top-level serialization structure + * @iommu_array_phys: Physical address of the first array of IOMMU HWs + * @iommu_domain_array_phys: Physical address of the first array of domains + * @device_array_phys: Physical address of the first array of devices + */ +struct iommu_flb_ser { + u64 version; + u64 iommu_array_phys; + u64 iommu_domain_array_phys; + u64 device_array_phys; +} __packed; + +#endif /* _LINUX_KHO_ABI_IOMMU_H */ --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EF5B4286D5E for ; Mon, 21 Sep 2026 00:48:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951726; cv=none; b=J/OfzgOBiaEsthaP/r3faV9ibAUWzaHzbYje5ObmXlpUxbYK4RgIHwZDKmI7aCqsNtPo4F7BBLBvEQ0rXW/1bvGClA1Chk1J5mxMakfLNifIGyMrGIK+1OdNssyE6cNnpb/UWsWm+5/ldX35YSeCHpkEhjXUwBHtc37SZKLS1M4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951726; c=relaxed/simple; bh=cH4LgWtd8WAbba4aDJwQQHyfDWm4RDvnu4B8RSUj3tE=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=j35/SCL9etRj6GZRfhzPXfGm9LnHiPInkQ24SvUsd2x5/E1rUCTdQEzFB5gxoecXn9/mnOtK1D4oLqzlDXWktx4s7mi7wibir5ABy8Jt9OljlAylc9nZrfSu39aIR+KRilQBuOhUeESmE0izhkSHy3NDmpSrtFkCajUEWIT7V2s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=LkwyINz8; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="LkwyINz8" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2db3b126c9fso33297255ad.0 for ; Sun, 20 Sep 2026 17:48:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951723; x=1790556523; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=pz0sP5OadudQTU/jOFxKypYztFVbcpWTHcp4vp3QfFc=; b=LkwyINz8zsoiC6aaUvYbiaFkiBmWiwFWRv3YCFu3+w8ZW4WHoXjgyglHsslaJbvYwl GUMiXayqXbqDc3pv08lga8wuQPuDpOkKnYPpZL3hAKxM377csVN92kiDP2vkw/GAcpFE FUQSXNDQaUApM/dyC103efA68CiLDljCyC+XoBD4VFh0KANy3FvxKKk+AQtVJ5mZUE08 lloFLD4+I7zcSX7gLGvYH89TXc3EeYb+yQjckrNp0bCnI6jONsmqan1P8ew+G987Uy81 p2ELWBMYDI4jW66trYDJRtzs8EEqc2BL+zLogghiJqX0piO2+v6luDyB6VH3DQw80VBN qHkQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951723; x=1790556523; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=pz0sP5OadudQTU/jOFxKypYztFVbcpWTHcp4vp3QfFc=; b=suNHVb5nqE1Yfze6RrYMpa1yaJHZWFDjaAmjxF9i+VZPwpH8uAdp94SWfckseIOz3O cLZpcxwrxPnOSu4p4ywCQoRH5DzMWze19MbaRNujtPBbMdpdug9HCfwKKxP3vIXZovLU zeRjqgyZOgbQXao+TdkorR2OeGBzBVQ+5LCnuJm5i1ysRKinQisMOPaw+vNRUx3SFUpn DmTNVn22rubDKJIFYKTNBH5nI82YMS1ccAcA9QpO0M0DOXKIZr7WiqJr2vXoPVfV1iLM JF7XWRgI0/oOugw8N53mRVlCRuHD8vj3vzLhVRaxv4ZaTQRVi0RhsgmI2v+mxn8gwpuC ySjA== X-Forwarded-Encrypted: i=1; AKwUvBzczSYxTMXOsrCG7tVPaMnnyDWfFPjaiC9yU0slr7xmzkGxu7kFJzyojHKWlE6mElIaJ+eky48nm2HPdJw=@vger.kernel.org X-Gm-Message-State: AFuF++nDflN+xyWPF9VgA6ZlCApS7kwKVoQP6hQb0BkHgfuczx20Ql9e 3s1umOkPWR794QcaXiY3w/xcZYxH9sISU1MWiitpyBxkwAqs51fa8S50w1zO44eKN78aA/1xhOa zV0LLVYlWRhI7iA== X-Received: from pli4.prod.google.com ([2002:a17:902:c104:b0:2df:3764:615]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:234d:b0:2da:e967:7953 with SMTP id d9443c01a7336-2ddb1b0d343mr144393205ad.12.1789951723018; Sun, 20 Sep 2026 17:48:43 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:19 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-4-skhawaja@google.com> Subject: [PATCH v5 03/18] iommu/pages: Add APIs to preserve/unpreserve/restore iommu pages From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Pranjal Shrivastava , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" IOMMU pages are allocated/freed using APIs using struct ioptdesc. For the proper preservation and restoration of ioptdesc add helper functions. Reviewed-by: Pranjal Shrivastava Signed-off-by: Samiullah Khawaja --- drivers/iommu/iommu-pages.c | 108 ++++++++++++++++++++++++++++++++++-- drivers/iommu/iommu-pages.h | 30 ++++++++++ 2 files changed, 134 insertions(+), 4 deletions(-) diff --git a/drivers/iommu/iommu-pages.c b/drivers/iommu/iommu-pages.c index 3bab175d8557..96cc935fb4e1 100644 --- a/drivers/iommu/iommu-pages.c +++ b/drivers/iommu/iommu-pages.c @@ -6,6 +6,7 @@ #include "iommu-pages.h" #include #include +#include #include =20 #define IOPTDESC_MATCH(pg_elm, elm) \ @@ -28,6 +29,13 @@ static inline size_t ioptdesc_mem_size(struct ioptdesc *= desc) return 1UL << (folio_order(ioptdesc_folio(desc)) + PAGE_SHIFT); } =20 +static inline void iommu_folio_update_stats(struct folio *folio, + unsigned long nr_pages) +{ + mod_node_page_state(folio_pgdat(folio), NR_IOMMU_PAGES, nr_pages); + lruvec_stat_mod_folio(folio, NR_SECONDARY_PAGETABLE, nr_pages); +} + /** * iommu_alloc_pages_node_sz - Allocate a zeroed page of a given size from * specific NUMA node @@ -80,8 +88,7 @@ void *iommu_alloc_pages_node_sz(int nid, gfp_t gfp, size_= t size) * rather large, i.e. multiple gigabytes in size. */ pgcnt =3D 1UL << order; - mod_node_page_state(folio_pgdat(folio), NR_IOMMU_PAGES, pgcnt); - lruvec_stat_mod_folio(folio, NR_SECONDARY_PAGETABLE, pgcnt); + iommu_folio_update_stats(folio, pgcnt); =20 return folio_address(folio); } @@ -95,8 +102,7 @@ static void __iommu_free_desc(struct ioptdesc *iopt) if (IOMMU_PAGES_USE_DMA_API) WARN_ON_ONCE(iopt->incoherent); =20 - mod_node_page_state(folio_pgdat(folio), NR_IOMMU_PAGES, -pgcnt); - lruvec_stat_mod_folio(folio, NR_SECONDARY_PAGETABLE, -pgcnt); + iommu_folio_update_stats(folio, -pgcnt); folio_put(folio); } =20 @@ -131,6 +137,100 @@ void iommu_put_pages_list(struct iommu_pages_list *li= st) } EXPORT_SYMBOL_GPL(iommu_put_pages_list); =20 +#if IS_ENABLED(CONFIG_IOMMU_LIVEUPDATE) +/** + * iommu_unpreserve_pages - Unpreserve pages that were preserved in KHO + * @virt: Virtual address of page to be unpreserved + */ +void iommu_unpreserve_pages(void *virt) +{ + kho_unpreserve_folio(ioptdesc_folio(virt_to_ioptdesc(virt))); +} +EXPORT_SYMBOL_GPL(iommu_unpreserve_pages); + +/** + * iommu_preserve_pages - Preserve pages using KHO + * @virt: Virtual address of page to preserve + * + * Returns 0 on success, negative error on failure + */ +int iommu_preserve_pages(void *virt) +{ + return kho_preserve_folio(ioptdesc_folio(virt_to_ioptdesc(virt))); +} +EXPORT_SYMBOL_GPL(iommu_preserve_pages); + +/** + * iommu_unpreserve_pages_list - Unpreserve a list of KHO preserved pages + * @list: List of pages to unpreserve + */ +void iommu_unpreserve_pages_list(struct iommu_pages_list *list) +{ + struct ioptdesc *iopt; + + list_for_each_entry(iopt, &list->pages, iopt_freelist_elm) + kho_unpreserve_folio(ioptdesc_folio(iopt)); +} +EXPORT_SYMBOL_GPL(iommu_unpreserve_pages_list); + +/** + * iommu_restore_pages - Restore pages that were preserved in KHO + * @phys: Physical address of page to restore + */ +void iommu_restore_pages(u64 phys) +{ + struct ioptdesc *iopt; + struct folio *folio; + unsigned long pgcnt; + unsigned int order; + + folio =3D kho_restore_folio(phys); + BUG_ON(!folio); + + iopt =3D folio_ioptdesc(folio); + + /* + * For the restored pages incoherent is set to false as these are not + * mapped using the DMA_API. The remapping of these pages using DMA_API + * is not needed as these are not going to be written to by the new + * kernel. + */ + iopt->incoherent =3D false; + + order =3D folio_order(folio); + pgcnt =3D 1UL << order; + iommu_folio_update_stats(folio, pgcnt); +} +EXPORT_SYMBOL_GPL(iommu_restore_pages); + +/** + * iommu_preserve_pages_list - Preserve list of pages using KHO + * @list: List of pages to preserve + * + * Returns 0 on success, negative error on failure + */ +int iommu_preserve_pages_list(struct iommu_pages_list *list) +{ + struct ioptdesc *iopt; + int ret; + + list_for_each_entry(iopt, &list->pages, iopt_freelist_elm) { + ret =3D kho_preserve_folio(ioptdesc_folio(iopt)); + if (ret) + goto err; + } + + return 0; + +err: + list_for_each_entry_continue_reverse(iopt, &list->pages, iopt_freelist_el= m) + kho_unpreserve_folio(ioptdesc_folio(iopt)); + + return ret; +} +EXPORT_SYMBOL_GPL(iommu_preserve_pages_list); +#endif + /** * iommu_pages_start_incoherent - Setup the page for cache incoherent oper= ation * @virt: The page to setup diff --git a/drivers/iommu/iommu-pages.h b/drivers/iommu/iommu-pages.h index e9e605b5fa3a..3c409ec3d170 100644 --- a/drivers/iommu/iommu-pages.h +++ b/drivers/iommu/iommu-pages.h @@ -53,6 +53,36 @@ void *iommu_alloc_pages_node_sz(int nid, gfp_t gfp, size= _t size); void iommu_free_pages(void *virt); void iommu_put_pages_list(struct iommu_pages_list *list); =20 +#if IS_ENABLED(CONFIG_IOMMU_LIVEUPDATE) +int iommu_preserve_pages(void *virt); +void iommu_unpreserve_pages(void *virt); +int iommu_preserve_pages_list(struct iommu_pages_list *list); +void iommu_unpreserve_pages_list(struct iommu_pages_list *list); +void iommu_restore_pages(u64 phys); +#else +static inline int iommu_preserve_pages(void *virt) +{ + return -EOPNOTSUPP; +} + +static inline void iommu_unpreserve_pages(void *virt) +{ +} + +static inline int iommu_preserve_pages_list(struct iommu_pages_list *list) +{ + return -EOPNOTSUPP; +} + +static inline void iommu_unpreserve_pages_list(struct iommu_pages_list *li= st) +{ +} + +static inline void iommu_restore_pages(u64 phys) +{ +} +#endif + /** * iommu_pages_list_add - add the page to a iommu_pages_list * @list: List to add the page to --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f197.google.com (mail-pl1-f197.google.com [209.85.214.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 212722DC79B for ; Mon, 21 Sep 2026 00:48:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.197 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951734; cv=none; b=sTAsNCGBqeXlmTbpF4+bdo6hLks7V983tKR2JUVKbZoA2EcFoX3KNdkWgZ89l6BzZmRp0Wsr4EMFdsxUfx2UV/LJUSiXHyOf6X+Shbu5b6dceSqNy+wwOFNZR3Mk0gG+70klnv5yB3e7FM77viWh51tJrGzj93H/5iBPyxvsISs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951734; c=relaxed/simple; bh=iOp397Tg1uyFDkJ+PjIYFgFLwFOnJQaPRudOX/K8Wi0=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=AfE2MiQ7InZdAVDEvVJKmS9IJSUenhpc2UoXN0wl1Z9blCVhlBMfyUybyMveIR5jLEDI2dqPL4rPGg0zKZ0RybZUOAvWOUnxZNrDUzVRFNI3gIKe7SoKeGeVolvPCa6senrQ5I57yqQu8hBJ3HkAoFO1Unsocs3WkRweOU+ztOM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=ry2JZp6i; arc=none smtp.client-ip=209.85.214.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="ry2JZp6i" Received: by mail-pl1-f197.google.com with SMTP id d9443c01a7336-2d6fb956002so34133305ad.1 for ; Sun, 20 Sep 2026 17:48:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951724; x=1790556524; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=2C1NA4Y2gkk+MrlRvGBVzGfZjWst0By6u+UYLSUEdKI=; b=ry2JZp6iVCehqsiTBcyUlo61VguXzjqGgWhm/03xGvoP0r0FTCUfcif1h07EwVk8MR ehgiK3lBvnfVQMUrGwQnslnBmTjIrPe8nntAdMsbO5M1WMtTRQNSwdjfAnsQ6uco/Wu+ zzUi1ZmBjt2xCKm8NzBWw+y6sk+3YCf4q2hDESjdi+jNC8tlA+F4PB7IRIG9Y1waNoYt z0I+G/TIO5UaRyqb23M47SlBaC8RON89F5lqvUKaQOi0D5g2yhyvD4YAuDMTfAt9brHM BR6mgbKOlAFEHn2txl1Niz7QsrxTkTAOj/TptGkBMWTz8wOoWhKTvGd4S8iGJzHDfHpx xXnA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951724; x=1790556524; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=2C1NA4Y2gkk+MrlRvGBVzGfZjWst0By6u+UYLSUEdKI=; b=C3jxjxVeQ1+DiDUon2wD+Ebm3cfXYR4WSSGG/Eto5WHLfcqrI5Gkuy+hywmW573eD2 TBmLBpV5M943JkunfJ3hEJ53HW7fC2C+t2VRzRy+q2q4vqZ8ojVTd7wvDSXRcQI5G49q tuUZJH90W5EDXt2vrydDV+CanNvxfXJWFErsa9e4cF54WMLfX3R2Q4rEn/6GQxtzQM89 VJR5HFfRSZBvjkTctbqcAWMheYVDWpTAScGqQjcL6K2VqDZzr7CxsjQKk54E+fBAIO4M ixMaMp1jJv7eQW+LCTSM5Aq/ZDH+XEWjz3H8Jf4OnWzvl3VcjOhgmTaGJKUlruqzic0X X13w== X-Forwarded-Encrypted: i=1; AKwUvBwqhp7TaEFoZ8VgE5bk16hfWbF1MvRXabfH4RzR+GkyhyAk70z2oGhK5l5G9eZG8t/TSRPxq9cPl6X8sVA=@vger.kernel.org X-Gm-Message-State: AFuF++lbMpZsCGQGsutZ3KDJ1FUHWrfxZBSFSWcxKcbmdMhtCeqpkjj1 ept/GmBnFn6igzJFuPbWdRQLcxtTlQ6hZKuT48XA4qWq8LEVQscZnhfiBUnGJ3UL4eYC5NPPm82 7kH19owvDGvMwYw== X-Received: from plmg19.prod.google.com ([2002:a17:903:3cd3:b0:2dd:b156:5a6e]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:b106:b0:2dd:c053:d73d with SMTP id d9443c01a7336-2ddc053d7camr58889585ad.36.1789951724292; Sun, 20 Sep 2026 17:48:44 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:20 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-5-skhawaja@google.com> Subject: [PATCH v5 04/18] iommupt: Implement preserve/unpreserve/restore callbacks From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add iommupt ops for presevation, unpresevation and restoration of iommu page tables for liveupdate. Use the existing page walker to preserve the ioptdesc of the top_table and the lower tables. Preserve top_level, VASZ and FEAT Sign Extended to restore the domain in the next kernel. On restore, the domain has only the preserved features enabled and all the other features are zeroed. This is ok since the restored domain is made immutable and can only be freed. A kunit test is added to verify that the IOMMU domain free can be done with trimmed features. Signed-off-by: Samiullah Khawaja --- drivers/iommu/generic_pt/iommu_pt.h | 152 ++++++++++++++++++++++ drivers/iommu/generic_pt/kunit_iommu_pt.h | 30 +++++ include/linux/generic_pt/iommu.h | 30 +++++ 3 files changed, 212 insertions(+) diff --git a/drivers/iommu/generic_pt/iommu_pt.h b/drivers/iommu/generic_pt= /iommu_pt.h index 07ec2b3ab986..2945ddb771d1 100644 --- a/drivers/iommu/generic_pt/iommu_pt.h +++ b/drivers/iommu/generic_pt/iommu_pt.h @@ -16,6 +16,7 @@ #include "../iommu-pages.h" #include #include +#include =20 enum { SW_BIT_CACHE_FLUSH_DONE =3D 0, @@ -1005,6 +1006,152 @@ static int NS(map_range)(struct pt_iommu *iommu_tab= le, dma_addr_t iova, return ret; } =20 +#ifdef CONFIG_IOMMU_LIVEUPDATE +static void NS(unpreserve)(struct pt_iommu *iommu_table, struct iommu_doma= in_ser *ser) +{ + struct pt_common *common =3D common_from_iommu(iommu_table); + struct pt_range range =3D pt_all_range(common); + struct pt_iommu_collect_args collect =3D { + .pending.free_list =3D IOMMU_PAGES_LIST_INIT( + collect.pending.free_list), + }; + + iommu_pages_list_add(&collect.pending.free_list, range.top_table); + + /* + * pt_walk_range() will never fail as check_mapped is not set in the + * collect args. + */ + pt_walk_range(&range, __collect_tables, &collect); + + iommu_unpreserve_pages_list(&collect.pending.free_list); +} + +static int NS(preserve)(struct pt_iommu *iommu_table, struct iommu_domain_= ser *ser) +{ + struct pt_common *common =3D common_from_iommu(iommu_table); + struct pt_range range =3D pt_all_range(common); + struct pt_iommu_collect_args collect =3D { + .pending.free_list =3D IOMMU_PAGES_LIST_INIT( + collect.pending.free_list), + }; + int ret; + + iommu_pages_list_add(&collect.pending.free_list, range.top_table); + + /* + * pt_walk_range() will never fail as check_mapped is not set in the + * collect args. + */ + pt_walk_range(&range, __collect_tables, &collect); + + ret =3D iommu_preserve_pages_list(&collect.pending.free_list); + if (ret) + return ret; + + ser->top_table_phys =3D virt_to_phys(range.top_table); + ser->top_level =3D range.top_level; + + /* + * VASZ and SIGN_EXTEND will be needed in next kernel for collector page + * table walk to restore and free pages. + * + * Use the max_vasz_lg2 from range, as that is the current one if + * DYNAMIC_TOP is supported. + */ + ser->vasz =3D range.max_vasz_lg2; + ser->sign_extend =3D pt_feature(common, PT_FEAT_SIGN_EXTEND); + + return 0; +} + +static int __restore_tables(struct pt_range *range, void *arg, + unsigned int level, struct pt_table_p *table) +{ + struct pt_state pts =3D pt_init(range, level, table); + int ret; + + for_each_pt_level_entry(&pts) { + if (pts.type =3D=3D PT_ENTRY_TABLE) { + iommu_restore_pages(virt_to_phys(pts.table_lower)); + + /* + * pt_descend can only fail if pts.table_lower is not + * init. So the if statement below is dead code. + */ + ret =3D pt_descend(&pts, arg, __restore_tables); + if (ret) + return ret; + } + } + + return 0; +} + +static void NS(deinit)(struct pt_iommu *iommu_table); +static const struct pt_iommu_ops NS(ops_immutable) =3D { + .deinit =3D NS(deinit), +}; + +static int NS(restore)(struct pt_iommu *iommu_table, struct iommu_domain_s= er *ser) +{ + struct pt_common *common =3D common_from_iommu(iommu_table); + struct pt_range range; + + if (ser->vasz > common->max_vasz_lg2 || + ser->top_level > PT_MAX_TOP_LEVEL) + return -EINVAL; + + /* + * Override the max_vasz_lg2 here to the one calculated using top_range + * in previous kernel. This makes sure that any later page walks use the + * clamped max_vasz_lg2 instead of using the full hardware maximum + * max_vasz_lg2. + */ + common->max_vasz_lg2 =3D ser->vasz; + + /* + * Restored page tables are strictly transient and only permitted to be + * destroyed via deinit() op. Because we only preserve user-assigned + * devices utilizing pass-through frameworks (VFIO / IOMMUFD), any + * concurrent or subsequent map/unmap operations on a restored domain + * are explicitly blocked at the subsystem boundary (e.g., via IOAS + * immutability). + */ + iommu_table->ops =3D &NS(ops_immutable); + + /* + * It is safe to override this here since this domain is immutable and + * can only be freed. This also means that for non-coherent pages DMA + * mappings will not be created using the DMA API. + * + * Note that the FMT specific bits are not cleared as they might be + * set/preserved/restored by the driver to handle special cases that are + * required for page walks. + */ + common->features &=3D ~GENMASK(PT_FEAT_FMT_START - 1, 0); + if (ser->sign_extend) + common->features |=3D BIT(PT_FEAT_SIGN_EXTEND); + + range =3D pt_all_range(common); + iommu_restore_pages(ser->top_table_phys); + + /* + * The old top_table can be freed here as it is not used for any IOMMU + * mappings yet, because the restore happens during device probe. + */ + iommu_pages_free_incoherent(range.top_table, + iommu_table->iommu_device); + + /* Set the restored top table */ + pt_top_set(common, phys_to_virt(ser->top_table_phys), ser->top_level); + + /* Restore all pages*/ + range =3D pt_all_range(common); + return pt_walk_range(&range, __restore_tables, NULL); +} +#endif + struct pt_unmap_args { struct iommupt_pending_gather pending; pt_vaddr_t unmapped; @@ -1188,6 +1335,11 @@ static const struct pt_iommu_ops NS(ops) =3D { #endif .get_info =3D NS(get_info), .deinit =3D NS(deinit), +#ifdef CONFIG_IOMMU_LIVEUPDATE + .preserve =3D NS(preserve), + .unpreserve =3D NS(unpreserve), + .restore =3D NS(restore), +#endif }; =20 static int pt_init_common(struct pt_common *common) diff --git a/drivers/iommu/generic_pt/kunit_iommu_pt.h b/drivers/iommu/gene= ric_pt/kunit_iommu_pt.h index ece1c9b8c55d..83bf18be8c1e 100644 --- a/drivers/iommu/generic_pt/kunit_iommu_pt.h +++ b/drivers/iommu/generic_pt/kunit_iommu_pt.h @@ -427,6 +427,35 @@ static void test_mixed(struct kunit *test) check_iova(test, start, oa, len); } =20 +static void test_restore_free(struct kunit *test) +{ + struct kunit_iommu_priv *priv =3D test->priv; + struct pt_range top_range =3D pt_top_range(priv->common); + u64 start =3D 0x3fe400ULL << 12; + u64 end =3D 0x4c0600ULL << 12; + pt_vaddr_t len =3D end - start; + + if (top_range.last_va <=3D start || sizeof(unsigned long) =3D=3D 4) + kunit_skip(test, "range is too small"); + if ((priv->safe_pgsize_bitmap & GENMASK(30, 21)) !=3D (BIT(30) | BIT(21))) + kunit_skip(test, "incompatible psize"); + + /* Map a large mixed range to populate multiple levels of page tables */ + do_map(test, start, start, len); + + /* + * Simulate a restored state by clearing all features except SIGN_EXTEND + * and DMA_INCOHERENT. This verifies that the generic page table free + * walker can correctly tear down a populated domain when other features + * are zeroed. Also set max_vasz_lg2 as done by the actual restore() + * function. + */ + priv->common->features &=3D (BIT(PT_FEAT_SIGN_EXTEND) | BIT(PT_FEAT_DMA_I= NCOHERENT)); + priv->common->max_vasz_lg2 =3D top_range.max_vasz_lg2; + + /* The domain will be freed when the test exits. */ +} + static struct kunit_case iommu_test_cases[] =3D { KUNIT_CASE_FMT(test_increase_level), KUNIT_CASE_FMT(test_map_simple), @@ -435,6 +464,7 @@ static struct kunit_case iommu_test_cases[] =3D { KUNIT_CASE_FMT(test_random_map), KUNIT_CASE_FMT(test_pgsize_boundary), KUNIT_CASE_FMT(test_mixed), + KUNIT_CASE_FMT(test_restore_free), {}, }; =20 diff --git a/include/linux/generic_pt/iommu.h b/include/linux/generic_pt/io= mmu.h index dd0edd02a48a..faa41d8032fe 100644 --- a/include/linux/generic_pt/iommu.h +++ b/include/linux/generic_pt/iommu.h @@ -13,6 +13,7 @@ struct iommu_iotlb_gather; struct pt_iommu_ops; struct pt_iommu_driver_ops; struct iommu_dirty_bitmap; +struct iommu_domain_ser; =20 /** * DOC: IOMMU Radix Page Table @@ -166,6 +167,35 @@ struct pt_iommu_ops { * table from all HW access and all caches. */ void (*deinit)(struct pt_iommu *iommu_table); + + /** + * @preserve: Preserve the iommu page table for liveupdate + * @iommu_table: Table to preserve + * @ser: Serialization struct to fill with preserved state + * + * Preserve iommu page table and the relevant state for liveupdate. The + * caller must make sure that the page table is not updated during and + * after preservation. + */ + int (*preserve)(struct pt_iommu *iommu_table, struct iommu_domain_ser *se= r); + + /** + * @unpreserve: Unpreserve the iommu page table + * @iommu_table: Table to unpreserve + * @ser: Serialization struct that contains preserved state + */ + void (*unpreserve)(struct pt_iommu *iommu_table, struct iommu_domain_ser = *ser); + + /** + * @restore: Restore the iommu page table after liveupdate + * @iommu_table: Table to restore the state into + * @ser: Serialization struct that contains preserved state + * + * The iommu_table is back filled with the restored state that was + * preserved in the serialization struct. + */ + int (*restore)(struct pt_iommu *iommu_table, struct iommu_domain_ser *ser= ); + }; =20 /** --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pg1-f200.google.com (mail-pg1-f200.google.com [209.85.215.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 33ECA2D9484 for ; Mon, 21 Sep 2026 00:48:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.200 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951729; cv=none; b=emgujGbP61RM4jOhcqUeeJx9+kGuJJcyo+6kWNPG8IvaS+0A9yI+/QB1jDsuYdV1QDtRuxif3PffjlI7+wEsdNBFuYMWAXZjFZA+KJ3YD/PNIHmw4PHxBJHDPISDUdTYXMTdXi5FTcHsVnGtofMj1kYdYTVjEu/2ZeaEDpiyiqI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951729; c=relaxed/simple; bh=nPezMDYp5xpL65qddw2WVC7a5eW4PH+MLRIAj/7Q7y8=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=MdmIpLasmV4q8Ax7QanveMT0iFRiheyW5Xy5zvb8ufCVedVPNvlmKla0d8d9Bq3QjQC7O71Fj5zovD/z4CPhsiQZ7U3Nw5L1T342B0wf2T8qBzE9IsE+Ki9tDUgbO/A+XBqhgsWxrXuFNpHDQeC2RZoxARrsEiB1rTdBputQeX0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=jMOKF+uw; arc=none smtp.client-ip=209.85.215.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="jMOKF+uw" Received: by mail-pg1-f200.google.com with SMTP id 41be03b00d2f7-cbedbd182f5so2166721a12.1 for ; Sun, 20 Sep 2026 17:48:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951725; x=1790556525; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=UKH4Gc6Z+bMWga98+CS78z+EWlxfzC75QI3yKAEUybQ=; b=jMOKF+uwEzepZrvfLeHgy1zdaudboCTuh/+Bq4lGaXGUy4wiWTelDPmLLn3FFpKUTG tNM85Yw2XpN46A/ZCIU5czo/18nQevRB5zJTeZxuNLaLZ4GS4jZKRccSCrgfCA2WXuyY sR82yWhbYhk0qljrxCF29Os5a+XI3kok3iJATVcIQFfwxdc9RnGRjS+zgEw+yaYP/hB7 64y3jFluFAQj6PthYHUbP9Ebgalw11EXmzMh5FyO5QvjiCpvdnpp6YiwIgUM+A+yoAVq eR4hv4vStXmziKaFDwG1KgH4tQkeivHAfKr3nRiEww0sNXwNiiMtqPAkrZbfHVwZl0/n LEDw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951725; x=1790556525; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=UKH4Gc6Z+bMWga98+CS78z+EWlxfzC75QI3yKAEUybQ=; b=LY9RThOJxRH+4/XUkBXctIr4/l8oRtsFywL+U4xOuLjEJf7XOqSMXnUxbOCCD+4zJF Iy5JfD0kYAvDwNxvykc3YzSNrvvs2Gjs/Ef7w6lq1dhlZjAsFm4wtB/k2i7wqt2ML7IY wEhpWVK0EELEKPQh5AFhUeKkShg1/qJZ49V9wpds5pBpAkzRBq/53hkbvogBMU/xzriT 3eK7B7Rt+6VckJmrG9jvDBnROQy8UQ6UeiEirGtjaFJQHKhEivbLyq3xbq8afCDRep+O k2MwY0WMrDlc1QQI0rdnyW9Q2AB0VIA4CpoGayIxHDABSmxe0hOEyi2D3733oBQjAwO7 zMjg== X-Forwarded-Encrypted: i=1; AKwUvBwpgadH/8V38NkuwdZqLybjuOL+5aSUTKom5YWhKuXbU3Ets8OLOBhnjDut4ElppIVE97NVdydSZPCPrqw=@vger.kernel.org X-Gm-Message-State: AFuF++nyXMi9Vb3gweYvrhyeyOm09F5htkAsVMCFRaJD+Vgk/ND9BTJn oQOR2IFJPZetBmfpIAMWtHATtaBAqn27d4YaZVyP1reI6A6cPkJCs9LmFbmqLzaRgNdg5pxqIf5 a1Uz9oQ65AEnb6w== X-Received: from plao7.prod.google.com ([2002:a17:903:3007:b0:2df:4e63:c417]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:3b86:b0:2dd:c100:251d with SMTP id d9443c01a7336-2ddc10025a3mr60411435ad.38.1789951725377; Sun, 20 Sep 2026 17:48:45 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:21 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-6-skhawaja@google.com> Subject: [PATCH v5 05/18] iommu: Implement IOMMU domain preservation From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Pranjal Shrivastava , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add IOMMU domain ops that can be implemented by the IOMMU drivers if they support IOMMU domain preservation across liveupdate. The new IOMMU domain preserve, unpreserve and restore APIs call these ops to perform respective live update operations. Reviewed-by: Pranjal Shrivastava Signed-off-by: Samiullah Khawaja --- drivers/iommu/liveupdate.c | 128 +++++++++++++++++++++++++++++++ include/linux/iommu-liveupdate.h | 11 +++ include/linux/iommu.h | 5 ++ 3 files changed, 144 insertions(+) diff --git a/drivers/iommu/liveupdate.c b/drivers/iommu/liveupdate.c index b644f4792532..be542efbb023 100644 --- a/drivers/iommu/liveupdate.c +++ b/drivers/iommu/liveupdate.c @@ -37,11 +37,15 @@ #define pr_fmt(fmt) "iommu: liveupdate: " fmt =20 #include +#include #include #include #include #include =20 +#define iommu_max_objs_per_page(_array) \ + ((PAGE_SIZE - sizeof(struct iommu_array_hdr_ser)) / sizeof((_array)->obje= cts[0])) + struct iommu_flb_obj { struct mutex lock; struct iommu_flb_ser *ser; @@ -251,3 +255,127 @@ void iommu_liveupdate_unregister_flb(struct liveupdat= e_file_handler *handler) liveupdate_unregister_flb(handler, &iommu_flb); } EXPORT_SYMBOL(iommu_liveupdate_unregister_flb); + +static int alloc_object_ser(void **curr_array_ptr, u64 max_objs) +{ + struct iommu_array_hdr_ser *curr_array =3D *curr_array_ptr; + struct iommu_array_hdr_ser *next_array; + + /* + * The objects marked as deleted are not reused to avoid traversal of + * linked-list and arrays. + */ + if (curr_array->nr_objects >=3D max_objs) { + next_array =3D kho_alloc_preserve(PAGE_SIZE); + if (IS_ERR(next_array)) + return PTR_ERR(next_array); + + curr_array->next_array_phys =3D virt_to_phys(next_array); + *curr_array_ptr =3D next_array; + curr_array =3D next_array; + } + + return curr_array->nr_objects++; +} + +static struct iommu_domain_ser *alloc_iommu_domain_ser(struct iommu_flb_ob= j *flb) +{ + int idx; + + idx =3D alloc_object_ser((void **) &flb->curr_domain_array, + iommu_max_objs_per_page(flb->curr_domain_array)); + if (idx < 0) + return ERR_PTR(idx); + + flb->curr_domain_array->objects[idx].hdr.ref_count =3D 1; + return &flb->curr_domain_array->objects[idx]; +} + +/** + * iommu_preserve_domain() - Preserve an IOMMU domain across live update + * @domain: Domain to preserve + * @ser: Pointer to receive the virtual serialized domain state handle + * + * Return: 0 on success, or negative error code. + */ +int iommu_preserve_domain(struct iommu_domain *domain, struct iommu_domain= _ser **ser) +{ + struct pt_iommu *pt =3D iommupt_from_domain(domain); + struct iommu_domain_ser *domain_ser; + struct iommu_flb_obj *flb_obj; + int ret; + + if (!pt || !pt->ops->preserve || !pt->ops->unpreserve) + return -EOPNOTSUPP; + + ret =3D liveupdate_flb_get_outgoing(&iommu_flb, (void **)&flb_obj); + if (ret) + return ret; + + mutex_lock(&flb_obj->lock); + if (domain->preserved_state) { + ret =3D -EBUSY; + goto out_unlock; + } + + domain_ser =3D alloc_iommu_domain_ser(flb_obj); + if (IS_ERR(domain_ser)) { + ret =3D PTR_ERR(domain_ser); + goto out_unlock; + } + + ret =3D pt->ops->preserve(pt, domain_ser); + if (ret) { + domain_ser->hdr.flags |=3D IOMMU_SER_FLAG_DELETED; + goto out_unlock; + } + + domain->preserved_state =3D domain_ser; + *ser =3D domain_ser; + ret =3D 0; +out_unlock: + mutex_unlock(&flb_obj->lock); + liveupdate_flb_put_outgoing(&iommu_flb); + return ret; +} +EXPORT_SYMBOL_GPL(iommu_preserve_domain); + +/** + * iommu_unpreserve_domain() - Unpreserve a preserved IOMMU domain + * @domain: Domain to unpreserve + */ +void iommu_unpreserve_domain(struct iommu_domain *domain) +{ + struct pt_iommu *pt =3D iommupt_from_domain(domain); + struct iommu_domain_ser *domain_ser; + struct iommu_flb_obj *flb_obj; + int ret; + + if (WARN_ON(!pt || !pt->ops->unpreserve)) + return; + + ret =3D liveupdate_flb_get_outgoing(&iommu_flb, (void **)&flb_obj); + if (WARN_ON(ret)) + return; + + mutex_lock(&flb_obj->lock); + if (!domain->preserved_state) + goto out_unlock; + + /* + * There is no check for attached devices here. The correctness relies + * on the Live Update Orchestrator's session lifecycle. All resources + * (iommufd, vfio devices) are preserved within a single session. If the + * session is torn down, the .unpreserve callbacks for all files will be + * invoked, ensuring a consistent cleanup without needing explicit + * refcounting for the serialized objects here. + */ + domain_ser =3D domain->preserved_state; + pt->ops->unpreserve(pt, domain_ser); + domain_ser->hdr.flags |=3D IOMMU_SER_FLAG_DELETED; + domain->preserved_state =3D NULL; +out_unlock: + mutex_unlock(&flb_obj->lock); + liveupdate_flb_put_outgoing(&iommu_flb); +} +EXPORT_SYMBOL_GPL(iommu_unpreserve_domain); diff --git a/include/linux/iommu-liveupdate.h b/include/linux/iommu-liveupd= ate.h index 4755ab3cd67a..caa9778eee2d 100644 --- a/include/linux/iommu-liveupdate.h +++ b/include/linux/iommu-liveupdate.h @@ -15,6 +15,8 @@ #ifdef CONFIG_IOMMU_LIVEUPDATE int iommu_liveupdate_register_flb(struct liveupdate_file_handler *handler); void iommu_liveupdate_unregister_flb(struct liveupdate_file_handler *handl= er); +int iommu_preserve_domain(struct iommu_domain *domain, struct iommu_domain= _ser **ser); +void iommu_unpreserve_domain(struct iommu_domain *domain); #else static inline int iommu_liveupdate_register_flb(struct liveupdate_file_han= dler *handler) { @@ -24,5 +26,14 @@ static inline int iommu_liveupdate_register_flb(struct l= iveupdate_file_handler * static inline void iommu_liveupdate_unregister_flb(struct liveupdate_file_= handler *handler) { } + +static inline int iommu_preserve_domain(struct iommu_domain *domain, struc= t iommu_domain_ser **ser) +{ + return -EOPNOTSUPP; +} + +static inline void iommu_unpreserve_domain(struct iommu_domain *domain) +{ +} #endif #endif /* _LINUX_IOMMU_LIVEUPDATE_H */ diff --git a/include/linux/iommu.h b/include/linux/iommu.h index ac43b8b93f14..26de40d5a98e 100644 --- a/include/linux/iommu.h +++ b/include/linux/iommu.h @@ -14,6 +14,7 @@ #include #include #include +#include #include =20 #define IOMMU_READ (1 << 0) @@ -249,6 +250,10 @@ struct iommu_domain { struct list_head next; }; }; + +#ifdef CONFIG_IOMMU_LIVEUPDATE + struct iommu_domain_ser *preserved_state; +#endif }; =20 static inline bool iommu_is_dma_domain(struct iommu_domain *domain) --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pg1-f198.google.com (mail-pg1-f198.google.com [209.85.215.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 640482DC783 for ; Mon, 21 Sep 2026 00:48:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951733; cv=none; b=HEjpthP5mxnMn8IeBaSU9iOzU0FQENhXXoK3lz383QWZuZqp5WR5KP0kCmZ9IKuwl9rqTKIuYVorfHxzPpXDV/gRSUzgibM3lVIiyCPBgIY/u0xDtNpNmTKP/ST2Ue2hpEt7kW4Fk3CoGj1cap/YbX7MzzxLUvTHFg0HrTQPPYE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951733; c=relaxed/simple; bh=KC381iw5UmgvllDgJ3YNjupbKUlYfiAYAjHYSVAumpI=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=UTqVyjARQbnJFU5HDlyeen7d5GK/hoL00Ku/+21CDP11FjyqRCX1vZF7a+QYI3ApjSCdnXfDCqPuh1nahQ1ZhHoHd78fwq7zdV1L2J6YRemxf1Ssg9T/9EXpO+bjNidn41vwkpkmuvjV00CpzqlhD8XJph8YMY/NHnh48WcVi6A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=KgVd6AcV; arc=none smtp.client-ip=209.85.215.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="KgVd6AcV" Received: by mail-pg1-f198.google.com with SMTP id 41be03b00d2f7-cc489e7a701so3739969a12.2 for ; Sun, 20 Sep 2026 17:48:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951726; x=1790556526; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=blp+qm96mUk79bE4u/LdP52s0pcty0khhzeGk8qCR7I=; b=KgVd6AcV4vmL1EnwsWcdr9UKRnVQx//uiQDV+WAX4QLuBJyQTsBCHsHwDhyroHIp+F A3EeCmpkgD2aZmmjhDK3P1VESt3C7QzvYHBrC7gNe5+DQ8K6Wsxlc7mAghLSdDMc7RP+ MWL/w7acWWmdkjI6GPw5KpASZMvHRXHW6d9ceQXQW6frBJPfU99bs6+em6z3M9KJAG3F V5NKO8MsWjXZMCHrSc3aaJUf6hfIkxr4/+QSdaImfNYzVn42JRYun0x6Cm90wEHQVSLi bijHBimetT4nlEgp3BwTRhHRhLp9BZMNe7lIfycYw2ghz5OynKyeAMeDZ/zZp573E5o9 sVpw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951726; x=1790556526; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=blp+qm96mUk79bE4u/LdP52s0pcty0khhzeGk8qCR7I=; b=MnJlKX/EyyZyBgu+6J516RjR0c/TESlLeMU2Cbnm76ZM7HJ4qGtpIlhDimWvC8Z6vW 7ETq3UQGX4bVR6W5vSSO/O50IOEDUyRsMOSFUtcSJN1SwobfvDawbgH19HdMl/YRk7/T zi/qXwJeN617IiaUpgCGAQxTYYbbJ0iKx0Zl8HAkfrbLgaYtE6bi7b1Azl5TP1Gwj+D/ UTvY1sNnSgwKKZ1sUoytMUiMKc3f+ar+ZDuV/m1kyiHfn+fistqagqEquLps7oaBmRvm Aogb5zgtChCxOsIIku1RnaHPpN66rAPkCwDTSQ2b9KpYVa9iNqCqosTR/YRgIVOnIsvR 32tg== X-Forwarded-Encrypted: i=1; AKwUvBzHffuwLpJZ+8b8M2dm8q0jTFRxracs1V117YeVBYMLrAhsrflV1KXZwyE6wwO4p7uWKDdcNQpZkwh3T2Q=@vger.kernel.org X-Gm-Message-State: AFuF++lG3JEHxjsC4M+KMDOHeybogk3qo18so91JgwiZuJfKgpEqwYMd k/teRDgEVPVRS5DBsqyTT8uv8UY2JLYb0aWYpKTVd6ZiIE+kSyJ15f3xG255sbMZP0YJx4JRjT5 gcCwK/1tsV7PJWw== X-Received: from pgdf8.prod.google.com ([2002:a05:6a02:5148:b0:cc4:c085:96d1]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:ce43:b0:3b4:7e2d:a3c2 with SMTP id adf61e73a8af0-3dd8c44e218mr14476854637.18.1789951726141; Sun, 20 Sep 2026 17:48:46 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:22 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-7-skhawaja@google.com> Subject: [PATCH v5 06/18] iommu: Implement device and IOMMU HW preservation From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add IOMMU ops to preserve/unpreserve the device state handled by the IOMMU driver. This is IOMMU specific state of the device that needs to be preserved for liveupdate. These ops can be implemented by the IOMMU drivers that support preservation of devices that have their IOMMU domains preserved. During device preservation the state of the associated IOMMU is also preserved as dependency. Signed-off-by: Samiullah Khawaja --- drivers/iommu/iommu-priv.h | 2 + drivers/iommu/iommu.c | 62 ++++++++- drivers/iommu/liveupdate.c | 212 +++++++++++++++++++++++++++++++ include/linux/iommu-liveupdate.h | 56 ++++++++ include/linux/iommu.h | 42 ++++++ 5 files changed, 373 insertions(+), 1 deletion(-) diff --git a/drivers/iommu/iommu-priv.h b/drivers/iommu/iommu-priv.h index aaffad5854fc..62612b1fc890 100644 --- a/drivers/iommu/iommu-priv.h +++ b/drivers/iommu/iommu-priv.h @@ -21,6 +21,8 @@ static inline const struct iommu_ops *dev_iommu_ops(struc= t device *dev) =20 void dev_iommu_free(struct device *dev); =20 +bool iommu_group_is_singleton(struct iommu_group *group); + const struct iommu_ops *iommu_ops_from_fwnode(const struct fwnode_handle *= fwnode); =20 static inline const struct iommu_ops *iommu_fwspec_ops(struct iommu_fwspec= *fwspec) diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c index cd1bca7ede9a..9c1ad8de3ad7 100644 --- a/drivers/iommu/iommu.c +++ b/drivers/iommu/iommu.c @@ -18,6 +18,7 @@ #include #include #include +#include #include #include #include @@ -328,6 +329,37 @@ void iommu_device_unregister(struct iommu_device *iomm= u) } EXPORT_SYMBOL_GPL(iommu_device_unregister); =20 +static int _iommu_for_each_dev_cb(struct device *dev, void *data) +{ + struct iommu_dev_iter *iter =3D data; + + if (dev->iommu && dev->iommu->iommu_dev =3D=3D iter->iommu) + return iter->fn(dev, iter->iommu, iter->arg); + + return 0; +} + +/** + * iommu_for_each_dev() - Iterate over all devices attached to an IOMMU + * @iter: Device iterator context + * + * Return: 0 on success, or negative error code. + */ +int iommu_for_each_dev(struct iommu_dev_iter *iter) +{ + int ret; + + for (int i =3D 0; i < ARRAY_SIZE(iommu_buses); i++) { + ret =3D bus_for_each_dev(iommu_buses[i], NULL, iter, + _iommu_for_each_dev_cb); + if (ret) + return ret; + } + + return 0; +} +EXPORT_SYMBOL_GPL(iommu_for_each_dev); + #if IS_ENABLED(CONFIG_IOMMUFD_TEST) void iommu_device_unregister_bus(struct iommu_device *iommu, const struct bus_type *bus, @@ -626,8 +658,8 @@ DEFINE_MUTEX(iommu_probe_device_lock); =20 static int __iommu_probe_device(struct device *dev, struct list_head *grou= p_list) { + struct group_device *gdev, *gdev2; struct iommu_group *group; - struct group_device *gdev; int ret; =20 /* @@ -661,6 +693,13 @@ static int __iommu_probe_device(struct device *dev, st= ruct list_head *group_list goto err_put_group; } =20 + for_each_group_device(group, gdev2) { + if (dev_iommu_preserved_state(gdev2->dev)) { + ret =3D -EBUSY; + goto err_free_gdev; + } + } + /* * The gdev must be in the list before calling * iommu_setup_default_domain() @@ -697,6 +736,7 @@ static int __iommu_probe_device(struct device *dev, str= uct list_head *group_list =20 err_remove_gdev: list_del(&gdev->list); +err_free_gdev: __iommu_group_free_device(group, gdev); err_put_group: iommu_deinit_device(dev); @@ -3543,6 +3583,26 @@ bool iommu_group_dma_owner_claimed(struct iommu_grou= p *group) } EXPORT_SYMBOL_GPL(iommu_group_dma_owner_claimed); =20 +/** + * iommu_group_is_singleton() - Query if group contains exactly one device + * @group: The group. + * + * Return: true if group contains exactly 1 device, false otherwise. + */ +bool iommu_group_is_singleton(struct iommu_group *group) +{ + bool ret; + + if (!group) + return false; + + mutex_lock(&group->mutex); + ret =3D (list_count_nodes(&group->devices) =3D=3D 1); + mutex_unlock(&group->mutex); + + return ret; +} + static void iommu_remove_dev_pasid(struct device *dev, ioasid_t pasid, struct iommu_domain *domain) { diff --git a/drivers/iommu/liveupdate.c b/drivers/iommu/liveupdate.c index be542efbb023..5f92a023ea65 100644 --- a/drivers/iommu/liveupdate.c +++ b/drivers/iommu/liveupdate.c @@ -42,6 +42,9 @@ #include #include #include +#include + +#include "iommu-priv.h" =20 #define iommu_max_objs_per_page(_array) \ ((PAGE_SIZE - sizeof(struct iommu_array_hdr_ser)) / sizeof((_array)->obje= cts[0])) @@ -379,3 +382,212 @@ void iommu_unpreserve_domain(struct iommu_domain *dom= ain) liveupdate_flb_put_outgoing(&iommu_flb); } EXPORT_SYMBOL_GPL(iommu_unpreserve_domain); + +static struct iommu_hw_ser *alloc_iommu_hw_ser(struct iommu_flb_obj *flb) +{ + int idx; + + idx =3D alloc_object_ser((void **)&flb->curr_iommu_array, + iommu_max_objs_per_page(flb->curr_iommu_array)); + if (idx < 0) + return ERR_PTR(idx); + + flb->curr_iommu_array->objects[idx].hdr.ref_count =3D 1; + return &flb->curr_iommu_array->objects[idx]; +} + +static int iommu_preserve_locked(struct iommu_device *iommu, + struct iommu_flb_obj *flb_obj) +{ + struct iommu_hw_ser *iommu_hw_ser; + int ret; + + if (!iommu->ops->preserve || !iommu->ops->unpreserve) + return -EOPNOTSUPP; + + lockdep_assert_held(&flb_obj->lock); + if (iommu->outgoing_preserved_state) { + iommu->outgoing_preserved_state->hdr.ref_count++; + return 0; + } + + iommu_hw_ser =3D alloc_iommu_hw_ser(flb_obj); + if (IS_ERR(iommu_hw_ser)) + return PTR_ERR(iommu_hw_ser); + + ret =3D iommu->ops->preserve(iommu, iommu_hw_ser); + if (ret) { + iommu_hw_ser->hdr.flags |=3D IOMMU_SER_FLAG_DELETED; + return ret; + } + + iommu->outgoing_preserved_state =3D iommu_hw_ser; + return ret; +} + +static void iommu_unpreserve_locked(struct iommu_device *iommu, + struct iommu_flb_obj *flb_obj) +{ + struct iommu_hw_ser *iommu_hw_ser =3D iommu->outgoing_preserved_state; + + lockdep_assert_held(&flb_obj->lock); + if (WARN_ON(!iommu_hw_ser)) + return; + + if (WARN_ON_ONCE(!iommu->ops->preserve || + !iommu->ops->unpreserve)) + return; + + iommu_hw_ser->hdr.ref_count--; + if (iommu_hw_ser->hdr.ref_count) + return; + + iommu->outgoing_preserved_state =3D NULL; + iommu->ops->unpreserve(iommu, iommu_hw_ser); + iommu_hw_ser->hdr.flags |=3D IOMMU_SER_FLAG_DELETED; +} + +static struct iommu_device_ser *alloc_iommu_device_ser(struct iommu_flb_ob= j *flb) +{ + int idx; + + idx =3D alloc_object_ser((void **)&flb->curr_device_array, + iommu_max_objs_per_page(flb->curr_device_array)); + if (idx < 0) + return ERR_PTR(idx); + + flb->curr_device_array->objects[idx].hdr.ref_count =3D 1; + return &flb->curr_device_array->objects[idx]; +} + +/** + * iommu_preserve_device() - Preserve device state across live update + * @domain: Associated IOMMU domain + * @dev: Device to preserve + * @dma_owner_token: Token to identify DMA owner of this device + * + * Return: 0 on success, or negative error code. + */ +int iommu_preserve_device(struct iommu_domain *domain, + struct device *dev, u64 dma_owner_token) +{ + struct iommu_device_ser *device_ser; + struct iommu_flb_obj *flb_obj; + struct dev_iommu *iommu; + struct pci_dev *pdev; + int ret; + + if (!dev_is_pci(dev)) + return -EOPNOTSUPP; + + if (!iommu_group_dma_owner_claimed(dev->iommu_group)) + return -EINVAL; + + pdev =3D to_pci_dev(dev); + iommu =3D dev->iommu; + if (!iommu->iommu_dev->ops->preserve_device || + !iommu->iommu_dev->ops->unpreserve_device || + !iommu->iommu_dev->ops->preserve || + !iommu->iommu_dev->ops->unpreserve) + return -EOPNOTSUPP; + + ret =3D liveupdate_flb_get_outgoing(&iommu_flb, (void **)&flb_obj); + if (ret) + return ret; + + mutex_lock(&flb_obj->lock); + if (!domain->preserved_state) { + ret =3D -EINVAL; + goto out_unlock; + } + + if (dev_iommu_preserved_state(dev)) { + ret =3D -EBUSY; + goto out_unlock; + } + + device_ser =3D alloc_iommu_device_ser(flb_obj); + if (IS_ERR(device_ser)) { + ret =3D PTR_ERR(device_ser); + goto out_unlock; + } + + ret =3D iommu_preserve_locked(iommu->iommu_dev, flb_obj); + if (ret) { + device_ser->hdr.flags |=3D IOMMU_SER_FLAG_DELETED; + goto out_unlock; + } + + device_ser->domain_iommu_ser.domain_phys =3D virt_to_phys(domain->preserv= ed_state); + device_ser->domain_iommu_ser.iommu_phys =3D virt_to_phys(iommu->iommu_dev= ->outgoing_preserved_state); + device_ser->devid =3D pci_dev_id(pdev); + device_ser->pci_domain_nr =3D pci_domain_nr(pdev->bus); + device_ser->dma_owner_token =3D dma_owner_token; + + ret =3D iommu->iommu_dev->ops->preserve_device(dev, device_ser); + if (!ret) { + WRITE_ONCE(dev->iommu->device_ser, device_ser); + + /* Validate that no sibling device was added concurrently */ + if (!iommu_group_is_singleton(dev->iommu_group)) { + iommu->iommu_dev->ops->unpreserve_device(dev, device_ser); + WRITE_ONCE(dev->iommu->device_ser, NULL); + ret =3D -EOPNOTSUPP; + } + } + + if (ret) { + device_ser->hdr.flags |=3D IOMMU_SER_FLAG_DELETED; + iommu_unpreserve_locked(iommu->iommu_dev, flb_obj); + goto out_unlock; + } + +out_unlock: + mutex_unlock(&flb_obj->lock); + liveupdate_flb_put_outgoing(&iommu_flb); + return ret; +} +EXPORT_SYMBOL_GPL(iommu_preserve_device); + +/** + * iommu_unpreserve_device() - Unpreserve device state + * @domain: Associated IOMMU domain + * @dev: Device to unpreserve + */ +void iommu_unpreserve_device(struct iommu_domain *domain, struct device *d= ev) +{ + struct iommu_device_ser *iommu_device_ser; + struct iommu_flb_obj *flb_obj; + struct dev_iommu *iommu; + int ret; + + if (!dev_is_pci(dev)) + return; + + if (!iommu_group_dma_owner_claimed(dev->iommu_group)) + return; + + iommu =3D dev->iommu; + if (WARN_ON(!iommu->iommu_dev->ops->unpreserve_device || + !iommu->iommu_dev->ops->unpreserve)) + return; + + ret =3D liveupdate_flb_get_outgoing(&iommu_flb, (void **)&flb_obj); + if (WARN_ON(ret)) + return; + + mutex_lock(&flb_obj->lock); + iommu_device_ser =3D dev_iommu_preserved_state(dev); + if (WARN_ON(!iommu_device_ser)) + goto out_unlock; + + iommu_device_ser->hdr.flags |=3D IOMMU_SER_FLAG_DELETED; + iommu->iommu_dev->ops->unpreserve_device(dev, iommu_device_ser); + WRITE_ONCE(dev->iommu->device_ser, NULL); + + iommu_unpreserve_locked(iommu->iommu_dev, flb_obj); +out_unlock: + mutex_unlock(&flb_obj->lock); + liveupdate_flb_put_outgoing(&iommu_flb); +} +EXPORT_SYMBOL_GPL(iommu_unpreserve_device); diff --git a/include/linux/iommu-liveupdate.h b/include/linux/iommu-liveupd= ate.h index caa9778eee2d..536bea09d064 100644 --- a/include/linux/iommu-liveupdate.h +++ b/include/linux/iommu-liveupdate.h @@ -8,6 +8,7 @@ #ifndef _LINUX_IOMMU_LIVEUPDATE_H #define _LINUX_IOMMU_LIVEUPDATE_H =20 +#include #include #include #include @@ -15,8 +16,43 @@ #ifdef CONFIG_IOMMU_LIVEUPDATE int iommu_liveupdate_register_flb(struct liveupdate_file_handler *handler); void iommu_liveupdate_unregister_flb(struct liveupdate_file_handler *handl= er); + +/** + * dev_iommu_preserved_state() - Get preserved state of a device + * @dev: Target device + * + * Return: Pointer to preserved device state, or NULL if not preserved. + */ +static inline void *dev_iommu_preserved_state(struct device *dev) +{ + struct iommu_device_ser *ser; + + if (!dev->iommu) + return NULL; + + ser =3D READ_ONCE(dev->iommu->device_ser); + if (ser && !(ser->hdr.flags & IOMMU_SER_FLAG_INCOMING)) + return ser; + + return NULL; +} + int iommu_preserve_domain(struct iommu_domain *domain, struct iommu_domain= _ser **ser); void iommu_unpreserve_domain(struct iommu_domain *domain); +int iommu_preserve_device(struct iommu_domain *domain, + struct device *dev, u64 dma_owner_token); +void iommu_unpreserve_device(struct iommu_domain *domain, struct device *d= ev); + +/** + * iommu_preserved_state() - Get preserved state of an IOMMU instance + * @iommu: IOMMU hardware instance + * + * Return: Pointer to preserved state, or NULL if not preserved. + */ +static inline void *iommu_preserved_state(struct iommu_device *iommu) +{ + return iommu->outgoing_preserved_state; +} #else static inline int iommu_liveupdate_register_flb(struct liveupdate_file_han= dler *handler) { @@ -27,6 +63,11 @@ static inline void iommu_liveupdate_unregister_flb(struc= t liveupdate_file_handle { } =20 +static inline void *dev_iommu_preserved_state(struct device *dev) +{ + return NULL; +} + static inline int iommu_preserve_domain(struct iommu_domain *domain, struc= t iommu_domain_ser **ser) { return -EOPNOTSUPP; @@ -35,5 +76,20 @@ static inline int iommu_preserve_domain(struct iommu_dom= ain *domain, struct iomm static inline void iommu_unpreserve_domain(struct iommu_domain *domain) { } + +static inline int iommu_preserve_device(struct iommu_domain *domain, + struct device *dev, u64 dma_owner_token) +{ + return -EOPNOTSUPP; +} + +static inline void iommu_unpreserve_device(struct iommu_domain *domain, st= ruct device *dev) +{ +} + +static inline void *iommu_preserved_state(struct iommu_device *iommu) +{ + return NULL; +} #endif #endif /* _LINUX_IOMMU_LIVEUPDATE_H */ diff --git a/include/linux/iommu.h b/include/linux/iommu.h index 26de40d5a98e..920abd777879 100644 --- a/include/linux/iommu.h +++ b/include/linux/iommu.h @@ -684,6 +684,10 @@ __iommu_copy_struct_to_user(const struct iommu_user_da= ta *dst_data, * resources shared/passed to user space IOMMU instance. Ass= ociate * it with a nesting @parent_domain. It is required for driv= er to * set @viommu->ops pointing to its own viommu_ops + * @preserve_device: Preserve state of a device for liveupdate. + * @unpreserve_device: Unpreserve state that was preserved earlier. + * @preserve: Preserve state of iommu translation hardware for liveupdate. + * @unpreserve: Unpreserve state of iommu that was preserved earlier. * @owner: Driver module providing these ops * @identity_domain: An always available, always attachable identity * translation. @@ -740,6 +744,13 @@ struct iommu_ops { struct iommu_domain *parent_domain, const struct iommu_user_data *user_data); =20 +#ifdef CONFIG_IOMMU_LIVEUPDATE + int (*preserve_device)(struct device *dev, struct iommu_device_ser *devic= e_ser); + void (*unpreserve_device)(struct device *dev, struct iommu_device_ser *de= vice_ser); + int (*preserve)(struct iommu_device *iommu, struct iommu_hw_ser *iommu_se= r); + void (*unpreserve)(struct iommu_device *iommu, struct iommu_hw_ser *iommu= _ser); +#endif + const struct iommu_domain_ops *default_domain_ops; struct module *owner; struct iommu_domain *identity_domain; @@ -827,6 +838,8 @@ struct iommu_domain_ops { * @singleton_group: Used internally for drivers that have only one group * @max_pasids: number of supported PASIDs * @ready: set once iommu_device_register() has completed successfully + * @outgoing_preserved_state: preserved iommu state of outgoing kernel for + * liveupdate. */ struct iommu_device { struct list_head list; @@ -836,6 +849,10 @@ struct iommu_device { struct iommu_group *singleton_group; u32 max_pasids; bool ready; + +#ifdef CONFIG_IOMMU_LIVEUPDATE + struct iommu_hw_ser *outgoing_preserved_state; +#endif }; =20 /** @@ -890,6 +907,9 @@ struct dev_iommu { u32 pci_32bit_workaround:1; u32 require_direct:1; u32 shadow_on_flush:1; +#ifdef CONFIG_IOMMU_LIVEUPDATE + struct iommu_device_ser *device_ser; +#endif }; =20 int iommu_device_register(struct iommu_device *iommu, @@ -1206,6 +1226,28 @@ static inline void *dev_iommu_priv_get(struct device= *dev) =20 void dev_iommu_priv_set(struct device *dev, void *priv); =20 +/** + * typedef iommu_dev_iter_fn - Callback for iterating IOMMU attached devic= es + * @dev: Attached device + * @iommu: IOMMU instance + * @arg: Private argument passed to iterator + * + * Return: 0 on success, or negative error code. + */ +typedef int (*iommu_dev_iter_fn)(struct device *dev, + struct iommu_device *iommu, void *arg); + +/** + * struct iommu_dev_iter - Iterator for devices attached to an IOMMU + */ +struct iommu_dev_iter { + struct iommu_device *iommu; + iommu_dev_iter_fn fn; + void *arg; +}; + +int iommu_for_each_dev(struct iommu_dev_iter *iter); + extern struct mutex iommu_probe_device_lock; int iommu_probe_device(struct device *dev); =20 --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EE18D2609E3 for ; Mon, 21 Sep 2026 00:48:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951737; cv=none; b=MdiRnEwT7bThK1COcsuGtDZuuY2hXO+NWp9TLCsKJPqOHPtMBQz9AfVIgmi5QhpVlj1KbrAoQprzxCH5nJlAnyrnGfPureSlwEClRSanS6328Aivj5ZoaNXNnXXaP9HRiFMuTqainCMCoibPZOa+4edeKY5jkid9/bTuw2tu4h4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951737; c=relaxed/simple; bh=3BUJ6LTwbjZSrKvTIK6/xvZavv/JN3y/YbZwrJiWz6Y=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=KQKZAVdD9tC6TYCe/kmNwlkaPAo71dN2pRTf/WuizSr1MCf1+fnCy/ycelIo1iBy/iR7CTYvPcI+f5Wv71fpM6s74mdc8Fu/u7b5nsX34N/HJL1UvhQc/oVeIsP1HFXpiElb0Zhf8Fejyx9441ckS+o3fb+yJQjmp5jFdh67pCc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=p8B91WZO; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="p8B91WZO" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2cec4226c70so31051065ad.1 for ; Sun, 20 Sep 2026 17:48:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951727; x=1790556527; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=oMgDfdy4n3gI7TbnZlQ1tz32BsddhUX0XW4jh9hXBH0=; b=p8B91WZOw3RVCs/xpUteQcHiEiR0t29HP4A0Z4ExDPqOFmBi0AzwJjgn3zGNGySLxV AG4j+R7r/s6FxWm3Polb8oydZID++eryKw+lfczFrufBkAJGPPXXLtKcRof+hPRUKvCI BIdZMeHUw6qAH8KkjNxu9cRp21Aa+HzFCSAtyq8U6Tk/raRNwQJDj4G11pj788WcNFI8 XJghqreSpr2wtcFuS+8Vqu1WE43dyXQFJoKgYlPGUftPF4eK9zvZaZxqZ8Yh0AP4Dga5 rDx4eIh4Xo8PfbhSSWQfIMAOHVmyMO2XXa7ugY5W9UPDSCk73fuvVcBvJj8renkvf3+Q Hfxw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951727; x=1790556527; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=oMgDfdy4n3gI7TbnZlQ1tz32BsddhUX0XW4jh9hXBH0=; b=0BxsujALJueuoRFQN7AX2snWg2rUH9jSYZuUluLLay8jv3L12dn8Mnb8z57Y71mFZK yRxBg0CCGDlVp9+rBWmaN9aogb8I3VEOlZARp34uuKNlB7B7YXeLr/40bVjxkT9w/kS/ JJYCHK+tyKCF+Msc9XTvDC1n+Cy02IHdlcrL0EoNi51kehKfSt/I5NC301dLAsA7JaXq c7Ba3QjwaZJu0th+Q7mvCfhRzbzTonTVfq1ob2Efn+/uWneBFG3jTZj+8MXVr6UKq8WX GKuaxgaE62ATiIyCZdVW6Lkf7QrA6HOgXCEB7Ln2C4+dGAawAxsOqqDofhvL5Mghe9lb boKg== X-Forwarded-Encrypted: i=1; AKwUvByEWhBYLDrHM2VGuCPe42HmmV35cm+5afCnklWlS9M7g0Kar9uLcHt77BBARo/2JjpTrs+p7g+0Yd+zDaM=@vger.kernel.org X-Gm-Message-State: AFuF++n0ajwZzvegr6emJqQG039C7KngEQt8AVwemWk4SjKcXBw7AaID 9lkr6zRTgYPn0/OMCJnMkwMu7iDCmY63rZjRLUiDGeWs58qk0kPKZoV1ufrKMEaJ/c1PqF7xqdg 1GV2vTWr62tNBig== X-Received: from pldy20.prod.google.com ([2002:a17:902:cad4:b0:2df:3fbc:5137]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:390d:b0:2dd:ad73:c98a with SMTP id d9443c01a7336-2ddb1bd6f42mr131665945ad.34.1789951726957; Sun, 20 Sep 2026 17:48:46 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:23 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-8-skhawaja@google.com> Subject: [PATCH v5 07/18] iommu/vt-d: Implement device and iommu preserve/unpreserve ops From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add implementation of the device and iommu preservation in a separate file. Also set the device and iommu preserve/unpreserve ops in the struct iommu_ops. Preservation is refused for devices where ATS is supported. Carrying an active ATS configuration across kexec would require adopting the ATS state in the PCI core and in the context entries, and handling the disable-to-enable, enable-to-disable and STU mismatch cases for a live device. Note that the preserved iommu root table, context table and device pasid table will need cleanup as devices do not go through the release flow during kexec. So the stale entries need to be removed during shutdown. This is done in the next commit. Signed-off-by: Samiullah Khawaja --- MAINTAINERS | 8 ++ drivers/iommu/intel/Makefile | 1 + drivers/iommu/intel/iommu.c | 9 +- drivers/iommu/intel/iommu.h | 13 ++ drivers/iommu/intel/liveupdate.c | 232 +++++++++++++++++++++++++++++++ include/linux/kho/abi/iommu.h | 24 ++++ 6 files changed, 285 insertions(+), 2 deletions(-) create mode 100644 drivers/iommu/intel/liveupdate.c diff --git a/MAINTAINERS b/MAINTAINERS index 2114b50412ee..ae6df95dc398 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -13220,6 +13220,14 @@ S: Supported T: git git://git.kernel.org/pub/scm/linux/kernel/git/iommu/linux.git F: drivers/iommu/intel/ =20 +INTEL IOMMU LIVEUPDATE (VT-d) +M: Samiullah Khawaja +M: Lu Baolu +L: iommu@lists.linux.dev +S: Maintained +T: git git://git.kernel.org/pub/scm/linux/kernel/git/iommu/linux.git +F: drivers/iommu/intel/liveupdate.c + INTEL IPU3 CSI-2 CIO2 DRIVER M: Yong Zhi M: Sakari Ailus diff --git a/drivers/iommu/intel/Makefile b/drivers/iommu/intel/Makefile index ada651c4a01b..d38fc101bc35 100644 --- a/drivers/iommu/intel/Makefile +++ b/drivers/iommu/intel/Makefile @@ -6,3 +6,4 @@ obj-$(CONFIG_INTEL_IOMMU_DEBUGFS) +=3D debugfs.o obj-$(CONFIG_INTEL_IOMMU_SVM) +=3D svm.o obj-$(CONFIG_IRQ_REMAP) +=3D irq_remapping.o obj-$(CONFIG_INTEL_IOMMU_PERF_EVENTS) +=3D perfmon.o +obj-$(CONFIG_IOMMU_LIVEUPDATE) +=3D liveupdate.o diff --git a/drivers/iommu/intel/iommu.c b/drivers/iommu/intel/iommu.c index 2e3b3ab216f8..4bfc2f173010 100644 --- a/drivers/iommu/intel/iommu.c +++ b/drivers/iommu/intel/iommu.c @@ -16,6 +16,7 @@ #include #include #include +#include #include #include #include @@ -58,8 +59,6 @@ static int rwbf_quirk; */ int intel_iommu_tboot_noforce; =20 -#define ROOT_ENTRY_NR (VTD_PAGE_SIZE/sizeof(struct root_entry)) - /* * Take a root_entry and return the Lower Context Table Pointer (LCTP) * if marked present. @@ -3961,6 +3960,12 @@ const struct iommu_ops intel_iommu_ops =3D { .is_attach_deferred =3D intel_iommu_is_attach_deferred, .def_domain_type =3D device_def_domain_type, .page_response =3D intel_iommu_page_response, +#ifdef CONFIG_IOMMU_LIVEUPDATE + .preserve_device =3D intel_iommu_preserve_device, + .unpreserve_device =3D intel_iommu_unpreserve_device, + .preserve =3D intel_iommu_preserve, + .unpreserve =3D intel_iommu_unpreserve, +#endif }; =20 static void quirk_iommu_igfx(struct pci_dev *dev) diff --git a/drivers/iommu/intel/iommu.h b/drivers/iommu/intel/iommu.h index 23dbe6c24439..785057236e7c 100644 --- a/drivers/iommu/intel/iommu.h +++ b/drivers/iommu/intel/iommu.h @@ -552,6 +552,8 @@ struct root_entry { u64 hi; }; =20 +#define ROOT_ENTRY_NR (VTD_PAGE_SIZE / sizeof(struct root_entry)) + /* * low 64 bits: * 0: present @@ -1296,6 +1298,17 @@ static inline int iopf_for_domain_replace(struct iom= mu_domain *new, return 0; } =20 +#ifdef CONFIG_IOMMU_LIVEUPDATE +int intel_iommu_preserve_device(struct device *dev, + struct iommu_device_ser *device_ser); +void intel_iommu_unpreserve_device(struct device *dev, + struct iommu_device_ser *device_ser); +int intel_iommu_preserve(struct iommu_device *iommu, + struct iommu_hw_ser *iommu_ser); +void intel_iommu_unpreserve(struct iommu_device *iommu, + struct iommu_hw_ser *iommu_ser); +#endif + #ifdef CONFIG_INTEL_IOMMU_SVM void intel_svm_check(struct intel_iommu *iommu); struct iommu_domain *intel_svm_domain_alloc(struct device *dev, diff --git a/drivers/iommu/intel/liveupdate.c b/drivers/iommu/intel/liveupd= ate.c new file mode 100644 index 000000000000..0837bb889fed --- /dev/null +++ b/drivers/iommu/intel/liveupdate.c @@ -0,0 +1,232 @@ +// SPDX-License-Identifier: GPL-2.0-only + +/* + * Copyright (C) 2026, Google LLC + * Author: Samiullah Khawaja + */ + +#define pr_fmt(fmt) "DMAR: liveupdate: " fmt + +#include +#include +#include +#include +#include + +#include "iommu.h" +#include "../iommu-pages.h" + +/* 2 tables per bus in scalable mode with upper table at odd bit */ +#define CONTEXT_TABLE_PRESERVED_BIT(bus, devfn) (((bus) << 1) + ((devfn) >= > 7)) +static bool is_context_table_preserved(struct intel_iommu *iommu, + struct iommu_hw_ser *ser, + u8 bus, u8 devfn) +{ + return test_bit(CONTEXT_TABLE_PRESERVED_BIT(bus, devfn), + (unsigned long *)&ser->intel.context_tables_bitmap[0]); +} + +static void unpreserve_context_table(struct intel_iommu *iommu, + struct iommu_hw_ser *ser, + u8 bus, u8 devfn) +{ + struct context_entry *context; + + /* + * In the Intel IOMMU driver, context tables are never freed once they + * are allocated during runtime, as they can be shared across multiple + * devices. So taking the iommu lock here to protect against concurrent + * allocations inside iommu_context_addr() should be enough. Once the + * address is read, it is safe to use it without holding the lock. + */ + spin_lock(&iommu->lock); + context =3D iommu_context_addr(iommu, bus, devfn, 0); + spin_unlock(&iommu->lock); + if (context && is_context_table_preserved(iommu, ser, bus, devfn)) { + iommu_unpreserve_pages(context); + clear_bit(CONTEXT_TABLE_PRESERVED_BIT(bus, devfn), + (unsigned long *)&ser->intel.context_tables_bitmap[0]); + } +} + +static int preserve_context_table(struct intel_iommu *iommu, + struct iommu_hw_ser *ser, + u8 bus, u8 devfn) +{ + struct context_entry *context; + int ret; + + spin_lock(&iommu->lock); + context =3D iommu_context_addr(iommu, bus, devfn, 0); + spin_unlock(&iommu->lock); + + /* + * Intel IOMMU context tables are never freed by the driver once + * allocated. It is safe to access the context pointer outside of the + * iommu->lock. + */ + if (context && !is_context_table_preserved(iommu, ser, bus, devfn)) { + ret =3D iommu_preserve_pages(context); + if (ret) + return ret; + + set_bit(CONTEXT_TABLE_PRESERVED_BIT(bus, devfn), + (unsigned long *)&ser->intel.context_tables_bitmap[0]); + } + + return 0; +} + +static void unpreserve_iommu_context_tables(struct intel_iommu *iommu, + struct iommu_hw_ser *ser) +{ + int i; + + for (i =3D 0; i < ROOT_ENTRY_NR; i++) { + unpreserve_context_table(iommu, ser, i, 0); + + if (!sm_supported(iommu)) + continue; + + unpreserve_context_table(iommu, ser, i, 0x80); + } +} + +static int preserve_iommu_context_tables(struct device_domain_info *info) +{ + struct iommu_hw_ser *iommu_ser; + struct intel_iommu *iommu; + int ret; + int i; + + /* IOMMU for this device should already preserved.*/ + iommu =3D info->iommu; + iommu_ser =3D iommu_preserved_state(&iommu->iommu); + if (!iommu_ser) + return -EINVAL; + + /* + * We could do preservation of context tables only for the bus of this + * device, but these devices can have PCI aliases, so context tables for + * those will also require preservation. Also unpreserve would require + * some kind of refcounting where the context table will only be + * unpreserved when the last device associated with it is unpreserved. + * + * This introduces unnecessary complication with minimum benefits as the + * unpreserved context tables will probably be recreated by the next + * kernel as these are all active devices. We follow simpler approach by + * just preserving the currently active context tables. + */ + for (i =3D 0; i < ROOT_ENTRY_NR; i++) { + ret =3D preserve_context_table(iommu, iommu_ser, i, 0); + if (ret) + return ret; + + if (!sm_supported(iommu)) + continue; + + ret =3D preserve_context_table(iommu, iommu_ser, i, 0x80); + if (ret) + return ret; + } + + return 0; +} + +/** + * intel_iommu_preserve_device() - Intel IOMMU callback to preserve device= state + * @dev: Target device + * @device_ser: Struct to populate with serialized device state + * + * Return: 0 on success, or negative error code. + */ +int intel_iommu_preserve_device(struct device *dev, + struct iommu_device_ser *device_ser) +{ + struct device_domain_info *info =3D dev_iommu_priv_get(dev); + int ret; + + if (!dev_is_pci(dev)) { + dev_err(dev, "Cannot preserve non-PCI device\n"); + return -EOPNOTSUPP; + } + + if (dev_is_real_dma_subdevice(dev)) + return -EOPNOTSUPP; + + if (!info || !info->domain) + return -EINVAL; + + if (info->ats_supported) + return -EOPNOTSUPP; + + ret =3D preserve_iommu_context_tables(info); + if (ret) + return ret; + + device_ser->domain_iommu_ser.attachment_id =3D domain_id_iommu(info->doma= in, + info->iommu); + return 0; +} + +/** + * intel_iommu_unpreserve_device() - Intel IOMMU callback to unpreserve de= vice state + * @dev: Target device + * @device_ser: Struct containing serialized device state + */ +void intel_iommu_unpreserve_device(struct device *dev, + struct iommu_device_ser *device_ser) +{ + /* + * The context tables preserved during device preservation, in the + * preserve_device() callback, might be shared with other devices, so + * those are unpreserved in the iommu unpreserve() callback. So this + * callback is kept empty. + * + * Once device PASID tables are preserved, the unpreservation of PASID + * tables will be added here. + */ +} + +/** + * intel_iommu_preserve() - Intel IOMMU callback to preserve hardware state + * @iommu_dev: Generic IOMMU device handle + * @ser: Struct to populate with serialized hardware state + * + * Return: 0 on success, or negative error code. + */ +int intel_iommu_preserve(struct iommu_device *iommu_dev, + struct iommu_hw_ser *ser) +{ + struct intel_iommu *iommu; + int ret; + + iommu =3D container_of(iommu_dev, struct intel_iommu, iommu); + + ret =3D iommu_preserve_pages(iommu->root_entry); + if (ret) + return ret; + + ser->intel.phys_addr =3D iommu->reg_phys; + ser->intel.root_table =3D __pa(iommu->root_entry); + ser->type =3D IOMMU_INTEL; + ser->token =3D ser->intel.phys_addr; + + return 0; +} + +/** + * intel_iommu_unpreserve() - Intel IOMMU callback to unpreserve hardware = state + * @iommu_dev: Generic IOMMU device handle + * @ser: Struct containing serialized hardware state + */ +void intel_iommu_unpreserve(struct iommu_device *iommu_dev, + struct iommu_hw_ser *ser) +{ + struct intel_iommu *iommu; + + iommu =3D container_of(iommu_dev, struct intel_iommu, iommu); + + unpreserve_iommu_context_tables(iommu, ser); + iommu_unpreserve_pages(iommu->root_entry); +} diff --git a/include/linux/kho/abi/iommu.h b/include/linux/kho/abi/iommu.h index aa42085409e5..5aaa29da6832 100644 --- a/include/linux/kho/abi/iommu.h +++ b/include/linux/kho/abi/iommu.h @@ -81,6 +81,7 @@ */ enum iommu_type_ser { IOMMU_INVALID, + IOMMU_INTEL, }; =20 #define IOMMU_SER_FLAG_DELETED (1 << 0) @@ -144,16 +145,39 @@ struct iommu_device_ser { struct iommu_dev_map_ser domain_iommu_ser; } __packed; =20 +/* There are maximum 256 buses, so maximum 512 context tables */ +#define VTD_PRESERVED_BITMAP_LONGS DIV_ROUND_UP(512, BITS_PER_LONG_LONG) + +/** + * struct iommu_intel_ser - Serialized state of an Intel IOMMU instance + * @restored: Whether IOMMU state is restored + * @phys_addr: Physical address of the IOMMU register base + * @root_table: Physical address of the root entry table + * @context_tables_bitmap: Bitmap representing the context tables that are + * preserved. + */ +struct iommu_intel_ser { + u8 restored; + u8 padding[7]; + u64 phys_addr; + u64 root_table; + u64 context_tables_bitmap[VTD_PRESERVED_BITMAP_LONGS]; +}; + /** * struct iommu_hw_ser - Serialized state of an IOMMU instance * @hdr: Common object header * @token: Unique token for the IOMMU * @type: IOMMU type serialized state belongs to + * @intel: Intel specific serialization data */ struct iommu_hw_ser { struct iommu_hdr_ser hdr; u64 token; u64 type; + union { + struct iommu_intel_ser intel; + }; } __packed; =20 /** --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pf1-f200.google.com (mail-pf1-f200.google.com [209.85.210.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8AC312E7622 for ; Mon, 21 Sep 2026 00:48:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.200 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951733; cv=none; b=pjTAug6I9NiRDHYtfHSYZyQI5y/MRjmK0jpol67wSQIfsr99rkqZDQ2DHukvTCZkgQXXYCN0eYUwXrJ0tUtR/FBq50KnPCuqitkLeLt1AL7vK3m54Qex62fBUoixql0+EfkvDcfXOmluzeluhFiY2pfGlmC8U87gslx64xQV7/E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951733; c=relaxed/simple; bh=ArO2m6pL7uO2PhZtrHs0jVCCkvGqo8M4uLzaVIqqEnM=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=oQmrhEdd1OQxShmdB9JZiQ+lop7KQwzTACBz9qw3IILybAJGfTycCmysniT7n5NhKlqU3ZOGYuHBOiHJniWidSy/7lBSpRpy+2tAe6Xdho4G++oXM3RR5SrAyKg+otP+KWghko5jWnuh7eTJ+ZQjvH8vlyEej32G+QuZjk0pQOI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=vVNC+Ez4; arc=none smtp.client-ip=209.85.210.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="vVNC+Ez4" Received: by mail-pf1-f200.google.com with SMTP id d2e1a72fcca58-86a74698972so2851517b3a.0 for ; Sun, 20 Sep 2026 17:48:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951728; x=1790556528; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=3DQXvWtZKOwtXJHttEiyqacBwQBVqgRl20BjrrSpLfU=; b=vVNC+Ez4r1IWqmRhCC3lN4gtFphKJAeVwbu7mfgIY8JpsKsyPMqIZPhqTiBn//4910 iXyIvTgfV7RLJbK/YlRPrZGt0Li9rihOzdm3Ix3Fz7Z5V4L6YMkAOs1LMjzk+aalZ71R u0SV0RSLEiiNQVaYa6TzvQhrIDwEJr9davEAPw2TxjcR6PoxFqKpQJyZjsjHteQhVyv6 SoSpBjNLfw750ILBJAEdhmgE4rI5W2Or2fwxkqS7s9s5qj7U+mDRPTWpAMX1eohLXMjn s73i0TI8aiqJCxQ11eQAVLzyhMNne2Y3Je5DDaYqqujd5qk4F8mHF7Iu3Htqtq/gQzE8 atsA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951728; x=1790556528; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=3DQXvWtZKOwtXJHttEiyqacBwQBVqgRl20BjrrSpLfU=; b=Mt5jIybZAyA+ahjZO6tNZ/Yf++r0XVrhW/2XVIX6li6VBETofEkrswu6UWFjjIRvgS sDbyYfB1dxbAl/D8+ekjH1uHHj2hBFKvfXNyx2Nt2wAUFNEmqmOd1lHIvHvW7z1Kk5sT WoLvxjk7viUSOLpRKIHhUx/lnLum1eQcRUy+gQiFrzOy/4ii+0wwfi63QFn2Oy8/tjed xuHNoyBU8Kgq4S7Y/v+5j5tAjeZoLzIlgupw1sagLqfy8rWfrNo+41LQhlI38xcJVYLO O5dEIMmItdYKStrjdTUiM82qSyC9Eu5WEp4SIbWCnRKBo0z28SNrz6GGtbsjozeQeQyB XGvA== X-Forwarded-Encrypted: i=1; AKwUvBw96jyLH46NNpte8j4llUVZCLT0CbbKaAjFsOOnMRjhLN8cF7jyMzqqu12BMXf5o+QM8x2Ql4faziCbaZs=@vger.kernel.org X-Gm-Message-State: AFuF++kkOFP/DnbIV8ETEWDZE+QG0WlseP7HrCteQI4oi4gaTTmKrPaM UiJ4ok3rCkg4rIe2ennaTkNEkqgUXvtmOyRivOVlvdcc9x2zfJ7KftvFlubn3mikuxX6bldEAZx Z60w5IV0PlMHkOg== X-Received: from pgvc11.prod.google.com ([2002:a65:618b:0:b0:cc1:c12a:92cf]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:464f:b0:84e:e741:174f with SMTP id d2e1a72fcca58-874db6f7494mr12164945b3a.7.1789951727653; Sun, 20 Sep 2026 17:48:47 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:24 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-9-skhawaja@google.com> Subject: [PATCH v5 08/18] iommu/vt-d: Clear unpreserved context entries during shutdown From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" During normal shutdown the iommu translation is disabled. Since the root table is preserved during live update, it needs to be cleaned up and the context entries of the unpreserved devices and root entries for the unpreserved context tables need to be cleared. This is required because during kexec reboot and shutdown the devices do not go through the release flow, so there might be stale entries in the root table and context tables. Also note that new unpreserved context tables for unpreserved devices might have been added after preservation, so the root table entries for unpreserved context tables also need to removed. Signed-off-by: Samiullah Khawaja --- drivers/iommu/intel/iommu.c | 15 +++- drivers/iommu/intel/iommu.h | 5 ++ drivers/iommu/intel/liveupdate.c | 139 +++++++++++++++++++++++++++++++ 3 files changed, 157 insertions(+), 2 deletions(-) diff --git a/drivers/iommu/intel/iommu.c b/drivers/iommu/intel/iommu.c index 4bfc2f173010..474c926172c5 100644 --- a/drivers/iommu/intel/iommu.c +++ b/drivers/iommu/intel/iommu.c @@ -1852,6 +1852,14 @@ static int iommu_suspend(void *data) =20 iommu_flush_all(); =20 + /* + * Note that IOMMU suspend doesn't affect live update. The state + * preserved during live update is not released and remains valid during + * suspend and reused during IOMMU resume. + * + * Also note deployment of suspend/resume and live updated use case + * should be mostly mutually exclusive. + */ for_each_active_iommu(iommu, drhd) { iommu_disable_translation(iommu); =20 @@ -2397,8 +2405,11 @@ void intel_iommu_shutdown(void) /* Disable PMRs explicitly here. */ iommu_disable_protect_mem_regions(iommu); =20 - /* Make sure the IOMMUs are switched off */ - iommu_disable_translation(iommu); + /* Make sure the IOMMUs are switched off if not preserved. */ + if (iommu_preserved_state(&iommu->iommu)) + clear_unpreserved_context_entries(iommu); + else + iommu_disable_translation(iommu); } } =20 diff --git a/drivers/iommu/intel/iommu.h b/drivers/iommu/intel/iommu.h index 785057236e7c..4feb5bd76b18 100644 --- a/drivers/iommu/intel/iommu.h +++ b/drivers/iommu/intel/iommu.h @@ -1307,6 +1307,11 @@ int intel_iommu_preserve(struct iommu_device *iommu, struct iommu_hw_ser *iommu_ser); void intel_iommu_unpreserve(struct iommu_device *iommu, struct iommu_hw_ser *iommu_ser); +void clear_unpreserved_context_entries(struct intel_iommu *iommu); +#else +static inline void clear_unpreserved_context_entries(struct intel_iommu *i= ommu) +{ +} #endif =20 #ifdef CONFIG_INTEL_IOMMU_SVM diff --git a/drivers/iommu/intel/liveupdate.c b/drivers/iommu/intel/liveupd= ate.c index 0837bb889fed..501dc0e9cc0a 100644 --- a/drivers/iommu/intel/liveupdate.c +++ b/drivers/iommu/intel/liveupdate.c @@ -77,6 +77,145 @@ static int preserve_context_table(struct intel_iommu *i= ommu, return 0; } =20 +static void clear_unpreserved_context_root_entries(struct intel_iommu *iom= mu, + struct iommu_hw_ser *ser) +{ + struct root_entry *root; + int i; + + /* + * Individual invalidations for each context table removal are not + * needed as we issue global invalidations later. + */ + for (i =3D 0; i < ROOT_ENTRY_NR; i++) { + root =3D &iommu->root_entry[i]; + + if (!is_context_table_preserved(iommu, ser, i, 0) && (root->lo & 1)) { + root->lo =3D 0; + __iommu_flush_cache(iommu, + &root->lo, + sizeof(root->lo)); + } + + if (!sm_supported(iommu)) + continue; + + if (!is_context_table_preserved(iommu, ser, i, 0x80) && (root->hi & 1)) { + root->hi =3D 0; + __iommu_flush_cache(iommu, + &root->hi, + sizeof(root->hi)); + } + } +} + +static void clear_unpreserved_context(struct device_domain_info *info, u8 = bus, u8 devfn) +{ + struct context_entry *context; + + /* + * This cleanup is done during shutdown, so it should be fine to only + * clear the entries here and issue one global invalidation later to + * invalidate all cleared entries. + * + * Note that the device IOTLB invalidation for unpreserved devices is + * skipped this way, but that should not be needed as the devices are + * quiesced at this point. This should improve the performance of the + * cleanup process and avoids any invalidation timeouts because drivers + * might have moved devices to D3 state. + * + * The iommu lock is not needed as the context tables are never removed + * and the new ones are not added during shutdown. + */ + context =3D iommu_context_addr(info->iommu, bus, devfn, 0); + if (context) { + /* + * Individual invalidations are not needed as we issue global + * invalidations later once all the cleanups are done. + */ + context_clear_present(context); + __iommu_flush_cache(info->iommu, context, sizeof(*context)); + context_clear_entry(context); + __iommu_flush_cache(info->iommu, context, sizeof(*context)); + } +} + +static int clear_unpreserved_alias_cb(struct pci_dev *pdev, u16 alias, voi= d *data) +{ + struct device_domain_info *info =3D data; + + clear_unpreserved_context(info, PCI_BUS_NUM(alias), alias & 0xff); + return 0; +} + +static int clear_unpreserve_context_entry_fn(struct device *dev, + struct iommu_device *iommu_dev, + void *arg) +{ + struct device_domain_info *info; + + info =3D dev_iommu_priv_get(dev); + if (!info) + return 0; + + if (dev_iommu_preserved_state(dev)) + return 0; + + if (dev_is_pci(dev)) + pci_for_each_dma_alias(to_pci_dev(dev), + clear_unpreserved_alias_cb, info); + else + clear_unpreserved_context(info, info->bus, info->devfn); + + return 0; +} + +/** + * clear_unpreserved_context_entries() - Clear context entries for unprese= rved devices + * @iommu: Target IOMMU + * + * Clear the context entries of unpreserved devices during shutdown before= kexec. + */ +void clear_unpreserved_context_entries(struct intel_iommu *iommu) +{ + struct iommu_dev_iter iter =3D { + .fn =3D clear_unpreserve_context_entry_fn, + .iommu =3D &iommu->iommu, + .arg =3D NULL, + + }; + + /* + * Clear context entries for unpreserved devices because during kexec + * reboot and shutdown the devices do not go through the release flow, + * so there might be stale entries in the root table and context tables. + * + * Note that the error can be ignored as the iterator function does not + * fail. + */ + iommu_for_each_dev(&iter); + + /* + * New unpreserved context tables for unpreserved devices might have + * been added after preservation, so the root table entries for + * unpreserved context tables need to be removed. + */ + clear_unpreserved_context_root_entries(iommu, + iommu_preserved_state(&iommu->iommu)); + + /* + * Some devices might not have teardown/detached properly depending on + * whether a proper device remove is done before kexec is triggered. + * Also unpreserved context tables and entries are removed during + * shutdown. So issue global invalidations to remove references to + * unpreserved tables and entries. + */ + iommu->flush.flush_context(iommu, 0, 0, 0, DMA_CCMD_GLOBAL_INVL); + if (sm_supported(iommu)) + qi_flush_pasid_cache(iommu, 0, QI_PC_GLOBAL, 0); + iommu->flush.flush_iotlb(iommu, 0, 0, 0, DMA_TLB_GLOBAL_FLUSH); +} + static void unpreserve_iommu_context_tables(struct intel_iommu *iommu, struct iommu_hw_ser *ser) { --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3A6282E2286 for ; Mon, 21 Sep 2026 00:48:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951735; cv=none; b=sJfciU8VpFxyHkWQcl5FgQzXZ5KjhflDNyp5M9xLuKxYSvr15XOblRilfk8XkhBQj65snX4ag21UcX/bvIU8LYmd69beXMMoNZN74LrGrReBl9mZZoHnfuE3PkTkjQZz/8jPliQJjCA5ISHQ9kNbBDcjGbXIhtJPZGiap/taE08= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951735; c=relaxed/simple; bh=Jv3iaex2JDtz3JgCAWlDmzXnCKQbniBZ71nQTC+bZ58=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=hQ3CHT4VwVFGwY+Gt0hUBraw3ZBuv3yv8UiFixvtQlSL9aYnEn2JBOCIqUQXoKoNDpO7Be7qiG9Poe6cRnwSjSBo68Rdnea3jtVQX1EyZmU1BU25x+Y67pdLmT5ikhk1vjABKJLPcHGJ5edTYCRS30MiGA9k6NOUHQFGOstyuYE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=DzxwO3rs; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="DzxwO3rs" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2d9336581a2so45130675ad.3 for ; Sun, 20 Sep 2026 17:48:49 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951728; x=1790556528; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=O5GxCZRn4tT/d2wdhRXzFHtH06EaWqTlMJbz+k/JmSo=; b=DzxwO3rs6fhBr2yKVbA5p9nNg6+NnQiBxtZz4w7ewqVc43Y9+Q67zU8cVgJpHaPtbS 3iQ9YyL6ZCUd/fNkwB2/SlbSFxWGTeuwnM+88nrQpgZ6hWpazDdSpl4Wy/rkEoNik0yn WYvyoW0ykthCbjuDDUJS2Rbjte1Ma/HV3ykOVmWOBAnD9/wsFQBF573Eaqbtj8QVUvpg xpWccbdG/1D7ItdY7WkZJKPXxuNtnF/xBDwdL50z4UKFoaX9XiR66RuvQgXTMM2KVGxN C3gm8WGA+DSn5YlgqgT2sG92cQhmnarCatRtFD404bHGHalN6Ix7encUsvC5MM0NkBMR 5hvQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951728; x=1790556528; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=O5GxCZRn4tT/d2wdhRXzFHtH06EaWqTlMJbz+k/JmSo=; b=SKXyUjtp4O1tBc53njp7xJgL2JqvvRfiuxS9GdpYykH1eLwYd9eFcSpsHc/7mWW8qu jbft3+2QFq18aptX8pZsUE/g6yrW8e2a6M70MVSlrRoAuz3Kxc4txdjeLLbYJyYytx+/ WS/+OsB27Wk0Mxh6qsKAU2FvbqTqLHvIBipxe7juzEWrtJtdtFDunsI/x2uJVpi+CMOu E/2rSFOpMhzbhGfIfobMVl1m+likGUN1f32Ju+OwFsT9yAW684yAfJwYcdMqClnG3ltf COC5lTX5YbMRMUmpvg7I4kYrJ06pLQiRJ0vh9qDDNrMZM1/aSB8q1tmXG4I4geKGDvwN hjAg== X-Forwarded-Encrypted: i=1; AKwUvBwPEWKZ2j7vIetehohkkLiJfJb5PsuLnt583y+T9DvhjEMgGKCc3XLoXKzecO93PSWWAFki7Kwe3IH/4QA=@vger.kernel.org X-Gm-Message-State: AFuF++mVPS4HA3q1ShEAlJA2NJb48iEQuhR+p24qLoc1k11q7adRXc8i ofapfRPewAfkWK8Aescw+4Qdc/2kiozJ1OcQHt6uHoorRDEgvQk0KOfis54u2j1yM9raffkNsJa SwvhPsdK/TX3jHQ== X-Received: from plas21.prod.google.com ([2002:a17:903:2015:b0:2df:4d76:bbc7]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:22cc:b0:2df:425b:3082 with SMTP id d9443c01a7336-2df425b3273mr28931155ad.57.1789951728436; Sun, 20 Sep 2026 17:48:48 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:25 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-10-skhawaja@google.com> Subject: [PATCH v5 09/18] iommu: Add APIs to get iommu and device preserved state From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Pranjal Shrivastava , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The preserved state of the device and IOMMU needs to be fetched during shutdown and boot in the next kernel. Add APIs that can be used to fetch the preserved state of a device and IOMMU. The APIs will only be used during shutdown and after liveupdate so no locking needed. Reviewed-by: Pranjal Shrivastava Signed-off-by: Samiullah Khawaja --- drivers/iommu/liveupdate.c | 125 +++++++++++++++++++++++++++++++ include/linux/iommu-liveupdate.h | 45 +++++++++++ 2 files changed, 170 insertions(+) diff --git a/drivers/iommu/liveupdate.c b/drivers/iommu/liveupdate.c index 5f92a023ea65..ff76a23d9583 100644 --- a/drivers/iommu/liveupdate.c +++ b/drivers/iommu/liveupdate.c @@ -49,6 +49,17 @@ #define iommu_max_objs_per_page(_array) \ ((PAGE_SIZE - sizeof(struct iommu_array_hdr_ser)) / sizeof((_array)->obje= cts[0])) =20 +#define iommu_liveupdate_for_each_obj(_arr, _obj, _idx) \ + for ((_idx) =3D 0, (_obj) =3D (_arr)->objects; \ + (_idx) < (_arr)->hdr.nr_objects; (_idx)++, (_obj)++) \ + if (((_obj)->hdr.flags & IOMMU_SER_FLAG_DELETED)) \ + continue; \ + else + +#define iommu_liveupdate_for_each_arr(_arr) \ + for (; (_arr); (_arr) =3D (_arr)->hdr.next_array_phys ? \ + phys_to_virt((_arr)->hdr.next_array_phys) : NULL) + struct iommu_flb_obj { struct mutex lock; struct iommu_flb_ser *ser; @@ -259,6 +270,120 @@ void iommu_liveupdate_unregister_flb(struct liveupdat= e_file_handler *handler) } EXPORT_SYMBOL(iommu_liveupdate_unregister_flb); =20 +/* + * iommu_liveupdate_flb_get_incoming() - Helper function to get FLB state + * @flb_objp: Pointer to get the restored FLB object + * + * Return: 0 if FLB state found and restored, error if no data found + */ +static int iommu_liveupdate_flb_get_incoming(struct iommu_flb_obj **flb_ob= jp) +{ + struct iommu_flb_obj *flb_obj; + int ret; + + ret =3D liveupdate_flb_get_incoming(&iommu_flb, (void **)flb_objp); + if (ret =3D=3D -ENODATA || ret =3D=3D -ENOENT || ret =3D=3D -EOPNOTSUPP) + return ret; + + if (ret) + goto err_fatal; + + flb_obj =3D *flb_objp; + mutex_lock(&flb_obj->lock); + + /* + * FLB version mismatch is considered fatal for security reasons for + * now. + */ + if (flb_obj->ser->version !=3D IOMMU_LUO_FLB_VERSION) + goto err_fatal; + + /* + * Array Phys of each type should be valid if the FLB was created for + * preservation. This is true even if no devices, iommus or domains were + * preserved. + */ + if (!flb_obj->ser->iommu_array_phys || + !flb_obj->ser->device_array_phys || + !flb_obj->ser->iommu_domain_array_phys) + goto err_fatal; + + mutex_unlock(&flb_obj->lock); + return 0; + +err_fatal: + panic("Failed to restore IOMMU Live Update FLB\n"); +} + +/** + * iommu_for_each_preserved_device() - Iterate on preserved devices + * @fn: Iterator function to call for each preserved device + * @arg: Argument to pass to iterator function + * + * Return: 0 on success, or an error. + */ +int iommu_for_each_preserved_device(iommu_preserved_device_iter_fn fn, + void *arg) +{ + struct iommu_device_array_ser *array; + struct iommu_device_ser *device_ser; + struct iommu_flb_obj *flb_obj; + int ret, idx; + + ret =3D iommu_liveupdate_flb_get_incoming(&flb_obj); + if (ret) + return ret; + + array =3D phys_to_virt(flb_obj->ser->device_array_phys); + iommu_liveupdate_for_each_arr(array) { + iommu_liveupdate_for_each_obj(array, device_ser, idx) { + ret =3D fn(device_ser, arg); + if (ret) + goto out; + } + } + +out: + liveupdate_flb_put_incoming(&iommu_flb); + return ret; +} +EXPORT_SYMBOL(iommu_for_each_preserved_device); + +/** + * iommu_get_preserved_data() - Get preserved data for an IOMMU HW + * @token: Token used to preserve this IOMMU HW + * @type: IOMMU type in preserved state + * + * Gets the preserved state of an IOMMU HW using token and the IOMMU type. + * + * Return: struct iommu_hw_ser on success, NULL if no preserved state foun= d. + */ +struct iommu_hw_ser *iommu_get_preserved_data(u64 token, enum iommu_type_s= er type) +{ + struct iommu_hw_ser *iommu_ser =3D NULL; + struct iommu_hw_array_ser *array; + struct iommu_flb_obj *flb_obj; + int ret, idx; + + ret =3D iommu_liveupdate_flb_get_incoming(&flb_obj); + if (ret) + return NULL; + + array =3D phys_to_virt(flb_obj->ser->iommu_array_phys); + iommu_liveupdate_for_each_arr(array) { + iommu_liveupdate_for_each_obj(array, iommu_ser, idx) { + if (iommu_ser->token =3D=3D token && iommu_ser->type =3D=3D type) + goto out; + } + } + + iommu_ser =3D NULL; +out: + liveupdate_flb_put_incoming(&iommu_flb); + return iommu_ser; +} +EXPORT_SYMBOL(iommu_get_preserved_data); + static int alloc_object_ser(void **curr_array_ptr, u64 max_objs) { struct iommu_array_hdr_ser *curr_array =3D *curr_array_ptr; diff --git a/include/linux/iommu-liveupdate.h b/include/linux/iommu-liveupd= ate.h index 536bea09d064..06f66cc8f526 100644 --- a/include/linux/iommu-liveupdate.h +++ b/include/linux/iommu-liveupdate.h @@ -13,6 +13,16 @@ #include #include =20 +/** + * iommu_preserved_device_iter_fn - Callback for iterating preserved devic= es + * @ser: Pointer to serialized device state + * @arg: User-defined context argument + * + * Return: 0 to continue iteration, or a negative error code to abort. + */ +typedef int (*iommu_preserved_device_iter_fn)(struct iommu_device_ser *ser, + void *arg); + #ifdef CONFIG_IOMMU_LIVEUPDATE int iommu_liveupdate_register_flb(struct liveupdate_file_handler *handler); void iommu_liveupdate_unregister_flb(struct liveupdate_file_handler *handl= er); @@ -37,6 +47,26 @@ static inline void *dev_iommu_preserved_state(struct dev= ice *dev) return NULL; } =20 +/** + * iommu_domain_restored_state() - Get restored state of an IOMMU domain + * @domain: Pointer to struct iommu_domain to get restored state of. + * + * Return: Pointer to restored state on success, NULL on error. + */ +static inline void *iommu_domain_restored_state(struct iommu_domain *domai= n) +{ + struct iommu_domain_ser *ser; + + ser =3D domain->preserved_state; + if (ser && (ser->hdr.flags & IOMMU_SER_FLAG_INCOMING)) + return ser; + + return NULL; +} + +int iommu_for_each_preserved_device(iommu_preserved_device_iter_fn fn, + void *arg); +struct iommu_hw_ser *iommu_get_preserved_data(u64 token, enum iommu_type_s= er type); int iommu_preserve_domain(struct iommu_domain *domain, struct iommu_domain= _ser **ser); void iommu_unpreserve_domain(struct iommu_domain *domain); int iommu_preserve_device(struct iommu_domain *domain, @@ -68,6 +98,21 @@ static inline void *dev_iommu_preserved_state(struct dev= ice *dev) return NULL; } =20 +static inline void *iommu_domain_restored_state(struct iommu_domain *domai= n) +{ + return NULL; +} + +static inline int iommu_for_each_preserved_device(iommu_preserved_device_i= ter_fn fn, void *arg) +{ + return -EOPNOTSUPP; +} + +static inline struct iommu_hw_ser *iommu_get_preserved_data(u64 token, enu= m iommu_type_ser type) +{ + return NULL; +} + static inline int iommu_preserve_domain(struct iommu_domain *domain, struc= t iommu_domain_ser **ser) { return -EOPNOTSUPP; --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f197.google.com (mail-pl1-f197.google.com [209.85.214.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3A7132E285C for ; Mon, 21 Sep 2026 00:48:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.197 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951734; cv=none; b=uwPRgQJX9LjDj6qlCbFUDKwD8unIm6oz9+dxBu28CEmhU/svC5knFPfuGmJBxT6P1UvopkWh+yTkI6O1UUd+zNxaPNHRczarXGikol96zg67FIdIbBIdQIrb+96U9jzgEGguuy6zxgu8SB/ad7HnvG93xAZQv5qyUw5sKy/nBqo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951734; c=relaxed/simple; bh=X89+9MCKR6N4rGgyZM3PSd3rBh+f/2+DypmElVBcadI=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=LKxSHmjukewYKPrmbY0Swk1ThxFRf3VM1fnEqiRd9fcNkH6crBdn/rohbbYbQn3Bugzcteti8KXZ3XQYCsxhAVwGDDtV25wBbuUj0hof2qGgGXc57FF9nqcMB9Qk+W++zrgULBmfTmhv+7cJ+qI3Snr2nkXwuIeo2WO+wipy8sI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Z0ZfCWyV; arc=none smtp.client-ip=209.85.214.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Z0ZfCWyV" Received: by mail-pl1-f197.google.com with SMTP id d9443c01a7336-2dc7337e2a7so31515165ad.0 for ; Sun, 20 Sep 2026 17:48:49 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951729; x=1790556529; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=Ah22k8fiMY4Ue4PB6FU/UzMwLJa6sW6U8ejSPLQHVqg=; b=Z0ZfCWyVOpkmSFBAz8qJfiH3xLdKV1Nt4aX2kngdDyRm56qS/70HolwdKZnT76q4sV Liuk+U1qztV/nlPkaGoydeEf+5vu4LCum2lsMdK+0PmYzQ0kL3oYRX1Mf8d+VS2bK5oI dYVdTYOWEwvmeEk/jFEpm24gsxgz+PsU+kdCpMkwjtLr60ICzLLIUvwsNKS/wU1SGb/o VUgzpEvhallJRzc8eO49CvERLeiRYrKiMVIruMhIKxNxRBCLLchzQkJ0cYLyz0oWLXji BxSgnF18MFYzz9hx1hH1ZYKoU18j8eHsmVA9l6tF+i4bZ03HHr9yPdCDge6q5CPmQHCI OikQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951729; x=1790556529; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Ah22k8fiMY4Ue4PB6FU/UzMwLJa6sW6U8ejSPLQHVqg=; b=t4cDUYkRBaDkJZ1dJtPMCAOFXJsIPURTBgcF0HPiZSYlLdHqCn+yb9WDb8VUxljp22 kNjKSY4AEvE3iRY3dO5YZjgjH3AMjE5He+W+sdmaEqZg4GkUObrNrWdXFeAaiCnuNEmM 5wBX8coPWUZaeUCr1Xts92sRqklRB/cxtuPQQal/ISQPOtOYEk8EtuJt+Yk+3YYOpWfU XaKFKr1sceWwI3U2ki8heKnsssA8/JnzgwzgBLD0WVoBc7GtnIazVoNV1XMnR8PV3H+R HDhR9HSQO6C51/kT9+k/uaAF5LMK3xEIDSm4LCyIAarM8ND2v/DyVsUocQABiV0seOL1 TKYA== X-Forwarded-Encrypted: i=1; AKwUvBx4xi61Q+SZoX9pLlGMq1PoCL+t1GaK1AfZd9CX9hV+yjytVa67s+heIrGqRBspqnFLfi8AQMAt6uF0qls=@vger.kernel.org X-Gm-Message-State: AFuF++mVt659aQLb8p1Yxhm9LJbWh33YeTgTWLE0GZezdD1czohk3jx6 A5oKXUO0rq4UYkpAyzPTIW+hnoCcFuf7+6SAHQSPPft75IyO4x3ZypRZ+mMR1KKaw8cndjYjddq jYrGbqx/NgkceRw== X-Received: from plg10.prod.google.com ([2002:a17:902:c24a:b0:2dd:2c40:5497]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:fc50:b0:2dd:c053:e0f6 with SMTP id d9443c01a7336-2ddc053e13emr81757465ad.37.1789951729140; Sun, 20 Sep 2026 17:48:49 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:26 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-11-skhawaja@google.com> Subject: [PATCH v5 10/18] iommu/vt-d: Restore IOMMU state and reclaimed domain ids From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" During boot fetch the preserved state of IOMMU unit and if found then restore the state. - Reuse the root_table that was preserved in the previous kernel. - Reclaim the domain ids of the preserved domains for each preserved devices so these are not acquired by another domain. Signed-off-by: Samiullah Khawaja --- drivers/iommu/intel/iommu.c | 111 +++++++++++++++++++------------ drivers/iommu/intel/iommu.h | 7 ++ drivers/iommu/intel/liveupdate.c | 79 ++++++++++++++++++++++ 3 files changed, 154 insertions(+), 43 deletions(-) diff --git a/drivers/iommu/intel/iommu.c b/drivers/iommu/intel/iommu.c index 474c926172c5..c5044e834337 100644 --- a/drivers/iommu/intel/iommu.c +++ b/drivers/iommu/intel/iommu.c @@ -987,28 +987,30 @@ static void iommu_disable_translation(struct intel_io= mmu *iommu) raw_spin_unlock_irqrestore(&iommu->register_lock, flag); } =20 -static void disable_dmar_iommu(struct intel_iommu *iommu) +static void release_dmar_iommu(struct intel_iommu *iommu) { - /* - * All iommu domains must have been detached from the devices, - * hence there should be no domain IDs in use. - */ - if (WARN_ON(!ida_is_empty(&iommu->domain_ida))) - return; + struct iommu_hw_ser *iommu_ser; =20 - if (iommu->gcmd & DMA_GCMD_TE) - iommu_disable_translation(iommu); -} + iommu_ser =3D iommu_get_preserved_data(iommu->reg_phys, IOMMU_INTEL); + if (!iommu_ser) { + /* + * All iommu domains must have been detached from the devices, + * hence there should be no domain IDs in use. + */ + WARN_ON(!ida_is_empty(&iommu->domain_ida)); + + if ((iommu->gcmd & DMA_GCMD_TE)) + iommu_disable_translation(iommu); + } =20 -static void free_dmar_iommu(struct intel_iommu *iommu) -{ if (iommu->copied_tables) { bitmap_free(iommu->copied_tables); iommu->copied_tables =3D NULL; } =20 - /* free context mapping */ - free_context_table(iommu); + /* free context mapping if there is no serialized state. */ + if (!iommu_ser) + free_context_table(iommu); =20 if (ecap_prs(iommu->ecap)) intel_iommu_finish_prq(iommu); @@ -1632,12 +1634,19 @@ static int copy_translation_tables(struct intel_iom= mu *iommu) =20 static int __init init_dmars(void) { + struct iommu_hw_ser *iommu_ser; struct dmar_drhd_unit *drhd; struct intel_iommu *iommu; int ret; =20 for_each_iommu(iommu, drhd) { + iommu_ser =3D iommu_get_preserved_data(iommu->reg_phys, IOMMU_INTEL); if (drhd->ignored) { + if (WARN_ON(iommu_ser)) { + ret =3D -EINVAL; + goto free_iommu; + } + iommu_disable_translation(iommu); continue; } @@ -1655,7 +1664,9 @@ static int __init init_dmars(void) } =20 intel_iommu_init_qi(iommu); - init_translation_status(iommu); + + if (!iommu_ser) + init_translation_status(iommu); =20 if (translation_pre_enabled(iommu) && !is_kdump_kernel()) { iommu_disable_translation(iommu); @@ -1664,14 +1675,18 @@ static int __init init_dmars(void) iommu->name); } =20 - /* - * TBD: - * we could share the same root & context tables - * among all IOMMU's. Need to Split it later. - */ - ret =3D iommu_alloc_root_entry(iommu); - if (ret) - goto free_iommu; + if (iommu_ser) { + intel_iommu_liveupdate_restore_root_table(iommu, iommu_ser); + } else { + /* + * TBD: + * we could share the same root & context tables + * among all IOMMU's. Need to Split it later. + */ + ret =3D iommu_alloc_root_entry(iommu); + if (ret) + goto free_iommu; + } =20 if (translation_pre_enabled(iommu)) { pr_info("Translation already enabled - trying to copy translation struc= tures\n"); @@ -1707,7 +1722,10 @@ static int __init init_dmars(void) */ for_each_active_iommu(iommu, drhd) { iommu_flush_write_buffer(iommu); - iommu_set_root_entry(iommu); + + iommu_ser =3D iommu_get_preserved_data(iommu->reg_phys, IOMMU_INTEL); + if (!iommu_ser) + iommu_set_root_entry(iommu); } =20 check_tylersburg_isoch(); @@ -1752,10 +1770,8 @@ static int __init init_dmars(void) return 0; =20 free_iommu: - for_each_active_iommu(iommu, drhd) { - disable_dmar_iommu(iommu); - free_dmar_iommu(iommu); - } + for_each_active_iommu(iommu, drhd) + release_dmar_iommu(iommu); =20 return ret; } @@ -2136,17 +2152,28 @@ int dmar_parse_one_satc(struct acpi_dmar_header *hd= r, void *arg) static int intel_iommu_add(struct dmar_drhd_unit *dmaru) { struct intel_iommu *iommu =3D dmaru->iommu; + struct iommu_hw_ser *iommu_ser; int ret; =20 + /* Use IOMMU HW unit MMIO base to identify the preserved state. */ + iommu_ser =3D iommu_get_preserved_data(iommu->reg_phys, IOMMU_INTEL); + /* * Disable translation if already enabled prior to OS handover. */ - if (iommu->gcmd & DMA_GCMD_TE) + if (!iommu_ser && iommu->gcmd & DMA_GCMD_TE) iommu_disable_translation(iommu); =20 - ret =3D iommu_alloc_root_entry(iommu); - if (ret) - goto out; + if (iommu_ser) { + if (WARN_ON(dmaru->ignored)) + return -EINVAL; + + intel_iommu_liveupdate_restore_root_table(iommu, iommu_ser); + } else { + ret =3D iommu_alloc_root_entry(iommu); + if (ret) + goto out; + } =20 intel_svm_check(iommu); =20 @@ -2165,23 +2192,23 @@ static int intel_iommu_add(struct dmar_drhd_unit *d= maru) if (ecap_prs(iommu->ecap)) { ret =3D intel_iommu_enable_prq(iommu); if (ret) - goto disable_iommu; + goto out; } =20 ret =3D dmar_set_interrupt(iommu); if (ret) - goto disable_iommu; + goto out; + + if (!iommu_ser) + iommu_set_root_entry(iommu); =20 - iommu_set_root_entry(iommu); iommu_enable_translation(iommu); =20 iommu_disable_protect_mem_regions(iommu); return 0; =20 -disable_iommu: - disable_dmar_iommu(iommu); out: - free_dmar_iommu(iommu); + release_dmar_iommu(iommu); return ret; } =20 @@ -2195,12 +2222,10 @@ int dmar_iommu_hotplug(struct dmar_drhd_unit *dmaru= , bool insert) if (iommu =3D=3D NULL) return -EINVAL; =20 - if (insert) { + if (insert) ret =3D intel_iommu_add(dmaru); - } else { - disable_dmar_iommu(iommu); - free_dmar_iommu(iommu); - } + else + release_dmar_iommu(iommu); =20 return ret; } diff --git a/drivers/iommu/intel/iommu.h b/drivers/iommu/intel/iommu.h index 4feb5bd76b18..3a2cb08c0ac1 100644 --- a/drivers/iommu/intel/iommu.h +++ b/drivers/iommu/intel/iommu.h @@ -1308,10 +1308,17 @@ int intel_iommu_preserve(struct iommu_device *iommu, void intel_iommu_unpreserve(struct iommu_device *iommu, struct iommu_hw_ser *iommu_ser); void clear_unpreserved_context_entries(struct intel_iommu *iommu); +void intel_iommu_liveupdate_restore_root_table(struct intel_iommu *iommu, + struct iommu_hw_ser *iommu_ser); #else static inline void clear_unpreserved_context_entries(struct intel_iommu *i= ommu) { } + +static inline void intel_iommu_liveupdate_restore_root_table(struct intel_= iommu *iommu, + struct iommu_hw_ser *iommu_ser) +{ +} #endif =20 #ifdef CONFIG_INTEL_IOMMU_SVM diff --git a/drivers/iommu/intel/liveupdate.c b/drivers/iommu/intel/liveupd= ate.c index 501dc0e9cc0a..c9e553379683 100644 --- a/drivers/iommu/intel/liveupdate.c +++ b/drivers/iommu/intel/liveupdate.c @@ -272,6 +272,85 @@ static int preserve_iommu_context_tables(struct device= _domain_info *info) return 0; } =20 +static void restore_iommu_context(struct intel_iommu *iommu) +{ + struct context_entry *context; + int i; + + for (i =3D 0; i < ROOT_ENTRY_NR; i++) { + context =3D iommu_context_addr(iommu, i, 0, 0); + if (context) + iommu_restore_pages(virt_to_phys(context)); + + if (!sm_supported(iommu)) + continue; + + context =3D iommu_context_addr(iommu, i, 0x80, 0); + if (context) + iommu_restore_pages(virt_to_phys(context)); + } +} + +static int _restore_used_domain_ids(struct iommu_device_ser *ser, void *ar= g) +{ + int id =3D ser->domain_iommu_ser.attachment_id; + struct iommu_hw_ser *iommu_hw_ser; + struct intel_iommu *iommu =3D arg; + int ret; + + if (WARN_ON(!ser->domain_iommu_ser.iommu_phys)) + return 0; + + iommu_hw_ser =3D phys_to_virt(ser->domain_iommu_ser.iommu_phys); + if (iommu_hw_ser->type !=3D IOMMU_INTEL) + return 0; + + /* Only allocate domain ID from associated IOMMU HW unit */ + if (iommu_hw_ser->intel.phys_addr !=3D iommu->reg_phys) + return 0; + + guard(mutex)(&iommu->did_lock); + + /* + * The domain IDs are reclaimed while the IOMMU HW unit is being + * restored and not registered with the IOMMU core. So if the ID already + * exists, it is safe to assume that a preserved device sharing the same + * domain ID reclaimed it. + */ + if (ida_exists(&iommu->domain_ida, id)) + return 0; + + ret =3D ida_alloc_range(&iommu->domain_ida, id, id, GFP_KERNEL); + if (ret < 0) + return ret; + + return 0; +} + +/** + * intel_iommu_liveupdate_restore_root_table() - Restore root table and re= claim domain IDs + * @iommu: Target IOMMU + * @iommu_ser: Serialized IOMMU hardware state from previous kernel + * + * Restores the preserved root table and context tables for the IOMMU hard= ware + * instance across Live Update, and reclaims all domain IDs previously all= ocated + * to preserved devices so they are not reused. + */ +void intel_iommu_liveupdate_restore_root_table(struct intel_iommu *iommu, + struct iommu_hw_ser *iommu_ser) +{ + if (!iommu_ser->intel.restored) + iommu_restore_pages(iommu_ser->intel.root_table); + + iommu->root_entry =3D __va(iommu_ser->intel.root_table); + + if (!iommu_ser->intel.restored) + restore_iommu_context(iommu); + + iommu_ser->intel.restored =3D 1; + BUG_ON(iommu_for_each_preserved_device(_restore_used_domain_ids, iommu)); +} + /** * intel_iommu_preserve_device() - Intel IOMMU callback to preserve device= state * @dev: Target device --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 615B12D1F40 for ; Mon, 21 Sep 2026 00:48:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951737; cv=none; b=dh/li1Sk9DF7O+DOlTZ/YKDqf5In7qONhkMtbkh03UwN65W+YXAUTgw5Yoc5q6P3+J2rEF0Uh/JHML8at4xzZydAFMzjOGBfp5gp5VjyMz44HHbV8yZKzrob1UqDZiNUueb5Zrtjrz6B4zW/1Leb9Oq/ICU2ZxNIsO/UPWXPg8o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951737; c=relaxed/simple; bh=gwAv+wRVw3t/fG51U9Ec6hw3eDnLX/9UJyVoiPX5g7Y=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=tmVw0/40hGqBI6QUSKnKyiu1i3/xtXkQ2I+wNJJWrE/11k6iTkcUU2bWSZFWncwLLzJen8XizvRhvf1UouCOvKye6gfboOFKzc/5IVjDcBs+oPu9FkUYfkVYrgK2rBqWQYCyTJB345MxDwKgHkBJ86PUUJgQeOjnloNFyIksglk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=nt24gl6w; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="nt24gl6w" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2db87f759c5so50275195ad.1 for ; Sun, 20 Sep 2026 17:48:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951730; x=1790556530; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=vboebmTOKLm+aioJh8qWJMYwNYCO/nAezbpt/kP9hZc=; b=nt24gl6wZbD1fimF8AFLJYzNIJZWHqrzYFhquDcnnLaBFOjRigrlEFycQIyihPPTSi HbeP2Du4r2/Bb+FLDKtRy/igL1S4jZJjlJXOpViepWbAQDs0FgGP4RM9LOwjeq1ojgFP QycM8nQjTbWO86gesHoOeTKRxRPIHCCZB+X3bsmjnSY6YfLvIn8nJvJ/odDZLLzIrKPs PZN7IB92+R3aUpXgy+nwDe5H1UzzubIwGK7Ue+FkpSTkFQ42dQ4PdSiuy5r1XpRsLNHt 3z7MsiEw2GySPIFAVVdzs5AN8QHcS8Be5lzV10qRheAXEiyPoecHxEfGEJnPcgesfI1f tjkg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951730; x=1790556530; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=vboebmTOKLm+aioJh8qWJMYwNYCO/nAezbpt/kP9hZc=; b=ID6N+oyMU0NkvYU/YnNNPibl9SbEh1xFjs6I5un6g2u9xby9FDgkq7GAVDq2PuWIZK n82rJpkqWX0eOrKF7w0M2VPLda3EZFIRVokaFgdhThEH60cykZ0vHvFaniQy26y0nrmx Qv0kFN1HD2/yMVrJA668uGrUtZ0bdl3XhWl4E3fvmyWI66lytYGwTxrt9dFmy8MMv3kv Y5xt3rRI4mt6dEC65qeatoNwskydn6Jgqe56V5UlNZ7OoRR24sWl/mDh29xwNfuEJuGn xZHNnGi92Jb37UtVcMwh8VeFJcwiIlscZ3J8VyC12f4fEzkebs8SqI2ZtVy9ZhcXvrQB IBkg== X-Forwarded-Encrypted: i=1; AKwUvBwJBbYOX9aLm6gt64jgj2i8LShZkIB6M1hVvpTclhjwkBmH9eJM37K3SAxpVPOx7JhjxL34KZv7GQfekb8=@vger.kernel.org X-Gm-Message-State: AFuF++nK7R2UnSGzn1mGBtSpQ3iP0Ol8uFBiJpymty+lRuliu+7VN2hd J36IBfUVi9XRirm+aKwcidbR1hZeQz/F1C3M5FSkqjqAAscaQTFDG6MJWchf0hXPZ35xA0UZWyq ++56h/q2yt2TW6Q== X-Received: from plly11.prod.google.com ([2002:a17:902:7c8b:b0:2dd:c0fa:93f7]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:fa5:b0:2dd:c100:b2c1 with SMTP id d9443c01a7336-2ddc100b33bmr78386365ad.44.1789951729909; Sun, 20 Sep 2026 17:48:49 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:27 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-12-skhawaja@google.com> Subject: [PATCH v5 11/18] iommu: Restore and reattach preserved domains to devices From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" During default domain setup, restore the preserved domains by restoring the page tables using restore() iommupt op. Associated the restored domain with the iommu group of the preserved device, and reattach the domain to the device. Signed-off-by: Samiullah Khawaja --- drivers/iommu/iommu.c | 116 +++++++++++++++++++++++++++++-- drivers/iommu/liveupdate.c | 116 +++++++++++++++++++++++++++++++ include/linux/iommu-liveupdate.h | 64 +++++++++++++++++ 3 files changed, 290 insertions(+), 6 deletions(-) diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c index 9c1ad8de3ad7..cca1eb332a43 100644 --- a/drivers/iommu/iommu.c +++ b/drivers/iommu/iommu.c @@ -158,6 +158,7 @@ static void __iommu_group_set_domain_nofail(struct iomm= u_group *group, WARN_ON(__iommu_group_set_domain_internal( group, new_domain, IOMMU_SET_DOMAIN_MUST_SUCCEED)); } +static int __iommu_group_alloc_blocking_domain(struct iommu_group *group); =20 static int iommu_setup_default_domain(struct iommu_group *group, int target_type); @@ -540,6 +541,10 @@ static int iommu_init_device(struct device *dev) goto err_free; } =20 +#ifdef CONFIG_IOMMU_LIVEUPDATE + iommu_init_device_preserved_data(dev); +#endif + iommu_dev =3D ops->probe_device(dev); if (IS_ERR(iommu_dev)) { ret =3D PTR_ERR(iommu_dev); @@ -604,7 +609,8 @@ static void iommu_deinit_device(struct device *dev) * Regardless, if a delayed attach never occurred, then the release * should still avoid touching any hardware configuration either. */ - if (!dev->iommu->attach_deferred && ops->release_domain) { + if (!dev->iommu->attach_deferred && ops->release_domain && + !dev_iommu_restored_state(dev)) { struct iommu_domain *release_domain =3D ops->release_domain; =20 /* @@ -694,7 +700,8 @@ static int __iommu_probe_device(struct device *dev, str= uct list_head *group_list } =20 for_each_group_device(group, gdev2) { - if (dev_iommu_preserved_state(gdev2->dev)) { + if (dev_iommu_preserved_state(gdev2->dev) || + dev_iommu_restored_state(gdev2->dev)) { ret =3D -EBUSY; goto err_free_gdev; } @@ -764,6 +771,27 @@ int iommu_probe_device(struct device *dev) return 0; } =20 +static void __iommu_group_remove_restored_device(struct iommu_group *group, + struct device *dev) +{ + struct iommu_device_ser *device_ser; + + lockdep_assert_held(&group->mutex); + device_ser =3D dev_iommu_restored_state(dev); + if (!device_ser) + return; + + if (!group->owner_cnt || group->owner !=3D device_ser) + return; + + if (group->owner_cnt > 1) { + group->owner_cnt--; + } else { + group->owner_cnt =3D 0; + group->owner =3D NULL; + } +} + static void __iommu_group_free_device(struct iommu_group *group, struct group_device *grp_dev) { @@ -775,13 +803,15 @@ static void __iommu_group_free_device(struct iommu_gr= oup *group, trace_remove_device_from_group(group->id, dev); =20 /* - * If the group has become empty then ownership must have been - * released, and the current domain must be set back to NULL or - * the default domain. + * If the group has become empty then ownership must have been released, + * and the current domain must be set back to NULL or the default + * domain. A restored domain remains attached to the restored device on + * removal. */ if (list_empty(&group->devices)) WARN_ON(group->owner_cnt || - group->domain !=3D group->default_domain); + (group->domain !=3D group->default_domain && + !iommu_domain_restored_state(group->domain))); =20 kfree(grp_dev->name); kfree(grp_dev); @@ -794,6 +824,7 @@ static void __iommu_group_remove_device(struct device *= dev) struct group_device *device; =20 mutex_lock(&group->mutex); + __iommu_group_remove_restored_device(group, dev); for_each_group_device(group, device) { if (device->dev !=3D dev) continue; @@ -2211,6 +2242,7 @@ static int __iommu_attach_device(struct iommu_domain = *domain, ret =3D domain->ops->attach_dev(domain, dev, old); if (ret) return ret; + dev->iommu->attach_deferred =3D 0; trace_attach_device_to_domain(dev); return 0; @@ -3175,6 +3207,62 @@ int iommu_fwspec_add_ids(struct device *dev, const u= 32 *ids, int num_ids) } EXPORT_SYMBOL_GPL(iommu_fwspec_add_ids); =20 +static struct device *__iommu_group_restored_device(struct iommu_group *gr= oup) +{ + struct group_device *gdev; + + lockdep_assert_held(&group->mutex); + for_each_group_device(group, gdev) { + if (!dev_is_pci(gdev->dev)) + continue; + + if (dev_iommu_restored_state(gdev->dev)) + return gdev->dev; + } + + return NULL; +} + +static int __iommu_group_restore_domain(struct iommu_group *group) +{ + struct iommu_device_ser *device_ser; + struct iommu_domain *domain; + struct device *dev; + void *owner; + int ret; + + lockdep_assert_held(&group->mutex); + if (group->domain) + return -EBUSY; + + dev =3D __iommu_group_restored_device(group); + device_ser =3D dev_iommu_restored_state(dev); + if (!device_ser) + return -ENOENT; + + ret =3D __iommu_group_alloc_blocking_domain(group); + if (ret) + return ret; + + domain =3D iommu_restore_domain(dev, device_ser, &owner); + if (WARN_ON(IS_ERR(domain))) + return PTR_ERR(domain); + + /* The restored domain is attached with the restored device. */ + ret =3D __iommu_group_set_domain(group, domain); + if (ret) + return ret; + + /* + * Ownership of groups with preserved devices is set during boot. These + * will be reclaimed later by the entity (iommufd) that preserved them. + */ + WARN_ON(group->owner); + group->owner =3D owner; + group->owner_cnt =3D 1; + return ret; +} + /** * iommu_setup_default_domain - Set the default_domain for the group * @group: Group to change @@ -3233,6 +3321,16 @@ static int iommu_setup_default_domain(struct iommu_g= roup *group, =20 /* We must set default_domain early for __iommu_device_set_domain */ group->default_domain =3D dom; + + /* Preserved devices need to be attached to the restore domain */ + if (__iommu_group_restored_device(group)) { + ret =3D __iommu_group_restore_domain(group); + if (ret) + goto err_restore_def_domain; + + goto out_free_old; + } + if (!group->domain) { /* * Drivers are not allowed to fail the first domain attach. @@ -4102,6 +4200,9 @@ int pci_dev_reset_iommu_prepare(struct pci_dev *pdev) if (!pci_ats_supported(pdev) || !dev_has_iommu(&pdev->dev)) return 0; =20 + if (dev_iommu_restored_state(&pdev->dev)) + return 0; + guard(mutex)(&group->mutex); =20 gdev =3D __dev_to_gdev(&pdev->dev); @@ -4213,6 +4314,9 @@ void pci_dev_reset_iommu_done(struct pci_dev *pdev) if (!pci_ats_supported(pdev) || !dev_has_iommu(&pdev->dev)) return; =20 + if (dev_iommu_restored_state(&pdev->dev)) + return; + guard(mutex)(&group->mutex); =20 gdev =3D __dev_to_gdev(&pdev->dev); diff --git a/drivers/iommu/liveupdate.c b/drivers/iommu/liveupdate.c index ff76a23d9583..4fa12b4efea5 100644 --- a/drivers/iommu/liveupdate.c +++ b/drivers/iommu/liveupdate.c @@ -716,3 +716,119 @@ void iommu_unpreserve_device(struct iommu_domain *dom= ain, struct device *dev) liveupdate_flb_put_outgoing(&iommu_flb); } EXPORT_SYMBOL_GPL(iommu_unpreserve_device); + +static inline bool match_device_ser(struct iommu_device_ser *match, + struct pci_dev *pdev) +{ + return match->devid =3D=3D pci_dev_id(pdev) && match->pci_domain_nr =3D= =3D pci_domain_nr(pdev->bus); +} + +/** + * iommu_init_device_preserved_data() - Initialize preserved state for dev= ice + * @dev: Target device + * + * Looks up incoming Live Update state for @dev and attaches it to the dev= ice if + * found. + */ +void iommu_init_device_preserved_data(struct device *dev) +{ + struct iommu_device_ser *device_ser =3D NULL; + struct iommu_device_array_ser *array; + struct iommu_flb_obj *flb_obj; + int ret, idx; + + if (!dev_is_pci(dev)) + return; + + ret =3D iommu_liveupdate_flb_get_incoming(&flb_obj); + if (ret) + return; + + mutex_lock(&flb_obj->lock); + array =3D phys_to_virt(flb_obj->ser->device_array_phys); + iommu_liveupdate_for_each_arr(array) { + iommu_liveupdate_for_each_obj(array, device_ser, idx) { + if (match_device_ser(device_ser, to_pci_dev(dev))) { + device_ser->hdr.flags |=3D IOMMU_SER_FLAG_INCOMING; + goto out; + } + } + } + + device_ser =3D NULL; +out: + WRITE_ONCE(dev->iommu->device_ser, device_ser); + mutex_unlock(&flb_obj->lock); + liveupdate_flb_put_incoming(&iommu_flb); +} +EXPORT_SYMBOL(iommu_init_device_preserved_data); + +/** + * iommu_restore_domain() - Restore a preserved domain for a device + * @dev: Target device + * @ser: Serialized device state + * @owner: Pointer to store group owner handle + * + * Restores or reuses a restored preserved domain for @dev from serialized= state + * @ser. + * + * Return: Restored iommu_domain pointer, or ERR_PTR. + */ +struct iommu_domain *iommu_restore_domain(struct device *dev, + struct iommu_device_ser *ser, + void **owner) +{ + struct iommu_domain_ser *domain_ser; + struct iommu_flb_obj *flb_obj; + struct iommu_domain *domain; + struct pt_iommu *pt; + int ret; + + ret =3D iommu_liveupdate_flb_get_incoming(&flb_obj); + if (ret) + return ERR_PTR(ret); + + mutex_lock(&flb_obj->lock); + + /* Preserved device should have a preserved domain */ + if (!ser->domain_iommu_ser.domain_phys) { + domain =3D ERR_PTR(-EINVAL); + goto out; + } + + domain_ser =3D phys_to_virt(ser->domain_iommu_ser.domain_phys); + if (domain_ser->restored_domain) { + *owner =3D ser; + domain =3D domain_ser->restored_domain; + goto out; + } + + domain_ser->hdr.flags |=3D IOMMU_SER_FLAG_INCOMING; + domain =3D iommu_paging_domain_alloc(dev); + if (IS_ERR(domain)) + goto out; + + pt =3D iommupt_from_domain(domain); + if (!pt) { + iommu_domain_free(domain); + domain =3D ERR_PTR(-EOPNOTSUPP); + goto out; + } + + ret =3D pt->ops->restore(pt, domain_ser); + if (ret) { + iommu_domain_free(domain); + domain =3D ERR_PTR(ret); + goto out; + } + + /* The device is owned by the preserved state. */ + *owner =3D ser; + domain->preserved_state =3D domain_ser; + domain_ser->restored_domain =3D domain; + +out: + mutex_unlock(&flb_obj->lock); + liveupdate_flb_put_incoming(&iommu_flb); + return domain; +} diff --git a/include/linux/iommu-liveupdate.h b/include/linux/iommu-liveupd= ate.h index 06f66cc8f526..3ff537ae081d 100644 --- a/include/linux/iommu-liveupdate.h +++ b/include/linux/iommu-liveupdate.h @@ -64,8 +64,51 @@ static inline void *iommu_domain_restored_state(struct i= ommu_domain *domain) return NULL; } =20 +/** + * dev_iommu_restored_state() - Get restored state of a device + * @dev: Target device + * + * Return: Restored state pointer or NULL. + */ +static inline void *dev_iommu_restored_state(struct device *dev) +{ + struct iommu_device_ser *ser; + + if (!dev->iommu) + return NULL; + + ser =3D READ_ONCE(dev->iommu->device_ser); + if (ser && (ser->hdr.flags & IOMMU_SER_FLAG_INCOMING)) + return ser; + + return NULL; +} + +/** + * dev_iommu_restore_did() - Get restored domain ID for a device + * @dev: Target device + * @domain: Target domain + * + * Fetches the domain ID preserved for @dev and @domain across Live Update. + * + * Return: Domain ID or -1 on error. + */ +static inline int dev_iommu_restore_did(struct device *dev, struct iommu_d= omain *domain) +{ + struct iommu_device_ser *ser =3D dev_iommu_restored_state(dev); + + if (ser && iommu_domain_restored_state(domain)) + return ser->domain_iommu_ser.attachment_id; + + return -1; +} + +struct iommu_domain *iommu_restore_domain(struct device *dev, + struct iommu_device_ser *ser, + void **owner); int iommu_for_each_preserved_device(iommu_preserved_device_iter_fn fn, void *arg); +void iommu_init_device_preserved_data(struct device *dev); struct iommu_hw_ser *iommu_get_preserved_data(u64 token, enum iommu_type_s= er type); int iommu_preserve_domain(struct iommu_domain *domain, struct iommu_domain= _ser **ser); void iommu_unpreserve_domain(struct iommu_domain *domain); @@ -98,16 +141,37 @@ static inline void *dev_iommu_preserved_state(struct d= evice *dev) return NULL; } =20 +static inline void *dev_iommu_restored_state(struct device *dev) +{ + return NULL; +} + +static inline int dev_iommu_restore_did(struct device *dev, struct iommu_d= omain *domain) +{ + return -1; +} + static inline void *iommu_domain_restored_state(struct iommu_domain *domai= n) { return NULL; } =20 +static inline struct iommu_domain *iommu_restore_domain(struct device *dev, + struct iommu_device_ser *ser, + void **owner) +{ + return NULL; +} + static inline int iommu_for_each_preserved_device(iommu_preserved_device_i= ter_fn fn, void *arg) { return -EOPNOTSUPP; } =20 +static inline void iommu_init_device_preserved_data(struct device *dev) +{ +} + static inline struct iommu_hw_ser *iommu_get_preserved_data(u64 token, enu= m iommu_type_ser type) { return NULL; --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8F0A1288505 for ; Mon, 21 Sep 2026 00:48:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951734; cv=none; b=rKlaJWt47bjw8YwW6fT0khpTx/Iqep7+gQEZYpSoWtB3nNaLtGRxijvECDelFVcTyGH6lvccZdLYvZfnGfQZg+e1rL5AoeKX7tRniNUv3JmZIMjbAUKp1+haFVrVUP0nuqaWBYqi9xu336ngj3dJUfmlSw9pHa81vTpf96vIYKs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951734; c=relaxed/simple; bh=zjw0Tnp8stathm9nlmJFxU761/AklxkMadgSeSWjsjE=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=uTQv7eHsLFfYT+dEev3+f9TSVVDwb+v3zWBLXUB2Rf+fsJEmv3rHASx67Nf74r2w/JupOaUkECE6+S6UfoKEcidlxN5JLdwCCWoVw4TP9OtwZd3h26rnPsuYqe4stlCpu2b7xxS4qRj2hM+rvzsROSKt9OtpoUMeYQhpOW4583Q= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=H0Nvy92a; arc=none smtp.client-ip=209.85.214.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="H0Nvy92a" Received: by mail-pl1-f198.google.com with SMTP id d9443c01a7336-2d63bad3d09so35784625ad.3 for ; Sun, 20 Sep 2026 17:48:51 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951731; x=1790556531; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=6I3SvizF5qvgyNRrC3PScbaNSFdWu6HfLifwVu0oFvU=; b=H0Nvy92aCgW3Sjm5Gl47CQu9WN0/YH/yuDp/IynJv4+mUwUuRUcSu5Mbe7kM2hIeX6 hDNFYr7iD59hzJEZRCPd6G03iHXbvwAmMyqiNlO/M8RkRkAzJUWvm61vzv5XbZooP81Y oHLJR33oavIiCeb3DS+GVVsP5iMuD2snricXhClth6vrO7+7fsksszt9hSeuwwu/c2BM 3M8fjS3N7tCGDRua5eQayAIqQqOg+OWDgvmpDBX0Eb86ZpUjU6xbjvVHCgRK5IjRTK9g L/aBTP2n9y5hZ+jMLyWeZfyEJk1HtsZ0pdrei2e0UFJD3slk0HnH6GIY/tmH/muSdC62 ZYcg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951731; x=1790556531; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=6I3SvizF5qvgyNRrC3PScbaNSFdWu6HfLifwVu0oFvU=; b=LIRchIkirmh+QxGeQcwYLstZzVQ48Gb/tsTV0oJAkg2qiNeeT1XllzASKaj6ucXy4L 3BPBYp1RzMba2/PBknHVPhtfWhWbBCkWqoHMEuGmCBkf6wfHADkaj/qr48Lxq3gkGi5i /x0QDolH6Xht99l6dsEJiD6He5ftChWUUuGPbW3Um7+s1WWVBfC8FAQ4fHZVGbACoDtv rbRF0w9qtCXEJjLUFN5i1nUluk8lbJMZRPjpp7G1RZxXifNRxwBo+YqHbk7eZRGLXe5A iumLknUf+6FQS+uOMj6r5M4E18msgPBh2Q8MUmwGMIyYiw919aVVIKX6o8ww028Grqui XIrQ== X-Forwarded-Encrypted: i=1; AKwUvBx9NUTzUgrmmbqAnpBbBwTfZiOefcOrTJdPiYnsDu1fPVhUv7X9+254hHLgAzb7d2dEmvns7DHKO5T6bHg=@vger.kernel.org X-Gm-Message-State: AFuF++kWsAN40jratP+4CO9g1EEi/3aX8qvhCUjaPCQ/ZxlGt+HrhCrf HAr6yxfcBvE09Bl5Pm6OAKsl4NKhFSG78ZYgppUyIpPU/MRuUtEtnVVXWZnOuob9YInU6YloC2i Ph+FBl3BCOvHzsQ== X-Received: from plmm2.prod.google.com ([2002:a17:902:c442:b0:2df:4477:df93]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:bf46:b0:2dd:c100:4b79 with SMTP id d9443c01a7336-2ddc1005efcmr55134685ad.48.1789951730815; Sun, 20 Sep 2026 17:48:50 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:28 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-13-skhawaja@google.com> Subject: [PATCH v5 12/18] iommu/vt-d: Handle reattach of the restored domain From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Reattach the restored domain to the preserved device using restored domain ID. While reattaching do not setup the context and PASID entries as those are preserved during liveupdate. Signed-off-by: Samiullah Khawaja --- drivers/iommu/intel/iommu.c | 32 +++++-- drivers/iommu/intel/iommu.h | 14 +++ drivers/iommu/intel/liveupdate.c | 159 +++++++++++++++++++++++++++++++ 3 files changed, 197 insertions(+), 8 deletions(-) diff --git a/drivers/iommu/intel/iommu.c b/drivers/iommu/intel/iommu.c index c5044e834337..6d3cbe0745c9 100644 --- a/drivers/iommu/intel/iommu.c +++ b/drivers/iommu/intel/iommu.c @@ -2819,6 +2819,9 @@ static int blocking_domain_attach_dev(struct iommu_do= main *domain, { struct device_domain_info *info =3D dev_iommu_priv_get(dev); =20 + if (dev_iommu_restored_state(dev)) + return -EBUSY; + iopf_for_domain_remove(info->domain ? &info->domain->domain : NULL, dev); device_block_translation(dev); return 0; @@ -3187,6 +3190,9 @@ static int intel_iommu_attach_device(struct iommu_dom= ain *domain, { int ret; =20 + if (dev_iommu_restored_state(dev)) + return intel_iommu_restore_device(domain, dev); + device_block_translation(dev); =20 ret =3D paging_domain_compatible(domain, dev); @@ -3317,7 +3323,8 @@ static struct iommu_device *intel_iommu_probe_device(= struct device *dev) info->iommu =3D iommu; RB_CLEAR_NODE(&info->node); if (dev_is_pci(dev)) { - if (ecap_dev_iotlb_support(iommu->ecap) && + if (!dev_iommu_restored_state(dev) && + ecap_dev_iotlb_support(iommu->ecap) && pci_ats_supported(pdev) && dmar_ats_supported(pdev, iommu)) { info->ats_supported =3D 1; @@ -3421,12 +3428,16 @@ static void intel_iommu_release_device(struct devic= e *dev) struct device_domain_info *info =3D dev_iommu_priv_get(dev); struct intel_iommu *iommu =3D info->iommu; =20 - iommu_disable_pci_pri(info); - iommu_disable_pci_ats(info); + if (!dev_iommu_restored_state(dev)) { + iommu_disable_pci_pri(info); + iommu_disable_pci_ats(info); =20 - if (info->pasid_enabled) { - pci_disable_pasid(to_pci_dev(dev)); - info->pasid_enabled =3D 0; + if (info->pasid_enabled) { + pci_disable_pasid(to_pci_dev(dev)); + info->pasid_enabled =3D 0; + } + } else { + intel_iommu_detach_restored_device(dev); } =20 mutex_lock(&iommu->iopf_lock); @@ -3434,11 +3445,13 @@ static void intel_iommu_release_device(struct devic= e *dev) device_rbtree_remove(info); mutex_unlock(&iommu->iopf_lock); =20 - if (sm_supported(iommu) && !dev_is_real_dma_subdevice(dev) && + if (!dev_iommu_restored_state(dev) && sm_supported(iommu) && + !dev_is_real_dma_subdevice(dev) && !context_copied(iommu, info->bus, info->devfn)) intel_pasid_teardown_sm_context(dev); =20 - intel_pasid_free_table(dev); + if (!dev_iommu_restored_state(dev)) + intel_pasid_free_table(dev); intel_iommu_debugfs_remove_dev(info); kfree(info); } @@ -3900,6 +3913,9 @@ static int identity_domain_attach_dev(struct iommu_do= main *domain, struct intel_iommu *iommu =3D info->iommu; int ret; =20 + if (dev_iommu_restored_state(dev)) + return -EBUSY; + device_block_translation(dev); =20 if (dev_is_real_dma_subdevice(dev)) diff --git a/drivers/iommu/intel/iommu.h b/drivers/iommu/intel/iommu.h index 3a2cb08c0ac1..25104644317c 100644 --- a/drivers/iommu/intel/iommu.h +++ b/drivers/iommu/intel/iommu.h @@ -1310,6 +1310,9 @@ void intel_iommu_unpreserve(struct iommu_device *iomm= u, void clear_unpreserved_context_entries(struct intel_iommu *iommu); void intel_iommu_liveupdate_restore_root_table(struct intel_iommu *iommu, struct iommu_hw_ser *iommu_ser); +int intel_iommu_restore_device(struct iommu_domain *domain, + struct device *dev); +int intel_iommu_detach_restored_device(struct device *dev); #else static inline void clear_unpreserved_context_entries(struct intel_iommu *i= ommu) { @@ -1319,6 +1322,17 @@ static inline void intel_iommu_liveupdate_restore_ro= ot_table(struct intel_iommu struct iommu_hw_ser *iommu_ser) { } + +static inline int intel_iommu_restore_device(struct iommu_domain *domain, + struct device *dev) +{ + return -EOPNOTSUPP; +} + +static inline int intel_iommu_detach_restored_device(struct device *dev) +{ + return -EOPNOTSUPP; +} #endif =20 #ifdef CONFIG_INTEL_IOMMU_SVM diff --git a/drivers/iommu/intel/liveupdate.c b/drivers/iommu/intel/liveupd= ate.c index c9e553379683..6e6707eefc8c 100644 --- a/drivers/iommu/intel/liveupdate.c +++ b/drivers/iommu/intel/liveupdate.c @@ -351,6 +351,165 @@ void intel_iommu_liveupdate_restore_root_table(struct= intel_iommu *iommu, BUG_ON(iommu_for_each_preserved_device(_restore_used_domain_ids, iommu)); } =20 +static void domain_detach_reattached_iommu(struct dmar_domain *domain, + struct intel_iommu *iommu) +{ + struct iommu_domain_info *info; + + guard(mutex)(&iommu->did_lock); + info =3D xa_load(&domain->iommu_array, iommu->seq_id); + if (--info->refcnt =3D=3D 0) { + xa_erase(&domain->iommu_array, iommu->seq_id); + kfree(info); + } +} + +static int domain_reattach_iommu(struct dmar_domain *domain, + struct intel_iommu *iommu, + struct iommu_device_ser *device_ser) +{ + struct iommu_domain_info *info, *curr; + struct iommu_domain_ser *domain_ser; + struct iommu_hw_ser *iommu_hw_ser; + int restored_did; + int ret; + + if (!iommu_domain_restored_state(&domain->domain)) + return -EINVAL; + + if (!device_ser->domain_iommu_ser.domain_phys || + !device_ser->domain_iommu_ser.iommu_phys) + return -EINVAL; + + domain_ser =3D phys_to_virt(device_ser->domain_iommu_ser.domain_phys); + if (domain_ser->restored_domain !=3D &domain->domain) + return -EINVAL; + + iommu_hw_ser =3D phys_to_virt(device_ser->domain_iommu_ser.iommu_phys); + if (iommu_hw_ser->type !=3D IOMMU_INTEL || + iommu_hw_ser->intel.phys_addr !=3D iommu->reg_phys) + return -EINVAL; + + restored_did =3D device_ser->domain_iommu_ser.attachment_id; + if (!ida_exists(&iommu->domain_ida, restored_did)) + return -EINVAL; + + info =3D kzalloc_obj(*info); + if (!info) + return -ENOMEM; + + guard(mutex)(&iommu->did_lock); + curr =3D xa_load(&domain->iommu_array, iommu->seq_id); + if (curr) { + curr->refcnt++; + kfree(info); + return 0; + } + + info->refcnt =3D 1; + info->did =3D restored_did; + info->iommu =3D iommu; + curr =3D xa_cmpxchg(&domain->iommu_array, iommu->seq_id, + NULL, info, GFP_KERNEL); + if (curr) { + ret =3D xa_err(curr) ? : -EBUSY; + goto err_unlock; + } + + return 0; + +err_unlock: + kfree(info); + return ret; +} + +/** + * intel_iommu_restore_device() - Restore device domain attachment after l= ive update + * @domain: Restored domain + * @dev: Restored device + * + * Return: 0 on success, or negative error code. + */ +int intel_iommu_restore_device(struct iommu_domain *domain, + struct device *dev) +{ + struct iommu_device_ser *device_ser =3D dev_iommu_restored_state(dev); + struct device_domain_info *info =3D dev_iommu_priv_get(dev); + struct dmar_domain *dmar_domain =3D to_dmar_domain(domain); + struct intel_iommu *iommu =3D info->iommu; + unsigned long flags; + int ret; + + if (!device_ser) + return -EINVAL; + + if (dev_is_real_dma_subdevice(dev)) + return -EOPNOTSUPP; + + ret =3D domain_reattach_iommu(dmar_domain, iommu, device_ser); + if (ret) + return ret; + + info->domain =3D dmar_domain; + info->domain_attached =3D true; + spin_lock_irqsave(&dmar_domain->lock, flags); + list_add(&info->link, &dmar_domain->devices); + spin_unlock_irqrestore(&dmar_domain->lock, flags); + + ret =3D cache_tag_assign_domain(dmar_domain, dev, IOMMU_NO_PASID); + if (ret) + goto err; + + ret =3D iopf_for_domain_set(domain, dev); + if (ret) + goto err; + + return 0; + +err: + /* + * Detach the restored domain from device and iommu on failure, but keep + * the hardware state intact. + */ + info->domain_attached =3D false; + cache_tag_unassign_domain(info->domain, dev, IOMMU_NO_PASID); + spin_lock_irqsave(&info->domain->lock, flags); + list_del(&info->link); + spin_unlock_irqrestore(&info->domain->lock, flags); + + domain_detach_reattached_iommu(info->domain, iommu); + info->domain =3D NULL; + return ret; +} + +int intel_iommu_detach_restored_device(struct device *dev) +{ + struct device_domain_info *info =3D dev_iommu_priv_get(dev); + struct intel_iommu *iommu =3D info->iommu; + struct iommu_domain *domain; + unsigned long flags; + + if (!info->domain_attached || !info->domain) + return -EINVAL; + + domain =3D &info->domain->domain; + if (!iommu_domain_restored_state(domain)) + return -EINVAL; + + iopf_for_domain_remove(domain, dev); + cache_tag_unassign_domain(info->domain, dev, IOMMU_NO_PASID); + info->domain_attached =3D false; + + spin_lock_irqsave(&info->domain->lock, flags); + list_del(&info->link); + spin_unlock_irqrestore(&info->domain->lock, flags); + + domain_detach_reattached_iommu(info->domain, iommu); + info->domain =3D NULL; + + return 0; +} + /** * intel_iommu_preserve_device() - Intel IOMMU callback to preserve device= state * @dev: Target device --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pf1-f198.google.com (mail-pf1-f198.google.com [209.85.210.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C4E4F2EB856 for ; Mon, 21 Sep 2026 00:48:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951736; cv=none; b=ESIH3YvIZZTQXYpyPClxobmOjq/sH6YWrv8jsX8gaiQX9B6iQzwVjReSCRLTslwTglTf1VlJA+EJeAnp5VrvlXkcfGOmOOXXaW+wgIvk6SZBueH0rO21ZMWUh2GtBgCsJhvSAEetSDHAN9mkHkOu8cxAl9IFpi9zMhYOxZFyE5c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951736; c=relaxed/simple; bh=bViJJwbrDWXP/RQSx7nYGasTuiaqPBrsxDARAbursxo=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=LjQOMuvHnCBVFF2jUVVzz2KOGTeavgtv5af7VTTzLb9aU9Hy0PMrgZLPAc5BRrkKyU2eyYCAtxu10C16jxjX7+RWQt9kYIwLEXYkmbfa6DDIb7zwxpmj395OZMF2FWd9qPSRrEVke2MmQCTQVbY0/C6SM2eNQMLk3mQMeMvJGWM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=VPhReC8D; arc=none smtp.client-ip=209.85.210.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="VPhReC8D" Received: by mail-pf1-f198.google.com with SMTP id d2e1a72fcca58-86a74698972so2851597b3a.0 for ; Sun, 20 Sep 2026 17:48:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951732; x=1790556532; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=dbOQml2BXHpzMYBrRxpebYsHTJLnApuit/nNvL/j2sY=; b=VPhReC8DkxkRkjb4l1QvHWyX/m8uAp8D6KnSVET9n+bpM9326E/5XahPN3C40XbJzq 0JfZp+rGsjiyb0G0BAzatdj4ArVG0G94lsxoMzaqFVYs4PmSCVvxNDT/WUujJQRWVgoV aOuyyVTFXmlgbJrtogGQIQnGltl9j5Ojmkrn9VHBjx5Qx81Kr4sK/s2vlQ8AKyrEgmMp Qtovcyql1MeUleggbqsGoN8vsLft9JW59eoqiH0rKbvuSzLPtxNsS4eQXP5zpAQL0Pgc gdqp6PI4pNcBULrhwKD1B2ivvWirWy7GcvKtaqBL0SfzaXVnn3RH9XBmyCVhxMskPnif yWkg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951732; x=1790556532; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=dbOQml2BXHpzMYBrRxpebYsHTJLnApuit/nNvL/j2sY=; b=i9B7kPcpuPjKttmIMGWwtBl/LSzwRbqhxrUyIuAPPkb6urEX15ydsLbqGrLzJ0gJeN xVZSa7WDucVbeDMesn4wGEOJDpytZKeQglQ1K3Pi2DxPr4t6G/rp0ntcYTW6ZPYYC2Ri WJytJ4Ggy3+0pwwCn2empOTRvA7KfJrsU+89KV0T54PVPDt8JPSqmYFrDlt2L6+fFD0g 3KOSPZPqmABO32cdjQ5VU/WHzXC6R+oQ6MnHU6ZqgNZpqqFoLIjyCr6tGhxUTyJMwt0j UNXNCbuAzh7cQKslOB8GP1bfrRkJywPvZFEL3u32JhhtxforvR+j/STsqWZ+dHYFV86g Ku5w== X-Forwarded-Encrypted: i=1; AKwUvBzGew+JkffeF5xxSb9QjI3Yx05H1BGjVmNQB5m39vEjFmT3IutNmXS1DYRGEfVbJLAj9bBr4c87SL5nLOI=@vger.kernel.org X-Gm-Message-State: AFuF++m2YQeYsvqIo7qmYGSLE0Pu3Lhd9ws2SHItVdn9/U7249pQQQP0 zfoixgeoXCS6gX/lts6gnKLBbIBc+v65uAVcIiFgBQYZR3EnPubTnGNJdXExLQlyvUZZMvT4yB6 U79IgwpFHD19SOQ== X-Received: from pgax32.prod.google.com ([2002:a05:6a02:2e60:b0:cc4:e5ec:35b9]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:2e83:b0:874:705d:f639 with SMTP id d2e1a72fcca58-874dd9f6be9mr12281098b3a.27.1789951731594; Sun, 20 Sep 2026 17:48:51 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:29 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-14-skhawaja@google.com> Subject: [PATCH v5 13/18] iommu/vt-d: Preserve PASID table of preserved device From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" In scalable mode the PASID table is used to fetch the io page tables. Preserve and restore the PASID table of the preserved devices. Signed-off-by: Samiullah Khawaja --- drivers/iommu/intel/liveupdate.c | 141 +++++++++++++++++++++++++++++-- drivers/iommu/intel/pasid.c | 10 ++- drivers/iommu/intel/pasid.h | 8 ++ include/linux/kho/abi/iommu.h | 17 ++++ 4 files changed, 169 insertions(+), 7 deletions(-) diff --git a/drivers/iommu/intel/liveupdate.c b/drivers/iommu/intel/liveupd= ate.c index 6e6707eefc8c..b9467fb7195f 100644 --- a/drivers/iommu/intel/liveupdate.c +++ b/drivers/iommu/intel/liveupdate.c @@ -14,6 +14,7 @@ #include =20 #include "iommu.h" +#include "pasid.h" #include "../iommu-pages.h" =20 /* 2 tables per bus in scalable mode with upper table at odd bit */ @@ -510,6 +511,69 @@ int intel_iommu_detach_restored_device(struct device *= dev) return 0; } =20 +enum pasid_lu_op { + PASID_LU_OP_PRESERVE =3D 1, + PASID_LU_OP_UNPRESERVE, + PASID_LU_OP_RESTORE, +}; + +static int pasid_lu_do_op(void *table, enum pasid_lu_op op) +{ + int ret =3D 0; + + switch (op) { + case PASID_LU_OP_PRESERVE: + ret =3D iommu_preserve_pages(table); + break; + case PASID_LU_OP_UNPRESERVE: + iommu_unpreserve_pages(table); + break; + case PASID_LU_OP_RESTORE: + iommu_restore_pages(virt_to_phys(table)); + break; + } + + return ret; +} + +static int pasid_lu_handle_pd(struct pasid_dir_entry *dir, + u32 max_pasid, enum pasid_lu_op op) +{ + int max_pde =3D max_pasid >> PASID_PDE_SHIFT; + struct pasid_entry *table; + int i, ret; + + for (i =3D 0; i < max_pde; i++) { + table =3D get_pasid_table_from_pde(&dir[i]); + if (!table) + continue; + + ret =3D pasid_lu_do_op(table, op); + if (ret) + goto err; + } + + ret =3D pasid_lu_do_op(dir, op); + if (ret) + goto err; + + return 0; + +err: + if (op !=3D PASID_LU_OP_PRESERVE) + return ret; + + while (i > 0) { + table =3D get_pasid_table_from_pde(&dir[--i]); + if (!table) + continue; + + pasid_lu_do_op(table, PASID_LU_OP_UNPRESERVE); + } + + return ret; +} + /** * intel_iommu_preserve_device() - Intel IOMMU callback to preserve device= state * @dev: Target device @@ -521,6 +585,7 @@ int intel_iommu_preserve_device(struct device *dev, struct iommu_device_ser *device_ser) { struct device_domain_info *info =3D dev_iommu_priv_get(dev); + struct pasid_table *pasid_table; int ret; =20 if (!dev_is_pci(dev)) { @@ -543,6 +608,22 @@ int intel_iommu_preserve_device(struct device *dev, =20 device_ser->domain_iommu_ser.attachment_id =3D domain_id_iommu(info->doma= in, info->iommu); + + if (!sm_supported(info->iommu)) + return 0; + + pasid_table =3D intel_pasid_get_table(dev); + if (!pasid_table) + return -EINVAL; + + ret =3D pasid_lu_handle_pd(pasid_table->table, + pasid_table->max_pasid, + PASID_LU_OP_PRESERVE); + if (ret) + return ret; + + device_ser->intel.pasid_table =3D virt_to_phys(pasid_table->table); + device_ser->intel.max_pasid =3D pasid_table->max_pasid; return 0; } =20 @@ -554,15 +635,33 @@ int intel_iommu_preserve_device(struct device *dev, void intel_iommu_unpreserve_device(struct device *dev, struct iommu_device_ser *device_ser) { + struct device_domain_info *info =3D dev_iommu_priv_get(dev); + struct pasid_table *pasid_table; + + if (!dev_is_pci(dev)) + return; + + if (!info) + return; + + if (!sm_supported(info->iommu)) + return; + /* * The context tables preserved during device preservation, in the * preserve_device() callback, might be shared with other devices, so - * those are unpreserved in the iommu unpreserve() callback. So this - * callback is kept empty. - * - * Once device PASID tables are preserved, the unpreservation of PASID - * tables will be added here. + * those are unpreserved in the iommu unpreserve() callback. */ + if (!device_ser->intel.pasid_table) + return; + + pasid_table =3D intel_pasid_get_table(dev); + if (!pasid_table) + return; + + pasid_lu_handle_pd(pasid_table->table, + pasid_table->max_pasid, + PASID_LU_OP_UNPRESERVE); } =20 /** @@ -607,3 +706,35 @@ void intel_iommu_unpreserve(struct iommu_device *iommu= _dev, unpreserve_iommu_context_tables(iommu, ser); iommu_unpreserve_pages(iommu->root_entry); } + +/** + * intel_pasid_restore_table() - Restore preserved PASID table for a device + * @dev: Restored device + * @max_pasid: Maximum supported PASID + * + * Return: Pointer to restored PASID table directory, or NULL if not prese= rved. + */ +void *intel_pasid_restore_table(struct device *dev, u64 max_pasid) +{ + struct iommu_device_ser *ser =3D dev_iommu_restored_state(dev); + + if (!ser || !ser->intel.pasid_table) + return NULL; + + /* + * MAX PASID of a device should not change as it is read from + * capabilities. + */ + BUG_ON(ser->intel.max_pasid !=3D max_pasid); + + if (ser->intel.restored) + goto out; + + BUG_ON(pasid_lu_handle_pd(phys_to_virt(ser->intel.pasid_table), + ser->intel.max_pasid, + PASID_LU_OP_RESTORE)); + ser->intel.restored =3D 1; + +out: + return phys_to_virt(ser->intel.pasid_table); +} diff --git a/drivers/iommu/intel/pasid.c b/drivers/iommu/intel/pasid.c index e4f24d3f19a6..59cc69383799 100644 --- a/drivers/iommu/intel/pasid.c +++ b/drivers/iommu/intel/pasid.c @@ -13,6 +13,7 @@ #include #include #include +#include #include #include #include @@ -60,8 +61,13 @@ int intel_pasid_alloc_table(struct device *dev) =20 size =3D max_pasid >> (PASID_PDE_SHIFT - 3); order =3D size ? get_order(size) : 0; - dir =3D iommu_alloc_pages_node_sz(info->iommu->node, GFP_KERNEL, - 1 << (order + PAGE_SHIFT)); + + max_pasid =3D 1 << (order + PAGE_SHIFT + 3); + if (dev_iommu_restored_state(dev)) + dir =3D intel_pasid_restore_table(dev, max_pasid); + else + dir =3D iommu_alloc_pages_node_sz(info->iommu->node, GFP_KERNEL, + 1 << (order + PAGE_SHIFT)); if (!dir) { kfree(pasid_table); return -ENOMEM; diff --git a/drivers/iommu/intel/pasid.h b/drivers/iommu/intel/pasid.h index 48d3bb6b68de..801768cdea16 100644 --- a/drivers/iommu/intel/pasid.h +++ b/drivers/iommu/intel/pasid.h @@ -301,6 +301,14 @@ static inline void pasid_set_eafe(struct pasid_entry *= pe) =20 extern unsigned int intel_pasid_max_id; int intel_pasid_alloc_table(struct device *dev); +#ifdef CONFIG_IOMMU_LIVEUPDATE +void *intel_pasid_restore_table(struct device *dev, u64 max_pasid); +#else +static inline void *intel_pasid_restore_table(struct device *dev, u64 max_= pasid) +{ + return NULL; +} +#endif void intel_pasid_free_table(struct device *dev); struct pasid_table *intel_pasid_get_table(struct device *dev); int intel_pasid_setup_first_level(struct intel_iommu *iommu, struct device= *dev, diff --git a/include/linux/kho/abi/iommu.h b/include/linux/kho/abi/iommu.h index 5aaa29da6832..308cd83fd3e3 100644 --- a/include/linux/kho/abi/iommu.h +++ b/include/linux/kho/abi/iommu.h @@ -129,6 +129,19 @@ struct iommu_dev_map_ser { u64 iommu_phys; } __packed; =20 +/** + * struct iommu_device_intel_ser - Intel specific state of serialized devi= ce + * @restored: Whether the device state is restored + * @pasid_table: Physical address of pasid table + * @max_pasid: Maximum supported pasid + */ +struct iommu_device_intel_ser { + u8 restored; + u8 padding[7]; + u64 pasid_table; + u64 max_pasid; +} __packed; + /** * struct iommu_device_ser - Serialized state of a device * @hdr: Common object header @@ -136,6 +149,7 @@ struct iommu_dev_map_ser { * @pci_domain_nr: PCI domain number * @dma_owner_token: Token to identify the DMA owner of this device * @domain_iommu_ser: Domain and IOMMU mapping + * @intel: Intel specific serialization data */ struct iommu_device_ser { struct iommu_hdr_ser hdr; @@ -143,6 +157,9 @@ struct iommu_device_ser { u32 pci_domain_nr; u64 dma_owner_token; struct iommu_dev_map_ser domain_iommu_ser; + union { + struct iommu_device_intel_ser intel; + }; } __packed; =20 /* There are maximum 256 buses, so maximum 512 context tables */ --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pf1-f200.google.com (mail-pf1-f200.google.com [209.85.210.200]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 00B792F0C7E for ; Mon, 21 Sep 2026 00:48:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.200 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951736; cv=none; b=W0XlYCaqfYT3udT/GthYN6ej40ClLDmKRiWjvK3hf/mudL7zyLB0bXajJkmnfuBOb0SL5zHvHQum4VesVhiG62lZHAAytQVs74wb78lLdTB+R5vyhKD/CYDzzm2PY9CNcti89gh0JFOlsepkjU7ez+na/HxGrhI0VvWA+DAhdsU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951736; c=relaxed/simple; bh=avPZ16PAm2KRLLs3zFSBl7jOowUvnuNq+fuMqUotd3A=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=bAtiItiyn3Fo0BFgLA+Zy3GL3/DoO6zJW3Gk/U1+Jgq00e9HVTeBdS1KTjmvDbsaJMYPuyiToJFUsIrvv8n65O2TZ37TWT+jBnyFSvwNgOS2hrPe/XOsBavcaJ9XpKwE6oJgUDuT9mNIPY1949Zvo+W8h7MFpt+nuxa12yopBmI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=BRfbpAcZ; arc=none smtp.client-ip=209.85.210.200 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="BRfbpAcZ" Received: by mail-pf1-f200.google.com with SMTP id d2e1a72fcca58-86a2639398cso5295365b3a.3 for ; Sun, 20 Sep 2026 17:48:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951733; x=1790556533; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=/sjf1DDcqeGOiFnNg/3dW38k0O9F/mHXt9LSFR/mRoI=; b=BRfbpAcZmOmJDGP4osSCg3PLLoHDJDWTaenp35KhmHipsl2mvDqOW0HJNs/Mz3dnm8 OuTGTXf7g0RDlyfXBs4N0Ojr8OQ5z4LQvslxytFh+xppBK75Hv2VdGPLY35vlU6HBHZJ ghCm3nYmcbctH6qIbfW413jxO9y1e1TtlNssbcpBXYrnLCPYOYmuK4DVLZoOMbmXmrRZ gZK4rwZbbbb3oI/Kl5nCLbIToLF6PBen8ll7VNflqmu0ASesIf6gEd26jbIyUrF1jPs2 xAyY+UNCQUM/RTXWMQimMTY4D/AwWHKdD7rID4c9C06Fj68Xvsze9Q5amlR/kMpMjQoM z2IA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951733; x=1790556533; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=/sjf1DDcqeGOiFnNg/3dW38k0O9F/mHXt9LSFR/mRoI=; b=f8t0V/ADkRD4zHzvdymwAF3vBy6HQhewhZXY53e24O4o2YlLu+OgXTaOwDCpHKS7Tx gxj4xICv4bhBcbPaTUxwgnVqZkAZbgGotDj+YhiMbuOGc5rLiY77WH0gLdYr7oHzfcjl MFJxTnD4bHAWtOfiUiN3u3vpS+Rr2ZPM7eey6aF2yu+X8snzNqoh188wORC4Otfuadlm GCesPoW59LL4N3qoK4E0WkMHcd91YoA1oeupQT7txd70EK+5a2otD4bdXbEF7KSrrftN Kpa2MMHdpYD+81O0dRk9082sTYQv0N/EjvOph3rLh16k3ApLdSUnEuqIQzuS+9jamb8A tN4g== X-Forwarded-Encrypted: i=1; AKwUvBympj6EPe5IDe5gZ4+VjDWdpXFrzISQx/sOiw1zp1lUASrRwZPUT8FcyHsUK3Ryfsewl3D8u43DgCrIf2Y=@vger.kernel.org X-Gm-Message-State: AFuF++mlXf0ExuhBb4abjv2DhZ0QAfwYDZS/y7sG56zA2oGV6OMEmuMp LGoRTl7pHzErcrXSqYDbyVXKgj+XQOZxYpYlViQq2/sPLEq/Ek5NEktc70nOM5XI8mG9mXUqDwb x/sn5bkI1MT5NVg== X-Received: from pgtk20.prod.google.com ([2002:a65:68d4:0:b0:cc5:1024:105e]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:1804:b0:852:131f:b9d2 with SMTP id d2e1a72fcca58-874db8eab76mr12887659b3a.2.1789951732364; Sun, 20 Sep 2026 17:48:52 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:30 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-15-skhawaja@google.com> Subject: [PATCH v5 14/18] iommufd: Implement ioctl to mark HWPT for preservation From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: YiFei Zhu , Pranjal Shrivastava , Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: YiFei Zhu Userspace provides a token to mark the HWPT for preservation. Note that this token is not the LUO token that is used to preserve the iommufd. Once all the required HWPT are marked for preservation, the user can preserve the iommufd into LUO. The iommufd will preserve the HWPTs that are marked for preservation. The marked HWPTs are tracked using a new XArray mark protected by a new liveupdate mutex. This mutex will also be used during iommufd preservation to protect against any race with the mark preserve ioctl. The HWPT token will be used during restore to identify this HWPT. The restoration logic is not implemented and will be added later. Reviewed-by: Pranjal Shrivastava Signed-off-by: YiFei Zhu Signed-off-by: Samiullah Khawaja --- MAINTAINERS | 1 + drivers/iommu/iommufd/Makefile | 1 + drivers/iommu/iommufd/iommufd_private.h | 17 ++++++ drivers/iommu/iommufd/liveupdate.c | 71 +++++++++++++++++++++++++ drivers/iommu/iommufd/main.c | 9 ++++ include/uapi/linux/iommufd.h | 27 ++++++++++ 6 files changed, 126 insertions(+) create mode 100644 drivers/iommu/iommufd/liveupdate.c diff --git a/MAINTAINERS b/MAINTAINERS index ae6df95dc398..d4848ec3a9d6 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -13728,6 +13728,7 @@ M: Samiullah Khawaja R: Pranjal Shrivastava L: iommu@lists.linux.dev S: Maintained +F: drivers/iommu/iommufd/liveupdate.c F: drivers/iommu/liveupdate.c F: include/linux/iommu-liveupdate.h F: include/linux/kho/abi/iommu.h diff --git a/drivers/iommu/iommufd/Makefile b/drivers/iommu/iommufd/Makefile index 67207914bb6e..cfd5b4b14ccb 100644 --- a/drivers/iommu/iommufd/Makefile +++ b/drivers/iommu/iommufd/Makefile @@ -12,6 +12,7 @@ iommufd-y :=3D \ =20 iommufd-$(CONFIG_IOMMUFD_NOIOMMU) +=3D hwpt_noiommu.o iommufd-$(CONFIG_IOMMUFD_TEST) +=3D selftest.o +iommufd-$(CONFIG_IOMMU_LIVEUPDATE) +=3D liveupdate.o =20 obj-$(CONFIG_IOMMUFD) +=3D iommufd.o obj-$(CONFIG_IOMMUFD_DRIVER) +=3D iova_bitmap.o diff --git a/drivers/iommu/iommufd/iommufd_private.h b/drivers/iommu/iommuf= d/iommufd_private.h index eb2e85b27e42..a33b32708afa 100644 --- a/drivers/iommu/iommufd/iommufd_private.h +++ b/drivers/iommu/iommufd/iommufd_private.h @@ -44,6 +44,11 @@ struct iommufd_ctx { struct file *file; struct xarray objects; struct xarray groups; +#ifdef CONFIG_IOMMU_LIVEUPDATE +#define IOMMUFD_OBJ_LIVEUPDATE_MARK XA_MARK_1 + /* @liveupdate_mutex: Protects the preservation of HWPTs. */ + struct mutex liveupdate_mutex; +#endif wait_queue_head_t destroy_wait; struct rw_semaphore ioas_creation_lock; struct maple_tree mt_mmap; @@ -392,6 +397,9 @@ struct iommufd_hwpt_paging { bool auto_domain : 1; bool enforce_cache_coherency : 1; bool nest_parent : 1; +#ifdef CONFIG_IOMMU_LIVEUPDATE + u64 liveupdate_token; +#endif /* Head at iommufd_ioas::hwpt_list */ struct list_head hwpt_item; struct iommufd_sw_msi_maps present_sw_msi; @@ -729,6 +737,15 @@ void iommufd_vdevice_abort(struct iommufd_object *obj); int iommufd_hw_queue_alloc_ioctl(struct iommufd_ucmd *ucmd); void iommufd_hw_queue_destroy(struct iommufd_object *obj); =20 +#ifdef CONFIG_IOMMU_LIVEUPDATE +int iommufd_hwpt_liveupdate_mark_preserve(struct iommufd_ucmd *ucmd); +#else +static inline int iommufd_hwpt_liveupdate_mark_preserve(struct iommufd_ucm= d *ucmd) +{ + return -ENOTTY; +} +#endif + #ifdef CONFIG_IOMMUFD_TEST int iommufd_test(struct iommufd_ucmd *ucmd); void iommufd_selftest_destroy(struct iommufd_object *obj); diff --git a/drivers/iommu/iommufd/liveupdate.c b/drivers/iommu/iommufd/liv= eupdate.c new file mode 100644 index 000000000000..96f01ee5a1e8 --- /dev/null +++ b/drivers/iommu/iommufd/liveupdate.c @@ -0,0 +1,71 @@ +// SPDX-License-Identifier: GPL-2.0-only + +/* + * Copyright (C) 2026, Google LLC + * Author: Samiullah Khawaja + */ + +#define pr_fmt(fmt) "iommufd: " fmt + +#include +#include +#include + +#include "iommufd_private.h" + +int iommufd_hwpt_liveupdate_mark_preserve(struct iommufd_ucmd *ucmd) +{ + struct iommu_hwpt_liveupdate_mark_preserve *cmd =3D ucmd->cmd; + struct iommufd_hwpt_paging *hwpt_target; + struct iommufd_hwpt_paging *hwpt_paging; + struct iommufd_ctx *ictx =3D ucmd->ictx; + struct iommufd_object *obj; + unsigned long index; + bool marked =3D false; + int rc =3D 0; + + hwpt_target =3D iommufd_get_hwpt_paging(ucmd, cmd->hwpt_id); + if (IS_ERR(hwpt_target)) + return PTR_ERR(hwpt_target); + + mutex_lock(&ictx->liveupdate_mutex); + + xa_lock(&ictx->objects); + + /* PRI use cases are not supported. */ + if (hwpt_target->common.fault) { + rc =3D -EOPNOTSUPP; + goto out_unlock; + } + + xa_for_each_marked(&ictx->objects, index, obj, IOMMUFD_OBJ_LIVEUPDATE_MAR= K) { + if (WARN_ON_ONCE(obj->type !=3D IOMMUFD_OBJ_HWPT_PAGING)) + continue; + + hwpt_paging =3D to_hwpt_paging(container_of(obj, struct iommufd_hw_paget= able, obj)); + + if (hwpt_paging =3D=3D hwpt_target) + marked =3D true; + + if (hwpt_paging->liveupdate_token =3D=3D cmd->hwpt_token) { + if (hwpt_paging =3D=3D hwpt_target) + goto out_unlock; + + rc =3D -EADDRINUSE; + goto out_unlock; + } + } + + __xa_set_mark(&ictx->objects, hwpt_target->common.obj.id, IOMMUFD_OBJ_LIV= EUPDATE_MARK); + + if (marked) + pr_warn_ratelimited("Overwriting HWPT liveupdate token from: %llu to %ll= u\n", + hwpt_target->liveupdate_token, cmd->hwpt_token); + hwpt_target->liveupdate_token =3D cmd->hwpt_token; + +out_unlock: + xa_unlock(&ictx->objects); + mutex_unlock(&ictx->liveupdate_mutex); + iommufd_put_object(ictx, &hwpt_target->common.obj); + return rc; +} diff --git a/drivers/iommu/iommufd/main.c b/drivers/iommu/iommufd/main.c index 9a921b153162..fe8610fd58e3 100644 --- a/drivers/iommu/iommufd/main.c +++ b/drivers/iommu/iommufd/main.c @@ -333,6 +333,9 @@ static int iommufd_fops_open(struct inode *inode, struc= t file *filp) init_rwsem(&ictx->ioas_creation_lock); xa_init_flags(&ictx->objects, XA_FLAGS_ALLOC1 | XA_FLAGS_ACCOUNT); xa_init(&ictx->groups); +#ifdef CONFIG_IOMMU_LIVEUPDATE + mutex_init(&ictx->liveupdate_mutex); +#endif ictx->file =3D filp; mt_init_flags(&ictx->mt_mmap, MT_FLAGS_ALLOC_RANGE); init_waitqueue_head(&ictx->destroy_wait); @@ -395,6 +398,9 @@ static int iommufd_fops_release(struct inode *inode, st= ruct file *filp) * iommufd_object_tombstone_user() */ xa_destroy(&ictx->objects); +#ifdef CONFIG_IOMMU_LIVEUPDATE + mutex_destroy(&ictx->liveupdate_mutex); +#endif =20 WARN_ON(!xa_empty(&ictx->groups)); =20 @@ -440,6 +446,7 @@ union ucmd_buffer { struct iommu_hwpt_alloc hwpt; struct iommu_hwpt_get_dirty_bitmap get_dirty_bitmap; struct iommu_hwpt_invalidate cache; + struct iommu_hwpt_liveupdate_mark_preserve mark_preserve; struct iommu_hwpt_set_dirty_tracking set_dirty_tracking; struct iommu_ioas_alloc alloc; struct iommu_ioas_allow_iovas allow_iovas; @@ -516,6 +523,8 @@ static const struct iommufd_ioctl_op iommufd_ioctl_ops[= ] =3D { __reserved), IOCTL_OP(IOMMU_VIOMMU_ALLOC, iommufd_viommu_alloc_ioctl, struct iommu_viommu_alloc, out_viommu_id), + IOCTL_OP(IOMMU_HWPT_LIVEUPDATE_MARK_PRESERVE, iommufd_hwpt_liveupdate_mar= k_preserve, + struct iommu_hwpt_liveupdate_mark_preserve, hwpt_token), #ifdef CONFIG_IOMMUFD_TEST IOCTL_OP(IOMMU_TEST_CMD, iommufd_test, struct iommu_test_cmd, last), #endif diff --git a/include/uapi/linux/iommufd.h b/include/uapi/linux/iommufd.h index 206fa667c782..34637712c873 100644 --- a/include/uapi/linux/iommufd.h +++ b/include/uapi/linux/iommufd.h @@ -58,6 +58,7 @@ enum { IOMMUFD_CMD_VEVENTQ_ALLOC =3D 0x93, IOMMUFD_CMD_HW_QUEUE_ALLOC =3D 0x94, IOMMUFD_CMD_IOAS_NOIOMMU_GET_PA =3D 0x95, + IOMMUFD_CMD_HWPT_LIVEUPDATE_MARK_PRESERVE =3D 0x96, }; =20 /** @@ -1390,4 +1391,30 @@ struct iommu_hw_queue_alloc { __aligned_u64 length; }; #define IOMMU_HW_QUEUE_ALLOC _IO(IOMMUFD_TYPE, IOMMUFD_CMD_HW_QUEUE_ALLOC) + +/** + * struct iommu_hwpt_liveupdate_mark_preserve - ioctl(IOMMU_HWPT_LIVEUPDAT= E_MARK_PRESERVE) + * @size: sizeof(struct iommu_hwpt_liveupdate_mark_preserve) + * @hwpt_id: Iommufd object ID of the target HWPT + * @hwpt_token: Token to identify this hwpt upon restore + * + * The target HWPT will be preserved during iommufd preservation. + * Only file-based memory mappings (e.g. memfd) are supported for HWPTs ma= rked + * for preservation. Mapping anonymous memory into a preserved HWPT will r= esult + * in a failure during the preservation phase. + * + * The hwpt_token is provided by userspace. If userspace enters a token + * already in use within this iommufd, -EADDRINUSE is returned from this i= octl. + * + * Note: There is no 'unmark' operation, so any HWPTs pooled in userspace = that + * are marked for preservation must be destroyed after use. + */ +struct iommu_hwpt_liveupdate_mark_preserve { + __u32 size; + __u32 hwpt_id; + __aligned_u64 hwpt_token; +}; +#define IOMMU_HWPT_LIVEUPDATE_MARK_PRESERVE \ + _IO(IOMMUFD_TYPE, IOMMUFD_CMD_HWPT_LIVEUPDATE_MARK_PRESERVE) + #endif --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9A2D3309EF2 for ; Mon, 21 Sep 2026 00:48:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951740; cv=none; b=KB5surxJ44Njv87BS+yLbGA2evJF5D2bZmYSqnxw5fysVX0IPKtE/7oXW6wF7EXB+RXABCQkujnwRzuxiUWKv9jWTL4QGb830GEUqP/nja/ajf7qkgp7oMBOwnXfafAv12LL0Sk/p6Pta/tTRPYTLL/4ZZiALnupe7M02pLZH04= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951740; c=relaxed/simple; bh=pqWxT3X1wBWecwwHNFQTmz/pTVpYDfQPUJrD79ZLb8w=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=jkFmJmU59EsMYS49lmPwZAOH3y/NQ76/BSBo12PUKM8z65/3B/M1inXOlt36e9fq3iLeN1f36pkJbUdN1bi7g7UI1V/HugLequhsIBiUvsh2F4Oj5/qo9VTfcnR59xBvOQIk+N3zaEeyffJoXlnn41+BtS2epXEohc7aRhUS9cc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=Y9yWSGPj; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="Y9yWSGPj" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2dc92350888so44556155ad.3 for ; Sun, 20 Sep 2026 17:48:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951733; x=1790556533; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=N0C1DXDyP/j3IlwYvr+0LSV2v1KgQLOMoBjsxRGMMe8=; b=Y9yWSGPjrS6F88SiEpNx+VDiySqsuQKMubP2jjUEcRaRfic1C0nxDaoaJD3qB4+nr/ KFzkLwk4xhS3sGfAM9A0KZnPcM28oS7L4AoghwgZPEkmo+lqkTBsSoZP3O+ZJS4FHFEg q5ZrL2N+pZbuB6wO72imvXOxj8GW5Q1buAj7Iw4H8e5EjiZPV0hTDFs9QxP0lld7UxIs Xpj7ORriIgSZuPj4EQr73B63ZQ81Rp+lN7EFm6Y6ixU0AL1ayWTMEbJw3tfsV/rI8KuI HbgNqzOna2GqSg7W4ICBmLvfyFqZKgOgRVJPd/fFQPTXt1QgD24SB9Z4DY3IfuRx+O4E G6Eg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951733; x=1790556533; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=N0C1DXDyP/j3IlwYvr+0LSV2v1KgQLOMoBjsxRGMMe8=; b=Ic/j5EH+oxv/5k17ZnfRedWDrqgzQCB8ODZc75GuD40N0MkYYnA13Nl74MdxPlUiPF +MxlnYXWbB83K9r3f3ySmXEqT/VK5Xn3DdYw7WIms4eNbEK5jHn7TAXsdMwuxfmw4haF eF50upT/5gL/gThlVlb2E3s8ogQU6TCWcm12DM7HjvmHak73QEU/MFnNQhCdK2277zk+ RYhcv1HephUx6TyEv6cLib3/Qf4b6XE7Cas9dunmvYG2gs+fM6lERE+gtMewqMdPXDFa 2R53Tbg/xKAZP39cIQBNVrpXhN762nQsx/k5KLuKVMRchKRdCdXICXjI7FSB+knH/dCv mk6w== X-Forwarded-Encrypted: i=1; AKwUvBxlTrdBHHmBFW/XwqJN/BAoihjmgWf5tWDjg+OZK8HKFabcg1Qm41+5vZCVau9PJPPiJZQiTypboTktF/g=@vger.kernel.org X-Gm-Message-State: AFuF++l+POLvkWcQZyFQZlafnu0RZ64lhtOihWOw2E47ucs+DzkIcUGA 4gGau9IjRZlR3oRAsbAOfX9jTzs10489roBbdFofrbcjQFjv967W08Pcys9NstMjc+8lIaNFK3k uo/MPkFmp4Fb8Xg== X-Received: from plmg19.prod.google.com ([2002:a17:903:3cd3:b0:2dd:b156:5a6e]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:2445:b0:2dd:c100:a5e9 with SMTP id d9443c01a7336-2ddc100a6c4mr73544375ad.61.1789951733301; Sun, 20 Sep 2026 17:48:53 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:31 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-16-skhawaja@google.com> Subject: [PATCH v5 15/18] iommufd: Persist iommu hardware pagetables for live update From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: YiFei Zhu , Samiullah Khawaja , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: YiFei Zhu Register iommufd with the LUO framework and implement the preserve and unpreserve ops to save marked HWPTs. To make sure mappings do not change during preserved state, add a liveupdate_immutable flag to IOAS. When an HWPT is preserved, its IOAS is marked immutable and any map/unmap attempts will fail with -EBUSY. This is synchronized using the domains_rwsem to prevent races with concurrent mapping operations. The preserve callback iterates over the marked HWPTs, verifies that the backing memory pages are preserved, and calls iommu_preserve_domain() to preserve the associated IOMMU domain. Signed-off-by: YiFei Zhu Signed-off-by: Samiullah Khawaja --- MAINTAINERS | 1 + drivers/iommu/iommufd/io_pagetable.c | 11 + drivers/iommu/iommufd/io_pagetable.h | 1 + drivers/iommu/iommufd/iommufd_private.h | 26 ++ drivers/iommu/iommufd/liveupdate.c | 308 ++++++++++++++++++++++++ drivers/iommu/iommufd/main.c | 10 +- drivers/iommu/iommufd/pages.c | 11 + include/linux/kho/abi/iommufd.h | 51 ++++ 8 files changed, 418 insertions(+), 1 deletion(-) create mode 100644 include/linux/kho/abi/iommufd.h diff --git a/MAINTAINERS b/MAINTAINERS index d4848ec3a9d6..058a227e99d9 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -13732,6 +13732,7 @@ F: drivers/iommu/iommufd/liveupdate.c F: drivers/iommu/liveupdate.c F: include/linux/iommu-liveupdate.h F: include/linux/kho/abi/iommu.h +F: include/linux/kho/abi/iommufd.h =20 IOMMUFD M: Jason Gunthorpe diff --git a/drivers/iommu/iommufd/io_pagetable.c b/drivers/iommu/iommufd/i= o_pagetable.c index 4e447ce74cf6..bad302e9e282 100644 --- a/drivers/iommu/iommufd/io_pagetable.c +++ b/drivers/iommu/iommufd/io_pagetable.c @@ -384,6 +384,11 @@ int iopt_map_pages(struct io_pagetable *iopt, struct l= ist_head *pages_list, return rc; =20 down_read(&iopt->domains_rwsem); + if (iopt_liveupdate_immutable(iopt)) { + rc =3D -EBUSY; + goto out_unlock_domains; + } + rc =3D iopt_fill_domains_pages(pages_list); if (rc) goto out_unlock_domains; @@ -755,6 +760,12 @@ static int iopt_unmap_iova_range(struct io_pagetable *= iopt, unsigned long start, again: down_read(&iopt->domains_rwsem); down_write(&iopt->iova_rwsem); + + if (iopt_liveupdate_immutable(iopt)) { + rc =3D -EBUSY; + goto out_unlock_iova; + } + while ((area =3D iopt_area_iter_first(iopt, start, last))) { unsigned long area_last =3D iopt_area_last_iova(area); unsigned long area_first =3D iopt_area_iova(area); diff --git a/drivers/iommu/iommufd/io_pagetable.h b/drivers/iommu/iommufd/i= o_pagetable.h index 27e3e311d395..207ff368d412 100644 --- a/drivers/iommu/iommufd/io_pagetable.h +++ b/drivers/iommu/iommufd/io_pagetable.h @@ -234,6 +234,7 @@ struct iopt_pages { struct { /* IOPT_ADDRESS_FILE */ struct file *file; unsigned long start; + u32 seals; }; /* IOPT_ADDRESS_DMABUF */ struct iopt_pages_dmabuf dmabuf; diff --git a/drivers/iommu/iommufd/iommufd_private.h b/drivers/iommu/iommuf= d/iommufd_private.h index a33b32708afa..a4ddc29ec5ae 100644 --- a/drivers/iommu/iommufd/iommufd_private.h +++ b/drivers/iommu/iommufd/iommufd_private.h @@ -98,6 +98,9 @@ struct io_pagetable { /* IOVA that cannot be allocated, struct iopt_reserved */ struct rb_root_cached reserved_itree; u8 disable_large_pages; +#ifdef CONFIG_IOMMU_LIVEUPDATE + u32 nr_preserved_domains; +#endif unsigned long iova_alignment; }; =20 @@ -398,6 +401,7 @@ struct iommufd_hwpt_paging { bool enforce_cache_coherency : 1; bool nest_parent : 1; #ifdef CONFIG_IOMMU_LIVEUPDATE + bool liveupdate_preserved; u64 liveupdate_token; #endif /* Head at iommufd_ioas::hwpt_list */ @@ -738,12 +742,34 @@ int iommufd_hw_queue_alloc_ioctl(struct iommufd_ucmd = *ucmd); void iommufd_hw_queue_destroy(struct iommufd_object *obj); =20 #ifdef CONFIG_IOMMU_LIVEUPDATE +int iommufd_liveupdate_register(void); +void iommufd_liveupdate_unregister(void); + int iommufd_hwpt_liveupdate_mark_preserve(struct iommufd_ucmd *ucmd); + +static inline bool iopt_liveupdate_immutable(const struct io_pagetable *io= pt) +{ + return iopt->nr_preserved_domains > 0; +} #else +static inline int iommufd_liveupdate_register(void) +{ + return 0; +} + +static inline void iommufd_liveupdate_unregister(void) +{ +} + static inline int iommufd_hwpt_liveupdate_mark_preserve(struct iommufd_ucm= d *ucmd) { return -ENOTTY; } + +static inline bool iopt_liveupdate_immutable(const struct io_pagetable *io= pt) +{ + return false; +} #endif =20 #ifdef CONFIG_IOMMUFD_TEST diff --git a/drivers/iommu/iommufd/liveupdate.c b/drivers/iommu/iommufd/liv= eupdate.c index 96f01ee5a1e8..423369e3a8af 100644 --- a/drivers/iommu/iommufd/liveupdate.c +++ b/drivers/iommu/iommufd/liveupdate.c @@ -9,9 +9,31 @@ =20 #include #include +#include +#include #include +#include +#include +#include =20 #include "iommufd_private.h" +#include "io_pagetable.h" + +static bool ioas_set_immutable(struct iommufd_ioas *ioas, bool set) +{ + bool was_immutable; + + down_write(&ioas->iopt.domains_rwsem); + was_immutable =3D ioas->iopt.nr_preserved_domains > 0; + if (set) + ioas->iopt.nr_preserved_domains++; + else if (!WARN_ON(!was_immutable)) + ioas->iopt.nr_preserved_domains--; + + up_write(&ioas->iopt.domains_rwsem); + + return was_immutable; +} =20 int iommufd_hwpt_liveupdate_mark_preserve(struct iommufd_ucmd *ucmd) { @@ -69,3 +91,289 @@ int iommufd_hwpt_liveupdate_mark_preserve(struct iommuf= d_ucmd *ucmd) iommufd_put_object(ictx, &hwpt_target->common.obj); return rc; } + +static int check_iopt_pages_preserved(struct liveupdate_session *s, + struct iommufd_hwpt_paging *hwpt) +{ + u32 req_seals =3D F_SEAL_SEAL | F_SEAL_GROW | F_SEAL_SHRINK; + struct iopt_area *area; + int ret =3D 0; + + down_read(&hwpt->ioas->iopt.iova_rwsem); + for (area =3D iopt_area_iter_first(&hwpt->ioas->iopt, 0, ULONG_MAX); area; + area =3D iopt_area_iter_next(area, 0, ULONG_MAX)) { + struct iopt_pages *pages =3D area->pages; + + if (!pages) + continue; + + /* Only allow file based mapping */ + if (pages->type !=3D IOPT_ADDRESS_FILE) { + ret =3D -EINVAL; + break; + } + + /* + * When this memory file was mapped it should be sealed and seal + * should be sealed. This means that since mapping was done the + * memory file was not grown or shrink and the pages being used + * until now remain pinned and preserved. + */ + if ((pages->seals & req_seals) !=3D req_seals) { + ret =3D -EINVAL; + break; + } + + /* Make sure that the file was preserved. */ + ret =3D liveupdate_get_token_outgoing(s, pages->file, NULL); + if (ret) + break; + } + up_read(&hwpt->ioas->iopt.iova_rwsem); + + return ret; +} + +static int iommufd_preserve_hwpt(struct iommufd_hwpt_paging *hwpt, + struct iommufd_hwpt_ser *hwpt_ser, + struct liveupdate_session *session) +{ + struct iommu_domain_ser *domain_ser; + bool was_immutable; + int rc; + + /* + * Make IOAS immutable so the DMA mappings do not change while + * the HWPT is preserved. Since one IOAS can have multiple + * HWPTs, if an error occurs this call needs to make the IOAS + * mutable again if it was the one that made it immutable. + */ + was_immutable =3D ioas_set_immutable(hwpt->ioas, true); + + if (!was_immutable) { + rc =3D check_iopt_pages_preserved(session, hwpt); + if (rc) + goto err; + } + + hwpt_ser->token =3D hwpt->liveupdate_token; + hwpt_ser->reclaimed =3D false; + + rc =3D iommu_preserve_domain(hwpt->common.domain, &domain_ser); + if (rc < 0) + goto err; + + hwpt_ser->domain_data =3D virt_to_phys(domain_ser); + return 0; + +err: + ioas_set_immutable(hwpt->ioas, false); + return rc; +} + +static void _iommufd_unpreserve(struct iommufd_ctx *ictx, + struct iommufd_ser *ser) +{ + struct iommufd_hwpt_paging *hwpt; + struct iommufd_object *obj; + unsigned long index; + + xa_lock(&ictx->objects); + xa_for_each_marked(&ictx->objects, index, obj, IOMMUFD_OBJ_LIVEUPDATE_MAR= K) { + if (obj->type !=3D IOMMUFD_OBJ_HWPT_PAGING) + continue; + + hwpt =3D to_hwpt_paging(container_of(obj, struct iommufd_hw_pagetable, o= bj)); + if (!hwpt->liveupdate_preserved) + continue; + + xa_unlock(&ictx->objects); + + iommu_unpreserve_domain(hwpt->common.domain); + ioas_set_immutable(hwpt->ioas, false); + + hwpt->liveupdate_preserved =3D false; + iommufd_put_object(ictx, obj); + + xa_lock(&ictx->objects); + } + xa_unlock(&ictx->objects); + + kho_unpreserve_free(ser); +} + +static int iommufd_liveupdate_preserve(struct liveupdate_file_op_args *arg= s) +{ + struct iommufd_ctx *ictx; + struct iommufd_hwpt_paging *hwpt; + struct iommufd_ser *iommufd_ser; + struct iommufd_object *obj; + unsigned int nr_hwpts; + unsigned long index; + unsigned int i; + void *mem; + int rc; + + ictx =3D iommufd_ctx_from_file(args->file); + if (IS_ERR(ictx)) + return PTR_ERR(ictx); + + mutex_lock(&ictx->liveupdate_mutex); + + /* Count the number of HWPTs to preserve */ + nr_hwpts =3D 0; + xa_lock(&ictx->objects); + xa_for_each_marked(&ictx->objects, index, obj, IOMMUFD_OBJ_LIVEUPDATE_MAR= K) { + if (obj->type !=3D IOMMUFD_OBJ_HWPT_PAGING) + continue; + + hwpt =3D to_hwpt_paging(container_of(obj, struct iommufd_hw_pagetable, o= bj)); + if (!hwpt->common.domain) { + rc =3D -EINVAL; + xa_unlock(&ictx->objects); + goto out_unlock; + } + nr_hwpts++; + } + xa_unlock(&ictx->objects); + + mem =3D kho_alloc_preserve(struct_size(iommufd_ser, + hwpt_array, nr_hwpts)); + if (IS_ERR(mem)) { + rc =3D PTR_ERR(mem); + goto out_unlock; + } + + iommufd_ser =3D mem; + iommufd_ser->nr_hwpts =3D nr_hwpts; + + /* Preserve HWPTs */ + i =3D 0; + xa_lock(&ictx->objects); + xa_for_each_marked(&ictx->objects, index, obj, IOMMUFD_OBJ_LIVEUPDATE_MAR= K) { + if (obj->type !=3D IOMMUFD_OBJ_HWPT_PAGING) + continue; + + if (!iommufd_lock_obj(obj)) { + rc =3D -ENOENT; + xa_unlock(&ictx->objects); + goto out_unpreserve; + } + + /* + * HWPT is locked so it will not be destroyed. The xarray lock + * can be released here before preserving the HWPT. + */ + xa_unlock(&ictx->objects); + hwpt =3D to_hwpt_paging(container_of(obj, struct iommufd_hw_pagetable, o= bj)); + rc =3D iommufd_preserve_hwpt(hwpt, &iommufd_ser->hwpt_array[i++], args->= session); + if (rc) { + iommufd_put_object(ictx, obj); + goto out_unpreserve; + } + + /* + * Mark the HWPT as successfully preserved. This is distinct + * from IOMMUFD_OBJ_LIVEUPDATE_MARK, which only indicates the + * userspace intent to preserve. + */ + hwpt->liveupdate_preserved =3D true; + xa_lock(&ictx->objects); + } + xa_unlock(&ictx->objects); + + /* Store the actual number of HWPTs that are preserved */ + iommufd_ser->nr_hwpts =3D i; + + args->serialized_data =3D virt_to_phys(iommufd_ser); + mutex_unlock(&ictx->liveupdate_mutex); + iommufd_ctx_put(ictx); + return 0; + +out_unpreserve: + _iommufd_unpreserve(ictx, iommufd_ser); +out_unlock: + mutex_unlock(&ictx->liveupdate_mutex); + iommufd_ctx_put(ictx); + return rc; +} + +static void iommufd_liveupdate_unpreserve(struct liveupdate_file_op_args *= args) +{ + struct iommufd_ctx *ictx; + + ictx =3D iommufd_ctx_from_file(args->file); + if (WARN_ON(IS_ERR(ictx))) + return; + + mutex_lock(&ictx->liveupdate_mutex); + _iommufd_unpreserve(ictx, phys_to_virt(args->serialized_data)); + mutex_unlock(&ictx->liveupdate_mutex); + + iommufd_ctx_put(ictx); +} + +static int iommufd_liveupdate_retrieve(struct liveupdate_file_op_args *arg= s) +{ + return -EOPNOTSUPP; +} + +static bool iommufd_liveupdate_can_finish(struct liveupdate_file_op_args *= args) +{ + return false; +} + +static void iommufd_liveupdate_finish(struct liveupdate_file_op_args *args) +{ +} + +static bool iommufd_liveupdate_can_preserve(struct liveupdate_file_handler= *handler, + struct file *file) +{ + struct iommufd_ctx *ictx =3D iommufd_ctx_from_file(file); + + if (IS_ERR(ictx)) + return false; + + iommufd_ctx_put(ictx); + return true; +} + +static struct liveupdate_file_ops iommufd_ser_file_ops =3D { + .can_preserve =3D iommufd_liveupdate_can_preserve, + .preserve =3D iommufd_liveupdate_preserve, + .unpreserve =3D iommufd_liveupdate_unpreserve, + .retrieve =3D iommufd_liveupdate_retrieve, + .can_finish =3D iommufd_liveupdate_can_finish, + .finish =3D iommufd_liveupdate_finish, +}; + +static struct liveupdate_file_handler iommufd_ser_handler =3D { + .compatible =3D IOMMUFD_LUO_COMPATIBLE, + .ops =3D &iommufd_ser_file_ops, +}; + +int iommufd_liveupdate_register(void) +{ + int ret; + + ret =3D liveupdate_register_file_handler(&iommufd_ser_handler); + if (ret) + goto out; + + ret =3D iommu_liveupdate_register_flb(&iommufd_ser_handler); + if (ret) + liveupdate_unregister_file_handler(&iommufd_ser_handler); + +out: + if (ret =3D=3D -EOPNOTSUPP) + return 0; + + return ret; +} + +void iommufd_liveupdate_unregister(void) +{ + iommu_liveupdate_unregister_flb(&iommufd_ser_handler); + liveupdate_unregister_file_handler(&iommufd_ser_handler); +} diff --git a/drivers/iommu/iommufd/main.c b/drivers/iommu/iommufd/main.c index fe8610fd58e3..86a578fb8e18 100644 --- a/drivers/iommu/iommufd/main.c +++ b/drivers/iommu/iommufd/main.c @@ -805,11 +805,18 @@ static int __init iommufd_init(void) if (ret) goto err_misc; } - ret =3D iommufd_test_init(); + + ret =3D iommufd_liveupdate_register(); if (ret) goto err_vfio_misc; + + ret =3D iommufd_test_init(); + if (ret) + goto err_liveupdate; return 0; =20 +err_liveupdate: + iommufd_liveupdate_unregister(); err_vfio_misc: if (IS_ENABLED(CONFIG_IOMMUFD_VFIO_CONTAINER)) misc_deregister(&vfio_misc_dev); @@ -821,6 +828,7 @@ static int __init iommufd_init(void) static void __exit iommufd_exit(void) { iommufd_test_exit(); + iommufd_liveupdate_unregister(); if (IS_ENABLED(CONFIG_IOMMUFD_VFIO_CONTAINER)) misc_deregister(&vfio_misc_dev); misc_deregister(&iommu_misc_dev); diff --git a/drivers/iommu/iommufd/pages.c b/drivers/iommu/iommufd/pages.c index 404f31d8f729..08e7b882e467 100644 --- a/drivers/iommu/iommufd/pages.c +++ b/drivers/iommu/iommufd/pages.c @@ -55,6 +55,7 @@ #include #include #include +#include #include =20 #include "double_span.h" @@ -1421,6 +1422,7 @@ struct iopt_pages *iopt_alloc_file_pages(struct file = *file, =20 { struct iopt_pages *pages; + int seals; =20 pages =3D iopt_alloc_pages(start_byte, length, writable); if (IS_ERR(pages)) @@ -1428,6 +1430,15 @@ struct iopt_pages *iopt_alloc_file_pages(struct file= *file, pages->file =3D get_file(file); pages->start =3D start - start_byte; pages->type =3D IOPT_ADDRESS_FILE; + + /* + * Get seals from the memfd during mapping to verify that these did not + * change before iommufd preservation. + */ + seals =3D memfd_get_seals(file); + if (seals > 0) + pages->seals =3D seals; + return pages; } =20 diff --git a/include/linux/kho/abi/iommufd.h b/include/linux/kho/abi/iommuf= d.h new file mode 100644 index 000000000000..e0c13b965cb9 --- /dev/null +++ b/include/linux/kho/abi/iommufd.h @@ -0,0 +1,51 @@ +/* SPDX-License-Identifier: GPL-2.0 */ + +/* + * Copyright (C) 2026, Google LLC + * Author: Samiullah Khawaja + */ + +#ifndef _LINUX_KHO_ABI_IOMMUFD_H +#define _LINUX_KHO_ABI_IOMMUFD_H + +#include +#include +#include + +/** + * DOC: IOMMUFD Live Update ABI + * + * This header defines the ABI for preserving the state of an IOMMUFD file + * across a kexec reboot using LUO. + * + * This interface is a contract. Any modification to any of the serializat= ion + * structs defined here constitutes a breaking change. Such changes require + * incrementing the version number in the IOMMUFD_LUO_COMPATIBLE string. + */ + +#define IOMMUFD_LUO_COMPATIBLE "iommufd-v1" + +/** + * struct iommu_hwpt_ser - IOMMUFD HWPT serialized state + * @domain_data: Physical address of the serialized state of associated do= main + * @token: User provided token + * @reclaimed: Whether the HWPT is reclaimed + */ +struct iommufd_hwpt_ser { + u64 domain_data; + u64 token; + u8 reclaimed; + u8 padding[7]; +} __packed; + +/** + * struct iommu_ser - IOMMUFD serialized state + * @nr_hwpts: Number of preserved HWPTs + * @hwpt_array: Array of serialized state of preserved HWPTs + */ +struct iommufd_ser { + u64 nr_hwpts; + struct iommufd_hwpt_ser hwpt_array[]; +} __packed; + +#endif /* _LINUX_KHO_ABI_IOMMUFD_H */ --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pg1-f199.google.com (mail-pg1-f199.google.com [209.85.215.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 214D52E7373 for ; Mon, 21 Sep 2026 00:48:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951739; cv=none; b=NcD7yjxR6emMTrE09qXx+JE2XzZEoyqv+hYf6V4ziR6ygmDdcp1Y8Sheouaij2gUeYRHIf90rLQrdInh6FVZVbHv2IY0SAMW/4rG0MkhozJ+z1uyqoQ2XTiIesL4aAET08bX5ELkmI6CS1vF5Z5nPvMz5B5l6AZr+ybRc61LaVc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951739; c=relaxed/simple; bh=g8GBgYiSWVkD7I1S+JEkWqmDKKAYErKNj/fk09dMpuk=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=gFaIMw0GmwO2ZCD9AFl4Sx/NfA01/UekmgOAdCgtMgk57A9pNtANNYIBQLkzR5fcc1cju7hf4o5zUSX9EZJgrEHO1by7zQnUsmNvQGquyt0mUzLIcfPFqMx9J5DysqnikKofdX4XUvSnybXUsF5eqIFIeUDl9EKWFCWIcoRyux0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=TSvzMSvy; arc=none smtp.client-ip=209.85.215.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="TSvzMSvy" Received: by mail-pg1-f199.google.com with SMTP id 41be03b00d2f7-cc4e496e5c3so1988728a12.0 for ; Sun, 20 Sep 2026 17:48:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951734; x=1790556534; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=QozPFrINpXQ5Eg89NebStO9QklWGOH83J1QrPXjzAKc=; b=TSvzMSvykjp7KIc4V6Y9j5VAyA5dquJjubYPRmvvdpuwZaWm5lnGE7W7KiLlWemKEk rkRukxVOXV2W2QsVgeu0zIMmkiBnbAtqK+XB8ftIRV+ZA4Pd2cCbInNclwkXpWLzMFCZ 3ImyOpPCBysFAJjv9fPruZMSOfqkXB0EBvi78f8FhoGCa9NzzWZC+oulvrRteUnFA5ja kK/wYG0Ew1fxxmF76qEey0Z6IexX0BM6ZCtBmxKjkoIVKUTyFFTIEvSQPMdEaG+aQ3Le 91JssYLkN8YfaEWCy9GWr/e+P3krpmFR7mFxz+67hW3MoexfE5k9EmeRWd68RgF8RZNO VB6g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951734; x=1790556534; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=QozPFrINpXQ5Eg89NebStO9QklWGOH83J1QrPXjzAKc=; b=Eebxj1SKk6ywROfC3qRbyG7neqvRtCvUs1+dRAet/vRFtVF/ZSIuUEoTpHWnVKZVx5 zqQWi7aLM0hlKnhtpo/3oeWyz7Y7+zullBg36LVaF2m/K1Zsbmh4NSGkiqkPHMMSVmRj 3wbY6P1ePLx5oqe1DI/IZ3e6p9s/LFwCVz36wlQljH8Wo6EHRb/FLALEyl8ee+CPNd+K MlKfh+40/bpA4Y4mLtbcKdeLSY2Lrp8faGpj7ol4CVLYD6CeHqo6akDJE5IvqGi1E5aq 4VGSmTdc70wBtOmVdg/4w6ISBiLAlRfvzJYsCqiYMTBwTd3ROtJBjDA+klbYSOhVXeo5 BLsA== X-Forwarded-Encrypted: i=1; AKwUvByVJsDtEDFP32IdRq3apw/khdiaYYq5YNY0vTbAtflFwhnzB5cmGoYRMT8XgnDaMni+1sXO4RJ35C9c/hk=@vger.kernel.org X-Gm-Message-State: AFuF++lUq8iuaiHHI0hBP97J5QBtKhzmFeFt1s3q1+IdO16w0U1cmXdZ OOl0cgz6Q2GozPvHRZqeIJOJKJ5+SkFevEi/QgAzMDqsQqr/RbiOrSXViUaOTpie1Q8RjNPA+LT Vy7q3XO/lnvIBRA== X-Received: from pgvm9.prod.google.com ([2002:a65:62c9:0:b0:cc5:284c:de35]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a20:2586:b0:3dd:85aa:452a with SMTP id adf61e73a8af0-3dd8c53aa87mr15058447637.37.1789951734021; Sun, 20 Sep 2026 17:48:54 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:32 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-17-skhawaja@google.com> Subject: [PATCH v5 16/18] iommufd: Add APIs to preserve/unpreserve a vfio cdev From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Pranjal Shrivastava , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add APIs that can be used to preserve and unpreserve a vfio cdev. Use the APIs exported by the IOMMU core to preserve/unpreserve device. The LUO token of the preserved iommufd is fetched and returned back to the caller as that can be used during restore to get the restored iommufd. Reviewed-by: Pranjal Shrivastava Signed-off-by: Samiullah Khawaja --- drivers/iommu/iommufd/device.c | 142 ++++++++++++++++++++++++ drivers/iommu/iommufd/iommufd_private.h | 6 + include/linux/iommufd.h | 29 +++++ 3 files changed, 177 insertions(+) diff --git a/drivers/iommu/iommufd/device.c b/drivers/iommu/iommufd/device.c index 5c4b06eda546..d4974238ba9a 100644 --- a/drivers/iommu/iommufd/device.c +++ b/drivers/iommu/iommufd/device.c @@ -2,6 +2,7 @@ /* Copyright (c) 2021-2022, NVIDIA CORPORATION & AFFILIATES */ #include +#include #include #include #include @@ -686,6 +687,10 @@ int iommufd_hw_pagetable_attach(struct iommufd_hw_page= table *hwpt, int rc; =20 mutex_lock(&igroup->lock); + if (iommufd_device_is_preserved(idev)) { + rc =3D -EBUSY; + goto err_unlock; + } =20 attach =3D xa_cmpxchg(&igroup->pasid_attach, pasid, NULL, XA_ZERO_ENTRY, GFP_KERNEL); @@ -1747,3 +1752,140 @@ int iommufd_get_hw_info(struct iommufd_ucmd *ucmd) iommufd_put_object(ucmd->ictx, &idev->obj); return rc; } + +#ifdef CONFIG_IOMMU_LIVEUPDATE +static bool _iommufd_device_has_pasid_attachments(struct iommufd_device *i= dev) +{ + struct iommufd_group *igroup =3D idev->igroup; + unsigned long start =3D IOMMU_NO_PASID; + + if (xa_find_after(&igroup->pasid_attach, + &start, UINT_MAX, XA_PRESENT)) + return true; + + return false; +} + +/** + * iommufd_device_preserve() - Preserve an iommufd device across live upda= te + * @s: Live update session + * @idev: Target iommufd device + * @iommufd_tokenp: Pointer to store outgoing iommufd token + * + * Return: 0 on success, or negative error code. + */ +int iommufd_device_preserve(struct liveupdate_session *s, + struct iommufd_device *idev, + u64 *iommufd_tokenp) +{ + struct iommufd_hwpt_paging *hwpt_paging; + struct iommufd_hw_pagetable *hwpt; + struct iommufd_attach *attach; + struct iommufd_group *igroup; + int ret; + + if (!idev) + return -EINVAL; + + igroup =3D idev->igroup; + mutex_lock(&igroup->lock); + if (idev->liveupdate_preserved) { + ret =3D -EBUSY; + goto out; + } + + if (_iommufd_device_has_pasid_attachments(idev)) { + ret =3D -EOPNOTSUPP; + goto out; + } + + attach =3D xa_load(&igroup->pasid_attach, IOMMU_NO_PASID); + if (!attach) { + ret =3D -ENOENT; + goto out; + } + + if (!xa_load(&attach->device_array, idev->obj.id)) { + ret =3D -ENOENT; + goto out; + } + + hwpt =3D attach->hwpt; + hwpt_paging =3D find_hwpt_paging(hwpt); + if (!hwpt_paging || !hwpt_paging->liveupdate_preserved) { + ret =3D -EINVAL; + goto out; + } + + ret =3D liveupdate_get_token_outgoing(s, idev->ictx->file, iommufd_tokenp= ); + if (ret) + goto out; + + ret =3D iommu_preserve_device(hwpt_paging->common.domain, + idev->dev, + *iommufd_tokenp); + + if (!ret) { + igroup->nr_liveupdate_preserved++; + idev->liveupdate_preserved =3D true; + } +out: + mutex_unlock(&igroup->lock); + return ret; +} +EXPORT_SYMBOL_NS_GPL(iommufd_device_preserve, "IOMMUFD"); + +/** + * iommufd_device_unpreserve() - Unpreserve an iommufd device + * @s: Live update session + * @idev: Target iommufd device + */ +void iommufd_device_unpreserve(struct liveupdate_session *s, + struct iommufd_device *idev) +{ + struct iommufd_hwpt_paging *hwpt_paging; + struct iommufd_hw_pagetable *hwpt; + struct iommufd_attach *attach; + struct iommufd_group *igroup; + + if (!idev) + return; + + igroup =3D idev->igroup; + mutex_lock(&igroup->lock); + if (!idev->liveupdate_preserved) + goto out; + + attach =3D xa_load(&igroup->pasid_attach, IOMMU_NO_PASID); + if (!attach) { + WARN(1, "IOMMU_NO_PASID attachment not found"); + goto out; + } + + hwpt =3D attach->hwpt; + hwpt_paging =3D find_hwpt_paging(hwpt); + if (!hwpt_paging || !hwpt_paging->liveupdate_preserved) { + WARN(1, "Attached domain is not preserved"); + goto out; + } + + iommu_unpreserve_device(hwpt_paging->common.domain, idev->dev); + igroup->nr_liveupdate_preserved--; + idev->liveupdate_preserved =3D false; +out: + mutex_unlock(&igroup->lock); +} +EXPORT_SYMBOL_NS_GPL(iommufd_device_unpreserve, "IOMMUFD"); + +/** + * iommufd_device_is_preserved() - Check if an iommufd device is preserved + * @idev: Target iommufd device + * + * Return: true if preserved, false otherwise. + */ +bool iommufd_device_is_preserved(struct iommufd_device *idev) +{ + return idev && idev->igroup && idev->igroup->nr_liveupdate_preserved; +} +EXPORT_SYMBOL_NS_GPL(iommufd_device_is_preserved, "IOMMUFD"); +#endif diff --git a/drivers/iommu/iommufd/iommufd_private.h b/drivers/iommu/iommuf= d/iommufd_private.h index a4ddc29ec5ae..131544626bf2 100644 --- a/drivers/iommu/iommufd/iommufd_private.h +++ b/drivers/iommu/iommufd/iommufd_private.h @@ -507,6 +507,9 @@ struct iommufd_group { struct xarray pasid_attach; struct iommufd_sw_msi_maps required_sw_msi; phys_addr_t sw_msi_start; +#ifdef CONFIG_IOMMU_LIVEUPDATE + int nr_liveupdate_preserved; +#endif }; =20 /* @@ -524,6 +527,9 @@ struct iommufd_device { bool enforce_cache_coherency; struct iommufd_vdevice *vdev; bool destroying; +#ifdef CONFIG_IOMMU_LIVEUPDATE + bool liveupdate_preserved; +#endif }; =20 static inline struct iommufd_device * diff --git a/include/linux/iommufd.h b/include/linux/iommufd.h index 6e7efe83bc5d..80382aceca18 100644 --- a/include/linux/iommufd.h +++ b/include/linux/iommufd.h @@ -9,6 +9,7 @@ #include #include #include +#include #include #include #include @@ -213,6 +214,15 @@ int iommufd_access_rw(struct iommufd_access *access, u= nsigned long iova, int iommufd_vfio_compat_ioas_get_id(struct iommufd_ctx *ictx, u32 *out_ioa= s_id); int iommufd_vfio_compat_ioas_create(struct iommufd_ctx *ictx); int iommufd_vfio_compat_set_no_iommu(struct iommufd_ctx *ictx); + +#ifdef CONFIG_IOMMU_LIVEUPDATE +int iommufd_device_preserve(struct liveupdate_session *s, + struct iommufd_device *idev, + u64 *iommufd_tokenp); +void iommufd_device_unpreserve(struct liveupdate_session *s, + struct iommufd_device *idev); +bool iommufd_device_is_preserved(struct iommufd_device *idev); +#endif #else /* !CONFIG_IOMMUFD */ static inline struct iommufd_ctx *iommufd_ctx_from_file(struct file *file) { @@ -397,4 +407,23 @@ static inline void iommufd_viommu_destroy_mmap(struct = iommufd_viommu *viommu, { _iommufd_destroy_mmap(viommu->ictx, &viommu->obj, offset); } + +#if !IS_ENABLED(CONFIG_IOMMU_LIVEUPDATE) || !IS_ENABLED(CONFIG_IOMMUFD) +static inline int iommufd_device_preserve(struct liveupdate_session *s, + struct iommufd_device *idev, + u64 *iommufd_tokenp) +{ + return 0; +} + +static inline void iommufd_device_unpreserve(struct liveupdate_session *s, + struct iommufd_device *idev) +{ +} + +static inline bool iommufd_device_is_preserved(struct iommufd_device *idev) +{ + return false; +} +#endif #endif --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f198.google.com (mail-pl1-f198.google.com [209.85.214.198]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EFEED30FC33 for ; Mon, 21 Sep 2026 00:48:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.198 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951739; cv=none; b=txY80azUA1e9tyos3dBYivzwKTLhrmxKmapNJUuAlVuFedxJmsPvuXC8V8XOz/fzncLdoglJWOJbFzJKEjXj93VwOgnOb5I5txyv2mAle/OhghKBKwsjjw+5JQrK5xtuVmDtsV3ni0Mkeaxhp2NzPd01bX1G5CL2NJazASUogr8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951739; c=relaxed/simple; bh=dNP3ug4zy6sMOnrcTiQ4lvN1PGyKC36bSNyoIpacE1w=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=FuSkNadVT3/9wXQ2cFRxdqifdFKuu0c6m3OVN++d7jGLSU1hFUfW6pgDwWzlgI10KP72bxQfJ969YKxGmvRnBAytDcuqxwhadLRPwwYWROh6YgEpgZ0l3Ti8xjK+FMb/LfnO4bIVb0PGr2xGvnKNuxiCPGCa9UNRInPDbqA3GB4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=FYgMLSf/; arc=none smtp.client-ip=209.85.214.198 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="FYgMLSf/" Received: by mail-pl1-f198.google.com with SMTP id d9443c01a7336-2db22383e8fso35757875ad.1 for ; Sun, 20 Sep 2026 17:48:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951735; x=1790556535; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=SnkGcHTtOCYviBMPyvl5P9OipihIRIFDETRmg21FMBU=; b=FYgMLSf/Stn94vSjB8j1YXTjXmiHtFUiFQl8BMww4hR4XevpEr6XGdekbdusE9GCMU WZu/MfEYnSscapdjv2O736hzY0VZQeVoWLh7g0T7M/IoSc1IyG4FFso18TMwmfj5gxcZ jLQejbI2FjYMwYO4fEHswRhb1o+YzwoBOtC5LrmO2eDkjVAmmNlN4ePZdx8UOoNnb2nP bbGtZa4VfXnB42tf/UJudjUJkKE+d2RlEPv51tx+Ds7wE0nN5R+FjihZ+8zGY6m+zb25 cUT94AUrr3Ke/w1L038WhVcjrbycRb/Z2yEZ/qKtjypVxX3bPCBONsp4a4RCth8T1ngn x/1A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951735; x=1790556535; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=SnkGcHTtOCYviBMPyvl5P9OipihIRIFDETRmg21FMBU=; b=0RTqkjVf3UTsXGuzf1dvAbfs9yZ2GWf4HBdB68Styc7bpjl/avxSGvI08QEuVxaQO/ yPVH0bKFKq5aBukX1DLJoLB28TeOtf5POnd12UVDGAEjM5Fx+9gfdRWRxWdpt2xfE7G+ wwpeiBS/HQph8Al43fvwUPRyhRBNKYAiq/R35jnAkQI+UFuN6Rhixm/86kqtfEeFUCLE 0+NC6F9QeiKekpURtOW6sOCubIiDAXl3tyO1hlUHmVVIVNu3pIgIHeulA0z43P0gVcRd 40EKig5JHv/MnKZxccqHUQmu35/v3dNcmdM9gQtXbArVDIaYQmWe1zjQYS5+x5pxWw9h gh/Q== X-Forwarded-Encrypted: i=1; AKwUvByZRxqxyxWEov2LZ61bNOspkzTvhYadTbp+M4V6wJ7CVR4Z0Yg62KBg4vJ8J0Pi/9mBszRxngPLREL0a6U=@vger.kernel.org X-Gm-Message-State: AFuF++kGhYN+MAZ4RkrRFM88Di8eJoy9Cn7ipp2ieQy3qC4k2QMidCAd 5QDQasMXPONB6eIPNQ//UDkwNKRGMUHz9V1TEvidfKfxPgNI3Hct8GjvlkXWJekKJXcUP6g0F2M CPQVJcOjgbFuMMQ== X-Received: from plpp18.prod.google.com ([2002:a17:902:c712:b0:2df:474d:a5c7]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:d487:b0:2dd:c053:b9c9 with SMTP id d9443c01a7336-2ddc053bab9mr64172125ad.26.1789951734758; Sun, 20 Sep 2026 17:48:54 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:33 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-18-skhawaja@google.com> Subject: [PATCH v5 17/18] vfio/pci: Preserve the iommufd state of the vfio cdev From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , Pranjal Shrivastava , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" If the vfio cdev is attached to an iommufd, preserve the state of the attached iommufd also. Basically preserve the iommu specific state of the device and also the attach iommu HW unit. Once the device and its iommufd attachment is preserved, it cannot be detached or attached to another IOAS until it is unpreserved. Reviewed-by: Pranjal Shrivastava Signed-off-by: Samiullah Khawaja --- drivers/vfio/device_cdev.c | 10 +++++++++ drivers/vfio/pci/vfio_pci_liveupdate.c | 28 ++++++++++++++++++++++++-- 2 files changed, 36 insertions(+), 2 deletions(-) diff --git a/drivers/vfio/device_cdev.c b/drivers/vfio/device_cdev.c index b818aab92f76..4eb7ffd489bd 100644 --- a/drivers/vfio/device_cdev.c +++ b/drivers/vfio/device_cdev.c @@ -284,6 +284,11 @@ int vfio_df_ioctl_attach_pt(struct vfio_device_file *d= f, } =20 mutex_lock(&device->dev_set->lock); + if (iommufd_device_is_preserved(device->iommufd_device)) { + ret =3D -EBUSY; + goto out_unlock; + } + if (attach.flags & VFIO_DEVICE_ATTACH_PASID) ret =3D device->ops->pasid_attach_ioas(device, attach.pasid, @@ -342,6 +347,11 @@ int vfio_df_ioctl_detach_pt(struct vfio_device_file *d= f, } =20 mutex_lock(&device->dev_set->lock); + if (iommufd_device_is_preserved(device->iommufd_device)) { + mutex_unlock(&device->dev_set->lock); + return -EBUSY; + } + if (detach.flags & VFIO_DEVICE_DETACH_PASID) device->ops->pasid_detach_ioas(device, detach.pasid); else diff --git a/drivers/vfio/pci/vfio_pci_liveupdate.c b/drivers/vfio/pci/vfio= _pci_liveupdate.c index f0ea37d98696..fb2141ee8942 100644 --- a/drivers/vfio/pci/vfio_pci_liveupdate.c +++ b/drivers/vfio/pci/vfio_pci_liveupdate.c @@ -106,6 +106,7 @@ =20 #include #include +#include #include #include #include @@ -114,6 +115,8 @@ =20 #include "vfio_pci_priv.h" =20 +MODULE_IMPORT_NS("IOMMUFD"); + static bool vfio_pci_liveupdate_can_preserve(struct liveupdate_file_handle= r *handler, struct file *file) { @@ -154,15 +157,25 @@ static int vfio_pci_liveupdate_preserve(struct liveup= date_file_op_args *args) struct vfio_pci_core_device_ser *ser; struct vfio_pci_core_device *vdev; struct pci_dev *pdev; - int ret; + u64 iommufd_token; + int ret =3D 0; =20 vdev =3D container_of(device, struct vfio_pci_core_device, vdev); pdev =3D vdev->pdev; =20 - ret =3D pci_liveupdate_preserve(pdev); + mutex_lock(&device->dev_set->lock); + ret =3D iommufd_device_preserve(args->session, + device->iommufd_device, + &iommufd_token); + mutex_unlock(&device->dev_set->lock); + if (ret) return ret; =20 + ret =3D pci_liveupdate_preserve(pdev); + if (ret) + goto err_iommufd_unpreserve; + ser =3D kho_alloc_preserve(sizeof(*ser)); if (IS_ERR(ser)) { ret =3D PTR_ERR(ser); @@ -177,6 +190,12 @@ static int vfio_pci_liveupdate_preserve(struct liveupd= ate_file_op_args *args) =20 err_unpreserve: pci_liveupdate_unpreserve(pdev); + +err_iommufd_unpreserve: + mutex_lock(&device->dev_set->lock); + iommufd_device_unpreserve(args->session, + device->iommufd_device); + mutex_unlock(&device->dev_set->lock); return ret; } =20 @@ -184,6 +203,11 @@ static void vfio_pci_liveupdate_unpreserve(struct live= update_file_op_args *args) { struct vfio_device *device =3D vfio_device_from_file(args->file); =20 + mutex_lock(&device->dev_set->lock); + iommufd_device_unpreserve(args->session, + device->iommufd_device); + mutex_unlock(&device->dev_set->lock); + pci_liveupdate_unpreserve(to_pci_dev(device->dev)); kho_unpreserve_free(phys_to_virt(args->serialized_data)); } --=20 2.55.0.1082.g2b9226bbc0-goog From nobody Thu Sep 24 20:02:47 2026 Received: from mail-pl1-f197.google.com (mail-pl1-f197.google.com [209.85.214.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4E47A2C3266 for ; Mon, 21 Sep 2026 00:48:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.197 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951740; cv=none; b=c9POCwpQwMxtXVwrwvai1I17YOLr4mcnbI1MuBsPXYXleQzZUS95YGOTmRGH6tiuXnLyTa1GNh6D7+BVJ+yya2a/lUlQPj7rhK2QbLoEmavgvHqF+H3kaaYiYbzpANCE+wcJOh9giT6szznpiy/bMXmH1+Ru3x4N+Kvxd8oSciw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789951740; c=relaxed/simple; bh=8am889ShSkVPusNK1kpFvgrxYeOGtjoGYqKXzuv+tR0=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=pB1r6PJic2F8gjaBQ9XowxNyEv08e0xNwpsdBkJh0cAlBW2/E/Fw6a3RVg5s6SmRWokkUxVwtsph9bphT4DdYRWYXXOXPmoJVS//SMTvg38fVgXAqxHKbCc5hQw/h+S1npT75YXj8CRuOkPKv/OuUWBKZUktBVvGKwwSaZ9FzoM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=NVm/v6WN; arc=none smtp.client-ip=209.85.214.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="NVm/v6WN" Received: by mail-pl1-f197.google.com with SMTP id d9443c01a7336-2d6fb956002so34133975ad.1 for ; Sun, 20 Sep 2026 17:48:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1789951736; x=1790556536; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=kArZK0FqrV6rmyoKKIup6IrA4GYYPnuNKWQqNG3i9W4=; b=NVm/v6WNOjrjaxaR90jnXqLPXurNM0XUv10oiKub1SUXNX4RcR2dC3XfMuhzQ2zBd1 /Cf1SMo8YOogVNpySIosyEw51OVodpd3609c/5craMY99th2fW1dPnoLj1iQq02DJ+Wm 7uuYShSa32T8D+j/sSehVBMQRVUFzRzybov9f7Y16SIGZtaL9Og/YQFzyQppodN3wZqL C+BMmNei83Tfk5ugpSuwo72rMh/+vS0wPf30GqMqugFrfXF3uji+cHC7wAgJCPnutXjC JR3Ph/EYeQU3J3NDYFziy4wQd2x5t8xcNsSu8bm0ZKV6if2NhM2k9JiZ58rv2qQswnhV W3Ow== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789951736; x=1790556536; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=kArZK0FqrV6rmyoKKIup6IrA4GYYPnuNKWQqNG3i9W4=; b=seFbQmy1BI63ogDtIGMFfxdTdqkhtlWm+9t+AKPA95POucViFdacYr/gkPPP3a0CUY 5HgV5ltJgEwUK6z7f9Yjo4/LEEFQW+ukLVp8zsZo/DhbF2pzHDeew+O3AKfQvb5fXeT2 /8zVhM8QrriX7UZeh3XmaVl2M5lyy8l3955NdWX0InN/FwkNVJu4xhmACf3K7OrYrR+H wWA+D3Tbrwug5r1+86K1otpaxQoEdT6w2cEqK8oxtph8QdHfZw2zCdLxDjBeDPjqneks cI8AJDxdpeCwp7PIi3cGxJVZ8e7RXJpO7YDta3GU8Ow1N3WON3Ltn8dwtaHKykPRTH1o w4WQ== X-Forwarded-Encrypted: i=1; AKwUvBwUzROpeF0vSQy9tqmMscHAen0WnYsGdlxRq+91fip95IwX23wh5DX8fMu5e1Ef1UCMRmAujO5uw95I1/w=@vger.kernel.org X-Gm-Message-State: AFuF++k6ApWHXYd1JiF7+9lrwJqN51zgNyClwWZiQ+zzdMtZbC7S9/i/ xwHbGtgoV3rdxHh51NOdEGRky058lJlUkXSgVoN/Dd6V1UDVVZRz+KrpbxrvMHi9yxfrq/agxKu y74RmXWaAjqyB/Q== X-Received: from plbd20.prod.google.com ([2002:a17:902:f154:b0:2df:3910:96c0]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:986:b0:2dd:c053:d73b with SMTP id d9443c01a7336-2ddc053d79dmr82308235ad.34.1789951735528; Sun, 20 Sep 2026 17:48:55 -0700 (PDT) Date: Mon, 21 Sep 2026 00:48:34 +0000 In-Reply-To: <20260921004834.2601285-1-skhawaja@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260921004834.2601285-1-skhawaja@google.com> X-Mailer: git-send-email 2.55.0.1082.g2b9226bbc0-goog Message-ID: <20260921004834.2601285-19-skhawaja@google.com> Subject: [PATCH v5 18/18] iommufd/selftest: Add test to verify iommufd preservation From: Samiullah Khawaja To: David Woodhouse , Lu Baolu , Joerg Roedel , Will Deacon , Jason Gunthorpe Cc: Samiullah Khawaja , YiFei Zhu , Robin Murphy , Kevin Tian , Alex Williamson , Shuah Khan , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, kvm@vger.kernel.org, Pratyush Yadav , Pasha Tatashin , David Matlack , Andrew Morton , Pranjal Shrivastava , Vipin Sharma Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Test iommufd preservation by setting up an iommufd and vfio cdev and preserve it across live update. Test takes VFIO cdev path of a device bound to vfio-pci driver and binds it to an iommufd being preserved. It also preserves the vfio cdev so the iommufd state associated with it is also preserved. The restore path is tested by restoring the preserved vfio cdev only. On restore, test verifies that the bind with a new iommufd fails as the device is attached to the restored IOMMU domain. Also the LUO session finish fails as the preserved iommufd is not restored. Signed-off-by: Samiullah Khawaja Signed-off-by: YiFei Zhu --- tools/testing/selftests/iommu/Makefile | 12 + .../iommu/iommufd_liveupdate_kexec_test.c | 241 ++++++++++++++++++ 2 files changed, 253 insertions(+) create mode 100644 tools/testing/selftests/iommu/iommufd_liveupdate_kexec_= test.c diff --git a/tools/testing/selftests/iommu/Makefile b/tools/testing/selftes= ts/iommu/Makefile index 84abeb2f0949..ab35e8b21580 100644 --- a/tools/testing/selftests/iommu/Makefile +++ b/tools/testing/selftests/iommu/Makefile @@ -7,4 +7,16 @@ TEST_GEN_PROGS :=3D TEST_GEN_PROGS +=3D iommufd TEST_GEN_PROGS +=3D iommufd_fail_nth =20 +TEST_GEN_PROGS_EXTENDED +=3D iommufd_liveupdate_kexec_test + include ../lib.mk +include ../liveupdate/lib/libliveupdate.mk + +CFLAGS +=3D -I$(top_srcdir)/tools/include +CFLAGS +=3D -MD +CFLAGS +=3D $(EXTRA_CFLAGS) + +$(TEST_GEN_PROGS_EXTENDED): %: %.o $(LIBLIVEUPDATE_O) + $(CC) $(CFLAGS) $(CPPFLAGS) $(LDFLAGS) $(TARGET_ARCH) $< $(LIBLIVEUPDATE_= O) $(LDLIBS) -o $@ + +EXTRA_CLEAN +=3D $(LIBLIVEUPDATE_O) diff --git a/tools/testing/selftests/iommu/iommufd_liveupdate_kexec_test.c = b/tools/testing/selftests/iommu/iommufd_liveupdate_kexec_test.c new file mode 100644 index 000000000000..26c7e15b6187 --- /dev/null +++ b/tools/testing/selftests/iommu/iommufd_liveupdate_kexec_test.c @@ -0,0 +1,241 @@ +// SPDX-License-Identifier: GPL-2.0-only + +/* + * Copyright (c) 2026, Google LLC. + * Samiullah Khawaja + */ + +#include +#include +#include +#include +#include +#include + +#define __EXPORTED_HEADERS__ +#include +#include +#include +#include +#include + +#include "../kselftest.h" + +#define ksft_assert(condition) \ + do { \ + if (!(condition)) \ + fail_exit("Failed: %s", #condition); \ + } while (0) + +static const char *device_cdev_path; +static char state_session[LIVEUPDATE_SESSION_NAME_LENGTH]; +static char iommufd_session[LIVEUPDATE_SESSION_NAME_LENGTH]; + +static const uint64_t STATE_TOKEN; +static const uint64_t IOMMUFD_TOKEN =3D 0x123456; +static const uint64_t CDEV_TOKEN =3D 0x654321; +static const uint64_t HWPT_TOKEN =3D 0x789012; +static const uint64_t MEMFD_TOKEN =3D 0x890123; + +static int open_cdev(const char *vfio_cdev_path) +{ + int cdev_fd; + + cdev_fd =3D open(vfio_cdev_path, O_RDWR); + if (cdev_fd < 0) + ksft_exit_skip("Failed to open VFIO cdev: %s\n", vfio_cdev_path); + + return cdev_fd; +} + +static int open_iommufd(void) +{ + int iommufd; + + iommufd =3D open("/dev/iommu", O_RDWR); + if (iommufd < 0) + ksft_exit_skip("Failed to open /dev/iommu. IOMMUFD support not enabled.\= n"); + + return iommufd; +} + +static int create_sealed_memfd(size_t size) +{ + int fd, ret; + + fd =3D memfd_create("buffer", MFD_ALLOW_SEALING); + if (fd < 0) + fail_exit("memfd_create failed"); + + ret =3D ftruncate(fd, size); + if (ret) + fail_exit("ftruncate failed"); + + ret =3D fcntl(fd, F_ADD_SEALS, F_SEAL_GROW | F_SEAL_SHRINK | F_SEAL_SEAL); + if (ret) + fail_exit("fcntl F_ADD_SEALS failed"); + + return fd; +} + +#define test_ioctl(fd, cmd, arg) \ + do { \ + if (ioctl(fd, cmd, arg)) \ + fail_exit("ioctl(%s) failed", #cmd); \ + } while (0) + +#define test_luo_session_preserve_fd(session, fd, token) \ + do { \ + if (luo_session_preserve_fd(session, fd, token)) \ + fail_exit("luo_session_preserve_fd(%s) failed", #token); \ + } while (0) + +#define test_luo_session_retrieve_fd(session, token) \ + ({ \ + int _fd =3D luo_session_retrieve_fd(session, token); \ + if (_fd < 0) \ + fail_exit("luo_session_retrieve_fd(%s) failed", #token); \ + _fd; \ + }) + +static void setup_iommufd(int iommufd, int memfd, int cdev_fd) +{ + struct vfio_device_bind_iommufd bind =3D { + .argsz =3D sizeof(bind), + .flags =3D 0, + .iommufd =3D iommufd, + }; + struct iommu_ioas_alloc alloc_data =3D { + .size =3D sizeof(alloc_data), + .flags =3D 0, + }; + struct iommu_hwpt_alloc hwpt_alloc =3D { + .size =3D sizeof(hwpt_alloc), + .flags =3D 0, + }; + struct vfio_device_attach_iommufd_pt attach_data =3D { + .argsz =3D sizeof(attach_data), + .flags =3D 0, + }; + struct iommu_hwpt_liveupdate_mark_preserve mark_preserve =3D { + .size =3D sizeof(mark_preserve), + .hwpt_token =3D HWPT_TOKEN, + }; + struct iommu_ioas_map_file map_file =3D { + .size =3D sizeof(map_file), + .length =3D SZ_1M, + .flags =3D IOMMU_IOAS_MAP_WRITEABLE | + IOMMU_IOAS_MAP_READABLE | + IOMMU_IOAS_MAP_FIXED_IOVA, + .iova =3D SZ_4G, + .fd =3D memfd, + .start =3D 0, + }; + + test_ioctl(cdev_fd, VFIO_DEVICE_BIND_IOMMUFD, &bind); + + test_ioctl(iommufd, IOMMU_IOAS_ALLOC, &alloc_data); + + hwpt_alloc.dev_id =3D bind.out_devid; + hwpt_alloc.pt_id =3D alloc_data.out_ioas_id; + test_ioctl(iommufd, IOMMU_HWPT_ALLOC, &hwpt_alloc); + + attach_data.pt_id =3D hwpt_alloc.out_hwpt_id; + test_ioctl(cdev_fd, VFIO_DEVICE_ATTACH_IOMMUFD_PT, &attach_data); + + map_file.ioas_id =3D alloc_data.out_ioas_id; + test_ioctl(iommufd, IOMMU_IOAS_MAP_FILE, &map_file); + + mark_preserve.hwpt_id =3D attach_data.pt_id; + test_ioctl(iommufd, IOMMU_HWPT_LIVEUPDATE_MARK_PRESERVE, &mark_preserve); +} + +static void before_kexec(int luo_fd) +{ + int iommufd, cdev_fd, memfd, session; + + create_state_file(luo_fd, state_session, STATE_TOKEN, /*next_stage=3D*/2); + + session =3D luo_create_session(luo_fd, iommufd_session); + if (session < 0) + fail_exit("luo_create_session failed"); + + iommufd =3D open_iommufd(); + memfd =3D create_sealed_memfd(SZ_1M); + cdev_fd =3D open_cdev(device_cdev_path); + + setup_iommufd(iommufd, memfd, cdev_fd); + + /* Cannot preserve cdev without iommufd */ + if (!luo_session_preserve_fd(session, cdev_fd, CDEV_TOKEN)) + fail_exit("Preserving cdev without iommufd should fail"); + + /* Cannot preserve iommufd without preserving memfd. */ + if (!luo_session_preserve_fd(session, iommufd, IOMMUFD_TOKEN)) + fail_exit("Preserving iommufd without memfd should fail"); + + test_luo_session_preserve_fd(session, memfd, MEMFD_TOKEN); + test_luo_session_preserve_fd(session, iommufd, IOMMUFD_TOKEN); + test_luo_session_preserve_fd(session, cdev_fd, CDEV_TOKEN); + + close(session); + session =3D luo_create_session(luo_fd, iommufd_session); + if (session < 0) + fail_exit("luo_create_session failed"); + + test_luo_session_preserve_fd(session, memfd, MEMFD_TOKEN); + test_luo_session_preserve_fd(session, iommufd, IOMMUFD_TOKEN); + test_luo_session_preserve_fd(session, cdev_fd, CDEV_TOKEN); + + close(luo_fd); + daemonize_and_wait(); +} + +static void after_kexec(int luo_fd, int state_session_fd) +{ + int iommufd, cdev_fd, session, stage; + struct vfio_device_bind_iommufd bind =3D { + .argsz =3D sizeof(bind), + .flags =3D 0, + }; + + restore_and_read_stage(state_session_fd, STATE_TOKEN, &stage); + ksft_assert(stage =3D=3D 2); + + session =3D luo_retrieve_session(luo_fd, iommufd_session); + if (session < 0) + fail_exit("luo_retrieve_session failed"); + + cdev_fd =3D test_luo_session_retrieve_fd(session, CDEV_TOKEN); + + iommufd =3D luo_session_retrieve_fd(session, IOMMUFD_TOKEN); + if (iommufd >=3D 0) + fail_exit("iommufd should not be retrievable yet"); + + iommufd =3D open_iommufd(); + + bind.iommufd =3D iommufd; + if (ioctl(cdev_fd, VFIO_DEVICE_BIND_IOMMUFD, &bind) =3D=3D 0 || errno != =3D EPERM) + fail_exit("Binding cdev to new iommufd should fail with EPERM"); + + /* Should fail */ + if (luo_session_finish(session) =3D=3D 0) + fail_exit("luo_session_finish should fail if iommufd is not restored"); + + close(iommufd); + close(cdev_fd); +} + +int main(int argc, char *argv[]) +{ + if (argc < 2) { + printf("Usage: %s \n", argv[0]); + return 1; + } + + device_cdev_path =3D argv[1]; + sprintf(iommufd_session, "iommufd-test-%s", "cdev"); + sprintf(state_session, "state-%s", "iommufd-cdev"); + + return luo_test(argc, argv, state_session, before_kexec, after_kexec); +} --=20 2.55.0.1082.g2b9226bbc0-goog