From nobody Fri Sep 25 06:46:58 2026 Received: from PH0PR06CU001.outbound.protection.outlook.com (mail-westus3azon11011035.outbound.protection.outlook.com [40.107.208.35]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8C90F3DA7D4 for ; Tue, 15 Sep 2026 21:19:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.107.208.35 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789507143; cv=fail; b=t32VcGzjXxZ29Zkb3Tl+dC73L9OSPy8QY/Z4oYLof5IU7+2elf3nTbn5ehE44CK4PNjj1KRUUpbTPEP458/YUmZ4FDJZRLUOYstt99hvjVyZDk+6bFc2ZmLf+eqehxW72jKrDCw27zb0wbFV+iwTKn3DzgkCluelFtuGb0MZgyc= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789507143; c=relaxed/simple; bh=XZLGSF3DGsqC8NJh7Ls6MycipeP+LXoiFH6t7FabU9I=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=JfmLmWjP7BsdycDxNcrhBNmOWfe58yXxBgnDWmikUpNGvv7BMVtgdDH4Apuk/Omel4sLRr130BS6wMOahv1HhWftw2IpchvJenEjuWqKIoo5jQ+rsIjXS6BWOhrl9oelHAn4gP/EHSY/XFUs0oUULYR3OtZUD3SRi1aHHz1iuGI= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com; spf=fail smtp.mailfrom=nvidia.com; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b=UnbxXm6q; arc=fail smtp.client-ip=40.107.208.35 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=nvidia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b="UnbxXm6q" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=u/JMssG9hbxZymMKcuD9kE4fC26MmZGyzlMPW3xznyzLVFDdY77knij7wj/HDaUuLlE/PtISxJc8Ms5TT5BPKAvUpFpsAKO1pQ6QU/iTVqFxg2qckf6s8Mw+dLJVemkj18fIZot50o/toLj5WJPkbx7sR+eT0iUGbDGCheemJPCGB4VptToTopXv7jNIz/ljRsmidKA7V7FyaQ3rq3jZXWVc8oqbbqw59zPMDIUMSxsHpmB/phZLl8A5nUNFTVgxnWkkGwkxBjsjfi4eJVd0LauPJ2+DuVpdbQnw6e08qGyzogrMUoBR02TRIoJcJvIEC0hM3jcGk2AKkuZzD5D4ug== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=EUg44PugxkZv++9OHjFtF49Wj5/UCQo/OA95IJ0S5ok=; b=I4rc0S7oF+ABCzHmYmImKf9T0mLag+cVJTx+HmUQxZSxfywSQJnwoag6R0Esq8HPXbI0usHJMh1zFh/l5deU7LrazRJIougKsoBURwb7tNe58y9rIqGFc7GkxQ43+rKYEIuSK0Vfp8qn9N/lW4T143BB+4v+96LTGCPUFCsN/yGJ65rvy+invKtLFtYOrMXMwTtVnAWM222PEpkSqpzGlljSCqV4PMIYyO/pmqT7XmnzlisS0n9gzMHfRwyCWgr3PI2jHG7rXn10vGrF9krL987p55MvQKyIw/pNW9cymQitONlrjDpCgaHWYaY1o1JHGG7MZW6j/oLFrG9DpzbWMw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.118.233) smtp.rcpttodomain=kernel.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=EUg44PugxkZv++9OHjFtF49Wj5/UCQo/OA95IJ0S5ok=; b=UnbxXm6qLS4A+cyilnkCdnVPg5Upav2R1xBdCsWFaJ6eP8/+wVKBj/m8/V+aj/vPKBmqXctHOpzLPvxVlGpbOIS9jKogumcTbOdsoWrCiP3VnqSYv5wryhtZGB2WvYN1nJRgFhl0Ss/DX4+rOybLOu32iHtXNM9eYgchPiKwfqfKM8wBEiB8CFVGOpKyfX2gNNbAWu2jJ6e9wMhiyL1iF7NUsRdhSuAyr0whAXtEm50OFy4xmZ/H1xsYyt5R8wqOWUYuDMLdX5bd7HzE+gEbllFRdIG0kqeixmZeVmD5UB1Ufj8p7RGPHGMaJ/c4GDPsOLjGyMvQybsnq0x28b+FPg== Received: from DM6PR08CA0066.namprd08.prod.outlook.com (2603:10b6:5:1e0::40) by LY0PR12MB120611.namprd12.prod.outlook.com (2603:10b6:408:3b9::21) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.428.9; Tue, 15 Sep 2026 21:18:51 +0000 Received: from BN3PEPF0000B372.namprd21.prod.outlook.com (2603:10b6:5:1e0:cafe::63) by DM6PR08CA0066.outlook.office365.com (2603:10b6:5:1e0::40) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.9 via Frontend Transport; Tue, 15 Sep 2026 21:18:50 +0000 X-MS-Exchange-Authentication-Results: mx.microsoft.com 1; spf=pass (sender IP is 216.228.118.233) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.118.233 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.118.233; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.118.233) by BN3PEPF0000B372.mail.protection.outlook.com (10.167.243.169) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.451.0 via Frontend Transport; Tue, 15 Sep 2026 21:18:50 +0000 Received: from drhqmail202.nvidia.com (10.126.190.181) by mail.nvidia.com (10.127.129.6) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.49; Tue, 15 Sep 2026 14:18:27 -0700 Received: from drhqmail201.nvidia.com (10.126.190.180) by drhqmail202.nvidia.com (10.126.190.181) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.49; Tue, 15 Sep 2026 14:18:27 -0700 Received: from inno-dell.home (10.127.8.9) by mail.nvidia.com (10.126.190.180) with Microsoft SMTP Server id 15.2.2562.49 via Frontend Transport; Tue, 15 Sep 2026 14:18:20 -0700 From: Zhi Wang To: , CC: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Zhi Wang Subject: [PATCH 1/2] gpu: nova-core: vgpu: export lifecycle operations to VFIO Date: Wed, 16 Sep 2026 00:18:09 +0300 Message-ID: <20260915211811.84790-2-zhiw@nvidia.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260915211811.84790-1-zhiw@nvidia.com> References: <20260915211811.84790-1-zhiw@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-NV-OnPremToCloud: ExternallySecured X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN3PEPF0000B372:EE_|LY0PR12MB120611:EE_ X-MS-Office365-Filtering-Correlation-Id: cc71e7df-67b9-4825-463d-08df136ef02e X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|36860700016|1800799024|82310400026|23010399003|376014|7416014|6133799003|18002099003|22082099003|13003099007|3023799007|10067099003|56012099006|11063799006; X-Microsoft-Antispam-Message-Info: S3XNnGt02bqBc9SebDjXPyRLAvHk1vc5NDMv9nZFXb8xixNbazBovuKzWeTjiGuSo8JnaZbqZrkcliJNUR/QfDtijKrHIiY3ZgD3Y1SscFJXbWCLTKc9yhzFu+WCXAEzABOjQoJuOuEYaPwoA5zygbal65j+miL8JgNZmFP+ZTzU1L1xoipeYMbSQoZBoEf0clDY6tHLBuhpIH+55qRY90jIxlnBytVN509eMvurLZ0dtAz6WFuhuo4XKonikmiFLgjlA8Ze4FsUorcwEmLaBrAiO+7NO/pCUIabuWG7K9wcywhkzYf6jNs87WGgzIOhcE6vClgq59iZKDp24aicvyBMWAe6jASRjp2tzTRpUFa8gDHMYil4OS1Ubfe67YYKvsecT9wZkHu9u/YesbsgZ1zAP0veg7NRVgacIWjvtyyxe4WSjtjSpwr2Pnsg3mglppAgRfKcj94sAXzYF0jzF0D22yGUKuNIO74TZHH+OCQTDGhME1U+xqKocv1IVycY6uH8Q6r3akhY5exr02GgoNz1zUwJAu+1XeoYbsvbbeKrVmftGo2mvfJXV4fDF4Ob/mrJkTLHf5x6vA++Zi0rAJUehXymiFL/ePHn7rGGxPJhpa6HxY+g55TpcHPivUtNtUzYCtJIRklhhh5M1th4VeHkDnf5AFpmFtPx37kHBkOqXMwLgXQXYAvAoDfKggDQxQJuiFf5UweDLdQ/Uu2SHQ== X-Forefront-Antispam-Report: CIP:216.228.118.233;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:mail.nvidia.com;PTR:dc7edge2.nvidia.com;CAT:NONE;SFS:(13230040)(36860700016)(1800799024)(82310400026)(23010399003)(376014)(7416014)(6133799003)(18002099003)(22082099003)(13003099007)(3023799007)(10067099003)(56012099006)(11063799006);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: vCM30+TCilbngPZZ9Nk+lrVhd22aPMAfj9wGMHnKjzmvAA4+nGAaD8N5zG5o/gDEJXhwjq101fyxgb2nJuUMxxgr/L9BMQCu79hi+1TIBKYUomQ0MavEIHAPErWl3/J8JL7meADFl6t3zm9TLiakma2fBa8sh8wfiZxq/Z+0c47M8aJS/huCoFi8FIvuDLY+03yJhMBCX/IJzvuz5nMU75996I6S79f1NjSlHaskz643zlaQERXOD+LvGRqStxUXt8AJndQul+keLoOPNIodOgA7kXMQB/cqEyIlxpa5JkzhB0Vy4u3eAK6s4OO9eQFFvYHlaSRKcYPCfo++CeJkgiUcsFmHNj2YheDlhIBNC108fkDGqlntKpRUmuQhd1BgrcugFUfxsA98ZE0n4Zk1rliqS3JJP0iDrYYmRezirUlPF7rjnnLOwMWxyl1gkkrc X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 15 Sep 2026 21:18:50.4627 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: cc71e7df-67b9-4825-463d-08df136ef02e X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a;Ip=[216.228.118.233];Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: BN3PEPF0000B372.namprd21.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: LY0PR12MB120611 Content-Type: text/plain; charset="utf-8" VFIO open queries the VF's assigned vGPU type, allocates its resources, and boots the GSP plugin. Close destroys the instance, and reset resets the plugin and scrubs guest VRAM. Expose these operations through NovaCoreVfApi, embedded in the PF's typed SR-IOV registration. Its C operations table borrows the API until VF removal. Keep instance locking, activation and teardown in VgpuManager. Leave the C descriptor empty when vGPU mode is disabled so VF drivers can select their ordinary PCI passthrough path. Co-developed-by: Alok Kumar Signed-off-by: Alok Kumar Signed-off-by: Zhi Wang --- MAINTAINERS | 1 + drivers/gpu/nova-core/driver.rs | 59 ++++-- drivers/gpu/nova-core/gpu.rs | 58 +++++- drivers/gpu/nova-core/vgpu.rs | 7 + drivers/gpu/nova-core/vgpu/commands.rs | 11 +- drivers/gpu/nova-core/vgpu/fw/commands.rs | 1 + drivers/gpu/nova-core/vgpu/instance.rs | 153 +++++++++++++- drivers/gpu/nova-core/vgpu/vgpu_api.rs | 235 ++++++++++++++++++++++ include/drm/nvidia_vgpu.h | 51 +++++ rust/bindings/bindings_helper.h | 1 + 10 files changed, 548 insertions(+), 29 deletions(-) create mode 100644 drivers/gpu/nova-core/vgpu/vgpu_api.rs create mode 100644 include/drm/nvidia_vgpu.h diff --git a/MAINTAINERS b/MAINTAINERS index 922cfddcb2dc..92d5c0424452 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -8383,6 +8383,7 @@ C: irc://irc.oftc.net/nouveau T: git https://gitlab.freedesktop.org/drm/rust/kernel.git drm-rust-next F: Documentation/gpu/nova/ F: drivers/gpu/nova-core/ +F: include/drm/nvidia_vgpu.h =20 DRM DRIVER FOR NVIDIA GPUS [RUST] M: Danilo Krummrich diff --git a/drivers/gpu/nova-core/driver.rs b/drivers/gpu/nova-core/driver= .rs index 1373417386a2..67ac3b348f38 100644 --- a/drivers/gpu/nova-core/driver.rs +++ b/drivers/gpu/nova-core/driver.rs @@ -30,6 +30,15 @@ }, // }; =20 +#[cfg(CONFIG_PCI_IOV)] +use kernel::types::ForLt; + +#[cfg(CONFIG_PCI_IOV)] +use crate::vgpu::vgpu_api::{ + NovaCoreVfAbi, + NovaCoreVfApi, // +}; + /// Counter for generating unique auxiliary device IDs. static AUXILIARY_ID_COUNTER: Atomic =3D Atomic::new(0); =20 @@ -38,7 +47,7 @@ pub(crate) struct NovaCore<'bound> { #[cfg(CONFIG_PCI_IOV)] #[allow(clippy::type_complexity)] #[pin] - _vf_registration: pci::VfRegistration<'bound, CovariantForLt!(())>, + _vf_registration: pci::VfRegistration<'bound, ForLt!(NovaCoreVfApi<'_>= )>, #[pin] pub(crate) gpu: Gpu<'bound>, bar: pci::Bar<'bound, BAR0_SIZE>, @@ -111,18 +120,15 @@ fn probe<'bound>( pin_init::pin_init_scope(move || { dev_dbg!(pdev, "Probe Nova Core GPU driver.\n"); =20 - Ok(try_pin_init!(NovaCore { - #[cfg(CONFIG_PCI_IOV)] - // SAFETY: - // - probe has exclusive access before SR-IOV can be enabl= ed; - // - the registration is pinned in driver data and is its = first field; - // - no other registration is created for this device; and - // - the PCI adapter uses managed SR-IOV. - _vf_registration <- unsafe { pci::VfRegistration::new(pdev= , Ok(())) }, - _: { - pdev.enable_device_mem()?; - pdev.set_master(); - }, + #[cfg(CONFIG_PCI_IOV)] + if pdev.is_virtfn() { + return Err(ENODEV); + } + + pdev.enable_device_mem()?; + pdev.set_master(); + + Ok(try_pin_init!(&this in NovaCore { bar: pdev.iomap_region_sized::(0, c"nova-core/b= ar0")?, bar1: { let bar1_idx =3D bar1_resource_index(pdev)?; @@ -149,6 +155,33 @@ fn probe<'bound>( // Run optional GPU selftests. #[cfg(CONFIG_NOVA_CORE_SELFTESTS)] _: { gpu.run_selftests(pdev) }, + #[cfg(CONFIG_PCI_IOV)] + _vf_registration <- { + // SAFETY: `gpu` is initialized at its pinned address = and + // outlives the registration and its borrowed API data. + let gpu =3D unsafe { &(*this.as_ptr()).gpu }; + let api =3D NovaCoreVfApi::new(gpu, pdev); + let publish_ffi =3D api.is_available(); + let data =3D Ok(api); + // SAFETY: Probe has exclusive access to the registrat= ion + // slot and no VFs are enabled before successful probe= . The + // PCI adapter uses managed SR-IOV, and the pinned dri= ver + // data drops this registration before its borrowed GP= U resources. + // Each branch initializes the supplied pinned slot ex= actly + // once and delegates error cleanup to its initializer. + unsafe { + pin_init::pin_init_from_closure(move |slot| { + if publish_ffi { + pin_init::raw_try_init( + slot, + pci::VfRegistration::new_ffi::(pdev, data), + ) + } else { + pin_init::raw_try_init(slot, pci::VfRegist= ration::new(pdev, data)) + } + }) + } + }, _reg: auxiliary::Registration::new( pdev.as_ref(), c"nova-drm", diff --git a/drivers/gpu/nova-core/gpu.rs b/drivers/gpu/nova-core/gpu.rs index 289cbd171250..f28cd4611633 100644 --- a/drivers/gpu/nova-core/gpu.rs +++ b/drivers/gpu/nova-core/gpu.rs @@ -1,6 +1,9 @@ // SPDX-License-Identifier: GPL-2.0 =20 -use core::ops::Range; +use core::{ + num::NonZero, + ops::Range, // +}; =20 use kernel::{ device, @@ -8,6 +11,7 @@ fmt, gpu::buddy::GpuBuddyParams, io::Io, + new_mutex, num::Bounded, pci, prelude::*, @@ -16,6 +20,7 @@ SizeConstants, SZ_4K, // }, + sync::Mutex, }; =20 use crate::{ @@ -33,6 +38,7 @@ fsp::Fsp, gsp::{ self, + cmdq::Cmdq, Gsp, GspBootContext, // }, @@ -326,7 +332,8 @@ pub(crate) struct Gpu<'gpu> { /// /// Must be kept declared *before* `gsp_resources`, so that its compon= ents are dropped while /// the GSP is still operational. - mm: GpuMm<'gpu>, + #[pin] + mm: Mutex>, /// BAR1 user interface for CPU access to GPU virtual memory. #[pin] bar_user: BarUser<'gpu>, @@ -380,6 +387,33 @@ fn drop(self: Pin<&mut Self>) { } =20 impl<'gpu> Gpu<'gpu> { + pub(crate) fn cmdq(&self) -> &Cmdq<'gpu> { + &self.gsp_resources.gsp.cmdq + } + + pub(crate) fn vgpu_manager(&self) -> Option<&VgpuManager<'gpu>> { + self.vgpu.as_ref().map(|vgpu| vgpu.as_ref().get_ref()) + } + + pub(crate) fn vgpu_total_vfs(&self) -> Option> { + match self.gsp_resources.vgpu_state { + VgpuState::Disabled =3D> None, + VgpuState::Enabled { total_vfs } =3D> Some(total_vfs), + } + } + + pub(crate) fn mm(&self) -> &Mutex> { + &self.mm + } + + pub(crate) fn bar_user(&self) -> &BarUser<'gpu> { + &self.bar_user + } + + pub(crate) fn bar0(&self) -> Bar0<'gpu> { + self.gsp_resources.bar + } + pub(crate) fn new<'a>( pdev: &'gpu pci::Device>, bar: Bar0<'gpu>, @@ -509,7 +543,7 @@ pub(crate) fn new<'a>( }, =20 // Create GPU memory manager owning memory management resource= s. - mm: { + mm <- { let info =3D &gsp_resources.boot_result.static_info; let usable_vram =3D info.usable_fb_regions.first().ok_or(E= NODEV)?; let buddy_params =3D GpuBuddyParams { @@ -518,12 +552,15 @@ pub(crate) fn new<'a>( chunk_size: Alignment::new::(), }; =20 - GpuMm::new( - bar, - gsp_resources.spec.chipset, - buddy_params, - VramAddress::from_raw(info.total_fb_end), - )? + new_mutex!( + GpuMm::new( + bar, + gsp_resources.spec.chipset, + buddy_params, + VramAddress::from_raw(info.total_fb_end), + )?, + "nova-core::gpu-mm", + ) }, =20 // Create BAR1 user interface for CPU access to GPU virtual me= mory. @@ -548,10 +585,11 @@ pub(crate) fn run_selftests(self: Pin<&mut Self>, pde= v: &pci::Device u32 { self.total_channels } =20 + /// Returns the live-instance registry. + fn instances(&self) -> &Mutex> { + &self.instances + } + fn fifo_engine_list(&self) -> &FifoEngineList { &self.fifo_engine_list } diff --git a/drivers/gpu/nova-core/vgpu/commands.rs b/drivers/gpu/nova-core= /vgpu/commands.rs index 63b6f0e3d862..2a9f7ec18ae6 100644 --- a/drivers/gpu/nova-core/vgpu/commands.rs +++ b/drivers/gpu/nova-core/vgpu/commands.rs @@ -67,7 +67,6 @@ }; =20 /// Query the vGPU type assigned to a VF by its DBDF. -#[expect(dead_code)] pub(super) fn query_assigned_vf_type(cmdq: &Cmdq<'_>, dbdf: Dbdf) -> Resul= t { let request =3D u64::from(dbdf.into_raw()).to_le_bytes(); let response =3D @@ -210,6 +209,16 @@ pub(super) fn set_plugin_bme( rpc.rpc_call_nvkv(dev, bar0, gfid, RpcMessage::UpdateBmeState, &bme) } =20 +/// Reset an active GSP plugin. +pub(super) fn reset_plugin( + dev: &device::Device, + bar: Bar0<'_>, + gfid: Gfid, + rpc: &mut PluginRpc<'_, '_>, +) -> Result { + rpc.rpc_call(dev, bar, gfid, RpcMessage::Reset, &[]) +} + /// Whether a failed allocation may still have transferred CHID ownership = to firmware. pub(super) enum CeUtilsAllocError { /// A matching firmware response explicitly rejected the allocation. diff --git a/drivers/gpu/nova-core/vgpu/fw/commands.rs b/drivers/gpu/nova-c= ore/vgpu/fw/commands.rs index aabbf0987daa..c5e5c1eb714e 100644 --- a/drivers/gpu/nova-core/vgpu/fw/commands.rs +++ b/drivers/gpu/nova-core/vgpu/fw/commands.rs @@ -33,6 +33,7 @@ pub(in crate::vgpu) enum RpcMessage { VersionNegotiation =3D bindings::MESSAGE_NV_VGPU_CPU_RPC_MSG_VERSION_N= EGOTIATION, SetupConfigParamsAndInit =3D bindings::MESSAGE_NV_VGPU_CPU_RPC_MSG_SET= UP_CONFIG_PARAMS_AND_INIT, + Reset =3D bindings::MESSAGE_NV_VGPU_CPU_RPC_MSG_RESET, UpdateBmeState =3D bindings::MESSAGE_NV_VGPU_CPU_RPC_MSG_UPDATE_BME_ST= ATE, } =20 diff --git a/drivers/gpu/nova-core/vgpu/instance.rs b/drivers/gpu/nova-core= /vgpu/instance.rs index 95d300803ef6..e1fa128ce70a 100644 --- a/drivers/gpu/nova-core/vgpu/instance.rs +++ b/drivers/gpu/nova-core/vgpu/instance.rs @@ -9,7 +9,8 @@ prelude::*, ptr::Alignment, sizes::SizeConstants, - str::CString, // + str::CString, + sync::Mutex, // time::{ delay::fsleep, Delta, @@ -56,6 +57,7 @@ free_ceutils, negotiate_plugin_version, query_vgpu_properties, + reset_plugin, send_bootload, send_cleanup, send_plugin_config, @@ -116,7 +118,6 @@ fn wait_plugin_ready( pub(super) struct Gfid(pub(super) u32); =20 /// Resource requirements and device identity for one vGPU type. -#[expect(dead_code)] pub(super) struct VgpuType { vgpu_type_id: u32, bar1_length: u64, @@ -132,6 +133,22 @@ pub(super) const fn vgpu_type_id(&self) -> u32 { self.vgpu_type_id } =20 + pub(super) const fn bar1_length(&self) -> u64 { + self.bar1_length + } + + pub(super) const fn pci_dev_id(&self) -> u32 { + self.pci_dev_id + } + + pub(super) const fn pci_subsys_id(&self) -> u32 { + self.pci_subsys_id + } + + pub(super) const fn fb_length(&self) -> u64 { + self.fb_length + } + fn from_properties(properties: &VgpuProperties) -> Self { Self { vgpu_type_id: properties.type_id, @@ -158,6 +175,7 @@ pub(super) struct VgpuInstance<'gpu> { pub(super) plugin_rpc: PluginRpc<'gpu, 'gpu>, ceutils: Option, initialized: bool, + active: bool, needs_teardown: bool, } =20 @@ -233,6 +251,7 @@ fn shutdown(&mut self, dev: &device::Device, cmdq: &Cmdq<'_>) -> send_shutdown(dev, cmdq, self.gfid)?; dev_dbg!(dev, "shutdown: gfid=3D{} stopped\n", self.gfid.0); } + self.active =3D false; Ok(()) } } @@ -245,7 +264,6 @@ pub(super) struct InstanceInfo { vm_pid: u32, } =20 -#[expect(dead_code)] impl InstanceInfo { pub(super) const fn new(gfid: Gfid, dbdf: Dbdf, vgpu_type: VgpuType, v= m_pid: u32) -> Self { Self { @@ -296,7 +314,6 @@ pub(super) struct VgpuInstances<'gpu> { vram_slots: Option, } =20 -#[expect(dead_code)] impl<'gpu> VgpuInstances<'gpu> { pub(super) const fn new() -> Self { Self { @@ -420,6 +437,7 @@ pub(super) fn allocate_instance( plugin_rpc: PluginRpc::new(comm), ceutils: None, initialized: false, + active: false, needs_teardown: false, }; // Register ownership before firmware work so an uncertain result = leaves @@ -503,6 +521,35 @@ pub(super) fn activate_instance( Ok(()) } =20 + /// Reset an active instance and scrub its guest VRAM. + pub(super) fn reset_instance( + &mut self, + dev: &device::Device, + cmdq: &Cmdq<'_>, + bar: Bar0<'_>, + bar_user: &BarUser<'gpu>, + mm: &mut GpuMm<'_>, + gfid: Gfid, + ) -> Result { + let instance =3D self + .instances + .iter_mut() + .find(|instance| instance.gfid =3D=3D gfid) + .ok_or(ENOENT)?; + if !instance.active { + return Err(EBUSY); + } + + reset_plugin(dev, bar, instance.gfid, &mut instance.plugin_rpc)?; + instance.ceutils.as_ref().ok_or(EINVAL)?.scrub_guest_fb( + dev, + cmdq, + bar_user, + mm, + &instance.vram_slot.fbmem, + ) + } + /// Stop the plugin and release the instance's firmware and host resou= rces. pub(super) fn destroy_instance( &mut self, @@ -534,8 +581,104 @@ pub(super) fn destroy_instance( } =20 /// Query and decode one vGPU type using the typed NVKV schema. -#[expect(dead_code)] pub(super) fn query_vgpu_type(cmdq: &Cmdq<'_>, type_id: u32) -> Result { let properties =3D query_vgpu_properties(cmdq, type_id)?; Ok(VgpuType::from_properties(&properties)) } + +/// Activate an instance already owned by the live-instance registry. +/// +/// If activation fails, attempt full teardown before returning the origin= al +/// error. +#[expect(clippy::too_many_arguments)] +fn activate_registered_instance<'gpu>( + instances: &mut VgpuInstances<'gpu>, + dev: &device::Device, + cmdq: &Cmdq<'_>, + bar: Bar0<'_>, + bar_user: &BarUser<'gpu>, + mm: &mut GpuMm<'_>, + gfid: Gfid, + vgpu: &VgpuManager<'gpu>, +) -> Result { + let index =3D instances + .instances + .iter() + .position(|instance| instance.gfid =3D=3D gfid) + .ok_or(EIO)?; + let activation_result =3D instances.activate_instance(dev, cmdq, bar, = gfid, vgpu); + + if let Err(original_error) =3D activation_result { + if let Err(cleanup_error) =3D instances.destroy_instance(dev, cmdq= , bar_user, mm, gfid) { + dev_err!( + dev, + "vgpu_open: cleanup failed for gfid=3D{} after activation = error {:?}: {:?}\n", + gfid.0, + original_error, + cleanup_error, + ); + } + return Err(original_error); + } + + instances.instances[index].active =3D true; + Ok(()) +} + +impl<'gpu> VgpuManager<'gpu> { + /// Allocate, register, and activate a vGPU instance. + /// + /// Keep the registry locked from allocation through activation or rol= lback + /// so duplicate checks and vGPU type limits remain stable. + pub(super) fn create_instance( + &self, + dev: &device::Device, + cmdq: &Cmdq<'_>, + bar: Bar0<'_>, + bar_user: &'gpu BarUser<'gpu>, + mm: &Mutex>, + info: InstanceInfo, + ) -> Result { + let mut instances =3D self.instances().lock(); + // Global vGPU lock order: instances -> MM -> BAR-user VMM. + let mut mm =3D mm.lock(); + let gfid =3D instances.allocate_instance(dev, cmdq, bar_user, &mut= mm, self, info)?; + + activate_registered_instance( + &mut instances, + dev, + cmdq, + bar, + bar_user, + &mut mm, + gfid, + self, + ) + } + pub(super) fn close_instance( + &self, + dev: &device::Device, + cmdq: &Cmdq<'_>, + bar_user: &BarUser<'gpu>, + mm: &Mutex>, + gfid: Gfid, + ) -> Result { + let mut instances =3D self.instances().lock(); + let mut mm =3D mm.lock(); + instances.destroy_instance(dev, cmdq, bar_user, &mut mm, gfid) + } + + pub(super) fn reset_instance( + &self, + dev: &device::Device, + cmdq: &Cmdq<'_>, + bar: Bar0<'_>, + bar_user: &BarUser<'gpu>, + mm: &Mutex>, + gfid: Gfid, + ) -> Result { + let mut instances =3D self.instances().lock(); + let mut mm =3D mm.lock(); + instances.reset_instance(dev, cmdq, bar, bar_user, &mut mm, gfid) + } +} diff --git a/drivers/gpu/nova-core/vgpu/vgpu_api.rs b/drivers/gpu/nova-core= /vgpu/vgpu_api.rs new file mode 100644 index 000000000000..c60e31ba2623 --- /dev/null +++ b/drivers/gpu/nova-core/vgpu/vgpu_api.rs @@ -0,0 +1,235 @@ +// SPDX-License-Identifier: GPL-2.0 +// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIA= TES. All rights reserved. + +//! Nova Core VF API and its C operations table. + +use core::num::NonZero; + +use kernel::{ + bindings, + device, + interop::ffi::{ + Abi, + Token, // + }, + pci, + prelude::*, + sync::Mutex, + types::ForLt, // +}; + +use crate::{ + driver::Bar0, + gpu::Gpu, + gsp::cmdq::Cmdq, + mm::{ + bar_user::BarUser, + GpuMm, // + }, + vgpu::instance::{ + query_vgpu_type, + Gfid, + InstanceInfo, + VgpuType, // + }, +}; + +use super::{ + commands::{ + query_assigned_vf_type, + Dbdf, // + }, + VgpuManager, // +}; + +#[repr(transparent)] +struct VgpuTypeInfo(kernel::bindings::nvidia_vgpu_type_info); + +impl VgpuTypeInfo { + fn from_vgpu_type(vgpu_type: &VgpuType) -> Self { + Self(kernel::bindings::nvidia_vgpu_type_info { + pci_dev_id: vgpu_type.pci_dev_id(), + pci_subsys_id: vgpu_type.pci_subsys_id(), + bar1_length: vgpu_type.bar1_length(), + }) + } +} + +/// PF-owned operations available while a VF driver is bound. +pub(crate) struct NovaCoreVfApi<'gpu> { + pdev: &'gpu pci::Device, + cmdq: &'gpu Cmdq<'gpu>, + bar: Bar0<'gpu>, + bar_user: &'gpu BarUser<'gpu>, + mm: &'gpu Mutex>, + vgpu: Option<&'gpu VgpuManager<'gpu>>, + total_vfs: Option>, +} + +impl<'gpu> NovaCoreVfApi<'gpu> { + pub(crate) fn new(gpu: &'gpu Gpu<'gpu>, pdev: &'gpu pci::Device) -> Self { + Self { + pdev, + cmdq: gpu.cmdq(), + bar: gpu.bar0(), + bar_user: gpu.bar_user(), + mm: gpu.mm(), + vgpu: gpu.vgpu_manager(), + total_vfs: gpu.vgpu_total_vfs(), + } + } + + pub(crate) fn is_available(&self) -> bool { + self.vgpu.is_some() + } + + fn gfid(&self, gfid: u32) -> Result { + let total_vfs =3D self.total_vfs.ok_or(ENODEV)?; + if gfid =3D=3D 0 || gfid > u32::from(total_vfs.get()) { + return Err(EINVAL); + } + + Ok(Gfid(gfid)) + } +} + +impl NovaCoreVfApi<'_> { + fn open_instance(&self, gfid: u32, sbdf: u32, vm_pid: u32) -> Result { + let dev =3D self.pdev.as_ref(); + let gfid =3D self.gfid(gfid)?; + let dbdf =3D Dbdf::from_raw(sbdf); + + dev_dbg!( + dev, + "vgpu_open: gfid=3D{} sbdf=3D{:#x}\n", + gfid.0, + dbdf.into_raw() + ); + + let bar =3D self.bar; + let cmdq =3D self.cmdq; + let vgpu =3D self.vgpu.ok_or(ENODEV)?; + + let type_id =3D query_assigned_vf_type(cmdq, dbdf)?; + dev_dbg!( + dev, + "vgpu_open: gfid=3D{} assigned type_id=3D{}\n", + gfid.0, + type_id + ); + + let vgpu_type =3D query_vgpu_type(cmdq, type_id)?; + dev_dbg!( + dev, + "vgpu_open: gfid=3D{} vgpu_type=3D{} fb_length=3D{:#x}\n", + gfid.0, + vgpu_type.vgpu_type_id(), + vgpu_type.fb_length() + ); + + let type_info =3D VgpuTypeInfo::from_vgpu_type(&vgpu_type); + vgpu.create_instance( + dev, + cmdq, + bar, + self.bar_user, + self.mm, + InstanceInfo::new(gfid, dbdf, vgpu_type, vm_pid), + )?; + + Ok(type_info) + } + + fn close_instance(&self, gfid: u32) -> Result { + let dev =3D self.pdev.as_ref(); + let gfid =3D self.gfid(gfid)?; + + dev_dbg!(dev, "vgpu_close: gfid=3D{}\n", gfid.0); + + let cmdq =3D self.cmdq; + let result =3D + self.vgpu + .ok_or(ENODEV)? + .close_instance(dev, cmdq, self.bar_user, self.mm, gfid); + if let Err(error) =3D result { + dev_err!(dev, "vgpu_close: gfid=3D{} failed: {:?}\n", gfid.0, = error); + } + result + } + + fn reset_instance(&self, gfid: u32) -> Result { + let dev =3D self.pdev.as_ref(); + let gfid =3D self.gfid(gfid)?; + + dev_dbg!(dev, "vgpu_reset: gfid=3D{}\n", gfid.0); + + let cmdq =3D self.cmdq; + self.vgpu.ok_or(ENODEV)?.reset_instance( + dev, + cmdq, + self.bar, + self.bar_user, + self.mm, + gfid, + )?; + + dev_dbg!(dev, "vgpu_reset: gfid=3D{} done\n", gfid.0); + Ok(()) + } +} + +#[kernel::macros::ffi_vtable(NVIDIA_VGPU_OPS: bindings::nvidia_vgpu_ops)] +impl NovaCoreVfApi<'_> { + /// Create and activate an instance, writing its vGPU type information= on success. + /// + /// # Safety + /// + /// If `type_info` is non-null, it must point to aligned, writable sto= rage + /// with no concurrent access for the duration of this call. + unsafe fn open( + self: Pin<&Self>, + gfid: core::ffi::c_uint, + sbdf: core::ffi::c_uint, + vm_pid: core::ffi::c_uint, + type_info: *mut bindings::nvidia_vgpu_type_info, + ) -> Result { + if type_info.is_null() { + return Err(EINVAL); + } + + let info =3D self.open_instance(gfid, sbdf, vm_pid)?; + // SAFETY: `type_info` is non-null and the caller guarantees align= ed, + // writable storage with no concurrent access. + unsafe { type_info.write(info.0) }; + Ok(()) + } + + /// Tear down an instance, retaining resources if firmware cleanup fai= ls. + fn close(self: Pin<&Self>, gfid: core::ffi::c_uint) { + let _ =3D self.close_instance(gfid); + } + + /// Reset an instance and scrub its guest VRAM. + fn reset(self: Pin<&Self>, gfid: core::ffi::c_uint) -> Result { + self.reset_instance(gfid) + } +} + +/// C ABI published for VF lifecycle calls. +pub(crate) struct NovaCoreVfAbi; + +// SAFETY: The C header defines the operations layout and ABI identity. The +// macro initializes that complete operations table, and every trampoline +// borrows the pinned `NovaCoreVfApi` published by `new_ffi`. +unsafe impl Abi for NovaCoreVfAbi { + type Context =3D ForLt!(NovaCoreVfApi<'_>); + type RawOps =3D bindings::nvidia_vgpu_ops; + + const OPS: &'static Self::RawOps =3D &NVIDIA_VGPU_OPS; + const TOKEN: Token =3D Token::new( + bindings::NVIDIA_VGPU_FFI_TOKEN_HIGH as u64, + bindings::NVIDIA_VGPU_FFI_TOKEN_LOW as u64, + ); + const ABI_MAJOR: u16 =3D bindings::NVIDIA_VGPU_FFI_ABI_MAJOR as u16; + const ABI_MINOR: u16 =3D bindings::NVIDIA_VGPU_FFI_ABI_MINOR as u16; +} diff --git a/include/drm/nvidia_vgpu.h b/include/drm/nvidia_vgpu.h new file mode 100644 index 000000000000..4a1579a9044d --- /dev/null +++ b/include/drm/nvidia_vgpu.h @@ -0,0 +1,51 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIA= TES. All rights reserved. +#ifndef __DRM_NVIDIA_VGPU_H__ +#define __DRM_NVIDIA_VGPU_H__ + +#include + +#define NVIDIA_VGPU_FFI_TOKEN_HIGH 0xa70c87381a844092ULL +#define NVIDIA_VGPU_FFI_TOKEN_LOW 0xaeedbe595835927cULL +#define NVIDIA_VGPU_FFI_ABI_MAJOR 1U +#define NVIDIA_VGPU_FFI_ABI_MINOR 0U + +/** + * struct nvidia_vgpu_type_info - vGPU type descriptor returned by open + * @pci_dev_id: PCI device ID to present to the guest + * @pci_subsys_id: PCI subsystem ID to present to the guest + * @bar1_length: BAR1 aperture size in MiB + */ +struct nvidia_vgpu_type_info { + u32 pci_dev_id; + u32 pci_subsys_id; + u64 bar1_length; +}; + +/** + * struct nvidia_vgpu_ops - PF operations for NVIDIA vGPU virtual functions + * @open: Create and activate an instance, returning its vGPU type informa= tion on success + * @close: Tear down an instance, retaining resources if firmware cleanup = fails + * @reset: Reset an instance and scrub its guest VRAM + * + * Each operation receives the context from the borrowed struct rust_ffi a= nd + * a Guest Function ID (VF index + 1). The context remains valid until the= VF + * driver is fully unbound, including its remove callback. The VF driver m= ust + * drain all calls before returning from remove or a failed probe. + * + * All operations may sleep. They must not acquire the PF device lock beca= use + * disabling SR-IOV removes VFs synchronously while holding that lock. + * + * @open receives the VF address encoded as (segment << 16) | (bus << 8) |= devfn + * and the VM process's thread-group ID. Its type_info argument must point= to + * writable storage with no concurrent access. @open and @reset return zero + * on success or a negative errno. + */ +struct nvidia_vgpu_ops { + int (*open)(const void *context, unsigned int gfid, unsigned int sbdf, + unsigned int vm_pid, struct nvidia_vgpu_type_info *type_info); + void (*close)(const void *context, unsigned int gfid); + int (*reset)(const void *context, unsigned int gfid); +}; + +#endif /* __DRM_NVIDIA_VGPU_H__ */ diff --git a/rust/bindings/bindings_helper.h b/rust/bindings/bindings_helpe= r.h index 2467305c4842..1cf79bec756a 100644 --- a/rust/bindings/bindings_helper.h +++ b/rust/bindings/bindings_helper.h @@ -37,6 +37,7 @@ #include #include #include +#include #include #include #include From nobody Fri Sep 25 06:46:58 2026 Received: from BYAPR05CU005.outbound.protection.outlook.com (mail-westusazon11010003.outbound.protection.outlook.com [52.101.85.3]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EAA1B48592B for ; Tue, 15 Sep 2026 21:19:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.85.3 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789507161; cv=fail; b=b52ORevm2nLt5R9B5jbLUqXbHMf3+hb/xfeff+p5IQIUxunrWpYK6qE3US8hri1enCQiwwMqnLwUOxML1Fjx6WnKc33bdkppp/BUmpoL6e1vEon8NVAlRzxe1dC5dXse3tLapDWDLH2BW+OKnA41P7fSovGNtOHE7c5EiHX477o= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789507161; c=relaxed/simple; bh=kFGm2v+bGm9i2GvMC0LyyGuPkBWSumeqNVx8T2z4ZNY=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=VYOAHOObfR57VOrKfB12KQMX/lYFMgEU7rLMHbdjsISfDpvkau0Ri1LXhgLriENGO9xOOw5NIxPPbp7Yo9LVZby6dCjc8qKarndzY6H+qVsu0asdTy8FjHU2Vnzv3UBDdRM1Q/HQY/WEqWLwUBKvp9r26IzvJk8VCI86adjLTBA= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com; spf=fail smtp.mailfrom=nvidia.com; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b=J7SU/IQi; arc=fail smtp.client-ip=52.101.85.3 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=nvidia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b="J7SU/IQi" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=gV9Otv4bjtYb+4TGWXEPV7+pTOExzft6JqiMqdG3tgsgoCkY0gGOmUCkqL9U2L/4JF29aigXFRdEMZ03KWNw+nWxuefa1feaE5kr/pQRcSJzSIL0XTFSHuROlREluhpViNQK6atO4dr/4cYhNIpnQBqiKQgb7kSvGxtkq05/dU0Li1POFcUfsv+MgY+ocjhhzYhAlInRuZKqUd6QckUJon65xeBKO4sMTZ46rS1p0jlUOY9I+FgeU77k2n/J/dtrW9YkcqWB02mxRULTA9cCvxcjJS3zjn8uyVFim3a/Jhp/Xfn/M7BGS9gq7lo5JyM9ZQ4T5uM3ys8jYOS9sHbkfA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=pItjb9yQrK1Xn9KMI7m/BfRYQvsm673YvQbU6bMLWSU=; b=UlGHtAUoTou6JqKMpy+MCW3ABo5tJ7hc4ZbOQScygnZ+o9w/Lzn8y1kCXhGDzKDc7MqhUOnPvnhFdkGmir+l8EHa1z2vL6na8hfkeVt/BQ5rJKaMEn4yku+aAH1BN7ZBJ/wRSyEnOsuDUGDqBy232NKANEZlHegVtzslQ4hl2WcS9Vf1DcxlL8+aMqtuI75NuF2yG8y7JKm6QTa3GqpInefVmympe6wmlRcrNo/y8PggZGK+wSXOj18tuHXsA1lI9iTIC23idEpDONCCjlw6RBIilyRWaPhT+S8LSb5VxNnrYpTQGH1bqSX8XVzvX/68pkAXwuGGLwFhROmArl+Jrw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.118.233) smtp.rcpttodomain=kernel.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=pItjb9yQrK1Xn9KMI7m/BfRYQvsm673YvQbU6bMLWSU=; b=J7SU/IQifAGHP5nzyg3EoiwkbIBTeCZFGPudJzAd5bpK9svpdqeJjdUJm3pffBsn0chH9gR6xv6O/Uw7V66fN4lR3nqkXTkrO+B+HA8ky5tzUMSo2FG49p+SUYoL7MLbTycakrlwy8FOsPZghYo2t36WOOI1M8iIkUYLYWNtYdviPEBYwVN29pzMNtdtd8bp6tFCIJi4Ow6DYwfm+feRUX03OIqftYi3v2Mws3920MbKP++/FZlWxPVgFIztLQMT8BidneYM1hjfCa7JIXHuNvwHFFOpN1XIHt34Gcz8WklMdazRuIhLrx0OJDQEsvsgPrsnRYI5m4mdFtB0CyrnGg== Received: from BN9PR03CA0057.namprd03.prod.outlook.com (2603:10b6:408:fb::32) by CH3PR12MB7691.namprd12.prod.outlook.com (2603:10b6:610:151::18) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.406.12; Tue, 15 Sep 2026 21:18:59 +0000 Received: from BN3PEPF0000B36D.namprd21.prod.outlook.com (2603:10b6:408:fb:cafe::85) by BN9PR03CA0057.outlook.office365.com (2603:10b6:408:fb::32) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.428.11 via Frontend Transport; Tue, 15 Sep 2026 21:18:59 +0000 X-MS-Exchange-Authentication-Results: mx.microsoft.com 1; spf=pass (sender IP is 216.228.118.233) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.118.233 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.118.233; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.118.233) by BN3PEPF0000B36D.mail.protection.outlook.com (10.167.243.164) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.451.0 via Frontend Transport; Tue, 15 Sep 2026 21:18:59 +0000 Received: from drhqmail202.nvidia.com (10.126.190.181) by mail.nvidia.com (10.127.129.6) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.49; Tue, 15 Sep 2026 14:18:35 -0700 Received: from drhqmail201.nvidia.com (10.126.190.180) by drhqmail202.nvidia.com (10.126.190.181) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.49; Tue, 15 Sep 2026 14:18:34 -0700 Received: from inno-dell.home (10.127.8.9) by mail.nvidia.com (10.126.190.180) with Microsoft SMTP Server id 15.2.2562.49 via Frontend Transport; Tue, 15 Sep 2026 14:18:27 -0700 From: Zhi Wang To: , CC: , , , , , , , , , , , , , , , , , , , , , , , , , , , , , Zhi Wang Subject: [PATCH 2/2] vfio/nvidia-vgpu: add the NVIDIA vGPU VFIO variant driver Date: Wed, 16 Sep 2026 00:18:10 +0300 Message-ID: <20260915211811.84790-3-zhiw@nvidia.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260915211811.84790-1-zhiw@nvidia.com> References: <20260915211811.84790-1-zhiw@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-NV-OnPremToCloud: ExternallySecured X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN3PEPF0000B36D:EE_|CH3PR12MB7691:EE_ X-MS-Office365-Filtering-Correlation-Id: b25db569-4693-4159-31b7-08df136ef545 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|1800799024|23010399003|376014|7416014|82310400026|36860700016|18002099003|3023799007|10067099003|22082099003|6133799003|11063799006|56012099006|5023799004; X-Microsoft-Antispam-Message-Info: /8PI75oVI7ohrgK0KDXzAsGNM9NwHsnaGZRh4SEK6uTl+sWjkqnS/LSQ9JBiMkhOGswcAX6D1+X8K/bZSPS0jfZ/JLKdfI4nVdrvGkO605haqyMCiYSr4D5O9GOmuRGH85gM/L5bSwKU+xszRV/lPzFGO7xniqsITD6tNmSOzSZvonoKGD0ehCMVHjfIBpbpEoyN74y6R7s+PB3fg4nIRyNsf8z5Afj0RaqRjIEpXo/qeotZOtXQHqSAy5gJwKZE4yuLP20XJvlZ8+PIDlTQNFYO2TlRLcAGd3dc07zyF9cdL/7HuJJXOrHBmstukbYOd6mPS7Z++Wo0MTz44VeilzJsAbOO5GHeyFyVn31OSEQzyXNVlIDsatMvRlvdJXpdHpRg1VCb/CQAsT/6IbFTYMb58hv+BYA3uzrT+VvzHU+Qky/A1OdTTQyiXANUNE4SiiP+46f81hOD4EmZqY2bPtLVQXnLBCif7dIOG2JALEhF67LBTySs6e2Wn3eBiO/VF0zOY1QjrNYu0loGEuQ69TwrGHN/6ie87BQWWDgJSa2aeKqLckXtpnz+RRJJiYI0XNVxDQSzoTbiy2teutbz5ghx5GcRIcFNaoAdyomwcS0LjhLKQWwG47KKBfD1ppE89WBkHauU4WX6jJ0OsUvAMudzYIvjXDgHVaO0rMUzTDjHZxr6mZgvDGY0G2aA8jOIzIHv7NxdICQKvHh6HoPS+g== X-Forefront-Antispam-Report: CIP:216.228.118.233;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:mail.nvidia.com;PTR:dc7edge2.nvidia.com;CAT:NONE;SFS:(13230040)(1800799024)(23010399003)(376014)(7416014)(82310400026)(36860700016)(18002099003)(3023799007)(10067099003)(22082099003)(6133799003)(11063799006)(56012099006)(5023799004);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: sFZcQvW91tRHuopNtTHThpKbRZkt4nE6HXWoWGvX0mxb4FgRoeTjRXUnt1tuAAoKYzkvRseZtuzXxVfPmuRKPsivpL3IBHuF8usCWFYmiccFu+zE+jrxBXwjetcMdvoMtmQSHRriXUvP+rWBkR+8iBLAH1+bAuDeVc7RKb3vNHDkVddUxxaFF9c0+HiClp43bSUYfDjEi7/QDlt3mGYanuh9qCc1wyyif3vYkM4SNt8DS5y298Wo+8upCsizESWUkDCCN6hMNhJIHNqa28gq7WtBcIOs83qvfH6iEW3ZX+e8qa2AhvOpsx+AN7CGa1isiwFCmMeUJlhtUzbWoomDoAOYt+8rSpxt5oAelEaEXpq0mY2tayud/PXcKe+eLQIRNQZLbWLWcvgg+j3wjxFzvvyrSaquYhclethdomQHjNjOV8gDKLJ0FuF7qD6O6G4Y X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 15 Sep 2026 21:18:59.0349 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: b25db569-4693-4159-31b7-08df136ef545 X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a;Ip=[216.228.118.233];Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: BN3PEPF0000B36D.namprd21.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: CH3PR12MB7691 Content-Type: text/plain; charset="utf-8" NVIDIA vGPU VFs require their open, reset, and close lifecycle to be coordinated with the PF-side nova-core driver. Add a VFIO PCI variant with an override alias for NVIDIA 10de:2bb5 devices with class/subclass 0302. Leave the programming-interface byte unmasked so the generated alias ends in i*, as required by libvirt VFIO variant matching. PFs share this ID with their VFs. Use ordinary PCI passthrough unless a VF can borrow a compatible, complete Nova lifecycle interface through the typed SR-IOV descriptor. Both paths use VFIO PCI core defaults and support PASID attachment; generic passthrough also retains the core DMA-BUF export operation. Provide a driver-local enable_sriov option for PFs bound to this driver. For vGPU instances, expose the firmware-selected device and subsystem IDs through configuration-space reads and constrain BAR1 region reporting, read, write, and mmap access to the assigned vGPU type's aperture. Coordinate firmware reset through PCI reset callbacks and serialize the entire reset window against instance open and close without retaining a mutex across a multi-device bus reset. Latch firmware reset failures until successful reopen, report them through the VFIO error IRQ and reset ioctls, and preserve core configuration-write return values. Signed-off-by: Zhi Wang --- drivers/vfio/pci/Kconfig | 2 + drivers/vfio/pci/Makefile | 2 + drivers/vfio/pci/nvidia-vgpu/Kconfig | 21 ++ drivers/vfio/pci/nvidia-vgpu/Makefile | 2 + drivers/vfio/pci/nvidia-vgpu/main.c | 509 ++++++++++++++++++++++++++ 5 files changed, 536 insertions(+) create mode 100644 drivers/vfio/pci/nvidia-vgpu/Kconfig create mode 100644 drivers/vfio/pci/nvidia-vgpu/Makefile create mode 100644 drivers/vfio/pci/nvidia-vgpu/main.c diff --git a/drivers/vfio/pci/Kconfig b/drivers/vfio/pci/Kconfig index 296bf01e185e..b48d8d1af42a 100644 --- a/drivers/vfio/pci/Kconfig +++ b/drivers/vfio/pci/Kconfig @@ -74,4 +74,6 @@ source "drivers/vfio/pci/qat/Kconfig" =20 source "drivers/vfio/pci/xe/Kconfig" =20 +source "drivers/vfio/pci/nvidia-vgpu/Kconfig" + endmenu diff --git a/drivers/vfio/pci/Makefile b/drivers/vfio/pci/Makefile index 6138f1bf241d..f3498e541555 100644 --- a/drivers/vfio/pci/Makefile +++ b/drivers/vfio/pci/Makefile @@ -24,3 +24,5 @@ obj-$(CONFIG_NVGRACE_GPU_VFIO_PCI) +=3D nvgrace-gpu/ obj-$(CONFIG_QAT_VFIO_PCI) +=3D qat/ =20 obj-$(CONFIG_XE_VFIO_PCI) +=3D xe/ + +obj-$(CONFIG_NVIDIA_VGPU_VFIO_PCI) +=3D nvidia-vgpu/ diff --git a/drivers/vfio/pci/nvidia-vgpu/Kconfig b/drivers/vfio/pci/nvidia= -vgpu/Kconfig new file mode 100644 index 000000000000..e7791fa79984 --- /dev/null +++ b/drivers/vfio/pci/nvidia-vgpu/Kconfig @@ -0,0 +1,21 @@ +# SPDX-License-Identifier: GPL-2.0-only +config NVIDIA_VGPU_VFIO_PCI + tristate "VFIO support for the NVIDIA vGPU" + depends on NOVA_CORE && PCI_IOV + select VFIO_PCI_CORE + help + This option enables VFIO (Virtual Function I/O) support for + NVIDIA virtual GPUs (vGPU). It allows the assignment of a virtual + GPU instance to userspace applications via VFIO, typically used + with hypervisors such as KVM and device emulators like QEMU. + + Devices without an available Nova vGPU interface use ordinary PCI + passthrough. Both paths use the VFIO PCI core defaults; vfio-pci + module options do not apply. The enable_sriov parameter of this + driver controls SR-IOV configuration for PFs bound to it. + + The NVIDIA vGPU allows a physical GPU to be partitioned into + multiple virtual GPUs, each of which can be passed to a virtual + machine as a PCI device using the standard VFIO infrastructure. + + If you don't know what to do here, say N. diff --git a/drivers/vfio/pci/nvidia-vgpu/Makefile b/drivers/vfio/pci/nvidi= a-vgpu/Makefile new file mode 100644 index 000000000000..193cc801a081 --- /dev/null +++ b/drivers/vfio/pci/nvidia-vgpu/Makefile @@ -0,0 +1,2 @@ +obj-$(CONFIG_NVIDIA_VGPU_VFIO_PCI) +=3D nvidia-vgpu-vfio-pci.o +nvidia-vgpu-vfio-pci-y :=3D main.o diff --git a/drivers/vfio/pci/nvidia-vgpu/main.c b/drivers/vfio/pci/nvidia-= vgpu/main.c new file mode 100644 index 000000000000..742979fb0ed5 --- /dev/null +++ b/drivers/vfio/pci/nvidia-vgpu/main.c @@ -0,0 +1,509 @@ +// SPDX-License-Identifier: GPL-2.0-only +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#define NVIDIA_VGPU_DRIVER_NAME "nvidia-vgpu-vfio-pci" + +static bool enable_sriov; +module_param(enable_sriov, bool, 0644); +MODULE_PARM_DESC(enable_sriov, "Enable SR-IOV configuration for PFs bound = to this driver"); + +static int nvidia_vgpu_fb_bar_index(struct pci_dev *pdev) +{ + if (pci_resource_flags(pdev, 0) & IORESOURCE_MEM_64) + return 2; + return 1; +} + +static const struct rust_ffi_token nvidia_vgpu_ffi_token =3D { + .high =3D NVIDIA_VGPU_FFI_TOKEN_HIGH, + .low =3D NVIDIA_VGPU_FFI_TOKEN_LOW, +}; + +struct nvidia_vgpu_pci_core_device { + struct vfio_pci_core_device core_device; + struct nvidia_vgpu_type_info type_info; + const struct nvidia_vgpu_ops *pf_ops; + const void *pf_context; + /* Protects lifecycle calls and the active/resetting state. */ + struct mutex instance_lock; + struct completion reset_completion; + unsigned int gfid; + int reset_error; + bool instance_active; + bool resetting; +}; + +/* Encode the VF's PCI segment:bus:device.function as a 32-bit address. */ +static inline unsigned int nvidia_vgpu_vf_sbdf(struct pci_dev *vf) +{ + return ((u32)pci_domain_nr(vf->bus) << 16) | pci_dev_id(vf); +} + +/* Keep open and close outside the complete prepare/PCI-reset/done window.= */ +static void nvidia_vgpu_lock_instance(struct nvidia_vgpu_pci_core_device *= nvdev) +{ + for (;;) { + mutex_lock(&nvdev->instance_lock); + if (!nvdev->resetting) + return; + mutex_unlock(&nvdev->instance_lock); + wait_for_completion(&nvdev->reset_completion); + } +} + +static int nvidia_vgpu_open_device(struct vfio_device *core_vdev) +{ + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_vdev, struct nvidia_vgpu_pci_core_device, + core_device.vdev); + struct pci_dev *vf =3D to_pci_dev(core_vdev->dev); + struct nvidia_vgpu_type_info type_info; + int ret; + + ret =3D vfio_pci_core_enable(&nvdev->core_device); + if (ret) + return ret; + + nvidia_vgpu_lock_instance(nvdev); + ret =3D nvdev->pf_ops->open(nvdev->pf_context, nvdev->gfid, + nvidia_vgpu_vf_sbdf(vf), task_tgid_nr(current), + &type_info); + if (!ret) { + nvdev->instance_active =3D true; + WRITE_ONCE(nvdev->reset_error, 0); + } + mutex_unlock(&nvdev->instance_lock); + if (ret) { + vfio_pci_core_disable(&nvdev->core_device); + return ret; + } + + nvdev->type_info =3D type_info; + put_unaligned_le16(type_info.pci_dev_id, + nvdev->core_device.vconfig + PCI_DEVICE_ID); + pci_dbg(vf, "vgpu open: dev_id=3D0x%x subsys_id=3D0x%x bar1_length=3D0x%l= lx\n", + type_info.pci_dev_id, type_info.pci_subsys_id, + type_info.bar1_length); + vfio_pci_core_finish_enable(&nvdev->core_device); + return 0; +} + +static void nvidia_vgpu_close_device(struct vfio_device *core_vdev) +{ + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_vdev, struct nvidia_vgpu_pci_core_device, + core_device.vdev); + + nvidia_vgpu_lock_instance(nvdev); + nvdev->instance_active =3D false; + nvdev->pf_ops->close(nvdev->pf_context, nvdev->gfid); + mutex_unlock(&nvdev->instance_lock); + vfio_pci_core_close_device(core_vdev); +} + +static int nvidia_vgpu_bar1_size(struct nvidia_vgpu_pci_core_device *nvdev, + u64 *size) +{ + struct pci_dev *pdev =3D nvdev->core_device.pdev; + u64 physical_size =3D pci_resource_len(pdev, nvidia_vgpu_fb_bar_index(pde= v)); + + /* The assigned vGPU type reports its VRAM aperture in MiB. */ + if (check_shl_overflow(nvdev->type_info.bar1_length, 20, size)) + return -EOVERFLOW; + + /* An unspecified or larger vGPU type aperture uses the physical limit. */ + if (!*size || *size > physical_size) + *size =3D physical_size; + + return 0; +} + +/* Enforce the same vGPU type aperture reported by GET_REGION_INFO. */ +static int nvidia_vgpu_bar1_access(struct vfio_device *core_vdev, + loff_t pos, size_t *count) +{ + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_vdev, struct nvidia_vgpu_pci_core_device, + core_device.vdev); + unsigned int index =3D VFIO_PCI_OFFSET_TO_INDEX(pos); + u64 size; + int ret; + + if (!*count || index !=3D nvidia_vgpu_fb_bar_index(nvdev->core_device.pde= v)) + return 0; + + ret =3D nvidia_vgpu_bar1_size(nvdev, &size); + if (ret) + return ret; + + pos &=3D VFIO_PCI_OFFSET_MASK; + if (pos >=3D size) + return -EINVAL; + + *count =3D min_t(u64, *count, size - pos); + return 0; +} + +static ssize_t nvidia_vgpu_read_config(struct vfio_device *core_vdev, + char __user *buf, size_t count, + loff_t *ppos) +{ + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_vdev, struct nvidia_vgpu_pci_core_device, + core_device.vdev); + loff_t pos =3D *ppos & VFIO_PCI_OFFSET_MASK; + loff_t next =3D *ppos, copy_offset; + size_t copy_count, register_offset; + __le16 value; + ssize_t ret; + + ret =3D vfio_pci_core_read(core_vdev, buf, count, &next); + if (ret <=3D 0) + return ret; + + /* The core reads the subsystem ID from physical configuration space. */ + if (vfio_pci_core_range_intersect_range(pos, ret, PCI_SUBSYSTEM_ID, + sizeof(value), ©_offset, + ©_count, ®ister_offset)) { + value =3D cpu_to_le16(nvdev->type_info.pci_subsys_id); + if (copy_to_user(buf + copy_offset, + (u8 *)&value + register_offset, copy_count)) + return -EFAULT; + } + + *ppos =3D next; + return ret; +} + +static ssize_t nvidia_vgpu_pci_read(struct vfio_device *core_vdev, + char __user *buf, size_t count, + loff_t *ppos) +{ + int ret; + + if (VFIO_PCI_OFFSET_TO_INDEX(*ppos) =3D=3D VFIO_PCI_CONFIG_REGION_INDEX) + return nvidia_vgpu_read_config(core_vdev, buf, count, ppos); + + ret =3D nvidia_vgpu_bar1_access(core_vdev, *ppos, &count); + if (ret) + return ret; + + return vfio_pci_core_read(core_vdev, buf, count, ppos); +} + +static ssize_t nvidia_vgpu_pci_write(struct vfio_device *core_vdev, + const char __user *buf, size_t count, + loff_t *ppos) +{ + int ret; + + ret =3D nvidia_vgpu_bar1_access(core_vdev, *ppos, &count); + if (ret) + return ret; + + /* Reset callbacks report a firmware FLR failure through the error IRQ. */ + return vfio_pci_core_write(core_vdev, buf, count, ppos); +} + +static int nvidia_vgpu_pci_mmap(struct vfio_device *core_vdev, + struct vm_area_struct *vma) +{ + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_vdev, struct nvidia_vgpu_pci_core_device, + core_device.vdev); + unsigned int index =3D vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHI= FT); + u64 size, req_start, req_len, end; + int ret; + + if (index !=3D nvidia_vgpu_fb_bar_index(nvdev->core_device.pdev)) + return vfio_pci_core_mmap(core_vdev, vma); + + ret =3D nvidia_vgpu_bar1_size(nvdev, &size); + if (ret) + return ret; + + req_start =3D (vma->vm_pgoff & (VFIO_PCI_OFFSET_MASK >> PAGE_SHIFT)) << P= AGE_SHIFT; + if (check_sub_overflow(vma->vm_end, vma->vm_start, &req_len) || + check_add_overflow(req_start, req_len, &end)) + return -EOVERFLOW; + if (end > size) + return -EINVAL; + + return vfio_pci_core_mmap(core_vdev, vma); +} + +static int nvidia_vgpu_get_region_info(struct vfio_device *core_vdev, + struct vfio_region_info *info, + struct vfio_info_cap *caps) +{ + struct pci_dev *pdev =3D to_pci_dev(core_vdev->dev); + int ret; + + ret =3D vfio_pci_ioctl_get_region_info(core_vdev, info, caps); + if (ret) + return ret; + + if (info->index =3D=3D nvidia_vgpu_fb_bar_index(pdev) && info->size) { + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_vdev, struct nvidia_vgpu_pci_core_device, + core_device.vdev); + u64 vgpu_bar1; + + ret =3D nvidia_vgpu_bar1_size(nvdev, &vgpu_bar1); + if (ret) + return ret; + + if (vgpu_bar1 && vgpu_bar1 < info->size) + info->size =3D vgpu_bar1; + } + + return 0; +} + +static long nvidia_vgpu_pci_ioctl(struct vfio_device *core_vdev, + unsigned int cmd, unsigned long arg) +{ + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_vdev, struct nvidia_vgpu_pci_core_device, + core_device.vdev); + bool reset =3D cmd =3D=3D VFIO_DEVICE_RESET || cmd =3D=3D VFIO_DEVICE_PCI= _HOT_RESET; + long ret; + + /* A failed firmware reset requires closing and reopening the instance. */ + if (reset && READ_ONCE(nvdev->reset_error)) + return READ_ONCE(nvdev->reset_error); + + ret =3D vfio_pci_core_ioctl(core_vdev, cmd, arg); + if (ret) + return ret; + + if (reset) + return READ_ONCE(nvdev->reset_error); + + return 0; +} + +static void nvidia_vgpu_reset_prepare(struct pci_dev *pdev) +{ + struct vfio_pci_core_device *core_device =3D dev_get_drvdata(&pdev->dev); + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_device, struct nvidia_vgpu_pci_core_device, + core_device); + int ret =3D 0; + + /* + * PCI holds the VF device lock, and VFIO may hold its memory lock. + * Open and close release instance_lock before entering vfio-pci-core; + * PF operations take only GPU locks, never the PF or VF device lock. + */ + mutex_lock(&nvdev->instance_lock); + nvdev->resetting =3D true; + reinit_completion(&nvdev->reset_completion); + if (nvdev->instance_active) + ret =3D nvdev->pf_ops->reset(nvdev->pf_context, nvdev->gfid); + if (ret && !nvdev->reset_error) + WRITE_ONCE(nvdev->reset_error, ret); + mutex_unlock(&nvdev->instance_lock); +} + +static void nvidia_vgpu_reset_done(struct pci_dev *pdev) +{ + struct vfio_pci_core_device *core_device =3D dev_get_drvdata(&pdev->dev); + struct nvidia_vgpu_pci_core_device *nvdev =3D + container_of(core_device, struct nvidia_vgpu_pci_core_device, + core_device); + int ret; + + mutex_lock(&nvdev->instance_lock); + ret =3D nvdev->reset_error; + /* PCI reset callbacks cannot return an error or abort the PCI reset. */ + if (ret) { + pci_err(pdev, "vGPU instance has a failed firmware reset: %d\n", ret); + vfio_pci_core_aer_err_detected(pdev, pci_channel_io_normal); + } + nvdev->resetting =3D false; + complete_all(&nvdev->reset_completion); + mutex_unlock(&nvdev->instance_lock); +} + +static const struct pci_error_handlers nvidia_vgpu_err_handlers =3D { + .error_detected =3D vfio_pci_core_aer_err_detected, + .reset_prepare =3D nvidia_vgpu_reset_prepare, + .reset_done =3D nvidia_vgpu_reset_done, +}; + +static const struct vfio_device_ops nvidia_vgpu_pci_ops =3D { + .name =3D NVIDIA_VGPU_DRIVER_NAME, + .init =3D vfio_pci_core_init_dev, + .release =3D vfio_pci_core_release_dev, + .open_device =3D nvidia_vgpu_open_device, + .close_device =3D nvidia_vgpu_close_device, + .ioctl =3D nvidia_vgpu_pci_ioctl, + .get_region_info_caps =3D nvidia_vgpu_get_region_info, + .device_feature =3D vfio_pci_core_ioctl_feature, + .read =3D nvidia_vgpu_pci_read, + .write =3D nvidia_vgpu_pci_write, + .mmap =3D nvidia_vgpu_pci_mmap, + .request =3D vfio_pci_core_request, + .match =3D vfio_pci_core_match, + .match_token_uuid =3D vfio_pci_core_match_token_uuid, + .bind_iommufd =3D vfio_iommufd_physical_bind, + .unbind_iommufd =3D vfio_iommufd_physical_unbind, + .attach_ioas =3D vfio_iommufd_physical_attach_ioas, + .detach_ioas =3D vfio_iommufd_physical_detach_ioas, + .pasid_attach_ioas =3D vfio_iommufd_physical_pasid_attach_ioas, + .pasid_detach_ioas =3D vfio_iommufd_physical_pasid_detach_ioas, +}; + +static int nvidia_vgpu_passthrough_open(struct vfio_device *core_vdev) +{ + struct vfio_pci_core_device *vdev =3D + container_of(core_vdev, struct vfio_pci_core_device, vdev); + int ret; + + ret =3D vfio_pci_core_enable(vdev); + if (ret) + return ret; + + vfio_pci_core_finish_enable(vdev); + return 0; +} + +/* PFs and VFs without a compatible Nova interface retain generic VFIO beh= avior. */ +static const struct vfio_device_ops nvidia_vgpu_passthrough_ops =3D { + .name =3D NVIDIA_VGPU_DRIVER_NAME, + .init =3D vfio_pci_core_init_dev, + .release =3D vfio_pci_core_release_dev, + .open_device =3D nvidia_vgpu_passthrough_open, + .close_device =3D vfio_pci_core_close_device, + .ioctl =3D vfio_pci_core_ioctl, + .get_region_info_caps =3D vfio_pci_ioctl_get_region_info, + .device_feature =3D vfio_pci_core_ioctl_feature, + .read =3D vfio_pci_core_read, + .write =3D vfio_pci_core_write, + .mmap =3D vfio_pci_core_mmap, + .request =3D vfio_pci_core_request, + .match =3D vfio_pci_core_match, + .match_token_uuid =3D vfio_pci_core_match_token_uuid, + .bind_iommufd =3D vfio_iommufd_physical_bind, + .unbind_iommufd =3D vfio_iommufd_physical_unbind, + .attach_ioas =3D vfio_iommufd_physical_attach_ioas, + .detach_ioas =3D vfio_iommufd_physical_detach_ioas, + .pasid_attach_ioas =3D vfio_iommufd_physical_pasid_attach_ioas, + .pasid_detach_ioas =3D vfio_iommufd_physical_pasid_detach_ioas, +}; + +static const struct vfio_pci_device_ops nvidia_vgpu_passthrough_dev_ops = =3D { + .get_dmabuf_phys =3D vfio_pci_core_get_dmabuf_phys, +}; + +static int nvidia_vgpu_pci_probe(struct pci_dev *pdev, + const struct pci_device_id *id) +{ + const struct vfio_device_ops *device_ops =3D &nvidia_vgpu_passthrough_ops; + struct nvidia_vgpu_pci_core_device *nvdev; + const struct nvidia_vgpu_ops *pf_ops =3D NULL; + const void *pf_context =3D NULL; + int vf_id =3D -1; + int ret; + + if (pdev->is_virtfn) { + const struct rust_ffi *ffi; + + ffi =3D pci_iov_borrow_rust_pf_data(pdev, &nvidia_vgpu_ffi_token, + NVIDIA_VGPU_FFI_ABI_MAJOR, + NVIDIA_VGPU_FFI_ABI_MINOR, + sizeof(*pf_ops)); + if (!IS_ERR(ffi)) { + const struct nvidia_vgpu_ops *ops =3D ffi->ops; + + if (ops->open && ops->close && ops->reset) { + pf_ops =3D ops; + pf_context =3D ffi->context; + device_ops =3D &nvidia_vgpu_pci_ops; + } + } + } + + if (pf_ops) { + vf_id =3D pci_iov_vf_id(pdev); + if (vf_id < 0) + return vf_id; + } + + nvdev =3D vfio_alloc_device(nvidia_vgpu_pci_core_device, core_device.vdev, + &pdev->dev, device_ops); + if (IS_ERR(nvdev)) + return PTR_ERR(nvdev); + + mutex_init(&nvdev->instance_lock); + init_completion(&nvdev->reset_completion); + nvdev->pf_ops =3D pf_ops; + nvdev->pf_context =3D pf_context; + nvdev->gfid =3D vf_id + 1; + if (!pf_ops) + nvdev->core_device.pci_ops =3D &nvidia_vgpu_passthrough_dev_ops; + dev_set_drvdata(&pdev->dev, &nvdev->core_device); + ret =3D vfio_pci_core_register_device(&nvdev->core_device); + if (ret) + goto out_put_vdev; + + return 0; + +out_put_vdev: + vfio_put_device(&nvdev->core_device.vdev); + return ret; +} + +static void nvidia_vgpu_pci_remove(struct pci_dev *pdev) +{ + struct vfio_pci_core_device *core_device =3D dev_get_drvdata(&pdev->dev); + + /* Drain all VFIO callbacks before the borrowed PF context expires. */ + vfio_pci_core_unregister_device(core_device); + vfio_put_device(&core_device->vdev); +} + +static const struct pci_device_id nvidia_vgpu_pci_table[] =3D { + /* RTX PRO 6000 Blackwell Server Edition PFs and VFs share this ID. */ + { PCI_DRIVER_OVERRIDE_DEVICE_VFIO(PCI_VENDOR_ID_NVIDIA, 0x2bb5), + .class =3D PCI_CLASS_DISPLAY_3D << 8, .class_mask =3D 0xffff00 }, + {} +}; +MODULE_DEVICE_TABLE(pci, nvidia_vgpu_pci_table); + +static int nvidia_vgpu_sriov_configure(struct pci_dev *pdev, int nr_virtfn) +{ + struct vfio_pci_core_device *vdev =3D dev_get_drvdata(&pdev->dev); + + if (!enable_sriov) + return -ENOENT; + + return vfio_pci_core_sriov_configure(vdev, nr_virtfn); +} + +static struct pci_driver nvidia_vgpu_pci_driver =3D { + .name =3D NVIDIA_VGPU_DRIVER_NAME, + .id_table =3D nvidia_vgpu_pci_table, + .probe =3D nvidia_vgpu_pci_probe, + .remove =3D nvidia_vgpu_pci_remove, + .sriov_configure =3D nvidia_vgpu_sriov_configure, + .err_handler =3D &nvidia_vgpu_err_handlers, + .driver_managed_dma =3D true, +}; +module_pci_driver(nvidia_vgpu_pci_driver); + +MODULE_DESCRIPTION("NVIDIA vGPU vfio-pci driver"); +MODULE_LICENSE("GPL");