From: Ankit Agrawal <ankita@nvidia.com>
NVIDIA's Grace based system have large GPU device memory. The device
memory is mapped as VM_PFNMAP in the VMM VMA. The nvgrace-gpu
module could make use of the huge PFNMAP support added in mm [1].
To achieve this, nvgrace-gpu module is updated to implement huge_fault ops.
The implementation establishes mapping according to the order request.
Note that if the PFN or the VMA address is unaligned to the order, the
mapping fallbacks to the PTE level.
Secondly, it is expected that the mapping not be re-established until
the GPU is ready post reset. Presence of the mappings during that time
could potentially leads to harmless corrected RAS events to be logged if
the CPU attempts to do speculative reads on the GPU memory on the Grace
systems.
It can take several seconds for the GPU to be ready. So it is desirable
that the time overlaps as much of the VM startup as possible to reduce
impact on the VM bootup time. The GPU readiness state is thus checked
on the first fault/huge_fault request which amortizes the GPU readiness
time. The GPU readiness is checked through BAR0 registers as is done
at the device probe.
Patch 1 updates the mapping mechanism to be done through faults.
Patch 2 splits the code to map at the various levels.
Patch 3 implements support for huge pfnmap.
Patch 4 move the code to map the BAR to a separate function.
Path 5-7 intercepts reset request and ensures that the GP is ready
before re-establishing the mapping after reset.
Applied over 6.18-rc6.
Link: https://lore.kernel.org/all/20240826204353.2228736-1-peterx@redhat.com/ [1]
Changelog:
v3:
- Moved the code for BAR mapping to a separate function.
- Added BAR0 mapping during open. Ensures BAR0 is mapped when registers
are checked. (Thanks Alex Williamson, Jason Gunthorpe for suggestion)
- Added check for GPU readiness on nvgrace_gpu_map_device_mem. (Thanks
Alex Williamson for the suggestion.
Link: https://lore.kernel.org/all/20251118074422.58081-1-ankita@nvidia.com/ [v2]
- Fixed build kernel warning
- subject text changes
- Rebased to 6.18-rc6.
Link: https://lore.kernel.org/all/20251117124159.3560-1-ankita@nvidia.com/ [v1]
Signed-off-by: Ankit Agrawal <ankita@nvidia.com>
Ankit Agrawal (7):
vfio/nvgrace-gpu: Use faults to map device memory
vfio: export function to map the VMA
vfio/nvgrace-gpu: Add support for huge pfnmap
vfio: export vfio_find_cap_start
vfio: move barmap to a separate function and export
vfio/nvgrace-gpu: split the code to wait for GPU ready
vfio/nvgrace-gpu: wait for the GPU mem to be ready
drivers/vfio/pci/nvgrace-gpu/main.c | 183 ++++++++++++++++++++++------
drivers/vfio/pci/vfio_pci_config.c | 3 +-
drivers/vfio/pci/vfio_pci_core.c | 84 ++++++++-----
include/linux/vfio_pci_core.h | 4 +
4 files changed, 207 insertions(+), 67 deletions(-)
--
2.34.1