From nobody Mon Sep 28 08:02:23 2026 Received: from mail-ej1-f71.google.com (mail-ej1-f71.google.com [209.85.218.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6791B43B3EA for ; Mon, 24 Aug 2026 15:29:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.71 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585380; cv=none; b=UCMWfl5Zu3gjtw7Z35CHFu+Ms2cltkiChlIdMIt8jRPtefgx0iLCe4VV1+qrCe1RvtpvMXD1nYuees5C1flntddLUooR4sTGppycnL0lz/xqcmLIUWWtLFqHYABDnSaJxpZPmOemAnFkrAndshB0pXrInitgagn+8kZoBGVePrI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585380; c=relaxed/simple; bh=wb3aYsExEyGsCbKmxC8jyjPeXblGtpkc9Lib1zp0eB4=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=ROWZW15a+jfmURoSQiosibkowXGRnMnotR+/w1VnOeo+Ozw6fwYvTad51P4q8CBHaCZPaQna90l5m7oSmuup3gN4gAwwwrfmvspQSNU5J1XSXgkmwrWPV28Ox9PUeMNvuJUbb3BLRSHcBfnpjfifXFkhTRtalVYkKqvHcUDyMYs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=g4fd5Xw3; arc=none smtp.client-ip=209.85.218.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="g4fd5Xw3" Received: by mail-ej1-f71.google.com with SMTP id a640c23a62f3a-c20e5890680so124434566b.1 for ; Mon, 24 Aug 2026 08:29:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787585376; x=1788190176; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=XxtYpttAkkFeNCs4heKfMrvlAQRBvH9LjM7sWbysPCs=; b=g4fd5Xw3tT6PsyB9wPLh4cY+UVPtPQKp+SUS9vm6SlnYKBo64ca3wWLVPMmLoYQzsE 5W/19tBu6I2RMP/iEwTP5maFjLeF/Crg9rmo6ayAtRiuDayooSRdtbqBCkdUfk1tA6XJ 9+62vFswcPs1sK9nZtXvJwKTwoYB1HwY/MZDXenF2aj7a9OMRIFxO3m2JGU/5hb43sA/ iYVIRdw7ldYyTJAvNON+mtX3c3frhiwAV3rVY1LS7BR9UA4xtPDoMoj4UIyWLJagCfrs j0stKsn6YGIqg/M1K3m2+M9QLc2TdYBe23YYhSO1z23tCJrSF7ZunGhoev1kef9eeu2E JP8Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787585376; x=1788190176; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=XxtYpttAkkFeNCs4heKfMrvlAQRBvH9LjM7sWbysPCs=; b=LDiivCQYY0JZFrfKGCjwxYrlP0Uy6IczeSuSf6pgjvFFUEGM0h/nieOIRYpHhd3pNN FDIms6YLBvFmMKpshURTKVK45RfCsOrtAu60MCaj27JOD4OsXEAjEl3XG5RUfTphgbUI oeXybqZNwAGqE4H03CqUJ3GEMuyaP6eokJuuxygk6AVS6XnJDW1JR7u84fDTOwZ8aMVg yCQ57IfyDf4txhIwYemZrj+Y5Cxr3VLrt4ERr/axEOJlPwcAc2Ada3dLOngF3DhnNhK2 +3aIS4VOgoirdJIcrppjShM1zLkZERLUOJ6eqqODcmkuw3itwknWjVD8Ok6jGUlAd5yX Wrdw== X-Forwarded-Encrypted: i=1; AHgh+RpFcKeDAYiIrMGF/W+CWkLcfDn0dLFpk+j4hItLezJrfRXaQjfmud1gPk/iOlpaH+xkC6bQLFuRlewvdZw=@vger.kernel.org X-Gm-Message-State: AFuF++nKfNmj96hULu68Z5dBqKLRNpvYs7hyZRQgw9duvvaQ97wmllO2 y/5LvFxNebVJ2QRTDn8fg/lx2DgmuNWmxCsHSF3Ff4lwqtTu+JOhNx8fAiJ0XowsvE5HrfTtO2c rm7tMcQ== X-Received: from ejcfb18.prod.google.com ([2002:a17:907:3a12:b0:c21:94a2:c168]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a17:907:c245:b0:c24:6f95:4f0e with SMTP id a640c23a62f3a-c2492719d94mr2282288466b.18.1787585376193; Mon, 24 Aug 2026 08:29:36 -0700 (PDT) Date: Mon, 24 Aug 2026 15:29:28 +0000 In-Reply-To: <20260824152932.1583506-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260615234220.3946885-1-lrizzo@google.com> <20260824152932.1583506-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.766.g2966f0265a-goog Message-ID: <20260824152932.1583506-2-lrizzo@google.com> Subject: [PATCH v2 1/5] swiotlb: enforce pool nareas and nslabs invariants From: Luigi Rizzo To: Marek Szyprowski , Robin Murphy , Willem de Bruijn , Kuniyuki Iwashima , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Luigi Rizzo , Luigi Rizzo Cc: Greg Kroah-Hartman , Dragos Tatulea , "Rafael J . Wysocki" , Andrew Morton , David Hildenbrand , netdev@vger.kernel.org, linux-mm@kvack.org, iommu@lists.linux.dev, driver-core@lists.linux.dev, linux-kernel@vger.kernel.org Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The SWIOTLB allocator relies on two runtime invariants across all pool initialization paths: 1. pool->nareas must always be a power of two so that a slot's area can be located efficiently via bitwise masking (index & (nareas - 1)) instead of integer division. 2. pool->nslabs must be a multiple of nareas * IO_TLB_SEGSIZE so that each area contains an integer multiple of IO_TLB_SEGSIZE (default 128) slots, preventing contiguous allocations from crossing area boundaries. Enforce these invariants consistently during early boot, pool initialization (swiotlb_init_io_tlb_pool), and restricted DMA pool setup. Fixes: 8ac04063354a ("swiotlb: reduce the number of areas to match actual m= emory pool size") Signed-off-by: Luigi Rizzo --- kernel/dma/swiotlb.c | 19 ++++++++++++++++--- 1 file changed, 16 insertions(+), 3 deletions(-) diff --git a/kernel/dma/swiotlb.c b/kernel/dma/swiotlb.c index 1abd3e6146f45..8e4bd9d47735a 100644 --- a/kernel/dma/swiotlb.c +++ b/kernel/dma/swiotlb.c @@ -33,6 +33,7 @@ #include #include #include +#include #include #include #include @@ -176,7 +177,7 @@ static void swiotlb_adjust_nareas(unsigned int nareas) static unsigned int limit_nareas(unsigned int nareas, unsigned long nslots) { if (nslots < nareas * IO_TLB_SEGSIZE) - return nslots / IO_TLB_SEGSIZE; + return rounddown_pow_of_two(nslots / IO_TLB_SEGSIZE); return nareas; } =20 @@ -269,7 +270,16 @@ static void swiotlb_init_io_tlb_pool(struct io_tlb_poo= l *mem, phys_addr_t start, unsigned long nslabs, bool late_alloc, unsigned int nareas) { void *vaddr =3D phys_to_virt(start); - unsigned long bytes =3D nslabs << IO_TLB_SHIFT, i; + unsigned long bytes, i; + + /* + * If we have multiple areas, ensure each area's size is a multiple of + * IO_TLB_SEGSIZE slots by aligning the total pool size down. + */ + if (nareas > 1) + nslabs =3D ALIGN_DOWN(nslabs, nareas * IO_TLB_SEGSIZE); + + bytes =3D nslabs << IO_TLB_SHIFT; =20 mem->nslabs =3D nslabs; mem->start =3D start; @@ -1813,7 +1823,10 @@ static int rmem_swiotlb_device_init(struct reserved_= mem *rmem, struct device *dev) { struct io_tlb_mem *mem =3D rmem->priv; - unsigned long nslabs =3D rmem->size >> IO_TLB_SHIFT; + unsigned long nslabs =3D round_down(rmem->size >> IO_TLB_SHIFT, IO_TLB_SE= GSIZE); + + if (!nslabs) + return -EINVAL; =20 /* Set Per-device io tlb area to one */ unsigned int nareas =3D 1; --=20 2.55.0.766.g2966f0265a-goog From nobody Mon Sep 28 08:02:23 2026 Received: from mail-ed1-f69.google.com (mail-ed1-f69.google.com [209.85.208.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A999043F4A2 for ; Mon, 24 Aug 2026 15:29:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585382; cv=none; b=J2g+BizxwQiVOvsMqePl7cDvOGFMiWWc3UUJiB5CK7O/4EPG4v9D0zmatc55MgcFwii99iLojEKDlp4mPP0neOgAy5oFq6Dr7MtJdfbzZvaS7vyIC6aYAaj+Lr2L8h45UsM1var2Fs2q6+tSdoUChWlFpIob5Ofs+PM0g/39idc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585382; c=relaxed/simple; bh=5NerewdcDfiL8OQ+QV5MZNyPF8loWtco0UZ0C++GWaM=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=pp/W3LQ0NTHXJXi06TF40Zb5Gwd5Bcv6y9FmXnm50FsrWPuht6VHD01C9NhKtIh/xQrctsfhqqF2gI1ikbj610lmNIwfLF6ksTb7Kf+8M831NpN/vpMNYBbPEUjWofDG9WStMp+08ERTxmMR1GVZpIT7wX5kYigcJpwwz28vCqg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=O8i/XY8+; arc=none smtp.client-ip=209.85.208.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="O8i/XY8+" Received: by mail-ed1-f69.google.com with SMTP id 4fb4d7f45d1cf-6a14f2e8c47so4327997a12.3 for ; Mon, 24 Aug 2026 08:29:39 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787585378; x=1788190178; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=a387GiIlsJuzM9aT5OA2RPOtHkxgW7YP7HmPN81inmc=; b=O8i/XY8+xls7Dym+G5C/g1wok3fECxQgnMc7a5hxNtIOQDNb5Wy8uj85coZ0syB4L4 bDlwwPSnw3Z27uCkumuIrjmMAfQIUHOHuVSMbzmOtV1YJSRnPWFLRpp5dmqyQ/LwkY0p OFh+74olWEvQ4rf+aufbTAGqiahSOLpLyxnC6j0oEAOKaBXKP2Z2QsdK+bTsWIhr/B6a THRWfsi6avRLJoE8L/q6woUmIS3YAqN5a/zZ/QTpf+v+Wa+KfDb/lRWuYO+v/657pDSj QD0Qd0zX6+c7cL8ETdHvgUfmBSaktKB28axnzvYe9fFY2/+UzHFFsbnjgIwDvQjaduI6 2brQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787585378; x=1788190178; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=a387GiIlsJuzM9aT5OA2RPOtHkxgW7YP7HmPN81inmc=; b=f5wsXjtFjThX1XSgzuUZb1MGsNzX2kSsDlLjANpIYDrK4NQAAJio1zPTbQKNVezlOE OGG488B2ic1Mut0ed2IEjulB+yxGbhHGzOajAjmXR8f6GtOsKkCLwm5Mm/IsAJflMMqR FbJLx1wXU5kUfIr3KRnGchzjfOONPen0zAPNrTd2w9jJRgh0IF1LFFtcdN8M9WBZ3TR4 Aybq0KXkzho+VBz6EpTeEMKTXTNPV6Ilz54G0/95vBTRLNqXiV/IErw+OC6DllEHlAAi Tzdm89XQnfq3l7nK60c07NcTGfHb/dQCl/tp6qpyVtcf8D3iUkMKocpiHXg94jDYwUiY eaqQ== X-Forwarded-Encrypted: i=1; AHgh+Rofd3xEvxOHQ0yVUfJ5NBB7zPX24xx8cKt73Ti01dUD4e+uNRiRtENCMvl/TPX8TYBru9sSqRdgD0ZwPqQ=@vger.kernel.org X-Gm-Message-State: AFuF++mUVo8R6TNqVm5pV+93qLgxA5wARvuadBpFGtdBLYS2LOQh3NNW hTB65udZDPq2F6mg5Wxk7KTYlLqQ7tumOFbVtKZYTJ+cn4mRQP/BI/iD2fPSQzQPacYlH5mmUfl G8derTQ== X-Received: from edxn5.prod.google.com ([2002:a05:6402:5c5:b0:6a1:4f77:f704]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6402:4613:b0:698:9e5e:5df8 with SMTP id 4fb4d7f45d1cf-6a42f1aca88mr27508443a12.7.1787585377536; Mon, 24 Aug 2026 08:29:37 -0700 (PDT) Date: Mon, 24 Aug 2026 15:29:29 +0000 In-Reply-To: <20260824152932.1583506-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260615234220.3946885-1-lrizzo@google.com> <20260824152932.1583506-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.766.g2966f0265a-goog Message-ID: <20260824152932.1583506-3-lrizzo@google.com> Subject: [PATCH v2 2/5] swiotlb/mm: Implement SWIOTLB nocopy page allocator From: Luigi Rizzo To: Marek Szyprowski , Robin Murphy , Willem de Bruijn , Kuniyuki Iwashima , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Luigi Rizzo , Luigi Rizzo Cc: Greg Kroah-Hartman , Dragos Tatulea , "Rafael J . Wysocki" , Andrew Morton , David Hildenbrand , netdev@vger.kernel.org, linux-mm@kvack.org, iommu@lists.linux.dev, driver-core@lists.linux.dev, linux-kernel@vger.kernel.org Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Introduce swiotlb_alloc_pages() and swiotlb_free_pages() to allocate and release compound pages directly from the default SWIOTLB pool. The allocator is restricted to slots in the static default pool, with caller-specified percentage limits on pool occupancy. This will be used for kernel data (e.g. socket buffers) in nocopy confidential computing. Signed-off-by: Luigi Rizzo --- include/linux/swiotlb.h | 24 ++++++ kernel/dma/swiotlb.c | 187 ++++++++++++++++++++++++++++++++++++++-- mm/page_alloc.c | 51 +++++++++++ 3 files changed, 255 insertions(+), 7 deletions(-) diff --git a/include/linux/swiotlb.h b/include/linux/swiotlb.h index 3dae0f592063e..4661e361c60e4 100644 --- a/include/linux/swiotlb.h +++ b/include/linux/swiotlb.h @@ -169,6 +169,23 @@ static inline struct io_tlb_pool *swiotlb_find_pool(st= ruct device *dev, return NULL; } =20 +bool swiotlb_pool_is_nocopy(struct io_tlb_pool *pool, phys_addr_t paddr); + +static inline bool swiotlb_addr_in_default_pool(struct device *dev, + phys_addr_t paddr) +{ + struct io_tlb_mem *mem =3D dev->dma_io_tlb_mem; + + return mem && paddr >=3D mem->defpool.start && paddr < mem->defpool.end; +} + +static inline bool swiotlb_is_nocopy_addr(struct device *dev, phys_addr_t = paddr) +{ + if (!swiotlb_addr_in_default_pool(dev, paddr)) + return false; + return swiotlb_pool_is_nocopy(&dev->dma_io_tlb_mem->defpool, paddr); +} + static inline bool is_swiotlb_force_bounce(struct device *dev) { struct io_tlb_mem *mem =3D dev->dma_io_tlb_mem; @@ -178,6 +195,13 @@ static inline bool is_swiotlb_force_bounce(struct devi= ce *dev) =20 void swiotlb_init(bool addressing_limited, unsigned int flags); void __init swiotlb_exit(void); +struct page *swiotlb_alloc_pages(struct device *dev, unsigned int order, g= fp_t gfp, + unsigned int percent); +bool swiotlb_free_pages(struct page *page, unsigned int order); +void swiotlb_nocopy_inc_ref(struct io_tlb_pool *pool, phys_addr_t phys); +void swiotlb_nocopy_dec_ref(struct io_tlb_pool *pool, phys_addr_t phys); +void swiotlb_prep_compound_page(struct page *page, unsigned int order); +void swiotlb_destroy_compound_page(struct page *page, unsigned int order); void swiotlb_dev_init(struct device *dev); size_t swiotlb_max_mapping_size(struct device *dev); bool is_swiotlb_allocated(void); diff --git a/kernel/dma/swiotlb.c b/kernel/dma/swiotlb.c index 8e4bd9d47735a..6b86a1e955fb4 100644 --- a/kernel/dma/swiotlb.c +++ b/kernel/dma/swiotlb.c @@ -66,17 +66,59 @@ /** * struct io_tlb_slot - IO TLB slot descriptor * @orig_addr: The original address corresponding to a mapped entry. + * @nocopy_refcnt: Lockless atomic refcount for Nocopy buffers. * @alloc_size: Size of the allocated buffer. * @list: The free list describing the number of free entries available * from each index. * @pad_slots: Number of preceding padding slots. Valid only in the first * allocated non-padding slot. + * @flags: Slot attributes (e.g. SWIOTLB_SLOT_NOCOPY for Nocopy buffers). + * + * The slot descriptor has states identified by @list and @flags (SWIOTLB_= SLOT_NOCOPY): + * + * 1. FREE (list > 0): + * Linear sweep free slot. + * + * 2. USED (list =3D=3D 0, SWIOTLB_SLOT_NOCOPY flag is NOT set in @flags): + * Allocated SWIOTLB bounce buffer. + * Fields used: @list, @pad_slots, @orig_addr, @alloc_size. + * + * 3. USED_NOCOPY (list =3D=3D 0, SWIOTLB_SLOT_NOCOPY flag is set in @flag= s): + * Allocated Nocopy SWIOTLB buffer. + * Fields used: @list, @nocopy_refcnt, @alloc_size. + */ +#define SWIOTLB_SLOT_NOCOPY BIT(0) + +/* + * SWIOTLB nocopy allocations (swiotlb_alloc_pages()) do not have an origi= nal + * physical address to bounce, but need to pass a caller-specified pool us= age + * limit (percentage) down to the area search logic. + * + * To avoid adding a parameter to swiotlb_find_slots(), swiotlb_search_are= a(), + * and swiotlb_search_pool_area(), the desired percentage (0..90) is encod= ed + * into the orig_addr parameter in the reserved high address range startin= g at + * INVALID_PHYS_ADDR (~0ULL). + * + * - NOCOPY_PCT_TO_ADDR(pct): Encodes a percentage into an orig_addr. + * - IS_SWIOTLB_NOCOPY(addr): Identifies a nocopy allocation request and + * restricts slot search to the static default pool. + * - NOCOPY_ADDR_TO_PCT(addr): Extracts the percentage to cap max_usable + * slots in swiotlb_search_pool_area(). */ +#define NOCOPY_PCT_MAX (90u) +#define NOCOPY_PCT_TO_ADDR(pct) (INVALID_PHYS_ADDR - min(pct, NOCOPY_PCT_= MAX)) +#define IS_SWIOTLB_NOCOPY(addr) ((addr) >=3D INVALID_PHYS_ADDR - NOCOPY_P= CT_MAX) +#define NOCOPY_ADDR_TO_PCT(addr) ((unsigned int)(INVALID_PHYS_ADDR - (addr= ))) + struct io_tlb_slot { - phys_addr_t orig_addr; + union { + phys_addr_t orig_addr; + atomic_t nocopy_refcnt; + }; size_t alloc_size; unsigned short list; unsigned short pad_slots; + unsigned int flags; }; =20 static bool swiotlb_force_bounce; @@ -300,6 +342,7 @@ static void swiotlb_init_io_tlb_pool(struct io_tlb_pool= *mem, phys_addr_t start, mem->slots[i].orig_addr =3D INVALID_PHYS_ADDR; mem->slots[i].alloc_size =3D 0; mem->slots[i].pad_slots =3D 0; + mem->slots[i].flags =3D 0; } =20 memset(vaddr, 0, bytes); @@ -869,12 +912,17 @@ static void swiotlb_bounce(struct device *dev, phys_a= ddr_t tlb_addr, size_t size enum dma_data_direction dir, struct io_tlb_pool *mem) { int index =3D (tlb_addr - mem->start) >> IO_TLB_SHIFT; - phys_addr_t orig_addr =3D mem->slots[index].orig_addr; size_t alloc_size =3D mem->slots[index].alloc_size; - unsigned long pfn =3D PFN_DOWN(orig_addr); unsigned char *vaddr =3D mem->vaddr + tlb_addr - mem->start; + phys_addr_t orig_addr; + unsigned long pfn; int tlb_offset; =20 + /* Nocopy swiotlb buffers do not need bouncing. */ + if (mem->slots[index].flags & SWIOTLB_SLOT_NOCOPY) + return; + + orig_addr =3D mem->slots[index].orig_addr; if (orig_addr =3D=3D INVALID_PHYS_ADDR) return; =20 @@ -904,6 +952,7 @@ static void swiotlb_bounce(struct device *dev, phys_add= r_t tlb_addr, size_t size size =3D alloc_size; } =20 + pfn =3D PFN_DOWN(orig_addr); if (PageHighMem(pfn_to_page(pfn))) { unsigned int offset =3D orig_addr & ~PAGE_MASK; struct page *page; @@ -1052,7 +1101,8 @@ static int swiotlb_search_pool_area(struct device *de= v, struct io_tlb_pool *pool unsigned long max_slots =3D get_max_slots(boundary_mask); unsigned int iotlb_align_mask =3D dma_get_min_align_mask(dev); unsigned int nslots =3D nr_slots(alloc_size), stride; - unsigned int offset =3D swiotlb_align_offset(dev, 0, orig_addr); + unsigned long max_usable =3D pool->area_nslabs; + unsigned int offset; unsigned int index, slots_checked, count =3D 0, i; unsigned long flags; unsigned int slot_base; @@ -1061,6 +1111,13 @@ static int swiotlb_search_pool_area(struct device *d= ev, struct io_tlb_pool *pool BUG_ON(!nslots); BUG_ON(area_index >=3D pool->nareas); =20 + if (IS_SWIOTLB_NOCOPY(orig_addr)) { + max_usable =3D (pool->area_nslabs * NOCOPY_ADDR_TO_PCT(orig_addr)) / 100; + orig_addr =3D 0; + } + + offset =3D swiotlb_align_offset(dev, 0, orig_addr); + /* * Historically, swiotlb allocations >=3D PAGE_SIZE were guaranteed to be * page-aligned in the absence of any other alignment requirements. @@ -1087,7 +1144,7 @@ static int swiotlb_search_pool_area(struct device *de= v, struct io_tlb_pool *pool stride =3D get_max_slots(max(alloc_align_mask, iotlb_align_mask)); =20 spin_lock_irqsave(&area->lock, flags); - if (unlikely(nslots > pool->area_nslabs - area->used)) + if (unlikely(area->used + nslots > max_usable)) goto not_found; =20 slot_base =3D area_index * pool->area_nslabs; @@ -1164,6 +1221,9 @@ static int swiotlb_search_pool_area(struct device *de= v, struct io_tlb_pool *pool * Search one memory area in all pools for a sequence of slots that match = the * allocation constraints. * + * If IS_SWIOTLB_NOCOPY(orig_addr) is true, the search is restricted to on= ly the + * default pool, which is what swiotlb_alloc_pages() is allowed to use. + * * Return: Index of the first allocated slot, or -1 on error. */ static int swiotlb_search_area(struct device *dev, int start_cpu, @@ -1177,6 +1237,9 @@ static int swiotlb_search_area(struct device *dev, in= t start_cpu, =20 rcu_read_lock(); list_for_each_entry_rcu(pool, &mem->pools, node) { + /* Only search the default pool (first in mem->pools) for nocopy allocat= ions. */ + if (IS_SWIOTLB_NOCOPY(orig_addr) && pool !=3D &mem->defpool) + break; if (cpu_offset >=3D pool->nareas) continue; area_index =3D (start_cpu + cpu_offset) & (pool->nareas - 1); @@ -1229,6 +1292,13 @@ static int swiotlb_find_slots(struct device *dev, ph= ys_addr_t orig_addr, goto found; } =20 + /* + * Passing a nocopy orig_addr restricts the search to only the + * default pool, so do not attempt dynamic pool expansion. + */ + if (IS_SWIOTLB_NOCOPY(orig_addr)) + return -1; + if (!mem->can_grow) return -1; =20 @@ -1468,11 +1538,16 @@ phys_addr_t swiotlb_tbl_map_single(struct device *d= ev, phys_addr_t orig_addr, return tlb_addr; } =20 +/* + * called with dev =3D=3D NULL from swiotlb_dealloc_pages(), in this case = force offset + * and align_mask to 0, pad_slots is also 0, and assume the pages come fro= m the + * default system pool. + */ static void swiotlb_release_slots(struct device *dev, phys_addr_t tlb_addr, struct io_tlb_pool *mem) { unsigned long flags; - unsigned int offset =3D swiotlb_align_offset(dev, 0, tlb_addr); + unsigned int offset =3D dev ? swiotlb_align_offset(dev, 0, tlb_addr) : 0; int index, nslots, aindex; struct io_tlb_area *area; int count, i; @@ -1506,6 +1581,7 @@ static void swiotlb_release_slots(struct device *dev,= phys_addr_t tlb_addr, mem->slots[i].orig_addr =3D INVALID_PHYS_ADDR; mem->slots[i].alloc_size =3D 0; mem->slots[i].pad_slots =3D 0; + mem->slots[i].flags =3D 0; } =20 /* @@ -1519,7 +1595,7 @@ static void swiotlb_release_slots(struct device *dev,= phys_addr_t tlb_addr, area->used -=3D nslots; spin_unlock_irqrestore(&area->lock, flags); =20 - dec_used(dev->dma_io_tlb_mem, nslots); + dec_used(dev ? dev->dma_io_tlb_mem : &io_tlb_default_mem, nslots); } =20 #ifdef CONFIG_SWIOTLB_DYNAMIC @@ -1912,3 +1988,100 @@ static const struct reserved_mem_ops rmem_swiotlb_o= ps =3D { =20 RESERVEDMEM_OF_DECLARE(dma, "restricted-dma-pool", &rmem_swiotlb_ops); #endif /* CONFIG_DMA_RESTRICTED_POOL */ + +static inline int swiotlb_nocopy_head_index(struct io_tlb_pool *pool, phys= _addr_t phys) +{ + return (page_to_phys(compound_head(phys_to_page(phys))) - pool->start) >>= IO_TLB_SHIFT; +} + +/** + * swiotlb_dealloc_pages() - Actually release Nocopy slots and page metada= ta + * @pool: SWIOTLB pool containing the buffer. + * @parent: Slot index of the buffer head. + */ +static void swiotlb_dealloc_pages(struct io_tlb_pool *pool, unsigned int p= arent) +{ + unsigned int order =3D get_order(pool->slots[parent].alloc_size); + phys_addr_t paddr =3D pool->start + (parent << IO_TLB_SHIFT); + struct page *head =3D phys_to_page(paddr); + + swiotlb_destroy_compound_page(head, order); + swiotlb_release_slots(NULL, paddr, pool); +} + +struct page *swiotlb_alloc_pages(struct device *dev, unsigned int order, + gfp_t gfp, unsigned int percent) +{ + struct io_tlb_pool *pool; + struct page *page; + int index, nslots, i; + + if (WARN_ON_ONCE(!dev || !dev->dma_io_tlb_mem)) + return NULL; + + if (dev->dma_io_tlb_mem !=3D &io_tlb_default_mem) + return NULL; + + index =3D swiotlb_find_slots(dev, NOCOPY_PCT_TO_ADDR(percent), + PAGE_SIZE << order, (PAGE_SIZE << order) - 1, + &pool); + if (index < 0) + return NULL; + + nslots =3D (PAGE_SIZE << order) >> IO_TLB_SHIFT; + page =3D phys_to_page(pool->start + (index << IO_TLB_SHIFT)); + swiotlb_prep_compound_page(page, order); + for (i =3D 0; i < nslots; i++) + pool->slots[index + i].flags |=3D SWIOTLB_SLOT_NOCOPY; + atomic_set(&pool->slots[index].nocopy_refcnt, 1); + return page; +} +EXPORT_SYMBOL(swiotlb_alloc_pages); + +bool swiotlb_free_pages(struct page *page, unsigned int order) +{ + struct io_tlb_mem *mem =3D &io_tlb_default_mem; + struct io_tlb_pool *pool =3D &mem->defpool; + struct page *head =3D compound_head(page); + unsigned int parent; + phys_addr_t paddr; + + paddr =3D page_to_phys(head); + if (paddr < pool->start || paddr >=3D pool->end) + return false; + + parent =3D swiotlb_nocopy_head_index(pool, paddr); + if (!(pool->slots[parent].flags & SWIOTLB_SLOT_NOCOPY)) + return false; + + if (atomic_dec_and_test(&pool->slots[parent].nocopy_refcnt)) + swiotlb_dealloc_pages(pool, parent); + + return true; +} +EXPORT_SYMBOL(swiotlb_free_pages); + +void swiotlb_nocopy_inc_ref(struct io_tlb_pool *pool, phys_addr_t phys) +{ + int head_idx =3D swiotlb_nocopy_head_index(pool, phys); + + atomic_inc(&pool->slots[head_idx].nocopy_refcnt); +} +EXPORT_SYMBOL(swiotlb_nocopy_inc_ref); + +void swiotlb_nocopy_dec_ref(struct io_tlb_pool *pool, phys_addr_t phys) +{ + int head_idx =3D swiotlb_nocopy_head_index(pool, phys); + + if (atomic_dec_and_test(&pool->slots[head_idx].nocopy_refcnt)) + swiotlb_dealloc_pages(pool, head_idx); +} +EXPORT_SYMBOL(swiotlb_nocopy_dec_ref); + +bool swiotlb_pool_is_nocopy(struct io_tlb_pool *pool, phys_addr_t paddr) +{ + int index =3D (paddr - pool->start) >> IO_TLB_SHIFT; + + return pool->slots[index].flags & SWIOTLB_SLOT_NOCOPY; +} +EXPORT_SYMBOL_GPL(swiotlb_pool_is_nocopy); diff --git a/mm/page_alloc.c b/mm/page_alloc.c index 083cbcb5bddec..ea148562a76d0 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -16,6 +16,7 @@ =20 #include #include +#include #include #include #include @@ -711,6 +712,56 @@ void prep_compound_page(struct page *page, unsigned in= t order) prep_compound_head(page, order); } =20 +#ifdef CONFIG_SWIOTLB +/* + * Prepare a SWIOTLB page (potentially compound). + * + * We explicitly initialize the head page refcount to 1 because recycled + * SWIOTLB pages might have a refcount of 0. + * + * If order > 0 (compound page), we must explicitly set all tail page + * refcounts to 0. This is because SWIOTLB pages might have a boot-default + * refcount of 1, but the core memory management subsystem expects tail pa= ges + * of a compound page to have a refcount of 0. + */ +void swiotlb_prep_compound_page(struct page *page, unsigned int order) +{ + init_page_count(page); + if (order > 0) { + for (int i =3D 1; i < (1 << order); i++) + set_page_count(page + i, 0); + prep_compound_page(page, order); + } +} + +/* + * Destroy a SWIOTLB compound page and restore page refcounts. + * + * When pages are returned to the SWIOTLB pool, we restore the refcount of + * all constituent pages (head and tails) to 1. This resets them to their + * clean boot-default state, ensuring they are ready for reuse either as + * individual order-0 pages or as part of a new compound allocation. + */ +void swiotlb_destroy_compound_page(struct page *page, unsigned int order) +{ + if (order > 0) { + struct folio *folio =3D (struct folio *)page; + + __ClearPageHead(page); + page[1].flags.f &=3D ~PAGE_FLAGS_SECOND; +#ifdef NR_PAGES_IN_LARGE_FOLIO + folio->_nr_pages =3D 0; +#endif + for (int i =3D 1; i < (1 << order); i++) { + page[i].mapping =3D NULL; + clear_compound_head(&page[i]); + set_page_count(page + i, 1); + } + } + set_page_count(page, 1); +} +#endif /* CONFIG_SWIOTLB */ + static inline void set_buddy_order(struct page *page, unsigned int order) { set_page_private(page, order); --=20 2.55.0.766.g2966f0265a-goog From nobody Mon Sep 28 08:02:23 2026 Received: from mail-ej1-f72.google.com (mail-ej1-f72.google.com [209.85.218.72]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8C53B43F4D9 for ; Mon, 24 Aug 2026 15:29:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.72 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585384; cv=none; b=qd15DfY/eOfNY0/cmkunkZHrSxZqyz8+6RBYD+oxL6GGvIQztvILAv65gCnw4km3BqlodU43Uts4ksrlcj9ETY6rrhe+0LXKwRbsdcwgALqqhOiftZ1uAn3nv8RHcZ1kJmQLswdZd24t8WLrFawPBlRS7wOncmgApqRtLGuIXxU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585384; c=relaxed/simple; bh=7RZvhwMCHTgYFqTlYPD6pPoVluRWmRUCimtR35dflb0=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=huCivDPCRdNlCqq5cNbq/4KVVMI73Aa1i8MuTtQu3HmoXfTOD3BvsxBkSA3FcBtoDnU9pCzYS5+wtI6VsFyvOHkDZR0VLkBrTKE13dT7Zv1AhePSylY1Umr2eUIrfEnYRw6QFIyyrjm15G/9X3jvpauCNJzrQKzG4WAJ3u78nNU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=B4ilhJ+T; arc=none smtp.client-ip=209.85.218.72 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="B4ilhJ+T" Received: by mail-ej1-f72.google.com with SMTP id a640c23a62f3a-c210c66672eso440376666b.2 for ; Mon, 24 Aug 2026 08:29:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787585379; x=1788190179; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=rIIxKJ/h4TdOX2PzM8xN9SJKh2HJ62O9uv+0COVNyBw=; b=B4ilhJ+TVc9EpVKWGZSOA42p/o3ClFtbInd7bdICsI9+WhK9rxFVy61eJrSce8u3w8 d107QflWvz95OJg54UbXWZitvJZCJzmPcC3ponJ1qfzqUbO/LB8M0i8I88Tc5y393nZR zrUGXXMNVoPEG67ssEzqGWhQR5ICOzvkPjGf0ahtFxkkirQTF3sPSkAalFZBQpPPrvOo Qh2O+IU3SX1NxH0dCflUyDpfru6ggrFIBQKiqQi12dzsI+qyOceQTglGwzU6TRkWL+73 O4nUDExUjR1y9lROPlXsLdqcP+/xj/07eC8O4d79WrpP89gF2paW3CvXtpKFsYUWI7BZ PRUQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787585379; x=1788190179; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rIIxKJ/h4TdOX2PzM8xN9SJKh2HJ62O9uv+0COVNyBw=; b=St2KwwNG+kJV5jOQvKBpCbLR9/lHe3c2Phkp5D1f08CwAQo6cW7qKrH8yK5BaMZYK8 sfHdt+wEQhUem2Tf/Gc1/DJRKnkDlcfeffpvWGSj/AbTqoX+2AhVk919lPUYsDMA2Bjc RZxmQy3MEA85/pvrEoTkqBtTwyyCq3OU6ok7z7HTH+Qo131TNx5ZO+2Tt/Xx4aj3Une3 VZztNhu96fa1y3YqnmGLzMV5pob16zwenzv46GqkZPNeV8ZOEw1sBs9jGE7tpe1QyEAQ q5R0/6Xa+/KPRAyeZjdFc5LtTQO56llXoGK5bSlhPE9GGcUnD8l9bFH6CAx6MgzioOQX xxDw== X-Forwarded-Encrypted: i=1; AHgh+Rrnrhx3sFHrelcNVGgOdIANCF8FPJN11NyU+/HLLG8/wtbJxUxVuV47K90ufSY5ijCOEonGwUol+0hqoLk=@vger.kernel.org X-Gm-Message-State: AFuF++mFrUXO/lETdvuA9V7pDXGdNRJzQfD01qROewrRA30U1BsYKvNw oFfNNptkxh1rJL6fHnSNkmf1yU++Gm0eo8+kunen6FrtPqhU0EPCzUrhEFOsqc2PhrPlxfqA4Xq CR6ixhw== X-Received: from ejzq1.prod.google.com ([2002:a17:906:b281:b0:c12:5229:7355]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a17:907:f509:b0:c19:49ee:19b9 with SMTP id a640c23a62f3a-c246a654f4emr3342635466b.15.1787585378467; Mon, 24 Aug 2026 08:29:38 -0700 (PDT) Date: Mon, 24 Aug 2026 15:29:30 +0000 In-Reply-To: <20260824152932.1583506-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260615234220.3946885-1-lrizzo@google.com> <20260824152932.1583506-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.766.g2966f0265a-goog Message-ID: <20260824152932.1583506-4-lrizzo@google.com> Subject: [PATCH v2 3/5] net/swiotlb: Track bounce device per socket From: Luigi Rizzo To: Marek Szyprowski , Robin Murphy , Willem de Bruijn , Kuniyuki Iwashima , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Luigi Rizzo , Luigi Rizzo Cc: Greg Kroah-Hartman , Dragos Tatulea , "Rafael J . Wysocki" , Andrew Morton , David Hildenbrand , netdev@vger.kernel.org, linux-mm@kvack.org, iommu@lists.linux.dev, driver-core@lists.linux.dev, linux-kernel@vger.kernel.org Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Record, on each network socket, the underlying hardware device that requests DMA mapping of tx packets. This is used for nocopy confidential computing, so that sockets can eventually allocate socket buffers directly from the swiotlb pools. Signed-off-by: Luigi Rizzo --- drivers/base/core.c | 1 + include/linux/netdevice.h | 21 +++++++++++ include/linux/swiotlb.h | 36 +++++++++++++++++++ include/net/sock.h | 46 ++++++++++++++++++++++++ kernel/dma/swiotlb.c | 73 +++++++++++++++++++++++++++++++++++++++ net/core/sock.c | 39 +++++++++++++++++++++ 6 files changed, 216 insertions(+) diff --git a/drivers/base/core.c b/drivers/base/core.c index 4c0c373998a19..091062228740d 100644 --- a/drivers/base/core.c +++ b/drivers/base/core.c @@ -3925,6 +3925,7 @@ void device_del(struct device *dev) unsigned int noio_flag; =20 device_lock(dev); + swiotlb_change_epoch(); kill_device(dev); device_unlock(dev); =20 diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h index 87cafc932e9e6..2457f4e464acf 100644 --- a/include/linux/netdevice.h +++ b/include/linux/netdevice.h @@ -5429,13 +5429,34 @@ static inline netdev_tx_t __netdev_start_xmit(const= struct net_device_ops *ops, return ops->ndo_start_xmit(skb, dev); } =20 +struct sock; + +#if defined(CONFIG_SWIOTLB) && !defined(CONFIG_PREEMPT_RT) +/* Per-CPU pointer to the socket currently performing transmission. Used + * to bridge the networking and DMA layers, allowing dma_map_page() to + * identify the socket originating the packet and apply SWIOTLB optimizati= ons. + */ +DECLARE_PER_CPU(struct sock *, current_tx_socket); +static inline struct sock *__save_current_tx_socket(struct sock *sk) +{ + struct sock *old_sk =3D this_cpu_read(current_tx_socket); + + this_cpu_write(current_tx_socket, sk); + return old_sk; +} +#else +static inline struct sock *__save_current_tx_socket(struct sock *sk) { ret= urn NULL; } +#endif + static inline netdev_tx_t netdev_start_xmit(struct sk_buff *skb, struct ne= t_device *dev, struct netdev_queue *txq, bool more) { + struct sock *old_sk =3D __save_current_tx_socket(skb->sk); const struct net_device_ops *ops =3D dev->netdev_ops; netdev_tx_t rc; =20 rc =3D __netdev_start_xmit(ops, skb, dev, more); + __save_current_tx_socket(old_sk); if (rc =3D=3D NETDEV_TX_OK) txq_trans_update(dev, txq); =20 diff --git a/include/linux/swiotlb.h b/include/linux/swiotlb.h index 4661e361c60e4..b1140db3cc397 100644 --- a/include/linux/swiotlb.h +++ b/include/linux/swiotlb.h @@ -202,6 +202,39 @@ void swiotlb_nocopy_inc_ref(struct io_tlb_pool *pool, = phys_addr_t phys); void swiotlb_nocopy_dec_ref(struct io_tlb_pool *pool, phys_addr_t phys); void swiotlb_prep_compound_page(struct page *page, unsigned int order); void swiotlb_destroy_compound_page(struct page *page, unsigned int order); +void swiotlb_safe_put_device(struct device *dev); + +/* Track epoch (number of delete operations) for leaf device info. */ +extern atomic_t global_device_epoch; + +static inline u32 swiotlb_dev_epoch(void) +{ + return atomic_read(&global_device_epoch); +} + +static inline void swiotlb_change_epoch(void) +{ + atomic_inc(&global_device_epoch); +} + +#if defined(CONFIG_NET) && !defined(CONFIG_PREEMPT_RT) +/* + * Track the socket for the currently transmitted packet, so the dma mappi= ng + * function can record there the leaf device if it needs bounce buffers. + */ +struct sock; +DECLARE_PER_CPU(struct sock *, current_tx_socket); +void sk_record_bounce_device(struct sock *sk, struct device *dev); +static inline void dma_learn_bounce_device(struct device *dev) +{ + struct sock *sk =3D this_cpu_read(current_tx_socket); + + if (sk) + sk_record_bounce_device(sk, dev); +} +#else +static inline void dma_learn_bounce_device(struct device *dev) {} +#endif void swiotlb_dev_init(struct device *dev); size_t swiotlb_max_mapping_size(struct device *dev); bool is_swiotlb_allocated(void); @@ -258,6 +291,9 @@ static inline phys_addr_t default_swiotlb_limit(void) { return 0; } +static inline void swiotlb_safe_put_device(struct device *dev) +{ +} #endif /* CONFIG_SWIOTLB */ =20 phys_addr_t swiotlb_tbl_map_single(struct device *hwdev, phys_addr_t phys, diff --git a/include/net/sock.h b/include/net/sock.h index 51185222aac29..39b5e81c7cc55 100644 --- a/include/net/sock.h +++ b/include/net/sock.h @@ -47,6 +47,7 @@ #include /* struct sk_buff */ #include #include +#include #include #include #include @@ -70,6 +71,14 @@ #include #include =20 +#if defined(CONFIG_SWIOTLB) && !defined(CONFIG_PREEMPT_RT) +struct sk_swiotlb_info { + struct device __rcu *dev; + u32 epoch; + unsigned long jiffies; +}; +#endif + /* * This structure really needs to be cleaned up. * Most of it is for TCP, and not used by any of @@ -602,8 +611,45 @@ struct sock { #if IS_ENABLED(CONFIG_PROVE_LOCKING) && IS_ENABLED(CONFIG_MODULES) struct module *sk_owner; #endif +#if defined(CONFIG_SWIOTLB) && !defined(CONFIG_PREEMPT_RT) + struct sk_swiotlb_info sk_swiotlb; +#endif }; =20 +#if defined(CONFIG_SWIOTLB) && !defined(CONFIG_PREEMPT_RT) +/* + * Clear bounce device on newly initialized or cloned sockets. + * Note: During socket cloning, sock_copy() performs a raw bitwise copy of + * the parent socket without incrementing the device refcount via get_devi= ce(). + * Therefore, we must zero sk_swiotlb.dev directly here without putting a + * reference. References are acquired solely by sk_record_bounce_device() = and + * released in sk_release_bounce_device(). + */ +static inline void sk_clear_bounce_device(struct sock *sk) +{ + rcu_assign_pointer(sk->sk_swiotlb.dev, NULL); +} + +/* + * Release any device reference acquired via sk_record_bounce_device() dur= ing + * socket transmission and clear the device pointer. Called during socket + * destruction (__sk_destruct). + */ +static inline void sk_release_bounce_device(struct sock *sk) +{ + struct device *dev; + + dev =3D rcu_dereference_raw(sk->sk_swiotlb.dev); + if (dev) { + swiotlb_safe_put_device(dev); + rcu_assign_pointer(sk->sk_swiotlb.dev, NULL); + } +} +#else +static inline void sk_clear_bounce_device(struct sock *sk) {} +static inline void sk_release_bounce_device(struct sock *sk) {} +#endif + struct sock_bh_locked { struct sock *sock; local_lock_t bh_lock; diff --git a/kernel/dma/swiotlb.c b/kernel/dma/swiotlb.c index 6b86a1e955fb4..d4a07a7c570e1 100644 --- a/kernel/dma/swiotlb.c +++ b/kernel/dma/swiotlb.c @@ -1683,6 +1683,8 @@ dma_addr_t swiotlb_map(struct device *dev, phys_addr_= t paddr, size_t size, phys_addr_t swiotlb_addr; dma_addr_t dma_addr; =20 + dma_learn_bounce_device(dev); + trace_swiotlb_bounced(dev, phys_to_dma(dev, paddr), size); =20 swiotlb_addr =3D swiotlb_tbl_map_single(dev, paddr, size, 0, dir, attrs); @@ -2085,3 +2087,74 @@ bool swiotlb_pool_is_nocopy(struct io_tlb_pool *pool= , phys_addr_t paddr) return pool->slots[index].flags & SWIOTLB_SLOT_NOCOPY; } EXPORT_SYMBOL_GPL(swiotlb_pool_is_nocopy); + +/* + * Dropping the reference to sk_swiotlb.dev must be done in two steps: + * + * 1. Readers inspect the pointer inside RCU critical sections without + * acquiring a reference. Use call_rcu() to wait for an RCU grace period + * to elapse so lockless in-flight readers finish accessing the device. + * + * 2. The RCU callback executes in atomic softirq context, but put_device() + * can block when releasing a device. Use schedule_work() to transition + * to sleepable process context where calling put_device() is safe. + */ +struct swiotlb_deferred_put { + struct rcu_head rcu; + struct work_struct work; + struct device *dev; +}; + +static void swiotlb_deferred_put_work(struct work_struct *work) +{ + struct swiotlb_deferred_put *dp =3D container_of(work, struct swiotlb_def= erred_put, work); + + /* Stage 2: Safely call put_device (can sleep) in process context */ + put_device(dp->dev); + kfree(dp); +} + +static void swiotlb_deferred_put_rcu(struct rcu_head *rcu) +{ + struct swiotlb_deferred_put *dp =3D container_of(rcu, struct swiotlb_defe= rred_put, rcu); + + /* RCU grace period has passed. Queue the work to do the actual put */ + schedule_work(&dp->work); +} + +/** + * swiotlb_safe_put_device() - Safely release device reference from atomic= /interrupt context + * @dev: The device structure to release. + * + * Enqueues a deferred put_device() call on a workqueue using GFP_ATOMIC. + * If memory allocation fails, the reference is leaked to avoid an immedia= te crash. + */ +void swiotlb_safe_put_device(struct device *dev) +{ + struct swiotlb_deferred_put *dp; + + if (!dev) + return; + + /* Lockless fast-path: if we are not the last reference, decrement is saf= e */ + if (refcount_dec_not_one(&dev->kobj.kref.refcount)) + return; + + /* + * On the last reference we must defer the final put_device() to task + * context because it will trigger device_release() which can sleep. + */ + dp =3D kmalloc_obj(*dp, GFP_ATOMIC); + if (dp) { + INIT_WORK(&dp->work, swiotlb_deferred_put_work); + dp->dev =3D dev; + /* Stage 1: Wait for RCU readers to finish */ + call_rcu(&dp->rcu, swiotlb_deferred_put_rcu); + } else { + pr_warn_ratelimited("swiotlb: failed to allocate deferred put, leaking d= evice ref\n"); + } +} +EXPORT_SYMBOL_GPL(swiotlb_safe_put_device); + +atomic_t global_device_epoch =3D ATOMIC_INIT(1); +EXPORT_SYMBOL(global_device_epoch); diff --git a/net/core/sock.c b/net/core/sock.c index 1ad41904db25b..ca3e08d3de141 100644 --- a/net/core/sock.c +++ b/net/core/sock.c @@ -103,6 +103,8 @@ #include #include #include +#include +#include #include #include #include @@ -152,6 +154,41 @@ =20 #include "dev.h" =20 +#if defined(CONFIG_SWIOTLB) && !defined(CONFIG_PREEMPT_RT) + +DEFINE_PER_CPU(struct sock *, current_tx_socket); +EXPORT_PER_CPU_SYMBOL(current_tx_socket); + +void sk_record_bounce_device(struct sock *sk, struct device *dev) +{ + struct device *old_dev; + + if (in_hardirq() || !sk_fullsock(sk) || sock_flag(sk, SOCK_ZEROCOPY)) + return; + + old_dev =3D rcu_dereference_protected(sk->sk_swiotlb.dev, 1); + + if (dev !=3D old_dev) { + /* Rate-limit updates to once per second to prevent bonding thrashing */ + if (old_dev && time_before(jiffies, sk->sk_swiotlb.jiffies + HZ)) + return; + + get_device(dev); + + /* Atomically swap in the new device and get the actual old one */ + old_dev =3D (struct device *)xchg((struct device __force **)&sk->sk_swio= tlb.dev, + (struct device __force *)dev); + + WRITE_ONCE(sk->sk_swiotlb.epoch, swiotlb_dev_epoch()); + sk->sk_swiotlb.jiffies =3D jiffies; + + /* Only drop the reference to the device we actually replaced */ + if (old_dev) + swiotlb_safe_put_device(old_dev); + } +} +EXPORT_SYMBOL(sk_record_bounce_device); +#endif static DEFINE_MUTEX(proto_list_mutex); static LIST_HEAD(proto_list); =20 @@ -2387,6 +2424,7 @@ static void __sk_destruct(struct rcu_head *head) __netns_tracker_free(net, &sk->ns_tracker, false); net_passive_dec(net); } + sk_release_bounce_device(sk); sk_prot_free(sk->sk_prot_creator, sk); } =20 @@ -2489,6 +2527,7 @@ struct sock *sk_clone(const struct sock *sk, const gf= p_t priority, goto out; =20 sock_copy(newsk, sk); + sk_clear_bounce_device(newsk); =20 newsk->sk_prot_creator =3D prot; #ifdef CONFIG_BPF_SYSCALL --=20 2.55.0.766.g2966f0265a-goog From nobody Mon Sep 28 08:02:23 2026 Received: from mail-ej1-f70.google.com (mail-ej1-f70.google.com [209.85.218.70]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BB687440A35 for ; Mon, 24 Aug 2026 15:29:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.70 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585385; cv=none; b=q6cy+ZTfhElyv5cvEDqtLRLNDUbV/DOnUezF7FPRrawTNq+CfV1DSca7iLv69V5oJU7KgukidBmnqRzHFLCQT+jDKvaGFAOH+pmhSmokSKKzFJRmrALqG+w1M0Q4l6KGpzupDUGwJVf72bCKoyqUgvdp9gEffW2A9wby4VaGxLs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585385; c=relaxed/simple; bh=+6KIs91WDIoo1NCHsUEn1NH39VCNgHI6xfWJXRyhn7c=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=qc3Psq2BHg+hG7xTjmvbHy+aKG9xgsxvrgOdn+fDdGMgm+L0KsbXHKrRvx9IRFdcreovjVw77Ta0sWQjOWHVRAYu1KqdFbhgG9PcnWiGi/tig29N4wcCsshADGeiHX/RRTg/Dtwej4aCHxP6qMr9PtOUeqekXNDauPKpq1SM1TE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=KpCPblvD; arc=none smtp.client-ip=209.85.218.70 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="KpCPblvD" Received: by mail-ej1-f70.google.com with SMTP id a640c23a62f3a-c167e032f03so288166266b.1 for ; Mon, 24 Aug 2026 08:29:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787585380; x=1788190180; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=+MUtQgSiF1UT5gfL8VJuu1JZ6DoMTlmedVjUdmZy7TI=; b=KpCPblvDJy2tMEq1IPjMdd04a+i8jlbqqVDFPGF45vrxkOF/sellLLxEB9ZUVyUK4C sN+CHykIbuQFIkRjDKMKWhajhWxTATcM/9wPezSbVx28+VhLI0PGZ6riA+N8J1coaqIX 7HmoujthLPmm+AHURxIzzckc8RXw6W9+zX1Vdr7qf/djDEbH1+za/+RBJYiwwKLoWIp1 fgvT6GREto2rKYxWd3CQLAq89mRfyGLq5bimM32xfHDMcJrSxUajFdOojq/N939GFks/ FZGsnwEnmIP12OgrIqE1+tshYhC/mO6VvT2rvv4EGZdPouHcHDu9MjNraQ5F0OwLoRhq VpJw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787585380; x=1788190180; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=+MUtQgSiF1UT5gfL8VJuu1JZ6DoMTlmedVjUdmZy7TI=; b=LCdYpwtqOPVGhd8M/F4bHVadmV0Jr60xGJqXo46ja+CyHP+r23bRz9dm07/aTx7Lly XEC8H3oR4+l+4yDvVvjULqOG1rdoEFeUEG9CKUHZM3x3t2drrqr536EB6HU17YugPlJG W3LRkvzZJCPAZRkZo83wpeIOeH8i5dh2dSW6QFCKDjRCsp10Ua9wVUKiSDn5pgt3DtB0 3rH0jJ4iIIFjBkj15NcCbgHVKijyPND8RwoLeB4qP+9BrFERG/t2Jyp9WEhRSGuWOPL5 nMHuKLMQdnlxxeTNLO0sWj8M/w/DxlVvtEPSE79sGUaZf3iOvU9+CGiPG/2OD8JtO+yL 1G4g== X-Forwarded-Encrypted: i=1; AHgh+RqjSHh3EmKxIu5fSaLuHJrFS2TpI4KJm9wA+HrU24YYhWUwyDM44MYL0CyH5nKvAao8Y0Jy5y41sAsD3ig=@vger.kernel.org X-Gm-Message-State: AFuF++mpjttLJWawijqOsAWHTup6xsN98MmmZkF6epvc2mQFi/hWpk+M TxaWN3vFWGaNYKYjveR6/3sesEJ6L6zGqEqAIjbKaoM1+Vw5Z4OsKthG3x35qc4bBL3yUkLWyG3 +XOd9AQ== X-Received: from ejdcu6.prod.google.com ([2002:a17:906:ba86:b0:c1f:97d8:cdd8]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a17:906:f59b:b0:c20:7cb1:9046 with SMTP id a640c23a62f3a-c24924b7508mr2143073766b.8.1787585380008; Mon, 24 Aug 2026 08:29:40 -0700 (PDT) Date: Mon, 24 Aug 2026 15:29:31 +0000 In-Reply-To: <20260824152932.1583506-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260615234220.3946885-1-lrizzo@google.com> <20260824152932.1583506-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.766.g2966f0265a-goog Message-ID: <20260824152932.1583506-5-lrizzo@google.com> Subject: [PATCH v2 4/5] net: Divert socket allocations to SWIOTLB for nocopy TX From: Luigi Rizzo To: Marek Szyprowski , Robin Murphy , Willem de Bruijn , Kuniyuki Iwashima , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Luigi Rizzo , Luigi Rizzo Cc: Greg Kroah-Hartman , Dragos Tatulea , "Rafael J . Wysocki" , Andrew Morton , David Hildenbrand , netdev@vger.kernel.org, linux-mm@kvack.org, iommu@lists.linux.dev, driver-core@lists.linux.dev, linux-kernel@vger.kernel.org Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Conditionally intercept socket buffer page allocations and direct them to the SWIOTLB page allocator when nocopy tx is active. This only happens when the swiotlb usage is below module parameter swiotlb.nocopy_tx_percent (default 0, range 0..90) A value of 0 disables the feature. Signed-off-by: Luigi Rizzo --- drivers/iommu/dma-iommu.c | 9 +++++- include/linux/skbuff.h | 7 ++++- include/linux/swiotlb.h | 2 ++ kernel/dma/direct.h | 11 +++++++ kernel/dma/swiotlb.c | 12 ++++++++ mm/page_alloc.c | 10 ++++++- net/core/sock.c | 62 ++++++++++++++++++++++++++++++++++----- 7 files changed, 102 insertions(+), 11 deletions(-) diff --git a/drivers/iommu/dma-iommu.c b/drivers/iommu/dma-iommu.c index 9a07eb39336eb..956d5e11b2896 100644 --- a/drivers/iommu/dma-iommu.c +++ b/drivers/iommu/dma-iommu.c @@ -1228,7 +1228,14 @@ dma_addr_t iommu_dma_map_phys(struct device *dev, ph= ys_addr_t phys, size_t size, * If both the physical buffer start address and size are page aligned, * we don't need to use a bounce page. */ - if (dev_use_swiotlb(dev, size, dir) && + bool is_nocopy =3D false; + + if (swiotlb_is_nocopy_addr(dev, phys)) { + swiotlb_nocopy_inc_ref(&dev->dma_io_tlb_mem->defpool, phys); + is_nocopy =3D true; + } + + if (!is_nocopy && dev_use_swiotlb(dev, size, dir) && iova_unaligned(iovad, phys, size)) { if (attrs & (DMA_ATTR_MMIO | DMA_ATTR_REQUIRE_COHERENT)) return DMA_MAPPING_ERROR; diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h index add0d282dea6e..d8f7041edc400 100644 --- a/include/linux/skbuff.h +++ b/include/linux/skbuff.h @@ -3786,7 +3786,12 @@ static inline void skb_frag_page_copy(skb_frag_t *fr= agto, fragto->netmem =3D fragfrom->netmem; } =20 -bool skb_page_frag_refill(unsigned int sz, struct page_frag *pfrag, gfp_t = prio); +/* nocopy swiotlb uses an additional non-null struct sock pointer. */ +bool __skb_page_frag_refill(unsigned int sz, struct page_frag *pfrag, gfp_= t prio, struct sock *sk); +static inline bool skb_page_frag_refill(unsigned int sz, struct page_frag = *pfrag, gfp_t prio) +{ + return __skb_page_frag_refill(sz, pfrag, prio, NULL); +} =20 /** * __skb_frag_dma_map - maps a paged fragment via the DMA API diff --git a/include/linux/swiotlb.h b/include/linux/swiotlb.h index b1140db3cc397..3baf52e6572d0 100644 --- a/include/linux/swiotlb.h +++ b/include/linux/swiotlb.h @@ -204,6 +204,8 @@ void swiotlb_prep_compound_page(struct page *page, unsi= gned int order); void swiotlb_destroy_compound_page(struct page *page, unsigned int order); void swiotlb_safe_put_device(struct device *dev); =20 +extern unsigned int nocopy_tx_percent; + /* Track epoch (number of delete operations) for leaf device info. */ extern atomic_t global_device_epoch; =20 diff --git a/kernel/dma/direct.h b/kernel/dma/direct.h index 7140c208c1238..21c65acb13823 100644 --- a/kernel/dma/direct.h +++ b/kernel/dma/direct.h @@ -88,6 +88,17 @@ static inline dma_addr_t dma_direct_map_phys(struct devi= ce *dev, { dma_addr_t dma_addr; =20 + if (swiotlb_is_nocopy_addr(dev, phys)) { + dma_addr_t unenc_addr =3D phys_to_dma_unencrypted(dev, phys); + + if (likely(dma_capable(dev, unenc_addr, size, true))) { + swiotlb_nocopy_inc_ref(&dev->dma_io_tlb_mem->defpool, phys); + if (!dev_is_dma_coherent(dev) && !(attrs & DMA_ATTR_SKIP_CPU_SYNC)) + arch_sync_dma_for_device(phys, size, dir); + return unenc_addr; + } + } + if (is_swiotlb_force_bounce(dev)) { if (!(attrs & DMA_ATTR_CC_SHARED)) { if (attrs & (DMA_ATTR_MMIO | DMA_ATTR_REQUIRE_COHERENT)) diff --git a/kernel/dma/swiotlb.c b/kernel/dma/swiotlb.c index d4a07a7c570e1..7b818a796ff96 100644 --- a/kernel/dma/swiotlb.c +++ b/kernel/dma/swiotlb.c @@ -63,6 +63,11 @@ */ #define IO_TLB_MIN_SLABS ((1<<20) >> IO_TLB_SHIFT) =20 +/* enable nocopy tx swiotlb and set the percentage of buffers allowed for = it. */ +unsigned int nocopy_tx_percent; +module_param(nocopy_tx_percent, uint, 0644); +MODULE_PARM_DESC(nocopy_tx_percent, "percentage of swiotlb buffer allowed = for nocopy tx"); + /** * struct io_tlb_slot - IO TLB slot descriptor * @orig_addr: The original address corresponding to a mapped entry. @@ -1640,6 +1645,13 @@ void __swiotlb_tbl_unmap_single(struct device *dev, = phys_addr_t tlb_addr, size_t mapping_size, enum dma_data_direction dir, unsigned long attrs, struct io_tlb_pool *pool) { + int index =3D (tlb_addr - pool->start) >> IO_TLB_SHIFT; + + if (pool->slots[index].flags & SWIOTLB_SLOT_NOCOPY) { + swiotlb_nocopy_dec_ref(pool, tlb_addr); + return; + } + /* * First, sync the memory before unmapping the entry */ diff --git a/mm/page_alloc.c b/mm/page_alloc.c index ea148562a76d0..32d5d630f9840 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -3002,9 +3002,14 @@ static void __free_frozen_pages(struct page *page, u= nsigned int order, { struct per_cpu_pages *pcp; struct zone *zone; - unsigned long pfn =3D page_to_pfn(page); + unsigned long pfn; int migratetype; =20 + if (unlikely(swiotlb_free_pages(page, order))) + return; + + pfn =3D page_to_pfn(page); + if (!pcp_allowed_order(order)) { __free_pages_ok(page, order, fpi_flags); return; @@ -3070,6 +3075,9 @@ void free_unref_folios(struct folio_batch *folios) unsigned long pfn =3D folio_pfn(folio); unsigned int order =3D folio_order(folio); =20 + if (unlikely(swiotlb_free_pages(&folio->page, order))) + continue; + if (!__free_pages_prepare(&folio->page, order, FPI_NONE)) continue; /* diff --git a/net/core/sock.c b/net/core/sock.c index ca3e08d3de141..ef40d1ff1de9f 100644 --- a/net/core/sock.c +++ b/net/core/sock.c @@ -188,6 +188,52 @@ void sk_record_bounce_device(struct sock *sk, struct d= evice *dev) } } EXPORT_SYMBOL(sk_record_bounce_device); + +/* + * Wrap alloc_pages in __skb_page_frag_refill(). If the socket's dma_devic= e requires + * SWIOTLB bounce buffering, divert allocation to the SWIOTLB slot allocat= or. + * This ensures the packet payload is written directly to a bounce buffer = from the start, + * enabling nocopy during driver DMA mapping. + */ +static inline struct page *alloc_any_pg(gfp_t gfp, unsigned int order, str= uct sock *sk) +{ + unsigned int pct =3D READ_ONCE(nocopy_tx_percent); + + if (sk && pct && !sock_flag(sk, SOCK_ZEROCOPY)) { + struct page *page =3D NULL; + bool release_dev =3D false; + struct device *dev; + + rcu_read_lock(); + dev =3D rcu_dereference(sk->sk_swiotlb.dev); + if (dev) { + /* + * The epoch check is just for cache invalidation, UAF is + * protected by the reference held in the sk. + */ + if (swiotlb_dev_epoch() !=3D READ_ONCE(sk->sk_swiotlb.epoch)) { + struct device __force **pdev =3D + (struct device __force **)&sk->sk_swiotlb.dev; + + release_dev =3D (cmpxchg(pdev, (struct device __force *)dev, + NULL) =3D=3D dev); + } else { + page =3D swiotlb_alloc_pages(dev, order, gfp, pct); + } + } + rcu_read_unlock(); + if (release_dev) + swiotlb_safe_put_device(dev); + if (page) + return page; + } + return alloc_pages(gfp, order); +} +#else +static inline struct page *alloc_any_pg(gfp_t gfp, unsigned int order, str= uct sock *sk) +{ + return alloc_pages(gfp, order); +} #endif static DEFINE_MUTEX(proto_list_mutex); static LIST_HEAD(proto_list); @@ -3213,7 +3259,7 @@ DEFINE_STATIC_KEY_FALSE(net_high_order_alloc_disable_= key); * no guarantee that allocations succeed. Therefore, @sz MUST be * less or equal than PAGE_SIZE. */ -bool skb_page_frag_refill(unsigned int sz, struct page_frag *pfrag, gfp_t = gfp) +bool __skb_page_frag_refill(unsigned int sz, struct page_frag *pfrag, gfp_= t gfp, struct sock *sk) { if (pfrag->page) { if (page_ref_count(pfrag->page) =3D=3D 1) { @@ -3229,27 +3275,27 @@ bool skb_page_frag_refill(unsigned int sz, struct p= age_frag *pfrag, gfp_t gfp) if (SKB_FRAG_PAGE_ORDER && !static_branch_unlikely(&net_high_order_alloc_disable_key)) { /* Avoid direct reclaim but allow kswapd to wake */ - pfrag->page =3D alloc_pages((gfp & ~__GFP_DIRECT_RECLAIM) | - __GFP_COMP | __GFP_NOWARN | - __GFP_NORETRY, - SKB_FRAG_PAGE_ORDER); + pfrag->page =3D alloc_any_pg((gfp & ~__GFP_DIRECT_RECLAIM) | + __GFP_COMP | __GFP_NOWARN | + __GFP_NORETRY, + SKB_FRAG_PAGE_ORDER, sk); if (likely(pfrag->page)) { pfrag->size =3D PAGE_SIZE << SKB_FRAG_PAGE_ORDER; return true; } } - pfrag->page =3D alloc_page(gfp); + pfrag->page =3D alloc_any_pg(gfp, 0, sk); if (likely(pfrag->page)) { pfrag->size =3D PAGE_SIZE; return true; } return false; } -EXPORT_SYMBOL(skb_page_frag_refill); +EXPORT_SYMBOL(__skb_page_frag_refill); =20 bool sk_page_frag_refill(struct sock *sk, struct page_frag *pfrag) { - if (likely(skb_page_frag_refill(32U, pfrag, sk->sk_allocation))) + if (likely(__skb_page_frag_refill(32U, pfrag, sk->sk_allocation, sk))) return true; =20 if (!sk->sk_bypass_prot_mem) --=20 2.55.0.766.g2966f0265a-goog From nobody Mon Sep 28 08:02:23 2026 Received: from mail-ed1-f70.google.com (mail-ed1-f70.google.com [209.85.208.70]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9939A442129 for ; Mon, 24 Aug 2026 15:29:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.70 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585385; cv=none; b=oplejXh/jtVQUdZSrJ4EZdbHGX7BdpajnrsPFQGsLEZ3aFE478J9xee50/5CGCuPQnTlGDkeiH/a8fUHsd0FQy0uCzp72zGShxUAJgHaUeDdYJBW2NEd/uVYULYc7VvrgaSZYGgOIl8rxXLNtvn8/zqxWMps7VchZDcgKuEC834= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787585385; c=relaxed/simple; bh=Y8jENIUNDeT2CaDx0lPDFpqJK/OjLgdPk5Qq7s+pqKQ=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=busZ1ngrnyz4pZ1ycF4+OSDARoHOdojcZDdK55pxSxzTmwUnrpdrEL871NWtZDYEfp0p04SYZ1dbzWDtWqzvlcGxciCoCwoF0iMjDm7R2GUj2NWd3Wy9SGKZbyxlXFGOZe56H+qLYcPQgNtNFs7E1GAWmPGB4fhKLaE/1xplpjk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=irk2HsTW; arc=none smtp.client-ip=209.85.208.70 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="irk2HsTW" Received: by mail-ed1-f70.google.com with SMTP id 4fb4d7f45d1cf-6a17bf3d42eso3965706a12.1 for ; Mon, 24 Aug 2026 08:29:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787585381; x=1788190181; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=JPV7mwpWvFlRgLJpo7McYGSN9ACQlDljrVAQMeggU7g=; b=irk2HsTWJIob0bXN0S/1VWBAdk/hIXUxQmHMNj/ZLztvvnPw+e4O7yiNJkdF4vwsnW QKpQvDuRRaqPu8LY5wScdbwQI/Z1FaEnzqfoPNo0g0kLxBdcXXikbnPySRuPPss12xzX kj6pd9ftfSSi18JSExUYEypmW623o5d2Fv523g9FHmfAtanZiDyssGNk3MY8lSF5I9PE EsBWLVBiVz1F2YPRkCDjj97pWQWYfEJTd5uVmpdvSWwEBcRmVkVk2u8fHA+K9FmHQiUv fvZ+ngiqC92nqsOd3AH7w6b9NuUiDj9rgLPNycqV2ijbP7ZxEWtfq5qv5mWzbSXFRjuQ +UfA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787585381; x=1788190181; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JPV7mwpWvFlRgLJpo7McYGSN9ACQlDljrVAQMeggU7g=; b=OULcE+QSCYaL940RBwJ6Mlt2kQ0tzA9caD4b7rDtjK59gop16RpzEchx+XmtJm5Xuv JGOLEyvskOXH6GNeHpnR/4dAEa++8qg2x5hdusznXNqKWBDhDT4mXQke52myvoTjw3Ow mPdrZ0qedKvyHyGwNPLG83dMZrMRwFjV0d35gJpUx/lLYec1vEsKiWS2DPN/KCFvNyko krUpg9AvAMXRzCp7DdjUr1kwDquyyvqk81jo/Pe9vG4tRhNaSy/CnNxL/rWo3iExhqqI MhkfFXu4YLuBqHFBQYrBdccgFNLK08JhmTpIELCqzM1wKAruY3ZbVltRbNAwRY9PaUkL jn0A== X-Forwarded-Encrypted: i=1; AHgh+RqLmxGDPUkZBd1G9ZrjfS0ufuOwEIyjjZmV1OW/4bSlRiGHOX7ynG6zhx2XevsUnH1+FhmfXrm3qkuQ2fM=@vger.kernel.org X-Gm-Message-State: AFuF++kuJ7BTf2ZKH8/sVjam9hwSCC2CwfuQ3PvFwmykZgqWHZl1I5qh Fqib6/gxmF42FFw5smebfTIuDCuEbm8kWQBaPLB/LBvvKTQrajTgrgIyRcZ75cOIJ86VxWhf/oW wqDSo0Q== X-Received: from edro3.prod.google.com ([2002:aa7:d3c3:0:b0:698:6d8c:dd8e]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6402:321a:b0:6a3:f3dc:d7f3 with SMTP id 4fb4d7f45d1cf-6a42f216053mr33518684a12.16.1787585381419; Mon, 24 Aug 2026 08:29:41 -0700 (PDT) Date: Mon, 24 Aug 2026 15:29:32 +0000 In-Reply-To: <20260824152932.1583506-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260615234220.3946885-1-lrizzo@google.com> <20260824152932.1583506-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.766.g2966f0265a-goog Message-ID: <20260824152932.1583506-6-lrizzo@google.com> Subject: [PATCH v2 5/5] swiotlb: Implement RX nocopy with fast recycling eviction From: Luigi Rizzo To: Marek Szyprowski , Robin Murphy , Willem de Bruijn , Kuniyuki Iwashima , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Luigi Rizzo , Luigi Rizzo Cc: Greg Kroah-Hartman , Dragos Tatulea , "Rafael J . Wysocki" , Andrew Morton , David Hildenbrand , netdev@vger.kernel.org, linux-mm@kvack.org, iommu@lists.linux.dev, driver-core@lists.linux.dev, linux-kernel@vger.kernel.org Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Conditionally divert receive buffer allocations in page_pool to the SWIOTLB page allocator. This only happens when swiotlb usage is below the threshold set by module parameter swiotlb.nocopy_rx_percent (default 0, range 0..90). A value of 0 disables the feature. To prevent existing DRAM or SWIOTLB pages from circulating indefinitely in the lockless receive ring after changing the parameter at runtime, __page_pool_put_page() checks residency against the active parameter state. Mismatched pages are immediately evicted back to their respective allocators, achieving rapid, lockless mode conversion across active network streams without requiring interface or queue resets. Signed-off-by: Luigi Rizzo --- include/linux/swiotlb.h | 1 + kernel/dma/swiotlb.c | 5 +++++ net/core/page_pool.c | 25 ++++++++++++++++++++++--- 3 files changed, 28 insertions(+), 3 deletions(-) diff --git a/include/linux/swiotlb.h b/include/linux/swiotlb.h index 3baf52e6572d0..f4597fd01c52d 100644 --- a/include/linux/swiotlb.h +++ b/include/linux/swiotlb.h @@ -205,6 +205,7 @@ void swiotlb_destroy_compound_page(struct page *page, u= nsigned int order); void swiotlb_safe_put_device(struct device *dev); =20 extern unsigned int nocopy_tx_percent; +extern unsigned int nocopy_rx_percent; =20 /* Track epoch (number of delete operations) for leaf device info. */ extern atomic_t global_device_epoch; diff --git a/kernel/dma/swiotlb.c b/kernel/dma/swiotlb.c index 7b818a796ff96..91f175c34a34e 100644 --- a/kernel/dma/swiotlb.c +++ b/kernel/dma/swiotlb.c @@ -129,6 +129,11 @@ struct io_tlb_slot { static bool swiotlb_force_bounce; static bool swiotlb_force_disable; =20 +/* enable nocopy rx swiotlb and set the percentage of buffers allowed for = it. */ +unsigned int nocopy_rx_percent; +module_param(nocopy_rx_percent, uint, 0644); +MODULE_PARM_DESC(nocopy_rx_percent, "percentage of swiotlb buffer allowed = for nocopy rx"); + #ifdef CONFIG_SWIOTLB_DYNAMIC =20 static void swiotlb_dyn_alloc(struct work_struct *work); diff --git a/net/core/page_pool.c b/net/core/page_pool.c index 50ee550fef73a..fe8839a7c70a8 100644 --- a/net/core/page_pool.c +++ b/net/core/page_pool.c @@ -19,6 +19,7 @@ =20 #include #include +#include #include #include /* for put_page() */ #include @@ -578,10 +579,16 @@ static bool page_pool_dma_map(struct page_pool *pool,= netmem_ref netmem, gfp_t g static struct page *__page_pool_alloc_page_order(struct page_pool *pool, gfp_t gfp) { + unsigned int pct =3D READ_ONCE(nocopy_rx_percent); struct page *page; =20 gfp |=3D __GFP_COMP; - page =3D alloc_pages_node(pool->p.nid, gfp, pool->p.order); + page =3D NULL; + if (pct && is_swiotlb_active(pool->p.dev)) + page =3D swiotlb_alloc_pages(pool->p.dev, pool->p.order, gfp, + pct); + if (!page) + page =3D alloc_pages_node(pool->p.nid, gfp, pool->p.order); if (unlikely(!page)) return NULL; =20 @@ -616,8 +623,9 @@ static noinline netmem_ref __page_pool_alloc_netmems_sl= ow(struct page_pool *pool if ((gfp & GFP_ATOMIC) =3D=3D GFP_ATOMIC) gfp |=3D __GFP_NOWARN; =20 - /* Don't support bulk alloc for high-order pages */ - if (unlikely(pp_order)) + /* Don't support bulk alloc for high-order pages or nocopy SWIOTLB */ + if (unlikely(pp_order || (READ_ONCE(nocopy_rx_percent) && + is_swiotlb_active(pool->p.dev)))) return page_to_netmem(__page_pool_alloc_page_order(pool, gfp)); =20 /* Unnecessary as alloc cache is empty, but guarantees zero count */ @@ -835,6 +843,17 @@ __page_pool_put_page(struct page_pool *pool, netmem_re= f netmem, { lockdep_assert_no_hardirq(); =20 + /* + * If runtime nocopy mode toggled, evict circulating buffers immediately + * back to their respective allocators rather than recycling them. + */ + if (unlikely(!netmem_is_net_iov(netmem) && + swiotlb_is_nocopy_addr(pool->p.dev, page_to_phys(netmem_to_page(net= mem))) !=3D + (READ_ONCE(nocopy_rx_percent) > 0))) { + page_pool_return_netmem(pool, netmem); + return 0; + } + /* This allocator is optimized for the XDP mode that uses * one-frame-per-page, but have fallbacks that act like the * regular page allocator APIs. --=20 2.55.0.766.g2966f0265a-goog