From nobody Mon Sep 28 20:13:05 2026 Received: from mail-pj1-f50.google.com (mail-pj1-f50.google.com [209.85.216.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C11991F63D9 for ; Mon, 17 Aug 2026 19:10:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786993807; cv=none; b=adiiKrTOry0XXLzx1caYn3oMHPCBMLnRyQjDDKnWRSJRGXQJnvTzAUTbYjXluW8QiBZs0lLTXjqtRQno6lpmkibMsgfX5oYi62UkFfkYFaL1q+feTz2rKBF61pqWVbdYcizY9nYHsJN4Ab1d+I0Ltqdq4EoONYDTdxLVFn86p60= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786993807; c=relaxed/simple; bh=MmTOyUrolNwdXixWEA61+iSEN1nJRaq75qaAf+1wVcw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Wh/kUSmkouKovpqFFKuI1E9zPMAjd23zGfIpSeaQhzxMocKhlksE5IVyD/DVvGVB1HZLGDFv80gLTIW5mhH8rWjman7y2tt+AgjC8DlbS4xFvSpeWO88ztr7RZfpoaZry3xt9A/2ZqTTxDBahzLam0p/l0rNMm2+jA8dqNig2uw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=nebusec.ai; spf=pass smtp.mailfrom=nebusec.ai; dkim=pass (2048-bit key) header.d=nebusec.ai header.i=@nebusec.ai header.b=djoDY2FS; arc=none smtp.client-ip=209.85.216.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=nebusec.ai Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=nebusec.ai Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=nebusec.ai header.i=@nebusec.ai header.b="djoDY2FS" Received: by mail-pj1-f50.google.com with SMTP id 98e67ed59e1d1-38e3efab7e0so210048a91.0 for ; Mon, 17 Aug 2026 12:10:05 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=nebusec.ai; s=google; t=1786993805; x=1787598605; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=c/2C+RYvjKIOLmlpG91ZI2qQFAEpzRihy5VcIrSXsdU=; b=djoDY2FSvFK7RnkZS5fyF5m4iaFcAqWvNOr02/AJLDNs0WzlvagJluBrwJ+5JZ6jKQ LWbw9KOvO63JslTr75enZib7ZwGuq3YOgIPhO2jT7qrQvF85Wxj8R7vb8fidE+Ddt80X peDzfqCaU3qFwhx+sPQWJA0yH1g18/XWz/8yAyYiiYMq81IFYAqTnsUdTrdRPvHbUj3K k9lYVx9aF3VsET7nW3zytnfvITDdzihg4QXHPBr2+DjBkm9WPq5grGQuNdMSDejKy/+C +YEM8eCH4NeQu7FpBi9mwmqGFB3JIlt7ukQaBmSHCYWrmptxRp/vzUi9hmwEwZedabK+ SPqA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786993805; x=1787598605; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=c/2C+RYvjKIOLmlpG91ZI2qQFAEpzRihy5VcIrSXsdU=; b=KD0+NHI56GmwcNc+ZdhWIJYeZX6eSHOA96E+hLycNEXYJARAdaiqgElKCCOic6Jmmu j78bvEVopkmYjAbeLNjHIpUWl+pMjnW7n73vqjgSM9JDx5Zirn40Ca1M/CeG5JCPX+BS I1kRrtSfwNmzFYYyg8mlPCTD9A8ALdbKh11AqE/v3Y3OAOLVbX25V1WFVQc+BiSHXhaH nilCNGiVLLH5cNhJSp/7QKdsOerolQGrTZe55ZLkWzq34ZtjX65YCNY/wRhppjGanuh8 oiVJRX8rGNzjVPzYItc02EnvBN8G2sQ1R11ojf7F39DT6RysGLGt24kGxZovnEH8ZP01 Yh1w== X-Forwarded-Encrypted: i=1; AHgh+Rpn89PM6GylDA7RFQgvwOaHj/gJXG9+0DTF35UA3qjjRaVDnEJmp309VVj9R7qrsxfuk48LxD0JHCPXDAs=@vger.kernel.org X-Gm-Message-State: AOJu0YwpJcd59KB5iLaLJNefT2QXhe5HiH7MVNBbqU4ivEz5gLGBG8qP ZQfbEhp4ao7GxJ8sRBWPSfCam6KMNW/w7oblmEUDuzv1UsjskE8PGFELTr0Easwfv5MV X-Gm-Gg: AR+sD12puYu/DU8Dp+42JPoT0hwEDAHdhWfzY8SfeKGpkHc3MI0ARS6qnIH3xCj+Z5U h8fKOVGrO1KxbtjalWqdOYLjQwI2AAzPK3KoU+wrlx1kg5CqtzL7VEYSk++Gu9FPoo4QFbQ+w8k q23etb7hCd4Ds6VWdtFKXDE3F8XOYEt1DD+6h0I4uK2Rn49pechhtjOQSznluWbmJRrwzBa4KUM XgO8Kj2nSQNJFFo8tKYzKofV5986n4SQ2NjcWAFXbr9R93ZbbziiC5YvexJlJiTrKfHwS2/z1Ih PHbgeJ8FAmy2Xyms/ACQkb4K7BPzQjMWZN7+zlPY65Vex4IO9ofL6fhRWADwZok/S87gd55V+GK nzkqs0HBPeSZuRDN8vMpvXNOYFEY3T7+o1m1MsGAcw/5um8B3ABBkaIvRDlAxWtHcNTqarbcbcm j1E8Pe/5Dj7tQN7UvRdbM4yBmnfwHverQTp8StGKT09sMBJY5YZnD63A== X-Received: by 2002:a17:90b:390c:b0:38f:1a29:7a82 with SMTP id 98e67ed59e1d1-395638840ccmr52847a91.6.1786993804800; Mon, 17 Aug 2026 12:10:04 -0700 (PDT) Received: from gmail.com ([115.196.71.116]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-39531e8904dsm6145831a91.8.2026.08.17.12.09.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 17 Aug 2026 12:10:04 -0700 (PDT) From: Zihan Xi To: netdev@vger.kernel.org Cc: steffen.klassert@secunet.com, herbert@gondor.apana.org.au, davem@davemloft.net, edumazet@google.com, kuba@kernel.org, pabeni@redhat.com, horms@kernel.org, linux-kernel@vger.kernel.org, Zihan Xi , stable@vger.kernel.org, Eyal Birger , Vega Subject: [PATCH ipsec v3 1/1] xfrm: bound nat keepalive state collection Date: Mon, 17 Aug 2026 19:09:56 +0000 Message-ID: <4fadde7d28791a4e1196664cec692fdebce99bb5.1786987905.git.zihanx@nebusec.ai> X-Mailer: git-send-email 2.47.3 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The v1 nat keepalive fix allocates a GFP_ATOMIC object for every state while collecting references for phase two. This makes the worker's temporary memory use depend on the number of states and lets -ENOMEM abort the scan. Replace the allocated list with a fixed-size batch. When the batch is full, return a private walk status so xfrm_state_walk() leaves a cursor; drain the references after the walk releases xfrm_state_lock and resume from the cursor. This bounds temporary memory use and avoids the allocation failure path. The v1 fix also moved nat_keepalive_send() out of the walk callback. Keep the phase-two drain BH-disabled, as required by local_lock_nested_bh() used by the keepalive sockets. Fixes: 763fe700b7c5 ("xfrm: avoid lock inversion in nat keepalive work") Cc: stable@vger.kernel.org Cc: Eyal Birger Reported-by: Vega Assisted-by: Codex:gpt-5.4 Signed-off-by: Zihan Xi --- Changes in v3: - send an incremental fix on top of ipsec/master as requested by Steffen Klassert instead of replacing the queued v1 - replace v1's unbounded GFP_ATOMIC list with a bounded batch and xfrm_state_walk() cursor resume - keep the phase-two drain BH-disabled for local_lock_nested_bh() - distinguish the template-only userspace command from the actual in-kernel validation integration, configuration, build, and boot steps - validate 17 concurrent states so the v3 worker fills its 16-entry batch and resumes from its cursor - recapture the complete original lockdep report and add its full batch addr2line mapping from the matching unstripped vmlinux - record the validation-kernel user-namespace limitation and the separate NETLINK_XFRM EINVAL result instead of implying that the template userspace command reproduced the bug - identify the crash log as evidence from a separate earlier root-cause reproducer revision, not output from the 17-state v3 batch PoC - v2 Link: https://lore.kernel.org/all/cover.1785861392.git.zihanx@nebuse= c.ai/ Changes in v2: - reroll on top of net cf6f8b29befb - replace the unbounded GFP_ATOMIC state list with a bounded batch - keep phase-two processing in BH-disabled context - clarify the validation permission model and reproducer scope - v1 Link: https://lore.kernel.org/all/cover.1784645321.git.xizh2024@lzu.= edu.cn/ net/xfrm/xfrm_nat_keepalive.c | 46 ++++++++++++++++------------------- 1 file changed, 21 insertions(+), 25 deletions(-) diff --git a/net/xfrm/xfrm_nat_keepalive.c b/net/xfrm/xfrm_nat_keepalive.c index 8679c68c10a1..5cd6d43164db 100644 --- a/net/xfrm/xfrm_nat_keepalive.c +++ b/net/xfrm/xfrm_nat_keepalive.c @@ -155,32 +155,30 @@ static void nat_keepalive_send(struct nat_keepalive *= ka) } } =20 +enum { + NAT_KEEPALIVE_BATCH_SIZE =3D 16, + NAT_KEEPALIVE_BATCH_FULL =3D 1, +}; + struct nat_keepalive_work_ctx { - struct list_head states; + struct xfrm_state *batch[NAT_KEEPALIVE_BATCH_SIZE]; + unsigned int nr; time64_t next_run; time64_t now; }; =20 -struct nat_keepalive_state { - struct list_head list; - struct xfrm_state *x; -}; - static int nat_keepalive_work_collect(struct xfrm_state *x, int count, voi= d *ptr) { struct nat_keepalive_work_ctx *ctx =3D ptr; - struct nat_keepalive_state *state; =20 if (!READ_ONCE(x->nat_keepalive_interval)) return 0; =20 - state =3D kmalloc_obj(*state, GFP_ATOMIC); - if (!state) - return -ENOMEM; + if (ctx->nr =3D=3D ARRAY_SIZE(ctx->batch)) + return NAT_KEEPALIVE_BATCH_FULL; =20 xfrm_state_hold(x); - state->x =3D x; - list_add_tail(&state->list, &ctx->states); + ctx->batch[ctx->nr++] =3D x; return 0; } =20 @@ -226,29 +224,27 @@ static void nat_keepalive_work_single(struct xfrm_sta= te *x, =20 static void nat_keepalive_work(struct work_struct *work) { - struct nat_keepalive_state *state, *tmp; struct nat_keepalive_work_ctx ctx; struct xfrm_state_walk walk; struct net *net; - int err; + int err, i; =20 - INIT_LIST_HEAD(&ctx.states); ctx.next_run =3D 0; ctx.now =3D ktime_get_real_seconds(); =20 net =3D container_of(work, struct net, xfrm.nat_keepalive_work.work); xfrm_state_walk_init(&walk, IPPROTO_ESP, NULL); - err =3D xfrm_state_walk(net, &walk, nat_keepalive_work_collect, &ctx); + do { + ctx.nr =3D 0; + err =3D xfrm_state_walk(net, &walk, nat_keepalive_work_collect, &ctx); + local_bh_disable(); + for (i =3D 0; i < ctx.nr; i++) { + nat_keepalive_work_single(ctx.batch[i], &ctx); + xfrm_state_put(ctx.batch[i]); + } + local_bh_enable(); + } while (err =3D=3D NAT_KEEPALIVE_BATCH_FULL); xfrm_state_walk_done(&walk, net); - list_for_each_entry_safe(state, tmp, &ctx.states, list) { - nat_keepalive_work_single(state->x, &ctx); - xfrm_state_put(state->x); - kfree(state); - } - if (err =3D=3D -ENOMEM) { - schedule_delayed_work(&net->xfrm.nat_keepalive_work, 0); - return; - } if (ctx.next_run) schedule_delayed_work(&net->xfrm.nat_keepalive_work, (ctx.next_run - ctx.now) * HZ); --=20 2.43.0