From nobody Mon Aug 24 20:41:02 2026 Received: from mail-pl1-f181.google.com (mail-pl1-f181.google.com [209.85.214.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C0F484756B0 for ; Wed, 29 Jul 2026 11:29:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.181 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785324579; cv=none; b=nX5IXaGVC1cCc3Lay78R1E/PUrcewhCiEh8omqM4A0Ji4uv5bSdaliqgliA0oDPfMFKvDpTVAIQ8Laboawvj2pNwVYj6QjEhp4HpPUcUnN5M3vvXNg6V4TTP+g6t6mZfAipgICkAY2GcRm0gBaSbAeAagm1RqftdiELJz8HP89o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785324579; c=relaxed/simple; bh=fONyvsGqe4281Xuw95Mc8exWac5SoFKftqs+5AIR3YU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=bVKmLFjSBKB/H4hoGQ4Pj87Bjj341PnUBhJvWvkRdkZUxuPXq1XBVaYm7GBoYwo5KL0aPJm0FjNsA/UUMkJO6cIymAxs9Nd6alBV+P2SICY/zKnDzqlkeEyioaSb+w+AB9TOTli2foINBBq0mJkVlP4Exjw9Xh8d3+X1jCZeOIQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=nebusec.ai; spf=pass smtp.mailfrom=nebusec.ai; dkim=pass (2048-bit key) header.d=nebusec.ai header.i=@nebusec.ai header.b=uwgzyxAf; arc=none smtp.client-ip=209.85.214.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=nebusec.ai Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=nebusec.ai Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=nebusec.ai header.i=@nebusec.ai header.b="uwgzyxAf" Received: by mail-pl1-f181.google.com with SMTP id d9443c01a7336-2ce87c7e3bbso9459125ad.1 for ; Wed, 29 Jul 2026 04:29:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=nebusec.ai; s=google; t=1785324577; x=1785929377; darn=lists.linux.dev; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=H0RSh7anhKhi3o3iqmDgvhVKZLB+ONGiikOgDOrEhr8=; b=uwgzyxAfORBD0oMym6waU2NOWLp3ddGhiqfssZpAuLZyBEwxJMGuClRwV1+NngYPVM v2VGXBdhioMOz1FVAKmrQDkoW4d/uZ1lkrYEjqAzcptqgv8CbcVabuV2Zd/tAfkTb4Ok AFxA0j4ul/hm7LSRieHcBgdq6msbqhtgPB192BO8uA4sMdhv9Ecmumo1cTyeQpFGyMSS Lx38/ThcV6kEOE99P1jdQydFjjO0KVKw/MdNyn6gWhdJJ+ORp5u9oBA2mZzixFdxymiE o0k9ghLITYdbPeyjga96Zd4xD+3801McOPO7qYBSCc/4FHn6Y/TJfAKrdJnlGF6uZeS2 jgqw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785324577; x=1785929377; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=H0RSh7anhKhi3o3iqmDgvhVKZLB+ONGiikOgDOrEhr8=; b=SuxwhS+W0EtK8H5o89kMrA691M9BBxg2CDeQgUe/7QDpexLY1LkWqQUfJG3OQ9iPZx 0XeDJXIaB8ttdEItDn7J8hVXOC81lnPL+pZf5TRKj80aR5HTSsPGUJyuE9hy6/vl0L1q XHZOMmuxQdoDihvfNnkMwdFanPHQiVg6t/nJOerQJhQoS2tKOAiiIFudTCG9QlFQY3gk EzjppunlPfHh0+pa6xPMVv45rPfPfLy/NUdJUaH82oPjyYJbIZ7DZgrs0YYRGCrEMJJ1 J0ie9g3m35j/Clr9+jEVuhMeF76sI/nXtR6tzoDZxoTmr0wt7Ofoes/KpjgExlJ1pxzU xPAA== X-Forwarded-Encrypted: i=1; AHgh+RqWC8MbR7tqlarD7SxgSJXwq88ccQHZNSbO87WAZQdR2HktJEcz7e8CjiOcEWzQ4bQms7Gm2Q==@lists.linux.dev X-Gm-Message-State: AOJu0YwnfxrpRM6CvoTVwL3o6g9EGtqZB6aLZeAIVAj6IToNddVyeqEn zZkbU5TdjIvfjqqu3dbEWvTzSlThi3GqCXhGc/u6v46ujUUUh3UTnNIf4XFDzRgl53xg X-Gm-Gg: AR+sD10l3bUB7obgd5GV8zOth53/g6IaXuaZUoEt+JsonlZOdhwVNxhxzw9MU/PFDtX XGQ/UrUueWkW+dzN1YE4gHMApcSVSjYf993kcOssmdfbhYqXVdC6qB5LRtbiWb6u3V1iQSPLICP 3ArHRtjtnzy2Tfa9HS1+okifAP/n8zMRIGkHp31S8e/oAZfhL7ufNSwmZbRw9ebW4/X/j9/aMcH UjxqYeMXUvIW8/5blD66ipr6T3YFr7EL0vccR55rTZPADvwbu3q2t/+TPSyIrPjFoRTEMxA0VBt dHkknuAG/uwTY9DGtcp/bq/w1wOqqgVYVWBXoa1LJkAeg4gNtKe7IbfqTD40NMvv16pS+6DjIGU 9qPk8Oo8wudwJi7buATRK6g2FrvSRlEU4+VcMJnVUL8TkdWWSeh/Lrx62XhvluwmDlccAhwE3OJ QOM3Ob7t3DUusGoyV0YT6QOKnqMpX3bWFs/mOLc0xxx0mlSK+a2Cuu39yMt24RvWhij2e6QgM= X-Received: by 2002:a17:902:ce83:b0:2bd:8395:fedd with SMTP id d9443c01a7336-2d015d874d1mr69608405ad.37.1785324576972; Wed, 29 Jul 2026 04:29:36 -0700 (PDT) Received: from localhost.localdomain ([115.192.250.185]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d022c55a7dsm9898505ad.84.2026.07.29.04.29.27 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 29 Jul 2026 04:29:36 -0700 (PDT) From: Zihan Xi To: netdev@vger.kernel.org, mptcp@lists.linux.dev Cc: davem@davemloft.net, edumazet@google.com, pabeni@redhat.com, horms@kernel.org, ncardwell@google.com, kuniyu@google.com, dsahern@kernel.org, kees@kernel.org, matttbe@kernel.org, martineau@kernel.org, geliang@kernel.org, fw@strlen.de, vega@nebusec.ai, zihanx@nebusec.ai Subject: [PATCH net 1/2] tcp: diag: fix unbounded bucket lock hold in tcp_diag_dump() Date: Wed, 29 Jul 2026 11:28:39 +0000 Message-ID: X-Mailer: git-send-email 2.47.3 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" inet_diag dumps execute attacker-controlled bytecode through inet_diag_bc_sk(). tcp_diag_dump() currently evaluates socket filters and runs that bytecode while holding the listener, bind and ehash bucket locks. A dump request can therefore force unbounded per-bucket lock hold by arranging for many sockets in the same bucket to fail the pre-bytecode filters, so the old 16-entry batching limit no longer bounds the locked walk itself. Under load this can trigger soft lockups and may escalate to a watchdog panic. Fix this by making each locked section collect only referenced sockets. Move all netns/family/port checks, inet_diag_bc_sk(), and fill work out of the bucket locks so the batch limit bounds raw bucket traversal rather than only filter hits. For listener, bind and ehash buckets, keep a referenced dump cursor so restarts resume after the previous socket instead of rescanning the bucket head under the same lock. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: Codex:gpt-5.4 Signed-off-by: Zihan Xi --- include/linux/inet_diag.h | 15 ++ net/ipv4/inet_diag.c | 13 ++ net/ipv4/tcp_diag.c | 316 +++++++++++++++++++++++++++----------- 3 files changed, 257 insertions(+), 87 deletions(-) diff --git a/include/linux/inet_diag.h b/include/linux/inet_diag.h index 704fd415c2b4..4859e77a28c7 100644 --- a/include/linux/inet_diag.h +++ b/include/linux/inet_diag.h @@ -6,6 +6,7 @@ #include =20 struct inet_hashinfo; +struct sock; =20 struct inet_diag_handler { struct module *owner; @@ -32,12 +33,24 @@ struct inet_diag_handler { }; =20 struct bpf_sk_storage_diag; + +enum inet_diag_dump_cursor_type { + INET_DIAG_DUMP_CURSOR_NONE, + INET_DIAG_DUMP_CURSOR_TCP_LISTEN, + INET_DIAG_DUMP_CURSOR_TCP_BIND, + INET_DIAG_DUMP_CURSOR_TCP_EHASH, + INET_DIAG_DUMP_CURSOR_MPTCP_LISTEN, +}; + struct inet_diag_dump_data { struct nlattr *req_nlas[__INET_DIAG_REQ_MAX]; #define inet_diag_nla_bc req_nlas[INET_DIAG_REQ_BYTECODE] #define inet_diag_nla_bpf_stgs req_nlas[INET_DIAG_REQ_SK_BPF_STORAGES] =20 struct bpf_sk_storage_diag *bpf_stg_diag; + struct sock *dump_cursor; + unsigned int dump_cursor_slot; + u8 dump_cursor_type; bool mark_needed; /* INET_DIAG_BC_MARK_COND present. */ #ifdef CONFIG_SOCK_CGROUP_DATA bool cgroup_needed; /* INET_DIAG_BC_CGROUP_COND present. */ @@ -53,6 +66,8 @@ int inet_sk_diag_fill(struct sock *sk, struct inet_connec= tion_sock *icsk, =20 int inet_diag_bc_sk(const struct inet_diag_dump_data *cb_data, struct sock= *sk); =20 +void inet_diag_dump_clear_cursor(struct inet_diag_dump_data *cb_data); + void inet_diag_msg_common_fill(struct inet_diag_msg *r, struct sock *sk); =20 static inline size_t inet_diag_msg_attrs_size(void) diff --git a/net/ipv4/inet_diag.c b/net/ipv4/inet_diag.c index 34b77aa87d0a..41148e880054 100644 --- a/net/ipv4/inet_diag.c +++ b/net/ipv4/inet_diag.c @@ -891,10 +891,23 @@ static int inet_diag_dump_start_compat(struct netlink= _callback *cb) return __inet_diag_dump_start(cb, sizeof(struct inet_diag_req)); } =20 +void inet_diag_dump_clear_cursor(struct inet_diag_dump_data *cb_data) +{ + if (!cb_data->dump_cursor) + return; + + sock_gen_put(cb_data->dump_cursor); + cb_data->dump_cursor =3D NULL; + cb_data->dump_cursor_slot =3D 0; + cb_data->dump_cursor_type =3D INET_DIAG_DUMP_CURSOR_NONE; +} +EXPORT_SYMBOL_GPL(inet_diag_dump_clear_cursor); + static int inet_diag_dump_done(struct netlink_callback *cb) { struct inet_diag_dump_data *cb_data =3D cb->data; =20 + inet_diag_dump_clear_cursor(cb_data); bpf_sk_storage_diag_free(cb_data->bpf_stg_diag); kfree(cb->data); =20 diff --git a/net/ipv4/tcp_diag.c b/net/ipv4/tcp_diag.c index ba1fdbe9807f..be0c22cc445b 100644 --- a/net/ipv4/tcp_diag.c +++ b/net/ipv4/tcp_diag.c @@ -285,6 +285,65 @@ static int sk_diag_fill(struct sock *sk, struct sk_buf= f *skb, net_admin); } =20 +/* Process a maximum of SKARR_SZ sockets at a time when walking hash bucke= ts + * with bh disabled. + */ +#define SKARR_SZ 16 + +static void tcp_diag_save_cursor(struct inet_diag_dump_data *cb_data, int = type, + unsigned int slot, struct sock *sk) +{ + sock_hold(sk); + inet_diag_dump_clear_cursor(cb_data); + cb_data->dump_cursor =3D sk; + cb_data->dump_cursor_slot =3D slot; + cb_data->dump_cursor_type =3D type; +} + +static bool tcp_diag_bind_collect_sock(struct sock *sk, struct sock **sk_a= rr, + int *num_arr, int *accum, int num) +{ + sock_hold(sk); + num_arr[*accum] =3D num; + sk_arr[*accum] =3D sk; + + return ++*accum =3D=3D SKARR_SZ; +} + +static bool tcp_diag_bind_collect_owners(struct hlist_head *owners, + struct sock **sk_arr, int *num_arr, + int *accum, int *num, int s_num) +{ + struct sock *sk; + + sk_for_each_bound(sk, owners) { + if (*num < s_num) { + (*num)++; + continue; + } + + if (tcp_diag_bind_collect_sock(sk, sk_arr, num_arr, accum, *num)) + return true; + (*num)++; + } + + return false; +} + +static bool tcp_diag_bind_collect_owners_continue(struct sock *sk, + struct sock **sk_arr, + int *num_arr, int *accum, + int *num) +{ + hlist_for_each_entry_continue(sk, sk_bind_node) { + if (tcp_diag_bind_collect_sock(sk, sk_arr, num_arr, accum, *num)) + return true; + (*num)++; + } + + return false; +} + static void twsk_build_assert(void) { BUILD_BUG_ON(offsetof(struct inet_timewait_sock, tw_family) !=3D @@ -335,8 +394,15 @@ static void tcp_diag_dump(struct sk_buff *skb, struct = netlink_callback *cb, for (i =3D s_i; i <=3D hashinfo->lhash2_mask; i++) { struct inet_listen_hashbucket *ilb; struct hlist_nulls_node *node; + struct sock *sk_arr[SKARR_SZ]; + struct sock *cursor; + int num_arr[SKARR_SZ]; + int idx, accum, res; + bool use_cursor; =20 +resume_listen_walk: num =3D 0; + accum =3D 0; ilb =3D &hashinfo->lhash2[i]; =20 if (hlist_nulls_empty(&ilb->nulls_head)) { @@ -344,52 +410,80 @@ static void tcp_diag_dump(struct sk_buff *skb, struct= netlink_callback *cb, continue; } spin_lock(&ilb->lock); - sk_nulls_for_each(sk, node, &ilb->nulls_head) { - struct inet_sock *inet =3D inet_sk(sk); + cursor =3D cb_data->dump_cursor; + use_cursor =3D cursor && + cb_data->dump_cursor_type =3D=3D + INET_DIAG_DUMP_CURSOR_TCP_LISTEN && + cb_data->dump_cursor_slot =3D=3D i && + !hlist_nulls_unhashed(&cursor->sk_nulls_node) && + cursor->sk_nulls_node.pprev !=3D LIST_POISON2; + node =3D use_cursor ? cursor->sk_nulls_node.next : + ilb->nulls_head.first; + hlist_nulls_for_each_entry_from(sk, node, sk_nulls_node) { + if (!use_cursor && num < s_num) + goto next_listen; =20 - if (!net_eq(sock_net(sk), net)) - continue; + sock_hold(sk); + num_arr[accum] =3D num; + sk_arr[accum] =3D sk; + if (++accum =3D=3D SKARR_SZ) + break; =20 - if (num < s_num) { - num++; - continue; - } +next_listen: + ++num; + } + spin_unlock(&ilb->lock); =20 + res =3D 0; + for (idx =3D 0; idx < accum; idx++) { + struct inet_sock *inet; + + sk =3D sk_arr[idx]; + if (!net_eq(sock_net(sk), net)) + goto processed_listen_sk; + + inet =3D inet_sk(sk); if (r->sdiag_family !=3D AF_UNSPEC && sk->sk_family !=3D r->sdiag_family) - goto next_listen; + goto processed_listen_sk; =20 if (r->id.idiag_sport !=3D inet->inet_sport && r->id.idiag_sport) - goto next_listen; - - if (!inet_diag_bc_sk(cb_data, sk)) - goto next_listen; + goto processed_listen_sk; =20 - if (inet_sk_diag_fill(sk, inet_csk(sk), skb, - cb, r, NLM_F_MULTI, - net_admin) < 0) { - spin_unlock(&ilb->lock); - goto done; + if (res >=3D 0 && inet_diag_bc_sk(cb_data, sk)) { + res =3D inet_sk_diag_fill(sk, inet_csk(sk), + skb, cb, r, NLM_F_MULTI, + net_admin); + if (res < 0) + num =3D num_arr[idx]; } +processed_listen_sk: + if (res >=3D 0) + tcp_diag_save_cursor(cb_data, + INET_DIAG_DUMP_CURSOR_TCP_LISTEN, + i, sk); + sock_put(sk); + } + if (res < 0) + goto done; =20 -next_listen: - ++num; + cond_resched(); + + if (accum =3D=3D SKARR_SZ) { + s_num =3D 0; + goto resume_listen_walk; } - spin_unlock(&ilb->lock); =20 + inet_diag_dump_clear_cursor(cb_data); s_num =3D 0; } skip_listen_ht: + inet_diag_dump_clear_cursor(cb_data); cb->args[0] =3D 1; s_i =3D num =3D s_num =3D 0; } =20 -/* Process a maximum of SKARR_SZ sockets at a time when walking hash bucke= ts - * with bh disabled. - */ -#define SKARR_SZ 16 - /* Dump bound but inactive (not listening, connecting, etc.) sockets */ if (cb->args[0] =3D=3D 1) { if (!(idiag_states & TCPF_BOUND_INACTIVE)) @@ -399,8 +493,10 @@ static void tcp_diag_dump(struct sk_buff *skb, struct = netlink_callback *cb, struct inet_bind_hashbucket *ibb; struct inet_bind2_bucket *tb2; struct sock *sk_arr[SKARR_SZ]; + struct sock *cursor; int num_arr[SKARR_SZ]; int idx, accum, res; + bool use_cursor; =20 resume_bind_walk: num =3D 0; @@ -412,34 +508,38 @@ static void tcp_diag_dump(struct sk_buff *skb, struct= netlink_callback *cb, continue; } spin_lock_bh(&ibb->lock); - inet_bind_bucket_for_each(tb2, &ibb->chain) { - if (!net_eq(ib2_net(tb2), net)) - continue; - - sk_for_each_bound(sk, &tb2->owners) { - struct inet_sock *inet =3D inet_sk(sk); - - if (num < s_num) - goto next_bind; - - if (sk->sk_state !=3D TCP_CLOSE || - !inet->inet_num) - goto next_bind; - - if (r->sdiag_family !=3D AF_UNSPEC && - r->sdiag_family !=3D sk->sk_family) - goto next_bind; - - if (!inet_diag_bc_sk(cb_data, sk)) - goto next_bind; - - sock_hold(sk); - num_arr[accum] =3D num; - sk_arr[accum] =3D sk; - if (++accum =3D=3D SKARR_SZ) + cursor =3D cb_data->dump_cursor; + use_cursor =3D cursor && + cb_data->dump_cursor_type =3D=3D + INET_DIAG_DUMP_CURSOR_TCP_BIND && + cb_data->dump_cursor_slot =3D=3D i && + !hlist_unhashed(&cursor->sk_bind_node) && + cursor->sk_bind_node.pprev !=3D LIST_POISON2 && + inet_csk(cursor)->icsk_bind2_hash; + if (use_cursor) { + tb2 =3D inet_csk(cursor)->icsk_bind2_hash; + sk =3D cursor; + if (tcp_diag_bind_collect_owners_continue(sk, sk_arr, + num_arr, + &accum, + &num)) + goto pause_bind_walk; + hlist_for_each_entry_continue(tb2, node) { + if (tcp_diag_bind_collect_owners(&tb2->owners, + sk_arr, + num_arr, + &accum, + &num, 0)) + goto pause_bind_walk; + } + } else { + inet_bind_bucket_for_each(tb2, &ibb->chain) { + if (tcp_diag_bind_collect_owners(&tb2->owners, + sk_arr, + num_arr, + &accum, + &num, s_num)) goto pause_bind_walk; -next_bind: - num++; } } pause_bind_walk: @@ -447,15 +547,33 @@ static void tcp_diag_dump(struct sk_buff *skb, struct= netlink_callback *cb, =20 res =3D 0; for (idx =3D 0; idx < accum; idx++) { - if (res >=3D 0) { - res =3D inet_sk_diag_fill(sk_arr[idx], - NULL, skb, cb, + struct inet_sock *inet; + + sk =3D sk_arr[idx]; + if (!net_eq(sock_net(sk), net)) + goto put_bind_sk; + + inet =3D inet_sk(sk); + if (sk->sk_state !=3D TCP_CLOSE || !inet->inet_num) + goto put_bind_sk; + + if (r->sdiag_family !=3D AF_UNSPEC && + r->sdiag_family !=3D sk->sk_family) + goto put_bind_sk; + + if (res >=3D 0 && inet_diag_bc_sk(cb_data, sk)) { + res =3D inet_sk_diag_fill(sk, NULL, skb, cb, r, NLM_F_MULTI, net_admin); if (res < 0) num =3D num_arr[idx]; } - sock_put(sk_arr[idx]); + if (res >=3D 0) + tcp_diag_save_cursor(cb_data, + INET_DIAG_DUMP_CURSOR_TCP_BIND, + i, sk); +put_bind_sk: + sock_put(sk); } if (res < 0) goto done; @@ -463,13 +581,15 @@ static void tcp_diag_dump(struct sk_buff *skb, struct= netlink_callback *cb, cond_resched(); =20 if (accum =3D=3D SKARR_SZ) { - s_num =3D num + 1; + s_num =3D 0; goto resume_bind_walk; } =20 + inet_diag_dump_clear_cursor(cb_data); s_num =3D 0; } skip_bind_ht: + inet_diag_dump_clear_cursor(cb_data); cb->args[0] =3D 2; s_i =3D num =3D s_num =3D 0; } @@ -482,42 +602,33 @@ static void tcp_diag_dump(struct sk_buff *skb, struct= netlink_callback *cb, spinlock_t *lock =3D inet_ehash_lockp(hashinfo, i); struct hlist_nulls_node *node; struct sock *sk_arr[SKARR_SZ]; + struct sock *cursor; int num_arr[SKARR_SZ]; int idx, accum, res; + bool use_cursor; =20 if (hlist_nulls_empty(&head->chain)) continue; =20 - if (i > s_i) + if (i > s_i) { + inet_diag_dump_clear_cursor(cb_data); s_num =3D 0; + } =20 next_chunk: num =3D 0; accum =3D 0; spin_lock_bh(lock); - sk_nulls_for_each(sk, node, &head->chain) { - int state; - - if (!net_eq(sock_net(sk), net)) - continue; - if (num < s_num) - goto next_normal; - state =3D (sk->sk_state =3D=3D TCP_TIME_WAIT) ? - READ_ONCE(inet_twsk(sk)->tw_substate) : sk->sk_state; - if (!(idiag_states & (1 << state))) - goto next_normal; - if (r->sdiag_family !=3D AF_UNSPEC && - sk->sk_family !=3D r->sdiag_family) - goto next_normal; - if (r->id.idiag_sport !=3D htons(READ_ONCE(sk->sk_num)) && - r->id.idiag_sport) - goto next_normal; - if (r->id.idiag_dport !=3D sk->sk_dport && - r->id.idiag_dport) - goto next_normal; - twsk_build_assert(); - - if (!inet_diag_bc_sk(cb_data, sk)) + cursor =3D cb_data->dump_cursor; + use_cursor =3D cursor && + cb_data->dump_cursor_type =3D=3D + INET_DIAG_DUMP_CURSOR_TCP_EHASH && + cb_data->dump_cursor_slot =3D=3D i && + !hlist_nulls_unhashed(&cursor->sk_nulls_node) && + cursor->sk_nulls_node.pprev !=3D LIST_POISON2; + node =3D use_cursor ? cursor->sk_nulls_node.next : head->chain.first; + hlist_nulls_for_each_entry_from(sk, node, sk_nulls_node) { + if (!use_cursor && num < s_num) goto next_normal; =20 if (!refcount_inc_not_zero(&sk->sk_refcnt)) @@ -534,13 +645,42 @@ static void tcp_diag_dump(struct sk_buff *skb, struct= netlink_callback *cb, =20 res =3D 0; for (idx =3D 0; idx < accum; idx++) { - if (res >=3D 0) { - res =3D sk_diag_fill(sk_arr[idx], skb, cb, r, - NLM_F_MULTI, net_admin); + int state; + + sk =3D sk_arr[idx]; + if (!net_eq(sock_net(sk), net)) + goto put_estab_sk; + + state =3D (sk->sk_state =3D=3D TCP_TIME_WAIT) ? + READ_ONCE(inet_twsk(sk)->tw_substate) : sk->sk_state; + if (!(idiag_states & (1 << state))) + goto put_estab_sk; + + if (r->sdiag_family !=3D AF_UNSPEC && + sk->sk_family !=3D r->sdiag_family) + goto put_estab_sk; + + if (r->id.idiag_sport !=3D htons(READ_ONCE(sk->sk_num)) && + r->id.idiag_sport) + goto put_estab_sk; + + if (r->id.idiag_dport !=3D sk->sk_dport && + r->id.idiag_dport) + goto put_estab_sk; + + twsk_build_assert(); + if (res >=3D 0 && inet_diag_bc_sk(cb_data, sk)) { + res =3D sk_diag_fill(sk, skb, cb, r, NLM_F_MULTI, + net_admin); if (res < 0) num =3D num_arr[idx]; } - sock_gen_put(sk_arr[idx]); + if (res >=3D 0) + tcp_diag_save_cursor(cb_data, + INET_DIAG_DUMP_CURSOR_TCP_EHASH, + i, sk); +put_estab_sk: + sock_gen_put(sk); } if (res < 0) break; @@ -548,9 +688,11 @@ static void tcp_diag_dump(struct sk_buff *skb, struct = netlink_callback *cb, cond_resched(); =20 if (accum =3D=3D SKARR_SZ) { - s_num =3D num + 1; + s_num =3D 0; goto next_chunk; } + + inet_diag_dump_clear_cursor(cb_data); } =20 done: --=20 2.43.0 From nobody Mon Aug 24 20:41:02 2026 Received: from mail-pl1-f178.google.com (mail-pl1-f178.google.com [209.85.214.178]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A46534749D8 for ; Wed, 29 Jul 2026 11:29:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.178 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785324584; cv=none; b=jA+Yg23i878ayzwfudQI/1uQrUIDOs4xAvhCLprom4+sloOrZYeaeZneiyD5PDCG+qTkqizX1pOwhlf1RqwFeH4WiAtJThyvOEPYSssLn+j9eHBVRwWpBH9rA6CMRCfOWS6s5JAFOedgioTvwAFZlu7mMHmDDzXsLGSQwmhOCEA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785324584; c=relaxed/simple; bh=9MhVi3tBuUVUiWdwEgz/IihUj3etazN9M4/4mSNlqdQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=TP4N06DtM/az8Kya9BaygX4gEfXD7pSzcnLTyu3BHXDfHosDp4MJjIH/YOMcay+OVYIdhJERkcnfl+f/1wWwkgFSyqZKqpHuSUwttu3ZogRzKlz0bH81HzAvxuYj8h5Auli68CoQft5TG6DcaxbnRNTIZZuBTl1Q0rdrypW+PMs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=nebusec.ai; spf=pass smtp.mailfrom=nebusec.ai; dkim=pass (2048-bit key) header.d=nebusec.ai header.i=@nebusec.ai header.b=e60V2GPj; arc=none smtp.client-ip=209.85.214.178 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=nebusec.ai Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=nebusec.ai Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=nebusec.ai header.i=@nebusec.ai header.b="e60V2GPj" Received: by mail-pl1-f178.google.com with SMTP id d9443c01a7336-2ce87c7e3bbso9459965ad.1 for ; Wed, 29 Jul 2026 04:29:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=nebusec.ai; s=google; t=1785324582; x=1785929382; darn=lists.linux.dev; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=VYe2iubu6ASdkv/skq7uS2fIda+uHnh2Q+ijL0hRfXc=; b=e60V2GPjkGrkKEEyUI8Wphlm34yO8b4+/UAizL9GVBgQlkE8FyLvx8EceAvqxXhyVo +PlnB+Zcww3hStzAPtatYkcvNEk0Sv2kNkGgM5R0uOxEnvV2g1j+gqKseFciGFLw0L6o zZ9KxLwqZuxwn9OBbncpcYi4So2zADS5OZ+ljQLVAwSk+MwGJ+zvm5EpYrhvG7ZRDIR/ uf8hPA+f6kL+iyDynVvdNvtot1uPtw5i6YRgz+Jm5K40qjO6C/fmQe83zpnXlXSmVcC5 5v718YjW5xk5KHLm+nP/qZNk4G0l50b7XPN/4nC7RNA00a2KQXNflMrDb9yWrN5MRqRo muGg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785324582; x=1785929382; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=VYe2iubu6ASdkv/skq7uS2fIda+uHnh2Q+ijL0hRfXc=; b=IWM5Ym53qnawtATrsAm8l0Cm/bA/Qo9kOQlR/P+Fw90fbWaR53tj7PIxJ+W7CSzFgB WGBFXCKRAT8Pas4Al1ZDSG4o2s4DqGq868ooswoWrevdlzByqGH6qsMdSvt6bZ6c+Tpw PLHUobIAtPJme5cEQQ1qHpUA4NLlY2udSG4q6xUHQxdYBY3mA9KaINyhrw+G3oXc12yY b6EJldmWDZaLMY4m1yVA3EXgwwaJZTOKUwUsFLsf55tSci9crfyiuh7uGo3LOiWHVhCD fQjXl3Sp0qjJJgzfJEyViJ1J60sTOls0PgM0+Cp9UQJfR+FTEzFsDGguwsHewjRkckfA Ulkw== X-Forwarded-Encrypted: i=1; AHgh+Ro4y4byXz+ntMdAd97xkQ7gF2gGon4nCLesG8gtRwInRFtnLNAQEA/RNP9RINBsCmzT1pHSPA==@lists.linux.dev X-Gm-Message-State: AOJu0Yxpqjslvzso4+TFWUFxBDaFapktCgX8mNEgl6Bki63keMu8jcwP OyDXN/wZT35KFdCbbkKncvG7dk8gGI3yT9IXeWDVd+F5EEYPdfml0GivAD+ZnvdV5Ksz X-Gm-Gg: AR+sD13PQRGx2k4Z+F1T++3PqG8eAJv1kO8a/DWygx683RTd1rVSCzFbNtrvyxzE68x DQi1ahaquCsp7hwtqxGfJOUObmn6X+p+jAY+cpz3wR+ATmaSAFpt7xlFHCC9CtQuIk6LrEL0feK T8HwmBRmj6ijwwC7zKpFiinL9kZNQNriTL7gRnzgYqOU60X0Cs65QhqxI0WSppMFWpaWT2cjaUH eL43MzsaEzdzHM51vWi3lAN5p6mNcTwuV3C/Tijws8WWCUPT0QNTEXJTX6RCEl9WKGW0OTnPhe0 a30CkghH9Zgot1xB3DZ5cbNKuyskb4Vh7suWhwSd+4EktEHo8WD1JlnsCSkHjZEeju3KhvZfRNn jhnwnhNyfePcZ2sk6G1M15hgFvALx/l+8EBgGZq3anAFJPdasGv1CBZHc5ghbqu65vPzlO4cAYD xM9JKNCheA5p3BkEBlctP5JmMr4yfqNp4I4YoauNFYpNA5/SHP2Lf7F+/EXT0iSkUh0wMYBOxf5 +u6XMx2xA== X-Received: by 2002:a17:903:1a2d:b0:2ca:d7a3:3723 with SMTP id d9443c01a7336-2d015d238bdmr70760425ad.29.1785324581971; Wed, 29 Jul 2026 04:29:41 -0700 (PDT) Received: from localhost.localdomain ([115.192.250.185]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2d022c55a7dsm9898505ad.84.2026.07.29.04.29.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 29 Jul 2026 04:29:41 -0700 (PDT) From: Zihan Xi To: netdev@vger.kernel.org, mptcp@lists.linux.dev Cc: davem@davemloft.net, edumazet@google.com, pabeni@redhat.com, horms@kernel.org, ncardwell@google.com, kuniyu@google.com, dsahern@kernel.org, kees@kernel.org, matttbe@kernel.org, martineau@kernel.org, geliang@kernel.org, fw@strlen.de, vega@nebusec.ai, zihanx@nebusec.ai Subject: [PATCH net 2/2] mptcp: diag: fix unbounded listener bucket lock hold Date: Wed, 29 Jul 2026 11:28:40 +0000 Message-ID: <4632ca2134dbd3947529180ba8b8ba4b467c2098.1785307984.git.zihanx@nebusec.ai> X-Mailer: git-send-email 2.47.3 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" MPTCP listener diag dumping reuses sk_diag_dump(), which executes inet_diag_bc_sk() before filling the netlink reply. The listener walk in mptcp_diag_dump_listeners() currently performs MPTCP listener selection and bytecode-triggering dump work while holding the listener bucket lock. A dump request can therefore force unbounded per-bucket lock hold by arranging for many listeners in the same bucket to miss the MPTCP or socket-attribute prefilters, so the old 16-entry batching limit no longer bounds the locked scan itself. Fix this by collecting only referenced listener sockets under the bucket lock. Re-check the MPTCP listener properties, obtain the parent MPTCP socket reference, and run sk_diag_dump() only after dropping the lock. Keep a referenced dump cursor so the next batch resumes after the previous listener instead of rescanning the bucket head under the listener lock. Fixes: 4fa39b701ce9 ("mptcp: listen diag dump support") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: Codex:gpt-5.4 Signed-off-by: Zihan Xi --- net/mptcp/mptcp_diag.c | 124 +++++++++++++++++++++++++++++------------ 1 file changed, 87 insertions(+), 37 deletions(-) diff --git a/net/mptcp/mptcp_diag.c b/net/mptcp/mptcp_diag.c index 136c2d05c0ee..981e64343039 100644 --- a/net/mptcp/mptcp_diag.c +++ b/net/mptcp/mptcp_diag.c @@ -6,12 +6,25 @@ * Author: Paolo Abeni */ =20 +/* Process a bounded number of listeners per bucket lock hold. */ +#define MPTCP_DIAG_BULK_SZ 16 + #include #include #include #include #include "protocol.h" =20 +static void mptcp_diag_save_cursor(struct inet_diag_dump_data *cb_data, + unsigned int slot, struct sock *sk) +{ + sock_hold(sk); + inet_diag_dump_clear_cursor(cb_data); + cb_data->dump_cursor =3D sk; + cb_data->dump_cursor_slot =3D slot; + cb_data->dump_cursor_type =3D INET_DIAG_DUMP_CURSOR_MPTCP_LISTEN; +} + static int sk_diag_dump(struct sock *sk, struct sk_buff *skb, struct netlink_callback *cb, const struct inet_diag_req_v2 *req, @@ -77,6 +90,7 @@ static void mptcp_diag_dump_listeners(struct sk_buff *skb= , struct netlink_callba bool net_admin) { struct mptcp_diag_ctx *diag_ctx =3D (void *)cb->ctx; + struct inet_diag_dump_data *cb_data =3D cb->data; struct net *net =3D sock_net(skb->sk); struct inet_hashinfo *hinfo; int i; @@ -86,62 +100,98 @@ static void mptcp_diag_dump_listeners(struct sk_buff *= skb, struct netlink_callba for (i =3D diag_ctx->l_slot; i <=3D hinfo->lhash2_mask; i++) { struct inet_listen_hashbucket *ilb; struct hlist_nulls_node *node; - struct sock *sk; - int num =3D 0; - + struct sock *tmp, *sk, *sk_arr[MPTCP_DIAG_BULK_SZ]; + struct sock *cursor; + int accum, idx, num, ret; + int num_arr[MPTCP_DIAG_BULK_SZ]; + bool use_cursor; + +resume_listen_walk: + num =3D 0; + accum =3D 0; ilb =3D &hinfo->lhash2[i]; + ret =3D 0; =20 rcu_read_lock(); spin_lock(&ilb->lock); - sk_nulls_for_each(sk, node, &ilb->nulls_head) { - const struct mptcp_subflow_context *ctx =3D mptcp_subflow_ctx(sk); - struct inet_sock *inet =3D inet_sk(sk); - int ret; - - if (num < diag_ctx->l_num) - goto next_listen; - - if (!ctx || strcmp(inet_csk(sk)->icsk_ulp_ops->name, "mptcp")) - goto next_listen; - - sk =3D ctx->conn; - if (!sk || !net_eq(sock_net(sk), net)) - goto next_listen; - - if (r->sdiag_family !=3D AF_UNSPEC && - sk->sk_family !=3D r->sdiag_family) - goto next_listen; - - if (r->id.idiag_sport !=3D inet->inet_sport && - r->id.idiag_sport) + cursor =3D cb_data->dump_cursor; + use_cursor =3D cursor && + cb_data->dump_cursor_type =3D=3D + INET_DIAG_DUMP_CURSOR_MPTCP_LISTEN && + cb_data->dump_cursor_slot =3D=3D i && + !hlist_nulls_unhashed(&cursor->sk_nulls_node) && + cursor->sk_nulls_node.pprev !=3D LIST_POISON2; + node =3D use_cursor ? cursor->sk_nulls_node.next : + ilb->nulls_head.first; + hlist_nulls_for_each_entry_from(sk, node, sk_nulls_node) { + if (!use_cursor && num < diag_ctx->l_num) goto next_listen; =20 if (!refcount_inc_not_zero(&sk->sk_refcnt)) goto next_listen; =20 - ret =3D sk_diag_dump(sk, skb, cb, r, net_admin); - - sock_put(sk); - - if (ret < 0) { - spin_unlock(&ilb->lock); - rcu_read_unlock(); - diag_ctx->l_slot =3D i; - diag_ctx->l_num =3D num; - return; - } - diag_ctx->l_num =3D num + 1; - num =3D 0; + num_arr[accum] =3D num; + sk_arr[accum] =3D sk; + if (++accum =3D=3D MPTCP_DIAG_BULK_SZ) + break; next_listen: ++num; } spin_unlock(&ilb->lock); rcu_read_unlock(); =20 + for (idx =3D 0; idx < accum; idx++) { + const struct tcp_ulp_ops *ulp_ops; + const struct mptcp_subflow_context *ctx; + struct inet_sock *inet; + + sk =3D sk_arr[idx]; + rcu_read_lock(); + ctx =3D mptcp_subflow_ctx(sk); + ulp_ops =3D READ_ONCE(inet_csk(sk)->icsk_ulp_ops); + inet =3D inet_sk(sk); + tmp =3D ctx ? ctx->conn : NULL; + if (!ctx || !ulp_ops || strcmp(ulp_ops->name, "mptcp") || + !tmp || !net_eq(sock_net(tmp), net) || + (r->sdiag_family !=3D AF_UNSPEC && + tmp->sk_family !=3D r->sdiag_family) || + (r->id.idiag_sport !=3D inet->inet_sport && + r->id.idiag_sport) || + !refcount_inc_not_zero(&tmp->sk_refcnt)) { + rcu_read_unlock(); + goto processed_listener_sk; + } + rcu_read_unlock(); + if (ret >=3D 0) { + ret =3D sk_diag_dump(tmp, skb, cb, r, net_admin); + if (ret < 0) + num =3D num_arr[idx]; + } + sock_put(tmp); +processed_listener_sk: + if (ret >=3D 0) + mptcp_diag_save_cursor(cb_data, i, sk); + sock_put(sk); + } + + if (ret < 0) { + diag_ctx->l_slot =3D i; + diag_ctx->l_num =3D num; + return; + } + cond_resched(); + + if (accum =3D=3D MPTCP_DIAG_BULK_SZ) { + diag_ctx->l_num =3D 0; + goto resume_listen_walk; + } + + inet_diag_dump_clear_cursor(cb_data); diag_ctx->l_num =3D 0; } =20 + inet_diag_dump_clear_cursor(cb_data); diag_ctx->l_num =3D 0; diag_ctx->l_slot =3D i; } --=20 2.43.0