From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5CD8837D12F for ; Tue, 1 Sep 2026 06:49:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245352; cv=none; b=S+y4UQDegKwpOSQznI547WmwMvTg272uWYQENbZYfZjz5m3+Vo8YQB3z6+p8tSzF9nyXAOSoLWDHW9y4c/UfaWJewPUq/XtuhyG/vsaXC+af6iAYZyUK1tAYjEJTY0S3eRGJz5xE5QFR0PI8msRJ7XDxIEqEyJU/8Rsk6SHiw9k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245352; c=relaxed/simple; bh=FrN1m+1+0xj6o07xVjvLzD/Ej9nGCJipIMfLIQk+ssA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lqhWY0MXV0N1hK65XlePtaUDkWBr6+xRc25gWmKg0yZxcGYS4rTGq0XlQ5eHrQfl7yl0GWUoAg36U7Lw93BZeLrCfQ1oJkbuICFCtNWDbzIVWn9Qju9VcjXsvaqu2yCwDETSQLEWYbtduKU4cG7RALk7TxVKpB8FgYy5ADPW6mw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=YRSKX59K; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="YRSKX59K" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1D4081F00A3E; Tue, 1 Sep 2026 06:49:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245351; bh=gvgAjBD224MkL7Aq/FJjeRX9j0R7Hx4PmrkMAF1KjSI=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=YRSKX59KTFOlRmD8lAw0P3tHhbVi+CPM3gYNJ18Df9kG1JNnABqbio6k5gzDsAzcS wsI1RBwcUVYSH/YfIBPLePcCcKiSBp9SmXm1iqXeojJmNtJdrnRfcIcnqNL8Cx2Hhy YNR3alRVOeBkxGmP6m5drBw/doaaFvRLZ//b/18H5UWKzXPAMfs4bM10Ii0BvgFIpb nttkZj0Ny82fkbK4QHghmjzQtOWHph3Y4BFiAEiPc/TRibGDPuFdoz+k3V9xt/pE5k 1OfaqJ+kX2+ssERQW6ktSPH/5ZMfFcGGiOf1T0iWlc9SWAbJO6NWcPZIy4ui4ObRN+ tcQvONBnr2thg== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Paolo Abeni , Geliang Tang Subject: [PATCH mptcp-next v13 01/10] mptcp: drop the mptcp_ooo_try_coalesce() helper Date: Tue, 1 Sep 2026 14:48:54 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Paolo Abeni It's used to save an additional comparison for in-order skbs, but is also a barrier to remove CB offset. Remove the helper, let __mptcp_try_coalesce() always perform the sequence check and remove duplicate checks from the callers. Co-developed-by: Geliang Tang Signed-off-by: Geliang Tang Signed-off-by: Paolo Abeni --- net/mptcp/protocol.c | 21 ++++++--------------- 1 file changed, 6 insertions(+), 15 deletions(-) diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index 96933a221eff..2e1df59817bd 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -167,7 +167,8 @@ static bool __mptcp_try_coalesce(struct sock *sk, struc= t sk_buff *to, { int limit =3D READ_ONCE(sk->sk_rcvbuf); =20 - if (unlikely(MPTCP_SKB_CB(to)->cant_coalesce) || + if (MPTCP_SKB_CB(from)->map_seq !=3D MPTCP_SKB_CB(to)->end_seq || + unlikely(MPTCP_SKB_CB(to)->cant_coalesce) || MPTCP_SKB_CB(from)->offset || ((to->len + from->len) > (limit >> 3)) || !skb_try_coalesce(to, from, fragstolen, delta)) @@ -200,15 +201,6 @@ static bool mptcp_try_coalesce(struct sock *sk, struct= sk_buff *to, return true; } =20 -static bool mptcp_ooo_try_coalesce(struct mptcp_sock *msk, struct sk_buff = *to, - struct sk_buff *from) -{ - if (MPTCP_SKB_CB(from)->map_seq !=3D MPTCP_SKB_CB(to)->end_seq) - return false; - - return mptcp_try_coalesce((struct sock *)msk, to, from); -} - /* "inspired" by tcp_rcvbuf_grow(), main difference: * - mptcp does not maintain a msk-level window clamp * - returns true when the receive buffer is actually updated @@ -348,7 +340,7 @@ static void mptcp_data_queue_ofo(struct mptcp_sock *msk= , struct sk_buff *skb) /* with 2 subflows, adding at end of ooo queue is quite likely * Use of ooo_last_skb avoids the O(Log(N)) rbtree lookup. */ - if (mptcp_ooo_try_coalesce(msk, msk->ooo_last_skb, skb)) { + if (mptcp_try_coalesce(sk, msk->ooo_last_skb, skb)) { MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_OFOMERGE); MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_OFOQUEUETAIL); return; @@ -394,7 +386,7 @@ static void mptcp_data_queue_ofo(struct mptcp_sock *msk= , struct sk_buff *skb) MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_DUPDATA); goto merge_right; } - } else if (mptcp_ooo_try_coalesce(msk, skb1, skb)) { + } else if (mptcp_try_coalesce(sk, skb1, skb)) { MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_OFOMERGE); return; } @@ -777,8 +769,7 @@ static void __mptcp_add_backlog(struct sock *sk, if (!list_empty(&msk->backlog_list)) tail =3D list_last_entry(&msk->backlog_list, struct sk_buff, list); =20 - if (tail && MPTCP_SKB_CB(skb)->map_seq =3D=3D MPTCP_SKB_CB(tail)->end_seq= && - ssk =3D=3D tail->sk && + if (tail && ssk =3D=3D tail->sk && __mptcp_try_coalesce(sk, tail, skb, &fragstolen, &delta)) { skb->truesize -=3D delta; kfree_skb_partial(skb, fragstolen); @@ -902,7 +893,7 @@ static bool __mptcp_ofo_queue(struct mptcp_sock *msk) =20 end_seq =3D MPTCP_SKB_CB(skb)->end_seq; tail =3D skb_peek_tail(&sk->sk_receive_queue); - if (!tail || !mptcp_ooo_try_coalesce(msk, tail, skb)) { + if (!tail || !mptcp_try_coalesce(sk, tail, skb)) { int delta =3D msk->ack_seq - MPTCP_SKB_CB(skb)->map_seq; =20 /* skip overlapping data, if any */ --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 133EF349CFD for ; Tue, 1 Sep 2026 06:49:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245354; cv=none; b=RJSi/+C1WrCdg29TUJHloZc/jjvDzOgT0+/cTKOHHpHcGy9OUnvbboGaFsmxh8rv31i0dktskNY/O9OTNAlkyD5P5tgpUFqiYn38iAkVFspKvQl0gciBxOcj34iVOjEKL77g6Q14nlBwIYN4QAKIUB4S+PHhOkhuQO3ygH5Y5TM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245354; c=relaxed/simple; bh=9OMSbXFhe//d6mbW8Sox0zazfzhsdPgEz/qtUVVZOd0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Mfbv3S4S/Q0WzF5HGK5LvJXpzFQxx+uS+XjVL9ohY2xT6LP8nONZXGQcIB0O7zXOaHea1Gyqx6UFpLecxFxkd5kgxyUXTO2PDA73QSa1nBgGwMS0I9YWjfr0WEbJEn/eEMJeGkTbfiTrxazK7x+pueHcmpBWnHZYcKEqFOtUXlA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=W3vPa9Bj; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="W3vPa9Bj" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 9AEA81F000E9; Tue, 1 Sep 2026 06:49:11 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245352; bh=T9LOnRHBXasHcD4dBNWnofo6M9l3vLJdULrvEDF68JM=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=W3vPa9Bj/sMBrkBTZhp8HjI7IaEPfte7hCYIZuiLxJG8FGr8UPiGKiMvhLmO28W9Z a0i1R0l13wtSVbQWAbN1z6T3gOZf8RlzBFpYFCX3gUiiSRYwHVy9fyMspx8F0zwO5a IVc2QAt1r2abRB4At5+QPpUXjc2e51T2xHt5ZxI9+Su9+V3T6l1XkCoQbSIOXvKT/2 ZP7Kl2VdVJYGp/JR9nRGTN0AP8VRPPgldAbRMa/N2hq94LqIIP1tTIuLZmMJqCI8It aDZOTvjSylIzDD0LIg0T/EWz0boXbiIVBI5Fm3TG9NT6fCxKWq+CH4eEf40uKdP8tD Lr0cK+FVfoo2g== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Paolo Abeni , Geliang Tang Subject: [PATCH mptcp-next v13 02/10] mptcp: drop the cant_coalesce CB field Date: Tue, 1 Sep 2026 14:48:55 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Paolo Abeni Such field is used to ensure in-sequence processing in case of fastopen. Instead let's perform synchronization of the fastopen skb sequence when the IASN becomes available with the 3rd ack. When the `cant_coalesce` field has been introduced, commit f03afb3aeb9d ("mptcp: drop __mptcp_fastopen_gen_msk_ackseq()") noted that updating the already queued skb for passive fastopen socket at 3rd ack time would be difficult and race prone. The main point is that such update don't need to be synchronously performed at 3rd ack time, but is sufficient to perform it before the next segment is introduced into the msk. To such extent, add an explicit test in __mptcp_move_skb(). Performance wise this trades a conditional in the fast path - in __mptcp_try_coalesce() - with a similar one in __mptcp_move_skb() and a couple more in slow paths. After this change the user-space will always observe consistent sequence numbers in the receive queue, even in the TFO dummy mapping case. There is still a potential race in mptcp_inq_hint() that will be addressed by a later patch in the series. Co-developed-by: Geliang Tang Signed-off-by: Geliang Tang Signed-off-by: Paolo Abeni --- net/mptcp/fastopen.c | 6 ++++-- net/mptcp/protocol.c | 40 ++++++++++++++++++++++++++++++++++++++-- net/mptcp/protocol.h | 4 +++- net/mptcp/subflow.c | 10 ++++++++++ 4 files changed, 55 insertions(+), 5 deletions(-) diff --git a/net/mptcp/fastopen.c b/net/mptcp/fastopen.c index f717750906ff..e18b31028e7e 100644 --- a/net/mptcp/fastopen.c +++ b/net/mptcp/fastopen.c @@ -9,6 +9,7 @@ void mptcp_fastopen_subflow_synack_set_params(struct mptcp_subflow_context= *subflow, struct request_sock *req) { + struct mptcp_sock *msk; struct sock *sk, *ssk; struct sk_buff *skb; struct tcp_sock *tp; @@ -49,15 +50,16 @@ void mptcp_fastopen_subflow_synack_set_params(struct mp= tcp_subflow_context *subf MPTCP_SKB_CB(skb)->end_seq =3D 0; MPTCP_SKB_CB(skb)->offset =3D 0; MPTCP_SKB_CB(skb)->has_rxtstamp =3D has_rxtstamp; - MPTCP_SKB_CB(skb)->cant_coalesce =3D 1; =20 mptcp_data_lock(sk); DEBUG_NET_WARN_ON_ONCE(sock_owned_by_user_nocheck(sk)); =20 + msk =3D mptcp_sk(sk); + msk->rcvd_dummy_seq =3D true; mptcp_borrow_fwdmem(sk, skb); skb_set_owner_r(skb, sk); __skb_queue_tail(&sk->sk_receive_queue, skb); - mptcp_sk(sk)->bytes_received +=3D skb->len; + msk->bytes_received +=3D skb->len; =20 sk->sk_data_ready(sk); =20 diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index 2e1df59817bd..904bf564b766 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -168,7 +168,6 @@ static bool __mptcp_try_coalesce(struct sock *sk, struc= t sk_buff *to, int limit =3D READ_ONCE(sk->sk_rcvbuf); =20 if (MPTCP_SKB_CB(from)->map_seq !=3D MPTCP_SKB_CB(to)->end_seq || - unlikely(MPTCP_SKB_CB(to)->cant_coalesce) || MPTCP_SKB_CB(from)->offset || ((to->len + from->len) > (limit >> 3)) || !skb_try_coalesce(to, from, fragstolen, delta)) @@ -430,7 +429,6 @@ static void mptcp_init_skb(struct sock *ssk, struct sk_= buff *skb, int offset, MPTCP_SKB_CB(skb)->end_seq =3D MPTCP_SKB_CB(skb)->map_seq + copy_len; MPTCP_SKB_CB(skb)->offset =3D offset; MPTCP_SKB_CB(skb)->has_rxtstamp =3D has_rxtstamp; - MPTCP_SKB_CB(skb)->cant_coalesce =3D 0; =20 __skb_unlink(skb, &ssk->sk_receive_queue); =20 @@ -438,6 +436,26 @@ static void mptcp_init_skb(struct sock *ssk, struct sk= _buff *skb, int offset, skb_dst_drop(skb); } =20 +void __mptcp_sync_rcv_sequence(struct sock *sk) +{ + struct mptcp_sock *msk =3D mptcp_sk(sk); + struct sk_buff *skb; + u32 offset; + + if (likely(!msk->rcvd_dummy_seq)) + return; + + /* User space can have already received the TFO skb. */ + msk->rcvd_dummy_seq =3D false; + skb =3D skb_peek_tail(&sk->sk_receive_queue); + if (!skb) + return; + + offset =3D MPTCP_SKB_CB(skb)->offset; + MPTCP_SKB_CB(skb)->map_seq =3D msk->ack_seq - skb->len + offset; + MPTCP_SKB_CB(skb)->end_seq =3D msk->ack_seq; +} + static bool __mptcp_move_skb(struct sock *sk, struct sk_buff *skb) { u64 copy_len =3D MPTCP_SKB_CB(skb)->end_seq - MPTCP_SKB_CB(skb)->map_seq; @@ -446,6 +464,17 @@ static bool __mptcp_move_skb(struct sock *sk, struct s= k_buff *skb) =20 mptcp_borrow_fwdmem(sk, skb); =20 + if (test_and_clear_bit(MPTCP_SYNC_SEQ, &msk->cb_flags)) { + /* Ensure we see the updated ack_seq after seeing the flag */ + smp_rmb(); + + /* Be sure to sync the eventual fastopen dummy mapping before + * any other skb lands into the msk. + */ + if (unlikely(msk->rcvd_dummy_seq)) + __mptcp_sync_rcv_sequence(sk); + } + if (MPTCP_SKB_CB(skb)->map_seq =3D=3D msk->ack_seq) { /* in sequence */ insert: @@ -3907,6 +3936,13 @@ static void mptcp_release_cb(struct sock *sk) __mptcp_error_report(sk); if (__test_and_clear_bit(MPTCP_SYNC_SNDBUF, &msk->cb_flags)) __mptcp_sync_sndbuf(sk); + if (test_and_clear_bit(MPTCP_SYNC_SEQ, &msk->cb_flags)) { + /* Ensure we see the updated ack_seq after seeing + * the flag + */ + smp_rmb(); + __mptcp_sync_rcv_sequence(sk); + } } } =20 diff --git a/net/mptcp/protocol.h b/net/mptcp/protocol.h index b3121c8c766b..d534f6cdabb1 100644 --- a/net/mptcp/protocol.h +++ b/net/mptcp/protocol.h @@ -126,13 +126,13 @@ #define MPTCP_FLUSH_JOIN_LIST 5 #define MPTCP_SYNC_STATE 6 #define MPTCP_SYNC_SNDBUF 7 +#define MPTCP_SYNC_SEQ 8 =20 struct mptcp_skb_cb { u64 map_seq; u64 end_seq; u32 offset; u8 has_rxtstamp; - u8 cant_coalesce; }; =20 #define MPTCP_SKB_CB(__skb) ((struct mptcp_skb_cb *)&((__skb)->cb[0])) @@ -313,6 +313,7 @@ struct mptcp_sock { u32 token; unsigned long flags; unsigned long cb_flags; + bool rcvd_dummy_seq; bool recovery; /* closing subflow write queue reinjected */ bool can_ack; bool fully_established; @@ -1172,6 +1173,7 @@ void mptcp_event_pm_listener(const struct sock *ssk, enum mptcp_event_type event); bool mptcp_userspace_pm_active(const struct mptcp_sock *msk); =20 +void __mptcp_sync_rcv_sequence(struct sock *sk); void mptcp_fastopen_subflow_synack_set_params(struct mptcp_subflow_context= *subflow, struct request_sock *req); int mptcp_pm_genl_fill_addr(struct sk_buff *msg, diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c index 2d7ccb01d234..ed3a48cf9c97 100644 --- a/net/mptcp/subflow.c +++ b/net/mptcp/subflow.c @@ -476,6 +476,8 @@ static void subflow_set_remote_key(struct mptcp_sock *m= sk, struct mptcp_subflow_context *subflow, const struct mptcp_options_received *mp_opt) { + struct sock *sk =3D (struct sock *)msk; + /* active MPC subflow will reach here multiple times: * at subflow_finish_connect() time and at 4th ack time */ @@ -494,6 +496,14 @@ static void subflow_set_remote_key(struct mptcp_sock *= msk, WRITE_ONCE(msk->ack_seq, subflow->iasn); WRITE_ONCE(msk->can_ack, true); atomic64_set(&msk->rcv_wnd_sent, subflow->iasn); + + if (!sock_owned_by_user(sk)) { + __mptcp_sync_rcv_sequence(sk); + } else { + /* Ensure ack_seq is visible before setting the flag */ + smp_wmb(); + set_bit(MPTCP_SYNC_SEQ, &msk->cb_flags); + } } =20 static void mptcp_propagate_state(struct sock *sk, struct sock *ssk, --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 07669367B9F for ; Tue, 1 Sep 2026 06:49:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245356; cv=none; b=Nm71oFnZ6IWwbWArOqrR6TrferJxzQCuppmY1fHkMsGB57ssiL7xfu+sDLRKXr434rV3Nhkxg41xkB9VyntKyOAO3ag02HFA3nEe4h/8g+dWj+P5Fk8nxVVTXuC1zS9jVnidfYCkz27egVFr0z3LyIHuqoeVfZg/I2vcHLy6mkw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245356; c=relaxed/simple; bh=uuKnczjTnJBlYjV8roqbfmHH960afg8HryEfs0/wrUI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=JjsZpIa29eRRBQVOq7Tr/kwV+Aa5hTI+czHeYE/6lBieaSj5sz2WOiPUwphNAZmyOTjltl75ZZMDrXXWromiwRWVcQK0WSAVOZ3Tz7FP8GahlJFgkc3XCeDwH6XwX8SRa3oBfDNuu/2Z3soJAQVHJFblTIMAr1+CO2+xGw8iIK4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=gZazCM+2; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="gZazCM+2" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 8E7011F00A3D; Tue, 1 Sep 2026 06:49:13 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245354; bh=LO2PlhofNFla1uGnsr6r2yAcz1fAh1a8AHmxd9MKKms=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=gZazCM+2zGy129KxnqSFNp/mVmtHqgLc4dPpM5vNB3VKMMez2IFyU/5zaRqjAFJoo IhuvCBpaDqI743uLiFvcNJu1TvAAsRgjrxTg39YUtR1LB4n0JvB1rlB5r2C5AtAY3B TVSixwJQZAdjb5NGNXNOwEVY9W0E/AoGQcawJYIXOKaPFQDlReurApqZb8cnWtdmZn 6Fv2sx8pl9Px+QucaG9l1nkUAv/UN55gaa8LNVhxKDHYejTYOIY8wO7jocfBjbXCqh c4Ini10cTmxR3oJybqL4Y8tjmORqxh1nkNGsonB+tudWOZakqDB1Nal5M7i0Bh0RxV 6nb7DhA0fYXIg== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Paolo Abeni , Geliang Tang Subject: [PATCH mptcp-next v13 03/10] mptcp: remove CB offset field Date: Tue, 1 Sep 2026 14:48:56 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Paolo Abeni Instead, use a new msk-level field to track the bytes already consumed inside each skb, carrying the amount of bytes already copied to user-space, alike what TCP is already doing. The newly introduce `copied_seq` field is always accessed under the msk socket lock, delegating the synchronization with IASN to the msk release CB, when the socket is owned by the user-space at remote key reception time. Such synchronization preserves any partial progress (copy) made on the TFO packet. Note that the explicit synchronization in __mptcp_move_skb() is needed to ensure that the TFO skb in the receive queue got its map_seq synched before the next skb lands into the receive queue when spooling the backlog at mptcp_release_cb() time, as the release CB synchronization will happen later. Prior to this patch, the TFO skb dummy mapping was always ignored, now it affects the `copied_seq` initial update: be sure to extends the sign correctly of such mapping initialization time. Overall this simplify a bit the __mptcp_recvmsg_mskq(), mptcp_inq_hint() and the __mptcp_move_skb() code and will also make possible the next patch. Initialize MPTCP sequence space to 0 in mptcp_propagate_state() when mp_opt is NULL, ensuring SKB's map_seq starts from 0 to match msk->copied_seq and prevent offset calculation underflow in fallback mode. Co-developed-by: Geliang Tang Signed-off-by: Geliang Tang Signed-off-by: Paolo Abeni --- net/mptcp/fastopen.c | 9 ++-- net/mptcp/protocol.c | 108 +++++++++++++++++++++---------------------- net/mptcp/protocol.h | 8 +++- net/mptcp/subflow.c | 9 ++++ 4 files changed, 74 insertions(+), 60 deletions(-) diff --git a/net/mptcp/fastopen.c b/net/mptcp/fastopen.c index e18b31028e7e..421a50a85547 100644 --- a/net/mptcp/fastopen.c +++ b/net/mptcp/fastopen.c @@ -45,10 +45,11 @@ void mptcp_fastopen_subflow_synack_set_params(struct mp= tcp_subflow_context *subf subflow->ssn_offset +=3D skb->len; has_rxtstamp =3D TCP_SKB_CB(skb)->has_rxtstamp; =20 - /* Only the sequence delta is relevant */ - MPTCP_SKB_CB(skb)->map_seq =3D -skb->len; + /* The TFO segment data sits before the IASN; before receiving + * the remote key, IASN is assumed being 0. + */ + MPTCP_SKB_CB(skb)->map_seq =3D -(u64)skb->len; MPTCP_SKB_CB(skb)->end_seq =3D 0; - MPTCP_SKB_CB(skb)->offset =3D 0; MPTCP_SKB_CB(skb)->has_rxtstamp =3D has_rxtstamp; =20 mptcp_data_lock(sk); @@ -56,6 +57,8 @@ void mptcp_fastopen_subflow_synack_set_params(struct mptc= p_subflow_context *subf =20 msk =3D mptcp_sk(sk); msk->rcvd_dummy_seq =3D true; + msk->copied_seq =3D MPTCP_SKB_CB(skb)->map_seq; + msk->tfo_skb_len =3D skb->len; mptcp_borrow_fwdmem(sk, skb); skb_set_owner_r(skb, sk); __skb_queue_tail(&sk->sk_receive_queue, skb); diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index 904bf564b766..67484be33b55 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -29,7 +29,7 @@ #include "protocol.h" #include "mib.h" =20 -static unsigned int mptcp_inq_hint(const struct sock *sk); +static unsigned int mptcp_inq_hint(struct sock *sk); =20 #define CREATE_TRACE_POINTS #include @@ -168,7 +168,6 @@ static bool __mptcp_try_coalesce(struct sock *sk, struc= t sk_buff *to, int limit =3D READ_ONCE(sk->sk_rcvbuf); =20 if (MPTCP_SKB_CB(from)->map_seq !=3D MPTCP_SKB_CB(to)->end_seq || - MPTCP_SKB_CB(from)->offset || ((to->len + from->len) > (limit >> 3)) || !skb_try_coalesce(to, from, fragstolen, delta)) return false; @@ -415,8 +414,7 @@ static void mptcp_data_queue_ofo(struct mptcp_sock *msk= , struct sk_buff *skb) skb_set_owner_r(skb, sk); } =20 -static void mptcp_init_skb(struct sock *ssk, struct sk_buff *skb, int offs= et, - int copy_len) +static void mptcp_init_skb(struct sock *ssk, struct sk_buff *skb, int offs= et) { struct mptcp_subflow_context *subflow =3D mptcp_subflow_ctx(ssk); bool has_rxtstamp =3D TCP_SKB_CB(skb)->has_rxtstamp; @@ -425,9 +423,9 @@ static void mptcp_init_skb(struct sock *ssk, struct sk_= buff *skb, int offset, * mptcp_subflow_get_mapped_dsn() is based on the current tp->copied_seq * value */ - MPTCP_SKB_CB(skb)->map_seq =3D mptcp_subflow_get_mapped_dsn(subflow); - MPTCP_SKB_CB(skb)->end_seq =3D MPTCP_SKB_CB(skb)->map_seq + copy_len; - MPTCP_SKB_CB(skb)->offset =3D offset; + MPTCP_SKB_CB(skb)->map_seq =3D mptcp_subflow_get_mapped_dsn(subflow) - + offset; + MPTCP_SKB_CB(skb)->end_seq =3D MPTCP_SKB_CB(skb)->map_seq + skb->len; MPTCP_SKB_CB(skb)->has_rxtstamp =3D has_rxtstamp; =20 __skb_unlink(skb, &ssk->sk_receive_queue); @@ -440,7 +438,6 @@ void __mptcp_sync_rcv_sequence(struct sock *sk) { struct mptcp_sock *msk =3D mptcp_sk(sk); struct sk_buff *skb; - u32 offset; =20 if (likely(!msk->rcvd_dummy_seq)) return; @@ -451,9 +448,8 @@ void __mptcp_sync_rcv_sequence(struct sock *sk) if (!skb) return; =20 - offset =3D MPTCP_SKB_CB(skb)->offset; - MPTCP_SKB_CB(skb)->map_seq =3D msk->ack_seq - skb->len + offset; - MPTCP_SKB_CB(skb)->end_seq =3D msk->ack_seq; + MPTCP_SKB_CB(skb)->map_seq =3D mptcp_iasn(msk) - skb->len; + MPTCP_SKB_CB(skb)->end_seq =3D MPTCP_SKB_CB(skb)->map_seq + skb->len; } =20 static bool __mptcp_move_skb(struct sock *sk, struct sk_buff *skb) @@ -467,6 +463,7 @@ static bool __mptcp_move_skb(struct sock *sk, struct sk= _buff *skb) if (test_and_clear_bit(MPTCP_SYNC_SEQ, &msk->cb_flags)) { /* Ensure we see the updated ack_seq after seeing the flag */ smp_rmb(); + msk->copied_seq +=3D mptcp_iasn(msk); =20 /* Be sure to sync the eventual fastopen dummy mapping before * any other skb lands into the msk. @@ -501,10 +498,6 @@ static bool __mptcp_move_skb(struct sock *sk, struct s= k_buff *skb) /* Partial packet */ if (after64(MPTCP_SKB_CB(skb)->end_seq, msk->ack_seq)) { copy_len =3D MPTCP_SKB_CB(skb)->end_seq - msk->ack_seq; - MPTCP_SKB_CB(skb)->offset +=3D msk->ack_seq - - MPTCP_SKB_CB(skb)->map_seq; - MPTCP_SKB_CB(skb)->map_seq +=3D msk->ack_seq - - MPTCP_SKB_CB(skb)->map_seq; goto insert; } =20 @@ -862,7 +855,7 @@ static bool __mptcp_move_skbs_from_subflow(struct mptcp= _sock *msk, if (offset < skb->len) { size_t len =3D skb->len - offset; =20 - mptcp_init_skb(ssk, skb, offset, len); + mptcp_init_skb(ssk, skb, offset); =20 if (own_msk) { mptcp_subflow_lend_fwdmem(subflow, skb); @@ -929,8 +922,6 @@ static bool __mptcp_ofo_queue(struct mptcp_sock *msk) pr_debug("uncoalesced seq=3D%llx ack seq=3D%llx delta=3D%d\n", MPTCP_SKB_CB(skb)->map_seq, msk->ack_seq, delta); - MPTCP_SKB_CB(skb)->offset +=3D delta; - MPTCP_SKB_CB(skb)->map_seq +=3D delta; __skb_queue_tail(&sk->sk_receive_queue, skb); } msk->bytes_received +=3D end_seq - msk->ack_seq; @@ -2211,33 +2202,24 @@ static void mptcp_eat_recv_skb(struct sock *sk, str= uct sk_buff *skb) } =20 static int __mptcp_recvmsg_mskq(struct sock *sk, struct msghdr *msg, - size_t len, int flags, int copied_total, + size_t len, int flags, u64 *seq, struct scm_timestamping_internal *tss, int *cmsg_flags, struct sk_buff **last) { struct mptcp_sock *msk =3D mptcp_sk(sk); struct sk_buff *skb, *tmp; - int total_data_len =3D 0; int copied =3D 0; =20 skb_queue_walk_safe(&sk->sk_receive_queue, skb, tmp) { - u32 delta, offset =3D MPTCP_SKB_CB(skb)->offset; + u64 offset =3D *seq - MPTCP_SKB_CB(skb)->map_seq; u32 data_len =3D skb->len - offset; u32 count; int err; =20 - if (flags & MSG_PEEK) { - /* skip already peeked skbs */ - if (total_data_len + data_len <=3D copied_total) { - total_data_len +=3D data_len; - *last =3D skb; - continue; - } - - /* skip the already peeked data in the current skb */ - delta =3D copied_total - total_data_len; - offset +=3D delta; - data_len -=3D delta; + /* Skip the already peeked data. */ + if (offset >=3D skb->len) { + *last =3D skb; + continue; } =20 count =3D min_t(size_t, len - copied, data_len); @@ -2256,14 +2238,12 @@ static int __mptcp_recvmsg_mskq(struct sock *sk, st= ruct msghdr *msg, } =20 copied +=3D count; + *seq +=3D count; =20 if (!(flags & MSG_PEEK)) { msk->bytes_consumed +=3D count; - if (count < data_len) { - MPTCP_SKB_CB(skb)->offset +=3D count; - MPTCP_SKB_CB(skb)->map_seq +=3D count; + if (count < data_len) break; - } =20 mptcp_eat_recv_skb(sk, skb); } else { @@ -2416,25 +2396,27 @@ static bool mptcp_move_skbs(struct sock *sk) return enqueued; } =20 -static unsigned int mptcp_inq_hint(const struct sock *sk) +static unsigned int mptcp_inq_hint(struct sock *sk) { const struct mptcp_sock *msk =3D mptcp_sk(sk); - const struct sk_buff *skb; - - skb =3D skb_peek(&sk->sk_receive_queue); - if (skb) { - u64 hint_val =3D READ_ONCE(msk->ack_seq) - MPTCP_SKB_CB(skb)->map_seq; - - if (hint_val >=3D INT_MAX) - return INT_MAX; + u64 hint_val; =20 - return (unsigned int)hint_val; + /* Avoid races vs ack_seq updates and MPTCP_SYNC_SEQ flag. */ + mptcp_data_lock(sk); + if (test_bit(MPTCP_SYNC_SEQ, &msk->cb_flags)) { + mptcp_data_unlock(sk); + return 0; } + hint_val =3D msk->ack_seq - msk->copied_seq; + mptcp_data_unlock(sk); + if (hint_val >=3D INT_MAX) + return INT_MAX; =20 - if (sk->sk_state =3D=3D TCP_CLOSE || (sk->sk_shutdown & RCV_SHUTDOWN)) + if (!hint_val && + (sk->sk_state =3D=3D TCP_CLOSE || (sk->sk_shutdown & RCV_SHUTDOWN))) return 1; =20 - return 0; + return (unsigned int)hint_val; } =20 static int mptcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len, @@ -2443,6 +2425,7 @@ static int mptcp_recvmsg(struct sock *sk, struct msgh= dr *msg, size_t len, struct mptcp_sock *msk =3D mptcp_sk(sk); struct scm_timestamping_internal tss; int copied =3D 0, cmsg_flags =3D 0; + u64 peek_seq, *seq; int target; long timeo; =20 @@ -2461,6 +2444,11 @@ static int mptcp_recvmsg(struct sock *sk, struct msg= hdr *msg, size_t len, =20 len =3D min_t(size_t, len, INT_MAX); target =3D sock_rcvlowat(sk, flags & MSG_WAITALL, len); + seq =3D &msk->copied_seq; + if (flags & MSG_PEEK) { + peek_seq =3D msk->copied_seq; + seq =3D &peek_seq; + } =20 if (unlikely(msk->recvmsg_inq)) cmsg_flags =3D MPTCP_CMSG_INQ; @@ -2470,7 +2458,7 @@ static int mptcp_recvmsg(struct sock *sk, struct msgh= dr *msg, size_t len, int err, bytes_read; =20 bytes_read =3D __mptcp_recvmsg_mskq(sk, msg, len - copied, flags, - copied, &tss, &cmsg_flags, + seq, &tss, &cmsg_flags, &last); if (unlikely(bytes_read < 0)) { if (!copied) @@ -2480,8 +2468,11 @@ static int mptcp_recvmsg(struct sock *sk, struct msg= hdr *msg, size_t len, =20 copied +=3D bytes_read; =20 - if (!list_empty(&msk->backlog_list) && mptcp_move_skbs(sk)) + if (!list_empty(&msk->backlog_list) && mptcp_move_skbs(sk)) { + if (flags & MSG_PEEK) + peek_seq =3D msk->copied_seq + copied; continue; + } =20 /* only the MPTCP socket status is relevant here. The exit * conditions mirror closely tcp_recvmsg() @@ -2525,6 +2516,10 @@ static int mptcp_recvmsg(struct sock *sk, struct msg= hdr *msg, size_t len, err =3D copied ? : err; goto out_err; } + + /* Recompute peek offset after eventual seq resync. */ + if (flags & MSG_PEEK) + peek_seq =3D msk->copied_seq + copied; } =20 mptcp_cleanup_rbuf(msk, copied); @@ -3707,11 +3702,13 @@ static int mptcp_disconnect(struct sock *sk, int fl= ags) msk->bytes_retrans =3D 0; msk->rcvspace_init =3D 0; msk->fastclosing =3D 0; + msk->tfo_skb_len =3D 0; mptcp_init_rtt_est(msk); =20 /* for fallback's sake */ WRITE_ONCE(msk->ack_seq, 0); atomic64_set(&msk->rcv_wnd_sent, 0); + msk->copied_seq =3D 0; =20 WRITE_ONCE(sk->sk_shutdown, 0); sk_error_report(sk); @@ -3941,6 +3938,7 @@ static void mptcp_release_cb(struct sock *sk) * the flag */ smp_rmb(); + msk->copied_seq +=3D mptcp_iasn(msk); __mptcp_sync_rcv_sequence(sk); } } @@ -4607,7 +4605,7 @@ static struct sk_buff *mptcp_recv_skb(struct sock *sk= , u32 *off) mptcp_move_skbs(sk); =20 while ((skb =3D skb_peek(&sk->sk_receive_queue)) !=3D NULL) { - offset =3D MPTCP_SKB_CB(skb)->offset; + offset =3D msk->copied_seq - MPTCP_SKB_CB(skb)->map_seq; if (offset < skb->len) { *off =3D offset; return skb; @@ -4649,11 +4647,9 @@ static int __mptcp_read_sock(struct sock *sk, read_d= escriptor_t *desc, copied +=3D count; =20 msk->bytes_consumed +=3D count; - if (count < data_len) { - MPTCP_SKB_CB(skb)->offset +=3D count; - MPTCP_SKB_CB(skb)->map_seq +=3D count; + msk->copied_seq +=3D count; + if (count < data_len) break; - } =20 mptcp_eat_recv_skb(sk, skb); if (!desc->count) diff --git a/net/mptcp/protocol.h b/net/mptcp/protocol.h index d534f6cdabb1..7186e0a55091 100644 --- a/net/mptcp/protocol.h +++ b/net/mptcp/protocol.h @@ -131,7 +131,6 @@ struct mptcp_skb_cb { u64 map_seq; u64 end_seq; - u32 offset; u8 has_rxtstamp; }; =20 @@ -292,6 +291,7 @@ struct mptcp_sock { u64 bytes_sent; u64 snd_nxt; u64 bytes_received; + u64 copied_seq; u64 ack_seq; atomic64_t rcv_wnd_sent; u64 rcv_data_fin_seq; @@ -311,6 +311,7 @@ struct mptcp_sock { u32 last_ack_recv; unsigned long timer_ival; u32 token; + u32 tfo_skb_len; unsigned long flags; unsigned long cb_flags; bool rcvd_dummy_seq; @@ -865,6 +866,11 @@ struct sock *mptcp_subflow_get_retrans(struct mptcp_so= ck *msk); int mptcp_sched_get_send(struct mptcp_sock *msk); int mptcp_sched_get_retrans(struct mptcp_sock *msk); =20 +static inline u64 mptcp_iasn(const struct mptcp_sock *msk) +{ + return msk->ack_seq - msk->bytes_received + msk->tfo_skb_len; +} + static inline u64 mptcp_data_avail(const struct mptcp_sock *msk) { return READ_ONCE(msk->bytes_received) - READ_ONCE(msk->bytes_consumed); diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c index ed3a48cf9c97..6b7a45710758 100644 --- a/net/mptcp/subflow.c +++ b/net/mptcp/subflow.c @@ -498,6 +498,8 @@ static void subflow_set_remote_key(struct mptcp_sock *m= sk, atomic64_set(&msk->rcv_wnd_sent, subflow->iasn); =20 if (!sock_owned_by_user(sk)) { + /* User space could have already read partially the TFO skb */ + msk->copied_seq +=3D subflow->iasn; __mptcp_sync_rcv_sequence(sk); } else { /* Ensure ack_seq is visible before setting the flag */ @@ -520,6 +522,13 @@ static void mptcp_propagate_state(struct sock *sk, str= uct sock *ssk, WRITE_ONCE(msk->snd_una, subflow->idsn + 1); WRITE_ONCE(msk->wnd_end, subflow->idsn + 1 + tcp_sk(ssk)->snd_wnd); subflow_set_remote_key(msk, subflow, mp_opt); + } else { + /* Fallback: initialize sequence space to 0 (no remote key) */ + subflow->map_seq =3D 0; + /* ensure mptcp_subflow_get_map_offset() returns 0 */ + subflow->map_subflow_seq =3D tcp_sk(ssk)->copied_seq - + subflow->ssn_offset; + WRITE_ONCE(msk->ack_seq, 0); } =20 if (!sock_owned_by_user(sk)) { --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 95D5E36A364 for ; Tue, 1 Sep 2026 06:49:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245358; cv=none; b=rwX0nV8Pjk+KMn8sgTfYh5kY3RCoLUEdv+KRnMfdLDj1OnXG26vGofXCdPPJ9Zy5F3a97ANkDdALtMRM3Ghs+tb4a/tZJtOqlesk/B8ZzyQbTMrFIPWXHvcB0JBDqh2GkTHGBFtPxSgYg4zxDa/OaETlIy7scOqVt7zs+vKOw4s= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245358; c=relaxed/simple; bh=dkbuwArDx7CJacUoWHVzrWVMbNNM/9FaItT1fm8Uv4g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=tMDu4jg69mVd3OK0hEtlPbAldIk7VtFzBKFbyT3QSK76pl5ZdhuJhSee1t1R6wVTrmsn11slq0dDvBG6VKU+KTCqNNOWKdXyIwDA0Nj/xednXhKIeVA0DxIA4WbaMBUb7LcBjmwUQ+M2Foq4XEGwx0nOtUPNv3ezBEzjWo3apjs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=hP1R3o3e; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="hP1R3o3e" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2B5AC1F000E9; Tue, 1 Sep 2026 06:49:14 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245356; bh=o488NE8OhwB7gs50ldtmy9DT+f7oWhGmMEKeXv4v2Wo=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=hP1R3o3eSO+sOlTW7EpbmtaNs9z36OsGnuA1ByZLTj9LkCzm9Jp0wtbSzERdA9Gh4 XWOLpQwc6nGhsbTEyc2qxONDd1PjgdL43XNgJd1K3op1TWLdkcZAO9yhlcIgN1AAx7 IauVL6vZnPrZeoHRlfQx1nz2o8G7rBZoZM0ZV3Wu9NjWvT8w3LDnxhHCQ6m+sLCDC/ /N3IMXFiMFKrA1CDvJ5Fya2HkJNPmQlaJC55KZ5HRfrqfmyb1E4vk/T/M1iDEp3BLp fpZIsp9LrRVaL53SEsi3gOURpGW8WwKE7PiZMNvo6w5tXPcrPPGtDBqdeoe6Z6Fn2Y qjT2afylWK5hw== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Paolo Abeni , Geliang Tang Subject: [PATCH mptcp-next v13 04/10] mptcp: sync mptcp skb cb layout with tcp one Date: Tue, 1 Sep 2026 14:48:57 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Paolo Abeni The MPTCP protocol uses a significantly different CB layout WRT TCP, as it includes different information and use 64 bits for the sequence numbers. As the msk-level rcvbuf buffer size is limited by the core socket code the INT_MAX; after validating the incoming skb vs the current receive window, we can safely use 32 bits for MPTCP-level sequence number. This allow updating the MPTCP CB layout so that fields with a corresponding TCP-level data use the same area inside the CB itself. Add build time check to ensure the latter invariant. Co-developed-by: Geliang Tang Signed-off-by: Geliang Tang Signed-off-by: Paolo Abeni --- net/mptcp/fastopen.c | 6 ++-- net/mptcp/protocol.c | 85 +++++++++++++++++++++++++++----------------- net/mptcp/protocol.h | 7 ++-- 3 files changed, 61 insertions(+), 37 deletions(-) diff --git a/net/mptcp/fastopen.c b/net/mptcp/fastopen.c index 421a50a85547..0d339dd454f9 100644 --- a/net/mptcp/fastopen.c +++ b/net/mptcp/fastopen.c @@ -48,8 +48,10 @@ void mptcp_fastopen_subflow_synack_set_params(struct mpt= cp_subflow_context *subf /* The TFO segment data sits before the IASN; before receiving * the remote key, IASN is assumed being 0. */ - MPTCP_SKB_CB(skb)->map_seq =3D -(u64)skb->len; + MPTCP_SKB_CB(skb)->map_seq64 =3D -(u64)skb->len; + MPTCP_SKB_CB(skb)->map_seq =3D MPTCP_SKB_CB(skb)->map_seq64; MPTCP_SKB_CB(skb)->end_seq =3D 0; + MPTCP_SKB_CB(skb)->flags =3D 0; MPTCP_SKB_CB(skb)->has_rxtstamp =3D has_rxtstamp; =20 mptcp_data_lock(sk); @@ -57,7 +59,7 @@ void mptcp_fastopen_subflow_synack_set_params(struct mptc= p_subflow_context *subf =20 msk =3D mptcp_sk(sk); msk->rcvd_dummy_seq =3D true; - msk->copied_seq =3D MPTCP_SKB_CB(skb)->map_seq; + msk->copied_seq =3D MPTCP_SKB_CB(skb)->map_seq64; msk->tfo_skb_len =3D skb->len; mptcp_borrow_fwdmem(sk, skb); skb_set_owner_r(skb, sk); diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index 67484be33b55..f0ac1ac750c3 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -172,7 +172,7 @@ static bool __mptcp_try_coalesce(struct sock *sk, struc= t sk_buff *to, !skb_try_coalesce(to, from, fragstolen, delta)) return false; =20 - pr_debug("colesced seq %llx into %llx new len %d new end seq %llx\n", + pr_debug("colesced seq %x into %x new len %d new end seq %x\n", MPTCP_SKB_CB(from)->map_seq, MPTCP_SKB_CB(to)->map_seq, to->len, MPTCP_SKB_CB(from)->end_seq); MPTCP_SKB_CB(to)->end_seq =3D MPTCP_SKB_CB(from)->end_seq; @@ -254,8 +254,8 @@ static void mptcp_prune_ofo_queue(struct sock *sk, struct sk_buff *skb =3D rb_to_skb(node); =20 /* Stop pruning if the incoming skb would land in OoO tail. */ - if (after64(MPTCP_SKB_CB(in_skb)->map_seq, - MPTCP_SKB_CB(skb)->map_seq)) + if (after(MPTCP_SKB_CB(in_skb)->map_seq, + MPTCP_SKB_CB(skb)->map_seq)) break; =20 pruned =3D true; @@ -301,15 +301,19 @@ static void mptcp_data_queue_ofo(struct mptcp_sock *m= sk, struct sk_buff *skb) { struct sock *sk =3D (struct sock *)msk; struct rb_node **p, *parent; - u64 seq, end_seq, max_seq; + u64 end_seq, max_seq; struct sk_buff *skb1; + u32 seq; =20 seq =3D MPTCP_SKB_CB(skb)->map_seq; - end_seq =3D MPTCP_SKB_CB(skb)->end_seq; + end_seq =3D MPTCP_SKB_CB(skb)->map_seq64 + skb->len; max_seq =3D atomic64_read(&msk->rcv_wnd_sent); =20 - pr_debug("msk=3D%p seq=3D%llx limit=3D%llx empty=3D%d\n", msk, seq, max_s= eq, + pr_debug("msk=3D%p seq=3D%x limit=3D%llx empty=3D%d\n", msk, seq, max_seq, RB_EMPTY_ROOT(&msk->out_of_order_queue)); + /* Use the full sequence space to perform the admission checks, to + * protect vs possible wrap-arounds. + */ if (after64(end_seq, max_seq)) { /* out of window */ mptcp_drop(sk, skb); @@ -345,7 +349,7 @@ static void mptcp_data_queue_ofo(struct mptcp_sock *msk= , struct sk_buff *skb) } =20 /* Can avoid an rbtree lookup if we are adding skb after ooo_last_skb */ - if (!before64(seq, MPTCP_SKB_CB(msk->ooo_last_skb)->end_seq)) { + if (!before(seq, MPTCP_SKB_CB(msk->ooo_last_skb)->end_seq)) { MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_OFOQUEUETAIL); parent =3D &msk->ooo_last_skb->rbnode; p =3D &parent->rb_right; @@ -357,18 +361,18 @@ static void mptcp_data_queue_ofo(struct mptcp_sock *m= sk, struct sk_buff *skb) while (*p) { parent =3D *p; skb1 =3D rb_to_skb(parent); - if (before64(seq, MPTCP_SKB_CB(skb1)->map_seq)) { + if (before(seq, MPTCP_SKB_CB(skb1)->map_seq)) { p =3D &parent->rb_left; continue; } - if (before64(seq, MPTCP_SKB_CB(skb1)->end_seq)) { - if (!after64(end_seq, MPTCP_SKB_CB(skb1)->end_seq)) { + if (before(seq, MPTCP_SKB_CB(skb1)->end_seq)) { + if (!after(end_seq, MPTCP_SKB_CB(skb1)->end_seq)) { /* All the bits are present. Drop. */ mptcp_drop(sk, skb); MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_DUPDATA); return; } - if (after64(seq, MPTCP_SKB_CB(skb1)->map_seq)) { + if (after(seq, MPTCP_SKB_CB(skb1)->map_seq)) { /* partial overlap: * | skb | * | skb1 | @@ -399,7 +403,7 @@ static void mptcp_data_queue_ofo(struct mptcp_sock *msk= , struct sk_buff *skb) merge_right: /* Remove other segments covered by skb. */ while ((skb1 =3D skb_rb_next(skb)) !=3D NULL) { - if (before64(end_seq, MPTCP_SKB_CB(skb1)->end_seq)) + if (before((u32)end_seq, MPTCP_SKB_CB(skb1)->end_seq)) break; rb_erase(&skb1->rbnode, &msk->out_of_order_queue); mptcp_drop(sk, skb1); @@ -421,11 +425,13 @@ static void mptcp_init_skb(struct sock *ssk, struct s= k_buff *skb, int offset) =20 /* the skb map_seq accounts for the skb offset: * mptcp_subflow_get_mapped_dsn() is based on the current tp->copied_seq - * value + * value; note that end seq number is only available in 32bits format. */ - MPTCP_SKB_CB(skb)->map_seq =3D mptcp_subflow_get_mapped_dsn(subflow) - - offset; + MPTCP_SKB_CB(skb)->map_seq64 =3D mptcp_subflow_get_mapped_dsn(subflow) - + offset; + MPTCP_SKB_CB(skb)->map_seq =3D (u32)MPTCP_SKB_CB(skb)->map_seq64; MPTCP_SKB_CB(skb)->end_seq =3D MPTCP_SKB_CB(skb)->map_seq + skb->len; + MPTCP_SKB_CB(skb)->flags =3D 0; MPTCP_SKB_CB(skb)->has_rxtstamp =3D has_rxtstamp; =20 __skb_unlink(skb, &ssk->sk_receive_queue); @@ -448,13 +454,14 @@ void __mptcp_sync_rcv_sequence(struct sock *sk) if (!skb) return; =20 - MPTCP_SKB_CB(skb)->map_seq =3D mptcp_iasn(msk) - skb->len; + MPTCP_SKB_CB(skb)->map_seq64 =3D mptcp_iasn(msk) - skb->len; + MPTCP_SKB_CB(skb)->map_seq =3D (u32)MPTCP_SKB_CB(skb)->map_seq64; MPTCP_SKB_CB(skb)->end_seq =3D MPTCP_SKB_CB(skb)->map_seq + skb->len; } =20 static bool __mptcp_move_skb(struct sock *sk, struct sk_buff *skb) { - u64 copy_len =3D MPTCP_SKB_CB(skb)->end_seq - MPTCP_SKB_CB(skb)->map_seq; + u32 copy_len =3D MPTCP_SKB_CB(skb)->end_seq - MPTCP_SKB_CB(skb)->map_seq; struct mptcp_sock *msk =3D mptcp_sk(sk); struct sk_buff *tail; =20 @@ -472,7 +479,7 @@ static bool __mptcp_move_skb(struct sock *sk, struct sk= _buff *skb) __mptcp_sync_rcv_sequence(sk); } =20 - if (MPTCP_SKB_CB(skb)->map_seq =3D=3D msk->ack_seq) { + if (MPTCP_SKB_CB(skb)->map_seq64 =3D=3D msk->ack_seq) { /* in sequence */ insert: if (!mptcp_try_rmem_schedule(sk, skb)) { @@ -490,14 +497,14 @@ static bool __mptcp_move_skb(struct sock *sk, struct = sk_buff *skb) skb_set_owner_r(skb, sk); __skb_queue_tail(&sk->sk_receive_queue, skb); return true; - } else if (after64(MPTCP_SKB_CB(skb)->map_seq, msk->ack_seq)) { + } else if (after64(MPTCP_SKB_CB(skb)->map_seq64, msk->ack_seq)) { mptcp_data_queue_ofo(msk, skb); return false; } =20 /* Partial packet */ - if (after64(MPTCP_SKB_CB(skb)->end_seq, msk->ack_seq)) { - copy_len =3D MPTCP_SKB_CB(skb)->end_seq - msk->ack_seq; + if (after64(MPTCP_SKB_CB(skb)->map_seq64 + skb->len, msk->ack_seq)) { + copy_len =3D MPTCP_SKB_CB(skb)->end_seq - (u32)msk->ack_seq; goto insert; } =20 @@ -892,40 +899,40 @@ static bool __mptcp_ofo_queue(struct mptcp_sock *msk) { struct sock *sk =3D (struct sock *)msk; struct sk_buff *skb, *tail; + u32 seq_delta, ack_seq; bool moved =3D false; struct rb_node *p; - u64 end_seq; =20 p =3D rb_first(&msk->out_of_order_queue); pr_debug("msk=3D%p empty=3D%d\n", msk, RB_EMPTY_ROOT(&msk->out_of_order_q= ueue)); while (p) { + ack_seq =3D msk->ack_seq; skb =3D rb_to_skb(p); - if (after64(MPTCP_SKB_CB(skb)->map_seq, msk->ack_seq)) + if (after(MPTCP_SKB_CB(skb)->map_seq, ack_seq)) break; =20 p =3D rb_next(p); rb_erase(&skb->rbnode, &msk->out_of_order_queue); =20 - if (unlikely(!after64(MPTCP_SKB_CB(skb)->end_seq, - msk->ack_seq))) { + if (unlikely(!after(MPTCP_SKB_CB(skb)->end_seq, ack_seq))) { mptcp_drop(sk, skb); MPTCP_INC_STATS(sock_net(sk), MPTCP_MIB_DUPDATA); continue; } =20 - end_seq =3D MPTCP_SKB_CB(skb)->end_seq; + seq_delta =3D MPTCP_SKB_CB(skb)->end_seq - ack_seq; tail =3D skb_peek_tail(&sk->sk_receive_queue); if (!tail || !mptcp_try_coalesce(sk, tail, skb)) { - int delta =3D msk->ack_seq - MPTCP_SKB_CB(skb)->map_seq; + int delta =3D ack_seq - MPTCP_SKB_CB(skb)->map_seq; =20 /* skip overlapping data, if any */ - pr_debug("uncoalesced seq=3D%llx ack seq=3D%llx delta=3D%d\n", - MPTCP_SKB_CB(skb)->map_seq, msk->ack_seq, + pr_debug("uncoalesced seq=3D%x ack seq=3D%x delta=3D%d\n", + MPTCP_SKB_CB(skb)->map_seq, ack_seq, delta); __skb_queue_tail(&sk->sk_receive_queue, skb); } - msk->bytes_received +=3D end_seq - msk->ack_seq; - WRITE_ONCE(msk->ack_seq, end_seq); + msk->bytes_received +=3D seq_delta; + WRITE_ONCE(msk->ack_seq, msk->ack_seq + seq_delta); moved =3D true; } return moved; @@ -2211,7 +2218,7 @@ static int __mptcp_recvmsg_mskq(struct sock *sk, stru= ct msghdr *msg, int copied =3D 0; =20 skb_queue_walk_safe(&sk->sk_receive_queue, skb, tmp) { - u64 offset =3D *seq - MPTCP_SKB_CB(skb)->map_seq; + u32 offset =3D (u32)(*seq) - MPTCP_SKB_CB(skb)->map_seq; u32 data_len =3D skb->len - offset; u32 count; int err; @@ -4605,7 +4612,7 @@ static struct sk_buff *mptcp_recv_skb(struct sock *sk= , u32 *off) mptcp_move_skbs(sk); =20 while ((skb =3D skb_peek(&sk->sk_receive_queue)) !=3D NULL) { - offset =3D msk->copied_seq - MPTCP_SKB_CB(skb)->map_seq; + offset =3D (u32)msk->copied_seq - MPTCP_SKB_CB(skb)->map_seq; if (offset < skb->len) { *off =3D offset; return skb; @@ -4856,11 +4863,23 @@ static int mptcp_napi_poll(struct napi_struct *napi= , int budget) return work_done; } =20 +#define CHK_CB_FIELD(mptcp_field, tcp_field) \ + ({ \ + BUILD_BUG_ON(offsetof(struct mptcp_skb_cb, mptcp_field) !=3D \ + offsetof(struct tcp_skb_cb, tcp_field)); \ + BUILD_BUG_ON(offsetofend(struct mptcp_skb_cb, mptcp_field) !=3D \ + offsetofend(struct tcp_skb_cb, tcp_field)); \ + }) + void __init mptcp_proto_init(void) { struct mptcp_delegated_action *delegated; int cpu; =20 + CHK_CB_FIELD(map_seq, seq); + CHK_CB_FIELD(end_seq, end_seq); + CHK_CB_FIELD(flags, tcp_flags); + mptcp_prot.h.hashinfo =3D tcp_prot.h.hashinfo; =20 if (percpu_counter_init(&mptcp_sockets_allocated, 0, GFP_KERNEL)) diff --git a/net/mptcp/protocol.h b/net/mptcp/protocol.h index 7186e0a55091..c57b92c2ab87 100644 --- a/net/mptcp/protocol.h +++ b/net/mptcp/protocol.h @@ -129,9 +129,12 @@ #define MPTCP_SYNC_SEQ 8 =20 struct mptcp_skb_cb { - u64 map_seq; - u64 end_seq; + u32 map_seq; + u32 end_seq; + u32 unused; + u16 flags; u8 has_rxtstamp; + u64 map_seq64; }; =20 #define MPTCP_SKB_CB(__skb) ((struct mptcp_skb_cb *)&((__skb)->cb[0])) --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5766D2C0F8C for ; Tue, 1 Sep 2026 06:49:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245359; cv=none; b=I7C6/OeTFbEMnfF85XYKlzfEBNIUuPsVwCHKC/ECm30r1Ax7Te1PJKAdZ53hlHwB04aOEON+dsBIsZuvU9NNV8hHyPXuOg6fLH5ChPayqheaRImkZycWuUYcfP2Lalb9jt0beN/x9w3Kay7Rx7qy61r4GOj5iv0ZjK/ielAE5TM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245359; c=relaxed/simple; bh=ru0gCESAkBiQczYlm28aOujREYLbQ9Znar8OG3T+x2w=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=YuZ7w078YuscNTU3OrYNLmz1impHSkJurWWFNf3j5d4pbf07deQfxV/n5VZVkqq8KFDGkW8dWYYh+tQKpSmLZ9PdXnHiRU9MezHr/leEd4asD+5Pw+OWf76dZDjIbPi9Cvxa1NZXMoYoBIZ6yFPGjFPEm2P8LHi+q2gl3yyl4PA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=nbzN75QM; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="nbzN75QM" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 00E1B1F00A3D; Tue, 1 Sep 2026 06:49:16 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245358; bh=x9QTIRdlxMUSfN4gACfy5q86ysQLtv06uJqMTTdT0kk=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=nbzN75QMBPqbu+FymZARgKgq4Q86khle4q319Oq7eYWr1IpOsexJQS9xDNirriK+G E30qLFcYkUDiNcK6yXrKBKC7NoLIGI9oP7Ks4erIhevlLGshajJjZ1BCFdC7Co572r JiQFjIzuymF+mb20hmeOlxYEftv6kCJ6wu++OpZYwuUi6IQunhrC1GZbd9r8M42rri mD/1vGM6NZpcI9whD36MEEeRAxDlaZirCpnvQaR6IjVhcgbTjzjFlkDxHsAdT4VtTs ZGVbC6JZePZHJSdsK+8M/zB2Mwm5bwABj1mO08qLgIwhV5t+NjDNzV6n+7jqH7aNXK YlYqapJNqhv4g== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Paolo Abeni , Geliang Tang Subject: [PATCH mptcp-next v13 05/10] mptcp: defer read_sock cleanup to mptcp_worker Date: Tue, 1 Sep 2026 14:48:58 +0800 Message-ID: <653b636299ddfa4e146aa0495aea9c2a644f61ae.1788245079.git.tanggeliang@kylinos.cn> X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Paolo Abeni When MPTCP carries TLS, the data path runs under mptcp_data_lock(). Reaching sk->sk_data_ready(sk) synchronously ends up at tls_strp_check_rcv() -> mptcp_recv_skb() -> mptcp_move_skbs(), which calls mptcp_data_lock() on the same sk and recurses on sk_lock.slock. The TLS path is not the only constraint: before the mptcp_recv_skb() calls, the TLS code would also reach __mptcp_read_sock(), which calls mptcp_rcv_space_adjust() and mptcp_cleanup_rbuf(). Both require holding the msk socket lock in process context, while the mptcp/TLS caller is in BH scope. Fix this by deferring sk->sk_data_ready(sk) to mptcp_worker() via a new MPTCP_WORK_READ_COMPLETE bit, reusing the existing mptcp_schedule_work()/ mptcp_cancel_work() infrastructure. The wakeup bit is consumed after the SOCK_DEAD && TCP_CLOSE destroy branch, so a socket that reaches the destroy path drops the pending wakeup rather than running it post-free. Co-developed-by: Geliang Tang Signed-off-by: Geliang Tang Signed-off-by: Paolo Abeni --- net/mptcp/protocol.c | 32 ++++++++++++++++++++++++-------- net/mptcp/protocol.h | 2 ++ 2 files changed, 26 insertions(+), 8 deletions(-) diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index f0ac1ac750c3..794d01afb031 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -3147,6 +3147,20 @@ static void mptcp_backlog_purge(struct sock *sk) sk_mem_reclaim(sk); } =20 +static void mptcp_read_complete(struct sock *sk) +{ + struct mptcp_sock *msk =3D mptcp_sk(sk); + int read_copied; + + mptcp_data_lock(sk); + read_copied =3D msk->read_copied; + msk->read_copied =3D 0; + mptcp_data_unlock(sk); + + mptcp_rcv_space_adjust(msk, read_copied); + mptcp_cleanup_rbuf(msk, read_copied); +} + static void mptcp_do_fastclose(struct sock *sk) { struct mptcp_subflow_context *subflow, *tmp; @@ -3222,6 +3236,9 @@ static void mptcp_worker(struct work_struct *work) if (test_and_clear_bit(MPTCP_WORK_RTX, &msk->flags)) __mptcp_retrans(sk); =20 + if (test_and_clear_bit(MPTCP_WORK_READ_COMPLETE, &msk->flags)) + mptcp_read_complete(sk); + fail_tout =3D msk->first ? READ_ONCE(mptcp_subflow_ctx(msk->first)->fail_= tout) : 0; if (fail_tout && time_after(jiffies, fail_tout)) mptcp_mp_fail_no_response(msk); @@ -4608,9 +4625,6 @@ static struct sk_buff *mptcp_recv_skb(struct sock *sk= , u32 *off) struct sk_buff *skb; u32 offset; =20 - if (!list_empty(&msk->backlog_list)) - mptcp_move_skbs(sk); - while ((skb =3D skb_peek(&sk->sk_receive_queue)) !=3D NULL) { offset =3D (u32)msk->copied_seq - MPTCP_SKB_CB(skb)->map_seq; if (offset < skb->len) { @@ -4625,6 +4639,7 @@ static struct sk_buff *mptcp_recv_skb(struct sock *sk= , u32 *off) /* * Note: * - It is assumed that the socket was locked by the caller. + * - Can be invoked in BH scope. */ static int __mptcp_read_sock(struct sock *sk, read_descriptor_t *desc, sk_read_actor_t recv_actor, bool noack) @@ -4634,8 +4649,6 @@ static int __mptcp_read_sock(struct sock *sk, read_de= scriptor_t *desc, int copied =3D 0; u32 offset; =20 - msk_owned_by_me(msk); - if (sk->sk_state =3D=3D TCP_LISTEN) return -ENOTCONN; while ((skb =3D mptcp_recv_skb(sk, &offset)) !=3D NULL) { @@ -4666,11 +4679,14 @@ static int __mptcp_read_sock(struct sock *sk, read_= descriptor_t *desc, if (noack) goto out; =20 - mptcp_rcv_space_adjust(msk, copied); - + /* The backlog flushing is only needed when some data is actually + * moved and will take place in the workers's release callback. + */ if (copied > 0) { mptcp_recv_skb(sk, &offset); - mptcp_cleanup_rbuf(msk, copied); + msk->read_copied +=3D copied; + set_bit(MPTCP_WORK_READ_COMPLETE, &msk->flags); + mptcp_schedule_work(sk); } out: return copied; diff --git a/net/mptcp/protocol.h b/net/mptcp/protocol.h index c57b92c2ab87..515b0afabbd1 100644 --- a/net/mptcp/protocol.h +++ b/net/mptcp/protocol.h @@ -117,6 +117,7 @@ #define MPTCP_FALLBACK_DONE 2 #define MPTCP_WORK_CLOSE_SUBFLOW 3 #define MPTCP_RTX_ENABLED 4 +#define MPTCP_WORK_READ_COMPLETE 5 =20 /* MPTCP socket release cb flags */ #define MPTCP_PUSH_PENDING 1 @@ -312,6 +313,7 @@ struct mptcp_sock { u32 last_data_sent; u32 last_data_recv; u32 last_ack_recv; + int read_copied; unsigned long timer_ival; u32 token; u32 tfo_skb_len; --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7567F2C0303 for ; Tue, 1 Sep 2026 06:49:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245361; cv=none; b=dxpmPctN/a78Fpsi1J4AgQ3969Cent/ufbbOUlqDMRuZtsz3ICrWhkri08HL7btcJHBaCNaGz9HHhnamotZ7X5d1XOKUhMg5bE+Y+kfAPUZR4EaS/59prO2R8Gg4Y9/WtEpKEZ/Vf7xv6+kUqv5qTxkl89SZcq/j6eAUfiwy0fI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245361; c=relaxed/simple; bh=wSOM4H+mIT+d4Hl8rHlDo8LVKIUVL3FZtwIaoXCkDIY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=m87FhRwsZqzvj6vsot2QHyoeqTkr7wpS9uRAKuTfu8JdV0Fj3zLZ32NzKJ2uE5IxCQKFOcp/M/UDnp6vFwy1Rc4H5/aX64asdxQ2mlS4qbPUFA6Gmhe5fsbnGUNifLn/LHpBoH7mrKhoJLWOW/TpwM0YOBygslTJ1ly8N6FamVY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=TZ7mxkHx; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="TZ7mxkHx" Received: by smtp.kernel.org (Postfix) with ESMTPSA id CA14D1F000E9; Tue, 1 Sep 2026 06:49:18 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245360; bh=lX9HCP3IvGlxXBluWcBrN7LEtNqKQJ4kwMx9MttjhfM=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=TZ7mxkHxQmnwgPXi1IKyb/GVhABV957oiHTYXCY50rRa5tfPHIOoFeT31N8qfWUWp 9iK74QOo8LVRJQjJCT0p+ru2E3/0sJZpjaS/QqN7IjcqOdCoua8Sy1cJHxncINIu4Z KFnFlTzfj6aXPevlb+l/jLXSmuzmtRangZE7sIm3UBQ8T3zFIR9rzeTcaLihX077I+ bNsW8f5QD66DKUl1XphoQ4xsmc/EknkhuBfu2r8eryvs/fOuWY1a9CQ0K5niR8mz6V FPzkFRs9nqbdteslVJYTCxSAtjrqVhHF+kDN2bSO6NYy3Nkx0aLBYsXOiRC+Efns/k DiPrJ5t98Zvog== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Geliang Tang , Zqiang Subject: [PATCH mptcp-next v13 06/10] mptcp: align FIN handling with TCP via SOCK_DONE Date: Tue, 1 Sep 2026 14:48:59 +0800 Message-ID: <837b36993bdfdeb79518c1b5bc1a919a2f37504c.1788245079.git.tanggeliang@kylinos.cn> X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Geliang Tang TCP uses SOCK_DONE to mark FIN reception because RCV_SHUTDOWN alone is ambiguous: it is also set by RST (via SHUTDOWN_MASK) without an ack_seq increment. mptcp_inq_hint() used the latter as a proxy for "FIN pending", causing over-count and potentially returning 0 to peek_len() on a RST-while-data-queued socket, hanging TLS/BPF strparser consumers. Mirror tcp_fin()/tcp_disconnect() in MPTCP: - Set SOCK_DONE in mptcp_check_data_fin() next to RCV_SHUTDOWN, strictly tied to the ack_seq +=3D1 that counts the FIN byte. - Test SOCK_DONE in mptcp_inq_hint() and mptcp_recvmsg() so the FIONREAD/ recvmsg FIN-handling only fires on a real DATA_FIN. - Reset SOCK_DONE in mptcp_disconnect() to mirror tcp_disconnect(). Memory ordering: use smp_mb() in mptcp_check_data_fin() to order the SHUTDOWN, rcv_data_fin, and ack_seq writes before the non-atomic sock_set_flag(SOCK_DONE) and the subsequent atomic operations in mptcp_set_state()/mptcp_close_wake_up(). Pair it with smp_rmb() in mptcp_inq_hint() so a reader never sees SOCK_DONE=3D1 without the matching ack_seq +=3D1. smp_mb() is required (rather than reusing the existing smp_mb__before_atomic()) because sock_set_flag() uses __set_bit() which is non-atomic: on LoongArch smp_mb__before_atomic() is just a compiler fence and a CPU could reorder the SOCK_DONE store past it. Co-developed-by: Zqiang Signed-off-by: Zqiang Signed-off-by: Geliang Tang --- net/mptcp/protocol.c | 14 +++++++++++--- 1 file changed, 11 insertions(+), 3 deletions(-) diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index 794d01afb031..f627270042c5 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -735,7 +735,8 @@ static void mptcp_check_data_fin(struct sock *sk) WRITE_ONCE(msk->rcv_data_fin, 0); =20 WRITE_ONCE(sk->sk_shutdown, sk->sk_shutdown | RCV_SHUTDOWN); - smp_mb__before_atomic(); /* SHUTDOWN must be visible first */ + smp_mb(); /* Order SHUTDOWN, ack_seq before SOCK_DONE. */ + sock_set_flag(sk, SOCK_DONE); =20 switch (sk->sk_state) { case TCP_ESTABLISHED: @@ -2406,8 +2407,12 @@ static bool mptcp_move_skbs(struct sock *sk) static unsigned int mptcp_inq_hint(struct sock *sk) { const struct mptcp_sock *msk =3D mptcp_sk(sk); + bool sock_done =3D sock_flag(sk, SOCK_DONE); u64 hint_val; =20 + /* Pair with smp_mb() in mptcp_check_data_fin(). */ + smp_rmb(); + /* Avoid races vs ack_seq updates and MPTCP_SYNC_SEQ flag. */ mptcp_data_lock(sk); if (test_bit(MPTCP_SYNC_SEQ, &msk->cb_flags)) { @@ -2419,8 +2424,7 @@ static unsigned int mptcp_inq_hint(struct sock *sk) if (hint_val >=3D INT_MAX) return INT_MAX; =20 - if (!hint_val && - (sk->sk_state =3D=3D TCP_CLOSE || (sk->sk_shutdown & RCV_SHUTDOWN))) + if (!hint_val && sock_done) return 1; =20 return (unsigned int)hint_val; @@ -2492,6 +2496,9 @@ static int mptcp_recvmsg(struct sock *sk, struct msgh= dr *msg, size_t len, !timeo) break; } else { + if (sock_flag(sk, SOCK_DONE)) + break; + if (sk->sk_err) { copied =3D sock_error(sk); break; @@ -3735,6 +3742,7 @@ static int mptcp_disconnect(struct sock *sk, int flag= s) msk->copied_seq =3D 0; =20 WRITE_ONCE(sk->sk_shutdown, 0); + sock_reset_flag(sk, SOCK_DONE); sk_error_report(sk); return 0; } --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6DE1D2C0303 for ; Tue, 1 Sep 2026 06:49:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245365; cv=none; b=JxSYJ517Znv/kyp5KYEZglCq1MCW6tZbH/CHvjTIm67SaR2ilnc8JOPxJ8Qj3scQY8b/Fd7azVYJb6CKATruSrS0XCxRCa2AEqSSXHTlTVHw7CZhkD8+FYY8ozNzIOdap72Ou4Gk3QKkhvSOtxh6FzRbaEI+COZs2nbsEN/Red0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245365; c=relaxed/simple; bh=61ly/NtZSVe8jUrEZ5W6PUqEc7SHyDrvkUUD0f7w6wc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Rr7hjzcH/yRozRb6VNL88ETMmBpi78ug5y0flnff6y2RY90ldoQtZAoyfPfddLr1RGH5Sq9cuMS/pxSivjPNNo6agVKvlEd0eIJNnJPypNJRGOOLnmoav1QzV0Q+DupynE2YZUMpGQD2J53PGDUPOmympQSkQ/0vLPAt5NEO0Is= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=NauutG5u; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="NauutG5u" Received: by smtp.kernel.org (Postfix) with ESMTPSA id B81091F00A3D; Tue, 1 Sep 2026 06:49:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245362; bh=cSpx4PAOfJB27zRYipeFrX9v4SDJpQo8KvZLe6G46N4=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=NauutG5ueyHetrevL3n+myMV1iHeJ/GvuqjOB80JjI8nFcstnNyql0FR90hSuVCQu lmf3t8CRpIiiN+HFZV0x4u1hSdyYOPlEdXrWvLZURhVgWNbkt7JlhhP6Ofi32Ge1XU yB8sjNkcgwKXwcmeyy69ZuHb2Ay4l8/3T9SI9QS5J6xMxLDMCznJDBIMBqzmxJqEy8 W77JdRxksPnmnYRYcR0gpqZSH3avtyJRRMBJA5BfSR7GLcjPFWBzagoFBt62WtPEJ9 CCY9CE1i/ei3+TIRXvpDd92of46yOewYZqnZ54fAUfI1uYWI7ew4KwUvrAGB656dV0 2MUlhtBEpg5wg== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Geliang Tang , Zqiang Subject: [PATCH mptcp-next v13 07/10] mptcp: implement peek_len for proto_ops Date: Tue, 1 Sep 2026 14:49:00 +0800 Message-ID: <14c573054c1427a0cb78cd24290b1c13b74ce9ef.1788245079.git.tanggeliang@kylinos.cn> X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Geliang Tang Add mptcp_inq() to compute the number of readable bytes at the MPTCP level. It derives the count from ack_seq - copied_seq, returns 0 while the connection is still handshaking (TCP_SYN_SENT/TCP_SYN_RECV), and subtracts 1 once a FIN has been received, mirroring tcp_inq(). The subtraction is needed because MPTCP's ack_seq is incremented by 1 in mptcp_check_data_fin() when the DATA_FIN is processed, consuming a sequence number without carrying any payload data. Without subtracting 1, peek_len would over-report by one byte, causing upper layers to wait for data that will never arrive. READ_ONCE() is used instead of mptcp_data_lock() for the u64 reads of ack_seq and copied_seq because mptcp_inq() will be called from TLS contexts with the socket lock held - taking the spinlock would deadlock. On 32-bit systems a torn read can produce a wrapped subtraction (close to 2^64), which is caught by the INT_MAX clamp. TLS callers use mptcp_inq() conservatively: over-reporting defers work to the next iteration, where recvmsg() provides the actual byte count. Wire mptcp_inq() into mptcp_peek_len() and assign .peek_len in both mptcp_stream_ops and mptcp_v6_stream_ops, so upper layers get the correct in-queue byte count for MPTCP connections. Co-developed-by: Zqiang Signed-off-by: Zqiang Signed-off-by: Geliang Tang --- net/mptcp/protocol.c | 37 +++++++++++++++++++++++++++++++++++++ 1 file changed, 37 insertions(+) diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index f627270042c5..45c05f896126 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -4819,6 +4819,41 @@ static ssize_t mptcp_splice_read(struct socket *sock= , loff_t *ppos, return ret; } =20 +static int mptcp_inq(struct sock *sk) +{ + const struct mptcp_sock *msk =3D mptcp_sk(sk); + int answ; + + if ((1 << sk->sk_state) & (TCPF_SYN_SENT | TCPF_SYN_RECV)) { + answ =3D 0; + } else if (test_bit(MPTCP_SYNC_SEQ, &msk->cb_flags)) { + answ =3D 0; + } else { + bool sock_done =3D sock_flag(sk, SOCK_DONE); + u64 hint_val; + + /* Pair with smp_mb() in mptcp_check_data_fin(). */ + smp_rmb(); + + hint_val =3D READ_ONCE(msk->ack_seq) - READ_ONCE(msk->copied_seq); + if (hint_val >=3D INT_MAX) + hint_val =3D INT_MAX; + + answ =3D (unsigned int)hint_val; + + /* Subtract 1, if FIN was received. Mirror tcp_inq() */ + if (answ && sock_done) + answ--; + } + + return answ; +} + +static int mptcp_peek_len(struct socket *sock) +{ + return mptcp_inq(sock->sk); +} + static const struct proto_ops mptcp_stream_ops =3D { .family =3D PF_INET, .owner =3D THIS_MODULE, @@ -4841,6 +4876,7 @@ static const struct proto_ops mptcp_stream_ops =3D { .set_rcvlowat =3D mptcp_set_rcvlowat, .read_sock =3D mptcp_read_sock, .splice_read =3D mptcp_splice_read, + .peek_len =3D mptcp_peek_len, }; =20 static struct inet_protosw mptcp_protosw =3D { @@ -4965,6 +5001,7 @@ static const struct proto_ops mptcp_v6_stream_ops =3D= { .set_rcvlowat =3D mptcp_set_rcvlowat, .read_sock =3D mptcp_read_sock, .splice_read =3D mptcp_splice_read, + .peek_len =3D mptcp_peek_len, }; =20 static struct proto mptcp_v6_prot; --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C336D32C8B for ; Tue, 1 Sep 2026 06:49:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245365; cv=none; b=iWtqWEr/065XeExMI2gwGxKpOQYJY9r597KrHmJdUpmRbjLAeHknz3PLvOGdGah86qj2EbolccmpEk6tZYeXzjdqVCWE5axL7ZqI6jQ9xXZSi+Qr90QQ49S6PesABR5S8ZJVPWzPoo0kCmzwpPYoZZl6FBPBfkWkkN/qPVj4jVU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245365; c=relaxed/simple; bh=Pu0H4YoOgtP1vxqDb3nVRZTX7V3SfwPE+YgP27XTeWk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=HG+bB6p9SPgFeGUvaap1S1CBFXvGUM2QgQJjwu0EF4ceuiIesxnnRBkX++jJy9C3XWEqmonzMarQeh+u1UPva9tqd4sEuhmQM+UiZ4gIPsPFnfujZ/CdVtCD54rOPde3+sPkELbQMVcrwuqe14suRO90TsniBzljwU8ZIK6BSWc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=IfGud6Cs; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="IfGud6Cs" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1A5191F00A3E; Tue, 1 Sep 2026 06:49:22 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245364; bh=Qzn1i3/hPBE3pM91RubroaCo6J+LW8UwaACM+5eEKRI=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=IfGud6Cs6KT09h5lfj182Fy62mOf6cono4GROjaats1gZqZ0EcDVvKpLsqHG9Nh90 06w7p/a6O2IgT30GCQg+JhT7b142Gmvxur8RjzoWNI/jDjTo2iUhdaA5EoJ6nWODnM uxiKZOgC6hrRKYrWdSJijacXr8QlmWCeplFRwkWxik9IG5FbIPON8sfNf4oQ2YxHgw Po/3qZUQECQenfhIvjTuePvNJC2Vmmvtj8hw0Fvju88lDerdpgKbJ5Z8T52fjRS7+r 9WdXj3RkTSnP2aN8zNMbA46s5/e/HklV3ZBYC1zo9BYbPJJ9uYAy6UVzGN9TDd+gCV p0i1cjYTJLC/A== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Geliang Tang Subject: [PATCH mptcp-next v13 08/10] mptcp: add sendmsg_locked to proto_ops Date: Tue, 1 Sep 2026 14:49:01 +0800 Message-ID: <5cb5300425ba59afb6b3a99041659bc3908bdf6a.1788245079.git.tanggeliang@kylinos.cn> X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Geliang Tang MPTCP currently provides a standard sendmsg() implementation which acquires and releases the socket lock internally. However, certain upper layers need to call the sendmsg method while the socket lock is already held. Split the existing mptcp_sendmsg() into mptcp_sendmsg_locked() which assumes the caller holds the socket lock, and a tiny wrapper mptcp_sendmsg() that acquires the lock and calls the locked version. Expose .sendmsg_locked in both mptcp_stream_ops and mptcp_v6_stream_ops. Signed-off-by: Geliang Tang --- net/mptcp/protocol.c | 18 ++++++++++++++---- 1 file changed, 14 insertions(+), 4 deletions(-) diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index 45c05f896126..2f8ee18ba6fe 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -2056,7 +2056,7 @@ static void mptcp_rps_record_subflows(const struct mp= tcp_sock *msk) } } =20 -static int mptcp_sendmsg(struct sock *sk, struct msghdr *msg, size_t len) +static int mptcp_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_= t len) { struct mptcp_sock *msk =3D mptcp_sk(sk); struct page_frag *pfrag; @@ -2068,8 +2068,6 @@ static int mptcp_sendmsg(struct sock *sk, struct msgh= dr *msg, size_t len) msg->msg_flags &=3D MSG_MORE | MSG_DONTWAIT | MSG_NOSIGNAL | MSG_FASTOPEN | MSG_EOR; =20 - lock_sock(sk); - mptcp_rps_record_subflows(msk); =20 if (unlikely(inet_test_bit(DEFER_CONNECT, sk) || @@ -2185,7 +2183,6 @@ static int mptcp_sendmsg(struct sock *sk, struct msgh= dr *msg, size_t len) } =20 out: - release_sock(sk); return copied; =20 do_error: @@ -2196,6 +2193,17 @@ static int mptcp_sendmsg(struct sock *sk, struct msg= hdr *msg, size_t len) goto out; } =20 +static int mptcp_sendmsg(struct sock *sk, struct msghdr *msg, size_t len) +{ + int ret; + + lock_sock(sk); + ret =3D mptcp_sendmsg_locked(sk, msg, len); + release_sock(sk); + + return ret; +} + static void mptcp_rcv_space_adjust(struct mptcp_sock *msk, int copied); =20 static void mptcp_eat_recv_skb(struct sock *sk, struct sk_buff *skb) @@ -4877,6 +4885,7 @@ static const struct proto_ops mptcp_stream_ops =3D { .read_sock =3D mptcp_read_sock, .splice_read =3D mptcp_splice_read, .peek_len =3D mptcp_peek_len, + .sendmsg_locked =3D mptcp_sendmsg_locked, }; =20 static struct inet_protosw mptcp_protosw =3D { @@ -5002,6 +5011,7 @@ static const struct proto_ops mptcp_v6_stream_ops =3D= { .read_sock =3D mptcp_read_sock, .splice_read =3D mptcp_splice_read, .peek_len =3D mptcp_peek_len, + .sendmsg_locked =3D mptcp_sendmsg_locked, }; =20 static struct proto mptcp_v6_prot; --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 55A7436EAAE for ; Tue, 1 Sep 2026 06:49:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245367; cv=none; b=iAAkduurnkmV7I6ibCL7TRFNZmUF4xTJYCUsXS+7wd4nXC3GCE22VcR7As3fgZzKNqtBLbf5qRI32DqRz4HwMTjIjnjpQmUPy23PKo/ozQwOKPXuHqRGw0osVNK7vA0UJTjdb56S9efPVbEtS294k5rDEYyDY4/M4fQCzFVElUc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245367; c=relaxed/simple; bh=dWfciBFnbFLxCR/Cb8OaHKUIjntSqwGMLlUYKRoL0Xw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=qbKpB2zdxs2mFH8i5ud/b3v/EDPslc3J492EPoTvBmaeU/CnPZjkzwnJXX385FvN5FKuolOBPUlOPHn5R9ds9pjc9qJN8GPx3BeknN1mkkVPdVOPQ3vCo7hcf6mAVUVZQXkhUqoFEu6IiyBeh3S4h19isuu4kgWvvaGQp9sXR6I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=iTKh5HEN; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="iTKh5HEN" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 581F01F000E9; Tue, 1 Sep 2026 06:49:25 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245366; bh=Kn0XXosUmxZlT15rsZYgnCYz5C5ZquIhEFVLPan5W9M=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=iTKh5HENpY97NeDzFqxJIaH2mUqGoh1PyD2UY6yUdBn5WaTUJBqG50BSpKQGEp12E 7VykEWEY0gyqV+YnhzsakOH+vpFhdcZvwWCQ3c1FxrZPiNf0+lzxkFC+8CIGu++Z+C T2IHufCTKsQujvUlLgao5HT65T/QDzA65HSHieROOixq5jPGVlkWBs1XnCSeWVlRSk VaD051hz8EEjn05RuJT8OuVM6YcP15UyKb0fn5SQgorNW0nWlrrPJJ4fn+B0i4jZYr pMMywFAnYgyF/EpsgFSXtIEtwQjkuF0e35pfZJUc6xeFaR6p/EI9gWIRbrEj5PLcWM r1ru5Pvs17D+g== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Geliang Tang Subject: [PATCH mptcp-next v13 09/10] mptcp: track app-limited state in mptcp_sendmsg Date: Tue, 1 Sep 2026 14:49:02 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Geliang Tang The application-limited accounting in TCP is updated by tcp_rate_check_app_limited(), which currently takes a struct sock * and internally calls tcp_sk(). MPTCP needs to apply the same accounting to each subflow individually - every subflow is an independent TCP socket with its own tp->app_limited / delivered state - so wrapping the call as a struct sock * -> tcp_sk() helper is awkward at the call site. Split the existing function: keep the logic as tcp_sock_rate_check_app_limited(struct tcp_sock *tp), and turn tcp_rate_check_app_limited(struct sock *) into a thin wrapper so the exported API is unchanged for other TCP users. Then add mptcp_sock_rate_check_app_limited() that walks every subflow of the mptcp_sock and runs tcp_sock_rate_check_app_limited() under each subflow's socket lock. Invoke it from mptcp_sendmsg() right after the send-side setup, so the delivery-rate app_limited state stays in sync with what the application actually has to send across each subflow. With this in place, TCP_INFO.tcpi_delivery_rate_app_limited is reported correctly for MPTCP connections instead of being left at 0. Signed-off-by: Geliang Tang --- include/net/tcp.h | 1 + net/ipv4/tcp.c | 9 +++++++-- net/mptcp/protocol.c | 18 ++++++++++++++++++ 3 files changed, 26 insertions(+), 2 deletions(-) diff --git a/include/net/tcp.h b/include/net/tcp.h index 436495ff2271..6a148c897bbe 100644 --- a/include/net/tcp.h +++ b/include/net/tcp.h @@ -848,6 +848,7 @@ static inline int tcp_bound_to_half_wnd(struct tcp_sock= *tp, int pktsize) =20 /* tcp.c */ void tcp_get_info(struct sock *, struct tcp_info *); +void tcp_sock_rate_check_app_limited(struct tcp_sock *tp); void tcp_rate_check_app_limited(struct sock *sk); =20 /* Read 'sendfile()'-style from a TCP socket */ diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c index c3d8046b552f..02f0d3091610 100644 --- a/net/ipv4/tcp.c +++ b/net/ipv4/tcp.c @@ -1096,9 +1096,9 @@ int tcp_sendmsg_fastopen(struct sock *sk, struct msgh= dr *msg, int *copied, } =20 /* If a gap is detected between sends, mark the socket application-limited= . */ -void tcp_rate_check_app_limited(struct sock *sk) +void tcp_sock_rate_check_app_limited(struct tcp_sock *tp) { - struct tcp_sock *tp =3D tcp_sk(sk); + struct sock *sk =3D (struct sock *)tp; =20 if (/* We have less than one packet to send. */ tp->write_seq - tp->snd_nxt < tp->mss_cache && @@ -1111,6 +1111,11 @@ void tcp_rate_check_app_limited(struct sock *sk) tp->app_limited =3D (tp->delivered + tcp_packets_in_flight(tp)) ? : 1; } + +void tcp_rate_check_app_limited(struct sock *sk) +{ + tcp_sock_rate_check_app_limited(tcp_sk(sk)); +} EXPORT_SYMBOL_GPL(tcp_rate_check_app_limited); =20 int tcp_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_t size) diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index 2f8ee18ba6fe..b00390a46eee 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -2056,6 +2056,21 @@ static void mptcp_rps_record_subflows(const struct m= ptcp_sock *msk) } } =20 +static void mptcp_rate_check_app_limited(struct sock *sk) +{ + struct mptcp_sock *msk =3D mptcp_sk(sk); + struct mptcp_subflow_context *subflow; + + mptcp_for_each_subflow(msk, subflow) { + struct sock *ssk =3D mptcp_subflow_tcp_sock(subflow); + bool slow; + + slow =3D lock_sock_fast_nested(ssk); + tcp_sock_rate_check_app_limited(tcp_sk(ssk)); + unlock_sock_fast(ssk, slow); + } +} + static int mptcp_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_= t len) { struct mptcp_sock *msk =3D mptcp_sk(sk); @@ -2084,6 +2099,9 @@ static int mptcp_sendmsg_locked(struct sock *sk, stru= ct msghdr *msg, size_t len) =20 timeo =3D sock_sndtimeo(sk, msg->msg_flags & MSG_DONTWAIT); =20 + /* is sending application-limited? */ + mptcp_rate_check_app_limited(sk); + if ((1 << sk->sk_state) & ~(TCPF_ESTABLISHED | TCPF_CLOSE_WAIT)) { ret =3D sk_stream_wait_connect(sk, &timeo); if (ret) --=20 2.53.0 From nobody Sat Sep 5 05:49:25 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EC9C74CCDD5 for ; Tue, 1 Sep 2026 06:49:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245369; cv=none; b=ngVHTdACL4hyy+rRhlJ8mbSb0U+q2co1zJtf8h1ydtXlZKDNzEDgnq9gGCGsvHkB+RNFd11+zxaRQkdMOnmd6xRcUO6iAcHd0TSpdd+y70AS19p8WE6oslgSjVy/ctr77ASy8YqOp3tz8TS/L4sKIBrYfK0IBnuLjFR2PDehV8Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788245369; c=relaxed/simple; bh=JxT3qeiSHuBS12KtGnTPLFRzVxKd0fvCXwqrRkI3nnM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=CBDiqVb9fNH330I+f0JAVtrv+66iabYOPnmHCYo2OQ8s7LoQVthggvrxC87dhH/ufqFtS1NTVBVYeARr9C8cpC0ov5IA/lS3A6oJdv32J4UVuQuY0TE7vAcUA3z8T2vo9YqsjYPgA2NnBUA5MK0NhfGDdsIZ7mb6Tb3xrcN7k4Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=kymqUGxI; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="kymqUGxI" Received: by smtp.kernel.org (Postfix) with ESMTPSA id BF7FA1F00A3F; Tue, 1 Sep 2026 06:49:26 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788245367; bh=xNNDwMcimuhvgxiiYj/oGzahS6NB1VtBfCB7+LyvfjA=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=kymqUGxIrMaS72pqft9Sxzf363HmYbbRagV2HQidtcEC8B4iLaxtTCW+f9HfP4aRX umlK0TEENyhggupmXNdHvNlfDs3H8y5s4A3uiALrkvupwbZdyb3J1r7Eqv24sUVpj/ EURfVG6crFF5i6QtEtaRWHen7Hss42Lf5pSBmSEr0btkLNWr6QnQVFswiS9nVIlYZa 8wh5H787EE9VGxuV4vlxg9r6/fj0S1CaKGPg7H8rv/+NSUYDZgWcRxNKa1JSI+gULG FuzoXmWoZiDs0qbf2J+WvU06DWK9eJIdUVbiJlOjSHtcO+iE6pyM3BwNgTllzmNF6H dznl6Kyfq6RWw== From: Geliang Tang To: mptcp@lists.linux.dev Cc: Geliang Tang Subject: [PATCH mptcp-next v13 10/10] selftests: mptcp: sockopt: check app_limited Date: Tue, 1 Sep 2026 14:49:03 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: mptcp@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Geliang Tang connect_one_server() in mptcp_sockopt exchanges only a few packets between the client and server, then closes the socket. After such a small transfer the application has nothing further to send, so the connection is, by definition, application-limited. Extend the TCP_INFO readback at the end of the function to assert s.tcp_info.tcpi_delivery_rate_app_limited =3D=3D 1. Without the preceding commit, mptcp_sendmsg() never updates the per-subflow app-limited state, and this field stays at 0 - the assertion would fail. With it in place, the value is forced to 1, turning this into a regression guard for the subflow-side application-limited accounting. Signed-off-by: Geliang Tang --- tools/testing/selftests/net/mptcp/mptcp_sockopt.c | 1 + 1 file changed, 1 insertion(+) diff --git a/tools/testing/selftests/net/mptcp/mptcp_sockopt.c b/tools/test= ing/selftests/net/mptcp/mptcp_sockopt.c index d68515b7903b..8d712bdb4325 100644 --- a/tools/testing/selftests/net/mptcp/mptcp_sockopt.c +++ b/tools/testing/selftests/net/mptcp/mptcp_sockopt.c @@ -640,6 +640,7 @@ static void connect_one_server(int fd, int pipefd) total +=3D 1; /* sequence advances due to FIN */ =20 assert(s.mptcpi_rcv_delta =3D=3D (uint64_t)total); + assert(s.tcp_info.tcpi_delivery_rate_app_limited =3D=3D 1); close(fd); } =20 --=20 2.53.0