From nobody Mon Sep 28 16:22:43 2026 Received: from mta0.migadu.com (out-236.mta0.migadu.com [91.218.175.236]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7C7FF38C407 for ; Thu, 20 Aug 2026 05:56:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.236 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787205374; cv=none; b=FG2zIbE3+bmtWbHFT0xy9fymAqEB2W713CTsnE1mgAgJzdm8es+jAhutqQZMdBSLwJ2JpHs2t3okqmzhBMhp1aQiJOoFKu54FfkscwD671rd27c1Mr76IM5vSEmS25hTFoF4tY0CebdtsbJj/ttSgXlsbfjTZl0St4tezRUbvWA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787205374; c=relaxed/simple; bh=Yyt0kO2zJ1JirMOwK56pHnm/TWr/UmxyxweM7ZJg9/s=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Sn4sstdy2D/EP51TnJ+xkbyscuklTl46nVZtdrf2tbfnrW1NulGBnn+s4vJ47UEFrH7ZOINiBQuQV64wBD6iyTQNUyixlO7CbHDWfTZ9vYcpcF9dy/p0o2e57mFiTV8owLEpQwEt3BOLjj93C44j5/+z5dB9ACD7H/JIedm1M/g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=E6hHe2YM; arc=none smtp.client-ip=91.218.175.236 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="E6hHe2YM" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=Yyt0kO2zJ1JirMOwK56pHnm/TWr/UmxyxweM7ZJg9/s=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787205369; v=1; x=1787810169; b=E6hHe2YMyWsjEZfishaYRt9b7JN5w5KY7G/y06JzFmeZcQPUgeIq/o8U5zsnmeWYNkfriftU ZKddQ5TXEdct2GtRlNzkcs0jiDaOue7rD9MPybkitRSelaJtFs3xcjzgcJ8P1UwqfrHi75AZoJk FdCx4/MHKhw1ZrZcQhB/Fqz8= X-Envelope-To: linux-kernel@vger.kernel.org Received: from [192.168.110.119] (216.236.36.152) by smtp.migadu.com with ESMTPS id 1234f702a62e50bb; Thu, 20 Aug 2026 05:56:09 +0000 X-Mizu-Trace-ID: 1234f702a62e50bb X-Migadu-Flow: FLOW_OUT From: Hangbin Liu Date: Thu, 20 Aug 2026 13:55:52 +0800 Subject: [PATCH net v4 1/2] bonding: convert unbalanced_load to per-cpu state Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260820-bond_overflow-v4-1-805ba0d3efb6@kylinos.cn> References: <20260820-bond_overflow-v4-0-805ba0d3efb6@kylinos.cn> In-Reply-To: <20260820-bond_overflow-v4-0-805ba0d3efb6@kylinos.cn> To: Jay Vosburgh , Andrew Lunn , "David S. Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Nikolay Aleksandrov Cc: Hangbin Liu , netdev@vger.kernel.org, linux-kernel@vger.kernel.org, Hangbin Liu X-Mailer: b4 0.14.3 From: Hangbin Liu A later patch widens the bonding TLB tx counters from u32 to u64. The unbalanced_load counter sits in the transmit hot path, and cross-CPU synchronization of a u64 would introduce measurable overhead. Convert unbalanced_load to a per-cpu counter first so that the subsequent widening only touches per-cpu data local to each CPU. Introduce struct unbalanced_load_stats to hold the per-cpu counter, and move the aggregation into a helper, reset_unbalanced_load(), which sums all per-cpu instances. Use the delta of current total load vs variable prev_total_unbalanced to calculate the loading. Signed-off-by: Hangbin Liu --- drivers/net/bonding/bond_alb.c | 37 ++++++++++++++++++++++++++++++------- include/net/bond_alb.h | 7 ++++++- 2 files changed, 36 insertions(+), 8 deletions(-) diff --git a/drivers/net/bonding/bond_alb.c b/drivers/net/bonding/bond_alb.c index 839f7482dc18..0afed2c39231 100644 --- a/drivers/net/bonding/bond_alb.c +++ b/drivers/net/bonding/bond_alb.c @@ -133,6 +133,10 @@ static int tlb_initialize(struct bonding *bond) if (!new_hashtbl) return -ENOMEM; =20 + bond_info->unbalanced_load =3D alloc_percpu(struct unbalanced_load_stats); + if (!bond_info->unbalanced_load) + goto out; + spin_lock_bh(&bond->mode_lock); =20 bond_info->tx_hashtbl =3D new_hashtbl; @@ -143,6 +147,10 @@ static int tlb_initialize(struct bonding *bond) spin_unlock_bh(&bond->mode_lock); =20 return 0; + +out: + kfree(new_hashtbl); + return -ENOMEM; } =20 /* Must be called only after all slaves have been released */ @@ -154,6 +162,8 @@ static void tlb_deinitialize(struct bonding *bond) =20 kfree(bond_info->tx_hashtbl); bond_info->tx_hashtbl =3D NULL; + free_percpu(bond_info->unbalanced_load); + bond_info->prev_total_unbalanced =3D 0; =20 spin_unlock_bh(&bond->mode_lock); } @@ -1345,7 +1355,7 @@ static netdev_tx_t bond_do_alb_xmit(struct sk_buff *s= kb, struct bonding *bond, /* unbalanced or unassigned, send through primary */ tx_slave =3D rcu_dereference(bond->curr_active_slave); if (bond->params.tlb_dynamic_lb) - bond_info->unbalanced_load +=3D skb->len; + this_cpu_add(bond_info->unbalanced_load->tx_bytes, skb->len); } =20 if (tx_slave && bond_slave_can_tx(tx_slave)) { @@ -1529,6 +1539,23 @@ netdev_tx_t bond_alb_xmit(struct sk_buff *skb, struc= t net_device *bond_dev) return bond_do_alb_xmit(skb, bond, tx_slave); } =20 +static u32 reset_unbalanced_load(struct alb_bond_info *bond_info) +{ + struct unbalanced_load_stats *p; + u32 delta, total_bytes =3D 0; + int i; + + for_each_possible_cpu(i) { + p =3D per_cpu_ptr(bond_info->unbalanced_load, i); + total_bytes +=3D READ_ONCE(p->tx_bytes); + } + + delta =3D total_bytes - bond_info->prev_total_unbalanced; + bond_info->prev_total_unbalanced =3D total_bytes; + + return delta / BOND_TLB_REBALANCE_INTERVAL; +} + void bond_alb_monitor(struct work_struct *work) { struct bonding *bond =3D container_of(work, struct bonding, @@ -1570,12 +1597,8 @@ void bond_alb_monitor(struct work_struct *work) if (atomic_read(&bond_info->tx_rebalance_counter) >=3D BOND_TLB_REBALANCE= _TICKS) { bond_for_each_slave_rcu(bond, slave, iter) { tlb_clear_slave(bond, slave, 1); - if (slave =3D=3D rcu_access_pointer(bond->curr_active_slave)) { - SLAVE_TLB_INFO(slave).load =3D - bond_info->unbalanced_load / - BOND_TLB_REBALANCE_INTERVAL; - bond_info->unbalanced_load =3D 0; - } + if (slave =3D=3D rcu_access_pointer(bond->curr_active_slave)) + SLAVE_TLB_INFO(slave).load =3D reset_unbalanced_load(bond_info); } atomic_set(&bond_info->tx_rebalance_counter, 0); } diff --git a/include/net/bond_alb.h b/include/net/bond_alb.h index e5945427f38d..6fb09b4fc7e2 100644 --- a/include/net/bond_alb.h +++ b/include/net/bond_alb.h @@ -123,9 +123,14 @@ struct tlb_slave_info { */ }; =20 +struct unbalanced_load_stats { + u32 tx_bytes; +}; + struct alb_bond_info { struct tlb_client_info *tx_hashtbl; /* Dynamically allocated */ - u32 unbalanced_load; + struct unbalanced_load_stats __percpu *unbalanced_load; + u32 prev_total_unbalanced; atomic_t tx_rebalance_counter; int lp_counter; /* -------- rlb parameters -------- */ --=20 2.55.0 From nobody Mon Sep 28 16:22:43 2026 Received: from mta0.migadu.com (out-245.mta0.migadu.com [91.218.175.245]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 544C53AD51C for ; Thu, 20 Aug 2026 05:56:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.245 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787205380; cv=none; b=fAZ+zWwVfkReSrYMz1IrXxdVezHjNYMcJyTIsemMRcXkycsZHSiMGyMnfCx/UGT1bin4R9X79iKwlPshbqY5FwsumASSIraVo/nC/YGL0/qIQ6eA2+MZbsrX7UqLFgfavnlAKbQJwuZea4vmx+MgmC63qhk+ai2ZnjPEyEZeirs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787205380; c=relaxed/simple; bh=TXPTgcyA9G2NBmGM3doHSytT8ovofAdzx8mGrJYfAXc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=r8OUDgu4hXDKKBkvVCxyFFub//tXsegwy9fflSFRPmSFfnaaF3Qu5rJ5YNzUr1YR1Ih/IfpbmmaTOpusp1KIoz6ucii38m8sFXQvMr0JvkStYierrd9dVJ58FtQ6WBxenVrsBxB5UnMTdqHocXZbyLxlq3ceqyk/JtHElJLva08= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=PmupPcS2; arc=none smtp.client-ip=91.218.175.245 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="PmupPcS2" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=TXPTgcyA9G2NBmGM3doHSytT8ovofAdzx8mGrJYfAXc=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787205376; v=1; x=1787810176; b=PmupPcS2J8qln7JwwcM0v3DWSuNGJjafdKFmFYthFS56i84mZLQ9b6EGSKRHbpR5C990Voko oKrSjAREk1z6kgtqkER4DFhQyCbYVF6d4VJpLHNvFOMl64BUy6u2zmr+MQa5gb9s5beob9CjRzM iwLJqxrZBdIa/XlcT4Ribepo= X-Envelope-To: linux-kernel@vger.kernel.org Received: from [192.168.110.119] (216.236.36.152) by smtp.migadu.com with ESMTPS id 622fa011b17667cf; Thu, 20 Aug 2026 05:56:16 +0000 X-Mizu-Trace-ID: 622fa011b17667cf X-Migadu-Flow: FLOW_OUT From: Hangbin Liu Date: Thu, 20 Aug 2026 13:55:53 +0800 Subject: [PATCH net v4 2/2] bonding: fix u32 overflow in compute_gap() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260820-bond_overflow-v4-2-805ba0d3efb6@kylinos.cn> References: <20260820-bond_overflow-v4-0-805ba0d3efb6@kylinos.cn> In-Reply-To: <20260820-bond_overflow-v4-0-805ba0d3efb6@kylinos.cn> To: Jay Vosburgh , Andrew Lunn , "David S. Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Nikolay Aleksandrov Cc: Hangbin Liu , netdev@vger.kernel.org, linux-kernel@vger.kernel.org, Hangbin Liu X-Mailer: b4 0.14.3 From: Hangbin Liu The TLB load-tracking fields tx_bytes, load_history, load, and unbalanced_load are all u32. At sustained throughput above ~3.2 Gbit/s over the 10-second rebalance interval the byte counters wrap, causing compute_gap() to produce incorrect gap values and mis-select slaves. Such speeds are common on modern NICs under heavy traffic. Widen these fields to u64. Use u64_stats_sync to protect the per-cpu unbalanced_load_stats against tearing on 32-bit architectures, and div_u64() for the 64-bit divisions. The tx_bytes and load_history are protected in spin_lock. Also protect the slave load writing in bond_alb_monitor() with spin_lock in case of tear on 32-bit. Rework compute_gap() to use u64 arithmetic throughout. Return 0 when the speed is unknown or the slave is already overloaded. Detected by AI code review. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Signed-off-by: Hangbin Liu --- drivers/net/bonding/bond_alb.c | 65 ++++++++++++++++++++++++++++++--------= ---- include/net/bond_alb.h | 11 +++---- 2 files changed, 53 insertions(+), 23 deletions(-) diff --git a/drivers/net/bonding/bond_alb.c b/drivers/net/bonding/bond_alb.c index 0afed2c39231..372db54803d3 100644 --- a/drivers/net/bonding/bond_alb.c +++ b/drivers/net/bonding/bond_alb.c @@ -6,6 +6,7 @@ #include #include #include +#include #include #include #include @@ -74,8 +75,8 @@ static inline u8 _simple_hash(const u8 *hash_start, int h= ash_size) static inline void tlb_init_table_entry(struct tlb_client_info *entry, int= save_load) { if (save_load) { - entry->load_history =3D 1 + entry->tx_bytes / - BOND_TLB_REBALANCE_INTERVAL; + entry->load_history =3D 1 + div_u64(entry->tx_bytes, + BOND_TLB_REBALANCE_INTERVAL); entry->tx_bytes =3D 0; } =20 @@ -133,7 +134,7 @@ static int tlb_initialize(struct bonding *bond) if (!new_hashtbl) return -ENOMEM; =20 - bond_info->unbalanced_load =3D alloc_percpu(struct unbalanced_load_stats); + bond_info->unbalanced_load =3D netdev_alloc_pcpu_stats(struct unbalanced_= load_stats); if (!bond_info->unbalanced_load) goto out; =20 @@ -168,27 +169,38 @@ static void tlb_deinitialize(struct bonding *bond) spin_unlock_bh(&bond->mode_lock); } =20 -static long long compute_gap(struct slave *slave) +static u64 compute_gap(struct slave *slave) { - return (s64) (slave->speed << 20) - /* Convert to Megabit per sec */ - (s64) (SLAVE_TLB_INFO(slave).load << 3); /* Bytes to bits */ + u64 slave_load =3D SLAVE_TLB_INFO(slave).load << 3; /* Bytes to bits */ + u32 raw_speed =3D READ_ONCE(slave->speed); + u64 speed =3D (u64)raw_speed << 20; /* Convert to bits per sec */ + + /* It's meaningless to compare gap on unknown speed NIC */ + if (raw_speed =3D=3D (u32)SPEED_UNKNOWN) + return 0; + + /* Skip slave which is over loaded */ + if (speed <=3D slave_load) + return 0; + + return speed - slave_load; } =20 static struct slave *tlb_get_least_loaded_slave(struct bonding *bond) { struct slave *slave, *least_loaded; struct list_head *iter; - long long max_gap; + u64 max_gap =3D 0; =20 least_loaded =3D NULL; - max_gap =3D LLONG_MIN; =20 /* Find the slave with the largest gap */ bond_for_each_slave_rcu(bond, slave, iter) { if (bond_slave_can_tx(slave)) { - long long gap =3D compute_gap(slave); + u64 gap =3D compute_gap(slave); =20 - if (max_gap < gap) { + /* Make sure we have one available slave */ + if (max_gap <=3D gap) { least_loaded =3D slave; max_gap =3D gap; } @@ -1354,8 +1366,14 @@ static netdev_tx_t bond_do_alb_xmit(struct sk_buff *= skb, struct bonding *bond, if (!tx_slave) { /* unbalanced or unassigned, send through primary */ tx_slave =3D rcu_dereference(bond->curr_active_slave); - if (bond->params.tlb_dynamic_lb) - this_cpu_add(bond_info->unbalanced_load->tx_bytes, skb->len); + if (bond->params.tlb_dynamic_lb) { + struct unbalanced_load_stats *pcpu_load; + + pcpu_load =3D this_cpu_ptr(bond_info->unbalanced_load); + u64_stats_update_begin(&pcpu_load->syncp); + u64_stats_add(&pcpu_load->tx_bytes, skb->len); + u64_stats_update_end(&pcpu_load->syncp); + } } =20 if (tx_slave && bond_slave_can_tx(tx_slave)) { @@ -1539,21 +1557,27 @@ netdev_tx_t bond_alb_xmit(struct sk_buff *skb, stru= ct net_device *bond_dev) return bond_do_alb_xmit(skb, bond, tx_slave); } =20 -static u32 reset_unbalanced_load(struct alb_bond_info *bond_info) +static u64 reset_unbalanced_load(struct alb_bond_info *bond_info) { + u64 delta, tx_bytes, total_bytes =3D 0; struct unbalanced_load_stats *p; - u32 delta, total_bytes =3D 0; + unsigned int start; int i; =20 for_each_possible_cpu(i) { p =3D per_cpu_ptr(bond_info->unbalanced_load, i); - total_bytes +=3D READ_ONCE(p->tx_bytes); + do { + start =3D u64_stats_fetch_begin(&p->syncp); + tx_bytes =3D u64_stats_read(&p->tx_bytes); + } while (u64_stats_fetch_retry(&p->syncp, start)); + + total_bytes +=3D tx_bytes; } =20 delta =3D total_bytes - bond_info->prev_total_unbalanced; bond_info->prev_total_unbalanced =3D total_bytes; =20 - return delta / BOND_TLB_REBALANCE_INTERVAL; + return div_u64(delta, BOND_TLB_REBALANCE_INTERVAL); } =20 void bond_alb_monitor(struct work_struct *work) @@ -1597,8 +1621,13 @@ void bond_alb_monitor(struct work_struct *work) if (atomic_read(&bond_info->tx_rebalance_counter) >=3D BOND_TLB_REBALANCE= _TICKS) { bond_for_each_slave_rcu(bond, slave, iter) { tlb_clear_slave(bond, slave, 1); - if (slave =3D=3D rcu_access_pointer(bond->curr_active_slave)) - SLAVE_TLB_INFO(slave).load =3D reset_unbalanced_load(bond_info); + if (slave =3D=3D rcu_access_pointer(bond->curr_active_slave)) { + u64 new_load =3D reset_unbalanced_load(bond_info); + + spin_lock_bh(&bond->mode_lock); + SLAVE_TLB_INFO(slave).load =3D new_load; + spin_unlock_bh(&bond->mode_lock); + } } atomic_set(&bond_info->tx_rebalance_counter, 0); } diff --git a/include/net/bond_alb.h b/include/net/bond_alb.h index 6fb09b4fc7e2..32f1981033e4 100644 --- a/include/net/bond_alb.h +++ b/include/net/bond_alb.h @@ -57,12 +57,12 @@ struct tlb_client_info { * packets to a Client that the Hash function * gave this entry index. */ - u32 tx_bytes; /* Each Client accumulates the BytesTx that + u64 tx_bytes; /* Each Client accumulates the BytesTx that * were transmitted to it, and after each * CallBack the LoadHistory is divided * by the balance interval */ - u32 load_history; /* This field contains the amount of Bytes + u64 load_history; /* This field contains the amount of Bytes * that were transmitted to this client by * the server on the previous balance * interval in Bps. @@ -118,19 +118,20 @@ struct tlb_slave_info { * are the entries that were assigned to use this * slave for transmit. */ - u32 load; /* Each slave sums the loadHistory of all clients + u64 load; /* Each slave sums the loadHistory of all clients * assigned to it */ }; =20 struct unbalanced_load_stats { - u32 tx_bytes; + u64_stats_t tx_bytes; + struct u64_stats_sync syncp; }; =20 struct alb_bond_info { struct tlb_client_info *tx_hashtbl; /* Dynamically allocated */ struct unbalanced_load_stats __percpu *unbalanced_load; - u32 prev_total_unbalanced; + u64 prev_total_unbalanced; atomic_t tx_rebalance_counter; int lp_counter; /* -------- rlb parameters -------- */ --=20 2.55.0