[PATCH net v2] nfc: nci: avoid unbounded skb allocation when max_pkt_payload_len is zero

Liu Chao posted 1 patch 6 days, 3 hours ago
net/nfc/nci/data.c | 12 ++++++++++--
1 file changed, 10 insertions(+), 2 deletions(-)
[PATCH net v2] nfc: nci: avoid unbounded skb allocation when max_pkt_payload_len is zero
Posted by Liu Chao 6 days, 3 hours ago
nci_queue_tx_data_frags() uses conn_info->max_pkt_payload_len as the
fragment size.  When that value is zero, frag_len is always zero and
total_len never decreases.  The loop then allocates skbs without
bound: none of them are freed inside the loop, they accumulate on
frags_q, and there is no cond_resched() in the loop body.  A single
sendmsg() can therefore consume all allocatable memory, and on
CONFIG_PREEMPT_NONE it occupies the CPU long enough to trip the
softlockup watchdog:

  watchdog: BUG: soft lockup - CPU#3 stuck for 26s! [kworker/3:1:57]
  Workqueue: events rawsock_tx_work [nfc]
  Call Trace:
   nci_send_data+0x1ca/0x6b0 [nci]
   nci_transceive+0xbb/0x170 [nci]
   rawsock_tx_work+0xb5/0x1a0 [nfc]

max_pkt_payload_len comes straight from controller-supplied fields,
with no check for zero:

  ntf.c: conn_info->max_pkt_payload_len = ntf.max_data_pkt_payload_size;
  rsp.c: conn_info->max_pkt_payload_len = rsp->max_ctrl_pkt_payload_len;

Reject the zero value in the fragmentation path rather than at the
assignment sites.  nci_queue_tx_data_frags() is the only place that
loops over the RF data path's conn_info, and nci_send_data() takes
the non-fragmenting branch only for skb->len <= max_pkt_payload_len,
which for a zero limit means empty skbs alone.  Validating on
assignment would not be sufficient either, because
nci_rf_disc_rsp_packet() allocates ndev->rf_conn_info with
devm_kzalloc(), so max_pkt_payload_len is already zero before any
notification arrives.

No legitimate configuration is affected: where the NCI spec does
mandate a zero Max Data Packet Payload Size -- the NFCEE Direct RF
Interface -- nci_rf_intf_activated_ntf_packet() takes the "goto
listen" shortcut, bypassing the assignment entirely.

Snapshot the field once with READ_ONCE() and use the snapshot for
both the check and the min_t() bound: the rx workqueue updates the
field without any lock held against this path, so without the
snapshot the check could validate a value the loop no longer
consumes.

nci_hci_send_data() also loops over the same field, but on a
different conn_info instance (ndev->hci_dev->conn_info) created by
nci_core_conn_create_rsp_packet(); a zero or one there underflows
the loop arithmetic and will be addressed in a separate patch.  This
guards the path carrying the reported bug.

Fixes: 6a2968aaf50c ("NFC: basic NCI protocol implementation")
Cc: stable@vger.kernel.org
Signed-off-by: Liu Chao <liuc63@xiaopeng.com>

---

v1 claimed that nci_queue_tx_data_frags() is "the only place that
loops" over max_pkt_payload_len.  That holds for the RF data path
(ndev->rf_conn_info) but was overstated as a blanket claim: the
Sashiko review of v1 pointed out that nci_hci_send_data() loops over
the same field as well, albeit on a separate conn_info instance
(ndev->hci_dev->conn_info) -- pre-existing and unchanged here; it
will be fixed separately.

The READ_ONCE() snapshot follows the same review: with the check and
the loop reading the field independently, a store from the rx
workqueue in between would let the loop spin on a value the check
had just rejected.
---
 net/nfc/nci/data.c | 12 ++++++++++--
 1 file changed, 10 insertions(+), 2 deletions(-)

diff --git a/net/nfc/nci/data.c b/net/nfc/nci/data.c
index 4253edea5..eeb5260c5 100644
--- a/net/nfc/nci/data.c
+++ b/net/nfc/nci/data.c
@@ -104,6 +104,7 @@ static int nci_queue_tx_data_frags(struct nci_dev *ndev,
 	struct sk_buff_head frags_q;
 	struct sk_buff *skb_frag;
 	int frag_len;
+	u8 max_len;
 	int rc = 0;
 
 	pr_debug("conn_id 0x%x, total_len %d\n", conn_id, total_len);
@@ -114,11 +115,18 @@ static int nci_queue_tx_data_frags(struct nci_dev *ndev,
 		goto exit;
 	}
 
+	/* the rx workqueue may update the field concurrently */
+	max_len = READ_ONCE(conn_info->max_pkt_payload_len);
+
+	if (!max_len) {
+		rc = -EPROTO;
+		goto exit;
+	}
+
 	__skb_queue_head_init(&frags_q);
 
 	while (total_len) {
-		frag_len =
-			min_t(int, total_len, conn_info->max_pkt_payload_len);
+		frag_len = min_t(int, total_len, max_len);
 
 		skb_frag = nci_skb_alloc(ndev,
 					 (NCI_DATA_HDR_SIZE + frag_len),
-- 
2.50.1
Re: [PATCH net v2] nfc: nci: avoid unbounded skb allocation when max_pkt_payload_len is zero
Posted by Simon Horman 1 day, 11 hours ago
On Sat, Sep 19, 2026 at 02:54:58AM +0800, Liu Chao wrote:
> nci_queue_tx_data_frags() uses conn_info->max_pkt_payload_len as the
> fragment size.  When that value is zero, frag_len is always zero and
> total_len never decreases.  The loop then allocates skbs without
> bound: none of them are freed inside the loop, they accumulate on
> frags_q, and there is no cond_resched() in the loop body.  A single
> sendmsg() can therefore consume all allocatable memory, and on
> CONFIG_PREEMPT_NONE it occupies the CPU long enough to trip the
> softlockup watchdog:
> 
>   watchdog: BUG: soft lockup - CPU#3 stuck for 26s! [kworker/3:1:57]
>   Workqueue: events rawsock_tx_work [nfc]
>   Call Trace:
>    nci_send_data+0x1ca/0x6b0 [nci]
>    nci_transceive+0xbb/0x170 [nci]
>    rawsock_tx_work+0xb5/0x1a0 [nfc]
> 
> max_pkt_payload_len comes straight from controller-supplied fields,
> with no check for zero:
> 
>   ntf.c: conn_info->max_pkt_payload_len = ntf.max_data_pkt_payload_size;
>   rsp.c: conn_info->max_pkt_payload_len = rsp->max_ctrl_pkt_payload_len;
> 
> Reject the zero value in the fragmentation path rather than at the
> assignment sites.  nci_queue_tx_data_frags() is the only place that
> loops over the RF data path's conn_info, and nci_send_data() takes
> the non-fragmenting branch only for skb->len <= max_pkt_payload_len,
> which for a zero limit means empty skbs alone.  Validating on
> assignment would not be sufficient either, because
> nci_rf_disc_rsp_packet() allocates ndev->rf_conn_info with
> devm_kzalloc(), so max_pkt_payload_len is already zero before any
> notification arrives.
> 
> No legitimate configuration is affected: where the NCI spec does
> mandate a zero Max Data Packet Payload Size -- the NFCEE Direct RF
> Interface -- nci_rf_intf_activated_ntf_packet() takes the "goto
> listen" shortcut, bypassing the assignment entirely.
> 
> Snapshot the field once with READ_ONCE() and use the snapshot for
> both the check and the min_t() bound: the rx workqueue updates the
> field without any lock held against this path, so without the
> snapshot the check could validate a value the loop no longer
> consumes.
> 
> nci_hci_send_data() also loops over the same field, but on a
> different conn_info instance (ndev->hci_dev->conn_info) created by
> nci_core_conn_create_rsp_packet(); a zero or one there underflows
> the loop arithmetic and will be addressed in a separate patch.  This
> guards the path carrying the reported bug.
> 
> Fixes: 6a2968aaf50c ("NFC: basic NCI protocol implementation")
> Cc: stable@vger.kernel.org
> Signed-off-by: Liu Chao <liuc63@xiaopeng.com>

A note on process: Please send new patch revisions in new email threads,
not as a response to an earlier revision. Thanks!

> 
> ---
> 
> v1 claimed that nci_queue_tx_data_frags() is "the only place that
> loops" over max_pkt_payload_len.  That holds for the RF data path
> (ndev->rf_conn_info) but was overstated as a blanket claim: the
> Sashiko review of v1 pointed out that nci_hci_send_data() loops over
> the same field as well, albeit on a separate conn_info instance
> (ndev->hci_dev->conn_info) -- pre-existing and unchanged here; it
> will be fixed separately.
> 
> The READ_ONCE() snapshot follows the same review: with the check and
> the loop reading the field independently, a store from the rx
> workqueue in between would let the loop spin on a value the check
> had just rejected.

There is a fresh AI-generated review of this patchset at
https://netdev-ai.bots.linux.dev/sashiko/#/message/20260918185209.2675909-1-liuc63%40xiaopeng.com

It raises a concern that the snapshot may not be broad enough.
Please take a look.

  The comment documents that the rx workqueue can update the field
  concurrently, but the only caller still reads the same field of the same
  conn_info with a plain load, and it is that read which decides whether
  fragmentation happens at all:

  net/nfc/nci/data.c:nci_send_data() {
	/* check if the packet need to be fragmented */
	if (skb->len <= conn_info->max_pkt_payload_len) {

  Should this read be part of the same snapshot?  Two things follow from
  leaving it as is.  The plain-versus-marked access pair on the field is
  still KCSAN-reportable, and the decision and the loop can consume
  different values: if nci_rf_intf_activated_ntf_packet() in
  net/nfc/nci/ntf.c lowers the limit (255 -> 64, say) right after
  nci_send_data() read the old value, nci_send_data() queues an
  unfragmented packet larger than the controller's current maximum, a
  transition the callee's snapshot cannot observe.  Would snapshotting once
  in nci_send_data() and passing the validated value down be preferable?


It also raises a concern regarding logging.  And while I do agree that the
logging in question should be rate limited, because it occurs on the data
path and is controlled by remote data, I am not convinced that it isn't a
pre-existing problem.

In any case, I think it should been addressed sooner or later.

  Does this new return value make the caller's log line reachable once per
  transmit attempt?

  net/nfc/nci/data.c:nci_send_data() {
		rc = nci_queue_tx_data_frags(ndev, conn_id, skb);
		if (rc) {
			pr_err("failed to fragment tx data packet\n");
			goto free_exit;

  Before the patch this input never returned from the helper, so the site
  was effectively unreachable for it.  With a controller reporting
  max_pkt_payload_len == 0, every sendmsg() on an NFC raw socket
  (rawsock_tx_work -> nfc_data_exchange -> nci_transceive -> nci_send_data)
  now logs a line, and so does every received HCP command, since
  nci_hci_cmd_received() answers each one with nci_hci_send_data(ndev, pipe,
  status, NULL, 0) -> nci_send_data().  Neither side is rate limited.
  Would pr_err_ratelimited(), or a ratelimited message at the new check
  itself, be a better fit for a controller-data-driven error path?

...