[PATCH] scsi: check that the tag map is still there in scsi_host_find_tag()

Yehyeong Lee posted 1 patch 1 month, 2 weeks ago
include/scsi/scsi_tcq.h | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
[PATCH] scsi: check that the tag map is still there in scsi_host_find_tag()
Posted by Yehyeong Lee 1 month, 2 weeks ago
scsi_host_find_tag() bounds the hardware queue index against
tag_set.nr_hw_queues and then dereferences tag_set.tags[hwq].
blk_mq_free_tag_set() clears the tags - __blk_mq_free_map_and_rqs()
sets each tags[i] to NULL and the array itself is freed afterwards -
but it never reduces nr_hw_queues, so the bound still passes and the
dereference is on NULL.

A driver that looks a tag up while its host is being removed therefore
faults.  ib_srp does: srp_remove_target() calls scsi_remove_host()
before it disconnects the target and destroys the queue pair, so an
SRP_RSP the initiator did not ask for reaches srp_process_rsp() after
the tag map is gone.

  [    8.800679] Oops: general protection fault, probably for non-canonical address 0xdffffc0000000001: 0000 [#1] SMP KASAN NOPTI
  [    8.802155] KASAN: null-ptr-deref in range [0x0000000000000008-0x000000000000000f]
  [    8.803134] CPU: 1 UID: 0 PID: 31 Comm: kworker/u8:1 Not tainted 7.2.0-rc5-PRIST2B-gf5098b6bae76-dirty #21 PREEMPT(lazy)
  [    8.804503] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
  [    8.805965] Workqueue: rxe_wq do_work
  [    8.806548] RIP: 0010:srp_recv_done+0x618/0x1aa0
  [    8.807184] Code: c1 e8 03 41 80 3c 30 00 0f 85 20 10 00 00 48 8b b0 38 01 00 00 48 8d 14 d6 48 be 00 00 00 00 00 fc ff df 48 89 d7 48 c1 ef 03 <80> 3c 37 00 0f 85 bb 0f 00 00 48 be 00 00 00 00 00 fc ff df 48 8b
  [    8.807759] ib_srpt DIAG2B: RDMA_CM_EVENT_DISCONNECTED posts=2693 ok=2689 flush=0 other=0
  [    8.809474] RSP: 0018:ffff88811b108d10 EFLAGS: 00010202
  [    8.809481] RAX: ffff888106098000 RBX: ffff88810603a180 RCX: 0000000000010006
  [    8.809484] RDX: 0000000000000008 RSI: dffffc0000000000 RDI: 0000000000000001
  [    8.809488] RBP: ffff888103baf360 R08: 1ffff11020c13027 R09: ffff888107598008
  [    8.809491] R10: ffff888106036048 R11: ffff888107598000 R12: ffff88810485c000
  [    8.809494] R13: ffff88810603a1ec R14: ffff8881060988a8 R15: 0000000000000006
  [    8.810670] ib_srpt DIAG2B: replay stopped posts=2693 ok=2689 flush=0 other=0
  [    8.811241] FS:  0000000000000000(0000) GS:ffff8881673a5000(0000) knlGS:0000000000000000
  [    8.811250] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
  [    8.812342] ib_srpt receiving failed for ioctx 0000000023edb106 with status 5
  [    8.813141] CR2: 000000000fcc7151 CR3: 0000000102313001 CR4: 0000000000770ef0
  [    8.813153] PKRU: 55555554
  [    8.813155] Call Trace:
  [    8.813159]  <IRQ>
  [    8.813163]  ? net_rx_action+0x349/0xfb0
  [    8.814224] ib_srpt receiving failed for ioctx 000000005729ebde with status 5
  [    8.815119]  ? __pfx_srp_recv_done+0x10/0x10
  [    8.815153]  ? rxe_poll_cq+0x253/0x3d0
  [    8.815161]  ? enqueue_task_fair+0x70f/0x2b60
  [    8.816188] ib_srpt receiving failed for ioctx 00000000a48ac180 with status 5
  [    8.816997]  __ib_process_cq+0xe1/0x390
  [    8.818042] ib_srpt receiving failed for ioctx 00000000e493e48a with status 5
  [    8.818769]  ib_poll_handler+0x6e/0x200
  [    8.819685] ib_srpt receiving failed for ioctx 00000000596851d8 with status 5
  [    8.820594]  irq_poll_softirq+0x1df/0x480
  [    8.820968] ib_srpt receiving failed for ioctx 00000000acb38618 with status 5
  [    8.821300]  ? __pfx_irq_poll_softirq+0x10/0x10
  [    8.821582] ib_srpt receiving failed for ioctx 00000000d2f29888 with status 5
  [    8.822094]  ? __pfx_sched_ttwu_pending+0x10/0x10
  [    8.823079] ib_srpt receiving failed for ioctx 00000000cd218270 with status 5
  [    8.823632]  handle_softirqs+0x18e/0x590
  [    8.824134] ib_srpt receiving failed for ioctx 00000000c49ed88c with status 5
  [    8.824694]  ? __pfx_handle_softirqs+0x10/0x10
  [    8.825607] ib_srpt receiving failed for ioctx 000000004d7feb6d with status 5
  [    8.826105]  do_softirq+0x3b/0x60
  [    8.826110]  </IRQ>
  [    8.827053] ib_srpt DIAG2B: RDMA_CM_EVENT_DISCONNECTED posts=2693 ok=2689 flush=0 other=0
  [    8.827515]  <TASK>
  [    8.838045]  __local_bh_enable_ip+0x61/0x70
  [    8.838594]  __alloc_skb+0x732/0x890
  [    8.839093]  ? _raw_spin_lock_irqsave+0x85/0xe0
  [    8.839790]  ? __pfx___alloc_skb+0x10/0x10
  [    8.840341]  ? _raw_read_unlock_irqrestore+0x16/0x50
  [    8.841007]  rxe_init_packet+0x16b/0x4f0
  [    8.841544]  prepare_ack_packet+0xb8/0x830
  [    8.842088]  rxe_receiver+0x499/0x9980
  [    8.842590]  ? __pfx_rxe_receiver+0x10/0x10
  [    8.843140]  ? rxe_completer+0x29e5/0x38c0
  [    8.843679]  ? pick_task_fair+0xbfc/0x19b0
  [    8.844226]  ? __pfx__raw_spin_lock_irqsave+0x10/0x10
  [    8.844884]  ? __pfx_rxe_receiver+0x10/0x10
  [    8.845440]  do_work+0x144/0x470
  [    8.845875]  process_one_work+0x633/0x1030
  [    8.846447]  ? assign_work+0x11d/0x370
  [    8.846972]  worker_thread+0x45b/0xd10
  [    8.847521]  ? __pfx_worker_thread+0x10/0x10
  [    8.848126]  kthread+0x2c6/0x3b0
  [    8.848592]  ? recalc_sigpending+0x15c/0x1e0
  [    8.849213]  ? __pfx_kthread+0x10/0x10
  [    8.849737]  ret_from_fork+0x36e/0x5a0
  [    8.850289]  ? __pfx_ret_from_fork+0x10/0x10
  [    8.850884]  ? __switch_to+0x572/0xdd0
  [    8.851430]  ? __pfx_kthread+0x10/0x10
  [    8.851962]  ret_from_fork_asm+0x1a/0x30
  [    8.852548]  </TASK>
  [    8.852872] Modules linked in: ib_srpt
  [    8.853455] ---[ end trace 0000000000000000 ]---

blk_mq_tagset_busy_iter() reads the same array and tests both the array
and the element before using them.  Do the same here.  Its SRCU section
covers the tags being freed; the tests cover them being cleared, which
is what faults above.

Fixes: 1ee8e889d946 ("scsi: add support for multiple hardware queues in scsi_(host_)find_tag")
Cc: stable@vger.kernel.org
Signed-off-by: Yehyeong Lee <yhlee@isslab.korea.ac.kr>
---
Measured over rxe with KASAN, with an SRP target that reposts an SRP_RSP
for a command it has already answered while I/O runs and the target is
deleted through sysfs: the report above appeared in 5 of 5 runs without
this patch and in none of 3 with it, on 7.2-rc5 with no other change.  A
conforming target is unaffected - the same 3 runs show no aborts and no
error completions, matching an unpatched kernel.
 include/scsi/scsi_tcq.h | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/include/scsi/scsi_tcq.h b/include/scsi/scsi_tcq.h
index ea7848e74d257..d62bae05d4e7d 100644
--- a/include/scsi/scsi_tcq.h
+++ b/include/scsi/scsi_tcq.h
@@ -29,7 +29,8 @@ static inline struct scsi_cmnd *scsi_host_find_tag(struct Scsi_Host *shost,
 		return NULL;
 
 	hwq = blk_mq_unique_tag_to_hwq(tag);
-	if (hwq < shost->tag_set.nr_hw_queues) {
+	if (hwq < shost->tag_set.nr_hw_queues && shost->tag_set.tags &&
+	    shost->tag_set.tags[hwq]) {
 		req = blk_mq_tag_to_rq(shost->tag_set.tags[hwq],
 					blk_mq_unique_tag_to_tag(tag));
 	}
-- 
2.43.0
Re: [PATCH] scsi: check that the tag map is still there in scsi_host_find_tag()
Posted by Bart Van Assche 1 month, 2 weeks ago
On 8/13/26 6:04 PM, Yehyeong Lee wrote:
> diff --git a/include/scsi/scsi_tcq.h b/include/scsi/scsi_tcq.h
> index ea7848e74d257..d62bae05d4e7d 100644
> --- a/include/scsi/scsi_tcq.h
> +++ b/include/scsi/scsi_tcq.h
> @@ -29,7 +29,8 @@ static inline struct scsi_cmnd *scsi_host_find_tag(struct Scsi_Host *shost,
>   		return NULL;
>   
>   	hwq = blk_mq_unique_tag_to_hwq(tag);
> -	if (hwq < shost->tag_set.nr_hw_queues) {
> +	if (hwq < shost->tag_set.nr_hw_queues && shost->tag_set.tags &&
> +	    shost->tag_set.tags[hwq]) {
>   		req = blk_mq_tag_to_rq(shost->tag_set.tags[hwq],
>   					blk_mq_unique_tag_to_tag(tag));
>   	}

Thanks for the detailed report but I think this is the wrong way to fix
the reported crash. Please help with testing this patch:


From: Bart Van Assche <bvanassche@acm.org>
Date: Fri, 14 Aug 2026 17:00:55 +0000
Subject: [PATCH] RDMA/srp: Fix srp_remove_target()

Remove all logical units before disconnecting the transport because one or
more SCSI commands may be submitted while removing logical units. Remove
the SCSI host after the transport has been disconnected because the code
that disconnects the transport needs resources that are freed by the code
that removes the SCSI host (SCSI host tag set). Remove the srp_rport_get()
and srp_rport_put() calls because the purpose of these calls was to keep
the rport until tl_err_work is cancelled. This patch fixes the following
crash:

Oops: general protection fault, probably for non-canonical address 
0xdffffc0000000001: 0000 [#1] SMP KASAN NOPTI
RIP: 0010:srp_recv_done+0x618/0x1aa0
  __ib_process_cq+0xe1/0x390
  ib_poll_handler+0x6e/0x200
  irq_poll_softirq+0x1df/0x480
  do_softirq+0x3b/0x60
  </IRQ>

Reported-by: Yehyeong Lee <yhlee@isslab.korea.ac.kr>
Signed-off-by: Bart Van Assche <bvanassche@acm.org>
---
  drivers/infiniband/ulp/srp/ib_srp.c | 14 ++++++++++----
  1 file changed, 10 insertions(+), 4 deletions(-)

diff --git a/drivers/infiniband/ulp/srp/ib_srp.c 
b/drivers/infiniband/ulp/srp/ib_srp.c
index acbd787de265..0b296b5715a8 100644
--- a/drivers/infiniband/ulp/srp/ib_srp.c
+++ b/drivers/infiniband/ulp/srp/ib_srp.c
@@ -1038,15 +1038,20 @@ static void srp_del_scsi_host_attr(struct 
Scsi_Host *shost)

  static void srp_remove_target(struct srp_target_port *target)
  {
+	struct scsi_device *sdev;
  	struct srp_rdma_ch *ch;
  	int i;

  	WARN_ON_ONCE(target->state != SRP_TARGET_REMOVED);

  	srp_del_scsi_host_attr(target->scsi_host);
-	srp_rport_get(target->rport);
-	srp_remove_host(target->scsi_host);
-	scsi_remove_host(target->scsi_host);
+	/*
+	 * Remove all logical units. This must happen before the
+	 * srp_disconnect_target() call because scsi_remove_device() may trigger
+	 * submission of SCSI commands. See also sd_shutdown().
+	 */
+	shost_for_each_device(sdev, target->scsi_host)
+		scsi_remove_device(sdev);
  	srp_stop_rport_timers(target->rport);
  	srp_disconnect_target(target);
  	kobj_ns_drop(KOBJ_NS_TYPE_NET, to_ns_common(target->net));
@@ -1055,7 +1060,8 @@ static void srp_remove_target(struct 
srp_target_port *target)
  		srp_free_ch_ib(target, ch);
  	}
  	cancel_work_sync(&target->tl_err_work);
-	srp_rport_put(target->rport);
+	srp_remove_host(target->scsi_host);
+	scsi_remove_host(target->scsi_host);
  	kfree(target->ch);
  	target->ch = NULL;
Re: [PATCH] scsi: check that the tag map is still there in scsi_host_find_tag()
Posted by Yehyeong Lee 1 month, 1 week ago
On 8/15/26 2:27 AM, Bart Van Assche wrote:
> Thanks for the detailed report but I think this is the wrong way to fix
> the reported crash. Please help with testing this patch:

Tested.  With the replay target that reposts an SRP_RSP for a command it
has already answered, the crash appears in 5 of 5 runs on an unpatched
kernel and in 0 of 5 with your patch.  The target still delivered about
2680 unsolicited responses per run either way, so what went away is the
window, not the traffic.

I have sent your patch as 2/2 of a v2 series on linux-rdma:

https://lore.kernel.org/linux-rdma/20260818035229.505098-1-yhlee@isslab.korea.ac.kr/

Martin, please drop this patch.  The lookup it guards is not reachable
once srp_remove_target() is fixed, and the review point that the check is
lockless stands.

Tested-by: Yehyeong Lee <yhlee@isslab.korea.ac.kr>

Best regards,

Yehyeong Lee