From nobody Sat Sep 26 21:59:52 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8FF3A355F35; Fri, 28 Aug 2026 21:47:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787953627; cv=none; b=WWeBArGzdynCmUxhof+urHXp8bPy8XqueATh4OHGhgpzW0EmFv6RNYsb+7iSZVQFI07IMTcJ/+x8De+YmXEeISVooG9+57qSXLmpR2w6ltjutYntsSC5q0QY1JQU7mv2at0wnekv2ao8+HM1F9nQgLxhiN5Xa9R3+hTcUNFGR30= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787953627; c=relaxed/simple; bh=eu7Cuj2YltU5X+0ytry+XH1NLU2ZDQNXc955su8ANfI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=YCB2VMC3n70FN2DPtjn1e3NM5mHi0vRIr7rNrXSauuaZ66mU2rincT5p+JV57+Lu5f/TeU+DPJ44QQBum1484kbipRyxzLKpYgfgiRCEnYwzpPM8EhkTn7KgVJtOtkU7pFKHjKGgmKESIM3Vbz4jeeIAQnSiYuuW0OUSH0GPW4M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=O4A111iO; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="O4A111iO" Received: from pps.filterd (m0356516.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 67SJVoNj2906535; Fri, 28 Aug 2026 21:47:00 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=JH/DxXvBVISBBL8hG Cj93r5pTsqPnB2I6ZLIVyPLFMA=; b=O4A111iOflip0RzTRJslV2TtDB4SEM6A9 gM7CPEI1AIEVoz+i3ybjL+4wB6TpltdpEcLqLACbs1uKaSKL9RMR2fv0qE7FruIk Gdf304UVb2zPjv4jAFAQUrXHu3TluaZ8fpPSVOpL40ixjfacOmCFxHuwJ8HiT856 +O+L2m+LldQXxWWEGnIMqY4YFm0VUBAfQnCEOUWF21v8IUIxAx+j32BClXCFuZ7V epD1TNqQzNJWqs9/Z+Pr5XRNJFc5b3/+SWb8ZfQcxri4VDaVkTpH5aQWY+D25rn2 idTPhpjTabqRn7mtAhWZWQPZucovt717yOZUMprC3VuxM7C3H97CA== Received: from ppma23.wdc07v.mail.ibm.com (5d.69.3da9.ip4.static.sl-reverse.com [169.61.105.93]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4g716jfc44-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 28 Aug 2026 21:47:00 +0000 (GMT) Received: from pps.filterd (ppma23.wdc07v.mail.ibm.com [127.0.0.1]) by ppma23.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 67SLfHiM012606; Fri, 28 Aug 2026 21:46:59 GMT Received: from smtprelay04.wdc07v.mail.ibm.com ([172.16.1.71]) by ppma23.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4g7qkhru3a-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 28 Aug 2026 21:46:59 +0000 (GMT) Received: from smtpav02.wdc07v.mail.ibm.com (smtpav02.wdc07v.mail.ibm.com [10.39.53.229]) by smtprelay04.wdc07v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 67SLkwOZ17039980 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 28 Aug 2026 21:46:58 GMT Received: from smtpav02.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 59ECA58059; Fri, 28 Aug 2026 21:46:58 +0000 (GMT) Received: from smtpav02.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 8E03C58058; Fri, 28 Aug 2026 21:46:56 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.33.14]) by smtpav02.wdc07v.mail.ibm.com (Postfix) with ESMTP; Fri, 28 Aug 2026 21:46:56 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, stable@vger.kernel.org Subject: [PATCH v4 1/4] s390/vfio-ap: Fix leak of pinned NIB and registered NISC in vfio_ap_irq_enable/disable() Date: Fri, 28 Aug 2026 17:46:50 -0400 Message-ID: <20260828214653.1087009-2-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260828214653.1087009-1-akrowiak@linux.ibm.com> References: <20260828214653.1087009-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Spam-Info: AW1haW4tMjYwODI4MDE4NyBTYWx0ZWRfX5rJ46Vey8ir+ ZVOF2xaQALI/9gSiZUl/9Vq9yJdviNaWRL9EPVF+xECSRkUkA9zW6kdaWx0swo3SxzpbR186NLH q37WSp4YTAwElNm86PMqoQ5LlWsmfoI= X-Proofpoint-GUID: t35JmfMFhvwedgJeAz-ODjdazHvKVenM X-Proofpoint-ORIG-GUID: t35JmfMFhvwedgJeAz-ODjdazHvKVenM X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODI4MDE4NyBTYWx0ZWRfX2JkLUbZWfljP SyXBeWn1lzMbxtu65MDhI/kWT0bqHxRqs3wBKik/yUeDLfwVaFU9ekt0Md2ZJcZIxlxoo7dWNR5 5nXJW/p4cf1Hx4qqiMOE8zeMzQZcUu+S+nfOnAma9m0tmpJN4DrIS0I8uhpf8s/ZqQ7AWJT4rdy zg7fGrsESBN+KQPgjQMSr/R+exOWkfjiGQlqjFwJ51/0o7eg8Ia8Bi0RQuGKCTehTxcRMQEmnH2 OBr0eYdfuzWMWPnaTgmHqwbrotVkR+IsJgV261sTdSIDkm/TpHqmF6b9upaXwbEuhM3Hfwc/MOo SxYgz2Qf/zi3EXIM9s1MY4G7dZtJJe6C00mkoj7il5Jx5Ez7svZhiibZ2VcMczwjiv3oOebuSRR lVPqYSeyl1LuDDXevc3q2TgdCtvFNu3RRWCEFttt3053/TU0qkUTcQ0+1chyzJeNSpgYSMrS7r3 mzpQGsgg+1Oh0cbGo0w== X-Authority-Analysis: v=2.4 cv=H7brBeYi c=1 sm=1 tr=0 ts=6a9201d4 cx=c_pps a=3Bg1Hr4SwmMryq2xdFQyZA==:117 a=3Bg1Hr4SwmMryq2xdFQyZA==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=Y2IxJ9c9Rs8Kov3niI8_:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=mhE2La7kI1ux2J5OmX4A:9 X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-28_06,2026-08-27_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 suspectscore=0 malwarescore=0 lowpriorityscore=0 impostorscore=0 spamscore=0 bulkscore=0 adultscore=0 priorityscore=1501 clxscore=1015 phishscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608280187 Content-Type: text/plain; charset="utf-8" The vfio_ap_irq_enable() and vfio_ap_disable() functions execute the PQAP(AQIC) instructions to enable/disable interrupts for an AP queue. A switch statement is used to examine the status response code returned from the instruction to determine whether it succeeded or failed and react accordingly. vfio_ap_irq_enable() ~~~~~~~~~~~~~~~~~~~~ For the default case, the vfio_ap_irq_disable function is invoked to disable interrupts for the queue and clean up the AQIC resources (i.e., unpin the NIB and unregister the NISC) that are stored with the vfio_ap_queue object. There are a number of problems with this: 1. Neither the q->saved_iova nor q->saved_isc has been set for the current AQIC call, so the AQIC resources - assuming those values have been previously set - will be the NIB and NISC resources from a prior call; the NIB and NISC from the current call are therefore leaked. 2. Interrupts may never have been enabled. Sending a disable instruction to a queue that the hardware just told you is in a bad state (CHECKSTOPPED, DECONFIGURED, Q_NOT_AVAIL) is at best wasted work and at worst generates a further WARN_ONCE from inside vfio_ap_irq_disable's own default. 3. The hardware just rejected the new ap_aqic() enable attempt with an unexpected status. Disabling a previously-working IRQ config - assuming that is even possible - as a reaction to a failed enable attempt does not make sense; it is actively destructive, tearing down something that was working for no valid reason. The fix is to unregister the NISC and an unpin the NIB used in the AQIC call in the default case of the switch statement and leave the AQIC resources stored with the vfio_ap_queue object alone. Another problem with vfio_ap_irq_enable() is that the old NIB stored in q->saved_iova from a previous AQIC are immediately freed - assuming they are stored - as soon as the AQIC response status returns AP_RESPONSE_NORMAL. The problem with this is, AQIC is an asynchronous operation; the only way to tell if it has completed is to check the I-bit (7) in the status returned from AQIC. This bit indicates whether interrupts are enabled (1) or disabled (0). The fix for this is to wait a specified period of time until the I-bit is set, similar to vfio_ap_wait_for_irqclear() - waits for I-bit=3D=3D0 - which is called from vfio_ap_irq_disable() to verify the operation has completed. If verification of I-bit =3D=3D 1 occurs within a specified period of time, freeing the old NIB (q->saved_iova) because once the I-bit is set, the new NIB is made available for queue interrupts and any old NIB will no longer be used. If the verification times out, then the new NIB will be freed and the old NIB will be allowed to leak which is preferable because the queue might still have in-flight DMA writes directed to the old NIB which would result in a use-after-free kernel crash. vfio_ap_irq_disable() ~~~~~~~~~~~~~~~~~~~~~ There are two problems with the way this function handles the response code returned from the PQAP(AQIC) instruction: 1. For response codes AP_RESPONSE_NORMAL or AP_RESPONSE_OTHERWISE_CHANGED, a call is made to vfio_ap_wait_for_irqclear() which waits for the I-bit (7) - indicates whether interrupts are enabled (1) or disabled (0) - to be cleared. That function does not return anything, so there is no way to determine whether it succeeded or not. The vfio_ap_irq_disable() function then frees the AQIC resources. This is a problem because the hardware may still write to the NIB resulting in a use-after-free kernel crash. The fix for this is to add a boolean return code from vfio_ap_wait_for_irqclear(). This will be checked in vfio_ap_irq_disable() and if clearing of the IR bit could not be verified, the AQIC resources will be allowed to leak. This is preferable to a kernel crash. 2. For response code AP_RESPONSE_INVALID_ADDRESS - indicates the NIB address passed to PQAP(AQIC) is not valid - as well as the default case, the vfio_ap_irq_disable() frees the AQIC resources. Since the AQIC disable was rejected, the IRQ is still enabled and the hardware still holds the NIB address, so freeing the NIB could result in a use-after-free kernel crash. The fix for this is to allow the AQIC resources to be leaked. This is preferable to a kernel crash. Fixes: ec89b55e3bce7 ("s390: ap: implement PAPQ AQIC interception in kernel= ") Cc: stable@vger.kernel.org Signed-off-by: Anthony Krowiak --- drivers/s390/crypto/vfio_ap_ops.c | 152 +++++++++++++++++++++++++----- 1 file changed, 129 insertions(+), 23 deletions(-) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index 940c0ff668be..24e93fb7f81a 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -30,6 +30,9 @@ #define AP_QUEUE_UNASSIGNED "unassigned" #define AP_QUEUE_IN_USE "in use" =20 +#define AP_IRQ_DISABLED 0 +#define AP_IRQ_ENABLED 1 + #define AP_RESET_INTERVAL 20 /* Reset sleep interval (20ms) */ =20 static int vfio_ap_mdev_reset_queues(struct ap_matrix_mdev *matrix_mdev); @@ -226,16 +229,27 @@ static struct vfio_ap_queue *vfio_ap_mdev_get_queue( } =20 /** - * vfio_ap_wait_for_irqclear - clears the IR bit or gives up after 5 tries - * @apqn: The AP Queue number + * vfio_ap_wait_for_irqstate - wait for the IR bit to reach the requested = state * - * Checks the IRQ bit for the status of this APQN using ap_tapq. - * Returns if the ap_tapq function succeeded and the bit is clear. - * Returns if ap_tapq function failed with invalid, deconfigured or - * checkstopped AP. - * Otherwise retries up to 5 times after waiting 20ms. + * @apqn: the APQN of the queue + * @ir: the expected state of the IR bit: AP_IRQ_DISABLED or AP_IRQ_ENABLED + * + * Repeatedly polls the AP queue status via PQAP(TAPQ) every 20ms until th= e IR + * bit matches @ir, the queue becomes non-operational, or 5 retries are + * exhausted. + * + * Because PQAP(AQIC) initiates an asynchronous process, a condition-code 0 + * completion does not guarantee the IR bit has reached the requested stat= e. + * The caller must use this function to confirm the state before proceedin= g. + * + * Return: + * - true if the IR bit matches @ir, or the AP is non-operational (in which + * case no further interrupts can be generated) + * + * - false if the IR bit still does not match @ir after all retries are + * exhausted */ -static void vfio_ap_wait_for_irqclear(int apqn) +static bool vfio_ap_wait_for_irqstate(int apqn, int ir) { struct ap_queue_status status; int retry =3D 5; @@ -245,8 +259,8 @@ static void vfio_ap_wait_for_irqclear(int apqn) switch (status.response_code) { case AP_RESPONSE_NORMAL: case AP_RESPONSE_RESET_IN_PROGRESS: - if (!status.irq_enabled) - return; + if (status.irq_enabled =3D=3D ir) + return true; fallthrough; case AP_RESPONSE_BUSY: msleep(20); @@ -257,12 +271,15 @@ static void vfio_ap_wait_for_irqclear(int apqn) default: WARN_ONCE(1, "%s: tapq rc %02x: %04x\n", __func__, status.response_code, apqn); - return; + return true; } } while (--retry); =20 - WARN_ONCE(1, "%s: tapq rc %02x: %04x could not clear IR bit\n", - __func__, status.response_code, apqn); + WARN_ONCE(1, "%s: tapq rc %02x: timed out waiting for interrupts %s for %= 02x.%04x\n", + __func__, status.response_code, + ir ? "enabled" : "disabled", + AP_QID_CARD(apqn), AP_QID_QUEUE(apqn)); + return false; } =20 /** @@ -317,8 +334,30 @@ static struct ap_queue_status vfio_ap_irq_disable(stru= ct vfio_ap_queue *q) switch (status.response_code) { case AP_RESPONSE_OTHERWISE_CHANGED: case AP_RESPONSE_NORMAL: - vfio_ap_wait_for_irqclear(q->apqn); - goto end_free; + /* + * AQIC disable was accepted (NORMAL), or the queue was + * already disabled or a prior async request is still + * completing (OTHERWISE_CHANGED). In both cases, we must + * wait until interrupt processing has been disabled + * before proceeding. + */ + if (vfio_ap_wait_for_irqstate(q->apqn, AP_IRQ_DISABLED)) + goto end_free; + /* + * Timed out waiting to confirm interrupts are disabled. + * If ap_aqic returned NORMAL, the guest would incorrectly + * interpret that as a successful disable and may free or + * reuse the NIB while hardware can still write to it. + * Zero the status word and set OTHERWISE_CHANGED to mimic + * what the hardware does for that response code. This + * signals to the guest that the reset operation did not + * complete. + */ + if (status.response_code =3D=3D AP_RESPONSE_NORMAL) { + memset(&status, 0, sizeof(status)); + status.response_code =3D AP_RESPONSE_OTHERWISE_CHANGED; + } + goto end_fail; case AP_RESPONSE_RESET_IN_PROGRESS: case AP_RESPONSE_BUSY: msleep(20); @@ -326,18 +365,46 @@ static struct ap_queue_status vfio_ap_irq_disable(str= uct vfio_ap_queue *q) case AP_RESPONSE_Q_NOT_AVAIL: case AP_RESPONSE_DECONFIGURED: case AP_RESPONSE_CHECKSTOPPED: + /* AP not operational; no further interrupts possible */ + WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, + status.response_code); + goto end_free; case AP_RESPONSE_INVALID_ADDRESS: default: - /* All cases in default means AP not operational */ + /* + * The AQIC disable was rejected; IRQ is still enabled + * and the hardware still holds the NIB address. Do not + * free resources. + */ WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, status.response_code); - goto end_free; + goto end_fail; } } while (retries--); =20 WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, status.response_code); + +end_fail: + /* + * We are here either because of a failure to verify that + * interrupts have been disabled, or because the AQIC instruction + * failed to disable them. The AQIC resources - the pinned NIB page + * and the registered guest ISC - cannot be freed here. The hardware + * may still write to the NIB; freeing the pinned page would result + * in a use-after-free kernel crash. The resources will therefore be + * leaked. This is preferable to a use-after-free. + */ + return status; + end_free: + /* + * This label is reached because the queue was successfully disabled, + * or because the queue is not operational, in which case interrupts + * can not be processed, so free the AQIC resources - the pinned NIB + * page and the registered guest ISC - used to enable interrupts + * so they will not be leaked. + */ vfio_ap_free_aqic_resources(q); return status; } @@ -432,6 +499,7 @@ static struct ap_queue_status vfio_ap_irq_enable(struct= vfio_ap_queue *q, struct kvm *kvm; phys_addr_t h_nib; dma_addr_t nib; + char *msg; int ret; =20 /* Verify that the notification indicator byte address is valid */ @@ -489,13 +557,49 @@ static struct ap_queue_status vfio_ap_irq_enable(stru= ct vfio_ap_queue *q, status =3D ap_aqic(q->apqn, aqic_gisa, h_nib); switch (status.response_code) { case AP_RESPONSE_NORMAL: - /* See if we did clear older IRQ configuration */ + /* + * AQIC initiates an asynchronous process; however, AP_RESPONSE_NORMAL + * does not guarantee the IR bit is set yet. Wait to confirm before + * committing the new NIB and freeing the old resources. + */ + if (!vfio_ap_wait_for_irqstate(q->apqn, AP_IRQ_ENABLED)) { + /* + * Timed out: the hardware may not have accepted the new + * NIB. Clean up the new resources and return + * OTHERWISE_CHANGED to signal the guest to retry. + */ + ret =3D kvm_s390_gisc_unregister(kvm, isc); + if (ret) { + msg =3D "%s: kvm_s390_gisc_unregister: rc=3D%d isc=3D%d, apqn=3D%#04x\= n"; + VFIO_AP_DBF_WARN(msg, __func__, ret, isc, q->apqn); + } + vfio_unpin_pages(&q->matrix_mdev->vdev, nib, 1); + memset(&status, 0, sizeof(status)); + status.response_code =3D AP_RESPONSE_OTHERWISE_CHANGED; + break; + } + /* + * IR bit confirmed set. AQIC enable initiates an asynchronous + * process; a CC=3D0 completion does not guarantee the process is + * done. If the queue was already enabled for interrupts, the old + * NIB (q->saved_iova) must not be freed until IR=3D1 is observed, + * because until then the hardware has not completed its + * transition to the new NIB. Now that IR=3D1 is confirmed, the + * new NIB is in use for this queue, so no interrupts can be made + * pending via any previously-registered NIB and the old + * resources can be safely freed. + */ vfio_ap_free_aqic_resources(q); q->saved_iova =3D nib; q->saved_isc =3D isc; break; case AP_RESPONSE_OTHERWISE_CHANGED: - /* We could not modify IRQ settings: clear new configuration */ + /* + * IRQ control is already set as requested or a prior async + * request has not yet completed; in either case, this response + * comes with CC=3D3 indicating the new NIB and ISC were not accepted by + * the hardware, so clean them up. + */ ret =3D kvm_s390_gisc_unregister(kvm, isc); if (ret) VFIO_AP_DBF_WARN("%s: kvm_s390_gisc_unregister: rc=3D%d isc=3D%d, apqn= =3D%#04x\n", @@ -503,9 +607,12 @@ static struct ap_queue_status vfio_ap_irq_enable(struc= t vfio_ap_queue *q, vfio_unpin_pages(&q->matrix_mdev->vdev, nib, 1); break; default: - pr_warn("%s: apqn %04x: response: %02x\n", __func__, q->apqn, - status.response_code); - vfio_ap_irq_disable(q); + /* We could not modify IRQ settings: clear new configuration */ + ret =3D kvm_s390_gisc_unregister(kvm, isc); + if (ret) + VFIO_AP_DBF_WARN("%s: kvm_s390_gisc_unregister: rc=3D%d isc=3D%d, apqn= =3D%#04x\n", + __func__, ret, isc, q->apqn); + vfio_unpin_pages(&q->matrix_mdev->vdev, nib, 1); break; } =20 @@ -635,7 +742,6 @@ static int handle_pqap(struct kvm_vcpu *vcpu) } =20 status =3D vcpu->run->s.regs.gprs[1]; - /* If IR bit(16) is set we enable the interrupt */ if ((status >> (63 - 16)) & 0x01) qstatus =3D vfio_ap_irq_enable(q, status & 0x07, vcpu); --=20 2.53.0 From nobody Sat Sep 26 21:59:52 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C265C395ACE; Fri, 28 Aug 2026 21:47:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787953628; cv=none; b=Z30xybe4XPyo97Mthm+p76BxUMHTbsjpUIcINTbDXCEznDZ7xDkWvSZPSAke/Zse83pK3rNEZQMPbHXWQrJf5968W3Oi0DlQF6a3lVE9/nVHtXmUVG87d7YhcjZSCTfEYsemWqLm/F/XBWaLmBhnNK+iaEX0U19uqI/RLLG4DVA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787953628; c=relaxed/simple; bh=9n9rN6YNjpDGUGOe2EgdWalJX6gfUJyOI8r0BUlgG0E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rHDE+NjhXJgjp02Q/axlquzaQJyHtib+dnvaRJRSMtdSpZScPr0wrf8xCtdZDQ6ajgiFwcxdjHH1OPMC6BwkGArlAV2g2/rWHW4cr1xgbk1hg7K12vA3DYLY4831FFnJijm9y+/S9gru0N3vUXp+RjH2r9PsB9JXPXLMbWHd4AI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=A9WpINGz; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="A9WpINGz" Received: from pps.filterd (m0353725.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 67SJVgiQ1887422; Fri, 28 Aug 2026 21:47:02 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=32AsIK+VwOzKpR/C0 2N8kqs1dqRcHjDNDTQ3mGrXoos=; b=A9WpINGzamOTZnaXpY8LK3YAKZ9MMiy2c pZNOl8+N5E04tzqSWVjnqc/bzkp6tZPegdwWWebptUfZypWeDm8VCgORtrAFikoN sjPV5givbHQpalDaCz00er3qgmELp8w2JO4NYEGJd1kdv4vbaej9NKkw611hqiqN qtOqIa75GYnweZHFvCpibidfEjJOjsCSzaOoGfPIvijcKckUacIGAhT5qPBXiasa b6sCV3GcKPxZ0YdVGM2fg1iREMrNUiz/Ul5Uxygrv1tGxDkbZXNTdNd8AWzZncNh tiAsAkunz78RsFUhhd829thUnBCZfW7vg6wXfeaLwJjxHv7ytUvog== Received: from ppma21.wdc07v.mail.ibm.com (5b.69.3da9.ip4.static.sl-reverse.com [169.61.105.91]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4g726f76mh-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 28 Aug 2026 21:47:02 +0000 (GMT) Received: from pps.filterd (ppma21.wdc07v.mail.ibm.com [127.0.0.1]) by ppma21.wdc07v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 67SLfIIB031302; Fri, 28 Aug 2026 21:47:01 GMT Received: from smtprelay06.wdc07v.mail.ibm.com ([172.16.1.73]) by ppma21.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4g7q3kgvs5-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 28 Aug 2026 21:47:01 +0000 (GMT) Received: from smtpav02.wdc07v.mail.ibm.com (smtpav02.wdc07v.mail.ibm.com [10.39.53.229]) by smtprelay06.wdc07v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 67SLl0H441353572 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 28 Aug 2026 21:47:00 GMT Received: from smtpav02.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 3EACB5805C; Fri, 28 Aug 2026 21:47:00 +0000 (GMT) Received: from smtpav02.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 86AD858058; Fri, 28 Aug 2026 21:46:58 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.33.14]) by smtpav02.wdc07v.mail.ibm.com (Postfix) with ESMTP; Fri, 28 Aug 2026 21:46:58 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, stable@vger.kernel.org Subject: [PATCH v4 2/4] s390/vfio-ap: Fix failure to release IRQ notification eventfd contexts Date: Fri, 28 Aug 2026 17:46:51 -0400 Message-ID: <20260828214653.1087009-3-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260828214653.1087009-1-akrowiak@linux.ibm.com> References: <20260828214653.1087009-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Authority-Analysis: v=2.4 cv=TfimcxQh c=1 sm=1 tr=0 ts=6a9201d6 cx=c_pps a=GFwsV6G8L6GxiO2Y/PsHdQ==:117 a=GFwsV6G8L6GxiO2Y/PsHdQ==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=V8glGbnc2Ofi9Qvn3v5h:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=fxeDLeT5AZ2qCvnLJ4gA:9 X-Proofpoint-Spam-Info: AW1haW4tMjYwODI4MDE4NyBTYWx0ZWRfX90ibCYhOSCa4 92LLCwKxWblKmU8swT0UD2L/pxLsGU7UI10y/sig/pBlLlVMOpvc1Dh6bXUPf4CE4x0CutZy4Zk jionJkgKPXTjB14WQj9z7HDBujXADVU= X-Proofpoint-GUID: Xws86833JIRxGks_zPP1_Z_-rYpJ1uVd X-Proofpoint-ORIG-GUID: Xws86833JIRxGks_zPP1_Z_-rYpJ1uVd X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODI4MDE4NyBTYWx0ZWRfX/sy7e0+yYQFj BGNGyZHyuGw6lCl3ow+RcMQEXAJCn+WHw3Sgjyl0q/Bn2nWent9AJX+oThQVJoxxgJOwJzDyGm3 vvAQdrdjbmSZXAYc1uvh+MbQITv/Jj9Xz/T9Ayu5/BDxMNRFQHpC+27Duh8DI5b66paZ9nDJxoz jCd5meCqYlbgfJYjyklw5LGmMjM8e3i4mO3Df1iCSiedbx+oiTMuJRANqYkJmvPlk5pCG6IWsdw mz2mnFuu1tR8jj7uBXMjm21mFq548l0xgslyMkDwd1j4mYTSoXbip0qUVKJ5uYM+RPED+XBfy6k 5myYogPpmy9BXw0tON//mHr9We/RLRqr+9MUi4HGAKBIesax5qdw+MhXqcmMQYLWi4cTCkK+7/d zMv2ksYxNxXWYV7QnsW2A4bkBRJ4Yabb8xL+HxFK4I81OmpINtwgx6qHgLB55kWt7bSEY2+PdwN 4BLe6pzTJzbY9+PltRQ== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-28_06,2026-08-27_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 suspectscore=0 spamscore=0 bulkscore=0 malwarescore=0 priorityscore=1501 lowpriorityscore=0 clxscore=1015 phishscore=0 impostorscore=0 adultscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608280187 Content-Type: text/plain; charset="utf-8" When userspace registers IRQ notification eventfds via the VFIO_DEVICE_SET_IRQS ioctl, vfio_ap_set_request_irq() and vfio_ap_set_cfg_change_irq() each call eventfd_ctx_fdget(), which takes a reference on the eventfd_ctx and stores it in matrix_mdev->req_trigger and matrix_mdev->cfg_chg_trigger respectively. These references are dropped only when userspace explicitly replaces or clears them via a subsequent SET_IRQS call. If the device is closed without that explicit teardown - because the guest exits, the VM process crashes, or the device file is simply closed - neither vfio_ap_mdev_close_device() nor the remove path releases these references. The eventfd_ctx backing objects and their associated file references therefore leak for the lifetime of the kernel. Fix this by introducing vfio_ap_mdev_release_eventfds() and calling it from vfio_ap_mdev_close_device() after vfio_ap_mdev_unset_kvm(). The VFIO core guarantees that close_device is called before vfio_unregister_group_dev() returns in the remove path, so fixing close_device is sufficient to cover both teardown paths. Note: ~~~~ The matrix_dev->mdevs lock must be held during the call to vfio_ap_mdev_release_eventfds(). There is a small window between the calls to vfio_ap_mdev_unset_kvm() which gets and releases the update locks and the acquisition of the matrix_dev->mdevs_lock mutex during which it is possible - although highly unlikely during normal operation - whereby a concurrent SET_IRQS call can get in. Taking matrix_dev->mdevs_lock around vfio_ap_mdev_release_eventfds() is sufficient to make this race-free. The SET_IRQS ioctl path writes req_trigger and cfg_chg_trigger only from vfio_ap_mdev_ioctl(), which holds mdevs_lock for its entire duration and always calls eventfd_ctx_put() on the previous value before storing the new one. Any number of concurrent SET_IRQS calls during the window between vfio_ap_mdev_unset_kvm() and the acquisition of mdevs_lock are therefore safe: each ioctl invocation puts the reference it found and installs a new one, leaving exactly one live reference in the field when it releases the lock. When release_eventfds subsequently acquires mdevs_lock it finds that single surviving reference and puts it. Conversely, a SET_IRQS call that loses the race and blocks on mdevs_lock will find the field NULL after release_eventfds finishes, take ownership of the reference it just created, and install it into a field that will never be read again - a transient leak. To close that final case, callers must ensure no new SET_IRQS ioctls can be issued after close_device() is called, which the VFIO core guarantees by releasing the device file before invoking close_device(). Fixes: bf48961f6f48e ("s390/vfio-ap: realize the VFIO_DEVICE_SET_IRQS ioctl= ") Cc: stable@vger.kernel.org Signed-off-by: Anthony Krowiak Reviewed-by: Matthew Rosato --- drivers/s390/crypto/vfio_ap_ops.c | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index 24e93fb7f81a..363d9e53e249 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -2167,12 +2167,28 @@ static int vfio_ap_mdev_open_device(struct vfio_dev= ice *vdev) return vfio_ap_mdev_set_kvm(matrix_mdev, vdev->kvm); } =20 +static void vfio_ap_mdev_release_eventfds(struct ap_matrix_mdev *matrix_md= ev) +{ + if (matrix_mdev->req_trigger) { + eventfd_ctx_put(matrix_mdev->req_trigger); + matrix_mdev->req_trigger =3D NULL; + } + if (matrix_mdev->cfg_chg_trigger) { + eventfd_ctx_put(matrix_mdev->cfg_chg_trigger); + matrix_mdev->cfg_chg_trigger =3D NULL; + } +} + static void vfio_ap_mdev_close_device(struct vfio_device *vdev) { struct ap_matrix_mdev *matrix_mdev =3D container_of(vdev, struct ap_matrix_mdev, vdev); =20 vfio_ap_mdev_unset_kvm(matrix_mdev); + + mutex_lock(&matrix_dev->mdevs_lock); + vfio_ap_mdev_release_eventfds(matrix_mdev); + mutex_unlock(&matrix_dev->mdevs_lock); } =20 static void vfio_ap_mdev_request(struct vfio_device *vdev, unsigned int co= unt) --=20 2.53.0 From nobody Sat Sep 26 21:59:52 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A4D8339A073; Fri, 28 Aug 2026 21:47:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787953631; cv=none; b=kcklftwMX27msiBfPmNrCzF6MFhelmrLrGXTEPu8AoJfgaF3eo6Wy5BjMJwTmuHDyoit12WLW0FEG12UiKjXHIdrr9j/Yom1HWG4pM223rsRtnO23tZ3uZv91VUmtLIZ3DNM4+VsJrB93g8TulPBj2+uPJCmLTs0y42EPOjI6fk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787953631; c=relaxed/simple; bh=5IP86hzRbbjhxxfQNEOYVuSqYDhdRiUR2Bns/Bl/xr4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=bHQwm3BVZEachhr/N24YnvPFzRYS1UDlwZ4EmDraYfBifXSh/WxcXevhz2VptrnvL/iRmASHK/2QVSVW4g3pjUo8pPByqODUDCVXk1c392ibBdPi/CeRltSC66WkZWM03mZpJKqwWXTnjgo6XO1FgxPS+1ouv6vVDRyf8clTSeo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=irKkpzT/; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="irKkpzT/" Received: from pps.filterd (m0356516.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 67SJVkGM2906405; Fri, 28 Aug 2026 21:47:04 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=6PvKm4sscKTZCn5rw plFou5k5vTrXQv6GtfdgHswfLE=; b=irKkpzT/nkU8PPQrKFl/O3fu15V0Ol+fG oUHiiQweq/iebdGx65RJMEUVzoVP9ujZa1lZEeD69fX1o+PVE4YToh+rQ0cJijHx FwcSp1BX4B24YTO+hsrHV7zMsLEtzq83FpuZg1a4GJtFvBZMuFCaFjsL6G6XdfkB d5nyQoT01Ng20Yd34ebtlBsQmpgla7GYMH6YB80w/sEUPwN/MNT8ABWwU79LH0Yt u9oz7WENB/LP3pIe87txOqfD/PA5Kb87ApzEs5lfDs0gwHeugFliKelOn6Vg0ZN1 Gmb4snIk+ka71EKSkmJb7r/SRisshf1KJ+GRTj9I6UNh3vCmOXtTA== Received: from ppma13.dal12v.mail.ibm.com (dd.9e.1632.ip4.static.sl-reverse.com [50.22.158.221]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4g716jfc4p-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 28 Aug 2026 21:47:04 +0000 (GMT) Received: from pps.filterd (ppma13.dal12v.mail.ibm.com [127.0.0.1]) by ppma13.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 67SLfGHi003244; Fri, 28 Aug 2026 21:47:03 GMT Received: from smtprelay03.dal12v.mail.ibm.com ([172.16.1.5]) by ppma13.dal12v.mail.ibm.com (PPS) with ESMTPS id 4g7rah0s01-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 28 Aug 2026 21:47:03 +0000 (GMT) Received: from smtpav02.wdc07v.mail.ibm.com (smtpav02.wdc07v.mail.ibm.com [10.39.53.229]) by smtprelay03.dal12v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 67SLl2k847972692 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 28 Aug 2026 21:47:02 GMT Received: from smtpav02.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 2270E58058; Fri, 28 Aug 2026 21:47:02 +0000 (GMT) Received: from smtpav02.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 6A9435805C; Fri, 28 Aug 2026 21:47:00 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.33.14]) by smtpav02.wdc07v.mail.ibm.com (Postfix) with ESMTP; Fri, 28 Aug 2026 21:47:00 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, stable@vger.kernel.org Subject: [PATCH v4 3/4] s390/vfio-ap: Fix unbounded loop in apq_reset_check() Date: Fri, 28 Aug 2026 17:46:52 -0400 Message-ID: <20260828214653.1087009-4-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260828214653.1087009-1-akrowiak@linux.ibm.com> References: <20260828214653.1087009-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Spam-Info: AW1haW4tMjYwODI4MDE4NyBTYWx0ZWRfXz5/qKkvooyEf j1NeskfGAIK5vS1J9OEWMvjHUOFdupzG0hVe9ADmv/doO4VKyxjb0drPsUfH2JE6/2YiR/OvpsF hVrzpAjTIR+ZQoOrZqmC8XBpr3a4VO0= X-Proofpoint-GUID: nNPGGOFxVeArsPRkcw1rfjp7axZFfYku X-Proofpoint-ORIG-GUID: nNPGGOFxVeArsPRkcw1rfjp7axZFfYku X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODI4MDE4NyBTYWx0ZWRfX5vBEwZX9zOT+ qn3c5rLMuR4Uqa4gkR4NcW0NsZO9nKDsGIAH9gNPnz4g5/VX/tGRlQBhSufW65ZWpcoSvfj/koJ O8r6ErsS18j7VP+e3KeqM9JRVuEM6wNOWIlJLjRhwif02H0+oauTLRgb/MauPBz3LuKxDsBlQP0 GYNnxGA+87ktUTSfuRdJxFG/i5wgycbIW1kzq1NrksZTO+zpk8Q5bQvUz7Ll8GXWNbb6q1lsZNV 5ki/sq1DQEC5+WqJ4ZCybbG2Z8c9QbMjzE5WB3JRzf4s1D5U23VpTXo09OWonA8BJTiXO0H1mSG 6xbIc/NdwxGskrrltSlBlLdh3hEA2/kMTIf4fas0exDquwBHwle85hmQCL0RuCuG0LX6ElV23Uj z5wP+iHJK5zy6YRK3EDephVD3vGw+Bu+Q8cWexjO7sHDqmMCO1MKFTyW7a/R2Q17MseWpftJRe0 WMdYlkvSXn1ER6FUA3g== X-Authority-Analysis: v=2.4 cv=H7brBeYi c=1 sm=1 tr=0 ts=6a9201d8 cx=c_pps a=AfN7/Ok6k8XGzOShvHwTGQ==:117 a=AfN7/Ok6k8XGzOShvHwTGQ==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=Y2IxJ9c9Rs8Kov3niI8_:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=QSCR_LEHqxABJDgvvGYA:9 X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-28_06,2026-08-27_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 suspectscore=0 malwarescore=0 lowpriorityscore=0 impostorscore=0 spamscore=0 bulkscore=0 adultscore=0 priorityscore=1501 clxscore=1015 phishscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608280187 Content-Type: text/plain; charset="utf-8" The apq_reset_check() worker polls ap_tapq() in a while(true) loop waiting for a queue reset to complete. When ap_tapq() returns AP_RESPONSE_BUSY or AP_RESPONSE_RESET_IN_PROGRESS, apq_status_check() returns -EBUSY and the loop continues after sleeping AP_RESET_MAX_WAIT (20ms). There is no upper bound on how many times the loop iterates, so if the hardware continuously returns a busy response the worker runs indefinitely. This is particularly harmful because several callers of vfio_ap_reset_queue() - such as vfio_ap_mdev_reset_queues(), vfio_ap_mdev_reset_qlist() and vfio_ap_mdev_remove_queue - call flush_work() on the queue's reset_work while holding one or more of the global matrix_dev locks (guests_lock, mdevs_lock) or the KVM lock. An indefinitely spinning worker permanently blocks access to all ap_matrix_mdev objects which could hang other guests that are using them. Fix this by introducing AP_RESET_MAX_WAIT (2000ms) and breaking out of the poll loop when elapsed time reaches that threshold. On timeout, if the apq_reset_check() on that last loop has determined that the reset has completed or that the queue is not operational, the AQIC resources associated with this queue - the pinned page containing the NIB and the registered guest ISC - will be cleared because a successful reset disables interrupts nor can interrupts be signaled from a non-operational queue. If the apq_reset_check() did not verify completion of the reset, the AQIC resources associated with this queue cannot be freed because the NIB is the active DMA target for AP interrupt delivery until the reset completes; freeing the pinned page while the hardware may still write to it would result in a use-after-free kernel crash. If the reset eventually completes, interrupts will be terminated, but the pinned NIB page and ISC registration will be leaked. This is preferable to either a use-after-free kernel crash or waiting indefinitely and blocking access to all mdevs, hanging the guests to which they are attached. Note that on timeout, q->reset_status will hold the status from the most recent reset operation so that callers inspecting q->reset_status.response_code after flush_work() will see the value and can return an appropriate return code. Fixes: dd174833e44e ("s390/vfio-ap: remove upper limit on wait for queue re= set to complete") Cc: stable@vger.kernel.org Signed-off-by: Anthony Krowiak --- drivers/s390/crypto/vfio_ap_ops.c | 40 ++++++++++++++++++++++++++++--- 1 file changed, 37 insertions(+), 3 deletions(-) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index 363d9e53e249..3f5b012be450 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -34,6 +34,7 @@ #define AP_IRQ_ENABLED 1 =20 #define AP_RESET_INTERVAL 20 /* Reset sleep interval (20ms) */ +#define AP_RESET_MAX_WAIT 2000 /* Maximum wait for reset (2000ms) */ =20 static int vfio_ap_mdev_reset_queues(struct ap_matrix_mdev *matrix_mdev); static int vfio_ap_mdev_reset_qlist(struct list_head *qlist); @@ -2067,6 +2068,37 @@ static void apq_reset_check(struct work_struct *rese= t_work) ret =3D apq_status_check(q->apqn, &status); if (ret =3D=3D -EIO) return; + if (elapsed >=3D AP_RESET_MAX_WAIT) { + /* + * If the status check determined that the reset completed + * successfully or the queue is not operational, clean up + * the AQIC resources because queue reset disables + * interrupts, or because interrupts are not possible on a + * non-operational queue. + */ + if (!ret) + goto done; + /* + * Timed out without being able to verify reset completed. + * + * The AQIC resources associated with this queue - the pinned page + * containing the NIB and the registered guest ISC - cannot be freed + * here. The NIB is the active DMA target for AP interrupt delivery + * until the reset completes; freeing the pinned page while the + * hardware may still write to it would result in a use-after-free + * kernel crash. + * + * If the reset eventually completes, interrupts will be terminated + * and the pinned NIB page and ISC registration will be leaked. This + * is preferable to either a use-after-free or waiting indefinitely: + * the caller of apq_reset_check() holds mdevs_lock while flush_work() + * blocks holds the matrix_dev->mdevs_lock mutex, which + * serializes access to all mdev objects system-wide, so blocking + * here would stall all other guests using AP queues. + */ + + return; + } if (ret =3D=3D -EBUSY) { pr_notice_ratelimited(WAIT_MSG, elapsed, AP_QID_CARD(q->apqn), @@ -2083,11 +2115,13 @@ static void apq_reset_check(struct work_struct *res= et_work) memcpy(&q->reset_status, &status, sizeof(status)); continue; } - if (q->saved_isc !=3D VFIO_AP_ISC_INVALID) - vfio_ap_free_aqic_resources(q); - break; + goto done; } } + +done: + if (q->saved_isc !=3D VFIO_AP_ISC_INVALID) + vfio_ap_free_aqic_resources(q); } =20 static void vfio_ap_mdev_reset_queue(struct vfio_ap_queue *q) --=20 2.53.0 From nobody Sat Sep 26 21:59:52 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8529839A7FE; Fri, 28 Aug 2026 21:47:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787953632; cv=none; b=Eqwncf3b1JTFjH00vyP+jKYd9GGwDU+XqrjVtTsL0QCJFVAp7Zq8YfZKMHhgBk9Dlc2LrgZPCW8f8DLfpP+UzQFqRSj2isKwoI62UJf3UqTXW1z2fz0x3gCo7IjxcbcZhH46zpLuW2Tcnn4MmmPWJRjG37wJ++Ehk0vfYz49zKQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787953632; c=relaxed/simple; bh=Se9FKOPpwlll3rGeNoVs4JCnXUH+ThM3MGOUnYeXV2Q=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=W40hQhaGNbqFepfO1yCpvoRd7cnW70jodw3OMZfuAFPoOfhdPWHmvW4pCHUTIT2JGYVSWPgy9Ie0N1EUOvt5Xzf7Y3Gkh/JV3r71AYgKq6pfn9V2NYExtGogb0CkhDE/Ddpps2flusq+HscOnWft+V2ESlcCnvfEwQedIAYGWMw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=B6XaHAx4; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="B6XaHAx4" Received: from pps.filterd (m0356516.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 67SJVimF2906391; Fri, 28 Aug 2026 21:47:06 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=CWXPw+tf4gRxPN3GA vBZzDnxuTBqj+yfz1VUEm4lCVU=; b=B6XaHAx4D6QKDWpNqM4jQESoXMmfjNpIO xBto/QikALOLDF2w1RL/4gZI0/6JBPxhubmPdoao4znPGrcC/y4J2ftWT+eEX506 beskpCLwcph+M4VprVXXwXfDlt/2P+GUdnJOoDeKM8bM2ZuMC4yq9uvcmFofdFyy 9QtJVv19zBLJZRKUfQeIlT6SXJ5b8iA+E7Fie2UD9Y9roNApiMR27fE6au43NxXH ZyE5LnyT3wWPR+6IFT8rSJ+SkTUNdLJxyZrdkFE4Pi8YJfTLk+ujFcekX08URrAf PsT4XPK/M8+Vo1i+Jgb/j2QpESlXsiDrj2r6Gctnue9TH0pCP9n/g== Received: from ppma13.dal12v.mail.ibm.com (dd.9e.1632.ip4.static.sl-reverse.com [50.22.158.221]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4g716jfc4s-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 28 Aug 2026 21:47:05 +0000 (GMT) Received: from pps.filterd (ppma13.dal12v.mail.ibm.com [127.0.0.1]) by ppma13.dal12v.mail.ibm.com (8.18.1.7/8.18.1.7) with ESMTP id 67SLfGHA003236; Fri, 28 Aug 2026 21:47:05 GMT Received: from smtprelay05.dal12v.mail.ibm.com ([172.16.1.7]) by ppma13.dal12v.mail.ibm.com (PPS) with ESMTPS id 4g7rah0s08-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 28 Aug 2026 21:47:05 +0000 (GMT) Received: from smtpav02.wdc07v.mail.ibm.com (smtpav02.wdc07v.mail.ibm.com [10.39.53.229]) by smtprelay05.dal12v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 67SLl45M33817336 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 28 Aug 2026 21:47:04 GMT Received: from smtpav02.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id F00B05805E; Fri, 28 Aug 2026 21:47:03 +0000 (GMT) Received: from smtpav02.wdc07v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 4E8585805B; Fri, 28 Aug 2026 21:47:02 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.33.14]) by smtpav02.wdc07v.mail.ibm.com (Postfix) with ESMTP; Fri, 28 Aug 2026 21:47:02 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com Subject: [PATCH v4 4/4] s390/vfio-ap: Use AP_DOMAINS for adm_add bitmap size in vfio_ap_mdev_cfg_add() Date: Fri, 28 Aug 2026 17:46:53 -0400 Message-ID: <20260828214653.1087009-5-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260828214653.1087009-1-akrowiak@linux.ibm.com> References: <20260828214653.1087009-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-Spam-Info: AW1haW4tMjYwODI4MDE4NyBTYWx0ZWRfX683J5DmJFkA1 aLWBuRG9wYF8yHNfFAzjX8v/DbTpxy9h++O45QwNt5OSt74GMX5Qj1IczUYBkJ7PzKVP2raulXL CRs4YF+5KaORKtG2Ft2SgD47caSWawM= X-Proofpoint-GUID: UxiT4iamRjZodyVhJoTQe_iWKEgcNJvM X-Proofpoint-ORIG-GUID: UxiT4iamRjZodyVhJoTQe_iWKEgcNJvM X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODI4MDE4NyBTYWx0ZWRfX9X2RvBpZOqZZ ZlS6u71sIcofwJGa0fSAj79D4BqjQq+tt78blC+XLueUAiFs8Aqr+0I/gzObWbJi6bW7ekTL2M7 EIhEt8EiX85c7vPKya43l6rtmiPMOdliBqMZPBlnwHTqRUcqo0MjeUiLob8Pcuoko9rvMtmyBeL EH3rA9DwNdSA8RurFNBd+RiG/0OtxMcl/GCvlyKHV3UhpvFLqLliBAwys+3ULso3B5uQTzfaq7A fF6//TOvj/LlodruaRumYn39STR24Ge/aZoFTkSbXb5x+DPE+Xq45lC141nZhdxNMlE9/Q0a4vY 8YWgqcoT/VKfe3ba4L+MQIwXRv2zn3vebxfB90N4iFxDw08fZW/rTrXa7KhMoO2Mh/mhTWhG0pE rP6e2tJ+bTHnZheZ6tVL7cJHbrYUDa8gQD34Z9OAMG20iyyeHEW+m3lz2bPFEGhhXoRwVcwBjWm NvS7Gl43i96/Totv0zA== X-Authority-Analysis: v=2.4 cv=H7brBeYi c=1 sm=1 tr=0 ts=6a9201da cx=c_pps a=AfN7/Ok6k8XGzOShvHwTGQ==:117 a=AfN7/Ok6k8XGzOShvHwTGQ==:17 a=Sv0fKeRqtYgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=Y2IxJ9c9Rs8Kov3niI8_:22 a=VnNF1IyMAAAA:8 a=gReD5cTbzgwmqpVxveEA:9 X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-28_06,2026-08-27_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 suspectscore=0 malwarescore=0 lowpriorityscore=0 impostorscore=0 spamscore=0 bulkscore=0 adultscore=0 priorityscore=1501 clxscore=1015 phishscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2608280187 Content-Type: text/plain; charset="utf-8" Domain and control domain bitmaps are sized by the AP_DOMAINS constant, not AP_DEVICES. The two constants are both 256 today so there is no functional impact, but using the wrong constant is inconsistent with every operation on aqm/adm bitmaps. Use AP_DOMAINS to keep the code consistent and correct in case the two constants ever diverge. Note: This patch was submitted in response to a sashiko review comment pointing out there are other functions besides vfio_ap_mdev_cfg_add(), so there are fixes included here for those also. The subject line was kept the same since this is in v2 of this patch. Signed-off-by: Anthony Krowiak --- drivers/s390/crypto/vfio_ap_ops.c | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index 3f5b012be450..d9a8844020bd 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -1517,7 +1517,7 @@ static void vfio_ap_mdev_hot_unplug_domain(struct ap_= matrix_mdev *matrix_mdev, { DECLARE_BITMAP(apqis, AP_DOMAINS); =20 - bitmap_zero(apqis, AP_DEVICES); + bitmap_zero(apqis, AP_DOMAINS); set_bit_inv(apqi, apqis); vfio_ap_mdev_hot_unplug_domains(matrix_mdev, apqis); } @@ -2853,11 +2853,11 @@ static void vfio_ap_mdev_on_cfg_remove(struct ap_co= nfig_info *cur_config_info, do_remove |=3D bitmap_andnot(aqrem, (unsigned long *)prev_config_info->aqm, (unsigned long *)cur_config_info->aqm, - AP_DEVICES); + AP_DOMAINS); do_remove |=3D bitmap_andnot(cdrem, (unsigned long *)prev_config_info->adm, (unsigned long *)cur_config_info->adm, - AP_DEVICES); + AP_DOMAINS); =20 if (do_remove) vfio_ap_mdev_cfg_remove(aprem, aqrem, cdrem); @@ -2968,7 +2968,7 @@ static void vfio_ap_mdev_cfg_add(unsigned long *apm_a= dd, unsigned long *aqm_add, bitmap_and(matrix_mdev->aqm_add, matrix_mdev->matrix.aqm, aqm_add, AP_DOMAINS); bitmap_and(matrix_mdev->adm_add, - matrix_mdev->matrix.adm, adm_add, AP_DEVICES); + matrix_mdev->matrix.adm, adm_add, AP_DOMAINS); =20 mutex_unlock(&matrix_dev->mdevs_lock); } --=20 2.53.0