From nobody Sat Sep 26 03:50:07 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E48F5470EB4; Fri, 25 Sep 2026 12:46:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340365; cv=none; b=di6H4KpqGhDnyZYgnaWUdkDIEeWihijGXXeoDrmEh5B0FMXPmtJdQcLd/EP07uSz5wRITOTUR+vlc3238SVX3V/nkfmYku9+ebdXJL1npuTY0QkdZKSIwK11LxeW08ySTP6ip9EljWOWjBv/p+QataAtH/KujgmsdiP95XLPFiE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340365; c=relaxed/simple; bh=cDepZvLCDsPrrmaXN8drU183B//q+1TemaLu15C21dc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=GuCC0xPq9S7U4eTMIM+JWe+Ic+4kns68MU0wQ91OF+9AlIGLV1J4yQc/gNJOF2/OYg45O8YTPZpoXw6lskpGC44yXe9kLi0fM8DRq+REQLjCxshlzuM0VsCzqAUs1dbm7d3tIuP7ZL3Z2Oya9bXP41ekw0wkUHbnxn6LPDb9HyM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=K5lWJPtM; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="K5lWJPtM" Received: from pps.filterd (m0353725.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 68P4a08h3083370; Fri, 25 Sep 2026 12:45:57 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=pp1; bh=dgWmgQ 6lGVDbj8O5Ol1WIPPIVcga50UMh0bJP6YcDRs=; b=K5lWJPtMJQewkOYeA2qayj aaOSuSLQQy4rHb/ccTLWgrDpIl+3LVvef6xe9D3bsrAFI/n7GRXUwznufwOF9QQG +PW7Ji475vT/p0yL1MrSjpuBxPCSqMlIIkgZsI7ES7oIs2kc1SnMP2/ovLZ82UDO nNPxVOL/VJHac0x0JKsKjUvC75Qf22Ojsrg8S4wn7AsLTD/7sjBkBDu/c2HD5ZjZ yk5gELymYMPdiXElE30IhZdX+rtEVjX95toYzuVq41yipFe2dc5TSVFdJk6Oym68 CmHrWWrIU4BRxRoO1jr80hHO7HK5SDa2NnefkVqbud54DbaXiPb6qZnluJwkd1OA == Received: from ppma23.wdc07v.mail.ibm.com (5d.69.3da9.ip4.static.sl-reverse.com [169.61.105.93]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4gskgqx0s7-1 (version=TLSv1.3 cipher=TLS_AES_256_GCM_SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:45:56 +0000 (GMT) Received: from pps.filterd (ppma23.wdc07v.mail.ibm.com [127.0.0.1]) by ppma23.wdc07v.mail.ibm.com (8.18.1.11/8.18.1.11) with ESMTP id 68PA457j3288221; Fri, 25 Sep 2026 12:45:56 GMT Received: from smtprelay05.wdc07v.mail.ibm.com ([172.16.1.72]) by ppma23.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4gvbe22e1b-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:45:56 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (smtpav04.dal12v.mail.ibm.com [10.241.53.103]) by smtprelay05.wdc07v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 68PCjsUG24248932 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 25 Sep 2026 12:45:55 GMT Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id BAFAB58062; Fri, 25 Sep 2026 12:45:54 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 60EAF5805A; Fri, 25 Sep 2026 12:45:53 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.86.59]) by smtpav04.dal12v.mail.ibm.com (Postfix) with ESMTP; Fri, 25 Sep 2026 12:45:53 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, freude@linux.ibm.com, stable@vger.kernel.org Subject: [PATCH v8 1/6] s390/vfio-ap: Fix leaks of pinned NIB and registered GISC Date: Fri, 25 Sep 2026 08:45:46 -0400 Message-ID: <20260925124551.665448-2-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260925124551.665448-1-akrowiak@linux.ibm.com> References: <20260925124551.665448-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Authority-Analysis: v=2.4 cv=G+OJgNk5 c=1 sm=1 tr=0 ts=6ab66d05 cx=c_pps a=3Bg1Hr4SwmMryq2xdFQyZA==:117 a=3Bg1Hr4SwmMryq2xdFQyZA==:17 a=IkcTkHD0fZMA:10 a=VdqzKS8jKosA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=V8glGbnc2Ofi9Qvn3v5h:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=eS_XXSIdeN0PBYrJ018A:9 a=HbXCwHrqLsxgwAln:21 a=3ZKOabzyN94A:10 a=QEXdDO2ut3YA:10 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX4XskuktfZiPB nE2vSr4sOOzTWTZ38Hh8v/nfbP5EDn8TV/+moIJ6+Sf+vG8LnCtiUgUmrx8fF9DJqxHKskwdINA bc6idxfVsLBHodnsk+587ouLqr5Q/gMnhnepv6mL1EunDBIQyi6oFSj4TJwd2mWQAc/TUEIKr/F qHS0wdKzjzPX2DI4mu19RuM2UEAJ6CXkJs2fTrbWyQmMwQAMt/an0rQMeJhfPil1GoS4aBI4Brd NNPtDlQR8GwnZVG0/9gqvYgyXMduuRtZIs5sqOAadFbolEZDx2frm7CtOgEqhjM7hFWDteOyFaq 8uIE8M1ywISARU49dMZccHMIDOJxI7lFjNjAxbLug4fzTS6OgbE6FnMBvcl3/aRlpk7IxT1RVoW iB/qsqbTtJrfbR3tm6NGF5HVN0BjgqMgzBczrzzK+HuqoeJmLQdKuNSoY8op7y5SRBQS9gFNMKU bUVL70ODfXBm/iC/qsQ== X-Proofpoint-ORIG-GUID: 4JW3jE_g9BST-fg2lWo2dyR6-dUQiQ7n X-Proofpoint-GUID: 4JW3jE_g9BST-fg2lWo2dyR6-dUQiQ7n X-Proofpoint-Spam-Info: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX4Wk56uLK8yaR MkWFzdZWmtiARsrGiJqQm2eskqselWkM9vcevYmKlediwT71+DPotcwG+TfLyJ1GHAZUkfJXw70 xfEL4UmZ2JfR/N0yciOpGq9HcMjhFrE= X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-09-25_02,2026-09-21_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 suspectscore=0 adultscore=0 phishscore=0 lowpriorityscore=0 impostorscore=0 bulkscore=0 priorityscore=1501 clxscore=1015 spamscore=0 malwarescore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2609040000 definitions=main-2609250050 Several code paths in the vfio_ap driver failed to free the AQIC resources =E2=80=94 the pinned guest NIB page and the registered guest ISC used to enable interrupts for a queue =E2=80=94 when a queue became unavail= able or when unexpected response codes were returned. This could cause memory exhaustion and depletion of KVM interrupt subclass registrations over time with repeated dynamic AP reconfiguration. On the other hand, there are situations whereby these resources must be intentionally leaked. If the page were unpinned and returned to the allocator, a subsequent wild DMA-write to that physical address would corrupt memory belonging to the new owner and could crash or compromise the host kernel. vfio_ap_mdev_reset_queue() ~~~~~~~~~~~~~~~~~~~~~~~~~~ AP_RESPONSE_Q_NOT_AVAIL (0x01) was not handled, causing it to fall through to the default case which only fired a WARN without calling vfio_ap_free_aqic_resources(). When ap_zapq() returns this response code the queue is physically unavailable and can no longer generate AP interrupts or DMA-write to the NIB. Add AP_RESPONSE_Q_NOT_AVAIL alongside AP_RESPONSE_DECONFIGURED and AP_RESPONSE_CHECKSTOPPED so that AQIC resources are freed immediately for all three non- operational cases. AP_RESPONSE_BUSY is not a valid response code for a PQAP(ZAPQ) instruction, so it is removed. This contradicts what is stated in the following: commit 411b0109daa52 ("s390/vfio-ap: wait for response code 05 to clear on = queue reset") A careful reading of the valid response code table in the architecture documentation, however, clearly shows that response code 05 is not valid for the PQAP-ZAPQ instruction. So in the apq_status_check() - which examines the response codes that can be returned from PQAP-TAPQ - if the response code is AP_RESPONSE_BUSY (05), it will return -EAGAIN which instructs the caller (apq_reset_check()) to re-issue the PQAP-ZAPQ instruction. vfio_ap_mdev_remove_queue() ~~~~~~~~~~~~~~~~~~~~~~~~~~~ When the AP bus scan detects a queue is no longer in the host AP configuration, vfio_ap_mdev_remove_queue() skips the ZAPQ since issuing it would return AP_RESPONSE_Q_NOT_AVAIL anyway. However, vfio_ap_free_aqic_resources() was also never called, leaking the pinned NIB page and registered guest ISC. Since the hardware is gone and can no longer DMA-write to the NIB, it is safe to call vfio_ap_free_aqic_resources() directly. Add an else branch to the host-config test_bit_inv guard to free AQIC resources when the queue is not in the host AP configuration. If the queue is not assigned to an mdev, vfio_ap_free_aqic_resources() is a no-op. apq_reset_check() ~~~~~~~~~~~~~~~~~ Keep in mind that apq_reset_check() - which calls apq_status_check() - is called on a work queue when the status response code from the ZAPQ is AP_RESPONSE_NORMAL, AP_RESPONSE_RESET_IN_PROGRESS, or AP_RESPONSE_STATE_CHANGE_IN_PROGRESS. If the apq_reset_check() times out before it can be verified that the asynchronous portion of the reset completed, the response code returned will be one of those above, so the status from the PQAP(TAPQ) is copied to q->reset_status so the queue is marked as not-passable and not passed through to a guest. ~~~~~~~~~~~~~~~~~~~~~~~~~~~ vfio_ap_irq_disable() ~~~~~~~~~~~~~~~~~~~~~ The valid response codes for PQAP(AQIC) include AP_RESPONSE_STATE_CHANGE_IN_PROGRESS (0x0a), AP_RESPONSE_INVALID_GISA (0x08), AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE (0x35), and AP_RESPONSE_ASSOC_FAILED (0x36), none of which were explicitly handled. All four fell through to the default case and jumped to end_free freed the AQIC resources. These each should be handled differently: * AP_RESPONSE_STATE_CHANGE_IN_PROGRESS indicates a transient condition, analogous to AP_RESPONSE_RESET_IN_PROGRESS and AP_RESPONSE_BUSY. It is added to the same case as the RESET_IN_PROGRESS AND RESPONSE_BUSY whereby the process sleeps for 20ms and the AQIC gets re-executed. * AP_RESPONSE_INVALID_GISA indicates the AQIC instruction was rejected due to an invalid GISA address. The disable did not take effect and the hardware still holds the NIB page address, so the AQIC resources must be leaked. If the NIB page were unpinned and returned to the allocator, a subsequent wild DMA-write to that physical address would corrupt memory belonging to the new owner and could crash or compromise the host kernel; so the AQIC resources must be intentionally leaked. * AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE and AP_RESPONSE_ASSOC_FAILED are asynchronous response codes from a previously executed association instruction. All subsequent AQIC calls will end with the asynchronous response code until the queue is reset; therefore the AQIC disable was not executed and the hardware still holds the NIB address. If the NIB page was unpinned and returned to the allocator, a subsequent wild DMA-write to that physical address would corrupt memory belonging to the new owner and could crash or compromise the host kernel; so the AQIC resources must be intentionally leaked. After a return from the vfio_ap_wait_for_clear(), vfio_ap_irq_disable() immediately frees the AQIC resources; however, vfio_ap_wait_for_clear() can return for a number of reasons which require a different response: * Timed out waiting for the I-bit (bit 7) in the AP queue status word returned from the TAPQ instruction to be cleared (indicates interrupts are disabled). In this case, the hardware still holds the NIB address. If the NIB page were unpinned and returned to the allocator, a subsequent wild DMA-write to that physical address would corrupt memory belonging to the new owner and could crash or compromise the host kernel; so the AQIC resources must be intentionally leaked. * The response code from TAPQ is one of AP_RESPONSE_NORMAL, AP_RESPONSE_Q_NOT_AVAIL, AP_RESPONSE_DECONFIGURED, OR AP_RESPONSE_CHECKSTOPPED. In this case, the response code indicates the interrupts are disabled or the queue is either not in the host's AP configuration or is not functional. In any case, it is safe to free the AQIC resources. * An invalid response code was returned from TAPQ. In this case, the hardware may still hold the NIB address. If so and the NIB page was unpinned and returned to the allocator, a subsequent wild DMA-write to that physical address would corrupt memory belonging to the new owner and could crash or compromise the host kernel; so the AQIC resources must be intentionally leaked. The solution here is to change the vfio_ap_wait_for_clear() to reply with a return code indicating the result, thus allowing vfio_ap_irq_disable() to respond accordingly: * 0: interrupt disablement is verified * -ENODEV: the queue is not functional * -ETIMEDOUT: the function timeout without verifying interrupts disabled * -EIO: an invalid response code was returned from TAPQ unmap_iova() ~~~~~~~~~~~~ Calls vfio_ap_irq_disable() to disable interrupts for the queue, but does not check the result. If the IRQ disable failed or could not be confirmed, then vfio_ap_irq_disable() leaks the NIB to prevent a wild DMA-write; however, the vfio core requires that NIB page to be unpinned before the dma_unmap returns, or it will BUG_ON after 10 re-notification rounds. The fix is to fall back to a bounded queue reset and zeroize (ZAPQ) which zeroizes the NIB pointer in the hardware, eliminating the DMA risk that justified the leak. Once the reset finishes or times out with a reset confirmed in-progress, the hardware no longer holds a reference to saved_iova and it is safe to unpin unconditionally. Fixes: ec89b55e3bce7 ("s390: ap: implement PAPQ AQIC interception in kernel= ") Cc: stable@vger.kernel.org Signed-off-by: Anthony Krowiak --- drivers/s390/crypto/vfio_ap_ops.c | 365 +++++++++++++++++++++++++----- 1 file changed, 304 insertions(+), 61 deletions(-) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index 940c0ff668be..6998bd0c88a2 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -31,6 +31,7 @@ #define AP_QUEUE_IN_USE "in use" =20 #define AP_RESET_INTERVAL 20 /* Reset sleep interval (20ms) */ +#define AP_RESET_MAX_WAIT 2000 /* Maximum wait for reset (2000ms) */ =20 static int vfio_ap_mdev_reset_queues(struct ap_matrix_mdev *matrix_mdev); static int vfio_ap_mdev_reset_qlist(struct list_head *qlist); @@ -226,27 +227,48 @@ static struct vfio_ap_queue *vfio_ap_mdev_get_queue( } =20 /** - * vfio_ap_wait_for_irqclear - clears the IR bit or gives up after 5 tries - * @apqn: The AP Queue number - * - * Checks the IRQ bit for the status of this APQN using ap_tapq. - * Returns if the ap_tapq function succeeded and the bit is clear. - * Returns if ap_tapq function failed with invalid, deconfigured or - * checkstopped AP. - * Otherwise retries up to 5 times after waiting 20ms. + * vfio_ap_wait_for_irqclear - wait for the IR bit to clear after a disable + * + * @apqn: the APQN of the queue + * @tapq_status: used to return the TAPQ status to the caller + * + * Repeatedly polls the AP queue status via PQAP(TAPQ) every 20ms until th= e IR + * bit is clear, the queue becomes non-operational, or 5 retries are exhau= sted. + * + * Because PQAP(AQIC) disable initiates an asynchronous process, a + * condition-code 0 completion does not guarantee the IR bit has been clea= red. + * The host must confirm IR=3D0 before unpinning the NIB page to avoid a w= ild + * DMA write to a freed page. + * + * Return: + * 0 if the IR bit is clear (i.e., interrupts are disabled). + * + * -ENODEV if the PQAP-TAPQ response code indicates the queue is not avail= able, + * is deconfigured, or is checkstopped (i.e., not operational). + * + * -ETIMEDOUT the function timed out before the IR bit was cleared, or TAPQ + * returned AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE or + * AP_RESPONSE_ASSOC_FAILED, which mean the instruction was not + * executed and will continue to be returned for all subsequent + * instructions except ZAPQ until the queue is reset. Since IR=3D0 + * cannot be confirmed, the NIB must be treated as a potential DMA + * target and leaked rather than freed. + * + * -EIO PQAP-TAPQ returned an invalid response code */ -static void vfio_ap_wait_for_irqclear(int apqn) +static int vfio_ap_wait_for_irqclear(int apqn, struct ap_queue_status *tap= q_status) { struct ap_queue_status status; int retry =3D 5; =20 do { status =3D ap_tapq(apqn, NULL); + memcpy(tapq_status, &status, sizeof(status)); switch (status.response_code) { case AP_RESPONSE_NORMAL: case AP_RESPONSE_RESET_IN_PROGRESS: if (!status.irq_enabled) - return; + return 0; fallthrough; case AP_RESPONSE_BUSY: msleep(20); @@ -254,15 +276,34 @@ static void vfio_ap_wait_for_irqclear(int apqn) case AP_RESPONSE_Q_NOT_AVAIL: case AP_RESPONSE_DECONFIGURED: case AP_RESPONSE_CHECKSTOPPED: + WARN_ONCE(1, "%s: tapq rc %02x: %04x\n", __func__, + status.response_code, apqn); + return -ENODEV; + case AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE: + case AP_RESPONSE_ASSOC_FAILED: + /* + * The TAPQ instruction was not executed. Executing TAPQ + * again will result in the same error until the queue + * is reset. We should't, however, reset the queue in + * this context, so log a warning and return -ETIMEDOUT + * since that would happen anyway if we continued to + * execute the TAPQ. + */ + WARN_ONCE(1, "%s: tapq rc %02x: %04x\n", __func__, + status.response_code, apqn); + return -ETIMEDOUT; default: WARN_ONCE(1, "%s: tapq rc %02x: %04x\n", __func__, status.response_code, apqn); - return; + return -EIO; } } while (--retry); =20 - WARN_ONCE(1, "%s: tapq rc %02x: %04x could not clear IR bit\n", - __func__, status.response_code, apqn); + WARN_ONCE(1, "%s: tapq rc %02x: timed out waiting for interrupts disabled= for %02x.%04x\n", + __func__, status.response_code, + AP_QID_CARD(apqn), AP_QID_QUEUE(apqn)); + + return -ETIMEDOUT; } =20 /** @@ -289,55 +330,139 @@ static void vfio_ap_free_aqic_resources(struct vfio_= ap_queue *q) } =20 /** - * vfio_ap_irq_disable - disables and clears an ap_queue interrupt - * @q: The vfio_ap_queue + * vfio_ap_irq_disable - disable interrupts for an AP queue + * @q: the vfio_ap_queue * - * Uses ap_aqic to disable the interruption and in case of success, reset - * in progress or IRQ disable command already proceeded: calls - * vfio_ap_wait_for_irqclear() to check for the IRQ bit to be clear - * and calls vfio_ap_free_aqic_resources() to free the resources associated - * with the AP interrupt handling. + * Issues PQAP(AQIC) to disable interrupts for the AP queue. On success + * (AP_RESPONSE_NORMAL or AP_RESPONSE_OTHERWISE_CHANGED), polls via + * vfio_ap_wait_for_irqclear() until the IR bit is confirmed clear before + * freeing the pinned NIB page and unregistering the guest ISC. This wait = is + * necessary because AQIC disable is asynchronous: freeing the NIB before = IR=3D0 + * is confirmed risks a wild DMA write to a freed host page. * - * In the case the AP is busy, or a reset is in progress, - * retries after 20ms, up to 5 times. + * Retries up to 5 times (with 20ms sleep) if the queue is busy or a reset= is + * in progress. * - * Returns if ap_aqic function failed with invalid, deconfigured or - * checkstopped AP. + * If the IR bit cannot be confirmed clear (timeout), the NIB page and gue= st + * ISC are intentionally leaked. If the page were unpinned and returned to= the + * allocator, a subsequent hardware DMA write to that physical address wou= ld + * corrupt memory belonging to a new owner =E2=80=94 a wild DMA write that= could crash + * or compromise the host kernel. + * + * If the queue is non-operational (deconfigured, checkstopped, not availa= ble), + * resources are freed immediately since the hardware can no longer write = to + * the NIB. * * Return: &struct ap_queue_status */ static struct ap_queue_status vfio_ap_irq_disable(struct vfio_ap_queue *q) { union ap_qirq_ctrl aqic_gisa =3D { .value =3D 0 }; - struct ap_queue_status status; - int retries =3D 5; + struct ap_queue_status status, tapq_status; + int retries =3D 5, ret; =20 do { status =3D ap_aqic(q->apqn, aqic_gisa, 0); switch (status.response_code) { case AP_RESPONSE_OTHERWISE_CHANGED: case AP_RESPONSE_NORMAL: - vfio_ap_wait_for_irqclear(q->apqn); - goto end_free; + /* + * AQIC disable was accepted (NORMAL), or the queue was + * already disabled or a prior async request is still + * completing (OTHERWISE_CHANGED). In both cases, we must + * wait until interrupt processing has been disabled + * before proceeding. + */ + ret =3D vfio_ap_wait_for_irqclear(q->apqn, &tapq_status); + if (ret =3D=3D 0 || ret =3D=3D -ENODEV) + goto end_free; + + if (ret =3D=3D -EIO) { + /* + * An unknown TAPQ response code was returned which + * indicates a bug or some type of hardware I/O issue. + * Since we don't know whether queue interrupts were + * disabled or not, the AQIC resources must be leaked. + */ + memcpy(&status, &tapq_status, sizeof(status)); + goto end_fail; + } + + /* Timed out waiting to confirm interrupts are disabled */ + if (tapq_status.response_code =3D=3D AP_RESPONSE_NORMAL || + tapq_status.response_code =3D=3D AP_RESPONSE_BUSY) { + /* + * If AQIC returned NORMAL, the guest would incorrectly + * interpret that as a successful disable and may free or + * reuse the NIB while hardware can still write to it. + * AP_RESPONSE_BUSY is not valid for PQAP-AQIC. + * + * Return OTHERWISE_CHANGED to signal to the guest that + * the disable interrupts operation did not complete. + */ + memset(&status, 0, sizeof(status)); + status.response_code =3D AP_RESPONSE_OTHERWISE_CHANGED; + } else { + /* + * For all other TAPQ response codes, + * return the TAPQ status directly since those + * codes are also valid for AQIC. + */ + memcpy(&status, &tapq_status, sizeof(status)); + } + goto end_fail; case AP_RESPONSE_RESET_IN_PROGRESS: case AP_RESPONSE_BUSY: + case AP_RESPONSE_STATE_CHANGE_IN_PROGRESS: msleep(20); break; case AP_RESPONSE_Q_NOT_AVAIL: case AP_RESPONSE_DECONFIGURED: case AP_RESPONSE_CHECKSTOPPED: + /* AP not operational; no further interrupts possible */ + WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, + status.response_code); + goto end_free; case AP_RESPONSE_INVALID_ADDRESS: + case AP_RESPONSE_INVALID_GISA: + case AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE: + case AP_RESPONSE_ASSOC_FAILED: default: - /* All cases in default means AP not operational */ + /* + * The AQIC disable was rejected; IRQ is still enabled + * and the hardware still holds the NIB address. Do not + * free resources. + */ WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, status.response_code); - goto end_free; + goto end_fail; } } while (retries--); =20 WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, status.response_code); + +end_fail: + /* + * We are here either because the AQIC instruction failed to disable + * interrupts, or because IR=3D0 could not be confirmed. In either case + * the NIB page and guest ISC cannot be freed: hardware may still write + * to the NIB, and unpinning the page would allow it to be reallocated + * to a new owner. A subsequent hardware DMA write to that physical + * address would corrupt the new owner's memory =E2=80=94 a wild DMA writ= e that + * could crash or compromise the host kernel. The resources are + * therefore intentionally leaked. + */ + return status; + end_free: + /* + * This label is reached because the queue was successfully disabled, + * or because the queue is not operational or not available, in which case + * interrupts can not be processed, so free the AQIC resources - the pinn= ed NIB + * page and the registered guest ISC - used to enable interrupts so they = will + * not be leaked. + */ vfio_ap_free_aqic_resources(q); return status; } @@ -401,22 +526,29 @@ static int ensure_nib_shared(unsigned long addr) } =20 /** - * vfio_ap_irq_enable - Enable Interruption for a APQN + * vfio_ap_irq_enable - enable interrupts for an AP queue on behalf of a g= uest * - * @q: the vfio_ap_queue holding AQIC parameters + * @q: the vfio_ap_queue for which interrupts are to be enabled * @isc: the guest ISC to register with the GIB interface - * @vcpu: the vcpu object containing the registers specifying the paramete= rs - * passed to the PQAP(AQIC) instruction. + * @vcpu: the vcpu whose registers contain the PQAP(AQIC) parameters * - * Pin the NIB saved in *q - * Register the guest ISC to GIB interface and retrieve the - * host ISC to issue the host side PQAP/AQIC + * Pins the guest NIB page, registers the guest ISC with the GIB to obtain= a + * host ISC, and reissues PQAP(AQIC) with the translated host-absolute NIB + * address and host ISC on behalf of the guest. * - * status.response_code may be set to AP_RESPONSE_INVALID_ADDRESS in case = the - * vfio_pin_pages or kvm_s390_gisc_register failed. + * The condition code and AP-queue status word returned by PQAP(AQIC) are + * reflected back to the guest as-is. IRQ state verification (polling until + * IR=3D1) is the responsibility of the guest AP bus, not the host. * - * Otherwise return the ap_queue_status returned by the ap_aqic(), - * all retry handling will be done by the guest. + * Resource management is based solely on whether hardware accepted the ne= w NIB: + * - AP_RESPONSE_NORMAL (CC=3D0): hardware accepted the new NIB; the old p= inned + * NIB page and registered guest ISC are freed and the new ones saved. + * - All other responses: hardware did not accept the new NIB; the newly p= inned + * page and registered ISC are freed and the previously saved resources = are + * left intact. + * + * AP_RESPONSE_INVALID_ADDRESS is returned if vfio_pin_pages() or + * kvm_s390_gisc_register() fails before the AQIC instruction is issued. * * Return: &struct ap_queue_status */ @@ -428,11 +560,10 @@ static struct ap_queue_status vfio_ap_irq_enable(stru= ct vfio_ap_queue *q, struct ap_queue_status status =3D {}; struct kvm_s390_gisa *gisa; struct page *h_page; - int nisc; + int nisc, ret; struct kvm *kvm; phys_addr_t h_nib; dma_addr_t nib; - int ret; =20 /* Verify that the notification indicator byte address is valid */ if (vfio_ap_validate_nib(vcpu, &nib)) { @@ -489,27 +620,46 @@ static struct ap_queue_status vfio_ap_irq_enable(stru= ct vfio_ap_queue *q, status =3D ap_aqic(q->apqn, aqic_gisa, h_nib); switch (status.response_code) { case AP_RESPONSE_NORMAL: - /* See if we did clear older IRQ configuration */ + /* + * Hardware accepted the new NIB address (CC=3D0). The old NIB and + * guest ISC are no longer used by hardware and can be freed. + * The new resources are saved for tracking and future teardown. + * + * IRQ state verification (polling until IR=3D1) is the + * responsibility of the guest AP bus, not the host. The + * condition code and status word are reflected back to the + * guest to respond to the PQAP-AQIC instruction. + */ vfio_ap_free_aqic_resources(q); q->saved_iova =3D nib; q->saved_isc =3D isc; break; case AP_RESPONSE_OTHERWISE_CHANGED: - /* We could not modify IRQ settings: clear new configuration */ + /* + * Interrupts are already enabled or a prior async request is still + * completing. Either way, the hardware's current NIB is the one + * saved in q->saved_iova/q->saved_isc =E2=80=94 not the newly prepared + * resources. Release the newly pinned page and registered ISC; + * leave saved_iova and saved_isc intact. + */ + fallthrough; + default: + /* + * Hardware did not accept the new NIB (CC=3D3 or error). The + * previously saved NIB and guest ISC remain active and must + * not be freed. Release the newly pinned page and registered + * ISC that were prepared for this (rejected) request. + */ ret =3D kvm_s390_gisc_unregister(kvm, isc); if (ret) VFIO_AP_DBF_WARN("%s: kvm_s390_gisc_unregister: rc=3D%d isc=3D%d, apqn= =3D%#04x\n", __func__, ret, isc, q->apqn); vfio_unpin_pages(&q->matrix_mdev->vdev, nib, 1); break; - default: - pr_warn("%s: apqn %04x: response: %02x\n", __func__, q->apqn, - status.response_code); - vfio_ap_irq_disable(q); - break; } =20 - if (status.response_code !=3D AP_RESPONSE_NORMAL) { + if (status.response_code !=3D AP_RESPONSE_NORMAL && + status.response_code !=3D AP_RESPONSE_OTHERWISE_CHANGED) { VFIO_AP_DBF_WARN("%s: PQAP(AQIC) failed with status=3D%#02x: " "zone=3D%#x, ir=3D%#x, gisc=3D%#x, f=3D%#x," "gisa=3D%#x, isc=3D%#x, apqn=3D%#04x\n", @@ -635,7 +785,6 @@ static int handle_pqap(struct kvm_vcpu *vcpu) } =20 status =3D vcpu->run->s.regs.gprs[1]; - /* If IR bit(16) is set we enable the interrupt */ if ((status >> (63 - 16)) & 0x01) qstatus =3D vfio_ap_irq_enable(q, status & 0x07, vcpu); @@ -1857,8 +2006,28 @@ static void unmap_iova(struct ap_matrix_mdev *matrix= _mdev, u64 iova, u64 length) int loop_cursor; =20 hash_for_each(qtable->queues, loop_cursor, q, mdev_qnode) { - if (q->saved_iova >=3D iova && q->saved_iova < iova + length) + if (q->saved_iova >=3D iova && q->saved_iova < iova + length) { vfio_ap_irq_disable(q); + /* + * If IRQ disable failed or IR=3D0 could not be confirmed, + * vfio_ap_irq_disable() intentionally leaks the NIB to + * prevent a wild DMA write. But vfio core requires the + * page to be unpinned before dma_unmap returns, or it + * will BUG_ON after 10 re-notification rounds. + * + * Fall back to a bounded queue reset. The ZAPQ zeroizes + * the NIB pointer in hardware, eliminating the DMA risk + * that justified the leak. Once the worker finishes (or + * times out with a reset confirmed in-progress), the + * hardware no longer holds a reference to saved_iova and + * it is safe to unpin unconditionally. + */ + if (q->saved_iova) { + vfio_ap_mdev_reset_queue(q); + flush_work(&q->reset_work); + vfio_ap_free_aqic_resources(q); + } + } } } =20 @@ -1919,22 +2088,57 @@ static int apq_status_check(int apqn, struct ap_que= ue_status *status) { switch (status->response_code) { case AP_RESPONSE_NORMAL: + /* + * This response code only indicates that the PQAP(ZAPQ) has + * been initiated. The following bit settings in the status + * returned from TAPQ must be verified to confirm that the + * asynchronous portion of the queue zeroization has completed. + */ + if (status->queue_empty && !status->replies_waiting && + !status->irq_enabled && !status->async) + return 0; + + /* Async zeroization still in progress; keep waiting */ + return -EBUSY; + case AP_RESPONSE_DECONFIGURED: case AP_RESPONSE_CHECKSTOPPED: - return 0; + /* + * The queue is non-operational: interrupts are not possible so + * AQIC resources can be safely freed. However, zeroization + * cannot be confirmed because all status bits are zeroed when + * these response codes are returned =E2=80=94 there is no way to + * distinguish a zeroized queue from one that has not been + * zeroized. Return -ENODEV to signal that AQIC resources should + * be freed but that zeroization has not been confirmed. + */ + return -ENODEV; + case AP_RESPONSE_RESET_IN_PROGRESS: - case AP_RESPONSE_BUSY: + /* + * A reset is in progress. It may be the reset we issued or one + * issued prior to ours; either way, once it completes the queue + * will be zeroized, so keep waiting. + */ return -EBUSY; + + case AP_RESPONSE_BUSY: case AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE: case AP_RESPONSE_ASSOC_FAILED: /* + * AP_RESPONSE_BUSY: + * The queue is busy with something unrelated to a reset and our + * ZAPQ was rejected outright. Re-issue the ZAPQ. + * + * AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE + * AP_RESPONSE_ASSOC_FAILED: * These asynchronous response codes indicate a PQAP(AAPQ) * instruction to associate a secret with the guest failed. All * subsequent AP instructions will end with the asynchronous - * response code until the AP queue is reset; so, let's return - * a value indicating a reset needs to be performed again. + * response code until the AP queue is reset. Re-issue the ZAPQ. */ return -EAGAIN; + default: WARN(true, "failed to verify reset of queue %02x.%04x: TAPQ rc=3D%u\n", @@ -1959,8 +2163,25 @@ static void apq_reset_check(struct work_struct *rese= t_work) elapsed +=3D AP_RESET_INTERVAL; status =3D ap_tapq(q->apqn, NULL); ret =3D apq_status_check(q->apqn, &status); - if (ret =3D=3D -EIO) + if (ret =3D=3D -EIO) { + /* + * TAPQ returned an invalid response code indicating a + * hardware or firmware bug. Since we cannot determine + * whether the queue can still DMA-write to the NIB, the + * AQIC resources are intentionally leaked lest the NIB + * page is reallocated to a new owner. A subsequent + * wild DMA-write to that physical address would corrupt + * the new owner's memory which could crash or compromise + * the host kernel. + * + * Record the TAPQ status so the queue is marked + * not-passable - _queue_passable() checks + * reset_status.response_code =3D=3D AP_RESPONSE_NORMAL - + * and the queue is not passed through to a guest. + */ + memcpy(&q->reset_status, &status, sizeof(status)); return; + } if (ret =3D=3D -EBUSY) { pr_notice_ratelimited(WAIT_MSG, elapsed, AP_QID_CARD(q->apqn), @@ -1977,8 +2198,12 @@ static void apq_reset_check(struct work_struct *rese= t_work) memcpy(&q->reset_status, &status, sizeof(status)); continue; } - if (q->saved_isc !=3D VFIO_AP_ISC_INVALID) - vfio_ap_free_aqic_resources(q); + /* + * We end up here when the ZAPQ has completed. ZAPQ + * disables interrupts, so the AQIC resources must be + * freed; otherwise they will be leaked. + */ + vfio_ap_free_aqic_resources(q); break; } } @@ -1995,18 +2220,27 @@ static void vfio_ap_mdev_reset_queue(struct vfio_ap= _queue *q) switch (status.response_code) { case AP_RESPONSE_NORMAL: case AP_RESPONSE_RESET_IN_PROGRESS: - case AP_RESPONSE_BUSY: case AP_RESPONSE_STATE_CHANGE_IN_PROGRESS: /* * Let's verify whether the ZAPQ completed successfully on a work queue. */ queue_work(system_long_wq, &q->reset_work); break; + case AP_RESPONSE_Q_NOT_AVAIL: case AP_RESPONSE_DECONFIGURED: case AP_RESPONSE_CHECKSTOPPED: vfio_ap_free_aqic_resources(q); break; default: + /* + * An invalid response code indicates a hardware or firmware bug. + * Since we cannot determine whether the queue can still + * DMA-write to the NIB, the AQIC resources are intentionally + * leaked lest the NIB page is reallocated to a new owner. A + * subsequent hardware wild DMA-write to that physical address + * would corrupt the new owner's memory which could crash or + * compromise the host kernel. + */ WARN(true, "PQAP/ZAPQ for %02x.%04x failed with invalid rc=3D%u\n", AP_QID_CARD(q->apqn), AP_QID_QUEUE(q->apqn), @@ -2534,6 +2768,15 @@ void vfio_ap_mdev_remove_queue(struct ap_device *apd= ev) test_bit_inv(apqi, (unsigned long *)matrix_dev->info.aqm)) { vfio_ap_mdev_reset_queue(q); flush_work(&q->reset_work); + } else { + /* + * The queue is no longer in the host's AP configuration. + * The hardware cannot DMA-write to the NIB, so it is safe + * to free the AQIC resources directly without issuing a + * ZAPQ. If the queue is not assigned to an mdev, + * vfio_ap_free_aqic_resources() is a no-op. + */ + vfio_ap_free_aqic_resources(q); } =20 done: --=20 2.53.0 From nobody Sat Sep 26 03:50:07 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4834D48094F; Fri, 25 Sep 2026 12:46:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340365; cv=none; b=QNAPLq5rG5yiyulAOLEtewZ6GTMpLNxRmyMiclc7hMBlRdST3/dqD9alpMrIthcHHlMDwRFM0nOjCj33XM6t3LOhQ91wO8+nH38B7uVovGjpfefyoPn38rEzRy7dgtyk4AoqFkNRyCQnkHnlc/kcD/3YBtuO+oFspm2tx3ImJBc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340365; c=relaxed/simple; bh=Lk8EuT5LJTb50lUlnLS5aoUz4i2T8zbmfo0SqD1myT0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=jWwT1xOniPfmz5MYdcLhjElU0lgXpxoEwRJcm6g6VN40YVYlPVHMVBEBghYoWzwwlcFnU/LJu67kTMsAatrH85rNQbaGpzdrChKq6KfPwtYYUzrtYfaNhQtztzSTjgRVCgrwT7pbvq3L8k5zLLry2ZDwVAE+hxicu1UcGtTH+2M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=A2ViYh4d; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="A2ViYh4d" Received: from pps.filterd (m0360072.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 68P4a7d8102472; Fri, 25 Sep 2026 12:45:59 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=DbJqklQ8fb8e6ZYgt DkSUGthyeop8YWWKKvxOzCS5p0=; b=A2ViYh4dX0drjXf/ZBJzYMNnM1aWthV70 HOKoTJSiVmmNv4YIpJbyY9uCkq+doBB4WXis9peqSS6k1lqE6LZ4N9ueuRUzoJVT xpg1SEslona3nsbtYsQGQACBQjrxe2yLBpsnCIhrq4/SnS1q5+oAVrQOTAKvhT3M YvGBrikAQzLYLlXXvaQbEWGyyn9w3WxZZ55ERP4o2l4BNMllUuntI3oeYarW83PP V3vDoWpSFZC/phveO1u7ymOAT3amB98u1ackjFciWVRXs/sf1r7AEA3P71pD2E3V 5slPFi9lY4IoFL13AD0YakM62Ob7oaK7rBPfOV5TyhB3LW/hErDrA== Received: from ppma11.dal12v.mail.ibm.com (db.9e.1632.ip4.static.sl-reverse.com [50.22.158.219]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4gskdvp2sf-1 (version=TLSv1.3 cipher=TLS_AES_256_GCM_SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:45:59 +0000 (GMT) Received: from pps.filterd (ppma11.dal12v.mail.ibm.com [127.0.0.1]) by ppma11.dal12v.mail.ibm.com (8.18.1.11/8.18.1.11) with ESMTP id 68P9jcCp3296986; Fri, 25 Sep 2026 12:45:58 GMT Received: from smtprelay07.wdc07v.mail.ibm.com ([172.16.1.74]) by ppma11.dal12v.mail.ibm.com (PPS) with ESMTPS id 4gvb6qted3-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:45:58 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (smtpav04.dal12v.mail.ibm.com [10.241.53.103]) by smtprelay07.wdc07v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 68PCjunB27919090 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 25 Sep 2026 12:45:56 GMT Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 2688F58066; Fri, 25 Sep 2026 12:45:56 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id D877458052; Fri, 25 Sep 2026 12:45:54 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.86.59]) by smtpav04.dal12v.mail.ibm.com (Postfix) with ESMTP; Fri, 25 Sep 2026 12:45:54 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, freude@linux.ibm.com, stable@vger.kernel.org Subject: [PATCH v8 2/6] s390/vfio-ap: Fix failure to release IRQ notification eventfd contexts Date: Fri, 25 Sep 2026 08:45:47 -0400 Message-ID: <20260925124551.665448-3-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260925124551.665448-1-akrowiak@linux.ibm.com> References: <20260925124551.665448-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-ORIG-GUID: xD_ypWxA8YCq8ypCD-BWkwq_EeJot2kQ X-Proofpoint-Spam-Info: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX2A/GDM0bCG3B ufXuHkWUK8m8XHNA1ZkJP0pmMNgZZkp3gvn/YNE3bYb3Y9z6SNH0QhzMzDR/dk5DJbzMMKu6Zb4 8FAYNlTGg9/ItEnxMi1vT/3GTRSoawY= X-Authority-Analysis: v=2.4 cv=FLiOVOos c=1 sm=1 tr=0 ts=6ab66d07 cx=c_pps a=aDMHemPKRhS1OARIsFnwRA==:117 a=aDMHemPKRhS1OARIsFnwRA==:17 a=VdqzKS8jKosA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=RzCfie-kr_QcCd8fBx8p:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=fxeDLeT5AZ2qCvnLJ4gA:9 a=+jEqtf1s3R9VXZ0wqowq2kgwd+I=:19 X-Proofpoint-GUID: xD_ypWxA8YCq8ypCD-BWkwq_EeJot2kQ X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX1odE1/uxL/OI +FzGuJ6wdKlGFy/c9Tv/dHteauTYmXlJlkQca0iq8nwIwyWd/lGA6LKPAslx8m6Z6mGZhu74OBr xUC2CZdN7pLtp+YGuCuqutqfLNOaDHXN74NfJjtSLQkkv8/Ej5aGZ8D8LYFUIE1x4a8oGdjVciB EHXRHaZcBoqbznPpLSia3NKnOSnDwZ4klTWQpTiBQ26bkLCZ2zCQR8U04vWIAa1Ib3mjzS76GZJ llyg0ipvlDgBKW1qPLJYw/0rnPFGfWSwdpZmG93BNmunaSb646gJnMVLrrPLIs8BfBUQR+oAKl/ UiSXl4l71nj0ftrcQgV0Yi3khU6HMe+Z4loCI/5GrKWYKEHgyymMK9ftqRSNCdSVJx9i4OlkiPS FQ/igWMtvzpjacXKHW+Ny8HVSi56xMVGYLxybDAM2G+9BuYJYntHFPnuBtFq1wEC0QhN7gPchdv 8/5w16lY41j++4jHyoA== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-09-25_02,2026-09-21_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 adultscore=0 impostorscore=0 phishscore=0 spamscore=0 clxscore=1015 suspectscore=0 bulkscore=0 lowpriorityscore=0 priorityscore=1501 malwarescore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2609040000 definitions=main-2609250050 Content-Type: text/plain; charset="utf-8" When userspace registers IRQ notification eventfds via the VFIO_DEVICE_SET_IRQS ioctl, vfio_ap_set_request_irq() and vfio_ap_set_cfg_change_irq() each call eventfd_ctx_fdget(), which takes a reference on the eventfd_ctx and stores it in matrix_mdev->req_trigger and matrix_mdev->cfg_chg_trigger respectively. These references are dropped only when userspace explicitly replaces or clears them via a subsequent SET_IRQS call. If the device is closed without that explicit teardown - because the guest exits, the VM process crashes, or the device file is simply closed - neither vfio_ap_mdev_close_device() nor the remove path releases these references. The eventfd_ctx backing objects and their associated file references therefore leak for the lifetime of the kernel. Fix this by introducing vfio_ap_mdev_release_eventfds() and calling it from vfio_ap_mdev_close_device() after vfio_ap_mdev_unset_kvm(). The VFIO core guarantees that close_device is called before vfio_unregister_group_dev() returns in the remove path, so fixing close_device is sufficient to cover both teardown paths. Note: ~~~~ The matrix_dev->mdevs lock must be held during the call to vfio_ap_mdev_release_eventfds(). There is a small window between the calls to vfio_ap_mdev_unset_kvm() which gets and releases the update locks and the acquisition of the matrix_dev->mdevs_lock mutex during which it is possible - although highly unlikely during normal operation - whereby a concurrent SET_IRQS call can get in. Taking matrix_dev->mdevs_lock around vfio_ap_mdev_release_eventfds() is sufficient to make this race-free. The SET_IRQS ioctl path writes req_trigger and cfg_chg_trigger only from vfio_ap_mdev_ioctl(), which holds mdevs_lock for its entire duration and always calls eventfd_ctx_put() on the previous value before storing the new one. Any number of concurrent SET_IRQS calls during the window between vfio_ap_mdev_unset_kvm() and the acquisition of mdevs_lock are therefore safe: each ioctl invocation puts the reference it found and installs a new one, leaving exactly one live reference in the field when it releases the lock. When release_eventfds subsequently acquires mdevs_lock it finds that single surviving reference and puts it. Conversely, a SET_IRQS call that loses the race and blocks on mdevs_lock will find the field NULL after release_eventfds finishes, take ownership of the reference it just created, and install it into a field that will never be read again - a transient leak. To close that final case, callers must ensure no new SET_IRQS ioctls can be issued after close_device() is called, which the VFIO core guarantees by releasing the device file before invoking close_device(). Fixes: bf48961f6f48e ("s390/vfio-ap: realize the VFIO_DEVICE_SET_IRQS ioctl= ") Cc: stable@vger.kernel.org Signed-off-by: Anthony Krowiak Reviewed-by: Matthew Rosato --- drivers/s390/crypto/vfio_ap_ops.c | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index 6998bd0c88a2..cbe2fb564a7e 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -2295,12 +2295,28 @@ static int vfio_ap_mdev_open_device(struct vfio_dev= ice *vdev) return vfio_ap_mdev_set_kvm(matrix_mdev, vdev->kvm); } =20 +static void vfio_ap_mdev_release_eventfds(struct ap_matrix_mdev *matrix_md= ev) +{ + if (matrix_mdev->req_trigger) { + eventfd_ctx_put(matrix_mdev->req_trigger); + matrix_mdev->req_trigger =3D NULL; + } + if (matrix_mdev->cfg_chg_trigger) { + eventfd_ctx_put(matrix_mdev->cfg_chg_trigger); + matrix_mdev->cfg_chg_trigger =3D NULL; + } +} + static void vfio_ap_mdev_close_device(struct vfio_device *vdev) { struct ap_matrix_mdev *matrix_mdev =3D container_of(vdev, struct ap_matrix_mdev, vdev); =20 vfio_ap_mdev_unset_kvm(matrix_mdev); + + mutex_lock(&matrix_dev->mdevs_lock); + vfio_ap_mdev_release_eventfds(matrix_mdev); + mutex_unlock(&matrix_dev->mdevs_lock); } =20 static void vfio_ap_mdev_request(struct vfio_device *vdev, unsigned int co= unt) --=20 2.53.0 From nobody Sat Sep 26 03:50:07 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6FE52443E56; Fri, 25 Sep 2026 12:46:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340367; cv=none; b=m8PJDCobhTkkgW/XsqJvcEr33D13oAjIkRbd+GzQbIXW8aix4h9PUgIZsgc69ZBHvxeIFXyzaTk7+Ae7Ww7OdKQoy4LYa97rhsGRCuER47l38Vf9BBIOFuf2vtm1OqQ5SJIorIw+np6ZY5/o2kvMK2dt6qKnhM1o8IAVO4WvFrE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340367; c=relaxed/simple; bh=y2Fed0C7y84JK0q/u00j/avv34yOALzuiq61AjFqUU8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=hl2gDIEajrLzZLinpUAhXRNfFhVMCpyKf3aBMHTAS4QZSdKRGHg0xtY56EhWya0SM6eAdrpaaEf1Ce+jS15OZFvJRYPIV9kWBbEVVRXOhlAiIrYQNqBTUEEv7KYqNDW5+LK/Qy7L2q4nYNgpk5Nmaw/+5LNGAKPltd78ojs9wuY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=bKfb3Fde; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="bKfb3Fde" Received: from pps.filterd (m0356517.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 68P4ac6u061841; Fri, 25 Sep 2026 12:45:59 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=03cZ3eQMUH71OX1oD x93QiyUJdNHzXpfkzyYZKdQa/Y=; b=bKfb3FdepzDQH+HPdr1s5b19u/9qxkidO LL5AeSLd2tMTJwecP/MbM615vUuCaQxDDlfWTSQWL8I8ydkVG2sroUHNRhva4lqA ngOO2L+RJZuCpO2OmmfYO2dnbfT/PmZ4A3uhcFboQg2gKB0LNCc6q8h2SKEyDGT3 Yo4pWVyYv4tS8llsODI1hwzgaBfgtBwPicJodTLJq24V9nqpBzvcWg0ELC1dvn6n pV8ERYcLysawlCudEc/AvXSjK6Zkc0KGBEplTSFKc5bFxREYF31HzL+/bUex3I8Y B72m56s4ELiFUE9iaOAo370hLRL/M/4CCatQ16hSUMd8GCY6jgQcA== Received: from ppma12.dal12v.mail.ibm.com (dc.9e.1632.ip4.static.sl-reverse.com [50.22.158.220]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4gskgsq096-1 (version=TLSv1.3 cipher=TLS_AES_256_GCM_SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:45:59 +0000 (GMT) Received: from pps.filterd (ppma12.dal12v.mail.ibm.com [127.0.0.1]) by ppma12.dal12v.mail.ibm.com (8.18.1.11/8.18.1.11) with ESMTP id 68P9uA9K3298386; Fri, 25 Sep 2026 12:45:58 GMT Received: from smtprelay04.dal12v.mail.ibm.com ([172.16.1.6]) by ppma12.dal12v.mail.ibm.com (PPS) with ESMTPS id 4gvb8k2e87-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:45:58 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (smtpav04.dal12v.mail.ibm.com [10.241.53.103]) by smtprelay04.dal12v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 68PCjv0j25035404 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 25 Sep 2026 12:45:57 GMT Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 83DCE58063; Fri, 25 Sep 2026 12:45:57 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 44BF558062; Fri, 25 Sep 2026 12:45:56 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.86.59]) by smtpav04.dal12v.mail.ibm.com (Postfix) with ESMTP; Fri, 25 Sep 2026 12:45:56 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, freude@linux.ibm.com, stable@vger.kernel.org Subject: [PATCH v8 3/6] s390/vfio-ap: Fix unbounded loop in apq_reset_check() Date: Fri, 25 Sep 2026 08:45:48 -0400 Message-ID: <20260925124551.665448-4-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260925124551.665448-1-akrowiak@linux.ibm.com> References: <20260925124551.665448-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Authority-Analysis: v=2.4 cv=V/XoQuni c=1 sm=1 tr=0 ts=6ab66d07 cx=c_pps a=bLidbwmWQ0KltjZqbj+ezA==:117 a=bLidbwmWQ0KltjZqbj+ezA==:17 a=VdqzKS8jKosA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=U7nrCbtTmkRpXpFmAIza:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=LCzJTBHBofSmWOAAFuYA:9 X-Proofpoint-ORIG-GUID: cJpxmXoIMX9A9XxvaD_DjVahBC1tx72L X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfXyOpqOkqrksbq s1krNSzt1i41fS5RDy2SmBq+zheyCZ6xR1gAoPaLfnAH0ht27Iu9DT3HM2iJaER4SVGQDtxg5vL ACOz3PGMlHqfTTyAN5OJ70L1OjV/h1grqW9lsI+0tnuT6pKPenffBAz0EBCr4VbQo9MEdNBsahL 4He1CB6mh1i7GmZduryUvKFoXLUAjT66CpTl15rHROQxMr1TC60LiWgJTulX1C141rsqbxQUqSM zUAZR2hE+17tjvSQrjAHJw2YiUq/G4faoFMI21hEvuJlZ65b7f3lZdLUKAHSMaMaT7pIgesfklF NlYJuozH0ZnC91+Dhfwq/h/fCWn+fCeyMTIR9aH/YzqhuXlro9p++lTy059Z6NKGQtBezYShfK+ TS8l8kYbBCKXkAS2zwQzSDoaEXnMdBJ5bnZ/hTNEn8LeeCbCiQzGCw/dP/B2EavNUcZn32JAhLr CbsaWBE8p0kjQo+vG2Q== X-Proofpoint-Spam-Info: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX0LU/obyMhv22 IPsemc20cFuXPpo7iws3ZB3/D2ImnQPVGXykyLHhocm2ElvNR2UCBIqSiZpcBLrW6q4+aE49xra UwnyR3ZW943HFiwJSEGCpZbVw8QwiAE= X-Proofpoint-GUID: cJpxmXoIMX9A9XxvaD_DjVahBC1tx72L X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-09-25_02,2026-09-21_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 priorityscore=1501 spamscore=0 malwarescore=0 clxscore=1015 phishscore=0 bulkscore=0 adultscore=0 lowpriorityscore=0 impostorscore=0 suspectscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2609040000 definitions=main-2609250050 Content-Type: text/plain; charset="utf-8" s390/vfio-ap: Fix unbounded loop in apq_reset_check() The apq_reset_check() worker polls ap_tapq() in a while(true) loop waiting for a queue reset to complete. When ap_tapq() returns AP_RESPONSE_BUSY, AP_RESPONSE_RESET_IN_PROGRESS, or AP_RESPONSE_BUSY, apq_status_check() returns -EBUSY and the loop continues after sleeping AP_RESET_INTERVAL (20ms). There is no upper bound on how many times the loop iterates, so if the hardware continuously returns a busy response the worker runs indefinitely. This is particularly harmful because several callers of vfio_ap_mdev_reset_queue() - such as vfio_ap_mdev_reset_queues(), vfio_ap_mdev_reset_qlist() and vfio_ap_mdev_remove_queue() - call flush_work() on the queue's reset_work while holding one or more of the global matrix_dev locks (guests_lock, mdevs_lock) or the KVM lock. An indefinitely spinning worker permanently blocks access to all ap_matrix_mdev objects, which could hang other guests that are using them. Fix this by introducing AP_RESET_MAX_WAIT (2000ms) and breaking out of the poll loop when elapsed time reaches that threshold. If apq_reset_check() times out before verifying completion of the reset, the AQIC resources associated with the queue cannot be freed. The NIB is the active DMA target for AP interrupt delivery until the reset completes; freeing the pinned page would allow it to be reallocated to a new owner. A subsequent hardware wild DMA-write to that physical address would corrupt the new owner's memory and could crash or compromise the host kernel. If the reset eventually completes, interrupts will be terminated, but the pinned NIB page and ISC registration will be leaked. This is preferable to a compromised kernel or kernel crash, or waiting indefinitely and blocking access to all mdevs, hanging the guests to which they are attached. On timeout, q->reset_status.response_code is set to AP_RESPONSE_RESET_IN_PROGRESS. This is used internally to signal that the reset did not complete, and ensures that if the queue is reset again, the re-issue logic in apq_reset_check() will re-issue the ZAPQ. This patch also fixes a bug whereby AP_RESPONSE_NORMAL (0) returned from PQAP(ZAPQ) was incorrectly treated as confirmation that the queue was zeroized. AP_RESPONSE_NORMAL only indicates that the ZAPQ was accepted; zeroization is performed asynchronously. To confirm completion, the following bits in the status word returned from PQAP(TAPQ) must all be verified: status->irq_enabled =3D=3D 0 status->queue_empty =3D=3D 1 status->replies_waiting =3D=3D 0 status->async =3D=3D 0 Fixes: dd174833e44e ("s390/vfio-ap: remove upper limit on wait for queue re= set to complete") Cc: stable@vger.kernel.org Signed-off-by: Anthony Krowiak --- drivers/s390/crypto/vfio_ap_ops.c | 134 ++++++++++++++++++++++++++---- 1 file changed, 118 insertions(+), 16 deletions(-) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index cbe2fb564a7e..47d4936fb9d7 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -2123,6 +2123,12 @@ static int apq_status_check(int apqn, struct ap_queu= e_status *status) return -EBUSY; =20 case AP_RESPONSE_BUSY: + /* + * The queue is busy with something unrelated to a reset and our + * ZAPQ was rejected outright. Re-issue the ZAPQ. + */ + return -EAGAIN; + case AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE: case AP_RESPONSE_ASSOC_FAILED: /* @@ -2148,8 +2154,59 @@ static int apq_status_check(int apqn, struct ap_queu= e_status *status) } } =20 +static void report_aqic_resource_leak(struct vfio_ap_queue *q) +{ + if (q->saved_isc !=3D VFIO_AP_ISC_INVALID || q->saved_iova) { + if (q->matrix_mdev) { + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "Reset timed out for APQN %02x.%04x: leaking AQIC resources (NIB= page & GISC) to prevent host crash\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn)); + } else { + pr_warn_ratelimited("Reset timed out for APQN %02x.%04x: leaking AQIC r= esources (NIB page & GISC) to prevent host crash\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn)); + } + } else { + if (q->matrix_mdev) { + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "Reset timed out for APQN %02x.%04x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn)); + } else { + pr_warn_ratelimited("Reset timed out for APQN %02x.%04x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn)); + } + } +} + #define WAIT_MSG "Waited %dms for reset of queue %02x.%04x (%u, %u, %u)" =20 +/** + * apq_reset_finalize - store final TAPQ status and free AQIC resources. + * @q: the vfio_ap_queue + * @status: the final AP queue status returned by PQAP(TAPQ) + * @ret: the return value from apq_status_check() + * + * Copies the full TAPQ status word to q->reset_status so that all status + * bits reflect the confirmed end state of the queue. If ret =3D=3D 0, + * zeroization was confirmed and the response code is overridden with + * AP_RESPONSE_NORMAL so that _queue_passable() returns true. For + * ret =3D=3D -ENODEV (DECONFIGURED or CHECKSTOPPED), the non-zero response + * code is left intact so _queue_passable() correctly returns false. + * AQIC resources are then freed. + */ +static void apq_reset_finalize(struct vfio_ap_queue *q, + struct ap_queue_status *status, int ret) +{ + memcpy(&q->reset_status, status, sizeof(*status)); + if (!ret) + q->reset_status.response_code =3D AP_RESPONSE_NORMAL; + + vfio_ap_free_aqic_resources(q); +} + static void apq_reset_check(struct work_struct *reset_work) { int ret =3D -EBUSY, elapsed =3D 0; @@ -2182,6 +2239,58 @@ static void apq_reset_check(struct work_struct *rese= t_work) memcpy(&q->reset_status, &status, sizeof(status)); return; } + + if (!ret || ret =3D=3D -ENODEV) { + /* + * Zeroization confirmed (ret =3D=3D 0): the TAPQ status bits + * indicate the async portion of the ZAPQ completed + * successfully. Free AQIC resources and return. + * + * Queue non-operational (ret =3D=3D -ENODEV): the queue is + * deconfigured or checkstopped; interrupts are not + * possible so AQIC resources can be safely freed. + * Zeroization cannot be confirmed in this state, but the + * queue cannot generate interrupts, so the NIB page is + * no longer a DMA target and it is safe to free it. + */ + apq_reset_finalize(q, &status, ret); + return; + } + + if (elapsed >=3D AP_RESET_MAX_WAIT) { + /* + * Timed out without being able to verify zapq completed. + * + * The AQIC resources associated with this queue - the pinned + * page containing the NIB and the registered guest ISC - + * cannot be freed here. The NIB is the active DMA target + * for AP interrupt delivery until the reset completes; + * freeing the pinned page while the hardware may still + * write to it would result in a wild DMA write that could + * corrupt host memory. + * + * If the reset eventually completes, interrupts will be + * terminated and the pinned NIB page and ISC registration + * will be leaked. This is preferable to either a wild DMA + * write or waiting indefinitely: flush_work() callers hold + * the matrix_dev->mdevs_lock mutex which serializes access + * to all mdev objects system-wide, so blocking here would + * hang all guests to which those mdevs are attached. + */ + report_aqic_resource_leak(q); + /* + * Zeroization could not be confirmed; set + * reset_status to AP_RESPONSE_RESET_IN_PROGRESS. + * This is used internally to signal that the reset + * did not complete, and ensures that if the queue + * is reset again, the re-issue logic in + * apq_reset_check() will re-issue the ZAPQ. + */ + q->reset_status.response_code =3D AP_RESPONSE_RESET_IN_PROGRESS; + + return; + } + if (ret =3D=3D -EBUSY) { pr_notice_ratelimited(WAIT_MSG, elapsed, AP_QID_CARD(q->apqn), @@ -2189,22 +2298,15 @@ static void apq_reset_check(struct work_struct *res= et_work) status.response_code, status.queue_empty, status.irq_enabled); - } else { - if (q->reset_status.response_code =3D=3D AP_RESPONSE_RESET_IN_PROGRESS = || - q->reset_status.response_code =3D=3D AP_RESPONSE_BUSY || - q->reset_status.response_code =3D=3D AP_RESPONSE_STATE_CHANGE_IN_PR= OGRESS || - ret =3D=3D -EAGAIN) { - status =3D ap_zapq(q->apqn, 0); - memcpy(&q->reset_status, &status, sizeof(status)); - continue; - } - /* - * We end up here when the ZAPQ has completed. ZAPQ - * disables interrupts, so the AQIC resources must be - * freed; otherwise they will be leaked. - */ - vfio_ap_free_aqic_resources(q); - break; + continue; + } + + if (ret =3D=3D -EAGAIN || + q->reset_status.response_code =3D=3D AP_RESPONSE_RESET_IN_PROGRESS || + q->reset_status.response_code =3D=3D AP_RESPONSE_STATE_CHANGE_IN_PRO= GRESS) { + status =3D ap_zapq(q->apqn, 0); + memcpy(&q->reset_status, &status, sizeof(status)); + elapsed =3D 0; } } } --=20 2.53.0 From nobody Sat Sep 26 03:50:07 2026 Received: from mx0a-001b2d01.pphosted.com (mx0a-001b2d01.pphosted.com [148.163.156.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 22B9549BD6F; Fri, 25 Sep 2026 12:46:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.156.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340368; cv=none; b=hy4MqNajtcIy+1jLjMBuJl65mEY9KyaCpyOeY+JPR19CrFL8sV5XvByvFUiJQKIgGp8RLN0m7YUMgFNIgKvtsK5LV2tsabuDmDpuo87yJzArBWmz5QyQmUR1fvE4v5jhAizCe4bzAGr5ySUi0+otY2ZVj+KQXXvzcUMUz3/Eabc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340368; c=relaxed/simple; bh=C7HhFeZjcRIDVvo6sITTqd90TodF9Mhzc6Lt6c8Rzcc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=VW48HYSSFVL7LTETON9wTKnDNv/Sk0jzL/lOYNzprwzm5glMYJ6Mh9GBuXvl2BdS1YGDiJsyE0oThbTJDzx6aSLguQWYxzybO6MJr74BIqGytJbRdyYUcVGi3/VReEcCCEMEfxGW5UQ90ot+2TX6WvhY8USqRoWeoZtEiDDSti4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=b2SiSogg; arc=none smtp.client-ip=148.163.156.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="b2SiSogg" Received: from pps.filterd (m0353729.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 68P4abPd2393903; Fri, 25 Sep 2026 12:46:01 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=uDxrYb9gnE5pqt4yu GCe9DFbRDF6vA95mkgEWKGvy5Y=; b=b2SiSogg+ziH/pEYg0h1TDTIocDNNKloj iHLPFUiPx+9uzO4DHn8ZLW1nSQez+dEmNFJLZ3220o7VVRCS1nYooFvQcpc4igQu dd6XsWkmoHWiYuqHeJzdH1dGiJhJFskPmhsC3szJDs5WXk2VGlJqjgv55r5EV/Xu nhlSexj5WXFtGFvpv5CKmELW6TH3lb7ZIkiKxZt4Y0hY1FHnAOY0dlsi4aUhapyr cfwivFM+FvNIknO4mJqYzlApY8gRwSFIqouq7gYOyawaW8xVa82pxKN0slydRdk3 G/39FgGg+tF1Vwx2ve5JqUyfDQF0CWjUJHnmGcqKLbnoFf154TJXQ== Received: from ppma22.wdc07v.mail.ibm.com (5c.69.3da9.ip4.static.sl-reverse.com [169.61.105.92]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4gske2729x-1 (version=TLSv1.3 cipher=TLS_AES_256_GCM_SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:46:01 +0000 (GMT) Received: from pps.filterd (ppma22.wdc07v.mail.ibm.com [127.0.0.1]) by ppma22.wdc07v.mail.ibm.com (8.18.1.11/8.18.1.11) with ESMTP id 68P9v56T3245755; Fri, 25 Sep 2026 12:46:00 GMT Received: from smtprelay05.dal12v.mail.ibm.com ([172.16.1.7]) by ppma22.wdc07v.mail.ibm.com (PPS) with ESMTPS id 4gvbu92aga-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:46:00 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (smtpav04.dal12v.mail.ibm.com [10.241.53.103]) by smtprelay05.dal12v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 68PCjwb933096418 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 25 Sep 2026 12:45:59 GMT Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id D08DE58056; Fri, 25 Sep 2026 12:45:58 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id A06C258052; Fri, 25 Sep 2026 12:45:57 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.86.59]) by smtpav04.dal12v.mail.ibm.com (Postfix) with ESMTP; Fri, 25 Sep 2026 12:45:57 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, freude@linux.ibm.com Subject: [PATCH v8 4/6] s390/vfio-ap: Use AP_DOMAINS for adm_add bitmap size in vfio_ap_mdev_cfg_add() Date: Fri, 25 Sep 2026 08:45:49 -0400 Message-ID: <20260925124551.665448-5-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260925124551.665448-1-akrowiak@linux.ibm.com> References: <20260925124551.665448-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-ORIG-GUID: lEAIDMZQ5hM48SiKcp6jxMFZKEokw-4U X-Authority-Analysis: v=2.4 cv=EOCTQFZC c=1 sm=1 tr=0 ts=6ab66d09 cx=c_pps a=5BHTudwdYE3Te8bg5FgnPg==:117 a=5BHTudwdYE3Te8bg5FgnPg==:17 a=VdqzKS8jKosA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=uAbxVGIbfxUO_5tXvNgY:22 a=VnNF1IyMAAAA:8 a=gReD5cTbzgwmqpVxveEA:9 X-Proofpoint-Spam-Info: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX3H9ZE68o5AiD OptErhFeMYiOiSvrmsfvVauEfpA2qziGFfDcc5GgNZ/VmsbE/Xr8VvKM7P6YC4bCKfi9+lp20IQ j+h0m0NCc7CS/YFCnAqyQqZ/XJS9UCw= X-Proofpoint-GUID: lEAIDMZQ5hM48SiKcp6jxMFZKEokw-4U X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX9RfUmzPaqJ98 eemRs36e/xokGTT/LXbYvhT+sTB1pzwTOsHJFzYYtuwQYusFJgG9/4LB+2VC4USIBUdNc/AcBhV iZ55g0Q4RvKK/VifnTX3hyQ0VAQdLdOLxRKL4WU5IKFg7xXWoCSd9UTKKgXCeaQyLSuT17FeBtB Qlu3Urt7Qp4xvZ55q3U33F2JJoLZ9oq6C3Lll3v1gz6Pakztvp4UbQ62adcUCYOKFPj7odq8g6u GmRLQmQPaxAFs8XuEjZ6XWyCVU/upkYntvHWuSaDmztI/YHc1Tqp/Tn7zss3BhPm73wXrUrPJeA hsn66aDnz4WGnVQBbgh8uPJX3sZSaC4Nr7q7r9687jEH0pU5DytsyLfa+x7OUh5/LzoxGg/cROg hJt0er5SZQ7fRPtJQAOOPGDIFmykzEK647pjF4no/C12eEQbng9KAQlWN/pnH3Ct5JRrfldmxoJ XGalo1J0Cb2sdbdyU+w== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-09-25_02,2026-09-21_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 bulkscore=0 phishscore=0 priorityscore=1501 clxscore=1015 spamscore=0 adultscore=0 impostorscore=0 lowpriorityscore=0 suspectscore=0 malwarescore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2609040000 definitions=main-2609250050 Content-Type: text/plain; charset="utf-8" Domain and control domain bitmaps are sized by the AP_DOMAINS constant, not AP_DEVICES. The two constants are both 256 today so there is no functional impact, but using the wrong constant is inconsistent with every operation on aqm/adm bitmaps. Use AP_DOMAINS to keep the code consistent and correct in case the two constants ever diverge. Note: This patch was submitted in response to a sashiko review comment pointing out there are other functions besides vfio_ap_mdev_cfg_add(), so there are fixes included here for those also. The subject line was kept the same since this is in v2 of this patch. Signed-off-by: Anthony Krowiak Reviewed-by: Jason J. Herne Reviewed-by: Matthew Rosato --- drivers/s390/crypto/vfio_ap_ops.c | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index 47d4936fb9d7..ffc2d8715bd9 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -1559,7 +1559,7 @@ static void vfio_ap_mdev_hot_unplug_domain(struct ap_= matrix_mdev *matrix_mdev, { DECLARE_BITMAP(apqis, AP_DOMAINS); =20 - bitmap_zero(apqis, AP_DEVICES); + bitmap_zero(apqis, AP_DOMAINS); set_bit_inv(apqi, apqis); vfio_ap_mdev_hot_unplug_domains(matrix_mdev, apqis); } @@ -3058,11 +3058,11 @@ static void vfio_ap_mdev_on_cfg_remove(struct ap_co= nfig_info *cur_config_info, do_remove |=3D bitmap_andnot(aqrem, (unsigned long *)prev_config_info->aqm, (unsigned long *)cur_config_info->aqm, - AP_DEVICES); + AP_DOMAINS); do_remove |=3D bitmap_andnot(cdrem, (unsigned long *)prev_config_info->adm, (unsigned long *)cur_config_info->adm, - AP_DEVICES); + AP_DOMAINS); =20 if (do_remove) vfio_ap_mdev_cfg_remove(aprem, aqrem, cdrem); @@ -3173,7 +3173,7 @@ static void vfio_ap_mdev_cfg_add(unsigned long *apm_a= dd, unsigned long *aqm_add, bitmap_and(matrix_mdev->aqm_add, matrix_mdev->matrix.aqm, aqm_add, AP_DOMAINS); bitmap_and(matrix_mdev->adm_add, - matrix_mdev->matrix.adm, adm_add, AP_DEVICES); + matrix_mdev->matrix.adm, adm_add, AP_DOMAINS); =20 mutex_unlock(&matrix_dev->mdevs_lock); } --=20 2.53.0 From nobody Sat Sep 26 03:50:07 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 25BB249E128; Fri, 25 Sep 2026 12:46:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340386; cv=none; b=pvh2YL8ZgYfAN4AW/42j2qAu0n55N9iftrE+1vMJz4SfqJlV4fGkA36MKH/UnCzKjbXSixaJom6XvdvvQP7jOAL7lEmhVeQTDjzNMyA1XiYLD9FlmsZceNYDOH6qSBiRc9I4hK3mqd/cpXQS3H43JOnhNsXspF73PTj26a1z1Qc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340386; c=relaxed/simple; bh=Je3ghUx5pd1RFGPL3NYEPLi0Hmp6ygvFQYm/tbRVR3Q=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=nPSocLM2AvtfuQ8LWDUiMehiwRgnJt8c7T56ddgWAZzVzUqr5jerClxTp6mzadvEaIAmkeaV//FOpoNrOPdZjl8zecu9k7qrbeBMpn8quNa3g6ofC5WfzO/DZ8lYnyEOTNyzc5jWoR/8mab9Q6k5VF7Q/fWwgecdNEAcgctoZIA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=KozZIiBw; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="KozZIiBw" Received: from pps.filterd (m0360072.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 68P4bGTE104300; Fri, 25 Sep 2026 12:46:02 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=q56NBarmd4MJcFXru O9PJfKc49PqBdL1Q7o8tn7iz5I=; b=KozZIiBwq+dbwLVaEv09uWhoprl1Wq9vg N1JoK9A7NJb3hh8Hgdh/yKcoto/TzE6QFg1h1D5qbM+2a4uTlnu2JJxBKzp/MXOr zgZz4asCUNr6xSz2yNfrBtB9shk8bIZYBOL2f7hXYh8wkFtG3oakqYU7cNGTxO2G xsQe+Cu8Vo/OwAx6bg/s0+YDTP2P2cWw+i67XvmiFvTTpwDOlBJ3Kwf+VdknGKdZ /2xSqLf+ZRp7UkiG5ywoEObXx9vpozAqyoLQ/depe1mBv9gSx0va7GPO0qPeEHXc 8avAjMbSR7z8HPXZUDBT4IQORqSeFE0ooIJszVcto+BAwXXCDtjfA== Received: from ppma12.dal12v.mail.ibm.com (dc.9e.1632.ip4.static.sl-reverse.com [50.22.158.220]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4gskdvp2sk-1 (version=TLSv1.3 cipher=TLS_AES_256_GCM_SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:46:02 +0000 (GMT) Received: from pps.filterd (ppma12.dal12v.mail.ibm.com [127.0.0.1]) by ppma12.dal12v.mail.ibm.com (8.18.1.11/8.18.1.11) with ESMTP id 68PA02Bh3298365; Fri, 25 Sep 2026 12:46:01 GMT Received: from smtprelay06.dal12v.mail.ibm.com ([172.16.1.8]) by ppma12.dal12v.mail.ibm.com (PPS) with ESMTPS id 4gvb8k2e8a-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:46:01 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (smtpav04.dal12v.mail.ibm.com [10.241.53.103]) by smtprelay06.dal12v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 68PCk04414680762 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 25 Sep 2026 12:46:00 GMT Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 8A1FE58063; Fri, 25 Sep 2026 12:46:00 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id EEEC85805A; Fri, 25 Sep 2026 12:45:58 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.86.59]) by smtpav04.dal12v.mail.ibm.com (Postfix) with ESMTP; Fri, 25 Sep 2026 12:45:58 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, freude@linux.ibm.com, stable@vger.kernel.org Subject: [PATCH v8 5/6] s390/vfio-ap: fix queue state leakage to guest and host Date: Fri, 25 Sep 2026 08:45:50 -0400 Message-ID: <20260925124551.665448-6-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260925124551.665448-1-akrowiak@linux.ibm.com> References: <20260925124551.665448-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-ORIG-GUID: vAC5mkCk9kG-K6wauqZkiQHj7LH_ij7D X-Proofpoint-Spam-Info: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfXxQOuitKeDCaC Y1MVn5hIW8Z84gmCGi3b8zdGtV4Ez23aO2UPuoUGTYpejrbGQGU1hdOZDDDyQn+7h42hoSwy3L2 mUbbfX7MKZM5xk8rqT9qrD1G2SSnSqg= X-Authority-Analysis: v=2.4 cv=FLiOVOos c=1 sm=1 tr=0 ts=6ab66d0a cx=c_pps a=bLidbwmWQ0KltjZqbj+ezA==:117 a=bLidbwmWQ0KltjZqbj+ezA==:17 a=VdqzKS8jKosA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=RzCfie-kr_QcCd8fBx8p:22 a=VwQbUJbxAAAA:8 a=VnNF1IyMAAAA:8 a=tMcIf9aicxS5gytnuLwA:9 X-Proofpoint-GUID: vAC5mkCk9kG-K6wauqZkiQHj7LH_ij7D X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX/JEhjYfQ43gg WOuuRZt0WqvsKNV1mnKSuosjl76ylk+xyuLEXOIs6KXnOighIPxf7g9C87FLxHsQeUGLUtCxpeo HsCQIEFS+RzcSwEsn0MigQSgBYg+zhQoOodwY0SUlit39gtZcJ9LBnxvVWiwTBHSvA6Tv6V3fUc 2MDUSLEC8zS3iN14y+XWVQuFASQc3gDQzm1meTtrepJL2z0Jh+hokxLjkkJAsENsSBqibtpKYx/ f3K2J9SZKD8za621/LSXfeJj4v63Jnk9C4zKxezCk9PPPmBdVVdP8t27+3lpYqtrGYbBJVDjDdx RkqVARo2RbQvFl5pnPE+7PT0c+dpS0RIhsjEmV0JdxmPheCOTZI5T5qftXbKfZk1EArXS7zF1f7 UiV61oN9p7C7HbzBA8YZZqAa9jyUgH7swYXopDS9dlfuTu5/MTVvM4hTwlE3qJ1o8g5AMSi/tXW IJjV7tsWg2tYsFA4V/w== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-09-25_02,2026-09-21_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 adultscore=0 impostorscore=0 phishscore=0 spamscore=0 clxscore=1015 suspectscore=0 bulkscore=0 lowpriorityscore=0 priorityscore=1501 malwarescore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2609040000 definitions=main-2609250050 Content-Type: text/plain; charset="utf-8" Commit dd174833e44e ("s390/vfio-ap: remove upper limit on wait for queue reset to complete") removed the upper bound on the wait for a queue reset to complete in apq_reset_check(), thus allowing the function to loop indefinitely. The reason given was to ensure both the security requirements and prevent resource leakage and corruption in the hypervisor. That is a legitimate concern; however, functions initiating the reset all hold the matrix_dev->mdevs_lock which guards access to all of the mdevs under the control of the vfio_ap device driver. Blocking of access prevents a system administrator from configuring the mdevs (i.e., assigning/unassigning adapters, domains and control domains via the mdev's sysfs interfaces) and may hang any guest that is started using one of the mdevs to supply its AP configuration. This patch limits the potential hang to the AP_RESET_MAX_WAIT (2000ms) timeout introduced in the preceding commit, and prevents leakage of queue state to a guest. _queue_passable() accepted AP_RESPONSE_DECONFIGURED and AP_RESPONSE_CHECKSTOPPED as passable states in addition to AP_RESPONSE_NORMAL. Neither DECONFIGURED nor CHECKSTOPPED confirms that the queue was zeroized; only AP_RESPONSE_NORMAL (0) does. A queue that is not confirmed zeroized must not be passed through to a guest, as it may contain key material from a previous guest or host operation. Since _queue_passable() now rejects queues that are check stopped or deconfigured, there is no way for a queue bound to the vfio_ap device driver to pass those through to a guest even if they are added back to the configuration or the reason they have check stopped has been fixed. To resolve this issue, a new on_queue_state_transition callback function is added to struct ap_driver which is invoked during the AP bus device scan when a queue device transitions from deconfigured to configured or check stopped to not check stopped and vice versa. The vfio_ap device driver provides an implementation that resets and zeroizes the queue when it transitions to configured or not check stopped and plugs it into the guest's AP configuration if the reset succeeds. To fix this, _queue_passable() is limited to returning true only when reset_status.response_code =3D=3D AP_RESPONSE_NORMAL. A new helper, apq_reset_finalize(), is introduced to ensure q->reset_status correctly reflects the confirmed end state of the queue. It copies the full TAPQ status word to q->reset_status and sets q->reset_status.response_code to AP_RESPONSE_NORMAL only when apq_status_check() returns 0, confirming zeroization via TAPQ status bit verification. For -ENODEV (DECONFIGURED or CHECKSTOPPED), the non-zero response code is left intact, ensuring _queue_passable() correctly returns false. Additionally, vfio_ap_mdev_probe_queue() now calls vfio_ap_mdev_reset_queue() and flush_work() at probe time to guarantee a clean queue before it can be assigned to a guest. Fixes: dd174833e44e ("s390/vfio-ap: remove upper limit on wait for queue re= set to complete") Cc: stable@vger.kernel.org Signed-off-by: Anthony Krowiak --- drivers/s390/crypto/ap_bus.c | 51 +++++++++++++++++++++ drivers/s390/crypto/ap_bus.h | 29 ++++++++++++ drivers/s390/crypto/vfio_ap_drv.c | 1 + drivers/s390/crypto/vfio_ap_ops.c | 64 ++++++++++++++++++++++----- drivers/s390/crypto/vfio_ap_private.h | 2 + 5 files changed, 137 insertions(+), 10 deletions(-) diff --git a/drivers/s390/crypto/ap_bus.c b/drivers/s390/crypto/ap_bus.c index d82df5b4e2db..a53cfad3543e 100644 --- a/drivers/s390/crypto/ap_bus.c +++ b/drivers/s390/crypto/ap_bus.c @@ -1979,6 +1979,28 @@ static inline void notify_scan_complete(void) __drv_notify_scan_complete); } =20 +/* Helper function for notify_config_changed */ +static int __drv_notify_qstate_transitioned(struct device_driver *drv, voi= d *data) +{ + struct ap_driver *ap_drv =3D to_ap_drv(drv); + struct ap_qstate_transition *qstate_trans =3D data; + + if (try_module_get(drv->owner)) { + if (ap_drv->on_qstate_transition) + ap_drv->on_qstate_transition(qstate_trans); + module_put(drv->owner); + } + + return 0; +} + +/* Notify all drivers about a queue state transition */ +static inline void notify_qstate_transitioned(struct ap_qstate_transition = *qstate_trans) +{ + bus_for_each_drv(&ap_bus_type, NULL, qstate_trans, + __drv_notify_qstate_transitioned); +} + /* * Helper function for ap_scan_bus(). * Remove card device and associated queue devices. @@ -1998,6 +2020,7 @@ static inline void ap_scan_rm_card_dev_and_queue_devs= (struct ap_card *ac) */ static inline void ap_scan_domains(struct ap_card *ac) { + struct ap_qstate_transition qstate_trans; struct ap_tapq_hwinfo hwinfo; bool decfg, chkstop; struct ap_queue *aq; @@ -2093,6 +2116,13 @@ static inline void ap_scan_domains(struct ap_card *a= c) spin_unlock_bh(&aq->lock); pr_debug("(%d,%d) queue dev checkstop on\n", ac->id, dom); + /* + * Notify drivers that the queue state has transitioned + * to checkstopped. + */ + qstate_trans.queue =3D aq; + qstate_trans.new_state =3D AP_QUEUE_CHKSTOP_ON; + notify_qstate_transitioned(&qstate_trans); /* 'receive' pending messages with -EAGAIN */ ap_flush_queue(aq); goto put_dev_and_continue; @@ -2104,6 +2134,13 @@ static inline void ap_scan_domains(struct ap_card *a= c) spin_unlock_bh(&aq->lock); pr_debug("(%d,%d) queue dev checkstop off\n", ac->id, dom); + /* + * Notify drivers that the queue state has transitioned + * to not checkstopped. + */ + qstate_trans.queue =3D aq; + qstate_trans.new_state =3D AP_QUEUE_CHKSTOP_OFF; + notify_qstate_transitioned(&qstate_trans); goto put_dev_and_continue; } /* config state change */ @@ -2117,6 +2154,13 @@ static inline void ap_scan_domains(struct ap_card *a= c) spin_unlock_bh(&aq->lock); pr_debug("(%d,%d) queue dev config off\n", ac->id, dom); + /* + * Notify drivers that the queue state has transitioned + * to deconfigured. + */ + qstate_trans.queue =3D aq; + qstate_trans.new_state =3D AP_QUEUE_CONFIG_OFF; + notify_qstate_transitioned(&qstate_trans); ap_send_config_uevent(&aq->ap_dev, aq->config); /* 'receive' pending messages with -EAGAIN */ ap_flush_queue(aq); @@ -2129,6 +2173,13 @@ static inline void ap_scan_domains(struct ap_card *a= c) spin_unlock_bh(&aq->lock); pr_debug("(%d,%d) queue dev config on\n", ac->id, dom); + /* + * Notify drivers that the queue state has transitioned + * to configured. + */ + qstate_trans.queue =3D aq; + qstate_trans.new_state =3D AP_QUEUE_CONFIG_ON; + notify_qstate_transitioned(&qstate_trans); ap_send_config_uevent(&aq->ap_dev, aq->config); goto put_dev_and_continue; } diff --git a/drivers/s390/crypto/ap_bus.h b/drivers/s390/crypto/ap_bus.h index fb4d678336e4..fd2c7be683e3 100644 --- a/drivers/s390/crypto/ap_bus.h +++ b/drivers/s390/crypto/ap_bus.h @@ -132,6 +132,29 @@ struct ap_message; */ #define AP_DRIVER_FLAG_DEFAULT 0x0001 =20 +/** + * ap_queue_state_transition: + * + * Used to notify a device driver that a queue state transition has occurr= ed. + * + * @queue: the queue device whose state transitioned + * @new_state: identifies the new state to which the queue transitioned: + * AP_QUEUE_CONFIG_ON: from deconfigured to configured + * AP_QUEUE_CONFIG_OFF: from configured to deconfigured + * AP_QUEUE_CHKSTOPPED_ON: from not checkstopped to checkstopped + * AP_QUEUE_CHKSTOPPED_OFF: from checkstopped to not checkstopped + */ +struct ap_qstate_transition { + struct ap_queue *queue; + + enum { + AP_QUEUE_CONFIG_ON, + AP_QUEUE_CONFIG_OFF, + AP_QUEUE_CHKSTOP_ON, + AP_QUEUE_CHKSTOP_OFF, + } new_state; +}; + struct ap_driver { struct device_driver driver; =20 @@ -155,6 +178,12 @@ struct ap_driver { void (*on_scan_complete)(struct ap_config_info *new_config_info, struct ap_config_info *old_config_info); =20 + /* + * Called during the ap bus scan when a queue state transition is + * detected. + */ + void (*on_qstate_transition)(struct ap_qstate_transition *qstate_trans); + struct ap_device_id *ids; unsigned int flags; }; diff --git a/drivers/s390/crypto/vfio_ap_drv.c b/drivers/s390/crypto/vfio_a= p_drv.c index 8e69ed286bb9..6a5d97fa9200 100644 --- a/drivers/s390/crypto/vfio_ap_drv.c +++ b/drivers/s390/crypto/vfio_ap_drv.c @@ -61,6 +61,7 @@ static struct ap_driver vfio_ap_drv =3D { .in_use =3D vfio_ap_mdev_resource_in_use, .on_config_changed =3D vfio_ap_on_cfg_changed, .on_scan_complete =3D vfio_ap_on_scan_complete, + .on_qstate_transition =3D vfio_ap_on_qstate_transition, .ids =3D ap_queue_ids, }; =20 diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index ffc2d8715bd9..cd4a436c4319 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -841,14 +841,14 @@ static bool _queue_passable(struct vfio_ap_queue *q) if (!q) return false; =20 - switch (q->reset_status.response_code) { - case AP_RESPONSE_NORMAL: - case AP_RESPONSE_DECONFIGURED: - case AP_RESPONSE_CHECKSTOPPED: - return true; - default: - return false; - } + /* + * A queue is only passable if zeroization was confirmed by + * apq_reset_check() via TAPQ status bit verification. This is + * indicated by reset_status.response_code =3D=3D AP_RESPONSE_NORMAL (0). + * This is to protect against leaking the internal state of the queue + * to the guest. + */ + return q->reset_status.response_code =3D=3D AP_RESPONSE_NORMAL; } =20 /* @@ -2809,8 +2809,9 @@ int vfio_ap_mdev_probe_queue(struct ap_device *apdev) =20 q->apqn =3D apqn; q->saved_isc =3D VFIO_AP_ISC_INVALID; - memset(&q->reset_status, 0, sizeof(q->reset_status)); INIT_WORK(&q->reset_work, apq_reset_check); + vfio_ap_mdev_reset_queue(q); + flush_work(&q->reset_work); =20 if (matrix_mdev) { vfio_ap_mdev_link_queue(matrix_mdev, q); @@ -2902,8 +2903,8 @@ void vfio_ap_mdev_remove_queue(struct ap_device *apde= v) vfio_ap_unlink_queue_fr_mdev(q); =20 dev_set_drvdata(&apdev->device, NULL); - kfree(q); release_update_locks_for_mdev(matrix_mdev); + kfree(q); } =20 /** @@ -3309,3 +3310,46 @@ void vfio_ap_on_scan_complete(struct ap_config_info = *new_config_info, =20 mutex_unlock(&matrix_dev->guests_lock); } + +/** + * vfio_ap_on_qstate_transition: + * + * AP bus callback notifying the vfio_ap device driver that the state of a + * queue has transitioned. + * + * @qstate_trans: the object containing a reference to the queue device an= d the + * state to which it transitioned. + */ +void vfio_ap_on_qstate_transition(struct ap_qstate_transition *qstate_tran= s) +{ + struct vfio_ap_queue *q =3D vfio_ap_find_queue(qstate_trans->queue->qid); + DECLARE_BITMAP(apm_filtered, AP_DEVICES); + + /* + * If the queue is not bound to the vfio_ap device driver, then it won't + * be passed through to a guest; so, no need to continue. + */ + if (!q) + return; + + get_update_locks_for_mdev(q->matrix_mdev); + + switch (qstate_trans->new_state) { + case AP_QUEUE_CONFIG_ON: + case AP_QUEUE_CHKSTOP_OFF: + vfio_ap_mdev_reset_queue(q); + flush_work(&q->reset_work); + + if (q->matrix_mdev) { + if (vfio_ap_mdev_filter_matrix(q->matrix_mdev, apm_filtered)) { + vfio_ap_mdev_update_guest_apcb(q->matrix_mdev); + reset_queues_for_apids(q->matrix_mdev, apm_filtered); + } + } + break; + default: + break; + } + + release_update_locks_for_mdev(q->matrix_mdev); +} diff --git a/drivers/s390/crypto/vfio_ap_private.h b/drivers/s390/crypto/vf= io_ap_private.h index 9bff666b0b35..c8b00465d258 100644 --- a/drivers/s390/crypto/vfio_ap_private.h +++ b/drivers/s390/crypto/vfio_ap_private.h @@ -165,4 +165,6 @@ void vfio_ap_on_cfg_changed(struct ap_config_info *new_= config_info, void vfio_ap_on_scan_complete(struct ap_config_info *new_config_info, struct ap_config_info *old_config_info); =20 +void vfio_ap_on_qstate_transition(struct ap_qstate_transition *qstate_tran= s); + #endif /* _VFIO_AP_PRIVATE_H_ */ --=20 2.53.0 From nobody Sat Sep 26 03:50:07 2026 Received: from mx0b-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BA0A549E141; Fri, 25 Sep 2026 12:46:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=148.163.158.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340385; cv=none; b=nReI5kvptUO+BIy1kklnDocVUZSlhukTU+qGrYDKga2g7v7MlCrp4BayNPvS8jGX6rocOwLxWIqbRi6dISjwzqaZ5dvdibLPTihcR8e9msM6vGnWQts5SF4wDtIBoPysic7KmRBI64RJCwT+i1+G8moXNc4atA4EZgXK6T+KOAQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790340385; c=relaxed/simple; bh=XQsChNS2X/278RaM72O+Vww+5glPuUrq+JG+2Bfg2Co=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=g7JdFeALA1cy1t17n/lXI3bi2OWUZtaX/F/y0fW+XGL5glcKUoDPtKo35L703PJ8R2LnSPl9ZcjofgfivDy2xGMGCxLIAE3m3fntO3g23jHquMYUvRNc4t5TMvoevCxhaKOnwdCN8p3BmMe/7Ul/+tz9879eBXgSDasZDJH9r9U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com; spf=pass smtp.mailfrom=linux.ibm.com; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b=pbPFJUHC; arc=none smtp.client-ip=148.163.158.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.ibm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ibm.com header.i=@ibm.com header.b="pbPFJUHC" Received: from pps.filterd (m0360072.ppops.net [127.0.0.1]) by mx0a-001b2d01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 68P4aiZU103383; Fri, 25 Sep 2026 12:46:04 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=cc :content-transfer-encoding:date:from:in-reply-to:message-id :mime-version:references:subject:to; s=pp1; bh=teAkEoOojljyG8oZ9 JAqfRA4CmODV2YS2q0KV2Fkr7U=; b=pbPFJUHCirQJ79vmfc+Ht0Z+A9S+lKqhg /wrflnxu3m3RScDBn2qP72kSkan6XiQqzwbWVHXEkSdWvemWtR2ioRK5w7PC6HMf pqvome7F60OHZOGhhU00cetsz9P7xGt7XEUvHsfB1Q9WfqKY82XdXstPvOy3qbUE AewtwCRP6CGv6HCNKf7baLOmWEAzAkxhENa941P68xlgHNlwjMK4ngLRA225HaQN KjzsvfTOQj+0tR4dTEMR44eoIx5G4sJMkinnvC7+NT12J0FOFHNIOHDSvjKsFKQB PNFke4uNcnvX2ZufSojLAYocp2F5aBmJ7pxBjdEaYCvAOiAoz9JPA== Received: from ppma12.dal12v.mail.ibm.com (dc.9e.1632.ip4.static.sl-reverse.com [50.22.158.220]) by mx0a-001b2d01.pphosted.com (PPS) with ESMTPS id 4gskdvp2su-1 (version=TLSv1.3 cipher=TLS_AES_256_GCM_SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:46:04 +0000 (GMT) Received: from pps.filterd (ppma12.dal12v.mail.ibm.com [127.0.0.1]) by ppma12.dal12v.mail.ibm.com (8.18.1.11/8.18.1.11) with ESMTP id 68PA1cSu3298414; Fri, 25 Sep 2026 12:46:03 GMT Received: from smtprelay01.wdc07v.mail.ibm.com ([172.16.1.68]) by ppma12.dal12v.mail.ibm.com (PPS) with ESMTPS id 4gvb8k2e8b-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Fri, 25 Sep 2026 12:46:03 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (smtpav04.dal12v.mail.ibm.com [10.241.53.103]) by smtprelay01.wdc07v.mail.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 68PCk2w41639260 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Fri, 25 Sep 2026 12:46:02 GMT Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id E2A4758063; Fri, 25 Sep 2026 12:46:01 +0000 (GMT) Received: from smtpav04.dal12v.mail.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id A816A5805A; Fri, 25 Sep 2026 12:46:00 +0000 (GMT) Received: from li-4c4c4544-004d-4810-8043-b7c04f423534.ibm.com.com (unknown [9.61.86.59]) by smtpav04.dal12v.mail.ibm.com (Postfix) with ESMTP; Fri, 25 Sep 2026 12:46:00 +0000 (GMT) From: Anthony Krowiak To: linux-s390@vger.kernel.org, linux-kernel@vger.kernel.org, kvm@vger.kernel.org Cc: jjherne@linux.ibm.com, borntraeger@de.ibm.com, mjrosato@linux.ibm.com, pasic@linux.ibm.com, alex@shazbot.org, kwankhede@nvidia.com, fiuczy@linux.ibm.com, pbonzini@redhat.com, frankja@linux.ibm.com, imbrenda@linux.ibm.com, agordeev@linux.ibm.com, hca@linux.ibm.com, gor@linux.ibm.com, freude@linux.ibm.com Subject: [PATCH v8 6/6] s390/vfio-ap: replace guest-reachable WARNs with ratelimited warnings Date: Fri, 25 Sep 2026 08:45:51 -0400 Message-ID: <20260925124551.665448-7-akrowiak@linux.ibm.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260925124551.665448-1-akrowiak@linux.ibm.com> References: <20260925124551.665448-1-akrowiak@linux.ibm.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-TM-AS-GCONF: 00 X-Proofpoint-ORIG-GUID: 16IOTR6tJGWphnoqp6-kmJ_W6W4EaowY X-Proofpoint-Spam-Info: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfXwXn+NaBK5Hnj yoUdussCobp/LV3f7oeP/wBDcVhFbtuNlD/f91GNFjeI4TGuleFd32A832GBWCNKvFzD02/Z8uC DAPGbMjhkFO2IOYmZi79ooKalvaZjUU= X-Authority-Analysis: v=2.4 cv=FLiOVOos c=1 sm=1 tr=0 ts=6ab66d0c cx=c_pps a=bLidbwmWQ0KltjZqbj+ezA==:117 a=bLidbwmWQ0KltjZqbj+ezA==:17 a=VdqzKS8jKosA:10 a=VkNPw1HP01LnGYTKEx00:22 a=RnoormkPH1_aCDwRdu11:22 a=RzCfie-kr_QcCd8fBx8p:22 a=VnNF1IyMAAAA:8 a=SMWzy2rGiXtFoilLqd4A:9 X-Proofpoint-GUID: 16IOTR6tJGWphnoqp6-kmJ_W6W4EaowY X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwOTI1MDA1MCBTYWx0ZWRfX9HgtoYLTUV0o pOrbV7Qyy+u0X5TogOYheB8okkBjMT+lk/4x1X1cmZAfcvytKLTkHYpTKsK2Hhu1GgP979FhIWa ypTg3jlD7gRcQ9WMIflQ/zyXBCZBEvWiBGSegLVgzXNZx7+bd+4PPbCqd1/jXC92b24f5sxgIrG M/UxEJXM0FELnA829GAH758xL4tBcuRhde1kTyp/7ptbC/kaMG+fRjUp2/InTsYVscpKCW9CzwO w9Ea21oTqb9wS85TtozTbpmi91wI/msf4+85+1RLcVdoj44QRxbsmEvbWREtVSoretWs3yyyZ1i qCwqNJfQixdCm2xnDeChqsPlYT72JyJ2UxZDfYil60N0+i/YD/hLlkdh6whYfRkxI1Vu/LbMgUz IXknOKkTncF7UxqI+g8uW1mLdmEd48v56ya/GZh8YdOivCCYqC5O82MYDbuRMGSRoiyki9/sU69 BJa6urgHO0cJFy75Kmg== X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-09-25_02,2026-09-21_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 adultscore=0 impostorscore=0 phishscore=0 spamscore=0 clxscore=1015 suspectscore=0 bulkscore=0 lowpriorityscore=0 priorityscore=1501 malwarescore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2609040000 definitions=main-2609250050 Content-Type: text/plain; charset="utf-8" WARN and WARN_ONCE macros in code paths reachable by a guest can be triggered repeatedly by a malicious or misbehaving guest, flooding the kernel log and potentially impacting system stability. Replace all WARN and WARN_ONCE calls reachable from the guest AP interrupt enable/disable and queue reset paths with ratelimited warning functions. When the queue is assigned to an mdev, dev_warn_ratelimited() is used so the mdev device name (which includes the UUID) appears in the message. Otherwise, pr_warn_ratelimited() is used. Five reporting functions are introduced: report_tapq_rc() - reports an invalid or unexpected response code from PQAP(TAPQ). Used in vfio_ap_wait_for_irqclear() and apq_status_check(). The signatures of both functions are changed to accept a struct vfio_ap_queue pointer instead of an apqn so the queue's mdev context is available for reporting. report_irqclear_timeout() - reports a timeout waiting for the IR bit to clear after a PQAP(AQIC) disable in vfio_ap_wait_for_irqclear(). report_aqic_disable_error() - reports a failed PQAP(AQIC) disable operation in vfio_ap_irq_disable(). Replaces three WARN_ONCE calls covering the non-operational queue, rejected disable, and retry exhaustion cases. report_zapq_rc() - reports an invalid response code from PQAP(ZAPQ) in vfio_ap_mdev_reset_queue(). report_gisc_unregister_failure() - reports a failure to unregister the guest ISC due to the fact that q->matrix_mdev or q->matrix_mdev->kvm is NULL. The VFIO_AP_DBF_WARN() calls in vfio_ap_irq_enable() and handle_pqap() are retained but augmented with companion dev_warn_ratelimited() or pr_warn_ratelimited() calls. VFIO_AP_DBF_WARN() writes only to the s390 debug feature ring buffer, which requires a sysadmin to know to look in /sys/kernel/debug/s390dbf/ to find the messages. The companion dmesg log entries ensure that warning conditions are immediately visible in the kernel log without requiring familiarity with the s390 debug feature infrastructure. Signed-off-by: Anthony Krowiak --- drivers/s390/crypto/vfio_ap_ops.c | 186 +++++++++++++++++++++++------- 1 file changed, 142 insertions(+), 44 deletions(-) diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_a= p_ops.c index cd4a436c4319..4b6e64daed25 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -226,6 +226,58 @@ static struct vfio_ap_queue *vfio_ap_mdev_get_queue( return NULL; } =20 +static void report_tapq_rc(struct vfio_ap_queue *q, u8 rc) +{ + if (q->matrix_mdev) + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(TAPQ) for %02x.%04x failed with invalid rc=3D%#02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), rc); + else + pr_warn_ratelimited("PQAP(TAPQ) for %02x.%04x failed with invalid rc=3D%= #02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), rc); +} + +static void report_irqclear_timeout(struct vfio_ap_queue *q, u8 rc) +{ + if (q->matrix_mdev) + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(TAPQ) timed out waiting for IRQ clear on %02x.%04x: rc=3D%#= 02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), rc); + else + pr_warn_ratelimited("PQAP(TAPQ) timed out waiting for IRQ clear on %02x.= %04x: rc=3D%#02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), rc); +} + +static void report_aqic_disable_error(struct vfio_ap_queue *q, u8 rc) +{ + if (q->matrix_mdev) + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(AQIC) disable for %02x.%04x failed with rc=3D%#02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), rc); + else + pr_warn_ratelimited("PQAP(AQIC) disable for %02x.%04x failed with rc=3D%= #02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), rc); +} + +static void report_zapq_rc(struct vfio_ap_queue *q, u8 rc) +{ + if (q->matrix_mdev) + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(ZAPQ) for %02x.%04x failed with invalid rc=3D%#02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), rc); + else + pr_warn_ratelimited("PQAP(ZAPQ) for %02x.%04x failed with invalid rc=3D%= #02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), rc); +} + /** * vfio_ap_wait_for_irqclear - wait for the IR bit to clear after a disable * @@ -256,14 +308,16 @@ static struct vfio_ap_queue *vfio_ap_mdev_get_queue( * * -EIO PQAP-TAPQ returned an invalid response code */ -static int vfio_ap_wait_for_irqclear(int apqn, struct ap_queue_status *tap= q_status) +static int vfio_ap_wait_for_irqclear(struct vfio_ap_queue *q, + struct ap_queue_status *tapq_status) { struct ap_queue_status status; int retry =3D 5; =20 do { - status =3D ap_tapq(apqn, NULL); + status =3D ap_tapq(q->apqn, NULL); memcpy(tapq_status, &status, sizeof(status)); + switch (status.response_code) { case AP_RESPONSE_NORMAL: case AP_RESPONSE_RESET_IN_PROGRESS: @@ -276,8 +330,6 @@ static int vfio_ap_wait_for_irqclear(int apqn, struct a= p_queue_status *tapq_stat case AP_RESPONSE_Q_NOT_AVAIL: case AP_RESPONSE_DECONFIGURED: case AP_RESPONSE_CHECKSTOPPED: - WARN_ONCE(1, "%s: tapq rc %02x: %04x\n", __func__, - status.response_code, apqn); return -ENODEV; case AP_RESPONSE_ASSOC_SECRET_NOT_UNIQUE: case AP_RESPONSE_ASSOC_FAILED: @@ -289,23 +341,53 @@ static int vfio_ap_wait_for_irqclear(int apqn, struct= ap_queue_status *tapq_stat * since that would happen anyway if we continued to * execute the TAPQ. */ - WARN_ONCE(1, "%s: tapq rc %02x: %04x\n", __func__, - status.response_code, apqn); + report_tapq_rc(q, status.response_code); return -ETIMEDOUT; default: - WARN_ONCE(1, "%s: tapq rc %02x: %04x\n", __func__, - status.response_code, apqn); + report_tapq_rc(q, status.response_code); return -EIO; } } while (--retry); =20 - WARN_ONCE(1, "%s: tapq rc %02x: timed out waiting for interrupts disabled= for %02x.%04x\n", - __func__, status.response_code, - AP_QID_CARD(apqn), AP_QID_QUEUE(apqn)); + report_irqclear_timeout(q, status.response_code); =20 return -ETIMEDOUT; } =20 +static void report_gisc_unregister_failure(struct vfio_ap_queue *q) +{ + if (q->matrix_mdev) { + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "APQN %02x.%04x: Failed to unregister guest ISC %c\n", + AP_QID_CARD(q->apqn), AP_QID_QUEUE(q->apqn), + q->saved_isc); + } else { + pr_warn_ratelimited("APQN %02x.%04x: Failed to unregister guest ISC %c\n= ", + AP_QID_CARD(q->apqn), AP_QID_QUEUE(q->apqn), + q->saved_isc); + } +} + +static bool verify_free_aqic_resourcers(struct vfio_ap_queue *q) +{ + bool verified =3D true; + + if (q->saved_isc !=3D VFIO_AP_ISC_INVALID && + !(q->matrix_mdev && q->matrix_mdev->kvm)) { + report_gisc_unregister_failure(q); + verified =3D false; + } + + if (q->saved_iova && !q->matrix_mdev) { + pr_warn_ratelimited("APQN %02x.%04x: Failed to unpin NIB page %08x\n", + AP_QID_CARD(q->apqn), AP_QID_QUEUE(q->apqn), + q->saved_iova); + verified =3D false; + } + + return verified; +} + /** * vfio_ap_free_aqic_resources - free vfio_ap_queue resources * @q: The vfio_ap_queue @@ -316,17 +398,13 @@ static int vfio_ap_wait_for_irqclear(int apqn, struct= ap_queue_status *tapq_stat */ static void vfio_ap_free_aqic_resources(struct vfio_ap_queue *q) { - if (!q) + if (!q || !verify_free_aqic_resourcers(q)) return; - if (q->saved_isc !=3D VFIO_AP_ISC_INVALID && - !WARN_ON(!(q->matrix_mdev && q->matrix_mdev->kvm))) { - kvm_s390_gisc_unregister(q->matrix_mdev->kvm, q->saved_isc); - q->saved_isc =3D VFIO_AP_ISC_INVALID; - } - if (q->saved_iova && !WARN_ON(!q->matrix_mdev)) { - vfio_unpin_pages(&q->matrix_mdev->vdev, q->saved_iova, 1); - q->saved_iova =3D 0; - } + + kvm_s390_gisc_unregister(q->matrix_mdev->kvm, q->saved_isc); + q->saved_isc =3D VFIO_AP_ISC_INVALID; + vfio_unpin_pages(&q->matrix_mdev->vdev, q->saved_iova, 1); + q->saved_iova =3D 0; } =20 /** @@ -373,7 +451,8 @@ static struct ap_queue_status vfio_ap_irq_disable(struc= t vfio_ap_queue *q) * wait until interrupt processing has been disabled * before proceeding. */ - ret =3D vfio_ap_wait_for_irqclear(q->apqn, &tapq_status); + ret =3D vfio_ap_wait_for_irqclear(q, &tapq_status); + if (ret =3D=3D 0 || ret =3D=3D -ENODEV) goto end_free; =20 @@ -420,8 +499,7 @@ static struct ap_queue_status vfio_ap_irq_disable(struc= t vfio_ap_queue *q) case AP_RESPONSE_DECONFIGURED: case AP_RESPONSE_CHECKSTOPPED: /* AP not operational; no further interrupts possible */ - WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, - status.response_code); + report_aqic_disable_error(q, status.response_code); goto end_free; case AP_RESPONSE_INVALID_ADDRESS: case AP_RESPONSE_INVALID_GISA: @@ -433,14 +511,12 @@ static struct ap_queue_status vfio_ap_irq_disable(str= uct vfio_ap_queue *q) * and the hardware still holds the NIB address. Do not * free resources. */ - WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, - status.response_code); + report_aqic_disable_error(q, status.response_code); goto end_fail; } } while (retries--); =20 - WARN_ONCE(1, "%s: ap_aqic status %d\n", __func__, - status.response_code); + report_aqic_disable_error(q, status.response_code); =20 end_fail: /* @@ -569,7 +645,10 @@ static struct ap_queue_status vfio_ap_irq_enable(struc= t vfio_ap_queue *q, if (vfio_ap_validate_nib(vcpu, &nib)) { VFIO_AP_DBF_WARN("%s: invalid NIB address: nib=3D%pad, apqn=3D%#04x\n", __func__, &nib, q->apqn); - + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(AQIC) enable for %02x.%04x: invalid NIB address %pad\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), &nib); status.response_code =3D AP_RESPONSE_INVALID_ADDRESS; return status; } @@ -584,7 +663,10 @@ static struct ap_queue_status vfio_ap_irq_enable(struc= t vfio_ap_queue *q, VFIO_AP_DBF_WARN("%s: vfio_pin_pages failed: rc=3D%d," "nib=3D%pad, apqn=3D%#04x\n", __func__, ret, &nib, q->apqn); - + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(AQIC) enable for %02x.%04x: vfio_pin_pages failed rc=3D%d\n= ", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), ret); status.response_code =3D AP_RESPONSE_INVALID_ADDRESS; return status; } @@ -607,7 +689,10 @@ static struct ap_queue_status vfio_ap_irq_enable(struc= t vfio_ap_queue *q, if (nisc < 0) { VFIO_AP_DBF_WARN("%s: gisc registration failed: nisc=3D%d, isc=3D%d, apq= n=3D%#04x\n", __func__, nisc, isc, q->apqn); - + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(AQIC) enable for %02x.%04x: GISC registration failed rc=3D%= d isc=3D%d\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), nisc, isc); vfio_unpin_pages(&q->matrix_mdev->vdev, nib, 1); status.response_code =3D AP_RESPONSE_INVALID_ADDRESS; return status; @@ -651,9 +736,14 @@ static struct ap_queue_status vfio_ap_irq_enable(struc= t vfio_ap_queue *q, * ISC that were prepared for this (rejected) request. */ ret =3D kvm_s390_gisc_unregister(kvm, isc); - if (ret) + if (ret) { VFIO_AP_DBF_WARN("%s: kvm_s390_gisc_unregister: rc=3D%d isc=3D%d, apqn= =3D%#04x\n", __func__, ret, isc, q->apqn); + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(AQIC) enable for %02x.%04x: GISC unregister failed rc=3D%d= isc=3D%d\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), ret, isc); + } vfio_unpin_pages(&q->matrix_mdev->vdev, nib, 1); break; } @@ -667,6 +757,11 @@ static struct ap_queue_status vfio_ap_irq_enable(struc= t vfio_ap_queue *q, aqic_gisa.zone, aqic_gisa.ir, aqic_gisa.gisc, aqic_gisa.gf, aqic_gisa.gisa, aqic_gisa.isc, q->apqn); + dev_warn_ratelimited(mdev_dev(q->matrix_mdev->mdev), + "PQAP(AQIC) enable for %02x.%04x failed with rc=3D%#02x\n", + AP_QID_CARD(q->apqn), + AP_QID_QUEUE(q->apqn), + status.response_code); } =20 return status; @@ -751,7 +846,8 @@ static int handle_pqap(struct kvm_vcpu *vcpu) if (!(vcpu->arch.sie_block->eca & ECA_AIV)) { VFIO_AP_DBF_WARN("%s: AIV facility not installed: apqn=3D0x%04x, eca=3D0= x%04x\n", __func__, apqn, vcpu->arch.sie_block->eca); - + pr_warn_ratelimited("PQAP(AQIC) for %02x.%04x: AIV facility not installe= d\n", + AP_QID_CARD(apqn), AP_QID_QUEUE(apqn)); return -EOPNOTSUPP; } =20 @@ -760,7 +856,8 @@ static int handle_pqap(struct kvm_vcpu *vcpu) if (!vcpu->kvm->arch.crypto.pqap_hook) { VFIO_AP_DBF_WARN("%s: PQAP(AQIC) hook not registered with the vfio_ap dr= iver: apqn=3D0x%04x\n", __func__, apqn); - + pr_warn_ratelimited("PQAP(AQIC) for %02x.%04x: hook not registered with = the vfio_ap driver\n", + AP_QID_CARD(apqn), AP_QID_QUEUE(apqn)); goto out_unlock; } =20 @@ -773,6 +870,9 @@ static int handle_pqap(struct kvm_vcpu *vcpu) VFIO_AP_DBF_WARN("%s: mdev %08lx-%04lx-%04lx-%04lx-%04lx%08lx not in use= : apqn=3D0x%04x\n", __func__, uuid[0], uuid[1], uuid[2], uuid[3], uuid[4], uuid[5], apqn); + dev_warn_ratelimited(mdev_dev(matrix_mdev->mdev), + "PQAP(AQIC) for %02x.%04x: mdev not in use\n", + AP_QID_CARD(apqn), AP_QID_QUEUE(apqn)); goto out_unlock; } =20 @@ -781,6 +881,9 @@ static int handle_pqap(struct kvm_vcpu *vcpu) VFIO_AP_DBF_WARN("%s: Queue %02x.%04x not bound to the vfio_ap driver\n", __func__, AP_QID_CARD(apqn), AP_QID_QUEUE(apqn)); + dev_warn_ratelimited(mdev_dev(matrix_mdev->mdev), + "PQAP(AQIC) for %02x.%04x: queue not bound to the vfio_ap driver\= n", + AP_QID_CARD(apqn), AP_QID_QUEUE(apqn)); goto out_unlock; } =20 @@ -2084,7 +2187,8 @@ static struct vfio_ap_queue *vfio_ap_find_queue(int a= pqn) return q; } =20 -static int apq_status_check(int apqn, struct ap_queue_status *status) +static int apq_status_check(struct vfio_ap_queue *q, + struct ap_queue_status *status) { switch (status->response_code) { case AP_RESPONSE_NORMAL: @@ -2146,10 +2250,7 @@ static int apq_status_check(int apqn, struct ap_queu= e_status *status) return -EAGAIN; =20 default: - WARN(true, - "failed to verify reset of queue %02x.%04x: TAPQ rc=3D%u\n", - AP_QID_CARD(apqn), AP_QID_QUEUE(apqn), - status->response_code); + report_tapq_rc(q, status->response_code); return -EIO; } } @@ -2219,7 +2320,7 @@ static void apq_reset_check(struct work_struct *reset= _work) msleep(AP_RESET_INTERVAL); elapsed +=3D AP_RESET_INTERVAL; status =3D ap_tapq(q->apqn, NULL); - ret =3D apq_status_check(q->apqn, &status); + ret =3D apq_status_check(q, &status); if (ret =3D=3D -EIO) { /* * TAPQ returned an invalid response code indicating a @@ -2343,10 +2444,7 @@ static void vfio_ap_mdev_reset_queue(struct vfio_ap_= queue *q) * would corrupt the new owner's memory which could crash or * compromise the host kernel. */ - WARN(true, - "PQAP/ZAPQ for %02x.%04x failed with invalid rc=3D%u\n", - AP_QID_CARD(q->apqn), AP_QID_QUEUE(q->apqn), - status.response_code); + report_zapq_rc(q, status.response_code); } } =20 --=20 2.53.0