From nobody Thu Sep 24 13:37:02 2026 Received: from mail-ej2-f12.google.com (mail-ej2-f12.google.com [74.125.228.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7928F53B322 for ; Wed, 23 Sep 2026 18:40:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.228.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790188822; cv=none; b=qCZovypkAaahwu/v/VKG3pYQPYGmROGJY1E4I1I/P9doL5bu6KxRL9JITq0gNADWIH96H1sIRya0LxHfHM75JFCBWxHP8EdRg+w2/1y2YiHIEhS7NYeiGhQoL8JaTJZNoLFW+yrsnaGF2gcK9dwjoIlcL5Q+1nQT5Asdu9S7hbI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790188822; c=relaxed/simple; bh=EvUQuL+9sEfu/yTDDsYynyczprpLsvtenWv1NJgsNFM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lx27+Ng+ETzlkutbJCWcRAbvxNL76anwuKUdSIWlPelchVmIiTftZWOFyEpysDEERUt2hKMN27+Z/53KqurTxFCWEDzsCL4ojQbes1+nI5Lj3Xa394ZkdZKO50VNv1IgPWykuQZVCxhYgOnchAjtCMkmOmpdW2B4u9J5Ef0M7Zw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=SXU0bD/a; arc=none smtp.client-ip=74.125.228.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="SXU0bD/a" Received: by mail-ej2-f12.google.com with SMTP id a640c23a62f3a-c254fa650c8so217702666b.0 for ; Wed, 23 Sep 2026 11:40:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790188812; x=1790793612; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JlVZpnAHAmHOvMabqqtARGu8SxbZU7qizTP5GLSJRM8=; b=SXU0bD/aaPLax4TXYMVFvLchdTACOZD3WoXSwdY7ltH24VEYxKen7mvixOWF0jzSJX 9P7aV6FJ9wOw4bt5pc338ZuTxUVw3avSvyEtaffVlW9Ahlsfmpp7wQWijsteVO7Nge+p DeWPfHA2Vh+Ev9tPZZOVtMEfVZ7niW+U4PfxU41tO3dECif6Mcw+iDLMhrZIwSgQdfKd zQZ9szqf8hKcX5MW5afpWAH+bUen2gnUc/NuUupTSyVM7nxBw9dCB+RfhzDMw8v+ecjX Q2iWHbDq67Nn0CJOkgg7UXz3ckGIsUVtaQSmD2HBWV+PW7aOTZ30KmWnN0RBcwhkHcqH h92A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790188812; x=1790793612; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=JlVZpnAHAmHOvMabqqtARGu8SxbZU7qizTP5GLSJRM8=; b=UqcIYe0xVuoJ+vXyeCZDHBtw+Pns4iT0s8/E4pOyRuET6jkCUFd3UkCUe+4m0uja5a HfsySyS+sMCqLbPtX3W/X+uFg4vuumSaBoZqw4aZIUkXCUdQ8pZ3h0wqwZ9swEWwr+DM a9mTnILKh3mVmzD03QQ/j/JkJguhmHwxHts9+NHQTdjiH3zJXNYho6xsftm2KMRl/BwF iFP8lspVURzsZLRqRjmSjHXPEiW8j8lTKsNSW1MZsaWF/xBKz42UuY9o5JAhHikM0RBu XJqewWCMFdo+tAGwZhDwXGNuBI6U2T4UhzQ6eTiu9hZNcSfwGgDtxSJ0fUZ0uImnksfS BoCw== X-Forwarded-Encrypted: i=1; AKwUvByxHrNS+W3B04QBy4j1d5q6c10L4lrRljdA+tGW+ykmrJFKuzVPuxRRMK1UnM7ejy4EeIxb+kfKwFYxMbw=@vger.kernel.org X-Gm-Message-State: AFuF++mi0Z+EF10n4bqKegl5OyC5x9CwdvQeteiVVAneOqQHnjEua4FJ akU2uGcnSpo8xRZyTUq9E9bxcOxwFLjB8uHhTqj2VZrsHqqmfIlvoEiJ X-Gm-Gg: AYBFou2XQ2lvxb+06x86GllqxVJBqPOkkMmOfP7azfg4sL9ffTVQRsUfV35qd7HcOvu C8HA6SWyLlChHymNrD7GxbOs6HiG9UZ2eriauEzIvFrIQxpp0E2zkQs0e/+R7bp47r6+5HUXYlc bHQekkBpiI3Zo9cIBrlR9rgoTKS4Ph3b4fQMOjzmcd5zz5r+RfIjFz/xIbPEW5Augmm5tJ3nRsm RmA3AJe6ZAXUlDXK0Jc6Oj++OatrQVaFAiQyA9J6kpZluYHAdMBVUQkW5fCkQwLWbfNhskJyKlc L7NESuyprUdtyj/8ezL4gjMqaB2sW/REWpk99B/KWn19OF1RRVOZiov9DQUEe0pA4esKEt/ub+k A5qOiIeqz+ibexXYn8hdI0ZIxF1JyAIkW4PQvodQLxbZ/XSWA43HP9Ex1oEvhYBzCHerEk/UnJi UCDSpxdGlMYgPRzR8lOBUzFCruyGIGF7K6e7rjRbUjZpXD+ptJ5u0hAlpjZhA/k6boWxtruCbhD CdofVHUyPMmKrxL2U+HGksC0sElaNSVmQnhcDUrb6mte8zW2HK/ZDQZefI= X-Received: by 2002:a17:907:c71b:b0:c29:53eb:9913 with SMTP id a640c23a62f3a-c2ac2531dafmr836366b.38.1790188812262; Wed, 23 Sep 2026 11:40:12 -0700 (PDT) Received: from dohko.chello.ie (188-141-5-72.dynamic.upc.ie. [188.141.5.72]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c2aae33faf4sm173417566b.3.2026.09.23.11.40.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 23 Sep 2026 11:40:11 -0700 (PDT) From: David Carlier To: netdev@vger.kernel.org Cc: sgarzare@redhat.com, bobbyeshleman@gmail.com, davem@davemloft.net, edumazet@google.com, kuba@kernel.org, pabeni@redhat.com, horms@kernel.org, virtualization@lists.linux.dev, linux-kernel@vger.kernel.org, David Carlier Subject: [PATCH net-next v2 1/2] vsock: report pending receive data to io_uring Date: Wed, 23 Sep 2026 19:40:07 +0100 Message-ID: <20260923184008.153541-2-devnexen@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260923184008.153541-1-devnexen@gmail.com> References: <20260923184008.153541-1-devnexen@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" AF_VSOCK stream receives never fill msghdr.msg_inq, so io_uring cannot set IORING_CQE_F_SOCK_NONEMPTY and retries a multishot receive even after the queue has been drained. Fill the hint at the common receive exit using the transport callback that SIOCINQ already uses, and report 1 once the connection is finished so the caller performs the receive which observes EOF, as TCP does after a FIN. A vsock loopback ping-pong with io_uring multishot receive drops entries into __vsock_connectible_recvmsg from 1.97 to 1.00 per delivered message, and receiver CPU time by about 3% (25 runs of 50000 messages, p=3D0.006). Signed-off-by: David Carlier --- net/vmw_vsock/af_vsock.c | 35 +++++++++++++++++++++++++++++++++++ 1 file changed, 35 insertions(+) diff --git a/net/vmw_vsock/af_vsock.c b/net/vmw_vsock/af_vsock.c index f840498b58af..20d6f9ca6a96 100644 --- a/net/vmw_vsock/af_vsock.c +++ b/net/vmw_vsock/af_vsock.c @@ -2543,6 +2543,35 @@ static int __vsock_seqpacket_recvmsg(struct sock *sk= , struct msghdr *msg, return err; } =20 +/* Bytes a following receive can consume, 1 if it would only see EOF, or -1 + * if the transport cannot tell. + * + * Called under the socket lock after a nonnegative stream receive, so a N= ULL + * transport implies SOCK_DONE. + */ +static int vsock_stream_inq_hint(struct sock *sk) +{ + struct vsock_sock *vsk =3D vsock_sk(sk); + s64 data; + + if ((sk->sk_shutdown & RCV_SHUTDOWN) || !vsk->transport || + (sock_flag(sk, SOCK_DONE) && sk->sk_state !=3D TCP_ESTABLISHED)) + return 1; + + data =3D vsock_stream_has_data(vsk); + if (data < 0) + return -1; + if (data > 0) + return min_t(s64, data, INT_MAX); + + /* Empty but finished: keep the caller reading so it sees EOF. */ + if (sock_flag(sk, SOCK_DONE) || + (READ_ONCE(vsk->peer_shutdown) & SEND_SHUTDOWN)) + return 1; + + return 0; +} + int __vsock_connectible_recvmsg(struct socket *sock, struct msghdr *msg, size_= t len, int flags) @@ -2606,6 +2635,12 @@ __vsock_connectible_recvmsg(struct socket *sock, str= uct msghdr *msg, size_t len, err =3D __vsock_seqpacket_recvmsg(sk, msg, len, flags); =20 out: + /* Seqpacket has_data counts messages, while io_uring treats msg_inq as + * a byte length when sizing retries, so only streams report a hint. + */ + if (msg->msg_get_inq && err >=3D 0 && sk->sk_type =3D=3D SOCK_STREAM) + msg->msg_inq =3D vsock_stream_inq_hint(sk); + release_sock(sk); return err; } --=20 2.55.0 From nobody Thu Sep 24 13:37:02 2026 Received: from mail-ej2-f12.google.com (mail-ej2-f12.google.com [74.125.228.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 84C9130C14A for ; Wed, 23 Sep 2026 18:40:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.228.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790188825; cv=none; b=mvt82UgsReiNgURUsUMMzPSa9OwETZOfm/qQHvLsIEbmMKut4HINdNGZUvjqHAyONyEbMKL3ITGfdFj/1nZGhxEiwIBTDdd0T1grp9ahQoOa5YfMpZUp92mVnWyPZfKjODaKNMzKCwBImy2ctk2y01BOuXV2K0yF0tEwF/4quM8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790188825; c=relaxed/simple; bh=Sq+z3t/1z1Igec9KM0/mPKnoW3zwgdtF59SRgFR4y4Y=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=DXzu4NUzmoBSNa7QvAcu0Tj7YwN4XfDEfHP5Hnz39LukDtRKgV5vsuA93+qr82SyKSOMlpbBgOzuRCi2JBNT+6OUkEFFr2xsSWHX8BY1pCTMeUwbYZmQk4ego/Ufh1kWPq2jyOgEfQh23LzIde3m0PIymOwyU2OY9fR9cxuZyic= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=gFxv32ES; arc=none smtp.client-ip=74.125.228.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="gFxv32ES" Received: by mail-ej2-f12.google.com with SMTP id a640c23a62f3a-c254f9f0b1eso204729066b.1 for ; Wed, 23 Sep 2026 11:40:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790188814; x=1790793614; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=b2R5khTlQ02M5fyLgYzXlFKQ5MLnD7NKcAMSs5jaNT8=; b=gFxv32ESasDqjRaJMafweb6pKF2Zq9pjEvlOXJNoPt9uQhM9aVfQVAhQyLlVEv89o8 rj4UZ451eYI3u9AfkDjEStEjA7DfiBuwcoRUwzd7QUOAKYrKNyzEfE1AlZYprjsOsL5m UI6lV5H4rI8ffiI//GaszoX69t5fyTxwYf6L3RBfApDvP1tZNS0c8WNJl0iDZEkg/U09 VVW9Of+e/uMIIYt/caKO8Lpm+TlsS2Ki6y6bQuD1WR1Ogwtd4nqb7l0hdWAHx4htt64T dWuCcuqzbkxAJe2XdvwDLmeYNcMOn9HbIBJHed7ezs8t5jkpM2+bImmHa1Jl4Sqo0rEA ChoA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790188814; x=1790793614; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=b2R5khTlQ02M5fyLgYzXlFKQ5MLnD7NKcAMSs5jaNT8=; b=D9WTYAKytZhpPliJpgZDj6BSybAD8ZjDDQ5TSiGgYC2RRjHv5CyEaQvzxkAnij/Vge 1WGhlY0Gd1sVohKG8ONnG9JuUP0DSdtiEJKjgk56FQVLlxeBfntuTvRF16zEzah2xOtm vyCpTvPZ8GD8156f8i0J5rOsVxiBqF4FsgZYxNjhWNOWunQBFYaKo1ADiouYyUT0KKed kowqnt7WBP9wupd5C/d6nq1jjFuoT0S+cCTKE+liZi65f6OyoPmnLSzjDP8v1G1cZv66 rD8+CRew1pd30ctj9vrYjeHvp3jTimWTdCVgFPi3Kip0jRNKcESAuJNse2Mphxi0f3yl 2wFA== X-Forwarded-Encrypted: i=1; AKwUvBxacjt5wtYLwhWSXU08+/LxGrwJcQUqruJtGelmvOGp1L1GiC+9pL40rx9TlomZMl5g8DQjg0YgKpyfth4=@vger.kernel.org X-Gm-Message-State: AFuF++lhLNZ/BQvDvMbuNT+Lg6mB+ml+fPj4P3ozDPy/WjvW43lK4PGX RrfWXeqU4X6AjkOOozTeDsXhdoAhsszXCl3XE3rxNKuTkdTKn533YGcV X-Gm-Gg: AYBFou2vITgKK5TpvjID1UbscaOfAWLHpckkvx/Xpj9lQuvvy1nxxJF62GfiZ1/E+pq Rc72ckAuIJbBpQSLe85PEgBKvfdtwt+Fwm2bB8b0izaEh2pc5baqF3NPFZ14J+E7TtqMeB9JT+Q hw7ElxTZxjcUaZ+QF3Yov1kZDu1vwznuEtox2AiO15/FZupjs3nLE7d1n62XNXsg+nfAtHGriZq 5AdhjP3S9TgszdWKeeijK2FGvF5gKkuNJy7ljvukZiy/34FqRmDsghXGwFBAbevcJejIgnfeBw0 28+GgXeSqbHVCRLfgjLmnr8dP+CRpBt9oOujKHLIRRPE1KI4gHEzswHqkKfF+twxOAFEfC2LtV+ njD6kL7p91yoO1rA3oqS2cfMWJPdyQXyD3E0JEyQtFoZ0bQG7pmR2zX5dMnqHkm4RSMblrKkGM0 uM/Ba6cIbMdHBrgAWh2H4ES0/ee9HeWN6UoKk9zySsxfzgXwIBu/zIuZ8+h74BbYEtZuT42TdqJ PNF2CdXDsn+tbLZ4dWGFvGBFv2aMHW23cg+5LFp9na8KLO3FD63ghZDrD7Rgrj9UD/dNA== X-Received: by 2002:a17:907:3e1e:b0:c29:f5d8:9c6e with SMTP id a640c23a62f3a-c2ac2235d42mr4902366b.29.1790188813415; Wed, 23 Sep 2026 11:40:13 -0700 (PDT) Received: from dohko.chello.ie (188-141-5-72.dynamic.upc.ie. [188.141.5.72]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c2aae33faf4sm173417566b.3.2026.09.23.11.40.12 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 23 Sep 2026 11:40:13 -0700 (PDT) From: David Carlier To: netdev@vger.kernel.org Cc: sgarzare@redhat.com, bobbyeshleman@gmail.com, davem@davemloft.net, edumazet@google.com, kuba@kernel.org, pabeni@redhat.com, horms@kernel.org, virtualization@lists.linux.dev, linux-kernel@vger.kernel.org, David Carlier Subject: [PATCH net-next v2 2/2] vsock/test: cover receive queue hints Date: Wed, 23 Sep 2026 19:40:08 +0100 Message-ID: <20260923184008.153541-3-devnexen@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260923184008.153541-1-devnexen@gmail.com> References: <20260923184008.153541-1-devnexen@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add io_uring receive tests checking IORING_CQE_F_SOCK_NONEMPTY across a partial receive, draining while the peer stays connected, EOF, a nonblocking receive on an empty queue, a zero-length request, and a multishot receive with provided buffers. Signed-off-by: David Carlier Reviewed-by: Bobby Eshleman --- tools/testing/vsock/util.c | 9 + tools/testing/vsock/util.h | 1 + tools/testing/vsock/vsock_uring_test.c | 391 +++++++++++++++++++++++++ 3 files changed, 401 insertions(+) diff --git a/tools/testing/vsock/util.c b/tools/testing/vsock/util.c index fe316b02a590..299a4e8ecb63 100644 --- a/tools/testing/vsock/util.c +++ b/tools/testing/vsock/util.c @@ -483,6 +483,15 @@ void recv_byte(int fd, int expected_ret, int flags) } } =20 +void expect_res(int res, int expected, const char *what) +{ + if (res !=3D expected) { + fprintf(stderr, "%s: expected %d, got %d\n", what, expected, + res); + exit(EXIT_FAILURE); + } +} + /* Run test cases. The program terminates if a failure occurs. */ void run_tests(const struct test_case *test_cases, const struct test_opts *opts) diff --git a/tools/testing/vsock/util.h b/tools/testing/vsock/util.h index bf633cde82b0..9dabb547021e 100644 --- a/tools/testing/vsock/util.h +++ b/tools/testing/vsock/util.h @@ -94,6 +94,7 @@ void send_buf(int fd, const void *buf, size_t len, int fl= ags, void recv_buf(int fd, void *buf, size_t len, int flags, ssize_t expected_r= et); void send_byte(int fd, int expected_ret, int flags); void recv_byte(int fd, int expected_ret, int flags); +void expect_res(int res, int expected, const char *what); void run_tests(const struct test_case *test_cases, const struct test_opts *opts); void list_tests(const struct test_case *test_cases); diff --git a/tools/testing/vsock/vsock_uring_test.c b/tools/testing/vsock/v= sock_uring_test.c index 5c3078969659..c6cc51db7adc 100644 --- a/tools/testing/vsock/vsock_uring_test.c +++ b/tools/testing/vsock/vsock_uring_test.c @@ -13,7 +13,10 @@ #include #include #include +#include #include +#include +#include #include =20 #include "util.h" @@ -28,6 +31,10 @@ =20 #define VSOCK_TEST_DATA_MAX_IOV 3 =20 +#define HINT_CHUNK_SIZE 4096 +#define HINT_BUF_GROUP 1 +#define HINT_BUF_ENTRIES 4 + struct vsock_io_uring_test { /* Number of valid elements in 'vecs'. */ int vecs_cnt; @@ -211,6 +218,365 @@ void test_stream_uring_msg_zc_client(const struct tes= t_opts *opts) vsock_io_uring_client(opts, &test_data_array[i], true); } =20 +struct uring_inq_ctx { + struct io_uring ring; + int fd; +}; + +static void inq_server_init(struct uring_inq_ctx *ctx, + const struct test_opts *opts) +{ + ctx->fd =3D vsock_stream_accept(VMADDR_CID_ANY, opts->peer_port, NULL); + if (ctx->fd < 0) { + perror("accept"); + exit(EXIT_FAILURE); + } + + if (io_uring_queue_init(RING_ENTRIES_NUM, &ctx->ring, 0)) + error(1, errno, "io_uring_queue_init"); +} + +static void inq_server_exit(struct uring_inq_ctx *ctx) +{ + io_uring_queue_exit(&ctx->ring); + close(ctx->fd); +} + +/* Submit a single receive and report both its result and its CQE flags. */ +static int inq_recv(struct uring_inq_ctx *ctx, void *buf, size_t len, + int flags, unsigned int *cflags) +{ + struct io_uring_sqe *sqe; + struct io_uring_cqe *cqe; + int res; + + sqe =3D io_uring_get_sqe(&ctx->ring); + io_uring_prep_recv(sqe, ctx->fd, buf, len, flags); + + if (io_uring_submit(&ctx->ring) !=3D 1) + error(1, errno, "io_uring_submit"); + + if (io_uring_wait_cqe(&ctx->ring, &cqe)) + error(1, errno, "io_uring_wait_cqe"); + + res =3D cqe->res; + *cflags =3D cqe->flags; + io_uring_cqe_seen(&ctx->ring, cqe); + + return res; +} + +static void expect_uring_nonempty(unsigned int cflags, bool expected, + const char *what) +{ + bool nonempty =3D !!(cflags & IORING_CQE_F_SOCK_NONEMPTY); + + if (nonempty !=3D expected) { + fprintf(stderr, "%s: expected SOCK_NONEMPTY %d, got %d\n", + what, expected, nonempty); + exit(EXIT_FAILURE); + } +} + +static void expect_uring_more(unsigned int cflags, bool expected, + const char *what) +{ + bool more =3D !!(cflags & IORING_CQE_F_MORE); + + if (more !=3D expected) { + fprintf(stderr, "%s: expected F_MORE %d, got %d\n", + what, expected, more); + exit(EXIT_FAILURE); + } +} + +/* Wait until the whole payload is queued, so the hint is deterministic. + * Return false if the test has to be skipped. + */ +static bool inq_wait_queued(int fd, int len) +{ + if (!vsock_ioctl_int(fd, SIOCINQ, len)) { + fprintf(stderr, "Test skipped, SIOCINQ not supported.\n"); + return false; + } + + return true; +} + +static void inq_send_chunks(const struct test_opts *opts, int chunks) +{ + char buf[HINT_CHUNK_SIZE]; + int fd, i; + + fd =3D vsock_stream_connect(opts->peer_cid, opts->peer_port); + if (fd < 0) { + perror("connect"); + exit(EXIT_FAILURE); + } + + memset(buf, 0xa5, sizeof(buf)); + for (i =3D 0; i < chunks; i++) + send_buf(fd, buf, sizeof(buf), 0, sizeof(buf)); + + control_writeln("SENT"); + control_expectln("DONE"); + close(fd); +} + +static void test_stream_uring_inq_client(const struct test_opts *opts) +{ + inq_send_chunks(opts, 2); +} + +static void test_stream_uring_inq_server(const struct test_opts *opts) +{ + char buf[HINT_CHUNK_SIZE]; + struct uring_inq_ctx ctx; + unsigned int cflags; + int res; + + inq_server_init(&ctx, opts); + + control_expectln("SENT"); + if (!inq_wait_queued(ctx.fd, 2 * HINT_CHUNK_SIZE)) + goto out; + + /* Data remains after this receive, so the flag must be set. */ + res =3D inq_recv(&ctx, buf, sizeof(buf), 0, &cflags); + expect_res(res, HINT_CHUNK_SIZE, "partial receive"); + expect_uring_nonempty(cflags, true, "partial receive"); + + /* This receive drains the queue while the peer stays connected. */ + res =3D inq_recv(&ctx, buf, sizeof(buf), 0, &cflags); + expect_res(res, HINT_CHUNK_SIZE, "draining receive"); + expect_uring_nonempty(cflags, false, "draining receive"); + +out: + control_writeln("DONE"); + inq_server_exit(&ctx); +} + +static void test_stream_uring_inq_eof_client(const struct test_opts *opts) +{ + char buf[HINT_CHUNK_SIZE]; + int fd; + + fd =3D vsock_stream_connect(opts->peer_cid, opts->peer_port); + if (fd < 0) { + perror("connect"); + exit(EXIT_FAILURE); + } + + memset(buf, 0x5a, sizeof(buf)); + send_buf(fd, buf, sizeof(buf), 0, sizeof(buf)); + control_writeln("SENT"); + + control_expectln("DRAINED"); + close(fd); + control_writeln("CLOSED"); +} + +static void test_stream_uring_inq_eof_server(const struct test_opts *opts) +{ + char buf[HINT_CHUNK_SIZE]; + struct uring_inq_ctx ctx; + unsigned int cflags; + int res; + + inq_server_init(&ctx, opts); + + control_expectln("SENT"); + if (!inq_wait_queued(ctx.fd, HINT_CHUNK_SIZE)) { + control_writeln("DRAINED"); + control_expectln("CLOSED"); + goto out; + } + + res =3D inq_recv(&ctx, buf, sizeof(buf), 0, &cflags); + expect_res(res, HINT_CHUNK_SIZE, "drain before EOF"); + expect_uring_nonempty(cflags, false, "drain before EOF"); + + control_writeln("DRAINED"); + control_expectln("CLOSED"); + + /* The queue is empty and the peer is gone. The hint stays non-zero + * so that this receive happens and reports EOF, as TCP does after a + * FIN. + */ + res =3D inq_recv(&ctx, buf, sizeof(buf), 0, &cflags); + expect_res(res, 0, "receive at EOF"); + expect_uring_nonempty(cflags, true, "receive at EOF"); + +out: + inq_server_exit(&ctx); +} + +static void test_stream_uring_inq_empty_client(const struct test_opts *opt= s) +{ + int fd; + + fd =3D vsock_stream_connect(opts->peer_cid, opts->peer_port); + if (fd < 0) { + perror("connect"); + exit(EXIT_FAILURE); + } + + control_writeln("READY"); + control_expectln("DONE"); + close(fd); +} + +static void test_stream_uring_inq_empty_server(const struct test_opts *opt= s) +{ + char buf[HINT_CHUNK_SIZE]; + struct uring_inq_ctx ctx; + unsigned int cflags; + int res; + + inq_server_init(&ctx, opts); + + control_expectln("READY"); + + /* A failed receive must not leave a stale positive hint. */ + res =3D inq_recv(&ctx, buf, sizeof(buf), MSG_DONTWAIT, &cflags); + expect_res(res, -EAGAIN, "empty nonblocking receive"); + expect_uring_nonempty(cflags, false, "empty nonblocking receive"); + + control_writeln("DONE"); + inq_server_exit(&ctx); +} + +static void test_stream_uring_inq_zerolen_client(const struct test_opts *o= pts) +{ + inq_send_chunks(opts, 1); +} + +static void test_stream_uring_inq_zerolen_server(const struct test_opts *o= pts) +{ + char buf[HINT_CHUNK_SIZE]; + struct uring_inq_ctx ctx; + unsigned int cflags; + int res; + + inq_server_init(&ctx, opts); + + control_expectln("SENT"); + if (!inq_wait_queued(ctx.fd, HINT_CHUNK_SIZE)) + goto out; + + /* A zero-length request is not an error and still describes the + * queue behind it. + */ + res =3D inq_recv(&ctx, buf, 0, 0, &cflags); + expect_res(res, 0, "zero-length receive"); + expect_uring_nonempty(cflags, true, "zero-length receive"); + +out: + control_writeln("DONE"); + inq_server_exit(&ctx); +} + +static void test_stream_uring_inq_mshot_client(const struct test_opts *opt= s) +{ + char buf[HINT_CHUNK_SIZE]; + int fd, i; + + fd =3D vsock_stream_connect(opts->peer_cid, opts->peer_port); + if (fd < 0) { + perror("connect"); + exit(EXIT_FAILURE); + } + + memset(buf, 0x3c, sizeof(buf)); + for (i =3D 0; i < 2; i++) + send_buf(fd, buf, sizeof(buf), 0, sizeof(buf)); + control_writeln("SENT"); + + control_expectln("DRAINED"); + close(fd); + control_writeln("CLOSED"); + + control_expectln("DONE"); +} + +static void test_stream_uring_inq_mshot_server(const struct test_opts *opt= s) +{ + static char bufs[HINT_BUF_ENTRIES][HINT_CHUNK_SIZE]; + struct io_uring_buf_ring *br; + struct uring_inq_ctx ctx; + struct io_uring_sqe *sqe; + struct io_uring_cqe *cqe; + int i, ret; + + inq_server_init(&ctx, opts); + + br =3D io_uring_setup_buf_ring(&ctx.ring, HINT_BUF_ENTRIES, + HINT_BUF_GROUP, 0, &ret); + if (!br) { + fprintf(stderr, "io_uring_setup_buf_ring: %d\n", ret); + exit(EXIT_FAILURE); + } + + for (i =3D 0; i < HINT_BUF_ENTRIES; i++) + io_uring_buf_ring_add(br, bufs[i], HINT_CHUNK_SIZE, i, + io_uring_buf_ring_mask(HINT_BUF_ENTRIES), + i); + io_uring_buf_ring_advance(br, HINT_BUF_ENTRIES); + + control_expectln("SENT"); + if (!inq_wait_queued(ctx.fd, 2 * HINT_CHUNK_SIZE)) { + control_writeln("DRAINED"); + control_expectln("CLOSED"); + goto out; + } + + /* Arm only once both chunks are queued, so every completion knows + * what is left behind it. + */ + sqe =3D io_uring_get_sqe(&ctx.ring); + io_uring_prep_recv_multishot(sqe, ctx.fd, NULL, 0, 0); + sqe->flags |=3D IOSQE_BUFFER_SELECT; + sqe->buf_group =3D HINT_BUF_GROUP; + + if (io_uring_submit(&ctx.ring) !=3D 1) + error(1, errno, "io_uring_submit"); + + /* A buffer holds one chunk, so the other one is still queued. */ + if (io_uring_wait_cqe(&ctx.ring, &cqe)) + error(1, errno, "io_uring_wait_cqe"); + + expect_res(cqe->res, HINT_CHUNK_SIZE, "multishot first chunk"); + expect_uring_nonempty(cqe->flags, true, "multishot first chunk"); + expect_uring_more(cqe->flags, true, "multishot first chunk"); + io_uring_cqe_seen(&ctx.ring, cqe); + + /* This completion drains the queue and keeps the request armed. */ + if (io_uring_wait_cqe(&ctx.ring, &cqe)) + error(1, errno, "io_uring_wait_cqe"); + + expect_res(cqe->res, HINT_CHUNK_SIZE, "multishot second chunk"); + expect_uring_nonempty(cqe->flags, false, "multishot second chunk"); + expect_uring_more(cqe->flags, true, "multishot second chunk"); + io_uring_cqe_seen(&ctx.ring, cqe); + + control_writeln("DRAINED"); + control_expectln("CLOSED"); + + /* EOF ends multishot regardless of the hint. */ + if (io_uring_wait_cqe(&ctx.ring, &cqe)) + error(1, errno, "io_uring_wait_cqe"); + + expect_res(cqe->res, 0, "multishot EOF"); + expect_uring_more(cqe->flags, false, "multishot EOF"); + io_uring_cqe_seen(&ctx.ring, cqe); + +out: + control_writeln("DONE"); + io_uring_free_buf_ring(&ctx.ring, br, HINT_BUF_ENTRIES, + HINT_BUF_GROUP); + inq_server_exit(&ctx); +} + static struct test_case test_cases[] =3D { { .name =3D "SOCK_STREAM io_uring test", @@ -222,6 +588,31 @@ static struct test_case test_cases[] =3D { .run_server =3D test_stream_uring_msg_zc_server, .run_client =3D test_stream_uring_msg_zc_client, }, + { + .name =3D "SOCK_STREAM io_uring receive queue hint", + .run_server =3D test_stream_uring_inq_server, + .run_client =3D test_stream_uring_inq_client, + }, + { + .name =3D "SOCK_STREAM io_uring receive hint at EOF", + .run_server =3D test_stream_uring_inq_eof_server, + .run_client =3D test_stream_uring_inq_eof_client, + }, + { + .name =3D "SOCK_STREAM io_uring receive hint on empty queue", + .run_server =3D test_stream_uring_inq_empty_server, + .run_client =3D test_stream_uring_inq_empty_client, + }, + { + .name =3D "SOCK_STREAM io_uring receive hint zero-length", + .run_server =3D test_stream_uring_inq_zerolen_server, + .run_client =3D test_stream_uring_inq_zerolen_client, + }, + { + .name =3D "SOCK_STREAM io_uring multishot receive hint", + .run_server =3D test_stream_uring_inq_mshot_server, + .run_client =3D test_stream_uring_inq_mshot_client, + }, {}, }; =20 --=20 2.55.0