From nobody Mon Jul 27 12:11:50 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1783549019; cv=none; d=zohomail.com; s=zohoarc; b=FrGpea4dtSh28MCSxbY5L+5JSsyEAyLs1Qi6py9PYknyghSqw0LruPraOxSlOFVaeJbIF1a2Nl1NeCw6Unh39tp/SB7BtsdHO6Qnue7cmbFe47a4VsOrNNTqcDtNlPTiGCVVUa4teasAmfqrxo6YxNhdQVPvVVSSSaveKNoyA1A= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1783549019; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=OXrws65aUNdQ91XaTH7yxpdPdrZyTi9Sy/5xyhQXqEY=; b=i/CLFHheqWI3m6h6z/85FisXAbmd4IPo0FTJcoB/rbRqfXM1Y09r40INZUZYQ3dAn6rshpaQR3eXtHqOK3jt8itT/K/21qTmh4nxUvn8LJLtSb10vkKFAEErzW2e+7ABK7a47mEKyzgmQlhgikRwvsgjDfLzbjifTO//o1/JXMQ= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1783549019276148.32920351259713; Wed, 8 Jul 2026 15:16:59 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1whaZ9-0008Eq-Fp; Wed, 08 Jul 2026 18:16:27 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1whaZ7-0008DY-M3 for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:25 -0400 Received: from mail-ed1-x52d.google.com ([2a00:1450:4864:20::52d]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1whaZ5-0004Fr-Vh for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:25 -0400 Received: by mail-ed1-x52d.google.com with SMTP id 4fb4d7f45d1cf-697564cb69eso2124167a12.0 for ; Wed, 08 Jul 2026 15:16:23 -0700 (PDT) Received: from dobby ([2a02:8109:a394:4800:2200:181c:eef9:4c12]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-69a19d786e7sm9059237a12.16.2026.07.08.15.16.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 15:16:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783548982; x=1784153782; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=OXrws65aUNdQ91XaTH7yxpdPdrZyTi9Sy/5xyhQXqEY=; b=f9vOmef8QLTeoyBEPmmCx4v3cm3rk9XNqbNj8FY1/3JSBXFjzbpXrSjPAIpr0E7o/i TIFTpNqG455W9F9mnximfETMAU2WAdpZy2xCdqP8OwPY9TPc8YR3g/vSus3yLJsz5GcU Mze4IfeTIWybUnyCGDMKEgp00Yuai7rV0obFmUpdGCcTXAgNMWxVJHnYZ2DXPfibLC9U 9zshofosY1TXOG5tQz/4XY+ijn1xeb+LO5Bmlv6iDW6Va3yYE0JKRL7KAohT5UC7e1CF vvSAy+ne7tk5qNs5JgoIOoc5itGvN9+t7Q1qBsOEdHcwDYD4mYKJGYWSBU+EwWRTOMnO WguQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783548982; x=1784153782; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=OXrws65aUNdQ91XaTH7yxpdPdrZyTi9Sy/5xyhQXqEY=; b=dgRg9TmJJaKCte9FYjADBobabXeCmGFhgFWLbfsKhUQ9IyAn8lP34dPZkj5jkI9mRl 4jZIqPx21A8CysKEsNw4K3yvfBp42lTAb7mGt5DokrYfWF0riEFFepQ9iTRz/N1YWN/y S/ghVJtzJPE+ett1Z1zuGMdtruCYZ5COgtHYPBQTqWtMjTXQkGmJ0NiAh2kd9NGiGNsZ CQKkOfVCgsseUab6oqof8daVsAkSkD38PcVUvXRYvCtOK8zKhA3q+EUWWMR5D/bWfdzm 25rJEAp8R//Io+N3bG73QYDn36m46q/dICyCPj140ig1CLT3UgAGD3FMn/uJXwYPL3tc DVpw== X-Gm-Message-State: AOJu0YyEG0Y0NE1jxlIGjGTpyuQvt5QPCydmXlPsmItuEP+WzAbPUzmY nUEMQjX+l2QSaKNweXO1oLQmfx5NxtMpeLiFFqmJAYYB0n355ynUsQ+zDHnU3LFYI+E= X-Gm-Gg: AfdE7ckEwM/dcLqC0gCWWlx6xOEdC3mM1Q9QEMSG9alyQwXsvc/DP78yLw4fHL8a+Nb a8ykcauzL/+An2pER/foDF7EFr1Sd02qAmUPacxmHqmrznuimGvgbEc+IV5AM6xGj1ZPxzC7qL7 x64CEZcW5pV7udScnofEw6qd2tUHrWsz7hAhzlfL3Ia3C+nd44a2fQPYuMi3tcl/o2ueTe3qbsI djXmgPQLzFOmyW4zggLCdprPPOhDckopnzwlYz7N7CUZxiwYN7PQ86MFFdSTqihE7mWGmwsCj87 VXAxwR/zFQlQMjI03kAVpS3U2UM4agbTL3TTKXwkCJvqgVKyEzukSU79TR5WOtdzYx9viVg3T3B QzRyW5C9DvfmJft460w4I78hzsf0ilUWd5ho4l3oXeyJCakKKcM+NOA7fDvSjrbVDdPM1ODZGWE zeP9Ql49o2k28KO3NjY8Aq7K+MwwRvYWduimNJfvhEqNoXTQ== X-Received: by 2002:a05:6402:3216:b0:69a:4372:72ed with SMTP id 4fb4d7f45d1cf-69ab4478397mr1979803a12.13.1783548982332; Wed, 08 Jul 2026 15:16:22 -0700 (PDT) From: Sam Li To: qemu-devel@nongnu.org Cc: dlemoal@kernel.org, Stefan Hajnoczi , Pierrick Bouvier , "Michael S. Tsirkin" , Hanna Reitz , cassel@kernel.org, Markus Armbruster , Kevin Wolf , Eric Blake , qemu-block@nongnu.org, Sam Li Subject: [PATCH v14 1/6] docs/qcow2: add the zoned format feature Date: Thu, 9 Jul 2026 00:16:10 +0200 Message-ID: <20260708221615.346155-2-faithilikerun@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260708221615.346155-1-faithilikerun@gmail.com> References: <20260708221615.346155-1-faithilikerun@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2a00:1450:4864:20::52d; envelope-from=faithilikerun@gmail.com; helo=mail-ed1-x52d.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1783549020320158500 Content-Type: text/plain; charset="utf-8" Add the specs for the zoned format feature of the qcow2 driver. The qcow2 file then can emulate real zoned devices, either passed through by virtio-blk device or NVMe ZNS drive to the guest given zoned information. Signed-off-by: Sam Li Reviewed-by: Stefan Hajnoczi Reviewed-by: Niklas Cassel --- docs/system/qemu-block-drivers.rst.inc | 51 ++++++++++++++++++++++++++ 1 file changed, 51 insertions(+) diff --git a/docs/system/qemu-block-drivers.rst.inc b/docs/system/qemu-bloc= k-drivers.rst.inc index 675daa72f9..532fbb1f99 100644 --- a/docs/system/qemu-block-drivers.rst.inc +++ b/docs/system/qemu-block-drivers.rst.inc @@ -172,6 +172,57 @@ This section describes each format and the options tha= t are supported for it. filename`` to check if the NOCOW flag is set or not (Capital 'C' is NOCOW flag). =20 + .. option:: zone.mode + + If this is set to ``host-managed``, the image is an emulated zoned + block device. This option is only valid for emulated zoned device file= s. + + .. option:: zone.size + + The size of a zone in bytes. The device is divided into zones of this + size with the exception of the last zone, which may be smaller. + Defaults to ``256 MiB``. + + .. option:: zone.capacity + + The initial capacity value, in bytes, for all zones. The capacity must + be less than or equal to zone size. If the last zone is smaller, then + its capacity is capped. Defaults to ``zone.size``. + + The zone capacity is per zone and may be different between zones in re= al + devices. QCow2 sets all zones to the same capacity. + + .. option:: zone.conventional_zones + + The number of conventional zones of the zoned device. Defaults to ``0`` + (all zones are sequential-write-required). + + .. option:: zone.max_active_zones + + The maximum allowed active zones (zones in the implicit open, explicit + open, or closed state). + + The max active zones must be less or equal to the number of SWR + (sequential write required) zones of the device. Defaults to ``0``, + meaning no limit. + + .. option:: zone.max_open_zones + + The maximum allowed open zones (zones in the implicit open or explicit + open state). The max open zones must not be larger than the max active + zones. Defaults to ``zone.max_active_zones`` if that is set, otherwise + ``0`` (no limit). + + If the limits of open zones or active zones are equal to the number of + SWR zones, then it is the same as having no limits. + + .. option:: zone.max_append_bytes + + The maximum number of bytes of a zone append request that can be issued + to the device. It must be 512-byte aligned and less than the zone + capacity. A value of ``0`` means that zone append requests are not + supported. Defaults to ``64 KiB``. + .. program:: image-formats .. option:: qed =20 --=20 2.53.0 From nobody Mon Jul 27 12:11:50 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1783549030; cv=none; d=zohomail.com; s=zohoarc; b=ZB/KzWEoGPzn+xzeUITCB4QCji8bjN4i+Cci+PqXmjgYmyo/xC+tlT2fIN808fIxFlECnkIVJp1JiqNxiGnd6WM0NcV9DNsF+a3W6vWjgFwTuh4osyvw/azNeabI6OZwnbfv1kNCeDwqWr6Ms3tVWWxRPVaqRcpOa7usBxQj83A= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1783549030; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=MazHJbKCHFojNo3ryzg4npnLeCDhiaoceFTMzGltIAA=; b=milWsIX+SRnM/aP627N88OvbRX9S75oQ2SB14g1uzqN6y1yXF+64C5B/6ShAyuCGlUkSiWHzP5tv/xYbAt1J15P/25XAfRDxvNtPUi/mIuT58AvnnzCLoDhsdTBrU+olGGNMRgQfT7cYfbSkVfatfVd9oIA5VSAO3b5QotpHO+Q= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1783549030583208.16412504011498; Wed, 8 Jul 2026 15:17:10 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1whaZB-0008F3-Bt; Wed, 08 Jul 2026 18:16:29 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1whaZA-0008Et-Nn for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:28 -0400 Received: from mail-ed1-x52f.google.com ([2a00:1450:4864:20::52f]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1whaZ9-0004GV-9K for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:28 -0400 Received: by mail-ed1-x52f.google.com with SMTP id 4fb4d7f45d1cf-698beff7178so2118882a12.3 for ; Wed, 08 Jul 2026 15:16:26 -0700 (PDT) Received: from dobby ([2a02:8109:a394:4800:2200:181c:eef9:4c12]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-69a19d786e7sm9059237a12.16.2026.07.08.15.16.22 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 15:16:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783548986; x=1784153786; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=MazHJbKCHFojNo3ryzg4npnLeCDhiaoceFTMzGltIAA=; b=EQ4MvM/flXEZDlXxfktd3YCinolziDUQsx68FhKk0iEONlNy801VEbX6JnIBwQHHtH 0U6wiF968iKgVshCFzkucQKLfwc26PMYNp6Xcx1J7oFieUR0firLFJ8/I43C9nF1/iqF +D/AL7LiqlZ5qtAfEWr65c+J2pecPFfaYRG7YT/yB7/maDx04fmN0nYm7Hkdi/uIGGjq 0X70cg/gPW2ZkrCGyBYBKPdavVMFSzL4N35KyG0ICAE4sTMSVZPBmOCXIBy3Mx30TQQ1 2Itdybh6PHERwHV+sMwzGuew2GPFmyEEm7zduj+OLC3Bh9TrgssqGHVei9X5ILQcY+yb Pq8w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783548986; x=1784153786; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=MazHJbKCHFojNo3ryzg4npnLeCDhiaoceFTMzGltIAA=; b=QfCeyzPifQLDML3xDWW6pMVcfk0vCXDMMhrmv6ULTqjrqp1udAxLLHqzUkMvpvqDZM 1Xbjm6ngsf6Cn/glXAaMp+ZutE/iOptOWfY3fawhlYeh78Yr1wY6rtjLpnIazV2jDIT7 W8mLjPwx6V7d84LMosTVHGMj0d7aKeRACNIrZWAX0DZZb4MyM9tcB1cO6a98p+HxLQbG T5IrgCuvKk+dUbrppXRT/nyBBGhlnQmIh1RCZWP2n72bolz/q7/qdwcL4gQfn5yAbrgQ tluS77/DMk3wjNYGfI8pysh/y2z2xchhVHDY82ZjP9b+2SiQ/CTJTR9kZUYPNoic6jNi Y2FQ== X-Gm-Message-State: AOJu0Yya/MfguJtc769C4KfTMqBRaRBaGTnKilbMzQ1okXEW/VJ/Lhdl 3USj36QyKOY0vGC1a9igVO75sGLDqpJ3nf4z+Zr7UUnyQGmfr3h4SVC6c7IH4m2O2kI= X-Gm-Gg: AfdE7cmDyeHKTU5gfFnhn7yxyE0wfxqkwiEI6WwAVBoeeXt5qTUz2c888/DP+VWcmjd rj2nJE6M4gRRnl8qrRUWx2EkepRTxZS/5og4gDDULtTFUVQleAsgY0LfjRgNcZwysq9u6rYw+Ih X9/f8snbUAoxSadQDZU5dwGOQ0VoFF7vQGM8YgevnehLcmxVpoMPxomTFIzxI0P/MYzK0VajI+K Vx5g+OQTI3CsLeksOiMATVjIuMylUSe3/ETuAl8QBdz6Y0KpGEPQhgV4ArZOR40NtP+uXTzHnVI OTytNlUzeohr7Etkgntv0tjpW0s7qRk35BWDgAyStmmUpXL6YD4SbohND1dKwPclbj/GalM9Uwi MS/OhTtjt+lQgGNpqi5Ehm3Ll4brKqV6g4NDjjnvHe0VJVTkm0A2ZzCS2akEPwMKAjWRxz7fYvR lXhFvvAC5pj76QP21zInn0vLKJKuJuAfS/Qbc5ptl633Yq5Q== X-Received: by 2002:a05:6402:4305:b0:698:585c:f07b with SMTP id 4fb4d7f45d1cf-69ab44634e8mr1960238a12.3.1783548985622; Wed, 08 Jul 2026 15:16:25 -0700 (PDT) From: Sam Li To: qemu-devel@nongnu.org Cc: dlemoal@kernel.org, Stefan Hajnoczi , Pierrick Bouvier , "Michael S. Tsirkin" , Hanna Reitz , cassel@kernel.org, Markus Armbruster , Kevin Wolf , Eric Blake , qemu-block@nongnu.org, Sam Li Subject: [PATCH v14 2/6] block: widen BlockLimits.zone_size to uint64_t Date: Thu, 9 Jul 2026 00:16:11 +0200 Message-ID: <20260708221615.346155-3-faithilikerun@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260708221615.346155-1-faithilikerun@gmail.com> References: <20260708221615.346155-1-faithilikerun@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2a00:1450:4864:20::52f; envelope-from=faithilikerun@gmail.com; helo=mail-ed1-x52f.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1783549032292158500 Content-Type: text/plain; charset="utf-8" The zone-size field in BlockLimits is currently uint32_t, capping expressible zone sizes at 4 GiB. Real zoned-device protocols like NVMe ZNS allow larger zones. Widen BlockLimits.zone_size to uint64_t to match. Signed-off-by: Sam Li Reviewed-by: Niklas Cassel --- block/file-posix.c | 2 +- include/block/block_int-common.h | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/block/file-posix.c b/block/file-posix.c index 3c985da94f..ddb159c58b 100644 --- a/block/file-posix.c +++ b/block/file-posix.c @@ -3600,7 +3600,7 @@ raw_co_zone_append(BlockDriverState *bs, =20 if (*offset & zone_size_mask) { error_report("sector offset %" PRId64 " is not aligned to zone siz= e " - "%" PRId32 "", *offset / 512, bs->bl.zone_size / 512); + "%" PRId64 "", *offset / 512, bs->bl.zone_size / 512); return -EINVAL; } =20 diff --git a/include/block/block_int-common.h b/include/block/block_int-com= mon.h index 147c08155f..7571ed9968 100644 --- a/include/block/block_int-common.h +++ b/include/block/block_int-common.h @@ -901,7 +901,7 @@ typedef struct BlockLimits { BlockZoneModel zoned; =20 /* zone size expressed in bytes */ - uint32_t zone_size; + uint64_t zone_size; =20 /* total number of zones */ uint32_t nr_zones; --=20 2.53.0 From nobody Mon Jul 27 12:11:50 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1783549037; cv=none; d=zohomail.com; s=zohoarc; b=lDLNHBl9iQH5KVNEm1lJ1H5JpNUH99z3T9dn05uvc/gmZzurVdd+lRksWkX1v/TCKk/FGM4ZMPyL9yHh3I26+FjIbnAMxD06CvbyRer/mQcvN1WHCfKMwf0g++cL9VlMMJs474mk1NJK0WhWZT44FC2r3+ENSIdWfiPz3oo/zME= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1783549037; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=4+LAAShg3RQTzyknUcVCFjp1ae7nguq4/KUqQYF4uAc=; b=afS7yduse93lFGFwgd1BHaxKf7dIep2dV5S71LgnMk/j6Bi/HoGpFNHJllWTYPtHom6X/keYiEYhdaSUMYdO6rbk7YT5KJElnkRaXoaQrSMIQDwWH2sFdMN3DwyUvV31vgd9wx/b4Alfa4Wl5mEdz5+v+HgisRekxoQQgG9+Md8= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1783549037354110.8146724899076; Wed, 8 Jul 2026 15:17:17 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1whaZJ-0008Jn-2w; Wed, 08 Jul 2026 18:16:37 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1whaZG-0008Ia-Rk for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:34 -0400 Received: from mail-ed1-x531.google.com ([2a00:1450:4864:20::531]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1whaZC-0004Gn-4b for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:34 -0400 Received: by mail-ed1-x531.google.com with SMTP id 4fb4d7f45d1cf-698beff7178so2118927a12.3 for ; Wed, 08 Jul 2026 15:16:29 -0700 (PDT) Received: from dobby ([2a02:8109:a394:4800:2200:181c:eef9:4c12]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-69a19d786e7sm9059237a12.16.2026.07.08.15.16.25 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 15:16:27 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783548989; x=1784153789; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=4+LAAShg3RQTzyknUcVCFjp1ae7nguq4/KUqQYF4uAc=; b=J8A0CIWIbe7JcGeIK8KG3tvyRHn+63UPAgYpRE5+R5P6Lrf7UKQmbLBYGavWA+PShq F/6tDGFTo6ashaGAj3bpcG+KkBZuAUu9sFnQOWqI3+nRdUKt2drSR5LfOsSTM5qlD2oP H7m3iG0beZTGm77zNDJzaIvpPnx+pcsz1on8A22H2aJBvSlk2MVzUuOqQerGuiVPkIG+ VgAn6wvfwBAo7ThGiBxwDJR1R5KpWvLpqze7ncBTK/z487PqFjgRYPIaLastbYwnClfd XRkByjAMogxTnuLCVrnVZ4CXYNKXZMTYffdjv9taFTvsAne04R/PdA70+sO+OMHiFeC/ u6MQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783548989; x=1784153789; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=4+LAAShg3RQTzyknUcVCFjp1ae7nguq4/KUqQYF4uAc=; b=kpzgqLDtx24rNqhi1Yl92AiDHNwpZkgNC11MLx1YzcFp8dYT+97arBvjing38CqDkC /JcKaK40N+su+s2hixHiNXCvHD//AIFOX3+yuRIqnOpmRrX1nFKewb22klvF20mNZbtF 6Rch1a4s00f9BnKTPW+hH99KI0ngtXUXYnTEZ+do2UBjbOKcGABLua74zzLiWZLkLO4H oPayLMFtiYM3F0Z2TdDU3/pjEIkACtuKGKDJ5VbCTjFia0bntuF3kMIfBVVujOU5tGt6 miNtQH35ZeX+IUp4c4eYQO8Z+SvAnmhls0eZqORM0d3Msz0M8C+dE/n1jE3mN7BFSc1x g/Zg== X-Gm-Message-State: AOJu0YxFrWpaDC/AszzIv4tWopQH1wSWOjEAzvweFHZK/N4LGv8HaDUV rc9rJmIJUktnnF2sYD9TCorWGMfuySs+VK/sYH/8ldOONSpIlq9ZATM8ML4P+zmSYX8= X-Gm-Gg: AfdE7cl2hMNnoQjjL+goIaa0creQtiChqK+32KsUn7zwN/zEm+wh65HrKinxF9GtrZC LUgtVdpC9MhUNrcbpb2DxQKHA6yZPWJrL20QQMW3oirQePTTAJl7Qa/rTxUKYNoTvoX1GjbbhrX hdCF+uOU2tNr9/hLSdUjLqAQ2b0i+SCvSoEUPadz4O28PS9Jm1mce/qzbQ+04cFz5/FSp4M0kAa iFPn65FH7VqNxycnsB+5VK/8eh8lBY3Oye5H3fBpwMgjBywhWWZf9n1Ds9jzYt1L8mQ6VWG05TA s1Kw9x3v3ARaaK5eJkILhEwk8pvWE4S1w4yiLN2/vzXilFYcm+LDPwCkCr7k1PFz2pOkZoWNq1I KhA5BxcvfdntJ5DZpDnoAYZULSPBQQ1vNDQQSJHE2i0KOWvbuvBKqP3kw+U0qFH6XDc1LwGHn5+ ksEEM71idITDkw2HBP7L/XZvSdzGcxRhyMCHdQE4jxdsiaHQ== X-Received: by 2002:a05:6402:1473:b0:69a:50f9:6048 with SMTP id 4fb4d7f45d1cf-69ab44b551dmr1873807a12.28.1783548988157; Wed, 08 Jul 2026 15:16:28 -0700 (PDT) From: Sam Li To: qemu-devel@nongnu.org Cc: dlemoal@kernel.org, Stefan Hajnoczi , Pierrick Bouvier , "Michael S. Tsirkin" , Hanna Reitz , cassel@kernel.org, Markus Armbruster , Kevin Wolf , Eric Blake , qemu-block@nongnu.org, Sam Li Subject: [PATCH v14 3/6] qcow2: add configurations for zoned format extension Date: Thu, 9 Jul 2026 00:16:12 +0200 Message-ID: <20260708221615.346155-4-faithilikerun@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260708221615.346155-1-faithilikerun@gmail.com> References: <20260708221615.346155-1-faithilikerun@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2a00:1450:4864:20::531; envelope-from=faithilikerun@gmail.com; helo=mail-ed1-x531.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1783549038416158500 Content-Type: text/plain; charset="utf-8" To configure the zoned format feature on the qcow2 driver, it requires settings as: the device size, zone model, zone size, zone capacity, number of conventional zones, limits on zone resources (max append bytes, max open zones, and max_active_zones). To create a qcow2 image with zoned format feature, use command like this: qemu-img create -f qcow2 zbc.qcow2 -o size=3D768M \ -o zone.size=3D64M -o zone.capacity=3D64M -o zone.conventional_zones=3D0 \ -o zone.max_append_bytes=3D4096 -o zone.max_open_zones=3D6 \ -o zone.max_active_zones=3D8 -o zone.mode=3Dhost-managed Signed-off-by: Sam Li Acked-by: Markus Armbruster # QAPI schema --- block/qcow2.c | 366 +++++++++++++++++- block/qcow2.h | 86 +++- docs/interop/qcow2.rst | 118 +++++- include/block/block_int-common.h | 13 + qapi/block-core.json | 91 ++++- tests/qemu-iotests/031.out | 8 +- tests/qemu-iotests/036.out | 4 +- tests/qemu-iotests/082.out | 126 ++++++ tests/qemu-iotests/303.out | 4 +- tests/qemu-iotests/tests/qcow2-encryption.out | 4 +- 10 files changed, 803 insertions(+), 17 deletions(-) diff --git a/block/qcow2.c b/block/qcow2.c index 19271b10a4..9873789f58 100644 --- a/block/qcow2.c +++ b/block/qcow2.c @@ -73,6 +73,7 @@ typedef struct { #define QCOW2_EXT_MAGIC_CRYPTO_HEADER 0x0537be77 #define QCOW2_EXT_MAGIC_BITMAPS 0x23852875 #define QCOW2_EXT_MAGIC_DATA_FILE 0x44415441 +#define QCOW2_EXT_MAGIC_ZONED_FORMAT 0x007a6264 =20 static int coroutine_fn qcow2_co_preadv_compressed(BlockDriverState *bs, @@ -194,6 +195,95 @@ qcow2_extract_crypto_opts(QemuOpts *opts, const char *= fmt, Error **errp) return cryptoopts_qdict; } =20 +/* + * Returns true if zone_opt is valid, false otherwise. + */ +static bool +qcow2_check_zone_options(Qcow2ZonedHeaderExtension *zone_opt) +{ + uint32_t sequential_zones; + + assert(zone_opt !=3D NULL); + + if (zone_opt->zoned !=3D QCOW2_Z_NONE && zone_opt->zoned !=3D QCOW2_Z_= HM) { + warn_report("Zoned extension header zoned field has unknown " + "value %" PRIu8, zone_opt->zoned); + return false; + } + + if (zone_opt->zone_size =3D=3D 0) { + warn_report("Zoned extension header zone_size field must not be 0"= ); + return false; + } + + if (!is_power_of_2(zone_opt->zone_size)) { + warn_report("Zoned extension header zone_size %" PRIu64 + "B is not a power of 2", zone_opt->zone_size); + return false; + } + + if (zone_opt->nr_zones > QCOW2_MAX_NR_ZONES) { + warn_report("Zoned extension header nr_zones %" PRIu32 + " exceeds maximum %u", + zone_opt->nr_zones, QCOW2_MAX_NR_ZONES); + return false; + } + + if (zone_opt->zone_capacity > zone_opt->zone_size) { + warn_report("zone capacity %" PRIu64 "B exceeds zone size " + "%" PRIu64 "B", zone_opt->zone_capacity, + zone_opt->zone_size); + return false; + } + + if (!QEMU_IS_ALIGNED(zone_opt->max_append_bytes, BDRV_SECTOR_SIZE)) { + warn_report("max append bytes %" PRIu32 "B is not aligned " + "to %" PRIu64, zone_opt->max_append_bytes, + (uint64_t)BDRV_SECTOR_SIZE); + return false; + } + + if ((uint64_t)zone_opt->max_append_bytes + BDRV_SECTOR_SIZE >=3D + zone_opt->zone_capacity) { + warn_report("max append bytes %" PRIu32 "B exceeds zone " + "capacity %" PRIu64 "B by more than block size", + zone_opt->max_append_bytes, zone_opt->zone_capacity); + return false; + } + + if (zone_opt->conventional_zones >=3D zone_opt->nr_zones) { + warn_report("Conventional_zones %" PRIu32 " exceeds " + "nr_zones %" PRIu32 ".", + zone_opt->conventional_zones, zone_opt->nr_zones); + return false; + } + + if (zone_opt->max_active_zones > zone_opt->nr_zones) { + warn_report("max_active_zones %" PRIu32 " exceeds nr_zones %" PRIu= 32 + ", clamping to nr_zones", + zone_opt->max_active_zones, zone_opt->nr_zones); + zone_opt->max_active_zones =3D zone_opt->nr_zones; + } + + sequential_zones =3D zone_opt->nr_zones - zone_opt->conventional_zones; + if (zone_opt->max_open_zones > sequential_zones) { + warn_report("max_open_zones %" PRIu32 " exceeds the number of SWR " + "zones, clamping to %" PRIu32, + zone_opt->max_open_zones, sequential_zones); + zone_opt->max_open_zones =3D sequential_zones; + } + if (zone_opt->max_active_zones !=3D 0 && + zone_opt->max_open_zones > zone_opt->max_active_zones) { + warn_report("max_open_zones %" PRIu32 " exceeds max_active_zones " + "%" PRIu32 ", clamping to max_active_zones", + zone_opt->max_open_zones, + zone_opt->max_active_zones); + zone_opt->max_open_zones =3D zone_opt->max_active_zones; + } + + return true; +} + /* * read qcow2 extension and fill bs * start reading from start_offset @@ -211,6 +301,7 @@ qcow2_read_extensions(BlockDriverState *bs, uint64_t st= art_offset, uint64_t offset; int ret; Qcow2BitmapHeaderExt bitmaps_ext; + Qcow2ZonedHeaderExtension zoned_ext; =20 if (need_update_header !=3D NULL) { *need_update_header =3D false; @@ -432,6 +523,83 @@ qcow2_read_extensions(BlockDriverState *bs, uint64_t s= tart_offset, break; } =20 + case QCOW2_EXT_MAGIC_ZONED_FORMAT: + { + if (ext.len !=3D sizeof(zoned_ext)) { + error_setg(errp, "zoned_ext: unexpected len=3D%" PRIu32 " " + "(expected %zu)", ext.len, sizeof(zoned_ext)); + return -EINVAL; + } + ret =3D bdrv_pread(bs->file, offset, ext.len, &zoned_ext, 0); + if (ret < 0) { + error_setg_errno(errp, -ret, "zoned_ext: " + "Could not read ext header"); + return ret; + } + + zoned_ext.zone_size =3D be64_to_cpu(zoned_ext.zone_size); + zoned_ext.zone_capacity =3D be64_to_cpu(zoned_ext.zone_capacit= y); + zoned_ext.conventional_zones =3D + be32_to_cpu(zoned_ext.conventional_zones); + zoned_ext.nr_zones =3D be32_to_cpu(zoned_ext.nr_zones); + zoned_ext.max_open_zones =3D be32_to_cpu(zoned_ext.max_open_zo= nes); + zoned_ext.max_active_zones =3D + be32_to_cpu(zoned_ext.max_active_zones); + zoned_ext.max_append_bytes =3D + be32_to_cpu(zoned_ext.max_append_bytes); + s->zoned_header =3D zoned_ext; + + /* validate the header first */ + if (!qcow2_check_zone_options(&zoned_ext)) { + error_setg(errp, "Invalid zoned extension header"); + return -EINVAL; + } + /* + * refuse to open broken images: reject untrusted total_sectors + * that would overflow when converted to bytes, then verify + * nr_zones matches the device size. + */ + if (bs->total_sectors < 0 || + (uint64_t)bs->total_sectors > UINT64_MAX / BDRV_SECTOR_SIZ= E) { + error_setg(errp, "Image size overflows when converted " + "to bytes"); + return -EINVAL; + } + if (zoned_ext.nr_zones !=3D DIV_ROUND_UP( + (uint64_t)bs->total_sectors * BDRV_SECTOR_SIZE, + zoned_ext.zone_size)) { + error_setg(errp, "Zoned extension header nr_zones field " + "is wrong"); + return -EINVAL; + } + + /* + * Reject a header that claims more zones than the file can ho= ld. + */ + { + int64_t file_length =3D bdrv_co_getlength(bs->file->bs); + uint64_t wp_table_size =3D + (uint64_t)zoned_ext.nr_zones * sizeof(uint64_t); + + if (file_length < 0) { + return file_length; + } + if (zoned_ext.zonedmeta_offset > (uint64_t)file_length || + (uint64_t)file_length - zoned_ext.zonedmeta_offset < + wp_table_size) { + error_setg(errp, "Zoned metadata region exceeds the " + "image file size"); + return -EINVAL; + } + } + +#ifdef DEBUG_EXT + printf("Qcow2: Got zoned format extension: " + "offset=3D%" PRIu64 "\n", offset); +#endif + break; + } + default: /* unknown magic - save it in case we need to rewrite the head= er */ /* If you add a new feature, make sure to also update the fast @@ -2068,6 +2236,25 @@ static void qcow2_refresh_limits(BlockDriverState *b= s, Error **errp) } bs->bl.pwrite_zeroes_alignment =3D s->subcluster_size; bs->bl.pdiscard_alignment =3D s->cluster_size; + + switch (s->zoned_header.zoned) { + case QCOW2_Z_HM: + bs->bl.zoned =3D BLK_Z_HM; + break; + case QCOW2_Z_NONE: + default: + bs->bl.zoned =3D BLK_Z_NONE; + break; + } + + bs->bl.nr_zones =3D s->zoned_header.nr_zones; + bs->bl.max_append_sectors =3D s->zoned_header.max_append_bytes + >> BDRV_SECTOR_BITS; + bs->bl.max_active_zones =3D s->zoned_header.max_active_zones; + bs->bl.max_open_zones =3D s->zoned_header.max_open_zones; + bs->bl.zone_size =3D s->zoned_header.zone_size; + bs->bl.zone_capacity =3D s->zoned_header.zone_capacity; + bs->bl.write_granularity =3D BDRV_SECTOR_SIZE; } =20 static int GRAPH_UNLOCKED @@ -3170,6 +3357,11 @@ int qcow2_update_header(BlockDriverState *bs) .bit =3D QCOW2_INCOMPAT_EXTL2_BITNR, .name =3D "extended L2 entries", }, + { + .type =3D QCOW2_FEAT_TYPE_INCOMPATIBLE, + .bit =3D QCOW2_INCOMPAT_ZONED_FORMAT_BITNR, + .name =3D "zoned format", + }, { .type =3D QCOW2_FEAT_TYPE_COMPATIBLE, .bit =3D QCOW2_COMPAT_LAZY_REFCOUNTS_BITNR, @@ -3215,6 +3407,31 @@ int qcow2_update_header(BlockDriverState *bs) buflen -=3D ret; } =20 + /* Zoned devices header extension */ + if (s->zoned_header.zoned =3D=3D QCOW2_Z_HM) { + Qcow2ZonedHeaderExtension zoned_header =3D { + .zoned =3D s->zoned_header.zoned, + .zone_size =3D cpu_to_be64(s->zoned_header.zone_size), + .zone_capacity =3D cpu_to_be64(s->zoned_header.zone_capac= ity), + .conventional_zones =3D + cpu_to_be32(s->zoned_header.conventional_zones), + .nr_zones =3D cpu_to_be32(s->zoned_header.nr_zones), + .max_open_zones =3D cpu_to_be32(s->zoned_header.max_open_z= ones), + .max_active_zones =3D + cpu_to_be32(s->zoned_header.max_active_zones), + .max_append_bytes =3D + cpu_to_be32(s->zoned_header.max_append_bytes) + }; + ret =3D header_ext_add(buf, QCOW2_EXT_MAGIC_ZONED_FORMAT, + &zoned_header, sizeof(zoned_header), + buflen); + if (ret < 0) { + goto fail; + } + buf +=3D ret; + buflen -=3D ret; + } + /* Keep unknown header extensions */ QLIST_FOREACH(uext, &s->unknown_header_ext, next) { ret =3D header_ext_add(buf, uext->magic, uext->data, uext->len, bu= flen); @@ -3589,6 +3806,8 @@ qcow2_co_create(BlockdevCreateOptions *create_options= , Error **errp) ERRP_GUARD(); BlockdevCreateOptionsQcow2 *qcow2_opts; QDict *options; + Qcow2ZoneCreateOptions *zone_struct; + Qcow2ZoneHostManaged *zone_host_managed; =20 /* * Open the image file and write a minimal qcow2 header. @@ -3615,6 +3834,8 @@ qcow2_co_create(BlockdevCreateOptions *create_options= , Error **errp) =20 assert(create_options->driver =3D=3D BLOCKDEV_DRIVER_QCOW2); qcow2_opts =3D &create_options->u.qcow2; + zone_struct =3D qcow2_opts->zone; + zone_host_managed =3D NULL; =20 bs =3D bdrv_co_open_blockdev_ref(qcow2_opts->file, errp); if (bs =3D=3D NULL) { @@ -3828,6 +4049,14 @@ qcow2_co_create(BlockdevCreateOptions *create_option= s, Error **errp) header->incompatible_features |=3D cpu_to_be64(QCOW2_INCOMPAT_DATA_FILE); } + if (zone_struct && zone_struct->mode =3D=3D QCOW2_ZONE_MODEL_HOST_MANA= GED) { + /* + * The incompatible bit must be set when the zone model is + * host-managed + */ + header->incompatible_features |=3D + cpu_to_be64(QCOW2_INCOMPAT_ZONED_FORMAT); + } if (qcow2_opts->data_file_raw) { header->autoclear_features |=3D cpu_to_be64(QCOW2_AUTOCLEAR_DATA_FILE_RAW); @@ -3885,10 +4114,9 @@ qcow2_co_create(BlockdevCreateOptions *create_option= s, Error **errp) bdrv_graph_co_rdlock(); ret =3D qcow2_alloc_clusters(blk_bs(blk), 3 * cluster_size); if (ret < 0) { - bdrv_graph_co_rdunlock(); error_setg_errno(errp, -ret, "Could not allocate clusters for qcow= 2 " "header and refcount table"); - goto out; + goto unlock; =20 } else if (ret !=3D 0) { error_report("Huh, first cluster in empty image is already in use?= "); @@ -3896,11 +4124,78 @@ qcow2_co_create(BlockdevCreateOptions *create_optio= ns, Error **errp) } =20 /* Set the external data file if necessary */ + BDRVQcow2State *s =3D blk_bs(blk)->opaque; if (data_bs) { - BDRVQcow2State *s =3D blk_bs(blk)->opaque; s->image_data_file =3D g_strdup(data_bs->filename); } =20 + if (zone_struct && zone_struct->mode =3D=3D QCOW2_ZONE_MODEL_HOST_MANA= GED) { + s->zoned_header.zoned =3D QCOW2_Z_HM; + zone_host_managed =3D &zone_struct->u.host_managed; + + if (zone_host_managed->has_size) { + s->zoned_header.zone_size =3D zone_host_managed->size; + } else { + s->zoned_header.zone_size =3D DEFAULT_ZONE_SIZE; + } + + if (s->zoned_header.zone_size =3D=3D 0) { + error_setg(errp, "Zoned extension header zone_size field " + "can not be 0"); + s->zoned_header.zoned =3D QCOW2_Z_NONE; + ret =3D -EINVAL; + goto unlock; + } + s->zoned_header.nr_zones =3D DIV_ROUND_UP(qcow2_opts->size, + s->zoned_header.zone_size); + + if (zone_host_managed->has_capacity) { + s->zoned_header.zone_capacity =3D zone_host_managed->capacity; + } else { + s->zoned_header.zone_capacity =3D s->zoned_header.zone_size; + } + + if (zone_host_managed->has_conventional_zones) { + s->zoned_header.conventional_zones =3D + zone_host_managed->conventional_zones; + } else { + s->zoned_header.conventional_zones =3D 0; + } + + if (zone_host_managed->has_max_active_zones) { + s->zoned_header.max_active_zones =3D + zone_host_managed->max_active_zones; + } else { + s->zoned_header.max_active_zones =3D 0; + } + + if (zone_host_managed->has_max_open_zones) { + s->zoned_header.max_open_zones =3D + zone_host_managed->max_open_zones; + } else if (zone_host_managed->has_max_active_zones) { + s->zoned_header.max_open_zones =3D + zone_host_managed->max_active_zones; + } else { + s->zoned_header.max_open_zones =3D 0; + } + + if (zone_host_managed->has_max_append_bytes) { + s->zoned_header.max_append_bytes =3D + zone_host_managed->max_append_bytes; + } else { + s->zoned_header.max_append_bytes =3D DEFAULT_ZONE_MAX_APPEND_B= YTES; + } + + if (!qcow2_check_zone_options(&s->zoned_header)) { + error_setg(errp, "Invalid zoned device options"); + s->zoned_header.zoned =3D QCOW2_Z_NONE; + ret =3D -EINVAL; + goto unlock; + } + } else { + s->zoned_header.zoned =3D QCOW2_Z_NONE; + } + /* Create a full header (including things like feature table) */ ret =3D qcow2_update_header(blk_bs(blk)); bdrv_graph_co_rdunlock(); @@ -3974,6 +4269,9 @@ qcow2_co_create(BlockdevCreateOptions *create_options= , Error **errp) } =20 ret =3D 0; + goto out; +unlock: + bdrv_graph_co_rdunlock(); out: blk_co_unref(blk); bdrv_co_unref(bs); @@ -4052,6 +4350,10 @@ qcow2_co_create_opts(BlockDriver *drv, const char *f= ilename, QemuOpts *opts, { BLOCK_OPT_COMPAT_LEVEL, "version" }, { BLOCK_OPT_DATA_FILE_RAW, "data-file-raw" }, { BLOCK_OPT_COMPRESSION_TYPE, "compression-type" }, + { BLOCK_OPT_CONVENTIONAL_ZONES, "zone.conventional-zones" }, + { BLOCK_OPT_MAX_OPEN_ZONES, "zone.max-open-zones" }, + { BLOCK_OPT_MAX_ACTIVE_ZONES, "zone.max-active-zones" }, + { BLOCK_OPT_MAX_APPEND_BYTES, "zone.max-append-bytes" }, { NULL, NULL }, }; =20 @@ -5448,6 +5750,29 @@ qcow2_get_specific_info(BlockDriverState *bs, Error = **errp) .data_file_raw =3D data_file_is_raw(bs), .compression_type =3D s->compression_type, }; + if (s->zoned_header.zoned =3D=3D QCOW2_Z_HM) { + ImageInfoSpecificQCow2Zoned *z =3D + g_new0(ImageInfoSpecificQCow2Zoned, 1); + *z =3D (ImageInfoSpecificQCow2Zoned){ + .mode =3D QCOW2_ZONE_MODEL_HOST_MANAGED, + .nr_zones =3D s->zoned_header.nr_zones, + .u.host_managed =3D { + .has_size =3D true, + .size =3D s->zoned_header.zone_size, + .has_capacity =3D true, + .capacity =3D s->zoned_header.zone_capacity, + .has_conventional_zones =3D true, + .conventional_zones =3D s->zoned_header.conventional_z= ones, + .has_max_open_zones =3D true, + .max_open_zones =3D s->zoned_header.max_open_zones, + .has_max_active_zones =3D true, + .max_active_zones =3D s->zoned_header.max_active_zones, + .has_max_append_bytes =3D true, + .max_append_bytes =3D s->zoned_header.max_append_bytes, + }, + }; + spec_info->u.qcow2.data->zone =3D z; + } } else { /* if this assertion fails, this probably means a new version was * added without having it covered here */ @@ -6271,6 +6596,41 @@ static QemuOptsList qcow2_create_opts =3D { .type =3D QEMU_OPT_BOOL, \ .help =3D "Assume the external data file already exists and " \ "do not overwrite it" \ + }, \ + { \ + .name =3D BLOCK_OPT_ZONE_MODEL, \ + .type =3D QEMU_OPT_STRING, \ + .help =3D "zone model modes, mode choice: host-managed", \ + }, \ + { \ + .name =3D BLOCK_OPT_ZONE_SIZE, \ + .type =3D QEMU_OPT_SIZE, \ + .help =3D "zone size", \ + }, \ + { \ + .name =3D BLOCK_OPT_ZONE_CAPACITY, \ + .type =3D QEMU_OPT_SIZE, \ + .help =3D "zone capacity", \ + }, \ + { \ + .name =3D BLOCK_OPT_CONVENTIONAL_ZONES, \ + .type =3D QEMU_OPT_NUMBER, \ + .help =3D "numbers of conventional zones", \ + }, \ + { \ + .name =3D BLOCK_OPT_MAX_APPEND_BYTES, \ + .type =3D QEMU_OPT_SIZE, \ + .help =3D "max append bytes", \ + }, \ + { \ + .name =3D BLOCK_OPT_MAX_ACTIVE_ZONES, \ + .type =3D QEMU_OPT_NUMBER, \ + .help =3D "max active zones", \ + }, \ + { \ + .name =3D BLOCK_OPT_MAX_OPEN_ZONES, \ + .type =3D QEMU_OPT_NUMBER, \ + .help =3D "max open zones", \ }, QCOW_COMMON_OPTIONS, { /* end of list */ } diff --git a/block/qcow2.h b/block/qcow2.h index ce517040c4..0defb84d0c 100644 --- a/block/qcow2.h +++ b/block/qcow2.h @@ -129,6 +129,12 @@ =20 #define DEFAULT_CLUSTER_SIZE 65536 =20 +#define DEFAULT_ZONE_SIZE (256 * MiB) +#define DEFAULT_ZONE_MAX_APPEND_BYTES (64 * KiB) + +#define QCOW2_Z_NONE 0 /* Regular block device */ +#define QCOW2_Z_HM 1 /* Host-managed zoned block device */ + #define QCOW2_OPT_DATA_FILE "data-file" #define QCOW2_OPT_LAZY_REFCOUNTS "lazy-refcounts" #define QCOW2_OPT_DISCARD_REQUEST "pass-discard-request" @@ -237,6 +243,63 @@ typedef struct Qcow2CryptoHeaderExtension { uint64_t length; } QEMU_PACKED Qcow2CryptoHeaderExtension; =20 +/* Maximum number of zones a zoned image can have. */ +#define QCOW2_MAX_NR_ZONES (1U << 20) + +typedef struct Qcow2ZonedHeaderExtension { + /* Zoned device attributes */ + uint8_t zoned; + uint8_t reserved[3]; + uint32_t nr_zones; + uint64_t zone_size; + uint64_t zone_capacity; + uint32_t conventional_zones; + uint32_t max_active_zones; + uint32_t max_open_zones; + uint32_t max_append_bytes; + uint64_t zonedmeta_offset; +} QEMU_PACKED Qcow2ZonedHeaderExtension; + +typedef struct Qcow2ZoneListEntry { + QTAILQ_ENTRY(Qcow2ZoneListEntry) exp_open_zone_entry; + QTAILQ_ENTRY(Qcow2ZoneListEntry) imp_open_zone_entry; + QTAILQ_ENTRY(Qcow2ZoneListEntry) closed_zone_entry; +} Qcow2ZoneListEntry; + +/* + * Per-zone tracking for write pointer (WP) advance. A submitter + * reserves an LBA range, performs the data write outside the zone + * lock, then waits until the persisted WP catches up to its end + * offset before returning success to the guest. A failure aborts + * all higher-LBA peers in the same zone. + */ +typedef enum Qcow2WPReqState { + /* data write submitted, awaiting completion */ + QCOW2_WP_REQ_INFLIGHT, + /* data write completed, persisted WP not yet advanced past lba + len = */ + QCOW2_WP_REQ_PENDING, + /* persisted WP covers this range, submitter may ack success */ + QCOW2_WP_REQ_RESOLVED, + /* data write failed, or a lower-LBA peer in the same zone failed */ + QCOW2_WP_REQ_ABORTED, +} Qcow2WPReqState; + +typedef struct Qcow2WPReq { + uint64_t lba; /* start offset of the data write = */ + uint64_t len; /* length in bytes */ + Qcow2WPReqState state; + /* submitter waits for the data write to reach a terminal state */ + CoQueue wait; + QTAILQ_ENTRY(Qcow2WPReq) entry; +} Qcow2WPReq; + +typedef struct Qcow2ZoneWPState { + /* INFLIGHT reqs */ + QTAILQ_HEAD(, Qcow2WPReq) in_flight; + /* PENDING reqs, sorted by lba */ + QTAILQ_HEAD(, Qcow2WPReq) completed_pending; +} Qcow2ZoneWPState; + typedef struct Qcow2UnknownHeaderExtension { uint32_t magic; uint32_t len; @@ -257,17 +320,20 @@ enum { QCOW2_INCOMPAT_DATA_FILE_BITNR =3D 2, QCOW2_INCOMPAT_COMPRESSION_BITNR =3D 3, QCOW2_INCOMPAT_EXTL2_BITNR =3D 4, + QCOW2_INCOMPAT_ZONED_FORMAT_BITNR =3D 5, QCOW2_INCOMPAT_DIRTY =3D 1 << QCOW2_INCOMPAT_DIRTY_BITNR, QCOW2_INCOMPAT_CORRUPT =3D 1 << QCOW2_INCOMPAT_CORRUPT_BITNR, QCOW2_INCOMPAT_DATA_FILE =3D 1 << QCOW2_INCOMPAT_DATA_FILE_BITN= R, QCOW2_INCOMPAT_COMPRESSION =3D 1 << QCOW2_INCOMPAT_COMPRESSION_BI= TNR, QCOW2_INCOMPAT_EXTL2 =3D 1 << QCOW2_INCOMPAT_EXTL2_BITNR, + QCOW2_INCOMPAT_ZONED_FORMAT =3D 1 << QCOW2_INCOMPAT_ZONED_FORMAT_B= ITNR, =20 QCOW2_INCOMPAT_MASK =3D QCOW2_INCOMPAT_DIRTY | QCOW2_INCOMPAT_CORRUPT | QCOW2_INCOMPAT_DATA_FILE | QCOW2_INCOMPAT_COMPRESSION - | QCOW2_INCOMPAT_EXTL2, + | QCOW2_INCOMPAT_EXTL2 + | QCOW2_INCOMPAT_ZONED_FORMAT, }; =20 /* Compatible feature bits */ @@ -426,6 +492,24 @@ typedef struct BDRVQcow2State { * is to convert the image with the desired compression type set. */ Qcow2CompressionType compression_type; + + /* States of zoned device */ + Qcow2ZonedHeaderExtension zoned_header; + QTAILQ_HEAD(, Qcow2ZoneListEntry) exp_open_zones; + QTAILQ_HEAD(, Qcow2ZoneListEntry) imp_open_zones; + QTAILQ_HEAD(, Qcow2ZoneListEntry) closed_zones; + Qcow2ZoneListEntry *zone_list_entries; + uint32_t nr_zones_exp_open; + uint32_t nr_zones_imp_open; + uint32_t nr_zones_closed; + + /* Per-zone in-memory WP tracking for zoned imgaes. */ + Qcow2ZoneWPState *zone_wp_state; + /* + * Cache holding the on-disk zonedmeta cluster(s) for zoned image. + * WP advances go through this cache. + */ + Qcow2Cache *wp_cache; } BDRVQcow2State; =20 typedef struct Qcow2COWRegion { diff --git a/docs/interop/qcow2.rst b/docs/interop/qcow2.rst index 5948591107..7eb6da09d0 100644 --- a/docs/interop/qcow2.rst +++ b/docs/interop/qcow2.rst @@ -128,7 +128,26 @@ the next fields through ``header_length``. allows subcluster-based allocation. See the Extended L2 Entries section for more detai= ls. =20 - Bits 5-63: Reserved (set to 0) + Bit 5: Zoned extension bit. If this bit is set th= en + the file is an emulated zoned device. The + zoned extension must be present. + Implementations that do not support zoned + emulation cannot open this file because it + generally only make sense to interpret the + data along with the zone information and + write pointers. + + It is unsafe when any qcow2 user without + knowing the zoned extension reads or edits + a file with the zoned extension. The write + pointer tracking can be corrupted when a + writer edits a file, like overwriting beyo= nd + the write pointer locations. Or a reader t= ries + to access a file without knowing write + pointers where the software setup will cau= se + invalid reads. + + Bits 6-63: Reserved (set to 0) =20 80 - 87: compatible_features Bitmask of compatible features. An implementation can @@ -259,6 +278,7 @@ be stored. Each extension has a structure like the foll= owing:: 0x23852875 - Bitmaps extension 0x0537be77 - Full disk encryption header pointer 0x44415441 - External data file name string + 0x007a6264 - Zoned extension other - Unknown header extension, can be safe= ly ignored =20 @@ -344,6 +364,102 @@ The fields of the bitmaps extension are:: Offset into the image file at which the bitmap directory starts. Must be aligned to a cluster boundary. =20 +Zoned extension +--------------- + +The zoned extension must be present if the incompatible_features bit 5 is = set, +and omitted when it is clear. It contains fields for emulating the zoned +storage model (https://zonedstorage.io/). Currently only the host-managed = zone +model is supported. + +The fields of the zoned extension are:: + + Byte 0: zoned + The byte represents the zoned model of the device. 0 is= for + a non-zoned device (all other information in this header + is ignored). 1 is for a host-managed device, which only + allows for sequential writes within each zone. Other + values may be added later, the implementation must refu= se + to open a device containing an unknown zone model. + + 1 - 3: Reserved, must be zero. + + 4 - 7: nr_zones + The number of zones. It is the sum of conventional zones + and sequential zones. The maximum value for nr_zones is + (2^32 - 1)/8 =3D 536870911. + + 8 - 15: zone_size + Total size of each zone, in bytes. The 64-bit field is = to + satisfy the virtio-blk zone_size range and emulate a re= ad + zoned device, whose maximum zone size can be as large as + 2TB. + + The value must be power of 2. Linux currently requires + the zone size to be a power of 2 number of LBAs. Qcow2 + following this is mainly to allow emulating a real + ZNS drive configuration. It is not relevant to the clus= ter + size. + + 16 - 23: zone_capacity + The number of writable bytes within the zones. The bytes + between zone capacity and zone size are unusable: reads + will return 0s and writes will fail. + + A zone capacity is always smaller or equal to the zone + size. It is for emulating a real ZNS drive configuratio= n, + which has the constraint of aligning to some hardware e= rase + block size. + + 24 - 27: conventional_zones + The number of conventional zones. The conventional zones + allow sequential writes and random writes. While the + sequential zones only allow sequential writes. + + 28 - 31: max_active_zones + The maximum allowed active zones (zones in the implicit + open, explicit open, or closed state). It cannot be lar= ger + than nr_zones. + + 32 - 35: max_open_zones + The maximum allowed open zones (zones in the implicit o= pen + or explicit open state). It cannot be larger than the n= umber + of SWR zones of the device, nor larger than max_active_= zones. + + If the limits of open zones or active zones are equal to + the total number of SWR zones, then it's the same as ha= ving + no limits therefore max open zones and max active zones= are + set to 0. + + 36 - 39: max_append_bytes + The maximum number of bytes of a zone append request th= at + can be issued to the device. It must be 512-byte aligned + and less than the zone capacity. + + A value of 0 means that the zone append requests are not + supported. + + 40 - 47: zonedmeta_offset + The offset of zoned metadata structure in the contained + image, in bytes. + +The zonedmeta clusters contain a table of the zone write pointers. Each en= try +is a 64-bit value encoding both the zone type and the zone's write pointer= :: + + Bit 0 - 62: Write pointer offset, in bytes, indicates the starting + point of the next write position in that zone. + For conventional zones the value is the zone's starting + offset. + + 63: Zone type. 0 =3D SWR; 1 =3D conventional. + + +Each zone's write pointer in the zonedmeta area is durable on disk only af= ter +the data it covers is durable. For a write to a SWR zone, the data is writ= ten +back to disk first, then the non-zoned qcow2 L1/L2/refcount metadata is wr= itten, +and finally the write pointers are written. This ordering ensures that aft= er +a power failure the on-disk write pointer never leads data being written. + Full disk encryption header pointer ----------------------------------- =20 diff --git a/include/block/block_int-common.h b/include/block/block_int-com= mon.h index 7571ed9968..349876bc65 100644 --- a/include/block/block_int-common.h +++ b/include/block/block_int-common.h @@ -57,6 +57,13 @@ #define BLOCK_OPT_COMPRESSION_TYPE "compression_type" #define BLOCK_OPT_EXTL2 "extended_l2" #define BLOCK_OPT_KEEP_DATA_FILE "keep_data_file" +#define BLOCK_OPT_ZONE_MODEL "zone.mode" +#define BLOCK_OPT_ZONE_SIZE "zone.size" +#define BLOCK_OPT_ZONE_CAPACITY "zone.capacity" +#define BLOCK_OPT_CONVENTIONAL_ZONES "zone.conventional_zones" +#define BLOCK_OPT_MAX_APPEND_BYTES "zone.max_append_bytes" +#define BLOCK_OPT_MAX_ACTIVE_ZONES "zone.max_active_zones" +#define BLOCK_OPT_MAX_OPEN_ZONES "zone.max_open_zones" =20 #define BLOCK_PROBE_BUF_SIZE 512 =20 @@ -903,6 +910,12 @@ typedef struct BlockLimits { /* zone size expressed in bytes */ uint64_t zone_size; =20 + /* + * The number of usable logical blocks within the zone, expressed + * in bytes. A zone capacity is smaller or equal to the zone size. + */ + uint64_t zone_capacity; + /* total number of zones */ uint32_t nr_zones; =20 diff --git a/qapi/block-core.json b/qapi/block-core.json index 1f87b07850..4542bfc8f5 100644 --- a/qapi/block-core.json +++ b/qapi/block-core.json @@ -93,6 +93,9 @@ # # @compression-type: the image cluster compression method (since 5.1) # +# @zone: zoned configuration; only set when the image was created +# with a zoned model (since 11.1) +# # Since: 1.7 ## { 'struct': 'ImageInfoSpecificQCow2', @@ -106,7 +109,8 @@ 'refcount-bits': 'int', '*encrypt': 'ImageInfoSpecificQCow2Encryption', '*bitmaps': ['Qcow2BitmapInfo'], - 'compression-type': 'Qcow2CompressionType' + 'compression-type': 'Qcow2CompressionType', + '*zone': 'ImageInfoSpecificQCow2Zoned' } } =20 ## @@ -5221,6 +5225,85 @@ { 'enum': 'Qcow2CompressionType', 'data': [ 'zlib', { 'name': 'zstd', 'if': 'CONFIG_ZSTD' } ] } =20 +## +# @Qcow2ZoneModel: +# +# Zoned device model used in qcow2 image file +# +# @host-managed: The host-managed model only allows sequential write +# over the device zones. +# +# Since: 11.1 +## +{ 'enum': 'Qcow2ZoneModel', + 'data': [ 'host-managed'] } + +## +# @Qcow2ZoneHostManaged: +# +# The host-managed zone model. It only allows sequential writes. +# +# @size: Total number of bytes within zones (default: 256 MB). +# +# @capacity: The usable space within each zone, in bytes. Always +# smaller than or equal to @size (default: same as @size). +# +# @conventional-zones: The number of conventional zones of the +# zoned device (default: 0). +# +# @max-open-zones: The maximal number of open zones. Must be less +# than or equal to the number of non-conventional zones (i.e. +# ``size / zone.size - @conventional-zones``), and less than or +# equal to @max-active-zones (default: 0, meaning no limit). +# +# @max-active-zones: The maximal number of zones in the implicit +# open, explicit open or closed state. It is less than or equal +# to the number of zones (default: 0, meaning no limit). +# +# @max-append-bytes: The maximal size in bytes of a zone-append +# request that can be issued to the device. Must be a multiple +# of 512, and less than @capacity (default: 64 KB). +# +# Since: 11.1 +## +{ 'struct': 'Qcow2ZoneHostManaged', + 'data': { '*size': 'size', + '*capacity': 'size', + '*conventional-zones': 'uint32', + '*max-open-zones': 'uint32', + '*max-active-zones': 'uint32', + '*max-append-bytes': 'size' } } + +## +# @Qcow2ZoneCreateOptions: +# +# Creation options for zoned qcow2 images. +# +# @mode: The zone model used by the image. +# +# Since: 11.1 +## +{ 'union': 'Qcow2ZoneCreateOptions', + 'base': { 'mode': 'Qcow2ZoneModel' }, + 'discriminator': 'mode', + 'data': { 'host-managed': 'Qcow2ZoneHostManaged' } } + +## +# @ImageInfoSpecificQCow2Zoned: +# +# @mode: The zone model used by the image. +# +# @nr-zones: total number of zones in the image, derived at create +# time from the size and the zone size. +# +# Since: 11.1 +## +{ 'union': 'ImageInfoSpecificQCow2Zoned', + 'base': { 'mode': 'Qcow2ZoneModel', + 'nr-zones': 'uint32' }, + 'discriminator': 'mode', + 'data': { 'host-managed': 'Qcow2ZoneHostManaged' } } + ## # @BlockdevCreateOptionsQcow2: # @@ -5263,6 +5346,9 @@ # @compression-type: The image cluster compression method # (default: zlib, since 5.1) # +# @zone: Options for zoned images. If absent, the device is not +# zoned. (since 11.1) +# # Since: 2.12 ## { 'struct': 'BlockdevCreateOptionsQcow2', @@ -5279,7 +5365,8 @@ '*preallocation': 'PreallocMode', '*lazy-refcounts': 'bool', '*refcount-bits': 'int', - '*compression-type':'Qcow2CompressionType' } } + '*compression-type':'Qcow2CompressionType', + '*zone': 'Qcow2ZoneCreateOptions' } } =20 ## # @BlockdevCreateOptionsQed: diff --git a/tests/qemu-iotests/031.out b/tests/qemu-iotests/031.out index 0054c2ed97..02a68b34f7 100644 --- a/tests/qemu-iotests/031.out +++ b/tests/qemu-iotests/031.out @@ -117,7 +117,7 @@ header_length 112 =20 Header extension: magic 0x6803f857 (Feature table) -length 384 +length 432 data =20 Header extension: @@ -150,7 +150,7 @@ header_length 112 =20 Header extension: magic 0x6803f857 (Feature table) -length 384 +length 432 data =20 Header extension: @@ -164,7 +164,7 @@ No errors were found on the image. =20 magic 0x514649fb version 3 -backing_file_offset 0x240 +backing_file_offset 0x270 backing_file_size 0x17 cluster_bits 16 size 67108864 @@ -188,7 +188,7 @@ data 'host_device' =20 Header extension: magic 0x6803f857 (Feature table) -length 384 +length 432 data =20 Header extension: diff --git a/tests/qemu-iotests/036.out b/tests/qemu-iotests/036.out index 1fa7cad28d..3ee310bb61 100644 --- a/tests/qemu-iotests/036.out +++ b/tests/qemu-iotests/036.out @@ -26,7 +26,7 @@ compatible_features [] autoclear_features [63] Header extension: magic 0x6803f857 (Feature table) -length 384 +length 432 data =20 =20 @@ -38,7 +38,7 @@ compatible_features [] autoclear_features [] Header extension: magic 0x6803f857 (Feature table) -length 384 +length 432 data =20 *** done diff --git a/tests/qemu-iotests/082.out b/tests/qemu-iotests/082.out index e0463815c6..d1c8bb2532 100644 --- a/tests/qemu-iotests/082.out +++ b/tests/qemu-iotests/082.out @@ -72,6 +72,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: create -f qcow2 -o ? TEST_DIR/t.qcow2 128M Supported options: @@ -99,6 +106,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: create -f qcow2 -o cluster_size=3D4k,help TEST_DIR/t.qcow2 128M Supported options: @@ -126,6 +140,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: create -f qcow2 -o cluster_size=3D4k,? TEST_DIR/t.qcow2 128M Supported options: @@ -153,6 +174,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: create -f qcow2 -o help,cluster_size=3D4k TEST_DIR/t.qcow2 128M Supported options: @@ -180,6 +208,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: create -f qcow2 -o ?,cluster_size=3D4k TEST_DIR/t.qcow2 128M Supported options: @@ -207,6 +242,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: create -f qcow2 -o cluster_size=3D4k -o help TEST_DIR/t.qcow2 128M Supported options: @@ -234,6 +276,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: create -f qcow2 -o cluster_size=3D4k -o ? TEST_DIR/t.qcow2 128M Supported options: @@ -261,6 +310,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: create -f qcow2 -u -o backing_file=3DTEST_DIR/t.qcow2,,help -F qc= ow2 TEST_DIR/t.qcow2 128M Formatting 'TEST_DIR/t.qcow2', fmt=3Dqcow2 cluster_size=3D65536 extended_l= 2=3Doff compression_type=3Dzlib size=3D134217728 backing_file=3DTEST_DIR/t.= qcow2,,help backing_fmt=3Dqcow2 lazy_refcounts=3Doff refcount_bits=3D16 @@ -301,6 +357,13 @@ Supported qcow2 options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 The protocol level may support further options. Specify the target filename to include those options. @@ -391,6 +454,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: convert -O qcow2 -o ? TEST_DIR/t.qcow2 TEST_DIR/t.qcow2.base Supported options: @@ -418,6 +488,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: convert -O qcow2 -o cluster_size=3D4k,help TEST_DIR/t.qcow2 TEST_= DIR/t.qcow2.base Supported options: @@ -445,6 +522,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: convert -O qcow2 -o cluster_size=3D4k,? TEST_DIR/t.qcow2 TEST_DIR= /t.qcow2.base Supported options: @@ -472,6 +556,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: convert -O qcow2 -o help,cluster_size=3D4k TEST_DIR/t.qcow2 TEST_= DIR/t.qcow2.base Supported options: @@ -499,6 +590,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: convert -O qcow2 -o ?,cluster_size=3D4k TEST_DIR/t.qcow2 TEST_DIR= /t.qcow2.base Supported options: @@ -526,6 +624,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: convert -O qcow2 -o cluster_size=3D4k -o help TEST_DIR/t.qcow2 TE= ST_DIR/t.qcow2.base Supported options: @@ -553,6 +658,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: convert -O qcow2 -o cluster_size=3D4k -o ? TEST_DIR/t.qcow2 TEST_= DIR/t.qcow2.base Supported options: @@ -580,6 +692,13 @@ Supported options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 Testing: convert -O qcow2 -o backing_fmt=3Dqcow2,backing_file=3DTEST_DIR/t= .qcow2,,help TEST_DIR/t.qcow2 TEST_DIR/t.qcow2.base qemu-img: Could not open 'TEST_DIR/t.qcow2.base': Could not open backing f= ile: Could not open 'TEST_DIR/t.qcow2,help': No such file or directory @@ -620,6 +739,13 @@ Supported qcow2 options: preallocation=3D - Preallocation mode (allowed values: off, meta= data, falloc, full) refcount_bits=3D - Width of a reference count entry in bits size=3D - Virtual disk size + zone.capacity=3D - zone capacity + zone.conventional_zones=3D - numbers of conventional zones + zone.max_active_zones=3D - max active zones + zone.max_append_bytes=3D - max append bytes + zone.max_open_zones=3D - max open zones + zone.mode=3D - zone model modes, mode choice: host-managed + zone.size=3D - zone size =20 The protocol level may support further options. Specify the target filename to include those options. diff --git a/tests/qemu-iotests/303.out b/tests/qemu-iotests/303.out index b3c70827b7..684a481cae 100644 --- a/tests/qemu-iotests/303.out +++ b/tests/qemu-iotests/303.out @@ -47,7 +47,7 @@ header_length 112 =20 Header extension: magic 0x6803f857 (Feature table) -length 384 +length 432 data =20 Header extension: @@ -133,7 +133,7 @@ wrote 1048576/1048576 bytes at offset 7340032 { "name": "Feature table", "magic": 1745090647, - "length": 384, + "length": 432, "data_str": "" }, { diff --git a/tests/qemu-iotests/tests/qcow2-encryption.out b/tests/qemu-iot= ests/tests/qcow2-encryption.out index 9b549dc2ab..8be3e57f6e 100644 --- a/tests/qemu-iotests/tests/qcow2-encryption.out +++ b/tests/qemu-iotests/tests/qcow2-encryption.out @@ -10,7 +10,7 @@ data =20 Header extension: magic 0x6803f857 (Feature table) -length 384 +length 432 data =20 image: TEST_DIR/t.IMGFMT @@ -24,7 +24,7 @@ No errors were found on the image. =20 Header extension: magic 0x6803f857 (Feature table) -length 384 +length 432 data =20 qemu-img: Could not open 'TEST_DIR/t.IMGFMT': Missing CRYPTO header for cr= ypt method 2 --=20 2.53.0 From nobody Mon Jul 27 12:11:50 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1783549043; cv=none; d=zohomail.com; s=zohoarc; b=HTQQfAll0+UhEodBMKDg1IY1JPN47vnG+GJBlwTouxP35k0gIx7jM4AV21nytmgRwQECC84+gbRyscCwHpPKVEnWx52NMweJINaGSJRF0tYB00UNXkmMIG9sAusytZ5dpskE4owtn3Lqj5HxsgluWQlt1l6wy9RYhjafXwcFlwc= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1783549043; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=rf0dRaa81KrSy0wpPsNeErRiWruk8QrKs+wx142BAg8=; b=Gn652fyJ2X4yM15xQMFqrUcFS40xrY0SSsZQUvIaXFFPqibwonJESxNVQq39GdOBZLeXCuLHiVBaLJXohWFV8SYjbNmSkZr4smGNk7xwlqqzbZY/m+78tSjoMQruP/7nJ6MdMnHZxulR23KcWT1egOLwLaAsgpZtyq51y7PB+ig= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1783549043679630.4778125766442; Wed, 8 Jul 2026 15:17:23 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1whaZJ-0008Ki-T4; Wed, 08 Jul 2026 18:16:37 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1whaZI-0008J1-5f for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:36 -0400 Received: from mail-ed1-x52a.google.com ([2a00:1450:4864:20::52a]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1whaZE-0004H4-9F for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:35 -0400 Received: by mail-ed1-x52a.google.com with SMTP id 4fb4d7f45d1cf-691c5776f95so410687a12.3 for ; Wed, 08 Jul 2026 15:16:31 -0700 (PDT) Received: from dobby ([2a02:8109:a394:4800:2200:181c:eef9:4c12]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-69a19d786e7sm9059237a12.16.2026.07.08.15.16.28 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 15:16:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783548991; x=1784153791; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=rf0dRaa81KrSy0wpPsNeErRiWruk8QrKs+wx142BAg8=; b=A4YQPRZ7WTLVEA8CiTqJVSzwlNCQV0uSDm2Sn1DSYeOGIkJa3yca8yYOwqy8LJmWdv wqjZZ3JQ7FDEVx+jTTMuU5WvB1BrZUfqiQfiqpdpl5pvnDUVUS5TRtRCncXuCUQRoOM0 TSCnyPAVXnatoV7OEaN+XDidVd51WM/3AicSjVTedZzraOHcgTYR8D3OJP8fdyL/6LMy WTnTp+jdtDqBdWk4pKoQ5WUjxMM0X30SH4jB3vIXIqhUeXK8EhmJhA/rr6LlyzuxSSdF Vm72KI9eZIWCytT9yZjpbdOaN7kKtAfF3OD7gsda8/wJl/AWzJ914r8SSV+NT0zaE0Iy x8AA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783548991; x=1784153791; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=rf0dRaa81KrSy0wpPsNeErRiWruk8QrKs+wx142BAg8=; b=EUnr2Z0hqoFQSP7SJhi+QEfwz89obBzhdokC5BetWQ5X/IzRjaVVL97jRhbiQopibJ 0Be5d3X9tDA+Z3q2Ugptq189OuSi1hlLuXEXqMdMygGihLvDUT6XMz1ahrH2s0Sn4ogh Wtbkr8Xh/0+TD0b6U3K+/OC9SuwbXpYF4Wf8h+XVt6yimABf3NyT2nkBHURHB0p3OLDI Gfelyd1jlBvU3TfIddSyGD/xRoD81x/AZ7BS7ZQBNGJKM2yVcOmCoohTtVhN5TuXPlCh 5sTlLt3tW0bx0rO7ajIIHHxn4jehWnGIT3h1exfkFlPOtduu05R84AcsUJttVzQfaSdL /Sxg== X-Gm-Message-State: AOJu0Yw4rxLJ2k0UzFWqxRQo7cKglMkK30Z4HULtGIP1rI6Evvp+4XUT jSM14pfk6c0c3zM6NDy9szurU4t7HeROTLkosQFm0xxEfC9N+4CclIcdfSELk/9K/3M= X-Gm-Gg: AfdE7ckrnM2Je8BpkBTAUjlZy2Fyl7kf4dGem/zbNbeNvFfA23AifjxxwB2K8jA9nM2 vjWmS/tzo7NTLCjRQQSHk97SNmU7pcdejVHP1awKH2kkEZg88x2sE4ZvK0fInwqptKbleE1H27T 5Qy3wBN2G0b+fCPUdOA/8CY/tDRq/hP4hMU1WG3yG+UjKcyQM17yfJVjW95FtgziHULRX+GWfVY E3Xjx01tyGkyFgPRsl5bn4Y1Do3mrwqRmo5g/CFiEGS9IWMWCkez7b/cPrDRprRKLTOmJ+F1M7a iUiExsuk048kwmf3d0P0J0XBnZfTZUrsxtxfR0zxKSVnuyuOsQXDrMAE2tln8xCOtV/WnrudxM4 1Dqj76MD+OYXkijKEIwVAdZIdeUelj8kBGU9LnqjP7s8184vg1hpjjW68ehNf0h8djMXnZJq2hU wlUEmR2zbLrFSfwrf02JsLozmkiwREtEUFLT8fAbycG0Qbaw== X-Received: by 2002:a05:6402:2345:b0:698:1504:e3d9 with SMTP id 4fb4d7f45d1cf-69ab44a23f9mr2196012a12.29.1783548990464; Wed, 08 Jul 2026 15:16:30 -0700 (PDT) From: Sam Li To: qemu-devel@nongnu.org Cc: dlemoal@kernel.org, Stefan Hajnoczi , Pierrick Bouvier , "Michael S. Tsirkin" , Hanna Reitz , cassel@kernel.org, Markus Armbruster , Kevin Wolf , Eric Blake , qemu-block@nongnu.org, Sam Li Subject: [PATCH v14 4/6] virtio-blk: do not merge writes across a zone boundary Date: Thu, 9 Jul 2026 00:16:13 +0200 Message-ID: <20260708221615.346155-5-faithilikerun@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260708221615.346155-1-faithilikerun@gmail.com> References: <20260708221615.346155-1-faithilikerun@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2a00:1450:4864:20::52a; envelope-from=faithilikerun@gmail.com; helo=mail-ed1-x52a.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1783549044303158500 Content-Type: text/plain; charset="utf-8" virtio_blk_submit_multireq() fuses adjacent in-zone requests into a single request. On a zoned backend, a merged zone append request that straddles a zone boundary is rejected by the device because each write must stay within a single zone. Add a bail condition to the merge coalescer: if combining the candidate request into the current batch would cross a zone boundary, flush the current batch and start a new one. Signed-off-by: Sam Li Reviewed-by: Stefan Hajnoczi Reviewed-by: Niklas Cassel --- block/block-backend.c | 11 +++++++++++ hw/block/virtio-blk.c | 22 +++++++++++++++++++++- include/system/block-backend-io.h | 1 + 3 files changed, 33 insertions(+), 1 deletion(-) diff --git a/block/block-backend.c b/block/block-backend.c index 37ba7e9fc4..049e70ddcb 100644 --- a/block/block-backend.c +++ b/block/block-backend.c @@ -2326,6 +2326,17 @@ uint32_t blk_get_request_alignment(BlockBackend *blk) return bs ? bs->bl.request_alignment : BDRV_SECTOR_SIZE; } =20 +/* + * Returns the zone size in bytes for a zoned backend, or 0 if @blk does + * not present zoned geometry. + */ +uint64_t blk_get_zone_size(BlockBackend *blk) +{ + BlockDriverState *bs =3D blk_bs(blk); + IO_CODE(); + return bs ? bs->bl.zone_size : 0; +} + /* Returns the optimal write zeroes alignment, in bytes; guaranteed nonzer= o */ uint32_t blk_get_pwrite_zeroes_alignment(BlockBackend *blk) { diff --git a/hw/block/virtio-blk.c b/hw/block/virtio-blk.c index 6b92066aff..f74cc3dbd0 100644 --- a/hw/block/virtio-blk.c +++ b/hw/block/virtio-blk.c @@ -294,6 +294,9 @@ static void virtio_blk_submit_multireq(VirtIOBlock *s, = MultiReqBuffer *mrb) int i =3D 0, start =3D 0, num_reqs =3D 0, niov =3D 0, nb_sectors =3D 0; uint32_t max_transfer; int64_t sector_num =3D 0; + uint64_t zone_size =3D blk_get_zone_size(s->blk); + bool zone_cross; + int64_t zone_sector, end_sector; =20 if (mrb->num_reqs =3D=3D 1) { submit_requests(s, mrb, 0, 1, -1); @@ -309,17 +312,34 @@ static void virtio_blk_submit_multireq(VirtIOBlock *s= , MultiReqBuffer *mrb) for (i =3D 0; i < mrb->num_reqs; i++) { VirtIOBlockReq *req =3D mrb->reqs[i]; if (num_reqs > 0) { + zone_cross =3D false; + + /* + * On zoned backends, a single backend write/read must not span + * a zone boundary. Bail out of merging if combining req into + * the current batch would straddle a zone. + */ + if (zone_size > 0) { + zone_sector =3D zone_size / BDRV_SECTOR_SIZE; + end_sector =3D req->sector_num + + req->qiov.size / BDRV_SECTOR_SIZE - 1; + zone_cross =3D (sector_num / zone_sector) !=3D + (end_sector / zone_sector); + } + /* * NOTE: We cannot merge the requests in below situations: * 1. requests are not sequential * 2. merge would exceed maximum number of IOVs * 3. merge would exceed maximum transfer length of backend de= vice + * 4. merge would cross a zone boundary on a zoned backend */ if (sector_num + nb_sectors !=3D req->sector_num || niov > blk_get_max_iov(s->blk) - req->qiov.niov || req->qiov.size > max_transfer || nb_sectors > (max_transfer - - req->qiov.size) / BDRV_SECTOR_SIZE) { + req->qiov.size) / BDRV_SECTOR_SIZE || + zone_cross) { submit_requests(s, mrb, start, num_reqs, niov); num_reqs =3D 0; } diff --git a/include/system/block-backend-io.h b/include/system/block-backe= nd-io.h index fd84723d9d..78d410b8df 100644 --- a/include/system/block-backend-io.h +++ b/include/system/block-backend-io.h @@ -121,6 +121,7 @@ uint32_t blk_get_request_alignment(BlockBackend *blk); uint32_t blk_get_pwrite_zeroes_alignment(BlockBackend *blk); uint32_t blk_get_max_transfer(BlockBackend *blk); uint64_t blk_get_max_hw_transfer(BlockBackend *blk); +uint64_t blk_get_zone_size(BlockBackend *blk); =20 int coroutine_fn blk_co_copy_range(BlockBackend *blk_in, int64_t off_in, BlockBackend *blk_out, int64_t off_out, --=20 2.53.0 From nobody Mon Jul 27 12:11:50 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1783549050; cv=none; d=zohomail.com; s=zohoarc; b=FPxpBL0xydkwr7Wo1GIwWaAVSM8Ja2RA67k+lOl1D68aouHNjWXxSkxzEWrrd5zAn4MtdtegyG8b95T7W1zNMxqKdPUhgfqyH6F73thZbZeUsuaxQay5C011qVb3WniGwcCUezewWU2ROJIvh3Ff94DajKvr+NcuBFXp3dkYjZ0= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1783549050; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=/q3R32EwpDpapU5qLyKPSVNTj/fyPBi6s6GAbIOHidc=; b=TUmq5MFeVlZis8krpE2XflB4jZTyysaM9ITFnelGuBLARomoYxXDBTIUTWOimO0/irofL5QYsgmyrZUKeh4owMOuNH9BRUZZ7CfB+AfS/eZ026C/J6r4wMqp4C5RrKatTxZa414TjN85C37Wj7/XleobdROH407vJ3+gnDtCiwA= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 178354905089886.97425288985892; Wed, 8 Jul 2026 15:17:30 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1whaZM-0008Me-9m; Wed, 08 Jul 2026 18:16:40 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1whaZL-0008M7-LM for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:39 -0400 Received: from mail-ed1-x535.google.com ([2a00:1450:4864:20::535]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1whaZG-0004HG-PO for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:39 -0400 Received: by mail-ed1-x535.google.com with SMTP id 4fb4d7f45d1cf-698e5859a3cso666477a12.0 for ; Wed, 08 Jul 2026 15:16:34 -0700 (PDT) Received: from dobby ([2a02:8109:a394:4800:2200:181c:eef9:4c12]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-69a19d786e7sm9059237a12.16.2026.07.08.15.16.30 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 15:16:32 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783548993; x=1784153793; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=/q3R32EwpDpapU5qLyKPSVNTj/fyPBi6s6GAbIOHidc=; b=E4uqCpyLMWR/1BbSNC9peHZjsyo6qSSfXbnXrHLFzbbJEY0MMmvumg4N57nMe+x394 f2sBD+xiLXZsDTrvCoTpjjgKvoMdibAyhwUePWkpDoGs4zqP2qG4TECv7QWcGjwTZLLB PmXGwzLwSC5Da22iJpHTZyLwXzKsN9AvXDA5Pv8qC4oQiDYFTV/k0NHTACGDvXJnmRal +8PUUbVd60w1ltMXrgQUcH+jUhClsWwBS3jlU8OgI3TOcnP4Kgklkw5A1RTlqWei57m4 C1eYqniRk+7RyevPOMooJL14bJ8/KONSTn0OTmegq/JECo/iE6J6KqiqA7esYMIq6SWT Ni+A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783548993; x=1784153793; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=/q3R32EwpDpapU5qLyKPSVNTj/fyPBi6s6GAbIOHidc=; b=jYKzolwB0TUSIgboN7GPOshPXh1jdnsCULNOfmdM6PirRHajEMUl8wZV89MCXV4rBw UGlpFPRGOgO2EAnk9ARIAGi7ofbZfGK2itR4jk/e10zYRzTow+t3fbrHjuOJ/oDae6yB wMr9udw/cVjEM8twW2bR+RmXT5Uftp5AvPetjeXdnv/fJCQK1/cnsMkU+lQgDn4K8qzL xj2BWGP7EqvUlbK8tvgL9BgggG4C9avgd2Df9dlWHI2mFQy2ipH9Siop5WQkGL2LiL0h Ms9RKziQIZedryhmEIB6FflDgBuN9t9oH+wCP4cqWbAfnFgUZMEcSpaMoUwGGZHYWzwp WUJA== X-Gm-Message-State: AOJu0YzVrycfOb9CGeQa2u+s5cOUCvTujkwEZQjDgNCWA94dQyRs09/5 VDQAm8L1CFz5OThOXUL6Cm2zjZ4uVXBw6myPlYiHH4b68+/VWmZgJu4sye951fOxIF0= X-Gm-Gg: AfdE7clLgBaWklU1Oi8iz09aUJz8BwVmwkS/VvT1d5JqwzKsIby8m9hpvdDW7HV73xe aIiPzHVXsorGaTLP2jLFIQ6tPo70lCnSLjDj/lEfo94AxwnytL7l3TMQKMsa5marJerKvO+UxV/ UyXMQ9ah0aIwBRTxvhivXX6nuOY30XXHwvI78xYI+k9canXA6Sc9uKuPGAJ5HBbFYqE3U0ibADU VWqUpRDUOKvmn6AXC67qAn9sCYxynwfJcqPlwWBWeVgEnnp8/ZUb6oB572fsaXuooyYfRpxFqsp GhgRlIyTnUQkcH5efMkny//O9dw17OqnqGeNXFlulFTYVa1UDGo39x1Etm/3YRZkh4xQERdoVRz RkNPf2jX/wCZ0Sw6NtTTAxObov4yrVYNV2qEhuGDvdVWRUrRUqyziEcxY8VeBy8OTXTbAlNd9Th /wntz7etpYLjzRd6FH/f1CjRpQNfF8bGu3Dw5lZCZHREmLLDywX/sM4GV8 X-Received: by 2002:a05:6402:2787:b0:698:3b7c:3b2e with SMTP id 4fb4d7f45d1cf-69c032a536bmr57794a12.16.1783548992886; Wed, 08 Jul 2026 15:16:32 -0700 (PDT) From: Sam Li To: qemu-devel@nongnu.org Cc: dlemoal@kernel.org, Stefan Hajnoczi , Pierrick Bouvier , "Michael S. Tsirkin" , Hanna Reitz , cassel@kernel.org, Markus Armbruster , Kevin Wolf , Eric Blake , qemu-block@nongnu.org, Sam Li Subject: [PATCH v14 5/6] qcow2: add zoned emulation capability Date: Thu, 9 Jul 2026 00:16:14 +0200 Message-ID: <20260708221615.346155-6-faithilikerun@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260708221615.346155-1-faithilikerun@gmail.com> References: <20260708221615.346155-1-faithilikerun@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2a00:1450:4864:20::535; envelope-from=faithilikerun@gmail.com; helo=mail-ed1-x535.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1783549052435158500 Content-Type: text/plain; charset="utf-8" By adding zone operations and zoned metadata, the zoned emulation capability enables full emulation support of zoned device using a qcow2 file. The zoned device metadata includes zone type, zoned device state and write pointer (WP) of each zone, which is stored to an array of unsigned integers. WP accessor (qcow2_rw_wp_at) routes reads and writes of an 8-byte WP slot through the write pointer cache. The write pointer cache is written to disk after the qcow2 metadata is written, thus guaranteeing that the write pointer is updated after the corresponding data is written. Per-completion cache flush is deferred. The WP cluster reaches disk on the next flush. Each zone of a zoned device makes state transitions following the zone state machine. The zone state machine mainly describes five states, IMPLICIT OPEN, EXPLICIT OPEN, FULL, EMPTY and CLOSED. READ ONLY and OFFLINE states will generally be affected by device internal events. The operations on zones cause corresponding state changing. Zoned devices have limits on zone resources, which put constraints on write operations on zones. It is managed by active zone queues following LRU policy. Signed-off-by: Sam Li --- block/qcow2-cache.c | 11 +- block/qcow2-refcount.c | 7 + block/qcow2.c | 1219 +++++++++++++++++++++++++++++++++++++++- block/qcow2.h | 2 + block/trace-events | 2 + 5 files changed, 1234 insertions(+), 7 deletions(-) diff --git a/block/qcow2-cache.c b/block/qcow2-cache.c index 23d9588b08..73f2be3641 100644 --- a/block/qcow2-cache.c +++ b/block/qcow2-cache.c @@ -275,14 +275,23 @@ int qcow2_cache_set_dependency(BlockDriverState *bs, = Qcow2Cache *c, { int ret; =20 + /* + * If the dependency graph is unchanged, nothing to do. This avoids + * a synchronous flush on every call below. + */ + if (c->depends =3D=3D dependency) { + return 0; + } + if (dependency->depends) { + /* Flatten any chain on the dependency side as a cycle-breaker. */ ret =3D qcow2_cache_flush_dependency(bs, dependency); if (ret < 0) { return ret; } } =20 - if (c->depends && (c->depends !=3D dependency)) { + if (c->depends) { ret =3D qcow2_cache_flush_dependency(bs, c); if (ret < 0) { return ret; diff --git a/block/qcow2-refcount.c b/block/qcow2-refcount.c index 6512cda407..f551726609 100644 --- a/block/qcow2-refcount.c +++ b/block/qcow2-refcount.c @@ -1239,6 +1239,13 @@ int qcow2_write_caches(BlockDriverState *bs) } } =20 + if (s->wp_cache) { + ret =3D qcow2_cache_write(bs, s->wp_cache); + if (ret < 0) { + return ret; + } + } + return 0; } =20 diff --git a/block/qcow2.c b/block/qcow2.c index 9873789f58..e6f0fb5681 100644 --- a/block/qcow2.c +++ b/block/qcow2.c @@ -195,6 +195,320 @@ qcow2_extract_crypto_opts(QemuOpts *opts, const char = *fmt, Error **errp) return cryptoopts_qdict; } =20 +#define QCOW2_WP_CONV_BIT (1ULL << 63) +#define QCOW2_ZT_IS_CONV(wp) ((wp) & QCOW2_WP_CONV_BIT) +#define QCOW2_GET_WP(wp) ((wp) & ~QCOW2_WP_CONV_BIT) + +/* + * To emulate a real zoned device, closed, empty and full states are + * preserved after a power cycle. The open states are in-memory and will + * be lost after closing the device. Read-only and offline states are + * device-internal events, which are not considered for simplicity. + * + * The qcow2 state lock must be held when called. + */ +static inline BlockZoneState qcow2_get_zone_state(BlockDriverState *bs, + uint32_t index) +{ + BDRVQcow2State *s =3D bs->opaque; + Qcow2ZoneListEntry *zone_entry =3D &s->zone_list_entries[index]; + uint64_t zone_wp =3D bs->wps->wp[index]; + uint64_t zone_start; + + if (QCOW2_ZT_IS_CONV(zone_wp)) { + return BLK_ZS_NOT_WP; + } + + if (QTAILQ_IN_USE(zone_entry, exp_open_zone_entry)) { + return BLK_ZS_EOPEN; + } + if (QTAILQ_IN_USE(zone_entry, imp_open_zone_entry)) { + return BLK_ZS_IOPEN; + } + + zone_start =3D (uint64_t)index * bs->bl.zone_size; + if (zone_wp =3D=3D zone_start) { + return BLK_ZS_EMPTY; + } + if (zone_wp >=3D zone_start + bs->bl.zone_capacity) { + return BLK_ZS_FULL; + } + if (zone_wp > zone_start) { + if (!QTAILQ_IN_USE(zone_entry, closed_zone_entry)) { + /* + * The number of closed zones is not always updated in time wh= en + * the device is closed. However, it only matters when doing + * zone report. Refresh the count and list of closed zones to + * provide correct zone states for zone report. + */ + QTAILQ_INSERT_HEAD(&s->closed_zones, zone_entry, closed_zone_e= ntry); + s->nr_zones_closed++; + } + return BLK_ZS_CLOSED; + } + return BLK_ZS_NOT_WP; +} + +/* The qcow2 state lock must be held when called */ +static void qcow2_rm_exp_open_zone(BDRVQcow2State *s, + uint32_t index) +{ + Qcow2ZoneListEntry *zone_entry =3D &s->zone_list_entries[index]; + + QTAILQ_REMOVE(&s->exp_open_zones, zone_entry, exp_open_zone_entry); + s->nr_zones_exp_open--; +} + +/* The qcow2 state lock must be held when called */ +static void qcow2_rm_imp_open_zone(BDRVQcow2State *s, + int32_t index) +{ + Qcow2ZoneListEntry *zone_entry; + if (index < 0) { + /* Apply LRU when the index is not specified. */ + zone_entry =3D QTAILQ_LAST(&s->imp_open_zones); + } else { + zone_entry =3D &s->zone_list_entries[index]; + } + + QTAILQ_REMOVE(&s->imp_open_zones, zone_entry, imp_open_zone_entry); + s->nr_zones_imp_open--; +} + +/* The qcow2 state lock must be held when called */ +static void qcow2_rm_open_zone(BDRVQcow2State *s, + uint32_t index) +{ + Qcow2ZoneListEntry *zone_entry =3D &s->zone_list_entries[index]; + + if (QTAILQ_IN_USE(zone_entry, exp_open_zone_entry)) { + qcow2_rm_exp_open_zone(s, index); + } else if (QTAILQ_IN_USE(zone_entry, imp_open_zone_entry)) { + qcow2_rm_imp_open_zone(s, index); + } +} + +/* The qcow2 state lock must be held when called */ +static void qcow2_rm_closed_zone(BDRVQcow2State *s, + uint32_t index) +{ + Qcow2ZoneListEntry *zone_entry =3D &s->zone_list_entries[index]; + + QTAILQ_REMOVE(&s->closed_zones, zone_entry, closed_zone_entry); + s->nr_zones_closed--; +} + +/* The qcow2 state lock must be held when called */ +static void qcow2_do_imp_open_zone(BDRVQcow2State *s, + uint32_t index, + BlockZoneState zs) +{ + Qcow2ZoneListEntry *zone_entry =3D &s->zone_list_entries[index]; + + switch (zs) { + case BLK_ZS_EMPTY: + break; + case BLK_ZS_CLOSED: + qcow2_rm_closed_zone(s, index); + break; + case BLK_ZS_IOPEN: + /* + * The LRU policy: update the zone that is most recently + * used to the head of the zone list + */ + if (zone_entry =3D=3D QTAILQ_FIRST(&s->imp_open_zones)) { + return; + } + QTAILQ_REMOVE(&s->imp_open_zones, zone_entry, imp_open_zone_entry); + s->nr_zones_imp_open--; + break; + default: + return; + } + + QTAILQ_INSERT_HEAD(&s->imp_open_zones, zone_entry, imp_open_zone_entry= ); + s->nr_zones_imp_open++; +} + +/* The qcow2 state lock must be held when called */ +static void qcow2_do_exp_open_zone(BDRVQcow2State *s, + uint32_t index) +{ + Qcow2ZoneListEntry *zone_entry =3D &s->zone_list_entries[index]; + + QTAILQ_INSERT_HEAD(&s->exp_open_zones, zone_entry, exp_open_zone_entry= ); + s->nr_zones_exp_open++; +} + +/* + * The list of zones is managed using an LRU policy: the last + * zone of the list is always the one that was least recently used + * for writing and is chosen as the zone to close to be able to + * implicitly open another zone. + * + * We can only close the open zones. The index is not specified + * when it is less than 0. + * + * The qcow2 state lock must be held when called. + */ +static void qcow2_do_close_zone(BlockDriverState *bs, + int32_t index, + BlockZoneState zs) +{ + BDRVQcow2State *s =3D bs->opaque; + Qcow2ZoneListEntry *zone_entry; + + if (index >=3D 0) { + zone_entry =3D &s->zone_list_entries[index]; + } else { + /* before removal of the last implicitly open zone */ + zone_entry =3D QTAILQ_LAST(&s->imp_open_zones); + } + + if (zs =3D=3D BLK_ZS_IOPEN) { + assert(zone_entry !=3D NULL); + qcow2_rm_imp_open_zone(s, index); + QTAILQ_INSERT_HEAD(&s->closed_zones, zone_entry, closed_zone_entry= ); + s->nr_zones_closed++; + return; + } + + if (index >=3D 0 && zs =3D=3D BLK_ZS_EOPEN) { + qcow2_rm_exp_open_zone(s, index); + /* + * The zone state changes when the zone is removed from the list of + * open zones (explicitly open -> empty). The closed zone list is + * refreshed during get_zone_state(). + */ + qcow2_get_zone_state(bs, index); + } +} + +/* + * Read/Write the new wp value for zone `index` through write pointer + * cache. Reads return the value currently held in the cache, which may + * be ahead of the on-disk value if the cache hasn't been flushed yet. + * Writes update the cache and mark the entry dirty. + */ +static int coroutine_fn GRAPH_RDLOCK +qcow2_rw_wp_at(BlockDriverState *bs, uint64_t *wp, + int32_t index, bool is_write) +{ + BDRVQcow2State *s =3D bs->opaque; + uint64_t wp_byte_off =3D sizeof(uint64_t) * index; + uint64_t cluster_file_off =3D + s->zoned_header.zonedmeta_offset + + (wp_byte_off & ~((uint64_t)s->cluster_size - 1)); + size_t off_in_cluster =3D wp_byte_off & (s->cluster_size - 1); + void *cluster_buf; + uint64_t *slot; + int ret; + + assert(s->wp_cache !=3D NULL); + + qemu_co_mutex_lock(&s->lock); + ret =3D qcow2_cache_get(bs, s->wp_cache, cluster_file_off, &cluster_bu= f); + if (ret < 0) { + qemu_co_mutex_unlock(&s->lock); + error_report("Failed to %s WP slot (zone %d): %s", + is_write ? "write" : "read", index, strerror(-ret)); + return ret; + } + + slot =3D (uint64_t *)((char *)cluster_buf + off_in_cluster); + if (is_write) { + *slot =3D cpu_to_be64(*wp); + qcow2_cache_entry_mark_dirty(s->wp_cache, cluster_buf); + } else { + *wp =3D be64_to_cpu(*slot); + } + qcow2_cache_put(s->wp_cache, &cluster_buf); + qemu_co_mutex_unlock(&s->lock); + + trace_qcow2_wp_tracking(index, *wp >> BDRV_SECTOR_BITS); + return 0; +} + +/* The qcow2 state lock must be held when called */ +static bool qcow2_can_activate_zone(BlockDriverState *bs) +{ + BDRVQcow2State *s =3D bs->opaque; + + /* When the max active zone is zero, there is no limit on active zones= */ + if (!s->zoned_header.max_active_zones) { + return true; + } + + /* Active zones are zones that are open or closed */ + return s->nr_zones_exp_open + s->nr_zones_imp_open + s->nr_zones_closed + < s->zoned_header.max_active_zones; +} + +/* + * This function manages open zones under active zones limit. It checks + * if a zone can transition to open state while maintaining max open and + * active zone limits. + * + * The qcow2 state lock must be held when called. + */ +static bool qcow2_can_open_zone(BlockDriverState *bs) +{ + BDRVQcow2State *s =3D bs->opaque; + + /* When the max open zone is zero, there is no limit on open zones */ + if (!s->zoned_header.max_open_zones) { + return true; + } + + /* + * The open zones are zones with the states of explicitly and + * implicitly open. + */ + if (s->nr_zones_imp_open + s->nr_zones_exp_open < + s->zoned_header.max_open_zones) { + return true; + } + + /* + * Zones are managed one at a time. Thus, the number of implicitly open + * zone can never be over the open zone limit. When the active zone li= mit + * is not reached, close only one implicitly open zone. + * + * Only an implicitly open zone can be evicted to make room. If every + * open zone is explicitly open, the open zone limit cannot be satisfi= ed. + */ + if (s->nr_zones_imp_open > 0 && qcow2_can_activate_zone(bs)) { + qcow2_do_close_zone(bs, -1, BLK_ZS_IOPEN); + trace_qcow2_imp_open_zones(0x23, s->nr_zones_imp_open); + return true; + } + return false; +} + +static inline int coroutine_fn GRAPH_RDLOCK +qcow2_refresh_zonedmeta(BlockDriverState *bs) +{ + int ret; + BDRVQcow2State *s =3D bs->opaque; + uint64_t wps_size =3D s->zoned_header.nr_zones * sizeof(uint64_t); + g_autofree uint64_t *temp =3D NULL; + + QEMU_BUILD_BUG_ON(!BDRV_ZT_IS_CONV(QCOW2_WP_CONV_BIT)); + + temp =3D g_new(uint64_t, s->zoned_header.nr_zones); + ret =3D bdrv_pread(bs->file, s->zoned_header.zonedmeta_offset, + wps_size, temp, 0); + if (ret < 0) { + error_report("Cannot read metadata"); + return ret; + } + + for (uint32_t i =3D 0; i < s->zoned_header.nr_zones; i++) { + bs->wps->wp[i] =3D be64_to_cpu(temp[i]); + } + return 0; +} + /* * Returns true if zone_opt is valid, false otherwise. */ @@ -547,6 +861,8 @@ qcow2_read_extensions(BlockDriverState *bs, uint64_t st= art_offset, be32_to_cpu(zoned_ext.max_active_zones); zoned_ext.max_append_bytes =3D be32_to_cpu(zoned_ext.max_append_bytes); + zoned_ext.zonedmeta_offset =3D + be64_to_cpu(zoned_ext.zonedmeta_offset); s->zoned_header =3D zoned_ext; =20 /* validate the header first */ @@ -593,6 +909,38 @@ qcow2_read_extensions(BlockDriverState *bs, uint64_t s= tart_offset, } } =20 + bs->wps =3D g_malloc(sizeof(BlockZoneWps) + + s->zoned_header.nr_zones * sizeof(uint64_t)); + ret =3D qcow2_refresh_zonedmeta(bs); + if (ret < 0) { + g_free(bs->wps); + bs->wps =3D NULL; + return ret; + } + + s->zone_list_entries =3D g_new0(Qcow2ZoneListEntry, + zoned_ext.nr_zones); + QTAILQ_INIT(&s->exp_open_zones); + QTAILQ_INIT(&s->imp_open_zones); + QTAILQ_INIT(&s->closed_zones); + qemu_co_mutex_init(&bs->wps->colock); + + s->zone_wp_state =3D g_new0(Qcow2ZoneWPState, zoned_ext.nr_zon= es); + for (uint32_t i =3D 0; i < zoned_ext.nr_zones; i++) { + QTAILQ_INIT(&s->zone_wp_state[i].in_flight); + QTAILQ_INIT(&s->zone_wp_state[i].completed_pending); + qemu_co_queue_init(&s->zone_wp_state[i].drain); + } + + s->wp_cache =3D qcow2_cache_create(bs, + DIV_ROUND_UP(zoned_ext.nr_zones * sizeof(uint64_t), + s->cluster_size), + s->cluster_size); + if (!s->wp_cache) { + error_setg(errp, "Could not allocate the write pointer cac= he"); + return -ENOMEM; + } + #ifdef DEBUG_EXT printf("Qcow2: Got zoned format extension: " "offset=3D%" PRIu64 "\n", offset); @@ -2946,21 +3294,292 @@ static coroutine_fn GRAPH_RDLOCK int qcow2_co_pwri= tev_task_entry(AioTask *task) t->l2meta); } =20 +/* + * Walk the per-zone completed_pending list while the head's LBA equals + * the cached persisted-WP, advancing persisted-WP and marking each + * covered request RESOLVED. On any advance, the new persisted-WP is + * written into the WP cache and depends-on qcow2 L2 table cache is + * established so write pointer is updated after the corresponding write + * completes. + * + * Caller holds bs->wps->colock. + */ static int coroutine_fn GRAPH_RDLOCK -qcow2_co_pwritev_part(BlockDriverState *bs, int64_t offset, int64_t bytes, - QEMUIOVector *qiov, size_t qiov_offset, - BdrvRequestFlags flags) +qcow2_wp_greedy_advance_locked(BlockDriverState *bs, uint32_t index) +{ + BDRVQcow2State *s =3D bs->opaque; + Qcow2ZoneWPState *zs =3D &s->zone_wp_state[index]; + uint64_t wp_byte_off =3D sizeof(uint64_t) * index; + uint64_t cluster_file_off =3D + s->zoned_header.zonedmeta_offset + + (wp_byte_off & ~((uint64_t)s->cluster_size - 1)); + size_t off_in_cluster =3D wp_byte_off & (s->cluster_size - 1); + void *cluster_buf; + uint64_t *slot; + uint64_t persisted_wp, new_persisted_wp; + Qcow2WPReq *head; + bool any_advanced =3D false; + int ret; + + qemu_co_mutex_lock(&s->lock); + ret =3D qcow2_cache_get(bs, s->wp_cache, cluster_file_off, &cluster_bu= f); + if (ret < 0) { + qemu_co_mutex_unlock(&s->lock); + return ret; + } + slot =3D (uint64_t *)((char *)cluster_buf + off_in_cluster); + persisted_wp =3D be64_to_cpu(*slot); + new_persisted_wp =3D persisted_wp; + + QTAILQ_FOREACH(head, &zs->completed_pending, entry) { + if (head->lba !=3D new_persisted_wp) { + break; + } + new_persisted_wp +=3D head->len; + any_advanced =3D true; + } + + if (!any_advanced) { + qcow2_cache_put(s->wp_cache, &cluster_buf); + qemu_co_mutex_unlock(&s->lock); + return 0; + } + + *slot =3D cpu_to_be64(new_persisted_wp); + qcow2_cache_entry_mark_dirty(s->wp_cache, cluster_buf); + qcow2_cache_put(s->wp_cache, &cluster_buf); + + qcow2_cache_set_dependency(bs, s->wp_cache, s->l2_table_cache); + qemu_co_mutex_unlock(&s->lock); + + while (!QTAILQ_EMPTY(&zs->completed_pending)) { + head =3D QTAILQ_FIRST(&zs->completed_pending); + if (head->lba >=3D new_persisted_wp) { + break; + } + QTAILQ_REMOVE(&zs->completed_pending, head, entry); + head->state =3D QCOW2_WP_REQ_RESOLVED; + qemu_co_queue_restart_all(&head->wait); + } + + return 0; +} + +/* + * Insert req into the zone's completed_pending list in ascending LBA + * order. Data writes may complete out of order; greedy_advance walks + * the list head while head->lba =3D=3D persisted_wp, so the list must + * stay sorted. + * + * Caller holds bs->wps->colock. + */ +static void +qcow2_wp_insert_pending_locked(Qcow2ZoneWPState *zs, Qcow2WPReq *req) +{ + Qcow2WPReq *iter; + + req->state =3D QCOW2_WP_REQ_PENDING; + + QTAILQ_FOREACH(iter, &zs->completed_pending, entry) { + if (iter->lba > req->lba) { + QTAILQ_INSERT_BEFORE(iter, req, entry); + return; + } + } + QTAILQ_INSERT_TAIL(&zs->completed_pending, req, entry); +} + +/* + * Peer requests with lba > failed_lba are marked ABORTED. Survivors + * with lba < failed_lba are left untouched. Logical WP is rolled back + * to the contiguous extent above persisted-WP that remains covered by + * survivor requests. + * + * Caller holds bs->wps->colock. + */ +static void +qcow2_wp_abort_higher_peers_locked(BlockDriverState *bs, uint32_t index, + uint64_t failed_lba) +{ + BDRVQcow2State *s =3D bs->opaque; + Qcow2ZoneWPState *zs =3D &s->zone_wp_state[index]; + Qcow2WPReq *r, *next; + uint64_t logical_wp; + + QTAILQ_FOREACH(r, &zs->in_flight, entry) { + if (r->lba > failed_lba && r->state !=3D QCOW2_WP_REQ_ABORTED) { + r->state =3D QCOW2_WP_REQ_ABORTED; + qemu_co_queue_restart_all(&r->wait); + } + } + QTAILQ_FOREACH_SAFE(r, &zs->completed_pending, entry, next) { + if (r->lba > failed_lba) { + r->state =3D QCOW2_WP_REQ_ABORTED; + QTAILQ_REMOVE(&zs->completed_pending, r, entry); + qemu_co_queue_restart_all(&r->wait); + } + } + + /* Rewind logical_wp to the highest survivor end_offset. */ + logical_wp =3D failed_lba; + QTAILQ_FOREACH(r, &zs->in_flight, entry) { + if (r->lba < failed_lba && r->lba + r->len > logical_wp) { + logical_wp =3D r->lba + r->len; + } + } + QTAILQ_FOREACH(r, &zs->completed_pending, entry) { + if (r->lba < failed_lba && r->lba + r->len > logical_wp) { + logical_wp =3D r->lba + r->len; + } + } + bs->wps->wp[index] =3D logical_wp; +} + +/* + * If it is an append write request, the offset pointer needs to be update= d to + * the wp value of that zone after the IO completion. The unique pointer is + * passed on to this function to prevent the value being changed in condit= ion of + * multiple concurrent writes. + * + * For host-managed zoned writes, the WP lock is acquired across the WP + * check and advance, with the qcow2 state lock nested for the zone + * state TAILQ traversals. + */ +static int coroutine_fn GRAPH_RDLOCK +qcow2_co_pwv_part(BlockDriverState *bs, int64_t *offset_ptr, int64_t bytes, + QEMUIOVector *qiov, size_t qiov_offset, bool is_append, + BdrvRequestFlags flags) { BDRVQcow2State *s =3D bs->opaque; int offset_in_cluster; int ret; unsigned int cur_bytes; /* number of sectors in current iteration */ uint64_t host_offset; + int64_t offset =3D *offset_ptr; QCowL2Meta *l2meta =3D NULL; AioTaskPool *aio =3D NULL; + int64_t start_offset, start_bytes; + BlockZoneState zs; + int64_t end_zone, end_offset; + uint64_t *wp; + int64_t zone_size =3D bs->bl.zone_size; + int64_t zone_capacity =3D bs->bl.zone_capacity; + int index =3D 0; + Qcow2WPReq req; + Qcow2WPReq *wp_req =3D NULL; =20 trace_qcow2_writev_start_req(qemu_coroutine_self(), offset, bytes); =20 + start_offset =3D offset; + start_bytes =3D bytes; + if (bs->bl.zoned =3D=3D BLK_Z_HM) { + index =3D start_offset / zone_size; + wp =3D &bs->wps->wp[index]; + qemu_co_mutex_lock(&bs->wps->colock); + if (QCOW2_ZT_IS_CONV(*wp)) { + qemu_co_mutex_unlock(&bs->wps->colock); + } else { + if (offset !=3D *wp && !is_append) { + /* The write offset must be equal to the zone write pointe= r */ + qemu_co_mutex_unlock(&bs->wps->colock); + error_report("Offset 0x%" PRIx64 " of regular writes must = be " + "equal to the zone write pointer 0x%" PRIx64 = "", + offset, *wp); + return -EIO; + } + + if (is_append) { + /* + * The offset of append write is the write pointer value of + * that zone. + */ + start_offset =3D *wp; + } + + end_offset =3D start_offset + start_bytes; + + qemu_co_mutex_lock(&s->lock); + + /* Only allow writes when there are zone resources left */ + zs =3D qcow2_get_zone_state(bs, index); + if (zs =3D=3D BLK_ZS_CLOSED || zs =3D=3D BLK_ZS_EMPTY) { + if (!qcow2_can_open_zone(bs)) { + qemu_co_mutex_unlock(&s->lock); + qemu_co_mutex_unlock(&bs->wps->colock); + error_report("no more open zones available"); + return -EINVAL; + } + } + + /* + * Align up (start_offset, zone_size), the start offset is not + * necessarily power of two. + */ + end_zone =3D index * zone_size + zone_capacity; + /* Write cannot exceed the zone capacity. */ + if (end_offset > end_zone) { + qemu_co_mutex_unlock(&s->lock); + qemu_co_mutex_unlock(&bs->wps->colock); + error_report("write exceeds zone capacity with end_offset:= " + "0x%" PRIx64 ", end_zone: 0x%" PRIx64, + end_offset / 512, end_zone / 512); + return -EIO; + } + + /* + * Real drives change states before it can write to the zone. = If + * the write fails, the zone state may have changed. + * + * The zone state transitions to implicit open when the origin= al + * state is empty or closed. When the wp reaches the end, the + * open states (explicit open, implicit open) become full. + */ + zs =3D qcow2_get_zone_state(bs, index); + if (end_offset =3D=3D end_zone) { + /* Reaching the zone capacity implies full state */ + qcow2_rm_open_zone(s, index); + trace_qcow2_imp_open_zones(0x24, + s->nr_zones_imp_open); + } else { + qcow2_do_imp_open_zone(s, index, zs); + trace_qcow2_imp_open_zones(0x24, + s->nr_zones_imp_open); + } + + qemu_co_mutex_unlock(&s->lock); + + /* + * Submission for a zone append write. The logical-WP is updat= ed + * while the on-disk WP is not touched. + */ + if (is_append) { + start_offset =3D *wp; + end_offset =3D start_offset + start_bytes; + end_zone =3D (uint64_t)index * zone_size + zone_capacity; + if (end_offset > end_zone) { + qemu_co_mutex_unlock(&bs->wps->colock); + error_report("append: end_offset 0x%" PRIx64 + " > end_zone 0x%" PRIx64, + end_offset, end_zone); + return -EIO; + } + *offset_ptr =3D start_offset; + offset =3D start_offset; + } + + wp_req =3D &req; + wp_req->lba =3D start_offset; + wp_req->len =3D start_bytes; + wp_req->state =3D QCOW2_WP_REQ_INFLIGHT; + qemu_co_queue_init(&wp_req->wait); + QTAILQ_INSERT_TAIL(&s->zone_wp_state[index].in_flight, + wp_req, entry); + + *wp =3D end_offset; + qemu_co_mutex_unlock(&bs->wps->colock); + } + } + while (bytes !=3D 0 && aio_task_pool_status(aio) =3D=3D 0) { =20 l2meta =3D NULL; @@ -3021,14 +3640,81 @@ fail_nometa: if (ret =3D=3D 0) { ret =3D aio_task_pool_status(aio); } - g_free(aio); + g_free(aio); + } + + if (wp_req !=3D NULL) { + qemu_co_mutex_lock(&bs->wps->colock); + QTAILQ_REMOVE(&s->zone_wp_state[index].in_flight, wp_req, entry); + if (QTAILQ_EMPTY(&s->zone_wp_state[index].in_flight)) { + qemu_co_queue_restart_all(&s->zone_wp_state[index].drain); + } + + if (wp_req->state =3D=3D QCOW2_WP_REQ_ABORTED) { + /* + * A peer's failure handler aborted us. Whether our data + * write itself succeeded or not, reject it. + */ + qemu_co_mutex_unlock(&bs->wps->colock); + ret =3D ret < 0 ? ret : -EIO; + goto wp_done; + } + + if (ret < 0) { + /* + * This req's data write failed. Higher-LBA peers (still + * in_flight or already completed_pending) are marked ABORTED, + * their waiters woken; logical-WP rewinds to the highest + * surviving end_offset below this LBA. The zone remains usabl= e. + */ + qcow2_wp_abort_higher_peers_locked(bs, index, wp_req->lba); + qemu_co_mutex_unlock(&bs->wps->colock); + goto wp_done; + } + + qcow2_wp_insert_pending_locked(&s->zone_wp_state[index], wp_req); + + ret =3D qcow2_wp_greedy_advance_locked(bs, index); + if (ret < 0) { + if (wp_req->state =3D=3D QCOW2_WP_REQ_PENDING) { + QTAILQ_REMOVE(&s->zone_wp_state[index].completed_pending, + wp_req, entry); + } + qcow2_wp_abort_higher_peers_locked(bs, index, wp_req->lba); + qemu_co_mutex_unlock(&bs->wps->colock); + goto wp_done; + } + + /* Block until this write pointer req is RESOLVED or ABORTED. */ + while (wp_req->state =3D=3D QCOW2_WP_REQ_PENDING) { + qemu_co_queue_wait(&wp_req->wait, &bs->wps->colock); + } + + if (wp_req->state =3D=3D QCOW2_WP_REQ_ABORTED) { + ret =3D -EIO; + } else { + assert(wp_req->state =3D=3D QCOW2_WP_REQ_RESOLVED); + ret =3D 0; + } + qemu_co_mutex_unlock(&bs->wps->colock); } =20 +wp_done: trace_qcow2_writev_done_req(qemu_coroutine_self(), ret); =20 return ret; } =20 +static int coroutine_fn GRAPH_RDLOCK +qcow2_co_pwritev_part(BlockDriverState *bs, int64_t offset, int64_t bytes, + QEMUIOVector *qiov, size_t qiov_offset, + BdrvRequestFlags flags) +{ + return qcow2_co_pwv_part(bs, &offset, bytes, qiov, qiov_offset, false, + flags); +} + + static int GRAPH_RDLOCK qcow2_inactivate(BlockDriverState *bs) { BDRVQcow2State *s =3D bs->opaque; @@ -3057,6 +3743,15 @@ static int GRAPH_RDLOCK qcow2_inactivate(BlockDriver= State *bs) strerror(-ret)); } =20 + if (s->wp_cache) { + ret =3D qcow2_cache_flush(bs, s->wp_cache); + if (ret) { + result =3D ret; + error_report("Failed to flush the WP cache: %s", + strerror(-ret)); + } + } + if (result =3D=3D 0) { qcow2_mark_clean(bs); } @@ -3064,6 +3759,25 @@ static int GRAPH_RDLOCK qcow2_inactivate(BlockDriver= State *bs) return result; } =20 +static void qcow2_do_close_all_zone(BDRVQcow2State *s) +{ + Qcow2ZoneListEntry *zone_entry, *next; + + QTAILQ_FOREACH_SAFE(zone_entry, &s->imp_open_zones, imp_open_zone_entr= y, + next) { + QTAILQ_REMOVE(&s->imp_open_zones, zone_entry, imp_open_zone_entry); + s->nr_zones_imp_open--; + } + + QTAILQ_FOREACH_SAFE(zone_entry, &s->exp_open_zones, exp_open_zone_entr= y, + next) { + QTAILQ_REMOVE(&s->exp_open_zones, zone_entry, exp_open_zone_entry); + s->nr_zones_exp_open--; + } + + assert(s->nr_zones_imp_open + s->nr_zones_exp_open =3D=3D 0); +} + static void coroutine_mixed_fn GRAPH_RDLOCK qcow2_do_close(BlockDriverState *bs, bool close_data_file) { @@ -3079,6 +3793,10 @@ qcow2_do_close(BlockDriverState *bs, bool close_data= _file) cache_clean_timer_del_and_wait(bs); qcow2_cache_destroy(s->l2_table_cache); qcow2_cache_destroy(s->refcount_block_cache); + if (s->wp_cache) { + qcow2_cache_destroy(s->wp_cache); + s->wp_cache =3D NULL; + } =20 qcrypto_block_free(s->crypto); s->crypto =3D NULL; @@ -3103,6 +3821,12 @@ qcow2_do_close(BlockDriverState *bs, bool close_data= _file) =20 qcow2_refcount_close(bs); qcow2_free_snapshots(bs); + qcow2_do_close_all_zone(s); + g_free(s->zone_list_entries); + s->zone_list_entries =3D NULL; + g_free(s->zone_wp_state); + s->zone_wp_state =3D NULL; + g_free(bs->wps); } =20 static void GRAPH_UNLOCKED qcow2_close(BlockDriverState *bs) @@ -3420,7 +4144,9 @@ int qcow2_update_header(BlockDriverState *bs) .max_active_zones =3D cpu_to_be32(s->zoned_header.max_active_zones), .max_append_bytes =3D - cpu_to_be32(s->zoned_header.max_append_bytes) + cpu_to_be32(s->zoned_header.max_append_bytes), + .zonedmeta_offset =3D + cpu_to_be64(s->zoned_header.zonedmeta_offset), }; ret =3D header_ext_add(buf, QCOW2_EXT_MAGIC_ZONED_FORMAT, &zoned_header, sizeof(zoned_header), @@ -3829,7 +4555,9 @@ qcow2_co_create(BlockdevCreateOptions *create_options= , Error **errp) int version; int refcount_order; uint64_t *refcount_table; - int ret; + uint64_t zoned_meta_size, zoned_clusterlen; + int64_t offset; + int ret, i; uint8_t compression_type =3D QCOW2_COMPRESSION_TYPE_ZLIB; =20 assert(create_options->driver =3D=3D BLOCKDEV_DRIVER_QCOW2); @@ -4192,6 +4920,48 @@ qcow2_co_create(BlockdevCreateOptions *create_option= s, Error **errp) ret =3D -EINVAL; goto unlock; } + + uint32_t nrz =3D s->zoned_header.nr_zones; + zoned_meta_size =3D sizeof(uint64_t) * nrz; + g_autofree uint64_t *meta =3D NULL; + meta =3D g_new0(uint64_t, nrz); + + for (i =3D 0; i < s->zoned_header.conventional_zones; ++i) { + meta[i] =3D i * s->zoned_header.zone_size; + meta[i] |=3D QCOW2_WP_CONV_BIT; + } + + for (; i < nrz; ++i) { + meta[i] =3D i * s->zoned_header.zone_size; + } + + offset =3D qcow2_alloc_clusters(blk_bs(blk), zoned_meta_size); + if (offset < 0) { + ret =3D offset; + error_setg_errno(errp, -ret, "Could not allocate clusters " + "for zoned metadata size"); + goto unlock; + } + s->zoned_header.zonedmeta_offset =3D offset; + + zoned_clusterlen =3D size_to_clusters(s, zoned_meta_size) + * s->cluster_size; + ret =3D qcow2_pre_write_overlap_check(blk_bs(blk), 0, offset, + zoned_clusterlen, false); + if (ret < 0) { + error_setg_errno(errp, -ret, "Overlap check failed for zoned " + "metadata"); + goto unlock; + } + for (i =3D 0; i < nrz; ++i) { + meta[i] =3D cpu_to_be64(meta[i]); + } + ret =3D bdrv_pwrite(blk_bs(blk)->file, offset, zoned_meta_size, me= ta, 0); + if (ret < 0) { + error_setg_errno(errp, -ret, "Could not write zoned metadata " + "to disk"); + goto unlock; + } } else { s->zoned_header.zoned =3D QCOW2_Z_NONE; } @@ -4597,6 +5367,439 @@ qcow2_co_pdiscard(BlockDriverState *bs, int64_t off= set, int64_t bytes) return ret; } =20 +static int coroutine_fn +qcow2_co_zone_report(BlockDriverState *bs, int64_t offset, + unsigned int *nr_zones, BlockZoneDescriptor *zones) +{ + BDRVQcow2State *s =3D bs->opaque; + uint64_t zone_size =3D s->zoned_header.zone_size; + uint64_t zone_capacity =3D s->zoned_header.zone_capacity; + int64_t capacity =3D bs->total_sectors << BDRV_SECTOR_BITS; + int64_t size =3D bs->bl.nr_zones * zone_size; + unsigned int nrz, si, i =3D 0; + + if (offset >=3D capacity) { + error_report("offset %" PRId64 " is equal to or greater than the " + "device capacity %" PRId64 "", offset, capacity); + return -EINVAL; + } + + nrz =3D ((*nr_zones) < bs->bl.nr_zones) ? (*nr_zones) : bs->bl.nr_zone= s; + si =3D offset / zone_size; /* Zone size cannot be 0 for zoned device */ + qemu_co_mutex_lock(&bs->wps->colock); + for (; i < nrz; ++i) { + if (i + si >=3D bs->bl.nr_zones) { + break; + } + + zones[i].start =3D (si + i) * zone_size; + + /* The last zone can be smaller than the zone size */ + if ((si + i + 1) =3D=3D bs->bl.nr_zones && size > capacity) { + uint32_t l =3D zone_size - (size - capacity); + zones[i].length =3D l; + zones[i].cap =3D MIN(zone_capacity, l); + } else { + zones[i].length =3D zone_size; + zones[i].cap =3D zone_capacity; + } + + uint64_t wp =3D bs->wps->wp[si + i]; + if (QCOW2_ZT_IS_CONV(wp)) { + zones[i].type =3D BLK_ZT_CONV; + zones[i].state =3D BLK_ZS_NOT_WP; + /* Clear masking bits */ + wp =3D QCOW2_GET_WP(wp); + } else { + zones[i].type =3D BLK_ZT_SWR; + qemu_co_mutex_lock(&s->lock); + zones[i].state =3D qcow2_get_zone_state(bs, si + i); + qemu_co_mutex_unlock(&s->lock); + } + zones[i].wp =3D wp; + } + qemu_co_mutex_unlock(&bs->wps->colock); + *nr_zones =3D i; + return 0; +} + +static int coroutine_fn GRAPH_RDLOCK +qcow2_open_zone(BlockDriverState *bs, uint32_t index) +{ + BDRVQcow2State *s =3D bs->opaque; + int ret; + + qemu_co_mutex_lock(&bs->wps->colock); + qemu_co_mutex_lock(&s->lock); + BlockZoneState zs =3D qcow2_get_zone_state(bs, index); + trace_qcow2_imp_open_zones(BLK_ZO_OPEN, s->nr_zones_imp_open); + + switch (zs) { + case BLK_ZS_EMPTY: + if (!qcow2_can_activate_zone(bs)) { + ret =3D -EBUSY; + goto unlock; + } + break; + case BLK_ZS_IOPEN: + qcow2_rm_imp_open_zone(s, index); + break; + case BLK_ZS_EOPEN: + qemu_co_mutex_unlock(&s->lock); + qemu_co_mutex_unlock(&bs->wps->colock); + return 0; + case BLK_ZS_CLOSED: + if (!qcow2_can_open_zone(bs)) { + ret =3D -EINVAL; + goto unlock; + } + qcow2_rm_closed_zone(s, index); + break; + case BLK_ZS_FULL: + break; + default: + ret =3D -EINVAL; + goto unlock; + } + + qcow2_do_exp_open_zone(s, index); + ret =3D 0; + +unlock: + qemu_co_mutex_unlock(&s->lock); + qemu_co_mutex_unlock(&bs->wps->colock); + return ret; +} + +static int qcow2_close_zone(BlockDriverState *bs, uint32_t index) +{ + BDRVQcow2State *s =3D bs->opaque; + int ret; + + qemu_co_mutex_lock(&bs->wps->colock); + qemu_co_mutex_lock(&s->lock); + BlockZoneState zs =3D qcow2_get_zone_state(bs, index); + + switch (zs) { + case BLK_ZS_EMPTY: + break; + case BLK_ZS_IOPEN: + break; + case BLK_ZS_EOPEN: + break; + case BLK_ZS_CLOSED: + /* Closing a closed zone is not an error */ + ret =3D 0; + goto unlock; + case BLK_ZS_FULL: + break; + default: + ret =3D -EINVAL; + goto unlock; + } + qcow2_do_close_zone(bs, index, zs); + ret =3D 0; + +unlock: + qemu_co_mutex_unlock(&s->lock); + qemu_co_mutex_unlock(&bs->wps->colock); + return ret; +} + +static int coroutine_fn GRAPH_RDLOCK +qcow2_finish_zone(BlockDriverState *bs, uint32_t index) +{ + BDRVQcow2State *s =3D bs->opaque; + int ret; + + qemu_co_mutex_lock(&bs->wps->colock); + uint64_t *wp =3D &bs->wps->wp[index]; + + /* Wait for any in-flight writes to this zone to drain. */ + if (s->zone_wp_state) { + while (!QTAILQ_EMPTY(&s->zone_wp_state[index].in_flight)) { + qemu_co_queue_wait(&s->zone_wp_state[index].drain, + &bs->wps->colock); + } + } + + qemu_co_mutex_lock(&s->lock); + BlockZoneState zs =3D qcow2_get_zone_state(bs, index); + + switch (zs) { + case BLK_ZS_EMPTY: + if (!qcow2_can_activate_zone(bs)) { + ret =3D -EBUSY; + goto unlock_state; + } + break; + case BLK_ZS_IOPEN: + qcow2_rm_imp_open_zone(s, index); + trace_qcow2_imp_open_zones(BLK_ZO_FINISH, s->nr_zones_imp_open); + break; + case BLK_ZS_EOPEN: + qcow2_rm_exp_open_zone(s, index); + break; + case BLK_ZS_CLOSED: + if (!qcow2_can_open_zone(bs)) { + ret =3D -EINVAL; + goto unlock_state; + } + qcow2_rm_closed_zone(s, index); + break; + case BLK_ZS_FULL: + ret =3D 0; + goto unlock_state; + default: + ret =3D -EINVAL; + goto unlock_state; + } + + qemu_co_mutex_unlock(&s->lock); + + uint64_t old_wp =3D *wp; + *wp =3D ((uint64_t)index + 1) * s->zoned_header.zone_size; + ret =3D qcow2_rw_wp_at(bs, wp, index, true); + if (ret < 0) { + /* Keep the in-memory write pointer consistent with the on-disk va= lue. */ + *wp =3D old_wp; + goto unlock; + } + + /* + * Flush the WP cache so the on-disk write pointer reflects the new st= ate + * on return. + */ + qemu_co_mutex_lock(&s->lock); + ret =3D qcow2_cache_flush(bs, s->wp_cache); + qemu_co_mutex_unlock(&s->lock); + goto unlock; + +unlock_state: + qemu_co_mutex_unlock(&s->lock); +unlock: + qemu_co_mutex_unlock(&bs->wps->colock); + return ret; +} + +static int coroutine_fn GRAPH_RDLOCK +qcow2_reset_zone(BlockDriverState *bs, uint32_t index, + int64_t len) +{ + BDRVQcow2State *s =3D bs->opaque; + unsigned int nrz =3D bs->bl.nr_zones; + int64_t zone_size =3D bs->bl.zone_size; + unsigned int n; + int ret =3D 0; + bool any_dirtied =3D false; + + qemu_co_mutex_lock(&bs->wps->colock); + uint64_t *wp =3D &bs->wps->wp[index]; + if (len =3D=3D bs->total_sectors << BDRV_SECTOR_BITS) { + n =3D nrz; + index =3D 0; + wp =3D &bs->wps->wp[0]; + } else { + n =3D len / zone_size; + } + + for (unsigned int i =3D 0; i < n; ++i) { + uint64_t *wp_i =3D (uint64_t *)(wp + i); + uint64_t wpi_v =3D *wp_i; + if (QCOW2_ZT_IS_CONV(wpi_v)) { + continue; + } + + /* Wait for any in-flight writes to this zone to drain. */ + if (s->zone_wp_state) { + while (!QTAILQ_EMPTY(&s->zone_wp_state[index + i].in_flight)) { + qemu_co_queue_wait(&s->zone_wp_state[index + i].drain, + &bs->wps->colock); + } + } + + qemu_co_mutex_lock(&s->lock); + BlockZoneState zs =3D qcow2_get_zone_state(bs, index + i); + switch (zs) { + case BLK_ZS_EMPTY: + break; + case BLK_ZS_IOPEN: + qcow2_rm_imp_open_zone(s, index + i); + trace_qcow2_imp_open_zones(BLK_ZO_RESET, s->nr_zones_imp_open); + break; + case BLK_ZS_EOPEN: + qcow2_rm_exp_open_zone(s, index + i); + break; + case BLK_ZS_CLOSED: + qcow2_rm_closed_zone(s, index + i); + break; + case BLK_ZS_FULL: + break; + default: + ret =3D -EINVAL; + qemu_co_mutex_unlock(&s->lock); + goto unlock; + } + qemu_co_mutex_unlock(&s->lock); + + if (zs =3D=3D BLK_ZS_EMPTY) { + continue; + } + + /* + * Zero the data extent first. Data write fires before the WP clus= ter + * hits disk. So the wp advance cannot become durable while stale = data + * is still readable. + */ + ret =3D qcow2_co_pwrite_zeroes(bs, (uint64_t)(index + i) * zone_si= ze, + zone_size, 0); + if (ret < 0) { + error_report("Failed to clear zone data at zone %u", + index + i); + goto unlock; + } + + qemu_co_mutex_lock(&s->lock); + qcow2_cache_depends_on_flush(s->wp_cache); + qemu_co_mutex_unlock(&s->lock); + + uint64_t old_wp =3D *wp_i; + *wp_i =3D (uint64_t)(index + i) * zone_size; + ret =3D qcow2_rw_wp_at(bs, wp_i, index + i, true); + if (ret < 0) { + /* Keep the in-memory write pointer consistent with the on-dis= k value. */ + *wp_i =3D old_wp; + goto unlock; + } + any_dirtied =3D true; + } + + if (any_dirtied) { + /* Single flush at the end. */ + qemu_co_mutex_lock(&s->lock); + ret =3D qcow2_cache_flush(bs, s->wp_cache); + qemu_co_mutex_unlock(&s->lock); + } + +unlock: + qemu_co_mutex_unlock(&bs->wps->colock); + return ret; +} + +static int coroutine_fn GRAPH_RDLOCK +qcow2_co_zone_mgmt(BlockDriverState *bs, BlockZoneOp op, + int64_t offset, int64_t len) +{ + BDRVQcow2State *s =3D bs->opaque; + int ret =3D 0; + int64_t capacity =3D bs->total_sectors << BDRV_SECTOR_BITS; + int64_t zone_size =3D s->zoned_header.zone_size; + int64_t zone_size_mask =3D zone_size - 1; + uint32_t index =3D offset / zone_size; + BlockZoneWps *wps =3D bs->wps; + + if (offset >=3D capacity) { + error_report("offset %" PRId64 " is equal to or greater than the " + "device capacity %" PRId64 "", offset, capacity); + return -EINVAL; + } + + if (offset & zone_size_mask) { + error_report("sector offset %" PRId64 " is not aligned to zone siz= e" + " %" PRId64 "", offset / 512, zone_size / 512); + return -EINVAL; + } + + /* The length must be zone-aligned and within the device capacity. */ + if ((len < capacity - offset && (len & zone_size_mask)) || + len > capacity - offset) { + error_report("number of sectors %" PRId64 " is not aligned to zone" + " size %" PRId64 "", len / 512, zone_size / 512); + return -EINVAL; + } + + qemu_co_mutex_lock(&wps->colock); + uint64_t wpv =3D wps->wp[index]; + qemu_co_mutex_unlock(&wps->colock); + + if (QCOW2_ZT_IS_CONV(wpv)) { + /* + * ZONE_RESET_ALL is a global operation that is allowed when the + * starting zone is conventional; the zone reset path itself skips + * conventional zones. + */ + if (op !=3D BLK_ZO_RESET || len !=3D capacity) { + error_report("zone mgmt operation 0x%x is not allowed on " + "a conventional zone", op); + return -EIO; + } + } + + switch (op) { + case BLK_ZO_OPEN: + ret =3D qcow2_open_zone(bs, index); + break; + case BLK_ZO_CLOSE: + ret =3D qcow2_close_zone(bs, index); + break; + case BLK_ZO_FINISH: + ret =3D qcow2_finish_zone(bs, index); + break; + case BLK_ZO_RESET: + ret =3D qcow2_reset_zone(bs, index, len); + break; + default: + error_report("Unsupported zone op: 0x%x", op); + ret =3D -ENOTSUP; + break; + } + return ret; +} + +static int coroutine_fn GRAPH_RDLOCK +qcow2_co_zone_append(BlockDriverState *bs, int64_t *offset, QEMUIOVector *= qiov, + BdrvRequestFlags flags) +{ + assert(flags =3D=3D 0); + int64_t capacity =3D bs->total_sectors << BDRV_SECTOR_BITS; + int64_t zone_size_mask =3D bs->bl.zone_size - 1; + int64_t iov_len =3D 0; + int64_t len =3D 0; + + if (*offset >=3D capacity) { + error_report("*offset %" PRId64 " is equal to or greater than the " + "device capacity %" PRId64 "", *offset, capacity); + return -ENOSPC; + } + + if (*offset & zone_size_mask) { + error_report("sector offset %" PRId64 " is not aligned to zone siz= e " + "%" PRId64 "", *offset / 512, bs->bl.zone_size / 512); + return -EINVAL; + } + + int64_t wg =3D bs->bl.write_granularity; + int64_t wg_mask =3D wg - 1; + for (int i =3D 0; i < qiov->niov; i++) { + iov_len =3D qiov->iov[i].iov_len; + if (iov_len & wg_mask) { + error_report("len of IOVector[%d] 0x%" PRIx64 " is not aligned= to " + "block size 0x%" PRIx64 "", i, iov_len, wg); + return -EINVAL; + } + } + len =3D qiov->size; + + if ((len >> BDRV_SECTOR_BITS) > bs->bl.max_append_sectors) { + error_report("len 0x%" PRIx64 " in sectors is greater than " + "max_append_sectors 0x%" PRIx32 "", + len >> BDRV_SECTOR_BITS, bs->bl.max_append_sectors); + return -EINVAL; + } + + return qcow2_co_pwv_part(bs, offset, len, qiov, 0, true, 0); +} + static int coroutine_fn GRAPH_RDLOCK qcow2_co_copy_range_from(BlockDriverState *bs, BdrvChild *src, int64_t src_offset, @@ -6686,6 +7889,10 @@ BlockDriver bdrv_qcow2 =3D { .bdrv_co_pwritev_compressed_part =3D qcow2_co_pwritev_compressed_pa= rt, .bdrv_make_empty =3D qcow2_make_empty, =20 + .bdrv_co_zone_report =3D qcow2_co_zone_report, + .bdrv_co_zone_mgmt =3D qcow2_co_zone_mgmt, + .bdrv_co_zone_append =3D qcow2_co_zone_append, + .bdrv_snapshot_create =3D qcow2_snapshot_create, .bdrv_snapshot_goto =3D qcow2_snapshot_goto, .bdrv_snapshot_delete =3D qcow2_snapshot_delete, diff --git a/block/qcow2.h b/block/qcow2.h index 0defb84d0c..9e7e09f6f8 100644 --- a/block/qcow2.h +++ b/block/qcow2.h @@ -298,6 +298,8 @@ typedef struct Qcow2ZoneWPState { QTAILQ_HEAD(, Qcow2WPReq) in_flight; /* PENDING reqs, sorted by lba */ QTAILQ_HEAD(, Qcow2WPReq) completed_pending; + /* zone mgmt ops wait here until in_flight has drained */ + CoQueue drain; } Qcow2ZoneWPState; =20 typedef struct Qcow2UnknownHeaderExtension { diff --git a/block/trace-events b/block/trace-events index 950c82d4b8..30a3e303ca 100644 --- a/block/trace-events +++ b/block/trace-events @@ -76,6 +76,8 @@ qcow2_writev_data(void *co, uint64_t offset) "co %p offse= t 0x%" PRIx64 qcow2_pwrite_zeroes_start_req(void *co, int64_t offset, int64_t bytes) "co= %p offset 0x%" PRIx64 " bytes %" PRId64 qcow2_pwrite_zeroes(void *co, int64_t offset, int64_t bytes) "co %p offset= 0x%" PRIx64 " bytes %" PRId64 qcow2_skip_cow(void *co, uint64_t offset, int nb_clusters) "co %p offset 0= x%" PRIx64 " nb_clusters %d" +qcow2_wp_tracking(int index, uint64_t wp) "wps[%d]: 0x%" PRIx64 +qcow2_imp_open_zones(uint8_t op, int nrz) "nr_imp_open_zones after op 0x%x= : %d" =20 # qcow2-cluster.c qcow2_alloc_clusters_offset(void *co, uint64_t offset, int bytes) "co %p o= ffset 0x%" PRIx64 " bytes %d" --=20 2.53.0 From nobody Mon Jul 27 12:11:50 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=gmail.com ARC-Seal: i=1; a=rsa-sha256; t=1783549064; cv=none; d=zohomail.com; s=zohoarc; b=iS4BCYJqCEXzWEg5bKIAewVF65WjlvVwJKT6qwTH9ckxJbkaVlliz6VTbL5myVRao2tNOOvylJOsjxUe23Vk90t4kp2EbqfxTGRWjwbhJGqfC//ZOXhPUSaHXemXXrFsOy5aHuAAuoy8jw2EzWki0Az30H2ckWtuqD6sUBGyvtI= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1783549064; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=OawgwAEBDe5M3zQW/bMmaejgW8J1szv+8A4bMXc0XjE=; b=ngD+NYZeSJ55Ql8vA+rSxImjJztNW3L++wS/H0EVqPVYzJvz/zvXkns1u/f+mff8iF2qP9LS6iNROhL+tbPSHOkpSw5xiAFebi3gRelm93OaGrDGddzC0h+q2g66mN2Ryv9sDKXDSdy8HcNTGQ6SaU8EWrydkrzxUVOcV1q6//g= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1783549064530174.34288773358003; Wed, 8 Jul 2026 15:17:44 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1whaZN-0008NI-On; Wed, 08 Jul 2026 18:16:41 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1whaZL-0008MM-VQ for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:40 -0400 Received: from mail-ed1-x52c.google.com ([2a00:1450:4864:20::52c]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1whaZI-0004Hg-Tc for qemu-devel@nongnu.org; Wed, 08 Jul 2026 18:16:39 -0400 Received: by mail-ed1-x52c.google.com with SMTP id 4fb4d7f45d1cf-698aa7ba320so741609a12.1 for ; Wed, 08 Jul 2026 15:16:36 -0700 (PDT) Received: from dobby ([2a02:8109:a394:4800:2200:181c:eef9:4c12]) by smtp.gmail.com with ESMTPSA id 4fb4d7f45d1cf-69a19d786e7sm9059237a12.16.2026.07.08.15.16.32 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 15:16:33 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783548995; x=1784153795; darn=nongnu.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=OawgwAEBDe5M3zQW/bMmaejgW8J1szv+8A4bMXc0XjE=; b=OgtX9rwJyZEFbnWZJzYFy6Qs7yrK7mfvtLpgb8P+udinVti75MgPgWGtmHrNRq78A4 r1LmZ4r5l3bpvCYvjCznJqYXsfd4Y1vcaEf2yqDkgR2eKjE+CVAXMpE0sDx0K2vuVQrF vjRwISpZ8O48n8V/UAPoL0h398ihUpZscq43WuXq9gHgUkASFTWON91g0CGxaCvfJyOL +0WAllop5UhzQAE/rJzmVYhrtv4z8odOlMj6uy2g7NDM41o7SlrB5b6099l8hCDSMMpS 2NvflzOaNh7Gq9SfuJwvRYW/QredVsKCVyxWJONZReJzeE1sES3/8bGHhYkhvWRmsSH7 8AYA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783548995; x=1784153795; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=OawgwAEBDe5M3zQW/bMmaejgW8J1szv+8A4bMXc0XjE=; b=g5J9PRAnoUPXdFDQ0oFFUq95MKDlejJ8zWe8Dg3logN/o0x7FtsDO+d2SWSfoIVCLq Vm8M++/+MM7qGLmKrmpwoXLKGXPqJo+rUdyi95bNu9BgcH+Fk5SZFQynMs+wmVVTpmIe j30kgdvVaeTFEEekLrZ2bHFeRcfUfNSP0DI51E7fYnKJQL72zAsNDBoMruQ1V2a+rGCH +JCZPkPccGfMhu5cdHUHg93O5FpHqzzDXBAauv4ZNy0FB3yt8aGsXdTW7/5la+9NlFOU LNupd6h+Sj5YMrWYIQLAQKgNFyXkOsAGFmkfaHHCGyZ3SP3voXRtYwjMtchFwT0R8s6b Jzwg== X-Gm-Message-State: AOJu0Yw9RnEbL39RNP+aDac7cEF+Zv+bZ0bgOEHMMMK/1bjTaA00Gzta t2q5wKwSMfFQZNh1+KZuJJDEmhBuL6YLNNELIDxHvt93qHUZcaUb7ZPgyhvjKiBJmsk= X-Gm-Gg: AfdE7cltrrtzpppL+XHefbPe08883IEaxn/pEBVq1+OC2VsgIhyEjUsmz+h0do8i2f1 BVn0ZyDfhXBjsfyTwziLiwfgIA+NHt/3otnZj/zDEJ4yBt5HsZoi2OCKoUvakY7YsJRMQ7t/78V iT3NhC8V+wsF2jLGbIcGb7Xu8VtU+Lp2gAK1eQRfl6ZM0YAZKlJU+pnemThkg7QAin59Aap8yPY xFpkoJ77d+JLlLXQHcNELKk6klSnoP8Zt+GJ20xcKgy3/P/1qS3x+aGl7aIF0Vh2YXptwg6Z9r6 BqHKzyEM8qdMVfGR3lVdj6LzmqMG2mIrxK5KP00Bm3d+AbrAAzcPfv1VmEIha+fzkWWviLNh/ci FXUqk0URGP114ZPmC4S+7iGTFsLRWzTB3JP5uS0K2/ghuUQvZ3dlQ2HdyJ7aQ1SkNxmsYq2hiL0 6+Pk4MTDTNA93SNt5XkoujyZL4HT4H5qgjssD/6aCgDIT0bQ== X-Received: by 2002:a05:6402:2681:b0:697:dc30:8fbf with SMTP id 4fb4d7f45d1cf-69c01b80dd6mr77550a12.6.1783548995166; Wed, 08 Jul 2026 15:16:35 -0700 (PDT) From: Sam Li To: qemu-devel@nongnu.org Cc: dlemoal@kernel.org, Stefan Hajnoczi , Pierrick Bouvier , "Michael S. Tsirkin" , Hanna Reitz , cassel@kernel.org, Markus Armbruster , Kevin Wolf , Eric Blake , qemu-block@nongnu.org, Sam Li Subject: [PATCH v14 6/6] iotests: test the zoned format feature for qcow2 file Date: Thu, 9 Jul 2026 00:16:15 +0200 Message-ID: <20260708221615.346155-7-faithilikerun@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260708221615.346155-1-faithilikerun@gmail.com> References: <20260708221615.346155-1-faithilikerun@gmail.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=2a00:1450:4864:20::52c; envelope-from=faithilikerun@gmail.com; helo=mail-ed1-x52c.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @gmail.com) X-ZM-MESSAGEID: 1783549066385158500 Content-Type: text/plain; charset="utf-8" The zoned format feature can be tested by: $ tests/qemu-iotests/check -qcow2 zoned-qcow2 Signed-off-by: Sam Li --- tests/qemu-iotests/tests/zoned-qcow2 | 209 +++++++++++++++++++++++ tests/qemu-iotests/tests/zoned-qcow2.out | 191 +++++++++++++++++++++ 2 files changed, 400 insertions(+) create mode 100755 tests/qemu-iotests/tests/zoned-qcow2 create mode 100644 tests/qemu-iotests/tests/zoned-qcow2.out diff --git a/tests/qemu-iotests/tests/zoned-qcow2 b/tests/qemu-iotests/test= s/zoned-qcow2 new file mode 100755 index 0000000000..d37100c8ab --- /dev/null +++ b/tests/qemu-iotests/tests/zoned-qcow2 @@ -0,0 +1,209 @@ +#!/usr/bin/env bash +# +# SPDX-License-Identifier: GPL-2.0-or-later +# +# Test zone management operations for qcow2 file. +# + +seq=3D"$(basename $0)" +echo "QA output created by $seq" +status=3D1 # failure is the default! + +file_name=3D"zbc.qcow2" +_cleanup() +{ + _cleanup_test_img + _rm_test_img "$file_name" +} +trap "_cleanup; exit \$status" 0 1 2 3 15 + +# get standard environment, filters and checks +. ../common.rc +. ../common.filter +. ../common.qemu + +# This test only runs on Linux hosts with qcow2 image files. +_supported_fmt qcow2 +_supported_proto file +_supported_os Linux + +echo +echo "=3D=3D=3D Initial image setup =3D=3D=3D" +echo + +$QEMU_IMG create -f qcow2 $file_name -o size=3D768M -o zone.size=3D64M -o \ +zone.capacity=3D64M -o zone.conventional_zones=3D0 -o zone.max_append_byte= s=3D32M \ +-o zone.max_open_zones=3D6 -o zone.max_active_zones=3D8 -o zone.mode=3Dhos= t-managed + +IMG=3D"--image-opts -n driver=3Dqcow2,file.driver=3Dfile,file.filename=3D$= file_name" +QEMU_IO_OPTIONS=3D$QEMU_IO_OPTIONS_NO_FMT + +echo +echo "=3D=3D=3D Testing a qcow2 img with zoned format =3D=3D=3D" +echo +echo "case 1: test zone operations one by one" + +echo "(1) report zones[0]:" +$QEMU_IO $IMG -c "zrp 0 1" +echo +echo "report zones[0~9]:" +$QEMU_IO $IMG -c "zrp 0 10" +echo +echo "report zones[-1]:" # zones[-1] dictates the last zone +$QEMU_IO $IMG -c "zrp 0x2C000000 2" # 0x2C000000 / 512 =3D 0x160000 +echo +echo +echo "(2) open zones[0], zones[1], zones[-1] then close, finish, reset:" +$QEMU_IO $IMG << EOF +zo 0 0x4000000 +zrp 0 1 +zo 0x4000000 0x4000000 +zrp 0x4000000 1 +zo 0x2C000000 0x4000000 +zrp 0x2C000000 2 +zc 0 0x4000000 +zrp 0 1 +zc 0x4000000 0x4000000 +zrp 0x4000000 1 +zc 0x2C000000 0x4000000 +zrp 0x2C000000 2 +zf 0 0x4000000 +zrp 0 1 +zf 64M 64M +zrp 0x4000000 2 +zf 0x2C000000 0x4000000 +zrp 0x2C000000 2 +zrs 0 0x4000000 +zrp 0 1 +zrs 0x4000000 0x4000000 +zrp 0x4000000 1 +zrs 0x2C000000 0x4000000 +zrp 0x2C000000 2 +EOF + +echo +echo "(3) append write with (4k, 8k) data" +$QEMU_IO $IMG -c "zrp 0 12" # the physical block size of the device is 4096 +echo "Append write zones[0], zones[1] twice" +$QEMU_IO $IMG << EOF +zap -p 0 0x1000 0x2000 +zrp 0 1 +zap -p 0 0x1000 0x2000 +zrp 0 1 +zap -p 0x4000000 0x1000 0x2000 +zrp 0x4000000 1 +zap -p 0x4000000 0x1000 0x2000 +zrp 0x4000000 1 +EOF + +echo +echo "Reset all:" +$QEMU_IO $IMG -c "zrp 0 12" -c "zrs 0 768M" -c "zrp 0 12" +echo +echo + +echo "case 2: test a sets of ops that works or not" +echo "(1) append write (4k, 4k) and then write to full" +$QEMU_IO $IMG << EOF +zap -p 0 0x1000 0x1000 +zrp 0 1 +zap -p 0 0x1000 0x1ffd000 +zap -p 0 0x1000000 0x1000000 +zrp 0 1 +EOF + +echo "Reset zones[0]:" +$QEMU_IO $IMG -c "zrs 0 64M" -c "zrp 0 1" + +echo "(2) write in zones[0], zones[3], zones[8], and then reset all" +$QEMU_IO $IMG << EOF +zap -p 0 0x1000 0x1000 +zap -p 0xc000000 0x1000 0x1000 +zap -p 0x20000000 0x1000 0x1000 +zrp 0 12 +zrs 0 768M +zrp 0 12 +EOF + +echo "case 3: test zone resource management" +echo "(1) write in zones[0], zones[1], zones[2] and then close it" +$QEMU_IO $IMG << EOF +zap -p 0 0x1000 0x1000 +zap -p 0x4000000 0x1000 0x1000 +zap -p 0x8000000 0x1000 0x1000 +zrp 0 12 +zc 0 64M +zc 0x4000000 64M +zc 0x8000000 64M +zrp 0 12 +EOF + +echo "(2) reset all after 3(1)" +$QEMU_IO $IMG << EOF +zrs 0 768M +zrp 0 12 +EOF + +echo +echo "case 4: WP cache crash consistency under concurrent appends" +echo "(1) concurrent writes to the same sequential zone (zone 5 @ 320M)" +# Three concurrent aio_writes at WP, WP+4K, WP+8K. Data writes are +# permitted to complete out of order +$QEMU_IO $IMG < qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x0, zcond:3,= [type: 2] +qemu-io> qemu-io> start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20000, = zcond:3, [type: 2] +qemu-io> qemu-io> start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000= , zcond:3, [type: 2] +qemu-io> qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x0, zcond:1,= [type: 2] +qemu-io> qemu-io> start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20000, = zcond:1, [type: 2] +qemu-io> qemu-io> start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000= , zcond:1, [type: 2] +qemu-io> qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x20000, zcon= d:14, [type: 2] +qemu-io> qemu-io> start: 0x20000, len 0x20000, cap 0x20000, wptr 0x40000, = zcond:14, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40000, zcond:1, [type: 2] +qemu-io> qemu-io> start: 0x160000, len 0x20000, cap 0x20000, wptr 0x180000= , zcond:14, [type: 2] +qemu-io> qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x0, zcond:1,= [type: 2] +qemu-io> qemu-io> start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20000, = zcond:1, [type: 2] +qemu-io> qemu-io> start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000= , zcond:1, [type: 2] +qemu-io> +(3) append write with (4k, 8k) data +start: 0x0, len 0x20000, cap 0x20000, wptr 0x0, zcond:1, [type: 2] +start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20000, zcond:1, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40000, zcond:1, [type: 2] +start: 0x60000, len 0x20000, cap 0x20000, wptr 0x60000, zcond:1, [type: 2] +start: 0x80000, len 0x20000, cap 0x20000, wptr 0x80000, zcond:1, [type: 2] +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, [type: 2] +start: 0x100000, len 0x20000, cap 0x20000, wptr 0x100000, zcond:1, [type: = 2] +start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120000, zcond:1, [type: = 2] +start: 0x140000, len 0x20000, cap 0x20000, wptr 0x140000, zcond:1, [type: = 2] +start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000, zcond:1, [type: = 2] +Append write zones[0], zones[1] twice +qemu-io> After zap done, the append sector is 0x0 +qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x18, zcond:2, [type: = 2] +qemu-io> After zap done, the append sector is 0x18 +qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x30, zcond:2, [type: = 2] +qemu-io> After zap done, the append sector is 0x20000 +qemu-io> start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20018, zcond:2, = [type: 2] +qemu-io> After zap done, the append sector is 0x20018 +qemu-io> start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20030, zcond:2, = [type: 2] +qemu-io> +Reset all: +start: 0x0, len 0x20000, cap 0x20000, wptr 0x30, zcond:4, [type: 2] +start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20030, zcond:4, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40000, zcond:1, [type: 2] +start: 0x60000, len 0x20000, cap 0x20000, wptr 0x60000, zcond:1, [type: 2] +start: 0x80000, len 0x20000, cap 0x20000, wptr 0x80000, zcond:1, [type: 2] +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, [type: 2] +start: 0x100000, len 0x20000, cap 0x20000, wptr 0x100000, zcond:1, [type: = 2] +start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120000, zcond:1, [type: = 2] +start: 0x140000, len 0x20000, cap 0x20000, wptr 0x140000, zcond:1, [type: = 2] +start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000, zcond:1, [type: = 2] +start: 0x0, len 0x20000, cap 0x20000, wptr 0x0, zcond:1, [type: 2] +start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20000, zcond:1, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40000, zcond:1, [type: 2] +start: 0x60000, len 0x20000, cap 0x20000, wptr 0x60000, zcond:1, [type: 2] +start: 0x80000, len 0x20000, cap 0x20000, wptr 0x80000, zcond:1, [type: 2] +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, [type: 2] +start: 0x100000, len 0x20000, cap 0x20000, wptr 0x100000, zcond:1, [type: = 2] +start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120000, zcond:1, [type: = 2] +start: 0x140000, len 0x20000, cap 0x20000, wptr 0x140000, zcond:1, [type: = 2] +start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000, zcond:1, [type: = 2] + + +case 2: test a sets of ops that works or not +(1) append write (4k, 4k) and then write to full +qemu-io> After zap done, the append sector is 0x0 +qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x10, zcond:2, [type: = 2] +qemu-io> After zap done, the append sector is 0x10 +qemu-io> After zap done, the append sector is 0x10000 +qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x20000, zcond:14, [ty= pe: 2] +qemu-io> Reset zones[0]: +start: 0x0, len 0x20000, cap 0x20000, wptr 0x0, zcond:1, [type: 2] +(2) write in zones[0], zones[3], zones[8], and then reset all +qemu-io> After zap done, the append sector is 0x0 +qemu-io> After zap done, the append sector is 0x60000 +qemu-io> After zap done, the append sector is 0x100000 +qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x10, zcond:2, [type: = 2] +start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20000, zcond:1, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40000, zcond:1, [type: 2] +start: 0x60000, len 0x20000, cap 0x20000, wptr 0x60010, zcond:2, [type: 2] +start: 0x80000, len 0x20000, cap 0x20000, wptr 0x80000, zcond:1, [type: 2] +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, [type: 2] +start: 0x100000, len 0x20000, cap 0x20000, wptr 0x100010, zcond:2, [type: = 2] +start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120000, zcond:1, [type: = 2] +start: 0x140000, len 0x20000, cap 0x20000, wptr 0x140000, zcond:1, [type: = 2] +start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000, zcond:1, [type: = 2] +qemu-io> qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x0, zcond:1,= [type: 2] +start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20000, zcond:1, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40000, zcond:1, [type: 2] +start: 0x60000, len 0x20000, cap 0x20000, wptr 0x60000, zcond:1, [type: 2] +start: 0x80000, len 0x20000, cap 0x20000, wptr 0x80000, zcond:1, [type: 2] +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, [type: 2] +start: 0x100000, len 0x20000, cap 0x20000, wptr 0x100000, zcond:1, [type: = 2] +start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120000, zcond:1, [type: = 2] +start: 0x140000, len 0x20000, cap 0x20000, wptr 0x140000, zcond:1, [type: = 2] +start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000, zcond:1, [type: = 2] +qemu-io> case 3: test zone resource management +(1) write in zones[0], zones[1], zones[2] and then close it +qemu-io> After zap done, the append sector is 0x0 +qemu-io> After zap done, the append sector is 0x20000 +qemu-io> After zap done, the append sector is 0x40000 +qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x10, zcond:2, [type: = 2] +start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20010, zcond:2, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40010, zcond:2, [type: 2] +start: 0x60000, len 0x20000, cap 0x20000, wptr 0x60000, zcond:1, [type: 2] +start: 0x80000, len 0x20000, cap 0x20000, wptr 0x80000, zcond:1, [type: 2] +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, [type: 2] +start: 0x100000, len 0x20000, cap 0x20000, wptr 0x100000, zcond:1, [type: = 2] +start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120000, zcond:1, [type: = 2] +start: 0x140000, len 0x20000, cap 0x20000, wptr 0x140000, zcond:1, [type: = 2] +start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000, zcond:1, [type: = 2] +qemu-io> qemu-io> qemu-io> qemu-io> start: 0x0, len 0x20000, cap 0x20000, = wptr 0x10, zcond:4, [type: 2] +start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20010, zcond:4, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40010, zcond:4, [type: 2] +start: 0x60000, len 0x20000, cap 0x20000, wptr 0x60000, zcond:1, [type: 2] +start: 0x80000, len 0x20000, cap 0x20000, wptr 0x80000, zcond:1, [type: 2] +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, [type: 2] +start: 0x100000, len 0x20000, cap 0x20000, wptr 0x100000, zcond:1, [type: = 2] +start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120000, zcond:1, [type: = 2] +start: 0x140000, len 0x20000, cap 0x20000, wptr 0x140000, zcond:1, [type: = 2] +start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000, zcond:1, [type: = 2] +qemu-io> (2) reset all after 3(1) +qemu-io> qemu-io> start: 0x0, len 0x20000, cap 0x20000, wptr 0x0, zcond:1,= [type: 2] +start: 0x20000, len 0x20000, cap 0x20000, wptr 0x20000, zcond:1, [type: 2] +start: 0x40000, len 0x20000, cap 0x20000, wptr 0x40000, zcond:1, [type: 2] +start: 0x60000, len 0x20000, cap 0x20000, wptr 0x60000, zcond:1, [type: 2] +start: 0x80000, len 0x20000, cap 0x20000, wptr 0x80000, zcond:1, [type: 2] +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, [type: 2] +start: 0x100000, len 0x20000, cap 0x20000, wptr 0x100000, zcond:1, [type: = 2] +start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120000, zcond:1, [type: = 2] +start: 0x140000, len 0x20000, cap 0x20000, wptr 0x140000, zcond:1, [type: = 2] +start: 0x160000, len 0x20000, cap 0x20000, wptr 0x160000, zcond:1, [type: = 2] +qemu-io> +case 4: WP cache crash consistency under concurrent appends +(1) concurrent writes to the same sequential zone (zone 5 @ 320M) +qemu-io> qemu-io> qemu-io> qemu-io> qemu-io> start: 0xa0000, len 0x20000, = cap 0x20000, wptr 0xa0018, zcond:2, [type: 2] +qemu-io> qemu-io> qemu-io> qemu-io> (2) concurrent writes to different seq= uential zones (zones 6, 7, 9) +qemu-io> qemu-io> qemu-io> qemu-io> qemu-io> start: 0xc0000, len 0x20000, = cap 0x20000, wptr 0xc0008, zcond:2, [type: 2] +qemu-io> start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0008, zcond:2, = [type: 2] +qemu-io> start: 0x120000, len 0x20000, cap 0x20000, wptr 0x120008, zcond:2= , [type: 2] +qemu-io> qemu-io> qemu-io> qemu-io> (3) reset zones with no in-flight writ= es (zones 5, 6) +start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, zcond:1, [type: 2] +start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, [type: 2] +(4) BLKRESETALL after concurrent writes +qemu-io> qemu-io> qemu-io> qemu-io> qemu-io> qemu-io> start: 0xa0000, len = 0x20000, cap 0x20000, wptr 0xa0008, zcond:2, [type: 2] +qemu-io> start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0008, zcond:2, = [type: 2] +qemu-io> start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0008, zcond:2, = [type: 2] +qemu-io> qemu-io> start: 0xa0000, len 0x20000, cap 0x20000, wptr 0xa0000, = zcond:1, [type: 2] +qemu-io> start: 0xc0000, len 0x20000, cap 0x20000, wptr 0xc0000, zcond:1, = [type: 2] +qemu-io> start: 0xe0000, len 0x20000, cap 0x20000, wptr 0xe0000, zcond:1, = [type: 2] +qemu-io> *** done --=20 2.53.0