From nobody Sat Sep 26 21:35:34 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=linux.alibaba.com ARC-Seal: i=1; a=rsa-sha256; t=1789454311; cv=none; d=zohomail.com; s=zohoarc; b=eb8Qj1iaDh/XhZWhD8Gm+aXSgox3YEEv+6A5UoSHi5iPOI84unE7WogHhXK5t2BppoLIS7Fmwhn9Ut2Hfek0lRBqZ6Q5Rd61aB3WOMybUznSb8NaaeLxD+M7Y2Oqg/STj2vcKK0ZrkcffICzGITkDCYpKVMVi5sLkzLodG7Ze9U= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1789454311; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=2jbStDiS254YvbhyEOUZC6f0WPv1hY7v8gUee1Q8C3c=; b=KCKJi4yaQ/u7sS7ch++soZQaLYaJ+QGSzqm/83XlyDkOizIswIh4RFvZ686uDCin5yfKXyjNFf5Y+4ammvgkyl9/Hj0NRVyJzbVzO6LINiUjqxro3xiEm4pUbd14/pT3x5INipvbrkAQJbP0wFT5CxqK8A2YCFWzMMrb8dcIt14= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1789454311166892.7796390062775; Mon, 14 Sep 2026 23:38:31 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x6Mnk-0006gp-U9; Tue, 15 Sep 2026 02:37:56 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x6Mnh-0006g1-Bo for qemu-devel@nongnu.org; Tue, 15 Sep 2026 02:37:53 -0400 Received: from [115.124.30.130] (helo=out30-130.freemail.mail.aliyun.com) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x6MnZ-0001Dg-AY for qemu-devel@nongnu.org; Tue, 15 Sep 2026 02:37:53 -0400 Received: from localhost(mailfrom:guobin@linux.alibaba.com fp:SMTPD_---0XB0Kx0r_1789454240 cluster:ay36) by smtp.aliyun-inc.com; Tue, 15 Sep 2026 14:37:21 +0800 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.alibaba.com; s=default; t=1789454241; h=From:To:Subject:Date:Message-ID:MIME-Version; bh=2jbStDiS254YvbhyEOUZC6f0WPv1hY7v8gUee1Q8C3c=; b=DZR4xPKtt4xiYhzhsARxDrmcfEPbLF4MZ3eskjNdZ+oZZLL1+rjZEJy6urhYPzzNvHbijD98KfMElEyzUKoHJd0e0c8rAu/L2IIB4Dd8wDK4mr2V8tfdl9PslKfxZwI6Vu0/EG6LYvgmGbGDornhN21H1ZkdRMLZA4SvNPCe40c= X-Alimail-AntiSpam: AC=PASS; BC=-1|-1; BR=01201311R181e4; CH=green; DM=||false|; DS=||; FP=0|-1|-1|-1|0|-1|-1|-1; HT=maildocker-contentspam033045133197; MF=guobin@linux.alibaba.com; NM=1; PH=DS; RN=4; SR=0; TI=SMTPD_---0XB0Kx0r_1789454240; From: Bin Guo To: qemu-devel@nongnu.org Cc: peterx@redhat.com, farosas@suse.de, pierrick.bouvier@oss.qualcomm.com Subject: [PATCH 1/2] docs/migration: document that postcopy can be used with multifd Date: Tue, 15 Sep 2026 14:37:18 +0800 Message-ID: <20260915063719.89031-2-guobin@linux.alibaba.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260915063719.89031-1-guobin@linux.alibaba.com> References: <20260915063719.89031-1-guobin@linux.alibaba.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Host-Lookup-Failed: Reverse DNS lookup failed for 115.124.30.130 (deferred) Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=115.124.30.130; envelope-from=guobin@linux.alibaba.com; helo=out30-130.freemail.mail.aliyun.com X-Spam_score_int: -166 X-Spam_score: -16.7 X-Spam_bar: ---------------- X-Spam_report: (-16.7 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, ENV_AND_HDR_SPF_MATCH=-0.5, RCVD_IN_DNSWL_NONE=-0.0001, RDNS_NONE=0.793, SPF_PASS=-0.001, T_SPF_HELO_TEMPERROR=0.01, UNPARSEABLE_RELAY=0.001, USER_IN_DEF_DKIM_WL=-7.5, USER_IN_DEF_SPF_WL=-7.5 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @linux.alibaba.com) X-ZM-MESSAGEID: 1789454313593158500 Content-Type: text/plain; charset="utf-8" Since commit e27418861288 ("migration: enable multifd and postcopy together") in the 10.1 release, the multifd and postcopy-ram capabilities are no longer mutually exclusive, but the postcopy documentation does not mention multifd at all, leaving users unaware that the two features can be combined. Add a short section describing how they interact: multifd carries the pages of the precopy phase, the channels are flushed and synced before the switchover (because loading a page in a multifd receive thread is not atomic with regard to a running vCPU), and the postcopy phase itself does not use the multifd channels. Signed-off-by: Bin Guo Acked-by: Fabiano Rosas Reviewed-by: Peter Xu --- docs/devel/migration/postcopy.rst | 26 ++++++++++++++++++++++++++ 1 file changed, 26 insertions(+) diff --git a/docs/devel/migration/postcopy.rst b/docs/devel/migration/postc= opy.rst index e319388d8f..570a3a1ce6 100644 --- a/docs/devel/migration/postcopy.rst +++ b/docs/devel/migration/postcopy.rst @@ -294,6 +294,32 @@ the background migration channel. Anyone who cares ab= out latencies of page faults during a postcopy migration should enable this feature. By default, it's not enabled. =20 +Postcopy with multifd +--------------------- + +The ``multifd`` capability can be enabled together with ``postcopy-ram`` +since the 10.1 QEMU release. The two features apply to different phases of +the migration: + + - During the precopy phase, guest pages are sent over the multifd channe= ls + as usual, so the initial RAM transfer can use all of them. + + - Just before switching to postcopy, the source flushes and syncs the + multifd channels. This guarantees that all the pages already queued on + those channels are loaded on the destination *before* the destination + CPUs are started, which is required because loading a page in a multifd + receive thread is not atomic with regard to a running vCPU. + + - During the postcopy phase the multifd channels are no longer used for + guest pages. Both the background stream and the pages requested by the + destination go through the background migration channel instead, or + through the preempt channel for the requested pages when postcopy + preemption is enabled. + +Consequently, multifd only speeds up the precopy phase of a postcopy +migration. The bandwidth available once postcopy has started is the same = as +without multifd. + Postcopy blocktime statistics ----------------------------- =20 --=20 2.50.1 (Apple Git-155) From nobody Sat Sep 26 21:35:34 2026 Delivered-To: importer@patchew.org Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass(p=none dis=none) header.from=linux.alibaba.com ARC-Seal: i=1; a=rsa-sha256; t=1789454311; cv=none; d=zohomail.com; s=zohoarc; b=jOlJZhvk+1ty5aj5EX0Aesj08aauLGmenFC3ZpHbfl1eI1OYleGBsJbPJroC9R84TB185qZQPIhD7xTDFcMiyaguNP7A1sKGUvA6eAyOR9yDwGMPK913pOnqaG3jFPWNP5+VWbAujIkfrmFvF1m10++u1ZhXU/ql+MjvSgiJsJQ= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1789454311; h=Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Archive:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=Qr52ItPjNoQFNh4q9cMQK9LkWdC9JRTZIHpsy1bQGyo=; b=FTWQH1DOAiKnSPnV8vGo/TY9CdSDVW7lgqWhHgwr1XIhDt2YNKvtIztM8KCWBwPSwrIKCXHILG3+CCqDFRxLLBn1wPtB+odjThTLHFN8YZxHhhgpObDHVad3IKkLp79IWRY0dCgaJaSK7oSmkXYw5CbeFREbmSjWFYTY00XNylg= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) smtp.mailfrom=qemu-devel-bounces+importer=patchew.org@nongnu.org; dmarc=pass header.from= (p=none dis=none) Return-Path: Received: from lists1p.gnu.org (lists1p.gnu.org [209.51.188.17]) by mx.zohomail.com with SMTPS id 1789454311000476.7761918126572; Mon, 14 Sep 2026 23:38:31 -0700 (PDT) Received: from localhost ([::1] helo=lists1p.gnu.org) by lists1p.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1x6Mnf-0006fW-6O; Tue, 15 Sep 2026 02:37:51 -0400 Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists1p.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x6Mnc-0006fC-GG for qemu-devel@nongnu.org; Tue, 15 Sep 2026 02:37:48 -0400 Received: from [115.124.30.113] (helo=out30-113.freemail.mail.aliyun.com) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1x6MnZ-0001Dv-1E for qemu-devel@nongnu.org; Tue, 15 Sep 2026 02:37:48 -0400 Received: from localhost(mailfrom:guobin@linux.alibaba.com fp:SMTPD_---0XB0TToh_1789454241 cluster:ay36) by smtp.aliyun-inc.com; Tue, 15 Sep 2026 14:37:22 +0800 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.alibaba.com; s=default; t=1789454243; h=From:To:Subject:Date:Message-ID:MIME-Version; bh=Qr52ItPjNoQFNh4q9cMQK9LkWdC9JRTZIHpsy1bQGyo=; b=M05nJSN7ic36x2cnodUAKQmfwQhxk3E8QSUZfkCtUS0cd1xqED68Mir4nAD2Y4IjBHowEXirdgUZLXN61uJBHtAG4fVRdyilF3UZCyPFu0Z2VpwcXO9x6NjIKc1aPc/JVFQst1Blzdta97w+T/A2kCHEYjc9ecT9Q0c0JT4ISVw= X-Alimail-AntiSpam: AC=PASS; BC=-1|-1; BR=01201311R411e4; CH=green; DM=||false|; DS=||; FP=0|-1|-1|-1|0|-1|-1|-1; HT=maildocker-contentspam033037033178; MF=guobin@linux.alibaba.com; NM=1; PH=DS; RN=4; SR=0; TI=SMTPD_---0XB0TToh_1789454241; From: Bin Guo To: qemu-devel@nongnu.org Cc: peterx@redhat.com, farosas@suse.de, pierrick.bouvier@oss.qualcomm.com Subject: [PATCH 2/2] migration/postcopy: account faults taken by memory sharing processes Date: Tue, 15 Sep 2026 14:37:19 +0800 Message-ID: <20260915063719.89031-3-guobin@linux.alibaba.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260915063719.89031-1-guobin@linux.alibaba.com> References: <20260915063719.89031-1-guobin@linux.alibaba.com> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Host-Lookup-Failed: Reverse DNS lookup failed for 115.124.30.113 (deferred) Received-SPF: pass (zohomail.com: domain of gnu.org designates 209.51.188.17 as permitted sender) client-ip=209.51.188.17; envelope-from=qemu-devel-bounces+importer=patchew.org@nongnu.org; helo=lists1p.gnu.org; Received-SPF: pass client-ip=115.124.30.113; envelope-from=guobin@linux.alibaba.com; helo=out30-113.freemail.mail.aliyun.com X-Spam_score_int: -166 X-Spam_score: -16.7 X-Spam_bar: ---------------- X-Spam_report: (-16.7 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, ENV_AND_HDR_SPF_MATCH=-0.5, RCVD_IN_DNSWL_NONE=-0.0001, RDNS_NONE=0.793, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, UNPARSEABLE_RELAY=0.001, USER_IN_DEF_DKIM_WL=-7.5, USER_IN_DEF_SPF_WL=-7.5 autolearn=no autolearn_force=no X-Spam_action: no action X-BeenThere: qemu-devel@nongnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: qemu development List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: qemu-devel-bounces+importer=patchew.org@nongnu.org Sender: qemu-devel-bounces+importer=patchew.org@nongnu.org X-ZohoMail-DKIM: pass (identity @linux.alibaba.com) X-ZM-MESSAGEID: 1789454313627158500 Content-Type: text/plain; charset="utf-8" Page faults taken by a process that shares guest memory with us (e.g. a vhost-user backend) are resolved over the migration stream just like our own, but they are not accounted for, so they are missing from the latency reports. There is no thread id to report them with: the client's userfaultfd is not required to enable UFFD_FEATURE_THREAD_ID, and libvhost-user does not, so the ptid in the fault message stays zero. Even when the feature is enabled, the reported ids belong to another process and can never match one of our vCPUs. Passing zero would not work either, as mark_postcopy_blocktime_begin() reads it as "no thread id" and returns early. Report those faults with POSTCOPY_TID_FOREIGN instead, which takes the same path as the faults of our own non-vCPU threads: no vCPU is involved, so the vCPU blocktime reports are left alone and only the latency of the fault is accounted. This also resolves the "TODO: support blocktime tracking" left in postcopy_request_shared_page(). Signed-off-by: Bin Guo --- migration/postcopy-ram.c | 26 ++++++++++++++++---------- migration/postcopy-ram.h | 10 ++++++++++ 2 files changed, 26 insertions(+), 10 deletions(-) diff --git a/migration/postcopy-ram.c b/migration/postcopy-ram.c index 5509cbdcb8..00e2a6fc62 100644 --- a/migration/postcopy-ram.c +++ b/migration/postcopy-ram.c @@ -1122,8 +1122,6 @@ int postcopy_request_shared_page(struct PostCopyFD *p= cfd, RAMBlock *rb, qemu_ram_get_idstr(rb), rb_offset); return postcopy_wake_shared(pcfd, client_addr, rb); } - /* TODO: support blocktime tracking */ - /* * The page will be placed by qemu_ufd_copy_ioctl(), which removes the * matching entry from mis->page_requested (and drops @@ -1133,10 +1131,16 @@ int postcopy_request_shared_page(struct PostCopyFD = *pcfd, RAMBlock *rb, * backend's address space and can never equal that host address, so t= he * removal would miss forever, leaking page_requested_count and hanging * postcopy teardown. + * + * We have no thread id to report for the client: its userfaultfd is n= ot + * required to enable UFFD_FEATURE_THREAD_ID, and even when it does the + * thread ids it reports belong to another process, so they could never + * match one of our vCPUs. Report the fault as a foreign one, so that= it + * still shows up in the latency reports. */ postcopy_request_page(mis, rb, aligned_rbo, (uint64_t)(uintptr_t)qemu_ram_get_host_addr(rb) + - aligned_rbo, 0); + aligned_rbo, POSTCOPY_TID_FOREIGN); return 0; } =20 @@ -1234,7 +1238,8 @@ bool try_mark_postcopy_blocktime_begin(MigrationIncom= ingState *mis, * blocking time. It's protected by @page_request_mutex. * * @addr: faulted host virtual address - * @ptid: faulted process thread id + * @ptid: faulted process thread id, or POSTCOPY_TID_FOREIGN when the fault + * was taken by a thread of another process * @rb: ramblock appropriate to addr */ void mark_postcopy_blocktime_begin(uintptr_t addr, uint32_t ptid, @@ -1256,7 +1261,7 @@ void mark_postcopy_blocktime_begin(uintptr_t addr, ui= nt32_t ptid, assert(!ramblock_recv_bitmap_test(rb, (void *)addr)); =20 current =3D get_current_ns(); - cpu =3D blocktime_get_vcpu(dc, ptid); + cpu =3D ptid =3D=3D POSTCOPY_TID_FOREIGN ? -1 : blocktime_get_vcpu(dc,= ptid); =20 if (cpu >=3D 0) { /* How many faults on this vCPU in total? */ @@ -1281,11 +1286,12 @@ void mark_postcopy_blocktime_begin(uintptr_t addr, = uint32_t ptid, } } else { /* - * For non-vCPU thread faults, we don't care about tid or cpu index - * or time the thread is blocked (e.g., a kworker trying to help - * KVM when async_pf=3Don is OK to be blocked and not affect guest - * responsiveness), but we care about latency. Track it with - * cpu=3D-1. + * For faults that did not come from a vCPU thread, we don't care + * about tid or cpu index or time the thread is blocked (e.g., a + * kworker trying to help KVM when async_pf=3Don is OK to be block= ed + * and not affect guest responsiveness; likewise for a thread of a + * process that shares guest memory with us), but we care about + * latency. Track it with cpu=3D-1. * * Note that this will NOT affect blocktime reports on vCPU being * blocked, but only about system-wide latency reports. diff --git a/migration/postcopy-ram.h b/migration/postcopy-ram.h index edba7b0240..028f774651 100644 --- a/migration/postcopy-ram.h +++ b/migration/postcopy-ram.h @@ -200,6 +200,16 @@ bool postcopy_is_paused(MigrationStatus status); bool try_mark_postcopy_blocktime_begin(MigrationIncomingState *mis, RAMBlock *rb, ram_addr_t start, uint64_t haddr, uint32_t tid); + +/* + * Thread id to report for a fault that did not come from one of our own + * threads, e.g. one taken by a process that shares guest memory with us. = It + * can never collide with a real thread id, as Linux pid values are far be= low + * this. See mark_postcopy_blocktime_begin() for how such faults are + * accounted. + */ +#define POSTCOPY_TID_FOREIGN ((uint32_t)-1) + void mark_postcopy_blocktime_begin(uintptr_t addr, uint32_t ptid, RAMBlock *rb); =20 --=20 2.50.1 (Apple Git-155)