From nobody Sat Sep 26 01:05:00 2026 Received: from mail-pj2-f1.google.com (mail-pj2-f1.google.com [74.125.227.129]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3B9A221A42D for ; Sun, 6 Sep 2026 12:23:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.129 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788697442; cv=none; b=hE1AdlZoB4k+L1OBCeHkSKqMvpV0mu2JBq6gIThDoyxjIFg0umnp2k+J25jMN+GuiicZIzI8Apo9L9qriBCNUCZpCPJTs+o3ixuXk4GOIprJzzbO5AyZhjTBGwjHLC0mKaO77mqieSGIJB1X/QyZa3TO7UzS2k9LptIPBL/lq6Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788697442; c=relaxed/simple; bh=i7T5MsATCuxq4jFi1F2opIb+g/TNVwOa5J8EBT+1tg0=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=dQY+KRQEMzDnjAraNgMcfDtza2OMLowMeBnRCTPbqXncB7L2TleJvgBNXfsoo1gJXE+lMl9bS23oGiQIEw7WLR/fgJiPL7S9xpNg0UlLhgW31mKGJsaqgwC662Kz/bmBcXcnaNIIeMJehw6hdtsFMDdByBwIiYvvLaY0FZpBCsM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=IxlBoVs2; arc=none smtp.client-ip=74.125.227.129 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="IxlBoVs2" Received: by mail-pj2-f1.google.com with SMTP id 98e67ed59e1d1-399213ef56cso1334021a91.1 for ; Sun, 06 Sep 2026 05:23:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788697439; x=1789302239; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=xaszmDPaAaT8VIgwEr8D3B1YhHgs6kESvSxQ+esjaWY=; b=IxlBoVs2Kp6Dt4Ma56OwYqWQ24XZ0nH7hrbzoiSZgyMRWdNWrPXd3Yoedt0N/Gr9cZ p6OTMS2bMeyM11mH9qtMIaBRGnJIPE5dhYBe0PNJ4gnCVkGyAKxAXXjF6y3Fu5O7aEKm ph+fI+VOnzS3igqV7ng3wPUibG6l/6THrgEzKUc6eSRZdN4XeMKZn9ugVlDyDnRqQHq6 nk9zzG8JZYrWzS+KFeQGdqNpd+mTMiaENhHYzt21z3KHihdk17Yszh6U2ypJZPWPfEaY 0e4YGh8XzVZ56/M52foFJ/WbRGXvKtnUY07Kje78UTK9zi04Ed9L2OPYm5fajBasUM5i 6WzQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788697439; x=1789302239; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=xaszmDPaAaT8VIgwEr8D3B1YhHgs6kESvSxQ+esjaWY=; b=bqtm0AObzbMVL8MWDCn9WI/vTDyKOWWMrmfLfldfFGF8gTuwRuph6EqBMA7jnay9Zn Affn9AaCBqECXwV1v2YJspaFj9B/IwawhnMf4X4uqMfAIynk20Nd3RJz6a2bqDyS0b7K 8vU3kQu2nQBJzGonl8BOU4wemfY+WMoEnAhV3GfI1k2bnIP84zYbae5vufHBnuBrwNHE N0RyAb/enL7UtpU+TaN99sUieeaj+l3nk418pW9MiBJqkLVo3TFTwKzkKsiA4xCf/xmG 9dl4bzQyIQ2WA8dUzouhg/D6eSdN5ZrqZupChVvwQq0z8uBHxKQCsvRiMITXEomOEf/t W8Xw== X-Forwarded-Encrypted: i=1; AKwUvBxrvKkM5uIspLDzi/ZIve57h46KLBKpghUDdxgWuLw4F5iWCXUgbgqcvQx0Phait5YkQhluXUG6xDsSWyg=@vger.kernel.org X-Gm-Message-State: AFuF++ldxaXVfEyTOmCV9Uk6s/WPO9+CYAjV3E8E1/MH2iaNVyj1TxhV kjko55c8IUpobLeCNDily+ez07fEuZte2jUhvZhL8ri7L44Qh9epvXFC X-Gm-Gg: AYBFou1KqDJzHKQ1kXj5COrde6s1UePmfzilvj75MrrKdJJs0vx2kQCetwvOaqu9dak hwQLxmhvgtOp5EI4U65Cd14O9Hvh7tUdiwL03Kpeq9tzNaLpTgfAF9GgTpLwjSj2Vo2SuclfR5Z /hFfmy7WenCTzWKqncObS44gZ41x9CJ/B1E/mzCXOP8wUckHj90V8rHqDxAWTZ06lX4uzbRC4iG 067d3GfcUxEN/GJ+W1L3OtvP8vXRbpgL2NXwvyFp3pVNarodNNblyvQ28kWEoTLEA1zh1JfuRv2 CKurLBj1GLpyXmQ3ZsCXFNsfSuNZtJQij+lFWTeCDYRjofdGzyEKoEeTXRbgwnbx+HblrAmuH2V sETqOvW6xy7o2js28hjQQypXnU+lJuAXMICcMQ4FQNTJCSkuMUFsYsRhbTDLwXnn6DIsvrHcZtL x24XWgQA5YR9MwRNFzabfC95C29rEqJyejC3vT7fMmPlaj7yHB1ga6D34d X-Received: by 2002:a17:90b:5346:b0:395:f0e8:9e13 with SMTP id 98e67ed59e1d1-39b26190fd9mr27161779a91.7.1788697438844; Sun, 06 Sep 2026 05:23:58 -0700 (PDT) Received: from nixos ([122.231.174.193]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-39b08bcd243sm21782207a91.5.2026.09.06.05.23.54 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 06 Sep 2026 05:23:58 -0700 (PDT) From: Thaumy Cheng To: Peter Zijlstra , Ingo Molnar , Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Alexander Shishkin , Jiri Olsa , Ian Rogers , Adrian Hunter , James Clark Cc: linux-perf-users@vger.kernel.org, linux-kernel@vger.kernel.org, Thaumy Cheng Subject: [PATCH] perf/core: Strengthen userpage update ordering Date: Sun, 6 Sep 2026 20:23:50 +0800 Message-ID: <20260906122350.24305-1-thaumy.love@gmail.com> X-Mailer: git-send-email 2.55.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The perf event mmap userpage uses a sequence counter to let userspace obtain a consistent snapshot of its data. The existing compiler barriers reflect the original self-monitoring use case, where the counter was updated and consumed on the same CPU. Some consumers also use the userpage only for time conversion and may read it from a CPU other than the one updating the event. For example, perf_read_tsc_conversion() reads the time conversion fields without constraining the caller to the event's CPU. Make the publication ordering explicit for such readers by replacing the compiler barriers around the userpage payload update with smp_wmb(). Document that cross-CPU readers of the time conversion fields must use read memory barriers and reject odd or changed sequence values. This does not change the UAPI layout or the values exposed to userspace. It strengthens the ordering guarantee for cross-CPU readers on weakly ordered architectures. Signed-off-by: Thaumy Cheng --- include/uapi/linux/perf_event.h | 6 ++++-- kernel/events/core.c | 6 ++++-- tools/include/uapi/linux/perf_event.h | 6 ++++-- tools/perf/design.txt | 6 ++++-- 4 files changed, 16 insertions(+), 8 deletions(-) diff --git a/include/uapi/linux/perf_event.h b/include/uapi/linux/perf_even= t.h index fd10aa8d697f..a7db00b9b455 100644 --- a/include/uapi/linux/perf_event.h +++ b/include/uapi/linux/perf_event.h @@ -629,8 +629,10 @@ struct perf_event_mmap_page { * barrier(); * } while (pc->lock !=3D seq); * - * NOTE: for obvious reason this only works on self-monitoring - * processes. + * NOTE: Reading the hardware counter as shown above only works for + * self-monitoring processes. A reader on another CPU may snapshot + * the time conversion fields, but must use rmb() around the field + * reads and retry if lock is odd or changes. */ __u32 lock; /* seqlock for synchronization */ __u32 index; /* hardware event identifier */ diff --git a/kernel/events/core.c b/kernel/events/core.c index 89b40e439717..72605776273f 100644 --- a/kernel/events/core.c +++ b/kernel/events/core.c @@ -6852,7 +6852,8 @@ void perf_event_update_userpage(struct perf_event *ev= ent) userpg =3D rb->user_page; =20 ++userpg->lock; - barrier(); + /* Publish the odd lock value before updating the payload. */ + smp_wmb(); userpg->index =3D perf_event_index(event); userpg->offset =3D perf_event_count(event, false); if (userpg->index) @@ -6866,7 +6867,8 @@ void perf_event_update_userpage(struct perf_event *ev= ent) =20 arch_perf_update_userpage(event, userpg, now); =20 - barrier(); + /* Publish the payload before the final lock update. */ + smp_wmb(); ++userpg->lock; preempt_enable(); unlock: diff --git a/tools/include/uapi/linux/perf_event.h b/tools/include/uapi/lin= ux/perf_event.h index fd10aa8d697f..a7db00b9b455 100644 --- a/tools/include/uapi/linux/perf_event.h +++ b/tools/include/uapi/linux/perf_event.h @@ -629,8 +629,10 @@ struct perf_event_mmap_page { * barrier(); * } while (pc->lock !=3D seq); * - * NOTE: for obvious reason this only works on self-monitoring - * processes. + * NOTE: Reading the hardware counter as shown above only works for + * self-monitoring processes. A reader on another CPU may snapshot + * the time conversion fields, but must use rmb() around the field + * reads and retry if lock is odd or changes. */ __u32 lock; /* seqlock for synchronization */ __u32 index; /* hardware event identifier */ diff --git a/tools/perf/design.txt b/tools/perf/design.txt index aa8cfeabb743..111afc90c442 100644 --- a/tools/perf/design.txt +++ b/tools/perf/design.txt @@ -316,8 +316,10 @@ struct perf_event_mmap_page { * barrier(); * } while (pc->lock !=3D seq); * - * NOTE: for obvious reason this only works on self-monitoring - * processes. + * NOTE: Reading the hardware counter as shown above only works for + * self-monitoring processes. A reader on another CPU may sn= apshot + * the time conversion fields, but must use rmb() around the= field + * reads and retry if lock is odd or changes. */ __u32 lock; /* seqlock for synchronization */ __u32 index; /* hardware counter identifier */ --=20 2.55.0