From nobody Sat Jul 25 16:48:59 2026 Received: from mail-pj1-f52.google.com (mail-pj1-f52.google.com [209.85.216.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A98BD417D6A for ; Wed, 15 Jul 2026 19:30:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784143830; cv=none; b=ABfIscTMfVxuGtr3ZnW0YffAUeoChEDnjsmqzcie7NF7LvJGZtXclSFIKIb39D5iGDNxwj3FkWD5xh1HQXx+55b1Xn8dAkyyiZ9W36J8Nz9osJzFGKPrzFgrqcy3OvjfRtRJtGA0+mavs+Gb0BnnQDeocnq43NQlfH9lqNcFekE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784143830; c=relaxed/simple; bh=/o2pQ6I0UNHS1txPe1Ba1QmitdEVq/qnXd8xOzN3DmI=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=VhXoJ+DjIXbtosbVd6I61GE1kxjlz5fNNCNh0eRZf+RjBqsr1PRCcuHo/87PfdO02rzzbXu3Yt1dVseua7SpZh0O7l+qaSuGeyJgUqbs0WyDcnQLUDfEcapi2KlJp+qknVByjmvyrMvhm3Exq5GMJ89s1/GJZy7jqJjGn/Vsxqs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=waychison.com; spf=pass smtp.mailfrom=waychison.com; dkim=pass (1024-bit key) header.d=waychison.com header.i=@waychison.com header.b=CCnFbz9c; arc=none smtp.client-ip=209.85.216.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=waychison.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=waychison.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=waychison.com header.i=@waychison.com header.b="CCnFbz9c" Received: by mail-pj1-f52.google.com with SMTP id 98e67ed59e1d1-38e1a9d9105so1975861a91.1 for ; Wed, 15 Jul 2026 12:30:27 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=waychison.com; s=google; t=1784143826; x=1784748626; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=k3p19RXCQA1WgE0TNasZ3+exAMqY6Cp1IX2RpB7+gGE=; b=CCnFbz9cAnqkw1I4PN/tYDRSCDtzOh/v55Hvz+ySX37+ibWMSofoHbjdtiOAc+zFue bRGWDortR0NEghs8bTfOWgGEjN0HVYiQ1HihmDF62IZmSjVSqqS08CVQTqxP9TKAWNHk L2f7s8pNJmLAWBqn7N/4jWiwVGD77/sJiG/9g= X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784143826; x=1784748626; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=k3p19RXCQA1WgE0TNasZ3+exAMqY6Cp1IX2RpB7+gGE=; b=P+JDJKWERu4XKBSQzWlJjot639ZtSM+rAtRHkkc4A6EsIiKnQ8Eb8mJxRgR8sKcjHJ lGnmzTdxXwE32HCcEE9nQAQcroJcHTwXvoz4qfWN1sJqoW/RWcyACwacwhMf1EyEIkLX HbufzxA+MtbqK6OYHWuvb2uYoNIzR3JMpIYxZ4iHMkkOsoQ7n/j8NeaPTRAdOfshgws+ lbTjMIQSUv8DbHZQD3fRBNjf8d/ibMGOUz0e5ZqhvsG52xiug6syJJ2nK3FsvP3/UsWn 7p3nwS7y8Jec9TENBbXQqFYbTXKMKjr9CMBWtNNUyYy3tIh2husMjWjZwo23U+i2K18L M5ng== X-Forwarded-Encrypted: i=1; AHgh+RrgIvfzJaARPni5XfNhz58aD3bmaDc5zc83RUFJD+Y7A9hAEih75KdC67vb7Wx9Z0mpCkomfra/YQ3kH8I=@vger.kernel.org X-Gm-Message-State: AOJu0Yx5OQP7T24K0JBsoZHwdLxn3FmdaQK3ZNncKZbSNt11F35H1xi5 LUUi6b9zKzqV37oysO9666I6GW5PkIBYzG/6vIMCljmo5WwLM+/ykY65TTWByZLgcvBKUXGQIhC CxOrjSUfA1O1+UFg= X-Gm-Gg: AfdE7cnGWolx1RlD8QXgFMgrbIQZgSpxlUF0obWQ8MOlLVGYyI4FUNwC4K777C7j0cn rtjxANirM87XQNaFvpRXloYWdA6YM3ZlB7lFBnBOfZFavYxCYrUR/VFqkvhj/4EK9IP9V3rQmmQ uJeHRSmraKgixNapAH+rxcvbrykhWYRZ9Y15zNF70q9/MXkzMAjvGcL0mF1hLAHwbjtzhFQXok9 DYPqsGpjEpq9nWrEs29yxCqr4BVniGbHzNqlnSaiZzksecHcpIU4YhI/1FkG9V+yNckeBbEfT+w qS0QYjH9KwKllBrKsl+tJkTu72Iw7ZPWyPMEKW20a0FrqOuYTVMqWSZEzFeZ7qj8iZvq3hjbAvh XKhf0Je94+O8l708wVVbQyW/qdy9DKtbrntmSUmzoLAlrx4YfAb7v8UrHogD4AMQFRxQzHEYQQt tIhm3ie1g= X-Received: by 2002:a17:90b:4a04:b0:387:e0bb:5802 with SMTP id 98e67ed59e1d1-38e2a12332amr3793336a91.41.1784143825821; Wed, 15 Jul 2026 12:30:25 -0700 (PDT) Received: from mike-yyz. ([184.175.42.134]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-38e39e1b7adsm84595a91.7.2026.07.15.12.30.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 15 Jul 2026 12:30:25 -0700 (PDT) From: Mike Waychison To: Jens Axboe , Usama Arif Cc: linux-block@vger.kernel.org, linux-kernel@vger.kernel.org, Mike Waychison , stable@vger.kernel.org Subject: [PATCH] block: fix race in blk_time_get_ns() returning 0 Date: Wed, 15 Jul 2026 15:29:50 -0400 Message-ID: <20260715192950.2488921-1-mike@waychison.com> X-Mailer: git-send-email 2.47.3 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" blk_time_get_ns() populates the per-plug cached timestamp and then returns it by re-reading the field: if (!plug->cur_ktime) { plug->cur_ktime =3D ktime_get_ns(); current->flags |=3D PF_BLOCK_TS; } return plug->cur_ktime; This is problematic when the compiler emits the final "return plug->cur_ktime" as a reload from memory, after PF_BLOCK_TS has already been set. Since the cached timestamp is now invalidated from finish_task_switch() (fad156c2af22 "block: invalidate cached plug timestamp after task switch"), a task preempted between setting PF_BLOCK_TS and that reload has plug->cur_ktime zeroed by blk_plug_invalidate_ts() when it is scheduled back in. The reload then returns 0. A 0 handed back here is stored as a start timestamp -- e.g. blk_account_io_start() writes it to rq->start_time_ns -- and later subtracted from "now". blk_account_io_done() then adds (now - 0), i.e. roughly the system uptime, to the per-group nsecs[] counters. On an otherwise idle, healthy device this appears as sudden ~uptime-sized jumps in the diskstats time fields (write_ticks/discard_ticks/time_in_queue). The solution is to be explicit in our reads and writes to this field that is preemption volatile. We also add a barrier() to ensure that any setting of PF_BLOCK_TS is ordered to happen after the cur_ktime update. This issue was discovered using AI-assisted kprobes looking for paths that were leaking zeroed timestamps in a live system, based on the observation that we were sometimes seeing uptime-sized jumps in kernel exported counters. This was flagged by NodeDiskIOSaturation prometheus alerts that started firing on all hosts post 7.1.3 kernel upgrade, due to node-exporter now exporting a nonsensical node_disk_io_time_weighted_seconds_total. Fixes: fad156c2af22 ("block: invalidate cached plug timestamp after task sw= itch") Cc: stable@vger.kernel.org Signed-off-by: Mike Waychison Assisted-by: Claude:claude-opus-4.8 --- block/blk.h | 13 ++++++++++--- 1 file changed, 10 insertions(+), 3 deletions(-) diff --git a/block/blk.h b/block/blk.h index b998a7761faf..17e03656ba9a 100644 --- a/block/blk.h +++ b/block/blk.h @@ -689,6 +689,7 @@ static inline int req_ref_read(struct request *req) static inline u64 blk_time_get_ns(void) { struct blk_plug *plug =3D current->plug; + u64 now; =20 if (!plug || !in_task()) return ktime_get_ns(); @@ -697,12 +698,18 @@ static inline u64 blk_time_get_ns(void) * 0 could very well be a valid time, but rather than flag "this is * a valid timestamp" separately, just accept that we'll do an extra * ktime_get_ns() if we just happen to get 0 as the current time. + * + * cur_ktime can be zeroed by pre-emption the moment PF_BLOCK_TS is set. */ - if (!plug->cur_ktime) { - plug->cur_ktime =3D ktime_get_ns(); + now =3D READ_ONCE(plug->cur_ktime); + if (!now) { + now =3D ktime_get_ns(); + WRITE_ONCE(plug->cur_ktime, now); + /* Ensure PF_BLOCK_TS is set after cur_ktime. */ + barrier(); current->flags |=3D PF_BLOCK_TS; } - return plug->cur_ktime; + return now; } =20 static inline ktime_t blk_time_get(void) --=20 2.47.3