From nobody Tue Sep 29 11:54:21 2026 Received: from mail-pg1-f197.google.com (mail-pg1-f197.google.com [209.85.215.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1C577308F03 for ; Sat, 8 Aug 2026 04:06:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.197 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786161979; cv=none; b=WEWDNj+Uh8Fquy7hlUdw59rTG3tF7qGmKu5tMl7IHfAEC5s5uzLFCKl8aPMXiAjHwYtK/WFWy4ZLU41cWj/IfyHLenrF4pov71TW4cfA2QpA3w/F6CPhagjzXYAir+5maTAGFfIUjVJ7ll5QRduPrEQPBtsco9rEWSPUiXNfjPE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786161979; c=relaxed/simple; bh=AiP5lQkQvr65OeITf1ZDc4QYsn+VgdaKtX4H6Wh88KQ=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=ENrgKdM5reSkUdomoe79wrGwIFIBp+dZNcy2Ya6VPShqmSMi9BmPyYDMKY/RD4C2UzkKWb2bqXk5VKZNYtV5fBhhceYyrBnjU04RrDnHlT7bE0y5y+FrisHP+IoeTHBLS2rbSDZI7tLtVNhSV/6sEwhq689sz3F2ikLo3OeVXoQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--ankitkap.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=EVrV2DPc; arc=none smtp.client-ip=209.85.215.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--ankitkap.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="EVrV2DPc" Received: by mail-pg1-f197.google.com with SMTP id 41be03b00d2f7-cbe9733fbf6so126221a12.2 for ; Fri, 07 Aug 2026 21:06:16 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786161976; x=1786766776; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=mGzS19pS03FkctEDmT6U8dvEyRhKg7nZqwMJ4coRBkk=; b=EVrV2DPcNJ+2cMfTKXx1rZ0bzb8AKcggxJuCek10rM/x1ddMjsjRlbPcPIFY+Abs9t JKb2IkikP+kejVhsxO3b+Mt21C4zjSwF13C7jf856k16PImwgsrAxFlMXRLFDlA5xo9N wMEfwtOMeERRtH8XYfR9gsAWJX38g1qVpfznTCucdm3JOxR/CKYRUtiF2CScKPKOX/qa XIXA1W7VuaJPl542D82ynDz2Ovz7joHsexENImze6Gu9FqEJWJoH9GSadmMzkZ0PPDCo r9QgytIrIS3cGV7uf7g8wLQWT/+FzEeykKrsQJymLt1mclPS5KXk0os2HGS5ZoLEc4ZV 66uw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786161976; x=1786766776; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=mGzS19pS03FkctEDmT6U8dvEyRhKg7nZqwMJ4coRBkk=; b=qN7cx8TB9GcHiWrivpLRth83DSD3QVnwNzAo6MdXpQLceh80n+JxneBSUQDFUsZEoH 7pVgZCapkIYeNmOwWIuuz+Qxeld7SKYpkkep5uIdHckZ0n9vAerCOvv67MhFLIZQW+3X iR83S7J9yNrai5fzzKmXy7WrP4h4GhDnQPrDCP2aQSc1J05TD9qNfkRaRqXZ0ZX5Edxk iNRL96THmCeZqJhUnZr1pQEOq2eqmtrtPLtZGdYJAFUvvHRa/uIu1VNL69i5cj0bgE9e ZsGOsyJ15u0sWQQ4Ih/2/hw+sfVlK03kqas+9iEOZ1W4rO1Ms4O4kFl+ZBhcpA+n5zbF U1gw== X-Forwarded-Encrypted: i=1; AHgh+RqJbyP08C/Alku/qs8gTHraTseJMEhiSAMMTD9rxgT4RS1+T4+G97RYSlJCtlNXB10t+CKd9AZFd6rnyQI=@vger.kernel.org X-Gm-Message-State: AOJu0Yyw0Dk9G9565dxOEqfY1caPGpV9ABRCITyt9VzDc6ohL/+XqE0b t0wtcriG5gfbEiJni9zitdQCN6wgSeGXVNaJPrqQqCOiRfL7gbzAnnRd62ZhGCFmde4gvBg5hnt cyfsEd6sn1X+NlA== X-Received: from pgct6.prod.google.com ([2002:a05:6a02:5286:b0:c85:98b1:c613]) (user=ankitkap job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:b91:b0:3bf:8de9:c647 with SMTP id adf61e73a8af0-3cb85ba8da6mr34695235637.0.1786161976140; Fri, 07 Aug 2026 21:06:16 -0700 (PDT) Date: Sat, 8 Aug 2026 04:05:48 +0000 In-Reply-To: <20260808040549.2778125-1-ankitkap@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260808040549.2778125-1-ankitkap@google.com> X-Mailer: git-send-email 2.55.0.679.g6767b8d81c-goog Message-ID: <20260808040549.2778125-2-ankitkap@google.com> Subject: [PATCH v3 1/2] bcache: track active bypass writes to fix read miss race From: Ankit Kapoor To: Coly Li , linux-bcache@vger.kernel.org Cc: Kent Overstreet , linux-kernel@vger.kernel.org, Ankit Kapoor Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" A race condition exists between a read cache miss and a bypass write due to either congestion or sequential bypass, which causes stale data to be cached when the read cache miss runs concurrently with a bypass write targeting the same sectors. If the read cache miss fetches data from the backing device before the write to the backing device finishes, stale data populates the cache. The root cause is that bcache currently executes btree key invalidation in parallel with (or prior to) writing the actual data payload to the backing device. Under this sequence, a concurrent read path can register a cache miss and insert a placeholder key. If the write's btree key invalidation completes before the read finishes fetching old data from the backing device, the read's subsequent key replacement will not detect a collision, allowing stale data to persist in the cache. Fix this by tracking active bypass writes and serializing cache invalidation. First, divide the backing device space into 32MB chunks and track concurrent bypass writes using refcounts. The tracking counters are stored in dynamically allocated pages backed by a dedicated mempool to prevent allocation failures under memory pressure, minimizing overall footprint (a single 4KB page supports 32GB of disk space using u32 counters). On a cache miss read, bcache checks if there are any active bypass writes overlapping the target sectors. If an active bypass write is detected, the read is forced to bypass the cache to ensure data consistency. Second, serialize the btree key invalidation (bch_data_insert) for bypass writes so that it executes in cached_dev_write_complete(), after the payload has been written to the backing device. This prevents early invalidation from opening a window where a concurrent read miss could fetch and re-cache stale data before the bypass write reaches the disk. Suggested-by: Coly Li Signed-off-by: Ankit Kapoor --- drivers/md/bcache/bcache.h | 35 ++++++++++ drivers/md/bcache/request.c | 131 +++++++++++++++++++++++++++++++++++- drivers/md/bcache/super.c | 39 +++++++++++ 3 files changed, 202 insertions(+), 3 deletions(-) diff --git a/drivers/md/bcache/bcache.h b/drivers/md/bcache/bcache.h index ec9ff9715081..2e50526a52fc 100644 --- a/drivers/md/bcache/bcache.h +++ b/drivers/md/bcache/bcache.h @@ -299,6 +299,12 @@ enum stop_on_failure { BCH_CACHED_DEV_STOP_MODE_MAX, }; =20 +struct bch_bypass_page { + u32 *counts; + unsigned int active; + spinlock_t lock; +}; + struct cached_dev { struct list_head list; struct bcache_device disk; @@ -407,8 +413,37 @@ struct cached_dev { */ #define BCH_WBRATE_UPDATE_MAX_SKIPS 15 unsigned int rate_update_retry; + + /* For tracking active bypass writes */ +#define BCH_BYPASS_CHUNK_SHIFT 16 /* 2^16 sectors =3D 32MB */ +#define BCH_BYPASS_PAGE_COUNTERS (PAGE_SIZE / sizeof(u32)) +#define BCH_BYPASS_PAGE_SHIFT (PAGE_SHIFT - 2) +#define BCH_BYPASS_PAGE_MASK ((1UL << BCH_BYPASS_PAGE_SHIFT) - 1) + struct bch_bypass_page *bypass_pages; + unsigned long bypass_num_pages; + mempool_t bypass_mempool; }; =20 +static inline unsigned long sector_to_bypass_chunk(sector_t sector) +{ + return sector >> BCH_BYPASS_CHUNK_SHIFT; +} + +static inline unsigned long bypass_chunk_to_page(unsigned long chunk) +{ + return chunk >> BCH_BYPASS_PAGE_SHIFT; +} + +static inline unsigned long bypass_chunk_to_offset(unsigned long chunk) +{ + return chunk & BCH_BYPASS_PAGE_MASK; +} + +static inline sector_t bypass_chunk_to_sector(unsigned long chunk) +{ + return (sector_t)chunk << BCH_BYPASS_CHUNK_SHIFT; +} + enum alloc_reserve { RESERVE_BTREE, RESERVE_PRIO, diff --git a/drivers/md/bcache/request.c b/drivers/md/bcache/request.c index 3fa3b13a410f..bfd28b68f499 100644 --- a/drivers/md/bcache/request.c +++ b/drivers/md/bcache/request.c @@ -492,6 +492,9 @@ struct search { struct block_device *orig_bdev; unsigned long start_time; =20 + sector_t bypass_sector; + unsigned int bypass_sectors; + struct btree_op op; struct data_insert_op iop; }; @@ -763,11 +766,16 @@ static inline struct search *search_alloc(struct bio = *bio, =20 /* Cached devices */ =20 +static void bch_bypass_write_end(struct cached_dev *dc, sector_t sector, u= nsigned int sectors); + static CLOSURE_CALLBACK(cached_dev_bio_complete) { closure_type(s, struct search, cl); struct cached_dev *dc =3D container_of(s->d, struct cached_dev, disk); =20 + if (s->iop.bypass && op_is_write(bio_op(s->orig_bio))) + bch_bypass_write_end(dc, s->bypass_sector, s->bypass_sectors); + cached_dev_put(dc); search_free(&cl->work); } @@ -830,6 +838,110 @@ static CLOSURE_CALLBACK(cached_dev_cache_miss_done) closure_put(&d->cl); } =20 +static void bch_bypass_write_start(struct cached_dev *dc, sector_t sector,= unsigned int sectors) +{ + unsigned long start_chunk =3D sector_to_bypass_chunk(sector); + unsigned long end_chunk =3D sector_to_bypass_chunk(sector + sectors - 1); + unsigned long end_pg_idx =3D bypass_chunk_to_page(end_chunk); + unsigned long chunk; + + if (WARN_ON_ONCE(end_pg_idx >=3D dc->bypass_num_pages)) + return; + + for (chunk =3D start_chunk; chunk <=3D end_chunk; chunk++) { + unsigned long pg_idx =3D bypass_chunk_to_page(chunk); + unsigned long pg_off =3D bypass_chunk_to_offset(chunk); + struct bch_bypass_page *pg =3D &dc->bypass_pages[pg_idx]; + u32 *new_counts; + u32 *dup_counts =3D NULL; + unsigned long flags; + + spin_lock_irqsave(&pg->lock, flags); + if (!pg->counts) { + spin_unlock_irqrestore(&pg->lock, flags); + new_counts =3D mempool_alloc(&dc->bypass_mempool, GFP_NOIO); + memset(new_counts, 0, PAGE_SIZE); + spin_lock_irqsave(&pg->lock, flags); + if (pg->counts) + dup_counts =3D new_counts; + else + pg->counts =3D new_counts; + } + pg->counts[pg_off]++; + pg->active++; + spin_unlock_irqrestore(&pg->lock, flags); + + if (dup_counts) + mempool_free(dup_counts, &dc->bypass_mempool); + } +} + +static void bch_bypass_write_end(struct cached_dev *dc, sector_t sector, u= nsigned int sectors) +{ + unsigned long start_chunk =3D sector_to_bypass_chunk(sector); + unsigned long end_chunk =3D sector_to_bypass_chunk(sector + sectors - 1); + unsigned long end_pg_idx =3D bypass_chunk_to_page(end_chunk); + unsigned long chunk; + + if (WARN_ON_ONCE(end_pg_idx >=3D dc->bypass_num_pages)) + return; + + for (chunk =3D start_chunk; chunk <=3D end_chunk; chunk++) { + unsigned long pg_idx =3D bypass_chunk_to_page(chunk); + unsigned long pg_off =3D bypass_chunk_to_offset(chunk); + struct bch_bypass_page *pg =3D &dc->bypass_pages[pg_idx]; + u32 *counts =3D NULL; + unsigned long flags; + + spin_lock_irqsave(&pg->lock, flags); + if (WARN_ON_ONCE(!pg->counts || !pg->counts[pg_off])) { + spin_unlock_irqrestore(&pg->lock, flags); + continue; + } + + pg->counts[pg_off]--; + pg->active--; + if (!pg->active) { + counts =3D pg->counts; + pg->counts =3D NULL; + } + spin_unlock_irqrestore(&pg->lock, flags); + + if (counts) + mempool_free(counts, &dc->bypass_mempool); + } +} + +static bool bch_has_active_bypass_writes(struct cached_dev *dc, sector_t s= ector, + unsigned int sectors) +{ + unsigned long start_chunk =3D sector_to_bypass_chunk(sector); + unsigned long end_chunk =3D sector_to_bypass_chunk(sector + sectors - 1); + unsigned long end_pg_idx =3D bypass_chunk_to_page(end_chunk); + unsigned long chunk; + bool has_active =3D false; + + if (WARN_ON_ONCE(end_pg_idx >=3D dc->bypass_num_pages)) + return false; + + for (chunk =3D start_chunk; chunk <=3D end_chunk; chunk++) { + unsigned long pg_idx =3D bypass_chunk_to_page(chunk); + unsigned long pg_off =3D bypass_chunk_to_offset(chunk); + struct bch_bypass_page *pg =3D &dc->bypass_pages[pg_idx]; + unsigned long flags; + + spin_lock_irqsave(&pg->lock, flags); + if (pg->counts && pg->counts[pg_off] > 0) { + has_active =3D true; + spin_unlock_irqrestore(&pg->lock, flags); + break; + } + spin_unlock_irqrestore(&pg->lock, flags); + } + + return has_active; +} + static CLOSURE_CALLBACK(cached_dev_read_done) { closure_type(s, struct search, cl); @@ -864,7 +976,9 @@ static CLOSURE_CALLBACK(cached_dev_read_done) bio_complete(s); =20 if (s->iop.bio && - !test_bit(CACHE_SET_STOPPING, &s->iop.c->flags)) { + !test_bit(CACHE_SET_STOPPING, &s->iop.c->flags) && + !bch_has_active_bypass_writes(dc, s->iop.bio->bi_iter.bi_sector, + bio_sectors(s->iop.bio))) { BUG_ON(!s->iop.replace); closure_call(&s->iop.cl, bch_data_insert, NULL, cl); } @@ -975,7 +1089,13 @@ static CLOSURE_CALLBACK(cached_dev_write_complete) struct cached_dev *dc =3D container_of(s->d, struct cached_dev, disk); =20 up_read_non_owner(&dc->writeback_lock); - cached_dev_bio_complete(&cl->work); + + if (s->iop.bypass) { + closure_call(&s->iop.cl, bch_data_insert, NULL, cl); + continue_at(cl, cached_dev_bio_complete, NULL); + } else { + cached_dev_bio_complete(&cl->work); + } } =20 static void cached_dev_write(struct cached_dev *dc, struct search *s) @@ -1018,6 +1138,10 @@ static void cached_dev_write(struct cached_dev *dc, = struct search *s) s->iop.bio =3D s->orig_bio; bio_get(s->iop.bio); =20 + s->bypass_sector =3D bio->bi_iter.bi_sector; + s->bypass_sectors =3D bio_sectors(bio); + bch_bypass_write_start(dc, s->bypass_sector, s->bypass_sectors); + if (bio_op(bio) =3D=3D REQ_OP_DISCARD && !bdev_max_discard_sectors(dc->bdev)) goto insert_data; @@ -1058,7 +1182,8 @@ static void cached_dev_write(struct cached_dev *dc, s= truct search *s) } =20 insert_data: - closure_call(&s->iop.cl, bch_data_insert, NULL, cl); + if (!s->iop.bypass) + closure_call(&s->iop.cl, bch_data_insert, NULL, cl); continue_at(cl, cached_dev_write_complete, NULL); } =20 diff --git a/drivers/md/bcache/super.c b/drivers/md/bcache/super.c index 97d9adb0bf96..b994483511de 100644 --- a/drivers/md/bcache/super.c +++ b/drivers/md/bcache/super.c @@ -1346,6 +1346,16 @@ void bch_cached_dev_release(struct kobject *kobj) { struct cached_dev *dc =3D container_of(kobj, struct cached_dev, disk.kobj); + if (dc->bypass_pages) { + unsigned long i; + + for (i =3D 0; i < dc->bypass_num_pages; i++) { + if (dc->bypass_pages[i].counts) + mempool_free(dc->bypass_pages[i].counts, &dc->bypass_mempool); + } + kvfree(dc->bypass_pages); + mempool_exit(&dc->bypass_mempool); + } kfree(dc); module_put(THIS_MODULE); } @@ -1407,6 +1417,31 @@ static CLOSURE_CALLBACK(cached_dev_flush) continue_at(cl, cached_dev_free, system_percpu_wq); } =20 +static int bch_cached_dev_bypass_init(struct cached_dev *dc, sector_t sect= ors) +{ + unsigned long chunks =3D + sector_to_bypass_chunk(sectors + (1UL << BCH_BYPASS_CHUNK_SHIFT) - 1); + unsigned long i; + + dc->bypass_num_pages =3D DIV_ROUND_UP(chunks, BCH_BYPASS_PAGE_COUNTERS); + dc->bypass_pages =3D kvcalloc(dc->bypass_num_pages, + sizeof(struct bch_bypass_page), + GFP_KERNEL); + if (!dc->bypass_pages) + return -ENOMEM; + + for (i =3D 0; i < dc->bypass_num_pages; i++) + spin_lock_init(&dc->bypass_pages[i].lock); + + if (mempool_init_kmalloc_pool(&dc->bypass_mempool, 16, PAGE_SIZE)) { + kvfree(dc->bypass_pages); + dc->bypass_pages =3D NULL; + return -ENOMEM; + } + + return 0; +} + static int cached_dev_init(struct cached_dev *dc, unsigned int block_size) { int ret; @@ -1447,6 +1482,10 @@ static int cached_dev_init(struct cached_dev *dc, un= signed int block_size) /* default to auto */ dc->stop_when_cache_set_failed =3D BCH_CACHED_DEV_STOP_AUTO; =20 + ret =3D bch_cached_dev_bypass_init(dc, bdev_nr_sectors(dc->bdev)); + if (ret) + return ret; + bch_cached_dev_request_init(dc); bch_cached_dev_writeback_init(dc); return 0; --=20 2.55.0.679.g6767b8d81c-goog From nobody Tue Sep 29 11:54:21 2026 Received: from mail-pl1-f199.google.com (mail-pl1-f199.google.com [209.85.214.199]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5A65B385D86 for ; Sat, 8 Aug 2026 04:06:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.199 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786161981; cv=none; b=kskX4koB8Xc03XMOcLP7KRXiU3uo/DY8W5UuDijuNz2Dtckrzm7eFs2XyRva+qgGbhh2/cFMsW5PWMhcvKVHuLPpWOWsKcn+4SO2O/P93Fvml7S0tAnd9zjvAXYuMvo7lfDh7RjQjvo32vx2+HKTyCPLY9b9SKNfq/EMf/25ogg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786161981; c=relaxed/simple; bh=xA+Anu1G01kalxTG8fE/ijwz0xdzhG/T7miN8n1j8/Q=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=dChh8sESDzksYg5D4G62NoRObJgBZkE/ZwBik/Wq778Bwat/C1S0R9YZz1v8zw7nwvKZQZ/G9IN5uuVGxVnX+qhEbh2fbcJvwsM0RjTSbNsDM/SKsS9/THFVo/OemNVrPpo5Uz8URDRdKQRbAVkExvvpen5btr6i2ouXWxRjTew= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--ankitkap.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=fY5zmUgL; arc=none smtp.client-ip=209.85.214.199 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--ankitkap.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="fY5zmUgL" Received: by mail-pl1-f199.google.com with SMTP id d9443c01a7336-2cea6a46766so5092405ad.0 for ; Fri, 07 Aug 2026 21:06:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786161980; x=1786766780; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=cp2CV40UiAeLycXOv282ow6mV6SoZ8nWxaGYowz7qXM=; b=fY5zmUgLOWZo3V6GvY3eBSCrhds82RzCJQDYAds0393xSRU/cA9sv6zI5sF+Cg1/7C HK3BAl1d79RLha36volF9XRK/pIgCnZcA87/x+2sO/f2ZpEls8jUwKy3KmeON6r1XNVj ZQ7XtNzBaLCi8waIGTmZ5J978Yh3r1ASMJC7P3D45GfpGRBFHpf4/aoLsw5UE0YpVMZf +z4wXsv+ILTvf78t4FZrNVy73mpSvZNPQeKXU1B0fHQBk3B3N+getQZ0laI2dcEDN97t Tjuu7t1BQaxO8C3+LhTO9ZBlctihLYupWrQeXLWhd4uruh4nHMqgxR14HysVe03cD6HK ExVg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786161980; x=1786766780; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=cp2CV40UiAeLycXOv282ow6mV6SoZ8nWxaGYowz7qXM=; b=P65QIoSVXdWFD/ic7xVUi8JswQV+9kiqXIb/g+FceUXW5ZFrQUDHAhVxbbWhVQN6/K gPqtbyHxDJDcW3U6fQbZwJFupQL2k/ruMGwkH0ysruLKCrQ9G25kysakbafN0c7kCSJD x057gwL3gWMOIFLQb3cAYEilA9bV/xtVeHRTBYq2jRTgJ0oIwwbzbywsnDoJO84TRVFq /3XLjbXuzBQqIx0iDpK6UT/a+coURCyNYksC3PPmM+xURW3S9eGfQxn0lqjmY23C5axO JiCjUCKFdwyddcdZxcR+NZpNM63/d6M7H2yUU+EMkC4C2LTmVT6111kZUfEeK/myyb+N RVpQ== X-Forwarded-Encrypted: i=1; AHgh+RpDZXml5CzKdV0SomcBlYiLBXUX7jUwp7z4T40jlX2mXxbViJB9eJorQlunmF6ve6R6T/B+1soLXe3rQW4=@vger.kernel.org X-Gm-Message-State: AOJu0Yz9Cwb/GtUu4oI8lBmerpXAAu9+eKcCg92vbcs4Pt5Vh7oEdfHU KXVUm+9XmVkAKsXTN9GTAUUPgKS5iNRjMX9KIkNF8OnT+MIMdxoh+UGGy+DYl6ydnRlnuRI7ScR nHw0FHcFZNVcAgw== X-Received: from plas19.prod.google.com ([2002:a17:903:2013:b0:2d0:34d:e973]) (user=ankitkap job=prod-delivery.src-stubby-dispatcher) by 2002:a17:902:e84b:b0:2c9:b396:1a55 with SMTP id d9443c01a7336-2d0ca76a3bamr304649095ad.12.1786161979362; Fri, 07 Aug 2026 21:06:19 -0700 (PDT) Date: Sat, 8 Aug 2026 04:05:49 +0000 In-Reply-To: <20260808040549.2778125-1-ankitkap@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260808040549.2778125-1-ankitkap@google.com> X-Mailer: git-send-email 2.55.0.679.g6767b8d81c-goog Message-ID: <20260808040549.2778125-3-ankitkap@google.com> Subject: [PATCH v3 2/2] bcache: inspect active bypass writes lock-free via RCU From: Ankit Kapoor To: Coly Li , linux-bcache@vger.kernel.org Cc: Kent Overstreet , linux-kernel@vger.kernel.org, Ankit Kapoor Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Checking for active bypass writes on the cache miss read path currently requires acquiring page-level spinlocks across the target sector range. While contention is low under normal operation, acquiring a lock in the latency-sensitive read path introduces unnecessary overhead. Optimize the read path by using RCU to inspect active bypass write counters lock-free. Writers continue to use page-level spinlocks to synchronize counter updates and allocations, while readers inspect the tracking array under rcu_read_lock(). Suggested-by: Coly Li Signed-off-by: Ankit Kapoor --- drivers/md/bcache/bcache.h | 14 +++++++-- drivers/md/bcache/request.c | 58 +++++++++++++++++++++++++------------ drivers/md/bcache/super.c | 9 ++++-- 3 files changed, 56 insertions(+), 25 deletions(-) diff --git a/drivers/md/bcache/bcache.h b/drivers/md/bcache/bcache.h index 2e50526a52fc..2ecff48d8902 100644 --- a/drivers/md/bcache/bcache.h +++ b/drivers/md/bcache/bcache.h @@ -299,10 +299,18 @@ enum stop_on_failure { BCH_CACHED_DEV_STOP_MODE_MAX, }; =20 +extern struct kmem_cache *bch_bypass_cache; + +struct bch_bypass_counts { + struct rcu_head rcu; + struct cached_dev *dc; + u32 counts[PAGE_SIZE / sizeof(u32)]; +}; + struct bch_bypass_page { - u32 *counts; - unsigned int active; - spinlock_t lock; + struct bch_bypass_counts __rcu *counts; + unsigned int active; + spinlock_t lock; }; =20 struct cached_dev { diff --git a/drivers/md/bcache/request.c b/drivers/md/bcache/request.c index bfd28b68f499..9221db669050 100644 --- a/drivers/md/bcache/request.c +++ b/drivers/md/bcache/request.c @@ -24,6 +24,7 @@ #define CUTOFF_CACHE_READA 90 =20 struct kmem_cache *bch_search_cache; +struct kmem_cache *bch_bypass_cache; =20 static CLOSURE_CALLBACK(bch_data_insert_start); =20 @@ -852,22 +853,28 @@ static void bch_bypass_write_start(struct cached_dev = *dc, sector_t sector, unsig unsigned long pg_idx =3D bypass_chunk_to_page(chunk); unsigned long pg_off =3D bypass_chunk_to_offset(chunk); struct bch_bypass_page *pg =3D &dc->bypass_pages[pg_idx]; - u32 *new_counts; - u32 *dup_counts =3D NULL; + struct bch_bypass_counts *new_counts; + struct bch_bypass_counts *dup_counts =3D NULL; + struct bch_bypass_counts *counts; unsigned long flags; =20 spin_lock_irqsave(&pg->lock, flags); - if (!pg->counts) { + counts =3D rcu_dereference_protected(pg->counts, lockdep_is_held(&pg->lo= ck)); + if (!counts) { spin_unlock_irqrestore(&pg->lock, flags); new_counts =3D mempool_alloc(&dc->bypass_mempool, GFP_NOIO); - memset(new_counts, 0, PAGE_SIZE); + memset(new_counts->counts, 0, PAGE_SIZE); + new_counts->dc =3D dc; spin_lock_irqsave(&pg->lock, flags); - if (pg->counts) + counts =3D rcu_dereference_protected(pg->counts, lockdep_is_held(&pg->l= ock)); + if (counts) { dup_counts =3D new_counts; - else - pg->counts =3D new_counts; + } else { + counts =3D new_counts; + rcu_assign_pointer(pg->counts, counts); + } } - pg->counts[pg_off]++; + WRITE_ONCE(counts->counts[pg_off], counts->counts[pg_off] + 1); pg->active++; spin_unlock_irqrestore(&pg->lock, flags); =20 @@ -876,6 +883,13 @@ static void bch_bypass_write_start(struct cached_dev *= dc, sector_t sector, unsig } } =20 +static void bch_bypass_counts_free_rcu(struct rcu_head *rcu) +{ + struct bch_bypass_counts *counts =3D container_of(rcu, struct bch_bypass_= counts, rcu); + + mempool_free(counts, &counts->dc->bypass_mempool); +} + static void bch_bypass_write_end(struct cached_dev *dc, sector_t sector, u= nsigned int sectors) { unsigned long start_chunk =3D sector_to_bypass_chunk(sector); @@ -890,25 +904,27 @@ static void bch_bypass_write_end(struct cached_dev *d= c, sector_t sector, unsigne unsigned long pg_idx =3D bypass_chunk_to_page(chunk); unsigned long pg_off =3D bypass_chunk_to_offset(chunk); struct bch_bypass_page *pg =3D &dc->bypass_pages[pg_idx]; - u32 *counts =3D NULL; + struct bch_bypass_counts *counts =3D NULL; + struct bch_bypass_counts *current_counts; unsigned long flags; =20 spin_lock_irqsave(&pg->lock, flags); - if (WARN_ON_ONCE(!pg->counts || !pg->counts[pg_off])) { + current_counts =3D rcu_dereference_protected(pg->counts, lockdep_is_held= (&pg->lock)); + if (WARN_ON_ONCE(!current_counts || !current_counts->counts[pg_off])) { spin_unlock_irqrestore(&pg->lock, flags); continue; } =20 - pg->counts[pg_off]--; + WRITE_ONCE(current_counts->counts[pg_off], current_counts->counts[pg_off= ] - 1); pg->active--; if (!pg->active) { - counts =3D pg->counts; - pg->counts =3D NULL; + counts =3D current_counts; + rcu_assign_pointer(pg->counts, NULL); } spin_unlock_irqrestore(&pg->lock, flags); =20 if (counts) - mempool_free(counts, &dc->bypass_mempool); + call_rcu(&counts->rcu, bch_bypass_counts_free_rcu); } } =20 @@ -924,20 +940,19 @@ static bool bch_has_active_bypass_writes(struct cache= d_dev *dc, sector_t sector, if (WARN_ON_ONCE(end_pg_idx >=3D dc->bypass_num_pages)) return false; =20 + rcu_read_lock(); for (chunk =3D start_chunk; chunk <=3D end_chunk; chunk++) { unsigned long pg_idx =3D bypass_chunk_to_page(chunk); unsigned long pg_off =3D bypass_chunk_to_offset(chunk); struct bch_bypass_page *pg =3D &dc->bypass_pages[pg_idx]; - unsigned long flags; + struct bch_bypass_counts *current_counts =3D rcu_dereference(pg->counts); =20 - spin_lock_irqsave(&pg->lock, flags); - if (pg->counts && pg->counts[pg_off] > 0) { + if (current_counts && READ_ONCE(current_counts->counts[pg_off]) > 0) { has_active =3D true; - spin_unlock_irqrestore(&pg->lock, flags); break; } - spin_unlock_irqrestore(&pg->lock, flags); } + rcu_read_unlock(); =20 return has_active; } @@ -1458,6 +1473,7 @@ void bch_flash_dev_request_init(struct bcache_device = *d) =20 void bch_request_exit(void) { + kmem_cache_destroy(bch_bypass_cache); kmem_cache_destroy(bch_search_cache); } =20 @@ -1467,5 +1483,9 @@ int __init bch_request_init(void) if (!bch_search_cache) return -ENOMEM; =20 + bch_bypass_cache =3D KMEM_CACHE(bch_bypass_counts, 0); + if (!bch_bypass_cache) + return -ENOMEM; + return 0; } diff --git a/drivers/md/bcache/super.c b/drivers/md/bcache/super.c index b994483511de..75aa6bb2e00b 100644 --- a/drivers/md/bcache/super.c +++ b/drivers/md/bcache/super.c @@ -1350,9 +1350,12 @@ void bch_cached_dev_release(struct kobject *kobj) unsigned long i; =20 for (i =3D 0; i < dc->bypass_num_pages; i++) { - if (dc->bypass_pages[i].counts) - mempool_free(dc->bypass_pages[i].counts, &dc->bypass_mempool); + struct bch_bypass_counts *counts =3D + rcu_dereference_protected(dc->bypass_pages[i].counts, 1); + if (counts) + mempool_free(counts, &dc->bypass_mempool); } + rcu_barrier(); kvfree(dc->bypass_pages); mempool_exit(&dc->bypass_mempool); } @@ -1433,7 +1436,7 @@ static int bch_cached_dev_bypass_init(struct cached_d= ev *dc, sector_t sectors) for (i =3D 0; i < dc->bypass_num_pages; i++) spin_lock_init(&dc->bypass_pages[i].lock); =20 - if (mempool_init_kmalloc_pool(&dc->bypass_mempool, 16, PAGE_SIZE)) { + if (mempool_init_slab_pool(&dc->bypass_mempool, 16, bch_bypass_cache)) { kvfree(dc->bypass_pages); dc->bypass_pages =3D NULL; return -ENOMEM; --=20 2.55.0.679.g6767b8d81c-goog