From nobody Mon Sep 28 20:05:16 2026 Received: from va-2-40.ptr.blmpb.com (va-2-40.ptr.blmpb.com [209.127.231.40]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0AA83378D6B for ; Tue, 18 Aug 2026 07:07:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.127.231.40 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036837; cv=none; b=Y0xoXVf284YRxOdN6yaWc6PbztijrVjli73bra/dGHMLbDgy0RyVcHrmIvY5k8geK3FKh+HObaboaE9NA47H3OSp8wM4g9DUNCys8727G4RCwXB5lpob7ijj2OjRts0bu3eYw7NkHKcxafI2DSlxpdHOGBW/mRLAdtH51vLXAFw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036837; c=relaxed/simple; bh=0bg/lByf+5Kj7BlZSA8bQ8x+MhSXaW87xNlm8QSGLR0=; h=To:From:Message-Id:References:Cc:Subject:Date:Mime-Version: In-Reply-To:Content-Type; b=DTDSl3uZ33uHmDPbf/rWLHFXc4V0FquAFIDMRI7cBKzVqp5AcOC9eo1zTpxEIuvsgVwruQOt/pee/oLaQlRqWLBaBQcRkDe1XbXUkPxuIN9LaKAoHhhIVN/dRHsr940ap74/UI248HvuBM6dkPW9Yg9bgy+aQEzQFnWnNHmuY0E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com; spf=pass smtp.mailfrom=fnnas.com; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b=XzJ4MGg9; arc=none smtp.client-ip=209.127.231.40 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=fnnas.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b="XzJ4MGg9" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; s=s1; d=fnnas-com.20200927.dkim.feishu.cn; t=1787036825; h=from:subject:mime-version:from:date:message-id:subject:to:cc: reply-to:content-type:mime-version:in-reply-to:message-id; bh=mEIJhFo42Is1jlZz91hZNN1RltsJcbPXgmbj+a6kffU=; b=XzJ4MGg9JZglFDng5okoKXrQXG6HXoPhBQ6mvG75hljcUjPe7LGh4a1tbLbUSn3laVkpYK 27OpMnXpQYLuKgn4p3uKa+xk8lksqqk/k6xEzGoV3skNE92qieXt9vhbCGL7bTKXZnto8+ WTZ+IFSz8R2XfTt5pCgFCUFzxNJHPqr8osf00QiXSUXIEsErHLTLrg90YL7roreh9DGaMl HiHTtIX6SMWuEvLweuC/vGzSMCvrME4k6svOYmmDkjBNUmDChRUp1tcWXmf37rbt0fdymD vxmdb6xECEuSKkmbn9KMTmeIg/hfZlLx7tkdcIzrCJBzNeg9IYdOt1VInAdZkg== To: , , From: "Chen Cheng" Message-Id: <20260818070646.1029149-2-chencheng@fnnas.com> References: <20260818070646.1029149-1-chencheng@fnnas.com> X-Lms-Return-Path: Received: from fedora ([183.34.162.43]) by smtp.feishu.cn with ESMTPS; Tue, 18 Aug 2026 15:06:59 +0800 Cc: , Content-Transfer-Encoding: quoted-printable X-Mailer: git-send-email 2.55.0 X-Original-From: chencheng@fnnas.com Subject: [RFC PATCH 1/5] md/raid1: balance reads across non-rotational disks Date: Tue, 18 Aug 2026 15:06:42 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 In-Reply-To: <20260818070646.1029149-1-chencheng@fnnas.com> Content-Type: text/plain; charset="utf-8" From: Chen Cheng choose_best_rdev() picks one disk for each read. min_pending is UINT_MAX; store it as unsigned int, like raid10. Current: 1. Sequential read: stay on the current disk. 2. Switch only if should_choose_next() is true. 3. should_choose_next() needs bdev_io_opt() > 0. 4. Non-sequential read: pick the disk with the lowest nr_pending. 5. The compare uses strict '>'. Same pending keeps the first disk. Problem: 1. Many client NVMe set optimal_io_size to 0. Then should_choose_next() never runs. One sequential stream stays on one disk. Why not use the idle disk? 2. Same for rot-only RAID1. One rot disk takes the whole stream. The other rot disk is idle. Why not use it? 3. Low-depth random reads often have the same pending. Why always stay on slot 0? Improve: 1. If a sequential disk already has pending I/O, do not return it at once. Let pending pick an idle disk. 2. On nonrot arrays, if pending is the same, rotate a start slot (0..raid_disks-1). 3. Only bump read_rr when the array has a nonrot member. Tested with fio libaio direct=3D1 (NVMe scheduler none, SATA scheduler mq-deadline): - 2x Predator GM9000 (optimal_io_size=3D0): a) 4k randread QD1 jobs=3D1: 0.084 GB/s, 100/0 -> 0.084 GB/s, 50/50 b) 1M read QD16 jobs=3D1: 7.031 -> 13.886 GB/s (+97%), 66/34 -> 50/50 c) 1M read QD16 jobs=3D2: 14.22 GB/s, 50/50 both sides - 4x Intel MEMPEK1J016GA (Optane pmem, optimal_io_size=3D0): 1M read QD16 jobs=3D1: ~0.85 GB/s on one member -> ~3.0+ GB/s, ~25% per disk - RAID1 of two then three TOSHIBA HDWG740: 2 disks, 1M read QD16 jobs=3D1: 0.294 GB/s, 100/0 -> 0.514 GB/s, 50/50 3 disks, 1M read QD16 jobs=3D1: 0.294 GB/s, 100/0 -> 0.630 GB/s, 33/33/33 2 and 3 disks, 4k randread QD1 jobs=3D1: 0.001 GB/s, all on one disk (rot-only random does not use read_rr) Signed-off-by: Chen Cheng --- drivers/md/raid1.c | 25 +++++++++++++++++++++---- drivers/md/raid1.h | 1 + 2 files changed, 22 insertions(+), 4 deletions(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index f0646fb24371..319b24bcab5b 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -780,27 +780,38 @@ static bool rdev_readable(struct md_rdev *rdev, struc= t r1bio *r1_bio) } =20 struct read_balance_ctl { sector_t closest_dist; int closest_dist_disk; - int min_pending; + unsigned int min_pending; int min_pending_disk; int sequential_disk; int readable_disks; }; =20 +static int raid1_rr_pos(int disk, int start, int n) +{ + return ((disk % n) - start + n) % n; +} + static int choose_best_rdev(struct r1conf *conf, struct r1bio *r1_bio) { int disk; + int rr_start =3D 0; + bool has_nonrot =3D READ_ONCE(conf->nonrot_disks); struct read_balance_ctl ctl =3D { .closest_dist_disk =3D -1, .closest_dist =3D MaxSector, .min_pending_disk =3D -1, .min_pending =3D UINT_MAX, .sequential_disk =3D -1, }; =20 + if (has_nonrot) + rr_start =3D (unsigned int)atomic_inc_return(&conf->read_rr) % + conf->raid_disks; + for (disk =3D 0 ; disk < conf->raid_disks * 2 ; disk++) { struct md_rdev *rdev; sector_t dist; unsigned int pending; =20 @@ -819,11 +830,11 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) dist =3D abs(r1_bio->sector - READ_ONCE(conf->mirrors[disk].head_position)); =20 /* Don't change to another disk for sequential reads */ if (is_sequential(conf, disk, r1_bio)) { - if (!should_choose_next(conf, disk)) + if (!should_choose_next(conf, disk) && !pending) return disk; =20 /* * Add 'pending' to avoid choosing this disk if * there is other idle disk. @@ -834,11 +845,16 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) * will be chosen. */ ctl.sequential_disk =3D disk; } =20 - if (ctl.min_pending > pending) { + if (ctl.min_pending > pending || + (has_nonrot && ctl.min_pending =3D=3D pending && + ctl.min_pending_disk >=3D 0 && + raid1_rr_pos(disk, rr_start, conf->raid_disks) < + raid1_rr_pos(ctl.min_pending_disk, rr_start, + conf->raid_disks))) { ctl.min_pending =3D pending; ctl.min_pending_disk =3D disk; } =20 if (ctl.closest_dist > dist) { @@ -859,11 +875,11 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) * non-rotational, choose the disk with less pending request even the * disk is rotational, which might/might not be optimal for raids with * mixed ratation/non-rotational disks depending on workload. */ if (ctl.min_pending_disk !=3D -1 && - (READ_ONCE(conf->nonrot_disks) || ctl.min_pending =3D=3D 0)) + (has_nonrot || ctl.min_pending =3D=3D 0)) return ctl.min_pending_disk; else return ctl.closest_dist_disk; } =20 @@ -3091,10 +3107,11 @@ static struct r1conf *setup_conf(struct mddev *mdde= v) goto abort; =20 err =3D -EINVAL; spin_lock_init(&conf->device_lock); conf->raid_disks =3D mddev->raid_disks; + atomic_set(&conf->read_rr, -1); rdev_for_each(rdev, mddev) { int disk_idx =3D rdev->raid_disk; =20 if (disk_idx >=3D conf->raid_disks || disk_idx < 0) continue; diff --git a/drivers/md/raid1.h b/drivers/md/raid1.h index c98d43a7ae99..d5de976d171d 100644 --- a/drivers/md/raid1.h +++ b/drivers/md/raid1.h @@ -54,10 +54,11 @@ struct r1conf { struct raid1_info *mirrors; /* twice 'raid_disks' to * allow for replacements. */ int raid_disks; int nonrot_disks; + atomic_t read_rr; =20 spinlock_t device_lock; =20 /* list of 'struct r1bio' that need to be processed by raid1d, * whether to retry a read, writeout a resync or recovery --=20 2.55.0 From nobody Mon Sep 28 20:05:16 2026 Received: from va-2-27.ptr.blmpb.com (va-2-27.ptr.blmpb.com [209.127.231.27]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 40F5026AE5 for ; Tue, 18 Aug 2026 07:07:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.127.231.27 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036839; cv=none; b=hQGUi66xyKwesvtebJ3Dg5SrAbHXFxynrVShmLaXi18HnLl5AohOYYxPiFvcsIzUgHneI+uk20InDCjoGPmX17vsRzVe1kK6L5IeSbNtcI7UxUYNzXDL8ECKtcujuB8suVZyPjNQ/8cqCaqfQiziepH4mFm53TklXSExc/8YreA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036839; c=relaxed/simple; bh=u4kQwNTlDL8++Cn5rG+WarjmGr/E4g2luoFJylHo45w=; h=Content-Type:Date:References:Cc:Subject:Mime-Version:To:From: Message-Id:In-Reply-To; b=a3X0vBF5RBvEd/v2wlcDppEjaJmCbtBP9zbm10AonUXbMnTvSsqCX/JSGcoCbZ3Zn3fswiC96ciaDwhIHg4lES4lZkIdnKrphemtKJHjwJna61QSJ8N1u2P+FDILmoc6JkxorNNgBXG78St74SorDSY4uJPf8xw8jzrYW7NmxA0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com; spf=pass smtp.mailfrom=fnnas.com; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b=dCl+m7PP; arc=none smtp.client-ip=209.127.231.27 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=fnnas.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b="dCl+m7PP" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; s=s1; d=fnnas-com.20200927.dkim.feishu.cn; t=1787036829; h=from:subject:mime-version:from:date:message-id:subject:to:cc: reply-to:content-type:mime-version:in-reply-to:message-id; bh=fUofpriDKKunrRStc3hcyDroS4uuxdeVc337dIz70K8=; b=dCl+m7PPYB6rO/WrR3roMMOHvNH14AzfhKChG5vRq6f5q6LRt7+EdZM6VjCke0mv2itUzt A+oqtnpLTFaKFiWfXXwvzR4N4CLysFu8Psf/spElKunSsnhk3rKF6B2eZMAmhzHHbu4bSj U+MOYbcSPpKwbf03uWeY43/2N76l8zgtfN8n4qh8vO9qtZSkJKDzYikP/JFLxmFQouox/g JHcoECDB/hW1/6jtdikwMpmZKh1JDwELsOMI+915xRQboPfb0cqk31tf871SF7c2m0Ds5e k1r+3UzH4rQkECtzw4hcv23tNvoLpc8TJFWgT4VtI6wDcaidvJdSakcE0cC/Iw== Date: Tue, 18 Aug 2026 15:06:43 +0800 X-Mailer: git-send-email 2.55.0 X-Original-From: chencheng@fnnas.com Received: from fedora ([183.34.162.43]) by smtp.feishu.cn with ESMTPS; Tue, 18 Aug 2026 15:07:02 +0800 References: <20260818070646.1029149-1-chencheng@fnnas.com> Cc: , Subject: [RFC PATCH 2/5] md/raid1: do not move nonrot reads onto a rot disk Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 To: , , From: "Chen Cheng" X-Lms-Return-Path: Message-Id: <20260818070646.1029149-3-chencheng@fnnas.com> Content-Transfer-Encoding: quoted-printable In-Reply-To: <20260818070646.1029149-1-chencheng@fnnas.com> Content-Type: text/plain; charset="utf-8" From: Chen Cheng The previous change: if a sequential disk already has pending I/O, the next read can go to an idle disk. Current: 1. A sequential disk can give the next read to an idle peer. 2. On a mixed array that peer can be a rot disk. 3. After a write, every disk has the same head_position. 4. Then every disk looks sequential. Problem: 1. The first sequential disk in slot order may be a rot disk. Then we return it and never see the nonrot disk. We want the nonrot disk. 2. When pending is the same, rot and nonrot share one round-robin. A rot disk can win. We want the nonrot disk. Improve: 1. If a nonrot disk is readable, do not stop on a sequential rot disk. 2. Remember sequential_disk once. A nonrot disk may replace a rot one. 3. If a nonrot disk is readable, do not keep a rot disk as the sequential fallback. 4. When pending is the same, prefer nonrot. Rotate only among nonrot disks. 5. Rot-only sequential reads still stay on the current disk when no peer is idle. Tested with fio libaio direct=3D1 (NVMe scheduler none, SATA scheduler mq-deadline), after a short write so both head_positions match: - RAID1 of Predator GM9000 + SATA HDD: 1M read QD16 jobs=3D1, nonrot first: 2.468 GB/s, 92/8 NVMe/HDD -> 7.112 GB/s, 100/0 NVMe 1M read QD16 jobs=3D1, rot first: 0.280 GB/s, 100/0 HDD -> 7.112 GB/s, 100/0 NVMe - RAID1 of Fanxiang S103Pro + SATA HDD: 1M read QD16 jobs=3D1, nonrot first: 0.562 GB/s, 100/0 Fanxiang both sides 1M read QD16 jobs=3D1, rot first: 0.278 GB/s, 100/0 HDD -> 0.562 GB/s, 100/0 Fanxiang Signed-off-by: Chen Cheng --- drivers/md/raid1.c | 42 +++++++++++++++++++++++++++++++++--------- 1 file changed, 33 insertions(+), 9 deletions(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index 319b24bcab5b..36520e48826f 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -784,17 +784,36 @@ struct read_balance_ctl { int closest_dist_disk; unsigned int min_pending; int min_pending_disk; int sequential_disk; int readable_disks; + bool min_pending_nonrot; + bool sequential_nonrot; }; =20 static int raid1_rr_pos(int disk, int start, int n) { return ((disk % n) - start + n) % n; } =20 +static bool is_better_disk(unsigned int pending, int disk, bool nonrot, + const struct read_balance_ctl *ctl, + int rr_start, int n) +{ + if (ctl->min_pending_disk < 0) + return true; + if (ctl->min_pending < pending) + return false; + if (ctl->min_pending > pending) + return true; + if (nonrot && !ctl->min_pending_nonrot) + return true; + return nonrot && ctl->min_pending_nonrot && + raid1_rr_pos(disk, rr_start, n) < + raid1_rr_pos(ctl->min_pending_disk, rr_start, n); +} + static int choose_best_rdev(struct r1conf *conf, struct r1bio *r1_bio) { int disk; int rr_start =3D 0; bool has_nonrot =3D READ_ONCE(conf->nonrot_disks); @@ -812,10 +831,11 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) =20 for (disk =3D 0 ; disk < conf->raid_disks * 2 ; disk++) { struct md_rdev *rdev; sector_t dist; unsigned int pending; + bool nonrot; =20 if (r1_bio->bios[disk] =3D=3D IO_BLOCKED) continue; =20 rdev =3D conf->mirrors[disk].rdev; @@ -827,14 +847,16 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) set_bit(R1BIO_FailFast, &r1_bio->state); =20 pending =3D atomic_read(&rdev->nr_pending); dist =3D abs(r1_bio->sector - READ_ONCE(conf->mirrors[disk].head_position)); + nonrot =3D test_bit(Nonrot, &rdev->flags); =20 /* Don't change to another disk for sequential reads */ if (is_sequential(conf, disk, r1_bio)) { - if (!should_choose_next(conf, disk) && !pending) + if (!should_choose_next(conf, disk) && !pending && + (nonrot || !has_nonrot)) return disk; =20 /* * Add 'pending' to avoid choosing this disk if * there is other idle disk. @@ -842,21 +864,22 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) pending++; /* * If there is no other idle disk, this disk * will be chosen. */ - ctl.sequential_disk =3D disk; + if (ctl.sequential_disk < 0 || + (nonrot && !ctl.sequential_nonrot)) { + ctl.sequential_disk =3D disk; + ctl.sequential_nonrot =3D nonrot; + } } =20 - if (ctl.min_pending > pending || - (has_nonrot && ctl.min_pending =3D=3D pending && - ctl.min_pending_disk >=3D 0 && - raid1_rr_pos(disk, rr_start, conf->raid_disks) < - raid1_rr_pos(ctl.min_pending_disk, rr_start, - conf->raid_disks))) { + if (is_better_disk(pending, disk, nonrot, &ctl, + rr_start, conf->raid_disks)) { ctl.min_pending =3D pending; ctl.min_pending_disk =3D disk; + ctl.min_pending_nonrot =3D nonrot; } =20 if (ctl.closest_dist > dist) { ctl.closest_dist =3D dist; ctl.closest_dist_disk =3D disk; @@ -865,11 +888,12 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) =20 /* * sequential IO size exceeds optimal iosize, however, there is no other * idle disk, so choose the sequential disk. */ - if (ctl.sequential_disk !=3D -1 && ctl.min_pending !=3D 0) + if (ctl.sequential_disk !=3D -1 && ctl.min_pending !=3D 0 && + (ctl.sequential_nonrot || !has_nonrot)) return ctl.sequential_disk; =20 /* * If all disks are rotational, choose the closest disk. If any disk is * non-rotational, choose the disk with less pending request even the --=20 2.55.0 From nobody Mon Sep 28 20:05:16 2026 Received: from va-2-28.ptr.blmpb.com (va-2-28.ptr.blmpb.com [209.127.231.28]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DAADC275AFD for ; Tue, 18 Aug 2026 07:07:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.127.231.28 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036840; cv=none; b=TDlINIV5j7lDm82EomS+MJV2UOH6yWmvdZsP8Q/98B+cNADNJAthhxI8ANCfN2ahryLOkc0CdCdgNJnP5C0tM1IAYEtH+qD3Y9cVb2eSQjlohxoyxW96e+w672jLv/GwaCBbsXUQR7dU12Y2oDXNjVwbexwsSHnJIC13xXBtllM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036840; c=relaxed/simple; bh=rv1clD//d0xcZ23jWUinlP3+s//4g6cIpMryZigiW6o=; h=Mime-Version:From:Subject:To:Message-Id:In-Reply-To:Cc:Date: Content-Type:References; b=i6C7WfRRI1WZbNtAm3WgLuMVpOxND/m74+QuyQgfjej7DbrMSdlTcHZIpkS86Nl8ATMufgWKRKA9jzv4IjQ77LqUb16PBNGLMmkEizjG3GGJ4rLDOn4CIMrYir+hMrk+3us5+pyX3Ekb2nqCRvJoja4OaX8gm6u+yMluFbm0hpk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com; spf=pass smtp.mailfrom=fnnas.com; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b=AGi+N2N2; arc=none smtp.client-ip=209.127.231.28 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=fnnas.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b="AGi+N2N2" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; s=s1; d=fnnas-com.20200927.dkim.feishu.cn; t=1787036827; h=from:subject:mime-version:from:date:message-id:subject:to:cc: reply-to:content-type:mime-version:in-reply-to:message-id; bh=ORx37DHw0St0qaQz+AO1Fr6HnBPdlA9eDeBpBmIdWJA=; b=AGi+N2N2qFsa+7gJP1uM+jjtiDPWOOz63nbmxiICzZN+IKgfOJzllGKHRieyxMCXh6bco3 Ob5jrSnsa6IWz0eEIiUbvzn5sUdZCenspVqF//CbGgPp/Mu0hGwB6Tsrhdx5UTmB7UKjMu N5oARy2Iy7Iw2vOWn5vKiAwQbds7vU4chtPeuJkchgABkNAGOI57PbM/86Hp/taYirp5Ol nEATAuDkfXBbFST27CcrZxLJ3urTsBt31ge6/3vC2Ux8kBLK21aFiRNjNuTncmtaVqP1gT 1Uonbbw7er5BeR9PjqC8SSLqvpAzsW4bBVC8g4rUdWESAkg+J0GMnjVnoxRgiA== Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Transfer-Encoding: quoted-printable From: "Chen Cheng" Subject: [RFC PATCH 3/5] md/raid1: do not send random reads to a rot disk To: , , Received: from fedora ([183.34.162.43]) by smtp.feishu.cn with ESMTPS; Tue, 18 Aug 2026 15:07:04 +0800 Message-Id: <20260818070646.1029149-4-chencheng@fnnas.com> In-Reply-To: <20260818070646.1029149-1-chencheng@fnnas.com> Cc: , Date: Tue, 18 Aug 2026 15:06:44 +0800 X-Mailer: git-send-email 2.55.0 X-Lms-Return-Path: X-Original-From: chencheng@fnnas.com References: <20260818070646.1029149-1-chencheng@fnnas.com> Content-Type: text/plain; charset="utf-8" From: Chen Cheng The previous change keeps sequential nonrot I/O off a rot disk. Current: 1. Sequential reads stay on nonrot when a nonrot disk is readable. 2. Random reads still pick the disk with the lowest nr_pending. 3. Rot disks also join that compare. Problem: 1. A nonrot disk is fast. It can have more pending I/O. 2. A rot disk is slow. It can have fewer pending I/O. 3. Then the next random read goes to the rot disk. 4. Why send a random 4k read to the slow disk? Improve: 1. If a nonrot disk is readable, do not use rot disks for min_pending. 2. Random reads stay on nonrot disks. 3. If no nonrot disk is readable, still pick a rot disk by head position. Tested with fio libaio direct=3D1 (NVMe scheduler none, SATA scheduler mq-deadline): - RAID1 of Predator GM9000 + SATA HDD: 4k randread QD1 jobs=3D1: 0.097 GB/s, 100/0 NVMe -> 0.084 GB/s, 100/0 NVMe. Both disks have pending 0, so stock already stayed on NVMe. 4k randread QD8 jobs=3D1: 0.363 GB/s, 99.6/0.4 NVMe/HDD, clat 87 us -> 0.671 GB/s, 100/0 NVMe, clat 46 us (-47%) 4k randread QD16 jobs=3D1: 0.668 GB/s, 99.7/0.3 NVMe/HDD, clat 95 us -> 1.215 GB/s, 100/0 NVMe, clat 52 us (-45%) - RAID1 of Fanxiang S103Pro + SATA HDD: 4k randread QD1 jobs=3D1: 0.077 GB/s, 100/0 Fanxiang -> 0.073 GB/s, 100/0 Fanxiang 4k randread QD8 jobs=3D1: 0.278 GB/s, 99.5/0.5 Fanxiang/HDD, clat 114 us -> 0.389 GB/s, 100/0 Fanxiang, clat 78 us (-31%) 4k randread QD16 jobs=3D1: 0.395 GB/s, 99.6/0.4 Fanxiang/HDD, clat 162 us -> 0.398 GB/s, 100/0 Fanxiang, clat 159 us (already at the Fanxiang limit) Signed-off-by: Chen Cheng --- drivers/md/raid1.c | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index 36520e48826f..523b55d42779 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -832,10 +832,11 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) for (disk =3D 0 ; disk < conf->raid_disks * 2 ; disk++) { struct md_rdev *rdev; sector_t dist; unsigned int pending; bool nonrot; + bool can_pick; =20 if (r1_bio->bios[disk] =3D=3D IO_BLOCKED) continue; =20 rdev =3D conf->mirrors[disk].rdev; @@ -848,15 +849,16 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) =20 pending =3D atomic_read(&rdev->nr_pending); dist =3D abs(r1_bio->sector - READ_ONCE(conf->mirrors[disk].head_position)); nonrot =3D test_bit(Nonrot, &rdev->flags); + can_pick =3D nonrot || !has_nonrot; =20 /* Don't change to another disk for sequential reads */ if (is_sequential(conf, disk, r1_bio)) { if (!should_choose_next(conf, disk) && !pending && - (nonrot || !has_nonrot)) + can_pick) return disk; =20 /* * Add 'pending' to avoid choosing this disk if * there is other idle disk. @@ -871,11 +873,12 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) ctl.sequential_disk =3D disk; ctl.sequential_nonrot =3D nonrot; } } =20 - if (is_better_disk(pending, disk, nonrot, &ctl, + if (can_pick && + is_better_disk(pending, disk, nonrot, &ctl, rr_start, conf->raid_disks)) { ctl.min_pending =3D pending; ctl.min_pending_disk =3D disk; ctl.min_pending_nonrot =3D nonrot; } --=20 2.55.0 From nobody Mon Sep 28 20:05:16 2026 Received: from va-2-27.ptr.blmpb.com (va-2-27.ptr.blmpb.com [209.127.231.27]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D2022339396 for ; Tue, 18 Aug 2026 07:07:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.127.231.27 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036841; cv=none; b=G0iTuauqgxbgMbnYfZRmjD+zQb3Tr66DjoNU2yrJEEsviLyxPjxBLFOXnf0o6nHUhO9gNqTVDSWDTcXdV3m8L4cUSoDfMkNlRuFYcPpZZRxLNailbCw/RzFLIl0Kjh3gJ4dIe5chQuZxSm+qidvlxJi/se5PzFRhgfu5GA4XkwU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036841; c=relaxed/simple; bh=szdIFnVZgylGKc2n0BYaZd+N/k33uRy5wb8YdUPKLKg=; h=Content-Type:To:Mime-Version:Subject:References:In-Reply-To:Cc: From:Date:Message-Id; b=erYizgz4xuWDhjQoCwMmJwNeGEVyebLkO0iFopxwsYkjAMV1rvaDBFmK8UHW6QMJxzobbOvzxEZ43PL1Af5mNacoJiS358XmdwIMwelRe5Ea64m/WSny9DlIcRWg+TcKmzkvEgV5/qtTWdwdUBrWrTaK53HhroXY6CPXDKeMZi0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com; spf=pass smtp.mailfrom=fnnas.com; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b=B4nZGQAr; arc=none smtp.client-ip=209.127.231.27 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=fnnas.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b="B4nZGQAr" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; s=s1; d=fnnas-com.20200927.dkim.feishu.cn; t=1787036829; h=from:subject:mime-version:from:date:message-id:subject:to:cc: reply-to:content-type:mime-version:in-reply-to:message-id; bh=GVV+VGR8AtIgAO3oYYtjKPhEqqtvJkvnoSHCoX8aQR0=; b=B4nZGQArrt8A+ISbZ+SoSG5ekm8/ZS8EUUfhzcN5zh9j6qAomOqMloTuEoiotFyPmXL/3j 6dySPqwopxv2K6oLQO98pA2/89R8fZ03a0phsPoB7KcyfoHUnAlY3GvFPJkNTfb6c1tj9N kRTODY3HcoedNLrbIiFObcfyHCK7seDHFHRc+03vZLr8/VvDGhHFJPPB36Udy6NUrLnmpy P5/9V/WtFKeGCrPMFlEDQcTXEOfhcAUyH+t9qptBisr+yvnvm4xRAap15gQYX+OhFIGYAq n4bhyoqv52PeSzw9n+vTuof4BN9isEdFpEudD4izg/HAPrawWCBIXCh8iK6oVg== X-Mailer: git-send-email 2.55.0 X-Original-From: chencheng@fnnas.com X-Lms-Return-Path: To: , , Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Subject: [RFC PATCH 4/5] md/raid1: use rot policy when no nonrot disk is readable References: <20260818070646.1029149-1-chencheng@fnnas.com> In-Reply-To: <20260818070646.1029149-1-chencheng@fnnas.com> Received: from fedora ([183.34.162.43]) by smtp.feishu.cn with ESMTPS; Tue, 18 Aug 2026 15:07:06 +0800 Cc: , From: "Chen Cheng" Date: Tue, 18 Aug 2026 15:06:45 +0800 Message-Id: <20260818070646.1029149-5-chencheng@fnnas.com> Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Chen Cheng has_nonrot selects mixed policy or rot-only policy. Current: 1. has_nonrot is true if conf->nonrot_disks > 0. 2. nonrot_disks counts every nonrot disk. 3. A Faulty disk, a rebuild disk, and a WriteMostly disk still count. Problem: 1. The array is NVMe + HDD. The NVMe fails. Only the HDD can take reads. 2. nonrot_disks is still 1. The code thinks this is a mixed array. 3. Sequential reads on the HDD do not stay on the HDD. Mixed policy will not keep a rot disk. 4. Every read still advances the nonrot round-robin, though no nonrot disk can take the read. Improve: 1. Look at disks that can take this read. 2. If none of them is nonrot, use the rot-only policy. Signed-off-by: Chen Cheng --- drivers/md/raid1.c | 20 +++++++++++++++++++- 1 file changed, 19 insertions(+), 1 deletion(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index 523b55d42779..f476d4dea4be 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -793,10 +793,28 @@ struct read_balance_ctl { static int raid1_rr_pos(int disk, int start, int n) { return ((disk % n) - start + n) % n; } =20 +static bool raid1_has_readable_nonrot(struct r1conf *conf, + struct r1bio *r1_bio) +{ + int disk; + + for (disk =3D 0; disk < conf->raid_disks * 2; disk++) { + struct md_rdev *rdev; + + if (r1_bio->bios[disk] =3D=3D IO_BLOCKED) + continue; + rdev =3D conf->mirrors[disk].rdev; + if (rdev_readable(rdev, r1_bio) && + test_bit(Nonrot, &rdev->flags)) + return true; + } + return false; +} + static bool is_better_disk(unsigned int pending, int disk, bool nonrot, const struct read_balance_ctl *ctl, int rr_start, int n) { if (ctl->min_pending_disk < 0) @@ -814,11 +832,11 @@ static bool is_better_disk(unsigned int pending, int = disk, bool nonrot, =20 static int choose_best_rdev(struct r1conf *conf, struct r1bio *r1_bio) { int disk; int rr_start =3D 0; - bool has_nonrot =3D READ_ONCE(conf->nonrot_disks); + bool has_nonrot =3D raid1_has_readable_nonrot(conf, r1_bio); struct read_balance_ctl ctl =3D { .closest_dist_disk =3D -1, .closest_dist =3D MaxSector, .min_pending_disk =3D -1, .min_pending =3D UINT_MAX, --=20 2.55.0 From nobody Mon Sep 28 20:05:16 2026 Received: from va-2-29.ptr.blmpb.com (va-2-29.ptr.blmpb.com [209.127.231.29]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B4C5A4B048C for ; Tue, 18 Aug 2026 07:09:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.127.231.29 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036961; cv=none; b=DUM6w+D0+IsqAVheqBpy9OKc8yx3eXJDrhQEueCrCMGgg44cseaLa9mD6LERUC8vWeNvKvQYWSiFJ49dI6Qimu4KiLbUL/sU9g0gPt7skamc/hQNYFtXs2kt6AVjS5Zyf+W8wAmY3TtEb5fdn7GLLRdORsvh/lvWivABqy804AA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787036961; c=relaxed/simple; bh=H9TJEGBnS2ZQozieVZAAkazrp0X0uS816ZTO/8oVRAU=; h=Date:Mime-Version:From:In-Reply-To:To:Subject:Message-Id: References:Content-Type:Cc; b=OTShPlUf+hJjeQm7m4uy44FZyj4LVtO/Af5zPi1vaN9OHL8XzfEoOfRVOtq86EycKcflUdn2Juur5kit0GCTCdjJzIZHSTHNas76UHjbr9rPTZqT7qQXmlB/znCsxOIQOnSiINdU2oI/NDMNS1Nv8Jc+lqlJLcKvRHTtYDmDCEc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com; spf=pass smtp.mailfrom=fnnas.com; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b=E/mn8jeO; arc=none smtp.client-ip=209.127.231.29 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=fnnas.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=fnnas.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=fnnas-com.20200927.dkim.feishu.cn header.i=@fnnas-com.20200927.dkim.feishu.cn header.b="E/mn8jeO" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; s=s1; d=fnnas-com.20200927.dkim.feishu.cn; t=1787036832; h=from:subject:mime-version:from:date:message-id:subject:to:cc: reply-to:content-type:mime-version:in-reply-to:message-id; bh=2iU1ricYfNDrpH6ArL1TkZe8uhXQBTjWWpyRZXWnbi8=; b=E/mn8jeOcPwfqXmUb9P2CKsxOIbbozGup4grkknYRYPQHC/aX8Bdq7k6kCS3Pm+WUdljMy NtHFox+7l8DmT8iQMu3ygg5LIR5esqrF3jXWlH3Dp2xblesCUVWDbY+nMewRliw+bhawei 80G0Qx7g2XgY1M9EsbxNbBO1oVffIg9ouFYUVKt8NLfwqoZZBphXfeHac9lcbahIwgqxYD 7WDEsSZyyplSyB+xqkvPWN9C3DTREEBFoUu1LQyxWecoErzQa3/IKduEzN7k5SO+1D1hQL L/AMc4Ss9sRkrRDlQfcGTYFT54nBwMhaxyfKddjI31Qa8P98eBOukC65ZFmAvg== Date: Tue, 18 Aug 2026 15:06:46 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 X-Mailer: git-send-email 2.55.0 X-Original-From: chencheng@fnnas.com Received: from fedora ([183.34.162.43]) by smtp.feishu.cn with ESMTPS; Tue, 18 Aug 2026 15:07:09 +0800 From: "Chen Cheng" In-Reply-To: <20260818070646.1029149-1-chencheng@fnnas.com> X-Lms-Return-Path: To: , , Subject: [RFC PATCH 5/5] md/raid1: clarify choose_best_rdev comments Message-Id: <20260818070646.1029149-6-chencheng@fnnas.com> Content-Transfer-Encoding: quoted-printable References: <20260818070646.1029149-1-chencheng@fnnas.com> Cc: , Content-Type: text/plain; charset="utf-8" From: Chen Cheng Write the rot/nonrot read policy in short comments next to the code. No functional change. Signed-off-by: Chen Cheng --- drivers/md/raid1.c | 48 +++++++++++++++++++++++++++++++++------------- 1 file changed, 35 insertions(+), 13 deletions(-) diff --git a/drivers/md/raid1.c b/drivers/md/raid1.c index f476d4dea4be..897eef3a022d 100644 --- a/drivers/md/raid1.c +++ b/drivers/md/raid1.c @@ -788,15 +788,17 @@ struct read_balance_ctl { int readable_disks; bool min_pending_nonrot; bool sequential_nonrot; }; =20 +/* Offset from rr start. Replacement uses the same slot as primary. */ static int raid1_rr_pos(int disk, int start, int n) { return ((disk % n) - start + n) % n; } =20 +/* True if some readable member is nonrot. */ static bool raid1_has_readable_nonrot(struct r1conf *conf, struct r1bio *r1_bio) { int disk; =20 @@ -811,10 +813,11 @@ static bool raid1_has_readable_nonrot(struct r1conf *= conf, return true; } return false; } =20 +/* Lower pending wins. Same pending: prefer nonrot, then rr order. */ static bool is_better_disk(unsigned int pending, int disk, bool nonrot, const struct read_balance_ctl *ctl, int rr_start, int n) { if (ctl->min_pending_disk < 0) @@ -828,10 +831,24 @@ static bool is_better_disk(unsigned int pending, int = disk, bool nonrot, return nonrot && ctl->min_pending_nonrot && raid1_rr_pos(disk, rr_start, n) < raid1_rr_pos(ctl->min_pending_disk, rr_start, n); } =20 +/* + * Choose a readable disk for this read. + * + * Prefer nonrot. Use rot only if no nonrot disk is readable. + * + * Sequential idle: keep this disk. Mixed array: do not keep a + * rot disk (after a write every disk looks sequential). + * Sequential busy: try an idle disk of the same class. If none + * is idle, keep the sequential disk. + * + * Else fewest pending I/Os among disks we may pick. Same + * pending: nonrot, then round-robin. Rot-only with no idle + * disk: closest head. + */ static int choose_best_rdev(struct r1conf *conf, struct r1bio *r1_bio) { int disk; int rr_start =3D 0; bool has_nonrot =3D raid1_has_readable_nonrot(conf, r1_bio); @@ -867,26 +884,32 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) =20 pending =3D atomic_read(&rdev->nr_pending); dist =3D abs(r1_bio->sector - READ_ONCE(conf->mirrors[disk].head_position)); nonrot =3D test_bit(Nonrot, &rdev->flags); + /* + * If a nonrot disk is readable, pick only nonrot. + * Else we must use rot. + */ can_pick =3D nonrot || !has_nonrot; =20 - /* Don't change to another disk for sequential reads */ + /* + * Idle sequential disk: return it now, unless + * should_choose_next() wants another disk. + * Mixed array: do not return a rot disk. + */ if (is_sequential(conf, disk, r1_bio)) { if (!should_choose_next(conf, disk) && !pending && can_pick) return disk; =20 - /* - * Add 'pending' to avoid choosing this disk if - * there is other idle disk. - */ + /* Make an idle disk win over this busy one. */ pending++; /* - * If there is no other idle disk, this disk - * will be chosen. + * Remember the first sequential disk. + * A nonrot disk may replace a rot disk. + * A later nonrot disk may not replace an earlier one. */ if (ctl.sequential_disk < 0 || (nonrot && !ctl.sequential_nonrot)) { ctl.sequential_disk =3D disk; ctl.sequential_nonrot =3D nonrot; @@ -906,22 +929,21 @@ static int choose_best_rdev(struct r1conf *conf, stru= ct r1bio *r1_bio) ctl.closest_dist_disk =3D disk; } } =20 /* - * sequential IO size exceeds optimal iosize, however, there is no other - * idle disk, so choose the sequential disk. + * Keep the sequential disk if no idle peer should take it. + * If a nonrot disk is readable: keep only a nonrot sequential disk. + * If not: an idle rot disk may take it. */ if (ctl.sequential_disk !=3D -1 && ctl.min_pending !=3D 0 && (ctl.sequential_nonrot || !has_nonrot)) return ctl.sequential_disk; =20 /* - * If all disks are rotational, choose the closest disk. If any disk is - * non-rotational, choose the disk with less pending request even the - * disk is rotational, which might/might not be optimal for raids with - * mixed ratation/non-rotational disks depending on workload. + * No readable nonrot disk: closest disk, unless some disk is idle. + * Some readable nonrot disk: that nonrot disk with fewest pending I/Os. */ if (ctl.min_pending_disk !=3D -1 && (has_nonrot || ctl.min_pending =3D=3D 0)) return ctl.min_pending_disk; else --=20 2.55.0