From nobody Thu Sep 24 13:47:13 2026 Received: from mail-pj2-f13.google.com (mail-pj2-f13.google.com [74.125.227.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6D8D830C60F for ; Thu, 24 Sep 2026 05:31:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790227902; cv=none; b=rsECHSYJXsK2fXWOel9t+ZdBExUVv4yR+xTPZuMkhv1jYbMTSamByLc3nSFnveYIbRh67fqW1UtgKAp0EdPTc+LT+TMt4XzKPFKPY7LqQ1ikNo/5DV/Ty5uECJHImaRpn+tTbvdMTeh2m6X0PRJP4Vb+zPY/umQF6QxBirfE1Z4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790227902; c=relaxed/simple; bh=IYNqVIY17mzknhWvzprnG6Kr7mEzAesMdvGwaC7MOn0=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=Lnbq+uvoUaVW0JO4UJhJjjRccPv2GMK0xCqmK9/EUweROWxsgMQqwuINzH21kCMw7YDxwAVU2ryrFepdXdadPhqdc3PvNYMGGu9/KtohfKav0VaMQIynOVts7uZRpV4sg3lrsfz0L8Hd83Z56Dupszfy3oLTl7HGWGp9mkEPloo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=EPqU7wq3; arc=none smtp.client-ip=74.125.227.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="EPqU7wq3" Received: by mail-pj2-f13.google.com with SMTP id 98e67ed59e1d1-39dacf053eeso922211a91.2 for ; Wed, 23 Sep 2026 22:31:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790227899; x=1790832699; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=hVsHh399mrI6w8J3GjDXw0fRpLh/58Rxk2DJlkQuAME=; b=EPqU7wq3utRRJdr3gUfjHjVsM5O7461S15CkWI/5nyAu6NkuIZjEU0QqVqxRbzhfdl 5ILiFBI7k/9EsXTGqCiwLM0sS+rKmR7/PYNh8ADLtksNFsk/2ISLseiSmXl/dGSCmEjh US2l3cvuXyXlwDmChwtLo3KBHsFVH3SvCqamEFyETmBzB+N7B+cKmdiyEDuZMr92UBNQ JVG+nUKqS0g6SN+VCcttt2YVemHgm2T9devAYoeMbAomMtonquynzV8DKxZCSXzTsd5g RhyTiSrlJRXpVbxdxuyyqFRxK++9mfqP24wMsnGPXSzuU4i0d+GTkHfFi2cTySuuIKoa 1s/Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790227899; x=1790832699; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=hVsHh399mrI6w8J3GjDXw0fRpLh/58Rxk2DJlkQuAME=; b=p5+u1r7qjYYFwE7DRNuV5Ay+/nrNgdK6aEYAJ8pO7pG1ofYJCUQNxRkzaCuD0HENpj CIMZ2V0Vkvv+okY4iS55zXSGsHT8GwUykYYrG8Q6Qs2xZNQJM7OZecNwX8M0snejCLGw gdPE1AcA4arQff4o8d1gHCDzK+XESLyd99pHdHjl/PacLZsJNXYRLHHpdDt2vD9WE6Il Ve68AtU68lsNRZnTAt+B2nhiBQA5cK0tYEpeQIZnQCXNhi3oAbw1TOveEdUUKyk7oOdS qlJ6y7UKFthdCwjlvA+/oU5bt/4MMNnHrCs6DeOk6IZlBiTc2Rbs5FYxen8qOfms3y8u j/dg== X-Forwarded-Encrypted: i=1; AKwUvBzIqAPnwlhY8DzXQsXADqAZF7iUe2mN+jZqhCfAyqe+Ke2lMKNTCODpgdst1YD/EtsN2s4VAuidMA354Vs=@vger.kernel.org X-Gm-Message-State: AFuF++kLs6E2Wm5eR1d5QhrVp5ZFoWKF2foBAO33GkoqTxcTA45ib3iu rzExa4hMd9gC4MnBD6kFPLSc72qM5kEphf5Keu50nIAl3mwDztSKZyuD X-Gm-Gg: AYBFou2TQL/Y43RZqwS0qitNAXu1Mz/X5aHatpXB3Lyg1KzcNcnd663FseUE8WYpSGA BHpHAtxN8lmMfBpGt0GYa5Y2VsnTIu1yD4u8f6UpMK4gs87EUaH+obkyljD/pKL5WmNBOf+8LOw aoTLAWmEHM7lG+ORfvI3twFc6qPr8GEIOn4aJIK90v/4jPlkCn5hp6x7HfnEUgOTRWf4PioaTim PlM/CQE8fyxCcUhGKRkupUjs7jt4nhIvfmYkeMfYqDw8Mt96AZHbDyj4s2AK6eHbBqPwyIAPlp5 EK26pMSRZU+TeKnjPNWpbW2XWr93e/DuKxyJkQ15tiRAjtMRft/Ebui78KAoGW9S94LFY6d3NdI ImAwhM4xXIjTcjIOxD9/WId8ilxA7tZoguBFv0zfa4G17jl1CpUiXO1qdbKrfqJlllwdyvkhn3O S0cHpFY/fF11K613t+006zUaycMUGhbcKE1u2PEa2pIurmoaUX9XSySvYBKYlXZLet4f/3/DG3T gUKJR1sIIlGCMxWuJjGxY5aBbjkXtOmbxhN09Vo5OCv8oSR2lqgl2VQQYsoWzpF5NPPS1lrbwkz B/ns X-Received: by 2002:a17:90b:4a08:b0:39e:6a7e:ee16 with SMTP id 98e67ed59e1d1-3a098da1d0fmr1184264a91.34.1790227894800; Wed, 23 Sep 2026 22:31:34 -0700 (PDT) Received: from yupeng-XPS-15-9520.. (c-73-169-192-12.hsd1.wa.comcast.net. [73.169.192.12]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-3a0974ec5d0sm2806276a91.4.2026.09.23.22.31.33 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 23 Sep 2026 22:31:34 -0700 (PDT) From: Peng Yu To: Christoph Hellwig , Sagi Grimberg , Chaitanya Kulkarni Cc: Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Josef Bacik , Jens Axboe , Maurizio Lombardi , cgroups@vger.kernel.org, linux-block@vger.kernel.org, linux-nvme@lists.infradead.org, linux-kernel@vger.kernel.org, Peng Yu Subject: [PATCH v4] nvmet: add cgroup_path to charge namespace I/O to a cgroup Date: Wed, 23 Sep 2026 22:31:21 -0700 Message-ID: <20260924053121.17703-1-yupeng0921@gmail.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Scenario: * Create multiple nvmet subsystems/namespaces. * The namespaces are backed by different LVM logical volumes. * Some of the logical volumes share the same physical volumes. * The subsystems are exported to different users. * We should provide each user a specific iops/bps quota, thus a noisy neighbor won't impact the performance of other logical volumes. Implementation: * Add a `cgroup_path` attribute under the nvmet namespace folder, e.g.: /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:bdev/namespaces= /1/cgroup_path * We can write a cgroup path to it, then that cgroup will control the IOs used by the namespace. * Write an empty string to clear it. * The cgroup_path could be modified when the namespace is disabled. * Support both block device backed namespaces and file backed namespaces. * Only the direct io scenario is supported, enabling a namespace that has both cgroup_path and buffered_io set fails with -EINVAL. Change since v2: * Fix the in ns->cgroup_path error free issue nvmet_ns_cgroup_path_show. * Use nvmet_blkcg_begin/end in nvmet_bdev_zmgmt_send_work. Change since v3: * Use cgroup_id instead of cgroup_path. Test: 1 Load kernel modules. sudo modprobe brd rd_nr=3D2 rd_size=3D1048576 sudo modprobe nvmet sudo modprobe nvmet_tcp sudo modprobe nvme_tcp 2 The two brd devices are used for a bdev backed namespace and a file backed namespace, remember their major/minor. lsblk --noheadings --nodeps --output NAME,MAJ:MIN,SIZE /dev/ram0 /dev/ram1 ram0 1:0 1G ram1 1:1 1G 3 Enable the io controller and create two cgroups, one for the bdev backed namespace, another for the file backed namespace. echo "+io" | sudo tee /sys/fs/cgroup/cgroup.subtree_control sudo mkdir -p /sys/fs/cgroup/nvmet-bdev sudo mkdir -p /sys/fs/cgroup/nvmet-file 4 Create a file system and a file on ram0. sudo mkfs.ext4 /dev/ram0 sudo mkdir -p /mnt/nvmet-file sudo mount -o noatime /dev/ram0 /mnt/nvmet-file sudo dd if=3D/dev/zero of=3D/mnt/nvmet-file/ns.img bs=3D1M count=3D512 ofla= g=3Ddirect 5 Create the bdev backed subsystem and namespace. sudo mkdir -p /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:bde= v/namespaces/1 echo 1 | sudo tee /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01= :bdev/attr_allow_any_host echo /dev/ram1 | sudo tee /sys/kernel/config/nvmet/subsystems/nqn.2026-09.i= o.test01:bdev/namespaces/1/device_path stat -c %i /sys/fs/cgroup/nvmet-bdev | sudo tee /sys/kernel/config/nvmet/su= bsystems/nqn.2026-09.io.test01:bdev/namespaces/1/cgroup_id echo 1 | sudo tee /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01= :bdev/namespaces/1/enable 6 Create the file backed subsystem and namespace. sudo mkdir -p /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:fil= e/namespaces/1 echo 1 | sudo tee /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01= :file/attr_allow_any_host echo /mnt/nvmet-file/ns.img | sudo tee /sys/kernel/config/nvmet/subsystems/= nqn.2026-09.io.test01:file/namespaces/1/device_path stat -c %i /sys/fs/cgroup/nvmet-file | sudo tee /sys/kernel/config/nvmet/su= bsystems/nqn.2026-09.io.test01:file/namespaces/1/cgroup_id echo 1 | sudo tee /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01= :file/namespaces/1/enable 7 Export the two subsystems to 127.0.0.1. sudo mkdir -p /sys/kernel/config/nvmet/ports/1 echo ipv4 | sudo tee /sys/kernel/config/nvmet/ports/1/addr_adrfam echo tcp | sudo tee /sys/kernel/config/nvmet/ports/1/addr_trtype echo 127.0.0.1 | sudo tee /sys/kernel/config/nvmet/ports/1/addr_traddr echo 4420 | sudo tee /sys/kernel/config/nvmet/ports/1/addr_trsvcid sudo ln -s /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:bdev /= sys/kernel/config/nvmet/ports/1/subsystems/nqn.2026-09.io.test01:bdev sudo ln -s /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:file /= sys/kernel/config/nvmet/ports/1/subsystems/nqn.2026-09.io.test01:file 8 Connect the two subsystems from 127.0.0.1. sudo nvme connect -t tcp -a 127.0.0.1 -s 4420 -n nqn.2026-09.io.test01:bdev sudo nvme connect -t tcp -a 127.0.0.1 -s 4420 -n nqn.2026-09.io.test01:file sudo udevadm settle 9 Set riops to 20k and run fio. echo "1:1 riops=3D20000 wiops=3Dmax rbps=3Dmax wbps=3Dmax" | sudo tee /sys/= fs/cgroup/nvmet-bdev/io.max echo "1:0 riops=3D20000 wiops=3Dmax rbps=3Dmax wbps=3Dmax" | sudo tee /sys/= fs/cgroup/nvmet-file/io.max sudo fio --name=3Dt --filename=3D/dev/nvme0n1 --rw=3Drandread --bs=3D4k --i= odepth=3D32 --numjobs=3D1 \ --direct=3D1 --ioengine=3Dlibaio --time_based --runtime=3D10 --ramp_time=3D= 3 \ --group_reporting 2>/dev/null | grep -E '^ +(read|write): IOPS' read: IOPS=3D20.0k, BW=3D78.1MiB/s (81.9MB/s)(781MiB/10001msec) sudo fio --name=3Dt --filename=3D/dev/nvme1n1 --rw=3Drandread --bs=3D4k --i= odepth=3D32 --numjobs=3D1 \ --direct=3D1 --ioengine=3Dlibaio --time_based --runtime=3D10 --ramp_time=3D= 3 \ --group_reporting 2>/dev/null | grep -E '^ +(read|write): IOPS' read: IOPS=3D20.0k, BW=3D78.1MiB/s (81.9MB/s)(781MiB/10001msec) 10 Set rbps to 100M and run fio. echo "1:1 riops=3Dmax wiops=3Dmax rbps=3D104857600 wbps=3Dmax" | sudo tee /= sys/fs/cgroup/nvmet-bdev/io.max echo "1:0 riops=3Dmax wiops=3Dmax rbps=3D104857600 wbps=3Dmax" | sudo tee /= sys/fs/cgroup/nvmet-file/io.max sudo fio --name=3Dt --filename=3D/dev/nvme0n1 --rw=3Dread --bs=3D1M --iodep= th=3D32 --numjobs=3D1 \ --direct=3D1 --ioengine=3Dlibaio --time_based --runtime=3D10 --ramp_time=3D= 3 \ --group_reporting 2>/dev/null | grep -E '^ +(read|write): IOPS' read: IOPS=3D97, BW=3D100MiB/s (105MB/s)(1033MiB/10304msec) sudo fio --name=3Dt --filename=3D/dev/nvme1n1 --rw=3Dread --bs=3D1M --iodep= th=3D32 --numjobs=3D1 \ --direct=3D1 --ioengine=3Dlibaio --time_based --runtime=3D10 --ramp_time=3D= 3 \ --group_reporting 2>/dev/null | grep -E '^ +(read|write): IOPS' read: IOPS=3D97, BW=3D100MiB/s (105MB/s)(1031MiB/10297msec) 11 Set wiops to 20k and run fio. echo "1:1 riops=3Dmax wiops=3D20000 rbps=3Dmax wbps=3Dmax" | sudo tee /sys/= fs/cgroup/nvmet-bdev/io.max echo "1:0 riops=3Dmax wiops=3D20000 rbps=3Dmax wbps=3Dmax" | sudo tee /sys/= fs/cgroup/nvmet-file/io.max sudo fio --name=3Dt --filename=3D/dev/nvme0n1 --rw=3Drandwrite --bs=3D4k --= iodepth=3D32 --numjobs=3D1 \ --direct=3D1 --ioengine=3Dlibaio --time_based --runtime=3D10 --ramp_time=3D= 3 \ --group_reporting 2>/dev/null | grep -E '^ +(read|write): IOPS' write: IOPS=3D20.0k, BW=3D78.1MiB/s (81.9MB/s)(781MiB/10001msec); 0 zone = resets sudo fio --name=3Dt --filename=3D/dev/nvme1n1 --rw=3Drandwrite --bs=3D4k --= iodepth=3D32 --numjobs=3D1 \ --direct=3D1 --ioengine=3Dlibaio --time_based --runtime=3D10 --ramp_time=3D= 3 \ --group_reporting 2>/dev/null | grep -E '^ +(read|write): IOPS' write: IOPS=3D20.0k, BW=3D78.1MiB/s (81.9MB/s)(781MiB/10001msec); 0 zone = resets 12 Set wbps to 100M and run fio. echo "1:1 riops=3Dmax wiops=3Dmax rbps=3Dmax wbps=3D104857600" | sudo tee /= sys/fs/cgroup/nvmet-bdev/io.max echo "1:0 riops=3Dmax wiops=3Dmax rbps=3Dmax wbps=3D104857600" | sudo tee /= sys/fs/cgroup/nvmet-file/io.max sudo fio --name=3Dt --filename=3D/dev/nvme0n1 --rw=3Dwrite --bs=3D1M --iode= pth=3D32 --numjobs=3D1 \ --direct=3D1 --ioengine=3Dlibaio --time_based --runtime=3D10 --ramp_time=3D= 3 \ --group_reporting 2>/dev/null | grep -E '^ +(read|write): IOPS' write: IOPS=3D97, BW=3D100MiB/s (105MB/s)(1031MiB/10305msec); 0 zone rese= ts sudo fio --name=3Dt --filename=3D/dev/nvme1n1 --rw=3Dwrite --bs=3D1M --iode= pth=3D32 --numjobs=3D1 \ --direct=3D1 --ioengine=3Dlibaio --time_based --runtime=3D10 --ramp_time=3D= 3 \ --group_reporting 2>/dev/null | grep -E '^ +(read|write): IOPS' write: IOPS=3D97, BW=3D100MiB/s (105MB/s)(1032MiB/10297msec); 0 zone rese= ts 13 Cleanup the environment. sudo nvme disconnect -n nqn.2026-09.io.test01:bdev sudo nvme disconnect -n nqn.2026-09.io.test01:file sudo udevadm settle sudo rm -f /sys/kernel/config/nvmet/ports/1/subsystems/nqn.2026-09.io.test0= 1:bdev sudo rm -f /sys/kernel/config/nvmet/ports/1/subsystems/nqn.2026-09.io.test0= 1:file sudo rmdir /sys/kernel/config/nvmet/ports/1 echo 0 | sudo tee /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01= :bdev/namespaces/1/enable sudo rmdir /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:bdev/n= amespaces/1 sudo rmdir /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:bdev echo 0 | sudo tee /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01= :file/namespaces/1/enable sudo rmdir /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:file/n= amespaces/1 sudo rmdir /sys/kernel/config/nvmet/subsystems/nqn.2026-09.io.test01:file sudo umount /mnt/nvmet-file sudo rmdir /mnt/nvmet-file sudo modprobe -r brd sudo rmdir /sys/fs/cgroup/nvmet-bdev sudo rmdir /sys/fs/cgroup/nvmet-file sudo rmmod nvme_tcp sudo rmmod nvmet_tcp sudo rmmod nvmet Signed-off-by: Peng Yu Assisted-by: Claude:claude-fable-5 [Claude Code] Assisted-by: Claude:claude-opus-5 [Claude Code] --- drivers/nvme/target/configfs.c | 47 ++++++++++++++++++++++ drivers/nvme/target/core.c | 65 +++++++++++++++++++++++++++++++ drivers/nvme/target/io-cmd-bdev.c | 18 ++++++++- drivers/nvme/target/io-cmd-file.c | 21 +++++++++- drivers/nvme/target/nvmet.h | 47 ++++++++++++++++++++++ drivers/nvme/target/zns.c | 5 +++ 6 files changed, 200 insertions(+), 3 deletions(-) diff --git a/drivers/nvme/target/configfs.c b/drivers/nvme/target/configfs.c index 6286e38436dd..cef832303d72 100644 --- a/drivers/nvme/target/configfs.c +++ b/drivers/nvme/target/configfs.c @@ -561,6 +561,50 @@ static ssize_t nvmet_ns_device_path_store(struct confi= g_item *item, =20 CONFIGFS_ATTR(nvmet_ns_, device_path); =20 +#ifdef CONFIG_BLK_CGROUP +static ssize_t nvmet_ns_cgroup_id_show(struct config_item *item, char *pag= e) +{ + struct nvmet_ns *ns =3D to_nvmet_ns(item); + struct nvmet_subsys *subsys =3D ns->subsys; + ssize_t ret; + + mutex_lock(&subsys->lock); + ret =3D snprintf(page, PAGE_SIZE, "%llu\n", ns->cgroup_id); + mutex_unlock(&subsys->lock); + return ret; +} + +static ssize_t nvmet_ns_cgroup_id_store(struct config_item *item, + const char *page, size_t count) +{ + struct nvmet_ns *ns =3D to_nvmet_ns(item); + struct nvmet_subsys *subsys =3D ns->subsys; + u64 cgroup_id; + int ret; + + ret =3D kstrtou64(page, 0, &cgroup_id); + if (ret) + return ret; + + mutex_lock(&subsys->lock); + + if (ns->enabled) { + ret =3D -EBUSY; + goto out_unlock; + } + + /* Writing 0 clears the association. */ + ns->cgroup_id =3D cgroup_id; + ret =3D count; + +out_unlock: + mutex_unlock(&subsys->lock); + return ret; +} + +CONFIGFS_ATTR(nvmet_ns_, cgroup_id); +#endif /* CONFIG_BLK_CGROUP */ + #ifdef CONFIG_PCI_P2PDMA static ssize_t nvmet_ns_p2pmem_show(struct config_item *item, char *page) { @@ -833,6 +877,9 @@ static struct configfs_attribute *nvmet_ns_attrs[] =3D { &nvmet_ns_attr_buffered_io, &nvmet_ns_attr_revalidate_size, &nvmet_ns_attr_resv_enable, +#ifdef CONFIG_BLK_CGROUP + &nvmet_ns_attr_cgroup_id, +#endif #ifdef CONFIG_PCI_P2PDMA &nvmet_ns_attr_p2pmem, #endif diff --git a/drivers/nvme/target/core.c b/drivers/nvme/target/core.c index 43871a8f56ca..58ef833f2cb0 100644 --- a/drivers/nvme/target/core.c +++ b/drivers/nvme/target/core.c @@ -4,6 +4,7 @@ * Copyright (c) 2015-2016 HGST, a Western Digital Company. */ #define pr_fmt(fmt) KBUILD_MODNAME ": " fmt +#include #include #include #include @@ -474,8 +475,68 @@ void nvmet_put_namespace(struct nvmet_ns *ns) percpu_ref_put(&ns->ref); } =20 +#ifdef CONFIG_BLK_CGROUP +static int nvmet_blkcg_ns_enable(struct nvmet_ns *ns) +{ + struct cgroup_subsys_state *css; + struct cgroup *cgrp; + + if (!ns->cgroup_id) + return 0; + + /* + * Buffered writes will be handled by a separate thread, + * these IOs have no namespace/cgroup information at that time, + * so we don't support buffered io. + */ + if (ns->buffered_io) { + pr_err("cgroup_id is not supported with buffered_io: %s\n", + ns->device_path); + return -EINVAL; + } + + cgrp =3D cgroup_get_from_id(ns->cgroup_id); + if (IS_ERR(cgrp)) { + pr_err("failed to resolve cgroup id %llu: %ld\n", + ns->cgroup_id, PTR_ERR(cgrp)); + return PTR_ERR(cgrp); + } + + css =3D cgroup_get_e_css(cgrp, &io_cgrp_subsys); + if (!css || css->cgroup !=3D cgrp) { + pr_err("the io controller is not enabled in cgroup %llu\n", + ns->cgroup_id); + if (css) + css_put(css); + cgroup_put(cgrp); + return -EINVAL; + } + ns->blkcg_css =3D css; + cgroup_put(cgrp); + return 0; +} + +static void nvmet_blkcg_ns_disable(struct nvmet_ns *ns) +{ + if (ns->blkcg_css) { + css_put(ns->blkcg_css); + ns->blkcg_css =3D NULL; + } +} +#else +static inline int nvmet_blkcg_ns_enable(struct nvmet_ns *ns) +{ + return 0; +} + +static inline void nvmet_blkcg_ns_disable(struct nvmet_ns *ns) +{ +} +#endif /* CONFIG_BLK_CGROUP */ + static void nvmet_ns_dev_disable(struct nvmet_ns *ns) { + nvmet_blkcg_ns_disable(ns); nvmet_bdev_ns_disable(ns); nvmet_file_ns_disable(ns); } @@ -602,6 +663,10 @@ int nvmet_ns_enable(struct nvmet_ns *ns) if (ret) goto out_unlock; =20 + ret =3D nvmet_blkcg_ns_enable(ns); + if (ret) + goto out_dev_disable; + ret =3D nvmet_p2pmem_ns_enable(ns); if (ret) goto out_dev_disable; diff --git a/drivers/nvme/target/io-cmd-bdev.c b/drivers/nvme/target/io-cmd= -bdev.c index f2d9e8901df4..0014e5ef2853 100644 --- a/drivers/nvme/target/io-cmd-bdev.c +++ b/drivers/nvme/target/io-cmd-bdev.c @@ -297,6 +297,7 @@ static void nvmet_bdev_execute_rw(struct nvmet_req *req) bio =3D bio_alloc(req->ns->bdev, bio_max_segs(sg_cnt), opf, GFP_KERNEL); } + nvmet_blkcg_set_bio(req->ns, bio); bio->bi_iter.bi_sector =3D sector; bio->bi_private =3D req; bio->bi_end_io =3D nvmet_bio_done; @@ -322,6 +323,7 @@ static void nvmet_bdev_execute_rw(struct nvmet_req *req) =20 bio =3D bio_alloc(req->ns->bdev, bio_max_segs(sg_cnt), opf, GFP_KERNEL); + nvmet_blkcg_set_bio(req->ns, bio); bio->bi_iter.bi_sector =3D sector; =20 bio_chain(bio, prev); @@ -358,6 +360,7 @@ static void nvmet_bdev_execute_flush(struct nvmet_req *= req) =20 bio_init(bio, req->ns->bdev, req->inline_bvec, ARRAY_SIZE(req->inline_bvec), REQ_OP_WRITE | REQ_PREFLUSH); + nvmet_blkcg_set_bio(req->ns, bio); bio->bi_private =3D req; bio->bi_end_io =3D nvmet_bio_done; =20 @@ -366,10 +369,16 @@ static void nvmet_bdev_execute_flush(struct nvmet_req= *req) =20 u16 nvmet_bdev_flush(struct nvmet_req *req) { + bool associated; + int ret; + if (!bdev_write_cache(req->ns->bdev)) return 0; =20 - if (blkdev_issue_flush(req->ns->bdev)) + associated =3D nvmet_blkcg_begin(req->ns); + ret =3D blkdev_issue_flush(req->ns->bdev); + nvmet_blkcg_end(associated); + if (ret) return NVME_SC_INTERNAL | NVME_STATUS_DNR; return 0; } @@ -380,9 +389,11 @@ static void nvmet_bdev_execute_discard(struct nvmet_re= q *req) struct nvme_dsm_range range; struct bio *bio =3D NULL; sector_t nr_sects; + bool associated; int i; u16 status =3D NVME_SC_SUCCESS; =20 + associated =3D nvmet_blkcg_begin(ns); for (i =3D 0; i <=3D le32_to_cpu(req->cmd->dsm.nr); i++) { status =3D nvmet_copy_from_sgl(req, i * sizeof(range), &range, sizeof(range)); @@ -394,6 +405,7 @@ static void nvmet_bdev_execute_discard(struct nvmet_req= *req) nvmet_lba_to_sect(ns, range.slba), nr_sects, GFP_KERNEL, &bio); } + nvmet_blkcg_end(associated); =20 if (bio) { bio->bi_private =3D req; @@ -431,6 +443,7 @@ static void nvmet_bdev_execute_write_zeroes(struct nvme= t_req *req) struct bio *bio =3D NULL; sector_t sector; sector_t nr_sector; + bool associated; int ret; =20 if (!nvmet_check_transfer_len(req, 0)) @@ -440,8 +453,11 @@ static void nvmet_bdev_execute_write_zeroes(struct nvm= et_req *req) nr_sector =3D (((sector_t)le16_to_cpu(write_zeroes->length) + 1) << (req->ns->blksize_shift - 9)); =20 + associated =3D nvmet_blkcg_begin(req->ns); ret =3D __blkdev_issue_zeroout(req->ns->bdev, sector, nr_sector, GFP_KERNEL, &bio, 0); + nvmet_blkcg_end(associated); + if (bio) { bio->bi_private =3D req; bio->bi_end_io =3D nvmet_bio_done; diff --git a/drivers/nvme/target/io-cmd-file.c b/drivers/nvme/target/io-cmd= -file.c index 0b22d183f927..570e65259ad8 100644 --- a/drivers/nvme/target/io-cmd-file.c +++ b/drivers/nvme/target/io-cmd-file.c @@ -79,6 +79,8 @@ static ssize_t nvmet_file_submit_bvec(struct nvmet_req *r= eq, loff_t pos, struct kiocb *iocb =3D &req->f.iocb; ssize_t (*call_iter)(struct kiocb *iocb, struct iov_iter *iter); struct iov_iter iter; + bool associated; + ssize_t ret; int rw; =20 if (req->cmd->rw.opcode =3D=3D nvme_cmd_write) { @@ -97,7 +99,10 @@ static ssize_t nvmet_file_submit_bvec(struct nvmet_req *= req, loff_t pos, iocb->ki_filp =3D req->ns->file; iocb->ki_flags =3D ki_flags | iocb->ki_filp->f_iocb_flags; =20 - return call_iter(iocb, &iter); + associated =3D nvmet_blkcg_begin(req->ns); + ret =3D call_iter(iocb, &iter); + nvmet_blkcg_end(associated); + return ret; } =20 static void nvmet_file_io_done(struct kiocb *iocb, long ret) @@ -251,7 +256,13 @@ static void nvmet_file_execute_rw(struct nvmet_req *re= q) =20 u16 nvmet_file_flush(struct nvmet_req *req) { - return errno_to_nvme_status(req, vfs_fsync(req->ns->file, 1)); + bool associated; + int ret; + + associated =3D nvmet_blkcg_begin(req->ns); + ret =3D vfs_fsync(req->ns->file, 1); + nvmet_blkcg_end(associated); + return errno_to_nvme_status(req, ret); } =20 static void nvmet_file_flush_work(struct work_struct *w) @@ -274,10 +285,12 @@ static void nvmet_file_execute_discard(struct nvmet_r= eq *req) int mode =3D FALLOC_FL_PUNCH_HOLE | FALLOC_FL_KEEP_SIZE; struct nvme_dsm_range range; loff_t offset, len; + bool associated; u16 status =3D 0; int ret; int i; =20 + associated =3D nvmet_blkcg_begin(req->ns); for (i =3D 0; i <=3D le32_to_cpu(req->cmd->dsm.nr); i++) { status =3D nvmet_copy_from_sgl(req, i * sizeof(range), &range, sizeof(range)); @@ -300,6 +313,7 @@ static void nvmet_file_execute_discard(struct nvmet_req= *req) break; } } + nvmet_blkcg_end(associated); =20 nvmet_req_complete(req, status); } @@ -336,6 +350,7 @@ static void nvmet_file_write_zeroes_work(struct work_st= ruct *w) int mode =3D FALLOC_FL_ZERO_RANGE | FALLOC_FL_KEEP_SIZE; loff_t offset; loff_t len; + bool associated; int ret; =20 offset =3D le64_to_cpu(write_zeroes->slba) << req->ns->blksize_shift; @@ -347,7 +362,9 @@ static void nvmet_file_write_zeroes_work(struct work_st= ruct *w) return; } =20 + associated =3D nvmet_blkcg_begin(req->ns); ret =3D vfs_fallocate(req->ns->file, mode, offset, len); + nvmet_blkcg_end(associated); nvmet_req_complete(req, ret < 0 ? errno_to_nvme_status(req, ret) : 0); } =20 diff --git a/drivers/nvme/target/nvmet.h b/drivers/nvme/target/nvmet.h index dbda55895f4f..62bbfcc4ea14 100644 --- a/drivers/nvme/target/nvmet.h +++ b/drivers/nvme/target/nvmet.h @@ -21,6 +21,7 @@ #include #include #include +#include =20 #define NVMET_DEFAULT_VS NVME_VS(2, 1, 0) =20 @@ -115,6 +116,16 @@ struct nvmet_ns { struct nvmet_subsys *subsys; const char *device_path; =20 +#ifdef CONFIG_BLK_CGROUP + u64 cgroup_id; + /* + * Resolved from ->cgroup_id when the namespace is enabled and + * released when it is disabled, so it has the same lifetime and + * visibility rules as ->bdev and ->file. + */ + struct cgroup_subsys_state *blkcg_css; +#endif + struct config_group device_group; struct config_group group; =20 @@ -732,6 +743,42 @@ void nvmet_bdev_execute_zone_mgmt_recv(struct nvmet_re= q *req); void nvmet_bdev_execute_zone_mgmt_send(struct nvmet_req *req); void nvmet_bdev_execute_zone_append(struct nvmet_req *req); =20 +#ifdef CONFIG_BLK_CGROUP +static inline void nvmet_blkcg_set_bio(struct nvmet_ns *ns, struct bio *bi= o) +{ + if (ns->blkcg_css) + bio_associate_blkg_from_css(bio, ns->blkcg_css); +} + +static inline bool nvmet_blkcg_begin(struct nvmet_ns *ns) +{ + if (!ns->blkcg_css || !in_task() || !(current->flags & PF_KTHREAD)) + return false; + + kthread_associate_blkcg(ns->blkcg_css); + return true; +} + +static inline void nvmet_blkcg_end(bool associated) +{ + if (associated) + kthread_associate_blkcg(NULL); +} +#else +static inline void nvmet_blkcg_set_bio(struct nvmet_ns *ns, struct bio *bi= o) +{ +} + +static inline bool nvmet_blkcg_begin(struct nvmet_ns *ns) +{ + return false; +} + +static inline void nvmet_blkcg_end(bool associated) +{ +} +#endif /* CONFIG_BLK_CGROUP */ + static inline u32 nvmet_rw_data_len(struct nvmet_req *req) { return ((u32)le16_to_cpu(req->cmd->rw.length) + 1) << diff --git a/drivers/nvme/target/zns.c b/drivers/nvme/target/zns.c index 23a17c02abee..a9836d44f5e6 100644 --- a/drivers/nvme/target/zns.c +++ b/drivers/nvme/target/zns.c @@ -480,8 +480,11 @@ static void nvmet_bdev_zmgmt_send_work(struct work_str= uct *w) struct block_device *bdev =3D req->ns->bdev; sector_t zone_sectors =3D bdev_zone_sectors(bdev); u16 status =3D NVME_SC_SUCCESS; + bool associated; int ret; =20 + associated =3D nvmet_blkcg_begin(req->ns); + if (op =3D=3D REQ_OP_LAST) { req->error_loc =3D offsetof(struct nvme_zone_mgmt_send_cmd, zsa); status =3D NVME_SC_ZONE_INVALID_TRANSITION | NVME_STATUS_DNR; @@ -511,6 +514,7 @@ static void nvmet_bdev_zmgmt_send_work(struct work_stru= ct *w) status =3D blkdev_zone_mgmt_errno_to_nvme_status(ret); =20 out: + nvmet_blkcg_end(associated); nvmet_req_complete(req, status); } =20 @@ -580,6 +584,7 @@ void nvmet_bdev_execute_zone_append(struct nvmet_req *r= eq) bio =3D bio_alloc(req->ns->bdev, req->sg_cnt, opf, GFP_KERNEL); } =20 + nvmet_blkcg_set_bio(req->ns, bio); bio->bi_end_io =3D nvmet_bdev_zone_append_bio_done; bio->bi_iter.bi_sector =3D sect; bio->bi_private =3D req; base-commit: d9cc476535d29a44df3f2aa3b14af9a833fddf90 --=20 2.53.0