From nobody Sat Jul 25 19:29:10 2026 Received: from mail-ot1-f52.google.com (mail-ot1-f52.google.com [209.85.210.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 23B9944CAE0 for ; Tue, 14 Jul 2026 13:07:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034447; cv=none; b=mBhlT1cdMitTgyNr0k8iD2Y/Z9ZMVCY+CVQfJzstjFH4sSVeIzjNhKtGTIo2jSjH9UviubIpaXPZfqqURk6VpCHk9gsvk7udUlpFEh1z9f8WMCsr3tDaQpUyhGhRm2H9hRFxrUK3D2uszhrGrmKQ8d3Nd3KHOy7KEq9GEqAr2Sg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034447; c=relaxed/simple; bh=K94VRwZaVExISaYH+2jjkBkd/38Sf/HsRBDFWlaligk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=suBRcXKeOBaNiYLza4xyoET/z2cBRn2a8FBZcNzdGL5CMGSRusdem658yPaVpBYhUpkjJb1lW/uxFnt3X3BAhJnNPjUuy2io8C4+dbHhL4cVGISzCTNCj6ueROF9b8VPa9Z0bD+FlxfMdMVj2OISkILDolpN6sHmaJXCyNL10ws= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=dy4k85cS; arc=none smtp.client-ip=209.85.210.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="dy4k85cS" Received: by mail-ot1-f52.google.com with SMTP id 46e09a7af769-7e9d7464b71so1392906a34.0 for ; Tue, 14 Jul 2026 06:07:24 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1784034444; x=1784639244; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=SD1E2b021BNT8K5FqPsuJFYVGZ1yPqvbXL/PhK3jvMg=; b=dy4k85cSxEadN3g3y8QfUXYwo9Z5DurmYafCkhfw3XU+puRhRJxX28CztpSHM0cVcS SgthW+M4ntmi06z1cSHm3BDpnyK6g9UX0GQE2H1VVsJ/G+yPYAzy4iunNYgaazY9UKqD V2jnP3kbXvDWD3wigJGJVqo/Ctf5NIFvn8sJ1/zsfQgZ66qKzr3YxakfaEGxqv69FRO3 QJtn+CsZF2VaJDz0xOBuEgMEyA3Y1xS1E1zFC3sDq8RhjdyjR0RlLLRzkXL/V6XNmsnQ VZqR5leG+NDUJegAkveAvzi6hwCimWSiPFjw1rjhJuDajqTSdLG+ILSIQomfwK6Y7w2/ o4aw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784034444; x=1784639244; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=SD1E2b021BNT8K5FqPsuJFYVGZ1yPqvbXL/PhK3jvMg=; b=jk+8RtQfIj0Zrn+H0XKjU88OabF5cnJb3Su+7Om70r7ZFrxCAi5y8H6YHLjREMhA3v LUat06u8Xbk4fuA2DcRd+nH8DWgPnqy6+4Us3TFGqEGGb3S31JKCv27Oa1o9QqC3Qcbi aneht2ZbAsfVz/MUo7hl2iZ8MPjvq7yctrhggHVBuxgNugCS8KrQz3AghQbQk0UYZZBo NtaU+7AEJjHXNs6fX24lS631Th+sDpc8paFWsFPnEH1/q5a3i42TjfytXe9GBorxC1Gv 8jgpBdSREgIepjPSoh9zY4OrHqSbUG2vtbO4uNg/2ZehNraOnmtVDaUzrugvXo4qxseP GXYQ== X-Forwarded-Encrypted: i=1; AFNElJ+SfyWpvYei5dhTEfWhGksdrFto1FMy4Ygu5I0OQvTPWkN90/zw7+R21d5TZHFjl/jAdDtdRJiVN3GumIQ=@vger.kernel.org X-Gm-Message-State: AOJu0YxdIP8z7Am6SXU5cREJ1TI8VPmtxppt9g3YwPCkuSlgGJJsmt5O Y5TS5VmH0dmNuIeEEA9+AHGAgkGkyEK3Rf0zkaT+HWW1X0QOt50uq9G+ndUAGFZp++Y= X-Gm-Gg: AfdE7ckEub50EzoTgRaazx4+1BJ+w63Dh5zs44g5JVlscloutH+0njBN3xR9y8xybOI wfd+aqoqOZU//Y2rFKLAtrUCdTRRZwG5oHrFAP1kuqelcRlR8sllEjUHRtryEVuHJsGT/haVi+w S+4rhkY6XFscw0E7ecCe5mY6B2D+ekCFLx+h91K2SOZwzHS8I/kuwISjoD2HCAiYHesI+FIhDZK S8tt+PUcRpDZ2+YWiB9jXplGno8JepXn+pXdYeVHQd2EupG1I5i4NNSERRYxyRnwaEI0irI2Ahh hmENa1o0Kw8Lam+1bFFTxWr38a55dKWyFMiquZX1a4exrmdhEPYDpoyIvCCDa1YrRfTEf69+e5n q2G153cPzDkV6rwL6ONOoeIUWEdD7l52REEgs9gxO9QI1PRyFbFsSZzzeOL4uWD6mxpNLZr3uuA HZCz+kjdZdnv52LYRm7tQjsmy68SdcXzzYI9A83MyXT6fCDtCESwlX1D6kSmwZLg== X-Received: by 2002:a05:6830:4182:b0:7e6:c9eb:535a with SMTP id 46e09a7af769-7ec4a768626mr991961a34.6.1784034443908; Tue, 14 Jul 2026 06:07:23 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([178.93.176.7]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7ebcab8efc3sm14657738a34.0.2026.07.14.06.07.13 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 14 Jul 2026 06:07:23 -0700 (PDT) From: Zhanpeng Zhang To: joro@8bytes.org, palmer@dabbelt.com, tony.luck@intel.com, reinette.chatre@intel.com, tomasz.jeznach@linux.dev Cc: will@kernel.org, robin.murphy@arm.com, fustini@kernel.org, pjw@kernel.org, aou@eecs.berkeley.edu, alex@ghiti.fr, Dave.Martin@arm.com, james.morse@arm.com, babu.moger@amd.com, corbet@lwn.net, shuah@kernel.org, jgg@ziepe.ca, kevin.tian@intel.com, cuiyunhui@bytedance.com, yuanzhu@bytedance.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, x86@kernel.org Subject: [RFC PATCH 1/7] iommu: Add group lookup by ID Date: Tue, 14 Jul 2026 21:06:51 +0800 Message-ID: <20260714130657.46963-2-zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> References: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add iommu_group_get_by_id() so callers can resolve an IOMMU group from the numeric ID used in /sys/kernel/iommu_groups. An ID lookup must keep the group object alive without also keeping an otherwise empty group active. Embed the devices kobject in struct iommu_group so its address remains valid until the parent group is released, and return a reference on the parent kobject to ID lookup callers. Add iommu_group_put_by_id() to release that reference and iommu_group_is_active() to detect when the devices kobject has become inactive. Serialize lookup against group teardown with iommu_group_kset_mutex and only return groups whose devices kobject still has a live reference. This prevents a concurrent lookup from dereferencing a stale child kobject while allowing external users to discard bindings to empty groups. Signed-off-by: Zhanpeng Zhang --- drivers/iommu/iommu.c | 107 +++++++++++++++++++++++++++++++++++++----- include/linux/iommu.h | 17 +++++++ 2 files changed, 113 insertions(+), 11 deletions(-) diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c index e8f13dcebbde..da269d10f6bf 100644 --- a/drivers/iommu/iommu.c +++ b/drivers/iommu/iommu.c @@ -40,6 +40,7 @@ #include "iommu-priv.h" =20 static struct kset *iommu_group_kset; +static DEFINE_MUTEX(iommu_group_kset_mutex); static DEFINE_IDA(iommu_group_ida); static DEFINE_IDA(iommu_global_pasid_ida); =20 @@ -52,7 +53,8 @@ enum { IOMMU_PASID_ARRAY_DOMAIN =3D 0, IOMMU_PASID_ARRAY_= HANDLE =3D 1 }; =20 struct iommu_group { struct kobject kobj; - struct kobject *devices_kobj; + /* Embedded so it remains addressable until the parent group is released.= */ + struct kobject devices_kobj; struct list_head devices; struct xarray pasid_array; struct mutex mutex; @@ -729,7 +731,7 @@ static void __iommu_group_free_device(struct iommu_grou= p *group, { struct device *dev =3D grp_dev->dev; =20 - sysfs_remove_link(group->devices_kobj, grp_dev->name); + sysfs_remove_link(&group->devices_kobj, grp_dev->name); sysfs_remove_link(&dev->kobj, "iommu_group"); =20 trace_remove_device_from_group(group->id, dev); @@ -1058,6 +1060,14 @@ static const struct kobj_type iommu_group_ktype =3D { .release =3D iommu_group_release, }; =20 +static void iommu_group_devices_release(struct kobject *kobj) +{ +} + +static const struct kobj_type iommu_group_devices_ktype =3D { + .release =3D iommu_group_devices_release, +}; + /** * iommu_group_alloc - Allocate a new group * @@ -1091,17 +1101,22 @@ struct iommu_group *iommu_group_alloc(void) } group->id =3D ret; =20 + mutex_lock(&iommu_group_kset_mutex); ret =3D kobject_init_and_add(&group->kobj, &iommu_group_ktype, NULL, "%d", group->id); if (ret) { kobject_put(&group->kobj); + mutex_unlock(&iommu_group_kset_mutex); return ERR_PTR(ret); } =20 - group->devices_kobj =3D kobject_create_and_add("devices", &group->kobj); - if (!group->devices_kobj) { + kobject_init(&group->devices_kobj, &iommu_group_devices_ktype); + ret =3D kobject_add(&group->devices_kobj, &group->kobj, "devices"); + if (ret) { + kobject_put(&group->devices_kobj); kobject_put(&group->kobj); /* triggers .release & free */ - return ERR_PTR(-ENOMEM); + mutex_unlock(&iommu_group_kset_mutex); + return ERR_PTR(ret); } =20 /* @@ -1114,15 +1129,18 @@ struct iommu_group *iommu_group_alloc(void) ret =3D iommu_group_create_file(group, &iommu_group_attr_reserved_regions); if (ret) { - kobject_put(group->devices_kobj); + kobject_put(&group->devices_kobj); + mutex_unlock(&iommu_group_kset_mutex); return ERR_PTR(ret); } =20 ret =3D iommu_group_create_file(group, &iommu_group_attr_type); if (ret) { - kobject_put(group->devices_kobj); + kobject_put(&group->devices_kobj); + mutex_unlock(&iommu_group_kset_mutex); return ERR_PTR(ret); } + mutex_unlock(&iommu_group_kset_mutex); =20 pr_debug("Allocated group %d\n", group->id); =20 @@ -1286,7 +1304,7 @@ static struct group_device *iommu_group_alloc_device(= struct iommu_group *group, goto err_remove_link; } =20 - ret =3D sysfs_create_link_nowarn(group->devices_kobj, + ret =3D sysfs_create_link_nowarn(&group->devices_kobj, &dev->kobj, device->name); if (ret) { if (ret =3D=3D -EEXIST && i >=3D 0) { @@ -1431,7 +1449,7 @@ struct iommu_group *iommu_group_get(struct device *de= v) struct iommu_group *group =3D dev->iommu_group; =20 if (group) - kobject_get(group->devices_kobj); + kobject_get(&group->devices_kobj); =20 return group; } @@ -1446,7 +1464,7 @@ EXPORT_SYMBOL_GPL(iommu_group_get); */ struct iommu_group *iommu_group_ref_get(struct iommu_group *group) { - kobject_get(group->devices_kobj); + kobject_get(&group->devices_kobj); return group; } EXPORT_SYMBOL_GPL(iommu_group_ref_get); @@ -1461,10 +1479,77 @@ EXPORT_SYMBOL_GPL(iommu_group_ref_get); void iommu_group_put(struct iommu_group *group) { if (group) - kobject_put(group->devices_kobj); + kobject_put(&group->devices_kobj); } EXPORT_SYMBOL_GPL(iommu_group_put); =20 +/** + * iommu_group_get_by_id - Lookup an IOMMU group by its sysfs ID + * @id: group ID matching /sys/kernel/iommu_groups/ + * + * Return a group with a reference on its parent kobject, or NULL if no ac= tive + * group exists for @id. The caller must release the returned group with + * iommu_group_put_by_id(). Keeping this reference does not keep an empty + * group's devices kobject active. + */ +struct iommu_group *iommu_group_get_by_id(int id) +{ + struct kobject *group_kobj; + struct iommu_group *group =3D NULL; + char name[12]; + + if (!iommu_group_kset || id < 0) + return NULL; + + snprintf(name, sizeof(name), "%d", id); + mutex_lock(&iommu_group_kset_mutex); + group_kobj =3D kset_find_obj(iommu_group_kset, name); + if (!group_kobj) + goto unlock; + + group =3D container_of(group_kobj, struct iommu_group, kobj); + if (!kobject_get_unless_zero(&group->devices_kobj)) { + kobject_put(group_kobj); + group =3D NULL; + goto unlock; + } + + kobject_put(&group->devices_kobj); +unlock: + mutex_unlock(&iommu_group_kset_mutex); + return group; +} +EXPORT_SYMBOL_GPL(iommu_group_get_by_id); + +/** + * iommu_group_put_by_id - Release a group returned by ID lookup + * @group: group returned by iommu_group_get_by_id() + */ +void iommu_group_put_by_id(struct iommu_group *group) +{ + if (group) + kobject_put(&group->kobj); +} +EXPORT_SYMBOL_GPL(iommu_group_put_by_id); + +/** + * iommu_group_is_active - Test whether a referenced group can accept devi= ces + * @group: referenced IOMMU group + * + * Return true while the devices kobject still has a live reference. Once = the + * group loses its last device and external device reference, it cannot be= come + * active again. + */ +bool iommu_group_is_active(struct iommu_group *group) +{ + if (!group || !kobject_get_unless_zero(&group->devices_kobj)) + return false; + + kobject_put(&group->devices_kobj); + return true; +} +EXPORT_SYMBOL_GPL(iommu_group_is_active); + /** * iommu_group_id - Return ID for a group * @group: the group to ID diff --git a/include/linux/iommu.h b/include/linux/iommu.h index d20aa6f6863a..e771b4a92f5b 100644 --- a/include/linux/iommu.h +++ b/include/linux/iommu.h @@ -989,6 +989,9 @@ extern void iommu_group_remove_device(struct device *de= v); extern int iommu_group_for_each_dev(struct iommu_group *group, void *data, int (*fn)(struct device *, void *)); extern struct iommu_group *iommu_group_get(struct device *dev); +struct iommu_group *iommu_group_get_by_id(int id); +void iommu_group_put_by_id(struct iommu_group *group); +bool iommu_group_is_active(struct iommu_group *group); extern struct iommu_group *iommu_group_ref_get(struct iommu_group *group); extern void iommu_group_put(struct iommu_group *group); =20 @@ -1401,6 +1404,20 @@ static inline struct iommu_group *iommu_group_get(st= ruct device *dev) return NULL; } =20 +static inline struct iommu_group *iommu_group_get_by_id(int id) +{ + return NULL; +} + +static inline void iommu_group_put_by_id(struct iommu_group *group) +{ +} + +static inline bool iommu_group_is_active(struct iommu_group *group) +{ + return false; +} + static inline void iommu_group_put(struct iommu_group *group) { } --=20 2.50.1 (Apple Git-155) From nobody Sat Jul 25 19:29:10 2026 Received: from mail-ot1-f43.google.com (mail-ot1-f43.google.com [209.85.210.43]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9CE7B46AF03 for ; Tue, 14 Jul 2026 13:07:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.43 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034458; cv=none; b=eiP3/UJCN/B4+furmrtFrlf1oVbH1X9pMesaLuIHRuvBw385oF0L4qS5vThYXdjdBu0dAUAjH4EwQY+OR87wayKUNbnkZVvT80boTCoENYxALfSbYhZyzNTGibp33LJGwtuKqd3QhnMU5RwlzJ2OkzIAsnojgaJ/0VBpODxVb7w= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034458; c=relaxed/simple; bh=zWI3HgS6VNdfgTHEihXvnUhJkpLl5h08ZrZ1Oj9J3pg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=AdzpBomYOTCMs0iAkLCQCJudfiS+JRqahIKLVCh++bVfjHuEleTxh/ZAqWIO4fPJLy6I8xizh9mzb+SdULGVpB2vN/KshsRAoA+ynfuZvelKlNu6IO4VuhQcM/AS6MnoSbO1AFwHi3JVX29teuXqF+bXahBDkUY5y7MVlJEMgR8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=CdQkj7xR; arc=none smtp.client-ip=209.85.210.43 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="CdQkj7xR" Received: by mail-ot1-f43.google.com with SMTP id 46e09a7af769-7e9ecd7216cso502862a34.3 for ; Tue, 14 Jul 2026 06:07:36 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1784034455; x=1784639255; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PgO28qqQ1xeK9Yvst4+iTUEqQXClh8oO6Jm7DHD/RdA=; b=CdQkj7xR+h9LHTxhEbV6qfSS3zfzryn9x5BR/nNpAEZx5yMWJzohFhnZ/CUN3Bw7P9 3Xlp0QADCcYAyANDygVONbE0PyV83y5r0jQppN1N+h0yHUNdud4dcK1hWEGal0qQSmKI IPnN6kNoWn978PIZvutNMwI+vO82gruj/jf2TwTbizC96PY+HPom7qKYDQEc97AZiNIs jcT+9dKzI5XHrqcnNhAW27+cU2HWOf4av+ClxhduDhSDPOF/yJmfV55c1G3jaYEgEbqF YoXvwWzE6VeIhloI8BtLX5IJstirWeSt74rPhtVls7A/Jv8pke37qZUPdDc2XhMvD/gB WzJg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784034455; x=1784639255; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=PgO28qqQ1xeK9Yvst4+iTUEqQXClh8oO6Jm7DHD/RdA=; b=m0loXGPK9dOFRwHTuvapwjVy3rSEvwaHcvoFfUvk49Vm8BFpvuMTrNyZ6mDivuEduZ 2mj6UjkU2cSV6Vlnuuim2mGvrLhIGv0EFso7ThbSV73IByqdbqx9I0GW11j4+QMlX8vd 04FHE3VFuAqi35TyrkXvHLdQz1PwdRVMIjlT1IyylRGOkZADYPvAtr5Re8dwtljTFxYV UVUzyvf3GTJsfcgeytFqdA266LqsuxsiQRl723mbQvJeGKstuJZyz2aDDzqLU/9CIRtU CdHmGinU3XwZDS3AvJk3dZZKeY6ypxUZjOzXo3t5CkGNTbNPwdAImvwIWPbsZE8FotXD po4A== X-Forwarded-Encrypted: i=1; AFNElJ/Qq4o7Nbdonni7hipb3QHEx3W2mQi9NIADD/zX8G/SiUX3SzeY5TZbVHISHEDLA8dM/9sc0PAof87jsxM=@vger.kernel.org X-Gm-Message-State: AOJu0Yz76OLKnilYe7o40Y2A2kSem+HRoeuow1dJ6yNo2gu2iu+HQrK1 cBEc1dfMkDIop0NXN2raU0l44oWWXTWP4ZcARjjoy/bsZkQ0zopS7e+VlC659HK4tJg= X-Gm-Gg: AfdE7cnZ5Eyl0Jun3fQA8f7gpwVP99zDlF0boUKi69O5eZ1vACeZHNx5/9NibQlk4vh 6P3AotXfLpjgdm8k4uNXck3Hy+hZWNJVKmaVkqTLAih+hfFE5I7TH5qmTA1PoXelkd9kzlAHZuU UTYRpzm6qapBt4AFnvOcgIHh+ru6fXbKDCgRZJHiL5DQ95EZEmBdLSTSIcmGUG88qRJUmrL4c2Q wOdoiH3Nbt8KtAvhPAT/N30j4enL9h79BvoMA3No2pt/PskiNYT/ImHmLCDyn4yyo8ANcf+fqi1 2rv1KEnSc6PNGGomqCGPDUL1SGt13TPpf2kqhVKvYLQCGWHYfywyhn+k7jFuCXGtiPiqyf9lrNX MeWqRUdlq6hbY+VrjBSDvi4tQ0zjYISWz4TWbbXIxU2D5aqce3XyPWTQYAYS8t621J8+sdEU4ZI DbCXdvQrG6ipYryQ9OxATJwOmREaJRaVjcWCN3dsC/gTAR/bYA9Qb7NhQUoh4wXg== X-Received: by 2002:a05:6830:f88:b0:7e9:fe5f:4406 with SMTP id 46e09a7af769-7ec0997528bmr8084846a34.30.1784034455345; Tue, 14 Jul 2026 06:07:35 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([178.93.176.7]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7ebcab8efc3sm14657738a34.0.2026.07.14.06.07.24 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 14 Jul 2026 06:07:34 -0700 (PDT) From: Zhanpeng Zhang To: joro@8bytes.org, palmer@dabbelt.com, tony.luck@intel.com, reinette.chatre@intel.com, tomasz.jeznach@linux.dev Cc: will@kernel.org, robin.murphy@arm.com, fustini@kernel.org, pjw@kernel.org, aou@eecs.berkeley.edu, alex@ghiti.fr, Dave.Martin@arm.com, james.morse@arm.com, babu.moger@amd.com, corbet@lwn.net, shuah@kernel.org, jgg@ziepe.ca, kevin.tian@intel.com, cuiyunhui@bytedance.com, yuanzhu@bytedance.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, x86@kernel.org Subject: [RFC PATCH 2/7] iommu: Add checked group device update helper Date: Tue, 14 Jul 2026 21:06:52 +0800 Message-ID: <20260714130657.46963-3-zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> References: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Some group-wide operations must validate every member before changing any device. Separate iommu_group_for_each_dev() calls cannot provide that guarantee because group membership may change between traversals. Add iommu_group_update_devices() to keep the group membership mutex held across a validation pass and a non-failing update pass. This provides all-or-none validation without exposing IOMMU group internals to callers. Signed-off-by: Zhanpeng Zhang --- drivers/iommu/iommu.c | 33 +++++++++++++++++++++++++++++++++ include/linux/iommu.h | 13 +++++++++++++ 2 files changed, 46 insertions(+) diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c index da269d10f6bf..9a6c4a7e7df6 100644 --- a/drivers/iommu/iommu.c +++ b/drivers/iommu/iommu.c @@ -1436,6 +1436,39 @@ int iommu_group_for_each_dev(struct iommu_group *gro= up, void *data, } EXPORT_SYMBOL_GPL(iommu_group_for_each_dev); =20 +/** + * iommu_group_update_devices - Check and update every device in a group + * @group: the group + * @data: caller data passed to both callbacks + * @check: validates whether one device can be updated + * @update: updates one device after every check has succeeded + * + * Keep group membership stable while first checking every device and then + * applying an update which cannot fail. No device is updated if a check f= ails. + */ +int iommu_group_update_devices(struct iommu_group *group, void *data, + int (*check)(struct device *, void *), + void (*update)(struct device *, void *)) +{ + struct group_device *device; + int ret =3D 0; + + mutex_lock(&group->mutex); + for_each_group_device(group, device) { + ret =3D check(device->dev, data); + if (ret) + goto unlock; + } + + for_each_group_device(group, device) + update(device->dev, data); + +unlock: + mutex_unlock(&group->mutex); + return ret; +} +EXPORT_SYMBOL_GPL(iommu_group_update_devices); + /** * iommu_group_get - Return the group for a device and increment reference * @dev: get the group that this device belongs to diff --git a/include/linux/iommu.h b/include/linux/iommu.h index e771b4a92f5b..befba0683e06 100644 --- a/include/linux/iommu.h +++ b/include/linux/iommu.h @@ -988,6 +988,9 @@ extern int iommu_group_add_device(struct iommu_group *g= roup, extern void iommu_group_remove_device(struct device *dev); extern int iommu_group_for_each_dev(struct iommu_group *group, void *data, int (*fn)(struct device *, void *)); +int iommu_group_update_devices(struct iommu_group *group, void *data, + int (*check)(struct device *, void *), + void (*update)(struct device *, void *)); extern struct iommu_group *iommu_group_get(struct device *dev); struct iommu_group *iommu_group_get_by_id(int id); void iommu_group_put_by_id(struct iommu_group *group); @@ -1399,6 +1402,16 @@ static inline int iommu_group_for_each_dev(struct io= mmu_group *group, return -ENODEV; } =20 +static inline int iommu_group_update_devices(struct iommu_group *group, + void *data, + int (*check)(struct device *, + void *), + void (*update)(struct device *, + void *)) +{ + return -ENODEV; +} + static inline struct iommu_group *iommu_group_get(struct device *dev) { return NULL; --=20 2.50.1 (Apple Git-155) From nobody Sat Jul 25 19:29:10 2026 Received: from mail-ot1-f47.google.com (mail-ot1-f47.google.com [209.85.210.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4F3FC47276D for ; Tue, 14 Jul 2026 13:07:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.47 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034470; cv=none; b=rw/v4Ip4AmnZZ6ztzcHS4vj5+s3NSLUrmabkx+ZRkbs1mK+BU587GRLWNPt51LDDAvX8rAPdioSrH6JDalHkS/unR/8lTmPAK2+CwXUuPxgUDYzjIb9JYprTIqgVY1nctLCZeshgyyal/56smQvP330f+u5Bc4Z15SNxlW2a3jk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034470; c=relaxed/simple; bh=lv5kWIUDJUDMWlaCosjBEeMFcK6y35rawRF3K3p6LiM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ciz9QoUM51ilY74ASNJt02E1xNANauhuYMK7/vLWOjlfesrFzHvyE6+5mC2VhVg5eJIdnELeu+WYsh43jIGingV9xDMof5k29If++GfDu81rbPVeKMdZ7LtEuqof0aGeicsrb/XMUbxo+AvsjTAv+VH9AO+AAlMdojSawhr2UoE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=U45udnMq; arc=none smtp.client-ip=209.85.210.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="U45udnMq" Received: by mail-ot1-f47.google.com with SMTP id 46e09a7af769-7e9f829d75aso610991a34.0 for ; Tue, 14 Jul 2026 06:07:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1784034467; x=1784639267; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=kK72RdH/1XE7usLQ2kyIcth7swn3eieDHyIRvr+BBH8=; b=U45udnMqaRZQYTEY2xcBi0M4W2vUhMTmAeF6XPlvn5POmzhX0S/qIAJm4JpRAYuGJZ Xza0KcZZb6L11iYH80cfE1EL61aN1GTIJzhJEUAlGvH7C9chaIMK7we+5yElY1Vw0S0A vNA60ayZANYgV7fOe4iCTfXrCL5tv3rGoOY2bLP/SoUVxoG8FT23+P8fTGmWBynlNFu9 b9vn3gTaXL6VI+5i7PgEB01YjYPKqfZk2SvR90qNulndqSgPgKstfQAWMfNxGjTY+KjL abAYHjPeb+t/xwLkuijg5mk7LF93U41yrHlBsjdIfc+sjyZ07YJ+AH3EDPB1R6Oj/wIo m9Ug== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784034467; x=1784639267; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=kK72RdH/1XE7usLQ2kyIcth7swn3eieDHyIRvr+BBH8=; b=ejpWSwh8QF4kPB3wlty+I75H6gx9yDznO2bO2Tn2UcecPvfMI9oETs5F1EtL9ZFs1+ nKQy6X95+56pLFVm5ILeurt8Sk8cDKOw4nVnsoL5sqnV0zgTA8LjlQ5aEMN85yl3WTCD PAh2m7WHiC0jF3uNXoDmFWNG5B++hvfAFsONp6G07NLR8CV3AuRTtnBqAjTeI/anQ8c8 2iJWWBJ4nXSWQn7HSYvCnTt6vP9o1Rg0vRGhK2m7anlF3amnSYepaRW9UgsoGrIXKgZa 4uAsOu/JkyKNBOnsv/tCGZZRcZy2u4uE1sAMjUCNuw2Js5aqMNmgb+bOAb5N7iN0Q740 e4CQ== X-Forwarded-Encrypted: i=1; AFNElJ9LSFgTisulN35yx+Cv5srjGx9NHOjGOdssE04MukoveRd26qYCyN1do6DSVoS5C7wh9U7SeLlow44pyDA=@vger.kernel.org X-Gm-Message-State: AOJu0Yyc0XzdO4xRoI+coGhVzjgbkm/nKzxXuAZ+d/pV4581BuxbKESc VELRfON/ISZNdkc8Eh3KFnzNBzsbOOXwnPIRujctTQkRGaLwMtz+FmlD5tdO7h+w1cw= X-Gm-Gg: AfdE7clRcOqQneHlLGzz68SnhGAm/2OE0uoS6w8awPyvaW7XX6KJMdXXiuG9ci44l1t p47DHtnsLCJcr72psvmR1284Be0tFco2MGqZPOYuzHhJpJdXGMMoBkr7jiYaChH4IimqGcGeyv9 ciJBKkjF5gzOssz1GdKJXqqm2JpFvKq6O7wm8g56Gtf52LY/IrKKbfXIgnV46Jp8tJQ2i393e+5 LoVC85TjfeI6/j8kFV5eZkZFgWt5lK35cyWaRWxit9KyNQZ8l4uAxXRODV+wbDo/zDnYON5zxZ8 u85iedfcjdLwP/pNyKT5qPt5XfzxZT5zlO5JsXOVxpMD9MJ1S3A0O+DPRT/2y4HO67rrC8vCWrq 7Tng4JUEPAtWB5EysrWIsPJ7zXcQcuy3y5vfB3HyKURlZejJ2C+0V+9Hi37q5gCm1lORmOTMP0P 2k7mHiswv++meN13bNGQIlmKt5GnJkOWgPAPWmLl9IeQvGes6NnTYy5xRV0YctbA== X-Received: by 2002:a05:6830:3492:b0:7e9:bd00:c6ad with SMTP id 46e09a7af769-7ec097e11c2mr8750217a34.16.1784034466935; Tue, 14 Jul 2026 06:07:46 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([178.93.176.7]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7ebcab8efc3sm14657738a34.0.2026.07.14.06.07.35 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 14 Jul 2026 06:07:46 -0700 (PDT) From: Zhanpeng Zhang To: joro@8bytes.org, palmer@dabbelt.com, tony.luck@intel.com, reinette.chatre@intel.com, tomasz.jeznach@linux.dev Cc: will@kernel.org, robin.murphy@arm.com, fustini@kernel.org, pjw@kernel.org, aou@eecs.berkeley.edu, alex@ghiti.fr, Dave.Martin@arm.com, james.morse@arm.com, babu.moger@amd.com, corbet@lwn.net, shuah@kernel.org, jgg@ziepe.ca, kevin.tian@intel.com, cuiyunhui@bytedance.com, yuanzhu@bytedance.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, x86@kernel.org Subject: [RFC PATCH 3/7] resctrl: Add a devices file for external requester assignment Date: Tue, 14 Jul 2026 21:06:53 +0800 Message-ID: <20260714130657.46963-4-zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> References: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Do not overload the resctrl tasks file with architecture-specific non-PID tokens. The tasks ABI remains a list of task IDs, while the new devices file carries external requesters assigned to a resctrl group. Add architecture hooks for assigning and showing those external objects, and reject group removal, reparenting, or pseudo-lock setup while devices are still assigned. The teardown path performs a best-effort reset to the default group. The devices file is an assignment interface only. It does not describe a new resource schema or domain; resource allocation and monitoring policy remain described by the existing schemata and info files. Signed-off-by: Zhanpeng Zhang --- Documentation/filesystems/resctrl.rst | 26 ++++ arch/Kconfig | 6 + fs/resctrl/rdtgroup.c | 206 +++++++++++++++++++++++++- include/linux/resctrl.h | 45 ++++++ 4 files changed, 280 insertions(+), 3 deletions(-) diff --git a/Documentation/filesystems/resctrl.rst b/Documentation/filesyst= ems/resctrl.rst index e4b66af55ffb..fce61d019114 100644 --- a/Documentation/filesystems/resctrl.rst +++ b/Documentation/filesystems/resctrl.rst @@ -581,6 +581,32 @@ All groups contain the following files: idle tasks. Instead, a CPU's idle task is always considered as a member of the group owning the CPU. =20 +"devices": + On architectures that support external requester assignment through + resctrl, reading this file shows the devices or device groups assigned + to this resource group. Writing an architecture-specific device token + moves that external requester to the group. Multiple tokens can be + separated by commas and are processed sequentially. A failure aborts the + write, but requesters moved before the failure remain in their new groups. + + On RISC-V, an IOMMU group is identified by the following token:: + + iommu_group: + + Each assigned IOMMU group is reported on a separate line when the file is + read. Writing the token to the root control group's devices file restores + the IOMMU group to the reserved default QoS IDs. Default assignments are + not listed in the root devices file. + + This file only controls external requester membership. Resource + allocation and monitoring policy remains described by schemata and + info files. + + Resource groups with assigned devices cannot be removed or reparented. + Move devices to another group first. + + Failures will be logged to /sys/fs/resctrl/info/last_cmd_status. + "cpus": Reading this file shows a bitmask of the logical CPUs owned by this group. Writing a mask to this file will add and remove diff --git a/arch/Kconfig b/arch/Kconfig index fa7507ac8e13..36a4f4fbb164 100644 --- a/arch/Kconfig +++ b/arch/Kconfig @@ -1615,6 +1615,12 @@ config ARCH_HAS_CPU_RESCTRL monitoring and control interfaces provided by the 'resctrl' filesystem (see RESCTRL_FS). =20 +config ARCH_HAS_RESCTRL_DEVICES + bool + help + An architecture selects this option to indicate that external + device objects can be attached to resctrl resource groups. + config HAVE_ARCH_COMPILER_H bool help diff --git a/fs/resctrl/rdtgroup.c b/fs/resctrl/rdtgroup.c index af2cbab14497..2e424c911049 100644 --- a/fs/resctrl/rdtgroup.c +++ b/fs/resctrl/rdtgroup.c @@ -99,13 +99,32 @@ void rdt_last_cmd_puts(const char *s) seq_buf_puts(&last_cmd_status, s); } =20 +void resctrl_last_cmd_puts(const char *s) +{ + rdt_last_cmd_puts(s); +} + +static void rdt_last_cmd_vprintf(const char *fmt, va_list ap) +{ + lockdep_assert_held(&rdtgroup_mutex); + seq_buf_vprintf(&last_cmd_status, fmt, ap); +} + void rdt_last_cmd_printf(const char *fmt, ...) { va_list ap; =20 va_start(ap, fmt); - lockdep_assert_held(&rdtgroup_mutex); - seq_buf_vprintf(&last_cmd_status, fmt, ap); + rdt_last_cmd_vprintf(fmt, ap); + va_end(ap); +} + +void resctrl_last_cmd_printf(const char *fmt, ...) +{ + va_list ap; + + va_start(ap, fmt); + rdt_last_cmd_vprintf(fmt, ap); va_end(ap); } =20 @@ -766,6 +785,151 @@ static int rdtgroup_move_task(pid_t pid, struct rdtgr= oup *rdtgrp, return ret; } =20 +static bool rdtgroup_effective_ids(struct rdtgroup *r, + struct resctrl_group_ids *ids) +{ + if (!r || !ids) + return false; + + if (r->type =3D=3D RDTMON_GROUP) + ids->closid =3D r->mon.parent->closid; + else if (r->type =3D=3D RDTCTRL_GROUP) + ids->closid =3D r->closid; + else + return false; + + ids->rmid =3D r->mon.rmid; + return true; +} + +#ifdef CONFIG_ARCH_HAS_RESCTRL_DEVICES +static void show_rdt_devices(struct rdtgroup *r, struct seq_file *s) +{ + struct resctrl_group_ids ids; + + if (rdtgroup_effective_ids(r, &ids)) + resctrl_arch_devices_show(s, ids); +} + +static int rdtgroup_reset_all_devices(struct rdtgroup *to) +{ + struct resctrl_group_ids default_ids; + + if (!to || !rdtgroup_effective_ids(to, &default_ids)) + return 0; + + return resctrl_arch_devices_reset_all(default_ids); +} + +static bool rdtgroup_arch_devices_assigned(struct rdtgroup *rdtgrp) +{ + struct resctrl_group_ids ids; + + if (!rdtgroup_effective_ids(rdtgrp, &ids)) + return false; + + return resctrl_arch_devices_assigned(ids); +} + +static bool rdtgroup_child_arch_devices_assigned(struct rdtgroup *rdtgrp) +{ + struct rdtgroup *crgrp; + + list_for_each_entry(crgrp, &rdtgrp->mon.crdtgrp_list, mon.crdtgrp_list) { + if (rdtgroup_arch_devices_assigned(crgrp)) + return true; + } + + return false; +} + +static int rdtgroup_reject_assigned_devices(struct rdtgroup *rdtgrp, + bool include_children, + const char *operation) +{ + if (!rdtgroup_arch_devices_assigned(rdtgrp) && + (!include_children || !rdtgroup_child_arch_devices_assigned(rdtgrp))) + return 0; + + rdt_last_cmd_printf("Move devices out before %s group\n", operation); + return -EBUSY; +} + +static ssize_t rdtgroup_devices_write(struct kernfs_open_file *of, + char *buf, size_t nbytes, loff_t off) +{ + struct resctrl_group_ids ids; + struct rdtgroup *rdtgrp; + char *tok; + int ret =3D 0; + + rdtgrp =3D rdtgroup_kn_lock_live(of->kn); + if (!rdtgrp) { + rdtgroup_kn_unlock(of->kn); + return -ENOENT; + } + rdt_last_cmd_clear(); + + if (rdtgrp->mode =3D=3D RDT_MODE_PSEUDO_LOCKED || + rdtgrp->mode =3D=3D RDT_MODE_PSEUDO_LOCKSETUP) { + ret =3D -EINVAL; + rdt_last_cmd_puts("Pseudo-locking in progress\n"); + goto unlock; + } + + if (!rdtgroup_effective_ids(rdtgrp, &ids)) { + ret =3D -EINVAL; + goto unlock; + } + + while ((tok =3D strsep(&buf, ","))) { + tok =3D strim(tok); + if (!*tok) { + rdt_last_cmd_puts("Device list parsing error\n"); + ret =3D -EINVAL; + break; + } + + ret =3D resctrl_arch_devices_write(tok, ids); + if (ret) + break; + } + +unlock: + rdtgroup_kn_unlock(of->kn); + + return ret ?: nbytes; +} + +static int rdtgroup_devices_show(struct kernfs_open_file *of, + struct seq_file *s, void *v) +{ + struct rdtgroup *rdtgrp; + int ret =3D 0; + + rdtgrp =3D rdtgroup_kn_lock_live(of->kn); + if (rdtgrp) + show_rdt_devices(rdtgrp, s); + else + ret =3D -ENOENT; + rdtgroup_kn_unlock(of->kn); + + return ret; +} +#else +static inline int rdtgroup_reset_all_devices(struct rdtgroup *to) +{ + return 0; +} + +static inline int rdtgroup_reject_assigned_devices(struct rdtgroup *rdtgrp, + bool include_children, + const char *operation) +{ + return 0; +} +#endif + static ssize_t rdtgroup_tasks_write(struct kernfs_open_file *of, char *buf, size_t nbytes, loff_t off) { @@ -1491,6 +1655,11 @@ static ssize_t rdtgroup_mode_write(struct kernfs_ope= n_file *of, rdtgrp->mode =3D RDT_MODE_EXCLUSIVE; } else if (IS_ENABLED(CONFIG_RESCTRL_FS_PSEUDO_LOCK) && !strcmp(buf, "pseudo-locksetup")) { + ret =3D rdtgroup_reject_assigned_devices(rdtgrp, true, + "entering pseudo-locksetup"); + if (ret) + goto out; + ret =3D rdtgroup_locksetup_enter(rdtgrp); if (ret) goto out; @@ -2067,6 +2236,16 @@ static struct rftype res_common_files[] =3D { .seq_show =3D rdtgroup_tasks_show, .fflags =3D RFTYPE_BASE, }, +#ifdef CONFIG_ARCH_HAS_RESCTRL_DEVICES + { + .name =3D "devices", + .mode =3D 0644, + .kf_ops =3D &rdtgroup_kf_single_ops, + .write =3D rdtgroup_devices_write, + .seq_show =3D rdtgroup_devices_show, + .fflags =3D RFTYPE_BASE, + }, +#endif { .name =3D "mon_hw_id", .mode =3D 0444, @@ -3061,6 +3240,8 @@ static void rmdir_all_sub(void) /* Move all tasks to the default resource group */ rdt_move_group_tasks(NULL, &rdtgroup_default, NULL); =20 + WARN_ON_ONCE(rdtgroup_reset_all_devices(&rdtgroup_default)); + list_for_each_entry_safe(rdtgrp, tmp, &rdt_all_groups, rdtgroup_list) { /* Free any child rmids */ free_all_child_rdtgrp(rdtgrp); @@ -3970,6 +4151,11 @@ static int rdtgroup_rmdir_mon(struct rdtgroup *rdtgr= p, cpumask_var_t tmpmask) struct rdtgroup *prdtgrp =3D rdtgrp->mon.parent; u32 closid, rmid; int cpu; + int ret; + + ret =3D rdtgroup_reject_assigned_devices(rdtgrp, false, "removing"); + if (ret) + return ret; =20 /* Give any tasks back to the parent group */ rdt_move_group_tasks(rdtgrp, prdtgrp, tmpmask); @@ -4020,6 +4206,11 @@ static int rdtgroup_rmdir_ctrl(struct rdtgroup *rdtg= rp, cpumask_var_t tmpmask) { u32 closid, rmid; int cpu; + int ret; + + ret =3D rdtgroup_reject_assigned_devices(rdtgrp, true, "removing"); + if (ret) + return ret; =20 /* Give any tasks back to the default group */ rdt_move_group_tasks(rdtgrp, &rdtgroup_default, tmpmask); @@ -4093,7 +4284,10 @@ static int rdtgroup_rmdir(struct kernfs_node *kn) rdtgrp !=3D &rdtgroup_default) { if (rdtgrp->mode =3D=3D RDT_MODE_PSEUDO_LOCKSETUP || rdtgrp->mode =3D=3D RDT_MODE_PSEUDO_LOCKED) { - ret =3D rdtgroup_ctrl_remove(rdtgrp); + ret =3D rdtgroup_reject_assigned_devices(rdtgrp, true, + "removing"); + if (!ret) + ret =3D rdtgroup_ctrl_remove(rdtgrp); } else { ret =3D rdtgroup_rmdir_ctrl(rdtgrp, tmpmask); } @@ -4210,6 +4404,12 @@ static int rdtgroup_rename(struct kernfs_node *kn, goto out; } =20 + if (rdtgrp->mon.parent !=3D new_prdtgrp) { + ret =3D rdtgroup_reject_assigned_devices(rdtgrp, false, "reparenting"); + if (ret) + goto out; + } + /* * Allocate the cpumask for use in mongrp_reparent() to avoid the * possibility of failing to allocate it after kernfs_rename() has diff --git a/include/linux/resctrl.h b/include/linux/resctrl.h index 73ff522448a0..a4b5c2d5e814 100644 --- a/include/linux/resctrl.h +++ b/include/linux/resctrl.h @@ -8,6 +8,8 @@ #include #include =20 +struct seq_file; + #ifdef CONFIG_ARCH_HAS_CPU_RESCTRL #include #endif @@ -18,6 +20,49 @@ =20 #define RESCTRL_PICK_ANY_CPU -1 =20 +/** + * struct resctrl_group_ids - resctrl control and monitoring IDs + * @closid: resource control class ID + * @rmid: resource monitoring ID + */ +struct resctrl_group_ids { + u32 closid; + u32 rmid; +}; + +void resctrl_last_cmd_puts(const char *s); +void resctrl_last_cmd_printf(const char *fmt, ...) __printf(1, 2); + +#ifdef CONFIG_ARCH_HAS_RESCTRL_DEVICES +int resctrl_arch_devices_write(char *tok, struct resctrl_group_ids ids); +void resctrl_arch_devices_show(struct seq_file *s, + struct resctrl_group_ids ids); +bool resctrl_arch_devices_assigned(struct resctrl_group_ids ids); +int resctrl_arch_devices_reset_all(struct resctrl_group_ids default_ids); +#else +static inline int resctrl_arch_devices_write(char *tok, + struct resctrl_group_ids ids) +{ + return -EOPNOTSUPP; +} + +static inline void resctrl_arch_devices_show(struct seq_file *s, + struct resctrl_group_ids ids) +{ +} + +static inline bool resctrl_arch_devices_assigned(struct resctrl_group_ids = ids) +{ + return false; +} + +static inline int +resctrl_arch_devices_reset_all(struct resctrl_group_ids default_ids) +{ + return 0; +} +#endif + #ifdef CONFIG_PROC_CPU_RESCTRL =20 int proc_resctrl_show(struct seq_file *m, --=20 2.50.1 (Apple Git-155) From nobody Sat Jul 25 19:29:10 2026 Received: from mail-ot1-f44.google.com (mail-ot1-f44.google.com [209.85.210.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 48AAE47278B for ; Tue, 14 Jul 2026 13:08:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.44 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034482; cv=none; b=YGNitfmGAthdguCggcdSlzIHWdcE5ni5DyjFGcVpcwPTvhCG1OO+3NaY/bV1CY5ikbboYKyFxOt/S8n16u7rfklXJXMnyHUvdbCPwqeBBPsK8z/p4PNnz10312dQoqY27lz4ZvJoZXKPAFXUK+PlVZW0QD6miGolcwsFNelU6PE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034482; c=relaxed/simple; bh=qsNLCYeMQllqAlZI5ZYlNbddWlxsI+AfKPo9caZQY2c=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=J0T6rZHt+m42s2rbKXfPxI03GM6G21ZripHambhyeoEOiw9/jfv/TCyJe39aUB3JndVpF0ENw11EOtt0dwtL8W6n4pPO75zIFx3VieVtHrSfqvKTwoNuGQHdah9FmRWQ4vGN1DMN2HGSM5FYdUs3uV9CY3TjxfNYekm2tvVtNm4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=M+67DoGQ; arc=none smtp.client-ip=209.85.210.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="M+67DoGQ" Received: by mail-ot1-f44.google.com with SMTP id 46e09a7af769-7eb61bbeb25so537045a34.1 for ; Tue, 14 Jul 2026 06:08:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1784034479; x=1784639279; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JpT9/sOjxA5cBGNsD1q8lvPgS5awws3sRPjnLJKoBtI=; b=M+67DoGQ2zaj+EWGWrH1WO6mgRguxmYHJ08P4bn2fRmjSNnkAamQuQ5GuCJ8fjPy5W O7y5/tQvHqABTw34AxVkpgdh+dbEN34+KRFUDBVOvCaYZ+Qus2nRQhkS1u2VnprEG7mX Qc6gl/4NRyAsFK6lrebi6p//Izukh4WM3RszfBluOtNd4VoRDNEmllXukJKUmX/wo49F 4nRG7msK3NqYgHCqpj11DAUbQ7xn3Bkw0JQ/koqh2iQ2h5b+KsHEk/qpfF5DvKidL5Yf gbH/BGWz0kERDRI0m/gwebvvkyln0qdIZtkqpqLsv8Mc4tzVOF9Fm7N4cSC2oRMUXQe0 Is9A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784034479; x=1784639279; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=JpT9/sOjxA5cBGNsD1q8lvPgS5awws3sRPjnLJKoBtI=; b=W059rQuJJwyXZt3E+DcFiOo/+fBYF9PJXbtj9cFio60tcBFtCdNrU2codVcHPn+nZn QGnmuE8/FWhV0G2oAtmq0yx4vtb4mCZLYU/1FOkbgBFkGTKp42trgRtY+WG1pJZD3r+X lTcH1JGbq/Vy7pvIZOiIPz3tOHZy9k1i1nRhmYweQO0jPsWWFFU7VwfXqYVD5WHvpcjJ jdwdyWXrpJ8OIvPYv4CYn1AMnCsjhozHDo2eZ4whVbGKxSeIJf49kZVfg+ILgypNV7QJ /JyJyvACwnIbEhDezdaHYbrd0hBL+cbjUxjdFc497wqYOamaMDeCNURI+/SOCEhGbPE4 QKXA== X-Forwarded-Encrypted: i=1; AFNElJ9Rq+U+NzNNR/YwtviSnAGReVKQw+z2saFbc2V/qgQyZDI94lCWjrNo8bdLe+73/sReG8b8We3v+Bkcp2Y=@vger.kernel.org X-Gm-Message-State: AOJu0Yx916vz+A8lsrxXN+6GoTjtsS5eXLvsyVWbeTD1eQdxmZcWoAoZ d5lWWu048FaV0SnxWd33C0mXsnfjo8dJit3rI38zxyE0rCPNZk0tuR/Cs+8U68mTMVY= X-Gm-Gg: AfdE7cnm16EbCEITgw5rr9yzd3neHXAXDWmektvKqkx/vPUc6OuJ6pcCuWFIA8bK5Xn n0a/4HCJ7PJmHfQp6aQVXI4RmihpUWOqbBUq5M3LaE3jKq6Bn1q9XEijMOxJ1p51isryOpuDvXC cXl4SvWsE53kzdhCQ0NjNH7DuAYN9Qs20jxT4/OaPDvVavjaVqURAg9Juf6kWiQyUQePstBNqKG Sz+PD+Ou50QsCMtN27PZo6NT3o8Bhe64J7ktDR1vmSeXHL2jhCme/CDVMt4DBFmyRrqxv9T/voM QLvwdssICCKXKIT+/ZN8LtqYNMvr+2SyVfMK9Vj+Y4fpdsxPkNW6ym1HCfzfBTbpBHr4Ctj5EEj aXZMosm81e6l24C53639Mqb325KfEVzel6irH9yz/2mI+UDJrFdTPKOxPTBtLBAkl8fRERcQ7xb gM1AiKauUBSBQNLxc6WJuG9S353X20Q0JGFAp/KVqG2JojjebY8QQnwc+euLrWcA== X-Received: by 2002:a05:6830:b13:b0:7e6:da40:b7fa with SMTP id 46e09a7af769-7ec0983a65bmr8871169a34.24.1784034478996; Tue, 14 Jul 2026 06:07:58 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([178.93.176.7]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7ebcab8efc3sm14657738a34.0.2026.07.14.06.07.47 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 14 Jul 2026 06:07:58 -0700 (PDT) From: Zhanpeng Zhang To: joro@8bytes.org, palmer@dabbelt.com, tony.luck@intel.com, reinette.chatre@intel.com, tomasz.jeznach@linux.dev Cc: will@kernel.org, robin.murphy@arm.com, fustini@kernel.org, pjw@kernel.org, aou@eecs.berkeley.edu, alex@ghiti.fr, Dave.Martin@arm.com, james.morse@arm.com, babu.moger@amd.com, corbet@lwn.net, shuah@kernel.org, jgg@ziepe.ca, kevin.tian@intel.com, cuiyunhui@bytedance.com, yuanzhu@bytedance.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, x86@kernel.org Subject: [RFC PATCH 4/7] iommu/riscv: Program QoS IDs for assigned groups Date: Tue, 14 Jul 2026 21:06:54 +0800 Message-ID: <20260714130657.46963-5-zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> References: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Program RCID and MCID for RISC-V IOMMU groups through the device context TA fields. The resctrl group assignment is per device group, so reject BARE mode where only the per-IOMMU iommu_qosid global default is available. Validate every group member, firmware ID, device context, field value, and QoS ID capability before changing hardware. Then update all members through the checked IOMMU group helper so a validation failure leaves the group unchanged. Serialize DC.ta changes with context setup under qosid_lock. Change only the RCID and MCID fields with ordinary accesses so fixed DDT mappings are not subject to atomic LR/SC operations, invalidate active device contexts after an update, and clear the IDs when a device is released. Signed-off-by: Zhanpeng Zhang --- arch/riscv/include/asm/qos.h | 16 +++ drivers/iommu/riscv/iommu-bits.h | 15 +++ drivers/iommu/riscv/iommu.c | 200 ++++++++++++++++++++++++++++++- drivers/iommu/riscv/iommu.h | 3 + 4 files changed, 232 insertions(+), 2 deletions(-) diff --git a/arch/riscv/include/asm/qos.h b/arch/riscv/include/asm/qos.h index cf19e8438bb9..daa758d4efff 100644 --- a/arch/riscv/include/asm/qos.h +++ b/arch/riscv/include/asm/qos.h @@ -2,7 +2,23 @@ #ifndef _ASM_RISCV_QOS_H #define _ASM_RISCV_QOS_H =20 +#include #include +#include + +struct iommu_group; + +#ifdef CONFIG_RISCV_IOMMU +int riscv_iommu_group_set_qosid(struct iommu_group *group, u32 rcid, + u32 mcid); +#else +static inline int riscv_iommu_group_set_qosid(struct iommu_group *group, + u32 rcid, u32 mcid) +{ + return -EOPNOTSUPP; +} + +#endif =20 #ifdef CONFIG_RISCV_ISA_SSQOSID =20 diff --git a/drivers/iommu/riscv/iommu-bits.h b/drivers/iommu/riscv/iommu-b= its.h index f2ef9bd3cde9..782de5c92727 100644 --- a/drivers/iommu/riscv/iommu-bits.h +++ b/drivers/iommu/riscv/iommu-bits.h @@ -63,6 +63,7 @@ #define RISCV_IOMMU_CAPABILITIES_PD8 BIT_ULL(38) #define RISCV_IOMMU_CAPABILITIES_PD17 BIT_ULL(39) #define RISCV_IOMMU_CAPABILITIES_PD20 BIT_ULL(40) +#define RISCV_IOMMU_CAPABILITIES_QOSID BIT_ULL(41) #define RISCV_IOMMU_CAPABILITIES_NL BIT_ULL(42) #define RISCV_IOMMU_CAPABILITIES_S BIT_ULL(43) =20 @@ -274,6 +275,14 @@ enum riscv_iommu_hpmevent_id { #define RISCV_IOMMU_TR_RESPONSE_SZ BIT_ULL(9) #define RISCV_IOMMU_TR_RESPONSE_PPN RISCV_IOMMU_PPN_FIELD =20 +/* 6.27 IOMMU QoS IDs for IOMMU-initiated requests (32bits) */ +#define RISCV_IOMMU_REG_IOMMU_QOSID 0x0270 +#define RISCV_IOMMU_IOMMU_QOSID_RCID GENMASK(11, 0) +#define RISCV_IOMMU_IOMMU_QOSID_MCID GENMASK(27, 16) + +#define RISCV_IOMMU_IOMMU_QOSID_RCID_SHIFT 0 +#define RISCV_IOMMU_IOMMU_QOSID_MCID_SHIFT 16 + /* 5.27 Interrupt cause to vector (64bits) */ #define RISCV_IOMMU_REG_ICVEC 0x02F8 #define RISCV_IOMMU_ICVEC_CIV GENMASK_ULL(3, 0) @@ -371,6 +380,12 @@ enum riscv_iommu_dc_iohgatp_modes { =20 /* Translation attributes fields */ #define RISCV_IOMMU_DC_TA_PSCID GENMASK_ULL(31, 12) +/* + * QoS IDs for translated device requests and IOMMU accesses with a + * device context (when capabilities.QOSID =3D=3D 1). + */ +#define RISCV_IOMMU_DC_TA_RCID GENMASK_ULL(51, 40) +#define RISCV_IOMMU_DC_TA_MCID GENMASK_ULL(63, 52) =20 /* First-stage context fields */ #define RISCV_IOMMU_DC_FSC_PPN RISCV_IOMMU_ATP_PPN_FIELD diff --git a/drivers/iommu/riscv/iommu.c b/drivers/iommu/riscv/iommu.c index cec3ddd7ab10..deab646bb1ea 100644 --- a/drivers/iommu/riscv/iommu.c +++ b/drivers/iommu/riscv/iommu.c @@ -48,6 +48,8 @@ static DEFINE_IDA(riscv_iommu_pscids); #define RISCV_IOMMU_MAX_PSCID (BIT(20) - 1) =20 +static const struct iommu_ops riscv_iommu_ops; + /* Device resource-managed allocations */ struct riscv_iommu_devres { void *addr; @@ -1091,6 +1093,28 @@ static void riscv_iommu_iotlb_inval(struct riscv_iom= mu_domain *domain, } =20 #define RISCV_IOMMU_FSC_BARE 0 +#define RISCV_IOMMU_DC_TA_QOSID \ + (RISCV_IOMMU_DC_TA_RCID | RISCV_IOMMU_DC_TA_MCID) + +static u64 riscv_iommu_qosid_ta(u32 rcid, u32 mcid) +{ + return FIELD_PREP(RISCV_IOMMU_DC_TA_RCID, rcid) | + FIELD_PREP(RISCV_IOMMU_DC_TA_MCID, mcid); +} + +static void riscv_iommu_dc_update_qosid(struct riscv_iommu_device *iommu, + struct riscv_iommu_dc *dc, + u32 rcid, u32 mcid) +{ + u64 qos_ta =3D riscv_iommu_qosid_ta(rcid, mcid); + u64 ta; + + lockdep_assert_held(&iommu->qosid_lock); + ta =3D READ_ONCE(dc->ta); + ta =3D (ta & ~RISCV_IOMMU_DC_TA_QOSID) | qos_ta; + WRITE_ONCE(dc->ta, ta); +} + /* * This function sends IOTINVAL commands as required by the RISC-V * IOMMU specification (Section 6.3.1 and 6.3.2 in 1.0 spec version) @@ -1202,12 +1226,23 @@ static void riscv_iommu_iodir_update(struct riscv_i= ommu_device *iommu, * is stored as DC_TC_V bit (both sharing the same location at BIT(0)). */ for (i =3D 0; i < fwspec->num_ids; i++) { + u64 dc_ta; + u64 ta_mask =3D RISCV_IOMMU_PC_TA_PSCID; + dc =3D riscv_iommu_get_dc(iommu, fwspec->ids[i]); tc =3D READ_ONCE(dc->tc); - tc |=3D ta & RISCV_IOMMU_DC_TC_V; + dc_ta =3D ta; + if (iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID) { + dc_ta |=3D READ_ONCE(dc->ta) & + (RISCV_IOMMU_DC_TA_RCID | + RISCV_IOMMU_DC_TA_MCID); + ta_mask |=3D RISCV_IOMMU_DC_TA_RCID | + RISCV_IOMMU_DC_TA_MCID; + } + tc |=3D dc_ta & RISCV_IOMMU_DC_TC_V; =20 WRITE_ONCE(dc->fsc, fsc); - WRITE_ONCE(dc->ta, ta & RISCV_IOMMU_PC_TA_PSCID); + WRITE_ONCE(dc->ta, dc_ta & ta_mask); /* Update device context, write TC.V as the last step. */ dma_wmb(); WRITE_ONCE(dc->tc, tc); @@ -1474,13 +1509,174 @@ static struct iommu_device *riscv_iommu_probe_devi= ce(struct device *dev) return &iommu->iommu; } =20 +static void riscv_iommu_qosid_invalidate_did(struct riscv_iommu_device *io= mmu, + unsigned int did) +{ + struct riscv_iommu_command cmd; + + riscv_iommu_cmd_iodir_inval_ddt(&cmd); + riscv_iommu_cmd_iodir_set_did(&cmd, did); + riscv_iommu_cmd_send(iommu, &cmd); +} + static void riscv_iommu_release_device(struct device *dev) { struct riscv_iommu_info *info =3D dev_iommu_priv_get(dev); + struct iommu_fwspec *fwspec =3D dev_iommu_fwspec_get(dev); + struct riscv_iommu_device *iommu =3D dev_to_iommu(dev); + bool sync_required =3D false; + unsigned int i; + + if (iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID) { + mutex_lock(&iommu->qosid_lock); + for (i =3D 0; fwspec && i < fwspec->num_ids; i++) { + struct riscv_iommu_dc *dc; + u64 tc; + + dc =3D riscv_iommu_get_dc(iommu, fwspec->ids[i]); + if (!dc) + continue; + + tc =3D READ_ONCE(dc->tc); + riscv_iommu_dc_update_qosid(iommu, dc, 0, 0); + if (!(tc & RISCV_IOMMU_DC_TC_V)) + continue; + + dma_wmb(); + riscv_iommu_qosid_invalidate_did(iommu, fwspec->ids[i]); + riscv_iommu_iodir_iotinval(iommu, false, dc->iohgatp, + dc, NULL); + sync_required =3D true; + } + + if (sync_required) + riscv_iommu_cmd_sync(iommu, + RISCV_IOMMU_IOTINVAL_TIMEOUT); + mutex_unlock(&iommu->qosid_lock); + } =20 kfree_rcu_mightsleep(info); } =20 +struct riscv_iommu_qosid_hw_ctx { + u32 rcid; + u32 mcid; + bool has_devices; + bool has_qosid; + bool reset; +}; + +static int riscv_iommu_qosid_validate_dev(struct device *dev, void *data) +{ + struct riscv_iommu_qosid_hw_ctx *ctx =3D data; + struct iommu_fwspec *fwspec =3D dev_iommu_fwspec_get(dev); + struct riscv_iommu_device *iommu; + unsigned int i; + + ctx->has_devices =3D true; + + if (!dev->iommu || !dev->iommu->iommu_dev || + dev->iommu->iommu_dev->ops !=3D &riscv_iommu_ops) + return -EOPNOTSUPP; + + if (!fwspec || !fwspec->num_ids) + return -ENODEV; + + iommu =3D dev_to_iommu(dev); + + if (!(iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID)) + return ctx->reset ? 0 : -EOPNOTSUPP; + + ctx->has_qosid =3D true; + + /* + * IOMMU group QoS is a per-device assignment. BARE mode only has the + * per-IOMMU iommu_qosid register, which is a global default rather + * than a safe target for moving an individual group between resctrl + * groups. + */ + if (iommu->ddt_mode <=3D RISCV_IOMMU_DDTP_IOMMU_MODE_BARE) + return -EOPNOTSUPP; + + for (i =3D 0; i < fwspec->num_ids; i++) { + if (!riscv_iommu_get_dc(iommu, fwspec->ids[i])) + return -ENODEV; + } + + return 0; +} + +static void riscv_iommu_qosid_apply_dev(struct device *dev, void *data) +{ + struct riscv_iommu_qosid_hw_ctx *ctx =3D data; + struct iommu_fwspec *fwspec =3D dev_iommu_fwspec_get(dev); + struct riscv_iommu_device *iommu; + bool sync_required =3D false; + unsigned int i; + + iommu =3D dev_to_iommu(dev); + if (!(iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID)) + return; + + mutex_lock(&iommu->qosid_lock); + for (i =3D 0; i < fwspec->num_ids; i++) { + struct riscv_iommu_dc *dc; + bool dc_is_valid; + u64 tc; + + dc =3D riscv_iommu_get_dc(iommu, fwspec->ids[i]); + if (WARN_ON_ONCE(!dc)) + continue; + + tc =3D READ_ONCE(dc->tc); + dc_is_valid =3D tc & RISCV_IOMMU_DC_TC_V; + + riscv_iommu_dc_update_qosid(iommu, dc, ctx->rcid, ctx->mcid); + dev_dbg(dev, "set QoS ID DC.ta did=3D%u rcid=3D%u mcid=3D%u\n", + fwspec->ids[i], ctx->rcid, ctx->mcid); + + if (dc_is_valid) { + dma_wmb(); + riscv_iommu_qosid_invalidate_did(iommu, fwspec->ids[i]); + riscv_iommu_iodir_iotinval(iommu, false, dc->iohgatp, + dc, NULL); + sync_required =3D true; + } + } + + if (sync_required) + riscv_iommu_cmd_sync(iommu, RISCV_IOMMU_IOTINVAL_TIMEOUT); + mutex_unlock(&iommu->qosid_lock); +} + +int riscv_iommu_group_set_qosid(struct iommu_group *group, u32 rcid, u32 m= cid) +{ + struct riscv_iommu_qosid_hw_ctx hw =3D { + .rcid =3D rcid, + .mcid =3D mcid, + .reset =3D !rcid && !mcid, + }; + int ret; + + if (rcid > FIELD_MAX(RISCV_IOMMU_DC_TA_RCID) || + mcid > FIELD_MAX(RISCV_IOMMU_DC_TA_MCID)) + return -ERANGE; + + ret =3D iommu_group_update_devices(group, &hw, + riscv_iommu_qosid_validate_dev, + riscv_iommu_qosid_apply_dev); + if (ret) + return ret; + if (!hw.has_devices) + return -ENODATA; + if (!hw.has_qosid && !hw.reset) + return -EOPNOTSUPP; + + pr_debug("set qosid: group=3D%d rcid=3D%u mcid=3D%u\n", + iommu_group_id(group), rcid, mcid); + return 0; +} + static const struct iommu_ops riscv_iommu_ops =3D { .of_xlate =3D riscv_iommu_of_xlate, .identity_domain =3D &riscv_iommu_identity_domain, diff --git a/drivers/iommu/riscv/iommu.h b/drivers/iommu/riscv/iommu.h index 46df79dd5495..2c57625637bf 100644 --- a/drivers/iommu/riscv/iommu.h +++ b/drivers/iommu/riscv/iommu.h @@ -19,6 +19,9 @@ =20 struct riscv_iommu_device; =20 +int riscv_iommu_group_set_qosid(struct iommu_group *group, u32 rcid, + u32 mcid); + struct riscv_iommu_queue { atomic_t prod; /* unbounded producer allocation index */ atomic_t head; /* unbounded shadow ring buffer consumer index */ --=20 2.50.1 (Apple Git-155) From nobody Sat Jul 25 19:29:10 2026 Received: from mail-ot1-f45.google.com (mail-ot1-f45.google.com [209.85.210.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2CAE347279B for ; Tue, 14 Jul 2026 13:08:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.45 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034494; cv=none; b=DyfiQsB1Si0VYUfqM5UOYpfqTWvYAuMe5dlZIqh8d+Ss1WVgXVfQ49XFBzievoPE9/puLPFSaP3j591cqiKgSupU2ebnCmutn9jPsdEY80GiaXYqaWWWL+4efTgyijyHMbKyHHRb0Z+iq5+gzBerjUB3GwYaXr8mYDvDiBFafcU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034494; c=relaxed/simple; bh=8nc2OVUHCKRxT0Ws6HOAWmw/VzTaZ0aIVfDSOxFg04M=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=tPvTOW3Le9AfIB/DblHhuNT82BBk6aZBktRJIFXNEwLMaWM7du1PLpc2PztcmPipxnA6LvRFJJ0EyDSvBHGsxTwrZVUFcX2xR/rA2S6nPmmjBjQKi+45FJq1AFhVm5KfezVptUOvxBA2JZK8XV6VhZtwnJTNC2sTh56U/6BR3ZA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=ZRdWoDgf; arc=none smtp.client-ip=209.85.210.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="ZRdWoDgf" Received: by mail-ot1-f45.google.com with SMTP id 46e09a7af769-7eb5bdb50fcso458764a34.1 for ; Tue, 14 Jul 2026 06:08:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1784034491; x=1784639291; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=tyPwfho2+gt5VglzPQUkATM5bIbB7AhJfPglM91knA4=; b=ZRdWoDgf5101ckJp+EfC9RoQvSMN8/RQV5Va1lyGKpYIqxuBWrYzecZffmq2hkLJCX ua18RsNNESGqnJxLBBiWPVlIcwwivywhQ4vy/rglmjX1e7YD3IuyU4PpT5FLDfqAAwRc CV/klMmhlN9QZibUeDxe6fWIslsHMa62qaN04BHAuE0nHLTiWIJ9LHZo5Tqt1z1ZLGel Jb9E4tVXT12XpBiMYeIp8m9g23IrlAPYm09m62M4ysPU4PSAlol7jUjqv8V7bfCSizDi TLQBUVsRqxoCfcchN0hIYZsUN4BGdQanofXi56n8ZbfS3CMzxmRUntZ4TiE2mH0AY04V zxvg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784034491; x=1784639291; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=tyPwfho2+gt5VglzPQUkATM5bIbB7AhJfPglM91knA4=; b=Xw+tGy0etxyY4N5OJ1bwjitU4BZcPvaSP7wukHMsESAhQgAIX5sLM0dVh1dV4bJ8kI eATIvWAX/YGNWMpqREZ8nfbAykKTQArBLiz0YYmGJ1iEIh7CtdKWXwTd8hUnEajVurc4 rbz9JlhJYuLejNXtGfsr4UGFofO6JSWts7CszboErYwy3Ds6mF1l1nMnDcbfkxAIr/3M q0wyv5bxjATqUqACD4HInPmA/PwEJfrjjZ/IzeJGP9BDfJFWn5EIC7GAk9ofQVxXWhMn CdDgBJ97eDQe0jyYFi/dLwbiHaI8EImdNAmFZz0tlpIVyfrwWbcUpVFcF0bZ7VKzb8mV z9jg== X-Forwarded-Encrypted: i=1; AFNElJ8zXuRKHh8gCeSeIMYWBG1CgMYgfccPzd2CiY9F/5HdP7/8reHwz/8H0UAHBV5+J4rCxJxCKTVHuR0m2Uc=@vger.kernel.org X-Gm-Message-State: AOJu0Ywlv+SHYdamxkDsEEF1sBs/EtbTyjXRl4d0CGoKdZ4TZhPmsdLh xh2aQXlRn8oRJxBvZnBwuz9J06ihsdAhIYGmF1bVetqxPVs9EhNKa5qlHndw9MvKfXo= X-Gm-Gg: AfdE7cmLElcO+vPH5G8Fw05kMhsZOF8RmODVotnPzbOuia/iUV+wzA5PCQHew3lrWTp 7ZhPjLNNKmbv35k6ZO/mN73vuKQd2JzA2aXeTNHouAqovMSCG5Me4PeqHyw5t+NWqvn8s2k4mI0 dO9c4iMJQ0mmrVIbfWPD7hejNY8bg3vVJzvOOq0Y+sidPHFpv8KGgrk71YlnUAaVxDh/TgiJvIo HKEKfMF1od69z9uHk1DsuxcDNzAYMllLyDhdVKPxiMMR4r0yUnqibVltC+VeNLIjOzuIruIF4Mv BVqrqrbiXk8KQ1Hh5QK0hLVhlh49HeFXJaSjzZL3Q1tz0tk2YQ5N4WLQb4JsTG0tjrQgVo1z9Wc WzLkso8yGJqND11DjKw0nRIRIF4vrYv5+pw58XlouWRBdJoUEwDXbbyPjJ9U+6RU4r5uLW+I9wB a+agFVHIYxKfoOhR+JlcXu6RLrlrf4VpFvHxzHdSRS8crF2gIsWXcaJ0Eiz7+VhA== X-Received: by 2002:a05:6830:82ad:b0:7eb:3af8:8c1a with SMTP id 46e09a7af769-7ec096ecdc5mr8483705a34.9.1784034490998; Tue, 14 Jul 2026 06:08:10 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([178.93.176.7]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7ebcab8efc3sm14657738a34.0.2026.07.14.06.07.59 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 14 Jul 2026 06:08:10 -0700 (PDT) From: Zhanpeng Zhang To: joro@8bytes.org, palmer@dabbelt.com, tony.luck@intel.com, reinette.chatre@intel.com, tomasz.jeznach@linux.dev Cc: will@kernel.org, robin.murphy@arm.com, fustini@kernel.org, pjw@kernel.org, aou@eecs.berkeley.edu, alex@ghiti.fr, Dave.Martin@arm.com, james.morse@arm.com, babu.moger@amd.com, corbet@lwn.net, shuah@kernel.org, jgg@ziepe.ca, kevin.tian@intel.com, cuiyunhui@bytedance.com, yuanzhu@bytedance.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, x86@kernel.org Subject: [RFC PATCH 5/7] iommu/riscv: Expose global QoS IDs in sysfs Date: Tue, 14 Jul 2026 21:06:55 +0800 Message-ID: <20260714130657.46963-6-zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> References: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The RISC-V IOMMU QoS extension provides iommu_qosid as a per-IOMMU global default tag. It is used for IOMMU-originated DDT, CQ, FQ, PQ, and MSI accesses, and for device-originated requests when DDTP is in BARE mode. Initialize iommu_qosid to RCID 0 and MCID 0 when the hardware advertises QOSID support. Preserve reserved and WPRI bits with read-modify-write, and use register readback to reject values which the WARL fields do not retain. Add a qosid attribute to the RISC-V IOMMU class device. Reading returns the current RCID and MCID values. Writing the documented 'rcid=3D mcid=3D' form updates both fields while preserving the other register bits. Keep this interface separate from resctrl group QoS. The sysfs attribute controls the IOMMU-wide default, while resctrl device assignment programs per-device DC.ta in translated modes. Signed-off-by: Zhanpeng Zhang --- .../ABI/testing/sysfs-class-iommu-riscv-iommu | 27 +++ MAINTAINERS | 10 ++ drivers/iommu/riscv/iommu.c | 159 +++++++++++++++++- drivers/iommu/riscv/iommu.h | 9 +- 4 files changed, 202 insertions(+), 3 deletions(-) create mode 100644 Documentation/ABI/testing/sysfs-class-iommu-riscv-iommu diff --git a/Documentation/ABI/testing/sysfs-class-iommu-riscv-iommu b/Docu= mentation/ABI/testing/sysfs-class-iommu-riscv-iommu new file mode 100644 index 000000000000..b0cd68997f17 --- /dev/null +++ b/Documentation/ABI/testing/sysfs-class-iommu-riscv-iommu @@ -0,0 +1,27 @@ +What: /sys/class/iommu//qosid +Date: June 2026 +KernelVersion: 6.18 +Contact: Zhanpeng Zhang +Description: + The RISC-V IOMMU global default QoS IDs for this IOMMU. + The file is present only when the IOMMU reports the QOSID + capability. + + Reading the file returns the RCID and MCID fields from the + iommu_qosid register: + + rcid=3D mcid=3D + + Writing the file updates the RCID and MCID fields while + preserving reserved/WPRI bits: + + rcid=3D mcid=3D + + Writes fail with ERANGE when either value cannot be represented + by the IOMMU. A successful write is verified by reading the WARL + fields back from the register. + + The iommu_qosid register is a per-IOMMU global default. It + tags IOMMU-originated DDT, CQ, FQ, PQ and MSI accesses, and + in BARE mode device-originated requests. It does not assign + per-device or per-IOMMU-group QoS IDs in translated modes. diff --git a/MAINTAINERS b/MAINTAINERS index 0b5d38b772e0..c59be02c8f02 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -23279,6 +23279,16 @@ T: git git://git.kernel.org/pub/scm/linux/kernel/g= it/iommu/linux.git F: Documentation/devicetree/bindings/iommu/riscv,iommu.yaml F: drivers/iommu/riscv/ =20 +RISC-V IOMMU QoS +M: Zhanpeng Zhang +R: Tomasz Jeznach +R: Drew Fustini +R: yunhui cui +L: iommu@lists.linux.dev +L: linux-riscv@lists.infradead.org +S: Maintained +F: Documentation/ABI/testing/sysfs-class-iommu-riscv-iommu + RISC-V MICROCHIP SUPPORT M: Conor Dooley M: Daire McNamara diff --git a/drivers/iommu/riscv/iommu.c b/drivers/iommu/riscv/iommu.c index deab646bb1ea..e85da9eef58e 100644 --- a/drivers/iommu/riscv/iommu.c +++ b/drivers/iommu/riscv/iommu.c @@ -1658,6 +1658,7 @@ int riscv_iommu_group_set_qosid(struct iommu_group *g= roup, u32 rcid, u32 mcid) }; int ret; =20 + /* Resctrl IDs are bounded by the system's reported controller counts. */ if (rcid > FIELD_MAX(RISCV_IOMMU_DC_TA_RCID) || mcid > FIELD_MAX(RISCV_IOMMU_DC_TA_MCID)) return -ERANGE; @@ -1688,9 +1689,154 @@ static const struct iommu_ops riscv_iommu_ops =3D { .release_device =3D riscv_iommu_release_device, }; =20 +static int riscv_iommu_set_default_qosid(struct riscv_iommu_device *iommu, + u32 rcid, u32 mcid) +{ + u32 old_qosid; + u32 qosid; + int ret =3D 0; + + if (!(iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID)) + return -EOPNOTSUPP; + + if (rcid > FIELD_MAX(RISCV_IOMMU_IOMMU_QOSID_RCID) || + mcid > FIELD_MAX(RISCV_IOMMU_IOMMU_QOSID_MCID)) + return -ERANGE; + + /* + * iommu_qosid is a per-IOMMU global default. It tags IOMMU-originated + * DDT/CQ/FQ/PQ and MSI accesses, and in BARE mode device-originated + * requests. Per-device group QoS is still handled separately through + * DC.ta. + */ + mutex_lock(&iommu->qosid_lock); + old_qosid =3D riscv_iommu_readl(iommu, RISCV_IOMMU_REG_IOMMU_QOSID); + qosid =3D old_qosid & ~(RISCV_IOMMU_IOMMU_QOSID_RCID | + RISCV_IOMMU_IOMMU_QOSID_MCID); + qosid |=3D FIELD_PREP(RISCV_IOMMU_IOMMU_QOSID_RCID, rcid) | + FIELD_PREP(RISCV_IOMMU_IOMMU_QOSID_MCID, mcid); + riscv_iommu_writel(iommu, RISCV_IOMMU_REG_IOMMU_QOSID, qosid); + + qosid =3D riscv_iommu_readl(iommu, RISCV_IOMMU_REG_IOMMU_QOSID); + if (FIELD_GET(RISCV_IOMMU_IOMMU_QOSID_RCID, qosid) !=3D rcid || + FIELD_GET(RISCV_IOMMU_IOMMU_QOSID_MCID, qosid) !=3D mcid) { + riscv_iommu_writel(iommu, RISCV_IOMMU_REG_IOMMU_QOSID, + old_qosid); + ret =3D -ERANGE; + } + mutex_unlock(&iommu->qosid_lock); + if (ret) + return ret; + + dev_dbg(iommu->dev, "set global QoS IDs rcid=3D%u mcid=3D%u\n", + (u32)FIELD_GET(RISCV_IOMMU_IOMMU_QOSID_RCID, qosid), + (u32)FIELD_GET(RISCV_IOMMU_IOMMU_QOSID_MCID, qosid)); + + return ret; +} + +static int riscv_iommu_get_default_qosid(struct riscv_iommu_device *iommu, + u32 *rcid, u32 *mcid) +{ + u32 qosid; + + if (!(iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID)) + return -EOPNOTSUPP; + + qosid =3D riscv_iommu_readl(iommu, RISCV_IOMMU_REG_IOMMU_QOSID); + if (rcid) + *rcid =3D FIELD_GET(RISCV_IOMMU_IOMMU_QOSID_RCID, qosid); + if (mcid) + *mcid =3D FIELD_GET(RISCV_IOMMU_IOMMU_QOSID_MCID, qosid); + + return 0; +} + +static int riscv_iommu_init_default_qosid(struct riscv_iommu_device *iommu) +{ + if (!(iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID)) + return 0; + + /* Avoid probing the live WARL fields with all-ones while in BARE mode. */ + return riscv_iommu_set_default_qosid(iommu, 0, 0); +} + +static struct riscv_iommu_device *dev_to_riscv_iommu(struct device *dev) +{ + struct iommu_device *iommu =3D dev_to_iommu_device(dev); + + return iommu ? container_of(iommu, struct riscv_iommu_device, iommu) : NU= LL; +} + +static ssize_t qosid_show(struct device *dev, struct device_attribute *att= r, + char *buf) +{ + struct riscv_iommu_device *iommu =3D dev_to_riscv_iommu(dev); + u32 rcid, mcid; + int ret; + + if (!iommu) + return -ENODEV; + + ret =3D riscv_iommu_get_default_qosid(iommu, &rcid, &mcid); + if (ret) + return ret; + + return sysfs_emit(buf, "rcid=3D%u mcid=3D%u\n", rcid, mcid); +} + +static ssize_t qosid_store(struct device *dev, struct device_attribute *at= tr, + const char *buf, size_t count) +{ + struct riscv_iommu_device *iommu =3D dev_to_riscv_iommu(dev); + char *args, *key, *value; + char *input; + u32 rcid, mcid; + int ret =3D -EINVAL; + + if (!iommu) + return -ENODEV; + + input =3D kstrdup(buf, GFP_KERNEL); + if (!input) + return -ENOMEM; + + args =3D strim(input); + args =3D next_arg(args, &key, &value); + if (!value || strcmp(key, "rcid") || kstrtou32(value, 10, &rcid)) + goto out; + + args =3D next_arg(args, &key, &value); + if (!value || strcmp(key, "mcid") || kstrtou32(value, 10, &mcid) || + *skip_spaces(args)) + goto out; + + ret =3D riscv_iommu_set_default_qosid(iommu, rcid, mcid); +out: + kfree(input); + return ret ? ret : count; +} + +static DEVICE_ATTR_RW(qosid); + +static struct attribute *riscv_iommu_attrs[] =3D { + &dev_attr_qosid.attr, + NULL, +}; + +static const struct attribute_group riscv_iommu_group =3D { + .attrs =3D riscv_iommu_attrs, +}; + +static const struct attribute_group *riscv_iommu_groups[] =3D { + &riscv_iommu_group, + NULL, +}; + static int riscv_iommu_init_check(struct riscv_iommu_device *iommu) { u64 ddtp; + int ret; =20 /* * Make sure the IOMMU is switched off or in pass-through mode during @@ -1721,6 +1867,10 @@ static int riscv_iommu_init_check(struct riscv_iommu= _device *iommu) return -EINVAL; } =20 + ret =3D riscv_iommu_init_default_qosid(iommu); + if (ret) + return ret; + /* * Distribute interrupt vectors, always use first vector for CIV. * At least one interrupt is required. Read back and verify. @@ -1753,10 +1903,12 @@ void riscv_iommu_remove(struct riscv_iommu_device *= iommu) =20 int riscv_iommu_init(struct riscv_iommu_device *iommu) { + const struct attribute_group **sysfs_groups =3D NULL; int rc; =20 RISCV_IOMMU_QUEUE_INIT(&iommu->cmdq, CQ); RISCV_IOMMU_QUEUE_INIT(&iommu->fltq, FQ); + mutex_init(&iommu->qosid_lock); =20 rc =3D riscv_iommu_init_check(iommu); if (rc) @@ -1788,8 +1940,11 @@ int riscv_iommu_init(struct riscv_iommu_device *iomm= u) if (rc) goto err_queue_disable; =20 - rc =3D iommu_device_sysfs_add(&iommu->iommu, NULL, NULL, "riscv-iommu@%s", - dev_name(iommu->dev)); + if (iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID) + sysfs_groups =3D riscv_iommu_groups; + + rc =3D iommu_device_sysfs_add(&iommu->iommu, NULL, sysfs_groups, + "riscv-iommu@%s", dev_name(iommu->dev)); if (rc) { dev_err_probe(iommu->dev, rc, "cannot register sysfs interface\n"); goto err_iodir_off; diff --git a/drivers/iommu/riscv/iommu.h b/drivers/iommu/riscv/iommu.h index 2c57625637bf..13ea67b7e42d 100644 --- a/drivers/iommu/riscv/iommu.h +++ b/drivers/iommu/riscv/iommu.h @@ -12,8 +12,12 @@ #define _RISCV_IOMMU_H_ =20 #include -#include #include +#include +#include +#ifdef CONFIG_RISCV_IOMMU_32BIT +#include +#endif =20 #include "iommu-bits.h" =20 @@ -50,6 +54,9 @@ struct riscv_iommu_device { u64 caps; u32 fctl; =20 + /* Serializes QoS updates to iommu_qosid and device contexts. */ + struct mutex qosid_lock; + /* available interrupt numbers, MSI or WSI */ unsigned int irqs[RISCV_IOMMU_INTR_COUNT]; unsigned int irqs_count; --=20 2.50.1 (Apple Git-155) From nobody Sat Jul 25 19:29:10 2026 Received: from mail-ot1-f50.google.com (mail-ot1-f50.google.com [209.85.210.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8968047A0C3 for ; Tue, 14 Jul 2026 13:08:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.50 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034508; cv=none; b=ZGfGCxq1tXAzgzMvfv9rGro1oAp8sk7j/N3SPOKh0Hkw90DAZ2JLHy/DguYK0myHhwO6/hl8QkZN79pakTV0Q5gkibHTGRx3slXdcNT+tTCX0VOz926t3z5946TSMSvKJAyai4B/1EUs/9tNvuLtQzuRnvIBK3/9oLXl6loCloM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034508; c=relaxed/simple; bh=88hFgKCg+bUuFow9gcc2rzIqGEA2+sQ8+uzIG1WjY5E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=OFgor0zNON0sRiZp61Cat7JNiwmLlJXVbWZfC36cx0kijPoyo1Z37OqNqAkEBHazS2XZJeMhmiJ4DD/CX61bwp8S4iM8kKDWqiYiICjfzMd+VFqZ4b3vnoo+3OB/PbOwNXvSD8zqaT4HCWWCEzIlpjdtJ1Ard9oYZdPDXQuoz4c= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=JhoqJraW; arc=none smtp.client-ip=209.85.210.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="JhoqJraW" Received: by mail-ot1-f50.google.com with SMTP id 46e09a7af769-7e9d7464b71so1393699a34.0 for ; Tue, 14 Jul 2026 06:08:25 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1784034504; x=1784639304; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=0MZi7cORh66e/+tuz0lS/dEufQyLW2u/I7nsyx5ACQE=; b=JhoqJraWwbSfQTOjo0oznYZNpmDNhAiy/r6HZRTPyF+CdqMfenm5c/A/JvYkH9psOR Vgf+90E88B2SXXxJDXYdss5dU8Ag2JW96MK0cPmP2JltbwLO4qX8YAKQlGnruIeGMr8T kp0JYI1zx3rgy4UeBzqEAUtVyDBa9O9081OyVKHSBYZ5KvbroPJz5rtvDy/+5MBk512H 0ENXJr6fZFxetNGDDUAI9l4dEcOO1APo+aXabh7T2GGc7ScVLNP68g+Oxqiyw31bOF4t B4G1PDvdwmiZMNE7QzsbJwwAqqAs6pwu0f3D43RxMoGqm6sqCJLme7utXF4Cn9mN6/on 91yA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784034504; x=1784639304; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=0MZi7cORh66e/+tuz0lS/dEufQyLW2u/I7nsyx5ACQE=; b=YwztX/NFtyUoPWvSUGVTXAh0PkeThgxQfltpMFJNGODoGjXl6EKxWjHjxekQ3yZaw7 j+E52ExbZ3u1YEDufqpFW/LgB5dFdIpZ6Q8Vyv9/hjXydkHmPeelIrB4FA2zjzArGPoK hdpN1blELjW6yOmm0nwCKkHLzJC7BGsInSJj27wdz2dae0mWPWOfAy3g3WP8VbVet9DR E6l/MC/qwBxOlpB7ImopPkWGVzj/5Q+bUc4eZZdEvnkvwPg5eKUcJ4F3pg0U0le+Nco+ S9YcD1Lyr5hNzfPuNwUhDMqBC/GeNVDiZZX8i8gaI4hZG4Xjv6EiFah0DD4mgVBWZOJf 1uNw== X-Forwarded-Encrypted: i=1; AFNElJ9sfnJX0DjmBtjMRuM6SglYxcvXmt8q9OTL9OcOCWrgAtitGaYLksw1Dn/k1d+9lit8lFmP4CsYdmhyKcI=@vger.kernel.org X-Gm-Message-State: AOJu0YxmDqmWQy3TlQcAgalUbwG0ngUOHOYq9snStEI9TODuTL3fM37a CiY1HZjjx+M+pDvJFA7FTPQ1meC47JV3kMfiSXnHkpX4XYtjw3+efyxQT317/2fT01o= X-Gm-Gg: AfdE7cnPfWVSw+aQFc6tNJP+4gsRM0EYYLDKdU4yt1pvy4jAgfLAoDwq4YIpk2UU90z nXYti5S2u1gQj9GrhQy6FHJ2VkMil10sMcrRr3aEedj5lyOTFb5PW8GpLrzpWIGCvgKxfgWagke Jit+yb9or+4Uq6rwQPADWTMREhtfnBo3phj5gZCvRCqjT/LcYABLVEiYrGN70jkKWQ+KQiFwwXQ ghC0gubltdZc8tlXwbF9EbjdMf8owpbZODnxm26vo20ZAAFs3eFu+Hef+ytMZuTWQPKaRAbV8GG UCnTBXkZX0iGqmxBTE01XKBqBCw/qch05QPq8jGoVX7a+7ny0NacCkf0CfDTW4yBBRk63DXSlnm ypX6DBJxYIGIjMZvy3MSc03ntB/Ydc6isfytkSxrqB1z6wib2exu8G8jm1zsDbgqvX9jcrWKa6l 1jaArSKgWUVdUpLUix5IIOQBYBWjqSeUdxg9Er1Ok2CwsYSB69E2j4RoZEsNdkiOTTLTFvrC1E X-Received: by 2002:a05:6830:3817:b0:7e1:cbe3:bb1b with SMTP id 46e09a7af769-7ec4a5cda29mr1176418a34.0.1784034503802; Tue, 14 Jul 2026 06:08:23 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([178.93.176.7]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7ebcab8efc3sm14657738a34.0.2026.07.14.06.08.11 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 14 Jul 2026 06:08:23 -0700 (PDT) From: Zhanpeng Zhang To: joro@8bytes.org, palmer@dabbelt.com, tony.luck@intel.com, reinette.chatre@intel.com, tomasz.jeznach@linux.dev Cc: will@kernel.org, robin.murphy@arm.com, fustini@kernel.org, pjw@kernel.org, aou@eecs.berkeley.edu, alex@ghiti.fr, Dave.Martin@arm.com, james.morse@arm.com, babu.moger@amd.com, corbet@lwn.net, shuah@kernel.org, jgg@ziepe.ca, kevin.tian@intel.com, cuiyunhui@bytedance.com, yuanzhu@bytedance.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, x86@kernel.org Subject: [RFC PATCH 6/7] riscv_cbqri: Assign IOMMU groups to resource groups Date: Tue, 14 Jul 2026 21:06:56 +0800 Message-ID: <20260714130657.46963-7-zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> References: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Allow the resctrl devices file to accept iommu_group: tokens on RISC-V and map the target resource group's closid and rmid values to RCID and MCID. Record non-default assignments in an RCU-protected binding list rather than inferring membership from hardware IDs, which may be shared. Keep a parent reference obtained by numeric group lookup so the binding does not keep an empty group's devices kobject active, and prune bindings after their group becomes inactive. Publish an explicit UPDATING state while hardware changes. Paging-domain attachment rejects that transient state. Static FSC=3DBare transitions preserve the current IDs so a mandatory release-domain attachment cannot fail. After the checked group update succeeds, publish the packed IDs. A validation failure restores the previous active state without partially changing hardware. Moving an IOMMU group to the default resource group resets its hardware state and removes the software binding. Device contexts created later for an assigned group inherit the IDs through the RCU lookup path. Signed-off-by: Zhanpeng Zhang --- MAINTAINERS | 1 + arch/riscv/include/asm/qos.h | 12 ++ drivers/iommu/riscv/iommu.c | 70 +++++++-- drivers/iommu/riscv/iommu.h | 16 +- drivers/resctrl/Kconfig | 6 + drivers/resctrl/Makefile | 1 + drivers/resctrl/cbqri_iommu.c | 276 ++++++++++++++++++++++++++++++++++ 7 files changed, 369 insertions(+), 13 deletions(-) create mode 100644 drivers/resctrl/cbqri_iommu.c diff --git a/MAINTAINERS b/MAINTAINERS index c59be02c8f02..162ad5a1780f 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -23288,6 +23288,7 @@ L: iommu@lists.linux.dev L: linux-riscv@lists.infradead.org S: Maintained F: Documentation/ABI/testing/sysfs-class-iommu-riscv-iommu +F: drivers/resctrl/cbqri_iommu.c =20 RISC-V MICROCHIP SUPPORT M: Conor Dooley diff --git a/arch/riscv/include/asm/qos.h b/arch/riscv/include/asm/qos.h index daa758d4efff..84a93739c6ff 100644 --- a/arch/riscv/include/asm/qos.h +++ b/arch/riscv/include/asm/qos.h @@ -8,6 +8,18 @@ =20 struct iommu_group; =20 +#ifdef CONFIG_RISCV_CBQRI_IOMMU +int riscv_resctrl_iommu_group_get_qosid(struct iommu_group *group, + u32 *rcid, u32 *mcid); +#else +static inline int +riscv_resctrl_iommu_group_get_qosid(struct iommu_group *group, u32 *rcid, + u32 *mcid) +{ + return -ENOENT; +} +#endif + #ifdef CONFIG_RISCV_IOMMU int riscv_iommu_group_set_qosid(struct iommu_group *group, u32 rcid, u32 mcid); diff --git a/drivers/iommu/riscv/iommu.c b/drivers/iommu/riscv/iommu.c index e85da9eef58e..1a7d8a582959 100644 --- a/drivers/iommu/riscv/iommu.c +++ b/drivers/iommu/riscv/iommu.c @@ -1188,16 +1188,50 @@ static void riscv_iommu_iodir_iotinval(struct riscv= _iommu_device *iommu, * device is not quiesced might be disruptive, potentially causing * interim translation faults. */ -static void riscv_iommu_iodir_update(struct riscv_iommu_device *iommu, - struct device *dev, u64 fsc, u64 ta) +static int riscv_iommu_iodir_update(struct riscv_iommu_device *iommu, + struct device *dev, u64 fsc, u64 ta) { struct iommu_fwspec *fwspec =3D dev_iommu_fwspec_get(dev); struct riscv_iommu_dc *dc; struct riscv_iommu_command cmd; bool sync_required =3D false; + u64 group_qos_ta =3D 0; + bool has_group_qos; + u32 group_rcid; + u32 group_mcid; u64 tc; + int ret =3D 0; + int qos_ret; int i; =20 + /* Serialize DC.ta updates with resctrl IOMMU group QoS changes. */ + mutex_lock(&iommu->qosid_lock); + + qos_ret =3D riscv_resctrl_iommu_group_get_qosid(dev->iommu_group, + &group_rcid, &group_mcid); + if (qos_ret =3D=3D -EAGAIN) { + /* Static domain transitions must not fail during device release. */ + if (fsc !=3D RISCV_IOMMU_FSC_BARE) { + ret =3D -EBUSY; + goto unlock; + } + qos_ret =3D -ENOENT; + } + has_group_qos =3D !qos_ret; + if (has_group_qos) { + if (!(iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID) || + iommu->ddt_mode <=3D RISCV_IOMMU_DDTP_IOMMU_MODE_BARE) { + ret =3D -EOPNOTSUPP; + goto unlock; + } + if (group_rcid > FIELD_MAX(RISCV_IOMMU_DC_TA_RCID) || + group_mcid > FIELD_MAX(RISCV_IOMMU_DC_TA_MCID)) { + ret =3D -ERANGE; + goto unlock; + } + group_qos_ta =3D riscv_iommu_qosid_ta(group_rcid, group_mcid); + } + for (i =3D 0; i < fwspec->num_ids; i++) { dc =3D riscv_iommu_get_dc(iommu, fwspec->ids[i]); tc =3D READ_ONCE(dc->tc); @@ -1233,11 +1267,12 @@ static void riscv_iommu_iodir_update(struct riscv_i= ommu_device *iommu, tc =3D READ_ONCE(dc->tc); dc_ta =3D ta; if (iommu->caps & RISCV_IOMMU_CAPABILITIES_QOSID) { - dc_ta |=3D READ_ONCE(dc->ta) & - (RISCV_IOMMU_DC_TA_RCID | - RISCV_IOMMU_DC_TA_MCID); - ta_mask |=3D RISCV_IOMMU_DC_TA_RCID | - RISCV_IOMMU_DC_TA_MCID; + if (has_group_qos) + dc_ta |=3D group_qos_ta; + else + dc_ta |=3D READ_ONCE(dc->ta) & + RISCV_IOMMU_DC_TA_QOSID; + ta_mask |=3D RISCV_IOMMU_DC_TA_QOSID; } tc |=3D dc_ta & RISCV_IOMMU_DC_TC_V; =20 @@ -1259,6 +1294,9 @@ static void riscv_iommu_iodir_update(struct riscv_iom= mu_device *iommu, } =20 riscv_iommu_cmd_sync(iommu, RISCV_IOMMU_IOTINVAL_TIMEOUT); +unlock: + mutex_unlock(&iommu->qosid_lock); + return ret; } =20 /* @@ -1324,6 +1362,7 @@ static int riscv_iommu_attach_paging_domain(struct io= mmu_domain *iommu_domain, struct riscv_iommu_info *info =3D dev_iommu_priv_get(dev); struct pt_iommu_riscv_64_hw_info pt_info; u64 fsc, ta; + int ret; =20 pt_iommu_riscv_64_hw_info(&domain->riscvpt, &pt_info); =20 @@ -1338,7 +1377,11 @@ static int riscv_iommu_attach_paging_domain(struct i= ommu_domain *iommu_domain, if (riscv_iommu_bond_link(domain, dev)) return -ENOMEM; =20 - riscv_iommu_iodir_update(iommu, dev, fsc, ta); + ret =3D riscv_iommu_iodir_update(iommu, dev, fsc, ta); + if (ret) { + riscv_iommu_bond_unlink(domain, dev); + return ret; + } riscv_iommu_bond_unlink(info->domain, dev); info->domain =3D domain; =20 @@ -1413,9 +1456,12 @@ static int riscv_iommu_attach_blocking_domain(struct= iommu_domain *iommu_domain, { struct riscv_iommu_device *iommu =3D dev_to_iommu(dev); struct riscv_iommu_info *info =3D dev_iommu_priv_get(dev); + int ret; =20 /* Make device context invalid, translation requests will fault w/ #258 */ - riscv_iommu_iodir_update(iommu, dev, RISCV_IOMMU_FSC_BARE, 0); + ret =3D riscv_iommu_iodir_update(iommu, dev, RISCV_IOMMU_FSC_BARE, 0); + if (ret) + return ret; riscv_iommu_bond_unlink(info->domain, dev); info->domain =3D NULL; =20 @@ -1435,8 +1481,12 @@ static int riscv_iommu_attach_identity_domain(struct= iommu_domain *iommu_domain, { struct riscv_iommu_device *iommu =3D dev_to_iommu(dev); struct riscv_iommu_info *info =3D dev_iommu_priv_get(dev); + u64 ta =3D RISCV_IOMMU_PC_TA_V; + int ret; =20 - riscv_iommu_iodir_update(iommu, dev, RISCV_IOMMU_FSC_BARE, RISCV_IOMMU_PC= _TA_V); + ret =3D riscv_iommu_iodir_update(iommu, dev, RISCV_IOMMU_FSC_BARE, ta); + if (ret) + return ret; riscv_iommu_bond_unlink(info->domain, dev); info->domain =3D NULL; =20 diff --git a/drivers/iommu/riscv/iommu.h b/drivers/iommu/riscv/iommu.h index 13ea67b7e42d..e7804ec65647 100644 --- a/drivers/iommu/riscv/iommu.h +++ b/drivers/iommu/riscv/iommu.h @@ -11,13 +11,11 @@ #ifndef _RISCV_IOMMU_H_ #define _RISCV_IOMMU_H_ =20 +#include #include #include #include #include -#ifdef CONFIG_RISCV_IOMMU_32BIT -#include -#endif =20 #include "iommu-bits.h" =20 @@ -26,6 +24,18 @@ struct riscv_iommu_device; int riscv_iommu_group_set_qosid(struct iommu_group *group, u32 rcid, u32 mcid); =20 +#ifdef CONFIG_RISCV_CBQRI_IOMMU +int riscv_resctrl_iommu_group_get_qosid(struct iommu_group *group, + u32 *rcid, u32 *mcid); +#else +static inline int +riscv_resctrl_iommu_group_get_qosid(struct iommu_group *group, u32 *rcid, + u32 *mcid) +{ + return -ENOENT; +} +#endif + struct riscv_iommu_queue { atomic_t prod; /* unbounded producer allocation index */ atomic_t head; /* unbounded shadow ring buffer consumer index */ diff --git a/drivers/resctrl/Kconfig b/drivers/resctrl/Kconfig index b7db6ff9d054..4e10c813e647 100644 --- a/drivers/resctrl/Kconfig +++ b/drivers/resctrl/Kconfig @@ -60,3 +60,9 @@ endif config RISCV_CBQRI_RESCTRL_FS bool default y if RISCV_CBQRI && RESCTRL_FS + +config RISCV_CBQRI_IOMMU + bool + depends on RISCV_CBQRI_RESCTRL_FS && RISCV_IOMMU + select ARCH_HAS_RESCTRL_DEVICES + default y diff --git a/drivers/resctrl/Makefile b/drivers/resctrl/Makefile index c8339113ef1f..148b2e28c6a3 100644 --- a/drivers/resctrl/Makefile +++ b/drivers/resctrl/Makefile @@ -8,3 +8,4 @@ obj-$(CONFIG_RISCV_CBQRI) +=3D cbqri.o cbqri-y +=3D cbqri_devices.o cbqri-$(CONFIG_RISCV_CBQRI_RESCTRL_FS) +=3D cbqri_resctrl.o cbqri-$(CONFIG_RISCV_CBQRI_CAPACITY) +=3D cbqri_capacity.o +cbqri-$(CONFIG_RISCV_CBQRI_IOMMU) +=3D cbqri_iommu.o diff --git a/drivers/resctrl/cbqri_iommu.c b/drivers/resctrl/cbqri_iommu.c new file mode 100644 index 000000000000..4086c32546bd --- /dev/null +++ b/drivers/resctrl/cbqri_iommu.c @@ -0,0 +1,276 @@ +// SPDX-License-Identifier: GPL-2.0-only + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include + +#define IOMMU_GROUP_TOKEN "iommu_group:" + +/* + * RISC-V IOMMU RCID and MCID fields are 12 bits each, so an active binding + * cannot contain U64_MAX. + */ +#define QOS_IOMMU_BINDING_UPDATING U64_MAX + +/* Records the resctrl QoS ID assignment for one IOMMU group. */ +struct qos_iommu_group_binding { + struct list_head node; + struct rcu_head rcu; + struct iommu_group *group; + int group_id; + /* Packed IDs, or QOS_IOMMU_BINDING_UPDATING during a hardware update. */ + u64 state; +}; + +/* + * resctrl owns IOMMU group membership bookkeeping. Do not infer membershi= p by + * reading RCID/MCID back from hardware: the same IDs may be shared by mul= tiple + * resctrl users. Writers are serialized by rdtgroup_mutex. IOMMU attach p= aths + * read the list under RCU so newly attached devices can inherit an existi= ng + * group assignment without taking locks in the opposite order to resctrl. + */ +static LIST_HEAD(qos_iommu_bindings); + +static u64 qos_iommu_group_ids_pack(struct resctrl_group_ids ids) +{ + return (u64)ids.closid << 32 | ids.rmid; +} + +static struct resctrl_group_ids qos_iommu_group_ids_unpack(u64 ids) +{ + return (struct resctrl_group_ids) { + .closid =3D upper_32_bits(ids), + .rmid =3D lower_32_bits(ids), + }; +} + +static bool qos_iommu_group_ids_match(u64 packed, struct resctrl_group_ids= ids) +{ + return packed !=3D QOS_IOMMU_BINDING_UPDATING && + packed =3D=3D qos_iommu_group_ids_pack(ids); +} + +static bool qos_iommu_group_ids_default(struct resctrl_group_ids ids) +{ + return ids.closid =3D=3D RESCTRL_RESERVED_CLOSID && + ids.rmid =3D=3D RESCTRL_RESERVED_RMID; +} + +static struct qos_iommu_group_binding * +qos_iommu_group_find_binding(int group_id) +{ + struct qos_iommu_group_binding *binding; + + list_for_each_entry(binding, &qos_iommu_bindings, node) { + if (binding->group_id =3D=3D group_id) + return binding; + } + + return NULL; +} + +static void qos_iommu_group_free_binding(struct qos_iommu_group_binding *b= inding) +{ + list_del_rcu(&binding->node); + iommu_group_put_by_id(binding->group); + kfree_rcu(binding, rcu); +} + +static void qos_iommu_group_prune_inactive(void) +{ + struct qos_iommu_group_binding *binding; + struct qos_iommu_group_binding *tmp; + + list_for_each_entry_safe(binding, tmp, &qos_iommu_bindings, node) { + if (!iommu_group_is_active(binding->group)) + qos_iommu_group_free_binding(binding); + } +} + +static int qos_iommu_group_set(struct iommu_group *group, + struct resctrl_group_ids ids) +{ + /* RISC-V IOMMU QoS maps closid->rcid and rmid->mcid 1:1. */ + return riscv_iommu_group_set_qosid(group, ids.closid, ids.rmid); +} + +int resctrl_arch_devices_write(char *tok, struct resctrl_group_ids ids) +{ + struct qos_iommu_group_binding *binding; + struct qos_iommu_group_binding *new_binding =3D NULL; + struct iommu_group *group; + u64 old_state =3D QOS_IOMMU_BINDING_UPDATING; + bool default_ids; + int group_id; + int ret; + + if (strncmp(tok, IOMMU_GROUP_TOKEN, strlen(IOMMU_GROUP_TOKEN))) { + resctrl_last_cmd_printf("Unsupported device token %s\n", tok); + return -EOPNOTSUPP; + } + + tok +=3D strlen(IOMMU_GROUP_TOKEN); + if (kstrtoint(tok, 10, &group_id) || group_id < 0) { + resctrl_last_cmd_printf("IOMMU group parsing error %s%s\n", + IOMMU_GROUP_TOKEN, tok); + return -EINVAL; + } + + if (!capable(CAP_SYS_ADMIN)) { + resctrl_last_cmd_puts("No permission to move IOMMU group\n"); + return -EPERM; + } + + default_ids =3D qos_iommu_group_ids_default(ids); + qos_iommu_group_prune_inactive(); + binding =3D qos_iommu_group_find_binding(group_id); + if (!binding) { + group =3D iommu_group_get_by_id(group_id); + if (!group) { + resctrl_last_cmd_printf("No IOMMU group %d\n", group_id); + return -ENOENT; + } + + if (default_ids) { + ret =3D qos_iommu_group_set(group, ids); + iommu_group_put_by_id(group); + if (ret) + resctrl_last_cmd_printf("Failed to reset IOMMU group %d QoS (%d)\n", + group_id, ret); + return ret; + } + + new_binding =3D kzalloc_obj(*new_binding, GFP_KERNEL); + if (!new_binding) { + iommu_group_put_by_id(group); + resctrl_last_cmd_puts("Unable to allocate IOMMU group QoS binding\n"); + return -ENOMEM; + } + binding =3D new_binding; + binding->group =3D group; + binding->group_id =3D group_id; + binding->state =3D QOS_IOMMU_BINDING_UPDATING; + list_add_tail_rcu(&binding->node, &qos_iommu_bindings); + } else { + group =3D binding->group; + old_state =3D READ_ONCE(binding->state); + WRITE_ONCE(binding->state, QOS_IOMMU_BINDING_UPDATING); + } + + ret =3D qos_iommu_group_set(group, ids); + if (ret) { + if (new_binding) + qos_iommu_group_free_binding(binding); + else + WRITE_ONCE(binding->state, old_state); + resctrl_last_cmd_printf("Failed to move IOMMU group %d QoS (%d)\n", + group_id, ret); + return ret; + } + + if (default_ids) + qos_iommu_group_free_binding(binding); + else + WRITE_ONCE(binding->state, qos_iommu_group_ids_pack(ids)); + + return 0; +} + +int riscv_resctrl_iommu_group_get_qosid(struct iommu_group *group, u32 *rc= id, + u32 *mcid) +{ + struct qos_iommu_group_binding *binding; + int ret =3D -ENOENT; + + if (!group) + return -ENODEV; + + rcu_read_lock(); + list_for_each_entry_rcu(binding, &qos_iommu_bindings, node) { + struct resctrl_group_ids ids; + u64 state; + + if (binding->group !=3D group) + continue; + state =3D READ_ONCE(binding->state); + if (state =3D=3D QOS_IOMMU_BINDING_UPDATING) { + ret =3D -EAGAIN; + break; + } + + ids =3D qos_iommu_group_ids_unpack(state); + if (rcid) + *rcid =3D ids.closid; + if (mcid) + *mcid =3D ids.rmid; + ret =3D 0; + break; + } + rcu_read_unlock(); + + return ret; +} + +void resctrl_arch_devices_show(struct seq_file *s, struct resctrl_group_id= s ids) +{ + struct qos_iommu_group_binding *binding; + + qos_iommu_group_prune_inactive(); + list_for_each_entry(binding, &qos_iommu_bindings, node) { + if (qos_iommu_group_ids_match(READ_ONCE(binding->state), ids)) + seq_printf(s, "iommu_group:%d\n", binding->group_id); + } +} + +bool resctrl_arch_devices_assigned(struct resctrl_group_ids ids) +{ + struct qos_iommu_group_binding *binding; + + qos_iommu_group_prune_inactive(); + list_for_each_entry(binding, &qos_iommu_bindings, node) { + if (qos_iommu_group_ids_match(READ_ONCE(binding->state), ids)) + return true; + } + + return false; +} + +int resctrl_arch_devices_reset_all(struct resctrl_group_ids default_ids) +{ + struct qos_iommu_group_binding *binding; + struct qos_iommu_group_binding *tmp; + u64 old_state; + int first_ret =3D 0; + int ret; + + qos_iommu_group_prune_inactive(); + list_for_each_entry_safe(binding, tmp, &qos_iommu_bindings, node) { + old_state =3D READ_ONCE(binding->state); + WRITE_ONCE(binding->state, QOS_IOMMU_BINDING_UPDATING); + ret =3D qos_iommu_group_set(binding->group, default_ids); + if (ret) { + WRITE_ONCE(binding->state, old_state); + if (!first_ret) + first_ret =3D ret; + resctrl_last_cmd_printf("Failed to reset IOMMU group %d QoS (%d)\n", + binding->group_id, ret); + continue; + } + + qos_iommu_group_free_binding(binding); + } + + return first_ret; +} --=20 2.50.1 (Apple Git-155) From nobody Sat Jul 25 19:29:10 2026 Received: from mail-ot1-f52.google.com (mail-ot1-f52.google.com [209.85.210.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BD969466B65 for ; Tue, 14 Jul 2026 13:08:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.52 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034520; cv=none; b=L4ylok6I8kev3KyrHyg5hJN8Baoz491ms0+ROffudxnxx8Pl3vuybQOcEUwqS88Cm/G8nNus0S8Mx7pmuKAi1nLMQRVTiV2RTD5PlxIAK2sC6g05LF+m9XZE9elVefkufiKDckWgTdN90aGVVk4SDsSstq7nQbEVo5n19Lx8zYE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784034520; c=relaxed/simple; bh=5Y43Y7fyHYZyPGuirU3OoYI1Mwv1k+B9YukPppo9vUM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=e+gMfxbmywpFOh7J2xT59xfqwoPg/UkXtTB7rBwALMp4w6aGsboPme9Fo3Zc2WZx91Dbh2haGreoZiPGb2dytbsL9QXJZ7LDd2U/cZHABjcTLCIslFErefRNy4XDez38uWJZRcwuqv2PlFAW71+nSZ6zN3+RrIR8MArX7AF0ALs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=aHJRZvhM; arc=none smtp.client-ip=209.85.210.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="aHJRZvhM" Received: by mail-ot1-f52.google.com with SMTP id 46e09a7af769-7ea9c6ea7deso2960845a34.3 for ; Tue, 14 Jul 2026 06:08:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1784034516; x=1784639316; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=emzLt0cm2On/W2Ip4YweG2bRFCgDp+q5PQrHqypf35c=; b=aHJRZvhMwBpG2iZPPm5KYoM8YtLkws5inCZR6EdnsR3rQ0+Z6gmsT/tHYgDJ32YRi1 cINQDkgdNMFKgqOdokd38eXmJQJWK+TH/BfUZpWdqh4tnaRyYZ06HMTon58bJaBFJ7nX Ye9MCh3EfNOubt8f8L3QRHWwnAR0wtGlQjoVaa0UkLQ1RzM3hhhKylCIj0MzF17tZxVV KWEDVRwEHbqprjnvJZVot/PsjiArG/ykDTvJV6YTHmzXynPCTPR3ZPsnnSvYIOAq2IuY YM0qc5u0G9Au1sW9Dgs+zYgzQFcyCp02nm0qiu9Ge0hju8eOVJJuwcjzmEnvUX42Y0ZW z3qg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784034516; x=1784639316; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=emzLt0cm2On/W2Ip4YweG2bRFCgDp+q5PQrHqypf35c=; b=ePWyHdadMRgB8+PVHbEKgsia6srXITagomW0ihndE30EveX0sDWfOXWtQqh9mKWXi6 dxM7XV/UowlaNPFOV+B8grNJ3uMt6yZmfs2/TWMNGryltizyhJkt5ndynuOHNc2ttmeZ fGtFWNECsglafIiOYPJjPFyL4LUR5CmnegRKk3z7WzCfgYQz+ONXO45yI73P8sr0Vokn NyLZ53WPlp3U8N3pQHiQJZT/JUCk2ZtoMkp2OiiHRWcft4PpuCO0QLGRnhRl6/KQmxXp pVYd7RGVFyzyg273DVGrjitmpmEgiICA33D1Ns3N/6hUqrsQtrY8Np7GGR3nYBn5OPeT qI1g== X-Forwarded-Encrypted: i=1; AFNElJ+h6NEG0X0TjxKAmTB5HRCaIBREXt24aaQYgcrdXfT/nuHFJwujHA6U7+FnPt8wQB4++2PYrMtz3TV+5vs=@vger.kernel.org X-Gm-Message-State: AOJu0Yz5Jbp+NJPknc/TICX7Or4MHsY4eGNehKJXdZfqnAvXtV4uZf8K 1BccExj5shaPCrdSXEYU4u5ToUf3hoCKI+gV5bNoU5KYclqRWogE4CSnd2XZnNd+Bfg= X-Gm-Gg: AfdE7ckgSqqi3dAxVbSn6Tm5aIeDX8Ea0ywP3r109lHoQbS3uOxbqe5736OmE3A8FE4 fnJUslyrVVvwcGtjCEdE5/C4qfm7AkQpBlPsF/j+wB5psK8eXeosQfABh/h7jBfgJZsf/bfv5LY edx8BC9iU6zUkjVNgCmohLJSuT0TWzxqDy2f0gXtOR2sTcPUqpyWr1nV0LCQ/GkdmyWcxSxKGI0 XbaTrUA3obWAE2qfSdZoyyA5cWAYCchAOLy1GI/cRmV6lRxr91QMECxYkNngW/gXNMlAoFjlqFt lsbo49Z1F9vedXgtXIhzHPAsEU/hG4bVNs+NNd3ABDqiIMiVlo3wun7+NqN7icgceNceKcOKEX7 l9Mub5YVj9qU7vgFe1rZufNQs4MsMjF4J0WL8Xb1orJsElENI06I+c9XUrIkUPlg3FgIBuTt1P6 xnaOIwmXXeETQFpPiVgSglFYAKVPHvGf9eojGH6jHSjnYvimQ6Mi/Zm2jmcfo7Vg== X-Received: by 2002:a05:6830:4884:b0:7e9:b537:102c with SMTP id 46e09a7af769-7ec4239b1acmr2074489a34.25.1784034516373; Tue, 14 Jul 2026 06:08:36 -0700 (PDT) Received: from FJ7FR2JRQ3.bytedance.net ([178.93.176.7]) by smtp.gmail.com with ESMTPSA id 46e09a7af769-7ebcab8efc3sm14657738a34.0.2026.07.14.06.08.24 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 14 Jul 2026 06:08:35 -0700 (PDT) From: Zhanpeng Zhang To: joro@8bytes.org, palmer@dabbelt.com, tony.luck@intel.com, reinette.chatre@intel.com, tomasz.jeznach@linux.dev Cc: will@kernel.org, robin.murphy@arm.com, fustini@kernel.org, pjw@kernel.org, aou@eecs.berkeley.edu, alex@ghiti.fr, Dave.Martin@arm.com, james.morse@arm.com, babu.moger@amd.com, corbet@lwn.net, shuah@kernel.org, jgg@ziepe.ca, kevin.tian@intel.com, cuiyunhui@bytedance.com, yuanzhu@bytedance.com, iommu@lists.linux.dev, linux-riscv@lists.infradead.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, x86@kernel.org Subject: [RFC PATCH 7/7] selftests/iommu: Add RISC-V IOMMU QoS smoke test Date: Tue, 14 Jul 2026 21:06:57 +0800 Message-ID: <20260714130657.46963-8-zhangzhanpeng.jasper@bytedance.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> References: <20260714130657.46963-1-zhangzhanpeng.jasper@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add a RISC-V IOMMU QoS smoke test for the resctrl devices assignment ABI and the per-IOMMU global qosid sysfs attribute. Probe for a usable IOMMU group with a non-default child-group assignment without disturbing groups already assigned by the system. Cover valid assignment and reset, devices-file readback, malformed and missing group IDs, strict decimal parsing, tasks-file rejection, group removal and pseudo-lock protection, and cleanup before child-group removal. When the global qosid attribute is present, cover its readback format, valid rewrite and restore, malformed input, missing separators, range checking, and integer overflow. Keep child-group coverage optional so an environment without allocatable resctrl resources can still exercise the root IOMMU QoS paths. Signed-off-by: Zhanpeng Zhang --- MAINTAINERS | 1 + tools/testing/selftests/iommu/Makefile | 2 + .../selftests/iommu/iommu_qos_smoke.sh | 649 ++++++++++++++++++ 3 files changed, 652 insertions(+) create mode 100755 tools/testing/selftests/iommu/iommu_qos_smoke.sh diff --git a/MAINTAINERS b/MAINTAINERS index 162ad5a1780f..9facf8ac79a2 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -23289,6 +23289,7 @@ L: linux-riscv@lists.infradead.org S: Maintained F: Documentation/ABI/testing/sysfs-class-iommu-riscv-iommu F: drivers/resctrl/cbqri_iommu.c +F: tools/testing/selftests/iommu/iommu_qos_smoke.sh =20 RISC-V MICROCHIP SUPPORT M: Conor Dooley diff --git a/tools/testing/selftests/iommu/Makefile b/tools/testing/selftes= ts/iommu/Makefile index 84abeb2f0949..2d22a10c17fc 100644 --- a/tools/testing/selftests/iommu/Makefile +++ b/tools/testing/selftests/iommu/Makefile @@ -7,4 +7,6 @@ TEST_GEN_PROGS :=3D TEST_GEN_PROGS +=3D iommufd TEST_GEN_PROGS +=3D iommufd_fail_nth =20 +TEST_PROGS +=3D iommu_qos_smoke.sh + include ../lib.mk diff --git a/tools/testing/selftests/iommu/iommu_qos_smoke.sh b/tools/testi= ng/selftests/iommu/iommu_qos_smoke.sh new file mode 100755 index 000000000000..1a81bca39460 --- /dev/null +++ b/tools/testing/selftests/iommu/iommu_qos_smoke.sh @@ -0,0 +1,649 @@ +#!/bin/sh +# SPDX-License-Identifier: GPL-2.0-only +# Smoke test for RISC-V IOMMU QoS resctrl integration. +# +# This intentionally tests the IOMMU group QoS control path: +# resctrl/devices -> iommu_group: parser -> IOMMU group QoS assignme= nt +# explicit reset to the default group before resctrl group removal +# and the per-IOMMU global default QoS ID sysfs path used by BARE mode. + +KSFT_PASS=3D0 +KSFT_FAIL=3D1 +KSFT_SKIP=3D4 + +RESCTRL=3D/sys/fs/resctrl +RESCTRL_FS=3Dresctrl +IOMMU_GROUPS=3D/sys/kernel/iommu_groups +LAST_STATUS=3D$RESCTRL/info/last_cmd_status +ROOT_DEVICES=3D$RESCTRL/devices +IOMMU_CLASS=3D/sys/class/iommu + +pass_count=3D0 +fail_count=3D0 +skip_count=3D0 +group=3D"" +child_group=3D"" +qosid_file=3D"" +tmp_dir=3D"" +lock_dir=3D"" +lock_taken=3D0 +mounted_by_test=3D0 +qosid_rcid=3D"" +qosid_mcid=3D"" + +log() +{ + printf '%s\n' "$*" +} + +pass() +{ + pass_count=3D$((pass_count + 1)) + log "PASS: $*" +} + +fail() +{ + fail_count=3D$((fail_count + 1)) + log "FAIL: $*" +} + +skip() +{ + log "SKIP: $*" + exit $KSFT_SKIP +} + +skip_optional() +{ + skip_count=3D$((skip_count + 1)) + log "SKIP: $*" +} + +cleanup() +{ + if [ -n "$child_group" ] && [ -d "$child_group" ]; then + if [ -n "$group" ]; then + printf '%s\n' "iommu_group:$group" > "$ROOT_DEVICES" 2>/dev/null || true + fi + rmdir "$child_group" 2>/dev/null || true + fi + + if [ "$mounted_by_test" =3D "1" ]; then + umount "$RESCTRL" 2>/dev/null || true + fi + + if [ "$lock_taken" =3D "1" ] && [ -n "$lock_dir" ]; then + rmdir "$lock_dir" 2>/dev/null || true + fi + + if [ -n "$tmp_dir" ]; then + rm -rf "$tmp_dir" + fi +} + +need_root() +{ + [ "$(id -u)" =3D "0" ] || skip "requires root" +} + +init_tmp_dir() +{ + tmp_dir=3D$(mktemp -d "${TMPDIR:-/tmp}/iommu_qos_smoke.XXXXXX") || + skip "failed to create temporary directory" +} + +take_global_lock() +{ + lock_dir=3D${TMPDIR:-/tmp}/iommu_qos_smoke.lock + + if mkdir "$lock_dir" 2>"$tmp_dir/lock.err"; then + lock_taken=3D1 + pass "acquired global IOMMU QoS test lock" + return + fi + + skip "another IOMMU QoS smoke test instance is running" +} + +ensure_resctrl_mounted() +{ + if ! grep -qw resctrl /proc/filesystems; then + skip "resctrl filesystem is not available" + fi + + mkdir -p "$RESCTRL" 2>/dev/null || true + + if grep -qs " $RESCTRL resctrl " /proc/mounts; then + pass "resctrl already mounted" + return + fi + + if mount -t resctrl "$RESCTRL_FS" "$RESCTRL" 2>"$tmp_dir/mount.err"; then + mounted_by_test=3D1 + pass "mounted resctrl" + else + log "mount error: $(cat "$tmp_dir/mount.err" 2>/dev/null)" + skip "failed to mount resctrl" + fi +} + +ensure_devices_file() +{ + [ -e "$ROOT_DEVICES" ] || skip "resctrl devices file is not available" + [ -w "$ROOT_DEVICES" ] || skip "$ROOT_DEVICES is not writable" +} + +create_child_group() +{ + child_group=3D$RESCTRL/iommu_qos_test_$$ + + if mkdir "$child_group" 2>"$tmp_dir/mkdir.err"; then + pass "created child resctrl group" + return + fi + + log "mkdir error: $(cat "$tmp_dir/mkdir.err" 2>/dev/null)" + child_group=3D"" + skip "failed to create a child resctrl group" +} + +iommu_group_is_assigned() +{ + candidate=3D$1 + + find "$RESCTRL" -type f -name devices \ + -exec grep -qx "iommu_group:$candidate" {} \; -print \ + 2>/dev/null | grep -q . +} + +pick_iommu_group() +{ + populated=3D0 + unsupported=3D0 + + [ -d "$IOMMU_GROUPS" ] || skip "no IOMMU groups sysfs directory" + + for path in "$IOMMU_GROUPS"/*; do + [ -d "$path" ] || continue + case "${path##*/}" in + *[!0-9]*|'') + continue + ;; + esac + + [ -d "$path/devices" ] || continue + if ! find "$path/devices" -mindepth 1 -maxdepth 1 2>/dev/null | + grep -q .; then + continue + fi + + populated=3D$((populated + 1)) + candidate=3D${path##*/} + if iommu_group_is_assigned "$candidate"; then + log "skip IOMMU group $candidate with an existing resctrl assignment" + continue + fi + + if printf '%s\n' "iommu_group:$candidate" > "$child_group/devices" \ + 2>"$tmp_dir/group_$candidate.err"; then + group=3D$candidate + if ! printf '%s\n' "iommu_group:$group" > "$ROOT_DEVICES" \ + 2>"$tmp_dir/group_${candidate}_reset.err"; then + log "reset error:" + cat "$tmp_dir/group_${candidate}_reset.err" 2>/dev/null + fail "failed to restore probed IOMMU group $group" + return 1 + fi + pass "selected group $group using non-default QoS assignment" + return 0 + fi + + if last_status_is_iommu_qos_unsupported; then + unsupported=3D$((unsupported + 1)) + continue + fi + + log "group $candidate write error:" + cat "$tmp_dir/group_$candidate.err" 2>/dev/null + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "IOMMU group $candidate failed QoS capability probing" + done + + [ "$populated" -gt 0 ] || skip "no populated IOMMU group found" + if [ "$unsupported" -eq "$populated" ]; then + skip "no populated IOMMU group supports QoS ID programming" + fi + + return 1 +} + +last_status_contains() +{ + pattern=3D$1 + + [ -f "$LAST_STATUS" ] || return 1 + grep -qi "$pattern" "$LAST_STATUS" +} + +last_status_is_iommu_qos_unsupported() +{ + [ -f "$LAST_STATUS" ] || return 1 + grep -Eiq \ + 'IOMMU group .* QoS \(-95\)|IOMMU QoSID.*not supported|not supported' \ + "$LAST_STATUS" +} + +pick_missing_iommu_group() +{ + max=3D0 + + for path in "$IOMMU_GROUPS"/*; do + [ -d "$path" ] || continue + id=3D${path##*/} + case "$id" in + *[!0-9]*|'') + continue + ;; + esac + if [ "$id" -gt "$max" ]; then + max=3D$id + fi + done + + if [ "$max" -ge 2147483647 ]; then + skip "cannot choose a missing IOMMU group id" + fi + + missing=3D$((max + 1)) + while [ -e "$IOMMU_GROUPS/$missing" ]; do + if [ "$missing" -ge 2147483647 ]; then + skip "cannot choose a missing IOMMU group id" + fi + missing=3D$((missing + 1)) + done +} + +pick_qosid_file() +{ + for file in "$IOMMU_CLASS"/*/qosid; do + [ -e "$file" ] || continue + qosid_file=3D$file + return 0 + done + + return 1 +} + +parse_qosid() +{ + line=3D$1 + + case "$line" in + rcid=3D[0-9]*" "mcid=3D[0-9]*) + qosid_rcid=3D${line%% *} + qosid_rcid=3D${qosid_rcid#rcid=3D} + qosid_mcid=3D${line##* } + qosid_mcid=3D${qosid_mcid#mcid=3D} + ;; + *) + return 1 + ;; + esac + + case "$qosid_rcid" in + *[!0-9]*|'') + return 1 + ;; + esac + case "$qosid_mcid" in + *[!0-9]*|'') + return 1 + ;; + esac + + return 0 +} + +test_malformed_group_rejected() +{ + if printf '%s\n' 'iommu_group:bad' > "$ROOT_DEVICES" \ + 2>"$tmp_dir/badtoken.err"; then + fail "malformed IOMMU group token was accepted" + return + fi + + if last_status_contains "IOMMU group parsing error"; then + pass "malformed IOMMU group token is rejected with useful status" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "malformed IOMMU group rejection status is unclear" + fi +} + +test_trailing_separator_rejected() +{ + if printf '%s\n' "iommu_group:$group," > "$ROOT_DEVICES" \ + 2>"$tmp_dir/trailing_separator.err"; then + fail "IOMMU group token with trailing separator was accepted" + return + fi + + if last_status_contains "Device list parsing error"; then + pass "trailing device-list separator is rejected" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "trailing separator rejection status is unclear" + fi +} + +test_missing_group_rejected() +{ + pick_missing_iommu_group + + if printf '%s\n' "iommu_group:$missing" > "$ROOT_DEVICES" \ + 2>"$tmp_dir/missing.err"; then + fail "missing IOMMU group was accepted" + return + fi + + if last_status_contains "No IOMMU group $missing"; then + pass "missing IOMMU group is rejected with useful status" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "missing IOMMU group rejection status is unclear" + fi +} + +test_non_decimal_group_rejected() +{ + if printf '%s\n' "iommu_group:0x$group" > "$ROOT_DEVICES" \ + 2>"$tmp_dir/non_decimal.err"; then + fail "non-decimal IOMMU group ID was accepted" + return + fi + + if last_status_contains "IOMMU group parsing error"; then + pass "non-decimal IOMMU group ID is rejected" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "non-decimal IOMMU group rejection status is unclear" + fi +} + +test_tasks_rejects_iommu_token() +{ + if printf '%s\n' "iommu_group:$group" > "$RESCTRL/tasks" \ + 2>"$tmp_dir/tasks_token.err"; then + fail "tasks accepted an IOMMU group token" + return + fi + + if last_status_contains "Task list parsing error pid iommu_group:$group";= then + pass "tasks rejects IOMMU group tokens" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "tasks rejection status is unclear" + fi +} + +test_global_qosid_sysfs() +{ + if ! pick_qosid_file; then + skip_optional "no RISC-V IOMMU qosid sysfs file is available" + return + fi + + if ! line=3D$(cat "$qosid_file" 2>"$tmp_dir/qosid_read.err"); then + log "read error: $(cat "$tmp_dir/qosid_read.err" 2>/dev/null)" + fail "failed to read IOMMU global qosid" + return + fi + + if parse_qosid "$line"; then + pass "read IOMMU global qosid from sysfs" + else + log "qosid content: $line" + fail "IOMMU global qosid has unexpected format" + return + fi + + if printf '%s\n' "$line" > "$qosid_file" \ + 2>"$tmp_dir/qosid_rewrite.err"; then + pass "IOMMU global qosid accepts readback format" + else + log "write error: $(cat "$tmp_dir/qosid_rewrite.err" 2>/dev/null)" + fail "failed to rewrite IOMMU global qosid readback" + return + fi + + if printf '%s\n' "bad" > "$qosid_file" 2>"$tmp_dir/qosid_bad.err"; then + fail "IOMMU global qosid accepted malformed input" + return + fi + pass "IOMMU global qosid rejects malformed input" + + if printf '%s\n' "rcid=3D0mcid=3D1" > "$qosid_file" \ + 2>"$tmp_dir/qosid_separator.err"; then + fail "IOMMU global qosid accepted missing token separator" + return + fi + pass "IOMMU global qosid requires a token separator" + + if printf '%s\n' "rcid=3D4096 mcid=3D0" > "$qosid_file" \ + 2>"$tmp_dir/qosid_range.err"; then + fail "IOMMU global qosid accepted out-of-range RCID" + return + fi + pass "IOMMU global qosid rejects out-of-range RCID" + + if printf '%s\n' "rcid=3D4294967296 mcid=3D0" > "$qosid_file" \ + 2>"$tmp_dir/qosid_overflow.err"; then + fail "IOMMU global qosid accepted overflowed RCID" + return + fi + pass "IOMMU global qosid rejects overflowed RCID" + + if printf '%s\n' "rcid=3D$qosid_rcid mcid=3D$qosid_mcid" > "$qosid_file" \ + 2>"$tmp_dir/qosid_restore.err"; then + pass "IOMMU global qosid accepts current RCID/MCID" + else + log "write error: $(cat "$tmp_dir/qosid_restore.err" 2>/dev/null)" + fail "failed to write current IOMMU global qosid" + return + fi + + if line=3D$(cat "$qosid_file" 2>/dev/null) && + [ "$line" =3D "rcid=3D$qosid_rcid mcid=3D$qosid_mcid" ]; then + pass "IOMMU global qosid readback matches restored value" + else + log "qosid content: $line" + fail "IOMMU global qosid readback mismatch" + fi +} + +test_valid_group_write() +{ + if printf '%s\n' "iommu_group:$group" > "$ROOT_DEVICES" \ + 2>"$tmp_dir/valid.err"; then + pass "write iommu_group:$group to root resctrl devices" + else + log "write error: $(cat "$tmp_dir/valid.err" 2>/dev/null)" + fail "valid IOMMU group write failed" + return + fi + + if last_status_contains ^ok; then + pass "last_cmd_status reports ok" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "last_cmd_status is not ok after valid write" + fi + + # The root control group represents the default QoS IDs. The kernel + # applies the reset but intentionally does not keep a software binding + # for default assignments, so the root devices file should stay empty. + if grep -qx "iommu_group:$group" "$ROOT_DEVICES"; then + log "devices content:" + cat "$ROOT_DEVICES" 2>/dev/null + fail "root devices unexpectedly shows default iommu_group:$group" + else + pass "root devices omits default iommu_group:$group" + fi + +} + +test_child_group_write() +{ + if printf '%s\n' "iommu_group:$group" > "$child_group/devices" \ + 2>"$tmp_dir/child.err"; then + pass "write iommu_group:$group to child resctrl group" + else + log "write error: $(cat "$tmp_dir/child.err" 2>/dev/null)" + fail "valid child IOMMU group write failed" + return + fi + + if last_status_contains ^ok; then + pass "last_cmd_status reports ok after child group write" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "last_cmd_status is not ok after child group write" + fi + + if grep -qx "iommu_group:$group" "$child_group/devices"; then + pass "child devices shows iommu_group:$group after set/get" + else + log "child devices content:" + cat "$child_group/devices" 2>/dev/null + fail "child devices does not show iommu_group:$group" + fi + + if grep -qx "iommu_group:$group" "$ROOT_DEVICES"; then + log "root devices content:" + cat "$ROOT_DEVICES" 2>/dev/null + fail "root devices still shows iommu_group:$group after child move" + else + pass "root devices no longer shows iommu_group:$group after child move" + fi + +} + +test_child_pseudo_locksetup_rejected() +{ + [ -d "$child_group" ] || return + + if [ ! -e "$child_group/mode" ]; then + skip_optional "resctrl mode file is not available" + return + fi + + if printf '%s\n' "pseudo-locksetup" > "$child_group/mode" \ + 2>"$tmp_dir/pseudo_locksetup.err"; then + printf '%s\n' "shareable" > "$child_group/mode" 2>/dev/null || true + fail "pseudo-locksetup accepted a group with assigned devices" + return + fi + + if last_status_contains \ + "Move devices out before entering pseudo-locksetup group"; then + pass "pseudo-locksetup rejects group with assigned devices" + elif last_status_contains "Unknown or unsupported mode"; then + skip_optional "pseudo-locksetup mode is not supported" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "pseudo-locksetup rejection status is unclear" + fi +} + +test_child_group_rmdir() +{ + [ -d "$child_group" ] || return + + if rmdir "$child_group" 2>"$tmp_dir/rmdir_busy.err"; then + fail "rmdir child resctrl group with IOMMU group was accepted" + # The child group no longer exists, so cleanup() cannot reset this. + if ! printf '%s\n' "iommu_group:$group" > "$ROOT_DEVICES" \ + 2>"$tmp_dir/rmdir_unexpected_move_root.err"; then + log "move error after unexpected rmdir:" + cat "$tmp_dir/rmdir_unexpected_move_root.err" 2>/dev/null + fi + child_group=3D"" + return + fi + + if last_status_contains "Move devices out before removing group"; then + pass "rmdir rejects child group while IOMMU group is assigned" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "rmdir rejection status is unclear" + fi + + if printf '%s\n' "iommu_group:$group" > "$ROOT_DEVICES" \ + 2>"$tmp_dir/rmdir_move_root.err"; then + pass "move iommu_group:$group back to root before rmdir" + else + log "move error: $(cat "$tmp_dir/rmdir_move_root.err" 2>/dev/null)" + fail "failed to move iommu_group:$group back to root" + return + fi + + if rmdir "$child_group" 2>"$tmp_dir/rmdir.err"; then + pass "rmdir child resctrl group after moving IOMMU group out" + else + log "rmdir error: $(cat "$tmp_dir/rmdir.err" 2>/dev/null)" + fail "failed to rmdir child resctrl group after moving IOMMU group out" + return + fi + child_group=3D"" + + if last_status_contains ^ok; then + pass "last_cmd_status reports ok after child group rmdir" + else + log "last_cmd_status: $(cat "$LAST_STATUS" 2>/dev/null)" + fail "last_cmd_status is not ok after child group rmdir" + fi + + if grep -qx "iommu_group:$group" "$ROOT_DEVICES"; then + log "root devices content:" + cat "$ROOT_DEVICES" 2>/dev/null + fail "root devices still shows iommu_group:$group after default move" + else + pass "root devices does not show iommu_group:$group after default move" + fi + +} + +trap cleanup EXIT +need_root +init_tmp_dir +take_global_lock +ensure_resctrl_mounted +ensure_devices_file +create_child_group +if ! pick_iommu_group; then + log "FAIL: $fail_count failure(s), $pass_count pass(es)" + exit $KSFT_FAIL +fi +test_malformed_group_rejected +test_trailing_separator_rejected +test_missing_group_rejected +test_non_decimal_group_rejected +test_tasks_rejects_iommu_token +test_global_qosid_sysfs +test_valid_group_write +test_child_group_write +test_child_pseudo_locksetup_rejected +test_child_group_rmdir + +if [ "$fail_count" -gt 0 ]; then + log "FAIL: $fail_count failure(s), $pass_count pass(es)" + exit $KSFT_FAIL +fi + +if [ "$skip_count" -gt 0 ]; then + log "PASS: $pass_count checks passed, $skip_count optional check(s) skipp= ed" +else + log "PASS: $pass_count checks passed" +fi +exit $KSFT_PASS --=20 2.50.1 (Apple Git-155)