From nobody Mon Sep 28 17:48:02 2026 Received: from mail-wm1-f69.google.com (mail-wm1-f69.google.com [209.85.128.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 56C7030675F for ; Wed, 19 Aug 2026 12:43:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143440; cv=none; b=OnkFqAcrbkx7tYuYfxcXL+cVSbmtu/0RlyOwqf5dUNEuy2AyM5m05zxyJl6m7IolQEQOAXVH17JdotK11Y0l0LuliYU0RiHB5JBZcesAWRBOSIjOEVRD2jdKUiGUqgimMUF8EYCPXO6joaLzwpxh6w9fFwWXloh9hEE84lHV1HY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143440; c=relaxed/simple; bh=yA4pjtDTrTuCsLf3b7pK2t0nVf64PN9VJJskTkM6KDc=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=AAvnitBAxz4BLc9owx3wZIBSJIzCWaX620hIEaXQtBtPm3G9ocYFarO4S96oQfZLePnwKgPXRL1EceS9Z+8eSYF+kNZZ0O7EotKaHSj50QMW6sVo+r8DpCu5KoETmhuXZrjPq28gGWBE4v1Nfr3DJdNUVwAk7W6fyPcNA1zZ/nk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=pDVTfqO6; arc=none smtp.client-ip=209.85.128.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="pDVTfqO6" Received: by mail-wm1-f69.google.com with SMTP id 5b1f17b1804b1-4954b1c6310so9116695e9.0 for ; Wed, 19 Aug 2026 05:43:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787143436; x=1787748236; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=13fKMAmOqnBUwt5r75YrAxL9KZ/M6jhE1nqTXhbmCgQ=; b=pDVTfqO6qu5cVRCd0yTSo9z5GT9/YwbXAsZahbry3oKBJWX7CEBuymv6yh/Z24Ff0C fQJU8dSlP4sF5YhmCZTi83Zh9HpfuiveK/1grghf4B431iui/GbaTN5q2ojpd2lNDu2b oGaYNB1HG5mhl2qeXWlyXf6m4Yx7KO+Ij0jjp5A7nuHbBpiKgL9sjJmkT4SwdWmDXHBF Enr7rAuHmoas/YsjHotQnrzOCjjelIkMB/Stc4khHV9Rq3YE+LY7Zgi7fgwasir1sbra zXFAsicqsKG3zeL5s5ZR2LJmnCqd9nJ3NLJSCiFydwEqpKFK7ceMJ3yJBf9PGWi7T19x ho5Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787143436; x=1787748236; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=13fKMAmOqnBUwt5r75YrAxL9KZ/M6jhE1nqTXhbmCgQ=; b=LOGs1vEIDsMGIxNZmxUGwBg9OhI6jZYPOvo/jDSsIv7ZefTpMhtXtiEeSpkNPAV6uA IZ0fxnZecUBWVufHcrHWN7jJByEMYmZJM7cQbEGTCd7ELaxAo9jhoAm7GQlVGXqTb8GH bEYaFctDnCdLEs+IMmF7tFqpn4z4l8EUSC+1xyNI3C4u0xE3RN9Q7TUoGKOpsZH+exxX OArAc3VKMiz7DQVG8D73y5GVmIAqgHKWsHUtX87zJzpUp0aWMNm1FDlKvGcth/qabTL8 +f3R0R96NQhPWrVaF0mFddzhe8ldrrWt2mReA/ESu4gmTyfX9RMnYj5fMhl0HlJFsQ6w gJPw== X-Gm-Message-State: AOJu0YxHQGOBJ0jT1VjpGpFsT/m/xZFdNkShoanHaFAdHHeOmj/Bpxan xa4pMmwRExZ0ah1L4YHMIJX/5TRFb8NUM8RZ1LhqlKPRw2G8Tla+G2LqbZnpjMX5b6qyCl7KJNT +t3kZGw== X-Received: from wrzq17.prod.google.com ([2002:a05:6000:1371:b0:47f:c79e:e0e3]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a05:600c:8b27:b0:493:c47f:3c55 with SMTP id 5b1f17b1804b1-499aa199601mr92606105e9.5.1787143436358; Wed, 19 Aug 2026 05:43:56 -0700 (PDT) Date: Wed, 19 Aug 2026 12:43:35 +0000 In-Reply-To: <20260819124341.4185621-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260819124341.4185621-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.737.g08866a6d13-goog Message-ID: <20260819124341.4185621-2-lrizzo@google.com> Subject: [PATCH v5 1/7] genirq: Add flags for software interrupt moderation. From: Luigi Rizzo To: Thomas Gleixner , Marc Zyngier , Luigi Rizzo , Paolo Abeni Cc: linux-kernel@vger.kernel.org, linux-pci@vger.kernel.org, Bjorn Helgaas , Luigi Rizzo Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add two flags to support software interrupt moderation: - IRQ_MODERATABLE is an irqdesc flag indicating that an interrupt supports moderation. This is a capability that can be set for eligible sources. - IRQD_MODERATED is an internal irqdata flag indicating that the interrupt is currently being moderated. This is a state flag. Signed-off-by: Luigi Rizzo --- include/linux/irq.h | 11 ++++++++++- kernel/irq/debugfs.c | 3 +++ kernel/irq/settings.h | 17 +++++++++++++++++ 3 files changed, 30 insertions(+), 1 deletion(-) diff --git a/include/linux/irq.h b/include/linux/irq.h index f485369b1b4f7..232df18af1b66 100644 --- a/include/linux/irq.h +++ b/include/linux/irq.h @@ -76,6 +76,7 @@ enum irqchip_irq_state; * IRQ_DISABLE_UNLAZY - Disable lazy irq disable * IRQ_HIDDEN - Don't show up in /proc/interrupts * IRQ_NO_DEBUG - Exclude from note_interrupt() debugging + * IRQ_MODERATABLE - Can do software interrupt moderation */ enum { IRQ_TYPE_NONE =3D 0x00000000, @@ -104,13 +105,14 @@ enum { IRQ_HIDDEN =3D (1 << 20), IRQ_NO_DEBUG =3D (1 << 21), IRQ_RESERVED =3D (1 << 22), + IRQ_MODERATABLE =3D (1 << 23), }; =20 #define IRQF_MODIFY_MASK \ (IRQ_TYPE_SENSE_MASK | IRQ_NOPROBE | IRQ_NOREQUEST | \ IRQ_NOAUTOEN | IRQ_LEVEL | IRQ_NO_BALANCING | \ IRQ_PER_CPU | IRQ_NESTED_THREAD | IRQ_NOTHREAD | IRQ_PER_CPU_DEVID | \ - IRQ_IS_POLLED | IRQ_DISABLE_UNLAZY | IRQ_HIDDEN) + IRQ_IS_POLLED | IRQ_DISABLE_UNLAZY | IRQ_HIDDEN | IRQ_MODERATABLE) =20 #define IRQ_NO_BALANCING_MASK (IRQ_PER_CPU | IRQ_NO_BALANCING) =20 @@ -224,6 +226,7 @@ struct irq_data { * irqchip have flag IRQCHIP_ENABLE_WAKEUP_ON_SUSPEND set. * IRQD_RESEND_WHEN_IN_PROGRESS - Interrupt may fire when already in progr= ess in which * case it must be resent at the next available opportunity. + * IRQD_MODERATED - Interrupt is currently moderated. */ enum { IRQD_TRIGGER_MASK =3D 0xf, @@ -249,6 +252,7 @@ enum { IRQD_AFFINITY_ON_ACTIVATE =3D BIT(28), IRQD_IRQ_ENABLED_ON_SUSPEND =3D BIT(29), IRQD_RESEND_WHEN_IN_PROGRESS =3D BIT(30), + IRQD_MODERATED =3D BIT(31), }; =20 #define __irqd_to_state(d) ACCESS_PRIVATE((d)->common, state_use_accessors) @@ -438,6 +442,11 @@ static inline bool irqd_needs_resend_when_in_progress(= struct irq_data *d) return __irqd_to_state(d) & IRQD_RESEND_WHEN_IN_PROGRESS; } =20 +static inline bool irqd_is_moderated(struct irq_data *d) +{ + return __irqd_to_state(d) & IRQD_MODERATED; +} + #undef __irqd_to_state =20 static inline irq_hw_number_t irqd_to_hwirq(struct irq_data *d) diff --git a/kernel/irq/debugfs.c b/kernel/irq/debugfs.c index 5c5ebaee35f2c..634f00c0f0f59 100644 --- a/kernel/irq/debugfs.c +++ b/kernel/irq/debugfs.c @@ -128,6 +128,8 @@ static const struct irq_bit_descr irqdata_states[] =3D { BIT_MASK_DESCR(IRQD_IRQ_ENABLED_ON_SUSPEND), =20 BIT_MASK_DESCR(IRQD_RESEND_WHEN_IN_PROGRESS), + + BIT_MASK_DESCR(IRQD_MODERATED), }; =20 static const struct irq_bit_descr irqdesc_states[] =3D { @@ -140,6 +142,7 @@ static const struct irq_bit_descr irqdesc_states[] =3D { BIT_MASK_DESCR(_IRQ_IS_POLLED), BIT_MASK_DESCR(_IRQ_DISABLE_UNLAZY), BIT_MASK_DESCR(_IRQ_HIDDEN), + BIT_MASK_DESCR(_IRQ_MODERATABLE), }; =20 static const struct irq_bit_descr irqdesc_istates[] =3D { diff --git a/kernel/irq/settings.h b/kernel/irq/settings.h index 0a0c027a5d342..1663bb46d3a99 100644 --- a/kernel/irq/settings.h +++ b/kernel/irq/settings.h @@ -19,6 +19,7 @@ enum { _IRQ_HIDDEN =3D IRQ_HIDDEN, _IRQ_NO_DEBUG =3D IRQ_NO_DEBUG, _IRQ_PROC_VALID =3D IRQ_RESERVED, + _IRQ_MODERATABLE =3D IRQ_MODERATABLE, _IRQF_MODIFY_MASK =3D IRQF_MODIFY_MASK, }; =20 @@ -36,6 +37,7 @@ enum { #define IRQ_HIDDEN GOT_YOU_MORON #define IRQ_NO_DEBUG GOT_YOU_MORON #define IRQ_RESERVED GOT_YOU_MORON +#define IRQ_MODERATABLE GOT_YOU_MORON #undef IRQF_MODIFY_MASK #define IRQF_MODIFY_MASK GOT_YOU_MORON =20 @@ -193,3 +195,18 @@ static inline void irq_settings_update_proc_valid(stru= ct irq_desc *desc, u32 set desc->status_use_accessors &=3D ~_IRQ_PROC_VALID; desc->status_use_accessors |=3D (set & _IRQ_PROC_VALID); } + +static inline bool irq_settings_moderatable(struct irq_desc *desc) +{ + return desc->status_use_accessors & _IRQ_MODERATABLE; +} + +static inline void irq_settings_set_moderatable(struct irq_desc *desc) +{ + desc->status_use_accessors |=3D _IRQ_MODERATABLE; +} + +static inline void irq_settings_clr_moderatable(struct irq_desc *desc) +{ + desc->status_use_accessors &=3D ~_IRQ_MODERATABLE; +} --=20 2.55.0.737.g08866a6d13-goog From nobody Mon Sep 28 17:48:02 2026 Received: from mail-ed1-f72.google.com (mail-ed1-f72.google.com [209.85.208.72]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B6B7B2D6E6C for ; Wed, 19 Aug 2026 12:43:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.72 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143441; cv=none; b=oz3FgxwE7SW0ZrT7ZEaXgvI/8qBJVpKSNl3ntXmCpJ9Map2xsnF0I0Fa9OmvBBIoz8FJ0acAoZ3ytLyPqi1Clw9WMsAf7QXsk8eANUyjOorzavFjElRbW51nhxQ3xmuB4DcXR9mcfPWCAIEDQxnIriPgOaOQuMJw8LsM1d/QLWw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143441; c=relaxed/simple; bh=kCVyzCwcZ6ZcbCcDFgtNaMRirAbiQLNUcJ9/50s9HXc=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=BaIWprS/k5DEz+7cQM3JZdxzcCZx3vHOv0IZuLxfBOqfg78W36lMt8CYko5vGMuahzSEC1iFgru2YjTA8T+RxWauwcqpuxdhBy9zi5EBXnaTrlDwLcFkWFQnxy+gKBu9Q+6LEdVKLekbka/688OzSkPIu83YhLaySa/mCIPNHno= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=HfXvdPcv; arc=none smtp.client-ip=209.85.208.72 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="HfXvdPcv" Received: by mail-ed1-f72.google.com with SMTP id 4fb4d7f45d1cf-6a18e266026so1203570a12.3 for ; Wed, 19 Aug 2026 05:43:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787143438; x=1787748238; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=lOZr8zYhx3EDw1Z6Z0FGemAoEG9B6UahqKdYX8gOB9w=; b=HfXvdPcv1hvD993PiSfBTFlJcjGhn2XxQ6+183Isa9kx0dgYr5KBteF4cw38MEqglg aVFyKJ7r8qUCTbTY8IX2gqTGmF2zX5kDyw1zPcNxBlyGY2zYxZ7BfzF2ek9wrhc0907+ 5vWgYELBfAxauXRWxLW2vYoWN4rs+YpCU4JTCkwQIM0TRZU6UtohVfLiLcsmlTe1mteQ cEuA5El44L+Fsm+c9q5i0+dIg22YspErfAJ6lgbDguwj89r18325F99uC6MC5eCsES6G h19KEq9HJ0SlWPLLde2MK23MhmUOZBHHdC4VTOyC/FJTDtqkeoRvLms9AnFki2+lUOx/ FEIw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787143438; x=1787748238; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=lOZr8zYhx3EDw1Z6Z0FGemAoEG9B6UahqKdYX8gOB9w=; b=ZhM2t+wHEC279zPg8taZBTXn6ilRfHHDvdCloe86nUuAS5eo4SxtfC9ZXq1FMSwTz3 F2Jx9k1Il5g8W3ZggIHpxo2kmgoX1k1CvupjXfxi62px8LqisdWx6quXBydlwcKfj6SG WjN2Os7ruxa0DhYqae2588lQzCBrCq3vxqZ8nIoEolvK9oWEBFmmUa5sOy6UybFp1MwF fOQ8iPcNOL4rFHg9/16Tj4bGS1FxZNIEISVZPwBjLROc8H32/+uDNFOfqcv1LU5b2pqX bQpNRNF79xcOm+8Z2cMAcNPt3ztI+YV0FaWP5liKv7lo2ZdPsPtntY+qrZ19LoTmtBrN mM1Q== X-Gm-Message-State: AOJu0YxHDX9AYzNQqf++Yweq/QJe48k6Pw7noBb3dvWnZ2REhQ8Ch+lO dULgoY+KjROOlykFykIhkpLfhhYwEamE6eaqHI41C7hpjtQV0wBLoive6wVzAFn99pmOe/D7s90 e0+0u2g== X-Received: from edoy17.prod.google.com ([2002:aa7:c251:0:b0:69a:a4ce:cc7d]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6402:1f4f:b0:6a1:2400:baea with SMTP id 4fb4d7f45d1cf-6a4032f8e7dmr2997486a12.10.1787143437426; Wed, 19 Aug 2026 05:43:57 -0700 (PDT) Date: Wed, 19 Aug 2026 12:43:36 +0000 In-Reply-To: <20260819124341.4185621-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260819124341.4185621-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.737.g08866a6d13-goog Message-ID: <20260819124341.4185621-3-lrizzo@google.com> Subject: [PATCH v5 2/7] genirq: Add GSIM infrastructure From: Luigi Rizzo To: Thomas Gleixner , Marc Zyngier , Luigi Rizzo , Paolo Abeni Cc: linux-kernel@vger.kernel.org, linux-pci@vger.kernel.org, Bjorn Helgaas , Luigi Rizzo Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Introduce the Kconfig option CONFIG_IRQ_SW_MODERATION and GSIM fields in struct irq_desc to support Global Software Interrupt Moderation (GSIM). Note: This commit only introduces the configuration option and basic data structures. The core moderation logic, procfs interface, and integration into the interrupt flow are implemented in subsequent commits. Enabling this option at this stage is a build-safe no-op. Signed-off-by: Luigi Rizzo --- include/linux/irqdesc.h | 12 ++++++++++++ kernel/irq/Kconfig | 11 +++++++++++ kernel/irq/internals.h | 8 ++++++++ kernel/irq/irqdesc.c | 1 + 4 files changed, 32 insertions(+) diff --git a/include/linux/irqdesc.h b/include/linux/irqdesc.h index 8080db17c1b1b..5ba130e224167 100644 --- a/include/linux/irqdesc.h +++ b/include/linux/irqdesc.h @@ -18,6 +18,16 @@ struct irq_desc; struct irq_domain; struct pt_regs; =20 +/** + * struct irq_desc_swmod - software interrupt moderation state + * @swmod_node: list head for per-CPU list of moderated irq_desc + */ +struct irq_desc_swmod { +#ifdef CONFIG_IRQ_SW_MODERATION + struct list_head swmod_node; +#endif +}; + /** * struct irqstat - interrupt statistics * @cnt: real-time interrupt count @@ -59,6 +69,7 @@ struct irq_redirect { * @threads_handled_last: comparator field for deferred spurious detection= of threaded handlers * @lock: locking for SMP * @redirect: Facility for redirecting interrupts via irq_work + * @swmod_state: software interrupt moderation state * @affinity_hint: hint to user space for preferred irq affinity * @affinity_notify: context for notification of affinity changes * @pending_mask: pending rebalanced interrupts @@ -95,6 +106,7 @@ struct irq_desc { atomic_t threads_handled; int threads_handled_last; raw_spinlock_t lock; + struct irq_desc_swmod swmod_state; struct cpumask *percpu_enabled; #ifdef CONFIG_SMP struct irq_redirect redirect; diff --git a/kernel/irq/Kconfig b/kernel/irq/Kconfig index 05cba4e16dad1..d8d556b4a279d 100644 --- a/kernel/irq/Kconfig +++ b/kernel/irq/Kconfig @@ -150,6 +150,17 @@ config IRQ_KUNIT_TEST =20 If unsure, say N. =20 +config IRQ_SW_MODERATION + bool "Enable Global Software Interrupt Moderation" + depends on PROC_FS + help + Enable Global Software Interrupt Moderation. + Uses a local timer to delay interrupts in configurable ways + and depending on various global system load indicators + and targets. + + If you don't know what to do here, say N. + endmenu =20 config GENERIC_IRQ_MULTI_HANDLER diff --git a/kernel/irq/internals.h b/kernel/irq/internals.h index 0ce21dd454047..716a7c2e3633a 100644 --- a/kernel/irq/internals.h +++ b/kernel/irq/internals.h @@ -395,3 +395,11 @@ static inline struct irq_data *irqd_get_parent_data(st= ruct irq_data *irqd) return NULL; #endif } +#ifdef CONFIG_IRQ_SW_MODERATION +static inline void irq_moderation_init_fields(struct irq_desc *desc) +{ + INIT_LIST_HEAD(&desc->swmod_state.swmod_node); +} +#else +static inline void irq_moderation_init_fields(struct irq_desc *desc) {} +#endif diff --git a/kernel/irq/irqdesc.c b/kernel/irq/irqdesc.c index 80ef4e27dcf47..e8488c9352d0c 100644 --- a/kernel/irq/irqdesc.c +++ b/kernel/irq/irqdesc.c @@ -120,6 +120,7 @@ static inline void free_masks(struct irq_desc *desc) { } static void desc_set_defaults(unsigned int irq, struct irq_desc *desc, int= node, const struct cpumask *affinity, struct module *owner) { + irq_moderation_init_fields(desc); desc->irq_common_data.handler_data =3D NULL; desc->irq_common_data.msi_desc =3D NULL; =20 --=20 2.55.0.737.g08866a6d13-goog From nobody Mon Sep 28 17:48:02 2026 Received: from mail-ej1-f69.google.com (mail-ej1-f69.google.com [209.85.218.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EA6592F745D for ; Wed, 19 Aug 2026 12:44:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143445; cv=none; b=LrbsJaqParB2V0kI/rPjmld04NY80fyjPeJ+nhdBaOtGBBRMYTpZxhkFENyqAx9Pc2ENoHvAyoKxlmw42RaCYJURN2TxP9501wDJh7Z3t1fLCeQ7mc0Afeh0W+6d3d02yrDNvRWbYOqyWCSeSj9/7kGbmlSFqUB0OxlB2Kwqm18= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143445; c=relaxed/simple; bh=IjTJqwhFGVDI4FVEfIUW6Az24jC7ENWkPSA/4XTgYZg=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=HfL8IPqirhAzwWJZJPs+LhAbbuJJEOaBueU1iEDMLCSzTTqQG9N1pGZGOVWZu/7QU+gYA7J9F6iy4chX7/Ip41HFRXvopZMZVx2mitGk+sBtPAPKCXpLDU8hX4BtHupUvPRejDOm75M89F7dwBYVtUSREJ+isnTGmMQ++6dSnbU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=OEspwffA; arc=none smtp.client-ip=209.85.218.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="OEspwffA" Received: by mail-ej1-f69.google.com with SMTP id a640c23a62f3a-c20262b5e10so104006166b.1 for ; Wed, 19 Aug 2026 05:44:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787143439; x=1787748239; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=aFxkMDZOQBabx4zg7nJClUru9l/HhQ40JvBa/Y+K1tQ=; b=OEspwffAypz5bcTCUi7q5twV0W+6CgUBaRbP8RzarOVFPw6IHDmMH6y0GOdU/fuC6D 2GXRcgO7Qf4SzIVXQSKCt9zZmh2hM6T33GOWZE4GOxAajcfqTLybRaSirj2spk5+afo3 SFzH2XAXVXn3bwsQ/QNKjDF8Dfw1I7KWR9bEZn1pFx+Iu9Sk2wbZe9BTXeQlvkg4Kzmg i2f8PugiI/k9SNWABgxbwGhlDDAafpLGKQQAemDRxkOE6sADX35l69YAaZMjeYHEHVdc qxDv/4180BMllDpATEuX4fYCC5pZNTggFhGTbFfPNyKSliX8e4bVQV9t+ThtAh30V5P+ 6j0A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787143439; x=1787748239; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=aFxkMDZOQBabx4zg7nJClUru9l/HhQ40JvBa/Y+K1tQ=; b=Fe7aLNGRwGNke1JLb6a85djNqjUP9YPoxqCYc/cFUlZLh1OlFYGFs+71rzSTlGEYcq JoS/o5wXBC2mlXWW3tbauJHQ3ulzEcQqpvGPb9ocewSALgXGV7n7tYRAET+7l3s+28RZ Fr/4EDxtLgcTGGc1qCcq0Ctf1Iwa5BzvACMk9QMV0pQZFxtayAiSScL1hucD9+5RZmli 0Y3QvfyWuP3RDbqLqeLvup5fqQOfSCCiUlXilxATRH6FB1qIZILMJv6rT1A/SHJXtT3D l7/hgvkhnbw1vVnO1bYqACJRaGQDivpu51Kfn/JZVM/W0ObcDbL+mCHspkGWmnx6EUiw 3wmw== X-Gm-Message-State: AOJu0Ywb35O916+r06KKl9vKtE59WBZFlz7uXlgSSFF4OZeverqPQWi6 NI/HQVTYzXk92sLkg5OXaa01My+fCDQiy2N81CDxsNXAEd7xt/KDhTcth7D0/IhCe+YfuIpICOJ SlbiHQQ== X-Received: from ejdb4.prod.google.com ([2002:a17:906:1504:b0:c15:d08f:480a]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a17:907:7b96:b0:c19:45df:1856 with SMTP id a640c23a62f3a-c242e249bc0mr225824566b.17.1787143438894; Wed, 19 Aug 2026 05:43:58 -0700 (PDT) Date: Wed, 19 Aug 2026 12:43:37 +0000 In-Reply-To: <20260819124341.4185621-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260819124341.4185621-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.737.g08866a6d13-goog Message-ID: <20260819124341.4185621-4-lrizzo@google.com> Subject: [PATCH v5 3/7] genirq: Implement core GSIM moderation logic From: Luigi Rizzo To: Thomas Gleixner , Marc Zyngier , Luigi Rizzo , Paolo Abeni Cc: linux-kernel@vger.kernel.org, linux-pci@vger.kernel.org, Bjorn Helgaas , Luigi Rizzo Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Introduce the core GSIM moderation logic, including full documentation of the architecture, per-CPU state tracking, basic list-draining functions and hooks for the interrupt handlers. Note: The architecture documentation describes how GSIM integrates with the interrupt flow handlers (e.g. handle_edge_irq, handle_fasteoi_irq). These handlers are modified and GSIM is integrated in subsequent commits. Signed-off-by: Luigi Rizzo --- kernel/irq/Makefile | 1 + kernel/irq/irq_moderation.c | 296 ++++++++++++++++++++++++++++++++++++ kernel/irq/irq_moderation.h | 93 +++++++++++ 3 files changed, 390 insertions(+) create mode 100644 kernel/irq/irq_moderation.c create mode 100644 kernel/irq/irq_moderation.h diff --git a/kernel/irq/Makefile b/kernel/irq/Makefile index 86a2e5ae08f9a..07c19d4697712 100644 --- a/kernel/irq/Makefile +++ b/kernel/irq/Makefile @@ -5,6 +5,7 @@ obj-$(CONFIG_GENERIC_IRQ_CHIP) +=3D generic-chip.o obj-$(CONFIG_GENERIC_IRQ_PROBE) +=3D autoprobe.o obj-$(CONFIG_IRQ_DOMAIN) +=3D irqdomain.o obj-$(CONFIG_IRQ_SIM) +=3D irq_sim.o +obj-$(CONFIG_IRQ_SW_MODERATION) +=3D irq_moderation.o obj-$(CONFIG_PROC_FS) +=3D proc.o obj-$(CONFIG_GENERIC_PENDING_IRQ) +=3D migration.o obj-$(CONFIG_GENERIC_IRQ_MIGRATION) +=3D cpuhotplug.o diff --git a/kernel/irq/irq_moderation.c b/kernel/irq/irq_moderation.c new file mode 100644 index 0000000000000..8f8893d952de0 --- /dev/null +++ b/kernel/irq/irq_moderation.c @@ -0,0 +1,296 @@ +// SPDX-License-Identifier: GPL-2.0 OR BSD-2-Clause + +/* + * Copyright (C) 2025-2026 Google LLC + * + * Global Software Interrupt Moderation (GSIM) core logic. + */ + +#include +#include +#include +#include +#include +#include + +#include "internals.h" +#include "irq_moderation.h" + +/* + * Global Software Interrupt Moderation (GSIM) + * + * Some platforms show reduced I/O performance when the total device inter= rupt + * rate across the entire platform becomes too high. To address the proble= m, + * GSIM runs after the handler to implement software interrupt moderation + * with programmable delay. + * + * =3D=3D=3D ARCHITECTURE =3D=3D=3D + * + * INTERRUPT HANDLING (for interrupt types that support moderation) + * - irq_start_moderation() runs under desc->lock right after the interrup= t handler. + * If the interrupt must be moderated, sets IRQD_IRQ_INPROGRESS and IRQD= _MODERATED, + * calls __disable_irq(), adds the irq_desc to a per-CPU list of moderat= ed interrupts, + * and starts a moderation timer if not yet active; + * - handle_xx_irq() is modified so that when called on a moderated irq_de= sc it + * calls mask_irq(), sets IRQS_PENDING and returns immediately; + * - the timer callback drains the moderation list: on each irq_desc it ac= quires + * desc->lock, and if desc->action !=3D NULL calls __enable_irq(), possi= bly calling + * the handler if IRQS_PENDING is set. + * + * INTERRUPT TEARDOWN + * It is protected by IRQD_IRQ_INPROGRESS and checking desc->action !=3D N= ULL. + * This works because free_irq() runs in two steps: + * - first clear desc->action (under lock), + * - then call synchronize_irq(), which blocks on IRQD_IRQ_INPROGRESS + * before freeing resources. + * When the moderation timer races with free_irq() we can have two cases: + * 1. timer runs before clearing desc->action. In this case __enable_irq() + * is valid and the subsequent free_irq() will complete as intended + * 2. desc->action is cleared before the timer runs. In this case synchron= ize_irq() + * will block until the timer expires (remember moderation delays are v= ery short, + * comparable to C-state exit times), __enable_irq() will not be run, + * and free_irq() will complete successfully. + * + * INTERRUPT MIGRATION + * It is protected by IRQD_IRQ_INPROGRESS that prevents running the handle= r on the + * new CPU while an interrupt is moderated. + * + * HOTPLUG + * During CPU shutdown, the kernel moves timers and reassigns interrupt af= finity + * to a new CPU. The easiest way and most robust way to guarantee that pen= ding + * events are handled correctly is to use a per-CPU "moderation_allowed" f= lag + * and hotplug callbacks on CPUHP_AP_ONLINE_DYN (some others are equally g= ood): + * - on setup, set the flag. That will allow interrupts to be moderated. + * - on shutdown, with interrupts disabled, 1. clear the flag thus prevent= ing + * more interrupts to be moderated on that CPU, 2. flush the list of mod= erated + * interrupts (as if the timer had fired), and 3. cancel the timer. + * This avoids depending with the internals of the up/down sequence. + * + * STATIC ENABLING + * GSIM is disabled by default (delay_ns =3D 0). To statically enable it: + * 1. Initialize irq_mod_params.delay_ns to a non-zero value (e.g., 100000= for 100us) + * in kernel/irq/irq_moderation.c. + * 2. Call irq_settings_set_moderatable(desc) during interrupt allocation + * (e.g., in __setup_irq() in kernel/irq/manage.c) for the desired inte= rrupts. + * + * SUSPEND & HIBERNATION + * During Suspend-to-RAM or Suspend-to-Disk (Hibernation), secondary CPUs = are + * taken offline, which triggers the CPU hotplug teardown and setup callba= cks. + * However, the boot processor is never taken offline via hotplug. + * + * To ensure the boot processor's GSIM state is safely drained/disabled be= fore + * suspend, and safely re-enabled after resume/restore, we register a PM n= otifier. + * + * The PM notifier handles all suspend, hibernation, and restore transitio= ns: + * - On prepare (*_PREPARE): runs mod_pm_prepare_cb() on all CPUs (via IPI) + * to clear the allowed flag and drain pending moderated interrupts, and + * then cancels GSIM timers on all online CPUs outside IPI context. + * - On resume/restore (POST_*): runs mod_pm_resume_cb() on all CPUs to + * safely re-allow moderation. We do NOT run cpu_setup_cb() here to avoid + * dangerous double-initialization of active hrtimers or list heads on + * already-online secondary CPUs. + */ + +/* + * GSIM parameters. Initialize delay_ns here to statically enable moderati= on + * (e.g. .delay_ns =3D 100000). + */ +struct irq_mod_params irq_mod_params ____cacheline_aligned; + +DEFINE_PER_CPU_ALIGNED(struct irq_mod_state, irq_mod_state); + +DEFINE_STATIC_KEY_FALSE(irq_moderation_enabled_key); + +static void update_enable_key(void) +{ + if (irq_mod_params.delay_ns !=3D 0) + static_branch_enable(&irq_moderation_enabled_key); + else + static_branch_disable(&irq_moderation_enabled_key); +} + +/* Actually start moderation. */ +bool irq_moderation_do_start(struct irq_desc *desc, struct irq_mod_state *= m) +{ + lockdep_assert_held(&desc->lock); + + if (!hrtimer_is_queued(&m->timer)) { + const unsigned int min_delay_ns =3D 10000; + const u64 slack_ns =3D 2000; + + /* Accumulate sleep time, no moderation if too small. */ + m->sleep_ns +=3D READ_ONCE(irq_mod_params.delay_ns); + if (m->sleep_ns < min_delay_ns) + return false; + /* We need moderation, start the timer. */ + m->timer_set++; + hrtimer_start_range_ns(&m->timer, ns_to_ktime(m->sleep_ns), + slack_ns, HRTIMER_MODE_REL_PINNED_HARD); + } + + /* + * Add to the timer list, set appropriate flags, and call + * __disable_irq() to prevent serving subsequent interrupts. + */ + m->enqueue++; + list_add(&desc->swmod_state.swmod_node, &m->descs); + irqd_set(&desc->irq_data, IRQD_IRQ_INPROGRESS | IRQD_MODERATED); + __disable_irq(desc); + return true; +} + +static void clean_moderation_state(struct irq_desc *desc) +{ + /* + * Clearing IRQD_IRQ_INPROGRESS allows synchronize_irq() to complete, + * signaling that GSIM teardown for this descriptor is finished. + */ + irqd_clear(&desc->irq_data, IRQD_IRQ_INPROGRESS | IRQD_MODERATED); + /* Only enable if action is set, protect against concurrent free_irq(). */ + if (desc->action) + __enable_irq(desc); +} + +/* Used on timer expiration or CPU shutdown. */ +static void drain_desc_list(struct irq_mod_state *m) +{ + struct irq_desc *desc, *next; + + /* Remove from list and enable interrupts back. */ + list_for_each_entry_safe(desc, next, &m->descs, swmod_state.swmod_node) { + guard(raw_spinlock)(&desc->lock); + list_del_init(&desc->swmod_state.swmod_node); + clean_moderation_state(desc); + } +} + +static enum hrtimer_restart timer_callback(struct hrtimer *timer) +{ + struct irq_mod_state *m =3D this_cpu_ptr(&irq_mod_state); + + lockdep_assert_irqs_disabled(); + + drain_desc_list(m); + /* Prepare to accumulate next moderation delay. */ + m->sleep_ns =3D 0; + return HRTIMER_NORESTART; +} + +/* Hotplug callback for setup. */ +static int cpu_setup_cb(unsigned int cpu) +{ + struct irq_mod_state *m =3D this_cpu_ptr(&irq_mod_state); + + hrtimer_setup(&m->timer, timer_callback, CLOCK_MONOTONIC, HRTIMER_MODE_RE= L_PINNED_HARD); + INIT_LIST_HEAD(&m->descs); + /* Ensure initialization is visible before setting the flag. */ + smp_store_release(&m->initialized, true); + m->moderation_allowed =3D true; + return 0; +} + +/* + * Hotplug callback for shutdown. + * Mark the CPU as offline for moderation, and drain the list of masked + * interrupts. Any subsequent interrupt on this CPU will not be + * moderated, but they will be on the new target. + */ +static int cpu_remove_cb(unsigned int cpu) +{ + struct irq_mod_state *m =3D this_cpu_ptr(&irq_mod_state); + + /* Protect the desc list, interrupts could modify it. */ + scoped_guard(irqsave) { + m->moderation_allowed =3D false; + drain_desc_list(m); + } + /* Run hrtimer_cancel() outside hardirq/IPI context. */ + hrtimer_cancel(&m->timer); + /* Ensure state is visible before clearing the flag. */ + smp_store_release(&m->initialized, false); + return 0; +} + +static void mod_pm_prepare_cb(void *arg) +{ + struct irq_mod_state *m =3D this_cpu_ptr(&irq_mod_state); + + /* Called via IPI so local interrupts are disabled. */ + if (mod_state_initialized(m)) { + m->moderation_allowed =3D false; + drain_desc_list(m); + } +} + +static void mod_pm_resume_cb(void *arg) +{ + struct irq_mod_state *m =3D this_cpu_ptr(&irq_mod_state); + + /* + * The hrtimer and list head are already initialized (either at boot + * or during hotplug CPU online). We must not re-initialize them here + * as they might already be active if devices resumed and fired + * interrupts before this notifier ran. + */ + if (mod_state_initialized(m)) + m->moderation_allowed =3D true; +} + +static int mod_pm_notifier_cb(struct notifier_block *nb, unsigned long eve= nt, void *unused) +{ + int cpu; + + switch (event) { + case PM_SUSPEND_PREPARE: + case PM_HIBERNATION_PREPARE: + case PM_RESTORE_PREPARE: + on_each_cpu(mod_pm_prepare_cb, NULL, 1); + /* Run hrtimer_cancel() outside hardirq/IPI context. */ + for_each_online_cpu(cpu) { + struct irq_mod_state *m =3D per_cpu_ptr(&irq_mod_state, cpu); + + if (mod_state_initialized(m)) + hrtimer_cancel(&m->timer); + } + break; + case PM_POST_SUSPEND: + case PM_POST_HIBERNATION: + case PM_POST_RESTORE: + on_each_cpu(mod_pm_resume_cb, NULL, 1); + break; + } + return NOTIFY_OK; +} + +struct notifier_block mod_nb =3D { + .notifier_call =3D mod_pm_notifier_cb, + .priority =3D 100, +}; + +static int __init init_irq_moderation(void) +{ + int cpuhp_state; + int ret; + + cpuhp_state =3D cpuhp_setup_state(CPUHP_AP_ONLINE_DYN, "sw_moderation", + cpu_setup_cb, cpu_remove_cb); + if (cpuhp_state < 0) { + pr_err("%s: Failed to setup hotplug notifier\n", __func__); + return cpuhp_state; + } + + ret =3D register_pm_notifier(&mod_nb); + if (ret < 0) { + pr_err("%s: Failed to register pm notifier\n", __func__); + goto cleanup; + } + + /* Enable if the defaults require it. */ + update_enable_key(); + return 0; + +cleanup: + cpuhp_remove_state(cpuhp_state); + return ret; +} +device_initcall(init_irq_moderation); diff --git a/kernel/irq/irq_moderation.h b/kernel/irq/irq_moderation.h new file mode 100644 index 0000000000000..8df350651cd7c --- /dev/null +++ b/kernel/irq/irq_moderation.h @@ -0,0 +1,93 @@ +/* SPDX-License-Identifier: GPL-2.0 OR BSD-2-Clause */ + +/* + * Copyright (C) 2025-2026 Google LLC + * + * Common data structures for Global Software Interrupt Moderation, GSIM + */ + +#ifndef _LINUX_IRQ_MODERATION_H +#define _LINUX_IRQ_MODERATION_H + +#ifdef CONFIG_IRQ_SW_MODERATION + +#include +#include +#include +#include + +/** + * struct irq_mod_params - configuration parameters + * @delay_ns: maximum delay + */ +struct irq_mod_params { + unsigned int delay_ns; +}; + +extern struct irq_mod_params irq_mod_params; + +/** + * struct irq_mod_state - per-CPU moderation state + * + * Used on every interrupt: + * @timer: moderation timer + * @initialized: true if hrtimer and list head are initialized + * @moderation_allowed: per-CPU flag, toggled during hotplug/suspend events + * @sleep_ns: accumulated time for actual delay + * + * Used once per moderation delay per interrupt source: + * @descs: list of moderated irq_desc on this CPU + * @enqueue: how many enqueue on the list + * + * Statistics + * @timer_set: how many timer_set calls + */ +struct irq_mod_state { + struct hrtimer timer; + bool initialized; + bool moderation_allowed; + unsigned int sleep_ns; + struct list_head descs; + unsigned int enqueue; + unsigned int timer_set; +}; + +DECLARE_PER_CPU_ALIGNED(struct irq_mod_state, irq_mod_state); + +static inline bool mod_state_initialized(struct irq_mod_state *m) +{ + /* + * There is no public API in the hrtimer or list subsystems to check + * if they are initialized. We use this flag to avoid dereferencing + * uninitialized pointers during CPU hotplug/suspend races. + */ + return smp_load_acquire(&m->initialized); +} + +extern struct static_key_false irq_moderation_enabled_key; + +bool irq_moderation_do_start(struct irq_desc *desc, struct irq_mod_state *= m); + +/* + * Call after running the handler, with lock held. If this source should be + * moderated, disable it, add to the timer list for this CPU and return tr= ue, + * and exit from handle_*_irq() without processing IRQS_PENDING, because + * that will happen when the moderation timer fires and calls __enable_irq= (). + */ +static inline bool irq_start_moderation(struct irq_desc *desc) +{ + struct irq_mod_state *m =3D this_cpu_ptr(&irq_mod_state); + + if (static_branch_unlikely(&irq_moderation_enabled_key) && + irq_settings_moderatable(desc) && + m->moderation_allowed) { + return irq_moderation_do_start(desc, m); + } + return false; +} + +#else /* CONFIG_IRQ_SW_MODERATION */ +static inline bool irq_start_moderation(struct irq_desc *desc) { return fa= lse; } +#endif /* CONFIG_IRQ_SW_MODERATION */ + +#endif /* _LINUX_IRQ_MODERATION_H */ --=20 2.55.0.737.g08866a6d13-goog From nobody Mon Sep 28 17:48:02 2026 Received: from mail-ej1-f71.google.com (mail-ej1-f71.google.com [209.85.218.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E183436A36E for ; Wed, 19 Aug 2026 12:44:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.71 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143445; cv=none; b=NZNt1ari/7KGW7jk2LhLaRUxSGyx/ob8jhuGMXS3ZmJ1V95pfw1q1W6Lhe5rJwfxK44F/iOYzyZxYAII3P2S4mrQnr4wljcyhL2NQAM5Mi8DYJeyBRW3Y0Zv1gLLFyzIwrzQOG6K9UXy7Wdmzoxa0srifcvDLifCFQyBTRPwkjA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143445; c=relaxed/simple; bh=Ox+d0B8XD1GDbofxCcqR2ba+haJM/n/mDwRsHvJc33c=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=Pg1ng9njAYQ0Cl70mdn0TgpZ/QOq1Q2FqcFwKp4aCrL4J0YqEyNU5y+1NdZ+ONsgtns3zEukMB5L5NscTxcQBVOiXYb0k1GqT9OcxU+sRhj35mUlLnjMTDBkFbxI+1ZKt8XNlmy58PMwoOJqy9AwEk6f46Lm9LWulOoQYkd7faU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=XWKhXt95; arc=none smtp.client-ip=209.85.218.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="XWKhXt95" Received: by mail-ej1-f71.google.com with SMTP id a640c23a62f3a-c1686e23b9cso73561466b.0 for ; Wed, 19 Aug 2026 05:44:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787143440; x=1787748240; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=lgslhZs5lHfUM6vsYvygJcRXohyJMB1GN9y8NyUjAaQ=; b=XWKhXt95SkQU+XYBRlYMs8s+PNWBwzk7/zYXLP3MouUU5h7mxSyUsczyOEIIDvhDIX YhP/U/wJFUjuWXzQrVGEttaf1XzV0JZQErpneHICyvgzmbv2lIFtrCpEtA2WlOnp+upr P8oNt2i+tcD7x+eKwyqQh1MJJfDma8N11lhkbbL7BaxJUATicHaUJrQ0nEuPTuCHqCrH 66yv2eUtuiRjSavb/d5+dWlzPmJBDf3l8TDipEO8+w6mL78hSySTyTl4VrwA2ClMoY9F 9xcBL5NzJ0FXBQtbhK3u3UsEpJcksJtdPcmjk6RLhoV17UmYmCUxwQXUg497MYWqjdC6 9DOg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787143440; x=1787748240; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=lgslhZs5lHfUM6vsYvygJcRXohyJMB1GN9y8NyUjAaQ=; b=DdFSpYHz9F1aBRFQ2Sx4cEx+oVB9vy/v3hYLrpkjySL9RXlnS4gzezU2zw7z7gjIMG PXOT9V2ecxxSd0+RI3zf7PDzXOIRZ9N0MGFdlQwHoFW6Bel+65FPaOfdgjpU+kVKgppo vb4aGp3xl17hX+khXx2dNcdujrFqaWWOIB9EgIblFyyr9wukzFWHaCK+3Rtmo+BBXM3S UuC1AHHHXmtcsD7WVcee5Z6wiJWQYEcH3qzeYMXUDidT+kc8z0/9jQKgnKjcNEz5hsP1 /RigkvEOG9GDbI5hr49yUCm+YtWfXTsClkJIjgb6F0VPiMN5K2FC9p2SVOtKFqsVa//W lreA== X-Gm-Message-State: AOJu0YzlBHOcWEEGa0GSlWQVqhX3nFjXshT0pwxp5xo//v/MSwJyb8rs xlVVn7aEuFkkDrqhWlhRRTK0dkiFb/84FRSCJ9+3/LjiIXVG3sYZJu9wXsciyfpaoP/EMSkDB0Y k5wvL3A== X-Received: from ejgi22.prod.google.com ([2002:a17:906:3c56:b0:c16:6d0e:5e64]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a17:907:6d08:b0:c20:5213:56f1 with SMTP id a640c23a62f3a-c240544ed8dmr308638566b.13.1787143439988; Wed, 19 Aug 2026 05:43:59 -0700 (PDT) Date: Wed, 19 Aug 2026 12:43:38 +0000 In-Reply-To: <20260819124341.4185621-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260819124341.4185621-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.737.g08866a6d13-goog Message-ID: <20260819124341.4185621-5-lrizzo@google.com> Subject: [PATCH v5 4/7] genirq: Integrate GSIM into interrupt flow From: Luigi Rizzo To: Thomas Gleixner , Marc Zyngier , Luigi Rizzo , Paolo Abeni Cc: linux-kernel@vger.kernel.org, linux-pci@vger.kernel.org, Bjorn Helgaas , Luigi Rizzo Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Hook GSIM into edge/fasteoi handlers and teardown. Specifically: - Hook GSIM into handle_edge_irq() and handle_fasteoi_irq(). - Introduce `irq_moderation_allow(desc, allow)` to validate if an IRQ is eligible for GSIM (non-oneshot, edge/fasteoi, single target, no bus lo= ck) and set/clear the IRQ_MODERATABLE flag. - Integrate GSIM setup in __setup_irq() using `irq_moderation_allow()`, defaulting to disabled (`use_moderation =3D false`). - Clear GSIM state in __free_irq() and __cleanup_nmi(). Note: At this stage, GSIM is disabled by default (use_moderation =3D false)= and cannot be enabled at runtime yet. Subsequent commits will introduce procfs configuration and a configurable default mode. For testing this commit, use_moderation can be manually set to true in the code. Signed-off-by: Luigi Rizzo --- kernel/irq/chip.c | 14 +++++++++++++ kernel/irq/internals.h | 8 +++++++ kernel/irq/irq_moderation.c | 42 +++++++++++++++++++++++++++++++++++++ kernel/irq/manage.c | 10 +++++++++ 4 files changed, 74 insertions(+) diff --git a/kernel/irq/chip.c b/kernel/irq/chip.c index de754db414d1d..04a30ba04e78b 100644 --- a/kernel/irq/chip.c +++ b/kernel/irq/chip.c @@ -20,6 +20,7 @@ #include =20 #include "internals.h" +#include "irq_moderation.h" =20 static irqreturn_t bad_chained_irq(int irq, void *dev_id) { @@ -500,6 +501,10 @@ static bool irq_can_handle_pm(struct irq_desc *desc) return false; } =20 + /* Moderated interrupts have IRQD_IRQ_INPROGRESS and need early return. */ + if (irqd_is_moderated(irqd)) + return false; + /* Check whether the interrupt is polled on another CPU */ if (unlikely(desc->istate & IRQS_POLL_INPROGRESS)) { if (WARN_ONCE(irq_poll_cpu =3D=3D smp_processor_id(), @@ -749,6 +754,10 @@ void handle_fasteoi_irq(struct irq_desc *desc) * handling the previous one - it may need to be resent. */ if (!irq_can_handle_pm(desc)) { + if (irqd_is_moderated(&desc->irq_data)) { + desc->istate |=3D IRQS_PENDING; + mask_irq(desc); + } if (irqd_needs_resend_when_in_progress(&desc->irq_data)) desc->istate |=3D IRQS_PENDING; cond_eoi_irq(chip, &desc->irq_data); @@ -769,6 +778,9 @@ void handle_fasteoi_irq(struct irq_desc *desc) =20 cond_unmask_eoi_irq(desc, chip); =20 + if (irq_start_moderation(desc)) + return; + /* * When the race described above happens this will resend the interrupt. */ @@ -858,6 +870,8 @@ void handle_edge_irq(struct irq_desc *desc) =20 handle_irq_event(desc); =20 + if (irq_start_moderation(desc)) + break; } while ((desc->istate & IRQS_PENDING) && !irqd_irq_disabled(&desc->irq_d= ata)); } EXPORT_SYMBOL(handle_edge_irq); diff --git a/kernel/irq/internals.h b/kernel/irq/internals.h index 716a7c2e3633a..705c3bfe28e12 100644 --- a/kernel/irq/internals.h +++ b/kernel/irq/internals.h @@ -400,6 +400,14 @@ static inline void irq_moderation_init_fields(struct i= rq_desc *desc) { INIT_LIST_HEAD(&desc->swmod_state.swmod_node); } + +int irq_moderation_allow(struct irq_desc *desc, bool allow); +bool irq_moderation_supported(struct irq_desc *desc); #else static inline void irq_moderation_init_fields(struct irq_desc *desc) {} +static inline int irq_moderation_allow(struct irq_desc *desc, bool allow) +{ + return allow ? -EOPNOTSUPP : 0; +} +static inline bool irq_moderation_supported(struct irq_desc *desc) { retur= n false; } #endif diff --git a/kernel/irq/irq_moderation.c b/kernel/irq/irq_moderation.c index 8f8893d952de0..2c75feb6634f3 100644 --- a/kernel/irq/irq_moderation.c +++ b/kernel/irq/irq_moderation.c @@ -294,3 +294,45 @@ static int __init init_irq_moderation(void) return ret; } device_initcall(init_irq_moderation); + +bool irq_moderation_supported(struct irq_desc *desc) +{ + struct irq_data *irqd =3D &desc->irq_data; + struct irq_chip *chip =3D irqd->chip; + + /* GSIM does not support shared interrupts */ + if (desc->action && desc->action->next) + return false; + + if (desc->istate & IRQS_ONESHOT) + return false; + if (irqd_is_level_type(irqd)) + return false; + if (!irqd_is_single_target(irqd)) + return false; + if (chip->irq_bus_lock || chip->irq_bus_sync_unlock) + return false; + if (!chip->irq_mask || !chip->irq_unmask) + return false; + if (desc->handle_irq !=3D handle_edge_irq && desc->handle_irq !=3D handle= _fasteoi_irq) + return false; + return true; +} + +int irq_moderation_allow(struct irq_desc *desc, bool allow) +{ + lockdep_assert_held(&desc->lock); + + if (!allow) { + irq_settings_clr_moderatable(desc); + return 0; + } + + if (!irq_moderation_supported(desc)) { + irq_settings_clr_moderatable(desc); + return -EOPNOTSUPP; + } + + irq_settings_set_moderatable(desc); + return 0; +} diff --git a/kernel/irq/manage.c b/kernel/irq/manage.c index 7eb07e3bdb4c2..ce92885e71323 100644 --- a/kernel/irq/manage.c +++ b/kernel/irq/manage.c @@ -1470,6 +1470,12 @@ static bool valid_percpu_irqaction(struct irqaction = *old, struct irqaction *new) static int __setup_irq(unsigned int irq, struct irq_desc *desc, struct irqaction *new) { + /* + * Choose the default moderation mode. + * Hardcoded to false for now; configurable defaults are added later. + * Set to true here to force-enable GSIM for testing this commit. + */ + bool use_moderation =3D false; struct irqaction *old, **old_ptr; unsigned long flags, thread_mask =3D 0; int ret, nested, shared =3D 0; @@ -1789,6 +1795,8 @@ __setup_irq(unsigned int irq, struct irq_desc *desc, = struct irqaction *new) =20 irq_pm_install_action(desc, new); =20 + irq_moderation_allow(desc, use_moderation); + /* Reset broken irq detection when installing new handler */ desc->irq_count =3D 0; desc->irqs_unhandled =3D 0; @@ -1897,6 +1905,7 @@ static struct irqaction *__free_irq(struct irq_desc *= desc, void *dev_id) /* If this was the last handler, shut down the IRQ line: */ if (!desc->action) { irq_settings_clr_disable_unlazy(desc); + irq_settings_clr_moderatable(desc); /* Only shutdown. Deactivate after synchronize_hardirq() */ irq_shutdown(desc); } @@ -2046,6 +2055,7 @@ static const void *__cleanup_nmi(unsigned int irq, st= ruct irq_desc *desc) desc->action =3D NULL; =20 irq_settings_clr_disable_unlazy(desc); + irq_settings_clr_moderatable(desc); irq_shutdown_and_deactivate(desc); } =20 --=20 2.55.0.737.g08866a6d13-goog From nobody Mon Sep 28 17:48:02 2026 Received: from mail-lf1-f70.google.com (mail-lf1-f70.google.com [209.85.167.70]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D2102306B3D for ; Wed, 19 Aug 2026 12:44:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.167.70 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143448; cv=none; b=p5EBy87gvrNn3uA8E4NFaqlgpUILnSc4Wkgp7/T7rheDDGjhhGcUGnSCqJkTyfXO9lnnS3rI/Rl66CzKACveoJX9skEP3gP05f6sO8Hluh6AIF4cpBnSJbWPc9oKRfK2/e/XyFHDqS0RifLbSGLSVldD9b2mbo6G2zAnoduj1jw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143448; c=relaxed/simple; bh=A1KBG7Ax/8AKi8Isyu+Y2QFb+3dSt59+RI7QmdoANfk=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=gN8Ndc51xHflla9BTkKvNkYuRytDFrPZu+alB/tQnM24FRDEl6+Nu2xhNRy2B7vdvMB+FkKhMLGaQjqknN3uu1fx1R2iqTLeHT8edhBSgVvRjtCzZ6YUxZGvqyKVjx25AFB5hk5ebUU7koaKcgnMiO+8x5puv/TrfnXA3zIffMs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=XdAtu4Tc; arc=none smtp.client-ip=209.85.167.70 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="XdAtu4Tc" Received: by mail-lf1-f70.google.com with SMTP id 2adb3069b0e04-5aeb54a4c5dso699771e87.2 for ; Wed, 19 Aug 2026 05:44:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787143442; x=1787748242; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=gxSlDtVCWpSOQOgYo4sL5Y8CxzWFyIGgo+drWq4mnP4=; b=XdAtu4TcgjJPO/wtC3iqVQTl0AGoTwhbPNvHXkoFBl6KNhNi2x7G5CG8/k7hG7dLvH T4cCRM27XLlxeyLo7lpq2lVZrPjT8nl03JOHNY9uTTvua07R/w9qWK49yK+ZecGMgCUa luEB+sf/clcSe1CznUBpXjE0zdoj4pArwxu9INlv3OIQDYm5jAyZCrR63uJiyLu2TlOa I77NSbtgxzAENfXBvQtuyiDhNzlgqMVfusURdyTrcIU0o4LEJj0M/ctMKZyN/LQoaCK1 Ob0Zozf7vLlMWadcUN2XsydBKQtK2oEDg2mMiuaOZOqya4u+sXHYBxx3oOinbl2xYjE2 jcvg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787143442; x=1787748242; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=gxSlDtVCWpSOQOgYo4sL5Y8CxzWFyIGgo+drWq4mnP4=; b=TKkhpiDkzDmcBsWtHC8waxqtDy+rMv8dB6c3pNkJkFn/8Ro8y6Lj9XdfOYITt4rU2s h2MDFO/+zzRvd1GpDcGm2TIrUXQhmNUt7ccBQ4Ke5Oc8iWM0yuo+TdlzCyC614we3Rl4 6FJF+Puwb4YHc65ZK2ZGoqR2pHaPxab5GNZxoHtLMrg4lngyTn9voIv6NmFjSPoxaEBN fkjH1YuyCWpDrRFoCXkgyX0aeP36NbHLf3tvgVwNKFdfdUoujK/l+BvAn1ba9ZVUTEMk mVYUrpweLA+D4fUSKrT2x2QtH7nC+Df5ABMAB89d4sHVu2hAS3Vk1QuR8OdKrCDOw+Ia 73Nw== X-Gm-Message-State: AOJu0Yw2xmlxqqX7TDRqDgdh0L7zXGJu8ewsngbxFcEI+NsqPh8A48kX CZdM9tw1OMgfZU6L0RgFktxYmgO6qhzjj8NgZEYPTjdFvJUE1ThX7dJQIlHBf0nAf7WOkvvr+Qk GxlZSew== X-Received: from lfhl6.prod.google.com ([2002:ac2:4306:0:b0:5b1:4ed1:9fa]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6512:a82:b0:5b2:e95a:491d with SMTP id 2adb3069b0e04-5b478bbe9e7mr1484145e87.9.1787143441463; Wed, 19 Aug 2026 05:44:01 -0700 (PDT) Date: Wed, 19 Aug 2026 12:43:39 +0000 In-Reply-To: <20260819124341.4185621-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260819124341.4185621-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.737.g08866a6d13-goog Message-ID: <20260819124341.4185621-6-lrizzo@google.com> Subject: [PATCH v5 5/7] genirq: Add GSIM user space configuration (procfs) From: Luigi Rizzo To: Thomas Gleixner , Marc Zyngier , Luigi Rizzo , Paolo Abeni Cc: linux-kernel@vger.kernel.org, linux-pci@vger.kernel.org, Bjorn Helgaas , Luigi Rizzo Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Introduce procfs interfaces to configure and monitor GSIM at runtime. This adds: - A global directory /proc/irq/sw_moderation/ containing: - delay_us: Read/write interface to set/get the maximum moderation delay in microseconds (defaults to 0, GSIM disabled). - stats: Read-only interface to monitor GSIM statistics per-CPU. - files /proc/irq/NN/allow_sw_moderation (where NN is the IRQ number) to individually allow/disallow moderation. Created only for interrupts that support moderation (e.g., edge-triggered, single-target, etc.). Signed-off-by: Luigi Rizzo --- kernel/irq/internals.h | 4 + kernel/irq/irq_moderation.c | 256 +++++++++++++++++++++++++++++++++++- kernel/irq/proc.c | 2 + 3 files changed, 261 insertions(+), 1 deletion(-) diff --git a/kernel/irq/internals.h b/kernel/irq/internals.h index 705c3bfe28e12..22a1926910f34 100644 --- a/kernel/irq/internals.h +++ b/kernel/irq/internals.h @@ -403,6 +403,8 @@ static inline void irq_moderation_init_fields(struct ir= q_desc *desc) =20 int irq_moderation_allow(struct irq_desc *desc, bool allow); bool irq_moderation_supported(struct irq_desc *desc); +void irq_moderation_procfs_add(struct irq_desc *desc, umode_t umode); +void irq_moderation_procfs_remove(struct irq_desc *desc); #else static inline void irq_moderation_init_fields(struct irq_desc *desc) {} static inline int irq_moderation_allow(struct irq_desc *desc, bool allow) @@ -410,4 +412,6 @@ static inline int irq_moderation_allow(struct irq_desc = *desc, bool allow) return allow ? -EOPNOTSUPP : 0; } static inline bool irq_moderation_supported(struct irq_desc *desc) { retur= n false; } +static inline void irq_moderation_procfs_add(struct irq_desc *desc, umode_= t umode) {} +static inline void irq_moderation_procfs_remove(struct irq_desc *desc) {} #endif diff --git a/kernel/irq/irq_moderation.c b/kernel/irq/irq_moderation.c index 2c75feb6634f3..1474d33455410 100644 --- a/kernel/irq/irq_moderation.c +++ b/kernel/irq/irq_moderation.c @@ -11,6 +11,8 @@ #include #include #include +#include +#include #include =20 #include "internals.h" @@ -24,6 +26,21 @@ * GSIM runs after the handler to implement software interrupt moderation * with programmable delay. * + * Configuration is done at runtime via procfs + * echo ${VALUE} > /proc/irq/sw_moderation/${NAME} + * + * Supported parameters: + * + * delay_us (default 0, suggested 100, 0 off, range 0-500) + * Maximum moderation delay. A reasonable range is 20-100. Higher va= lues + * can be useful if the hardirq handler has long runtimes. + * + * Moderation is allowed/disallowed dynamically for individual interrupts = with + * echo 1 > /proc/irq/NN/allow_sw_moderation # use 0 to disallow + * + * Monitoring of per-cpu and global statistics is available via procfs + * cat /proc/irq/sw_moderation/stats + * * =3D=3D=3D ARCHITECTURE =3D=3D=3D * * INTERRUPT HANDLING (for interrupt types that support moderation) @@ -101,6 +118,8 @@ DEFINE_PER_CPU_ALIGNED(struct irq_mod_state, irq_mod_st= ate); =20 DEFINE_STATIC_KEY_FALSE(irq_moderation_enabled_key); =20 +static DEFINE_MUTEX(swmod_mutex); + static void update_enable_key(void) { if (irq_mod_params.delay_ns !=3D 0) @@ -139,6 +158,185 @@ bool irq_moderation_do_start(struct irq_desc *desc, s= truct irq_mod_state *m) return true; } =20 +/* + * struct var_info - target and limits for parameters + * @ptr: pointer to the value, NULL if not used. + * @min: minimum value allowed + * @max: maximum value allowed + * @scale: scale factor between procfs and internal. + */ +struct var_info { + unsigned int *ptr; + unsigned int min; + unsigned int max; + unsigned int scale; +}; + +/* + * struct swmod_procfs_entry - description for procfs entries and paramete= r limits + * @name: name in procfs. If NULL, the entry is only for limit checks. + * @wr: write handler for procfs. NULL if readonly + * @rd: read handler for procfs. + * @var: variable address and limits, if used. + */ +struct swmod_procfs_entry { + const char *name; + ssize_t (*wr)(struct var_info *n, const char __user *s, size_t count); + void (*rd)(struct seq_file *p); + struct var_info var; +}; + +static ssize_t swmod_wr(struct var_info *v, const char __user *s, size_t c= ount) +{ + unsigned int value; + int ret; + + ret =3D kstrtouint_from_user(s, count, 0, &value); + if (ret) + return ret; + if (value < v->min || value > v->max) + return -ERANGE; + WRITE_ONCE(*v->ptr, value * v->scale); + return count; +} + +static void swmod_rd(struct seq_file *p) +{ + struct swmod_procfs_entry *n =3D p->private; + + seq_printf(p, "%u\n", *n->var.ptr / n->var.scale); +} + +static ssize_t swmod_wr_delay(struct var_info *v, const char __user *s, si= ze_t count) +{ + ssize_t ret =3D swmod_wr(v, s, count); + + if (ret >=3D 0) + update_enable_key(); + return ret; +} + +#define HEAD_FMT "%5s %8s %11s %11s\n" +#define BODY_FMT "%5u %8u %11u %11u\n" + +/* Print statistics */ +static void rd_stats(struct seq_file *p) +{ + unsigned int delay_ns =3D READ_ONCE(irq_mod_params.delay_ns); + int cpu; + + if (delay_ns =3D=3D 0) + return; + seq_printf(p, HEAD_FMT, + "# CPU", "delay_ns", "timer_set", "enqueue"); + + for_each_possible_cpu(cpu) { + /* Copy statistics, will only use some unsigned int values; races ok. */ + struct irq_mod_state cur =3D data_race(*per_cpu_ptr(&irq_mod_state, cpu)= ); + + seq_printf(p, BODY_FMT, + cpu, + delay_ns, + cur.timer_set, + cur.enqueue); + } + + seq_printf(p, "\n" + "delay_us %lu\n", + delay_ns / NSEC_PER_USEC); +} + +static int param_show(struct seq_file *p, void *v) +{ + struct swmod_procfs_entry *n =3D p->private; + + n->rd(p); + return 0; +} + +static int param_open(struct inode *inode, struct file *file) +{ + return single_open(file, param_show, pde_data(inode)); +} + +static ssize_t param_write(struct file *f, const char __user *buf, size_t = count, loff_t *ppos) +{ + struct swmod_procfs_entry *n =3D (struct swmod_procfs_entry *)pde_data(fi= le_inode(f)); + ssize_t ret; + + if (!n->wr) + return -EINVAL; + mutex_lock(&swmod_mutex); + ret =3D n->wr(&n->var, buf, count); + mutex_unlock(&swmod_mutex); + return ret; +} + +static const struct proc_ops param_ops =3D { + .proc_open =3D param_open, + .proc_read =3D seq_read, + .proc_lseek =3D seq_lseek, + .proc_release =3D single_release, + .proc_write =3D param_write, +}; + +/* Handlers for /proc/irq/NN/allow_sw_moderation */ +static int allow_flag_show(struct seq_file *p, void *v) +{ + struct irq_desc *desc =3D irq_to_desc((long)p->private); + + if (!desc) + return -ENODEV; + + seq_puts(p, irq_settings_moderatable(desc) ? "on\n" : "off\n"); + return 0; +} + + +static ssize_t allow_flag_write(struct file *f, const char __user *buf, si= ze_t count, loff_t *ppos) +{ + struct irq_desc *desc =3D irq_to_desc((long)pde_data(file_inode(f))); + bool allow; + int ret; + + if (!desc) + return -ENODEV; + + ret =3D kstrtobool_from_user(buf, count, &allow); + + if (!ret) { + guard(raw_spinlock_irq)(&desc->lock); + ret =3D irq_moderation_allow(desc, allow); + } + return ret ? : count; +} + +static int allow_flag_open(struct inode *inode, struct file *file) +{ + return single_open(file, allow_flag_show, pde_data(inode)); +} + +static const struct proc_ops allow_flag_ops =3D { + .proc_open =3D allow_flag_open, + .proc_read =3D seq_read, + .proc_lseek =3D seq_lseek, + .proc_release =3D single_release, + .proc_write =3D allow_flag_write, +}; + +void irq_moderation_procfs_add(struct irq_desc *desc, umode_t umode) +{ + if (!irq_moderation_supported(desc)) + return; + proc_create_data("allow_sw_moderation", umode, desc->dir, + &allow_flag_ops, (void *)(long)desc->irq_data.irq); +} + +void irq_moderation_procfs_remove(struct irq_desc *desc) +{ + remove_proc_entry("allow_sw_moderation", desc->dir); +} + static void clean_moderation_state(struct irq_desc *desc) { /* @@ -267,10 +465,41 @@ struct notifier_block mod_nb =3D { .priority =3D 100, }; =20 +/* Helper to initialize the struct var_info. */ +#define SET_VAR(_ptr, _min, _max, _scale) \ + { .ptr =3D (_ptr), .min =3D (_min), .max =3D (_max), .scale =3D (_scale),= } + +static struct swmod_procfs_entry procfs_entries[] =3D { + { + .name =3D "delay_us", + .wr =3D swmod_wr_delay, + .rd =3D swmod_rd, + .var =3D SET_VAR(&irq_mod_params.delay_ns, 0, 500, NSEC_PER_USEC), + }, + { + .name =3D "stats", + .rd =3D rd_stats, + }, +}; + static int __init init_irq_moderation(void) { + struct proc_dir_entry *dir; int cpuhp_state; - int ret; + int i, ret; + + for (i =3D 0; i < ARRAY_SIZE(procfs_entries); i++) { + struct var_info *v =3D &procfs_entries[i].var; + + if (!v->ptr) + continue; + if (*v->ptr >=3D v->min * v->scale && *v->ptr <=3D v->max * v->scale) + continue; + pr_err("%s: Parameter %s: value %u out of bounds [%u,%u]\n", + __func__, procfs_entries[i].name ? : "no-name", + *v->ptr / v->scale, v->min, v->max); + return -ERANGE; + } =20 cpuhp_state =3D cpuhp_setup_state(CPUHP_AP_ONLINE_DYN, "sw_moderation", cpu_setup_cb, cpu_remove_cb); @@ -285,10 +514,35 @@ static int __init init_irq_moderation(void) goto cleanup; } =20 + /* Safe because /proc/irq is created earlier, in kernel_init_freeable(). = */ + dir =3D proc_mkdir("irq/sw_moderation", NULL); + if (!dir) { + pr_err("%s: Failed to create procfs directory\n", __func__); + goto cleanup_1; + } + for (i =3D 0; i < ARRAY_SIZE(procfs_entries); i++) { + struct swmod_procfs_entry *n =3D &procfs_entries[i]; + + if (!n->name || proc_create_data(n->name, n->wr ? 0644 : 0444, dir, &par= am_ops, n)) + continue; + pr_err("%s: Failed to create procfs entry %s\n", __func__, n->name); + for (i--; i >=3D 0; i--) { + n =3D &procfs_entries[i]; + if (n->name) + remove_proc_entry(n->name, dir); + } + remove_proc_entry("irq/sw_moderation", NULL); + goto cleanup_1; + } + /* Enable if the defaults require it. */ update_enable_key(); return 0; =20 +cleanup_1: + ret =3D -ENOMEM; + unregister_pm_notifier(&mod_nb); + cleanup: cpuhp_remove_state(cpuhp_state); return ret; diff --git a/kernel/irq/proc.c b/kernel/irq/proc.c index 1b835725f7b1c..edffb30efab58 100644 --- a/kernel/irq/proc.c +++ b/kernel/irq/proc.c @@ -379,6 +379,7 @@ void register_irq_proc(unsigned int irq, struct irq_des= c *desc) irq_effective_aff_list_proc_show, irqp); # endif #endif + irq_moderation_procfs_add(desc, 0644); proc_create_single_data("spurious", 0444, desc->dir, irq_spurious_proc_show, (void *)(long)irq); =20 @@ -400,6 +401,7 @@ void unregister_irq_proc(unsigned int irq, struct irq_d= esc *desc) remove_proc_entry("effective_affinity_list", desc->dir); # endif #endif + irq_moderation_procfs_remove(desc); remove_proc_entry("spurious", desc->dir); =20 snprintf(name, MAX_NAMELEN, "%u", irq); --=20 2.55.0.737.g08866a6d13-goog From nobody Mon Sep 28 17:48:02 2026 Received: from mail-ej1-f69.google.com (mail-ej1-f69.google.com [209.85.218.69]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CB74D30676C for ; Wed, 19 Aug 2026 12:44:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.69 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143449; cv=none; b=D+HOPX4FyS6aAq3DDuc3Y7PPoPuRgixrY/y0r6RiZkzvHafVuwlSwb0oiPYZVwFluwBvROVQRf0YySZ0RreZbl45pyuUg6lQB7/NAL4ei/Mk7FPn99oPLc+50/UQoURc9VZoQ4PyEl6vAh2CS17YCi6mJTfSvj9iJ1FGlJvMFwg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143449; c=relaxed/simple; bh=ovoju7LxqUNwF1NpKHUrgY7Zv0XswIrL5zt3gNzoZn8=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=OsXl5r6sjAdeA5RK7Yvc8YtjXUe5C9K259B1QuVvew9wNqL+nhhcRQrtNEqGyaBhqBqCw4nwMEVw0QLKEo6eNR2TaAAtn9XIDCHPKUtWyrphJ45/4NKbK8mTxbyiaMOLLhWcoi23H7OtyIW/HX+sdIUH4bYIMDa6s8IMHvT+JN8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=KDuBCmoM; arc=none smtp.client-ip=209.85.218.69 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="KDuBCmoM" Received: by mail-ej1-f69.google.com with SMTP id a640c23a62f3a-c16740eb587so73993766b.1 for ; Wed, 19 Aug 2026 05:44:05 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787143443; x=1787748243; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=PonVhEuAEKT0E4a1yU7TgphmO+Mvdl6szG3C0USyxek=; b=KDuBCmoM/7V2rAWB5STghO9OvKnC5Jf0UkLa7BkmOnCTWo+TM3NlwkJ2rxVVPoYSRJ RVKAX15+xv+SBRFOCzZQKvCr+wBMTsGZcZM0LYhawFiMxsQapWfuvlmFV1gkuc2l1uzY wbBML9KOr5L/fZiTw8P5EJc0kTgv6wNm2hCwBDnB4TEehnLAiSvh6J82Nu7zofwKZ8h8 tmp9q9x/VidzmdcNZNWGBG6shkfrZ816JsXAyxubyL8sR3XEbv3bkbA9PFmPNSRfoYNR 7Kv7S2RUSZL/21MO8DL2Bvx2JoleZzv6I486mtb/KJJNrIFg7NIN+eC7ssa2i2S3C6RF TYMA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787143443; x=1787748243; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=PonVhEuAEKT0E4a1yU7TgphmO+Mvdl6szG3C0USyxek=; b=q0U4uIL7kYPcfsC+sYya1BjxN6okS/+tKPVDc4sfz8vmZ4dTVrbTMJtyNCeWF2PoLO no1Mr4jjKTJSFle4R8EWuihgY/35s7l9GxKUpgTnpt2s+fHz6IZRlIM1af/bvq0tGXSP ySqw+3XbuIyO/XJWB6bMK45LoC0HJRq2L52CzncJxKxaZtuuQvAYngd2JXOpMzuNlKu0 31QdNxoy7WldEWls/041mmeDNYlvIzg8KF2TcHRQai5vz9pWO1G6WP8Afj+Dv2GZ449k 6iAJD7Ll+WegE61KBC6pd2ohT7us5l1mUhmmVDNChBept7o5tM+UMRAwdSXE+7uubc9v jkyw== X-Gm-Message-State: AOJu0YxizugrXp+3mK87tt/L1ruRA66DCp+SBa8iMzRSyP9szxfBQaO9 PFkJ5oNK9XQiOFj6vpyPxwP565k+mj8cPyVja6bHv/13QYdNjm1TRBPtCY9oJqBMpDEUwf//UBc Z5yd13w== X-Received: from ejbjp20.prod.google.com ([2002:a17:906:f754:b0:c21:148:79d2]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6938:a088:10b0:c24:47b8:ae5b with SMTP id a640c23a62f3a-c2447b8af3dmr85868166b.22.1787143443037; Wed, 19 Aug 2026 05:44:03 -0700 (PDT) Date: Wed, 19 Aug 2026 12:43:40 +0000 In-Reply-To: <20260819124341.4185621-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260819124341.4185621-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.737.g08866a6d13-goog Message-ID: <20260819124341.4185621-7-lrizzo@google.com> Subject: [PATCH v5 6/7] genirq: Adaptive Global Software Interrupt Moderation (GSIM). From: Luigi Rizzo To: Thomas Gleixner , Marc Zyngier , Luigi Rizzo , Paolo Abeni Cc: linux-kernel@vger.kernel.org, linux-pci@vger.kernel.org, Bjorn Helgaas , Luigi Rizzo Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" GSIM helps keeping system load (as defined later) under control, but the use of a fixed moderation delay causes unnecessary latency when the system load is already low. We focus on two types of system load related to interrupts: - total device interrupt rates. Some platforms show severe I/O performance degradation with more than 1-2 Mintr/s across the entire system. This affects especially server-class hardware with hundreds of interrupt sources (NIC or SSD queues from many physical or virtualized devices). - percentage of time spent in hardirq. This can become problematic in presence of large numbers of interrupt sources hitting individual CPUs (for instance, as a result of attempts to isolate application CPUs from interrupt load). Make GSIM adaptive by measuring total and per-CPU interrupt rates, as well as time spent in hardirq on each CPU. The metrics are compared to configurable targets, and each CPU adjusts the moderation delay up or down depending on the result, using multiplicative increase/decrease. Configuration of moderation parameters is done via procfs echo ${VALUE} > /proc/irq/sw_moderation/${NAME} Parameters are: delay_us (0 off, range 0-500) Maximum moderation delay in microseconds. target_intr_rate (0 off, range 0-50000000) Total rate above which moderation should trigger. hardirq_percent (0 off,range 0-100) Percentage of time spent in hardirq above which moderation should trigger. update_ms (range 1-100, default 5) How often the metrics should be computed and moderation delay updated. When target_intr_rate and hardirq_percent are both 0, GSIM uses delay_us as fixed moderation delay. Otherwise, the delay is dynamically adjusted up or down, independently on each CPU, based on how the total and per-CPU metrics compare with the targets. Provided that delay_us suffices to bring the metrics within the target, the control loop will dynamically converge to the minimum actual moderation delay to stay within the target. PERFORMANCE BENEFITS: The tests below demonstrate how adaptive moderation allows improved throughput at high load (same as fixed moderation) and low latency ad moderate load (same as no moderation) without having to hand-tune the system based on load. We run experiments on one x86 platform with 8 SSD (64 queues) capable of an aggregate of approximately 14M IOPS running fio with variable number of SSD devices (1 or 8), threads per disk (1 or 200), and IODEPTH (from 1 to 128 per thread) the actual command is below: ${FIO} --name=3Dread_IOPs_test ${DEVICES} --iodepth=3D${IODEPTH} --numjobs= =3D${JOBS} \ --bs=3D4K --rw=3Drandread --filesize=3D1000G --ioengine=3Dlibaio --dir= ect=3D1 \ --verify=3D0 --randrepeat=3D0 --time_based=3D1 --runtime=3D600s \ --cpus-allowed=3D1-119 --cpus_allowed_policy=3Dsplit --group_reporting= =3D1 For each configuration we test three moderations settings: - OFF: delay_us=3D0 - FIXED: delay_us=3D200 target_intr_rate=3D0 hardirq_percent=3D0 - ADAPTIVE: delay_us=3D200 target_intr_rate=3D1000000 hardirq_percent=3D70 The first set of measurements is for ONE DISK, ONE THREAD. At low IODEPTH the throughput is latency bound, and moderation is not necessary. A fixed moderation delay dominates the latency hence reducing throughput; adaptive moderation avoids the problem. As the IODEPTH increases, the system becomes I/O bound, and even the fixed moderation delay does not harm. Overall: adaptive moderation is better than fixed moderation and at least as good as moderation off. ------ OFF -------- ------ FIXED ------ ----- ADAPTIVE ---- IODEPTH IOPS p50 p90 p99 IOPS p50 p90 p99 IOPS p50 p90 p99 ------------------- ------------------- ------------------- 1: 12K 78 88 94 . 5K 208 210 215 . 12K 78 88 96 8: 94K 83 91 110 . 38K 210 212 221 . 94K 83 91 105 32: 423K 72 85 124 . 150K 210 219 235 . 424K 72 85 124 128: 698K 180 200 243 . 513K 251 306 347 . 718K 174 194 239 A second set of measurements is with one disk and 200 threads. The system is I/O bound but without significant interrupt overhead. All three scenarios are basically equivalent. --------- OFF -------- ------- FIXED ------- ------ ADAPTIVE -= ---- IODEPTH IOPS p50 p90 p99 IOPS p50 p90 p99 IOPS p50 p90 = p99 ---------------------- --------------------- -----------------= ---- 1: 1581K 110 174 281 . 933K 208 223 363 . 1556K 114 176 = 277 8: 1768K 889 1516 1926 . 1768K 848 1516 2147 . 1768K 889 1516 = 1942 32: 1768K 3589 5735 7701 . 1768K 3589 5735 7635 . 1768K 3589 5735 = 7504 128: 1768K 14ms 24ms 31ms . 1768K 14ms 24ms 29ms . 1768K 14ms 24ms = 30ms Finally, we have one set of measurements with 8 disks and 200 threads per disk, all running on socket 0. The system would be I/O bound (and CPU/latency bound at low IODEPTH), but this platform is unable to cope with= the high global interrupt rate and so moderation is really necessary to hit the disk limits. As we see below, adaptive moderation gives more than 2X higher throughput at meaningful iodepths, and even latency is much better. The only case where we see a small regression is with iodepth=3D1, because the high interrupt rate triggers the control loop to increase the moderation delay. --------- OFF -------- -------- FIXED ------- ------ ADAPTIVE = ------ IODEPTH IOPS p50 p90 p99 IOPS p50 p90 p99 IOPS p50 p9= 0 p99 ---------------------- ---------------------- ----------------= ------ 1: 2304K 82 94 128 . 1030K 208 277 293 . 1842K 97 14= 9 188 8: 5240K 128 938 1680 . 7500K 208 233 343 . 10000K 151 21= 0 281 32: 5251K 206 3621 3949 . 12300K 184 1106 5407 . 12100K 184 113= 9 5407 128: 5330K 4228 12ms 17ms . 13800K 1123 4883 7373 . 13800K 1074 488= 3 7635 Finally, here are experiments indicating how throughput is affected by the various parameters (with 8 disks and 200 threads). IOPS vs delay_us (target_intr_rate =3D 0, hardirq_percent=3D0) delay_us 0 50 100 150 200 250 IODEPTH 1 2300 1860 1580 1300 1034 1066 IODEPTH 8 5254 9936 9645 8818 7500 6150 IODEPTH 32 5250 11300 13800 13800 13800 13800 IODEPTH 128 5900 13600 13900 13900 13900 13900 IOPS vs target_intr_rate (delay_us =3D 200, hardirq_percent=3D0, iodepth 12= 8) value 250K 500K 750K 1000K 1250K 1500K 1750K 2000K socket0 13900 13900 13900 13800 13800 12900 11000 8808 both sockets 13900 13900 13900 13800 8600 8000 6900 6400 hardirq_percent (delay_us =3D 200, target_intr_rate=3D0, iodepth 128) hardirq% 1 10 20 30 40 50 60 70 KIOPS 13900 13800 13300 12100 10400 8600 7400 6500 Signed-off-by: Luigi Rizzo --- kernel/irq/irq_moderation.c | 391 ++++++++++++++++++++++++++++++++++-- kernel/irq/irq_moderation.h | 63 ++++++ 2 files changed, 442 insertions(+), 12 deletions(-) diff --git a/kernel/irq/irq_moderation.c b/kernel/irq/irq_moderation.c index 1474d33455410..74bd33c786c90 100644 --- a/kernel/irq/irq_moderation.c +++ b/kernel/irq/irq_moderation.c @@ -9,6 +9,9 @@ #include #include #include +#include +#include +#include #include #include #include @@ -23,8 +26,9 @@ * * Some platforms show reduced I/O performance when the total device inter= rupt * rate across the entire platform becomes too high. To address the proble= m, - * GSIM runs after the handler to implement software interrupt moderation - * with programmable delay. + * GSIM runs after the handler to measure global and per-CPU interrupt rat= es, + * compares them with configurable targets, and implements independent, pe= r-CPU + * software moderation delays. * * Configuration is done at runtime via procfs * echo ${VALUE} > /proc/irq/sw_moderation/${NAME} @@ -35,6 +39,19 @@ * Maximum moderation delay. A reasonable range is 20-100. Higher va= lues * can be useful if the hardirq handler has long runtimes. * + * target_intr_rate (default 0, suggested 1000000, 0 off, range 0-500000= 00) + * The total interrupt rate above which moderation kicks in. + * Not particularly critical, a value in the 500K-1M range is usuall= y ok. + * + * hardirq_percent (default 0, suggested 70, 0 off, range 0-100) + * The hardirq percentage above which moderation kicks in. + * 50-90 is a reasonable range. + * + * FIXED MODERATION mode requires target_intr_rate=3D0, hardirq_perc= ent=3D0 + * + * update_ms (default 5, range 1-100) + * How often the load is measured and moderation delay updated. + * * Moderation is allowed/disallowed dynamically for individual interrupts = with * echo 1 > /proc/irq/NN/allow_sw_moderation # use 0 to disallow * @@ -112,7 +129,47 @@ * GSIM parameters. Initialize delay_ns here to statically enable moderati= on * (e.g. .delay_ns =3D 100000). */ -struct irq_mod_params irq_mod_params ____cacheline_aligned; + +/* + * Recommended values for the adaptive control loop. + * + * update_ns is documented earlier and can be modified via procfs + * + * The following two allow fine tuning of the control loop and should not + * be modified unless there is good understanding of their impact on the + * stability of the controller. + * + * scale_cpus (default 150, range 50-1000) + * Small update_ms may lead to underestimate the number of CPUs + * simultaneously handling interrupts, and the opposite can happen + * with very large values. This parameter may help correct the value, + * though it is not recommended to modify the default unless there are + * very strong reasons. + * + * increase_divisor (default 8, range 8-128) + * This is base parameter used for multiplicative increase/decrease. + */ + +#define MIN_SCALING_DIVISOR 8 + +struct irq_mod_params irq_mod_params ____cacheline_aligned =3D { + .update_ns =3D 5 * NSEC_PER_MSEC, + .scale_cpus =3D 150, + .increase_divisor =3D MIN_SCALING_DIVISOR, + .seq =3D SEQCNT_ZERO(irq_mod_params.seq), +}; + +/* + * Accumulator for total interrupt and active CPUs, updated by all active + * CPUs on each epoch (update_ns or more). + * @total_intrs: running count of total interrupts + * @total_cpus: running count of total active CPUs + */ +struct irq_mod_counters { + atomic_t total_intrs; + atomic_t total_cpus; +}; +static struct irq_mod_counters irq_mod_counters; =20 DEFINE_PER_CPU_ALIGNED(struct irq_mod_state, irq_mod_state); =20 @@ -122,12 +179,215 @@ static DEFINE_MUTEX(swmod_mutex); =20 static void update_enable_key(void) { + lockdep_assert_held(&swmod_mutex); + if (irq_mod_params.delay_ns !=3D 0) static_branch_enable(&irq_moderation_enabled_key); else static_branch_disable(&irq_moderation_enabled_key); } =20 +/* Slow path functions for interrupt moderation. */ + +/* + * Compute smoothed average between old and cur. 'steps' is used + * to approximate applying the smoothing multiple times. + */ +static inline unsigned int smooth_avg(unsigned int old, unsigned int cur, = unsigned int steps) +{ + const unsigned int smooth_factor =3D 64; + u64 sum; + + steps =3D min(steps, smooth_factor - 1); + sum =3D (u64)(smooth_factor - steps) * old + (u64)steps * cur; + return div_u64(sum, smooth_factor); +} + +/* Measure and assess time spent in hardirq. */ +static inline bool hardirq_high(struct irq_mod_state *m, unsigned int hard= irq_percent, + u64 epoch_ns) +{ + bool above_threshold; + u64 irqtime, cur; + + if (!IS_ENABLED(CONFIG_IRQ_TIME_ACCOUNTING)) + return false; + + cur =3D kcpustat_this_cpu->cpustat[CPUTIME_IRQ]; + irqtime =3D cur - m->last_irqtime; + m->last_irqtime =3D cur; + + if (hardirq_percent =3D=3D 0) + return false; + + above_threshold =3D irqtime * 100 > epoch_ns * hardirq_percent; + m->hardirq_high +=3D above_threshold; + return above_threshold; +} + +/* Measure and assess total and per-CPU interrupt rates. */ +static inline bool irqrate_high(struct irq_mod_state *m, unsigned int targ= et_rate, + unsigned int steps, u64 epoch_ns, + unsigned int update_ns, unsigned int scale_cpus) +{ + unsigned int global_intr_rate, local_intr_rate, delta_intrs, tmp; + bool local_rate_high, global_rate_high; + u64 num_local, num_global, num_cpus; + /* Use unsigned long to avoid overflow in intermediate results. */ + unsigned long active_cpus; + u32 denom =3D epoch_ns; + int shift =3D 0; + + num_local =3D (u64)m->intr_count * NSEC_PER_SEC; + /* Scale denominator so we can avoid 64-bit division. */ + if (unlikely(epoch_ns > U32_MAX)) { + shift =3D fls64(epoch_ns) - 32; + denom =3D epoch_ns >> shift; + num_local >>=3D shift; + } + + local_intr_rate =3D div_u64(num_local, denom); + + /* Accumulate global counter and compute global interrupt rate. */ + tmp =3D atomic_add_return(m->intr_count, &irq_mod_counters.total_intrs); + m->intr_count =3D 0; + delta_intrs =3D tmp - m->last_total_intrs; + m->last_total_intrs =3D tmp; + num_global =3D (u64)delta_intrs * NSEC_PER_SEC; + if (unlikely(shift)) + num_global >>=3D shift; + global_intr_rate =3D div_u64(num_global, denom); + + /* + * Count how many CPUs handled interrupts in the last epoch, needed + * to determine the per-CPU target (target_rate / active_cpus). + * Each active CPU increments the global counter approximately every + * update_ns. Scale the value by (update_ns / epoch_ns) to get the + * correct value. Also apply rounding and make sure active_cpus > 0. + */ + tmp =3D atomic_add_return(1, &irq_mod_counters.total_cpus); + active_cpus =3D tmp - m->last_total_cpus; + m->last_total_cpus =3D tmp; + num_cpus =3D (u64)active_cpus * update_ns + (epoch_ns / 2); + if (unlikely(shift)) + num_cpus >>=3D shift; + active_cpus =3D div_u64(num_cpus, denom); + if (active_cpus < 1) + active_cpus =3D 1; + + /* Compare with global and per-CPU targets. */ + global_rate_high =3D global_intr_rate > target_rate; + + /* + * Short epochs may lead to underestimate the number of active CPUs. + * Apply a scaling factor to compensate. This may make the controller + * a bit more aggressive but does not harm system throughput. + */ + local_rate_high =3D (u64)local_intr_rate * active_cpus * + scale_cpus > (u64)target_rate * 100; + + /* Statistics. */ + m->global_intr_rate =3D smooth_avg(m->global_intr_rate, global_intr_rate,= steps); + m->local_intr_rate =3D smooth_avg(m->local_intr_rate, local_intr_rate, st= eps); + m->scaled_cpu_count =3D smooth_avg(m->scaled_cpu_count, active_cpus * 256= , steps); + + if (target_rate =3D=3D 0) + return false; + + m->local_irq_high +=3D local_rate_high; + m->global_irq_high +=3D global_rate_high; + + /* Moderate on this CPU only if both global and local rates are high. */ + return global_rate_high && local_rate_high; +} + +/* Periodic adjustment, called once per epoch. */ +void irq_moderation_update_epoch(struct irq_mod_state *m, u64 epoch_ns) +{ + unsigned int hardirq_percent, target_rate, delay_ns, update_ns; + unsigned int increase_divisor, scale_cpus; + const unsigned int min_delay_ns =3D 500; + bool above_target =3D false; + unsigned int steps, seq; + + do { + seq =3D read_seqcount_begin(&irq_mod_params.seq); + hardirq_percent =3D READ_ONCE(irq_mod_params.hardirq_percent); + target_rate =3D READ_ONCE(irq_mod_params.target_intr_rate); + delay_ns =3D READ_ONCE(irq_mod_params.delay_ns); + update_ns =3D READ_ONCE(irq_mod_params.update_ns); + increase_divisor =3D READ_ONCE(irq_mod_params.increase_divisor); + scale_cpus =3D READ_ONCE(irq_mod_params.scale_cpus); + } while (read_seqcount_retry(&irq_mod_params.seq, seq)); + + /* + * If one parameter changes, set the moderation delay to max, and rely + * on the adaptive mechanism to adjust it down if necessary. + * Otherwise the system may be stuck with an interrupt rate that is + * already below the threshold because of bus congestion (one of the + * problems that GSIM is trying to address), and the controller would + * have no signal react. Starting from a high value gives it a chance + * to converge if parameters allow it. + */ + if (seq !=3D m->seq) { + m->seq =3D seq; + m->mod_ns =3D delay_ns; + m->intr_count =3D 0; + m->last_total_intrs =3D atomic_read(&irq_mod_counters.total_intrs); + m->last_total_cpus =3D atomic_read(&irq_mod_counters.total_cpus); + if (IS_ENABLED(CONFIG_IRQ_TIME_ACCOUNTING)) + m->last_irqtime =3D kcpustat_this_cpu->cpustat[CPUTIME_IRQ]; + return; + } + + if (target_rate =3D=3D 0 && hardirq_percent =3D=3D 0) { + /* Use fixed moderation delay. */ + m->mod_ns =3D delay_ns; + m->global_intr_rate =3D 0; + m->local_intr_rate =3D 0; + m->scaled_cpu_count =3D 0; + return; + } + + /* + * The controller wants to scale the delay mod_ns by (1 + 1/D) every "upd= ate_ns". + * Since we operate every epoch_ns >=3D update_ns, the formula becomes + * mod_ns =3D mod_ns * ((1 + 1/D) ** (epoch_ns / update_ns)) + * which we approximate with "mod_ns =3D mod_ns * (1 + steps/D)" + * where "steps =3D epoch_ns / update_ns" clamped to a value < D. + */ + steps =3D (unsigned int)clamp_t(u64, div_u64(epoch_ns, update_ns), + 1ULL, (u64)(MIN_SCALING_DIVISOR - 1u)); + + if (irqrate_high(m, target_rate, steps, epoch_ns, update_ns, scale_cpus)) + above_target =3D true; + + if (hardirq_high(m, hardirq_percent, epoch_ns)) + above_target =3D true; + + /* + * Controller: adjust delay with exponential increase or decrease. + * + * Following standard practices, we increase fast (smaller divisor) to + * aggressively slow down when the interrupt rate goes up, but decrease + * slowly (larger divisor) to reduce the chance of load spikes as the + * delay goes down. + */ + if (above_target) { + /* Make sure the value is large enough for the exponential to grow. */ + if (m->mod_ns < min_delay_ns) + m->mod_ns =3D min_delay_ns; + m->mod_ns +=3D m->mod_ns * steps / increase_divisor; + if (m->mod_ns > delay_ns) + m->mod_ns =3D delay_ns; + } else { + m->mod_ns -=3D m->mod_ns * steps / (2 * increase_divisor); + /* Round down to 0 values that are too small to bother. */ + if (m->mod_ns < min_delay_ns) + m->mod_ns =3D 0; + } +} + /* Actually start moderation. */ bool irq_moderation_do_start(struct irq_desc *desc, struct irq_mod_state *= m) { @@ -138,7 +398,7 @@ bool irq_moderation_do_start(struct irq_desc *desc, str= uct irq_mod_state *m) const u64 slack_ns =3D 2000; =20 /* Accumulate sleep time, no moderation if too small. */ - m->sleep_ns +=3D READ_ONCE(irq_mod_params.delay_ns); + m->sleep_ns +=3D m->mod_ns; if (m->sleep_ns < min_delay_ns) return false; /* We need moderation, start the timer. */ @@ -153,6 +413,11 @@ bool irq_moderation_do_start(struct irq_desc *desc, st= ruct irq_mod_state *m) */ m->enqueue++; list_add(&desc->swmod_state.swmod_node, &m->descs); + /* + * Set IRQD_IRQ_INPROGRESS so that synchronize_irq() called during + * free_irq() will block until the timer drains this descriptor from + * the moderation list. + */ irqd_set(&desc->irq_data, IRQD_IRQ_INPROGRESS | IRQD_MODERATED); __disable_irq(desc); return true; @@ -186,17 +451,31 @@ struct swmod_procfs_entry { struct var_info var; }; =20 +static void write_param(unsigned int *ptr, unsigned int value) +{ + unsigned long flags; + + local_irq_save(flags); + write_seqcount_begin(&irq_mod_params.seq); + WRITE_ONCE(*ptr, value); + write_seqcount_end(&irq_mod_params.seq); + local_irq_restore(flags); +} + static ssize_t swmod_wr(struct var_info *v, const char __user *s, size_t c= ount) { unsigned int value; int ret; =20 + lockdep_assert_held(&swmod_mutex); + ret =3D kstrtouint_from_user(s, count, 0, &value); if (ret) return ret; if (value < v->min || value > v->max) return -ERANGE; - WRITE_ONCE(*v->ptr, value * v->scale); + write_param(v->ptr, value * v->scale); + return count; } =20 @@ -216,34 +495,87 @@ static ssize_t swmod_wr_delay(struct var_info *v, con= st char __user *s, size_t c return ret; } =20 -#define HEAD_FMT "%5s %8s %11s %11s\n" -#define BODY_FMT "%5u %8u %11u %11u\n" +#define HEAD_FMT "%5s %8s %10s %4s %8s %11s %11s %11s %11s %11s\n" +#define BODY_FMT "%5u %8u %10u %4u %8u %11u %11u %11u %11u %11u\n" =20 /* Print statistics */ static void rd_stats(struct seq_file *p) { unsigned int delay_ns =3D READ_ONCE(irq_mod_params.delay_ns); - int cpu; + unsigned long global_intr_rate =3D 0, global_irq_high =3D 0; + unsigned long local_irq_high =3D 0, hardirq_high =3D 0; + int recent_epoch_limit, cpu, active_cpus =3D 0; =20 if (delay_ns =3D=3D 0) return; seq_printf(p, HEAD_FMT, - "# CPU", "delay_ns", "timer_set", "enqueue"); + "# CPU", "irq/s", "loc_irq/s", "cpus", "delay_ns", + "irq_hi", "loc_irq_hi", "hardirq_hi", "timer_set", + "enqueue"); + + /* + * Accumulate/print only entries updated within ~20-30s. The high 32 bits + * of timestamps give ~4s resolution, so we can use them without the need + * for 64bit atomics (because epoch_start_ns is updated concurrently). + */ + recent_epoch_limit =3D (ktime_get_ns() - 20ULL * NSEC_PER_SEC) >> 32; =20 for_each_possible_cpu(cpu) { /* Copy statistics, will only use some unsigned int values; races ok. */ struct irq_mod_state cur =3D data_race(*per_cpu_ptr(&irq_mod_state, cpu)= ); =20 + if (cur.epoch_start_ns && (int)(cur.epoch_start_ns >> 32) >=3D recent_ep= och_limit) { + /* Recent entry, accumulate in global rate. */ + active_cpus++; + global_intr_rate +=3D cur.global_intr_rate; + } else { + /* Stale entries, print as 0. */ + cur.global_intr_rate =3D 0; + cur.local_intr_rate =3D 0; + cur.scaled_cpu_count =3D 0; + cur.mod_ns =3D 0; + } + + global_irq_high +=3D cur.global_irq_high; + local_irq_high +=3D cur.local_irq_high; + hardirq_high +=3D cur.hardirq_high; + seq_printf(p, BODY_FMT, cpu, - delay_ns, + cur.global_intr_rate, + cur.local_intr_rate, + (cur.scaled_cpu_count + 128) / 256, + cur.mod_ns, + cur.global_irq_high, + cur.local_irq_high, + cur.hardirq_high, cur.timer_set, cur.enqueue); } =20 seq_printf(p, "\n" - "delay_us %lu\n", - delay_ns / NSEC_PER_USEC); + "delay_us %lu\n" + "target_intr_rate %u\n" + "hardirq_percent %u\n" + "update_ms %ld\n" + "scale_cpus %u\n", + delay_ns / NSEC_PER_USEC, + READ_ONCE(irq_mod_params.target_intr_rate), + READ_ONCE(irq_mod_params.hardirq_percent), + READ_ONCE(irq_mod_params.update_ns) / NSEC_PER_MSEC, + READ_ONCE(irq_mod_params.scale_cpus)); + + seq_printf(p, + "intr_rate %lu\n" + "irq_high %lu\n" + "my_irq_high %lu\n" + "hardirq_percent_high %lu\n" + "total_interrupts %u\n" + "total_cpus %u\n", + active_cpus ? global_intr_rate / active_cpus : 0, + global_irq_high, local_irq_high, hardirq_high, + atomic_read(&irq_mod_counters.total_intrs), + atomic_read(&irq_mod_counters.total_cpus)); } =20 static int param_show(struct seq_file *p, void *v) @@ -476,15 +808,41 @@ static struct swmod_procfs_entry procfs_entries[] =3D= { .rd =3D swmod_rd, .var =3D SET_VAR(&irq_mod_params.delay_ns, 0, 500, NSEC_PER_USEC), }, + { + .name =3D "target_intr_rate", + .wr =3D swmod_wr, + .rd =3D swmod_rd, + .var =3D SET_VAR(&irq_mod_params.target_intr_rate, 0, 50000000, 1), + }, + { + .name =3D "hardirq_percent", + .wr =3D swmod_wr, + .rd =3D swmod_rd, + .var =3D SET_VAR(&irq_mod_params.hardirq_percent, 0, 100, 1), + }, + { + .name =3D "update_ms", + .wr =3D swmod_wr, + .rd =3D swmod_rd, + .var =3D SET_VAR(&irq_mod_params.update_ns, 1, 100, NSEC_PER_MSEC), + }, { .name =3D "stats", .rd =3D rd_stats, }, + /* The next parameters have no procfs entries, only range validation. */ + { + .var =3D SET_VAR(&irq_mod_params.increase_divisor, MIN_SCALING_DIVISOR, = 128, 1), + }, + { + .var =3D SET_VAR(&irq_mod_params.scale_cpus, 50, 1000, 1), + }, }; =20 static int __init init_irq_moderation(void) { struct proc_dir_entry *dir; + unsigned long flags; int cpuhp_state; int i, ret; =20 @@ -535,8 +893,17 @@ static int __init init_irq_moderation(void) goto cleanup_1; } =20 + /* Increment sequence counter so per-CPU m->seq (0) mismatches on epoch 1= */ + local_irq_save(flags); + write_seqcount_begin(&irq_mod_params.seq); + write_seqcount_end(&irq_mod_params.seq); + local_irq_restore(flags); + /* Enable if the defaults require it. */ + /* Acquire swmod_mutex to satisfy lockdep assertions. */ + mutex_lock(&swmod_mutex); update_enable_key(); + mutex_unlock(&swmod_mutex); return 0; =20 cleanup_1: diff --git a/kernel/irq/irq_moderation.h b/kernel/irq/irq_moderation.h index 8df350651cd7c..007c7c8646698 100644 --- a/kernel/irq/irq_moderation.h +++ b/kernel/irq/irq_moderation.h @@ -15,13 +15,26 @@ #include #include #include +#include =20 /** * struct irq_mod_params - configuration parameters * @delay_ns: maximum delay + * @target_intr_rate: target maximum interrupt rate + * @hardirq_percent: target maximum hardirq percentage + * @update_ns: how often to update delay/rate/fraction (epoch duration) + * @increase_divisor: constant for multiplicative increase/decrease of del= ay + * @scale_cpus: (percent) scale factor to estimate active CPUs + * @seq: incremented every time parameters change */ struct irq_mod_params { unsigned int delay_ns; + unsigned int target_intr_rate; + unsigned int hardirq_percent; + unsigned int update_ns; + unsigned int increase_divisor; + unsigned int scale_cpus; + seqcount_t seq; }; =20 extern struct irq_mod_params irq_mod_params; @@ -34,22 +47,50 @@ extern struct irq_mod_params irq_mod_params; * @initialized: true if hrtimer and list head are initialized * @moderation_allowed: per-CPU flag, toggled during hotplug/suspend events * @sleep_ns: accumulated time for actual delay + * @mod_ns: dynamically computed moderation delay + * @intr_count: interrupt counter + * @epoch_start_ns: start time of current epoch * * Used once per moderation delay per interrupt source: * @descs: list of moderated irq_desc on this CPU * @enqueue: how many enqueue on the list * + * Used once per epoch: + * @seq: latest seq from irq_mod_info + * @last_total_intrs: from irq_mod_info + * @last_total_cpus: from irq_mod_info + * @last_irqtime: from cpustat[CPUTIME_IRQ] + * * Statistics + * @global_intr_rate: smoothed global interrupt rate + * @local_intr_rate: smoothed interrupt rate for this CPU * @timer_set: how many timer_set calls + * @scaled_cpu_count: smoothed CPU count (scaled) + * @global_irq_high: how many times global irq rate was above threshold + * @local_irq_high: how many times local irq rate was above threshold + * @hardirq_high: how many times local hardirq_percent was above threshold */ struct irq_mod_state { struct hrtimer timer; bool initialized; bool moderation_allowed; unsigned int sleep_ns; + unsigned int mod_ns; + unsigned int intr_count; + u64 epoch_start_ns; struct list_head descs; unsigned int enqueue; + unsigned int seq; + unsigned int last_total_intrs; + unsigned int last_total_cpus; + u64 last_irqtime; + unsigned int global_intr_rate; + unsigned int local_intr_rate; unsigned int timer_set; + unsigned int scaled_cpu_count; + unsigned int global_irq_high; + unsigned int local_irq_high; + unsigned int hardirq_high; }; =20 DECLARE_PER_CPU_ALIGNED(struct irq_mod_state, irq_mod_state); @@ -67,6 +108,26 @@ static inline bool mod_state_initialized(struct irq_mod= _state *m) extern struct static_key_false irq_moderation_enabled_key; =20 bool irq_moderation_do_start(struct irq_desc *desc, struct irq_mod_state *= m); +void irq_moderation_update_epoch(struct irq_mod_state *m, u64 epoch_ns); + +static inline void check_epoch(struct irq_mod_state *m) +{ + const unsigned int slack_ns =3D 5000; + u64 now, epoch_ns; + + /* Don't check too often, fetching time is moderately expensive. */ + if ((m->intr_count & 0xf) !=3D 0) + return; + now =3D ktime_get_ns(); + epoch_ns =3D now - m->epoch_start_ns; + + /* Run approximately every update_ns, a little bit early is ok. */ + if (epoch_ns < READ_ONCE(irq_mod_params.update_ns) - slack_ns) + return; + WRITE_ONCE(m->epoch_start_ns, now); + /* Do the expensive processing. */ + irq_moderation_update_epoch(m, epoch_ns); +} =20 /* * Call after running the handler, with lock held. If this source should be @@ -81,6 +142,8 @@ static inline bool irq_start_moderation(struct irq_desc = *desc) if (static_branch_unlikely(&irq_moderation_enabled_key) && irq_settings_moderatable(desc) && m->moderation_allowed) { + m->intr_count++; + check_epoch(m); return irq_moderation_do_start(desc, m); } return false; --=20 2.55.0.737.g08866a6d13-goog From nobody Mon Sep 28 17:48:02 2026 Received: from mail-ed1-f71.google.com (mail-ed1-f71.google.com [209.85.208.71]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E0BF1433E7E for ; Wed, 19 Aug 2026 12:44:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.71 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143448; cv=none; b=XNCP8VQgKZBlq4fFxtiPqJHfmxzpfpKHRKDrEpfO3CKtVUWoZJs8JxkP6UeXbniEnfPjICPgEZ+hqzbk0IwYl2hu1K9fEurSj4A4BmyxI8zdlmAOONsCXDkKwZa113C4bRQe7Koc2uMEyiUN6gliDOOAY5Uto+npdS/cQ0XuL9Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787143448; c=relaxed/simple; bh=W407ChCnhP/v+ZGkJQPC1qqf5wL2ia/l4iQolLSkd9I=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=mE7WZsQY6KLn+fY63d7GCR6A7EmtIJA8WOYRPj8b1fQoTzaj9sIaWKIS7W1KWsBg2TSLNBRZ4shKPYZV41CfXlse0vB/nq78qNQ4/bpa3pCyZBP1lWad9iicJkTee42Nh0AnifPDWeD+l1Ir4Lu8huq9FY+9pAlqMiEjqiOgRPw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=UirvIaOL; arc=none smtp.client-ip=209.85.208.71 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--lrizzo.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="UirvIaOL" Received: by mail-ed1-f71.google.com with SMTP id 4fb4d7f45d1cf-6a36683ee94so962923a12.0 for ; Wed, 19 Aug 2026 05:44:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787143445; x=1787748245; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=9C0y7TvBKpzQlhXfTgiBd4YtOIaPrxO67kIoX8b8oug=; b=UirvIaOLD54sFxmQAWnlN7jkz06e24VOwYrb2HObsx92DT7actarRvyhDZLxNe5fW5 YKlRCFcPhs6xE8L39k/6oNNG/zYNGPpSAhp/MXr8IWliiko46fJafk3ZrycNGkgpMtzS D5lcSU5Qqs+/JX5lsTq6czx42BHVfS92KiyU5HI2AZYgcpaIcCzoraCQsZsZRt+dIhOy M0W73KCsDP52ptASZq3vck6LSE21zsV/LWfgqfTrQiD6nbEoiYCofgMYg4LlVc0rcRXX IphYghAX29QZb/l22NtcnaUZTs3a/9RoRyYzvovHzDFKsRDsC/APY9VDR+4u8/+eYVd8 k9+g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787143445; x=1787748245; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=9C0y7TvBKpzQlhXfTgiBd4YtOIaPrxO67kIoX8b8oug=; b=crVyQqbHylBCb/277CxOcobYrpUu8+AS8+Gqoy0sxkKIg5yX3bvoOWCaMhLpqzOpDP d/kf4NgVh0VGTWxXM84Zm6uBg5xr2L1Yi/Edh079nEfClntt8EeRL6G1U8u21B4Y5D2q kUAQypLVwqKn3pwmjnSLpurRkxn3T4DRRWGusxPmI5gsRA43KISd2a5qHoX8tJBbU4VU bi9OuT+4g+iIbDE+aRUFyke91zOg0xVgFhjbRXt0vj3ImxOzkRsmEWRIhoUtKLZcJD57 XHRFocG90pNc3MCQMjwcZFpXDEqN8aeU/+ti4WJghqoDt2iOtyJqqlRY7ouKjoE+7fSX a9xQ== X-Gm-Message-State: AOJu0YzLvqEWqumUHJuMwt8BkDNfUHHhS0o9ZBCPPLdmLJlz5Me6f5eK CmlXqNYj6J3ZKZ7trNmzVRcmnELgFuuJdHXP6nJJxGKC/LlTjLFUhRNhaJx2DT+jwgu8lCxQM67 0Zck6qQ== X-Received: from edsn13.prod.google.com ([2002:aa7:c78d:0:b0:6a3:8183:78d5]) (user=lrizzo job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6402:5298:b0:698:6084:db7f with SMTP id 4fb4d7f45d1cf-6a4032f21c3mr2807598a12.10.1787143444457; Wed, 19 Aug 2026 05:44:04 -0700 (PDT) Date: Wed, 19 Aug 2026 12:43:41 +0000 In-Reply-To: <20260819124341.4185621-1-lrizzo@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260819124341.4185621-1-lrizzo@google.com> X-Mailer: git-send-email 2.55.0.737.g08866a6d13-goog Message-ID: <20260819124341.4185621-8-lrizzo@google.com> Subject: [PATCH v5 7/7] PCI/MSI: re-enable conditional parent mask/unmask with sw moderation From: Luigi Rizzo To: Thomas Gleixner , Marc Zyngier , Luigi Rizzo , Paolo Abeni Cc: linux-kernel@vger.kernel.org, linux-pci@vger.kernel.org, Bjorn Helgaas , Luigi Rizzo Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Recent commits removed masking of interrupts at the PCIe level when the masking is also done at the parent (e.g. on ARM with GIC). This optimization saves the PCIe write, and especially the expensive readback, but causes the device to still send MSI interrupts that are ignored. As a result, it very harmful on platforms (many, for multiple architectures) that are unable to deal with too many MSI interrupts, and defeats GSIM (CONFIG_IRQ_SW_MODERATION), which is esplicitely designed to address this problem, but depends on masking interrupts at the PCIe level. Revert to the old behavior when GSIM (CONFIG_IRQ_SW_MODERATION) is compiled in. Signed-off-by: Luigi Rizzo --- drivers/irqchip/irq-msi-lib.c | 8 ++++++++ drivers/pci/msi/irqdomain.c | 20 ++++++++++++++++++++ 2 files changed, 28 insertions(+) diff --git a/drivers/irqchip/irq-msi-lib.c b/drivers/irqchip/irq-msi-lib.c index 45e0ed3134ce1..453a5d5a2c106 100644 --- a/drivers/irqchip/irq-msi-lib.c +++ b/drivers/irqchip/irq-msi-lib.c @@ -116,6 +116,14 @@ bool msi_lib_init_dev_msi_info(struct device *dev, str= uct irq_domain *domain, if (!chip->irq_set_affinity && !(info->flags & MSI_FLAG_NO_AFFINITY)) chip->irq_set_affinity =3D msi_domain_set_affinity; =20 + /* + * The problem addressed by software moderation depends on actually + * masking at the PCI level. irq_mask() can be optimized separately + * for the common cases outside the move/teardown paths. + */ + if (IS_ENABLED(CONFIG_IRQ_SW_MODERATION)) + return true; + /* * If the parent domain insists on being in charge of masking, obey * blindly. The interrupt is un-masked at the PCI level on startup diff --git a/drivers/pci/msi/irqdomain.c b/drivers/pci/msi/irqdomain.c index 6e65f0f44112e..44b022fee7f1a 100644 --- a/drivers/pci/msi/irqdomain.c +++ b/drivers/pci/msi/irqdomain.c @@ -80,6 +80,22 @@ static unsigned int cond_startup_parent(struct irq_data = *data) return 0; } =20 +static __always_inline void cond_mask_parent(struct irq_data *data) +{ + struct msi_domain_info *info =3D data->domain->host_data; + + if (unlikely(info->flags & MSI_FLAG_PCI_MSI_MASK_PARENT)) + irq_chip_mask_parent(data); +} + +static __always_inline void cond_unmask_parent(struct irq_data *data) +{ + struct msi_domain_info *info =3D data->domain->host_data; + + if (unlikely(info->flags & MSI_FLAG_PCI_MSI_MASK_PARENT)) + irq_chip_unmask_parent(data); +} + static void pci_irq_shutdown_msi(struct irq_data *data) { struct msi_desc *desc =3D irq_data_get_msi_desc(data); @@ -102,12 +118,14 @@ static void pci_irq_mask_msi(struct irq_data *data) struct msi_desc *desc =3D irq_data_get_msi_desc(data); =20 pci_msi_mask(desc, BIT(data->irq - desc->irq)); + cond_mask_parent(data); } =20 static void pci_irq_unmask_msi(struct irq_data *data) { struct msi_desc *desc =3D irq_data_get_msi_desc(data); =20 + cond_unmask_parent(data); pci_msi_unmask(desc, BIT(data->irq - desc->irq)); } =20 @@ -160,10 +178,12 @@ static unsigned int pci_irq_startup_msix(struct irq_d= ata *data) static void pci_irq_mask_msix(struct irq_data *data) { pci_msix_mask(irq_data_get_msi_desc(data)); + cond_mask_parent(data); } =20 static void pci_irq_unmask_msix(struct irq_data *data) { + cond_unmask_parent(data); pci_msix_unmask(irq_data_get_msi_desc(data)); } =20 --=20 2.55.0.737.g08866a6d13-goog