[PATCH v2] x86/ucode: Work around Granite Rapids erraturm GNR98

Andrew Cooper posted 1 patch 2 weeks, 2 days ago
There is a newer version of this series
xen/arch/x86/cpu/microcode/intel.c | 40 ++++++++++++++++++++++++++++++
1 file changed, 40 insertions(+)
[PATCH v2] x86/ucode: Work around Granite Rapids erraturm GNR98
Posted by Andrew Cooper 2 weeks, 2 days ago
Block loads which are known to hang the system.

Signed-off-by: Andrew Cooper <andrew.cooper3@citrix.com>
---
CC: Jan Beulich <jbeulich@suse.com>
CC: Roger Pau Monné <roger@xenproject.org>
CC: Teddy Astie <teddy.astie@vates.tech>

A more complete solution is in the works, but it's taken 4 months to get this
much published...

v2:
 * Correct the sign of the cpu_sig->rev check.
 * Expand the comment to explain why we are not following what GNR98 says.
---
 xen/arch/x86/cpu/microcode/intel.c | 40 ++++++++++++++++++++++++++++++
 1 file changed, 40 insertions(+)

diff --git a/xen/arch/x86/cpu/microcode/intel.c b/xen/arch/x86/cpu/microcode/intel.c
index c45b00c6b033..dc21a89aa47b 100644
--- a/xen/arch/x86/cpu/microcode/intel.c
+++ b/xen/arch/x86/cpu/microcode/intel.c
@@ -27,6 +27,7 @@
 #include <xen/string.h>
 #include <xen/xmalloc.h>
 
+#include <asm/intel-family.h>
 #include <asm/msr.h>
 #include <asm/processor.h>
 #include <asm/system.h>
@@ -273,6 +274,44 @@ static bool microcode_fits_cpu(const struct microcode_patch *mc)
     return false;
 }
 
+static bool microcode_safe_to_load(const struct microcode_patch *mc)
+{
+    struct cpu_signature *cpu_sig = &this_cpu(cpu_sig);
+
+    /*
+     * Treat pre-production as always safe - anyone using pre-production
+     * microcode knows what they are doing, and can keep any resulting pieces.
+     */
+    if ( (int)cpu_sig->rev < 0 || mc->rev < 0 )
+        return true;
+
+    /*
+     * GNR98 states that Granite Rapids systems hang when loading new ucode on
+     * sufficiently old firmware.  GNR101 retroactively declares that one
+     * ucode had incorrect min_rev fields, in light of discovering GNR98.
+     *
+     * Both are incomplete statements of the problem.
+     *
+     * At the time of writing (August 2026), the believed safe sequence is:
+     *   0x01000370 -> [0x01000380...0x010003f3] -> 0x01000405 -> any later
+     *
+     * Disallow known-unsafe loads while permitting believed-safe loads.  For
+     * GNR, this allows multi-hop loading to get up to the latest.
+     */
+    if ( boot_cpu_data.vfm == INTEL_GRANITERAPIDS_X &&
+         boot_cpu_data.stepping == 1 && (cpu_sig->pf & 0x95) &&
+         ((cpu_sig->rev < 0x01000380 && mc->rev >= 0x01000405) ||
+          (cpu_sig->rev < 0x01000405 && mc->rev >  0x01000405)) )
+    {
+        printk_once(XENLOG_WARNING
+                    "microcode: Granite Rapids erratum GNR98 detected.  Skipping ucode 0x%08x\n"
+                    "microcode: Firmware update recommended\n", mc->rev);
+        return false;
+    }
+
+    return true;
+}
+
 static int cf_check intel_compare(
     const struct microcode_patch *old, const struct microcode_patch *new)
 {
@@ -365,6 +404,7 @@ static struct microcode_patch *cf_check intel_ucode_parse(
          * one with higher revision.
          */
         if ( microcode_fits_cpu(mc) &&
+             microcode_safe_to_load(mc) &&
              (!saved || compare_revisions(saved->rev, mc->rev) == NEW_UCODE) )
             saved = mc;
 
-- 
2.39.5


Re: [PATCH v2] x86/ucode: Work around Granite Rapids erraturm GNR98
Posted by Jan Beulich 2 weeks, 2 days ago
On 07.09.2026 23:32, Andrew Cooper wrote:
> @@ -273,6 +274,44 @@ static bool microcode_fits_cpu(const struct microcode_patch *mc)
>      return false;
>  }
>  
> +static bool microcode_safe_to_load(const struct microcode_patch *mc)
> +{
> +    struct cpu_signature *cpu_sig = &this_cpu(cpu_sig);
> +
> +    /*
> +     * Treat pre-production as always safe - anyone using pre-production
> +     * microcode knows what they are doing, and can keep any resulting pieces.
> +     */
> +    if ( (int)cpu_sig->rev < 0 || mc->rev < 0 )
> +        return true;
> +
> +    /*
> +     * GNR98 states that Granite Rapids systems hang when loading new ucode on
> +     * sufficiently old firmware.  GNR101 retroactively declares that one
> +     * ucode had incorrect min_rev fields, in light of discovering GNR98.

We still have no min_rev field, so imo a reference to it wants some
clarification. Really I first meant to ask why there's no use of that field,
to merely make that one exception.

> +     * Both are incomplete statements of the problem.
> +     *
> +     * At the time of writing (August 2026), the believed safe sequence is:
> +     *   0x01000370 -> [0x01000380...0x010003f3] -> 0x01000405 -> any later
> +     *
> +     * Disallow known-unsafe loads while permitting believed-safe loads.  For
> +     * GNR, this allows multi-hop loading to get up to the latest.
> +     */
> +    if ( boot_cpu_data.vfm == INTEL_GRANITERAPIDS_X &&
> +         boot_cpu_data.stepping == 1 && (cpu_sig->pf & 0x95) &&
> +         ((cpu_sig->rev < 0x01000380 && mc->rev >= 0x01000405) ||

This is odd: The lhs of && uses the lower bound of the inner permitted
range, while the rhs of the && doesn't use the upper one. If it's intended
that way, I think this also needs clarifying in the comment. Otherwise imo
lhs and rhs better would be consistent in this regard.

> +          (cpu_sig->rev < 0x01000405 && mc->rev >  0x01000405)) )

This one, otoh, fully matches the comment.

> +    {
> +        printk_once(XENLOG_WARNING
> +                    "microcode: Granite Rapids erratum GNR98 detected.  Skipping ucode 0x%08x\n"
> +                    "microcode: Firmware update recommended\n", mc->rev);

I think a 2nd XENLOG_WARNING is wanted after the inner \n (or none at all,
to use the default for both).

Jan
Re: [PATCH v2] x86/ucode: Work around Granite Rapids erraturm GNR98
Posted by Andrew Cooper 2 weeks, 2 days ago
On 08/09/2026 7:26 am, Jan Beulich wrote:
> On 07.09.2026 23:32, Andrew Cooper wrote:
>> @@ -273,6 +274,44 @@ static bool microcode_fits_cpu(const struct microcode_patch *mc)
>>      return false;
>>  }
>>  
>> +static bool microcode_safe_to_load(const struct microcode_patch *mc)
>> +{
>> +    struct cpu_signature *cpu_sig = &this_cpu(cpu_sig);
>> +
>> +    /*
>> +     * Treat pre-production as always safe - anyone using pre-production
>> +     * microcode knows what they are doing, and can keep any resulting pieces.
>> +     */
>> +    if ( (int)cpu_sig->rev < 0 || mc->rev < 0 )
>> +        return true;
>> +
>> +    /*
>> +     * GNR98 states that Granite Rapids systems hang when loading new ucode on
>> +     * sufficiently old firmware.  GNR101 retroactively declares that one
>> +     * ucode had incorrect min_rev fields, in light of discovering GNR98.
> We still have no min_rev field, so imo a reference to it wants some
> clarification.

No, I don't think so.  The fact Xen has no min_rev field (yet) has no
baring on the wording of GNR101.

This patch stops Xen hanging with the real ucode which has existed in
the world for 6 months.

I am still waiting on Intel to publish new blobs (including corrected
min_rev fields) before the multi-hop update can be made to work.  I have
no ETA on this.

>  Really I first meant to ask why there's no use of that field,
> to merely make that one exception.
>
>> +     * Both are incomplete statements of the problem.
>> +     *
>> +     * At the time of writing (August 2026), the believed safe sequence is:
>> +     *   0x01000370 -> [0x01000380...0x010003f3] -> 0x01000405 -> any later
>> +     *
>> +     * Disallow known-unsafe loads while permitting believed-safe loads.  For
>> +     * GNR, this allows multi-hop loading to get up to the latest.
>> +     */
>> +    if ( boot_cpu_data.vfm == INTEL_GRANITERAPIDS_X &&
>> +         boot_cpu_data.stepping == 1 && (cpu_sig->pf & 0x95) &&
>> +         ((cpu_sig->rev < 0x01000380 && mc->rev >= 0x01000405) ||
> This is odd: The lhs of && uses the lower bound of the inner permitted
> range, while the rhs of the && doesn't use the upper one. If it's intended
> that way, I think this also needs clarifying in the comment. Otherwise imo
> lhs and rhs better would be consistent in this regard.

It is intentional.  Furthermore, it is the only coherent way of
expressing the sequence as given.

I'm not writing a comment explaining why it's a good idea to use the
same boundary numerals between the comment and the code.  It goes
without saying.

I'm also not interested about pureness concerns about inner vs outer
bounds.  I can't see a change here that won't make it worse.

>
>> +          (cpu_sig->rev < 0x01000405 && mc->rev >  0x01000405)) )
> This one, otoh, fully matches the comment.
>
>> +    {
>> +        printk_once(XENLOG_WARNING
>> +                    "microcode: Granite Rapids erratum GNR98 detected.  Skipping ucode 0x%08x\n"
>> +                    "microcode: Firmware update recommended\n", mc->rev);
> I think a 2nd XENLOG_WARNING is wanted after the inner \n (or none at all,
> to use the default for both).

Fine, fixed up locally.

~Andrew

Re: [PATCH v2] x86/ucode: Work around Granite Rapids erraturm GNR98
Posted by Jan Beulich 2 weeks, 2 days ago
On 08.09.2026 12:08, Andrew Cooper wrote:
> On 08/09/2026 7:26 am, Jan Beulich wrote:
>> On 07.09.2026 23:32, Andrew Cooper wrote:
>>> @@ -273,6 +274,44 @@ static bool microcode_fits_cpu(const struct microcode_patch *mc)
>>>      return false;
>>>  }
>>>  
>>> +static bool microcode_safe_to_load(const struct microcode_patch *mc)
>>> +{
>>> +    struct cpu_signature *cpu_sig = &this_cpu(cpu_sig);
>>> +
>>> +    /*
>>> +     * Treat pre-production as always safe - anyone using pre-production
>>> +     * microcode knows what they are doing, and can keep any resulting pieces.
>>> +     */
>>> +    if ( (int)cpu_sig->rev < 0 || mc->rev < 0 )
>>> +        return true;
>>> +
>>> +    /*
>>> +     * GNR98 states that Granite Rapids systems hang when loading new ucode on
>>> +     * sufficiently old firmware.  GNR101 retroactively declares that one
>>> +     * ucode had incorrect min_rev fields, in light of discovering GNR98.
>> We still have no min_rev field, so imo a reference to it wants some
>> clarification.
> 
> No, I don't think so.  The fact Xen has no min_rev field (yet) has no
> baring on the wording of GNR101.

The wording there is "Minimum Runtime Microcode Update Revision". As long
as we don't have a field of the name, how can such a comment be unambiguous?

>>> +     * Both are incomplete statements of the problem.
>>> +     *
>>> +     * At the time of writing (August 2026), the believed safe sequence is:
>>> +     *   0x01000370 -> [0x01000380...0x010003f3] -> 0x01000405 -> any later
>>> +     *
>>> +     * Disallow known-unsafe loads while permitting believed-safe loads.  For
>>> +     * GNR, this allows multi-hop loading to get up to the latest.
>>> +     */
>>> +    if ( boot_cpu_data.vfm == INTEL_GRANITERAPIDS_X &&
>>> +         boot_cpu_data.stepping == 1 && (cpu_sig->pf & 0x95) &&
>>> +         ((cpu_sig->rev < 0x01000380 && mc->rev >= 0x01000405) ||
>> This is odd: The lhs of && uses the lower bound of the inner permitted
>> range, while the rhs of the && doesn't use the upper one. If it's intended
>> that way, I think this also needs clarifying in the comment. Otherwise imo
>> lhs and rhs better would be consistent in this regard.
> 
> It is intentional.  Furthermore, it is the only coherent way of
> expressing the sequence as given.
> 
> I'm not writing a comment explaining why it's a good idea to use the
> same boundary numerals between the comment and the code.  It goes
> without saying.

But that's the problem - code and comment are not (obviously) in sync.

> I'm also not interested about pureness concerns about inner vs outer
> bounds.  I can't see a change here that won't make it worse.

So why are

         ((cpu_sig->rev <= 0x01000370 && mc->rev >= 0x01000405) ||

and

         ((cpu_sig->rev < 0x01000380 && mc->rev > 0x010003f3) ||

both worse? Both 0x01000370 ... 0x01000380 and 0x010003f3 ... 0x01000405
are discontiguous, which is bad enough. If then you pick apparently
randomly (and in any event inconsistently) from the possible boundaries,
how can that help the situation (and be "the only coherent way of
expressing the sequence as given")?

Even if it's merely a matter of us disagreeing on what "consistent" or
"coherent" would be here, doesn't me as the first reader not easily
spotting the "coherency" you claim already indicate there is a
(possible) issue? (For context: Whichever way the expressions are going
to end up, they'll make implicit statements on the gaps, i.e. on the
non-public 0x01000371 ... 0x0100037f and 0x010003f4 ... 0x01000404. You
may say that doesn't matter, because of their non-publicness, but that
won't make that implicit statement go away.)

Jan

Re: [PATCH v2] x86/ucode: Work around Granite Rapids erraturm GNR98
Posted by Andrew Cooper 2 weeks, 2 days ago
On 08/09/2026 11:55 am, Jan Beulich wrote:
> On 08.09.2026 12:08, Andrew Cooper wrote:
>> On 08/09/2026 7:26 am, Jan Beulich wrote:
>>> On 07.09.2026 23:32, Andrew Cooper wrote:
>>>> @@ -273,6 +274,44 @@ static bool microcode_fits_cpu(const struct microcode_patch *mc)
>>>>      return false;
>>>>  }
>>>>  
>>>> +static bool microcode_safe_to_load(const struct microcode_patch *mc)
>>>> +{
>>>> +    struct cpu_signature *cpu_sig = &this_cpu(cpu_sig);
>>>> +
>>>> +    /*
>>>> +     * Treat pre-production as always safe - anyone using pre-production
>>>> +     * microcode knows what they are doing, and can keep any resulting pieces.
>>>> +     */
>>>> +    if ( (int)cpu_sig->rev < 0 || mc->rev < 0 )
>>>> +        return true;
>>>> +
>>>> +    /*
>>>> +     * GNR98 states that Granite Rapids systems hang when loading new ucode on
>>>> +     * sufficiently old firmware.  GNR101 retroactively declares that one
>>>> +     * ucode had incorrect min_rev fields, in light of discovering GNR98.
>>> We still have no min_rev field, so imo a reference to it wants some
>>> clarification.
>> No, I don't think so.  The fact Xen has no min_rev field (yet) has no
>> baring on the wording of GNR101.
> The wording there is "Minimum Runtime Microcode Update Revision". As long
> as we don't have a field of the name, how can such a comment be unambiguous?

Fine, I'll say "minimum revision field", but the whole name is (and
always has been) silly.

>
>>>> +     * Both are incomplete statements of the problem.
>>>> +     *
>>>> +     * At the time of writing (August 2026), the believed safe sequence is:
>>>> +     *   0x01000370 -> [0x01000380...0x010003f3] -> 0x01000405 -> any later
>>>> +     *
>>>> +     * Disallow known-unsafe loads while permitting believed-safe loads.  For
>>>> +     * GNR, this allows multi-hop loading to get up to the latest.
>>>> +     */
>>>> +    if ( boot_cpu_data.vfm == INTEL_GRANITERAPIDS_X &&
>>>> +         boot_cpu_data.stepping == 1 && (cpu_sig->pf & 0x95) &&
>>>> +         ((cpu_sig->rev < 0x01000380 && mc->rev >= 0x01000405) ||
>>> This is odd: The lhs of && uses the lower bound of the inner permitted
>>> range, while the rhs of the && doesn't use the upper one. If it's intended
>>> that way, I think this also needs clarifying in the comment. Otherwise imo
>>> lhs and rhs better would be consistent in this regard.
>> It is intentional.  Furthermore, it is the only coherent way of
>> expressing the sequence as given.
>>
>> I'm not writing a comment explaining why it's a good idea to use the
>> same boundary numerals between the comment and the code.  It goes
>> without saying.
> But that's the problem - code and comment are not (obviously) in sync.
>
>> I'm also not interested about pureness concerns about inner vs outer
>> bounds.  I can't see a change here that won't make it worse.
> So why are
>
>          ((cpu_sig->rev <= 0x01000370 && mc->rev >= 0x01000405) ||
>
> and
>
>          ((cpu_sig->rev < 0x01000380 && mc->rev > 0x010003f3) ||
>
> both worse?

Because they are both buggy.  They fail to exclude some unsafe cases.

If it's not obvious, I do know more than I can say publicly.

~Andrew

Re: [PATCH v2] x86/ucode: Work around Granite Rapids erraturm GNR98
Posted by Jan Beulich 2 weeks, 1 day ago
On 08.09.2026 19:10, Andrew Cooper wrote:
> On 08/09/2026 11:55 am, Jan Beulich wrote:
>> On 08.09.2026 12:08, Andrew Cooper wrote:
>>> On 08/09/2026 7:26 am, Jan Beulich wrote:
>>>> On 07.09.2026 23:32, Andrew Cooper wrote:
>>>>> @@ -273,6 +274,44 @@ static bool microcode_fits_cpu(const struct microcode_patch *mc)
>>>>>      return false;
>>>>>  }
>>>>>  
>>>>> +static bool microcode_safe_to_load(const struct microcode_patch *mc)
>>>>> +{
>>>>> +    struct cpu_signature *cpu_sig = &this_cpu(cpu_sig);
>>>>> +
>>>>> +    /*
>>>>> +     * Treat pre-production as always safe - anyone using pre-production
>>>>> +     * microcode knows what they are doing, and can keep any resulting pieces.
>>>>> +     */
>>>>> +    if ( (int)cpu_sig->rev < 0 || mc->rev < 0 )
>>>>> +        return true;
>>>>> +
>>>>> +    /*
>>>>> +     * GNR98 states that Granite Rapids systems hang when loading new ucode on
>>>>> +     * sufficiently old firmware.  GNR101 retroactively declares that one
>>>>> +     * ucode had incorrect min_rev fields, in light of discovering GNR98.
>>>> We still have no min_rev field, so imo a reference to it wants some
>>>> clarification.
>>> No, I don't think so.  The fact Xen has no min_rev field (yet) has no
>>> baring on the wording of GNR101.
>> The wording there is "Minimum Runtime Microcode Update Revision". As long
>> as we don't have a field of the name, how can such a comment be unambiguous?
> 
> Fine, I'll say "minimum revision field", but the whole name is (and
> always has been) silly.
> 
>>
>>>>> +     * Both are incomplete statements of the problem.
>>>>> +     *
>>>>> +     * At the time of writing (August 2026), the believed safe sequence is:
>>>>> +     *   0x01000370 -> [0x01000380...0x010003f3] -> 0x01000405 -> any later
>>>>> +     *
>>>>> +     * Disallow known-unsafe loads while permitting believed-safe loads.  For
>>>>> +     * GNR, this allows multi-hop loading to get up to the latest.
>>>>> +     */
>>>>> +    if ( boot_cpu_data.vfm == INTEL_GRANITERAPIDS_X &&
>>>>> +         boot_cpu_data.stepping == 1 && (cpu_sig->pf & 0x95) &&
>>>>> +         ((cpu_sig->rev < 0x01000380 && mc->rev >= 0x01000405) ||
>>>> This is odd: The lhs of && uses the lower bound of the inner permitted
>>>> range, while the rhs of the && doesn't use the upper one. If it's intended
>>>> that way, I think this also needs clarifying in the comment. Otherwise imo
>>>> lhs and rhs better would be consistent in this regard.
>>> It is intentional.  Furthermore, it is the only coherent way of
>>> expressing the sequence as given.
>>>
>>> I'm not writing a comment explaining why it's a good idea to use the
>>> same boundary numerals between the comment and the code.  It goes
>>> without saying.
>> But that's the problem - code and comment are not (obviously) in sync.
>>
>>> I'm also not interested about pureness concerns about inner vs outer
>>> bounds.  I can't see a change here that won't make it worse.
>> So why are
>>
>>          ((cpu_sig->rev <= 0x01000370 && mc->rev >= 0x01000405) ||
>>
>> and
>>
>>          ((cpu_sig->rev < 0x01000380 && mc->rev > 0x010003f3) ||
>>
>> both worse?
> 
> Because they are both buggy.  They fail to exclude some unsafe cases.
> 
> If it's not obvious, I do know more than I can say publicly.

Which I had in mind as a possible option. Problem being that with
incomplete information it's of questionable value to offer an ack
on such changes. Here you go:
Acked-by: Jan Beulich <jbeulich@suse.com>

Jan