[RFC PATCH 0/3] cxl: Auto-create a region for Type-2 memdev attach

Richard Cheng posted 3 patches 1 month, 3 weeks ago
drivers/cxl/core/region.c      | 422 +++++++++++++++++++++++++++++----
tools/testing/cxl/test/accel.c |   7 +
tools/testing/cxl/test/cxl.c   |  61 ++++-
3 files changed, 439 insertions(+), 51 deletions(-)
[RFC PATCH 0/3] cxl: Auto-create a region for Type-2 memdev attach
Posted by Richard Cheng 1 month, 3 weeks ago
A Type-2 accelerator driver calls devm_cxl_probe_mem() to register its
memdev and get an HPA range, but today only if FW already committed a
region. Real accelerators can have usable memory with no committed decoder,
so get nothing.

If no FW region is mapped, the core picks the device's unused manual
DEVMEM decoder and a compatible x1 Type-2 RAM root, allocates the full
volatile DPA and HPA, commits the decoder, and returns the range. Unbind
resets and removes it. Strict first cut with single decoder, IW=1, minimum
granularity, first-compatible root.

The design intent is that the provider F_LOCKs its region against userspace
but must reset its own software region on detach. A plain flag would also
allow reset on generic kill/delete paths, so we thread a reset context
through teardown and commit rollback. devm_cxl_probe_mem() may now commit
decoders.

Testing result is in the following.
- Built clean with clang/LLVM on arm64
- cxl_test, type2_test=1. accel0 takes the unchanged attach path. accel1
  drives auto_create -> a committed 512 MB RAM region. The test asserts
  the 512 MB HPA range. committed state and 256 byte granularity confirmed
  via sysfs.
- Unbind tears the region down with no orphaned decoder, rebind re-creates
  a fresh committed region.
- Mock test only. Real accelerators whose FW commits a decoder take the
  attach path, and vfio-cxl binds only FW-committed devices, so auto-create
  has no real-HW caller yet.

Best regards,
Richard Cheng.

Richard Cheng (3):
  cxl/region: Reset software-created regions on memdev detach
  cxl/region: Auto-create a region for memdev attach
  cxl/test: Exercise Type-2 automatic region creation

 drivers/cxl/core/region.c      | 422 +++++++++++++++++++++++++++++----
 tools/testing/cxl/test/accel.c |   7 +
 tools/testing/cxl/test/cxl.c   |  61 ++++-
 3 files changed, 439 insertions(+), 51 deletions(-)


base-commit: 1c6b4ceafc3b994871c29340e0c1ddb0af5800e7
-- 
2.43.0
Re: [RFC PATCH 0/3] cxl: Auto-create a region for Type-2 memdev attach
Posted by Alejandro Lucero Palau 1 month, 2 weeks ago
Hi Richard,


Some comments below.


Thanks!


On 8/5/26 08:40, Richard Cheng wrote:
> A Type-2 accelerator driver calls devm_cxl_probe_mem() to register its
> memdev and get an HPA range, but today only if FW already committed a
> region. Real accelerators can have usable memory with no committed decoder,
> so get nothing.


The previous paragraph describes the current situation and the next one 
is about what the patchset tries to address. Maybe to explicitly make 
the difference would help people not so used to the subject.


> If no FW region is mapped, the core picks the device's unused manual
> DEVMEM decoder and a compatible x1 Type-2 RAM root


I had to look for this x1 reference ... and I would say it creates 
confusion. At least it does to me. Not sure if you meant interleaving, 
because I do not think you are referring to link lanes here ...


> , allocates the full
> volatile DPA and HPA, commits the decoder, and returns the range.


This is something requiring discussion or clarification. I think it 
would make sense the provider/driver specifying a DPA size instead of 
using the default full size. I think there is a good reason for this 
non-default size use: why would the kernel create a region from a CXL 
Type2 device using the full DPA size when the FW/BIOS did not do so?


This leads us to wondering why the FW/BIOS would not do so, the use 
case. Current Intel/AMD BIOS (I think you have the aim at ARM servers) 
are not allowing this case ... for a Type2 device having all the bits in 
place. If something requires to be specifically configured, would not 
the driver do so before using the CXL mem? If this logic makes sense, 
the auto-creation should not be the way to go.


The first 15 Type2 basic support patchset versions supported the case of 
a driver specifying the size for the cxl region to be created. And it 
was through a specific API call after the memdev was created. Last Type2 
patchset and the functionality finally merged only supported the case of 
auto-create regions from committed decoders, and using this final 
agreement for region attachment by the Type2 memdev/driver. I think it 
makes sense in that supported case to have the auto-create region but I 
can not see the reason for the case you are addressing now.


>   Unbind
> resets and removes it. Strict first cut with single decoder, IW=1, minimum
> granularity, first-compatible root.


I'm lost here.


>
> The design intent is that the provider F_LOCKs its region against userspace
> but must reset its own software region on detach. A plain flag would also
> allow reset on generic kill/delete paths, so we thread a reset context
> through teardown and commit rollback. devm_cxl_probe_mem() may now commit
> decoders.


If there is a real use case for this auto-create region from 
non-committed decoders, I think your patchset makes sense. But I'm 
afraid we need to discuss this further.


Thank you,

Alejandro.


> Testing result is in the following.
> - Built clean with clang/LLVM on arm64
> - cxl_test, type2_test=1. accel0 takes the unchanged attach path. accel1
>    drives auto_create -> a committed 512 MB RAM region. The test asserts
>    the 512 MB HPA range. committed state and 256 byte granularity confirmed
>    via sysfs.
> - Unbind tears the region down with no orphaned decoder, rebind re-creates
>    a fresh committed region.
> - Mock test only. Real accelerators whose FW commits a decoder take the
>    attach path, and vfio-cxl binds only FW-committed devices, so auto-create
>    has no real-HW caller yet.
>
> Best regards,
> Richard Cheng.
>
> Richard Cheng (3):
>    cxl/region: Reset software-created regions on memdev detach
>    cxl/region: Auto-create a region for memdev attach
>    cxl/test: Exercise Type-2 automatic region creation
>
>   drivers/cxl/core/region.c      | 422 +++++++++++++++++++++++++++++----
>   tools/testing/cxl/test/accel.c |   7 +
>   tools/testing/cxl/test/cxl.c   |  61 ++++-
>   3 files changed, 439 insertions(+), 51 deletions(-)
>
>
> base-commit: 1c6b4ceafc3b994871c29340e0c1ddb0af5800e7
Re: [RFC PATCH 0/3] cxl: Auto-create a region for Type-2 memdev attach
Posted by Richard Cheng 1 month, 1 week ago
On Wed, Aug 12, 2026 at 10:58:02AM +0800, Alejandro Lucero Palau wrote:
> Hi Richard,
> 
> 
> Some comments below.
> 
> 
> Thanks!
> 
> 
> On 8/5/26 08:40, Richard Cheng wrote:
> > A Type-2 accelerator driver calls devm_cxl_probe_mem() to register its
> > memdev and get an HPA range, but today only if FW already committed a
> > region. Real accelerators can have usable memory with no committed decoder,
> > so get nothing.
> 
> 
> The previous paragraph describes the current situation and the next one is
> about what the patchset tries to address. Maybe to explicitly make the
> difference would help people not so used to the subject.
> 
> 
> > If no FW region is mapped, the core picks the device's unused manual
> > DEVMEM decoder and a compatible x1 Type-2 RAM root
> 
> 
> I had to look for this x1 reference ... and I would say it creates
> confusion. At least it does to me. Not sure if you meant interleaving,
> because I do not think you are referring to link lanes here ...
> 
> 
> > , allocates the full
> > volatile DPA and HPA, commits the decoder, and returns the range.
> 
> 
> This is something requiring discussion or clarification. I think it would
> make sense the provider/driver specifying a DPA size instead of using the
> default full size. I think there is a good reason for this non-default size
> use: why would the kernel create a region from a CXL Type2 device using the
> full DPA size when the FW/BIOS did not do so?
> 
> 
> This leads us to wondering why the FW/BIOS would not do so, the use case.
> Current Intel/AMD BIOS (I think you have the aim at ARM servers) are not
> allowing this case ... for a Type2 device having all the bits in place. If
> something requires to be specifically configured, would not the driver do so
> before using the CXL mem? If this logic makes sense, the auto-creation
> should not be the way to go.
> 
> 
> The first 15 Type2 basic support patchset versions supported the case of a
> driver specifying the size for the cxl region to be created. And it was
> through a specific API call after the memdev was created. Last Type2
> patchset and the functionality finally merged only supported the case of
> auto-create regions from committed decoders, and using this final agreement
> for region attachment by the Type2 memdev/driver. I think it makes sense in
> that supported case to have the auto-create region but I can not see the
> reason for the case you are addressing now.
> 
> 
> >   Unbind
> > resets and removes it. Strict first cut with single decoder, IW=1, minimum
> > granularity, first-compatible root.
> 
> 
> I'm lost here.
> 
> 
> > 
> > The design intent is that the provider F_LOCKs its region against userspace
> > but must reset its own software region on detach. A plain flag would also
> > allow reset on generic kill/delete paths, so we thread a reset context
> > through teardown and commit rollback. devm_cxl_probe_mem() may now commit
> > decoders.
> 
> 
> If there is a real use case for this auto-create region from non-committed
> decoders, I think your patchset makes sense. But I'm afraid we need to
> discuss this further.
> 
> 
> Thank you,
> 
> Alejandro.
> 
>

Hi Alejandro,

Thanks for the review and explanation. I've read them all.

I think you are right that this RFC doesn't currently have a production platform
where the system FW publishes a Type-2 CFMWS but leaves the EP decoder
uncommitted.

However, the config appears to be permitted by the CXL model. A CFMWS describes
a FW-established root HPA window and the restrictions governing its use,
including Type-2 v.s. Type-3 and volatile v.s. PMEM. The CFMWS def also
describes OSPM assigning HPA ranges from those windows to discovered CXL.mem
devices [1].

The Linux CXL doc similarly states that only root decoders are required to be
programmed during probe. Switch and EP decoder may remain available for runtime
programming when the platform supports it [2].

I raise the RFC intended for the question of how Linux should support that
architecturally permitted config.

The cxl_test config added in patch 3/3 constructs this scenario synthetically.
This demonstrates the proposed kernel behavior, but I agree I don't know
whether there exists a deployed FW scenario.


And I agree that devm_cxl_probe_mem() shouldn't silently change from
"attach to a FW-established region" into "allocate resources and program a new
region". Those operations should have different semantics and ownership
expectations.

I am planning to rebase onto cxl/nexxt and rework the proposal as the following,
please take a look and see if that matches your imagination or not.
* Keep devm_cxl_probe_mem() behavior unchanged for FW-committed regions
* Make region creation an explicit request from the accelerator provider,
  rather than an automatic fallback during memdev attach
* Have the provider specify the required size. CXL core shouldn't assume
  that it maybe consume the entire volatile DPA partition as you mentioned.
* Separate the reusable region-provisioning mechanism from the initial Type-2
  policy.
* The common mechanism should handle HPA/DPA allocation, decoder-path
  construction, commit , rollback and managed teardown.
* The initial type-2 caller would constrain that to volatile DEVMEM, IW=1
  and a provider-requested size.

Oh and I'll replace "x1" with "IW=1" and explain the initial decoder,
root-selection and granularity restriction more clearly.

How does that sound to you ?

[1]: https://computeexpresslink.org/wp-content/uploads/2024/02/CEDT_ECN_1.0A_Eval.pdf
[2]: https://docs.kernel.org/driver-api/cxl/linux/cxl-driver.html#runtime-programming

Best regards,
Richard Cheng.
 
> > Testing result is in the following.
> > - Built clean with clang/LLVM on arm64
> > - cxl_test, type2_test=1. accel0 takes the unchanged attach path. accel1
> >    drives auto_create -> a committed 512 MB RAM region. The test asserts
> >    the 512 MB HPA range. committed state and 256 byte granularity confirmed
> >    via sysfs.
> > - Unbind tears the region down with no orphaned decoder, rebind re-creates
> >    a fresh committed region.
> > - Mock test only. Real accelerators whose FW commits a decoder take the
> >    attach path, and vfio-cxl binds only FW-committed devices, so auto-create
> >    has no real-HW caller yet.
> > 
> > Best regards,
> > Richard Cheng.
> > 
> > Richard Cheng (3):
> >    cxl/region: Reset software-created regions on memdev detach
> >    cxl/region: Auto-create a region for memdev attach
> >    cxl/test: Exercise Type-2 automatic region creation
> > 
> >   drivers/cxl/core/region.c      | 422 +++++++++++++++++++++++++++++----
> >   tools/testing/cxl/test/accel.c |   7 +
> >   tools/testing/cxl/test/cxl.c   |  61 ++++-
> >   3 files changed, 439 insertions(+), 51 deletions(-)
> > 
> > 
> > base-commit: 1c6b4ceafc3b994871c29340e0c1ddb0af5800e7
Re: [RFC PATCH 0/3] cxl: Auto-create a region for Type-2 memdev attach
Posted by Lucero Palau, Alejandro 1 month, 1 week ago
Hi Richard,

On 20/08/2026 10:41, Richard Cheng wrote:
> On Wed, Aug 12, 2026 at 10:58:02AM +0800, Alejandro Lucero Palau wrote:
>> Hi Richard,
>>
>>
>> Some comments below.
>>
>>
>> Thanks!
>>
>>
>> On 8/5/26 08:40, Richard Cheng wrote:


<snip>

> Hi Alejandro,
>
> Thanks for the review and explanation. I've read them all.
>
> I think you are right that this RFC doesn't currently have a production platform
> where the system FW publishes a Type-2 CFMWS but leaves the EP decoder
> uncommitted.
>
> However, the config appears to be permitted by the CXL model. A CFMWS describes
> a FW-established root HPA window and the restrictions governing its use,
> including Type-2 v.s. Type-3 and volatile v.s. PMEM. The CFMWS def also
> describes OSPM assigning HPA ranges from those windows to discovered CXL.mem
> devices [1].


Right. I'm not saying this should not be supported, just pointing out 
the use case does not make sense with current BIOS functionality. I 
think BIOS will/could support a config option for just leaving a Type2 
HDM uncommitted, but then why the kernel should do the same a default 
BIOS config would do?


>
> The Linux CXL doc similarly states that only root decoders are required to be
> programmed during probe. Switch and EP decoder may remain available for runtime
> programming when the platform supports it [2].


Tangential to this discussion, but I have problems with this assertion. 
Any switch or EP HDM programming will need a root port HDM programming 
as well. Not sure which root decoders will need to be programmed at boot 
time: a CFMWS is "programmed" by the BIOS and root decoders will need to 
be programmed as well for any Type2/switch found with an enabled link.


>
> I raise the RFC intended for the question of how Linux should support that
> architecturally permitted config.
>
> The cxl_test config added in patch 3/3 constructs this scenario synthetically.
> This demonstrates the proposed kernel behavior, but I agree I don't know
> whether there exists a deployed FW scenario.
>
>
> And I agree that devm_cxl_probe_mem() shouldn't silently change from
> "attach to a FW-established region" into "allocate resources and program a new
> region". Those operations should have different semantics and ownership
> expectations.


Glad with the consensus :-)


>
> I am planning to rebase onto cxl/nexxt and rework the proposal as the following,
> please take a look and see if that matches your imagination or not.
> * Keep devm_cxl_probe_mem() behavior unchanged for FW-committed regions
> * Make region creation an explicit request from the accelerator provider,
>    rather than an automatic fallback during memdev attach
> * Have the provider specify the required size. CXL core shouldn't assume
>    that it maybe consume the entire volatile DPA partition as you mentioned.
> * Separate the reusable region-provisioning mechanism from the initial Type-2
>    policy.
> * The common mechanism should handle HPA/DPA allocation, decoder-path
>    construction, commit , rollback and managed teardown.
> * The initial type-2 caller would constrain that to volatile DEVMEM, IW=1
>    and a provider-requested size.
>
> Oh and I'll replace "x1" with "IW=1" and explain the initial decoder,
> root-selection and granularity restriction more clearly.
>
> How does that sound to you ?


It sounds perfect!


FWIW, you likely saw Gregory's comment (discord) on this work requiring 
the support for PMEM or at least the awareness PMEM support will need to 
use same interface. His opinion and mine came from Dan's vision on this, 
and your work will be the base for such PMEM support. I do not have an 
impending reason for working on this PMEM support, but I am really 
interested in how Type3 PMEMs can leverage CXL.mem for improving storage 
needs, and currently reading/thinking about all this ...


Thanks!


> [1]: https://computeexpresslink.org/wp-content/uploads/2024/02/CEDT_ECN_1.0A_Eval.pdf
> [2]: https://docs.kernel.org/driver-api/cxl/linux/cxl-driver.html#runtime-programming
>
> Best regards,
> Richard Cheng.
>   
>>> Testing result is in the following.
>>> - Built clean with clang/LLVM on arm64
>>> - cxl_test, type2_test=1. accel0 takes the unchanged attach path. accel1
>>>     drives auto_create -> a committed 512 MB RAM region. The test asserts
>>>     the 512 MB HPA range. committed state and 256 byte granularity confirmed
>>>     via sysfs.
>>> - Unbind tears the region down with no orphaned decoder, rebind re-creates
>>>     a fresh committed region.
>>> - Mock test only. Real accelerators whose FW commits a decoder take the
>>>     attach path, and vfio-cxl binds only FW-committed devices, so auto-create
>>>     has no real-HW caller yet.
>>>
>>> Best regards,
>>> Richard Cheng.
>>>
>>> Richard Cheng (3):
>>>     cxl/region: Reset software-created regions on memdev detach
>>>     cxl/region: Auto-create a region for memdev attach
>>>     cxl/test: Exercise Type-2 automatic region creation
>>>
>>>    drivers/cxl/core/region.c      | 422 +++++++++++++++++++++++++++++----
>>>    tools/testing/cxl/test/accel.c |   7 +
>>>    tools/testing/cxl/test/cxl.c   |  61 ++++-
>>>    3 files changed, 439 insertions(+), 51 deletions(-)
>>>
>>>
>>> base-commit: 1c6b4ceafc3b994871c29340e0c1ddb0af5800e7
Re: [RFC PATCH 0/3] cxl: Auto-create a region for Type-2 memdev attach
Posted by Richard Cheng 1 month ago
On Tue, Aug 25, 2026 at 10:50:06AM +0800, Lucero Palau, Alejandro wrote:
> Hi Richard,
> 
> On 20/08/2026 10:41, Richard Cheng wrote:
> > On Wed, Aug 12, 2026 at 10:58:02AM +0800, Alejandro Lucero Palau wrote:
> > > Hi Richard,
> > > 
> > > 
> > > Some comments below.
> > > 
> > > 
> > > Thanks!
> > > 
> > > 
> > > On 8/5/26 08:40, Richard Cheng wrote:
> 
> 
> <snip>
> 
> > Hi Alejandro,
> > 
> > Thanks for the review and explanation. I've read them all.
> > 
> > I think you are right that this RFC doesn't currently have a production platform
> > where the system FW publishes a Type-2 CFMWS but leaves the EP decoder
> > uncommitted.
> > 
> > However, the config appears to be permitted by the CXL model. A CFMWS describes
> > a FW-established root HPA window and the restrictions governing its use,
> > including Type-2 v.s. Type-3 and volatile v.s. PMEM. The CFMWS def also
> > describes OSPM assigning HPA ranges from those windows to discovered CXL.mem
> > devices [1].
> 
> 
> Right. I'm not saying this should not be supported, just pointing out the
> use case does not make sense with current BIOS functionality. I think BIOS
> will/could support a config option for just leaving a Type2 HDM uncommitted,
> but then why the kernel should do the same a default BIOS config would do?
>

Agreed. The kernel shouldn't recreate the config that BIOS would normally provide.
And that's why I think we should move region createion out of devm_cxl_probe_mem().
That helper should discover and attach to an already committed region.

If FW leaves the decoders unconfigured intentionally , a driver may explicitly request a region
and provide the size it needs.

In my mind the new model should be
- FW-committed config is only discovered and attached
- an uncommitted config remains untouched unless a driver explicitly requests it
- CXL core supplieds the allocation, validation, programming, accounting and teardown mechnism
- the requesting driver owns the policy and the use of the region

So far I think PMEM reconstruction form label data maybe be a potential use case, though it's not
provided in kernel right now, we can work on it in the future, or work on that one first and we'll
continue the auto-create region part.

What do you think ?

 
> 
> > 
> > The Linux CXL doc similarly states that only root decoders are required to be
> > programmed during probe. Switch and EP decoder may remain available for runtime
> > programming when the platform supports it [2].
> 
> 
> Tangential to this discussion, but I have problems with this assertion. Any
> switch or EP HDM programming will need a root port HDM programming as well.
> Not sure which root decoders will need to be programmed at boot time: a
> CFMWS is "programmed" by the BIOS and root decoders will need to be
> programmed as well for any Type2/switch found with an enabled link.
> 
> 

Here root decoder I mean the logical object created from a CFMWS. I didn't mean switch or EP
decoder can be programmed independently.

My point was that FW establishes the CFMWS window, while the platform may leave downstream decoder
path for later programming.

> > 
> > I raise the RFC intended for the question of how Linux should support that
> > architecturally permitted config.
> > 
> > The cxl_test config added in patch 3/3 constructs this scenario synthetically.
> > This demonstrates the proposed kernel behavior, but I agree I don't know
> > whether there exists a deployed FW scenario.
> > 
> > 
> > And I agree that devm_cxl_probe_mem() shouldn't silently change from
> > "attach to a FW-established region" into "allocate resources and program a new
> > region". Those operations should have different semantics and ownership
> > expectations.
> 
> 
> Glad with the consensus :-)
> 
> 
> > 
> > I am planning to rebase onto cxl/nexxt and rework the proposal as the following,
> > please take a look and see if that matches your imagination or not.
> > * Keep devm_cxl_probe_mem() behavior unchanged for FW-committed regions
> > * Make region creation an explicit request from the accelerator provider,
> >    rather than an automatic fallback during memdev attach
> > * Have the provider specify the required size. CXL core shouldn't assume
> >    that it maybe consume the entire volatile DPA partition as you mentioned.
> > * Separate the reusable region-provisioning mechanism from the initial Type-2
> >    policy.
> > * The common mechanism should handle HPA/DPA allocation, decoder-path
> >    construction, commit , rollback and managed teardown.
> > * The initial type-2 caller would constrain that to volatile DEVMEM, IW=1
> >    and a provider-requested size.
> > 
> > Oh and I'll replace "x1" with "IW=1" and explain the initial decoder,
> > root-selection and granularity restriction more clearly.
> > 
> > How does that sound to you ?
> 
> 
> It sounds perfect!
> 
> 
> FWIW, you likely saw Gregory's comment (discord) on this work requiring the
> support for PMEM or at least the awareness PMEM support will need to use
> same interface. His opinion and mine came from Dan's vision on this, and
> your work will be the base for such PMEM support. I do not have an impending
> reason for working on this PMEM support, but I am really interested in how
> Type3 PMEMs can leverage CXL.mem for improving storage needs, and currently
> reading/thinking about all this ...
> 
> 
> Thanks!
> 

Hmmm for this part I have no idea for now, I'll study more and discuss with you guys.

Best regards,
Richard Cheng.

> 
> > [1]: https://computeexpresslink.org/wp-content/uploads/2024/02/CEDT_ECN_1.0A_Eval.pdf
> > [2]: https://docs.kernel.org/driver-api/cxl/linux/cxl-driver.html#runtime-programming
> > 
> > Best regards,
> > Richard Cheng.
> > > > Testing result is in the following.
> > > > - Built clean with clang/LLVM on arm64
> > > > - cxl_test, type2_test=1. accel0 takes the unchanged attach path. accel1
> > > >     drives auto_create -> a committed 512 MB RAM region. The test asserts
> > > >     the 512 MB HPA range. committed state and 256 byte granularity confirmed
> > > >     via sysfs.
> > > > - Unbind tears the region down with no orphaned decoder, rebind re-creates
> > > >     a fresh committed region.
> > > > - Mock test only. Real accelerators whose FW commits a decoder take the
> > > >     attach path, and vfio-cxl binds only FW-committed devices, so auto-create
> > > >     has no real-HW caller yet.
> > > > 
> > > > Best regards,
> > > > Richard Cheng.
> > > > 
> > > > Richard Cheng (3):
> > > >     cxl/region: Reset software-created regions on memdev detach
> > > >     cxl/region: Auto-create a region for memdev attach
> > > >     cxl/test: Exercise Type-2 automatic region creation
> > > > 
> > > >    drivers/cxl/core/region.c      | 422 +++++++++++++++++++++++++++++----
> > > >    tools/testing/cxl/test/accel.c |   7 +
> > > >    tools/testing/cxl/test/cxl.c   |  61 ++++-
> > > >    3 files changed, 439 insertions(+), 51 deletions(-)
> > > > 
> > > > 
> > > > base-commit: 1c6b4ceafc3b994871c29340e0c1ddb0af5800e7
Re: [RFC PATCH 0/3] cxl: Auto-create a region for Type-2 memdev attach
Posted by Jonathan Cameron 1 week, 3 days ago
On Mon, 31 Aug 2026 16:52:29 +0800
Richard Cheng <icheng@nvidia.com> wrote:

> On Tue, Aug 25, 2026 at 10:50:06AM +0800, Lucero Palau, Alejandro wrote:
> > Hi Richard,
> > 
> > On 20/08/2026 10:41, Richard Cheng wrote:  
> > > On Wed, Aug 12, 2026 at 10:58:02AM +0800, Alejandro Lucero Palau wrote:  
> > > > Hi Richard,
> > > > 
> > > > 
> > > > Some comments below.
> > > > 
> > > > 
> > > > Thanks!
> > > > 
> > > > 
> > > > On 8/5/26 08:40, Richard Cheng wrote:  
> > 
> > 
> > <snip>
> >   
> > > Hi Alejandro,
> > > 
> > > Thanks for the review and explanation. I've read them all.
> > > 
> > > I think you are right that this RFC doesn't currently have a production platform
> > > where the system FW publishes a Type-2 CFMWS but leaves the EP decoder
> > > uncommitted.
> > > 
> > > However, the config appears to be permitted by the CXL model. A CFMWS describes
> > > a FW-established root HPA window and the restrictions governing its use,
> > > including Type-2 v.s. Type-3 and volatile v.s. PMEM. The CFMWS def also
> > > describes OSPM assigning HPA ranges from those windows to discovered CXL.mem
> > > devices [1].  
> > 
> > 
> > Right. I'm not saying this should not be supported, just pointing out the
> > use case does not make sense with current BIOS functionality. I think BIOS
> > will/could support a config option for just leaving a Type2 HDM uncommitted,
> > but then why the kernel should do the same a default BIOS config would do?
> >  
> 
> Agreed. The kernel shouldn't recreate the config that BIOS would normally provide.

I'm a bit lost. Both BIOS doing nothing beyond cfmws as a design decision and
hotplug (where bios isn't in the loop) require this sort of flow.

Sure both might not be what you happen to have today but they are both
very much real usecases!

Jonathan

> And that's why I think we should move region createion out of devm_cxl_probe_mem().
> That helper should discover and attach to an already committed region.
> 
> If FW leaves the decoders unconfigured intentionally , a driver may explicitly request a region
> and provide the size it needs.
> 
> In my mind the new model should be
> - FW-committed config is only discovered and attached
> - an uncommitted config remains untouched unless a driver explicitly requests it
> - CXL core supplieds the allocation, validation, programming, accounting and teardown mechnism
> - the requesting driver owns the policy and the use of the region
> 
> So far I think PMEM reconstruction form label data maybe be a potential use case, though it's not
> provided in kernel right now, we can work on it in the future, or work on that one first and we'll
> continue the auto-create region part.
> 
> What do you think ?
> 
>  
> >   
> > > 
> > > The Linux CXL doc similarly states that only root decoders are required to be
> > > programmed during probe. Switch and EP decoder may remain available for runtime
> > > programming when the platform supports it [2].  
> > 
> > 
> > Tangential to this discussion, but I have problems with this assertion. Any
> > switch or EP HDM programming will need a root port HDM programming as well.
> > Not sure which root decoders will need to be programmed at boot time: a
> > CFMWS is "programmed" by the BIOS and root decoders will need to be
> > programmed as well for any Type2/switch found with an enabled link.
> > 
> >   
> 
> Here root decoder I mean the logical object created from a CFMWS. I didn't mean switch or EP
> decoder can be programmed independently.
> 
> My point was that FW establishes the CFMWS window, while the platform may leave downstream decoder
> path for later programming.
> 
> > > 
> > > I raise the RFC intended for the question of how Linux should support that
> > > architecturally permitted config.
> > > 
> > > The cxl_test config added in patch 3/3 constructs this scenario synthetically.
> > > This demonstrates the proposed kernel behavior, but I agree I don't know
> > > whether there exists a deployed FW scenario.
> > > 
> > > 
> > > And I agree that devm_cxl_probe_mem() shouldn't silently change from
> > > "attach to a FW-established region" into "allocate resources and program a new
> > > region". Those operations should have different semantics and ownership
> > > expectations.  
> > 
> > 
> > Glad with the consensus :-)
> > 
> >   
> > > 
> > > I am planning to rebase onto cxl/nexxt and rework the proposal as the following,
> > > please take a look and see if that matches your imagination or not.
> > > * Keep devm_cxl_probe_mem() behavior unchanged for FW-committed regions
> > > * Make region creation an explicit request from the accelerator provider,
> > >    rather than an automatic fallback during memdev attach
> > > * Have the provider specify the required size. CXL core shouldn't assume
> > >    that it maybe consume the entire volatile DPA partition as you mentioned.
> > > * Separate the reusable region-provisioning mechanism from the initial Type-2
> > >    policy.
> > > * The common mechanism should handle HPA/DPA allocation, decoder-path
> > >    construction, commit , rollback and managed teardown.
> > > * The initial type-2 caller would constrain that to volatile DEVMEM, IW=1
> > >    and a provider-requested size.
> > > 
> > > Oh and I'll replace "x1" with "IW=1" and explain the initial decoder,
> > > root-selection and granularity restriction more clearly.
> > > 
> > > How does that sound to you ?  
> > 
> > 
> > It sounds perfect!
> > 
> > 
> > FWIW, you likely saw Gregory's comment (discord) on this work requiring the
> > support for PMEM or at least the awareness PMEM support will need to use
> > same interface. His opinion and mine came from Dan's vision on this, and
> > your work will be the base for such PMEM support. I do not have an impending
> > reason for working on this PMEM support, but I am really interested in how
> > Type3 PMEMs can leverage CXL.mem for improving storage needs, and currently
> > reading/thinking about all this ...
> > 
> > 
> > Thanks!
> >   
> 
> Hmmm for this part I have no idea for now, I'll study more and discuss with you guys.
> 
> Best regards,
> Richard Cheng.
> 
> >   
> > > [1]: https://computeexpresslink.org/wp-content/uploads/2024/02/CEDT_ECN_1.0A_Eval.pdf
> > > [2]: https://docs.kernel.org/driver-api/cxl/linux/cxl-driver.html#runtime-programming
> > > 
> > > Best regards,
> > > Richard Cheng.  
> > > > > Testing result is in the following.
> > > > > - Built clean with clang/LLVM on arm64
> > > > > - cxl_test, type2_test=1. accel0 takes the unchanged attach path. accel1
> > > > >     drives auto_create -> a committed 512 MB RAM region. The test asserts
> > > > >     the 512 MB HPA range. committed state and 256 byte granularity confirmed
> > > > >     via sysfs.
> > > > > - Unbind tears the region down with no orphaned decoder, rebind re-creates
> > > > >     a fresh committed region.
> > > > > - Mock test only. Real accelerators whose FW commits a decoder take the
> > > > >     attach path, and vfio-cxl binds only FW-committed devices, so auto-create
> > > > >     has no real-HW caller yet.
> > > > > 
> > > > > Best regards,
> > > > > Richard Cheng.
> > > > > 
> > > > > Richard Cheng (3):
> > > > >     cxl/region: Reset software-created regions on memdev detach
> > > > >     cxl/region: Auto-create a region for memdev attach
> > > > >     cxl/test: Exercise Type-2 automatic region creation
> > > > > 
> > > > >    drivers/cxl/core/region.c      | 422 +++++++++++++++++++++++++++++----
> > > > >    tools/testing/cxl/test/accel.c |   7 +
> > > > >    tools/testing/cxl/test/cxl.c   |  61 ++++-
> > > > >    3 files changed, 439 insertions(+), 51 deletions(-)
> > > > > 
> > > > > 
> > > > > base-commit: 1c6b4ceafc3b994871c29340e0c1ddb0af5800e7