[PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT

Bert Karwatzki posted 1 patch 2 months ago
There is a newer version of this series
drivers/gpu/drm/amd/display/dc/core/dc_stream.c  | 2 +-
drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 4 ++--
2 files changed, 3 insertions(+), 3 deletions(-)
[PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Bert Karwatzki 2 months ago
On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
converted to rt_mutex. dc_create_plane_state() can be called while
inside an FPU-guarded region, resuling in "scheduling while atomic"
errors on PREEMPT_RT kernels.
 Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
Also fix the error path in dc_create_stream_for_sink().

Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/

Signed-off-by: Bert Karwatzki <spasswolf@web.de>
---
 drivers/gpu/drm/amd/display/dc/core/dc_stream.c  | 2 +-
 drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 4 ++--
 2 files changed, 3 insertions(+), 3 deletions(-)

diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
index cdcf140bc1bb..accad9e20e88 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
@@ -229,7 +229,7 @@ struct dc_stream_state *dc_create_stream_for_sink(
 
 fail:
 	if (stream)
-		kfree(stream);
+		DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));
 
 	return NULL;
 }
diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
index 88e825a6582c..d5c6427796b6 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
@@ -85,8 +85,8 @@ uint8_t  dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
  ******************************************************************************/
 struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
 {
-	struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
-							  GFP_ATOMIC);
+	struct dc_plane_state *plane_state;
+	DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
 
 	if (NULL == plane_state)
 		return NULL;
-- 
2.53.0


Rebased to next-20260729+. In these version the allocation of
update_scratch has been removed from dc_create_stream_for_sink().

Bert Karwatzki
Re: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by kernel test robot 1 month, 4 weeks ago
Hi Bert,

kernel test robot noticed the following build errors:

[auto build test ERROR on next-20260804]
[cannot apply to drm-misc/drm-misc-next linus/master v7.2-rc6 v7.2-rc5 v7.2-rc4 v7.2-rc6]
[If your patch is applied to the wrong git tree, kindly drop us a note.
And when submitting patch, we suggest to use '--base' as documented in
https://git-scm.com/docs/git-format-patch#_base_tree_information]

url:    https://github.com/intel-lab-lkp/linux/commits/Bert-Karwatzki/drm-amd-display-fix-usage-of-DC_FPU_-BEGIN-END-with-PREEMPT_RT/20260805-170330
base:   next-20260804
patch link:    https://lore.kernel.org/r/20260801071724.12998-1-spasswolf%40web.de
patch subject: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
config: alpha-allyesconfig (https://download.01.org/0day-ci/archive/20260806/202608061233.eufnR5Qm-lkp@intel.com/config)
compiler: alpha-linux-gcc (GCC) 16.1.0
reproduce (this is a W=1 build): (https://download.01.org/0day-ci/archive/20260806/202608061233.eufnR5Qm-lkp@intel.com/reproduce)

If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <lkp@intel.com>
| Closes: https://lore.kernel.org/oe-kbuild-all/202608061233.eufnR5Qm-lkp@intel.com/

All errors (new ones prefixed by >>):

   drivers/gpu/drm/amd/amdgpu/../display/dc/core/dc_surface.c: In function 'dc_create_plane_state':
>> drivers/gpu/drm/amd/amdgpu/../display/dc/core/dc_surface.c:89:9: error: implicit declaration of function 'DC_RUN_WITH_PREEMPTION_ENABLED' [-Wimplicit-function-declaration]
      89 |         DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
         |         ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~


vim +/DC_RUN_WITH_PREEMPTION_ENABLED +89 drivers/gpu/drm/amd/amdgpu/../display/dc/core/dc_surface.c

    82	
    83	/*******************************************************************************
    84	 * Public functions
    85	 ******************************************************************************/
    86	struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
    87	{
    88		struct dc_plane_state *plane_state;
  > 89		DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
    90	
    91		if (NULL == plane_state)
    92			return NULL;
    93	
    94		kref_init(&plane_state->refcount);
    95		dc_plane_construct(dc->ctx, plane_state);
    96	
    97		return plane_state;
    98	}
    99	

--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
[PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Bert Karwatzki 1 month, 3 weeks ago
On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
converted to rt_mutex. dc_create_plane_state() can be called while
inside an FPU-guarded region, resuling in "scheduling while atomic"
errors on PREEMPT_RT kernels.
 Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
Also fix the error path in dc_create_stream_for_sink().

Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
Signed-off-by: Bert Karwatzki <spasswolf@web.de>
---
 drivers/gpu/drm/amd/display/dc/core/dc_stream.c  | 5 +++--
 drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 5 +++--
 2 files changed, 6 insertions(+), 4 deletions(-)

diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
index 7666cdc78f4e..a5a304a3f802 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
@@ -233,8 +233,9 @@ struct dc_stream_state *dc_create_stream_for_sink(
 
 fail:
 	if (stream) {
-		kfree(stream->update_scratch);
-		kfree(stream);
+		if (stream->update_scratch)
+			DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream->update_scratch));
+		DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));
 	}
 
 	return NULL;
diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
index 72845fc788f3..04982673ffbc 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
@@ -33,6 +33,7 @@
 #include "dpp.h"
 
 #include "dc_plane_priv.h"
+#include "dc_fpu.h"
 
 /*******************************************************************************
  * Private functions
@@ -86,8 +87,8 @@ uint8_t  dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
  ******************************************************************************/
 struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
 {
-	struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
-							  GFP_ATOMIC);
+	struct dc_plane_state *plane_state;
+	DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
 
 	if (NULL == plane_state)
 		return NULL;
-- 
2.55.0

Fix for stable with the old error handling in
dc_create_stream_for_sink() and compile fix form arm64/clang.

Bert Karwatzki
Re: [PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Sebastian Andrzej Siewior 1 month ago
On 2026-08-07 14:49:42 [+0200], Bert Karwatzki wrote:
> On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> converted to rt_mutex. dc_create_plane_state() can be called while
> inside an FPU-guarded region, resuling in "scheduling while atomic"
> errors on PREEMPT_RT kernels.
>  Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> Also fix the error path in dc_create_stream_for_sink().
> 
> Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> Signed-off-by: Bert Karwatzki <spasswolf@web.de>

Thank you Bert. Got this somewhere in the meantime?

Could the FPU regions with disabled preemption be limited to where we
have actually have FPU usage in way that you don't have to worry when it
is needed to use DC_RUN_WITH_PREEMPTION_ENABLED() and when not? Also the
dc_fpu_begin()/ end() can nest and if they do the usage of
DC_RUN_WITH_PREEMPTION_ENABLED() is futile, isn't it?

The first usage kernel_fpu_begin() saves the FPU state to the user task.
kernel_fpu_end() does not restore it. Therefore the subsequent
invocation of kernel_fpu_begin() is cheaper.

Sebastian
Re: [PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Bert Karwatzki 1 month ago
Am Donnerstag, dem 27.08.2026 um 14:49 +0200 schrieb Sebastian Andrzej Siewior:
> On 2026-08-07 14:49:42 [+0200], Bert Karwatzki wrote:
> > On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> > converted to rt_mutex. dc_create_plane_state() can be called while
> > inside an FPU-guarded region, resuling in "scheduling while atomic"
> > errors on PREEMPT_RT kernels.
> >  Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> > Also fix the error path in dc_create_stream_for_sink().
> > 
> > Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> > Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> > Signed-off-by: Bert Karwatzki <spasswolf@web.de>
> 
> Thank you Bert. Got this somewhere in the meantime?

No, this is seems to be stuck somewhere ...

> 
> Could the FPU regions with disabled preemption be limited to where we
> have actually have FPU usage in way that you don't have to worry when it
> is needed to use DC_RUN_WITH_PREEMPTION_ENABLED() and when not? Also the
> dc_fpu_begin()/ end() can nest and if they do the usage of
> DC_RUN_WITH_PREEMPTION_ENABLED() is futile, isn't it?
> 

I'm not very familiar with the amd display engine, and it's a lot of code,
but there I think there's some room for improvement, e.g. dml2_destroy():

dml2_destroy is called from dc_state_free() from within DC_FP_{START,END}()
and then basically just calls vfree with DC_RUN_WITH_PREEMPTION_ENABLED().
(directly or through dml21_destroy()). Here both the DC_FP_*() tags and
DC_RUN_WITH_PREEMPTION_ENABLED() could be dropped I think (not tested, yet)

Regarding nesting DC_FP_*()s: When I was searching for the cause of the
kernel panics I monitored every DC_FP_{START,END}() with printk(), and I
don't think I encountered nesting (the logs are lost unfortunately)


> The first usage kernel_fpu_begin() saves the FPU state to the user task.
> kernel_fpu_end() does not restore it. Therefore the subsequent
> invocation of kernel_fpu_begin() is cheaper.
> 
> Sebastian

Bert Karwatzki
Re: [PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Sebastian Andrzej Siewior 1 month ago
On 2026-08-27 23:07:08 [+0200], Bert Karwatzki wrote:
> Am Donnerstag, dem 27.08.2026 um 14:49 +0200 schrieb Sebastian Andrzej Siewior:
> > On 2026-08-07 14:49:42 [+0200], Bert Karwatzki wrote:
> > > On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> > > converted to rt_mutex. dc_create_plane_state() can be called while
> > > inside an FPU-guarded region, resuling in "scheduling while atomic"
> > > errors on PREEMPT_RT kernels.
> > >  Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> > > Also fix the error path in dc_create_stream_for_sink().
> > > 
> > > Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> > > Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> > > Signed-off-by: Bert Karwatzki <spasswolf@web.de>
> > 
> > Thank you Bert. Got this somewhere in the meantime?
> 
> No, this is seems to be stuck somewhere ...

halp!

> > 
> > Could the FPU regions with disabled preemption be limited to where we
> > have actually have FPU usage in way that you don't have to worry when it
> > is needed to use DC_RUN_WITH_PREEMPTION_ENABLED() and when not? Also the
> > dc_fpu_begin()/ end() can nest and if they do the usage of
> > DC_RUN_WITH_PREEMPTION_ENABLED() is futile, isn't it?
> > 
> 
> I'm not very familiar with the amd display engine, and it's a lot of code,
> but there I think there's some room for improvement, e.g. dml2_destroy():
> 
> dml2_destroy is called from dc_state_free() from within DC_FP_{START,END}()
> and then basically just calls vfree with DC_RUN_WITH_PREEMPTION_ENABLED().
> (directly or through dml21_destroy()). Here both the DC_FP_*() tags and
> DC_RUN_WITH_PREEMPTION_ENABLED() could be dropped I think (not tested, yet)
> 
> Regarding nesting DC_FP_*()s: When I was searching for the cause of the
> kernel panics I monitored every DC_FP_{START,END}() with printk(), and I
> don't think I encountered nesting (the logs are lost unfortunately)

I am not saying it is nesting, just the START/STOP API is able to deal
this. Otherwise kernel_fpu_being/end() could be used directly. And
DC_RUN_WITH_PREEMPTION_ENABLED() is not able to deal with nesting.

On PREEMPT_RT it could be bad for the max observed latency if this nests
or has otherwise an unbound runtime. For crypto there an maximum amount
of input data which can be processed within a single FPU disabled
section. The "cheaper" second kernel_fpu_begin() makes it possible.

> > The first usage kernel_fpu_begin() saves the FPU state to the user task.
> > kernel_fpu_end() does not restore it. Therefore the subsequent
> > invocation of kernel_fpu_begin() is cheaper.
> > 
> > Sebastian
> 
> Bert Karwatzki

Sebastian
Re: [PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Bert Karwatzki 1 month ago
> 
> > > 
> > > Could the FPU regions with disabled preemption be limited to where we
> > > have actually have FPU usage in way that you don't have to worry when it
> > > is needed to use DC_RUN_WITH_PREEMPTION_ENABLED() and when not? Also the
> > > dc_fpu_begin()/ end() can nest and if they do the usage of
> > > DC_RUN_WITH_PREEMPTION_ENABLED() is futile, isn't it?
> > > 
> > 
> > I'm not very familiar with the amd display engine, and it's a lot of code,
> > but there I think there's some room for improvement, e.g. dml2_destroy():
> > 
> > dml2_destroy is called from dc_state_free() from within DC_FP_{START,END}()
> > and then basically just calls vfree with DC_RUN_WITH_PREEMPTION_ENABLED().
> > (directly or through dml21_destroy()). Here both the DC_FP_*() tags and
> > DC_RUN_WITH_PREEMPTION_ENABLED() could be dropped I think (not tested, yet)
> > 

Now, I've tested dropping some DC_FP_START,END() and DC_RUN_WITH_PREEMPTION_ENABLED().
(Patch for next-20260828+ with the fix above applied)
This works without error so far.
The function dml2_create() and dml2_create_copy() are moved to a non-FPU file so they
do not accidently use FPU instruction (they do not use FPU instruction on x86_64 when
compiled with gcc-16 but other architectures and compilers could probably use them).

diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_state.c b/drivers/gpu/drm/amd/display/dc/core/dc_state.c
index 666212cac105..709c37d2072f 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_state.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_state.c
@@ -209,9 +209,7 @@ struct dc_state *dc_state_create(struct dc *dc, struct dc_state_create_params *p
 	bool status;
 
 	if (dc->debug.using_dml2) {
-		DC_FP_START();
 		status = dml2_create(dc, &dc->dml2_options, &state->bw_ctx.dml2);
-		DC_FP_END();
 
 		if (!status) {
 			dc_state_release(state);
@@ -221,9 +219,7 @@ struct dc_state *dc_state_create(struct dc *dc, struct dc_state_create_params *p
 		if (dc->caps.dcmode_power_limits_present) {
 			bool dc_power_status;
 
-			DC_FP_START();
 			dc_power_status = dml2_create(dc, &dc->dml2_dc_power_options, &state->bw_ctx.dml2_dc_power_source);
-			DC_FP_END();
 
 			if (!dc_power_status) {
 				dc_state_release(state);
@@ -251,17 +247,13 @@ void dc_state_copy(struct dc_state *dst_state, struct dc_state *src_state)
 #ifdef CONFIG_DRM_AMD_DC_FP
 	dst_state->bw_ctx.dml2 = dst_dml2;
 	if (src_state->bw_ctx.dml2) {
-		DC_FP_START();
 		dml2_copy(dst_state->bw_ctx.dml2, src_state->bw_ctx.dml2);
-		DC_FP_END();
 	}
 
 	dst_state->bw_ctx.dml2_dc_power_source = dst_dml2_dc_power_source;
 
 	if (src_state->bw_ctx.dml2_dc_power_source) {
-		DC_FP_START();
 		dml2_copy(dst_state->bw_ctx.dml2_dc_power_source, src_state->bw_ctx.dml2_dc_power_source);
-		DC_FP_END();
 	}
 #endif // CONFIG_DRM_AMD_DC_FP
 	/* context refcount should not be overridden */
@@ -285,9 +277,7 @@ struct dc_state *dc_state_create_copy(struct dc_state *src_state)
 	new_state->bw_ctx.dml2_dc_power_source = NULL;
 
 	if (src_state->bw_ctx.dml2) {
-		DC_FP_START();
 		status = dml2_create_copy(&new_state->bw_ctx.dml2, src_state->bw_ctx.dml2);
-		DC_FP_END();
 
 		if (!status) {
 			dc_state_release(new_state);
@@ -297,10 +287,8 @@ struct dc_state *dc_state_create_copy(struct dc_state *src_state)
 
 
 	if (src_state->bw_ctx.dml2_dc_power_source) {
-		DC_FP_START();
 		status = dml2_create_copy(&new_state->bw_ctx.dml2_dc_power_source,
 					  src_state->bw_ctx.dml2_dc_power_source);
-		DC_FP_END();
 
 		if (!status) {
 			dc_state_release(new_state);
@@ -389,13 +377,11 @@ static void dc_state_free(struct kref *kref)
 	dc_state_destruct(state);
 
 #ifdef CONFIG_DRM_AMD_DC_FP
-	DC_FP_START();
 	dml2_destroy(state->bw_ctx.dml2);
 	state->bw_ctx.dml2 = 0;
 
 	dml2_destroy(state->bw_ctx.dml2_dc_power_source);
 	state->bw_ctx.dml2_dc_power_source = 0;
-	DC_FP_END();
 #endif
 
 	kvfree(state);
diff --git a/drivers/gpu/drm/amd/display/dc/dml2_0/dml21/dml21_wrapper.c b/drivers/gpu/drm/amd/display/dc/dml2_0/dml21/dml21_wrapper.c
index 8bed59e976d1..29e5cce51b99 100644
--- a/drivers/gpu/drm/amd/display/dc/dml2_0/dml21/dml21_wrapper.c
+++ b/drivers/gpu/drm/amd/display/dc/dml2_0/dml21/dml21_wrapper.c
@@ -23,11 +23,11 @@
 
 static bool dml21_allocate_memory(struct dml2_context **dml_ctx)
 {
-	DC_RUN_WITH_PREEMPTION_ENABLED(*dml_ctx = vzalloc(sizeof(struct dml2_context)));
+	*dml_ctx = vzalloc(sizeof(struct dml2_context));
 	if (!(*dml_ctx))
 		return false;
 
-	DC_RUN_WITH_PREEMPTION_ENABLED((*dml_ctx)->v21.dml_init.dml2_instance = vzalloc(sizeof(struct dml2_instance)));
+	(*dml_ctx)->v21.dml_init.dml2_instance = vzalloc(sizeof(struct dml2_instance));
 	if (!((*dml_ctx)->v21.dml_init.dml2_instance))
 		return false;
 
@@ -37,7 +37,7 @@ static bool dml21_allocate_memory(struct dml2_context **dml_ctx)
 	(*dml_ctx)->v21.mode_support.display_config = &(*dml_ctx)->v21.display_config;
 	(*dml_ctx)->v21.mode_programming.display_config = (*dml_ctx)->v21.mode_support.display_config;
 
-	DC_RUN_WITH_PREEMPTION_ENABLED((*dml_ctx)->v21.mode_programming.programming = vzalloc(sizeof(struct dml2_display_cfg_programming)));
+	(*dml_ctx)->v21.mode_programming.programming = vzalloc(sizeof(struct dml2_display_cfg_programming));
 
 	if (!((*dml_ctx)->v21.mode_programming.programming))
 		return false;
@@ -51,15 +51,17 @@ bool dml21_create(const struct dc *in_dc, struct dml2_context **dml_ctx, const s
 	if (!dml21_allocate_memory(dml_ctx))
 		return false;
 
+	DC_FP_START();
 	dml21_init(in_dc, *dml_ctx, config);
+	DC_FP_END();
 
 	return true;
 }
 
 void dml21_destroy(struct dml2_context *dml2)
 {
-	DC_RUN_WITH_PREEMPTION_ENABLED(vfree(dml2->v21.dml_init.dml2_instance));
-	DC_RUN_WITH_PREEMPTION_ENABLED(vfree(dml2->v21.mode_programming.programming));
+	vfree(dml2->v21.dml_init.dml2_instance);
+	vfree(dml2->v21.mode_programming.programming);
 }
 
 void dml21_copy(struct dml2_context *dst_dml_ctx,
@@ -88,7 +90,9 @@ void dml21_copy(struct dml2_context *dst_dml_ctx,
 	dst_dml_ctx->v21.mode_programming.programming = dst_dml2_programming;
 
 	/* need to initialize copied instance for internal references to be correct */
+	DC_FP_START();
 	dml2_initialize_instance(&dst_dml_ctx->v21.dml_init);
+	DC_FP_END();
 }
 
 bool dml21_create_copy(struct dml2_context **dst_dml_ctx,
diff --git a/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper.c b/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper.c
index 1772e74349c7..570da14fb1a7 100644
--- a/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper.c
+++ b/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper.c
@@ -21,7 +21,7 @@ struct dml2_context *dml2_allocate_memory(void)
 {
 	struct dml2_context *dml2;
 
-	DC_RUN_WITH_PREEMPTION_ENABLED(dml2 = vzalloc(sizeof(struct dml2_context)));
+	dml2 = vzalloc(sizeof(struct dml2_context));
 	return dml2;
 }
 bool dml2_validate(const struct dc *in_dc, struct dc_state *context, struct dml2_context *dml2,
@@ -51,7 +51,9 @@ bool dml2_validate(const struct dc *in_dc, struct dc_state *context, struct dml2
 static void dml2_init(const struct dc *in_dc, const struct dml2_configuration_options *config, struct dml2_context **dml2)
 {
 	if ((in_dc->debug.using_dml21) && (in_dc->ctx->dce_version >= DCN_VERSION_4_01)) {
+		DC_FP_START();
 		dml21_reinit(in_dc, *dml2, config);
+		DC_FP_END();
 		return;
 	}
 
@@ -82,11 +84,13 @@ static void dml2_init(const struct dc *in_dc, const struct dml2_configuration_op
 		break;
 	}
 
+	DC_FP_START();
 	initialize_dml2_ip_params(*dml2, in_dc, &(*dml2)->v20.dml_core_ctx.ip);
 
 	initialize_dml2_soc_bbox(*dml2, in_dc, &(*dml2)->v20.dml_core_ctx.soc);
 
 	initialize_dml2_soc_states(*dml2, in_dc, &(*dml2)->v20.dml_core_ctx.soc, &(*dml2)->v20.dml_core_ctx.states);
+	DC_FP_END();
 
 }
 
@@ -116,17 +120,52 @@ void dml2_destroy(struct dml2_context *dml2)
 	if (dml2->architecture == dml2_architecture_21)
 		dml21_destroy(dml2);
 
-	DC_RUN_WITH_PREEMPTION_ENABLED(vfree(dml2));
+	vfree(dml2);
 }
 
 void dml2_reinit(const struct dc *in_dc,
 				 const struct dml2_configuration_options *config,
 				 struct dml2_context **dml2)
 {
+	/*
 	if ((in_dc->debug.using_dml21) && (in_dc->ctx->dce_version >= DCN_VERSION_4_01)) {
+		DC_FP_START();
 		dml21_reinit(in_dc, *dml2, config);
+		DC_FP_END();
 		return;
 	}
+	*/
 
 	dml2_init(in_dc, config, dml2);
 }
+
+/* Moved here from dml2_wrapper_fpu.c */
+void dml2_copy(struct dml2_context *dst_dml2,
+	struct dml2_context *src_dml2)
+{
+	if (src_dml2->architecture == dml2_architecture_21) {
+		dml21_copy(dst_dml2, src_dml2);
+		return;
+	}
+	/* copy Mode Lib Ctx */
+	memcpy(dst_dml2, src_dml2, sizeof(struct dml2_context));
+}
+
+/* Moved here from dml2_wrapper_fpu.c */
+bool dml2_create_copy(struct dml2_context **dst_dml2,
+	struct dml2_context *src_dml2)
+{
+	if (src_dml2->architecture == dml2_architecture_21)
+		return dml21_create_copy(dst_dml2, src_dml2);
+	/* Allocate Mode Lib Ctx */
+	*dst_dml2 = dml2_allocate_memory();
+
+	if (!(*dst_dml2))
+		return false;
+
+	/* copy Mode Lib Ctx */
+	dml2_copy(*dst_dml2, src_dml2);
+
+	return true;
+}
+
diff --git a/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper_fpu.c b/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper_fpu.c
index a14e3004a7b7..b590d58ad3a9 100644
--- a/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper_fpu.c
+++ b/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper_fpu.c
@@ -561,31 +561,3 @@ void dml2_prepare_mcache_programming(struct dc *in_dc, struct dc_state *context,
 		dml21_prepare_mcache_programming(in_dc, context, dml2);
 }
 
-void dml2_copy(struct dml2_context *dst_dml2,
-	struct dml2_context *src_dml2)
-{
-	if (src_dml2->architecture == dml2_architecture_21) {
-		dml21_copy(dst_dml2, src_dml2);
-		return;
-	}
-	/* copy Mode Lib Ctx */
-	memcpy(dst_dml2, src_dml2, sizeof(struct dml2_context));
-}
-
-bool dml2_create_copy(struct dml2_context **dst_dml2,
-	struct dml2_context *src_dml2)
-{
-	if (src_dml2->architecture == dml2_architecture_21)
-		return dml21_create_copy(dst_dml2, src_dml2);
-	/* Allocate Mode Lib Ctx */
-	*dst_dml2 = dml2_allocate_memory();
-
-	if (!(*dst_dml2))
-		return false;
-
-	/* copy Mode Lib Ctx */
-	dml2_copy(*dst_dml2, src_dml2);
-
-	return true;
-}
-


Bert Karwatzki
Re: [PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Greg KH 1 month, 3 weeks ago
On Fri, Aug 07, 2026 at 02:49:42PM +0200, Bert Karwatzki wrote:
> On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> converted to rt_mutex. dc_create_plane_state() can be called while
> inside an FPU-guarded region, resuling in "scheduling while atomic"
> errors on PREEMPT_RT kernels.
>  Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> Also fix the error path in dc_create_stream_for_sink().
> 
> Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> Signed-off-by: Bert Karwatzki <spasswolf@web.de>
> ---
>  drivers/gpu/drm/amd/display/dc/core/dc_stream.c  | 5 +++--
>  drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 5 +++--
>  2 files changed, 6 insertions(+), 4 deletions(-)
> 
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> index 7666cdc78f4e..a5a304a3f802 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> @@ -233,8 +233,9 @@ struct dc_stream_state *dc_create_stream_for_sink(
>  
>  fail:
>  	if (stream) {
> -		kfree(stream->update_scratch);
> -		kfree(stream);
> +		if (stream->update_scratch)
> +			DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream->update_scratch));
> +		DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));
>  	}
>  
>  	return NULL;
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> index 72845fc788f3..04982673ffbc 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> @@ -33,6 +33,7 @@
>  #include "dpp.h"
>  
>  #include "dc_plane_priv.h"
> +#include "dc_fpu.h"
>  
>  /*******************************************************************************
>   * Private functions
> @@ -86,8 +87,8 @@ uint8_t  dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
>   ******************************************************************************/
>  struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
>  {
> -	struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
> -							  GFP_ATOMIC);
> +	struct dc_plane_state *plane_state;
> +	DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
>  
>  	if (NULL == plane_state)
>  		return NULL;
> -- 
> 2.55.0
> 
> Fix for stable with the old error handling in
> dc_create_stream_for_sink() and compile fix form arm64/clang.
> 
> Bert Karwatzki
> 

<formletter>

This is not the correct way to submit patches for inclusion in the
stable kernel tree.  Please read:
    https://www.kernel.org/doc/html/latest/process/stable-kernel-rules.html
for how to do this properly.

</formletter>
Re: [PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by sashiko-bot@kernel.org 1 month, 3 weeks ago
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:

Pre-existing issues:
- [High] Incomplete Fix: Normal deallocation paths (dc_stream_free and dc_plane_state_free) still sleep in atomic context.
- [High] DC_RUN_WITH_PREEMPTION_ENABLED fails to re-enable preemption if FPU blocks are nested.
- [High] DC_RUN_WITH_PREEMPTION_ENABLED corrupts callee-saved FPU registers on ARM64.
--

--- Patch [1]: [PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT ---
commit e67f6942f0899807d82dcd2276705097e892a34b
Author: Bert Karwatzki <spasswolf@web.de>

drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT

This commit wraps deallocations in dc_create_stream_for_sink and
dc_create_plane_state with DC_RUN_WITH_PREEMPTION_ENABLED. This prevents
"scheduling while atomic" panics on PREEMPT_RT kernels when sleeping
allocations are called from within FPU-guarded regions.

Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/

> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> index 7666cdc78f4ed..a5a304a3f8029 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c

[ ... ]

> @@ -233,8 +233,9 @@ struct dc_stream_state *dc_create_stream_for_sink(
>  
>  fail:
>  	if (stream) {
> -		kfree(stream->update_scratch);
> -		kfree(stream);
> +		if (stream->update_scratch)
> +			DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream->update_scratch));
> +		DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));

[Severity: High]
This isn't a bug introduced by this patch, but will the normal deallocation
paths still sleep in atomic context?

When dc_state_release_phantom_streams_and_planes() is called from
dcn32_internal_validate_bw(), it executes under DC_FP_START where
preemption is disabled. That validation path drops references to phantom
streams and planes, which synchronously invokes dc_stream_free() and
dc_plane_state_free().

Those destructors unconditionally call kfree() and kvfree(). On PREEMPT_RT,
slab allocators use sleepable rt_mutexes, which will still cause a kernel
panic from scheduling while atomic. Is it possible to apply this fix to
those deallocation paths as well?

>  	}
>  
>  	return NULL;
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> index 72845fc788f3f..04982673ffbc9 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c

[ ... ]

> @@ -86,8 +87,8 @@ uint8_t  dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
>   ******************************************************************************/
>  struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
>  {
> -	struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
> -							  GFP_ATOMIC);
> +	struct dc_plane_state *plane_state;
> +	DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));

[Severity: High]
This isn't a bug introduced by this patch, but does
DC_RUN_WITH_PREEMPTION_ENABLED() fail to re-enable preemption if FPU blocks
are nested?

In dc_fpu_begin(), preempt_disable() is called unconditionally on every
entry, meaning preempt_count increases with each nesting level. However,
DC_RUN_WITH_PREEMPTION_ENABLED() calls DC_FP_END() exactly once, which
only decrements the recursion depth and calls preempt_enable() once.

If the FPU recursion depth is 2 or more, preemption remains disabled and
the wrapped kvzalloc_obj() will still crash on PREEMPT_RT.


[Severity: High]
This is a pre-existing issue, but can DC_RUN_WITH_PREEMPTION_ENABLED()
corrupt callee-saved FPU registers on ARM64 when used here?

A caller compiled with FPU support, such as
dml2_mall_phantom_streams_and_planes_create(), has live variables in
callee-saved FPU registers (v8-v15 on ARM64). It calls
dc_state_create_phantom_plane() which routes to dc_create_plane_state().

Since dc_create_plane_state() is compiled without FPU support, it doesn't
save v8-v15 to the stack. When DC_RUN_WITH_PREEMPTION_ENABLED() calls
DC_FP_END(), it invokes kernel_neon_end() and enables preemption.

If a context switch occurs while preemption is enabled, the OS scheduler
will not save the thread's FPU state since the kernel has yielded it. The
v8-v15 registers are permanently clobbered, and execution returns to the
caller with garbage FPU values.

>  
>  	if (NULL == plane_state)
>  		return NULL;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260807124944.154232-1-spasswolf@web.de?part=1
Re: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by kernel test robot 1 month, 4 weeks ago
Hi Bert,

kernel test robot noticed the following build errors:

[auto build test ERROR on next-20260804]
[cannot apply to drm-misc/drm-misc-next linus/master v7.2-rc6 v7.2-rc5 v7.2-rc4 v7.2-rc6]
[If your patch is applied to the wrong git tree, kindly drop us a note.
And when submitting patch, we suggest to use '--base' as documented in
https://git-scm.com/docs/git-format-patch#_base_tree_information]

url:    https://github.com/intel-lab-lkp/linux/commits/Bert-Karwatzki/drm-amd-display-fix-usage-of-DC_FPU_-BEGIN-END-with-PREEMPT_RT/20260805-170330
base:   next-20260804
patch link:    https://lore.kernel.org/r/20260801071724.12998-1-spasswolf%40web.de
patch subject: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
config: loongarch-defconfig (https://download.01.org/0day-ci/archive/20260806/202608061154.QtrglsMo-lkp@intel.com/config)
compiler: clang version 24.0.0git (https://github.com/llvm/llvm-project bacfe2950f8218268fcc0a8765644ea0c15f0360)
reproduce (this is a W=1 build): (https://download.01.org/0day-ci/archive/20260806/202608061154.QtrglsMo-lkp@intel.com/reproduce)

If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <lkp@intel.com>
| Closes: https://lore.kernel.org/oe-kbuild-all/202608061154.QtrglsMo-lkp@intel.com/

All errors (new ones prefixed by >>):

>> drivers/gpu/drm/amd/amdgpu/../display/dc/core/dc_surface.c:89:2: error: call to undeclared function 'DC_RUN_WITH_PREEMPTION_ENABLED'; ISO C99 and later do not support implicit function declarations [-Wimplicit-function-declaration]
      89 |         DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
         |         ^
   1 error generated.


vim +/DC_RUN_WITH_PREEMPTION_ENABLED +89 drivers/gpu/drm/amd/amdgpu/../display/dc/core/dc_surface.c

    82	
    83	/*******************************************************************************
    84	 * Public functions
    85	 ******************************************************************************/
    86	struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
    87	{
    88		struct dc_plane_state *plane_state;
  > 89		DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
    90	
    91		if (NULL == plane_state)
    92			return NULL;
    93	
    94		kref_init(&plane_state->refcount);
    95		dc_plane_construct(dc->ctx, plane_state);
    96	
    97		return plane_state;
    98	}
    99	

--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
Re: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Mikhail Gavrilov 2 months ago
On Sat, Aug 1, 2026 at 12:17 PM Bert Karwatzki <spasswolf@web.de> wrote:
>
> On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> converted to rt_mutex. dc_create_plane_state() can be called while
> inside an FPU-guarded region, resuling in "scheduling while atomic"
> errors on PREEMPT_RT kernels.
>  Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> Also fix the error path in dc_create_stream_for_sink().
>
> Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
>
> Signed-off-by: Bert Karwatzki <spasswolf@web.de>
> ---
>  drivers/gpu/drm/amd/display/dc/core/dc_stream.c  | 2 +-
>  drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 4 ++--
>  2 files changed, 3 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> index cdcf140bc1bb..accad9e20e88 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> @@ -229,7 +229,7 @@ struct dc_stream_state *dc_create_stream_for_sink(
>
>  fail:
>         if (stream)
> -               kfree(stream);
> +               DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));
>
>         return NULL;
>  }
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> index 88e825a6582c..d5c6427796b6 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> @@ -85,8 +85,8 @@ uint8_t  dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
>   ******************************************************************************/
>  struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
>  {
> -       struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
> -                                                         GFP_ATOMIC);
> +       struct dc_plane_state *plane_state;
> +       DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));


This hunk breaks the build when CONFIG_DRM_AMD_DC_FP=n.

dc_surface.c gets DC_RUN_WITH_PREEMPTION_ENABLED only indirectly, via
dc.h -> dc_types.h -> os_types.h, and there the include is guarded:

#if defined(CONFIG_DRM_AMD_DC_FP)
#include "amdgpu_dm/dc_fpu.h"
#endif

With DC_FP=n nothing defines the macro in that file.  dc_stream.c is not
affected because it includes dc_fpu.h directly.

This is not only exotic architectures: DRM_AMD_DC has

select DRM_AMD_DC_FP if ARCH_HAS_KERNEL_FPU_SUPPORT && \
!(CC_IS_CLANG && (ARM64 || LOONGARCH || RISCV))

so an arm64 clang build is enough.  On amd-staging-drm-next with your
patch applied:

  $ make LLVM=1 ARCH=arm64 allmodconfig
  $ make LLVM=1 ARCH=arm64 drivers/gpu/drm/amd/amdgpu/
  dc_surface.c:89:2: error: call to undeclared function
    'DC_RUN_WITH_PREEMPTION_ENABLED'; ISO C99 and later do not support
    implicit function declarations [-Wimplicit-function-declaration]
     89 |  DC_RUN_WITH_PREEMPTION_ENABLED(plane_state =
kvzalloc_obj(*plane_state, GFP_ATOMIC));

A plain x86_64 build (DC_FP=y) is clean, which is presumably why it did
not show up for you.

Adding

#include "dc_fpu.h"

to dc_surface.c fixes it, and matches what dc_stream.c already does.
Please do not copy the local fallback that sits below that include in
dc_stream.c:

#if !defined(DC_RUN_WITH_PREEMPTION_ENABLED)
#define DC_RUN_WITH_PREEMPTION_ENABLED(code) code
#endif

It is dead code there - dc_fpu.h defines the macro in every branch - and
in a new file it would compile cleanly while quietly doing nothing.

One unrelated note: you kept Cc: stable # v7.1, but this version is
rebased onto next-20260729, where the update_scratch allocation is gone
from dc_create_stream_for_sink().  It does not apply to current mainline
either, so older trees will need a separate backport - worth saying so
in the commit message.

With the include added I can give this a Tested-by on dcn32 (RX 7900
XTX), which exercises the DML1 phantom-plane path rather than the dml21
one you hit.

-- 
Best Regards,
Mikhail Gavrilov.
[PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by Bert Karwatzki 1 month, 3 weeks ago
On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
converted to rt_mutex. dc_create_plane_state() can be called while
inside an FPU-guarded region, resuling in "scheduling while atomic"
errors on PREEMPT_RT kernels.
 Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
Also fix the error path in dc_create_stream_for_sink().

Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/

Signed-off-by: Bert Karwatzki <spasswolf@web.de>
---
 drivers/gpu/drm/amd/display/dc/core/dc_stream.c  | 2 +-
 drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 5 +++--
 2 files changed, 4 insertions(+), 3 deletions(-)

diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
index cdcf140bc1bb..accad9e20e88 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
@@ -229,7 +229,7 @@ struct dc_stream_state *dc_create_stream_for_sink(
 
 fail:
 	if (stream)
-		kfree(stream);
+		DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));
 
 	return NULL;
 }
diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
index 88e825a6582c..f4eb021711cb 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
@@ -33,6 +33,7 @@
 #include "dpp.h"
 
 #include "dc_plane_priv.h"
+#include "dc_fpu.h"
 
 /*******************************************************************************
  * Private functions
@@ -85,8 +86,8 @@ uint8_t  dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
  ******************************************************************************/
 struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
 {
-	struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
-							  GFP_ATOMIC);
+	struct dc_plane_state *plane_state;
+	DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
 
 	if (NULL == plane_state)
 		return NULL;
-- 
2.55.0

Added compile fix (by Mikhail Gavrilov) for CONFIG_DRM_AMD_DC_FP=n to
the patch for linux-next. Compile tested with arm64/clang on debian sid.

Bert Karwatzki
Re: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by sashiko-bot@kernel.org 1 month, 3 weeks ago
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:

New issues:
- [Critical] DC_RUN_WITH_PREEMPTION_ENABLED() silently clobbers callee-saved FPU/SIMD registers on architectures like ARM64.
- [High] DC_RUN_WITH_PREEMPTION_ENABLED() fails to enable preemption if the FPU recursion depth is greater than 1.

Pre-existing issues:
- [High] The teardown paths for streams and planes call kfree() and kvfree() inside FPU-guarded regions without dropping preemption.
--

--- Patch [1]: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT ---
commit ca829d4b6abc0445e9b58703c783e610dbd244a1
Author: Bert Karwatzki <spasswolf@web.de>

drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT

This commit addresses "scheduling while atomic" warnings on PREEMPT_RT by
wrapping memory allocations in DC_RUN_WITH_PREEMPTION_ENABLED() in the
plane and stream creation paths.

Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/

> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c

[ ... ]

> @@ -208,13 +208,13 @@ struct dc_stream_state *dc_create_stream_for_sink(
>  	if (sink == NULL)
>  		goto fail;
>  
> -	stream = kzalloc_obj(struct dc_stream_state, GFP_ATOMIC);
> +	DC_RUN_WITH_PREEMPTION_ENABLED(stream = kzalloc_obj(struct dc_stream_state, GFP_ATOMIC));

[Severity: Critical]
Does DC_RUN_WITH_PREEMPTION_ENABLED() silently clobber callee-saved FPU/SIMD
registers on architectures like ARM64?

The AMD display math library relies on FPU calculations. On ARM64, the ABI
designates certain SIMD registers (e.g., d8-d15 or v8-v15) as callee-saved.
When the display math library calls dc_create_stream_for_sink() from within
an FPU-guarded region, the C compiler expects these callee-saved registers
to be preserved.

When DC_RUN_WITH_PREEMPTION_ENABLED() calls DC_FP_END(), it invokes
kernel_neon_end(). This enables preemption and clears the TIF_KERNEL_FPSTATE
flag, but leaves TIF_FOREIGN_FPSTATE set. If a context switch occurs during
the wrapped kzalloc_obj() allocation, fpsimd_thread_switch() checks
TIF_FOREIGN_FPSTATE and skips saving the hardware FPU registers, erroneously
assuming they do not contain valid task state. The hardware FPU registers are
then clobbered by other tasks. When control returns to the math library, its
local floating-point variables stored in callee-saved registers will contain
garbage.

> @@ -229,7 +229,7 @@ struct dc_stream_state *dc_create_stream_for_sink(
>  
>  fail:
>  	if (stream)
> -		kfree(stream);
> +		DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));

[Severity: High]
This is a pre-existing issue, but do the teardown paths for streams and planes
still call kfree() and kvfree() inside FPU-guarded regions without dropping
preemption?

For example, phantom planes and streams are destroyed during bandwidth
validation by calling dc_state_release_phantom_streams_and_planes(), such as
from dcn32_internal_validate_bw(). This validation executes entirely inside an
FPU-guarded region created by DC_FP_START() in dcn32_validate_bandwidth(),
which disables preemption.

When the last reference is dropped, the teardown paths like dc_stream_free()
and dc_plane_state_free() execute synchronously. Since they call kfree() and
kvfree() directly without being wrapped in DC_RUN_WITH_PREEMPTION_ENABLED(),
these allocations are freed while preemption is disabled, which can trigger
the exact PREEMPT_RT bug this patch aims to fix.

>  	return NULL;
>  }

> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c

[ ... ]

> @@ -85,8 +86,8 @@ uint8_t  dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
>   ******************************************************************************/
>  struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
>  {
> -	struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
> -							  GFP_ATOMIC);
> +	struct dc_plane_state *plane_state;
> +	DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));

[Severity: High]
Does DC_RUN_WITH_PREEMPTION_ENABLED() fail to enable preemption if the FPU
recursion depth is greater than 1?

The FPU protection in the AMD display driver supports nesting via the
fpu_recursion_depth counter. DC_FP_END() only invokes kernel_fpu_end() (and
thus preempt_enable()) if the recursion depth decrements to exactly 0.

If dc_create_plane_state() or dc_create_stream_for_sink() is invoked from a
nested FPU-guarded region where the depth is 2 or more, the
DC_RUN_WITH_PREEMPTION_ENABLED() macro will merely decrement the depth
counter by 1 but leave preemption disabled. The subsequent call to
kvzalloc_obj() or kfree() will then execute with preemption still disabled,
which will still trigger a scheduling while atomic panic on PREEMPT_RT.

>  	if (NULL == plane_state)
>  		return NULL;
>

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260807125832.157856-1-spasswolf@web.de?part=1
Re: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
Posted by sashiko-bot@kernel.org 2 months ago
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

Pre-existing issues:
- [High] The patch wraps allocations and error-path frees in DC_RUN_WITH_PREEMPTION_ENABLED() to fix a 'scheduling while atomic' bug on PREEMPT_RT, but misses the normal destruction paths which are called from the exact same FPU-guarded contexts.
- [Medium] dc_state_create_phantom_stream() leaks a phantom stream reference if tracking fails.
--

--- Patch [1]: [PATCH] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT ---
commit 333ce260a6d2d044dc84469ac2172b5e7123c941
Author: Bert Karwatzki <spasswolf@web.de>

drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT

On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
converted to rt_mutex. dc_create_plane_state() can be called while
inside an FPU-guarded region, resuling in "scheduling while atomic"
errors on PREEMPT_RT kernels.
Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
Also fix the error path in dc_create_stream_for_sink().

Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/

> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> index cdcf140bc1bbd..accad9e20e88a 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> @@ -229,7 +229,7 @@ struct dc_stream_state *dc_create_stream_for_sink(
>  
>  fail:
>  	if (stream)
> -		kfree(stream);
> +		DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));

[Severity: High]
This isn't a bug introduced by this patch, but does this fix need to cover the
normal destruction paths as well?

The patch wraps allocations and error-path frees to prevent scheduling while
atomic errors on PREEMPT_RT kernels. However, normal destruction paths are
also called from the exact same FPU-guarded contexts.

For example, when dropping phantom streams inside an FPU-guarded block during
DML bandwidth validation:

dcn32_internal_validate_bw()
   dc_state_release_phantom_streams_and_planes()
      dc_stream_free()
         kfree(stream);
      dc_plane_state_free()
         kvfree(plane_state);

Should dc_stream_free() and dc_plane_state_free() also wrap their frees with
DC_RUN_WITH_PREEMPTION_ENABLED() to avoid triggering the same panic?

[Severity: Medium]
This is a pre-existing issue, but does dc_state_create_phantom_stream() in
drivers/gpu/drm/amd/display/dc/core/dc_state.c leak a phantom stream
reference if tracking fails?

If dc_state_track_phantom_stream() fails, for example because the
MAX_PHANTOM_PIPES limit is reached, it returns false. This return value is
ignored by dc_state_create_phantom_stream(). The stream is added to the main
array but not tracked in the phantom array.

When the state is destroyed, dc_state_destruct() only drops the reference
from the main array, permanently leaking the initial creation reference.

>  
>  	return NULL;
>  }

[ ... ]

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260801071724.12998-1-spasswolf@web.de?part=1