drivers/gpu/drm/amd/display/dc/core/dc_stream.c | 5 +++-- drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 5 +++-- 2 files changed, 6 insertions(+), 4 deletions(-)
On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
converted to rt_mutex. dc_create_plane_state() can be called while
inside an FPU-guarded region, resuling in "scheduling while atomic"
errors on PREEMPT_RT kernels.
Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
Also fix the error path in dc_create_stream_for_sink().
Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
Signed-off-by: Bert Karwatzki <spasswolf@web.de>
---
drivers/gpu/drm/amd/display/dc/core/dc_stream.c | 5 +++--
drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 5 +++--
2 files changed, 6 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
index 7666cdc78f4e..a5a304a3f802 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
@@ -233,8 +233,9 @@ struct dc_stream_state *dc_create_stream_for_sink(
fail:
if (stream) {
- kfree(stream->update_scratch);
- kfree(stream);
+ if (stream->update_scratch)
+ DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream->update_scratch));
+ DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));
}
return NULL;
diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
index 72845fc788f3..04982673ffbc 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
@@ -33,6 +33,7 @@
#include "dpp.h"
#include "dc_plane_priv.h"
+#include "dc_fpu.h"
/*******************************************************************************
* Private functions
@@ -86,8 +87,8 @@ uint8_t dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
******************************************************************************/
struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
{
- struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
- GFP_ATOMIC);
+ struct dc_plane_state *plane_state;
+ DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
if (NULL == plane_state)
return NULL;
--
2.55.0
Fix for stable with the old error handling in
dc_create_stream_for_sink() and compile fix form arm64/clang.
Bert Karwatzki
On 2026-08-07 14:49:42 [+0200], Bert Karwatzki wrote:
> On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> converted to rt_mutex. dc_create_plane_state() can be called while
> inside an FPU-guarded region, resuling in "scheduling while atomic"
> errors on PREEMPT_RT kernels.
> Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> Also fix the error path in dc_create_stream_for_sink().
>
> Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> Signed-off-by: Bert Karwatzki <spasswolf@web.de>
Thank you Bert. Got this somewhere in the meantime?
Could the FPU regions with disabled preemption be limited to where we
have actually have FPU usage in way that you don't have to worry when it
is needed to use DC_RUN_WITH_PREEMPTION_ENABLED() and when not? Also the
dc_fpu_begin()/ end() can nest and if they do the usage of
DC_RUN_WITH_PREEMPTION_ENABLED() is futile, isn't it?
The first usage kernel_fpu_begin() saves the FPU state to the user task.
kernel_fpu_end() does not restore it. Therefore the subsequent
invocation of kernel_fpu_begin() is cheaper.
Sebastian
Am Donnerstag, dem 27.08.2026 um 14:49 +0200 schrieb Sebastian Andrzej Siewior:
> On 2026-08-07 14:49:42 [+0200], Bert Karwatzki wrote:
> > On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> > converted to rt_mutex. dc_create_plane_state() can be called while
> > inside an FPU-guarded region, resuling in "scheduling while atomic"
> > errors on PREEMPT_RT kernels.
> > Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> > Also fix the error path in dc_create_stream_for_sink().
> >
> > Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> > Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> > Signed-off-by: Bert Karwatzki <spasswolf@web.de>
>
> Thank you Bert. Got this somewhere in the meantime?
No, this is seems to be stuck somewhere ...
>
> Could the FPU regions with disabled preemption be limited to where we
> have actually have FPU usage in way that you don't have to worry when it
> is needed to use DC_RUN_WITH_PREEMPTION_ENABLED() and when not? Also the
> dc_fpu_begin()/ end() can nest and if they do the usage of
> DC_RUN_WITH_PREEMPTION_ENABLED() is futile, isn't it?
>
I'm not very familiar with the amd display engine, and it's a lot of code,
but there I think there's some room for improvement, e.g. dml2_destroy():
dml2_destroy is called from dc_state_free() from within DC_FP_{START,END}()
and then basically just calls vfree with DC_RUN_WITH_PREEMPTION_ENABLED().
(directly or through dml21_destroy()). Here both the DC_FP_*() tags and
DC_RUN_WITH_PREEMPTION_ENABLED() could be dropped I think (not tested, yet)
Regarding nesting DC_FP_*()s: When I was searching for the cause of the
kernel panics I monitored every DC_FP_{START,END}() with printk(), and I
don't think I encountered nesting (the logs are lost unfortunately)
> The first usage kernel_fpu_begin() saves the FPU state to the user task.
> kernel_fpu_end() does not restore it. Therefore the subsequent
> invocation of kernel_fpu_begin() is cheaper.
>
> Sebastian
Bert Karwatzki
On 2026-08-27 23:07:08 [+0200], Bert Karwatzki wrote:
> Am Donnerstag, dem 27.08.2026 um 14:49 +0200 schrieb Sebastian Andrzej Siewior:
> > On 2026-08-07 14:49:42 [+0200], Bert Karwatzki wrote:
> > > On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> > > converted to rt_mutex. dc_create_plane_state() can be called while
> > > inside an FPU-guarded region, resuling in "scheduling while atomic"
> > > errors on PREEMPT_RT kernels.
> > > Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> > > Also fix the error path in dc_create_stream_for_sink().
> > >
> > > Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> > > Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> > > Signed-off-by: Bert Karwatzki <spasswolf@web.de>
> >
> > Thank you Bert. Got this somewhere in the meantime?
>
> No, this is seems to be stuck somewhere ...
halp!
> >
> > Could the FPU regions with disabled preemption be limited to where we
> > have actually have FPU usage in way that you don't have to worry when it
> > is needed to use DC_RUN_WITH_PREEMPTION_ENABLED() and when not? Also the
> > dc_fpu_begin()/ end() can nest and if they do the usage of
> > DC_RUN_WITH_PREEMPTION_ENABLED() is futile, isn't it?
> >
>
> I'm not very familiar with the amd display engine, and it's a lot of code,
> but there I think there's some room for improvement, e.g. dml2_destroy():
>
> dml2_destroy is called from dc_state_free() from within DC_FP_{START,END}()
> and then basically just calls vfree with DC_RUN_WITH_PREEMPTION_ENABLED().
> (directly or through dml21_destroy()). Here both the DC_FP_*() tags and
> DC_RUN_WITH_PREEMPTION_ENABLED() could be dropped I think (not tested, yet)
>
> Regarding nesting DC_FP_*()s: When I was searching for the cause of the
> kernel panics I monitored every DC_FP_{START,END}() with printk(), and I
> don't think I encountered nesting (the logs are lost unfortunately)
I am not saying it is nesting, just the START/STOP API is able to deal
this. Otherwise kernel_fpu_being/end() could be used directly. And
DC_RUN_WITH_PREEMPTION_ENABLED() is not able to deal with nesting.
On PREEMPT_RT it could be bad for the max observed latency if this nests
or has otherwise an unbound runtime. For crypto there an maximum amount
of input data which can be processed within a single FPU disabled
section. The "cheaper" second kernel_fpu_begin() makes it possible.
> > The first usage kernel_fpu_begin() saves the FPU state to the user task.
> > kernel_fpu_end() does not restore it. Therefore the subsequent
> > invocation of kernel_fpu_begin() is cheaper.
> >
> > Sebastian
>
> Bert Karwatzki
Sebastian
>
> > >
> > > Could the FPU regions with disabled preemption be limited to where we
> > > have actually have FPU usage in way that you don't have to worry when it
> > > is needed to use DC_RUN_WITH_PREEMPTION_ENABLED() and when not? Also the
> > > dc_fpu_begin()/ end() can nest and if they do the usage of
> > > DC_RUN_WITH_PREEMPTION_ENABLED() is futile, isn't it?
> > >
> >
> > I'm not very familiar with the amd display engine, and it's a lot of code,
> > but there I think there's some room for improvement, e.g. dml2_destroy():
> >
> > dml2_destroy is called from dc_state_free() from within DC_FP_{START,END}()
> > and then basically just calls vfree with DC_RUN_WITH_PREEMPTION_ENABLED().
> > (directly or through dml21_destroy()). Here both the DC_FP_*() tags and
> > DC_RUN_WITH_PREEMPTION_ENABLED() could be dropped I think (not tested, yet)
> >
Now, I've tested dropping some DC_FP_START,END() and DC_RUN_WITH_PREEMPTION_ENABLED().
(Patch for next-20260828+ with the fix above applied)
This works without error so far.
The function dml2_create() and dml2_create_copy() are moved to a non-FPU file so they
do not accidently use FPU instruction (they do not use FPU instruction on x86_64 when
compiled with gcc-16 but other architectures and compilers could probably use them).
diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_state.c b/drivers/gpu/drm/amd/display/dc/core/dc_state.c
index 666212cac105..709c37d2072f 100644
--- a/drivers/gpu/drm/amd/display/dc/core/dc_state.c
+++ b/drivers/gpu/drm/amd/display/dc/core/dc_state.c
@@ -209,9 +209,7 @@ struct dc_state *dc_state_create(struct dc *dc, struct dc_state_create_params *p
bool status;
if (dc->debug.using_dml2) {
- DC_FP_START();
status = dml2_create(dc, &dc->dml2_options, &state->bw_ctx.dml2);
- DC_FP_END();
if (!status) {
dc_state_release(state);
@@ -221,9 +219,7 @@ struct dc_state *dc_state_create(struct dc *dc, struct dc_state_create_params *p
if (dc->caps.dcmode_power_limits_present) {
bool dc_power_status;
- DC_FP_START();
dc_power_status = dml2_create(dc, &dc->dml2_dc_power_options, &state->bw_ctx.dml2_dc_power_source);
- DC_FP_END();
if (!dc_power_status) {
dc_state_release(state);
@@ -251,17 +247,13 @@ void dc_state_copy(struct dc_state *dst_state, struct dc_state *src_state)
#ifdef CONFIG_DRM_AMD_DC_FP
dst_state->bw_ctx.dml2 = dst_dml2;
if (src_state->bw_ctx.dml2) {
- DC_FP_START();
dml2_copy(dst_state->bw_ctx.dml2, src_state->bw_ctx.dml2);
- DC_FP_END();
}
dst_state->bw_ctx.dml2_dc_power_source = dst_dml2_dc_power_source;
if (src_state->bw_ctx.dml2_dc_power_source) {
- DC_FP_START();
dml2_copy(dst_state->bw_ctx.dml2_dc_power_source, src_state->bw_ctx.dml2_dc_power_source);
- DC_FP_END();
}
#endif // CONFIG_DRM_AMD_DC_FP
/* context refcount should not be overridden */
@@ -285,9 +277,7 @@ struct dc_state *dc_state_create_copy(struct dc_state *src_state)
new_state->bw_ctx.dml2_dc_power_source = NULL;
if (src_state->bw_ctx.dml2) {
- DC_FP_START();
status = dml2_create_copy(&new_state->bw_ctx.dml2, src_state->bw_ctx.dml2);
- DC_FP_END();
if (!status) {
dc_state_release(new_state);
@@ -297,10 +287,8 @@ struct dc_state *dc_state_create_copy(struct dc_state *src_state)
if (src_state->bw_ctx.dml2_dc_power_source) {
- DC_FP_START();
status = dml2_create_copy(&new_state->bw_ctx.dml2_dc_power_source,
src_state->bw_ctx.dml2_dc_power_source);
- DC_FP_END();
if (!status) {
dc_state_release(new_state);
@@ -389,13 +377,11 @@ static void dc_state_free(struct kref *kref)
dc_state_destruct(state);
#ifdef CONFIG_DRM_AMD_DC_FP
- DC_FP_START();
dml2_destroy(state->bw_ctx.dml2);
state->bw_ctx.dml2 = 0;
dml2_destroy(state->bw_ctx.dml2_dc_power_source);
state->bw_ctx.dml2_dc_power_source = 0;
- DC_FP_END();
#endif
kvfree(state);
diff --git a/drivers/gpu/drm/amd/display/dc/dml2_0/dml21/dml21_wrapper.c b/drivers/gpu/drm/amd/display/dc/dml2_0/dml21/dml21_wrapper.c
index 8bed59e976d1..29e5cce51b99 100644
--- a/drivers/gpu/drm/amd/display/dc/dml2_0/dml21/dml21_wrapper.c
+++ b/drivers/gpu/drm/amd/display/dc/dml2_0/dml21/dml21_wrapper.c
@@ -23,11 +23,11 @@
static bool dml21_allocate_memory(struct dml2_context **dml_ctx)
{
- DC_RUN_WITH_PREEMPTION_ENABLED(*dml_ctx = vzalloc(sizeof(struct dml2_context)));
+ *dml_ctx = vzalloc(sizeof(struct dml2_context));
if (!(*dml_ctx))
return false;
- DC_RUN_WITH_PREEMPTION_ENABLED((*dml_ctx)->v21.dml_init.dml2_instance = vzalloc(sizeof(struct dml2_instance)));
+ (*dml_ctx)->v21.dml_init.dml2_instance = vzalloc(sizeof(struct dml2_instance));
if (!((*dml_ctx)->v21.dml_init.dml2_instance))
return false;
@@ -37,7 +37,7 @@ static bool dml21_allocate_memory(struct dml2_context **dml_ctx)
(*dml_ctx)->v21.mode_support.display_config = &(*dml_ctx)->v21.display_config;
(*dml_ctx)->v21.mode_programming.display_config = (*dml_ctx)->v21.mode_support.display_config;
- DC_RUN_WITH_PREEMPTION_ENABLED((*dml_ctx)->v21.mode_programming.programming = vzalloc(sizeof(struct dml2_display_cfg_programming)));
+ (*dml_ctx)->v21.mode_programming.programming = vzalloc(sizeof(struct dml2_display_cfg_programming));
if (!((*dml_ctx)->v21.mode_programming.programming))
return false;
@@ -51,15 +51,17 @@ bool dml21_create(const struct dc *in_dc, struct dml2_context **dml_ctx, const s
if (!dml21_allocate_memory(dml_ctx))
return false;
+ DC_FP_START();
dml21_init(in_dc, *dml_ctx, config);
+ DC_FP_END();
return true;
}
void dml21_destroy(struct dml2_context *dml2)
{
- DC_RUN_WITH_PREEMPTION_ENABLED(vfree(dml2->v21.dml_init.dml2_instance));
- DC_RUN_WITH_PREEMPTION_ENABLED(vfree(dml2->v21.mode_programming.programming));
+ vfree(dml2->v21.dml_init.dml2_instance);
+ vfree(dml2->v21.mode_programming.programming);
}
void dml21_copy(struct dml2_context *dst_dml_ctx,
@@ -88,7 +90,9 @@ void dml21_copy(struct dml2_context *dst_dml_ctx,
dst_dml_ctx->v21.mode_programming.programming = dst_dml2_programming;
/* need to initialize copied instance for internal references to be correct */
+ DC_FP_START();
dml2_initialize_instance(&dst_dml_ctx->v21.dml_init);
+ DC_FP_END();
}
bool dml21_create_copy(struct dml2_context **dst_dml_ctx,
diff --git a/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper.c b/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper.c
index 1772e74349c7..570da14fb1a7 100644
--- a/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper.c
+++ b/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper.c
@@ -21,7 +21,7 @@ struct dml2_context *dml2_allocate_memory(void)
{
struct dml2_context *dml2;
- DC_RUN_WITH_PREEMPTION_ENABLED(dml2 = vzalloc(sizeof(struct dml2_context)));
+ dml2 = vzalloc(sizeof(struct dml2_context));
return dml2;
}
bool dml2_validate(const struct dc *in_dc, struct dc_state *context, struct dml2_context *dml2,
@@ -51,7 +51,9 @@ bool dml2_validate(const struct dc *in_dc, struct dc_state *context, struct dml2
static void dml2_init(const struct dc *in_dc, const struct dml2_configuration_options *config, struct dml2_context **dml2)
{
if ((in_dc->debug.using_dml21) && (in_dc->ctx->dce_version >= DCN_VERSION_4_01)) {
+ DC_FP_START();
dml21_reinit(in_dc, *dml2, config);
+ DC_FP_END();
return;
}
@@ -82,11 +84,13 @@ static void dml2_init(const struct dc *in_dc, const struct dml2_configuration_op
break;
}
+ DC_FP_START();
initialize_dml2_ip_params(*dml2, in_dc, &(*dml2)->v20.dml_core_ctx.ip);
initialize_dml2_soc_bbox(*dml2, in_dc, &(*dml2)->v20.dml_core_ctx.soc);
initialize_dml2_soc_states(*dml2, in_dc, &(*dml2)->v20.dml_core_ctx.soc, &(*dml2)->v20.dml_core_ctx.states);
+ DC_FP_END();
}
@@ -116,17 +120,52 @@ void dml2_destroy(struct dml2_context *dml2)
if (dml2->architecture == dml2_architecture_21)
dml21_destroy(dml2);
- DC_RUN_WITH_PREEMPTION_ENABLED(vfree(dml2));
+ vfree(dml2);
}
void dml2_reinit(const struct dc *in_dc,
const struct dml2_configuration_options *config,
struct dml2_context **dml2)
{
+ /*
if ((in_dc->debug.using_dml21) && (in_dc->ctx->dce_version >= DCN_VERSION_4_01)) {
+ DC_FP_START();
dml21_reinit(in_dc, *dml2, config);
+ DC_FP_END();
return;
}
+ */
dml2_init(in_dc, config, dml2);
}
+
+/* Moved here from dml2_wrapper_fpu.c */
+void dml2_copy(struct dml2_context *dst_dml2,
+ struct dml2_context *src_dml2)
+{
+ if (src_dml2->architecture == dml2_architecture_21) {
+ dml21_copy(dst_dml2, src_dml2);
+ return;
+ }
+ /* copy Mode Lib Ctx */
+ memcpy(dst_dml2, src_dml2, sizeof(struct dml2_context));
+}
+
+/* Moved here from dml2_wrapper_fpu.c */
+bool dml2_create_copy(struct dml2_context **dst_dml2,
+ struct dml2_context *src_dml2)
+{
+ if (src_dml2->architecture == dml2_architecture_21)
+ return dml21_create_copy(dst_dml2, src_dml2);
+ /* Allocate Mode Lib Ctx */
+ *dst_dml2 = dml2_allocate_memory();
+
+ if (!(*dst_dml2))
+ return false;
+
+ /* copy Mode Lib Ctx */
+ dml2_copy(*dst_dml2, src_dml2);
+
+ return true;
+}
+
diff --git a/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper_fpu.c b/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper_fpu.c
index a14e3004a7b7..b590d58ad3a9 100644
--- a/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper_fpu.c
+++ b/drivers/gpu/drm/amd/display/dc/dml2_0/dml2_wrapper_fpu.c
@@ -561,31 +561,3 @@ void dml2_prepare_mcache_programming(struct dc *in_dc, struct dc_state *context,
dml21_prepare_mcache_programming(in_dc, context, dml2);
}
-void dml2_copy(struct dml2_context *dst_dml2,
- struct dml2_context *src_dml2)
-{
- if (src_dml2->architecture == dml2_architecture_21) {
- dml21_copy(dst_dml2, src_dml2);
- return;
- }
- /* copy Mode Lib Ctx */
- memcpy(dst_dml2, src_dml2, sizeof(struct dml2_context));
-}
-
-bool dml2_create_copy(struct dml2_context **dst_dml2,
- struct dml2_context *src_dml2)
-{
- if (src_dml2->architecture == dml2_architecture_21)
- return dml21_create_copy(dst_dml2, src_dml2);
- /* Allocate Mode Lib Ctx */
- *dst_dml2 = dml2_allocate_memory();
-
- if (!(*dst_dml2))
- return false;
-
- /* copy Mode Lib Ctx */
- dml2_copy(*dst_dml2, src_dml2);
-
- return true;
-}
-
Bert Karwatzki
On Fri, Aug 07, 2026 at 02:49:42PM +0200, Bert Karwatzki wrote:
> On PREEMPT_RT kernels kvzalloc_obj() can sleep because spin_lock is
> converted to rt_mutex. dc_create_plane_state() can be called while
> inside an FPU-guarded region, resuling in "scheduling while atomic"
> errors on PREEMPT_RT kernels.
> Fix this by calling kvzalloc_obj() with DC_RUN_WITH_PREEMPTION_ENABLED().
> Also fix the error path in dc_create_stream_for_sink().
>
> Fixes: 3539437f354b ("drm/amd/display: Move FPU Guards From DML To DC - Part 1")
> Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> Signed-off-by: Bert Karwatzki <spasswolf@web.de>
> ---
> drivers/gpu/drm/amd/display/dc/core/dc_stream.c | 5 +++--
> drivers/gpu/drm/amd/display/dc/core/dc_surface.c | 5 +++--
> 2 files changed, 6 insertions(+), 4 deletions(-)
>
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> index 7666cdc78f4e..a5a304a3f802 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> @@ -233,8 +233,9 @@ struct dc_stream_state *dc_create_stream_for_sink(
>
> fail:
> if (stream) {
> - kfree(stream->update_scratch);
> - kfree(stream);
> + if (stream->update_scratch)
> + DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream->update_scratch));
> + DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));
> }
>
> return NULL;
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> index 72845fc788f3..04982673ffbc 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> @@ -33,6 +33,7 @@
> #include "dpp.h"
>
> #include "dc_plane_priv.h"
> +#include "dc_fpu.h"
>
> /*******************************************************************************
> * Private functions
> @@ -86,8 +87,8 @@ uint8_t dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
> ******************************************************************************/
> struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
> {
> - struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
> - GFP_ATOMIC);
> + struct dc_plane_state *plane_state;
> + DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
>
> if (NULL == plane_state)
> return NULL;
> --
> 2.55.0
>
> Fix for stable with the old error handling in
> dc_create_stream_for_sink() and compile fix form arm64/clang.
>
> Bert Karwatzki
>
<formletter>
This is not the correct way to submit patches for inclusion in the
stable kernel tree. Please read:
https://www.kernel.org/doc/html/latest/process/stable-kernel-rules.html
for how to do this properly.
</formletter>
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
Pre-existing issues:
- [High] Incomplete Fix: Normal deallocation paths (dc_stream_free and dc_plane_state_free) still sleep in atomic context.
- [High] DC_RUN_WITH_PREEMPTION_ENABLED fails to re-enable preemption if FPU blocks are nested.
- [High] DC_RUN_WITH_PREEMPTION_ENABLED corrupts callee-saved FPU registers on ARM64.
--
--- Patch [1]: [PATCH v7.2-rc6] drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT ---
commit e67f6942f0899807d82dcd2276705097e892a34b
Author: Bert Karwatzki <spasswolf@web.de>
drm/amd/display: fix usage of DC_FPU_{BEGIN,END} with PREEMPT_RT
This commit wraps deallocations in dc_create_stream_for_sink and
dc_create_plane_state with DC_RUN_WITH_PREEMPTION_ENABLED. This prevents
"scheduling while atomic" panics on PREEMPT_RT kernels when sleeping
allocations are called from within FPU-guarded regions.
Link: https://lore.kernel.org/lkml/20260723123449.6494-1-spasswolf@web.de/
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> index 7666cdc78f4ed..a5a304a3f8029 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_stream.c
[ ... ]
> @@ -233,8 +233,9 @@ struct dc_stream_state *dc_create_stream_for_sink(
>
> fail:
> if (stream) {
> - kfree(stream->update_scratch);
> - kfree(stream);
> + if (stream->update_scratch)
> + DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream->update_scratch));
> + DC_RUN_WITH_PREEMPTION_ENABLED(kfree(stream));
[Severity: High]
This isn't a bug introduced by this patch, but will the normal deallocation
paths still sleep in atomic context?
When dc_state_release_phantom_streams_and_planes() is called from
dcn32_internal_validate_bw(), it executes under DC_FP_START where
preemption is disabled. That validation path drops references to phantom
streams and planes, which synchronously invokes dc_stream_free() and
dc_plane_state_free().
Those destructors unconditionally call kfree() and kvfree(). On PREEMPT_RT,
slab allocators use sleepable rt_mutexes, which will still cause a kernel
panic from scheduling while atomic. Is it possible to apply this fix to
those deallocation paths as well?
> }
>
> return NULL;
> diff --git a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> index 72845fc788f3f..04982673ffbc9 100644
> --- a/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
> +++ b/drivers/gpu/drm/amd/display/dc/core/dc_surface.c
[ ... ]
> @@ -86,8 +87,8 @@ uint8_t dc_plane_get_pipe_mask(struct dc_state *dc_state, const struct dc_plane
> ******************************************************************************/
> struct dc_plane_state *dc_create_plane_state(const struct dc *dc)
> {
> - struct dc_plane_state *plane_state = kvzalloc_obj(*plane_state,
> - GFP_ATOMIC);
> + struct dc_plane_state *plane_state;
> + DC_RUN_WITH_PREEMPTION_ENABLED(plane_state = kvzalloc_obj(*plane_state, GFP_ATOMIC));
[Severity: High]
This isn't a bug introduced by this patch, but does
DC_RUN_WITH_PREEMPTION_ENABLED() fail to re-enable preemption if FPU blocks
are nested?
In dc_fpu_begin(), preempt_disable() is called unconditionally on every
entry, meaning preempt_count increases with each nesting level. However,
DC_RUN_WITH_PREEMPTION_ENABLED() calls DC_FP_END() exactly once, which
only decrements the recursion depth and calls preempt_enable() once.
If the FPU recursion depth is 2 or more, preemption remains disabled and
the wrapped kvzalloc_obj() will still crash on PREEMPT_RT.
[Severity: High]
This is a pre-existing issue, but can DC_RUN_WITH_PREEMPTION_ENABLED()
corrupt callee-saved FPU registers on ARM64 when used here?
A caller compiled with FPU support, such as
dml2_mall_phantom_streams_and_planes_create(), has live variables in
callee-saved FPU registers (v8-v15 on ARM64). It calls
dc_state_create_phantom_plane() which routes to dc_create_plane_state().
Since dc_create_plane_state() is compiled without FPU support, it doesn't
save v8-v15 to the stack. When DC_RUN_WITH_PREEMPTION_ENABLED() calls
DC_FP_END(), it invokes kernel_neon_end() and enables preemption.
If a context switch occurs while preemption is enabled, the OS scheduler
will not save the thread's FPU state since the kernel has yielded it. The
v8-v15 registers are permanently clobbered, and execution returns to the
caller with garbage FPU values.
>
> if (NULL == plane_state)
> return NULL;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260807124944.154232-1-spasswolf@web.de?part=1
© 2016 - 2026 Red Hat, Inc.