arch/s390/boot/ipl_parm.c | 19 ++++++++++++++++++- 1 file changed, 18 insertions(+), 1 deletion(-)
Users may accidentally add multi-byte UTF-8 characters to zipl.conf
parmline, for example, by copying snippets containing non-breaking
spaces (\xC2\xA0) from web pages.
The kernel will then interpret the entire command line as EBCDIC,
making it unusable. Distinguish this situation from the legitimate
EBCDIC conversion by looking for non-printable characters and issue
a warning.
Signed-off-by: Ilya Leoshkevich <iii@linux.ibm.com>
---
arch/s390/boot/ipl_parm.c | 19 ++++++++++++++++++-
1 file changed, 18 insertions(+), 1 deletion(-)
diff --git a/arch/s390/boot/ipl_parm.c b/arch/s390/boot/ipl_parm.c
index 6bc950b92be76..71c0c26d56bba 100644
--- a/arch/s390/boot/ipl_parm.c
+++ b/arch/s390/boot/ipl_parm.c
@@ -173,12 +173,29 @@ static inline int has_ebcdic_char(const char *str)
return 0;
}
+static inline int has_nonprintable_char(const char *str)
+{
+ int i;
+
+ for (i = 0; str[i]; i++) {
+ unsigned char c = (unsigned char)str[i];
+
+ /* isprint() is Latin-1, and we need ASCII here */
+ if (c < 0x20 || c > 0x7e)
+ return 1;
+ }
+ return 0;
+}
+
void setup_boot_command_line(void)
{
parmarea.command_line[COMMAND_LINE_SIZE - 1] = 0;
/* convert arch command line to ascii if necessary */
- if (has_ebcdic_char(parmarea.command_line))
+ if (has_ebcdic_char(parmarea.command_line)) {
EBCASC(parmarea.command_line, COMMAND_LINE_SIZE);
+ if (has_nonprintable_char(parmarea.command_line))
+ boot_warn("Kernel command line was treated as EBCDIC, but contains non-printable characters\n");
+ }
/* copy arch command line */
strscpy(early_command_line, strim(parmarea.command_line));
--
2.55.0
On Tue, 25 Aug 2026 17:08:08 +0200
Ilya Leoshkevich <iii@linux.ibm.com> wrote:
> Users may accidentally add multi-byte UTF-8 characters to zipl.conf
> parmline, for example, by copying snippets containing non-breaking
> spaces (\xC2\xA0) from web pages.
>
> The kernel will then interpret the entire command line as EBCDIC,
> making it unusable. Distinguish this situation from the legitimate
> EBCDIC conversion by looking for non-printable characters and issue
> a warning.
Would it be better to check for the entire line being printable ebcdic?
All of EBCDIC a-zA-Z0-9 have the 0x80 bit set and most of 0x20..0x7f
are invalid or control characters (or punctuation).
David
>
> Signed-off-by: Ilya Leoshkevich <iii@linux.ibm.com>
> ---
> arch/s390/boot/ipl_parm.c | 19 ++++++++++++++++++-
> 1 file changed, 18 insertions(+), 1 deletion(-)
>
> diff --git a/arch/s390/boot/ipl_parm.c b/arch/s390/boot/ipl_parm.c
> index 6bc950b92be76..71c0c26d56bba 100644
> --- a/arch/s390/boot/ipl_parm.c
> +++ b/arch/s390/boot/ipl_parm.c
> @@ -173,12 +173,29 @@ static inline int has_ebcdic_char(const char *str)
> return 0;
> }
>
> +static inline int has_nonprintable_char(const char *str)
> +{
> + int i;
> +
> + for (i = 0; str[i]; i++) {
> + unsigned char c = (unsigned char)str[i];
> +
> + /* isprint() is Latin-1, and we need ASCII here */
> + if (c < 0x20 || c > 0x7e)
> + return 1;
> + }
> + return 0;
> +}
> +
> void setup_boot_command_line(void)
> {
> parmarea.command_line[COMMAND_LINE_SIZE - 1] = 0;
> /* convert arch command line to ascii if necessary */
> - if (has_ebcdic_char(parmarea.command_line))
> + if (has_ebcdic_char(parmarea.command_line)) {
> EBCASC(parmarea.command_line, COMMAND_LINE_SIZE);
> + if (has_nonprintable_char(parmarea.command_line))
> + boot_warn("Kernel command line was treated as EBCDIC, but contains non-printable characters\n");
> + }
> /* copy arch command line */
> strscpy(early_command_line, strim(parmarea.command_line));
>
On 8/26/26 15:51, David Laight wrote: > On Tue, 25 Aug 2026 17:08:08 +0200 > Ilya Leoshkevich <iii@linux.ibm.com> wrote: > >> Users may accidentally add multi-byte UTF-8 characters to zipl.conf >> parmline, for example, by copying snippets containing non-breaking >> spaces (\xC2\xA0) from web pages. >> >> The kernel will then interpret the entire command line as EBCDIC, >> making it unusable. Distinguish this situation from the legitimate >> EBCDIC conversion by looking for non-printable characters and issue >> a warning. > > Would it be better to check for the entire line being printable ebcdic? > All of EBCDIC a-zA-Z0-9 have the 0x80 bit set and most of 0x20..0x7f > are invalid or control characters (or punctuation). > > David I actually started with that, but this required introducing a new _ctype-like table (unfortunately it's not as simple as checking a couple ranges), so I decided against that and took a shortcut via ASCII. [...]
On Wed, 26 Aug 2026 16:08:50 +0200 Ilya Leoshkevich <iii@linux.ibm.com> wrote: > On 8/26/26 15:51, David Laight wrote: > > On Tue, 25 Aug 2026 17:08:08 +0200 > > Ilya Leoshkevich <iii@linux.ibm.com> wrote: > > > >> Users may accidentally add multi-byte UTF-8 characters to zipl.conf > >> parmline, for example, by copying snippets containing non-breaking > >> spaces (\xC2\xA0) from web pages. > >> > >> The kernel will then interpret the entire command line as EBCDIC, > >> making it unusable. Distinguish this situation from the legitimate > >> EBCDIC conversion by looking for non-printable characters and issue > >> a warning. > > > > Would it be better to check for the entire line being printable ebcdic? > > All of EBCDIC a-zA-Z0-9 have the 0x80 bit set and most of 0x20..0x7f > > are invalid or control characters (or punctuation). > > > > David > > I actually started with that, but this required introducing a new > _ctype-like table (unfortunately it's not as simple as checking a > couple ranges), so I decided against that and took a shortcut via > ASCII. Could you get the conversion function to return an error if it found invalid EBCDIC characters? If there is a single UTF8 character (eg non-breaking space) you really want to treat the line as ASCII. Actually you could count the number of characters with the 0x80 bit set. If more than 1/2 assume EBCDIC (all of 0-9a-zA-Z have the bit set). (I didn't realise anyone still used EBCDIC. I guess the unix implementation(s) use ASCII (otherwise too much code is broken) but the old IBM OS uses EBCDIC. I worked for ICL for a while, their old 1900 series (from the early 1970s) used 6bit characters (4 in a 24bit word) that were ACSII codes 32-95. The replacement 2900 series (very late 1970s) used EBCDIC internally (I guess because IBM used it...) but all the peripherals were ASCII.) David > > [...]
On 8/27/26 10:32, David Laight wrote: > On Wed, 26 Aug 2026 16:08:50 +0200 > Ilya Leoshkevich <iii@linux.ibm.com> wrote: > >> On 8/26/26 15:51, David Laight wrote: >>> On Tue, 25 Aug 2026 17:08:08 +0200 >>> Ilya Leoshkevich <iii@linux.ibm.com> wrote: >>> >>>> Users may accidentally add multi-byte UTF-8 characters to zipl.conf >>>> parmline, for example, by copying snippets containing non-breaking >>>> spaces (\xC2\xA0) from web pages. >>>> >>>> The kernel will then interpret the entire command line as EBCDIC, >>>> making it unusable. Distinguish this situation from the legitimate >>>> EBCDIC conversion by looking for non-printable characters and issue >>>> a warning. >>> >>> Would it be better to check for the entire line being printable ebcdic? >>> All of EBCDIC a-zA-Z0-9 have the 0x80 bit set and most of 0x20..0x7f >>> are invalid or control characters (or punctuation). >>> >>> David >> >> I actually started with that, but this required introducing a new >> _ctype-like table (unfortunately it's not as simple as checking a >> couple ranges), so I decided against that and took a shortcut via >> ASCII. > > Could you get the conversion function to return an error if it found > invalid EBCDIC characters? > If there is a single UTF8 character (eg non-breaking space) you really > want to treat the line as ASCII. > Actually you could count the number of characters with the 0x80 bit set. > If more than 1/2 assume EBCDIC (all of 0-9a-zA-Z have the bit set). I also considered that, but setting any threshold feels arbitrary and will probably fail for punctuation-heavy command lines. I also didn't want to make it a hard fail, because I'm not certain that I know all uses cases. Perhaps there are people who wants umlauts and what not? It would be bad to break whatever they are doing. But at the same time it was very painful to debug the issue, so I settled for the compromise: add a warning that will be helpful to 99.9% users and will only mildly annoy the 0.1% umlaut users. I just had an off-list discussion with Heiko and we think about going with your first proposal for v3: a new _ctype table for EBCDIC for determining whether characters are printable. The overhead from this is not bad as I thought it would be. > (I didn't realise anyone still used EBCDIC. > I guess the unix implementation(s) use ASCII (otherwise too much code > is broken) but the old IBM OS uses EBCDIC. > I worked for ICL for a while, their old 1900 series (from the early > 1970s) used 6bit characters (4 in a 24bit word) that were ACSII codes > 32-95. The replacement 2900 series (very late 1970s) used EBCDIC internally > (I guess because IBM used it...) but all the peripherals were ASCII.) Linux on s390 still uses it for interfacing with traditional IBM hypervisors (z/VM and PR/SM), which are very much alive and used today. > David > >> >> [...] >
On Tue, Aug 25, 2026 at 05:08:08PM +0200, Ilya Leoshkevich wrote:
> Users may accidentally add multi-byte UTF-8 characters to zipl.conf
> parmline, for example, by copying snippets containing non-breaking
> spaces (\xC2\xA0) from web pages.
>
> The kernel will then interpret the entire command line as EBCDIC,
> making it unusable. Distinguish this situation from the legitimate
> EBCDIC conversion by looking for non-printable characters and issue
> a warning.
>
> Signed-off-by: Ilya Leoshkevich <iii@linux.ibm.com>
> ---
> arch/s390/boot/ipl_parm.c | 19 ++++++++++++++++++-
> 1 file changed, 18 insertions(+), 1 deletion(-)
>
> diff --git a/arch/s390/boot/ipl_parm.c b/arch/s390/boot/ipl_parm.c
> index 6bc950b92be76..71c0c26d56bba 100644
> --- a/arch/s390/boot/ipl_parm.c
> +++ b/arch/s390/boot/ipl_parm.c
> @@ -173,12 +173,29 @@ static inline int has_ebcdic_char(const char *str)
> return 0;
> }
>
> +static inline int has_nonprintable_char(const char *str)
> +{
> + int i;
> +
> + for (i = 0; str[i]; i++) {
> + unsigned char c = (unsigned char)str[i];
> +
> + /* isprint() is Latin-1, and we need ASCII here */
> + if (c < 0x20 || c > 0x7e)
> + return 1;
Hm, I guess the comment refers to a different implementation than the
kernel internal one? Since isprint() (see include/linux/ctype.h) is
true for exactly the range you open-coded, as far as I can tell.
Furthermore kernel command line parsing also allows for all sorts of
spaces, tabs, and line feeds (see e.g. next_arg()). So I guess the
above should be changed (and shortened :) ) to something like:
static inline int has_nonprintable_char(const char *str)
{
for (int i = 0; str[i]; i++) {
if (isprint(str[i]) || isspace(str[i]))
return 1;
}
return 0;
}
On 8/26/26 11:29, Heiko Carstens wrote:
> On Tue, Aug 25, 2026 at 05:08:08PM +0200, Ilya Leoshkevich wrote:
>> Users may accidentally add multi-byte UTF-8 characters to zipl.conf
>> parmline, for example, by copying snippets containing non-breaking
>> spaces (\xC2\xA0) from web pages.
>>
>> The kernel will then interpret the entire command line as EBCDIC,
>> making it unusable. Distinguish this situation from the legitimate
>> EBCDIC conversion by looking for non-printable characters and issue
>> a warning.
>>
>> Signed-off-by: Ilya Leoshkevich <iii@linux.ibm.com>
>> ---
>> arch/s390/boot/ipl_parm.c | 19 ++++++++++++++++++-
>> 1 file changed, 18 insertions(+), 1 deletion(-)
>>
>> diff --git a/arch/s390/boot/ipl_parm.c b/arch/s390/boot/ipl_parm.c
>> index 6bc950b92be76..71c0c26d56bba 100644
>> --- a/arch/s390/boot/ipl_parm.c
>> +++ b/arch/s390/boot/ipl_parm.c
>> @@ -173,12 +173,29 @@ static inline int has_ebcdic_char(const char *str)
>> return 0;
>> }
>>
>> +static inline int has_nonprintable_char(const char *str)
>> +{
>> + int i;
>> +
>> + for (i = 0; str[i]; i++) {
>> + unsigned char c = (unsigned char)str[i];
>> +
>> + /* isprint() is Latin-1, and we need ASCII here */
>> + if (c < 0x20 || c > 0x7e)
>> + return 1;
>
> Hm, I guess the comment refers to a different implementation than the
> kernel internal one? Since isprint() (see include/linux/ctype.h) is
> true for exactly the range you open-coded, as far as I can tell.
Unfortunately isprint() matches some extra ASCII upper-half characters,
e.g.:
const unsigned char _ctype[] = {
[...]
_P,_P,_P,_P,_P,_P,_P,_P,_P,_P,_P,_P,_P,_P,_P,_P, /* 176-191 */
> Furthermore kernel command line parsing also allows for all sorts of
> spaces, tabs, and line feeds (see e.g. next_arg()). So I guess the
> above should be changed (and shortened :) ) to something like:
I completely forgot about newlines, thanks!
> static inline int has_nonprintable_char(const char *str)
> {
> for (int i = 0; str[i]; i++) {
> if (isprint(str[i]) || isspace(str[i]))
> return 1;
> }
> return 0;
> }
© 2016 - 2026 Red Hat, Inc.