.gitignore | 1 + Documentation/kbuild/kbuild.rst | 13 + Documentation/kbuild/reproducible-builds.rst | 16 + Kbuild | 5 + Makefile | 49 +- arch/arm64/kernel/pi/Makefile | 2 +- arch/riscv/kernel/pi/Makefile | 2 +- arch/x86/Kconfig | 17 + arch/x86/Makefile | 12 +- arch/x86/boot/Makefile | 2 +- arch/x86/boot/compressed/Makefile | 2 +- drivers/firmware/efi/libstub/Makefile | 2 +- include/asm-generic/vmlinux.lds.h | 2 +- include/linux/vermagic.h | 2 +- init/Kconfig | 84 +++ rust/Makefile | 5 + scripts/Makefile | 4 +- scripts/Makefile.build | 27 +- scripts/Makefile.lib | 7 +- scripts/Makefile.modfinal | 36 +- scripts/Makefile.modpost | 2 +- scripts/Makefile.vmlinux | 27 +- scripts/Makefile.warn | 28 +- scripts/basic/.gitignore | 1 + scripts/basic/Makefile | 2 +- scripts/basic/depcheck.c | 442 +++++++++++++++ scripts/check-function-names.sh | 3 +- scripts/elf-parse.c | 48 +- scripts/elf-parse.h | 19 + scripts/kallsyms-sysmap.c | 269 ++++++++++ scripts/kallsyms.c | 387 ++++++++++---- scripts/kallsyms.h | 44 ++ scripts/link-vmlinux.sh | 25 +- scripts/mksysmap | 94 ---- scripts/mod/.gitignore | 1 + scripts/mod/Makefile | 9 + scripts/mod/modpost.c | 757 ++++++++++++++++++++------ scripts/mod/modpost.h | 2 + scripts/mod/module-offsets.c | 35 ++ scripts/mod/sumversion.c | 57 +- scripts/tags.sh | 5 +- tools/objtool/Makefile | 2 +- tools/objtool/check.c | 767 ++++++++++++++++++++------- tools/objtool/elf.c | 242 ++++++++- tools/objtool/include/objtool/elf.h | 4 + tools/objtool/include/objtool/objtool.h | 3 +- tools/objtool/objtool.c | 1 - 47 files changed, 2884 insertions(+), 682 deletions(-)
A typical kernel build consists of a frustratingly large amount of time
spent stuck in single-threaded bottlenecks.
It turns out that there's a lot we can do about this and doing so
significantly impacts kernel build times.
This series makes allmodconfig builds up to 36% faster, incremental builds
up to ~70% faster, and noop builds up to ~90% faster.
Builds are faster across the board on every device I tested.
Machines used for perf testing:
* Threadripper - x86, AMD Threadripper 9980X, 64 cores, 128 threads
* EPYC - x86, 2 socket EPYC 9754, 256 cores, 512 threads
* M2 - arm64, 2022 M2 macbook pro, 8 cores, 8 threads
(4 perf, 4 efficiency)
Cutting to the chase:
== allmodconfig FULL build ==
before after delta
----------------------------------
Threadripper, gcc 345.3s 275.2s -70.1s (-20%)
Threadripper, clang 345.5s 265.7s -79.8s (-23%)
EPYC, gcc 188.6s 121.1s -67.5s (-36%)
EPYC, clang 261.7s 184.6s -77.1s (-29%)
== allmodconfig INCREMENTAL build ==
before after delta
----------------------------------
Threadripper, gcc 46.4s 15.3s -31.1s (-67%)
Threadripper, clang 44.4s 15.2s -29.1s (-66%)
EPYC, gcc 82.0s 24.3s -57.7s (-70%)
EPYC, clang 77.6s 24.3s -53.3s (-69%)
== allmodconfig NO-OP build ==
before after delta
----------------------------------
Threadripper, gcc 12.97s 1.18s -11.79s (-91%)
Threadripper, clang 13.84s 1.64s -12.20s (-88%)
EPYC, gcc 25.92s 1.52s -24.40s (-94%)
EPYC, clang 27.58s 2.17s -25.41s (-92%)
== defconfig FULL build ==
before after delta
----------------------------------
Threadripper, gcc 35.3s 28.6s -6.7s (-19%)
Threadripper, clang 33.5s 26.1s -7.4s (-22%)
EPYC, gcc 31.2s 20.6s -10.6s (-34%)
EPYC, clang 38.6s 32.2s -6.4s (-17%)
M2, gcc 564.0s 512.4s -51.6s (-9%)
M2, clang 616.1s 569.4s -46.7s (-8%)
== defconfig INCREMENTAL build ==
before after delta
----------------------------------
Threadripper, gcc 11.6s 5.5s -6.1s (-53%)
Threadripper, clang 11.6s 5.0s -6.6s (-57%)
EPYC, gcc 19.3s 9.2s -10.1s (-52%)
EPYC, clang 20.1s 8.6s -11.5s (-57%)
M2, gcc 18.6s 9.9s -8.7s (-47%)
M2, clang 18.6s 8.2s -10.4s (-56%)
== defconfig NO-OP build ==
before after delta
----------------------------------
Threadripper, gcc 0.96s 0.47s -0.49s (-51%)
Threadripper, clang 1.21s 0.55s -0.66s (-55%)
EPYC, gcc 1.47s 0.66s -0.81s (-55%)
EPYC, clang 1.87s 0.80s -1.07s (-57%)
M2, gcc 5.59s 1.69s -3.90s (-70%)
M2, clang 6.62s 1.74s -4.88s (-74%)
Further performance numbers are provided for each commit giving a sense of
what each contributes to the final result.
== Testing ==
Beyond x86, allmodconfig was built with the series for arm64, arm, riscv,
powerpc64, s390 and loongarch, and for arm64, arm, s390 and loongarch the
System.map is identical to what the previous shell mksysmap produces for
the same vmlinux.
Also tested were parisc64 and m68k build as far as mainline lets them
(a driver's static assertion and an undefined-symbol check on parisc, gcc
16 internal compiler errors on m68k, none of it from this series).
Kernels for x86, arm64, arm, riscv, loongarch, powerpc64, s390, m68k and
parisc64 all boot under qemu, every text symbol of System.map is in
/proc/kallsyms at the relocated address, and a module loads and unloads.
The s390, arm and loongarch kernels also pass the kallsyms selftest.
An x86 kernel with CONFIG_MODVERSIONS, CONFIG_EXTENDED_MODVERSIONS and
CONFIG_MODULE_SRCVERSION_ALL boots, loads and unloads modules. External
modules build against both in-tree and O= builds, and every commit builds
on x86 defconfig.
Build times are the best of several runs, no unexpected errors or
warnings were seen.
While some aspects of the build process have been changed, all tooling
should function identically to before.
== LLM usage ==
An LLM was used to first determine where the bottlenecks were then to
figure out how to improve them.
It generated a lot of code, much of it hideous.
I extensively audited and rewrote a lot of it, and heavily edited commit
messages, the cover letter and comments.x
The LLM has also orchestrated build runs, testing, debugging and analysis.
I have manually checked for correctness in both build and running kernels
generated with this series applied.
Performance improvements were also verified manually.
Since an LLM was used extensively, each commit carries an Assisted-by tag.
== What was changed? ==
Fundamentally the series improves build times by parallelising
single-threaded tasks as much as possible and improving the efficiency of
code used in the build process.
kbuild, kallsyms, modpost, objtool, mksysmap and the rust build system were
all updated as part of this change.
Nothing too controversial was included. There are further improvements that
could be made, but they would either by very invasive (large scale C header
changes) or generate diminishing returns.
== Patches ==
mksysmap (1, 2): A couple of bugfixes the rest of the series relies
upon.
kallsyms (3, 4): Some efficiency improvements through use of a cache.
kbuild (5, 6): Don't sort nm output unnecessarily, do not include
relocations in the kallsyms trial links.
elf-parse (7): Section flags, symbol binding and a read-only mapping,
for the next patch.
kallsyms (8): Don't use nm, read the ELF symbol table directly.
kbuild (9): Make .modinfo an INFO section.
kbuild (10, 11): Implement a cache to track objects, check dependency
timestamps more efficiently.
kbuild (12): Probe compiler, linker flags once at top of build.
modpost (13, 14): Hash module sources a file at a time not a byte at a
time, also improve performance through use of a
cache.
modules (15, 16): Emit module descriptors as assembly (*.mod.S) rather
than C (*.mod.c), reducing CPU seconds taken by 10x to
perform the task. Also shard module finalisation
rather than running 10's of thousands of tiny runs.
modpost (17): Compute srcversions in parallel using a thread pool.
objtool (18, 19): Avoid polluting the reloc hash with millions of DWARF
relocations. Remember dead-ends and decode large
objects over multiple threads. Size instruction hash
to the code rather than hard-code it.
rust (20-22): Parallelise the frontend, set correct dependencies and
build crates in parallel with C code, retaining the
requirement that rust/ crates are built first.
kbuild (23): Default to using a parallel implementation of gzip
(pigz) if available on the system.
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
---
Lorenzo Stoakes (ARM) (23):
scripts/mksysmap: drop the MODULE_INFO() symbols from kallsyms
scripts/mksysmap: fix escape of '$' in the __pi_ pattern
kallsyms: index symbols by token to speed up table compression
kallsyms: output binary data to speed output and kallsyms assembly
kbuild: do not sort nm output where the order is irrelevant
kbuild: only emit vmlinux relocations when required
elf-parse: add section flags, symbol binding and a read-only mapping
kallsyms: reimplement mksysmap in C
kbuild: do not allocate .modinfo in vmlinux
kbuild: cache list, composite object state per object
kbuild: implement and use depcheck to check dependency timestamps
kbuild: avoid re-running compiler and linker probes
modpost: hash module source per-file, not per-byte
modpost: cache section relocation mismatch state
modpost: emit module descriptors as assembly
kbuild: batch module finalisation
modpost: perform srcversion hashing in parallel
objtool: cache relocations and function dead end state, do less work
objtool: decode instructions and resolve branch targets in parallel
kbuild: rust: parallelise rustc front end
rust: make exports.o depend on the headers generated for it
kbuild: build rust crates in parallel with the rest of the build
kbuild: use pigz for gzip compression if available
.gitignore | 1 +
Documentation/kbuild/kbuild.rst | 13 +
Documentation/kbuild/reproducible-builds.rst | 16 +
Kbuild | 5 +
Makefile | 49 +-
arch/arm64/kernel/pi/Makefile | 2 +-
arch/riscv/kernel/pi/Makefile | 2 +-
arch/x86/Kconfig | 17 +
arch/x86/Makefile | 12 +-
arch/x86/boot/Makefile | 2 +-
arch/x86/boot/compressed/Makefile | 2 +-
drivers/firmware/efi/libstub/Makefile | 2 +-
include/asm-generic/vmlinux.lds.h | 2 +-
include/linux/vermagic.h | 2 +-
init/Kconfig | 84 +++
rust/Makefile | 5 +
scripts/Makefile | 4 +-
scripts/Makefile.build | 27 +-
scripts/Makefile.lib | 7 +-
scripts/Makefile.modfinal | 36 +-
scripts/Makefile.modpost | 2 +-
scripts/Makefile.vmlinux | 27 +-
scripts/Makefile.warn | 28 +-
scripts/basic/.gitignore | 1 +
scripts/basic/Makefile | 2 +-
scripts/basic/depcheck.c | 442 +++++++++++++++
scripts/check-function-names.sh | 3 +-
scripts/elf-parse.c | 48 +-
scripts/elf-parse.h | 19 +
scripts/kallsyms-sysmap.c | 269 ++++++++++
scripts/kallsyms.c | 387 ++++++++++----
scripts/kallsyms.h | 44 ++
scripts/link-vmlinux.sh | 25 +-
scripts/mksysmap | 94 ----
scripts/mod/.gitignore | 1 +
scripts/mod/Makefile | 9 +
scripts/mod/modpost.c | 757 ++++++++++++++++++++------
scripts/mod/modpost.h | 2 +
scripts/mod/module-offsets.c | 35 ++
scripts/mod/sumversion.c | 57 +-
scripts/tags.sh | 5 +-
tools/objtool/Makefile | 2 +-
tools/objtool/check.c | 767 ++++++++++++++++++++-------
tools/objtool/elf.c | 242 ++++++++-
tools/objtool/include/objtool/elf.h | 4 +
tools/objtool/include/objtool/objtool.h | 3 +-
tools/objtool/objtool.c | 1 -
47 files changed, 2884 insertions(+), 682 deletions(-)
---
base-commit: 28924df2a08f440c73991b83028032c901de2ae4
change-id: 20260904-build-speedup-25e11a00b3d0
Best regards,
--
Lorenzo Stoakes (ARM) <ljs@kernel.org>
> A typical kernel build consists of a frustratingly large amount of time > spent stuck in single-threaded bottlenecks. > > It turns out that there's a lot we can do about this and doing so > significantly impacts kernel build times. > > This series makes allmodconfig builds up to 36% faster, incremental builds > up to ~70% faster, and noop builds up to ~90% faster. > > Builds are faster across the board on every device I tested. > > Machines used for perf testing: > > * Threadripper - x86, AMD Threadripper 9980X, 64 cores, 128 threads > * EPYC - x86, 2 socket EPYC 9754, 256 cores, 512 threads > * M2 - arm64, 2022 M2 macbook pro, 8 cores, 8 threads > (4 perf, 4 efficiency) > > Cutting to the chase: > > == allmodconfig FULL build == > > before after delta > ---------------------------------- > Threadripper, gcc 345.3s 275.2s -70.1s (-20%) > Threadripper, clang 345.5s 265.7s -79.8s (-23%) > EPYC, gcc 188.6s 121.1s -67.5s (-36%) > EPYC, clang 261.7s 184.6s -77.1s (-29%) > > == allmodconfig INCREMENTAL build == > > before after delta > ---------------------------------- > Threadripper, gcc 46.4s 15.3s -31.1s (-67%) > Threadripper, clang 44.4s 15.2s -29.1s (-66%) > EPYC, gcc 82.0s 24.3s -57.7s (-70%) > EPYC, clang 77.6s 24.3s -53.3s (-69%) > > == allmodconfig NO-OP build == > > before after delta > ---------------------------------- > Threadripper, gcc 12.97s 1.18s -11.79s (-91%) > Threadripper, clang 13.84s 1.64s -12.20s (-88%) > EPYC, gcc 25.92s 1.52s -24.40s (-94%) > EPYC, clang 27.58s 2.17s -25.41s (-92%) > > == defconfig FULL build == > > before after delta > ---------------------------------- > Threadripper, gcc 35.3s 28.6s -6.7s (-19%) > Threadripper, clang 33.5s 26.1s -7.4s (-22%) > EPYC, gcc 31.2s 20.6s -10.6s (-34%) > EPYC, clang 38.6s 32.2s -6.4s (-17%) > M2, gcc 564.0s 512.4s -51.6s (-9%) > M2, clang 616.1s 569.4s -46.7s (-8%) > > == defconfig INCREMENTAL build == > > before after delta > ---------------------------------- > Threadripper, gcc 11.6s 5.5s -6.1s (-53%) > Threadripper, clang 11.6s 5.0s -6.6s (-57%) > EPYC, gcc 19.3s 9.2s -10.1s (-52%) > EPYC, clang 20.1s 8.6s -11.5s (-57%) > M2, gcc 18.6s 9.9s -8.7s (-47%) > M2, clang 18.6s 8.2s -10.4s (-56%) > > == defconfig NO-OP build == > > before after delta > ---------------------------------- > Threadripper, gcc 0.96s 0.47s -0.49s (-51%) > Threadripper, clang 1.21s 0.55s -0.66s (-55%) > EPYC, gcc 1.47s 0.66s -0.81s (-55%) > EPYC, clang 1.87s 0.80s -1.07s (-57%) > M2, gcc 5.59s 1.69s -3.90s (-70%) > M2, clang 6.62s 1.74s -4.88s (-74%) > > Further performance numbers are provided for each commit giving a sense of > what each contributes to the final result. > > == Testing == > > Beyond x86, allmodconfig was built with the series for arm64, arm, riscv, > powerpc64, s390 and loongarch, and for arm64, arm, s390 and loongarch the > System.map is identical to what the previous shell mksysmap produces for > the same vmlinux. > > Also tested were parisc64 and m68k build as far as mainline lets them > (a driver's static assertion and an undefined-symbol check on parisc, gcc > 16 internal compiler errors on m68k, none of it from this series). > > Kernels for x86, arm64, arm, riscv, loongarch, powerpc64, s390, m68k and > parisc64 all boot under qemu, every text symbol of System.map is in > /proc/kallsyms at the relocated address, and a module loads and unloads. > The s390, arm and loongarch kernels also pass the kallsyms selftest. > > An x86 kernel with CONFIG_MODVERSIONS, CONFIG_EXTENDED_MODVERSIONS and > CONFIG_MODULE_SRCVERSION_ALL boots, loads and unloads modules. External > modules build against both in-tree and O= builds, and every commit builds > on x86 defconfig. > > Build times are the best of several runs, no unexpected errors or > warnings were seen. > > While some aspects of the build process have been changed, all tooling > should function identically to before. I appreciate all of the testing that you have done! I will test this side by side on a couple of my own machines to see what results I get but I do like those results. > == LLM usage == > > An LLM was used to first determine where the bottlenecks were then to > figure out how to improve them. > > It generated a lot of code, much of it hideous. > > I extensively audited and rewrote a lot of it, and heavily edited commit > messages, the cover letter and comments.x > > The LLM has also orchestrated build runs, testing, debugging and analysis. > > I have manually checked for correctness in both build and running kernels > generated with this series applied. > > Performance improvements were also verified manually. > > Since an LLM was used extensively, each commit carries an Assisted-by tag. Thank you for calling all of this out as well. > == What was changed? == > > Fundamentally the series improves build times by parallelising > single-threaded tasks as much as possible and improving the efficiency of > code used in the build process. > > kbuild, kallsyms, modpost, objtool, mksysmap and the rust build system were > all updated as part of this change. > > Nothing too controversial was included. There are further improvements that > could be made, but they would either by very invasive (large scale C header > changes) or generate diminishing returns. > > == Patches == > > mksysmap (1, 2): A couple of bugfixes the rest of the series relies > upon. As I note in other threads of this review, I think we should take these to Linus now, they seem Obviously CorrectTM and it helps chunk out the series. I have just reviewed a few of the low hanging fruit patches. I will try to take a look at the rest over the next couple of weeks but I might not get to them until after Plumbers. I would greatly appreciate if there are others who are more knowledgeable in these areas who could review these changes, as I don't know too much about some of the corners of Kbuild yet. -- Cheers, Nathan
On Wed, Sep 09, 2026 at 09:19:12PM -0700, Nathan Chancellor wrote: > > A typical kernel build consists of a frustratingly large amount of time > > spent stuck in single-threaded bottlenecks. > > > > It turns out that there's a lot we can do about this and doing so > > significantly impacts kernel build times. > > > > This series makes allmodconfig builds up to 36% faster, incremental builds > > up to ~70% faster, and noop builds up to ~90% faster. > > > > Builds are faster across the board on every device I tested. > > > > Machines used for perf testing: > > > > * Threadripper - x86, AMD Threadripper 9980X, 64 cores, 128 threads > > * EPYC - x86, 2 socket EPYC 9754, 256 cores, 512 threads > > * M2 - arm64, 2022 M2 macbook pro, 8 cores, 8 threads > > (4 perf, 4 efficiency) > > > > Cutting to the chase: > > > > == allmodconfig FULL build == > > > > before after delta > > ---------------------------------- > > Threadripper, gcc 345.3s 275.2s -70.1s (-20%) > > Threadripper, clang 345.5s 265.7s -79.8s (-23%) > > EPYC, gcc 188.6s 121.1s -67.5s (-36%) > > EPYC, clang 261.7s 184.6s -77.1s (-29%) > > > > == allmodconfig INCREMENTAL build == > > > > before after delta > > ---------------------------------- > > Threadripper, gcc 46.4s 15.3s -31.1s (-67%) > > Threadripper, clang 44.4s 15.2s -29.1s (-66%) > > EPYC, gcc 82.0s 24.3s -57.7s (-70%) > > EPYC, clang 77.6s 24.3s -53.3s (-69%) > > > > == allmodconfig NO-OP build == > > > > before after delta > > ---------------------------------- > > Threadripper, gcc 12.97s 1.18s -11.79s (-91%) > > Threadripper, clang 13.84s 1.64s -12.20s (-88%) > > EPYC, gcc 25.92s 1.52s -24.40s (-94%) > > EPYC, clang 27.58s 2.17s -25.41s (-92%) > > > > == defconfig FULL build == > > > > before after delta > > ---------------------------------- > > Threadripper, gcc 35.3s 28.6s -6.7s (-19%) > > Threadripper, clang 33.5s 26.1s -7.4s (-22%) > > EPYC, gcc 31.2s 20.6s -10.6s (-34%) > > EPYC, clang 38.6s 32.2s -6.4s (-17%) > > M2, gcc 564.0s 512.4s -51.6s (-9%) > > M2, clang 616.1s 569.4s -46.7s (-8%) > > > > == defconfig INCREMENTAL build == > > > > before after delta > > ---------------------------------- > > Threadripper, gcc 11.6s 5.5s -6.1s (-53%) > > Threadripper, clang 11.6s 5.0s -6.6s (-57%) > > EPYC, gcc 19.3s 9.2s -10.1s (-52%) > > EPYC, clang 20.1s 8.6s -11.5s (-57%) > > M2, gcc 18.6s 9.9s -8.7s (-47%) > > M2, clang 18.6s 8.2s -10.4s (-56%) > > > > == defconfig NO-OP build == > > > > before after delta > > ---------------------------------- > > Threadripper, gcc 0.96s 0.47s -0.49s (-51%) > > Threadripper, clang 1.21s 0.55s -0.66s (-55%) > > EPYC, gcc 1.47s 0.66s -0.81s (-55%) > > EPYC, clang 1.87s 0.80s -1.07s (-57%) > > M2, gcc 5.59s 1.69s -3.90s (-70%) > > M2, clang 6.62s 1.74s -4.88s (-74%) > > > > Further performance numbers are provided for each commit giving a sense of > > what each contributes to the final result. > > > > == Testing == > > > > Beyond x86, allmodconfig was built with the series for arm64, arm, riscv, > > powerpc64, s390 and loongarch, and for arm64, arm, s390 and loongarch the > > System.map is identical to what the previous shell mksysmap produces for > > the same vmlinux. > > > > Also tested were parisc64 and m68k build as far as mainline lets them > > (a driver's static assertion and an undefined-symbol check on parisc, gcc > > 16 internal compiler errors on m68k, none of it from this series). > > > > Kernels for x86, arm64, arm, riscv, loongarch, powerpc64, s390, m68k and > > parisc64 all boot under qemu, every text symbol of System.map is in > > /proc/kallsyms at the relocated address, and a module loads and unloads. > > The s390, arm and loongarch kernels also pass the kallsyms selftest. > > > > An x86 kernel with CONFIG_MODVERSIONS, CONFIG_EXTENDED_MODVERSIONS and > > CONFIG_MODULE_SRCVERSION_ALL boots, loads and unloads modules. External > > modules build against both in-tree and O= builds, and every commit builds > > on x86 defconfig. > > > > Build times are the best of several runs, no unexpected errors or > > warnings were seen. > > > > While some aspects of the build process have been changed, all tooling > > should function identically to before. > > I appreciate all of the testing that you have done! I will test this > side by side on a couple of my own machines to see what results I get > but I do like those results. Ack thanks :) > > > == LLM usage == > > > > An LLM was used to first determine where the bottlenecks were then to > > figure out how to improve them. > > > > It generated a lot of code, much of it hideous. > > > > I extensively audited and rewrote a lot of it, and heavily edited commit > > messages, the cover letter and comments.x > > > > The LLM has also orchestrated build runs, testing, debugging and analysis. > > > > I have manually checked for correctness in both build and running kernels > > generated with this series applied. > > > > Performance improvements were also verified manually. > > > > Since an LLM was used extensively, each commit carries an Assisted-by tag. > > Thank you for calling all of this out as well. Of course, I am on the receiving end of a flood of LLM patches in mm, so it's important to me to be transparent and consistent. I repeatedly tell people - audit what it makes with a human pass, fix up awful code/comments/commit msgs, make sure you understand it + on you to make upstreamable + obviously ack that you used it. So what's good for the goose is good for the gander :) > > > == What was changed? == > > > > Fundamentally the series improves build times by parallelising > > single-threaded tasks as much as possible and improving the efficiency of > > code used in the build process. > > > > kbuild, kallsyms, modpost, objtool, mksysmap and the rust build system were > > all updated as part of this change. > > > > Nothing too controversial was included. There are further improvements that > > could be made, but they would either by very invasive (large scale C header > > changes) or generate diminishing returns. > > > > == Patches == > > > > mksysmap (1, 2): A couple of bugfixes the rest of the series relies > > upon. > > As I note in other threads of this review, I think we should take these > to Linus now, they seem Obviously CorrectTM and it helps chunk out the > series. Ack thanks. I will continue to include them in respins just to enforce the ordering in the meantime, but it should all come out in the wash I think? I can add something in the cover letter too about these. > > I have just reviewed a few of the low hanging fruit patches. I will try > to take a look at the rest over the next couple of weeks but I might not > get to them until after Plumbers. I would greatly appreciate if there > are others who are more knowledgeable in these areas who could review > these changes, as I don't know too much about some of the corners of > Kbuild yet. Thanks so much for taking a look! I'll be at plumbers so feel free to come say hi + discuss any of this in person. And thanks to everybody else reviewing also! I've replied to the more straightforward stuff, will take a deeper look at Linus's comments + anything else that needs a deeper think + sashiko reports a little later before sending a v2. > > -- > Cheers, > Nathan > -- Cheers, Lorenzo
On 9/8/26 13:55, Lorenzo Stoakes (ARM) wrote: > A typical kernel build consists of a frustratingly large amount of time > spent stuck in single-threaded bottlenecks. > > It turns out that there's a lot we can do about this and doing so > significantly impacts kernel build times. > > This series makes allmodconfig builds up to 36% faster, incremental builds > up to ~70% faster, and noop builds up to ~90% faster. > > Builds are faster across the board on every device I tested. > > Machines used for perf testing: > > * Threadripper - x86, AMD Threadripper 9980X, 64 cores, 128 threads > * EPYC - x86, 2 socket EPYC 9754, 256 cores, 512 threads > * M2 - arm64, 2022 M2 macbook pro, 8 cores, 8 threads > (4 perf, 4 efficiency) My AMD EPYC 9454P server shows an improvement in the order of 13% using the arm64's defconfig and make -j97. Thanks for doing this! -- Florian
On Wed, Sep 09, 2026 at 02:58:59PM -0700, Florian Fainelli wrote: > On 9/8/26 13:55, Lorenzo Stoakes (ARM) wrote: > > A typical kernel build consists of a frustratingly large amount of time > > spent stuck in single-threaded bottlenecks. > > > > It turns out that there's a lot we can do about this and doing so > > significantly impacts kernel build times. > > > > This series makes allmodconfig builds up to 36% faster, incremental builds > > up to ~70% faster, and noop builds up to ~90% faster. > > > > Builds are faster across the board on every device I tested. > > > > Machines used for perf testing: > > > > * Threadripper - x86, AMD Threadripper 9980X, 64 cores, 128 threads > > * EPYC - x86, 2 socket EPYC 9754, 256 cores, 512 threads > > * M2 - arm64, 2022 M2 macbook pro, 8 cores, 8 threads > > (4 perf, 4 efficiency) > > My AMD EPYC 9454P server shows an improvement in the order of 13% using the > arm64's defconfig and make -j97. Thanks for doing this! Nice thanks! Good to confirm others are seeing the benefits :) > -- > Florian > -- Cheers, Lorenzo
On Fri, Sep 11, 2026 at 12:28:34PM +0100, Lorenzo Stoakes (ARM) wrote: > On Wed, Sep 09, 2026 at 02:58:59PM -0700, Florian Fainelli wrote: > > On 9/8/26 13:55, Lorenzo Stoakes (ARM) wrote: > > > A typical kernel build consists of a frustratingly large amount of time > > > spent stuck in single-threaded bottlenecks. > > > > > > It turns out that there's a lot we can do about this and doing so > > > significantly impacts kernel build times. > > > > > > This series makes allmodconfig builds up to 36% faster, incremental builds > > > up to ~70% faster, and noop builds up to ~90% faster. > > > > > > Builds are faster across the board on every device I tested. > > > > > > Machines used for perf testing: > > > > > > * Threadripper - x86, AMD Threadripper 9980X, 64 cores, 128 threads > > > * EPYC - x86, 2 socket EPYC 9754, 256 cores, 512 threads > > > * M2 - arm64, 2022 M2 macbook pro, 8 cores, 8 threads > > > (4 perf, 4 efficiency) > > > > My AMD EPYC 9454P server shows an improvement in the order of 13% using the > > arm64's defconfig and make -j97. Thanks for doing this! > > Nice thanks! Good to confirm others are seeing the benefits :) I see similar gains on an 80-core Ampere Altra system 6h 21m 42s -> 5h 31m 15s (-13.22%) and a 32-core AMD system. 3h 38m 55s -> 3h 13m 18s (-11.7%) The build matrices were different between machines but it is good to see similar numbers across different environments. -- Cheers, Nathan
On Fri, Sep 11, 2026 at 11:44:37PM -0700, Nathan Chancellor wrote: > On Fri, Sep 11, 2026 at 12:28:34PM +0100, Lorenzo Stoakes (ARM) wrote: > > On Wed, Sep 09, 2026 at 02:58:59PM -0700, Florian Fainelli wrote: > > > On 9/8/26 13:55, Lorenzo Stoakes (ARM) wrote: > > > > A typical kernel build consists of a frustratingly large amount of time > > > > spent stuck in single-threaded bottlenecks. > > > > > > > > It turns out that there's a lot we can do about this and doing so > > > > significantly impacts kernel build times. > > > > > > > > This series makes allmodconfig builds up to 36% faster, incremental builds > > > > up to ~70% faster, and noop builds up to ~90% faster. > > > > > > > > Builds are faster across the board on every device I tested. > > > > > > > > Machines used for perf testing: > > > > > > > > * Threadripper - x86, AMD Threadripper 9980X, 64 cores, 128 threads > > > > * EPYC - x86, 2 socket EPYC 9754, 256 cores, 512 threads > > > > * M2 - arm64, 2022 M2 macbook pro, 8 cores, 8 threads > > > > (4 perf, 4 efficiency) > > > > > > My AMD EPYC 9454P server shows an improvement in the order of 13% using the > > > arm64's defconfig and make -j97. Thanks for doing this! > > > > Nice thanks! Good to confirm others are seeing the benefits :) > > I see similar gains on an 80-core Ampere Altra system > > 6h 21m 42s -> 5h 31m 15s (-13.22%) > > and a 32-core AMD system. > > 3h 38m 55s -> 3h 13m 18s (-11.7%) > > The build matrices were different between machines but it is good to see > similar numbers across different environments. Nice, yes for sure! > > -- > Cheers, > Nathan -- Cheers, Lorenzo
On Tue, 8 Sept 2026 at 13:55, Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote:
>
> A typical kernel build consists of a frustratingly large amount of time
.. you're preaching to the choir.
> == allmodconfig FULL build ==
>
> before after delta
> ----------------------------------
> Threadripper, gcc 345.3s 275.2s -70.1s (-20%)
Well, that's certainly very encouraging. We should do this. I sent out
some emails as just reactions to individual patches, but I have no
numbers to back any of those emails up, and I never actually applied
this series to my tree - they were all based on reading the patches
themselves.
That said:
> An LLM was used to first determine where the bottlenecks were then to
> figure out how to improve them.
>
> It generated a lot of code, much of it hideous.
This made me scared to look at the patches originally, and I held off
in fear that the patches would be horrible and this build time
improvement would be hugely controversial garbage code.
But:
> I extensively audited and rewrote a lot of it, and heavily edited commit
> messages, the cover letter and comments.
None of the patches look at all horrible to me. You clearly excised
the hideous parts. All of my reactions were of the type "this could
probably be taken _further_" rather than me throwing my hands up in
disgust.
So the Rust parallelism thing clearly isn't ready based on feedback
from that quarter, but the rest looked safe and innocuous. I'm all for
merging this, although it should obviously go in through the right
channels. Mostly the kbuild tree, although some of it clearly would be
other cases - the objtool change in particular is fairly substantial
and needsobjtool people to approve. It didn't look all that
contentious, but still..
Anyway, I'd love for this all to go in. Build times are a pet peeve of mine.
Linus
On Wed, Sep 09, 2026 at 08:37:10AM -0700, Linus Torvalds wrote: > On Tue, 8 Sept 2026 at 13:55, Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote: > > > > A typical kernel build consists of a frustratingly large amount of time > > .. you're preaching to the choir. :) Yeah, it's something that's annoyed me for a long time too. > > > == allmodconfig FULL build == > > > > before after delta > > ---------------------------------- > > Threadripper, gcc 345.3s 275.2s -70.1s (-20%) > > Well, that's certainly very encouraging. We should do this. I sent out > some emails as just reactions to individual patches, but I have no > numbers to back any of those emails up, and I never actually applied > this series to my tree - they were all based on reading the patches > themselves. Thanks, will have a read through those! I do recommend trying them locally ;) It makes a surprising difference to defconfig even. (Though you need pigz installed for the full effect) > > That said: > > > An LLM was used to first determine where the bottlenecks were then to > > figure out how to improve them. > > > > It generated a lot of code, much of it hideous. > > This made me scared to look at the patches originally, and I held off > in fear that the patches would be horrible and this build time > improvement would be hugely controversial garbage code. > > But: > > > I extensively audited and rewrote a lot of it, and heavily edited commit > > messages, the cover letter and comments. > > None of the patches look at all horrible to me. You clearly excised > the hideous parts. All of my reactions were of the type "this could > probably be taken _further_" rather than me throwing my hands up in > disgust. Thanks! :) I made sure to audit everything it came up with patch-by-patch, got it to rework the horrors then rewrote large chunks of what it came up with. > > So the Rust parallelism thing clearly isn't ready based on feedback > from that quarter, but the rest looked safe and innocuous. I'm all for Sure, happy to drop anything like that. Though based on what Miguel said it _might_ be ok perhaps with a nightly gate on it? Will see what he/Bjorn say. > merging this, although it should obviously go in through the right > channels. Mostly the kbuild tree, although some of it clearly would be > other cases - the objtool change in particular is fairly substantial > and needsobjtool people to approve. It didn't look all that > contentious, but still.. Agreed, it definitely needs some eyes on it, especially the trickier bits. > > Anyway, I'd love for this all to go in. Build times are a pet peeve of mine. Yeah, it'd be nice to see this benefit everyone! Am happy to iterate on it as needed to get it upstream. > > Linus -- Cheers, Lorenzo
On Tue, Sep 8, 2026 at 1:55 PM Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote: > > A typical kernel build consists of a frustratingly large amount of time > spent stuck in single-threaded bottlenecks. > > It turns out that there's a lot we can do about this and doing so > significantly impacts kernel build times. > Fundamentally the series improves build times by parallelising > single-threaded tasks as much as possible and improving the efficiency of > code used in the build process. Cool! I have not reviewed this patch series, but am supportive of solving the problem, since I do recall this being an issue; i.e. it always annoyed me that despite have a 128 core machine, much of the build is sequential! Having all those cores and not saturating them during a build was always something I noticed (usually have system monitor always showing system load, always) but never had time to look into. It would be cool if your LLM could generate any visualizations of what's going on during a build? Seeing what's serial vs parallelized in perhaps useful. Particularly when you can compare the before vs after. -- Thanks, ~Nick Desaulniers
On Tue, Sep 08, 2026 at 02:06:18PM -0700, Nick Desaulniers wrote:
> On Tue, Sep 8, 2026 at 1:55 PM Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote:
> >
> > A typical kernel build consists of a frustratingly large amount of time
> > spent stuck in single-threaded bottlenecks.
> >
> > It turns out that there's a lot we can do about this and doing so
> > significantly impacts kernel build times.
>
>
> > Fundamentally the series improves build times by parallelising
> > single-threaded tasks as much as possible and improving the efficiency of
> > code used in the build process.
>
> Cool! I have not reviewed this patch series, but am supportive of
> solving the problem, since I do recall this being an issue;
Thanks :)
>
> i.e. it always annoyed me that despite have a 128 core machine, much
> of the build is sequential! Having all those cores and not saturating
> them during a build was always something I noticed (usually have
> system monitor always showing system load, always) but never had time
> to look into.
Yeah this was exactly the motivation.
I've seen single-threaded stalls endlessly in kernel builds, even on smaller
boxes, and thought that there must be a lot of low-hanging fruit there.
Obviously I have not enough time to do anything extra (TM) so never dug
into it.
Which made it a great candidate for LLM usage (though with a LOT of taming,
reviewing, auditing, rewriting of code because I _still_ have to be able to
maintain my changes and there's still a bar there :)
>
> It would be cool if your LLM could generate any visualizations of
> what's going on during a build? Seeing what's serial vs parallelized
> in perhaps useful. Particularly when you can compare the before vs
> after.
Great idea!
I asked it to do exactly that, attached. Came out quite nice and sums it up
nicely :)
It also provided an ASCII version:
x86 allmodconfig, clang, 128 threads. One character per interval: height = CPU utilisation,
letter beneath = what was running (C compile, L link, O objtool, P modpost, K kallsyms, Z compress, m make, . idle)
== clean build (4s per character) ==
before ▁▁▁▁▁▂█████████████████████████████████████████████████████████▃▁▁▁▁▁▁▁███████████▇▅▁▁ 343.6s
mCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCOOOOPPPCCCCCCCCCCCCCLL
after ▁▆████████████████████████████████████████████████████████▇▁▁▁▄██▆▁ 265.6s
CCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCOOCCCCC
== touch mm/vma.c (0.5s per character) ==
before ▁▁▁█▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ 45.7s
mCmCCmmmCLOOOOOOOOOOOOOOOOOOOOmPPPPPPPPPPPPPPPPPPPCLmKKKCCCmmKKKCCCLmmmmmmCCZZZZZZZZZZZZZZC
after ▁▁▅▁▁▁▁▁▁▁▁▁▁▁▁▁▁▂▁▁▁▁▁▁▁▁▁▁▁▁▂▁ 15.7s
mmCCCLOOOOOOOOOPPPPLKCKCLmmmmmCC
== no-op make (0.2s per character) ==
before ▁▁▁▁▁▁▁███▂▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁ 14.4s
CCCCmmmCCCmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmmm
after ▁▁▁▁▄▆▄▁ 2.0s
CCmmmCmm
> --
> Thanks,
> ~Nick Desaulniers
--
Cheers, Lorenzo
On Wed, Sep 9, 2026 at 7:17 AM Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote: > > On Tue, Sep 08, 2026 at 02:06:18PM -0700, Nick Desaulniers wrote: > > It would be cool if your LLM could generate any visualizations of > > what's going on during a build? Seeing what's serial vs parallelized > > in perhaps useful. Particularly when you can compare the before vs > > after. > > Great idea! > > I asked it to do exactly that, attached. Came out quite nice and sums it up > nicely :) > > It also provided an ASCII version: unreadable, but the pictures you attached are worth a thousand words. Particularly exciting are the speed ups to the usual edit+compile+run loop. For the blue graph (the top one, first), any idea what that dip in the middle is? That's always what annoyed me. -- Thanks, ~Nick Desaulniers
On Wed, Sep 09, 2026 at 03:09:10PM -0700, Nick Desaulniers wrote: > On Wed, Sep 9, 2026 at 7:17 AM Lorenzo Stoakes (ARM) <ljs@kernel.org> wrote: > > > > On Tue, Sep 08, 2026 at 02:06:18PM -0700, Nick Desaulniers wrote: > > > It would be cool if your LLM could generate any visualizations of > > > what's going on during a build? Seeing what's serial vs parallelized > > > in perhaps useful. Particularly when you can compare the before vs > > > after. > > > > Great idea! > > > > I asked it to do exactly that, attached. Came out quite nice and sums it up > > nicely :) > > > > It also provided an ASCII version: > > unreadable, but the pictures you attached are worth a thousand words. :) Yeah I didn't think the ASCII was amazing either but fun that it tried ;) and yeah I was surprised at how nice the graphic came out...! > Particularly exciting are the speed ups to the usual edit+compile+run > loop. Yes! I never really expected that out of this! I mean I didn't think there'd be such a big speed up in general, though I expected at least to shave off a big chunk of the delta between compile + full build on a defconfig clean build. But the incremental/noop build improvements are really important to day-to-day use. > > For the blue graph (the top one, first), any idea what that dip in the > middle is? That's always what annoyed me. That's ld -r of vmlinux.o, then modpost and objtool which are critical path blockers on the rest of the build. Before the series that was 23 seconds for allmodconfig on the threadripper, now 7s :) There's still more objtool stuff that could be done that I opted not to do yet, it takes 6s to run so dominating. > -- > Thanks, > ~Nick Desaulniers -- Cheers, Lorenzo
© 2016 - 2026 Red Hat, Inc.