fs/coredump.c | 178 +++++++-- include/linux/coredump.h | 7 + include/uapi/linux/coredump.h | 65 +++- tools/include/uapi/linux/coredump.h | 65 +++- .../coredump/coredump_socket_protocol_test.c | 415 +++++++++++++++++++-- tools/testing/selftests/coredump/coredump_test.h | 9 +- .../selftests/coredump/coredump_test_helpers.c | 200 +++++++++- 7 files changed, 849 insertions(+), 90 deletions(-)
A coredump generated via the coredump socket ends up transferring
zeroed data when a mapping contains holes. For a large process that
maps a bunch of data that's wasting a ton of work.
Jacob ran into this and Josef has bitched^wcomplained about this to me
before. I dislike the coredump_filter bit solution in [1] which stops
each PT_LOAD at the last populated page.
The problem is real though. I don't think coredump_filter is where we
need to solve this. That mask says which kinds of memory to include and
it propagates across fork and exec, whereas what is being selected here
is an encoding mechanism.
I also think that the usermodehelper - may it swiftly die - isn't really
salvagable for this and it's not the future anyway. The coredump socket
already has a handshake for stuff like this.
I always had an idea how this would look like but punted on it back
then. So here it is.
A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get
the coredump as a plain byte stream but as a sequence of frames. Each
one a struct coredump_frame_header followed by what it describes. A data
frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero
frames are sent for unpopulated mappings. They only indicate how many
zero bytes need to be written and to not include data. Reassembling the
frames gives back the same coredump. A debugger and everything else
still see an ordinary core file and nothing outside the coredump server
has to learn anything.
Numbers from the selftest in patch 11, on a kernel built from this
series:
- a process with 128 threads: 1740014 bytes on the socket for a
coredump of 1075150848 bytes
- a 256MB mapping with one page touched: 170542 bytes on the socket for
a coredump of 268890112 bytes
- the same 256MB mapping with COREDUMP_HEADER alone: 270993024 bytes on
the socket, so the framing overhead itself is under one percent
The first one is the interesting case. Almost all of it is thread stacks.
All stacks are 8MB reservations that are nearly all holes. And they are
holes in the middle of the dump rather than at the end.
Link: https://lore.kernel.org/all/20260731171336.2255844-1-jalalonde@meta.com [1]
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Christian Brauner (11):
selftests/coredump: discard the right amount after the coredump request
selftests/coredump: collapse the expected request check into the helper
coredump: pin the protocol struct sizes
coredump: move the negotiated mask into struct coredump_params
coredump: deduplicate the to_skip flush
coredump: add COREDUMP_HEADER to the coredump socket protocol
coredump: add COREDUMP_SPARSE to the coredump socket protocol
tools: sync coredump.h header
coredump: frame the coredump when COREDUMP_HEADER is negotiated
coredump: describe the holes when COREDUMP_SPARSE is negotiated
selftests/coredump: test COREDUMP_HEADER and COREDUMP_SPARSE
fs/coredump.c | 178 +++++++--
include/linux/coredump.h | 7 +
include/uapi/linux/coredump.h | 65 +++-
tools/include/uapi/linux/coredump.h | 65 +++-
.../coredump/coredump_socket_protocol_test.c | 415 +++++++++++++++++++--
tools/testing/selftests/coredump/coredump_test.h | 9 +-
.../selftests/coredump/coredump_test_helpers.c | 200 +++++++++-
7 files changed, 849 insertions(+), 90 deletions(-)
---
base-commit: db2ddb87143519e20a95aa36c60b36107b736a58
change-id: 20260811-work-coredump-sparse-18177d77b014
On Tue, Aug 11, 2026 at 05:27:21PM +0200, Christian Brauner wrote: > A coredump generated via the coredump socket ends up transferring > zeroed data when a mapping contains holes. For a large process that > maps a bunch of data that's wasting a ton of work. > > Jacob ran into this and Josef has bitched^wcomplained about this to me > before. I dislike the coredump_filter bit solution in [1] which stops > each PT_LOAD at the last populated page. > > The problem is real though. I don't think coredump_filter is where we > need to solve this. That mask says which kinds of memory to include and > it propagates across fork and exec, whereas what is being selected here > is an encoding mechanism. > > I also think that the usermodehelper - may it swiftly die - isn't really > salvagable for this and it's not the future anyway. The coredump socket > already has a handshake for stuff like this. > > I always had an idea how this would look like but punted on it back > then. So here it is. > > A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get > the coredump as a plain byte stream but as a sequence of frames. Each > one a struct coredump_frame_header followed by what it describes. A data > frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero > frames are sent for unpopulated mappings. They only indicate how many > zero bytes need to be written and to not include data. Reassembling the > frames gives back the same coredump. A debugger and everything else > still see an ordinary core file and nothing outside the coredump server > has to learn anything. Hey, Christian, I proposed pretty much this exact solution to Jacob, so thank you for writing it :) There are a couple of reasons we still wanted to explore the coredump_filter solution: 1. Our core dumper application is not really prepared to run as a daemon that listens on a socket, having been written to be a transient usermode helper. But thinking about it more, maybe that's something we could paper over with systemd socket activation? 2. More importantly, we sometimes write core dumps to disk and sometimes upload them to blob storage. For the former, this approach of sending holes over the socket is great. For the latter, we'd now need to wrap the dump in some sort of container supporting sparseness that all consumers then need to reassemble. The coredump_filter approach doesn't require any changes in that pipeline. To be transparent, I still prefer the sparse socket approach, but Jacob has different contraints that I'd love to have addressed: it's a trade-off of more work on the core dump server side and all of its consumers vs. in the debug tooling side, which is mostly already there. Thanks, Omar
On Thu, Aug 13, 2026 at 10:08:36AM -0700, Omar Sandoval wrote: > On Tue, Aug 11, 2026 at 05:27:21PM +0200, Christian Brauner wrote: > > A coredump generated via the coredump socket ends up transferring > > zeroed data when a mapping contains holes. For a large process that > > maps a bunch of data that's wasting a ton of work. > > > > Jacob ran into this and Josef has bitched^wcomplained about this to me > > before. I dislike the coredump_filter bit solution in [1] which stops > > each PT_LOAD at the last populated page. > > > > The problem is real though. I don't think coredump_filter is where we > > need to solve this. That mask says which kinds of memory to include and > > it propagates across fork and exec, whereas what is being selected here > > is an encoding mechanism. > > > > I also think that the usermodehelper - may it swiftly die - isn't really > > salvagable for this and it's not the future anyway. The coredump socket > > already has a handshake for stuff like this. > > > > I always had an idea how this would look like but punted on it back > > then. So here it is. > > > > A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get > > the coredump as a plain byte stream but as a sequence of frames. Each > > one a struct coredump_frame_header followed by what it describes. A data > > frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero > > frames are sent for unpopulated mappings. They only indicate how many > > zero bytes need to be written and to not include data. Reassembling the > > frames gives back the same coredump. A debugger and everything else > > still see an ordinary core file and nothing outside the coredump server > > has to learn anything. > > Hey, Christian, > > I proposed pretty much this exact solution to Jacob, so thank you for > writing it :) > > There are a couple of reasons we still wanted to explore the > coredump_filter solution: > > 1. Our core dumper application is not really prepared to run as a daemon > that listens on a socket, having been written to be a transient > usermode helper. But thinking about it more, maybe that's something > we could paper over with systemd socket activation? > 2. More importantly, we sometimes write core dumps to disk and sometimes > upload them to blob storage. For the former, this approach of sending > holes over the socket is great. For the latter, we'd now need to wrap > the dump in some sort of container supporting sparseness that all > consumers then need to reassemble. The coredump_filter approach > doesn't require any changes in that pipeline. > > To be transparent, I still prefer the sparse socket approach, but Jacob > has different contraints that I'd love to have addressed: it's a > trade-off of more work on the core dump server side and all of its > consumers vs. in the debug tooling side, which is mostly already there. Hey Omar, Good to hear from you. I think the coredump_filter solution is just the wrong approach. It is a hack and in a part of coredumping that is riddled with bolted on hacks already. And I don't think the argument that changing this in userspace is somehow not possible holds. It not being wanted is another thing. This is a new feature and as such it is our responsibility to expose it in a way that is correct, maintainable, and clean. I don't think that this is possible in the usermodehelper case without doing really hacky things. Plus, I really want usermodehelpers to go away. It really is a CVE machine. Most of the ~40 CVEs in the last 10+ years have been in the handlers. Mainly because the interface is really not nice and prone to security issues. They are spawned from a kthread with kernel level privileges that userspace then needs to drop. So really, changing userspace is a long-term self-service. :) And yes, there doesn't have to be anything long-running. They can just be socket activated. The pull request that's up against systemd splits this out into systemd-coredump though. So PID 1 itself can also be coredumped. That daemon is long-running but it's only job is to spawn workers on incoming connections and maintain backpressure. The listen fd is in the fdstore and so survives restarts. > consumers then need to reassemble. The coredump_filter approach > doesn't require any changes in that pipeline. We could add COREDUMP_TRAILING_ZERO or whatever so you get that truncated behavior over the socket from the earlier patch. While that will get rid of smuggling non-memory type bits into the filter and won't have the fork/exec inheritance bugs that the other patch has it still has other issues. It will always have to accept an O(hole / PAGE_SIZE) walk under the mmap write lock before ever writing a coredump. So a JVM starting with a 64 gb heap reservation or a go program with a large reservation means millions of probes before the dump even starts. The selftest in this series dumpds with 256 MB. That amounts to 65536 probes. This seems just wrong when we know the holes when we emit the dump. For default 8 mb pthread stacks that are in a single anon vma 4 × 8 mb = 32 mb == 8196 pages only 8 of the 8196 pages are resident. Most stacks grow down and glibc parks the tcb at the top. That in turn means every hole in a stack is a leading hole. The trailing zero hack cannot handle that at all wasting a bunch of space. For the blob storage thing you can just upload the header last. So stream the data as the records arrive and compute the corrected phdr table at the end. If you're using a temporary file or whatever it becomes even easier to rewrite the header.
On Tue, Aug 11, 2026 at 5:27 PM Christian Brauner <brauner@kernel.org> wrote: > A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get > the coredump as a plain byte stream but as a sequence of frames. Each > one a struct coredump_frame_header followed by what it describes. A data > frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero > frames are sent for unpopulated mappings. They only indicate how many > zero bytes need to be written and to not include data. Reassembling the > frames gives back the same coredump. A debugger and everything else > still see an ordinary core file and nothing outside the coredump server > has to learn anything. Hmm... I think what you're doing is probably the easiest way to do this in practice. I guess some design alternatives would be: 1. (overengineered, not generically useful enough): If the transport was a pipe (which already has the concept of different types of pipe buffers) instead of a unix domain socket, we could introduce a special representation for zero-filled pipe buffers and some API for receiving zeroed holes through lseek(pipefd, 0, SEEK_DATA), but that's probably not sufficiently useful for stuff other than core dumping to be worth the effort. 2. (somewhat overengineered) With some refactoring, we could maybe do something like /proc/kcore and create a seekable virtual coredump file that we send over the socket via SCM_RIGHTS? 3. We could leave the userspace memory dump out of the core dump data, and let userspace take care of filling out the memory contents using /proc/$pid/pagemap and /proc/$pid/mem? That would also avoid task switches and SKB allocations, and probably reduce the number of data copies involved in this by 1.
On Tue, Aug 11, 2026 at 9:07 PM Jann Horn <jannh@google.com> wrote: > On Tue, Aug 11, 2026 at 5:27 PM Christian Brauner <brauner@kernel.org> wrote: > > A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get > > the coredump as a plain byte stream but as a sequence of frames. Each > > one a struct coredump_frame_header followed by what it describes. A data > > frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero > > frames are sent for unpopulated mappings. They only indicate how many > > zero bytes need to be written and to not include data. Reassembling the > > frames gives back the same coredump. A debugger and everything else > > still see an ordinary core file and nothing outside the coredump server > > has to learn anything. > > Hmm... > > I think what you're doing is probably the easiest way to do this in > practice. I guess some design alternatives would be: > > 1. (overengineered, not generically useful enough): If the transport > was a pipe (which already has the concept of different types of pipe > buffers) instead of a unix domain socket, we could introduce a special > representation for zero-filled pipe buffers and some API for receiving > zeroed holes through lseek(pipefd, 0, SEEK_DATA), but that's probably > not sufficiently useful for stuff other than core dumping to be worth > the effort. > 2. (somewhat overengineered) With some refactoring, we could maybe do > something like /proc/kcore and create a seekable virtual coredump file > that we send over the socket via SCM_RIGHTS? > 3. We could leave the userspace memory dump out of the core dump data, > and let userspace take care of filling out the memory contents using > /proc/$pid/pagemap and /proc/$pid/mem? That would also avoid task > switches and SKB allocations, and probably reduce the number of data > copies involved in this by 1. I also wonder what userspace actually does with this data - does userspace just want to write it to disk, potentially after compressing it? Or does userspace actually do some fancy parsing of the data stream to extract stack memory or something like that? Or does userspace buffer the whole thing into RAM and then process it from there (it kinda looks like systemd tries to do that but I might be reading this wrong)? Anyway, I've looked through your code and it does look fine to me.
On Tue, Aug 11, 2026 at 09:39:12PM +0200, Jann Horn wrote: > On Tue, Aug 11, 2026 at 9:07 PM Jann Horn <jannh@google.com> wrote: > > On Tue, Aug 11, 2026 at 5:27 PM Christian Brauner <brauner@kernel.org> wrote: > > > A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get > > > the coredump as a plain byte stream but as a sequence of frames. Each > > > one a struct coredump_frame_header followed by what it describes. A data > > > frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero > > > frames are sent for unpopulated mappings. They only indicate how many > > > zero bytes need to be written and to not include data. Reassembling the > > > frames gives back the same coredump. A debugger and everything else > > > still see an ordinary core file and nothing outside the coredump server > > > has to learn anything. > > > > Hmm... > > > > I think what you're doing is probably the easiest way to do this in > > practice. I guess some design alternatives would be: > > > > 1. (overengineered, not generically useful enough): If the transport > > was a pipe (which already has the concept of different types of pipe > > buffers) instead of a unix domain socket, we could introduce a special > > representation for zero-filled pipe buffers and some API for receiving > > zeroed holes through lseek(pipefd, 0, SEEK_DATA), but that's probably > > not sufficiently useful for stuff other than core dumping to be worth > > the effort. > > 2. (somewhat overengineered) With some refactoring, we could maybe do > > something like /proc/kcore and create a seekable virtual coredump file > > that we send over the socket via SCM_RIGHTS? > > 3. We could leave the userspace memory dump out of the core dump data, > > and let userspace take care of filling out the memory contents using > > /proc/$pid/pagemap and /proc/$pid/mem? That would also avoid task > > switches and SKB allocations, and probably reduce the number of data > > copies involved in this by 1. > > I also wonder what userspace actually does with this data - does > userspace just want to write it to disk, potentially after compressing > it? Or does userspace actually do some fancy parsing of the data That would be the most straightforward use-case, yes. > stream to extract stack memory or something like that? Or does > userspace buffer the whole thing into RAM and then process it from > there (it kinda looks like systemd tries to do that but I might be > reading this wrong)? I'm not sure about that. I think it's writing it to disk and then parsing it. But it also wranges it into a socket to forward to containers or services. Note that systemd has a pull request for the coredump socket up: https://github.com/systemd/systemd/pull/43330 This should kill the usermodehelper soon on kernels that support the socket. > Anyway, I've looked through your code and it does look fine to me. Thanks.
© 2016 - 2026 Red Hat, Inc.