[PATCH 00/11] coredump: allow to create sparse coredumps on the coredump socket

Christian Brauner posted 11 patches 1 month, 2 weeks ago
There is a newer version of this series
fs/coredump.c                                      | 178 +++++++--
include/linux/coredump.h                           |   7 +
include/uapi/linux/coredump.h                      |  65 +++-
tools/include/uapi/linux/coredump.h                |  65 +++-
.../coredump/coredump_socket_protocol_test.c       | 415 +++++++++++++++++++--
tools/testing/selftests/coredump/coredump_test.h   |   9 +-
.../selftests/coredump/coredump_test_helpers.c     | 200 +++++++++-
7 files changed, 849 insertions(+), 90 deletions(-)
[PATCH 00/11] coredump: allow to create sparse coredumps on the coredump socket
Posted by Christian Brauner 1 month, 2 weeks ago
A coredump generated via the coredump socket ends up transferring
zeroed data when a mapping contains holes. For a large process that
maps a bunch of data that's wasting a ton of work.

Jacob ran into this and Josef has bitched^wcomplained about this to me
before. I dislike the coredump_filter bit solution in [1] which stops
each PT_LOAD at the last populated page.

The problem is real though. I don't think coredump_filter is where we
need to solve this. That mask says which kinds of memory to include and
it propagates across fork and exec, whereas what is being selected here
is an encoding mechanism.

I also think that the usermodehelper - may it swiftly die - isn't really
salvagable for this and it's not the future anyway. The coredump socket
already has a handshake for stuff like this.

I always had an idea how this would look like but punted on it back
then. So here it is.

A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get
the coredump as a plain byte stream but as a sequence of frames. Each
one a struct coredump_frame_header followed by what it describes. A data
frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero
frames are sent for unpopulated mappings. They only indicate how many
zero bytes need to be written and to not include data. Reassembling the
frames gives back the same coredump. A debugger and everything else
still see an ordinary core file and nothing outside the coredump server
has to learn anything.

Numbers from the selftest in patch 11, on a kernel built from this
series:

- a process with 128 threads: 1740014 bytes on the socket for a
  coredump of 1075150848 bytes
- a 256MB mapping with one page touched: 170542 bytes on the socket for
  a coredump of 268890112 bytes
- the same 256MB mapping with COREDUMP_HEADER alone: 270993024 bytes on
  the socket, so the framing overhead itself is under one percent

The first one is the interesting case. Almost all of it is thread stacks.
All stacks are 8MB reservations that are nearly all holes. And they are
holes in the middle of the dump rather than at the end.

Link: https://lore.kernel.org/all/20260731171336.2255844-1-jalalonde@meta.com [1]

Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
Christian Brauner (11):
      selftests/coredump: discard the right amount after the coredump request
      selftests/coredump: collapse the expected request check into the helper
      coredump: pin the protocol struct sizes
      coredump: move the negotiated mask into struct coredump_params
      coredump: deduplicate the to_skip flush
      coredump: add COREDUMP_HEADER to the coredump socket protocol
      coredump: add COREDUMP_SPARSE to the coredump socket protocol
      tools: sync coredump.h header
      coredump: frame the coredump when COREDUMP_HEADER is negotiated
      coredump: describe the holes when COREDUMP_SPARSE is negotiated
      selftests/coredump: test COREDUMP_HEADER and COREDUMP_SPARSE

 fs/coredump.c                                      | 178 +++++++--
 include/linux/coredump.h                           |   7 +
 include/uapi/linux/coredump.h                      |  65 +++-
 tools/include/uapi/linux/coredump.h                |  65 +++-
 .../coredump/coredump_socket_protocol_test.c       | 415 +++++++++++++++++++--
 tools/testing/selftests/coredump/coredump_test.h   |   9 +-
 .../selftests/coredump/coredump_test_helpers.c     | 200 +++++++++-
 7 files changed, 849 insertions(+), 90 deletions(-)
---
base-commit: db2ddb87143519e20a95aa36c60b36107b736a58
change-id: 20260811-work-coredump-sparse-18177d77b014
Re: [PATCH 00/11] coredump: allow to create sparse coredumps on the coredump socket
Posted by Omar Sandoval 1 month, 2 weeks ago
On Tue, Aug 11, 2026 at 05:27:21PM +0200, Christian Brauner wrote:
> A coredump generated via the coredump socket ends up transferring
> zeroed data when a mapping contains holes. For a large process that
> maps a bunch of data that's wasting a ton of work.
> 
> Jacob ran into this and Josef has bitched^wcomplained about this to me
> before. I dislike the coredump_filter bit solution in [1] which stops
> each PT_LOAD at the last populated page.
> 
> The problem is real though. I don't think coredump_filter is where we
> need to solve this. That mask says which kinds of memory to include and
> it propagates across fork and exec, whereas what is being selected here
> is an encoding mechanism.
> 
> I also think that the usermodehelper - may it swiftly die - isn't really
> salvagable for this and it's not the future anyway. The coredump socket
> already has a handshake for stuff like this.
> 
> I always had an idea how this would look like but punted on it back
> then. So here it is.
> 
> A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get
> the coredump as a plain byte stream but as a sequence of frames. Each
> one a struct coredump_frame_header followed by what it describes. A data
> frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero
> frames are sent for unpopulated mappings. They only indicate how many
> zero bytes need to be written and to not include data. Reassembling the
> frames gives back the same coredump. A debugger and everything else
> still see an ordinary core file and nothing outside the coredump server
> has to learn anything.

Hey, Christian,

I proposed pretty much this exact solution to Jacob, so thank you for
writing it :)

There are a couple of reasons we still wanted to explore the
coredump_filter solution:

1. Our core dumper application is not really prepared to run as a daemon
   that listens on a socket, having been written to be a transient
   usermode helper. But thinking about it more, maybe that's something
   we could paper over with systemd socket activation?
2. More importantly, we sometimes write core dumps to disk and sometimes
   upload them to blob storage. For the former, this approach of sending
   holes over the socket is great. For the latter, we'd now need to wrap
   the dump in some sort of container supporting sparseness that all
   consumers then need to reassemble. The coredump_filter approach
   doesn't require any changes in that pipeline.

To be transparent, I still prefer the sparse socket approach, but Jacob
has different contraints that I'd love to have addressed: it's a
trade-off of more work on the core dump server side and all of its
consumers vs. in the debug tooling side, which is mostly already there.

Thanks,
Omar
Re: [PATCH 00/11] coredump: allow to create sparse coredumps on the coredump socket
Posted by Christian Brauner 1 month, 2 weeks ago
On Thu, Aug 13, 2026 at 10:08:36AM -0700, Omar Sandoval wrote:
> On Tue, Aug 11, 2026 at 05:27:21PM +0200, Christian Brauner wrote:
> > A coredump generated via the coredump socket ends up transferring
> > zeroed data when a mapping contains holes. For a large process that
> > maps a bunch of data that's wasting a ton of work.
> > 
> > Jacob ran into this and Josef has bitched^wcomplained about this to me
> > before. I dislike the coredump_filter bit solution in [1] which stops
> > each PT_LOAD at the last populated page.
> > 
> > The problem is real though. I don't think coredump_filter is where we
> > need to solve this. That mask says which kinds of memory to include and
> > it propagates across fork and exec, whereas what is being selected here
> > is an encoding mechanism.
> > 
> > I also think that the usermodehelper - may it swiftly die - isn't really
> > salvagable for this and it's not the future anyway. The coredump socket
> > already has a handshake for stuff like this.
> > 
> > I always had an idea how this would look like but punted on it back
> > then. So here it is.
> > 
> > A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get
> > the coredump as a plain byte stream but as a sequence of frames. Each
> > one a struct coredump_frame_header followed by what it describes. A data
> > frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero
> > frames are sent for unpopulated mappings. They only indicate how many
> > zero bytes need to be written and to not include data. Reassembling the
> > frames gives back the same coredump. A debugger and everything else
> > still see an ordinary core file and nothing outside the coredump server
> > has to learn anything.
> 
> Hey, Christian,
> 
> I proposed pretty much this exact solution to Jacob, so thank you for
> writing it :)
> 
> There are a couple of reasons we still wanted to explore the
> coredump_filter solution:
> 
> 1. Our core dumper application is not really prepared to run as a daemon
>    that listens on a socket, having been written to be a transient
>    usermode helper. But thinking about it more, maybe that's something
>    we could paper over with systemd socket activation?
> 2. More importantly, we sometimes write core dumps to disk and sometimes
>    upload them to blob storage. For the former, this approach of sending
>    holes over the socket is great. For the latter, we'd now need to wrap
>    the dump in some sort of container supporting sparseness that all
>    consumers then need to reassemble. The coredump_filter approach
>    doesn't require any changes in that pipeline.
> 
> To be transparent, I still prefer the sparse socket approach, but Jacob
> has different contraints that I'd love to have addressed: it's a
> trade-off of more work on the core dump server side and all of its
> consumers vs. in the debug tooling side, which is mostly already there.

Hey Omar,

Good to hear from you.

I think the coredump_filter solution is just the wrong approach. It is
a hack and in a part of coredumping that is riddled with bolted on hacks
already.

And I don't think the argument that changing this in userspace is
somehow not possible holds. It not being wanted is another thing.

This is a new feature and as such it is our responsibility to expose it
in a way that is correct, maintainable, and clean. I don't think that
this is possible in the usermodehelper case without doing really hacky
things.

Plus, I really want usermodehelpers to go away. It really is a CVE
machine. Most of the ~40 CVEs in the last 10+ years have been in the
handlers. Mainly because the interface is really not nice and prone to
security issues. They are spawned from a kthread with kernel level
privileges that userspace then needs to drop. So really, changing
userspace is a long-term self-service. :)

And yes, there doesn't have to be anything long-running. They can just
be socket activated.

The pull request that's up against systemd splits this out into
systemd-coredump though. So PID 1 itself can also be coredumped. That
daemon is long-running but it's only job is to spawn workers on incoming
connections and maintain backpressure. The listen fd is in the fdstore
and so survives restarts.

>    consumers then need to reassemble. The coredump_filter approach
>    doesn't require any changes in that pipeline.

We could add COREDUMP_TRAILING_ZERO or whatever so you get that
truncated behavior over the socket from the earlier patch. While that
will get rid of smuggling non-memory type bits into the filter and
won't have the fork/exec inheritance bugs that the other patch has it
still has other issues.

It will always have to accept an O(hole / PAGE_SIZE) walk under the mmap
write lock before ever writing a coredump. So a JVM starting with a 64
gb heap reservation or a go program with a large reservation means
millions of probes before the dump even starts. The selftest in this
series dumpds with 256 MB. That amounts to 65536 probes. This seems just
wrong when we know the holes when we emit the dump.

For default 8 mb pthread stacks that are in a single anon vma 4 × 8 mb =
32 mb == 8196 pages only 8 of the 8196 pages are resident. Most stacks
grow down and glibc parks the tcb at the top. That in turn means every
hole in a stack is a leading hole. The trailing zero hack cannot handle
that at all wasting a bunch of space.

For the blob storage thing you can just upload the header last. So
stream the data as the records arrive and compute the corrected phdr
table at the end. If you're using a temporary file or whatever it
becomes even easier to rewrite the header.
Re: [PATCH 00/11] coredump: allow to create sparse coredumps on the coredump socket
Posted by Jann Horn 1 month, 2 weeks ago
On Tue, Aug 11, 2026 at 5:27 PM Christian Brauner <brauner@kernel.org> wrote:
> A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get
> the coredump as a plain byte stream but as a sequence of frames. Each
> one a struct coredump_frame_header followed by what it describes. A data
> frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero
> frames are sent for unpopulated mappings. They only indicate how many
> zero bytes need to be written and to not include data. Reassembling the
> frames gives back the same coredump. A debugger and everything else
> still see an ordinary core file and nothing outside the coredump server
> has to learn anything.

Hmm...

I think what you're doing is probably the easiest way to do this in
practice. I guess some design alternatives would be:

1. (overengineered, not generically useful enough): If the transport
was a pipe (which already has the concept of different types of pipe
buffers) instead of a unix domain socket, we could introduce a special
representation for zero-filled pipe buffers and some API for receiving
zeroed holes through lseek(pipefd, 0, SEEK_DATA), but that's probably
not sufficiently useful for stuff other than core dumping to be worth
the effort.
2. (somewhat overengineered) With some refactoring, we could maybe do
something like /proc/kcore and create a seekable virtual coredump file
that we send over the socket via SCM_RIGHTS?
3. We could leave the userspace memory dump out of the core dump data,
and let userspace take care of filling out the memory contents using
/proc/$pid/pagemap and /proc/$pid/mem? That would also avoid task
switches and SKB allocations, and probably reduce the number of data
copies involved in this by 1.
Re: [PATCH 00/11] coredump: allow to create sparse coredumps on the coredump socket
Posted by Jann Horn 1 month, 2 weeks ago
On Tue, Aug 11, 2026 at 9:07 PM Jann Horn <jannh@google.com> wrote:
> On Tue, Aug 11, 2026 at 5:27 PM Christian Brauner <brauner@kernel.org> wrote:
> > A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get
> > the coredump as a plain byte stream but as a sequence of frames. Each
> > one a struct coredump_frame_header followed by what it describes. A data
> > frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero
> > frames are sent for unpopulated mappings. They only indicate how many
> > zero bytes need to be written and to not include data. Reassembling the
> > frames gives back the same coredump. A debugger and everything else
> > still see an ordinary core file and nothing outside the coredump server
> > has to learn anything.
>
> Hmm...
>
> I think what you're doing is probably the easiest way to do this in
> practice. I guess some design alternatives would be:
>
> 1. (overengineered, not generically useful enough): If the transport
> was a pipe (which already has the concept of different types of pipe
> buffers) instead of a unix domain socket, we could introduce a special
> representation for zero-filled pipe buffers and some API for receiving
> zeroed holes through lseek(pipefd, 0, SEEK_DATA), but that's probably
> not sufficiently useful for stuff other than core dumping to be worth
> the effort.
> 2. (somewhat overengineered) With some refactoring, we could maybe do
> something like /proc/kcore and create a seekable virtual coredump file
> that we send over the socket via SCM_RIGHTS?
> 3. We could leave the userspace memory dump out of the core dump data,
> and let userspace take care of filling out the memory contents using
> /proc/$pid/pagemap and /proc/$pid/mem? That would also avoid task
> switches and SKB allocations, and probably reduce the number of data
> copies involved in this by 1.

I also wonder what userspace actually does with this data - does
userspace just want to write it to disk, potentially after compressing
it? Or does userspace actually do some fancy parsing of the data
stream to extract stack memory or something like that? Or does
userspace buffer the whole thing into RAM and then process it from
there (it kinda looks like systemd tries to do that but I might be
reading this wrong)?

Anyway, I've looked through your code and it does look fine to me.
Re: [PATCH 00/11] coredump: allow to create sparse coredumps on the coredump socket
Posted by Christian Brauner 1 month, 2 weeks ago
On Tue, Aug 11, 2026 at 09:39:12PM +0200, Jann Horn wrote:
> On Tue, Aug 11, 2026 at 9:07 PM Jann Horn <jannh@google.com> wrote:
> > On Tue, Aug 11, 2026 at 5:27 PM Christian Brauner <brauner@kernel.org> wrote:
> > > A server that raises COREDUMP_HEADER in coredump_ack->mask doesn't get
> > > the coredump as a plain byte stream but as a sequence of frames. Each
> > > one a struct coredump_frame_header followed by what it describes. A data
> > > frame carries its bytes. If a server also raises COREDUMP_SPARSE, zero
> > > frames are sent for unpopulated mappings. They only indicate how many
> > > zero bytes need to be written and to not include data. Reassembling the
> > > frames gives back the same coredump. A debugger and everything else
> > > still see an ordinary core file and nothing outside the coredump server
> > > has to learn anything.
> >
> > Hmm...
> >
> > I think what you're doing is probably the easiest way to do this in
> > practice. I guess some design alternatives would be:
> >
> > 1. (overengineered, not generically useful enough): If the transport
> > was a pipe (which already has the concept of different types of pipe
> > buffers) instead of a unix domain socket, we could introduce a special
> > representation for zero-filled pipe buffers and some API for receiving
> > zeroed holes through lseek(pipefd, 0, SEEK_DATA), but that's probably
> > not sufficiently useful for stuff other than core dumping to be worth
> > the effort.
> > 2. (somewhat overengineered) With some refactoring, we could maybe do
> > something like /proc/kcore and create a seekable virtual coredump file
> > that we send over the socket via SCM_RIGHTS?
> > 3. We could leave the userspace memory dump out of the core dump data,
> > and let userspace take care of filling out the memory contents using
> > /proc/$pid/pagemap and /proc/$pid/mem? That would also avoid task
> > switches and SKB allocations, and probably reduce the number of data
> > copies involved in this by 1.
> 
> I also wonder what userspace actually does with this data - does
> userspace just want to write it to disk, potentially after compressing
> it? Or does userspace actually do some fancy parsing of the data

That would be the most straightforward use-case, yes.

> stream to extract stack memory or something like that? Or does
> userspace buffer the whole thing into RAM and then process it from
> there (it kinda looks like systemd tries to do that but I might be
> reading this wrong)?

I'm not sure about that. I think it's writing it to disk and then
parsing it. But it also wranges it into a socket to forward to
containers or services.

Note that systemd has a pull request for the coredump socket up:

https://github.com/systemd/systemd/pull/43330

This should kill the usermodehelper soon on kernels that support the
socket.

> Anyway, I've looked through your code and it does look fine to me.

Thanks.