include/linux/igmp.h | 3 ++- include/linux/net.h | 32 ++++++++++++++++++++++++++++++++ net/ipv4/igmp.c | 23 ++++++++++++++--------- net/ipv4/ip_sockglue.c | 10 +++++++++- net/socket.c | 4 ++-- 5 files changed, 59 insertions(+), 13 deletions(-)
This series continues the migration of our protocols to sockopt_t, as
described in [1].
There are some protocols that use optlen as the header size, and the real
buffer size comes from a field inside the header. This means a bug, given
that optlen should be the full buffer size, but there are indications [2]
that we have programs that use the bad behaviour above, and we want to
avoid breaking them (or, honestly, avoid being cursed by Linus).
That said, create a quirk helper that preserves the same behaviour, even
using sockopt_t. The way to do it is simple:
1) Only do it for userspace callers (ubuf), otherwise a bug here will
corrupt the kernel instead of a simple SIGSEGV.
2) Expand optval mid-air based on the header field.
IP_MSFILTER is the first user, and the smallest one: a single caller, no
compat variant, and a reply written front to back. MCAST_MSFILTER on ipv4
and ipv6 comes next, and TCP_AO_GET_KEYS has the same shape. Do this
quirk on IP_MSFILTER to make sure the dynamic is ok, so, we can expand
it later.
None of this is meant to change what userspace sees.
Link: https://lore.kernel.org/all/20260401-getsockopt-v2-0-611df6771aff@debian.org/ [1]
Link: https://lore.kernel.org/all/20260806-mcast_fix-v1-0-bed0a5518e57@debian.org/ [2]
Signed-off-by: Breno Leitao <leitao@debian.org>
---
Changes in v2:
- sockopt_expand_out() measures the request against opt->optlen instead
of the iterator's remaining count, and asserts that nothing has been
written through iter_out yet. Comparing against the remaining count
made the verdict depend on how far a callback had already got, and a
late call would rewind the write cursor to the head of optval.
- Cap the requested size at INT_MAX, since optlen and the getsockopt ABI
are int.
- do_ip_getsockopt() writes optlen back only when ip_mc_msfget()
succeeded. Without the guard, a read-only optlen turned -EINVAL,
-ENODEV and -EADDRNOTAVAIL into -EFAULT.
- Name IP_MSFILTER in patch 1, document the WARN_ON_ONCE() and that no
in-tree path reaches it, and mention sockptr_to_sockopt() losing its
static.
- Link to v1: https://patch.msgid.link/20260910-getsockopt_phase6-v1-0-e681e102d5b8@debian.org
---
Breno Leitao (2):
net: add sockopt_expand_out()
ipv4: igmp: convert ip_mc_msfget() to sockopt_t
include/linux/igmp.h | 3 ++-
include/linux/net.h | 32 ++++++++++++++++++++++++++++++++
net/ipv4/igmp.c | 23 ++++++++++++++---------
net/ipv4/ip_sockglue.c | 10 +++++++++-
net/socket.c | 4 ++--
5 files changed, 59 insertions(+), 13 deletions(-)
---
base-commit: 548b86839f7fb819a4d6c83b71c73ec378d24275
change-id: 20260909-getsockopt_phase6-c7c96a21b062
Best regards,
--
Breno Leitao <leitao@debian.org>
On 09/14, Breno Leitao wrote: > This series continues the migration of our protocols to sockopt_t, as > described in [1]. > > There are some protocols that use optlen as the header size, and the real > buffer size comes from a field inside the header. This means a bug, given > that optlen should be the full buffer size, but there are indications [2] > that we have programs that use the bad behaviour above, and we want to > avoid breaking them (or, honestly, avoid being cursed by Linus). > > That said, create a quirk helper that preserves the same behaviour, even > using sockopt_t. The way to do it is simple: > > 1) Only do it for userspace callers (ubuf), otherwise a bug here will > corrupt the kernel instead of a simple SIGSEGV. > 2) Expand optval mid-air based on the header field. > > IP_MSFILTER is the first user, and the smallest one: a single caller, no > compat variant, and a reply written front to back. MCAST_MSFILTER on ipv4 > and ipv6 comes next, and TCP_AO_GET_KEYS has the same shape. Do this > quirk on IP_MSFILTER to make sure the dynamic is ok, so, we can expand > it later. > > None of this is meant to change what userspace sees. > > Link: https://lore.kernel.org/all/20260401-getsockopt-v2-0-611df6771aff@debian.org/ [1] > Link: https://lore.kernel.org/all/20260806-mcast_fix-v1-0-bed0a5518e57@debian.org/ [2] > > Signed-off-by: Breno Leitao <leitao@debian.org> > --- > Changes in v2: > - sockopt_expand_out() measures the request against opt->optlen instead > of the iterator's remaining count, and asserts that nothing has been > written through iter_out yet. Comparing against the remaining count > made the verdict depend on how far a callback had already got, and a > late call would rewind the write cursor to the head of optval. > - Cap the requested size at INT_MAX, since optlen and the getsockopt ABI > are int. > - do_ip_getsockopt() writes optlen back only when ip_mc_msfget() > succeeded. Without the guard, a read-only optlen turned -EINVAL, > -ENODEV and -EADDRNOTAVAIL into -EFAULT. > - Name IP_MSFILTER in patch 1, document the WARN_ON_ONCE() and that no > in-tree path reaches it, and mention sockptr_to_sockopt() losing its > static. > - Link to v1: https://patch.msgid.link/20260910-getsockopt_phase6-v1-0-e681e102d5b8@debian.org Acked-by: Stanislav Fomichev <sdf@fomichev.me>
On Mon, 14 Sep 2026 05:20:06 -0700 Breno Leitao <leitao@debian.org> wrote: ... > and ipv6 comes next, and TCP_AO_GET_KEYS has the same shape. That is the sockopt I couldn't remember last time. It is really horrid and shouldn't have been allowed. David
© 2016 - 2026 Red Hat, Inc.