[PATCH net-next v2 0/2] net: a sockopt_t quirk for the options that write past optlen

Breno Leitao posted 2 patches 1 week, 3 days ago
include/linux/igmp.h   |  3 ++-
include/linux/net.h    | 32 ++++++++++++++++++++++++++++++++
net/ipv4/igmp.c        | 23 ++++++++++++++---------
net/ipv4/ip_sockglue.c | 10 +++++++++-
net/socket.c           |  4 ++--
5 files changed, 59 insertions(+), 13 deletions(-)
[PATCH net-next v2 0/2] net: a sockopt_t quirk for the options that write past optlen
Posted by Breno Leitao 1 week, 3 days ago
This series continues the migration of our protocols to sockopt_t, as
described in [1].

There are some protocols that use optlen as the header size, and the real
buffer size comes from a field inside the header. This means a bug, given
that optlen should be the full buffer size, but there are indications [2]
that we have programs that use the bad behaviour above, and we want to
avoid breaking them (or, honestly, avoid being cursed by Linus).

That said, create a quirk helper that preserves the same behaviour, even
using sockopt_t. The way to do it is simple:

1) Only do it for userspace callers (ubuf), otherwise a bug here will
   corrupt the kernel instead of a simple SIGSEGV.
2) Expand optval mid-air based on the header field.

IP_MSFILTER is the first user, and the smallest one: a single caller, no
compat variant, and a reply written front to back. MCAST_MSFILTER on ipv4
and ipv6 comes next, and TCP_AO_GET_KEYS has the same shape. Do this
quirk on IP_MSFILTER to make sure the dynamic is ok, so, we can expand
it later.

None of this is meant to change what userspace sees.

Link: https://lore.kernel.org/all/20260401-getsockopt-v2-0-611df6771aff@debian.org/ [1]
Link: https://lore.kernel.org/all/20260806-mcast_fix-v1-0-bed0a5518e57@debian.org/ [2]

Signed-off-by: Breno Leitao <leitao@debian.org>
---
Changes in v2:
- sockopt_expand_out() measures the request against opt->optlen instead
  of the iterator's remaining count, and asserts that nothing has been
  written through iter_out yet. Comparing against the remaining count
  made the verdict depend on how far a callback had already got, and a
  late call would rewind the write cursor to the head of optval.
- Cap the requested size at INT_MAX, since optlen and the getsockopt ABI
  are int.
- do_ip_getsockopt() writes optlen back only when ip_mc_msfget()
  succeeded. Without the guard, a read-only optlen turned -EINVAL,
  -ENODEV and -EADDRNOTAVAIL into -EFAULT.
- Name IP_MSFILTER in patch 1, document the WARN_ON_ONCE() and that no
  in-tree path reaches it, and mention sockptr_to_sockopt() losing its
  static.
- Link to v1: https://patch.msgid.link/20260910-getsockopt_phase6-v1-0-e681e102d5b8@debian.org

---
Breno Leitao (2):
      net: add sockopt_expand_out()
      ipv4: igmp: convert ip_mc_msfget() to sockopt_t

 include/linux/igmp.h   |  3 ++-
 include/linux/net.h    | 32 ++++++++++++++++++++++++++++++++
 net/ipv4/igmp.c        | 23 ++++++++++++++---------
 net/ipv4/ip_sockglue.c | 10 +++++++++-
 net/socket.c           |  4 ++--
 5 files changed, 59 insertions(+), 13 deletions(-)
---
base-commit: 548b86839f7fb819a4d6c83b71c73ec378d24275
change-id: 20260909-getsockopt_phase6-c7c96a21b062

Best regards,
--  
Breno Leitao <leitao@debian.org>
Re: [PATCH net-next v2 0/2] net: a sockopt_t quirk for the options that write past optlen
Posted by Stanislav Fomichev 1 week, 3 days ago
On 09/14, Breno Leitao wrote:
> This series continues the migration of our protocols to sockopt_t, as
> described in [1].
> 
> There are some protocols that use optlen as the header size, and the real
> buffer size comes from a field inside the header. This means a bug, given
> that optlen should be the full buffer size, but there are indications [2]
> that we have programs that use the bad behaviour above, and we want to
> avoid breaking them (or, honestly, avoid being cursed by Linus).
> 
> That said, create a quirk helper that preserves the same behaviour, even
> using sockopt_t. The way to do it is simple:
> 
> 1) Only do it for userspace callers (ubuf), otherwise a bug here will
>    corrupt the kernel instead of a simple SIGSEGV.
> 2) Expand optval mid-air based on the header field.
> 
> IP_MSFILTER is the first user, and the smallest one: a single caller, no
> compat variant, and a reply written front to back. MCAST_MSFILTER on ipv4
> and ipv6 comes next, and TCP_AO_GET_KEYS has the same shape. Do this
> quirk on IP_MSFILTER to make sure the dynamic is ok, so, we can expand
> it later.
> 
> None of this is meant to change what userspace sees.
> 
> Link: https://lore.kernel.org/all/20260401-getsockopt-v2-0-611df6771aff@debian.org/ [1]
> Link: https://lore.kernel.org/all/20260806-mcast_fix-v1-0-bed0a5518e57@debian.org/ [2]
> 
> Signed-off-by: Breno Leitao <leitao@debian.org>
> ---
> Changes in v2:
> - sockopt_expand_out() measures the request against opt->optlen instead
>   of the iterator's remaining count, and asserts that nothing has been
>   written through iter_out yet. Comparing against the remaining count
>   made the verdict depend on how far a callback had already got, and a
>   late call would rewind the write cursor to the head of optval.
> - Cap the requested size at INT_MAX, since optlen and the getsockopt ABI
>   are int.
> - do_ip_getsockopt() writes optlen back only when ip_mc_msfget()
>   succeeded. Without the guard, a read-only optlen turned -EINVAL,
>   -ENODEV and -EADDRNOTAVAIL into -EFAULT.
> - Name IP_MSFILTER in patch 1, document the WARN_ON_ONCE() and that no
>   in-tree path reaches it, and mention sockptr_to_sockopt() losing its
>   static.
> - Link to v1: https://patch.msgid.link/20260910-getsockopt_phase6-v1-0-e681e102d5b8@debian.org

Acked-by: Stanislav Fomichev <sdf@fomichev.me>
Re: [PATCH net-next v2 0/2] net: a sockopt_t quirk for the options that write past optlen
Posted by David Laight 1 week, 3 days ago
On Mon, 14 Sep 2026 05:20:06 -0700
Breno Leitao <leitao@debian.org> wrote:

...
> and ipv6 comes next, and TCP_AO_GET_KEYS has the same shape.

That is the sockopt I couldn't remember last time.

It is really horrid and shouldn't have been allowed.

David