[PATCH bpf-next v4 0/2] bpf: htab: Reduce memory use of hash maps

T.J. Mercier posted 2 patches 1 month, 2 weeks ago
There is a newer version of this series
kernel/bpf/hashtab.c                          | 418 ++++++++++++------
kernel/bpf/map_in_map.c                       |  13 +
kernel/bpf/map_in_map.h                       |   2 +
.../selftests/bpf/progs/map_ptr_kern.c        |   2 +-
4 files changed, 289 insertions(+), 146 deletions(-)
[PATCH bpf-next v4 0/2] bpf: htab: Reduce memory use of hash maps
Posted by T.J. Mercier 1 month, 2 weeks ago
Memory is expensive and scarce these days. This series reduces the
memory use of BPF hash maps by eliminating the per-element overheads
below. This saves up to 50% of per-element memory use for standard and
PCPU hash maps. The memory use of LRU hash maps is unaffected.

Map Type & Configuration            | Old size | New size | Savings
------------------------------------|----------|----------|--------
Standard (key <= 8 B, val <= 8 B)   |   64 B   |   32 B   | 50.0%
Per-CPU (prealloc) (key <= 8 B)     |   64 B   |   32 B   | 50.0%
Per-CPU (non-prealloc) (key <= 8 B) |   64 B   |   40 B   | 37.5%
LRU (Any key/value size)            |    -     |    -     | 00.0%

1) Unused LRU / PCPU fields in standard and PCPU hash maps (patch 1)
struct htab_elem is used for all hash map types, and includes fields
that are not always used (bpf_lru_node, ptr_to_pptr). For standard
(non-LRU, non-PCPU) hash maps the 24 bytes for the bpf_lru_node (union)
are entirely overhead and can be eliminated. Non-preallocated PCPU maps
only need the 8 byte ptr_to_pptr which is currently unioned with the
unneeded 24 byte bpf_lru_node, so 16 bytes of overhead can be
eliminated. Preallocated PCPU maps don't need ptr_to_pptr, so 24 bytes
of overhead can be saved.

2) Hash caching for small keys (patch 2)
For hash maps with small key sizes (<= word size), comparing keys only
requires a single instruction. Currently the 4 byte hash value (8 byte
aligned and padded) is used for this, but offers no performance
advantage in this case and can be eliminated.

The implementation splits htab_elem into dedicated structures for the
different map types (htab_elem_lru, htab_elem_pcpu, htab_elem) which
share a common initial sequence (struct htab_node), but contain
additional map-type specific fields where necessary. This means the
placement of the key for each element varies with the map type, and
key_offset is added to bpf_htab for this purpose.

While using key_offset and conditional hash checks adds new pointer
dereferences and branching during element traversal,
run_bench_htab_mem.sh shows no significant performance regression across
10 runs on my 3995WX.

Benchmark (all in kops/sec)  | Avg. Before | StDev | Avg. After | StDev
-----------------------------|-------------|-------|------------|------
prealloc overwrite           | 115.11      | 4.10  | 115.45     | 5.24
prealloc batch_add_batch_del | 127.14      | 4.32  | 127.06     | 2.32
prealloc add_del_on_diff_cpu | 23.22       | 0.93  | 22.91      | 1.60
normal overwrite             | 78.52       | 3.05  | 80.40      | 3.25
normal batch_add_batch_del   | 45.37       | 0.69  | 47.71      | 0.66
normal add_del_on_diff_cpu   | 12.02       | 0.73  | 12.48      | 0.70

---
Changes in v4:
Removed inline from new functions per BPF CI (netdev/source_inline).

From Mykyta Yatsenko:
Factor out duplicate lookup_elem code into __lookup_elem_raw.
Use offsetof instead of sizeof for key_offset assignments in
htab_map_alloc (patch 1).
Eliminate branching and htab_elem casting in htab_elem_hash /
htab_elem_set_hash.

Changes in v3:
From Sashiko on torn reads/writes:
Use a local unsigned long and READ_ONCE / WRITE_ONCE instead of memcmp /
memcpy for atomic key comparisons for hashless elements.

Changes in v2:
Make maximum key_size for !has_hash depend on word size for atomicity
on 32-bit.

From Mykyta Yatsenko:
Put the htab_elem* common initial sequence in its own struct (htab_node)
and reuse it across all element types that share it. Eliminate
associated BUILD_BUG_ON additions.
Replace both the hash and key fields with data[].
Store has_hash in struct bpf_htab, and avoid per-element reads of it.# Please edit the description for the branch

T.J. Mercier (2):
  bpf: htab: Split htab_elem_lru and htab_elem_pcpu off of htab_elem
  bpf: htab: Reduce elem_size by 8 bytes for small key sizes

 kernel/bpf/hashtab.c                          | 418 ++++++++++++------
 kernel/bpf/map_in_map.c                       |  13 +
 kernel/bpf/map_in_map.h                       |   2 +
 .../selftests/bpf/progs/map_ptr_kern.c        |   2 +-
 4 files changed, 289 insertions(+), 146 deletions(-)


base-commit: cfce77b63375dac81d53f2f85593c548415206b7
-- 
2.55.0.691.gc56d675ccc-goog
Re: [PATCH bpf-next v4 0/2] bpf: htab: Reduce memory use of hash maps
Posted by Mykyta Yatsenko 1 week, 4 days ago
On 8/13/26 12:19 AM, T.J. Mercier wrote:
> Memory is expensive and scarce these days. This series reduces the
> memory use of BPF hash maps by eliminating the per-element overheads
> below. This saves up to 50% of per-element memory use for standard and
> PCPU hash maps. The memory use of LRU hash maps is unaffected.
> 
> Map Type & Configuration            | Old size | New size | Savings
> ------------------------------------|----------|----------|--------
> Standard (key <= 8 B, val <= 8 B)   |   64 B   |   32 B   | 50.0%
> Per-CPU (prealloc) (key <= 8 B)     |   64 B   |   32 B   | 50.0%
> Per-CPU (non-prealloc) (key <= 8 B) |   64 B   |   40 B   | 37.5%
> LRU (Any key/value size)            |    -     |    -     | 00.0%
> 

T.J. are you still interested landing this? Maybe respin the series?
Alexei was away back when you sent this.

> 1) Unused LRU / PCPU fields in standard and PCPU hash maps (patch 1)
> struct htab_elem is used for all hash map types, and includes fields
> that are not always used (bpf_lru_node, ptr_to_pptr). For standard
> (non-LRU, non-PCPU) hash maps the 24 bytes for the bpf_lru_node (union)
> are entirely overhead and can be eliminated. Non-preallocated PCPU maps
> only need the 8 byte ptr_to_pptr which is currently unioned with the
> unneeded 24 byte bpf_lru_node, so 16 bytes of overhead can be
> eliminated. Preallocated PCPU maps don't need ptr_to_pptr, so 24 bytes
> of overhead can be saved.
> 
> 2) Hash caching for small keys (patch 2)
> For hash maps with small key sizes (<= word size), comparing keys only
> requires a single instruction. Currently the 4 byte hash value (8 byte
> aligned and padded) is used for this, but offers no performance
> advantage in this case and can be eliminated.
> 
> The implementation splits htab_elem into dedicated structures for the
> different map types (htab_elem_lru, htab_elem_pcpu, htab_elem) which
> share a common initial sequence (struct htab_node), but contain
> additional map-type specific fields where necessary. This means the
> placement of the key for each element varies with the map type, and
> key_offset is added to bpf_htab for this purpose.
> 
> While using key_offset and conditional hash checks adds new pointer
> dereferences and branching during element traversal,
> run_bench_htab_mem.sh shows no significant performance regression across
> 10 runs on my 3995WX.
> 
> Benchmark (all in kops/sec)  | Avg. Before | StDev | Avg. After | StDev
> -----------------------------|-------------|-------|------------|------
> prealloc overwrite           | 115.11      | 4.10  | 115.45     | 5.24
> prealloc batch_add_batch_del | 127.14      | 4.32  | 127.06     | 2.32
> prealloc add_del_on_diff_cpu | 23.22       | 0.93  | 22.91      | 1.60
> normal overwrite             | 78.52       | 3.05  | 80.40      | 3.25
> normal batch_add_batch_del   | 45.37       | 0.69  | 47.71      | 0.66
> normal add_del_on_diff_cpu   | 12.02       | 0.73  | 12.48      | 0.70
> 
> ---
> Changes in v4:
> Removed inline from new functions per BPF CI (netdev/source_inline).
> 
> From Mykyta Yatsenko:
> Factor out duplicate lookup_elem code into __lookup_elem_raw.
> Use offsetof instead of sizeof for key_offset assignments in
> htab_map_alloc (patch 1).
> Eliminate branching and htab_elem casting in htab_elem_hash /
> htab_elem_set_hash.
> 
> Changes in v3:
> From Sashiko on torn reads/writes:
> Use a local unsigned long and READ_ONCE / WRITE_ONCE instead of memcmp /
> memcpy for atomic key comparisons for hashless elements.
> 
> Changes in v2:
> Make maximum key_size for !has_hash depend on word size for atomicity
> on 32-bit.
> 
> From Mykyta Yatsenko:
> Put the htab_elem* common initial sequence in its own struct (htab_node)
> and reuse it across all element types that share it. Eliminate
> associated BUILD_BUG_ON additions.
> Replace both the hash and key fields with data[].
> Store has_hash in struct bpf_htab, and avoid per-element reads of it.# Please edit the description for the branch
> 
> T.J. Mercier (2):
>   bpf: htab: Split htab_elem_lru and htab_elem_pcpu off of htab_elem
>   bpf: htab: Reduce elem_size by 8 bytes for small key sizes
> 
>  kernel/bpf/hashtab.c                          | 418 ++++++++++++------
>  kernel/bpf/map_in_map.c                       |  13 +
>  kernel/bpf/map_in_map.h                       |   2 +
>  .../selftests/bpf/progs/map_ptr_kern.c        |   2 +-
>  4 files changed, 289 insertions(+), 146 deletions(-)
> 
> 
> base-commit: cfce77b63375dac81d53f2f85593c548415206b7
Re: [PATCH bpf-next v4 0/2] bpf: htab: Reduce memory use of hash maps
Posted by T.J. Mercier 1 week, 4 days ago
On Thu, Sep 17, 2026 at 9:10 AM Mykyta Yatsenko
<mykyta.yatsenko5@gmail.com> wrote:
>
> On 8/13/26 12:19 AM, T.J. Mercier wrote:
> > Memory is expensive and scarce these days. This series reduces the
> > memory use of BPF hash maps by eliminating the per-element overheads
> > below. This saves up to 50% of per-element memory use for standard and
> > PCPU hash maps. The memory use of LRU hash maps is unaffected.
> >
> > Map Type & Configuration            | Old size | New size | Savings
> > ------------------------------------|----------|----------|--------
> > Standard (key <= 8 B, val <= 8 B)   |   64 B   |   32 B   | 50.0%
> > Per-CPU (prealloc) (key <= 8 B)     |   64 B   |   32 B   | 50.0%
> > Per-CPU (non-prealloc) (key <= 8 B) |   64 B   |   40 B   | 37.5%
> > LRU (Any key/value size)            |    -     |    -     | 00.0%
> >
>
> T.J. are you still interested landing this? Maybe respin the series?
> Alexei was away back when you sent this.

Hi Mykyta, thanks for following up. Yes, I'd still like to land it. I
took a break during the merge window which coincided with a vacation,
where I broke my collarbone on a mountain bike jump gone wrong. Then
surgery and recovery, and I'm still catching up on everything from
while I was out. I plan to rebase this and send it out again before I
head out for pre-LPC travel next week.


> > 1) Unused LRU / PCPU fields in standard and PCPU hash maps (patch 1)
> > struct htab_elem is used for all hash map types, and includes fields
> > that are not always used (bpf_lru_node, ptr_to_pptr). For standard
> > (non-LRU, non-PCPU) hash maps the 24 bytes for the bpf_lru_node (union)
> > are entirely overhead and can be eliminated. Non-preallocated PCPU maps
> > only need the 8 byte ptr_to_pptr which is currently unioned with the
> > unneeded 24 byte bpf_lru_node, so 16 bytes of overhead can be
> > eliminated. Preallocated PCPU maps don't need ptr_to_pptr, so 24 bytes
> > of overhead can be saved.
> >
> > 2) Hash caching for small keys (patch 2)
> > For hash maps with small key sizes (<= word size), comparing keys only
> > requires a single instruction. Currently the 4 byte hash value (8 byte
> > aligned and padded) is used for this, but offers no performance
> > advantage in this case and can be eliminated.
> >
> > The implementation splits htab_elem into dedicated structures for the
> > different map types (htab_elem_lru, htab_elem_pcpu, htab_elem) which
> > share a common initial sequence (struct htab_node), but contain
> > additional map-type specific fields where necessary. This means the
> > placement of the key for each element varies with the map type, and
> > key_offset is added to bpf_htab for this purpose.
> >
> > While using key_offset and conditional hash checks adds new pointer
> > dereferences and branching during element traversal,
> > run_bench_htab_mem.sh shows no significant performance regression across
> > 10 runs on my 3995WX.
> >
> > Benchmark (all in kops/sec)  | Avg. Before | StDev | Avg. After | StDev
> > -----------------------------|-------------|-------|------------|------
> > prealloc overwrite           | 115.11      | 4.10  | 115.45     | 5.24
> > prealloc batch_add_batch_del | 127.14      | 4.32  | 127.06     | 2.32
> > prealloc add_del_on_diff_cpu | 23.22       | 0.93  | 22.91      | 1.60
> > normal overwrite             | 78.52       | 3.05  | 80.40      | 3.25
> > normal batch_add_batch_del   | 45.37       | 0.69  | 47.71      | 0.66
> > normal add_del_on_diff_cpu   | 12.02       | 0.73  | 12.48      | 0.70
> >
> > ---
> > Changes in v4:
> > Removed inline from new functions per BPF CI (netdev/source_inline).
> >
> > From Mykyta Yatsenko:
> > Factor out duplicate lookup_elem code into __lookup_elem_raw.
> > Use offsetof instead of sizeof for key_offset assignments in
> > htab_map_alloc (patch 1).
> > Eliminate branching and htab_elem casting in htab_elem_hash /
> > htab_elem_set_hash.
> >
> > Changes in v3:
> > From Sashiko on torn reads/writes:
> > Use a local unsigned long and READ_ONCE / WRITE_ONCE instead of memcmp /
> > memcpy for atomic key comparisons for hashless elements.
> >
> > Changes in v2:
> > Make maximum key_size for !has_hash depend on word size for atomicity
> > on 32-bit.
> >
> > From Mykyta Yatsenko:
> > Put the htab_elem* common initial sequence in its own struct (htab_node)
> > and reuse it across all element types that share it. Eliminate
> > associated BUILD_BUG_ON additions.
> > Replace both the hash and key fields with data[].
> > Store has_hash in struct bpf_htab, and avoid per-element reads of it.# Please edit the description for the branch
> >
> > T.J. Mercier (2):
> >   bpf: htab: Split htab_elem_lru and htab_elem_pcpu off of htab_elem
> >   bpf: htab: Reduce elem_size by 8 bytes for small key sizes
> >
> >  kernel/bpf/hashtab.c                          | 418 ++++++++++++------
> >  kernel/bpf/map_in_map.c                       |  13 +
> >  kernel/bpf/map_in_map.h                       |   2 +
> >  .../selftests/bpf/progs/map_ptr_kern.c        |   2 +-
> >  4 files changed, 289 insertions(+), 146 deletions(-)
> >
> >
> > base-commit: cfce77b63375dac81d53f2f85593c548415206b7
>
Re: [PATCH bpf-next v4 0/2] bpf: htab: Reduce memory use of hash maps
Posted by Andrii Nakryiko 1 week, 4 days ago
On Thu, Sep 17, 2026 at 9:31 AM T.J. Mercier <tjmercier@google.com> wrote:
>
> On Thu, Sep 17, 2026 at 9:10 AM Mykyta Yatsenko
> <mykyta.yatsenko5@gmail.com> wrote:
> >
> > On 8/13/26 12:19 AM, T.J. Mercier wrote:
> > > Memory is expensive and scarce these days. This series reduces the
> > > memory use of BPF hash maps by eliminating the per-element overheads
> > > below. This saves up to 50% of per-element memory use for standard and
> > > PCPU hash maps. The memory use of LRU hash maps is unaffected.
> > >
> > > Map Type & Configuration            | Old size | New size | Savings
> > > ------------------------------------|----------|----------|--------
> > > Standard (key <= 8 B, val <= 8 B)   |   64 B   |   32 B   | 50.0%
> > > Per-CPU (prealloc) (key <= 8 B)     |   64 B   |   32 B   | 50.0%
> > > Per-CPU (non-prealloc) (key <= 8 B) |   64 B   |   40 B   | 37.5%
> > > LRU (Any key/value size)            |    -     |    -     | 00.0%
> > >
> >
> > T.J. are you still interested landing this? Maybe respin the series?
> > Alexei was away back when you sent this.
>
> Hi Mykyta, thanks for following up. Yes, I'd still like to land it. I
> took a break during the merge window which coincided with a vacation,
> where I broke my collarbone on a mountain bike jump gone wrong. Then
> surgery and recovery, and I'm still catching up on everything from

oh, wow, hope you'll heal fast and well!

> while I was out. I plan to rebase this and send it out again before I
> head out for pre-LPC travel next week.

have you considered splitting lru and non-lru flavors of hashtable
before doing this optimization? I'm wondering if it will allow to
clean up some parts of it, while also making map struct itself smaller
for non-lru map (there is that LRU-specific piece in the union which
artificially blows up the size of any hash map).

I am a bit concerned about that key offset, even if the microbenchmark
doesn't show much difference. What if we put hash itself before
per-element header, so that key/value are always at the same offset.
For cases where key size > 8 we'll just access hash at
addrof(htab_elem) - 8, while smaller key sizes will just directly
compare keys.

Just some high level thoughts, haven't really coded any of that, so
hard to tell upfront if it's worth doing.

>
>
> > > 1) Unused LRU / PCPU fields in standard and PCPU hash maps (patch 1)
> > > struct htab_elem is used for all hash map types, and includes fields
> > > that are not always used (bpf_lru_node, ptr_to_pptr). For standard
> > > (non-LRU, non-PCPU) hash maps the 24 bytes for the bpf_lru_node (union)
> > > are entirely overhead and can be eliminated. Non-preallocated PCPU maps
> > > only need the 8 byte ptr_to_pptr which is currently unioned with the
> > > unneeded 24 byte bpf_lru_node, so 16 bytes of overhead can be
> > > eliminated. Preallocated PCPU maps don't need ptr_to_pptr, so 24 bytes
> > > of overhead can be saved.
> > >
> > > 2) Hash caching for small keys (patch 2)
> > > For hash maps with small key sizes (<= word size), comparing keys only
> > > requires a single instruction. Currently the 4 byte hash value (8 byte
> > > aligned and padded) is used for this, but offers no performance
> > > advantage in this case and can be eliminated.
> > >
> > > The implementation splits htab_elem into dedicated structures for the
> > > different map types (htab_elem_lru, htab_elem_pcpu, htab_elem) which
> > > share a common initial sequence (struct htab_node), but contain
> > > additional map-type specific fields where necessary. This means the
> > > placement of the key for each element varies with the map type, and
> > > key_offset is added to bpf_htab for this purpose.
> > >
> > > While using key_offset and conditional hash checks adds new pointer
> > > dereferences and branching during element traversal,
> > > run_bench_htab_mem.sh shows no significant performance regression across
> > > 10 runs on my 3995WX.
> > >
> > > Benchmark (all in kops/sec)  | Avg. Before | StDev | Avg. After | StDev
> > > -----------------------------|-------------|-------|------------|------
> > > prealloc overwrite           | 115.11      | 4.10  | 115.45     | 5.24
> > > prealloc batch_add_batch_del | 127.14      | 4.32  | 127.06     | 2.32
> > > prealloc add_del_on_diff_cpu | 23.22       | 0.93  | 22.91      | 1.60
> > > normal overwrite             | 78.52       | 3.05  | 80.40      | 3.25
> > > normal batch_add_batch_del   | 45.37       | 0.69  | 47.71      | 0.66
> > > normal add_del_on_diff_cpu   | 12.02       | 0.73  | 12.48      | 0.70
> > >
> > > ---
> > > Changes in v4:
> > > Removed inline from new functions per BPF CI (netdev/source_inline).
> > >
> > > From Mykyta Yatsenko:
> > > Factor out duplicate lookup_elem code into __lookup_elem_raw.
> > > Use offsetof instead of sizeof for key_offset assignments in
> > > htab_map_alloc (patch 1).
> > > Eliminate branching and htab_elem casting in htab_elem_hash /
> > > htab_elem_set_hash.
> > >
> > > Changes in v3:
> > > From Sashiko on torn reads/writes:
> > > Use a local unsigned long and READ_ONCE / WRITE_ONCE instead of memcmp /
> > > memcpy for atomic key comparisons for hashless elements.
> > >
> > > Changes in v2:
> > > Make maximum key_size for !has_hash depend on word size for atomicity
> > > on 32-bit.
> > >
> > > From Mykyta Yatsenko:
> > > Put the htab_elem* common initial sequence in its own struct (htab_node)
> > > and reuse it across all element types that share it. Eliminate
> > > associated BUILD_BUG_ON additions.
> > > Replace both the hash and key fields with data[].
> > > Store has_hash in struct bpf_htab, and avoid per-element reads of it.# Please edit the description for the branch
> > >
> > > T.J. Mercier (2):
> > >   bpf: htab: Split htab_elem_lru and htab_elem_pcpu off of htab_elem
> > >   bpf: htab: Reduce elem_size by 8 bytes for small key sizes
> > >
> > >  kernel/bpf/hashtab.c                          | 418 ++++++++++++------
> > >  kernel/bpf/map_in_map.c                       |  13 +
> > >  kernel/bpf/map_in_map.h                       |   2 +
> > >  .../selftests/bpf/progs/map_ptr_kern.c        |   2 +-
> > >  4 files changed, 289 insertions(+), 146 deletions(-)
> > >
> > >
> > > base-commit: cfce77b63375dac81d53f2f85593c548415206b7
> >
Re: [PATCH bpf-next v4 0/2] bpf: htab: Reduce memory use of hash maps
Posted by T.J. Mercier 1 week, 4 days ago
On Thu, Sep 17, 2026 at 2:21 PM Andrii Nakryiko
<andrii.nakryiko@gmail.com> wrote:
>
> On Thu, Sep 17, 2026 at 9:31 AM T.J. Mercier <tjmercier@google.com> wrote:
> >
> > On Thu, Sep 17, 2026 at 9:10 AM Mykyta Yatsenko
> > <mykyta.yatsenko5@gmail.com> wrote:
> > >
> > > On 8/13/26 12:19 AM, T.J. Mercier wrote:
> > > > Memory is expensive and scarce these days. This series reduces the
> > > > memory use of BPF hash maps by eliminating the per-element overheads
> > > > below. This saves up to 50% of per-element memory use for standard and
> > > > PCPU hash maps. The memory use of LRU hash maps is unaffected.
> > > >
> > > > Map Type & Configuration            | Old size | New size | Savings
> > > > ------------------------------------|----------|----------|--------
> > > > Standard (key <= 8 B, val <= 8 B)   |   64 B   |   32 B   | 50.0%
> > > > Per-CPU (prealloc) (key <= 8 B)     |   64 B   |   32 B   | 50.0%
> > > > Per-CPU (non-prealloc) (key <= 8 B) |   64 B   |   40 B   | 37.5%
> > > > LRU (Any key/value size)            |    -     |    -     | 00.0%
> > > >
> > >
> > > T.J. are you still interested landing this? Maybe respin the series?
> > > Alexei was away back when you sent this.
> >
> > Hi Mykyta, thanks for following up. Yes, I'd still like to land it. I
> > took a break during the merge window which coincided with a vacation,
> > where I broke my collarbone on a mountain bike jump gone wrong. Then
> > surgery and recovery, and I'm still catching up on everything from
>
> oh, wow, hope you'll heal fast and well!

Thanks! The surgery already made a huge difference, so now I'm waiting
for bone to grow and it'll be a few months before I'm back on a bike.

> > while I was out. I plan to rebase this and send it out again before I
> > head out for pre-LPC travel next week.
>
> have you considered splitting lru and non-lru flavors of hashtable
> before doing this optimization? I'm wondering if it will allow to
> clean up some parts of it, while also making map struct itself smaller
> for non-lru map (there is that LRU-specific piece in the union which
> artificially blows up the size of any hash map).
>
> I am a bit concerned about that key offset, even if the microbenchmark
> doesn't show much difference. What if we put hash itself before
> per-element header, so that key/value are always at the same offset.
> For cases where key size > 8 we'll just access hash at
> addrof(htab_elem) - 8, while smaller key sizes will just directly
> compare keys.

This is an interesting idea and I think it will work. Let me try this
first, and then I will take a look at splitting out LRU hashtables
aftewards.

> Just some high level thoughts, haven't really coded any of that, so
> hard to tell upfront if it's worth doing.
>
> >
> >
> > > > 1) Unused LRU / PCPU fields in standard and PCPU hash maps (patch 1)
> > > > struct htab_elem is used for all hash map types, and includes fields
> > > > that are not always used (bpf_lru_node, ptr_to_pptr). For standard
> > > > (non-LRU, non-PCPU) hash maps the 24 bytes for the bpf_lru_node (union)
> > > > are entirely overhead and can be eliminated. Non-preallocated PCPU maps
> > > > only need the 8 byte ptr_to_pptr which is currently unioned with the
> > > > unneeded 24 byte bpf_lru_node, so 16 bytes of overhead can be
> > > > eliminated. Preallocated PCPU maps don't need ptr_to_pptr, so 24 bytes
> > > > of overhead can be saved.
> > > >
> > > > 2) Hash caching for small keys (patch 2)
> > > > For hash maps with small key sizes (<= word size), comparing keys only
> > > > requires a single instruction. Currently the 4 byte hash value (8 byte
> > > > aligned and padded) is used for this, but offers no performance
> > > > advantage in this case and can be eliminated.
> > > >
> > > > The implementation splits htab_elem into dedicated structures for the
> > > > different map types (htab_elem_lru, htab_elem_pcpu, htab_elem) which
> > > > share a common initial sequence (struct htab_node), but contain
> > > > additional map-type specific fields where necessary. This means the
> > > > placement of the key for each element varies with the map type, and
> > > > key_offset is added to bpf_htab for this purpose.
> > > >
> > > > While using key_offset and conditional hash checks adds new pointer
> > > > dereferences and branching during element traversal,
> > > > run_bench_htab_mem.sh shows no significant performance regression across
> > > > 10 runs on my 3995WX.
> > > >
> > > > Benchmark (all in kops/sec)  | Avg. Before | StDev | Avg. After | StDev
> > > > -----------------------------|-------------|-------|------------|------
> > > > prealloc overwrite           | 115.11      | 4.10  | 115.45     | 5.24
> > > > prealloc batch_add_batch_del | 127.14      | 4.32  | 127.06     | 2.32
> > > > prealloc add_del_on_diff_cpu | 23.22       | 0.93  | 22.91      | 1.60
> > > > normal overwrite             | 78.52       | 3.05  | 80.40      | 3.25
> > > > normal batch_add_batch_del   | 45.37       | 0.69  | 47.71      | 0.66
> > > > normal add_del_on_diff_cpu   | 12.02       | 0.73  | 12.48      | 0.70
> > > >
> > > > ---
> > > > Changes in v4:
> > > > Removed inline from new functions per BPF CI (netdev/source_inline).
> > > >
> > > > From Mykyta Yatsenko:
> > > > Factor out duplicate lookup_elem code into __lookup_elem_raw.
> > > > Use offsetof instead of sizeof for key_offset assignments in
> > > > htab_map_alloc (patch 1).
> > > > Eliminate branching and htab_elem casting in htab_elem_hash /
> > > > htab_elem_set_hash.
> > > >
> > > > Changes in v3:
> > > > From Sashiko on torn reads/writes:
> > > > Use a local unsigned long and READ_ONCE / WRITE_ONCE instead of memcmp /
> > > > memcpy for atomic key comparisons for hashless elements.
> > > >
> > > > Changes in v2:
> > > > Make maximum key_size for !has_hash depend on word size for atomicity
> > > > on 32-bit.
> > > >
> > > > From Mykyta Yatsenko:
> > > > Put the htab_elem* common initial sequence in its own struct (htab_node)
> > > > and reuse it across all element types that share it. Eliminate
> > > > associated BUILD_BUG_ON additions.
> > > > Replace both the hash and key fields with data[].
> > > > Store has_hash in struct bpf_htab, and avoid per-element reads of it.# Please edit the description for the branch
> > > >
> > > > T.J. Mercier (2):
> > > >   bpf: htab: Split htab_elem_lru and htab_elem_pcpu off of htab_elem
> > > >   bpf: htab: Reduce elem_size by 8 bytes for small key sizes
> > > >
> > > >  kernel/bpf/hashtab.c                          | 418 ++++++++++++------
> > > >  kernel/bpf/map_in_map.c                       |  13 +
> > > >  kernel/bpf/map_in_map.h                       |   2 +
> > > >  .../selftests/bpf/progs/map_ptr_kern.c        |   2 +-
> > > >  4 files changed, 289 insertions(+), 146 deletions(-)
> > > >
> > > >
> > > > base-commit: cfce77b63375dac81d53f2f85593c548415206b7
> > >