[PATCH v3 0/2] LoongArch: BPF: per-CPU address resolution and timed may_goto

George Guo posted 2 patches 1 month, 4 weeks ago
arch/loongarch/include/asm/inst.h       |  1 +
arch/loongarch/net/Makefile             |  2 +-
arch/loongarch/net/bpf_jit.c            | 27 +++++++++++++-
arch/loongarch/net/bpf_timed_may_goto.S | 47 +++++++++++++++++++++++++
4 files changed, 75 insertions(+), 2 deletions(-)
create mode 100644 arch/loongarch/net/bpf_timed_may_goto.S
[PATCH v3 0/2] LoongArch: BPF: per-CPU address resolution and timed may_goto
Posted by George Guo 1 month, 4 weeks ago
These are the first two patches of the LoongArch BPF JIT feature work,
sent on their own to make review easier. The remaining pieces
(per-program private stacks, exceptions/bpf_throw, sign-extending loads
and atomics on arena pointers, and the matching selftests) will follow in
a later series.

Both patches are independent of each other. They apply on top of Chenguang
Zhao's "LoongArch bpf kptr xchg inline support" v4 series:

  https://lore.kernel.org/all/20260729022837.355549-1-chenguang.zhao@linux.dev/

Patch 1 adds the internal-only BPF MOV that resolves a per-CPU offset to
the current CPU's address, advertised through
bpf_jit_supports_percpu_insn(). LoongArch keeps the current CPU's per-CPU
base in $r21 (__my_cpu_offset), so the resolution is a single add.
Exercised by the cpumask and percpu_alloc selftests.

Patch 2 implements arch_bpf_timed_may_goto() and advertises it, so the
verifier lowers may_goto to its timed variant, which bounds the loop by a
wall-clock timeout kept in a per-loop stack slot. It uses a custom calling
convention: the count/timestamp stack offset is passed in BPF_REG_AX and
returned there, so the JIT call path skips the usual "BPF_REG_0 = return
value" move for this helper. Exercised by the iters selftests.

Selftest results (all PASS):

  Feature                          Test(s)
  --------------------------------------------------------------
  1  internal-only MOV (percpu)     cpumask, percpu_alloc
  2  timed may_goto                 iters

$ sudo ./test_progs -t cpumask
...
Summary: 1/35 PASSED, 0 SKIPPED, 0 FAILED

$ sudo ./test_progs -t percpu_alloc
...
Summary: 1/18 PASSED, 0 SKIPPED, 0 FAILED

$ sudo ./test_progs -t iters
...
Summary: 1/93 PASSED, 0 SKIPPED, 0 FAILED

Based on loongarch-next:
  https://git.kernel.org/pub/scm/linux/kernel/git/chenhuacai/linux-loongson.git/log/?h=loongarch-next

These two patches are split from the full v2 series (11 patches):
  https://lore.kernel.org/all/20260702022322.51033-1-dongtai.guo@linux.dev/

George Guo (2):
  LoongArch: BPF: Support internal-only MOV to resolve per-CPU addrs
  LoongArch: BPF: Add timed may_goto support

 arch/loongarch/include/asm/inst.h       |  1 +
 arch/loongarch/net/Makefile             |  2 +-
 arch/loongarch/net/bpf_jit.c            | 27 +++++++++++++-
 arch/loongarch/net/bpf_timed_may_goto.S | 47 +++++++++++++++++++++++++
 4 files changed, 75 insertions(+), 2 deletions(-)
 create mode 100644 arch/loongarch/net/bpf_timed_may_goto.S

-- 
2.53.0
Re: [PATCH v3 0/2] LoongArch: BPF: per-CPU address resolution and timed may_goto
Posted by Tiezhu Yang 1 month, 3 weeks ago
On 2026/8/4 下午11:39, George Guo wrote:
> These are the first two patches of the LoongArch BPF JIT feature work,
> sent on their own to make review easier. The remaining pieces
> (per-program private stacks, exceptions/bpf_throw, sign-extending loads
> and atomics on arena pointers, and the matching selftests) will follow in
> a later series.

The patches apply cleanly on top of Chenguang's series.

The following selftests passed on LoongArch:

   # ./test_progs -t cpumask
   # ./test_progs -t percpu_alloc
   # ./test_progs -t iters

Acked-by: Tiezhu Yang <yangtiezhu@loongson.cn>
Tested-by: Tiezhu Yang <yangtiezhu@loongson.cn>

Thanks,
Tiezhu

Re: [PATCH v3 0/2] LoongArch: BPF: per-CPU address resolution and timed may_goto
Posted by Huacai Chen 1 month, 3 weeks ago
On Thu, Aug 6, 2026 at 8:23 PM Tiezhu Yang <yangtiezhu@loongson.cn> wrote:
>
> On 2026/8/4 下午11:39, George Guo wrote:
> > These are the first two patches of the LoongArch BPF JIT feature work,
> > sent on their own to make review easier. The remaining pieces
> > (per-program private stacks, exceptions/bpf_throw, sign-extending loads
> > and atomics on arena pointers, and the matching selftests) will follow in
> > a later series.
>
> The patches apply cleanly on top of Chenguang's series.
>
> The following selftests passed on LoongArch:
>
>    # ./test_progs -t cpumask
>    # ./test_progs -t percpu_alloc
>    # ./test_progs -t iters
>
> Acked-by: Tiezhu Yang <yangtiezhu@loongson.cn>
> Tested-by: Tiezhu Yang <yangtiezhu@loongson.cn>
Applied, thanks.

Huacai

>
> Thanks,
> Tiezhu
>
Re: [PATCH v3 0/2] LoongArch: BPF: per-CPU address resolution and timed may_goto
Posted by Huacai Chen 1 month, 3 weeks ago
Hi, George,

In V1 it is independent,

In V2 it is part of
https://lore.kernel.org/loongarch/20260702022322.51033-1-dongtai.guo@linux.dev/T/#t

In V3 it is independent again.

What do you want? You can ask Tiezhu and Hengqi to review your
patches, but don't do such confusing things.


Huacai

On Tue, Aug 4, 2026 at 11:46 PM George Guo <dongtai.guo@linux.dev> wrote:
>
> These are the first two patches of the LoongArch BPF JIT feature work,
> sent on their own to make review easier. The remaining pieces
> (per-program private stacks, exceptions/bpf_throw, sign-extending loads
> and atomics on arena pointers, and the matching selftests) will follow in
> a later series.
>
> Both patches are independent of each other. They apply on top of Chenguang
> Zhao's "LoongArch bpf kptr xchg inline support" v4 series:
>
>   https://lore.kernel.org/all/20260729022837.355549-1-chenguang.zhao@linux.dev/
>
> Patch 1 adds the internal-only BPF MOV that resolves a per-CPU offset to
> the current CPU's address, advertised through
> bpf_jit_supports_percpu_insn(). LoongArch keeps the current CPU's per-CPU
> base in $r21 (__my_cpu_offset), so the resolution is a single add.
> Exercised by the cpumask and percpu_alloc selftests.
>
> Patch 2 implements arch_bpf_timed_may_goto() and advertises it, so the
> verifier lowers may_goto to its timed variant, which bounds the loop by a
> wall-clock timeout kept in a per-loop stack slot. It uses a custom calling
> convention: the count/timestamp stack offset is passed in BPF_REG_AX and
> returned there, so the JIT call path skips the usual "BPF_REG_0 = return
> value" move for this helper. Exercised by the iters selftests.
>
> Selftest results (all PASS):
>
>   Feature                          Test(s)
>   --------------------------------------------------------------
>   1  internal-only MOV (percpu)     cpumask, percpu_alloc
>   2  timed may_goto                 iters
>
> $ sudo ./test_progs -t cpumask
> ...
> Summary: 1/35 PASSED, 0 SKIPPED, 0 FAILED
>
> $ sudo ./test_progs -t percpu_alloc
> ...
> Summary: 1/18 PASSED, 0 SKIPPED, 0 FAILED
>
> $ sudo ./test_progs -t iters
> ...
> Summary: 1/93 PASSED, 0 SKIPPED, 0 FAILED
>
> Based on loongarch-next:
>   https://git.kernel.org/pub/scm/linux/kernel/git/chenhuacai/linux-loongson.git/log/?h=loongarch-next
>
> These two patches are split from the full v2 series (11 patches):
>   https://lore.kernel.org/all/20260702022322.51033-1-dongtai.guo@linux.dev/
>
> George Guo (2):
>   LoongArch: BPF: Support internal-only MOV to resolve per-CPU addrs
>   LoongArch: BPF: Add timed may_goto support
>
>  arch/loongarch/include/asm/inst.h       |  1 +
>  arch/loongarch/net/Makefile             |  2 +-
>  arch/loongarch/net/bpf_jit.c            | 27 +++++++++++++-
>  arch/loongarch/net/bpf_timed_may_goto.S | 47 +++++++++++++++++++++++++
>  4 files changed, 75 insertions(+), 2 deletions(-)
>  create mode 100644 arch/loongarch/net/bpf_timed_may_goto.S
>
> --
> 2.53.0
>
Re: [PATCH v3 0/2] LoongArch: BPF: per-CPU address resolution and timed may_goto
Posted by George Guo 1 month, 3 weeks ago
Hi Huacai,

On Wed, Aug 5, 2026 at 3:11 PM, Huacai Chen <chenhuacai@kernel.org> wrote:
> In V1 it is independent,
> In V2 it is part of [...] the 11-patch series
> In V3 it is independent again.
>
> What do you want?

You are right that the v1 -> v2 -> v3 shape looks inconsistent. The plan
has always been incremental: Tiezhu suggested sending this work as small,
stacked series -- A first, then B rebased on A, then C on B -- so each
round stays small enough to review carefully. My v2 bundled all 11 patches
to show the full picture, but that was too large to review effectively, so
v3 goes back to the incremental approach.

Concretely, this 2-patch series (per-CPU MOV + timed may_goto) is A. The
remaining pieces -- private stack, exceptions/bpf_throw, arena LDSX and
atomics, and the matching selftests -- will follow as separate series
stacked on top of it. The structure stays incremental from here.

> You can ask Tiezhu and Hengqi to review your patches, but don't do such
> confusing things.

Understood. I will keep the split stable and ask Tiezhu and Hengqi to
review this batch.

Thanks,
George