Skip to content

Publish UL slots through a native gNB helper instead of a jbpf copy - #8

Draft
aferaudo wants to merge 19 commits into
mainfrom
tests/native-memcpy
Draft

aferaudo wants to merge 19 commits into
mainfrom
tests/native-memcpy

Conversation

@aferaudo

@aferaudo aferaudo commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

The UL slot codelet no longer copies the resource grid through jbpf. It calls a native helper that the gNB registers at runtime, and that helper converts the grid and writes it straight into /e3_ran_buffers. Only a 56 B descriptor goes through jbpf.

Before (main): uplink_slot_collect copied the whole 4-port grid (733,824 B per slot) into a jbpf output map with a chunked eBPF memcpy. The ubpf JIT emits scalar moves, so this was about 166 µs per slot. The controller then converted bf16 → fp16 and wrote the row into shared memory: two bulk passes per slot, one of them in eBPF.

After: the codelet calls jbpf_e3_publish_slot (helper id 33). The helper does a single AVX2+F16C bf16 → fp16 pass, with scale, directly into the shared-memory ring: about 65 µs per slot. The codelet then sends a 56 B e3_slot_desc (buffer and row index, flags, geometry, timestamps). The controller no longer touches IQ bytes. It forwards the descriptor as the L1-KPM indication, and the dApp reads the row.

Timings were measured on an OpenShift pod: the copy cost on 2026-06-29 and the publish cost on 2026-08-14.

How it works

  • Helpers, implemented in the gNB. e3_attach (id 32) and e3_publish_slot (id 33) live in OCUDU (ocudu_janus/e3/jbpf_e3_publish.cpp, branch e3). They are registered with jbpf_register_helper_function before any codelet loads, so jbpf itself needs no fork for them.
  • Shared contract in codelets/include/:
    • jbpf_e3_ids.h: helper and program-type ids.
    • jbpf_e3_slot_api.h: e3_shm_cfg, e3_slot_sel and e3_slot_desc.
    • e3_slot_convert.h: the convert itself.
    • OCUDU keeps byte-identical copies and fails its configure if they drift.
  • Region ownership. The controller creates /e3_ran_buffers in the NVIDIA SharedMemoryHeader layout (2 buffers × N rows) and pushes e3_shm_cfg (name, epoch, geometry, cbf16_scale) to the codelet's shm_in. The helper attaches O_RDWR without creating the region and checks the geometry against the header. If the source is outside the hook's window, it refuses rather than truncating.
  • Slot selection. e3_slot_sel (slot_mask, sfn_mod, sfn_offset) arrives on sel_in and is applied before any data moves. A non-selected slot costs a few comparisons on the PHY RX thread. The controller sends a "publish all" selector on the first batch.
  • shm.writer: gnb | controller enforces a single writer. The shipped codelet only supports gnb; controller remains for the legacy copying codelet.

Other changes

  • Configuration file. The controller now takes a single --config <yaml> argument instead of about 20 CLI flags. Unknown keys are an error, and the geometry, shm size and scale are validated. Two configs are provided: configs/e3_controller.yaml (JSON on 5555–5557) and configs/e3_controller_asn1.yaml (APER on 9990/9991/9999).
  • Offline codelet verifier (codelets/verifier/e3_verifier_cli.cpp). It knows both program types and the helper prototypes. The codelet Makefiles refuse to install a .o that fails verification. The gNB itself does not verify at load time.
  • Latency instrumentation.
    • latrec stage records on the slot path (src/e3sm/l1_kpm/l1_kpm_trace.h), enabled with ./build.sh --latrec.
    • bench/bench_stage_recording.cpp.
    • A CI workflow that builds with latrec on and off.
  • Monotonic clock. jbpf_patches/jbpf_monotonic_time.patch moves jbpf_time_get_ns() to CLOCK_MONOTONIC so that codelet, gNB, controller and dApp stamps share one clock. apply_jbpf_patches.sh applies it, and build.sh runs that script.
  • libe3 0.0.6 → 0.2.0 (E3Controller: bump libe3 submodule to latest tag 0.2.0 #7).
  • RF=1 eCPRI codelet. Adds compile-time slot, symbol and eAxC filters, and the matching updates to the Spectrum SM. Spectrum-ConfigControl is no longer compiled, which drops the BOOLEAN runtime dependency.

Deployment notes / breaking changes

  • gNB requirement. Needs a gNB built from OCUDU e3 with -DENABLE_JBPF=ON, so the helpers exist. Its jbpf must also have jbpf_monotonic_time.patch applied, or the codelet timestamps won't be comparable with gnb_ts_ns.
  • Start order. Start the controller first, then the gNB, then the dApp. Restart the gNB after any controller restart (see the first follow-up below).
  • Removed CLI flags. Launch scripts must switch to --config.
  • Descriptor size. e3_slot_desc is now 56 B, so the codelet and controller must be rebuilt together. The indication sent to the dApp is unchanged.
  • libe3 0.2.0 wire change. It changes the APER encoding (wider id ranges, sequenceId), so ASN.1 dApps must also use libe3 0.2.0. JSON is unaffected.

Known issues (follow-ups)

  • The gNB keeps its first attach. The epoch is fixed at 1, so the helper keeps the scale and mapping from its first attach for the gNB's whole lifetime. A changed cbf16_scale, or a restarted controller (which recreates the region), is not picked up until the gNB restarts.
  • Re-subscribe after a full unsubscribe. SlotIqPipeline::stop() doesn't reset shm_cfg_sent_, so the selector is likely not resent to the reloaded codelet.
  • libe3 log level is hard-coded to TRACE (src/e3_controller.cpp).
  • Failed SM start. If the codelet fails to load, the subscription is still acknowledged positively.
  • Header drift. The OCUDU copy of jbpf_e3_slot_api.h differs from this one in comments only, but the drift check fails the OCUDU configure until the copies are synced.

Testing

Tested on an OpenShift pod: Foxconn RU, n78, 100 MHz, 4T4R, OCUDU e3 gNB, adaptive_cpu dApp over JSON. In one run, 5941 slots were published and 1 was dropped (the first slot, before the attach completed), with no SPSC drops. The controller was built against libe3 0.2.0.

🤖 Generated with Claude Code

aferaudo and others added 19 commits July 8, 2026 12:39
The per-slot stage CSV opened an ofstream, wrote a row and flushed it,
once per slot, inside the sample handler. A synchronous formatted write
plus a flush on the slot path costs far more than the intervals it was
recording, so every number it produced included the cost of producing
it. Measured: ~1840 ns per slot against ~111 ns for the five records
that replace it, or roughly 17x.

It also documented a mechanism that does not exist: README and two
source comments described downstream send stages as joinable through a
libe3 `--pub-stages-log` option. There is no such option in libe3 on
any branch, so the send side was simply unmeasured.

Stamp the same boundaries into latrec instead - one clock read and four
stores into an mmap-backed ring, no syscall, allocation, formatting,
lock or I/O on the slot path. The mapping onto the shared stage catalog
lives in one header, l1_kpm_trace.h, which the handler and the bench
both use, so what CI measures is the code that runs on the radio:

  A1  RECORD_BEGIN -> PROCESS_BEGIN          the data recording
  A2  PROCESS_BEGIN -> ENCODE_E3SM_BEGIN     getting it to the SM
  A3  ENCODE_E3SM_BEGIN -> ENCODE_E3SM_DONE  E3SM encode
      ENCODE_E3SM_DONE -> WAIT_ENTER         the emit tail

A1 and A2 are recorded here, and the gNB needs no instrumentation for
it. Both of A1's boundaries are already on the wire when the slot
arrives: gnb_ts_ns is stamped on the last symbol with the grid complete
and nothing copied, and codelet_ts_ns just before the codelet submits.
So the controller back-dates those two stamps rather than the RAN
keeping a ring of its own. That is sound because the whole chain reads
CLOCK_MONOTONIC now - the hook, jbpf_time_get_ns() and the dispatcher
poll - which is latrec's own clock, so there is no domain conversion.

What still has to hold is the ring's own invariant: a single-writer log
whose t_ns ascends, whose one permitted descent libe3's reader treats as
the wrap point and silently rotates at. Program order gives a wide
margin (the jbpf hook is a synchronous inline call, and A1 is tens of
microseconds against a >=500 us slot spacing), but a pipeline stall or
several sectors interleaving on one ring would break it. So the floor is
enforced per thread and the count of enforced stamps is reported, in the
drop CSV and the shutdown summary - a non-zero count means A1 is
understated for that many slots.

Four sub-hops stay recoverable from aux payloads, which is what keeps
jbpf dispatch separable from the data movement, and preserves the
distinction the codelet's entry/exit timestamp split was added for.

Correct the two timestamp comments. codelet_ts_ns was documented as a
"codelet entry timestamp" whose interval with gnb_ts_ns was "jbpf
invocation + verifier path overhead". It is stamped last, not at entry;
the interval is A1, the data recording itself, dominated by the convert
and the row write; and there is no runtime verifier cost at all, since
the gNB loads through ubpf whose JIT emits no bounds checks and
verification is an offline build-time gate. gnb_ts_ns said
CLOCK_REALTIME in three places and no longer is.

Replace the CSV's option with logging.latrec_dir, which is placement
only: whether anything is recorded is decided when libe3 is built, by
-DLIBE3_ENABLE_LATREC (./build.sh --latrec). Keep the throttled
drop-accounting CSV - it is aggregate and <=1/s, so it is not the thing
that perturbed. The legacy eCPRI SM keeps its own stage CSV under its
own key; it is a different data path and has the same problem, to be
fixed when that path is next touched.

Move the libe3 pin to 22f3918 (0.1.2), a fast-forward. Its message-id
work surfaces the assigned id to a dApp calling send_control/send_report
but not on the Service Model emit path, so origin_seq remains the only
cross-component join.

Drop the BOOLEAN skeleton staging from build.sh. It existed to have
libe3 compile asn_DEF_BOOLEAN for Spectrum-ConfigControl; that type is
no longer compiled and nothing references the symbol, so the staging had
nothing left to supply while still aborting the build on any host
without asn1c's reference skeletons.

Add CI, which this repository had none of. It builds both ways, asserts
that an untraced binary links no latrec symbols and carries no ring
registry, runs the recording-cost bench, and converts a capture with
libe3's own tool to check the ring role maps to the intended component,
the source leg is complete, the emit-tail hop column materialises, and
nothing wrapped.
Co-authored-by: ocudu <ocudu@localhost>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants