Conversation
The per-slot stage CSV opened an ofstream, wrote a row and flushed it,
once per slot, inside the sample handler. A synchronous formatted write
plus a flush on the slot path costs far more than the intervals it was
recording, so every number it produced included the cost of producing
it. Measured: ~1840 ns per slot against ~111 ns for the five records
that replace it, or roughly 17x.
It also documented a mechanism that does not exist: README and two
source comments described downstream send stages as joinable through a
libe3 `--pub-stages-log` option. There is no such option in libe3 on
any branch, so the send side was simply unmeasured.
Stamp the same boundaries into latrec instead - one clock read and four
stores into an mmap-backed ring, no syscall, allocation, formatting,
lock or I/O on the slot path. The mapping onto the shared stage catalog
lives in one header, l1_kpm_trace.h, which the handler and the bench
both use, so what CI measures is the code that runs on the radio:
A1 RECORD_BEGIN -> PROCESS_BEGIN the data recording
A2 PROCESS_BEGIN -> ENCODE_E3SM_BEGIN getting it to the SM
A3 ENCODE_E3SM_BEGIN -> ENCODE_E3SM_DONE E3SM encode
ENCODE_E3SM_DONE -> WAIT_ENTER the emit tail
A1 and A2 are recorded here, and the gNB needs no instrumentation for
it. Both of A1's boundaries are already on the wire when the slot
arrives: gnb_ts_ns is stamped on the last symbol with the grid complete
and nothing copied, and codelet_ts_ns just before the codelet submits.
So the controller back-dates those two stamps rather than the RAN
keeping a ring of its own. That is sound because the whole chain reads
CLOCK_MONOTONIC now - the hook, jbpf_time_get_ns() and the dispatcher
poll - which is latrec's own clock, so there is no domain conversion.
What still has to hold is the ring's own invariant: a single-writer log
whose t_ns ascends, whose one permitted descent libe3's reader treats as
the wrap point and silently rotates at. Program order gives a wide
margin (the jbpf hook is a synchronous inline call, and A1 is tens of
microseconds against a >=500 us slot spacing), but a pipeline stall or
several sectors interleaving on one ring would break it. So the floor is
enforced per thread and the count of enforced stamps is reported, in the
drop CSV and the shutdown summary - a non-zero count means A1 is
understated for that many slots.
Four sub-hops stay recoverable from aux payloads, which is what keeps
jbpf dispatch separable from the data movement, and preserves the
distinction the codelet's entry/exit timestamp split was added for.
Correct the two timestamp comments. codelet_ts_ns was documented as a
"codelet entry timestamp" whose interval with gnb_ts_ns was "jbpf
invocation + verifier path overhead". It is stamped last, not at entry;
the interval is A1, the data recording itself, dominated by the convert
and the row write; and there is no runtime verifier cost at all, since
the gNB loads through ubpf whose JIT emits no bounds checks and
verification is an offline build-time gate. gnb_ts_ns said
CLOCK_REALTIME in three places and no longer is.
Replace the CSV's option with logging.latrec_dir, which is placement
only: whether anything is recorded is decided when libe3 is built, by
-DLIBE3_ENABLE_LATREC (./build.sh --latrec). Keep the throttled
drop-accounting CSV - it is aggregate and <=1/s, so it is not the thing
that perturbed. The legacy eCPRI SM keeps its own stage CSV under its
own key; it is a different data path and has the same problem, to be
fixed when that path is next touched.
Move the libe3 pin to 22f3918 (0.1.2), a fast-forward. Its message-id
work surfaces the assigned id to a dApp calling send_control/send_report
but not on the Service Model emit path, so origin_seq remains the only
cross-component join.
Drop the BOOLEAN skeleton staging from build.sh. It existed to have
libe3 compile asn_DEF_BOOLEAN for Spectrum-ConfigControl; that type is
no longer compiled and nothing references the symbol, so the staging had
nothing left to supply while still aborting the build on any host
without asn1c's reference skeletons.
Add CI, which this repository had none of. It builds both ways, asserts
that an untraced binary links no latrec symbols and carries no ring
registry, runs the recording-cost bench, and converts a capture with
libe3's own tool to check the ring role maps to the intended component,
the source leg is complete, the emit-tail hop column materialises, and
nothing wrapped.
Co-authored-by: ocudu <ocudu@localhost>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The UL slot codelet no longer copies the resource grid through jbpf. It calls a native helper that the gNB registers at runtime, and that helper converts the grid and writes it straight into
/e3_ran_buffers. Only a 56 B descriptor goes through jbpf.Before (
main):uplink_slot_collectcopied the whole 4-port grid (733,824 B per slot) into a jbpf output map with a chunked eBPF memcpy. The ubpf JIT emits scalar moves, so this was about 166 µs per slot. The controller then converted bf16 → fp16 and wrote the row into shared memory: two bulk passes per slot, one of them in eBPF.After: the codelet calls
jbpf_e3_publish_slot(helper id 33). The helper does a single AVX2+F16C bf16 → fp16 pass, with scale, directly into the shared-memory ring: about 65 µs per slot. The codelet then sends a 56 Be3_slot_desc(buffer and row index, flags, geometry, timestamps). The controller no longer touches IQ bytes. It forwards the descriptor as the L1-KPM indication, and the dApp reads the row.Timings were measured on an OpenShift pod: the copy cost on 2026-06-29 and the publish cost on 2026-08-14.
How it works
e3_attach(id 32) ande3_publish_slot(id 33) live in OCUDU (ocudu_janus/e3/jbpf_e3_publish.cpp, branche3). They are registered withjbpf_register_helper_functionbefore any codelet loads, so jbpf itself needs no fork for them.codelets/include/:jbpf_e3_ids.h: helper and program-type ids.jbpf_e3_slot_api.h:e3_shm_cfg,e3_slot_selande3_slot_desc.e3_slot_convert.h: the convert itself./e3_ran_buffersin the NVIDIASharedMemoryHeaderlayout (2 buffers × N rows) and pushese3_shm_cfg(name, epoch, geometry,cbf16_scale) to the codelet'sshm_in. The helper attachesO_RDWRwithout creating the region and checks the geometry against the header. If the source is outside the hook's window, it refuses rather than truncating.e3_slot_sel(slot_mask,sfn_mod,sfn_offset) arrives onsel_inand is applied before any data moves. A non-selected slot costs a few comparisons on the PHY RX thread. The controller sends a "publish all" selector on the first batch.shm.writer: gnb | controllerenforces a single writer. The shipped codelet only supportsgnb;controllerremains for the legacy copying codelet.Other changes
--config <yaml>argument instead of about 20 CLI flags. Unknown keys are an error, and the geometry, shm size and scale are validated. Two configs are provided:configs/e3_controller.yaml(JSON on 5555–5557) andconfigs/e3_controller_asn1.yaml(APER on 9990/9991/9999).codelets/verifier/e3_verifier_cli.cpp). It knows both program types and the helper prototypes. The codelet Makefiles refuse to install a.othat fails verification. The gNB itself does not verify at load time.src/e3sm/l1_kpm/l1_kpm_trace.h), enabled with./build.sh --latrec.bench/bench_stage_recording.cpp.jbpf_patches/jbpf_monotonic_time.patchmovesjbpf_time_get_ns()to CLOCK_MONOTONIC so that codelet, gNB, controller and dApp stamps share one clock.apply_jbpf_patches.shapplies it, andbuild.shruns that script.Spectrum-ConfigControlis no longer compiled, which drops the BOOLEAN runtime dependency.Deployment notes / breaking changes
e3with-DENABLE_JBPF=ON, so the helpers exist. Its jbpf must also havejbpf_monotonic_time.patchapplied, or the codelet timestamps won't be comparable withgnb_ts_ns.--config.e3_slot_descis now 56 B, so the codelet and controller must be rebuilt together. The indication sent to the dApp is unchanged.sequenceId), so ASN.1 dApps must also use libe3 0.2.0. JSON is unaffected.Known issues (follow-ups)
cbf16_scale, or a restarted controller (which recreates the region), is not picked up until the gNB restarts.SlotIqPipeline::stop()doesn't resetshm_cfg_sent_, so the selector is likely not resent to the reloaded codelet.src/e3_controller.cpp).jbpf_e3_slot_api.hdiffers from this one in comments only, but the drift check fails the OCUDU configure until the copies are synced.Testing
Tested on an OpenShift pod: Foxconn RU, n78, 100 MHz, 4T4R, OCUDU
e3gNB, adaptive_cpu dApp over JSON. In one run, 5941 slots were published and 1 was dropped (the first slot, before the attach completed), with no SPSC drops. The controller was built against libe3 0.2.0.🤖 Generated with Claude Code