Decision
bee_size_multiplier (N) stays a single global value, shared across every lab and video. We are not tuning a separate N per lab.
This answers the open question from the Aug 24 update (commit 587b07e).
Rationale
WaggleNet is meant to be a one-model-fits-all solution. Every parameter we make lab-specific is a parameter the end user has to determine for their own setup, and the per-lab comb-cell calibration we already ask for is at — arguably past — the limit of what a working biologist can and wants to do before running a detector. Adding a second, per-lab quantity that can only be found by a hyperparameter search over labelled data would put the method out of reach for exactly the users we are building it for.
So: the per-lab quantity stays the one thing that is measurable without labels (BEE_LENGTH_FRACTION, from the comb-cell annotator), and the tuned quantity stays global.
Follow-ups needed before we lean on the current N
Decision
bee_size_multiplier(N) stays a single global value, shared across every lab and video. We are not tuning a separate N per lab.This answers the open question from the Aug 24 update (commit 587b07e).
Rationale
WaggleNet is meant to be a one-model-fits-all solution. Every parameter we make lab-specific is a parameter the end user has to determine for their own setup, and the per-lab comb-cell calibration we already ask for is at — arguably past — the limit of what a working biologist can and wants to do before running a detector. Adding a second, per-lab quantity that can only be found by a hyperparameter search over labelled data would put the method out of reach for exactly the users we are building it for.
So: the per-lab quantity stays the one thing that is measurable without labels (
BEE_LENGTH_FRACTION, from the comb-cell annotator), and the tuned quantity stays global.Follow-ups needed before we lean on the current N
bee_n_mindefaults to 0.5 (viewer/server.py, search range 0.5–20.0). An optimum landing on a boundary usually means the search was clipped — re-run with a floor around 0.05 and confirm 0.5 is a real optimum and not a wall.compute_dance_level_metricsmatches withpos_thresholdsas a plain fraction of frame width (0.02–0.10). At 0.02 that is ~0.9 bee-lengths for berlin but ~0.26 for nieh, so nieh dances are substantially easier to match. We now calibrate the clustering in bee units while the metric stays in frame-width units — the search is being scored against a yardstick with a built-in per-lab bias.eval.max_dets(5 → 10, which feeds the dance-level pipeline viackpt_eval.py). The clean ablation is new params, bee-size off vs new params, bee-size on.BEE_LENGTH_FRACTIONmeasurement base.output/projected_summary.csvshowsn_source_videos = 1per lab, with only the top resolution tier measured and the rest derived by scaling. If a global N rides on these constants, they should rest on more than one video each.