What
Get more labelled dances from the Sharoni lab (Avi Gabai), but not by having them annotate from scratch. Instead: run the current model over the unlabelled footage, load the predictions into the viewer dashboard, and have the annotator correct them.
Why
Annotating waggle runs from a blank slate is slow and is the reason our labelled set has grown so little. If the model is already reasonably good — and from what we have seen it is — then reviewing and fixing predictions is a much cheaper interaction than producing them, and it scales with model quality instead of fighting it.
The dashboard already has the pieces: GT annotation editing, per-video inspection, and re-clustering with adjustable post-processing.
Depends on
Model quality on Sharoni footage specifically. If precision is poor there, correcting predictions is more work than annotating fresh, and this approach backfires. Establish that first — a per-lab recall/precision breakdown at the current checkpoint is the go/no-go.
Note that the Sharoni 30-frame annotation offset (earlier milestone #14) has never been conclusively resolved. If existing Sharoni GT is misaligned, that has to be settled before we ask for more labels against the same convention.
Done when
What
Get more labelled dances from the Sharoni lab (Avi Gabai), but not by having them annotate from scratch. Instead: run the current model over the unlabelled footage, load the predictions into the viewer dashboard, and have the annotator correct them.
Why
Annotating waggle runs from a blank slate is slow and is the reason our labelled set has grown so little. If the model is already reasonably good — and from what we have seen it is — then reviewing and fixing predictions is a much cheaper interaction than producing them, and it scales with model quality instead of fighting it.
The dashboard already has the pieces: GT annotation editing, per-video inspection, and re-clustering with adjustable post-processing.
Depends on
Model quality on Sharoni footage specifically. If precision is poor there, correcting predictions is more work than annotating fresh, and this approach backfires. Establish that first — a per-lab recall/precision breakdown at the current checkpoint is the go/no-go.
Note that the Sharoni 30-frame annotation offset (earlier milestone #14) has never been conclusively resolved. If existing Sharoni GT is misaligned, that has to be settled before we ask for more labels against the same convention.
Done when