Skip to content

feat(active-learning): add FLORA active-learning operator - #1364

Open
EmileAydar wants to merge 3 commits into
IBM:mainfrom
EmileAydar:feat/flora-active-learning
Open

EmileAydar wants to merge 3 commits into
IBM:mainfrom
EmileAydar:feat/flora-active-learning

Conversation

@EmileAydar

@EmileAydar EmileAydar commented Aug 19, 2026 •

Copy link
Copy Markdown

This PR builds on PR #1361 (please merge that first). This PR adds FLORA,
the second operator in the ado-active-learning plugin: a pointwise
random-forest disagreement acquisition operator for finite Discovery Spaces.

Where PKH (this plugin's first operator) chooses entities to make the
labelled sample representative of the full pool, FLORA chooses entities to
make the fitted model as predictively accurate as possible. They're
deliberately complementary: two objectives for the same "what should I
measure next" problem, sharing the same plumbing.

Why add this to ADO

See PR #1361 for the general motivation.
FLORA specifically covers the case PKH doesn't: when the goal isn't a
representative sample but the lowest-error predictive model over the entity
pool. Shipping both together, with the same parameter shape and installation
path, lets ADO users pick the acquisition objective that matches their actual
goal rather than only having one option.

How it works

FLORA reuses the same _PredictiveSampleSelector/_SequentialSelectionState
machinery PKH uses (from active_learning/regression/_shared.py), swapping
in a different acquisition score.

After one deterministic pilot (by default: 5 randomly sampled entities), if no labels are already available, FLORA repeats:

  1. Fit or reuse a RandomForestRegressor on the labels collected so far, refitting according to the schedule below.
  2. For every pool entity, compute the variance across the individual trees'
    predictions - a proxy for how uncertain the forest is there - and average
    that disagreement within each leaf.
  3. Select the unmeasured entity with the largest mean disagreement-based leaf-gain score across trees: each leaf contribution grows with its disagreement and pool count and decreases as its labelled count grows.
  4. Wait for ADO to record that entity's measurement, then repeat.

The forest is refit on a log-tempered expanding schedule, so refits become less
frequent as the labelled sample grows (vs. PKH's fixed epochLength).

Example

uv sync --group test
ado get operators   # flora now appears alongside pkh

ado create space -f plugins/operators/active_learning/examples/discoveryspace.yaml --new-sample-store
ado create operation -f plugins/operators/active_learning/examples/operation_flora.yaml --use-latest space
ado show measurements operation --use-latest --property-format target

Non-Breaking Changes

Purely additive: a new ado.operators entry point (flora) inside the
already-registered ado-active-learning plugin. Does not modify PKH's
behavior or any other operator.

Emile Aydar added 2 commits August 19, 2026 11:58
Add the ado-active-learning plugin implementing Predictive Kernel Herding
(PKH), a regression-based active-learning operator for finite Discovery
Spaces, and register it as a uv workspace member.

Signed-off-by: Emile Aydar <emile.aydar@ibm.com>
Add FLORA, a pointwise random-forest disagreement acquisition operator,
alongside PKH in the ado-active-learning plugin. FLORA selects the entity
expected to most reduce a fitted model's predictive risk, complementing
PKH's representativeness-focused acquisition.

Signed-off-by: Emile Aydar <emile.aydar@ibm.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(active-learning): add FLORA active-learning operator

1 participant