Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions .github/workflows/pytest.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,28 @@ name: Run Tests
on: [push]

jobs:
unit:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
- name: set up Python
uses: actions/setup-python@v5
with:
python-version: "3.10"
- name: set up uv
uses: astral-sh/setup-uv@v3
- name: install dependencies
run: |
uv sync --extra dev --frozen
- name: test_unit
# no Learning Loop and no secrets needed, so this suite gates the slow ones
run: |
uv run python -m pytest learning_loop_node/tests/unit -v

pytest_3_10:
needs:
- unit
runs-on: ubuntu-latest
timeout-minutes: 30
strategy:
Expand Down Expand Up @@ -91,6 +112,7 @@ jobs:

slack:
needs:
- unit
- pytest_3_10
- pytest_3_13
if: always() # also execute when pytest fails
Expand Down
156 changes: 156 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,156 @@
# AI Agent Guidelines for the Learning Loop Node Library

`learning_loop_node` is the **public Python library** every node uses to talk to the
[Learning Loop](https://learning-loop.ai). Four node types build on it: Trainer, Detector,
Annotator and Converter. It is published to PyPI, so its public surface is an API that the
Learning Loop backend and every node repository depends on — `../yolov5_node` is the public one.

For coding standards see [CONTRIBUTING.md](CONTRIBUTING.md). [README.md](README.md) documents the
environment variables, the node types and how to write a node against them.

## Layout

- `learning_loop_node/` — the library itself: `node.py` and `rest.py` for the shared node
machinery, `trainer/`, `detector/`, `annotation/` for the per-type base logic, `data_classes/`
and `enums/` for the wire types, `loop_communication.py` and `data_exchanger.py` for the
loop-facing HTTP and socket.io traffic.
- `learning_loop_node/testing/` — test helpers that **ship in the wheel**, so a node repository
can import them: `TestingTrainerLogic`, `TestingDetectorLogic`, the dummy detections, the
`condition` poller and the `fixtures` pytest plugin. Anything reusable belongs here, not in
`tests/`.
- `learning_loop_node/tests/` — the library's own suites (`unit`, `annotator`, `detector`,
`trainer`, `general`), excluded from the wheel. Only `unit` runs without a Learning Loop; it
covers the pure helpers such as `detector/postprocess.py` and `detector/geometry.py`.
- `mock_trainer/`, `mock_detector/`, `mock_annotator/` — reference implementations with their own
tests. They are what `loop`'s CI runs against, so they are also the best template for a new node.
- `demo_segmentation_tool/` — a worked annotator example.

## Architecture

Every node is a `FastAPI` subclass (`node.py`): its lifespan connects to the loop and starts a
`repeat_loop` that calls the subclass' `on_repeat` every `repeat_loop_cycle_sec` (5 s) — that loop,
not an event handler, is what drives status reporting, model updates and training continuation.
Subclasses implement `on_startup`, `on_shutdown`, `on_repeat` and `register_sio_events`.

Two channels lead to the loop and both are needed: `LoopCommunicator` (httpx, login cookies, retry
on 401/429) for the REST API, and a socket.io *client* for status updates and loop-issued commands.
`DataExchanger` sits on top of the communicator and moves images and model zips.

- **Trainer** — `TrainerLogicGeneric._training_loop` is a state machine over `TrainerState`
(`enums/trainer.py`): download data → download base model → train → sync confusion matrix →
upload model → detect → upload detections → cleanup. `_perform_state` wraps each step: an
ordinary exception records the error and rewinds to the previous state (retried on the next
cycle), a `CriticalError` jumps to `ReadyForCleanup`. Every transition is persisted through
`LastTrainingIO`, so `try_continue_run_if_incomplete` resumes an interrupted training after a
restart. The loop starts a training via the `begin_training` sio event. A concrete trainer
implements `_train`, `_do_detections`, `_get_new_best_training_state`, `_on_metrics_published`,
`_get_latest_model_files` and `_clear_training_data`; `TrainerLogic` adds an `Executor` for
trainers that shell out to a training process.
- **Detector** — the exception to the pattern: constructed with `needs_login=False, needs_sio=False`,
so it has no sio client to the loop. It *hosts* a socket.io server for its own clients and polls
`/{org}/projects/{project}/deployment/target` over REST in `on_repeat` instead. `_DetectorState`
(`_Initializing` / `_Updating` / `_ActiveDetector`) models the model swap: download to
`models/<version>`, build a `DetectorLogic` through the factory, then swap atomically so the old
model keeps serving until the new one is ready (unless `EXCLUSIVE_MODEL_BUILD` frees VRAM first).
`OperationMode` gates whether updates may happen at all. Detections flow through
`RelevanceFilter`, which writes selected images to the `Outbox` on disk; a separate upload process
drains it.
- **Annotator** — thin: it forwards the loop frontend's `handle_user_input` events into
`AnnotatorLogic` and keeps a per-frontend history.

`trainer/subprocess.py`, `trainer/batch_size.py` and `trainer/metrics.py` hold the parts of a
trainer that are not framework-specific. `iterator_cpu_bound` runs a training generator in a
spawned process and yields its progress through a `maxsize=1` queue, so the event loop stays
responsive, CUDA state stays out of the node process, and the training can never run more than
one item ahead of the bookkeeping. `find_batch_size` probes for the largest power-of-two batch
that fits, around a `fits` predicate the trainer supplies — none of it imports torch, so a node
brings its own way of running a step. `macro_f1` scores the confusion matrix
`_get_new_best_training_state` returns.

`helpers/entrypoint.py` holds what every node's `main.py` repeats: `node_parser` builds a
configargparse parser with `--host`/`--port`, `run_node` starts uvicorn. A setting is a flag
*and* an environment variable from one declaration — `--conf-threshold` reads
`CONF_THRESHOLD`. The exception is `--host`/`--port`, which read `NODE_HOST`/`NODE_PORT`: the
bare `HOST` already means the loop's address, and a node adopting it would hand it to uvicorn
and fail to bind. A node that used to require a prefix passes `legacy_env_prefix`, and the
prefixed names keep working with a warning.

`detector/postprocess.py` and `detector/geometry.py` hold the parts of a detector that do *not*
depend on the model: confidence filtering, per-class NMS, box/point clipping, and turning
predictions into the loop's dataclasses. A node should import them rather than write its own —
every node repository had grown its own drifting copy, which is why they live here. Note the two
containers: `to_image_metadata` builds what a **detector** node reports, `to_detections` what a
**trainer**'s auto-detection pass reports. Both go through one routine, so the two paths cannot
drift apart again.

All node state lives under `GLOBALS.data_folder` (`DATA_FOLDER`, default `/data`): `uuids.json`
(the node uuid is derived from its name and reused across restarts), `models/` plus the
`current_model` symlink, `outbox/`, and the per-project training folders.

## Running and testing

The `unit` suite is self-contained — no Learning Loop, no credentials, no network — so it is the
one to run while iterating, and the one CI gates the others on:

```bash
python -m pytest learning_loop_node/tests/unit -v
```

Every other suite talks to a real Learning Loop instance and reads its credentials from a local
`.env` (`LOOP_HOST`, `LOOP_USERNAME`, `LOOP_PASSWORD`). Without a reachable loop those cannot pass —
do not treat their failure as a regression you introduced.

```bash
./run_tests.sh # all suites
./run_tests.sh <filter> # passed to pytest as -k
```

Each suite carries its own `pytest.ini` (that is where `asyncio_mode = auto` comes from), so always
run pytest with a path inside one suite — a bare `pytest` from the repository root picks up no
config and the async tests error out:

```bash
python -m pytest learning_loop_node/tests/trainer -v # one suite
python -m pytest learning_loop_node/tests/trainer/test_errors.py -v # one file
python -m pytest learning_loop_node/tests/trainer -v -k <test_name> # one test
```

An autouse fixture repoints `GLOBALS.data_folder` at `/tmp/learning_loop_lib_data` and wipes it
around every test, so tests never touch `/data`. It lives in `learning_loop_node/testing/fixtures.py`
and every suite pulls it in with

```python
from ...testing.fixtures import clear_loggers, data_folder # noqa: F401
```

The `general` suite generates and deletes a real `zauberzeug/pytest_nodelib_general` project on the
loop; the detector suite starts the node in a forked uvicorn process on `GLOBALS.detector_port`.

Every fixture that creates or deletes a project calls `assert_not_production_loop()` first, and
`run_tests.sh` refuses to start without `LOOP_HOST`. `LoopCommunicator` defaults to
`learning-loop.ai` — production — so a forgotten `.env` would otherwise aim those fixtures at real
customer data. CI is unaffected: the workflows pin `preview.learning-loop.ai`.

There is no `.pre-commit-config.yaml` here and no ruff in the project environment. Lint with:

```bash
uvx ruff check .
```

A clean tree already reports several hundred ruff findings, so a clean run is not a reachable goal.
Compare the count on the files you touched, before and after.

`.github/workflows/pytest.yml` runs the suites (the `unit` job first, without secrets),
`publish.yml` releases to PyPI on a tagged release.

## Working in this repository

- **This is a library — declaring a dependency is part of its API.** Before removing or loosening
one, grep every consuming repository checked out beside this one for the package: a consumer that
imports it without declaring it inherits it from here and breaks when it goes away.
- **Renaming or reshaping anything exported** breaks those repositories. Say so in the pull request
and check whether a companion change is needed there.
- `../loop` checks this repository out as its `nodes` symlink, so a local change is visible to a
local loop immediately — but only a released version reaches CI and production.
- Bump `version` in `pyproject.toml` for a release; the trainer nodes pin the library version in
their image tags (`A.B.C-nlvX.Y.Z`).
1 change: 1 addition & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
@AGENTS.md
66 changes: 49 additions & 17 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -1,21 +1,35 @@
# Contributing to Learning Loop Node

## Data Team Standards

## Linting

We use ruff and [pre-commit](https://github.com/pre-commit/pre-commit) to make sure the coding style is enforced.
You first need to install pre-commit and the corresponding git commit hooks by running the following commands:
We use [ruff](https://docs.astral.sh/ruff/) to enforce the coding style.
Its rules live in the `pyproject.toml` of the respective (sub-)project.

Repositories that ship a `.pre-commit-config.yaml` run ruff through
[pre-commit](https://github.com/pre-commit/pre-commit).
Install the git hooks once:

```bash
python3 -m pip install pre-commit
pre-commit install
```

After that you can make sure your code satisfies the coding style by running the following command:
After that you can check the whole repository with:

```bash
pre-commit run --all-files
```

Repositories without a `.pre-commit-config.yaml` run ruff directly:

```bash
uvx ruff check .
```

The repository-specific part of this file names the exact command wherever it differs.

## Style Guide

### 1. Single or double quotes
Expand All @@ -26,8 +40,8 @@ Double quotes may also be used for strings if the string contains a single quote
### 2. String formatting

We use f-strings if possible (due to their readability).
Logs shell use the lazy formatting provided by the logging (performance optimization).
When diverging from above rules, please provide a short comment with `# NOTE: ...` to explaining the reason.
Logs shall use the lazy formatting provided by logging (performance optimization).
When diverging from above rules, please provide a short comment with `# NOTE: ...` to explain the reason.

### 3. Line continuation

Expand Down Expand Up @@ -72,7 +86,7 @@ def parse_config(path: Path) -> Config | None:
...
```

## 6. Ordering: Important things first
### 6. Ordering: Important things first

We want to have main classes and functions at the top of the file, while helper functions and classes should be placed below. This allows to quickly understand the main purpose of the file without having to scroll through a lot of code.

Expand Down Expand Up @@ -101,23 +115,23 @@ class ReportGenerator:
...
```

## 7. Docstrings and type hints
### 7. Docstrings and type hints

We use the reStructuredText (reST) or Sphinx-style docstring format.
We don't declare types in the docstring but use type hints.
While type hints are mandatory, parameter descriptions and docstrings should only
be used if the function and parameter names are not self-explanatory.
Examples are listetd below:
Examples are listed below:

### One-line Docstring without parameters
#### One-line Docstring without parameters

```python
def greet() -> None:
"""Print a friendly greeting."""
print("Hello!")
print('Hello!')
```

### One-line Docstring with parameters
#### One-line Docstring with parameters

```python
def square(x: int | float) -> int | float:
Expand All @@ -129,7 +143,7 @@ def square(x: int | float) -> int | float:
return x * x
```

### Multi-line Docstring without parameters
#### Multi-line Docstring without parameters

```python
def get_timestamp() -> str:
Expand All @@ -140,10 +154,10 @@ def get_timestamp() -> str:
as "YYYY-MM-DD HH:MM:SS". It can be useful for logging or
displaying time information in user interfaces.
"""
return datetime.now().strftime("%Y-%m-%d %H:%M:%S")
return datetime.now().strftime('%Y-%m-%d %H:%M:%S')
```

### Multi-line Docstring with parameters
#### Multi-line Docstring with parameters

```python
def save_data(data: str | bytes, filename: str, overwrite: bool = False) -> None:
Expand All @@ -157,12 +171,30 @@ def save_data(data: str | bytes, filename: str, overwrite: bool = False) -> None
:param data: The content to be saved. Can be text or bytes.
:param filename: Path to the destination file.
:param overwrite: Whether to overwrite an existing file.
:return: True if the data was saved successfully, False otherwise.
:raises FileExistsError: If the file exists and overwrite is False.
"""
if os.path.exists(filename) and not overwrite:
raise FileExistsError(f"{filename} already exists.")
mode = "wb" if isinstance(data, bytes) else "w"
raise FileExistsError(f'{filename} already exists.')
mode = 'wb' if isinstance(data, bytes) else 'w'
with open(filename, mode) as f:
f.write(data)
```

### 8. Relative imports within a package

Inside a package we import other modules of the same package with relative imports, not with the package name.
This way the package can be renamed or moved without touching its imports.
We use the shortest path: siblings with a single dot, no going up and back down (`from ..ui.numpad import ...` in a module that lives in `ui/` itself).
Code outside the package — scripts, tests — has no choice and uses absolute imports.

```python
# preferred (in app/ui/scan_view.py)
from .. import config
from ..system import System
from .numpad import Numpad

# avoid
from app import config
from app.system import System
from app.ui.numpad import Numpad
```
46 changes: 42 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,11 +36,49 @@ You can configure connection to our Learning Loop by specifying the following en

Note that organization and project IDs are always lower case and may differ from the names in the Learning Loop which can have uppercase letters.

#### Testing
#### Testing your own node

We use github actions for CI. Tests can also be executed locally by running
`LOOP_HOST=XXXXXXXX LOOP_USERNAME=XXXXXXXX LOOP_PASSWORD=XXXXXXXX python -m pytest -v`
from learning_loop_node/learning_loop_node
`learning_loop_node.testing` ships with the library, so a node repository can test itself without
a Learning Loop:

```bash
pip install learning_loop_node[testing]
```

```python
# tests/conftest.py
pytest_plugins = ['learning_loop_node.testing.fixtures']
```

That gives every test an autouse `data_folder` fixture, which repoints `GLOBALS.data_folder` at
`/tmp/learning_loop_lib_data` and wipes it around each test, so a test never touches `/data`. The
path is shared, so do not run two such suites at once. The module also provides:

| | |
| --- | --- |
| `TestingTrainerLogic` | a `TrainerLogic` that trains a sleeping subprocess — drive the state machine without a framework |
| `TestingDetectorLogic`, `TestingDetectorFactory` | a detector that returns fixed detections |
| `get_dummy_detections()`, `get_dummy_metadata()` | one detection of every type |
| `condition(...)` | await a predicate with a timeout |
| `assert_training_state(...)`, `create_active_training_file(...)` | drive and assert the trainer state machine |
| `assert_not_production_loop()` | call this first in any fixture that creates or deletes a project |

`assert_not_production_loop()` matters because `LoopCommunicator` falls back to `learning-loop.ai`
when neither `LOOP_HOST` nor `HOST` is set — a forgotten `.env` would otherwise point a destructive
fixture at production.

#### Testing this library

We use github actions for CI. Locally, the `unit` suite needs nothing at all:

```bash
python -m pytest learning_loop_node/tests/unit -v
```

The other suites need a reachable Learning Loop and its credentials in a local `.env`
(`LOOP_HOST`, `LOOP_USERNAME`, `LOOP_PASSWORD`); `./run_tests.sh` runs them all and refuses to
start without `LOOP_HOST`. Each suite carries its own `pytest.ini`, so always pass a path inside
one suite — a bare `pytest` from the repository root picks up no config.

## Detector Node

Expand Down
Loading
Loading