Skip to content

Share the framework-agnostic parts of a trainer - #91

Closed
klangenk wants to merge 1 commit into
share-node-entrypointfrom
share-trainer-helpers
Closed

klangenk wants to merge 1 commit into
share-node-entrypointfrom
share-trainer-helpers

Conversation

@klangenk

Copy link
Copy Markdown
Contributor

Motivation

Three things every trainer needs, none of which depend on the training framework, all of which existed only inside dfine_node — or in worse form elsewhere:

  • Running a training out of process. A trainer that trains in-process must not stall the node, and CUDA state must stay out of the node process. dfine_node solved this with a generator running in a spawned process; it is pure stdlib and nothing about it is D-FINE.
  • Choosing a batch size. Three repositories had three different implementations: a real probe (dfine_node), an estimate parsed out of torchinfo's summary text (yolov5_node), and a doubling loop catching bare RuntimeError (classification_node).
  • Scoring the confusion matrix. Implemented three times over two different data shapes.

Stacked on #90 (itself on #89); review those first.

Implementation

  • trainer/subprocess.py — iterator_cpu_bound, moved from dfine_node. The maxsize=1 queue is the interesting part: the producer can never run more than one item ahead of the bookkeeping, which is what lets a trainer alternate between two model files and know the one it is copying is not being rewritten.
  • trainer/batch_size.py — find_batch_size and the power-of-two arithmetic around it. Nothing here imports torch: the search is pure, and the trainer supplies a fits predicate that runs a real step. That is what lets a node with any framework use it.
  • trainer/metrics.py — macro_f1, category_f1, confusion_matrix_from_counts, over the {'tp','fp','fn'} shape _get_new_best_training_state returns.
  • TrainerLogic._get_executor_error_from_log gains a default implementation — yolov5_node and classification_node had byte-identical copies. It stops being abstract, so existing overrides keep working.
  • 66 new tests; the unit suite is now 197 and still runs in ~1.5s with no Learning Loop.

is_out_of_memory is not incidental

classification_node's probe catches bare RuntimeError and treats it as "too big", so a shape mismatch or a failed assert silently ends the search and the training continues at whatever size it had reached. Allocation failures do not all arrive as a framework's dedicated error type — cuDNN and cuBLAS workspaces raise a plain RuntimeError — so the message heuristics have to live somewhere. Now they live here, once.

Python version

dfine_node targets 3.12 and used PEP 695 type parameters (def f[T](...)). This library supports 3.10, so the generics are rewritten as TypeVar and ParamSpec. Caught by ruff, not at runtime — worth knowing for the phases still to come.

Deliberately not included

  • SubprocessTrainerLogic, the third trainer base class the plan sketches. With dfine_node as the only consumer, its progress protocol would be invented from one example, and yolov5_node's executor-and-log-scraping shape may not fit it. Better designed against two.
  • cuda.py (free_cuda_memory, usable_memory_bytes, limit_cuda_memory). These genuinely need torch, which would mean an optional extra and a module this repository's CI cannot exercise. Left in dfine_node until a second node wants it.

⬛ claude-opus-5[1m] · 500k tokens · $12.10

Three things every trainer needs, none of which depend on the training
framework, all of which existed only inside dfine_node or in worse form
elsewhere:

iterator_cpu_bound runs a training generator in a spawned process. A trainer
that trains in-process needs this - the event loop must stay responsive and
CUDA state must stay out of the node - and it is pure stdlib.

find_batch_size probes for the largest power-of-two batch that fits, around a
fits predicate the trainer supplies. The search is pure arithmetic; only the
predicate needs a framework, so nothing here imports torch. Three repositories
had three different implementations of this: a real probe, an estimate parsed
out of torchinfo's summary text, and a doubling loop catching bare
RuntimeError. is_out_of_memory exists because of that last one - catching
RuntimeError treats every crash as a full card.

macro_f1 scores the confusion matrix the loop stores, which was implemented
three times over two different data shapes.

_get_executor_error_from_log gets a default implementation, because
yolov5_node and classification_node had byte-identical copies of it.

The PEP 695 type parameters the generator used are rewritten as TypeVar and
ParamSpec: this library supports Python 3.10 and dfine_node targets 3.12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@klangenk
klangenk force-pushed the share-node-entrypoint branch from e1bd034 to fa8505c Compare August 21, 2026 06:42
@klangenk
klangenk force-pushed the share-trainer-helpers branch from 78fe5ba to f3b3ca9 Compare August 21, 2026 06:42
@klangenk

Copy link
Copy Markdown
Contributor Author

Superseded by #94, which combines all three phases into one PR per repository. The commits are unchanged and still readable in order.

@klangenk klangenk closed this Aug 21, 2026
@klangenk
klangenk deleted the share-trainer-helpers branch August 21, 2026 14:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant