Conversation
Three things every trainer needs, none of which depend on the training framework, all of which existed only inside dfine_node or in worse form elsewhere: iterator_cpu_bound runs a training generator in a spawned process. A trainer that trains in-process needs this - the event loop must stay responsive and CUDA state must stay out of the node - and it is pure stdlib. find_batch_size probes for the largest power-of-two batch that fits, around a fits predicate the trainer supplies. The search is pure arithmetic; only the predicate needs a framework, so nothing here imports torch. Three repositories had three different implementations of this: a real probe, an estimate parsed out of torchinfo's summary text, and a doubling loop catching bare RuntimeError. is_out_of_memory exists because of that last one - catching RuntimeError treats every crash as a full card. macro_f1 scores the confusion matrix the loop stores, which was implemented three times over two different data shapes. _get_executor_error_from_log gets a default implementation, because yolov5_node and classification_node had byte-identical copies of it. The PEP 695 type parameters the generator used are rewritten as TypeVar and ParamSpec: this library supports Python 3.10 and dfine_node targets 3.12. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
klangenk
force-pushed
the
share-node-entrypoint
branch
from
August 21, 2026 06:42
e1bd034 to
fa8505c
Compare
klangenk
force-pushed
the
share-trainer-helpers
branch
from
August 21, 2026 06:42
78fe5ba to
f3b3ca9
Compare
Contributor
Author
|
Superseded by #94, which combines all three phases into one PR per repository. The commits are unchanged and still readable in order. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Three things every trainer needs, none of which depend on the training framework, all of which existed only inside
dfine_node— or in worse form elsewhere:dfine_nodesolved this with a generator running in a spawned process; it is pure stdlib and nothing about it is D-FINE.dfine_node), an estimate parsed out oftorchinfo's summary text (yolov5_node), and a doubling loop catching bareRuntimeError(classification_node).Stacked on #90 (itself on #89); review those first.
Implementation
trainer/subprocess.py—iterator_cpu_bound, moved fromdfine_node. Themaxsize=1queue is the interesting part: the producer can never run more than one item ahead of the bookkeeping, which is what lets a trainer alternate between two model files and know the one it is copying is not being rewritten.trainer/batch_size.py—find_batch_sizeand the power-of-two arithmetic around it. Nothing here imports torch: the search is pure, and the trainer supplies afitspredicate that runs a real step. That is what lets a node with any framework use it.trainer/metrics.py—macro_f1,category_f1,confusion_matrix_from_counts, over the{'tp','fp','fn'}shape_get_new_best_training_statereturns.TrainerLogic._get_executor_error_from_loggains a default implementation —yolov5_nodeandclassification_nodehad byte-identical copies. It stops being abstract, so existing overrides keep working.is_out_of_memoryis not incidentalclassification_node's probe catches bareRuntimeErrorand treats it as "too big", so a shape mismatch or a failed assert silently ends the search and the training continues at whatever size it had reached. Allocation failures do not all arrive as a framework's dedicated error type — cuDNN and cuBLAS workspaces raise a plainRuntimeError— so the message heuristics have to live somewhere. Now they live here, once.Python version
dfine_nodetargets 3.12 and used PEP 695 type parameters (def f[T](...)). This library supports 3.10, so the generics are rewritten asTypeVarandParamSpec. Caught by ruff, not at runtime — worth knowing for the phases still to come.Deliberately not included
SubprocessTrainerLogic, the third trainer base class the plan sketches. Withdfine_nodeas the only consumer, its progress protocol would be invented from one example, andyolov5_node's executor-and-log-scraping shape may not fit it. Better designed against two.cuda.py(free_cuda_memory,usable_memory_bytes,limit_cuda_memory). These genuinely need torch, which would mean an optional extra and a module this repository's CI cannot exercise. Left indfine_nodeuntil a second node wants it.⬛ claude-opus-5[1m] · 500k tokens · $12.10