Log PPO/CISPO clip fraction from the trainer loss - #57
Open
micahtyong wants to merge 1 commit into
Open
micahtyong wants to merge 1 commit into
micahtyong wants to merge 1 commit into
Conversation
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Contributor
|
I'll fix CI failures and address comments from users with write access that start with 'DevinAI' or '@devin'.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Neither Lilo nor Miles emitted a clip ratio, so there was no way to check whether the two stacks take equally on-policy updates (the question behind the Lilo-vs-Miles reward and throughput comparison). This adds
clip_fraction:mean(plus itsclipped_tokens:sum/loss_tokens:sumnumerator and denominator) toForwardBackwardOutput.metricsfor both the Megatron path and the Miles-runtime path. Loss values and gradients are unchanged — the count is derived from tensors the loss already computes, and is detached.A token counts as clipped only where clipping actually bites, not merely where the ratio leaves
[low, high]; with a negative advantage the out-of-band branch is often the larger one and still contributes normal gradient:For CISPO, where the clamped coefficient is detached and multiplies
logprobs, the equivalent condition iscoefficient != ratio.detach(). Losses that never clip (cross entropy, importance sampling, DRO) emit no clip keys at all, so the metric's presence means "this loss clips"._lossnow returns a third value (the clipped-token count). It is accumulated per job alongsideloss/tokensthrough the existing distributed payload merge, so it aggregates across TP/CP/DP like the metrics next to it. The Miles-runtime path forwards the per-datumclipped_tokens/loss_tokensthat the companion Miles change returns and aggregates them per client the same way.loss_tokenscounts masked (eligible) tokens, soclip_fractionis comparable across the two backends and across runs with different batch shapes.Testing
uv run pytest -q— 518 passed; the 4 failures intests/scoped/test_lifecycle.pyare pre-existing and fail identically onmain. New tests cover PPO/CISPO clipped-token counting, the losses that must report nothing, and exclusion of zero-mask tokens.Link to Devin session: https://modal.devinenterprise.com/sessions/e95e74695d884d4f8535471bae0d7eee
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/e95e74695d884d4f8535471bae0d7eee?variant=devin
Requested by: @micahtyong