Skip to content

Log PPO/CISPO clip fraction from the trainer loss - #57

Open
micahtyong wants to merge 1 commit into
mainfrom
devin/1790101346-clip-fraction-metrics
Open

micahtyong wants to merge 1 commit into
mainfrom
devin/1790101346-clip-fraction-metrics

Conversation

@micahtyong

Copy link
Copy Markdown
Contributor

Summary

Neither Lilo nor Miles emitted a clip ratio, so there was no way to check whether the two stacks take equally on-policy updates (the question behind the Lilo-vs-Miles reward and throughput comparison). This adds clip_fraction:mean (plus its clipped_tokens:sum / loss_tokens:sum numerator and denominator) to ForwardBackwardOutput.metrics for both the Megatron path and the Miles-runtime path. Loss values and gradients are unchanged — the count is derived from tensors the loss already computes, and is detached.

A token counts as clipped only where clipping actually bites, not merely where the ratio leaves [low, high]; with a negative advantage the out-of-band branch is often the larger one and still contributes normal gradient:

unclipped_objective = probability_ratio * advantages
clipped_objective = clamp(probability_ratio, low, high) * advantages
objective = torch.minimum(unclipped_objective, clipped_objective)   # unchanged
clipped_tokens = count(clipped_objective < unclipped_objective, where mask != 0)

For CISPO, where the clamped coefficient is detached and multiplies logprobs, the equivalent condition is coefficient != ratio.detach(). Losses that never clip (cross entropy, importance sampling, DRO) emit no clip keys at all, so the metric's presence means "this loss clips".

_loss now returns a third value (the clipped-token count). It is accumulated per job alongside loss/tokens through the existing distributed payload merge, so it aggregates across TP/CP/DP like the metrics next to it. The Miles-runtime path forwards the per-datum clipped_tokens/loss_tokens that the companion Miles change returns and aggregates them per client the same way.

loss_tokens counts masked (eligible) tokens, so clip_fraction is comparable across the two backends and across runs with different batch shapes.

Testing

uv run pytest -q — 518 passed; the 4 failures in tests/scoped/test_lifecycle.py are pre-existing and fail identically on main. New tests cover PPO/CISPO clipped-token counting, the losses that must report nothing, and exclusion of zero-mask tokens.

Link to Devin session: https://modal.devinenterprise.com/sessions/e95e74695d884d4f8535471bae0d7eee
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/e95e74695d884d4f8535471bae0d7eee?variant=devin
Requested by: @micahtyong

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access that start with 'DevinAI' or '@devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant