graddiff e2e: Arm A (Lilo --adv-weighting sample_mean) + Arm B (Miles calculate_per_token_loss propagated) launchers - #43
Closed
micahtyong wants to merge 29 commits into
Closed
micahtyong wants to merge 29 commits into
micahtyong wants to merge 29 commits into
Conversation
Contributor
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
… client Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…p hook Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tep entry Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…atch runs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ch remote) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…h remote) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…erride Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…rough miles_arm Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
devin-ai-integration
Bot
force-pushed
the
devin/1789691530-graddiff-step0
branch
from
September 18, 2026 17:03
530b3a9 to
d82e3f3
Compare
…t; re-add LongRLVR client scripts (dropped from main by #31) with pinned-prompt cycling; lilo-dp2 client launcher Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
devin-ai-integration
Bot
force-pushed
the
devin/1789750213-adv-weighting
branch
from
September 18, 2026 17:04
dfe9882 to
0402235
Compare
devin-ai-integration
Bot
changed the base branch from
devin/1789691530-graddiff-step0
to
main
September 18, 2026 17:04
…ora/bridge.py patch propagating calculate_per_token_loss Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Contributor
Author
|
not checking this in. keeping around for posterity |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Discriminating runs for the A6 finding (#41: Miles' effective objective is a per-sample mean because
calculate_per_token_lossnever reaches the Bridge model config, Lilo's client is token-mean). Stacked on #41's branch (both rebased ontomainafter #30 merged).The LongRLVR client scripts (
run_longrlvr_lilo_lora.py,longrlvr_dataset.py,longrlvr_comparison_common.py) were removed frommainby #31 ("kept locally"), so this PR re-adds them as they were on #30's branch plus the changes below — this is what the analysis child should reuse.Arm A — Lilo with Miles-style weighting
scripts/run_longrlvr_lilo_lora.py— new--adv-weighting {token_mean,sample_mean}(defaulttoken_mean= czzwpqau behaviour), applied after the existing std-normalisation and--per-token-loss-scale:T = total action tokens in the batch, n = #datums, len_i = action tokens of datum i, so Σ weights over the batch is unchanged (same global scale / lr as czzwpqau). The server loss is a plain token sum, so this is exactly Miles' 1/len_i weighting. Logged as
cmp/adv_weighting.scripts/longrlvr_dataset.py—_load_rows_from_filecycles the pinned file instead of raising whensteps × source_groupsexceeds its rows (30 × 24 = 720 > 360); the second epoch repeats the same order.scripts/graddiff/e2e_lilo_client.py— detached Modal launcher reproducing the czzwpqau client container against the deployedlilo-dp2:Arm B — Miles with the flag propagated
scripts/graddiff/miles_pertoken_bridge.patch— Mileslora/bridge.py:_setup_lora_model_via_bridgecopies a hand-picked subset ofargsonto the Bridge provider and omittedcalculate_per_token_loss; the patch addsprovider.calculate_per_token_loss = args.calculate_per_token_lossand a rank-0 log line[lora/bridge] model_config.calculate_per_token_loss=%s (args=%s).scripts/graddiff/e2e_miles_pertoken.py— training-gymMilesRecipelauncher reproducing slim-hill's 9B/16k config on the pinned prompts (sha-checked, written to parquet in order) with the patched local Miles overlaid. Only delta vs slim-hill: imageradixark/miles:dev-202609050049(the defaultdev-202608120325ships an sglang withoutgated_launch_port, so training-gym's rollout-cell gate never binds).Results (W&B group
graddiff-e2e)model_config.calculate_per_token_loss=Trueon all 8 ranks)Neither arm moves toward the other framework: the sample-mean vs token-mean weighting difference does not explain the 9B reward gap.
Link to Devin session: https://modal.devinenterprise.com/sessions/2e0133f9b0ae4c42b93f4d4bb1c01d92
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/2e0133f9b0ae4c42b93f4d4bb1c01d92?variant=devin
Requested by: @micahtyong