Skip to content

graddiff e2e: Arm A (Lilo --adv-weighting sample_mean) + Arm B (Miles calculate_per_token_loss propagated) launchers - #43

Closed
micahtyong wants to merge 29 commits into
mainfrom
devin/1789750213-adv-weighting
Closed

micahtyong wants to merge 29 commits into
mainfrom
devin/1789750213-adv-weighting

Conversation

@micahtyong

@micahtyong micahtyong commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Discriminating runs for the A6 finding (#41: Miles' effective objective is a per-sample mean because calculate_per_token_loss never reaches the Bridge model config, Lilo's client is token-mean). Stacked on #41's branch (both rebased onto main after #30 merged).

The LongRLVR client scripts (run_longrlvr_lilo_lora.py, longrlvr_dataset.py, longrlvr_comparison_common.py) were removed from main by #31 ("kept locally"), so this PR re-adds them as they were on #30's branch plus the changes below — this is what the analysis child should reuse.

Arm A — Lilo with Miles-style weighting

scripts/run_longrlvr_lilo_lora.py — new --adv-weighting {token_mean,sample_mean} (default token_mean = czzwpqau behaviour), applied after the existing std-normalisation and --per-token-loss-scale:

# token_mean:  adv_{i,t} = A_i / T
# sample_mean: adv_{i,t} = A_i / T * T/(n*len_i) = A_i / (n*len_i)
advantages_P = [adv * (total_tokens / (num_trajectories * lens_G)) ...]

T = total action tokens in the batch, n = #datums, len_i = action tokens of datum i, so Σ weights over the batch is unchanged (same global scale / lr as czzwpqau). The server loss is a plain token sum, so this is exactly Miles' 1/len_i weighting. Logged as cmp/adv_weighting.

scripts/longrlvr_dataset.py — _load_rows_from_file cycles the pinned file instead of raising when steps × source_groups exceeds its rows (30 × 24 = 720 > 360); the second epoch repeats the same order.

scripts/graddiff/e2e_lilo_client.py — detached Modal launcher reproducing the czzwpqau client container against the deployed lilo-dp2:

uv run --with modal python -m modal run --detach -e micah-dev scripts/graddiff/e2e_lilo_client.py \
  --run-name lilo-9b-16k-samplemean-30 --steps 30 --adv-weighting sample_mean --wandb-group graddiff-e2e

Arm B — Miles with the flag propagated

scripts/graddiff/miles_pertoken_bridge.patch — Miles lora/bridge.py: _setup_lora_model_via_bridge copies a hand-picked subset of args onto the Bridge provider and omitted calculate_per_token_loss; the patch adds provider.calculate_per_token_loss = args.calculate_per_token_loss and a rank-0 log line [lora/bridge] model_config.calculate_per_token_loss=%s (args=%s).
scripts/graddiff/e2e_miles_pertoken.py — training-gym MilesRecipe launcher reproducing slim-hill's 9B/16k config on the pinned prompts (sha-checked, written to parquet in order) with the patched local Miles overlaid. Only delta vs slim-hill: image radixark/miles:dev-202609050049 (the default dev-202608120325 ships an sglang without gated_launch_port, so training-gym's rollout-cell gate never binds).

Results (W&B group graddiff-e2e)

run weighting reward mean steps 0–14 steps 15–29 final resp len
Arm A Lilo eniyg4jn sample-mean 0.634 0.809 394
Arm B Miles woolen-weapon-6c8af473e1b8 (model_config.calculate_per_token_loss=True on all 8 ranks) token-mean 0.290 0.327 521
Lilo czzwpqau (control) token-mean 0.607 — 338
Miles slim-hill-371ae1cfcb8e (control) sample-mean (in effect) 0.440 — 340

Neither arm moves toward the other framework: the sample-mean vs token-mean weighting difference does not explain the 9B reward gap.

Link to Devin session: https://modal.devinenterprise.com/sessions/2e0133f9b0ae4c42b93f4d4bb1c01d92
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/2e0133f9b0ae4c42b93f4d4bb1c01d92?variant=devin
Requested by: @micahtyong

@devin-ai-integration

Copy link
Copy Markdown
Contributor

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR that start with 'DevinAI' or '@devin'.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

micahtyong and others added 27 commits September 18, 2026 17:03
… client

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…p hook

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tep entry

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…atch runs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ch remote)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…h remote)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…erride

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…rough miles_arm

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789691530-graddiff-step0 branch from 530b3a9 to d82e3f3 Compare September 18, 2026 17:03
…t; re-add LongRLVR client scripts (dropped from main by #31) with pinned-prompt cycling; lilo-dp2 client launcher

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789750213-adv-weighting branch from dfe9882 to 0402235 Compare September 18, 2026 17:04
@devin-ai-integration
devin-ai-integration Bot changed the base branch from devin/1789691530-graddiff-step0 to main September 18, 2026 17:04
…ora/bridge.py patch propagating calculate_per_token_loss

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title graddiff e2e: --adv-weighting sample_mean for the LongRLVR Lilo client (Arm A) graddiff e2e: Arm A (Lilo --adv-weighting sample_mean) + Arm B (Miles calculate_per_token_loss propagated) launchers Sep 18, 2026
@micahtyong

Copy link
Copy Markdown
Contributor Author

not checking this in. keeping around for posterity

@micahtyong micahtyong closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant