[Lilo] Miles LoRA backend: add support for context parallelism; check in validated Qwen3.8-27B 16k–128k definitions - #31
Conversation
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
8822f69 to
0737bab
Compare
4d65526 to
f3f1326
Compare
de9c452 to
bad001c
Compare
|
/devin review |
There was a problem hiding this comment.
Most, if not all, of the changes in this file are temporary. The miles runtime currently has no support for context parallelism, so we're handrolling a lot of the key methods ourselves. Per the RadixArk roadmap here, they do want to add CP support.
I'm having a Devin agent put some candidate upstream PRs in parallel, but I still would like to check in these methods as they are validated e2e here + Engram may start using this tomorrow.
…O_APP_NAME to trainers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…r steps Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…grades to H200 anyway) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ontainers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ontainers Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…32 logits upcast) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ne-miles-27b runs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tring Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…, evict retrieved results eagerly Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…143b31 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Completed results were dropped as soon as they had been handed out once and the per-model cap was exceeded, so a retried retrieve_future (for example after the control plane fails to reach the engine and returns 503) resolved to a LOST request and killed the training client. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
d95e8cc to
17b8b1a
Compare
…t; re-add LongRLVR client scripts (dropped from main by #31) with pinned-prompt cycling; lilo-dp2 client launcher Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Merging now esp since this is validated e2e; can address follow-ups after |
…P changes) Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Summary
Lilo side of the Lilo-vs-raw-Miles Qwen3.8-27B LoRA (r32) long-context comparison: adds context parallelism to the Miles LoRA backend, plus three 27B definitions (16k / 64k / 128k, one 8×H200 node). Stacked on #30 (DP>1 +
LILO_APP_NAME). Experiment scripts and the markdown report were removed from this PR (kept locally); results live in this W&B report. The unvalidated 256k definition and multi-node trainer support (Ray cluster bring-up,actor_num_nodes, cross-node checkpoint volume sync) moved to a stacked follow-up PR (#39).Context parallelism on the Miles LoRA path (
miles_config.py,miles_runtime/actor.py,runtime.py)MilesBackendConfig.context_parallel_size(default 1):data_parallel_size = world // (tp*cp), validated, emitted as--context-parallel-size, recorded in checkpoint topology metadata. Checkpoints written before this PR lack the key;_validate_checkpointtreats a missingcontext_parallel_sizeas 1 (same asdata_parallel_size).slice_log_prob_with_cpon rewards/masks), but its Tinker loss path (loss_hub/tinker_losses.py) is not: it zips the local logprob shard against the client's full-length per-tokenadvantages/loss_weights, hence Miles' "no multi-LoRA with CP>1" guard.actor.pypatches, at import:_gather_tinker_logprobs_across_cp:_pad_local_shardplaces each rank's shard at its global response positions (via Miles'get_logits_and_tokens_offset_with_cp), then all-reduce a detached copy —padded + (summed - padded.detach())— so every rank sees full-response logprobs but gradient flows only through its own shard.slice_log_prob_with_cp→ identity for the Tinker path (tensors are now full-length on every rank).tinker_lossesCP-aware and lift the multi-LoRA+CP guard) and drop the Lilo patches once the deployed Miles image has it; the patch here is what let the 64k/128k runs above happen.runtime.py::_allow_context_parallel_multi_lorabypasses that guard._preserve_advantages_in_dp_shardspatch is gone: Miles main packagesadvantagesinto DP shards since radixark/milesada143b31(2026-09-16), and Lilo resolvesMILES_REF=main.LiloMilesTrainRayActoralso disables inherited Qwen MTP heads for LoRA (they inject an auxiliary loss even for zero-weight datums) and printslilo_memorypeak-memory lines after each op.Engine (
engine/server.py)max_resultsonce the client has retrieved them (_ModelState.retrieved), so a slow client behind a long step can't lose an unreadforward_backwardresult. Eviction (_evict_retrieved) runs both on completion and on retrieval, so retrieved results don't linger when no later op completes.result_retention_s(default 900 s) after it finishes, and the count cap is nowmax_results=1024(fb results are small). "Retrieved once + cap exceeded" was enough to forget a result that the client still needed, so a retriedretrieve_future— e.g. after the control plane momentarily can't reach the engine and answers503— resolved toLOSTand killed a 128k run mid-training.Definitions / app
qwen3_8_27b_miles_lora_{16k,64k,128k}.py,GPU_TYPE="H200"; 64k+ areCATALOG_VISIBLE=False.deployment.pyforwardsLILO_APP_NAMEinto the trainer env so thelilo-27bapp is independent of prod.Not covered by tests
The CP logprob gather/unslice patches bind Miles private symbols and are exercised only by the 64k/128k runs above, not by unit tests (running them needs Megatron + a CP group).
Link to Devin session: https://modal.devinenterprise.com/sessions/58fc0f92cbec4882bf83bb464c8e3d22
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/58fc0f92cbec4882bf83bb464c8e3d22?variant=devin
Requested by: @micahtyong