Skip to content

[Lilo] Miles LoRA backend: add support for context parallelism; check in validated Qwen3.8-27B 16k–128k definitions - #31

Merged
micahtyong merged 29 commits into
mainfrom
devin/1789604737-lilo-27b-longcontext
Sep 18, 2026
Merged

micahtyong merged 29 commits into
mainfrom
devin/1789604737-lilo-27b-longcontext

Conversation

@micahtyong

@micahtyong micahtyong commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Lilo side of the Lilo-vs-raw-Miles Qwen3.8-27B LoRA (r32) long-context comparison: adds context parallelism to the Miles LoRA backend, plus three 27B definitions (16k / 64k / 128k, one 8×H200 node). Stacked on #30 (DP>1 + LILO_APP_NAME). Experiment scripts and the markdown report were removed from this PR (kept locally); results live in this W&B report. The unvalidated 256k definition and multi-node trainer support (Ray cluster bring-up, actor_num_nodes, cross-node checkpoint volume sync) moved to a stacked follow-up PR (#39).

ctx trainer rollout Lilo run raw Miles run
16k 8×H200 TP4×DP2 H200 TP1 ×8 uwcvqxpb vv9l76wt
64k 8×H200 TP4×CP2 H200 TP1 ×8 h2yt8e23 5uvc8lks
128k 8×H200 TP2×CP4 H200 TP2 ×4 t9feh1se nakynr4k

Context parallelism on the Miles LoRA path (miles_config.py, miles_runtime/actor.py, runtime.py)

  • MilesBackendConfig.context_parallel_size (default 1): data_parallel_size = world // (tp*cp), validated, emitted as --context-parallel-size, recorded in checkpoint topology metadata. Checkpoints written before this PR lack the key; _validate_checkpoint treats a missing context_parallel_size as 1 (same as data_parallel_size).
  • Miles' native losses are CP-aware (local zigzag logprob shard + slice_log_prob_with_cp on rewards/masks), but its Tinker loss path (loss_hub/tinker_losses.py) is not: it zips the local logprob shard against the client's full-length per-token advantages/loss_weights, hence Miles' "no multi-LoRA with CP>1" guard. actor.py patches, at import:
    • _gather_tinker_logprobs_across_cp: _pad_local_shard places each rank's shard at its global response positions (via Miles' get_logits_and_tokens_offset_with_cp), then all-reduce a detached copy — padded + (summed - padded.detach()) — so every rank sees full-response logprobs but gradient flows only through its own shard.
    • slice_log_prob_with_cp → identity for the Tinker path (tensors are now full-length on every rank).
    • To be upstreamed to Miles: CP support for the Tinker/multi-LoRA loss path is an open item on the Miles roadmap (Multi-LoRA / Tinker API roadmap radixark/miles#3284, Stage 2 "Context parallel support"). We will upstream this as a native change (make tinker_losses CP-aware and lift the multi-LoRA+CP guard) and drop the Lilo patches once the deployed Miles image has it; the patch here is what let the 64k/128k runs above happen.
  • runtime.py::_allow_context_parallel_multi_lora bypasses that guard.
  • The earlier _preserve_advantages_in_dp_shards patch is gone: Miles main packages advantages into DP shards since radixark/miles ada143b31 (2026-09-16), and Lilo resolves MILES_REF=main.
  • LiloMilesTrainRayActor also disables inherited Qwen MTP heads for LoRA (they inject an auxiliary loss even for zero-weight datums) and prints lilo_memory peak-memory lines after each op.

Engine (engine/server.py)

  • Completed results are only evicted past max_results once the client has retrieved them (_ModelState.retrieved), so a slow client behind a long step can't lose an unread forward_backward result. Eviction (_evict_retrieved) runs both on completion and on retrieval, so retrieved results don't linger when no later op completes.
  • Retrieval alone no longer makes a result droppable: a completed result is kept for result_retention_s (default 900 s) after it finishes, and the count cap is now max_results=1024 (fb results are small). "Retrieved once + cap exceeded" was enough to forget a result that the client still needed, so a retried retrieve_future — e.g. after the control plane momentarily can't reach the engine and answers 503 — resolved to LOST and killed a 128k run mid-training.

Definitions / app

  • qwen3_8_27b_miles_lora_{16k,64k,128k}.py, GPU_TYPE="H200"; 64k+ are CATALOG_VISIBLE=False.
  • deployment.py forwards LILO_APP_NAME into the trainer env so the lilo-27b app is independent of prod.

Not covered by tests

The CP logprob gather/unslice patches bind Miles private symbols and are exercised only by the 64k/128k runs above, not by unit tests (running them needs Megatron + a CP group).

Link to Devin session: https://modal.devinenterprise.com/sessions/58fc0f92cbec4882bf83bb464c8e3d22
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/58fc0f92cbec4882bf83bb464c8e3d22?variant=devin
Requested by: @micahtyong


Devin Review

@devin-ai-integration

Copy link
Copy Markdown
Contributor

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR that start with 'DevinAI' or '@devin'.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

@devin-ai-integration devin-ai-integration Bot changed the title Qwen3.8-27B Miles LoRA long-context ramp (lilo-27b app) Lilo (Miles-backend) Qwen3.8-27B LoRA long-context ramp (lilo-27b app) Sep 17, 2026
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789585776-longrlvr-lilo-lora-16k branch from 8822f69 to 0737bab Compare September 17, 2026 22:48
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789604737-lilo-27b-longcontext branch from 4d65526 to f3f1326 Compare September 17, 2026 22:55
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789604737-lilo-27b-longcontext branch from de9c452 to bad001c Compare September 18, 2026 03:19
@devin-ai-integration devin-ai-integration Bot changed the title Lilo (Miles-backend) Qwen3.8-27B LoRA long-context ramp (lilo-27b app) Miles LoRA backend: context parallelism + multi-node trainers; Qwen3.8-27B 16k–256k definitions Sep 18, 2026
@micahtyong
micahtyong marked this pull request as ready for review September 18, 2026 03:31
@micahtyong

Copy link
Copy Markdown
Contributor Author

/devin review

@devin-ai-integration

Copy link
Copy Markdown
Contributor

Starting Devin Review.

Devin Review

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 5 potential issues.

Devin Review

Comment thread src/lilo/engine/server.py
Comment thread src/lilo/backends/miles_config.py Outdated
Comment thread src/lilo/backends/miles_lora.py
Comment thread src/lilo/backends/miles_runtime/actor.py Outdated
Comment thread src/lilo/providers/modal/ray_cluster.py Outdated
@devin-ai-integration devin-ai-integration Bot changed the title Miles LoRA backend: context parallelism + multi-node trainers; Qwen3.8-27B 16k–256k definitions Miles LoRA backend: context parallelism; Qwen3.8-27B 16k–128k definitions Sep 18, 2026

@micahtyong micahtyong Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Most, if not all, of the changes in this file are temporary. The miles runtime currently has no support for context parallelism, so we're handrolling a lot of the key methods ourselves. Per the RadixArk roadmap here, they do want to add CP support.

I'm having a Devin agent put some candidate upstream PRs in parallel, but I still would like to check in these methods as they are validated e2e here + Engram may start using this tomorrow.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validated here

@micahtyong micahtyong changed the title Miles LoRA backend: context parallelism; Qwen3.8-27B 16k–128k definitions [Lilo] Miles LoRA backend: add support for context parallelism; check in validated Qwen3.8-27B 16k–128k definitions Sep 18, 2026
@devin-ai-integration
devin-ai-integration Bot changed the base branch from devin/1789585776-longrlvr-lilo-lora-16k to main September 18, 2026 16:55
micahtyong and others added 11 commits September 18, 2026 16:55
…O_APP_NAME to trainers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…r steps

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…grades to H200 anyway)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ontainers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ontainers

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…32 logits upcast)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
micahtyong and others added 18 commits September 18, 2026 16:55
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ne-miles-27b runs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tring

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…, evict retrieved results eagerly

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…143b31

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Completed results were dropped as soon as they had been handed out once and the per-model cap was exceeded, so a retried retrieve_future (for example after the control plane fails to reach the engine and returns 503) resolved to a LOST request and killed the training client.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789604737-lilo-27b-longcontext branch from d95e8cc to 17b8b1a Compare September 18, 2026 16:55
devin-ai-integration Bot added a commit that referenced this pull request Sep 18, 2026
…t; re-add LongRLVR client scripts (dropped from main by #31) with pinned-prompt cycling; lilo-dp2 client launcher

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@micahtyong

Copy link
Copy Markdown
Contributor Author

Merging now esp since this is validated e2e; can address follow-ups after

@micahtyong
micahtyong merged commit 8b77a28 into main Sep 18, 2026
2 checks passed
devin-ai-integration Bot added a commit that referenced this pull request Sep 18, 2026
…P changes)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant