Skip to content

[Lilo] Miles LoRA backend: DP>1 support + validated Qwen3.5-9B 16k on LongRLVR - #30

Merged
micahtyong merged 16 commits into
mainfrom
devin/1789585776-longrlvr-lilo-lora-16k
Sep 18, 2026
Merged

micahtyong merged 16 commits into
mainfrom
devin/1789585776-longrlvr-lilo-lora-16k

Conversation

@micahtyong

@micahtyong micahtyong commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Backend + definition changes needed for the Lilo-vs-Miles 9B LongRLVR parity runs (results in the W&B report Validating Multi-client LoRA on Lilo). This PR is the library surface those runs validated.

Miles LoRA backend: allow data parallelism (world_size % tp == 0 instead of == tp)

  • MilesBackendConfig.data_parallel_size = world_size // tensor_model_parallel_size.
  • Miles shards a forward_backward batch across DP ranks, so ragged batches are padded rather than rejected:
    pad_slot_rows(slot_rows, multiple=dp)  # appends zero-weight 2-token rows
    # pad row: {"tokens": t[:2], "target_len": 1, "target_tokens": t[1:2], "weights": [0.0], ...}
    MilesCommandBackend.forward_backward trims the padded outputs back to the real row count.
  • Checkpoint metadata topology now records data_parallel_size; _validate_checkpoint treats a missing key in pre-existing checkpoints as 1, so old TP-only checkpoints still restore on an unchanged deployment while DP mismatches are rejected.

Definitions

  • qwen3_5_9b_miles_lora_16k (8×H100, TP8, catalog-visible) and qwen3_5_9b_miles_lora_16k_dp2 (TP4×DP2, hidden) — identical except id/TP/visibility; the DP2 one is what the parity runs used (81 s/step vs 107 s on TP8).
  • serve.py/app.py: APP_NAME = os.environ.get("LILO_APP_NAME", "lilo") so a second app (e.g. lilo-dp2) can be deployed alongside prod without code changes.

Tests: DP config acceptance/rejection, pad/trim behavior, topology metadata, legacy-checkpoint compatibility.

Link to Devin session: https://modal.devinenterprise.com/sessions/8d6804dd317c41329d77eb2c427dda6f
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/8d6804dd317c41329d77eb2c427dda6f?variant=devin
Requested by: @micahtyong


Devin Review

@devin-ai-integration

Copy link
Copy Markdown
Contributor

🤖 Devin AI Engineer

I'll be helping with this pull request! Here's what you should know:

✅ I will automatically:

  • Address comments on this PR that start with 'DevinAI' or '@devin'.
  • Look at CI failures and help fix them

Note: I can only respond to comments from users who have write access to this repository.

⚙️ Control Options:

  • Disable automatic comment, CI, and merge conflict monitoring

micahtyong and others added 14 commits September 17, 2026 22:43
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…APP_NAME override

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…A backend

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…den definitions

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1789585776-longrlvr-lilo-lora-16k branch from 8822f69 to 0737bab Compare September 17, 2026 22:48
@devin-ai-integration
devin-ai-integration Bot changed the base branch from codex/deterministic-training to main September 17, 2026 22:55
@devin-ai-integration
devin-ai-integration Bot marked this pull request as ready for review September 17, 2026 22:55
@micahtyong

Copy link
Copy Markdown
Contributor Author

/devin review

@devin-ai-integration

Copy link
Copy Markdown
Contributor

Starting Devin Review.

Devin Review

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 8 potential issues.

Devin Review

Comment thread src/lilo/backends/miles_runtime/data.py Outdated
Comment thread src/lilo/backends/miles_lora.py
Comment thread scripts/log_miles_baseline_to_wandb.py Outdated
Comment thread scripts/longrlvr_dataset.py Outdated
Comment thread scripts/longrlvr_dataset.py Outdated
Comment thread src/lilo/backends/miles_runtime/data.py
Comment thread scripts/longrlvr_dataset.py Outdated
Comment thread scripts/run_longrlvr_lilo_lora.py Outdated
micahtyong and others added 2 commits September 18, 2026 02:21
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… topology

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration devin-ai-integration Bot changed the title LongRLVR-16k Qwen3.5-9B LoRA recipe + comparison metrics for Lilo vs Miles parity runs Miles LoRA backend: DP>1 support + Qwen3.5-9B 16k definitions (validated on LongRLVR Lilo-vs-Miles runs) Sep 18, 2026
slot_rows: tuple[tuple[int, dict[str, Any]], ...],
multiple: int,
) -> tuple[tuple[int, dict[str, Any]], ...]:
"""Pad to a multiple of multiple with zero-weight rows for Miles DP sharding."""

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When DP > 1 (two replicas of the model), Miles will split the batch into one shard per DP rank. Since every rank must receive at least one row, this method is used to ensure the number of rows is divisible by DP. Without this, a DP rank would hang

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This config is validated with an actual e2e miles vs lilo run here

@micahtyong micahtyong changed the title Miles LoRA backend: DP>1 support + Qwen3.5-9B 16k definitions (validated on LongRLVR Lilo-vs-Miles runs) [Lilo] Miles LoRA backend: DP>1 support + validated Qwen3.5-9B 16k on LongRLVR Sep 18, 2026
@micahtyong

Copy link
Copy Markdown
Contributor Author

Going to merge this in @kailash109. Next one is #31 which is where we added context parallelism support for up to 128K.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant