Skip to content

Remove forward_backward batch cap from 256k Qwen3.8-27B definition - #64

Merged
micahtyong merged 1 commit into
mainfrom
devin/1790189933-remove-256k-fb-cap
Sep 23, 2026
Merged

micahtyong merged 1 commit into
mainfrom
devin/1790189933-remove-256k-fb-cap

Conversation

@micahtyong

Copy link
Copy Markdown
Contributor

Summary

Drops MAX_FORWARD_BACKWARD_BATCH = 1 from qwen3_8_27b_miles_lora_256k.py so the engine uses its default (uncapped) coalescing: a step's 128 SDK chunks become one forward_backward op with batch 128 instead of 128 single-datum ops.

The cap was added in #39 as a guard against a suspected process-group wedge from long collective chains across both nodes. A 3-step probe on the matched 2-node TP2×CP8 topology with the cap removed ran clean (forward_backward_calls = 1/step, mean_batch = 128, no NCCL timeout/wedge/OOM), so the guard isn't needed. Performance is neutral: trainer-side forward_backward_s ≈ 2256 s in both configs; client train time 2470 s vs 2404 s capped over a 3-step sample.

Probe: https://wandb.ai/modal-labs/miles-lora-longcontext/runs/lilo-27b-256k-2n-nocap-3step-1790117905-1790117910

Link to Devin session: https://modal.devinenterprise.com/sessions/d017b320e7e1430aadae6c51e49c3029
Open in Devin Desktop: https://modal.devinenterprise.com/desktop/session/d017b320e7e1430aadae6c51e49c3029?variant=devin
Requested by: @micahtyong

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access that start with 'DevinAI' or '@devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

@micahtyong
micahtyong merged commit a7e2841 into main Sep 23, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant