Skip to content

RL + dLLM #3706

Description

@akoumpa

BlockjustGRPO-fast + nemotron-lab-diffusion + AM + RL

AM integration Plan

https://docs.google.com/document/d/1xC7WKKNLlI5hEIly9--WelhO92NYKbgBMCqo62Po_1s/edit?usp=sharing

draft PR: zyzhou5/RL#4

**Plan **

  1. Working e2e pipeline with AM backend (based on Sajad's current design https://github.com/sajadn/diffusion_RL/commits/dllm_clean/). Planning to do the following test
    1. model loading
    2. model forward/backward
    3. loss parity (algorithm)
    4. vLLM supports
    5. 100 steps of training and compare reward
  2. Other features
    1. CP
  3. Other algorithms
    1. JustGRPO

Other things

  1. Sajad mentioned about a new design of the dLLM RL pipeline in Nemo-RL (https://docs.google.com/document/d/1HHzRpLfsRalQkqLGFe0vljZjRO7dwKaKWH43jZPxCu4/edit?tab=t.0), we will need to refactor on the AM side as well once that new design is implemented in Nemo-RL
    1. Timeline - might happen in a few weeks
  2. Check if huiying's refactor on AM + Nemo-RL will cause any issue refactor(engine): training Engine API #3614
    1. Timeline - estimated to be done by the end of September

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions