Welcome 🖐️ I am an agent trainer at Nvidia.
My research covers agentic post-training data, infra, algo, recipe and harness, collectively optimized as a single problem.
Selected work (as First Author) >>
🎨 Skill2Env | Reinforcing agents with collective Skills. [Code] [Dataset]
🔳 Linex | The minimal all-in-one infra for Agentic RL. [Code]
⚡ FlashREINFOCE | Solving instability from async RL policy drifts and token credit mis-assignent. [Paper]
⭐ Polar | Agentic RL on ANY harnesses at scale (first in its field). [Code] [Paper]
🧠 NanoGPX | Clean collection of modern LLM architectures (RoPE, GQA, RMSNorm, MoE, SSM, etc.) in nanoGPT style. [Code]
🤖 Gentopia & GentPool | An Agent [Framework] & [Platform].
🚀 ReWOO | Token-efficient harness via decoupling reasoning from observation. [Code] [Paper]




