Skip to content
View billxbf's full-sized avatar
☕
☕

Highlights

  • Pro

Organizations

@Gentopia-AI

Block or report billxbf

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
billxbf/README.md

Welcome 🖐️ I am an agent trainer at Nvidia.

My research covers agentic post-training data, infra, algo, recipe and harness, collectively optimized as a single problem.


Selected work (as First Author) >>

🎨 Skill2Env | Reinforcing agents with collective Skills. [Code] [Dataset]

🔳 Linex | The minimal all-in-one infra for Agentic RL. [Code]

⚡ FlashREINFOCE | Solving instability from async RL policy drifts and token credit mis-assignent. [Paper]

⭐ Polar | Agentic RL on ANY harnesses at scale (first in its field). [Code] [Paper]

🧠 NanoGPX | Clean collection of modern LLM architectures (RoPE, GQA, RMSNorm, MoE, SSM, etc.) in nanoGPT style. [Code]

🤖 Gentopia & GentPool | An Agent [Framework] & [Platform].

🚀 ReWOO | Token-efficient harness via decoupling reasoning from observation. [Code] [Paper]

Pinned Loading

  1. NVlabs/Skill2Env NVlabs/Skill2Env Public

    Reinforcing Agents with Collective Skills

    Python 156 17

  2. NVIDIA-NeMo/ProRL-Agent-Server NVIDIA-NeMo/ProRL-Agent-Server Public

    Agentic RL on Any Harness at Scale

    Python 853 92

  3. Linex Linex Public

    The One & Minimal Framework for Agentic RL.

    Python 2

  4. ReWOO ReWOO Public

    Decoupling Reasoning from Observations for Efficient Augmented Language Models

    Python 944 84

  5. Gentopia-AI/Gentopia Gentopia-AI/Gentopia Public

    Build Hierarchical Autonomous Agents through Config. Collaborative Growth of Specialized Agents.

    Python 329 42

  6. nanoGPX nanoGPX Public

    Forked from karpathy/nanoGPT

    Clean implementation of modern LLM recipes in nanoGPT style (RoPE, GQA, RMSNorm, MoE, SSM, etc.)

    Python 3