Skip to content
jjjabcdPublic

About

Agentic Multi-objective Molecular Optimization via Dynamic Routing of Property-specific Editor Networks

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

MORSE: Agentic Multi-objective Molecular Optimization via Dynamic Routing of Property-specific Editor Networks

AI4Sci Korea 2026 — Extended Abstract
September 28–October 1, 2026 · Seoul Dragon City, Seoul, Republic of Korea

MORSE overview

Getting Started

1. Clone the Repository

Property-specific Editor checkpoints are distributed through Git LFS. Install Git LFS before cloning the repository.

On Ubuntu or Debian:

sudo apt update
sudo apt install git-lfs

Alternatively, install Git LFS from conda-forge in an active Conda environment:

conda install -c conda-forge git-lfs

Initialize Git LFS and clone MORSE:

git lfs install
git clone https://github.com/jjjabcd/MORSE.git
cd MORSE

2. Environment Setup

Create a Python 3.11 Conda environment and install MORSE with all property scorers:

conda create -n morse python=3.11 -y
conda activate morse
pip install -e .

3. Download Checkpoints and Models

Download the Editor checkpoints from the project root:

git lfs pull --include="ckpts/editor/**/*.pt"

The released Router checkpoints are regular Git files and do not require an additional download.

Download the DRD2 classifier:

bash scripts/fetch_drd2_model.sh

Download the LLM state encoder from Qwen/Qwen2.5-7B-Instruct:

python scripts/fetch_encoder.py \
  --model Qwen/Qwen2.5-7B-Instruct \
  --cache-dir ckpts/llm

The BBBP scorer uses ADMET-AI and may download its model automatically on first use.

Checkpoint structure
ckpts/
├── editor/
│   ├── foundation/
│   │   ├── foundation_final.pt
│   │   ├── configs.csv
│   │   └── char2idx.csv
│   └── adapters/
│       ├── qed_lora.pt
│       ├── drd2_lora.pt
│       ├── plogp_lora.pt
│       └── bbbp_lora.pt
├── router/
│   ├── bdp/best.pt
│   ├── bdq/best.pt
│   └── bpq/best.pt
└── llm/                       # Local Hugging Face cache

4. Training

MORSE trains one Router for each objective combination. Run one of the following scripts:

# BBBP + DRD2 + penalized logP
bash scripts/bdp_train.sh

# BBBP + DRD2 + QED
bash scripts/bdq_train.sh

# BBBP + penalized logP + QED
bash scripts/bpq_train.sh
Training arguments

Arguments supplied after a training script override the corresponding values in the task config.

Argument Value Description
--data_dir Directory path Router training dataset directory
--checkpoints Directory path Editor checkpoint root
--encoder_model Hugging Face model ID or local path Frozen LLM state encoder
--llm_cache_dir Directory path LLM download/cache directory
--total_episodes Positive integer Number of training episodes
--batch_episodes Positive integer Episodes collected per training step
--max_length Positive integer Maximum encoder token length
--max_timesteps Positive integer Maximum Router steps per episode
--K Positive integer Number of Editor candidates when Editor beam search is disabled
--editor_num_beams Positive integer Editor decoding beam width; 1 uses sampling
--min_similarity Float in [0, 1] Similarity threshold used by the reward
--gamma Float in [0, 1] Reward discount factor
--lr Positive float Router learning rate
--grad_clip Positive float Gradient clipping norm
--value_coef Non-negative float Value loss coefficient
--entropy_coef Non-negative float Entropy bonus coefficient
--dropout Float in [0, 1) Router dropout probability
--policy_arch no_numeric or original Router policy input architecture
--ppo_epochs Positive integer PPO update epochs per collected batch
--clip_eps Positive float PPO clipping epsilon
--ppo_buffer_size Positive integer PPO replay-buffer capacity in training steps
--seed Integer Random seed
--router_device cuda:N, cpu, or auto LLM encoder and Router device
--qed_editor cuda:N or cpu QED Editor device
--drd2_editor cuda:N or cpu DRD2 Editor device
--plogp_editor cuda:N or cpu Penalized-logP Editor device
--bbbp_editor cuda:N or cpu BBBP Editor device
--bbbp_device GPU index such as 0, or cpu BBBP scorer device
--resume Path to last.pt Resume an interrupted training run
--save_trajectories Flag; no value Save timestep-level training trajectories
--wandb Flag; no value Enable Weights & Biases logging

The root entry point can also be used directly:

python train.py --config config/bdp.json

5. Evaluation

Evaluate the released Router checkpoint using sampled routing or greedy routing.

Sampled routing (the default in the released configs):

bash scripts/bdp_eval.sh
bash scripts/bdq_eval.sh
bash scripts/bpq_eval.sh

Greedy routing:

bash scripts/bdq_eval.sh --greedy

When --input is omitted, the evaluation script automatically evaluates the matching seen and unseen splits in data/TEST_multi_prop/.

The root entry point can also be used directly:

python evaluate.py \
  --config config/bdq.json \
  --checkpoint ckpts/router/bdq/best.pt

Output Directory Structure

Router training artifacts are saved under:

outputs/router/{TASK}/{ENCODER_NAME}/{TRAJECTORY_MODE}/

For the released configuration, the directory has the following structure:

outputs/router/bdp/Qwen2.5-7B-Instruct/none/
├── best.pt                    # Best Router policy
├── last.pt                    # Checkpoint for resuming training
├── config.json                # Resolved run configuration
├── train_history.csv          # Training metrics
└── eval/
    ├── sampled/               # Categorical sampling results
    └── greedy/                # Greedy routing results

Contact

For questions or issues, please contact:

About

Agentic Multi-objective Molecular Optimization via Dynamic Routing of Property-specific Editor Networks

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages