MORSE: Agentic Multi-objective Molecular Optimization via Dynamic Routing of Property-specific Editor Networks
AI4Sci Korea 2026 — Extended Abstract
September 28–October 1, 2026 · Seoul Dragon City, Seoul, Republic of Korea
Property-specific Editor checkpoints are distributed through Git LFS. Install Git LFS before cloning the repository.
On Ubuntu or Debian:
sudo apt update
sudo apt install git-lfsAlternatively, install Git LFS from conda-forge in an active Conda environment:
conda install -c conda-forge git-lfsInitialize Git LFS and clone MORSE:
git lfs install
git clone https://github.com/jjjabcd/MORSE.git
cd MORSECreate a Python 3.11 Conda environment and install MORSE with all property scorers:
conda create -n morse python=3.11 -y
conda activate morse
pip install -e .Download the Editor checkpoints from the project root:
git lfs pull --include="ckpts/editor/**/*.pt"The released Router checkpoints are regular Git files and do not require an additional download.
Download the DRD2 classifier:
bash scripts/fetch_drd2_model.shDownload the LLM state encoder from Qwen/Qwen2.5-7B-Instruct:
python scripts/fetch_encoder.py \
--model Qwen/Qwen2.5-7B-Instruct \
--cache-dir ckpts/llmThe BBBP scorer uses ADMET-AI and may download its model automatically on first use.
Checkpoint structure
ckpts/
├── editor/
│ ├── foundation/
│ │ ├── foundation_final.pt
│ │ ├── configs.csv
│ │ └── char2idx.csv
│ └── adapters/
│ ├── qed_lora.pt
│ ├── drd2_lora.pt
│ ├── plogp_lora.pt
│ └── bbbp_lora.pt
├── router/
│ ├── bdp/best.pt
│ ├── bdq/best.pt
│ └── bpq/best.pt
└── llm/ # Local Hugging Face cache
MORSE trains one Router for each objective combination. Run one of the following scripts:
# BBBP + DRD2 + penalized logP
bash scripts/bdp_train.sh
# BBBP + DRD2 + QED
bash scripts/bdq_train.sh
# BBBP + penalized logP + QED
bash scripts/bpq_train.shTraining arguments
Arguments supplied after a training script override the corresponding values in the task config.
| Argument | Value | Description |
|---|---|---|
--data_dir |
Directory path | Router training dataset directory |
--checkpoints |
Directory path | Editor checkpoint root |
--encoder_model |
Hugging Face model ID or local path | Frozen LLM state encoder |
--llm_cache_dir |
Directory path | LLM download/cache directory |
--total_episodes |
Positive integer | Number of training episodes |
--batch_episodes |
Positive integer | Episodes collected per training step |
--max_length |
Positive integer | Maximum encoder token length |
--max_timesteps |
Positive integer | Maximum Router steps per episode |
--K |
Positive integer | Number of Editor candidates when Editor beam search is disabled |
--editor_num_beams |
Positive integer | Editor decoding beam width; 1 uses sampling |
--min_similarity |
Float in [0, 1] |
Similarity threshold used by the reward |
--gamma |
Float in [0, 1] |
Reward discount factor |
--lr |
Positive float | Router learning rate |
--grad_clip |
Positive float | Gradient clipping norm |
--value_coef |
Non-negative float | Value loss coefficient |
--entropy_coef |
Non-negative float | Entropy bonus coefficient |
--dropout |
Float in [0, 1) |
Router dropout probability |
--policy_arch |
no_numeric or original |
Router policy input architecture |
--ppo_epochs |
Positive integer | PPO update epochs per collected batch |
--clip_eps |
Positive float | PPO clipping epsilon |
--ppo_buffer_size |
Positive integer | PPO replay-buffer capacity in training steps |
--seed |
Integer | Random seed |
--router_device |
cuda:N, cpu, or auto |
LLM encoder and Router device |
--qed_editor |
cuda:N or cpu |
QED Editor device |
--drd2_editor |
cuda:N or cpu |
DRD2 Editor device |
--plogp_editor |
cuda:N or cpu |
Penalized-logP Editor device |
--bbbp_editor |
cuda:N or cpu |
BBBP Editor device |
--bbbp_device |
GPU index such as 0, or cpu |
BBBP scorer device |
--resume |
Path to last.pt |
Resume an interrupted training run |
--save_trajectories |
Flag; no value | Save timestep-level training trajectories |
--wandb |
Flag; no value | Enable Weights & Biases logging |
The root entry point can also be used directly:
python train.py --config config/bdp.jsonEvaluate the released Router checkpoint using sampled routing or greedy routing.
Sampled routing (the default in the released configs):
bash scripts/bdp_eval.sh
bash scripts/bdq_eval.sh
bash scripts/bpq_eval.shGreedy routing:
bash scripts/bdq_eval.sh --greedyWhen --input is omitted, the evaluation script automatically evaluates the matching seen and unseen splits in data/TEST_multi_prop/.
The root entry point can also be used directly:
python evaluate.py \
--config config/bdq.json \
--checkpoint ckpts/router/bdq/best.ptRouter training artifacts are saved under:
outputs/router/{TASK}/{ENCODER_NAME}/{TRAJECTORY_MODE}/
For the released configuration, the directory has the following structure:
outputs/router/bdp/Qwen2.5-7B-Instruct/none/
├── best.pt # Best Router policy
├── last.pt # Checkpoint for resuming training
├── config.json # Resolved run configuration
├── train_history.csv # Training metrics
└── eval/
├── sampled/ # Categorical sampling results
└── greedy/ # Greedy routing results
For questions or issues, please contact:
- Email: rlawlsgurjh@gmail.com
