LEAD is a deep-learning multi-task framework designed to accurately predict drug targets for novel compounds. The architecture features a dynamic Autoencoder (with multi-decoder latent reconstruction) and a dual-discriminator that calibrates classification confidence and manages out-of-distribution (OOD) risk in novel chemical spaces.
To clone, set up, and deploy LEAD locally or on a remote cluster, follow these steps:
git clone https://github.com/mitkeng/leadk.git
cd leadkIt is highly recommended to use an isolated conda or virtualenv space to prevent dependency collisions.
# Creating and activating conda environment
conda create -n lead_env python=3.10 -y
conda activate lead_envOr using virtualenv:
python3 -m venv lead_env
source lead_env/bin/activateEnsure dependencies are updated and installed securely:
pip install --upgrade pip
pip install -r requirements.txtThis application is split into two primary modular pipelines: Training (train.py) and Inference (predict.py).
Use this module to retrain your autoencoder networks and adversarial discriminators on custom tabular feature profiles.
--train_pos(Required): String. Path to the CSV file representing active/positive compound configurations.--train_neg(Required): String. Path to the CSV file representing negative/unlabelled background configurations.--save_dir(Optional): String. Output folder to write model checkpoints and transforms (Default:models).--epochs(Optional): Integer. Maximum number of training epochs (Default:250).--model_lr(Optional): Float. Adjust learning rate parameters for model gradients (Default:0.0003).--decoders(Optional): Integer. The number of independent reconstruction decoders for your feature space (Default:5).--adv_weight(Optional): Float. Adversarial classification loss regularization scale parameter (Default:0.03).--seed(Optional): Integer. Anchor global seed parameters for perfect result reproduction (Default:42).
python train.py \
--train_pos "Cancer_Drug_Train_Sanitized.csv" \
--train_neg "Cancer_Drug_Test_Negative.csv" \
--save_dir "models" \
--epochs 250 \
--model_lr 0.0003 \
--decoders 5Predict kinase target affinities for newly designed target molecules with active dual-discriminator logit shielding.
--input(Required): String. Path to input screening CSV containing custom descriptor values.--checkpoint(Optional): String. Target model parameters checkpoint path (Default:models/lead_checkpoint.pth).--transforms(Optional): String. Saved preprocessing transform configurations path (Default:models/transforms.pkl).--output(Optional): String. Filepath to store the prediction log results (Default:predictions_output.csv).--penalty_max(Optional): Float. Maximum temperature penalty factor allowed (Default:6.0).
python predict.py \
--input "Cancer_Drug_Test_Strictly_Novel.csv" \
--checkpoint "models/lead_checkpoint.pth" \
--output "Test_Prediction_Logs.csv" \
--penalty_max 6.0