______ _____ ______ _____ ____ _____ ______ | ____| | __ \| ____/ ____/ __ \| __ \| ____| | |__ _ _ _ __ ___| | | | |__ | | | | | | | | | |__ | __| | | | '_ \ / __| | | | __|| | | | | | | | | __| | | | |_| | | | | (__| |__| | |___| |___| |__| | |__| | |____ |_| \__,_|_| |_|\___|_____/|______\_____\____/|_____/|______|
FuncDECODE is a deep-learning framework for estimating the relative contribution of each cell type to a selected functional program from mixed transcriptomic profiles. It combines reference-derived pseudo-bulk data, functional-program activity, domain adaptation, cell-type-specific distribution priors, and attention-based feature interaction in a unified workflow.
single-cell reference ── functional scoring ── pseudo-bulk learning ── domain adaptation ── contribution prediction
Overview of the FuncDECODE framework.
Given a pathway or gene program of interest, FuncDECODE predicts how its activity is distributed across the cell types represented in a mixed sample. Its target is the program-specific relative contribution vector rather than cell-type abundance.
The framework consists of three stages:
-
Functional pseudo-bulk generation
Single-cell expression profiles are scored for functional programs and repeatedly sampled to construct pseudo-bulk mixtures. Each mixture is paired with the relative contribution of every cell type to each program. -
Domain-adaptive representation learning
An encoder, predictor, and domain discriminator learn target-relevant representations while reducing the discrepancy between reference-derived and target expression profiles. -
Prior-guided contribution prediction
Cell-type-specific functional distribution summaries are projected and fused with expression-derived features. A CLS-token transformer captures interactions among cell types before predicting their relative functional contributions.
FuncDECODE/
├── data/
│ ├── data_process.py # Functional scoring and pseudo-bulk generation
│ └── PBMC/ # PBMC input and generated data
├── exp/
│ └── exp_utils.py # End-to-end training and evaluation utilities
├── model/
│ ├── FuncDECODE_stage2.py # Domain-adaptive encoder training
│ ├── FuncDECODE_stage3.py # Prior-guided FuncDECODE model
│ └── utils.py # Data loaders, prediction, and metrics
├── Tutorials/
│ └── FuncDECODE_PBMC_tutorial.ipynb # Complete PBMC workflow
├── save_models/ # Trained checkpoints
├── res/ # Predictions, histories, and metrics
├── environment.yml
└── fig.png
FuncDECODE is developed with Python 3.10 and PyTorch. CUDA is used automatically when available; CPU execution is also supported.
cd FuncDECODE
conda env create -f environment.yml
conda activate FuncDECODE
python -m ipykernel install --user --name FuncDECODE --display-name "Python (FuncDECODE)"The environment includes the packages required for the complete workflow, including PyTorch, Scanpy, GSEApy, pySCENIC, and ctxcore.
The recommended entry point is:
Tutorials/FuncDECODE_PBMC_tutorial.ipynb
Select the Python (FuncDECODE) kernel and run the notebook from top to bottom. The tutorial performs the complete experiment:
- loads and preprocesses the PBMC single-cell dataset;
- identifies marker genes and performs KEGG enrichment;
- selects the ten most significant eligible pathways;
- computes cell-level pathway activity with AUCell;
- generates 6,000 training and 1,000 test pseudo-bulk samples;
- calculates cell-type-specific pathway distribution priors;
- trains the domain-adaptation and FuncDECODE stages for all ten pathways;
- saves checkpoints, predictions, ground truth, training histories, and metrics.
Data note: FuncDECODE only requires the normalized pseudo-bulk dataset,
PBMC_norm.pkl. The tutorial intentionally does not generate the unused non-normalized or Scaden-specific datasets.
The PBMC example starts from data/PBMC/PBMC.h5ad. The AnnData object must contain:
| Location | Field | Description |
|---|---|---|
adata.X |
— | Cell-by-gene expression matrix |
adata.obs |
str_labels |
Cell-type annotation |
adata.obs |
batch |
Reference/target batch assignment |
adata.var |
gene_symbols |
Gene symbols used for enrichment and AUCell |
For another dataset, update the corresponding keys and paths in the configuration section of the tutorial.
The full PBMC workflow writes reproducible artifacts to the following locations:
| Output | Location |
|---|---|
| Normalized pseudo-bulk data | data/PBMC/PBMC_norm.pkl |
| Functional distribution priors | data/PBMC/*_dist_feats.csv |
| Trained model checkpoints | save_models/PBMC_FuncDECODE/*.pt |
| Pathway panel and configuration | res/PBMC/FuncDECODE_full_training/ |
| Predictions and ground truth | res/PBMC/FuncDECODE_full_training/*.csv |
| Training histories | res/PBMC/FuncDECODE_full_training/FuncDECODE_train_loss_*.csv |
| Overall and cell-type metrics | res/PBMC/FuncDECODE_full_training/FuncDECODE_*metrics.csv |
Each pathway is trained independently and receives its own checkpoint and result files. The consolidated table FuncDECODE_training_summary.csv records the pathway, evaluation scores, validation performance, model path, prediction path, and training parameters.
FuncDECODE reports three complementary metrics:
| Metric | Interpretation | Preferred direction |
|---|---|---|
| CCC | Agreement between predicted and true contributions | Higher is better |
| RMSE | Magnitude of the prediction error | Lower is better |
| Pearson correlation | Linear association between prediction and truth | Higher is better |
Metrics are calculated both across all predictions and separately for each pathway–cell-type pair.
Complete experimental records are archived on Zenodo: https://zenodo.org/records/21253075.
The FuncDECODE manuscript and citation information will be added upon publication.
For questions, bug reports, or feature requests, please open an issue in this repository.
