Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

GFP fluorescence prediction

Reproducible prediction of log10 fluorescence from protein sequence. We compared one-hot baselines, mean-pooled ESM-2 regressors, and three CNNs on full residue embeddings, added Jannis CNN on native one-hot inputs, and tested transfer between GFP landscapes.

Candidate model: Kermut

Kermut is being evaluated as an uncertainty-aware GP baseline. The published model requires ESM-2 650M embeddings, zero-shot scores, ProteinMPNN features and one reference structure; our existing ESM caches are 150M/640-dimensional. Moreover, its structure kernel assumes variants of one reference protein, so cross-ortholog transfer requires a separately labelled adaptation. The integration audit and experiment plan state which comparisons can be called Kermut and which must be called ESM-GP (Kermut-inspired). Live preprocessing and run status.

Experimental story

The experiments answer four questions in sequence. Throughout, the target test set is frozen, is shared by all compared methods, and is never available for training or acquisition.

Results navigation

The atlas covers all 8 transfer directions, 7 main training/acquisition scenarios, 4 models and 5 seeds. Filenames state the source, target, selection method, round and number of target sequences added.

Stage What enters training after the initial fit? Target sequences added?
Transfer without AL Nothing No
Source-pool AL transfer Sequences selected from the original source-protein training pool No
Target adaptation Sequences selected from a separate labelled pool of the target protein/domain Yes: 96 per round for 10 rounds
One-shot target-adaptation control 960 target-pool sequences added once Yes: 960 once, no AL rounds

1. Establish the base models without active learning

We first compare Aubin, Linear, MLP and CNN on OHE without iterative acquisition. These fits establish predictive accuracy on natural cgreGFP and the initial transferability between proteins and between natural cgreGFP and artificial peaks.

Transfer without AL between GFP proteins

Transfer without AL between natural cgreGFP and artificial peaks

2. Ask whether active selection improves source-only transfer

Here AL selects additional records from the source training pool. It does not add any sequence from the target protein or target domain. Each trajectory starts at 10% of the 14,709-record source budget and reaches 14,709 source records over ten rounds. The target test remains untouched.

Non-AL versus source-pool AL between GFP proteins

Non-AL versus source-pool AL between natural and artificial cgreGFP

3. Test the acquisition rule and its uncertainty term

Random and Fancy arms use identical initial source records, round budgets and frozen tests. Fancy ranks the source pool by 0.62 × scaled distance + 0.38 × scaled uncertainty. For Aubin and Linear, uncertainty is variance across a five-member bootstrap ensemble; for MLP it is variance across five independently initialized fits. All 120/120 ensemble-uncertainty trajectories are complete. This experiment still adds zero target sequences.

Ensemble uncertainty between GFP proteins

Ensemble uncertainty between natural cgreGFP and artificial peaks

4. Add target sequences: target adaptation

This is the experiment in which target data actually enter training. Starting from a source-trained model, the iterative arms add 96 selected target-pool sequences after every round for 10 rounds, for a total of 960. The matched one-shot controls add 960 target-pool sequences once and have no acquisition rounds. Random and Fancy use the same initial fit, total target-label budget, validation records and frozen target test.

All four models are shown together below. Aubin, Linear and MLP are complete. CNN protein-to-protein adaptation is also complete; every row uses all five seeds.

Target adaptation for all four models between GFP proteins

All four models are complete for natural cgreGFP ↔ artificial peaks:

Target adaptation for all four models between natural and artificial cgreGFP

The density panels below show the same cgreGFP → amacGFP iterative-Fancy trajectory after adding the first 96 target sequences and after all ten rounds (960 target sequences). Curves use the frozen amacGFP test; predicted densities pool five model seeds, while the true test density is shown once.

True and predicted densities after adding 96 target sequences

True and predicted densities after adding 960 target sequences

For the matched non-AL control, 960 target sequences are sampled randomly and added in a single fit. There is no Fancy score and there are no iterative refitting/acquisition rounds:

True and predicted densities after one-shot random addition of 960 target sequences

The same one-shot random control for transfer between natural cgreGFP and its artificial peaks:

Natural cgreGFP to artificial peaks after one-shot random addition

Artificial peaks to natural cgreGFP after one-shot random addition

The complete diagnostic atlas contains every direction and all seven main conditions, with densities and dotplots in separate folders:

Complete result index

Experiment What was done Tables Main figures
cgreGFP benchmark Fixed 60/20/20 holdout and 10-fold CV; 24,516 sequences; 121 ESM fits + 11 Jannis OHE fits + sequence baselines Report · CSV CNN predictions · Mean-ESM predictions · Accuracy/runtime
Transfer before AL Aubin, Linear, MLP and CNN on OHE; seven ortholog mixtures and four natural/artificial directions; 5 seeds Report · CSV GFP proteins · Artificial peaks
Transfer with AL Same models, seeds, budgets and frozen tests; 10 acquisition rounds; 160/160 trajectories Results · CSV · Protocol GFP proteins · Non-AL vs AL · Artificial peaks
Random vs acquisition score ("Fancy") transfer CNN on OHE; matched initial sets, budgets, seeds and frozen tests Live status and exact method · Frozen protocol GFP proteins — available after completion · Natural/artificial peaks — available after completion
Random vs acquisition score transfer — Linear 40 acquisition-score and 40 matched random CPU trajectories; 5 seeds Status · Frozen protocol GFP proteins · Natural/artificial peaks
Ensemble uncertainty for non-CNN models Aubin, Linear and MLP; 5-member ensembles; matched Random/Fancy comparison; 120/120 trajectories Results and method · Protocol GFP proteins · Natural/artificial peaks
Target-protein adaptation 96 target sequences × 10 or 960 once; fancy/random; 6 directed protein pairs; 3 CPU models; 5 seeds; 90/90 jobs Results · CSV · Design Four-arm comparison
Natural cgreGFP/artificial-peak adaptation Same 96 × 10 versus 960-once design in both directions; 3 CPU models; 5 seeds; 30/30 jobs Results · CSV · Design Four-arm comparison
CNN target-protein adaptation CNN on OHE; 96 × 10 versus 960 once; fancy/random; 6 directed protein pairs; 5 seeds; 30/30 complete Results · CSV Four-arm comparison
CNN natural/artificial adaptation CNN on OHE; same four arms in both directions; 5 seeds; 10/10 complete Results · CSV Four-arm comparison
Detailed result notes and archived intermediate figures

Selected models: cgre → cgre scatter panel shows individual saved predictions from all three cgre-only control fits. PDF · All seeds and formats · Metrics.

amacGFP + ppluGFP → cgreGFP scatter panel shows the four selected models across all three seeds on the same 4,904 held-out cgreGFP sequences (metrics).

Best measured accuracy: Jannis CNN on OHE (CV RMSE 0.227, R² 0.912, Spearman 0.900). Efficient default: Aubin 1–10–1 on OHE (RMSE 0.234, R² 0.908): its holdout fit takes 65 seconds on CPU versus 20.7 minutes on GPU for Jannis OHE. Deep MLP has nearly the same ranking score (0.899) at lower cost. Small differences are descriptive, not claims of statistical significance.

The transfer experiment compares Aubin, Linear, MLP and CNN on OHE over five seeds. All models use aligned one-hot inputs and identical source budgets.

Transferability of models between GFP proteins

Transferability between natural cgreGFP and artificial peaks

Transfer uses 14,709 train / 4,903 validation records per fit, with fixed 20% target holdouts: 4,904 natural or 5,382 artificial sequences. Sources contribute equally in ortholog mixtures. Error bars show SD across five seeds; validation uses training sources only. Artificial → artificial tests within the four represented peaks, not an unseen peak. The natural test was used in earlier comparisons, so these transfer results are exploratory.

Active-learning transfer

All 160 five-seed trajectories for Aubin, Linear, MLP and CNN on OHE are complete. Each trajectory has an initial fit plus ten acquisition rounds and ends at the same 14,709-label budget as the non-AL transfer benchmark. AL and non-AL use the exact same frozen test records. Bars show the mean and error bars show sample SD across seeds 42–46.

Matched model selection without and with AL on frozen natural cgreGFP test:

Non-AL versus AL model comparison

Transfer between GFP proteins after AL:

AL transfer between GFP proteins

Matched non-AL and AL transfer between GFP proteins:

Non-AL versus AL transfer between GFP proteins

Transfer between natural cgreGFP and artificial peaks after AL:

AL transfer between natural cgreGFP and artificial peaks

Matched non-AL and AL transfer between natural cgreGFP and artificial peaks:

Non-AL versus AL peak transfer

Target-protein adaptation

The next experiment adds 96 labelled target sequences over ten iterative rounds or adds 960 sequences once. It compares random and acquisition-score selection for all six directed natural-protein pairs and both natural/artificial directions.

Natural-protein target adaptation

Natural/artificial target adaptation

Random versus acquisition score ("Fancy") transfer

This is a separate five-seed CNN-on-OHE experiment for both GFP-protein transfer and natural cgreGFP/artificial-peak transfer. Random and acquisition-score arms share the initial labelled records, label budgets, validation sets and frozen tests. Current status and exact method.

The acquisition arm uses 0.62 × scaled hidden-space distance + 0.38 × scaled MC-dropout variance, followed by global descending-score selection. It does not run the student's dense SpectralClustering/3-mer stage, whose pool affinity matrix is quadratic at this scale. The labels and documentation state this explicitly.

The following links will resolve to the final figures immediately after all 40 random trajectories finish and the finalizer validates both arms:

Random versus acquisition score transfer — CPU models

This five-seed CPU experiment repeats the same comparison for Aubin, Linear and MLP for GFP-protein transfer and natural cgreGFP/artificial-peak transfer. The acquisition-score and matched random trajectories are complete for all three models. Every arm shares initial labelled records, label budgets, validation sets and frozen tests.

Aubin, Linear and MLP contain no dropout, so their uncertainty term is zero. Their acquisition score is therefore the scaled distance from the labelled training set in each model's projected hidden representation. As in the CNN experiment, selection uses global descending scores and does not run the dense SpectralClustering/3-mer stage.

We retain those completed distance-only trajectories as the reproducible first version. A versioned five-member ensemble experiment recalculates the three CPU models with uncertainty included. Aubin and Linear use bootstrap ensembles; MLP uses independent initializations. In each case the pool score is 0.62 × scaled distance + 0.38 × scaled ensemble prediction variance, while the primary model alone evaluates the unchanged frozen test. See the exact method and reproduction commands.

Final ensemble-uncertainty comparisons:

Aubin, Linear and MLP between GFP proteins

Aubin, Linear and MLP between natural cgreGFP and artificial peaks

The earlier Linear-only figures remain available for provenance:

Linear Random versus acquisition score between GFP proteins

Random versus acquisition-score sampling between natural cgreGFP and artificial peaks:

Linear Random versus acquisition score between natural cgreGFP and artificial peaks

Distance-only three-model figure paths:

R² is reported in the result tables alongside correlations: 1 is perfect, 0 matches the test-mean predictor, and negative values indicate worse squared error.

Quick start

Experiment 1 — base-model comparison without active learning: the supplied draft describes a random 80/20 cgre split but omits several parameters and its promised appendix. The reproduction audit and exact commands distinguish published facts, legacy-code assumptions, the existing ten-fold validation and the new five-model 80/20 comparison. Do not present the 1%-initial-label AL pilot as reproduction of the manuscript's fully supervised Pearson 0.943 result. Results · CSV · Figure.

The initial seed-42 result is retained as an important no-AL baseline; seeds 42–46 quantify its variation.

Experiment 2 — acquisition versus random sampling: matched seeds compare the manuscript acquisition score with nested random subsets at equal label budgets. Both use fixed 20% tests and report R², Pearson, Spearman and Kendall tau. See the protocol and reproduction commands. Results · CSV · Figure.

git clone https://github.com/kalininalab/gfp_project.git
cd gfp_project
conda env create -f environment.yml
conda activate gfp-benchmark
python -m unittest discover -s tests -v
python scripts/prepare_data.py
python scripts/train_aubin.py --gene cgreGFP --output results/aubin_cgreGFP_seed42

Run from the repository root. Create the environment and prepared data once; the preparer refuses to overwrite existing data. The required experimental input files are included. Training and plotting do not require the foreign repositories. The architecture comparison against Jannis's original source is an optional test and skips when that repository is absent.

To reproduce… Instructions
One-hot baselines, 10-fold splits, ESM embeddings, CNNs, tables and prediction plots Running models
Training mixtures and natural/artificial transfer figures Transfer procedure and commands
Transfer with active learning, round diagnostics and paired figures Transfer active-learning protocol
Iterative and one-shot target-protein adaptation Target-adaptation protocol
Mutation indexing, target scaling, exclusions and Aubin reference audit Data and baselines
CNN architectures and fixes to the original training code CNN differences and bug fixes

The environment is pinned in environment.yml. ESM uses the frozen facebook/esm2_t30_150M_UR50D checkpoint; full CNN workloads need an NVIDIA GPU. Use the documented HTCondor jobs for expensive runs. Condor wrappers contain our cluster paths and interpreter: adjust them for another installation.

Repository layout

  • scripts/: preparation, models, evaluation, ESM and transfer pipelines.
  • tests/: indexing, splits, label alignment, architecture and checkpoint checks.
  • condor/: job wrappers, queue lists and the transfer completion workflow.
  • data/: experimental inputs; generated data/processed/baseline_v1/ is ignored.
  • results/: selected reports, CSV tables and PNG/PDF/SVG figures are committed; generated per-fit predictions, checkpoints and logs remain local.
  • esm_embeddings/: generated feature caches, ignored because of their size.

On this installation, ignored large artifacts belong under /data/users/akolchina/gfp_project_artifacts; compatibility symlinks may keep the paths above usable. They are never repository inputs.

master_thesis_jaca00001/ and Orthologous_GFP_Fitness_Peaks/ are external reference repositories and are deliberately excluded. The reusable Jannis architecture is implemented in our model module; the original CNN script is preserved as CNN/model_legacy.py for reference, not as a training entry point.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages