Reproducible prediction of log10 fluorescence from protein sequence. We compared one-hot baselines, mean-pooled ESM-2 regressors, and three CNNs on full residue embeddings, added Jannis CNN on native one-hot inputs, and tested transfer between GFP landscapes.
Kermut is being evaluated as an uncertainty-aware GP baseline. The published model requires ESM-2 650M embeddings, zero-shot scores, ProteinMPNN features and one reference structure; our existing ESM caches are 150M/640-dimensional. Moreover, its structure kernel assumes variants of one reference protein, so cross-ortholog transfer requires a separately labelled adaptation. The integration audit and experiment plan state which comparisons can be called Kermut and which must be called ESM-GP (Kermut-inspired). Live preprocessing and run status.
The experiments answer four questions in sequence. Throughout, the target test set is frozen, is shared by all compared methods, and is never available for training or acquisition.
- Complete target-adaptation report: R² and other metrics, bar plots and links to prediction diagnostics.
- Scenario-by-scenario diagnostic atlas
- All 56 true/predicted distribution panels
- All 56 true-vs-predicted dotplot panels
The atlas covers all 8 transfer directions, 7 main training/acquisition scenarios, 4 models and 5 seeds. Filenames state the source, target, selection method, round and number of target sequences added.
| Stage | What enters training after the initial fit? | Target sequences added? |
|---|---|---|
| Transfer without AL | Nothing | No |
| Source-pool AL transfer | Sequences selected from the original source-protein training pool | No |
| Target adaptation | Sequences selected from a separate labelled pool of the target protein/domain | Yes: 96 per round for 10 rounds |
| One-shot target-adaptation control | 960 target-pool sequences added once | Yes: 960 once, no AL rounds |
We first compare Aubin, Linear, MLP and CNN on OHE without iterative acquisition. These fits establish predictive accuracy on natural cgreGFP and the initial transferability between proteins and between natural cgreGFP and artificial peaks.
Here AL selects additional records from the source training pool. It does not add any sequence from the target protein or target domain. Each trajectory starts at 10% of the 14,709-record source budget and reaches 14,709 source records over ten rounds. The target test remains untouched.
Random and Fancy arms use identical initial source records, round budgets and
frozen tests. Fancy ranks the source pool by
0.62 × scaled distance + 0.38 × scaled uncertainty. For Aubin and Linear,
uncertainty is variance across a five-member bootstrap ensemble; for MLP it is
variance across five independently initialized fits. All 120/120
ensemble-uncertainty trajectories are complete. This experiment still adds
zero target sequences.
This is the experiment in which target data actually enter training. Starting from a source-trained model, the iterative arms add 96 selected target-pool sequences after every round for 10 rounds, for a total of 960. The matched one-shot controls add 960 target-pool sequences once and have no acquisition rounds. Random and Fancy use the same initial fit, total target-label budget, validation records and frozen target test.
All four models are shown together below. Aubin, Linear and MLP are complete. CNN protein-to-protein adaptation is also complete; every row uses all five seeds.
All four models are complete for natural cgreGFP ↔ artificial peaks:
The density panels below show the same cgreGFP → amacGFP iterative-Fancy
trajectory after adding the first 96 target sequences and after all ten rounds
(960 target sequences). Curves use the frozen amacGFP test; predicted densities
pool five model seeds, while the true test density is shown once.
For the matched non-AL control, 960 target sequences are sampled randomly and added in a single fit. There is no Fancy score and there are no iterative refitting/acquisition rounds:
The same one-shot random control for transfer between natural cgreGFP and its artificial peaks:
The complete diagnostic atlas contains every direction and all seven main conditions, with densities and dotplots in separate folders:
- Browse all 56 distribution panels
- Browse all 56 true-vs-predicted dotplot panels
- Scenario-by-scenario index
| Experiment | What was done | Tables | Main figures |
|---|---|---|---|
| cgreGFP benchmark | Fixed 60/20/20 holdout and 10-fold CV; 24,516 sequences; 121 ESM fits + 11 Jannis OHE fits + sequence baselines | Report · CSV | CNN predictions · Mean-ESM predictions · Accuracy/runtime |
| Transfer before AL | Aubin, Linear, MLP and CNN on OHE; seven ortholog mixtures and four natural/artificial directions; 5 seeds | Report · CSV | GFP proteins · Artificial peaks |
| Transfer with AL | Same models, seeds, budgets and frozen tests; 10 acquisition rounds; 160/160 trajectories | Results · CSV · Protocol | GFP proteins · Non-AL vs AL · Artificial peaks |
| Random vs acquisition score ("Fancy") transfer | CNN on OHE; matched initial sets, budgets, seeds and frozen tests | Live status and exact method · Frozen protocol | GFP proteins — available after completion · Natural/artificial peaks — available after completion |
| Random vs acquisition score transfer — Linear | 40 acquisition-score and 40 matched random CPU trajectories; 5 seeds | Status · Frozen protocol | GFP proteins · Natural/artificial peaks |
| Ensemble uncertainty for non-CNN models | Aubin, Linear and MLP; 5-member ensembles; matched Random/Fancy comparison; 120/120 trajectories | Results and method · Protocol | GFP proteins · Natural/artificial peaks |
| Target-protein adaptation | 96 target sequences × 10 or 960 once; fancy/random; 6 directed protein pairs; 3 CPU models; 5 seeds; 90/90 jobs | Results · CSV · Design | Four-arm comparison |
| Natural cgreGFP/artificial-peak adaptation | Same 96 × 10 versus 960-once design in both directions; 3 CPU models; 5 seeds; 30/30 jobs | Results · CSV · Design | Four-arm comparison |
| CNN target-protein adaptation | CNN on OHE; 96 × 10 versus 960 once; fancy/random; 6 directed protein pairs; 5 seeds; 30/30 complete | Results · CSV | Four-arm comparison |
| CNN natural/artificial adaptation | CNN on OHE; same four arms in both directions; 5 seeds; 10/10 complete | Results · CSV | Four-arm comparison |
Detailed result notes and archived intermediate figures
Selected models: cgre → cgre scatter panel shows individual saved predictions from all three cgre-only control fits. PDF · All seeds and formats · Metrics.
amacGFP + ppluGFP → cgreGFP scatter panel shows the four selected models across all three seeds on the same 4,904 held-out cgreGFP sequences (metrics).
Best measured accuracy: Jannis CNN on OHE (CV RMSE 0.227, R² 0.912, Spearman 0.900). Efficient default: Aubin 1–10–1 on OHE (RMSE 0.234, R² 0.908): its holdout fit takes 65 seconds on CPU versus 20.7 minutes on GPU for Jannis OHE. Deep MLP has nearly the same ranking score (0.899) at lower cost. Small differences are descriptive, not claims of statistical significance.
The transfer experiment compares Aubin, Linear, MLP and CNN on OHE over five seeds. All models use aligned one-hot inputs and identical source budgets.
Transfer uses 14,709 train / 4,903 validation records per fit, with fixed 20% target holdouts: 4,904 natural or 5,382 artificial sequences. Sources contribute equally in ortholog mixtures. Error bars show SD across five seeds; validation uses training sources only. Artificial → artificial tests within the four represented peaks, not an unseen peak. The natural test was used in earlier comparisons, so these transfer results are exploratory.
All 160 five-seed trajectories for Aubin, Linear, MLP and CNN on OHE are complete. Each trajectory has an initial fit plus ten acquisition rounds and ends at the same 14,709-label budget as the non-AL transfer benchmark. AL and non-AL use the exact same frozen test records. Bars show the mean and error bars show sample SD across seeds 42–46.
Matched model selection without and with AL on frozen natural cgreGFP test:
Transfer between GFP proteins after AL:
Matched non-AL and AL transfer between GFP proteins:
Transfer between natural cgreGFP and artificial peaks after AL:
Matched non-AL and AL transfer between natural cgreGFP and artificial peaks:
The next experiment adds 96 labelled target sequences over ten iterative rounds or adds 960 sequences once. It compares random and acquisition-score selection for all six directed natural-protein pairs and both natural/artificial directions.
This is a separate five-seed CNN-on-OHE experiment for both GFP-protein transfer and natural cgreGFP/artificial-peak transfer. Random and acquisition-score arms share the initial labelled records, label budgets, validation sets and frozen tests. Current status and exact method.
The acquisition arm uses 0.62 × scaled hidden-space distance + 0.38 × scaled MC-dropout variance, followed by global descending-score selection. It does not
run the student's dense SpectralClustering/3-mer stage, whose pool affinity
matrix is quadratic at this scale. The labels and documentation state this
explicitly.
The following links will resolve to the final figures immediately after all 40 random trajectories finish and the finalizer validates both arms:
- Random vs acquisition score between GFP proteins
- Random vs acquisition score between natural cgreGFP and artificial peaks
This five-seed CPU experiment repeats the same comparison for Aubin, Linear and MLP for GFP-protein transfer and natural cgreGFP/artificial-peak transfer. The acquisition-score and matched random trajectories are complete for all three models. Every arm shares initial labelled records, label budgets, validation sets and frozen tests.
Aubin, Linear and MLP contain no dropout, so their uncertainty term is zero. Their acquisition score is therefore the scaled distance from the labelled training set in each model's projected hidden representation. As in the CNN experiment, selection uses global descending scores and does not run the dense SpectralClustering/3-mer stage.
We retain those completed distance-only trajectories as the reproducible first
version. A versioned five-member ensemble experiment recalculates the
three CPU models with uncertainty included. Aubin and Linear use bootstrap
ensembles; MLP uses independent initializations. In each case the pool score is
0.62 × scaled distance + 0.38 × scaled ensemble prediction variance, while
the primary model alone evaluates the unchanged frozen test. See the
exact method and reproduction commands.
Final ensemble-uncertainty comparisons:
The earlier Linear-only figures remain available for provenance:
Random versus acquisition-score sampling between natural cgreGFP and artificial peaks:
Distance-only three-model figure paths:
- Aubin, Linear and MLP between GFP proteins
- Aubin, Linear and MLP between natural cgreGFP and artificial peaks
R² is reported in the result tables alongside correlations: 1 is perfect, 0 matches the test-mean predictor, and negative values indicate worse squared error.
Experiment 1 — base-model comparison without active learning: the supplied draft describes a random 80/20 cgre split but omits several parameters and its promised appendix. The reproduction audit and exact commands distinguish published facts, legacy-code assumptions, the existing ten-fold validation and the new five-model 80/20 comparison. Do not present the 1%-initial-label AL pilot as reproduction of the manuscript's fully supervised Pearson 0.943 result. Results · CSV · Figure.
The initial seed-42 result is retained as an important no-AL baseline; seeds 42–46 quantify its variation.
Experiment 2 — acquisition versus random sampling: matched seeds compare the manuscript acquisition score with nested random subsets at equal label budgets. Both use fixed 20% tests and report R², Pearson, Spearman and Kendall tau. See the protocol and reproduction commands. Results · CSV · Figure.
git clone https://github.com/kalininalab/gfp_project.git
cd gfp_project
conda env create -f environment.yml
conda activate gfp-benchmark
python -m unittest discover -s tests -v
python scripts/prepare_data.py
python scripts/train_aubin.py --gene cgreGFP --output results/aubin_cgreGFP_seed42Run from the repository root. Create the environment and prepared data once; the preparer refuses to overwrite existing data. The required experimental input files are included. Training and plotting do not require the foreign repositories. The architecture comparison against Jannis's original source is an optional test and skips when that repository is absent.
| To reproduce… | Instructions |
|---|---|
| One-hot baselines, 10-fold splits, ESM embeddings, CNNs, tables and prediction plots | Running models |
| Training mixtures and natural/artificial transfer figures | Transfer procedure and commands |
| Transfer with active learning, round diagnostics and paired figures | Transfer active-learning protocol |
| Iterative and one-shot target-protein adaptation | Target-adaptation protocol |
| Mutation indexing, target scaling, exclusions and Aubin reference audit | Data and baselines |
| CNN architectures and fixes to the original training code | CNN differences and bug fixes |
The environment is pinned in environment.yml. ESM uses the
frozen facebook/esm2_t30_150M_UR50D checkpoint; full CNN workloads need an NVIDIA
GPU. Use the documented HTCondor jobs for expensive runs. Condor wrappers contain
our cluster paths and interpreter: adjust them for another installation.
scripts/: preparation, models, evaluation, ESM and transfer pipelines.tests/: indexing, splits, label alignment, architecture and checkpoint checks.condor/: job wrappers, queue lists and the transfer completion workflow.data/: experimental inputs; generateddata/processed/baseline_v1/is ignored.results/: selected reports, CSV tables and PNG/PDF/SVG figures are committed; generated per-fit predictions, checkpoints and logs remain local.esm_embeddings/: generated feature caches, ignored because of their size.
On this installation, ignored large artifacts belong under
/data/users/akolchina/gfp_project_artifacts; compatibility symlinks may keep
the paths above usable. They are never repository inputs.
master_thesis_jaca00001/ and Orthologous_GFP_Fitness_Peaks/ are external
reference repositories and are deliberately excluded. The reusable Jannis
architecture is implemented in our model module; the original CNN script is
preserved as CNN/model_legacy.py for reference, not as a training entry point.



















