Tier: D (capstone-scale) · Type: composite · Category: Statistical Analysis
What
The full pipeline, end to end: assemble datasets meeting the inclusion criteria → filter and CLR-transform → per-dataset covariate-adjusted effect sizes → random-effects pooling with heterogeneity → FDR correction → forest plots and a reportable results table. Parameterized by the covariate, so age, sex, BMI, or a new one are the same protocol with different inputs.
Why it matters
This is the "apply to new data" deliverable for the cMD side. Running it against a newer cMD release with the covariate set to age should reproduce the paper's age results; running it with a new covariate is a new analysis with no new methods work. It is also the proof that the Tier D atomic protocols compose.
Source material
waldronlab/curatedMetagenomicDataAnalyses — vignettes/Age_metaanalysis_vignette.Rmd and vignettes/Sex_metaanalysis_vignette.Rmd as the two reference implementations; python_tools/metaanalyze.py, one_feature_forest_plot.py, plot_meta_analysis.py, draw_figure_with_ma.py; cMD3_paper_analyses/all_command_lines_py.sh SECTION 1
- Paper: 10.1038/s41467-025-66888-1
Scope
In: the ordering of component protocols and the hand-off between them; the binary-vs-continuous branch (SMD or Fisher-Z partial correlation); which feature types are run (species, genus, pathways, KOs) and whether multiple testing is applied within or across them; the required contents of the results report.
Out: every step a component protocol already specifies.
Frontmatter starting point
type: "composite"
category: "Statistical Analysis"
citation: "10.1038/s41467-025-66888-1"
protocols_used:
- name: "cmd-cohort-assembly"
- name: "clr-transformation"
- name: "per-dataset-smd-covariate-adjusted"
- name: "partial-correlation-fisher-z"
- name: "random-effects-meta-analysis-pm"
Acceptance criteria
Cite the method's origin, not its users
PROTOCOL_STANDARD.md is explicit: an atomic protocol carries "strictly 1 citation... corresponding
to the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
- Candidate DOIs in this issue are leads, not answers. Anything marked VERIFY has not been checked.
- Some methods predate modern citation practice or have no single identifiable origin. If that is
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue in waldronlab/agent-protocol-standard — the standard may need a way to express
"classical method, no primary source".
- If you cannot name one paper that proposed everything the protocol does, it is more than one
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read CONTRIBUTING.md and
PROTOCOL_STANDARD.md.
The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-variance
protocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR:
git clone https://github.com/waldronlab/agent-protocol-standard.git
Rscript agent-protocol-standard/scripts/validate-protocol.R protocols
Tier: D (capstone-scale) · Type:
composite· Category: Statistical AnalysisWhat
The full pipeline, end to end: assemble datasets meeting the inclusion criteria → filter and CLR-transform → per-dataset covariate-adjusted effect sizes → random-effects pooling with heterogeneity → FDR correction → forest plots and a reportable results table. Parameterized by the covariate, so age, sex, BMI, or a new one are the same protocol with different inputs.
Why it matters
This is the "apply to new data" deliverable for the cMD side. Running it against a newer cMD release with the covariate set to
ageshould reproduce the paper's age results; running it with a new covariate is a new analysis with no new methods work. It is also the proof that the Tier D atomic protocols compose.Source material
waldronlab/curatedMetagenomicDataAnalyses—vignettes/Age_metaanalysis_vignette.Rmdandvignettes/Sex_metaanalysis_vignette.Rmdas the two reference implementations;python_tools/metaanalyze.py,one_feature_forest_plot.py,plot_meta_analysis.py,draw_figure_with_ma.py;cMD3_paper_analyses/all_command_lines_py.shSECTION 1Scope
In: the ordering of component protocols and the hand-off between them; the binary-vs-continuous branch (SMD or Fisher-Z partial correlation); which feature types are run (species, genus, pathways, KOs) and whether multiple testing is applied within or across them; the required contents of the results report.
Out: every step a component protocol already specifies.
Frontmatter starting point
Acceptance criteria
Cite the method's origin, not its users
PROTOCOL_STANDARD.mdis explicit: an atomic protocol carries "strictly 1 citation... correspondingto the primary literature where the method was originally published." Find the paper that proposed
the method. Do not cite a paper that merely applied it — including the BugSigDB and curatedMetagenomicData
papers, which are the source of the analysis these protocols were extracted from but almost never the
source of the method.
Tracing a method back to its first publication is real work, and it is part of the task, not a
formality. Three things to expect:
genuinely the case, say so in the pull request rather than reaching for a convenient recent paper.
Raise it as an issue in
waldronlab/agent-protocol-standard— the standard may need a way to express"classical method, no primary source".
protocol. That test has now split four protocols out of this batch: enrichment into three methods,
filtering from transformation, LODO from random forest, and PERMANOVA from ANOSIM.
Where the lab's own paper genuinely did propose the method — the oral-to-gut score, and LODO
cross-validation in Pasolli et al. 2016 — citing it is correct. That is the exception, not the pattern.
Before you start
Read
CONTRIBUTING.mdandPROTOCOL_STANDARD.md.The format is defined in the standard repo, not this one. Protocols are prose, not code: they say what to do
and why, precisely enough that two people — or two agents, in two languages — get the same answer. The existing
independent-filtering-varianceprotocol is the model to imitate for tone and level of detail.
Validate locally before opening the PR: