Modular metadata schema components for documenting geochemical analytical Methods and Datasets. Built using the OGC Building Blocks pattern.
The scheme involves three components:
-
A Technique-Aligned protocol (TAPP) that defines a analytical procedure, including kinds of samples used, target analytes, instruments used, sample preparation, analysis workflow and data reduction. In the TAPP definition, some of these might be specified as fixed, some might have default values, and some are expected to be specified a the individual session level. The fixed properties are the necessary properties that define the TAPP. There are also properties that apply as the analytical session (or 'analysis event') level, and properties that are specific to the description of individual analytes. The authoritative protocol definition is in an Excel workbook. For discussion purposes, the label 'property' is used for properties in the TAPP that are fixed, and 'parameter' for properties that may be adjusted at the session level. Parameters may have default values specified in the TAPP definition.
-
A building block JSON schema specific to the protocol. This protocol definition object is registered in a protocol registry and accessible via its URI. The TAPP definition is referenced as a measurementTechnique in dataset metadata.
-
A technique-specific 'detail' building block JSON schema that defines the parameters that may be assigned values at the individual dataset level. There is one detail block per technique, at
_sources/techniqueProfile/<GROUP>/<TECH>/detail/(GROUP=geochemProfileoradaProfile), not a single 'details' file. Two kinds exist and they compose at different nodes: a dataset-root block overlaysschema:Dataset, while a hasPart-item block pinsada:componentTypeand overlays a distribution part - seeagents.mdfor the split and the grep that re-derives it. The content of this schema is included in the schema for dataset instances to create a metadata schema for Datasets conforming to the profile. Session-level and per-analyte parameters are defined once in a registered parameter registry (parameterValues) and referenced from the detail blocks by URI, so a parameter can be reused across detail definitions; the references are resolved inline into the published resolved schema.
_sources/ has three top-level areas: the shared base schemas, the shared registries, and one directory per analytical technique.
Most of
_sources/is generated, despite the name. Of its 242schema.yaml, about 169 are written by a generator — thetapp/,detail/andprofile/schemas from the TAPP tables and sidecars, the composition modules, the shared registries — along with all 242resolvedSchema.json, all 242<name>Schema.json, and the generated examples. What you edit directly isBaseSchema/*and theadaProfile/profiles. Editing a generated file instead of its generator is the most common way to lose work here;CLAUDE.mdhas the per-directory table of what writes what.
_sources/
BaseSchema/ 18 shared BBs: geochemProduct (domain-neutral base),
adaProduct (extends it), tappDefinition, instrument,
laboratory, the file-type blocks (image, imageMap,
tabularData, dataCube, collection, document,
supDocImage, otherFile, files), structuredData,
spatialRegistration, creativeWork, stringArray
registry/ five catalogs keyed by what the column identifies
targetSpeciesColumns/ registered BB: PropertyValueSpecification $defs,
one per per-species column (258)
monitoredPropertyColumns/ registered BB: $defs for the monitored-property
table (mass, cup, edge, X-ray line) (63)
reportedPropertyColumns/ registered BB: $defs for reported-quantity columns (1)
parameterTemplates/ registered BB: PropertyValueSpecification $defs,
editable params (414)
parameterValues/ registered BB: schema:PropertyValue $defs,
fixed values (1185)
vocab/ catalog: schema:DefinedTermSet files,
by @id not $ref (451)
techniqueProfile/ one directory per technique (91), under two roots:
geochemProfile/ the 59 TAPP-aware techniques
<TECH>/tapp/ the TAPP definition for that technique (59 techniques)
<TECH>/detail/ per-dataset analysis-instance detail (59 techniques)
<TECH>/profile/ path-driven product profile: geochemProduct +
detail + TAPP linkage (29 techniques)
<TECH>/profile-ada/ generic product profile, written by the
TAPP tooling (9 techniques)
adaProfile/ the other 32 techniques, untouched by the TAPP work
<TECH>/profile-ada/ generic product profile: adaProduct +
componentType constraints only (31 techniques)
<TECH>/detail/ instrument-detail stub (14 techniques)
Profile directory names are not profile names. EPMA/profile-ada publishes adaEPMA; SEM/profile publishes adaSEMFull. A profile's canonical name is the schema:subjectOf.dcterms:conformsTo const inside its own schema — read it from there rather than inferring from the path.
Shared building blocks. The product profile is split into two layers:
geochemProductis the domain-neutral base product profile, composing the CDIF v1.1 profile schemas viaallOf. It carries no ADA-specific requirements: its distribution has an optionalschema:additionalType(drawn from the componentType vocabulary), not a requiredada:componentType.adaProductextendsgeochemProduct(viaallOf: [$ref geochemProduct, …]) with the ADA/SAMIS overlays: technique types, instrument/lab/sample, and a requiredada:componentTypeon each distribution. Everything ADA-specific lives here, sogeochemProductstays reusable outside ADA.
The composed CDIF v1.1 profiles (via geochemProduct):
cdifCore— core metadata propertiescdifDataDescription— variableMeasured with DDI-CDI extensions,@idrequirementcdifProvenance—prov:wasGeneratedByprovenance activitiescdifManifest— archive distribution withhasPartcomponent files (wascdifArchiveDistributionin CDIF ≤1.0). Applied conditionally: theif/thenfires only when aschema:distributionitem carriesschema:Collectionin its@type, so a monolithic single-file distribution isn't held to the manifest rules.
Two BBs extend CDIF core BBs:
- instrument — extends core CDIF instrument; requires
schema:additionalType(at least one entry, e.g.nxs:BaseClass/NXinstrumentor a technique term likeada:EPMAInstrument) - laboratory — extends core CDIF spatialExtent (
schema:Placewithnxs:BaseClass/NXsourceinadditionalType)
tappDefinition is documented in its own section below.
targetSpeciesColumns, monitoredPropertyColumns, reportedPropertyColumns, parameterTemplates, and parameterValues are each a registered type-library building block (bblock.json with isTypeLibrary: true): every entry lives as a named $def in the catalog's schema.yaml, and TAPP / detail blocks reference them by URI fragment ($ref: …/<catalog>/schema.yaml#/$defs/<name>). Because they are registered, the OGC bblocks annotate step resolves those refs locally via the register and inlines them into resolvedSchema.json. This matters: a loose helper file (a plain <name>.json not inside a registered BB) is instead fetched from the published gh-pages URL, which 404s on moved or unpublished paths (process-bblocks.yml sets skip-pages: true, so gh-pages never auto-updates) — that fragility is why the catalogs were promoted to registered BBs. vocab/ is the exception: it stays a plain catalog of schema:DefinedTermSet files because it is referenced only by JSON-LD @id (schema:inDefinedTermSet), never by $ref, so the annotate step never fetches it.
The catalogs are shared dictionary resources — multiple TAPPs $ref the same $defs when their definitions match. share_or_write_catalog lets a TAPP regen overwrite its own entries (matched by $id ownership) but errors out on a collision with an entry originated by a different TAPP, so a new TAPP either reuses identical catalog entries or surfaces a renaming requirement.
parameterTemplates holds editable parameters (a PropertyValueSpecification with a default the analyst may override); parameterValues holds fixed protocol values (a schema:PropertyValue). That split — specification vs value — is how read-only-ness is expressed; ada:methodParameters was retired repo-wide in favour of schema:additionalProperty.
All 59 techniques under geochemProfile/ have a tapp/ and a detail/. 29 of them also publish a path-driven profile/ — the set registered in build_profile.PROFILES.
tapp/— the protocol definition. ExtendstappDefinitionviaallOfwith technique-specific top-levelada:properties,schema:additionalProperty[]entries, andada:targetSpeciesTemplate.ada:targetSpeciesColumnsconstraints referencing the registry catalogs.detail/— the per-dataset analysis instance. Placement is not uniform, and does not track whether the technique is path-driven. Seven overlay theschema:Datasetroot (analyst contributor, session dates, sample, funding, per-analysis parameter values): Basemap, EPMA, Geochron, SEM, SEM-Composition, Solution-Q-ICPMS, Solution-SF-ICPMS. The other eighteen pinada:componentTypeand overlay aschema:distribution.hasPartitem: ARGT, DSC, EAIRMS, ICPOES, L2MS, LA-ICPMS, LAF, NanoIR, NanoSIMS, PSFD, QRIS, SEM-FIBSEM, SEM-Imaging, SLS, TEM, VNMIR, XCT, XRD. Consumers cannot assume one placement.profile/— path-driven product profile: bases on the domain-neutralgeochemProduct+ thedetailblock +prov:usednarrowed to that technique's TAPP + the technique'sada:componentTypeenum onhasPart(the profile layers the ADA componentType constraint on top of the ADA-agnostic base).profile-ada/— the generic product profile: bases onadaProduct+ada:componentTypeconstraints only, no TAPP linkage or detail block.
Base-selection rule.
geochemProfile/<TECH>/profile/(generic, path-driven) →geochemProduct;geochemProfile/<TECH>/profile-ada/(the 4 ADA variants) and alladaProfile/<TECH>/profile-ada/→adaProduct. Principle: geochem→geochemProduct, ada→adaProduct.
A dataset instance selects between the two profile variants by how it references its protocol: a bare {"@id": …} node reference in schema:measurementTechnique targets the path-driven profile, an inline schema:DefinedTerm targets the generic one.
A module defines a shared field once instead of once per technique.
_sources/BaseSchema/modules/* is generated from the sidecars in
docs/modules/Module_*.schemapaths.csv — never hand-edit a module.
Across the 16 technique tapp/ schemas, 651 parameter slots are
composed from a module against 231 minted per technique (73%), ranging
from 100% for Solution-MC-ICPMS down to 7% for Lab-XCT. The spread tracks
which modules exist: the 2026-09 delivery added ICPMS, CollisionCell
and CompositionQC, which took the ICP-MS families from ~25% to 78–100%,
while electron-beam and tomography still have no family module.
That percentage counts parameters — entries under
schema:additionalProperty — not all properties, and it ignores the
module ROOT $defs a technique composes, so it understates what a module
supplies. 100% is not reachable: of the 79 distinct fields no module
covers, 50 are carried by exactly one table and a module needs two
consumers. See MODULE_CONSOLIDATION_STATUS.md for the full definition and
the ~85% ceiling.
- docs/modules/MODULE_CONSOLIDATION_STATUS.md — current state: what was measured and how to reproduce it, what is decided, what is open, and the next drafting pass. Start here.
agents.md§"Module composition" — how composition actually works, including why parameters compose differently from structural fields and why some duplication is correct.docs/modules/draft/— 8 provisionalDraft_Module_*.csv, all electron-beam. Ours and provisional; the library's modules are Ruolin's to author. The six ICP-MS drafts were adopted upstream in the 2026-09 delivery and deleted from here; eight of their fields were not taken up and are recorded indocs/upstream-requests.md§1.
One measurement warning: do not count duplication in
_sources/registry/. Those catalogues are per-technique by
construction, so they report ~87% duplication whatever composition does.
Measure the tapp/ schemas.
Each archive hasPart item carries an ada:componentType (a single string like ada:EPMAImageMap) that classifies the file. The term list is governed by a vocabulary, and two schema layers add per-context constraints:
Governing vocabulary. registry/vocab/componentType.json is a SKOS ConceptScheme (@id: ada:vocab/componentType, the ~22 universal cross-technique terms). The base products reference it by annotation only — the universalComponentType $def in geochemProduct (and duplicated in adaProduct) is {type: string, schema:inDefinedTermSet: "ada:vocab/componentType"} with no inline enum, so at the base layer any string validates and conformance to the vocabulary is advisory (SHACL-checkable), not hard-enforced by JSON Schema. geochemProduct exposes the vocab as an optional schema:additionalType; adaProduct requires it as ada:componentType.
-
File type ↔ componentType mapping — each file-type building block (
image,imageMap,tabularData,collection,dataCube,document,supDocImage,otherFile) declares a sealedenumof valid componentType values. The enum is derived from the Components worksheet ofamds-ldeo/metadata/ADA-AnalyticalMethodsAndAttributes.xlsx(the canonical mapping; columnscomponentType/FileType/isSupplement). E.g.ada:EPMAImageMapis valid only on parts whose@typeincludesada:imageMap. -
Profile-level constraint — a technique profile's
schema:distribution.items.schema:hasPart.itemsuses a schema-levelanyOfwith three kinds of branch: (a)$reftogeochemProduct/schema.yaml#/$defs/universalComponentTypeBranch(factored once, used everywhere) for universal componentTypes; (b) inline string-enum for technique-specific componentTypes; (c) for techniques whosedetail/block is the older hasPart-item kind (XRD, ARGT, DSC, …), a$refto that detail schema, which pinsada:componentTypeto its technique consts and contributes detail-specific sibling properties (e.g.ada:geometry) flat on the hasPart item — not nested inside componentType. Path-driven profiles do not use branch (c): their detail block overlays the dataset root instead, andhasPartgets only branches (a) and (b).
Keeping the layers in sync. Because the base layer is annotation-only, JSON-Schema validation no longer catches componentType drift on its own. python tools/check_componentType.py restores that check: it fails if a universal vocab term is missing from the enum cache, if a base schema stops annotating the vocab @id, or if any ada:componentType used in an example is not a known term (universal vocab ∪ enum cache ∪ per-technique profile enums). Run it after touching the worksheet, the vocab, or example componentTypes.
After editing the Components worksheet:
python tools/apply_componentType_enums.py --refresh \
--xlsx ../../amds-ldeo/metadata/ADA-AnalyticalMethodsAndAttributes.xlsx
python tools/regenerate_schema_json.py
python tools/resolve_schema.py --all
python tools/validate_examples.py
python tools/check_componentType.py # confirm vocab / cache / schemas / examples agree
The cached mapping at tools/componentType_enum_cache.json is committed so the apply step works on a fresh clone without spreadsheet access. Worksheet FileType values map to file-type BBs as: image→image (or supDocImage when isSupplement=supplement), imageMap→imageMap, tabularData→tabularData, archive→collection, dataCube→dataCube, document→document, video/otherFile→otherFile, plus the document | image and document | tabularData splits.
This repository imports shared schema.org and CDIF property building blocks from metadataBuildingBlocks via the OGC Building Blocks import mechanism. All external references use absolute URLs (https://cross-domain-interoperability-framework.github.io/metadataBuildingBlocks/_sources/...).
Four workflows live in .github/workflows/. Three report a check on every pull request and are
required in branch protection on main:
| check | workflow | what it does |
|---|---|---|
Validate and annotate (no pages) |
validate-branch.yml |
the full OGC postprocess (validate + annotate + build register/tests), no Pages deploy |
Regenerate and diff |
check-schema-drift.yml |
runs tools/regenerate.py and fails if any committed artifact moves |
Regenerate twice and compare |
check-determinism.yml |
regenerates under two different PYTHONHASHSEEDs and compares, so a generator cannot vary by run |
process-bblocks.yml is the fourth. It runs on pushes to main and produces the published
build/ tree.
docs/SILENT_SUCCESS.md catalogues the incidents behind this setup — six cases where something in this pipeline reported success while not doing its job, what exposed each one, and the rule that follows. Read it before changing a check or a trigger.
Two consequences worth knowing before editing any of them:
- A required workflow must not carry a
paths:filter. A required check that never runs is pending forever rather than skipped, and blocks the pull request indefinitely. - Auto-merge waits only on required checks. It will merge past a failing check that is not required.
build/ is committed, because deploy-viewer.yml reads build/register.json and
build/tests/report.json out of main and publishes the tree to GitHub Pages without
regenerating them.
Branch protection declines a direct push from the postprocess, and GitHub offers no way to exempt it — the GitHub Actions app can only be a bypass actor on an organization ruleset, while a repository ruleset accepts only deploy-key and repository-role bypasses. So the generated output arrives the same way every other change does, as a pull request:
process-bblocks.ymlforce-resets thebblocks-buildbranch tomain;- the reusable OGC postprocess runs against that branch and commits its output there;
- a pull request is opened from it and auto-merge is armed, so it lands once the three required checks pass.
Nothing bypasses protection: the regenerated build/ is reviewed by the same checks as authored
source. Step 3 uses a fine-grained PAT (BBLOCKS_PR_TOKEN, Contents:read + PullRequests:write on
this repository only) for one reason — a pull request opened by GITHUB_TOKEN triggers no
workflows, so its required checks would never report and it could never merge.
bblocks-buildis machine-owned. It is force-reset on every postprocess run. Do not branch from it, commit to it, or base work on it.
Browse the building blocks at: https://amds-ldeo.github.io/geochemBuildingBlocks/
tools/_tapp_lib.write_profile_ada_companions(<block dir>) writes the four files a profile-ada
block ships beside its schema — context.jsonld, description.md, rules.shacl and
examples.yaml. They used to come from generate_profiles.py, now blocked because its SCHEMA
template emits the retired object-form ada:componentType; _tapp_lib took the schema over and
nothing took the companions, so five blocks added after the deprecation shipped without them. It
reads each block's own schema.yaml and bblock.json rather than a registry, so a block that
exists is describable whether or not anyone registered it. examples.yaml is written only when an
example*.json exists — emitting a ref: to a missing file would satisfy the audit while pointing
at nothing.
tools/build_html_views.py renders two page types into build/htmlViews/: a dataset record
page for a product profile instance, and a TAPP definition page for the protocol it names. The
dataset page follows the record's own prov:used TAPP reference through to that TAPP's page.
Keyed values render as a grid: one row per member of the keyset the defines: row declares
(each target species, each monitored property), one column per property keyed to that set.
python tools/build_html_views.py --source examples --all # the 97 schema examples
python tools/build_html_views.py --source ada2 --limit 50 # real ADA holdings, public.json_table
python tools/build_html_views.py --source ada2 --doi <doi> # one record
python tools/build_html_views.py --all --no-tapp-pages
--source ada2 reads public.json_table using ADA_NAME / DB_2024_USER / DB_2024_PASSWORD /
DB_2024_HOST / DB_2024_PORT — the same variables the metadata loaders use, not a
PGSERVICEFILE or .pgpass entry.
Output lands in build/htmlViews/, which is not committed. The rest of build/ — the annotated schemas, OAS3 downcompiles, register.json, bblocks.jsonld/.ttl and test report — IS committed, generated by the OGC postprocess in CI; see Continuous integration.
docs/TAPP-schema-generation-workflow.md is the authoritative walkthrough — written for three audiences (workbook author, pipeline maintainer, form builder) with a flowchart of the whole path from spreadsheet to validated schema. Read it first; the summary here is orientation only.
TAPP source = the
tapp/git submodule (amds-ldeo/tapp). The TAPP tables and modules live in that submodule, not in this repo;tools/tapp_source.py:current_delivery()resolves totapp/(falling back to any inlineTAPPS<date>/drop). Clone withgit clone --recursive, or rungit submodule update --initin an existing checkout, before regenerating. Pin/bump the delivery by updating the submodule commit — a deliberate act, because.gitmodulessetsupdate = noneto stop the OGC postprocess workflow advancing the pointer on its own. The committed schemas are built from the pinned revision (af3f7bc, adopted 2026-09-02).
One upstream-authored TAPP table per technique (a CSV in the tapp/ submodule's Current TAPPs/) drives everything downstream. Nothing generated should ever be hand-edited — fix the table (upstream, in amds-ldeo/tapp) or a tool and regenerate.
Regenerate through tools/regenerate.py. It runs the nine stages in dependency order, which is load-bearing:
python tools/regenerate.py # everything, in order
python tools/regenerate.py --tapp semTAPP # one technique (shared stages still run)
python tools/regenerate.py --dry-run # print the plan, run nothing
python tools/regenerate.py --from resolve # resume at a stage
The order is a dependency chain, not a checklist, and getting it wrong fails SILENTLY. Both known instances produced a green
validate_examples, because dropping a constraint only makes a schema more permissive — so no example can ever detect it. Modules before simplify:simplify_sidecarsblanks a technique row when a module covers the field, and deciding that against module BBs not rebuilt since their sidecars changed deletedLimit of Quantification (LOQ) Methodfrom nine ICP-MS schemas (2026-09-03). Resolve beforebuild_profile's second pass:build_profilebackfills its examples'variableMeasuredentries by readingprofile/resolvedSchema.json, so run too early it reads the previous one.
The stages, if you need to drive them individually:
python tools/bootstrap_schemapaths.py <XLSX> # 0. seed/refresh the schema-path sidecar
python tools/build_module_bb.py --write # 1. module BBs + docs/modules/emitted.json
python tools/simplify_sidecars.py --write # 2. blank rows a module now covers
python tools/build_tapp.py <TAPP_NAME> # 3. registry catalogs + vocab
python tools/build_pathdriven.py <TAPP_NAME> # 4. tapp/ + detail/ schemas from the sidecar
python tools/build_profile.py <TAPP_NAME> # 5. profile/ schema
python tools/resolve_schema.py --all # 6. resolvedSchema.json everywhere
python tools/build_profile.py <TAPP_NAME> # 7. again — backfill example variables
python tools/build_tapp_examples.py <TAPP_NAME> # 8. publication-derived example*.json
python tools/regenerate_schema_json.py # 9. *Schema.json mirrors
python tools/validate_examples.py # then verify
docs/modules/emitted.jsonrecords what the built module$defsactually carry, written bybuild_module_bb --writein the same run that writes the schemas.module_composition.plan()reads it rather than re-deriving coverage from the module sidecars — the sidecar says where a field should go, the built$defsays where it did, and the two diverge whenever a module BB is stale. Reading the manifest makes that fail closed: a stale build yields a stale manifest that agrees with it, so coverage is under-reported and a technique keeps its own row instead of losing the field. Never hand-edit it.
Do not skip step 5.
build_pathdrivendoes not rebuild the publication examples, so a sidecar change moves the schema while they keep the placement they were last generated with — and nothing complains untilvalidate_examplesruns, where it reads as a schema bug rather than a stale artifact.build_tapp_examplesalso PRUNES examples whose publication column has gone from the table; without that, a narrowed table leaves orphaned files that keep being validated.
Step 6 blocker lifted (2026-08-19). Through mid-2026-08
resolve_schema.py --alldegraded its output, because it fetches upstream CDIF$refs from the published mbb gh-pages and that copy carried a dangling$ref: '#/$defs/id-reference'— leaving temp-dir$commentstamps and droppingcdifConceptOrTermOrStringdefs. CDIF now publishesobjectReferenceand the resolver runs clean (6be59b52regenerated against it). If a resolve ever produces temp-dir$comments again, the cause is the same class of stale-gh-pages drift; the local-mbb decouple workaround is in agents.md.
The schema-path sidecar docs/<workbook>.schemapaths.csv is the source of truth for the workbook → schema mapping: one row per (Metadata Item → canonical schema path), with a Source column marking each path authored (human-set, preserved verbatim across re-seeds), inferred (bootstrap's best guess), keyed (routed from the table's Keyed By), module (a composition module owns the placement, so the path is deliberately blank), or flagged (needs a path). A dual-homed editable parameter is two rows — its TAPP default and its detail value. tools/schemapath_io.py reads and writes it; tools/normalize_schema_paths.py canonicalises selector names; the grammar is specified in docs/SCHEMA_PATH_GRAMMAR.md, and docs/README.md explains the sidecars and the guides around them.
tools/build_dataset_template.py <tapp-instance.json> [out.xlsx] generates an xlsx data-entry template from a TAPP instance — columns from targetSpeciesColumns, one row per entry in ada:defaultTargetSpecies.
Superseded drivers.
build_TAPP_from_spreadsheet.pyandbuild_detail_BB.pywere the earlier impl-tag/tier-matrix route and now delegate tobuild_tapp.pyfor epma;build_profile_BB.pyscaffolded the oldprofiles/geochemProfiles/layout.generate_profiles.pyis deprecated and refuses to run without--force-deprecated— its template emits the old object-formada:componentType. Use the path-driven pipeline above for new work.
python tools/interpret_pub_analytes.py # preview only (review files)
python tools/interpret_pub_analytes.py --apply # also rewrite source xlsx
Reads publication columns whose analyte axis isn't explicitly populated and infers it from rows 48 / 59 / 64 (Halogen Correction / Primary Calibration Standard / Typical Detection Limit). Default-mode outputs:
docs/TAPP_EPMA_filled-interp.xlsx— side workbook with each<pub>-interpcolumn inserted right after its source pub for side-by-side review.build/interp-review/example<epmaTAPP|detailEPMA>-<pub>-interp.json— paired review JSON instances built from the inferred data.
With --apply, additionally rewrites rows 32 / 40 / 59 / 64 of each inferred pub column in docs/TAPP_EPMA_filled.xlsx to the pipe-delim convention. After migration, the regular pipeline (build_TAPP_from_spreadsheet.py etc.) reproduces the same rich examples directly from the source — no interp loop needed.
Detection-limit values keep their full text per element (e.g. "SiO2: 0.02 wt%", "<0.03 wt% for TiO2") so context isn't lost in the migration.
tools/resolve_schema.py— resolve all$refinto a structuredresolvedSchema.json($defs+ internal$ref, recursion-safe and ~88–90% smaller than the old fully-inlined form, which is no longer emitted;--structuredis now a no-op). This is the file downstream validators read — the old*StructuredSchema.jsonoutput is gone.tools/regenerate_schema_json.py— generate *Schema.json from schema.yaml sources (YAML→JSON + ref rewrite)tools/schema_path_parser.py/schema_path_emitter.py/normalize_schema_paths.py/bootstrap_schemapaths.py/schemapath_io.py— the schema-path layer (parse a canonical path, materialise the nested structure it implies, canonicalise selector names, seed and read the CSV sidecar)tools/generate_profiles.py— deprecated, refuses to run without--force-deprecated; its template emits the old object-formada:componentType.--liststill works for reference.
tools/audit_building_blocks.py— comprehensive audit: file completeness, schema consistency, resolvedSchema freshness (via the structured resolver), SHACL coverage.isTypeLibraryBBs (reusable$defslibraries with no instantiable root class, e.g.stringArray,parameterValues) are exempt from the standalone-example and SHACL-NodeShape requirements.tools/audit_shacl_coverage.py— check SHACL rules cover all schema.yaml properties; reports missing/extra shapestools/validate_examples.py— validate example JSON files against resolved schemastools/validate_instance.py— profile-aware validation of ADA metadata instancestools/compare_schemas.py— detect drift between schema.yaml and *Schema.jsontools/constraint_census.py— count what the resolved schemas constrain and fail when a constraint disappears. Deleting a restriction only makes a schema more permissive, sovalidate_examplesstays green through a loss; this countsrequirednames,enummembers,const,$reftargets, closed objects and branch counts against the committed baseline indocs/constraint_census.json.--writeto record an intended change,--explain <block>for a breakdown.tools/validate_counterexamples.py— assert that instances which must fail still do.validate_examplesproves valid instances validate; it cannot prove invalid ones do not, and that is the direction constraints go missing. 13 cases indocs/counterexamples.json, each a single mutation of a real example that must be rejected, for a named reason.
tools/download_ecl_methods.py— download analytical method Excel workbooks from the EarthChem Library. Reads methods list from Google Sheets, downloads available workbooks. Supports--dry-run,--output-dir.
tools/augment_register.py— add resolvedSchema URLs to build/register.json for the viewertools/generate_custom_report.py— generate HTML validation report with granular SHACL severity breakdowntools/cors_server.py— local HTTP server with CORS headers for testing the viewer
resolve_schema.py and regenerate_schema_json.py are synced from the canonical copies in metadataBuildingBlocks/tools/. Do not edit locally — update the canonical copy and run python tools/sync_resolve_schema.py --apply from the metadataBuildingBlocks repo. The audit, validation, and report tools were also sourced from that repository.
The tappDefinition building block at _sources/BaseSchema/tappDefinition/ defines a registry-backed Technique-Aligned Protocol Profile (TAPP) definition schema (v3). Was previously methodDefinition. A TAPP definition is modeled as a prov:Plan + cdi:Activity + schema:Action + ada:TAPPDefinition + bios:LabProtocol — all five required in @type.
A TAPP definition is a plan — a reusable procedure that prescribes an analysis — not the analysis event itself. This distinction resolves an apparent conflict with the CDIF provenance model and drives how instrument/tool/reagent fields are placed.
- Two PROV roles. The analysis occurrence is a
prov:Activity— it lives inadaProduct.prov:wasGeneratedBy[](aprov:Activity+schema:Action, followingcdifDataType/cdifProvActivity). That activity references the TAPP as one of itsprov:usedentities (prov:wasGeneratedBy[].prov:used[] → tappDefinition, alongside the actual instrument). The TAPP is therefore a used entity, and in PROV terms a plan used by an activity is aprov:Plan— henceprov:Planin the TAPP@type. cdi:Activityvsprov:Activity.prov:Activity(W3C PROV) is an occurrence — something that happened, thatprov:used/prov:generatedentities.cdi:Activity(DDI-CDI process model) is a design-level description of a process/method — reusable, plan-like. The TAPP usescdi:Activity(which aligns withprov:Plan) because it describes a method; it is not typedprov:Activity. The TAPP'sschema:actionProcess(aschema:HowToofcdi:Activitysteps) is likewise a plan.- Why instrument/tool/reagent are direct properties (no
prov:usedon the TAPP). IncdifProvActivity, an activity's instruments areprov:used[].schema:instrumententities — because an occurrence uses them. A plan does not "use" entities in the provenance sense; it specifies resources. So the TAPP carriesschema:instrument,bios:computationalTool,bios:reagentas direct properties (the BioschemasLabProtocolconvention), and has noprov:used. Theprov:usedpattern operates one level up, on theprov:ActivityinadaProduct.prov:wasGeneratedBy, which uses both the actual instrument and this plan. - Division of labour. The TAPP (plan) fixes the reproducible aspects of the method; the analysis instance leaves the rest to
adaProduct.prov:wasGeneratedByand the technique'stechniqueProfile/geochemProfile/<TECH>/detail/block (per-dataset values). Instrument-type terms populateschema:category(a controlled-vocabularyschema:DefinedTerm); standalone-vs-schema:hasPartplacement of sub-components is a per-field decision recorded in the schema-path sidecar.
Every generated profile/ example emits the link the section above describes, as a reference
inside the analysis activity's prov:used:
{ "@id": "ex:labxctTAPP-P0",
"@type": ["prov:Entity", "prov:Plan", "ada:TAPPDefinition"] }Three things about that shape are load-bearing, and each was arrived at by a failure:
- It is a reference, not an inlined plan. The TAPP is a separate document with its own
@id; copying it whole into every record would duplicate it hundreds of times and leave nothing to navigate to. prov:Entityleads the@type. Baseprov:usedadmits a typed item only through its "inlineprov:Entity"anyOfbranch, which keys on@typecontainingprov:Entity; the bare{@id}branch isadditionalProperties: falseand rejects@typeoutright. PROV-O agrees —prov:Planis a subclass ofprov:Entity— so this is the correct assertion, not a workaround.geochemProduct's TAPP conditional is guarded byschema:name, not by@typealone. It pins an inline TAPP to the fulltappDefinitionschema. Keyed on@typealone it also fired on every reference and failed it on the four properties a reference does not carry.schema:nameseparates the two: the TAPP schema requires it, and a{@id, @type}reference never has it.build_profile._schema()emits the same guard in each technique overlay, so the two layers agree.
Why it matters beyond navigation. The profile's prov:used conditional — the one that pins
that technique's TAPP constraints — keys on this entry. A record that never names its procedure
leaves the conditional with nothing to fire on, so the constraints are silently absent and the
record validates clean. That is the same class of silent failure as a mis-named workflow step.
Ordering, inside build_profile. _name_procedure() runs after _fill_required(). A
reference is complete by construction — {@id, @type} and nothing else — but the sentinel and
typing passes cannot tell that from an object they are meant to finish. Run before them, the
reference came back with schema:instrument: "missing" and four {"@id": "nil:missing"} members
padded into its @type, failing 58 of 97 examples.
The four source-derived profile examples (exampleadaEPMA-UAZ-20260131 and -points,
exampleadaLAMCICPMSUPb-Sundell2021, exampleadaSolutionMCICPMS-ETHZ-20240903) are not
regenerated by this pipeline and carry no link. Adding one would assert which TAPP a real published
analysis followed, which their sources do not say.
schema:instrument.schema:hasPart[additionalType 'Collector'] carries two properties, defined on
the instrument building block so every technique with a Collector part inherits them:
ada:collectorConfiguration |
the assignment as the source states it — free text |
ada:collectors |
the collector table; schema:name is the label other properties reference |
They are not the same information twice. The string is the claim; the table is the reading of it.
The table uses the N=1 fallback that TAPP-keyed-values-design.md Decision 6 settles for every
keyed axis — "rather than admit two shapes, always emit the table: a declaration that does not parse
into members yields a one-row table whose row key is the text as written". One member per position
when the string parses; one member carrying the text when it does not, with the source fields in
schema:additionalProperty as name/value pairs. Parsing these strings to individual collectors is
not tractable in general — they are written for a person to read — so N=1 is the expected case.
An N=1 member names no cup and so makes no per-cup claim, which matters because the resistor values are attested per mass, not per cup.
This comes from our instrument representation rather than a delivered TAPP table, so it is not
generated from a sidecar row. ada:collectorConfiguration used to be registered as a keyed table and
served as the container for eight unrelated items while declaring itself Text (free); those seven
others are now ordinary schema:additionalProperty entries on the Collector.
- TAPP identity (top level) —
schema:name,schema:identifier(DOI),schema:version,schema:measurementTechnique(an array ofschema:DefinedTerm),schema:object(target materials),schema:instrument(one instrument or an array when the method uses several, e.g. LA-ICP-MS = ablation system + ICP-MS),schema:location(laboratory/facility — wasada:laboratory),bios:computationalTool,bios:reagent,schema:creator(wasschema:agent),schema:relatedLink,schema:funding - Standard workflow (
schema:actionProcess) — aschema:HowTocontaining orderedcdi:Activity+schema:Actionsteps: sample preparation, calibration, data acquisition, data processing, quality control. Exactly one step must be namedSample preparationand carrybios:LabProcessinschema:additionalType. - Parameters (
schema:additionalProperty, top level and per step — replaces the retiredada:methodParameters) — each entry is one of two shapes:MethodParameter, aschema:PropertyValueSpecificationfor an editable parameter:schema:defaultValueplusschema:valueRequired,schema:minValue/maxValue,schema:inDefinedTermSet, and the requiredada:fieldScope(method/session/element) andada:dataType(string/number/integer/boolean/date/uri)MethodParameterValue, aschema:PropertyValuefor a read-only parameter, carrying the fixed protocol value inschema:value
- Analyte template (
ada:targetSpeciesTemplate) — per-element column definitions (alsoPropertyValueSpecification) and default analyte rows. Exactly one column must be theTargetSpeciesIdentifierColumn:schema:valueName=analyte, pinned toada:dataType: string,readonlyValue: true,valueRequired: true,ada:tier: M. - Quality metrics (
dqv:hasQualityMeasurement) — at method level and on workflow steps @context— required, and theschema/ada/cdiprefixes are pinned to exact values (noteschemaishttp://schema.org/, not https)
Example files use the sibling example<bbName>-<variant>.json pattern (validated by tools/validate_examples.py):
exampletappDefinition-concord-glass-v1-0-6.json— EPMA WDS tephra glass (Concord University)exampletappDefinition-nmnh-spinel-oxybar-v1.json— EPMA WDS spinel oxybarometry (Smithsonian NMNH)exampletappDefinition-uoc-laicpms-glass-v1.json— LA-ICP-MS volcanic glass trace elements (University of Cologne)
Each technique's tapp/, detail/, and profile/ directories carry their own paired publication-derived examples (exampleepmaTAPP-P0.json, exampledetailEPMA-P0.json, exampleepmaProfile.json, …).
- W3C PROV-O —
prov:Plan(the TAPP is a plan; the analysis occurrence is aprov:ActivityinadaProduct.prov:wasGeneratedBythat references the plan viaprov:used) - Bioschemas —
bios:LabProtocol,bios:LabProcess,bios:computationalTool,bios:reagent - DDI-CDI —
cdi:Activity(design-level process description) for workflow steps - W3C DQV —
dqv:hasQualityMeasurementfor quality metrics - schema.org —
PropertyValueSpecificationfor parameter definitions,Action/HowTo/HowToStepfor workflow