Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

ADA Geochemistry Building Blocks

Modular metadata schema components for documenting geochemical analytical Methods and Datasets. Built using the OGC Building Blocks pattern.

The scheme involves three components:

  1. A Technique-Aligned protocol (TAPP) that defines a analytical procedure, including kinds of samples used, target analytes, instruments used, sample preparation, analysis workflow and data reduction. In the TAPP definition, some of these might be specified as fixed, some might have default values, and some are expected to be specified a the individual session level. The fixed properties are the necessary properties that define the TAPP. There are also properties that apply as the analytical session (or 'analysis event') level, and properties that are specific to the description of individual analytes. The authoritative protocol definition is in an Excel workbook. For discussion purposes, the label 'property' is used for properties in the TAPP that are fixed, and 'parameter' for properties that may be adjusted at the session level. Parameters may have default values specified in the TAPP definition.

  2. A building block JSON schema specific to the protocol. This protocol definition object is registered in a protocol registry and accessible via its URI. The TAPP definition is referenced as a measurementTechnique in dataset metadata.

  3. A technique-specific 'detail' building block JSON schema that defines the parameters that may be assigned values at the individual dataset level. There is one detail block per technique, at _sources/techniqueProfile/<GROUP>/<TECH>/detail/ (GROUP = geochemProfile or adaProfile), not a single 'details' file. Two kinds exist and they compose at different nodes: a dataset-root block overlays schema:Dataset, while a hasPart-item block pins ada:componentType and overlays a distribution part - see agents.md for the split and the grep that re-derives it. The content of this schema is included in the schema for dataset instances to create a metadata schema for Datasets conforming to the profile. Session-level and per-analyte parameters are defined once in a registered parameter registry (parameterValues) and referenced from the detail blocks by URI, so a parameter can be reused across detail definitions; the references are resolved inline into the published resolved schema.

Structure

_sources/ has three top-level areas: the shared base schemas, the shared registries, and one directory per analytical technique.

Most of _sources/ is generated, despite the name. Of its 242 schema.yaml, about 169 are written by a generator — the tapp/, detail/ and profile/ schemas from the TAPP tables and sidecars, the composition modules, the shared registries — along with all 242 resolvedSchema.json, all 242 <name>Schema.json, and the generated examples. What you edit directly is BaseSchema/* and the adaProfile/ profiles. Editing a generated file instead of its generator is the most common way to lose work here; CLAUDE.md has the per-directory table of what writes what.

_sources/
  BaseSchema/           18 shared BBs: geochemProduct (domain-neutral base),
                        adaProduct (extends it), tappDefinition, instrument,
                        laboratory, the file-type blocks (image, imageMap,
                        tabularData, dataCube, collection, document,
                        supDocImage, otherFile, files), structuredData,
                        spatialRegistration, creativeWork, stringArray
  registry/           five catalogs keyed by what the column identifies
    targetSpeciesColumns/      registered BB: PropertyValueSpecification $defs,
                               one per per-species column                        (258)
    monitoredPropertyColumns/  registered BB: $defs for the monitored-property
                               table (mass, cup, edge, X-ray line)                (63)
    reportedPropertyColumns/   registered BB: $defs for reported-quantity columns   (1)
    parameterTemplates/        registered BB: PropertyValueSpecification $defs,
                               editable params                                   (414)
    parameterValues/           registered BB: schema:PropertyValue $defs,
                               fixed values                                     (1185)
    vocab/                     catalog: schema:DefinedTermSet files,
                               by @id not $ref                                   (451)
  techniqueProfile/     one directory per technique (91), under two roots:
    geochemProfile/     the 59 TAPP-aware techniques
      <TECH>/tapp/      the TAPP definition for that technique      (59 techniques)
      <TECH>/detail/    per-dataset analysis-instance detail        (59 techniques)
      <TECH>/profile/   path-driven product profile: geochemProduct +
                        detail + TAPP linkage                        (29 techniques)
      <TECH>/profile-ada/ generic product profile, written by the
                        TAPP tooling                                  (9 techniques)
    adaProfile/         the other 32 techniques, untouched by the TAPP work
      <TECH>/profile-ada/ generic product profile: adaProduct +
                        componentType constraints only               (31 techniques)
      <TECH>/detail/    instrument-detail stub                       (14 techniques)

Profile directory names are not profile names. EPMA/profile-ada publishes adaEPMA; SEM/profile publishes adaSEMFull. A profile's canonical name is the schema:subjectOf.dcterms:conformsTo const inside its own schema — read it from there rather than inferring from the path.

BaseSchema

Shared building blocks. The product profile is split into two layers:

  • geochemProduct is the domain-neutral base product profile, composing the CDIF v1.1 profile schemas via allOf. It carries no ADA-specific requirements: its distribution has an optional schema:additionalType (drawn from the componentType vocabulary), not a required ada:componentType.
  • adaProduct extends geochemProduct (via allOf: [$ref geochemProduct, …]) with the ADA/SAMIS overlays: technique types, instrument/lab/sample, and a required ada:componentType on each distribution. Everything ADA-specific lives here, so geochemProduct stays reusable outside ADA.

The composed CDIF v1.1 profiles (via geochemProduct):

  • cdifCore — core metadata properties
  • cdifDataDescription — variableMeasured with DDI-CDI extensions, @id requirement
  • cdifProvenance — prov:wasGeneratedBy provenance activities
  • cdifManifest — archive distribution with hasPart component files (was cdifArchiveDistribution in CDIF ≤1.0). Applied conditionally: the if/then fires only when a schema:distribution item carries schema:Collection in its @type, so a monolithic single-file distribution isn't held to the manifest rules.

Two BBs extend CDIF core BBs:

  • instrument — extends core CDIF instrument; requires schema:additionalType (at least one entry, e.g. nxs:BaseClass/NXinstrument or a technique term like ada:EPMAInstrument)
  • laboratory — extends core CDIF spatialExtent (schema:Place with nxs:BaseClass/NXsource in additionalType)

tappDefinition is documented in its own section below.

registry (shared catalogs)

targetSpeciesColumns, monitoredPropertyColumns, reportedPropertyColumns, parameterTemplates, and parameterValues are each a registered type-library building block (bblock.json with isTypeLibrary: true): every entry lives as a named $def in the catalog's schema.yaml, and TAPP / detail blocks reference them by URI fragment ($ref: …/<catalog>/schema.yaml#/$defs/<name>). Because they are registered, the OGC bblocks annotate step resolves those refs locally via the register and inlines them into resolvedSchema.json. This matters: a loose helper file (a plain <name>.json not inside a registered BB) is instead fetched from the published gh-pages URL, which 404s on moved or unpublished paths (process-bblocks.yml sets skip-pages: true, so gh-pages never auto-updates) — that fragility is why the catalogs were promoted to registered BBs. vocab/ is the exception: it stays a plain catalog of schema:DefinedTermSet files because it is referenced only by JSON-LD @id (schema:inDefinedTermSet), never by $ref, so the annotate step never fetches it.

The catalogs are shared dictionary resources — multiple TAPPs $ref the same $defs when their definitions match. share_or_write_catalog lets a TAPP regen overwrite its own entries (matched by $id ownership) but errors out on a collision with an entry originated by a different TAPP, so a new TAPP either reuses identical catalog entries or surfaces a renaming requirement.

parameterTemplates holds editable parameters (a PropertyValueSpecification with a default the analyst may override); parameterValues holds fixed protocol values (a schema:PropertyValue). That split — specification vs value — is how read-only-ness is expressed; ada:methodParameters was retired repo-wide in favour of schema:additionalProperty.

techniqueProfile/geochemProfile/<TECH>

All 59 techniques under geochemProfile/ have a tapp/ and a detail/. 29 of them also publish a path-driven profile/ — the set registered in build_profile.PROFILES.

  • tapp/ — the protocol definition. Extends tappDefinition via allOf with technique-specific top-level ada: properties, schema:additionalProperty[] entries, and ada:targetSpeciesTemplate.ada:targetSpeciesColumns constraints referencing the registry catalogs.
  • detail/ — the per-dataset analysis instance. Placement is not uniform, and does not track whether the technique is path-driven. Seven overlay the schema:Dataset root (analyst contributor, session dates, sample, funding, per-analysis parameter values): Basemap, EPMA, Geochron, SEM, SEM-Composition, Solution-Q-ICPMS, Solution-SF-ICPMS. The other eighteen pin ada:componentType and overlay a schema:distribution.hasPart item: ARGT, DSC, EAIRMS, ICPOES, L2MS, LA-ICPMS, LAF, NanoIR, NanoSIMS, PSFD, QRIS, SEM-FIBSEM, SEM-Imaging, SLS, TEM, VNMIR, XCT, XRD. Consumers cannot assume one placement.
  • profile/ — path-driven product profile: bases on the domain-neutral geochemProduct + the detail block + prov:used narrowed to that technique's TAPP + the technique's ada:componentType enum on hasPart (the profile layers the ADA componentType constraint on top of the ADA-agnostic base).
  • profile-ada/ — the generic product profile: bases on adaProduct + ada:componentType constraints only, no TAPP linkage or detail block.

Base-selection rule. geochemProfile/<TECH>/profile/ (generic, path-driven) → geochemProduct; geochemProfile/<TECH>/profile-ada/ (the 4 ADA variants) and all adaProfile/<TECH>/profile-ada/ → adaProduct. Principle: geochem→geochemProduct, ada→adaProduct.

A dataset instance selects between the two profile variants by how it references its protocol: a bare {"@id": …} node reference in schema:measurementTechnique targets the path-driven profile, an inline schema:DefinedTerm targets the generic one.

Module composition

A module defines a shared field once instead of once per technique. _sources/BaseSchema/modules/* is generated from the sidecars in docs/modules/Module_*.schemapaths.csv — never hand-edit a module.

Across the 16 technique tapp/ schemas, 651 parameter slots are composed from a module against 231 minted per technique (73%), ranging from 100% for Solution-MC-ICPMS down to 7% for Lab-XCT. The spread tracks which modules exist: the 2026-09 delivery added ICPMS, CollisionCell and CompositionQC, which took the ICP-MS families from ~25% to 78–100%, while electron-beam and tomography still have no family module.

That percentage counts parameters — entries under schema:additionalProperty — not all properties, and it ignores the module ROOT $defs a technique composes, so it understates what a module supplies. 100% is not reachable: of the 79 distinct fields no module covers, 50 are carried by exactly one table and a module needs two consumers. See MODULE_CONSOLIDATION_STATUS.md for the full definition and the ~85% ceiling.

  • docs/modules/MODULE_CONSOLIDATION_STATUS.md — current state: what was measured and how to reproduce it, what is decided, what is open, and the next drafting pass. Start here.
  • agents.md §"Module composition" — how composition actually works, including why parameters compose differently from structural fields and why some duplication is correct.
  • docs/modules/draft/ — 8 provisional Draft_Module_*.csv, all electron-beam. Ours and provisional; the library's modules are Ruolin's to author. The six ICP-MS drafts were adopted upstream in the 2026-09 delivery and deleted from here; eight of their fields were not taken up and are recorded in docs/upstream-requests.md §1.

One measurement warning: do not count duplication in _sources/registry/. Those catalogues are per-technique by construction, so they report ~87% duplication whatever composition does. Measure the tapp/ schemas.

componentType architecture

Each archive hasPart item carries an ada:componentType (a single string like ada:EPMAImageMap) that classifies the file. The term list is governed by a vocabulary, and two schema layers add per-context constraints:

Governing vocabulary. registry/vocab/componentType.json is a SKOS ConceptScheme (@id: ada:vocab/componentType, the ~22 universal cross-technique terms). The base products reference it by annotation only — the universalComponentType $def in geochemProduct (and duplicated in adaProduct) is {type: string, schema:inDefinedTermSet: "ada:vocab/componentType"} with no inline enum, so at the base layer any string validates and conformance to the vocabulary is advisory (SHACL-checkable), not hard-enforced by JSON Schema. geochemProduct exposes the vocab as an optional schema:additionalType; adaProduct requires it as ada:componentType.

  1. File type ↔ componentType mapping — each file-type building block (image, imageMap, tabularData, collection, dataCube, document, supDocImage, otherFile) declares a sealed enum of valid componentType values. The enum is derived from the Components worksheet of amds-ldeo/metadata/ADA-AnalyticalMethodsAndAttributes.xlsx (the canonical mapping; columns componentType / FileType / isSupplement). E.g. ada:EPMAImageMap is valid only on parts whose @type includes ada:imageMap.

  2. Profile-level constraint — a technique profile's schema:distribution.items.schema:hasPart.items uses a schema-level anyOf with three kinds of branch: (a) $ref to geochemProduct/schema.yaml#/$defs/universalComponentTypeBranch (factored once, used everywhere) for universal componentTypes; (b) inline string-enum for technique-specific componentTypes; (c) for techniques whose detail/ block is the older hasPart-item kind (XRD, ARGT, DSC, …), a $ref to that detail schema, which pins ada:componentType to its technique consts and contributes detail-specific sibling properties (e.g. ada:geometry) flat on the hasPart item — not nested inside componentType. Path-driven profiles do not use branch (c): their detail block overlays the dataset root instead, and hasPart gets only branches (a) and (b).

Keeping the layers in sync. Because the base layer is annotation-only, JSON-Schema validation no longer catches componentType drift on its own. python tools/check_componentType.py restores that check: it fails if a universal vocab term is missing from the enum cache, if a base schema stops annotating the vocab @id, or if any ada:componentType used in an example is not a known term (universal vocab ∪ enum cache ∪ per-technique profile enums). Run it after touching the worksheet, the vocab, or example componentTypes.

Refreshing the mapping

After editing the Components worksheet:

python tools/apply_componentType_enums.py --refresh \
    --xlsx ../../amds-ldeo/metadata/ADA-AnalyticalMethodsAndAttributes.xlsx
python tools/regenerate_schema_json.py
python tools/resolve_schema.py --all
python tools/validate_examples.py
python tools/check_componentType.py    # confirm vocab / cache / schemas / examples agree

The cached mapping at tools/componentType_enum_cache.json is committed so the apply step works on a fresh clone without spreadsheet access. Worksheet FileType values map to file-type BBs as: image→image (or supDocImage when isSupplement=supplement), imageMap→imageMap, tabularData→tabularData, archive→collection, dataCube→dataCube, document→document, video/otherFile→otherFile, plus the document | image and document | tabularData splits.

Cross-repo imports

This repository imports shared schema.org and CDIF property building blocks from metadataBuildingBlocks via the OGC Building Blocks import mechanism. All external references use absolute URLs (https://cross-domain-interoperability-framework.github.io/metadataBuildingBlocks/_sources/...).

Continuous integration

Four workflows live in .github/workflows/. Three report a check on every pull request and are required in branch protection on main:

check workflow what it does
Validate and annotate (no pages) validate-branch.yml the full OGC postprocess (validate + annotate + build register/tests), no Pages deploy
Regenerate and diff check-schema-drift.yml runs tools/regenerate.py and fails if any committed artifact moves
Regenerate twice and compare check-determinism.yml regenerates under two different PYTHONHASHSEEDs and compares, so a generator cannot vary by run

process-bblocks.yml is the fourth. It runs on pushes to main and produces the published build/ tree.

docs/SILENT_SUCCESS.md catalogues the incidents behind this setup — six cases where something in this pipeline reported success while not doing its job, what exposed each one, and the rule that follows. Read it before changing a check or a trigger.

Two consequences worth knowing before editing any of them:

  • A required workflow must not carry a paths: filter. A required check that never runs is pending forever rather than skipped, and blocks the pull request indefinitely.
  • Auto-merge waits only on required checks. It will merge past a failing check that is not required.

How build/ reaches main

build/ is committed, because deploy-viewer.yml reads build/register.json and build/tests/report.json out of main and publishes the tree to GitHub Pages without regenerating them.

Branch protection declines a direct push from the postprocess, and GitHub offers no way to exempt it — the GitHub Actions app can only be a bypass actor on an organization ruleset, while a repository ruleset accepts only deploy-key and repository-role bypasses. So the generated output arrives the same way every other change does, as a pull request:

  1. process-bblocks.yml force-resets the bblocks-build branch to main;
  2. the reusable OGC postprocess runs against that branch and commits its output there;
  3. a pull request is opened from it and auto-merge is armed, so it lands once the three required checks pass.

Nothing bypasses protection: the regenerated build/ is reviewed by the same checks as authored source. Step 3 uses a fine-grained PAT (BBLOCKS_PR_TOKEN, Contents:read + PullRequests:write on this repository only) for one reason — a pull request opened by GITHUB_TOKEN triggers no workflows, so its required checks would never report and it could never merge.

bblocks-build is machine-owned. It is force-reset on every postprocess run. Do not branch from it, commit to it, or base work on it.

Viewer

Browse the building blocks at: https://amds-ldeo.github.io/geochemBuildingBlocks/

Human-readable record and TAPP pages

tools/_tapp_lib.write_profile_ada_companions(<block dir>) writes the four files a profile-ada block ships beside its schema — context.jsonld, description.md, rules.shacl and examples.yaml. They used to come from generate_profiles.py, now blocked because its SCHEMA template emits the retired object-form ada:componentType; _tapp_lib took the schema over and nothing took the companions, so five blocks added after the deprecation shipped without them. It reads each block's own schema.yaml and bblock.json rather than a registry, so a block that exists is describable whether or not anyone registered it. examples.yaml is written only when an example*.json exists — emitting a ref: to a missing file would satisfy the audit while pointing at nothing.

tools/build_html_views.py renders two page types into build/htmlViews/: a dataset record page for a product profile instance, and a TAPP definition page for the protocol it names. The dataset page follows the record's own prov:used TAPP reference through to that TAPP's page.

Keyed values render as a grid: one row per member of the keyset the defines: row declares (each target species, each monitored property), one column per property keyed to that set.

python tools/build_html_views.py --source examples --all      # the 97 schema examples
python tools/build_html_views.py --source ada2 --limit 50     # real ADA holdings, public.json_table
python tools/build_html_views.py --source ada2 --doi <doi>    # one record
python tools/build_html_views.py --all --no-tapp-pages

--source ada2 reads public.json_table using ADA_NAME / DB_2024_USER / DB_2024_PASSWORD / DB_2024_HOST / DB_2024_PORT — the same variables the metadata loaders use, not a PGSERVICEFILE or .pgpass entry.

Output lands in build/htmlViews/, which is not committed. The rest of build/ — the annotated schemas, OAS3 downcompiles, register.json, bblocks.jsonld/.ttl and test report — IS committed, generated by the OGC postprocess in CI; see Continuous integration.

Tools

TAPP / detail / profile generation pipeline

docs/TAPP-schema-generation-workflow.md is the authoritative walkthrough — written for three audiences (workbook author, pipeline maintainer, form builder) with a flowchart of the whole path from spreadsheet to validated schema. Read it first; the summary here is orientation only.

TAPP source = the tapp/ git submodule (amds-ldeo/tapp). The TAPP tables and modules live in that submodule, not in this repo; tools/tapp_source.py:current_delivery() resolves to tapp/ (falling back to any inline TAPPS<date>/ drop). Clone with git clone --recursive, or run git submodule update --init in an existing checkout, before regenerating. Pin/bump the delivery by updating the submodule commit — a deliberate act, because .gitmodules sets update = none to stop the OGC postprocess workflow advancing the pointer on its own. The committed schemas are built from the pinned revision (af3f7bc, adopted 2026-09-02).

One upstream-authored TAPP table per technique (a CSV in the tapp/ submodule's Current TAPPs/) drives everything downstream. Nothing generated should ever be hand-edited — fix the table (upstream, in amds-ldeo/tapp) or a tool and regenerate.

Regenerate through tools/regenerate.py. It runs the nine stages in dependency order, which is load-bearing:

python tools/regenerate.py                 # everything, in order
python tools/regenerate.py --tapp semTAPP  # one technique (shared stages still run)
python tools/regenerate.py --dry-run       # print the plan, run nothing
python tools/regenerate.py --from resolve  # resume at a stage

The order is a dependency chain, not a checklist, and getting it wrong fails SILENTLY. Both known instances produced a green validate_examples, because dropping a constraint only makes a schema more permissive — so no example can ever detect it. Modules before simplify: simplify_sidecars blanks a technique row when a module covers the field, and deciding that against module BBs not rebuilt since their sidecars changed deleted Limit of Quantification (LOQ) Method from nine ICP-MS schemas (2026-09-03). Resolve before build_profile's second pass: build_profile backfills its examples' variableMeasured entries by reading profile/resolvedSchema.json, so run too early it reads the previous one.

The stages, if you need to drive them individually:

python tools/bootstrap_schemapaths.py  <XLSX>        # 0. seed/refresh the schema-path sidecar
python tools/build_module_bb.py --write              # 1. module BBs + docs/modules/emitted.json
python tools/simplify_sidecars.py --write            # 2. blank rows a module now covers
python tools/build_tapp.py             <TAPP_NAME>   # 3. registry catalogs + vocab
python tools/build_pathdriven.py       <TAPP_NAME>   # 4. tapp/ + detail/ schemas from the sidecar
python tools/build_profile.py          <TAPP_NAME>   # 5. profile/ schema
python tools/resolve_schema.py --all                 # 6. resolvedSchema.json everywhere
python tools/build_profile.py          <TAPP_NAME>   # 7. again — backfill example variables
python tools/build_tapp_examples.py    <TAPP_NAME>   # 8. publication-derived example*.json
python tools/regenerate_schema_json.py               # 9. *Schema.json mirrors
python tools/validate_examples.py                    # then verify

docs/modules/emitted.json records what the built module $defs actually carry, written by build_module_bb --write in the same run that writes the schemas. module_composition.plan() reads it rather than re-deriving coverage from the module sidecars — the sidecar says where a field should go, the built $def says where it did, and the two diverge whenever a module BB is stale. Reading the manifest makes that fail closed: a stale build yields a stale manifest that agrees with it, so coverage is under-reported and a technique keeps its own row instead of losing the field. Never hand-edit it.

Do not skip step 5. build_pathdriven does not rebuild the publication examples, so a sidecar change moves the schema while they keep the placement they were last generated with — and nothing complains until validate_examples runs, where it reads as a schema bug rather than a stale artifact. build_tapp_examples also PRUNES examples whose publication column has gone from the table; without that, a narrowed table leaves orphaned files that keep being validated.

Step 6 blocker lifted (2026-08-19). Through mid-2026-08 resolve_schema.py --all degraded its output, because it fetches upstream CDIF $refs from the published mbb gh-pages and that copy carried a dangling $ref: '#/$defs/id-reference' — leaving temp-dir $comment stamps and dropping cdifConceptOrTermOrString defs. CDIF now publishes objectReference and the resolver runs clean (6be59b52 regenerated against it). If a resolve ever produces temp-dir $comments again, the cause is the same class of stale-gh-pages drift; the local-mbb decouple workaround is in agents.md.

The schema-path sidecar docs/<workbook>.schemapaths.csv is the source of truth for the workbook → schema mapping: one row per (Metadata Item → canonical schema path), with a Source column marking each path authored (human-set, preserved verbatim across re-seeds), inferred (bootstrap's best guess), keyed (routed from the table's Keyed By), module (a composition module owns the placement, so the path is deliberately blank), or flagged (needs a path). A dual-homed editable parameter is two rows — its TAPP default and its detail value. tools/schemapath_io.py reads and writes it; tools/normalize_schema_paths.py canonicalises selector names; the grammar is specified in docs/SCHEMA_PATH_GRAMMAR.md, and docs/README.md explains the sidecars and the guides around them.

tools/build_dataset_template.py <tapp-instance.json> [out.xlsx] generates an xlsx data-entry template from a TAPP instance — columns from targetSpeciesColumns, one row per entry in ada:defaultTargetSpecies.

Superseded drivers. build_TAPP_from_spreadsheet.py and build_detail_BB.py were the earlier impl-tag/tier-matrix route and now delegate to build_tapp.py for epma; build_profile_BB.py scaffolded the old profiles/geochemProfiles/ layout. generate_profiles.py is deprecated and refuses to run without --force-deprecated — its template emits the old object-form ada:componentType. Use the path-driven pipeline above for new work.

Publication migration helper

python tools/interpret_pub_analytes.py            # preview only (review files)
python tools/interpret_pub_analytes.py --apply    # also rewrite source xlsx

Reads publication columns whose analyte axis isn't explicitly populated and infers it from rows 48 / 59 / 64 (Halogen Correction / Primary Calibration Standard / Typical Detection Limit). Default-mode outputs:

  • docs/TAPP_EPMA_filled-interp.xlsx — side workbook with each <pub>-interp column inserted right after its source pub for side-by-side review.
  • build/interp-review/example<epmaTAPP|detailEPMA>-<pub>-interp.json — paired review JSON instances built from the inferred data.

With --apply, additionally rewrites rows 32 / 40 / 59 / 64 of each inferred pub column in docs/TAPP_EPMA_filled.xlsx to the pipe-delim convention. After migration, the regular pipeline (build_TAPP_from_spreadsheet.py etc.) reproduces the same rich examples directly from the source — no interp loop needed.

Detection-limit values keep their full text per element (e.g. "SiO2: 0.02 wt%", "<0.03 wt% for TiO2") so context isn't lost in the migration.

Schema generation and resolution

  • tools/resolve_schema.py — resolve all $ref into a structured resolvedSchema.json ($defs + internal $ref, recursion-safe and ~88–90% smaller than the old fully-inlined form, which is no longer emitted; --structured is now a no-op). This is the file downstream validators read — the old *StructuredSchema.json output is gone.
  • tools/regenerate_schema_json.py — generate *Schema.json from schema.yaml sources (YAML→JSON + ref rewrite)
  • tools/schema_path_parser.py / schema_path_emitter.py / normalize_schema_paths.py / bootstrap_schemapaths.py / schemapath_io.py — the schema-path layer (parse a canonical path, materialise the nested structure it implies, canonicalise selector names, seed and read the CSV sidecar)
  • tools/generate_profiles.py — deprecated, refuses to run without --force-deprecated; its template emits the old object-form ada:componentType. --list still works for reference.

Validation and auditing

  • tools/audit_building_blocks.py — comprehensive audit: file completeness, schema consistency, resolvedSchema freshness (via the structured resolver), SHACL coverage. isTypeLibrary BBs (reusable $defs libraries with no instantiable root class, e.g. stringArray, parameterValues) are exempt from the standalone-example and SHACL-NodeShape requirements.
  • tools/audit_shacl_coverage.py — check SHACL rules cover all schema.yaml properties; reports missing/extra shapes
  • tools/validate_examples.py — validate example JSON files against resolved schemas
  • tools/validate_instance.py — profile-aware validation of ADA metadata instances
  • tools/compare_schemas.py — detect drift between schema.yaml and *Schema.json
  • tools/constraint_census.py — count what the resolved schemas constrain and fail when a constraint disappears. Deleting a restriction only makes a schema more permissive, so validate_examples stays green through a loss; this counts required names, enum members, const, $ref targets, closed objects and branch counts against the committed baseline in docs/constraint_census.json. --write to record an intended change, --explain <block> for a breakdown.
  • tools/validate_counterexamples.py — assert that instances which must fail still do. validate_examples proves valid instances validate; it cannot prove invalid ones do not, and that is the direction constraints go missing. 13 cases in docs/counterexamples.json, each a single mutation of a real example that must be rejected, for a named reason.

Data collection

  • tools/download_ecl_methods.py — download analytical method Excel workbooks from the EarthChem Library. Reads methods list from Google Sheets, downloads available workbooks. Supports --dry-run, --output-dir.

Build and deployment support

  • tools/augment_register.py — add resolvedSchema URLs to build/register.json for the viewer
  • tools/generate_custom_report.py — generate HTML validation report with granular SHACL severity breakdown
  • tools/cors_server.py — local HTTP server with CORS headers for testing the viewer

Tool provenance

resolve_schema.py and regenerate_schema_json.py are synced from the canonical copies in metadataBuildingBlocks/tools/. Do not edit locally — update the canonical copy and run python tools/sync_resolve_schema.py --apply from the metadataBuildingBlocks repo. The audit, validation, and report tools were also sourced from that repository.

TAPP Definition Building Block

The tappDefinition building block at _sources/BaseSchema/tappDefinition/ defines a registry-backed Technique-Aligned Protocol Profile (TAPP) definition schema (v3). Was previously methodDefinition. A TAPP definition is modeled as a prov:Plan + cdi:Activity + schema:Action + ada:TAPPDefinition + bios:LabProtocol — all five required in @type.

The TAPP is a plan, not an occurrence (PROV alignment)

A TAPP definition is a plan — a reusable procedure that prescribes an analysis — not the analysis event itself. This distinction resolves an apparent conflict with the CDIF provenance model and drives how instrument/tool/reagent fields are placed.

  • Two PROV roles. The analysis occurrence is a prov:Activity — it lives in adaProduct.prov:wasGeneratedBy[] (a prov:Activity + schema:Action, following cdifDataType/cdifProvActivity). That activity references the TAPP as one of its prov:used entities (prov:wasGeneratedBy[].prov:used[] → tappDefinition, alongside the actual instrument). The TAPP is therefore a used entity, and in PROV terms a plan used by an activity is a prov:Plan — hence prov:Plan in the TAPP @type.
  • cdi:Activity vs prov:Activity. prov:Activity (W3C PROV) is an occurrence — something that happened, that prov:used/prov:generated entities. cdi:Activity (DDI-CDI process model) is a design-level description of a process/method — reusable, plan-like. The TAPP uses cdi:Activity (which aligns with prov:Plan) because it describes a method; it is not typed prov:Activity. The TAPP's schema:actionProcess (a schema:HowTo of cdi:Activity steps) is likewise a plan.
  • Why instrument/tool/reagent are direct properties (no prov:used on the TAPP). In cdifProvActivity, an activity's instruments are prov:used[].schema:instrument entities — because an occurrence uses them. A plan does not "use" entities in the provenance sense; it specifies resources. So the TAPP carries schema:instrument, bios:computationalTool, bios:reagent as direct properties (the Bioschemas LabProtocol convention), and has no prov:used. The prov:used pattern operates one level up, on the prov:Activity in adaProduct.prov:wasGeneratedBy, which uses both the actual instrument and this plan.
  • Division of labour. The TAPP (plan) fixes the reproducible aspects of the method; the analysis instance leaves the rest to adaProduct.prov:wasGeneratedBy and the technique's techniqueProfile/geochemProfile/<TECH>/detail/ block (per-dataset values). Instrument-type terms populate schema:category (a controlled-vocabulary schema:DefinedTerm); standalone-vs-schema:hasPart placement of sub-components is a per-field decision recorded in the schema-path sidecar.

How a record names the TAPP it followed

Every generated profile/ example emits the link the section above describes, as a reference inside the analysis activity's prov:used:

{ "@id": "ex:labxctTAPP-P0",
  "@type": ["prov:Entity", "prov:Plan", "ada:TAPPDefinition"] }

Three things about that shape are load-bearing, and each was arrived at by a failure:

  • It is a reference, not an inlined plan. The TAPP is a separate document with its own @id; copying it whole into every record would duplicate it hundreds of times and leave nothing to navigate to.
  • prov:Entity leads the @type. Base prov:used admits a typed item only through its "inline prov:Entity" anyOf branch, which keys on @type containing prov:Entity; the bare {@id} branch is additionalProperties: false and rejects @type outright. PROV-O agrees — prov:Plan is a subclass of prov:Entity — so this is the correct assertion, not a workaround.
  • geochemProduct's TAPP conditional is guarded by schema:name, not by @type alone. It pins an inline TAPP to the full tappDefinition schema. Keyed on @type alone it also fired on every reference and failed it on the four properties a reference does not carry. schema:name separates the two: the TAPP schema requires it, and a {@id, @type} reference never has it. build_profile._schema() emits the same guard in each technique overlay, so the two layers agree.

Why it matters beyond navigation. The profile's prov:used conditional — the one that pins that technique's TAPP constraints — keys on this entry. A record that never names its procedure leaves the conditional with nothing to fire on, so the constraints are silently absent and the record validates clean. That is the same class of silent failure as a mis-named workflow step.

Ordering, inside build_profile. _name_procedure() runs after _fill_required(). A reference is complete by construction — {@id, @type} and nothing else — but the sentinel and typing passes cannot tell that from an object they are meant to finish. Run before them, the reference came back with schema:instrument: "missing" and four {"@id": "nil:missing"} members padded into its @type, failing 58 of 97 examples.

The four source-derived profile examples (exampleadaEPMA-UAZ-20260131 and -points, exampleadaLAMCICPMSUPb-Sundell2021, exampleadaSolutionMCICPMS-ETHZ-20240903) are not regenerated by this pipeline and carry no link. Adding one would assert which TAPP a real published analysis followed, which their sources do not say.

How a collector is described (E1)

schema:instrument.schema:hasPart[additionalType 'Collector'] carries two properties, defined on the instrument building block so every technique with a Collector part inherits them:

ada:collectorConfiguration the assignment as the source states it — free text
ada:collectors the collector table; schema:name is the label other properties reference

They are not the same information twice. The string is the claim; the table is the reading of it.

The table uses the N=1 fallback that TAPP-keyed-values-design.md Decision 6 settles for every keyed axis — "rather than admit two shapes, always emit the table: a declaration that does not parse into members yields a one-row table whose row key is the text as written". One member per position when the string parses; one member carrying the text when it does not, with the source fields in schema:additionalProperty as name/value pairs. Parsing these strings to individual collectors is not tractable in general — they are written for a person to read — so N=1 is the expected case.

An N=1 member names no cup and so makes no per-cup claim, which matters because the resistor values are attested per mass, not per cup.

This comes from our instrument representation rather than a delivered TAPP table, so it is not generated from a sidecar row. ada:collectorConfiguration used to be registered as a keyed table and served as the container for eight unrelated items while declaring itself Text (free); those seven others are now ordinary schema:additionalProperty entries on the Collector.

Structure

  • TAPP identity (top level) — schema:name, schema:identifier (DOI), schema:version, schema:measurementTechnique (an array of schema:DefinedTerm), schema:object (target materials), schema:instrument (one instrument or an array when the method uses several, e.g. LA-ICP-MS = ablation system + ICP-MS), schema:location (laboratory/facility — was ada:laboratory), bios:computationalTool, bios:reagent, schema:creator (was schema:agent), schema:relatedLink, schema:funding
  • Standard workflow (schema:actionProcess) — a schema:HowTo containing ordered cdi:Activity + schema:Action steps: sample preparation, calibration, data acquisition, data processing, quality control. Exactly one step must be named Sample preparation and carry bios:LabProcess in schema:additionalType.
  • Parameters (schema:additionalProperty, top level and per step — replaces the retired ada:methodParameters) — each entry is one of two shapes:
    • MethodParameter, a schema:PropertyValueSpecification for an editable parameter: schema:defaultValue plus schema:valueRequired, schema:minValue/maxValue, schema:inDefinedTermSet, and the required ada:fieldScope (method/session/element) and ada:dataType (string/number/integer/boolean/date/uri)
    • MethodParameterValue, a schema:PropertyValue for a read-only parameter, carrying the fixed protocol value in schema:value
  • Analyte template (ada:targetSpeciesTemplate) — per-element column definitions (also PropertyValueSpecification) and default analyte rows. Exactly one column must be the TargetSpeciesIdentifierColumn: schema:valueName = analyte, pinned to ada:dataType: string, readonlyValue: true, valueRequired: true, ada:tier: M.
  • Quality metrics (dqv:hasQualityMeasurement) — at method level and on workflow steps
  • @context — required, and the schema / ada / cdi prefixes are pinned to exact values (note schema is http://schema.org/, not https)

Examples

Example files use the sibling example<bbName>-<variant>.json pattern (validated by tools/validate_examples.py):

  • exampletappDefinition-concord-glass-v1-0-6.json — EPMA WDS tephra glass (Concord University)
  • exampletappDefinition-nmnh-spinel-oxybar-v1.json — EPMA WDS spinel oxybarometry (Smithsonian NMNH)
  • exampletappDefinition-uoc-laicpms-glass-v1.json — LA-ICP-MS volcanic glass trace elements (University of Cologne)

Each technique's tapp/, detail/, and profile/ directories carry their own paired publication-derived examples (exampleepmaTAPP-P0.json, exampledetailEPMA-P0.json, exampleepmaProfile.json, …).

Vocabularies used

  • W3C PROV-O — prov:Plan (the TAPP is a plan; the analysis occurrence is a prov:Activity in adaProduct.prov:wasGeneratedBy that references the plan via prov:used)
  • Bioschemas — bios:LabProtocol, bios:LabProcess, bios:computationalTool, bios:reagent
  • DDI-CDI — cdi:Activity (design-level process description) for workflow steps
  • W3C DQV — dqv:hasQualityMeasurement for quality metrics
  • schema.org — PropertyValueSpecification for parameter definitions, Action/HowTo/HowToStep for workflow

License

Apache 2.0

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages