Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
d9cc9bd
feat(SOF-8051): notebook that uploads an SPM run
VsevolodX Sep 21, 2026
be4b3c6
fix(SOF-8051): the upload notebook carries its script and one host
VsevolodX Sep 21, 2026
fc34ff1
fix(SOF-8051): uploader follows the placement model — a sample set pe…
VsevolodX Sep 21, 2026
c1e6e31
fix(SOF-8051): uploader review round; the notebook says wafer
VsevolodX Sep 21, 2026
bc4b50d
fix(SOF-8051): --emit-example thins the curves
VsevolodX Sep 22, 2026
0477bdb
fix(SOF-8051): the uploader takes a physical id and leaves sets bare
VsevolodX Sep 22, 2026
2fa6494
fix(SOF-8051): sample names from the given physical id
VsevolodX Sep 22, 2026
44b459b
fix(SOF-8051): PHYSICAL_ID has no default — the human names the piece
VsevolodX Sep 22, 2026
9600f40
feat(SOF-8051): notebook that uploads NLR's data for one piece
VsevolodX Sep 22, 2026
a092282
fix(SOF-8051): the notebooks point at alphafilm.mat3ra.com, where the…
VsevolodX Sep 22, 2026
b314ad1
fix(SOF-8051): the install cell brings mat3ra-notebooks-utils, which …
VsevolodX Sep 22, 2026
e603a48
fix(SOF-8051): the NLR path validates, and its inputs are checked bef…
VsevolodX Sep 22, 2026
0b17199
fix(SOF-8051): an existing sample set takes metadata it does not have…
VsevolodX Sep 22, 2026
588d8b8
fix(SOF-8051): a duplicate pad is an error, and a set takes records i…
VsevolodX Sep 22, 2026
31e7b6a
refactor(SOF-8051): one upload path, and `files` is the list of what …
VsevolodX Sep 22, 2026
7dedd1a
fix(SOF-8051): an image keeps the path that makes its name unique
VsevolodX Sep 22, 2026
8be3766
refactor(SOF-8051): the uploader is the uploader; each instrument rea…
VsevolodX Sep 22, 2026
3cb193b
refactor(SOF-8051): the uploader takes documents; parsing, serializin…
VsevolodX Sep 22, 2026
54d86cc
fix(SOF-8051): pin the standata branch the instruments are registered on
VsevolodX Sep 22, 2026
0a6eb9c
feat(SOF-8051): one file states exactly what to install, and the exam…
VsevolodX Sep 22, 2026
85a7867
update cleanup
VsevolodX Sep 22, 2026
ef1b1df
fix(SOF-8051): pin the versions, drop the machinery that asked which …
VsevolodX Sep 22, 2026
0294e9e
fix(SOF-8051): validate against the schema, not a generated model
VsevolodX Sep 22, 2026
ec90ca6
fix(SOF-8051): the API token leaves the notebook, and three review po…
VsevolodX Sep 22, 2026
7f9b805
fix(SOF-8051): put back the notebook cells that were yours
VsevolodX Sep 22, 2026
d11829e
clean
VsevolodX Sep 22, 2026
7b9606c
clean
VsevolodX Sep 22, 2026
e9475ed
revert(SOF-8051): the notebooks do not go in the docs
VsevolodX Sep 22, 2026
52f6be2
feat(SOF-8051): the Library - the piece every Sample Set on it belong…
VsevolodX Oct 6, 2026
0587f3f
feat(SOF-8051): topography metrics already on the platform become pro…
VsevolodX Oct 6, 2026
5298b52
fix(SOF-8051): three review findings on the Library upload
VsevolodX Oct 6, 2026
c8d8f7f
SOF-8051: NLR pads inferred from the I-V file; wafer dimensions and f…
VsevolodX Oct 6, 2026
545372e
SOF-8051: film thickness in um, as the XRF file gives it
VsevolodX Oct 6, 2026
5feffec
SOF-8051: README - XRF is on the bare-film grid, I-V on the pads
VsevolodX Oct 6, 2026
0cafb2b
SOF-8051: placeholder instrument names; the deliveries do not name th…
VsevolodX Oct 6, 2026
8b5feb1
SOF-8051: NLR photographs are the Library's, not the XRF grid set's
VsevolodX Oct 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
79 changes: 79 additions & 0 deletions examples/measurement/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
# Uploading a measurement run

Data a lab delivers becomes Samples, Measurements and Properties on the platform in two steps that stay apart:
**parse** reads one lab's delivery and writes a *run document*; **upload** takes run documents and sends them.
Nothing in the uploader knows what an instrument is, and no parser talks to the platform.

```
delivery ──► parse_utk.py ──► run document ──► upload_run.py ──► platform
parse_nlr.py (parsed/*.json)
```

## Install

```bash
python -m venv venv && . venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
```

api-client, standata and esse are pinned to the branch carrying the REST endpoints, the instrument registry
entries and the Sample and Measurement schemas. They become version pins when it releases.

## Credentials

The uploader reads them from the environment. Either an API token from the platform's Preferences page:

```bash
export ACCOUNT_ID=... AUTH_TOKEN=...
export MAT3RA_HOST=alphafilm.mat3ra.com # or pass --host
```

or, if you signed in through the browser, `OIDC_ACCESS_TOKEN`.

## Parse

Each parser reads one lab's delivery. Point it at the folder as delivered and give it the identifier written
on the physical piece — every Sample carries it, and it is how the piece is found again later.

```bash
# UTK: an Asylum SPM run folder (summary.json or recipe.json + records/ + loops/)
python parse_utk.py ~/data/From_UTK --physical-id PDAC_COM5_01448 --out parsed

# NLR: an XRF map of the bare film on a grid, and a DC I-V sweep over the Pt pads patterned afterwards
python parse_nlr.py ~/data/From_NLR --physical-id PDAC_COM5_01448 \
--xrf-instrument instrument-2 --iv-instrument instrument-3 --out parsed
```

`parsed/` now holds a run document per run, plus any file a parser derived. Read it — it is the whole upload,
in JSON, before anything is sent. `run_document.py` states the shape.

## Upload

```bash
python upload_run.py parsed/*.json --account <your account slug> --files records
```

- `--files` picks which file groups go up: `records` (the per-measurement JSONs, the delivered tables, the
photographs) and `loops` (the raw arrays and plots — thousands of files, tens of minutes). `--files` with no
value uploads none and keeps the raw records in each measurement's metadata instead.
- `--dry-run` validates the documents against the ESSE schemas and stops.
- Re-running is safe: sets are found by name, samples and measurements by label, properties by what they belong
to. A second pass creates nothing.

Then open the platform's Measurements tab: the run is there as a set, one measurement per sample.

## Adding a lab

Write a parser. It reads whatever that lab ships and returns the dict `run_document.py` describes — one sample
per measured position, one measurement per sample, the properties each measurement produced. Take the
measurement's workflow from standata (`standata_workflow(application, name)`); if the instrument is not in that
registry yet, add it there — `mat3ra/standata`, `assets/applications/` and `assets/workflows/` — rather than
building a workflow in Python, so the platform resolves it the same way it resolves a job's.

Nothing in `upload_run.py` changes.

## The notebooks

`upload_spm_run.ipynb` and `upload_nlr_data.ipynb` run the same three steps with the same code, for people who
would rather not use a terminal. They install the pins themselves; restart the kernel after that cell if either
package was already imported.
83 changes: 83 additions & 0 deletions examples/measurement/derive_topography.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
#!/usr/bin/env python3
"""AFM topography metrics already stored on the platform -> the properties they are.

UTK's uploader kept each site's image metrics in the measurement's metadata (`frames[].imageMetrics`)
under names of its own. This reads them back and posts the ESSE properties they correspond to, so the
Results tab shows them and they are comparable with any other instrument's. Only the unambiguous
metrics are mapped; peak-to-valley, kurtosis and correlation length wait for UTK to say how they were
computed. Idempotent: a property already present on a measurement is not posted again.

derive_topography.py <measurement set id> [--dry-run]
"""
import argparse, os, sys, urllib.parse

from mat3ra.api_client import APIClient

from upload_run import account_id, base_url, find, holder

# imageMetrics key -> the ESSE property it is. Metres in, metres out.
METRICS = {
"rq_m": {"name": "areal_surface_texture", "parameter": "Sq", "units": "m"},
"ra_m": {"name": "areal_surface_texture", "parameter": "Sa", "units": "m"},
"grain_radius_median_m": {"name": "grain_size", "statistic": "median", "units": "m"},
"grain_radius_mean_m": {"name": "grain_size", "statistic": "mean", "units": "m"},
"grain_radius_std_m": {"name": "grain_size", "statistic": "std", "units": "m"},
"grain_radius_iqr_m": {"name": "grain_size", "statistic": "iqr", "units": "m"},
"grain_coverage": {"name": "grain_coverage"},
}


def properties_of(measurement):
"""The properties one topography measurement's stored metrics stand for."""
unit_id = measurement["workflow"]["subworkflows"][0]["units"][0]["flowchartId"]
out = []
for frame in (measurement.get("metadata") or {}).get("frames", []):
for key, shape in METRICS.items():
if key in frame.get("imageMetrics", {}):
out.append((unit_id, dict(shape, value=frame["imageMetrics"][key])))
return out


def derive(client, set_id, dry_run):
owner_id = client.my_account.id
measurements, skip = [], 0
while True: # the list route pages at 20 whatever the limit
page = client.measurements.list({"inSet._id": set_id, "isEntitySet": {"$ne": True}, "owner._id": owner_id},
{"limit": 20, "skip": skip})
measurements += page
skip += 20
if len(page) < 20:
break
posted = present = 0
for m in measurements:
for unit_id, prop in properties_of(m):
selector = {"source.info.origin._id": m["_id"], "data.name": prop["name"]}
for key in ("parameter", "statistic"):
if key in prop:
selector[f"data.{key}"] = prop[key]
if find(client.properties, selector, owner_id, 1):
present += 1
continue
if not dry_run:
client.properties.create(dict(holder(prop, m["_id"], m["_sample"]["_id"], unit_id, 0), owner={"_id": owner_id}))
posted += 1
print(f"{len(measurements)} measurements: {posted} properties {'to post' if dry_run else 'posted'}, {present} already present")


def main():
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("set_id", help="the topography Measurement Set")
ap.add_argument("--host", default=os.environ.get("MAT3RA_HOST", "localhost:3000"))
ap.add_argument("--account", help="slug of the account the data belongs to")
ap.add_argument("--dry-run", action="store_true", help="count what would be posted and stop")
a = ap.parse_args()
url = urllib.parse.urlsplit(base_url(a.host))
address = {"host": url.hostname, "port": url.port or (443 if url.scheme == "https" else 80), "secure": url.scheme == "https"}
client = APIClient.authenticate(**address)
if a.account:
client = APIClient.authenticate(account_id=account_id(client, a.account), **address)
derive(client, a.set_id, a.dry_run)


if __name__ == "__main__":
sys.exit(main())
144 changes: 144 additions & 0 deletions examples/measurement/parse_nlr.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,144 @@
"""NLR's delivery for one piece as platform documents: the Library (the piece itself, with its layout and
synthesis), the XRF map as a Sample Set of grid points, and the DC I-V sweep as a Sample Set of pads.

Ad hoc parser for SOF-8050: it reads the tab-separated files NLR ships and nothing else.
"""
import argparse
import json
from pathlib import Path

from mat3ra.standata.workflows import WorkflowStandata

from run_document import serialize


def standata_workflow(application_name, workflow_name):
"""The procedure the instrument runs, from the standata registry — the entry the platform resolves
a job's workflow through."""
workflow = WorkflowStandata.find_by_application_and_name(application_name, workflow_name)
if workflow is None:
raise LookupError(f"standata has no '{workflow_name}' workflow for {application_name}")
return workflow


def unit_id(workflow):
"""The execution unit a property of this workflow comes from."""
return workflow["subworkflows"][0]["units"][0]["flowchartId"]

FRAME = {"origin": "wafer corner", "axes": "x, y", "units": "mm"}
DIMENSIONS = {"shape": "square", "side": 50.8, "units": "mm"} # a 2-inch substrate
IV_COLUMNS = 11 # the I-V file lists its pads row by row, eleven to a row
XRF_APPLICATION = "xrf-mapper" # the standata application whose workflow this run records


def read_columns(path):
"""Every line of a tab-separated file after its header, split into its cells."""
return [line.split("\t") for line in Path(path).read_text().splitlines()[1:] if line.strip()]


IV_APPLICATION = "probe-station"


def parse_nlr(folder, physical_id, xrf_instrument, iv_instrument, description="", deposition=()):
"""NLR's delivery for one piece as two run documents sharing one Library: the XRF map, measured on the bare
film at the grid points, and the DC I-V sweep, measured on the Pt pads patterned afterwards. The pads are the
ones NLR probed: the I-V file lists them row by row, IV_COLUMNS to a row, so each pad gets its row and
column; where each pad sits on the wafer, and its size, come with NLR's pattern and are written onto the
same pads by label."""
folder = Path(folder)
grid_file = sorted(folder.rglob("*xrf_grid.txt"))[0]
volts_file, amps_file = sorted(folder.rglob("IV_Volts.txt"))[0], sorted(folder.rglob("IV_Amps.txt"))[0]
run_name = grid_file.stem
images = [] # the photograph is of the piece, not of the XRF grid: it goes on the Library page
sample_set = {"name": run_name, "entitySetType": "ordered", "metadata": {}}
xrf_run_name = f"{run_name} XRF"
xrf_workflow = standata_workflow(XRF_APPLICATION, "XRF Grid Map")
xrf_unit_id = unit_id(xrf_workflow)
grid = read_columns(grid_file)
synthesis = []
for f in deposition: # a file may hold one record or a list of them
record = json.loads(Path(f).read_text())
synthesis.extend(record if isinstance(record, list) else [record])
library = {"physicalId": physical_id, "name": physical_id, "description": description,
"entitySetType": "unordered",
"metadata": {"dimensions": DIMENSIONS, "frame": FRAME, "synthesis": synthesis}}
samples, xrf_measurements, xrf_properties = {}, {}, []
for row, column, x_mm, y_mm, thickness_um, aluminium_at_pct, scandium_at_pct in grid:
label = f"r{int(row)}c{int(column)}"
# two rows for one pad would overwrite each other here and leave the row counts below
# agreeing against a dictionary that has already lost an entry
if label in samples:
raise SystemExit(f"{grid_file.name}: pad {label} appears twice")
samples[label] = {"name": f"{physical_id} {label}", "label": label, "physicalId": physical_id,
"position": {"coordinates": [float(x_mm), float(y_mm)], "units": "mm"},
"metadata": {"row": int(row), "column": int(column)}}
xrf_measurements[label] = {"name": f"{xrf_run_name} {label}", "_sample": None, "workflow": xrf_workflow,
"setup": {"name": xrf_instrument}, "status": "finished", "_records": [],
"metadata": {"row": int(row), "column": int(column), "thickness_um": float(thickness_um),
"al_at_pct": float(aluminium_at_pct), "sc_at_pct": float(scandium_at_pct)}}
# NLR's columns become what ESSE already defines: composition is one elemental_ratio per
# element, a fraction, not a property named after the element
xrf_properties += [(label, xrf_unit_id, {"name": "film_thickness", "value": float(thickness_um), "units": "um"}, 0),
(label, xrf_unit_id, {"name": "elemental_ratio", "element": "Al", "value": float(aluminium_at_pct) / 100}, 0),
(label, xrf_unit_id, {"name": "elemental_ratio", "element": "Sc", "value": float(scandium_at_pct) / 100}, 0)]
iv_run_name = f"{physical_id} DC IV"
iv_workflow = standata_workflow(IV_APPLICATION, "DC I-V Sweep")
iv_unit_id = unit_id(iv_workflow)
volts = [[float(v) for v in cells] for cells in read_columns(volts_file)]
amps = [[float(a) for a in cells] for cells in read_columns(amps_file)]
if len(volts) != len(amps):
raise SystemExit(f"{volts_file.name}/{amps_file.name}: {len(volts)} and {len(amps)} rows")
if len(volts) % IV_COLUMNS:
raise SystemExit(f"{len(volts)} I-V rows do not fill rows of {IV_COLUMNS} pads")
for row, (bias_row, current_row) in enumerate(zip(volts, amps)):
if len(bias_row) != len(current_row):
raise SystemExit(f"row {row}: {len(bias_row)} bias points but {len(current_row)} current points")
# the pads NLR probed, by their row and column in the file; positions and sizes come with NLR's pattern
layout = [{"label": f"pad_r{i // IV_COLUMNS}c{i % IV_COLUMNS:02d}", "row": i // IV_COLUMNS, "column": i % IV_COLUMNS,
"position": None, "extent": None, "stack": None} for i in range(len(volts))]
library["metadata"]["layout"] = layout
iv_setup = {"name": iv_instrument, "settings": {"v_min": min(volts[0]), "v_max": max(volts[0]), "points": len(volts[0])}}
pads, iv_measurements, iv_properties = {}, {}, []
for index, (pad, bias, current) in enumerate(zip(layout, volts, amps)):
label = pad["label"]
pads[label] = {"name": f"{physical_id} {label}", "label": label, "physicalId": physical_id,
"metadata": {"site": "pad", "row": pad["row"], "column": pad["column"]}}
iv_measurements[label] = {"name": f"{iv_run_name} {label}", "_sample": None, "workflow": iv_workflow,
"setup": iv_setup, "status": "finished", "_records": [], "metadata": {"row_index": index}}
iv_properties.append((label, iv_unit_id, {"name": "current_voltage_curve", "xAxis": {"label": "voltage", "units": "V"},
"yAxis": {"label": "current", "units": "A"},
"xDataArray": bias, "yDataSeries": [current]}, 0))
return [{"physicalId": physical_id, "library": library, "run": xrf_run_name, "sample_set": sample_set, "images": images,
"samples": samples, "measurement_set": {"name": xrf_run_name, "entitySetType": "ordered", "metadata": {}},
"measurements": xrf_measurements, "files": {}, "set_files": [(grid_file.name, grid_file)],
"records": grid, "properties": xrf_properties},
{"physicalId": physical_id, "library": library, "run": iv_run_name,
"sample_set": {"name": f"{physical_id} pads", "entitySetType": "ordered", "metadata": {}}, "images": [],
"samples": pads, "measurement_set": {"name": iv_run_name, "entitySetType": "ordered", "metadata": {}},
"measurements": iv_measurements, "files": {},
"set_files": [(volts_file.name, volts_file), (amps_file.name, amps_file)],
"records": volts, "properties": iv_properties}]



def main():
"""Read NLR's delivery and write its run document."""
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("folder")
ap.add_argument("--physical-id", required=True, help="the identifier written on the physical piece, e.g. PDAC_COM5_01448")
ap.add_argument("--xrf-instrument", required=True, help="identity of the machine the grid was mapped on")
ap.add_argument("--iv-instrument", required=True, help="identity of the machine the sweep was measured on")
ap.add_argument("--description", default="", help="what the piece is, for the Library")
ap.add_argument("--deposition", nargs="*", default=[], metavar="JSON", help="deposition record(s) for the Library's synthesis")
ap.add_argument("--out", default="parsed", help="directory for the run documents (default: parsed/)")
a = ap.parse_args()

for parsed in parse_nlr(a.folder, a.physical_id, a.xrf_instrument, a.iv_instrument, a.description, a.deposition):
print(f"{parsed['physicalId']}: {len(parsed['samples'])} samples · run {parsed['run']}: "
f"{len(parsed['measurements'])} measurements · {len(parsed['records'])} rows -> "
f"{len(parsed['set_files'])} files · {len(parsed['images'])} image(s) · {len(parsed['properties'])} properties")
print("run document:", serialize(parsed, a.out, name=parsed["run"].replace(" ", "_") + ".json"))


if __name__ == "__main__":
main()
Loading
Loading