Flat metadata namespace and semantic owners#
PhenoTypic writes every generated scientific metadata field as
Metadata_<Label>. The prefix is short and predictable for people entering CSV
headers, while the public enum that owns a label keeps its semantic category
available to code:
Public enum |
Example member |
Emitted header |
REMBI module |
|---|---|---|---|
|
|
|
ImageData |
|
|
|
Study |
|
|
|
Study |
|
|
|
Biosample |
|
|
|
Biosample |
|
|
|
SpecimenPreparation |
|
|
|
Biosample |
|
|
|
SpecimenPreparation |
|
|
|
ImageAcquisition |
All nine enums inherit phenotypic.schema.MetadataInfo, whose
category() is "Metadata". Category strings therefore cannot distinguish
owners. Use enum identity and the ownership APIs instead:
from phenotypic.schema import GENETIC, MetadataInfo
from phenotypic.sdk_ import (
metadata_member_for_header,
metadata_owner_for_header,
)
member = metadata_member_for_header("Metadata_Strain")
assert member is GENETIC.STRAIN
assert metadata_owner_for_header("Metadata_Strain") is GENETIC
assert "Metadata_Strain" in GENETIC.header_set()
assert issubclass(metadata_owner_for_header("Strain"), MetadataInfo)
The lookup functions also accept a bare known label such as Strain and an
exact historical spelling such as MetadataGenetic_Strain. Corresponding
metadata_member_for_label() and metadata_owner_for_label() helpers are
available when the input is conceptually a label. Unknown metadata remains
valid but has no owner and routes to REMBI Uncategorized.
Use phenotypic.sdk_.is_metadata_header() to identify the metadata family.
It recognizes canonical Metadata_* headers and the finite exact set of
historical headers. It intentionally rejects arbitrary lookalikes such as
MetadataFoo_Bar. Use phenotypic.sdk_.normalize_metadata_columns() at a
DataFrame ingress boundary. It returns a copy, canonicalizes bare and historical
headers, and coalesces duplicate aliases only when their dtypes are compatible
and their overlapping non-null values agree.
See also
To attach metadata to a run and emit deliverables/rembi.yaml, see
Describe a run’s metadata with REMBI.
Column order#
Measurement sheets order the metadata front block by a bench-scientist narrative, not by spelling or REMBI module:
Identity: fields owned by
SAMPLE, thenPLATEStrain: fields owned by
GENETICCondition: fields owned by
CONDITION, thenCULTUREDesign and provenance: fields owned by
EXPERIMENT,STUDY, thenACQUISITION
Unknown Metadata_* fields trail the known owner blocks. Within an owner,
columns follow enum declaration order. The IMAGE block is per-image
provenance and appears after measurements, before the per-object info block.
REMBI classification is a separate provenance axis and does not drive this
presentation order.
Compatibility and migration#
Stored-header compatibility is permanent. Ordinary reads normalize exact old
headers such as MetadataImage_ImageName and MetadataCulture_Time in memory
without rewriting their source. The previous Python enum class names remain as
one-transition-release aliases: importing one emits DeprecationWarning, and
the alias is absent from phenotypic.schema.__all__ and schema discovery. New
code should use IMAGE, CULTURE, and the other canonical owner names.
The CLI’s deliverables/metadata.csv is a byte-exact startup snapshot of the
configured input CSV. It is immutable provenance rather than a generated schema
table, so it may retain historical headers. Readers normalize it in memory;
generated measurement, analysis, QC, and REMBI outputs remain canonical.
--mode migrate is the sole metadata migration boundary for a run bundle. It
preflights bundle-owned authoritative legacy data, refuses blocked plans, and
applies only the fingerprinted plan. --mode recompile performs neither that
preflight nor migration. Legacy storage and external per-image measurement
authority must be migrated explicitly before recompile; readable historical
header spellings continue to normalize in memory.
Migration uses source and plan fingerprints plus a prepared/applied receipt
journal. CSV, parquet, typed pipeline JSON, and HDF metadata attributes are
replaced atomically. HDF migration writes and validates a sibling copy, changes
only metadata attributes and the metadata-schema marker, and leaves the HDF
layout schema_version, arrays, grid state, and unrelated attributes alone.
Receipts make an interrupted bundle migration resumable and can be passed to
rollback_metadata_migration().
An external file supplied with --metadata is always read-only during
recompile: PhenoTypic normalizes its contents in memory without writing to the
source, and copies its original bytes to the immutable bundle-owned
deliverables/metadata.csv provenance snapshot. To mutate a standalone
external file, call migrate_metadata_file() explicitly.
from pathlib import Path
from phenotypic.sdk_ import (
migrate_metadata_bundle,
migrate_metadata_file,
preflight_metadata_schema,
rollback_metadata_migration,
)
source = Path("metadata.csv")
report = preflight_metadata_schema(source)
result = migrate_metadata_file(
source,
expected_source_fingerprint=report.source_fingerprint,
)
bundle_report = preflight_metadata_schema(Path("out"))
bundle_result = migrate_metadata_bundle(
Path("out"),
expected_plan_fingerprint=bundle_report.plan_fingerprint,
)
if bundle_result.receipt_path is not None:
rollback_metadata_migration(bundle_result.receipt_path)
Always inspect status, conflicts, and the proposed header maps in a
preflight report before invoking standalone migration. A blocked report never
mutates its target.