Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 12 additions & 6 deletions packages/zarr-metadata/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ Two layers and an optional integration:
specifications, plus types for [`zarr-extensions`](https://github.com/zarr-developers/zarr-extensions/)
and a few widely-used-but-unspecified entities (e.g. consolidated metadata).
- **Document models** (`zarr_metadata.model`): canonical frozen-dataclass
models of whole metadata documents, with structural validators, loc-aware
models of whole metadata documents, with validators, loc-aware
parsers, and store-key (de)serialization. A document produced by `to_json`
shares no mutable state with the model that produced it.
- **Optional Pydantic integration** (`zarr_metadata.pydantic`, requires
Expand Down Expand Up @@ -57,10 +57,16 @@ members that the strict model parser rejects.
The model validators enforce the declared document structure and a small set
of context-free consistency rules, including fixed format literals, finite
JSON numbers, non-negative dimensions, non-empty v3 codec pipelines, and one
`dimension_names` entry per array dimension. They do not interpret extension
names or configurations, resolve codec pipelines, or decide whether a data
type, chunk grid, codec, or storage transformer is supported. Those decisions
belong to consumer implementations.
`dimension_names` entry per array dimension. In a v3 document they also read
each extension point -- the data type, chunk grid, chunk key encoding, each
codec and each storage transformer -- through the definition that claims its
name in a scope, `CORE_AND_EXTENSIONS` unless a `context` is passed: a
configuration its definition refuses is refused, and a key it does not
declare is reported as `unknown_key`. A name nothing in the scope claims is
left unjudged, and whether to support it is the consumer's decision. The
validators do not judge fields against each other: a fill value against its
data type, a codec against the array it is handed, a chunk grid against the
shape.

The Pydantic integration's generated JSON Schemas express independently
checkable document structure and field constraints, but they are not a
Expand All @@ -75,7 +81,7 @@ versus `shape`. Consumers should run the model parser after schema validation.
At minimum, this library supports what Zarr-Python needs: the complete
Zarr v2 and v3 specs, consolidated metadata, and a subset of the metadata
defined in `zarr-extensions`. We are generally open to contributions that
add types, models, or structural validation for Zarr metadata with a
add types, models, or validation for Zarr metadata with a
published spec.

Runtime array behavior is out of scope: nothing here encodes or decodes
Expand Down
13 changes: 13 additions & 0 deletions packages/zarr-metadata/changes/4436.feature.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
**Breaking:** the v3 validators read each extension point through its
definition: `validate_array_metadata_v3` judges the data type, the chunk
grid, the chunk key encoding, each codec and each storage transformer
against the definition that claims its name in a scope,
`CORE_AND_EXTENSIONS` unless a `context` is passed, so a gzip `level` of
99, or a key a codec does not declare, is a problem located in the
configuration, and a document holding one, accepted before, is refused.
So do the `is_*` and `parse_*` beside it and the v3 group validators,
with each array an inline `consolidated_metadata` holds, and the v3
model classes' `from_json`, `from_key_value` and `to_key_value` take the
same `context`, so a model read in a scope is written in it. The
pydantic types read in `CORE_AND_EXTENSIONS`. A name nothing in scope
claims is left unjudged.
2 changes: 1 addition & 1 deletion packages/zarr-metadata/docs/api/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ title: API reference
The package is organized to mirror the structure of the Zarr specifications:

- [`zarr_metadata.model`](model.md) — frozen-dataclass document models,
structural validators, loc-aware parsers, and the `UNSET` sentinel
validators, loc-aware parsers, and the `UNSET` sentinel
- [`zarr_metadata.pydantic`](pydantic.md) — optional Pydantic field types
over the models
- [`zarr_metadata.typed_json`](typed_json.md) — `check`, which type-checks
Expand Down
18 changes: 12 additions & 6 deletions packages/zarr-metadata/docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ closely model the content of the Zarr specifications, such as:
[zarr-extensions](https://github.com/zarr-developers/zarr-extensions/) and a
few widely-used-but-unspecified entities (e.g. consolidated metadata).
- **Document models** ([`zarr_metadata.model`](api/model.md)): canonical
frozen-dataclass models of whole metadata documents, with structural
frozen-dataclass models of whole metadata documents, with
validators, loc-aware parsers, and store-key (de)serialization. A document
produced by `to_json` shares no mutable state with the model that produced
it.
Expand Down Expand Up @@ -72,17 +72,23 @@ members that the strict model parser rejects.
The model validators enforce the declared document structure and a small set
of context-free consistency rules, including fixed format literals, finite
JSON numbers, non-negative dimensions, non-empty v3 codec pipelines, and one
`dimension_names` entry per array dimension. They do not interpret extension
names or configurations, resolve codec pipelines, or decide whether a data
type, chunk grid, codec, or storage transformer is supported. Those decisions
belong to consumer implementations.
`dimension_names` entry per array dimension. In a v3 document they also read
each extension point -- the data type, chunk grid, chunk key encoding, each
codec and each storage transformer -- through the definition that claims its
name in a scope, `CORE_AND_EXTENSIONS` unless a `context` is passed: a
configuration its definition refuses is refused, and a key it does not
declare is reported as `unknown_key`. A name nothing in the scope claims is
left unjudged, and whether to support it is the consumer's decision. The
validators do not judge fields against each other: a fill value against its
data type, a codec against the array it is handed, a chunk grid against the
shape.

## Scope

At minimum, this library supports what Zarr-Python needs: the complete
Zarr v2 and v3 specs, consolidated metadata, and a subset of the metadata
defined in `zarr-extensions`. We are generally open to contributions that
add types, models, or structural validation for Zarr metadata with a
add types, models, or validation for Zarr metadata with a
published spec.

Runtime array behavior is out of scope: nothing here encodes or decodes
Expand Down
19 changes: 11 additions & 8 deletions packages/zarr-metadata/src/zarr_metadata/model/__init__.py
Original file line number Diff line number Diff line change
@@ -1,14 +1,17 @@
"""In-memory models for Zarr metadata documents.

Models are frozen dataclasses that hold a canonical, semantically lossless
representation of the JSON documents; they never interpret extension points
(codecs, chunk grids, data types). Validators check JSON structure, not domain validity.
Each document concept gets a `validate_*` function returning every problem
found (a `list[ValidationProblem]`, each with a machine-readable `kind`), an
`is_*` type guard, and a `parse_*` function that narrows or raises
`MetadataValidationError`. Model `from_json` / `from_key_value` constructors
raise `MetadataValidationError` for every ingestion failure, including
missing store keys and undecodable bytes.
representation of the JSON documents. Validators check a document's JSON
structure and, in a v3 document, read each extension point (codecs, chunk
grids, data types, ...) through the definition that claims its name in a
scope, `CORE_AND_EXTENSIONS` unless a `context` is passed; they do not judge
fields against each other. Each document concept gets a `validate_*`
function returning every problem found (a tuple of `ValidationProblem`, each
with a machine-readable `kind`), an `is_*` type guard, and a `parse_*`
function that narrows or raises `MetadataValidationError`. Model
`from_json` / `from_key_value` constructors raise `MetadataValidationError`
for every ingestion failure, including missing store keys and undecodable
bytes, and the v3 ones take the same `context`.
"""

from zarr_metadata._json import (
Expand Down
30 changes: 21 additions & 9 deletions packages/zarr-metadata/src/zarr_metadata/model/_array.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@
from zarr_metadata.v2.array import ZARR_V2_ARRAY_METADATA_STORE_KEY
from zarr_metadata.v2.attributes import ZARR_V2_ATTRIBUTES_STORE_KEY
from zarr_metadata.v3._common import parse_metadata_field_v3
from zarr_metadata.v3._registry import CORE_AND_EXTENSIONS
from zarr_metadata.v3.array import ZARR_V3_ARRAY_METADATA_STORE_KEY

if TYPE_CHECKING:
Expand All @@ -40,6 +41,7 @@
from zarr_metadata.v2.attributes import ZarrV2AttributesStoreKey
from zarr_metadata.v2.codec import ZarrV2CodecMetadata
from zarr_metadata.v3._common import ZarrV3MetadataFieldJSON
from zarr_metadata.v3._registry import Context
from zarr_metadata.v3.array import (
ZarrV3ArrayMetadataJSON,
ZarrV3ArrayMetadataStoreKey,
Expand Down Expand Up @@ -158,8 +160,11 @@ class ZarrV3ArrayMetadata:
content for an array. Extension points (`data_type`, `chunk_grid`,
`chunk_key_encoding`, `codecs`, `storage_transformers`) are held as
`ZarrV3MetadataField` values (currently `ZarrV3NamedConfig` name,
configuration, and obligation records) and are never interpreted;
`fill_value` is held verbatim in its JSON form. Equivalent extension
configuration, and obligation records). `from_json` and
`from_key_value` read each through the definition that claims its name
in a scope -- `CORE_AND_EXTENSIONS` unless a `context` is passed -- and
the model holds what they read as written; `fill_value` is held
verbatim in its JSON form. Equivalent extension
spellings normalize to shorthand strings when configuration is empty and
understanding is required.
"""
Expand Down Expand Up @@ -276,9 +281,11 @@ def to_json(self) -> ZarrV3ArrayMetadataJSON:
return out

@classmethod
def from_json(cls, data: object) -> ZarrV3ArrayMetadata:
def from_json(
cls, data: object, *, context: Context = CORE_AND_EXTENSIONS
) -> ZarrV3ArrayMetadata:
# A read model shares no mutable state with what it read.
parsed = copy.deepcopy(parse_array_metadata_v3(data))
parsed = copy.deepcopy(parse_array_metadata_v3(data, context=context))
extra_fields: dict[str, ZarrV3ExtensionField] = {
k: v for k, v in parsed.items() if k not in ARRAY_METADATA_STANDARD_KEYS_V3
}
Expand Down Expand Up @@ -310,15 +317,20 @@ def must_understand_fields(self) -> dict[str, ZarrV3ExtensionField]:
return must_understand_subset(self.extra_fields)

@classmethod
def from_key_value(cls, mapping: Mapping[StoreKey, bytes]) -> ZarrV3ArrayMetadata:
return cls.from_json(load_store_json(mapping, ZARR_V3_ARRAY_METADATA_STORE_KEY))
def from_key_value(
cls, mapping: Mapping[StoreKey, bytes], *, context: Context = CORE_AND_EXTENSIONS
) -> ZarrV3ArrayMetadata:
return cls.from_json(
load_store_json(mapping, ZARR_V3_ARRAY_METADATA_STORE_KEY), context=context
)

def to_key_value(
self, *, indent: int | str | None = None
self, *, indent: int | str | None = None, context: Context = CORE_AND_EXTENSIONS
) -> Mapping[ZarrV3ArrayMetadataStoreKey, bytes]:
# A model built by hand is not validated: its document is written only
# if it reads as `from_json` reads one, and every problem is raised.
document = parse_array_metadata_v3(self.to_json())
# if it reads as `from_json` reads one in `context`, and every problem
# is raised.
document = parse_array_metadata_v3(self.to_json(), context=context)
return {ZARR_V3_ARRAY_METADATA_STORE_KEY: dump_store_json(document, indent=indent)}


Expand Down
40 changes: 26 additions & 14 deletions packages/zarr-metadata/src/zarr_metadata/model/_group.py
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,7 @@
from zarr_metadata.v2.attributes import ZARR_V2_ATTRIBUTES_STORE_KEY
from zarr_metadata.v2.consolidated import ZARR_V2_CONSOLIDATED_METADATA_STORE_KEY
from zarr_metadata.v2.group import ZARR_V2_GROUP_METADATA_STORE_KEY
from zarr_metadata.v3._registry import CORE_AND_EXTENSIONS
from zarr_metadata.v3.consolidated import ZARR_V3_CONSOLIDATED_METADATA_KEY
from zarr_metadata.v3.group import ZARR_V3_GROUP_METADATA_STORE_KEY

Expand All @@ -42,6 +43,7 @@
from zarr_metadata.v2.attributes import ZarrV2AttributesStoreKey
from zarr_metadata.v2.consolidated import ZarrV2ConsolidatedMetadataStoreKey
from zarr_metadata.v2.group import ZarrV2GroupMetadataJSON, ZarrV2GroupMetadataStoreKey
from zarr_metadata.v3._registry import Context
from zarr_metadata.v3.array import ZarrV3ExtensionField
from zarr_metadata.v3.consolidated import ZarrV3ConsolidatedMetadataJSON
from zarr_metadata.v3.group import ZarrV3GroupMetadataJSON, ZarrV3GroupMetadataStoreKey
Expand Down Expand Up @@ -72,8 +74,9 @@ class ZarrV3GroupMetadata:

A canonical, semantically lossless representation of the `zarr.json`
content for a group. The `consolidated_metadata` reference-implementation
convention is modeled as a typed field holding thin child models; every
other unknown top-level key lands in `extra_fields` verbatim.
convention is modeled as a typed field holding thin child models, each
array read in the scope the group is read in; every other unknown
top-level key lands in `extra_fields` verbatim.
"""

zarr_format: Literal[3] = field(default=3, init=False)
Expand Down Expand Up @@ -135,9 +138,11 @@ def to_json(self) -> ZarrV3GroupMetadataJSON:
return out

@classmethod
def from_json(cls, data: object) -> ZarrV3GroupMetadata:
def from_json(
cls, data: object, *, context: Context = CORE_AND_EXTENSIONS
) -> ZarrV3GroupMetadata:
# A read model shares no mutable state with what it read.
parsed = copy.deepcopy(parse_group_metadata_v3(data))
parsed = copy.deepcopy(parse_group_metadata_v3(data, context=context))
consolidated_raw: object = parsed.get(ZARR_V3_CONSOLIDATED_METADATA_KEY, UNSET)
consolidated: ZarrV3ConsolidatedMetadata | UNSET
if consolidated_raw is UNSET or consolidated_raw is None:
Expand All @@ -146,7 +151,7 @@ def from_json(cls, data: object) -> ZarrV3GroupMetadata:
# absence and never written back — repaired, not preserved.
consolidated = UNSET
else:
consolidated = ZarrV3ConsolidatedMetadata.from_json(consolidated_raw)
consolidated = ZarrV3ConsolidatedMetadata.from_json(consolidated_raw, context=context)
extra_fields: dict[str, ZarrV3ExtensionField] = {
k: v
for k, v in parsed.items()
Expand All @@ -171,15 +176,20 @@ def must_understand_fields(self) -> dict[str, ZarrV3ExtensionField]:
return must_understand_subset(self.extra_fields)

@classmethod
def from_key_value(cls, mapping: Mapping[StoreKey, bytes]) -> ZarrV3GroupMetadata:
return cls.from_json(load_store_json(mapping, ZARR_V3_GROUP_METADATA_STORE_KEY))
def from_key_value(
cls, mapping: Mapping[StoreKey, bytes], *, context: Context = CORE_AND_EXTENSIONS
) -> ZarrV3GroupMetadata:
return cls.from_json(
load_store_json(mapping, ZARR_V3_GROUP_METADATA_STORE_KEY), context=context
)

def to_key_value(
self, *, indent: int | str | None = None
self, *, indent: int | str | None = None, context: Context = CORE_AND_EXTENSIONS
) -> Mapping[ZarrV3GroupMetadataStoreKey, bytes]:
# A model built by hand is not validated: its document is written only
# if it reads as `from_json` reads one, and every problem is raised.
document = parse_group_metadata_v3(self.to_json())
# if it reads as `from_json` reads one in `context`, and every problem
# is raised.
document = parse_group_metadata_v3(self.to_json(), context=context)
return {ZARR_V3_GROUP_METADATA_STORE_KEY: dump_store_json(document, indent=indent)}


Expand Down Expand Up @@ -221,19 +231,21 @@ def to_json(self) -> ZarrV3ConsolidatedMetadataJSON:
}

@classmethod
def from_json(cls, data: object) -> ZarrV3ConsolidatedMetadata:
def from_json(
cls, data: object, *, context: Context = CORE_AND_EXTENSIONS
) -> ZarrV3ConsolidatedMetadata:
normalized = arrays_to_tuples(data)
problems = validate_consolidated_metadata_v3(normalized)
problems = validate_consolidated_metadata_v3(normalized, context=context)
if len(problems) != 0:
raise MetadataValidationError(problems)
env = cast("Mapping[str, object]", normalized)
entries: dict[str, ZarrV3ArrayMetadata | ZarrV3GroupMetadata] = {}
for key, entry in cast("Mapping[str, object]", env["metadata"]).items():
node_type = cast("Mapping[str, object]", entry).get("node_type")
if node_type == "array":
entries[key] = ZarrV3ArrayMetadata.from_json(entry)
entries[key] = ZarrV3ArrayMetadata.from_json(entry, context=context)
else:
entries[key] = ZarrV3GroupMetadata.from_json(entry)
entries[key] = ZarrV3GroupMetadata.from_json(entry, context=context)
return cls(metadata=entries)


Expand Down
Loading
Loading