Skip to content

ENH: virtual byte-range access to Sigmet/IRIS RAW — VirtualiZarr parser + zarr v3 codec #394

Description

@aladinor

What

Add virtual access to Vaisala Sigmet/IRIS RAW volumes: a SigmetParser (VirtualiZarr Parser) that indexes a RAW file into a ManifestStore of byte-range references (one chunk = one whole sweep), and a decode-only zarr v3 codec (sigmet-sweep) that demultiplexes one moment from the sweep's byte span at read time. No data is copied — reading yields the same xarray.DataTree as open_iris_datatree, straight from the original files (local or S3):

import xarray as xr
from obspec_utils.registry import ObjectStoreRegistry
from obstore.store import S3Store
from xradar.io.virtual import SigmetParser   # proposed home

# IDEAM's public bucket (anonymous access)
registry = ObjectStoreRegistry({
    "s3://s3-radaresideam/": S3Store(
        bucket="s3-radaresideam", region="us-east-1", skip_signature=True,
    )
})

# parse one IRIS RAW volume into a virtual (ManifestArray-backed) DataTree
vtree = SigmetParser()(
    url="s3://s3-radaresideam/l2_data/2026/01/15/Guaviare/GUA260115000109.RAWNBJT",
    registry=registry,
).to_virtual_datatree()

vtree["sweep_0"]["DBZH"]
# <xarray.DataArray 'DBZH' (azimuth: 360, range: 747)>
# ManifestArray<shape=(360, 747), dtype=uint8, chunks=(360, 747)>
# Attributes: units: dBZ, scale_factor: 0.5, add_offset: -32.0,
#             sigmet_data_type: DB_DBZ, ...

vtree["sweep_0"]["DBZH"].data.metadata.codecs
# [{'name': 'sigmet-sweep',
#   'configuration': {'moment_index': 1, 'ndatatypes': 12,
#                     'sort_rays': False, 'pad_missing_rays': False}}]

vtree["sweep_0"]["DBZH"].data.manifest.dict()    # {chunk_key: {path, offset, length}}
# {'0.0': {'path': 's3://s3-radaresideam/l2_data/2026/01/15/Guaviare/GUA260115000109.RAWNBJT',
#          'offset': 12288, 'length': 2070528}}

# one chunk = one whole sweep, and EVERY moment of the sweep references the
# SAME byte span — the sigmet-sweep codec demultiplexes its own moment_index
# out of the interleaved, RLE-compressed rays at read time:
vtree["sweep_0"]["ZDR"].data.manifest.dict() == vtree["sweep_0"]["DBZH"].data.manifest.dict()
# True
vtree["sweep_1"]["DBZH"].data.manifest.dict()["0.0"]
# {'path': 's3://.../GUA260115000109.RAWNBJT', 'offset': 2082816, 'length': 1929216}

# reading decodes on demand through the codec — same tree as open_iris_datatree:
dt = xr.open_datatree(SigmetParser()(url, registry), engine="zarr")

(Output above is real, from the file shown.)

A working implementation exists and is validated in production, building multi-year icechunk archives of Colombia's IDEAM radar network. It is entirely my own work and would be contributed under xradar's MIT license.

Why xradar

  • The codec decides who can read the data. A store of byte-range references is unreadable wherever the codec doesn't resolve. Registered as a zarr.codecs entry point in xradar, any environment with xradar + zarr can open these stores — no explicit import, and no virtualizarr needed on the read side.
  • It ends code duplication. The parser's byte offsets and mapping/scaling tables were generated from xradar's own iris.py structures (INGEST_HEADER, SIGMET_DATA_TYPES, iris_mapping). Living in xradar they become imports/derivations instead of copies — verified: all 17 header offsets reproduce from the struct dicts, and the CF scaling matches decode_array exactly.
  • xradar is the parity oracle. The test gate compares the virtual decode bin-for-bin against the eager IRIS backend; in-repo, it protects both implementations at once.

Validated on 21 real files — 6 IDEAM radars (all tasks, per-sweep fragments, same-timestamp companions) plus the two existing open-radar-data IRIS fixtures — bit-identical to the eager decode, including the all-16-bit + extended-header + missing-ray case (SUR210819000227.RAWKPJV).

How

  • New xradar/io/virtual/ subpackage: pure-numpy format walker, SigmetSweepCodec, SigmetParser, plus shared manifest helpers .
  • Codec ships in the base install (entry point sigmet-sweep; imports only zarr + numpy, and zarr does not become a runtime dependency — the module is only imported when zarr itself resolves the entry point). Parser sits behind a new xradar[virtual] extra (virtualizarr, which brings obstore/obspec-utils).
  • Two small PRs: (1) format walker + codec + entry point + tests; (2) parser + parity tests + a docs notebook.
  • Tests run on the existing fixtures (cor-main131125105503.RAW2049, SUR210819000227.RAWKPJV) — no new test data required.
  • One prerequisite: bump zarr>=2,<3 → zarr>=3 in requirements_dev.txt/environment.yml. The pin dates to the zarr 3.0.0 release window (Feb 2025); the repo's only zarr use is the to_zarr round-trip in one notebook, which runs unchanged on zarr v3.

Expected outcome

  • pip install xradar zarr reads any virtual Sigmet store — e.g. icechunk archives referencing the public IDEAM bucket (s3://s3-radaresideam) — at ~0.4% of the storage cost of a re-encoded copy.
  • One authoritative Sigmet decode in openradar, with a permanent virtual ↔ eager parity gate.
  • Optional follow-up: an open_iris_virtual_datatree(...) convenience mirroring open_iris_datatree.

Questions for maintainers

  1. Placement and naming OK? (xradar/io/virtual/, extra named virtual)
  2. Any objection to the zarr>=3 dev bump?
  3. Zarr codec names share one global namespace (stores reference codecs by name in their metadata), and the community records third-party names in zarr-developers/zarr-extensions to prevent collisions and document their configuration. Should we submit a short sigmet-sweep spec there as part of this effort?

CC @mgrover1, @kmuehlbauer, @syedhamidali

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions