Skip to content

Partial sweep= open renames groups by position, mislabelling elevations #397

Description

@aladinor

What happens

Opening a subset of sweeps names the returned groups by position, not by sweep identity. Requesting sweep=[0, 2, 8] returns sweep_0, sweep_1, sweep_2 — so the 25° tilt arrives named sweep_2, whose elevation in the full volume is 1.0°.

_attach_sweep_groups names from the enumeration index:

# xradar/io/backends/common.py:99-101
for i, sw in enumerate(sweeps):
    sw = sw.drop_vars(_STATION_VARS, errors="ignore").drop_attrs(deep=False)
    dtree[f"sweep_{i}"] = xr.DataTree(sw)

The sweep docstring says only "Sweep number(s) to extract" — nothing about renaming. The data is fine; each sweep still carries its true sweep_number, only the group name is wrong.

This matters because it fails silently. Writing a subset into an elevation-indexed schema (a fixed sweep→elevation template) files measurements at the wrong tilt with no error.

Reproducer

Dependencies inline (PEP 723) — save as mre.py, run uv run mre.py. Reads one public IRIS volume from the IDEAM open-data bucket anonymously; no malformed input involved — the same file opens perfectly with a full open.

# /// script
# requires-python = ">=3.11"
# dependencies = [
#     "xradar",
#     "s3fs",
# ]
# ///
import fsspec
import xradar
from xradar.io import open_iris_datatree

KEY = "s3-radaresideam/l2_data/2023/02/07/Tablazo/TAB230207000003.RAW3D38"
data = fsspec.filesystem("s3", anon=True).cat(KEY)

print(f"xradar {xradar.__version__}\n")


def show(tree, title):
    print(title)
    names = sorted(
        (k for k in tree.children if k.startswith("sweep_")),
        key=lambda s: int(s.split("_")[1]),
    )
    for name in names:
        ds = tree[name].to_dataset()
        print(
            f"  {name:9s} sweep_fixed_angle={float(ds['sweep_fixed_angle'].values):5.1f}"
            f"   sweep_number={int(ds['sweep_number'].values)}"
        )
    print()
    return {n: float(tree[n].to_dataset()["sweep_fixed_angle"].values) for n in names}


truth = show(open_iris_datatree(data), "full open (reference):")
subset = show(open_iris_datatree(data, sweep=[0, 2, 8]), "sweep=[0, 2, 8]:")

print("sweep_N name -> elevation, full open vs subset:")
for name in subset:
    ok = "ok" if truth[name] == subset[name] else "MISLABELLED"
    print(f"  {name:9s} {truth[name]:5.1f}  vs {subset[name]:5.1f}   {ok}")

Output on xradar 0.12.0:

xradar 0.12.0

full open (reference):
  sweep_0   sweep_fixed_angle=  0.0   sweep_number=0
  sweep_1   sweep_fixed_angle=  0.5   sweep_number=1
  sweep_2   sweep_fixed_angle=  1.0   sweep_number=2
  sweep_3   sweep_fixed_angle=  2.0   sweep_number=3
  sweep_4   sweep_fixed_angle=  4.5   sweep_number=4
  sweep_5   sweep_fixed_angle=  7.0   sweep_number=5
  sweep_6   sweep_fixed_angle= 10.0   sweep_number=6
  sweep_7   sweep_fixed_angle= 15.0   sweep_number=7
  sweep_8   sweep_fixed_angle= 25.0   sweep_number=8

sweep=[0, 2, 8]:
  sweep_0   sweep_fixed_angle=  0.0   sweep_number=0
  sweep_1   sweep_fixed_angle=  1.0   sweep_number=2
  sweep_2   sweep_fixed_angle= 25.0   sweep_number=8

sweep_N name -> elevation, full open vs subset:
  sweep_0     0.0  vs   0.0   ok
  sweep_1     0.5  vs   1.0   MISLABELLED
  sweep_2     1.0  vs  25.0   MISLABELLED

Expected

Group names preserve sweep identity — sweep=[0, 2, 8] yields sweep_0, sweep_2, sweep_8 — or, if positional naming is intentional, the docstring should say so, since the current behaviour is indistinguishable from correct output until someone checks sweep_number.

Secondary (same code path, lower priority)

A truncated sweep whose times fall back to cftime (dates before the 1582 reform) makes root assembly a mixed-type reduction:

# xradar/io/backends/common.py:139  in _assign_root
time_coverage_start = min(times)
TypeError: '>' not supported between instances of 'int' and 'cftime._cftime.DatetimeGregorian'

The input really is malformed there, so raising is defensible — but the message names neither the sweep nor the cause, and IrisArrayWrapper.__init__ (iris.py:3767) has already warned sweep_N empty or corrupted for that very sweep. Surfacing that instead would be a big diagnostic win. Encountered on 55 of 9,994 IDEAM Tablazo volumes, e.g. s3-radaresideam/l2_data/2023/02/19/Tablazo/TAB230219165503.RAW4WHB, whose sweep_7 holds one ray stamped year 0001 while sweeps 0-6 and 8 are clean.

Working around that one by re-opening with only the good sweeps is what surfaced the naming bug above.

Versions

xradar 0.12.0 (also 0.12.1.dev9+gd723857f2), Python 3.12, Linux.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions