Replies: 4 comments 3 replies
|
this is very cool work! I'd like to understand this point better:
am I right in inferring that, where xarray today can only express these coordinates by creating dummy arrays, your implementation here instead declares the same data via the JSON attributes? because that's a very welcome change. |
|
Thanks for the praise! The CF Conventions require that data variables (arrays) store their coordinates along each of their dimensions in coordinate variables (more arrays). This also applies to boundary values, additional coordinate sets, parametric vertical coordinates, scalar axes, etc. XArray implements that paradigm so CF-encoded netCDF files can be read in "as-is". The two principal conventions that I referenced do not need these additional arrays with coordinate values. The The {
"zarr_format": 3,
"node_type": "array",
"shape": [512, 256, 31046],
...
"dimension_names": ["lon", "lat", "time"],
"attributes": {
"cs": {
"name": "cs",
"crs": [ -- 3 types of CRS objects, follows OGC standard
{
"type": "planar", -- First CRS object: planar axes
"axes": {
"lon": {
"coordinates": [
{
"values": { -- lon is regular so just start and step get recorded
"regular": [0, 0.703125]
},
"unit": "degrees", -- A true unit, not the CF-style "degrees_east"
"boundaries": { -- Same with boundary values
"regular": [-0.3515625, 0.3515625]
}
}
],
"abbreviation": "X", -- OGC-style axis descriptors rather than CF "axis"
"direction": "EAST",
"attributes": { -- CF-style attributes are included
"standard_name": "longitude",
"long_name": "Longitude"
}
},
"lat": {
"coordinates": [
{
"values": {
"external": {
"ref": { -- lat is irregular so use an external array referenced to elsewhere in the store
"node": "../../../../coords/lat"
}
}
},
"unit": "degrees",
"boundaries": {
"external": {
"ref": { -- Boundary values are irregular too: same external referencing
"node": "../../../../coords/lat_bounds"
}
}
}
}
],
"abbreviation": "Y",
"direction": "NORTH"
}
}
},
{
"type": "temporal", -- Next CRS object, also fully compressed
"axes": {
"time": {
"coordinates": [
{
"values": {
"regular": [60265.5, 1]
},
"time": {
"unit": "days",
"epoch": "1850-01-01",
"calendar": "proleptic_gregorian"
},
"boundaries": {
"regular": [-0.5, 0.5]
}
}
],
"abbreviation": "T",
"direction": "FUTURE",
"attributes": {
"long_name": "time",
"standard_name": "time"
}
}
}
},
{
"type": "vertical", -- Next CRS object - scalar axis for height, fully inlined
"axes": {
"height": {
"coordinates": [
{
"values": {
"explicit": 2
},
"unit": "m"
}
],
"abbreviation": "Z",
"direction": "UP",
"attributes": {
"standard_name": "height",
"long_name": "height"
}
}
}
}
]
},
"zarr_conventions": ["cs", "ref"],
"standard_name": "air_temperature", -- CF-style attributes of the data variable
"long_name": "Daily Maximum Near-Surface Air Temperature",
"comment": "maximum near-surface (usually, 2 meter) air temperature (add cell_method attribute 'time: max')",
"history": "2026-09-14 09:51:33 CEST: Zarr array created with R package ncdfCF0.8.2.9000; 2020-12-24T12:18:40Z altered by CMOR: Treated scalar dimension: 'height'. 2020-12-24T12:18:40Z altered by CMOR: Reordered dimensions, original order: lat lon time.",
"cell_measures": { -- External reference to another object
"ref": {
"node": "../../../fx/areacella"
}
}
}
}This arrangement results in a smaller download volume but crucially also in a much reduced number of HTTPS GET requests when retrieved over the internet: the one XArray needs to know nothing about this whole arrangement: everything is handled in the |
|
you're exactly right that #11541 is going in the same direction, except I believe that the backends should care only about I/O and decoding array data, while the coders are supposed to use arrays and attributes to construct in-memory data structures (mainly indexes for now, but could also be changing array types or other things). The current draft is basically ready (minus documentation and tests), but could benefit from someone writing coders and providing feedback. For example, is the split between dataset and variable coders useful, or would we rather want to be able to define the precise order the coders are applied? In a next step I would try to figure out how to have the coders interact with the backend (if at all). The idea being that some backends like zarr are capable of storing certain dtypes natively (e.g. datetime, timedelta) while others like grib or netcdf are more limited. |
|
I am not convinced that we need a new layer in the form of coders. The I/O in the backend tends to be marginal compared with the decoding/encoding logic. Furthermore, the codec may need to interact with the underlying data format in ways that may not be obvious or actionable right now, so how to define an interface that can deal with that? XArray defines a backend with a declared interface. Why not leave everything else to the backend? It is very likely that the backend will make the distinction between low-level I/O calls and codec logic anyway, so why formalise this (which may impose limitations on what a backend can do)? I am, however, very sympathetic to the direction of the PR: move the codec out of the XArray core (and into the backend), so perhaps the PR can be recast along those lines: the backend is composed of one or more coders that use low-level I/O to construct On your next step: Your dtypes are Python-specific and I would argue for a more open ecosystem view on things (e.g. Julia, R) where it comes to storage. I live mostly in R and from that perspective some Python tools look rather "odd" (time encoding for CF data comes to mind, as opposed to R's I am by no means a XArray expert, nor have I seen a discussion or issue leading up to the coders PR (which I have studied), so I am happy to be educated. |
Uh oh!
There was an error while loading. Please reload this page.
Background
XArray was originally developed to ingest netCDF-3 files whose data was encoded using the CF Metadata Conventions (CF henceforth). When opening a netCDF-3 file, XArray will only "see" the objects located in the root group, from which it will construct a
DataArray. For the new hierarchical structures in the netCDF-4 file format, XArray added the functionality of opening any group in the hierarchy, while maintaining its single-group view on the file: related objects that compose aDataArraystill need to be contained within a single group. ADataTreetakes this one step further by traversing the entire hierarchy of the file and building a corresponding tree ofDataTreenodes, but everyDataArrayis still built from what is contained in a singleDataTreenode.With the emergence of cloud-based data sets that are more complex and larger by the day, this paradigm based on a (for all practical purposes) pre-internet file format from 1997 is no longer fit for purpose. Compounding this is the CF decoding and encoding machinery in the XArray core code (meaning: not in a backend). Effectively, this restricts XArray to using CF-encoded data, or something that is made to resemble it. While that may be defensible from the perspective of the original intent of XArray, it is nevertheless unfortunate that a tool that is so widely used today is not supporting a wider community of users and their data formats.
Moving beyond netCDF (and CF)
Over the years there have been multiple proposals / attempts to reduce the dependency of XArray on the netCDF structure and CF encoding, with limited apparent progress to date. (The same can be said of GeoZarr, which started out with the objective of netCDF/CF alignment in 2020, resulting in no progress made by the end of 2025.) Some prior issues, among various other ones:
There is good reason, however, to revive and pursue those proposals and attempts now.
The CF Conventions Workshop 2026 was held 21-24 September. There was much interest to discuss Zarr and GeoZarr during the workshop and @maxrjones delivered a presentation on GeoZarr principles and approaches, while this author demonstrated a live Zarr store of 4 scientific variables spanning a complete 85-year CMIP6 scenario at daily resolution with all CF attributes and secondary arrays preserved, encoded using 2 GeoZarr conventions, and not duplicating any data. The Zarr store was purposely made hierarchical to mimic the structure of CMIP6 data and to highlight the cross-group reference resolution requirement:
This data set was built in R, but it can also be opened in XArray as a
DataTreeor aDataSetusing the/gr/day/ssp245/r1i1p1f1group and the 4 resultingDataArrayinstances will have coordinates assigned to each of their 3 dimensional axes even though only latitude coordinates and boundary values are explicitly included in a separate group (more on this further down).Perhaps most interestingly, during the same workshop, a breakout session looked at separating the semantics and data model of CF out from its netCDF encoding, with a view of enabling use of the CF conventions with other data formats. In this context Zarr was most often mentioned but other data formats would equally benefit from such an encoding-free description of what CF is and what it wants to achieve. This approach was supported by the workshop participants and work will continue on the cf-conventions GitHub organization to enumerate the actual semantics and data model of CF, before seeking endorsement from the larger CF community through its usual consensus-based decision-making.
Where does that leave XArray?
The world is changing and so should XArray. Larger data stores are not well-served by the "flat" group-based view. CF is looking beyond netCDF and its encoding in other data formats is very likely going to require different decode-assembly code paths, and there may be multiple such format-encoding combinations, as well as data formats that use entirely different ways to label its array dimensions.
XArray is actually in a very good starting position, having the backend mechanism well defined. The major change that XArray has to make is to move all of the manipulation of CF attributes from the core into the default backend. Other backends may then be defined and simply pass of a
DataTreeorDataArrayto the core, with all coordinates, attributes and secondary arrays set. The live demonstration at the CF Workshop using XArray did just that, using thexgroupbackend.Show and Tell: The
xgroupbackendxarray-zarr-xgroupis aBackendEntrypointfor ingesting Zarr data encoded with GeoZarr conventions, registered asengine="xgroup":DataArray.StoreBackendEntrypoint.A real-world example, as presented at the CF Workshop — EC-Earth3-CC CMIP6, 65GB assembled from 69 individual netCDF files, four variables with latitude coordinates in a separate group and other coordinate values and user attributes built from GeoZarr attributes. To emphasise once more: the Zarr store contains no arrays for
lon,lon_bounds,time,time_boundsandheight- these are constructed from the attributes in the GeoZarrcsconvention and attached to theDataArrayby thexgroupbackend. Not shown here: all coordinates and other objects maintain the attributes that described them in the original netCDF files.All arrays are lazy. No chunk data is read during
open_dataset().The repo contains a number of synthetic Zarr stores that may be used for testing. A good testing pattern would be to compare opening any of these stores using the default backend and then with
engine="xgroup".GeoZarr convention support
Coordinate resolution is delegated to convention handlers in the
xgroupbackend. A Zarr array declares its conventions via thezarr_conventionsattribute (per the zarr-conventions spec). The backend detects the declared conventions and calls the appropriate handlers.Currently implemented conventions:
A
principalconvention provides coordinates to be attached to the dimensions of the array while theserviceconventions provide additional features.Third-party conventions may be registered with
xgroupvia Python entry points:Workarounds required for XArray's CF machinery
Implementing this backend exposed several places where XArray's CF decoding machinery assumes conventions live in the core rather than in backends. These required workarounds that would be unnecessary with a cleaner backend/codec separation — directly relevant to #11541:
encoding["coordinates"]XArray only promotes a variable to a coordinate if its name appears in a coordinates attribute on a data variable, or in
encoding["coordinates"]. Variables resolved by our convention handlers and added to the variable dict are not promoted automatically, even when they are 1D variables whose name exactly matches their single dimension. We work around this by settingencoding["coordinates"]on each data variable:The
all()guard inconventions.pydiscards the entire coordinates attribute if any single referenced name is absent from the variable dict. This means one unresolvable out-of-group reference causes all coordinates to be lost with no warning. Our workaround is to resolve all references before XArray sees the variable dict, so theall()guard always passes — but the guard itself is the wrong design.StoreBackendEntrypointapplies CF decoding (time, mask/scale, coordinates) to everything returned by a backend. For GeoZarr conventions this is mostly harmless, but it means backends cannot opt out of CF decoding for variables whose convention explicitly defines their own decoding semantics.Relation to other efforts
The custom coders PR #11541 is exactly the right direction. Once coders are pluggable, a GeoZarr backend can:
cs,spatial, etc.csconvention.encoding["coordinates"]hacksIn the meantime,
xarray-zarr-xgroupdemonstrates that the full functionality is achievable today as aBackendEntrypoint, with the workarounds documented above.Invitation to test
Install and try it on any hierarchical Zarr store:
Feedback welcome on:
The package specification and architecture documentation are in the repository.
P.S.
Pinging some people who have been active in similar issues / discussions:
@keewis @kmuehlbauer @TomNicholas @d-v-b @shoyer @rabernat @dcherian
All reactions