Repository navigation
Make it as easy as possible to add new data types #359
Description
Activity
After a code review, a change of this kind is harder than I thought, as the directory segments need to know the data format in order to know the file extension to use to find the files inside.
However, having JPEG and NetCDF in the pipeline is still an opportunity for a refactoring to make it significantly easier to add data formats. Ideally, one should be able to just add an entry to
DataFormatinarki/defs.h, recompile, and add some python code to do the scanning.We are not there yet, but that should be achievable without too much work and without breaking compatibility. I'll work on it
- changed the title
[-]New data type: file[/-][+]Make it as easy as possible to add new data types[/+]on Dec 18, 2025 - added 12 commits that reference this issue
on Dec 18, 2025 3 remaining items
- added 12 commits that reference this issue
on Dec 20, 2025 This should now be done: has a version of arkimet with this change been deployed in production? If scanning still works after the refactoring, then the only thing that is needed is for me to write a documentation guide on how to add a new file format
- added a commit that references this issue
on Jan 28, 2026 I have now added
doc/howtos/new-data-format.rst
In a conversation about the zarr file format I realised that arkimet doesn't need to know if a file is a JPEG, a NetCDF, an HDF5, zarr, and so on, as long as it's a data format that is stored in a single standalone file that arkimet shouldn't store concatenated like GRIB and BUFR.
I'm therefore proposing to introduce a BLOB file format named "FILE" identifying such data types.
Scanning can be delegated to Python: a python function can get a
Pathobject, look at its extension and decide wheter to return a Metadata or delegate scanning to the standard arkimet.This is worth doing before deploying JPEG and NetCDF support to production, so that they can be handled under the FILE scheme without needing to plan a transition