Skip to content

E57 dataset - #101

Merged
harry7557558 merged 7 commits into
harry7557558:masterfrom
vheun:e57-dataset
Oct 5, 2026
Merged

harry7557558 merged 7 commits into
harry7557558:masterfrom
vheun:e57-dataset

Conversation

@vheun

@vheun vheun commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Summary

Laser scanners that also take photos have already registered every image. This adds a way to turn an E57 scan straight into a training dataset (images, poses and a seed point cloud) with no SfM.

  • CLI: spirula e57 <scan.e57> [<folder>] [--points <n>|all] [--no-depth] [--no-pinhole] [--no-panoramas] [--overwrite] [--info]
  • GUI: File > Create Dataset from E57..., a button on Home, or drop an .e57 on the window.

The output is a transforms.json dataset, one of the three layouts the trainer already reads, so it opens and trains like any other.

What's in it

Reading and converting

  • src/data/E57Reader.{h,cpp}: a standalone E57 reader with no libE57Format or Xerces. It reads the XML section, the paged binary sections, bitPackCodec point records across packets, and image blobs. Invalid points and points at the scanner origin are dropped. Colour is normalised, and a scan without colour is seeded grey from its intensity.
  • src/app/E57Dataset.{h,cpp}: pinhole images become PINHOLE cameras and full-sphere panoramas become EQUIRECTANGULAR cameras. Cylindrical, partial-sphere and uncalibrated images are skipped, with the reason logged. Poses and intrinsics are written in full double precision.
  • Alignment check: every conversion projects the scan's coloured points into a few images of each kind under all 24 axis turns. It keeps the assumed orientation unless another one correlates clearly better; in that case it corrects the cameras and says so.
  • Seed points: voxel-thinned to --points (default 500,000), or all of them with --points all ("Use every point"). They are written as a double-precision PLY so geo-referenced scans keep millimetre precision.
  • Depth and normal maps (src/app/ScanDepth.{h,cpp}): rendered from the laser points for every image, as 16-bit millimetres and normals facing the camera. They feed the trainer's existing supervision: normals are on by default, depth needs --depth-supervision-weight. --no-depth leaves them out.
  • gauge.txt marks the frame as oriented and metric, since E57 is in metres and scanners level their frame.

GUI (E57Runner, and the E57 screen in GuiApp)

  • A summary of the file as soon as it is picked, and a 3D preview of the scan's cameras and a million of its points. The preview is the dataset screen's viewport in preview mode.
  • Masking reuses the dataset screen's code on this page: the SAM model and prompt, stencils, "Try the mask" and the correction editor.
    • The prompt starts from the 360-camera preset, and Create waits until there is a prompt or a click.
    • Masking runs as part of the job, after the conversion.
    • If masking fails or is cancelled, its partial masks are removed and the dataset is left unmasked.
  • Progress uses the dataset screen's step row (Frames > Scans > Depth > Seed points > Masks) and its film reels: frames, photo/normal/depth rows, and masks. The reels read the child process's output folder on their own thread, because scanner photos can be tens of megapixels.

Changes to shared code

  • Nerfstudio datasets now read gauge.txt. read_gauge moved from ColmapParser.cpp to dsparse and is called by both parsers. A Nerfstudio dataset without a gauge.txt behaves as before.
  • OutputWatch moved from GeometryRunner.cpp into FilmReel.{h,cpp} and gained an optional fresh_only flag. The geometry step behaves as before.
  • The video dataset page's step row, view tabs and masking panel are now shared helpers: draw_run_steps, draw_view_tabs and draw_masking_options(MaskingPanel). No behaviour change is intended.
  • write_ply_points has a new optional double_xyz parameter; the default is unchanged.
  • New i18n catalog src/i18n/catalog/E57.h with all 13 languages, plus menu, Home and tool-list entries in Gui.h and Cli.h.
  • Docs: an "E57 laser scans" section in docs/datasets.md, a README bullet, and the AGENTS.md repo map.

No new dependencies.

Testing

  • New tests, all passing:
    • e57_reader_test: synthetic E57 files with multi-packet and empty-packet vectors, JPEG and PNG blobs, and a truncated file.
    • e57_dataset_test: the camera round trip through transforms.json stays within 0.002 px, and the alignment check turns a wrongly oriented photo back.
    • scan_depth_test: a plane, a sparse patch in front of it, and a sphere around a panorama.
  • Project checks pass: comment budget, i18n, font coverage, Windows macro names, private paths and comment citations. The translated progress lines the GUI reads back were checked to parse correctly in all 13 languages (a one-off script, not a committed test).
  • Real scans: a Leica BLK360 (pinhole cube faces) and an XGRIDS Lixel (panoramas) on macOS with the Vulkan backend, and a Matterport Pro3 on Windows.
  • GUI end to end on the BLK360 scan: conversion, depth maps, the preview, SAM 3 masking (cancelled after a few frames), and the cancel path.

Not tested: a masking run to completion (SAM 3 takes about 9 s per frame on the Mac used), a video dataset run after the shared-helper refactor, Linux, and the CUDA backend.

Known limitations

  • The downward face of a tripod station can be a black disc where the scanner painted out itself and the tripod. Nothing in the masking stack targets a fixed region in only some images.
  • Upward faces are sky with no points behind them; background_mode sh is the setting for outdoor scans.
  • A tripod scan has as many viewpoints as it has stations, far fewer than a walk-around video.
  • Cylindrical and partial-sphere images are skipped.

vheun and others added 3 commits September 25, 2026 17:27
Adds `spirula e57` and File > Create Dataset from E57: a scan's registered
images, poses and point cloud written out as a transforms.json dataset, with
no reconstruction.

- data/E57Reader: standalone E57 reader (XML, paged binary sections,
  bitPackCodec point records, image blobs).
- app/E57Dataset: pinhole and panorama cameras, voxel-thinned seed points
  (or all of them), a check of the photos' alignment against the scan's
  colours, gauge.txt (oriented, metric).
- app/ScanDepth: depth and normal maps rendered from the laser points.
- GUI: an E57 screen with a 3D preview of the scan, the dataset screen's
  masking (prompt seeded from the 360-camera preset) run as part of the job,
  and its step row and film reels as progress. Masking that fails or is
  cancelled leaves the dataset unmasked.
- Shared: OutputWatch moved into FilmReel, draw_masking_options and the
  step/view helpers made reusable, gauge.txt read for Nerfstudio datasets
  too, double-precision PLY seed points.
- Tests: e57_reader_test, e57_dataset_test, scan_depth_test.
@vheun

vheun commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

Related to #99

@harry7557558 harry7557558 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for contributing! I tuned it on a number of publicly available datasets and made some improvements to the alignment, as well as some UI changes, and I think it's now ready to merge. If you notice regression on your datasets, please let me know.

@Myarcer

Myarcer commented Oct 7, 2026

Copy link
Copy Markdown

Stupid Question, but are XGRID Camera Datasets supported? I didnt find any info specifically on what settings or how to import them.

@harry7557558

Copy link
Copy Markdown
Owner

@Myarcer I've only tested on two XGRIDS datasets, both private ones shared by users for testing only. The datasets are in the following format:

output/
├── developer_data/
│   ├── perspective
│   ├── raw/
│   │   ├── images
│   │   └── sparse
│   └── map_remov.las
├── nCore_data/
│   └── (similar to developer_data)
└── mesh/, etc.

If your dataset is the same case, drag and drop the output folder into the home screen should bring you to a dataset creation screen, and click "Update dataset" should give you depth/normal maps as well as a denser point cloud written into output/developer_data/raw. You can also go to dataset creation screen and load an existing COLMAP reconstruction (one with images and sparse subfolders) as well as the LAS point cloud, tick the "The dataset is already in the scan's frame` checkbox, and then click "Update dataset". If neither works for you, and your dataset is in a different format, consider sharing a minimum dataset with me so I can take a look and add support.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants