Skip to content

Fix the CSCS pipeline: numcodecs from conda-forge, a SLURM time limit that applies, one test job per precision - #197

Merged
chrisfinlay merged 2 commits into
mainfrom
cscs-numcodecs-from-conda
Sep 2, 2026
Merged

chrisfinlay merged 2 commits into
mainfrom
cscs-numcodecs-from-conda

Conversation

@chrisfinlay

@chrisfinlay chrisfinlay commented Sep 2, 2026 •

Copy link
Copy Markdown
Collaborator

Two independent faults have kept the CSCS pipeline red. This PR fixes both, and splits the test stage so each precision is its own job.

1. Build stage: numcodecs no longer compiles

Both build_job variants fail since the pipeline for 4c2f725 (2 Sep, 19:11): cuda 12, cuda 13. The previous pipeline (cc88ef2, 18:28) built fine.

pip install -e ".[test,cudaN]" compiles numcodecs 0.15.1 from its sdist and gcc rejects a flag it is handed:

Building wheel for numcodecs (pyproject.toml): finished with status 'error'
gcc: error: unrecognized command-line option ‘-partition=none’; did you mean ‘-flto-partition=none’?

Two things combine:

  • numcodecs has no aarch64 wheel below 0.16 (0.16.0 was the first release with one), and tabsim -> zarr<3 caps numcodecs below 0.16 (Investigate lifting the zarr<3.0.0 bound once dask-ms supports zarr 3 #120). So on the GH200 nodes pip has always built it from source. That is why build-essential was in the image.
  • conda-forge published a new build of Python 3.12.14 on 1 Sep (h94ad73b_2_cpython, replacing ha505bbe_1_cpython). The feedstock strips its own LTO flags out of the sysconfig CFLAGS that extension builds inherit. Build 2 moved that into a new install_base.sh whose gcc detection (${CC} =~ .*gcc.*) does not match there, so only the clang-style -flto gets stripped: -flto-partition=none loses its prefix and becomes -partition=none, and -fuse-linker-plugin -ffat-lto-objects stay in, exactly as the CI compile line shows. I checked the package contents directly: build 2's _sysconfigdata contains -partition=none, build 1's does not. Upstream has the fix open for 3.13 and 3.14 but nothing for 3.12 yet.

Every C-extension sdist built against this Python fails the same way, so this is not specific to numcodecs; it is just the only sdist we build.

Fix: install numcodecs from conda-forge in the same step as python-casacore, pinned <0.16 so it satisfies zarr's bound and pip does not try to replace it. conda-forge has linux-aarch64 builds of 0.15.1 for py312. With that, nothing in the image is compiled any more, so build-essential goes too. This holds regardless of what conda-forge does next.

Verified on arm64 in Docker before pushing (the exact conda Python build reproduces the CI error; a mirror of the Dockerfile builds, reports numcodecs already satisfied, builds no C extension, imports and round-trips blosc, and has no gcc), and then by the first pipeline on this PR, where both build jobs passed and the log shows Requirement already satisfied: numcodecs ... (0.15.1) with only the pure-Python tabascal and asciitree wheels built.

2. Test stage: every SLURM job gets 15 minutes, not 30

Across the 78 pipelines from 17 Aug to 2 Sep (last 40 main commits plus last 40 PR heads), 40 of the 51 failed test jobs ended with

error: *** STEP ... CANCELLED AT ... DUE TO TIME LIMIT ***
ERROR: Job failed: exit status 137

at 905 to 940 s of job time. Three perf jobs died the same way. Every job log prints the submitted options and all of them say #SBATCH --time=00:15:00.

The 30-minute SLURM_TIMELIMIT added in #93 sits in the pipeline-global variables: block. CSCS's .f7t-runner template sets SLURM_TIMELIMIT: '00:15:00' in a job-level variables: block, and GitLab gives job-level variables inherited through extends precedence over global ones. So our value has never reached SLURM.

It only started to matter on 28 Aug: the suite grew from 799 tests on 17 Aug to 2113 today and each job ran it twice, so pytest alone went from about 190 s + 390 s to about 350 s + 530 s. With container start and imports on top the job crossed 15 minutes. Since then the JAX-latest job (the slower one) has timed out in every pipeline and the JAX 0.6.0 job flips between pass and kill by a few seconds.

Fix: set SLURM_TIMELIMIT at job level on test_job and .perf_job_template, and drop the ineffective global one.

3. One test job per JAX build and precision

The matrix is now JAX_VERSION x PRECISION, four jobs named e.g. test_job: [latest, double], each running one pytest invocation. A failure then names its precision directly instead of being buried in the second half of a combined log, and each job runs 6 to 9 minutes instead of 15+, so the time limit has real headroom again.

…mpiler

Both CSCS build jobs fail since conda-forge's Python 3.12.14 build 2: its
sysconfig CFLAGS carry a mangled "-partition=none", so every C-extension
sdist built against it dies in gcc. numcodecs is the one sdist we build,
because zarr<3 caps it below 0.16 and no earlier release has an aarch64
wheel. Take it from conda-forge alongside python-casacore instead; with
that nothing in the image is compiled, so build-essential goes too.
…recision

Every SLURM job has been submitted with --time=00:15:00: CSCS's .f7t-runner
template sets SLURM_TIMELIMIT at job level, which overrides the
pipeline-global value added in #93. Since 28 Aug the two-precision test job
needs more than that and is killed mid-run. Set the limit at job level on
test_job and the perf template, and split the test matrix into JAX build x
precision so each job runs one pytest invocation and a failure names its
precision.
@chrisfinlay chrisfinlay changed the title Install numcodecs from conda-forge in the CSCS image, and drop the compiler Fix the CSCS pipeline: numcodecs from conda-forge, a SLURM time limit that applies, one test job per precision Sep 2, 2026
@chrisfinlay
chrisfinlay merged commit eac55a3 into main Sep 2, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant