Fix the CSCS pipeline: numcodecs from conda-forge, a SLURM time limit that applies, one test job per precision - #197
Merged
Conversation
…mpiler Both CSCS build jobs fail since conda-forge's Python 3.12.14 build 2: its sysconfig CFLAGS carry a mangled "-partition=none", so every C-extension sdist built against it dies in gcc. numcodecs is the one sdist we build, because zarr<3 caps it below 0.16 and no earlier release has an aarch64 wheel. Take it from conda-forge alongside python-casacore instead; with that nothing in the image is compiled, so build-essential goes too.
…recision Every SLURM job has been submitted with --time=00:15:00: CSCS's .f7t-runner template sets SLURM_TIMELIMIT at job level, which overrides the pipeline-global value added in #93. Since 28 Aug the two-precision test job needs more than that and is killed mid-run. Set the limit at job level on test_job and the perf template, and split the test matrix into JAX build x precision so each job runs one pytest invocation and a failure names its precision.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two independent faults have kept the CSCS pipeline red. This PR fixes both, and splits the test stage so each precision is its own job.
1. Build stage: numcodecs no longer compiles
Both
build_jobvariants fail since the pipeline for 4c2f725 (2 Sep, 19:11): cuda 12, cuda 13. The previous pipeline (cc88ef2, 18:28) built fine.pip install -e ".[test,cudaN]"compiles numcodecs 0.15.1 from its sdist and gcc rejects a flag it is handed:Two things combine:
build-essentialwas in the image.h94ad73b_2_cpython, replacingha505bbe_1_cpython). The feedstock strips its own LTO flags out of thesysconfigCFLAGS that extension builds inherit. Build 2 moved that into a newinstall_base.shwhose gcc detection (${CC} =~ .*gcc.*) does not match there, so only the clang-style-fltogets stripped:-flto-partition=noneloses its prefix and becomes-partition=none, and-fuse-linker-plugin -ffat-lto-objectsstay in, exactly as the CI compile line shows. I checked the package contents directly: build 2's_sysconfigdatacontains-partition=none, build 1's does not. Upstream has the fix open for 3.13 and 3.14 but nothing for 3.12 yet.Every C-extension sdist built against this Python fails the same way, so this is not specific to numcodecs; it is just the only sdist we build.
Fix: install numcodecs from conda-forge in the same step as python-casacore, pinned
<0.16so it satisfies zarr's bound and pip does not try to replace it. conda-forge has linux-aarch64 builds of 0.15.1 for py312. With that, nothing in the image is compiled any more, sobuild-essentialgoes too. This holds regardless of what conda-forge does next.Verified on arm64 in Docker before pushing (the exact conda Python build reproduces the CI error; a mirror of the Dockerfile builds, reports numcodecs already satisfied, builds no C extension, imports and round-trips blosc, and has no gcc), and then by the first pipeline on this PR, where both build jobs passed and the log shows
Requirement already satisfied: numcodecs ... (0.15.1)with only the pure-Python tabascal and asciitree wheels built.2. Test stage: every SLURM job gets 15 minutes, not 30
Across the 78 pipelines from 17 Aug to 2 Sep (last 40 main commits plus last 40 PR heads), 40 of the 51 failed test jobs ended with
at 905 to 940 s of job time. Three perf jobs died the same way. Every job log prints the submitted options and all of them say
#SBATCH --time=00:15:00.The 30-minute
SLURM_TIMELIMITadded in #93 sits in the pipeline-globalvariables:block. CSCS's.f7t-runnertemplate setsSLURM_TIMELIMIT: '00:15:00'in a job-levelvariables:block, and GitLab gives job-level variables inherited throughextendsprecedence over global ones. So our value has never reached SLURM.It only started to matter on 28 Aug: the suite grew from 799 tests on 17 Aug to 2113 today and each job ran it twice, so pytest alone went from about 190 s + 390 s to about 350 s + 530 s. With container start and imports on top the job crossed 15 minutes. Since then the JAX-latest job (the slower one) has timed out in every pipeline and the JAX 0.6.0 job flips between pass and kill by a few seconds.
Fix: set
SLURM_TIMELIMITat job level ontest_joband.perf_job_template, and drop the ineffective global one.3. One test job per JAX build and precision
The matrix is now
JAX_VERSIONxPRECISION, four jobs named e.g.test_job: [latest, double], each running one pytest invocation. A failure then names its precision directly instead of being buried in the second half of a combined log, and each job runs 6 to 9 minutes instead of 15+, so the time limit has real headroom again.