Skip to content

Commit ed50c83

Browse files
committed
extensive improvement of validation suite and general cleaning
1 parent 309e67f commit ed50c83

50 files changed

Lines changed: 8605 additions & 230 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.
Lines changed: 70 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,70 @@
1+
name: validation-mutation
2+
3+
# The mutation score is the measured coverage of the semantic validation suite.
4+
# The host job runs on every push because it needs no Docker; the container job
5+
# is the authority, replaying the same catalogue under the engines that ship in
6+
# the image.
7+
8+
on:
9+
push:
10+
branches:
11+
- "**"
12+
pull_request:
13+
workflow_dispatch:
14+
15+
jobs:
16+
mutation-score:
17+
name: mutation score (rdflib)
18+
runs-on: ubuntu-latest
19+
steps:
20+
- name: Checkout
21+
uses: actions/checkout@v4
22+
23+
- name: Setup Python
24+
uses: actions/setup-python@v5
25+
with:
26+
python-version: "3.11"
27+
28+
- name: Install test dependencies
29+
run: python -m pip install --upgrade pip rdflib
30+
31+
- name: Run the mutation harness
32+
env:
33+
VCF_RDFIZER_MUTATION_REPORT: mutation-score.json
34+
run: python -m unittest test.test_validation_mutation_unit -v
35+
36+
- name: Report the score
37+
run: |
38+
python - <<'PY'
39+
import json
40+
report = json.load(open("mutation-score.json"))
41+
print(f"mutation score: {report['detected']}/{report['total']} "
42+
f"({report['score']:.0%})")
43+
for item in report["mutations"]:
44+
if not item["detected"]:
45+
print(f" undetected: {item['id']} ({item['representation']}) "
46+
f"- {item['knownUndetected']}")
47+
PY
48+
49+
- name: Upload the score
50+
uses: actions/upload-artifact@v4
51+
with:
52+
name: mutation-score
53+
path: mutation-score.json
54+
55+
engine-agreement:
56+
name: cross-engine query agreement
57+
runs-on: ubuntu-latest
58+
# The image build is slow, so this gates merges rather than every push.
59+
if: github.event_name == 'pull_request' || github.event_name == 'workflow_dispatch'
60+
steps:
61+
- name: Checkout
62+
uses: actions/checkout@v4
63+
64+
- name: Build the image
65+
run: docker build -t vcf-rdfizer:mutation-ci .
66+
67+
- name: Every validation query must agree across Comunica and QLever
68+
run: |
69+
docker run --rm -v "$PWD:/repo:ro" vcf-rdfizer:mutation-ci \
70+
/opt/pycottas-venv/bin/python /repo/test/cross_engine_agreement.py

‎ACKNOWLEDGEMENTS.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -15,7 +15,7 @@ adapted for the Acknowledgements section:
1515
> This work was carried out at KNoWS, IDLab, Ghent University - imec, as part
1616
> of the doctoral research of Elias Crum. The authors thank the maintainers of
1717
> the open-source components on which VCF-RDFizer builds - RMLStreamer,
18-
> hdt-cpp, hdtc, pycottas, Comunica, bcftools, and cyvcf2.
18+
> hdt-cpp, hdtc, pycottas, Comunica, QLever, pySHACL, bcftools, and cyvcf2.
1919
2020
If the work is co-funded by a specific project or grant, add that project name
2121
and grant number to the sentence above and record it in the table at the top of

‎Dockerfile‎

Lines changed: 56 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,11 @@
11
ARG RMLSTREAMER_VERSION=2.5.0
22
ARG HDTC_VERSION=1.1.0
33
ARG COMUNICA_VERSION=5.3.0
4+
# QLever is an optional second SPARQL engine for validation. Its binaries are
5+
# copied from the upstream published image rather than built here: compiling
6+
# QLever needs a large C++ toolchain and would dominate this image's build.
7+
# Pin a digest or release tag here to make validation runs reproducible.
8+
ARG QLEVER_IMAGE=adfreiburg/qlever:latest
49

510
FROM eclipse-temurin:11-jre AS build-hdt-cpp
611

@@ -59,6 +64,10 @@ RUN cargo build --locked --release \
5964
&& cp LICENSE /opt/third_party_licenses/HDTC.LICENSE
6065

6166

67+
# Named stage so the runtime image can COPY QLever's binaries out of it.
68+
FROM ${QLEVER_IMAGE} AS qlever
69+
70+
6271
FROM eclipse-temurin:11-jre
6372

6473
ARG RMLSTREAMER_VERSION
@@ -94,7 +103,8 @@ RUN python3 -m venv /opt/pycottas-venv \
94103
duckdb==1.5.5 \
95104
pyarrow==22.0.0 \
96105
numpy==2.4.6 \
97-
cyvcf2==0.34.0
106+
cyvcf2==0.34.0 \
107+
pyshacl==0.30.1
98108

99109
RUN npm install --global "@comunica/query-sparql-file@${COMUNICA_VERSION}"
100110

@@ -110,6 +120,33 @@ COPY --from=build-hdt-cpp /usr/local/lib/libhdt* /usr/local/lib/
110120
COPY --from=build-hdt-cpp /opt/third_party_licenses/ /usr/share/licenses/vcf-rdfizer/
111121
COPY --from=build-hdtc /opt/hdtc/target/release/hdtc /usr/local/bin/hdtc
112122
COPY --from=build-hdtc /opt/third_party_licenses/ /usr/share/licenses/vcf-rdfizer/
123+
124+
# Optional QLever SPARQL engine (--validation-engine qlever). Comunica remains
125+
# the default, so an image whose QLever binaries turn out to be unusable is
126+
# still fully functional; the validator reports the reason instead of failing
127+
# obscurely. Pin a different tag or digest with
128+
# --build-arg QLEVER_IMAGE=adfreiburg/qlever:<tag>.
129+
#
130+
# QLever's image is built on a different Ubuntu release than this one, so the
131+
# binaries alone are not enough: their Boost, ICU, jemalloc and io_uring
132+
# sonames are release-specific and absent here. Those libraries travel with the
133+
# binaries into a private directory that only QLever's own processes are
134+
# pointed at, so they cannot shadow anything the rest of the image links
135+
# against. (glibc itself is not copied - it is backward compatible, and this
136+
# base is newer than QLever's.)
137+
COPY --from=qlever /qlever/qlever-index /opt/qlever/bin/qlever-index
138+
COPY --from=qlever /qlever/qlever-server /opt/qlever/bin/qlever-server
139+
COPY --from=qlever \
140+
/lib/x86_64-linux-gnu/libboost_iostreams.so.1.83.0 \
141+
/lib/x86_64-linux-gnu/libboost_program_options.so.1.83.0 \
142+
/lib/x86_64-linux-gnu/libboost_url.so.1.83.0 \
143+
/lib/x86_64-linux-gnu/libgomp.so.1 \
144+
/lib/x86_64-linux-gnu/libicudata.so.74 \
145+
/lib/x86_64-linux-gnu/libicui18n.so.74 \
146+
/lib/x86_64-linux-gnu/libicuuc.so.74 \
147+
/lib/x86_64-linux-gnu/libjemalloc.so.2 \
148+
/lib/x86_64-linux-gnu/liburing.so.2 \
149+
/opt/qlever/lib/
113150
COPY THIRD_PARTY_NOTICES.md /usr/share/licenses/vcf-rdfizer/THIRD_PARTY_NOTICES.md
114151
COPY src/*.sh /opt/vcf-rdfizer/
115152
COPY src/*.py /opt/vcf-rdfizer/
@@ -121,7 +158,22 @@ COPY vcf_rdfizer_gzip.py /opt/vcf-rdfizer/
121158
RUN chmod +x /opt/vcf-rdfizer/*.sh \
122159
&& chmod +x /usr/local/bin/rdf2hdt \
123160
&& chmod +x /usr/local/bin/hdt2rdf \
124-
&& chmod +x /usr/local/bin/hdtc
161+
&& chmod +x /usr/local/bin/hdtc \
162+
&& chmod +x /opt/qlever/bin/qlever-index /opt/qlever/bin/qlever-server
163+
164+
# QLever's binaries come from a different base image, so record at build time
165+
# whether they actually link here. The validator reads this marker to explain
166+
# an unavailable engine instead of surfacing a bare "not found".
167+
RUN set -eu; \
168+
status="ok"; \
169+
for binary in qlever-index qlever-server; do \
170+
missing="$(LD_LIBRARY_PATH=/opt/qlever/lib ldd "/opt/qlever/bin/$binary" 2>&1 | grep 'not found' || true)"; \
171+
if [ -n "$missing" ]; then \
172+
status="QLever binary $binary has unresolved shared libraries in this base image: $missing"; \
173+
echo "WARNING: $status" >&2; \
174+
fi; \
175+
done; \
176+
printf '%s\n' "$status" > /opt/vcf-rdfizer/qlever-status.txt
125177

126178
ENV RMLSTREAMER_JAR=/opt/rmlstreamer/RMLStreamer-v${RMLSTREAMER_VERSION}-standalone.jar
127179
ENV JAR=/opt/rmlstreamer/RMLStreamer-v${RMLSTREAMER_VERSION}-standalone.jar
@@ -132,6 +184,8 @@ ENV COTTAS_MERGE_BATCH_ROWS=2048
132184
ENV RDF2HDT_BIN=/usr/local/bin/rdf2hdt
133185
ENV HDT2RDF_BIN=/usr/local/bin/hdt2rdf
134186
ENV COTTAS_PYTHON_BIN=/opt/pycottas-venv/bin/python
187+
ENV QLEVER_INDEX_BUILDER_BIN=/opt/qlever/bin/qlever-index
188+
ENV QLEVER_SERVER_BIN=/opt/qlever/bin/qlever-server
135189
ENV LD_LIBRARY_PATH=/usr/local/lib
136190

137191
# COTTAS creates a temporary DuckDB database in the container working

‎README.md‎

Lines changed: 94 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -19,6 +19,12 @@ VCF-RDFizer is a Docker-first CLI wrapper for:
1919

2020
The VCF-RDFizer vocabulary is available at [https://w3id.org/vcf-rdfizer/vocab#](https://w3id.org/vcf-rdfizer/vocab#).
2121

22+
This README is the task-oriented reference. For how the tool works, why it is
23+
built that way, and where it stops working, see the documentation set in
24+
**[`docs/`](docs/README.md)** - starting with
25+
[Architecture](docs/architecture.md) and, before you rely on the output,
26+
[Limitations](docs/limitations.md).
27+
2228
## Requirements
2329

2430
- Python 3.10+
@@ -134,6 +140,59 @@ stage summary at `stages/validation/<dataset-id>.json`. See [Semantic VCF/RDF
134140
validation](docs/validation.md) for query definitions, preflight checks, result
135141
statuses, and cleanup evidence.
136142

143+
144+
### Artifacts and engines
145+
146+
`--rdf` accepts any artifact the pipeline produces - `.nt`, `.nt.gz`, `.nt.br`,
147+
`.hdt`, `.cottas`, `.cottas.gz`, `.cottas.br`. A compressed or indexed artifact
148+
is decoded back to N-Triples **inside the container** and then put through the
149+
full semantic suite, which proves it decodes to a graph that still reproduces
150+
every VCF summary - stronger than the triple-count round-trip that runs during
151+
compression.
152+
153+
In full mode, `--validate-artifacts {aggregate,hdt,cottas,all}` chooses which
154+
produced artifacts to check; each is validated independently with its own
155+
report:
156+
157+
```bash
158+
vcf-rdfizer --mode full -i ./cohort.vcf.gz \
159+
--rdf-storage-mode space-optimized --representations hdt,cottas \
160+
--validate --validate-artifacts all -o ./results
161+
```
162+
163+
`--validation-engine {comunica,qlever}` selects the SPARQL backend. Comunica
164+
(default) queries the file in memory; [QLever](https://github.com/ad-freiburg/qlever)
165+
builds an on-disk index inside the container and serves it, which is what makes
166+
cohort-scale graphs queryable. Both answer identical queries, so the choice is
167+
never semantic, and every report records which engine ran.
168+
169+
```bash
170+
vcf-rdfizer --mode validation -i ./cohort.vcf.gz --rdf ./results/cohort/cohort.hdt \
171+
--validation-engine qlever --qlever-memory-gb 32 -o ./validation-results
172+
```
173+
174+
Tuning: `--qlever-memory-gb`, `--qlever-port`, `--qlever-startup-timeout`,
175+
`--validation-query-timeout`, and repeatable `--qlever-index-arg` /
176+
`--qlever-server-arg` escape hatches.
177+
178+
Validation runs three independent layers: exact aggregate comparison against
179+
the VCF, a predicate/class census plus per-record and per-value identity
180+
digests, and — with `--shacl-shapes` — SHACL conformance against the
181+
vocabulary's published shapes.
182+
183+
```bash
184+
vcf-rdfizer --mode validation -i ./cohort.vcf.gz --rdf ./results/cohort/cohort.nt.gz \
185+
--shacl-shapes ./vocabulary/shacl/vcf-rdfizer-vocabulary.shacl.ttl \
186+
--strict-conformance -o ./validation-results
187+
```
188+
189+
> **What a PASS means.** Coverage is measured, not asserted: a mutation harness
190+
> corrupts a correct graph in 36 named ways and records which are detected
191+
> (currently **64/66**). See [`docs/vcf-coverage.md`](docs/vcf-coverage.md) for
192+
> the element-by-element matrix and the remaining gaps, and
193+
> [`docs/validation-methodology.md`](docs/validation-methodology.md) for how the
194+
> number is produced.
195+
137196
## Compression Plan
138197

139198
Compression is configured as three independent decisions:
@@ -204,6 +263,23 @@ host filesystem.
204263
- `--remove-rdf-storage-output` explicitly remove the aggregate `.nt`/`.nt.gz` after successful compression
205264
- `-e, --estimate-size` preflight size estimate
206265

266+
## VCF Coverage
267+
268+
Full-mode conversion covers every VCF column. Three options control how much
269+
structure is emitted; all default to the richer form.
270+
271+
| Option | Effect |
272+
| --- | --- |
273+
| `--sample-representation {expanded,condensed}` | Genotype shape (see below) |
274+
| `--info-representation {structured,raw}` | `structured` adds one `vcfr:InfoFieldValue` per record and key, with typed values, alongside `vcfr:infoRaw` |
275+
| `--header-representation {structured,basic}` | `structured` types each `##` line with its vocabulary subclass and lifts FILTER/ALT/contig attributes into their own properties |
276+
277+
QUAL is always emitted, typed `xsd:decimal` or `vcfr:Null`, because the
278+
published SHACL shape requires a datatype that depends on the value.
279+
280+
[`docs/vcf-coverage.md`](docs/vcf-coverage.md) maps every VCF element to its RDF
281+
terms and to the validation check that covers it.
282+
207283
## Sample Representation Modes
208284

209285
Full mode has exactly two explicit sample workflows. There is no automatic
@@ -853,10 +929,25 @@ in `vcf_rdfizer.py` append it directly to the aggregate, because the equivalent
853929
RML maps would first have to materialize variants x samples (x FORMAT keys)
854930
helper TSV rows. See [`docs/sample-representation-guide.md`](docs/sample-representation-guide.md).
855931

856-
Further reading:
932+
Further reading: **[`docs/`](docs/README.md) is the in-depth documentation set** -
933+
how each part of the tool works, why, and where it stops working.
934+
935+
| Document | Covers |
936+
| --- | --- |
937+
| [Architecture](docs/architecture.md) | Host/container split, failure policy, pinned toolchain |
938+
| [Conversion](docs/conversion.md) | VCF -> TSV -> RDF, stage by stage |
939+
| [Representations](docs/representations.md) | Compression, HDT/COTTAS, chunking, round-trip checks |
940+
| [Output and metrics](docs/output-and-metrics.md) | Run layout, reports, progress, exit codes |
941+
| [Custom RML mappings](docs/rml-mappings.md) | The `--rules` contract in full |
942+
| [Sample representations](docs/sample-representation-guide.md) | Expanded vs condensed genotype shapes |
943+
| [Validation](docs/validation.md) | The semantic suite, and what it does not test |
944+
| [Validation methodology](docs/validation-methodology.md) | How coverage is measured, not asserted |
945+
| [VCF coverage matrix](docs/vcf-coverage.md) | Element by element, with the mutation that proves each row |
946+
| [CLI reference](docs/cli-reference.md) | Every flag, with constraints and interactions |
947+
| [Limitations](docs/limitations.md) | Everything the tool cannot do, in one place |
948+
| [Roadmap](docs/roadmap.md) | Planned work, known defects, and rejected options |
949+
| [Data linking design](docs/datalinking-design.md) | Proposal: a plug-in system for external links |
857950

858-
- [`docs/validation.md`](docs/validation.md) - semantic validation design and query set
859-
- [`docs/sample-representation-guide.md`](docs/sample-representation-guide.md) - emitted genotype shapes
860951
- [`changelog.md`](changelog.md) - dated change history
861952
- [`ACKNOWLEDGEMENTS.md`](ACKNOWLEDGEMENTS.md) - funding and attribution
862953

0 commit comments

Comments
 (0)