Lenny has three test layers: frontend component tests (React Testing Library, run on every PR), backend unit tests (fast, mocked, run on every PR), and integration tests (real HAPI FHIR + PostgreSQL, run via Docker).
Coverage floor: 70% (enforced by CI). app/main.py and app/services/worker.py are excluded (startup/lifecycle code not suited for unit testing).
Run from frontend/:
cd frontend && npm testUses React Testing Library (@testing-library/react, @testing-library/user-event, @testing-library/jest-dom) via react-scripts test. Tests exercise rendered component behavior — clicks, input changes, conditional rendering — not implementation details.
Test files:
| File | What it covers |
|---|---|
src/components/PeriodPicker.test.js |
Year dropdown, custom date toggle, onChange wiring |
Run from the repo root:
cd backend && python -m pytest tests/ --ignore=tests/integration -vRun with coverage:
cd backend && python -m pytest tests/ --ignore=tests/integration --cov=app --cov-report=term-missing -vThese tests use an in-memory SQLite database and mock all FHIR service calls. They complete in ~15 seconds and require no external infrastructure.
Test files:
| File | What it covers |
|---|---|
test_routes_health.py |
Health check endpoints |
test_routes_jobs.py |
Job creation and status endpoints |
test_routes_measures.py |
Measure listing endpoints |
test_routes_results.py |
Result inspection endpoints |
test_routes_settings.py |
CDR configuration endpoints |
test_routes_validation.py |
Bundle upload and validation run endpoints |
test_services_bundle_loader.py |
Startup connectathon bundle loading |
test_services_fhir_client.py |
FHIR client utilities, DataRequirementsStrategy, BatchQueryStrategy |
test_services_orchestrator.py |
Job orchestration service |
test_services_validation.py |
Bundle triage, population extraction, comparison, measure ID resolution (canonical and relative refs), ValueSet compose patching, canonical-URL clash guard (_assert_no_canonical_url_clash) |
test_cors.py |
CORS middleware behavior |
Integration tests spin up real HAPI FHIR instances and a PostgreSQL database via Docker.
Prerequisites:
# Pre-pull the HAPI image on a new machine (~2.5 GB — do this before your first run)
docker pull hapiproject/hapi:v8.8.0-1
# Start test infrastructure
docker compose -f docker-compose.test.yml up -d
# Wait ~60 seconds for HAPI to initializeRun:
./scripts/run-integration-tests.shThe script waits for HAPI health checks, runs the tests, and tears down containers automatically. If you pass explicit test files, directories, or node IDs, it runs only those targets; option-only invocations still default to the full integration suite.
# Run only the full workflow suite
./scripts/run-integration-tests.sh tests/integration/test_full_workflow.pyIntegration tests are marked @pytest.mark.integration and live in backend/tests/integration/.
Integration test files:
| File | What it covers |
|---|---|
test_fhir_operations.py |
Core FHIR client operations against a live HAPI instance; stays in the PR-gate integration smoke suite |
test_smart_load.py |
triage_test_bundle + bundle loader against live HAPI; stays in the PR-gate integration smoke suite |
test_golden_measures.py |
End-to-end $evaluate-measure against golden bundles; part of the nightly source-of-truth job |
test_connectathon_measures.py |
Parametrized per-test-case run across all connectathon bundles; skipped on the PR gate and included in the nightly source-of-truth job |
test_full_workflow.py |
Full-stack pipeline covering job orchestration → measure eval → result storage; skipped on the PR gate and run in its own clean nightly job |
test_full_jobs_pipeline.py |
Runs all connectathon measures through Lenny's Jobs API and asserts per-patient numerator/denominator outputs match ground-truth expected populations from the connectathon bundles; closes the gap between test_connectathon_measures.py (bypasses Lenny) and test_full_workflow.py (no count assertions); requires prebaked HAPI images; skipped on the PR gate and run in its own nightly jobs-pipeline-validation job |
The nightly connectathon-measures.yml workflow runs once per night and on manual dispatch. It uses four independent clean jobs (PR #203 + jobs-pipeline-validation):
- Bundle Loader Test (vanilla HAPI) — exercises the runtime IG load + bundle replay path against
hapiproject/hapi:v8.8.0-1. - Connectathon Eval (pre-baked HAPI) — runs the connectathon source-of-truth suite (
test_golden_measures.py+test_connectathon_measures.py) against the prebaked GHCR images. - Full Workflow (clean nightly run) — runs
test_full_workflow.pyby itself so it does not inherit the connectathon-loaded CDR state. - Jobs Pipeline Validation (pre-baked HAPI) — runs
test_full_jobs_pipeline.pyagainst the prebaked GHCR images; validates Lenny's orchestration layer end-to-end by asserting per-patient numerator/denominator outputs against ground-truth connectathon populations.
Manually triggering the nightly workflow (Actions → Connectathon Measures → Run workflow) is appropriate before merging measure-engine, HAPI-bump, or bundle-format changes, in addition to the automatic nightly run.
If /jobs reports lower initial population counts than the direct-HAPI integration harness, work the ladder below in order — each step rules out one cause class. The two remaining env vars (PATIENT_DATA_STRATEGY, VALUESET_RELOAD_MODE) are the levers; isolate one at a time.
Step 1 — rule out async-indexing. Use the triage rule in the Recurring bug: HAPI async-indexing race section of CLAUDE.md: read the resource directly (GET /{Type}/{id}) and compare to what the search endpoint returns. If the direct read has data and the search doesn't, async-indexing is still the cause — synchronization.strategy=sync is the structural fix; investigate why it isn't taking effect (check docker-compose config, container restart). If direct read and search agree, async-indexing is not the issue; continue to Step 2.
Step 2 — confirm the data-acquisition path. Set PATIENT_DATA_STRATEGY=data_requirements (or batch_query if you were on data_requirements), restart, and re-run. If counts shift, Lenny is fetching different resource sets between strategies — investigate which strategy is missing resources the measure CQL requires.
Step 3 — confirm ValueSet expansion. Set VALUESET_RELOAD_MODE=remap, restart, and re-run. If counts move, a ValueSet expansion is the culprit (typically a code-system version mismatch between bundle and HAPI's terminology cache).
Background. HAPI expands large ValueSets asynchronously after bundle upload, so Lenny waits for $expand readiness before jobs can run against freshly uploaded Connectathon bundles. Index consistency (POST/PUT blocking until Lucene is refreshed) is handled by synchronization.strategy=sync on both HAPI services — the Python-side compensator was removed in PR #214. The full HAPI consistency model lives in the Recurring bug: HAPI async-indexing race section of CLAUDE.md.
backend/tests/integration/test_connectathon_measures.py runs one parametrized pytest case per test-case MeasureReport found in the connectathon bundles. The set of measures under test — and which ones trigger strict assertions — is controlled entirely by seed/connectathon-bundles/manifest.json.
Each entry in manifest.json has a "strict" field. All current entries have "strict": true. Entries with "expected_test_cases": 0 are definition-only bundles (no patient-level test cases) and produce no parametrized test cases; they are silently skipped.
{
"id": "CMS122FHIRDiabetesAssessGreaterThan9Percent",
"bundle_file": "CMS122FHIRDiabetesAssessGreaterThan9Percent-bundle.json",
"expected_test_cases": 56,
"strict": true,
"known_issues": []
}The STRICT_STU6 environment variable controls how population-count mismatches are handled during test runs:
| Value | Behavior |
|---|---|
STRICT_STU6=1 |
Hard-fail on any population mismatch or HTTP error. |
STRICT_STU6=0 (current CI default) |
Log the mismatch as a warning and mark the test as skipped rather than failed. Use this while onboarding a new measure whose CQL is known to diverge from MADiE. |
# Strict mode
STRICT_STU6=1 ./scripts/run-integration-tests.sh
# Soft mode — mismatches warn instead of fail
STRICT_STU6=0 ./scripts/run-integration-tests.shCI (.github/workflows/pr-checks.yml) currently sets STRICT_STU6=0 during the rollout period. Flip to STRICT_STU6=1 once all connectathon measures pass cleanly in CI. See docs/connectathon-measures-status.md for the current per-measure state.
scripts/connectathon-rehearsal.sh is an operator-level end-to-end smoke test for the full connectathon workflow. Run it before a connectathon event or after a significant infrastructure change.
What it does (six steps):
- Cold-starts Docker services (
docker compose down -v && up -d) unless--no-restartis passed. - Polls
GET /healthuntil all services (db, measure engine, CDR) report connected — up to 5 minutes. - Asserts all measures listed in
manifest.jsonare loaded viaGET /measures. - Triggers
$evaluate-measurefor the first measure withexpected_test_cases > 0, polls until complete, and prints population counts. - Prints a full status table: one row per manifest measure showing
loaded | evaluated | populations_match | notes. - Exits nonzero if any row fails; exits zero on full pass.
Usage:
# Full cold-start rehearsal (tears down volumes — wipes HAPI data)
./scripts/connectathon-rehearsal.sh
# Rehearsal against already-running containers (no restart)
./scripts/connectathon-rehearsal.sh --no-restartAll output is written to the console and appended to rehearsal.log in the repo root. The log file is gitignored and safe to keep between runs as a timing history.
When to run: before any connectathon event, after pulling a new HAPI image, or after updating measure bundles. The --no-restart flag is useful for quick re-checks after a code fix without waiting for HAPI to reload IGs from scratch.
Golden tests validate the full measure evaluation pipeline: a test bundle is loaded into HAPI, $evaluate-measure is called, and the MeasureReport structure is verified.
Location: backend/tests/integration/golden/
Each test case is a directory with a bundle.json:
tests/integration/golden/
basic-measure/
bundle.json <- FHIR transaction bundle with Library, Measure, Patient
To add a new golden test case, create a new directory with a bundle.json:
mkdir -p backend/tests/integration/golden/my-measure
# Write bundle.json with Library, Measure, and Patient resources
# Use unique IDs (e.g. "my-measure-001") to avoid conflicts with seed dataThe test runner (test_golden_measures.py) discovers and runs all bundles automatically.
Golden tests currently assert structural correctness only:
- Response is a
MeasureReport - At least one population group is present
- Patient reference is present
Exact population count assertions should be added once HAPI evaluation behavior is confirmed stable on CI runners.
Every PR to main triggers .github/workflows/pr-checks.yml with 5 jobs:
| Job | What it does | Fails if |
|---|---|---|
| Unit Tests + Coverage | Runs all unit tests with --cov-fail-under=70 |
Any test fails OR coverage < 70% |
| Lint | ruff check + ruff format --check |
Any lint or formatting violation |
| Integration Tests | Spins up Docker containers, runs integration smoke coverage while excluding test_golden_measures.py, test_connectathon_measures.py, and test_full_workflow.py (all moved out of the PR gate). STRICT_STU6=0 during rollout. |
Any remaining integration test fails |
| Script Security Lint | Rejects set -x in prod scripts and password expansion in echo/printf; runs bash unit tests |
Any script violates the rules |
| Frontend Build | npm ci && npm run build |
Build fails |
All 5 jobs must pass before a PR can merge.
The nightly connectathon-measures.yml workflow adds a once-nightly source-of-truth run with clean runner isolation:
source-of-truthrunstest_golden_measures.pyandtest_connectathon_measures.py.full-workflowrunstest_full_workflow.pyin a separate clean runner.jobs-pipeline-validationrunstest_full_jobs_pipeline.pyagainst prebaked HAPI, validating per-patient population correctness end-to-end through the Jobs API.
Trigger it manually from the Actions tab (Run workflow) before merging any change that touches the measure evaluation pipeline, FHIR data flow, or job orchestration.
Coverage is configured in backend/.coveragerc:
[run]
source = app
omit =
app/main.py
app/services/worker.pyExcluded files and rationale:
app/main.py— FastAPI app startup/lifespan, not suited for unit testingapp/services/worker.py— Background job orchestration runner, not suited for unit testing
Check coverage locally:
cd backend && python -m pytest tests/ --ignore=tests/integration --cov=app --cov-report=term-missingThe TOTAL line at the bottom shows the effective coverage percentage.