Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion .github/workflows/fern-docs-ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ on:
description: Publish an isolated hosted preview from this branch
type: choice
default: none
options: [none, edition]
options: [none, edition, canonical]

permissions:
contents: read
Expand Down Expand Up @@ -108,6 +108,7 @@ jobs:
FERN_TOKEN: ${{ secrets.DOCS_FERN_TOKEN }}
DOCS_PREVIEW_ID: nvcf-${{ github.run_id }}
DOCS_PREVIEW_SKIP_COMMENT: '1'
DOCS_PREVIEW_CONFIG: ${{ inputs.preview == 'canonical' && 'fern/docs.yml' || '' }}
run: |
set -euo pipefail
./tools/ci/preview-docs | tee "$RUNNER_TEMP/docs-preview.log"
Expand Down
4 changes: 2 additions & 2 deletions docs/edition-manifest.json
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
{
"schema_version": 1,
"docs_edition": {
"version": "1.0.3",
"version": "1.0.4",
"status": "development",
"previous_version": "1.0.2",
"previous_version": "1.0.3",
"change": "patch"
},
"stacks": {
Expand Down
255 changes: 255 additions & 0 deletions docs/overview/release-notes/0.6.1-to-0.6.2-upgrade.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,255 @@
# Upgrade from NVCF Self-Managed 0.6.1 to 0.6.2

The 0.6.2 Helmfile bundle patch refreshes OpenBao's JWT plugin catalog before the
larger 1.0.1 upgrade. The patch changes the OpenBao chart from 0.30.25 to
0.32.6 and keeps the 0.6.1 runtime versions. Complete these steps and validate
the deployment before starting the [0.6.2 to 1.0.1 upgrade](./0.6.1-to-1.0.1-upgrade.md).

## 1. Record a healthy source deployment

Start with a healthy 0.6.1 deployment. Every source Helm release must be
`deployed`, every OpenBao member initialized and unsealed, and every Cassandra
migration ledger clean. Resolve degraded health before changing charts.

Use Helm 3 and Helmfile 1.1.x. Set the bundle paths, environment, and intended
Kubernetes context. Use a new private directory for backups:

```bash
set -euo pipefail
export SOURCE_CONTROL=<absolute-path-to-0.6.1-bundle>
export TARGET_CONTROL=<new-absolute-path-for-0.6.2-bundle>
export CONTROL_KUBECONFIG=<control-plane-kubeconfig>
export CONTROL_CONTEXT=<control-plane-context>
export HELMFILE_ENV=<existing-environment-name>
export BACKUP_DIR=<new-private-backup-directory>

umask 077
mkdir -p "$BACKUP_DIR"
kcp() {
kubectl --kubeconfig "$CONTROL_KUBECONFIG" \
--context "$CONTROL_CONTEXT" "$@"
}
test "$(kubectl --kubeconfig "$CONTROL_KUBECONFIG" config current-context)" \
= "$CONTROL_CONTEXT"
kcp cluster-info
helm --kubeconfig "$CONTROL_KUBECONFIG" --kube-context "$CONTROL_CONTEXT" \
list --all-namespaces > "$BACKUP_DIR/helm-releases.txt"
kcp -n vault-system get statefulset openbao-server -o yaml \
> "$BACKUP_DIR/openbao-statefulset.yaml"
kcp -n vault-system get pods -o json > "$BACKUP_DIR/openbao-pods.json"
```

Record the chart versions, image digests, replica counts, storage settings,
plugin catalog digest, JWT mount configuration, and injector webhook settings.
Keep any intentional server or migration image override in this baseline.

Record an existing function and deployment, their cluster ID, worker pod UID,
container IDs, and restart counts. Invoke that function through the existing
authentication flow, then keep a timestamped probe running during the patch.
Use the same function afterward to check that the upgrade retained the
original deployment.

## 2. Back up state

Save the source environment values and secrets securely. From an authenticated
operator connection, take an OpenBao Raft snapshot:

```bash
bao operator raft snapshot save "$BACKUP_DIR/openbao.snap"
```

Take a Cassandra snapshot on every member using your site's backup procedure.
Save the migration ledgers, database topology, and persistent-volume identities.
Keep the source bundle and mirrored images with the recovery baseline so a
restore does not depend on a later registry lookup. Verify your restore
procedure before continuing. Do not put tokens, unseal material, or rendered
secrets into shared logs.

## 3. Prepare the patch bundle

Obtain the `nvcf-self-managed-stack` 0.6.2 Helmfile bundle through your
NVCF artifact distribution channel. Confirm access to that exact version before
scheduling the upgrade. Compare its contents and source commit with the
[stable 0.6.2 release](https://github.com/NVIDIA/nvcf/releases/tag/deploy/stacks/self-managed/v0.6.2)
and its
[artifact inventory](https://github.com/NVIDIA/nvcf/releases/download/deploy/stacks/self-managed/v0.6.2/nvcf-self-managed-stack-inventory.json).
The recorded source commit is `55803ad7cb647de45cf29415cfd3766225534617`.
Verify the bundle checksum supplied by your distribution channel, then extract
it into `$TARGET_CONTROL`; keep `$SOURCE_CONTROL` intact for recovery.

Use the [Image Mirroring](../image-mirroring.md) process with the versions in
that inventory. The current documentation manifest describes the 1.0.1 stack;
do not substitute its artifact versions for this legacy patch.

Start with the target templates and carry your settings into these files:

```text
$TARGET_CONTROL/environments/$HELMFILE_ENV.yaml
$TARGET_CONTROL/secrets/$HELMFILE_ENV-secrets.yaml
```

Preserve registry and pull-secret settings, endpoints, certificates, storage
classes and sizes, replica counts, scheduling constraints, Cassandra passwords,
OpenBao configuration, and injector webhook scope. Keep enabled optional
components and existing secret and cluster identities.

Keep the target `global.yaml.gotmpl`. It contains the runtime pinning fix, so
copying the source file over it would remove part of the patch. Mirror required
artifacts at the exact versions listed in the 0.6.2 inventory.

## 4. Review the rendered change

Render the OpenBao release into the private backup directory and inspect its
diff before applying it:

```bash
cd "$TARGET_CONTROL"
HELMFILE_ENV="$HELMFILE_ENV" helmfile --environment default \
--kubeconfig "$CONTROL_KUBECONFIG" --kube-context "$CONTROL_CONTEXT" \
--selector name=openbao-server template > "$BACKUP_DIR/openbao-target.yaml"
HELMFILE_ENV="$HELMFILE_ENV" helmfile --environment default \
--kubeconfig "$CONTROL_KUBECONFIG" --kube-context "$CONTROL_CONTEXT" \
--selector name=openbao-server diff --suppress-secrets
```

Check these pins against the recorded source overrides:

| Component | Target version |
| --- | --- |
| OpenBao chart | 0.32.6 |
| Server, injected agent, auto-unseal sidecar | 2.5.5-nv-1.3.1 |
| OpenBao migrations | 0.16.2 |
| Injector | 1.7.4 |

The StatefulSet must keep `OnDelete`. The catalog refresh Job must run before
migrations. Review changes to RBAC, hooks, webhook configuration, probes, and
pod templates. Stop if the diff changes a volume, identity, runtime image, or
credential that you did not intend to change.

## 5. Upgrade OpenBao

Both releases use migration image 0.16.2. The image tag does not show whether
Helm ran a new migration Job, so record the old Job UID:

```bash
old_migration_uid=$(kcp -n vault-system get job openbao-server-migrations \
-o jsonpath='{.metadata.uid}' --ignore-not-found)
```

Open a second terminal with the same environment variables and `kcp` helper.
Watch the hook Jobs and collect their logs while the sync runs. Helm can delete
a successful refresh Job, so capture its completion when it happens.

Run the bundle's selected sync in the first terminal:

```bash
make -C "$TARGET_CONTROL" install \
HELMFILE_ENV="$HELMFILE_ENV" \
KUBECONFIG_FILE="$CONTROL_KUBECONFIG" \
HELMFILE_SELECTOR=name=openbao-server
```

Check the hooks in this order:

1. `openbao-server-refresh-jwt-plugin-catalog` registers the digest from the
retained server image and completes successfully.
1. A new `openbao-server-migrations` Job starts after refresh completes. Its
UID must differ from `$old_migration_uid`.
1. The new migration Job and Helm sync complete successfully.

Catalog registration does not reload plugins in running servers. If the source
catalog digest already matches the retained server image and the server pod
template is unchanged, wait for migrations without rotating servers. This was
the path exercised in the healthy-source rehearsal.

A stale catalog or changed server image or pod template needs a separately
validated rotation and recovery sequence. Stop rather than assuming the
healthy-source result covers that case. If that sequence calls for rotation,
handle one server at a time, standbys first and the active member last. A
single-member deployment has downtime during rotation.

For each server selected for rotation:

```bash
export POD=<one-standby-pod>
old_pod_uid=$(kcp -n vault-system get pod "$POD" -o jsonpath='{.metadata.uid}')
kcp -n vault-system delete pod "$POD" --wait=true &&
kcp -n vault-system wait "pod/$POD" --for=create --timeout=5m &&
kcp -n vault-system wait "pod/$POD" --for=condition=Ready --timeout=10m &&
kcp -n vault-system exec "$POD" -c openbao -- sh -c \
'BAO_ADDR=http://127.0.0.1:8200 bao status' || {
echo 'Stop: replacement OpenBao pod did not pass health checks' >&2
exit 1
}
new_pod_uid=$(kcp -n vault-system get pod "$POD" -o jsonpath='{.metadata.uid}')
test -n "$new_pod_uid"
test "$new_pod_uid" != "$old_pod_uid"
```

Require a new pod UID, `Initialized true`, `Sealed false`, the expected image,
and healthy Raft membership before rotating another member. Readiness alone
does not prove that auto-unseal has finished. Recheck which member is active
before the last rotation. Never delete the StatefulSet, persistent-volume
claims, or multiple members together.

Confirm the migration Job and release status:

```bash
kcp -n vault-system wait job/openbao-server-migrations \
--for=condition=Complete --timeout=15m
helm --kubeconfig "$CONTROL_KUBECONFIG" --kube-context "$CONTROL_CONTEXT" \
status openbao-server -n vault-system
```

## 6. Validate the deployment

Check that every OpenBao member is healthy and unsealed, the catalog digest
matches the expected plugin, and the new migration Job succeeded. Then:

1. Exercise the existing JWT-backed authentication flow and invoke the original
function again.
1. Compare the original function, deployment, cluster, database records, and
worker identity with the baseline.
1. Admit a new test workload within the injector's configured webhook scope
and verify that it receives its secrets.
1. Review a full stack diff. After the OpenBao checks pass, apply the remaining
expected changes with the bundle's normal `make install` command.
1. Repeat the selected OpenBao sync and the full stack sync to check
idempotency. Repeat the health, JWT, injector, and invocation checks.

Record failed probes even if the final checks pass. Keep the 0.6.2 values and
backups as the source baseline for the [1.0.1 upgrade](./0.6.1-to-1.0.1-upgrade.md).

## Rehearsal results

A rehearsal upgraded released 0.6.1 to commit
`55803ad7cb647de45cf29415cfd3766225534617`, later published as stable 0.6.2,
on isolated ARM64 k3d. It used
three OpenBao members, one Cassandra member, and fake GPUs on one host.
Optional add-ons and observability were disabled, and local registry and
resource settings were adapted for the test.

Refresh completed before the new migration Job started. The server templates
and runtime images were unchanged, and all three members kept their pod
identities without rotation. Fresh JWT issuance, new secret injection, full
syncs, and repeat syncs passed. Existing credentials, PVCs, function and
deployment records, clean Cassandra ledgers, and worker identities were retained.
Each of the function-read, deployment-read, and invocation probes recorded
218 successes with no failures.

This evidence covers the healthy-source path at that commit. It does not cover
stale-catalog recovery, backup restoration, optional-component upgrades, or
production failure domains.

## Stop and recover

Stop on a failed hook, sealed member, dirty migration ledger, JWT or injector
failure, missing record, or unexpected identity change. Save Helm history and
hook logs before retrying. Do not skip hooks to force a successful release
status.

The original deployment and bundle are the recovery baseline before state
changes. Once a hook changes OpenBao state, a chart rollback alone is not a
complete recovery plan. Validate the catalog and runtime combination before
retrying, or restore the coordinated backup with your site's tested procedure.
Preserve all persistent volumes and original credentials.
Loading
Loading