diff --git a/README.md b/README.md index 5e190b6..6ef5ee1 100644 --- a/README.md +++ b/README.md @@ -4,8 +4,9 @@ [Flyte 2](https://github.com/flyteorg/flyte) lets you write batch workflows as plain async Python. [Armada](https://github.com/armadaproject/armada) schedules millions of jobs a day across many -Kubernetes clusters, with fair-share, gang scheduling, and preemption. `armada-flyte` connects the -two: your Flyte task runs as an Armada job, with one line of config and no new API to learn. +Kubernetes clusters, with fair-share, gang scheduling (running a group of pods all-or-nothing), and +preemption. `armada-flyte` connects the two: your Flyte task runs as an Armada job, with one line of +config and no new API to learn. ## The whole integration @@ -25,10 +26,11 @@ async def greet(name: str) -> str: return f"hello {name}, from an Armada pod" # runs in an Armada-scheduled pod ``` -A stock `@env.task` and one `plugin_config` line. Fan out with `asyncio.gather`, pass dataclasses -between tasks, gang-schedule a group: it is all just Flyte, running on Armada. +The `queue` names the Armada queue the job goes into, its fair-share bucket. A stock `@env.task` and +one `plugin_config` line. Fan out with `asyncio.gather`, pass dataclasses between tasks, or gang-schedule +a group all-or-nothing with `armada_flyte.Gang`: it is all just Flyte, running on Armada. -The connector submits to the Armada at `ARMADA_URL` (default `localhost:50051`). Point it at a +The connector submits to Armada at `ARMADA_URL` (default `localhost:50051`). Point it at a remote cluster by setting that env var, or in code: ```python @@ -41,18 +43,22 @@ never lands in the control plane. See [docs/getting-started.md](docs/getting-sta ## See it run -With a local Armada cluster up, the [demo](demo/) stands up a Flyte backend and the connector in one -command, then you submit the task: +`./hack/up.sh` brings up Armada and Flyte in one local kind cluster (see the +[quickstart](demo/README.md)). Submitting a task then returns its typed result to your terminal, and the +run appears in the Flyte UI: ```console -$ ./demo/setup.sh $ ./.venv/bin/python examples/hello.py submitted run rxc4nspfkjqr5px6q9nj - UI: http://localhost:30080/v2/.../runs/rxc4nspfkjqr5px6q9nj + Flyte console: http://localhost:5001/v2/.../runs/rxc4nspfkjqr5px6q9nj + Armada Lookout: http://localhost:30000 + +hello armada, from an Armada pod ``` -The run shows up in the Flyte UI, scheduled and executed by Armada. See -[getting started](docs/getting-started.md) for the walkthrough. +This is an integration between two systems, so there is real setup, but `./hack/up.sh` does it in one +command: the [quickstart](demo/README.md) covers what it stands up and how to run this. From there, +[examples/](examples/) ladder from `hello` up to a gang inside a DAG. ## Why both @@ -62,7 +68,7 @@ The run shows up in the Flyte UI, scheduled and executed by Armada. See | The Flyte console: runs, lineage, logs | Fair-share between queues, gang scheduling, preemption | | Local execution for fast iteration | Battle-tested at millions of jobs a day | -You keep Flyte's authoring and console; Armada does the scheduling. No rewrite, no second SDK. +You keep Flyte's authoring and console, and Armada does the scheduling. No rewrite, no second SDK. ## How it works @@ -83,12 +89,14 @@ Flyte UI. ## Where to go next -- **Run it locally.** [demo/](demo/) stands up the backend and connector in one command (the - [See it run](#see-it-run) commands above). +- **Run it locally.** `./hack/up.sh` brings up Armada, Flyte, and the connector in one kind cluster; + the [quickstart](demo/README.md) covers it. - **Write tasks.** Start from [examples/hello.py](examples/hello.py), then [examples/](examples/) for fan-out, gang scheduling, and a gang inside a DAG. - **Run against your own backend.** [Getting started](docs/getting-started.md) covers installing the - connector, running it as a service, and building the task image. + connector, running it as a service, and building the task image. The examples read `FLYTE_ENDPOINT` + (the Flyte API) and `FLYTE_UI_BASE` (the console link they print) from the environment, both + defaulting to the local devbox, so point them at your backend there. - **Understand the internals.** [How it works](docs/architecture.md): the connector, state mapping, and gang scheduling. - **Deploy the connector as a service.** [deploy/](deploy/). diff --git a/demo/README.md b/demo/README.md index a7d8117..8e66757 100644 --- a/demo/README.md +++ b/demo/README.md @@ -1,106 +1,145 @@ -# Local quickstart +# Quickstart -Run a Flyte 2 backend in the same Kind cluster Armada already schedules onto, then submit a task. -This is the one-command way to see the connector work end to end. +`armada-flyte` runs Flyte tasks as Armada jobs, so trying it means bringing up both. `./hack/up.sh` +does that in one command: a local kind cluster with Armada (operator, dependencies, and the component +CRs), Traefik, the Flyte 2 backend, the connector, and the `flyte` queue, all wired together. ``` - host Kind cluster "armada-test" - ┌────────────┐ localhost:30080 ┌─────────────────────────────────────────┐ - │ Flyte CLI │ ─────────────────▶│ flyte-binary (API + console + TaskAction │ - │ (examples) │ localhost:30900 │ reconciler) │ - ├────────────┤ ─────────────────▶│ postgres (metadata + runs) │ - │ connector │◀──── :8000 ───────│ minio (blob store) │ - │ c0 │ ──┐ │ armada- pods ──in-cluster──▶ minio │ - └────────────┘ │ :50051 └─────────────────────────────────────────┘ - └──▶ Armada control plane (docker, from `dev:full`) + host kind cluster "armada" + ┌──────────┐ localhost:30080 ┌───────────────────────────────────────────┐ + │ examples │ ─────────────────▶│ flyte-binary (API + console) │ + │ (client) │ │ armada-flyte-connector ──▶ armada-server │ + │ browser │ localhost:5001 │ minio (blob store), postgres (metadata) │ + │ │ ─────────────────▶│ armada (operator): server, scheduler, │ + └──────────┘ │ executor ──▶ job pods │ + └───────────────────────────────────────────┘ ``` -`flyte-binary`, minio, and postgres run in the cluster. The connector `c0` runs on the host and is -the only process bridging Flyte and Armada. The Armada pods reach minio in-cluster at -`minio.flyte:9000`, and `setup.sh` port-forwards the Flyte API and minio to `localhost` so the client -reaches them without any kind port mapping. +Everything runs in the cluster. `hack/kind-config.yaml` maps every host-facing port to a NodePort, so +the client (the examples) and your browser reach the API and console straight from the host with no +port-forwards. Every hop between components is cluster DNS. ## Prerequisites -1. **Armada**, with a real executor against the Kind cluster and the `flyte` queue created. Clone - [armada](https://github.com/armadaproject/armada), then: - ``` - go run github.com/magefile/mage@v1.17.2 dev:full # kind "armada-test" + full stack - go run cmd/armadactl/main.go create queue flyte --armadaUrl localhost:50051 - ``` - Wait for the executor to log `Reporting current free resource` before submitting. -2. **This repo's venv** (an arm64 Python on Apple Silicon, since Flyte's `obstore` wheel has no x86 build): - ``` - python3.11 -m venv .venv && ./.venv/bin/pip install -e . - ``` -3. `docker`, `helm`, `kubectl`, `kind` on PATH. - -`setup.sh` pulls the published flyte-binary chart from the flyteorg helm repo and a stock upstream -`cr.flyte.org/flyteorg/flyte-binary-v2` image, so no Flyte checkout is needed. The pinned image is the -first commit that registers the connector-service plugin in the v2 executor -([flyteorg/flyte#7565](https://github.com/flyteorg/flyte/pull/7565)). Earlier stock builds do not have -it. - -## Run +- `docker` (running), `helm`, `kubectl`, `kind`, and `python3` (3.10+) on your PATH. `up.sh` checks + these before it does any work. +- `armadactl` is fetched to `~/bin` if it is not already on your PATH; add `~/bin` to your PATH if the + command is not found afterwards. + +## Bring it up ``` -export KUBECONFIG=/.kube/external/config # the kubeconfig `dev:full` writes -./demo/setup.sh +./hack/up.sh +./.venv/bin/python examples/hello.py ``` -`setup.sh` deploys minio + postgres, installs the flyte-binary chart, builds and loads the task -image, starts the connector, and port-forwards the Flyte API and minio to `localhost`. Then submit an -example. It points at `localhost:30080` and targets queue `flyte`: +`up.sh` creates the kind cluster; installs Armada and its dependencies (cert-manager, Pulsar, Postgres, +Redis) via the operator and creates the `flyte` queue; builds the repo's virtualenv; then deploys +Traefik and the Flyte 2 backend (`flyte-binary`), builds and loads the task image, and deploys the +connector. The +first run pulls a lot of images before it settles. It waits for Armada to register executor capacity, so +when it prints `Devbox up` the integration is ready to submit: ``` -./.venv/bin/python examples/hello.py +Devbox up. Run an example: + ./.venv/bin/python examples/hello.py + ./.venv/bin/python examples/dag.py ``` -Open the printed `UI:` link to watch Armada schedule the pod and record the typed result. Other -examples work the same way (`examples/fanout.py`, `examples/gang.py`, `examples/dag.py`). +`hello.py` prints its result to your terminal (`hello armada, from an Armada pod`), and the run appears +in the Flyte console at the link it prints. That is the whole loop: a Flyte task, scheduled and run by +Armada. -## Files +## The UIs -| File | What it is | -|------|-----------| -| `setup.sh` | One-command stand-up of the backend + connector. Idempotent. | -| `minio.yaml` | Blob store. Pods reach it in-cluster. The host reaches it via a port-forward `setup.sh` starts. | -| `postgres.yaml` | Flyte's `flyte` (metadata) and `runs` (run-graph) databases. | -| `flyte-binary-values.yaml` | Chart overrides. `__HOST_IP__` (the connector endpoint) is filled in by `setup.sh`. | +Each example prints two links: -## Overrides +- **Flyte console** (`http://localhost:5001/v2`) - the run graph, status, and pod logs. Served by + Traefik through the flyte-binary chart's own ingress. +- **Armada Lookout** (`http://localhost:30000`) - the job and pod as Armada sees them. + +Per-task input/output values do not render in the Flyte console yet. The Flyte 2 console is pre-release +and its I/O panel is incomplete upstream. The values are recorded correctly: each example prints its +result, and the data is in the blob store. + +## Work through the examples + +Once `hello.py` runs, the [examples](../examples/) build up in order: -`setup.sh` reads these env vars: `KIND_CLUSTER` (default `armada-test`), `FLYTE_CHART` (default -`flyteorg/flyte-binary`, or a local directory for an unreleased chart), `FLYTE_CHART_VERSION` -(default `v2.0.27`), `FLYTE_IMAGE` (default the pinned stock `flyte-binary-v2` build), `TASK_IMAGE` -(default `armada-flyte-task:v1`), `ARMADA_URL` (default `localhost:50051`), `HOST_IP` (auto-detected). +1. [`function.py`](../examples/function.py) - one task doing real work (a Black-Scholes price). +2. [`fanout.py`](../examples/fanout.py) - a typed dataclass through a parallel fan-out / fan-in. +3. [`gang.py`](../examples/gang.py) - N co-scheduled workers as one Armada gang (all-or-nothing). +4. [`dag.py`](../examples/dag.py) - the full shape: generate a dataset, run a gang over it, aggregate. + +Each runs the same way: `./.venv/bin/python examples/.py`. ## Teardown ``` -helm -n flyte uninstall flyte-binary -kubectl delete namespace flyte -pkill -f "bin/c0 --port 8000" -pkill -f "port-forward svc/flyte-binary-http" -pkill -f "port-forward svc/minio" +./hack/down.sh ``` -The kind cluster and Armada come down with `go run github.com/magefile/mage@v1.17.2 dev:fullDown` in -the armada repo. +`down.sh` deletes the kind cluster and clears the Flyte client's upload cache. The cache is keyed by the +endpoint (`localhost:30080`, the same for every devbox), so clearing it here keeps the next `up.sh` from +reusing entries that point at the deleted cluster's minio and 404-ing on their code bundle. + +## Overrides + +`up.sh` reads `KIND_CLUSTER` (default `armada`). `demo/setup.sh`, which it calls for the Flyte side, +reads `FLYTE_CHART` (default `flyteorg/flyte-binary`, or a local directory for an unreleased chart), +`FLYTE_CHART_VERSION` (default `v2.0.27`), `FLYTE_IMAGE` (default the pinned stock `flyte-binary-v2` +build; this only pre-loads the image into kind — the tag the pod runs is pinned in +`flyte-binary-values.yaml`, so change both together), `TASK_IMAGE` (default `armada-flyte-task:v1`), +and `CONNECTOR_IMAGE` (default +`gresearch/armada-flyte-connector:0.2.0`). The pinned Flyte image is the first that registers the +connector plugin in the executor +([flyteorg/flyte#7565](https://github.com/flyteorg/flyte/pull/7565)). + +## Manual setup without the devbox script + +`up.sh` is the supported path. If you want the steps by hand, or to submit to an Armada cluster you +already run, the pieces are: + +1. **Armada.** `hack/setup-armada.sh` installs Armada into the current kind context (operator, + dependencies, and the CRs in `hack/armada/`). The stock + [armada-operator](https://github.com/armadaproject/armada-operator) `make kind-all` also works, but + its kind config maps only ports 30000-30002, so the Flyte console on 5001 will not be reachable from + the host - create the cluster from `hack/kind-config.yaml` for that. +2. **Flyte.** `./demo/setup.sh` deploys Traefik, the `flyte-binary` backend (with its minio blob store + and postgres), builds the task image, and deploys the connector. +3. **Queue.** `armadactl create queue flyte`. + +## Files + +| File | What it is | +|------|-----------| +| `setup.sh` | Stands up the Flyte side (Traefik + backend + connector) in the cluster. Idempotent. | +| `minio.yaml` | Blob store. Pods reach it in-cluster; the host reaches it via the kind NodePort mapping. | +| `postgres.yaml` | Flyte's `flyte` (metadata) and `runs` (run-graph) databases. | +| `flyte-binary-values.yaml` | Chart overrides: storage, the chart ingress, and the connector routing. | ## Troubleshooting -- **`no connector found for task type [armada]`**: the connector config did not reach the running +- **`no connector found for task type [armada]`**: the connector routing did not reach the running config. It must live under `configuration.inline.plugins.connector-service` with - `supportedTaskTypes: [armada]`. The chart's top-level `configuration.connectorService` is not wired - in. Confirm with `kubectl -n flyte get cm flyte-binary-config -o yaml | grep -A4 connector-service`. -- **`TaskAction` stuck `Queued`**: Armada has no executor yet, or the `flyte` queue is missing. - Check the connector log (`/tmp/armada-flyte-c0.log`) and `armadactl get queue flyte`. -- **`localhost:30080` connection refused**: the port-forward is not running. Re-run `setup.sh`, or - start it by hand: `kubectl -n flyte port-forward svc/flyte-binary-http 30080:8090`. + `supportedTaskTypes: [armada]`; the chart's top-level `configuration.connectorService` is not wired in. + Confirm with `kubectl -n flyte get cm flyte-binary-config -o yaml | grep -A4 connector-service`. +- **Task stuck `Queued`**: Armada has no executor yet, or the `flyte` queue is missing. Check + `kubectl -n armada get pods` and `armadactl get queue flyte`. +- **A gang (`gang.py` / `dag.py`) is `REJECTED` or stuck `Queued`**: on your own Armada, the gang's + node-uniformity label must be tracked by the executor and indexed by the scheduler, or the scheduler + cannot place the gang. Confirm `kubernetes.io/hostname` is in the executor's + `kubernetes.trackedNodeLabels` and the scheduler's `scheduling.indexedNodeLabels`. The devbox CRs + (`hack/armada/armada-crs.yaml`) set both already. Also make sure the whole gang fits: with + `kubernetes.io/hostname` uniformity all members land on one node, so the sum of their requests must + fit that node's free capacity. +- **`localhost:30080` connection refused**: the cluster is not up, or was created without + `hack/kind-config.yaml`, so the NodePort is not mapped to the host. Bring the devbox up with + `./hack/up.sh`. +- **Flyte console shows "No actions found"**: open the exact `http://localhost:5001/v2/...` link the + example prints. The console reaches the API through Traefik on the same origin. - **Pod `Completed` but the run fails resolving outputs**: the pod could not reach minio in-cluster. - Confirm the connector was started with `FLYTE_BLOB_ENDPOINT=http://minio.flyte.svc.cluster.local:9000` - and that the minio credentials match `minio.yaml`. -- **Pod fails with a 404 on `fast.tar.gz` after a re-run**: you recreated the blob store (a - fresh minio) but the client cached the previous upload and skipped re-uploading the code bundle to - the new one. Clear the cache and rerun: `rm -f ~/.flyte/local-cache/cache.db`. + Confirm the connector's `FLYTE_BLOB_ENDPOINT` is `http://minio.flyte.svc.cluster.local:9000` and the + minio credentials match `minio.yaml`. +- **Pod fails with a 404 on `fast.tar.gz` after a re-run**: you recreated the blob store but the + client cached the previous upload. Tear down with `./hack/down.sh`, which clears the cache. diff --git a/demo/flyte-binary-values.yaml b/demo/flyte-binary-values.yaml index 9a380ba..40628cb 100644 --- a/demo/flyte-binary-values.yaml +++ b/demo/flyte-binary-values.yaml @@ -1,12 +1,7 @@ -# flyte-binary (v2) values for the local quickstart. +# flyte-binary (v2) chart values for the local quickstart. # -# Derived from the flyte-devbox config, with the bundled rustfs/postgres swapped for the external -# minio.yaml and postgres.yaml in this directory, and the connector pointed at the host-run armada -# connector (c0). -# -# The __HOST_IP__ placeholder (in configuration.inline.plugins.connector-service.defaultConnector.endpoint) -# is filled at install time with the host's LAN IP (setup.sh). The flyte-binary pods dial the connector, -# which is a process on the host, so they need its LAN address rather than a cluster DNS name. +# Points storage at the minio.yaml and postgres.yaml in this directory, and routes the `armada` task type +# to the in-cluster armada-flyte-connector Service, which the flyte-binary pods reach over cluster DNS. fullnameOverride: flyte-binary @@ -24,12 +19,6 @@ deployment: tag: 16-alpine pullPolicy: IfNotPresent -console: - image: - repository: ghcr.io/unionai-oss/flyteconsole-v2 - tag: latest - pullPolicy: IfNotPresent - configuration: database: postgres: @@ -81,21 +70,21 @@ configuration: gpu: 0 # Signed-URL endpoint the Flyte client (on the host) uses to upload fast-register bundles and read - # outputs. setup.sh makes localhost:30900 reach minio, through a kind port mapping or a port-forward. + # outputs. localhost:30900 reaches minio through the kind NodePort mapping. storage: signedURL: stowConfigOverride: endpoint: http://localhost:30900 plugins: - # The connector that armada tasks are delegated to. Runs on the host (c0), reachable from in-cluster - # Flyte at the host LAN IP. supportedTaskTypes is what makes the plugin match the `armada` task - # type to this connector. Without it, Flyte reports "no connector found for task type [armada]". - # This must live under inline.plugins because the chart's top-level configuration.connectorService - # is not wired into the config. + # The connector that armada tasks are delegated to. Runs in-cluster as the armada-flyte-connector + # Deployment/Service (deploy/kubernetes/connector.yaml), reached over cluster DNS. supportedTaskTypes + # is what makes the plugin match the `armada` task type to this connector. Without it, Flyte reports + # "no connector found for task type [armada]". This must live under inline.plugins because the chart's + # top-level configuration.connectorService is not wired into the config. connector-service: defaultConnector: - endpoint: "dns:///__HOST_IP__:8000" + endpoint: "armada-flyte-connector.flyte.svc.cluster.local:8000" insecure: true supportedTaskTypes: - armada @@ -119,27 +108,27 @@ configuration: k8sConfig: namespace: flyte -# Route the "armada" task type to the connector, and keep the stock plugins for the rest. +# Route the "armada" task type to the connector. Helm merges this map over the chart's defaults, so +# the stock plugin list and the container/sidecar routes stay as shipped. enabled_plugins: tasks: task-plugins: - enabled-plugins: - - container - - sidecar - - connector-service - - echo default-for-task-types: - container: container - sidecar: sidecar - container_array: k8s-array armada: connector-service -# Expose the Flyte API on a fixed NodePort. setup.sh port-forwards localhost:30080 to it. +# Expose the Flyte API on a fixed NodePort, reachable at localhost:30080 via the kind mapping, where +# the examples submit. service: type: NodePort nodePorts: http: 30080 +# Serve the console (/v2) and the flyteidl2 API through the chart's own ingress, so the browser reaches +# both from one origin. Traefik (the ingress controller setup.sh installs) routes them. +ingress: + create: true + ingressClassName: traefik + # Flyte needs broad RBAC in a dev sandbox to create per-namespace resources. rbac: extraRules: diff --git a/demo/setup.sh b/demo/setup.sh index 2d4d214..4f01e92 100755 --- a/demo/setup.sh +++ b/demo/setup.sh @@ -1,63 +1,59 @@ #!/usr/bin/env bash -# Stand up a Flyte 2 backend in the same Kind cluster Armada already uses, and wire the connector. +# Stand up the Flyte 2 side in the kind cluster Armada runs in, and wire the connector. # # What it does: # 1. deploys minio (blob store) + postgres (Flyte metadata) into the kind cluster, -# 2. installs the flyte-binary (v2) chart, pointed at that minio/postgres and at the host connector, -# 3. builds the task image from this checkout and loads it into the cluster, -# 4. starts the connector (c0) on the host, pointed at Armada and the in-cluster minio, -# 5. makes the Flyte API and blob store reachable from the host (port-forward, or a kind port map). +# 2. installs Traefik (the ingress controller that serves the Flyte console), +# 3. installs the flyte-binary (v2) chart, pointed at that minio/postgres and the in-cluster connector, +# 4. builds the task image from this checkout and loads it into the cluster, +# 5. deploys the connector in-cluster, pointed at the in-cluster Armada and minio. +# The host reaches the API, blob store, and console through the kind-config.yaml NodePort mappings. # # Prerequisites (see README.md): -# - Armada up with a real executor against the kind cluster $KIND_CLUSTER, and queue "flyte" created -# (in the armada repo: `go run github.com/magefile/mage@v1.17.2 dev:full`, then create queue "flyte") -# - this repo's venv built (`python3.11 -m venv .venv && ./.venv/bin/pip install -e .`) +# - An Armada cluster in the kind cluster $KIND_CLUSTER, with a queue "flyte" (hack/setup-armada.sh +# installs it). This installs the Flyte side into the same cluster. # - docker, helm, kubectl, kind on PATH +# (Running the examples afterwards needs this repo's venv; hack/up.sh builds it.) set -euo pipefail cd "$(dirname "$0")" ROOT="$(cd .. && pwd)" -KIND_CLUSTER="${KIND_CLUSTER:-armada-test}" +KIND_CLUSTER="${KIND_CLUSTER:-armada}" # Stock upstream Flyte 2 binary at the first commit that registers the connector-service plugin -# (flyteorg/flyte#7565). Keep this in sync with the image tag in flyte-binary-values.yaml. +# (flyteorg/flyte#7565). This var only controls the kind preload below; the tag the pod actually +# runs is pinned in flyte-binary-values.yaml, so change the two together. FLYTE_IMAGE="${FLYTE_IMAGE:-cr.flyte.org/flyteorg/flyte-binary-v2:sha-d9e0ebe97be436c7c03c13a8243d3b399d1729e7}" TASK_IMAGE="${TASK_IMAGE:-armada-flyte-task:v1}" +# The connector image tag. The devbox builds it locally from deploy/Dockerfile (below) so it always +# runs the current source and needs no registry. CI publishes this same name on a GitHub release. +CONNECTOR_IMAGE="${CONNECTOR_IMAGE:-gresearch/armada-flyte-connector:0.2.0}" # Published flyte-binary chart from the flyteorg helm repo. Set FLYTE_CHART to a local path (e.g. a # flyteorg/flyte checkout) to use an unreleased chart. A local path skips the version pin. FLYTE_CHART="${FLYTE_CHART:-flyteorg/flyte-binary}" FLYTE_CHART_VERSION="${FLYTE_CHART_VERSION:-v2.0.27}" -HOST_IP="${HOST_IP:-$(ipconfig getifaddr en0 2>/dev/null || hostname -I | awk '{print $1}')}" -C0="$ROOT/.venv/bin/c0" -# Make a host port reach an in-cluster service. If something already serves it (a kind port mapping -# or an earlier forward), reuse that. Otherwise start a background kubectl port-forward. -ensure_host_port() { - local port=$1 svc=$2 target=$3 - if nc -z localhost "$port" 2>/dev/null; then - echo " localhost:$port already reachable (reusing)" - return - fi - nohup kubectl -n flyte port-forward "svc/$svc" "$port:$target" >"/tmp/af-pf-$port.log" 2>&1 & - for _ in $(seq 1 15); do - nc -z localhost "$port" 2>/dev/null && break - sleep 1 - done - echo " port-forward localhost:$port -> $svc:$target" -} - -echo "==> using kind cluster '$KIND_CLUSTER', host IP $HOST_IP" -# The armada Kind target writes a repo-local kubeconfig (.kube/external/config) rather than merging -# into ~/.kube/config. Honour an already-set KUBECONFIG. Otherwise select the kind context. -if [ -z "${KUBECONFIG:-}" ]; then - kubectl config use-context "kind-${KIND_CLUSTER}" >/dev/null -fi +echo "==> using kind cluster '$KIND_CLUSTER'" +# Target this cluster's kind context explicitly on every kubectl/helm call, so the script never switches +# the user's current context and never installs into whatever cluster KUBECONFIG happens to point at. +CTX="kind-${KIND_CLUSTER}" +kc() { kubectl --context "$CTX" "$@"; } echo "==> 1/5 blob store + metadata database" -kubectl apply -f minio.yaml -f postgres.yaml >/dev/null -kubectl -n flyte rollout status deploy/postgres --timeout=120s >/dev/null -kubectl -n flyte rollout status deploy/minio --timeout=120s >/dev/null +kc apply -f minio.yaml -f postgres.yaml >/dev/null +kc -n flyte rollout status deploy/postgres --timeout=120s >/dev/null +kc -n flyte rollout status deploy/minio --timeout=120s >/dev/null -echo "==> 2/5 install flyte-binary (v2)" +echo "==> 2/5 ingress controller (traefik)" +# Traefik serves the flyte-binary chart's ingress, which routes /v2 to the console and the flyteidl2 +# API paths to flyte-binary-http, same origin. Its web entrypoint is a NodePort the kind config maps +# to localhost:5001, so the console is browsed there. +helm repo add traefik https://traefik.github.io/charts >/dev/null 2>&1 || true +helm repo update traefik >/dev/null 2>&1 || true +helm --kube-context "$CTX" upgrade --install traefik traefik/traefik -n traefik --create-namespace \ + --set service.type=NodePort --set ports.web.nodePort=30500 >/dev/null +kc -n traefik rollout status deploy/traefik --timeout=120s >/dev/null + +echo "==> 3/5 install flyte-binary (v2)" # Preload the backend image so the pod uses it without a registry round trip. Falls through to a # normal pull if the image is not in local docker (the default stock image is pulled from cr.flyte.org). kind load docker-image "$FLYTE_IMAGE" --name "$KIND_CLUSTER" >/dev/null 2>&1 || true @@ -69,18 +65,16 @@ if [ ! -d "$FLYTE_CHART" ]; then helm repo update flyteorg >/dev/null 2>&1 || true chart_version_arg=(--version "$FLYTE_CHART_VERSION") fi -rendered="$(mktemp)" -sed "s/__HOST_IP__/${HOST_IP}/g" flyte-binary-values.yaml > "$rendered" -helm upgrade --install flyte-binary "$FLYTE_CHART" "${chart_version_arg[@]}" -n flyte -f "$rendered" >/dev/null -rm -f "$rendered" -kubectl -n flyte rollout status deploy/flyte-binary --timeout=180s >/dev/null +helm --kube-context "$CTX" upgrade --install flyte-binary "$FLYTE_CHART" "${chart_version_arg[@]}" -n flyte -f flyte-binary-values.yaml >/dev/null +kc -n flyte rollout status deploy/flyte-binary --timeout=180s >/dev/null +kc -n flyte rollout status deploy/flyte-binary-console --timeout=120s >/dev/null -echo "==> 3/5 task image -> $TASK_IMAGE" +echo "==> 4/5 task image -> $TASK_IMAGE" build="$(mktemp -d)" mkdir -p "$build/pkg" cp -R "$ROOT/pyproject.toml" "$ROOT/README.md" "$ROOT/src" "$build/pkg/" cat > "$build/Dockerfile" </dev/null rm -rf "$build" kind load docker-image "$TASK_IMAGE" --name "$KIND_CLUSTER" >/dev/null -echo "==> 4/5 connector (c0) on the host" -# The connector hands each Armada pod the in-cluster minio address, which the pod reaches directly -# through cluster DNS. Nothing on the job path depends on a host-published port. -if lsof -nP -iTCP:8000 -sTCP:LISTEN >/dev/null 2>&1; then - echo " connector already on :8000 (reusing; 'pkill -f \"bin/c0 --port 8000\"' to restart)" -else - ARMADA_URL="${ARMADA_URL:-localhost:50051}" \ - FLYTE_BLOB_ENDPOINT="http://minio.flyte.svc.cluster.local:9000" \ - FLYTE_BLOB_ACCESS_KEY=minio FLYTE_BLOB_SECRET_KEY=minio12345 \ - nohup "$C0" --port 8000 --prometheus_port 9099 >/tmp/armada-flyte-c0.log 2>&1 & - for _ in $(seq 1 15); do - grep -aq "armada (0)" /tmp/armada-flyte-c0.log 2>/dev/null && break - sleep 1 - done - echo " connector ready (log: /tmp/armada-flyte-c0.log)" -fi - -echo "==> 5/5 host access to the Flyte API and blob store" -# The Flyte client (the examples) and signed-URL uploads run on the host, so they need localhost to -# reach the in-cluster API and minio. These forwards run in the background until you kill them. -ensure_host_port 30080 flyte-binary-http 8090 -ensure_host_port 30900 minio 9000 +echo "==> 5/5 connector (build + deploy in-cluster)" +# Build the connector from source and load it into kind, so the devbox always runs the current source +# with no registry. connector.yaml pulls this image name only IfNotPresent, so the loaded build is used. +docker build -t "$CONNECTOR_IMAGE" -f "$ROOT/deploy/Dockerfile" "$ROOT" >/dev/null +kind load docker-image "$CONNECTOR_IMAGE" --name "$KIND_CLUSTER" >/dev/null +kc apply -f "$ROOT/deploy/kubernetes/connector.yaml" >/dev/null +kc -n flyte rollout status deploy/armada-flyte-connector --timeout=120s >/dev/null # Flyte's connector-service plugin does not retry a failed CreateTask, so the first task submitted # must not race the backend's connection to the connector. The backend polls the connector for its # task types every ~10s. Wait until it logs a successful discovery of the armada connector before # declaring the backend ready, otherwise the first submit can fail with "connection refused". echo "==> waiting for the backend to reach the connector" -fb=$(kubectl -n flyte get pods -l app.kubernetes.io/name=flyte-binary --field-selector=status.phase=Running -o name 2>/dev/null | head -1) -for _ in $(seq 1 45); do - kubectl -n flyte logs "$fb" -c flyte --since=25s 2>/dev/null | grep -aq "supports the following task types: \[armada\]" && break - sleep 2 -done +fb=$(kc -n flyte get pods -l app.kubernetes.io/name=flyte-binary --field-selector=status.phase=Running -o name 2>/dev/null | head -1) +discovered="" +if [ -n "$fb" ]; then + for _ in $(seq 1 45); do + if kc -n flyte logs "$fb" -c flyte --since=25s 2>/dev/null | grep -aq "supports the following task types: \[armada\]"; then + discovered=1 + break + fi + sleep 2 + done +fi +if [ -z "$discovered" ]; then + echo " warning: the backend has not logged connector discovery yet; the first submit may fail" >&2 + echo " with 'connection refused'. If it does, wait a few seconds and resubmit." >&2 +fi echo -echo "Flyte backend up. UI: http://localhost:30080/v2 API: localhost:30080" +echo "Flyte backend up." +echo " API (examples submit here): localhost:30080" +echo " Flyte console (run graph): http://localhost:5001/v2" +echo " Armada Lookout (job status): http://localhost:30000" echo "Submit an example: $ROOT/.venv/bin/python $ROOT/examples/hello.py" diff --git a/examples/README.md b/examples/README.md index 665e8e4..6bb38ac 100644 --- a/examples/README.md +++ b/examples/README.md @@ -1,10 +1,11 @@ # Examples Write ordinary Flyte 2 Python. Each `@env.task` runs in an Armada-scheduled pod. The only -Armada-specific line is `plugin_config=ArmadaConfig(queue=...)`; everything else (resources, +Armada-specific line is `plugin_config=ArmadaConfig(queue=...)`. Everything else (resources, chaining, fan-out, typed data) is stock Flyte. -Five examples, in order: +Five examples that build up in order. Start at the top and work down: each adds one idea, ending with +a gang inside a DAG. | File | Shows | Expected output | | --- | --- | --- | @@ -16,14 +17,16 @@ Five examples, in order: ## Run one -The runner submits the example through the Flyte backend, so the run shows up in the Flyte UI: +The runner submits the example through the Flyte backend. Each example prints its typed result and a +link to the run in the Flyte UI: ```bash ./.venv/bin/python examples/hello.py ``` -Pass any example as the argument. Prerequisite: a running Armada cluster and a Flyte 2 backend (see -[../docs/getting-started.md](../docs/getting-started.md)). +Run any example the same way. Prerequisite: a Flyte 2 backend with the connector. The +[quickstart](../demo/) walks the setup end to end, or point at your own +([../docs/getting-started.md](../docs/getting-started.md)). ## What you write @@ -45,4 +48,14 @@ async def greet(name: str) -> str: Resources are required (Armada rejects a job without them). Declare them with `flyte.Resources` on the environment, or per task with `@env.task(resources=...)`. Need a GPU? `flyte.Resources(gpu=1)`. -`_runner.py` is the shared helper the examples call to run; it is not an example itself. +`_runner.py` is the shared helper the examples call to run. It is not an example itself. + +To gang-schedule a group of tasks all-or-nothing, add them to an `armada_flyte.Gang` and `await` it; +the gang id and cardinality are derived from the members, so there is nothing to keep in sync: + +```python +from armada_flyte import Gang + +# one worker per item, all co-scheduled or none (see gang.py). dag.py shows the add()/run() form +parts = await Gang.map(worker, range(4), node_uniformity_label="kubernetes.io/hostname") +``` diff --git a/examples/_runner.py b/examples/_runner.py index b6af293..0f6d3d7 100644 --- a/examples/_runner.py +++ b/examples/_runner.py @@ -2,20 +2,34 @@ block only, so the task pod (which imports the example module to load the task, but never runs __main__) never needs to import this file. -It submits the example through the Flyte backend (via ``flyte.init`` at localhost:30080) and waits -for the typed result. The connector routes the task to Armada. Adjust the endpoint for your backend. +It submits the example through the Flyte backend (``flyte.init`` at ``$FLYTE_ENDPOINT``, default +``localhost:30080``, where the demo's kind NodePort mapping exposes the API) and waits for the typed result. The +connector routes the task to Armada. The example prints the result, and this helper also prints a +link to the run in the Flyte UI. """ from __future__ import annotations +import os + import flyte import flyte.remote +PROJECT = "flytesnacks" +DOMAIN = "development" + +# The demo serves the Flyte console at http://localhost:5001/v2 and Armada's Lookout at +# http://localhost:30000. Override FLYTE_UI_BASE / ARMADA_LOOKOUT for a backend that is not the demo. +UI_BASE = os.environ.get("FLYTE_UI_BASE", "http://localhost:5001/v2").rstrip("/") +LOOKOUT = os.environ.get("ARMADA_LOOKOUT", "http://localhost:30000") + def run(entrypoint, **inputs): """Submit entrypoint through the Flyte backend, returning its first output.""" - flyte.init(endpoint="localhost:30080", insecure=True, project="flytesnacks", domain="development") + endpoint = os.environ.get("FLYTE_ENDPOINT", "localhost:30080") + flyte.init(endpoint=endpoint, insecure=True, project=PROJECT, domain=DOMAIN) r = flyte.run(entrypoint, **inputs) - print(f"\nsubmitted run {r.name}\n UI: {r.url}") + url = f"{UI_BASE}/domain/{DOMAIN}/project/{PROJECT}/runs/{r.name}" + print(f"\nsubmitted run {r.name}\n Flyte console: {url}\n Armada Lookout: {LOOKOUT}") r.wait() return flyte.remote.Run.get(r.name).outputs()[0] diff --git a/examples/fanout.py b/examples/fanout.py index d56d29e..0f5659e 100644 --- a/examples/fanout.py +++ b/examples/fanout.py @@ -1,4 +1,4 @@ -"""Various: a typed multi-stage pipeline with a parallel fan-out / fan-in. +"""Typed multi-stage pipeline with a parallel fan-out / fan-in. generate(n) produce n numbers split into S shards diff --git a/examples/function.py b/examples/function.py index 66fcc6a..aed216a 100644 --- a/examples/function.py +++ b/examples/function.py @@ -1,4 +1,4 @@ -"""Simple: one @env.task that runs in an Armada pod. Write normal typed Python, it runs on Armada. +"""One @env.task that runs in an Armada pod. Write normal typed Python, it runs on Armada. The only Armada-specific line is plugin_config=ArmadaConfig(queue=...). Resources are declared the stock-Flyte way via flyte.Resources. Run: diff --git a/examples/hello.py b/examples/hello.py index 5fc10be..a4ad06a 100644 --- a/examples/hello.py +++ b/examples/hello.py @@ -1,6 +1,9 @@ """Hello world: the smallest Armada task. One @env.task, one line of Armada config. ./.venv/bin/python examples/hello.py # runs on Armada, shows in the Flyte UI + +Next: function.py does real work, then fanout.py, gang.py, and dag.py build up to a gang inside a DAG. +See examples/README.md for the ordered tour. """ from __future__ import annotations diff --git a/hack/armada/armada-crs.yaml b/hack/armada/armada-crs.yaml new file mode 100644 index 0000000..1e4aaac --- /dev/null +++ b/hack/armada/armada-crs.yaml @@ -0,0 +1,213 @@ +apiVersion: install.armadaproject.io/v1alpha1 +kind: ArmadaServer +metadata: + name: armada-server + namespace: armada +spec: + pulsarInit: false + ingress: + ingressClass: "nginx" + replicas: 1 + image: + repository: gresearch/armada-server + tag: latest + applicationConfig: + httpNodePort: 30001 + grpcNodePort: 30002 + schedulerApiConnection: + armadaUrl: "armada-scheduler.armada.svc.cluster.local:50051" + forceNoTls: true + corsAllowedOrigins: + - "http://localhost:3000" + - "http://localhost:8089" + - "http://localhost:10000" + - "http://localhost:30000" + auth: + anonymousAuth: true + permissionGroupMapping: + submit_any_jobs: ["everyone"] + create_queue: ["everyone"] + delete_queue: ["everyone"] + cancel_any_jobs: ["everyone"] + reprioritize_any_jobs: ["everyone"] + watch_all_events: ["everyone"] + eventsApiRedis: + addrs: + - redis-ha.data.svc.cluster.local:6379 + postgres: + connection: + host: postgresql.data.svc.cluster.local + port: 5432 + user: postgres + password: psw + dbname: lookout + sslmode: disable + queryapi: + postgres: + connection: + host: postgresql.data.svc.cluster.local + port: 5432 + user: postgres + password: psw + dbname: lookout + sslmode: disable + pulsar: + URL: pulsar://pulsar-broker.data.svc.cluster.local:6650 + redis: + addrs: + - redis-ha.data.svc.cluster.local:6379 +--- +apiVersion: install.armadaproject.io/v1alpha1 +kind: Executor +metadata: + name: armada-executor + namespace: armada +spec: + image: + repository: gresearch/armada-executor + tag: latest + applicationConfig: + executorApiConnection: + armadaUrl: armada-scheduler.armada.svc.cluster.local:50051 + forceNoTls: true + metric: + port: 9001 + # Report this node label to the scheduler, so it can be used as a gang node-uniformity label. + kubernetes: + trackedNodeLabels: + - kubernetes.io/hostname + # Claim spare capacity quickly so a fresh cluster becomes schedulable in seconds, not ~a minute. + task: + allocateSpareClusterCapacityInterval: 2s +--- +apiVersion: install.armadaproject.io/v1alpha1 +kind: Lookout +metadata: + name: armada-lookout + namespace: armada +spec: + replicas: 1 + ingress: + ingressClass: "nginx" + image: + repository: gresearch/armada-lookout + tag: latest + prometheus: + enabled: false + applicationConfig: + httpNodePort: 30000 + # See https://github.com/armadaproject/armada/blob/master/config/lookoutv2/config.yaml + # for the full list of configuration options. + apiPort: 8080 + corsAllowedOrigins: + - "http://localhost" + uiConfig: + armadaApiBaseUrl: "http://localhost:30001" + postgres: + connection: + host: postgresql.data.svc.cluster.local + port: 5432 + user: postgres + password: psw + dbname: lookout + sslmode: disable +--- +apiVersion: install.armadaproject.io/v1alpha1 +kind: LookoutIngester +metadata: + name: armada-lookout-ingester + namespace: armada +spec: + image: + repository: gresearch/armada-lookout-ingester + tag: latest + applicationConfig: + postgres: + connection: + host: postgresql.data.svc.cluster.local + port: 5432 + user: postgres + password: psw + dbname: lookout + sslmode: disable + pulsar: + URL: pulsar://pulsar-broker.data.svc.cluster.local:6650 +--- +apiVersion: install.armadaproject.io/v1alpha1 +kind: Scheduler +metadata: + name: armada-scheduler + namespace: armada +spec: + replicas: 1 + ingress: + ingressClass: "nginx" + image: + repository: gresearch/armada-scheduler + tag: latest + applicationConfig: + grpc: + port: 50051 + # Schedule often so a fresh cluster becomes schedulable in seconds, not ~a minute. + schedulePeriod: 5s + # Index this node label, so it can be used as a gang node-uniformity label when placing gangs. + scheduling: + indexedNodeLabels: + - kubernetes.io/hostname + executorUpdateFrequency: 5s + auth: + anonymousAuth: true + permissionGroupMapping: + execute_jobs: ["everyone"] + armadaApi: + armadaUrl: armada-server.armada.svc.cluster.local:50051 + forceNoTls: true + pulsar: + URL: pulsar://pulsar-broker.data.svc.cluster.local:6650 + postgres: + connection: + host: postgresql.data.svc.cluster.local + port: 5432 + user: postgres + password: psw + dbname: scheduler + sslmode: disable +--- +apiVersion: install.armadaproject.io/v1alpha1 +kind: SchedulerIngester +metadata: + name: armada-scheduler-ingester + namespace: armada +spec: + replicas: 1 + image: + repository: gresearch/armada-scheduler-ingester + tag: latest + applicationConfig: + postgres: + connection: + host: postgresql.data.svc.cluster.local + port: 5432 + user: postgres + password: psw + dbname: scheduler + sslmode: disable + pulsar: + URL: pulsar://pulsar-broker.data.svc.cluster.local:6650 +--- +apiVersion: install.armadaproject.io/v1alpha1 +kind: EventIngester +metadata: + name: armada-event-ingester + namespace: armada +spec: + replicas: 1 + image: + repository: gresearch/armada-event-ingester + tag: latest + applicationConfig: + pulsar: + URL: pulsar://pulsar-broker.data.svc.cluster.local:6650 + redis: + addrs: + - redis-ha.data.svc.cluster.local:6379 diff --git a/hack/armada/postgres.values.yaml b/hack/armada/postgres.values.yaml new file mode 100644 index 0000000..e9c7d00 --- /dev/null +++ b/hack/armada/postgres.values.yaml @@ -0,0 +1,9 @@ +settings: + superuserPassword: + value: "psw" + +customScripts: + init-databases.sql: | + -- Create databases for Armada components + CREATE DATABASE scheduler; + CREATE DATABASE lookout; diff --git a/hack/armada/priority-class.yaml b/hack/armada/priority-class.yaml new file mode 100644 index 0000000..1016e23 --- /dev/null +++ b/hack/armada/priority-class.yaml @@ -0,0 +1,7 @@ +apiVersion: scheduling.k8s.io/v1 +kind: PriorityClass +metadata: + name: armada-default +value: 1000 +globalDefault: false +description: "This priority class should be as a default for Armada jobs." diff --git a/hack/armada/pulsar.values.yaml b/hack/armada/pulsar.values.yaml new file mode 100644 index 0000000..f6b4fba --- /dev/null +++ b/hack/armada/pulsar.values.yaml @@ -0,0 +1,76 @@ +volumes: + persistence: false + +affinity: + anti_affinity: false + +monitoring: + prometheus: false + +components: + zookeeper: true + bookkeeper: true + broker: true + proxy: true + autorecovery: false + functions: false + toolset: true + pulsar_manager: false + +zookeeper: + replicaCount: 1 + podMonitor: + enabled: false + +bookkeeper: + replicaCount: 1 + podMonitor: + enabled: false + configData: + # minimal memory use for bookkeeper + # https://bookkeeper.apache.org/docs/reference/config#db-ledger-storage-settings + dbStorage_writeCacheMaxSizeMb: "32" + dbStorage_readAheadCacheMaxSizeMb: "32" + dbStorage_rocksDB_writeBufferSizeMB: "8" + dbStorage_rocksDB_blockCacheSize: "8388608" + # make bookie work with large disks having little percentage disk space left + diskUsageThreshold: "0.999" + +broker: + replicaCount: 1 + podMonitor: + enabled: false + configData: + ## Enable `autoSkipNonRecoverableData` since bookkeeper is running + ## without persistence + autoSkipNonRecoverableData: "true" + # storage settings + managedLedgerDefaultEnsembleSize: "1" + managedLedgerDefaultWriteQuorum: "1" + managedLedgerDefaultAckQuorum: "1" + +proxy: + replicaCount: 1 + podMonitor: + enabled: false + +autorecovery: + podMonitor: + enabled: false + +victoria-metrics-k8s-stack: + enabled: false + victoria-metrics-operator: + enabled: false + crds: + plain: false + vmsingle: + enabled: false + vmagent: + enabled: false + vmalert: + enabled: false + alertmanager: + enabled: false + grafana: + enabled: false diff --git a/hack/armada/redis.values.yaml b/hack/armada/redis.values.yaml new file mode 100644 index 0000000..38ce5e4 --- /dev/null +++ b/hack/armada/redis.values.yaml @@ -0,0 +1,2 @@ +replicas: 2 +hardAntiAffinity: false diff --git a/hack/down.sh b/hack/down.sh new file mode 100755 index 0000000..36d2161 --- /dev/null +++ b/hack/down.sh @@ -0,0 +1,13 @@ +#!/usr/bin/env bash +# Tear the devbox down: delete the kind cluster, and drop the Flyte client's upload cache. +# +# The cache is keyed by endpoint (localhost:30080), which is the same for every devbox, so without +# clearing it here the next `up.sh` would reuse cache entries that point at this cluster's now-deleted +# minio, and the first example submit would 404 on its code bundle. +set -euo pipefail +cd "$(dirname "$0")/.." +CLUSTER="${KIND_CLUSTER:-armada}" + +kind delete cluster --name "$CLUSTER" 2>/dev/null || true +rm -rf "$HOME/.flyte/local-cache" +echo "Devbox down." diff --git a/hack/kind-config.yaml b/hack/kind-config.yaml new file mode 100644 index 0000000..9651206 --- /dev/null +++ b/hack/kind-config.yaml @@ -0,0 +1,32 @@ +# Kind cluster for the armada-flyte quickstart. +# +# Based on the armada-operator quickstart kind config (the 30000/30001/30002 Armada NodePorts), with the +# extra host ports the Flyte side needs. Everything the host reaches is a NodePort mapped straight to +# localhost here, so the demo needs no port-forwards. +kind: Cluster +apiVersion: kind.x-k8s.io/v1alpha4 +nodes: +- role: control-plane + extraPortMappings: + # --- Armada (from armada-operator) --- + - containerPort: 30000 # Lookout UI + hostPort: 30000 + protocol: TCP + - containerPort: 30001 # Armada Server REST API + hostPort: 30001 + protocol: TCP + - containerPort: 30002 # Armada Server gRPC API (armadactl) + hostPort: 30002 + protocol: TCP + # --- Flyte --- + - containerPort: 30080 # Flyte API gRPC NodePort (the examples client submits here) + hostPort: 30080 + protocol: TCP + - containerPort: 30900 # MinIO NodePort (the client's signed-URL uploads) + hostPort: 30900 + protocol: TCP + # --- Ingress (Flyte console + API served same-origin) --- + - containerPort: 30500 # Traefik web entrypoint NodePort -> browse the console at localhost:5001 + hostPort: 5001 + protocol: TCP +- role: worker diff --git a/hack/setup-armada.sh b/hack/setup-armada.sh new file mode 100755 index 0000000..a890128 --- /dev/null +++ b/hack/setup-armada.sh @@ -0,0 +1,129 @@ +#!/usr/bin/env bash +# Install Armada into the current kind cluster: the armada-operator (upstream Helm chart), its +# dependencies (Pulsar, Redis, Postgres via their upstream charts), and the Armada component CRs from +# armada/ (vendored from the operator's quickstart, with the gang node-label config baked in and +# Prometheus disabled). Mirrors the armada-operator quickstart's install steps, minus the cluster +# creation, so it runs against a cluster created from hack/kind-config.yaml. +set -euo pipefail +cd "$(dirname "$0")" + +# Derive the kube context from the cluster name up.sh passes, so KIND_CLUSTER=foo targets kind-foo. +CTX="${KUBECTL_CONTEXT:-kind-${KIND_CLUSTER:-armada}}" +# armadactl's endpoint. A dedicated override, not ARMADA_URL, which is the connector's gRPC endpoint +# (:50051) documented elsewhere and would point armadactl at the wrong address. +ARMADACTL_URL="${ARMADACTL_URL:-localhost:30002}" +kc() { kubectl --context "$CTX" "$@"; } + +# retry_until : run cmd every 5s until it exits 0 (or its stderr says +# "already exists"), or the timeout elapses. Mirrors the armada-operator e2e test's stabilization wait. +retry_until() { + local timeout=$1 desc=$2; shift 2 + local start ef; start=$(date +%s); ef=$(mktemp) + until "$@" 2>"$ef"; do + grep -qiE 'already exists' "$ef" && { rm -f "$ef"; return 0; } + if [ $(( $(date +%s) - start )) -ge "$timeout" ]; then + echo " timed out after ${timeout}s waiting for $desc" >&2; rm -f "$ef"; return 1 + fi + sleep 5 + done + rm -f "$ef" +} + +# wait_ready