Skip to content

Add launch skill: scenario to a verified, monitored Spice.ai Cloud project - #29

Merged
lukekim merged 1 commit into
trunkfrom
lukekim/launch-skill
Oct 2, 2026
Merged

lukekim merged 1 commit into
trunkfrom
lukekim/launch-skill

Conversation

@lukekim

@lukekim lukekim commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds launch, a skill that takes a plain-language scenario, such as "I need to serve data to agents from Snowflake, Postgres, and S3, using OpenAI for models", to a Spice.ai Cloud project that is designed, deployed, tested end to end, monitored with alerts, and handed off with a runbook. It is meant for demos, POCs, and greenfield production projects.

The agent turns the scenario into a spicepod, using references/scenarios.md and examples/spicepod.agents.yaml. scripts/spice-launch.sh then does the mechanical parts and checks its own results:

Command What it does
preflight CLI and runtime versions, the Cloud credential and its role, plan limits (including whether CPU and memory can be resized), regions, and which secret references are available locally
local Optional local smoke run. It skips itself when some secrets exist only in Cloud
create Creates a managed project (--kind set), or forks a base project so linked org secrets carry over, and resets the channel to stable
secrets Stores each referenced secret through the Management API. Values never appear in output or on a command line
deploy Rejects local and private-network sources up front. It watches the new instance and stops early on an unresolved secret, a dataset stuck in Error, or a model that fails to load; otherwise it would wait out a rollout that stays in_progress while the old version serves. It then confirms the endpoint serves the new spicepod
verify Checks rows from every dataset and view, a model answer grounded in the data (--ask/--expect), embeddings and search, an MCP session across 8 tool calls, an OpenAI-SDK-style request, and a p50/p99 latency baseline
monitors, fire-drill Creates the demo, POC, or production alert set in place, then fires a temporary monitor to prove an alert is delivered
handoff Writes RUNBOOK.md and AGENT-CONNECT.md, regenerating only the text between markers
status, teardown One-shot health check; deletes only what the helper created, and only with --yes

The skill is self-contained: it mentions other skills but does not depend on their helpers. README, AGENTS.md (skill graph), and setup's "Where to go next" link it.

Platform behavior the skill handles

Each item was reproduced against live Spice.ai Cloud projects on runtime 2.3.2-enterprise-models. They are described by symptom and workaround, as AGENTS.md asks:

Symptom Evidence What the skill does
/v1/mcp returns 403 Host header is not allowed on Cloud 403 with the default setting and with the public hostname listed; 200 with "*" Recommends runtime.mcp.allowed_hosts: ["*"] (the project API key still guards the endpoint); deploy lints for it
More than one replica breaks MCP sessions with 404 Session not found 15 sessions × 5 calls: 31 of 75 calls lost on 2 replicas, 0 of 90 on 1 replica Uses 1 replica for MCP workloads; verify checks 8 calls in one session; deploy warns
OpenAI SDK clients get 502 from non-streaming /v1/chat/completions Fails whenever Accept-Encoding asks for gzip, deflate, or zstd: openai-python 3.23.0 and openai-node 7.27.0 return 502 by default and work with identity. Streaming, /v1/sql, /v1/nsql, /v1/search, and /v1/mcp are unaffected verify reports a WARN; AGENT-CONNECT.md snippets set Accept-Encoding: identity
A stuck rollout leaves the old version serving, and /v1/ready still says ready Observed with a missing secret, a model with a bad key, and a dataset in Error deploy diagnoses the new instance from its logs and per-instance dataset status
Org secrets resolve only when linked to the project; a fork copies the links The runtime's startup check reports not found in [env] on a fresh project and all N resolved on a fork create --base
Cloud rejects a scalar embeddings or full-text-search row_id that spice validate accepts 400 Invalid spicepod configuration when the deployment starts References use lists; deploy adds a hint
CPU and memory alerts fire about 15–20 min after a redeploy, naming the replaced instance Seen in 4 test projects; each resolved within about 5 min Documented in monitoring.md and the runbook template
spice cloud logs answers 404 for 1–2.5 min after a deploy Seen on redeploys deploy captures the secrets check during the rollout and retries

Testing

Repo gates: make check check-versions test-distribution release-preview all pass. The packaged ZIP contains the skill with the helper executable (0755) and excludes evals/.

Live end-to-end runs in a test organization, every project torn down afterwards:

  • Credential-free path (public S3 data): created, deployed in 18 s, and passed 5/5 verify checks. Latency p99 was 303 ms over 50 uncached queries, client-observed. Monitors were created and stayed idempotent on re-run, the fire drill was delivered in 171 s, and teardown left nothing behind.
  • Credentialed path (forked base, Postgres plus an OpenAI model): secrets resolved and 10/10 checks passed. Asked to count rows with SQL, the model answered 42858, matching SELECT COUNT(*).
  • Fixes from the eval, each verified live:
    • The secrets check reports resolved on a redeploy; it previously returned null.
    • A commented-out secret reference is ignored.
    • local skips when secrets are Cloud-only.
    • verify --expect checks the answer.
    • A monitors re-run kept an edited description while updating targets.
    • The slower fire drill fired in 142 s, and no regular monitor fired.
    • Regenerating the hand-off docs kept a hand-written note.
    • The generated claude mcp add --scope user snippet reported ✔ Connected from an unrelated directory.

Skill eval (skill-creator; 3 scenarios, run with and without the skill, each deploying a real project):

Scenario With skill Without skill
Agents over Snowflake (pending) + Postgres + S3, OpenAI, production alerts 14/14 checks; 1,542 s; 97 tool calls 11/14; 3,510 s; 168 tool calls
Retail demo on public sample data, docs search, Claude Code steps 10/10; 1,189 s; 89 tool calls 10/10; 3,519 s; 214 tool calls
Greenfield Postgres analytics over MCP, production, alert proof, runbook 10/11; 3,103 s; 161 tool calls 10/11; 6,674 s; 235 tool calls

Most checks ran against the live projects: rows from each source, model answers matched to SQL ground truth, an 8-call MCP session, active monitors and their targets, and alert delivery. Both configurations reached working deployments. The differences were in what the user is handed and how the run behaves:

  • The with-skill OpenAI snippets ran verbatim. One baseline's snippet returns 502 as written.
  • One baseline handed off no runbook.
  • Three runs, including one with the skill, searched the user's mailbox to confirm an alert without being asked. The skill now forbids that and asks the user instead.

All six runs were interrupted once by an account usage limit and resumed, so the timings cover the resumed segment only. The eval workspace is not committed. The with-skill runs also found the defects fixed in this PR, including the row_id and secrets-check issues.

Follow-ups

  • Re-run the full eval on this version, then optimize the skill description for triggering.
  • Report upstream: the MCP host check on Cloud, MCP sessions across replicas, the gzip 502, and the post-deploy false alerts.

…oject

`launch` takes a plain-language scenario ("serve Snowflake, Postgres, and S3
data to agents with OpenAI") to a Spice.ai Cloud project that is designed,
deployed, verified, monitored, and handed off, for demos, POCs, and greenfield
production.

scripts/spice-launch.sh does the mechanical parts and checks its own results:

- create: a managed project (--kind set), or a fork of a base project so linked
  org secrets carry over; reset to the stable channel.
- secrets: store each referenced secret through the Management API, reading
  values from the environment or .env files, never printing them or passing
  them on a command line.
- deploy: watch the new instance and stop early on an unresolved secret, a
  dataset stuck in Error, or a model that fails to load, instead of waiting out
  a rollout that stays in_progress while the old version serves; then confirm
  the endpoint serves the new spicepod. Rejects local and private-network
  sources first.
- verify: rows from every dataset and view, a model answer grounded in the
  data (--ask/--expect), embeddings and search, an MCP session across several
  tool calls, an OpenAI-SDK-style request, and a p50/p99 latency baseline.
- monitors and fire-drill: the profile's alert set, updated in place, and an
  alert proven delivered.
- handoff: RUNBOOK.md and AGENT-CONNECT.md, regenerated only between markers.

The references cover scenario design, monitoring, and production operation,
including behavior found in live testing: Cloud MCP needs
runtime.mcp.allowed_hosts ["*"], more than one replica breaks MCP sessions,
OpenAI SDK clients need Accept-Encoding: identity for non-streaming chat
completions, Cloud's schema needs search row_ids as lists, and resizing needs
private compute.

README, AGENTS.md, and setup link the new skill.
@lukekim lukekim self-assigned this Oct 2, 2026
@lukekim
lukekim merged commit b65d5a2 into trunk Oct 2, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant