Repository navigation
Add launch skill: scenario to a verified, monitored Spice.ai Cloud project - #29
Merged
Merged
Conversation
…oject
`launch` takes a plain-language scenario ("serve Snowflake, Postgres, and S3
data to agents with OpenAI") to a Spice.ai Cloud project that is designed,
deployed, verified, monitored, and handed off, for demos, POCs, and greenfield
production.
scripts/spice-launch.sh does the mechanical parts and checks its own results:
- create: a managed project (--kind set), or a fork of a base project so linked
org secrets carry over; reset to the stable channel.
- secrets: store each referenced secret through the Management API, reading
values from the environment or .env files, never printing them or passing
them on a command line.
- deploy: watch the new instance and stop early on an unresolved secret, a
dataset stuck in Error, or a model that fails to load, instead of waiting out
a rollout that stays in_progress while the old version serves; then confirm
the endpoint serves the new spicepod. Rejects local and private-network
sources first.
- verify: rows from every dataset and view, a model answer grounded in the
data (--ask/--expect), embeddings and search, an MCP session across several
tool calls, an OpenAI-SDK-style request, and a p50/p99 latency baseline.
- monitors and fire-drill: the profile's alert set, updated in place, and an
alert proven delivered.
- handoff: RUNBOOK.md and AGENT-CONNECT.md, regenerated only between markers.
The references cover scenario design, monitoring, and production operation,
including behavior found in live testing: Cloud MCP needs
runtime.mcp.allowed_hosts ["*"], more than one replica breaks MCP sessions,
OpenAI SDK clients need Accept-Encoding: identity for non-streaming chat
completions, Cloud's schema needs search row_ids as lists, and resizing needs
private compute.
README, AGENTS.md, and setup link the new skill.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
launch, a skill that takes a plain-language scenario, such as "I need to serve data to agents from Snowflake, Postgres, and S3, using OpenAI for models", to a Spice.ai Cloud project that is designed, deployed, tested end to end, monitored with alerts, and handed off with a runbook. It is meant for demos, POCs, and greenfield production projects.The agent turns the scenario into a spicepod, using
references/scenarios.mdandexamples/spicepod.agents.yaml.scripts/spice-launch.shthen does the mechanical parts and checks its own results:preflightlocalcreate--kind set), or forks a base project so linked org secrets carry over, and resets the channel to stablesecretsdeployin_progresswhile the old version serves. It then confirms the endpoint serves the new spicepodverify--ask/--expect), embeddings and search, an MCP session across 8 tool calls, an OpenAI-SDK-style request, and a p50/p99 latency baselinemonitors,fire-drillhandoffRUNBOOK.mdandAGENT-CONNECT.md, regenerating only the text between markersstatus,teardown--yesThe skill is self-contained: it mentions other skills but does not depend on their helpers. README, AGENTS.md (skill graph), and setup's "Where to go next" link it.
Platform behavior the skill handles
Each item was reproduced against live Spice.ai Cloud projects on runtime
2.3.2-enterprise-models. They are described by symptom and workaround, as AGENTS.md asks:/v1/mcpreturns403 Host header is not allowedon Cloud"*"runtime.mcp.allowed_hosts: ["*"](the project API key still guards the endpoint);deploylints for it404 Session not foundverifychecks 8 calls in one session;deploywarns502from non-streaming/v1/chat/completionsAccept-Encodingasks for gzip, deflate, or zstd: openai-python 3.23.0 and openai-node 7.27.0 return 502 by default and work withidentity. Streaming,/v1/sql,/v1/nsql,/v1/search, and/v1/mcpare unaffectedverifyreports a WARN;AGENT-CONNECT.mdsnippets setAccept-Encoding: identity/v1/readystill saysreadydeploydiagnoses the new instance from its logs and per-instance dataset statusnot found in [env]on a fresh project andall N resolvedon a forkcreate --baserow_idthatspice validateaccepts400 Invalid spicepod configurationwhen the deployment startsdeployadds a hintspice cloud logsanswers 404 for 1–2.5 min after a deploydeploycaptures the secrets check during the rollout and retriesTesting
Repo gates:
make check check-versions test-distribution release-previewall pass. The packaged ZIP contains the skill with the helper executable (0755) and excludesevals/.Live end-to-end runs in a test organization, every project torn down afterwards:
42858, matchingSELECT COUNT(*).resolvedon a redeploy; it previously returnednull.localskips when secrets are Cloud-only.verify --expectchecks the answer.claude mcp add --scope usersnippet reported ✔ Connected from an unrelated directory.Skill eval (skill-creator; 3 scenarios, run with and without the skill, each deploying a real project):
Most checks ran against the live projects: rows from each source, model answers matched to SQL ground truth, an 8-call MCP session, active monitors and their targets, and alert delivery. Both configurations reached working deployments. The differences were in what the user is handed and how the run behaves:
All six runs were interrupted once by an account usage limit and resumed, so the timings cover the resumed segment only. The eval workspace is not committed. The with-skill runs also found the defects fixed in this PR, including the
row_idand secrets-check issues.Follow-ups