Skip to content

Make setup a closed loop from empty directory to verified SQL rows - #28

Merged
lukekim merged 1 commit into
trunkfrom
lukim/setup-closed-loop
Oct 1, 2026
Merged

lukekim merged 1 commit into
trunkfrom
lukim/setup-closed-loop

Conversation

@lukekim

@lukekim lukekim commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Why

The setup Quick Start ended at an empty spice init, then spice run and show tables. spice validate reports OK for a spicepod with zero datasets, and also for one whose data file doesn't exist, so an agent could declare success with nothing to query. Live runs with the previous skill reproduced this. Haiku 4.5 reported "Spice is now set up ✅" on a pod with datasets=0, using SELECT 1 … 'Hello Spice!' as proof. In other runs it bounced off /v1/sql twice with JSON bodies, and ran pkill -f spiced. That killed another runtime holding port 8090, and the agent still reported success.

What changed

  • setup is the single entry point that takes an agent from an empty directory to rows from a real dataset. The new skills/setup/scripts/spice-local.sh covers preflight, init, start, ready, verify, and stop, and a 20-row sample CSV ships with it. The helper:
    • seeds a dataset: the sample, or --data pointing at the user's own file;
    • starts spice run in the background on free loopback ports;
    • waits for /v1/ready, then checks that each dataset is Ready and that a query returns rows;
    • returns a report with the directory, PID, ports, log path, dataset status, and sample rows;
    • never installs software, and refuses to start when spice version reports no runtime, because spice run would download one;
    • detects another runtime already holding 8090 and moves to free ports;
    • stops only the PID it started.
  • The skill defines what "working" means and explains the anti-patterns: stopping after spice init, treating spice validate as proof, pkill spiced, JSON /v1/sql bodies without parameters, and metrics on the first run (a taken metrics port makes the runtime exit). It also explains a failed user-run install whose download URL contains /download//. That happens when the installer can't read the latest release tag, usually because of GitHub's anonymous rate limit.
  • /v1/sql request bodies are now documented in setup, sql, cache, cookbook, and cloud. A raw SQL body works. A JSON body needs "parameters" ([] when unused) and exactly Content-Type: application/json. {"sql": "..."} alone returns 400 Invalid JSON: missing field 'parameters', and application/json; charset=utf-8 sends the JSON text to the SQL parser.
  • Secret placeholders use ${ store:KEY } everywhere. The secrets skill notes that spacing is optional: the runtime's lexer ignores whitespace, which I confirmed at runtime for inline and whole-value references.
  • spicepod's Quick Start is a credential-free local CSV, with PostgreSQL and OpenAI as the next example. It also notes that the spiceai/quickstart pod on Spicerack is still version: v1 and loads taxi data from public S3.
  • connectors and datasets each end with a verify section: /v1/ready, then dataset status, then SELECT … LIMIT 1.
  • AGENTS.md gets the skill graph and a list of workflow-skill regressions to check for. README leads with "Start here" and a short install table; the per-agent install details are collapsed, not removed. docs/publishing.md discloses the third helper to reviewers.
  • Setup evals add three prompts that require a seeded dataset, readiness, and rows. Five patterns that rejected correct answers over wording (for example "on its own" instead of "auto…") were widened without changing what they test.

Deliberately not included

  • No installer automation. Installation stays user-managed, as set in Prepare spiceai plugin and automate marketplace distribution #27 and enforced by validate_plugin.py, so there are no install scripts and no brew/pin/upgrade fallbacks. One early draft named Homebrew and spice install, and Opus then added installer commands to 1 of 2 answers. After rewording, 0 of 4 did.
  • No separate end-to-end skill. The loop lives in setup, where agents already land. Skills install one at a time, so a second skill would have to duplicate the helper or depend on another skill's files.
  • No runtime bug details. The installer's empty-tag behavior should be fixed upstream in install/install.sh by failing on an empty tag and using curl -f. The skill only describes the symptom and points to the docs.

Verification

  • make check test-distribution release-preview passes. The release preview packages spice-local.sh as executable, and no evals ship.

  • I ran the helper against a real v2.3.2 runtime in these cases:

    • the happy path;
    • an existing empty pod;
    • user data in a path with spaces;
    • a decoy runtime holding 8090 (it moves to 18090/15051 and leaves the decoy running);
    • a missing runtime (it refuses, and nothing is installed);
    • a dataset path that matches no files (it reports the dataset's error);
    • a runtime that exits after validate passes (a bad TLS path);
    • stop and restart on the same port.
  • Live end-to-end evals used isolated claude -p sessions with only the plugin loaded and a real v2.3.2 runtime on PATH. They compared this branch with trunk on three scenarios: the prompt "Set up Spice in ./my_app and show me a SQL result", the user's own CSV with a question to answer, and an existing project while another runtime holds port 8090.

    Model This branch trunk
    Opus 5.5 27/27 (100%), 16.6 s 22/27 (81%), 22.3 s
    Haiku 4.5 27/27 (100%), 24.5 s 14/27 (52%), 84.1 s
  • On the eight setup text evals this branch scores 35/35, and trunk scores 31/35 on the same prompts.

The setup Quick Start ended at an empty `spice init`, which `spice validate`
reports as OK with zero datasets, so agents declared success with nothing to
query. setup is now the single entry point for getting Spice working:

- Add scripts/spice-local.sh (preflight, init, start, ready, verify, stop) and
  a bundled 20-row sample CSV. It seeds a dataset, starts `spice run` in the
  background on free loopback ports, waits for /v1/ready, checks each dataset
  is Ready and a query returns rows, and reports directory, PID, ports, log,
  and sample rows. It never installs software, refuses to start when the
  runtime is missing, detects another runtime holding 8090, and stops only
  the PID it started.
- Define what "working" means and the anti-patterns: stopping after
  `spice init`, treating `spice validate` as proof, `pkill spiced`, JSON
  /v1/sql bodies without `parameters`, metrics on the first run.
- Document /v1/sql request bodies in setup, sql, cache, cookbook, and cloud.
- Use `${ store:KEY }` consistently; note that spacing is optional.
- Make spicepod's Quick Start a credential-free local CSV, and add verify
  sections to connectors and datasets.
- Lead AGENTS.md and README with the skill graph; list workflow-skill
  regressions in AGENTS.md; update the reviewer disclosure for the new helper.
- Strengthen setup evals to require a seeded dataset, readiness, and rows.
@lukekim
lukekim merged commit 997e696 into trunk Oct 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant