A staff-presented booth demo. The thesis: tokens are the atomic unit of the AI economy, so every workflow decision is a spending decision — and putting Gemini 3.5 Flash on the orchestration backbone of a multi-agent system (coupling in a frontier model only where reasoning earns it) is dramatically cheaper.
All numbers are computed deterministically from a bundled pricing config and realistic per-node token counts — no live API calls, so it works on flaky booth wifi and is repeatable every time.
- Anatomy of a request — one call exploded into the five token-consumption variables. Toggle RAG, history, tool definitions, a reasoning model, and prompt caching and watch the cost move.
- The multi-agent workflow (hero) — three strategies raced side by side: all-frontier, routed (Flash backbone + selective frontier), and all-Flash, each broken down by the five variables. Switch scenarios (deep-research vs audio pipeline) and the coupled frontier model (Claude Opus 4.8 vs GPT-5.6 Sol).
- Forecast at fleet scale — per-workflow economics × customers × volume, with a Jevons-paradox toggle (cheaper unit cost pulls more usage, so the aggregate can still rise).
Navigate with the Act tabs, the arrow keys, or number keys 1/2/3.
Each shows up as a colored band in the stacked bars:
| Band | Variable | How it's modeled |
|---|---|---|
| System prompt | Standing instructions per call | sys input tokens × input price |
| Context & memory | Retrieved docs, history, tool defs | ctx input tokens × input price |
| Model reasoning | Larger/reasoning models emit more tokens | thinking tokens × output price (per-model) |
| Output length | Answer length | out output tokens × output price |
| Retry & orchestration | Failed calls, validation, agent-to-agent chatter | inter-agent tokens + expected retries |
Model selection (variable 3) is both the reasoning band and the price multiplier applied to every band — which is why switching the whole config rescales the entire stack.
src/
data/pricing.json # per-1M-token rates for all 5 models — edit this the morning of
data/scenarios.ts # the multi-agent workflows (deep-research, audio pipeline)
lib/cost.ts # the cost engine: token counts + pricing -> 5 cost bands
components/ # Act1Anatomy, Act2CostRace, Act3Forecast, StackedBar
scripts/verify.ts # sanity-check the totals from the CLI (not shipped)
server.js # zero-dep static server for production (PORT-aware)
Dockerfile # multi-stage: node build -> static serve
deploy.sh # one-command Cloud Run deploy
npm install
npm run dev # http://localhost:5173npx esbuild scripts/verify.ts --bundle --platform=node --format=esm --outfile=/tmp/v.mjs && node /tmp/v.mjsnpm run build # -> dist/
npm start # node server.js, serves dist/ on $PORT (default 8080)gcloud auth login
gcloud config set project YOUR_PROJECT
./deploy.sh # or: SERVICE=tokenomics-demo REGION=europe-west1 ./deploy.shAll rates live in src/data/pricing.json (USD per 1M
tokens). Re-verify at the provider pages linked in that file's sources block,
edit the numbers, and rebuild. Current figures as of 2026-07-21:
| Model | Input | Output | Cached in |
|---|---|---|---|
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.15 |
| Gemini 3.1 Pro (preview) | $2.00 | $12.00 | $0.20 |
| Claude Sonnet 4.6 | $3.00 | $15.00 | $0.30 |
| Claude Opus 4.8 | $5.00 | $25.00 | $0.50 |
| GPT-5.6 Sol | $5.00 | $30.00 | $0.50 |
Figures are a model, not a guarantee — token consumption is non-linear and forecasts are directional.