Skip to content

Repository files navigation

Tokenomics — the atomic unit of the AI economy

A staff-presented booth demo. The thesis: tokens are the atomic unit of the AI economy, so every workflow decision is a spending decision — and putting Gemini 3.5 Flash on the orchestration backbone of a multi-agent system (coupling in a frontier model only where reasoning earns it) is dramatically cheaper.

All numbers are computed deterministically from a bundled pricing config and realistic per-node token counts — no live API calls, so it works on flaky booth wifi and is repeatable every time.

The three acts

  1. Anatomy of a request — one call exploded into the five token-consumption variables. Toggle RAG, history, tool definitions, a reasoning model, and prompt caching and watch the cost move.
  2. The multi-agent workflow (hero) — three strategies raced side by side: all-frontier, routed (Flash backbone + selective frontier), and all-Flash, each broken down by the five variables. Switch scenarios (deep-research vs audio pipeline) and the coupled frontier model (Claude Opus 4.8 vs GPT-5.6 Sol).
  3. Forecast at fleet scale — per-workflow economics × customers × volume, with a Jevons-paradox toggle (cheaper unit cost pulls more usage, so the aggregate can still rise).

Navigate with the Act tabs, the arrow keys, or number keys 1/2/3.

The five variables (from the FinOps token-economics breakdown)

Each shows up as a colored band in the stacked bars:

Band Variable How it's modeled
System prompt Standing instructions per call sys input tokens × input price
Context & memory Retrieved docs, history, tool defs ctx input tokens × input price
Model reasoning Larger/reasoning models emit more tokens thinking tokens × output price (per-model)
Output length Answer length out output tokens × output price
Retry & orchestration Failed calls, validation, agent-to-agent chatter inter-agent tokens + expected retries

Model selection (variable 3) is both the reasoning band and the price multiplier applied to every band — which is why switching the whole config rescales the entire stack.

Project layout

src/
  data/pricing.json     # per-1M-token rates for all 5 models — edit this the morning of
  data/scenarios.ts     # the multi-agent workflows (deep-research, audio pipeline)
  lib/cost.ts           # the cost engine: token counts + pricing -> 5 cost bands
  components/            # Act1Anatomy, Act2CostRace, Act3Forecast, StackedBar
scripts/verify.ts       # sanity-check the totals from the CLI (not shipped)
server.js               # zero-dep static server for production (PORT-aware)
Dockerfile              # multi-stage: node build -> static serve
deploy.sh               # one-command Cloud Run deploy

Run locally

npm install
npm run dev        # http://localhost:5173

Verify the numbers

npx esbuild scripts/verify.ts --bundle --platform=node --format=esm --outfile=/tmp/v.mjs && node /tmp/v.mjs

Build + serve production

npm run build      # -> dist/
npm start          # node server.js, serves dist/ on $PORT (default 8080)

Deploy to Cloud Run

gcloud auth login
gcloud config set project YOUR_PROJECT
./deploy.sh        # or: SERVICE=tokenomics-demo REGION=europe-west1 ./deploy.sh

Updating pricing before the summit

All rates live in src/data/pricing.json (USD per 1M tokens). Re-verify at the provider pages linked in that file's sources block, edit the numbers, and rebuild. Current figures as of 2026-07-21:

Model Input Output Cached in
Gemini 3.5 Flash $1.50 $9.00 $0.15
Gemini 3.1 Pro (preview) $2.00 $12.00 $0.20
Claude Sonnet 4.6 $3.00 $15.00 $0.30
Claude Opus 4.8 $5.00 $25.00 $0.50
GPT-5.6 Sol $5.00 $30.00 $0.50

Figures are a model, not a guarantee — token consumption is non-linear and forecasts are directional.

About

tokenomics-demo-version1

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages