Skip to content

perf: xum server first load downloads and parses about 9.8 MB of JavaScript #5971

Description

@ThomasK33

xum server first load downloads and parses about 9.8 MB of JavaScript

Problem: in browser mode the first page loads 9,802 KiB of JavaScript (1,945 KiB brotli as served after #5953). Most of it is not needed for first paint. A Lighthouse audit workspace measured mobile LCP at about 12 s and Performance 54 (first-run page) and 50 (seeded page). Desktop is 87-88.

Measured on origin/main 27e888d36a (production build, chunk sizes and source-map attribution; plan lane, not rerun by me):

Chunk Raw KiB br KiB Contents
main 8,665 1,630 the app and most libraries
__vite-browser-external-* 659 174 the ghostty-web terminal (its own stub name plus inlined WASM)
API 448 130 react-dom, zod, oRPC client
DesktopPanel 30 10 VNC panel

Large modules not needed for first paint: @shikijs/langs (Streamdown's unused grammars), the terminal, Mermaid, Settings, recharts, Lottie, KaTeX. lottie-web calls direct eval, which stops esbuild from renaming top-level names in the whole main chunk (about 1,572 KiB raw).

Plan (one PR at a time, each at most about 300 changed lines)

PR Change Δ raw KiB Δ br KiB
PR0a First-load bytes script, must-stay-lazy list, startup mark 0 0
PR0b Desktop cold-start A/B runner (CI, xvfb-run) 0 0
PR1 Lazy Lottie loading animation (adds the shared lazy wrapper) −1,572 −175
PR2 Lazy terminal −640 −168
PR3 Lazy Mermaid −471 −108
PR4 Lazy Settings −350 −74
PR5 Stub Streamdown's unused Shiki grammars at build time −1,052 −77
PR6 Lazy Analytics and charts −501 −116
PR7 Lazy right-sidebar panels except Costs and Review −188 −41
PR8 CI first-load budget 0 0

Total: 9,802 → 5,027 KiB raw (−49%), 1,945 → 1,186 KiB brotli (−39%). Estimate: desktop Lighthouse likely reaches 90. Mobile LCP stays about 6-8 s, so mobile 90 needs a later phase (tool cards, KaTeX, models.json, a route-level split).

Gates per product PR: first-load bytes at or below the predicted value + 2%, no lazy chunk requested before first paint on both pages, desktop startup within +5% if a same-code A/A on the CI runner resolves it (else a deterministic request-log check), React Compiler count unchanged, Lighthouse runs as information. If a PR saves less than 90% of its predicted bytes, the remaining PRs are re-planned.

Plan copy: /home/coder/perf-runs/t3/plan.md on the perf owner's host.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high

Activity

  1. self-assigned this
    on Oct 9, 2026
  2. ThomasK33 commented on Oct 9, 2026

    @ThomasK33
    MemberAuthor

    Inputs for the mobile phase (after PR7)

    From the Lighthouse audit workspace, main 56cf11fa3c, run 20261009T212810Z, mobile preset (150 ms RTT, 1,638 kbps, 4x CPU). Not rerun by me.

    1. The LCP element is React-rendered text: on the first-run page, the onboarding dialog paragraph; on the seeded page, a code line in the chat. So LCP waits for the whole JS boot. FCP is already 3.3 s.
    2. At 1,638 kbps, the planned 1,186 KiB brotli after PR7 is about 5.9 s of transfer alone. Mobile LCP near 2.5 s needs the critical-path JS at roughly 300 KiB brotli or less, or LCP content that paints without waiting for the app bundle.
    3. Mobile TBT is 490-540 ms. A good score needs under about 200 ms.
    4. CLS is already good, so mobile 90 needs both LCP and TBT fixed. Fixing one alone tops out around 75-80.

    Decision point: after PR7 and its re-audit, I plan the mobile phase (candidates in the plan's "Phase 2" list, plus a route-level split or a prerendered shell) as its own issue.


    Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high

  3. ThomasK33 commented on Oct 10, 2026

    @ThomasK33
    MemberAuthor

    T3 desktop A/A (protocol D2): the half-width is 0.96%, which is at most 2.5%, so D2 gates the later T3 PRs.

    • Run: https://github.com/coder/xum/actions/runs/38015899054 (Desktop Cold Start, base = head = 9c6e4c3e3b, 20 pairs, 2 warm-ups per arm, 44 launches).
    • Mean per-pair delta -0.04%, one-sided 95% half-width 0.96%, upper bound 0.92%. The tolerance is +5%.
    • Median xum:app-shell-ready startTime: base 868.2 ms, head 869.1 ms.
    • The values come from the uploaded cold-start.json, and a recompute from its per-launch values matches.

    The first A/A run (https://github.com/coder/xum/actions/runs/38013815349) gave mean -0.27% and half-width 1.03%. Its artifact had stderr mixed in (Ubuntu 22.04 xvfb-run runs its command with 2>&1), which #6000 fixed. Runner: #5997.


    Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high

  4. ThomasK33 commented on Oct 10, 2026

    @ThomasK33
    MemberAuthor

    T3 Phase 1 status (perf owner). PR7 merged as #6081 (ea19a44101), so PR1 to PR7 are all on main.

    raw KiB brotli KiB
    Plan start 9,803 1,945
    After PR7, served (auditor branch run 20261010T193051Z) — 1,181.6
    Plan target after PR7 5,027 1,186

    PR7 follow-ups, both moving into PR8:

    1. The first-load check forbids only DesktopPanel. PR8's must-stay-lazy list adds the other lazy panels: Instructions, Output, Browser, DevTools, Goal, Memory, Workflows, Timeline and Artifacts.
    2. A lazy panel is blank until its chunk arrives on first open. Dogfood at 1440 px did not show it. Watch it in the PR8 dogfood.

    Next: PR8 (CI first-load budget), then the decisions on F1 (Review panel split) and on the Phase 2 order.


    Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $682.08

  5. ThomasK33 commented on Oct 10, 2026

    @ThomasK33
    MemberAuthor

    T3 status (perf owner): PR8, the CI first-load budget (#6084), is parked. Phase 1 is not complete. Codex rounds 1-4 each found one new case of what the check counts. The park reason and the restart conditions are on #6084 (issuecomment-6101945657). PR1-PR7 stay merged. Without PR8, no CI check holds the first-load bytes or keeps the lazy modules lazy, so the auditor's served-bytes record is the only watch until PR8 restarts.


    Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $689.62

  6. ThomasK33 commented on Oct 11, 2026

    @ThomasK33
    MemberAuthor

    T3 Phase 2 plan accepted (perf owner). Plan: /home/coder/perf-runs/t3/phase2-plan.md (on the perf host).

    Where the time goes today (exploratory runs: throttled HTTP, the service worker blocked, WebSockets unthrottled): useful content appears after 10.5 s on the first-run page and 11.1 s on a seeded workspace page. 74–78% of that time is the download of 1,190 KiB brotli of JavaScript, all of which must arrive before React renders anything.

    Tranche 1 (one PR at a time):

    1. PR9: measurement tooling (a boot-timeline harness with an A/A qualification, and a desktop transcript-ready metric). No product change.
    2. PR10: boot without the fade animation and without motion (−35 KiB).
    3. PR11: lazy-load the workspace, project and scratch pages, with a per-URL modulepreload that the server adds on /workspace/ URLs (first load about 506 KiB brotli).
    4. PR12: leaf imports instead of the @/common/orpc/schemas barrel.
    5. PR13: brotli quality 11 for large chunks, and a Latin subset of the font.

    Decisions:

    • D1, gates: useful content and interaction readiness gate separately, through the PR9 harness. Bytes, the desktop cold start (D2) and Lighthouse also gate. Desktop transcript readiness must pass its own A/A before PR11 depends on it. If it fails, the result is inconclusive, not information. One exception is declared now: in PR11, the workspace page's simulated Lighthouse LCP may rise. That applies only with request evidence, a paired improvement in useful content, no interaction regression, workspace page bytes within the declared bound, and no desktop transcript regression. Lighthouse runs block the service worker and start from a fresh profile, with a new baseline.
    • D2: stop after PR13 and judge the measured benefit before any larger step.
    • D3: F1 (split ReviewPanel) stays a research item. Its move-only size exception needs a measured payoff first.
    • D4: restart PR8 (🤖 ci: fail PRs that grow first-load JS over the recorded budget #6084) after PR11, but only after reproducing the bypass through this repo's Vite build. It keeps its four-round history and needs a fresh GO.

    Honest limit: tranche 1 targets earlier useful content, not a mobile Lighthouse score of 90. The plan predicts about 74–81 on the first-run page and about 56–66 on the workspace page. A score of 90 needs a larger architecture step (a startup split of the client, or a server-rendered first view). That decision goes to the owner after tranche 1, with measured costs.


    Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $713.86

  7. ThomasK33 commented on Oct 11, 2026

    @ThomasK33
    MemberAuthor

    T3 Phase 2 plan amended (perf owner, recorded before any new measurement). Scope is reduced.

    Parked: PR9a #6097 and PR9b #6098, the custom measurement harness.

    • PR9a: Codex rounds 1–3 each found the same class of defect. The head-only setup launch leaks head state into the seed: browser caches, the OS file cache, then a config.json migration.
    • PR9b: the fixed version exceeds its 480-line ceiling (531). It also found a contract problem: on the workspace page, the only LCP entry is the syntax-highlight repaint, about 0.87 s after the code line is readable.
    • Their merge permissions are revoked. Their evidence and review counts stay as they are. Restarting them is a separate measurement-design step, which needs neutral seeding, an explicitly defined paint signal, and fresh qualification.

    Blocked: PR11 (route split + per-URL preload). It needs validated useful-content, interaction and desktop transcript-readiness gates, which the parked harness was meant to provide. app-shell-ready cannot replace transcript readiness.

    Continuing, one at a time, as isolated improvements (not a completion of Phase 2): PR10 (boot without the fade and without motion), then PR12 (barrel leaf imports), then PR13 (brotli quality 11 for large chunks, and a Latin font subset). The plan's savings came from stacked prototypes, so each PR measures against its own merge-base. Its acceptance criteria are declared before its gate runs. The route-split predictions are not carried forward.

    Gates for PR10, PR12 and PR13:

    • Deterministic first-load bytes, with a forbid entry for each module the PR removes.
    • The auditor's v2 Lighthouse setup (service worker disabled, fresh profiles, service-worker proof per run) for both base and head, at 3 runs per page per preset. Mobile and desktop LCP must not be worse than base by more than 10%, and TBT by more than max(10%, 50 ms), on both pages.
    • The desktop cold-start D2 runner: one-sided 95% upper bound of at most +5%.
    • make static-check (React Compiler 23/24) and the sibling tests.
    • User-visible behavior:
      • PR10: actual content and working controls, with no new blank interval and no lost focus.
      • PR12: unchanged schema behavior and request construction.
      • PR13: non-Latin glyph coverage and font fallback preserved, and the build-time cost of compression recorded.
    • An earlier loading screen or a changed LCP candidate is not proof of earlier useful content.

    Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $725.45

  8. ThomasK33 commented on Oct 11, 2026

    @ThomasK33
    MemberAuthor

    PR10 amendment (perf owner, before any gate run).

    Finding: on current main, a scratch prototype of the motion-free boot saves only 0.9 KiB brotli (1,189.7 → 1,188.8 KiB). framer-motion stays on the first load through ChatPane → MessageRenderer → ToolMessage → getToolComponent → WorkflowRunToolCall → WorkflowTimeline → WorkflowPhaseFlow. The plan's −35 KiB was measured on top of the route split (PR11), which is blocked. That prediction was wrong for current main.

    New scope: PR10 removes the startup fade only. AnimatePresence mode="wait" holds the app's mount until the loader's 0.4 s exit animation ends. Reduced-motion users already skip it.

    Criteria for PR10:

    • Bytes: no growth (≤ base + 0.5 KiB brotli). No motion forbid entries.
    • Measured benefit is required. The desktop cold start (D2, 20 pairs, base = merge-base) must show an improvement, with the upper bound of the paired delta below 0. Lighthouse and dogfood timing are reported too. If no measurement shows a gain, PR10 stops unmerged. A pure change in how the boot looks does not merge as a perf PR.
    • Unchanged: v2 Lighthouse for base and head (LCP ≤ +10%, TBT ≤ +max(10%, 50 ms) on both pages), no new blank interval, no lost focus, dogfood with reduced motion on and off, the one-main-landmark rule, and static-check.

    Lesson for PR12 and PR13: each lane measures a scratch prototype on current main before it declares its byte criteria.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions