Repository navigation
🤖 perf: update zod to 4.6.5 - #6078
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Dogfood evidence for the head build (
Video of the whole run (108 s): dogfood.mp4Generated with |
|
@codex review |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 368e36b5a2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
fe67d1b to
c8262fd
Compare
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c8262fd925
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Raise the zod floor to ^4.6.5 and dedupe the MCP packages onto the root copy. Refs #6071
zod >= 4.5 emits closed tuples as prefixItems + items:false, which json-schema-to-typescript renders as never[]. Emit draft-7 for result types. Refs #6071
Workflow journals skip a line whose timestamp lacks seconds (zod >= 4.5 rejects minute precision) and keep the run readable. Lenient userPreferences loading drops __proto__ keys and invalid record entries without adopting a prototype. Refs #6071
….5 does zod >= 4.5 `.strict()` reports an own `__proto__` key as unrecognized (colinhacks/zod#6221). The streaming readability evidence still skipped that key in strict contexts, so it called rows readable that the real parse rejects. Refs #6071
zod >= 4.5 counts string bounds in code points. The streamed evidence kept maxLength + 1 UTF-16 units, so a description of 1,025 emoji was cut to about 512 code points and looked readable. Keep 2 * (maxLength + 1) units, which always hold one code point past the bound, and test the bound against the real parse for ASCII, emoji and mixed input. Refs #6071
zod >= 4.5 rejects a zoned datetime without seconds ("2026-05-29T00:01Z"),
which zod 4.4 accepted. A run.json with one became unreadable, and
createRunIfAbsent then treated the run as half-created and deleted it.
Workflow timestamps and the evaluation attemptDeadlineAt now use zod's
documented union of both precisions. Writers still use toISOString.
Refs #6071
b876091 to
21fc89c
Compare
|
@codex review Perf owner: this head is a rebase of the approved b876091 onto main (with #6080 and #6087) and has no code changes. The merge-base-relative diff is byte-identical. The automatic review on this head completed with no findings but left no approval, so I am requesting one explicit review. Review ledger for this PR: Codex runs on 368e36b, c8262fd (automatic plus manual), b876091 and 21fc89c (automatic), so this request is run 6 of the six-round limit. If it finds anything or completes without an approval, the PR stops unmerged and I record the blocker. No further request will follow. |
|
Codex Review: Didn't find any major issues. Can't wait for the next one! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |






Refs #6071. Depends on #6080 and #6087 (both merged): this branch is rebased onto them.
Summary
This PR updates zod from 4.4.3 to 4.6.5 and raises the
package.jsonfloor to^4.6.5. zod 4.5 and 4.6 build schemas much faster. Most of the first-loadmain-*.jstask inxum serveris the evaluation of the entry module graph, and about half of that is zod schema construction.The PR has six commits:
perf: update zod to 4.6.5:package.jsonandbun.lock. The lockfile keeps one root zod copy. The nested 4.6.1 copies under the MCP SDK packages are gone.fix: keep tuple result types exact in code_execution declarations: one line insrc/node/services/ptc/typeGenerator.ts, plus a regression test. Without it, the bump turns theattach_fileresultvalueintonever[]in the code_execution type declarations.tests: pin persisted-data parsing across the zod 4.6 changes: two behavioral tests for persisted data that the new parse rules touch.fix: fail history rows with a strict-object __proto__ key, as zod 4.5 does: one line insrc/node/services/historyMessageEvidence.ts. That streaming readability check copies zod's strict-object rules by hand, and it still let an own__proto__key through. The existing differential testvalidates actual workflow variants, strict fields, and unbounded JSON valuesfailed on the bump without it.fix: keep a code-point-safe prefix in streamed history evidence: the same check keptmaxLength + 1UTF-16 units of each string, but zod 4.6 counts string bounds in code points. A description of 1,025 emoji was cut to about 512 code points and looked readable. It now keeps2 * (maxLength + 1)units, and a differential test compares it with the real parse at 1,024 and 1,025 code points for ASCII, emoji and mixed input (Codex finding, round 1).fix: keep minute-precision workflow timestamps readable under zod 4.6:IsoDateTimeSchemainorpc/schemas/workflow.tsandattemptDeadlineAtintypes/evaluation.tsnow use zod's documented union of both precisions (offset: true, with and withoutprecision: -1). Without it, arun.jsonwith a zoned timestamp without seconds became unreadable, andcreateRunIfAbsentthen treated a deterministic child run as half-created and deleted its directory (Codex finding, round 2). Writers still usetoISOString().Performance
Local Lighthouse gate: pass. Head's desktop eval median is 86.5 ms, against 160.05 ms on base: 73.6 ms lower (the gate needs at least 50 ms). Every head run stays under the bounds: desktop at most 88.1 ms observed and 88 ms reported (bound 120 ms), mobile at most 363 ms reported (bound 480 ms). No run was disqualified.
D2 Desktop Cold Start: pass. On the current head (rebased onto #6087, no code changes): run 38086092991, base
cd6aa8c7c6(merge-base), head21fc89cbee, 20 pairs. Mean change -5.17%, one-sided 95% upper bound -4.48% (gate: at most +5%), half-width 0.69%. MedianmarkMs: base 765.2, head 724.0. Earlier heads: run 38083040151 onb876091e48(mean -4.90%, upper bound -4.17%), run 38080882296 onc8262fd925(mean -4.79%, upper bound -4.24%) and run 38073922951 on368e36b5a2(mean -6.14%, upper bound -4.66%).Lighthouse auditor branch run (information, not a gate): head
21fc89cbeemerged with main0390895cc2, compared with the auditor's maincd6aa8c7c6cycle.0390895cc2adds only #6082 (bug-bash test tooling, docs and one generated skill-content file). Lighthouse 13.5.0, 3 runs per page and form factor, all 12 head runs host-qualified. The host load average was higher than in earlier runs (13-24 on 32 CPUs, both sides). All 12 head runs are under the bounds. On main, 11 of 12 runs are over (seeded mobile run 1 was 449 ms).main-*.jstask, head (ms)LCP is unchanged. Served first-load JS grows by 7,496 bytes (brotli), almost all in the
API-*.jschunk (+7,269), the same as on earlier heads. The auditor attributes this to zod 4.6.5 itself. Neither of us verified that.Lighthouse method and every run
Setup: Lighthouse 13.5.0 against
xum server --no-auth --host 127.0.0.1underenv -i, with a freshXUM_ROOTandHOMEandXUM_DISABLE_TELEMETRY=1, on the first-run page. Base and head ran as two servers at the same time, alternating run by run (8 desktop and 4 mobile each). Every build and measurement block ran under the shared host lock. "obsEval" is themain-*.jsevaluation task in the trace. "LH main" is the task length Lighthouse reports (mobile is simulated with 4x CPU throttling).I wrote the host qualification rule before any candidate run. A run is disqualified only by host data: Lighthouse failed, the auditor's Chrome overlapped, CPU pressure (
/proc/pressure/cpusome avg10) above 10, or a load average above 16 (32 CPUs). The retry budget was one extra interleaved pair per preset. It was not needed.A/A on base (two servers, same build): 0 of 24 runs disqualified. Desktop medians 155.8 and 155.15 ms (limit: within 10 ms). Mobile observed medians 154.65 and 157.45 ms (limit: within 40 ms). The host qualified.
Candidate, base vs head:
Base desktop run 6 had a host stall (observed 385.8 ms, a Layout inside the task, pressure 3.69 after the run). It is under the disqualification limits, so it stays in the record. It changes neither the base median nor any head result.
zod changelog audit (4.4.3 to 4.6.5)
I read the release notes for 4.5.0 through 4.6.5. Each row is a behavior change, Xum's use of the affected API, and the disposition.
z.iso.datetime()with a zone requires seconds (#6457)IsoDateTimeSchemainorpc/schemas/workflow.ts(about 24 fields),evaluation.tsattemptDeadlineAt,sessionTape.tsstartedAt. Every producer writestoISOString().run.json, journal lines and evaluation deadlines with minute precision stay readable, as on zod 4.4.sessionTape.tsstartedAtkeeps the new rule: only the recorder writes it, withtoISOString(), and a rejected header only skips that tape..min()/.max()/.length()count code points (#6441).min(2),.min(5)),refinement.tssha256.length(64)(hex)..max()only loosens. The.min()bounds apply to model text and are tiny.agents.get,agentSkills.getinputs). Regex-keyed records in config.unrecognized_keysno longer aborts the object (#2200).strict()objects, including tool inputs.__proto__is always stripped, and.strict()reports an own__proto__key (#6386, #6221)userPreferencesrecords.evaluation.tsalready rejects__proto__.__proto__key. Commit 4 makeshistoryMessageEvidenceagree with the real parse. New config-loading test pins that lenient preferences drop the key without adopting a prototype.items: false/additionalItems: false(#6194)json-schema-to-typescriptreads only draft-7 tuples, so result types now usetarget: "draft-7".session_history.offset_chars(.int().min(0)).minimumchanges from-MAX_SAFE_INTEGERto0, which is what runtime already enforced.type: [X, "null"]) instead ofanyOf.nullish()tool inputs per route.error(#6519)z.configor error-map use..optionsdrop reverse mappings (#6542)z.enumover a TypeScript enum..minLengthand others) become prototype getters (#6554). Methods live on the prototype (4.5).src/cli/proxifyOrpc.tsclones a schema with a spread.proxifyOrpctests pass, and the CLI help comparison over 357 commands shows only help-text changes.catchandprefaultfixes (#6192, #6440, #6587)lenient()inuserPreferences.ts,.optional().catch(undefined)inmessage.tsandtelemetry.ts.chat.jsonlrows, real config files and damaged variants of them.Provider-facing schema comparison
I dumped every tool definition through the real provider conversion paths (
getToolsForModeland AI SDKprepareToolsfor OpenAI Responses, OpenAI Chat, Anthropic and Google), the oRPC OpenAPI document, and the code_execution type declarations, on base (4.4.3) and head (4.6.5). After I normalizeanyOf: [{type: X}, {type: "null"}]totype: [X, "null"], these differences remain:anyOfbecomes a type arraysession_history.offset_charsminimum-9007199254740991 becomes 0minItems,maxItems,additionalItems: falseanyOfof two strings: the seconds form (withformat: date-time) and a minutes-only patternitems: falserequiredgainsscopez.preprocessnow reports requiredscopeagents.getandagentSkills.getallOffold into one objectxum api agents get --helpoutput now lists real flags instead of--input [json]attach_fileresultvalue: unknown[]becomes the exact tuple unionptc-types-diffParse changes that touch persisted data
2026-05-29T00:01Z): zod 4.5 rejects them. Commit 6 keeps them readable in workflow records, as zod 4.4 did. The newWorkflowRunStore.test.tscases check that a minute-precisionrun.jsonstays readable throughgetRunandgetRunStatusSnapshot, thatcreateRunIfAbsentkeeps the existing run directory and journal, and that minute-precision journal lines stay. Aworkflow.test.tscase coversattemptDeadlineAt. All of them fail onc8262fd925(before commit 6) and pass with it. Xum writes every timestamp withtoISOString(), and none of the 5,580 workflow files on this host had minute precision.__proto__key. This only affects model tool calls, not persisted data..min()and.length()count code points.The lenient
userPreferencestest passes on base and on head. It guards the__proto__and record changes: the loader keeps dropping only the bad entries, and no__proto__key becomes a prototype.Provider routes
Tool input schemas now carry nullable fields as
type: [X, "null"](22 fields per route). I checked every route Xum supports:claude-haiku-5-5accepted the head tool set and returned a validsession_historycall with nulls.gpt-6-lunaaccepted the head tool set and returned a valid call.parametersJsonSchemagemini-3.8-flash(direct API) accepted all 39 functions, including the 22 type arrays, and returned a valid call. Google's structured-output docs accept{"type": ["string","null"]}, andFunctionDeclaration.parametersJsonSchematakes JSON Schema.@ai-sdk/xaidropsadditionalProperties: falseonlyollama:route)typearray to the equivalentanyOfat the model boundarytypeonly as a string. With zod 4.6.5 on this branch, the captured/api/chatbody for the realollama:tool set has 39 tools and 0 type arrays. The customopenai-compatibleroute (for example an old Ollama/v1endpoint, which decodes the same Go structs) gets the same rewrite since #6087. With zod 4.6.5 on this branch, a customopenai-compatibleprovider with the real tool set sends 39 tools and 0 type arrays.@ai-sdk/amazon-bedrockpasses the schema through astoolSpec.inputSchema.jsonstrict. I have no Bedrock key, and I did not check how each Bedrock model family (Nova, Llama, Mistral) handles type arrays.@openrouter/ai-sdk-providerpasses the schema throughXum sets no tool-level
stricton any of these routes, so no restricted strict-mode dialect applies. Live probe cost: one request per provider per side, about 34k input and 0.6k output tokens per side in total. I estimate under $0.10. I did not check invoices.Validation
b876091e48:make static-checkandmake typecheckpass (React Compiler line stays at 23/24). The workflow (store, runner, schema, evaluation), evidence, PTC, config andproviderModelFactorysuites pass: 1,301 tests in 36 files.make teston368e36b5a2(before the code-point fix and the rebase): 24,793 pass, 3 fail. The 3 failures also fail on base4a181066f4on this host: 2taskGitPatchEnginetests (the globalinit.templateDirpoints to a missing directory, so.git/infodoes not exist) andBackupRepoCache > does not transfer 200 commits…(5 s timeout under load). The one "unhandled error between tests" comes from that timed-out test.node_modules/zod) on base and on head.typeValidator.test.tscase fails on the bump without commit 2 (Property 'type' does not exist on type 'never') and passes with it. The generated-type comparison base vs head differs only in theattach_filetuple.WorkflowRunStore.test.tsandworkflow.test.tsminute-precision cases fail onc8262fd925and pass with commit 6. With only theworkflow.tschange reverted, the store cases fail. With only theevaluation.tschange reverted, theattemptDeadlineAtcase fails.historyMessageEvidence.test.ts: 10 of 10 pass on head with commits 4 and 5. With commit 4 on base (zod 4.4.3), the suite fails, so the fix belongs with the bump. Without commit 5, the new code-point test fails.xum <cmd> --helpfor 357 CLI commands, base vs head: only help text changes. Nullable flags read "type: string or null", andagents getandagent-skills getlist real flags instead of--input [json].Dogfood
Head build,
env -iserver on 127.0.0.1 withXUM_MOCK_AI=1(a fake Anthropic key on a dead loopback port, as the bug-bash seed does, because the composer needs a configured provider):agent_skill_listcard and aworkflow_runcard to the workspace'schat.jsonl, and restarted. Both cards render. The expanded skill card groups the skills by scope.Screenshots at 1440 px and 390 px and a video of the whole run are in a comment below.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$23.89