Context
graphistrygpt PR #3591 (feat(notebook): perceived-perf streaming + live token meter) added a live in-progress draft element — a draft=True TextElement the agent streams into the answer slot while it works, before the final answer card. On the public chat routes this element was streamed to every client, so Response.text grabbed the provisional draft instead of the final card.
What changed server-side
The public FastAPI chat routes now gate that stream behind a default-off opt-in flag, include_reasoning (graphistry/graphistrygpt PR #3591, graphistrygpt/api/routes/runners.py):
/chat/ and /chat_singleshot/ — query param or JSON body field include_reasoning (default false)
/chat_upload/ — include_reasoning form field (default false)
Default off means: by default clients now receive only the final answer card, no per-token draft chatter. So the earlier .text regression is resolved without any client change — a default louie-py client no longer needs to filter drafts.
Ask for louie-py
- Confirm the default path is clean. With
include_reasoning unset, the stream carries no draft=True TextElements, so Response.text / text_elements should already point at the card. Add a regression test asserting the draft is absent by default.
- Add an opt-in. Expose a client option (e.g.
include_reasoning=True on the chat / single-shot / upload calls) that forwards the flag, and surface the streamed draft/reasoning TextElements to callers who want live progress (kept distinct from the final answer so .text still resolves to the card).
Notes
- On the wire there is exactly one extra artifact — the
draft=True TextElement. Reasoning is not a separate emitted message; it's a token-count channel on token_flow. So this is a single flag, not two.
- The notebook renders drafts via the BFF delta path, unaffected by this flag.
Ref: graphistry/graphistrygpt#3591
Context
graphistrygpt PR #3591 (
feat(notebook): perceived-perf streaming + live token meter) added a live in-progress draft element — adraft=TrueTextElementthe agent streams into the answer slot while it works, before the final answer card. On the public chat routes this element was streamed to every client, soResponse.textgrabbed the provisional draft instead of the final card.What changed server-side
The public FastAPI chat routes now gate that stream behind a default-off opt-in flag,
include_reasoning(graphistry/graphistrygpt PR #3591,graphistrygpt/api/routes/runners.py):/chat/and/chat_singleshot/— query param or JSON body fieldinclude_reasoning(defaultfalse)/chat_upload/—include_reasoningform field (defaultfalse)Default off means: by default clients now receive only the final answer card, no per-token draft chatter. So the earlier
.textregression is resolved without any client change — a default louie-py client no longer needs to filter drafts.Ask for louie-py
include_reasoningunset, the stream carries nodraft=TrueTextElements, soResponse.text/text_elementsshould already point at the card. Add a regression test asserting the draft is absent by default.include_reasoning=Trueon the chat / single-shot / upload calls) that forwards the flag, and surface the streamed draft/reasoningTextElements to callers who want live progress (kept distinct from the final answer so.textstill resolves to the card).Notes
draft=TrueTextElement. Reasoning is not a separate emitted message; it's a token-count channel ontoken_flow. So this is a single flag, not two.Ref: graphistry/graphistrygpt#3591