You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Proposal: a Result Service to read execution results without a computing unit
#8945
Follow-up to #5363. I'd like to propose moving result retrieval out of the computing unit into a new global service, and to hear objections or alternatives before opening sub-issues.
The problem
Execution results (result tables, runtime statistics, console logs) are Iceberg tables in the deployment's storage, and they outlive the computing unit that produced them. Reading them does not: the endpoints behind the result panel (result paging over the workflow websocket, /api/executions/{wid}/stats/{eid}, result export) are hosted by ComputingUnitMaster, and the gateway routes them to the CU pod. Once the CU is gone there is no route, and the panel stays empty even though the data is still there.
With per-user warehouses on (#8713), those results now sit in a warehouse the user owns, which makes not being able to see them harder to justify.
Why a service, and not Lakekeeper
Reading a result does not need the engine. Paging is a DB lookup for the result URI followed by an Iceberg read: load the table from Lakekeeper, read the Parquet files from S3. Only the deployment topology ties it to the CU.
Lakekeeper can't take this over. It is a catalog, not a query engine: its API serves table metadata and request signing, and has no endpoint that returns rows. Server-side scan planning is not implemented, and even in the Iceberg REST spec it returns a list of files, not data. Something on our side has to hold an Iceberg client and read the files.
Proposal
A new global result-service that serves the results of executions that have finished (completed, failed or killed): paging, export and runtime statistics, and later their deletion under a retention policy. It plays the same role for Lakekeeper that file-service plays for LakeFS: a Texera service in front of a catalog, translating Texera concepts (execution, operator) into catalog ones (warehouse, namespace, table) and enforcing Texera's access checks.
Today
Proposed
Execution running
Updates pushed by the engine over the websocket; paging over the websocket to the CU
Unchanged — the CU is alive while it runs
Execution finished
Paging over the websocket to the CU
HTTP to result-service; no CU needed
Reads also have to name the execution. Today the CU resolves "the latest execution of this workflow on this CU" for every page; without a CU that question has no answer. The panel instead resolves which execution to show once — by default the user's latest run of the workflow — and pages it by id.
A local prototype confirmed the shape: with every CU terminated and the CU process stopped, reopening a workflow showed the last run's results and paged through them, without opening a websocket.
Cleanup is affected too
Results are deleted from the CU today: when the last user leaves a workflow (after 30 seconds) and when a re-run starts on the same CU. So whether a result survives depends on how the user left. Once reads leave the CU, deletion should follow a retention policy in the same service rather than CU events.
Plan
Result-service with read endpoints; the frontend reads finished results over HTTP; the CU stops deleting results when the last user closes the workflow.
Retention in result-service, replacing the CU-side deletion.
User-initiated deletion moves into the same service.
Comments welcome, especially from anyone who knows of other places that assume a CU is alive when results are read.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Follow-up to #5363. I'd like to propose moving result retrieval out of the computing unit into a new global service, and to hear objections or alternatives before opening sub-issues.
The problem
Execution results (result tables, runtime statistics, console logs) are Iceberg tables in the deployment's storage, and they outlive the computing unit that produced them. Reading them does not: the endpoints behind the result panel (result paging over the workflow websocket,
/api/executions/{wid}/stats/{eid}, result export) are hosted byComputingUnitMaster, and the gateway routes them to the CU pod. Once the CU is gone there is no route, and the panel stays empty even though the data is still there.With per-user warehouses on (#8713), those results now sit in a warehouse the user owns, which makes not being able to see them harder to justify.
Why a service, and not Lakekeeper
Reading a result does not need the engine. Paging is a DB lookup for the result URI followed by an Iceberg read: load the table from Lakekeeper, read the Parquet files from S3. Only the deployment topology ties it to the CU.
Lakekeeper can't take this over. It is a catalog, not a query engine: its API serves table metadata and request signing, and has no endpoint that returns rows. Server-side scan planning is not implemented, and even in the Iceberg REST spec it returns a list of files, not data. Something on our side has to hold an Iceberg client and read the files.
Proposal
A new global result-service that serves the results of executions that have finished (completed, failed or killed): paging, export and runtime statistics, and later their deletion under a retention policy. It plays the same role for Lakekeeper that
file-serviceplays for LakeFS: a Texera service in front of a catalog, translating Texera concepts (execution, operator) into catalog ones (warehouse, namespace, table) and enforcing Texera's access checks.Reads also have to name the execution. Today the CU resolves "the latest execution of this workflow on this CU" for every page; without a CU that question has no answer. The panel instead resolves which execution to show once — by default the user's latest run of the workflow — and pages it by id.
A local prototype confirmed the shape: with every CU terminated and the CU process stopped, reopening a workflow showed the last run's results and paged through them, without opening a websocket.
Cleanup is affected too
Results are deleted from the CU today: when the last user leaves a workflow (after 30 seconds) and when a re-run starts on the same CU. So whether a result survives depends on how the user left. Once reads leave the CU, deletion should follow a retention policy in the same service rather than CU events.
Plan
Comments welcome, especially from anyone who knows of other places that assume a CU is alive when results are read.
All reactions