Repository navigation
Add a CloudNativePG workspace: fleet triage, protection evidence, declarations and composed CNPG detail - #1921
Conversation
… composed Cluster summary - GET /api/cnpg/workspace: every CNPG kind plus owner-validated instance Pods, authorized per kind (namespaced kinds fall back per namespace; ClusterImageCatalog needs a cluster-scope list), with coverage states, CNPG issues from the Issues engine and the no-schedule audit finding withheld where coverage is missing. - k8s-ui: buildCNPGFleet derives per-cluster facts (instances, replication, protection as separate schedule/destination/last-backup/WAL/recovery-window/ restore-validation facts, declarations) without inventing values: replication lag is unknown without runtime data, restore validation is never green, and unreadable kinds read "No access" rather than none. - ResourcesSidebar gains optional category workspaces (destinations above a collapsible Resource kinds block, grouped by API group). - WorkloadView gains an optional composed summary: Overview shows it and the kind's renderer moves to a Spec & status tab. - /cnpg: Overview fleet with Needs attention / All, problem-category chips, search, namespace chip and a URL-backed drawer (?drawer=kind:group:ns:name).
…ens and composed summaries Screens (/cnpg/protection, /declarations, /pooling, /operator): - Protection keeps schedule, destination, last successful backup (with its source), WAL archiving, recovery window and restore validation as separate facts; ObjectStore upload health is labelled inferred with its evidence. - Declarations groups Databases, Publications, Subscriptions and managed roles by cluster, separating applied, not applied and pending, with controller errors and GitOps source. - Pooling lists Poolers with readiness; connection pressure reads "Not measured" until PgBouncer metrics are read. - Operator shows operator/plugin workloads and versions, image catalogs and their visible users, and operator configuration (Secret contents never read). Composed Overview summaries (Spec & status keeps the existing renderers) for Backup, ScheduledBackup, ObjectStore, Database, Publication, Subscription, Pooler and image catalogs; drawer trail back-link on workspace screens. GET /api/cnpg/operator discovers operator and plugin Deployments and their config references, gated per namespace on list deployments/services and get configmaps; it does not follow the namespace view filter. Workspace hardening from review: - issues are read flat, so Pod evidence only reaches callers with Pod access; - instance Pods must be controller-owned by the visible Cluster's UID; - partial coverage names denied namespaces only when the caller supplied the candidates, and always reports the allowed ones; - unobserved managed roles are pending, unreadable schedules or Poolers are unknown rather than absent, restore evidence requires a ready restored cluster resolved by exact reference, and an unreadable Cluster list is not reported as "no clusters".
…ctivity
- Every CNPG kind now has a CNPG-framed full detail at /cnpg/<plural>/<ns|_>/<name>,
reached from row Open, drawer expand and a redirect from /workload/...: the
workspace sidebar stays highlighted with the object nested under its home, a
return link names the page the drilldown came from (history state, kept
across tab changes), and a crumb names the object's place.
- Detail URLs pin the kube context (ctx=, added on first view). After a context
switch the page says the object is not in the new context and offers
Switch back / Go to Overview instead of loading a same-named object.
- A Cluster adds Protection (its recovery evidence), Activity (replacing
Timeline) and merged instance Logs; WorkloadView gains additive extraTabs and
the logs viewer an initialPods selection.
- GET /api/cnpg/clusters/{ns}/{name}/logs (+/stream): logs from controller-owned
instance Pods, gated on get clusters, list pods and get pods/log, with CNPG's
JSON lines parsed into level/logger/message.
- GET /api/cnpg/clusters/{ns}/{name}/activity: timeline events for the Cluster,
its instance Pods and CNPG objects attributed to it, per-kind gated. The
timeline now records the owning cluster (cnpg.io/cluster label or
spec.cluster.name) at ingestion so deleted Backups and declarations stay
attributed; earlier history is reported as incomplete.
- The drawer trail back link moves to the drawer shell so it survives hops to
Pods and Secrets.
Review fixes on the workspace screens and summaries: empty states follow each
kind's coverage, Declarations separates not applied from pending, GitOps
provenance reads "not recorded" instead of "applied directly", ObjectStore no
longer asserts recoverability, a Backup's destination inferred from the
Cluster's current config is labelled, schedule runs require the owner UID,
partial catalog coverage keeps the known users, and CNPG's six-field cron is
shown verbatim.
Docs: docs/cnpg.md (destinations, navigation, certainty table, access).
… redirects - Activity rows from Kubernetes Events now also require list events in the namespace; they carry their subject's kind, so the subject check alone let Event messages through. - The /workload redirect to a CNPG detail keeps ctx, pod and every other parameter, and maps a Cluster's timeline tab to Activity. - The drawer trail back link keeps the page's return state. - Activity describes its earliest linked child event as a recorded boundary, not a completeness guarantee.
PR Summary by QodoAdd a read-only CloudNativePG workspace and composed object details
AI Description
Diagram
High-Level Assessment
Files changed (60)
|
Code Review by Qodo
1.
|
…severity - Replication and ObjectStore-sourced facts read "No access" when Pods or ObjectStores are unreadable, instead of asserting absence. - Restore evidence only counts barman-cloud plugin recovery sources. - Inferred ObjectStore health is "Accepting uploads" only when every user cluster reports archiving; the failed-backup window uses completion time. - Log level detection reads CloudNativePG's nested PostgreSQL severity, and live stream lines keep their instance role label. - A restarted instance log stream resumes from the last delivered line instead of replaying its tail; activity Pod rows must match the live Cluster's UID; activity and stream failures are logged. - Workspace components use the shared Tooltip rather than native titles.
Conflicts: - CLAUDE.md: main condensed the endpoint list; the CloudNativePG endpoints move to docs/cnpg.md#api behind a one-line pointer. - useLogBuffer.ts: main moved level detection to utils/log-level.ts; the CloudNativePG record.error_severity rule moves with it into selectLevelField.
CloudNativePG's record.error_severity is one of PostgreSQL's severity names; another logger's numeric or custom record field no longer overrides the line's own level.
…nreadable Backups - PostgreSQL writes every DEBUGn level as DEBUG in error_severity. - Restore validation says the Backup a recovery names couldn't be read, instead of None recorded, when Backups are denied in the namespace.
Recovery replays archived WAL past the latest base backup, and no status reports the newest recoverable point, so the window shows its earliest point and reaches the newest archived WAL. It reads not advancing only while WAL archiving fails; a failed base backup no longer implies it.
Recovery replays archived WAL past the last base backup, so a failed backup makes recovery slower, not shorter; only archiving stops it.
The failing-backups banner names servers whose cluster also stopped archiving and drops the claim that recovery reaches the newest WAL, and the row keeps its archiving-stopped note while backups fail.
With failures and no success the store holds nothing to restore from, and retention isn't shown to move the earliest point.
The plugin keeps a successful backup when it fails to update the ObjectStore status, so an absent success time leaves recoverability unestablished rather than proving there is nothing to restore.
Status can miss a success whose update failed, so a failure is stated as recorded after the last recorded success, and an entry without timestamps reads as a recovery point not reported.
It is internal planning; docs/cnpg.md stops pointing at it.
…workspace (#1969) ## Summary The CloudNativePG workspace (#1921, #1922) carried a generic layer under CNPG names: the facts and problem list on every summary, per-kind coverage, the reviewed-action contract, grants, Prometheus series attribution, and screen layout borrowed from Capacity's internals. This PR moves that layer into shared, integration-neutral modules, so the next workspace-style integration (Velero is the likely one) imports it instead of copying CNPG. Nothing here is released yet, so the three CNPG wire shapes that every future workspace would copy are fixed now: - grants are structured objects, not sentences; - "not cached by Radar" is told apart from "no access"; - each object's GitOps manager comes from the server. It also adds the version-skew gate the CNPG endpoints were missing. Only mechanisms whose interface is already evident in today's code are shared: each has two or more consumers, or is a contract every workspace must follow. The rest is listed below as not shared yet. ## What changed ### Shared UI (`@skyhook-io/k8s-ui`) - **`components/facts/`**, for any surface that shows observed values (single-kind renderers included): - `Fact`, with `FactGrid`/`FactRow`/`FactValue`/`FactSource` - `CertaintyGlyph`, moved from Capacity; `CapacityCertainty` stays as the same type - `ManagedByText` - **`components/problems/`**: `WorkspaceProblem<Category>` and `ProblemCallout`/`ProblemList`/`ProblemMeta`, which take the workspace's `rootKind` instead of a hard-coded `'Cluster'`, plus `OpenIssueContext`. - **`ui/FoldSection`**: `SectionHeading`, `FoldSection`, `FoldSummary`. - **Elsewhere in k8s-ui**: - `ui/RefLink` now uses the existing `ResourceRef`. - `toneTextClass` and `worseTone` live in `ui/status-tone`. - `utils/grant` adds `Grant`, `formatGrant` and `grantParts`. - `issueReasonTitle` gives a reason one title on both the Issues page and the workspace. - **Kept CNPG-specific:** categories, ordering, wording and builders. CNPG binds `WorkspaceProblem` to its own categories. - **Badges:** CNPG badges now use `healthToSeverity` like the rest of the app. - **Buttons:** secondary buttons use a new `.btn-secondary` class (documented in DESIGN.md) instead of ten hand-rolled class strings. - **Sidebar:** `SidebarCategoryDestination.countLowerBound` renders a count over partly read data as `≥N`, and its zero as unknown rather than none. ### App (`web/`) - **`components/workspace/`** holds the screen primitives Capacity and CNPG share: `ScreenBody`, `ScreenEmptyState`, `Notice`, `Segments`, `FilterChips`, `SectionTable` and its table classes, `RefreshFailedNotice`, `GrantText`, and the text helpers. CNPG no longer imports from `capacity/shared.tsx`. - **`api/actions.ts`** is the client half of the action contract: - the `ActionCapability` and `ActionRequest` types - `actionErrorCode`, which an integration widens with its own codes - `actionOutcomeLocked` and `actionCompleted` - `capabilityReason`, now the one "why is this disabled" wording - **Utilities:** `utils/drawer-trail.ts` holds the `?drawer=` codec, and `utils/page-links.ts` the back label and the subject-filtered Issues link. - **Kind table:** CNPG's detail kinds derive kind and group from the workspace kind table. ### Server - **`internal/server/actions.go`**: the reviewed-action contract. - `ActionCapability`, `ActionRequest`, `decodeActionRequest` and `decodeActionParams` - refusals: `changedAction`, `blockedAction`, `partialAction` and `writeActionError` - the permission decision: `grantPermission` and `capabilityVerdict`, with a real SelfSubjectAccessReview in local mode - `mergePatchAtVersion` - **Grants:** `Grant` lives in `internal/auth` and carries `In(ns)` and `String()`. - **Kind access:** `kind_access.go` covers per-kind access and coverage (`readWorkspaceKind`, `typedKindScope`, `KindCoverage`); `cache_scope.go` has `namespacesWithinCache`. - **Other shared helpers:** - `read_source.go`: one `ReadSource` shape for HA, recovery and storage - `fanout.go`: the bounded per-namespace fan-out - `internal/prometheus/series_scope.go`: `SeriesIsolation`, already used by the PVC usage charts - **Kept CNPG-specific:** facts, guards, runners and the extra refusal codes. - **Exported manager detection:** `pkg/topology` now exports `ManagedByFromMeta`. This is additive. ### Wire changes (CNPG endpoints, unreleased) - **`grant`** is `{verb, group?, resource, subresource?, namespace?}` (no namespace means cluster-wide) on every capability, read source, chart and report item. The wording is unchanged and comes from `Grant.String()` / `formatGrant`. - **Kind coverage** gains: - a state `uncached`, for a scope Radar's cache does not cover at all; - `uncachedNamespaces`, named under the same rule as `deniedNamespaces`. Before, a namespace the caller can read but Radar does not cache was reported as denied, and the UI said "No access". It now says "Radar does not cache Pods in db". - **`/api/cnpg/workspace`** gains `managedBy`, keyed `Kind/ns/name`. - "Declared in" now uses the server's manager detection and links to the Argo CD application or Flux object. - Before, the client parsed labels itself. That misnamed Argo CD applications living outside Argo CD's namespace (`ns_app`) and ignored the Flux namespace label. ### Version skew New `FeatureCapabilities` flags: `cnpgWorkspace` (every `/api/cnpg/*` route except the two image-catalog lookups that predate it) and `gitopsWriteEvidence`. Each has a `radarFeatures` entry. Every CNPG query, mutation, log stream, report download and the write-evidence query goes through `useRadarFeature`. Mutations are never retried. On a Radar without the workspace: - the sidebar offers no workspace destinations; - `/workload` links and drawer Expand stay on the standard views; - a Cluster's Logs tab shows its Pods' logs; - a workspace screen opened directly says it needs a newer Radar. ### Vocabulary On screen these areas are named by their subject. The block above a sidebar category's kinds reads **Views**, and the copy says "CloudNativePG views" and "CloudNativePG data", never "workspace", because Radar Hub uses "workspace" for the customer's account. "Workspace" remains the internal term for an integration with several kinds and its own screens (the integration guide, `SidebarCategoryWorkspace`, `/api/cnpg/workspace`). ### Docs - **DESIGN.md:** the rules for unknown, partial and denied values are written once. `capacity.md` and `cnpg.md` link to them and keep their per-value tables. - **`docs/INTEGRATION_GUIDE.md`:** a new "Workspace integrations" chapter covering when a workspace is warranted, the pieces to import, and what is not shared yet. CLAUDE.md points to it. ## Not shared yet These have one consumer, and their interface should come from the second: - the workspace registry / App wiring - the operation tracker - the generic action runner - the fixed-path `pods/proxy` reader - the merged log stream - the report bundle - operator diagnosis - a TTL-memo helper Rollouts' capabilities answer "allowed" in local mode without asking the apiserver. That is a real bug, but Rollouts returns bare booleans and hides denied actions, so it needs its own PR. ## Public surface - **k8s-ui:** exports that exist on main are unchanged. Renamed exports were all added in #1921/#1922. `CapacityCertainty` is kept. - **`pkg/`:** one additive export (`topology.ManagedByFromMeta`). - **`pkg/capacityapi` v1alpha1:** unchanged. - **radar-app:** surface unchanged. ## Testing - **Go:** `go build ./...`; `go test ./...` in both modules. The `cmd/desktop` env test is flaky and passes alone. `gofmt -l` is clean. - **Frontend:** - type checks pass for web and k8s-ui; - vitest: k8s-ui 236 files / 4,369 tests, web 182 files / 2,028 tests; - eslint: 0 errors, and no new warnings. - **Live on the kind CloudNativePG demo, full access and a view-only identity:** - every CNPG page rendered with structured grants and no "[object Object]" or "undefined"; - a namespaced Argo CD tracking ID on a demo Cluster showed "Declared in: Argo CD application argocd/payments", linked; - with the capability flags removed from `/api/capabilities` (an older Radar behind Radar Hub), the sidebar workspace, redirects, composed summaries and actions fell back to the standard views, and `/cnpg` showed the upgrade note. Stacked on #1922 (which is stacked on #1921). Merge order: #1921, #1922, then this PR, before a release ships `/api/cnpg/*`. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches CNPG API shapes, RBAC/coverage semantics, and cluster write paths via the shared action layer; behavior is intended to be equivalent with clearer grant and cache messaging. > > **Overview** > Extracts the generic workspace layer that CloudNativePG had under CNPG-specific names into shared modules, so future integrations (e.g. Capacity-style screens) import one contract instead of copying. > > **Server:** Adds `internal/auth.Grant` (structured RBAC on the wire), `internal/server/actions.go` (reviewed `ActionRequest`, capabilities, 409/partial errors, `mergePatchAtVersion`), `cache_scope.go` / `kind_access` patterns, and `prometheus/series_scope.go` (shared Prometheus isolation). CNPG handlers now use these; history/PVC denial fields expose `*Grant` instead of strings. `FeatureCapabilities` adds `cnpgWorkspace` and `gitopsWriteEvidence`. CNPG workspace coverage gains `uncached` / `uncachedNamespaces` (distinct from denied) and `managedBy` from server-side GitOps detection. > > **UI:** New k8s-ui `facts/`, `problems/`, fold sections, `formatGrant`, and app `components/workspace/` plus `api/actions.ts`. CNPG is rewired to structured grants and shared action types; secondary buttons use `.btn-secondary`. > > **Docs:** DESIGN.md documents unknown/partial/denied values once; INTEGRATION_GUIDE adds a workspace-integrations chapter; CLAUDE.md points builders at it. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 5af6ed1. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->
# Conflicts: # packages/k8s-ui/src/components/logs/WorkloadLogsViewer.tsx
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 2 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 135d8fd. Configure here.
…health inference A CNPG Cluster's instance Pods are its children, not related Pods, so the Logs tab no longer depends on them. Protection's Destinations column now uses the same inference as the ObjectStore summary, so recovery windows left by clusters that no longer use the store don't read as failures.

Summary
CloudNativePG support in Radar today is ten separate CRD lists and renderers. Answering "which Postgres cluster needs attention, why, and what do I look at next" means stitching Clusters, Backups, ObjectStores, declarations and Pods together by hand. This adds a CloudNativePG workspace inside Resources that does that stitching. It covers fleet triage, recovery evidence, declared vs reconciled state, pooling and the operator. Every CNPG object gets one composed detail, whichever way you reach it.
The workspace is read-only. Every value it shows is something the cluster actually reports, labelled with its source. When something isn't reported it says so, never zero, "none" or green. The full contract is in
docs/cnpg.md.What changed
Workspace screens (
/cnpg/*, inside Resources)The Resources sidebar's CloudNativePG group gains a Workspace block above the exact kinds. Kinds sit under a collapsible "Resource kinds" block, grouped by API group.
Overview is the cluster fleet, defaulting to Needs attention. It shows instance pills, replication, protection, declarations, PG version and the top problem, with problem-category chips, search and a namespace chip. Each row inspects in the drawer and has explicit Logs and Open → actions.
Protection keeps these as separate facts per cluster, each with its source:
It also lists failed backups from the last 7 days, destinations (ObjectStore upload health, labelled inferred with its evidence) and schedules. Restore validation is never green: Kubernetes records no restore tests.
Declarations lists Databases, Publications, Subscriptions and managed roles by cluster. It separates applied, not applied and pending, and shows controller errors and the GitOps source ("not recorded" when there isn't one).
Pooling lists Poolers and the clusters they front. Connection pressure reads "Not measured".
Operator shows operator and plugin Deployments with versions and readiness, image catalogs and their visible users, and operator configuration. Secret contents are never read.
One composed detail per object
WorkloadViewgains an optional composed summary. For CNPG kinds, Overview is the summary (facts grouped by task, e.g. Backup → Outcome / Relationships) and the existing renderer moves to a Spec & status tab, so nothing repeats./cnpg/<plural>/<ns|_>/<name>. You reach it from Open, from drawer expand, or by redirect from/workload/.... It keeps the workspace sidebar with the object nested under its home destination.record.error_severityis read by the sharedutils/log-level.ts, so every log viewer levels CNPG lines the same way).?pod=preselects an instance.docs/cnpg.md#api;CLAUDE.mdpoints there.Navigation
?drawer=kind:group:ns:name. Links inside it build a trail with a "← previous object" link in the drawer shell.ctx=<kube context>, added on first view. After a context switch the page says " is not in ", with Switch back and Go to Overview, instead of opening a same-named object from another cluster.Backend
GET /api/cnpg/workspacereturns every CNPG kind plus instance Pods, authorized per kind, with coverage states:full,partial(including the allowed namespaces),denied,syncing,errorornotInstalled.ClusterImageCatalogneeds a cluster-scopelist.cnpgNoDeclarativeBackupaudit finding. Both are withheld where coverage is missing.GET /api/cnpg/operatordiscovers operator and plugin workloads and their config references.get configmaps; Secrets are never read.GET /api/cnpg/clusters/{ns}/{name}/logs(plus/logs/stream) returns merged logs from owner-validated instance Pods. It is gated onget clusters,list podsandget pods/log.GET /api/cnpg/clusters/{ns}/{name}/activityreturns timeline events. Each kind is gated onlist <kind>; Event rows additionally needlist events.ExtractLabelsnow retainscnpg.io/cluster(orspec.cluster.name) on CNPG objects and Pods, so deleted Backups and declarations stay attributed. No storage schema change.Shared package (
@skyhook-io/k8s-ui), all additiveResourcesSidebar: optionalcategoryWorkspaces, with a pass-throughsidebarCategoryWorkspacesonResourcesView.WorkloadView: optionalrenderSummaryandextraTabs, plus a'spec'value inWorkloadTabType.WorkloadLogsViewer: optionalinitialPods.components/cnpg:buildCNPGFleet, the relation helpers and the summaries.Consumers that pass none of these props see no change. That includes Radar Hub, which only needs to opt in if it wants the workspace.
Worth reviewing closely
internal/server/cnpg_workspace.go: coverage, issue withholding, and not naming namespaces.docs/cnpg.md, andpackages/k8s-ui/src/components/cnpg/workspace.ts+relations.ts, which implement it.?drawer=two-way sync inweb/src/components/cnpg/CNPGView.tsx.App.tsx, which preservesctxon/cnpg/routes.Not included (needs a product/RBAC decision first)
pods/proxy) and/or Prometheus. Replication reads "lag unknown" until then.These ship in #1922, stacked on this PR.
Testing
main; CI is green on the head.make tscpasses.go test ./internal/server/ ./internal/timeline/ ./internal/k8s/andpkg/timelinepass, including new handler tests for:make test: one unrelatedcmd/desktopenv test (TestEnrichEnvPrecedenceAndDiagnostics) failed under the full parallel run and passes on its own; nothing incmd/changed.make cnpg-demoand walked these paths:Note
Medium Risk
New RBAC-gated read APIs (workspace per-kind coverage, pod logs, timeline activity) and log streaming increase authorization surface area, though behavior is read-only and heavily tested for partial/denied responses.
Overview
Adds a read-only CloudNativePG workspace under Resources (
/cnpg/*) that stitches fleet triage, protection evidence, declarations, pooling, and operator views from one per-kind API payload, with composed per-kind detail pages (Overview summary + Spec & status) and Cluster-specific Protection, Activity, and merged instance Logs.Backend: New
GET /api/cnpg/workspace(per-kind coverage, issues, audit, owner-validated instance Pods),GET /api/cnpg/operator(ignores namespace view filter; ConfigMaps only withget), cluster logs (snapshot + SSE with CNPG JSON parsing) and activity (timeline with deleted-child attribution). Routes are registered inserver.go; timeline ingestion retainscnpg.io/clusterfor attribution.Frontend (
@skyhook-io/k8s-ui): Newcomponents/cnpg(fleet derivation, relations, certainty-aware summaries), optionalWorkloadViewsummary/extra tabs, sidebar workspace block, drawer/ctxnavigation patterns documented indocs/cnpg.md(also linked fromCLAUDE.mdandintegrations.md).Tests: Extensive server tests for coverage, RBAC withholding, logs/activity/stream behavior; vitest for CNPG log level detection.
Reviewed by Cursor Bugbot for commit e860514. Bugbot is set up for automated code reviews on this repo. Configure here.