Skip to content

Add read-only support for LLMInferenceService - #198

Merged
juliusvonkohout merged 3 commits into
kserve:masterfrom
LogicalGuy77:llm-isvc
Aug 18, 2026
Merged

juliusvonkohout merged 3 commits into
kserve:masterfrom
LogicalGuy77:llm-isvc

Conversation

@LogicalGuy77

Copy link
Copy Markdown
Contributor

Implements #175. The web application could already display InferenceService and InferenceGraph custom resources but had no visibility into LLMInferenceService, forcing anyone deploying a large language model to fall back to kubectl for basic questions such as whether the resource was accepted, how far the controller had progressed, and what topology and parallelism the specification actually requested. This change is scoped to reading; the create, edit and delete workflows are follow ups.

Assisted-By: Claude noreply@anthropic.com

@juliusvonkohout

Copy link
Copy Markdown
Contributor

@danish9039

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds read-only LLMInferenceService visibility for #175.

Changes:

  • Adds version-tolerant backend routes and permissions.
  • Adds list/detail pages with status, topology, events, and YAML.
  • Adds utility, backend, and Cypress tests.

Reviewed changes

Copilot reviewed 21 out of 21 changed files in this pull request and generated 5 comments.

Show a summary per file
File Description
manifests/kustomize/base/cluster-role.yaml Grants read access.
frontend/src/app/types/kfserving/llm-inference-service.ts Defines resource types.
frontend/src/app/types/backend.ts Extends response types.
frontend/src/app/shared/llm-inference-service.utils.ts Parses display values.
frontend/src/app/shared/llm-inference-service.utils.spec.ts Tests parsing utilities.
frontend/src/app/services/backend.service.ts Adds client methods.
frontend/src/app/pages/llm-inference-service/llm-inference-service.module.ts Declares pages.
frontend/src/app/pages/llm-inference-service/llm-inference-service.component.ts Implements listing logic.
frontend/src/app/pages/llm-inference-service/llm-inference-service.component.html Renders the list.
frontend/src/app/pages/llm-inference-service/llm-details/llm-details.component.ts Loads resource details.
frontend/src/app/pages/llm-inference-service/llm-details/llm-details.component.scss Styles details.
frontend/src/app/pages/llm-inference-service/llm-details/llm-details.component.html Renders details.
frontend/src/app/pages/llm-inference-service/config.ts Configures table columns.
frontend/src/app/pages/index/index.component.ts Adds navigation.
frontend/src/app/app.module.ts Registers the module.
frontend/src/app/app-routing.module.ts Adds routes.
frontend/cypress/e2e/llm-inference-service.cy.ts Tests browser flows.
frontend/__mocks__/kubeflow.ts Aligns mock typing.
backend/apps/common/versions.py Defines supported versions.
backend/apps/common/routes/get.py Adds read endpoints.
backend/apps/common/routes/get_test.py Tests version detection.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread frontend/src/app/shared/llm-inference-service.utils.ts
Comment thread frontend/src/app/shared/llm-inference-service.utils.ts
Comment thread frontend/src/app/pages/llm-inference-service/llm-details/llm-details.component.ts Outdated
Comment thread frontend/src/app/types/kfserving/llm-inference-service.ts Outdated
@juliusvonkohout

Copy link
Copy Markdown
Contributor

Please do a rebase to master (not merge) for a clear commit history.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 23 out of 23 changed files in this pull request and generated 1 comment.

Suppressed comments (5)

frontend/src/app/shared/llm-inference-service.utils.ts:33

  • This treats an omitted local worker/prefill block as single-node even when baseRefs inherits a multi-node or disaggregated workload. That makes the advertised “effective topology” incorrect for configuration-driven services. Derive reconciled topology from status.workloads when available, and report an indeterminate/delegated state rather than “Single node” when referenced configurations have not been observed.
  const hasWorker = !!spec?.worker;
  const hasPrefill = !!spec?.prefill;

frontend/src/app/types/kfserving/llm-inference-service.ts:83

  • prefill is itself a workload specification with independent worker, replicas, parallelism, and scaling, but typing it as an opaque object means every new summary ignores those requested settings. A disaggregated service can therefore display only its decode configuration and even miss that its prefill workload is multi-node. Model a reusable workload interface and render decode and prefill settings separately.
  template?: K8sObject;
  worker?: K8sObject;
  prefill?: K8sObject;
  scaling?: LLMInferenceServiceScaling;

frontend/src/app/shared/llm-inference-service.utils.ts:239

  • A Kubernetes condition may contain a message without a reason. In that valid case this produces the user-facing text undefined: <message>. Prefix the message only when reason is present.
      message: `${failed.reason}: ${failed.message}`,

frontend/src/app/pages/llm-inference-service/llm-inference-service.component.ts:138

  • The parameter name a does not communicate that it contains the emitted table action. Rename it to actionEvent so the handler remains understandable without relying on surrounding context.
  public reactToAction(a: ActionEvent) {
    const llmInferenceService = a.data as LLMInferenceServiceIR;

frontend/src/app/pages/llm-inference-service/llm-details/llm-details.component.html:56

  • The details page exposes only the primary status.url; it never renders status.addresses, although the resource reports multiple reachable endpoints and #175 explicitly requires URL visibility. Secondary or origin-specific endpoints are therefore hidden. Render and deduplicate all reported addresses alongside the primary URL.
            <lib-details-list-item
              key="URL"
              *ngIf="llmInferenceService.status?.url"
            >
              <a

Signed-off-by: Harshit Nayan <harshitacademia@gmail.com>
…ing info

Signed-off-by: Harshit Nayan <harshitacademia@gmail.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 23 out of 23 changed files in this pull request and generated no new comments.

Suppressed comments (5)

frontend/src/app/pages/llm-inference-service/llm-details/llm-details.component.ts:77

  • Angular reuses this component when only route parameters change, but this branch leaves detailsLoaded and the previous object intact and does not cancel in-flight object/event requests. The page can therefore show the old service under the new route name, and a slower response for the old route can overwrite the new result. Reset the view and switch/cancel the request stream when parameters change.
    this.paramsSubscription = this.route.params.subscribe(params => {
      this.namespaceService.updateSelectedNamespace(params.namespace);

      this.serviceName = params.name;
      this.namespace = params.namespace;

      // Initial load before starting polling
      this.getBackendObjects();

frontend/src/app/shared/llm-inference-service.utils.ts:33

  • This defaults to Single node whenever worker and prefill are absent from the service's inline spec, but baseRefs can inject either field during KServe's configuration merge. A service inheriting a multi-node or disaggregated configuration is therefore mislabeled, contrary to #175's effective-topology requirement. Resolve the referenced configuration chain (or use observed workload status where available) before deriving topology; an unresolved inherited topology must not be reported as single-node.
  const hasWorker = !!spec?.worker;
  const hasPrefill = !!spec?.prefill;

frontend/src/app/types/kfserving/llm-inference-service.ts:83

  • KServe defines prefill as a full workload specification, so it can independently contain worker, replicas, scaling, and parallelism. Typing it loosely and only summarizing top-level workload fields makes those requested prefill settings invisible and can also misclassify a distributed prefill workload. Model the nested workload explicitly and render decode and prefill values separately.
  template?: K8sObject;
  worker?: K8sObject;
  prefill?: K8sObject;
  scaling?: LLMInferenceServiceScaling;

frontend/src/app/pages/llm-inference-service/llm-inference-service.component.ts:138

  • The single-letter parameter a makes the action-handling flow unnecessarily context-dependent. Rename it to actionEvent so each field access is self-explanatory.
  public reactToAction(a: ActionEvent) {
    const llmInferenceService = a.data as LLMInferenceServiceIR;

frontend/src/app/pages/llm-inference-service/llm-details/llm-details.component.html:56

  • The details page renders only the primary status.url and ignores the already-modeled status.addresses. KServe can report multiple cluster-local/private/public addresses, and #175 explicitly requires URLs to be visible, so users still need raw YAML for the additional endpoints. Render every reported address, while retaining url as the primary endpoint.
            <lib-details-list-item
              key="URL"
              *ngIf="llmInferenceService.status?.url"
            >
              <a

…dling

Signed-off-by: Harshit Nayan <harshitacademia@gmail.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 31 out of 31 changed files in this pull request and generated no new comments.

Suppressed comments (3)

frontend/src/app/shared/llm-inference-service.utils.ts:46

  • This returns the locally declared topology before considering observed workloads, so it can report the requested fragment instead of the effective merged topology required by #175. For example, if the local specification declares prefill and a baseRef contributes a decode worker, topologyFromSpec returns Disaggregated even when status.workloads.primary is a LeaderWorkerSet. Combine the local and observed evidence rather than returning the first non-empty summary, and cover mixed local/inherited topology in the tests.
  const fromSpec = topologyFromSpec(llmInferenceService?.spec);
  if (fromSpec) {
    return fromSpec;

frontend/src/app/shared/llm-inference-service.utils.ts:93

  • The observed prefill workload can also be a LeaderWorkerSet (for example when prefill.worker is inherited through a base configuration). Checking only primary.kind classifies that effective topology as merely Disaggregated. Include both observed workloads when detecting multi-node topology.
  const hasWorker = workloads.primary?.kind === 'LeaderWorkerSet';

frontend/cypress/support/sse-mock.ts:1

  • The newly exported types spell the initialism as Sse, which is inconsistent with the established SSEService/SSE naming in frontend/src/app/services/sse.service.ts:15. Rename the exported utility types to SSEWatchEventType, SSEWatchEvent, and SSEMockOptions so the public test API uses one clear form.
export type SseWatchEventType =

@juliusvonkohout
juliusvonkohout merged commit b3e2c61 into kserve:master Aug 18, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants