Skip to content

Scope product-lifecycle and update-advisor eval prompts - #1

Merged
jhadvig merged 1 commit into
jhadvig:jhadvig/ota-2024-cluster-update-evalsfrom
jrangelramos:scope-eval-system-prompts
Jul 9, 2026
Merged

Scope product-lifecycle and update-advisor eval prompts#1
jhadvig merged 1 commit into
jhadvig:jhadvig/ota-2024-cluster-update-evalsfrom
jrangelramos:scope-eval-system-prompts

Conversation

@jrangelramos

Copy link
Copy Markdown
Collaborator

Summary

  • product-lifecycle: Replace copy-pasted broad "upgrade advisor" persona with a narrowly scoped, single-skill prompt (matches the pattern used by find-token, kubernetes-docs, openshift-docs evals)
  • update-advisor: Rewrite with its own identity and an explicit decision matrix, removing the duplication
  • test_cases: Decouple product-lifecycle queries from CVO proposal framing so they exercise the skill directly
  • Remove dangling OWNERS symlink from product-lifecycle eval

Why

Both eval system prompts were character-for-character identical. This made it impossible to isolate what each eval was testing — a product-lifecycle failure could have been caused by the broad persona routing through another skill. Each eval should test one thing.

Test plan

  • Run product-lifecycle eval and verify test cases pass against the live Product Life Cycle API
  • Run update-advisor eval and verify test cases still pass with the new prompt
  • Confirm no other evals are affected

🤖 Generated with Claude Code

Both eval system prompts were identical copies of a broad "OpenShift
upgrade advisor" persona. This made it impossible to isolate what each
eval was actually testing.

- product-lifecycle: narrow to single-skill prompt focused on the
  Product Life Cycle API (mirrors find-token/kubernetes-docs pattern)
- update-advisor: rewrite with its own identity and decision matrix
- test_cases: decouple queries from CVO proposal framing so they
  exercise the product-lifecycle skill directly
- Remove dangling OWNERS symlink from product-lifecycle eval

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@jrangelramos

jrangelramos commented Jul 9, 2026

Copy link
Copy Markdown
Collaborator Author

Product lifecycle manual test exercise - All tests pass

🔍 Click to expand - Product Lifecycle Test Log
Starting provider containers...
  claude: port 18080 (container 8645bf88ee03)
Waiting for servers...
  claude: ready

Running evals...
/usr/lib/python3.14/site-packages/pytest_asyncio/plugin.py:211: PytestDeprecationWarning: The configuration option "asyncio_default_fixture_loop_scope" is unset.
The event loop scope for asynchronous fixtures will default to the fixture caching scope. Future versions of pytest-asyncio will default the loop scope for asynchronous fixtures to function scope. Set the default fixture loop scope explicitly in order to avoid unexpected behavior in the future. Valid fixture loop scopes are: "function", "class", "module", "package", "session"

  warnings.warn(PytestDeprecationWarning(_DEFAULT_FIXTURE_LOOP_SCOPE_UNSET))
=============================================================================================== test session starts ================================================================================================
platform linux -- Python 3.14.3, pytest-8.3.5, pluggy-1.6.0 -- /usr/bin/python3
cachedir: .pytest_cache
rootdir: /home/jeramos/pixaa/test/test5-rename/agentic-skills/evals
configfile: pytest.ini
plugins: anyio-4.12.1, asyncio-1.1.0, xdist-3.7.0, timeout-2.4.0
asyncio: mode=Mode.AUTO, asyncio_default_fixture_loop_scope=None, asyncio_default_test_loop_scope=function
timeout: 600.0s
timeout method: signal
timeout func_only: False
collected 228 items / 186 deselected / 42 selected                                                                                                                                                                 

evals/skills/test_eval.py::test_skill[claude-product-lifecycle-plc_proposal_olm_batch_check] PASSED                                                                                                          [  2%]
evals/skills/test_eval.py::test_skill[claude-product-lifecycle-plc_cluster_logging_supported] PASSED                                                                                                         [  4%]
evals/skills/test_eval.py::test_skill[claude-product-lifecycle-plc_web_terminal_compat_check] PASSED                                                                                                         [  7%]
evals/skills/test_eval.py::test_skill[claude-product-lifecycle-plc_compliance_operator_status] PASSED                                                                                                        [  9%]
evals/skills/test_eval.py::test_skill[claude-product-lifecycle-plc_ocp_platform_status] PASSED                                                                                                               [ 11%]
evals/skills/test_eval.py::test_skill[claude-product-lifecycle-plc_ocp_old_version_extended] PASSED                                                                                                          [ 14%]
evals/skills/test_eval.py::test_skill[claude-product-lifecycle-plc_batch_known_operators_only] PASSED                                                                                                        [ 16%]
evals/skills/test_eval.py::test_skill[gemini-product-lifecycle-plc_proposal_olm_batch_check] SKIPPED (No server for gemini)                                                                                  [ 19%]
evals/skills/test_eval.py::test_skill[gemini-product-lifecycle-plc_cluster_logging_supported] SKIPPED (No server for gemini)                                                                                 [ 21%]
evals/skills/test_eval.py::test_skill[gemini-product-lifecycle-plc_web_terminal_compat_check] SKIPPED (No server for gemini)                                                                                 [ 23%]
evals/skills/test_eval.py::test_skill[gemini-product-lifecycle-plc_compliance_operator_status] SKIPPED (No server for gemini)                                                                                [ 26%]
evals/skills/test_eval.py::test_skill[gemini-product-lifecycle-plc_ocp_platform_status] SKIPPED (No server for gemini)                                                                                       [ 28%]
evals/skills/test_eval.py::test_skill[gemini-product-lifecycle-plc_ocp_old_version_extended] SKIPPED (No server for gemini)                                                                                  [ 30%]
evals/skills/test_eval.py::test_skill[gemini-product-lifecycle-plc_batch_known_operators_only] SKIPPED (No server for gemini)                                                                                [ 33%]
evals/skills/test_eval.py::test_skill[openai-product-lifecycle-plc_proposal_olm_batch_check] SKIPPED (No server for openai)                                                                                  [ 35%]
evals/skills/test_eval.py::test_skill[openai-product-lifecycle-plc_cluster_logging_supported] SKIPPED (No server for openai)                                                                                 [ 38%]
evals/skills/test_eval.py::test_skill[openai-product-lifecycle-plc_web_terminal_compat_check] SKIPPED (No server for openai)                                                                                 [ 40%]
evals/skills/test_eval.py::test_skill[openai-product-lifecycle-plc_compliance_operator_status] SKIPPED (No server for openai)                                                                                [ 42%]
evals/skills/test_eval.py::test_skill[openai-product-lifecycle-plc_ocp_platform_status] SKIPPED (No server for openai)                                                                                       [ 45%]
evals/skills/test_eval.py::test_skill[openai-product-lifecycle-plc_ocp_old_version_extended] SKIPPED (No server for openai)                                                                                  [ 47%]
evals/skills/test_eval.py::test_skill[openai-product-lifecycle-plc_batch_known_operators_only] SKIPPED (No server for openai)                                                                                [ 50%]
evals/skills/test_eval.py::test_skill[deepagents-claude-product-lifecycle-plc_proposal_olm_batch_check] SKIPPED (No server for deepagents-claude)                                                            [ 52%]
evals/skills/test_eval.py::test_skill[deepagents-claude-product-lifecycle-plc_cluster_logging_supported] SKIPPED (No server for deepagents-claude)                                                           [ 54%]
evals/skills/test_eval.py::test_skill[deepagents-claude-product-lifecycle-plc_web_terminal_compat_check] SKIPPED (No server for deepagents-claude)                                                           [ 57%]
evals/skills/test_eval.py::test_skill[deepagents-claude-product-lifecycle-plc_compliance_operator_status] SKIPPED (No server for deepagents-claude)                                                          [ 59%]
evals/skills/test_eval.py::test_skill[deepagents-claude-product-lifecycle-plc_ocp_platform_status] SKIPPED (No server for deepagents-claude)                                                                 [ 61%]
evals/skills/test_eval.py::test_skill[deepagents-claude-product-lifecycle-plc_ocp_old_version_extended] SKIPPED (No server for deepagents-claude)                                                            [ 64%]
evals/skills/test_eval.py::test_skill[deepagents-claude-product-lifecycle-plc_batch_known_operators_only] SKIPPED (No server for deepagents-claude)                                                          [ 66%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-product-lifecycle-plc_proposal_olm_batch_check] SKIPPED (No server for deepagents-gemini)                                                            [ 69%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-product-lifecycle-plc_cluster_logging_supported] SKIPPED (No server for deepagents-gemini)                                                           [ 71%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-product-lifecycle-plc_web_terminal_compat_check] SKIPPED (No server for deepagents-gemini)                                                           [ 73%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-product-lifecycle-plc_compliance_operator_status] SKIPPED (No server for deepagents-gemini)                                                          [ 76%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-product-lifecycle-plc_ocp_platform_status] SKIPPED (No server for deepagents-gemini)                                                                 [ 78%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-product-lifecycle-plc_ocp_old_version_extended] SKIPPED (No server for deepagents-gemini)                                                            [ 80%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-product-lifecycle-plc_batch_known_operators_only] SKIPPED (No server for deepagents-gemini)                                                          [ 83%]
evals/skills/test_eval.py::test_skill[deepagents-openai-product-lifecycle-plc_proposal_olm_batch_check] SKIPPED (No server for deepagents-openai)                                                            [ 85%]
evals/skills/test_eval.py::test_skill[deepagents-openai-product-lifecycle-plc_cluster_logging_supported] SKIPPED (No server for deepagents-openai)                                                           [ 88%]
evals/skills/test_eval.py::test_skill[deepagents-openai-product-lifecycle-plc_web_terminal_compat_check] SKIPPED (No server for deepagents-openai)                                                           [ 90%]
evals/skills/test_eval.py::test_skill[deepagents-openai-product-lifecycle-plc_compliance_operator_status] SKIPPED (No server for deepagents-openai)                                                          [ 92%]
evals/skills/test_eval.py::test_skill[deepagents-openai-product-lifecycle-plc_ocp_platform_status] SKIPPED (No server for deepagents-openai)                                                                 [ 95%]
evals/skills/test_eval.py::test_skill[deepagents-openai-product-lifecycle-plc_ocp_old_version_extended] SKIPPED (No server for deepagents-openai)                                                            [ 97%]
evals/skills/test_eval.py::test_skill[deepagents-openai-product-lifecycle-plc_batch_known_operators_only] SKIPPED (No server for deepagents-openai)                                                          [100%]

============================================================================ 7 passed, 35 skipped, 186 deselected in 677.59s (0:11:17) =============================================================================
eval-claude

🔍 Click to expand - Product Lifecycle Sandbox container logs
$ podman logs -f 8645bf88ee03
INFO lightspeed_agentic.config: Resolved LIGHTSPEED_PROVIDER=anthropic → SDK=claude
INFO lightspeed_agentic.app: Starting app (sdk=claude, model=claude-opus-4-6, LIGHTSPEED_MODEL=, audit=False)
INFO:     Started server process [2]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://0.0.0.0:8080 (Press CTRL+C to quit)
INFO:     192.168.0.30:43716 - "GET /health HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=f461c89863883865c6899f9baaa52ab6)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to use the product-lifecycle skill to look up lifecycle data for 5 operators against OCP 4.21. Let me invoke that skill.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "product-lifecycle", "args": "Look up lifecycle data for each of these 5 operators against OCP 4.21:\n\n1. compliance-operator v1.9.0 (channel: stable, package: compliance-operator)\n2. cluster-logging v6.5.1 (channel: stable-6.5, package: cluster-logging)\n3. devworkspace-operator v0.41.0)
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to look up lifecycle data for 5 operators against OCP 4.21 using the Product Life Cycle API. Let me follow the skill's guidance:

1. First, do a batch lookup with `?name=OpenShift` to cover most operators
2. Then do individual lookups for any not found

Let me start by doing the batch query and individual queries for each operator. I'll use the approach recommended in the skill - fetch `?name=OpenShift` once, then make individual calls for operators not found.

Let me run the queries in parallel.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=OpenShift\" | jq -r '[.data[] | select(.is_operator) | {name, package, is_operator}]'", "description": "Batch fetch all OpenShift operator products from PLC API", "timeout": 30000})
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=compliance\" | jq -r '[.data[] | {name, package, is_operator}]'", "description": "Search PLC API for compliance operator", "timeout": 30000})
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=devworkspace\" | jq -r '[.data[] | {name, package, is_operator}]'", "description": "Search PLC API for devworkspace operator", "timeout": 30000})
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=web+terminal\" | jq -r '[.data[] | {name, package, is_operator}]'", "description": "Search PLC API for web-terminal operator", "timeout": 30000})
INFO lightspeed_agentic: [provider:run] thinking: Now I know which operators have PLC data:
1. compliance-operator - FOUND (package: compliance-operator)
2. cluster-logging - FOUND (package: cluster-logging, name: "logging for Red Hat OpenShift")
3. devworkspace-operator - NOT FOUND
4. openshift-pipelines-operator-rh - Need to check, not in the OpenShift batch. Let me search specifically.
5. web-terminal - FOUND (package: web-terminal)

I still need to search for openshift-pipelines-operator-rh. Let me also get the detailed version/compatibility data for the ones that were found.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=pipelines\" | jq -r '[.data[] | {name, package, is_operator}]'", "description": "Search PLC API for pipelines operator", "timeout": 30000})
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=compliance\" | jq -r --arg pkg \"compliance-operator\" --arg target \"4.21\" '.data[] | select(.package == $pkg) | \"\\(.name) (package: \\(.package))\", (.versions[] | .name as $ver | .type as $type | (.opensh)
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=logging+for+Red+Hat+OpenShift\" | jq -r --arg pkg \"cluster-logging\" --arg target \"4.21\" '.data[] | select(.package == $pkg) | \"\\(.name) (package: \\(.package))\", (.versions[] | .name as $ver | .type as $)
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=web+terminal\" | jq -r --arg pkg \"web-terminal\" --arg target \"4.21\" '.data[] | select(.package == $pkg) | \"\\(.name) (package: \\(.package))\", (.versions[] | .name as $ver | .type as $type | (.openshift_c)
INFO lightspeed_agentic: [provider:run] thinking: Now I have all the data. Let me also check if openshift-pipelines has a different name. The search for "pipelines" returned nothing. Let me try a broader search.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=Tekton\" | jq -r '[.data[] | {name, package, is_operator}]'", "description": "Search PLC API for pipelines under alternate name Tekton", "timeout": 30000})
INFO lightspeed_agentic: [provider:run] thinking: OK, so the results are clear. Let me now note that for web-terminal, the user asked about v1.16.0 but the PLC API only shows up to v1.15. This means v1.16 isn't listed yet but the product itself does have lifecycle data. Let me compile the full report.

Summary:
1. **compliance-operator** v1.9.0 — **HAS DATA**. v1.9 is Full Support, compatible with OCP 4.21.
2. **cluster-logging** v6.5.1 — **HAS DATA**. v6.5 is Full Support, compatible with OCP 4.21.
3. **devworkspace-operator** v0.41.0 — **NO DATA**. Not found in PLC API.
4. **openshift-pipelines-operator-rh** v1.22.0 — **NO DATA**. Not found in PLC API.
5. **web-terminal** v1.16.0 — **HAS DATA**. Product exists but v1.16 not yet listed; latest is v1.15 (Full Support). The product itself has lifecycle data though.

Out of the five products checked, three have lifecycle data available while two don't appear in the PLC API at all.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"olm_check_ran": true, "operators_checked": 5, "operators_with_lifecycle_data": 3, "operators_without_lifecycle_data": 2})
INFO lightspeed_agentic: [provider:run] result: cost=$0.3455, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"olm_check_ran": true, "operators_checked": 5, "operators_with_lifecycle_data": 3, "operators_without_lifecycle_data": 2}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.3455
INFO:     192.168.0.30:43732 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=f8e8a81ac1ee9b6d027af7306f27ee2a)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to look up the lifecycle status of cluster-logging using the product-lifecycle skill. Let me invoke that skill.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "product-lifecycle", "args": "Look up the product \"cluster-logging\" (package name: cluster-logging) in the Red Hat Product Life Cycle API. Report whether it was found, its current lifecycle status, and whether it is compatible with OCP 4.21."})
INFO lightspeed_agentic: [provider:run] thinking: The skill has loaded instructions for querying the Red Hat Product Life Cycle API. Let me now query the API for the cluster-logging package.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=logging+for+Red+Hat+OpenShift\" | jq -r --arg pkg \"cluster-logging\" --arg target \"4.21\" '\n    [.data[] | select(.package == $pkg)] |\n    if length == 0 then \"No Product Life Cycle entry with package=\\($)
INFO lightspeed_agentic: [provider:run] thinking: The API found the product. The installed version is v6.5.1, which maps to the 6.5 version line. Let me report the findings.

- **Product found**: Yes — "logging for Red Hat OpenShift" (package: cluster-logging)
- **Version 6.5**: Full Support
- **OCP 4.21 compatible**: YES

So the status is "supported" (Full Support), it was found, and it is compatible with OCP 4.21.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"product_found": true, "status": "supported", "ocp_compatible": true})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2342, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"product_found": true, "status": "supported", "ocp_compatible": true}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2342
INFO:     192.168.0.30:45774 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=6828efbc1aa5bc304e1d3596e2f4bdee)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to check the lifecycle status of the web-terminal operator for compatibility with OCP 4.21. Let me use the product-lifecycle skill to look this up.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "product-lifecycle", "args": "Check if the Red Hat product \"web-terminal\" (operator package name: web-terminal) has any version that is both \"supported\" (full support phase) and compatible with OCP 4.21. Report all versions found, their support phases, and OCP version compatibility."})
INFO lightspeed_agentic: [provider:run] thinking: Let me query the Red Hat Product Life Cycle API for the web-terminal product.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=web+terminal\" | jq .", "description": "Query PLC API for web-terminal product"})
INFO lightspeed_agentic: [provider:run] thinking: Let me analyze the results. The product "web terminal operator" (package: web-terminal) was found. Now let me check each version for OCP 4.21 compatibility and Full Support status.

Looking at all versions and their `openshift_compatibility`:

| Version | Support Status | OCP Compatibility |
|---------|---------------|-------------------|
| 1.15 | Full Support | 4.20 |
| 1.14 | Maintenance Support | 4.19 |
| 1.13 | Maintenance Support | 4.18 |
| 1.12 | End of life | 4.17 |
| 1.11 | Extended Support | 4.16 |
| 1.10 | End of life | 4.15 |
| 1.9 | End of life | 4.14 |
| 1.8 | End of life | 4.13 |
| 1.7 | End of life | 4.12 |
| 1.6 | End of life | 4.11 |
| 1.5 | End of life | 4.10 |

None of the versions support OCP 4.21 — the highest compatible version is 4.20, which is only available in version 1.15 with Full Support status.

The installed version 1.16.0 doesn't appear in the lifecycle data at all, suggesting it's either very recent or not yet tracked by the API.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"product_found": true, "has_supported_version_for_421": false})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2851, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"product_found": true, "has_supported_version_for_421": false}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2851
INFO:     192.168.0.30:51078 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=37d59e35ff37942b5b841f2098c0fc2d)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to check the lifecycle status of compliance-operator v1.9 using the product-lifecycle skill.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "product-lifecycle", "args": "Check the lifecycle status of compliance-operator version 1.9 (package name: compliance-operator). Report whether the product was found in the Red Hat Product Life Cycle API and its current support phase."})
INFO lightspeed_agentic: [provider:run] thinking: Let me query the Red Hat Product Life Cycle API for the compliance-operator package.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=compliance\" | jq -r --arg pkg \"compliance-operator\" '\n  [.data[] | select(.package == $pkg)] |\n  if length == 0 then \"NO_MATCH\"\n  else .[0] |\n    \"PRODUCT: \\(.name)\",\n    \"PACKAGE: \\(.package // )
INFO lightspeed_agentic: [provider:run] thinking: The API returned data for compliance-operator. Version 1.9 has type "Full Support". Let me report this.

The status should be "supported" since the type is "Full Support".
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"product_found": true, "status": "supported"})
INFO:     192.168.0.30:56932 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [provider:run] result: cost=$0.2286, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"product_found": true, "status": "supported"}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2286
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=247462ecaca4e3a7f50fb953ac8a11a2)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to check the lifecycle status of OCP 4.21 using the product-lifecycle skill. Let me invoke that skill.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "product-lifecycle", "args": "Look up \"Red Hat OpenShift Container Platform\" version 4.21 in the Red Hat Product Life Cycle API. Report whether the product was found, its current lifecycle phase/status, and key dates (general availability, end of full support, end of maintenance, end of )
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to look up "Red Hat OpenShift Container Platform" version 4.21 in the Red Hat Product Life Cycle API. Let me query the API using the instructions from the skill.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=OpenShift+Container+Platform\" | jq -r '.data[] | select(.name == \"Red Hat OpenShift Container Platform\") | {name: .name, versions: [.versions[] | select(.name == \"4.21\")]}'", "description": "Query Red Hat )
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"product_found": true, "status": "supported"})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2344, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"product_found": true, "status": "supported"}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2344
INFO:     192.168.0.30:46154 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=544d53c02541d0205420832e0c2487f1)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to check the lifecycle status of Red Hat OpenShift Container Platform version 4.14 using the product-lifecycle skill.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "product-lifecycle", "args": "Check the lifecycle status of Red Hat OpenShift Container Platform version 4.14. Report its current support phase and any relevant end-of-life dates."})
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to check the lifecycle status of OCP 4.14. Let me query the Red Hat Product Life Cycle API for OpenShift Container Platform and then filter for version 4.14.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=OpenShift+Container+Platform\" | jq -r '.data[] | select(.name == \"Red Hat OpenShift Container Platform\") | .versions[] | select(.name == \"4.14\")'", "description": "Query Red Hat Product Life Cycle API for )
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"product_found": true, "status": "extended"})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2318, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"product_found": true, "status": "extended"}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2318
INFO:     192.168.0.30:55206 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=be270f25d82f8c5b683d96275b701d82)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: The user wants me to check the lifecycle status of two operators against OCP 4.21 using the product-lifecycle skill. Let me invoke that skill.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "product-lifecycle", "args": "Check lifecycle status for these 2 operators against OCP 4.21: 1. cluster-logging (package: cluster-logging) 2. web-terminal (package: web-terminal). For each operator, look up its lifecycle data and report whether both were found in the API and whether any of)
INFO lightspeed_agentic: [provider:run] thinking: Let me query the Red Hat Product Life Cycle API for both operators. I'll start by looking up both operators.
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=logging\" \\\n  | jq -r --arg pkg \"cluster-logging\" --arg target \"4.21\" '\n    [.data[] | select(.package == $pkg)] |\n    if length == 0 then \"No Product Life Cycle entry with package=\\($pkg)\"\n    else)
INFO lightspeed_agentic: [provider:run] tool_use: Bash({"command": "curl -s \"https://access.redhat.com/product-life-cycles/api/v1/products?name=web+terminal\" \\\n  | jq -r --arg pkg \"web-terminal\" --arg target \"4.21\" '\n    [.data[] | select(.package == $pkg)] |\n    if length == 0 then \"No Product Life Cycle entry with package=\\($pkg)\"\n    el)
INFO lightspeed_agentic: [provider:run] thinking: Let me analyze the results:

**cluster-logging (package: cluster-logging)**
- Found in API: YES — "logging for Red Hat OpenShift"
- OCP 4.21 compatible versions:
  - Version 6.5 — Full Support ✅
  - Version 6.4 — Full Support ✅
- None of the OCP 4.21-compatible versions are end-of-life. The EOL versions (5.x, 6.1) are NOT compatible with OCP 4.21.

**web-terminal (package: web-terminal)**
- Found in API: YES — "web terminal operator"
- No versions list OCP 4.21 in their compatibility matrix, so there are no compatible versions to evaluate for EOL status.

For cluster-logging, the OCP 4.21-compatible versions (6.5, 6.4) are both in Full Support — none are EOL. For web-terminal, there are zero OCP 4.21-compatible versions listed at all. So the answer is `any_eol` = false — no OCP 4.21-compatible versions are end-of-life.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"check_completed": true, "both_found": true, "any_eol": false})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2816, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"check_completed": true, "both_found": true, "any_eol": false}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2816
INFO:     192.168.0.30:34748 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO:     Shutting down
INFO:     Waiting for application shutdown.
INFO:     Application shutdown complete.
INFO:     Finished server process [2]

@jrangelramos

Copy link
Copy Markdown
Collaborator Author

All Update Adivisory evals tests pass

🔍 Click to expand - Update Advisory Test Log
bash evals/run.sh -k "update-advisor" --timeout=600
Starting provider containers...
  claude: port 18080 (container 3927ef7b37aa)
Waiting for servers...
  claude: ready

Running evals...
/usr/lib/python3.14/site-packages/pytest_asyncio/plugin.py:211: PytestDeprecationWarning: The configuration option "asyncio_default_fixture_loop_scope" is unset.
The event loop scope for asynchronous fixtures will default to the fixture caching scope. Future versions of pytest-asyncio will default the loop scope for asynchronous fixtures to function scope. Set the default fixture loop scope explicitly in order to avoid unexpected behavior in the future. Valid fixture loop scopes are: "function", "class", "module", "package", "session"

  warnings.warn(PytestDeprecationWarning(_DEFAULT_FIXTURE_LOOP_SCOPE_UNSET))
=============================================================================================== test session starts ================================================================================================
platform linux -- Python 3.14.3, pytest-8.3.5, pluggy-1.6.0 -- /usr/bin/python3
cachedir: .pytest_cache
rootdir: /home/jeramos/pixaa/test/test5-rename/agentic-skills/evals
configfile: pytest.ini
plugins: anyio-4.12.1, asyncio-1.1.0, xdist-3.7.0, timeout-2.4.0
asyncio: mode=Mode.AUTO, asyncio_default_fixture_loop_scope=None, asyncio_default_test_loop_scope=function
timeout: 600.0s
timeout method: signal
timeout func_only: False
collected 228 items / 198 deselected / 30 selected                                                                                                                                                                 

evals/skills/test_eval.py::test_skill[claude-update-advisor-advisor_healthy_cluster_recommend] PASSED                                                                                                        [  3%]
evals/skills/test_eval.py::test_skill[claude-update-advisor-advisor_degraded_operator_warn] PASSED                                                                                                           [  6%]
evals/skills/test_eval.py::test_skill[claude-update-advisor-advisor_api_deprecation_block] PASSED                                                                                                            [ 10%]
evals/skills/test_eval.py::test_skill[claude-update-advisor-advisor_etcd_unhealthy_block] PASSED                                                                                                             [ 13%]
evals/skills/test_eval.py::test_skill[claude-update-advisor-advisor_errored_checks_escalate] PASSED                                                                                                          [ 16%]
evals/skills/test_eval.py::test_skill[gemini-update-advisor-advisor_healthy_cluster_recommend] SKIPPED (No server for gemini)                                                                                [ 20%]
evals/skills/test_eval.py::test_skill[gemini-update-advisor-advisor_degraded_operator_warn] SKIPPED (No server for gemini)                                                                                   [ 23%]
evals/skills/test_eval.py::test_skill[gemini-update-advisor-advisor_api_deprecation_block] SKIPPED (No server for gemini)                                                                                    [ 26%]
evals/skills/test_eval.py::test_skill[gemini-update-advisor-advisor_etcd_unhealthy_block] SKIPPED (No server for gemini)                                                                                     [ 30%]
evals/skills/test_eval.py::test_skill[gemini-update-advisor-advisor_errored_checks_escalate] SKIPPED (No server for gemini)                                                                                  [ 33%]
evals/skills/test_eval.py::test_skill[openai-update-advisor-advisor_healthy_cluster_recommend] SKIPPED (No server for openai)                                                                                [ 36%]
evals/skills/test_eval.py::test_skill[openai-update-advisor-advisor_degraded_operator_warn] SKIPPED (No server for openai)                                                                                   [ 40%]
evals/skills/test_eval.py::test_skill[openai-update-advisor-advisor_api_deprecation_block] SKIPPED (No server for openai)                                                                                    [ 43%]
evals/skills/test_eval.py::test_skill[openai-update-advisor-advisor_etcd_unhealthy_block] SKIPPED (No server for openai)                                                                                     [ 46%]
evals/skills/test_eval.py::test_skill[openai-update-advisor-advisor_errored_checks_escalate] SKIPPED (No server for openai)                                                                                  [ 50%]
evals/skills/test_eval.py::test_skill[deepagents-claude-update-advisor-advisor_healthy_cluster_recommend] SKIPPED (No server for deepagents-claude)                                                          [ 53%]
evals/skills/test_eval.py::test_skill[deepagents-claude-update-advisor-advisor_degraded_operator_warn] SKIPPED (No server for deepagents-claude)                                                             [ 56%]
evals/skills/test_eval.py::test_skill[deepagents-claude-update-advisor-advisor_api_deprecation_block] SKIPPED (No server for deepagents-claude)                                                              [ 60%]
evals/skills/test_eval.py::test_skill[deepagents-claude-update-advisor-advisor_etcd_unhealthy_block] SKIPPED (No server for deepagents-claude)                                                               [ 63%]
evals/skills/test_eval.py::test_skill[deepagents-claude-update-advisor-advisor_errored_checks_escalate] SKIPPED (No server for deepagents-claude)                                                            [ 66%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-update-advisor-advisor_healthy_cluster_recommend] SKIPPED (No server for deepagents-gemini)                                                          [ 70%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-update-advisor-advisor_degraded_operator_warn] SKIPPED (No server for deepagents-gemini)                                                             [ 73%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-update-advisor-advisor_api_deprecation_block] SKIPPED (No server for deepagents-gemini)                                                              [ 76%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-update-advisor-advisor_etcd_unhealthy_block] SKIPPED (No server for deepagents-gemini)                                                               [ 80%]
evals/skills/test_eval.py::test_skill[deepagents-gemini-update-advisor-advisor_errored_checks_escalate] SKIPPED (No server for deepagents-gemini)                                                            [ 83%]
evals/skills/test_eval.py::test_skill[deepagents-openai-update-advisor-advisor_healthy_cluster_recommend] SKIPPED (No server for deepagents-openai)                                                          [ 86%]
evals/skills/test_eval.py::test_skill[deepagents-openai-update-advisor-advisor_degraded_operator_warn] SKIPPED (No server for deepagents-openai)                                                             [ 90%]
evals/skills/test_eval.py::test_skill[deepagents-openai-update-advisor-advisor_api_deprecation_block] SKIPPED (No server for deepagents-openai)                                                              [ 93%]
evals/skills/test_eval.py::test_skill[deepagents-openai-update-advisor-advisor_etcd_unhealthy_block] SKIPPED (No server for deepagents-openai)                                                               [ 96%]
evals/skills/test_eval.py::test_skill[deepagents-openai-update-advisor-advisor_errored_checks_escalate] SKIPPED (No server for deepagents-openai)                                                            [100%]

============================================================================ 5 passed, 25 skipped, 198 deselected in 418.04s (0:06:58) =============================================================================
eval-claude
🔍 Click to expand - Update Advisory Sandbox container logs
$ podman logs -f eval-claude 
INFO lightspeed_agentic.config: Resolved LIGHTSPEED_PROVIDER=anthropic → SDK=claude
INFO lightspeed_agentic.app: Starting app (sdk=claude, model=claude-opus-4-6, LIGHTSPEED_MODEL=, audit=False)
INFO:     Started server process [2]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://0.0.0.0:8080 (Press CTRL+C to quit)
INFO:     192.168.0.30:32940 - "GET /health HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=0a0dd2af533256f707f804b9964a09e9)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: Let me analyze the cluster readiness data systematically.

## Parsing the JSON

**Current version:** 4.21.5
**Target version:** 4.21.16
**Channel:** stable-4.21
**Update type:** Patch (z-stream)

## Check-by-Check Analysis

1. **cluster_conditions** — `_status: ok`
   - Available: True
   - Progressing: False (not currently updating)
   - Upgradeable: True (reason: AsExpected)
   - No update in progress
   - → No blockers, no warnings

2. **operator_health** — `_status: ok`
   - All 34 operators healthy with master and worker control plane nodes fully ready and operational
   - → No blockers, no warnings

3. **api_deprecations** — `_status: ok`
   - No deprecation issues detected
   - → No blockers, no warnings

4. **node_capacity** — `_status: ok`
   - All 6 nodes are ready and available for scheduling
   - → No blockers, no warnings

5. **pdb_drain** — `_status: ok`
   - 21 pod disruption budgets configured with no blocking constraints
   - → No blockers, no warnings

6. **etcd_health** — `_status: ok`
   - All 3 etcd members are healthy and the operator is available and upgradeable
   - → No blockers, no warnings

7. **network** — `_status: ok`
   - Using OVNKubernetes with Intermediate TLS profile
   - → No blockers, no warnings

8. **crd_compat** — `_status: ok`
   - 160 custom resources with no version compatibility issues
   - → No blockers, no warnings

9. **olm_operator_lifecycle** — `_status: ok`
   - No OLM operators installed, so no compatibility concerns to address
   - → No blockers, no warnings

All 9 health checks are passing with no errors detected.

The full assessment took 2.5 seconds to complete. With zero blockers and zero warnings across all checks, the upgrade is safe to proceed.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"decision": "recommend", "blockers_found": 0, "warnings_found": 0})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2109, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"decision": "recommend", "blockers_found": 0, "warnings_found": 0}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2109
INFO:     192.168.0.30:44164 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=9e77bf2630241790ad3c55fc762ab5ed)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: Let me analyze the cluster readiness data and apply the update-advisor skill's decision framework.

Let me parse the JSON data:

1. **Current version**: 4.21.5
2. **Target version**: 4.21.16
3. **Channel**: stable-4.21
4. **Update type**: Patch (within same minor version)

Let me go through each check:

### 1. Cluster Conditions - OK
- Available: True
- Progressing: False
- Upgradeable: True (reason: AsExpected)
- No update in progress
- Recent history: 4.21.5 Completed
### 2. Operator Health - OK (but has a degraded operator)
- **Degraded operators: 1** - `authentication` operator is Degraded (reason: OAuthFlaky, message: "oauth pods intermittently failing")
- Not available: 0
- Not upgradeable: 0
- MCP summary: 0 degraded, 0 paused, 0 updating
- All MCPs (master and worker) are healthy
- Total operators: 34, degraded: 1

### 3. API Deprecations - OK
- Blockers: 0, Warnings: 0

### 4. Node Capacity -
OK - all 6 nodes are ready with no scheduling issues.

### 5. PDB Drain - OK
- 21 PDBs configured, none blocking upgrades

### 6. etcd Health - OK
- All 3 members healthy, operator available and not degraded

### 7. Network - OK
- Using OVNKubernetes networking

### 8. CRD Compatibility - OK
- 160 CRDs present with no version conflicts

### 9. OLM Operator Lifecycle - OK
- No operators pending upgrades

**Blockers**: None - no operators are unupgradeable, no blocking deprecations or PDBs, etcd is healthy, and all nodes are ready.

**Warnings**: The authentication operator is degraded due to intermittent OAuth pod failures. While this isn't blocking the upgrade, it's worth investigating since these failures could escalate during the upgrade process. Based on the decision matrix with zero blockers and one warning, the recommendation is to proceed with caution.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "update-advisor", "args": "Evaluate cluster upgrade readiness from OCP 4.21.5 to 4.21.16 (patch update, stable-4.21 channel). Key findings from readiness data: 1 degraded operator (authentication, reason OAuthFlaky, oauth pods intermittently failing), 0 not-available operators, 0 not-upgra)
INFO lightspeed_agentic: [provider:run] thinking: Now I have the update-advisor skill's classification rules loaded. Let me analyze the cluster readiness data systematically.

Let me go through each check:

### 1. Cluster Conditions — ✅ OK
- Available: True
- Progressing: False (no update in progress)
- Upgradeable: True (reason: AsExpected)
- No update in progress
- Classification: No blocker, no warning

### 2. Operator Health — ⚠️ Warning
- 1 degraded operator: `authentication` (Degraded=True, reason: OAuthFlaky, "oauth pods intermittently failing")
- The classification rules flag any degraded operator as a warning, and I have no unavailable or non-upgradeable operators
- Both master and worker MCPs are healthy and not paused or updating
- Classification: **1 warning** from the degraded authentication operator

### 3. API Deprecations — ✅ OK
- No blockers or warnings detected
- Classification: No issues

### 4. Node Capacity — ✅ OK
- All 6 nodes are ready with no unschedulable nodes

, and there's sufficient headroom across the cluster
- Classification: No issues

### 5. PDB Drain — ✅ OK
- 21 pod disruption budgets with no blocking issues
- Classification: No issues

### 6. etcd Health — ✅ OK
- All 3 etcd members are healthy and running, operator is available and not degraded
- Classification: No issues

### 7. Network — ✅ OK
- Using OVNKubernetes networking
- Classification: No issues

### 8. CRD Compatibility — ✅ OK
- 160 custom resource definitions with no version conflicts
- Classification: No issues

### 9. OLM Operator Lifecycle — ✅ OK
- No OLM-managed operators being tracked
- Classification: No issues

### Summary: No blockers detected, but there's one warning about the authentication operator being degraded, so I'm marking this as a warning-level issue overall.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"decision": "warn", "blockers_found": 0, "degraded_operator_detected": true})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2646, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"decision": "warn", "blockers_found": 0, "degraded_operator_detected": true}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2646
INFO:     192.168.0.30:54302 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=f59061cda0ae0a23bb555ff24361d564)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: Let me analyze the cluster readiness data carefully.

Current version: OCP 4.21.5
Target version: OCP 4.21.16
Channel: stable-4.21
Update type: Patch (within same minor)
Update path: Recommended

Let me go through each check:

1. **cluster_conditions** - OK
   - Available: True
   - Progressing: False (not updating)
   - Upgradeable: True
   - No update in progress
   - Summary: Good

2. **operator_health** - OK
   - 34 operators, 0 degraded, 0 not available, 0 not upgradeable
   - All control plane nodes ready, worker nodes ready, nothing paused or degraded
   - Summary: Good

3. **api_deprecations** - OK (status ok but has findings)
   - One blocker: `flowschemas.v1beta3.flowcontrol.apiserver.k8s.io` is being removed in 4.21.16 with 1250 requests still using it
   - One warning: `cronjobs.v1beta1.batch` deprecated since 4.20 with 42 requests
   - Summary: 1 blocker, 1 warning

The flowschemas API removal is critical since there are 1250 active requests hitting it in the target version. The rest of the checks are passing: nodes are all ready, PDBs aren't blocking drains, etcd is healthy across all three members, networking is using OVNKubernetes, CRDs have no version conflicts, and the operator lifecycle is stable.

So I have one hard blocker from the API compatibility check—the flowschemas issue means I need to use the update-advisor tool to get a proper assessment before proceeding.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "update-advisor", "args": "Assess upgrade readiness for OCP 4.21.5 → 4.21.16 (patch, stable-4.21 channel). Key findings from cluster readiness data: 1) API deprecation BLOCKER: flowschemas.v1beta3.flowcontrol.apiserver.k8s.io is removed in 4.21.16 with 1250 active requests. 2) API deprecat)
INFO lightspeed_agentic: [provider:run] thinking: Now let me analyze the cluster readiness data systematically according to the update-advisor skill's framework.

Let me walk through each check:

### 1. Cluster Conditions
- Available: True ("Done applying 4.21.5")
- Progressing: False ("Cluster version is 4.21.5")
- Upgradeable: True (reason: "AsExpected")
- No update in progress
- **Finding: PASS** — no blocker, no warning

### 2. Operator Health
- 34 total operators
- 0 degraded, 0 not available, 0 not upgradeable
- MCPs are all ready across master and worker nodes with no paused or degraded states
- **Finding: PASS** — no blocker, no warning

### 3. API Deprecations
- `flowschemas.v1beta3.flowcontrol.apiserver.k8s.io` is a blocker — it's been removed in 4.21.16 and has 1250 active requests, which means workloads are actively using an API that won't exist in the target version
- There are also warning-level APIs like `cronjobs.v1beta1.batch` that need attention — deprecated in 4.20 with 42 requests
- **Findings: 1 BLOCKER, 1 WARNING**

### 4. Node Capacity
All 6 nodes are ready with no unschedulable ones, so there's plenty of headroom for the upgrade.
- **Finding: PASS**

### 5. PDB Drain
21 Pod Disruption Budgets are in place with none blocking the drain process.
- **Finding: PASS**

### 6. etcd Health
The etcd cluster has all 3 members healthy and running, with the operator available and not degraded.
- **Finding: PASS**

### 7. Network
The cluster is using OVNKubernetes for networking

.
- **Finding: PASS**

### 8. CRD Compatibility
160 CRDs are present with no version compatibility issues detected.
- **Finding: PASS**

### 9. OLM Operator Lifecycle
No OLM operators are installed or have issues.
- **Finding: PASS**

### Overall Assessment
All 9 checks came back clean with complete data. The blockers and warnings are clear: one removed API still in active use and one deprecated API with 42 active requests. Since there's at least one blocker, the upgrade should be blocked until that's resolved.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"decision": "block", "blockers_found": 1, "api_blocker_detected": true})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2759, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"decision": "block", "blockers_found": 1, "api_blocker_detected": true}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2759
INFO:     192.168.0.30:46914 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=e56075a03cf162111ffe3fbf8f85dfad)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: Let me analyze the cluster readiness data carefully.

Let me parse through each check:

1. **cluster_conditions**: OK
   - Available: True
   - Progressing: False
   - Upgradeable: True
   - No update in progress
   - ✅ No issues

2. **operator_health**: OK
   - 34 operators, 0 degraded, 0 not available, 0 not upgradeable
   - MCPs: master (3/3 ready), worker (3/3 ready), no degraded/paused/updating
   - ✅ No issues
   
3. **api_deprecations**: OK
   - 0 blockers, 0 warnings
   - ✅ No issues

4. **node_capacity**: OK
   - 6/6 nodes ready, 0 unschedulable
   - ✅ No issues

5. **pdb_drain**: OK
   - 21 PDBs, 0 blocking
   - ✅ No issues

6. **etcd_health**: This is where I'm seeing a problem
   - 3 total members but only 2 are healthy
   - etcd-master-2 is not ready, which is a critical issue for cluster stability
   - The operator shows degraded status with "MemberDown" as the reason, and this explicitly qualifies as a blocker

7. **network**: OK
   - OVNKubernetes, not SDN
   - ✅ No issues

8. **crd_compat**: OK
   - 160 CRDs, 0 version issues
   - ✅ No issues

9. **olm_operator_lifecycle**: OK
   - 0 operators tracked, no issues
   - ✅ No issues

The etcd member failure is the blocker I need to flag here.

So I've identified one blocker: the unhealthy etcd-master-2 member. Now I need to invoke the update-advisor skill as instructed to get guidance on how to proceed.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "update-advisor", "args": "Evaluate upgrade readiness for OCP 4.21.5 → 4.21.16 (patch, stable-4.21). Key finding: etcd-master-2 is not ready (healthy_members=2/3, operator Degraded=True, reason=MemberDown). All other checks pass: 34 operators healthy, 6/6 nodes ready, 0 blocking PDBs, 0 AP)
INFO lightspeed_agentic: [provider:run] thinking: Now I have the update-advisor skill's classification table. Let me apply it to the cluster readiness data.

From the classification table:
- **etcd health**: "Blocker if... Any member unhealthy"

etcd-master-2 has `ready: false`, and the etcd operator is `Degraded: True` with reason `MemberDown`. This is clearly a **blocker** per the skill's classification table.

Let me walk through all checks systematically:

1. **Cluster conditions**: Available=True, Progressing=False, Upgradeable=True → ✅ No issue
2. **Operator health**: All operators healthy with no degradation → ✅ No issue
3. **MachineConfigPool**: All pools ready with no paused or degraded states → ✅ No issue
4. **API deprecations**: No blockers or warnings detected → ✅ No issue
5. **Node capacity**: All 6 nodes ready and schedulable → ✅ No issue
6. **PDB config**: No blocking pod disruption budgets → ✅ No issue
7. **etcd health**: etcd-master-2 is down with a member failure

, operator marked as degraded → ❌ **BLOCKER**
8. **Network**: Using OVNKubernetes networking → ✅ No issue
9. **CRD compatibility**: No version conflicts → ✅ No issue
10. **OLM operator lifecycle**: All operators functioning normally → ✅ No issue

I've identified one critical blocker preventing the upgrade: the etcd cluster has a failed member that needs to be recovered before proceeding.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"decision": "block", "etcd_issue_detected": true, "unhealthy_member": "etcd-master-2"})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2619, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"decision": "block", "etcd_issue_detected": true, "unhealthy_member": "etcd-master-2"}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2619
INFO:     192.168.0.30:35462 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO lightspeed_agentic: [agent] Starting query (model=claude-opus-4-6, provider=claude, trace_id=3a038ba358eff7ae55cce93bd6644509)
INFO claude_agent_sdk._internal.transport.subprocess_cli: Using bundled Claude Code CLI: /opt/app-root/lib64/python3.12/site-packages/claude_agent_sdk/_bundled/claude
INFO lightspeed_agentic: [provider:run] thinking: Let me analyze the cluster readiness data carefully.

Parsing the JSON:

- **current_version**: 4.21.5
- **target_version**: 4.21.16
- **Total checks**: 9
- **Checks OK**: 2
- **Checks errored**: 7

Let me go through each check:

1. **cluster_conditions** - `_status: "ok"` - Cluster is Available, not Progressing, Upgradeable is True. ✅
2. **operator_health** - `_status: "error"` - "connection refused" ❌
3. **api_deprecations** - `_status: "error"` - "connection refused" ❌
4. **node_capacity** - `_status: "error"` - "connection refused" ❌
5. **pdb_drain** - `_status: "error"` - "connection refused" ❌
6. **etcd_health** - `_status: "error"` - "connection refused" ❌
7. **network** - `_status: "ok"` - OVNKubernetes, not SDN ✅
8. **crd_compat** - `_status: "error"` - Still checking the remaining failures...

Most of the checks are hitting connection refused errors, which points to a connectivity problem with the cluster API. Since I can't get reliable data from 7 of 9 checks, I need to escalate this rather than make an unreliable assessment. Let me start by invoking the update-advisor skill.
INFO lightspeed_agentic: [provider:run] tool_use: Skill({"skill": "update-advisor", "args": "Evaluate cluster upgrade readiness from OCP 4.21.5 to 4.21.16 (patch, stable-4.21 channel). 7 of 9 checks errored with \"connection refused\" after 60s timeouts. Only cluster_conditions (ok, upgradeable=True) and network (ok, OVNKubernetes) succeeded. Errored che)
INFO lightspeed_agentic: [provider:run] thinking: Now I have the update-advisor skill instructions. Let me analyze the cluster readiness data systematically.

## Analysis

### Data Completeness
- Total checks: 9
- Checks OK: 2 (cluster_conditions, network)
- Checks errored: 7 (operator_health, api_deprecations, node_capacity, pdb_drain, etcd_health, crd_compat, olm_operator_lifecycle)

All 7 errored checks failed with "connection refused" after 60-second timeouts. This is a systemic connectivity failure to the cluster API server.

### Checks that passed:
The cluster_conditions check shows the cluster is at version 4.21.5 with Available=True and Upgradeable=True, indicating it's ready from a status perspective. The network check confirms OVNKubernetes is in use, which isn't a deprecated SDN option.

### Checks that errored:
The remaining seven checks all failed due to connectivity issues—operator_health can't be verified (critical for update safety), api_deprecations can't be checked (critical for workload compatibility), node_capacity can't be assessed (important for rolling updates), and the others similarly blocked by the same connection failures.

### Decision
With 7 of 9 checks unable to complete, I'm looking at the decision matrix: when blockers can't be assessed, the recommendation is to escalate. The connection refused errors mean I can't verify operator health, etcd health, API deprecations, node capacity, PDB configurations, CRD compatibility, or OLM operator lifecycle—all critical areas for a safe upgrade. This clearly falls into "unable to assess" territory, so the decision is to escalate.
INFO lightspeed_agentic: [provider:run] tool_use: StructuredOutput({"decision": "escalate", "checks_errored": 7})
INFO lightspeed_agentic: [provider:run] result: cost=$0.2545, tokens=0
INFO lightspeed_agentic: [provider:run] output: {"decision": "escalate", "checks_errored": 7}
INFO lightspeed_agentic: [agent] query complete: success=True, cost=$0.2545
INFO:     192.168.0.30:37558 - "POST /v1/agent/run HTTP/1.1" 200 OK
INFO:     Shutting down
INFO:     Waiting for application shutdown.
INFO:     Application shutdown complete.
INFO:     Finished server process [2]

@jrangelramos

Copy link
Copy Markdown
Collaborator Author

This changes should address failures on eval exeuctions of openshift#34

@jhadvig jhadvig left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@jhadvig
jhadvig merged commit 040140f into jhadvig:jhadvig/ota-2024-cluster-update-evals Jul 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants