Skip to content

✨ feat: introduce NodeReadinessEvaluation (NRE) CRD and controller - #345

Open
Karthik-K-N wants to merge 1 commit into
kubernetes-sigs:mainfrom
Karthik-K-N:feat-nre
Open

✨ feat: introduce NodeReadinessEvaluation (NRE) CRD and controller#345
Karthik-K-N wants to merge 1 commit into
kubernetes-sigs:mainfrom
Karthik-K-N:feat-nre

Conversation

@Karthik-K-N

@Karthik-K-N Karthik-K-N commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Description

Introduces the NodeReadinessEvaluation (NRE) custom resource and its dedicated reconciler. NRE provides a per-node, read-only mirror of the evaluated state of every applicable NodeReadinessRule, giving operators a single object to kubectl get nre <node-name> to see the full readiness picture of any node without having to cross-reference multiple rule statuses.

Discussion document: https://docs.google.com/document/d/1DOP1G6i__nQN8qbhlSvdswkUquMeSKQY4LU554AbcgM/edit?usp=sharing

Feature Flag

The controller is opt-in via: --enable-node-readiness-evaluation

Related Issue

Type of Change

Testing

Checklist

  • make test passes
  • make lint passes

Does this PR introduce a user-facing change?


Doc #(issue)

@netlify

netlify Bot commented Aug 4, 2026

Copy link
Copy Markdown

Deploy Preview for node-readiness-controller ready!

Name Link
🔨 Latest commit e33efec
🔍 Latest deploy log https://app.netlify.com/projects/node-readiness-controller/deploys/6a8c6c0edd3ef7000852ad58
😎 Deploy Preview https://deploy-preview-345--node-readiness-controller.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@kubernetes-prow

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: Karthik-K-N
Once this PR has been reviewed and has the lgtm label, please assign sergeykanzhelev for approval. For more information see the Code Review Process.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@kubernetes-prow
kubernetes-prow Bot requested a review from dchen1107 August 4, 2026 08:44
@kubernetes-prow kubernetes-prow Bot added the cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. label Aug 4, 2026
@kubernetes-prow
kubernetes-prow Bot requested a review from tallclair August 4, 2026 08:44
@kubernetes-prow kubernetes-prow Bot added the size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. label Aug 4, 2026
@Karthik-K-N
Karthik-K-N marked this pull request as draft August 5, 2026 12:18
@kubernetes-prow kubernetes-prow Bot added the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Aug 5, 2026
@ajaysundark
ajaysundark self-requested a review August 8, 2026 05:24
// - Node objects (conditions, taints, labels)
// - NodeReadinessRule objects (enqueues all nodes matching the changed rule)
func (r *NodeReadinessEvaluationReconciler) SetupWithManager(mgr ctrl.Manager) error {
return ctrl.NewControllerManagedBy(mgr).

@ajaysundark ajaysundark Aug 9, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Karthik-K-N IIUC, this would trigger two parallel reconciliations for node updates (NodeReconciler and NREReconciler) and one is handling the output of the other. I wonder if a separate controller for NRE is a better fitting pattern for this or should be bridged into NodeReconciler's watch?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes I thought about that and I missed to point out why I chose this way, here are my thoughts

  1. Initially tried combining both into the NodeReconciler, but then when I checked with best practices of writing controller, if we do so we will be merging two different operations into one, Node controller adding/removing taints,NRE just storing the result, One failure should not cause requeue for other and I think NodeReconcile should be as quick as possible as it affects the workloads.
  2. Followed the existing pattern of having separate controllers like NRR, Node and NRE
  3. Even if there is race and NRE reconciles first and later the Node, Since the NRE also watches for condition, taint and label, if something updated by the Node Rec then eventually it triggers the NRE. We will be eventually consistent.

These are the thoughts and I lean towards keeping them separate, but let me know what do you think.

@ajaysundark ajaysundark Aug 24, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My concern with another reconciler is that we are reintroducing a shared (Node) state again between two controllers. looking at buildRuleEvaluation it is very similar to evaluateRuleForNode. reevaluating the node status between two controllers with caches seem like a bad idea (and that seems like repeating what we did and trying to change with rule.status). :/

I saw #343 emitting individual nodeStatusDelta in rule-reconciliation, which could make adding NRE even cheaper than before. Happy to discuss more / hear your thoughts further on this.

@DsThakurRawat

Copy link
Copy Markdown
Contributor

i tested this branch and found two behaviours worth knowing about, plus one interaction with #315. all three are reproduced with runnable tests, happy to share them.

first, a rule-creation window that never heals. the NRE fan-out fires on the rule Create event (GenerationChangedPredicate only filters updates), but the shared ruleCache is only populated on the rule controller's second reconcile pass, because the first one adds the finalizer and returns RequeueAfter one second. so the NRE reconciles for every matching node run against a cache that cannot contain the new rule yet. for rules whose evaluation then changes a node, the taint write re-triggers the node watch and everything heals. but for a rule whose conditions are already satisfied everywhere, nothing mutates any node, the rule's status patch doesn't bump generation, heartbeats don't pass the node predicate (it compares only condition type to status), and informer resyncs are filtered as no-ops. the NRE just permanently misses the rule until some unrelated node change. deletion is fine, reconcileDelete empties the cache before dropping the finalizer. i think this is the concrete version of the question ajaysundark raised above: the reconciler's correctness depends on another controller's queue having run first. listing rules from the informer inside Reconcile instead of reading the private cache would remove the ordering dependency entirely.

second, this branch merges cleanly with #315, and the two disagree once combined. evaluateRuleForNode there honours conditionPolicy anyOf, while buildRuleEvaluation here ANDs every condition unconditionally. an anyOf rule with one of two conditions satisfied ends up enforced as satisfied (no taint) while the NRE for the same node reports it Unmatched. a shared evaluation helper honouring GetConditionPolicy() in both paths would keep them from drifting.

small one: nothing currently produces RuleStatusError, so the Errors count, State Pending, and the Evaluated False condition are unreachable in this iteration. fine if that's intentional groundwork, might deserve a TODO.

tejassinghbhati added a commit to tejassinghbhati/nrc-explorer that referenced this pull request Aug 14, 2026
The controller side of the NRE proposal is still a draft, so nothing
populates these objects yet. Writing them here makes the per-node
datasource measurable now, since the payload is decided by the schema
and by how many rules apply to a node.

Refs kubernetes-sigs/node-readiness-controller#345
@ajaysundark

Copy link
Copy Markdown
Contributor

@Karthik-K-N is this ready for review?

@Karthik-K-N
Karthik-K-N marked this pull request as ready for review August 21, 2026 04:35
@kubernetes-prow kubernetes-prow Bot removed the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Aug 21, 2026
@Karthik-K-N

Copy link
Copy Markdown
Contributor Author

@Karthik-K-N is this ready for review?

yes , its ready for review

@DsThakurRawat

Copy link
Copy Markdown
Contributor

The rule-creation window I flagged here on Aug 11 has a second victim now, and this one is worse because it's the reporting surface. Walked the sequence on your head 7c0ecc42: a rule Create passes GenerationChangedPredicate into this controller too, ruleToNodeRequests enqueues the matching nodes immediately, but the NRE reconciler reads getApplicableRulesForNode, which serves only the private ruleCache, and the cache isn't populated until the rule controller's second reconcile, because the first one just adds the finalizer and returns RequeueAfter one second (nodereadinessrule_controller.go:118-125, cache write at :139). So for a rule whose conditions are already satisfied on every matching node and whose taint is absent everywhere, the NRE usually reconciles against the empty cache and writes state: Ready with zero rule entries. Nothing corrects it after that: evaluation mutates no node, so no node event fires, and resync updates die on your equality predicates in the For(). Until something else happens to touch that node, the NRE keeps telling any reader (#327's history story, dashboards) that no rule applies to it while one actively does.

The Aug 11 fix suggestion covers both consumers, and it got cheaper to justify: have this controller read rules through the informer instead of the private cache, and the ordering dependency disappears rather than getting patched around.

Separately, I reproduced prow's lint failure locally so it can be named precisely: gocritic ifElseChain on the four-way taint transition at nodereadinessevaluation_controller.go:290, and unparam's unused ctx in the test's sharedSetup at :91. Both trivial.

What I checked that holds: rule deletions do reach this controller (in controller-runtime v0.24.1 a nil DeleteFunc defaults to passing, so GenerationChangedPredicate only filters updates), and your D1 shows the prune works once a reconcile runs; the CEL immutability on nodeSelector makes mapping by selector safe across rule updates; node deletion is covered by the ownerReference; and the single-writer separation is real, this reconciler writes nothing but its own status. One RBAC nit: the marker asks delete on nodereadinessevaluations but nothing in the code ever calls it, GC does that work.


for _, r := range status.Rules {
switch r.RuleStatus {
case readinessv1alpha1.RuleStatusMatched:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Matched / Unmatched rules on status is bit unclear to me. Sorry I missed to get this at the doc, will look into it again.

@Karthik-K-N Karthik-K-N Aug 24, 2026

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So the RuleStatus field helps to understand "whether node satisfy every condition listed in the rule"
Matched - All rule.Spec.Conditions are satisfied - ideally no taint
Unmatched - Any condition is not satisfied - there is taint on node

Do you think we should use better field name to convey this?

Instead of Matched should I change that to RuleStatusSatisfied means the rule is satisfied by the node?


status.Summary = readinessv1alpha1.EvaluationSummary{
MatchedRules: &matched,
UnmatchedRules: &unmatched,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

does 'unmatched' (node-selector escaped?) add value to the per-node status? As a node-owner I would less-likely be interested in rules that are not affecting my node.

Imagine a large cluster with many different node-pools. Setting the status of a node in pool-A for other rules associated with every other pools may be rather confusing to the user.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

True, I think we can drop the UnmatchedRules field.

// SetupWithManager wires the reconciler to watch:
// - Node objects (conditions, taints, labels)
// - NodeReadinessRule objects (enqueues all nodes matching the changed rule)
func (r *NodeReadinessEvaluationReconciler) SetupWithManager(mgr ctrl.Manager) error {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I expect, adding another controller will also increase the API interaction volume to a larger extent at scale.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think if we can finalize on this aspect, then other things will be easy to tackle
Since we use controller-runtime and it uses a Shared Informer Cache, additional controller does not increase API server watch traffic or our memory footprint. but It does duplicate the reconcile queue (may be CPU usage) and small window of race, but it provides less blast radius and separate of concern, we have two controller doing different things.
Also I think we should consdier at what scale level we can expect this latency of having new controller

@ajaysundark ajaysundark left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for looking into this @Karthik-K-N!

Left some comments, mostly on the need for separate controller loop for NRE. We could find sometime to align on the implementation plan to avoid adding to your work and delaying this.

@ajaysundark

Copy link
Copy Markdown
Contributor

NRE reconciler reads getApplicableRulesForNode, which serves only the private ruleCache, and the cache isn't populated until the rule controller's second reconcile

Thanks for the callout @DsThakurRawat. agree, there are some risks with a managed cache and two controllers reconciling at two different snapshots of the cached data. we need to carefully handle the race-conditions.

@kubernetes-prow

Copy link
Copy Markdown

@Karthik-K-N: The following test failed, say /retest to rerun all failed tests or /retest-required to rerun all mandatory failed tests:

Test name Commit Details Required Rerun command
pull-node-readiness-controller-test-e2e e33efec link true /test pull-node-readiness-controller-test-e2e

Full PR test history. Your PR dashboard. Please help us cut down on flakes by linking to an open issue when you hit one in your PR.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@yindia

yindia commented Aug 25, 2026

Copy link
Copy Markdown

@ajaysundark On the cache-ordering window: reading rules via an informer List in Reconcile, instead of the private ruleCache, closes it. The fan-out fires on the rule watch event, so the informer already has the new rule, whereas ruleCache isn't written until RuleReconciler's second pass. It also drops the ordering dependency, with no need to remove ruleCache itself. wdyt ?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants