Skip to content

chore: define default deployment scope and tune operator resource limits #407

Description

@bjosv

The operator's resource limits in config/manager/manager.yaml are the untouched kubebuilder scaffold defaults, with a TODO(user) comment. They are not based on measured profiles:

resources:
  limits:
    cpu: 500m
    memory: 128Mi
  requests:
    cpu: 10m
    memory: 64Mi

Profiling against v0.5.0 shows the memory limit may be too tight depending on the deployment scope, and the CPU request is too low for active reconciliation. The current 128 Mi limit OOMKills during scaling beyond ~20 shards.

What needs to happen

  1. Define the default deployment scope — the maximum cluster size (shards, replicas, number of clusters) the operator must handle out of the box without manual resource tuning. This is the baseline all default limits are sized against.

  2. Tune resource limits and requests to match that scope, based on the profiling data.

  3. Document guidance for users who exceed the default scope (e.g. "if scaling beyond N shards, raise the operator memory limit to X").

Profiling summary (v0.5.0)

The binding constraint is transient allocation during slot rebalancing, not steady-state pod count.

Steady state (no scaling)

workload pods peak RSS outcome
40 clusters, 3 shards each 240 83 MB healthy
100 clusters, 3 shards each 498 124 MB healthy

Scaling (active slot rebalancing)

workload pods peak RSS outcome
1 cluster, 1 to 3 shards 6 60 MB healthy
1 cluster, 1 to 23 shards 46 138 MB OOMKilled
4 clusters, 1 to 19 shards 152 120 MB OOMKilled at shard 19
10 clusters, 1 to 16 shards 320 158 MB OOMKilled at shard 16, host CPU saturated

At the possible default scope (10 clusters × 5 shards = 100 pods), peak RSS was 90 MB with 128 Mi limit — comfortable headroom.

Key observations:

  • Fixed baseline is ~60 MB (Go runtime, informer cache). Nearly half the 128 Mi budget is consumed before any workload.
  • Holding pods is cheap (~0.075 MB per pod). Rebalancing is expensive: heap hit 109 MB at just 42 pods during slot migration.
  • 128 Mi tolerates roughly 20 shards while actively scaling. The GC sawtooth is ~20 MB, so any limit needs that much headroom above measured peaks.
  • CPU request of 10m (0.01 cores) is effectively zero; the operator sustains 0.02–0.04 cores at 3 shards and 0.2–0.3 cores during active rebalancing at 15–18 shards, so it gets no meaningful scheduling guarantee under node pressure.

MaxConcurrentReconciles

Other operators increase MaxConcurrentReconciles to improve multi-cluster throughput. Testing shows this is not beneficial at the current memory budget. With concurrency=2, OOM occurred earlier (shard 14 vs 19 at concurrency=1). With concurrency=10, OOM occurred immediately before any work started. Parallel reconciles stack heap allocations and amplify them through cascading watch events.

The per-reconcile bottleneck is serial I/O (GetClusterState connects to every node one by one), not queueing. Parallelizing GetClusterState internally would improve single-reconcile speed and is a prerequisite before MaxConcurrentReconciles > 1 becomes worthwhile. Without that, increasing concurrency just trades memory for no wall-time gain.

If concurrency is raised in the future, memory limits must be raised proportionally (roughly 2x memory per doubling of concurrency during scaling workloads).

Cluster state memory cost

GetClusterState connects to every node and parses CLUSTER NODES output from each. For a 23-shard/1-replica cluster (46 nodes), that is 46 connections each returning 46 lines of topology. This is the dominant source of transient heap allocation during reconciliation and the reason rebalancing is so much more expensive than steady state — each rebalance pass re-scrapes the full topology.

Potential optimizations to keep in mind when calibrating resource limits:

  • Store only the fields needed from CLUSTER NODES (node ID, address, role, slots) rather than retaining the full raw output per node
  • Reduce rebalanceSlotBatchSize (currently 400) to lower peak allocation per rebalance pass at the cost of slower rebalancing

Any of these would lower the memory ceiling and allow tighter resource limits for the same deployment scope.

Suggested starting point

Once the default scope is defined, a reasonable starting point (covering moderate scaling):

A possible default scope is ~10 clusters, 3–5 shards each, 1 replica (up to ~100 pods). Measured peak for this range is ~80–90 MB RSS. 256 Mi provides comfortable headroom including GC sawtooth and room to slightly exceed the default scope without manual tuning.

resources:
  limits:
    cpu: 500m
    memory: 256Mi
  requests:
    cpu: 100m
    memory: 128Mi
  • memory limit 256 Mi: baseline is 60 MB, GC sawtooth is 20 MB, scaling to ~10–15 shards adds another 50–80 MB transiently. 256 Mi gives headroom for moderate scaling. 128 Mi OOMs at ~20 shards.
  • memory request 128 Mi: reflects the actual baseline (60 MB) + GC sawtooth (20 MB) + minimal working set. Gives the scheduler an honest signal.
  • cpu request 100m (0.1 cores): measured 0.02–0.04 at steady state, 0.2–0.3 during scaling. Guarantees enough for light reconciliation without over-reserving.
  • cpu limit 500m: unchanged, never the bottleneck in any test.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions