The operator's resource limits in config/manager/manager.yaml are the untouched kubebuilder scaffold defaults, with a TODO(user) comment. They are not based on measured profiles:
resources:
limits:
cpu: 500m
memory: 128Mi
requests:
cpu: 10m
memory: 64Mi
Profiling against v0.5.0 shows the memory limit may be too tight depending on the deployment scope, and the CPU request is too low for active reconciliation. The current 128 Mi limit OOMKills during scaling beyond ~20 shards.
What needs to happen
-
Define the default deployment scope — the maximum cluster size (shards, replicas, number of clusters) the operator must handle out of the box without manual resource tuning. This is the baseline all default limits are sized against.
-
Tune resource limits and requests to match that scope, based on the profiling data.
-
Document guidance for users who exceed the default scope (e.g. "if scaling beyond N shards, raise the operator memory limit to X").
Profiling summary (v0.5.0)
The binding constraint is transient allocation during slot rebalancing, not steady-state pod count.
Steady state (no scaling)
| workload |
pods |
peak RSS |
outcome |
| 40 clusters, 3 shards each |
240 |
83 MB |
healthy |
| 100 clusters, 3 shards each |
498 |
124 MB |
healthy |
Scaling (active slot rebalancing)
| workload |
pods |
peak RSS |
outcome |
| 1 cluster, 1 to 3 shards |
6 |
60 MB |
healthy |
| 1 cluster, 1 to 23 shards |
46 |
138 MB |
OOMKilled |
| 4 clusters, 1 to 19 shards |
152 |
120 MB |
OOMKilled at shard 19 |
| 10 clusters, 1 to 16 shards |
320 |
158 MB |
OOMKilled at shard 16, host CPU saturated |
At the possible default scope (10 clusters × 5 shards = 100 pods), peak RSS was 90 MB with 128 Mi limit — comfortable headroom.
Key observations:
- Fixed baseline is ~60 MB (Go runtime, informer cache). Nearly half the 128 Mi budget is consumed before any workload.
- Holding pods is cheap (~0.075 MB per pod). Rebalancing is expensive: heap hit 109 MB at just 42 pods during slot migration.
- 128 Mi tolerates roughly 20 shards while actively scaling. The GC sawtooth is ~20 MB, so any limit needs that much headroom above measured peaks.
- CPU request of 10m (0.01 cores) is effectively zero; the operator sustains 0.02–0.04 cores at 3 shards and 0.2–0.3 cores during active rebalancing at 15–18 shards, so it gets no meaningful scheduling guarantee under node pressure.
MaxConcurrentReconciles
Other operators increase MaxConcurrentReconciles to improve multi-cluster throughput. Testing shows this is not beneficial at the current memory budget. With concurrency=2, OOM occurred earlier (shard 14 vs 19 at concurrency=1). With concurrency=10, OOM occurred immediately before any work started. Parallel reconciles stack heap allocations and amplify them through cascading watch events.
The per-reconcile bottleneck is serial I/O (GetClusterState connects to every node one by one), not queueing. Parallelizing GetClusterState internally would improve single-reconcile speed and is a prerequisite before MaxConcurrentReconciles > 1 becomes worthwhile. Without that, increasing concurrency just trades memory for no wall-time gain.
If concurrency is raised in the future, memory limits must be raised proportionally (roughly 2x memory per doubling of concurrency during scaling workloads).
Cluster state memory cost
GetClusterState connects to every node and parses CLUSTER NODES output from each. For a 23-shard/1-replica cluster (46 nodes), that is 46 connections each returning 46 lines of topology. This is the dominant source of transient heap allocation during reconciliation and the reason rebalancing is so much more expensive than steady state — each rebalance pass re-scrapes the full topology.
Potential optimizations to keep in mind when calibrating resource limits:
- Store only the fields needed from
CLUSTER NODES (node ID, address, role, slots) rather than retaining the full raw output per node
- Reduce
rebalanceSlotBatchSize (currently 400) to lower peak allocation per rebalance pass at the cost of slower rebalancing
Any of these would lower the memory ceiling and allow tighter resource limits for the same deployment scope.
Suggested starting point
Once the default scope is defined, a reasonable starting point (covering moderate scaling):
A possible default scope is ~10 clusters, 3–5 shards each, 1 replica (up to ~100 pods). Measured peak for this range is ~80–90 MB RSS. 256 Mi provides comfortable headroom including GC sawtooth and room to slightly exceed the default scope without manual tuning.
resources:
limits:
cpu: 500m
memory: 256Mi
requests:
cpu: 100m
memory: 128Mi
- memory limit 256 Mi: baseline is 60 MB, GC sawtooth is 20 MB, scaling to ~10–15 shards adds another 50–80 MB transiently. 256 Mi gives headroom for moderate scaling. 128 Mi OOMs at ~20 shards.
- memory request 128 Mi: reflects the actual baseline (60 MB) + GC sawtooth (20 MB) + minimal working set. Gives the scheduler an honest signal.
- cpu request 100m (0.1 cores): measured 0.02–0.04 at steady state, 0.2–0.3 during scaling. Guarantees enough for light reconciliation without over-reserving.
- cpu limit 500m: unchanged, never the bottleneck in any test.
The operator's resource limits in
config/manager/manager.yamlare the untouched kubebuilder scaffold defaults, with aTODO(user)comment. They are not based on measured profiles:Profiling against v0.5.0 shows the memory limit may be too tight depending on the deployment scope, and the CPU request is too low for active reconciliation. The current 128 Mi limit
OOMKillsduring scaling beyond ~20 shards.What needs to happen
Define the default deployment scope — the maximum cluster size (shards, replicas, number of clusters) the operator must handle out of the box without manual resource tuning. This is the baseline all default limits are sized against.
Tune resource limits and requests to match that scope, based on the profiling data.
Document guidance for users who exceed the default scope (e.g. "if scaling beyond N shards, raise the operator memory limit to X").
Profiling summary (v0.5.0)
The binding constraint is transient allocation during slot rebalancing, not steady-state pod count.
Steady state (no scaling)
Scaling (active slot rebalancing)
At the possible default scope (10 clusters × 5 shards = 100 pods), peak RSS was 90 MB with 128 Mi limit — comfortable headroom.
Key observations:
MaxConcurrentReconciles
Other operators increase
MaxConcurrentReconcilesto improve multi-cluster throughput. Testing shows this is not beneficial at the current memory budget. With concurrency=2, OOM occurred earlier (shard 14 vs 19 at concurrency=1). With concurrency=10, OOM occurred immediately before any work started. Parallel reconciles stack heap allocations and amplify them through cascading watch events.The per-reconcile bottleneck is serial I/O (
GetClusterStateconnects to every node one by one), not queueing. ParallelizingGetClusterStateinternally would improve single-reconcile speed and is a prerequisite beforeMaxConcurrentReconciles > 1becomes worthwhile. Without that, increasing concurrency just trades memory for no wall-time gain.If concurrency is raised in the future, memory limits must be raised proportionally (roughly 2x memory per doubling of concurrency during scaling workloads).
Cluster state memory cost
GetClusterStateconnects to every node and parsesCLUSTER NODESoutput from each. For a 23-shard/1-replica cluster (46 nodes), that is 46 connections each returning 46 lines of topology. This is the dominant source of transient heap allocation during reconciliation and the reason rebalancing is so much more expensive than steady state — each rebalance pass re-scrapes the full topology.Potential optimizations to keep in mind when calibrating resource limits:
CLUSTER NODES(node ID, address, role, slots) rather than retaining the full raw output per noderebalanceSlotBatchSize(currently 400) to lower peak allocation per rebalance pass at the cost of slower rebalancingAny of these would lower the memory ceiling and allow tighter resource limits for the same deployment scope.
Suggested starting point
Once the default scope is defined, a reasonable starting point (covering moderate scaling):
A possible default scope is ~10 clusters, 3–5 shards each, 1 replica (up to ~100 pods). Measured peak for this range is ~80–90 MB RSS. 256 Mi provides comfortable headroom including GC sawtooth and room to slightly exceed the default scope without manual tuning.