CPU-based
Scale up when average CPU exceeds your target (e.g. 70%).
Horizontal autoscaling adds replicas to your deployment under load and removes them when traffic drops, based on live CPU and memory usage. Turn it on for any web or worker deployment and Ownkube owns the replica count from then on.
Enable it from the deployment Settings tab in your dashboard.
Autoscaling watches your containers’ average CPU and memory usage. When usage crosses a target threshold, it adds replicas; when usage drops back below the threshold, it removes them. Scaling reacts within seconds of a sustained change.
CPU-based
Scale up when average CPU exceeds your target (e.g. 70%).
Memory-based
Scale up when average memory exceeds your target (e.g. 80%).
You can use one signal, both, or neither. Autoscaling acts on whichever threshold is crossed first.
For a web service that sees variable traffic through the day:
| Setting | Value |
|---|---|
| Min replicas | 2 |
| Max replicas | 10 |
| Target CPU utilization | 70% |
| Target memory utilization | 80% |
| CPU request | 250m |
| Memory request | 256Mi |
This keeps at least two replicas always warm, bursts up to ten under load, and scales back down once traffic drops.
Once autoscaling is enabled, the fixed replica count field is ignored. Autoscaling owns replica count from that point forward. Turn autoscaling off to return to a manual replica count.
On Ownkube-hosted compute, the capacity underneath your replicas scales automatically. There’s nothing to configure: you set the replica range and Ownkube handles the rest.