Skip to content

Autoscaling Your App

Horizontal Pod Autoscaler (HPA) changes the number of running pods in response to average CPU and, optionally, memory utilization. Use it for stateless workloads whose demand changes over time. This guide is about capacity control; it does not change your rollout strategy. See the deployment guide for rolling-update settings.

HPA needs a running app, resource requests, and cluster metrics. Satusky configures CPU and memory requests from satusky.toml or the matching 1ctl deploy flags. The cluster must expose the Kubernetes resource-metrics API:

Terminal window
kubectl get apiservice v1beta1.metrics.k8s.io
kubectl top pods -n <your-organization-namespace>

Available=True on the API service and pod metrics from kubectl top are the prerequisites for utilization-based scaling. If either check is unavailable, the HPA object can exist but cannot make informed scaling decisions.

1ctl deploy --wait persists the HPA intent and reconciles an owned autoscaling/v2 HorizontalPodAutoscaler. Still run the rendered-contract check below before relying on autoscaling: it confirms the final bounds, target metrics, and any Kubernetes condition that prevents scaling.

Use the supported [hpa] section alongside an explicit CPU request. This example maintains one to three API pods, aiming for 60% average CPU utilization. Memory scaling is disabled when memory_target is 0.

[app]
name = "my-api"
port = 8080
cpu_request = "250m"
cpu_limit = "1"
memory = "256Mi"
[hpa]
enabled = true
min_replicas = 1
max_replicas = 3
cpu_target = 60
memory_target = 0

Deploy or update the app with that file:

Terminal window
1ctl deploy --config satusky.toml --wait

CLI flags are useful for a one-off update and take precedence over values in the file:

Terminal window
1ctl deploy --config satusky.toml \
--hpa \
--hpa-min-replicas 1 \
--hpa-max-replicas 3 \
--hpa-cpu-target 60 \
--hpa-memory-target 0 \
--wait

The public defaults when HPA is enabled are min 1, max 10, and CPU target 80. --hpa-memory-target 0 means no memory metric is added. HPA min replicas must not exceed max replicas.

Set the app and namespace once, then inspect both the control-plane record and the Kubernetes object:

Terminal window
APP=my-api
NAMESPACE=<your-organization-namespace>
1ctl app status "$APP"
kubectl -n "$NAMESPACE" get hpa "${APP}-hpa" -o json | jq '{
minReplicas: .spec.minReplicas,
maxReplicas: .spec.maxReplicas,
metrics: .spec.metrics,
currentReplicas: .status.currentReplicas,
desiredReplicas: .status.desiredReplicas,
conditions: .status.conditions
}'

When the object exists, the HPA targets Deployment/$APP. Its conditions are the authoritative explanation of whether it can scale. In particular, check for AbleToScale=True and ScalingActive=True; a metrics error or missing resource request will appear in the condition message.

Terminal window
kubectl -n "$NAMESPACE" describe hpa "${APP}-hpa"
kubectl -n "$NAMESPACE" get deployment "$APP" -o jsonpath='{.spec.template.spec.containers[0].resources.requests}{"\n"}'

Only after the HPA object exists does it own the Deployment replica count. Do not use 1ctl app scale for that app; update the HPA bounds through satusky.toml or the HPA flags and deploy the change.

Set enabled = false in the same [hpa] section, keep an explicit app.replicas value, and deploy the update:

[app]
replicas = 1
[hpa]
enabled = false
Terminal window
1ctl deploy --config satusky.toml --wait
kubectl -n "$NAMESPACE" get hpa "${APP}-hpa"
kubectl -n "$NAMESPACE" rollout status "deployment/$APP" --timeout=2m

The HPA command should return NotFound, while the Deployment remains present at the configured replica count. The platform removes only its owned HPA; do not delete the HPA manually while it is still enabled in the deployment configuration.

Run a bounded CPU test only in a non-production deployment that you are authorized to disturb. This example generates CPU in one app pod for two minutes, prints HPA status while it runs, and always stops the workload.

Terminal window
POD=$(kubectl -n "$NAMESPACE" get pod -l "app=$APP" \
-o jsonpath='{.items[0].metadata.name}')
kubectl -n "$NAMESPACE" exec "$POD" -- sh -c '
yes > /dev/null & load_pid=$!
trap "kill $load_pid" EXIT
sleep 120
'

In another terminal while that command is running:

Terminal window
kubectl -n "$NAMESPACE" get hpa "${APP}-hpa" --watch
kubectl -n "$NAMESPACE" get pods -l "app=$APP" --watch

Do not treat a particular scale-up or scale-down duration as a product guarantee. Kubernetes evaluates metrics periodically and applies its configured behavior; inspect the HPA conditions and events for the observed decision:

Terminal window
kubectl -n "$NAMESPACE" get events \
--field-selector involvedObject.kind=HorizontalPodAutoscaler,involvedObject.name="${APP}-hpa" \
--sort-by=.lastTimestamp

After the load ends, continue watching until the desired replica count returns to the configured minimum.

CPU-only scaling is a good first configuration. Add memory only when increased memory use is a reliable signal that more replicas will relieve load:

[hpa]
enabled = true
min_replicas = 2
max_replicas = 10
cpu_target = 70
memory_target = 80

Both metrics are resource-utilization percentages calculated against the container requests. Raising cpu_request or memory changes the denominator, so retest the thresholds after changing resource sizing.

HPA changes pod count. Vertical Pod Autoscaler (VPA) changes resource requests. Do not combine active HPA with VPA Auto: the CLI rejects that combination because both controllers would affect utilization behavior. HPA with VPA Off or Initial is supported, but validate it on your workload before using it in production.

A PodDisruptionBudget protects against voluntary disruption such as maintenance; it is not an autoscaler and it does not replace HPA capacity. Configure it separately when your minimum replica count needs availability protection.

Symptom Check Likely correction
TARGETS shows <unknown> kubectl -n "$NAMESPACE" describe hpa "${APP}-hpa" Restore metrics-server access and verify resource requests are present.
Deploy completes but the HPA is NotFound 1ctl app status "$APP", then kubectl -n "$NAMESPACE" get hpa "${APP}-hpa" Wait for reconciliation to finish; if it remains absent, retain the deployment ID and inspect the reconciliation error.
HPA does not scale above the minimum HPA conditions, events, and kubectl top pods Confirm utilization exceeds the configured target and max_replicas is high enough.
A manual scale command is refused 1ctl app scale output Change min_replicas or max_replicas, then deploy.
Scale behavior is surprising after changing requests HPA target and Deployment requests Re-evaluate CPU/memory thresholds because utilization is request-relative.

HPA configuration and disabling are reconciled through the deployment intent. Verify the rendered HPA and its conditions after enabling, and verify that the owned HPA is absent while the Deployment stays healthy after disabling. Do not delete ${APP}-hpa manually while it remains enabled in the deployment configuration.