Autoscaling Your App
Horizontal Pod Autoscaler (HPA) changes the number of running pods in response to average CPU and, optionally, memory utilization. Use it for stateless workloads whose demand changes over time. This guide is about capacity control; it does not change your rollout strategy. See the deployment guide for rolling-update settings.
Before you enable HPA
Section titled “Before you enable HPA”HPA needs a running app, resource requests, and cluster metrics. Satusky configures CPU and memory requests from satusky.toml or the matching 1ctl deploy flags. The cluster must expose the Kubernetes resource-metrics API:
kubectl get apiservice v1beta1.metrics.k8s.iokubectl top pods -n <your-organization-namespace>Available=True on the API service and pod metrics from kubectl top are the prerequisites for utilization-based scaling. If either check is unavailable, the HPA object can exist but cannot make informed scaling decisions.
1ctl deploy --wait persists the HPA intent and reconciles an owned autoscaling/v2 HorizontalPodAutoscaler. Still run the rendered-contract check below before relying on autoscaling: it confirms the final bounds, target metrics, and any Kubernetes condition that prevents scaling.
Configure HPA in satusky.toml
Section titled “Configure HPA in satusky.toml”Use the supported [hpa] section alongside an explicit CPU request. This example maintains one to three API pods, aiming for 60% average CPU utilization. Memory scaling is disabled when memory_target is 0.
[app] name = "my-api" port = 8080 cpu_request = "250m" cpu_limit = "1" memory = "256Mi"
[hpa] enabled = true min_replicas = 1 max_replicas = 3 cpu_target = 60 memory_target = 0Deploy or update the app with that file:
1ctl deploy --config satusky.toml --waitCLI flags are useful for a one-off update and take precedence over values in the file:
1ctl deploy --config satusky.toml \ --hpa \ --hpa-min-replicas 1 \ --hpa-max-replicas 3 \ --hpa-cpu-target 60 \ --hpa-memory-target 0 \ --waitThe public defaults when HPA is enabled are min 1, max 10, and CPU target 80. --hpa-memory-target 0 means no memory metric is added. HPA min replicas must not exceed max replicas.
Verify the rendered contract
Section titled “Verify the rendered contract”Set the app and namespace once, then inspect both the control-plane record and the Kubernetes object:
APP=my-apiNAMESPACE=<your-organization-namespace>
1ctl app status "$APP"
kubectl -n "$NAMESPACE" get hpa "${APP}-hpa" -o json | jq '{ minReplicas: .spec.minReplicas, maxReplicas: .spec.maxReplicas, metrics: .spec.metrics, currentReplicas: .status.currentReplicas, desiredReplicas: .status.desiredReplicas, conditions: .status.conditions}'When the object exists, the HPA targets Deployment/$APP. Its conditions are the authoritative explanation of whether it can scale. In particular, check for AbleToScale=True and ScalingActive=True; a metrics error or missing resource request will appear in the condition message.
kubectl -n "$NAMESPACE" describe hpa "${APP}-hpa"kubectl -n "$NAMESPACE" get deployment "$APP" -o jsonpath='{.spec.template.spec.containers[0].resources.requests}{"\n"}'Only after the HPA object exists does it own the Deployment replica count. Do not use 1ctl app scale for that app; update the HPA bounds through satusky.toml or the HPA flags and deploy the change.
Disable HPA safely
Section titled “Disable HPA safely”Set enabled = false in the same [hpa] section, keep an explicit app.replicas value, and deploy the update:
[app] replicas = 1
[hpa] enabled = false1ctl deploy --config satusky.toml --waitkubectl -n "$NAMESPACE" get hpa "${APP}-hpa"kubectl -n "$NAMESPACE" rollout status "deployment/$APP" --timeout=2mThe HPA command should return NotFound, while the Deployment remains present at the configured replica count. The platform removes only its owned HPA; do not delete the HPA manually while it is still enabled in the deployment configuration.
Observe an intentional scale test
Section titled “Observe an intentional scale test”Run a bounded CPU test only in a non-production deployment that you are authorized to disturb. This example generates CPU in one app pod for two minutes, prints HPA status while it runs, and always stops the workload.
POD=$(kubectl -n "$NAMESPACE" get pod -l "app=$APP" \ -o jsonpath='{.items[0].metadata.name}')
kubectl -n "$NAMESPACE" exec "$POD" -- sh -c ' yes > /dev/null & load_pid=$! trap "kill $load_pid" EXIT sleep 120'In another terminal while that command is running:
kubectl -n "$NAMESPACE" get hpa "${APP}-hpa" --watchkubectl -n "$NAMESPACE" get pods -l "app=$APP" --watchDo not treat a particular scale-up or scale-down duration as a product guarantee. Kubernetes evaluates metrics periodically and applies its configured behavior; inspect the HPA conditions and events for the observed decision:
kubectl -n "$NAMESPACE" get events \ --field-selector involvedObject.kind=HorizontalPodAutoscaler,involvedObject.name="${APP}-hpa" \ --sort-by=.lastTimestampAfter the load ends, continue watching until the desired replica count returns to the configured minimum.
CPU and memory metrics
Section titled “CPU and memory metrics”CPU-only scaling is a good first configuration. Add memory only when increased memory use is a reliable signal that more replicas will relieve load:
[hpa] enabled = true min_replicas = 2 max_replicas = 10 cpu_target = 70 memory_target = 80Both metrics are resource-utilization percentages calculated against the container requests. Raising cpu_request or memory changes the denominator, so retest the thresholds after changing resource sizing.
HPA, VPA, and availability controls
Section titled “HPA, VPA, and availability controls”HPA changes pod count. Vertical Pod Autoscaler (VPA) changes resource requests. Do not combine active HPA with VPA Auto: the CLI rejects that combination because both controllers would affect utilization behavior. HPA with VPA Off or Initial is supported, but validate it on your workload before using it in production.
A PodDisruptionBudget protects against voluntary disruption such as maintenance; it is not an autoscaler and it does not replace HPA capacity. Configure it separately when your minimum replica count needs availability protection.
Troubleshooting
Section titled “Troubleshooting”| Symptom | Check | Likely correction |
|---|---|---|
TARGETS shows <unknown> |
kubectl -n "$NAMESPACE" describe hpa "${APP}-hpa" |
Restore metrics-server access and verify resource requests are present. |
Deploy completes but the HPA is NotFound |
1ctl app status "$APP", then kubectl -n "$NAMESPACE" get hpa "${APP}-hpa" |
Wait for reconciliation to finish; if it remains absent, retain the deployment ID and inspect the reconciliation error. |
| HPA does not scale above the minimum | HPA conditions, events, and kubectl top pods |
Confirm utilization exceeds the configured target and max_replicas is high enough. |
| A manual scale command is refused | 1ctl app scale output |
Change min_replicas or max_replicas, then deploy. |
| Scale behavior is surprising after changing requests | HPA target and Deployment requests | Re-evaluate CPU/memory thresholds because utilization is request-relative. |
HPA configuration and disabling are reconciled through the deployment intent. Verify the rendered HPA and its conditions after enabling, and verify that the owned HPA is absent while the Deployment stays healthy after disabling. Do not delete ${APP}-hpa manually while it remains enabled in the deployment configuration.