Skip to content

Troubleshooting

Use this guide after a deployment has been submitted. It is a diagnostic runbook, not another deployment tutorial: identify the failing layer first, then make the smallest recovery change.

Set the application name and derive its deployment ID and namespace from the CLI instead of guessing them:

APP=my-api
DEPLOYMENT_ID=$(1ctl -o json app get "$APP" | jq -r '.deployment_id')
NAMESPACE=$(1ctl -o json app get "$APP" | jq -r '.namespace')
1ctl app status "$APP"
kubectl -n "$NAMESPACE" get deployment,service
kubectl -n "$NAMESPACE" get pod -l "app=$APP"

The first command reports the platform record. The kubectl commands report the live workload and Service. When they disagree, use pod status and Kubernetes events to diagnose the workload, then capture both results for support.

Do not run a namespace-wide destructive command while troubleshooting. Every recovery command below takes the application name or a selected pod.

The failure occurs while 1ctl deploy is printing Docker build output, before an image is pushed or a new workload appears.

The build error is in the deploy output; application logs do not exist yet. Read the first failed Docker instruction and its preceding context. Common examples are:

  • COPY names a file outside the submitted build context.
  • A pinned Python package has no compatible distribution.
  • A system library required by a Python wheel is missing.
  • CMD names a module or executable that the image does not contain.

Run the deploy again only after changing the affected Dockerfile, requirements file, or build context:

1ctl deploy

Do not use app restart for a build failure. Restart reuses the current image and cannot include source changes.

kubectl -n "$NAMESPACE" get pod -l "app=$APP" -o wide
kubectl -n "$NAMESPACE" describe pod <pod-name>

Read the Events section of describe. A pod that is Pending may be waiting for a node, image pull, volume attach, or an init container. A pod in Init:n/m has started its setup sequence but has not reached the application container. These are different failures; do not change the application port until the event identifies a networking or probe problem.

For a normal Deployment, wait for the declared rollout:

kubectl -n "$NAMESPACE" rollout status "deployment/$APP" --timeout=2m

If the deployment has zero available replicas after the timeout, keep the pod event output and inspect the application logs before redeploying.

kubectl -n "$NAMESPACE" get pod -l "app=$APP"
kubectl -n "$NAMESPACE" describe pod <pod-name>
1ctl logs --app "$APP" --tail 100

The last log lines normally contain the actionable failure: a missing dependency, invalid application setting, permission error, failed database connection, or a process that exits immediately. If the container has restarted, inspect the prior attempt too:

kubectl -n "$NAMESPACE" logs <pod-name> --previous

Treat logs as potentially sensitive. Do not paste API keys, connection strings, tokens, or full configuration into tickets or chat.

  • Fix code, a Dockerfile, or dependencies, then run 1ctl deploy.
  • If only a non-secret runtime setting changed, update it with 1ctl config create and wait for application status.
  • If a required credential changed, update it with 1ctl secret create and wait for application status.

Use the application-specific flags rather than relying on the current directory:

1ctl config create --app "$APP" --env LOG_LEVEL=info
1ctl secret create --app "$APP" --kv API_KEY="$API_KEY"
1ctl app status "$APP" --watch

Never put the literal secret value in source control, a shell history shared with others, or a support request.

To confirm that a configuration or secret record exists without printing a secret value:

1ctl config list --app "$APP"
1ctl secret list --app "$APP"

Compare all four port declarations: the application record, the Deployment container port, the Service targetPort, and the HTTPRoute backend port. They must match the port on which the process listens. A workload can report Running while its public URL returns 503 when a later image update changes the container port but leaves a stale Service or route port. Keep the deployment and report the deployment ID plus these four values rather than repeatedly redeploying.

The Service can exist even when no ready pod matches its selector. Check all three layers:

kubectl -n "$NAMESPACE" get service "$APP"
kubectl -n "$NAMESPACE" get endpointslice -l "kubernetes.io/service-name=$APP"
kubectl -n "$NAMESPACE" get pod -l "app=$APP"

An EndpointSlice with no ready addresses means the Service cannot send traffic anywhere. Fix pod readiness or the Service selector; do not add a second Service with the same public intent.

For a short-lived, private diagnostic from a cluster operator workstation:

kubectl -n "$NAMESPACE" port-forward "service/$APP" 18000:<app-port>
# In another terminal:
curl -fsS http://127.0.0.1:18000/health

Stop port-forwarding with Ctrl+C. A successful port-forward proves the Service and application path only; it does not prove DNS, TLS, or Gateway routing.

SatuSky public routing uses the Gateway API. Verify the platform-owned HTTPRoute; do not add a competing Ingress while diagnosing:

kubectl -n "$NAMESPACE" get httproute "$APP-route"

Then run the focused doctor command. Passing a deployment ID is safer than a namespace-wide doctor invocation:

1ctl doctor --deployment-id "$DEPLOYMENT_ID" --health-path /health

Doctor reports route attachment, DNS propagation, TLS ownership, reachability, and an optional smoke result. It exits non-zero when the requested smoke path does not return a successful HTTP status.

Once Doctor reports a route and DNS are ready, test the public endpoint:

APP_URL=$(1ctl -o json app get "$APP" | jq -r '.domain')
curl -fsS "$APP_URL/health"

Interpret the layers in order:

Result Likely layer Next check
HTTPRoute is absent Route reconciliation 1ctl app status and Doctor output
DNS lookup fails DNS condition is not verified 1ctl app status, then wait before changing app code
DNS works but HTTP fails Gateway, Service, or pod EndpointSlice, pod readiness, and logs
Port-forward works but public URL fails Gateway/DNS/TLS HTTPRoute and Doctor output

Capture useful logs without waiting forever

Section titled “Capture useful logs without waiting forever”

For a bounded snapshot:

1ctl logs --app "$APP" --tail 100

For live logs while reproducing one request:

1ctl logs stream --deployment-id "$DEPLOYMENT_ID"

Stop the stream with Ctrl+C after the request completes. If the CLI cannot resolve the app name, use the deployment ID shown by 1ctl -o json app get instead of guessing a namespace.

When the application is no longer needed, delete only that application:

1ctl app delete "$APP" --yes

Confirm its platform record is gone:

1ctl app get "$APP"

The expected result is a non-zero response confirming that the application is absent. If Kubernetes resources remain after the platform deletion completes, record their exact names, labels, and the deployment ID before escalating. Do not delete shared namespace resources by label or name pattern.

Provide the smallest useful evidence set:

  • The deployment ID and application name.
  • The failing command and its exact non-secret error.
  • Deployment and pod status, including the relevant Events section.
  • A bounded log excerpt with secrets redacted.
  • Doctor output for the deployment ID.
  • Whether the EndpointSlice had ready addresses and whether the HTTPRoute existed.

This separates build, scheduling, workload, Service, and Gateway failures so the correct owner can act without reproducing the entire deployment.