ML Model API
This guide deploys a small, deterministic CPU-only model API from source. It uses a fixed two-feature linear model so the complete path can be tested without downloading model weights, using a paid inference API, or supplying a token at build or runtime.
The example is intentionally small, but the operational constraints are the same for a real packaged model: keep the model artifact in the image, bound request size and per-pod concurrency, give loading time its own startup probe, and size resources from measurements rather than from the model file’s size alone.
Prerequisites
Section titled “Prerequisites”1ctlinstalled and authenticated.- A unique, lowercase application name. Names are shared in an organization namespace.
kubectlaccess is optional, but lets you inspect the applied probes and resources.
The commands use ml-model-api-demo. Replace it with a unique name before deploying.
1. Create a self-contained model service
Section titled “1. Create a self-contained model service”Create the project:
mkdir ml-model-api-democd ml-model-api-demoCreate go.mod:
module ml-model-api-demo
go 1.23Create main.go:
package main
import ( "encoding/json" "errors" "log" "net/http" "time")
const ( maxRequestBytes = 1024 maxConcurrentPredictions = 4)
var predictionSlots = make(chan struct{}, maxConcurrentPredictions)
type predictRequest struct { Features []float64 `json:"features"`}
type predictResponse struct { Class string `json:"class"` Score float64 `json:"score"`}
// score is a fixed, two-feature linear model. Treat these weights as the// model artifact for this example: they are versioned with the service and// need no runtime download.func score(features []float64) (float64, error) { if len(features) != 2 { return 0, errors.New("features must contain exactly two numbers") } return 0.2 + 0.6*features[0] - 0.4*features[1], nil}
func predict(w http.ResponseWriter, r *http.Request) { if r.Method != http.MethodPost { w.Header().Set("Allow", http.MethodPost) http.Error(w, "method not allowed", http.StatusMethodNotAllowed) return }
// Four in-flight predictions per replica is a deliberate CPU and memory // bound. Return overload to the caller instead of building an unbounded queue. select { case predictionSlots <- struct{}{}: defer func() { <-predictionSlots }() default: http.Error(w, "inference queue is full", http.StatusTooManyRequests) return }
r.Body = http.MaxBytesReader(w, r.Body, maxRequestBytes) defer r.Body.Close()
var request predictRequest decoder := json.NewDecoder(r.Body) decoder.DisallowUnknownFields() if err := decoder.Decode(&request); err != nil { var maxBytesErr *http.MaxBytesError if errors.As(err, &maxBytesErr) { http.Error(w, "request body exceeds 1024 bytes", http.StatusRequestEntityTooLarge) return } http.Error(w, "invalid JSON request", http.StatusBadRequest) return }
value, err := score(request.Features) if err != nil { http.Error(w, err.Error(), http.StatusBadRequest) return }
class := "negative" if value >= 0 { class = "positive" } w.Header().Set("Content-Type", "application/json") _ = json.NewEncoder(w).Encode(predictResponse{Class: class, Score: value})}
func health(w http.ResponseWriter, _ *http.Request) { w.Header().Set("Content-Type", "application/json") _ = json.NewEncoder(w).Encode(map[string]string{"status": "ok"})}
func main() { mux := http.NewServeMux() mux.HandleFunc("/healthz", health) mux.HandleFunc("/readyz", health) mux.HandleFunc("/predict", predict)
server := &http.Server{ Addr: ":8080", Handler: mux, ReadHeaderTimeout: 2 * time.Second, ReadTimeout: 5 * time.Second, WriteTimeout: 5 * time.Second, IdleTimeout: 30 * time.Second, } log.Printf("deterministic CPU model API listening on %s", server.Addr) log.Fatal(server.ListenAndServe())}/predict accepts only a JSON body up to 1024 bytes and exactly two numeric features. Its response has only a class and a score, so the response is also bounded. Invalid feature counts return 400; an oversized valid JSON request returns 413; and a full in-process queue returns 429.
The four-slot channel is the per-replica concurrency limit. A real model should use a limit based on its measured peak memory and CPU time: with two replicas, this configuration permits at most eight simultaneous inferences. Do not use an unbounded in-memory queue for slow CPU inference. Test the overload path locally before release; it is deliberately independent of the platform routing check.
2. Package the model image
Section titled “2. Package the model image”Create Dockerfile:
FROM golang:1.23-alpine AS buildWORKDIR /srcCOPY go.mod main.go ./RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags='-s -w' -o /model-api .
FROM alpine:3.20RUN adduser -D -H -u 10001 appUSER appCOPY --from=build /model-api /model-apiEXPOSE 8080ENTRYPOINT ["/model-api"]This is a multi-stage image: the Go toolchain stays out of the runtime image, and the service runs as an unprivileged user. SatuSky’s source builder publishes both linux/amd64 and linux/arm64 variants, including the ARM64 variant used by an ARM worker.
For a production model, include the immutable artifact in the build context and copy it into the final image, for example:
COPY model/model-v2026-07-27.json /app/model/model.jsonRecord the artifact version and checksum in your release process. Do not fetch weights at container startup: readiness should reflect a locally available, validated artifact, not an external download that can fail or change. Keep raw data and experiments out of the build context, but do not accidentally ignore the model file you intend to ship:
.git.envdata/notebooks/*.ipynb*.ckpt3. Configure probes and resources
Section titled “3. Configure probes and resources”Create satusky.toml:
[app] name = "ml-model-api-demo" port = 8080 cpu_request = "100m" cpu_limit = "250m" memory = "256Mi" replicas = 1
[build] dockerfile = "Dockerfile"
[checks] health_path = "/healthz"
[checks.startup] initial_delay_seconds = 0 timeout_seconds = 2 period_seconds = 3 failure_threshold = 10
[checks.startup.http_get] path = "/healthz" port = 8080
[checks.readiness] timeout_seconds = 2 period_seconds = 3 failure_threshold = 3
[checks.readiness.http_get] path = "/readyz" port = 8080
[checks.liveness] timeout_seconds = 2 period_seconds = 5 failure_threshold = 3
[checks.liveness.http_get] path = "/healthz" port = 8080The lightweight example was verified with 100m requested CPU, a 250m CPU limit, and a 256Mi memory limit. Start a real model from its measured peak resident memory with its planned batch size and four in-flight requests; model-file size alone is not a memory estimate. Increase the startup probe budget if artifact deserialization or model warm-up needs longer than 30 seconds (10 × 3s).
Use /readyz only after the artifact has loaded and passed its validation. Keep /healthz inexpensive: it is for process liveness, not a test inference.
4. Deploy and confirm multi-architecture output
Section titled “4. Deploy and confirm multi-architecture output”Run the deployment from the directory containing satusky.toml:
1ctl deploy --waitThe source build output should include both platforms:
[build] Building for platforms: linux/amd64,linux/arm64...Image architecture: linux/amd64,linux/arm64Save the namespace and deployment ID only through the current application record:
APP=ml-model-api-demoAPP_JSON="$(1ctl -o json app get "$APP")"NAMESPACE="$(printf '%s' "$APP_JSON" | jq -r '.namespace')"DEPLOYMENT_ID="$(printf '%s' "$APP_JSON" | jq -r '.deployment_id')"APP_URL="$(printf '%s' "$APP_JSON" | jq -r '.domain')"printf '%s\n' "$APP_URL"Confirm readiness, the generated Gateway API route, and the effective container settings:
1ctl app status "$APP"kubectl -n "$NAMESPACE" rollout status "deployment/$APP" --timeout=2mkubectl -n "$NAMESPACE" get httproute "$APP-route"kubectl -n "$NAMESPACE" get deployment "$APP" -o json | jq '{ image: .spec.template.spec.containers[0].image, resources: .spec.template.spec.containers[0].resources, startup: .spec.template.spec.containers[0].startupProbe.httpGet.path, readiness: .spec.template.spec.containers[0].readinessProbe.httpGet.path, liveness: .spec.template.spec.containers[0].livenessProbe.httpGet.path, nodeSelector: .spec.template.spec.nodeSelector}'For a multi-architecture image, an empty nodeSelector is expected: Kubernetes selects the matching AMD64 or ARM64 variant on the chosen worker. The probe paths should be /healthz, /readyz, and /healthz, and the resources should show the configured CPU limits and memory limit.
5. Call the public HTTPS API
Section titled “5. Call the public HTTPS API”Wait for 1ctl app status to report DNS condition: verified for the
generated hostname, then retry the public request rather than creating an
Ingress or HTTPRoute yourself. A hostname reservation or a non-verified DNS
condition does not prove the public target is correct:
curl --retry 24 --retry-delay 5 --retry-all-errors \ --fail --silent --show-error \ -X POST "$APP_URL/predict" \ -H 'content-type: application/json' \ --data '{"features":[1,0]}'Expected response:
{"class":"positive","score":0.8}Exercise the input boundary too:
curl --silent --show-error -o /dev/null -w '%{http_code}\n' \ -X POST "$APP_URL/predict" \ -H 'content-type: application/json' \ --data '{"features":[1]}'This returns 400. A valid JSON request larger than 1024 bytes returns 413. Treat 429 as back-pressure: retry with exponential backoff or send work to a durable queue outside the request path.
Inspect startup and release logs by the immutable deployment ID:
1ctl logs --deployment-id "$DEPLOYMENT_ID" --tail 50You should see deterministic CPU model API listening on :8080. For a real model, log the artifact version and successful load time, never the model input or secrets.
6. Clean up only this example
Section titled “6. Clean up only this example”When finished, delete the application created by this guide:
1ctl app delete "$APP" --yeskubectl -n "$NAMESPACE" get deployment,service,httproute | grep "$APP" || trueThe final command should return no resources for this example. It does not affect other applications in the namespace.
Validation scope
Section titled “Validation scope”The example’s deterministic score, invalid-feature response, oversized-body response, and four-slot overload path are suitable local unit-test cases. The concurrency gate is application code, so verify its overload behavior in your own build before releasing.
The live deployment verification covers:
- A CPU-only deterministic model without an external model host, token, or paid API.
- Cloud-built
linux/amd64andlinux/arm64image variants. - Separate startup, readiness, and liveness probes with the configured CPU and memory limits.
- A platform-owned Gateway API route serving the bounded prediction response over public HTTPS.
Next steps
Section titled “Next steps”- Environment Configuration for model-serving credentials and non-sensitive settings.
- Autoscaling for measured, traffic-driven replica scaling.
- CI/CD Integration for reproducible image releases.