Skip to content

ML Model API

This guide deploys a small, deterministic CPU-only model API from source. It uses a fixed two-feature linear model so the complete path can be tested without downloading model weights, using a paid inference API, or supplying a token at build or runtime.

The example is intentionally small, but the operational constraints are the same for a real packaged model: keep the model artifact in the image, bound request size and per-pod concurrency, give loading time its own startup probe, and size resources from measurements rather than from the model file’s size alone.

  • 1ctl installed and authenticated.
  • A unique, lowercase application name. Names are shared in an organization namespace.
  • kubectl access is optional, but lets you inspect the applied probes and resources.

The commands use ml-model-api-demo. Replace it with a unique name before deploying.

Create the project:

Terminal window
mkdir ml-model-api-demo
cd ml-model-api-demo

Create go.mod:

module ml-model-api-demo
go 1.23

Create main.go:

package main
import (
"encoding/json"
"errors"
"log"
"net/http"
"time"
)
const (
maxRequestBytes = 1024
maxConcurrentPredictions = 4
)
var predictionSlots = make(chan struct{}, maxConcurrentPredictions)
type predictRequest struct {
Features []float64 `json:"features"`
}
type predictResponse struct {
Class string `json:"class"`
Score float64 `json:"score"`
}
// score is a fixed, two-feature linear model. Treat these weights as the
// model artifact for this example: they are versioned with the service and
// need no runtime download.
func score(features []float64) (float64, error) {
if len(features) != 2 {
return 0, errors.New("features must contain exactly two numbers")
}
return 0.2 + 0.6*features[0] - 0.4*features[1], nil
}
func predict(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodPost {
w.Header().Set("Allow", http.MethodPost)
http.Error(w, "method not allowed", http.StatusMethodNotAllowed)
return
}
// Four in-flight predictions per replica is a deliberate CPU and memory
// bound. Return overload to the caller instead of building an unbounded queue.
select {
case predictionSlots <- struct{}{}:
defer func() { <-predictionSlots }()
default:
http.Error(w, "inference queue is full", http.StatusTooManyRequests)
return
}
r.Body = http.MaxBytesReader(w, r.Body, maxRequestBytes)
defer r.Body.Close()
var request predictRequest
decoder := json.NewDecoder(r.Body)
decoder.DisallowUnknownFields()
if err := decoder.Decode(&request); err != nil {
var maxBytesErr *http.MaxBytesError
if errors.As(err, &maxBytesErr) {
http.Error(w, "request body exceeds 1024 bytes", http.StatusRequestEntityTooLarge)
return
}
http.Error(w, "invalid JSON request", http.StatusBadRequest)
return
}
value, err := score(request.Features)
if err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
class := "negative"
if value >= 0 {
class = "positive"
}
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(predictResponse{Class: class, Score: value})
}
func health(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(map[string]string{"status": "ok"})
}
func main() {
mux := http.NewServeMux()
mux.HandleFunc("/healthz", health)
mux.HandleFunc("/readyz", health)
mux.HandleFunc("/predict", predict)
server := &http.Server{
Addr: ":8080",
Handler: mux,
ReadHeaderTimeout: 2 * time.Second,
ReadTimeout: 5 * time.Second,
WriteTimeout: 5 * time.Second,
IdleTimeout: 30 * time.Second,
}
log.Printf("deterministic CPU model API listening on %s", server.Addr)
log.Fatal(server.ListenAndServe())
}

/predict accepts only a JSON body up to 1024 bytes and exactly two numeric features. Its response has only a class and a score, so the response is also bounded. Invalid feature counts return 400; an oversized valid JSON request returns 413; and a full in-process queue returns 429.

The four-slot channel is the per-replica concurrency limit. A real model should use a limit based on its measured peak memory and CPU time: with two replicas, this configuration permits at most eight simultaneous inferences. Do not use an unbounded in-memory queue for slow CPU inference. Test the overload path locally before release; it is deliberately independent of the platform routing check.

Create Dockerfile:

FROM golang:1.23-alpine AS build
WORKDIR /src
COPY go.mod main.go ./
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags='-s -w' -o /model-api .
FROM alpine:3.20
RUN adduser -D -H -u 10001 app
USER app
COPY --from=build /model-api /model-api
EXPOSE 8080
ENTRYPOINT ["/model-api"]

This is a multi-stage image: the Go toolchain stays out of the runtime image, and the service runs as an unprivileged user. SatuSky’s source builder publishes both linux/amd64 and linux/arm64 variants, including the ARM64 variant used by an ARM worker.

For a production model, include the immutable artifact in the build context and copy it into the final image, for example:

COPY model/model-v2026-07-27.json /app/model/model.json

Record the artifact version and checksum in your release process. Do not fetch weights at container startup: readiness should reflect a locally available, validated artifact, not an external download that can fail or change. Keep raw data and experiments out of the build context, but do not accidentally ignore the model file you intend to ship:

.git
.env
data/
notebooks/
*.ipynb
*.ckpt

Create satusky.toml:

[app]
name = "ml-model-api-demo"
port = 8080
cpu_request = "100m"
cpu_limit = "250m"
memory = "256Mi"
replicas = 1
[build]
dockerfile = "Dockerfile"
[checks]
health_path = "/healthz"
[checks.startup]
initial_delay_seconds = 0
timeout_seconds = 2
period_seconds = 3
failure_threshold = 10
[checks.startup.http_get]
path = "/healthz"
port = 8080
[checks.readiness]
timeout_seconds = 2
period_seconds = 3
failure_threshold = 3
[checks.readiness.http_get]
path = "/readyz"
port = 8080
[checks.liveness]
timeout_seconds = 2
period_seconds = 5
failure_threshold = 3
[checks.liveness.http_get]
path = "/healthz"
port = 8080

The lightweight example was verified with 100m requested CPU, a 250m CPU limit, and a 256Mi memory limit. Start a real model from its measured peak resident memory with its planned batch size and four in-flight requests; model-file size alone is not a memory estimate. Increase the startup probe budget if artifact deserialization or model warm-up needs longer than 30 seconds (10 × 3s).

Use /readyz only after the artifact has loaded and passed its validation. Keep /healthz inexpensive: it is for process liveness, not a test inference.

4. Deploy and confirm multi-architecture output

Section titled “4. Deploy and confirm multi-architecture output”

Run the deployment from the directory containing satusky.toml:

Terminal window
1ctl deploy --wait

The source build output should include both platforms:

[build] Building for platforms: linux/amd64,linux/arm64
...
Image architecture: linux/amd64,linux/arm64

Save the namespace and deployment ID only through the current application record:

Terminal window
APP=ml-model-api-demo
APP_JSON="$(1ctl -o json app get "$APP")"
NAMESPACE="$(printf '%s' "$APP_JSON" | jq -r '.namespace')"
DEPLOYMENT_ID="$(printf '%s' "$APP_JSON" | jq -r '.deployment_id')"
APP_URL="$(printf '%s' "$APP_JSON" | jq -r '.domain')"
printf '%s\n' "$APP_URL"

Confirm readiness, the generated Gateway API route, and the effective container settings:

Terminal window
1ctl app status "$APP"
kubectl -n "$NAMESPACE" rollout status "deployment/$APP" --timeout=2m
kubectl -n "$NAMESPACE" get httproute "$APP-route"
kubectl -n "$NAMESPACE" get deployment "$APP" -o json | jq '{
image: .spec.template.spec.containers[0].image,
resources: .spec.template.spec.containers[0].resources,
startup: .spec.template.spec.containers[0].startupProbe.httpGet.path,
readiness: .spec.template.spec.containers[0].readinessProbe.httpGet.path,
liveness: .spec.template.spec.containers[0].livenessProbe.httpGet.path,
nodeSelector: .spec.template.spec.nodeSelector
}'

For a multi-architecture image, an empty nodeSelector is expected: Kubernetes selects the matching AMD64 or ARM64 variant on the chosen worker. The probe paths should be /healthz, /readyz, and /healthz, and the resources should show the configured CPU limits and memory limit.

Wait for 1ctl app status to report DNS condition: verified for the generated hostname, then retry the public request rather than creating an Ingress or HTTPRoute yourself. A hostname reservation or a non-verified DNS condition does not prove the public target is correct:

Terminal window
curl --retry 24 --retry-delay 5 --retry-all-errors \
--fail --silent --show-error \
-X POST "$APP_URL/predict" \
-H 'content-type: application/json' \
--data '{"features":[1,0]}'

Expected response:

{"class":"positive","score":0.8}

Exercise the input boundary too:

Terminal window
curl --silent --show-error -o /dev/null -w '%{http_code}\n' \
-X POST "$APP_URL/predict" \
-H 'content-type: application/json' \
--data '{"features":[1]}'

This returns 400. A valid JSON request larger than 1024 bytes returns 413. Treat 429 as back-pressure: retry with exponential backoff or send work to a durable queue outside the request path.

Inspect startup and release logs by the immutable deployment ID:

Terminal window
1ctl logs --deployment-id "$DEPLOYMENT_ID" --tail 50

You should see deterministic CPU model API listening on :8080. For a real model, log the artifact version and successful load time, never the model input or secrets.

When finished, delete the application created by this guide:

Terminal window
1ctl app delete "$APP" --yes
kubectl -n "$NAMESPACE" get deployment,service,httproute | grep "$APP" || true

The final command should return no resources for this example. It does not affect other applications in the namespace.

The example’s deterministic score, invalid-feature response, oversized-body response, and four-slot overload path are suitable local unit-test cases. The concurrency gate is application code, so verify its overload behavior in your own build before releasing.

The live deployment verification covers:

  • A CPU-only deterministic model without an external model host, token, or paid API.
  • Cloud-built linux/amd64 and linux/arm64 image variants.
  • Separate startup, readiness, and liveness probes with the configured CPU and memory limits.
  • A platform-owned Gateway API route serving the bounded prediction response over public HTTPS.