Skip to content

Observability

Observability in Satusky spans three different things that are easy to confuse:

Layer Question
Workload Are pods built, scheduled, and ready?
Endpoint Can users reach the public hostname over valid HTTPS?
Machine Is the underlying capacity healthy and usable?

A complete operator view needs all three.

Source Current use
Kubernetes watch / informer data deployment status and pod readiness
Kubernetes metrics API CPU / memory observations
Prometheus application, network, and machine/network time-series
Hubble-derived metrics paths deployment network quality / latency inputs
Talos API low-level machine resources and state
SideroLink machine connectivity and discovery
WebSockets live deployment, machine, log, and notification streams
PostgreSQL historical metrics, billing records, persisted metadata
Audit log actor and resource history for control-plane operations

The backend and CLI expose several useful observability pieces:

  • deployment live-status WebSockets,
  • machine status WebSockets,
  • deployment metrics endpoints,
  • Prometheus-backed application and network metrics,
  • machine health jobs,
  • 1ctl machine inspect for inventory, hardware, labels, and Talos status,
  • 1ctl machine logs and machine events for bounded diagnostics,
  • 1ctl audit list and audit get for control-plane history.

The gap is less “no observability exists” and more “the user-facing model is not yet unified.”

status = concise current state
metrics = changing measurements over time
logs = emitted events / text streams
events = lifecycle and system transitions
check = cross-plane validation, especially for domains

That means:

  • app status should remain workload-focused,
  • domain checks should own public endpoint diagnostics,
  • machine commands should own fleet health and point-in-time diagnostics,
  • audit commands should explain who changed control-plane resources, not replace logs or metrics,
  • dashboards may compose them, but the architecture should not blur them.
Terminal window
1ctl machine inspect MACHINE_ID
1ctl machine logs MACHINE_ID --source kubernetes --tail 100 --since 10m
1ctl machine events MACHINE_ID --tail 50
1ctl audit list --limit 20
1ctl audit get AUDIT_ID

Machine diagnostics can return a successful empty state. Audit list and detail support JSON output for automation; filters include action and user ID. Audit access is permission-scoped and does not expose secret values.

A deployment can be:

  • pod-ready but publicly unreachable,
  • publicly routable but backed by a degraded node,
  • healthy in one cluster and unhealthy in another,
  • consuming resources normally while billing state is impaired.

The platform should preserve these dimensions instead of collapsing them into one opaque “healthy” bit.

Gap Target
Machine diagnostics are bounded queries rather than a unified live stream. First-class machine status and metrics workflows.
Public endpoint readiness is not yet reported with the same rigor as pod readiness. Route/DNS/TLS/HTTP checks become standard.
Metrics sources are rich but scattered. One documented observability model with clear source-of-truth boundaries.
Historical billing observations and live operational metrics can be conflated. Distinguish accounting records from live telemetry.