Skip to main content

Monitoring Guide

Monitor Stem as a distributed, at-least-once system. Metrics and signals are telemetry, not exactly-once accounting or business truth.

Symptom → check → remedy​

SymptomCheckRemedy
Queue depth or enqueue-to-start latency risesWorker heartbeat, concurrency, routing, and broker latencyRestore workers, correct routing, or add capacity; do not blindly increase retries
Retry rate risestask-retry payloads, downstream errors, and retry policyRepair the dependency or bound/reduce retries
DLQ volume risesSample entries, error class, and payload/schema versionFix handler/deployment compatibility, then replay selected entries
Heartbeats stopProcess health, namespace, broker connectivity, and heartbeat intervalReplace the worker after checking for in-flight work
Scheduler drift or missed runsSchedule store, lock ownership, and broker publish errorsRepair the store/lock path and reconcile schedules

Metrics, logs, and signals​

Configure ObservabilityConfig and StemMetrics in application code. Supported environment names include STEM_METRIC_EXPORTERS, STEM_OTLP_ENDPOINT, STEM_HEARTBEAT_INTERVAL, STEM_WORKER_NAMESPACE, STEM_SIGNALS_ENABLED, and STEM_SIGNALS_DISABLED. Exporters are not automatic integrations.

Subscribe to StemSignals.taskRetry, taskSucceeded, and taskFailed, and include task ID, task name, queue, and run ID in application logs. Signals are in-process notifications and can be lost on crash; persist audit data separately.

CLI probes​

The optional stem_cli package exposes inspection commands through an adapter-backed application context:

stem observe metrics
stem observe queues
stem observe workers
stem worker stats

Run stem <command> --help for flags in the installed version. Protect mutating control commands and dashboards with normal operator authentication and TLS. See Observe & Operate and CLI control.