Skip to main content

Observability

What you will learn​

  • How to use the health endpoints for liveness and readiness probes
  • How to scrape Prometheus metrics from a running worker
  • How to emit CloudEvents during workflow execution
  • How to control log verbosity
  • How correlationId is added to logs automatically

Health checks​

Zigflow exposes two dedicated health endpoints while the worker is running:

EndpointPurpose
GET /livezLiveness: returns 200 OK when the process is running
GET /readyzReadiness: returns 200 OK when the worker is connected and polling
GET /healthBackwards-compatible alias for /readyz

Change the listen address with --health-listen-address (default 0.0.0.0:3000).

Use /livez for liveness probes and /readyz for readiness probes. Use /health if you have an existing integration that cannot be updated.

Kubernetes​

The Helm chart configures liveness and readiness probes automatically using /livez and /readyz. No additional configuration is needed for standard deployments.

Docker Compose​

healthcheck:
test: ["CMD-SHELL", "curl -f http://localhost:3000/readyz || exit 1"]
interval: 10s
timeout: 5s
retries: 3

Prometheus metrics​

Zigflow exposes Prometheus metrics at:

http://localhost:9090/metrics

Change the address with --metrics-listen-address (default 0.0.0.0:9090).

Use --metrics-prefix to add a prefix to all metric names.

CloudEvents metrics​

MetricLabelsDescription
zigflow_events_emitted_totalclient, typeTotal events emitted per client and event type
zigflow_events_undelivered_totalclient, typeEvents that failed to deliver
zigflow_events_emit_duration_secondsclient, typeEmit duration histogram

These metrics are only populated when a CloudEvents configuration is active.

Scraping in Kubernetes​

When using the Helm chart, add a Prometheus scrape annotation to the pod:

values.yaml
podAnnotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9090"
prometheus.io/path: "/metrics"

CloudEvents​

tip

For full configuration options, event structure, file output examples and debugging guidance, see Debugging workflows.

Zigflow can emit CloudEvents v1.0 at key points during workflow execution. This is the primary mechanism for real-time observability into running workflows.

Enable it by providing a configuration file:

zigflow run -f workflow.yaml --cloudevents-config ./cloudevents.yaml

Configuration file​

cloudevents.yaml
clients:
- name: file-logger
protocol: file
target: ./tmp/events

- name: http-sink
protocol: http
target: "{{ .env.HTTP_EVENTS_URL }}"
options:
timeout: 5s
method: POST

The target field supports Go template syntax with access to environment variables via .env.

Supported protocols​

ProtocolTarget formatNotes
fileDirectory pathEvents written as YAML, one file per workflow execution
httpHTTP URLEvents sent as POST requests

Event types​

Event typeEmitted when
dev.zigflow.workflow.startedA workflow execution begins
dev.zigflow.workflow.completedA workflow execution completes successfully
dev.zigflow.task.startedA task begins
dev.zigflow.task.retriedA task is retried after failure
dev.zigflow.task.cancelledA task is cancelled
dev.zigflow.task.faultedA task fails
dev.zigflow.task.completedA task completes successfully
dev.zigflow.iteration.completedA task iteration completes

Important notes​

  • CloudEvents are emitted on a best-effort basis.
  • A failed event delivery does not fail the workflow.
  • Event emission must remain within Temporal's determinism constraints.
  • For high-throughput workflows, set appropriate HTTP timeouts to avoid workflow delays.

Logging​

Control log verbosity with --log-level:

zigflow run -f workflow.yaml --log-level debug

Valid values: trace, debug, info, warn, error. The code default is info. If the LOG_LEVEL environment variable is set, that value takes precedence.

Logs are structured JSON, written to stderr.

Correlation IDs in logs​

tip

For how propagated values reach the workflow in the first place, see Context Propagation.

Zigflow gives one propagated value special treatment. If the caller propagates a key named correlationId and its value is a string, Zigflow adds it as a structured correlationId field to every workflow and activity log entry.

You do not need to add it to individual log statements.

Given a caller that propagates:

{
"correlationId": "abc-123",
"tenantId": "acme"
}

Log entries from that workflow execution carry correlationId=abc-123 automatically. The tenantId value is still available to the workflow as ${ $propagated.tenantId }, but it is not added to log entries. Read it like any other value:

- recordTenant:
set:
tenantId: ${ $propagated.tenantId }

Rules​

ConditionResult
correlationId is present and is a stringAdded to every log entry as correlationId
correlationId is absentNo correlation field is added
correlationId is present but is not a stringNo correlation field is added

correlationId is optional. A workflow runs normally without it.

Zigflow does not convert a non-string correlationId to a string. If you want the automatic field, propagate a string. If you propagate a number, the value remains readable as ${ $propagated.correlationId } but is not logged automatically.

Why only correlationId​

Only correlationId is auto-logged because it is the one value whose purpose is to be correlated across log lines. Logging arbitrary propagated data by default would risk writing sensitive values into logs and adding high cardinality fields to every entry.

All propagated values remain available through $propagated whether or not they are logged. To surface another one, read it into the workflow with set or export, where it becomes visible in Temporal history and in CloudEvents.