Skip to main content

Observability

Operator and developer reference for ReelBolt's telemetry (work package F9): what the three services export, how to switch it on, what a span or metric may carry, and which dashboards to build. The backend decision is Grafana Cloud's EU free tier over OTLP (decision D19, design/cloud-provider-adr.md); nothing in the code is specific to Grafana, so any OTLP/HTTP collector works. Secrets handling is in secrets.md.

Turning export on​

Export is off by default. Each service exports traces and metrics over OTLP/HTTP only when OTEL_EXPORTER_OTLP_ENDPOINT is set to a non-blank value (the per-signal OTEL_EXPORTER_OTLP_TRACES_ENDPOINT and OTEL_EXPORTER_OTLP_METRICS_ENDPOINT also count). With none set, no exporter, span processor or extra instrumentation is registered, the global providers stay the no-op defaults, and startup is identical to a build without telemetry. This holds for the Go API (api/telemetry), the Inference API and the engine (ReelBolt.Shared.Observability), and each has a test pinning it.

The exporters read the standard variables themselves:

VariablePurpose
OTEL_EXPORTER_OTLP_ENDPOINTThe collector base URL. Blank disables everything.
OTEL_EXPORTER_OTLP_PROTOCOLhttp/protobuf (the cloud compose default, and what Grafana Cloud's OTLP gateway expects).
OTEL_EXPORTER_OTLP_HEADERSAuth, for Grafana Cloud Authorization=Basic <base64 instanceId:token>. A secret.
OTEL_DEPLOYMENT_ENVIRONMENTBecomes deployment.environment (falls back to ASPNETCORE_ENVIRONMENT).
OTEL_SERVICE_NAME, OTEL_RESOURCE_ATTRIBUTESStandard overrides, honoured by every service.

The cloud compose layer (infra/compose/docker-compose.cloud.yml) passes all of these to go-api, inference and workflow-engine with empty defaults; the values go in the per-host .env (env.control.example, env.engine.example). To enable it, create a Grafana Cloud stack in an EU region, create an access policy with metrics:write and traces:write, base64-encode <instanceId>:<token>, and put the endpoint and header in both hosts' .env.

What is exported​

Resource attributes on every signal: service.name (ReelBolt.Inference.Api, ReelBolt.WorkflowEngine, ReelBolt.GoApi), service.version and deployment.environment.

Traces: ASP.NET Core and HttpClient spans, EF Core spans (statement text is not recorded, which is the instrumentation's default), MassTransit publish and consume spans (the MassTransit source, which carries traceparent through RabbitMQ headers so a request in the Inference API and the execution it queues are one trace), the engine's own ReelBolt.WorkflowEngine source (ExecuteWorkflow and the step spans), and otelhttp server spans in the Go API.

Metrics: ASP.NET Core and HttpClient runtime metrics, the MassTransit meter, the engine's reelbolt.workflows.active, reelbolt.workflows.completed, reelbolt.step.duration_ms and the decision-gate histograms, and InferenceTenancyDiagnostics.

What is not wired yet: nginx does not start or forward a trace context (it would need the nginx otel module), so a browser request's trace begins at the first service. The runner gateway and the cloud worker are not instrumented, so a trace currently ends at the engine. Both belong with the runner-fabric work.

Tenant tagging and cardinality​

Every span can carry reelbolt.org_id, the organization the work ran for. Traces are the right place for it: a trace is one request, so the high cardinality costs nothing. It is set from:

  • the Inference API's ITenantContext (the validated token's organization), stamped when the span ends because an ASP.NET Core span starts before the token is validated;
  • the engine's execution context (WorkflowExecutionContext.OrganizationId), plus an explicit tag on the ExecuteWorkflow span;
  • the Go API's authentication middleware, after the organization membership check.

Metrics never carry an organization id or any other per-tenant value. The only tenant-derived metric dimensions are plan_tier and compute_target, both normalized to a closed set before use (ReelBoltTelemetry.NormalizePlanTier collapses unknown plan keys to other, and compute target is cloud, runner or unknown). They are attached to reelbolt.workflows.completed and reelbolt.step.duration_ms. Today compute_target is always cloud, because an execution row does not yet record where its media work ran; the dimension becomes meaningful when the runner fabric writes it. Per-organization numbers (spend, usage, refunds) belong in the usage ledger, not in metrics, and Grafana reads them from Postgres.

Spans carry no user content: no prompts, model output, file names or request bodies are added by this work, in line with the data-processing note in the cloud ADR.

Dashboards to build​

Build these in Grafana once data is flowing. The first four read OTLP metrics; the last two read Postgres through a read-only data source, because the usage ledger is the source of truth for money.

DashboardSourcePanels
Queue depth and slotsRabbitMQ metrics from the broker host (or CloudAMQP) plus reelbolt.workflows.activeMessages ready on workflow-execution, active executions against WorkflowEngine:MaxConcurrency, executions started per minute.
Step latencyreelbolt.step.duration_ms by step.type and agent.typep50 and p95 per step type, split by plan_tier; alert when VideoCompile or Agent p95 doubles against its 7-day median.
Outcomesreelbolt.workflows.completed by status and plan_tierPassed, failed and cancelled per hour, failure ratio, with the ExecutionRecoverySweeper interruptions called out.
API healthASP.NET Core and otelhttp request metricsRequest rate, 5xx ratio and p95 per service.
LLM spend per hourusage ledger (B2), SUM(cost_usd) grouped by hour and by providerSpend per hour against the daily budget caps, top providers, cost per completed execution.
Refundsusage ledger refund rowsRefund count and credits per hour, refunds as a share of executions, split by failure reason.

Alerts worth setting at the start: no reelbolt.workflows.completed datapoint for 30 minutes while the queue is non-empty, failure ratio above 20% for 15 minutes, and queue depth growing while active executions sit at the cap (a capacity signal for the autoscaler).

Adding telemetry​

Register new ActivitySource or Meter names in the service's own registration, and pass any extra meter to AddReelBoltTelemetry. Tag spans freely, but a new metric dimension must come from a closed set; if a value could be an id, a name or free text, it is a span attribute. The shared helpers live in inference/src/ReelBolt.Shared/Observability/ReelBoltTelemetry.cs and api/telemetry/telemetry.go.