Observability
Operator and developer reference for ReelBolt's telemetry (work package F9): what the three
services export, how to switch it on, what a span or metric may carry, and which dashboards to
build. The backend decision is Grafana Cloud's EU free tier over OTLP (decision D19,
design/cloud-provider-adr.md); nothing in the code is specific to
Grafana, so any OTLP/HTTP collector works. Secrets handling is in secrets.md.
Turning export on
Export is off by default. Each service exports traces and metrics over OTLP/HTTP only when
OTEL_EXPORTER_OTLP_ENDPOINT is set to a non-blank value (the per-signal
OTEL_EXPORTER_OTLP_TRACES_ENDPOINT and OTEL_EXPORTER_OTLP_METRICS_ENDPOINT also count). With
none set, no exporter, span processor or extra instrumentation is registered, the global providers
stay the no-op defaults, and startup is identical to a build without telemetry. This holds for the
Go API (api/telemetry), the Inference API and the engine (ReelBolt.Shared.Observability), and
each has a test pinning it.
The exporters read the standard variables themselves:
| Variable | Purpose |
|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT | The collector base URL. Blank disables everything. |
OTEL_EXPORTER_OTLP_PROTOCOL | http/protobuf (the cloud compose default, and what Grafana Cloud's OTLP gateway expects). |
OTEL_EXPORTER_OTLP_HEADERS | Auth, for Grafana Cloud Authorization=Basic <base64 instanceId:token>. A secret. |
OTEL_DEPLOYMENT_ENVIRONMENT | Becomes deployment.environment (falls back to ASPNETCORE_ENVIRONMENT). |
OTEL_SERVICE_NAME, OTEL_RESOURCE_ATTRIBUTES | Standard overrides, honoured by every service. |
The cloud compose layer (infra/compose/docker-compose.cloud.yml) passes all of these to go-api,
inference and workflow-engine with empty defaults; the values go in the per-host .env
(env.control.example, env.engine.example). To enable it, create a Grafana Cloud stack in an EU
region, create an access policy with metrics:write and traces:write, base64-encode
<instanceId>:<token>, and put the endpoint and header in both hosts' .env.
What is exported
Resource attributes on every signal: service.name (ReelBolt.Inference.Api,
ReelBolt.WorkflowEngine, ReelBolt.GoApi), service.version and deployment.environment.
Traces: ASP.NET Core and HttpClient spans, EF Core spans (statement text is not recorded, which
is the instrumentation's default), MassTransit publish and consume spans (the MassTransit
source, which carries traceparent through RabbitMQ headers so a request in the Inference API and
the execution it queues are one trace), the engine's own ReelBolt.WorkflowEngine source
(ExecuteWorkflow and the step spans), and otelhttp server spans in the Go API.
Metrics: ASP.NET Core and HttpClient runtime metrics, the MassTransit meter, the engine's
reelbolt.workflows.active, reelbolt.workflows.completed, reelbolt.step.duration_ms and the
decision-gate histograms, and InferenceTenancyDiagnostics.
What is not wired yet: nginx does not start or forward a trace context (it would need the nginx otel module), so a browser request's trace begins at the first service. The runner gateway and the cloud worker are not instrumented, so a trace currently ends at the engine. Both belong with the runner-fabric work.
Tenant tagging and cardinality
Every span can carry reelbolt.org_id, the organization the work ran for. Traces are the right
place for it: a trace is one request, so the high cardinality costs nothing. It is set from:
- the Inference API's
ITenantContext(the validated token's organization), stamped when the span ends because an ASP.NET Core span starts before the token is validated; - the engine's execution context (
WorkflowExecutionContext.OrganizationId), plus an explicit tag on theExecuteWorkflowspan; - the Go API's authentication middleware, after the organization membership check.
Metrics never carry an organization id or any other per-tenant value. The only tenant-derived
metric dimensions are plan_tier and compute_target, both normalized to a closed set before
use (ReelBoltTelemetry.NormalizePlanTier collapses unknown plan keys to other, and compute
target is cloud, runner or unknown). They are attached to reelbolt.workflows.completed and
reelbolt.step.duration_ms. Today compute_target is always cloud, because an execution row
does not yet record where its media work ran; the dimension becomes meaningful when the runner
fabric writes it. Per-organization numbers (spend, usage, refunds) belong in the usage ledger, not
in metrics, and Grafana reads them from Postgres.
Spans carry no user content: no prompts, model output, file names or request bodies are added by this work, in line with the data-processing note in the cloud ADR.
Dashboards to build
Build these in Grafana once data is flowing. The first four read OTLP metrics; the last two read Postgres through a read-only data source, because the usage ledger is the source of truth for money.
| Dashboard | Source | Panels |
|---|---|---|
| Queue depth and slots | RabbitMQ metrics from the broker host (or CloudAMQP) plus reelbolt.workflows.active | Messages ready on workflow-execution, active executions against WorkflowEngine:MaxConcurrency, executions started per minute. |
| Step latency | reelbolt.step.duration_ms by step.type and agent.type | p50 and p95 per step type, split by plan_tier; alert when VideoCompile or Agent p95 doubles against its 7-day median. |
| Outcomes | reelbolt.workflows.completed by status and plan_tier | Passed, failed and cancelled per hour, failure ratio, with the ExecutionRecoverySweeper interruptions called out. |
| API health | ASP.NET Core and otelhttp request metrics | Request rate, 5xx ratio and p95 per service. |
| LLM spend per hour | usage ledger (B2), SUM(cost_usd) grouped by hour and by provider | Spend per hour against the daily budget caps, top providers, cost per completed execution. |
| Refunds | usage ledger refund rows | Refund count and credits per hour, refunds as a share of executions, split by failure reason. |
Alerts worth setting at the start: no reelbolt.workflows.completed datapoint for 30 minutes
while the queue is non-empty, failure ratio above 20% for 15 minutes, and queue depth growing
while active executions sit at the cap (a capacity signal for the autoscaler).
Adding telemetry
Register new ActivitySource or Meter names in the service's own registration, and pass any
extra meter to AddReelBoltTelemetry. Tag spans freely, but a new metric dimension must come from
a closed set; if a value could be an id, a name or free text, it is a span attribute. The shared
helpers live in inference/src/ReelBolt.Shared/Observability/ReelBoltTelemetry.cs and
api/telemetry/telemetry.go.