Skip to main content

Platform Assistant + MCP Server

ReelBolt supplies a conversational platform assistant, served by the Inference API and mounted over the whole authenticated dashboard, plus an independent Model Context Protocol (MCP) server. The assistant manages projects, files, workflows, agents and inference providers through a closed, server-authorised tool set. The MCP server is deliberately narrower: it stays the seven workflow tools it has always been, for external CLI clients.


Table of Contents​


Architecture overview​

┌──────────────────────────────────────────────────────────────────────┐
│ Inference API │
├──────────────────────────────────┬───────────────────────────────────┤
│ AssistantChatController │ AssistantThreadsController │
│ POST /api/v1/assistant/chat │ /api/v1/assistant/threads │
│ SSE stream of one reply │ caller-scoped thread CRUD │
│ ▲ │ ▲ │
│ │ runs │ │ persists │
│ ┌────────┴──────────────────────┴───────────────────────────────┐ │
│ │ AssistantAgent │ │
│ │ └── AssistantToolProvider (45 tools, composed per request) │ │
│ │ ├── 7 core provider tools │ │
│ │ ├── AssistantProjectTools (8) │ │
│ │ ├── AssistantFileTools (5) │ │
│ │ ├── AssistantWorkflowTools (6) │ │
│ │ ├── AssistantKnowledgeTools(3) │ │
│ │ ├── AssistantStyleTools (3) │ │
│ │ └── AssistantAdminTools (13, gated: 10 platform + 3 │ │
│ │ organization) │ │
│ └────────────────────────────────────────────────────────────────┘ │
│ ▲ │
├──────────────────────────┼────────────────────────────────────────────┤
│ Services │ │
│ WorkflowEditDiffService │ (ProposeWorkflow/ApplyWorkflow) │
│ IConfirmationTokenService (HMAC-SHA256, one key per deployment) │
│ IAssistantExecutionGuard (rolling-hour rate limit) │
│ InferenceApiDbContext (assistant_threads, assistant_messages) │
└──────────────────────────────────────────────────────────────────────┘
▼
PostgreSQL Database
(assistant_threads, assistant_messages,
Project, WorkflowDefinition, WorkflowStep,
WorkflowExecution, AgentDefinition,
InferenceProvider)

The Inference API owns the assistant end to end: the transport, the conversation store, the tool implementations and the check on who is allowed to call them. Nothing the model produces is an authorisation — every tool re-resolves the project, the workflow and the caller from the database on the call that is executing.

The MCP server is a separate Node.js process that talks to the workflow REST endpoints, not to the assistant. See The MCP server.


The platform assistant​

The assistant is one agent (AssistantAgent, an AgentType.Assistant built-in) reachable from every authenticated page. It answers questions about projects, files, workflows, agents and inference providers, and it changes them — but only through a propose-then-apply pair carrying a server-minted confirmation token, never on the strength of an argument the model supplied.

Its system prompt (AssistantAgent.DefaultPrompt, duplicated verbatim in DatabaseSeeder.BuiltInAgents and pinned by AssistantPromptConsistencyTests) names every registered tool, states how many there are, and frames everything a tool returns as untrusted DATA the model is reading, never as an instruction to follow — the same framing this codebase uses for media-derived content (see CLAUDE.md § Agent Types).


Chat transport​

POST /api/v1/assistant/chat
Authorization: Bearer <jwt>
Content-Type: application/json

{
"threadId": "<guid>",
"message": "delete the stale drafts workflow",
"projectId": "<guid>",
"workflowId": "<guid>"
}
FieldRequiredMeaning
threadIdyesThe server-persisted thread this turn belongs to.
messageyesThe user's new message. Prior turns come from the thread, not from the request.
projectIdnoThe project the caller currently has open.
workflowIdnoThe workflow the caller currently has open.

projectId is per-call context, not route state​

The assistant was previously mounted per project, with the project as a route parameter. That route is deleted. It could not express the case the shipped client actually has: a page with no project at all, where the assistant must still open and answer. A route parameter cannot say "no project", so context moved into the body.

  • A page with no project is a normal state. projectId is null, the request is accepted, and project-scoped tools answer with an actionable JSON error ({"error": "no project in context"}) that the model can act on rather than an exception.
  • A project the caller does not own is treated as ABSENT, not rejected. The controller checks p.OwnerId == userId and, if it fails, drops the context and proceeds. Rejecting would confirm to a stranger that a guessed project id is real; the thread keeps working either way.
  • workflowId is validated by the same rule, and additionally must belong to the project in context. It is injected into the first user turn as a [WORKFLOW CONTEXT] block, so it stays visible across the whole multi-turn exchange.

Responses​

Before any frame is written:

  • 404 with {"error": "Thread not found"} when the thread does not exist or belongs to another user — never 403, which would confirm it exists.
  • 400 with {"error": "Message must not be empty"} when message is blank.

The user's turn is persisted before the response stream starts, so a refused request records nothing and a dropped stream still records what the user said.

On success the response is text/event-stream (Cache-Control: no-cache, Connection: keep-alive) carrying the ui-message-stream protocol. Each frame is data: <json> followed by a blank line, and the frames arrive in this order:

FramePayloadWhen
start{"type":"start","messageId":"<guid>"}Always, first.
text-start{"type":"text-start","id":"<guid>"}Always, second.
text-delta{"type":"text-delta","textDelta":"<the reply>"}Always. The whole reply arrives in this one frame.
text-end{"type":"text-end"}Always. The assistant's turn is persisted before this frame, so a stream that dies afterwards still leaves a complete record.
data-workflowDiff{"type":"data-workflowDiff","data":{…WorkflowEditDiff…}}Only when a workflow was proposed during this turn.
data-toolStatus{"type":"data-toolStatus","data":{"caption":"<plain words>"}}Zero or more, between text-start and text-delta: one when each tool starts ("Searching the documentation") and one when it returns ("Thinking", or "Docs index still building — answering without it"), so the panel's label always says what is true now.
finish{"type":"finish","finishReason":"<reason>","usage":{"inputTokens":<n>,"outputTokens":<n>}}Always. <reason> is the provider's own finish reason mapped onto the AI SDK vocabulary (stop, length, content-filter, tool-calls, other, or stop when the provider reported none) — never a hardcoded stop.
[DONE]the literal data: [DONE]Always, last.

Payloads are serialized camelCase with nulls omitted. The frame set and its order are a published contract — the client parses by type.

Reply length, and why a truncated reply says so​

An assistant turn is bounded twice, from opposite ends, and both bounds announce themselves.

The model's own ceiling is set before the call. AssistantChatLimits.MaxOutputTokens resolves one number per turn and sends it as ChatOptions.MaxOutputTokens:

  • On a provider that draws its reasoning from the same budget as the visible answer — Anthropic (thinking.type=adaptive) and Gemini (native reasoning_effort mapping) — the ceiling is the configured answer allowance plus ReasoningHeadroomTokens (8,192). The configured value is a promise about the answer, not about the thinking that comes first, so headroom is added on top rather than taken out of it. Default answer allowance: 16,384, i.e. 24,576 on the wire.
  • On every other kind — and when the provider cannot be resolved — no ceiling is sent unless the deployment explicitly configures one. An OpenAI-compatible row may point at a self-hosted vLLM server whose context window is whatever its operator passed to --max-model-len; a ceiling ReelBolt invented there could turn a working turn into a 400.
  • Anthropic is the exception by necessity: the Messages API requires max_tokens on every request, so the only choice there is between the SDK's built-in number and ours.

When the provider answers with the length stop reason, the reply is marked (AssistantChatLimits.ModelTruncationMarker) before it is persisted or streamed — a half sentence streamed as a finished answer is the defect that marker exists to close. The marker names the ceiling ReelBolt actually sent, or omits the number when it sent none.

The character cap runs after the call and behaves as before: an assistant reply over Assistant:MaxMessageChars is truncated to the cap with its own marker, because it lands in assistant_messages and is replayed into every later prompt of that thread. When both bounds bite, the character marker is what the reader ends on — the bound that actually shortened the visible text is the one that gets to say so.


Conversation threads​

Conversations are persisted server-side. The client no longer sends history: it sends one new turn, and the server loads the thread's prior turns in sequence order.

MethodPathDescription
GET/api/v1/assistant/threadsThe caller's non-archived threads, most recently updated first.
POST/api/v1/assistant/threadsCreates a thread; body { projectId?, title? }, both optional. Returns the new id.
GET/api/v1/assistant/threads/{threadId:guid}One thread with its full, ordered history.
PATCH/api/v1/assistant/threads/{threadId:guid}Renames and/or archives; body { title?, archived? }. A null field is left unchanged, so a rename never implicitly un-archives.
DELETE/api/v1/assistant/threads/{threadId:guid}Deletes the thread and its messages.

The endpoints are caller-scoped, not route-scoped — the assistant is mounted globally, so there is no project in the URL to hang a conversation off. Every lookup filters on UserId inside the LINQ predicate, so the ownership filter is applied by the database before anything is materialized.

A thread id that exists but belongs to another user is 404, never 403. A 403 would confirm the thread exists to a user with no right to know that. The same rule applies to the chat transport above.

POST /api/v1/assistant/threads drops a project the caller does not own rather than rejecting it, so a page handing the assistant a stale or foreign project id cannot make conversation history unopenable.

Storage​

AddAssistantThreads (Migrations/20261002142456_AddAssistantThreads.cs) creates two tables:

  • assistant_threads — id, user_id (FK to application_users, cascade delete), nullable project_id, title, created_at, updated_at, nullable archived_at.
  • assistant_messages — id, thread_id (FK to assistant_threads, cascade delete), role (user/assistant/system/tool), content, sequence, nullable metadata_json (jsonb), created_at.

sequence is 1-based and unique per thread (unique index IX_assistant_messages_thread_id_sequence), assigned as max(sequence) + 1 when a turn is appended. The assistant's turn is written with userTurn.Sequence + 1, so the user and assistant halves of one exchange always take consecutive numbers. Appending a turn also touches the thread's updated_at, which is what orders the thread list by real activity.


Unanswered turns and turns that outlive the browser​

A thread whose newest message is the user's looks the same in the database whether its reply is still being written or its request died. Customers saw the second case as a question that silently disappeared after they navigated away. The server now tells the two apart and keeps answering when it can:

  • A turn outlives its browser. ChatStream runs the model on its own cancellation token, not the request's. When the client disconnects (panel closed, navigation, reload) the turn keeps running for up to Assistant:DetachedTurnSeconds (default 300) and still persists its reply; writes to the dead stream are skipped instead of thrown. Past that bound the turn is abandoned and logged.
  • In-flight turns are visible. AssistantTurnRegistry (process-wide) marks a thread while a turn on it runs. GET /api/v1/assistant/threads adds awaitingReply (newest message is the user's and nothing is answering it) and replyInProgress to each summary; GET …/threads/{id} adds replyInProgress. Both are appended with defaults, so older clients are unaffected.
  • "Ask again" never duplicates the question. When a turn arrives whose text equals a trailing, unanswered user message that is not in flight, the server reuses that message instead of storing the question a second time (and replaying it twice into every later prompt). A double-send while a turn is still running appends as before.
  • The panel says which. The History control is a labelled button; rows read "No reply yet" or "Still answering…"; opening the assistant shows a "has no reply yet — Open" hint for a waiting thread other than the current one; and a restored thread ending in an unanswered message shows either "Still working on your last message" (with Check again) or "No reply yet" with Ask again, which starts a run whose parent is that same message.

Background work never blocks an upload or the assistant​

Uploading a 233-file repository took ~15 minutes, and the first assistant question after it waited minutes too. Two shared resources were the cause, and both are now protected.

Uploads never wait for summarization. IBackgroundTaskQueue<T> has no awaiting enqueue any more: the upload, content-update and assistant file-write paths persist SummaryStatus.Pending and offer the file with the non-blocking TryQueue. A full queue leaves the row Pending. FileSummarizationService runs a sweeper shortly after startup and then every FileSummarization:SweepIntervalSeconds that re-offers Pending user files (oldest first, never video/audio, never engine-written agent/output files) and resets rows a previous process left in Processing for longer than StuckProcessingMinutes. The queue tracks every id from enqueue until the worker completes it, so a sweep can never queue or run the same file twice.

The person at the keyboard goes first. Summarization runs MaxConcurrency workers (default 1) on the chat provider, and each yields to an in-flight assistant turn for up to InteractiveYieldSeconds before its LLM call. Per-file indexing and the docs indexer yield likewise on the embedding server (see platform-docs-search.md). Yields are always bounded, so a long chat slows background work down but never stops it.

Key (FileSummarization:*)DefaultMeaning
MaxConcurrency1Summarization LLM calls at once (1..8).
QueueCapacity100In-memory queue size; a full queue never blocks anything.
SweepIntervalSeconds45How often Pending/stuck rows are re-offered.
InitialSweepDelaySeconds10First sweep after startup.
StuckProcessingMinutes15When a Processing row from an earlier process is re-queued.
InteractiveYieldSeconds60Longest wait for an assistant turn before one summarization call.

The tool surface​

The assistant has exactly 45 tools. AssistantToolProvider.GetTools composes them from seven lists, and AssistantPromptConsistencyTests asserts the prompt against the registration itself — both directions, and the count — so a tool that is registered but undescribed, or described but not registered, fails the suite.

GroupCountSource
Core provider tools7AssistantToolProvider
Project8AssistantProjectTools
File5AssistantFileTools
Workflow6AssistantWorkflowTools
Knowledge3AssistantKnowledgeTools
Styles3AssistantStyleTools
Admin (gated)13AssistantAdminTools
Total45

Every registered name, marked read-only or Propose/Apply:

Core provider tools (7)

ToolKind
ListAgentsread-only — agents available to this project (built-in plus this owner's custom)
ListStepSchemasread-only — JSON Schema for every step type's config
ListProjectFilesread-only — this project's files
ValidateWorkflowread-only — validates proposed steps, writes nothing
ProposeWorkflowPropose — diff + confirmationToken, writes nothing
ApplyWorkflowApply — the write half; see the token carve-out below
ExecuteWorkflowwrites — starts an execution, subject to the rolling-hour rate limit; takes regenerateVideoClips (bool?) — for a workflow with a VideoGenerate step, omitting it makes the tool refuse and return needsDecision plus the Max-Spend upper bound so the assistant asks the user (reuse = free, regenerate = paid) before re-calling; optional styleId / stylePresetKey (one of them) and outputFormat run it in an editing style and shape, resolved by RunStyleResolver before the paid-clips question and before a rate-limit slot is spent (see editing-styles.md)

Project (8) — ListProjects and GetProject (read-only); ProposeCreateProject / ApplyCreateProject, ProposeUpdateProject / ApplyUpdateProject, ProposeDeleteProject / ApplyDeleteProject (three Propose/Apply pairs).

File (5) — ReadProjectFile (read-only); ProposeWriteProjectFile / ApplyWriteProjectFile, ProposeDeleteProjectFile / ApplyDeleteProjectFile (two Propose/Apply pairs).

Workflow (6) — ListWorkflows, GetWorkflow (a workflow's complete steps in the exact CreateWorkflowStepRequest shape ValidateWorkflow/ProposeWorkflow accept, so it can copy a workflow or change one setting), ListWorkflowExecutions and GetWorkflowExecution (run history; per-step status, error, cache reuse, whether a video was produced, and an output preview truncated to 1500 characters) — all read-only and ownership-checked like every project-scoped tool; ProposeDeleteWorkflow / ApplyDeleteWorkflow.

Knowledge (3) — SearchPlatformDocs (semantic search over the shipped /docs tree, see platform-docs-search.md), ListWorkflowTemplates and GetWorkflowTemplate (the built-in template catalog; the latter returns ready-to-propose steps with each agent type resolved to the built-in agent row's id, exactly as applying the template from the UI does). All read-only and need no project in context.

Styles (3) — ListStyles (read-only: the built-in presets, the caller's saved styles usable in the project in context, and the output formats; needs no project) and ProposeSaveStyle / ApplySaveStyle (create a saved style, or replace one the caller owns; the confirmation payload binds the target's id and current name plus the normalized name, description, project pin and sanitized style, recomputed from the database on apply). Another user's style is reported as not found. See editing-styles.md "Assistant tools for styles".

Documentation is also injected. Besides the tool, every turn's system prompt carries the documentation sections most similar to the user's latest message (PlatformDocsContextBuilder, best-effort, bounded by a timeout and a character budget, never persisted to the thread). See platform-docs-search.md.

Plain language. The prompt's "How you talk" section tells the assistant to speak to users plainly — no step type identifiers, agent type names, config keys, JSON or ids in replies unless the user asks — and its "Knowing what ReelBolt can do" section tells it to search the documentation before declaring anything impossible and to start from a built-in template when one fits.

Admin (13, in two authorities) — the tools registered by AssistantAdminTools. They do not share one "admin" bit: ten are platform-scoped and three are organization-scoped, and the two gates are independent (see Per-call authority). TestInferenceProvider runs the same connection test as the admin page's Test button (both go through InferenceProviderTester: same unsupported-combination and private-endpoint guards, same per-capability probe, same persisted LastTestAt/LastTestOk/LastTestError); it is not a Propose/Apply pair because it changes no configuration, only the last-test columns, and its result's error text is key-scrubbed. ListInferenceProviders, ListAgentsAdmin and GetDecisionCalibration are read-only (the last reads the WorkflowEngine's calibration report for the caller by forwarding the caller's own bearer token, so the engine's admin check applies to the same identity, and it only ever recommends — it has no write counterpart); ProposeCreateInferenceProvider / ApplyCreateInferenceProvider, ProposeUpdateInferenceProvider / ApplyUpdateInferenceProvider, ProposeDeleteInferenceProvider / ApplyDeleteInferenceProvider, ProposeAssignAgentSkills / ApplyAssignAgentSkills (four Propose/Apply pairs). ListSkills and the skill pair read and write one organization's own agents, so they are the three organization-scoped tools; the other ten are platform-scoped.

That is 19 read-only tools, TestInferenceProvider, 12 Propose/Apply pairs (24 tools) and ExecuteWorkflow — 45 in all.

No sandbox, shell, filesystem or render tool is ever granted. The only tool that writes file content (ApplyWriteProjectFile) goes through IFileStorageService.UploadAsync, never the filesystem.

Scope guard​

AssistantAdminTools is also the boundary of what the assistant deliberately does not do: user-management tools (create/update/delete user, password reset) are absent, blocked on decision memo DM-016, and workflow-service admin actions are absent because the WorkflowEngine exposes read-only status only (GET /api/v1/workflow-engine/status and GET /api/v1/workflow-engine/skills/status, both admin-only).


Per-call authority: platform and organization​

Authority is per call and server-derived; it is never cached.

  • AssistantToolProvider holds the request-scoped ICurrentUser and reads it inside the tool body that is running. It does not hold a user id, an admin boolean, or any other derived authority in a field, a closure, a Lazy or a static.
  • AssistantToolContext carries the acting UserId and the caller's page context (ProjectId, WorkflowId) and deliberately carries no authority — no IsAdmin, no role, no permission set. An admin claim that travelled with the conversation would let a prompt-injected instruction inside an admin's session outlive the privilege that authorised it.
  • The page context is not a grant either: every tool that uses ProjectId re-resolves the project from the database and re-checks project.OwnerId == UserId on every call.

Two authorities, never one blurred "admin" bit​

The assistant's 13 gated tools are exactly the tools registered by AssistantAdminTools, and they are not gated on one privilege. Which authority governs depends on what the tool acts on, and the two are independent:

AuthorityRead asToolsWhat they act on
PlatformICurrentUser.IsPlatformAdmin10 — the inference-provider tools (ListInferenceProviders, TestInferenceProvider, and Propose/Apply for create, update and delete), ListAgentsAdmin and GetDecisionCalibrationThe platform-managed rows — inference_providers.organization_id IS NULL — the deployment's own resources, plus the platform-wide calibration report.
OrganizationOrgAuthority.IsOrganizationAdmin(ICurrentUser.OrgRole) — Owner or Admin of the active organization3 — ListSkills, ProposeAssignAgentSkills, ApplyAssignAgentSkillsOne organization's own custom agents and their skill assignment.

Two consequences follow, and both are intentional:

  • Being an organization's Owner grants nothing on the platform tools. The provider tools act on platform-managed rows only, so a workspace's own bring-your-own provider is as unreachable through them as a nonexistent one — the same answer /api/v1/inference-providers gives for it. Before A11 those tools read through the EF tenant filter, which also admits the caller's own organization's rows, so a platform admin could edit or delete a workspace's own provider and clear that workspace's default as a side effect.
  • Being a platform admin grants nothing on the organization tools. A platform admin who is only a Member of their active workspace is refused on the skill tools exactly as any other Member is, and needs no platform claim at all if they are an Owner or Admin there.

Across both, existence never leaks: another organization's custom agent answers {"error":"agent not found"} — the same answer a nonexistent id gets — whatever authority the caller holds.

The gated subset is not maintained by hand: AssistantToolProvider.GetAdminToolNames filters the provider's own registered list through the same AssistantAdminTools.GetTools list the registration comes from, so a new gated tool is covered the moment it is registered and a tool that forgets its gate fails the test rather than escaping it.

The gate itself:

  • The authority is re-read from the request scope inside every gated tool body, and the gate is the first executed statement of every tool method. It is never memoized in a field, a Lazy, a closure or a static, and no tool infers authority from an earlier tool result or from the conversation — an authority claim that travelled with the conversation would let a prompt-injected instruction inside an admin's session outlive the privilege that authorised it. ICurrentUser reads IsPlatformAdmin (the platformAdmin claim, itself re-validated against the database on every request by the JWT bearer's membership validator) and OrgRole (the role in the active organization, likewise re-read from the membership row, never from the token). Two identities that differ only by authority get different outcomes from the same provider instance.
  • A refusal writes nothing. The refusal is inert: the caller gets {"error": "admin privileges required"} and the database is untouched. AssistantAdminToolTests asserts the DbContext, not just the response string, because a tool that returns an error after writing is still a privilege-escalation bug.
  • Secrets never come back out. A provider's API key may be an input (create/update — it is the admin's own write), but no result ever contains the ciphertext or a decrypted key: reads return only the redaction vocabulary hasApiKey / apiKeyLastFour.

Confirmation tokens​

Every mutating tool is backed by IConfirmationTokenService (Services/Assistant/ConfirmationTokenService.cs), registered as a singleton — one key per deployment, shared by every surface, because a per-scope instance would derive its own key and reject tokens its sibling had just minted.

Propose* derives a canonical payload, mints a token with Compute(operation, canonicalPayload, userId) and shows the model a summary. Apply* re-derives the same payload from the same sources and asks Verify(operation, canonicalPayload, userId, presented) to check it. A token that does not verify produces a JSON error and writes nothing — the payload, not the model's claim about it, is what authorises the write.

The MAC​

Compute returns lowercase hex HMAC-SHA256 over a length-prefixed message:

<operation.Length>:<operation>\n
<canonicalPayload.Length>:<canonicalPayload>\n
<userId as "D">

The explicit lengths matter: without them the operation name and the payload could be re-split at their separator, and a token minted for one pair would verify against another.

Verify recomputes the token and compares with CryptographicOperations.FixedTimeEquals — not string equality — so a byte-by-byte probe of a token learns nothing from response timing. It returns false for an empty or length-mismatched token rather than throwing, and fails closed on any exception in the computation path.

The key​

Resolved lazily, on first use, so constructing the service (at DI validation time, or in a test) never touches the filesystem. In order:

  1. Assistant:ConfirmationTokenKey — an operator-supplied, Data-Protection-protected base64 key. If it is set but cannot be decrypted with the current key ring, the service logs a warning and falls back rather than silently keeping a second key: tokens minted under the configured key are rejected, which is recoverable by re-proposing, whereas a silent divergence would be invisible.
  2. Assistant:ConfirmationTokenKeyFile — otherwise the persisted key is read from here.
  3. First run — 32 random bytes are generated, protected with ISecretProtector, and persisted with a write-then-rename so a crash mid-write cannot leave a half-written key. The default path is <DataProtection:KeysPath>/assistant-confirmation-token, and DataProtection:KeysPath defaults to /keys — the shared dpkeys volume both services already mount — so every replica and every restart derives the same key. If the key cannot be persisted, the service degrades to a process-local key rather than refusing to start; the only consequence is that a token proposed before a restart is rejected after it.

The key is never hardcoded, never logged and never returned; only its role in the MAC is observable.

What a canonical payload binds​

  • Destructive operations bind the target's id AND its exact current name, and the name is re-read from the database at apply time. A confirmation the user saw for a project called "Archive 2024" therefore cannot be honoured against a project that has since been renamed — the recomputed payload no longer matches and the apply is refused instead of destroying something the user was never shown. This applies to ProposeDeleteProject, ProposeDeleteWorkflow, ProposeDeleteProjectFile, ProposeDeleteInferenceProvider and ProposeAssignAgentSkills.
  • ProposeWriteProjectFile binds a hash of the content, so any edit to the body between propose and apply invalidates the token.
  • A provider create/update payload carries a SHA-256 fingerprint of the submitted API key, not the key, so the plaintext never enters a payload the model could be shown. An empty string means "clear the stored key" and a null means "leave it unchanged", and those two are distinct in the payload.

The one unkeyed operation — an honest asymmetry​

ApplyWorkflow is a compatibility carve-out, not a template. It predates this service, and its token is pinned byte-for-byte, because a token proposed before a deploy has to stay acceptable after it. The pin is documented where it is implemented — ConfirmationTokenService's own class remarks and the ApplyWorkflowOperation constant's doc comment, against the keyless digest WorkflowEditDiffService has always produced — and no decision record states it. ConfirmationTokenService.Compute therefore reproduces the legacy unkeyed SHA-256 digest of the canonical payload — no key, no user — for that one operation name (ConfirmationTokenService.ApplyWorkflowOperation) and for nothing else. Every other operation is HMAC-SHA256 over a keyed, user-bound message. New mutating tools must not copy it.


The web surface​

The shipped client is one assistant, mounted once for the whole application.

Mounting and trigger​

  • Mounted once in web/app/(app)/layout.tsx, which wraps the app shell in AssistantShellProvider. The provider renders AssistantDrawer unconditionally — never behind opened && — because Mantine is what hides a closed drawer, and the drawer's own runtime provider has to stay mounted while it is closed. That is what keeps the conversation from being discarded every time the panel is shut to look at the canvas.
  • Triggered from AppShell's header (web/components/shell/AppShell.tsx), by AssistantTrigger next to the mobile Burger. The trigger lives in AppShell rather than in a page header so every authenticated page reaches it, including the majority that render no page header at all.

Shape​

AssistantDrawerShell (web/components/assistant/AssistantDrawer.shell.tsx) follows the viewport:

ViewportShape
DesktopLeft drawer — position="left", size={560}
Below Mantine's sm breakpointBottom sheet — useMediaQuery('(max-width: 48em)', false) → position="bottom", size="85%"

The set of confirmation tokens already handed to the page lives in the drawer shell for the lifetime of the thread, and a token is forwarded at most once: reopening the drawer remounts every message, each message's diff extractor re-emits its proposal, and a proposal that was applied, discarded or superseded must not come back.

There is exactly one assistant​

The workflow page's page-local assistant is deleted — it no longer mounts a drawer of its own. The page opens the shell's assistant and contributes to it:

  • <AssistantPageContext projectId? workflowId? /> tells the shell which project and workflow the page on screen is about. It renders nothing, and it clears the context on unmount so the next page never inherits this one's project.
  • useAssistantPageContext({ onWorkflowDiff, currentSteps, review }) hands over the handler that applies proposals, the steps a diff is drawn against, and the page's own review tray, for as long as the page is mounted. Proposals arrive in the data-workflowDiff SSE frame.

Attachments​

The composer's attach control does not send files through the chat stream. It uploads through the project's existing file endpoint — uploadFile in web/lib/api/files.ts — so an attachment is an ordinary ProjectFile the rest of the platform already understands, and then sends an ordinary user turn naming the new project file id. The turn is persisted like any other. With no project in context the control is disabled and says so out loud ("Open a project to attach files"), because there is nowhere to put the file.

Client API​

web/lib/api/assistant.ts holds the transport helpers: getAssistantChatStreamUrl() returns /api/v1/assistant/chat and takes no arguments (the old getAssistantChatStreamUrl(projectId) threw on a non-GUID and could not be called from a project-less page), alongside listThreads, createThread, getThread, renameThread and deleteThread.


Assistant execution rate limiting​

The IAssistantExecutionGuard enforces a rolling-hour cap per user + project:

  • Config key: Assistant:MaxExecutionsPerHour (default 10)
  • Window: last 60 minutes from now
  • Counted: WorkflowExecution rows where Initiator == ExecutionInitiator.Assistant and InitiatedByUserId == userId and ProjectId == projectId and StartedAt >= (now - 1 hour)

When the assistant calls ExecuteWorkflow, the guard checks how many assistant-initiated executions the user has started in the past hour. If below the limit, the execution is allowed. If at or above, the tool returns a rate-limit error. The REST path (POST /api/v1/projects/{projectId}/workflows/{workflowId}/assistant/execute) returns 429 for the same condition. This protects against token-spend exhaustion: a single user can run at most 10 assistant workflows per hour per project.

The guard uses StartedAt (when the execution actually began) rather than CreatedAt, and counts only rows where StartedAt is not null, because queued executions waiting for a concurrency slot should not count toward the limit until they start.

The count check and the insert run under a lock — a process-local semaphore keyed on (userId, projectId) plus a PostgreSQL advisory lock — so two concurrent calls cannot both observe count < max and both proceed.


The MCP server​

The MCP server (mcp/src/server.ts, built in mcp/Dockerfile) runs as an independent, stateless Node.js service. It is NOT proxied by nginx and provides both HTTP and stdio transports:

  • Stdio mode (default): spawned by a CLI client as a child process, reads REELBOLT_MCP_TOKEN once from environment at startup
  • HTTP mode (when MCP_TRANSPORT=http): listens on MCP_PORT (default 3002), expects Authorization: Bearer <jwt> header on every POST

Both modes verify the token once at the transport boundary (never as a tool-call argument a model could omit or forge).

It is deliberately still the seven workflow tools​

The MCP server does not follow the assistant's route change and is not a mirror of the 45-tool platform assistant. It remains the seven workflow tools it has always been:

const tools = {
listAgents: createListAgentsTool(serverConfig),
listStepSchemas: createListStepSchemasTool(serverConfig),
listProjectFiles: createListProjectFilesTool(serverConfig),
validateWorkflow: createValidateWorkflowTool(serverConfig),
proposeWorkflow: createProposeWorkflowTool(serverConfig),
applyWorkflow: createApplyWorkflowTool(serverConfig),
executeWorkflow: createExecuteWorkflowTool(serverConfig),
};

Those tools wrap the still-live, project-scoped workflow REST surface — not the assistant's /api/v1/assistant/* routes:

MCP toolInference API endpoint it calls
validate_workflowPOST /api/v1/projects/{projectId}/workflows/assistant/validate
propose_workflowPOST /api/v1/projects/{projectId}/workflows/assistant/propose
apply_workflowPOST /api/v1/projects/{projectId}/workflows/assistant/apply
execute_workflow (request.regenerateVideoClips, default false = reuse clips)POST /api/v1/projects/{projectId}/workflows/{workflowId}/assistant/execute

WorkflowAssistantController still serves all four, and the project-scoped shape is right for a CLI client, which always knows which project it is working in — the "no project" case that forced the browser transport into the request body does not arise there. Bringing the MCP server up to the assistant's full surface is deferred, together with the assistant's own user-management tools, on DM-016; the WorkflowEngine meanwhile exposes read-only status only (GET /api/v1/workflow-engine/status, GET /api/v1/workflow-engine/skills/status).


Step configuration schemas​

The StepConfigSchemaExporter (in inference/src/ReelBolt.Shared/Workflows/StepConfigSchemaExporter.cs) uses reflection to scan the ReelBolt.Shared.Workflows namespace for sealed records ending in StepConfig. For each StepType enum value, it reflectively generates a JSON Schema using .NET's JsonSchemaExporter and exposes it via:

  • Inference API: GET /api/v1/workflow-step-schemas (list all) and GET /api/v1/workflow-step-schemas/{stepType} (single)
  • MCP: the ListStepSchemas tool calls the Inference API endpoint and returns the result

The schema describes what a step's config JSON must look like. This contract is consumed by the frontend workflow builder component (workflow-builder-ux) to:

  1. Validate user-built steps client-side before proposing them
  2. Render dynamic forms for each step type's config fields
  3. Display helpful errors when the user's config doesn't match the schema

Authentication​

The assistant and MCP server use the same ReelBolt JWT the Go API issues (HS256, includes sub userId, email, and the organization claims org, orgRole and platformAdmin; isAdmin is still emitted as an alias of platformAdmin for one release). No new IdP, no separate permission model. The same JWT that authenticates to the Inference API (projects, workflows, files) authenticates to the assistant and MCP server. All three services share JWT_SIGNING_KEY in the environment.

Inference API (/api/v1/assistant/*)​

Standard [Authorize] attribute on AssistantChatController and AssistantThreadsController — JWT extracted from Authorization: Bearer <token> header by ASP.NET Core's authentication middleware. Authentication establishes who the caller is, never the whole authorisation: ownership is checked per call (thread UserId, the project's owning organization, the workflow's project), and the platform/organization authority is re-read per call inside each gated tool body.

MCP Server​

Stdio mode: Read REELBOLT_MCP_TOKEN once at startup, verify it with verifyReelBoltJwt() (checks issuer, audience, signature, algorithm = HS256 only). If invalid/missing, exit with status 1.

HTTP mode: Extract Authorization: Bearer <token> from the request header, verify at the transport boundary. If invalid/missing, return 401 Unauthorized.

Organization scope of a token​

A ReelBolt token is scoped to the organization it was minted in: the org claim names the active organization and orgRole the caller's role there at issue time, while platformAdmin (alias isAdmin) is platform-wide authority. The MCP server's verifyReelBoltJwt() accepts tokens with or without isAdmin (platformAdmin wins when both are present, and a token carrying neither is a non-admin) and exposes org, orgRole and platformAdmin on its typed auth context; org/orgRole are absent on legacy tokens minted before organizations existed, which keep working. These values are informational: the bearer token is forwarded to the Inference API unchanged, and the Inference API re-reads membership and applies org scoping itself. To work in a different organization from the CLI or an external client, mint a new token by switching organization in the dashboard (or via the Go API's switch-org endpoint) and restart the stdio server with the new REELBOLT_MCP_TOKEN, or send the new bearer token in HTTP mode; an existing token never changes organization.


Deployment​

In docker-compose​

The MCP service is defined in docker-compose.yml:

mcp:
build:
context: ./mcp
restart: unless-stopped
environment:
JWT_SIGNING_KEY: ${JWT_SIGNING_KEY}
JWT_ISSUER: ${JWT_ISSUER:-reelbolt-api}
JWT_AUDIENCE: ${JWT_AUDIENCE:-reelbolt-inference}
INFERENCE_API_URL: http://inference:8080
PORT: ${MCP_PORT:-3002}
MCP_TRANSPORT: http
depends_on:
inference:
condition: service_healthy
networks:
- reelbolt

The MCP service:

  • Is NOT proxied by nginx. It has no entry in nginx/locations.conf — by deliberate design. The Inference API's /api/v1/assistant/* endpoints (proxied by nginx) are the web-facing chat surface for browsers. The MCP server is supplementary: for CLI tools, external services, and programmatic workflows.
  • Listens on port 3002 (configurable via MCP_PORT), internal to the compose network only. Not published to the host unless explicitly configured.
  • Depends on the Inference API being healthy, since every MCP tool call proxies through it.

To use externally, either expose the HTTP port (ports: ["0.0.0.0:3002:3002"]) or spawn it as a subprocess with stdio transport.

Inference API settings the assistant reads​

Config keyDefaultNotes
Assistant:MaxExecutionsPerHour10Rolling-hour cap on assistant-initiated executions, per user + project.
Assistant:MaxOutputTokens16384Output-token allowance for one assistant turn's visible answer. ReasoningHeadroomTokens (8,192) is added on top on a provider that charges thinking to the same ceiling, so an Anthropic turn is sent 24,576 by default and a configured 4,096 is sent as 12,288. Only sent at all — on a kind other than Anthropic — when this key is set.
Assistant:MaxMessageChars32000Both halves of a turn: a user message over it is refused with a 400; an assistant reply over it is truncated with a visible marker.
Assistant:MaxReplayMessages40How many of a thread's newest persisted turns are replayed into the next prompt.
Assistant:DetachedTurnSeconds300How long a turn keeps running after the browser that asked for it disconnected.
Assistant:ConfirmationTokenKey(unset)Operator override: a Data-Protection-protected base64 HMAC key. When unset, the key is generated and persisted.
Assistant:ConfirmationTokenKeyFile<DataProtection:KeysPath>/assistant-confirmation-tokenWhere the generated key is persisted.
DataProtection:KeysPath/keysThe shared dpkeys volume, mounted at /keys in both inference services so either can decrypt what the other wrote.
Agents:Assistant:SystemPrompt(unset)Overrides the assistant's built-in system prompt.

Environment variables (MCP)​

VariableDefaultNotes
JWT_SIGNING_KEY(required)Shared with Go API and Inference API; must be >= 32 characters
JWT_ISSUERreelbolt-apiMust match the Go API's JWT:Issuer
JWT_AUDIENCEreelbolt-inferenceMust match the Go API's JWT:Audience
INFERENCE_API_URLhttp://localhost:3001Base URL of the Inference API (used by MCP tools)
MCP_TRANSPORTstdioSet to http to use HTTP instead
PORT3002Port to listen on (HTTP mode only)

Why the MCP server is not proxied by nginx​

Nginx proxies the Inference API's public REST endpoints (/api/v1/*) so clients can reach them from browsers and external networks with TLS termination and cookie↔header translation. The MCP server is not for browsers:

  • Browsers cannot maintain stdio connections — stdio transport is for CLI tools spawned as child processes only
  • HTTP mode is for programmatic clients (external assistants, APIs) that already have authentication machinery, not browser sessions needing cookie handling
  • nginx's strength is authentication at the request boundary — MCP's model is a bearer token verified once at transport startup (stdio) or per request (HTTP), not cookies + header translation

The Inference API's assistant endpoints (/api/v1/assistant/chat and /api/v1/assistant/threads) fall under nginx's location /api/v1/ block, which forwards to the Inference API, so browsers can reach the chat UI. If you need HTTP-based external access to the MCP server, expose its port explicitly or use a separate TLS terminator.

Page awareness, titles and docking (QA 2026-10)​

  • Docking. The drawer opens on the RIGHT (bottom sheet on phones); the header trigger sits in the right-hand header group.
  • Who and where. Every chat turn carries pageContext (path, kind, title, documentTitle, tab, entityIds, optional extra), derived client-side from the route (components/assistant/pageInfo.ts). The server treats it as an untrusted hint: AssistantSessionContext bounds every field, allowlists ids/paths, strips control characters and brackets, caps extra at about 2 KB, and appends the resulting block to the agent's SYSTEM prompt for that turn only (never persisted, so it cannot go stale). Identity (id, email, display name, admin) is read server-side from ICurrentUser and the user row; admin authority is still re-checked inside each gated tool.
  • Welcome state is page-aware (welcomeFor(kind)): a greeting plus 2 to 4 suggestion chips per page kind, generic fallback.
  • Thread titles. The first message (truncated) names a thread immediately; after the first exchange a short separate model call (10 s budget, AssistantThreadTitler) replaces it, persisted on the thread and announced with a data-threadTitle stream part. The client reloads the history list when a reply settles. Failure keeps the truncated title.
  • History list. Grouped Today / Yesterday / Earlier, active thread highlighted, relative times, search (shown above 6 threads), inline rename and confirm-delete, empty state.
  • Opening with a prompt. useAssistantShell().openWithPrompt(prompt, extraContext?, { send? }) (see docs/screenshots/qa-2026-10/NOTES-assistant.md).