Platform documentation search (assistant RAG)
The platform assistant answers questions about what ReelBolt can do from this documentation, not just from the step JSON Schemas. Before this, it only had the schemas and reasoned its way to confidently wrong answers — for example that nothing renders Remotion, or that a tracked screen insert has no content source — because a schema says what a field accepts, not what it is for or how steps combine.
What gets indexed
docs/user-guide/** (audience User) and the top-level developer docs docs/*.md (audience
Developer). archive/, design/, marketing/, screenshots/ and README.md are excluded — the same
set the docs site excludes.
The tree ships inside the Inference API image at /app/platform-docs
(COPY --from=platform-docs, a named build context pointing at ./docs; see the Dockerfile). A plain
docker build without that context ships an empty tree and search reports itself unavailable instead of
the build failing.
Chunking
PlatformDocsCorpus splits each file at ## headings (never inside a fenced code block), strips front
matter, and splits any section longer than MaxChunkChars at paragraph breaks. Each chunk carries the
page title, the section heading, its audience and its docs-site URL (/docs/... or
/docs/developer/...#anchor). What is embedded is "{title} — {section}\n\n{text}", so a short section
still lands near questions about its page. This is why the writing guideline is that every ## section
must stand on its own.
Index
QdrantPlatformDocsIndex keeps one global collection per (embedding model, corpus version):
platformdocs_{model-slug}_{model-hash8}_{corpus-hash12}
- The corpus hash covers every chunk plus a chunker version, so any doc edit produces a new collection.
- The model is part of the name for the same reason project-file collections key on it (embeddings.md): vectors from different models are not comparable.
- A new collection is built in full, then superseded
platformdocs_*collections are dropped. - "Already indexed" is a name lookup plus an exact point count; a partial collection from a crashed attempt is deleted and rebuilt.
- The
platformdocs_prefix can never collide with a project collection (those start with a 32-hex project id).
PlatformDocsIndexingService (hosted) runs EnsureIndexedAsync 15 s after startup and then every
RecheckIntervalSeconds (default 600; at most 120 s after a failure). That makes adding or switching the
default Embedding provider at runtime take effect without a restart. With no Embedding provider
configured, nothing is indexed and both retrieval paths below degrade to "unavailable".
Retrieval: injected and on demand
- Injected per turn.
PlatformDocsContextBuilderembeds the user's latest message, takes the topInjectTopKsections aboveInjectMinScore(user-guide sections get a small ranking boost), caps the block atInjectMaxChars, and appends it to the assistant's system prompt next to the session context block — never to the persisted thread, so it is recomputed every turn and never replayed stale. It is bounded byInjectTimeoutSecondsand degrades to nothing on any failure. - On demand. The
SearchPlatformDocstool lets the model search with its own query when the injected sections do not cover the question. The system prompt tells it never to declare something impossible without searching first. The tool is bounded bySearchTimeoutSeconds(see below).
Retrieved text is documentation, not user instruction; the injected block says so explicitly, matching the assistant's "tool output is data" rule.
Configuration (PlatformDocs:*, Inference API)
| Key | Default | Meaning |
|---|---|---|
Enabled | true | Master switch (indexing, tool, injection). Compose: PLATFORM_DOCS_ENABLED. |
Path | /app/platform-docs | Root of the shipped Markdown tree. |
RecheckIntervalSeconds | 600 | How often the indexer re-checks corpus/model. |
MaxChunkChars | 2400 | Longest chunk before a section is split. |
InjectTopK | 4 | Sections injected per turn (0 disables injection only). |
InjectMaxChars | 7000 | Character budget of the injected block. |
InjectTimeoutSeconds | 5 | Hard bound on per-turn retrieval. |
InjectMinScore | 0.35 | Cosine floor for injection (the tool ignores it). |
SearchTimeoutSeconds | 15 | Hard bound on one SearchPlatformDocs tool call. |
The search tool never holds a turn
SearchPlatformDocs embeds the model's query on the same embedding server that re-embeds the corpus
after a deploy and every uploaded project file. With no bound, a busy (often CPU-only) server left the
assistant on "Searching the documentation…" for minutes — reproduced twice with 60 s+ waits, and once
with no answer at all. Three things now keep the turn moving:
- A bound, with the injection path's semantics. One tool call gets
SearchTimeoutSeconds(default 15, a little more thanInjectTimeoutSecondsbecause the model asked explicitly). On expiry — and whenever the index is not built yet or is unreachable — the tool returns{ "available": false, "status": "timeout" | "still_indexing" | "unavailable", "reason", "guidance" }, whereguidancetells the model to answer from what it knows, say it could not check the documentation this time, and not retry in the same turn. Search itself never waits on the indexer's lock: it reads whichever collection was last confirmed complete. - A truthful label. Every tool reports a caption when it starts (
data-toolStatus) and another when it returns: "Thinking" normally, or "Docs index still building — answering without it" when the search came back unavailable. Before this the label kept naming the search for as long as the model then took. - Interactive priority. Between embedding batches the docs indexer waits (at most 15 s per batch)
while an assistant turn is in flight, as does per-file indexing (at most 10 s per chunk) and file
summarization on the chat provider (at most
FileSummarization:InteractiveYieldSeconds). The count of turns in flight isAssistantActivityMonitor, a process-wide singleton. See platform-assistant-mcp.md.
What leaves the host
The user's message (truncated to 2000 characters) is embedded through the default Embedding provider on
every assistant turn, and the docs chunks are embedded once per corpus version. With the local
embeddings container nothing leaves the host; with a hosted embedding provider, the user's message
does.
Re-indexing after a documentation change
A changed document produces a new corpus hash and therefore a new collection, but two things keep the assistant useful while it builds:
- Searches keep answering from the previous collection of the same Embedding model until the new one is complete — slightly stale, never blind. Before this, every deploy that touched a doc left the assistant answering "the documentation is still being indexed" for as long as a full re-embed took on the CPU embedding server (several minutes for ~700 chunks), and it fell back to reading whatever it could find in the project's own files.
- Unchanged chunks are not re-embedded. Every point stores
embedHash, a hash of exactly the text that was embedded; the new collection reuses the previous collection's vector for every chunk whose hash is unchanged, so a docs edit embeds only the edited sections. Collections written beforeembedHashexisted are re-embedded in full once.