Embeddings (semantic file search)
Semantic search over project files (POST /api/v1/projects/{id}/files/search, the engine's
search_project_files tool) embeds file chunks and queries into Qdrant. The embedding backend is an
inference-provider capability, Embedding.
Resolution
ResolveEmbeddingAsync: explicit id -> the single enabled default Embedding row -> legacy
AzureOpenAI:Endpoint / AzureOpenAI:ApiKey (+ VectorSearch:EmbeddingDeployment, default
text-embedding-3-large) -> none. With none, search answers indexNotReady: true (no 500).
Kinds: AzureOpenAI, OpenAICompatible (text-embeddings-inference, vLLM, Ollama, LiteLLM).
Others are rejected with a 400. DeepSeek documents no embeddings endpoint (only
deepseek-flash/deepseek-v4-pro chat models; api-docs.deepseek.com, 2026-10). Anthropic has none either.
Local model
docker compose up -d embeddings starts text-embeddings-inference (CPU) serving
Qwen/Qwen3-Embedding-0.6B (1024 dimensions) on 127.0.0.1:${EMBEDDINGS_PORT:-9003}. First start
downloads the weights into the embeddings-hf-cache volume. In /app/admin/inference-providers add:
Kind OpenAICompatible, Capability Embedding, endpoint http://embeddings:80/v1, model
Qwen/Qwen3-Embedding-0.6B, default on, then Test connection (expects vector length 1024).
Index layout and switching models
Collections are named {projectId:N}_{model-slug}_{hash8}_{dimension}, one per project, model and
dimension, so models/dimensions never mix. Switching the default provider leaves the old collection
untouched and unsearched; the new one is empty until files are reindexed
(POST /api/v1/projects/{id}/files/reindex, or POST .../files/{fileId}/reindex). Per-file deletes
clear the file from every collection of the project, including the pre-capability bare
{projectId:N} collection. Orphaned collections can be dropped from Qdrant manually.
Qdrant API key and TLS
The Inference API can talk to a hosted Qdrant (for example Qdrant Cloud) as well as the bundled
compose Qdrant. Set VectorSearch__QdrantApiKey (compose: VECTOR_QDRANT_API_KEY) together with an
https:// VectorSearch__QdrantUrl (compose: VECTOR_QDRANT_URL); the gRPC port is
VectorSearch__QdrantGrpcPort. The key is passed to every Qdrant client, for both file search and the
platform docs index. A key with a non-https URL is refused: the Inference API fails at startup with a
message naming both settings, so the key is never sent unencrypted. With no key (the default) behaviour
is unchanged and plain http://qdrant:6333 keeps working. The index can always be rebuilt from the
database by reindexing, so a new Qdrant cluster starts empty and fills via
POST /api/v1/projects/{id}/files/reindex.
Hosted ASR and embeddings
A cloud deployment does not run the whisper and embeddings containers. Instead the Inference API
seeds platform provider rows (no organization, so every organization can use them) from the
PlatformProviders configuration array at startup. The seeder creates or updates a row by Name,
never deletes a row, never replaces a stored key with an empty value, and does nothing when the
section is empty (self-host is unchanged). Each entry passes through the same validation as the
platform admin API, so an unsupported kind/capability pair or a private endpoint is logged and
skipped without blocking the other entries or startup.
Each entry takes Name, Kind, Capability, Endpoint, Model, ApiKey, IsDefault,
Enabled (default true) and TimeoutSeconds. As environment variables:
PlatformProviders__0__Name, PlatformProviders__0__Endpoint, and so on. The cloud control compose
file wires two slots from PLATFORM_ASR_* and PLATFORM_EMBEDDING_* (see
infra/compose/env.control.example); a blank *_NAME leaves the slot unconfigured.
- ASR:
Kind=OpenAICompatible,Capability=Transcription, an endpoint speaking the OpenAI/v1/audio/transcriptionsAPI, and the hosted model name. - Embeddings:
Kind=OpenAICompatibleorAzureOpenAIonly (the Embedding allowlist),Capability=Embedding. - Keys: the
env:sentinel is accepted only for kinds that support ambient credentials (Anthropic, Gemini, DeepSeek); for ASR and embeddings supply the key itself from a secret (docs/secrets.md). The seeder refusesenv:for any other kind. - Dimension warning: collections are keyed per project, model and dimension. The existing
collections are Qwen3-Embedding-0.6B at 1024 dimensions. Pick a hosted model that outputs 1024
dimensions to keep them, but note that the collection name also contains the model name, so a
different model always starts an empty collection. Either way, reindex after switching
(
POST /api/v1/projects/{id}/files/reindex), or file search reportsindexNotReady. The seeder does not reindex.