Skip to main content

Embeddings (semantic file search)

Semantic search over project files (POST /api/v1/projects/{id}/files/search, the engine's search_project_files tool) embeds file chunks and queries into Qdrant. The embedding backend is an inference-provider capability, Embedding.

Resolution​

ResolveEmbeddingAsync: explicit id -> the single enabled default Embedding row -> legacy AzureOpenAI:Endpoint / AzureOpenAI:ApiKey (+ VectorSearch:EmbeddingDeployment, default text-embedding-3-large) -> none. With none, search answers indexNotReady: true (no 500).

Kinds: AzureOpenAI, OpenAICompatible (text-embeddings-inference, vLLM, Ollama, LiteLLM). Others are rejected with a 400. DeepSeek documents no embeddings endpoint (only deepseek-flash/deepseek-v4-pro chat models; api-docs.deepseek.com, 2026-10). Anthropic has none either.

Local model​

docker compose up -d embeddings starts text-embeddings-inference (CPU) serving Qwen/Qwen3-Embedding-0.6B (1024 dimensions) on 127.0.0.1:${EMBEDDINGS_PORT:-9003}. First start downloads the weights into the embeddings-hf-cache volume. In /app/admin/inference-providers add: Kind OpenAICompatible, Capability Embedding, endpoint http://embeddings:80/v1, model Qwen/Qwen3-Embedding-0.6B, default on, then Test connection (expects vector length 1024).

Index layout and switching models​

Collections are named {projectId:N}_{model-slug}_{hash8}_{dimension}, one per project, model and dimension, so models/dimensions never mix. Switching the default provider leaves the old collection untouched and unsearched; the new one is empty until files are reindexed (POST /api/v1/projects/{id}/files/reindex, or POST .../files/{fileId}/reindex). Per-file deletes clear the file from every collection of the project, including the pre-capability bare {projectId:N} collection. Orphaned collections can be dropped from Qdrant manually.

Qdrant API key and TLS​

The Inference API can talk to a hosted Qdrant (for example Qdrant Cloud) as well as the bundled compose Qdrant. Set VectorSearch__QdrantApiKey (compose: VECTOR_QDRANT_API_KEY) together with an https:// VectorSearch__QdrantUrl (compose: VECTOR_QDRANT_URL); the gRPC port is VectorSearch__QdrantGrpcPort. The key is passed to every Qdrant client, for both file search and the platform docs index. A key with a non-https URL is refused: the Inference API fails at startup with a message naming both settings, so the key is never sent unencrypted. With no key (the default) behaviour is unchanged and plain http://qdrant:6333 keeps working. The index can always be rebuilt from the database by reindexing, so a new Qdrant cluster starts empty and fills via POST /api/v1/projects/{id}/files/reindex.

Hosted ASR and embeddings​

A cloud deployment does not run the whisper and embeddings containers. Instead the Inference API seeds platform provider rows (no organization, so every organization can use them) from the PlatformProviders configuration array at startup. The seeder creates or updates a row by Name, never deletes a row, never replaces a stored key with an empty value, and does nothing when the section is empty (self-host is unchanged). Each entry passes through the same validation as the platform admin API, so an unsupported kind/capability pair or a private endpoint is logged and skipped without blocking the other entries or startup.

Each entry takes Name, Kind, Capability, Endpoint, Model, ApiKey, IsDefault, Enabled (default true) and TimeoutSeconds. As environment variables: PlatformProviders__0__Name, PlatformProviders__0__Endpoint, and so on. The cloud control compose file wires two slots from PLATFORM_ASR_* and PLATFORM_EMBEDDING_* (see infra/compose/env.control.example); a blank *_NAME leaves the slot unconfigured.

  • ASR: Kind=OpenAICompatible, Capability=Transcription, an endpoint speaking the OpenAI /v1/audio/transcriptions API, and the hosted model name.
  • Embeddings: Kind=OpenAICompatible or AzureOpenAI only (the Embedding allowlist), Capability=Embedding.
  • Keys: the env: sentinel is accepted only for kinds that support ambient credentials (Anthropic, Gemini, DeepSeek); for ASR and embeddings supply the key itself from a secret (docs/secrets.md). The seeder refuses env: for any other kind.
  • Dimension warning: collections are keyed per project, model and dimension. The existing collections are Qwen3-Embedding-0.6B at 1024 dimensions. Pick a hosted model that outputs 1024 dimensions to keep them, but note that the collection name also contains the model name, so a different model always starts an empty collection. Either way, reindex after switching (POST /api/v1/projects/{id}/files/reindex), or file search reports indexNotReady. The seeder does not reindex.