DeepSeek inference provider
ReelBolt can run any agent on DeepSeek by registering an inference_providers row of kind
DeepSeek. This document covers how that provider is wired, how to obtain an API key,
and supported capabilities.
What this is, and what it deliberately is not
This is DeepSeek as a model backend accessed through DeepSeek's official OpenAI
compatibility layer (https://api.deepseek.com/ or https://api.deepseek.com/v1).
ReelBolt already has an agent harness: ReelBoltAgentBase + AgentStepExecutor +
ToolGroupCatalog + the room infrastructure supply the agent loop, structured output,
retry-with-feedback, tool scoping, step caching, and a container sandbox.
Rather than introducing any third-party or conflicting SDK that could introduce dependencies on
Microsoft.Extensions.AI.Abstractions, ReelBolt routes DeepSeek requests through its
verified and unified OpenAI client (client.GetChatClient(model).AsIChatClient()).
DeepSeek's OpenAI compatibility endpoint natively supports:
- Standard chat completions and streaming
- Vision / multimodal content parts (
deepseek-flashaccepts images) - Function calling / tool invocation
- Reasoning models (e.g. DeepSeek-R1 / DeepSeek-V3 / deepseek-reasoner / deepseek-flash)
The integration is cleanly contained as one arm of ChatClientFactory.Build — all agents remain
completely provider-agnostic.
How to get an API key
- Navigate to the DeepSeek Platform.
- Sign in or register an account.
- Go to API Keys and generate a new key (
sk-...). - Copy the generated API key.
Configuring a provider in ReelBolt
Admin → Inference Providers → New, or POST /api/v1/inference-providers:
| Field | Value |
|---|---|
| Kind | DeepSeek |
| Capability | Chat or Vision (not Transcription — see below) |
| Base URL | https://api.deepseek.com (prefilled automatically) |
| Model | e.g. deepseek-chat, deepseek-reasoner, deepseek-flash |
| API key | Paste your API key, or use the env: sentinel (see below) |
| Timeout | Defaults to 300s |
Set the row as IsDefault for Chat to route all agents through DeepSeek, or assign it as a
per-agent override with PUT /api/v1/agents/{id}/inference-provider to test DeepSeek on a single step.
The two credential modes
ChatClientFactory.BuildDeepSeek branches on the stored key:
- The literal
env:— ReelBolt readsDEEPSEEK_API_KEYfrom the container's environment variables. Export the variable in.env, storeenv:as the key in the database, and no secret is persisted in the database. - Direct API Key (
sk-...) — The API key is stored encrypted at rest using ASP.NET Core Data Protection (ISecretProtector) on the shareddpkeysvolume (/keys).
An empty key is an error, not mode 1. An empty or whitespace-only key throws
InvalidOperationException rather than silently falling back to ambient credentials. This prevents
misconfigured rows or failed decrypts from silently routing billing to the host's account.
Capabilities
Chat: Fully supported. Serves agent reasoning and structured completions.Vision: Fully supported. DeepSeek models such asdeepseek-flashprovide vision / multimodal capabilities;VideoAnalyzeshot captioning works out of the box using inline base64 JPEG images.Transcription: Rejected with HTTP 400 at the API boundary on create, update, and test. DeepSeek does not expose an OpenAI-compatible speech-to-text (/audio/transcriptions) endpoint.TranscriptionClientFactorythrows an explanatoryNotSupportedExceptionas defense in depth. Continue using the bundledwhisperservice (or any OpenAI-compatible ASR provider) for transcription.
Vision calls and the JSON-schema response format
DeepSeek answers a JSON-schema response_format with HTTP 400 ("This response_format type is
unavailable now"). ReelBolt's vision callers — VideoAnalyze shot captioning and the object locator
that seeds TrackObjects — go through StructuredChatFallback (WorkflowEngine), which retries a
rejected request with plain JSON mode and then with no format, since every vision prompt already
demands strict JSON and parses defensively. The fallback remembers the format a provider
accepted, keyed by the provider row's id plus its model name, for 12 hours per engine process, so a
DeepSeek vision row pays the rejected first attempt once rather than on every call. A provider that
accepts the schema format leaves nothing remembered and behaves exactly as before; if the remembered
format is later rejected too, the call falls back to the next plainer one and remembers that; and
after 12 hours the schema format is tried again, so a provider that adds support is picked up
without a restart.
Reasoning Models & Token Headroom
DeepSeek reasoning models (such as deepseek-reasoner) generate chain-of-thought tokens before
emitting answer content. Because reasoning tokens count against max output tokens, the provider
connectivity ping (POST /api/v1/inference-providers/test) allocates 1,024 tokens for DeepSeek
(compared to 128 tokens for non-thinking endpoints), ensuring test pings do not get truncated mid-thought.
No Schema Migration
InferenceProviderKind is persisted as a string (.HasConversion<string>() in both DbContexts), so
adding DeepSeek requires no EF Core migration.