DeepSeek shipped the production build of V4 Pro overnight with zero fanfare — the API docs just quietly switched over. The model id is still deepseek-v4-pro, but the fingerprint flipped to fp_v4pro_20260812, ending a 111-day preview. Same architecture, very different model: a post-training overhaul that turns V4 Pro into a genuinely usable agent model, plus first-class image reasoning — at unchanged prices.
What changed in the 0813 build
The skeleton didn’t move: 1.6T total params, 49B active (MoE), 1M context, 384K max output tokens. The gains come entirely from post-training, and they land exactly where the preview was weakest — long-horizon tool use and code modification:
- DeepSWE (DeepSeek’s software-engineering benchmark): 12.8 → 62.7, roughly a 4.9x jump
- Cybergym: 52.7 → 83.3
- Terminal Bench: 72.1 → 87.9
- DSBench-Hard: more than doubled, to 67.2
The other headline: native image reasoning. The preview was text-only; the release build can analyze screenshots, diagrams, and mixed documents inside the same DeepThink reasoning flow — no separate VLM hop needed.
Calling it
Same OpenAI-compatible endpoint, same model name. Vision works through the standard image_url message part:
curl https://api.deepseek.com/chat/completions \
-H "Authorization: Bearer $DEEPSEEK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-pro",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Walk me through this error screenshot and suggest a fix."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}
]
}],
"max_tokens": 8192
}'
The thinking-mode gotcha (read this before long codegen)
DeepThink reasoning tokens count against max_tokens, and on complex generation tasks they can eat 80%+ of your output budget. In one real test, a 16,384-token budget burned 13,394 tokens on “thinking” — the HTML output truncated mid-file. Re-running with thinking disabled produced a complete 715-line implementation in 54 seconds.
Practical rule: for long structured outputs (code, configs, docs), either turn thinking off when the task doesn’t need deep deliberation, or raise max_tokens well past your expected output size. With a 384K ceiling there’s plenty of headroom.
Where it fits in your stack
- High-volume agent loops: the benchmark jump is real — multi-step tool use, long-horizon coding, environment manipulation now hold up against frontier closed models. At ¥3/M input tokens (roughly 1/10 of Fable 5, 1/3 of Kimi K3), it’s the default for cost-sensitive agent pipelines.
- Still not the pure-reasoning king: HLE (no tools) scores 42.7 vs Opus 4.8’s 49.8 and Fable 5’s 53.3. For hard reasoning or the very toughest SWE tasks, the closed frontier models still edge it out — pay for them only when you need that last mile.
- Price window: DeepSeek announced on Aug 6 that API pricing will rise “significantly” soon. The 0813 build is still at preview pricing — lock in long-running workloads now and budget for the increase.
- UI/screenshot agents: with native vision in the same reasoning flow, you can drop the separate VLM routing layer for screenshot-understanding tasks.