DeepSeek V4 Pro API: What Changed and How to Use It

DeepSeek shipped the production build of V4 Pro overnight with zero fanfare — the API docs just quietly switched over. The model id is still deepseek-v4-pro, but the fingerprint flipped to fp_v4pro_20260812, ending a 111-day preview. Same architecture, very different model: a post-training overhaul that turns V4 Pro into a genuinely usable agent model, plus first-class image reasoning — at unchanged prices.

What changed in the 0813 build

The skeleton didn’t move: 1.6T total params, 49B active (MoE), 1M context, 384K max output tokens. The gains come entirely from post-training, and they land exactly where the preview was weakest — long-horizon tool use and code modification:

  • DeepSWE (DeepSeek’s software-engineering benchmark): 12.8 → 62.7, roughly a 4.9x jump
  • Cybergym: 52.7 → 83.3
  • Terminal Bench: 72.1 → 87.9
  • DSBench-Hard: more than doubled, to 67.2

The other headline: native image reasoning. The preview was text-only; the release build can analyze screenshots, diagrams, and mixed documents inside the same DeepThink reasoning flow — no separate VLM hop needed.

Calling it

Same OpenAI-compatible endpoint, same model name. Vision works through the standard image_url message part:

curl https://api.deepseek.com/chat/completions \
  -H "Authorization: Bearer $DEEPSEEK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-pro",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Walk me through this error screenshot and suggest a fix."},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}
      ]
    }],
    "max_tokens": 8192
  }'

The thinking-mode gotcha (read this before long codegen)

DeepThink reasoning tokens count against max_tokens, and on complex generation tasks they can eat 80%+ of your output budget. In one real test, a 16,384-token budget burned 13,394 tokens on “thinking” — the HTML output truncated mid-file. Re-running with thinking disabled produced a complete 715-line implementation in 54 seconds.

Practical rule: for long structured outputs (code, configs, docs), either turn thinking off when the task doesn’t need deep deliberation, or raise max_tokens well past your expected output size. With a 384K ceiling there’s plenty of headroom.

Where it fits in your stack

  • High-volume agent loops: the benchmark jump is real — multi-step tool use, long-horizon coding, environment manipulation now hold up against frontier closed models. At ¥3/M input tokens (roughly 1/10 of Fable 5, 1/3 of Kimi K3), it’s the default for cost-sensitive agent pipelines.
  • Still not the pure-reasoning king: HLE (no tools) scores 42.7 vs Opus 4.8’s 49.8 and Fable 5’s 53.3. For hard reasoning or the very toughest SWE tasks, the closed frontier models still edge it out — pay for them only when you need that last mile.
  • Price window: DeepSeek announced on Aug 6 that API pricing will rise “significantly” soon. The 0813 build is still at preview pricing — lock in long-running workloads now and budget for the increase.
  • UI/screenshot agents: with native vision in the same reasoning flow, you can drop the separate VLM routing layer for screenshot-understanding tasks.

Resources

Related News