DeepSeek Harness: Run a Coding Agent with One Command

On August 13, DeepSeek shipped two things in one night: the V4 Pro flagship and DeepSeek Harness v0.1, an MIT-licensed agent runtime now in public developer preview. The model gets the headlines, but the harness is the more interesting release for anyone who builds agents. It is the same runtime DeepSeek used to benchmark its own coding agents, and it is built around a deliberately radical premise: everything is a plugin.

Why the harness layer matters

An agent is not a model. The model is the brain, but the harness — the engineering shell that lets it read files, call tools, manage context, retry after failures, and keep working for hours — is what makes it useful on real tasks. Claude Code proved this layer is commercially valuable; DeepSeek just open-sourced it. If you have been treating coding agents as black boxes, Harness is the chance to see (and rebuild) the whole box.

Quickstart: one command

Prerequisites: Node.js 18+. That is it. The npm package starts a Web UI that runs on your own machine:

# Start the Web UI (defaults to http://127.0.0.1:3080)
npx @deepseek-ai/dsh web

Open the URL in a browser. The left panel lists workspaces — each maps to a local project directory — the middle is the chat, and every tool call is laid out on a timeline. Prefer running from source?

git clone https://github.com/deepseek-ai/deepseek-harness.git
cd deepseek-harness
pnpm install
pnpm run build
pnpm dsh web

Four built-in modes

  • Standard — the full agent: file editing, shell, web and file search, skills, planning, goal management, subagents, workflows.
  • Code (PTC) — tools are exposed through a Code Mode SDK; the model writes a TypeScript program and merges what would be a dozen tool round-trips into one execution.
  • Minimal — just a persistent bash and a file editor. Official use: benchmarking. Strip the shell away and measure the raw model.
  • Creator — author your own mode and preset. The other three are just preset configs anyway.

Everything is a plugin

Under the hood is Cordis, a microkernel that came out of the Koishi chatbot framework ecosystem. Models, tools, skills, sessions, sandboxes, storage, loops, scheduling — even the UI you are looking at — are all plugins. Two properties make this practical:

  • Reversible side effects: every side effect a plugin registers is tracked and rolled back on uninstall. Plugins hot-swap without restarting the process — no leaks, no garbage left behind.
  • Discoverability: tag your repo with the dsh-plugin topic and it shows up in community directories like awesome-dsh-plugins (288 repos within a day of launch).

Useful starter plugins: dsh-plan-execute routes planning to a reasoning model and execution to a cheaper one (the bill drops immediately); dsh-vision bridges any OpenAI-compatible vision model so a text-only DeepSeek model can see images; sandbox isolation is a choice too — swap in sandbox-micro, sandbox-mxc, or sandbox-nono.

Every run is traceable

The Trajectory feature writes everything the model sees — system prompt, reasoning, tool calls and results, subagent scheduling, context injections — into an append-only session log. You can inspect each entry by source, then resume, fork, search, or replay any run. Debugging an agent becomes a replay problem instead of a guessing game.

Practical advice

  • Expect breakage. The README warns in all caps: THERE WILL BE COMPATIBILITY-BREAKING CHANGES. Pin your version if you build plugins.
  • Think about the sandbox. Plugins touch your shell and filesystem, so pick an isolation plugin and treat the attack surface seriously.
  • Benchmark with Minimal mode. It removes the harness variable from evals, so you compare models, not shells.
  • Pair it with V4 Pro. The same launch added native OpenAI Responses API support and a big agentic benchmark jump — see the DeepSeek V4 Pro API guide for integration details and token-budget gotchas.

Resources

Related News