DeepSeek Harness Goes Open Source: Everything Is a Plugin

DeepSeek made two announcements in one day. Early this morning it released the DeepSeek V4 Pro flagship model; by noon, the DeepSeek Harness developer preview hit GitHub. The model got the headlines, but the harness is the story that will age better.

DeepSeek Harness is not another model and not another API client. It is an SDK and application framework for building, running and extending AI agents. The repo already has more than 230 workspace members, with code spread across packages/, apps/, examples/, python/, native/ and vendor/. File system access, terminals, subprocesses, PTYs, language servers, web browsing, skills, subagents, workflows, planning modes, session persistence, credentials and telemetry — nearly every capability has its own package.

The most striking design statement is the project’s core slogan: everything is a plugin. Even the agent loop itself is a plugin.

A breadboard, not a finished computer

If a typical agent project is a pre-assembled computer, DeepSeek Harness is an unusually large breadboard. Models, tools, interfaces, storage, security policies and context management can all be plugged in — and pulled out.

The project sits on the Cordis microkernel. A running Harness instance is essentially a Cordis Context: packages register services, events and capabilities, and a config file assembles them into a working agent.

packages/core/ handles the basics — what a session is, how system prompts are assembled, how tools register and get called, how an agent is created, and how one turn flows from user input to model request to tool execution to final answer. Around the core sit capability packages: packages/llm/ for model adapters and streaming, packages/shell/, subprocess/ and terminal/ for commands and process trees, packages/fs/ for file operations, packages/lsp/ for semantic code navigation, packages/web/ for search and fetching, packages/skill/ for reusable skills, and packages/subagent/ plus packages/workflow/ to extend a single agent into a multi-agent system.

The architecture has a deliberate three-layer split: interface, implementation and consumer. Take Bash as an example. The interface defines what “execute a command” means; the local implementation actually spawns processes; the model-facing toolkit turns that capability into a schema and results the model can understand. If a local shell later needs to become a remote container, a cloud sandbox or an enterprise execution platform, you swap the implementation layer — without rewriting the model tools or the agent loop.

This is framework thinking. It makes the early repo look bloated, but it also means DeepSeek’s goal is not a finished product only its own team can maintain. Different deployments can swap models, storage, security policies, tools — even the agent loop itself.

cordis.yml: one codebase, many agents

The plugin architecture lands for developers through cordis.yml. The config file lists plugin names, stable IDs and parameters, and decides which capabilities an agent actually has.

The same code assembles into completely different product shapes. Add the DeepSeek LLM adapter, file system, Bash and a TUI, and you get a terminal coding agent. Swap the interface for the Web plugin and it becomes a browser app. Use the headless entry point and it takes a task, runs model-and-tool rounds, prints an answer and exits. Point an ACP or JSON-RPC front door at it, and it becomes an automation service other programs can drive.

Configs also support overlays. TUI and Web UI can share a base config and layer their own interface plugins and parameters on top, with personal config at the last level. Deployers don’t copy the whole config tree — they patch specific plugins. One caveat worth flagging: a config patch replaces the target plugin’s entire config rather than deep-merging. Add one new field and existing API keys or base URLs can silently disappear. It’s explicit, but not intuitive on first use. For a deeper look at what DeepSeek’s harness push means, see our earlier analysis of why DeepSeek stopped being a model-only company.

The agent loop as traffic rules

Many early agent projects boil down to a few lines: send messages to the model, execute tool calls when they come back, return results, repeat until the model outputs text. DeepSeek Harness does that too, but it splits the process into a strict lifecycle.

One user input opens a Turn; a Turn contains multiple Steps, each corresponding to one model request plus its tool execution. Before a request, the system assembles a stable system prompt, current runtime environment, tool schemas and session messages. After it, streaming chunks, full messages, tool calls, tool results and end reasons all flow into an event stream.

Tools are not invoked the moment a name is returned. They pass through pre-policies, irreversible safety guards, actual execution, post-processing, content tidy-up and result notification. Allow or deny, timeout, retry, metrics and extra context can hook into different stages of the pipeline. A tool can declare itself concurrency-safe for certain parameter shapes, and the scheduler runs consecutive read-only tasks in parallel; when it hits a state-mutating or uncertain call, it treats it as a barrier and runs it exclusively after earlier tasks finish.

This looks like installing air-traffic control on a country road — until an agent is searching ten files at once, running tests, accepting mid-run user commands and allowing cancellation at any moment. Then those rules turn from over-engineering into the thing you wish you’d had when writing the postmortem.

The project also takes mid-run messages seriously. New content from a user while the agent works can be a queued task or a steering instruction. The system distinguishes queued messages, injected context and steering, and confirms whether a steering instruction actually reached a model request. It doesn’t just care that the message was received; it cares which step of the model’s thinking saw it.

Session Log: the single source of truth

Another design worth calling out: anything the model sees must be reconstructable from the log. User messages, runtime context, model request info, streaming output, tool calls and results, compression events, permission switches and cancellation reasons all enter an append-only session stream as events. The UI, persistence, resume, fork, telemetry and replay should not each maintain a “mostly correct” copy of state — they derive from one event source.

This solves a genuinely painful problem in agent systems: when a task goes wrong, can you actually know what the model saw at that moment? If you only keep the final chat text, you lose what matters — workspace state injected before the request, truncated tool results, an automatic model-routing switch, a user changing direction mid-stream. Harness saves enough at request boundaries to rebuild messages, and keeps raw streaming chunks so the UI and replay stay consistent.

Session persistence is itself a plugin, with JSONL and SQLite backends. Resume continues a session in place; Fork derives a new session from a defined history boundary. For developers, this gives a unified foundation for debugging, evaluation, auditing and automation.

From one agent to a crowd

Harness ships with subagent and workflow capabilities built in. A main agent can delegate to subagents — freshly created instances, forks from a completed boundary of an existing session, or external subprocesses connected over ACP. Scope design matters here: each agent has its own context layer and sees specific tools, prompts and commands. One subagent can be restricted to search and analysis while another is allowed to modify files.

The 36Kr team’s demo is a useful data point. They asked a V4-Flash-configured Harness to build a first-person zombie shooter with one prompt and no mid-run intervention; 30-odd minutes later they had a playable — if rough — game. The same team ran the same prompt style on Codex with GPT-5.6 sol-xhigh for a 3D animation task and got noticeably worse results, despite the larger model. The harness, not the model, carried the work.

Why this matters

Harness is becoming the new battleground in AI infrastructure. Codex, Claude Code, OpenClaw and now DeepSeek Harness are all fighting in this layer, but they’re fighting different battles. Most are closed products optimized for one experience; DeepSeek chose to open-source the assembly method itself. That’s the structural shift: agents stop being applications you buy and become assemblies you compose — with an open microkernel, swappable implementations and an auditable event log as the foundation. We traced the first signs of this in why the harness is the new battleground in AI, and this piece goes deeper into the architecture.

The open-source bet also has a strategic edge. If the harness layer becomes the interface between frontier models and real work, whoever owns the default assembly gets distribution — for models, for tools, for cloud compute. By open-sourcing it the same day as the V4 Pro release, DeepSeek is saying the model is only half the product — the same strategic call MiniMax made when it argued model companies must build their own harness.

What to do about it

Developers should clone the repo and run a session — even the headless entry point is enough to feel the difference between an agent and an assembled agent. Deployers evaluating agent infrastructure should test the three-layer split: can you swap storage, security policy or execution backend without touching the loop? Security and audit teams should look at the Session Log principle — an append-only event source is exactly what compliance wants. And anyone building on top of agent frameworks should treat “everything is a plugin” as the new default expectation, because the closed-app era of agents just got a deadline.

Related News