DeepSeek Elastic Compute (DSec) is the sandbox infrastructure behind DeepSeek’s agent training, evaluation, and data preparation from V3.2 through V4.1. Agents need stateful environments where they can inspect code, edit files, install dependencies, run tests, and start services across many turns. DSec is built for bursts of sandbox creation, long periods of low CPU use with memory still resident, diverse environments, and training jobs that may be preempted.

This is an English translation of the substantive technical content of a DeepSeek article published in Chinese on September 29, 2026. The accompanying technical report provides the full methods and results.

One interface, four execution backends

Different tasks need different isolation levels, operating-system features, and execution costs. DSec exposes four backends through a single Python SDK, libdsec:

BackendIntended use
FnCallReuses pre-created containers for short tasks such as online evaluation.
ContainerOffers fast startup and high density for common software-engineering and tool-use tasks.
MicroVMGives security-sensitive tasks a stronger isolation boundary.
Full VMProvides a complete operating system for graphical interfaces, rendering, Android, and similar applications.

Composable environment layers

At scale, environment construction and updates become a problem in their own right. In one production week in 2026, the container backend used 11,266 base images, 102,171 workspaces, and hundreds of toolkits. Rebuilding monolithic images whenever any one component changes would repeat a great deal of work.

DSec splits each environment into independently versioned layers: a base image with the operating system and basic software, a workspace with task code and dependencies, and a toolkit such as the DeepSeek Harness. It stores these parts as EROFS images and combines them at sandbox creation with OverlayFS. EROFS also separates metadata from data and supports deduplication across images. Updating a toolkit or workspace therefore requires rebuilding only the affected layer.

Load image data on demand

DeepSeek found that running sandboxes actually accessed only 4.2% to 13.3% of the total bytes in their images. DSec keeps image data on the 3FS distributed file system, downloads the needed EROFS metadata locally, and fetches file data as it is read.

In an experiment that created 8,192 containers in a burst, on-demand loading reduced task completion time from more than 60 minutes to about 35 minutes compared with pulling complete images. That was about a 1.71× speedup and cut disk writes by roughly 57%. In a separate workspace-supply experiment, mounting an EROFS layer instead of unpacking a tar.gz archive in each sandbox reduced completion time from 79 to 45 minutes and disk writes to about 1/5.5 of the original amount.

Manage resources at high density

During agent training, sandboxes often wait for the model to generate its next action. About 90% of sandboxes averaged no more than 5% of their requested CPU capacity, while their memory had to remain available to preserve files and processes. This supports CPU overcommit exceeding 50× in production.

For MicroVMs on one host, DSec uses virtio-pmem and DAX so they can share a single host page cache. Enabling that mechanism alone lowered peak host memory use by 40.2% relative to the baseline in DeepSeek’s experiment. Memory reclamation with DAMON and balloon free-page reporting alone reduced time-integrated host memory consumption by 21.2%; combining the mechanisms produced the lowest overall memory use.

At high density, DSec also gives latency-sensitive work priority over tasks with looser timing requirements and uses core scheduling to reduce interference between hyperthreads on the same physical core. When colocated tasks consumed 50% of a node’s CPU capacity, scheduling changes reduced the latency increase for sensitive tasks from 45.2% to 17.3%, relative to a no-interference baseline.

Decouple agent rollouts from GPU training

In reinforcement learning, an agent must interact with its sandbox over several turns to produce a rollout. Earlier, the agent loop lived in the same Pod as the GPU training job. If that job was preempted, the sandbox survived but the loop that advanced the interaction stopped. Recovery required replaying command logs to reconcile the trainer’s saved progress with the sandbox’s actual state.

Starting with DeepSeek-V4.1, DSec moved the execution logic outside the preemptible GPU pool. An agent sandbox runs the agent framework and toolkit, while a worker container manages sandboxes and advances interactions. Together they retain the rollout’s progress and environment state. A GPU job can therefore resume an interrupted rollout without reconstructing it from command logs.

Let agents build their own environments

Agent training and evaluation require many combinations of binary dependencies, repositories, harnesses, and grading scripts. DSec uses the same platform to run the agents that construct these environments and the agents that later train in them, keeping construction and runtime conditions aligned.

Its pack_diff mechanism lets an agent request an incremental snapshot of a sandbox and later restore it as a new sandbox. A snapshot at step k can also branch a rollout into multiple explorations from the same state. Branches share read-only layers and store only their changes, avoiding a replay of the first k steps. MicroVM snapshots can preserve memory and process state; container snapshots currently focus mainly on disk state.

Bound agent behavior

DeepSeek reports agents trying to read leftover answers, forge RPC requests, overwrite /bin/bash to inject commands, and use XFS_IOC_SWAPEXT to bypass access controls. Reward-seeking agents can exploit any available shortcut, damaging both evaluations and their environments.

DSec uses AppArmor to restrict file and socket access even when an agent has administrator privileges. It also uses eBPF to enforce a network allowlist per sandbox, including addresses, ports, and protocols. These controls reduce some risks, but DeepSeek says there is still no general defense against destructive behavior such as triggering kernel defects.

Production scale

DSec scales through shards. Each shard contains about 160 servers, around 30,000 CPU cores, and 250 TB of memory. One shard serves about three million sandboxes per day, reaches more than 380,000 concurrent sandboxes, and can create over 5,000 sandboxes per second. Multiple production shards support millions of simultaneously running sandboxes.

DeepSeek’s broader point is that large-scale agent training depends on reliable, diverse execution environments as much as on the training loop itself. DSec treats those environments as shared, stateful infrastructure that can be composed, scheduled, snapshotted, and secured at scale.

Source: DeepSeek, “DeepSeek 弹性计算 (DSec):面向大规模 Agent 训练的沙盒基础设施,” September 29, 2026. Technical report on arXiv.