Agents playing and developing games
DeepSeek Flash uses tools to play a native Amoris sailing game and modify TypeScript gameplay systems. The model makes decisions between simulation steps; the Rust host runs the game. The gameplay demonstration and development benchmark measure different tasks.
Native gameplay
The player receives a project-level tool projection with a 30 m observation range. It cannot call the developer's arbitrary world edits, script tools or snapshots through this projection. The boat moves through the engine's native Boat physics, driven by wind, sail, keel and rudder; tool calls submit intentions and advance fixed ticks.
| Player tool | Purpose |
|---|---|
observe |
Read the boat's player-facing state and nearby objects within the range limit |
helm |
Set sailing intentions for the next simulation steps |
navigate |
Select a nearby destination for the game's helmsman |
take |
Request a nearby parcel collection under the game's reach rule |
wait |
Advance a bounded number of simulation ticks |
navigate delegates continuous steering to the disclosed game-side Helm executor.
DeepSeek chooses destinations and other actions; it does not compute every rudder update.
This demonstration uses structured observations, rather than a vision model playing from
rendered pixels. The range filter is implemented by the sample's tool projection. Engine-wide
player authorization and pixel-based occlusion are still pending.
The three 2,048-token pretests collected 1/4, 1/4 and 2/4; every response limit was recorded. With low effort and an 8,192-token cap, the three formal trials collected 2/4, 4/4 and 4/4 (one length stop, two clean completions). A separate non-thinking, 4,096-token condition collected 2/4, 2/4 and 3/4 (transport error, request limit, and a clean but incomplete result). Conditions are reported separately.
The selected video is the first successful low-effort trial, repetition 2: 4/4 cargo, 7 worth, 1,577 simulated ticks. Every action, response and final world hash matched a fresh native replay. It is encoded at 30 fps from two-tick samples; recorded playback speed is separate from model wall time or rendering performance.
Game-development benchmark
The frozen courier-v1 experiment has six tasks × three independent trials = 18 runs:
two bug fixes and four gameplay features. Each trial starts from a fresh candidate project
with one defect or missing implementation. The agent reads, writes, applies and checks
scripts through the native host's developer API. It may edit only the specified module.
| Task | Kind | Required behavior | Passed / trials |
|---|---|---|---|
| Collection range | Bug fix | Inclusive horizontal radius 3 and vertical reach 2 | 3 / 3 |
| Parcel value scoring | Bug fix | Score each parcel's actual value exactly once, including zero | 3 / 3 |
| Collection victory | Feature | Persist and emit one victory for a completed nonempty objective | 3 / 3 |
| Dash cooldown | Feature | Four-times movement and a component-based five-tick cooldown | 3 / 3 |
| Score milestone | Feature | One score-ten milestone per courier, preserved across reload | 3 / 3 |
| Damage and defeat | Feature | Armor-reduced damage, clamped health and one defeat event | 3 / 3 |
18/18 candidate patches passed strict grading. All also passed types, native integrity, the five other feature suites and file scope. 15/18 sessions ended cleanly; three provider disconnects occurred after valid patches and remain separate in the results.
Median agent wall time: 86.95 s. Median native tool calls: 58. Overall prompt-cache hit fraction across reported usage: 97.29%. Recorded developer calls cost at most $0.416942 at the published peak rates, plus unsettled requests.
Every trial, final patch and tool trace · Readable result tables.
A trial passes only if its feature, all five unrelated feature suites, TypeScript checking, file scope and native determinism/fork/replay/reload checks pass. The golden controls exercise 63 behavioral assertions through real component state and emitted events. Checks include positive, negative and boundary cases, forced reload during dash cooldown, and an explicit Play fork that must leave the edit world unchanged after Stop.
Prompts and grader checks are frozen before model calls. The agent cannot read the grader, golden source, other trials or saved evidence. Every initial control fails its intended feature while passing types, native integrity and unrelated features; restoring the golden target module passes the entire suite. Controls establish that the test can distinguish a working implementation; they are not model successes.
All 18 scheduled trials belong in the results, including refusals, provider errors, timeouts and budget stops. Feature-only success is a diagnostic, not a substitute for the strict all-or-nothing result. These are short ECS gameplay tasks in one small game, with three repetitions each; they do not establish a general coding ranking or a comparison with other models.
Model, tools and cost accounting
The requested alias is deepseek-flash; DeepSeek currently maps it to
DeepSeek-V4.1-Flash. Model aliases can change, so each run records the request name and
provider response metadata. See DeepSeek's model and pricing documentation.
The runner uses a direct Chat Completions tool loop, following DeepSeek's tool-calling protocol. The client executes approved native tools and returns their results to the model. This path was chosen because the inspected OpenCode setup has no suitable per-request budget hook for this experiment.
Gameplay and all development trials share one $5 budget ledger. It checks the budget
before each provider request. Credentials are supplied through an environment variable or
an external key file; they are never embedded in game projects or public evidence. Public
records retain actions, tool results and usage totals without reasoning_content.
The accounting uses documented peak rates as a conservative upper bound: per million tokens,
$0.30 cache-miss input, $0.006 cache-hit input and $1.20 output, as checked on
5 October 2026. The ledger's token-based upper bound is not an invoice; billing_cost_usd
remains null unless a run-specific provider bill is available.
Pricing and billing rules.
All 18 development runs and all nine gameplay trials shared the authorized $5 ceiling.
Reported usage gives $0.694152 as the known peak-price
cost upper bound. 4 interrupted requests have no returned usage;
the ledger retains $1.277952 of conservative reservations.
The total committed upper bound is $1.972104, below $5.
The exact provider bill is unavailable and remains null. No automatic retries discard
these unknown charges.
Reproduce and inspect
The frozen protocol, task prompts and hashes, control results and native grader define the development experiment independently of the provider runner.
python3 tools/eval/agent_dev_bench.py list
python3 tools/eval/agent_dev_bench.py selftest \
--pocket ./target/release/pocket --output out/agent-dev-controls
python3 tools/eval/agent_dev_bench.py prepare --task dash-cooldown \
--project /absolute/fresh/candidate
python3 tools/eval/agent_dev_bench.py grade --task dash-cooldown \
--project /absolute/fresh/candidate --pocket ./target/release/pocket \
--output /absolute/fresh/grade
Use fresh candidate paths and run grades sequentially. The script refuses to overwrite an existing candidate. Provider runs additionally need the frozen prompts, tool restrictions, trial limits and shared budget ledger; the grader itself makes no model calls.
For the common native command surface, see CLI and MCP. For the underlying stateless systems and generated declarations, see Script API.
Run the real agents with one shared ledger (the credential file stays outside the repo):
python3 tools/showcase_assets.py --projects
python3 tools/eval/run_agent_showcase.py --mode development \
--credential-file /external/deepseek-key.txt --repeats 3 \
--output out/agent-showcase/development-v1
python3 tools/eval/run_agent_showcase.py --mode gameplay \
--credential-file /external/deepseek-key.txt --repeats 3 \
--game-output-tokens 8192 --game-effort low \
--output out/agent-showcase/gameplay-v2
A different run needs fresh output directories. The default shared ledger keeps the combined ceiling at $5; reusing it does not reset previous charges or reservations.