Native agent showcase evidence
Snapshot generated 2026-10-05T09:18:50.302067+00:00. Only completed agent summaries are counted.
Developer trials: 18 observed / 18 planned; 18 independently graded. Gameplay trials: 6 formal observed / 6 formal planned; 3 pretests retained separately.
Behavioral acceptance and agent completion are reported separately. Every recorded trial, including timeouts and transport errors, remains in the denominator.
| Task | n | Behavioral pass | Types | Integrity | Regressions | Scope | Wall mean / median (s) | API mean / median | Tool mean / median |
|---|---|---|---|---|---|---|---|---|---|
| Repair collection radius | 3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 71.97 / 67.77 | 32.67 / 32.00 | 49.67 / 49.00 |
| Repair parcel value scoring | 3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 88.78 / 93.48 | 33.33 / 34.00 | 57.00 / 59.00 |
| Implement collection victory | 3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 101.90 / 101.83 | 40.67 / 41.00 | 66.33 / 66.00 |
| Implement persistent dash cooldown | 3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 84.47 / 84.51 | 33.33 / 35.00 | 53.33 / 53.00 |
| Implement once-only score milestone | 3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 81.83 / 88.48 | 36.33 / 39.00 | 61.00 / 68.00 |
| Implement armored damage and defeat | 3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 96.81 / 98.90 | 42.00 / 41.00 | 66.00 / 68.00 |
| Task | Prompt tokens mean / median | Cache-hit tokens mean / median | Output tokens mean / median | Known cache hit | Peak-rate known cost upper | Unknown reserved |
|---|---|---|---|---|---|---|
| Repair collection radius | 438480.00 / 406214.00 | 425214.00 / 393597.00 | 9709.00 / 10405.00 | 96.97% | $0.054545652 | $0 |
| Repair parcel value scoring | 594017.00 / 477111.00 | 577322.67 / 464256.00 | 11194.67 / 11247.00 | 97.19% | $0.065717508 | $0.319488 |
| Implement collection victory | 899912.67 / 967723.00 | 877781.33 / 943232.00 | 12166.67 / 12419.00 | 97.54% | $0.079518264 | $0.319488 |
| Implement persistent dash cooldown | 801801.33 / 839915.00 | 777898.67 / 815488.00 | 12052.00 / 11902.00 | 97.02% | $0.078901776 | $0 |
| Implement once-only score milestone | 597069.00 / 616528.00 | 579285.33 / 601088.00 | 9777.67 / 11061.00 | 97.02% | $0.061632036 | $0.319488 |
| Implement armored damage and defeat | 796792.67 / 795731.00 | 778581.33 / 779392.00 | 12839.33 / 13326.00 | 97.71% | $0.076626264 | $0 |
Known provider-token totals use exact response usage; unknown requests contribute no invented tokens.
Known peak-rate cost upper bound across the shared ledger: $0.694151928. Unknown/unsettled request reservations: $1.277952 (4 requests). Actual provider invoice: unknown.
Reservations protect the shared budget and are not measured spend.
Trial outcomes
| Trial | Behavior | Agent finish | API / tool calls | Peak-rate known cost upper | Unknown reserved |
|---|---|---|---|---|---|
| development-v1/collection-range-1 | pass | stop | 25 / 43 | $0.018351840 | $0 |
| development-v1/collection-range-2 | pass | stop | 41 / 57 | $0.021012330 | $0 |
| development-v1/collection-range-3 | pass | stop | 32 / 49 | $0.015181482 | $0 |
| development-v1/score-values-1 | pass | transport_or_tool_error | 31 / 59 | $0.020138436 | $0.319488 |
| development-v1/score-values-2 | pass | stop | 34 / 60 | $0.028428324 | $0 |
| development-v1/score-values-3 | pass | stop | 35 / 52 | $0.017150748 | $0 |
| development-v1/objective-victory-1 | pass | stop | 33 / 66 | $0.021621168 | $0 |
| development-v1/objective-victory-2 | pass | transport_or_tool_error | 41 / 63 | $0.027939492 | $0.319488 |
| development-v1/objective-victory-3 | pass | stop | 48 / 70 | $0.029957604 | $0 |
| development-v1/dash-cooldown-1 | pass | stop | 29 / 53 | $0.025493316 | $0 |
| development-v1/dash-cooldown-2 | pass | stop | 35 / 55 | $0.026438628 | $0 |
| development-v1/dash-cooldown-3 | pass | stop | 36 / 52 | $0.026969832 | $0 |
| development-v1/milestone-event-1 | pass | stop | 39 / 68 | $0.022220928 | $0 |
| development-v1/milestone-event-2 | pass | stop | 48 / 72 | $0.023038020 | $0 |
| development-v1/milestone-event-3 | pass | transport_or_tool_error | 22 / 43 | $0.016373088 | $0.319488 |
| development-v1/damage-and-defeat-1 | pass | stop | 41 / 68 | $0.029784444 | $0 |
| development-v1/damage-and-defeat-2 | pass | stop | 37 / 57 | $0.019616568 | $0 |
| development-v1/damage-and-defeat-3 | pass | stop | 48 / 73 | $0.027225252 | $0 |
Gameplay
| Protocol / effort / max output | Phase | n | Native success | Wall mean / median (s) | API mean / median | Output tokens mean / median | Known peak upper | Unknown reserved |
|---|---|---|---|---|---|---|---|---|
| gameplay-v1 / low / 2048 | pretest | 3 | 0/3 | 84.72 / 69.44 | 30.00 / 27.00 | 11069.67 / 9060.00 | $0.055232280 | $0 |
| gameplay-v2 / low / 8192 | formal | 3 | 2/3 | 184.16 / 156.12 | 59.00 / 62.00 | 25166.00 / 17061.00 | $0.130148040 | $0 |
| gameplay-v3 / none / 4096 | formal | 3 | 0/3 | 151.58 / 144.95 | 84.67 / 89.00 | 10536.33 / 10132.00 | $0.091830108 | $0.319488 |
- gameplay-v1/episode-1 (pretest, max output 2048): 1 / 4 collected; native success False; agent finish
length; 18 API calls and 17 tool calls. - gameplay-v1/episode-2 (pretest, max output 2048): 1 / 4 collected; native success False; agent finish
length; 27 API calls and 26 tool calls. - gameplay-v1/episode-3 (pretest, max output 2048): 2 / 4 collected; native success False; agent finish
length; 45 API calls and 44 tool calls. - gameplay-v2/episode-1 (formal, max output 8192): 2 / 4 collected; native success False; agent finish
length; 45 API calls and 44 tool calls. - gameplay-v2/episode-2 (formal, max output 8192): 4 / 4 collected; native success True; agent finish
stop; 70 API calls and 69 tool calls. - gameplay-v2/episode-3 (formal, max output 8192): 4 / 4 collected; native success True; agent finish
stop; 62 API calls and 61 tool calls. - gameplay-v3/episode-1 (formal, max output 4096): 2 / 4 collected; native success False; agent finish
transport_or_tool_error; 75 API calls and 74 tool calls. - gameplay-v3/episode-2 (formal, max output 4096): 2 / 4 collected; native success False; agent finish
max_requests; 90 API calls and 90 tool calls. - gameplay-v3/episode-3 (formal, max output 4096): 3 / 4 collected; native success False; agent finish
stop; 89 API calls and 88 tool calls.
Media candidate: gameplay-v2/episode-2. Highest recorded native collection score; ties first in protocol/repetition order. Illustrative selected run; all recorded trials and conditions remain visible.
Frozen protocol
- development-v1:
courier-v1; efforthigh. - Fixture SHA-256:
126bf315b27a4bc1f6bc6282988606c48ab946dd5d9a343a75064a98d0ef7456. - Grader SHA-256:
9f95bc39e540208b50d5cdd7a59913a9e77d5ce889712c0a0748ff8bc6e094e8. - gameplay-v1: fixture
0d48ff5eaff3fbfa5a23b71e5883c811b176bf938f2a9726a4e6b583629ccfd0; gateway hash recorded at run:not recorded. - gameplay-v2: fixture
0d48ff5eaff3fbfa5a23b71e5883c811b176bf938f2a9726a4e6b583629ccfd0; gateway hash recorded at run:not recorded. - gameplay-v3: fixture
0d48ff5eaff3fbfa5a23b71e5883c811b176bf938f2a9726a4e6b583629ccfd0; gateway hash recorded at run:not recorded. - Gateway source snapshot SHA-256:
61c728abb00e500c2af60fc3bcdb5bcad7527b10a11e8e3c2416c81ddde10e1e. File at report generation; not a substitute for a pre-run implementation hash.
Actual response model aliases: deepseek-flash
- Six fixed development tasks and three planned sailing runs are a small, engine-specific study.
- Behavior is graded independently from model finish status; a transport interruption can leave a passing patch.
- Gameplay uses a project-level restricted gateway and a disclosed native Helm executor; this is not global player authorization.
- The model never receives the grader, golden files, private native state or shared budget ledger.
- Tool traces are optional bounded excerpts; their omission or truncation does not remove trials from statistics.
- The 2048-token gameplay-v1 pretests, low-thinking 8192-token gameplay-v2 runs and non-thinking 4096-token gameplay-v3 runs are separate frozen conditions; their success rates and performance averages are not pooled.
- Any highest-score media selection is disclosed and does not replace the complete trial table.