One reference · Three agent setups

The Drowned Court

An experiment from the team behind OhMyUnicorn Forge.

Can a smaller model directing fresh Astra builders reach the same visual quality as Astra working alone?

Reference: a red-cloaked traveler crossing a flooded, overgrown gothic courtyard
The fixed reference image. Generated before either run; this is the target, not a browser render.
11 September 2026 · Solo model comparison

A full DeepSeek rerun

New experiment: DeepSeek High vs Maximum, run in parallel — side-by-side scenes, elapsed time and API costs ↗

DeepSeek V4.1 Flash takes the original solo Astra challenge: the same reference, untouched seed, three visual passes and a 45-minute allowance.

DeepSeek wrote the scene, viewed its screenshots and chose every refinement. All recorded worker responses identify deepseek-flash. The shared Forge harness supplied rules, skills, tools and coordination. There were no other model builders or visual directors.

Original controls for these snapshots: click / tap to move, drag to orbit, scroll to zoom. The keyboard follow-up below belongs to the earlier A/B/C builds.

DeepSeek produced a working scene at much lower model cost; Astra retained stronger visual detail. DeepSeek captures the flooded court, lantern warmth, a red traveler and reflections. Its broad slabs, plain arch walls, sparse vines and oversized lily pads look closer to a blockout. Astra’s masonry, arch mouldings, foliage and traveler are more articulated. This is an unblinded visual judgment from one comparison, not a general model ranking.

All six common state-based interaction checks passed in every DeepSeek pass. Those checks did not notice that the traveler was absent from the first two renders; DeepSeek added it in pass three. Extra validation exposed a negative first-frame time step and a drifting reset camera, both repaired by DeepSeek. The final reset, pointer/touch and animation checks passed. Public verification is linked below.

Three passes, preserved

DeepSeek pass 1 browser capture
Pass 1 · 4m 47s
Explore this pass ↗
DeepSeek pass 2 browser capture
Pass 2 · 6m 45s
Explore this pass ↗
DeepSeek pass 3 browser capture
Pass 3 · 20m 46s
Explore this pass ↗

Time and model usage

PassOriginal solo AstraFull DeepSeek
Pass 14m 04s to capture
473.9K in / 8.5K out
$1.127
4m 47s to capture
695.2K in / 32.9K out
$0.0557
Pass 27m 09s to capture
549.3K in / 9.8K out
$1.298
6m 45s to capture
1,070.6K in / 13.8K out
$0.0249
Pass 312m 09s to capture
1,090.1K in / 4.5K out
$1.512
20m 46s to capture
4,570.5K in / 25.4K out
$0.0625

Times are cumulative from each lane’s launch to screenshot completion, an upper bound on first visible rendering. Tokens and costs are incremental, including each pass’s self-review. Input includes cache hits. Pass boundaries follow the next pass’s first edit; pass three includes final verification and repairs. Its capture time uses the accepted recapture after those repairs.

DeepSeek total: 6,336.2K in / 72.1K out · $0.1430. Finishing outside the pass cells: $0.0000. Root setup, accounting, verification and publication are separate: $10.244 API equivalent through 2026-09-11 09:24:10 UTC.

Conditions, pricing and limits

The source seed and reference hashes match B. Both use Three.js 0.180.0, Vite 7.1.5, the same rendering host and the unchanged 1536 × 1024 browser probe. Both requested low effort; effort levels are provider-specific. Astra used the native Codex lane; DeepSeek used Claude Code’s tool runtime through its official API adapter. These are single exploratory runs with shared task conditions, not identical model interfaces or a general benchmark.

DeepSeek authored all geometry and canvas textures in code. It did not use Blender, generated textures, third-party art or a finished asset kit; an image-generation tool was not exposed to this lane. The inherited reference is a generated illustration, not a measured in-engine target. Neither result establishes reference-level realism. Captures use SwiftShader software rendering; hardware frame rate is unmeasured.

DeepSeek pricing uses the 10 September provider announcement: peak rates of $0.30 fresh input, $0.006 cached input and $1.20 output per million tokens. This weekday run occurred in the 06:00–10:00 UTC peak window. Usage is deduplicated by provider response ID; the runtime’s Claude-price estimate is ignored. Estimates are not invoices. Tool fees, image generation and server compute are excluded. No subscription-quota saving is inferred.

Root overhead includes this rerun only, excluding the earlier DeepSeek integration. Its final response after the ledger cutoff is excluded. DeepSeek omitted the first two source/build archives; root recovered exact logged file states and rebuilt them with the same dependencies. Its written status included inaccurate timestamps, so those were rejected. DeepSeek’s total includes additional reset/mobile verification absent from the recorded original B timing. Host load was not controlled. Original development probes logged a favicon 404; that is retained in the evidence.

What we’re comparing

Shared starting conditionsMeasured separately
Reference, seed, Three.js version, tools, hardware and time allowanceVisual match, working controls, browser errors and elapsed time
Astra builder at low reasoning effortNative session token usage, including the coordinator

These are three exploratory runs, not a benchmark. Software-rendered captures do not establish hardware frame rate. Shared preparation and publication time are reported separately.

Protocol and starting point

All lanes started from the same minimal interactive scene and the same reference image, using Three.js 0.180.0. Each received up to three visual passes and a 45-minute execution allowance. A used Luna with an Astra relay; B used Astra alone; C used Luna with direct fresh-context Astra workers. Runs were sequential.

The shared starting scene: a gray plane and a rust-colored capsule
The shared starting scene.

What A and B showed

Astra alone finished sooner; Luna directing fresh Astra workers had a lower API-equivalent model cost. A completed in 31m 28s; B in 15m 25s. Both produced coherent, explorable Gothic dioramas. Neither approached the reference’s realism.

In the final captures, B has more visible stone texture, a more articulated traveler and stronger warm lamp reflections. A has denser ground vegetation and more architectural tracery. Both retain regular masonry, geometric distant ruins and conspicuous repeating water patterns. This is an unblinded judgment from one pair, not evidence of a general model winner.

Both final A/B builds pass readiness, movement, camera follow, orbit, zoom and animation checks. Lane C passes all six checks in every completed pass after a parent-owned follow-timing repair in pass three; its reset report records zero camera, target and player drift. C’s static production probe is also green. Hardware frame rate was not measured.

Runtime evidence and experiment limits

Final development captures: A 95 draws / 611,385 triangles; B 61 draws / 377,670 triangles. SwiftShader median frame intervals were 1,183 ms and 1,917 ms respectively, on different scenes. These software-rendered measurements are diagnostic, not playable hardware FPS.

The harness did not expose worker-spawning tools inside Luna’s child session. Root relayed Luna’s briefs to fresh Astra workers and returned their results. That added a startup detour and manual handoffs. The topology was therefore Luna-directed with root relay, not a fully nested autonomous coordinator. B did not inspect A’s work. Both used the same target, untouched seed, tools, maximum three passes and 45-minute allowance.

The original scene validation used a hash-identical static copy on the rendering host. This public edition opens without a login; its page, scene routes and CTA are checked on the public URL. The scene bundles are unchanged.

A production verification · B production verification · A full production probe · B full production probe

A source archive · B source archive · C source archive

Three passes: time, tokens & usage

Run sequentially: A, then B, then C. Time is cumulative from each lane’s start to its first captured render for that pass. Tokens and costs are incremental per pass, including review. K = thousand tokens; input includes cached input and output includes reasoning.

PassLane A · Luna + AstraLane B · Astra aloneLane C · Full Luna direction
Pass 111m 32s from lane start
Luna: 2,484.1K in / 18.7K out · $0.090
Astra: 147.8K in / 6.7K out · $0.702
Pass total $0.792
Account quota used: 38% → 38%*
4m 04s from lane start
Astra: 473.9K in / 8.5K out · $1.127
Pass total $1.127
Account quota used: 39% → 39%*
7m 50s from lane start
Luna: 5,095.8K in / 20.9K out · $0.142
Astra: 344.1K in / 9.9K out · $1.169
Pass total $1.311
Account quota used: 40% → 40%*
Pass 220m 13s from lane start
Luna: 1,059.9K in / 6.8K out · $0.041
Astra: 214.7K in / 3.7K out · $0.611
Pass total $0.651
Account quota used: 38% → 39%*
7m 09s from lane start
Astra: 549.3K in / 9.8K out · $1.298
Pass total $1.298
Account quota used: 39% → 39%*
18m 35s from lane start
Luna: 1,764.6K in / 9.8K out · $0.059
Astra: 334.9K in / 4.3K out · $0.709
Pass total $0.768
Account quota used: 40% → 40%*
Pass 328m 27s from lane start
Luna: 695.0K in / 4.7K out · $0.025
Astra: 333.6K in / 4.8K out · $0.991
Pass total $1.016
Account quota used: 39% → 39%*
12m 09s from lane start
Astra: 1,090.1K in / 4.5K out · $1.512
Pass total $1.512
Account quota used: 39% → 40%*
26m 28s from lane start
Luna: 3,163.3K in / 10.4K out · $0.092
Astra: 359.8K in / 5.5K out · $0.844
Pass total $0.936
Account quota used: 40% → 40%*

Subscription usage: what the logs actually show

*These are account-wide meter readings, not the usage of an individual lane. The recorded 7-day allowance was 38% → 39% during A, 39% → 40% during B, and 40% → 41% during C (C’s increase occurred after pass 3, during finishing). The readings are rounded to whole percentages, and concurrent coordinator work consumed the same allowance.

Exact subscription usage by lane, pass or model is not measurable from these records. An unchanged reading does not mean zero usage. We cannot use this meter to verify the original post’s 80–90% subscription saving or convert API dollars into subscription percentage.

Observed subscription meter log · How Codex subscription usage works

Lane totals: A $2.459 · B $3.937 · C $3.181 USD. C includes $0.029 of director setup and $0.136 of director publication/reporting outside the three pass cells.

C used approximately 19% fewer API-equivalent model dollars than B. The original post’s 80–90% claim concerns subscription quota; the available account-wide readings cannot establish lane usage or reproduce that saving. C reproduces direct Luna-to-fresh-Astra orchestration. The visual workflow differs from the original demo, so this is not a full reproduction of its conditions.

Timing, pricing and separate overhead

A/B accounting cutoff: 15:18:44 UTC. Updated C and follow-up accounting cutoff: 2026-09-10T17:06:41.718610Z. C began 16:24:26 UTC. First screenshot completion is an upper bound on first rendering, not instrumented first paint. C’s later reset-verification timestamps remain in the detailed ledger.

Standard API equivalent, not an invoice. Per million tokens: Astra $10 input / $1 cached / $50 output; Luna $0.20 / $0.02 / $1.20. Cache writes and long-context multipliers follow Astra pricing and Luna pricing.

Original shared root work: $15.820. Additional root Astra follow-up/setup/observation/publication: 5,556.7K input / 42.5K output, $9.414. Separate native Luna harness preflight: 54.1K input / 0.5K output, $0.003. All recorded language-model work across the experiment and follow-ups: $34.814. These shared costs are outside the lane comparison. Image generation, tool fees and server compute are unpriced and excluded; responses after the cutoff are excluded.

A/B exact pass figures · C exact pass figures · C usage and time log · Separate follow-up overhead · Original full usage · A/B request log

Keyboard controls added after the experiment: arrow keys and WASD / ZQSD move the character. The recorded captures, timings and lane costs describe the original runs. Control update source manifest · Keyboard checks · Repair time, tokens & cost.