11 September 2026 · DeepSeek V4.1 Flash · Parallel effort comparison

Does more thinking improve the scene?

High and Maximum start from the same seed and reference, with the same tools, a requested three-round limit and a 45-minute allowance each.

Higher effort did not close the gap to Astra. High is cleaner than Maximum in this sample, but both fall short in geometry, materials, water and lighting.

Click an image for the full-resolution capture.

DeepSeek · High

High effort final scene: tiled stone platforms around mostly opaque teal water, a red traveler and foliage-covered walls.

23m 25s · $0.0930 model cost

Cleaner than Maximum in this sample, but water lacks architectural reflections, the masonry texture looks tiled, the foliage is scattered, and the geometry and lighting remain crude beside Astra.

Explore this scene ↗

DeepSeek · Maximum

Maximum effort final scene: a dark stone courtyard, harsh water glare and large gray polygonal shapes floating above it.

15m 00s · $0.0781 model cost

Large floating gray polygons, harsh white water glare, repetitive masonry and obstructive walls remain. Maximum did not deliver better visual quality than High in this run.

Explore this scene ↗

Click or tap to move, drag to orbit, scroll to zoom. These are the original experiment controls.

View the fixed referenceThe fixed target: a flooded Gothic courtyard, moss-covered stonework and a red traveler in warm sunlight.

Earlier baselines

These runs happened separately. The camera actions and viewport match, but each model chose its own scene composition.

DeepSeek · Low

Earlier Low effort scene with slab arches and misplaced mirror-like reflections.

22m 47s · $0.1430 model cost

Simple slab architecture, weak surface detail and misplaced mirror-like shapes. The earlier run that motivated this effort comparison.

Explore this scene ↗

Astra · Low

Astra baseline with articulated Gothic arches, varied wet masonry, attached ivy and restrained warm lantern light.

15m 26s · $3.9368 model cost

Visibly stronger arch geometry, varied wet masonry, attached foliage, coherent reflections and restrained lantern lighting. Neither new DeepSeek run matches this result.

Explore this scene ↗

Elapsed time and pricing

RunElapsedModel cost¹At DeepSeek peak rates²Output tokens
DeepSeek Low (earlier)22m 47s$0.1430$0.143072,106
DeepSeek High23m 25s$0.0930$0.186199,047
DeepSeek Maximum15m 00s$0.0781$0.156294,511
Astra Low (earlier solo B)15m 26s$3.936822,728

High and Maximum launched 1.2 seconds apart. Both finished after 23m 25s of combined wall-clock time. Each lane’s time includes its setup, implementation, captures, self-review and finishing checks.

¹ Estimated API cost from provider response usage, deduplicated by response ID; Astra is an API equivalent, not a subscription invoice. High and Maximum used off-peak rates; Low used peak rates. ² The common peak-rate column removes this time-of-day discount from the DeepSeek comparison.

Recorded root setup, observation, accounting, verification and publication: $9.9085 API equivalent through 2026-09-11T10:37:22.435605+00:00. This overhead is shared by the two new runs and excluded from their model-cost cells. Server compute, tool fees and responses after that cutoff are excluded.

Pricing sources and accounting

DeepSeek per million tokens: off-peak $0.15 fresh input / $0.003 cached input / $0.60 output; peak $0.30 / $0.006 / $1.20. Provider announcement · Pricing table. Thinking content was present; the returned thinking-token detail counter is unreliable and is not treated as evidence that reasoning was disabled.

Astra baseline accounting and its pricing source are preserved in the original experiment. Costs here exclude earlier model integration and earlier diagnostic conversations. Root overhead sums surviving native usage records: one non-usage event was truncated during local disk exhaustion. The DeepSeek worker ledgers completed beforehand and match runtime totals.

Recorded rounds and checks

DeepSeek High: Round 1 (12m 17s) · Round 2 (14m 08s) · Round 3 (20m 49s)

DeepSeek Maximum: Round 1 (7m 03s) · Round 2 (9m 55s) · Round 3 (12m 13s)

Times above are cumulative to each saved capture’s completion. High overwrote captures within round labels after further corrections, so these are saved rounds, not a strict count of visual iterations.

Both deployed final bundles pass all six original movement/orbit/zoom/animation checks and the additional mobile touch/layout checks, with no JavaScript page errors. High passes exact reset equality; Maximum fails it (camera drift 4.92 scene units). Passing controls does not establish visual fidelity.

One sample per setting, with shared host load and software rendering. High repeated visual corrections within round labels, so the requested three-round limit was not a strict count of iterations. Both workers omitted earlier source snapshots; exact successful edits were replayed and rebuilt, and both final bundles were verified byte-for-byte against the workers’ originals. No root scene edits or visual direction were supplied. Earlier Astra ran separately and its timing did not include equivalent reset/mobile checks. See the protocol for development-cache and tool-access limits.