I gave a local AI one prompt. 1h 45m later it had built a voxel pagoda garden

Local AI experiment

I gave a local agent one hard prompt, “Create voxel pagoda garden scene”, and went about my evening. An hour and 45 minutes later it had built a complete, interactive 3D Japanese garden in a single HTML file. All of it ran on one gaming desktop.

Live demo, running right here. Drag to orbit, scroll to zoom, press D for dusk. Open it full screen for the best view.

1h 45mwall-clock, start to last edit
~180tool calls (estimated)
34.9 KBone self-contained index.html
1context compaction, survived
0console errors on live load

What it built

The scene is a 72 by 72 block island of roughly 27,000 voxels, generated from a seed so you can regrow a different garden with one click. It has a multi-roof pagoda on a stone plinth, a koi pond with an open-arch red bridge, a waterfall, cherry, maple, pine and bamboo, stone lanterns, a torii gate and stepping-stone paths. On top of the geometry sit falling petals, swimming koi and a day-to-dusk cycle with a custom sky shader, stars, fireflies and lanterns that light up at night.

It also ships its own controls: Day/Dusk, Petals, Auto-rotate, two camera presets, Regrow and a PNG export. The model was not asked for any of that. It was asked for a scene.

The setup

Hardware

A normal desktop: Core i5-13600K, 64 GB DDR4 and a single RTX 5070 Ti with 16 GB of VRAM. No cloud fallback was needed.

How the time was spent

Time (EEST)Phase
22:22 – 22:30Project folder, HTML template and a generator script
22:30 – 23:30Scene build and iteration: pagoda, garden, trees, water, bridge, waterfall
23:30 – 00:02Atmosphere polish: sky shader, sun sprite, fog and dusk color grading
00:02 – 00:06Verification passes: pond and bridge close-ups, voxel probes
00:06 – 00:07Final polish: open-arch bridge, taller waterfall, horizon haze

Times are reconstructed from file timestamps, because no start stamp was taken. October 11 to 12, 2026.

The part I find most interesting: it checked its own work

Most of those 180 tool calls were not writing code. They were a loop, repeated over and over:

patch the file→headless test→screenshot→local vision review→patch again

The checks were concrete. A syntax check, headless generation across three seeds, a live browser load, and programmatic probes of exact voxel coordinates to confirm that the bridge deck, arch, railing, waterfall and koi really exist where they should. The screenshots were reviewed by the vision encoder of the same local model.

The run also crossed the 120K context window once and was compacted. It carried on cleanly because the verification results had been written to files in the workspace and could be re-probed, not just remembered in the conversation. That is a habit worth copying when you work with any long-running agent.

What is not perfect

  • The far horizon of the lake is a little too crisp at dusk.
  • The lantern lights pool visibly if you zoom in close at night.
  • A programmatic probe proves a thing is there, not that it looks right. I would still look at the final scene with my own eyes, and so should you.
  • No token counters were captured during the run, so there is no throughput figure here. This is one task on one run, a demonstration and not a benchmark.

Links

Published by

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.