Local AI experiment
I gave a local agent one hard prompt, “Create voxel pagoda garden scene”, and went about my evening. An hour and 45 minutes later it had built a complete, interactive 3D Japanese garden in a single HTML file. All of it ran on one gaming desktop.
Live demo, running right here. Drag to orbit, scroll to zoom, press D for dusk. Open it full screen for the best view.
What it built
The scene is a 72 by 72 block island of roughly 27,000 voxels, generated from a seed so you can regrow a different garden with one click. It has a multi-roof pagoda on a stone plinth, a koi pond with an open-arch red bridge, a waterfall, cherry, maple, pine and bamboo, stone lanterns, a torii gate and stepping-stone paths. On top of the geometry sit falling petals, swimming koi and a day-to-dusk cycle with a custom sky shader, stars, fireflies and lanterns that light up at night.
It also ships its own controls: Day/Dusk, Petals, Auto-rotate, two camera presets, Regrow and a PNG export. The model was not asked for any of that. It was asked for a scene.
The setup
Hardware
A normal desktop: Core i5-13600K, 64 GB DDR4 and a single RTX 5070 Ti with 16 GB of VRAM. No cloud fallback was needed.
How the time was spent
| Time (EEST) | Phase |
|---|---|
| 22:22 – 22:30 | Project folder, HTML template and a generator script |
| 22:30 – 23:30 | Scene build and iteration: pagoda, garden, trees, water, bridge, waterfall |
| 23:30 – 00:02 | Atmosphere polish: sky shader, sun sprite, fog and dusk color grading |
| 00:02 – 00:06 | Verification passes: pond and bridge close-ups, voxel probes |
| 00:06 – 00:07 | Final polish: open-arch bridge, taller waterfall, horizon haze |
Times are reconstructed from file timestamps, because no start stamp was taken. October 11 to 12, 2026.
The part I find most interesting: it checked its own work
Most of those 180 tool calls were not writing code. They were a loop, repeated over and over:
The checks were concrete. A syntax check, headless generation across three seeds, a live browser load, and programmatic probes of exact voxel coordinates to confirm that the bridge deck, arch, railing, waterfall and koi really exist where they should. The screenshots were reviewed by the vision encoder of the same local model.
The run also crossed the 120K context window once and was compacted. It carried on cleanly because the verification results had been written to files in the workspace and could be re-probed, not just remembered in the conversation. That is a habit worth copying when you work with any long-running agent.
What is not perfect
- The far horizon of the lake is a little too crisp at dusk.
- The lantern lights pool visibly if you zoom in close at night.
- A programmatic probe proves a thing is there, not that it looks right. I would still look at the final scene with my own eyes, and so should you.
- No token counters were captured during the run, so there is no throughput figure here. This is one task on one run, a demonstration and not a benchmark.