I tested 10 model/harness combinations on the same Three.js task
alvins82.github.io
alvins82.github.io
I don't think any of these examples are using things like tone mapping so they're stuck in sRGB (AgX or ACES look much better), they're not using the node materials (good for programmatic texture implementation), and they're not doing anything cool like baking shadow environments or using post-processing effect.
They're nice, but I think they're showing how far behind AI models are on this sort of project rather than how good they are.
The models _can_ do it, but you need to ask for the right things. Most people don't, they'll usually blame the browser for being slow or ugly when they can't break through the THREE demo page wall.
A larger THREE.js project starts to look more and more like a game engine, so you pick and choose the parts you need. There's a ton of open source libs, most of the heavy components the big players use are open source, things like physics, mesh optimization.
Key AI-specific parts might be:
- a harness (so the agent can drive the thing)
- authoring pipeline (so you can bake/optimize assets)
- some sort of coherent renderer architecture (what are your assets, your passes, what's your shader graph).
Without some fundamentals here you are on the short road to falling off the cliff of tech debt and the AI will gladly drive you off of it until you ask for an expensive rewrite.
> would you create models independently
Yes. Pretty much any proven gamedev/asset pipeline is something frontier models are good at. Procedural systems, asset store, free content, Blender, Meshy.
Literally anything except "make a THREE.js scene" -- which is not a technique used in gamedev, beyond throwaway prototypes or demos. Which is what you will get if you ask for a THREE.js scene.
I found the remarkable https://babylonjs.com/lite-demos/
Both the regular and lite demos are pretty amazing for browserware!
There are thousands, if not millions now, of people working on the problem, a vast majority of humanity text output used for training, an enormous infrastructure composed of the most complex human made device (the processor), a gigantic amount of energy running it, for an attempt that spanned decades, if not century, heck if not millennia if you go as far as considering Antikythera.
It might be impressive but it's anything but a miracle when we consider how much effort was poured into this.
Genuinely happy with some of the Qwen 3.8 results (especially since I can run that model at Q8).
Interesting to see how much better (at this task) Pi (OMP) is over Opencode as a harness.
I’d love to see a few more with outcomes that are as easy to judge but less subjective.
I’ve got a toy project going to make a fun to watch battle simulator where an LLM (or two if playing vs) has to write programs that control multiple bots (each with their own line of sight and limited battle context) that have to coordinate and fight alongside each other. Goal is to have the LLM update the code based on current situations maybe 5-10 times in a 5 min simulated battle. Exploring even allow the bots to request new programming and score based on number of reprogram steps.
I suppose the only thing that the prompt asks for is the cinematic view, and honestly they all kinda fail on the “subtle volumetric-style fog planes”, none of them have more fog when you get farther from a light source.
Also your last paragraph sounds like the setup for a late 80s scifi movie...
Ideally also finding somehow (not sure what would be the right away) what is publicly available before running the test. It's quite a different outcome if there are competitions, e.g. js13k, live code examples from books, even templates, on specific that topic. Visually here the results looks very very similar to the point that I can't help but wonder if it's the result from the short yet relatively descriptive prompt or because some template was always found and relied on.
I wanted a powerful GUI+harness setup for open models so I could use/test as they came out.
But I am annoyed at these GUIs implementing features I don’t care about. I want them to just wrap my harness and forward it to my iPhone, but they can’t help themselves from feature creep.
https://github.com/stablyai/orca
I did have some issues getting it installed on a headless server. I sort of gave up and installed the instance that I use as the remote server on a Debian + xfce machine I had laying around.
If you don't have anything working check the console, maybe a WebGL issue.
I have been building something like this with human voting, elo and ranking on multiple scenes and many models.
From all the examples I've seen, Astra does it really well, and I suspect it's because they wanted to attract game designers, so they trained the model more on 3D, animation libraries, etc.
Or, my standards are lower. It's sometimes hard to tell in these discussions whether people are talking about getting production quality results, or stuff that's good enough for a one off blog post.
GLM 5.3 Flash Max had an interesting showing. Its Codex version was bad [0], and it completed in 9 minutes. The OpenCode version was much richer [1] and more detailed, completed in 20 minutes. And the OMP version was arguably the most complete [2], completing in 30 minutes.
This is probably the strongest argument for the effect of a harness, and I'd be interested to learn the differences in the prompts and tools between these three.
[0]: https://alvins82.github.io/hangar-harness-model-tests/hangar...
[1]: https://alvins82.github.io/hangar-harness-model-tests/hangar...
[2]: https://alvins82.github.io/hangar-harness-model-tests/hangar...
Also I think Astra looks the best and has the best functionality. Also shocked how much better GLM is on the Non Codex harnesses. Didn’t think it would make such a difference.
Would be nice if you could include cost in the table
How different are the results between multiple runs of the same setup?
Would love to see Claude and Gemini as well. (And it would be nice to include the total cost in the table.)
Basically it can run these mini programs where each input might be another toolcall, so it can run without waiting for whole LLM response and ready the parameters async.