Fable 5.1 World Modeling
github.com
github.com
Few things:
1. Opus 5 is just as good for this tbh and cheaper.
2. They don't generate optimized 3d models. They have high poly count for simple geometries.
3. A better approach I've utilized for game ready assets is to use the model to create low poly silhouettes of the 3d model and then bake textures that include a lot more details like windows, doors etc with tools like Meshy. Can post a tutorial if there's interest.
The models need some more RL to be able to do this autonomously.
There is.
I posted this previously and people wanted more info so here is how I built it, with a deeper dive into some film scene recreation I worked on focusing on a scene from Apocalypto:
https://banagale.com/cinematic-canvas-ai-film-animation.htm
I’d love to improve the look of these things.
The car started from a low-poly DeLorean glTF. I had the model rebuild the geometry into inline Canvas code rather than just dropping a model into three.js.
The flight itself is procedural: position/rotation, hover, pitch, wobble, flames, camera, etc. are all driven by code and can be scrubbed/debugged.
That link has a link to an interactive page where you can adjust the car’s path and when it reaches 88mph if you want to see.
It is not mobile friendly atm, though.
I would like to apply textures to the people and scenery in the gauntlet scene. Just haven’t had the time or tokens, as getting motion and pathing and camera angle right seemed the mvp.
It depends on the fidelity, there's a threshold slightly over "low poly" that Opus cannot get over. Once you get into creating foliage l-systems, or physical mob animation, or a house with realistic appliances, Opus is under the convergence threshold for a world model no matter how much time you give it, it will flail around and say it's done even though it's nowhere close to plausible. Fable takes forever but knows how to zoom in and out on the abstractions.
> They don't generate optimized 3d models. They have high poly count for simple geometries.
By default LLMs will do the quick prototype thing and slap some primitives into a THREE.js scene. Which is probably correct because most users don't know or care. But if prompted Fable will connect the manifolds/sculpt with a mesher, instance out the pieces, etc. and make you an efficient art pipeline. You can make it optimized. Just not in 10 minutes.
> A better approach I've utilized for game ready assets is to use the model to create low poly silhouettes
You can also let Meshy make the geo and then fit it back to the world with PCA + BB's. Diffusion models are still 10x better at making physically plausible geometries than LLMs, plus they are faster all things considered. One technique I've found that works well is socketing; let the LLM generate the high level structure with blockouts for sockets, then slot in the higher poly objects that fit. Which is closer to the professional approach for world design.
Can you expand on this?
I've tried several AI 3d generators that claim low-poly output, and they're all far off. None of them decide what belongs in the mesh vs. what should be baked in with the texture, so the poly budget lands in the wrong places.
This kind of decomposition helps to get the art composed the way you want it, though it won't help with a runtime poly budget.
But if you are intentionally doing low-poly style, it is doubtful you are (at least technologically) limited by having too many vertices on assets; you're limited by draws and submits and rendering architecture. You can push millions of animated polys with bells and whistles, postprocessing, physics etc, in the browser, at 90 FPS, on a macbook. You can have thousands of objects, but you have to cull/instance/share materials/atlas/etc, you can't just do `new THREE.Mesh` for each object -- which is what the LLMs will naively do by default unless you ask otherwise.
If resource size is a problem, meshes compress/quantize well with things like meshopt/DRACO.
I read it as "play, warp, act"
Part of it is probably struggle to understand the DXF and mixing up "wall lines" with "dimension/helper lines" (I tried also feeding PDFs)
Anyways, if someone has an idea how to improve 3D modelling with AI or how to make understand 2D drawings to understand rooms/walls/square meters, etc - would be nice.
I'd love to see a sample .fbx as a wireframe, because as far as I've tried that model and basically any other model under the roof that runs on less than 96GB VRAM and can spit out 3D, have had absolutely horrible outputs that are barely usable for anything else than lightweight prototype renders. I'd love to be proven wrong though, would help a ton with building higher fidelity prototypes of various things.
Fake edit: there is a longer video here 1 min long: https://github.com/PhiloLabs/fable51-worlds/blob/main/union-...
I would be especially curious to see the NPC person/car logic and if they're on rails or what, that's a pretty good NPC density for a demo.
Here ya go: - https://github.com/PhiloLabs/fable51-worlds/blob/main/union-... - https://github.com/PhiloLabs/fable51-worlds/blob/main/union-...
https://github.com/PhiloLabs/fable51-worlds/blob/main/union-...
https://github.com/PhiloLabs/fable51-worlds/blob/main/union-...
Uncharted 4: https://youtu.be/NedDxIGQVs0&t=147
Coco: https://youtu.be/nl_JkjgHfFU&t=10
Up: https://www.youtube.com/watch?v=JHfLGgOs6gY
And, eventually you can just run it through Nvidia DLSS 7 for better hair physics and textures. ;)
I wonder how high fidelity we can get these views using just ThreeJS and Fable5.1 iteration cycles.
Here is my Palisades Tahoe world that I made in a similar way: https://ski-red-dog-face.vercel.app/
There's a link to the presentation at SIGGRAPH - https://youtu.be/N2_lb77gKQ8.
Pretty damn compelling.
So, I wonder when is Anthropic (or OpenAI (or... NAME_IT)) going to release one single stable super-working library that does... just anything, that we can then reuse the way these guys demonstrate somebody model's caps, while actually standing on the shoulder of giants.
Because they do stand on the shoulder of gigantic work done by Three.js team. Same goes for demos based on D3, imgui, etc.
Classical algorithmic approaches for data to CAD are quite unreliable heck commercial CAD software still struggles turning scanned 2D plans into sketches with any degree of reliability.
Something of this scale would normally require 1000’s and 1000’s of man hours in manual modeling, placement and quality control.
Some of the latest prompt to CAD agents and models I’ve seen are the most impressive use of AI I’ve seen in a long time.
At least for me this is rather exiting since it gives me the ability to turn many more ideas into reality as a hobbyist rather than spending the entire weekend in Fusion 360.
Begs the question, an earnest one: what software would be above that bar?
Opus has been great at building the game engine but it does struggle with world building for me.
Any tips welcomed!
check source code + more worlds soon
> Geometry is derived from OpenStreetMap (ODbL) and USGS 3DEP (public domain).
This is probably how I would have tackled this, saves on time, especially if Fable just writes code to convert OSM data to a reasonable ThreeJS model.
~2 hour (extensive subagents usage), total ~8M tokens, ~$33 under API
I am your biggest hater.
But god dammit if I’m not also your biggest respecter.
reaches out for a handshake
You guys are onto something with this one.
Keep going.
You crazy bastards!
You crazy fuckin bastards you hear me?! Hahahahah
WOO!
WOO!
Or where can I explore the demo app? It seems like it's just a static page!
All you need to do is create the gh-pages branch! And then you have it at https://philolabs.github.io/fable51-worlds/kyoto-higashiyama
It’s not enough to call it just an image model though. Not quite video either because it’s so specific to rendering images of the world.
Thus, “world model”?
The problem I have with “world” is it should imply so much more depth. To me this is actually a “POV image model” or “First-person perspective model” (maybe we call this an FPM).
There could also be TPMs (third-person models).
This maps better to modern game and 3D dev phrasing, where “world model” might imply 3D object collections spatially organized in a cohesive file format - complete with interactivity, audio, and all the baseline constituents of what one might call a “world”.
FPMs and TPMs, think about it. Reserve POVMs for photographic realism.
I guess “world” works if interactivity and sound etc. are considered other modalities.
Video kinda has the same issue (for me) where the audio part of the video is not part of the model.
When I think “world” I think more than just the visuals. Like a “film” model to me would imply sound and more than just the video.
We typically call 3D art “models” too, so I guess it can’t be called a model model, idk the word “world” doesn’t seem right to me