Mach Engine: The Future of Graphics (With Zig)
devlog.hexops.com
devlog.hexops.com
* glm -> zig has native vector types with SIMD support and I think in the long-term these could be much nicer.
* bgfx is truly awesome, but I'm keen to is WebGPU for the reasons outlined in the post.
* imgui - I do think it's the best thing out there currently, but I have something up my sleeve for a much better UI library built around vector graphics and have done a lot of research in this area.
1. GPUs are actually _ridiculously and stupidly_ excellent at natively rendering Bézier curves if you "limit" yourself to quadratic curves.
2. Most of the reasons designers / end users hate working with Bézier curves is because of the absolutely terrible control point manipulation that cubic curves require.
3. GPUs cannot natively render cubic curves (look at any of the papers on this and you'll find they do CPU-side translation of cubic -> quadratic before sending to the GPU in the best case) which harms perf substantially and makes animation of curves expensive. Plus handling the intersection rules is incredibly complex.
...so maybe using the 'mathematically pure' cubic curves is not actually so great after all. :) if only we had tools and formats (not SVG) that worked on isolated quadratic curves (without care for intersection) - then GPU rendering implementations wouldn't need to be absolutely horridly complex (nvidia et. al), or terrible approximations (valve SDF), we could literally send a vector graphic model directly to the GPU and have it render (and animate!) at incredible speeds, with easier to understand tools.
At least - that's the idea.
Eliminating cubics eliminates some trickier problems, but I don't think it fully solves the issues.
Basically just build your graphics around non-overlapping convex and concave quadratic curves represented as triangles. The trick is in building tooling that makes that a pleasant experience (and do any required reductions of more complex shapes -> isolated quadratic curves at tool time, not at render time.)
Existing approaches try to fit what we have (SVGs, complex polygons) onto what GPUs can do. This approach tries to take what GPUs _can do_ and make that competitive with what we have (SVGs, etc.)
To me, that suggests the main distinction between source and output formats comes from how/whether they go about decimating the "erasure" data, since the final result produces much more complex non-overlapped paths. The attractiveness of the Bezier usually isn't in the initial definition of a shape - it just happens to be great at reproduction of multiple curves cut and spliced together. So it has the feel of a premature optimization that probably should be reconsidered in making more ergonomic tooling.
Loop-Blinn natively renders cubic curves in the same way as quadratic curves (implicitization), with some preprocessing to classify cubic curves based on number of inflection points (but this is not the same as converting to quadratic). But I don't think the complexity is really worth it when you can just approximate cubic Béziers with quadratic Béziers.
> we could literally send a vector graphic model directly to the GPU and have it render (and animate!) at incredible speeds, with easier to understand tools.
You can still submit high-level display lists to the GPU with an all-quadratic pipeline. Just add a compute shader step that runs before your pipeline that converts your high-level vector display lists to whatever primitives are easiest for you to render.
You may want to look at Raph Levien's piet-gpu, which implements a lot of this idea. Google's Spinel is also a good resource, though the code is impenetrable.
> The overlap removal and triangulation is a one-time preprocess that takes place on the CPU
All I'm suggesting is merely to ditch that process, ditch the idea that you can have intersecting lines, and instead design tools around convex and concave isolated triangles[0] (quadratic curves) so you end up composing shapes like the ones from the Loop-Blinn paper[1] directly, with your tools aiding you in composing them rather than more complex shapes (cubic curves, intersecting segments, etc.) being in the actual model one needs to render.
I don't think any of this is truly novel from a research POV or anything, it's just a different approach that I think can in practice lead to far better real-world application. It's considering the whole picture from designer -> application instead of the more limited (and much harder) "start with SVGs, how do we render them on GPUs?" question.
It's been a year since I looked at the Loop-Blinn paper, I had forgotten they did this. That's not a good excuse, I should've shut up or spent the time to actually look it up before commenting :)
Regardless of what they do there, hopefully my general idea comes across clearly from my comment above: create tools that allow users to directly compose/animate convex/concave quadratic curves (which render trivially on a GPU as triangles with a fragment shader), and leave dealing with overlaps to the user + tools.
How do you plan on doing text rendering? I feel like that’s often a challenge with cross-platform UI frameworks. Electron/Chromium does this well. Often QT apps don’t handle pixel ratio correctly, or the text looks non-native.
What about differentials? How do you get a change of normal, how do you get a curvature tensor? Without a continuous 2nd derivative?
Edit: In e.g. product design like e.g. cars, people want curvature continuity, i.e. C2 at minimum for their surfaces, because on machined surfaces, you see any curvature discontinuities as edges in reflected light... that's why degree 3 is kind of an industry standard.
For example, if you had 100 PointA objects and 100 PointB objects and you wanted to compute the midpoint of all of them in an output array, you would use SIMD vectors to represent, e.g. 10 X values at a time, and 10 Y values at a time, and then do the math on that, rather than to represent each Point as a separate vector.
But that's just something I heard, of course real world experience would be interesting to hear about.
https://deplinenoise.files.wordpress.com/2015/03/gdc2015_afr...
This is why Pathfinder uses horizontal SIMD, incidentally: most of the heavy number crunching is in super branchy code (e.g. clipping, or recursive de Casteljau subdivision) where there just isn't enough batch data available to make good use of vertical SIMD. Compute shader is going to be faster for any heavy-duty batch processing that is non-branchy. The game code I've written is similar.
I guess I can imagine things like physics simulations where you could have a lot of batch data processing on CPU, but even then, running that on GPU is attractive…
On the newest Intel cpus, a vector double precision multiply is 4 cycles. A horizontal add is 6 cycles. There is a single instruction dot product, which is 9 cycles.
Something like a dot product for fp can only ever be 1 cycle in a cpu that has a very low clockspeed. (Or is a barrel processor or something.) There is simply too much to do.
Anyways I checked it, it's DPPS, and it is a 25 clock cycle latency operation, with a throughput that yields an amortized 6 clock cycle cost.
So, at least in this particular application, horizontal SIMD didn't make an enormous difference, but a 10% boost is nice.
https://github.com/floooh/sokol-zig
This is quite a bit slimmer than using one of the WebGPU libraries like wgpu-rs or Dawn, because sokol-gfx doesn't need an integrated shader cross-compiler (instead translation from a common shader source to the backend-specific shader formats happens offline).
Eventually I'd also like to support the Android, iOS and WASM backends from Zig (currently this only works from C/C++, for instance here are the samples compiled to WASM: https://floooh.github.io/sokol-html5/)
And finally, I'd also like to become completely independent from system SDKs for easier cross-compilation, but with a different approach: by automatically extracting and directly including only the required declarations right into the source files instead of bundling them into separate git repositories (similar to how GL loaders for Windows work).
you might also be interested to know about stub dynamic libraries, e.g. `.tbd` files on MacOS. That's what is in the Mach SDK repos primarily[0], though I'm sure you could slim them down some more with a concrete list of symbols you depend on.
[0] https://github.com/hexops/sdk-macos-11.3/blob/main/root/Syst...
Does it implement an interface like Raylib? Is it a scene graph? How about an entity component system? Will game logic and graphical representation be tightly bound or kept separate?
None of these things are necessarily good, bad, required or forbidden, but they're fairly important to know before you dive in.
In general, I have a preference of providing (1) solid groundwork, (2) libraries, and (3) extensive tooling. We're only at phase 1/3 :)
I don't want to _force_ one way of interacting with this, either. You're free to take the work I've done on making GLFW easy to use and build on top of that (several people are), same with the upcoming WebGPU work.
But yes, there will eventually be a graphical editor, component system, and "obvious" ways to structure your code if that's not something you want to choose for yourself.
It's great to see how the web is giving back to the native platforms.
Is that confirmed? I cant search anything official and Windows 11 was suppose to be the moment they could do DX13 ( DX12 Ultimate was a naming madness ). But they didn't.