Doubling Mono’s Float Speed
tirania.org
tirania.org
Seems like it could make an awesome shader-free animation if you translated the entire worlds position by a ridiculously large (increasing in value) float.
I cant even imagine programming at FB and amazon scales.
(see https://developer.twitter.com/en/docs/basics/twitter-ids)
They can in a few edge cases, but in most cases they don’t.
Simplifying many things, to render a model, a game engine uploads 2 things to GPU:
1. Model’s vertex buffer + index buffer. The vertices are in mesh’s own coordinate system, and most 3D designers don’t design their meshes placed 100km away from origin.
2. A single 4x4 matrix, containing ( world * view * projection ) transform. World transforms the model from local to world coordinate system, view from world to camera related, and projection from camera related to 2D screen coordinates + depth.
If you’re 200km far from origin, and are looking at a model near the camera, world transform will contain large values because you’re very far, view transform will also contain large values because camera’s also very far, but multiplied together they won’t have very large values, because model is near the camera.
And if your model is very far from the camera, so the ( world * view * projection ) transform contains huge values, you won’t notice precision degradation, because the whole model will occupy a single pixel at most.
Also if you’ll do nothing and ignore precision issues, numerical errors made while calculating the WVP matrix won’t cause such voxelization of models. A model can be slightly misplaced, maybe jittery between frames, but the shape will stay fine.
This is a good point; this problem is more prone to show up in a ray tracer than a rasterizer, since rasterizers have to apply the camera transform to the geometry, and ray tracers don't.
It's pretty easy to see this problem while using Maya though. Z-buffer resolution in the editor drops off from the origin.
We might see this issue crop up with increasing frequency as more and more people use GPUs for ray tracing...
[1] https://www.gamasutra.com/view/feature/131393/a_realtime_pro...
Unless the field of view is sufficiently narrow?
Good catch.
If the projection matrix is extremely non uniform (e.g. orthographic with very large Z size and very small XY size, or perspective one with FOV angle very close to zero), the described problem can still be encountered.
But I did mentioned “a few edge cases”, this is one of them.
It's cool that the pbrt renders hold up in voxel form, like the model's still solid and the shadows don't freak out or anything.
The problem is fairly well known in film & games production. Artists, especially world designers, all know to model things near the origin and not far away because precision drops as you move away. They will also sometimes avoid modeling small things in small units like millimeters even though they might prefer it, because the units dictate how big your floats get, which in turn determines how fast you lose precision.
Here's the voxel prediction chart: https://en.m.wikipedia.org/wiki/IEEE_754#/media/File%3AIEEE7...
Scaling model from meters to mm only changes precision in model’s units, but the precision in real millimeters stays the same, so the units don’t matter.
I think they avoid millimeters because down the art pipeline people who’ll use the models expect meters, e.g. to place the stuff in a metric world without extra scale involved.
The units do matter though, since the absolute float value is what determines your precision.
The author of this article translated by 200k units, so the precision loss is relative in this case. But artists might need to translate something 100 meters, so their precision loss depends completely on their choice of units.
If an artist needs to translate a model 100 meters off center, and the model is in meters, the graph you’ve linked says the precision will be 10^(-5) model’s units, which is 0.01mm. If the same model is in mm, the graph says the precision for 100*1000 will be 10^(-2) model’s units, which is the same 0.01mm.
As you see, it's independent on the choice of units.
Edit: this translates to 320m if you want 0.01mm precision.
The number's the roughly the same even if you'll model in km or imperial units.
Certainly 64-bit fixed point would be "sufficient".
Actually, Minecraft used to have some interesting float-precision-induced artifacts once you got far from the origin:
My understanding of what causes the Far Lands isn't that great, but I think it's caused by one of six or eight shorts overflowing in the generation algorithm.
Edit: It seems I needed to read more of the Minecraft Wiki link for the Farlands. https://minecraft.gamepedia.com/Far_Lands#Cause covers it much better than I could.
If any one is interested the source code is available here: https://gitlab.com/youngwerth/ray-tracer/settings/repository
Compiled using "swift build -c release".
With this sort of speed difference my guess is that there is either excess memory allocation going on, pointer hopping when you think you are using something by value on the stack, or both.
My guess is that it is either the program being written slightly differently or that swift is doing some sort of indirection under the hood.
Isn't the best to find out to look at the assembly? How else can you know what the compiler and optimizer are doing?
If swift is creating some variables on the heap and/or creating virtual tables for inheritance (I don't know much about it) then you don't need to look at the assembly to know that you are creating indirection or doing too many heap allocations.
However I did not really care about performance here (just a toy project), so I haven't really delved into that side of it.
EDIT: I get a 404 on your git link
Here's the correct one: https://gitlab.com/youngwerth/ray-tracer/tree/performance
Decimal in .NET is 128bit, range (-7.9 x 10^28 to 7.9 x 10^28) / (10^0 to 10^28), 28-29 significant digits
If you are trying to write a super-fast raytracer or game or some such, you will get the best results by writing SIMD manually. Otherwise, you are benchmarking how well autovectorization works in the various compilers and JITs you are testing.
Now, specifically here it looks like Mono had a bunch of other problems they had to fix to get to the right ballpark (not using the right data type, etc.), which is what the blogpost focuses on. And it's nice to see speedups for C# code there.
Still, though, if you need maximal perf, raw SIMD is necessary. Comparing to a C++ version with SIMD might have been interesting, for example. (Likely the reason Burst is "faster than C++" is that it happens to autovectorize that code better.)
If you read the linked blog about the ray tracer, the C# ports are just blog 3 in a series of 8 that does actually conclude with two posts on SIMD - https://aras-p.info/blog/2018/03/28/Daily-Pathtracer-Part-0-...
> In Mono, decades ago, we made the mistake of performing all 32-bit float computations as 64-bit floats while still storing the data in 32-bit locations. (...) Applications did pay a heavier price for the extra computation time, but [in the 2003 era] Mono was mostly used for Linux desktop application, serving HTTP pages and some server processes, so floating point performance was never an issue we faced day to day. (...) Nowadays, Games, 3D applications image processing, VR, AR and machine learning have made floating point operations a more common data type in modern applications. When it rains, it pours, and this is no exception. Floats are no longer your friendly data type that you sprinkle in a few places in your code, here and there. They come in an avalanche and there is no place to hide. There are so many of them, and they won’t stop coming at you.
The raytracer is just a good performance test.
First of all, this is Burden of Proof fallacy: the onus is on you to prove this statement right, not on us to prove you wrong.
Second of all, nobody has been trying to prove you wrong because you did not actually say that floating point performance does not matter in real world code. You may have had it in mind, but you cannot blame others for not picking up on something you did not communicate in the first place.
What you did say was "correctness > speed", which is not the same thing. Furthermore, while this statement is true it needs a context to be applied to, which you have to give. Without further justification by you why using float32 operations for float32 data types would reduce correctness, it is a hollow truism.
This is the same reasoning that tanked the Cyrix 6x86.
If it were true, why don't we just ditch hardware floating point altogether and just emulate it with integer arithmetic instead? I'm sure chip manufacturers would appreciate having the die-space back.
The article does say that "it was a real application", which is a bit of a stretch.