A Shader Trick
the-witness.net
the-witness.net
At first glance The Witness is a puzzle game, made annoying by the fact that you have to walk everywhere instead of clicking "Next Puzzle". But then you occasionally see things in the environment that are just really pleasing. An example that springs to mind is a bunch of broken metal in a window[0], and some branches on a tree, that line up with the sun (which doesn't move) to cast a shadow of a woman sitting underneath a tree[1].
And there's loads of this stuff. And noticing it doesn't contribute anything at all to the apparent objective of the game (except where it does!), but it adds so much that you just wouldn't get if it were the simple puzzle game that it initially appears to be.
I really like this game.
If you're keen to this then you will find it surprising how few games work this way.
Games are part art, what's more to understand?
There's some fun image out there I can't exactly recall, but it's basically of an interview with some game devs, and one of them saying something like "We wanted to have a pot with flowers on the table here, and when you come back to this room later, the color of the flowers changes. No gameplay consequences. But we cut it for time. Why did we want it? Because we thought it'd be cool."
But in a lot of games there's tons of stuff like that which doesn't get cut. Gamedev is full of those "because it'd be cool" or other vague artistic reasons (as opposed to business reasons or researched game testing reasons) for something to be there, or something to be polished well beyond reason.
[1]https://minecraft.fandom.com/wiki/Bedrock_Edition_distance_e...
[2]https://minecraft.fandom.com/wiki/Java_Edition_distance_effe...
Note that one Youtuber KurtJMac has been walking to the Farlands since 2011. He's been raising money for charity as part of those videos/streams. He's roughly 40% of the way there.
Some game engines solve the jittery rendering issues by moving the world relative to the player/camera, rather than the opposite (though it may be more intuitive). This way all your shaders work in nice accurate low floats relative to the camera at the origin, no matter the player location in the world.
But to implement worldgen with relative coordinates would be much more complex.
Here's how the terrain normally looks, zoomed out with the grid on:
https://rezmason.github.io/excel_97_egg/?o=bq
And here's how it looks 8589990 units from the origin, with sanitizing off:
https://rezmason.github.io/excel_97_egg/?o=bq&sanitizePositi...
There's two weird phenomena happening there: the lack of camera position precision causes the terrain to shift left and right, but a little bit further, you can see that all the terrain quads collapse in one dimension for some reason.
Someone also made a game out of this, called Floating Point Leviathan:
For embedded or low resource computing, sin/cos may be expensive, so I use a table based fixed point version. I pick the table to have size power of 2, making lots of things easier. Then to make time wrap, I use a large power of 2, which is exactly the same as this trick, with base 10 replaced with base 2 (and using fixed point math).
You also hit problems where delta times can go negative, so those also need to be max time aware. In short, I always make a timing module, it tracks time (and stretches it as needed), and doles out a few things used everywhere: a delta frame time, a large time (say 64 bits as ns for 584 year wraparound), and a capped time (say 16 or 24 bits) to use in places where you know the wrap amount and still give space for computations not to overflow.
As far as I know the Hypnocube never repeats nor does anything flicker at any time due to bad wraparounds. But that took work to ensure.
A Rust trait with an associated const would help with this:
trait TimeDependentFn {
const PERIOD: f32;
type Output;
fn call(&self, dt: f32) -> Self::Output;
}so you're stuck doing it from scratch every pixel. that's fine, shaders are fast.
at most it might make sense to calculate a truncated time globally per frame and provide that as a uniform.
Then instead of time deltas, we could work with fractions of the animation period.
When switching from float to fraction, we can remove code that handles float's edge cases: +inf, -inf, NaN, -0.0, non-canonical encodings, and loss of precision.
Why limit one’s self to absolutely having to describe it in a single integer value? Why not some wrapper around a handful of integer values that can handle a much bigger max?
Would this have a meaningful performance cost?
Double precision operations are much slower on GPUs. This can get very bad indeed for certain optimizations, like LUTs. 32 bit integers can accumulate much more time delta than floats without precision errors, but have similar problems.
You can pass in an integer and convert it to a float, but that doesn't really solve any problems. The accumulated time is being used in functions that noticeably change over a dozen milliseconds. The total accumulation is simply too large to for floats to represent with that precision; you would also need to convert most of the math surrounding the time uniform.
It is a much better solution to limit the time uniform. The periodic functions depending on it are sensitive to millisecond changes and loop every 100-10000 milliseconds; there's no reason for time to ever be much larger than 10000.
For example, take a look at this looping technique: https://blender.stackexchange.com/a/195316/60486
Not only it isn't trivial, but also it requires a 4D texture. You may have a nice effect going on that uses a texture with less dimensions and looping it will be ever harder... Or ugly.
Just convert the integer into a float before passing it into the shader. For periodic effects, apply the appropriate modulo. Fog doesn't change very rapidly, so if wrapping is a pain, you could just accept the loss in precision. You could round the value as part of the conversion so that the precision doesn't change over the range. With 1 second precision, you should be good for a few months with a 32 bit float.
That solves no problems at all. The number needs to stay an integer until after it is fed to a periodic function, which will restore it to a small enough number to be precisely represented as a float.
> Fog doesn't change very rapidly, so if wrapping is a pain, you could just accept the loss in precision. You could round the value as part of the conversion so that the precision doesn't change over the range.
That doesn't help either. Unless you round the time uniform CPU side- sending a counter that is incremented once per second to the fog shader, and different counter incremented once per millisecond to the shimmer shader- you're still sending a giant number to a function that is periodic over a 1,000,000x smaller window. Precision errors will still cause wildly varying outputs.
Sending separate uniforms only solves the problem for very, very slowly varying functions.
> With 1 second precision, you should be good for a few months with a 32 bit float.
Human vision is exceptionally well-tuned for noticing sudden changes, even relatively subtle ones. A gradual change over 1 second can be hundreds of times larger than a sudden change before it's noticeable.
``` float periodic_animation_sec = (now_ns % periodic_animation_period_ns) * 1e-9f; float fog_sec = now_ns / 1000000000; // or use a power of 2 ```
Regarding the acceptable level of approximation for human perception, I think my point still stands. My assumption is that the frequency content of fog is low enough that pixel colors won't change appreciably over the course of a second. Want 30Hz? That will give you a few days before the precision degrades. 10Hz? About a week or so. Or, find a different solution, like using the CPU to reset the state of the shader every so often. Figure out how to do that seamlessly, or just make sure there's an in-game sunny day every few hours.
If you aren't able to constrain the design or make assumptions about frequencies, then it seems like you have to instead parameterize time as some sort of tuple so that you have more bits, such as the sum of a course absolute and a fine offset, and write the shader to cope with it. That's how time APIs in many operating systems work, where there's integer seconds and integer micro or nanoseconds.
2. It's not exactly cheap, and I don't think compilers put much effort into making sure you actually retain that precision. I have only really used floats in shaders though, and don't know what would happen.
3. I'm pretty sure that in practice you'd lose out on a lot of hardware-accelerated functions, doing trig and interpolations with multiple messy conversions. It's also possible you'd fuck up some compiler optimizations.
It is correct that precision is very important here, but, a millisecond is way too coarse: at 120fps, a millisecond is 1/8 of the frame time, and you'd get horrible jitter.
10 Hz is already a slow strobe light; 1 ms deltas means that each flash is within ~1% of the correct color.
If you're doing something like raymarching in the pixel shader, then you might want sub-millisecond resolution. In 99% of normal shaders, I don't think so. That kind of precision comes into play more with moving objects, where a tiny time delta can mean the difference between a pixel being completely lit or completely dark. Even then though, bad time resolution is just as likely to manifest as motion blur or something.
If your formula involves Π, multiplying it by an integer will produce a float, with its precision problems.
Though you could define pi as 31415 integer... Though if Hwillis is right and originally an integer would overflow after ~50 days, now, being 10000 larger, it would overflow after ~7 minutes.
And finally, if you use sine or cosine, which take radians as input, any whole number passed to them (which probably are at that point converted to floats, but let's assume they aren't) will be a multiple of 57.2957795° expressed in degrees. Almost a sixth of full rotation is way too big of a step for any smooth transition.
Out of curiosity I decided to check if multiplying 1 radian could result with a very big, but visually (due to wrapping around 360°) only a little step:
print(min((180/pi*i % 360, i) for i in range(1,100000)))
Apparently 19 radians is ~1088.62° (mod 360° =~ 8.62°)44 radians is ~2521.01° (mod 360° =~ 1.01°)
377 radians is ~21600.509° (mod 360° =~ 0.509°)
710 radians is ~40680.00345° (mod 360° =~ 0.00345°)
The last result is surprisingly good, but isn't it a spoonful of honey in a barrel of tar? :)
BTW, changing min to max in the Python script will also give useful results (close to 360 rather than close to 0). Worse results for low multipliers but better results near the end of the range.
seeing it delivered as a clever trick in a blog post makes me wonder if i'm more competent than i thought, or if everyone else is generally less competent than i thought.
> How do we ensure that, easily, in a way that people don't have to think about?
The relation is that the time should switch from 10^x to 0, and in the shader the number of digits after a dot should be no more than x.
This is probably true for everyone.