How to Implement an FPS Counter
vplesko.com
vplesko.com
spf_avg = alpha * cur_spf + (1.0f - alpha) * spf_avg;
For alpha value you can use the formula:
alpha = 2/(n+1)
Which will give smoothing comparable to an n sample moving average. This is the same formula used for n-day exponential moving averages for stocks.
As the article points out, this is a sample based window which is not as good as a time based window, but it's also dead simple to implement.
Edit:
Just spit balling because I haven't thought about this problem in a while and asked AI to give a foruma for EMA with variable duration events, so take it with a grain of salt. Maybe for a time based window you could use a dynamic alpha (forgetting factor) with the following formula:
alpha = 1-e^-(cur_spf/window_secs)
You can also dynamically calculate the filter coefficient based on the current delta time, which makes sure that the smoothing behavior is independent from the current framerate.
So...T seconds have elapsed since your last frame, and you have a new data point you'd like to incorporate into your EMA (and, for this problem, that data point is T itself). You keep e^-(bT) of the old data, 1 minus that of the new data, and you're done. Alpha is indeed 1-e^-(b * cur_spf) for some constant b, just like the AI said.
Which b do you choose though? I usually prefer to think of it in terms of half lives. Let H be your desired half life, and set 1-e^-(bH) equal to 1/2. You get b=log(2)/H. That's similar to the AI answer, but it rescales window_secs into a parameter you can actually start to reason about. The AI answer gives you a 1/e life instead, which is a less comfortable constant to mentally process.
That answer also naturally generalizes to other "time" axes or denominators. Suppose you want to EMA using discrete events rather than wall-clock time. Replace T in your dynamic alpha with the discrete count of events you haven't updated the EMA with yet (e.g., if something happens every time tick you would usually take T=1, but you can increase that based on a few skipped frames or whatever if you'd like).
As a fun fact, you can tailor this sort of stateless solution to have almost any decay property you'd like. Start with your favorite ODE satisfying a few properties, and this equation falls out as the step a discrete solver would take to approximate the ODE.
Wouldn't it feel like 10fps for 0.1s only? I agree it's a good thing to measure, I think it's called "stutter" usually, but I'm not sure you can say "it feels like 10 fps" since its for such a small moment.
FPS based on the median of a moving window is good if you want perceived frame rate, which rejects extreme outliers.
FPS based on the average of a moving window is good if you want statistical mean frame rate.
This.
The other two are not good for anything gaming as any framerate inconsistency breaks the experience. A stable frame rate is much more important than a high one.
Considering stability is the name of the game:
1. If it’s changing quickly in the ones position because it’s not stable, that’s immediately apparent
2. If it’s changing quickly in the significant digits, specifically hundreds or tens position, that’s immediately apparent
Which is largely all the info a user cares about anyways. It gives them the three important states:
1. Stable
2. Mildly unstable
3. Severely unstable
That is, if the number is entirely unreadable with this strategy, it’s because the game is entirely unplayable.
After all that, people can talk about averaging methods, but there's a lot to be done before what this blog is talking about is even available.
The reason solving just-in-time rendering is important is because queue priority is not actually supported by most drivers. Some extensions can give us global priority for the process, not real priority for queues. The right way then to avoid workload A from causing workload B to miss a latch is to put workload A into the idle time that would exist from running B just in time. This is itself a luxury based on the fact that workload B is lightweight enough that its own uncertainty can only rarely exceed the latch deadline.
At least on VRR displays, making B a bit late has much less dire consequences, but driving refresh from the application needs exclusive access to the display, and not all compositors want to provide this.
Please do reach out if it seems like I'm only still catching up. I'm sure someone knows a decent way to get sub-millisecond just-in-time rendering accuracy without watching the phase suddenly double on FRR. Ping https://github.com/positron-solutions/mutate and we can get in touch.
* If you compute the value as the amount of data in last chunk (usually constant except for the very last chunk) divided by time taken to receive that chunk, then it's like "FPS based on the latest frame". This can result in misleading metrics because you only update _after_ the chunk is received so slowdowns are not reported in real time. If your recv size is small, the number may also bounce around too much.
* If you show it as cumulative download / cumulative time, then it's similar to last N frames as N->\infty. This doesn't really tell you what you care about which is the "current" speed.
const auto now = SDL_GetPerformanceCounter();
static auto prior = now;
static const auto frequency = static_cast<double>(SDL_GetPerformanceFrequency());
const auto delta = std::min(static_cast<float>(static_cast<double>(now - prior) / frequency), 1.f / 30.f);
prior = now;
static auto tick = now;
static auto frames = 0;
++frames;
const auto elapsed = static_cast<double>(now - tick) / frequency;
if (elapsed >= 1.0) [[unlikely]] {
const auto fps = frames / elapsed;
const auto memory = lua_gc(L, LUA_GCCOUNT, 0);
std::println("{:.1f} {}KB", fps, memory);
frames = 0;
tick = now;
}But that's just the nerd in me talking. The article is great!
The measured frame duration will have jitter up to 1 or even 2 milliseconds for various 'external reasons' even when your per-frame-work fits comfortably into the vsync-interval each single frame. Using an extremly precise timer doesn't help much unfortunately, it will just very precisely measure the externally introduced jitter which your code has absolutely no control over :)
What you are measuring is basically the time distance between when the operating system decides to schedule your per-frame workload. But OS schedulers (usually) don't know about vsync, and they don't care about being one or two milliseconds late, and this may introduce micro-stutter when naively using a measured frame time directly for 'driving the game logic'.
For instance if the previous frame was a 'long' frame, but the current frame will be 'short' because of scheduling jitter, you'll overshoot and introduce visible micro-stuttering, because the rendered frames will still be displayed at the fixed vsync-interval (I'm conveniently ignoring vsync-off or variable-refresh-rate scenarios).
The measurement jitter may be caused by other reasons too, e.g. on web browsers all time sources have reduced precision since Sprectre/Meltdown, but thankfully the resulting jitter goes both ways and averaging/filtering over enough frames gives you back the exact refresh interval (for instance 8.333 or 16.667 milliseconds even when the time source only has millisecond precision).
On some 3D APIs you can also query the 'presentation timestamp', but so far I only found the timestamp provided by CADisplayLink on macOS and iOS to be completely jitter-free.
I also found an EMA filter (Exponential Moving Average) more useful than a simple sliding window average (which I used before in sokol_app.h). A properly tuned EMA filter reacts quicker and 'less harshly' to frame duration changes (like moving the render window to a display with different refresh rate), it's also has less implementation complexity because it doesn't require a ring buffer of previous frame durations.
TL;DR: proper frame timing for games is a surprisingly complex topic because desktop operating systems are usually not "tuned" for game workloads.
Also see the "classic" blog post about frame timing jitter:
https://medium.com/@alen.ladavac/the-elusive-frame-timing-16...