.NET can generate the code, but your game won't pass certification if it runs machine code generated by something other than the officially provided C/C++ tool chain.
1,079 karma · joined August 1, 2022
.NET can generate the code, but your game won't pass certification if it runs machine code generated by something other than the officially provided C/C++ tool chain.
True, but for the problems that Git LFS solves P4 is usually the other option. P4 handles large binaries very well, caveat being the rest of P4. In some projects the caveats are worth it. I work in video games and our in house engine has an unfiltered checkout size over a 1TB.
I've found LFS very brittle when bouncing between branches and moving around in history. There's been many times I've had my local check out blow up through rebases or mistakes I've made. Many times the recovery is to nuke and re-clone. With a 1TB repo that's not a good option.
It's been a while since I've used it, and I'm sure some of my problems were skill issues, but I never found it worth the pain at any scale I've tried to use it.
I don't really see what the article is talking about as the future, but if P4 or Git LFS is the best we can do as a species then we're doomed. All VCS options suck for one reason or another, I hope we don't stop trying to make something better. If only to save me from perforce.
Separating the hot path into a prefix before calling into a separate cold function should still generate better code. Your prefix only needs to allocate registers and stack space for just that single path. You would only pay the spilling costs in the old code off the hot path rather than every instruction. And I would expect the branch prediction accuracy of that prefix check to be higher than having the hot and cold paths all dispatching through the same tree of branches.
However it's speculation until you measure so I could be wrong.
Good article in any case, I enjoyed reading along.
The old, enum based value type used a single big match statement to dispatch between all possible type combinations. Their assembler output looks like the match gets compiled to something like a big stack of nested if statements.
The new code uses an explicit fast path check with a dispatch into a tagged 'cold' path when the common case isn't hit. The generated code is a single upfront branch for the fast path that exits immediately, with a dispatch into the slow path in a separate function.
This would be contributing significantly to the performance improvements. The old path requires taking several branches even on the hot path. The new code has a single, highly predictable branch that skips all the messy dispatch for the other types.
This could have been implemented for the enum based value type, and I would expect to see a jump in performance there too even without the new compact value type. There will be a much higher branch predictor hit rate with the explicit fast path.
It's depressing how bad the drivers are. They have 0 market pressure to improve the quality, unlike PC where the whiniest group on the internet (gamers on reddit) will crucify you for driver problems. I've spent too much time investigating and implementing workarounds. Though it does pay my bills...
And then Google has the gall to charge basically flagship prices for their under powered junk.
Personally I don't buy flagship phones because I don't use their full compute capabilities anyway, but a Pixel 11 would cost 2x what I paid for my Pixel 8. For 0 GPU performance improvement.
An add instruction on an AMD GPU adds two scalar values. If they're in a VGPR then each thread will add two values unique to that thread. A SIMD ISA as is common on a CPU is different because an add instruction explicitly adds a vector of values. xmm1 stores 128-bits of data. VGPR[1] stores 32-bits of data vectored over 32-64 threads in a thread group.
Without special instructions a thread can't access the VGPR values stored in other threads.
If you want to get more pedantic you also need to look at your target hardware and their specific micro-architectural quirks and features to get the best performance. AMD specifically benefits a lot from exploiting the scalar unit over the vector unit, you save loads of register file space if you can keep data in SGPRs over VGPRs. There's lots of traps you can fall into where you can load data from buffers into SGPRs but they get promoted to VGPRs because the scalar unit lacks an opcode for like one math operation you did to the value somewhere.
While each lane isn't truly a thread because it doesn't have its own PC the programming model definitely tries to make it seem that way. The threads can terminate at different points too. And again, the ISA isn't a vector ISA. Your register values are scalar.
If you go look at AMD's ISA docs (they're public) you'll see you don't have the equivalent of a __mm256 register like on x86. Each 'thread' just deals with single scalar values like int32 of float32. The hardware, however, groups 32 or 64 threads together which all run the same program and runs them together. Each 'thread' loosely maps to a SIMD lane. The SIMD is implicit, not explicit.
The main difference is that the 'SIMD' execution is somewhat opaque to the program. You just write plain scalar code and the hardware model dispatches it efficiently to SIMD execution units. It's not really an abstraction because to extract maximum performance you have to understand how it works. You can use this kind of programming model on a CPU too, Intel did it with [0] ISPC. It's a C-like language that has execution semantics similar to GPU shader languages but compiles to regular CPU code, and maps threads to your CPUs SIMD lanes like a GPU.
Vast majority of titles use deferred rendering, and lighting is done off screen too. Usually the only thing done to the "screen buffer" is a final post-process pass or a copy.
Bots can teach basics, but many of these games rely on abusing game mechanics in ways the developers didn't think of in development so it's impossible to teach. Or things are physically difficult to pull off.
Movement shooters especially have this problem because a core part of the game is abusing the movement system and abilities to out maneuver other players. You can explain the how, but executing it all automatically is a real skill that requires practice. I adore Titanfall 2. I sunk hundreds of hours into it. It's really something else when you get the movement down, it all happens so fast you can't think about it you just have to do it. There's no way to teach that with a tutorial, it's muscle memory and perseverance.
Titanfall 2 has all of these things, campaign, skirmish with bots and PvE, but the skill floor and ceiling is still brutal. Whether it failed or not is a question of how long you expected the game to last, but it still succumbed to the same fate. The more experienced players pulled the ladder up behind them and the matchmaker didn't have enough players to give everyone a fair match.
As long as you can draw 2D triangles there's nothing stopping you from drawing a "3D" scene with a software transform pipeline. Depth sorting gets fun without a z-buffer though.
Exposing the hardware primitives directly is not an abstraction, so in reality all you end up having to do is implement your own Vulkan instead if you want to support more than 1 GPU. And your app won't work on new GPUs without shipping an update.
If you want a common API to target all GPUs you get Vulkan again. GPU hardware is too diverse to go much lower than Vulkan without just moving the hardware abstraction into your app instead.
Apple could've chosen to offer no tools and leave it to the community to build their own. I think that's reasonable. I don't expect Microsoft to bring MSVC to Linux.
But Apple take the explicit position that you must purchase a Mac in order to ship apps for iOS.
I have nothing against the hardware. I despise how artificial the problem is. It's not even like Apple just refuse to provide tools and leave it up to the community. They force you to use a Mac in the terms of service.
For context I work on the mobile team for a large AAA game engine. This is going to colour my opinion because we're very Windows centric, but my experience wouldn't exist if it weren't for Apple. I need to sync code on two different machines and build half the project on Windows and the other half on the Mac. If we could just build and debug the iPhone on Windows none of this mess would be needed.
We don't develop for MacOS, we develop for iOS. Just debugging the app requires several hoops across multiple machines because of Apple's policy here. Just because a base spec Mac Mini is cheap doesn't make this whole problem any less stupid. I don't need a vendor's special-sauce computer to build for consoles. Cross compilation and remote debugging has existed for decades, this problem should not exist.
Calling it more affordable because Mac Minis aren't junk anymore is really just accepting Apple's ridiculous requirement because "I guess it could be worse right?".
Forcing people onto Macs is such a pathetic way of extracting rent from developers and makes workflows worse. No you can't just build and debug your app from your Windows or Linux workstation because we want you to buy hardware you don't need. Sorry we only support Xcode on the latest version of MacOS, how else will your build machine roll out of the OS support window and make you replace a functioning machine?
It's especially nice when you work on projects that aren't iOS/Apple native and you don't get Mac workstations. You end up with this wonderful mess of shared machines and half assed tools because you just need the bare minimum to push a build to a phone.
This really gets to the core of what I think Rust is about, you can add compiler checked constraints to your APIs that your C and C++ code can't. It's up to you to use them effectively. Rust's ability to keep your safe code safe is a measure of the language, but also your architecture. The buck has to stop somewhere for the language to prove safety, Rust lets you decide rather than the language itself.
It happens to also benefit the Linux gaming crowd, but it's still ultimately self-interest driving the work. The engineers doing the work are probably doing it for the altruistic reasons, but ultimately Valve is writing the cheques.
Any sane router also uses a firewall for IPv6. A correctly configured router will deny inbound traffic for both v4 and v6. You are not less secure on IPv6.
- Wavefront: AMD, comes from their hardware naming
- Warp: Nvidia, comes from their hardware naming for largely the same concept
Both of these were implementation detail until Microsoft and Khronos enshrined them in the shader programming model independent of the hardware implementation so you get
- Subgroup: Khronos' name for the abstract model that maps to the hardware
- Wave: Microsoft's name for the same
They all describe mostly the same thing so they all get used and you get the naming mess. Doesn't help that you'll have the API spec use wave/subgroup, but the vendor profilers will use warp/wavefront in the names of their hardware counters.
fwiw: no idea on the other anime adaptions quality
Platform bugs, build issues, distro differences, implicitly relying on behavior of Windows. It's not just "use Linux API", there's a lot of effort to ship properly. Lots of effort for a tiny user base. There's more users now, but proton is probably a better target than native Linux for games.
You can run KDE but depending on the app and containerization you open you'll get a Qt environment, a Qt environment that doesn't respect the system theme, random GTK apps that don't follow the system theme, random GTK apps that only follow a light/dark mode toggle. The GTK apps render their own window decorations too. Sometimes the cursor will change size and theme depending on the window it's on top of.