HNHacker News
TopNewBestAskShowJobs

MindSpunk

1,079 karma · joined August 1, 2022

submissionscomments
MindSpunk··on Using any C++ library in Godot
It's not up to .NET AOT, it's up to Xbox and Sony.

.NET can generate the code, but your game won't pass certification if it runs machine code generated by something other than the officially provided C/C++ tool chain.

MindSpunk··on Steam Frame starts at $1059
Australian sticker price includes 10% GST so the correct way to compare is to take the tax free US price, convert that to AUD and add 10%. Applying that process yields $1633 AUD so the sticker price we get is actually better than the US price.
MindSpunk··on Version control second coming
> Git and P4 are not meant to be directly competitive here.

True, but for the problems that Git LFS solves P4 is usually the other option. P4 handles large binaries very well, caveat being the rest of P4. In some projects the caveats are worth it. I work in video games and our in house engine has an unfiltered checkout size over a 1TB.

I've found LFS very brittle when bouncing between branches and moving around in history. There's been many times I've had my local check out blow up through rebases or mistakes I've made. Many times the recovery is to nuke and re-clone. With a 1TB repo that's not a good option.

It's been a while since I've used it, and I'm sure some of my problems were skill issues, but I never found it worth the pain at any scale I've tried to use it.

MindSpunk··on Version control second coming
From my experience Git LFS is extremely brittle and will break your local check out if you so much as breath on it wrong. Perforce is an expensive solution to the problems of LFS, and I've yet to find a workflow as powerful as my git workflows for managing code.

I don't really see what the article is talking about as the future, but if P4 or Git LFS is the best we can do as a species then we're doomed. All VCS options suck for one reason or another, I hope we don't stop trying to make something better. If only to save me from perforce.

MindSpunk··on Replacing a Rust Enum with a 64-Bit Word Made My Interpreter 17% Faster
Absolutely, I agree. I suspect that explicit branch for the hot path is doing a lot too.

Separating the hot path into a prefix before calling into a separate cold function should still generate better code. Your prefix only needs to allocate registers and stack space for just that single path. You would only pay the spilling costs in the old code off the hot path rather than every instruction. And I would expect the branch prediction accuracy of that prefix check to be higher than having the hot and cold paths all dispatching through the same tree of branches.

However it's speculation until you measure so I could be wrong.

Good article in any case, I enjoyed reading along.

MindSpunk··on Replacing a Rust Enum with a 64-Bit Word Made My Interpreter 17% Faster
I'm not convinced the performance benefits are entirely the result of the more compact object representation. It definitely would help, but looking at the code snippets the author provides for the add instruction there's an important structural change that would be making a huge difference.

The old, enum based value type used a single big match statement to dispatch between all possible type combinations. Their assembler output looks like the match gets compiled to something like a big stack of nested if statements.

The new code uses an explicit fast path check with a dispatch into a tagged 'cold' path when the common case isn't hit. The generated code is a single upfront branch for the fast path that exits immediately, with a dispatch into the slow path in a separate function.

This would be contributing significantly to the performance improvements. The old path requires taking several branches even on the hot path. The new code has a single, highly predictable branch that skips all the messy dispatch for the other types.

This could have been implemented for the enum based value type, and I would expect to see a jump in performance there too even without the new compact value type. There will be a much higher branch predictor hit rate with the explicit fast path.

MindSpunk··on Arm Mali G2-Ultra NX GPU: desktop-class mobile gameplay with AI-native graphics
Since when do Samsung ship driver updates? I'm being a bit hyperbolic but from my experience Samsung don't ship them very often either. To Google's credit, the Pixels get them quite frequently. Samsung is still better than most other Android vendors, who are somehow even worse.

It's depressing how bad the drivers are. They have 0 market pressure to improve the quality, unlike PC where the whiniest group on the internet (gamers on reddit) will crucify you for driver problems. I've spent too much time investigating and implementing workarounds. Though it does pay my bills...

MindSpunk··on Arm Mali G2-Ultra NX GPU: desktop-class mobile gameplay with AI-native graphics
The Pixel 10 and 11 GPUs are offensively bad. Not only are the Power VR GPUs slow, and less capable feature wise than Mali and Adreno (Xclipse is still around too!), but the tile Google put in the Pixel 11 is actually an _older_ design than what they put in the Pixel 10. They just clocked it higher to make up the performance difference.

And then Google has the gall to charge basically flagship prices for their under powered junk.

Personally I don't buy flagship phones because I don't use their full compute capabilities anyway, but a Pixel 11 would cost 2x what I paid for my Pixel 8. For 0 GPU performance improvement.

MindSpunk··on An ongoing 3D-printer AGPL violation
They cost 3x as much
MindSpunk··on Rust SIMD on the GPU
A VGPR is not the same thing as a vector register like in SSE4 or AVX. Each addressed register contains a single 32-bit value. A VGPR differs from an SGPR in that each thread in a thread group can have a different value in that register. An SGPR will have a uniform value shared with all threads in a group.

An add instruction on an AMD GPU adds two scalar values. If they're in a VGPR then each thread will add two values unique to that thread. A SIMD ISA as is common on a CPU is different because an add instruction explicitly adds a vector of values. xmm1 stores 128-bits of data. VGPR[1] stores 32-bits of data vectored over 32-64 threads in a thread group.

Without special instructions a thread can't access the VGPR values stored in other threads.

MindSpunk··on Rust SIMD on the GPU
Yes, of course writing naive code assuming each lane in a thread group is a real thread is going to cause problems, but I didn't feel like I needed to go into that level of detail replying to someone just learning about GPU internals. I tried to cover this loosely by mentioning how you need to know how it works for maximum performance.

If you want to get more pedantic you also need to look at your target hardware and their specific micro-architectural quirks and features to get the best performance. AMD specifically benefits a lot from exploiting the scalar unit over the vector unit, you save loads of register file space if you can keep data in SGPRs over VGPRs. There's lots of traps you can fall into where you can load data from buffers into SGPRs but they get promoted to VGPRs because the scalar unit lacks an opcode for like one math operation you did to the value somewhere.

While each lane isn't truly a thread because it doesn't have its own PC the programming model definitely tries to make it seem that way. The threads can terminate at different points too. And again, the ISA isn't a vector ISA. Your register values are scalar.

MindSpunk··on Rust SIMD on the GPU
It's not really obvious unless you go in depth of the details on modern GPU architecture. GPUs aren't really SIMD, they're SIMT (single instruction multiple thread). The silicon looks a lot like SIMD, but the programming model is different.

If you go look at AMD's ISA docs (they're public) you'll see you don't have the equivalent of a __mm256 register like on x86. Each 'thread' just deals with single scalar values like int32 of float32. The hardware, however, groups 32 or 64 threads together which all run the same program and runs them together. Each 'thread' loosely maps to a SIMD lane. The SIMD is implicit, not explicit.

The main difference is that the 'SIMD' execution is somewhat opaque to the program. You just write plain scalar code and the hardware model dispatches it efficiently to SIMD execution units. It's not really an abstraction because to extract maximum performance you have to understand how it works. You can use this kind of programming model on a CPU too, Intel did it with [0] ISPC. It's a C-like language that has execution semantics similar to GPU shader languages but compiles to regular CPU code, and maps threads to your CPUs SIMD lanes like a GPU.

[0] https://ispc.github.io/

MindSpunk··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
Most games haven't rendered directly into the "screen buffer" for 15-20 years.

Vast majority of titles use deferred rendering, and lighting is done off screen too. Usually the only thing done to the "screen buffer" is a final post-process pass or a copy.

MindSpunk··on Quake – 30th Anniversary Update
You can't actually ship a game on the porting toolkit without breaching the license. It's intended by Apple to aid porting to use Mac's native APIs one component at a time, but the license forbids redistribution or commercial use. The community can hack away with it however they want because Apple isn't going to chase every individual person, but a business can't ship it unless they want a nice call with Apple's legal team.
MindSpunk··on Quake – 30th Anniversary Update
That's largely what matchmaking aims to solve, assuming you have enough players you can get a reasonably balanced match to drop people into.

Bots can teach basics, but many of these games rely on abusing game mechanics in ways the developers didn't think of in development so it's impossible to teach. Or things are physically difficult to pull off.

Movement shooters especially have this problem because a core part of the game is abusing the movement system and abilities to out maneuver other players. You can explain the how, but executing it all automatically is a real skill that requires practice. I adore Titanfall 2. I sunk hundreds of hours into it. It's really something else when you get the movement down, it all happens so fast you can't think about it you just have to do it. There's no way to teach that with a tutorial, it's muscle memory and perseverance.

Titanfall 2 has all of these things, campaign, skirmish with bots and PvE, but the skill floor and ceiling is still brutal. Whether it failed or not is a question of how long you expected the game to last, but it still succumbed to the same fate. The more experienced players pulled the ladder up behind them and the matchmaker didn't have enough players to give everyone a fair match.

MindSpunk··on Minecraft: Java Edition now uses SDL3
If you want to get pedantic the hardware rasterizer in your GPU doesn't rasterize "3D triangles" either. Any notion of '3D' comes from how you project your vertices when transforming them through your coordinate spaces. Though the perspective divide needed to make perspective projection work is baked in, but otherwise NDC is a 2D space. Z only exists for the depth buffer/depth test and perspective correct interpolation. The "shape" of the triangle on screen isn't affected by it.

As long as you can draw 2D triangles there's nothing stopping you from drawing a "3D" scene with a software transform pipeline. Depth sorting gets fun without a z-buffer though.

MindSpunk··on Writing a bindless GPU abstraction layer
You could try and do this today, to an extent. You could skip Vulkan and talk to the kernel driver in your app directly. vulkan-1.dll is regular user-mode code, there's nothing special about it. Nobody does this because it's a colossal amount of work for a small pay off. Your code becomes completely un-portable.

Exposing the hardware primitives directly is not an abstraction, so in reality all you end up having to do is implement your own Vulkan instead if you want to support more than 1 GPU. And your app won't work on new GPUs without shipping an update.

If you want a common API to target all GPUs you get Vulkan again. GPU hardware is too diverse to go much lower than Vulkan without just moving the hardware abstraction into your app instead.

MindSpunk··on Building and shipping Mac and iOS apps without opening Xcode
The Apple developer terms of service require all app submissions to come from a Mac.

Apple could've chosen to offer no tools and leave it to the community to build their own. I think that's reasonable. I don't expect Microsoft to bring MSVC to Linux.

But Apple take the explicit position that you must purchase a Mac in order to ship apps for iOS.

MindSpunk··on Building and shipping Mac and iOS apps without opening Xcode
I'm not really sure how any machine is more affordable than the $0 it should cost because you shouldn't need a Mac. It's absolutely true that you can get by with a base spec Mac Mini, but often you can't with bigger projects and the cost skyrockets because Apple prices memory like kidneys (current market conditions aside, they've always done it).

I have nothing against the hardware. I despise how artificial the problem is. It's not even like Apple just refuse to provide tools and leave it up to the community. They force you to use a Mac in the terms of service.

For context I work on the mobile team for a large AAA game engine. This is going to colour my opinion because we're very Windows centric, but my experience wouldn't exist if it weren't for Apple. I need to sync code on two different machines and build half the project on Windows and the other half on the Mac. If we could just build and debug the iPhone on Windows none of this mess would be needed.

We don't develop for MacOS, we develop for iOS. Just debugging the app requires several hoops across multiple machines because of Apple's policy here. Just because a base spec Mac Mini is cheap doesn't make this whole problem any less stupid. I don't need a vendor's special-sauce computer to build for consoles. Cross compilation and remote debugging has existed for decades, this problem should not exist.

Calling it more affordable because Mac Minis aren't junk anymore is really just accepting Apple's ridiculous requirement because "I guess it could be worse right?".

MindSpunk··on Building and shipping Mac and iOS apps without opening Xcode
Why should you need a Mac to build for iOS at all? What makes Apple so special here? Cross-compilation toolchains aren't forbidden alien technology.

Forcing people onto Macs is such a pathetic way of extracting rent from developers and makes workflows worse. No you can't just build and debug your app from your Windows or Linux workstation because we want you to buy hardware you don't need. Sorry we only support Xcode on the latest version of MacOS, how else will your build machine roll out of the OS support window and make you replace a functioning machine?

It's especially nice when you work on projects that aren't iOS/Apple native and you don't get Mac workstations. You end up with this wonderful mess of shared machines and half assed tools because you just need the bare minimum to push a build to a phone.

MindSpunk··on Rewriting Bun in Rust
Curious why you'd move from C# to Rust. C# has you covered mostly for memory safety so I would guess performance or lots of shared memory across threads?
MindSpunk··on OpenBSD has a use-after-free allowing local privilege escalation to root
This is the answer I think. The correctness of your safe code is dependent on the diligence of the unsafe code except for the most simple cases. A kernel is going to have a pretty high unsafe to safe ratio compared to most usermode apps.

This really gets to the core of what I think Rust is about, you can add compiler checked constraints to your APIs that your C and C++ code can't. It's up to you to use them effectively. Rust's ability to keep your safe code safe is a measure of the language, but also your architecture. The buck has to stop somewhere for the language to prove safety, Rust lets you decide rather than the language itself.

MindSpunk··on Qualcomm Linux 2.0
Valve is just hedging against Microsoft having a big red button to kill Steam. They've built their kingdom on top of Microsoft, and Microsoft would love to have it for themselves I'm sure. It's in Valve's best interest to divorce themselves from Windows to protect themselves from Microsoft.

It happens to also benefit the Linux gaming crowd, but it's still ultimately self-interest driving the work. The engineers doing the work are probably doing it for the altruistic reasons, but ultimately Valve is writing the cheques.

MindSpunk··on Physical disc production ending in Jan 2028 for new games on PlayStation
Reminder that Valve's liberal refund policy only exists because they were sued by the Australian government.
MindSpunk··on Minecraft: Java Edition 26.2, the first version with Vulkan 1.2
Consoles ban JITs too.
MindSpunk··on Why is IPv6 so complicated?
NAT is not a security device. A firewall, which will be part of any sane router's NAT implementation, is a security device. NAT is not a firewall, but is often part of one.

Any sane router also uses a firewall for IPv6. A correctly configured router will deny inbound traffic for both v4 and v6. You are not less secure on IPv6.

MindSpunk··on Rust Threads on the GPU
All the names for waves come from different hardware and software vendors adopting names for the same or similar concept.

- Wavefront: AMD, comes from their hardware naming

- Warp: Nvidia, comes from their hardware naming for largely the same concept

Both of these were implementation detail until Microsoft and Khronos enshrined them in the shader programming model independent of the hardware implementation so you get

- Subgroup: Khronos' name for the abstract model that maps to the hardware

- Wave: Microsoft's name for the same

They all describe mostly the same thing so they all get used and you get the naming mess. Doesn't help that you'll have the API spec use wave/subgroup, but the vendor profilers will use warp/wavefront in the names of their hardware counters.

MindSpunk··on The Mechanics of Steins Gate (2023) [pdf]
Not recommending the Steins;Gate anime adaption is pretty wild, it's an incredibly highly rated Anime series. The story telling language of a VN and an Anime are very different so it's no surprise they don't perfectly capture the complexities of the other medium. They don't have to be the same to be worth watching.

fwiw: no idea on the other anime adaptions quality

MindSpunk··on Wine 11 rewrites how Linux runs Windows games at kernel with massive speed gains
The hard part of Linux ports isn't the first 90% (Using the Linux APIs). It's the second 90%.

Platform bugs, build issues, distro differences, implicitly relying on behavior of Windows. It's not just "use Linux API", there's a lot of effort to ship properly. Lots of effort for a tiny user base. There's more users now, but proton is probably a better target than native Linux for games.

MindSpunk··on Our commitment to Windows quality
As if you don't get a jumble of UI frameworks on Linux too.

You can run KDE but depending on the app and containerization you open you'll get a Qt environment, a Qt environment that doesn't respect the system theme, random GTK apps that don't follow the system theme, random GTK apps that only follow a light/dark mode toggle. The GTK apps render their own window decorations too. Sometimes the cursor will change size and theme depending on the window it's on top of.

Page 1 of 9Next →