HNHacker News
TopNewBestAskShowJobs

exDM69

10,584 karma · joined February 2, 2010

submissionscomments
exDM69··on Iran's internet blackout may become permanent, with access for elites only
The GPS jamming maps are based on commercial air traffic flying in the area.

While that gives some ideas of how widespread the jamming is, it won't give accurate information about the range (air traffic avoids areas with jamming) of the interference or any information from places where there is no commercial air traffic (war zones, etc).

exDM69··on Iran's internet blackout may become permanent, with access for elites only
Starlink was also blocked by radio frequency interference.

Granted that can't possibly cover the entire area of the country.

exDM69··on An Experimental Approach to Printf in HLSL
Using printf and debuggers are complementary.

Finding which pixel to debug, or just dumping some info from the pixel under mouse cursor (for example) is better done with a simple printf. Then you can pick up the offending pixel/vertex/mesh/compute in the debugger if you still need it.

You get both, a debugger and printf related tooling in Renderdoc and it's better than either of those alone.

I've been writing a lot of GPU code over the past few years (and the few decades before it) and shader printf has been a huge productivity booster.

exDM69··on An Experimental Approach to Printf in HLSL
Using printf in shaders is awesome, it makes a huge difference when writing and debugging shaders. Vulkan and GLSL (and Slang) have a usable printf out of the box, but HLSL and D3D do not.

Afaik the way it works in Vulkan is that all the string formatting is actually done on the CPU. The GPU writes only writes the data to buffers with structs based on the format string.

All the shader prints are captured by tools such as Renderdoc, so you can easily find the vertex or pixel that printed something and then replay the shader execution in a debugger.

I only wish that we would've had this 20 years ago, it would have saved me so much time, effort and frustration.

exDM69··on Lessons from Hash Table Merging
Maybe this would be a suitable application for "Fibonacci hashing" [0][1], which is a trick to assign a hash table bucket from a hash value. Instead of just taking the modulo with the hash table size, it first multiplies the hash with a constant value 2^64/phi where phi is the golden ratio, and then takes the modulo.

There may be better constants than 2^64/phi, perhaps some large prime number with roughly equal number of one and zero bits could also work.

This will prevent bucket collisions on hash table resizing that may lead to "accidentally quadratic" behavior [2], while not requiring rehashing with a different salt.

I didn't do detailed analysis on whether it helps on hash table merging too, but I think it would.

[0] https://probablydance.com/2018/06/16/fibonacci-hashing-the-o... [1] https://news.ycombinator.com/item?id=43677122 [2] https://accidentallyquadratic.tumblr.com/post/153545455987/r...

exDM69··on Vector graphics on GPU
> why is it not trivial to add a path stage as an alternative to the vertex stage?

Because paths, unlike triangles are not fixed size or have screen space locality. Paths consist of multiple contours of segments, typically cubic bezier curves and a winding rule.

You can't draw one segment out of a contour on the screen and continue to the next one, let alone do them in parallel. A vertical line segment on the left hand side going bottom to top of your screen will make every pixel to the right of it "inside" the path, but if there's another line segment going top to bottom somewhere the pixel and it's outside again.

You need to evaluate the winding rule for every curve segment on every pixel and sum it up.

By contrast, all the pixels inside the triangle are also inside the bounding box of the triangle and the inside/outside test for a pixel is trivially simple.

There are at least four popular approaches to GPU vector graphics:

1) Loop-Blinn: Use CPU to tessellate the path to triangles on the inside and on the edges of the paths. Use a special shader with some tricks to evaluate a bezier curve for the triangles on the edges.

2) Stencil then cover: For each line segment in a tessellated curve, draw a rectangle that extends to the left edge of the contour and use two sided stencil function to add +1 or -1 to the stencil buffer. Draw another rectangle on top of the whole path and set the stencil test to draw only where the stencil buffer is non-zero (or even/odd) according to the winding rule.

3) Draw a rectangle with a special shader that evaluates all the curves in a path, and use a spatial data structure to skip some. Useful for fonts and quadratic bezier curves, not full vector graphics. Much faster than the other methods for simple and small (pixel size) filled paths. Example: Lengyel's method / Slug library.

4) Compute based methods such as the one in this article or Raph Levien's work: use a grid based system with tessellated line segments to limit the number of curves that have to be evaluated per pixel.

Now this is only filling paths, which is the easy part. Stroking paths is much more difficult. Full SVG support has both and much more.

> In fact, you could likely use the geometry stage to create arbitrarily dense vertices based on path data passed to the shader without needing any new GPU features.

Geometry shaders are commonly used with stencil-then-cover to avoid a CPU preprocessing step.

But none of the GPU geometry stages (geometry, tessellation or mesh shaders) are powerful enough to deal with all the corner cases of tessellating vector graphics paths, self intersections, cusps, holes, degenerate curves etc. It's not a very parallel friendly problem.

> Why is this not done?

As I've described here: all of these ideas have been done with varying degrees of success.

> Is the CPU render still faster than these options?

No, the fastest methods are a combination of CPU preprocessing for the difficult geometry problems and GPU for blasting out the pixels.

exDM69··on Vector graphics on GPU
> NV_path_rendering solved this in 2011.

By no means is this a solved problem.

NV_path_rendering is an implementation of "stencil then cover" method with a lot of CPU preprocessing.

It's also only available on OpenGL, not on any other graphics API.

The STC method scales very badly with increasing resolutions as it is using a lot of fill rate and memory bandwidth.

It's mostly using GPU fixed function units (rasterizer and stencil test), leaving the "shader cores" practically idle.

There's a lot of room for improvement to get more performance and better GPU utilization.

exDM69··on Rust's Block Pattern
Yes, sadly this isn't a part of standard C or C++.

It is available as a language extension in Clang and GCC and widely used (e.g. by the Linux kernel).

Unfortunately it is not supported by the third major compiler out there so many projects can't or don't want to use it.

exDM69··on No Graphics API
> tons of drivers that dont implement the extensions that really improve things.

This isn't really the case, at least on desktop side.

All three desktop GPU vendors support Vulkan 1.4 (or most of the features via extensions) on all major platforms even on really old hardware (e.g. Intel Skylake is 10+ years old and has all the latest Vulkan features). Even Apple + MoltenVK is pretty good.

Even mobile GPU vendors have pretty good support in their latest drivers.

The biggest issue is that Android consumer devices don't get GPU driver updates so they're not available to the general public.

exDM69··on Young journalists expose Russian-linked vessels off the Dutch and German coast
Anything involving people on the ground is just too slow.

It takes radars, interceptor drones, sensor networks, etc. Stuff like this is in active development but not widely deployed yet.

exDM69··on Young journalists expose Russian-linked vessels off the Dutch and German coast
I've seen my fair share of frontline combat videos from Ukraine.

The hard part isn't shooting a drone when it is in shotgun range. It's getting the shooter close enough to the drone to have a chance of taking the shot in the first place.

For example the drones mentioned in the article can fly at 2.5km altitude at 140km/h.

exDM69··on Young journalists expose Russian-linked vessels off the Dutch and German coast
The drones here aren't your neighbor's kids' quadrotors. Some sightings over airports have been large (>2m) fixed wing aircraft travelling at 200 km/h. Even the quads are pretty fast. And they can appear out of nowhere, taking off from the ground near the target.

Shooting them down from the ground is next to impossible. They don't hover around waiting for someone to come by with a shotgun in their hand, catching them by land (ie. chasing them in a car) is not feasible.

Just to give an idea how hard it is to hit airborne targets from the ground with traditional guns: I once spent an afternoon shooting at a slow moving fixed wing target drone with tracer rounds from a 12.7mm anti-aircraft machine gun. There were about 50 of us taking turns, each with a few hundred rounds to shoot at the damn thing and the target aircraft didn't get a single hit.

My guess is that the drones are conducting signals intelligence, listening to radar signals and radio comms around sensitive installations (airports, military bases) and surveying the response time to a sighting.

exDM69··on Rust in the kernel is no longer experimental
> Rust has native SIMD support

std::simd is nightly only.

> while in standard C there is no way to express that.

In ISO Standard C(++) there's no SIMD.

But in practice C vector extensions are available in Clang and GCC which are very similar to Rust std::simd (can use normal arithmetic operations).

Unless you're talking about CPU specific intrinsics, which are available to in both languages (core::arch intrinsics vs. xmmintrin.h) in all big compilers.

exDM69··on The C++ standard for the F-35 Fighter Jet [video]
> If the DoD enforces the requirement for Ada, Universities, job training centers, and companies will follow

DoD did enforce a requirement for Ada but universities and others did not follow.

The JSF C++ guidelines were created for circumventing the DoD Ada mandate (as discussed in the video).

exDM69··on Applets are officially gone, but Java in the browser is better
Correct me if I'm wrong but during this timeframe (circa 2005), Java was not open source at all. OpenJDK was announced in 2006 and first release was 2008, by which time the days Java in the browser were more or less over.
exDM69··on Applets are officially gone, but Java in the browser is better
Exactly.

Java was so buggy and had so many security issues about 20 years ago that my local authorities gave a security advisory to not install it at all in end user/home computers. That finally forced the hand of some banks to stop using it for online banking apps.

Flash also had a long run of security issues.

exDM69··on Eurydice: a Rust to C compiler
The project itself is cool and useful but the motivating example of crypto (primitives?) isn't great.

Cryptography is already difficult to write in high level languages without introducing side channels via timing, branch predictor, caches etc.

Cryptography while going through two high level compilers, especially when the code was not designed and written to do so is an exercise fraught with peril.

Tbf, this is just nitpicking about the article, not the project itself

exDM69··on Inside Rust's std and parking_lot mutexes – who wins?
Hi, I don't have public examples to share but I can give an explanation of a simple scenario.

I have a container of resources, e.g. textures. When the GPU wants to use them, CPU will lease them until a point of time in the future denoted by a value (u64) of a GPU timeline semaphore. The handle and value of the semaphore is added to a list guarded by a mutex. Then GPU work is kicked off and the GPU will increment semaphore to that value when done.

In the Drop implementation of the container, we need to wait until all semaphores reach their respective value before freeing resources, and do so even if some thread panicked while holding the lock guarding the list. This is where I use .unwrap_or_else to get the list from the poison value.

It's not infeasible to try to catch any errors and propagate them when the lock is grabbed. But this is mostly for OOM and asserts that are not expected to fire. The ergonomics would be worse if the "lease" function would be fallible.

This said, I would not object to poisoning being made optional.

exDM69··on Inside Rust's std and parking_lot mutexes – who wins?
I've used recovering from poisoned state in impl Drop in quite a few places.

In my case it's usually waiting for the GPU to finish some asynchronous work that's been spun up by CPU threads that may have panicked while holding the lock. This is necessary to avoid freeing resources that the GPU may still be using.

I usually prefix this with `if !std::thread::panicking() {}`, so I don't end up waiting (possibly forever) if I'm already cleaning up after a panic.

exDM69··on Inside Rust's std and parking_lot mutexes – who wins?
Worth noting that this is not `std::mutex` or `parking_lot::mutex` as discussed in the article, but `tokio::sync::Mutex` in cancellable async code.
exDM69··on Inside Rust's std and parking_lot mutexes – who wins?
I disagree, lock poisoning is a good way of improving correctness of concurrent code in case of fatal errors. As demonstrated by the benchmarks in this article, it's not very expensive for typical use cases.

In 99% of the cases where one thread has panic'd while holding a lock, you want to panic the thread that attempts to grab the lock. The contents of anything inside the lock is very much undefined and continuing will lead to unpredictable results. So most of the time you just want:

    let guard = mutex.lock().expect("poisoned");
The last 1% is when you want to clean up something even if a panic has occured. This is usually in a impl Drop situation. It's not much more verbose either, just:

    let guard = mutex.lock().unwrap_or_else(|poison| poison.into_inner());
What is painful is trying to propagate the poison value as an error using `?`. In that case you're probably better off using a match expression because the usual `.into()` will not play nice with common error handling crates (thiserror, anyhow) or need to implement `From` manually for the error types and drop the contents of the poison error before propagating.

This might be the case for long running server processes where you have n:m threading with long running threads and want to keep processing other requests even if one request fails. Although in that case you probably want (or your framework provides) some kind of robustness with `catch_unwind` that will log the errors, respond with HTTP 500 or whatever and then resume. Because that's needed to catch panics from non-mutex related code.

exDM69··on The state of SIMD in Rust in 2025
> A f32x16 version would also be faster on 256b hardware, but spill in SSE. For Zen5 you probably want to use f32x32.

Yeah, exceeding native vector width is kinda just adding another round of loop unrolling. Sometimes it helps, sometime it doesn't. This is probably mostly about register pressure.

And architecture specific benchmarking is required if you want to get most performance out of it.

> I'd prefer if std::simd would encurage relative to native SIMD width scaling (and support scalable SIMD ISAs).

It is possible to write width-generic SIMD code (ie. have vector width as generic parameter) in Rust std::simd (or C++ templates and vector extensions) and make it relative to native vector width (albeit you need to explicitly define that).

In my problem domain (computer graphics etc) the vector width is often mandated by the task at hand (e.g. 2d vs 3d). It's often not about doing something on an array of size N. This does not lead to optimal HW utilization, but it's convenient and still a lot faster than scalar code.

Scalable SIMD ISAs are kind of a new thing, so not sure how well current std::simd or C vector extensions (or LLVM IR SIMD ops) map to the HW. Maybe they would be better served by another kind of API? I don't really know, haven't had the privilege of writing any scalable vector code yet.

What I'm trying to say is IMO std::simd works well enough and should probably be stabilized (almost) as is, barring any show stopper issues. It's already useful and has been for many years.

exDM69··on The state of SIMD in Rust in 2025
> Getting maximum performance out of SIMD requires rolling your own code with intrinsics

Not disagreeing with this statement in general, but with std::simd I can get 80% of the performance with 20% of the effort compared to intrinsics.

For the last 20%, there's a zero cost fallback to intrinsics when you need it.

exDM69··on The state of SIMD in Rust in 2025
Despite all of these issues you mention, std::simd is perfectly usable in the state it is in today in nightly Rust.

I've written thousands and thousands of lines of Rust SIMD code over the last ~4 years and it's, in my opinion, a pretty nice way of doing SIMD code that is portable.

I don't know about the specific issues in stabilization, but the API has been relatively stable, although there were some breaking changes a few years ago.

Maybe you can't extract 100% of your CPUs capabilities using it, but I don't find that a problem because there's a zero-cost fallback to CPU-specific intrinsics when necessary.

I recently wrote some computer graphics code and I could get really nice performance (~20x my scalar code, 5x from just a naive translation). And the same codebase can be compiled to AVX2, SSE2 and ARM NEON. It uses f32x8's (256b vector width), which are not available on SSE or NEON, but the compiler can split those vectors. The f32x8 version was faster than f32x4 even on 128b hardware. I would've needed to painstakingly port this codebase to each CPU, so it was at least a 3x reduction in lines of code (and more in programmer time).

exDM69··on The state of SIMD in Rust in 2025
In my experience, compiling C with -ffast-math will tremendously improve floating point autovectorization and optimizations to SIMD (C vector extensions, which are similar to Rust std::simd) code in general.

This obviously has a lot of caveats, and should only be enabled on a per function or per file basis.

Unfortunately Rust does not currently have options for adjusting per-function compiler optimization parameters. This is possible in some C compilers using function attributes.

exDM69··on Things you can do with diodes
Two more from the world of analog music/guitar electronics:

1) Ring modulator: https://en.wikipedia.org/wiki/Ring_modulation

A device used to multiply two analog signals in time domain. Best known for the sound of the Daleks in the original 1960s Doctor Who series. Has some applications outside of music and sound effects. If you can find those old fashioned audio transformers, this effect does not require a power source.

2) Diode clipper: https://en.wikipedia.org/wiki/Clipper_(electronics)

Two diodes in parallel with opposite polarities. Clips the incoming AC signal to a +/- diode threshold voltage. Put a high voltage gain amplifier stage in front of it and you get the classic electric guitar distortion tone you know and love. Allegedly works best with germanium-unobtainium diodes. In their absence, using two different kinds of diodes can also have pleasant tonal qualities.

exDM69··on How to stop Linux threads cleanly
> To kill the thread, set the stop flag and cond_signal the condvar

This is a race condition. When you "spin" on a condition variable, the stop flag you check must be guarded by the same mutex you give to cond_wait.

See this article for a thorough explanation:

https://zeux.io/2024/03/23/condvars-atomic/

exDM69··on OpenGL: Mesh shaders in the current year
> I am sadly aware, but I won't switch until the complexity is fixed

It pretty much is by now if you can use Vulkan 1.4 (or even 1.3). It's a pretty lean and mean API once you've got it bootstrapped.

There's still a lot of setup code to get off the ground (device enumeration, extensions and features, swapchain setup, pipeline layouts), but beyond that Vulkan is much nicer to work with than OpenGL. Just gotta get past the initial hurdle.

exDM69··on How to create an OS from scratch
Maybe worth venturing into the embedded land? There are some pretty cool Rust embedded projects (e.g. embassy) which are sort of best of both worlds. You get to do low level hardware tinkering (which arguably requires a systems programming language) and you get a sort of batteries included environment where you can use all the nice Rust high level features (memory safety, async, etc).

Easier to get started with than OSdev and less gruesome legacy hardware details to study to get stuff done.

Next time I need some lights blinking or actuators actuating I'm gonna do it with Rust and rp2040+.

exDM69··on How to create an OS from scratch
The osdev wiki is the best source for this and they have the bare bones tutorials with all the linker scripts, bootstrap code, makefiles, etc you might need. They have examples for multiboot and UEFI (and BIOS boot sector).

For Rust there's this popular series: https://os.phil-opp.com/

If I were to start a new OS project I would do it in Rust too but as awesome as Rust is, I can't recommend doing a bare metal project as your first foray into the language. Learning two or more things at once doesn't work for me.

← PreviousPage 2 of 34Next →