HNHacker News
TopNewBestAskShowJobs

Const-me

6,400 karma · joined March 2, 2015

meet.hn/city/42.4303762,18.6988104/Tivat

http://const.me/

submissionscomments
Const-me··on SDF vs. MSDF vs. Slug: GPU Text Rendering
“What makes glyphs hard” Another reason is hinting. Traditional text renderers like FreeType are aware of the pixel grid and they adjust the curves slightly, snapping them to that grid. For all methods in the article quite hard to do on GPUs.

“Chinese, Japanese, and Korean have tens of thousands of glyphs, and baking all of them at several sizes is a memory disaster” One possible solution is dynamic atlas built on CPU for visible glyphs only.

“The distance-field panels notch, where interpolating between stored samples no longer matches the true curve” Can’t it be fixed in the shader, using screen-space derivatives of the SDF? I think in theory, SDF value for pixel center combined with screen-space gradient vector of that number delivers enough data to compute partial coverage for the pixels on the edge.

Const-me··on If we do not stop to help each other, what do we become?
> Can you name anything you've learned from an LLM?

Not GP, but here’s some of the things I have learned in the last couple weeks from LLMs while asking LLMs to review parts of my code and documentation. I wrote the code and documentation without LLMs.

Linux user permissions is a fake because by design, the owner has permissions to write these 3 bits. A good way to reliably deploy a readonly file is owner=root, assign the read permission to the group.

chown command follows links by default, -h switch disables that.

In Linux, resetting environment variable in a long-lived server process doesn’t erase much data because /proc/{pid}/environ is a snapshot taken on process startup.

If my server crashes (or manually killed by a root to test things on staging VM) and Alpine’s supervise-daemon restarts it, the start_pre() function from the OpenRC service script won’t be called again. A fix is use OpenRC to run a launcher shell script as root, in that script manually downgrade the account with `exec su-exec svc:svc /the/server/elf`

This last one is irrelevant to my use cases but still interesting. AES GCM may fail catastrophically when different message encrypted with same password and nonce. For 12 bytes of random nonce, the stuff is only safe up to approximately 4E+9 messages using the same password, beyond that birthday paradox ruins things.

Const-me··on Reverse-engineering the Intel 8087's tangent algorithm: more than CORDIC
I remember I once wanted to compute tangent of fp64 vectors. Here’s what I did.

https://github.com/Const-me/AvxMath/blob/master/AvxMath/AvxM...

Const-me··on Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
See remark on the slide 11: https://www.servethehome.com/micron-evolving-memory-architec... That presentation is by Micron.
Const-me··on OpenAI buys smartphone camera maker Glass Imaging for $300M
I wonder why hasn’t the OP forked postmarketOS or some other Linux?

I totally get why they wanted a custom userland, built it myself for embedded devices. But IME Linux kernel is not terribly bad these days.

Const-me··on Anthropic researcher believes more than 10% chance AI 'could kill all humans'
I think nuclear winter is a fake invented by cold war propaganda. If tomorrow we detonate all nukes available, the amount of dust in the atmosphere won’t even approach a large volcanic eruption.

Volcanic eruptions occasionally inject literally cubic kilometres of stuff high into the atmosphere. That’s insufficient to cause a global winter, only a year without a summer: https://en.wikipedia.org/wiki/Year_Without_a_Summer

Const-me··on Nvidia agrees to acquire Hugging Face for $13B
For the last few months, I’ve been working on a vendor-agnostic inference library. Have 3 backends so far: legacy D3D11 for compatibility, D3D12, and Vulkan 1.3. I have reasons to believe nVidia deliberately crippling Vulkan API for their consumer GPUs. Couple examples to be specific.

nVidia driver sets quite low number for VkPhysicalDeviceLimits::maxTexelBufferElements. I don’t think that’s a hardware limit because I have D3D12 backend doing the same thing on the same hardware.

Vulkan performance is not great on nVidia. On all AMD cards I am testing, Vulkan 1.3 is the fastest backend. On nVidia however, D3D12 is faster despite tensor cores (WMMA / wave matrix multiply accumulate / cooperative matrices) are only available through Vulkan, D3D12 is using shader cores exclusively.

Const-me··on SIMD in the 90s: Programming Intel's Pentium MMX
I would add than SSE1 and SSE2 are now required parts of AMD64 instruction set. All 64-bit PC processors are required to support them both. For that reason, modern compilers are ignoring x87 FPU when building 64-bit binaries. Instead, they compile all float and double arithmetic into SSE1 and SSE2 instructions, respectively.
Const-me··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
Before Vista, we did not have window previews in task bar and alt+tab. We didn’t have a good multimedia framework based on the hardware codecs: media foundation arriving with Vista wasn’t a coincidence. It was hard to capture and encode desktop for screen recording and presentations: despite MS only added desktop duplication API in Win8, technically Vista and Win7 graphics stacks could already do that, MS just neglected. Also, these aero translucency visuals in Vista and Win7 were nice, until Win8 ruined everything.

All that stuff would be hard to impossible with the older GDI architecture and no desktop compositor process.

Const-me··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
“aren't truly using a fully shared memory pool” I think with AMD iGPUs I have here (GCN 5.1 and RDNA3 generations) it’s actually unified, at least on Windows 10. The difference between the reserved portion and the rest of the memory is cosmetic.

From Vulkan API POV, the reserved portion has device local and multi instance heap flags, the main heap doesn’t. However, the memory type is identical across all heaps, all of them have device local, host visible and host coherent property flags. And I can confirm VMA library from Vulkan SDK successfully allocates way more device visible memory than the size of that reserved portion.

Const-me··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
There’s a reason why all modern desktop environments are designed the same way: power efficiency when multitasking.

Imagine you have 3 windows visible at the same time: a videogame rendering at the refresh rate of the display 144 Hz, a video player rendering frames at 30 Hz, and a text editor rendering blinking cursor at 2 Hz. Because the videogame wants to deliver frames at 144 Hz, the desktop compositor has to deliver the entire desktop at 144 Hz. Asking the video player and especially the text editor to also deliver frames at that frequency would be wasteful. Irrelevant for desktops with fast discrete GPUs, but directly translates to battery drain on laptops.

Const-me··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
True, but many modern computers are using unified memory. On such systems all memory is almost equal, despite often reported differently.

For example, on my 5 years old laptop with integrated AMD GPU, windows 10 calculator in default state uses 33 MB system RAM, 9.6 MB dedicated VRAM. Maximized to FullHD screen, same app uses 36 MB system RAM, 13 MB dedicated VRAM. Maybe the OS counts VRAM as active private working set, maybe the app uses more than 1 buffer.

Regardless of the reason, it’s IMO unrealistic to expect a modern GUI app to consume less memory than required for the frame buffer for its window.

Const-me··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
10 MB is not too bad for a GUI app. If the app is full screen, display is FullHD and has 8 bit depth, that's almost 8 MB memory for the back buffer alone. Enable HDR and pixels become 8 bytes RGBA16_Float instead of 4 bytes BGRA8_Unorm, twice as much memory.
Const-me··on Is it time for a new Embedded Linux build system?
I agree. In the past, I have successfully used Debian and Alpine for embedded. Never needed to compile OS kernels or standard DLLs, other people already did and published in these package repositories.
Const-me··on Windows 11 New Media Player Uses 3.5x More RAM, Charges for Popular Video Codecs
Why is that HEVC video extension is required?

As a part of the user-mode half of the GPU driver, GPU vendors ship media foundation transform DLLs to use HEVC hardware codecs. Don’t AMD, Intel and nVidia already pay patent royalties? I expect them to include into price of the GPUs with hardware support i.e. all of them made in the last decade.

Const-me··on Satellite reveals immense scale of GPS signal tampering
> typical crystals in the 10-100MHz range

I think most quarts watches oscillate at 32 kHz = 2^15 Hz, high precision quartz watches at 8.4 MHz = 2^23 Hz.

> The actual problem is stability over temperature

Apparently, designers of these watches compensating for that somehow: https://en.wikipedia.org/wiki/Quartz_clock#Thermal_compensat...

> benefits from being kept at constant-ish body temperature

Some people take off their watches every day before going to sleep.

These high-end quartz oscillators are probably too expensive to use in commodity computers. Still, the cost shouldn’t look too bad when compared to a price or an airplane, marine vessel, or most military equipment.

Const-me··on Satellite reveals immense scale of GPS signal tampering
> A typical PC clock is +/-100ppm. After 1 hour that's 0.36s

Are you confident in these numbers? They add up to 52 minutes of drift/year.

Good modern quartz watches specify 5 seconds/year drift, almost 3 orders of magnitude better.

Const-me··on What every coder should know about gamma (2016)
> On which image does the gradation appear more even? It’s the second one!

Can’t reproduce. Tested on two monitors on my desk, designer-targeted Benq and cheap laptop. On the Benq, darkest 3 segments are indistinguishable, the 4-th one barely distinguishable. On the laptop, darkest 4 segments are indistinguishable, the 5-th barely distinguishable. However, on the “emitted light intensity” all bars are clearly visible.

> Image resizing

“Unsurprisingly, C gives the correct result” On my computers B very similar to A, just a tiny bit darker. While the “correct” C result is a lot lighter than A.

Also from the same section:

> B the result of resizing the pattern by 50% directly in sRGB-space (using bicubic interpolation)

Bicubic interpolation is only applicable when enraging images; downsampling is very different problem from interpolation.

Const-me··on Hetzner Price Adjustment
“Is it because of government regulations, do we need to deregulate?”

Insufficient law enforcement. The same memory manufacturers already broke antimonopoly laws in the past, pleaded guilty. Apparently the fines were too small for these companies to care, and the people responsible were promoted instead of being punished. More information: https://en.wikipedia.org/wiki/DRAM_price_fixing_scandal

Const-me··on Making Graphics Like it's 1993
On modern processors, floating point addition often has equal performance to floating point multiplication. For example, on AMD Zen4 it’s 3 cycles latency and 0.5 cycles throughput.

I’m not sure that trick going to work in the context of computer graphics. To transform vectors or multiply matrices you need a mix of multiplications and additions, or an equivalent sequence of FMAs.

Const-me··on Poland to jail online streamers of violent crimes and cruelty for up to 5 years
Might be jurisdiction. Let’s say a person who is not a Polish citizen committing and broadcasting a crime outside of Poland, then trying to enter Poland. IANAL but I think this law sends that person to jail as long as the video is accessible from inside Poland.
Const-me··on Making Graphics Like it's 1993
> let alone more performant

Not anymore. On modern hardware, the only operation where integers win is single cycle add/sub. For the rest of operations (multiplication, division, square roots, etc.) floating point is faster, sometimes by a lot.

Const-me··on Lies we tell ourselves about email addresses
Good article. Worth noting C# standard library handles most of that complexity, no regular expressions required. Call System.Net.Mail.MailAddress.TryCreate, if successful read Address property to find the normalised address.
Const-me··on Premature optimization is fun sometimes
Cool trick, but personally I don’t trust C bitfields. When I need something like that, I usually create C++ class or C# structure with a single private uint64 field, and public methods to extract or manipulate the logical fields.

Because the class/structure only has a single uint64 field, the compilers are likely to pass value in a single general-purpose register. I believe that’s unlikely to happen for a structure with bit fields.

If you target AVX2 or newer you also have BMI1 and BMI2, intrinsics like bextr and bzhi are probably faster than whatever codes compilers are generating for bit fields.

Binary compatibility of bit fields is a moot point, using them at the API surface across compilers or languages is not ideal. A structure with a single uint64 field is very compatible.

Const-me··on Ask HN: What was your "oh shit" moment with GenAI?
None so far. When I try to use these language models in the primary areas of my expertise like SIMD or GPGPU they fail to do any good. When I ask them to implement some general-purpose stuff, the output is too low quality to be useful in my software.

Still, find them incredibly useful for code review (despite unable to write good C++ or C#, smart enough to detect issues there), also dealing with technologies outside of my area of expertise like Python or web stuff.

Const-me··on Dav2d
> Performance should not be priority #1. Security should be.

For a web browser, or a server in a bank, sure. For anything else, questionable.

> adding a sandbox around a memory-unsafe codec is going to be way more expensive

In modern world, overhead of strong sandboxes is surprisingly small. A nuclear but most reliable option is hardware assisted VM. On modern computers with SLAT and virtualized IO the overhead for most use cases is negligible. If you want something lighter weight, can use a multi-user nature of all modern OS kernels and isolate into a separate process with restricted permissions. Sandboxing overhead is approximately zero.

Const-me··on What it takes to transpose a matrix
The AVX2 SIMD version is not ideal. Too many instructions, and it needs constant vectors. I would rather do it like that https://godbolt.org/z/cn6YKbfYd
Const-me··on Making deep learning go brrrr from first principles (2022)
> decode (GEMV) is memory bound

Decode with batch size 1 is GEMV. Batching makes the decode GEMM too.

Const-me··on Making deep learning go brrrr from first principles (2022)
> Most of those FLOPS are constrained by memory bandwidth

I believe inference with large enough batch size is almost always compute bound, simply due to algorithmic complexity.

Each step of tiled matric multiplication with square tiles of size N^2 takes O(N^2) memory loads and O(N^3) compute operations. With N = 32 or 64, you will likely saturate compute even on iGPUs with DDR4 or DDR5 memory pretending to be VRAM.

Const-me··on Ollama is now powered by MLX on Apple Silicon in preview
While data centres indeed have awesome internet connectivity, don’t forget the bandwidth is shared by all clients using a particular server.

If you have 100 mbit/sec internet connection at home, a computer in a data centre has 10 gbit/sec, but the server is serving 200 concurrent clients — your bandwidth is twice as fast.

Page 1 of 34Next →