HNHacker News
TopNewBestAskShowJobs

pocak

29 karma · joined April 25, 2018

submissionscomments
pocak··on Intel's Battlemage Architecture
There is, it's Druid. Intel announced the first four codenames in 2021.

> [...] first generation, based on the Xe HPG microarchitecture, codenamed Alchemist (formerly known as DG2). Intel also revealed the code names of future generations under the Arc brand: Battlemage, Celestial and Druid.

https://www.intel.com/content/www/us/en/newsroom/news/introd...

pocak··on FuryGpu – Custom PCIe FPGA GPU
In the post about the texture unit, that ROM table for mip level address offsets seems to use quite a bit of space. Have you considered making the mip base addresses a part of the texture spec instead?
pocak··on The Apple GPU and the impossible bug
Right, I had the article's bunny test program on my mind, which looks like it has only one pass.

In OpenGL, the driver would have to scan the following commands to see if it can discard the depth data. If it doesn't see the depth buffer get cleared, it has to be conservative and save the data. I assume mobile GPU drivers in general do make the effort to do this optimization, as the bandwidth savings are significant.

In Vulkan, the application explicitly specifies which attachment (i.e. stencil, depth, color buffer) must be persisted at the end of a render pass, and which need not. So that maps nicely to the "final render flush program".

The quote is about Metal, though, which I'm not familiar with, but a sibling comment points out it's similar to Vulkan in this aspect.

So that leaves me wondering: did Rosenzweig happen to only try Metal apps that always use MTLStoreAction.store in passes that overflow the TVB, or is the Metal driver skipping a useful optimization, or neither? E.g. because the hardware has another control for this?

pocak··on The Apple GPU and the impossible bug
I don't understand why the programs are the same. The partial render store program has to write out both the color and the depth buffer, while the final render store should only write out color and throw away depth.
pocak··on The Apple GPU and the impossible bug
That's what I thought, too, until I saw ARM's Hot Chips 2016 slides. Page 24 shows that they write transformed positions to RAM, and later write varyings to RAM. That's for Bifrost, but it's implied Midgard is the same, except it doesn't filter out vertices from culled primitives.

That makes me wonder whether the other GPUs with position-only shading - Intel and Adreno - do the same.

As for PowerVR, I've never seen them described as position-only shaders - I think they've always done full vertex processing upfront.

edit: slides are at https://old.hotchips.org/wp-content/uploads/hc_archives/hc28...

pocak··on How Facebook encodes videos
In the article, cost is cpu time, and benefit is file size reduction multiplied by number of times watched.
pocak··on Why mmap is faster than system calls (2019)
Linux allocates page tables lazily, and fills them lazily. The only upfront work is to mark the virtual address range as valid and associated with the file. I'd expect mapping giant files to be fast enough to not need windowing.
pocak··on Reasonably priced color e-ink display
Only one row at a time has voltage applied to it. In one update, the image is scanned out multiple times, so it appears as if all pixels were changing simultaneously (and perhaps they do, if the electrodes have significant capacitance.)

With fewer rows to update, each row gets a push more often, flipping the grains faster.

pocak··on No Paint
Took me a while to get past the home screen. I kept clicking 'paint' at the bottom and 'no' at the top.
pocak··on Intel Discloses Lakefield CPUs Specifications
Thread migration only costs on the order of 100 microseconds, including the effect of cold caches. If you keep the AVX thread on the big core for at least 100 milliseconds at a time, you only lose ~0.2% performance.
pocak··on Intel Discloses Lakefield CPUs Specifications
Migrate back to the little core if the thread hasn't used AVX in a while.

Linux already tracks how long ago a task used AVX-512. I assume the same mechanism could be used to track AVX as well.

https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git/...

pocak··on How to trim video clips instantly without reencoding
It works by applying a low-pass filter on the camera's motion, so a sudden kick turns into a slow, low amplitude bob.

Perhaps you just need to increase the smoothing parameter. I found the default of 10 way too low, and needed around 50-100 for my very shaky 50 fps home videos.

pocak··on OpenBSD chief de Raadt says no easy fix for new Intel CPU bug
If you give two threads on one core to one VM, the VM could attack the hypervisor. It could have one thread run the TLB probe loop while the other does something that's handled by the hypervisor.
pocak··on AMD Tackles Coming “Chiplet” Revolution With New Chip Network Scheme
Are newer Ryzen 1000 series CPUs still made with the B1 die?

I thought steppings were for fixing errata, and once a new revision is qualified, the old one is no longer manufactured.

pocak··on Sleep Scientist Warns Against Walking Through Life 'In an Underslept State'
There's a nine digit number in the middle of the normal URL. You can take a text-only URL and replace the identifier at the end.

Shame on them for not making all articles reachable from the text home page, and for not redirecting to the article from the choice page.

https://text.npr.org/s.php?sId=558058812

pocak··on TSMC Kicks Off Volume Production of 7nm Chips
If your chips aren't very large, you can fit multiple copies in the reticle. E.g., if the litho machine has a 40mm by 50mm reticle, and you're making a 10mm by 10mm die, you can expose 20 dice at a time. It's usually worth it to go out to the edges even if you only get a few complete chips from the exposure. I can't find a wafer picture for Nvidia's GV100, but presumably a giant chip like that uses the whole reticle, so I don't expect to see partial chips on its wafer.