HNHacker News
TopNewBestAskShowJobs

feffe

114 karma · joined February 16, 2021

submissionscomments
feffe··on Atari ST in daily use since 1985 [video]
But computers use a lot more power nowadays. The early home computers were all passively cooled. My Atari ST power supply actually broke down on me so I had to buy a new one from Great Britain. It's a quite small affair with quite minimal venting through the grills at the top of the chassis.
feffe··on Resource efficient Thread Pools with Zig
The effect is the same. If someone touch the cache line, it's evicted from all other caches that has it, triggering a cache miss when other cores touch it. Everyone knows this. I just think it's a bit depressing if you try to optimize message passing on many-core CPUs you'll realize you can't make it fast. No-one has been able to make it fast (I've checked Intel message passing code as well). If you get more than 20 Mevents through some kind of shared queue, you are lucky. That is slow if you compare to how many instructions a CPU can retire.

So all these event loop frameworks try to implement an async behaviour ontop of a hardware mechanism that is synchronous in it's core. The main issue is that the lowest software construct, the queue used to communicate is built on this cache line ping pong match. What the software want it still a queue, you could still send pointers if the memory they point to has been committed to memory when the receiving core see them. Just remove the inefficient atomic operations way of synchronizing the data between CPUs. Send them using some kind of network pipe instead :-) As you say the memory subsystem is really asynchronous.

I'm convinces this is coming, it should just have arrived 10 years ago...

feffe··on Resource efficient Thread Pools with Zig
Yes, I'm probably wrong on the semantics of lock free. I should read up on that. My point was that maybe you can gain a factor of 2 by using atomic ops in a "lockless" fashion compared to mutexes but it still scales extremely poorly compared to what you could do if the hardware had real async mechanisms that did not rely on cache coherency. The cache line that is subjected to an atomic operation will be evicted from all other CPUs cache lines, back and forth.

A simple example is just letting all CPUs in a multi core CPU doing an atomic add at a memory address. The scaling will be exactly 1 with the number of available cores. I realize this is very of topic wrt to this cool article an implementation in zig. It's just that this problem can't really be solved in an efficient manner with todays hardware.

feffe··on Resource efficient Thread Pools with Zig
It's a lock associated with the cache line that the atomic operation operates on. This is because they are built on top of the cache coherency mechanism, a synchronous blocking operation on the hardware level to implement an asynchronous mechanism on the software level.

The big frustration with today's multi core CPUs is that there's simply not efficient way to communicate using message passing mechanisms. I think this is something the hardware guys should focus on :-) Provide an async mechanism to communicate between cores, not relying on the cache coherency.

feffe··on Resource efficient Thread Pools with Zig
A pet peeve of mine is calling something using atomic operations lock less. It's very common but once it click that atomic operations are just locks on the instruction level it feels very wrong to call most algorithms lock less. There's still the same concerns to avoid contention when using atomic operations as with regular mutexes. A mutex lock in the uncontended case never enters the kernel and is just an atomic operation... The main strategy regardless of if using locks or atomics directly is to avoid contention.

The only truly lock less algorithm that I'm aware of (for most CPU archs) is RCU in the read case.

feffe··on In Go, pointers (mostly) don't go with slices in practice
This is how I think about Go slices (may help other understand them).

A slice itself is just a window into a backing array of fixed size. The slice carries three data members. The pointer to the backing array and its remaining capacity and the length of the slice data.

Typically slices are passed around by value but you can take their address and modify a "shared" slice.

The built-in append() returns a new slice by value.

What happens is simply that when appending data to a slice and there is no room in the backing array, a new backing array is allocated that the returned slice points into. The old "input" slice to append is still intact and if some code has access to it, it will look at data stored in the old backing array.

I've constructed similar utility types in C and find them quite convenient. It's very convenient to have the distinction between the backing memory (array) and a slice viewing a portion of it instead of just a dynamic array.

feffe··on Away from Exceptions: Errors as Values
The article is not about Go, although that was my guess before reading it as well.
feffe··on Why are hyperlinks blue?
Maybe because blue is the most visible color against a white background, except for black? Like blue and black are the two most common pen ink colors.
feffe··on JPG vs. RAW: Get it right the first time (2009)
JPEG is quite old by now. The reason I shoot RAW is mostly to adjust contrast, saturation, white balance and recover some highlight and shadow details.

If there where some lossy image format that recorded the image in YUV (or some similar color space) and recorded information beyond the black and white point to be able to recover some details and additionally supported some standardized transforms that could be adjusted on top of the base image I think the need for RAW would diminish. Such transforms would be to set the white point, black point, gamma curve, saturation, white balance. If the standard transform data could be edited in an image editing program, the lossy YUV data would not need to be rewritten causing an image degradation.

feffe··on GNU Coding Standards: Writing Robust Programs
The core dump contains information for fault analysis.

There is another aspect, C and C++ are not memory safe languages so the bug may not be particular logical (i.e. some kind of memory corruption). In these cases I actually prefer something to the like of __builtin_trap instead of abort. Calling any code after an invalid invariant has been detected clobbers registers and may make it impossible to investigate the state at the time of the "crash". Some features of modern optimizing compilers make this even worse, such as abort being marked as "noreturn".

feffe··on Commander X16
I think so. Although I had a C64, I cut my teeth on programming the Atari ST in assembler. The 68k is so much nicer to program than the 6502. The graphics and sound capabilities of the x16 is more in line with a turbo charged Atari ST or Amiga 500. That would be my dream retro station. Still I'll buy a x16 if it ever materializes, could be fun to play with some raster effects and sprites again.
feffe··on New x86 micro-op vulnerability breaks all known Spectre defenses
The memory latency hiding also works with 2-way SMT. I worked on a networking software doing per packet session lookup in large hash tables. SMT with a Sandybridge core in this application gave 40% better performance which is higher than usually mentioned. So for memory bound (as in cache misses) applications, SMT is a boon.
feffe··on Go is not an easy language
I don't think this is a good example of proving the point that Go is difficult to read. Isn't it expected to internalize the semantics of basic language primitives when learning a new language, or does people just jump in guessing what different language constructs do? IMHO you learn this the first week when reading up on the language features.

range returns index and elements by value. The last example does what you asks of it, it's like complaining something is not returned by reference when it's not. Your mistake. Perhaps some linter could give warnings for it.

feffe··on Most M1 Macs appear to have a serious SSD wear defect
These are my stats from 2 identical sticks of 250 GB Samsung SSD bought in the summer of 2018. So yea, Apple probably want to tweak some settings. I use my computer a lot, Linux more than the Windows install which is mostly for gaming.

Windows 10: Data Units Read: 11 332 875 [5,80 TB] Data Units Written: 8 173 369 [4,18 TB]

Linux: Data Units Read: 5 837 296 [2,98 TB] Data Units Written: 4 891 744 [2,50 TB]

← PreviousPage 3 of 3