HNHacker News
TopNewBestAskShowJobs

evmar

6,850 karma · joined June 7, 2011

https://neugierig.org

evan.martin@gmail.com

My posts on HN: https://news.ycombinator.com/from?site=neugierig.org

submissionscomments
evmar··on Everybody’s home. No one’s coming over
I moved from a rich city social life to a lonely suburban one (I have regrets, there are complicating factors). I met a fellow dad who mentioned a book as we started chatting and I was excited (someone who reads!).

As we talked about this very isolation problem he told me about his neighbors who kept inviting him over and offering to watch the kids for him and how he thought it was just weird and couldn’t trust them. It broke my heart — I could just imagine the neighbors trying to be good to their community and that very behavior is met with distrust.

(I read half the book he mentioned and asked him about it the next time I saw him. Turns out he hadn’t read the book, just heard about it on a podcast. My quest to meet someone who reads continues.)

evmar··on Transit rewards
A trick I learned from an east bay friend is that (when possible) the bus across the bridge is way nicer than the BART under the bay, especially around sunset.
evmar··on It took a year to ship WebAssembly in Anubis
I think the Rust feature you’re looking for regarding recompiling the standard library is called “build-std”, that should be enough for you to search for it. (For similar reasons you also need that flag if you are trying to use Rust to build multithreaded wasm binaries, so it might come up for you!)
evmar··on Visualizing Rust's Vtables: How dyn Trait Works In Memory
In my own journey of discovery I found https://cheats.rs/ very helpful, and in particular its "memory layout" section has visualizations. (No affiliation with the site, just a happy reader!)
evmar··on Classical chess ranks 562nd of 960 starting positions after 460,800 games
Totally agree. I saw the idea of “Internet Kessler syndrome” recently and I can’t stop thinking about how it captures this.
evmar··on Faster Than Ninja
I did some exploration of this idea in a followup build system! See https://neugierig.org/software/blog/2022/03/n2.html . (It's not really production-ready.)
evmar··on Faster Than Ninja
Thanks for saying this! I am close enough to it that I mostly remember all of the bad decisions I made that are now unfixable, haha.
evmar··on Faster Than Ninja
Yes, I don’t remember the details but vaguely remember that CMake tends to group things together that could in principle be made more parallel, as you mention with generates files in a library. On the other hand if build2 makes it easier for authors to express these kinds of patterns without the serialization then I count that as a win for build2!
evmar··on Faster Than Ninja
[Ninja author here] Nice post, cool to see the deep dive! I also appreciate the details on how they produced their numbers.

As they observe, Ninja gets to be fast mostly by cheating: it avoids a lot of work by saying many things are just out of scope for Ninja to do, and that means it is a useful a target to race against. (Funny thing: when I wrote Ninja I was misremembering how fast an earlier build system was so I kept trying to make it faster. So don't treat it as a lower bound, I just made it up!)

I comment here to say I find the explanation for 'why' in this post unsatisfying. They mention three design decisions.

The first one is a criticism of CMake, not Ninja (?), so I don't think it can be why. I might have misunderstood?

The second reason given is doing some work like header dependencies in multiple threads. This is the most plausible reason to me but it still feels unlikely. It's a very small amount of work: the post mentions 300 compiles, so maybe parsing 300 small text files?

The third is that they run the compiler up front an additional time to gather headers, which is strictly more work than Ninja. There is some hand waving about file access patterns but I am skeptical; if the end-to-end build time is 3 seconds then the project is small enough to all fit in kernel caches. They also mention doing other things like invoking the compiler to get version information. This seems like it would dwarf any performance gain from number 2.

Maybe it's just my own curiosity, I think this post would be better if it had a better explanation for the reason. I'm not disputing the result, I just think the result should make you suspicious that something else is going on, and you might learn something from that! You could for example explore whether it's the header dependency thing by profiling the Ninja invocation and seeing if it's waiting for CPU or waiting for tasks to execute.

(If I had to guess without looking at any of the involved code, I would predict it's something about how CMake generates the build, like it introduces serialization in a place where build2 is parallel, or it adds some extra build steps like gathering the current git hash into a header file or something.)

evmar··on The great blogging collapse: What happened to 100 successful blogs?
If you imagine Google's job is to present useful information, these blogs that are maximizing cash while simulating usefulness are exactly the sorts of things I would hope Google to want to filter out.

(I don't think Google's often capricious ranking changes really succeed at this, but the outcomes in this post seems like something hypothetically good?)

evmar··on I hate compilers
A better solution might be to use https://github.com/evanw/polywasm to run the original wasm in place.
evmar··on The experience of rendering Arabic typography and its technical debt
One thing I sometimes think about when I think about text layout problems is how the text we use also has a bunch of complexities that we can take for granted.

Think of variable width characters and kerning and ligatures and hyphenation and justification. Imagine computers had been won by a CJK language, which have none of these problems. You could imagine a similar article about how exotic and difficult English layout is.

evmar··on Britain Became as Poor as Mississippi
Thanks, I'd love to add them! Do you have a good source for this data? I did a quick look at the site you linked above and I'm not sure whether it has numbers for GDP or landmass for these regions.
evmar··on Britain’s output per person is now only just above that of Mississippi
I was curious about comparisons like the ones you're making between US states and EU countries and made this little app, maybe you'll find it useful!

https://evmar.github.io/states/

evmar··on Theseus: Translating Win32 to WASM
This is my second emulator, and in my first I picked a name more like that and regretted it. A thing I now appreciate about emulators is that it's common to increase scope -- like this one already supports non-wasm output, and I am tinkering with adding support for DOS executables as well, which means the name 'win2wasm' would already become obsolete!
evmar··on Theseus: Translating Win32 to WASM
Thanks a lot for this, I will put it on my list to investigate.

I've gone in circles a few times with how to think about image buffer management because I also support DirectDraw, which is designed to be backed by accelerated graphics, with operations like scaling bitblit. (Currently the Theseus implementation uses a shared "Surface" type as the backing store for both GDI Windows and DirectX Surfaces.)

It's a bit complicated by a few things. (1) DirectX surfaces can be "locked" to access as pixel buffers, so any accelerated surface indirection I guess would need to be able to copy pixels back down into emulator memory. Which I guess I could just implement. (2) There's a bunch of different modes for operations like bitblit like setting a color key for transparency that I can't implement with the canvas API, so I think I'd need to use GL shaders if I want acceleration, not just canvas.

evmar··on Theseus: Translating Win32 to WASM
Wow, thanks for the link, that is perfect timing! I submitted my post and my own feedback on the discussion: https://github.com/WebAssembly/shared-everything-threads/dis...
evmar··on Theseus: Translating Win32 to WASM
[post author] I looked into this API but wasn't sure how to make good use of it. I may have misunderstood, maybe you could help!

The doc you linked has two forms of use, sync and async.

For sync: it seems the idea is for the worker to render into an OffscreenCanvas, then postMessage an ImageBitmap created with transferToImageBitmap from worker to main thread for drawing. It seems like it would need to allocate a new bitmap for each frame. Currently Theseus puts the pixel data in shared memory and the main thread copies it out (required to create an ImageData), which at least in principle could reuse the copy buffer (though it currently doesn't), which seems better? https://github.com/evmar/theseus/blob/a5a849dbcf8046a2d1837a...

For async: in this the idea is have the worker render into an OffscreenCanvas linked to the on-screen one. But it seems to get an OffscreenCanvas in a worker, the main thread canvas must .transferControlToOffscreen() it to the worker. Under the current synchronization model[1] the only time the worker can receive a message is during startup, because the rest of the time it's deep in its own wasm call stacks. This means that if the worker needs to resize its canvas and then paint to it, it's stuck.

[1] I wrote "current" because after writing this post I learned about JSPI which might help with this.

evmar··on Theseus: Translating Win32 to WASM
Gosh, I think that means even when your code is wholly running on workers (where you would be able to use the atomic wait mentioned in the comment), it still will busy loop, doesn't it? At least it's within the allocator and not in the general implementation of Mutex... I think?
evmar··on The Third Hard Problem
One nice tool for analyzing maps as a tree is as a dominator trees. I wrote a bit about it here: https://neugierig.org/software/blog/2023/07/dominator.html
evmar··on Deterministic Fully-Static Whole-Binary Translation Without Heuristics
The translator I made is only hobbyist quality, but I just have a big table that says “if you indirect jmp to address X then the associated block is at location Y”.

This is slower than a direct jmp (which doesn’t use the table) but also indirect jumps were slower in the original program to begin with and typically don’t occur in performance-critical loops.

evmar··on Theseus, a Static Windows Emulator
Do you have any notes or other artifacts from your recompiler? I’d love to learn more.
evmar··on Theseus, a Static Windows Emulator
Yes, I agree that there is little harm in gathering too much code. I have tried out just scanning data memory for values that refer to addresses within the region marked as code and disassembling from those points, as well as scanning the instructions I traverse for any immediate values in the same range.
evmar··on Theseus, a Static Windows Emulator
Slow progress is fine, it took me like two years to get where I got! (Not that I was working on it full time or anything, but also there were just many false starts and I had no idea what I was doing...)
evmar··on Theseus, a Static Windows Emulator
Looking through Wikipedia at least, it's not exactly clear to me. They have separate pages for 'binary recompiler' and 'binary translation' that link to each other, and the latter is more about going between architectures (which is the main objective here).
evmar··on Theseus, a Static Windows Emulator
[post author] I went down some similar paths in retrowin32, though 32-bit x86 is likely easier.

I was also surprised by how much goop there is between startup and main. In retrowin32 I just implemented it all, though I wonder how much I could get away with not running it in the Theseus replace-some-parts model.

I mostly relied on my own x86 emulator, but I also implemented the thunking between 64-bit and 32-bit mode just to see how it was. It definitely was some asm but once I wrapped my head around it it wasn't so bad, check out the 'trans64' and 'trans32' snippets in https://github.com/evmar/retrowin32/blob/ffd8665795ae6c6bdd7... for I believe all of it. One reframing that helped me (after a few false starts) was to put as much code as possible in my high-level language and just use asm to bridge to it.

evmar··on Theseus, a Static Windows Emulator
It depends on what results you’re expecting! Relative to an interpreter, even the simplest unoptimized translation of that code is already significantly more efficient code at runtime.

In the post there is a godbolt link showing a compiler inlining a simple add, but a real implementation of x86 add would be much more complex.

I have read other projects where the authors put some effort into getting exactly the machine code they wanted out. For example, maybe you want the virtual regs.eax to actually exist in a machine register, and one way you might be able to convince a compiler to do that is by passing it around as a function parameter everywhere instead as a struct element. I have not investigated this myself.

evmar··on How to make Firefox builds 17% faster
I see, I might be confused by the terminology. "clobber" to me suggests intentionally trying to throw away cached results (clobbering what you have), but it sounds like you might just use it to mean builds where you don't have any existing build state already present.
evmar··on How to make Firefox builds 17% faster
I’m not too familiar with Firefox builds. Why are clobber builds common? At first glance it seems weird to add a cache around your build system vs fixing your build system.
evmar··on Ninja is a small build system with a focus on speed
To be honest, it's not clear to me why other systems are not faster. Ninja is relatively straightforward but also not too clever.

Now that I think about it, I did write more about some of the performance stuff we did here: https://aosabook.org/en/posa/ninja.html Looking back over that, I guess we did do some lower-level optimization work. I think a lot of it was just coming at it from a performance mindset.

Page 1 of 27Next →