HNHacker News
TopNewBestAskShowJobs

abainbridge

1,994 karma · joined October 5, 2012

submissionscomments
abainbridge··on Benchmarking Malloc with Doom 3
> I would have measured the time directly inside of the Doom allocator

Yeah, that'd be much better. It'd still only tell us how the allocator performed when having to share cache, memory B/W etc with Doom 3 though. Performance of similar allocations in a different application could be different. Even changing the screen resolution of Doom 3 might be enough to change the performance of the allocator. To get anywhere we need to understand what all the hardware components that alter performance are, the limits of their various resources and the penalties for running out of those resources. I normally find it is easier to remove the allocs from my program's main loop than to do this kind of analysis. In fact, I think this analysis is impossible. What happens when some background process like a virus checker decides to run? Doing less work in your program is almost always better. Moving the allocations out of the main loop normally makes the ownership easier to understand too, so helps to minimise code complexity.

abainbridge··on Show HN: Micro LZMA decoder (x86 assembly code golf)
I guess it would be a nice option, if I was trying to compress something that was already small, so the extra 2kb was significant.
abainbridge··on Show HN: Micro LZMA decoder (x86 assembly code golf)
I wonder how this compares to the one in UPX.
abainbridge··on Benchmarking Malloc with Doom 3
"My first attempt at a benchmark involved allocating and freeing blocks of random size. Twitter friends correctly scolded me and said that's not good enough. I need real data with real allocation patterns and sizes."

"The goal is to create a "journal" of memory operations. It should record malloc and free operations with their inputs and outputs. Then the journal can be replayed with different allocators to compare performance."

I don't think that is sufficient either. You need to mix the malloc/free with the other work, because the other work knocks your allocator data structures out of cache and pollutes your TLB and branch predictor. And loads up the dram interface (eg by the gpu fetching data from it). Etc etc.

abainbridge··on In praise of the humble Sheffield stand
We're talking about this type of pipe cutter, right?

https://www.screwfix.com/p/irwin-record-handicutter-15-45mm-...

They're much cheaper than an angle grinder, and fit in your pocket. I've used one to cut through the barrel of a d-lock with ease. I don't buy d-locks with cylindrical barrels any more. I learned this trick when thieves stole my friends bike this way :-(

abainbridge··on Modern programming languages require generics
Good point about #ifdef. I'm not sure what other tools are more appropriate though. I went through a phase of moving target specific code into separate files instead of using #ifdef. But that makes the build system more complex and makes it harder to eyeball the difference between two targets' implementation of the same function. This is just an intrinsically complex problem and not one I blame the preprocessor for. As always, aiming for minimum total complexity is the goal. #ifdef is a useful tool in that persuit.

That said, I'd be interested to hear about other solutions that are better. I guess Zig's comptime is a candidate, but I see that as only a superficial improvement (not a criticism - I can't think of anything it could do better).

abainbridge··on Modern programming languages require generics
I think the case against the C preprocessor is overblown. What's the worst problem C macros have caused you? I've been programming C for 27 years and I can't really think of them causing any bugs in my programs. The worst thing for me is that #defines don't show up in the debug info.
abainbridge··on Earn $200K by fuzzing for a weekend: Part 1
Can someone explain the CVS 2021 46102 bug? Various sources tell me it was an integer overflow in some Rust. The faulty line was:

    let addr = (sym.st_value + refd_pa) as u64;
I guess, the + is evaluated using 32-bit arithmetic and then it is cast to a u64, and thus overflow is possible? And in release mode, Rust doesn't trap integer overflow.

Shouldn't something as critical as the EBF compiler be trapping integer overflow?

abainbridge··on Ask HN: I'm interested in so many disciplines, but what can I do with that?
I can't even read light fiction. I read the first Harry Potter book once. To give a feeling of how hard reading is for me: it took me longer than working out how car understeer and oversteer physics work from first principles and making a simulation in C++ [1].

[1] Admittedly crap 2D car sim https://github.com/abainbridge/car_sim

abainbridge··on Ask HN: I'm interested in so many disciplines, but what can I do with that?
I feel that you can assign a separate IQ-like score to many different components of human intelligence. I score badly on book comprehension relative to other components.

I find it easier to concentrate for an hour or two on a video than a book. I'm probably slightly dyslexic or something. At university I made sure I never missed a lecture but hardly ever read a book. I did above averagely well.

This is a disadvantage though. I have friends who can read books in a couple of hours and absorb the content. I'd have to sit through more hours of video to get the same content, even watching the videos at 2x speed.

Wild speculation: being intellectual used to require being good at book comprehension. Maybe only 10% of the population are good at that. Maybe some of the remaining 90% are still clever. If so, video content could allow those book-comprehension-limited people to become intellectual. That'd be nice.

abainbridge··on Swapping two Numbers without Temporary Variables
That uses 4 registers and a stack!
abainbridge··on Addressing Criticism of RISC-V Microprocessors
I'm guess that your assembly code is RISC-V with the Zba extension. Is the non-Zba version worse than Arm64?

Compiling your function with Godbolt, I get:

  RISC-V (no Zba) Clang - 7 instructions - https://godbolt.org/z/7znnrzxKq
  Arm64 Clang -           7 instructions - https://godbolt.org/z/Trv8scxad
Annoyingly I can't see the code size for the Arm64 case because no output is generated if I tick the "Compile to binary" option in "Output". I have to use GCC instead:

  RISC-V (no Zba) Clang - 20 bytes - https://godbolt.org/z/eWfPaorcj
  Arm64 GCC             - 24 bytes - https://godbolt.org/z/bzsPzov5h
abainbridge··on Lapce – Fast open-source code editor
> if you resize it you move the whole text

What's wrong with moving the whole text?

abainbridge··on Bugs in Hello World
Came here to say the same. Zig gets it right. https://ziglang.org/documentation/master/#Hello-World
abainbridge··on Algorithms for Modern Hardware
Actually, I thought of another thing: when optimizing some module, I often find myself making a test harness that puts a lot of load on the module in question. Then I tweak it until it is fast. The problem is that I've optimized it in a state where the data is likely to be cached and the branch predictor has learned how the branches go. But then when you put that module back into a full program, the rest of the program might have overwritten our module's state in the cache and branch predictor, so the performance we get is much lower. In which case, Getting Accurate Results requires flushing the cache and branch prediction state in our test harness. I'd be interested to see some ideas on how to do that.
abainbridge··on Algorithms for Modern Hardware
Yeah, I should have said I thought the Getting Accurate Results section was excellent and only missing that one thing!
abainbridge··on Algorithms for Modern Hardware
I think the profiling section could benefit from a section describing the sort of profiling noise described in this video: https://youtu.be/r-TLSBdHe1A?t=436

The noise source described is that changes to the code typically alter alignments of code and/or data and that is often the cause of the performance difference measured, not the change to the algorithm. The reason this is noise is because unrelated changes to the code might make it change again.

IIRC, the speaker suggests improving compiler/profiler tools to randomise these alignments and benchmark many times.

abainbridge··on Game Boy Wordle clone: How to compress 12972 five-letter words to 17871 bytes
How would that work? If I want to store "apple", "apply", "apron" the bytes in memory would be "appleyron". How would I know where the second word ends?
abainbridge··on U.S. corn-based ethanol worse for the climate than gasoline
I thought Ethanol has a very high octane rating (~110-120 RON) and therefore would reduce knocking compared to regular 85-95 RON fuels. I'm sure it has plenty of other problems though.
abainbridge··on Nim vs Rust Benchmarks
OK, I did "for i in {1..10}; do o//examples/rusage.com o//examples/hello.com; done" and the lowest wall time I got was 1,140us, which is a lot more than your 90us for some reason. There's probably something wrong with my setup. Maybe it takes my machine that long to come out of some P-state or C-state sleep mode I don't understand.

Then for my musl-gcc hello world, I did, "for i in {1..10}; do o//examples/rusage.com ../hello" and the lowest wall time I got was 120us. But most of the runs were closer to 350us.

So, my attempt to answer your, "Why does hello world take 1.8ms to run?" question is: a) they timed it with "time" which measures lots of overhead, b) there's something wrong with their setup that means they are mostly timing the time taken for a CPU core to wake up.

abainbridge··on Nim vs Rust Benchmarks
How would you time this? I tried a simple C hello world and got a similar ~2ms run time. But I was just running "time ./hello" at a Bash prompt, which a) only has 1ms resolution and b) probably measures some shell overhead. Running it under "strace -tt" so I could see where the slow bits were made it 10x slower :-(

Edit: FWIW, building with "gcc foo.c -o foo -O2 -s -Wall --static" was slightly faster than without the --static. And "musl-gcc foo.c -o foo -O2 -s -Wall --static" was slightly better again. Maybe 0.5ms improvement in total. I'm on Ubuntu 20 x64, gcc 9, Intel Xeon W-2133.

abainbridge··on Nim vs Rust Benchmarks
I wouldn't have down voted this comment, but I guess people are objecting because using C as the backend doesn't really tell us anything about performance. If the C code the Nim compiler generates is inefficient, the resulting program will still be slow. I imagine it is difficult to convert a Nim program into efficient C.
abainbridge··on Do things, tell people (2012)
> I'm not sure what nomenclature I could use besides "data format"

Yeah, I don't know either. I see that the Wikipedia page for JSON describes it as a "data interchange format". Maybe that helps. I'd probably side-step the problem by having your home page start with "Like Protocol Buffers but more secure", followed immediately by the example you just gave me.

abainbridge··on Do things, tell people (2012)
I skimmed your site and didn't really get it. Maybe that's what everyone else does. In case it helps, here are the thoughts that occurred to me, in the order they occurred.

0. To me "data format" makes me think of things like PNG, or the DWARF debug symbols format, or the MPEG transport stream, not general purpose text formats like JSON and XML (I'm not a web person).

1. A secure data format. OK, weird. I thought it was always the programs/libraries that dealt with the data formats that were guilty of the security bugs, not the format itself.

2. All the bullet points under simple and efficient are already true for all the data formats I know/use/care about. So already I've dismissed your project as interesting - ie the "simple" and "efficient" are already solved and I don't believe "secure" is a real problem. But I'm aware I might be jumping to conclusions, so I keep reading.

3. I skipped straight to "Security - Protecting your data". I thought the sentence explaining why security matters was superfluous - your audience already knows why security matters.

4. "The existing ad-hoc data formats are too loosely defined to be secure, and can't be fixed because they're not versioned". OK, this looks like the meat. I click the link.

5. "There are many vectors that attackers could take advantage of when they control the data your system is receiving, the most common of which are induced data loss, field omission, key collisions, and exploitation of algorithmic complexity". If the data is from an attacker, data loss and field omission sound like good things - I don't want their data or fields because they are an attacker.

6. '"change user" command with a group of admin\U+D800'. I'm still confused. It still sounds like the "admin\U+D800" string is processed by my program. Your data format doesn't know whether that is a valid string or not. Telling your data format is no easier than telling my program. Oh! Maybe it is. Because you only need to specify it in the data format, not in every program/library that implements the format. Is this the point of the system?

abainbridge··on Who wrote this shit?
More than once I've googled how to do something new and found the answer on Stackoverflow, written by me. I haven't even written that many answers on Stackoverflow.
abainbridge··on How to design a house to last 1000 years
York Minster in England is mostly 1000 years old. There are some nice stories about how it handled a serious fire in 1984 here, (starting about 18 minutes in) https://www.bbc.co.uk/sounds/play/m0007pws. eg, the 2000 year old Roman drainage system got used for the first time in 1000 years because lots of water got inside the building from the fire hoses.

When rebuilding it, they had to decide whether to remake it in the same way it was originally, or to use a modern approach. They decided to use oak beams again because of lack of evidence about what happens to steel structures after hundreds of years. But then they couldn't find any oak trees big enough.

abainbridge··on QOI – The “Quite OK Image Format” for fast, lossless image compression
XPM and gzip are still not that simple. QOI is much simpler.
abainbridge··on South Africa’s omicron coronavirus outbreak subsides as fast as it grew
Comparing SA with itself, omicron looks less severe than delta.
abainbridge··on Three loud climate warning signals
That would only work on some. The skeptic I know acknowledges that the world is getting warmer but doesn't believe this is a problem. I do believe it is a problem and find it hard to convince him.
abainbridge··on Lapce – Fast and Powerful Code Editor written in Rust
It could be very extensible without being written in web technologies. For example, an IDE I worked on was written in C++ with Qt and had an embedded Python interpreter to allow extensions to be written easily. It also provided a .dll/.so extension mechanism for stuff that wanted more performance than Python could give.

It used 10-20 megs of RAM and loaded in about 2 seconds on the computers of the era (about 10-15 years ago).

It wasn't very good, but that's beside the point :-) It wasn't bad because of the technology stack it was built from.

← PreviousPage 5 of 18Next →