HNHacker News
TopNewBestAskShowJobs

jstimpfle

4,223 karma · joined May 21, 2015

http://jstimpfle.de
submissionscomments
jstimpfle··on Fast and Hard Code
How do you judge "better" if you don't understand the LLM output?
jstimpfle··on The Two Factions of C++ (2024)
Yes -- I can see good use for runtimes. For example, compiler can autogenerate good runtime error messages. Thinking about it, debuggers make use of reflection. Debuginfo formats have some kind of reflection built in.
jstimpfle··on The Two Factions of C++ (2024)
Reflection is a joke. De/serializing arbitrary C++ structs is ill-defined. When you need serialization, even lots of it (e.g. silly JSON), I think you're better off writing your own framework where you can be clear about data formats and transformation rules.
jstimpfle··on To save C, we must save ABI (2022)
I challenge you to find one random person making that claim and to present it with a straight face. What kind of ghosts are you fighting?
jstimpfle··on To save C, we must save ABI (2022)
Who uses Ada95 or whatever? You are fighting strawmans, nobody has made the claims you imply. My personal opinion is just that low-level access is essential to make interesting and performant programs. Object-type fluff doesn't help with that, it's getting in the way.
jstimpfle··on To save C, we must save ABI (2022)
Casting integers to pointers is implementation-defined as far as I know. Even if weren't, I'm not convinced that you have to interpret C's address space as flat just because it is finite or because pointers are representable as integers. In any case, machine's address spaces are flat (the physical memory mapped into them not so much), and working with real machines is what I'm interested in.
jstimpfle··on To save C, we must save ABI (2022)
Mind you, the "you can use this other way as you see fit" idea often isn't practical (like combining GC'ed and non-GC'ed parts). You generally want a whole codebase to be structured according to shared idioms. Otherwise the interfacing cost becomes too high.

I have doubts that you can program easily in a C-style way in C# without adding lots of annotations everywhere in many places. But don't know, maybe I'm wrong, I did a search for a simple C-style arena allocator in C#, and it looked acceptable, it was quite close. The most annoying thing was maybe keyword boilerplate.

jstimpfle··on To save C, we must save ABI (2022)
It isn't actually that flat in the spec, though modern machines' address spaces are. So in a sense it is merely an accident of a specific implementation that you can smash stacks.
jstimpfle··on To save C, we must save ABI (2022)
But I assume that both on compiler and on machine, the evaluation is still conforming to the semantics of the C abstract machine?
jstimpfle··on To save C, we must save ABI (2022)
So do you want to "rewrite" some C code in C++ to think you made a point? I think you should do C# or Java.

What about you do xxHash? Should be quite basic, not a lot of complicated structures. https://github.com/Cyan4973/xxHash/blob/dev/xxhash.h

Or what about you do an audio or video codec? Or an operating system?

Not going to paste any of my own code, because any non-trivial stuff is hundreds to thousands of lines. But one more example (that I recently did myself): Create a block allocator (power of two blocks) with bookkeeping in shadow memory (administered in individually committed zones representing virtual memory regions of 64 MB (2^26)). Any used memory has bookkeeping support for being sub-allocated at any and all levels up from 64 KB (2^16) to 64 MB (2^26), and even higher (by joining committed regions). Individual blocks are collected (using intrinsic linking, because no memory allocation) in a hierarchy of pools of same-sized chunks that have the same parent, and can be recursively sub-allocated on any smaller chosen power-of-2 level, and finally be consumed in linear fashion (arenas). Blocks are pooled with a moderate retain policy (watermark system) to allow subsystems to almost completely avoid any system calls and avoid inter-thread synchronisation. The memory overhead must be below 1% even though it's totally flexible (as said has metadata for all levels from 64 KB up).

The bookkeeping should function on 32-bit systems (small virtual space, occupancy range from megabytes to 3 GB) as well 64-bit systems (2^48-2^57 bytes of virtual address space, occupancy range from megabytes to hundreds of gigabytes) with reasonable overhead compared to actual usage.

This requires intrusively linked lists, occupancy bitmasks, bit-counting and bit-prefix counting, OS syscall access (virtual memory), pointer arithmetic (alignment needed to address shadow bookkeeping memory) and thread synchronisation. The reference code is >> 95% pure ISO C++11 (could be C99 with few changes), with a little platform code glued in. It works on Windows but it could be ported to Linux in a few hours. It supports a mostly-immediate-mode GUI with hundreds of thousands (maybe millions?) of small variable-sized allocations per second. Allocation has almost completely disappeared from the CPU profile, well below 1% of CPU usage.

jstimpfle··on To save C, we must save ABI (2022)
You are again misreading even the most clearly put statement. Compared to e.g. Javascript, C is "closer" to the hardware, gives you "more control" of it. It would be completely ridiculous to deny this fact.

And if you move to e.g. C# / Java or similar, if you squint, and you try to be a smart-arse, then you could deny that C is closer to the hardware than C#, because C# probably has everything you need to control it, to the same degree that C allows you to. But if you work in these languages for a while, and look at the code that you ended up producing, then again you will absolutely find that it would be ridiculous to not admit that C gives you better control.

And you could even extend this to Rust, because the language encourages you to use high-level prefabricated components. It discourages you from doing low-level things, at least a little bit I think (I'm not a Rust user).

I think what you are doing all the time, is you are being a smart-arse, nothing else. What interesting low-level performant things have you actually programmed lately?

jstimpfle··on To save C, we must save ABI (2022)
Many other languages only have one compiler available to start with.

Each additional compiler supported by a project means variance in functionality and thus additional work for the project. That work could make the codebase more robust. Or it could be a ton of useless work. Or anything in between. Depends on the context of the project.

jstimpfle··on To save C, we must save ABI (2022)
_You_ do that. All the time. And then you fight these strawmans.
jstimpfle··on GNU Hurd News 2026-Q2
Are you saying there is cache thrashing because callers often sercice rheid own requests themselves? If you don't want to service requests in the same thread, doesnt it mean you have to spend entire core(s) for running the kernel?
jstimpfle··on The PImpl idiom and the C++26 std:indirect type
> I at least showed you cppreference so you can look up the data structures and their guarantees.

Why do you show this to me??? Don't you think I know it?

> But you did blame the STL for concurrency bugs so there must have been something.

I explained the problem at your request: I pointed out that this was a bug I introduced myself, but yes, I blamed it on STL (and its complexity). Again, it was abstractly a "concurrency" problem, and it did manifest when using multiple threads, but the problem was not due to missing mutex nor reference counting. Instead it was because of a kind of iterator invalidation that I had not expected at the time of banging out some shitty iterator code. The uniform iterator abstraction made it arguably way easier to miss.

> You used C++'s dequeue, wouldn't that be boilerplate by this bizarre definition?

yes absolutely, std::deque is a super bad offender, in many ways. I advise against using it. More than against using STL in general, although I don't recommend that either.

> I think you might have misunderstood that the reference counting is for anything returned from a data structure so that it can see that something is being used and not modify it. The reference counts of the returned object are actually pointers to the internal reference counts in the data structure, like checking out a library book.

No I have not misunderstood anything. Again you're coming back to your arrogant pattern. There are many ways to implement reference counting. When doing it manually instead of with e.g. std::shared_ptr, it's quite common to embed the count inside the object, not make it a separate allocation.

> This is not how I would do a queue though and not how the queues I linked work. They copy data in and out and are best used for small data. Large amounts of data can be handled in a different way by a different structure.

There are many types of queues. There are queues that buffer two elements, there are queues that buffer millions of elements. There are queues that get persisted (like a database). There are queues that are ephemeral. There are queues that have multiple produces and/or consumers, there are queues with only a single producer/consumer. There are queues that get locked. There are queues that get accessed with atomics only. There are queues that store elements directly. There are queues that store elements buffered in chunks or packets... "Queue" typically implies FIFO but not always.

> The C style allocation of structs to pointers then allocation of the underlying data is two allocations and double indirection.

But I rarely don't do that. And that's not implied by "C style" at all. And importantly, structure (pointer indirection) doesn't imply allocation strategy.

> Well.. we all get bit by standard library assumptions from time to time and need to read the docs, but it just isn't a concurrency problem with the STL.

I have explained the issue at length. So please stop repeating made-up contradictions.

> Claims without evidence unfortunately. The fast concurrent queues I linked are great and using destructors to keep track of reference counts is great. Both are minimal.

Claims without evidence unfortunately... Except, it's quite evident that there is a lot of code in them and it's hard to find out how anything works because of that. How would I even evaluate if the queue is doing what I need? That queue functionality implemented here should probably be a tenth of that code (!).

> I would say inserting resource management manually into every function is boilerplate.

Good, because I don't do that at all. And I criticize that RAII is a system that sneaks in resource management _implicitly_ everywhere, which is not visible in the source code. That's why I prefer C-style: making it explicit, allowing me to find the optimal structure that avoids unnecessary fluff in the first place.

jstimpfle··on The PImpl idiom and the C++26 std:indirect type
Actually, not half a dozen, more like hundreds if you count all the tiny ad-hoc stuff too (many of them similar or the same).
jstimpfle··on Qwen3.8-Max: A New Bar for Coding and Cowork
At this point, what isn't much of an estimation anymore is that you are here to explain things that nobody's asked for, and like to assume people around you have been waiting for your pearls of wisdom. You're unable to realize when this isn't so. You're so full of yourself that you don't notice.

Also, in this particular thread, you started by wrongly making a correction of something that was clearly not an error, nor a wrong use of words, not even a misspelling. The problem is you started to post a correction before realizing that you didn't read right. That happens when one is more eager to boast one's own greatness than one is interested in the topic at hand. The result is that in this thread, you were clearly, as a rule of thumb, well, 100% wrong.

jstimpfle··on The PImpl idiom and the C++26 std:indirect type
> I gave you great information

"Great" is quite debatable. In any case, nothing I hadn't already known.

> You aren't going to notice all your concurrency bugs without threads.

True, but my problem was neither proper locking / thread safety, nor reference counting.

You still felt the need to explain to me because you don't realize the problem isn't that I don't understand what you say. The problem is that you don't understand / don't want to accept what I say, and you prefer assuming I'm talking out of my ass.

> Lots of people get a lot of good out of them.

Well if they don't want to create and understand their own but instead prefer to invite tons of unnecessary boilerplate to the point where you can't find the actual functionality -- good for them.

> You might want to benchmark and test those bad boys thoroughly if you think you can hold a raw pointer into a data structure that can change from other threads

I DO NOT THINK THAT. Why do you keep implying that my thinking is wrong? That is so arrogant of you.

Reference counting (how you keep something alive) is completely orthogonal to the queue's functionality. In my case, the queue was used as a "global" kind of object, so no reference counting needed.

> Also don't forget that allocations can lock and that your double allocations of the struct and data in a data structure can amplify that.

In general I avoid unnecessary allocations, where did I imply making "double allocations"? What I argued is that indirection may not be as bad as you think, may in fact be the correct way to make your program both more maintainable and more performant.

I try to organize memory allocation upfront to keep memory local to subsystems, which reduces or avoids contention in many cases (for example there might be only a single thread doing allocations for a subsystem at a time).

> Thanks, but I haven't made the same assumptions about raw pointers in concurrent data structures then blamed the STL, so I haven't had the bugs that you're talking about here

You're arguing all the time for just buying into stuff as a cargo cult, I'm only trying to describe how much weight all this ceremony introduces, which makes it more painful to maintain, makes it more likely to introduce bugs, and harder to find bugs. Don't explain basic C++ stuff to me. I understand it. What I'm saying is that this is not the best way to write things at all. There's a lot of "abstraction" slop that brings more downsides than upsides.

But I'm sure you never run into this type of problem... Good for you!

jstimpfle··on The PImpl idiom and the C++26 std:indirect type
Serious question, are you an AI programmed to be annoying?

> This isn't a problem with the standard library because a std::dequeue or any other core data structure doesn't make any promises about concurrency.

Dude, I KNOW I need to handle concurrency myself. But I'd contend the point that it isn't a problem with the STL: It is a bug (that I introduced myself) that I had to deal with because of complexity, or rather because non-obvious behaviour, because bullshit boilerplate.

> If you have an underlying data structure that is being used from multiple threads, you can't hold on to raw pointers into the data structure. There is no way for other threads to know that it can't be changed, moved, freed or invalidated.

This is totally irrelevant because if you paid attention, the problem wasn't even threads. It was concurrency, more abstractly. Iterator invalidation based on the "manifested" order of execution.

But anyway, you want to jump to reference counting. I'd say you can absolutely hold on to raw pointers from multiple threads, it entirely depends on what you do. If the threads have unpredictable lifetimes, then yes, some form of reference counting is indicated.

But when you know that isn't the case, then it isn't the case and you probably don't need reference counting.

> I hope it isn't lost on you that the reference counting approach is much easier to do with a destructor, since the reference count can be incremented before it is given to you from the API and decremented automatically when it goes out of scope.

Except when you're passing around stuff and have to duplicate or move references, and have to use APIs that receive pre-incremented or un-incremented pointers. In some cases your data structures might even be so messy that you end up with cycles.

I have my scars from making my own COM pointer classes with copy and move semantics, and also from using "official" COM pointer classes. After a couple of iterations I've decided to cut all the boilerplate and C++ ceremony that doesn't do anything, and get rid of ugly method wrappers that are a pain to step through in the debugger, and stopped clinging to a cargo cult which simply leaves you with harder to detect bugs.

You heard right, I'm back to completely manual reference counting (and only counting where I _have_ to), somehow the code is much shorter and easily understandable, I got back control over what happens. Have been able to keep atomic ops at a minimum, with RAII superfluous ops can happen easily. (Remember Chromium's 25000 copies per keystroke bug?) And there has only been a single instance where I introduced a leak, that was immediately pointed out by the D3D11 debug layer. I'm doing this approach for my second project already and have found it to work great.

There is no solution except good understanding of what you do, and good code structure that expresses this understanding. Generic "RAII" type understanding is rarely helpful IMO, you give up control and sometimes end up throwing hands in the AIIR and hope it will not break.

> https://github.com/cameron314/concurrentqueue

Thanks for the pointers to what is probably 5K lines of C++ boilerplate. But I have written half a dozen concurrent queues myself, locking and lock-free ones. Some in less than a hundred lines. Also one in ~2K lines, that was for a longer-term project where the queue needs to safely persist to disk every couple of milliseconds, while ingesting millions of messages per second and billions of bytes per second (was hitting the ~2GB/s that I could get out of my flash drive).

If you want an approachable source that leaves out the fluff, I'd recommend 1024cores by Dmitry Vyukov (only issue is formatting).

jstimpfle··on The PImpl idiom and the C++26 std:indirect type
You have repeatedly proven, and continue to do so, also by way of your exchanges with other commenters, that you're not asking out of curiosity. After all the previous comments we've exchanged, your line of asking was, "Which part of the STL are you expecting to be thread safe?" i.e. you were continuing to assume that I was somehow naive or uneducated. I did not assume anything to be thread safe in the way you imply (I used a simple mutex based approach to protect accesses).

If you'd been asking in good faith, the question would have been, "how did the concurrency problems look like"?

To which I'm going to answer, one of the bugs I hit was due to unexpected invalidation of a std::deque iterator. This came from being mislead to use std::deque as a quick & dirty implementation of a producer-consumer queue, and keeping iterators to track the read and write positions. Almost nobody has actually used std::deque (I hadn't either) but there is a common understanding (perhaps misunderstanding) that it is something like a chunk-queue. That vague understanding led me to believe that I can (and should, to avoid O(n) random access ) keep iterators after write operations. And using them that way did work for quite some time, I only hit confusing issues later.

(Actually random access is specified to be O(1) but this is even less widely known and makes std::deque a quite arcane data structure).

The problem with an abstract iterator interface here is that it doesn't help understanding what std::deque actually is. In case of std::deque, keeping read and write cursors works mostly fine, but it stops working (for example) if the read cursor pointed to the current end (was equal to deque::end) and the deque gets an append, which will invalidate the old end (read) cursor.

This is a good example of the complexity we have to deal with if we don't want to write a simple straightforward solution from scratch (chunk list) but instead code against something that we don't understand well. Not trying to use the STL but instead doing straightforward low level code would have made potential pitfalls more clear, and would have made bug search easier. It would have required less work to get the code to a working and maintainable state.

Another problem with std::deque is that the sizes of the chunks are not specified. They vary wildly between implementations, such that you can in practice get no performance guarantees from using std::deque, unless committing to a specific STL implementation (which is rarely practical). In fact, it is not even specified that deque uses something like chunks internally. It's too abstract to be useful.

jstimpfle··on The PImpl idiom and the C++26 std:indirect type
Try harder, Sherlock. You're very close to proving that I assumed the STL was "thread safe". You've almost got me.
jstimpfle··on The PImpl idiom and the C++26 std:indirect type
Why don't you read again what I said instead of continuing to make implications?
jstimpfle··on The PImpl idiom and the C++26 std:indirect type
Where did I imply that I am expecting any part of the STL to be thread safe?
jstimpfle··on Claude Code uses Bun written in Rust now
You are insufferable. Get lost.
jstimpfle··on A shell colon does nothing. Use it anyway
Javascript
jstimpfle··on The PImpl idiom and the C++26 std:indirect type
I've found that RAII and the stuff you have to buy into in order to use RAII come with more downsides than upsides once you scale beyond high level programs that try to get done a lot with very few lines.

> how plain C helps with this?

By staying out of the way and providing everything of what you actually need in the end. That is assuming a detail oriented approach where you deeply think about, and want to be flexible about, the organization of what your program should do. As programs grow into large architectures, and as programs get more performance conscious, they also get more detail oriented like that, and they tend to opt out of unflexible high level language features.

jstimpfle··on The PImpl idiom and the C++26 std:indirect type
I understand C++ as well, or better, than most (or all) of my peers, and certainly betters than people on here thinking they need to explain to me how RAII works. Do you want to argue that C++ RAII / objects stuff isn't complex and doesn't put considerably restrictions on how you design your app, then maybe you should reconsider.

I would argue that if cleaning up resources properly is among the hard problems, or among the most error-prone problems in your code, then maybe you're problems aren't that hard or complex after all.

I'm currently working on a distributed caching system and on real time voxel geometry boolean operations simulation (on GPU), both on the scale of >= 10^9. Is that "complex" enough? Both are done in C++. C++ helps exactly 0 in achieving any of these things (as opposed to using plain C), well the one help is I don't have to type 'struct' all the time.

In fact, in one of these projects I was pushed to use STL initially. I'm now working on getting rid of the last of them because we have had concurrency bugs and performance problems from using them. The code was not obvious and using STL containers (std::deque is very bad specifically) meant the actual runtime characteristics depend on which STL implementation is being compiled in. It would have been easier to just do straightforward obvious manual code.

jstimpfle··on Understanding the Odin programming language
You don't understand even the first things of what I say. That's because you apply beginner level concepts and understand to argue, and are not looking for nuance or deeper understanding at all.
jstimpfle··on Understanding the Odin programming language
You don't get it. That's because you still don't understand some basic things about the language. There is no method call here on a pointer expression. There is only a call on a value expression.
jstimpfle··on The PImpl idiom and the C++26 std:indirect type
Well, the best way to prevent mistakes is to make everything super complicated and damn hard to do. The best way to prevent mistakes is to approach it to do the essential stuff and avoid the fluffy stuff.
← PreviousPage 2 of 34Next →