4,852 karma · joined August 4, 2018
The screenshots in the Steam page do look impressive though! As an outsider, I could mistake this for a recent Civ game. Congrats on making it here!
For extracting the fractal from the residual stream, did I understand it correctly as follows: You repeatedly sample the transformer, each time recording the actual internal state of the HMM and the (higher-dimensional) residual stream. Then you perform a linear regression to obtain a projection matrix from residual stream vector to HMM state vector.
If so, then doesn't that risk "finding" something that isn't necessarily there? While I think/agree that the structure of the mixed state representation is obviously represented in the transformer in this case, in general I don't think that, strictly speaking, finding a particular kind of structure when projecting transformer "state" into known world "state" is proof that the transformer models the world states and its beliefs about the world states in that same way. Think "correlation is not causation". Maybe this is splitting hairs (because, in effect, what does it matter how exactly the transformer "works" when we can "see" the expected mixed state structure inside it), but I am slightly concerned that we introduce our knowledge of the world through the linear regression.
Like, consider a world with two indistinguishable states (among others), and a predictor that (noisily) models those two with just one equivalent state. Wouldn't the linear regression/projection of predictor states into world states risk "discovering" the two world states in the predictor, which don't actually exist there in isolation at all?
Again, I'm not doubting the conclusions/explanation of how, in the article, that transformer models that world. I am only hypothesizing that, for more complex examples with more "messy" worlds, looking for the best projection into the known world states is dangerous: It presupposes that the world states form a true subspace of the residual stream states (or equivalent).
Would be happy to be convinced that there is something deeper that I'm missing here. :)
Comma operator I would agree is really only used for shenanigans. But operator= is as crucial as the copy/move constructors for ownership, so I'm not sure how you picture this...
> apparently in C++ a structure cannot really have zero length
Yes it can: https://en.cppreference.com/w/cpp/language/attributes/no_uni...
> I think it should be possible for your own structure or union type to overload any operators that are not already defined
I see that you are trying to solve the problem of "when reading this code, I don't know if these are the built-in operators or custom ones". I would be interested in learning about situations in which this was a problem for you, because I have not made that experience myself. From my perspective, you need to know the types of x and y to know in the first place which operations are possible - default or custom - so this rule does not really buy you much. But again, I don't think I understand your problem well enough.
FWIW, you can define a custom operator-> and operator* (they have to return pointers, or types that themselves have operator->/operator* which are then applied recursively). I am not convinced that mixing pointer-related semantics into e.g. object assignment is a sufficiently elegant solution from a conceptual perspective though.
Sure, I agree, you should carefully evaluate which language best serves your needs, and what tools in a language you should use or stay away from. But some game dev ranting about which C++ features they consider useless (without even articulating their background or constraints) is worthless imo. Especially if this is the credit they give to their name:
> After over 20 years working with C/C++ I finally got clear idea how header files need to be organized.
[1] Yes, that FAQ parody is now very outdated, seeing as it precedes even C++11. I can believe that it was on-point at the time of writing, but C++ grew up since then.
Is there anyone who is seriously contesting the "it is intended to cure/prevent a disease, therefore it is a drug, therefore it needs FDA approval to be sold legally" line of reasoning?
More humoristically: https://xkcd.com/2475/ and https://xkcd.com/2530/
Less humoristically: https://en.wikipedia.org/wiki/Thalidomide
However, I don't know why you are comparing a single billionaire vs a single X kid household. Like, the number of each (or even of private jets) are not even _remotely_ in the same ballpark. Which is why "number of kids" is not at all a strange place to focus on environmental impact, but "billionaire lifestyle choices" is.
I would like to also highlight this article by Joel Spolsky: https://www.joelonsoftware.com/2006/04/11/the-development-ab... (It's concerned with the everything else in a company, but it applies equally to the senior vs junior responsibilities imo).
If you are under the impression that "writing good code" is the be-all-end-all of being a software engineer, and that being a senior means you get to enjoy it more, then you're probably working in a company that is really good at maintaining this part of the abstraction layer.
That's because the more senior you get, the more you inevitably get to be part of the system that allows more junior devs to keep their mind on writing good code. That means dealing with the meetings, timelines, uncertainties, bug triage, maintenance, documentation, compliance and customer/expectation management. So they don't have to.
Simpler example: Dreams.
(The other models being only partially able to source good references is unsurprising/"unfair" on a technical level, but that's not relevant for assessing their safety.)
"We assume 64 bit overflow is not going to happen because nobody can store that many bytes" could be valid if the existence of those bytes was required for reaching this code. But if user input can lead to UB being triggered here, fixing the code is indeed prudent, even if everyone were fully convinced that current compilers are not outsmarting themselves.
The effect on the actual inlining optimization could be anything between "none", "the compiler might weigh it differently purely due to linkage" and "it is actually taken as a hint", plus a similar set of considerations for link-time optimization. A much stronger inlining signal to the compiler (specifically for functions that are only used in the file they are defined in) is to define them in an anonymous namespace.
The original intent behind marking the function `inline` could reasonably have been an attempt at actually achieving inlining. A measurable benefit is not obvious (but also not impossible).
Gimp and stb_image_resize results are subtly different from others, even for Bilinear. Looks like they do convert source sRGB image to linear, do filtering there, and convert back to sRGB.
You can definitely see that as a problem with bright lines (e.g. collar highlight in einar*.png) being darker in the upscaled Blender versions.[0]: https://aras-p.info/img/misc/upsample_filter_comp_2024/
- If there's a flaw in your application code that causes a crash (as is the motivating example in the essay), then restoring the entire program into the state it was in just before the crash happened would just cause it to crash again ad infinitum. Sure, this model helps against "my VM instance got preempted", but that's a pretty different category of "crash" (and also notably unrelated to supervision trees).
- "External" state (like an API endpoint being down/returning gibberish) can be part of the reason why your program got into a bad state. In fact, that's disproportionately likely, since external "weirdness" is comparatively hard to cover exhaustively in tests. In such a situation, the suggested computation model would never be able to recover even when restarted, because it would forever retain the (bad) API response. Effectively, this is just caching all the non-pure effects of your computation, and we all know about cache invalidation being a hard problem...
[1] I did find out so far that my WiFi setup occasionally "clumps" packets, causing like 10 packets to hit a given ESP32 at an instant (instead of a few ms apart) - not great, but should not be disastrous. However, this seems to cause the ESP32 WiFi stack to just slow to a crawl: It responds to pings much slower (like, in 100+ ms range) (to my surprise it actually responds to PING requests out of the box in the first place...) and/or doesn't really process any more packets in general if I continue sending at the same rate as normal. But backing off on the packet stream usually gets it back on track, strangely enough. This also happens if I do nothing in the main loop except clear packets as they come in, so it's not in my code.
Goodbye!
Again, I am not objecting to the general advice you give, but maybe you can take a step back and appreciate that not every code base and C++ use case falls into the range of what you have seen and interacted with so far. There is no doubt that a vast majority of C++ code bases would benefit immensely from improving compile-time encapsulation, but that is not the be-all-end-all of solving long compilation times, and it does not give you grounds for dismissing concerns of people who have done that and still face different problems.
> You instantiate what you need, you move your template code into submodules and out of interface headers, and you're set.
You seem to have completely missed my point. I implore you to consider for a second that I might be familiar with what you are trying to tell me, and that there are indeed complications beyond the basic techniques you advocate for. Let me illustrate with an example.
https://godbolt.org/z/5j43WrM68
Here we have a very simple class that uses a few std library types. It is a template, but say that we know that it is only valid to use it with the shown basic arithmetic types (note 1). The class implementation and according explicit template instantiation definitions have therefore been moved to a separate TU (not shown) and we only have the shown code in the header. This is a simple application of those "basic C++ techniques" you mention.
What happens to the compilation time if we enable the explicit template instantiation declarations in the header? Measuring on godbolt is noisy, but by repeatedly changing the source file (e.g. just adding spaces) you can quickly get a bunch of measurements. The lowest value of 10-20 measurements is going to be a reasonable proxy for real-world compilation performance. And if you perform these measurements with and without the ETIDeclarations enabled, you will notice that they cause this short piece of code to take 100-200 ms longer to compile. Mind you, these are explicit template instantiation declarations - they tell the compiler that it does not have to generate code for these template instantiations. However, they do still force the compiler to instantiate all the declarations of all involved transitive templates. Explicit-template-instantiation-declarating X<int> requires the compiler to (internally, in some abstract way) note down the existence of every last std::vector<int>::const_reverse_iterator::operator-=(size_t) and so on (note 2). That's a lot of nested templates that it needs to (abstractly) "declare", even if it never has to actually generate code for them, which explains why the mere presence of the ETIDeclarations slows down compilation measurably - in every single TU that includes this header!
Note 1: It is, in my experience, relatively rare that the set of valid template arguments is closed and not open - in a way, an open set of permissible types is the whole point of writing templates. And such a templates intentionally being provided across a module boundary is also all but rare for lower level libraries in a bigger code base.
Note 2: It's easy to underestimate the amount of template nesting in e.g. the standard library too. Here's a single typedef in std::vector (with two further levels of typedefs expanded):
using const_reverse_iterator = _STD reverse_iterator<_Vector_const_iterator<_Vector_val<conditional_t<_Is_simple_alloc_v<_Alty>, _Simple_types<_Ty>,
_Vec_iter_types<_Ty, size_type, difference_type, pointer, const_pointer, _Ty&, const _Ty&>>>>>;However, I will still challenge you on your build time/object size claim. Take a codebase structured around lazy tasks (i.e. every function is a coroutine), factor in runtime parallelism/SIMD dispatch, multiply with loop abstractions (think std::ranges) and you get frighteningly close to linker limits in no time (which, of course, also implies minutes of compilation per source file). Each of those pieces is templates through and through, and there's no explicit template instantiation way out of any of them.
Also, I have to (sadly) point out that explicit template instantiation declarations only take you so far. Yes, they reduce compiler time spent in codegen, but also cause the compiler to create all required transitive (!) declarations in every TU, which will quickly offset those gains when talking about template types that have lots of small member functions.
It feels a bit like the error correction exists mainly so you can embed fancy logos in your QR code... Avoiding damage to some parts but not others seems unaligned with real world scenarios...
> What if an open source project is used directly by consumers, and causes them harm? The public policy is clear: they must be compensated.
It's expressly not clear what the implications here are, according to the article.