Static arrays are the best vectors
mynameistrez.github.io
mynameistrez.github.io
It may compile, but it might not link. With very large static arrays, in the link step, you might get something like this:
my_program.o: In function `my_function':
my_program.c:(.text+0xf38): relocation truncated to fit: R_X86_64_PC32 against `.bss'
It's due to the linker assuming statically allocated things can be addressed with a 4-byte pointer (I encountered this on x86_64 arch). Using malloc to allocate the array will fix this.That being said, using malloc once to allocate a very large array will get the same dynamic allocation of backing RAM as described in the article on linux, assuming overcommit is enabled.
I don't know what Windows does.
Windows does not allow overcommit. The commit limit is the total size of RAM + the size of the page file, attempting to commit more than that will fail.
I haven’t run it myself but I’ve heard others mention that you can prob get several hundred MB to a GB or so but shouldn’t count in getting near the 4GB limit.
Nonetheless, that’s plenty for my use case and random hacking.
(I’m actually somewhat conflicted about TFA’s idea, for what it’s worth, I just don’t think freeing memory is the most problematic part of it.)
And you save on having to write some kind of allocator for it.
Most programmers may not have touched those APIs but that doesn't mean that they didn't have a need for them.
To a degree, I agree with the premise of the post, we can preallocate huge chunks, and actual resource consumption will only happen on demand when the memory is first used. But note that there is still a resource whose allocation we can't delay: virtual address space, which must get acquired immediately obviously. When trying to be compatible with 32-bit machines (less than 4GB of virtual address space available), preallocating huge chunks is not practical.
As to using static globals, I'm not sure it's the best idea. One can use mmap() (or Virtual Alloc() on Windows) to the same effect. That might be better from an architectural view, with regards to encapsulation and such, and allows a little more dynamicity, making multiple separate (but still large) allocations and such. It's possible to give debugger-visible names to these allocations as well, for example by way of the memfd_create() API.
An error indicates the user's fault: they need to correct their input before retrying.
An apology indicates the program's fault: either logic for what had been incorrectly thought to be a rare^2 case is missing, or the program has only been configured to deal with N foos, and the user's input requires more. In the former case, the user needs to get a programmer fix before retrying, but in the latter, maybe they can just reconfigure and rebuild (even better, restart with different env vars?) then retry.
The relevant point of this is that when the user is likely to be another programmer, "reconfigure" can be as simple as editing a static allocation.
However, many programs require a bit more planning how to deal with unfortunate situations, at least without entirely restarting. It's good to not prevent repairing an unfortunate situation from the start by baking in certain constants that can't be practically assumed.
Better than editing a static allocation, is if the program can cope with that situation automatically in some way. Next better thing, have it runtime configurable (e.g. command line arguments, no recompilation needed). I don't see a huge benefit to prefer global arrays over memory maps generally.
Use request at creation to set a capacity for 99.9% of use cases. Still have realloc there to stop any sneaky buffer overflows.
Have your cake and eat it too!
But if we had let the vector grow naturally it would probably not make such a big realloc()
This would mean a more complicated malloc, where you ask for both an actual memory allocation, and an amount of address space to be allocated. And will still require you to have some sense of the maximum size you might need in a given vector. Also, it forces memory allocations to be multiples of the page size, which isn't always ideal.
[0] https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/cheri...
As many (or most) of C++ features, std::array is a logical idea with terrible ergonomics, bordering to unusable. While they are potentially a little safer to use than plain C arrays (but compilers could easily add runtime checking for plain C arrays too), they are much less ergonomic, both in verbosity of declaration and errors.
They can't replace an init-allocated std::vector either, because they have a size fixed at compile time.
And unique_ptr itself is TERRIBLE to use too. If RAII is your thing, then still very often, it's totally worth the boilerplate to wrap individual objects in type-specific explicit "Holder classes", instead of relying on std::unique_ptr which has terrible ergonomics too, again both for declaration and use.
Oh, and if you're using a free-standing destructor function for the managed object (which is often a very good idea because exposing the class definition in typical C++ style leads to TERRIBLE TERRIBLE compile times), you'll have to bake in a function pointer that is carried by unique_ptr at runtime (I suspect that is because they couldn't make the function a templated parameter because in C++, this leads to ambiguitiy if it is a static inline function declared in a header (terrible!)). So you're paying for all the all the drawbacks of the templated type and get none of the benefits (except for saving a few lines of boilerplate).
TERRIBLE.
And now you propose combining these two terrible things? How would you use this?
#include <memory>
#include <array>
struct Foo
{
~Foo()
{
printf("Bye!\n");
}
};
int main()
{
// std::unique_ptr<std::array<Foo, 42>> foo = std::make_unique<std::array<Foo, 42>>(); // my fingers hurt and my eyes bleed
auto foo = std::make_unique<std::array<Foo, 42>>(); // fingers hurt a little less but my brain will melt on next compiler error
foo->at(0); // valid
foo[0]; // error. TERRIBLE!
return 0;
}
I wouldn't be so upset about your advice if I didn't know you often don't validate your claims and "references", but keep on suggesting and criticising existing practice that has evolved to certain points (and keeps evolving!) for a reason.> running in places like CERN and Nokia Networking infrastrure
Is there no bad code in these places? And, as an avid follower of your posts, how long ago has this been? 15 years? And now you're selling people ideas on HN instead?
You couldn't even be bothered to actually use type alias to make the example simpler, or take into consideration build configurations, only to make the typical C rant.
The C++ modules that you've been advocating for years as a workaround for slow compile times still haven't arrived in my (and most everybody else's) practice, sorry.
I do use precompiled headers in some settings to work around slowness, which does help a lot for build performance but is again terrible (breaks modularity, it's basically a choice: #include everything vs have slow build, with little possibility for making a tradeoff).
There's a lot of problems that can be solved just fine with an upfront defined fixed memory layout in a global static variable.
Dismissing globals is basically the same problem as sweeping generalizations like "goto considered harmful", "avoid for-loops" or "almost always auto" - such statements can serve as advice for newbies, but following the advice blindly can also lead to less elegant and less maintainable solutions.
...or from a function that calls itself, directly or indirectly. This automatically becomes possible basically whenever the function accesses a user-supplied callback. So you'll run into issues unless you have full control over your callees, or only access the global array in logically-atomic patterns.
(To give an ancient example: to enable function calls, PDP-8 programs allocated a word of static memory before the top of each subroutine to store the return address. However, if a subroutine tried to recurse into itself, then the old return address would be overwritten, and it would never be able to return properly. Supposedly this could result in very gnarly errors to debug.)
Programming languages should directly support a data structure which is guaranteed to have this behavior. Not just that it allocates a new array which is 4k bigger and the copies it over. But one which grows by having the OS map another 4k block to have the addresses past the old end of the array.
I made a small experiment with 1-million int array, which has a single element written to at index 100000, and then force a core dump by writing to a dereferenced int pointer which is initialized as 0.
The resulting core file is almost 4gigabytes.
https://www.merriam-webster.com/grammar/people-vs-persons
Then again, I'm not a language prescriptivist. I post the above only to demonstrate that the use of "persons" is hardly unusual.