Sane C++ Libraries
github.com
github.com
> Unplanned Features:
> SharedPtr
> UniquePtr
> In Principles there is a rule that discourages allocations of large number of tiny objects and also creating systems with unclear or shared memory ownership. For this reason this library is missing Smart Pointers.
I don’t like that at all. I take the common view that all heap objects should at least be allocated via smart pointers. Doing so is safer and easier and usually zero-overhead. After allocation, it may be necessary to pass those objects via raw pointers/references, but smart pointers should be used where appropriate.
So while I agree that it’s undesirable to allocate “large numbers of tiny objects”, I would want smart pointers as long as there’s any dynamic allocation at all.
Regarding UniquePtr<T> I used to have one but I later on decided to remove it.
https://github.com/Pagghiu/SaneCppLibraries/commit/9149e28
However, that being said the library is lean enough so that you can still use it with smart pointers provided by any other library (including the standard one) if that's your preference.
Is the point of having a kitchen-sink library like this not that you dont have to reach for a 3rdparty library for things that you need 'all the time'?
Certainly, not everyone needs it.
...but, not everyone needs threads either. Not everyone needs an http server; and yet, if you have an application framework that provides them, when you do need them, it saves you reaching for yet-another-dependency.
Was that not the point from the beginning?
unique_ptr is a fundamental primitive for many, as you see from some other frameworks (1), and implementation is not always either a) trivial, or b) as simple as 'just use std::unique_ptr'.
This does seem like a very opinionated decision with reasonably unclear justification; perfectly fair, you're certainly not beholden to anyone to implement features just because they want you to, but I think it's difficult to argue there's not concrete use for something like this, in a way that aligns with the project principals.
I would go so far as to argue that:
> Do not allocate many tiny objects with their own lifetime (and *probably unclear or shared ownership*)
Is hostile to not having a unique pointer.
[1] - eg. https://github.com/EpicGames/UnrealEngine/blob/release/Engin..., https://github.com/electronicarts/EASTL/blob/master/include/...
This style is also more cache friendly if you are going to be looping through the elements.
I only really use C++ for a toy game engine right now and in that codebase I don’t use any smart pointers and most objects/functions get passed references to their object dependencies. I classify objects into groups where each groups ownership is very clear. So its owner is responsible for maintaining the memory and any raw pointers can always be assumed to be borrowed references. I use handles then the underlying objects lifetime might differ from whatever is holding a handle to it. Short lived objects are kept trivial and allocated from stack/bump allocators or pools and reset at well defined times (every frame, end of level, etc)
I’m much happier this way than when I used smart pointers or when I had less well defined memory ownership.
I have not been working (yet) on custom allocators, but that's on the roadmap: https://pagghiu.github.io/SaneCppLibraries/library_container...
(Not really a c++ expert, but that's my understanding; someone more knowledgeable can correct me).
For the rare case of porting software with unclear ownership, I use a `dumb_ptr` template with allocation and deallocation methods. Since this is header-only it naturally avoids the poisoning.
In particular, the `vector` method mentioned elsewhere is completely broken since objects move and thus you can't keep weak/borrowed references to them. If you use indices you give up on all ownership and are probably using global variables, ick. Please just write a proper pool allocator if that's what you want (possibly using generational references to implement weak).
Regarding the std::vector method, you may have a very loosely coupled system where a bunch of T1's enter into a pipeline and come out as T2's. For this use case, std::vector<T1> and std::vector<T2> are great. On the other hand, if you need to create an object and hand it off to someone else with no knowledge of how long they will need to hold onto it, then std::shared_ptr could be a good option.
In the in-between you have entity component systems that do the type of index tracking you mention so that identities are decoupled from memory location, allowing objects to move. I didn't understand your point about global variables and why they are necessary to implement this type of system. I also didn't understand how this gives up on all ownership. The owner would be the system that maintains the index to memory location mapping.
It is deeply integrated into the C++11 memory model. The compiler has to know about the semantics of the type to make sure it doesn’t reorder operations around it.
> No C++ Standard Library / Exceptions / RTTI
In some vanilla app code exceptions are fine, but they introduce nasty edge cases in the kinds of systems code architectures C++ is mostly used for these days. Additionally, they don’t solve an urgent problem in practice that would strongly incentive someone to use them, so it the price of admission is rather steep for minimal benefit.
I really wish that everyone that plagues C++ library fragmentation with disabled this, disabled that, just stick to C and leave C++ community alone, so that we can fully enjoy the language as designed.
I guess we need something else everyone can agree upon.
Until then, the computing foundations will kept be being written in C++, until some big name decides to rewrite LLVM, GCC and VC++ into something else.
It is trivial to “guarantee” the underlying memory and is idiomatic for a lot of software that cares about performance or reliability. That is code anyone can write if they care. There isn’t much that can go wrong with memory allocation if you are not allocating memory from the system. No one is requiring C++ developers to poorly duct tape a bunch of rubbish STL together and call it an app. That simply isn’t something you see much in the hardcore systems domains where C++ is the tool of choice.
Somehow, mission-critical software is routinely written in C++ without exceptions and it works just fine. Error states are a normal part of all code, no exceptions required since obviously many languages don’t have them. And no, the alternative is not a segfault. C++ is designed to work just fine without exceptions. The language allows you to bring your own error/exception handling models with minimal overhead, same way you can import alternative ownership/safety models.
An error and an exception are the same concept. Yes, it’s true that not using the std::exception class or any template named “exception” is exception free code. But you’re just lying to yourself. An error check or an assert or verify not <condition> is a “catch”.
One large open source project that is using Exception free C++ is SerenityOS for example https://github.com/SerenityOS/serenity
The best way to write proper exception free C++ is not to use the C++ Standard Library.
Only in C++ land do developers delusion themselves into thinking their way isn’t this way. That exception free means just simply not doing try/catch. jandrewrogers made a good argument below about memory safety and allocations in regards to mission-critical code but even in that scenario, underlying memory can be manipulated by MITM or other conditions that could cause corruption or segfaults in allocator pages.
I think that handling errors with ErrorOr<T,E> or similar techniques (I use something similar too) is very different from exception handling.
My main problems with exception handling are:
1. It's not zero-overhead (brings in RTTI often) 2. You can't know if a function throws something by looking at its signature 3. You don't know what types exceptions a function can throw 4. It doesn't force users to handle errors that can happen, leading tomore "happy path" style code
Something like ErrorOr or similar with [[nodiscard]] ticks all the above 4 points.
ErrorOr is a better design for things to be deterministic. The point I was trying to make is that exceptions are errors and error handling and exception handling (while implemented differently) are essentially the same thing. In C++ std::exception is “exceptions” and everything else is “pretty error”. It doesn’t matter how the house is decorated. Exceptions = Errors = Oopsies = NotIntendedState
I'm not going to talk about errorOr specifically because I don't know how it's implemented but rather your premise.
Exceptions are the modern goto, the try catch may be in the caller function. Or it could be 3 inherented classes away (this is hyperbole, I'm not sure if it would actually work) with so much indirection it would be impossible to follow the code flow other than to step in a debugger. So in fact it's worse than a goto since at least a goto had a label.
I'm not for or against exceptions just pointing out that a result or optional type is in no way at all similar to an exception.
On another point C++ exceptions are notoriously inefficient as well. There ware valid reasons to want to be exception free even from a code style perspective.
Observer pattern is fine, because there is still a link between the execution flow; basically any time you register a callback you're using the observer pattern.
Pubsub adds a middleware, typically in the form of a message passing framework, so that both publisher and subscriber aren't necessarily aware of each others existence. The publisher throws its data into the framework hoping someone finds it useful, and the subscriber listens for those messages hoping someone is publishing them. Think of it like multithread asynchronous goto of your execution flow, except with a lot more boilerplate.
How are pubsub systems a modern goto? Genuinely curious since I don't work with them directly that often
It seems like you know C++ pretty well, so I think you could probably come up with some reasons if you gave it a try.
The obvious thing that comes to my mind is because you can't rely on an undesirable, unknown, or potentially nonexistent implementation of the STL for your target platform.
As for cookies: who am I to dictate what/how someone should bake?
Otherwise good luck porting that exception free code across multiple OSes and C++ compilers, not only the big three.
There is a "turn code into object file using compiler" C++, that is much different from "committee vision" C++, and until recently (probably until initializer lists), -fno-exceptions -fno-rtti -ffreestanding did not emmit standard library symbols, except maybe static initialization.
Current C++ is being glued to library like crazy, look at continuations...
Here-in lies the problem. This only says “no std::exception” “no throw” but nothing is stopping you from writing your own template for error handling. An exception is by definition a “unexpected condition” and we all have to handle exceptions in some form or fashion. Whether that’s with error codes, macro checks, or the like. That’s my point. Exception free code doesn’t exist. Safety guarantees are compile time at best. If this weren’t true, things like memory manipulation wouldn’t be possible. But it is. CheatEngine exists. Funnest thing is to read a var, set a var, and then check if the set var equals the set value or if it equals the old value. You just detected a memory pin cheat.
This is all in an effort to bring down compile times, avoiding including anything from the standard because you can never know if including <atomic> is bringing 10K lines of code in your header.
I will probably provide an optional USE_STANDARD_HEADERS flag someday to allow including a few standard things, including atomics, to avoid doing things wrong on compilerS that are not tested enough (as I clearly can't test every compiler).
The stuff he considers complex is mostly complex because it handles a lot of corner cases, or else it has back compatibility constraints he will be dealing with as well.
Good luck to him though. Doing stuff for fun is the best!
And yes, complex stuff sometimes tries to handle "everyone's use case" but if you can limit yourself to 95% of use cases, your code suddenly become a lot simpler. The backward compatibility consideration holds true as well.
For example, I have been creating an Async Library (plus a few other things like the FileSystemWatcher etc.) that cover a good portion of what is done in libuv. Of course libuv code handles A TON more of edge cases and has a lot of compatibility constraints, but with a lot less code I can provide enough functionality to satisfy a lot of use cases. Not all use cases, but a lot of use cases.
Thanks for the good luck! I am definitively doing it just for fun :)
I realize people think of this as a difference between c and c++, but if you go to compile C code with a thread_local variable and throw -nostdlib you’re going to have a bad day. Same for atomics, sometimes complex numbers, floating point exceptions, even receiving arguments and other core language features require some crt code. Removing the runtime library guts the implementation regardless of your language. The question is, does your language provide ways to deal with this? Rust has core, c has alternate ad-hoc embedded libc implementations and a history of bootstrapping implementations long enough to have them be well understood, c++ has/will have free-standing.
I used this for JSON last time I wrote any C++ a few years ago and it still seems popular. It seemed sane enough to me.
Overview: https://www.boost.org/doc/libs/1_84_0/libs/json/doc/html/jso...
Benchmarks: https://www.boost.org/doc/libs/1_84_0/libs/json/doc/html/jso...
Parsing Options for non-standard JSON: https://www.boost.org/doc/libs/1_84_0/libs/json/doc/html/jso...
This alone already rules them out as sane on my book.
Any company that worships that style guide is one I am happily never going to work on.
What macros are bothering you the most?
(the community version of Qt is LGPL, I have a small business licence)
Declarative build definitions are generally much easier to work with and scale than using an imperative language.
https://en.wikipedia.org/wiki/CMake#CMakeLists.txt
Most "declarative" build systems are not actually what they "declare" to be. I've seen too many DSLs introducing half backed imperative concepts here and there to do _if_ and _for_ constructs or function calls, redoing the same as imperative languages but poorly.
I like to market it as an "alternative world" where the C++ stdlib is more a platform abstraction library focused on carrying practical tasks like networking, Async I/O, HTTP etc.
It's also definitively placing itself in the middle between unsafe C and bloated C++.
I love well written C libraries, like the sokol or stb libraries.
https://pagghiu.github.io/SaneCppLibraries/library_algorithm...
Hopefully it will get expanded with more useful algorithms, it has not been a priority in the first releases cycle.
It's also often the correct choice for something like a collision partition for a physics sim where elements are most likely in the same sort position each frame.
POCO started 20 years ago, so it has a slight advantage :)
The documentation here states:
https://pagghiu.github.io/SaneCppLibraries/library_reflectio...
Note Reflection uses more complex C++ constructs compared to other libraries in this repository. To limit the issue, effort has been spent trying not to use obscure C++ meta-programming techniques. The library uses only template partial specialization and constexpr.
TinyEngine is good, but doesn't build on windows yet
What containers, beside Vector<T> (and Map<K,V> made with Vector) + variants would you like to see the most?
HashMap and proper Map<K,V> are already on the roadmap https://pagghiu.github.io/SaneCppLibraries/library_container...
Stack can be easily created with Vector.
I think Queue is pretty specialized, but I will think about it.
[a]: Obviously, people still write their own collections/containers in C#, but they tend to only do so for very specific/performance-sensitive circumstances.
Whereas due to C++'s "never break compatibility" decision, the standard library has progressively decayed over time. It has become a bloated, rotting dinosaur where even the slowest of interpreted languages can comfortably beat several of its aspects. (Ex: std::regex is pathetic and pitiful, vector<bool> triggers laughter, substandard maps, etc). Considering that C++ thumps its chest and loudly proclaims its superb performance, this has now become a sad joke.
In the natural world, a species that cannot adapt to new circumstances and never discards undesirable characteristics simply perishes.
The C++ Standard Committee has firmly and unequivocally decided that the C++ language should mirror the same approach and limp down the road, carrying the full-weight of its sins for all its journey, until it falls into oblivion.
Also breaking everything a couple of times is why most corporations are nowadays stuck in a Python 2 / 3 parallel world in .NET Framework / .NET Core , or in Xamarin.Forms / MAUI, UWP / WinUI,...
One that I find particularly annoying lately: if you work on a project that makes heavy use of the STL (or, really, any heavily templated library written in a similar style, with a focus on "ergonomics" :-( ), you'll quickly find that backtraces for debug builds can easily be 50-100 levels deep with most stack frames just consisting of incomprehensible layers of abstraction which get optimized away. So, debug builds are totally useless, and build times are typically long. Contrast that with something written in "C style" or simply making far fewer use of C++ features, written in a more straightline style, with fewer levels of abstraction: builds will be fast, stack traces will be small and can easily map directly onto the concepts relevant to the program or library, and debug builds are useful once more. Night and day difference.
Maybe as an example just compare a library like Eigen to a similar imaginary library written in C.
Eigen leans heavily on C++ features to reduce line count and make something look visually more mathematical or maybe more MATLAB-style in the name of ergonomics; and obviously on the backend it relies on the behavior of the C++ compiler to simplify the elaborate template expressions and make the code efficient. If you ever try to step through code that uses Eigen in a debug build or examine a stack trace from inside Eigen where an exception has originated, the situation is not pretty! Contrast this with what will happen in the imagined C-style library. If all the heavy lifting happens in BLAS or LAPACK and the C library is basically a thin library to make things a little more automatic and easier to manage, the stack traces will be short and each stack frame will be easy to digest at a glance because of names like "mat_mat_mul" or similar, instead of "mat<double, double, 4, allocator<...>> const & operator*(mat<double, double... 159 more characters of template gobbledygook that makes the eyes glaze over)". The former will also be significantly faster to compile.
Anyway, I guess I just don't see how the latter STL/Eigen/whatever-style approach to things is an improvement over the far simpler idiot style of doing things.
I'm not trying to be argumentative here, I just want to clarify my point as much as possible.
My point is that one can have a lot of C++ working at scale. Reading code is very important part of large scale projects. Dropping debug builds is a reasonable tradeoff.
I would also argue that keeping it so that you can always easily run your project "REPL style" from a debugger exerts a very favorable downward force on all the complexities that make large projects hard to maintain.
Some of the common issues: static allocation or lack thereof, requiring default constructible classes, initializing memory you are going to overwrite anyway, inability to be used in some metaprogramming contexts, suboptimal allocation behavior, etc. The STL is opinionated but unfortunately that opinion dates to a time when C++ was primarily used for ordinary app development, not high-performance code.
Like many, I maintain my own C++ "standard library" that is much better designed for the kinds of software I tend to work on (database kernels and data infrastructure, mostly).
That's actually not true, though I certainly don't fault you for believing it :-) but there are definitely more things to quibble about around vector if you're serious about performance. As an example, try writing a can_fit(n) function, which tells you whether the vector can fit n elements without reallocating. Observe the performance difference between (a) a smart manual version, (b) a naive manual version, and (c) the only STL version: https://godbolt.org/z/88sfM1sxW
#include <vector>
template<class T>
struct Vec { T *b, *e, *f; };
template<class T>
bool can_fit_fast(Vec<T> const &v, size_t n) {
return reinterpret_cast<char*>(v.f) - reinterpret_cast<char*>(v.e) >= n * sizeof(T);
}
template<class T>
bool can_fit(Vec<T> const &v, size_t n) {
return v.f - v.e >= n;
}
template<class T>
bool can_fit(std::vector<T> const &v, size_t n) {
return v.capacity() - v.size() >= n;
}
struct S { size_t a[3]; };
template bool can_fit_fast(Vec<S> const &, size_t);
template bool can_fit(Vec<S> const &, size_t);
template bool can_fit(std::vector<S> const &, size_t); template<typename T,
class vector {
private:
pointer __begin_;
pointer __end_;
__compressed_pair<pointer, allocator_type> __end_cap_;
public:
constexpr const pointer& __end_cap() const noexcept {
return this->__end_cap_.first();
}
constexpr size_type size() const noexcept {
return static_cast<size_type>(this->__end_ - this->__begin_);
}
constexpr size_type capacity() const noexcept {
return static_cast<size_type>(__end_cap() - this->__begin_);
}
};
So if the functions are fully inlined we should end up with template<class T>
bool can_fit(std::vector<T> const &v, size_t n) {
return static_cast<size_t>(v.__end_cap_.first() - v.__begin_) - static_cast<size_t>(v.__end_ - v.__begin_);
}
At least algebraically (and ignoring casts) it should be equivalent to v.__end_cap_.first() - v.__end_, which is more or less what the manual implementations do. Maybe the optimizer can't make that transformation for some reason or another (overflow and/or not knowing the relationship between the pointers involved, maybe)?If you change can_fit(Vec<S>) to:
return (v.f - v.b) - (v.e - v.b) >= n;
You end up with code that looks pretty similar to the can_fit(std::vector<S>) overload (the same for clang, a bit different for GCC), so it does seem it might be something about the extra pointer math that can't be safely reduced, and the casts aren't really relevant.(I'm also a bit surprised that can_fit_fast produces different assembly than the Vec can_fit overload)
(There's also a secondary issue here, which is that sizeof(S) isn't a power of 2. That's what's introducing multiplications, instead of bit shifts. You might not see as drastic of a difference if sizeof(S) == alignof(S).)
I feel the division by sizeof(T) shouldn't matter that much, since the compiler knows it has pointers to T so I don't think the divisions would have remainders. I want to say pointer overflow and arithmetic on pointers to different objects (allocations?) should also be UB, so I suppose that might clear up most obstacles? I think I'm still missing something...
Does make me wonder how frequently this pattern might pop up elsewhere if it does turn out to be optimizable.
> Does make me wonder how frequently this pattern might pop up elsewhere if it does turn out to be optimizable.
Probably a fair bit, but as I mentioned, it might break a lot of code too, because there's too much code in the wild doing illegal things with pointers (like shoving random state into the lower bits, etc.). Or not... the Clang folks would probably know better.
Maybe this could be a good way to jump into messing with LLVM...
Out of curiosity, how much of a performance difference did you observe in practice when you made this optimization?
(Although, note I didn't claim "the STL is slow"... that's painting with a much broader stroke than I made.)
Speaking of realism, putting these in quickbench seems to confirm that the differences between them are not material, and that the STL version is in fact the quickest, but they are all essentially free. There's not a way to make a realistic microbenchmark for this, for the same reason that it doesn't feel like a real-world performance issue.
By the way clang does a much better job here: https://quick-bench.com/q/XcKK782d-7A6YHbiBRTlOnIRnPY
Your fallacy here is assuming that that just because N is a variable, therefore N is large. N can easily be 0, 1, 2, etc... it's a variable because it's not a fixed compile time value, not because it's necessarily large. (This isn't complexity analysis!)
> Speaking of realism, putting these in quickbench seems to confirm that the differences between them are not material, and that the STL version is in fact the quickest,
Your benchmark is what's unrealistic, not my example.
I'm guessing you didn't look at the disassembly (?) because your STL version is using SSE2 instructions (vectorizing?), which should tell you something funny is going on, because this isn't vectorizable, and it's not like we have floating-point here (which uses SSE2 by default).
Notice you're just doing arithmetic on the pointers repeatedly. You're not actually modifying them. It sure looks like Clang is noticing this and vectorizing your (completely useless) math. This is as far from "realistic" as you could possibly make the benchmark. Nobody would do this and then throw away the result. They would actually try to modify the vector in between.
I don't have the energy to play with your example, but I would suggest playing around with it more and trying to disprove your position before assuming you've proven anything. Benchmarks are notoriously easy to get wrong.
> So it doesn't feel like a realistic use case.
I only knew of this example because I've literally had to optimize this before. Not everyone has seen every performance problem; clearly you hadn't run across this issue before. That's fine, but that doesn't mean reality is limited to what you've seen.
I really recommend not replying with this sentiment in the future. This sort of denial of people's reality (sadly too common in the performance space) just turns people off from the conversation, and makes them immediately discredit (or just ignore/disengage from) everything you say afterward. Which is a shame, because this kind of a conversation can be a learning experience for both sides, but it can't do that if you just cement yourself in your own position and assume others are painting a false reality.
Admittedly that's not a performance issue, but it's annoying.
Vector<bool> is a little weird if you are just starting with C++, but it does have major performance benefits in its niche, and it came from the 1990s so we can be generous in overlooking its rough edges.
Not really. Unless you have a very specific set of performance requirements, STL containers are usually more than enough. And we have third party libraries such as Abseil/Boost to cover the major gaps in the rest. I do see some legitimate cases to write own container libraries. But for many cases people don't really measure their primary workload before writing such libraries, instead they just write it because they can write (and it's fun).
I mean: i remember the days, when each and every project had its own string class.
Every library having its own version of common data structure is unfortunately something that C++ programmers can't really seem to agree on :)
There are special cases, and there are engineering politics, but it's all basically fine.
Example: want to measure the length of a UTF-8 code point in a string and did not synchronize the call? Well, too bad, now you might have corrupted the global mutable state it relies on! (the fact that you can make such a trivial piece of code have two!!! points of thread-safety failure one is C locale and another is transcoder still refuses to leave my mind)
C++ has its place, but something new could not possibly displace it soon enough, even with C it feels like libraries fit together more easily.
But I see the point in it helping C++'s unusual longevity as well
(also sorry my initial comment came off like ragging on your library, it wasn't meant that way, it was more of a commentary on the overall state of the C++ ecosystem, so I appreciate people with a slightly broader view like yours!)
No problem at all for your initial comment! I share similar sentiment, I've found a lot easier in the past glueing C libraries to do something more than trying to integrate a C++ library for the exact reasons you're describing...