Sadly I don't think the tooling exists to enforce something like that automatically, because C++ is so hard to parse (a separate issue to memory safety). But when/if metaclasses finally land, it should be possible to write a "fixed" class/struct keyword that default initialises all fields.
(Oh, some background before I share that monstrosity with you: the project is a virtual machine I designed for a class I taught recently, and I experimented in implementing it in modern C++ and had a bit too much fun trying to see how much of it I could encode in the type system, which is why it will probably take a really long time to compile. Here is some code that actually uses that header, for reference: https://github.com/regular-vm/emulator, and I should note that I have internally changed some of the architecture slightly do accommodate an assembler that I have put aside for now.)
struct Instruction { explicit Instruction(int encoding) { } };
using T = Instruction &&;
int encoding = 0;
auto &&result = T(encoding);
There are a few things that went wrong here, but the final red flag to notice is that you should never call a constructor directly with 1 argument directly, because it's just a different syntax for a C-style cast, which we know is dangerous due to its bypassing of safety checks—and this is true regardless of whether we're dealing with C++ constructs (like classes and constructors and such), which I think might be what you're realizing now. (This is poor C++ design, but it's old and people know to avoid them syntactically just like C-style casts. It might be nice to have a warning for it too.) Rather, when you're passing a single argument, you want to write one of these syntaxes: T result(encoding); // option 1
auto &&result = static_cast<T>(encoding); // option 2
With these, you receive an error, e.g.: error: invalid static_cast from type 'int' to type 'T' {aka 'Instruction&&'}
I believe this is because the code requires 2 conversions to occur at once, which is an error because (surprise!) it's a generally unsafe thing to do: (a) conversion of int to Instruction, and (b) conversion of Instruction to Instruction&&.Now you bypassed this by using the uniform initialization syntax (i.e. braces). I'm going to go out on a limb here and say that was another mistake, even though it "solved" your problem here: despite the widespread use, brace initializers are, in my experience, not a good thing, and it's unfortunate that people embraced them (ha) with open arms, and similarly goes with emplace_back() and some other things which I'll address below. The syntactic convenience they provide is just too minor compared to the issues they introduce or obscure. And in this case, they indeed actually introduce a new issue if you use them like 'T result{encoding}': they allow 2 casts to occur at once. I find that incredibly dangerous, and I think it should be at least a warning if not an outright error like before. It seems like a C++ design flaw to me, but in any case, maybe someone should get compiler writers to add a warning for this.
Anyway, let's move on. If you're reading this, you're probably noticing that you ended up with Instruction&& in the first place—that's probably not what you wanted, or at least not what you should've wanted.
And that's where we get to the heart of the issue: your real problem is that you used decltype. If I saw that during code review, I would force you to change it—and using it on declval is just adding more fuel to the ember.
The reality—which unfortunately you do not see people acknowledging—is that decltype, auto, uniform initialization syntax, emplace_back, etc. are all dangerous, and harder to reason about than they look. I think it's unfortunate that the C++ committee encouraged people to use them so much, and I think they're overused to an insane degree. People who love them for their nice syntax don't go out of their way to figure out their pitfalls, but I almost never use any of them unless I absolutely need to. Most problems that they solve (one notable exception being 'auto' with lambdas) were quite elegantly solved in C++03 using typedefs. The only caveat was that you had to give up on the idea of minimizing keystrokes. It's quite a realization when you realize that solves so many of your problems.
Anyway, I'm not trying to blame these on you. Obviously these are blamable on C++, and we could (and should) have more warnings for them. Rather, my main message is that you can avoid these problems (even if they're other people's faults) syntactically—and locally—if you don't try to embrace the absolute "latest and greatest" in C++. IMHO you should only use the newer features if they solve an actual semantic problem for you, not merely because they minimize your typing.
In fact, I think is a huge mistake people make with software in general, and here, C++ in particular. They feel if something is old then it must be bad and you have to do everything in a new way. But if you stick with what works and start caring less about people looking down on you for using "old" syntax just because it's old, you'll find a lot of the old C++03 patterns (typename Pair::first_type, etc.) are actually robust to the problems that the newer ones introduce. People just don't realize this because they're more verbose than they would like, and we're in an era where doing things old style looks bad for no good reason.
Oh also, one last thing: aside from avoiding decltype and using the equivalent of a C-style cast, one more thing that helps you avoid this is to avoid overusing templates. They also obscure what's going on, like here. Not to mention the slow compilation speed and lack of independent compilability. Those are also overused (and I see their appeal) but they're often unnecessary and make code statically difficult to reason about.
Not only is it possible, I think it is very likely, although your explanation is something I can follow along with and very much appreciated.
> There are a few things that went wrong here, but the final red flag to notice is that you should never call a constructor directly with 1 argument directly, because it's just a different syntax for a C-style cast, which we know is dangerous due to its bypassing of safety checks—and this is true regardless of whether we're dealing with C++ constructs (like classes and constructors and such), which I think might be what you're realizing now.
Huh, interesting, I actually did not realize this. Is there a way to do this safely without creating an extra lvalue? I take it that there is no “extra explicit” keyword I can add to prevent this kind of accidental call, is there?
> I'm going to go out on a limb here and say that was another mistake, even though it "solved" your problem here: despite the widespread use, brace initializers are, in my experience, not a good thing
Yeah, I am not really a fan of them either :( Even I know of a bunch of caveats about them and C++ initialization is an extremely complicated topic…
> If you're reading this, you're probably noticing that you ended up with Instruction&& in the first place—that's probably not what you wanted, or at least not what you should've wanted.
No, but as you observed that it “works out” at some point in the pipeline so I obviously did not care to really figure out if this was what I wanted or not.
> And that's where we get to the heart of the issue: your real problem is that you used decltype. If I saw that during code review, I would force you to change it—and using it on declval is just adding more fuel to the ember.
Somewhat strangely, C++ seems like the only language where I would even consider to use such a construct. I think every other language just erases their types or simplifies them so you can be comfortable writing something like “Iterator i = collection.start” or “int size = collection.count” whereas in C++ you have some generic distance_type and it feels dirty to just work with a size_t or whatever you know the thing to be.
> IMHO you should only use the newer features if they solve an actual semantic problem for you, not merely because they minimize your typing.
A good point, but I would like to just mention that this was clearly an experiment in trying out the “latest and greatest” ;)
> Not to mention the slow compilation speed and lack of independent compilability.
Wait, you’re telling me my 100 line program shouldn’t take a dozen seconds to compile?!
The static_cast<T>(arg) syntax I used does exactly this! It's what you should use pretty much everywhere instead of T(arg). If it's too much typing, yeah unfortunately it is, though life is a lot easier if you can e.g. bind 'sc' to expand to it in your editor.
> No, but as you observed that it “works out” at some point in the pipeline so I obviously did not care to really figure out if this was what I wanted or not.
Yeah... sadly C++ is just about the 2nd-to-last last language you should deal with like that. The last probably being C. :-) Pro tip that might make it easier to avoid this: use typedefs very liberally. They help you avoid auto/decltype/etc. and are quite robust. (At least if your reviewers let you. If they don't, they probably haven't learned it the hard way yet.)
> Somewhat strangely, C++ seems like the only language where I would even consider to use such a construct. I think every other language just erases their types or simplifies them so you can be comfortable writing something like “Iterator i = collection.start” or “int size = collection.count” whereas in C++ you have some generic distance_type and it feels dirty to just work with a size_t or whatever you know the thing to be.
Those languages break too actually. Go Google "binary search bug" (with quotes). For example in C# there's Length and LongLength, which is dirty. When what they really need is just a native int. Another C++ tip: almost every 'int' or 'unsigned int' you ever deal with should be size_t or ptrdiff_t, because at some point or another it's probably an array index. It's very rare for that not to be the case; the only case I can think of off the top of my head is a logarithm (i.e. the shift amount in a bit-shift expression) or a timestamp (long long). Unless you're writing a generic STL-like container or allocator type (in which case, best of luck...), you won't need to care about difference_type or size_type.
> A good point, but I would like to just mention that this was clearly an experiment in trying out the “latest and greatest” ;)
Yeah ;) just keep it confined to experiments!
It's something I don't see Rust proponents address. It's easy to build a straw man argument of Rust vs. old-school C, but it's a more natural path to go from C to modern C++ than to go from C to Rust. You get to keep your compiler, build system, tools, libraries and indeed existing code.
See this for an example: https://news.ycombinator.com/item?id=21681395
https://people.gnome.org/~federico/blog/exposing-c-and-rust-...
And yet we see the same memory-safety bugs come up time and time again in supposedly-modern C++ programs. So either this modern way is not enough to avoid these classes of bugs, or people very quickly fall to the temptation to use unsafe constructs due to performance or just because it's easier to write.
We're humans. If there's an easier way to do something, even if it's less safe, we'll invariably do it sometimes. I like that Rust makes it harder to do so, and makes you explicitly say that you want to do something unsafe, which I imagine deters a lot of people from going down those paths. And when someone writes a memory-safety bug in unsafe Rust, they get much more egg on their face than if they were to write the same bug in C++.
The memory management issues in my C++ programs (recently, real-time audio stuff) are where I explicitly decide not to use modern C++ / automated memory management, and write my own allocators that rely on malloc/free or some variant (aligned_alloc, etc.)
In Rust I suppose I would just use an unsafe block and have the exact same issues.
Note that good unit testing, assertions and sanitizers generally take care of the issue.
Unit tests, assertions, and sanitizers are nice, but they demonstrably don't work; Chrome and Firefox use all three. They have hundreds of thousands or millions of tests, assertions on every other line, and all kinds of compile-time sanitizers but they still have hundreds of memory safety problems a year.
We built a new project in all "modern C++". It is 100% shared_ptr, unique_ptr, std::string, RAII, etc. It initially targeted C++17 specifically to get all the "modern C++" goodness.
It segfaults. It segfaults all the time. It is entirely routine for us to run a new build through the CI process and find segfaults. We fuzz it and find dozens of segfaults. Segfaults because of uninitialized memory. Segfaults because dereferencing pointers. Segfaults because running off the end of arrays. Segfaults because trusting input from the outside world ("the length of this payload is X bytes").
This is where the "modern C++" people tell me we must be doing it wrong. But the reality is that "modern C++" isn't as safe or as foolproof as the advocates say it is. But don't take my word for it - this whole thread is about Google people coming to the same conclusion.
Meanwhile I can throw a new dev at Rust and watch them go from zero to works in a week or so, and their code doesn't segfault, doesn't panic, and actually does what it is supposed to do the first time. Code reviews are easy because I don't have to ponder the memory safety and correctness of every line of code. Reasoning about unwrap() is trivial. Finding unsafe {} is trivial (and removing it is also usually easy).
And then one day I found Rust, and all those problems went away. I can now write fearless code, and I don't have to endure the stench of rotting bodies anymore.
True story.
Any progress on a C++ to Rust converter? Not a "transpiler". Something with enough smarts to figure out when to use native Rust arrays, not "offsets" to imitate pointer arithmetic. I'm surprised that one of the big C++ users, like Google, doesn't have a group doing that.
The guy who maintains it said in the reddit thread[2] about this same topic that the Google people have been sending him good PRs, which is presumably related to integrating Rust into Chrome.
[1] https://crates.io/crates/cxx [2] https://reddit.com/r/rust/comments/gpdorw/the_chromium_proje...
There are some non-lint type things that would help in a safe mode. These all need type information; they're not just syntax.
- Can't keep a raw pointer. If you create one, it has to have local scope and cannot be copied to an outer scope. This is like a borrow in Rust, and limits the lifetime of the pointer. Most uses of raw pointers involve calling legacy code, and don't need much lifetime. Most trouble with pointers involves them outliving the thing to which they point.
- Can't read into or memcopy into any type that is not fully mapped. That is, all bit values have to be valid. Char OK, int OK, enum not OK, pointer not OK. This is better than prohibiting binary reads or memcopy, because programmers will not be tempted to bypass it.
- Casts into non fully mapped types are prohibited. If you need to convert something to a non fully mapped type, it requires a constructor, with checking.
That gives a sense of the general idea. Do enough analysis to see if something iffy is safe, and prohibit the cases which are not easy to show safe.
I see this claim all the time, but can you give some examples of large C++ projects that don't constantly struggle with memory safety issues? (And are looking for them, of course.)
It only works when everyone plays balls and doesn't do C style coding, ever.
Usually that can work in small teams with security minded individuals but it isn't a given.
And then there is the little fact that on most surveys, the amount of answers referring any kind of static analysis tooling are usually around 50%.
What C++ has definitely going for it, is having the type system tools to write much safer code than plain old C.
However its copy-paste compatibility with C is also what hinders any attempt to force people to actually only use those better features.
The only way to fix this is having systems programming languages being adopted that aren't C at copy-paste level.
Other would be some kind of Safe C, but both WG14 and C community in general have voted against such improvements.
If that’s the problem, and solving that would solve all of C++s memory-issues, why have no-one made a compiler option to simply make that code illegal? A -EUNSAFE or whatever?
C++ was also born at Bell Labs, and due to that, all major C compiler vendors quickly started shipping C++ on their boxes as well.
If you take away copy-paste compatibility you might be better off doing D, C# or whatever safe variant already exists.
Which is what many of us have done, to move to type safe languages, and only use C and C++ at the boundaries, in small pockets of unsafe code.
In fact if you look at mobile OSes, that is the reality for app developers, C and C++ are no longer the full stack languages they were 20 years ago, rather used for the kernel, drivers, compositor and shading languages, but everything else happens in safer languages.
And the SDKs only allow you to write libraries, not full applications.
Naturally are clever developers that subvert the workflow and transform the libraries into the actual application.
One problem is that C++ can directly include C headers of the operating system, while languages that aren't copy-paste compatible with C have to create some kind of wrapper library for them... this is a major reason for the success of C++, but it also makes improvements of this kind far harder to deploy in practice.
So if anyone is claiming you could write real-world safe C++ if you wanted to, they are making a false claim then?
Once you have that to build on top of, such a compiler flag could make sense... if it were possible in C++, which I'm not sure about.
See for example this criticism of one such effort: https://robert.ocallahan.org/2016/06/safe-c-subset-is-vapour...
It’s obviously still too early to declare a winner, but to me this sounds like a turtle slowly but surely overtaking a rabbit.
This is never acknowledged by the C++ people: Idiomatic C++ is not suitable for formal proofs, if you don't believe me, ask Xavier Leroy.
https://news.ycombinator.com/item?id=23290030
I would argue that the closer you stay to C while using the good features of C++, the safer the code is.
And obviously well written and debugged C code is nearly always more robust than C++ code. But for ideological reasons most people here are unable to acknowledge that, perhaps because they cannot do it.
- implicit conversions
- decays from enums to integers
- implicit conversiosn from integers to unexisting enums
- no bounds checking
- implicit conversions between pointers and arrays
- no proper way to ensure a given array length is valid as part of a function parameter
- null terminated strings, that occasionally aren't terminated
- abusing null terminated strings with clever algorithms, e.g. strtok()
- the preprocessor
- const that isn't really const
- variable arguments that require getting the macro type arguments
- UB explored to the last possibility of code optimization
- no safe way to deal with output parameters
- typedef don't introduce strong typing
All of that came from C, not C++.
Manageable in C. Integer conversions are only a tiny fraction of all the other implicitness in C++.
> decays from enums to integers
Compiler warns. Recent real compilers like gcc even have exhaustiveness checks like OCaml.
> implicit conversions from integers to unexisting enums.
Compiler warns.
- no bounds checking
Reason about that and implement your own scheme. Or prove.
> implicit conversions between pointers and arrays
Have not seen a single bug due to that in more than 1000000 lines of C.
> null terminated strings, that occasionally aren't terminated
Have not seen a bug due to that, this is the canonical example of an overblown hypothetical threat.
> abusing null terminated strings with clever algorithms, e.g. strtok()
Prove the algorithm or don't use it. Hint: As far as proofs are concerned, NUL terminated strings are like Lisp lists terminated with NIL, hence a well-founded data structure that is easily amenable to proofs (unlike C++ constructs).
- the preprocessor
Rarely introduces anything and is still required for C++, especially in sane test suites.
I don't think all that came from C, things like typedef being an alias rather than a separate type seem much older.
You are again just throwing dirt at C, mocking all people who write actually robust buzzword free software.
You are ignoring that C code is much easier for formal proofs that C++ code (the kind that you advocate).
Which happen to be written in C with several layers of code review and static analysis.
So by your reasoning those 32 years have not happened, in spite of being so easy to prevent exploits in C code.
Yeah, right.
All while their own industrial strength C++-OS (according to you) never has any exploits.
The fact that you are singling out Linux shows that you are only interested in throwing dirt.
I wonder why.
Here is a little tip for you, Microsoft has been acknowledging security issues with C and C++ since the XP SP2 days.
Which is why Windows happens to have plenty of mitigations that only recently FOSS UNIX clones are catching up to.
Yet they have come public that hasn't been enough, hence the migration effort away from C, enforcing programming guidelines with C++ and coming up with plans to migrate to safer systems programming languages.
Guess which OS vendor is now having first party support for writing GUIs in Rust?
But I can also rephrase what Oracle, Apple and Google have stated in the same vein regarding OS security.
Or maybe you prefer the statements of an UNIX hero instead?
I am always baffled by the people that supposedly write modern C++ professionally and constantly have memory safety issues. Most serious projects won’t hire you if you aren’t capable of writing memory safe code in your sleep, it is a basic skill.
The reality is that there isn’t much opportunity for memory safety issues to occur anyway, the type system and scheduler do most of the heavy lifting. Similarly, concurrency safety isn’t much of an issue because threads barely interact. Most high-performance server software looks this way these days.
Bugs tend to be of the boring logic variety that can happen in any programming language.
> Similarly, concurrency safety isn’t much of an issue because threads barely interact.
I think the domain you're working in isn't as susceptible to memory safety issues, but that doesn't mean they're not present. It also sounds like they domain you're working in is trivially parallelized if threads "barely interact", which limits your exposure to those memory and data race issues.
If you gave an adversary access to the API of your kernel however, how long do you think before they found a use-after-free, double-free, stack or heap overflow, etc? Days, weeks, or hours?
If you haven't run a fuzzer yet, I wouldn't be so confident.
No one can know if it is bug free, that is impractical. But it also isn’t like this is a weekend hobby project either. Most of the bugs that get out are in unimportant peripheral code and integrations.
I'm quite curious what tools you're using to formally verify your C++ code if you are. My understanding is that in general you can't, which is why msan/asan exist, to get a first approximation of verifying things that can't be formally verified for most C++ code.
(In general, I'm dubious of these claims that "Most serious projects won’t hire you if you aren’t capable of writing memory safe code in your sleep", because I think if you asked the majority of the members of the C++ committee if they could do that, they'd say no).
In many of these systems it is standard practice to generate arithmetically limited types pervasively. This is almost transparent in C++17. While it is possible to verify much of this at compile-time in theory, it almost never is because it isn't worth the effort (C++20 may start to change this) and testing at runtime has proven to be nearly as good. People underestimate what is possible with the C++ type infrastructure in this regard.
This type of software design was originally done because it allows for exceptional performance but has become popular for safety reasons. It uniquely allows you to make guarantees about runtime behavior under diverse adversarial workloads that would otherwise be difficult to make.
Bugs in practice tend to occur at the interface with third-party code, which requires dropping out of any internal type system, or in the form of performance anomalies due to unexpected hardware behaviors interacting with the scheduler design. Logic bugs in the core bits tend to be found in testing.
Presuming that this style works equally well for all software seems presumptuous and perhaps naive, does it not?
Sadly as codebases get more complex, that vision gets further and further from reality.
Safer languages are better tools because they make obvious when you're stepping out of the safe zone, which is not the case with C++, even modern releases.
The “sort of” is what matters though:
- With std::unique_ptr, you trade use-after-free for use-after-move. Rarer, but still a threat.
- std::shared_ptr is costly which makes it unsuitable for some uses, and you can still have data rave with it if you don't use it correctly.
https://github.com/Microsoft/GSL
Stroustrup talk from 2016: https://www.youtube.com/watch?v=JtMPGwA3MzQ