Chromium project finds that 70% of security defects are memory safety problems
chromium.org
chromium.org
Interestingly enough Microsoft arrived at the same number. I think it really stresses how hard it is to reason about memory manually. I'm still surprised how much stuff is being written in non-safe languages even if people could get away with a managed language.
One thing that I just recently noticed when scrolling through some github repos is how much new software in the linux ecosystem is still written in C despite it being probably avoidable, like flatpak.
Real dynamic linking on that (dlopen-style) can be a real nightmare, specially between compiler versions.
You can also add that Rust do not have a runtime, which is in many cases an advantage, but also tend to make libraries bigger.
Though the average one is probably harder, yes. I'd wager that it's mostly lack of care or motivation though, made a bit worse by common language features.
C++ has a stable ABI. Fragile, but stable. Or doing things like updating libQt on your Linux distribution without recompiling half of the world would be close to impossible :)
I'm not aware of any plan to not stabilize Rust's ABI, it just hasn't happened yet. It's completely fair to label that a deal-breaker for using Rust, but trying to draw a hypothetical box around it with the label "its ABI will be difficult to create and/or use" seems a bit unwarranted.
If you're using C or C++ "for performance" and not using a profiler in day-to-day development, you don't need to be using C or C++ and should be using a memory-safe language instead.
There is certainly a need for the industry (and hobbyists) to take stock of both what is necessary and what is just desired. I just don’t want to see the bar for using a language that doesn’t handle all aspects of memory usage and performance to be restricted to soft real-time applications and infrastructure projects. Digging into low level programming can be fun and rewarding.
My problem is essentially that even if I am not absolutely concerned with maximum performance across memory usage, binary size, CPU efficiency, bandwidth, etc. I still truly enjoy the options for control (or the illusion of it) that is provided by C. I enjoy memory layout design, allocator design, cache-efficiency considerations and the like; while at the same time, I don’t have a huge love of malloc and free or having to track down segfaults. I think ‘performance-by-default’ is a viable language design goal and want more of it in my tools.
I keep posting comments mentioning my pet language project, mostly in hopes that when I see it on my threads display I can continue to shame myself into getting it released on time, and this is one of those. These kinds of concerns motivated me to design a language for personal use that gives me what I want, but doesn’t require (but can) allow me to deal with other things. I enjoy using Coq, Haskell, Idris, Lisp, ATS, and Clean. But those language deprive me of things I really do enjoy. So I’m going for a low level language (in the Perlis sense) that has an experimental type theory. This will certainly allow for a memory safe subset and is the kind of thing I want to see more of from others.
The adage is named after Niklaus Wirth, who discussed it in his 1995 article "A Plea for Lean Software".[1][2]"
We could do better languages. But it just takes insane resources to compete with the others ecosystem. IDE's, the Pandas, the format on save.
I've been a professional C++-developer in the past and one of the great things is the tooling and the just sheer knowledge of my past C++ colleagues. They know what happens in the kernel, they know the performance optimisation tricks, they know a lot, because the knowledge they gather just have a longer lifespan.
Ask a JS-developer and they will tell you all the web frameworks that came before React and their quirks...
C++ developers may be smart people because you have to be smart to do anything in C++ - imagine if the same brainpower that's being spent tracking memory usage and pointer/reference distinctions could be put into your actual application logic instead.
However: I'd argue that accidentally quadratic algorithms are easier to hide in a concise language. Writing out a quadratic loop explicitly takes space, and that space alone might make people pay more attention than some subtle implicit language construct. Either way, the most common source of unintended quadratic (or higher) behavior is helper functions and library calls.
The other thing to keep in mind when it comes to algorithms is that cache behavior and therefore memory layout matters a lot for performance on modern hardware. Managed languages really stand in the way of optimizing memory layout, which can be a systematic performance disadvantage compared to C++.
I do hope we get some more innovation in the design space occupied by Rust, where you get fairly explicit control over memory layout, but still have statically checked memory safety guarantees.
I disagree. When every loop is full of cruft around setting up the iterators, it's easy to drift past what's actually happening. In a language where looping over a list takes a single syntactic token, it's a lot more obvious when you've nested several such loops.
> The other thing to keep in mind when it comes to algorithms is that cache behavior and therefore memory layout matters a lot for performance on modern hardware. Managed languages really stand in the way of optimizing memory layout, which can be a systematic performance disadvantage compared to C++.
C++ doesn't really make cache behaviour clear either though. I agree that we need better tooling for handling those aspects of high-performance code, but they actually need to come from somewhere lower-level than C++.
The real problems tend to come from where the quadratic behaviour doesn't come from nested loops, but from library calls. The canonical example of this is building up a string with successive string concatenation in C.
As for cache behaviour, C++ allows you to control memory layout, which is really what's required there, while most managed languages don't give you that control at all.
We live in a fallen world. In a large enterprise codebase there will almost certainly be parts that aren't indented correctly. And even if everything is indented perfectly, the sheer amount of stuff in a C++ codebase makes everything far, far less obvious.
Correct me if I'm wrong, but I doubt you've ever written a program in modern C++?
It's not easy to see what algorithms you're using in a C++ codebase, because most of the lines of code are taken up micromanaging details that are broadly irrelevant. Yes, C++ makes it easy to tell whether you're using 8 bytes or 16 in this one datastructure. But you drown in a sea of those details and lose track of whether you're creating 10 or 10,000 instances of it.
As for algorithms, I honestly don't know what you mean. They're all documented online with respective big-O running times[1]. If you're talking about making unintended copies of things, then yes, C++ does expect you to know what's going on... it's the whole point of the language. If that's too much for you then don't use it, but that doesn't make it a bad language (I'm not denying it has some hair-pulling moments) Use std::move() when appropriate.
[1] https://en.cppreference.com/w/cpp/algorithm/sort (see 'Complexity' section.
Then yes, I work on modern C++ codebases.
> As for algorithms, I honestly don't know what you mean. They're all documented online with respective big-O running times[1].
Nontrivial programming tends to involve implementing, or even inventing, some algorithms yourself.
If performance is on your mind constantly, why wouldn't you choose the one with the least restrictions on what you can achieve?
It's not like those people would find joy in being locked into the JVM instruction-set.
Yet fairly modest requirements like 4000 concurrenct connections or a 16gb cache can be hard to achieve.
We almost exclusivly write for the JVM. I see systems with 20 - 30 - 40 jvms with huge CPU and RAM demands, to meet modest perf requirements.
Sometimes I do wonder if we shouldn't write a little more carefully written C and a lot less memory safe Java.
I would lay money that the language isn't your real bottleneck. Switching languages might save you a factor of 2, using better algorithms or datastructures can save you a factor of 1000 or more. How much profiling do you do?
That's in my mind wrong for at least two reasons:
- If you are making building blocks (libraries) for other languages, you have to use C or C++ (maybe Rust soon). They are currently the only languages that can be bridge to the rest of the world (Python, Ruby, JS) without loosing your mind.
The main reason is that they are without GC, meaning deterministic destruction of object, meaning easy to interface with language with GC.
- When you aim for high performance you don't profile day 1, that's a complete waste of time. You will never transform every of your function in a critical kernel. You profile when you reach a performance bottleneck and you need it.
Sure (though there are often other ways to achieve the thing you actually want to do). But in that case you're not doing it "for performance".
> When you aim for high performance you don't profile day 1, that's a complete waste of time. You will never transform every of your function in a critical kernel. You profile when you reach a performance bottleneck and you need it.
In that case your performance requirements are not extreme enough to justify using C/C++ for your whole application. Write it in a safer language, when you hit performance issues profile and optimise, and maybe drop into C/C++ for those few "critical kernels" in the unlikely event that it turns out you actually need to.
Again : no. That is in my experience both wrong and over-idealist.
For many applications, the overhead in term of dev time of writing bindings for every of your compute kernels + the pleasure of debugging the problems associated with them and heterogeneous build chains is generally several order of magnitude higher in term of man-hour-cost than just doing your program entirely in C/C++/rust.
There is an other aspects generally ignored:
- Theory say that often 80% of the compute time is consumed in 20% of the code. That's often wrong, and many HPC simulators do not have any kernel taking more than 3-7% of total run time. Consequently, everything might one day need to be optimised.
- Many performance critical algorithm are state of art and evolve. Meaning your innocent little function in "memory managed" language might become tomorrow a new bottleneck. And you do not want to have to rewrite that all the time.
A lot of new devs are wrongly scared of manual memory management where it became a non-problem with RAII in C++[1-2]x or the borrow checker in Rust.
And generally the ones that are scared are the one that do not use it.
The mental overhead with memory in C++ does not come so much with the lifetime of object, it comes mainly with HOW to use efficiently YOUR memory: aligned object in memory, cache effect, indirection, cost of polymorphism, allocation, etc, etc.
All these aspects, you do not think about them in memory managed language because you can not: You have no control over it.
And that's also why they bite you in the face in term of performance, generally much more than the 2x you quoted before. Just a remember, a cache miss and it's ~200 cycle you loose
> - Many performance critical algorithm are state of art and evolve. Meaning your innocent little function in "memory managed" language might become tomorrow a new bottleneck. And you do not want to have to rewrite that all the time.
You can't have it both ways. If it's really common for everything to become a performance bottleneck, it's worth profiling from the start so that you avoid having major pitfalls anywhere. If it's rare and exceptional, FFI for those cases is fine.
> The mental overhead with memory in C++ does not come so much with the lifetime of object, it comes mainly with HOW to use efficiently YOUR memory: aligned object in memory, cache effect, indirection, cost of polymorphism, allocation, etc, etc.
> All these aspects, you do not think about them in memory managed language because you can not: You have no control over it.
> And that's also why they bite you in the face in term of performance, generally much more than the 2x you quoted before. Just a remember, a cache miss and it's ~200 cycle you loose
On the contrary. Plenty of people do those kind of things in, say, Java. They require knowing about compiler internals and using unsupported hints, or even bypassing parts of the compiler. But so does controlling these things in in C++.
You can have it both way, code evolve. It is pretty common in performance critical code that a minor, almost never call function in one scenario, become a performance critical bottleneck in an other. If you already played with large scaled simulation software, this happened almost every week depending of your inputs on what you are interested to simulate.
> On the contrary. Plenty of people do those kind of things in, say, Java. They require knowing about compiler internals and using unsupported hints, or even bypassing parts of the compiler
No it's not again. I have been developing in C++ for 15 years, including in the HPC world and I (close to) never had to touch a compiler internal. The language gives you what you need for performances, you do not need to play with that.
At the opposite JIT compilers like V8 or Java are monster of complexity very sensible to side effect [1] and controlling things like "Does this data fit in my L2 cache ?" in them is close to impossible because even things like "Where are my data and what are there size ?" is an hard question.
Once again, there is theory and there is practice.
Theory is what you say. Practice is that 98% of performance critical software in HPC, game industry, physics and High Frequency Trading is in C++/C (maybe Rust soon). And this is why.
And game programming isn't normally "performance critical" so much as OS-less and embedded. Sometimes there's real-time constraints, but very few parts of a AAA game are the inner rendering loop that has real time requirements. Mort of it if boring stuff that could be (and often is!) written in python or more often Lua.
In which case you're in the "worth profiling from day 1" world. It's much easier to work on the performance of code when you're already working on it and have it in your head - particularly in a verbose language like C++ where it takes a relatively long time to comprehend existing code - so if there's a decent chance that the performance of this code is going to be important in the future, profiling as you write saves you time overall.
> No it's not again. I have been developing in C++ for 15 years, including in the HPC world and I (close to) never had to touch a compiler internal. The language gives you what you need for performances, you do not need to play with that.
I said be aware of, not touch. If you weren't doing things like memory alignment pragmas then I guess your performance requirements were never so stringent. Fact is that a Java program that's fitting its data into L2 or avoiding cache line aliasing will blow a C++ program that isn't out of the water.
> Theory is what you say. Practice is that 98% of performance critical software in HPC, game industry, physics and High Frequency Trading is in C++/C (maybe Rust soon). And this is why.
HPC/physics follow questionable development practices in a lot of areas, and the games industry follows questionable everything practices. HFT uses a lot of Java and even higher level languages. C++ survives because people are rewarded for being seen to put a lot of effort into performance, and are not rewarded for avoiding bugs.
That's pure bullshit.
C++ survives because, even in 2020, it does the job.
Most people criticizing C++ are still stuck in there mind with C++98 and its quirks.
C++ evolved and modern C++ is at least as productive as Java or C# when used correctly.
That's why he is still actively uses and continue to grow.
This message just translate at best, your feelings (as a Java/scala developers ?). And you allow yourself to insult both the HPC Industry and the game industry without even providing metrics based on your "feelings".
You're claiming this in a thread about how a flagship project from one of the biggest names in the industry found that 70% of their security bugs were things that wouldn't have happened in Java or C#. Your statement may be true, but only for a kind of "correctly" that doesn't actually exist in practice.
> That's why he is still actively uses and continue to grow.
Where are you getting those stats?
> This message just translate at best, your feelings (as a Java/scala developers ?). And you allow yourself to insult both the HPC Industry and the game industry without even providing metrics based on your "feelings".
Would you defend either of those industries as a haven of good coding practices? Do you believe that they have fewer bugs, make better use of up-to-date tools, make more data-driven decisions, than other parts of the industry? I'm repeating a reputation rather than a specific metric, sure, but does anyone actually dispute that reputation?
Which is a project which is born in 2008, and still ship codes from the 90's. That also include one of the most optimized (meaning complex) piece of code world wide: V8.
You have nothing that come even close to the complexity, usability and popularity of Chrome in both Java and C# world. Ironically, even Microsoft uses a C++-core in its software, including MS Office and Edge. Maybe you should reflect on that.
> Where are you getting those stats?
https://tjpalmer.github.io/languish/#y=stars&names=java%2Cc%...
You're welcome.
>Would you defend either of those industries as a haven of good coding practices?
Every industry has domain driven standards in term of coding practice. They all have their reasons based on deadlines, usages, iteration cycle, developer backgrounds, safety.
Pretending that one culture is superior to the other is both pretentious and let appear a bad misunderstanding of the world we are in.
Now this is my last comment on this thread. I do not think you are open to any discussion.
Back in 2008 C++ advocates were saying the same thing: all those errors are only in old codebases, modern C++ doesn't have those problems. At what point should we stop believing it?
> You have nothing that come even close to the complexity, usability and popularity of Chrome in both Java and C# world.
Nonsense. There are dozens of more complex, more usable, and more popular systems written in Java and C#.
> Ironically, even Microsoft uses a C++-core in its software, including MS Office and Edge.
In the older projects that they're most conservative about, yes. Large companies change slowly. Doesn't mean what they're doing today is wise.
> Every industry has domain driven standards in term of coding practice. They all have their reasons based on deadlines, usages, iteration cycle, developer backgrounds, safety.
Which is to say that good development practices will be a lesser or greater priority level in different industries.
For one, not all software is like Chromium. If you look at something like OpenSSH, the vast majority of their security holes have nothing to do with memory safety and are just logic bugs (often caused by features that somebody added that aren't core to the basic SSH experience, e.g. code that interfaces with X11 or something) or protocol weaknesses. (http://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=openssh)
The other effect is that in practice, memory-safe languages can come with security baggage of their own. If you look at the zillions of security holes in something like Rails or Wordpress or Django, a fair portion of them relate to an attacker's ability to invoke sophisticated-but-unintended behaviors that are more likely to be hiding in these managed languages (and their support libraries) than in something like C++. E.g. CVE-2013-0156 or CVE-2013-0277 or apparently any current Python or Ruby program that, even today, calls yaml.load on untrusted input. That kind of "security hole from unwanted latent functionality" is less likely to exist in C++. (I realize this is a contrarian view not shared by the vast majority of PL/security experts, but the ones I hang out with seem to interpret "memory-safe language" to mean "expert-written Haskell or, if you want to slum it, Rust" and are not thinking "random Ruby/Python/JavaScript/enterprise Java".)
Not to mention the countless high-profile security holes that have nothing to do with memory safety, e.g. Shellshock, "goto fail", Lucky Thirteen, BEAST, CRIME, POODLE, FREAK, Logjam, etc. Or bugs that are very relevant to Chromium but aren't really Chromium's fault and probably don't appear on their own list, e.g. Meltdown/Spectre etc.
In practice what I’ve found is that people prefer to deal in absolutes. A large reduction in a category of bugs isn’t enough, it needs to be eliminated altogether, they say. If it’s still possible in a contrived example, what’s the point of investing in switching?
A thing can happen zero times (never). An example might be 1+1 =0. You could get really unlucky with cosmic rays or something, but really adding two registers, both containing 1, that result isn't going to be zero unless there's some sort of hardware failure.
A thing can happen once. an example might be, you can delete a file once. There are ways to get unlucky, of corse, but once it's unlinked it's gone.
if it happens more than once, you really should probably think about an unbounded number of times. now days, that sorta means 2^64 times. There will be bugs when something overflows int64, but I hope you get the gist.
The parent comment is talking about invariants you can use in an algorithm.
I think you might be worried about python vs C or something along those lines. Really, it should all work with a pencil and paper. Which is obviously going to be slower than pushing around electrons. But if you can find those invariants, 0-1-many, you can make a better algorithm. If you're stuck with a pencil it'll still make that faster.
Risk calculations are far broader than infosec and don't deserve the dismissiveness you seem to be casting towards them. Risk calculations are the core of business. Almost every decision a business makes is a risk calculation; every action has an opportunity cost if it isn't intrinsically risky, and actions with certainty are very rare.
(For the avoidance of doubt, I believe that use of memory-unsafe languages should be avoided if reasonably possible, but there are still plenty of reasonable reasons to use C, C++ etc. instead.)
I feel this thread is getting out of hand. I initially replied to say why, in general, the kind of thinking that makes you unsatisfied with reductions but not eliminations of concerns, is common among programmers - because it's a sound heuristic. Reducing is good, but eliminating is better.
As for having more time to write a fast program... that's funny. If you want a fast program on something JVM based you're pretty much going to be spending the majority of your time writing things in a way where the GC plays as little role as possible.
Sorry, this is not hand wavy nonsense. And what you are providing is called annectodal evidence.
Also, your universe seems to consist only of the JVM as a memory-safe alternative to C++. Yes, there are a lot of bloated, badly performing programs implemented in Java. However this isn't a given. Yes, some design decisions for Java introduce the risk of bloat, but you can avoid them at much less effort (and risk) than memory corruptions and new features like the value classes are reducing the bloat quite a bit. But still, the JVM is extremely high performance, so for speed in surprisingly many cases, it often beats C++. Virtual method calls are just better optimizable at run time, the Java JIT creates excellent code and Java has some of the best garbage collectors, so at really dynamic memory loads, it beats any manual management by a wide marging.
And of course, there is a whole world beyond Java as alternatives. Rust has been explicitly designed to excel at tasks C++ traditionally shines for, while giving your full safety.
Rust is certainly interesting and it's on my radar. I wonder though, when it comes to having it in use in anger if its guarantees turn out to be over sold, just like the JVM's safety claims were. Time will tell.
Edit: I tend to focus on comapring against the JVM because pretty much any framework you use on The Cloud is JVM based. I'm of the opinion that there are cost savings to be had if these were ported to more appropriate languages, hence the Cassandra vs Scylla comparison. The money saved was 'noticeable'.
[1] https://github.com/LMAX-Exchange/disruptor [2] http://mishadoff.com/blog/java-magic-part-4-sun-dot-misc-dot...
Java 15 just accepted the JEP for native memory management, yet another stepping stone for having value types support.
If the cadence continues, Java will eventually have all the features that it should have had in 1996, had Sun properly taken into consideration languages like Modula-3 and Eiffel.
Which you can get today in a language like Swift, C#, Nim or D, productivity of GC, type safe, while having the language features to do C++ like resource management.
You'll never actually see anyone tout the benefits of using GC for this, though, because the performance characteristics of persistent data structures are so horrendous compared to mutable ones, no one actually uses them in C++.
Of course, repeatedly allocating and freeing is poor for performance. Cache/pre-allocate when you can. This goes for managed languages too.
Stack is also very hot in cache. Memory that GC is handing allocations from is not.
I recently ported some of Cassandra code from Java to stack-only Rust and I got ~25x performance improvement, most from avoiding heap allocations and GC.
Both Java and C# may be somewhat slower, but the maintainability and freedom from memory management issues more than makes up for this.
Any engineer worth his salt will take this into account.
If you unlucky and given GC is not well optimized for given workload, then memory usage overhead can be huge. Just a few days ago had such problem with Go (HeapIdle grew until OOM).
Go GC is less mature than Java GC, but Java GC is not free too - sometimes you need to spend a lot of time time to tune GC settings or to optimize code to avoid GC problems.
In my case case it would be faster to use malloc/free than to spend time fighting with GC (looking for a workarounds).
On average GC saves development time and allows to avoid memory management bugs, but in some cases overhead is big and developers have to spend more time, not less.
https://www.ptc.com/en/products/developer-tools/perc
https://www.aicas.com/cms/en/JamaicaVM
One of the Java vendors acquired by PTC, Aonix, used to sell real time Java implementations for military deployments, including weapons controls and targeting systems.
You don't want a couple of ms pause when playing with such "toys".
If you mean the `synchronized` keyword, then correct, they don't have that. Most languages do not have that concept. C++ does have mutexes and has had them since C++11 (nearly a decade). C++ also has as part of the language spec the concept of threads, again, there since C++11.
If you meant co-routines, then C++ just added them with C++20.
Or do you mean something different like green threads (ala go)?
C++ certainly has concurrency constructs and has been expanding them since C++11.
The entire design of C++ is to enable efficient libraries for these kinds of things to be built. And it does, and has, and will continue to for decades more.
And ?
C (even Fortran) had threads and was used to create high performance program with high degree of parallelism before any "concurrent" modern language was even born.
You can not agree if you want. But fact are there, 98% of programs running on the biggest "parallel" machines nowadays (supercomputers) are C, C++ or Fortran.
You don't need to be "designed" concurrent to be efficient at it. The same way you do not need to be designed "Cloud-native (bullshit)" to run on a virtual machine.
We have to give programmers, with different backgrounds & training, tools to write high performant code in their everyday job. Many of the tools we use today are not designed for that. We are stuck in a mental model 50 years old that is no longer true.
Here are some interesting stackoverflow answers. You are of course free to dismiss these answers as anecdotal.
So much that in the 90's there were vendors that could make a business out of selling specialized memory allocators.
https://inf.ethz.ch/personal/wirth/ProjectOberon/Sources/Ker...
The malloc() exposed in a ISO C standard library implementation isn't necessarily related to whatever means the OS does memory allocation.
What C and C++ have going for them is 40 years of investment in optimizing compilers, to detriment of other languages, and abuse of UB in optimizers.
Now with shared compiler backends becoming a mainstream thing across most OSes, that advantage is getting thinner.
[1] handles could be unboxed to expose a pointer. But unboxing had machinery to detect and trap on bad handles
Edit: that static allocation was essentially a centralized store. Everything repeats itself.
https://floooh.github.io/2018/06/17/handles-vs-pointers.html
You've already been greyed out, but I agree 100% with what you said.
That's my main beef with this kind of bi-weekly discussion where people always put C and C++ in the same bag.
I rarely, if ever, make memory management errors in C. The model can't be more simple. If you don't explicitly allocate, you don't have to care about it (in the general case, it is on the stack and will disappear when exiting a function, anyway you also don't have to care about those implementation details, just do as if it disappeared with scope). If you allocate something (on the heap), then you free it later. Malloc(), free(), end of story. Okay, don't keep pointers to areas you will free or realloc, of course. Any function which returns pointers tomemory chunks are always to be treated as malloc'ed, those objects can never be on the stack and never have a fancy automatic management of any sort.
Now almost each piece of code I wrote in Objective-C contained memory errors/leaks. Be cause 1. I didn't grasp the more complicated, less deterministic (so to speak) models and the jargon well 2. you (well, I) never know right away which type of management was used for the object that some function returned, you have to check the doc and pray it is clearly written 2.b you have to keep that information in mind until the moment when you'd like to release the said object in order to release what should be released and not release what is/will be/has been automatically released now/later/sooner/whoknows.
And I don't even talk about the abomination that is Objective-C++, where I drowned in muddy waters.
This is where I make the case for having knowledge of asm/C/C++. The JavaScript developers that started on JavaScript or, maybe, Ruby have no clue what the hell is going on in the language they are working in. I see this all the time. From senior and lead developers, too.
They have no mental model for how pointers work. Which means they have no concept of what an object is, on a memory-organization level. They aren't always aware of when mutation will occur, and they are entirely too reliant on operations that are inefficient. In Lisp, this would be the developer than uses CONS everywhere. There is no awareness of the allocations behind the mechanisms. There is premature optimization and then there is not driving into the damn pothole in the first place! One of these is F1 racing and the other is basic driving skills.
I believe this may also explain why so many developers are so bad at using git. Git is entirely based on pointers. There is an elegance and simplicity to git that is lost on many people.
In the old days people usually chose to become a programmer for a reason, today it is like any other job.
Just like you can be a very effective C or C++ developer without knowing anything about css (which in my opinion is far more difficult to understand than when and how stack vs heap allocations are made)
Yes, exactly.
https://kkimdev.github.io/posts/2019/04/22/Rust-Compile-Time...
Of course, you can bypass this by using 'unsafe' code, where you are explicitly telling the compiler that you have manually verified Rust's safety guarantees.
This is why a significant portion of Rust's community gets agitated with (possibly unnecessary) use of 'unsafe' in popular libraries.
Considering many of the software problems we face have some connection to thread-safety, Rust is the way to go.
Threading seemed a good idea that went wrong.
Then you are back into the same synchronization issues as in every other language.
I'm trying to learn Android and some libraries give zero indication of what thread your callbacks will be called on.
Maybe there's a convention somewhere that says "On Android, you can't assume anything about callbacks, so always assume you're on some anonymous worker thread and lock everything or dispatch to a thread you own"
But I have not come across it yet.
You cannot assume anything about process or thread lifetimes on Android, as the simple act of rotating the screen will restart your application, and it can be killed at any moment and restarted later, due to memory pressure, or because the use has switched applications.
So whatever was the state of your application can be completely different when the callback is supposed to be invoked later on.
Does that actually exist? (I mean in the general case, not just a few carefully chosen benchmarks).
I find this very hard to believe.
I don't claim that Java is always faster than C++, that would be silly and there are plenty of Java programs in the wild which proves that it isn't. But there are quite some tasks at which Java is indeed faster.
However, in my experience, using virtual dispatch is a relatively rare ocurrence in C++ (compared to the vast majority of method calls). On the other hand, on Java, most of _everything else_ is indirect and has overhead: All objects are allocated on the heap, primitives (int) are often objects (Integer) where they ought not to be, all objects have 16 bytes of overhead, etc.
But the JVM will convert those heap allocations to stack allocations! And it will realize those Integers are used as int and remove the overhead! And it will realize you're not using the information on the header of every object!
Perhaps in synthetic benchmarks, but in real programs, where there are an almost infinite amount of code paths, dumb 'data transfer' objects are common and things need to be modularized, the JVM is forced to assume the worst case can happen (even if you as an human can prove that it won't happen) and inhibit those optimizations. And now you have indirect accesses everywhere, memory overhead (=cache trashing) everywhere, the runtime can't vectorize that tight due to Integers, etc.
In fact, I can't think of any domain where there is heavy competition and where high performance is a determining factor where Java has won to C/C++. In browsers, it certainly has not.
The number of instructions is much less important than what they are doing. In the case of virtual dispatch, it's doing a memory lookup. If that memory is in cache it could be relatively inexpensive but not guaranteed. However, if you have to hit main memory then things are much slower.
> But the JVM will convert those heap allocations to stack allocations!
I was surprised recently to learn this is not the case. (at least, not with hotspot) The JVM will try and "scalarize" things (pull the fields out of the object which may push them onto the stack) but it won't actually allocate a full object on the stack (OpenJ9 will, but I don't see people using that very often).
It is also somewhat bad at doing the Integer to int conversion. That is mainly because the Integer::valueOf method will break things (EA has a hard time realizing this is a non-escaped value). Simple code like
Integer a = 1;
a++;
can screw up the current analysis and end up in heap allocations.There is current work to try and make these things better, but it hasn't landed yet (AFAIK).
> In fact, I can't think of any domain where there is heavy competition and where high performance is a determining factor where Java has won to C/C++. In browsers, it certainly has not.
I think the realm where the JVM can potentially beat C++ is, funnily, work that requires a lot of memory. The thing that the JVM memory model has going for it is that heap allocations are relatively cheap compared to heap allocations in C++. If you have a ton of tiny short lived object allocations then the JVM will do a great job at managing them for you.
Except C++ also has many other options besides malloc() each of these individual objects.
Of course you normally try to avoid writing C++ like that.
For example, if you use `std::vector` or `std::unique_ptr` you would be deallocating memory when those go out of scope. The JVM might actually never do that, e.g., if the program terminates before sufficient memory pressure arises.
Writing Rust code that leaks a `std::vector` is trivial, but doing the same in C++ actually requires some skill.
auto* v = new std::vector<char>(); vector<int> foo {...};
and heap allocate a vector on the heap with `new` (without using a smart pointer), and then avoiding your linters warnings about this (e.g. clang-tidy).If this is common in your place of work, I truly pity you.
Also, trading a call to `free` for a second call to malloc, and a second pointer indirection to the vector elements, isn't a very effective way of improving performance to beat the JVM. The whole idea behind leaking memory is doing less operations, not more :D
In Rust, leaking "the right way" is trivial (mem::forget is safe), but in C++, leaking the stack allocated vector probably requires putting it behind an union or aligned_storage or similar.
Some things, namely array indexing and RefCell borrowing, have unavoidable extra runtime overhead in the default safe usage, but it's unlikely that it is significant (often array bounds checks are either essential and thus needed in C++ as well or optimized out by LLVM in Rust, and RefCell usage is generally rare), and you can use unsafe unchecked operations as well.
You have a type for UnauthenticatedUser and one for AuthenicatedUser. You have an authenticate() method to convert from one to the other.
I'll wait.
authenticate() : role1 | role2
The type system is a tool to reduce the space of invalid inputs.
Here's an example: create a type, such as "RoleAuthorization". Make it a required parameter for access-controlled methods, either through the constructor of the parent object or explicitly in the method.
In order to create a RoleAuthorization, you must first pass the user's Role object in an authorization method.
If a method requires authorization, the compiler will complain if you don't first check the current user's role through the authorization method.
> I'll wait.
If you're sincerely interested in discussing something, antagonizing someone this way (and implying that you are infallible) really doesn't help.
If you just wanted to make what you believe is a statement of fact, just make it.
If you don't want to discuss the issue at all and would instead prefer to be unchallenged on this topic, why write about it in a comments section?
Maybe type systems could, in some cases, prevent this type of problem. You could construct a system where performing a password authentication returns an object of type Authorized, and performing operations requires passing in objects of type Authorized. I don't know how often this is useful in practice though.
So we would have to make sure that type safety is ultimately backed by a hardware based encryption module, but that's just not where we are today.
I'm not sure this is a sensible way to compare - Heartbleed is a memory safety bug and is a bigger bug than the rest of the ones you've listed combined. The various terrifying nameless iOS exploits that have had actual real-world, documented usage are not clever crypto protocol bugs with catchy names, either.
This is why we have to abandon C, not only is the language unsafe, it has also driven hardware manufacturing to a bad state with a reinforcing feedback loop, the more C we write the more hardware manufacturers want to make you believe that we still are working on PDP-11.
But here is a great article
I mean, supposing that Chromium was rewritten in a safer language, have we any reason to believe that these memory issues would be replaced by a similar number of non-memory-safety issues?
>That kind of "security hole from unwanted latent functionality" is less likely to exist in C++.
Why? That seems like a non-sequitur to me. For me the real difference between the projects you list is that Chromium and OpenSSH are not, as you put it "random Ruby/Python/JavaScript/enterprise Java", they are heavily audited and have a lot of resources allocated into preventing the very issues you mention. Comparing OpenSSH to Wordpress and attributing their respective security track record mainly to the languages they use is fairly absurd IMO. If you're clueless enough not to properly sanitize untrusted input can I really believe that you'd manage to write safe C code?
This thread is oddly reminiscent of discussions around gun control. Because something is not 100% effective doesn't mean that it's not valuable. Although I guess we do need these unsafe languages in case we ever need to overthrow a tyrannical government... no wait, I'm getting confused.
I have heard people say "we rely on exploits to be able to exercise software freedom on locked-down platforms" and "if games were written in Rust, a lot of the glitches speedrunners use would go away and we'd lose our hobby" so... you're not always off.
Also you assert it is absurd to compare WordPress and openssh In the context of this argument why? They are both widely used software with very high stakes on their reliability and security. They show very well the contrast that even though one is memory managed and the other is not, that does not stop both having serious bugs. Actually it shows memory management is a don't care variable. On the other hand if chromium finds their project has much more memory management issues than other issues, then yes: memory management for the functionality of chromium seems to be a relevant facto upon which improvements can be done.
I meant that input sanitization is generally an easier problem to solve than writing safe C. If you can't do the former reliably, I'm certain you wouldn't be able to do the latter.
Input sanitization should generally be handled very close to the interface, once you went through this step you should expect to only deal with safe data. Memory safety covers the entire application, if some non-critical debugging function 10 level deep in the call stack messes up when it computes the size of a message buffer you can have a remote code execution vulnerability.
>Also you assert it is absurd to compare WordPress and openssh In the context of this argument why? They are both widely used software with very high stakes on their reliability and security.
Wordpress is a complicated ecosystem with multiple plugins that each can introduce security problems. Its attack surface is also incredibly large since it's generally public facing and anybody can interact with a lot of the features by default. That coupled with the fact that it's extremely popular make it a very good target for attackers. It's also a relatively fast moving target since the web changes fairly fast and new features have to be implemented regularly.
OpenSSH does mostly one thing and does it well. Its surface of attack is a lot smaller and it's vastly easier to audit and test. Its development is also overseen by the OpenBSD developers, who are famously uncompromising regarding safety. The feature set is very stable and changes very slowly.
>Actually it shows memory management is a don't care variable.
The absence of evidence is not the evidence of absence. How much development effort went into making sure that these issues didn't exist? How many thousands of hours auditing the code and making sure it did what people wanted? In a safer language this time could've been spent auditing other parts of the codebase or implementing other features.
I'd also like to point out that one of the biggest OpenSSH vulnerabilities ever found (in the Debian version of OpenSSH where the maintainer heavily reduced the generated key entropy by mistake) was indirectly caused by a memory issue since the reason for the patch was a false-positive returned by Valgrind regarding use of uninitialized memory.
There are tons of comments, including here on HN on a regular basis, which one could use to conclude that WordPress should be more secure than anything written in C no matter the quality. And yet, it isn't true in that instance. So it's fine to advocate rust or whatever, but let's not make that other, more dangerous conclusion.
I'm not sure which point you are trying to make, but yes, 100%-70% = 30%, i.e., there are many security vulnerabilities in Firefox, Chrome and Windows that are not attributed to memory safety by these projects, and preventing memory unsafety wouldn't remove all of them (at most "only" 70% of them).
Nobody is claiming here that fixing memory unsafety in software fixes hardware bugs, nor anything about OpenSSH or other projects, and your anecdotes do not show how many security vulnerabilities are caused in those projects due to memory unsafety.
Both of those methods still allow leaks, which is why Rust encourages RAII. [1] Are there other structured lifetimes we can get compilers to enforce for us, like they enforce certain invariants about flow control using control structures?
[1] https://users.rust-lang.org/t/memory-leaks-in-rust/18187/2
http://huonw.github.io/blog/2016/04/memory-leaks-are-memory-...
The distinction between unsafe and managed languages does not make any sense. You can for example have a safe C implementation. No need to move to another language.
Does the amount of unsafe implementations matter?
And when talking about C++, since 2011 there are smart pointers to help developers managing memory kind of automatically.
I believe most of the problems are due the people typing on those keyboards such bad code, especially when it comes to "smart" code.
In 2020 we have plenty of tools supporting the developer job (memory sanitizers, static analysis tools, linting, profilers, etc.): The big problem is failing to / ignoring to use such tools.
You are directly rewarded for "new" features (i.e. plagiarizing or rewriting old ones and selling them as new).
So I don't think the engineers are particularly talented, and the incentives are not the same as for SQLite.
The reason SQLite can test so thoroughly is that their requirements never change. I have tried to browse the web with SQLite but so far no luck.
Why do we see the version of Chrome reaching 100 at a sustained pace, instead of seeing something like Chrome 8.X or 9.Y?
The pressure is all about introducing new features (and thus, more bugs!) while fixing the existing bugs slowly.
And we should stop thinking that all the brilliant minds are working at FAANG, because it isn't true.
C is syntactically C++ being a subset of C++ and I've seen many programs which claim to be C++ but are actually programmed using C methodologies.
It's a case of "When in Rome, do as the Romans do".
> Over the course of its lifetime, there have been 69 security bugs in Firefox’s style component. If we’d had a time machine and could have written this component in Rust from the start, 51 (73.9%) of these bugs would not have been possible
Also interesting on the topic of memory safety
https://hacks.mozilla.org/2019/01/fearless-security-memory-s...
[1] https://hacks.mozilla.org/2019/02/rewriting-a-browser-compon...
At Firefox they already knew the problem they were trying to solve using Rust, they already wrote that software, already discovered many of the overlooked complications involved in writing a modern browser, so, in conclusion, even a rewrite in plain C would have solved many of the bugs.
The simple operation of rewriting the same software with previous knowledge of how it works usually leads to simpler code (at the cost of developers time)
Firefox is also the entity that invented Rust, so it's in their best interest to publicize it as "the final weapon" against bug, but "if we had used Rust from the beginning these bugs would not have been possible" is just wishful thinking.
Rust itself could not be there without the browser war and the pressure that contemporary web puts on software that runs it
And it'been C/C++ that has driven us there
C/C++ or it's developers couldn't deliver, so they went a very risky direction that led to success.
I think it's clever, it was risky but it seems to be paying off.
They didn't think just a rewrite in C would be enough, they didn't think any other existing language would be sufficient, and then they went off to design Rust. So the statement "Firefox is also the entity that invented Rust" kind of misses the point.
And that if I just took off my rose-tinted glasses, I'd realize my Rust code is buggy, unsafe, slow, and hard to maintain, and the only reason I'm using Rust is because of hype.
It was pretty hyped when I started using it, in 2015.
I didn't get the impression.
I understood that Firefox talks about their success in rewriting in Rust because it's their language, they control it and are the major sponsor and user.
I don't think Google or MS or any other company heavily involved in crafting programming languages for their own purposes will ever go that route for some of their core software, because they can't control the language and if they tried they would get the blame for trying.
Google and Microsoft are both already using Rust for real products.
There will never be Rust in Chrome or in Office.
Mozilla controls Rust because it's the largest Rust user.
Just like Google controls Go, even though it's opensource.
You said it yourself "the only real way to get a job working on Rust was to work at Mozilla"
There is a branch in the repository right now trying it out. Rust is also used in ChromeOS.
> Mozilla controls Rust because it's the largest Rust user.
Mozilla is not the largest Rust user, nor does the largest user control the language. Governance is consensus-based, and anyone is eligible to join.
> You said it yourself "the only real way to get a job working on Rust was to work at Mozilla"
I may have said that a long, long time ago, but it's not true today. The Rust team at Mozilla has been shrinking, and other companies have been letting folks work on Rust as part of their job.
And volunteers are like, 10x-25x more numerous than people who are paid to do so.
ChromeOS is a dead product.
> Mozilla is not the largest Rust user
Who is then?
> Governance is consensus-based, and anyone is eligible to join.
Same thing Google says about Go.
If Google is really interested in Rust I don't believe they will let the community keep the governance, I might be wrong, but better safe than sorry.
Is Google trying to take control over Rust?
You're not even aknowledging the fact that there could be different opinions on the matter, if I was you I wouldn't play the card "you're dismissing evidence".
You work on Rust, that's a fact, I'm giving you credit for it.
Can you say it doesn't affect your judgement at all?
Office will eventually make use of it.
That's good to know! It appears Microsoft can be both bad and good.
> Office will eventually make use of it.
I have to see it before I believe it.
Office team is not "developers who make third party Windows apps so Microsoft indirectly profit from them"
There was a time when MS embraced Java, we've never seen it in Office though.
And I bet you weren't reverse engineering Office to discover which ActiveX were implemented in J++ instead of VB 6.
so what?
MS also ships a complete Linux distro now.
> Java has had several talks at Build and has parity with .NET on Azure SDKs, Office doesn't dictate all business lines.
And Linux is the most installed OS on Azure...
I can buy milk from my butcher, but his core product is still meat.
you still fail to see the difference between what they offer to potential clients and what they use internally.
They are expanding the offer but are still a software house in the end.
It also means that MS is using its weight on free (as in free speech) technologies, like Google has done before with other OSS projects and we all know how it ended.
> Office doesn't dictate all business lines
Obviously, it doesn't.
It only generates 33% of the revenues and 39% of the operative margins.
The second largest segment for revenues, behind computing (mainly HW), the first for margins.
Cloud comes third - and last - for revenues and second for profits with a pretty strong growth - less than 2018 but still strong -, but keep in mind that they include the Office 365 online offer and Gaming cloud in the segment.
> And I bet you weren't reverse engineering Office to discover which ActiveX were implemented in J++ instead of VB 6.
Would you bet ten millions dollar on it?
The Rust re-write was the third attempt; the first two were in C++ and failed.
Have you contributed any code?
AFAIK most of Firefox is still written in C++, JavaScript and C [1]
The JavaScript interpreter, which is particularly important to users perceived speed, is entirely written in C++
Stylo has been the default CSS parser starting from the beginning of 2018.
It's good that Rust could have avoided them, but is it a fair comparison?
I think that when at Firefox they started to think about a new architecture to better enable parallelism they began improving considerably, Rust is only a part of that.
I've watched it several times over the past 2 years
And I've read the posts about writing and HTML engine in Rust when they first came out in 2014
https://limpet.net/mbrubeck/2014/08/08/toy-layout-engine-1.h...
and ported them to Elixir and still use them in my programming lessons
> Rust was key to the success.
For them
It's important to specify that Rust, built by Firefox, lead to a Firefox success.
Just like Dart, created by Google, is the language of choice for Flutter, also created by Google.
I know you've been working at Mozilla to work on Rust and I believe Rust is very good, but I also think Mozilla could have used other languages, there were a few that could led them to success, but they understood this are times where the "means of production" aren't the machines but engineers tools, and creating a programming language is the best way to control part of that world.
I haven't worked there in a year and a half.
Anyway I wasn't implying anything bad, just that you worked for years at Mozilla on Rust and it's like asking Anders Hejlsberg if C# enabled Microsoft to do things that have failed before with C++ or if TypeScript is better than vanilla JavScript.
which ones? If they tried 2 or 3 times with C++ and the Rust one succeeded, what other information would you need to have to convince you that Rust was the differentiating factor? It seems like you just don't want to admit that Rust was the key to their success in the project, even when you have someone who was there telling you that it was.
We aren't going to get research study levels of replication on large projects like this, so I don't know what standard you're looking for here.
The fact that Chrome is doing just fine without it?
> t seems like you just don't want to admit that Rust was the key to their success in the projec
It seems like you are trying a classic ad personam, I agree that Rust was one of the changing factor, I also wrote it, but just for Firefox, not in general.
Which is the original point of this sub-thread.
> We aren't going to get research study levels of replication on large projects like this
I don't think Firefox is the only large project out there. nor the largest.
middle age is a time of strict typing.
old age is all dangling references and memory safety issues.
> That doesn't make any sense at all
I'm sure it's not true—70% is way too high—but the real number isn't 0% as you might expect.
In particular, (most? all?) non-Rust languages that claim or appear to be memory-safe can have data races. In Go for example, those can be exploitable. [1]
And Rust is memory-safe...in safe code, with a bug-free compiler. Real programs have some unsafe code in their transitive dependencies and are compiled with the real, buggy compiler. [2] The percentage of security bugs in Rust code that are due to memory safety problems is more than 0%.
[1] https://blog.stalkr.net/2015/04/golang-data-races-to-break-m...
[2] https://github.com/rust-lang/rust/issues?q=is%3Aopen+is%3Ais...
They all seem to misbehave on a single nightly build, affect only a very narrow and unsupported target group, be related to a bug in a unsupported release of LLVM, or simply aren’t reproducible. I think the percentage of security bugs in Rust related to memory safety, based off that list, is _effectively_ 0%. I can’t find an issue in the list that seems like it would impact Rust programs that people write today or on targets that people deploy Rust code where memory safety matters.
Still, it was cool that it was fixed so quickly.
https://blog.rust-lang.org/2020/02/27/Rust-1.41.1.html#a-sou...
I don't think soundness/security flaws due to compiler bugs are common, but they qualitatively can happen. I think that might be getting forgotten when folks are puzzled about the idea of memory safety bugs in memory-safe languages.
If I make an array of 3 integers that are not supposed to be used yet but I accidentaly access one of them, then that is a memory-safety bug because I accessed unallocated memory. The functions that languages provide to "allocate" memory don't define what allocating memory means. They just provide a tool that you can use to keep track of what memory you are using.
What memory is allocated or not is relative and there are multiple levels of allocation. If we zoom out a bit we can even call C a memory-safe language because all memory accesses in C must be to allocated memory and the ones that are not will kill the process (allocated here as in allocated to the process by the operating system).
So while this may be unsurprising to those paying attention to Rust and memory safety, it's still relevant to a lot of software and a great confirmation of Rust's importance.
(I also don't really think Java is that much slower but substitute a more precise guess and my statement stands. And I think Java and other GCed languages really do use ~2X as much RAM which is also unacceptable.)
Using Java over C++ would have meant a factor of 2 up-front performance cost. But it would also have meant significantly less time debugging, easier testing, faster iteration, better automated refactoring support... all of which would have added up to being able to spend more development effort finding the kind of algorithmic improvements that give you those 1000x speedups. I'm not at all convinced that the end result would have been slower.
Over sufficiently long runs, throughput is easily within the factor of 2 you mentioned. But over short timescales, thousands of times slower is completely plausible. Some C++ CLI application might run in 5 ms where the equivalent Java program takes 5 seconds.
IIRC there was some article recently challenging the assumption that Java programs commonly reach the optimized steady state at all. I can't find it though.
[1] Except V8 of course on the user-supplied JavaScript. I'd call that quite different than optimizing all Chrome's own code.
In the context of something like Chrome that's already willing to implement a custom process model, I don't think that there are business requirements where that kind of large overhead is unavoidable. Running an individual unix-style JVM process can indeed perform very poorly on the default settings. But that's not the only way to build a web browser.
And the gap is getting _wider_ as Java is an increasingly bad fit for modern CPUs due to the heavy pointer chasing nature of it. It desperately needs value types to stay competitive in a performance battle.
Chromium started in 2008, at which point there were plenty of mature cross-platform high-performance memory-safe languages that would have been suitable (e.g. OCaml or Haskell).
A quick check of Wikipedia shows that they derived it from an existing small talk implementation.
Just the assembler files, for what it's worth:
https://web.archive.org/web/20100722105022/http://code.googl...
I will concede that a branch of GHC being utilized specifically for such a task could have been modified enough in the intervening years to enable a viable browser by 2020. But I really don’t know if it could be done with standard GHC.
Short answer: yes. Long answer: profiling and diagnosing performance issues is still kind of a black art, but for large codebases I've seen Haskell rewrites outperform C++ significantly. You need talented Haskell developers, but surely Google should be able to find those. Most of the tasks that I can think of in a web browser seem like things that Haskell is ideally suited to - parsing, data transformation, rule-based logic - what is it that makes you think it would be unsuitable?
From there, benchmarking could reveal if it is adequate for production use. Perhaps some parts could be rewritten in a lower level language.
For example, assume v8-ish semantics:
let mut cell = allocate(0.0);
let counter : &mut f64 = cell.as_mut_ref();
*counter += 1;
stuff();
*counter -= 1;
Here `stuff()` may trigger further allocations which result in moving `cell`, leaving `counter` dangling.Without any more specific knowledge, I bet the way a generational garbage collector moves data around really messes with Rust's view of the world...
There are also escape hatches, between unsafe leaking and the RefCell type, to provide “interior mutability”, i.e. mutable aliasing via an without the memory safety issues which come with this. Note I have only had to use the unsafe bits when passing memory up to a c stack.
I don’t believe there is a tracing GC implementation in the standard library, although I’ve heard plenty of discussion and you can find some simple implementations on crates.io. However, I would assume there would need to be some language changes to properly account for assumptions on how the Drop (free) functionality fires, not to mention significant backend work for tracing to work.
If you allocate something new, Rust pretends it is independent, assuming that malloc won't muck with the other objects it knows about.
What if we relax that assumption? Allocations may perturb other allocations - that's how moving GC heaps work! Can that be modeled usefully in Rust?
Rust has the capability to do things like execute load barriers on access, so it could still work when objects have been moved around.
&mut T means that there is only one pointer (the &mut T) that will be used to access some memory. You can have as many other pointers to that memory as you want, as long as you don't invalidate the first one.
This code is safe:
let mut foo = 13_i32;
let x: &mut i32 = &mut foo; // First &mut
{
let y: &mut i32 = unsafe { std::mem::transmute(x as *mut i32) }; // Second &mut T (copy)
*y = 42; // Write through second
}
dbg!(*x); // Read through first
You can run this under miri here: https://play.rust-lang.org/?version=stable&mode=release&edit...Getting shared mutability as a user of GC is basically the same as getting shared mutability in any other scenario in Rust. You use `UnsafeCell<T>` or a safe wrapper around it, like `Mutex` or `Cell`. You can get a feel for this even in idiomatic Rust using `Rc` instead of a tracing GC.
When you start looking at other methods of GC things get trickier. Even with `UnsafeCell<T>` the GC can't, in general, know whether something is borrowed (mutably or immutably) by the mutator, so the problem to be solved is finding ways to enforce that there are no references into the heap at all, when collection needs to happen.
But when you think about it, this is actually not so different from the usual problem of GC safepoints. The problem is the same- enforce that there are no un-rooted pointers into the heap, when collection needs to happen.
And it turns out there are some tricks you can do with lifetimes to get borrowck to enforce safepoints for you, and it could even be made pretty ergonomic in the future.
Some early discussion of the problem space:
* http://blog.pnkfx.org/blog/categories/gc/
* https://manishearth.github.io/blog/2015/09/01/designing-a-gc...
* https://manishearth.github.io/blog/2016/08/18/gc-support-in-...
A design from Servo, and a paper to go with it: https://github.com/asajeffrey/josephine
A series of blogposts describing a more recent attempt: https://boats.gitlab.io/blog/post/shifgrethor-i/
The most recent attempt I know of: https://github.com/kyren/gc-arena/
I think gc-arena is fairly similar to josephine- both treat the GC heap itself as a container-like object, requiring `&` or `&mut` access to dereference a GC pointer. Gc-arena also suggests that we could use generator yield points as GC safepoints- they have exactly the required lifetime properties. So you might (to borrow Python syntax) `yield from fn_that_allocates_gc_memory()`, with the compiler ensuring you don't hold any references into the heap across that call.
Both gc-arena and josephine appear to take the "unidirectional" approach, where there is no recursive re-entrance into the GC heap. For example gc-arena will not collect while the heap is being mutated, which is Rust-friendly but also impractical.
GC'd languages typically require reentrancy into managed code (managed -> native -> managed), and this has historically been the source of type confusion and other security vulnerabilities.
This is where the generator stuff I mentioned comes in- I think the approach is much more practical than it seems.
At the end of the day, any GC language needs to ensure there are no live un-rooted GC pointers at points where collection takes place. So any realistic Rust/GC integration will need to solve the same problem, if it is to retain memory safety.
And the best way (so far) to ensure all references into the heap have disappeared is to make sure they are all derived from a function parameter with an unconstrained lifetime, because then the mutator knows nothing except that it might go away when it returns. (This is "generative lifetimes" from the links.)
The real trick then is to make it look like the mutator is returning, without actually forcing it to exit. A yield point in a generator can do this, even without ever actually yielding. So you can manually insert collection points in the middle of a Rust mutator, and you can have a "managed -> native -> managed" stack where the native frame thinks its managed callee can yield across it, and borrowck will ensure anything derived from that reference parameter is gone.
Of course this is a long way off from Rust today, where generators aren't available on stable (though perhaps you could hack something together with async/await), but does look like it should eventually be possible to have reentrant heap access and memory safety across Rust and a GC language.
What do you think Chromium developers have been doing for the last 10 years, sitting on their hands? My understanding is that the main reason why Google OSS-Fuzz exists is Chromium. The problem with this approach is that it evidently (see OP) doesn't work to find all the bugs before release.
For example, initially Chromium had no support from isolating iframes from the main document. Everything was in Blink. But isolating iframes required extremely complex code in the main process to properly forward all DOM events between process. Then that code had to be made asynchronous so web pages could not block main UI.
Then OS interfaces to GPU and it’s architecture rules out having GPU accelerated graphics and decoders in a per-site process. So that must be put in single a GPU process. Surely it is heavy sandboxed. But, as with the recent Networking process, it is still a shared component with very complex C++ code and fat platform libraries in C/Objective-C for interfacing with GPU.
So at the end the original idea of writing everything in C++ and using sandboxed processes to make memory-safety bugs harmless just does not work. One needs a memory-safe language.
The article mentioned various library-based approaches. But those, as having such foreign C++ usage, are becoming essentially domain-specific-languages that one needs to learn. That leads to boilerplating as the host language is not well-suited for that.
It doesn't work in Firefox with its small block of Rust code, doesn't work in Chrome nor Safari. Because it's the dumbest fucking idea ever.
But sure, at this point we don't have anything to lose by rewriting it in Rust/Swift/whatever.
[0]: https://drewdevault.com/2016/11/24/Electron-considered-harmf...
Web apps are anyway nonsense both from a privacy (zero privacy) and security PoV (single high-value target, foreign jurisdiction, crappy browsers, byzantine tools, hopelessly overcomplicated architectures etc).
Case in point: Firefox. They've been rewriting parts of it in Rust for the last 6 years or so and they've only managed a little over 10%.
Just because you're a good driver doesn't mean you should forgo a seat belt!
For instance, perl has this construct, to make your programs better:
a program foo.pl:
$c = 1;
runs fine. Now turn on strict: use strict;
$c = 1;
and $ perl ./foo.pl
Global symbol "$c" requires explicit package name at ./foo.pl line 2.
Execution of ./foo.pl aborted due to compilation errors.
you must declare $c: my $c = 1;
the idea would be... write your program and you can opt-in to various constructs.NOTE: my first example didnt' work.
use strict;
$a = 1;
won't throw an error. Turns out the variables "$a" and "$b" are exempt from "use strict;" ha. wow.So make this function functional with no side effects.
This function doesn't (overtly) use pointers.
etc...
At various jobs, we've done things like this with warnings.
Get a source file to compile warning free, then turn on -Werror (this is done via compile flags, a different mechanism)
You can't. C is a fundamentally defective language because it doesn't have the "memory slice"/"array" construct.
Of course you can work around it, but if you eliminate pointers then you can't do much with C. Heck you can't even print "Hello World" with it (well, ok, technically you can but not in the usual way).
But still, I believe you can't do much with C if you "shut down" pointers completely. Because strings are pointers in the end.
For example it had a carveout so that "$a" and "$b" didn't have to be declared (I think they are known variable names for sort?)
Why not allow:
printf("Hello world\n");
but turn off: acpi_ut_repair_name(&(*converted_name)[j]);
or if (atomic_read(&(*per_cpu_ptr(sdd->sds, cpu))->ref))That doesn't mean it will operate without a GC mechanism per se, but it can operate with a "zero" GC that never collects garbage. I know this isn't what you meant, but thought it was an interesting point.
More generally I don't think you can really "patch" a language to make it safe. It's going to leak unsafety all over the place. Consider a simple `printf("%s", 12);` which is obviously broken but doesn't really trigger any kind of memory unsafety at the call site. You might argue that "%s" is a pointer in disguise, but if you ban C strings and arrays you won't have a lot left to work with (besides in this case the string pointer is perfectly valid, it's the format string that's incorrect).
Of course, it's slightly dated, but a lot of the points feel truly 'eternal'. All of the SoK papers from Oakland are good, but this one is great.
"//base is already getting into shape for spatial memory safety"
https://docs.google.com/presentation/d/15Zwb53JcncHfEwHpnG_P...
WTF (Web Template Framework) is the similar library from WebKit, that also exists in Blink, though the two are quite different these days.
Is this the result of 'open use of pointers' that inevitably causes problems here and there?
Or is this problem due to the fact is just impossible to use even hard idioms to avoid such problems? (Or hard to enforce them?)
On some small-to-midsized projects I've seen narrow and rigorously enforced rules around pointers that seemed to work well. So what's up?
Many people underestimate what is possible using C++ these days. Complex for sure, but also very powerful.
And that's the problem. There are very few true C++ experts, and most people will run into these issues as a matter of course.
This just feels like a reduction to the usual: if you are perfect, you will write perfect code, and never have memory-safety bugs. Thanks, but I'm not perfect, and I'd rather write in a language with a compiler that rejects programs with memory-safety bugs.
And yes, that does mean sometimes it'll reject some programs that are perfectly ok, but the Rust borrow checker is getting better all the time, and sometimes you just have to accept being hamstrung a little for the greater good.
Some of these idioms are not zero-cost though. My understanding is that to prevent use-after-free you basically can't use bare pointers or references, at all, ever. You need shared ownership everywhere (std::shared_ptr, not 0 cost), a garbage collector (In c++ that's probably another smart pointer type, not 0 cost), or additional metadata like the lifetimes fed into Rust's borrow checker.
Based on my reading here on HN, I think Chrome has a reputation of using modern C++ features extensively to try to improve its memory safety, but it's really hard to do in C++.
Edit: I realized this answer sounds like "yes idiom will fix it". To directly answer the question of "can you use idioms to fix this", no you can't. Evidence shows we screw it up. It's broken and you can't bolt on fixes to make it work.
However, while lifetime management with smart pointers is certainly important, it doesn't help with non-owning pointers between objects.
"In theory, yes - in practice, no."
You can use static analysis, vendor-specific annotations, runtime tools like clang's various sanitizers or valgrind, thorough code audits, use smart pointers instead of raw, have clear ownership semantics, follow MIRSA C guidelines, NASA guidelines, etc etc etc... but eventually your project will grow to the point where you'll botch an edge case, and the number of botched edge cases all your coworkers manage to add up will turn into a statistic.
Even if me and all my coworkers were all godlings incapable of making mistakes, we'd still end up debugging plenty of memory safety issues - in second/third party code, sometimes without source code, and writing bug reports / workarounds as a result.
"Today we were unlucky, but remember we only have to be lucky once. You will have to be lucky always."
Of course, one could make all the back references weak, but that turns out to also significantly hinder understanding of the code. It ends up being really easy to lose all assertions about when any given object that's weakly reference is actually supposed to be live, and you end up with code that's littered with checks everywhere "just in case".
The problem is, we haven't had a real application programming language since the demise of Pascal / Delphi. And Java / C# have been able to pick-up some of the slack, but not nearly enough as there are still tons of software being written in unsafe languages.
That is just like programming in D with extra steps.
And then replace it with... a language that is exactly as performant, is available on every known platform and has replacements for all the existing libraries (including all the ones written in C!).
We got stuck with the worst of all possible worlds--long compile times, long debugging times, long bugtails, horrible security, rickety and hard to refactor, ugly code. Oh, but it's fast. Fast and broken. Wonderful. Wonderful. Grandma is so much happier with her 15% faster something or other that spends 95% of its time idle waiting for user input. Oh, but that 15% of 5%...man that 0.8% is keeping me up at night! We forgot to zoom out to see that people need and depend on reliable software and utterly failed at that. It's as if we removed seatbelts, airbags, windshields, and mounted knives in random places in cars because we thought everyone wanted to go fast. No other disciple in all of engineering is so wrong with their priorities. God, society should be hopping mad at us for being so hostile to them.
[1] Obviously, performance differences between "fast" languages and "safe" languages are so variable that we might as well be talking about comparing Fords to Ferraris--without specifying whether we are talking racecars (Ford makes some fast ones)! But 15% is a number we see often in the JVM world. After years in this field, I think the fastest safe language is not more than 15% off the fastest unsafe language, across the wide spectrum. Sometimes, you can get absolutely the same performance for the innermost hot loop out of a safe language vs unsafe language in some situations, you can get 2x slower. But people get terrified they'll never figure out why, that there is some kind of hidden, unknowable "language cost" they'll never get offer. Which is of course hogwash. Every program can be tuned and improved. Offer tools to find and remove bottlenecks and stop being sloppy with allocating memory.
It's not a question of simply retiring C++.. it has to be replaced with something. Even safe languages frequently rely on libraries implemented in C or C++ as part of their runtime.
Even if you were willing to pay a 100% performance penalty (and there are plenty of places where I'd be fine with that), it's still a massive undertaking.
If C++ is only kept around because of a few libraries needed to implement the runtime of safe languages, then I would consider that some kind of victory.
Safer languages are better tools because they make obvious when you're stepping out of the safe zone, which is not the case with C++, even modern releases.
I see this claim all the time, but can you give some examples of large C++ projects that don't constantly struggle with memory safety issues? (And are looking for them, of course.)
It only works when everyone plays balls and doesn't do C style coding, ever.
Usually that can work in small teams with security minded individuals but it isn't a given.
And then there is the little fact that on most surveys, the amount of answers referring any kind of static analysis tooling are usually around 50%.
What C++ has definitely going for it, is having the type system tools to write much safer code than plain old C.
However its copy-paste compatibility with C is also what hinders any attempt to force people to actually only use those better features.
The only way to fix this is having systems programming languages being adopted that aren't C at copy-paste level.
Other would be some kind of Safe C, but both WG14 and C community in general have voted against such improvements.
If that’s the problem, and solving that would solve all of C++s memory-issues, why have no-one made a compiler option to simply make that code illegal? A -EUNSAFE or whatever?
C++ was also born at Bell Labs, and due to that, all major C compiler vendors quickly started shipping C++ on their boxes as well.
If you take away copy-paste compatibility you might be better off doing D, C# or whatever safe variant already exists.
Which is what many of us have done, to move to type safe languages, and only use C and C++ at the boundaries, in small pockets of unsafe code.
In fact if you look at mobile OSes, that is the reality for app developers, C and C++ are no longer the full stack languages they were 20 years ago, rather used for the kernel, drivers, compositor and shading languages, but everything else happens in safer languages.
And the SDKs only allow you to write libraries, not full applications.
Naturally are clever developers that subvert the workflow and transform the libraries into the actual application.
One problem is that C++ can directly include C headers of the operating system, while languages that aren't copy-paste compatible with C have to create some kind of wrapper library for them... this is a major reason for the success of C++, but it also makes improvements of this kind far harder to deploy in practice.
So if anyone is claiming you could write real-world safe C++ if you wanted to, they are making a false claim then?
Once you have that to build on top of, such a compiler flag could make sense... if it were possible in C++, which I'm not sure about.
See for example this criticism of one such effort: https://robert.ocallahan.org/2016/06/safe-c-subset-is-vapour...
It’s obviously still too early to declare a winner, but to me this sounds like a turtle slowly but surely overtaking a rabbit.
This is never acknowledged by the C++ people: Idiomatic C++ is not suitable for formal proofs, if you don't believe me, ask Xavier Leroy.
https://news.ycombinator.com/item?id=23290030
I would argue that the closer you stay to C while using the good features of C++, the safer the code is.
And obviously well written and debugged C code is nearly always more robust than C++ code. But for ideological reasons most people here are unable to acknowledge that, perhaps because they cannot do it.
- implicit conversions
- decays from enums to integers
- implicit conversiosn from integers to unexisting enums
- no bounds checking
- implicit conversions between pointers and arrays
- no proper way to ensure a given array length is valid as part of a function parameter
- null terminated strings, that occasionally aren't terminated
- abusing null terminated strings with clever algorithms, e.g. strtok()
- the preprocessor
- const that isn't really const
- variable arguments that require getting the macro type arguments
- UB explored to the last possibility of code optimization
- no safe way to deal with output parameters
- typedef don't introduce strong typing
All of that came from C, not C++.
Manageable in C. Integer conversions are only a tiny fraction of all the other implicitness in C++.
> decays from enums to integers
Compiler warns. Recent real compilers like gcc even have exhaustiveness checks like OCaml.
> implicit conversions from integers to unexisting enums.
Compiler warns.
- no bounds checking
Reason about that and implement your own scheme. Or prove.
> implicit conversions between pointers and arrays
Have not seen a single bug due to that in more than 1000000 lines of C.
> null terminated strings, that occasionally aren't terminated
Have not seen a bug due to that, this is the canonical example of an overblown hypothetical threat.
> abusing null terminated strings with clever algorithms, e.g. strtok()
Prove the algorithm or don't use it. Hint: As far as proofs are concerned, NUL terminated strings are like Lisp lists terminated with NIL, hence a well-founded data structure that is easily amenable to proofs (unlike C++ constructs).
- the preprocessor
Rarely introduces anything and is still required for C++, especially in sane test suites.
I don't think all that came from C, things like typedef being an alias rather than a separate type seem much older.
You are again just throwing dirt at C, mocking all people who write actually robust buzzword free software.
You are ignoring that C code is much easier for formal proofs that C++ code (the kind that you advocate).
Which happen to be written in C with several layers of code review and static analysis.
So by your reasoning those 32 years have not happened, in spite of being so easy to prevent exploits in C code.
Yeah, right.
All while their own industrial strength C++-OS (according to you) never has any exploits.
The fact that you are singling out Linux shows that you are only interested in throwing dirt.
I wonder why.
Here is a little tip for you, Microsoft has been acknowledging security issues with C and C++ since the XP SP2 days.
Which is why Windows happens to have plenty of mitigations that only recently FOSS UNIX clones are catching up to.
Yet they have come public that hasn't been enough, hence the migration effort away from C, enforcing programming guidelines with C++ and coming up with plans to migrate to safer systems programming languages.
Guess which OS vendor is now having first party support for writing GUIs in Rust?
But I can also rephrase what Oracle, Apple and Google have stated in the same vein regarding OS security.
Or maybe you prefer the statements of an UNIX hero instead?
I am always baffled by the people that supposedly write modern C++ professionally and constantly have memory safety issues. Most serious projects won’t hire you if you aren’t capable of writing memory safe code in your sleep, it is a basic skill.
The reality is that there isn’t much opportunity for memory safety issues to occur anyway, the type system and scheduler do most of the heavy lifting. Similarly, concurrency safety isn’t much of an issue because threads barely interact. Most high-performance server software looks this way these days.
Bugs tend to be of the boring logic variety that can happen in any programming language.
> Similarly, concurrency safety isn’t much of an issue because threads barely interact.
I think the domain you're working in isn't as susceptible to memory safety issues, but that doesn't mean they're not present. It also sounds like they domain you're working in is trivially parallelized if threads "barely interact", which limits your exposure to those memory and data race issues.
If you gave an adversary access to the API of your kernel however, how long do you think before they found a use-after-free, double-free, stack or heap overflow, etc? Days, weeks, or hours?
If you haven't run a fuzzer yet, I wouldn't be so confident.
No one can know if it is bug free, that is impractical. But it also isn’t like this is a weekend hobby project either. Most of the bugs that get out are in unimportant peripheral code and integrations.
I'm quite curious what tools you're using to formally verify your C++ code if you are. My understanding is that in general you can't, which is why msan/asan exist, to get a first approximation of verifying things that can't be formally verified for most C++ code.
(In general, I'm dubious of these claims that "Most serious projects won’t hire you if you aren’t capable of writing memory safe code in your sleep", because I think if you asked the majority of the members of the C++ committee if they could do that, they'd say no).
In many of these systems it is standard practice to generate arithmetically limited types pervasively. This is almost transparent in C++17. While it is possible to verify much of this at compile-time in theory, it almost never is because it isn't worth the effort (C++20 may start to change this) and testing at runtime has proven to be nearly as good. People underestimate what is possible with the C++ type infrastructure in this regard.
This type of software design was originally done because it allows for exceptional performance but has become popular for safety reasons. It uniquely allows you to make guarantees about runtime behavior under diverse adversarial workloads that would otherwise be difficult to make.
Bugs in practice tend to occur at the interface with third-party code, which requires dropping out of any internal type system, or in the form of performance anomalies due to unexpected hardware behaviors interacting with the scheduler design. Logic bugs in the core bits tend to be found in testing.
Presuming that this style works equally well for all software seems presumptuous and perhaps naive, does it not?
Sadly as codebases get more complex, that vision gets further and further from reality.
Sadly I don't think the tooling exists to enforce something like that automatically, because C++ is so hard to parse (a separate issue to memory safety). But when/if metaclasses finally land, it should be possible to write a "fixed" class/struct keyword that default initialises all fields.
(Oh, some background before I share that monstrosity with you: the project is a virtual machine I designed for a class I taught recently, and I experimented in implementing it in modern C++ and had a bit too much fun trying to see how much of it I could encode in the type system, which is why it will probably take a really long time to compile. Here is some code that actually uses that header, for reference: https://github.com/regular-vm/emulator, and I should note that I have internally changed some of the architecture slightly do accommodate an assembler that I have put aside for now.)
struct Instruction { explicit Instruction(int encoding) { } };
using T = Instruction &&;
int encoding = 0;
auto &&result = T(encoding);
There are a few things that went wrong here, but the final red flag to notice is that you should never call a constructor directly with 1 argument directly, because it's just a different syntax for a C-style cast, which we know is dangerous due to its bypassing of safety checks—and this is true regardless of whether we're dealing with C++ constructs (like classes and constructors and such), which I think might be what you're realizing now. (This is poor C++ design, but it's old and people know to avoid them syntactically just like C-style casts. It might be nice to have a warning for it too.) Rather, when you're passing a single argument, you want to write one of these syntaxes: T result(encoding); // option 1
auto &&result = static_cast<T>(encoding); // option 2
With these, you receive an error, e.g.: error: invalid static_cast from type 'int' to type 'T' {aka 'Instruction&&'}
I believe this is because the code requires 2 conversions to occur at once, which is an error because (surprise!) it's a generally unsafe thing to do: (a) conversion of int to Instruction, and (b) conversion of Instruction to Instruction&&.Now you bypassed this by using the uniform initialization syntax (i.e. braces). I'm going to go out on a limb here and say that was another mistake, even though it "solved" your problem here: despite the widespread use, brace initializers are, in my experience, not a good thing, and it's unfortunate that people embraced them (ha) with open arms, and similarly goes with emplace_back() and some other things which I'll address below. The syntactic convenience they provide is just too minor compared to the issues they introduce or obscure. And in this case, they indeed actually introduce a new issue if you use them like 'T result{encoding}': they allow 2 casts to occur at once. I find that incredibly dangerous, and I think it should be at least a warning if not an outright error like before. It seems like a C++ design flaw to me, but in any case, maybe someone should get compiler writers to add a warning for this.
Anyway, let's move on. If you're reading this, you're probably noticing that you ended up with Instruction&& in the first place—that's probably not what you wanted, or at least not what you should've wanted.
And that's where we get to the heart of the issue: your real problem is that you used decltype. If I saw that during code review, I would force you to change it—and using it on declval is just adding more fuel to the ember.
The reality—which unfortunately you do not see people acknowledging—is that decltype, auto, uniform initialization syntax, emplace_back, etc. are all dangerous, and harder to reason about than they look. I think it's unfortunate that the C++ committee encouraged people to use them so much, and I think they're overused to an insane degree. People who love them for their nice syntax don't go out of their way to figure out their pitfalls, but I almost never use any of them unless I absolutely need to. Most problems that they solve (one notable exception being 'auto' with lambdas) were quite elegantly solved in C++03 using typedefs. The only caveat was that you had to give up on the idea of minimizing keystrokes. It's quite a realization when you realize that solves so many of your problems.
Anyway, I'm not trying to blame these on you. Obviously these are blamable on C++, and we could (and should) have more warnings for them. Rather, my main message is that you can avoid these problems (even if they're other people's faults) syntactically—and locally—if you don't try to embrace the absolute "latest and greatest" in C++. IMHO you should only use the newer features if they solve an actual semantic problem for you, not merely because they minimize your typing.
In fact, I think is a huge mistake people make with software in general, and here, C++ in particular. They feel if something is old then it must be bad and you have to do everything in a new way. But if you stick with what works and start caring less about people looking down on you for using "old" syntax just because it's old, you'll find a lot of the old C++03 patterns (typename Pair::first_type, etc.) are actually robust to the problems that the newer ones introduce. People just don't realize this because they're more verbose than they would like, and we're in an era where doing things old style looks bad for no good reason.
Oh also, one last thing: aside from avoiding decltype and using the equivalent of a C-style cast, one more thing that helps you avoid this is to avoid overusing templates. They also obscure what's going on, like here. Not to mention the slow compilation speed and lack of independent compilability. Those are also overused (and I see their appeal) but they're often unnecessary and make code statically difficult to reason about.
Not only is it possible, I think it is very likely, although your explanation is something I can follow along with and very much appreciated.
> There are a few things that went wrong here, but the final red flag to notice is that you should never call a constructor directly with 1 argument directly, because it's just a different syntax for a C-style cast, which we know is dangerous due to its bypassing of safety checks—and this is true regardless of whether we're dealing with C++ constructs (like classes and constructors and such), which I think might be what you're realizing now.
Huh, interesting, I actually did not realize this. Is there a way to do this safely without creating an extra lvalue? I take it that there is no “extra explicit” keyword I can add to prevent this kind of accidental call, is there?
> I'm going to go out on a limb here and say that was another mistake, even though it "solved" your problem here: despite the widespread use, brace initializers are, in my experience, not a good thing
Yeah, I am not really a fan of them either :( Even I know of a bunch of caveats about them and C++ initialization is an extremely complicated topic…
> If you're reading this, you're probably noticing that you ended up with Instruction&& in the first place—that's probably not what you wanted, or at least not what you should've wanted.
No, but as you observed that it “works out” at some point in the pipeline so I obviously did not care to really figure out if this was what I wanted or not.
> And that's where we get to the heart of the issue: your real problem is that you used decltype. If I saw that during code review, I would force you to change it—and using it on declval is just adding more fuel to the ember.
Somewhat strangely, C++ seems like the only language where I would even consider to use such a construct. I think every other language just erases their types or simplifies them so you can be comfortable writing something like “Iterator i = collection.start” or “int size = collection.count” whereas in C++ you have some generic distance_type and it feels dirty to just work with a size_t or whatever you know the thing to be.
> IMHO you should only use the newer features if they solve an actual semantic problem for you, not merely because they minimize your typing.
A good point, but I would like to just mention that this was clearly an experiment in trying out the “latest and greatest” ;)
> Not to mention the slow compilation speed and lack of independent compilability.
Wait, you’re telling me my 100 line program shouldn’t take a dozen seconds to compile?!
The static_cast<T>(arg) syntax I used does exactly this! It's what you should use pretty much everywhere instead of T(arg). If it's too much typing, yeah unfortunately it is, though life is a lot easier if you can e.g. bind 'sc' to expand to it in your editor.
> No, but as you observed that it “works out” at some point in the pipeline so I obviously did not care to really figure out if this was what I wanted or not.
Yeah... sadly C++ is just about the 2nd-to-last last language you should deal with like that. The last probably being C. :-) Pro tip that might make it easier to avoid this: use typedefs very liberally. They help you avoid auto/decltype/etc. and are quite robust. (At least if your reviewers let you. If they don't, they probably haven't learned it the hard way yet.)
> Somewhat strangely, C++ seems like the only language where I would even consider to use such a construct. I think every other language just erases their types or simplifies them so you can be comfortable writing something like “Iterator i = collection.start” or “int size = collection.count” whereas in C++ you have some generic distance_type and it feels dirty to just work with a size_t or whatever you know the thing to be.
Those languages break too actually. Go Google "binary search bug" (with quotes). For example in C# there's Length and LongLength, which is dirty. When what they really need is just a native int. Another C++ tip: almost every 'int' or 'unsigned int' you ever deal with should be size_t or ptrdiff_t, because at some point or another it's probably an array index. It's very rare for that not to be the case; the only case I can think of off the top of my head is a logarithm (i.e. the shift amount in a bit-shift expression) or a timestamp (long long). Unless you're writing a generic STL-like container or allocator type (in which case, best of luck...), you won't need to care about difference_type or size_type.
> A good point, but I would like to just mention that this was clearly an experiment in trying out the “latest and greatest” ;)
Yeah ;) just keep it confined to experiments!
https://github.com/Microsoft/GSL
Stroustrup talk from 2016: https://www.youtube.com/watch?v=JtMPGwA3MzQ
It's something I don't see Rust proponents address. It's easy to build a straw man argument of Rust vs. old-school C, but it's a more natural path to go from C to modern C++ than to go from C to Rust. You get to keep your compiler, build system, tools, libraries and indeed existing code.
See this for an example: https://news.ycombinator.com/item?id=21681395
https://people.gnome.org/~federico/blog/exposing-c-and-rust-...
And yet we see the same memory-safety bugs come up time and time again in supposedly-modern C++ programs. So either this modern way is not enough to avoid these classes of bugs, or people very quickly fall to the temptation to use unsafe constructs due to performance or just because it's easier to write.
We're humans. If there's an easier way to do something, even if it's less safe, we'll invariably do it sometimes. I like that Rust makes it harder to do so, and makes you explicitly say that you want to do something unsafe, which I imagine deters a lot of people from going down those paths. And when someone writes a memory-safety bug in unsafe Rust, they get much more egg on their face than if they were to write the same bug in C++.
The memory management issues in my C++ programs (recently, real-time audio stuff) are where I explicitly decide not to use modern C++ / automated memory management, and write my own allocators that rely on malloc/free or some variant (aligned_alloc, etc.)
In Rust I suppose I would just use an unsafe block and have the exact same issues.
Note that good unit testing, assertions and sanitizers generally take care of the issue.
Unit tests, assertions, and sanitizers are nice, but they demonstrably don't work; Chrome and Firefox use all three. They have hundreds of thousands or millions of tests, assertions on every other line, and all kinds of compile-time sanitizers but they still have hundreds of memory safety problems a year.
We built a new project in all "modern C++". It is 100% shared_ptr, unique_ptr, std::string, RAII, etc. It initially targeted C++17 specifically to get all the "modern C++" goodness.
It segfaults. It segfaults all the time. It is entirely routine for us to run a new build through the CI process and find segfaults. We fuzz it and find dozens of segfaults. Segfaults because of uninitialized memory. Segfaults because dereferencing pointers. Segfaults because running off the end of arrays. Segfaults because trusting input from the outside world ("the length of this payload is X bytes").
This is where the "modern C++" people tell me we must be doing it wrong. But the reality is that "modern C++" isn't as safe or as foolproof as the advocates say it is. But don't take my word for it - this whole thread is about Google people coming to the same conclusion.
Meanwhile I can throw a new dev at Rust and watch them go from zero to works in a week or so, and their code doesn't segfault, doesn't panic, and actually does what it is supposed to do the first time. Code reviews are easy because I don't have to ponder the memory safety and correctness of every line of code. Reasoning about unwrap() is trivial. Finding unsafe {} is trivial (and removing it is also usually easy).
And then one day I found Rust, and all those problems went away. I can now write fearless code, and I don't have to endure the stench of rotting bodies anymore.
True story.
Any progress on a C++ to Rust converter? Not a "transpiler". Something with enough smarts to figure out when to use native Rust arrays, not "offsets" to imitate pointer arithmetic. I'm surprised that one of the big C++ users, like Google, doesn't have a group doing that.
The guy who maintains it said in the reddit thread[2] about this same topic that the Google people have been sending him good PRs, which is presumably related to integrating Rust into Chrome.
[1] https://crates.io/crates/cxx [2] https://reddit.com/r/rust/comments/gpdorw/the_chromium_proje...
There are some non-lint type things that would help in a safe mode. These all need type information; they're not just syntax.
- Can't keep a raw pointer. If you create one, it has to have local scope and cannot be copied to an outer scope. This is like a borrow in Rust, and limits the lifetime of the pointer. Most uses of raw pointers involve calling legacy code, and don't need much lifetime. Most trouble with pointers involves them outliving the thing to which they point.
- Can't read into or memcopy into any type that is not fully mapped. That is, all bit values have to be valid. Char OK, int OK, enum not OK, pointer not OK. This is better than prohibiting binary reads or memcopy, because programmers will not be tempted to bypass it.
- Casts into non fully mapped types are prohibited. If you need to convert something to a non fully mapped type, it requires a constructor, with checking.
That gives a sense of the general idea. Do enough analysis to see if something iffy is safe, and prohibit the cases which are not easy to show safe.
The “sort of” is what matters though:
- With std::unique_ptr, you trade use-after-free for use-after-move. Rarer, but still a threat.
- std::shared_ptr is costly which makes it unsuitable for some uses, and you can still have data rave with it if you don't use it correctly.
Every time there is a security issue, someone can see your naked pictures.
GC is one way to reduce the number of memory safety issues, but there are often tricky interactions between GC and non-GC code. Another issue with GC is it's often much harder to reason about object lifetimes.
However, it is possible, especially with resources of Google and other C++ shops to develop an advanced lint, which could be just a ripoff of compile time borrow checking from Rust (taking into account complexity of c++ memory model).
Flagging simple things like use after free would already be a big deal.
Nowadays it probably should be a plugin for clang or something similar instead of lint.
Like writing the program we already wrote....again.
We take this option off the table. Cannot rewrite. Definitely not in a safe language! Surely we have ceased to dream?
> Like writing the program we already wrote....again.
That's not named courage, that's named stupidity. Specially when the program you talk about is several millions lines of code and the results of years en engineering of entire teams.
> We take this option off the table. Cannot rewrite. Definitely not in a safe language! Surely we have ceased to dream?
There is people that dreams and there is the one that code.
Evolution are (almost) always preferrable to Revolution.
In this case, Evolution can mean:
- Rewrite progressively the security sensible part in Rust mixed with the legacy (aka the Mozilla way)
- Develop the tooling to make C++ safer/ or even better safe. That has already been done for a subset in the aeronautic industry
- Work with the comitte to make C++ itself safer, which Google is already doing.
Forever, never, ever, ever?
Switch, right now, is not a good idea, but NEVER switch is the cobol curse.
And not forget:
C/C++ cause BILLONS of damage and costs in this industry. Is WELL KNOW, for DECADES that is the wrong tool for the job.
Yet, the persisten myth that "lets not rewrite to something better" hold everything back.
Imagine how stupid will any of us look if a customer ask us for build a new version of the software (or port to the web, or to mobile or cloud) and we say "nope, sorry. Is not good. Not rewrite sir!"
Or if the software WE build have be caught, OFTEN with severe security and reliability bugs and we answer equally "we can't do anything better, is not a option".
Dumb, right?
Why rewrite is ok, normal and good for others, but not us?
And why the excuse that the MOST PROFITABLES COMPANIES IN THE WORLD (Apple, Google, MS, ...) can't do rewrite, is too costly.... and then the rest PAY the cost anyway, with not solution on sight...
For example, Android source code is also not a very good example of modern usage of C++ security features.
After 10 years, the NDK keeps having a C only API space, with all the memory corruption issues that it entails, even though the actual implementations are in C++ or Java (via JNI).
So their workaround, starting with Android 11, is requiring hardware memory tagging in all ARM devices, while having kernel fuzzing support for other platforms lacking such hardware capabilities.
Instead, they have been postponing it since NDK was introduced in Android 2.0 until after the introduction of native packages, which will only arrive with Android 11.
There are open tickets related to it.
You can reduce the attack surface, but never eliminate it.
I'd rather work on something nicer for less money.
For example, I hardly do any C++ nowadays, it is however what I get to use to integrate with OS native libraries and GPGPU shaders.
Rewriting everything is also not an option. Imagine rebuilding your house with modern materials, because old ones may have some shortcomings. Who will pay for that?
Wait, then what kind of devs would have enough expertise?
> in a project as large as Chrome the vast majority of programmers will not have the necessary expertise.
Then how do we staff such projects? Hire only people with a track record of never writing such bugs? It’s not possible.
So … you’re talking about Google not being able to pay enough or have a good enough reputation to hire developers who want to work on one of the highest-impact codebases in the world, not to mention contributing to the standards process and popular tools used to improve security for the entire community. Who realistically should look at that and say “no problem, we’ll do better!”?