Use-After-Freedom: MiraclePtr
security.googleblog.com
security.googleblog.com
> graph showing that 90% of high-severity bugs in the "Browser" part of the Chromium code are memory safety errors
I'm going to bite the bullet here and mention the elephant in the room: the vast majority of these errors are virtually absent from Rust codebases.
Logic errors happen in every language, but as the article shows, the majority of safety exploits come from memory safety errors, not logic errors, and memory-safe languages exist.
I realize that Chromium is a pretty big codebase and they're not going to just switch to a different language, but at some point I feel like we should be acknowledging that maintaining a large-scale C++ codebase in a domain where exploits are highly sought by intelligence agencies is extremely unsafe, and solutions like this are just another layer of band-aids.
(Though, to be clear, the amount of effort Google poured into defense-in-depth for Chromium and achieving an acceptable level of safety in practice is nothing short of amazing)
But if Google, Microsoft, and Apple can't find enough good programmers for their flagship most security-sensitive projects, that's still a terrible outlook for the C++ language. It's already hard to hire good programmers, and these flawless C++ programmers are harder to find than True Scotsmen!
Systemic solutions to recurring problems are good.
Nice excuse. But just wholly untrue. Most of us have not had any of these faults in years. Zero. I cannot even remember the last time I had one. Or even saw one.
It is just possible that it is particularly harder in the hackneyed dialect of C++ that Google coders are forced to use, but that is doubtful. The fault does seem to be particularly associated with Google, although old Mozilla code seems similarly afflicted.
It is probably connected with fooling with actual naked pointers, and transferring ownership according to more-or-less documented conventions, as opposed to never needing to document it because it is simply never done, and never needs to be done.
The skill level of Google's in-house coders, as for Microsoft's, has seen plenty of boosterism, but the numbers don't lie. Neither place is a habitual producer of code to brag about. Having worked there should not count in your favor in an interview.
It's not like Google is uniquely bad at this. The Linux Kernel also suffers from memory safety error, and Google's Project Zero keeps fidning them in other projects.
"Scale of your codebase" is a clever excuse, but it doesn't wash. The bigger your codebase, the less cowboy-ism you want. But Google seems to particularly foster it.
Insisting everybody else in the world has all your problems too can make your management failures seem excusable. But management failure can happen anywhere. It is evidently happening at Google. Software faults are a symptom. I don't know how they can fix the management problem that washed up as billions of lines of bad code.
I do know pretending won't help.
Google can be assumed to know the hole they are in. The mistake is thinking everybody else is in it too.
Which project of similar size doesn't have bugs? You seem to imply most don't, which is empirically untrue.
https://scholar.google.com/scholar?hl=en&as_sdt=0%2C15&q=bug...
And, if you really want to restrict to your goalpost moving subset, then go ahead and demonstrate that the above general trends are not trends in your subset. Because the literature on it disagrees with you.
So, to defend your claim "The mistake is thinking everybody else is in it too", care to show where "everybody else" is free from these bugs?
https://scholar.google.com/scholar?hl=en&as_sdt=0%2C15&q=use...
https://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=use+after+f...
This is the very first line from the very first post in the thread:
>> It’s hard, if not impossible, to avoid use-after-frees in a non-trivial codebase.
I just provided thousands of examples where it is true. To offset all that evidence can you provide, say, millions of examples where zero (your phrase) use-after-free bugs have been found?
>You reveal that you never had any intention of engaging honestly.
Honest engaging would be for you to provide evidence for a claim you made when asked. I provided extremely solid counter evidence from the result of many, many research teams.
So, care to honestly engage about your claim?
Bugs scale with the size of the codebase, for every entity on earth.
But zero times W is zero.
You could try to claim they are in all programs, but your selection-bias slip is showing. Programs they are not in (e.g. mine) are nowhere in your list. You can have no idea about the number of such programs. To pretend you do, as Google has done, is not honest.
Of course you could show that sampling is an invalid way to gather estimates of frequency, and rewrite all of statistics, but I suspect your high level C++ perfection is more time useful to you.
Me, on the other hand, trusts the aggregation of researchers over an internet commenter, even if they do have half of all comments in a topic.
Of course your mystery programs no one else can see provides you a way to claim such sampling is not representative. But since you made the claim, it is up to you to provide evidence.
I get that you have none except your anecdotal self-claims, which is certainly selection bias.
For example:
>Programs they are not in (e.g. mine) are nowhere in your list
I could point out you do mot have proof they are not in it; you just have not found one. Current tech on whole program correctness provers don't yet scale to codebases of this size, and absent that, you do not know your few programs do not have them, no matter how much you claim otherwise or try to code otherwise.
"To pretend as you do ... is not honest."
Also trying to implement certain algorithms and datastructures _efficiently_ often unfortunately calls for raw pointers. And writing the most used web browser certainly calls for trying to optimize things. If you think differently, please tell me how you would implement something as simple as a mutable tree, where each knows about it's parent.
The principle about trees, if you must use one -- they are very slow to walk -- is that you either use it, or mutate it, not both at the same time.
Generally, instead of pointers to nodes as native objects, it is much better to have all the nodes in a more array-like container, and represent edges with indices. Then none own any others, and you can visit all the nodes, if you must, without following any edges at all. (Graphs, likewise.)
You might keep all the nodes with one property in one vector, and the rest in a different one. If they have some unique identifier, that could be their index.
Doing whatever your freshman Data Structures class taught, forever, is no way to code for a multi-billion dollar company.
Now think about the fact that API provided by browser to JS is actually stateful. And you have to make a lot of things mutable. That is an external constraint that a lot of things in the browser world are formulated as a mutable OO-model.
Also think about the fact that storing everything in a vector means that you get a massive fragmentation issues if your use pattern (which is dictated by actual downloaded content) for some reason requires allocating N items and then freeing 0.9N.
Also think about the fact that whenever your array overflows and you need to resize it, you have to a) make a massive allocation from the OS, which might be quite janky because large allocations usually incur context switch + TLB misses for new pages b) you temporary use 2x memory.
Also note the fact that if you reference everything by index, you effecitevly have a raw pointer. The only win here is that it is guaranteed to point to the right kind of memory, but you may very well reference a newely allocated node in the same spot.
It is easy to preach "the right" approach until you meet real world.
If size varies much, a deque-like structure would work better than a naïve vector.
You can always do the dumb thing if you are in a hurry, but then you have a dumb thing.
Deque-like data-structure obviously has 2 memory references on each access instead of one with regulare pointers or with indices into a vectos. Didn't you want to have a more efficient datastructure?
You didn't address the fragmentation issue either.
Just to clarify things: I am not saying that using data oriented design is bad in any way. I am saying that proposal to use it specifically to prevent memory issues in self-referential mutable data-structures is nonsense.
Data structure choices are dictated by requirements. The type system doesn't demand pointers. When the type system accurately represents relationships, there is no need for users of a data structure to be careful to obey rules. Operations that compile at all are correct by construction. (Anything-oriented design is code smell.)
Any "fragmentation issue" is off-topic: if something is an "issue", you do something differently so it isn't anymore.
Of course. Unfortunately sticking everything into an array and pointing to things via integers instead of raw pointers doesn't provide you any help from typesystem. My whole point is that there is a huge class of datastructures and algorithms which can not be reperented in a type-safe way without automatic manual memory management. You either have to be accurate when writing your algorithms, or introduce runtime costs for validity checks.
You get pretty much 0 support from your compiler (for imperative language without GC) if you are not dealing with hierarchical datastructure without backreferences. It does not really matter how you reperesent it. No language can model that at type level.
You can either encapsulate it into some abstraction and prove (and/or hope) that it is implemented correctly. This is a common way to approach the problem of type-system no helping you (for example, pretty much all foundational Rust libraries are unsafe underneath to get around borrow checker; STL in C++ lives and breathes with raw pointers in the implementation of succint datastructures like maps, etc). If you can't encapsulate it fully, you try to be accurate OR use some helpers to deal with the complexity (shared and weak ptr, this raw_ptr that article is talking etc).
> Any "fragmentation issue" is off-topic: if something is an "issue", you do something differently so it isn't anymore.
You were selling pushing everything into an array as a solution to not using pointers. I tried to give an example why this might be not the solution that you want. Specifically in browsers.
> (Anything-oriented design is code smell.)
Data-oriented design is a name for what you were proposing, I just brought it up for easier reference of your point. I.e. Modelling your data as a set of tables with relations between them.
And, if you think that a type system can be no help with some family of data structures, it can only because you are unfamiliar with powerful type systems.
Rust coders rely on the compiler to protect them from themselves. C++ coders rely on the type system and libraries in the same way. The language is powerful enough to enable building libraries that can only be used correctly; or that it is much more work to use incorrectly, so no one is tempted.
I am. Would you please point me at some things to study?
The best I can imagine is languages like F* (https://www.fstar-lang.org/, and more specifically the Low* subset for dealing with memory) which make you prove everything. And thus are such a PITA to program in that it make sense to code only the most security-demanding things.
Dependant types are also not able to protect you from memory bugs in self-referential datastructures AFAIK.
I am open to reading about research in PL. But if you point me at some mainstream (or just production-ready) languages/runtimes it would be even better.
Thanks.
So reading up on what can now be done with C++ concepts should give you some ideas. For a more abstract treatment, read about Haskell's type classes.
Neither C++ now Rust can help you self-referential datastructures through types. In both you have either give up any help from compiler. Or pay some price for runtime checks (via shared_ptr/weak_ptr in C++ or their friends Rc/Weak in rust). Exactly like the ones described in the article.
Concepts, Rust traits, Haskell Type classes provide close to 0 help with memory management. These are a typesafe way to write generic code.
Things which do help with memory management at the language level (with 0 runtime costs) are Rust's borrow checker (and more generally(?) linear/affine types), F* proofs, and dependant types (you can prove index bounds with them to some extent). All of them are useful, but can't really model all interesting programs (i.e. you have to sidestep typesystem and just tell compiler "I know that I am right").
General graphs are, as I noted, better handled by decoupling ownership from connectivity, as is almost always quite easy.
Let me reiterate once again. Decopling of ownership from connectivity does not solve any correctness problems. Instead of "use-after-free" you get "use-of-object-in-some-random-state" which might be even worse, since no tooling like address sanitizer or Valgrind will help you.
> Runtime assistance coupled via the type system is fully as legitimate as GC. Your example with "up" pointers in a pointer-linked tree (or DAG) structure is very easily handled that way.
> Runtime assistance coupled via the type system is fully as legitimate as GC.
You do understand that this particular article is exactly about it? Runtime assistance comes at a cost, though. Actually very measurable if you for some reason want to get max performance. (atomic refcounting is expensive).
More generally, any memory management approach is fully legitimate, including fully manual a-la C. But rigorous refcounting in C++ and Rust comes at a very high cost, and to make it as useable as tracing GC you have to actually implement cycle detection which makes it even more costly. So people generally cheat, and for some sorts of performance-sensitive code, that makes a lot of sense.
Fully manual memory management has been demonstrated not to be legitimate, for reasons well covered already. It is just the only choice in weak languages.
...because people coding for multi-billion dollar companies usually need to deal with concurrency if not outright parallelism. So people working at such companies don't code like first-year CS students (avoiding the sexist term) under an outdated curriculum. The real way to fail in that kind of environment is to keep giving advice suited only to a now-purely-pedagogical model of computation as a single sequential stream.
Sexist term... "freshman"?
You reveal that you had no intention of engaging honestly.
Mozilla invented a language to get around it.
In my opinion it managed to improve the field of programming in a huge way.
That said as long as C++ code relies on C code or programmer invalidates memory safety invariant somewhere you are just as safe as C code. I.e. not much.
Rust components are in Firefox, in Linux kernel, and others.
Bolting on security after the fact, will never be as effective as building it from the ground up.
And relying on humans to just not make mistakes is a fool's errand.
> Most of us have not had any of these faults in years. Zero. I cannot even remember the last time I had one. Or even saw one.
Implying it's an issue of just not making mistakes/getting better people/etc.
So rather than relying on a strict borrow checker or a GC runtime, you rely that people won't make mistakes.
That I never have use-after-free or other memory errors is prima facie evidence that their absence is not a matter of "not making mistakes". We all make mistakes. Their entire non-occurrence must be attributed to other reasons.
The main effect of your aggressive propaganda efforts is to destroy your credibility, and that of everybody who apes them. Congratulations.
I find it interesting that people often like to project their own thoughts onto others. When that happens it causes a sharp contrast between my own and their own reflected opinion:
- First off. I never said that borrow checker is only way to ensure static correctness (assuming you don't consider GC static correctness), it's just the most well known, both to me and to people on HN at large. Other approach like Cyclone are mostly research languages.
- Second you never proved your code had no user-after-free. Just claimed it was.
- Third, propaganda? I just use it and consider more safety better than nothing. It's just a hobby for me. So I have no idea where that came up?
None of the existing security assessment agencies has your foresight, big business opportunity lies ahead.
Google's problem is the same as most big corporations': Everybody who works there knows Google doesn't give a rat's ass about them. They, very sensibly, don't give a rat's ass about Google. Code quality slips.
Microsoft's additional problem has been that fixing or preventing bugs was very early on recognized throughout as a waste of time and company money: in a monopoly market, a bug can hardly ever cause any reduction in revenue. Google's actual customers are likewise insulated from bugs that don't affect ad presentation or billing. Even those rarely affect Google's revenue in any detectable way.
So, Google's overriding concern is uptime, because when they are down or slow, revenue falls off in exact proportion.