Memory Safe Languages in Android 13
security.googleblog.com
security.googleblog.com
Defenders of C/C++ frequently note that memory safety bugs aren't a significant percentage of the total bug count, and argue that this means it's not worth the hassle of switching to a new language.
Google's data suggests that while this is true, almost all severe vulnerabilities are related to memory safety. Their switch to memory-safe languages has led to a dramatic decrease in critical-severity and remotely-exploitable vulnerabilities even as the total number of vulnerabilities has remained steady.
On top of that, memory safety vulnerabilities are disproportionately high severity: "Memory safety vulnerabilities disproportionately represent our most severe vulnerabilities. In 2022, despite only representing 36% of vulnerabilities in the security bulletin (NOTE: down to 36% from 65% because of moving from C++ to Rust and other memory safe languages), memory-safety vulnerabilities accounted for 86% of our critical severity security vulnerabilities, our highest rating, and 89% of our remotely exploitable vulnerabilities. Over the past few years, memory safety vulnerabilities have accounted for 78% of confirmed exploited “in-the-wild” vulnerabilities on Android devices."
Imagine if in any other field, a process or technology were developed that cuts the number of high-severity issues in half.
For example, a modification to the standard anesthesia protocols that demonstrably reduces anesthesia-related fatalities by 50% in clinical practice.
And now imagine, in reaction to this revolutionary development, thousands of anesthesiologists publicly said things like "what matters is not the technique but the skill of the physician", "good anesthesiologists don't make mistakes like that in the first place", "but this new technique takes 1%-3% longer than the previous one" or similar.
Utterly unthinkable, isn't it?
Yet in software engineering, this is exactly what has been happening every day for more than a decade.
The attitude "it's just computers, nothing truly important like medicine" might have been viable 40 years ago, but it certainly isn't anymore.
The earth is still spinning, despite so many things not working everywhere. That's not just in software. In every system, there's relatively few single points of failure. As you zoom out, failure points disappear and new ones appear.
Sure, still sometimes someone notices a critical security flaw that's been in there for a while, and that had found its way into large parts of the infrastructure already. (The last one I heard of was not a memory vulnerability).
https://qz.com/646467/how-one-programmer-broke-the-internet-...
If this breakage is an argument for anything, it is against depending on lots of code that you don't even know. The last Rust projects I tried to build all had on the order of 500 transitive dependencies, by the way.
Software is not terribly different from that. The cost of replacing foundational software is tremendous, much higher than just "adopting a new protocol".
Of course, new software should still take heed of this and try to improve.
This is the exact equivalent of physicians continuing to use unsafe medical procedures, and what's worse, many of those engineers defend their dangerous practices by claiming there is no real danger in the first place if the programmer is "smart enough".
I've tried a lot of languages in the past, and am currently not willing to dive into a whole new ecosystem, re-learn all the best practices, and unlearn what's worked very well for me with sometimes no good replacement. Best practice for Rust seems to be to avoid linked lists for example. I don't know what's the safe replacement for memset and memcpy to anonymous data structures (void pointers -- generic code) but I have a sense that it is more painful. The general recommendations seem to be to switch to different, more complicated datastructures with more failure modes, or to put boxes and Arcs around things.
I don't think this fits me - I like the feeling of understanding what I'm doing, and once in a while coming up with something that compiles and runs fast and robustly. If you can do that in Rust, good for you.
Another part is that it still seems easier to interface with existing ecosystems in C. I tried writing some Win32 Rust code once in an evening and I have to admit I failed. Maybe I picked the wrong bindings library or whatever. At this point in my life, I have little patience to spend my time like this.
I appreciate the work that is being done, and I feel it's not unlikely that at some point we all will switch. At this point though I feel I'm way more productive staying in my current habitat. And that's not only on me, but also that the ecosystem and developed practices likely are not quite ready for a complete switch.
Comparing the investment to simply washing hands and putting on gloves or following a checklist as a surgeon feels unfair to me.
The equivalent of a generic memcpy is probably something like a .clone() call on a generic type that implements Clone.
Probably not the Clone trait but the Copy trait.
If you type "memcpy" into the documentation search, rustdoc will point you to
* https://doc.rust-lang.org/stable/std/primitive.slice.html#me...
* https://doc.rust-lang.org/stable/std/intrinsics/fn.copy_nono... (Though it should really point to the reexport at https://doc.rust-lang.org/stable/std/ptr/fn.copy_nonoverlapp... )
The latter will also mention
Implementing linked list in any language that is not memory-safe is challenging because of safety issues. Rust just points it out.
I don't feel that I am less in control, just that all of that needs to be put in code and not just go "okay, I know this part don't need a lock coz I will never call it concurrently" and hope for best.
Even writing for embedded (as in no os, tens of kilobytes of RAM microcontrollers) haven't been too bad althought I haven't managed to convince borrow checker to borrow non-contigous block of bits from a register yet... althought that's what unsafe{} is for after all
> Comparing the investment to simply washing hands and putting on gloves or following a checklist as a surgeon feels unfair to me.
The closer one would be "read that 300 pages of how to do stuff safely and apply it". Once you get into good habits it's not a problem but investment is there
> New projects written in C are being started every day.
But it is mostly about existing software, even if its not about replacing existing software. I write C++ code every day. I hate it, and I'd rather not. But I use C++ libraries written by my teammates, and reference implementations in prior C++ work in our past projects, and have a set-up C++ toolchain, and my company even has all sorts of written C++ docs and style guides and linters and macros and ... and ...
Even if we wanted to stop, there's so much extra stuff to consider. I get your point, <<why start a new thing knowing there's a better way?>> and I would want to stop using c++ but most software isn't one-and-done like a surgery. It is a continuous commitment and ongoing operation, it is the tools used, and the knowledge learned, and the libraries built.
> This is the exact equivalent of physicians continuing to use unsafe medical procedures
Plenty of platforms don't support rust. Just because you've improved knee surgery, doesn't mean it works on an elbow... yet. And sometimes you still gotta perform surgery on elbows.
> if the programmer is "smart enough".
Can't defend this, but tbh I've never heard it.
> Can't defend this, but tbh I've never heard it.
Take a look at this comment: https://news.ycombinator.com/item?id=33824934
> Modern C++ has many memory safety features. If a company has learned that its people fail to use them, then bad for them.
It is a somewhat common attitude in this type of thread.
Except the safety features are often a lot easier to use than the original C isms that tend to cause the most issues. So it is less an issue of smart and more one of bad habits. C strings instead of std::string, plain arrays instead of std::vector, implicit ownership instead of smart pointers, ... .
The first Google Style Guide for C++ I ever came across in the wild espoused C with classes, had "standard" in scare quotes and banned most of boost for encouraging functional programming. I almost threw a fit when someone unironically tried to push that POS at work because "Google", it was entirely nonsensical especially given that we made heavy use of boost and math libraries with operator overloading.
Especially since Google has a ton of automated tools to perform tests and analysis on code, and enforces certain behavior before you can merge your code in. Something that is probably missing from a smaller organization that’s simply adopting their style guide. Also probably missing is googles alternative stdlib they use.
This is not about C++ developers using old school C approaches. The language is fundamentally dangerous. There were huge efforts within Chrome to get all memory managed by smart pointers even more powerful than those offered by the standard, and they still have UAFs all the time.
It’s hacker news. People say all sorts of inflammatory crap here. I discount anything said here, especially in response to an article about said topic.
If my c++ writing coworker shared that opinion with me, or someone at a conference, that’s be different.
> Modern C++ has many memory safety features. If a company has learned that its people fail to use them, then bad for them
I will admit I vaguely agree with this. My company has all sorts of tools that perform basic checks for memory safety and a style guide that is very opinionated. Just because we can’t switch easily doesn’t mean we can’t try to improve as an org.
I don’t avow the belief that any program is truly smart enough/too dumb to make bugs, but I do think the organization maintaining the code has a responsibility to improve. Especially a large organization. Whether that’s better code review processes, automated tooling, or even soft-banning the use of certain unsafe practices.
> Plenty of platforms don't support rust. Just because you've improved knee surgery, doesn't mean it works on an elbow... yet. And sometimes you still gotta perform surgery on elbows.
That's a non-argument, obviously if it isn't even applicable it's not a discussion. Also most of what people code on does support Rust
>> if the programmer is "smart enough".
>Can't defend this, but tbh I've never heard it.
Programmer is never smart enough. Decades of bugs showed that
And why do you think it's wrong, per se?
> This is the exact equivalent of physicians continuing to use unsafe medical procedures
You mean with different safety tradeoffs?
Because what you propose is exactly like forcing very expert physicians to switch to a procedure they are novice of and that it's not been battle tested like the old one, that proved to be very effective in most cases.
It's the same reason why patients prefer to be treated with established procedures and to undergo experimental treatments they need to sign a document that proves their informed consent.
The issue is not "bridge is suboptimally designed" or "bridge will need some extra maintenance because something started to break".
The memory safety issues are "some cars randomly explode when passing that bridge" or "when driver breaks 6 seconds after entering the bridge, every other driver dies.
> Software is not terribly different from that.
It is MASSIVELY different, especially anywhere near anything security-related. Most buildings don't have a group of people with hammers trying to find a weak point that will never happen in any actual use conditions and then hit that weakpoint in every similar building built in every place of the world.
Please don't make horribly useless comparisons like that
All the time! And the stuff we don't tear down, we retrofit.
No. Ignaz Semmelweis faced it in the 1800s for daring to suggest (what we know know as germs) made people sick and hand washing could drastically reduce medical complications. He was able to prove it too.
By the end was locked up in an asylum for his ‘crimes’.
Want more recent?
How many stories have you heard of instruments or gauze or whatever left in surgical patients? Of operating on the wrong part or person?
People are fallible. But checklists help a ton. We know that for sure. Why is aviation obsessed with following them? Because they work to increase safety.
Surgeons have resisted them. I don’t know the current state of it, but they were making that exact “good surgeons don’t need it” argument. I remember it being a plot point on an episode of a medical drama (ER? Or maybe Grey’s Anatomy).
There are probably tons of other examples in other fields.
Surgeons resisting checklists is new to me. Do you have a reference other than a TV show? My understanding until now was that checklists are extensively used in medicine.
> Despite all the evidence, Gawande admits that even he was skeptical that using a checklist in everyday practice would help to save the lives of his patients.
> "I didn't expect it," Gawande says with a chuckle. "It's massively improved the kind of results that I'm getting. When we implemented this checklist in eight other hospitals, I started using it because I didn't want to be a hypocrite. But hey, I'm at Harvard, did I need a checklist? No."
https://www.npr.org/2010/01/05/122226184/atul-gawandes-check...
Not sure if this supports the reluctance idea since it says 93% use checklists, but most surgeons don't think it improves safety.
> Of the 353 survey respondents, 93.6% use SSCs and 62.6% would want one used in their own child’s operation, but only 54.7% felt that checklists improve patient safety.
The part you quoted from your second link directly contradicts your claim that most surgeons don't think it improves safety (even if the 54.7% is among the 93.6% that use it that still gives 51% that think it improves safety).
> Yet in software engineering, this is exactly what has been happening every day for more than a decade.
Consider what happened to poor Semmelweis when he discovered in the middle of the 19th century that hand-washing improved medical outcomes: doctors of the time were too proud to accept his results and drove him out of the profession (and to his early grave) rather than change their practices.
All these people who insist on continuing to write new C programs have the same mentality as those doctors who refused to wash their hands.
A bit much, eh? I write programs in whatever I want as an hobby. Why does that make me a murderer? Everyone needs to chill about this.
Not really. Rust is the first major attempt to achieve c/c++ performance & capability while also being safe.
Prior to that nearly every memory safe language came with crippling tradeoffs, and that is why people rejected them. Especially as performance and efficiency took center stage again with the massive increase in battery powered devices & the general plateau of single core CPU performance over the last decade
This says more about those C++ defenders.
Now Google has put this to the test and has the data to prove it. We should not allow the worlds technology security to be held hostage by a group of people too lazy to adapt with the times.
Nitpick, this is not quite true. Memory safe languages are what you should be using in contexts where security and reliability are critical. This is generally the case but there are some contexts where other concerns are genuinely more important. Of course when this is necessary is often misrepresented but these cases do exist.
Raw memory access is something you normally need to create vulnerabilities, not to exploit them :)
It's an ergonomics thing, not a "can't" issue. There is a reason I called out Swift in particular—their "unsafe" APIs are so horrid to use that they make you regret doing unsafe things in the first place. Plus, throw in FFI and now you've got an even worse problem because not only are you forced to use the unsafe types, but often a lot of the critical APIs you need to interface with (in *OS exploitation, mach is the worst offender) have such funky types due to their generic nature that you have to go through half a dozen different conversions to get access to the underlying data.
https://github.com/hellman/libformatstr
You can do something like this, no need to work with raw memory.
Not everybody is writing security-critical code. For some things productivity and time-to-market is more important and security is not enough of a concern to justify dealing with a language with horrible compile times and a self-righteous, dogmatic community.
After memory safe languages become a bit more battle tested, C/++ needs to be regulated like asbestos.
There are some things that the language is good at, but at some other things (I think a lot actually) it isn't. At the very least, I wouldn't recommend it for writing a video decoder.
In nearly 10 years of professional Haskell I've never come across a practical situation where extensions that were mutually incompatible. Can you name any extensions that are incompatible in a way that actually matters in practice?
Some of these things would make working in the language pretty painful. I remember trying some library to toy around with a basic 400 line OpenGL program. It always needed 8 seconds to rebuild at the time. I don't recall and probably didn't understand why, but I suppose it has to do with some extra type or template hackery in the library that would just overcomplicate everything, probably even at the outset.
What remains is the feeling of, I can't do this thing yet in the type system, so take that extension. Oh wow. Now I can't do this other thing that I also need, the only fix is another extension (if it's available yet). I couldn't get around of this feeling of constantly having to hack around the language.
In my experience you need to have a very good overview and understanding of all the available tricks and extensions to be able to navigate your way around the language and not paint yourself into a corner. Maybe I have the wrong personality, the wrong motivations, or am just not smart enough. Obviously I'm not you, and am not Edward Kmett (who would resort to lots of GHC specific hacks and drop down to C++ as well).
This is a fair criticism. Compile times are slow.
Extension confusion is not a fair criticism. Extensions typically remove restrictions. They don't create incompatible languages.
Granted, it's a bit annoying to have to turn them all on, one by one. These days one should just enable GHC2021 and forget about language extensions. That saves one having to be Edward Kmett, or from bothering to think about language extensions at all.
If you write in C and your program is complex enough, you will spend a lot of time just chasing segfaults and concurrency bugs and getting it to work the first time you write it. If your systems programming is in userspace, that's sort of fine. But if you're in kernel or on bare metal, the cost of debugging goes up by an order of magnitude. There's no debuggers on bare metal, and nobody can tell you how your program crashed—the device just stops responding, that's it.
That's why safe languages make even more sense in these restricted environments. If you get twice fewer memory safety bugs while you're getting your program to work, that can reduce your development time like 5-6 times.
So now you need to make your interrupts talk to your Java objects? Is this any safer?
Is it easier to get a VM running in your kernel (probably no mean feat to do that in the first place) and you'll never get any concurrency bugs? And if you reduce memory bugs by half, those will be easier to debug?
And you won't be annoyed that you can't guarantee to be able to link objects in queues (because of allocation failure) and access them with generic code to copy data, link/unlink them, and so on? You're fine to pay for callbacks and interfaces everywhere, both in terms of runtime and performance as as maintenance headaches?
I'm asking incredulously, but seriously. Because frankly I've never looked at a project like MirageOS or whatever. But given real world evidence of what has survived, I don't see why you should assume I'm joking.
That you can't debug a "bare metal" kernel isn't quite true, either. But sure, the more complex a system becomes, the more contemplation it requires to figure out problems. This is universally true, but you can't simply discuss complexity away. And adding complex object models on top without consideration doesn't make your task easier just like that.
Are we really talking about the “price” of managed code, when C code is chock-full of linked list data structures? A list of boxed objects is more cache-friendly than that. It is simply not always that performance sensitive to begin with (e.g. does it really matter if it queries the available display connectors in 0.0001s or 10x that?)
Regarding MirageOS, they actually achieved better performance by using a managed language than contemporary C OSs for select tasks. This is possible due to context switches being very expensive, and a managed env can get away with (some?) of those.
Context switching can be expensive, but I don't see what's specific about managed envs about that. Fundamentally you have to have trust in code, and need to have the hardware support to enforce authenticity of the trusted code in order to avoid context switching. A different, promising development to reduce context switches is CPUs growing more and more cores, and more and more kernel resources being available through io_uring and similar async interfaces.
How did we start talking about Java and VMs? Is this some sort of strawman argument?
> But sure, the more complex a system becomes, the more contemplation it requires to figure out problems
Strawman again? I wasn't talking about complexity. I said that the more "systems" your programming is, the higher is the cost of memory safety bugs.
Is it a strawman if my comment was in response to "managed languages"?
> Strawman again? I wasn't talking about complexity. I said that the more "systems" your programming is, the higher is the cost of memory safety bugs.
For clarity, you spoke about the cost of debugging of memory bugs. And I said, it's universally true that programs are harder to debug the more "systems" they get. The reason is that it typically isn't sufficient to simply trace an individual thread anymore. "Logical tasks" are served over a number of event handlers executed in various (OS) threads.
It's not first and foremost a refutal of what you said. But an observation that I even placed in opposition to my other statement that it's not quite true that you can't use debuggers with kernels. FWIW and so on. I don't get why you are calling "strawman" repeatedly, and don't get the aggressive tone of your comments.
Who would choose C/C++ and not something like Go in that situation?
I think whether there "should" be a law making you liable could depend on the details of the exploit.
If you get exploited via rowhammer, I don't think anyone would blame you. It would be unreasonable if every small business running a website could be sued if they didn't defend against electromagnetic interference within the RAM.
However, if you're Apple and say -- you could get pwned because someone clicked a button to register version 9000 on the public npm/pypi registry (https://medium.com/@alex.birsan/dependency-confusion-4a5d60f...) -- maybe I agree there's an argument for some accountability there :)
Computing is the only industry, where people accept to live with tainted goods instead of forcing whoever sold them to pay back, cover for their damage or whatever.
We already have high integrity computing, digital stores with returns, consulting with warranty clauses, and some countries are finally waking up that computing shouldn't be a special snowflake.
https://www.twobirds.com/en/insights/2021/germany/the-german...
I agree if there's a high social cost to a breach then the government should punish those involved. Also, the security of your software depends on your threat model and which threats are in scope and you're willing to invest in protecting against. The tradeoff is ease of development and velocity. So maybe such laws will incentive this process differently, and maybe it's a worthwhile change.
I look at computing as a big experiment. Personally, I am very careful to use trustworthy services and don't depend on software for anything critical (besides banking, but luckily FDIC). Most people don't take the same precautions and rely very heavily. It's obviously critical infrastructure at this point. Maybe it's time to stop thinking of it as an experiment, and maybe these laws make sense.
I don't like the concept for emotional reasons; to me it's sad and signals another step towards the end of the golden age of the internet.
Worst Rust feature by far.
I understand how Rust solves some problems and these are indeed very important ones. But it still is a constraint that has to prove itself.
C is horrible, out of the question. Really dated too without thousands of band aids. C++ has millions of those. But why not start re-implementing stuff in Rust if that is so close to your heart?
We can also reimplement everything in JavaScript. It is memory safe too. Wait, where is the enthusiasm now?
Look at js package stats - that's exactly what's happening. Many of apps and packages created today in js would be created in C/C++ few years ago. People who learn programming today don't know what C/C++ is. If they need something low-level, it's Rust.
No matter what people do, it's never enough for the critics. Many people criticize Rust fans for "rewriting it in Rust"!
> how about just doing it
People absolutely are. That's what the article is about, even.
Defenders of C might note that Android is java and IOS is not and compare the security of those two systems and say clearly memory-safe is focusing on the wrong thing. This is equally true but no more valid an argument.
The one that really bothers me in all these language-booster discussions (that we should and need to have) is the functional programming formal verification claims. We have no ssl library written in a memory-safe, functional language that has been proven correct that has dominated the space. Heartbleed wasn't yesterday.
I look down the list here: https://en.wikipedia.org/wiki/Comparison_of_TLS_implementati...
And I think something is not being discussed as far as replacing memory unsafe languages of critical security infrastructure. What is it?
Heartbleed didn't affect https://hackage.haskell.org/package/tls even though it isn't formally verified.
Why hasn't this really good result meant _everybody_ now uses that library by default and has to justify using something else?
There is something here not being discussed, what is it?
There is something here, at least one thing, that seems to be dominating outcomes, and is not being discussed.
Nobody has even a half-suggestion of what it might be and that is not making it (or them) go away as problems that are not being solved.
And I don't know about "used at scale" but we use it in production for Pernosco.
You will find at the bottom of that page C implementations of curve25519 that are proofed and derived from F* and Coq. Curve25519 is a relatively simple implementation and only one part of any system that uses it. As you can see in both of these implementations the papers recognize a team of contributors each - this should provide some insight as to the cost of such work. That doesn't make it unimportant, it just makes it rare.
Wireguard seems to be written almost entirely in C, is that right?
1) The most cost-effective way to eliminate the majority of memory bugs is to just start writing all new code in a memory-safe language. If you were going to write new code anyway, you may as well do it safely.
2) Going back and re-writing existing code that doesn't need to be changed may solve latent memory bugs, but it will likely introduce other regressions that could be worse for security or for user experience. If code doesn't need to change, it's often better to leave it as is.
Not that a rewrite is never called for, but it's not necessarily the best course of action by any metric (even when neglecting the cost).
People tend to forget that they don't have to live with those problems, but they still have a cost
Defenders of C++ argue that there's no reason to change the language, because new features around safety guarantees are being introduced into every C++ standard starting from C++11 at a remarkable pace, so remarkable that compilers implement them faster than the existing adoption rate. And the adoption rate speaks volumes about existing capacity to port/rewrite big codebases in entirely new stacks. The new stacks also tend to have fewer custom static code quality analyzers from third-party vendors, and they are used a lot in mission-critical C++ codebases.
Are these static code quality analyzers detecting code quality problems that Rust and company are also vulnerable to? Or are they mostly looking out for the hundreds of legacy footguns that C++ still officially supports?
If they’re saying that C++ can’t be saved, maybe they’re worth listening to.
It might be true, but it also sounds like an appeal to authority. I suspect there also might be voices that are being silenced or aren't given a similar platform to speak up and provide an alternative viewpoint on the matter within the same organisation, because <team budget/political reasons why>. After all, there are greenfield projects that are being started in C++20 and people are enthusiastic about their prospects. I wouldn't just blindly dismiss their reasons in favour of Google ones.
That said, they won't catch a ton of memory and thread safety issues. You'll need tests with 100% coverage for that. Or you could just write it in rust and the compiler will catch it.
And full agree, all what cppcheck does imo should have long gone into the warning suite, and Werror and Wall should be the default..
If the authority comes along with a bunch of well researched and documented data from experiences in the real world… that seems worth listening to.
It’s no longer an appeal to authority. It’s just looking at evidence.
You’re merely reading what you want between the lines.
We continue to invest in tools to improve the safety of our C/C++. Over the past few releases we’ve introduced the Scudo hardened allocator, HWASAN, GWP-ASAN, and KFENCE on production Android devices. We’ve also increased our fuzzing coverage on our existing code base. Vulnerabilities found using these tools contributed both to prevention of vulnerabilities in new code as well as vulnerabilities found in old code that are included in the above evaluation. These are important tools, and critically important for our C/C++ code. However, these alone do not account for the large shift in vulnerabilities that we’re seeing, and other projects that have deployed these technologies have not seen a major shift in their vulnerability composition. We believe Android’s ongoing shift from memory-unsafe to memory-safe languages is a major factor.Modern C++ has many memory safety features. If a company has learned that its people fail to use them, then bad for them.
Of course, there are languages that abstract memory safety to the point that they eliminate those types of mistakes. But languages are tools for a job, and only some tools are applicable where C++ is applicable. We should not bury C++ prematurely before answering the question - "what else is as fast and efficient as to replace it for OOP?" And if a project doesn't need fast and efficient code, then why is it using C or C++ in the first place?
Overall, selecting the correct tool for a job is more important than figuring out which tool is better in some abstract way.
This recapitulates an argument at least as old as C89. You can probably find Usenet posts deploying it to argue against the adoption of strncpy, because if people don't know how to use sizeof and strlen, then bad for them.
C++ nowadays can be used in a very memory-safe way without much effort. In my professional experience, memory leaks and corruptions are sporadic in modern C++ code and common in old-style pre-C++ 11 code.
That's why I'm a bit skeptical of this article from Google. It seems reasonable that Android has quite a lot of pre-C++ 11 code. And the article seems to lump two very different approaches to memory safety in pre and post-C++ 11 style programming.
Another approximation is to look at the Android source tree to see what proportion of it is as old as you assume in your argument. There are 431 results for a search `"Copyright 200" filepath:.\.cpp`. There are 7465 results for `"Copyright 201" filepath:.\.cpp`. 4152 results for `"Copyright 202" filepath:.\.cpp`. 220 for 2012, 354 for 2011. If you exclude tests the ratio is even less favorable for your theory.
In case you're wondering the project policy is to add a copyright header at the time the file is created, they do not update years in headers arbitrarily. As a spot check the first file matching "Copyright 200" that wasn't just essentially C code wrapped in extern "C" was: external/angle/src/libANGLE/Config.cpp. This file contains the use of std::make_pair.
You can perform these searches yourself here: https://cs.android.com/search?q=%22Copyright%20200%22%20file...
I've looked at many "Copyright 201" and "Copyright 202" headers. I needed to see more use of C++ 11 or equivalent smart pointers or containers to say that this codebase uses modern C++ memory safety features. Other modern C++ features (like std::make_pair that you mention) are easier to spot.
I expect this codebase to have many memory safety issues. It may not pass code review in a company/team that expects their people to use modern C++ memory safety features. After seeing it, I'm more convinced that the reason Google has so many problems with C++ in Android really is because they don't insist their engineers use modern C++ (or equivalent in-house containers/pointers).
Here's another insightful pair of searches:
" std::make_" filepath:.*\.cpp
"delete " filepath:.*\.cpp
The style guide, C++ readability, and general code review has all but banned raw "new" for years and years. You can find plenty of CVEs where the root cause is a UAF on a managed object.
Thanks for the context about UAF. I am curious about this. Much of my C++ experience comes from working with in-house reimplementations of std/stl, so my question might be a bit stupid, but how is use after freed of an obj managed by smart pointers so prevalent? Should the smart pointer not be nulled after the object is destroyed? Maybe you have a good example CVE? Are these cases of using the raw pointer in the smart pointer without checking it first?
I'm not sure what you are going for here.
The way this often happens is there is some module that owns an object with a unique_ptr and references to that object are used elsewhere. But the ownership of the object is complicated so a bug sneaks in where a non-owning reference to the object gets dereferenced after the unique_ptr is deleted. You can prevent this by having literally everything use shared_ptr for everything but that sucks for lots of reasons.
You can avoid multi-ownership problems of shared ptrs with weak pointer member variables (which only need to be turned into shared in a given {} scope). Some other problems can be solved by marking objects as pending kill without destroying them immediately and ensuring all threads finish access before actual deletion.
Unreal Engine uses both weak pointers and object marking in a global object array. It also uses GC but that's besides the point.
Would the same approach to modern memory management not help Android?
We've got like 30 years of people insisting that it really is possible to write safe C and C++ programs if you just follow the One True Way (TM) and its never been the case. Each new One True Way helps, but it sure as hell doesn't solve the problem altogether.
(And sure, it doesn't apply to every niche yet, but it sure applies to a lot of them)
And a good workman put his old/obsolete/dangerous/etc tool behind when something better show up.
The bad workman, instead, continue blaming his tools, when the problem is that he CONTINUE using bad tools, anyway!.
P.D: I learn about mechanical engineering. Get rid of bad tools fast is like key around that...
Are you serious? A bad workman blames his tools, because workmen are reponsible for their tools. A large part of being a good workman is identifying what tools are good and using them.
And C++ is a terrible tool for any task where you are not forced to use it because of existing libraries. All the memory safety features of modern C++ are a tiny, almost vanishingly small step in the right direction.
> "what else is as fast and efficient as to replace it for OOP?" And if a project doesn't need fast and efficient code, then why is it using C or C++ in the first place?
If you need fast and efficient code, why on earth would you be doing OOP?
As I said in the comment to which you are responding, "selecting the correct tool for a job is more important than figuring out which tool is better in some abstract way."
> C++ is a terrible tool for any task where you are not forced to use it
Many game developers, OS developers, and massive hardware-software makers doing embedded programming who use C++ would disagree. What would you say to them?
> If you need fast and efficient code, why on earth would you be doing OOP?
For small projects, I could agree. What would your recommended alternative be for massive codebases in large tech companies that need fast and efficient code?
P.S. Please read https://news.ycombinator.com/newsguidelines.html about snarky comments. Thanks.
The poster wrote very clearly:
> where you are not forced to use it
If one is forced, there's obviously no option.
The idea is not that C/C++ should be replaced right now, rather, that devs finally understand that C/C++ should not be used where possible.
I actually see this pattern used by some, who defend C/++: "C/++" should be deprecated" - "No, it's impossible to eliminate C/++ today".
Deprecation is not elimination. Linux started introducing it, and Google is doing as well, so it can be done gradually.
>Are you serious? A bad workman blames his tools, because workmen are reponsible for their tools. A large part of being a good workman is identifying what tools are good and using them.
Also C/C++ made into real life tools would be OHSA violation on OHSA violation in real world
We have an answer: Rust. It's no longer premature, bury it.
See how many reference types are there, how async is handled and the underspecified unsafe semantic.
For higher level tasks, I prefer a language with GC like go or java. Rust can work with references counting, but it don’t mix well with the larger ecosystem. For lower level task, the underspecified unsafe model make it worse than C aliasing problem
2, & and &mut. What else?
I would love to know about any projects that do OOP well in Rust.
Concepts like sugar syntax and syntax salt do exist
Approaches to problems may vary by languages because lang environment shapes its users in some ways
Google's one of the worst C++ shops because their code standard basically forbids using modern C++, and their C++ is more like 90s Java than modern C++. It's no wonder they want to get away from it.
I write C++ at Google, and it encourages use of modern C++ features, and many things you see adopted in std have roots in our libraries.
I'm curious what you think Google prevents us from using and why you think our C++ is like 90s Java.
https://abseil.io/tips has a lot of our philosophies and abseil is chunks of our internal libraries published externally.
I think it stems from the fact that Google was slow at making C++11 available internally so there was a period of time where the rest of the world was using smart pointers and we couldn't. That may have just solidified a "Google uses old C++" meme out in the wild despite it being wildly out of date.
Modern C++ is indeed a huge upgrade on what came before and with a good amount of static and dynamic analysis the state of low-level programming is much better now, but there really is no reason for new programs to be written with these. Besides the bottom of the stack, managed languages are more than fast enough for nigh everything.
And yet, there is a lot of enthusiasm (at least here on HN) for web development in Rust...
Someone can think it's fine to write a web application in java/python/js and Rust.
It seems plausible that splashy projects in new languages are better for careers than grinding through "stable" codebases using "boring" engineering practices.
I also gather that Google has a challenge, possibly for similar reasons, keeping their third party dependencies updated and up to standards. A lot of those are written in C and C++, probably.
One of the people most involved in the systems described in the blog post that are used to harden the C++ side of things is a L9 here.
Designing the ultimate everything sanitizer with zero performance overhead would surely be impressive even at Google. Especially if it was actually adopted across the org.
And I assure you that, despite the memes, code health efforts do end up with promos here. The org responsible for third_party and large scale code health had above average promo rates for ages.
The same question comes up when an existing system is rewritten from language A to language B and big performance gains are seen. The language could be the big cause, but so could the extra engineering effort itself -- updated design, fresh attention to the requirements, etc.
Reducing defects is one of the main reasons (others being maintainability, readability, better integration, and similar) for refactoring and rewriting code. There's usually not enough time/money to do it, especially for large codebases.
I quite like rewriting parts of a codebase to modernize it, and I have often closed tons of bugs in a short time this way. It is definitely effective. But not as cost-effective as deprioritizing bugs into "won't fix" territory, which is what many companies like to do.
Yes, things have gotten better. Smart pointers are a godsend. Sanitizers are a godsend. Various static analysis tools work pretty well.
But even codebases that adopt all of these things religiously still are riddled with security vulns.
#RustForTheWin
Well, first of all, this is said but not proven.
But it's easy to prove that memory safety bugs are not a significant percentage of the total number of bugs, even Google agrees.
Vulnerabilities are not the same thing as bugs, a vulnerability like spectre or meltdown are not due to a bug in the software, have an ubiquitous immediate impact on 100% of the devices and are much harder to fix or mitigate, sometimes it's could even prove impossible.
The same bias can be explained using the exact same words used in the article
"Despite most of the existing code in Android being in C/C++, most of Android’s API surface is implemented in Java. This means that Java is disproportionately represented in the OS’s attack surface that is reachable by apps."
It can be read as: of course most of the vulnerabilities are due to memory safety bugs, it's much harder to gain root privileges exploiting a bug on the colors of a specific element of the UI, assuming it would be possible.
It can also be read as: most of the userland software is based on Java, which is memory safe by default, assuming there are no bugs in the implementation of the JVM, which is entirely not Java.
Given that, the problems become
- rewriting the entire ecosystem in memory safe languages requires rewriting everything from scratch, which is a task that even Google will have huge problems to complete (reminder: Google is the number one killer of its own projects) in reasonable time or without wasting more money that it's worth on it. Is an half complete not battle tested complete rewrite actually safer? Historical data says it usually isn't.
- are the user actually safer when memory is safe? I mean, memory safety bugs gave us jailbreaking for locked devices, memory safe languages gave us bugs like CVE-2021-44832
I wouldn't classify the issue as black/white, there's a lot of grey to be considered.
Using language-independent bug example in discussion about language-caused bug vectors isn't exactly honest.
Rust would stop Heartbleed for example, and that was one of huge vulnerabilities
using a low budget project with few developers maintaining one of the most used libraries in the whole World as an example of non memory safe languages perils is not exactly honest.
Heartbleed could have been easily fixed if the companies profiting from using OpenSSL donated a few more eyes to look at the code.
Similarly to what happened to Log4j bug, which had an enormous impact, similar to the heartbleed one and affected a fully memory safe language.
Why? That's the situation of enormous amount of code people use.
Anyway heartbleed was discovered after many years, let's wait the same many years and see what kind of bugs we'll find in code written today with different languages.
Full disclaimer: I do not write C code since long time ago and have no intention of going back, but dogmatic programmers that believe in "saviours" are a real mistery to me.
The same people forgetting to free a resource are the same people that will forget to sanitize some input, meaning all of us make mistakes and will keep making them, in any language and those mistakes will be abused by some malevolent actor.
Google problems are not everyone's problems.
Goggle's solutions to problems are not everyone's solutions to the same problems.
Assuming that what Google says is applicable everywhere is at best naive.
.. unless said language makes making those mistakes difficult or impossible. Sanitizing input for example has not been an issue for me for decades, as every framework I used handles that by default, I'd have to work extra hard to make a mistake there.
> Google problems are not everyone's problems.
In this case they are. Not only memory safety is an issue for many codebases that have at least some C somewhere, but also because Google products are used by milions.
> Goggle's solutions to problems are not everyone's solutions to the same problems. > Assuming that what Google says is applicable everywhere is at best naive
Any other time I'd agree with you, but I don't see anything Google-specific here.
You're missing the point [1]
(or I was unclear)
Yes, improvements in neuro surgery can save lives, but the bulk of preventable deaths it's in human mistakes [2] that are almost impossible to make impossible.
Just like the majority of the bugs are not prevented using rust, just a minority of them, which are also arguably the hardest to find and exploit, while a SQL injection can be exploid by a script kiddie with average IQ.
[1] https://portswigger.net/daily-swig/mastodon-users-vulnerable...
[2] The three risk factors most commonly leading to preventable death in the population of the United States are smoking, high blood pressure, and being overweight.
"For more than a decade, memory safety vulnerabilities have consistently represented more than 65% of vulnerabilities". That's not minority. We are getting into minority territory now because of Rust and other memory-safe languages.
> while a SQL injection can be exploid by a script kiddie with average IQ.
Use any framework and it's solved problem.
Even more evidence that the negative performance impact of bounds checking is minimal, nay, it can even be positive.
But if "security" isn't remotely a concern for a given project (like almost anything graphics / gaming related), this is not at all evidence for changing anything. It could be that Rust's optimizer eliminates the bounds checking so regularly as to be a moot point, but this isn't saying anything of the sort. It's saying that the cost, whatever it was, was judged to be worth paying for the improved security for these projects
I'm sure it's possible to construct such a thing, but I cannot imagine it ever being common enough to show up on any sort of head to head comparison.
The question isn't "does Rust have bad bounds checking optimizations" but rather "what is this mythical heavily-bounds-checked C code that the compiler can't optimize away?"
Then people insist on wanting to replace every x[i] in prod with x.get_unchecked(i) only to learn that, not only was that indexing not slowing the code down (the branch is perfectly predictable in a correct program!), but actually any difference is so in the noise that the random perturbation is worse (or that the asserts were actually adding extra facts for more profitable optimizations in llvm).
There is definitely specific hot loops with weird access patterns where it can be high impact but those are the exception, not the rule, as the Android team demonstrated.
I don't know that I've ever actually manually indexed an array over years of using Rust.
> Can you point to any such bounds check in C
Now, in modern C you can say you don't have aliasing, but you're probably wrong and so there's a high risk when you do that you now get "impossible" bugs because you swore to the optimiser that if X changes, Y is unaffected, then created a situation where that wasn't true and now your program has no defined meaning, which is going to be tricky to debug. So, on the whole C programmers do not use this, indeed in places like the Linux kernel they even turn off the C standard's very minimal aliasing rules (which forbid aliasing objects of different types), they just don't trust themselves.
Like yeah there's aliasing changes, but in "idiomatic" C/C++ how is that getting you bounds checking that's not being optimized away fairly consistently?
Gaming platforms have gotten a lot less lenient over time, and with pretty much every game these days having online components, "security isn't remotely a concern" has become a lot less true.
When C++ dies (which in 20 years it will, and I wasn't that hopeful 20 years ago), people will look in the same bewilderment at excuses made for its insane behavior as they look now at mid-20th-century arguments against high-level languages (and assemblers before them - like Mel said, "you never know where it will put things so you'd have to use separate constants.")
That's too optimistic. C++ would easily die in 20 years if it didn't already have 30+ years of still-active legacy that can't easily be converted or rewritten.
I've recently even had to start new projects in C++ because platforms I depend on demand it or because I have to interface with existing code and libraries that still only exist as C++. I'm not a fan of the language by any means, but I'll eat my shoe if it's "dead" in 20 years for anything except maybe greenfield development.
What if the better analogy is updating building codes in Manhattan?
But anyway, in my experience a lot of the crashes are often due to quickly hacked together plugins for the various DCCs written for artists, that don't have good error checking or testing, and it's not completely clear to me how that situation's going to improve that much with something like Rust, if the same programmer time constraints are going to exist in writing them: i.e. I think it's very likely people will just unwrap() their way to getting things to compile instead of correctly handling errors, so it will be the same situation from the artists' perspective: technically it may be a panic rather than a segfault, but from the artists' perspective, it will likely be identical and take the DCC down.
In my experiences with university HPC clusters, security is very important because you have a lot of young students with no Unix experience accessing the resources. We've had real compromises of individual research machines because of this.
This happens all the time at research universities, but it's not always public. In one public example from my uni, hackers from China compromised a research machine, which was used to attack IT infrastructure, which lead to PII including SSNs being compromised.
I'm talking about the actual HPC algorithm code heavily priortising performance (or in some cases memory efficiency), at the expense of pretty much everything else (other than correctness, obviously).
Students writing code are not prioritizing security or performance. (I've seen FEM analysis written in Matlab, large neural-networks written in nearly-pure Python, etc.) The real 'performance' priority is human time, at the cost of everything else. To this extent, extra security "for free" from memory safety is nice.
There are exceptions, of course. The 2012 AlexNet breakthrough was a result of performance-engineering, for example. But generally speaking, publish-or-perish rewards neither optimizing performance nor optimizing security.
So, students will be installing Docker images (which have super user privileges), sudo running bash scripts, sudo installing pip or npm packages. I've seen students replace libraries (including CUDA) with modded binary blobs from researchers from other universities. All to save time in pursuit of ~~interesting~~ publishable results.
These are horrible things I've seen during my time in academia. We (should) do virtualization, jails, firewalls, etc. to insulate the rest of us from these horrible things. (I'd add "keep machines offline", but that's rare, and even rarer because of the pandemic.) This insulation is imperfect, and many of those imperfections are due to memory safety flaws.
Rust is starting to make inroads in HPC/scientific computing. The libraries have a ways to go for widespread end-to-end adoption, but to give a concrete example, a current project has drastically beaten OpenBLAS across a suite of matrix factorizations. It was developed over a few months by one person with much less arch-specific or unsafe code. (The library is on GitHub/crates.io, but the author isn't ready for a public announcement so I won't link it yet.) Expect to see lots more Rust in HPC over the next few years.
I do think Rust will make inroads, but more because of better WASM toolchain,so loading data into the browser is significantly easier than with JS (e.g. https://crates.io/crates/moc).
Many times when I've helped some researcher make their code run on a cluster I have discovered that the code crashes at runtime if bounds checking is enabled. The usual response is that "this can't be a problem because we've (or someone else) published papers with results computed with this program". Sorry sunshine, this isn't how it works. Maybe the corruption is entirely benign, but how can you tell?
And so you save the overhead of the extra refcounting or the re-entrant locks (though that would be unsound anyway), and you can safely use anything which works off of references.
These games would go on sale once a year or whatever and attract new players, and people would post warnings in the steam forums and whatnot to try to stop people from being effected by the issue, but I am sure some people either didn't listen or didn't notice the warnings.
I haven't looked into the issue in a few years at this point, but it's very possible that the games are still unpatched and being listed in the store to this day.
Anything that connects to the internet needs to be strongly concerned about security!
But especially desktop linux.. every random bash script could encrypt your documents, or leak out your browser cache, do whatever it wants..
Though that makes non-memory safe code more reliable in the case of a bit flip. (this is not a serious advantage - only some bit flips can be prevented this way)
Unchecked Access (unsafe pointers): http://www.ada-auth.org/standards/22rm/html/RM-13-10.html#I5...
Unchecked Deallocations ("free"): http://www.ada-auth.org/standards/22rm/html/RM-13-11-2.html#...
Unchecked type conversions (unsafe casting): http://www.ada-auth.org/standards/22rm/html/RM-13-9.html#I57...
Ada has FFI:
C / C++: http://www.ada-auth.org/standards/22rm/html/RM-B-3.html
COBOL: http://www.ada-auth.org/standards/22rm/html/RM-B-4.html
Fortran: http://www.ada-auth.org/standards/22rm/html/RM-B-5.html
Ada has pragmas to both enable more security measures or relax security measures:
http://www.ada-auth.org/standards/22rm/html/RM-L.html
Just like in unsafe Rust, sometimes in Ada you need to turn off some security features or tell the compiler "I know what I am doing for this part" when interfacing with some hardware or similar low-level stuff.
No language can prevent a person from allocating a writable buffer, then reusing it without cleaning it. Do that on a server, and you have step 1 to a security vulnerability.
If requests to allocate memory come faster than the garbage can be collected.
Or a data container holding many/large references that will never be used. The difference between that and a lost pointer in C are moot in a practical sense.
All of these _can_ be prevented. But it's programmer care, rather than the language, that prevents them. Hence, the term "memory safe" is inaccurate. "Memory safer" would be more accurate, but far less catchy.
1. We love to simplify matters down to black & white thinking with absolutist statements.
2. Attention spans are short (and probably getting shorter), so we try very hard to be pithy.
3. General seeming statements are actually narrower than they appear.
4. When taking a statement at its literal absolutist meaning leads you to an absurd conclusion, you're "supposed" to use your own judgment to interpret it imprecisely rather than ridiculously.
"memory safety" fits these criteria pretty well, especially the third point. Clearly, you really can't have a programming language be practical/general-purpose while simultaneously being completely and totally "memory safe." It's just ridiculous given our current predominant operating systems and architectures. The pithiness of "memory safety" relies on you, dear reader, knowing that and interpreting "memory safety" as something more reasonable than that.
> No language can prevent a person from allocating a writable buffer, then reusing it without cleaning it. Do that on a server, and you have step 1 to a security vulnerability.
This is a good example of (4), where you interpret something generally, but it's actually much narrower. When folks say "memory safety," they are not referring to the problem you speak of here. The problem you speak of might be a vulnerability, but it is not, in and of itself, something that would be recognized as memory safety. A memory safety bug could lead to the circumstances you describe, but it is not necessary. (Some people like to claim that other people think memory safety is the only kind of safety that matters, but few people with any credibility actually espouse that view as far as I'm aware. But it's important to call out: if you fixed every single memory safety issue ever, you would not fix every single security or vulnerability issue.)
The important life lesson here is that jargon is abound, and a good skill to pick up is knowing when to recognize it. If we go around interpreting every very literally, it's going to be a bad time.
We could also stubbornly demand that everyone use crystal clear, unambiguous, precise and accurate terms all of the time everywhere so that nobody ever gets confused about anything ever again. But of course, I'm quite certain that is simply not possible.
> which makes it easier for developers to compartmentalize code to achieve memory-safety
The problem here is that this is incomplete. Many many many languages have achieved this before Rust. Where Rust is (somewhat although not entirely) unique is bringing this compartmentalization into a context that (mostly) lacks a runtime and garbage collection.
I have no problems calling Rust a "memory safe language" precisely because I have no problems calling Java or Python "memory safe languages." What matters isn't whether the language is "entirely" memory safe. What matters is what its default is. C and C++ are by default unsafe everywhere. Rust, Java, Python and many others are all safe by default everywhere. This notion is, IMO, synonymous with the more pithy "memory safe language."
Yep, ctypes is part of the stdlib and lets you corrupt the VM on the fly. Fun stuff like changing the value of cached integers and everything.
But ctypes being a terrifying pain in the ass, people tread very carefully around it. Cffi’s a lot better though it requires an external package. At the end of the day I think I’d be more enclined to bind through pyo3 or cython than write C in python (which is what ctypes has you do without even what little type system C has, to say nothing of -Wall -Weverything).
I'm not sure how much people treading carefully actually translates into safety in practice.
CPython in particular has ad-hoc refcounting semantics where references can either be borrowed or stolen and you have to carefully verify both the documentation and implementation of functions you call because it's the wild west and nothing can be trusted: https://docs.python.org/3.9/c-api/intro.html#reference-count...
This ad-hoc borrowed vs stolen references convention bleeds into cffi as well. If you annotate an FFI function as returning `py_object`, cffi assumes that the reference is stolen and thus won't increment the ref count. However, if that same function instead returns a `struct` containing a `py_object`, cffi assumes the reference is borrowed and will increment the ref count instead.
So a harmless looking refactoring that changes a directly returned `py_object` into a composite `struct` containing a `py_object` is now a memory leak.
Memory leaks aren't so bad (even Rust treats them as safe after the leakpocalypse [1] [2]). It's when you go the other way and treat what should have been a borrowed reference as stolen that real bad things happen.
Here's a quick demo that deallocates the `None` singleton:
Python 3.9.13 (main, May 17 2022, 14:19:07)
[GCC 11.3.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import sys
>>> sys.getrefcount(None)
4584
>>> import ctypes
>>> ctypes.pythonapi.Py_DecRef.argtypes = [ctypes.py_object]
>>> for i in range(5000):
... ctypes.pythonapi.Py_DecRef(None)
...
0
0
0
0
0
[snip]
Fatal Python error: none_dealloc: deallocating None
Python runtime state: initialized
Current thread 0x00007f28b22b7740 (most recent call first):
File "<stdin>", line 2 in <module>
fish: Job 1, 'python3' terminated by signal SIGABRT (Abort)
[1]: https://rust-lang.github.io/rfcs/1066-safe-mem-forget.html
[2] https://cglab.ca/~abeinges/blah/everyone-poops/As I said, you can trivially corrupt the VM through ctypes. However I don't think I've ever seen anyone wilfully interact with the VM for reasons other than shit and giggles.
The few uses of ctypes I've seen were actual FFI (interacting with native libraries), and IME it's rare enough and alien enough that people tread quite carefully around that. I've actually seen a lot less care with the native library on the other side of the FFI call than on the FFI call itself (I've had to point issues with that just this morning during a code review, if anything the ctypes call was over-protected, otoh the update to the so's source had multiple major issues).
I think it's fair to call Rust a memory safe language. But I don't think it's on the same tier as a fully managed language like python.
It's not like you need `unsafe` in Rust to build every data structure. I build oodles of data structures on top of the fundamental primitives provided by std without using any `unsafe` explicitly whatsoever.
And it is not at all uncommon to write application code in Rust that doesn't utter `unsafe` at all. Even ripgrep has almost none of it. At the "application" level it has exactly two uses: one related to PCRE2 shenanigans and one related to the use of file backed memory maps. Both of those things are optional.
Then there's another whole perspective here, which is that if you're using Rust in the first place, there's a non-trivial chance you're working on something "low level" that might require `unsafe`. Where as with Python you probably aren't doing "low level" work and just don't care much about perf within certain contexts. That has less (albeit not "nothing") to do with the design of the languages and more to do with the problems you're trying to solve.
To be clear, I am not saying you're definitely wrong. But as someone who has written many tens of thousands of lines of both Rust and Python, I would put them on the same or very very close level in terms of memory safety personally. Certainly within the same tier.
You can't write any code at all in Python or Java without relying on unsafe operations. Both of them have their runtimes written in C/C++.
So based off of this unusual line of reasoning, Rust is strictly more memory safe than either of those as it's at least possible to have a Rust program without any unsafe code. That program will be of questionable value, sure, but it can at least exist at all whereas it can't for Python or Java.
We are discussing the languages themselves, not any particular implementation.
Technically, pypy is a Python runtime written in Python.
There are Python interpreters written in other languages: there is one in Rust, and there is Jython and IronPython.
At some level of the stack some Assembly or compiler intrisics are needed, not at every line of code.
Also bootstrapting a language always requiring using its subset for low level layers, apparently not an issue that many parts of C, C++ cannot be implemented only with what ISO provides on the standard.
This isn't a meaningful distinction, in the end. Hardware is unsafe too. Real production CPUs have bugs in them which lead to cache lines becoming corrupted, address translations being wrong, branches going to the wrong place, etc. under extremely weird conditions. But, in the end, we don't really do much about it because we trust that it probably won't impact us since we assume the people who built the SoCs or those who wrote the standard library did a good enough job.
Interpreters and the rest of the VM is a different beast, while they also have to be bootstrapped from some unsafe language one way or another, they are usually written in a much more expert, security- and correctness oriented way than your average program. So while they can and do have bugs, they are exceptionally well tested and, well, I wouldn’t expect the JVM to die out under my program the same way you don’t really expect the kernel to freeze either. This is also true of Rust stdlibs, I assume, but is it true of third party libs?
This isn't possible. Eventually you are sitting at a block of memory and need to write the allocator. Maybe (like python) your allocator is written in C and you hide it, but there is always something that isn't memory safe sitting under your language.
You could write a language for an actual Turing machine which since it has infinite memory is by definition memory safe. However as soon as you need to run on real hardware you have to work with something unsafe.
You can of course prove a memory allocator is correct, but it would still have to use unsafe in rust. I supposed you could them implement this alloator in hardware, and make rust use that - but since this doesn't seem like it will happen I'm going with all languages have unsafe somewhere at the bottom.
The article was about languages being used to implement Android. Clearly, no, you can't have an entirely memory safe language that can be used to implement Android, for the reason you said. But there's a wide gap between "practical for doing useful work of any kind" and "practical for implementing Android".
Then, "entirely". What's "entirely"? Entirely until you get to library calls? Entirely until you get to OS calls? Entirely including the OS? If you include the OS then again, you are right for the reason you said. But if you exclude the OS, I'm not so certain.
Although I did use the weasel word "practical" to narrow the field. If you don't limit yourself to general purpose languages, then I'm sure you can find one that is "entirely" safe.
Playing devil's advocate, I can think of at least one language which has no escape hatches: Javascript running within a web page.
The hard part comes at allowing it to do something useful, but only the parts I believe should be able to. E.g. plugging in file system access to our brainfuck interpreter will make it quite unsafe. Node for example does have C FFI.
Whatever can be compiled to BPF meets this requirement. The price though is that it wouldn't be very useful.
> Perhaps this is related to the problem domain.
Yes, I included that possibility in my comment here: https://news.ycombinator.com/item?id=33821787
A Rust library for some sort of mathematical modelling might well need no unsafe at all, while a Java library for controlling some hardware might soon turn into JNI talking to some C++ code and oops you're unsafe.
In C# you need to reach for unsafe to do some of the stuff Rust can just do safely anyway. Did you know a C# struct with an array of 8 ints in it, doesn't actually have the eight ints baked inside the struct? It was easier in the CLR not to do that, so they didn't. Which means C# structs which look like a compact single object that surely lives in a single cache line don't actually do that in safe C#. You need unsafe.
You beed it for that feature. It is questionable whether you really want to mandate a special memory layout (because you can’t really do that even in Rust, you don’t have explicit control of struct alignments, paddings, order(!) )
Is there anything else that the various repr options don’t give you? My team at work does OS dev in Rust, and haven’t ever run into cases where Rust can’t do what we need it to do in these cases.
Well, my specific case is writing a fast interpreter in Rust, where I would like to use elements like a stackframe from both inline asm and proper Rust code. In my first iteration I chose a dynamically sized u64 array, wrapped in a safe API, because I couldn’t be more specific. But even with known size elements the best I can do - to my knowledge - is Layout? Or just a raw pointer and a wrapper with helper functions, as otherwise I can’t modify the object in question from both places.
It’s sort of tough because I am only familiar in passing with the patterns in that type of code, but Layout is an allocator API, so I’m not 100% sure why it would be used here. I’d guess that if I was doing something like this, I’d be casting it to and from a struct that’s defined correctly. This is one area where stuff is a little simpler than C, thanks to the lack of TBAA, though many projects do turn that off.
https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
In actual native code produced by RyuJit, you don't need to worry about cache lines for single instances, because the struct might not even exist at all, the Jit having mapped fields into CPU registers instead.
When it matters, like the struct being part of an array, use StructLayout.
https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
You main issue was how structures arrange their fields.
Also regarding arrays and structs, as of C# 7 you can use fixed to declare static arrays inside structs, however these structs need to be marked as unsafe.
That is exactly what I was talking about.
However there are actually good reasons for it to be unsafe, although it is debatable if that alone should be it.
One is due to the interactions with the GC, in case it moves the data and there are references to its elements, and stack size.
One way to get around it is to use AoS instead of SoA, which is any the best option if performance is the ultimate goal.
This has also advantages in that you don't need to allocate the struct in a coherent memory block. Edge case of course, but there are domains where this is relevant.
There was an allocation bug once because unsafe code needs to be allocated consecutively but most memory checks that only returned available memory failed to account for fragmented memory.
The high level idea of my original rebuke was this idea that Rust was somehow lesser because it isn't "entirely" memory safe, and that its purpose was to divide safe from unsafe. But that really misses some very big points, because the programming language implementations used to build programs virtually everywhere are similarly not "entirely" memory safe, and many many many languages before Rust divided safe from unsafe.
Notice how I modified my rebuke to include your caveat. Does my point change? Does the strength of my rebuttal change? Does anything materially change at all, other than using yet more word vomit to account for caveat? No, I don't think there's anything materially different other than more words.
I tried to sidestep all of this by using the weasel word "practical." So next time I'll just say, "any practical non-sandboxed programming language." You might still chide me for confusing "programming language" with "implementation of programming language," but I've never much cared for that semantic because the ambiguity is almost always obviously resolvable from the context.
> It’s not very useful to talk about the memory safety of languages as a whole without looking at specific implementations.
Not sure I would agree with this, but it probably depends on what you mean. We can meaningfully discuss the memory safety properties of the programming languages (not just the implementations) of Rust, C and C++. I think you have to still acknowledge the practical realities of any particular implementation that others will use to build real programs, but I contend you need not do so more than what the language design does on its own already. Because languages aren't designed in a vacuum. Even if you can build an abstract machine, for example, C was not designed to be an abstract machine. It was designed to get stuff done in the real world, and the real world influenced that design. Same for Rust.
Things like CHERI will potentially change this conversation quite a bit. I was even thinking about it when I wrote my original comment in this thread. But I think it is, at present, covered by the weasel word "practical." It isn't practical to use CHERI yet, as far as I know.
Memory safety is, as you have already mentioned, not black and white: I wouldn't even put it on an axis, because that suggests the scale is one-dimensional, and I don't even think it is practical to discuss it in that context. I prefer to categorize languages (for a definition of "language") in a couple of rough groups where most of them hang out.
In the first group is C and C++ as you're typically used to it, where pretty much every operation can do something unsafe and there's really no safe subset of the language, much less safety by default.
The second group is the "safe by default" languages like Rust or Python or Java, were you can write functional programs in the entirely safe subset (which is usually the default). This is where things get more complicated, though, because what the unsafe bits look like differ. Some give you language-level constructs to do unsafe things, such as Rust (with unsafe) and Java (with sun.misc.Unsafe or whatever). I think CPython technically also falls here because of some weird implementation choices where you can corrupt memory, but it's really more of being in the other category where you can do unsafe things via FFI and external interfaces. That's kind of where most Lua implementations live, or nodejs stuff.
Then you have the things which (usually intentionally) do not give you any of these things. That's JavaScript or WebAssembly in a browser. The final stop in this line is where you start placing significant limits to what the language itself can do, such as eBPF running in the kernel, or domain-specific parsers like WUFFS.
I've been pretty sloppy with what I call a "language" here, because you can always take a programming language and slap memory safety on it: though not trivial, you can sandbox it, pick some subsets of it, put in hardware, etc. (FWIW CHERI doesn't actually make C/C++ completely memory safe, it just helps.) And going the other way is pretty easy, you just add features to let programs mess with the execution environment.
I get that the comment that you're replying to is trying to well acktually you and I agree with the rest of your response, but the takeaway I have here is "you [the commenter you were responding to originally are coming in with a definition of memory safety, yes in this context Rust does have these escape hatches and this is what they do, but in vernacular it is safe because this is how we typically evaluate languages for this sort of thing". Which, again, is like 90% of what you wrote already, I just think that it is probably worth bringing up that there is a pretty common environment for a popular language that actually takes things a step further than this, with whatever tradeoffs that entails. Not really a disagreement, just a "hey I think this is worth mentioning".
The JVM has well-defined bad execution as well, e.g. data racing is well-defined. Safe Rust does prevent data races statically, but if they do happen due to a bad unsafe block, you are entirely on your own. While memory safety can abruptly stop both processes, FFI is very rare in Java, it is an almost completely pure platform being Java all the ways down, so in my experience the former is safer from this aspect.
Sounds hard to do and you also haven't accounted for what problems are being solved in each language. I can pretty much decide to never ever use `unsafe` again, but I'll be leaving perf on the table. If I were writing Java, I would probably be fine with that. But I'm working on interesting problems that want the most perf possible, and so U do very occasionally justify `unsafe` when writing Rust.
I’m just saying that corrupting the heap is much easier with Rust than with Java, and there is no coming back from heap corruption on a process basis, while most exceptional cases are recoverable by the JVM (hence the cliff analogy).
And Java can have surprisingly good performance, especially in multi-threaded code that has a non-predictable allocation pattern (where ARC is just not too good) — if you want significant performance improvements you really have to go down the inline asm road, which you can do from anywhere.
Now we've come full circle. I recommend you go back and read my initial comment in this thread and the comment I was responding to. You've veered far off course from there into waters in which we likely have very little disagreement of any consequence.
> And Java can have surprisingly good performance
Show me a regex engine written in Java that can compete with my own, RE2, PCRE2 or one of a number of production grade regex engines written in C, C++ or Rust. I'm not aware of any.
That Java can "have surprising good performance" is not a statement I'd ever disagree with in general terms. That has absolutely zero to do with anything I've written in this thread (or elsewhere, ever).
Will all due respect, I think you've lost the script here.
> Where Rust is (somewhat although not entirely) unique is bringing this compartmentalization into a context that (mostly) lacks a runtime and garbage collection
I just think that the model of “breaking down” is different between the two platforms and that might matter for some use cases.
> But it's not completely obvious to me that you're correct.
:-)
Which is to say, I don't know you're wrong. But it's a pretty subtle thing that requires a careful survey. And likely discussion of lots of concrete examples. It's far more nuanced than the thing I was responding to originally (not to you), which was this wrong-headed notion that Rust isn't memory safe "entirely." Because once you go down that path, the entire notion of "memory safety" starts to unravel. That is, of course Rust isn't "entirely" memory safe. Pretty much nothing practical actually is in the first place. I tried to force this issue by asking for counter-examples. The only good one I got was Javascript in browser, but that basically falls under the category of "programs in a strictly controlled sandbox" rather than "programming language" IMO.
I think this comment of mine might also be helpful, which reflects a bit on terms like "memory safety" and why they are a tricky but very common type of phenomenon: https://news.ycombinator.com/item?id=33825307
I think this has always been the goal, but it wasn't obvious at the outset that it would be achievable. The fact that we now have empirical evidence in real shipping products is significant.
But i am completely baffled by arguments that count number of unsafe blocks or code lines. Like this:
>the number of unsafe sections is a small fraction of the total code size
Code execution combinatorial effects makes number of sections or code size completely useless metrics to judge security. They do help mechanical part of auditing security in sense that they help to locate things. But locating things was never enough to judge if security is there.
I'm still thinking about how we could integrate something like that in a language or the languages package manager. I'm unsure if it's possible.
It didn't on its own, but it is worth noting that with type-safe languages, you can protect yourself from this by encoding that invariant into the type system.
Using Rust as an example, take a &std::path::Path (or &camino::Utf8Path or whatever) in your public API; have custom InternalPathBuf and InternalPath types that perform validation to ensure they aren't using relative paths to "break out" during construction, and then pass those around in your internal API. Bingo bango, now there's no way (short of transmuting, an `unsafe` operation) to pass invalid paths to the functions that hit the filesystem without a compile error. No redundant runtime checks required, and no need for you as the developer to keep track of which codepaths have already validated a Path and which haven't.
I'm sure you already know this, and I would imagine that Java can do the same, but it's a big step above languages like Python where you can do whatever you want to anything you want.
EDIT: lol while I was typing this you made a post about the same thing below.
I guess it is also an often used square peg that fits really nicely in the square Rust typesystem hole (no, not talking about those typed holes, haskellers).
"../" attacks are also just way less of an issue when you shove your programs into minimal containers, which at this point is more or less standard practice.
- great standard library, especially all the iter methods. having 'obscure' stuff like `try_for_each` just makes me so happy as a dev
- unit tests built into the lang
- tooling is great
- docs are top notch
The memory safety aspect is... sometimes helpful, sometimes irritating. I prefer zig solution (BYO allocator, special one for testing that reports errors) over rusts, which simplifies a lot of stuff and lets you make cyclical data structures without a lot of hoop jumping.
However, I appreciate that Rust tries to achieve safety through improving program correctness, not merely crashing sooner.
Of course Rust has run-time panics (hasn't solved the halting problem yet), but it also has many patterns catching problems at compile-time. For example, a hardened allocator can detect use-after-free bugs when they happen, but borrow checking can prevent them from existing in the code in the first place.
That was my concern that I had when I started learning the language. I think it is true, but I was surprised how quickly I managed to get used to the Rust way.
I mean I've improved but I still have a lot of "wtf" moments. End up doing a lot of deep copying and heap allocations just to shut the compiler up, which makes me question just how "zero cost" the safety is.
Of course, that's just "often," not nearly "always." There are many valid programs that the borrow checker struggles with, and in those cases, the "Rust way" is just needlessly convoluted. There are many problems for which the most common answer is "just store all this data in a list and use list indices like pointers." (Thankfully there has been a lot of progress on a new borrow checker that's less easily startled.)
It's because Rust hits the spot for MANY different domains and people. System programmers, functional programmers, backend, db engineers, GUI, devs from the formal verification/mission critical camp, OS devs, graphics and even frontend with WASM.
For me personally, the thing I miss in Rust is better const fn/comptime/metaprogramming (like in C++). Without them, I find writing generic high-perf data structures quite tedious (and proc macros just aren't a suitable replacement for generics)
Another big one is while the language supports lambdas and closures, it’s highly recommended to not perform any currying.
traverse_ :: Foldable t => (a -> ExceptT e IO b) -> t a -> ExceptT e IO ()
Of course I'm not saying Rust should have this (there are good reasons why this isn't a good fit for Rust, see http://smallcultfollowing.com/babysteps/blog/2016/11/09/asso...) but it really caught my eye that you called `try_for_each` obscure; it could have been an everyday function if it had a bit more expressiveness.And it's possible the lack of higher kinded types, forcing reimplementation, actually gave the opportunity to look for frindlier names.
There are some ways to get around it, by simulation and using ‘as’ but it isn’t particularly safe.
No experience with haskell - sometimes I try and learn it and give up when stack complains about version conflicts - but regardless I try not to get too many hangups about functions not being super generic, just that they are there when I need them.
- It has the best parts of C (code generation is predictable, no mandatory extra thread for GC or similar, interoperates very well with C code)
- It also has some of the niceties that were popularized after C (type inference, an equivalent of unions that isn't terrible, macros that are not terrible, functional features)
- The thread safety stuff. Async has issues, but the core thread safety constructs (send, sync, mutex vs atomic, etc) are just so well designed that rust is by far my favorite language for writing synchronous, shared-memory threaded code. It's quite possible to mess it up, but it's far easier to keep that system in my head than any comparable one that I've used. Of course shared-nothing is even easier, and rust lets you do that too!
I don't use it for everything, but when I do encounter a thing rust is good at, it's just a treat.
and
> As we noted in the original announcement, our goal is not to convert existing C/C++ to Rust, but rather to shift development of new code to memory safe languages over time.
For those working on C/C++ code bases, how does the incremental addition of Rust work in practice? What strategies and tools come in handy? Also, what are some good case studies where Rust was introduced to add new features, but still needed to work within a system mainly composed of C/C++? I've seen some of the first steps with Linux, but I'm thinking of a project where more substantial additions have been made.
The most famous example of that is AFAIK Firefox. There's a (perhaps outdated) list at https://wiki.mozilla.org/Oxidation#Rust_Components of places where Rust was used to replace some component in Firefox; the most famous of these are probably the CSS style calculation which came from the Servo project (https://hacks.mozilla.org/2017/08/inside-a-super-fast-css-en...), and a renderer also from the Servo project (https://hacks.mozilla.org/2017/10/the-whole-web-at-maximum-f...).
I can tell you that bad programmers can write bad code in Rust.
Of course. The question rust users seem to put forth is that bad programmers write _better_ (not good) rust code than C code.
That bad programmers write bad code is to expected. That is, after all, a likely explanation for why they're bad.
Rust has no null/nil pointers. Instead, it has nullable types (Option<T>), which are harder to misuse. If you want to "dereference" a Option, you either use unwrap() (which panics) or you handle the None case properly. While `.unwrap().foo()` is still just as crashy as an unguarded `x->foo()` in other languages, a code review can catch misuse of .unwrap() much more readily than a missing null/nil check.
As another example, the Send/Sync traits provide strong guarantees about thread safety; using these traits properly lets you make data structures that cannot be misused in a multithreaded program. Although deadlock is still possible (hello, halting problem), many types of data races are simply eliminated.
> The behavior is undefined if *this does not contain a value.
I'll take crashes over "probably does the wrong thing" any day.
I think it is also worth noting that unwrap itself is much safer than an unguarded `x->foo()`, since it doesn't involve undefined behavior. Meanwhile an unguarded `x->foo()` allows compiler to assume that x is never null and consequently do things that the programmer wouldn't expect. The recent post on undefined behavior [1] had a (to me) quite shocking example of that [2]. So while bad Rust programmers misusing unwrap get crashes, they won't get nasal demons.
[1] https://news.ycombinator.com/item?id=33771922
[2] https://kristerw.blogspot.com/2017/09/why-undefined-behavior...
Brave is definitely not aiming to convert the entire browser to Rust, but it's increasingly chosen for new development.
[1] https://cxx.rs/
Many informal rules of C libraries like: "don't call read() after close()", "pointer must never be null", or "keep this data alive for as long as the handle is in use" in Rust can be expressed using the type system, and enforced at compile time. Cleanup can be automated. This catches bugs in user code and shields the library form misuse.
An informal rule would be “don’t use gets” before it was deprecated, or “don’t use the str* functions”, or “don’t pass a user-provided format to printf”.
Perhaps this is an obvious point, but I found it interesting .
This probably also applies to the Linux kernal work in Rust too; if only new device drivers get written in Rust and the core remains C, it could still be a major win for kernel security.
> how does the incremental addition of Rust work in practice?
Probably a good starting point: https://firefox-source-docs.mozilla.org/writing-rust-code/cp...
I think folks who write languages should have a typographer on their team because something like this:
use std::collections::HashMap
Is a typographic nightmare. While I understand “form follows function”, it’s tough to be excited to program in something like this.
:o)
It’s the difference that makes the difference between Rust-like and C++-like gibberish and poetry.
But that use statement is not a good example of it, and if that's the biggest criticism then Rust is faring fantastically well.
I.e
use Std.String;
…. let string = String(…);
Dot is way better
https://www.doverbooks.co.uk/point-and-line-to-plane
Consider h::i::j::k versus h.i.j.k
Two :’s is an extremely loud combination of visual elements compared to the subtle point. :: drags the eye away from the content and says “look at me oscillate”
In addition, humans group similar visual elements together so a combination of anything::doesnt::matter::what::between::clumps it is impossible to escape the common pattern and the eye jumps between the ::’s. Therefore it takes double the cognitive load to read.
What’s worse, the eye gets trapped within each :: because it’s a combination of four dots, which naturally creates an implied circular pattern which draws the viewer in further.
Compare to something like: use HashMap from std.collections
Or… more obviously the python import statements.
> Two :’s is an extremely loud combination of visual elements compared to the subtle point. :: drags the eye away from the content and says “look at me oscillate”
> In addition, humans group similar visual elements together so a combination of anything::doesnt::matter::what::between::clumps it is impossible to escape the common pattern and the eye jumps between the ::’s. Therefore it takes double the cognitive load to read.
It is true that the colons take up more space and your example looks good on HN.
But this problem will be immediately solved by syntax coloring. It's just never going to come up.
Rust is safe but it’s at least as ugly as C++.
Guido got a lot right with Python because it’s clean and easy to jump in. Minimal cognitive load. But still very powerful.
I'm curious, are you dyslexic? Seeing letters move is generally something that people with dyslexia complain about.
I feel like Rust has similar pitfalls in that vein—instead of defining the user experience first, they went for a symbol that is not used elsewhere—making it easier on the engineers while satisfying the “functional requirement” of namespaces.
So—Rust will not be the language-to-end-all-languages because it is not beautiful enough. Perhaps it is almost there in functionality, but i foresee a problem with any language’s longevity unless it’s literally perfectly thought-out user experience.
I’d love to jump on the hype train with Rust but it’s not really that exciting. It doesn’t really feel all that natural. Well-thought-out user experiences should feel natural through-and-through.
That said, there are a ton of ways that auditing for vulnerabilities is just as easy or way easier. In particular, most tooling for C/C++ can be applied to rust - fuzzers and sanitizers, for example. Additionally, one only has to "grep for unsafe" and work from there to find Rust vulns, which largely amounts to "what are the assertions for this unsafe block, are they complete, are they held?".
Most of the graphs here are about new code.
[1]: https://security.googleblog.com/2021/04/rust-in-android-plat...
https://dwrensha.github.io/capnproto-rust/2022/11/30/out_of_...
Another explanation is that fuzzing is generally expensive and it's easier to focus on unsafe parts of Rust programs than whole C++ programs.
> As Android migrates away from C/C++ to Java/Kotlin/Rust, we expect the number of memory safety vulnerabilities to continue to fall. Here’s to a future where memory corruption bugs on Android are rare!
(Not that it's easy to tell from the pie chart, which colors one orange and the other red, in similar shades. Do better, guys.)
To meet the goals of improving security, stability, and quality Android-wide, we need to be able to use Rust anywhere in the codebase that native code is required. We’re implementing userspace HALs in Rust. We’re adding support for Rust in Trusted Applications. We’ve migrated VM firmware in the Android Virtualization Framework to Rust.
With support for Rust landing in Linux 6.1 we’re excited to bring memory-safety to the kernel, starting with kernel drivers."
The push for Rust to replace C and C++ continues.
This is confirmation of the NSA recommendation to discontinue using C and C++:
https://media.defense.gov/2022/Nov/10/2003112742/-1/-1/0/CSI...
If you do not know Rust, now is the time to learn:
https://doc.rust-lang.org/book/
Here are some fun holiday-themed exercises to learn with:
I plan on doing Advent of Code in Rust this year.
The article states:
"We continue to invest in tools to improve the safety of our C/C++. ... Vulnerabilities found using these tools contributed both to prevention of vulnerabilities in new code as well as vulnerabilities found in old code that are included in the above evaluation."
"These are important tools, and critically important for our C/C++ code. However, these alone do not account for the large shift in vulnerabilities that we’re seeing, and other projects that have deployed these technologies have not seen a major shift in their vulnerability composition. We believe Android’s ongoing shift from memory-unsafe to memory-safe languages is a major factor."
"our Rust code is proving to be significantly safer than pure C/C++ implementations."
C++ is great if you do not have alternatives.
If you can use Rust instead of C / C++, you should because of the greater safety it offers with similar performance characteristics.
That is well and good for legacy systems that cannot / will not be updated.
Android is not a legacy system and is adopting Rust and seeing security benefits.
Linux is not a legacy system and is slowly adopting Rust and may see greater use with time.
Time will tell, but so far the future is optimistic.
There will continue to be c/c++ applications that will need new features, new bug fixes and new development for decades.
Some of the most fundamental code on the planet for the most critical systems are written in those languages. Also a lot of video game development leans heavily on c++ in particular and that's not going to change very quickly. Big or small enterprise hardware vendors won't change to new languages overnight, abandoning potentially decades of system building experience and supporting processes. There's a cost to even being able to run rust alongside your c/c++ code. Some places may never switch or enable interoperability even for new products.
Learning c/c++ will still be lucrative and a good idea today.
> Linux is not a legacy system and is slowly adopting Rust and may see greater use with time.
100% USDA Grade-A Certified Prime BS, as shown by both projects' ages (15+, 30+ years) and maybe the definition of legacy. If you consider legacy to be something you have to keep a certain way so that it functions a certain way, then both are basically legacy technologies. What isn't legacy? I would say projects like Fuchsia and Rust itself aren't legacy because you can move fast and break things there. This also means stuff like C++ is legacy, so there's that too.
How should 'legacy' be defined?
If the software is actively developed and gets new features regularly, I would not consider it legacy.
If the software only gets maintenance / bug fix updates or is not updated, but is still used in production, then it is legacy software.
>> I would say projects like Fuchsia and Rust itself aren't legacy because you can move fast and break things
Fuchsia is not widely used (yet?). I am not sure how much "moving fast and breaking things" happens in Fuchsia.
Rust has the concept of editions (https://doc.rust-lang.org/edition-guide/editions/index.html) and spends significant effort to maintain backward compatibility.
You do not have to "move fast and break things" to not be legacy.
If this hasn't happened for PHP, maybe because there are fewer critical infra projects written in it. If it's just another web service, that can be rewritten (unless it's at the scale of Facebook). Whereas the tens of thousands of firmware code bases and other critical infra - we won't even think about them for decades as long as they continue to work (or appear to). Let's check back in 2038!
Nice, I'm doing it in chatGPT this year [0].
Many of the higher level constructs are then in Java/Kotlin, though, like the builtin apps, system ui, activity & window managers, etc...
I was very enthusiastic about Rust about 4 years ago but decided to learn Swift instead as a new “cool non-Lisp language” to learn. Seeing an recent article on writing Python in Rust made me wish that I had chosen differently.
I found this article made me feel better about security. I expect that increasing use of memory safe languages will lock down public infrastructure and software, making everyone safer.
$ swiftc -
import Dispatch
var array = [Int]()
DispatchQueue.global().async {
while true {
array.append(Int.random(in: .min...(.max)))
}
}
DispatchQueue.global().async {
while true {
let range = 0..<array.count
guard !range.isEmpty else {
continue
}
array.remove(at: Int.random(in: range))
}
}
dispatchMain()
$ ./main
Segmentation fault: 11
$ ./main
main(71894,0x16f103000) malloc: double free for ptr 0x13e808200
main(71894,0x16f103000) malloc: *** set a breakpoint in malloc_error_break to debug
Abort trap: 6
$I hope this helps to clear up any confusion.
I’m not looking for an explanation; I understand the theory. I’m just not super familiar with Swift — I’m looking for an example program that has a (potential) data race which leads to memory y safety.
When comparing the safety of Swift and Rust, it is important to note that both languages employ strict type systems and compile-time checks to prevent common memory-related bugs such as buffer overflows and access violations. However, the way they handle memory safety in multi-threaded environments differs.
In Swift, memory is managed using automatic reference counting (ARC), which tracks and manages the lifetime of objects in order to prevent memory leaks. However, ARC is only effective in single-threaded environments. In a multi-threaded environment, multiple threads can access the same memory simultaneously, which can lead to race conditions and other memory-related issues. In order to avoid these problems in a multi-threaded Swift environment, developers must use synchronization techniques such as locks or atomic operations to ensure that memory is accessed in a safe and predictable way.
Rust, on the other hand, uses a borrowing and ownership model to ensure memory safety in both single-threaded and multi-threaded environments. In this model, any value in Rust has a single owner, and the owner is responsible for managing the value's lifetime. When a value is borrowed, the borrower is given a temporary, read-only view of the value, and the original owner continues to be responsible for managing the value's lifetime. This ensures that memory is always accessed in a safe and predictable way, even in a multi-threaded environment.
Overall, both Swift and Rust offer strong guarantees of memory safety, but Rust's borrowing and ownership model may provide better guarantees in a multi-threaded environment.
Memory safety, oh so sweet
Reducing vulnerabilities, a feat
Google switches to Rust
Security improved, a must
Memory corruption, a thing of the past
Safe coding, a task at last
I don't need to program in an environment where a systems programming language is required, so for me a languages debugability and introspective capabilities are the most important. So far I think Pharo hits the sweet spot for me, but it still has a long way to go documentation wise before I think it will see any sort of adoption. I'm rooting for it though.
As someone who has been in the Rust community since a couple of months after 1.0, I only see growth, year after year. At the start the community was small. Nowadays even the OSS part of it, which is a subset of the total community, is so big you can't have an overview of it any more. Even for the Rust project itself it's hard to have an overview over everything that's happening (but it's still possible to some degree).
crates.io downloads keep growing exponentially, although recently growth rate dipped from 2× to 1.9× (not sure if that's a trend or noise). Ratio of weekday to weekend downloads keeps going up, so presumably professional adoption of Rust keeps growing.
Something like Rust which can ahead of time verify that these issues aren't present is a better solution imo. Maybe then we can use this as a last resort check that the usages of unsafe didn't cause issues.
They were warned many years ago about ownership/concurrency/compile speed, they didn't listen at all
A vision alone doesn't matter, you need skilled engineers and a dedicated team who understand what "taste" for great things is
What could have been the Swift for android will end up just being the "java" alternative, i suspect they'll get rid of the JVM, or whatever the tech is, altogether and focus on Rust for whatever they plan next (maybe Fuschia), maybe there is still hope for Swift, we'll see
It would require rearchitecting everything to work with separate processes and IPC, and that’s got some overhead(s).
It also has memory safety issues around concurrency, and its limited allowance for abstractions makes papering over those complicated.
In Rust, overflow is not undefined behavior (neither signed nor unsigned). Today, arithmetic wraps in release mode (panics in debug mode), but it may panic in release mode in the future.
The other point is that, since overflow wraps, that may indeed result in logic bugs. And in theory that logic bug could be used to lead to exploit, such as a buffer overrun. But Rust uses bounds checking by default everywhere, so the worst case you'll usually end up with here is a panic or a DoS vulnerability. Not great, but probably preferable to other types of vulns that grant access to sensitive data.
So I think two things would need to happen, at minimum, to get a non-DoS vulnerability here. You'd need to find a bug that caused arithmetic to overflow where it otherwise shouldn't, and you need to find some way to connect that to an explicitly `unsafe` unchecked access into memory.
Also Rust allows enabling overflow checking in release, it’s a perfectly valid configuration.
https://android.googlesource.com/platform/build/soong/+/refs...
Not an expert, but a file named "global.go" makes me confident :).
Yes, it's particularly a dangerous problem when doing FFI. Eg you have a `data: &[u8]` and you need to call `extern "C" fn foo(data: *const u8` and `data_len_in_bits: usize)`, you could write `data.len() * 8` for the second parameter and Rust wouldn't stop you.
It's also annoying that the only way to do checked arithmetic is replace all the simple arithmetic operators with `.checked_*()?` function calls. libstd doesn't have a `std::num::Checked<T>` like it has `Wrapping` that would implement all the arithmetic traits in terms of `checked_*` and return `Result<>`. But at least it's not hard to do yourself. ( https://github.com/Arnavion/terminal/blob/475d917377eea43b52... )
https://doc.rust-lang.org/std/primitive.u32.html#method.chec...
So this is not an issue at all.
For comparison, in Swift normal arithmetic operations check for overflow.
[1] https://doc.rust-lang.org/src/alloc/vec/mod.rs.html#1773
Until "the nukes are on our side", I don't think this should be celebrated or praised.
If the Rust community can be insufferable when trying to sell Rust at any and every opportunity, Lord have mercy when they see any kind of success because the preening and puffing will have no end. This blog post will live forever in the annals of Rust.
On the other hand I’m not surprised to see that it’s written by the security team and not a development team. Same pattern as with Microsoft: security specialists tend to get obsessed by their subject matter (just like many Rust programmers), end up hating C and C++ and see them as the source of all problems (just like many Rust programmers).
In reality, C and C++ were the saviors of Android and Google reluctantly had to conjure the NDK to make up for a severe lack of performance compared to the iPhone. Note how even in the current Android they mention improving the UI performance and having less “jank”. Without C++ Android would have been DoA, but it doesn’t look like there will be a blog post thanking C++ any time soon :-)
I can’t take seriously reports about the benefits of a tool that don’t mention any of the downsides. When the devs post a blog cursing the troubles they had getting Java and C++ and Rust to work together and pitying themselves for having to read both C++ and Rust, then we’ll be getting closer to the truth.
What should they do instead, write an eulogy to C++?
They should have the dev teams write about the good, the bad and the ugly of using Rust instead of having the security team write a sanitized rainbows and sunshine story.
This here was a blog for managers.
As an Android developer though, I have to be the one bitter old man yelling at cloud. I was spoiled by Java and Kotlin to the point where I cannot look at Rust and think it's a nice modern language.
Yes, I will start a language war today. Rust is like a truck driver who tried to make a race car that doesn't blow up. It's safe but it does not look refined. I wish you knew how weird Rust looks to me. I tried to learn it 3 times and had to give up after all the WTFs. I said it in the past, and I'll say it again:
* What the hell kind of language choice is to force everyone to type "#[derive(Debug)]" for annotations? I'd rather write past the end of an array and have someone steal my Bitcoin wallet than press "shift 3 bracket shift 9 shift 0 bracket" at the top of my structs. What's wrong with @? Nothing. @derive(debug). 2java4u? Ok, [derive debug] then.
* What's with the ' everywhere? Not the quote, the piece of dirt on your display. Why are our displays so dirty?
* Going for super short keywords "let", "fn", "mod", but then, "let mut" could have been "var". When is typing speed the bottleneck in writing software where you can't take the time to type out "module" or even 'function'? Come on now.
* print! now! fast! it's! a! macro! why! are! we! yelling!
* Passing "self" as the first argument was bullshit in Python, and it's bullshit in Rust too. Don't look at me like that - the compiler can inject it as the first parameter without requiring you to type it in.
* What does "::" do that "." can't? That's right. Nothing. All hail D.
D got language design right. Syntax better than Rust. CTFE better than Rust. Templates better than Rust. Marketing worse than Rust, though.
I'll try to learn Rust a 4th time now to stay relevant in the Android world, because I have bills to pay, but I want you to know that I blame each and everyone of you weirdos for not holding language designers to a higher bar.
Edit: if you downvote, reply with a link to the last compiler you wrote.
let mut line = String::new();
stdin().read_line(&mut line)?;
println!("input: {}", line);
I suspect that there _have_ been some changes since you last looked - for example, the ? operator lets you propagate errors in a more compact way. let mut line = String::new();
stdin().read_line(&mut line)?;
This is what I'm talking about. This is so ugly and clunky and needlessly verbose like Java compared to C++ where you can do std::string input;
std::cin >> input;
and have it Just Work. Why can't Rust do something similar?Here's another example. C++ lets you do this:
long input;
std::cin >> input;
process(input);
and this is very convenient! It's much shorter than the Rust code for doing the same. It's also wrong! If the input cannot be read, or it cannot be parsed as an integer, `input` now contains undefined content. You'd need to remember to check the contents, or end up processing undefined memory values. You cannot make this error in Rust without a lot of contortions (e.g. unsafe).Some people will say, I just want to get stuff done and not worry about all this safety junk. As someone who works in security and have exploited bugs in C/C++ programs, this is pretty much why we wound up with so many CVEs.
Are constructions like these really the root cause of CVEs? As long as you do a correct check, then you're all good. If Rust can do the same thing in fewer lines, then that may or may not be a good thing because some people might want a lower degree of abstraction or a greater one.
“An issue was discovered in slicer69 doas before 6.2 on certain platforms other than OpenBSD. On platforms without strtonum(3), sscanf was used without checking for error cases. Instead, the uninitialized variable errstr was checked and in some cases returned success even if sscanf failed. The result was that, instead of reporting that the supplied username or group name did not exist, it would execute the command as root.”
Fail to check that your input parsed correctly, wind up with privilege escalation to root.
char buf[10];
std::cin >> buf;
std::cout << buf << std::endl;
g++ -Wall -std=c++2a (sorry, I’ve only got GCC 9), no warnings or errors, happily overruns the stack when fed more than 10 characters. This code is essentially equivalent to the C gets(buf) (splitting on whitespace instead of newline), but gets is deprecated and generates big warnings on most modern C compilers. Yet, operator>>(char *) does not.Sure, you can write secure code in C++. You can write secure code in C too, even if that language is “super unsafe”. But neither language will stop you from writing insecure code either.
It’s been really eye opening as I do security analysis of C++ projects how many foot guns are in a ton of codebases that now need so much scaffolding to handle properly, after they’ve done damage.
With Rust, my first pass is usually the right pass.
And `std::cin` does not read a line from the input, you need `getline` for that.
Nope, that would be a static function on your struct. By using self you let the compiler know that you want to use an instance function.
I prefer an explicit `this` or `self` being passed, so that I know where it's defined. Something close enough, Kotlin uses `it` in lambdas. If you're nesting your lambdas, it can be tricky to review the code when just staring at it (when the IDE isn't helping you with types).
Rust could have taken the C++ route here: `fn static foo()`, `fn foo()`, `fn mut foo()`, `fn &mut foo()`, etc. but I feel like the explicit `self` is very clear and easy to understand.
`fn Box foo()`? `fn Arc foo()`? It also requires more parser lookahead.
The `mut` case is also very odd, as it's not part of the function API (it just configures the binding of the internal local). Plus if self was implicit it likely would need to be a keyword, so that wouldn't be a usecase at all anymore.
I regularly get frustrated but I don’t really see an easy alternative, it’s be a lot of work having it work let alone contextually (e.g. from pub functions for pub fields but not for private types / fields), and how would you even express the borrows / holes?
impl Foo {
// static
fn foo() {}
}
impl Foo {
// instance, owned
fn foo(self) {}
}
impl Foo {
// instance, borrowed
fn foo(&self) {}
}
impl Foo {
// instance, boxed
fn foo(self: Box<Self>) {}
}
impl Foo {
// instance, refcounted
fn foo(self: Rc<Self>) {}
} impl Foo {
// static
fn foo() {}
}
Why would I need to use self to tell the compiler about an instance method? Surely the Rust compiler is smart enough to detect this case and complain if the call site is ambiguous.So your suggestion is to remove static methods from the language?
> Why would I need to use self to tell the compiler about an instance method? Surely the Rust compiler is smart enough to detect this case and complain if the call site is ambiguous.
Detect what case? There is a dozen and eventually an infinite number of potential instance call ABIs.
And the Rust compiler is generally very much on the “refuse to guess” side of the fence (hence no global type inference), so you have to tell it what it should expect, `self`, `&mut self`, and `&self` have rather different requirements, impacts, and capabilities to say nothing of the rest.
Don't have anything against Rust though, except the syntax can be a bit obtuse.
When I choose languages, I choose them based on semantics, tooling, community/ecosystem, UX, and then maybe syntax
I really dislike both the semantics and syntax of python. However, the scientific computing libraries (pandas, numpy, scipy, scikit-learn, matplotlib, seaborn, etc) are just too productive to turn away from for quick data analysis tasks. If these libraries didn't exist, I wouldn't touch python with a 500-foot pole.
Lamborghini originally made tractors.
> press "shift 3 bracket shift 9 shift 0 bracket" at the top of my structs
Just use an IDE and you can press alt+enter. Same thing you do in Java when writing getters and setters for your AbstractFactoryBeans.
> "let mut" could have been "var"
The "mut" does not belong to the "let", it belongs to the variable name.
let (read_only, mut read_write) = (1, 2);
> Passing "self" as the first argument was bullshitExplicit is better than implicit. I quite like the names in the argument list to match the set of names available in the function, rather than having an extra keyword (`static`) that you add to remove a function parameter.
Not only that, but in Rust it actually has a meaning. Self can be &self or &mut self, which impacts the calling conventions.
The alternative would be fully typing anything but owned method calls which would just not have a self parameter, which would be incredibly weird.
It also shows that as in Python (and really even more so) "methods" are little more than syntactic sugar for functions. It would also require a new and separate syntax to declare non-instance methods.
The aesthetics of syntax is a personal matter, but when the critique of a language focuses exclusively on its syntax, it tells me that the critique is skin deep. Semantics are way more important to what code "feels" like to write.
Why not @derive? @ was a reserved token for something else before 1.0 and today is still used in patterns. Now it's too late to change.
The ' lifetime syntax was borrowed from another language. Some way of differentiating types and lifetimes is necessary, ' is not any worse than most others we could have chosen.
We try to make things that are common and safe terse, and things that are uncommon and potentially problematic more verbose. ? is common and safe, .unwrap() is less common and potentially problematic. Mutable bindings are not exactly unidiomatic, but mildly discouraged.
println! is a macro because it 1) is a compiler intrinsic to do compile time magic like the recent addition of capturing bindings directly in the formatting string and 2) it takes a variable number of arguments. Differentiating between macros and function calls is important if you want to get a sense for what the code you're reading can do. You can choose another way of differentiating them, but whatever you choose will be subjective and has to mesh well with the rest of the language.
Explicit self makes it easy syntax to differentiate between associated functions (part of the type) and methods (part of the instance), while also making very clear when you're accessing the current instance's data. And because Rust cares about mutability and ownership, you still need to communicate the differences between self, &self and &mut self. It also provides syntactic space for arbitrary self types: fn foo(self: Pin<&mut Self>)
this post have nothing to do with writing apps for android. They are talking about developing the android system itself. They are writing new system code in Rust instead of C++. Userspace code is still enoraged to be written in Kotlin.
edit2: wow mods hard at work, already removed one dissenting post.
How is any language supposed to grow if you're only allowed to use it once everybody does? Somebody has to be first (and google isn't, not by a long shot)
So you claim to have a better metric, feel free to share it?
site broken and wont let me reply any more: "shipped into production" is meaningless, compilers are directly related to language community and complexity.
It sounds like you didn't read the blog post. The whole point is that they are not rewriting code, but writing new code in memory-safe languages (including "niche" languages like Java). This strategy is paying off with fewer severe vulnerabilities due to memory safety bugs.
where did you get the impression they are not rewriting code in rust? because the author cited some other blog post saying they should focus on new code and not rewriting it? How does that translate to "The whole point is that they are not rewriting code," ???
Obviously they're going to be rewriting whatever they see fit. If YOU actually read the article you would have seen they have "keystore2" listed as new rust code. I WONDER what was keystore1 implemented in? Nice try, but no cookie.
Yes, one is named "keystore2" but you don't seem to actually know anything about it (nor do I). It seems doubtful to me that they would rewrite a perfectly working component in Rust just for fun. Most likely the original keystore, whatever language it was written in, was not meeting its requirements and needed to be replaced.
Some people are posting the article around the Internet as evidence that Rust solved security. Instead, there are many other memory safe languages around and there has been thousands in the past. Additionally, many security issues are not due to memory safety. Please keep that in mind when making comparisons.
EDIT: of course, 2 minutes and I'm downvoted down to -2
And even in those toy level benchmarks (like https://benchmarksgame-team.pages.debian.net/benchmarksgame/... ), there's a clear divide between C/C++/Rust & everything else.
Anyhow the whole HN conversation is now a trainwreck where any dissenting opinion is downvoted so there's no point to continue.
Rust and C++ don't have those overheads from the start, and then on top of that their compilers can do whole program optimization because they have plenty of time and resources to do so.
- D, Nim: neither provide strong memory safety guarantees. D has a safe subset, but it's not default (see https://news.ycombinator.com/item?id=12391720). Nim allows you to turn checks off at runtime. Both Nim and D generally use a GC.
- Java/Kotlin are not low-level: requires GC, and not often used for low-level systems programming. Great languages, but different niche.
Ada is primarily used for low-level systems programming and embedded systems:
https://learn.adacore.com/courses/intro-to-ada/chapters/intr...
Substantial portions of the Ada standard deal with systems programming and low-level representation on hardware:
Ada is used for low-level programming plenty.
> Nim allows you to turn checks off at runtime.
That's not relevant.
> Both Nim and D generally use a GC.
No, Nim now uses deterministic reference counting ARC and provides safety similar to Rust.
One of the extra advantages with Ada is that you can decide to utilize the SPARK variant and do formal proofs on software components that are more critical. I invite you and and others to watch the presentation by Altrans where they talk about 30+ years of SPARK usage to help develop ultra low defect software: https://youtu.be/VKPmBWTxW6M
That doesn't make you a martyr.
> Some people are posting the article around the Internet as evidence that Rust solved security.
OK, so show me. Let's see all these people declaring that this is "evidence that Rust solved security."
I think it's far more likely that you're exaggerating what folks are saying to make them appear far more unreasonable than they actually are. But we can't know that unless you actually show your work, now can we?
Yes, and in fact they write about how they are replacing C/C++ with Kotlin or Java. All three are memory safe languages. But Rust is the only low level language of the three. This is Rust's killer feature. Of course, JavaScript is memory safe and so is Lua. But they are high level languages.
> Additionally, many security issues are not due to memory safety.
I think it depends on the domain the program is in. If you are a program managing some security property, like e.g. TLS encryption, then virtually any bug you have might be a security issue. Same goes for OS kernels. But if you are e.g. an image decoder library, then there is little security impact if your output is a weirdly colored image. But memory safety issues might still be a problem for your decoder library.
Overall the number being cited is 70/30, as in, 70% of security issues are memory safety ones, and 30% are ones which are not about memory safety.
https://www.zdnet.com/article/microsoft-70-percent-of-all-se...
So you can reduce the number of security issues by 70%, isn't that great?
- https://www.infoq.com/news/2021/11/rudra-rust-safety/
- https://cve.report/vendor/rust-lang
And you can find online similar links and research on the topic.
All it takes is that you use that specific piece of code at the wrong time, and that’s it: your system which you once believed to be safe is not anymore safe. And you know what? You don’t even know what a sanitizer is, because they told you that Rust is a memory safe language and you don’t need anything else than the rust compiler - this is how I find 99% of today’s articles about Rust.
Sure, today being Rust less used than C/C++ will clearly show less CVEs.
The question is … once it’s in the wild (and by this I mean “major” libs or serious shit done in Rust) how can you prevent such bugs without a runtime sanitizer?
Beware that such CVEs can affect Rust Standard library as well, not just an unknown library.
Again, there is no way of escaping bad programming and bad practices, no matter which language you use. Someone might argue that in Python or Java you can still do bad stuff, however, the likelihood is way lower, especially because most of your Java libraries will very likely be written in Java - unless you know it’s using JNI under the hood etc.
They found 250 bugs after analyzing the entirety of rust and all of it's packages.
There are probably more memory safety bugs in the python standard library and top 50k packages.
Essentially, the Rust standard is that more or less every line of the C++ standard is a CVE.
Hell, the Rust project releases CVE for things other languages literally just shrug about e.g. https://blog.rust-lang.org/2022/01/20/cve-2022-21658.html
The C++ people just go "lol nothing in std::filesystem is safe we don't give a shit". The spec pretty much says it's UB to have other programs interact with the filesystem: http://eel.is/c++draft/fs.race.behavior#1.sentence-2
That said things that happen in the language or standard libraries like the one you linked, are often (but not always!) filed by the project.
> It could appear that these results undermine the belief that Rust safety model represents an improvement over other languages, e.g. C++, but this would not be correct, say the researchers behind Rudra, who still consider Rust safety a supreme improvement.
They found ~100 security issues in 45k packages. That's clearly better than C++.
Does Rust provide a way to check if you’re using unsafe code? What if I want to disable that? If I need to make a mission critical software I need to be aware of what I am deploying. If on the other hand we want Rust to be the new JS for backend, then yes so be it, we improved over c++, well done.
Note that the half Rust would prevent tends to be less impactful, still a CVE, but the exploit is less impactful to end users.
We did a big improvement, but why can’t we disable “unsafe”?
That would leave absolutely no margin for such errors.
Because there are things which literally can not be safe, and rust will not let you get away with pretending.
Can you post maybe a good article explaining this?
If you want to go further, you can disable unsafe in a crate by adding #[forbid(unsafe)].
And if you need more control than that, there's probably tooling out there that will help depending on what exactly you need.
https://github.com/rustsec/rustsec/tree/main/cargo-audit
Cool, so it's possible to exclude dependencies which include unsafe stuff! That's awesome.
See, this is the kind of stuff I was looking for.
From the perspective of a team writing new Rust code:
1) Don't allow unsafe (you can have an easy code search for this) 2) Forbid unsafe cargos
Finally: how do you catch unsafe in the standard library?
This can be accomplished with cargo vet (https://mozilla.github.io/cargo-vet/how-it-works.html?highli...)
I'm a c++ guy interested in Rust, my understanding is Unsafe lets me design custom high performance interfaces that do weird pointer tricks (which is needed to interface to the C and C++98 interfaces I work with), and it is on me to bounds check. Then the users of my interface don't use unsafe because I did all the nasty parts and their code is easy to write well.
If you mean "writing unsafe in your codebase yourself," this isn't borne out by the numbers.
If you mean "depending on unsafe somewhere in your dependencies" then 100% of Rust code needs unsafe, just like any other language. Interacting with hardware, many operating systems' APIs, these aren't created in a way to guarantee it, and you need to interact with them to do anything.
You… almost literally did?
> I am happy that we are moving towards a future where there are less memory bugs, but… are we really?
I don't understand this comment, to be honest. What does it add to the conversation?
I’m happy to believe GP didn’t say exactly what they intended to or simply misremembered what exactly they said previously. It’s certainly something I’ve done before in an online conversation. And when it’s happened to me I’ve appreciated having it pointed out explicitly so I could clarify what my thoughts actually were. Often it’s because I misstated my opinion without realizing.
Either that’s the case here (what I’m choosing to believe) and it gives GP an opportunity to explain further or GP is engaging in this discussion in bad faith. In either case, I don’t see the downside.
So, when I wrote the message I was really physically and mentally tired, I wrote from the phone which unfortunately doesn't always help me write 100% consistent sentences.
Finally I just used wrong wording. As I mentioned in another comment, the improvement is clear and it's the new way forward. In my opinion it simply doesn't solve all memory safety issues, and while with c and c++ you know unfortunately what you get, with Rust or Swift you might get a false sense of 100% safety where there is not.
What it does do is drastically cut down the number of places you have to deeply audit.
It is - as long as you don't use unsafe. Which is very rare, so we've made huge progress here already. Validation for cases where unsafe is necessary is needed and welcomed, but doesn't change the fact that 99% safe Rust is much better than 100% unsafe C/C++.
> Again, there is no way of escaping bad programming and bad practices, no matter which language you use.
Escape entirely maybe not. Eliminate most of it - definitely.
That's what I tried to show in the previous post: it's just not true, because if you rely even on an unsafe function from the standard library affected by a CVE, you're just fucked as if it was in C, C++ or other languages. The only difference is that you don't know what's happening under the hood and you feel "safe" because that's how they sold the language to you. Until you get hacked and "hey we didn't know...".
You just use that "unsafe zip/unzip" function you find in the docs, and with maybe with the wrong input (a weird filename or so) it happens to create undefined behavior that opens up the door to vulnerabilities.
I really believe that Rust, Swift and lots of new[er] programming languages improve in terms of memory safety (at least on the high level), however, we need to admit that they are being sold for what they are not -> memory safe. When I read memory safe it means that it cannot happen at all, and therefore I don't have to think about that kind of stuff when I write a program.
It'd be more honest to say: it's memory safe until:
- you use unsafe
- the libs you rely on use unsafe (and go figure once you start to pull a lib that pulls another lib etc.)
- the standard library functions you use use unsafe
It's an important remark, because the next generations of programmers will build programs based on a false assumption - that they don't have to worry about certain types of errors, while it's just not true.
EDIT: maybe it's just my definition of memory safe too "strict", not sure. Memory safe to me means -> it just can't happen. That's it.
I've never needed to use unsafe in my Rust applications, which means that my code is 100% memory safe (by Rust definition of course). I expect libraries I am using to provide safe abstractions over unsafe, if there is any there. Sure, there were bugs (including Rust stdlib), but such small scale effectively means that this problems disappeared comparing to languages that aren't safe and that means all I need to do is maybe update some dependency once a year. It's that rare, so I will call it safe and solved problem. We can move to next one (and there are many).
> It'd be more honest to say: it's memory safe until:
Which is what most Rust introductions I've seen will say. I don't think anyone misunderstood this part.
And this is really the only meaning for "memory safe" that can apply to anything that runs on real world hardware and operating systems. It's always conditional on the correctness of the compiler checks, the runtime, and in Rust the unsafe blocks.
An important factor here is that, because this scheme makes the trusted/human-verified part of the codebase a bit more extensible than usual, a lot of things that would typically live in the compiler or runtime get moved into libraries, just as a matter of flexibility and architecture.
So as long as people understand that these guarantees are conditional on the correctness of that trusted code, "memory safe" is a pretty reasonable name for what's going on. And I don't think people really expect to be immune to bugs in their compiler or runtime just because the language is "safe"!
I would really ask you to rethink what people think and expect. Maybe you work for a company with lots of people who understand this kind of things. I am more used to people who just jump on the "oh this lib does this, let me use it".
Everywhere I read only Rust is memory safe, you don't have to worry about anything (just don't use safe). As I mentioned before, your code might be implicitly vulnerable because of other libs, and if today it's not maybe an issue, it might be in 5 years when you have tons of libs out there rewritten in Rust and "hey they need to use unsafe code, and I can't exclude them".
For me it's important to have a proper definition and I am not happy about all the marketing around it. I still believe it provides great improvements for the code you own, so you can't mess up with pointers, use after free, and weird things like that.
The lowest common denominator of every team in the world simply cannot be our target audience for every technical term we use. There is a minimum level of background people need to learn before they can be effective, and I don't believe we can get around that just by using maximally-pedantic language everywhere all the time.
If you are actually encountering misleading Rust materials, I certainly support any efforts to clarify them, but in my experience people are actually pretty good about that already. TFA here has an entire section on `unsafe` in Rust, for instance.
It is helpful to have some way to refer to this approach to language design, and "memory safe" (like "type safe") has a long history with a precise definition. But perhaps you have some alternative "forum thread friendly" term in mind that you would prefer over "memory safe"?
I'd suggest rethinking this definition. You are always at the mercy of the lower-levels of your system, both in the runtime and the compiler. By this definition, nothing is memory-safe, since it is possible for your code in a "memory safe" language to encounter memory-safety issues in its runtime or the operating system.
Memory-safe means the bug won't happen in YOUR CODE. If your safe Rust code calls a library that uses unsafe incorrectly, yes you can encounter a memory-safety issue. But the memory safety bug is in the library, not your code. It is even possible to encounter a memory-safety issue in Rust if you never use unsafe, but in that case the memory safety bug is in the Rust compiler, not your code.
I am trying not to be extreme here, and of course I agree we objectively can't guarantee 100% safety, especially what's happening outside our code. However, importing a library happens within our code, it's in our executable, even if you don't write 90% of what's inside.
As far as I know, the JVM checks at runtime for overflows, dangling pointers, etc. So in theory there I only have to worry about using a good operating system that doesn't do funny things, because I know that the runtime has my back (at some cost, of course).
> Memory-safe means the bug won't happen in YOUR CODE.
So what happens when Rust replaces a major component in some important framework or library and your team chooses to use it because "it's written in Rust therefore it's safe to use", but actually it's vulnerable to some weird CVE because of some unsafe calls under the hood? The only way to find it is by using some sanitizer at runtime, or to exclude the dependency, but hey can you really rewrite that major library in a safe[r] manner?
It doesn't. Overflows wrap around in Java.
> dangling pointers
Only to managed memory. You may still allocate something through Unsafe and you're on your own, just as in Rust. And some popular Java libraries do use Unsafe / FFI under the hood.
No, you have to worry about your Java runtime too. The JRE has hundreds of thousands of lines of C++ and can have memory safety issues:
https://www.cvedetails.com/vulnerability-list/vendor_id-5/op...
> So what happens when Rust replaces a major component in some important framework or library ... but actually it's vulnerable to some weird CVE because of some unsafe calls under the hood?
You have a memory safety issue. Shit happens. If you are using Java, your runtime or OS can have a memory safety bug. If you are using Rust, the compiler, a library you consume, or the OS can have a memory safety bug.
The purpose of Rust is to reduce the surface area where such bugs can hide. In Rust, that surface area is the compiler and unsafe code. Well-written Rust minimizes unsafe code and audits it carefully, precisely because it is understood that this is where memory safety issues are most likely to hide.
The goal of Rust is to enable these low level components, like the JVM, to be written in a memory-safe language too.