This makes me a little worried about non-memory management security bugs. Rust is definitely helpful in eliminating one major category of bugs, but the belief that this is sufficient for users to not have to worry about security is misguided.
This makes me a little worried about non-memory management security bugs. Rust is definitely helpful in eliminating one major category of bugs, but the belief that this is sufficient for users to not have to worry about security is misguided.
For example, in GC'd/RC'd languages, if we have several UserAccount instances and a bunch of long-running operations on them, any particular long-running operation will just hold a reference to the UserAccount itself and modify it at will without any confusion.
In Rust, the borrow checker often doesn't let us hold a reference from a long-running operation in practice, so we work around it by putting all UserAccount instances into a Vec<UserAccount>, and have our long-running operations refer to it via an index. However, we might erase and re-use a spot in that Vec, meaning the index now refers to some other UserAccount.
If the operation uses that "dangling index", it can lead to leaking sensitive data to the wrong users, or data loss.
When using Rust, one has to use discipline to avoid this bug: use IDs into a hash map, or generational indices, or Rc<RefCell<T>>. Each has its own performance hit, but that hit can be worth it to prevent privacy bugs.
In the GC'd/RC'd language, this would still be a bug, but it wouldn't cause any mixups between different users' data.
I'm not saying we should always use GC'd or RC'd languages for privacy-sensitive purposes, but one should be aware of the particular shortcomings of their tools and have proper practices and oversight in place to mitigate them.
The perf hit of generational indices/arenas is minimal, and the cost of Rc<RefCell<T>> is still lower than complex GCs without JIT.
> In the GC'd/RC'd language, this would still be a bug, but it wouldn't cause any mixups between different users' data.
I've seen that exact bug in systems written in all kinds of languages.
And for what is worth, keeping a Vec<UserAccount> in memory only works on single instance services, anything beyond that and you'd have to deal with cache invalidation as well.
The most obvious/naive solution along the lines you spelled out (in rust but really for any language) would be a hash map of ids/values, otherwise Rc and maybe some weak references.
I think we can do better than saying that indexes into Vecs are "non-idiomatic" in applications... such advice could remove much of Rust's performance advantage, and make folks wonder why we're not just using GC.
Perhaps we could instead say that if one finds themselves reusing slots in a Vec, they should instead use generational indices, hash maps, or Rc<RefCell<T>>, depending on their use case and what kind of performance overhead they'd prefer.
Rust's focus on making things explicit has a tendency to make programmers want to remove every clone and allocation and make a mess of lifetimes and references in the process.
Just use Arc<Mutex<UserAccount>>. Clone freely. Box things. It makes programming so much easier, and performance will still match or exceed other languages. GC languages do the same things, but implicitly.
Except the confusion of data races and having multiple concurrent writers more generally. Not a hypothetical: I've worked in a large C# code base where other people had decided it was fine to just pass a bunch of references around to different long running processes, and sure enough, they ended up stomping over each other's assumptions in really dangerous ways.
Unless of course you're actually controlling access to the data somehow (mutex / read/write lock), in which case you can just use _exactly the same pattern_ in Rust... so this whole thing seems like a bit of a red herring.
> [...] so we work around it by putting all UserAccount instances into a Vec<UserAccount> [...]
No, "we" don't. That's one particular (bad) pattern you could choose, and I wouldn't even say it's an obvious alternative. If in your hypothetical alternative programming language you would have just kept a reference to the data (via GC or ref counting, as you said) then why not do exactly the same thing? `Arc` is a thing. It works just fine.
This sounds like a case of trying to come up with convoluted solutions to simple problems and thereby doing something unnecessarily bad that nobody made you do.
Rust certainly has its warts... but this isn't one of them. Rust doesn't make you do what you're describing, and you could equally choose to do the same bad design in your preferred GC language.
We agree it's not the best solution. And it's easy for us to say that now, after I've spelled out why. You'd be surprised how many people don't know that this can be a problem.
Also, it amuses me that having a simple index into a Vec would be a "convoluted solution". It's the easiest solution of all the alternatives. It can also be risky for privacy.
Compare that to a GC'd language, where the easiest solution (just hold a reference) doesn't introduce privacy risks.
I wonder if this error could be preventable by Rust's type system so that you can't have indexes that fall out of sync with the underlying Vec. It definitely wouldn't be backward compatible, so it would have to a new type built on Vecs. Although like another reply to your comment points out, keeping indexes is unidiomatic (in many languages) and would probably cause other bugs, which would hopefully prompt a developer to rethink it, so maybe an abstraction on top of Vec isn't needed.
Doing it with an indexed Vec is basically re-inventing your own memory management system on top of the native one, which as you point out can get very contrived and error-prone. Because then you also have to re-invent allocation/freeing, removal of holes, etc etc.
People sometimes do this in VMs/interpreters where they really do need custom/"unsafe"[1] memory management, which makes sense, but it's definitely not needed for application code like this
[1] Of course it's still memory-safe, but it's more fallible in terms of panics and bugs, as you've pointed out
I would rather advise: don't reuse indices, even if that is the simplest solution that complies with the borrow checker. When one finds themselves reusing like that, that's when to turn to other more expensive approaches such as Rc.
> that's when to turn to other more expensive approaches such as Rc
If you're saying this was done just as an optimization... all I can say is, I hope you benchmarked first. As estebank pointed out, Rc is very fast: https://news.ycombinator.com/item?id=32240478. It can even be faster than mark-and-sweep in some situations. In fact the Swift language only uses reference-counting, not mark-and-sweep, at the language level.
If you profiled and found that the Vec approach solved some performance problem you had with reference-counting, then so be it. But I would be surprised if it meaningfully helped, and shocked if it helped enough to outweigh the extra complexity.
It is now also possible to specify “no debug info” for the release build in the build config.
It also affects obfuscation and licensing in commercial products. As you can map most functions to a file based off of this handling and find code to target. It's even worse, as if you were to strip the string and replace them all with giberish, it's still possible to map each file by xref's and find related functions. With how signature are in RE tools, you can automatically detect crypto functions and then map that every other function in the file, finding licensing checks easily.
The compiler should really respect privacy, and while the ticket is ongoing, it's been ongoing for 6 years and they just made it a bug last year, with no progress. This is a huge blocker for businesses or privacy related domains which instantly makes the language a no-touch.
Using CI/CD can still give forensic indicators as the paths are still in the binary, even if they arn't your host machine path. In authoritarian countries where people have to worry about security services, this is a big threat which can get people killed. It also leaks information, even if you don't view the information as valuable, it's still fingerprinting.
Pure paranoia. Find me one instance ever where anything like this has happened.
C:\Users\zhouh\.cargo\registry\src\github.com-1ecc6299db9ec823\aho-corasick-0.7.18\src\ahocorasick.rs C:\Users\zhouh\.cargo\registry\src\github.com-1ecc6299db9ec823\aho-corasick-0.7.18\src\classes.rs
I also know what version of the library he uses which means I can pwn anyone using this if a exploit comes out. And not only this library, I know the version for every library he's using. Like Chrono-0.4.19 from this string C:\Users\zhouh\.cargo\registry\src\github.com-1ecc6299db9ec823\chrono-0.4.19\src\sys\windows.rsSystemTimeToTzSpecificLocalTime failed with:
In strings there's actually 304 unique strings with his username in it. I know every file, I know every library. There's nothing stopping a malicious actor from downloading popular rust server binaries and using strings to determine if they have an exploit. And alot of software DOESN'T upgrade their libraries from vulnerable versions, especially if unmaintained. It's a big issue.
It's a real shame that Rust ignores this. And the fact that my original comment was so lambasted against really shows the amature nature of Rust users which don't align with real world organizational goals.