Improving Interoperability Between Rust and C++
security.googleblog.com
security.googleblog.com
What will the money actually fund? What are the conditions and KPIs of the grant? Are they putting their weight behind any preferred approach?
Edit: The RF announcement includes more information: https://foundation.rust-lang.org/news/google-contributes-1m-...
"The Rust Foundation’s first Interop Initiative task will be to draft a scope of work proposal for discussion amongst our team members, the Rust Project Leadership Council, Rust Project stakeholders, relevant Rust Foundation member organizations, and its board.
[...] Recommendations will likely include the hiring of one or more Interop Initiative engineers and may include the provision of resources towards expanding on existing interoperability work, build system integration, using AI for C++ to Rust conversion, or some combination of all these. The Foundation will engage the appropriate stakeholders across the Rust Project and its member base to review the proposal and carve the path forward for this important work."
$1 mil is less than a rounding error for an organization like that. I'd love to know how much they've burned on the numerous projects that occupy their product graveyard, and imagine a future where that was devoted to the growth and education surrounding core language development in every language - Rust, Go, C++, Python, etc.
Spending money requires a certain amount of structure and process, and it takes a while to adapt to a money windfall. Going from a little money to a lot changes an org very significantly.
Granted, I also haven't looked much into the Rust Foundation, but I'm going to assume a $1 mio is significant for them. Google can always donate more down the road when the investment pays off and is digested usefully.
(Edit: The Rust Foundation had an income of $2.89 mio in 2022 and $2.56 mio in 2023 according to their annual reports. $1 mio sounds about right for what they should be able to usefully handle, from the gut.)
I'm not affiliated with Google in any way.
I think it's great that Google has been embracing Rust, but this feels hyperbolic to me.
A quick search shows they spent almost $40 billion last year on development expenses. I'd like to imagine the alternate reality where even just $100mil of that was spread out across various core language organizations.
Yes, these are important projects. It is a big deal. But there’s way more to Google than these two things.
That said, the OP has clarified that they didn't mean that Rust was this thing, but that language development in general is.
That's 4 $200k/yr engineers working on this full time and then 1 $200k/yr manager to keep them on track.
There is a group of highly intelligent people who want to go into lucrative careers that won't pick CS in the EU, and they do in the US because of the pay.
I think it's pretty apparent that, while there is fantastic talent in the EU, the density isn't as high as it in in the US.
Back 5-10 years ago Google tended to calculate opportunities in SWE-years. Roughly $300K/engineer.
So I'm asking for 5 engineers for 3 years.
I'm a bit disappointed that cxx gets all the glory and nobody likes the approach of the cpp crate.
If you can replace
void qsort(void base[.size * .nmemb], size_t nmemb, size_t size,
int (*compar)(const void [.size], const void [.size]))
with extern "C" fn qsort<erase T>(base: *const T, nmemb: usize, compar: fn(*const T, *const T) -> std::ffi::c_int)
such that `qsort` is still single a regular C function, not something that corresponds to many separate C functions, many good things are possible.In particular for C++, being able to do
struct MyVTable<erase T> {
virt_method_0: fn(*const T) -> bool,
virt_method_2: fn(*const T, usize) -> (usize, usize),
};
struct MyType<erase R> {
vtable: *const MyVTable<MyType<R>>,
rest_of_fields: R,
};
Opens a lot of doors. People while whine because it is exposing the C++ ABI in the Rust API, but that's fine. That's what `private` is there to fix. pub unsafe fn qsort<E, C: FnMut(&E, &E) -> core::cmp::Ordering>(elements: &mut [E], mut comp: C)
In rust as currently stands: https://play.rust-lang.org/?version=stable&mode=debug&editio...On the other hand, both this wrapper and yours are counterproductive if the element size is dynamic (e.g. perhaps you're dealing with some nonsense like:)
struct ITableColumn {
virtual ~ITableColumn() {}
virtual void* base() = 0;
virtual size_t stride() const = 0;
virtual int (*)(void*, void*) get_comparison_func() const = 0;
}; extern "C" { fn qsort(ptr: *mut c_void, count: size_t, size: size_t, comp: extern "C" fn(*const c_void, *const c_void) -> c_int); }
And the call: unsafe { qsort(ptr, len, size_of::<E>(), comp_wrapper::<E, C>) };
Are both hidden away in the body of the wrapper fn.Your abstract declaration still wouldn't be safe. Remaining unsafety includes:
• possible dangling pointers
• possible incorrect lengths
• possible libc bugs with ZST elements
• undefined behavior if the sort fn misbehaves - (recently reported as a security issue against glibc because that being UB is dumb even if allowed by the standard: https://news.ycombinator.com/item?id=39264396 )
This is why I call it a "half measure".
I also fail to see how your proposed abstract declaration would simplify either my wrapper, or other code that would actually bother to use the raw FFI definition in any significant way. This is why I further call it "unnecessary". It also fails to specify which underlying FFI parameter size_of::<T>() would actually be passed into.
> You are writing some unsafe code with safe wrapper.
My wrapper remains `unsafe` as well, but it ameliorates everything it reasonably can.
> These are not the same.
No, but my wrapper demonstrates an actual use case of your raw FFI definition... and shows the actual concerns of surrounding code that aren't significantly helped by your declaration.
The second example with the vtables is the point (i.e. complicated data structures). qsort is just a simple example to introduce the concept.
https://microsoft.github.io/windows-docs-rs/doc/windows/Win3...
Which can then be converted to a refcounted smart pointer:
https://microsoft.github.io/windows-docs-rs/doc/windows/Win3...
All driven by win32 sdk parsing and metadata.
But suppose we want to roll our own, because we tend to prefer `winapi` but it lacks definition. That's not too terrible either:
• https://github.com/MaulingMonkey/thindx-xaudio2/blob/master/...
• https://github.com/MaulingMonkey/thindx-xaudio2/blob/master/...
• https://github.com/MaulingMonkey/thindx-xaudio2/blob/master/...
I could more heavily lean on my macros ala `windows`, but I went the route of manual control for better doc comments, more explicit control of thread safety traits to match the existing C++ codebase, etc.
Is there some pointer casting? Yes. Is it annoying or likely to be what breaks? No. What is annoying?
• Stacked borrows and narrowing spatial provenance ( https://github.com/retep998/winapi-rs/issues/1025 - this can be "solved" by sticking to pointers ala `windows`, or by choosing a different provenance model like rustc might be doing?)
• Guarding against panics unwinding over an FFI boundary. This is at least being worked on, but remains unfinished ( https://rust-lang.github.io/rfcs/2945-c-unwind-abi.html )
• Edge case ABI weirdness specific to C++ methods ( https://devblogs.microsoft.com/oldnewthing/20220113-00/?p=10... , https://github.com/retep998/winapi-rs/issues/523 )
Rust/WinRT has plenty of old open tickets into that regard, and given the team's track record on how they messed the developer experience with C++/WinRT and then gave up on it, I am not too keen investing into Rust/WinRT.
Let's just agree to disagree on this.
Without trait bounds there isn't much you can do with it anyway besides generic code like this. I guess the question is 'does Rust emit' or elide an unused generic function with unused interior types?
The other major use case is if you've got something like a Vec<T> where several of the helper methods can be erasable. But it's already possible to do this with helper functions--the standard library takes advantage of it--so it's not entirely clear that you need a language-level feature for this use case.
Would it be a stack allocated, owning version of that, that can be moved? A pure `dyn Trait`?
struct DynFoo<erase T> {
vtable: &mut VTableFoo<T>
ptr: &mut T
}
This is the underlying feature. `dyn` is a bit of an ad-hoc hack around not having this feature.erased: just ignore the type variables for compilation. This is like the typescript -> javascript compilation. (Also how Java, Scala, Haskell, OCaml, etc. type paramters work)
Existential: "there exists a type such that..." see https://en.wikipedia.org/wiki/Existential_quantification https://en.wikipedia.org/wiki/System_F
C++ stores references to dictionaries from/with the things they describe, that's "existential"
Haskell passes in extra parmeters that are dictionaries and doesn't (except for user-written code that does) keep around extra references to them. That's more like "universal" than "existential".
Yes, that's true, but it's not type safe, and I don't really want to program that way.
The point of erased types is to be able to put in the language invariants that hitherto are only in our heads, but unless you have a typed-based mental model for what makes vtables/dictionaries safe (say, based on experience in other languages) you won't have ideas you can't write down, and thus you won't know what is missing.
I wish there was a good blog-sized resource I could point you to that did explain it instead, but I don't know of any off-hand.
I guess, try to wrote `dyn Trait` things (like a heterogenous collection) without `dyn Trait`, and then when it doesn't work without unsafe Rust, you will see what I mean.
Yes, that's true, but it's not type safe, and I don't really want to program that way.
We had re-implemented a few mathematical and curve finding functions and even with the newest python they were uncomfortably slow.
PyO3 made it so easy to implement a runtime switchable rust library exposed to python, it was almost unnerving. By writing those handful of functions in rust we got something like a 30x speedup.
Note this was a small project with limited funding for r&d so the level of effort for performance speedup was really nice.
I recently did use it for a project, and it was quite a lot of time and pain guessing what would work and what wouldn't.
Hopefully, the language will develop further and then reach ISO standardization - which Java cannot as it remains proprietary.
The JDK has a bunch of garbage in the JDK that people shouldn't use. That stuff has to remain due to strong backwards compatibility guarantees.
Rust's approach means that non-standard stuff that we end up realizing is a mistake can quietly die off.
Now, should it be larger? Probably. I'd prefer if rustlang had a standard async/await implementation rather than leaving it up to the ecosystem. But I don't think rust needs, for example, a gui api (like swing or awt) in the standard library.
I'd prefer if rust was managed a bit more closely to the way java is managed today. Java doesn't pull in new apis willy nilly, but the ones they do pull in end up being things that have broad appeal and utility. Rust could take a look at common crates in the ecosystem and start pulling those in to the standard lib.
The "default" is tokio, and I think even the tokio authors would agree that it's neither suitable nor ready to be part of the stdlib.
OTOH, pollster[1] (or something like it) should be pulled into the standard library. Not particularly useful for anybody who wants async, but super useful for anybody who doesn't want async but wants to use a 3rd party crate that includes async.
Surely you can see how one is worse than the other?
What you are describing is a problem, no doubt about it, but it's not as if this isn't something that can't come up in javaland. For example, dealing with a project that uses both Gson and Jackson. Or dealing with a project that's mixed together Netty with Apache http.
And when you can ignore these dead libs very much depends on what the other parts of the lib you are dealing with. For example, you might never interact with `Enumeration` or `Vector` yet there are parts of the Swing api that expose those.
That is not actually a problem. Let it stay there, it hurts nobody and may actually occasionally help people.
The massive feature set of the JDK bloats the size of every container shipping with the JVM. It pumps up the requirements for metaspace. And it negatively impacts JVM startup time.
I'm not saying the JVM needs to remove every dated API. But things like JNDI, for example, are not only dangerous to use (that was the root cause of log4j2's big vulnerability), they are massive feature sets that pretty much nobody wants to use.
Java was designed in an era where we thought having thinish clients streaming jars/classes from central network servers was probably a good idea. When we thought the JVM could be more than just a language VM, it could be an operating system. A lot of those concepts simply don't apply to modern jvm dev or even jvm dev that's happened in the last 20 years.
And, to be clear, the JDK was not wrong in bringing along all these libraries. After all, in the 90s it's not like we really had a great story around community package development. That was an era where devs routinely downloaded their dependencies manually. Clever devs even had curl scripts wired into ant or make to do that job.
I have been thinking about this for the past few years, and I think I'd go so far as to make the standard library "not special." That is, make it work similar to the edition system: cargo new would add a dependency for the latest std at the time, build-std would just be transparent, and you treat it like any other crate.
That said, I am sure there are a zillion issues with this today, especially in Rust itself, but if I were to make a Rust++, this is one of the things I'd consider trying to give a shot.
If you start versioning the stdlib now you need a bridge between the Option from std1.0 and the Option from std2.0. This will become confusing quickly for programs that use libraries that depend on different versions of the std.
We could forbid that situation, but then we have an ecosystem split
Many (all?) languages were successful before being fully specified, and many still are not (cough Python cough) yet that has not impacted their adoption, popularity, or effectiveness.
1: https://ferrous-systems.com/blog/officially-qualified-ferroc...
The first reason is because the stdlib is the only thing you can always count on being there. People are often in situations where they can't download library packages due to security procedures, and have to rely on just the stdlib. People like to complain about urllib2 still being in Python even though it's not really used any more, but I've been in situations where urllib2 was the only thing available and I was damn glad it existed.
The second reason is because the way Rust does things is horribly confusing. What's the best crate to use for X? If you are a regular in the community you probably know, but a newbie is going to have no idea which of the many available options to pick. Whereas something within the stdlib is always a reasonable choice, even if it isn't the best choice.
I really hate that the Rust community in general is so dogmatic about this topic. Having so much functionality outside the stdlib makes the language worse, not better.
MAYBE a Ruby-style "standard dump," but I don't see why such a thing would ever be useful.
Rust is yet to have at least two fully working implementations, and language specification is ongoing.
Compare through how many breaking changes even high-quality ecosystems crates have gone through in the last few years.
Of course, in a company of Google’s size, I’m sure different groups are pushing all reasonable options.
Is it, on average, safer?
https://cxx.rs/binding/cxxstring.html
It's because Rust doesn't have move constructors. I doubt they're going to add move constructors to Rust so that seems kind of unsolvable which is a shame.