Ballista: Distributed compute platform implemented in Rust using Apache Arrow
github.com
github.com
In reality, all the engineering and optimization time is behind the implementations for the google internal languages, and even the python protobuf implementation is pretty bad.
Protobuf makes some stunningly bad decisions like using varints, etc that you shouldn't make the immediate assumption "google has tons of great engineers, google uses protobuf for everything internally, therefore, protobuf is a good foundation to build my new thing on top of"
In reality, path dependence and the (amazing) internal tooling ecosystem at google both play a huge part of why they use protobuf so extensively.
(Grpc is a little overly complicated to be a universal recommendation, but I could believe it's a good choice for Arrow Flight. But it seems like they didn't do grpc + arrow or grpc + flatbuffer + arrow in the hopes that "dumb" grpc + protobuf implementations would be able to still benefit. In my opinion, grpc implementations are so coupled, there's no reason to make this unnecessary concession to protobufs)
That being said, the fact that that particular tradeoff was considered to be good for Google doesn't mean it actually is, or that it's applicable to one's application.
Most of what the CPUs at Google are doing is just copying fields from one protobuffer to another.
Most of them are doing nothing more than copying data from the database into an http stream.
(sung to this tune https://www.youtube.com/watch?v=iENQXIQ8wH0 )
Side note: It really is incredible what happened in the early days of computing when memory and computation were limited. How much care was taken in the precise layout of memory or even the timing of a calculation was insane.
Webapps are dumb middleware that pipes data from the database into an http stream - but it needs to determine which database calls to invoke and sanitize all the incoming junk.
we still need to do semantic validation, but w/ arrow, now we do that in bulk and on gpus :)
It isn't really the majority these days, is it?
It's very possible that will change over time
(My implicit assumption here is that a project like Arrow Flight wants a cross-language, widely used foundation for their protocol, and there's not a ton of things that fit that bill. But depending on your application's needs, implementing a language-specific rpc system is perfectly acceptable, and may have even better ergonomics. Rust and Python both have a plethora of mono-lingual rpc frameworks)
Here's an article making the argument for it https://blog.spaceuptech.com/posts/why-we-moved-from-grpc-to...
And much as we're considering GraphQL for some services as work... I'm not sure I buy it as an RPC framework. I suppose it has about the same appeal as SOAP for that purpose.
Flight is interesting to us bc they're thinking through parallel i/o. I'm guessing grpc overhead there is more tolerable, though yeah, I'd be curious. I had similar initial reservations, esp as we were happy to remove slow google protobuf stack stuff as part of our switch to arrow :)
The use of Arrow to support multiple programming languages is also a great concept. Other distributed computing engines have ended up tied to the JVM (Spark, Presto, Kafka) as a way of avoiding serialization/deserialization costs when you go across a language boundary. Arrow is a really elegant solution, as long as you're willing to batch up operations.
https://blogs.oracle.com/javamagazine/understanding-the-jdks...
One can do both.
I would love to have something more resource efficient than Spark on JVM, but Delta Engine isn't there yet.
Probably got cut because of maximum title length but important nonetheless.
And even without value types, there are off heap allocations, also if "Python" libraries can be actually written in C, so can Java ones, without loosing the plus of the ecosystem.
While the database engine of most RDMS servers is written in C and C++, anc it won't change given the history behind the code, the bulk of the code is written in some form of managed SQL, with IBM, Oracle and Microsoft also allowing for Java and .NET code.
Finally, https://www.efinancialcareers.co.uk/news/2020/11/low-latency...
Once upon a time, anyone knew that high performance systems were naturally only viable in Assembly unless proven otherwise.
Or how C++ wasn't ever to be a thing in game development, C was the king of console SDKs, after years trying to take Assembly's place on 16 bit platforms.
> Essentially, we use a contrived form of Java that avoids all the Java constructs that make things go slow. We only use the constructs that are fast and efficient, and we avoid all the garbage
> The only problem with low latency Java is that most experience Java programmers struggle with the new paradigm. "A lot of people who program in Java are used to working in an environment where latency isn't a criteria," says Lawrey.
So, the best Java developers would be former C/C++ developers. That's hardly a ringing endorsement for the language. Look at LMAX's Disruptor, for example. It's hardly Java since it gets its performance from use of sun.misc.UnSafe.
Java does give you quite good IDEs though. That's about it.
So, the best C developers would be former Assembly developers. That's hardly a ringing endorsement for the language. Look at XYZ game, for example. It's hardly C since it gets its performance from use of inline Assembly.
C does give you quite good shell utilities though. That's about it.
C++ is never going to be safe by default, unless ISO is willing to do a Python 3.
Rust still needs to improve its usability story against everything that is available on the JVM, and compile times, oh boy.
While such projects are welcomed, when value types finally land, the argumentation would need to be upgraded.
And as someone that has to track down memory issues in C and C++ projects, not having a GC doesn't mean the memory usage is deterministic.
I'm curious to see how this evolves as there are a number of motivated folks working on similar efforts such as Vega. I for one would welcome a mature rust based distributed compute platform.
> With the release of Apache Arrow 3.0.0 there are many breaking changes in the Rust implementation and as a result it has been necessary to comment out much of the code in this repository and gradually get each Rust module working again with the 3.0.0 release.
This appears to be an issue with Arrows implementation hitting a new major version and the Rust libraries not yet being compatible with the newest versions of Arrow. That's not something specific to the Rust ecosystem. It's not like a new version of Rust broke this project.
But even if it had, maintenance is always hard and the health of a project is better measured by how long it takes to be working with new, major, stable versions after widespread community adoption of those new, major, stable versions.
I don't know if Arrow 3.0 is the most commonly used implementation-- it may not have even reached that milestone.