Rewriting a high performance vector database in Rust
pinecone.io
pinecone.io
This is a common experience and I'm still surprised by the choice I constantly see to use a managed-memory languages to build a database - one of a very small set of special cases where having full control over the memory layout might just be a reasonable thing to want. In this universe (absent doing something completely absurd) it's not algorithmic complexity but managing data locality in the cache hierarchy (e.g. reading things from L3 vs main memory vs disk) that makes things orders of magnitude faster, especially if you're in the realm of doing things like SIMD operations to speed things up.
Perhaps there's some level of suck we're willing to tolerate for all the other benefits you get, but I've been noticing a pattern of "align things just so at the higher level and hope they mostly turn out the way you want at the lower level" (e.g. also with the Apache java-y databases like hadoop / hbase / cassandra which I guess were mostly supposed to derive their total throughput from massive scale rather than per-node performance) which is a bit funny.
But also it seems like part of Rust's promise was "low level but make it high level" which seems to be succeeding (zero-cost abstractions and whatnot), so I imagine this will get better over time - having not attempted a project like this myself, I'm not sure what the limitations you'd run up against are in terms of laying things out in memory in a favorable way - I imagine the kind of massive manually managed arena allocations and ad-hoc pointers going everywhere that one normally does doesn't really fly.
D, Nim, C#, Swift, not to count all of those that existed since Xerox PARC days.
I don't think garbage collection is in the top 3 causes of why Python is slow.
The reason for slowness - is the weak dynamic type system of Python.
Every single instruction need to be type checked at runtime and thus making everything slow.
Compare to C#/Java which have GC but both are amazingly fast, because these languages have stricter type system. If you add JIT on top (which can selectively replace MSIL/java opcodes with native machine instructions) and it makes perf on par with natively compiled languages like C++.
Python is structurally typed, that makes it dynamic, but it is not weak as there is no type coercion.
> Every single instruction need to be type checked at runtime and thus making everything slow.
This is also wrong, python does not type check anything, not in the "regular" manner of typechecking. It relies on structural typing, if it quacks like a duck, then it is treated like a duck.
In fact, PyPy is an argument against your position as it still allows the same (more or less) behaviour that python has while operating a lot faster due to JIT.
Python doesn't have the luxury of compiling that C# and Java, nor is the VM intended to be high performing.
Fixed it.
So, they took some DBMS (which is probably not so hot in terms of performance) and rewrote it in Rust. Surely possible, possibly useful, but not much to write home about if one is interested in DBMS performance.
What stood out for me in the article is him saying that it's difficult getting developers with experience in both Python and C++.
So, I wonder, if his in-house devs could pick up Rust that they previously couldn't write, why does he think he can not hire a good programmer and charge him to learn the stack the company uses. Why must they employ someone that already writes Python or C++.
Is Rust such a straight-forward language that people new to the language can write a very performant programme
This is seen as an important thing to improve because of the slogan "A language empowering everyone to build reliable and efficient software". If Rust isn't empowering everyone because say, 10% of people who try to pick it up can't get anywhere that's not good enough.
This is because inexperienced Rust programmers are relatively harmless. Noob mistakes won't compile, rather than running into dangerous gotchas. You can tell noobs not to use `unsafe` (and there are ways to enforce that), and mostly they'll just write inefficient or non-idiomatic code, but the code will be free from data races and memory corruption.
The strictness of the Rust compiler is quite the opposite of something like the C++ Core Guidelines where the majority of the rules aren't enforced by the compiler, and have to be in the programmers' head first.
Noobs make lifetime errors and fight the Rust compiler, but imagine working with a compiler that doesn't tell you when you have lifetime errors.
You seem to paint the landscape as full of tools and imply that they're used. Either they're insufficient or they're often under utilized, simply due to the number of bugs we see. No?
Static analysis tools have much harder job analyzing C++ (aliasing and escape analysis are way harder, and static analysis of thread-safety is basically impossible due to lack of thread-safety info in the type system). The results are a trade-off between being sparse or having false positives.
The sanitizers only catch issues they can observe at run time, and that relies on having sufficient test and fuzz coverage. Some data races are incredibly hard to reproduce, and might depend on a timing difference that won't happen in your test harness.
OTOH Rust proves absence of these issues by construction, at compile time.
It's like a difference between dynamically-typed and statically-typed languages. Sure, you can fuzz type errors out of JS or Python, but in statically-typed languages such errors are eliminated entirely at compile time. Rust extends this experience to more classes of errors.
Rust just takes the other side of the trade-off, and will reject valid programs. Hence why the unsafe keyword exists, and why tools like Miri (https://github.com/rust-lang/miri) exist specifically for rust.
Are we still talking about ease of add novice Rust programmers to a project?
It's obvious that Correct programs compile, Not Correct programs result in a compiler error, but what do we do with shrug emoji ? Rust says those go in the "Not Correct" pile and get a compiler error. C++ says - really, I'm not making this up - that they go in the Correct pile.
There's an immediate short term consequence, but also, after decades, a long term consequence that's arguably worse. Short term, C++ programmers can't know if their non-trivial program is Correct. It might be complete horse shit, the compiler won't necessarily tell them.
Long term, C++ grows worse and worse because there is no reason to shrink that shrug emoji category. All the programs in that category act as though they're fine, so there's no pressure whatsoever to improve the language standard, the compilers, etc.
This has always been such a weird claim to me because it's not clear to me what's meant by "valid". My instinct is that it would be hard to define valid in a way without making most languages either accept some invalid programs or reject some valid ones (the exception being defining "valid" as "anything that language X accepts", but that wouldn't really say anything interesting about Rust). Sure, you could draw the line so that Rust is the only one that rejects valid programs, but is it worse to be the only language that rejects some valid programs than to be a language that accepts invalid ones? The alternative is that all languages accept some valid programs, at which point there's nothing specific about Rust that's worse until you start specifying which valid programs each language rejects. If there's a way to define "valid" so that Rust is the only one that rejects valid programs but other languages accept all valid and reject all invalid programs, I think it's non-obvious enough that it should be explicitly stated.
Here is an example where a variable is definitely initialized, but it will be rejected.
fn main() {
let foo;
if true {
foo = 1;
}
println!("{}", foo);
}At the faang I worked at, some small portion of servers ran the sanitizers in prod, so you’re not reliant on test coverage nearly so much for catching rare issues.
Your milage may vary, but I think Rust offers a well considered step into the right direction. Stupid and dangerous code of the kind every developer will produce once in a while just won't compile in Rust. You cannot forget to run a check, you can't hide behind not knowing a tool. You can't ignore the warning of you want a running program.
That is not nothing.
It's not feasible to have a developer compile for 3 different platforms in two configurations with and without sanitizers, and core guidelines checkers for every change, and these tools take so long it's a huge cost.
I recently spent some time fixing performance issues in some novice Rust code, and while the code was pretty clearly written by someone new to systems programming it still all worked fine - https://jackson.dev/post/rust-coreutils-dd/
UNSAFE {
// TODO: Verify all the lines, all the time, are ok
// Just like you do testing, documentation, security and all that
// ok?
#include <iostream>
using namespace std;
int main() {
// YOUR CODE
}
}> What are the main correctness risks in C++ if you just never use a raw pointer?
All the code on C/C++ IS a correctness "risks". Only constant, manual inspection could(maybe) say otherwise.
What Rust gives is significant reduction of the risks.
std::string s(s);
To be fair, compilers will warn you about this nowadays. But when I converted a C codebase to C++ 20 years ago they didn't.
IIRC, references can also refer to de-allocated memory. Also, if you don't pass-by-reference or pointer, you can literally "slice" the dynamic doohickies off your instance so your AlbinoCat behaves like a Cat because all that extra special stuff is gone as far as the function is concerned.
This is just off the top of my head after not working with C++ for 20 years. I'm sure with all the new features it's gained over the past 20 years theres whole new exciting ways to blow your leg off.
And even then, my understanding is that raw pointers are still intended to be used in Modern C++: they're there for when you don't want to transfer ownership.
- lifetimes are hard. The core guidelines recommend using span and string view, but correct usage of those isnt straightforward:
std::string_view make(const std::string& in)
{
return in;
}
Is not safe, at all. Span is also super dangerous: std::vector<int> vec = {1, 2};
std::span<int> sp = vec;
vec.push-back(1);
Lambda capture semantics are another area where you might end up surprised by the behaviour too, (and by surprised, I mean you'll have memory issues).Then there's all the normal stuff that still exists like slicing, lack of bounds checking, resource leaks because of improper inheritance use, data races, use-after-free issues. Sure these can be caught by a sanitizer with enough effort, but they still exist.
if you use shared ptr for this:
1. shared ptrs aren't made for this use case, they're for shared ownership 2. why using C++ at all if it means reducing its performance to an (atomatocally) performance counted language? Rust allows for much more performance with its safe borrowed references.
That also not what the CPP core guidelines, that are supposed to define modern C++, prescribe: For general use, take T* or T& arguments rather than smart pointers[0]. This rule incurs the risk of use after free if the returned reference is kept for too long by the caller. The rules for reference validity are non trivial (eg you returned a reference to an item of a vector, if you push into your vector you might invalidate the reference), so this is a significant source of bugs even in modern C++.
if you don't use shared ptr or references/pointers for this i'm curious as to which mechanism you're using to observe non owning data
[0]: https://isocpp.github.io/CppCoreGuidelines/CppCoreGuidelines...
- C++ is an unusually large language. And it has many historic footguns, requiring a higher level of vigilance and code review. If I were starting a brand new project today, I wouldn't try to build a team of C++ programmers.
- Untyped Python becomes more difficult to refactor and maintain once you reach 50k to 80k lines on a group project. Typed Python, however, scales nicely beyond this size.
- Rust is a "medium-sized" language. It requires developers to learn more than Go or Python does, but less than C++. And Rust has far fewer traps for the unwary and the reckless than C++. Rust's tooling is also very good in many areas.
- It's tempting to split a project into a fast "core" language, and high-level "glue" language. There are real advantages to this. (Which is why I've done it on one recent project!) But this also comes with costs: everyone needs to be fairly good at two languages, and switch back and forth. And you pay a tax at the boundary.
If I were building a brand new database (and a team to maintain it), I'd actually be strongly tempted to use Rust exclusively. But this is partly because databases rarely have a "business logic" layer that changes constantly, so there's less need for a high level scripting language.
But with a different team or different constraints, C++ could also be the right choice.
Why does everyone needs to be good at both languages? You can seperate and have the core people writing efficient low level code - and you have higher level scripting/gluing code.
You will get hot loops in Python, because a Rust programmer wasn't around, and Rust programmers implementing whole complex business logic in Rust behind a single `do_it()` Python call.
If it is doing things right, then the rust people don't do complex buisness logic, because it is not assigned to them and they would not even have the details.
And if the python people were too eager and have core stuff implemented and it is affecting performance, than you can always reimplement it low level.
It all depends on the project of course, of what would be the best mix.
The part of Python (usually) is that you don't need to be "good at it" it you aren't trying to write super polymorphic core that runs super efficient computations like scipy. If you have a fast core engine for the innner loop, a slow Python management layer is plenty fast.
Could you say more about what tools and practices make this possible, beyond simply adding type annotations in your code? Asking for a friend.
> It's tempting to split a project into a fast "core" language, and high-level "glue" language.
I did this with C++ and Boost Python back in the day and loved the experience. I wonder if Rust will someday get a high-level language for writing applications and scripts on top of a Rust codebase, like Boost Python for C++ or Tcl for C.
Python with type annotations works really well with type checkers like Mypy, along with LSP servers, and both of those integrate with most development environments.
Using a Python-oriented IDE like Pycharm with type annotated Python also allows for better refactoring options. It reduces the uncertainty and guesswork an IDE's static analyzer must engage in for even basic features you'd take for granted with IDE and statically typed languages.
In practice, developers don't have to keep what can be a massively complex application running in their heads to modify code accurately. A nicely typed project makes it easy to see exactly what types of data get passed around and modified. Before gradual typing, you'd have to backtrack to all of a function's call sites to understand exactly what kind of data it takes and returns. With gradual typing, you can just look at types and rely on Mypy to ensure the right data is passed around.
> I did this with C++ and Boost Python back in the day and loved the experience. I wonder if Rust will someday get a high-level language for writing applications and scripts on top of a Rust codebase, like Boost Python for C++ or Tcl for C.
I haven't used Boost Python, but there are some options for Rust and Python that work well and seem to suit this use case like PyO3.
Only software that does not have to run correctly (prototypes, personal hobby projects) can get away with a non-static type system.
When I have to pick a tool and I see it is written in Python I will have a look for an alternative if possible. Because I know it will have many bugs: some known, lots hidden.
More like they had difficulty finding cheap experienced c++/python devs.
Yes. At least that was my experience coming to Rust from a primarily JavaScript background. My first Rust program wasn't as fast as it could have been (I know a lot more about optimising programs now than I did then), but it was still ~30x faster than the JavaScript and Pythons versions my company had attempted to write first. That's without putting all that much effort into optimising it.
You might need to define what is "very performant".
I come from C++ performance work side in games. Typical C++ is ~10 times slower than optimized C++. And optimized C++ is sometimes ~2-10 times slower than what game hardware can do. What limits many projects is that cost of going to 'next level' of performance is also nonlinear. Where is on this scale Rust code written by new to the language?
I could see huge benefits from Kafka and Cassandra being re-written in Rust.
See also, their install dependencies script.
https://github.com/redpanda-data/redpanda/blob/dev/install-d...
I'm not necessarily advocating it [1] but the parent's claim was that those programs could benefit from memory safety, thread safety and better concurrency and C++ does not deliver along that axis.
[1] https://www.joelonsoftware.com/2000/04/06/things-you-should-...
Prior systems I had built with facebook folly (c++ lib) and had also written my own eventing systems in the past, but the real value is having seastar being battle tested since 2016. Largely it has been the right decision for us as redpanda for it's young age has benefited from the stability of seastar.
Rust rewrite of cassandra, if it could reach feature parity with 4.0, would be a good thing. (C++ port already sortof exists with ScyllaDB but they haven't reached feature parity yet)
We've been past the "it's the shiny new thing" phase for a while.
https://www.pinecone.io/learn/vector-database/
...was less than informative.
ABCABCABCABC
Vector databases store them like this: AAAABBBBCCCC
This allows faster queries if you just need one (or a few) columns, because unrelated columns don’t have to be processed at all. Caches are more efficient, vector CPU instructions can be used, etc…The downside is that random single row access is more expensive because a row has to be reassembled from many locations.
Row store: “table id; row id; column id”
Column store: “table id; column id; row id”
The difference is the order used to sort the data in storage.In reality, most column store systems use a more complex partitioning scheme with row groups and the like…
Example:
Let's suppose I've downloaded all the Gutemberg library books.
I can feed a transformer like Bert or GPT-3 to calculate the embeddings of these books.
These embeddings will represent in vector form (an array of numbers of fixed size) the meaning of these books.
I can save these vectors in this database and this database can then calculate the distance between these vectors, so basically how closely related they are in terms of topic.
If I query this database with the embedding of a sentence like "Love story between teenagers from 2 enemy families in Italy, they die at the end", hopefully the best result will be Romeo & Juliette.
I'm no ML person so take my comment with a grain of salt in terms of how well it works. But in theory, that's the goal.
Wonder what the commit was that caused a more than 2x regression and got a fix instead of an undo.
So you decided on a language that makes it even harder to find experienced developers?
It is a lot more similar to JS than C++ is.
That is strange as I have experienced the opposite. I've written all three languages and I've noticed that JS patterns don't translate well to Rust. Many C++ patterns translate well to Rust (albeit after a bit of borrow checker fighting).
Thoughts?
How many epochs will exist in 30 years?
If you look at what [changes] were introduced by the [2018] and [2021] editions, they weren't as earth shattering as some might think:
2018:
- Module system changes
- Mandatory associated fn argument names
- dyn, async, await and try are now keywords
- You can't write `let s = libc::getenv(k.as_ptr()) as const _;` anymore, instead needing `let s = libc::getenv(k.as_ptr()) as const libc::c_char;` (Method dispatch for raw pointers to inference variables)
2021:
- TryInto, TryFrom and IntoIterator added to the prelude
- cargo dependency resolver changes
- [1, 2, 3].into_iterator() now works
- `|| a.x + 1` now captures `a.x`, not `a`
- Small technically backwards incompatible change to the panic macro formatting string
- any_identifier#, any_identifier"...", and any_identifier'...' are now reserved syntax
- Some previously existing warnings are now errors
- a | b is now matched in pat macro rules
[changes]: https://doc.rust-lang.org/edition-guide
[2018]: https://doc.rust-lang.org/edition-guide/rust-2018/index.html
[2021]: https://doc.rust-lang.org/edition-guide/rust-2021/index.html
Edit: it just dawned on me that by "epoch" you might not have meant the edition mechanism, which was at some point referred to as epochs and have so far happened every three years, but rather "how many iterations of idiomatic Rust code will there be in 30 years". If that was the original intent, you would also consider things like the introduction of the ? operator, or the upcoming let pat = expr else {}, or match ergonomics, or the likely deref patterns, or any number of features that on isolation might not be huge, but that can materially impact what idiomatic code looks like. I personally believe that that kind of iteration and evolution of a language is good and necessary. As long as forwards compatibility is maintained, and that forward compat doesn't hinder the future design space, making things better over time is a great thing!
This isn't unique to C++ as people like to criticize, compare C# 11, Java 19, Python 3.10, C23... with everything in between down to their initial versions.
The first two are also strengths of C++, and for the third the article says that "Rust is async, and Tokio is the one of the most popular async providers ... However, it’s not great for running CPU intensive workloads, like with Pinecone." Puzzling.
We looked at and compared several languages - Go, Java, C++, and Rust. We knew that C++ was harder to scale and maintain high quality as you build a dev team; that Java doesn’t provide the flexibility and systems programming language we needed; and that Go is also a garbage collected language. This left us with Rust. With Rust, the pros around performance, memory management, and ease of use outweighed the cons of it not yet being a very established language.
In other words, they wanted to unify the programming languages and evaluated several. Rust won out of those for performance reasons.The article is a short recap of a 40 minute video. The video has more context and explains the intentions much better than the web page.
They show a graph of performance over time as the rewrite progressed. There were some small optimisations and problems, a few big regressions, and then a huge improvement that was maintained. Looks like the rewrite process made the database perform significantly better. There's nothing on how much this was caused by the language switch itself, but that's functionally impossible: nobody is rewriting their application twice to see what rewrite is better.
Agreed, and hypothetically the 2nd rewrite should still be better than the first. So the language would have to make it significantly worse to outweigh the yet again experience in improving things.
To be clear though i'm not stating that every rewrite is assured to be better. However a carefully considered rewrite has a much easier time making decisions learned from any warts discovered in previous implementations. God knows there's always some warts.
As a Rust fanatic, i wouldn't expect Rust itself to be due to the performance gains. It's not expected to be faster than C/C++ typically. Just comparable.
Makes me think the eng lead just wanted to do Rust, and made up a rationalization.
As someone who was heavy in C++14/17 it's much nicer in Go/Rust-land.
As much as I love Bjarne, I'm not coming back.
But yes, it is a given that C++/Rust is faster than Python unless there is some fundamental algorithmic foolishness done by the C++/Rust programmer. People with industry experience know why that is a given.
Before Rust became viable, rewrites were done in Go.
From the archives:
- Rewriting a large production system in Go https://news.ycombinator.com/item?id=6234736 (2013)
- How We Moved Our API From Ruby to Go https://news.ycombinator.com/item?id=9693743 (2015)
- Matrix and Riot Confirmed as the Basis for France’s Secure Instant Messenger App https://news.ycombinator.com/item?id=16938545 (2018)
- Toward Vagrant 3.0 https://news.ycombinator.com/item?id=27476676 (2021)
- I’m porting the TypeScript type checker tsc to Go https://news.ycombinator.com/item?id=30074414 (2022)
There's a reason why folks take the time to rewrite things in Rust. No matter how good you are at C/C++ you will encounter bugs that you would not have if you had written it in Rust.
People keep forgetting C++ has 30 years of being deployed in production.
Rust is 2022 is like using C++ in 1990's in terms of ecosystem.
Nice as a language consumer, but a bummer that building new languages and getting some attention is so much harder than it used to be.
Yes, it's a problem for that language you plan on creating (try specializing into a niche). But it's not something that should impact Rust's adoption.
You mean rustls? https://github.com/rustls/rustls
Regardless, I'm surprised you haven't heard of rustls - https://github.com/rustls/rustls
Then go around for Khronos, NVidia, Microsoft, Sony, Nintendo, Unreal, Godot,.... to support Rust on their SDKs.
That's not exactly fair to C++ because entire categories of dev tools (like build systems, package managers, IDEs, debuggers, version control, static analyzers, etc.) have matured after C++ did. And let's not forget that when C++ was new, most libraries were proprietary licensed and paid for, whereas today almost all libs are open source. And those improvements (along with general size of the programming community) mean that a trendy language today is going to develop and mature a lot faster than a formerly trendy language did 30 years ago.
IMO Rust in 2022 is a lot closer to java in 2005 than C++ in the 1990s.
Where is the Rust IDE that is half as capable as C++ Builder, MPW/Metrowerks, Visual Age, Zortech?
Considering all features they offered across the board in the box, not only code completion.
As for Visual Studio, yeah it is great in all aspects, except having nothing else beyond MFC to offer on the GUI department, WinUI is still a mess after UWP.
In any case your examples are for C++ IDEs, reinforcing my case of C++ tooling versus Rust.
C++ isn't boring technology, either. If you just want to deliver value, I'd recommend Java.
Not much of an issue if you stick to the stable subset of the language, and libraries that work within that subset.
That kinda feels like saying Linux is too crazy because new apps get made for Linux frequently.
You can use the same part of the language tomorrow that you used today. Nothing is changing out from under you. If you're afraid of libraries, don't use them. You'd have the same problem in any ecosystem that is new, no?
Apps are okay, but other parts of userland that roll out breaking changes on a regular basis are definitely a problem [1] [2] [3]. Even if they aren't technically part of the kernel, they are usually used with it to provide a complete working system, and they break stuff all the time.
[1]: https://lwn.net/Articles/904892/
In my experience, Rust delivers by far the fewest number of bugs in production out of any mainstream language. It gets the fundamentals right like nothing before it. &, &mut, Send and Sync take care of many classes of bugs in the inner loop of productivity.
I agree with you though, so is Rust. The less boring areas imo these days aren't languages (at least none i see), as all the good languages are boring. Zig for example, is pretty mundane too.
The older i get the more i value confidence in a product. Confidence that it won't crash at runtime. Confidence that i won't be bugged over the weekend. etc