MongoDB acquires Voyage AI
investors.mongodb.com
investors.mongodb.com
Postgres is hard, you have to learn SQL. SQL is hard and mean.
Mongo means we can just dump everyone into a magic box and worry about it later.No tables to create.
But their is little time, we need to ship our CRUD APP NOW! No one on the team knows SQL!
I'm actually using Postgres via Supabase for my current project, but I would probably never use straight up Postgres.
It has supported this since 9.4: https://www.postgresql.org/docs/current/datatype-json.html
It's easier to get started with.
But why is that the top priority?
Maintainability? Secondary. Security? Secondary. Data-integrity/correctness? Secondary.
Writing code and creating good software requires a lot of mental clarity and effort; that fact is never going to change, not even with AI.
Firebase by and almost every NoSql technology is based upon this.
I've been playing with CloudNativePG recently and adding replicas is easy as can be, they automatically sync up and join the cluster without you thinking about it.
Way nicer than the bare-vm ansible setup I used at my last company.
On average an AWS availability zone tends to suffer at least one failure a year. Some are disclosed. Many are not. And so that database you are running on a single instance will die.
Question is do you want to do something about it or just suffer the outage.
That being said there are plenty of ways to shard Postgres that are free, e.g. Citus. It's also questionable whether many need sharding. You can go a long way with simply a replica.
Postgres also has plenty of its own strengths. For one, you can get a managed solution without being locked into MongoDB the company.
And history has not been nice to startups like this continuing their products over the long term.
It's why unless it is built-in and supported it's not feasible for most to depend on it.
Microsoft does not make money supporting Citus.
MongDB is basically a pile of JSON in comparison, no matter how much you distribute and scale it.
See https://jepsen.io/analyses for how MongoDB has a tradition of incorrect claims and losing your data.
Distributed databases are not easy. Just saying "it is web scale" doesn't make it so.
1. That PgSQL also has issues in jepsen tests?
2. of any distributed DB which doesn't have jepsen issues?
3. It is configurable behavior for MongoDB: can it lose data and work fast, or work slower and do not lose data. There is no issues of unintentional data loss in most recent(5yo) jepsen report for MongoDB.
Your second point seems to imply that everything has issues, so using MongoDB is fine. But there are various kinds of problems. Take a look at the report for RethinkDB, for example, and compare the issues found there to the MongoDB problems.
RethinkDB doesn't support cross document transactions, problem solved lol
MongoDB defects were, let's say, somewhat more severe
[2.4.3] "In this post, we’ll see MongoDB drop a phenomenal amount of data."
[2.6.7] "Mongo’s consistency model is broken by design: not only can “strictly consistent” reads see stale versions of documents, but they can also return garbage data from writes that never should have occurred. [...] almost all write concern levels allow data loss.
[3.6.4] "with MongoDB’s default consistency levels, CC sessions fail to provide the claimed invariants"
[4.2.6] "even at the strongest levels of read and write concern, it failed to preserve snapshot isolation. Instead, Jepsen observed read skew, cyclic information flow, duplicate writes, and internal consistency violations"
let's not pretend that Mongo is a reliable database please. Fast? likely. But if you value your data, don't use it.
* why you are referring on 12yo reports for very early MongoDB version?
If you wish to have a more recent MongoDB report, Jepsen is available for hire, from what I understand.
b) MongoDB does one thing well. JSON documents. If your domain model is built around that then nothing is faster. Seriously nothing. You can do tuple updates on complex structures at speeds that cripple PostgreSQL in seconds.
c) Nobody who is architecting systems ever thinks this way. It is never MongoDB or PostgreSQL. They specialise in different things and have different strengths. It is far more common to see both deployed.
Statements like yours are meaningless when you aren't specific about the operations, schema, access patterns etc.
If you have a single server, relational use case then PostgreSQL is great. But like all technology it's not great at everything.
In all seriousness, calling Postgres’ scalability “not-negotiable for most use cases” is wild.
"built-in, supported, proven scalability and high availability"
PostgreSQL does not have any of this. It's only good for a single server instance which isn't really enough in a cloud world where instances are largely ephemeral.
Excuse me? I do enterprise apps, along with most of the developers I know. We run like 100 transactions per second and can easily survive hours of planned downtime.
It's 2025, computers are really fast. I barely need a database, but ACID makes transaction processing so much easier.
granted, the failures were pretty minor, especially compared to previous reports (like the first one [1], that was a fun read), but they still had bad defaults back then (and maybe still do)
I would not trust anything MongoDB says without independent confirmation
Although I would point out:
> scalability [...] no company would even bother with PostgreSQL at all
In my experience, you can get pretty far with Postgresql on a beefy server, and when combined with monitoring, pg_stat_statements and application level caching (e.g. the user for this given request, instead of fetching that data on every layer of the request handling), certainly enough most businesses/organisations out there.
In MongoDB the situation is different. You have to deal with the bare minimum of a database. But in return your data design has much higher horizontal scalability survivability.
In the initial phase of your startup, choose MongoDB. It's easier to start and evolve in earlier stages. And later on, if you feel the need and have resources to scale PostgreSQL, move your data there.
Whenever a fly.io post about sqlite ends up in here, there are a scary amount of comments about using sqlite in way more scenarios than it should be.
But mainly because management hasn't worked out how to cancel their licenses without breaking their budgets.
SQLite is a lean and clean tool, it's very much a candidate for being inserted into all manner of contexts.
What beggars belief is the overly complicated, inefficient, rats nests of trendy software that developers actually string together to get things done, totally unaware of how they are implemented or meant to work.
By comparison using SQLite outside of its "blessed (by who?) use cases" is very practical.
I have personally been incredibly impressed with the way MongoDB has "looked out" for my company over the past year. I'll try to be satisfactorily specific their with privacy in mind, so this may come out a bit fuzzy. Their technical teams, both overseas and US, have produced some of the most thorough, detailed recommendations I have ever seen and their communication/followup was excellent. I've run many designs and ideas by their team, out of habit at this point, and have always been pleased with the response. They remember who I am. All of this is really unusual for a company at my growth stage.
My use case requires full technical depth on text searching and vectorization; I use every aspect of Atlas Search that is available. A downside of building "bleeding edge" is that my tooling needs always seem to be just inches beyond what is available, so just about every release seems to have something that is "for me." It's hard to say if my feedback specifically has an impact on their roadmap - but they really do seem to build things I want. I think they reported ~50% better performance on bulkWrite() in the 8 release, but it was closer to 500% for my use case.
Speaking of, this acquisition is like providence for me, because I've shared my various solutions with them for synchronously vectorizing "stuff" for use with local LLMs. It's a reasonably hard technical problem without a lot of consensus on standards; I think a lot of people believe there are standards, yet any discussion will quickly devolve into something like the "postgres/mongo" conversations you see here (I won't be visiting that topic).
I strongly agree with the "this should be a database level feature" approach they are taking here; that's certainly how my brain wants to think about it and currently I have to do a lot of "glue"-ing to make it work the way I require.
I hope they win.
Last MongoDB is still slower than MongoDB 3.4. An almost 10-year old release. For both reads and writes.
We have internally a benchmark with MongoDB 8.x, but same pattern of disappointing results.
8 is waaay faster in the real world. It's not really comparable. Your micro benchmark is comparing the few nanoseconds of the heavier query planner, but in the real world that query planner gives real benefits. Not to mention aggregations, memory management improvements, and improvements when your working set size is very large/larger than memory.
Here's another dataset about performance regression doing `$inc`s as fast as possible on the same object.
Mongo 3.4.24:
332,037 stats update in 100s. (3,321 stats updates per s)
Mongo 8.0.4: 287,553 stats update in 100s. (2,876 stats updates per s)
(higher is better)Do you often do that many independent $incs (or any query) in a single second? I have gotten much better performance by using `BulkWrite` to do a bunch of small updates in a batch.
To go to a specific example from the "Driver Benchmark" on the link from your first reply:
client[:users].insert_one(name: Digest::MD5.hexdigest(index.to_s))
I notice in this specific example that there's no separation of the hashing from the query timing. so I might try to do the hashing first then time just the inserts. I would also a batch of `insertOne`s and then do a bulk write so I'm making much fewer queries. I will often pick some random size like 1,000 queries or so and do the `bulkWrite ` when I have accumulated that many queries, have surpassed some time (like if it has been more than 0.5s since the last update) or if there's no more items to process. Additionally if the order of the inserts doesn't matter using `ordered: false` can provide additional speedup.For me the limiting factor is mostly around the performance of BulkWrite. I haven't hit any performance bottlenecks there that would merit benchmarking different ways to use it, but I would mostly be trying to fine tune things like how to group the items in a BulkWrite for optimal performance if I did.
Even in the case of one-off queries it almost always feels faster on 7+ than earlier versions. As I mentioned the one bottleneck we hit with migration was that we had some queries where we were querying on fields that were not properly indexed and in those cases performance tanked horribly to the point where some queries actually stopped working. However, once we added an index the queries were always faster than on the old version. When we did hit problems, it took only a few minutes to figure out what to index then everything was fine. We didn't have to make changes to our application or the queries themselves to fix any issues we had.
Setup a clone of prod and build a tool to replay your traffic to it.
I have lots of data from datadog and ops manager but not going to take the time to publish ATM.
I just moved a 4tb deployment from 3.2 to 7. It cut max query time by about half. I actually went to instances with half the cpus, too (although I switched from ebs to ssds).
It was for us. API calls that need to aggregate stats on same ID. We found a way around, but it would not have been an issue if MongoDB 8 was like 2x faster.
> I just moved a 4tb deployment from 3.2 to 7. It cut max query time by about half. I actually went to instances with half the cpus, too (although I switched from ebs to ssds).
Just single-core performance improvement in the last 10-year might explain your outperformance.
Nope, after migration max query time was still over a minute in some cases. What makes the biggest difference is performance tuning. After a week or so of index tuning, I got max index time below 6s. If Mongo makes each query take 2ms instead of 1ms, it literally doesn't matter to that customer or their customers, since it's just noise at that point. The old instances were M5s, so not that old.
The point is that the few nanoseconds difference you're measuring is not what you spend the most time on, usually.
Also you mentioned write performance. If you set journal commit interval to 500ms or something, then you can easily beat the old 3.2 write speeds, since if you're using 3.2 you probably don't care that much about journal commit intervals anyway ^_^
On the tech side ... no idea what Mongo's plan is ... their embedding model is not SOTA, does not even outperform the open ones out there, and reranking is a dead end in 2025.
I think the value is on Voyage's team, their user base and having a vision that aligned with Mongo's.
Congrats!
With embedding/generative models becoming better with time, the need for a rerank step will be optimized away.
Are you referring to the MTEB leaderboard? It's widely believed many of those test datasets are considered during the training of most open-source text embedding models, hence why you see novel + private benchmarks discussed in many launch blogs that don't exclusively refer to MTEB. There are problems there, and it would be great to see more folks in the search benchmark dataset production space like what Marqo AI has done in recent months.
Also what makes you say reranking is dead? Mongo doesn't provide it out of the box but many other search providers like ES, Pinecone, Opensearch do so it must provide some value to their customers? Maybe you're saying it's overrated in terms of how many apps actually need it?
disclosure: I work on vector search at Mongo
Yes, my comment leans more towards that, rather than suggesting is useless.
All of which are just data storage/retrieval mechanics and custom types. This isn't adding some omnipotent AI agent to run/manage/optimize your DB or otherwise turn it into some blackbox gizmo.
I see GenAI as a stop gap solution at best, not really optimal for any problems; and AGI is a major distraction from finding good solutions to important problems.
The wild goose chase to apply GenAI to everything has serious consequences.
People are so excited about the fact that a computer can sort of drive a car that they don't even stop to consider that a human driver that randomly fails the same way would never get a license, and rightly so.
So excited about the fact that a computer can sort of write functional code that they don't stop to consider that any human developer that fails randomly the same way would never get a job, and rightly so.
We're already applying it to weapons/warfare, which is obviously a very bad idea.
I'm sure the technology will improve, but never to the point where it's reliable. It will fail randomly less often, but the magnitude of its failures isn't going anywhere.
It was starting to get expensive though, so we were experimenting with other document stores (dynamodb was being trialled, since we were already AWS for most things, just around the time I left)
thats the reason for using Atlas
Great products are in my opinion a function of skill and care. The only benefit a "popular" tool or language gets you is a greater developer pool for hiring.
There’s probably some extremely specific niche use-cases where it works well, but I suspect they’re pretty few-and-far-between.
Also, on some things, it allows us to pivot much faster. And now with the help of LLMs, writing "Aggregation Pipelines" are very fast.
Maintaining apps built on MongoDB is soul killing.
https://seekingalpha.com/news/4412466-mongodb-acquires-voyag...
https://gist.github.com/cpursley/c8fb81fe8a7e5df038158bdfe0f...
What am I missing? Are Mongo users simply front end folks who didn't have time to learn basic SQL or back end architecture?
also, I'd challenge your thinking - ultimately the goal is to solve problems. you don't necessarily need SQL, or relations for that matter. that being said, naively modeling your stuff in mongodb (or other things like dynamodb) will cause you severe pain...
what's also true, which people forget, is naively modeling your stuff with a relational database will also cause you pain. as they sometimes say, normalize until it hurts, and then denormalize to scale and make it work
the amount of places I've seen that skip the second part and have extremely normalized databases makes me cringe. it's like people think joins are free...
I just ran into a greenfield project where the dev reached for Mongo, and didn't have a good technical reason for it beyond "I'm handing documents". Probably wasn't aware of alternatives. FWIW Postgres would've been a great fit for it, they were modeling research publications.
"MongoDB ships with horizontal sharding out-of-the-box, has idiomatic and well-maintained drivers for pretty much every language you could want (no C library re-use), is reasonably vendor-neutral and can be run locally, and the data modeling it encourages is both preferential for some people as well as pushes users to avoid patterns that don't scale very well with other models. Whether these things are important to you is a different question, but there is a lot to like that alternatives may not have answers for. If you currently or plan on spending > 10K per month on your database, I think MongoDB is one of the strongest choices out there."
I have also run Postgres at very large scale. Postgres' JSONB has some serious performance drawbacks that don't matter if you don't plan on spending a lot of money to run your database, but MongoDB does solve those problems. This new documentdb extension from Microsoft may solve some of the pain, but this is some very rough code if you browse around, and Postgres extensions are quite painful to use over the long term.
The reality is that it is not possible to run vanilla Postgres at scale. It's possible to fix its issues with third party solutions or cobbling together your own setup, but it takes a lot of effort and knowledge to ensure you've done things correctly. It's true that many people never reach that scale, but if you do, you're willing to spend a lot of money on something that works well.
Maybe it's better than it was, but my experience with Mongodb a decade ago is that that horizontal sharding didn't work very well. We constantly ran into data corruption and performance issues with rebalancing the shards. So much so that we had a company party to celebrate moving off of Mongodb.
So before the Apple Watch was released.
Why is this relevant today ? Technology changes very quickly.
SQL or rather just some schema based database has a ton of advantages. Besides speed, there is a huge benefit for developers to be able to look at a schema and see how the relationships in the data work. Mongodb usually involves looking at a de facto schema, but with fewer guarantees on types relations or existence, then trawling code for how its used.
Now I could possible spend a bunch of time and do the same thing with open source dbs - but why? I have a small team and stuff to deliver. Atlas allows me to do it fast.
This is why startups like to get into a single supply chain contract with an enterprise - it’s extremely hard to get it setup, but once done very easy to reuse the template.
Just because Postgres is good doesn't mean other things can't also be good (and better for some use cases).
Or it could just be an expansion to their cloud offering. In a lot of cases embedding models just need to be 'good enough' and cheap and/or convenient is a winning GTM approach.
Unless you can get to a point where you can make these models small enough that basically sit in the DB layer of an application...