Databricks Strikes $1.3B Deal for Generative AI Startup MosaicML
wsj.com
wsj.com
Databricks may have its thunder stolen by snowflake. But the true AI boom happening right now, benefits many data product vendors.
The basic requirement for enterprises to use LLMs is having their data in order, which basically requires a cloud data warehouse. It is simply responsible to profit off the hype for a established data company.
I'm curious why you would say this. The main use cases for LLMs involve things like customer service chatbots, knowledgebase search, document summarization, co-pilot/code generation, and content generation for product descriptions and marketing emails. The main data sources would be things like document repositories on Sharepoint and transcripts of customer calls. Not a bunch of historical sales data or financial data sitting in the typical data warehouse. I think there's a big misconception about this. Data sources for Generative AI are very different than data sources for 2013 era data science projects (which perhaps not coincidentally is what led to the development of things like Databricks).
Snowflake stores data in S3 in Snowflake's various AWS accounts. Egress fees are necessary when performing anything outside of Snowflake.
Databricks operates on data stored in your S3 on your AWS account(s). Databricks also runs on your compute contained within your AWS account(s).
Both approaches have valid use cases.
Egress fees are only if you are unloading into a different region/cloud. If your data is in SF AWS us-east-1, and you unload to your S3 in us-east-1, there is no egress fee.
> Snowflake charges a per-byte fee for data egress when users transfer data from a Snowflake account into a different region on the same cloud platform or into a completely different cloud platform. Data transfers within the same region are free.
https://docs.snowflake.com/en/user-guide/cost-understanding-...
Usually, no. But in some cases, yes.
If you * need * it, you find a human to talk to (sales or connect from your network).
While Databricks definitely started with Spark, and it’s still a significant foundation, there’s done much more on top and around it. For instance:
- MLFlow for ML lifecycle management from experiments to real-time ML serving;
- Delta Storage format, extending parquet to leverage cloud storage and enable efficient updates and very fast access;
- SQL Warehouses, which expose Databricks as a SQL engine for Analytics;
- Jobs/Workflows, which is one of the most used orchestration engines in the world;
- Unity Catalog, which will replace (at least in Databricks) Hive Metastore for metadata, access control, lineage and data governance tooling in general.
And now LLMs, on top of the data and ML capabilities mentioned above, much extended by the Mosaic deal (still to be approved).
The interesting thing is that yes, Databricks can do the whole “Data Warehousing” thing, but it can also do very large scale streaming, machine learning, process unstructured data like text, audio or images, support BI applications, etc - all accessing the same data with compatible tooling. So, it’s a full blown multi-workload data platform for any kind of use case and company size.
One can argue that most components are open-source and can be deployed independently - and Databricks has open-sourced Spark, MLFlow and Delta. It’s just that most companies simply don’t have enough (if any) staff with skills to deploy and operate all these things, let alone as one integrated platform. With Databricks, I’m used to deliver a demo where it takes me about 20 minutes from having a new cloud account to be running data workloads against a cluster or SQL, with all the functionality above.
As such, they’ll be forced to do a lot more services oriented work rather than product / platform oriented work, because it pays well. Their sales team is also excellent. I see a similar fate for them as Cloudera.
Do you have some data to support this? This is a pretty bold claim.
Source: https://www.bloomberg.com/news/articles/2023-06-13/databrick...
> “corporate leaders under pressure to get their data ready for AI”. That has nothing to do with LLMs
I agree that its buzzwordy and a little abstract. But also in my experience, getting "data ready for AI" is actually the primary constraint many orgs have with respect to using LLMs in an enterprise context. Their data is not stored in a way thats easy to tokenize/label/embed for effective training or fine-tuning. And you could argue the preprocessing actually should the easy part for non-ML devs to tackle, as its primarily a software engineering problem. Yet, its still the thing that keeps many folks from getting started (before you even get to attempt to tackle the inference problem).
It may surprise HN to learn, but most companies don't have top-tier technical talent.
Consequently, how do you do {cutting edge thing board and investors are demanding company do} with the people you have/can afford? Use an offering that decreases the necessary skill level by providing powerful prebuilt components.
From this perspective, LLM integration sounds like a perfect fit. It's something every company is being asked to have a plan on, but one few are technically staffed to execute on their own.
But this is still the "language of hype". LLM integration to do what? I'm not saying they're not doing anything useful, just that this is not a way to describe either a productivity-enhancing technology (feature) or an actual value proposition (benefit) in a meaningful way.
It depends who you're selling to. If customers are asking for a way to deal with LLMs, then providing that is a perfectly fine value proposition. The "do what" is going to depend on the customer.
What the end deliverable of the plan is doesn't matter: effectively it's the company demonstrating it can "do AI," so that if it becomes existentially necessary, the company's ability to execute is already known.
Absent demonstration, hell, the company might not even be retaining useful historical data. (E.g. theft data at a major US retailer)
I'm not even saying it's unmitigatedly a negative thing. When the question "What are we going to actually do with XML?" came up, many had good ideas, and useful tech investments started getting past boards that wouldn't otherwise have. -- So there was a lot of "collateral goodness".
But people also started ripping out databases and purpose-built DSLs from existing applications just so they could replace them with something XML-based.
As a way of making tech investments, it's weak. If it's 1999 and I'm asked to invest in a tech company that makes a damned good XML parser and nothing else, I'm well advised to not invest. And if I'm a CTO in a non-tech business asked to start replacing database servers with XML files, I'm well advised to not do it.
Running your own Spark, especially on prem, is a lot of work. Most companies would prefer to just provide their data and let someone else handle the query engine.
The parent is right however that Databricks has a feature store (tokenization) but it's not simple to set up and just getting content in and out of it is a major pain point right now.
Databrick cost is stupid high, like snowflake. Any companies looking to chop budget would put those contracts up first.
Like, I'm at company 3 that pays for databricks.
Amount of those companies that use Spark = 0.
I've given up complaining now.
It's really depressing, tbh but I've made my peace with it now.
Most of these startups (AI and others) have to offer a compelling product before even being notable.
Besides, AWS top level doesn’t care if you use sagemaker or not. There’s a premium but if you’re still using EC2 via another startup, they’re still capturing lions share of value.
Mosaic is one of the better providers for this. AWS is nowhere near ready at this current point in time, it is pretty much a "dumb" infra provider in large LLM training at this point. (Of course they won't be standing still and will prob acquire that capability one way or another)
For SMBs, the value would be in using the LLM to generate responses to customer Q&A/search queries. But these companies aren't going to integrate some external third party service, they'll only use it if it's already baked into their CMS - Wordpress/Shopify/Wix/etc. I just don't see who the final consumer for this product would be.
Why not? Every larger "cloud" company seems to be randomly buying 1 at the moment to offer "AI" so might get some good deals. This is clearly 1 of them - panic buy.
YMMV. Sometimes a LORA is fine, but sometimes a full finetune is necessary for higher quality output.
That being said, backwards pass free training keeps making more and more progress. Seems like a short matter of time before it becomes practical.
I just fine tuned a ~30b parameter model on my 2x 3090s to check it out. It worked fantastically. I should be able to fine tune up-to 65b parameter models locally but wanted to get my dataset right on a smaller model before trying.
It seems to me that the vast majority of these people would be better off just doing semantic search with their documents chunked, run through an embeddings process, and stored in a vector database, with the search queries and results then run through an LLM at the final step to create an actual "answer". For applications where this is not practical, I agree that LoRA should be the next approach. I have a hard time believing that the future is everyone training their own domain specific LLMs from the ground up.
This is likely mostly stock deal, backed by databricks equity that, like most other private companies, is illiquid and has taken a nose dive in value in the last couple of years.
I don’t know the exact split, but I’d imagine it is 300m cash, 1B in databricks stock at some last round private valuation that’s likely untested
Databricks last raises at 38B valuation at 1B ARR near peak valuation hype in August 2021.
Assuming they can get 10x and their ARR has doubled since then, they’re down maybe 50 percent in valuation.
Internally they say they are marking themselves down to about 34B.
Is the deal using databricks shares valued at 38B? 34B? 20B? Who knows. I hope it’s the latter.
And I hope that whenever Mosaic folks get to sell their databricks stocks that it is worth more than that
However, a year later, their SaaS has been discontinued. The open-source repository is now stagnant, with hundreds of unresolved merge requests. On a positive note, they recently shared the repository with some other open-source maintainers, so there's hope that Redash will be reborn.
Yeah, the root of the problem (in my opinion) was that Redash has sooo many Python dependencies (due to supporting so many databases and similar) that it's become a real hairball source code wise to keep them all playing together nicely.
Especially as over time (say) library package FOO has some security vulnerability reported that gets fixed in a new release... but the dependencies of the newer release are too new to work with (say) package BAR. Times that by 50 and it's a real pita.
Simultaneously to that, the Redash team got busy with their work at Databricks (mostly not Redash related). Then the automatic CircleCI checks on PRs started failing (ugh), etc.
---
But, as @bratao mentioned above that's all getting worked through now.
Admin and maintainer permissions have been given to a group of dev volunteers / known Redash enthusiasts. CI is working again now (as of last night), and we're currently untangling the dependency hairball.
It's likely to be a few weeks (minimum I guess) before any new official releases are ready, but it will happen. :)
I had a recruiter reach out from DB late last year. I wasn't looking for new work but it seemed like it'd be worth a chat. I had to move my schedule around to fit it in and then get up extra early to be prepared for the chat...only the recruiter never showed. Didn't follow up about missing a meeting they had scheduled. It's been crickets. That is enough of a lack of professionalism really stood out as a red flag. Hard pass.
So often we see these "success" stories and we don't hear exactly how the "regular joes" made out. And it can color our perception of the startup gamble, without real evidence.
Mosaic's 7B (and 30B?) models have the same issue, and 7B kinda paled in comparison to LLaMA 7B... But maybe it would be better if finetuned?
To me, its kinda baffling that Mosaic didn't work on adding highly quantized inference to the popular frameworks.
Here's MPT-30B running in 4-bit precision on CPU :) https://twitter.com/abacaj/status/1673133443339763712?s=20
[1] https://huggingface.co/replit/replit-code-v1-3b [2] https://www.youtube.com/watch?v=roEKOzxilq4
Estimates of Databricks revenue is $400M and “valuation” of $36B, which would never hold today
Rough math plugging in public #s and comments here:
- All stock deal at Aug 2021 val of 38B (1B ARR)
- Assume rev doubled to 2B (which may even be aggressive)
- SAAS multiples are down 6x since Aug 2021
- 38B x 2 / 6 = $12.7B
- 12.7B / 38B * 1.3B = 434M = effective price
- Assume 100M to pref stock
--> Comes out to 334M, with a chunk of that (1/3? 1/4?) potentially subject to earn out
I think it's a good accretive deal because Databricks already has a good number of enterprise relationships and Mosaics key offering has been training software with their own optimizations that increase gpu efficiency.
This is not a very big company and hasn't been around for very long. 1.3B is a lot of money
That's a lot of people you have to replace with one model--even if you get the cost basis down. Do those costs include the period refitting?