> “corporate leaders under pressure to get their data ready for AI”. That has nothing to do with LLMs
I agree that its buzzwordy and a little abstract. But also in my experience, getting "data ready for AI" is actually the primary constraint many orgs have with respect to using LLMs in an enterprise context. Their data is not stored in a way thats easy to tokenize/label/embed for effective training or fine-tuning. And you could argue the preprocessing actually should the easy part for non-ML devs to tackle, as its primarily a software engineering problem. Yet, its still the thing that keeps many folks from getting started (before you even get to attempt to tackle the inference problem).
It may surprise HN to learn, but most companies don't have top-tier technical talent.
Consequently, how do you do {cutting edge thing board and investors are demanding company do} with the people you have/can afford? Use an offering that decreases the necessary skill level by providing powerful prebuilt components.
From this perspective, LLM integration sounds like a perfect fit. It's something every company is being asked to have a plan on, but one few are technically staffed to execute on their own.
Running your own Spark, especially on prem, is a lot of work. Most companies would prefer to just provide their data and let someone else handle the query engine.
The parent is right however that Databricks has a feature store (tokenization) but it's not simple to set up and just getting content in and out of it is a major pain point right now.
But this is still the "language of hype". LLM integration to do what? I'm not saying they're not doing anything useful, just that this is not a way to describe either a productivity-enhancing technology (feature) or an actual value proposition (benefit) in a meaningful way.
What the end deliverable of the plan is doesn't matter: effectively it's the company demonstrating it can "do AI," so that if it becomes existentially necessary, the company's ability to execute is already known.
Absent demonstration, hell, the company might not even be retaining useful historical data. (E.g. theft data at a major US retailer)
I'm not even saying it's unmitigatedly a negative thing. When the question "What are we going to actually do with XML?" came up, many had good ideas, and useful tech investments started getting past boards that wouldn't otherwise have. -- So there was a lot of "collateral goodness".
But people also started ripping out databases and purpose-built DSLs from existing applications just so they could replace them with something XML-based.
As a way of making tech investments, it's weak. If it's 1999 and I'm asked to invest in a tech company that makes a damned good XML parser and nothing else, I'm well advised to not invest. And if I'm a CTO in a non-tech business asked to start replacing database servers with XML files, I'm well advised to not do it.
It depends who you're selling to. If customers are asking for a way to deal with LLMs, then providing that is a perfectly fine value proposition. The "do what" is going to depend on the customer.
Databricks may have its thunder stolen by snowflake. But the true AI boom happening right now, benefits many data product vendors.
The basic requirement for enterprises to use LLMs is having their data in order, which basically requires a cloud data warehouse. It is simply responsible to profit off the hype for a established data company.
I'm curious why you would say this. The main use cases for LLMs involve things like customer service chatbots, knowledgebase search, document summarization, co-pilot/code generation, and content generation for product descriptions and marketing emails. The main data sources would be things like document repositories on Sharepoint and transcripts of customer calls. Not a bunch of historical sales data or financial data sitting in the typical data warehouse. I think there's a big misconception about this. Data sources for Generative AI are very different than data sources for 2013 era data science projects (which perhaps not coincidentally is what led to the development of things like Databricks).
While Databricks definitely started with Spark, and it’s still a significant foundation, there’s done much more on top and around it. For instance:
- MLFlow for ML lifecycle management from experiments to real-time ML serving;
- Delta Storage format, extending parquet to leverage cloud storage and enable efficient updates and very fast access;
- SQL Warehouses, which expose Databricks as a SQL engine for Analytics;
- Jobs/Workflows, which is one of the most used orchestration engines in the world;
- Unity Catalog, which will replace (at least in Databricks) Hive Metastore for metadata, access control, lineage and data governance tooling in general.
And now LLMs, on top of the data and ML capabilities mentioned above, much extended by the Mosaic deal (still to be approved).
The interesting thing is that yes, Databricks can do the whole “Data Warehousing” thing, but it can also do very large scale streaming, machine learning, process unstructured data like text, audio or images, support BI applications, etc - all accessing the same data with compatible tooling. So, it’s a full blown multi-workload data platform for any kind of use case and company size.
One can argue that most components are open-source and can be deployed independently - and Databricks has open-sourced Spark, MLFlow and Delta. It’s just that most companies simply don’t have enough (if any) staff with skills to deploy and operate all these things, let alone as one integrated platform. With Databricks, I’m used to deliver a demo where it takes me about 20 minutes from having a new cloud account to be running data workloads against a cluster or SQL, with all the functionality above.
Snowflake stores data in S3 in Snowflake's various AWS accounts. Egress fees are necessary when performing anything outside of Snowflake.
Databricks operates on data stored in your S3 on your AWS account(s). Databricks also runs on your compute contained within your AWS account(s).
Both approaches have valid use cases.
Egress fees are only if you are unloading into a different region/cloud. If your data is in SF AWS us-east-1, and you unload to your S3 in us-east-1, there is no egress fee.
> Snowflake charges a per-byte fee for data egress when users transfer data from a Snowflake account into a different region on the same cloud platform or into a completely different cloud platform. Data transfers within the same region are free.
https://docs.snowflake.com/en/user-guide/cost-understanding-...
Usually, no. But in some cases, yes.
If you * need * it, you find a human to talk to (sales or connect from your network).
Databrick cost is stupid high, like snowflake. Any companies looking to chop budget would put those contracts up first.
Like, I'm at company 3 that pays for databricks.
Amount of those companies that use Spark = 0.
I've given up complaining now.
It's really depressing, tbh but I've made my peace with it now.
Most of these startups (AI and others) have to offer a compelling product before even being notable.
Besides, AWS top level doesn’t care if you use sagemaker or not. There’s a premium but if you’re still using EC2 via another startup, they’re still capturing lions share of value.
Mosaic is one of the better providers for this. AWS is nowhere near ready at this current point in time, it is pretty much a "dumb" infra provider in large LLM training at this point. (Of course they won't be standing still and will prob acquire that capability one way or another)
As such, they’ll be forced to do a lot more services oriented work rather than product / platform oriented work, because it pays well. Their sales team is also excellent. I see a similar fate for them as Cloudera.
Source: https://www.bloomberg.com/news/articles/2023-06-13/databrick...
Do you have some data to support this? This is a pretty bold claim.