Machine Learning Engineer Guide: Feature Store vs. Data Warehouse
logicalclocks.com
logicalclocks.com
The fact that it happens to be kinda buzz worthy is a collateral aspect: everything that answers what some people wonder and that is not yet answered plainly, is.
(and I mean, the first on the front page at this very second has : "We hacked apple" in the title.)
It's not so much about data itself as an attempt to solve a communications and coordination organizational problem; you decouple sources of data and consumers of data (not the technical systems/databases, but the people and organizational units) to a 'hub-and-spoke' model where the providers of data just supply raw data without getting into a multinational project that takes a year just to identify the potential stakeholders for that data throughout a distributed organization with tens of thousands of employees.
From the end-user standpoint that's not very useful. But that's why you have data marts that normalize the raw data into a standardized format. Ultimately though the raw data needs to remain the single source of truth. If you skip that step, and only store the post-normalized format, you're likely to run into problems down the road. This could either be because you want to change the normalization format. Or you want to use some aspect of the data that isn't captured in the normalized form. Or even you discover a pre-existing bug that affected all the previous post-ETL data.
That's kind of what I understand as well, but the data science folks pitched it in a slightly more positive way, like, "Please don't limit us just to the data you have time to nicely structure and validate. We want all of it. It doesn't matter if a column is getting truncated to three characters or columns are mislabeled or there are amounts in dollars and euros mixed together; we can still get signal out of it."
If it's something else, it's the antithesis of data science.
Today you can, but then when the app owner drops a column (or worse, stop populating it!) in a month that signal will break, and the data lake maintainer will be in the unenviable position of navigating the completely undocumented dependency.
I really understand what you are saying... I know the hype problem, I lived it. It makes me both frustrated and sad - because the hype is annoying, there is a lot of vaporware - but there is also something real that is happening too which is part of the story of the evolution of data architecture/infrastructure. My strong advice is, being open-minded is helpful - learn and take what was good/real and leave behind the stigma/hype. Something real and useful happened in terms of architecture, so take the benefits - but of course, don't compromise on delivering real, working solutions.
Regarding "data lakehouse", I struggle with the buzzword term also, but once again, I recommend looking at what is good/real and ignoring the stigma of the buzzwords. One way of looking at it is the literal translation - a data warehouse made from the components used to make data lakes. To be honest, it is a marketing term, but it is also an architectural pattern we had even before the term existed - for example - you could use a data warehouse product such as Vertica, and back it by HDFS - guess what, there's a data "lakehouse".. and most of the big database vendors can do this trick now - a full traditional data warehouse engine sitting on top of lake storage infrastructure.
There are several "real" use cases for data lakes. Precursor architectures could be seen to be "operational data stores" [1]. Data lakes are real, they are one approach to solving some problems.
These use cases could include: 1) raw, long term storage of large volumes of diversely structured data for staging/historical purposes; 2) data discovery/exploration of this data to identify patterns/models and relationships (this is both an AI and analytics use case for power users and BI/analytics/data scientists, etc.); and 3) an opportunity to change the paradigm of traditional ETL - instead of pulling from sources, one way you look at it is, allowing many diverse/distributed sources to push their data into the lake for powering analytics, exploration, AI model building, etc. It makes sense as part of lambda/kappa architecture as well - some of the "push" in can come from streaming sources as well.
Use case #1 is very much a "data infrastructure" kind of use case that we do anyway in data warehouses - especially those that do ELT (vs. ETL) - staging databases. If you want an architecture that actually helps make some sense of such a use case of data lakes more formally, one could look into Dan Lindstedt's data vault architecture [2]. While data vault modelling doesn't necessarily require "data lakes", the "raw" part of the data vault architecture use case overlaps nicely with data lakes.
Some people treat this as an excuse to throw whatever they want into it, without any organization or standardization... and the consequences of that are quite predictable.
But it doesn't have to be that way. You can accumulate diverse and large data sets, in cheap cloud storage, while knowing what everything is and where you can find it.
As a trivial example, let's say you have a typical OLTP database (or perhaps many), with useful data that is, unfortunately, mutable. You can store entire copies of those tables in your data lake for pennies a day, giving you the ability to recall a transactionally-consistent view of that data from various past times. This is something we've always been able to do using traditional tools, the difference is that storing the data in a "data lake" (i.e. S3) is orders-of-magnitude cheaper.
Another major use case, perhaps the most significant one, is storing the raw ingested data -- e.g. from telemetry collection, 3rd party exports, etc -- along with each stage of its transformation. By preserving the original input, along with all intermediate outputs, no information is ever lost. If a buggy transformation is discovered it no longer means your output is irrevocably corrupted, fixed results can be re-computed from wherever in the transform pipeline the bug manifested. And again, this was always possible, a data lake just makes it cheap enough to actually do on a large scale.
Having nicely organized data is perfect, but I'd rather fetch the data myself from a pile of messy data instead of dealing with all these organizational nightmare.
https://medium.com/data-for-ai - several blogs discussing feature stores for machine learning
https://www.logicalclocks.com/blog/feature-store-the-missing... - in depth explanation
featurestore.org - list of all feature stores available and in production
I really believe that all Enterprise platform software should have an open-source version if it is to have a meaningful effect on how people work (in this case, how Data Scientists and Data Engineers work together).
To sum up, assuming that the online feature store is some sort of a cache, how do you know which objects to place in the cache?
Generally I don’t view a “feature store” as being a simple lake of raw data. To me a feature store is a place where features known to have signal in a ML model are stored. It has structure, “clean” data (to the extent possible) and some pre-computed elements (e.g. aggregates over some commonly used time windows) to facilitate efficiency in the end to end ML pipeline.
I’d also argue a feature store is not complete unless there are tools and infrastructure to make it relatively easy to provision and access new features: that is, it should cover the issues you raise.
My question is about the operational aspect of online vs offline stores.
I.e. what happen if the features needed for a prediction are not in the ONLINE store?
https://databricks.com/session_eu20/real-time-feature-engine...
Quoting some from the article: "Data warehouses are used primarily by business analysts for interactive querying and for generating historical reports/dashboards on the business. Feature stores are used by both data scientists and by the online/batch applications, and they are fed data by feature pipelines, typically written in Python or Scala/Java. Also, Data warehouses mostly stores data in relational tables, whereas a Feature Store stores it as numerical and categorical features and outputs tensors and/or vectors for training or serving.
This company is trying to make a distinction between Online Data (real time streaming with low latency), no joins, key/store and a more traditional batch processing, OLAP type configurations.
but modern data warehouses can support both.
https://www.snowflake.com/streaming-data/
I think this is an effort to segment the data warehousing market and provide new names for things that already exist and providing a vocabulary to users who may not be familiar with a company's existing datawarehouse solutions
To be more concrete, Feast is an open-source Feature Store built on BigQuery and originally BigTable. But the latency of BigTable for the OLTP workload was too high for GoJEK (feature lookup is just one part of making a prediction), so they switched to Redis. Redis PK lookups are a couple of ms, on average, compared with 10+ ms for BigTable. What is the latency of a PK lookup on snowflake? It ain't a millisecond or two. On MySQL Cluster (NDB), our online feature store, PK lookups return in sub-ms latency on dedicated hardware.
Disclaimer: My company supports ClickHouse.
> isn't this just a data warehouse [DW]?
My understanding is that the addition of a RowStore to the DW/ColumnStore addresses the training phase of Machine Learning.
I come from a Data Engineering background. I struggled with the Data Science centric terminology. Uber's 2017 post [1] was helpful in establishing the motivations and terminology of their Michelangelo machine learning (ML) platform. The main distinguishing feature of a Feature Store seems to be that it supports efficient batch downloading of row data that is used as the training dataset. The discussion made more sense once I figured out that feature refers to a column or data field.
Figure 4 in the Logical Clocks whitepaper Hopsworks Feature Store [2] helped me understand the architecture better. The architecture appears to be what I would call a Hybrid Transactional/Analytical Processing (HTAP) [3] engine with MySQL cluster acting as the RowStore and Apache Hive acting as the ColumnStore. I'm assuming that the Hopsworks Feature Store periodically merges the MySQL updates into Hive and also provides a mechanism to perform federated queries.
The use of Hive seems outdated (vs. say Presto) and I wonder if the use of MySQL is required compared to directly accessing column oriented files like ORC/Parquet/Kudu.
[1] https://eng.uber.com/michelangelo-machine-learning-platform/
[2] (PDF) https://uploads-ssl.webflow.com/5e6f7cd3ee7f51d539a4da0b/5ef...
[3] https://en.wikipedia.org/wiki/Hybrid_transactional/analytica...
We also automatically synchronize the metadata to elasticsearch, so you can do free-text search for features, files, extended metadata, etc. So, you can search up to PBs of files/dirs/tables/features/extended-metadata in milliseconds.
You can find more info here on the research behind it:
"They could derive historical insights into the business using BI tools."
A lot of work in some traditional BI tools (e.g. Oracle Hyperion) is actually about collecting and aggregating forward looking data (within forecasts, budgets, scenarios) that are then compared with actuals.
Appreciate the comparison table: Data Warehouse vs Feature Store. Might stea.. re-use that.
I would challenge the assumption that "the Data Warehouse is an input to the Feature Store" though. I'm more inclined towards having a first stage of data cleanup/modeling that could be reused (as input) for both DWH and FS instead.
Respectfully, let's keep ML to be Machine Learning. ;)
Allas, I'll stop shouting "get off my lawn" and let the young ones make their own mistakes...