Quoting some from the article: "Data warehouses are used primarily by business analysts for interactive querying and for generating historical reports/dashboards on the business. Feature stores are used by both data scientists and by the online/batch applications, and they are fed data by feature pipelines, typically written in Python or Scala/Java. Also, Data warehouses mostly stores data in relational tables, whereas a Feature Store stores it as numerical and categorical features and outputs tensors and/or vectors for training or serving.
This company is trying to make a distinction between Online Data (real time streaming with low latency), no joins, key/store and a more traditional batch processing, OLAP type configurations.
but modern data warehouses can support both.
https://www.snowflake.com/streaming-data/
I think this is an effort to segment the data warehousing market and provide new names for things that already exist and providing a vocabulary to users who may not be familiar with a company's existing datawarehouse solutions
To be more concrete, Feast is an open-source Feature Store built on BigQuery and originally BigTable. But the latency of BigTable for the OLTP workload was too high for GoJEK (feature lookup is just one part of making a prediction), so they switched to Redis. Redis PK lookups are a couple of ms, on average, compared with 10+ ms for BigTable. What is the latency of a PK lookup on snowflake? It ain't a millisecond or two. On MySQL Cluster (NDB), our online feature store, PK lookups return in sub-ms latency on dedicated hardware.
Disclaimer: My company supports ClickHouse.
> isn't this just a data warehouse [DW]?
My understanding is that the addition of a RowStore to the DW/ColumnStore addresses the training phase of Machine Learning.
I come from a Data Engineering background. I struggled with the Data Science centric terminology. Uber's 2017 post [1] was helpful in establishing the motivations and terminology of their Michelangelo machine learning (ML) platform. The main distinguishing feature of a Feature Store seems to be that it supports efficient batch downloading of row data that is used as the training dataset. The discussion made more sense once I figured out that feature refers to a column or data field.
Figure 4 in the Logical Clocks whitepaper Hopsworks Feature Store [2] helped me understand the architecture better. The architecture appears to be what I would call a Hybrid Transactional/Analytical Processing (HTAP) [3] engine with MySQL cluster acting as the RowStore and Apache Hive acting as the ColumnStore. I'm assuming that the Hopsworks Feature Store periodically merges the MySQL updates into Hive and also provides a mechanism to perform federated queries.
The use of Hive seems outdated (vs. say Presto) and I wonder if the use of MySQL is required compared to directly accessing column oriented files like ORC/Parquet/Kudu.
[1] https://eng.uber.com/michelangelo-machine-learning-platform/
[2] (PDF) https://uploads-ssl.webflow.com/5e6f7cd3ee7f51d539a4da0b/5ef...
[3] https://en.wikipedia.org/wiki/Hybrid_transactional/analytica...
We also automatically synchronize the metadata to elasticsearch, so you can do free-text search for features, files, extended metadata, etc. So, you can search up to PBs of files/dirs/tables/features/extended-metadata in milliseconds.
You can find more info here on the research behind it:
"They could derive historical insights into the business using BI tools."
A lot of work in some traditional BI tools (e.g. Oracle Hyperion) is actually about collecting and aggregating forward looking data (within forecasts, budgets, scenarios) that are then compared with actuals.
To sum up, assuming that the online feature store is some sort of a cache, how do you know which objects to place in the cache?
Generally I don’t view a “feature store” as being a simple lake of raw data. To me a feature store is a place where features known to have signal in a ML model are stored. It has structure, “clean” data (to the extent possible) and some pre-computed elements (e.g. aggregates over some commonly used time windows) to facilitate efficiency in the end to end ML pipeline.
I’d also argue a feature store is not complete unless there are tools and infrastructure to make it relatively easy to provision and access new features: that is, it should cover the issues you raise.
My question is about the operational aspect of online vs offline stores.
I.e. what happen if the features needed for a prediction are not in the ONLINE store?
https://databricks.com/session_eu20/real-time-feature-engine...