The Big Dictionary of MLOps
hopsworks.ai
hopsworks.ai
Data leakage is I think when data IN the training dataset regardless how they ended up there contained information which were part of the output and will not be available at inference time, thus enhancing the model's performance.
Example: Estimating crop yield earnings in $ from sat photos and getting a historical dataset that included crop type. Crop type should be unavailable at inference time so it should also be detected from the sat photo although it is not the target variable.
Whether it is externally or internally provided does not capture the essence of the problem.
Now, for your example - if the crop was a feature used when training, your model inference input should also require that you input the crop type. When you write your inference pipeline, you will realize that - oh, that feature is not available at inference time. I should retrain my model without the feature. It's not really data leakage - it's incorrect use of features when training a model - that model should never make it to production. So, maybe it's an edge case of data leakage.
Feature data leakage is when you include a feature when training a model, but that feature is not available at inference time. You should, hopefully, discover feature data leakage when you try to write your inference pipeline and realize the feature is not available. The use of a feature store, that informs you whether your feature will be available for online models or only offline models, should help prevent you from training models with feature data leakage.
AI cannot escape replicating the training material and characteristics of its creators.
Is it not a hallucination we are creating “artificial intelligence”? Human consciousness took millions of years of galactic churn. Is convincing our senses via disembodied metal and plastic we can unplug creating intelligence? Seems pretty dumb if it can’t avoid it’s own death from unplugging.