Forum Discussion
Topic: Feature Engineering & Data Leakage
- 2 months ago
What is Data Leakage?Data leakage happens when information from outside the training dataset is used to create the machine learning model. This extra data contains information about the target variable that would not actually be available during real-world predictions.Why is it a Problem?False High Performance: The model looks incredibly accurate during testing (e.g., 98% accuracy).Real-World Failure: The model fails completely when deployed because it no longer has access to that "secret" leaked data.Overoptimistic Results: It creates an illusion of a perfect model that cannot perform in production.Common CausesMixing Train and Test Data: Performing data preprocessing (like normalization or scaling) on the entire dataset before splitting it into training and testing sets.Including Future Data: Using features that happen after the target event occurs (e.g., using "hospital stay duration" to predict if a patient has a specific disease).Duplicate Rows: Having the exact same data points in both the training set and the validation set.How to Prevent ItSplit Data First: Always separate your data into training and testing sets before doing any preprocessing, scaling, or imputation.Check Timestamps: Ensure all your training features occur chronologically before the target variable you want to predict.Drop Predictive After-the-Fact Features: Exclude any variable that is updated or created after the target event has taken place.
- 2 months ago
The Problem: It is called Data Leakage.The Cause: The model learns information from the future that would not be available during real prediction.The Result: It gives misleading performance results.Prevention and DetectionCheck all feature sources.Use a proper train-test split based on time.Remove future-based features entirely.Validate the machine learning pipeline carefully.
Data Leakage in ML Pipelines
What it is: Data leakage happens when a model has access, during training, to information it wouldn't realistically have at prediction time. It's useful to split this into two flavors, since they're detected and fixed differently:
- Target leakage - a feature is derived from or correlated with the outcome after the outcome occurred (e.g., a days_since_last_login field that's only populated once a customer has already churned).
- Train-test contamination - information from the test/validation set bleeds into training (e.g., normalizing before splitting, or duplicate records across splits).
Your churn example is target leakage: the feature encodes a consequence of churn, not a cause.
Why it distorts performance
The model latches onto a feature that's essentially a proxy for the label itself, rather than learning genuine predictive signal. This produces a shortcut, not a pattern - accuracy looks great (98%+) because the model is nearly reading the answer key, but that shortcut vanishes in production since the leaked feature simply isn't computable at inference time. The result is a sharp, often embarrassing performance cliff post-deployment.
How to detect it
- Timestamp every feature. For each one, ask: "was this value knowable at the moment I need to make the prediction?" If not, cut it.
- Check feature-target correlation. A single feature with near-perfect correlation to the label is a red flag, not a gift.
- Audit feature provenance. Trace each feature back to its source table/pipeline step - leakage often hides in joins against tables that get updated post-event.
- Be suspicious of accuracy that's implausibly high for the problem's inherent difficulty. Real-world churn, fraud, and default models rarely clear 90%+ without leakage.
- Use time-based (chronological) splits, not random splits, for any temporally-ordered problem. Random splits let future rows "teach" the model about the past.
- Ablation test: drop the suspicious feature and retrain. If accuracy collapses from 98% to something like 75%, that feature was almost certainly leaking.
How to prevent it
- Construct features using a strict point-in-time cutoff - only data available before the prediction timestamp.
- Maintain data lineage, so every feature's origin and computation time is traceable.
- Use a feature store with point-in-time correctness (e.g., Feast, Tecton) so training and serving pull from consistent historical snapshots.
- Add automated leakage checks in the pipeline - e.g., flag any feature with correlation above a threshold to the target for manual review.
- Loop in domain experts to confirm real-world availability of each feature at inference time.
- Shadow-test the full pipeline in a production-like environment before shipping, to catch leakage that only surfaces under real serving conditions.
Key takeaway: A model at 75% accuracy that generalizes is worth more than one at 98% that's silently reading the label. Leakage is the single most common reason ML models perform well in notebooks and fail in production - treating it as a first-class pipeline concern (not just a modeling detail) is what separates production-grade ML from a Kaggle submission.