Forum Discussion
Topic: Feature Engineering & Data Leakage
Data Leakage in ML Pipelines
What it is: Data leakage happens when a model has access, during training, to information it wouldn't realistically have at prediction time. It's useful to split this into two flavors, since they're detected and fixed differently:
- Target leakage - a feature is derived from or correlated with the outcome after the outcome occurred (e.g., a days_since_last_login field that's only populated once a customer has already churned).
- Train-test contamination - information from the test/validation set bleeds into training (e.g., normalizing before splitting, or duplicate records across splits).
Your churn example is target leakage: the feature encodes a consequence of churn, not a cause.
Why it distorts performance
The model latches onto a feature that's essentially a proxy for the label itself, rather than learning genuine predictive signal. This produces a shortcut, not a pattern - accuracy looks great (98%+) because the model is nearly reading the answer key, but that shortcut vanishes in production since the leaked feature simply isn't computable at inference time. The result is a sharp, often embarrassing performance cliff post-deployment.
How to detect it
- Timestamp every feature. For each one, ask: "was this value knowable at the moment I need to make the prediction?" If not, cut it.
- Check feature-target correlation. A single feature with near-perfect correlation to the label is a red flag, not a gift.
- Audit feature provenance. Trace each feature back to its source table/pipeline step - leakage often hides in joins against tables that get updated post-event.
- Be suspicious of accuracy that's implausibly high for the problem's inherent difficulty. Real-world churn, fraud, and default models rarely clear 90%+ without leakage.
- Use time-based (chronological) splits, not random splits, for any temporally-ordered problem. Random splits let future rows "teach" the model about the past.
- Ablation test: drop the suspicious feature and retrain. If accuracy collapses from 98% to something like 75%, that feature was almost certainly leaking.
How to prevent it
- Construct features using a strict point-in-time cutoff - only data available before the prediction timestamp.
- Maintain data lineage, so every feature's origin and computation time is traceable.
- Use a feature store with point-in-time correctness (e.g., Feast, Tecton) so training and serving pull from consistent historical snapshots.
- Add automated leakage checks in the pipeline - e.g., flag any feature with correlation above a threshold to the target for manual review.
- Loop in domain experts to confirm real-world availability of each feature at inference time.
- Shadow-test the full pipeline in a production-like environment before shipping, to catch leakage that only surfaces under real serving conditions.
Key takeaway: A model at 75% accuracy that generalizes is worth more than one at 98% that's silently reading the label. Leakage is the single most common reason ML models perform well in notebooks and fail in production - treating it as a first-class pipeline concern (not just a modeling detail) is what separates production-grade ML from a Kaggle submission.