Subscribe

Machine Learning Data Preparation

A comprehensive, production-grade guide to data selection, preprocessing, scaling, categorical encoding, and leakage prevention in machine learning pipelines.

AJ
Atul Jha Systems & AI Researcher
•
18 min read
•
w := w - α · ∇L(w) ŷ = X · w + b
Machine Learning
18m
Machine Learning Data Preparation
VERIFIED BLUEPRINT SYS // ML-REV2

1. Why Data Preparation Dictates Model Performance

In production machine learning systems, model architectures and hyperparameter tuning rarely account for more than 10% to 15% of performance variations. The remaining 85% is governed by data quality, feature representation, and pipeline hygiene.

Raw operational data is messy: it contains unrecorded events, system timestamps with irregular clock skew, sensor noise, categorical labels with high cardinality, and non-Gaussian statistical distributions. Feeding raw or improperly transformed inputs into gradient descent optimizers or tree-based splits degrades convergence rates and introduces silent failures.

+-------------------------------------------------------------------------------+
|                      THE PRODUCTION DATA PREPARATION PIPELINE                |
|                                                                               |
|  [Raw Data Sources]                                                           |
|          |                                                                    |
|          v                                                                    |
|  [Phase 1: Data Selection] -----> Class Imbalance & Temporal Splits          |
|          |                                                                    |
|          v                                                                    |
|  [Phase 2: Cleaning & Imputation] -> Drop Dupes, Iterative Impute, IQR Outliers|
|          |                                                                    |
|          v                                                                    |
|  [Phase 3: Transformations] ----> Scaling, One-Hot/Target Encoding, Cyclical  |
|          |                                                                    |
|          v                                                                    |
|  [Phase 4: Leakage Isolation] --> Fit ONLY on Train -> Transform Test/Prod    |
+-------------------------------------------------------------------------------+

2. Phase 1: Strategic Data Selection & Problem Formulation

Data selection is not merely collecting as many records as possible; it is identifying the subset of data that accurately reflects the production distribution at inference time.

Temporal Splitting vs. Random Splitting

If your data has any temporal sequence (user clicks, sensor streams, transactions), random k-fold cross-validation is an anti-pattern. Random splits cause future information to leak into past predictions:

import pandas as pd
import numpy as np

def temporal_train_test_split(df: pd.DataFrame, time_col: str, train_ratio: float = 0.8):
    """
    Splits chronological data along a strict temporal boundary to prevent lookahead leakage.
    """
    df_sorted = df.sort_values(by=time_col).reset_index(drop=True)
    split_idx = int(len(df_sorted) * train_ratio)
    
    train_df = df_sorted.iloc[:split_idx].copy()
    test_df = df_sorted.iloc[split_idx:].copy()
    
    return train_df, test_df

Addressing Severe Class Imbalance

For anomaly detection and fraud classification, skewed class distributions (>99:1> 99:1) require informed sampling strategies:

  1. Never downsample your test set: Your test set must mirror the authentic production distribution.
  2. Stratified Splitting: Use StratifiedKFold or train_test_split(..., stratify=y) to guarantee that rare classes are proportionally preserved across all folds.
  3. Synthetic Resampling (SMOTE / ADASYN): Apply oversampling only to the training partitions during cross-validation loops, never before splitting.

3. Phase 2: Systematic Preprocessing & Imputation

Preprocessing transforms corrupt, incomplete, or noisy records into consistent matrix representations.

Handling Missing Values with Discipline

Dropping missing rows indiscriminately reduces sample size and introduces selection bias when values are Missing Not at Random (MNAR).

StrategyWhen to UsePitfall / Drawback
Median / Mode ImputationBaseline pipelines, low missingness (\<2\< 2%)Distorts feature variance and suppresses covariance
KNN ImputationFeature relationships are nonlinear, moderate dataset sizeHigh inference latency (O(N⋅D)O(N \cdot D) per sample)
Iterative Imputer (MICE)Multivariable dependencies, high-stakes tabular modelsComputationally heavy during training phase
Missingness IndicatorThe fact that a value is missing is informativeDoubles feature dimension for sparse inputs
from sklearn.impute import SimpleImputer, IterativeImputer
from sklearn.compose import ColumnTransformer

# Production strategy: Impute numerical features using median + add missing indicator
num_imputer = SimpleImputer(strategy='median', add_indicator=True)
cat_imputer = SimpleImputer(strategy='most_frequent')

Robust Outlier Detection

Outliers can severely distort linear boundaries and distance metrics. The Interquartile Range (IQR) technique filters extreme values without assuming a Gaussian distribution:

  • IQR Calculation: IQR = Q3 - Q1
  • Lower Outlier Fence: Q1 - 1.5 * IQR
  • Upper Outlier Fence: Q3 + 1.5 * IQR
def cap_outliers_iqr(series: pd.Series, factor: float = 1.5) -> pd.Series:
    """Caps numerical outliers using Tukey's fences to avoid losing rows."""
    q25 = series.quantile(0.25)
    q75 = series.quantile(0.75)
    iqr = q75 - q25
    lower_limit = q25 - (factor * iqr)
    upper_limit = q75 + (factor * iqr)
    return series.clip(lower=lower_limit, upper=upper_limit)

4. Phase 3: Feature Transformation & Scaling

Feature transformations adjust scale, variance, and representation so that models can learn optimal weights efficiently.

1. Scaling Numerical Features

Neural networks, Support Vector Machines, Logistic Regression, and distance-based clustering algorithms require uniform scales. Decision trees and Random Forests are scale-invariant, but scaling remains good practice for pipeline interoperability.

  • StandardScaler: Scales to mean 0, variance 1. Ideal for Gaussian-distributed features.
  • RobustScaler: Centers using the median and scales using IQR. Ideal when unpruned anomalies exist.
  • MinMaxScaler: Scales bounds strictly to [0, 1]. Mandatory for bounded image pixels or bounded neural activations.

2. Categorical Encodings

  • One-Hot Encoding: Best for nominal categories with low cardinality (< 15 unique levels).
  • Target / Frequency Encoding: Best for high-cardinality features (e.g. zip codes, device IDs). Ensure smoothing (m-estimate) is applied to prevent target leakage.

3. Trigonometric Cyclical Encodings

For features like hour of day (0 to 23) or month (1 to 12), linear values fail because 23:00 and 00:00 are adjacent, yet separated by 23 units. We project them onto a 2D unit circle:

  • Sine coordinate: x_sin = sin(2 * pi * t / T)
  • Cosine coordinate: x_cos = cos(2 * pi * t / T)
def encode_cyclical_feature(df: pd.DataFrame, col: str, period: int):
    """Encodes periodic time cycles into continuous orthogonal sine/cosine features."""
    df[f'{col}_sin'] = np.sin(2 * np.pi * df[col] / period)
    df[f'{col}_cos'] = np.cos(2 * np.pi * df[col] / period)
    return df.drop(columns=[col])

5. End-to-End Production Pipeline Implementation

The following complete, runnable Python script constructs a leak-proof preprocessing and classification pipeline using scikit-learn:

import numpy as np
import pandas as pd
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder, RobustScaler
from sklearn.impute import SimpleImputer
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import classification_report

def build_production_pipeline(numerical_cols: list[str], categorical_cols: list[str]) -> Pipeline:
    """
    Constructs an encapsulated, leak-free Scikit-Learn data preparation pipeline.
    """
    # 1. Numerical branch: Imputation -> Robust Scaling
    numeric_transformer = Pipeline(steps=[
        ('imputer', SimpleImputer(strategy='median', add_indicator=True)),
        ('scaler', RobustScaler()),
    ])

    # 2. Categorical branch: Frequent imputation -> One-Hot encoding
    categorical_transformer = Pipeline(steps=[
        ('imputer', SimpleImputer(strategy='most_frequent')),
        ('onehot', OneHotEncoder(handle_unknown='ignore', sparse_output=False)),
    ])

    # 3. Combine column transformations
    preprocessor = ColumnTransformer(
        transformers=[
            ('num', numeric_transformer, numerical_cols),
            ('cat', categorical_transformer, categorical_cols),
        ],
        remainder='drop',
    )

    # 4. Bind preprocessor with estimator
    pipeline = Pipeline(steps=[
        ('preprocessor', preprocessor),
        ('classifier', HistGradientBoostingClassifier(random_state=42)),
    ])

    return pipeline

# Demonstration
if __name__ == '__main__':
    # Synthesize sample dataset
    X_raw, y = make_classification(n_samples=2000, n_features=6, random_state=42)
    df = pd.DataFrame(X_raw, columns=['num_1', 'num_2', 'num_3', 'num_4', 'cat_1', 'cat_2'])
    
    # Introduce discrete categories and missingness
    df['cat_1'] = pd.cut(df['cat_1'], bins=3, labels=['low', 'med', 'high']).astype(str)
    df['cat_2'] = pd.cut(df['cat_2'], bins=2, labels=['type_a', 'type_b']).astype(str)
    df.loc[::10, 'num_1'] = np.nan
    
    # Train / Test split
    X_train, X_test, y_train, y_test = train_test_split(
        df, y, test_size=0.2, random_state=42, stratify=y
    )
    
    pipeline = build_production_pipeline(
        numerical_cols=['num_1', 'num_2', 'num_3', 'num_4'],
        categorical_cols=['cat_1', 'cat_2'],
    )
    
    # Fit strictly on train; transform and evaluate on test
    pipeline.fit(X_train, y_train)
    predictions = pipeline.predict(X_test)
    
    print(classification_report(y_test, predictions))

6. The Golden Rule: Preventing Data Leakage

Data leakage is the silent killer of predictive performance. When data leakage occurs, validation metrics are unrealistically high (9999%+ accuracy), but real-world performance collapses in production.

WRONG (LEAKAGE):
  Raw Dataset ---> [Fit Scaler on Entire Dataset] ---> [Split Train / Test] ❌

CORRECT (ISOLATED):
  Raw Dataset ---> [Split Train / Test]
                         |
                         +---> Train ---> [Fit Scaler] ---> [Transform Train]
                         |                      |
                         +---> Test  ----------> (Only Transform with Train Scaler) ✅

Critical Leakage Safeguards

  1. Never fit a scaler on test data: Summary statistics (mean, variance, min, max) must be derived exclusively from the training partition.
  2. Encapsulate in Pipelines: Never perform preprocessing manually in standalone Jupyter notebook cells. Always wrap transforms in a Pipeline or ColumnTransformer.
  3. Target Leakage: Ensure features do not contain future state data (e.g. including refund_timestamp when predicting is_fraudulent_purchase).

7. Production Readiness Checklist

Before promoting any machine learning data preparation pipeline to staging or production, verify this checklist:

  • [x] Temporal Validation: Time-series datasets use rolling window or cutoff date splits rather than uniform random sampling.
  • [x] Leakage Isolation: Every transformer fits on training sets only; validation and test sets are strictly transformed.
  • [x] Missingness Policy: Missing indicators are retained if missingness conveys signal; imputation strategy is justified.
  • [x] Category Drift Defense: Categorical encoders specify handle_unknown='ignore' or fallback to an unknown bucket for novel production labels.
  • [x] Memory & Types: Integer and float types are downcast (float32 vs float64) to prevent memory spikes on inference workers.
  • [x] Serialization Artifacts: Pipeline is serialized as a single immutable artifact (joblib.dump / ONNX) ensuring identical preprocessing across training and inference.

Interactive Code Lab: Machine Learning Data Preparation

Machine LearningMatched to lesson

Solves optimal weights w = (X^T X)^-1 X^T y from scratch.

Labs:
Ordinary Least Squares Closed-Form Normal Equation
Pyodide Wasm
Interactive Challenge: Modify code inputs, click Run Code to execute live in WebAssembly.
+25 XP Reward
Wasm Terminal Output

Click Run Code to execute this algorithm in the browser sandbox.

How did you find this blueprint?

Tap a reaction to share instant feedback with the engineering team.

Frequently Asked Questions

Frequently Asked Questions

Architecture

Feature Engineering

Citations & Recommended References

References

AJ
Written by
Atul Jha

AI Researcher and Systems Engineer focusing on production machine learning pipelines, transformer architectures, and performant Python runtime internals.

View Profile →
Share Blueprint:

Discussion & Community Thoughts

0%
Notification