Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Feature Engineering (7 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Spotting the Leak in a Target Encoding Pipeline Permalink →

A colleague encodes a merchant_id column (40,000 distinct values) for a fraud model like this, run on the full training set before the train/test split:

merchant_rate = df.groupby("merchant_id")["is_fraud"].transform("mean")
df["merchant_te"] = merchant_rate

Cross-validated AUC on this feature alone is 0.97. In production, AUC drops to 0.71.

  1. Explain exactly why this code leaks the label, using a merchant with 3 transactions as a concrete example.
  2. Rewrite the encoding so it is safe to use inside 5-fold cross-validation.
  3. Why does smoothing (shrinking toward the global rate) matter here even after you fix the leak?

Share this question

Advanced Open Pro

Point-in-Time Correctness for a Churn Model

Unlock this question →
Intermediate Open Pro

Choosing Encodings by Cardinality and Model Family

Unlock this question →
Beginner Open Pro

Cyclical Encoding for a Ride-Demand Model

Unlock this question →
Intermediate Open Pro

Missing Values: Indicator, Impute, or Let the Model Handle It

Unlock this question →
Advanced Open Pro

Feature Selection Inside vs Outside Cross-Validation

Unlock this question →
Intermediate Open Free

Does Standardizing 'income' Change a Gradient-Boosted Tree's Predictions? Permalink →

A data scientist standardizes every numeric feature (z-score: subtract the mean, divide by the standard deviation) before fitting two models on the same data: a logistic regression and a gradient-boosted tree ensemble (LightGBM). Their reasoning: "scaling is always good practice, it can only help."

For the logistic regression, this is true — unscaled features with very different ranges distort the effective regularization strength per coefficient and can slow optimizer convergence. For the gradient-boosted tree model specifically, what effect does standardizing a single numeric feature (a strictly monotonic transform: same relative order, just a different scale) have on the trained model's actual predictions?

Share this question

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.