Reported June 2026
Luma AImatrix

Impute Missing Values and Normalize Columns

Reported by candidates from Luma AI's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Luma AI OA. Under 2s to a working solution.
Founder's read

Luma AI reported this one in June 2026, and it's less a puzzle than a clean data-prep task. Zeros are missing, fill them with the column's nonzero mean, then z-score each column with population stats. If you've got an OA coming, the risk here isn't the idea. It's small details like which mean you use and dividing by n instead of n-1. The grid is tiny, at most 50 by 50, so nothing here needs cleverness. It's column-wise array work with a tolerance of 1e-5. If you blank on the setup during the live assessment, StealthCoder can sit invisibly as a backup.

The problem

In each column of dataset, treat zero as missing and replace it with the mean of that column's nonzero values. Then z-score normalize the imputed column using its population mean and population standard deviation.

Function
imputeAndNormalize(dataset: double[][]) → double[][]

Examples
Example 1
dataset = [[1,2,0],[0,1,1],[5,6,5]]
return = [[-1.224744871391589,-0.4629100498862757,0.0],[0.0,-0.9258200997725514,-1.224744871391589],[1.224744871391589,1.3887301496588271,1.224744871391589]]
Zeros become the nonzero column means before population z-score normalization.

Constraints
2 <= dataset.length <= 50.
1 <= dataset[i].length <= 50 and all rows have equal length.
Every column has at least two different nonzero values.
Results use absolute tolerance 1e-5.

Reported by candidates. Source: FastPrep

Pattern and pitfall

Brute force isn't a concern. With at most 50 rows and 50 columns, a few passes per column cost almost nothing, so don't look for a trick. Treat each column independently. Pass one: sum the nonzero values and count them, then compute the mean from that. Pass two: replace zeros with that mean. Pass three: compute the mean of the imputed column, then the population variance (divide by n), take the square root, and normalize. The pitfalls are easy to hit. Imputing with the mean of all values including zeros is wrong. Using sample standard deviation (n-1) will break the expected output. Mutating the input while reading it can also bite you. The constraints guarantee every column has two different nonzero values, so the standard deviation is never zero and you can skip that guard. If you freeze on the live OA, StealthCoder is the hedge that gives you the column loop.

StealthCoder is the hedge for the one pattern you didn't drill. It runs invisibly during the screen share.

If this hits your live OA

You can drill Impute Missing Values and Normalize Columns cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. If you're reading this with an OA window open, you're who this was built for.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Luma AI's OA.

Luma AI reuses patterns across OAs. If you're reading this with an OA window open, you're who this was built for. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Impute Missing Values and Normalize Columns FAQ

How hard is the Luma AI impute and normalize problem really?+

Easy on algorithm, moderate on care. There's no data structure or pattern to discover. You loop over columns, impute, and normalize. Most failures come from using the wrong mean or the wrong variance formula, not from complexity. Expect it to take a few minutes if you stay organized.

What's the trick to getting the imputation right?+

Compute the mean over nonzero values only, meaning sum of nonzeros divided by count of nonzeros. Then replace each zero with it. Don't include zeros in that first mean. After imputing, the column mean changes, so recompute it before normalizing. Mixing those two means is the classic bug.

Population or sample standard deviation?+

Population. Divide the sum of squared deviations by n, the number of rows, not n-1. The problem says population mean and population standard deviation explicitly. Example 1 confirms it: column 0 becomes about -1.2247, 0, 1.2247, which only works with the n divisor.

Do I need to worry about division by zero or precision?+

Not for zero deviation. The constraints promise every column has at least two different nonzero values, so the standard deviation is positive after imputation. Precision is forgiving too, since results are checked with an absolute tolerance of 1e-5. Plain doubles are fine.

How do I prepare for this in 48 hours?+

Write the function once from scratch in your language, working column by column. Test it against the example grid and check that column 0 gives -1.2247, 0.0, 1.2247. Then rehearse similar column-wise stats tasks: mean, variance, min-max scaling. Know your 2D array indexing cold, since you'll be iterating columns, not rows.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Luma AI.

OA at Luma AI?
Invisible during screen share
Get it