Linear Layer Forward and Backward Passes
Reported by candidates from Meta's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
Meta reportedly put a linear layer forward and backward pass in front of candidates in September 2026, and the naive version breaks on one detail. It looks like ML trivia, but it's plain array math with a strict output shape. Five rows, ragged lengths, no rounding, no parameter updates. If you blank on the gradient algebra, the whole answer is wrong even when the loop is clean. StealthCoder is the safety net that runs invisibly during the live OA if your mind goes empty on the chain rule. The rest of this page is the script.
The problem
Implement the forward and backward passes of a single-output linear layer under half mean-squared error. For each sample i, compute prediction[i] = bias + sum_j(features[i][j] * weights[j]). The scalar loss is sum_i((prediction[i] - targets[i])^2) / (2 * B), where B is the sample count. Return a matrix whose rows may have different lengths: Row 0: all predictions. Row 1: the single loss value. Row 2: the gradient with respect to the weights. Row 3: the single gradient with respect to the bias. Row 4 + i: the gradient with respect to feature row i. Return analytic, unrounded values and do not update the parameters. Function linearLayerForwardBackward(features: double[][], weights: double[], bias: double, targets: double[]) → double[][] Examples Example 1 features = [[1,2],[3,4]] weights = [2,-1] bias = 0.5 targets = [0,1] return = [[0.5,2.5],[0.625],[2.5,3.5],[1.0],[0.5,-0.25],[1.5,-0.75]] The prediction errors are 0.5 and 1.5. Dividing by the batch size gives output gradients 0.25 and 0.75, from which every returned gradient follows. Example 2 features = [[1,-1]] weights = [0,0] bias = 0 targets = [0] return = [[0],[0],[0,0],[0],[0,0]] A zero prediction equal to the target produces zero loss and zero gradients. Example 3 features = [[2]] weights = [3] bias = -1 targets = [1] return = [[5],[8],[8],[4],[12]] The error is 4, so the half-squared loss is 8, the weight gradient is 4 * 2 = 8, and the feature gradient is 4 * 3 = 12. Constraints 1 <= features.length <= 64. 1 <= features[i].length <= 64, and the matrix is rectangular. weights.length = features[i].length and targets.length = features.length. All input values are finite and lie in [-10, 10].
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick is one quantity: err[i] = prediction[i] - targets[i]. Divide by B once to get d[i] = err[i] / B. Then everything falls out. Weight gradient j is sum over i of d[i] * features[i][j]. Bias gradient is sum of d[i]. Feature gradient for row i, column j is d[i] * weights[j]. Loss is sum of err squared, divided by 2B. The common pitfall is dividing by B in the wrong place, or forgetting the 2 that cancels against the square's derivative. Check Example 3: B is 1, error 4, loss 8, weight grad 8, feature grad 12. The other trap is the return shape. Rows 1 and 3 hold a single value each, so wrap them in arrays. Row 4+i must be a fresh array per sample. If you freeze, StealthCoder can hand you the working solution in the live OA. Don't round anything.
StealthCoder is the hedge for the one pattern you didn't drill. It runs invisibly during the screen share.
You can drill Linear Layer Forward and Backward Passes cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. If you're reading this with an OA window open, you're who this was built for.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass Meta's OA.
Meta reuses patterns across OAs. If you're reading this with an OA window open, you're who this was built for. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Linear Layer Forward and Backward Passes FAQ
How hard is the Meta linear layer problem really?+
Easy on algorithms, annoying on bookkeeping. It's nested loops and three formulas. The difficulty is getting the batch scaling right and building a ragged output matrix with exactly 4 + B rows. Verify against Example 1 by hand before you submit.
What's the trick to the backward pass?+
Compute d[i] = (prediction[i] - targets[i]) / B once. Weight gradient is the sum of d[i] times features[i][j]. Bias gradient is the sum of d[i]. Feature gradient is d[i] times weights[j]. No other derivation is needed.
What edge case breaks a naive solution?+
Output shape and scaling. Rows 1 and 3 must be one-element arrays, not bare numbers. Dividing by B instead of 2B in the loss, or applying 1/B twice in the gradients, also breaks it. Example 2 with all zeros catches sloppy initialization.
Should I round the outputs?+
No. The problem says return analytic, unrounded values and don't update the parameters. Use doubles throughout and return raw results. Rounding can fail the comparison, and mutating weights or bias is explicitly not wanted.
How do I prepare in 48 hours?+
Write the function from scratch twice. Use the three examples as tests, especially Example 1 with its 5-plus rows. Practice allocating a jagged double matrix in your language. Memorize the formulas for loss, weight gradient, bias gradient, and feature gradient.