Reported September 2026
Metamath

Linear Layer Forward and Backward Passes

Reported by candidates from Meta's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Meta OA. Under 2s to a working solution.
Founder's read

Meta reportedly put a linear layer forward and backward pass in front of candidates in September 2026, and the naive version breaks on one detail. It looks like ML trivia, but it's plain array math with a strict output shape. Five rows, ragged lengths, no rounding, no parameter updates. If you blank on the gradient algebra, the whole answer is wrong even when the loop is clean. StealthCoder is the safety net that runs invisibly during the live OA if your mind goes empty on the chain rule. The rest of this page is the script.

The problem

Implement the forward and backward passes of a single-output linear layer under half mean-squared error.
For each sample i, compute prediction[i] = bias + sum_j(features[i][j] * weights[j]). The scalar loss is sum_i((prediction[i] - targets[i])^2) / (2 * B), where B is the sample count.
Return a matrix whose rows may have different lengths:
Row 0: all predictions.
Row 1: the single loss value.
Row 2: the gradient with respect to the weights.
Row 3: the single gradient with respect to the bias.
Row 4 + i: the gradient with respect to feature row i.
Return analytic, unrounded values and do not update the parameters.

Function
linearLayerForwardBackward(features: double[][], weights: double[], bias: double, targets: double[]) → double[][]

Examples
Example 1
features = [[1,2],[3,4]]
weights = [2,-1]
bias = 0.5
targets = [0,1]
return = [[0.5,2.5],[0.625],[2.5,3.5],[1.0],[0.5,-0.25],[1.5,-0.75]]
The prediction errors are 0.5 and 1.5. Dividing by the batch size gives output gradients 0.25 and 0.75, from which every returned gradient follows.
Example 2
features = [[1,-1]]
weights = [0,0]
bias = 0
targets = [0]
return = [[0],[0],[0,0],[0],[0,0]]
A zero prediction equal to the target produces zero loss and zero gradients.
Example 3
features = [[2]]
weights = [3]
bias = -1
targets = [1]
return = [[5],[8],[8],[4],[12]]
The error is 4, so the half-squared loss is 8, the weight gradient is 4 * 2 = 8, and the feature gradient is 4 * 3 = 12.

Constraints
1 <= features.length <= 64.
1 <= features[i].length <= 64, and the matrix is rectangular.
weights.length = features[i].length and targets.length = features.length.
All input values are finite and lie in [-10, 10].

Reported by candidates. Source: FastPrep

Pattern and pitfall

The trick is one quantity: err[i] = prediction[i] - targets[i]. Divide by B once to get d[i] = err[i] / B. Then everything falls out. Weight gradient j is sum over i of d[i] * features[i][j]. Bias gradient is sum of d[i]. Feature gradient for row i, column j is d[i] * weights[j]. Loss is sum of err squared, divided by 2B. The common pitfall is dividing by B in the wrong place, or forgetting the 2 that cancels against the square's derivative. Check Example 3: B is 1, error 4, loss 8, weight grad 8, feature grad 12. The other trap is the return shape. Rows 1 and 3 hold a single value each, so wrap them in arrays. Row 4+i must be a fresh array per sample. If you freeze, StealthCoder can hand you the working solution in the live OA. Don't round anything.

StealthCoder is the hedge for the one pattern you didn't drill. It runs invisibly during the screen share.

If this hits your live OA

You can drill Linear Layer Forward and Backward Passes cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. If you're reading this with an OA window open, you're who this was built for.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Meta's OA.

Meta reuses patterns across OAs. If you're reading this with an OA window open, you're who this was built for. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Linear Layer Forward and Backward Passes FAQ

How hard is the Meta linear layer problem really?+

Easy on algorithms, annoying on bookkeeping. It's nested loops and three formulas. The difficulty is getting the batch scaling right and building a ragged output matrix with exactly 4 + B rows. Verify against Example 1 by hand before you submit.

What's the trick to the backward pass?+

Compute d[i] = (prediction[i] - targets[i]) / B once. Weight gradient is the sum of d[i] times features[i][j]. Bias gradient is the sum of d[i]. Feature gradient is d[i] times weights[j]. No other derivation is needed.

What edge case breaks a naive solution?+

Output shape and scaling. Rows 1 and 3 must be one-element arrays, not bare numbers. Dividing by B instead of 2B in the loss, or applying 1/B twice in the gradients, also breaks it. Example 2 with all zeros catches sloppy initialization.

Should I round the outputs?+

No. The problem says return analytic, unrounded values and don't update the parameters. Use doubles throughout and return raw results. Rounding can fail the comparison, and mutating weights or bias is explicitly not wanted.

How do I prepare in 48 hours?+

Write the function from scratch twice. Use the three examples as tests, especially Example 1 with its 5-plus rows. Practice allocating a jagged double matrix in your language. Memorize the formulas for loss, weight gradient, bias gradient, and feature gradient.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Meta.

OA at Meta?
Invisible during screen share
Get it