Stochastic Gradient Descent with Momentum
Reported by candidates from Meta's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
Meta reported this one in September 2026, and it's a different animal from the usual graph and DP fare. It's a plain simulation of SGD with momentum. The solution hinges on two flat arrays, one for weights and one for velocity, updated in place for each gradient row. If you've got an OA coming, the math looks scary but it's two lines of arithmetic inside a loop. The real work is not mutating the input and returning the right two-row shape. If your brain freezes on ML jargon, StealthCoder is the quiet safety net running during the live OA.
The problem
Simulate stochastic gradient descent with momentum for one parameter vector. Start with zero velocity. For each gradient row in order, update every coordinate using: velocity = momentum * velocity + gradient weight = weight - learningRate * velocity Return a two-row matrix: the final weights followed by the final velocity. Do not modify the input arrays. Function sgdWithMomentum(initialWeights: double[], gradients: double[][], learningRate: double, momentum: double) → double[][] Examples Example 1 initialWeights = [1,2] gradients = [[0.5,-1],[1,1]] learningRate = 0.1 momentum = 0.9 return = [[0.805,2.09],[1.45,0.1]] After the first step the velocity is [0.5,-1] and weights are [0.95,2.1]. The second velocity is [1.45,0.1]. Example 2 initialWeights = [0] gradients = [[2],[2]] learningRate = 0.5 momentum = 0 return = [[-2],[2]] With zero momentum, each step uses only the current gradient. Example 3 initialWeights = [3,-4] gradients = [[0,0]] learningRate = 0.25 momentum = 0.5 return = [[3,-4],[0,0]] A zero gradient and zero initial velocity leave the weights unchanged. Constraints 1 <= initialWeights.length <= 100. 1 <= gradients.length <= 1000, and every gradient row has the parameter dimension. Every initial weight and gradient value is finite and has absolute value at most 10^6. 0 < learningRate <= 1 and 0 <= momentum < 1.
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick is to ignore the ML framing. Copy initialWeights into a new array, create a velocity array of zeros with the same length, then loop over each gradient row. For every coordinate, set velocity = momentum * velocity + gradient first, then weight = weight - learningRate * velocity using the NEW velocity. Order matters here. Updating the weight with the old velocity breaks Example 1. Return [weights, velocity] as a two-row matrix. Pitfalls: mutating the input array, which the problem forbids, and mixing up the order of updates. Also watch float output, since 0.805 style values may differ by tiny epsilons, so don't round unless asked. Complexity is O(n * d) time and O(d) space, trivial against the constraints. If you blank on the setup during the live OA, StealthCoder can surface this loop instantly, but you can honestly write it from memory.
Drill it cold or hedge it with StealthCoder. Either way, don't walk into the OA hoping you remember the trick.
You can drill Stochastic Gradient Descent with Momentum cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Made for the candidate who got the OA invite this morning and has 72 hours, not six months.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass Meta's OA.
Meta reuses patterns across OAs. Made for the candidate who got the OA invite this morning and has 72 hours, not six months. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Stochastic Gradient Descent with Momentum FAQ
How hard is the Meta SGD with momentum question really?+
Easy once you read past the ML terms. It's a simulation with two update lines inside a nested loop. No tricky algorithm, no data structure beyond arrays. Most of the risk is careless mistakes like wrong update order or mutating the input.
What's the trick to getting the update order right?+
Update velocity first, using the old velocity times momentum plus the gradient. Then update the weight using that new velocity. Check Example 1: after step one, velocity is [0.5,-1] and weights are [0.95,2.1]. If your numbers match, your order is right.
Do I need to worry about modifying the input arrays?+
Yes, the problem says not to. Clone initialWeights into a fresh array before looping, and never write into the gradients rows. Create a new velocity array of zeros. Return a new two-row matrix containing the final weights and the final velocity.
Is this simulation pattern still asked in 2026 OAs?+
This Meta report from September 2026 says yes. Simulation and ML-flavored implementation questions show up alongside classic algorithm problems. They reward careful reading over clever tricks, so expect more of them rather than fewer.
How do I prepare for this in 48 hours?+
Write the solution once from scratch and run all three examples by hand. Check the zero momentum case and the zero gradient case. Then practice a couple of other plain simulation problems so you stay calm when the prompt looks unfamiliar but the logic is simple.