Gradient Descent Linear Regression from Scratch
Reported by candidates from TikTok's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
TikTok reported this one in September 2026, and it's not a typical array-and-hash-map OA. You implement batch gradient descent for linear regression with no libraries, and the whole thing hinges on plain arrays: a 2D array of features, a 1D array of targets, and a 1D array of thetas that you update in place. You're handed the initial bias and thetas, so nothing is random. If you've got the OA invite and the formulas look scary, relax. It's a loop with arithmetic. StealthCoder is the invisible hedge if you blank on the update order during the live OA.
The problem
Your task is to implement parts of the Gradient Descent optimization algorithm from scratch (i.e., without importing any libraries or packages). You will apply this algorithm for linear regression, finding the coefficients for the following equation: y = b + θ₁x₁ + θ₂x₂ +... + θₘxₘ As a reminder, Gradient Descent comprises following steps: Randomly set initial values of bias b and thetas θ~i. Calculate predicted value ŷ. Calculate partial derivative of a cost function with respect to bias.d_b = ∂/∂b (1/n) Σᵢ₌₁ⁿ(yᵢ - ŷᵢ)² = -2/n Σᵢ₌₁ⁿ(yᵢ - ŷᵢ) Calculate partial derivatives of a cost function with respect to thetas.d_θₖ = ∂/∂θₖ (1/n) Σᵢ₌₁ⁿ(yᵢ - ŷᵢ)² = -2/n Σᵢ₌₁ⁿxᵢᵏ(yᵢ - ŷᵢ) Update bias and thetas (α is a learning rate).b = b - αd_b θₖ = θₖ - αd_θₖ Repeat steps 2-5 for the iterations times. FastPrep execution adapter The source image is cropped before the complete callable and output contract. For deterministic grading, FastPrep receives the initial bias and theta values that represent the source's randomly set initial state, then returns the final bias followed by the final coefficients. This adapter does not replace the source requirement to begin from initialized values. Function fitLinearRegression(xTrain: double[][], yTrain: double[], initialBias: double, initialThetas: double[], learningRate: double, iterations: int) → double[] Examples Example 1 xTrain = [[1], [2]] yTrain = [3, 5] initialBias = 0 initialThetas = [0] learningRate = 0.1 iterations = 1 return = [0.8, 1.3] FastPrep-authored runnable example (not shown in the source image): The initial predictions are both 0, so the residuals are -3 and -5. Therefore dBias = -8 and dTheta[0] = -13. One update with learning rate 0.1 produces bias 0.8 and coefficient 1.3. Constraints FastPrep execution-adapter constraints (not shown in the source image): xTrain.length == yTrain.length and the arrays are non-empty. Every row of xTrain has exactly initialThetas.length features. learningRate > 0 and iterations >= 0. All inputs and expected final parameters are finite numbers.
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick is the update timing. Each iteration, compute predictions for every row using the current bias and thetas, then compute dBias = -2/n * sum(y - yhat) and each dTheta[k] = -2/n * sum(x[i][k] * (y[i] - yhat[i])). Only after all gradients are done do you update bias and thetas. That's the pitfall: updating thetas mid-loop changes predictions and gives wrong numbers. Compute all gradients first, then apply them. Check with the example: residuals -3 and -5, dBias = -8, dTheta = -13, so you get 0.8 and 1.3. Return bias first, then the thetas. Handle iterations = 0 by returning the initial values. Watch the 2/n factor and sign. If the live OA freezes you on indexing, StealthCoder can give you the skeleton while you sanity-check against the example.
The honest play: practice the pattern, and have StealthCoder ready for the one you didn't see coming.
You can drill Gradient Descent Linear Regression from Scratch cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built for the candidate who saw this exact problem leak two days before his OA and wondered if anyone had a play.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass TikTok's OA.
TikTok reuses patterns across OAs. Built for the candidate who saw this exact problem leak two days before his OA and wondered if anyone had a play. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Gradient Descent Linear Regression from Scratch FAQ
How hard is this TikTok gradient descent question really?+
Easy on algorithms, medium on care. There's no clever data structure, just nested loops over rows and features. Most failures come from sign errors, forgetting the 2/n factor, or updating parameters before all gradients are computed. Walk the one-iteration example by hand and you'll catch all of it.
What's the trick to getting the right numbers?+
Use batch updates. Per iteration, compute predictions from the current parameters, accumulate the bias gradient and every theta gradient in separate variables, then update everything at once. Never mutate thetas while you're still summing gradients, or later features see shifted predictions.
Do I need to initialize random values?+
No. The function signature gives you initialBias and initialThetas, so grading is deterministic. Start from those values, copy the thetas array so you don't mutate the input, and run the loop. Zero iterations should return the initial values as is.
What should the return array look like?+
A double array with the final bias at index 0, followed by the final thetas in order. For the example with one feature, you return [0.8, 1.3]. Length is thetas length plus one. Double-check the order before submitting, since swapping it silently fails every test.
How do I prepare for this in 48 hours?+
Write the loop once from memory, then test it on the example by hand: residuals -3 and -5, dBias -8, dTheta -13. Practice a two-feature case too, so your inner loop over features is solid. That covers about everything this problem can throw at you.