Reported September 2026
Amazonhash table

Group-Relative Policy Advantages

Reported by candidates from Amazon's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Amazon OA. Under 2s to a working solution.
Founder's read

Amazon reportedly served this one in September 2026, and it looks friendlier than it is. Group-Relative Policy Advantages is a hash-table grouping problem wearing an ML costume. You combine two reward arrays, bucket by group ID, then z-score each bucket on its own. The first-attempt killer is dividing by zero when a group's rewards are all equal, or normalizing across the whole array instead of per group. If you have the OA in a day or two, this is a 20-minute problem once you see it. StealthCoder sits invisibly on your screen as a safety net if you blank mid-assessment.

The problem

Compute group-relative advantages for sampled policy responses. Response i belongs to groupIds[i] and receives the integer reward correctnessRewards[i] + formatRewards[i].
Within each group, compute the population mean and population standard deviation of the combined rewards. The advantage of a response is (reward - mean) / standardDeviation. When a group's standard deviation is zero, every response in that group has advantage zero.
Return the advantages in original response order, formatted to exactly six decimal places.

Function
groupRelativeAdvantages(groupIds: int[], correctnessRewards: int[], formatRewards: int[]) → String[]

Examples
Example 1
groupIds = [0,0,1,1,1]
correctnessRewards = [1,0,1,1,0]
formatRewards = [1,0,0,1,0]
return = ["1.000000","-1.000000","0.000000","1.224745","-1.224745"]
Each prompt group is normalized independently after correctness and format rewards are combined.
Example 2
groupIds = [7,7]
correctnessRewards = [1,1]
formatRewards = [0,0]
return = ["0.000000","0.000000"]
Equal rewards have zero variance, so both normalized advantages are zero.

Constraints
2 <= groupIds.length <= 100000
All three arrays have the same length.
Each group ID appears at least twice.
-100 <= correctnessRewards[i], formatRewards[i] <= 100

Reported by candidates. Source: FastPrep

Pattern and pitfall

The trick is two passes. First pass: sum rewards, sum of squares and count per group ID in a hash map. Second pass: compute mean, then population variance as sumSq/n - mean^2, take the square root, and emit (r - mean)/std for each index in original order. The pitfall is the zero case. Floating point can leave variance as a tiny negative or tiny positive number when rewards are equal, so compute variance from integer sums exactly, or clamp at zero and check std == 0 before dividing. Use population variance, not sample (divide by n, not n-1). Format each value with exactly six decimals, and watch for negative zero printing as -0.000000. Add 0.0 or special-case zero. With 100000 elements, O(n) is plenty. If you freeze on the formatting or the zero edge during the live OA, StealthCoder is the hedge that hands you a clean solution.

If this hits your live OA and you blank, StealthCoder solves it in seconds, invisible to the proctor.

If this hits your live OA

You can drill Group-Relative Policy Advantages cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Amazon's OA.

Amazon reuses patterns across OAs. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Group-Relative Policy Advantages FAQ

What's the trick in Group-Relative Policy Advantages?+

Group by ID with a hash map, then normalize each group independently using population mean and standard deviation. Combine correctness and format rewards first. Track sum, sum of squares and count per group so you only need two linear passes over the data.

How do I avoid the divide-by-zero bug?+

Compute variance from integer sums: (n*sumSq - sum^2) / n^2. If the numerator is zero, std is zero and every advantage in that group is 0. Checking this with integers avoids floating point noise that makes equal rewards look like a tiny nonzero variance.

Population or sample standard deviation?+

Population. Divide by n, not n-1. Example 1 confirms it: group 0 has rewards 2 and 0, mean 1, population std 1, giving advantages 1 and -1. Sample std would give a different value and fail the expected output.

Why might my output show -0.000000?+

A tiny negative value or negative zero rounds to -0.000000 when formatted. If std is zero, output the literal string 0.000000. Otherwise, after computing the value, add 0.0 or check whether the rounded result is zero before formatting.

How hard is this really, and how do I prep in 48 hours?+

It's easy to medium. The logic is simple, the traps are edge cases and formatting. Write it once from scratch, test the two given examples, then test a group with all equal rewards and a group with negative rewards. That covers nearly every failure.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Amazon.

OA at Amazon?
Invisible during screen share
Get it