Reported September 2026
Amazonsimulation

Grouped-Query Attention with a KV Cache

Reported by candidates from Amazon's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Amazon OA. Under 2s to a working solution.
Founder's read

The mistake that sinks a first attempt on this Amazon OA, reported in September 2026, is mixing up which key-value head each query head reads from. It's grouped-query attention with a KV cache, tagged as a design problem, but it's really careful simulation. You get query rows, cached keys and values, and new rows to append. Then you run softmax attention per head and format to six decimals. If you blank on the head mapping, StealthCoder is the safety net that runs invisibly during the live assessment and hands you the structure. The math is easy. The indexing is where people bleed time.

The problem

Compute one projection-free grouped-query scaled dot-product attention pass using an existing key-value cache plus newly generated key and value rows.
query has width queryHeads * headDim. Every cached or new key and value row has width kvHeads * headDim. Append the new rows after the cached rows for the attention calculation. queryHeads is divisible by kvHeads; each consecutive group of queryHeads / kvHeads query heads shares one key-value head.
For each query row and query head, compute scaled dot products against all cached and new key rows, divide by sqrt(headDim), apply maximum-subtracted softmax, and combine the corresponding value-head slices. Concatenate the query-head outputs and return each row as comma-separated values formatted to exactly six decimal places.

Function
groupedQueryAttentionWithCache(query: int[][], cachedKey: int[][], cachedValue: int[][], newKey: int[][], newValue: int[][], queryHeads: int, kvHeads: int) → String[]

Examples
Example 1
query = [[1]]
cachedKey = []
cachedValue = []
newKey = [[0],[1]]
newValue = [[10],[20]]
queryHeads = 1
kvHeads = 1
return = ["17.310586"]
With one head and no cache, this is ordinary scaled dot-product attention.
Example 2
query = [[1,0]]
cachedKey = [[1]]
cachedValue = [[4]]
newKey = [[0]]
newValue = [[8]]
queryHeads = 2
kvHeads = 1
return = ["5.075766,6.000000"]
Two query heads share one key-value head and attend over one cached and one new token.
Example 3
query = [[1,0],[0,1]]
cachedKey = [[1,0]]
cachedValue = [[10,20]]
newKey = [[0,1]]
newValue = [[30,40]]
queryHeads = 2
kvHeads = 2
return = ["15.378828,30.000000","20.000000,34.621172"]
Equal query and key-value head counts recover ordinary multi-head attention while still using the cache.

Constraints
1 <= query rows, cached rows + new rows <= 40
1 <= kvHeads <= queryHeads <= 16 and queryHeads is divisible by kvHeads.
1 <= headDim <= 16; all matrices are rectangular with the stated widths and cached/new key-value row counts match.
newKey and newValue are nonempty; cached matrices may be empty.
Every entry is an integer from -100 through 100.

Reported by candidates. Source: FastPrep

Pattern and pitfall

The trick is one integer: group size = queryHeads / kvHeads, and query head h uses kv head h / group size (integer division). Concatenate cachedKey and newKey (cache first), same for values. Then for each query row and head, slice headDim values, dot with the matching slice of every key row, divide by sqrt(headDim), subtract the max, exponentiate, normalize, and take the weighted sum of the value slices. The common pitfall is slicing the kv head with the query head index, which breaks whenever kvHeads is less than queryHeads. Others: forgetting cached matrices can be empty, skipping max subtraction, and printing with the wrong precision. Use a fixed six-decimal format and join with commas. Sizes are tiny, so plain nested loops are fine. If the indexing scrambles under pressure, StealthCoder can give you a clean reference to check against.

StealthCoder is the hedge for the one pattern you didn't drill. It runs invisibly during the screen share.

If this hits your live OA

You can drill Grouped-Query Attention with a KV Cache cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. If you're reading this with an OA window open, you're who this was built for.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Amazon's OA.

Amazon reuses patterns across OAs. If you're reading this with an OA window open, you're who this was built for. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Grouped-Query Attention with a KV Cache FAQ

How hard is the grouped-query attention OA really?+

Conceptually easy, mechanically fiddly. There's no clever algorithm, just head indexing, slicing, softmax, and formatting. Constraints are tiny (at most 40 rows, 16 heads), so brute-force loops pass. Most failures come from off-by-one slicing or wrong precision, not from complexity.

What's the one trick to get right?+

Map each query head to its kv head with queryHead / (queryHeads / kvHeads) using integer division. Consecutive query heads share a kv head. Then slice keys and values by kvHead * headDim, not by the query head index.

Do I need to handle an empty cache?+

Yes. Cached matrices may be empty, as Example 1 shows. Just concatenate the cached rows with the new rows and the empty case falls out naturally. Newly generated key and value rows are always nonempty, so you never softmax over zero keys.

Why subtract the max before softmax?+

It keeps exponentials from overflowing and the problem explicitly asks for it. Compute all scores for a head, find the max, exponentiate score minus max, divide by the sum. The result is mathematically the same but numerically safe.

How do I prepare for this in 48 hours?+

Write attention by hand once for a single head, then add the group mapping. Check it against Examples 2 and 3. Practice the output format too: six decimals, commas between heads, one string per query row. Don't spend time on optimizations, since nothing here needs them.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Amazon.

OA at Amazon?
Invisible during screen share
Get it