Reported January 2021
Scale AIsorting

Adjacent Order-Statistic Gap Distributions

Reported by candidates from Scale AI's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Scale AI OA. Under 2s to a working solution.
Founder's read

Scale AI sent this one around in January 2021, and it looks harder than it is. Up to 100000 rows, ten values each, and you need two gaps per row. The brute-force fear is real if you think about sorting the whole matrix, but you never do that. Each row is its own tiny sort. It's a sorting problem dressed up in statistics language, with a visualization footnote you can ignore. If the OA clock is ticking and you blank on the indexing, StealthCoder runs invisibly as a safety net. Most people won't need it here.

The problem

For this exercise, the random draws are supplied as a matrix samples. Each row contains ten finite values in [0,1].
For every row:
Sort a copy in ascending order.
Compute gapA = sorted[4] - sorted[3], the gap between the fourth and fifth order statistics.
Compute gapB = sorted[5] - sorted[4], the gap between the fifth and sixth order statistics.
Return a double[][] with two rows. The first row contains every gapA; the second contains every gapB. Preserve the input-row order, and do not mutate samples.
The returned arrays are the two empirical distributions. Visualization and hypothesis testing are follow-up discussion topics rather than additional judged outputs.

Function
adjacentGapDistributions(samples: double[][]) → double[][]

Examples
Example 1
samples = [[0.0,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9]]
return = [[0.1],[0.1]]
The sorted row is unchanged. Both adjacent gaps around the fifth order statistic equal 0.1.
Example 2
samples = [[0.9,0.1,0.4,0.8,0.2,0.7,0.3,0.6,0.0,0.5],[0.5,0.5,0.5,0.5,0.5,0.5,0.5,0.5,0.5,0.5]]
return = [[0.1,0.0],[0.1,0.0]]
The first row sorts to tenths from 0.0 through 0.9. The all-equal row has zero gaps.
Example 3
samples = [[0.0,0.0,0.0,0.0,0.25,0.75,1.0,1.0,1.0,1.0]]
return = [[0.25],[0.5]]
The fourth value is 0.0, the fifth is 0.25, and the sixth is 0.75.

Constraints
1 <= samples.length <= 100000.
samples[i].length == 10.
Every value is finite and lies in [0,1].
Return gaps in input-row order.
Answers within 10^-9 of the correct values are accepted.

Reported by candidates. Source: FastPrep

Pattern and pitfall

The trick is that the row length is fixed at 10. Sort a copy of each row, then read sorted[3], sorted[4], sorted[5]. gapA is sorted[4] minus sorted[3]. gapB is sorted[5] minus sorted[4]. Total work is 100000 rows times a constant-size sort, so it's effectively linear in the number of rows. Brute force isn't a threat. The real pitfalls are small. Don't mutate the input, so clone each row before sorting. Remember the indices are zero-based, so the fourth order statistic is index 3. Return two rows, not one pair per sample. Write into two preallocated arrays of length samples.length to keep order. Ties give 0.0 gaps, which is fine. The 10^-9 tolerance means plain double subtraction passes. If the indexing slips under pressure, StealthCoder is the hedge on the live OA, but the logic fits in about eight lines.

If you see this problem in your OA tomorrow, the play is to recognize the pattern in 30 seconds. StealthCoder buys you that recognition.

If this hits your live OA

You can drill Adjacent Order-Statistic Gap Distributions cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who passed his OA cold and still thinks the filter is broken.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Scale AI's OA.

Scale AI reuses patterns across OAs. Built by an Amazon engineer who passed his OA cold and still thinks the filter is broken. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Adjacent Order-Statistic Gap Distributions FAQ

How hard is the Scale AI adjacent gap problem really?+

Easy. It's a per-row sort plus two subtractions. The long statistics wording is noise. If you can sort a copy of an array and index it correctly, you can finish this fast. Most of your time goes to reading the spec and checking the output shape.

What's the trick to avoid timeouts?+

Sort each row separately, never the whole matrix. Each row has exactly 10 values, so the sort is constant cost. Across 100000 rows you're doing roughly a million element operations total. Nothing fancy is needed, no heaps and no selection algorithms.

Which indices do I use for the fourth, fifth, and sixth order statistics?+

Zero-based, so they're sorted[3], sorted[4], and sorted[5]. gapA is sorted[4] minus sorted[3]. gapB is sorted[5] minus sorted[4]. Off-by-one here is the most common way to fail the examples, so verify against Example 1 where both gaps are 0.1.

Do I need to worry about mutating the input or output order?+

Yes on both. Copy each row before sorting, because the spec says not to mutate samples. Fill the result at index i for row i so order is preserved. The return is a 2 by n array, first row all gapA values, second row all gapB values.

How do I prepare for this in 48 hours?+

Don't over-prep. Write a quick version in your language of choice, test it on the three examples, and check the all-equal row gives zeros. Practice the clone-then-sort idiom and building a 2 by n result array. That covers everything this problem tests.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Scale AI.

OA at Scale AI?
Invisible during screen share
Get it