Reported October 2026
Yahoosorting

Rank Unseen Items by Embedding Similarity

Reported by candidates from Yahoo's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Yahoo OA. Under 2s to a working solution.
Founder's read

Yahoo's October 2026 OA has a problem that sounds like machine learning and isn't. Strip the embedding talk and you're left with a filter, a dot product per item, and a sort with a tie-break. If you've got an invite in your inbox, this is the kind of question you can finish fast if you stay calm. It's tagged sorting, and that's the right read. The only real risk is sloppy handling of ties and seen items. StealthCoder sits invisibly on your screen as a safety net if you blank mid-assessment, but the logic here is short enough to hold in your head.

The problem

You are given one user embedding and n item embeddings of the same dimension. The similarity of an item is the dot product of its embedding with the user embedding.
Exclude every item index in seenItemIndexes. Return all remaining item indices sorted by descending similarity. Break an exact score tie by the smaller item index.

Function
rankUnseenItems(userEmbedding: double[], itemEmbeddings: double[][], seenItemIndexes: int[]) → int[]

Examples
Example 1
userEmbedding = [1.0,2.0]
itemEmbeddings = [[1.0,0.0],[0.0,2.0],[2.0,1.0],[-1.0,0.0]]
seenItemIndexes = [1]
return = [2,0,3]
After excluding item 1, the dot products are 4, 1, and -1 for indices 2, 0, and 3.
Example 2
userEmbedding = [1.0,1.0]
itemEmbeddings = [[2.0,0.0],[0.0,2.0],[1.0,1.0]]
seenItemIndexes = []
return = [0,1,2]
All three scores are 2, so indices determine the order.

Constraints
1 <= itemEmbeddings.length <= 100000.
1 <= userEmbedding.length <= 200, and every item embedding has that length.
All coordinates are finite doubles with absolute value at most 10^6.
Seen indices are distinct and valid.

Reported by candidates. Source: FastPrep

Pattern and pitfall

What it really reduces to: mark the seen indices in a hash set or boolean array, compute each remaining item's dot product with the user vector, then sort by score descending and index ascending. That's O(n*d + n log n), and with n up to 100000 and d up to 200 it's about 20 million multiply-adds, which is fine. The pitfalls are small but costly. Don't sort a copy of the embeddings, sort index/score pairs. Don't negate scores in a way that breaks the tie-break. Compare doubles directly, since the problem says exact ties use the smaller index. Don't scan the seen array for every item, that's quadratic. Watch out for returning an int array, not the scores. If the comparator gets fiddly mid-assessment, StealthCoder is the hedge that can hand you a clean version while you're live. Otherwise, write it yourself.

StealthCoder is the hedge for the one pattern you didn't drill. It runs invisibly during the screen share.

If this hits your live OA

You can drill Rank Unseen Items by Embedding Similarity cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. If you're reading this with an OA window open, you're who this was built for.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Yahoo's OA.

Yahoo reuses patterns across OAs. If you're reading this with an OA window open, you're who this was built for. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Rank Unseen Items by Embedding Similarity FAQ

What's the actual trick in this Yahoo OA problem?+

There isn't a deep trick. It's filter, score, sort. Put seen indices in a set, compute dot products for the rest, then sort by score descending with index ascending as the tie-break. The work is in getting the comparator right, not in finding a clever algorithm.

How hard is it really?+

Easy to low-medium. The logic fits in about fifteen lines. Most failures come from comparator mistakes, using a slow seen-check, or sorting the wrong structure. If you know your language's custom sort, you can finish it quickly.

How do I handle ties between equal scores?+

Sort by score descending, and when two scores are exactly equal, sort by index ascending. Example 2 shows this: all scores are 2, so the answer is [0,1,2]. Keep the original index with each score so the tie-break is trivial.

Will the sizes cause a timeout?+

Not if you're efficient. With up to 100000 items and 200 dimensions, computing all dot products is about 20 million operations, and the sort is n log n. Avoid checking seen items with a linear scan per item, use a set or boolean array.

How do I prep for this in 48 hours?+

Practice sorting index/score pairs with a custom comparator in your language, and write a dot product loop. Then test the two examples by hand, including negative scores and an empty seen list. That covers nearly everything this problem can throw at you.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Yahoo.

OA at Yahoo?
Invisible during screen share
Get it