Evaluate Noisy Annotators
Reported by candidates from OpenAI's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
OpenAI reported this one in September 2026, and the input size is the first thing to read. Up to 500 items and 200 annotators looks tiny, but a naive positive-versus-negative pair loop per annotator is still around 250 pairs times 200 annotators. That's fine, but only if you code it cleanly. The task is F1 plus AUROC per annotator, formatted to four decimals. It's a counting problem dressed up as ML evaluation, and the traps are in the edge cases. If you blank during the live OA, StealthCoder is the silent safety net that reads the screen and hands you the solution.
The problem
Given binary ground-truth labels and one row of integer confidence scores per annotator, return each annotator's F1 score and AUROC. A score of at least 50 predicts label 1. F1 is 2TP / (2TP + FP + FN); return 0 when its denominator is zero. AUROC compares every positive item with every negative item: a higher positive score earns 1, a tie earns 0.5, and a lower score earns 0. Return 0 if either class is absent. Format each annotator as f1,auc with exactly four digits after the decimal point, in input row order. Function evaluateAnnotators(truth: int[], scores: int[][]) → String[] Examples Example 1 truth = [1,0,1] scores = [[90,20,70],[40,30,60]] return = ["1.0000,1.0000","0.6667,1.0000"] The first annotator is perfect; the second misses one positive but ranks both positives above the negative. Example 2 truth = [1,0] scores = [[50,50]] return = ["0.6667,0.5000"] Both items are predicted positive, while an equal-score pair counts as one half in AUROC. Constraints 1 <= truth.length <= 500. 1 <= scores.length <= 200; every score row has truth.length entries. Truth entries are 0 or 1; scores are between 0 and 100.
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick is that each annotator is independent, so you loop over rows and compute two numbers. For F1, threshold at 50, count TP, FP and FN, then return 0 if 2TP+FP+FN is zero. For AUROC, count positives and negatives first. If either is zero, return 0. Otherwise compare every positive score with every negative score: win adds 1, tie adds 0.5. Divide by positives times negatives. With 500 items the worst case is about 62,500 pairs per annotator, times 200, roughly 12 million comparisons. That passes. You can sort and use ranks for O(n log n), but it's optional. The pitfalls are the tie handling, the score of exactly 50 counting as positive, and formatting with exactly four digits. Use fixed-point formatting, not string trimming. StealthCoder is your hedge if the formatting or the tie logic slips under time pressure.
If this hits your live OA and you blank, StealthCoder solves it in seconds, invisible to the proctor.
You can drill Evaluate Noisy Annotators cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass OpenAI's OA.
OpenAI reuses patterns across OAs. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Evaluate Noisy Annotators FAQ
How hard is Evaluate Noisy Annotators really?+
Easy to medium. There's no clever data structure. You count, compare and format. Most people lose points on edge cases like a zero F1 denominator, a missing class in AUROC, or the tie rule, not on the algorithm itself.
Do I need an O(n log n) AUROC?+
No. With truth.length up to 500 and up to 200 annotators, the pairwise comparison is about 12 million operations in the worst case. That's fine. Rank-based AUROC works too, but it adds tie-handling risk for no real gain here.
What's the trick with the threshold and ties?+
A score of at least 50 predicts 1, so exactly 50 is positive. Example 2 shows it: both items predicted positive gives F1 of 0.6667. For AUROC, equal positive and negative scores add 0.5, not 0 or 1.
What edge cases should I test before submitting?+
Test all-zero truth, all-one truth, and an annotator who predicts nothing positive. In these cases F1's denominator can be zero, or AUROC has a missing class. Both must return 0, and the output must still show four decimals, like 0.0000.
How do I prepare for this in 48 hours?+
Write the function once from scratch. Compute TP, FP, FN, then the pair loop, then format with fixed four decimals. Run both examples by hand. Practice the formatting in your language, since rounding and trailing zeros cause most wrong answers.