Reported September 2026
Harveysorting

Highlight and Rank Overlapping Citations

Reported by candidates from Harvey's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Harvey OA. Under 2s to a working solution.
Founder's read

Harvey reported this one in September 2026, and the first attempt usually dies on the same bug: merging spans that only touch. "Highlight and Rank Overlapping Citations" looks like string formatting, but it's really interval merging with a ranking step bolted on. You find every word-level match of every source phrase, merge the overlapping spans, wrap each region in yellow tags, then append citations sorted by global frequency. If you've got an OA invite for this, expect the edge cases to hurt more than the idea. StealthCoder is the safety net if you blank mid-assessment.

The problem

Given an output string text and a list of source phrases sources, find every exact word-level occurrence of every source phrase in text. Words contain only letters and digits, and adjacent words are separated by exactly one space. Matching is case-sensitive and must use complete words, so network does not match a word inside networking.
Each occurrence covers a half-open span of words. Merge all occurrences whose spans overlap, directly or transitively, into a maximal highlighted region. Merely adjacent spans do not overlap. Wrap each merged region in <yellow> and </yellow>.
After each highlighted region, append one citation [i] for every distinct source index i that has an occurrence participating in that region. Order those indices by the total number of exact occurrences of that source across the entire text, from greatest to least. Break equal-frequency ties by smaller source index.
Keep exactly one space between the original words and inserted highlighted regions. If no source phrase occurs, return text unchanged.

Function
highlightAndRankCitations(text: String, sources: String[]) → String

Examples
Example 1
text = "the quick brown fox"
sources = ["the quick","quick brown"]
return = "<yellow>the quick brown</yellow>[0][1] fox"
The phrases cover word spans [0, 2) and [1, 3). They overlap and merge into the quick brown. Each source occurs once, so the index tie is resolved as [0][1].
Example 2
text = "red blue green red blue red"
sources = ["red blue","blue green","red"]
return = "<yellow>red blue green</yellow>[2][0][1] <yellow>red blue</yellow>[2][0] <yellow>red</yellow>[2]"
Across the whole text, source 2 occurs three times, source 0 twice, and source 1 once. Those totals determine citation order inside every merged region. The last one-word occurrence is adjacent to, but does not overlap, the preceding phrase.
Example 3
text = "aa aaab x"
sources = ["aa","aaab","aa aaab"]
return = "<yellow>aa aaab</yellow>[0][1][2] x"
Word-level matching keeps aa separate from aaab. The two single-word matches overlap the two-word source span, so all three belong to one region. Their global counts tie.

Constraints
1 <= text.length <= 100000
1 <= sources.length <= 2000
1 <= sources[i].length
text and every sources[i] contain only letters, digits, and single spaces between words.
The total number of exact source occurrences in text is at most 200000.

Reported by candidates. Source: FastPrep

Pattern and pitfall

Split text into words and turn each source into a word list. Find every occurrence at word boundaries, not by raw substring search, so network never matches inside networking. Record each as a half-open span [start, end) with its source index, and count total occurrences per source across the whole text. Sort spans by start, then sweep and merge only when next.start < currentEnd. Strict less-than is the pitfall: adjacent spans like [2,3) and [3,4) must stay separate, as Example 2 shows. For each merged region, collect the distinct source indices, then sort by count descending and index ascending. Keep the count from the whole text, not the region. Watch the 200000 occurrence cap: matching naively can be slow, so hash the word sequences or use a trie or rolling hash. If the live OA freezes you, StealthCoder can supply the merge sweep while you handle the formatting details.

The honest play: practice the pattern, and have StealthCoder ready for the one you didn't see coming.

If this hits your live OA

You can drill Highlight and Rank Overlapping Citations cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built for the candidate who saw this exact problem leak two days before his OA and wondered if anyone had a play.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Harvey's OA.

Harvey reuses patterns across OAs. Built for the candidate who saw this exact problem leak two days before his OA and wondered if anyone had a play. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Highlight and Rank Overlapping Citations FAQ

What's the trick in the Harvey highlight and citation problem?+

Treat it as interval merging on word indices. Collect every match as a half-open span, sort by start, and merge only when the next start is strictly less than the current end. Then rank citations using global counts, not counts inside the region.

Why do adjacent spans stay separate?+

The spec says merely adjacent spans don't overlap. With half-open spans, [2,3) and [3,4) touch but share no word. Merge on start < end, never start <= end. Example 2 shows this with the last single word red.

How do I avoid matching inside longer words?+

Work at the word level. Split text on spaces and compare sequences of whole words, never raw substrings. Then network can't match networking, and aa stays separate from aaab, exactly as Example 3 requires.

How should I order the citation indices?+

Count total exact occurrences of each source across the entire text first. For each region, take the distinct source indices, sort by count descending, and break ties by smaller index. Don't recount per region, since that gives the wrong order.

How do I prepare for this in 48 hours?+

Practice merge intervals until the sweep is automatic, then write a word-level phrase matcher. Hand-trace the three examples, especially the adjacent case and the tie case. Test duplicate sources and a text with no matches, which must return unchanged.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Harvey.

OA at Harvey?
Invisible during screen share
Get it