Highlight Matching Phrases
Reported by candidates from Harvey's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
Harvey reported this one in September 2026, and the mistake that sinks a first attempt is finding only the first occurrence of each phrase. Highlight Matching Phrases looks like string search, but the real task is interval merging. Overlapping hits like "aa" inside "aaaa" must be found at every start index, then fused with adjacent ones into one <mark> block. If you have an OA invite for this, expect the edge cases to be the whole test. StealthCoder sits invisibly on your screen as a safety net if you blank mid-assessment, but the logic below is short enough to carry in your head.
The problem
Given a sentence sentence and an array of literal phrases phrases, highlight every character that belongs to at least one occurrence of any phrase. Matching is case-sensitive. Search for every occurrence, including overlapping occurrences. Merge overlapping or directly adjacent matched ranges, then wrap each merged range with <mark> and </mark>. Preserve every unmatched character exactly. If no phrase occurs, return the original sentence. Function highlightMatches(sentence: String, phrases: String[]) → String Examples Example 1 sentence = "The quick brown fox jumps over the lazy dog." phrases = ["quick brown","brown fox jumps"] return = "The <mark>quick brown fox jumps</mark> over the lazy dog." The two occurrences overlap on brown, so their ranges form one highlighted segment. Example 2 sentence = "aaaa and aa" phrases = ["aa","aaa"] return = "<mark>aaaa</mark> and <mark>aa</mark>" All overlapping occurrences in the first word merge into one range. The final occurrence is separated by unmatched text. Example 3 sentence = "Case Sensitive" phrases = ["case","missing"] return = "Case Sensitive" Lowercase case does not match uppercase Case, so the sentence is unchanged. Constraints 1 <= sentence.length <= 2 * 10^4 1 <= phrases.length <= 200 1 <= phrases[i].length <= 200 The sentence and phrases contain printable ASCII characters other than < and >.
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick: mark characters first, wrap later. Build a boolean array the length of the sentence. For each phrase, scan every start index (use indexOf from i+1 after each hit, not after the match end) and set the covered positions to true. Then walk the array and emit <mark> when you enter a true run and </mark> when you leave it. Adjacent ranges merge for free because they form one continuous run of true values. The pitfall is skipping ahead by phrase length, which drops overlaps like "aaa" in "aaaa". Worst case is about 200 phrases times 20,000 positions times 200 characters, so naive comparison is fine, but indexOf is cleaner. Alternatively collect intervals, sort, and merge with a touch-or-overlap rule. StealthCoder is your hedge in the live OA if the off-by-one on adjacent ranges trips you up. Test Example 2 by hand before submitting.
Drill it cold or hedge it with StealthCoder. Either way, don't walk into the OA hoping you remember the trick.
You can drill Highlight Matching Phrases cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Made for the candidate who got the OA invite this morning and has 72 hours, not six months.
Get StealthCoderRelated leaked OAs
This OA pattern shows up on LeetCode as add bold tag in string. If you have time before the OA, drill that.
You've seen the question.
Make sure you actually pass Harvey's OA.
Harvey reuses patterns across OAs. Made for the candidate who got the OA invite this morning and has 72 hours, not six months. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Highlight Matching Phrases FAQ
What's the trick in Highlight Matching Phrases?+
Don't build the output while searching. Mark every matched character in a boolean array, then do one pass to wrap consecutive true runs in <mark> tags. Merging overlapping and adjacent ranges happens automatically because they become one continuous run.
Why does "aaaa" with phrases "aa" and "aaa" give one block?+
Matches start at indexes 0, 1, 2 for "aa" and 0, 1 for "aaa". Together they cover all four characters in one continuous run. You only get that if you check every start index, not skip past each match.
How hard is this really for the Harvey OA?+
Medium at most. The algorithm is simple, but the details bite: overlaps, adjacency, case sensitivity, and the no-match case returning the original sentence. Candidates who rush the first attempt usually fail on overlapping occurrences.
Do I need a trie or Aho-Corasick?+
No. With sentence length up to 20,000 and 200 phrases of up to 200 characters, brute-force search per phrase fits comfortably. Aho-Corasick is valid but adds bug risk for no real gain here.
How do I prepare for this in 48 hours?+
Write the boolean-marking version from scratch twice. Then write the interval-sort-merge version. Test on the three examples plus an adjacent case like phrases "ab" and "cd" on "abcd", which should yield one merged mark block.