Sequential Phrase Search Across Documents
Reported by candidates from Confluent's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
Confluent reportedly put a phrase search problem in front of candidates in September 2026, and the input sizes look friendly enough that people will overthink it. You get up to 1000 documents, each up to 1000 characters, and a phrase up to 200. That's small. The hinted pattern is binary search, but this is really a contiguous word-sequence match on strings. If you've got the OA in a day or two, know the shape before you open the editor. And if you blank mid-assessment, StealthCoder runs invisibly as a safety net.
The problem
Given an array of text documents and a query phrase, return the zero-based indices of all documents that contain the phrase. A phrase matches when all of its words appear as one contiguous sequence of complete words in the same order. Words are lowercase English-letter tokens separated by one space. Include each matching document index once, in increasing order. Function findDocumentsWithPhrase(documents: String[], phrase: String) → int[] Examples Example 1 documents = ["distributed systems need careful testing","careful distributed systems testing","distributed reliable systems"] phrase = "distributed systems" return = [0,1] The first two documents contain distributed systems contiguously. The third separates the two words with reliable. Example 2 documents = ["alpha beta gamma","beta alpha gamma","alpha gamma beta"] phrase = "alpha beta" return = [0] Only document 0 contains the two query words next to each other in the requested order. Constraints 1 <= documents.length <= 1000. 1 <= documents[i].length <= 1000. 1 <= phrase.length <= 200. Every document and phrase contains lowercase English words separated by exactly one space, with no leading or trailing space.
Reported by candidates. Source: FastPrep
Pattern and pitfall
Do the math first. 1000 documents times 1000 characters is about a million characters total. Checking the phrase at each word offset costs at most 200 characters per compare, so worst case is still a few hundred million trivial ops at the very extreme, and in practice far less because documents are short. Brute force passes. Split each document into words, split the phrase into words, then slide the phrase over the document word by word and compare. Or pad with spaces and use a substring check: (" " + doc + " ").contains(" " + phrase + " "). The padding is the trick. Without it, "systems" matches inside "ecosystems". That's the pitfall. Don't reach for binary search, nothing here is sorted. Collect indices in loop order and they're already increasing. If you freeze during the live OA, StealthCoder is the hedge that reads the problem and hands you the padded-substring solution.
Drill it cold or hedge it with StealthCoder. Either way, don't walk into the OA hoping you remember the trick.
You can drill Sequential Phrase Search Across Documents cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Made for the candidate who got the OA invite this morning and has 72 hours, not six months.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass Confluent's OA.
Confluent reuses patterns across OAs. Made for the candidate who got the OA invite this morning and has 72 hours, not six months. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Sequential Phrase Search Across Documents FAQ
What's the trick in the Confluent phrase search problem?+
Match on whole words only. Either pad both the document and the phrase with a leading and trailing space and use a substring check, or split into word arrays and slide the phrase across. Plain contains() on raw strings will falsely match partial words like ecosystems.
Is binary search actually needed here?+
No. The hint says binary search, but nothing in the input is sorted and the answer is a filter over documents. A linear scan per document is enough. Indices come out in increasing order naturally because you iterate documents from 0 upward.
Will brute force pass the constraints?+
Yes. There are at most 1000 documents of at most 1000 characters, and the phrase is at most 200 characters. Scanning every document once with a substring or word-window compare is comfortably fast. No indexing or hashing structure is required.
What edge cases should I test before submitting?+
Test a phrase that's a single word, a phrase equal to the whole document, a phrase longer than the document, and a word that's only a substring of another word. Also confirm the words appear in the same order, as in example 2 where beta alpha doesn't match.
How do I prep for this in 48 hours?+
Write the padded-substring version and the word-window version once each, from memory. Run both on the two examples. Then practice returning index lists in order. It's a string matching warmup, so an hour is plenty, and the rest of your time is better spent on harder OA questions.