Word Search Across Documents
Reported by candidates from Confluent's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
This Confluent OA, reported in September 2026, looks like a search problem and tempts you toward something fancy. It isn't. Strip the wrapper and it's a token-match scan: for each document, ask whether the query shows up as a whole word. The hinted pattern says binary-search, but nothing here is sorted, so don't force it. If you blank on the day, StealthCoder runs invisibly on your screen and can hand you the clean version while you keep your head. Read this first and you probably won't need it.
The problem
Given an array of text documents and a query word, return the zero-based indices of all documents that contain the query as a complete word. Words are lowercase English-letter tokens separated by one space. Matching is exact: a query does not match part of a longer token. Include each matching document index once, in increasing order. Function findDocumentsWithWord(documents: String[], word: String) → int[] Examples Example 1 documents = ["red blue green","bluebird flies","green blue"] word = "blue" return = [0,2] Documents 0 and 2 contain token blue. Token bluebird is not an exact match. Example 2 documents = ["alpha beta","gamma delta","alpha"] word = "omega" return = [] No document contains the token omega. Constraints 1 <= documents.length <= 1000. 1 <= documents[i].length <= 1000. 1 <= word.length <= 50. Every document contains lowercase English words separated by exactly one space, with no leading or trailing space. word contains only lowercase English letters.
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick is exact token matching. Split each document on a single space and compare every token to the query with equality. That's it. Documents are processed in order, so the output indices come out already increasing. Break out of the token loop on the first hit so each index is added once. The classic pitfall is calling contains or indexOf on the raw string. That makes blue match bluebird and fails Example 1. Another trap is a manual substring check that forgets word boundaries. With 1000 documents of up to 1000 characters, a full scan is about a million character operations, so no index or binary search is needed. Don't build a hash map of words unless you want it. If the live OA throws you off and you freeze on the edge cases, StealthCoder is the hedge that reads the prompt and gives you a working answer without the proctor seeing it.
If this hits your live OA and you blank, StealthCoder solves it in seconds, invisible to the proctor.
You can drill Word Search Across Documents cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass Confluent's OA.
Confluent reuses patterns across OAs. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Word Search Across Documents FAQ
What's the real trick in Word Search Across Documents?+
Match whole tokens, not substrings. Split each document by a space and compare each token to the query with exact equality. Using contains or indexOf is the bug, because blue would match bluebird. Everything else is just collecting indices in order.
Is binary search actually needed here?+
No. The hint suggests it, but the documents aren't sorted and nothing is monotonic. A linear scan over every document and token fits easily inside the constraints of 1000 documents at up to 1000 characters each. Forcing binary search adds risk for zero gain.
What's the time complexity I should state?+
O(total characters across all documents), since you look at each character about once when splitting and comparing. Extra space is O(length of one document) for the split tokens, plus the output list. That's plenty fast for the stated limits.
Which edge cases should I test before submitting?+
Test a query that's a prefix of a longer token, like blue and bluebird. Test a query that's a suffix of a token. Test a single-word document, a document where the word appears twice (index only once), and a query with no matches, which should return an empty array.
How do I prepare for this in 48 hours?+
Practice a few string-splitting and exact-match problems, and get comfortable with your language's split and equality calls. Write this one from scratch twice and run both examples by hand. It's an easy problem, so the risk is sloppy substring matching, not difficulty.