SimpleRDD Word Count
Reported by candidates from DatologyAI's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
The DatologyAI OA reported in September 2026 dresses a hash map problem up in Spark vocabulary. "SimpleRDD Word Count" really reduces to this: split on whitespace, count tokens, and keep first-seen order. If you've got an invite and 48 hours, this is a gift, as long as you don't trip on the edge cases. Empty sentences, tabs, and case sensitivity are where people lose points. StealthCoder sits invisibly on your screen during the live OA as a safety net if you blank on the ordering detail, but you probably won't need it for this one.
The problem
Use the semantics of a small RDD pipeline to count words across a batch of sentences. Split each sentence on one or more whitespace characters, ignore empty tokens, and treat words as case-sensitive. Return one row [word, count] for each distinct word. Rows must follow the order in which each word first appears while scanning the sentences from left to right. Function simpleRddWordCount(sentences: String[]) → String[][] Examples Example 1 sentences = ["hello world","hello spark","world hello"] return = [["hello","3"],["world","2"],["spark","1"]] Rows retain first-occurrence order. Example 2 sentences = ["a a\tb","","b c"] return = [["a","2"],["b","2"],["c","1"]] Runs of spaces and tabs are delimiters and empty sentences add no words. Example 3 sentences = ["Data data DATA"] return = [["Data","1"],["data","1"],["DATA","1"]] Word equality is case-sensitive. Constraints 0 <= sentences.length <= 100000. The total number of characters is at most 1000000. Sentences contain printable ASCII characters and whitespace.
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick is an insertion-ordered hash map. Scan each sentence, split on one or more whitespace characters, drop empty tokens, and increment a count per word. Use a LinkedHashMap in Java or a regular dict in modern Python, which preserves insertion order. Then emit [word, count] with the count as a string. The pitfalls are small but costly. Splitting on a single space leaves empty tokens, so use a regex like \s+ or a manual scan and filter blanks. Don't lowercase anything, since Data, data, and DATA are three separate words. Don't sort the output either. The order is first appearance, not alphabetical or by frequency. With up to 1,000,000 characters, one linear pass is plenty. If you freeze on the ordering or the string conversion mid-OA, StealthCoder can hand you the clean solution in real time without the proctor seeing it.
If you see this problem in your OA tomorrow, the play is to recognize the pattern in 30 seconds. StealthCoder buys you that recognition.
You can drill SimpleRDD Word Count cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who passed his OA cold and still thinks the filter is broken.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass DatologyAI's OA.
DatologyAI reuses patterns across OAs. Built by an Amazon engineer who passed his OA cold and still thinks the filter is broken. Works on HackerRank, CodeSignal, CoderPad, and Karat.
SimpleRDD Word Count FAQ
How hard is the SimpleRDD Word Count problem really?+
Easy. It's a frequency count with one twist, preserving first-occurrence order. If you know hash maps and how to split strings on whitespace, you can finish in 15 minutes. The difficulty is in the details, not the algorithm.
What's the trick to getting the output order right?+
Use an insertion-ordered map. LinkedHashMap in Java, dict in Python 3.7+, or a Map in JavaScript all work. Alternatively keep a separate list of words appended the first time you see each one. Never sort the result.
How should I split sentences with tabs and multiple spaces?+
Split on one or more whitespace characters, such as the regex \s+, and discard empty strings. Leading whitespace can produce an empty first token, and an empty sentence produces nothing. Filtering blanks handles both cases cleanly.
Is case sensitivity a trap here?+
Yes. Example 3 shows Data, data, and DATA counted separately. Don't call lowercase or normalize anything. Use the raw token as the map key and you're fine.
How do I prepare for this in 48 hours?+
Write the solution once from scratch in your OA language. Test it on the three examples plus an empty array and a list of only empty strings. Confirm counts are returned as strings, not integers. That covers nearly every way this problem can go wrong.