Spam Text Classification
Reported by candidates from Salesforce's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
Salesforce reported this one in August 2025, and the detail that matters is buried in the statement: a text is spam at two spam-word occurrences, and every occurrence counts, not every distinct word. That's a hash set plus a tokenizing pass, nothing exotic. The trap is the constraints, because spamWords can hit 10^5 entries. If you have an OA invite in your inbox, the job is to avoid the nested-loop version. StealthCoder sits invisibly on your screen as a safety net if you blank mid-assessment, but this one is learnable tonight.
The problem
A classification system evaluates whether each text is spam using a list of spam words. A text is labeled "spam" if it contains at least two spam-word occurrences. Each occurrence of a spam word counts toward the total, and matching is case-sensitive. Given an array texts and an array spamWords, return an array where each element is either "spam" or "not_spam" for the corresponding text. Function classifyTexts(texts: String[], spamWords: String[]) → String[] Examples Example 1 texts = ["This is a limited offer just for you","Win cash now! Click here to claim your prize","Hello friend, just checking in","Congratulations! You have won a free gift"] spamWords = ["offer","cash","Click","prize","Congratulations","free"] return = ["not_spam","spam","not_spam","spam"] The first and third texts contain fewer than two spam-word occurrences. The second and fourth texts contain at least two. Constraints 1 <= texts.length <= 10^3 1 <= spamWords.length <= 10^5 1 <= text.length <= 10^5 1 <= spamWord.length <= 10^5 The combined length of all spam words is at most 10^7.
Reported by candidates. Source: FastPrep
Pattern and pitfall
Put spamWords in a hash set so each lookup is O(1). For each text, split on spaces and count tokens that appear in the set. Stop early once the count hits 2, then label the text spam or not_spam. Matching is case-sensitive, so don't lowercase anything. The pitfall is punctuation. In the example, "cash" is followed by a space but "now!" and "Congratulations!" carry punctuation, and "Congratulations!" must still count as a match in the sample output. So you need to strip or split on non-letter characters, not just spaces. Check the example carefully against whatever tokenization you pick. Also never loop over all spam words per text, since 10^3 texts times 10^5 words is far too slow. The set approach runs in roughly linear time in total text length. If the tokenizing rule trips you up live, StealthCoder is the hedge that reads the statement and gives you a working version.
The honest play: practice the pattern, and have StealthCoder ready for the one you didn't see coming.
You can drill Spam Text Classification cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built for the candidate who saw this exact problem leak two days before his OA and wondered if anyone had a play.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass Salesforce's OA.
Salesforce reuses patterns across OAs. Built for the candidate who saw this exact problem leak two days before his OA and wondered if anyone had a play. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Spam Text Classification FAQ
What's the trick in Spam Text Classification?+
Load spamWords into a hash set, tokenize each text, and count tokens found in the set. Stop at 2 and mark it spam. It's a lookup problem, so avoiding a scan of the whole word list per text is the whole point.
How hard is this Salesforce OA question really?+
Easy to medium. The logic is simple, but the large constraints punish brute force. The only real friction is deciding how to tokenize text with punctuation, like the exclamation marks in the sample. Get that right and it's quick.
Does case matter when matching spam words?+
Yes. The statement says matching is case-sensitive. So "Click" matches but "click" doesn't. Don't lowercase the texts or the spam words, or you'll count occurrences that shouldn't count and flip some labels.
Do repeated occurrences of the same spam word count twice?+
Yes. Each occurrence counts toward the total, so a text with the same spam word twice is spam. Count tokens, not distinct words. Don't dedupe into a set of found words, or you'll mislabel those texts.
How do I prepare for this in 48 hours?+
Write the set-plus-tokenizer solution once from scratch and test it against the sample. Then practice tokenizing strings with punctuation in your language. That covers the main failure points. Spend the remaining time on other hash-table and string problems.