Wildcard Bad-Word Filter
Reported by candidates from Sentry's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
The Sentry OA reported in August 2023 looks like a string filter, but one edge case wrecks the naive version: patterns made only of asterisks. Sentry's Wildcard Bad-Word Filter hands you a list of patterns with edge wildcards and a message full of punctuation and odd whitespace. You have to mask only the letters of matching words and leave everything else untouched. It's tokenizing plus pattern matching, nothing exotic. But the details bite. If you're taking this in the next day or two, know the trap before you start. StealthCoder sits invisibly on your screen as a safety net if you blank mid-assessment.
The problem
Replace bad words in a message while preserving every punctuation mark and whitespace character. A candidate word is a maximal run of ASCII letters. Matching ignores case. A pattern without * must equal the whole word. A leading * matches any prefix, a trailing * matches any suffix, and both match any containing word. Only edge wildcards are allowed; patterns containing no letters are ignored. Replace every letter of a matching word with *. Function filterBadWords(badWords: String[], message: String) → String Examples Example 1 badWords = ["jerk*","*lame*"] message = "Stop it, you jerkwad! I remain blameless!" return = "Stop it, you *******! I remain *********!" The prefix and contains patterns match while punctuation remains. Example 2 badWords = ["bad","*tail","head*"] message = "BAD cocktail headwind badge" return = "*** ******** ******** badge" Matching ignores case and exact bad does not match badge. Example 3 badWords = ["*","**"] message = "keep all!" return = "keep all!" Wildcard-only patterns are ignored and repeated spaces stay unchanged. Constraints 1 <= badWords.length <= 1000. 1 <= message.length <= 100000. Patterns contain ASCII letters and optional edge * characters; the message contains printable ASCII and whitespace.
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick is to split the message into maximal runs of ASCII letters and copy every non-letter character through unchanged. Preprocess the patterns once: lowercase them, strip the edge stars, and record the type as exact, prefix, suffix, or contains. If the stripped core is empty, drop the pattern. That's the Example 3 trap, where a pattern like * or ** would match every word if you forget it. For each word, lowercase it and test it against each rule with equals, startsWith, endsWith, or contains. With 1000 patterns and a message up to 100000 characters, that's fine, but build the output with a list or builder, not repeated string concatenation. Another pitfall is treating bad as matching badge. Exact means the whole word. StealthCoder is your hedge in the live OA if the tokenizing logic slips away under pressure, but the pattern is simple enough to hold in your head.
Drill it cold or hedge it with StealthCoder. Either way, don't walk into the OA hoping you remember the trick.
You can drill Wildcard Bad-Word Filter cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Made for the candidate who got the OA invite this morning and has 72 hours, not six months.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass Sentry's OA.
Sentry reuses patterns across OAs. Made for the candidate who got the OA invite this morning and has 72 hours, not six months. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Wildcard Bad-Word Filter FAQ
What's the trick in the Sentry Wildcard Bad-Word Filter?+
Tokenize the message into maximal letter runs, keep every other character as is, and normalize patterns first. Strip the edge stars, lowercase, and classify as exact, prefix, suffix, or contains. Ignore any pattern whose core has no letters. That one rule covers the nastiest edge case.
How hard is this problem really?+
Easy to medium. There's no fancy algorithm, just careful string handling. Most failures come from edge cases: star-only patterns, case handling, and exact patterns wrongly matching longer words like badge. If you've written a tokenizer before, it's quick.
Do I need a trie or regex for this?+
No. Brute force is enough. With up to 1000 patterns, each word is checked against each rule with simple string methods. Regex adds escaping risk and no benefit. A trie is overkill given the suffix and contains cases.
How should I handle punctuation and repeated spaces?+
Don't touch them. Walk the message once, and when you hit a letter, extend to the end of the run, decide whether to mask it, and append the result. Any non-letter gets appended directly. That preserves every space and symbol exactly.
How do I prepare for this in 48 hours?+
Write a small tokenizer and test it on the three examples, especially the star-only case and the bad versus badge case. Then try a message with leading punctuation, trailing punctuation, and mixed case. Two hours of that covers it.