Validate Six-Column Business CSV Rows
Reported by candidates from Stripe's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
Stripe reported this one in September 2026, and the input size is the first thing to read. Up to 100000 rows and 2 * 10^6 total characters means you can't rescan the word lists for every token. It's a string parsing and validation problem with four ordered rules, and the whole game is hash sets plus careful edge cases. Build the sets once, walk each row once, stop at the first failure. If you blank on the rule 4 math during the live OA, StealthCoder runs invisibly as a safety net and shows you the working structure.
The problem
rows[0] is a six-column header. Validate every later comma-separated row and return one label per data row. Fields do not contain commas, quoting, or escapes. Apply these rules in order and report only the first failure: RULE_1: exactly six trimmed fields, each nonempty. RULE_2: the trimmed fifth field has length from 5 through 31 inclusive. RULE_3: no whitespace token in the second field is a case-insensitive member of forbiddenWords. RULE_4: remove case-insensitive ignoreWords from the distinct lowercase whitespace tokens of field two. At least half of the remaining words, rounded up, must appear in the distinct lowercase non-ignored tokens of field four or field five. If no word remains from field two, this rule fails. Return VALID or INVALID:RULE_X. Function validateBusinessRows(rows: String[], forbiddenWords: String[], ignoreWords: String[]) → String[] Examples Example 1 rows = ["id,name,type,description,code,status","1,Acme LLC,vendor,Acme services,12345,active","2,Bad Corp,vendor,Other,12345,active"] forbiddenWords = ["bad"] ignoreWords = ["llc","corp"] return = ["VALID","INVALID:RULE_3"] The first row has an overlapping word after ignoring LLC; the second fails the forbidden-word rule first. Example 2 rows = ["h1,h2,h3,h4,h5,h6","1,Alpha Beta,x,beta market,alpha,ok","1,Alpha Beta,x,none,alpha,ok","1,x,a,12345,ok"] forbiddenWords = [] ignoreWords = [] return = ["VALID","VALID","INVALID:RULE_1"] Rules are checked in order, and both valid rows meet the one-of-two overlap threshold. Example 3 rows = ["a,b,c,d,e,f","1,LLC LLC,x,llc,llc12,ok"] forbiddenWords = [] ignoreWords = ["llc"] return = ["INVALID:RULE_4"] A second field with no non-ignored word fails rule four. Constraints 2 <= rows.length <= 100000. The header has exactly six comma-separated fields. Each row and word-list entry has at most 1000 characters. The combined input length is at most 2 * 10^6.
Reported by candidates. Source: FastPrep
Pattern and pitfall
The trick is preprocessing. Lowercase forbiddenWords and ignoreWords into two hash sets once. For each data row, split on commas with a limit that keeps empty trailing fields, so "1,x,a,12345,ok" gives five fields and fails RULE_1. Trim each field. Then check length of field five, then scan field two tokens against the forbidden set. For RULE_4, build the distinct lowercase non-ignored token set from field two. If it's empty, fail. Build a combined set from fields four and five, skipping ignored words. Count matches and compare against ceil(n/2), which is (n+1)/2 in integer math. Pitfalls: splitting on whitespace with multiple spaces, forgetting distinctness, and checking rules out of order. Rule 3 uses raw tokens, rule 4 uses ignore-filtered ones. Total work is linear in input size, which is what the constraints demand. StealthCoder is your hedge if the ordering or rounding trips you live.
If this hits your live OA and you blank, StealthCoder solves it in seconds, invisible to the proctor.
You can drill Validate Six-Column Business CSV Rows cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it.
Get StealthCoderRelated leaked OAs
You've seen the question.
Make sure you actually pass Stripe's OA.
Stripe reuses patterns across OAs. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Validate Six-Column Business CSV Rows FAQ
What's the trick in this Stripe OA question?+
Precompute lowercase hash sets for forbiddenWords and ignoreWords, then process each row once. Check rules in order and return on the first failure. Rule 4 is the only one with real logic: distinct non-ignored tokens, ceil half threshold, empty set fails.
How do I split rows without losing empty fields?+
Use a split that keeps trailing empty strings, like split(",", -1) in Java. Otherwise "a,b,c,d,e," looks like five fields instead of six with an empty one. Either way it fails RULE_1, but the count and emptiness check must be exact.
How do I compute the at-least-half rounded up threshold?+
With n remaining words from field two, the required matches are (n + 1) / 2 using integer division. One word needs one match. Three words need two. Count distinct field-two words found in the combined field four and five set.
Does brute force pass the constraints?+
Not safely. With 100000 rows and word lists up to the full input size, scanning lists per token blows up. Hash sets give constant-time lookups, so total work stays linear in the 2 * 10^6 characters.
How do I prepare for this in 48 hours?+
Write the solution once from scratch with the four rules as separate small checks. Test the three examples, plus a row with extra spaces, a row with five fields, and a field two made only of ignored words. Those cases catch nearly every bug.