Reported April 2022
Bloombergsorting

Sort by Variable-Length Alphabet Tokens

Reported by candidates from Bloomberg's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Bloomberg OA. Under 2s to a working solution.
Founder's read

The mistake that sinks a first attempt on this Bloomberg OA question, reported April 2022, is sorting the words as plain strings. The alphabet has a custom order and the symbols have different lengths, so ordinary string comparison gives wrong answers. It's a sorting problem wearing a tokenizer costume. You map each word to a sequence of symbol ranks, then sort by those sequences. If you blank on the tokenizing step during the live OA, StealthCoder runs invisibly on your desktop and gives you a working solution as a safety net.

The problem

alphabet lists language symbols from smallest to largest; each symbol is a nonempty string and lengths may vary. Every input word has exactly one tokenization into these symbols.
Sort words lexicographically by their token-rank sequences. If one sequence is a prefix of another, the shorter word comes first.

Function
sortTokenAlphabetWords(alphabet: String[], words: String[]) → String[]

Examples
Example 1
alphabet = ["ba","aa","cb","abc","d","dd"]
words = ["cbba","abccb","aaba","baaa"]
return = ["baaa","aaba","cbba","abccb"]
The unique token sequences begin with ranks 0,1,2,3 respectively.

Constraints
Every word is uniquely tokenizable.
Total input length is at most 10^5.

Reported by candidates. Source: FastPrep

Pattern and pitfall

The trick is two steps. First, tokenize each word into rank numbers. Put the alphabet symbols in a hash map from string to rank, then walk the word and grab the symbol that matches at the current position. The problem guarantees exactly one tokenization, so you don't need backtracking or DP, but you can still try each possible length at each position. Second, sort the words by their rank lists. Lists compare lexicographically and a shorter prefix sorts first, which matches the required rule in most languages for free. The pitfall is comparing raw characters, which ignores the custom order. Another is a slow scan against every symbol at each position. With total input up to 10^5, build a trie or hash the symbols and cap the substring length at the longest symbol. Total work stays near O(n log n) plus tokenizing. StealthCoder is your hedge if the tokenizer logic slips under pressure.

Memorize the pattern. If you can't, run StealthCoder. The proctor sees the IDE. They don't see what's behind it.

If this hits your live OA

You can drill Sort by Variable-Length Alphabet Tokens cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Made by an engineer who treats the OA as theater. If yours is tonight, you don't have time to grind. You have time to hedge.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Bloomberg's OA.

Bloomberg reuses patterns across OAs. Made by an engineer who treats the OA as theater. If yours is tonight, you don't have time to grind. You have time to hedge. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Sort by Variable-Length Alphabet Tokens FAQ

What's the trick in the Bloomberg variable-length alphabet sort?+

Convert every word into a list of symbol ranks, then sort words by those lists. Raw string comparison fails because the alphabet order is custom and symbols vary in length. Once each word is a rank sequence, the lexicographic comparison with shorter-prefix-first is standard.

How do I tokenize a word efficiently?+

Store the alphabet in a hash map from symbol to rank. At each position, try substrings up to the longest symbol length and take the one in the map. A trie also works. Since tokenization is guaranteed unique, no backtracking is needed.

How hard is this problem really?+

Easy to medium. The sorting idea is simple. The risk is implementation: off-by-one errors in tokenizing and forgetting the prefix rule. Total input is at most 10^5, so an accidental quadratic scan is the real danger.

Does the prefix rule need special handling?+

Usually not. If you sort lists of integers, most languages compare element by element and treat a shorter list that is a prefix as smaller. In a language without that behavior, write a comparator that checks length after the shared elements match.

How do I prepare for this in 48 hours?+

Practice sorting with custom keys and comparators, and write a small tokenizer using a map or trie. Then run the sample by hand: alphabet ranks, token sequences, final order. Check edge cases like one-symbol words and words that are prefixes of others.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Bloomberg.

OA at Bloomberg?
Invisible during screen share
Get it