Reported August 2021
Benchlinghash table

Generate Proteins from Contiguous Named Sequences

Reported by candidates from Benchling's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Benchling OA. Under 2s to a working solution.
Founder's read

The mistake that sinks a first attempt on this Benchling OA, reported in August 2021, is treating the recipe like a sorted interval merge. It isn't. You get named ranges, ordered recipes, and one rule: each sequence must start exactly where the last one ended. It's a hash map lookup plus a linear scan per recipe. Easy on paper, but the edge cases (empty recipe, unknown name) quietly kill submissions. If you blank during the live assessment, StealthCoder runs invisibly as a safety net while you work the problem.

The problem

Each named DNA sequence occupies a half-open interval [start,end). Each protein recipe lists sequence names in order. A recipe is valid when every sequence begins exactly where the previous one ends.
Return each valid protein in recipe order as name:start:end.
For this exercise, assume a recipe is invalid and skipped when it is empty or references any unknown sequence name. A valid recipe must be nonempty, reference only known sequences, and have matching endpoints at every adjacent join.

Function
generateNamedProteins(sequenceNames: String[], ranges: int[][], recipes: String[][], proteinNames: String[]) → String[]

Examples
Example 1
sequenceNames = ["AC","BC","PQ","XY","AB","BT"]
ranges = [[5,15],[3,20],[15,22],[22,35],[20,32],[9,13]]
recipes = [["AC","PQ"],["AC","PQ","XY"],["BC","AB"],["BT","AC"]]
proteinNames = ["P1","P2","P3","P4"]
return = ["P1:5:22","P2:5:35","P3:3:32"]
The first three recipes are contiguous; BT ends at 13 while AC starts at 5.

Constraints
Names are unique.
sequenceNames.length == ranges.length.
recipes.length == proteinNames.length.
All interval bounds are non-negative and each start is smaller than its end.

Reported by candidates. Source: FastPrep

Pattern and pitfall

Build a hash map from sequence name to its [start,end) pair. For each recipe, skip it if it's empty. Look up the first name, and if it's missing, skip. Record start as the first sequence's start. Then walk the rest, checking each name exists and that its start equals the previous end. Any mismatch invalidates the whole recipe. If you finish clean, emit proteinName:start:lastEnd. The common pitfall is sorting or merging intervals, which breaks the order the recipe defines. Another is emitting a partial result before you've validated the full chain. Validate first, then format. Also keep output in recipe order, not sorted by name. Complexity is O(total recipe length) after the O(n) map build. In the example, BT ends at 13 but AC starts at 5, so P4 drops out. StealthCoder is the hedge if you freeze on the live OA, but the logic here is short enough to write cold.

If this hits your live OA and you blank, StealthCoder solves it in seconds, invisible to the proctor.

If this hits your live OA

You can drill Generate Proteins from Contiguous Named Sequences cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Benchling's OA.

Benchling reuses patterns across OAs. Built by an Amazon engineer who would have shipped this the night before his JPMorgan OA if he'd had it. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Generate Proteins from Contiguous Named Sequences FAQ

What's the trick in the Benchling protein generation problem?+

Hash map from name to range, then one pass per recipe. Check that each sequence's start equals the previous sequence's end. Don't sort or merge anything. The recipe order is the order you validate in, and it's the order you output in.

How hard is this OA really?+

Easy to medium. The algorithm is simple, but the validity rules trip people up. Empty recipes, unknown names, and a mismatch in the middle of a chain all need to skip the whole recipe. Most failures come from missed edge cases, not hard logic.

What edge cases should I test?+

An empty recipe, a recipe with an unknown name anywhere in it, a single-sequence recipe (valid, since there are no joins), and a mismatch at the last join only. Also check that output keeps recipe order and skips invalid proteins without leaving gaps.

What's the time complexity I should aim for?+

O(n + m) where n is the number of sequences and m is the total length of all recipes. Building the map is O(n). Each recipe is scanned once with O(1) lookups. Anything slower, like searching the sequence list per name, is a red flag.

How do I prepare for this in 48 hours?+

Write this one from scratch twice with a map and a loop. Then do a few interval-chaining and validation problems with string output formatting. Practice the name:start:end formatting too, since off-by-one on the end value is easy to make.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Benchling.

OA at Benchling?
Invisible during screen share
Get it