Finite Same-Origin Web Crawler

Reported by candidates from Opendoor's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Opendoor OA. Under 2s to a working solution.
Founder's read

Opendoor reported this one in July 2026, and it looks like a web crawler but it's really a graph traversal with a string-parsing chore bolted on. You get a corpus of pages, a start URL, and you walk links that share an origin. If your OA lands in the next day or two, the good news is there's no networking, no concurrency, just BFS or DFS over a hash map. StealthCoder sits invisibly on your screen as a safety net if you blank on the href parsing mid-assessment. But the logic here is small enough to own before you sit down.

The problem

You are given parallel arrays urls and html. The string html[i] is the retrieved body for the absolute URL urls[i]. Starting at startUrl, crawl every reachable page in this supplied corpus that has the same origin as startUrl.
A link is an exact lowercase href attribute written as href="URL" or href='URL'. Attribute values do not contain their surrounding quote character.
Every link URL is absolute and starts with http:// or https://.
The origin is the scheme and authority before the first path, query, or fragment character. Host matching is case-sensitive because every supplied URL is already canonical.
Remove a URL fragment beginning with # before testing or visiting the link.
Follow a link only when its fragment-free URL has the same origin and appears in urls.
Visit each URL at most once. Return all visited URLs in lexicographic order.

Function
crawlPages(urls: String[], html: String[], startUrl: String) → String[]

Examples
Example 1
urls = ["https://homes.test/","https://homes.test/a","https://homes.test/b","https://other.test/x"]
html = ["<a href='https://homes.test/a'>A</a><a href='https://other.test/x'>X</a>","<a href=\"https://homes.test/b\">B</a>","<a href='https://homes.test/a#top'>A</a>",""]
startUrl = "https://homes.test/"
return = ["https://homes.test/","https://homes.test/a","https://homes.test/b"]
The crawler reaches pages /a and /b, ignores the external-origin page, and does not revisit /a through its fragment.
Example 2
urls = ["http://site.test/root","http://site.test/kept"]
html = ["<a href=\"http://site.test/missing\">Missing</a>",""]
startUrl = "http://site.test/root"
return = ["http://site.test/root"]
The linked page is absent from the supplied corpus, so only the starting page is visited.

Constraints
1 <= urls.length = html.length <= 5000
All strings in urls are unique canonical absolute HTTP or HTTPS URLs.
startUrl appears in urls.
The total length of all strings in html is at most 200000.
Every recognized href value is a nonempty absolute HTTP or HTTPS URL.

Reported by candidates. Source: FastPrep

Pattern and pitfall

The reduction: build a map from URL to html, compute the origin of startUrl, then run BFS or DFS with a visited set. Each page's links are nodes reachable by an edge. The only real work is extraction. Scan the html for href=" or href=' and read up to the matching closing quote. A regex like href=("|')(.*?)\1 does it cleanly. Then strip everything from the first # onward, compare the origin, and check the cleaned URL exists in the map. Origin is the text before the first /, ?, or # after the scheme's //. The classic pitfall is checking the origin before removing the fragment, or forgetting that a fragment-only variant maps to an already visited page. Sort the visited set at the end. Total html is 200000 characters, so linear scanning is fine. If the regex or the origin slicing slips under pressure, StealthCoder is the hedge during the live OA.

Drill it cold or hedge it with StealthCoder. Either way, don't walk into the OA hoping you remember the trick.

If this hits your live OA

You can drill Finite Same-Origin Web Crawler cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Made for the candidate who got the OA invite this morning and has 72 hours, not six months.

Get StealthCoder
⏵ The honest play

You've seen the question. Make sure you actually pass Opendoor's OA.

Opendoor reuses patterns across OAs. Made for the candidate who got the OA invite this morning and has 72 hours, not six months. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Finite Same-Origin Web Crawler FAQ

How hard is the Finite Same-Origin Web Crawler really?+

Easy to medium. The graph part is a textbook BFS with a visited set. The friction is parsing hrefs with two quote styles and slicing the origin correctly. If you've done any graph traversal, the logic takes minutes. The edge cases are what cost time.

What's the trick to extracting links?+

Find href=" or href=' and read until the same quote character comes back. A regex with a backreference handles both styles. Values never contain their own quote, so no escaping logic is needed. Then cut at the first # before doing anything else with the URL.

How do I compute the origin of a URL?+

Find the end of the scheme with ://, then look for the first /, ?, or # after that point. Everything before it is the origin. If none exist, the whole string is the origin. Compare it as an exact string since hosts are case-sensitive here.

What mistakes fail hidden tests?+

Not stripping the fragment before the corpus lookup, so /a#top gets dropped. Revisiting pages and looping on cycles. Following links to pages that aren't in urls. Forgetting to sort the output lexicographically. Also make sure the start URL itself is always included in the result.

How do I prepare for this in 48 hours?+

Write a BFS over a dictionary graph from memory, then write the href regex and origin slicer separately and test on both examples. Add a cyclic link case and a fragment duplicate case. That covers nearly everything this problem can throw at you.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Opendoor.

OA at Opendoor?
Invisible during screen share
Get it