Crawl Same-Hostname URLs Concurrently
Reported by candidates from Anthropic's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.
Anthropic reported this one in September 2026, and the title sounds scarier than the problem. "Concurrently" makes you think threads, locks and race conditions. Strip that away and it's a plain graph traversal with a visited set, a hostname filter and a sorted return. The workerCount parameter is mostly flavor, since the output is the same however you schedule the work. If you've got an Anthropic OA coming, expect the difficulty to live in the string handling, not the parallelism. StealthCoder sits invisibly on your screen as a safety net if you blank mid-assessment, but the shape of this problem is simple enough to hold in your head.
The problem
Crawl the finite web graph described by pageUrls and linkLists, beginning at startUrl. The string linkLists[i] contains the absolute URLs returned when pageUrls[i] is parsed, separated by single spaces; an empty string means no links. A normalized URL absent from pageUrls has no outgoing links. Before any comparison or lookup, remove the first # and everything after it. Visit only normalized URLs whose hostname exactly equals the normalized start URL's hostname. Different schemes may still have the same hostname. Never visit a normalized URL more than once, including through fragments or cycles. Use exactly workerCount workers for the crawl. Return every visited normalized URL in lexicographic order. Function crawlSameHostname(startUrl: String, pageUrls: String[], linkLists: String[], workerCount: int) → String[] Examples Example 1 startUrl = "https://docs.acme.com/root#intro" pageUrls = ["https://docs.acme.com/root","https://docs.acme.com/a","https://docs.acme.com/b"] linkLists = ["https://docs.acme.com/a#one https://docs.acme.com/a#two https://other.com/out","https://docs.acme.com/b#part https://docs.acme.com/root#back","https://docs.acme.com/a"] workerCount = 4 return = ["https://docs.acme.com/a","https://docs.acme.com/b","https://docs.acme.com/root"] Both fragments of /a normalize to one URL. The cycle through /root and /a cannot create another visit, and other.com is excluded. Example 2 startUrl = "https://api.site.com/home" pageUrls = ["https://api.site.com/home","http://api.site.com/v1"] linkLists = ["http://api.site.com/v1 https://api.site.com/unlisted#top","https://api.site.com/home#again"] workerCount = 2 return = ["http://api.site.com/v1","https://api.site.com/home","https://api.site.com/unlisted"] The scheme does not affect hostname equality. The unlisted URL is visited but contributes no outgoing links. Example 3 startUrl = "https://solo.example.com" pageUrls = ["https://known.example.com/page"] linkLists = ["https://solo.example.com/hidden"] workerCount = 1 return = ["https://solo.example.com"] An unlisted start page is still visited. Its parser result is empty, so the crawl ends immediately. Constraints 1 <= pageUrls.length == linkLists.length <= 120. The input contains at most 480 links in total. 1 <= workerCount <= 4. Every URL is an absolute http or https URL of at most 300 characters, contains no whitespace, and has a lowercase ASCII hostname with no credentials or explicit port. Entries of pageUrls are fragment-free and unique. Paths and query strings remain case-sensitive; only fragments are removed.
Reported by candidates. Source: FastPrep
Pattern and pitfall
It reduces to BFS or DFS over a map from URL to its link list. Normalize first: cut at the first # and drop the rest. Parse the hostname as the text between :// and the next /, ?, or end of string. Compare hostnames only, never schemes. Build a dict from pageUrls to split linkLists, and use an empty list for any URL not in the dict. Add a URL to visited when you enqueue it, not when you pop it, so cycles and duplicate fragments can't double-visit. Sort at the end. The pitfalls: forgetting to normalize the start URL, treating an unlisted URL as unvisitable, and splitting an empty string into one empty token. Filter empties. If you do implement real threads, guard the visited set with a lock, but a sequential simulation returns the same answer. StealthCoder is the hedge if the hostname parsing trips you up live.
StealthCoder is the hedge for the one pattern you didn't drill. It runs invisibly during the screen share.
You can drill Crawl Same-Hostname URLs Concurrently cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. If you're reading this with an OA window open, you're who this was built for.
Get StealthCoderRelated leaked OAs
This OA pattern shows up on LeetCode as web crawler. If you have time before the OA, drill that.
You've seen the question.
Make sure you actually pass Anthropic's OA.
Anthropic reuses patterns across OAs. If you're reading this with an OA window open, you're who this was built for. Works on HackerRank, CodeSignal, CoderPad, and Karat.
Crawl Same-Hostname URLs Concurrently FAQ
What's the trick in the Anthropic crawler problem?+
Normalize before everything, then run a standard BFS with a visited set. Compare hostnames only, ignoring scheme. Treat URLs missing from pageUrls as pages with zero links, and still include them in the output if they match the hostname.
Do I actually need real threads for workerCount?+
The output is a sorted set of visited URLs, so it doesn't depend on scheduling. A sequential BFS gives the correct answer. If you do use threads, protect the visited set with a lock so no URL is added twice.
How do I extract the hostname safely?+
Strip the fragment, then take everything after "://" up to the first /, ?, or end of string. Constraints guarantee lowercase hostnames with no credentials or ports, so no extra cleanup is needed. Compare the result exactly to the start hostname.
What edge cases should I test?+
An unlisted start URL, which is still visited. Links with an empty linkLists entry. Two fragments of the same page. Cycles between pages. Same hostname with different schemes. A URL with a query string, which stays case-sensitive and distinct.
How hard is this really, and how do I prep in 48 hours?+
It's easy to medium. Inputs are tiny, at most 120 pages and 480 links, so efficiency isn't the concern. Spend your time on graph traversal with a visited set and on string splitting. Write the normalize and hostname helpers first, then the BFS.