Reported September 2026
Harveybreadth first search

Find Duplicate Files in a Filesystem

Reported by candidates from Harvey's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Harvey OA. Under 2s to a working solution.
Founder's read

Harvey's September 2026 OA hands you a fake filesystem and asks for duplicate files, and the input size (2000 entries, 200000 total bytes) tells you brute-force pairwise comparison isn't the point. The real test is the traversal rules. It's a design-flavored simulation: BFS from root, resolve symlinks, skip broken links and cycles, dedupe by resolved file, then group by content. Nothing here is hard algorithmically. It's easy to get one rule wrong. If you blank mid-assessment, StealthCoder is the silent backup running on your screen.

The problem

Implement a deterministic snapshot adapter for find_dups(root). Given a starting directory root and a filesystem snapshot entries, traverse reachable entries with iterative breadth-first search and return groups of duplicate regular files.
Each entry is [path, kind, target, content, readable]:
kind is D for a directory, F for a regular file, or L for a symbolic link.
A symbolic link stores its destination path in target. Other entries use an empty target.
A regular file stores its exact byte content in content. Other entries use empty content.
readable is 1 or 0. An unreadable entry is skipped.
Directory children are the entries whose normalized parent path is that directory. Process children in lexicographic path order. Resolve symbolic-link chains; skip a broken link or a link cycle. Visit each resolved directory at most once, so directory links cannot make traversal infinite.
Include each resolved regular file at most once. If several reachable paths resolve to the same file, represent it by the lexicographically smallest path. Group different files only when their contents are byte-for-byte equal. A memory-bounded implementation may bucket by byte length, compute SHA-256 with a fixed one-megabyte buffer, and confirm equal-hash candidates byte-for-byte.
Return only groups containing at least two files. Sort paths inside each group lexicographically, then sort groups by their first path.

Function
findDuplicateFiles(root: String, entries: String[][]) → String[][]

Examples
Example 1
root = "/"
entries = [["/","D","","","1"],["/a.txt","F","","same","1"],["/b.txt","F","","same","1"],["/c.txt","F","","other","1"]]
return = [["/a.txt","/b.txt"]]
The two readable files with content same form one duplicate group. The third file has different content.
Example 2
root = "/r"
entries = [["/r","D","","","1"],["/r/a","D","","","1"],["/r/a/one","F","","x","1"],["/r/two","F","","x","1"],["/r/a/back","L","/r","","1"],["/r/alias","L","/r/a/one","","1"]]
return = [["/r/a/one","/r/two"]]
The directory link back to /r is bounded by the visited-directory set. The link /r/alias and /r/a/one identify the same file, so that file appears once under the smaller path.
Example 3
root = "/r"
entries = [["/r","D","","","1"],["/r/a","F","","same","1"],["/r/b","F","","same","0"]]
return = []
The unreadable file is skipped, leaving no content shared by two reachable files.

Constraints
1 <= entries.length <= 2000.
Every entry contains exactly five strings.
Every path is unique, normalized, absolute, and contains no trailing slash except /.
root names a readable directory entry.
Every symbolic-link destination is a normalized absolute path.
The sum of regular-file content lengths is at most 200000.
readable is exactly 0 or 1.

Reported by candidates. Source: FastPrep

Pattern and pitfall

The trick is to separate three jobs. First, a resolve function: follow link targets through a map of path to entry, track a seen set to catch cycles, return null on a missing target or unreadable entry. Second, BFS with a queue and a visited set of resolved directories, sorting children by path before enqueueing. Third, a map from resolved file path to the smallest reachable path that points to it, then group by content. Bucketing by length first, then hashing, then comparing bytes is optional at 200000 total bytes, but a dict keyed on content works fine. The common pitfalls: forgetting that a link to a file and the file itself are the same file, treating an unreadable link target as valid, and sorting groups by something other than the first path. Check example 2 by hand. StealthCoder is your hedge if the link-resolution edge cases tangle you live during the OA.

StealthCoder is the hedge for the one pattern you didn't drill. It runs invisibly during the screen share.

If this hits your live OA

You can drill Find Duplicate Files in a Filesystem cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. If you're reading this with an OA window open, you're who this was built for.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Harvey's OA.

Harvey reuses patterns across OAs. If you're reading this with an OA window open, you're who this was built for. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Find Duplicate Files in a Filesystem FAQ

How hard is the Harvey duplicate files OA really?+

Medium on paper, fiddly in practice. The algorithms are BFS, a hash map and sorting. The difficulty is the rule count: symlink chains, cycles, unreadable entries, visited directories and smallest-path dedupe. Miss one rule and a hidden test fails.

What's the trick to symlink handling?+

Write one resolve(path) helper with a seen set. Loop while the entry is a link, jump to its target, and return null if the target is missing, unreadable or already seen. Everything else then works on resolved paths only.

Do I need SHA-256 and a one-megabyte buffer?+

The statement says a memory-bounded implementation may do that, so it's optional. Total content is capped at 200000 characters, so grouping by the content string directly is correct and simpler. Bucket by length first if you want to match the described approach.

How do I handle two paths pointing to the same file?+

Keep a map from resolved file path to the smallest reachable path that resolves to it. Since you process children in sorted order, update the map with a min check anyway. Then group only those unique files by content.

How do I prepare in 48 hours?+

Practice writing BFS with a visited set, then write a symlink resolver with cycle detection. Hand-trace the three examples, especially example 2 with the back link and alias. Finally, check output ordering: paths inside groups, then groups by first path.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Harvey.

OA at Harvey?
Invisible during screen share
Get it