Reported September 2026
Anthropichash table

Find Duplicate File Groups

Reported by candidates from Anthropic's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Anthropic OA. Under 2s to a working solution.
Founder's read

The mistake that sinks a first attempt on this Anthropic OA, reported in September 2026, is trusting a hash and skipping the real comparison. The task looks easy: group file paths by identical content and return the groups with two or more members. It's a hash-table problem with a sorting finish. The traps are in the details: empty files count as equal, paths sort by raw ASCII, and groups sort by first path. If you blank on the grouping or the ordering, StealthCoder runs invisibly during the live assessment as a safety net and gives you the solution in real time.

The problem

A filesystem traversal has produced two equal-length arrays. Entry i contains the unique absolute path of one regular file and that file's complete content.
Two files are duplicates exactly when their complete contents are equal. Return every group containing at least two duplicate paths. Sort paths inside a group lexicographically, then sort groups by their first path. Every path contains printable ASCII characters, so lexicographic order compares their ASCII character values from left to right. Return an empty array when no duplicate group exists.
An implementation may use content hashes to reduce comparison work, but exact content equality remains authoritative.

Function
findDuplicateFileGroups(paths: String[], contents: String[]) → String[][]

Examples
Example 1
paths = ["/a/x.txt","/b/y.txt","/c/z.txt","/d/w.txt"]
contents = ["red","blue","red","blue"]
return = [["/a/x.txt","/c/z.txt"],["/b/y.txt","/d/w.txt"]]
The two red files form the first group, and the two blue files form the second. Group order follows each group's first sorted path.
Example 2
paths = ["/a","/b","/c"]
contents = ["one","two","three"]
return = []
Every content value is unique, so there is no duplicate group.
Example 3
paths = ["/zero-b","/zero-a","/filled"]
contents = ["","","x"]
return = [["/zero-a","/zero-b"]]
Empty files are valid and equal. Their paths are sorted inside the returned group.

Constraints
0 <= paths.length <= 5000.
paths.length == contents.length.
Paths are unique absolute paths with length from 1 to 100 and contain only printable ASCII characters (code points 32 through 126).
File content may be empty.
The combined number of characters across all paths and contents is at most 500000.

Reported by candidates. Source: FastPrep

Pattern and pitfall

The core move is a hash map from the full content string to a list of paths. Use the content itself as the key. Languages hash strings internally and still check equality on collisions, so exact equality stays authoritative, which is what the problem demands. Don't key on a custom checksum or a length-plus-prefix shortcut. Two different files can collide, and that's the pitfall. After grouping, drop any list with fewer than two paths. Sort each list with plain ASCII comparison, not locale-aware compare. Then sort the groups by their first path. Since paths are unique, there are no ties. Empty content is just a valid key, so don't filter it out. Total input is capped at 500000 characters, so one pass plus sorting is comfortably fast. If the ordering rules trip you mid-assessment, StealthCoder is the hedge that gets you a clean version fast.

The honest play: practice the pattern, and have StealthCoder ready for the one you didn't see coming.

If this hits your live OA

You can drill Find Duplicate File Groups cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built for the candidate who saw this exact problem leak two days before his OA and wondered if anyone had a play.

Get StealthCoder

Related leaked OAs

⏵ The honest play

You've seen the question. Make sure you actually pass Anthropic's OA.

Anthropic reuses patterns across OAs. Built for the candidate who saw this exact problem leak two days before his OA and wondered if anyone had a play. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Find Duplicate File Groups FAQ

What's the trick in Find Duplicate File Groups?+

Map each full content string to a list of paths, keep lists with at least two entries, sort each list, then sort the groups by first path. The real trick is using the content itself as the key so equality is exact, not a weak custom hash.

How hard is this really?+

Easy to medium. The algorithm is a single hash map pass. Most failed attempts come from edge cases: empty files being equal, forgetting to sort inside groups, or sorting groups by the wrong element. Read the output rules twice before coding.

Do empty files count as duplicates?+

Yes. Example 3 shows two empty files forming a group, sorted by path. Don't skip empty strings when building the map. An empty string is a perfectly valid key and should be grouped like any other content.

How should I sort the paths?+

Use plain lexicographic ASCII comparison, comparing character codes left to right. Avoid locale-aware comparators, which can reorder punctuation and case differently. Default string sort in most languages is already correct for printable ASCII.

How do I prepare for this in 48 hours?+

Write the group-by-key pattern from memory twice, once with a map of lists and once with sorting at the end. Then test the three examples plus an empty input array. That covers nearly every edge this problem can throw at you.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Anthropic.

OA at Anthropic?
Invisible during screen share
Get it