Reported February 2026
Postmanhash table

Group Duplicate Files by Content

Reported by candidates from Postman's online assessment. Pattern, common pitfall, and the honest play if you blank under the timer.

Get StealthCoderRuns invisibly during the live Postman OA. Under 2s to a working solution.
Founder's read

The checksum line in the Postman problem statement is the detail that matters. Postman reported this one in February 2026, and it's a hash-table grouping problem dressed up as a filesystem task. You get parallel arrays of paths and contents, and you return comma-joined groups of files with exactly equal contents. Singletons get dropped. Sorting happens twice. None of it is hard, but the output rules are where people lose points. If you blank on the exact shape during the live OA, StealthCoder runs invisibly as a safety net and reads the problem for you.

The problem

You are given parallel arrays paths and contents, where contents[i] is the complete content of the file at paths[i].
Group files whose contents are exactly equal. Return one string for every group containing at least two files. Within a group, sort paths lexicographically and join them with a comma. Sort the returned group strings lexicographically.
A production filesystem implementation may use a content checksum to find candidates, but it must verify exact content equality before declaring duplicates so checksum collisions cannot merge different files.

Function
groupDuplicateFiles(paths: String[], contents: String[]) → String[]

Examples
Example 1
paths = ["/a/x.txt","/b/y.txt","/c/z.txt"]
contents = ["hello","hello","bye"]
return = ["/a/x.txt,/b/y.txt"]
The first two files have exactly equal contents. The third file is unique.
Example 2
paths = ["b","a","d","c"]
contents = ["1","2","1","2"]
return = ["a,c","b,d"]
Paths are sorted inside each duplicate group, and the group strings are sorted before returning.

Constraints
1 ≤ paths.length = contents.length ≤ 10^5
Every path is unique and non-empty.
The total length of all paths and contents is at most 10^6.
Paths and contents may contain spaces, but paths do not contain commas.

Reported by candidates. Source: FastPrep

Pattern and pitfall

The trick is a hash map from content string to a list of paths. Walk both arrays once, append each path under its content key. Then keep only lists with two or more entries, sort each list lexicographically, join with a comma, and sort the final list of strings. The checksum note is a distraction for your code. Using the full content string as the map key already gives exact equality, so collisions can't merge different files. Don't hash to an int and trust it. The common pitfalls are forgetting to sort inside each group, sorting only the groups, and including singletons. Total input is capped at 10^6 characters, so keying on the raw string is fine. Sorting cost is bounded by the total path length times log. If you freeze on the output formatting, StealthCoder is the hedge during the live OA, but the logic fits in about ten lines.

If you see this problem in your OA tomorrow, the play is to recognize the pattern in 30 seconds. StealthCoder buys you that recognition.

If this hits your live OA

You can drill Group Duplicate Files by Content cold, or you can hedge it. StealthCoder runs invisibly during screen share and surfaces a working solution in under 2 seconds. The proctor sees the IDE. They don't see what's behind it. Built by an Amazon engineer who passed his OA cold and still thinks the filter is broken.

Get StealthCoder

Related leaked OAs

⏵ Practice the LeetCode equivalent

This OA pattern shows up on LeetCode as find duplicate file in system. If you have time before the OA, drill that.

⏵ The honest play

You've seen the question. Make sure you actually pass Postman's OA.

Postman reuses patterns across OAs. Built by an Amazon engineer who passed his OA cold and still thinks the filter is broken. Works on HackerRank, CodeSignal, CoderPad, and Karat.

Group Duplicate Files by Content FAQ

What's the trick in the Postman group duplicate files problem?+

Use a hash map keyed by the full content string, with a list of paths as the value. Dictionary key equality is exact, so you get the collision-safe check the statement asks for. Then filter groups of size two or more, sort, join, and sort again.

Do I need to implement a checksum?+

No. The statement says a production system may use one, but it must verify exact equality anyway. Keying your map on the raw content already does that. Adding a checksum only adds a bug surface and no benefit at these input sizes.

What are the easy mistakes on this one?+

Returning singleton files, sorting only the final list and not the paths inside each group, and joining with the wrong separator. Also watch for sorting group strings versus sorting groups of lists. Sort the joined strings lexicographically at the end.

How hard is this really?+

Easy to medium. The algorithm is a single pass with a hash map. The difficulty is in following the output spec exactly. With up to 10^5 files and 10^6 total characters, a straightforward approach runs comfortably.

How do I prepare in 48 hours?+

Write the solution once from scratch in your language. Test it on both examples, especially the second one where order flips. Then practice two or three other group-by-key problems so the map-of-lists pattern is automatic before the OA.

Problem reported by candidates from a real Online Assessment. Sourced from a publicly-available candidate-aggregated repository. Not affiliated with Postman.

OA at Postman?
Invisible during screen share
Get it