Verifying Downloads

How to find duplicate files by hash on Mac

Sort a list of digests and you get the one thing no amount of squinting at names, dates and sizes will give you: the exact sets of files on this Mac that hold identical bytes. Rocket Hash produces that list from a single pass over the folders, and this page covers what to do with it — including which copy to keep.

Four places claim to hold your photographs: the library itself, a folder called Photos copy, something called Desktop stuff (old Mac), and an external drive from 2019 that may or may not be a subset of the other three. Several thousand files in there exist twice over, and the Finder cannot tell you which ones, because every field it offers to sort by describes the file rather than its contents.

Names are the worst of them: a copy gets a suffix, a download gets (1), an import renames by date, and two files with completely different names are routinely the same bytes. Dates are nearly as bad, since copying a file can reset one timestamp while preserving the other. Size is the only honest field, and it only ever rules things out — thousands of unrelated files share a size, and every empty file in the world shares one.

A digest has none of those problems. It is derived from the contents and nothing else, so two files with the same digest hold the same bytes no matter what they are called or where they live. This page is the workflow that turns that into a list you can act on: hash everything once, export the list, and sort it so the duplicates line up next to each other.

Where the work happens

Nothing here hunts for duplicates on your behalf and then offers to tidy up. Hashing produces the one column that makes duplicates obvious, and sorting the exported list does the grouping. The deleting stays a decision you make with your eyes open, which given what is usually in these folders is the right division of labor.

Why the digest is the field that settles it

Two files with the same SHA-256 are the same file, for every practical purpose: no two different inputs sharing a SHA-256 digest have ever been found, by accident in anybody’s library or on purpose in a lab. Treating equal digests as equal contents across two hundred thousand files is sound in a way that sorting by name never is.

It does not hold for every algorithm, and this is the one job where the choice genuinely bites. Grouping by digest means you are relying on accidental collisions being impossible at your scale, and that depends on how many values the algorithm has:

How many files it takes before an accidental digest collision becomes likely, by algorithm
AlgorithmDigestFiles before an accidental collision is likely
CRC328 hex charactersAbout 77,000 — a single photo library
MD532Around 18 quintillion
SHA-25664More files than will ever exist

That first row is the trap. CRC32 has only 4.3 billion possible values, and the birthday problem puts a 50/50 chance of two unrelated files colliding at around 77,000 of them — so on a real library it will confidently group files that have nothing to do with each other. It is an excellent error-detecting code doing the job it was built for, which is not this one; what CRC32 is actually used for explains the difference. MD5 is statistically fine here; the only reason to prefer SHA-256 is that a deliberately crafted MD5 pair is cheap to make, which matters if the files came from somewhere you do not control — is MD5 still safe has the rest. The run is one pass over the disk whichever you pick, so there is nothing to gain from the weaker guarantee.

Hash everything once

Everything has to be in the same run, because being a duplicate is a relationship rather than a property: a file hashed on Tuesday and another hashed on Thursday tell you nothing at all until both lines are sitting in one list. That makes this a job for the Files tool in Rocket Hash rather than for anything that takes one file at a time.

  1. Drop every folder in at the same time

    Drag all the candidates into Files together — the suspect folder, the copy of the suspect folder, the external drive. Each file arrives as a row showing its name, the folder it came from and its size, and the status bar totals the queue, so you can check the file count and the byte total against what you meant to drop in before a single byte is read. A hundred thousand files is a normal size for this job rather than an extreme one, and memory stays flat at around 16 MB however many there are, because files are streamed rather than loaded.

  2. Use one algorithm for the whole run

    Put the Algorithm control on SHA-256 so every row in the run is comparable with every other. Each file is read exactly once and all eight digests come out of that single pass, so there is no speed penalty for curiosity — but a single algorithm is what you want in the exported list, because the sort you are about to do compares one column.

  3. Let it run, and watch the counter

    The status bar summarizes the queue on one side and progress on the other, with live throughput and an estimate of what is left. Expect this to be the slow part: deduplicating by hash means reading every byte of every file, and a folder of many small files runs well below the rate a single large file would, because each one costs an open and a close. A long run can be paused and picked up from the same byte.

  4. Export the list

    Use the export control when the run finishes. What you get is one line per file with the digest as the first column, which is exactly the shape the next section needs: the sort groups duplicates only because the digest comes first. Exporting a manifest covers the details, and hashing many files at once covers managing a queue this size.

Let the duplicates group themselves

Because the digest is the first field on every line, sorting the exported list sorts it by content — and any digest appearing more than once marks a set of files holding identical bytes. Sort on that first column in whatever you read text in, and every copy lands next to its siblings: one digest, then every path those bytes are living at.

Read the shape before the detail. A handful of heavily repeated digests means a few big wins: twelve copies of one 4 GB video is ten minutes of work and a lot of space back. A long tail of pairs is a filing problem that no delete key fixes.

For a small suspect set you do not need the export at all. Every row in Files carries its own digest, middle-truncated so the rows stay readable, and two files holding the same bytes show the same characters at both ends — with a dozen files dropped in together, the duplicates are visible in the window. The copy button on a row puts the full digest on the clipboard.

Two exported lists can be compared instead of one, which covers volumes that are never mounted together: hash the laptop this week, the drive next week, and set the two lists side by side. By then you are comparing 64-character strings rather than files, so nothing has to be plugged in.

What counts as a duplicate, and what does not

This method finds byte-for-byte identical files and nothing else, which is narrower than what people often mean by “duplicate”. It will not group:

  • The same photograph exported twice. Two exports of one raw file differ in their embedded timestamps and encoder output even at identical settings, so the bytes differ and so do the digests. Hashing a photo or video file goes into why an export never matches its source.
  • A resized, re-compressed or rotated copy. Visually the same image; numerically a different file.
  • The same document re-saved. Many applications rewrite internal metadata on every save, so an untouched round trip through the app can produce different bytes.
  • The same song with different tags. Tags live inside the file, so editing one changes the digest while the audio is untouched.

For the narrow job — finding exact copies so you can reclaim space or collapse three archive folders into one — that precision is the feature. If what you want is visually similar images, hashing is the wrong instrument entirely.

Confirming a pair byte for byte

A line in a list is a claim, and for a file you cannot replace it is worth having that claim checked by something that reads both files again. Verify in File vs. File mode does that: two drop wells side by side with a swap control between them, the keeper in one and the candidate in the other. The verdict appears underneath as a green seal and a sentence — “Files are identical.”, with “Verified byte for byte with SHA-256.” below it.

Each well has a Replace… link, which makes the mode practical for a set rather than a pair: leave the copy you intend to keep in place and replace the other side as you work down the paths that shared its digest. Comparing two files covers the route on its own.

Before you delete anything

A list of duplicate sets is not a list of things to delete, and the gap between the two is where people lose files. Four habits make the difference.

  1. Decide which copy is canonical by path, not by name. The one in the organized library wins; the one in Desktop stuff (old Mac) loses, regardless of which has the tidier filename.
  2. Check what points at the file. Photo libraries, video projects, page layouts and build systems all reference files by path. A duplicate can be the copy something else depends on.
  3. Move, do not delete. Everything you have decided against goes to one holding folder for a month. If nothing has broken by then, empty it.
  4. Confirm the pair before you act on a big one. For anything irreplaceable, take the verdict from Verify rather than from a line in a list, as above.
The classic way to lose everything

A script that deletes every file whose digest appears more than once deletes all the copies, including the one you were keeping. Any automated pass must keep one member of each set by an explicit rule, and the cheapest way to get that wrong is to write it at the end of a long day.

Troubleshooting

Two files I know are identical have different digests

Then they are not identical, and the difference is usually invisible in the application you look at them through. Embedded metadata is the most common culprit — EXIF fields, an ID3 tag, a PDF’s internal modification date — followed by one copy having been re-saved or re-exported rather than copied. A partial copy will also do it, and is worth ruling out by comparing the two sizes first.

Thousands of duplicates that are all tiny files

Those are real duplicates and useless ones: .DS_Store files, empty placeholders, identical stub icons, build artifacts. Every empty file on earth has the same SHA-256 — e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 — so they all group perfectly and none of it means anything. Filter by size before you read the results, since nothing under a few kilobytes is worth anyone’s attention.

I deleted duplicates and freed no space

You were looking at links or clones rather than second copies. A hard link is two names for one file; an APFS clone — which a plain Finder copy on the same volume can create — shares its storage invisibly until one side is modified. In both cases the digests matched honestly, because the bytes really are identical: there was only ever one set of them to reclaim. The tell is that the volume’s free space does not move when one name goes away, which is worth knowing before a cleanup is planned around a number.

The run is going to take all afternoon

Hashing a terabyte means reading a terabyte, and there is no shortcut inside the algorithm — at roughly 2 GB/s on Apple Silicon, SHA-256 would be finished with that terabyte in about eight minutes, and no ordinary disk will feed it that fast. The shortcut is outside it: a file whose size is unique cannot possibly have a duplicate. Open the candidates in a Finder window in list view, sort by Size, and the repeated numbers show you where duplicates can exist at all — hash those, rather than everything you own. A long run can also be paused mid-file and resumed from the same byte.

The same file is listed twice in one run

That means the same file went into the queue twice — typically because you dropped in both a folder and something inside it, or two aliases resolving to one place. It is harmless but it inflates the duplicate count, so check the overlap in what you dropped before you trust a dramatic number. Hashing a folder covers how a folder run walks what is underneath it.

Frequently asked questions

How do I find duplicate files on a Mac without anything deleting them for me?

Hash everything once, export the results as a list of digests and paths, then sort that list by the digest column — any digest appearing more than once marks a set of byte-identical files. That gives you the findings without anything acting on them, which is the right order for a folder full of photographs. Deleting stays a separate, deliberate step.

Can two different files have the same hash?

For SHA-256, no pair has ever been found and none is expected, so equal digests can be treated as equal contents. For CRC32 it happens readily by accident — with only 4.3 billion possible values, a set of around 77,000 files already has a coin-flip chance of containing an unrelated pair that collides. That is why the algorithm you pick matters more for deduplication than for almost anything else.

Is it faster to compare files by size before hashing them?

Much, and it is the standard trick: a file whose size is unique in the set cannot have a duplicate, so it never needs reading. A Finder window in list view, sorted by Size, shows you which sizes repeat; hash only the files that share one with at least one other. On a large library that removes the great majority of the reading.

Will hashing find duplicate photos that look the same but are not identical?

No. Two exports of the same raw file differ in their embedded timestamps and encoder output, so their bytes differ and their digests differ — and a resized or re-compressed copy is a completely different file numerically. Hashing finds exact copies with total precision, which is the opposite of what perceptual similarity matching does.

Why did deleting duplicates not free up any disk space?

Because the copies were sharing their storage. A hard link is two names for one file, and an APFS clone — which a Finder copy on the same volume can create — shares its data invisibly until one side is modified. In both cases the digests match honestly and there was never a second copy of the bytes to reclaim, so the free space does not move when one of the names goes away.

Can I use CRC32 to find duplicate files?

No — it is the one algorithm that actively fails at this. Eight hexadecimal characters give 4.3 billion possible values, and the birthday problem puts a 50/50 chance of two unrelated files colliding at roughly 77,000 files, which is an ordinary photo library. The result is files reported as duplicates that have nothing in common. Use SHA-256; the disk is the bottleneck, so nothing is saved by choosing something shorter.