Files & Folders

How to hash a ZIP archive on Mac

Every ZIP already carries a CRC32 for each file inside it, so a publisher's SHA-256 next to the download looks redundant until you know what each one covers. Rocket Hash computes the second kind — and this page explains why the two answer different questions, and which one you need.

A ZIP file is unusual among downloads: it already contains checksums. Every entry inside it carries a CRC32 of its own contents, and the thing that unpacks the archive checks them on the way past. So when a download page also publishes a SHA-256 next to the .zip, a fair question is what the second one adds.

Quite a lot, and they are not competing. The CRC32s inside answer “did each file come out of the archive the way it went in?” The SHA-256 of the archive file answers “is this the archive the publisher actually built?” One is an error check performed by your own tooling, the other is a claim somebody else made that you can test.

This page is about knowing which one you are running. The mechanics of pasting a published value and reading a verdict are covered in verifying a download; what follows is the part specific to archives.

Check the archive before you unpack it

Unpacking is the point at which contents touch your disk. The SHA-256 comparison takes seconds and happens while the archive is still one inert file.

Two checksums live in every ZIP

The CRC32 stored inside a ZIP compared with a SHA-256 of the archive file
CRC32 inside the archiveSHA-256 of the .zip
Length8 hex characters, one per entry64 hex characters, one for the file
CoversEach entry's uncompressed contentsEvery byte of the archive: headers, compressed streams, directory, comment
Computed byWhatever made the archive; rechecked by whatever unpacks itYou, against a value the publisher posted somewhere
CatchesAccidental damage, and tells you which entry is badAny change at all, down to a single bit, anywhere in the file
Survives forgeryNo. CRC32 is linear; a value can be forced on purposeYes. No method of constructing a SHA-256 match is known

The first column is genuinely useful and is not security. CRC32 was designed to catch the kind of damage transmission and storage inflict — a burst of flipped bits, a truncated write — and it does that job at enormous speed with eight characters. It is not a cryptographic hash and was never meant to be one, which is a distinction worth holding on to — type a sentence into the CRC32 calculator here, then change a single character of it, and you can watch all eight characters turn over as you type.

Verify an archive, then prove its contents

  1. Hash the archive before you unpack it

    Drop the .zip into the Files tool exactly as it arrived — not a re-saved copy, not the expanded folder. One row appears with its size and digest. If the size already disagrees with the download page, stop here; the rest is wasted effort on a truncated file.

  2. Compare it with the published value

    Use the Verify tool's Against a Checksum mode and paste the publisher's string, and you get a sentence rather than two lines of hex to read across. The algorithm is detected from the checksum itself, so there is nothing to declare and nothing to select.

  3. Let the internal CRC check run as you unpack

    Now expand it. Whatever does the expanding — the Finder, for almost everybody — recomputes each entry's CRC32 as it decompresses and complains if one does not match what the archive claims. That is a second, weaker net underneath the first, and it costs you nothing.

  4. Hash the extracted files if you need to compare contents

    If the real question is whether two copies of the same material are the same — your extraction against a colleague's, or last year's archive against this year's — do it on the extracted files, not on the archives. Drop the unpacked folder in and you get a digest per file, which is the comparison you actually wanted. Hashing a folder covers what that list does and does not mean.

Why zipping the same files twice gives a different digest

This is the single most common surprise with archives, and it is not a bug in anything. Compress the same folder twice and the two .zip files almost certainly have different SHA-256 digests, because a ZIP is not a pure function of its contents. At least five other things get written into it:

  • Timestamps. Each entry records a modification date, so an archive built on Tuesday differs from one built on Wednesday even if nothing inside changed.
  • Entry order. The order files were added is the order they are stored, and that depends on how the archive was made.
  • Compression choices. A different compression level — or a different implementation of the same algorithm — produces different compressed bytes from identical input.
  • Mac metadata. When files carry extended attributes or resource forks, archives made on macOS commonly include a parallel __MACOSX tree of ._ entries to hold them. Archives made elsewhere do not.
  • Odds and ends. Whether folders get their own entries, whether an archive comment was set, which extra header fields the tool chose to write.

So the digest of a .zip is a fingerprint of one particular act of archiving, not of the material inside. Two archives with different digests may hold byte-identical files; two archives with the same digest are the same file, full stop. The practical rule: to compare contents, compare the extracted files, ideally through a manifest — see checking files against a manifest.

What the file row tells you about an archive

An archive dropped into the Files tool becomes one row: the file name in bold, the folder it came from underneath it — ~/Downloads, usually — the size, and the digest, middle-truncated so that 64 characters fit on a line. The status bar along the bottom summarizes the batch on the left and the progress on the right, so a row with no digest against it yet is a read still running rather than an archive that failed.

Two controls earn their keep on archives. The Algorithm control in the toolbar, marked with a #, decides which digest the rows show, which is the answer to a publisher who posted an MD5 rather than a SHA-256. The chevron at the end of a row expands that one file into all eight algorithms at once — worth doing when you are about to publish an archive and do not yet know which digest the people downloading it will ask for. Either way the archive is read once: a single pass over the disk produces every digest requested, so eight of them cost what one costs.

The export button turns those rows into a text file in the two-column shape checksum tools have read for thirty years — digest, two spaces, filename. An export covering an archive and the files that came out of it looks like this:

SHA256SUMS
e37287b0ddbda7edb55845acc6de4244cd82e14df5824bdd71007a452a7cbaab  release-2026.zip
2898d907fba3a01bbe9e078e73fb9c20c33a69fc322d062c61e826f2798e3621  release-2026/CHANGELOG.md
987c6e569e2a352a6909764d5453afd773d0d7e77cb82fbd6f58747afbfd1f71  release-2026/app.bin

Read that as a record rather than a recipe. The first line is the claim you can hold against the value the publisher posted; the lines beneath it are what the extracted contents were at the moment you looked, which is the half that still means something in a year. How to read a SHA256SUMS file takes one of those lines apart field by field, including the asterisk that sometimes sits in front of a filename.

One thing a row cannot do, and it is worth being precise about it. An archive that expands without a single complaint has told you something real: every entry decompressed to what the archive claimed for it, so nothing rotted on the drive and nothing was truncated in transit. It has not told you the archive is the one the publisher built, because anybody who rebuilt it wrote fresh, perfectly valid CRC32s at the same time. Only the digest of the whole file, set against a value from somewhere other than the archive, answers that.

What a matching archive checksum does not tell you

It does not tell you the contents are safe. A verified archive is the archive the publisher built, nothing more; if what they built contains something you would rather not run, the digest matched all the way in. It also says nothing about how enormous the contents are once expanded, which is worth a thought before you unpack something unfamiliar onto a nearly full disk.

And it is only as good as where the value came from. A checksum printed on the same page, served over the same connection, by the same party that served you the archive proves the transfer was clean and nothing else. For a .dmg there is at least a signature in the picture as well — hashing a DMG covers how that interacts with what Gatekeeper already checked. For a plain .zip, there is usually nothing but the checksum, so where you read it matters.

Troubleshooting

Unzipping reports a CRC error

One entry decompressed to something other than what the archive says it should be, which means the archive is damaged rather than merely awkward. Check the archive's SHA-256 against the published value before anything else: if that also fails, you have an incomplete or corrupted download and the fix is to fetch it again. If the SHA-256 matches and a CRC still fails, the published archive is itself broken — which happens, and is worth telling the publisher.

The archive will not open at all

A ZIP keeps its index at the end of the file, so a truncated download is not a partly usable archive — it is an unreadable one, and nothing will get as far as checking a CRC. Compare the file's size with the figure on the download page; a short file is a short download. When a checksum does not match covers the rest of the causes in order.

I rebuilt the ZIP and the checksum changed

Expected, for all the reasons above — the per-entry modification times alone guarantee it, so two archives built a minute apart from identical files will not agree. If you need a reproducible archive, that is a build-tooling problem rather than a hashing one. If you only need to prove the files are the same, drop both unpacked folders into the Files tool and compare the per-file digests.

The publisher only gives an MD5 for the archive

Use it — then be precise about what you got. MD5 collisions have been trivially constructible since 2004, and SHA-1's were demonstrated publicly in 2017 with SHAttered, but both of those are an attacker who gets to choose both files. Taking a digest somebody posted years ago and building a different archive that matches it is a second-preimage attack, and no practical one exists against either function. So an old published MD5 still tells you your download is the publisher's archive; what it cannot do is make a new file's digest trustworthy on its own. Is MD5 still safe draws that line properly.

The numbers inside the archive are only eight characters

That is correct for CRC32: 32 bits, eight hex characters, one per entry. It is not a truncated SHA-256 and there is no way to lengthen it. If you want a digest of an archive you can quote to somebody, it is the SHA-256 of the .zip you want — 64 characters, computed over the whole file.

Frequently asked questions

Should I hash the ZIP file or the files inside it?

Hash the .zip when you are checking it against a value a publisher posted — that is what their value describes. Hash the extracted files when the question is whether two sets of contents are the same, because two archives built from identical files usually have different digests. Both are valid; they answer different questions.

What is the CRC32 inside a ZIP file for?

It lets whatever unpacks the archive confirm each entry came out as it went in. The archive stores an 8-character CRC32 of every entry's uncompressed contents, and the extractor recomputes it while decompressing. It is excellent at catching accidental corruption and it tells you which entry is bad, which a single digest of the whole file cannot.

Why do two ZIP files with the same contents have different checksums?

Because a ZIP records more than its contents: per-entry modification times, the order the files were added, the compression level, any Mac metadata stored alongside, and whatever extra header fields the tool wrote. Change any of those and every byte downstream changes. A ZIP's digest fingerprints one act of archiving, not the material inside.

Is the CRC32 check enough to prove a ZIP has not been tampered with?

No. CRC32 is an error-detecting code, not a cryptographic hash: it is linear, and a specific value can be forced deliberately. Anyone who modified an archive would also write fresh CRC32s that check out perfectly. For tampering you need a digest of the whole archive compared against a value from a source you trust.

Do I need to unzip a file before checking its checksum?

No, and you should not. The published checksum describes the archive as downloaded, so expanding it first throws away the thing you were going to check. Verify the .zip, then unpack it.

How do I verify a ZIP downloaded from GitHub?

Check release assets — the files a maintainer uploaded — against whatever checksums they published with the release. Source archives that a site generates on demand are a different matter: they are built when you ask for them, so their bytes are not guaranteed to be identical every time, and a digest you recorded last year may not reproduce. Treat those as convenience downloads rather than as fixed artifacts.

Does renaming a .zip change its checksum?

No. A digest is computed from the file's contents, and the name is not part of the contents. Renaming, moving or copying an archive leaves its SHA-256 exactly as it was — which is why a checksum can identify a file somebody has renamed.