Verifying Downloads

What to do when a checksum does not match

A no-match verdict is a fact, not a diagnosis: it tells you the bytes differ and nothing whatsoever about why. Nine situations produce one, eight of them are dull, and the dull ones are overwhelmingly the likely ones — so work through them in the order they actually happen, with the file and the published value in front of you in Rocket Hash, before you conclude that somebody did this to you on purpose.

You found the checksum, you dropped the file in, and the answer came back negative. That verdict is a statement about two strings and nothing more. It says the bytes differ somewhere, and it says nothing about where, by how much, or whose fault it is.

Nine situations produce that outcome, and they are nowhere near equally likely. An interrupted download and the wrong file account for more mismatches than the other seven put together, and both are ruled out in about fifteen seconds. The cause everybody jumps to — somebody replaced the file — is ninth, and it only becomes the answer once the other eight are gone.

So the order matters more than the list. Below is the two-minute route that eliminates most of them at once, then all nine ranked by how often each turns out to be the explanation, then what to do if you reach the bottom. If you have not run the comparison itself yet, checking a file’s checksum on a Mac covers the mechanics.

Cheapest check first

Compare the size of your copy with the size the publisher lists. It takes five seconds, needs no tools at all, and catches the single most common cause outright.

What a no-match verdict actually tells you

A digest is a function of every byte in the file, so a difference anywhere produces a different digest. Flip one bit in a 5 GB image and about half the bits of its SHA-256 flip with it, which means very nearly all 64 hexadecimal characters come out different. There is no gradient: two digests are equal, or they look unrelated, with nothing in between for you to interpret.

Two consequences follow. The first is that a mismatch is never the algorithm’s fault — the same bytes always produce the same digest, on any machine, in any decade, so if two digests differ then two different sets of bytes went in. The second is that the digest cannot tell you how bad the difference is: a missing final kilobyte and a completely different file look identical from where you are standing. So the work ahead is diagnostic, not cryptographic: you are working out which of the two things you compared is not what you assumed it was.

The two-minute route

Five checks in this order, because each one costs more than the one before it. Most mismatches do not survive the second.

  1. Compare the file size

    Get the byte count of your copy — the Finder’s Get Info panel will tell you, and so will the size on the file’s row in the Files tool — and hold it against the size on the download page. An interrupted transfer is almost always visibly short, and the extreme case has a signature worth memorizing: a zero-byte file has a SHA-256 of e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855, the digest of nothing at all.

  2. Confirm it is the file you think it is

    Read the whole filename and the folder it is sitting in, not the first few characters of it. A (1) before the extension, a .part or .crdownload suffix, and last quarter’s build still in ~/Downloads are all extremely good at producing an alarming mismatch for an entirely boring reason.

  3. Count the characters on both sides

    If the published value and yours are different lengths, they are different algorithms and the comparison was never going to succeed: 32 characters is MD5, 40 is SHA-1, 64 is SHA-256, 128 is SHA-512. If they are the same length you are probably fine, with two traps — 64 characters is SHA-256 or SHA3-256, and 128 is SHA-512 or SHA3-512, and each pair is two unrelated functions rather than two versions of one. Working out which algorithm a checksum uses settles the ambiguous lengths.

  4. Copy both values again

    Go back to the original page and re-copy the published digest rather than reusing the one in the chat message, the ticket or the notes file where it has had an opportunity to pick up a line break. A trailing space, a non-breaking space lifted out of a styled web page, or a digest that wrapped across two lines in a PDF will all produce a no-match for a value that is otherwise perfectly correct.

  5. Download it again from the source

    Fetch the file a second time, from the publisher’s own site rather than a mirror or a link somebody forwarded, and hash the new copy. This is the step that splits the nine causes into two groups. If the second copy matches, the first download was damaged and you are finished. If both copies share a digest that still disagrees with the published value, your side of this is clean — the file being served is not the file the checksum describes, which is a different and more serious situation.

Nine causes, in order of likelihood

Ranked by frequency rather than by drama. The first four cover the overwhelming majority.

Causes of a checksum mismatch, most likely first, with the evidence that identifies each one
#CauseWhat gives it away
1The download stopped shortYour copy is smaller than the published size
2You hashed a different fileA (1), a .part, or an older build in the same folder
3The checksum is for a different buildVersion, edition or architecture differs — arm64 against amd64 is the classic
4The two values are different algorithmsDifferent lengths, or both 64 characters and one of them is SHA-3
5The paste carried something invisibleTwo values that look identical character for character still disagree
6The published checksum is out of dateThe release was re-rolled; the directory listing disagrees with the announcement
7The file was still being writtenA download in progress, a syncing folder, a log, a mounted disk image
8Storage or transfer corrupted it quietlyIt matched when it arrived and does not now
9The file is not the published fileTwo independent downloads agree with each other and not with the checksum

The download stopped short

A transfer can end early and still look finished. A captive portal that re-authenticated mid-stream, a disk that filled up, a server that closed the connection at 99%, a resume that picked up from the wrong offset — in every one of those cases you are left with a file that has a plausible name, a plausible icon and a missing tail, and a browser that has already shown you a tick.

Re-download from the publisher’s own page rather than resuming the copy you have, then drop the new file into the Files tool and read the size on its row against the size the publisher lists. A short file is visible there before you compare a single hexadecimal character, which is why this step costs seconds rather than minutes.

The checksum belongs to a different build

This is the cause that wastes the most time, because nothing about it looks wrong. Point releases are the usual culprit — two images several months apart can carry the same headline version number — followed closely by architecture, where an Apple Silicon Mac wants an arm64 image and the value you copied was for amd64. Editions, language variants and “latest” links that quietly moved on do the rest.

The structural fix is to take the checksum from the same directory as the image rather than from an announcement, a wiki or a mirror’s front page: publishers put the two side by side precisely so the list cannot describe a different build from the file beside it. Verifying a Linux ISO shows what that looks like in practice.

The paste carried something invisible

Hexadecimal is made entirely of visible characters, which makes it easy to forget how much invisible material a clipboard can carry. A trailing newline, a non-breaking space from a styled page, a zero-width character from a documentation site, a soft hyphen where a PDF broke the line: none of them show up, and every one of them makes two identical-looking strings unequal.

You do not have to identify which one it was. Re-copy the value, and if it still fails, retype the last eight characters by hand — if that fixes it, the paste was carrying a passenger and the digest itself was always correct. Case, by contrast, is never the cause — A3F0 and a3f0 are one value typed two ways, and a comparison that trips on that is comparing text.

Quiet corruption between the disk and you

A file that matched once and does not now is a different story from a download. The candidates are a failing SSD, a bad cable, a flaky network share, or an archive that has sat untouched on a drive for six years. This is the case checksums were invented for, and it is the reason a manifest you made last year is worth more than it looked like at the time.

Re-copy from the original if you still have one, re-hash, and compare the two digests with each other. If the fresh copy matches your old record and the file on the drive does not, you have found a dying drive rather than a verification problem. Verifying a file after a transfer covers doing this deliberately rather than by accident.

When you have ruled the other eight out

If two independent downloads agree with each other and disagree with the published checksum, stop treating it as a puzzle. Whether this is an attack, a broken mirror or a publisher who forgot to update a page, all three are handled the same way.

  • Do not run it, open it, or unpack it to have a look. Unpacking an archive means running your own parser on somebody else’s data, which is not the neutral act it feels like.
  • Find the checksum somewhere else. The project’s repository, its release notes, a package manager’s metadata, a signed manifest, the same page over a different network. Two independent sources that agree tell you which side of your comparison is wrong.
  • Prefer a signature where one exists. A checksum published beside a file proves the transfer was clean; a signature binds that checksum to a key, which is the part that survives a mirror being compromised. What a checksum proves about tampering is the longer version.
  • Tell the publisher. A stale checksum or a bad mirror is common, fixable, and invisible to them until somebody says so.

Mismatches on files you hashed yourself

When the value you are comparing against is one you produced — a manifest from last month, a digest noted before a move — the nine causes above re-rank completely. Nobody tampered with anything, the file legitimately changed, and three traps account for most of it. Archive formats store modification times inside themselves, so re-zipping identical contents tomorrow produces a different digest. Re-exporting a photo or a video re-encodes it, which means new bytes even when the picture is indistinguishable. And anything with an open handle hands you a digest for a state that lasted milliseconds — hashing a file that is still being written goes into that one.

Troubleshooting

The ends match but the verdict is still no match

You are probably comparing two abbreviations. Interfaces that display long digests shorten them in the middle with an ellipsis, so two genuinely different values can be identical at both visible ends. Copy both values out in full and compare those, not what is on screen — and this is the argument for letting something else do the comparing. Rocket Hash answers with a sentence and a seal rather than two strings for you to scan, which removes the failure mode instead of asking you to be careful.

If you have checked both values in full and they really are identical character for character, the difference is whitespace you cannot see. Go back to the fourth step.

The same file matches on one machine and not another

Then the two machines do not have the same file, or the two tools are not reading the same bytes. Sync services are the usual explanation — a placeholder that has not downloaded yet, or one machine holding a conflicted copy. Hash both copies with the same algorithm and compare them to each other rather than to the published value; if those two differ it is a file problem, not a tool problem. Why two tools give different hashes covers the cases where it really is the tool.

The publisher lists two checksums and neither matches

Look at which filename each line names. Releases routinely publish digests for the compressed download and for the image inside it, or for an installer and for the archive containing the installer, and hashing the wrong member of that pair gives you two failures instead of one. Reading a SHA256SUMS file covers the two-column format and what the asterisk before a filename means.

There is no published checksum to compare against

Then you cannot verify the download at all, only compare copies of it. Two downloads over two different networks producing the same digest is weak evidence, but it rules out a damaged transfer. The more useful move is forward-looking: record digests for the files you care about while you still have good reason to think they are intact, because a fingerprint taken today is what makes next year’s mismatch mean anything.

Frequently asked questions

Does a checksum mismatch mean the file has a virus?

No. A mismatch means the bytes differ from the ones the publisher measured, and the overwhelmingly likely reason is an incomplete download or the wrong file, not malware. It is still a reason not to open the file until you know which: download it again from the publisher’s own site and compare the two copies with each other, because two damaged downloads almost never produce the same digest.

Can I still install a file if the checksum does not match?

Not sensibly. A mismatch you cannot explain means you are about to run something that is not what the publisher shipped, and the two most likely causes — a truncated download and the wrong build — both produce software that misbehaves in ways that are hard to diagnose later. Re-downloading costs a few minutes.

Why does my download have a different checksum every time?

Usually because something in the path is modifying it: a proxy, a filtering appliance, or an installer stub that a vendor generates per download and therefore genuinely differs each time. If the digest changes between two downloads of the same static file, the published checksum cannot apply to either, and the file should be fetched from the publisher’s own host on a different network before you trust it.

How many characters have to be different for a checksum to count as a mismatch?

One. Hash comparison is exact, and in practice the question never comes up: changing a single bit of the input flips about half the bits of the digest, so nearly every hexadecimal character changes at once. Two digests that differ in only one or two characters almost always mean a transcription error — a mistyped value, a dropped character, a digest copied out of a PDF — because a damaged file changes nearly all 64 at once.

Could two different files have the same SHA-256 by accident?

Collisions have to exist in principle, because there are more possible files than there are 256-bit digests — but no two inputs sharing a SHA-256 digest have ever been found, and chance will not hand you one. That is why a mismatch is never a false alarm from the algorithm’s side: if the digests differ, the bytes differ. What broke for MD5 and SHA-1 is collision resistance, meaning an attacker can now construct such a pair deliberately, which is a separate matter from stumbling into one.

The checksum matches but macOS still says the app is damaged. What now?

Those are two separate checks and the second one is not about your bytes. A checksum proves the file is the file the publisher measured; Gatekeeper is asking whether the code is signed and notarized by a developer Apple recognizes, which an unsigned or unnotarized build will fail even when it is perfectly intact. A matching checksum means the download is fine and the question has moved on to code signing.