What is a checksum, and when do you actually need one?
Underneath every serious download there is a row of hexadecimal that almost nobody clicks, and a sentence saying you should check it. The thirty seconds it takes to drop the file into Rocket Hash and read the verdict is the whole difference between assuming a file arrived intact and knowing it. What follows is why that row is there, and the honest limits of what it can tell you.
Ubuntu publishes one under every ISO. GitHub release pages carry them next to the binaries. Your colleague pasted one into a message along with the file, because someone once told them to. Sixty-four characters of hexadecimal, no explanation, and no indication of what you are supposed to do with it.
A checksum is a short value computed from every byte of a file, in such a way that two different files essentially never produce the same one. It is a fingerprint: far too small to reconstruct the file from, but specific enough that matching fingerprints mean matching files. That is the entire idea, and most of the confusion around it comes from expecting it to carry more meaning than that.
This page is about what the fingerprint genuinely establishes, what people routinely believe it establishes and it does not, and the handful of moments in an ordinary week when computing one is worth doing. If you already have a file and a published value and simply want the verdict, checking a file's checksum on a Mac is the practical route and takes about fifteen seconds.
A matching checksum proves the bytes you have are the bytes somebody measured; it says nothing whatsoever about who that somebody was, or whether the bytes were good to begin with.
What the number actually is
A hash function is a piece of arithmetic that takes input of any length and returns output of one fixed length. SHA-256 returns 256 bits — 32 bytes, written as 64 hexadecimal characters — whether you hand it an empty file or a terabyte of video. The value is not a sample, a summary or a compression of the file; every byte of input affects every character of output.
Two properties make the result useful as a fingerprint rather than a curiosity. It is deterministic: the same bytes always give the same digest, on any machine, in any decade, which is why a value published in 2014 still tests a file today. And it avalanches: change one bit anywhere and about half the bits of the output flip, which in hexadecimal leaves almost none of the 64 characters as they were. There is no gradient, no partial credit, and no sense in which two digests can be nearly equal.
You can verify this yourself in a few seconds, because the standard fixes the answers. The SHA-256 of an empty input is e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 and the SHA-256 of the three letters abc is ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad. Every correct implementation in the world produces exactly those, because NIST publishes them as test vectors. Type the three letters into the free SHA-256 generator here, or into anything else that claims to do this, and you have a thirty-second honesty test for the tool — the one piece of software in the chain you cannot check by comparison.
Checksum, hash or digest?
In casual use these are the same word, and on a download page “checksum” nearly always means a SHA-256 digest. The distinction is still worth having, because it explains why some of these values are trustworthy and others are explicitly not.
- A checksum, strictly, is any short value used to detect accidental corruption — including weak ones such as a parity bit or a CRC. Its job is catching damage, and it was never designed to resist anybody.
- A cryptographic hash is built to resist a person who is deliberately trying to defeat it. SHA-256, SHA-512 and the SHA-3 family are in this category; MD5 and SHA-1 were, and no longer fully are.
- A digest is simply the output value, whichever function produced it. “The digest” and “the hash” are used interchangeably for the 64 characters themselves.
CRC32 is the one you will meet that is firmly in the first category: eight hexadecimal characters, built into ZIP, PNG and Ethernet, excellent at spotting a bad cable, and straightforward to forge on purpose. What CRC32 is actually used for explains why something that weak is still everywhere, and why that is fine.
What a checksum proves
Exactly one thing: that a set of bytes is unchanged since the digest was computed. Everything else you might want from it is built on top of that, and the strength of the claim depends on which of two different guarantees you are relying on — a distinction that decides when an old algorithm is still fine and when it is useless.
Preimage resistance means that given a digest, nobody can work backwards to a file that produces it. This is why publishing a digest is safe: it leaks nothing about the contents. It is also why “SHA-256 decryption” sites are nonsense — the function discards information, so there is nothing to reverse, and those sites are simply looking your value up in a table of digests of common inputs.
Collision resistance means that nobody can construct two files with the same digest, even when they get to choose both files freely. This is the guarantee that lets a digest stand in for identity, and it is the one that breaks first. MD5 collisions have been constructible since 2004 and are now trivial; SHA-1 collisions became a public fact in 2017, when two different PDF files with an identical SHA-1 were published under the name SHAttered.
Both MD5 and SHA-1 are still preimage-resistant in practice, and a cosmic ray cannot choose its corruption — so for catching a damaged copy they remain perfectly serviceable. The moment a human gets to author the file you are checking, collision resistance is the property you need, and those two no longer have it.
That is the whole of the MD5 argument, compressed: it depends on whether your threat is an accident or a person. Whether MD5 is still safe takes the same question apart at length, with the cases where using it is a reasonable engineering decision rather than negligence.
What it does not prove
It does not prove where the file came from. Anyone who modifies a file can compute a perfectly valid checksum of the modified version and publish that one instead. A digest on the same page as the download, served by the same machine over the same connection, only establishes that the copy survived the journey — if the page was lying, the checksum is lying in exactly the same voice. Turning integrity into origin takes a signature, which binds a digest to a key; detecting tampering covers what that buys you and what it does not.
It does not prove the contents are harmless. A malicious installer has a checksum too, and it matches, because the file is intact. Integrity and safety are unrelated questions.
And it does not tell you where two files differ, or by how much. A digest is a single verdict with no detail behind it, which is precisely why a one-byte change and a completely different file look identical from the outside: both simply say “no match”. When you need to know what changed, you want a diff; when you need to know whether anything changed, you want this.
The four moments it earns its keep
Before you run something you downloaded
An installer, a firmware image, an ISO you are about to write to a USB stick. This is the classic case and the one the publishers are anticipating when they list a value. A mismatch here is almost always an incomplete download rather than an attack, but you would rather find out before the thing is running with your password. Verifying a download on a Mac has the route, including where the real checksum tends to live.
After something crossed a cable or a network
A progress bar reaching the end is a claim about a transfer, not a claim about the result. USB cables are flaky, SMB shares drop connections, cheap flash media lies about what it wrote, and none of those failures come with an error message. Hash the file at both ends and the question is settled in one comparison — and if the two disagree, you have learned something genuinely useful before you deleted the original.
When something has been sitting in storage
Files that nobody opens can still go bad. Drives develop unreadable sectors, controllers fail in small ways, and an archive you have not touched for four years is the place you are least likely to notice. Recording digests while you know the files are good turns that into a check you can run on demand, which is the difference between owning a backup and hoping for one. Verifying a backup is about exactly this.
When you need to know two things are the same thing
Names lie, timestamps lie, and file sizes collide constantly — two different photographs can easily be the same number of bytes. A digest is the only field that settles the question, which is why it is the basis of every sane duplicate finder and of content-addressed storage in general. Finding duplicates by hash uses that property directly: hash a folder, sort the list, and the duplicates group themselves.
When you do not need one
Plenty of the time, and pretending otherwise would be dishonest. Anything installed from the App Store has already been signed and checked before it runs, and so has a notarized app you downloaded from a developer's own site — macOS did that verification without asking you, and it is a stronger check than a checksum because it involves a signature. A file you just created on the machine you are sitting at does not need fingerprinting either.
The other case is subtler: if the only available checksum comes from the same place as the file and you have no independent way to confirm it, computing it tells you the download finished cleanly and nothing more. That is a real and useful thing to learn — most mismatches in the world are truncated downloads — but it is worth being clear with yourself about which question you just answered.
Troubleshooting
The page lists a checksum but not which algorithm
Count the characters. Eight is CRC32, 32 is MD5, 40 is SHA-1, 64 is SHA-256 or SHA3-256, 96 is SHA-384, and 128 is SHA-512 or SHA3-512. Only two lengths are ambiguous, and both resolve the same way: an unlabeled 64-character value is SHA-256 and an unlabeled 128-character one is SHA-512, because a publisher using the SHA-3 family almost always says so. If it still does not match, compute both and see which one lands.
I have two different checksums for the same file
If they are different lengths, they are different algorithms and both are probably correct — publishers often list an MD5 alongside a SHA-256 so that older tooling can still check something. Match the one whose length corresponds to the function you are computing, and prefer the SHA-256 if you have the choice.
Mine does not match and I do not know why
Work through the mundane causes before the alarming one, because the mundane ones account for nearly all of it: an unfinished download, the wrong file in a folder of similar files, a checksum for a different build or a different architecture, or a value that lost a few characters on its way through a copy and paste. What to do when a checksum does not match orders them by likelihood.
There is a signature file next to the checksum
Then the publisher has done the job properly, and the two files are answering different questions. The checksum file lists digests for the release; the detached signature beside it — often ending in .gpg or .asc — proves the checksum file itself was written by whoever holds the key. Checking the first without the second still leaves you trusting the page that served them both.
Frequently asked questions
What is a checksum in simple terms?
A short string of characters calculated from every byte of a file, which acts as a fingerprint for it. If you compute it on your copy and get the same string the publisher listed, your copy holds exactly the same bytes as theirs. If you get a different string, something about your copy differs — even if the difference is a single bit.
What is the difference between a checksum and a hash?
In everyday use, none — both refer to the value and to the function that produced it. Strictly, a checksum is any short value for detecting accidental damage, including weak ones like CRC32, while a cryptographic hash such as SHA-256 is designed to resist somebody deliberately attacking it. On a download page, “checksum” almost always means a SHA-256 digest.
Can two different files have the same checksum?
In principle yes, because there are infinitely many possible files and only a finite number of digests. In practice it depends on the algorithm: no two different inputs with the same SHA-256 have ever been found, while for MD5 and SHA-1 pairs can now be constructed deliberately. Accidental collisions with a 256-bit digest are not something you will encounter.
Does a checksum tell me who made a file?
No, and this is the most common misunderstanding about them. Anyone who alters a file can publish a matching checksum for the altered version, so the digest only proves the bytes are unchanged since someone measured them. Tying a file to an identity requires a digital signature, which is a separate mechanism built on top of the digest.
How long is a checksum supposed to be?
It depends entirely on the algorithm, and the length is how you identify which one was used: 8 hexadecimal characters for CRC32, 32 for MD5, 40 for SHA-1, 64 for SHA-256, 96 for SHA-384 and 128 for SHA-512. Digest lengths in full has the whole table.
Do I need to check the checksum of every download?
No. Apps from the App Store, and notarized apps from a developer's site, are already verified by macOS through a signature, which is a stronger check. Checksums are worth your time for operating system images, firmware, anything you are about to run with administrator rights, and anything where the publisher has gone to the trouble of listing one.
Is a checksum the same thing as encryption?
No. Encryption is reversible by design — that is the point of holding a key. Hashing throws information away permanently and has no key, so there is nothing to decrypt and no way back to the original file. A digest is a measurement of data, not a protected version of it.