How to check that a hashing tool is telling the truth
You are about to keep or delete a download on the strength of sixty-four characters a program handed you, and a wrong digest looks exactly like a right one. Hash functions happen to be the easiest software on earth to audit: one correct answer per input, published decades ago, identical on every machine ever built. Here is how to hold Rocket Hash — or anything else that claims to do this — to it in under a minute.
Nothing about a digest tells you whether it is correct. It is sixty-four characters of hexadecimal whether the implementation is flawless or off by one bit in the padding, and you would make the same decision either way: keep the installer, delete the installer, ship the release, pull the release. The tool doing the counting is the one thing in the chain you cannot check by comparing it against something else — unless you know what the answer is supposed to be in advance.
Which, for hash functions, you do. They are deterministic to a degree almost no other software is: no randomness, no salt, no configuration, no locale, no version drift. The SHA-256 of the three letters abc was fixed when the standard was published and will be the same value in a century. That makes a single published input and output into a complete claim about an implementation, and it means you never have to take a vendor's word for anything.
So: two inputs whose answers were settled decades ago, three more algorithms checked on the same screen, then the one cross-check no published answer can give you, and an honest account of what the whole exercise does not prove. If what you actually want is to compare a file against a checksum somebody published, that is a different job and checking a file's checksum is the page for it.
SHA-256 of nothing at all is e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855. SHA-256 of abc is ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad. Anything that disagrees with either is not computing SHA-256, whatever the label on it says.
Why a known-answer test works
A test vector is not a convention somebody agreed to. It is the algorithm's own definition, applied by hand to a short input. FIPS 180-4 defines SHA-1 and the SHA-2 family and FIPS 202 defines SHA-3; NIST publishes the worked examples, with the intermediate values, separately from the standards themselves. Its validation program goes further again — thousands of inputs covering every message length that behaves differently.
Because the function is deterministic, those sets are not samples. There is no configuration under which a correct implementation produces a different answer, so one mismatch is not a discrepancy to be investigated. It is a bug — and a tool that gets a published vector wrong was handing you worthless digests all along.
Every algorithm in Rocket Hash is checked against those vectors and cross-checked against independent reference implementations before it ships — which is the right thing for a developer to claim and exactly the sort of claim you should not have to believe. The following takes about as long as reading this paragraph.
Three known answers, then one cross-check
-
Hash the three letters abc
Type
abcinto the Text tool — free, nothing to buy — and check that the byte count underneath reads 3 bytes before you look at anything else. The SHA-256 row must readba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad. Long digests are shortened in the middle on screen, so click the row to copy the whole thing and compare it properly rather than trusting the first eight characters. That one input exercises the text path, the UTF-8 encoding and the padding of a message shorter than a single block. -
Hash a file with nothing in it
Now drop a zero-byte file into Files: the answer must be
e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855. This is a better test than it looks, because it goes through the file-reading code rather than the text field — different code entirely — and because zero bytes is the edge case careless implementations get wrong, returning an error, an empty string, or the digest of a badly constructed padding block. Any zero-byte file will do, and most Macs already have a few lying about: a stalled download, a placeholder that arrived with a project, a log that was rotated and never written to again. Sort a Finder window by size and they collect at one end of the list. -
Check a second algorithm on the same screen
Implementations are written separately even inside one program, so SHA-256 coming out right is evidence about SHA-256 and about nothing else. The Text tool settles several at once for free, because every algorithm's answer for
abcis on screen together: MD5 must read900150983cd24fb0d6963f7d28e17f72, SHA-1 must reada9993e364706816aba3e25717850c26c9cd0d89d, and CRC32, under the CHECKSUM heading, must read352441c2. That is four published answers from one three-letter input, and hashing text as you type explains the groups the rows are sorted into. -
Cross-check the file you actually care about
Every test so far fits inside a single 64-byte block, which leaves the code that loops block after block, counts the message length and joins one read of the disk to the next completely untouched. NIST publishes a vector for that as well — one million letter
a, whose SHA-256 iscdc76e5c9914fb9281a1c7e284d73e67f1809a48a497200e046d39ccc7112cd0— but known answers only ever cover inputs somebody thought of in advance, and your 4 GB disk image is not one of them. What covers that is a second opinion on the same file: the checksum a publisher prints beside the download, the line for that file inside a manifest the sender shipped, or a digest a colleague reads off whatever they already use. Hash the file in Files and set the two side by side, or paste their value into Verify in Against a Checksum mode and let it work the algorithm out of the string. Two separately written implementations agreeing on an arbitrary 4 GB input says more than any published vector does.
The four together take a couple of minutes, and what you have at the end is not a developer's promise but four answers you checked yourself. Copy All earns its keep here: it lifts every algorithm at once as one labeled block, which is easier to set beside a page of published vectors than eight separate copies are.
The digest length is a free sanity check
Before you check a single vector you can catch an entire class of mislabeling by counting characters, because each algorithm has exactly one output size and there is no negotiating it.
| Algorithm | Hex characters | Bits |
|---|---|---|
| CRC32 | 8 | 32 |
| MD5 | 32 | 128 |
| SHA-1 | 40 | 160 |
| SHA-256, SHA3-256 | 64 | 256 |
| SHA-384 | 96 | 384 |
| SHA-512, SHA3-512 | 128 | 512 |
A tool labeling something SHA-512 and handing you 64 characters is giving you either SHA-256 under the wrong name or a truncated value, and both are disqualifying. The same arithmetic works in the other direction when you have a checksum and no idea what made it, which is what telling which algorithm a checksum uses is for.
CRC32 deserves its own warning, because it is the one where leading zeros get eaten. Eight characters, always: the CRC32 of abc is 352441c2, of an empty input it is 00000000, and the standard's own check value — the nine characters 123456789 — is cbf43926. A tool that prints 0 for the empty input is formatting a number rather than a checksum, and its output will not line up with anything else.
What a known-answer test does not prove
Three limits, and the first one accounts for almost every real-world failure.
- It tests the arithmetic, not the plumbing. A tool can compute SHA-256 perfectly and still hand you the digest of the wrong bytes — a partially downloaded copy, last month's version sitting in the same folder, a
(1)duplicate, the file a symlink points at rather than the one you meant. This is overwhelmingly the common case, and no vector catches it. The defense is to read the name, the containing folder and the size on screen before you read the digest. - It cannot out-think a dishonest tool. Software that wanted to deceive you could recognize the famous vectors, answer them correctly, and lie about your file. Nothing you run on known inputs closes that hole. What closes it is two unrelated implementations agreeing on your input, which is why the cross-check above is not a nicety for anything that matters.
- It says nothing about whether the algorithm is worth using. A flawlessly correct MD5 implementation is still MD5, and MD5 stopped resisting collisions in 2004 — building a colliding pair is now seconds of work. SHA-1 went the same way in 2017, via the SHAttered result and its pair of PDFs sharing one digest. Correct arithmetic on a broken function is a correct answer to the wrong question.
That last point is worth one more sentence, because the two properties get conflated constantly. What MD5 and SHA-1 have lost is collision resistance: an attacker who controls both files can construct a pair that share a digest. What they have not lost is preimage resistance — given a digest and nothing else, nobody can produce a file that hashes to it — or second preimage resistance, which is the one that matters here: given a file that already exists, nobody can construct a different file carrying the same digest. That is why both remain perfectly serviceable for spotting accidental corruption and useless as evidence that a file is the one somebody intended you to have. Whether MD5 is still safe and whether SHA-1 is still safe each take one of them in full.
If the published checksum came from the same page, over the same connection, as the download, then a perfect match proves the transfer was clean and nothing else. Checking a file has not been tampered with is where that gap gets closed.
When two tools disagree about your file
Start from the assumption that both are right and that you fed them different bytes, because that is usually exactly what happened. A trailing newline is the single most common cause: the SHA-256 of abc followed by a newline begins edeaaff3…, which looks nothing like the vector you were expecting, and most ways of getting three letters into a file add that fourth byte without telling you. The byte count under the text field settles it — three characters typed in should read 3 bytes.
After that: the file was still downloading when one of the two read it, the two were pointed at different copies of it, or one hashed the text of a filename rather than the file. Read the name, the folder and the size on the row before you read the digest, give both the same file in the same minute, and the disagreement usually evaporates — why two tools give different hashes works through the rest in order of frequency.
Troubleshooting
It gets abc right but not my file
Then the arithmetic is fine and the input is not, which narrows things usefully. Check the size the tool reports against the size the Finder reports, and the folder it names against the folder you meant. A digest computed correctly over a half-finished download is a correct digest of a file you do not want.
My digest is in capital letters
Case is presentation, not content: BA7816BF and ba7816bf are the same value, because hexadecimal digits do not care. Some tools and some publishers use uppercase. Compare without regard to case, and never conclude a mismatch from it.
The Text tool and the file do not agree
Type the contents of a short text file into Text, hash the file itself in Files, and the two digests usually differ. That is correct rather than broken. Almost every text file ends with a newline you cannot see, files written on Windows end each line with two bytes instead of one, and some editors add a byte-order mark at the front. If you want the file's digest, hash the file; if you want the digest of the characters, type them. They are two different inputs that merely look alike.
The tool reports a different algorithm than I asked for
Count the characters before you investigate anything else. Thirty-two is MD5, forty is SHA-1, sixty-four is SHA-256 or SHA3-256, ninety-six is SHA-384, a hundred and twenty-eight is SHA-512 or SHA3-512. Note that SHA-256 and SHA3-256 are the same length and completely different functions, so length narrows it without settling it: check against the vector.
Should I test every algorithm, or just SHA-256?
Test the one you are going to rely on. Four of them have short vectors you can check on one screen, which covers most needs, but the SHA-2, SHA-3, CHECKSUM and LEGACY groups are four different pieces of arithmetic, and a correct SHA-256 says nothing whatever about SHA3-512. If a release check rests on SHA-512, spend the extra ten seconds looking up SHA-512's own published vector.
Frequently asked questions
How do I know a hash calculator is accurate?
Hash it something whose answer is already published and compare. The SHA-256 of abc is ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad and the SHA-256 of an empty file is e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855. Because the function is deterministic, any tool that disagrees with either is broken — there is no configuration under which a correct implementation gives a different answer.
What is a hash test vector?
A published input paired with the output a correct implementation must produce. They come from the standards that define the algorithms — FIPS 180-4 for SHA-1 and SHA-2, FIPS 202 for SHA-3 — and NIST's validation program publishes much larger sets covering every message length that behaves differently. Checking against them is called a known-answer test, and it is the only way to audit a hash tool without trusting anyone.
What is the SHA-256 of an empty file?
e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855, every time, on every machine. A hash function produces a fixed-length output from any input including no input at all, so zero bytes has a digest just like anything else. It is a useful test because a tool that returns an error or an empty string for it has a bug in the code path that handles lengths.
Can two hash tools give different results for the same file?
Not if both are correct and both read the same bytes — and the second condition is the one that fails. A trailing newline, a download that was still in progress, or two copies of a file in the same folder will do it. Why two tools give different hashes lists the causes in order of how often each turns out to be the culprit.
Does passing NIST test vectors mean a hash tool is secure?
It means the arithmetic is right, which is a different claim. The tool can still hash the wrong file, and the algorithm itself can be unfit regardless — a perfectly correct MD5 implementation is still MD5, which stopped resisting collisions in 2004. Correct arithmetic on a broken function gives you a correct answer to the wrong question.
Is a local app more trustworthy than an online hash generator?
For the arithmetic, neither has an advantage: test both against the same vectors and believe the results. The difference is what happens to the file. A page can upload it, and you have to inspect network traffic to know that it did not, whereas an app with no network capability at all has nothing to inspect. A page that only ever hashes text you typed sidesteps the question entirely, because there is no file for it to send. Whether it is safe to hash files on a website covers how to tell.