Choosing an Algorithm

Which hash algorithm should I use?

Eight names in a list and no indication which row you are supposed to read. Rocket Hash shows all of them for the same input at once, which makes the question easy to avoid and does not make it go away — so here is the decision itself, as four questions asked in the order that settles it fastest. Most people are done after the first one.

You are looking at a dropdown with eight algorithm names in it, or a form field labeled Hash type, or a release checklist that says “publish a checksum” and then stops. The list is not sorted by quality, which is the first unhelpful thing about it. The longest number in it is not the best one, which is the second.

What follows is a decision procedure rather than a league table. Each question either ends the matter or passes you to the next, and they are ordered so that the most common situation is resolved immediately. None of it requires you to know how any of these functions work on the inside.

If you already know which two candidates you are weighing, MD5 against SHA-256 and SHA-256 against SHA-512 go further on those pairs than a page covering all eight can.

If you read nothing else

SHA-256, unless one of the four questions below tells you otherwise. It is the default the rest of the world already settled on, and agreeing with it costs you nothing.

Question one: are you matching a value somebody else produced?

If so, you have no choice to make and you are finished. Two digests are comparable only when they came out of the same function, so the publisher's choice is your choice. An excellent SHA-256 of a file will never equal a perfectly good MD5 of the same file, and the fact that one algorithm is stronger than the other does not make the comparison work.

Nearly always you can tell which one they used by counting characters: 8 is CRC32, 32 is MD5, 40 is SHA-1, 64 is SHA-256 or SHA3-256, 96 is SHA-384, and 128 is SHA-512 or SHA3-512. The only real ambiguity is between a SHA-2 and a SHA-3 of the same width, and published checksums are SHA-2 overwhelmingly more often. Identifying an algorithm from its digest covers the awkward cases, including what to do when the label on the page contradicts the length of the value underneath it.

This question ends the decision for most people who arrive here. “Which algorithm should I use” is usually “which algorithm did they use”, and the answer is printed above the value you were given.

Question two: are you guarding against an accident or against a person?

This is the question that does the real work when the choice is genuinely yours, and the two cases have different answers.

Accidents are truncated downloads, a cable that drops a byte, a disk going soft after four years in a drawer, a copy that finished early. Damage of that kind cannot aim. It has no way to arrange for the broken version to produce the digest the good version produced, so every algorithm on the list catches it — including the ones no longer considered secure, and including CRC32, which is already doing exactly this job inside every ZIP archive and PNG file you own.

People are a different problem, because an attacker picks the bytes. Once somebody is choosing the input on purpose, you need collision resistance: the guarantee that nobody can produce two different files with one digest between them. MD5 lost that guarantee in 2004 and SHA-1 lost it publicly in 2017, and neither has got it back. SHA-256 has it, as do SHA-384, SHA-512 and both SHA-3 sizes.

The test that separates the two cases is not a technical one. Ask whether anybody who has touched the file would benefit from two versions of it being indistinguishable. If the file is yours, or came from a publisher you are trusting for other reasons anyway, the answer is no and a legacy algorithm is merely old. If it came from a stranger and a matching digest is the thing that convinces you to run it, the answer is yes, and the honest state of MD5 is worth five minutes before you rely on it.

Question three: is something already deciding for you?

Frequently, and it is cheaper to find out now than after you have written the code. The usual suspects:

  • A specification or a procurement document. A security profile, an audit requirement, a protocol definition. These are not arguments to win; match the document exactly, output size included.
  • A database column. A field declared 32 characters wide was built for MD5, and a SHA-256 will not fit in it. The fix is to widen the column, never to shorten the digest.
  • An API. Object storage services that want a content digest at upload time often still ask for MD5, because that is what the interface was defined with a decade ago.
  • The other side's tooling. Somebody working from a command line has shasum, which covers SHA-1 and the SHA-2 family and has nothing at all for SHA-3. Choosing SHA3-256 for a manifest another person has to check is choosing a second conversation about installing software.

When something upstream has already decided, the only job left is to write down which algorithm it was, next to the value, so that nobody downstream has to guess.

Question four: is the thing you are hashing a secret?

If the input is a password, an API token, a PIN or anything else an attacker would like to recover, then none of the eight is the right answer — not even the strongest. A general-purpose hash is built to be fast, and when the input comes from a small or guessable set, fast is precisely the flaw: the digest becomes a target to guess against, and a modern graphics card guesses billions of times a second. Password storage wants a deliberately slow, salted function: bcrypt, scrypt or Argon2. Why SHA-256 is the wrong tool for passwords has the full reasoning.

The same trap catches small inputs that are not secret so much as personal. Hashing an email address or a phone number does not anonymize it, because there are few enough possible phone numbers that anybody can hash all of them and look yours up in the list. And if what you actually want is proof that a message came from someone holding a shared key, the primitive for that is HMAC, which uses a hash inside it rather than in place of it.

The eight, and the job each one is for

Assuming the four questions have not already answered you, here is the whole list with the reason each row exists.

The eight algorithms, their digest lengths, and the situation each one is appropriate for
AlgorithmHex charactersReach for it when
SHA-25664Always, unless a row below applies. The default everything else assumes.
SHA-38496A certificate or protocol asks for it, or you want SHA-512's internals without its length-extension weakness.
SHA-512128Something specified 512 bits, or you hash many small inputs in memory and measured a gain.
SHA3-25664You need a construction unrelated to SHA-2, usually because a policy says so.
SHA3-512128The same requirement, at 512 bits.
CRC328Catching accidental corruption inside a pipeline, cheaply, where nobody is attacking you.
MD532A legacy value you did not choose and cannot change.
SHA-140The same, plus reading Git object IDs, which are SHA-1 for historical reasons.

Cases that change the answer

Three situations move the default, and all three are narrower than the internet suggests.

You are hashing a secret together with a message. SHA-256 and SHA-512 both expose their internal state in their output, so given one digest and the length of the input, somebody can append bytes and compute the digest of the longer version without ever knowing what the original said. That is length extension. It is irrelevant to file checksums and it matters a great deal the moment a digest is acting as authentication. SHA-384 is immune because its output is truncated, and SHA-3 is immune because it is not built the same way.

You want your eggs in two baskets. Every member of SHA-2 shares a single internal design, so picking a wider variant of it is not diversification. SHA-3 comes from an entirely different mechanism, which is the real reason it was standardized — insurance, not an upgrade. What the sponge construction does differently covers what that buys and what it costs you in compatibility.

You need a short identifier rather than a checksum. Truncating a digest on purpose is legitimate and normal; Git does it every time it shows you seven characters. Do the arithmetic before choosing a length, though, because collisions turn up at roughly the square root of the space: 8 hex characters start colliding somewhere around 77,000 items, and 16 hex characters hold out to about 5 billion. Take the leading characters of a SHA-256 rather than reaching for a shorter algorithm, so you can lengthen the identifier later without rehashing anything.

Things that look like they should decide it

Four considerations that feel decisive and are not.

  • Digest length. Collision resistance is half the output size, so SHA-256 gives you 128 bits of it — around 340 undecillion attempts, a figure no present or planned machine reaches. Doubling something already unreachable changes nothing about your download, which is why the length of a digest is useful for telling algorithms apart and useless for ranking them.
  • Which one is newest. SHA-3 was published in 2015 and SHA-256 in 2001, and the older of the two is not the weaker of the two. Neither has a known weakness that affects anybody.
  • Speed. Hashing a file on a Mac is limited by how fast the disk hands over bytes, not by the arithmetic. The algorithms do differ, by a factor the storage hides completely.
  • How it was described to you. “Military grade” and “bank level” are not properties of a hash function. Collision resistance, preimage resistance and output size are.

There is a useful consequence of that third point: you rarely have to commit. The Files tool in Rocket Hash produces all eight digests from a single read of a file, and does the same for every file under a folder, because one streaming pass can feed every algorithm at once. When you do not know which value somebody will eventually ask you for, generate the lot and keep them. For a string rather than a file — a token, a config line, a value in a ticket — the free calculators here will settle it in a browser tab instead, one algorithm to a page: SHA-256, MD5 and CRC32 among them, each computed on your own machine as you type.

Troubleshooting

The instructions just say “SHA-2”

SHA-2 is a family, not a function: SHA-224, SHA-256, SHA-384 and SHA-512 are all SHA-2. Written without a number it means SHA-256 nearly every time. If there is a sample value anywhere in the documentation, count its characters and believe the count over the prose.

Two systems support different algorithms

Take the intersection and pick the strongest member of it, which in practice means SHA-256 almost always, because it is the one thing everything has. If the intersection is empty, or holds nothing but MD5, publish both values instead of arguing: one pass over the files produces them together, so the second digest costs nothing, and the weaker one can be retired later without breaking anybody.

I already used the wrong one everywhere

A digest cannot be converted from one algorithm to another, so the only route is back to the original files — which you need to have anyway. Produce both for a transition period, label which is which, and stop publishing the old one on a date you announce in advance. Existing MD5 manifests for files you produced yourself remain perfectly good at detecting corruption meanwhile, so there is no emergency here.

The form will not accept my digest

It was built for a shorter one, and a 32-character limit is a field designed around MD5. Never trim a digest to fit a box: a truncated SHA-256 is not a shorter SHA-256, it is a value nothing else will reproduce. Widen the field, or produce the algorithm the field was built for with your eyes open about what that means.

Nobody will tell me which to use

Choose SHA-256 and move on. It is the answer you would have arrived at by asking, and the one the next person will assume. Then record the choice beside the value, in whatever file or page carries it — a digest with no algorithm written next to it is the reason somebody will be reading this page in two years.

Frequently asked questions

What is the most secure hash algorithm?

As a bare answer, SHA3-512 and SHA-512 have the largest margins. The question rarely decides anything, though, because no attack reaches SHA-256 either and every algorithm in current use is already far beyond what any machine can search. The real trade-off is between what is secure and what the other side can actually verify, and SHA-256 wins that comparison more or less always.

Is SHA-256 enough, or should I use something stronger?

For checksums, manifests, downloads and signatures it is enough, and nothing within reach of any attacker threatens it. Go past it only when a specification names something else, when you want a construction unrelated to SHA-2, or when a digest will sit next to a secret — and in that last case the answer is SHA-384 or a SHA-3 size rather than a longer SHA-2.

Is CRC32 good enough for checking files?

Good enough for accidental damage and for nothing else. Eight hexadecimal characters are ample to notice that a transfer truncated or a disk went bad, which is why every ZIP archive and PNG file already carries one internally. CRC32 is also straightforward to forge deliberately, so it proves nothing at all about a file whose contents somebody else chose.

Which hash algorithm is fastest?

CRC32, by a wide margin, because it is not a cryptographic hash at all. Among the real hashes the ranking depends on your processor more than on the algorithm, and it seldom matters: hashing a file is limited by disk throughput, and a single pass can feed every algorithm at once, so asking for eight digests does not cost eight times the wait.

What hash algorithm should I use for file integrity?

SHA-256. It is what publishers use, what verification tools assume, and what somebody else will be able to check with whatever tooling they already have. Reserve CRC32 for accidental damage inside a system you control, and avoid MD5 and SHA-1 wherever a stranger chose the bytes you are checking.

Do I need to switch from SHA-2 to SHA-3?

No. SHA-3 was standardized as a structurally different alternative, not as the replacement for a broken predecessor, and SHA-2 remains entirely sound. Move to SHA-3 when a policy requires it, or when you deliberately want a construction that would survive a break in SHA-2 — and expect to meet tools that cannot verify it.

Which hash should I use for storing passwords?

None of these. Password storage needs a slow, salted password-hashing function — Argon2, scrypt or bcrypt — precisely because general-purpose hashes are fast enough to guess against at scale. A salted SHA-256 is better than a bare SHA-256 and still the wrong tool for the job.