MD5 vs SHA-256: what actually changed
A 32-character checksum on a download page that has not been redesigned since 2011, or a column somebody sized at 32 characters in a database you inherited. Rocket Hash will hand you either value for the same file from a single read of it — what it cannot tell you is whether the shorter one is still doing its job, and that turns out to depend less on the algorithm than on who chose the bytes.
Somebody has told you to move off MD5 and has not said what breaks if you do not. The swap fixes one failure, completely, and leaves everything else exactly where it was: SHA-256 does not make a digest harder to reverse, does not make your hashed email addresses private, and does not retroactively invalidate the MD5 manifest sitting in your archive folder.
So this page is an audit of the swap: what you gain, what you do not, which tasks genuinely need the change, and what to do about the MD5 values you are already carrying. For the history of how MD5 got into this state, and whether you can keep using it, the honest answer on MD5 is the page to read next; if what you want is the shortest route to a decision for a new project, four questions settle it.
Use SHA-256 anywhere a stranger chose the bytes you are checking; MD5 is still perfectly good at noticing that a file you already trust got damaged on the way to you.
The same three letters, both ways
Start with something you can reproduce. The letters abc hash to 900150983cd24fb0d6963f7d28e17f72 under MD5 and to ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad under SHA-256. The first is 32 hexadecimal characters and has been published since 1992, when Ron Rivest specified MD5; the second is 64 and has been published since 2001, when NIST standardized the SHA-2 family. Neither value resembles the other, and neither resembles the three letters that produced it; the only difference you can see between them is length.
Everything else the two have in common is the part people assume is at stake and is not. Both are deterministic: the same bytes always produce the same digest. Both avalanche, so a single flipped bit changes roughly half the output. Both read the file once and produce a fixed-size value regardless of input size. Both are free, fast, and available everywhere. The length difference is the visible difference and the least important one.
What the swap buys you
One thing: the assurance that nobody can construct two different inputs with the same digest. That is collision resistance, and MD5 has not had it since 2004. Producing a pair of files that differ however you like and share an MD5 now takes ordinary hardware an unremarkable amount of time, and such pairs have been used in real attacks on real systems rather than only in papers.
SHA-256 has that assurance intact with no credible challenge to it. Finding a collision by brute force means working through a 39-digit number of attempts, and no shortcut to that is known. SHA-1, for reference, sits between the two: collisions were demonstrated publicly in 2017 and it should be treated as retired, which the SHA-1 page covers in full.
Notice what that property is about. It concerns somebody who controls both files and wants them to look alike. That is the whole of what changed, and it is the reason the answer to “is MD5 fine?” is a question in return: fine for what?
What the swap does not buy you
Three things, and each one is a belief worth dropping.
It does not make the digest less reversible. MD5 is still preimage resistant: given a digest and nothing else, no practical method of producing an input for it beats guessing candidates one at a time, and the best published attack shaves about five bits off a 128-bit search — a saving nobody can run. The sites offering to “decrypt” an MD5 are looking the value up in a table of strings people commonly hash, and SHA-256 falls to precisely the same trick, because the weakness there is the small set of likely inputs rather than the function. Hash a six-digit PIN with SHA-256 and it is recovered in under a second; that is why passwords need a slow, salted function and not a better general-purpose hash.
It does not make your existing MD5 records worthless. The other property MD5 has kept is second-preimage resistance: given a file that already exists, with a digest that was recorded before anybody went looking, building a different file that matches that digest is not practical. An MD5 manifest you generated last year over files you already held still detects modification as well as it did then. What it cannot do is vouch for a file that arrived afterward from someone with an interest in the outcome.
It does not turn a checksum into a signature. Neither algorithm tells you who made the file. If an attacker replaces both the download and the digest published next to it, a SHA-256 match proves only that you received the file the page is currently serving. That gap closes with a signature, not with a longer hash.
Even a perfect MD5 would be too short
Set the cryptanalysis aside for a moment, because there is a second, independent reason MD5 was always going to be retired. A 128-bit digest gives you only 64 bits of collision resistance, because the birthday bound halves it: roughly 18 quintillion random attempts are enough to expect a collision with no cleverness at all.
That is not a theoretical figure. The SHA-1 collision announced in 2017 consumed work on that same order of magnitude, and it was carried out, paid for and published — which tells you a 64-bit barrier is inside the reach of a well-funded team rather than beyond it. SHA-256 gives you 128 bits of the same resistance, and that gap between 64 and 128 is not a doubling in difficulty but a squaring of it.
“Nobody would bother attacking us” is a judgment about attackers rather than about arithmetic, and it ages badly. “We only hash internal files” is sound right up to the first file that came in from outside.
Where it matters and where it does not
The honest version of this comparison is task by task, not algorithm by algorithm.
| What you are doing | MD5 | Why |
|---|---|---|
| Confirming a download from a publisher you already trust arrived intact | Fine | Truncation and corruption cannot aim at a digest |
| Checking a file a stranger sent you against a digest the same stranger sent you | No | They control both halves of the comparison |
| Spotting bit rot in your own archive against a manifest you made | Fine | The digest predates any interest in forging it |
| Deduplicating files you produced yourself | Fine | Nothing in the set was built to collide |
| Deduplicating or content-addressing files other people upload | No | A user can upload a known colliding pair on purpose |
| Proving a contract or a log has not been altered | No | Whoever wrote it could have prepared two versions |
| Storing a password | Neither | Both are far too fast; this needs Argon2, scrypt or bcrypt |
The deduplication rows deserve a note, because they look identical and are not. Colliding MD5 pairs are published and easy to obtain, so a store that indexes other people's uploads by MD5 can be made to treat two genuinely different files as one, quietly, on request. Over your own photo library the same index is completely safe. If you are finding duplicates by digest on a disk of your own, the algorithm is not what decides the result.
Three reasons people still reach for MD5
Only one of them survives contact with the details.
- “It is faster.” Less than you would think, and often not at all. Processors carry instructions that implement SHA-256's round function in hardware and have nothing equivalent for MD5, which closes the gap and on many machines reverses it. Where MD5 can still win is millions of short strings in memory, and that is also where the strings most often came from a user — the one case MD5 is no longer fit for.
- “It is shorter.” True, and legitimate when the constraint is a field width, a filename or a log line rather than a security argument. Just be clear which one you are making — shortening a value for convenience is fine, and shortening the claim you make about it is not.
- “It is what the other end uses.” The good reason, and it ends the discussion, because digests from different functions cannot be compared at all. If the published value is 32 characters you produce 32 characters. Counting the characters is how you settle which one you are dealing with.
What to do with the MD5s you already have
Nothing urgent, and nothing that involves converting anything — a digest holds no trace of the input beyond its own bits, so there is no path from an MD5 to a SHA-256 except back through the original file. That makes the migration mechanical rather than clever.
- Widen the field first. A 32-character column, filename convention or API parameter has to take 64 before you write a single SHA-256 into it, and truncating the new value to fit is worse than not migrating.
- Produce both for a transition period. One read of a file in Rocket Hash feeds every algorithm at once, so carrying both values costs storage rather than time, and nobody's existing tooling breaks while you switch.
- Label every value with its algorithm. Most of the confusion in this subject is a bare string of hex with no word next to it.
- Keep the old MD5 records. Until the SHA-256 pass is finished they remain a working corruption check for files you already hold — just not evidence of where those files came from.
If you only want a quick look at one of the two values, the MD5 generator on this site runs in a browser tab with nothing uploaded. For the pair, producing an MD5 on a Mac is the page: expand a file's row and the MD5 under LEGACY and the SHA-256 under SHA-2 sit one above the other, both out of the same single read of the file, which is what you want when the point is to record them together.
Troubleshooting
The MD5 matches but the SHA-256 does not
Look for the boring explanation, because it is almost certainly there: two published values taken from different builds or different versions of the page, a checksum copied with a character missing, or two different files on your disk. The exciting explanation — a file deliberately built to keep one digest while changing the other — would require a practical second-preimage attack on MD5, which nobody has. Re-download and re-check both values from a single source.
I have an MD5 and I need a SHA-256
Then you need the file, and no tool or service can help you without it. This is the one genuinely irreversible step in the whole subject: hashing discards the input. If the file is gone and only the MD5 survives, what you have is a record that a particular file once existed, not something you can upgrade.
The only checksum published is an MD5
Check it anyway — it still catches a truncated or corrupted download, which is the failure you are overwhelmingly most likely to hit. Then decide separately whether you trust the source, because that question was never answered by the checksum's algorithm. Downloading the same file from a second mirror and comparing the two copies directly is a cheap way to raise your confidence without any published value at all.
One value is uppercase and the other is lowercase
They are the same digest. Hexadecimal is case-insensitive, so 900150983CD2 and 900150983cd2 describe identical bytes; only the formatting convention differs between tools. Compare the two values case-insensitively, or lowercase both before comparing, and never treat a case difference as a mismatch.
A scanner flagged MD5 in a file integrity check
It is flagging the algorithm rather than analyzing your use of it, which is reasonable of it — a tool cannot tell from the outside whether an adversary picks your inputs. Write down which of the rows in the table above your case matches, fix the ones that need fixing, and document the rest. “Corruption detection over files we generate” is an answer an auditor can accept; “it has always worked” is not.
Frequently asked questions
What is the difference between MD5 and SHA-256?
MD5 produces 32 hexadecimal characters and SHA-256 produces 64, but the difference that matters is that MD5 is no longer collision resistant: two different files with the same MD5 can be built deliberately, and have been. SHA-256 still has that property intact. For detecting accidental corruption both work equally well; for proving a file is the one somebody promised you, only SHA-256 does.
Is SHA-256 more secure than MD5?
Yes, and specifically in one way: nobody can construct two inputs that share a SHA-256 digest, whereas doing that with MD5 is routine. SHA-256 is not meaningfully better at resisting attempts to recover the original input, because MD5 was never broken in that direction. If you are choosing for anything new, choose SHA-256.
Can I convert an MD5 hash to SHA-256?
No. A digest keeps nothing of the input except its own bits, so there is no computation that turns one into the other — the only route is to hash the original file again with the new algorithm. If you no longer have the file, the MD5 cannot be upgraded by any tool or service.
Can an MD5 and a SHA-256 of the same file be compared?
Never. Two digests are only comparable when the same function produced them, so a correct MD5 and a correct SHA-256 of one identical file will look nothing alike, and that is not a sign of a problem. Count the characters of the value you were given — 32 means produce an MD5, 64 means produce a SHA-256.
Is MD5 harder or easier to crack than SHA-256?
For recovering an input from a digest, the two are in the same position: neither can be reversed, and both fall quickly when the input is short or predictable, because an attacker simply hashes candidates until one matches. The “cracking” that works on MD5 and not on SHA-256 is collision building, which is about creating two matching files rather than breaking into one.
Should I use MD5 or SHA-256 to check a download?
Use whichever the publisher published, since the two values cannot be compared with each other. When both are offered, take the SHA-256: it rules out a deliberately substituted file as well as a corrupted one, and an MD5 only rules out the corruption.
Is it worth replacing MD5 with SHA-256 in an existing system?
It is worth it wherever the inputs come from outside — uploads, submitted documents, files from partners — and it is rarely urgent where you generate the data yourself and are only watching for corruption. Do it by widening the field to 64 characters first, then writing both digests during a transition, since there is no way to derive the new value from the old one.