Choosing an Algorithm

Is SHA-1 still safe to use?

A function that was publicly and dramatically broken in 2017 is still naming every object in every Git repository you have ever cloned. That combination is not an oversight, and it is not denial — it is what happens when an attack destroys one property of a hash and leaves the others intact. Rocket Hash keeps SHA-1 under a heading marked LEGACY for exactly that reason: sometimes you still need it, and it helps to know why.

Forty characters of hexadecimal, and somebody wants a ruling on them. Usually it is a scanner that has flagged SHA-1 in a pipeline, a firmware archive whose download page offers nothing else, or the fact that git log is full of values produced by an algorithm everyone agrees is broken.

The answer in one sentence: SHA-1 is unsafe anywhere an adversary gets to write the input, because manufacturing two files with the same SHA-1 costs tens of thousands of dollars rather than a national budget — and it remains perfectly good at detecting accidental damage and at pinning bytes you produced yourself. The property that fell is collision resistance. Preimage resistance never did, and the gap between them explains every apparent contradiction on this page. MD5 went through the same thing thirteen years earlier and ended up in the same awkward position.

What follows is what the 2017 collision actually produced, why the 2020 result was the one that should have worried people, the reassuring things the attacks did not touch, the three fixes that do not work, and why Git survived all of it.

The test, in two questions

Did an adversary get to choose the bytes? And did they get to choose them before the digest was written down? Two noes and your SHA-1 is still doing useful work. One yes and it is decoration.

What SHAttered actually produced

On 23 February 2017, researchers at CWI Amsterdam and Google published two PDF files. They displayed different content, they were the same size, and they had the same SHA-1. The attack was named SHAttered and it ended a twelve-year argument: the 2005 cryptanalysis that first put a SHA-1 collision below brute force had remained theoretical for long enough that plenty of people had stopped expecting anyone to finish the job.

The numbers are worth knowing because they are the reason the result was believed rather than debated. It took roughly nine quintillion SHA-1 computations — about 2 to the power of 63 — which Google described as 6,500 CPU-years plus 110 GPU-years of work. Rented rather than owned, that was somewhere around 110,000 dollars of cloud time at 2017 prices. Expensive, entirely purchasable, and roughly a hundred thousand times cheaper than the 2 to the power of 80 a 160-bit digest is supposed to cost.

What it was not is a way to forge an arbitrary document. SHAttered was an identical-prefix collision: the two files had to begin with the same bytes, then contain a pair of carefully crafted near-identical blocks, then continue identically to the end. The researchers worked inside that constraint with a trick — both PDFs drew an image from the same embedded JPEG, and the colliding blocks decided which of two pictures that JPEG rendered. Clever, convincing, and still a long way from being able to collide two documents you did not design together.

The 2020 result was the dangerous one

Three years later Gaëtan Leurent and Thomas Peyrin published the attack that removed the constraint. Their paper, titled “SHA-1 is a Shambles”, produced the first chosen-prefix collision for SHA-1 in January 2020, for about 45,000 dollars of rented GPU time — less than half the cost of the weaker 2017 result.

Chosen-prefix is the step change, and it is worth being precise about why. With an identical-prefix collision, an attacker must author both files around a shared skeleton; with a chosen-prefix collision, the attacker takes two arbitrary beginnings that were written independently and for different purposes — two contracts, two certificates, two executables, two public keys — and computes bridging blocks that make the digests agree from there on. To prove the point, Leurent and Peyrin used theirs to forge a PGP identity certification, transferring a signature from one key to a key belonging to somebody else entirely.

This is precisely the escalation MD5 went through between 2004 and 2008, and it ended the same way: the moment chosen-prefix collisions are affordable, a digest stops being usable as a statement about which document you are holding.

The price only ever falls

2 to the power of 69 in 2005, 110,000 dollars in 2017, 45,000 dollars in 2020 — cryptanalysis shaves exponents and hardware gets cheaper, and neither direction has ever reversed. “Too expensive to bother attacking” is a statement with a date on it, not a property of an algorithm.

What the attacks did not break

Nothing published touches preimage resistance. There is still no practical method for taking a SHA-1 digest and producing any input that matches it; the published preimage work applies to versions of SHA-1 with rounds removed, and the full function remains out of reach. Second preimages are just as safe: given a file somebody else made, nobody can construct a different file with that file's digest.

Three useful consequences fall straight out of that.

  • A SHA-1 you recorded for a file you made still pins those bytes. If you hashed your own release, your own archive or your own master copy, a mismatch six months later still means something changed, because altering the file while preserving the digest would require a second preimage. The collision attack needs both files built together, which nobody did on your behalf.
  • HMAC-SHA1 is not forgeable. HMAC's security argument does not depend on the collision resistance of the hash inside it, and the key the attacker does not have is doing the work. This is why a great many old APIs, signed URLs and one-time-password schemes went on functioning through 2017 without an emergency.
  • Accidental corruption is still caught perfectly. A truncated download or a failing cable cannot construct anything. Damage with no agenda has no route to a matching digest, which is why the 40-character value on a firmware page is worth the twenty seconds — read what a digest proves about tampering for where that stops being enough.

So the decision rule is about authorship, not about dates. SHA-1 fails when the person who wrote the file benefits from two versions of it matching. It holds when the bytes came from you, or from somebody with no motive and no opportunity to have prepared a pair in advance.

The fixes that do not work

Three workarounds get proposed every time this comes up, and none of them does what people hope.

  • Hashing twice. If two inputs collide under SHA-1 they produce the same digest, and feeding that identical digest through SHA-1 a second time produces the same answer again. Double hashing cannot undo a collision — it is the same collision, one step later. It does defeat length extension, which is a different problem with a different fix.
  • Publishing a SHA-1 and an MD5 side by side. Intuition says two broken hashes are harder to break at once. Antoine Joux showed in 2004 that concatenating iterated hash functions buys far less than that: multicollisions in this kind of construction are cheap enough that the combination is barely stronger than the stronger of the two components. Two weak hashes make a weak hash.
  • Shortening the digest to something readable. Abbreviating makes it worse, not safer, and dramatically so. Git's short hashes exist for human conversation, and large repositories have already had to lengthen them because accidental collisions started appearing — seven characters is a nickname, not an identifier.

What works is a different function. SHA-256 for almost everything, with SHA-3 as a structurally unrelated alternative where a specification calls for it; which hash algorithm to use has the rest of the decision. Two digests from two unrelated designs are genuinely worth more than two from the same lineage, which is the one version of the “publish both” idea that holds up.

Why Git is the exception

Git names every blob, tree, commit and tag by the SHA-1 of its contents, which sounds like the worst possible place for a broken hash. Three things have kept it standing.

First, Git has not used plain SHA-1 since 2017. It ships a hardened variant — collision-detecting SHA-1, from Marc Stevens and Dan Shumow — which watches the function's internal state for the particular disturbance pattern every known collision attack has to produce, and refuses the input when it sees one. It costs a little speed and it rejects both published collisions outright. Feeding a SHAttered PDF to Git does not produce a quiet collision; it produces a refusal.

Second, the attack is not passive. To exploit a Git collision you must get your crafted object accepted into a repository other people rely on, having already placed its benign twin there — a far higher bar than intercepting a download. That does not make it imaginary, and it was the specific worry for hosting services, but it is not something that happens to you by accident.

Third, the project is migrating anyway, and for the honest reason. Signed commits and signed tags put a signature over an object name, so collision resistance does matter to Git's authenticity story even if its integrity story survives. A SHA-256 object format exists and works, but a SHA-256 repository cannot yet interoperate with a SHA-1 one, which is why almost nobody has moved. That is an ecosystem problem rather than a cryptographic one, and it is the clearest illustration of the real cost of a broken hash: not a catastrophe, a decade of migration.

The deadline, and who is already enforcing it

Most of the genuinely dangerous uses of SHA-1 were taken away from you by other people years ago. Certificate authorities were required to stop issuing SHA-1 certificates for TLS from the start of 2016, and the major browsers began rejecting them outright in 2017, so your Mac will not negotiate a connection secured that way whatever you do. Code signing moved on at the same time.

The formal position is older than the attack. NIST disallowed SHA-1 for generating new digital signatures at the end of 2013, and in December 2022 it announced the algorithm will be retired completely by 31 December 2030, with federal systems expected to have finished the transition by then. If something you own still depends on SHA-1, that is a dated obligation rather than a suggestion.

What is left, then, is the long tail nobody can switch off for you: internal scripts, vendor download tables, firmware archives, BitTorrent version 1, and a decade of automation written around 2012. Verifying a download covers the case you will hit first, where a page offers a SHA-1 and nothing stronger. The rest is the ruling above, applied one case at a time — and when you need the digest itself, producing a SHA-1 on a Mac covers the mechanics, including telling a file's SHA-1 apart from the other 40-character values floating around.

Troubleshooting

A scanner flagged SHA-1 in our Git tooling

Separate the two kinds of finding before you open a ticket. Object names that Git itself produced are not yours to change and are protected by the collision detection described above; that finding is noise and the only real fix is Git's own migration. Code you wrote that computes a SHA-1 over content somebody else supplied — an upload, a webhook payload, a third-party artifact — and then treats a match as proof of identity is a genuine finding, and it is usually a one-line change to SHA-256.

Somebody says SHA-1 is fine because there is no preimage attack

The premise is true and the conclusion does not follow. A signature is computed over a digest, so an attacker who can produce two documents with one digest gets a signature on the second for free the moment you sign the first — no preimage required. Preimage resistance is what protects a digest you recorded for a file that already existed; collision resistance is what protects you from documents that were written to match each other. Different jobs, and only one of them is still covered.

We have a SHA-1 for an archive somebody else gave us

Then the digest is a weaker claim than it looks, and nothing you do now can strengthen it retroactively. If the sender prepared a colliding pair before handing it over, a match tells you only that you hold one of them. In practice the question is whether anyone had a motive years ago; if the archive matters, the right move is to re-verify it against the source and record a SHA-256 of what you actually have, so that from today onward the fingerprint is one you trust.

We cannot drop SHA-1 because another system requires it

Then keep producing it and stop relying on it. Treat the SHA-1 as an identifier the other system needs, record a SHA-256 of the same bytes on your side, and make the SHA-256 the value any security decision is made against. Computing both costs nothing extra, because Rocket Hash produces all eight digests from one read of the file, and the SHA-1 generator on this site is enough for a one-off if you only need to check a single value.

Frequently asked questions

What did the SHAttered attack actually do?

In February 2017, researchers at CWI Amsterdam and Google published two PDF files with different visible content and the same SHA-1 digest. It took about 2 to the power of 63 computations — described at the time as 6,500 CPU-years plus 110 GPU-years, or roughly 110,000 dollars of rented cloud time. It was the first demonstrated SHA-1 collision, twelve years after the attack was first shown to be theoretically possible.

How much does a SHA-1 collision cost now?

The 2017 collision was estimated at around 110,000 dollars of rented computing time, and by January 2020 a stronger chosen-prefix collision — the kind that can collide two documents written independently — was produced for about 45,000 dollars. Both figures fall over time as hardware gets cheaper and cryptanalysis improves, so any claim that an attack is uneconomical should be read as a statement about today.

Is it safe to use SHA-1 to check a download?

It reliably tells you the download arrived intact, because corruption cannot construct a matching digest. It does not establish that the file is the one the publisher meant to give you, since anyone who built both files could have made them share a digest. If the publisher offers a SHA-256 as well, use that; if SHA-1 is all there is, check it and treat the result as evidence about the transfer rather than about the file's origin.

Can somebody create two Git commits with the same hash?

Not with the published attacks, because Git has used a collision-detecting variant of SHA-1 since 2017 that recognizes the internal pattern those attacks must produce and rejects the input. An exploit would also require getting both the benign and the crafted object into a repository other people trust. Git is nonetheless migrating its object format to SHA-256, since signatures are computed over object names.

Is HMAC-SHA1 broken too?

No. HMAC's security does not rest on the collision resistance of the hash inside it, and an attacker without the key cannot compute the inner value at all, so the 2017 and 2020 collisions did not make HMAC-SHA1 forgeable. It is a deprecated choice rather than a dangerous one: keep it where you must for compatibility, and do not pick it for anything new.

When will SHA-1 be removed completely?

NIST disallowed SHA-1 for generating new digital signatures at the end of 2013 and announced in December 2022 that it will retire the algorithm entirely by 31 December 2030. Certificate authorities stopped issuing SHA-1 TLS certificates at the start of 2016 and browsers rejected them from 2017, so in the places that mattered most it has already gone.