How to hash a photo or video file on Mac
Record a digest for a photograph or a video once and you can prove, years later, that a copy is the same file down to the byte. Rocket Hash does that in one pass — and this page also explains the flip side, which is why re-exporting the same shot never gives you the same digest twice.
You have one set of originals off a camera card, a working copy on the laptop, and an archive copy on a drive in a drawer. Nobody has opened the archive in eighteen months. The question you want answered is narrow and concrete: are those three still the same files, or has one of them quietly changed?
A digest answers that and almost nothing else. Hash the original once, keep the value, and any time afterwards you can prove a copy is the same file down to the last byte — which is the strongest claim available about a photograph or a video, and weaker than people assume. It says nothing about whether the image was edited before you recorded the digest, and it will not recognize the same shot exported twice.
That second half is why this page exists. Re-export the same clip with the same settings and you get a different digest; add a keyword to a JPEG and you get a different digest; the picture on screen is identical in both cases. None of that is a fault, and all of it is predictable once you know what the digest is actually looking at.
A hash reads bytes. It has no idea what a photograph is, so “the same image” and “the same file” are two different claims and only the second one is testable here.
The bytes are not the picture
A JPEG is not a grid of pixels on disk. It is a container: a header, EXIF metadata with the camera model and the exposure and possibly your coordinates, maybe an embedded thumbnail and a color profile, then the entropy-coded image data. A digest covers all of it equally. Change the GPS tag and leave every pixel untouched and you have a new file as far as SHA-256 is concerned: about half the 256 bits flip, which leaves almost none of the 64 characters where they were.
The reverse trips people up more. Two files that produce identical pictures — the same photo exported at quality 100 by two different applications — have completely unrelated digests, and no amount of staring at them reveals that the images match. Measuring visual similarity is the job of a perceptual hash, which is a different kind of function built for a different purpose, and none of the eight algorithms here is one. SHA-256, SHA-512, MD5 and the rest are deliberately built so that a one-bit change scrambles the output; that is precisely the property that makes them useless for “looks like”.
So use a digest for what it is good at: proving a specific file traveled from A to B intact, and catching the day it stops being intact.
Fingerprint a master and check the copies later
-
Hash the originals before you do anything else
Do this while the files are still the camera's output — before an import, before a rename, before any application has had the chance to write a rating into them. Drop the card's folder, or the freshly offloaded folder, into the Files tool with the Algorithm control on SHA-256. Every file underneath gets a row of its own.
-
Export the list and store it somewhere else
The export button turns the finished run into a
shasum-compatible manifest. Keep that file off the disk it describes — a manifest that dies with the drive it was protecting has not protected anything. Exporting a manifest covers the format and where to keep it. -
Check the copy before you erase the card
Hash the folder you just copied, export a second manifest from it, and compare the two lists line by line. This is the one moment in the whole workflow where a mismatch is cheap, because the original is still in your hand. After the card is formatted, a bad copy is the only copy. Hashing files on an external drive covers reading the card in place.
-
Re-check the archive once a year
Run the same folder again, export a fresh manifest and diff it against the one you filed away. Files whose digests changed without your having touched them are the interesting ones — that is what silent corruption looks like from the outside. Verifying a backup goes through that cycle properly.
Why an export never matches its source
Because an export is a new file built by an encoder, and encoders are not required to be deterministic between versions, settings or applications. Even a change nobody would call an edit rewrites enough of the container to change everything.
| What you did | Digest changes? | Picture changes? |
|---|---|---|
| Copied it to another disk | No | No |
| Renamed or moved it | No | No |
| Added a keyword, rating or caption | Yes | No |
| Stripped the GPS tag before sharing | Yes | No |
| Rotated it in the Finder | Yes | Only the orientation |
| Exported a JPEG at quality 100 | Yes | Not visibly |
| Rewrapped a MOV as an MP4 | Yes | No — same frames |
| Made a proxy or transcoded it | Yes | Slightly |
Rows three to eight are the ones that generate support questions. The useful takeaway is that a digest mismatch between a master and an export is the normal, expected result, and a digest match between two files you produced separately would be the surprising one. If you want to compare a master against its export, compare the images with your eyes; the hash has nothing to contribute.
RAW files, sidecars and pairs
This is where the archive workflow gets tidy rather than messy. Editors generally refuse to write into a proprietary RAW file — a .CR3, .NEF or .ARW — and put your adjustments, keywords and ratings into a separate .xmp sidecar beside it. The consequence is excellent: the RAW's digest stays valid for the life of the file no matter how much you edit, and the sidecar, which changes constantly, is a few kilobytes you can simply expect to change.
DNG behaves differently, because metadata can be written inside it — so a DNG's digest does move when you rate or keyword it. Shoot RAW+JPEG and you have two files per frame with two unrelated digests, both of which belong in the manifest. An iPhone Live Photo is likewise a still and a short movie, two files, two digests; copy one and leave the other and a manifest check will tell you so.
It shows a file has not changed since you recorded the value. It cannot show the photograph is unmanipulated, or that it came from the camera it claims — that needs a signature, which is a different mechanism entirely.
Video files and the time they take
Video is the same problem with bigger numbers. A two-hour ProRes master costs no more memory than a 2 MB JPEG, because the bytes are read in sequence and then let go, so the only real constraint on a 400 GB project folder is how fast the drive can hand it over. Expect a little over an hour for 500 GB on a USB hard disk and around ten minutes on a fast external SSD, and see hashing a very large file for planning a run at that scale.
One habit worth forming: hash the camera original, not the version your editor has conformed, optimized or rewrapped. Editing applications legitimately rewrite container metadata, and a digest taken after that point describes the application's file rather than the camera's.
Which algorithm for a photo archive
SHA-256 unless something else makes the decision for you. It is 64 characters, it is what the rest of the world publishes, and nobody has ever produced two files with the same SHA-256 digest.
Plenty of preservation and broadcast workflows specify MD5, and if yours does, use it rather than fighting the tooling. The threat an archive actually faces is a disk going bad over a decade on a shelf, not somebody building a forged negative, and 128 bits of MD5 catches a flipped bit as reliably as 256 bits of SHA-256 does. Where it stops being enough is the moment a file arrives from outside your own shelves — a frame a client sends back, footage bought from a library — because then the question changes from “has this decayed?” to “is this what they say it is?”, and MD5 cannot answer the second one. MD5 versus SHA-256 sets out exactly which promise broke, and detecting tampering covers the case where that distinction decides the outcome.
CRC32 does not belong here at all — eight characters is too few to be an identifier for a collection of any size, and it is trivial to force a particular value on purpose.
Troubleshooting
The file in my photo library does not match the one I imported
A library is a managed store, not a folder of your files, and what it hands back depends on how you ask. An export of the unmodified original should return the bytes that went in; an ordinary export re-encodes, writes fresh metadata, and will not match anything. Test it once with a single file rather than taking anyone's word for it: hash the file you import, export it back out, hash that, and you will know exactly what your library does.
Two identical-looking photos have completely different digests
Expected. They are different files — different metadata, a different encoder, a different quality setting, a different embedded thumbnail. A cryptographic hash is designed to make a tiny difference produce a totally different output, so “completely different” carries no information about how large the underlying change was. For grouping actual byte-identical duplicates, finding duplicates by hash is the right route; for grouping similar-looking shots, no hash will help you.
A digest changed and I did not touch the file
Something else did, or the storage did. Check the file's size and modification date first: a changed date usually means an application wrote metadata — a rating, a keyword, a face tag, an orientation fix — while the same size and same date with a different digest is a much more serious signal and suggests the drive. Open the file and look at it. Corruption in a JPEG typically shows as color-shifted or gray blocks from the damaged point down, and in video as a glitch lasting a second or two.
The copy on the NAS does not match the card
Compare the two sizes first; if they differ, the transfer has not finished and there is nothing to investigate yet. Once the sizes agree, check you are comparing like with like: a sidecar is not its RAW, a Live Photo is a still plus a movie, and some copy tools strip or relocate Mac metadata on the way across. Hashing writes nothing at either end, so it is safe to rerun both sides as many times as you need — see does hashing change a file.
Hashing thousands of photos is slower than one big video
It will be, and it is not the algorithm: ten thousand 4 MB JPEGs are ten thousand separate files to open and close, where one 40 GB movie is one. Two things make a photo archive worse than that arithmetic suggests. A RAW+JPEG card carries two files per frame and a Live Photo carries two per shot, so the count you are really hashing is often double the number of pictures you took. And the card reader is usually the slowest link in the chain, so the same files will hash several times faster once they are on your internal disk — budget for the card being slow rather than assuming something is wrong with it.
Frequently asked questions
Will hashing a photo or video change or damage it?
No. Hashing reads every byte in order and writes none of them back: no pixel, no metadata tag, no modification date. That is what makes it safe to point at camera originals, a client's drive or an archive you do not want disturbed.
Why do two exports of the same video have different checksums?
Because each export is a new file produced by an encoder, and encoders are not obliged to produce identical bytes twice. Container metadata, timestamps, the exact encoder build and any setting that differs all change the output. The frames can look identical while the files share nothing — a mismatch between a master and an export is the expected result, not a fault.
Can a checksum tell me whether two photos are the same image?
Only in the strict sense of being the same file. Two byte-identical copies give the same digest; two visually identical images saved by different applications give completely unrelated ones. Judging whether two pictures look alike is the job of a perceptual hash, which is a different kind of function — SHA-256 and friends are built so that any difference at all scrambles the result.
Which hash should I use for a photo or video archive?
SHA-256 unless an existing workflow dictates otherwise. It is the value everyone else publishes, and no two inputs with the same SHA-256 digest have ever been found. If your preservation tooling specifies MD5, that is defensible for detecting decay — nobody can construct a second file matching a digest you recorded — but it is not enough to prove a file someone hands you today is genuine.
Does editing a photo's keywords or metadata change its checksum?
Yes, whenever the change is written into the file itself, which is the usual case for JPEG, TIFF and DNG. With proprietary RAW formats most editors write to a separate .xmp sidecar instead, so the RAW's digest survives any amount of editing and only the sidecar changes. That is a large part of why sidecars exist.
How can I prove a photo has not been edited?
You cannot, in general — you can only prove it has not changed since you recorded its digest. A hash taken today says nothing about what happened to the file yesterday, and nothing about where it came from. Proving origin requires a signature over the digest, which binds it to a key rather than merely to the bytes.
How long does it take to checksum 500 GB of video on a Mac?
It is set by the drive, not by the algorithm. Reckon on a little over an hour from a USB hard disk at around 120 MB/s, roughly ten minutes from a fast external SSD, and a few minutes from an internal Apple SSD. SHA-256 manages approximately 2 GB/s on Apple Silicon, which almost no storage can keep up with.