What happens if a file changes while it is being hashed?
A hash is computed by reading a file from one end to the other, and nothing on macOS stops another program writing to it while that happens. You still get a perfectly normal-looking digest back — it just may not describe any version of the file that has ever existed. Here is what Rocket Hash is really reading when the file is moving, how to notice, and how to get a value somebody else can reproduce.
You hashed a log file, or a download that was not quite finished, or a database an application still had open. Nothing failed. There was no error, no warning, and 64 clean hexadecimal characters at the end of it. Then you ran it again and got a different answer.
That is the whole problem in one sentence: hashing a moving file does not break, it lies. The digest is computed over whatever bytes were at each position at the moment the read arrived there, so if the file changed behind the read, the value you got back describes a mixture — part of it from before the change, part from after. Often that mixture corresponds to no state the file was ever in, which means nobody can reproduce your digest, including you.
This page covers what the read actually captures in each of the three ways a file can change, which files this happens to, how to detect it after the fact, and the four ways to get a digest that holds still. If the file in question is a download, verifying a download has the shorter answer: wait for it to finish.
Publish one and the next person to check it gets a mismatch, then has to work out whether they have been attacked. You have not recorded a fact about the file; you have planted a false alarm.
What the read actually captures
A hash is a single pass. The file is opened, read from the beginning to the end in order, and each chunk is folded into a small internal state that becomes the digest at the end. There is no second look. Whatever was at byte 900,000,000 when the read reached byte 900,000,000 is what went in.
macOS does not stop the writer. Hashing takes no exclusive claim on a file — it is an ordinary read, which is precisely why it is safe to point at a master copy — and the locking that exists on a Mac is advisory, meaning it only works between programs that have agreed to respect it. An application with the file open carries on writing, entirely unaware that anybody is reading.
What you end up with depends on how the writer changes the file, and the three cases behave very differently.
| How the file changed | What your digest describes | Can anyone reproduce it? |
|---|---|---|
| Appended to, as a log is | A genuine prefix of the file — a real state it passed through | Only if you also recorded the exact byte count you stopped at |
| Rewritten in place, as a database is | A blend of two eras, old bytes after the read position and new bytes before it | No — that combination never existed on disk |
| Replaced wholesale, as an app's save does | The old version, cleanly and completely | No — the file you hashed no longer has a name |
The first row is the benign one and the most misunderstood. If bytes are only ever added to the end, everything you read really was the file's content at some moment, so the digest is valid — for a file of exactly that length. Since nothing recorded the length, the value is a true statement about a file nobody can identify.
The third row is the one that catches people out. Many Mac applications save by writing a new file and renaming it over the old one, which is what makes saving atomic and crash-safe. If that happens while you are reading, your open handle keeps reading the original, which now exists only for as long as you hold it open. You get a flawless digest of a file that has already been replaced, and hashing the same path a second later gives a different, equally flawless answer.
The files this happens to
- A download in progress. While it runs, a browser writes to a temporary name — something ending in
.crdownload,.partor.download— and renames it when it is done. Hash it early and you get a digest of however much had arrived. This is the most common cause of a download that “does not match” when nothing is wrong with it at all. - Logs and anything with a tail. These append continuously, so a digest is a statement about a length rather than a file. Fine to do, useless to share.
- An open database. Worse than it looks, because the main file may not even be the whole database. SQLite, in the write-ahead-logging mode most Mac apps run it in, keeps recently committed transactions in a second file beside the database, so hashing the main file on its own can give you a stable, reproducible digest of a database that is missing data it has already accepted.
- Anything inside a sync folder. iCloud Drive, Dropbox and the rest can rewrite a file underneath you when another machine changes it, and can be materializing an evicted file at the same time you are reading it.
- A running virtual machine's disk image, an open mail store, a Photos library, a document with autosave on. All of them are live, all of them are large enough that hashing takes long enough to matter.
A digest is computed over the file's data only. Renaming it, moving it, or adding a colored tag or a comment leaves the digest identical, because none of that lives in the bytes.
How to tell it happened
There is no alarm, so you check deliberately. Four checks, in order of how much they are worth.
Hash it twice. Two different digests for the same path is proof the file moved. Two identical digests is good evidence it did not, but not proof — something could have changed and changed back, or changed after the second read. For anything you intend to publish, this is the cheapest sanity check available, and in the Files tool it is two gestures: the reset button empties the list, and dropping the same file in again reads it from the disk a second time.
Copy it first, then compare the copy with the original. The copy has stopped moving and the original has not, so a disagreement between them is movement caught in the act rather than inferred afterwards. Agreement is softer evidence, for the same reason two matching reads are: both could have landed between writes. Put the copy in one well of the Verify tool's File vs. File mode and the live file in the other, and you get a sentence instead of two rows of hex — “Files are identical.”, with the algorithm it used underneath, or the plain statement that they are not. Comparing two files covers the mode in full.
Compare the size before and after. A file row in the Files tool shows the size it saw, and the status bar sums it for the whole batch; if the Finder now reports a different number, you hashed a different file than the one sitting there. For a download, compare against the size on the publisher's page — a short file is a truncated file, and the digest was never going to match.
Compare the modification time too — with the caveat that it is the weakest of the four. The Finder will tell you when a file was last written, before the run and again after it, but plenty of tools preserve or restore a modification date while changing the contents, and two writes inside the same second are indistinguishable at one-second resolution. A changed timestamp is strong evidence; an unchanged one is barely evidence at all.
How to get a digest that holds still
Four options, cheapest first. The right one depends on whether you can stop the writer.
- Wait. For a download, the only correct answer. The browser renames the file off its temporary name when it finishes, which is your signal that the bytes have stopped arriving.
- Stop the writer. Quit the application, stop the service, close the database. A digest of a file nobody is touching is the only kind worth putting in an email, and quitting an app takes five seconds.
- Copy it, then hash the copy. The subtlety here is worth understanding: the copy can itself be a blend, because copying races with the writer exactly as hashing does. What the copy gives you is something that stops moving. Your digest then genuinely describes the copy — so pair the value with the copy, archive the copy, and send the copy. It is the wrong move if you needed a statement about the original.
- Freeze the filesystem underneath it. APFS snapshots are read-only and point-in-time, which is exactly the guarantee the other three lack. Time Machine takes local snapshots of the startup volume as a matter of course, and a snapshot can be mounted as a read-only volume — the file inside it cannot change while you read it, so dropping that version into the Files tool gives you a digest of one instant while the live file carries on moving.
For a database, none of the four beats asking the application itself. A database engine knows what a consistent state is and your filesystem does not, so an export, or the engine's own backup routine, produces a file that is both stable and complete — then hash that. The same logic applies to a backup set: see verifying a backup for checking an archive whose contents are supposed to have stopped changing.
Why long runs make this likelier
A 40 MB installer is read in well under a second, and the window for anything to interfere is too narrow to worry about. A 900 GB disk image on an external drive takes an hour, and an hour is long enough for a backup agent, a sync client or your own forgetfulness to touch the file. The exposure scales with the duration, which is the one respect in which hashing a very large file is genuinely different from hashing a small one.
Pausing a run has the same implication. Rocket Hash can resume from the exact byte it stopped at, which is what makes a long job survivable — but resuming necessarily assumes the bytes already read have not changed in the meantime. For a file on an archive disk that nothing else touches, that assumption is safe. For a live file you paused overnight, start again rather than resume, and pausing and resuming covers the rest of the mechanics.
Troubleshooting
Two runs on the same file give different digests
Then the file changed between or during them, and the question is only what changed it. Look for the obvious writers first: a download still running, a sync client, an app with the document open, a backup restoring over the top. Check the size and modification time, and if both are steady while the digest still moves, something is rewriting the contents in place.
The digest is stable but nobody else can reproduce it
Then you are both hashing steady files — just not the same one. A prefix of a log, a database without its write-ahead log, a version replaced by a save since you started, or two genuinely different builds. Exchange the file sizes before you exchange digests; a size mismatch identifies the problem immediately, and why two tools give different hashes covers the causes that have nothing to do with movement.
The file is smaller than the download page says
It is incomplete, and no digest computed from it will ever match. Let the download finish or start it again. A partial file is the single most common reason a checksum fails, well ahead of anything sinister — what to do when a checksum does not match takes the rest from most likely downwards.
The hash changed but I did not touch the file
Something else did. Autosave, a sync client reconciling with another machine, a backup agent writing a restored copy, or an application rewriting metadata inside the file's own format — a media app updating embedded tags changes the data bytes, even though it looks like it only edited a label. Finder tags and renames, by contrast, never change a digest.
Can hashing a file that is being written damage it?
No. Hashing only ever reads, so the worst outcome is a digest that describes nothing useful — the file itself is no more at risk than if you had opened a window to look at it. Whether hashing changes a file has the full answer, which is the reassuring one.
Frequently asked questions
What happens if a file changes while it is being hashed?
You get a digest with no error and no warning, computed over whatever bytes were at each position when the read reached them. If the file was appended to, that value describes a genuine earlier state of a length you did not record. If it was rewritten in place, the value blends old and new content and describes a version of the file that never existed, so nobody can reproduce it.
Can I hash a file that is still downloading?
You can, and the result is meaningless — it covers only the part that had arrived. Wait until the browser renames the file off its temporary .crdownload, .part or .download name, then hash it. If the digest does not match the published one, check the file size against the download page first; a truncated file is by far the most common cause.
Why do I get a different hash every time I hash the same file?
Because the file is not the same file each time. Something is writing to it: a download in progress, a log being appended to, a sync client, an app with the document open, or an application saving by replacing the file wholesale. Identical bytes always produce an identical digest, so a value that moves is telling you the bytes are moving.
Does hashing lock a file or stop other programs writing to it?
No. Hashing is an ordinary read and takes no exclusive claim, and the file locking macOS offers is advisory — it only binds programs that choose to cooperate. That is what makes hashing safe to point at originals and master copies, and also why nothing prevents a writer from changing a file mid-read.
How do I hash a database or a file that is always open?
Ask the application for a consistent copy rather than hashing the live file. An engine knows what a complete state looks like and the filesystem does not — hashing a SQLite database on its own can miss committed transactions still sitting in its write-ahead log. Export or back it up properly, then hash the export. Failing that, quit the application, or take an APFS snapshot and hash the frozen view.
Is a checksum of a log file useful?
Only with a length attached. Because a log is appended to, the digest you compute is a valid digest of the file as it stood at whatever size you happened to stop at, which nobody can reproduce without knowing that size. Record the byte count alongside the digest, or copy the log to a fixed file and hash that instead.
Can a file change without its modification date changing?
Yes. Tools can preserve or restore a modification time while altering the contents, and two writes inside the same second are indistinguishable at one-second resolution. A changed timestamp is strong evidence that something happened; an unchanged one is weak evidence that nothing did. Hashing the file twice is a better check.