Hashing Text

How to hash Unicode and emoji text on Mac

A hash function cannot see letters, only bytes — so the interesting question is how many bytes your text turned into, and why that number is not the one you expected. The byte count under the text field in Rocket Hash answers it live, and once you know what each possible surplus byte means, a mismatch stops being a mystery and becomes a reading.

You and somebody else hashed what you both believe is the same string and got two different digests. Or a test that passes on Linux fails on your Mac, and the only thing it touches is one accented character. Or an emoji copied from a phone refuses to match the same emoji typed in a browser.

None of these is a bug, and none of them is about the algorithm. A hash function is handed a sequence of bytes; text becomes bytes only after an encoding has been applied; and two strings that are indistinguishable on screen can be entirely different sequences of bytes. When they are, the digests have no reason to agree, because they are answers to different questions.

What follows is the arithmetic that turns characters into bytes, so you can look at a byte count, compare it with the number of characters you think you typed, and say what the difference is made of. Hashing text on a Mac covers the basic procedure and the commonest culprits, and why two tools disagree covers the causes that have nothing to do with encoding. This page is the detailed version of the encoding half.

The digest is a fact about bytes

Not about characters, not about meaning, and not about how the text renders in the font you happen to be using. Every mismatch below follows from that one sentence.

How many bytes a character costs

Essentially everything modern encodes text as UTF-8, and UTF-8 is variable width: a character costs one to four bytes depending on where it sits in Unicode. This is the whole foundation, and it is four rows long.

How many UTF-8 bytes each range of Unicode code points requires
Code pointsUTF-8 bytes eachWhat lives there
U+0000–U+007F1a–z, digits, ASCII punctuation, space, newline
U+0080–U+07FF2Accented Latin, Greek, Cyrillic, Hebrew, Arabic, combining marks
U+0800–U+FFFF3CJK, curly quotes, zero-width characters, variation selectors
U+10000 and above4Emoji, musical notation, the rarer scripts

So hello is five characters and five bytes, and they coincide only because it is ASCII. café is four characters and five bytes. A single rocket is one character and four bytes. The byte count shown under the text field is the number the digest is actually computed over, which makes it the only measurement on screen worth trusting.

Find the difference between two strings

  1. Hash the first version and note the count

    Open Text in the sidebar and paste the first of your two strings. Read the byte count beneath the field — 23 bytes, or whatever it says — and write it down next to the number of characters you believe the string contains.

  2. Clear the field and paste the second

    Use the small ⊗ button at the top-right corner of the field rather than selecting and deleting, which can leave a stray character behind. Paste the second version and read its byte count. Everything recomputes as you type, so both numbers arrive without you pressing anything.

  3. Subtract

    The difference between the two counts is the size of your problem, and the next section translates it. A difference of three bytes with nothing visibly different means something invisible; a difference of one per accented letter means the two strings spell that letter differently.

  4. Bisect to locate it

    When the counts differ but you cannot see where, delete the back half of the string and watch the number. If the surplus is still there, the culprit is in the front half; if it vanished, it was in the back. Four or five rounds of that will find a zero-width character in a sentence.

What a surprising byte count is telling you

This is the table worth keeping. Compare the byte count against the characters you believe you have, and read off what the surplus is made of.

Diagnosing the cause of an unexpected byte count
The byte count isAlmost certainly
Equal to the character countPlain ASCII — look elsewhere for the mismatch
One higher per accented letterNormal: the precomposed single code point
Two higher per accented letterDecomposed: a plain letter plus a separate combining mark
Three higher, with nothing new on screenA zero-width character, a byte-order mark, a variation selector or a direction mark
One higher where a space should beA non-breaking space standing in for an ordinary one
Exactly double, for ASCII textUTF-16, not UTF-8
Three higher per emojiNormal: four bytes for one character — but eight bytes for one picture means a skin tone, and eleven or more means a joined sequence

The second and third rows are the same character written two legal ways, and they are the single most common cause of this whole problem on a Mac. Unicode lets é be one code point, or a plain e followed by an invisible combining acute accent. The composed spelling is called NFC, the decomposed one NFD, and they render identically in every font you own. macOS has historically preferred the decomposed form for filenames, which is why a name copied out of the Finder can refuse to match the same name typed on Linux.

Emoji are sequences, not characters

An emoji is very often several code points pretending to be one picture, and every one of them goes into the digest. These counts are worth a minute of your attention, because they explain nearly every emoji mismatch.

Code point and UTF-8 byte counts for several emoji
On screenCode pointsUTF-8 bytesWhat is in there
🚀14One code point, above U+FFFF
❤13The bare symbol, from the older part of Unicode
❤️26The same symbol plus a selector asking for the color form
👍🏽28The hand, plus a skin-tone modifier
🇬🇧28Two regional-indicator letters — a flag is GB in disguise
👩‍💻311A woman, a zero-width joiner, a laptop
👨‍👩‍👧‍👦725Four people and three joiners holding them together

The two hearts are the pair that causes arguments: they are the same symbol, three bytes apart, and which one you get depends on the keyboard or the web page the character came from. The family is the one to remember for scale — one thing on screen, seven code points, twenty-five bytes. A string containing emoji has three different lengths: what you see, what your programming language reports, and what gets hashed. Only the third one decides the digest.

When the byte counts agree but the digests do not

The byte count narrows things down; it does not prove anything. Equal counts with different digests means you have the same number of different bytes, and there are three realistic ways to get there.

  • Case. Check this first and feel no shame about it. Café and café are the same length and nowhere near the same digest.
  • Lookalike characters of equal width. A Greek omicron and a Cyrillic о are both two bytes and both look like an o to you. Substitutions like that survive a byte count intact, which is exactly why they are used in phishing domains.
  • Combining marks in a different order. A letter carrying two accents — ệ, say — can be written with the marks in either sequence. Both spellings are five bytes, both render identically, and only one of them is canonical, which is exactly the problem normalization exists to solve.

Should you normalize before hashing?

It depends entirely on what the digest is for, and the two answers point in opposite directions.

Verifying a file: never. The digest has to describe the bytes as they sit on disk, surplus byte-order mark and all, because that is what the person on the other end will hash. Normalizing the content first would change the file and produce a digest for something that does not exist.

Comparing text that humans typed: yes, on both sides. If two systems have to agree on a digest for a name, an address or a search term, they have to agree on a normalization form first. NFC is the usual choice — it is what the web recommends and what most text is already in. Apply it, write down that you applied it, and make sure the other end does the same.

Note that no hash function will do this for you, and no hashing tool can make the decision on your behalf. SHA-256 given two spellings of the same word returns two digests because it received two inputs; it has no concept of text and no opinion about Unicode. Normalization is a decision the two ends of a protocol make before any hashing happens. What the Text tool gives you is the evidence that the decision was never made — two byte counts that disagree by one per accented letter — and hashing text on a Mac covers the rest of getting a string into the field intact.

Never normalize a file you are verifying

Byte-order marks, line endings and composition are part of the file. A digest computed over a tidied-up copy will not match the one the publisher calculated, and the mismatch will look exactly like corruption.

When you need the bytes to stop moving

A text field is a live thing, and some routes between applications rewrite composition on the way through. That is fine while you are diagnosing — the byte count moves with the text and tells you what happened — but it is no good at all when the answer has to survive being sent to somebody else. At that point, stop working with the text and start working with files.

Save each version as its own file and drop both into Files. Each one becomes a row: the name in bold, the folder it came from underneath it, the size, and the digest. That size is the byte count you were reading under the text field, now attached to something nothing can silently renormalize, and the disclosure chevron at the end of the row opens every algorithm for that file out of the single pass already made over it.

If what you want is a verdict rather than two numbers to compare, Verify in File vs. File mode takes the two files in two drop wells, with a swap control between them, and answers in a sentence under a green seal. That matters more here than anywhere else: two spellings of one word are exactly the case a person comparing 64 hexadecimal characters by eye will wave through. Comparing two files on a Mac covers that route in full.

Troubleshooting

The byte count is higher than the characters I typed

For anything outside ASCII that is expected, not a fault. An accented letter costs two bytes, a CJK character three, an emoji four. Work out what the count should be from the width table above, and only start worrying when the real number exceeds your prediction.

Copying the text somewhere else made the mismatch disappear

Some routes between applications rewrite text on the way through, including normalizing its composition, so a round trip through the clipboard or through one particular field can silently change the bytes. Treat that as the diagnosis rather than the cure: whichever version you need to be able to prove later belongs in a file, not in a text field.

The same emoji hashes differently on two devices

Compare the byte counts before anything else. One side almost certainly added a variation selector, a skin-tone modifier or both, which are invisible by design. Three extra bytes is a selector; four extra is a modifier; eleven or more for one picture means a joined sequence.

My string matches but the filename does not

Filenames pass through the filesystem, which has its own habits about composition, so the text inside a file and the name on the outside can be different byte strings while printing identically. When a manifest disagrees about a filename rather than a digest, this is usually why — see how a SHA256SUMS file is laid out for what the two columns actually contain.

I need to prove which version I hashed

Record the byte count alongside the digest. It costs you a few characters, it is the one piece of metadata that distinguishes two spellings of the same word after the fact, and it makes a mismatch diagnosable by somebody who was not there. Copying digests out covers what else is worth sending with one.

Frequently asked questions

Why do two identical-looking strings have different hashes?

Because they are different bytes. The three usual reasons are a character spelled two legal ways — é as one code point or as e plus a combining accent — an invisible character such as a zero-width space or a byte-order mark, and one side using a different encoding entirely. Comparing the byte counts of the two strings tells you which of the three you are looking at.

Should I normalize text before hashing it?

Yes when you are comparing text that people typed, and never when you are verifying a file. For text, pick NFC, apply it on both sides, and record that you did — two systems cannot agree on a digest for human-entered text unless they first agree on a normalization form. For a file, the digest must describe the bytes exactly as they are, including anything you would have tidied up.

Why does the same emoji give a different hash on different devices?

Because most emoji are sequences rather than single characters, and different keyboards emit different sequences. A heart may or may not carry a variation selector after it, which is three extra bytes; a hand may or may not carry a skin-tone modifier, which is four more. Nothing about the difference is visible, but the byte count shows it immediately.

Why does my hash differ from the one a Windows tool produced?

Very likely because the Windows side encoded the string as UTF-16, which a lot of Microsoft documentation simply calls “Unicode”. In UTF-16 every ASCII character carries a zero byte beside it, so a five-letter word becomes ten bytes instead of five and the digest can never match the UTF-8 one. The giveaway is a byte count exactly double the character count.

How do I find an invisible character in a string?

Count the bytes and compare that with the number of characters you expect, then bisect. Delete half the string, see whether the surplus bytes went with it, and repeat on whichever half kept them. A zero-width space, a byte-order mark and a direction mark all cost three bytes; a non-breaking space costs two where an ordinary space costs one.

Does SHA-256 work on non-English text?

Yes, and it treats it no differently — SHA-256 consumes bytes and has no idea whether they represent Japanese, Arabic or an emoji. The only requirement is that both ends encode the text the same way, which in practice means UTF-8. You can confirm any of this in the SHA-256 generator, which encodes what you type as UTF-8 before hashing it.