Daily CryptogramA free daily cryptogram puzzle paired with a historical almanac.
Editorial

Letter Frequency: The Cryptogram Solver's Sharpest Tool

The technique at the heart of cryptogram solving is over a thousand years old and requires nothing but the ability to count. Described first by the Baghdad polymath al-Kindi in the ninth century, frequency analysis rests on a simple observation: a substitution cipher conceals the identity of each letter but not how often it is used. Count the symbols, compare the counts against what English normally does, and the puzzle begins to give itself away.

The order worth memorizing

The approximate frequencies of letters in ordinary English prose, in descending order, are as follows. Percentages vary a little between sources depending on the body of text sampled, but the ranking is stable.

  • E, about 12.7 percent
  • T, about 9.1 percent
  • A, about 8.2 percent
  • O, about 7.5 percent
  • I, about 7.0 percent
  • N, about 6.7 percent
  • S, about 6.3 percent
  • H, about 6.1 percent
  • R, about 6.0 percent
  • D, about 4.3 percent
  • L, about 4.0 percent

Together those eleven letters account for roughly three quarters of all the letters you will ever read. The traditional mnemonic for the top group, inherited from the days of hand-set type, is ETAOIN SHRDLU, which was simply the order of the letter magazines on a Linotype keyboard.

At the other end, the rare letters are worth knowing too, because a symbol that appears exactly once in a puzzle is unlikely to be E and quite likely to be one of these:

  • V, about 1.0 percent
  • K, about 0.8 percent
  • J, X, Q and Z, all under 0.2 percent each

What to do with a count

The practical procedure is unglamorous. Before entering a single letter, go through the puzzle and tally how many times each symbol appears. Write the tallies down. Then look at the top of your list.

If one symbol clearly leads the count, E is your first hypothesis. If two symbols are close at the top, you are probably looking at E and T in some order, and the tie will be broken by position rather than by count, which we come to below. The symbol leading the count in a typical cryptogram of ten to fifteen words is E perhaps half the time, and one of E, T, A, O, I, N, S, H, R the great majority of the time.

Crucially, frequency analysis is a ranking tool, not an identification tool. It tells you which candidates to test first. It does not tell you the answer, and treating it as if it does is the most common way beginners go wrong with it.

Why short puzzles betray the frequencies

Here is the limitation that matters most for a daily cryptogram. The published frequencies describe large bodies of text. A cryptogram is a single quotation, typically ten to fifteen words and sixty to ninety letters. That is a tiny sample, and small samples are noisy.

A quotation about zoology may be thick with Z. A quotation with the word QUALITY in it has a Q where you least expect one. A short sentence can easily contain more S than E. The shorter the puzzle, the less you should trust the raw counts and the more you should lean on structure: word lengths, apostrophes, doubled letters, and the position of symbols within words.

This is why frequency analysis alone rarely finishes a cryptogram, and why the guides that present it as the whole method leave beginners frustrated. Use it to generate an ordered list of hypotheses. Use structure to confirm or kill them.

Position is more informative than raw count

A symbol''s frequency tells you something. Where it sits tells you more. Certain letters have strong positional preferences that survive encryption intact.

  • T, A, O, S, W and I begin words far more often than they end them
  • E, S, T, D, N and Y end words far more often than they begin them
  • E is extremely common as a final letter, and a short word ending in your most frequent symbol is good evidence that symbol is E
  • H appears constantly in second position, because of TH, and almost never at the end of a word
  • U appears in second position with suspicious regularity when the first letter is Q

So if you have two symbols tied at the top of your count, and one of them repeatedly ends words while the other repeatedly starts them, you are very likely looking at E and T respectively. The count could not separate them. Position did.

Pairs and triples

Letters do not occur independently, and the statistics of adjacent pairs, called digraphs, are sharper than those of single letters. The most common digraphs in English, roughly in order, are TH, HE, IN, ER, AN, RE, ON, EN, AT, ND, and OR.

TH deserves special attention. It is the most common pair in the language, and in combination with the extremely common word THE it means that a symbol appearing immediately before your candidate H is almost certainly T. This mutual reinforcement, where T supports H and H supports T, is what makes THE such a reliable first foothold.

The common triples, or trigraphs, are fewer and even more useful: THE, AND, ING, ION, ENT, HER, FOR, THA, NTH and INT. Notice that three of those, THE, AND and ING, are the exact patterns recommended as first tests in any solving guide. That is not a coincidence; it is frequency analysis applied at the level of groups rather than single letters.

Vowels hide in plain sight

There is a neat trick for separating vowels from consonants without identifying either. Vowels are promiscuous: they sit next to a very wide variety of other letters. Consonants are choosier, and some are extremely restricted.

So for each symbol, count not how often it appears but how many distinct symbols it appears adjacent to. The symbols with the widest variety of neighbours are your vowels. This works because English permits almost any consonant to precede or follow a vowel, while consonant clusters are heavily constrained. In a puzzle where the raw frequencies are misleadingly flat, this diversity test often still separates the alphabet cleanly into two groups, and knowing which five or six symbols are vowels is enormous progress.

A related signal: every English word contains a vowel, with a handful of exceptions such as WHY and RHYTHM. So any word in which you have found no vowel candidate should make you suspicious of your assignments.

Al-Kindi''s method, unchanged

What is remarkable about all of this is how little it has changed. Al-Kindi''s instruction was to find a long text in the same language as the message, count how often each letter appears, rank them, then rank the symbols of the ciphertext and line the two lists up. Adjust where the fit is poor. That is, almost word for word, the advice in every modern solving guide, including this one.

The reason it endures is that it attacks the right thing. A simple substitution cipher has an astronomically large number of possible keys, and none of that matters, because the cipher leaves the statistical fingerprint of the language completely intact. Al-Kindi understood in the ninth century something that took cipher designers several more centuries to accept: security does not come from the size of the keyspace. It comes from destroying the structure an attacker can measure.

For a solver, that is good news. The structure is still there, in every puzzle, waiting to be counted.

Related Reading

Put your skills to the test

Ready to apply your cryptanalysis techniques? Try solving today's cipher or explore the fascinating origins of common phrases.