The question
What can actually be known about the Quran's unexplained letter openings?
STARTQURANالقرآن الكريمInstrument 04Language & limits
الحروف المقطعة
Measure the disconnected openings, publish the claims that fail beside those that hold, and stop before measurement becomes decoding.
Letters instrument
What can actually be known about the Quran's unexplained letter openings?
Inventory, frequency against a control, orthography, what immediately follows, rhyme, and semantic clustering.
No statistic on this page decodes the letters. Their meaning remains unsettled, and that unresolved boundary is part of the result.
Exactly 14 of the alphabet's 28 letters are ever used this way: half of it, lit below. 26 of the 29 suras are Meccan and 3 Medinan. In 20 places the letters are counted as a verse of their own; elsewhere they open a longer first verse.
Four stretches of consecutive suras share an opening. The longest is the best known: seven suras in a row, 40 to 46, all beginning Ha Mim.
A run is counted on the first two letters, which is why Ash-Shura (42) belongs to the Ha Mim seven even though it alone continues, in a second verse, with Ayn Sin Qaf.
The oldest observation about the letters is not statistical. It is that what comes immediately after them is, nearly always, the Book speaking about itself. That is checkable, so we checked it.
The running text after the letters is searched for five terms: kitab, qur'an, ayat, a form of nazzala, and adh-dhikr. Prefix matching on normalised text, so al-kitab and kitab both count. The window is one verse-unit of running text after the letters end: the rest of verse 1 where the letters open a longer verse, otherwise the first verse past them, which is verse 2 in every sura but Ash-Shura, whose letters occupy two verses. A second count widens it by one more verse. Misses are listed with the hits.
23 of 29 mention the Book, its signs, or its sending down in the very next verse-unit. Widening the window by a single verse makes it 24.
The 5 that still do not open elsewhere: Maryam on the mercy of a Lord to His servant, Al-Ankabut on a question about being tried, Ar-Rum on news of a battle, Al-Qalam on an oath by the pen. Ash-Shura deserves naming as a near miss rather than a miss: the verse after its letters does speak of revelation, but with the verb awha, which is not one of the five terms this rule searches for. Widening the term list would move it, which is exactly why the term list is printed above rather than left implicit.
الٓر تِلْكَ ءَايَٰتُ ٱلْكِتَٰبِ ٱلْحَكِيمِ
Alif, Lam, Ra. These are the verses of the wise Book
matched: al-kitab (the Book), ayat (signs, verses)
ذَٰلِكَ ٱلْكِتَٰبُ لَا رَيْبَ فِيهِ هُدًى لِّلْمُتَّقِينَ
This is the Book about which there is no doubt, a guidance for those conscious of Allah -
matched: al-kitab (the Book)
نٓ وَٱلْقَلَمِ وَمَا يَسْطُرُونَ
Nun. By the pen and what they inscribe,
no match
The claim that circulates most widely online is that each sura's opening letters occur unusually often inside that sura. It is a testable claim, and it is the kind of claim that survives mainly because no one runs the control.
For each sura that opens with letters, every one of the 28 Arabic letters is scored by its share of that sura's letters divided by its share of the whole book, then ranked. Rank 1 is the most over-represented letter in the sura. If the opening letters were unusually frequent, their ranks would cluster near 1. Pure chance puts them at 14.5.
Barely there. The opening letters do lean high in their own suras, by 1.82 of a rank off chance, but the same sets lean high by 0.35 in suras that never received them, so most of the effect is a fact about Arabic rather than about the muqatta'at. The plain ratio version of the claim fares worse still: the opening letters average x1.041 of their book-wide share, and the control averages x1.057, which is higher. And in only 2 of the 29 suras is the single most over-represented letter one of that sura's own.
What the test does turn up is a handful of genuinely striking individual cases, which is not the same thing as a pattern:
And a matching handful run the other way. Ha ranks 23 of 28 for enrichment in Ad-Dukhan, a sura that opens with it, and Qaf ranks 25 in Ash-Shura, which also opens with it. Both lists are in the data; publishing only the first would be how this claim got its reputation.
The question people arrive with. Another script, another language, a key that classifies the sura it opens. Three parts of it can be answered with the text and the measurements already on this site, and one part cannot be answered here at all.
They are ordinary Arabic letters at ordinary Arabic codepoints, the same ones the rest of the book is written in. There is no second alphabet hiding in the file. What the vocalised text does add is a single mark, the maddah, and it is not sprinkled at random:
Across all 29 openings there is not one exception. And the split is not arbitrary: it is exactly the recitation rule. The 8 marked letters are the ones whose spelled names run to three letters and take the long elongation. The 6 unmarked are alif and the five whose names are two letters long. So the text does encode something about these letters, and what it encodes is how to say them.
Worth being exact about what that is and is not. These marks belong to the later vocalisation layer, not to the earliest skeletal script, which wrote these letters plain like everything else. They are a record of how the letters were recited, carried in the orthography. They are not a cipher, and nothing about them is concealed.
The strongest form of the idea: if Alif Lam Mim were a label for a kind of content, the six suras carrying it should resemble each other more than lettered suras do in general.
Each sura is the mean of its verse embeddings, normalised, with the letter verses themselves excluded. A family's score is the mean cosine over every pair of suras that share an opening. It is compared against three baselines: every pair among the 29 lettered suras, every pair among 29 suras with no letters matched to them one for one on revelation period and verse count, and every pair among all 114. All pairs, no sampling, no RNG.
Something real, but not the thing the question asks about. Suras that carry letters do resemble one another more than 29 suras matched to them on period and length that carry none, 0.9899 against 0.9812. But knowing WHICH letters adds almost nothing on top of that: two families sit slightly above the lettered baseline and two sit below it. The letters mark out a group of suras that have something in common. They do not sort that group into kinds.
A sound-level version of the same idea, tested against the prosody measured for the Rhythm: 9 of the 27 testable suras rhyme on a consonant that is one of their own opening letters, against 0.177 for the same openings scored against suras that never received them. On its face that is a p of 0.0374.
Do not believe it, and here is why, which is more useful than the p. Every one of the 9 hits is Mim or Nun. Those are the two commonest rhyme consonants in the book, carrying 60 of the 88 suras with a consonantal rhyme between them, and Mim alone appears in 17 of the 29 openings. Two common things coincide at the rate two common things coincide. This is also the third test on this page, and a p of 0.0374 does not survive being the third test.
The hits, for anyone who wants to check: 3, 30, 31, 40, 41, 42, 45, 46, 68.
Whether the letters come from another language or an older scribal practice. Proposals exist and some are serious: abbreviations of words or names, a conjecture that they were initials of the scribes who held the early codices (its own author later left it), derivations from Syriac or Hebrew usage. None has manuscript evidence or an internal argument that has persuaded the field, and testing any of them needs comparative corpora this site does not have. They are reported here as positions people hold, which is the same footing as the section below. A page that measured what it could and then guessed at the rest would have wasted the measuring.
The one genuinely startling result was not looked for. Every verse in the book was embedded by meaning and grouped by an algorithm that was never told these verses were unusual, never given the list, and never shown a sura boundary. It sorted the book into 120 groups, and one of them came out like this:
Both directions are exact. The group is the letter verses, and the letter verses are the group. It also caught the split in Ash-Shura, placing 42:2 with the rest without being told the two halves belonged together.
The honest reading of this is narrow, and worth saying plainly: it shows the letters are unlike everything else in the book, in a way a machine measuring meaning can detect. It does not show what they mean. A group of verses that share a form will cluster on that form; that is what clustering does. The finding is the sharpness of the boundary, not a decoding. You can see them in the Universe as a knot of stars sitting apart from the whole cloud.
Positions, not conclusions. These are held by people of knowledge and reported here as their positions; the site does not adjudicate between them.
Every figure on this page is computed by scripts/build-letters.mjs from the Tanzil text this site ships, and the script is deterministic: the same text gives the same numbers. Letter counts fold the hamza carriers into plain alif, final ya into ya, and ta marbuta into ha, or every alif statistic would be wrong by the rate of hamzated words. The maddah test is the one place the vocalised text is read instead of the stripped one, because the stripped text has no diacritics to find: running it on the wrong corpus reported all fourteen letters as unmarked, a tidy null that was purely an artefact. Of the seven things measured here, three came out negative and are published as such. No numerological claim about multiples is made, tested, or endorsed on this page. The letters remain undeciphered, and nothing here should be read as decoding them.