Instrument 04Language & limits

The opening letters

الحروف المقطعة

Measure the disconnected openings, publish the claims that fail beside those that hold, and stop before measurement becomes decoding.

29sura openings

Letters instrument

01 · asks

The question

What can actually be known about the Quran's unexplained letter openings?

02 · reads

The evidence

Inventory, frequency against a control, orthography, what immediately follows, rhyme, and semantic clustering.

03 · stops

The boundary

No statistic on this page decodes the letters. Their meaning remains unsettled, and that unresolved boundary is part of the result.

The inventory

Exactly 14 of the alphabet's 28 letters are ever used this way: half of it, lit below. 26 of the 29 suras are Meccan and 3 Medinan. In 20 places the letters are counted as a verse of their own; elsewhere they open a longer first verse.

ابتثجحخدذرزسشصضطظعغفقكلمنهوي

The combinations

حمHa (ha') Mim404143444546
المAlif Lam Mim2329303132
الرAlif Lam Ra1011121415
طسمTa Sin Mim2628
صSad38
قQaf50
نNun68
طهTa Ha (haa')20
طسTa Sin27
يسYa Sin36
المصAlif Lam Mim Sad7
المرAlif Lam Mim Ra13
كهيعصKaf Ha (haa') Ya Ayn Sad19
حمعسقHa (ha') Mim Ayn Sin Qaf42

Runs

Four stretches of consecutive suras share an opening. The longest is the best known: seven suras in a row, 40 to 46, all beginning Ha Mim.

  • ال Alif Lam: suras 10 to 15, 6 in a rowالرالرالرالمرالرالر
  • طس Ta Sin: suras 26 to 28, 3 in a rowطسمطسطسم
  • ال Alif Lam: suras 29 to 32, 4 in a rowالمالمالمالم
  • حم Ha (ha') Mim: suras 40 to 46, 7 in a rowحمحمحمعسقحمحمحمحم

A run is counted on the first two letters, which is why Ash-Shura (42) belongs to the Ha Mim seven even though it alone continues, in a second verse, with Ayn Sin Qaf.

What follows them

The oldest observation about the letters is not statistical. It is that what comes immediately after them is, nearly always, the Book speaking about itself. That is checkable, so we checked it.

The rule

The running text after the letters is searched for five terms: kitab, qur'an, ayat, a form of nazzala, and adh-dhikr. Prefix matching on normalised text, so al-kitab and kitab both count. The window is one verse-unit of running text after the letters end: the rest of verse 1 where the letters open a longer verse, otherwise the first verse past them, which is verse 2 in every sura but Ash-Shura, whose letters occupy two verses. A second count widens it by one more verse. Misses are listed with the hits.

23 of 29 mention the Book, its signs, or its sending down in the very next verse-unit. Widening the window by a single verse makes it 24.

The 5 that still do not open elsewhere: Maryam on the mercy of a Lord to His servant, Al-Ankabut on a question about being tried, Ar-Rum on news of a battle, Al-Qalam on an oath by the pen. Ash-Shura deserves naming as a near miss rather than a miss: the verse after its letters does speak of revelation, but with the verb awha, which is not one of the five terms this rule searches for. Widening the term list would move it, which is exactly why the term list is printed above rather than left implicit.

Yunus 10:1the letters open a longer verse

الٓر تِلْكَ ءَايَٰتُ ٱلْكِتَٰبِ ٱلْحَكِيمِ

Alif, Lam, Ra. These are the verses of the wise Book

matched: al-kitab (the Book), ayat (signs, verses)

Al-Baqara 2:2the letters stand alone, and the sentence begins in verse 2

ذَٰلِكَ ٱلْكِتَٰبُ لَا رَيْبَ فِيهِ هُدًى لِّلْمُتَّقِينَ

This is the Book about which there is no doubt, a guidance for those conscious of Allah -

matched: al-kitab (the Book)

Al-Qalam 68:1one of the six with no mention in the immediate window

نٓ وَٱلْقَلَمِ وَمَا يَسْطُرُونَ

Nun. By the pen and what they inscribe,

no match

The frequency claim, tested

The claim that circulates most widely online is that each sura's opening letters occur unusually often inside that sura. It is a testable claim, and it is the kind of claim that survives mainly because no one runs the control.

The rule

For each sura that opens with letters, every one of the 28 Arabic letters is scored by its share of that sura's letters divided by its share of the whole book, then ranked. Rank 1 is the most over-represented letter in the sura. If the opening letters were unusually frequent, their ranks would cluster near 1. Pure chance puts them at 14.5.

12.68
mean rank of the opening letters in their own sura
14.5
what pure chance gives
14.15
the same letter sets in the 85 suras that never received them

The verdict

Barely there. The opening letters do lean high in their own suras, by 1.82 of a rank off chance, but the same sets lean high by 0.35 in suras that never received them, so most of the effect is a fact about Arabic rather than about the muqatta'at. The plain ratio version of the claim fares worse still: the opening letters average x1.041 of their book-wide share, and the control averages x1.057, which is higher. And in only 2 of the 29 suras is the single most over-represented letter one of that sura's own.

What the test does turn up is a handful of genuinely striking individual cases, which is not the same thing as a pattern:

قQaf in Qaaf
3.78% of the sura's letters against 2.14% of the book's: x1.77, rank 2 of 28
طTa in Ash-Shu'araa
0.59% of the sura's letters against 0.39% of the book's: x1.52, rank 1 of 28
صSad in Saad
0.95% of the sura's letters against 0.63% of the book's: x1.51, rank 3 of 28
طTa in An-Naml
0.57% of the sura's letters against 0.39% of the book's: x1.46, rank 1 of 28

And a matching handful run the other way. Ha ranks 23 of 28 for enrichment in Ad-Dukhan, a sura that opens with it, and Qaf ranks 25 in Ash-Shura, which also opens with it. Both lists are in the data; publishing only the first would be how this claim got its reputation.

Is it a code?

The question people arrive with. Another script, another language, a key that classifies the sura it opens. Three parts of it can be answered with the text and the measurements already on this site, and one part cannot be answered here at all.

The script is the same script

They are ordinary Arabic letters at ordinary Arabic codepoints, the same ones the rest of the book is written in. There is no second alphabet hiding in the file. What the vocalised text does add is a single mark, the maddah, and it is not sprinkled at random:

س ص ع ق ك ل م ن
8 letters that carry it every single time
ا ح ر ط ه ي
6 that never carry it

Across all 29 openings there is not one exception. And the split is not arbitrary: it is exactly the recitation rule. The 8 marked letters are the ones whose spelled names run to three letters and take the long elongation. The 6 unmarked are alif and the five whose names are two letters long. So the text does encode something about these letters, and what it encodes is how to say them.

Worth being exact about what that is and is not. These marks belong to the later vocalisation layer, not to the earliest skeletal script, which wrote these letters plain like everything else. They are a record of how the letters were recited, carried in the orthography. They are not a cipher, and nothing about them is concealed.

Does a shared opening mean a shared subject?

The strongest form of the idea: if Alif Lam Mim were a label for a kind of content, the six suras carrying it should resemble each other more than lettered suras do in general.

The rule

Each sura is the mean of its verse embeddings, normalised, with the letter verses themselves excluded. A family's score is the mean cosine over every pair of suras that share an opening. It is compared against three baselines: every pair among the 29 lettered suras, every pair among 29 suras with no letters matched to them one for one on revelation period and verse count, and every pair among all 114. All pairs, no sampling, no RNG.

حم family, 6 suras0.9918
الم family, 6 suras0.994
الر family, 5 suras0.9879
طسم family, 2 suras0.9868
all 29 lettered suras0.9899
29 suras with no letters, matched on period and length0.9812
all 114 suras0.9695

The verdict

Something real, but not the thing the question asks about. Suras that carry letters do resemble one another more than 29 suras matched to them on period and length that carry none, 0.9899 against 0.9812. But knowing WHICH letters adds almost nothing on top of that: two families sit slightly above the lettered baseline and two sit below it. The letters mark out a group of suras that have something in common. They do not sort that group into kinds.

Does the opening echo the sura's rhyme?

A sound-level version of the same idea, tested against the prosody measured for the Rhythm: 9 of the 27 testable suras rhyme on a consonant that is one of their own opening letters, against 0.177 for the same openings scored against suras that never received them. On its face that is a p of 0.0374.

The verdict

Do not believe it, and here is why, which is more useful than the p. Every one of the 9 hits is Mim or Nun. Those are the two commonest rhyme consonants in the book, carrying 60 of the 88 suras with a consonantal rhyme between them, and Mim alone appears in 17 of the 29 openings. Two common things coincide at the rate two common things coincide. This is also the third test on this page, and a p of 0.0374 does not survive being the third test.

The hits, for anyone who wants to check: 3, 30, 31, 40, 41, 42, 45, 46, 68.

What this site cannot test

Whether the letters come from another language or an older scribal practice. Proposals exist and some are serious: abbreviations of words or names, a conjecture that they were initials of the scribes who held the early codices (its own author later left it), derivations from Syriac or Hebrew usage. None has manuscript evidence or an internal argument that has persuaded the field, and testing any of them needs comparative corpora this site does not have. They are reported here as positions people hold, which is the same footing as the section below. A page that measured what it could and then guessed at the rest would have wasted the measuring.

What the clustering found

The one genuinely startling result was not looked for. Every verse in the book was embedded by meaning and grouped by an algorithm that was never told these verses were unusual, never given the list, and never shown a sura boundary. It sorted the book into 120 groups, and one of them came out like this:

20 / 20
of its members are letter verses
20 / 20
of the letter verses are in it
0.9639
internal cohesion, among the highest in the book

Both directions are exact. The group is the letter verses, and the letter verses are the group. It also caught the split in Ash-Shura, placing 42:2 with the rest without being told the two halves belonged together.

The honest reading of this is narrow, and worth saying plainly: it shows the letters are unlike everything else in the book, in a way a machine measuring meaning can detect. It does not show what they mean. A group of verses that share a form will cluster on that form; that is what clustering does. The finding is the sharpness of the boundary, not a decoding. You can see them in the Universe as a knot of stars sitting apart from the whole cloud.

What the commentators hold

Positions, not conclusions. These are held by people of knowledge and reported here as their positions; the site does not adjudicate between them.

  • Their meaning belongs to God. The oldest and most widely held position, reported from several of the Companions: the letters are among what God alone knows, and the reciter says them as they were given.
  • They are a challenge. The book is made of these same letters, which every Arab already possessed, and it is placed beside a standing invitation to produce its like. On this reading the letters point at the raw material and let the contrast speak.
  • They are names. Of the sura, or of the Quran itself, or abbreviations of divine names. Two suras are known to this day by their letters alone: Ta Ha and Ya Sin.
  • They arrest attention. A sound with no meaning, before a text with nothing but meaning, at the moment recitation begins.

Honesty note

Every figure on this page is computed by scripts/build-letters.mjs from the Tanzil text this site ships, and the script is deterministic: the same text gives the same numbers. Letter counts fold the hamza carriers into plain alif, final ya into ya, and ta marbuta into ha, or every alif statistic would be wrong by the rate of hamzated words. The maddah test is the one place the vocalised text is read instead of the stripped one, because the stripped text has no diacritics to find: running it on the wrong corpus reported all fourteen letters as unmarked, a tidy null that was purely an artefact. Of the seven things measured here, three came out negative and are published as such. No numerological claim about multiples is made, tested, or endorsed on this page. The letters remain undeciphered, and nothing here should be read as decoding them.