Indus Valley Script · structural analysis01 August 2026
Undeciphered Scripts · c. 2600–1900 BCE
Short Texts, Not Absent Structure
Linear B, cut to Indus dimensions, reproduces the statistical profile
that has been used to argue the Indus script is not writing
Verified findings · updated August 2026
Indus script signs show strong positional grammar
Analysis of 2,536 inscriptions reveals that specific signs overwhelmingly occupy the opening and closing positions of the text. One sign appears first in 69 percent of occurrences, while others reach up to 89 percent. Similarly, a distinct set of signs dominates the closing position, with frequencies up to 88 percent. This pattern holds regardless of the sign's identity, indicating a fixed structural order within the script.
The inscriptions are short, with a mean of 4.4 signs, which limits the depth of structural analysis. However, the positional preferences are statistically significant and robust against random baselines. This finding confirms a consistent grammatical framework in the Indus script, offering a clearer understanding of its organization without assigning meaning to individual signs.
How confident: The positional patterns are highly significant, with effect sizes reaching z-scores of up to 38, confirming a strong and reliable structural regularity in the script.
These are structural findings, not decipherments. No sign has been assigned a sound or meaning; these scripts remain undeciphered.
2536 inscriptions11,135 sign tokens591 signs01 August 2026
A stamp seal and its modern impression. The row of signs along the top is the
script; the animal below it is the motif. Our measurements show the two channels carry
almost no information about each other. Public domain, released under CC0.
The Indus script has resisted decipherment for a century. A recurring argument holds that
it never encoded language at all: its texts are too short, its dependencies too local, its
structure too thin. That argument has never been tested against the obvious control, which is
a script we can already read, cut to the same dimensions.
We did that. Linear B, deciphered in 1952 and known to encode Greek, was subsampled to the
Indus token count and length distribution and run through the identical procedure, alongside
proto-cuneiform, which is attested notation that was not yet writing. On the
information carried purely by sign order, Indus scores 0.901 bits and Linear B 0.913.
Proto-cuneiform manages 0.405. Both Indus and Linear B lose almost all dependency
beyond adjacent signs, which is the observation most often cited as evidence against Indus
being writing.
The absence of long-range structure in the Indus corpus is therefore a consequence of
short texts, not a property of the script. This does not show that the Indus script is
writing, and nothing here reads a single sign. It shows that the strongest statistical
argument that it is not does not survive its own control.
Guided tour · 11 minutes · read by the author
Start here. The page scrolls itself as it goes.
14 stops
IThe People Who Wrote It
Between roughly 2600 and 1900 BCE, while the pyramids at Giza were going up and the cities
of Sumer were arguing over irrigation, the largest civilization of the Bronze Age occupied what
is now Pakistan and northwest India. It covered more than 1.25 million square kilometres, an
area comparable to Mesopotamia and Egypt put together, and held perhaps five million people
across more than a thousand known settlements. Almost nobody outside archaeology can name it.
It was forgotten so completely that it had to be rediscovered in the 1920s, and we still do not
know what its people called themselves.
Their cities were planned before they were built. Mohenjo-daro and Harappa were laid out on
street grids, with fired brick housing, private wells, bathing platforms, and covered drains
running beneath the streets to carry waste out of the residential blocks. Domestic sanitation
at that standard would not be seen again in Europe for well over three thousand years.
Dholavira had reservoirs cut into rock. Lothal had a basin most archaeologists read as a
working dock. None of it looks improvised.
The startling thing is not any single building. It is the agreement. Harappan merchants
weighed goods with cubes of chert cut to a binary series, each step doubling the one before it.
The smallest is about 0.856 grams; the commonest, sixteen steps along, is about 13.7 grams. The
same series turns up at sites hundreds of kilometres apart, and the spread between individual
weights runs to only a few per cent, which looks like the limit of what their stoneworking
could hold rather than a limit of what they were aiming at. Their fired bricks were made to
consistent proportions across the same range. A trader could carry a weight from one city and
have it balance, correctly, against a weight cut in another one eight hundred kilometres
away.
The hard part was never the stonecutting. It was the agreement. Holding a
measurement standard synchronized across a territory that size, for something like six
centuries, with no telegraph and no post, is an information problem before it is a technical
one. Every merchant, brickmaker and inspector had to be working from the same convention, and
the convention had to survive being copied by hand, generation after generation, across a
thousand settlements. Modern states need bureaux and legislation to manage this. We do not know
how the Indus did it, and we have not found the palace or the king that would explain it.
Which is the next strange thing. No royal tombs have been identified, no building can be
called a palace with confidence, and the art does not celebrate rulers or conquest, which sets
the Indus apart from every other Bronze Age state we can read about. How much weight that
absence can carry is actively argued, and it is treated far too casually in popular accounts of
this civilization. But the contrast is real, and it is part of why the script matters: the
texts are the one channel that might say who was in charge, and we cannot read them.
They did write. Roughly four thousand inscribed objects survive, mostly stamp seals in
steatite, along with pottery, copper tablets and moulded terracotta. The corpus measured on
this page holds 2,536 inscriptions and 11,135 signs in total, which is about the length of a
magazine article. Three quarters of the texts run to five signs or fewer. The longest is
seventeen. Whatever the script was for, it was not for literature, and nearly everything in
this chapter comes from what these people built and buried rather than from anything they
wrote, because nobody can read a word of it.
One question gets asked often enough to be worth answering with evidence rather than
sentiment, which is whether a system this consistent was really theirs. It was. The script
carries the fingerprints of individual human hands. Inscriptions sharing long runs of signs
cluster by site, 59.1 per cent of them from the same place against 36.5 per cent expected by
chance, and when that same measurement is calibrated on Linear B the families it finds track
individual scribes rather than document types. Sign shapes drift into variants the way
handwriting does. That is apprentices copying masters in regional workshops. A system handed
down ready made, from anywhere, would be uniform everywhere, and it would not have accents.
By about 1900 BCE the cities were being left. The reasons are still argued and were probably
several at once, with a weakening monsoon and shifting river courses prominent among them. The
script went with them. It was in use for something like seven centuries, by people who could
hold a standard across a subcontinent, and then it stopped, and no one has read it since. What
follows is an attempt to measure what kind of thing it was, without pretending to know what it
says.
IIThe Control That Was Missing
Every published structural analysis of the Indus corpus, including our own until now, has
compared it against systems chosen to be non-linguistic: heraldry, ownership marks, tallies,
invented baselines. None has compared it against a known writing system measured at the same
size. Without that, a result showing Indus to be statistically impoverished says nothing,
because nobody had established what real writing looks like when you only ever get four
signs at a time.
Three corpora, all attested, all openly published. Every one cut to the Indus token count
and the Indus length distribution, sampled without replacement and averaged over repeated
draws.
Measure
Indus
Linear B
Proto-cuneiform
texts / tokens
1,902 / 8,839
1,893 / 8,841
1,907 / 8,710
mean length
4.65
4.67
4.57
information in sign order alone
0.901
0.913
0.405
dependency at distance 1
+0.901
+0.912
+0.404
dependency at distance 2
+0.077
+0.095
−0.118
dependency at distance 3
−0.106
−0.054
−0.421
recurring three-sign blocks
0.235
0.330
0.046
vocabulary growth
0.488
0.294
0.431
signs occurring once
0.364
0.193
0.210
Table 1. Three attested corpora at identical dimensions. Linear B is deciphered
administrative writing in a known language. Proto-cuneiform is attested notation that
records commodities and quantities without encoding speech.
The two highlighted rows carry the argument. On information carried by sign order
alone, Indus and Linear B are within 1.3 per cent of each other. And both collapse
to almost nothing at distance two, which is precisely the observation that has been offered
as evidence that Indus cannot be language. A script we have read for seventy years does the
same thing at this text length.
Proto-cuneiform, which really is a notation rather than a script, sits at less than half
the Indus figure and turns negative at distance two. Whatever the Indus script is, it is not
behaving like that.
What this does and does not show
It does not show that the Indus script encodes language. It shows that one specific
argument that it does not, the argument from missing long-range structure, fails its own
control. That is a narrower claim and it is the one the data supports.
One limitation matters and is not hidden here. The arms are matched on text count,
length and token count, but they cannot be matched on inventory size: Indus has 591 sign
types where Linear B offers 211, because Linear B does not contain 591 signs. A larger
inventory means sparser data and more estimator bias, which plausibly explains part of the
remaining gap in the total figures. It does not explain the ordering result or the decay
pattern, which are the two rows the argument rests on.
IIIThe Corpus
Every figure on this page comes from excavated inscriptions in the published
digitisation of the Corpus of Indus Seals and Inscriptions. No reconstructed,
simulated or generated text enters the analysis at any point. This is worth stating
plainly because the alternative is common in this field and hard to detect from outside.
Inscriptions
2536
1902 after collapsing duplicate seal impressions
Sign tokens
11,135
mean 4.39 signs per inscription, maximum 17
Distinct signs
591
34% occur exactly once
Sites
10
Mohenjo-daro, Harappa, Lothal, Dholavira, Kalibangan and others
Table 2. Corpus composition.
IVThe Finding
At this corpus size every measure of structure is inflated. With 591 distinct signs and
fewer than nine thousand tokens, the table of sign pairs is about two per cent occupied,
and in that regime a corpus shuffled into complete meaninglessness still returns a large
positive score. A raw figure therefore means nothing. What follows is in every case the
margin over a matched control: the same corpus, its structure destroyed, measured
the same way.
Dependency that survives the control
33%
Adjacent signs in this corpus measure 3.169 bits of mutual
information. The same signs re-dealt at random measure 2.136. Two thirds of the raw figure
is sampling bias; the remaining 1.033 bits is the finding. Matched English prose
returns +0.579.
Measure
Observed
Control
Excess
z
Sequential dependency
order carries information
3.169
2.136
1.033
71.5
Order alone
controlling for co-occurrence
3.169
2.269
0.900
64.9
Boundary closure
negative = inscriptions are closed
2.226
3.089
-0.863
-18.9
Directionality
opening vs closing sign sets
0.831
0.206
0.625
61.7
Positional fixity
signs hold their region
0.660
0.478
0.182
35.1
Table 3. Structural measures on the deduplicated corpus, each against its
matched control. Boundary closure is negative by construction of the test: sign pairs
spanning two different objects share less information than chance.
The second row is the one that rules out a weaker reading. Scrambling the order of signs
within an inscription, while keeping exactly which signs co-occur, still destroys
0.900 bits. The information is in the sequence, not merely in the company.
The boundary result is the cleanest in the study. Whatever an inscription says, it
finishes saying it: dependency across the join between two artifacts falls below
chance. Each object carries one complete statement, and the corpus is not the scattered
fragments of a longer running text.
VInscription Families
A separate question, and one that does not depend on any of the entropy machinery: do
inscriptions repeat each other? A workshop, a single scribe, or a fixed formula should
leave relatives behind, sharing long runs of signs more often than chance permits.
Shared run
Pairs observed
Expected by chance
Effect size
three or more signs
4,064
147
h = 0.08
four or more signs
467
1
h = 0.03
Table 4. Inscription pairs sharing an identical run of signs, against a control
preserving sign frequencies and inscription lengths. Effect size is Cohen's h on the rate.
Four hundred and sixty-seven pairs of inscriptions share a run of four or more identical
signs, where the control predicts one. That is a real relationship and not a marginal one.
The effect size says the other half of it. Those 467 pairs are three hundredths of one
per cent of all possible pairs, so on the rate the effect is negligible. The corpus
contains a small number of genuine families, not a corpus-wide system of formulae.
Both halves of that sentence come from the same measurement, and reporting only the first
would misrepresent it.
The longest shared runs reach seven signs, on pairs of inscriptions that are otherwise
unrelated objects. Those are the best candidates in the corpus for two artifacts carrying
the same text, and they are a natural target for anyone with access to the originals.
The families are local
If these families are the work of individual hands, they should betray it physically. A
single workshop stands in one place, draws on one stock of material and makes one kind of
object, so two inscriptions from the same hand ought to agree about where they were found
far more often than two inscriptions picked at random. If instead the shared runs are a
standard formula circulating across the whole civilisation, they should be scattered.
They are not scattered. Pairs sharing a run of four signs come from the same site
59.1% of the time against 36.5% expected by chance, and that holds after
near-identical inscriptions are collapsed at edit distance two. Agreement on object type
and material runs the same way but weakens once pairs from different sites are compared
directly, which says the object and the material are largely following the geography rather
than adding to it. Whatever these repeated sequences are, they are local.
That is what a workshop looks like and it is not what a civilisation-wide formula looks
like.
The position of the repeated run points the same way. When the same four signs turn up in
two inscriptions, they sit at the same place in both 59.3% of the time, against 37.0% expected,
a ratio of 1.60. A fixed formula occupies a fixed slot, so that ratio is a measure of how
formulaic a script is. The comparison is the informative part: Linear B reaches 2.10 and Ur III
Sumerian 3.43 on the same measurement. Indus repeated sequences are the least
positionally fixed of the writing systems tested here. Whatever recurs in this
corpus recurs more freely than the recurring material in two readable administrative
archives, which is the opposite of what the objection that these are rote formulae predicts.
The comparison corpora were size-matched but not near-deduplicated, so this is
suggestive rather than settled.
Why this is reported with two numbers
Significance and effect size answer different questions, and at this corpus size they
disagree constantly. The four-sign result carries a z of 375, which is enormous, next to
an effect size that is negligible. Quoting either alone would be misleading. Every
headline measure on this page has been rescored the same way, and the two strongest of
them turn out to be medium effects rather than large ones.
VIThe Controls
A structural result is uninterpretable until you know what other systems score. The
identical analysis was therefore run on four control corpora, each matched to the Indus
corpus on text count, length distribution, token count and vocabulary size. Without that
matching one measures corpus size rather than writing system, which is the error
that makes cross-corpus comparisons in this field unreliable.
Corpus
Sequential
Boundary
Positional
Directional
Indus script
the real corpus
+1.033
-0.863
+0.181
+0.624
Natural language
real text, matched
+0.579
+0.304
+0.000
-0.008
Administrative records
modelled
+0.537
-1.965
+0.354
+0.758
Ownership marks
modelled
-0.468
-1.433
+0.221
+0.766
Random signs
structure floor
-0.007
+0.083
-0.007
-0.006
Table 5. Matched controls. Read the last row first: randomly ordered signs score
zero on everything, which is the check that the method is calibrated and not manufacturing
structure.
What separates Indus
Sequential dependency, and only sequential dependency. Indus scores above every
control including real natural language, while the ownership-mark model scores
negative: emblems carry no information in their ordering. This is a
genuine constraint on any account that treats the script as a set of marks.
What does not
Directionality and positional fixity. Both are strong in Indus and stronger still in
the administrative and emblem models. They establish that inscriptions are ordered and
read in a fixed direction; they cannot tell writing from record-keeping. Any argument
resting on them alone is unsafe. That includes ours, until this test was run.
Where that leaves the question
Indus sits cleanly with neither family. It is high on sequential dependency like a
language, discrete at its boundaries like an administrative record, and positionally rigid
like both. On the profile as a whole it is intermediate, which is independently the
same conclusion reached by a separate 2026 analysis of this corpus by different means.
The pure emblem hypothesis is the one that struggles. It predicts no sequential
dependency, and the Indus corpus has more of it than English prose.
An attack on the central result, and what happened to it
The sharpest objection this work has faced is that it is one slot deep. The Indus opening
is unusually restricted, with five signs covering well over half of all first positions. If
that single restricted position is generating the sequential dependency, then the whole
resemblance to Greek and Sumerian is an artifact of one slot and the central result should
be withdrawn. The objection is testable in the bluntest possible way: delete the opening
sign from every inscription in all five corpora and measure again.
Corpus
Intact
Opening removed
Closing removed
Random interior removed
Indus
0.900
0.727
0.861
0.454
Linear B
0.922
0.809
0.734
0.461
Ur III Sumerian
0.908
0.741
0.974
0.435
Proto-Elamite
0.769
0.824
0.591
0.419
Proto-cuneiform
0.417
0.201
0.409
0.196
Table 6. Information carried by sign order after ablating one position. Every
corpus is cut to the same token count and length distribution. Removing any sign costs
information simply by shortening the text, which is what the last column measures.
The objection does not hold. Removing the opening costs Indus 0.173, while
removing an arbitrary interior sign costs 0.446, more than twice as much. The restricted
opening is not where the dependency lives. More to the point, the readable scripts behave
the same way: Linear B and Ur III Sumerian both lose far more from an interior deletion than
from an edge one, in the same proportion. If the Indus figure were an artifact of its
opening, Indus would have separated from them under this test. It did not.
One corpus does break the pattern, and it breaks it usefully. Proto-Elamite gains
information when its opening is removed, because proto-Elamite is the one system in the panel
that locks its final position rather than its first. That is an independent confirmation the
measurement is sensitive to where a script's constraints actually sit, rather than returning
the same shape regardless.
This test was not our idea. It came out of an adversarial pass in which the machinery was
asked, repeatedly and from deliberately unrelated disciplines, for the computation most
likely to destroy something we currently believe.
A limitation we have not solved
The natural-language control is running prose cut into short segments, so it has no
genuine text beginnings or endings. That makes it a poor comparator for directionality in
particular, and it is why no claim is made that Indus directionality is language-like,
only that it is real. A control built from genuinely short complete texts would settle it.
We have not built one.
VIIWhat Survived
The findings were attacked before they were published. The strongest objection was that
duplicate seal impressions, the same seal pressed many times or mass-produced copies
of one design, could manufacture all of these results from nothing. Collapsing every
repeated inscription removes 634 texts. Each measure was then recomputed on the reduced
corpus, and separately at each major site, so that no finding rests on pooling two
traditions that might differ.
Excess over control
All 2,536
Dedup.
Mohenjo-daro
Harappa
Sequential dependency
1.396
1.033
0.932
0.746
Order alone
1.073
0.900
0.837
0.692
Boundary closure
-0.848
-0.863
-1.210
-1.140
Directionality
0.687
0.625
0.605
0.554
Positional fixity
0.213
0.182
0.199
0.219
Grammatical categories (withdrawn, z)
2.61
1.47
1.56
2.35
Table 7. Every headline measure across corpus variants. The struck row is the
withdrawn result described below.
The effects persist after deduplication and hold independently at Mohenjo-daro and
Harappa, two sites some six hundred kilometres apart, showing the same structure.
How many independent inscriptions are there really?
Collapsing only identical inscriptions is the weak version of this control. Indus seals
were also copied, and near-identical texts are not independent observations either. Grouping
inscriptions that differ by one or two sign edits reduces 1,902 distinct sequences to 1,141
and then to 558. Intervals computed on the raw count are therefore too narrow, which is a
problem for most published statistics on this corpus, and was for ours.
Measure (excess, with z)
1,902 exact
1,141 at edit 1
558 at edit 2
Sequential dependency
+1.033 71.5
+0.875 54.5
+0.665 34.0
Directionality
+0.625 61.7
+0.561 45.5
+0.464 26.7
Positional fixity
+0.182 35.1
+0.151 24.8
+0.125 18.0
Opening vs closing classes
z 6.0
z 4.6
z 0.5
Table 8. Every headline measure recomputed on progressively stricter definitions
of an independent inscription. The struck row is the one that does not survive.
The three main results hold on 558 inscriptions, under a quarter of what we started with.
That is the number they should be judged on.
Is the headline number an artifact of the estimator?
This is the sharpest objection available and it has been made before, against earlier
entropy work on this corpus. At 11,135 tokens across 591 sign types the table of sign pairs
is 0.84 per cent occupied, 2,919 filled cells out of 349,281. In that regime
the naive maximum-likelihood estimator overstates dependency badly, and a figure quoted
without correction cannot be trusted.
Our figure is already a margin over a permutation control computed with the same
estimator, which cancels most of that bias because the control shares the sample size,
alphabet and marginal frequencies. It does not cancel it exactly. So the same quantity was
recomputed under three published bias-corrected estimators, each applied to the observed
data and to the control alike.
Estimator
Observed
Control
Excess
95% interval
Maximum likelihood
3.169
2.136
+1.033
+0.873 to +0.978
Miller-Madow
2.971
1.778
+1.192
+1.014 to +1.124
Chao-Shen
3.001
1.566
+1.435
+1.370 to +1.511
Chao-Wang-Jost
2.538
0.861
+1.677
+1.495 to +1.742
Table 9. The same excess under four estimators. Intervals come from
subsampling inscriptions without replacement, which is the correct resampling unit here:
drawing them with replacement reintroduces the duplicate structure the analysis exists to
control for.
Correction moves the result upward, not downward. The excess rises from
1.033 to 1.677 bits as the correction gets stronger, because the correction reduces the
control far more than it reduces the observed data. The shuffled corpus is dominated by
sign pairs seen once, which is exactly the situation these estimators exist to fix. The
uncorrected figure we publish is therefore the conservative one, and every estimator tested
puts the interval clear of zero.
A result we withdrew
An earlier pass found that signs sort into grammatical categories which then line up in
order along the inscription. It was the most interesting result of the project and the
closest thing to a grammar. It did not survive. Once duplicate impressions were collapsed
the effect fell to chance, and it failed to reproduce consistently across sites. The
duplicates had been manufacturing it.
It is recorded here because a finding that dies under its own control test is worth
more to other researchers than one that was never tested, and because anyone running
similar analyses on this corpus should expect the same trap.
What was left when we looked again
Withdrawing a result is not the same as showing there is nothing there. The original test asked one broad question (do a dozen induced categories arrange themselves in order?) and spent all its evidence on it. Three narrower tests were run afterwards on the deduplicated corpus, each able to fail independently.
Second-pass test
Result
Against
Cluster stability under resampling
0.266
0.049
Category model vs full sign model (bits, held-out)
+0.0284
t = 1.89
Opening vs closing sign contexts
z = 6.02
p < 0.001
Table 10. Second-pass tests, deduplicated corpus only.
The categories are more stable under resampling than chance, but not by much. Compressing 591 signs into a dozen categories costs nothing measurable on unseen text. The interval crosses zero, so the honest reading is no detectable loss, not an improvement.
The third test held up at first. The signs that open inscriptions and those that close them differ in the company they keep, and not merely in where they sit: the comparison uses each sign's neighbours with position information removed. Twenty-one openers and forty-eight closers form two distributionally distinct groups at z = 6.02.
It does not survive the harder duplicate control. Collapsing near-identical inscriptions as well as exact ones takes it to z = 4.6 at edit distance one, and to z = 0.5 at edit distance two, which is nothing at all. We are leaving it here as a reported result rather than a claim: it may be real and merely under-powered once the corpus is cut to genuinely independent inscriptions, or it may be another artifact of repeated seals. On this corpus we cannot tell.
VIIIWhat This Settles
Supported
Sign order carries information, and more of it than in running prose
Inscriptions are complete, self-contained statements
The script is directional, with distinct opening and closing positions
Signs hold fixed positions rather than floating freely
The same system is in use at sites hundreds of kilometres apart
Not established
What any inscription says
What language, if any, underlies the script
The phonetic or semantic value of any sign
Whether the script is logographic, syllabic or mixed
Any connection to a known language family
Whether signs group into a full system of grammatical categories though opening and closing signs do form two distinct classes
This is not a decipherment
No one has deciphered the Indus script, and this work does not either. Claims to have
read it appear regularly and are not supported by evidence of this kind; neither is
anything here. Not one inscription is translated. One sign opens roughly a third of the
corpus and we do not know what it means.
The strongest honest statement the data supports is this: the Indus script carries
ordered, self-contained, positionally organised structure, and more dependency between
successive signs than English prose. That is a real constraint, and the emblem and
ownership-mark readings have to answer it. It is not the same as showing the script
encodes a language, and we are not showing that.
IXThe Picture and the Object
Every Indus seal carries an image, most often a one-horned bull, sometimes an
elephant, a rhinoceros, a tree or a human figure, set above the row of signs. The
natural reading is that the signs caption the picture. If so, the two channels must share
information, and the image would become a weak bilingual: a known concept standing beside
an unknown word.
Channel
Observed
Control
Excess
z
Iconographic motif
0.2978
0.2814
+0.0164
2.0
Object type
0.2085
0.1437
+0.0648
14.0
Material
0.1889
0.1525
+0.0364
7.1
Table 11. Information shared between the sign string and three non-textual
channels, deduplicated corpus. Labels were permuted across objects to build each
control.
The caption reading is not supported. Before deduplication the motif
appears to carry substantial information about the signs; after repeated seal impressions
are collapsed, almost all of it disappears. What remains does not clear the bar once the
three channels tested here are accounted for. The picture and the text are, as far as this
measurement goes, independent channels. Whatever the signs record, it is not a
description of the animal above them.
The object itself is a different story. What a piece is, whether seal, tablet,
tag or potsherd, genuinely conditions which signs appear on it, and survives
deduplication comfortably. Text length tracks it too: seals average 5.06 signs, tablets 3.70,
pottery 2.78. The corpus is not one uniform practice, and analyses that pool it are averaging
over at least two.
Whether it could be a list of names
A standing objection to the whole enterprise is that these are ownership marks,
personal names, lineages and titles, rather than text. That hypothesis makes a sharp prediction.
A corpus of one-off identifiers never settles down: every new object introduces new signs,
and the vocabulary grows about as fast as the corpus itself. Reused language does the
opposite, because a finite lexicon is being drawn on repeatedly.
Corpus
Vocabulary growth
Reading
Indus script
0.488
finite vocabulary, reused
English prose (matched)
0.415
finite vocabulary, reused
Indus, seals only
0.506
Indus, tablets only
0.581
Table 12. Heaps' law exponent. Near 0.5, a finite vocabulary is being reused;
approaching 1.0, it is not.
Indus sits at 0.488, against 0.415 for matched English prose.
Vocabulary growth is slightly faster, consistent with more one-off signs, but nowhere
near the behaviour of a list of unique identifiers. The sign inventory is genuinely
reused. Whatever the signs are, they are drawn from a working vocabulary rather
than minted per object, and the pure name-list reading has to account for that.
Whether they are accounts
The third standing reading is that these are administrative records: so many measures of
grain, so many head of cattle. That is what the neighbouring systems are. It also makes a
prediction that can be tested without reading a single sign, because two of the comparison
corpora label their numerals explicitly, and a numeral does not behave like a word.
In proto-cuneiform and proto-Elamite, numerals account for 44% and 42% of all sign tokens
respectively. They cluster in one region of the text. Above all they travel in company:
where a numeral appears at all, another appears in the same short text 32% of the time,
because an account that records one quantity almost always records a second. That
combination is a signature, and it can be learned where the answer is known and then hunted
where it is not.
A third corpus was added later and agrees. Linear A, undeciphered but certainly writing,
keeps its numerals in a separate Unicode block so they identify themselves without any
interpretation; its numerals co-occur at 71%, higher than either Mesopotamian system. Three
independent corpora, two of them undeciphered, all behave the same way.
Nothing in the Indus corpus matches it. Scoring every sufficiently common sign against
the learned profile, the closest candidates reach a co-occurrence rate between 2% and 7%,
against roughly 32% for genuine numerals. No sign in this corpus behaves like a
number. Either these inscriptions do not record quantities at all, which would sit
comfortably with objects used to mark ownership rather than to tally goods, or they record
them in a way that shares no statistical property with how every neighbouring bureaucracy
did it. The result is negative, and it closes off the one anchor that decipherment of a
related script has historically depended on. Proto-Elamite numerals were read long before
anything else in proto-Elamite was. That door is not open here.
XThe Hand and the Stone
Everything so far has treated an inscription as a string of symbols. It is also a
physical act: a person cutting into a piece of steatite about the size of a postage stamp,
who had to decide, before the first cut, how much would fit. The source database records
width, height and thickness for 2,325 seals, and those measurements can be set against the
texts they carry.
Longer inscriptions sit on larger objects. Across 1,607 pieces with
recorded dimensions, text length and seal width move together at rho 0.445. That number on its
own means little, because it could be nothing more than a cataloguing convention: if long
texts belong on big copper tablets and short ones on small clay tags, the correlation is a
fact about object classes rather than about people.
Test
n
Correlation
z
Object type and material held fixed
lengths permuted within class
1,607
0.445
15.4
Face area rather than width
the whole writing surface
1,607
0.411
13.1
Steatite seals alone
one object, one material
1,030
0.406
13.0
After collapsing near-duplicates
edit distance two
481
0.398
8.6
Table 13. Text length against seal width. The stratified control permutes
lengths only within object type and material, so every class convention is preserved in
the null and only the pairing within a class is destroyed. Steatite seals alone hold
object, material and workshop tradition fixed simultaneously.
It survives all of it. Holding object type and material fixed, the effect is unchanged.
Restricted to steatite seals alone, one object in one material, it is unchanged again.
It holds separately at Mohenjo-daro and at Harappa, and it is still present after
near-identical inscriptions are collapsed at edit distance two, which is the gate that
killed five earlier findings in this project. The one place it is absent is clay, where
the correlation is 0.038 on 63 pieces.
Whether the text was planned before the first cut
Two quite different craftsmen produce that same correlation. One knows the text before
he starts and picks a blank large enough to hold it, so the signs stay the same size and
the stone grows with the message. The other starts carving on whatever is to hand and
squeezes as he runs out of room, so the stone stays the same and the signs shrink. Only the
first is planning, and the two can be separated by asking how much of a longer text is
absorbed by a larger surface.
Writing the width of the surface against the number of signs as a power law, the
exponent answers the question directly. An exponent of one means every extra sign was paid
for with proportionally more stone. An exponent of zero means none of it was, and every
extra sign was paid for by cutting smaller.
Corpus cut
n
Exponent
95% interval
z
Steatite seals
1,030
0.252
[0.224, 0.282]
12.4
Tablets
322
0.376
[0.306, 0.450]
7.2
All objects
1,607
0.312
[0.287, 0.341]
17.4
Near-duplicates collapsed, distance one
980
0.372
[0.335, 0.409]
12.9
Near-duplicates collapsed, distance two
481
0.457
[0.397, 0.524]
8.8
Table 14. Allometric exponent of writing surface against text length. Intervals
by subsampling without replacement. An earlier version of this measurement used signs per
millimetre against length, which is not a valid statistic here: length appears on both
sides of it, so it is positive by construction whatever the carvers did.
The answer is both, and mostly the second. On steatite seals the
exponent is 0.25, rising to 0.46 once near-duplicates are collapsed. Somewhere between a
quarter and a half of a longer text was accommodated by choosing a bigger seal; the
remainder was accommodated by carving the signs smaller. A carver with ten signs to fit
did not simply reach for a stone twice the size. He reached for one slightly larger and
cut finer.
That is a modest observation and it is worth being clear about what it does and does not
support. It does not show the signs are language. What it does show is that the length of
the message was known, at least approximately, before the carving began, because the
surface was chosen with reference to it. A sequence produced without foresight, one mark
at a time until the maker stopped, has no reason to correlate with the size of the object
it is on. This is the first result in this work that comes from the objects rather than
from the symbols, and it is independent of every information-theoretic argument above.
A result that did not survive
The same measurement initially separated the two great cities: Mohenjo-daro at 0.188 and
Harappa at 0.476, with non-overlapping intervals, which would have meant the two capitals
had measurably different workshop habits. It does not hold. The excavated assemblages
differ sharply, 290 tablets and 224 seals at one site against 781 seals and 47 tags at the other, and once object type is held fixed
the separation weakens; once the width and length distributions are matched as well, the
gap is -0.033 and the significance is gone entirely. What looked like a difference between
two cities was a difference between two excavations. It is recorded here rather than
deleted, because the reason it failed is more useful than the result would have been.
XIThe Frame
This page has said since its first version that the opening slot of an Indus inscription
is sharply restricted. That is a statement about one position. The question never asked was
whether the other positions are restricted too, and the answer changes what kind of object
an inscription is.
The difficulty is that position has to be measured carefully. Inscriptions here average
under five signs, so if position is expressed as a fraction of the way through a text, a
sign that happens to occur in short inscriptions is handed a middle position by the
arithmetic with no scribe involved. The test therefore has to be run inside a single
inscription length, where every slot exists and every slot is equally available, against a
control that shuffles each inscription's own signs and so holds length, membership and every
sign frequency exactly fixed.
Corpus
4 signs
5 signs
6 signs
7 signs
Bound to a slot
Indus
21/27
29/34
24/27
19/26
82%
Indus, near-duplicates collapsed
7/12
23/31
22/26
20/26
76%
Linear B
14/34
17/39
6/34
7/25
33%
Ur III Sumerian
22/29
23/27
20/26
14/21
77%
Proto-Elamite
30/35
31/46
26/32
13/16
78%
Proto-cuneiform
24/29
29/40
23/32
15/26
72%
Table 15. Signs bound to a particular slot, out of signs common enough to test,
within inscriptions of exactly that length. Benjamini-Hochberg correction applied at each
length. All corpora cut to the Indus token count and length distribution.
Indus inscriptions are strongly slotted: 82% of testable signs are bound to a
position, and 71% survive after near-identical inscriptions are collapsed. This is
not the opening restriction in another guise. Delete the first sign of every inscription and
sweep again, and the anchors are still there.
The frame has named parts
What makes this a frame rather than a statistic is that individual signs hold specific
slots, and hold them across every inscription length. The demonstration is in the two columns
below. As an inscription grows from four signs to seven, a sign anchored to the front keeps
its position counted from the front while its position counted from the back slides; a sign
anchored to the end does the exact reverse. No artifact of length or of measurement produces
two groups of signs moving in opposite directions.
Sign
Slot counted from the front
Slot counted from the back
Anchored to
Share of its uses
sign 740
1, 1, 1, 1
3, 4, 5, 6
first slot
68%
sign 2
3, 4, 5, 6
1, 1, 1, 1
penultimate slot
57%
sign 60
3, 4, 5, 6
1, 1, 1, 1
penultimate slot
83%
sign 820
4, 5, 6, 7
0, 0, 0, 0
last slot
77%
sign 520
1, 1, 1, 1
3, 4, 5, 6
first slot
82%
sign 741
3, 4, 5
2, 2, 2
2 from the end
59%
sign 817
4, 5, 6
0, 0, 0
last slot
90%
sign 861
5, 6, 7
0, 0, 0
last slot
70%
Table 16. The strongest slot-bound Indus signs at inscription lengths four
through seven, deduplicated corpus. The front-anchored and back-anchored groups move in
opposite directions as the inscription lengthens, which is the internal control.
So the corpus has a first-position set, a last-position set and a penultimate-position
set, each occupied by particular signs that stay in their part of the inscription whatever
its length. An Indus inscription behaves less like a sentence and more like a filled
form. That is a structural claim, not a reading, but it is a useful one: it says
where to look. If one slot admits only a handful of signs, those signs are a field type, and
field types are the first thing anyone reads in a bureaucratic script.
The control this panel was missing
Every comparison so far has set Indus against scripts that can be read, which invites the
reply that readability itself is doing the work. There is one script that answers it. Linear
A is undeciphered, its language unidentified, and nobody doubts it is writing, because Linear
B was adapted from it sign by sign and Linear B is Greek. It is the only available example of
an unreadable but unquestionably real writing system, and it belongs beside Indus more than
anything else in this panel does.
Corpus
Status
Signs tested
Bound to a slot
Indus
undeciphered, unknown
228
61%
Linear A
undeciphered, certainly writing
23
26%
Linear B
deciphered, Greek
245
23%
Proto-cuneiform
attested non-language notation
276
41%
Table 17. Slot binding measured at each corpus's own well-populated inscription
lengths, since Linear A tablets run about three times longer than Indus inscriptions.
Absolute rates depend on how often a sign must occur to be testable; the ORDERING does
not, and the ordering is what is being read here.
Linear A sits with Linear B, and Indus sits above both. The two Aegean
scripts agree with each other, which is what their shared ancestry predicts. The Linear A
sample is small, 23 signs testable against 228 for Indus, so this is a difference of degree
rather than a settled quantity. It is also, as the next section shows, a comparison between
the wrong kinds of object.
The comparison that was on our own disk
Every corpus named so far is an archive of tablets. Indus is a corpus of seals. This work
has itself shown that the physical object shapes the text, so a valid comparison must hold object type constant: seals against seals, not seals against tablet archives.
Doing this properly required no new acquisition. The CDLI catalogue, held here since the
beginning of this work, records object type, and 17,024 of its records are seals rather than
tablets. Of those, 6,365 are Ur III, which is 2100 to 2000 BC and therefore contemporary with
the mature Indus period. They are the same kind of object, from the same centuries, serving
the same administrative purpose. And they can be read.
What they say is a formula: personal name, son of personal name, servant of a god or
king, often with a title. That is precisely what Indus seals have been supposed to carry
for a century, which makes this the one place where the frame machinery can be checked
against a known answer. The prediction was written down before the test ran.
Token
Meaning
Predicted position
Position found
Share of its uses
dumu
son of
middle
middle
57%
arad2
servant of
late
penultimate slot
75%
arad2-zu
your servant
last
last slot
98%
dub-sar
scribe, a title
early
slot 2
79%
ensi2
governor, a title
early
slot 2
70%
lugal
king
late
middle
45%
Table 18. The slot-binding sweep run on Ur III seal legends, which are readable.
It was given no information about Sumerian. Predictions were fixed in advance from the
published formula.
It recovers the formula unprompted. Personal names take the first slot,
titles take the second, the word for "son of" sits in the middle, and "your servant" takes
the last slot in 98 per cent of its appearances. The machinery that found the Indus frame
finds the correct frame in a language it was told nothing about. That is the validation the
Indus result needed, and it passes.
One prediction is only half met, and it is left in the table rather than dropped.
Lugal, "king", was predicted late and comes out middling. The reason is visible once
the legends are read: it appears both as the final royal name and inside compound names
earlier in the line, so it genuinely occupies two roles. A test that scored five out of five
would be more suspicious than one that scores four.
Then the comparison itself, with both corpora put through the identical near-duplicate
collapse, because Ur III seals repeat heavily for the same reason Indus seals do: one
official's seal was rolled onto many tablets.
Corpus
Signs tested
Bound to a slot
Ur III seal legends, words
61
97%
Ur III seal legends, signs
142
75%
Indus inscriptions
95
76%
Table 19. Slot binding on seals rather than tablets, every corpus deduplicated
identically. Sumerian is written as sign sequences within words and words within lines, and
nobody knows which of those an Indus sign corresponds to, so both granularities are given
and the answer is bracketed rather than assumed.
Read at sign granularity, Indus is slotted to the same degree as contemporary
readable seal legends: 76% against 75%. Read at word granularity, the same Ur III
corpus scores 97 per cent and Indus sits well below it. Both rows are in the table above and
neither can be preferred, because nobody knows whether an Indus sign answers to a Sumerian
sign or to a Sumerian word. The word reading is the one whose mean legend length matches an
Indus inscription almost exactly, 4.7 tokens against 4.65; the sign reading spreads the same
formula over roughly twice as many slots, which is why the figure falls. Neither of those is
a reason to choose.
So the result is an interval rather than a point: Indus is slotted at or below the
level of readable seal legends, matching them exactly at one end of the interval and falling
short at the other. An earlier version of this page reported only the matching end.
Three further seal corpora extracted since, Early Old Babylonian, Old Babylonian and Old
Assyrian, all score 100 per cent at word granularity, so the upper end of the interval is not
a peculiarity of the Ur III chancery but what seal legends do across six centuries. Set beside
tablet archives Indus looked like an outlier; set beside the same object from the same
centuries it is ordinary or slightly less templated, and the earlier reading on this page has
been corrected accordingly.
This does not say Indus seals carry names and titles. It says that if they did, they would
look like this, and that nothing in their positional structure argues against it. That is the
closest structural analogue this work has found for what an Indus inscription might be, and
the hypothesis it favours is the oldest and least exciting one in the field.
What this does not mean, which matters more than what it does
The obvious temptation is to read a strongly slotted script as an un-language, and the
panel forbids it. Ur III Sumerian is a fully deciphered language and it scores
77% on this measurement, essentially the same as Indus. Linear B is also a fully deciphered
language and it scores 33%, less than half as much. Two readable languages sit at
opposite ends of the scale.
Slot binding therefore does not separate writing from non-writing at all. It separates
formulaic text from free text, and it puts Indus with administrative Sumerian rather than
with the more varied Linear B tablets. That is a real finding about what these objects are
for. It is not evidence either way about whether the signs encode speech, and anyone
quoting the Indus figure without the Ur III figure beside it is misusing it.
XIIHow Close Is This To A Decipherment
Reading an unknown script is not one problem but a sequence of them, and the later ones
are far harder than the earlier ones. Setting them out in order makes it possible to say
where this work actually sits, and where the field sits, without either being flattered.
✓
Machine-readable corpus2,536 inscriptions in hand. The fuller ICIT corpus is held behind a permission request.
field: done
✓
Corpus is not randomLong established in the literature, and reconfirmed here against shuffled controls.
field: done
✓
Ordered, directional, closed unitsSign order carries information, and inscriptions do not run on across objects.
field: done
✓
Not a pure emblem systemModelled ownership marks score negative on sequential dependency where Indus scores above English prose. The non-linguistic case still has serious advocates.
field: contested
✓
Vocabulary is reused, not minted per objectHeaps exponent 0.488 against 0.415 for matched English. Argues against a pure list of names.
field: open
~
Sign classes and grammarPositional classes are established: particular signs hold the first, penultimate and final slots across every inscription length, and survive deduplication. The twelve-category grammar and the opener-against-closer split were both withdrawn. A positional frame is not a grammar, and nobody has demonstrated a grammar.
field: contested
Word or morpheme segmentationNo agreed way to divide Indus strings into units exists.
field: contested
Language family identifiedDravidian, Indo-Aryan and isolate all have advocates. None is established.
field: contested
Phonetic values for signsA full Sanskrit mapping has been claimed and is not independently confirmed. Auditing it is open work.
field: claimed
Read an arbitrary unseen inscriptionThis is the actual bar for decipherment. Nobody is here.
field: not achieved
Survives independent verificationWould require everything above to hold up under outside scrutiny.
field: not achieved
5 of 11
rungs cleared here, one partial, which is roughly 50% of the way
up. That figure should be read with care. The rungs are not equal in size, and the remaining
ones are much larger than the ones already climbed.
The honest position
Nobody has verifiably cleared rung nine or beyond, ourselves included. Rung nine has
been claimed more than once and never independently confirmed. The last three rungs are
what the word decipherment actually refers to, and no one is standing on them.
What this work adds is confined to rungs four and five, and to narrowing rung six. That
is a genuine contribution to a hundred-year-old problem and it is also a long way from
reading the script. Both halves of that sentence matter.
XIIIWhat Is New Here, And What Is Not
A structural analysis of this corpus published in 2026 by Ashish Nair, How
Non-Linguistic Is the Indus Sign System? A Synthetic-Baseline Scorecard
(arXiv:2604.17828), works from the same underlying digitisation and computes several of the
same quantities. Anything here has to be read against it, so the overlap is set out first
rather than left for a reader to find.
Independently replicated
These were computed here before that paper was read, from different sign identifiers and
separate code. The agreement is close enough to be worth stating plainly.
Quantity
Nair 2026
This work
Zipf rank-frequency slope
-1.492
-1.489
Zipf fit R squared
0.956
0.957
Hapax rate
33.2%
33.7%
Mean inscription length
4.42
4.39
Exact duplicate rate
24%
25%
Table 20. Independent agreement on shared quantities.
Extends existing work
The duplicate-structure problem was identified in that paper at the level of exact
repeats. Here it is carried further: inscriptions differing by one or two sign edits are also
not independent observations, which takes 1,902 distinct sequences down to 1,141 and then
558, and every headline measure is reported at all three levels. Separately, the
bias-correction result in section IV appears to run against expectation, since correcting a
null-referenced statistic increases the margin rather than shrinking it.
Not previously attempted, as far as we can establish
No dispersion measure appears to have been applied to this corpus. We applied one, asking
whether signs divide into a frequent evenly-spread class and a rare bursty class, which is
how function words separate from content words in every natural language. They do
not. The apparent split is 94 per cent explained by frequency alone, and once each
sign is compared against a null matched to its own frequency, not one of 153 signs survives
correction. This is a negative result and it is the clearest thing in this work that nobody
had checked.
Withdrawn
Five results looked strong and did not survive their own controls: a twelve-category
induced grammar, a claim that those categories recovered most of the script's predictability,
a correlation between the picture on a seal and the signs above it, the closed-class split
above, and a distinction between opening and closing sign classes. Each is described where it
arose. None is claimed.
The weakness we have not fixed
The comparison against English prose in section III uses continuous text cut into short
segments. It is matched on text count, length distribution, token count and vocabulary size,
but it is not inscriptional writing and it has no genuine text beginnings or endings. Until
the same battery is run against a known script of comparable genre and identical sample
size, such as Linear B tablet lines, the claim that Indus exceeds real writing on sequential
dependency should be read as provisional.
This is the same gap as in the prior work, which uses no linguistic control at all. It is
the next thing being built here, and it is the test on which this argument should be
judged.
XIVMethods
The implementation is not published. The protocol is, in enough detail to be attacked.
Corpus
2,536 inscriptions from the published digitisation of the Corpus of Indus Seals and
Inscriptions: 11,135 sign tokens, 591 sign types, mean length 4.39, ten sites. Analysis
runs on the deduplicated set unless stated. Object type, material and iconographic motif are
carried as covariates. No reconstructed or generated text is used anywhere.
Statistics and controls
Every quantity is reported as a margin over a matched null, never as a raw value. Two
nulls are used: a global shuffle preserving sign frequencies and length distribution, and a
within-inscription shuffle preserving which signs co-occur while destroying their order.
Nulls run at 1,000 to 5,000 iterations. Significance is a z-score against the null
distribution, and effect size is reported alongside it.
Estimator bias
At 0.84 per cent occupancy of the sign-pair table, maximum-likelihood entropy is strongly
upward-biased. Referencing every statistic to a null computed with the same estimator removes
that bias to first order. Because it does not remove it exactly, the principal result is also
reported under Miller-Madow, Chao-Shen and Chao-Wang-Jost corrections.
Independence
Seals were copied, so inscriptions are not independent draws. Results are recomputed after
collapsing exact duplicates and then near-duplicates at edit distance one and two, and
separately at Mohenjo-daro and Harappa. Intervals are computed by a resampling scheme chosen
so that the duplicate structure the analysis exists to control for is not reintroduced.
Comparison corpora
Controls are attested, published corpora rather than invented baselines, each cut to the
dimensions of the Indus corpus and run through the identical procedure. They include
deciphered writing, attested notation that is not writing, and a randomly ordered arm whose
approximately zero result on every measure is the calibration check. The corpora used are
named in the chapters above.
Discipline
Findings are put to an adversarial review instructed to refute rather than confirm them,
and any result that fails deduplication, per-site replication or a matched null is withdrawn
and recorded rather than removed. Predictions are fixed before the tests that decide them,
and multiple comparisons are corrected for.
Availability
Corpus provenance is public and named above, so the inputs are independently obtainable.
Researchers wishing to compare results against their own measurements, or to see the figures
behind any table here, are welcome to make contact.