Subject · Text

The Indus script

A hundred and fifty years, hundreds of claimed decipherments, and an argument about whether the thing is writing at all.

Most undeciphered scripts are undeciphered because nobody knows the language behind them. The Indus script has that problem and a worse one on top: a serious body of scholarship argues it is not a script.

That is what makes it the best test case in this series. Everywhere else, a machine method is asked what an object says. Here it was asked something prior and stranger: whether the object is the kind of thing that says anything.

What survives

The Indus civilisation flourished in South Asia from roughly 2600 to 1900 BCE. It left about 4,000 inscriptions, on seals, miniature tablets, pottery, stoneware, copper plates, tools and weapons. The symbol inventory is estimated at around 400 distinct signs.

Now the number that constrains everything else. The inscriptions are tiny. Rao and colleagues give an average length of five signs, with the longest on a single surface running to 17. Sproat, describing the standard Mahadevan corpus specifically, gives about 7,000 sign occurrences in total, roughly 400 types, and an average inscription of about 4.5 symbols.

Published

The entire corpus is around 7,000 signs. That is shorter than this page. Every statistical argument about the Indus script, on both sides, is being made on a dataset that would fit in a long article, and no amount of method compensates for that.

Rao, R. P. N. et al. (2010), Computational Linguistics 36(4), pages 795 to 805, and Sproat, R. (2010), Computational Linguistics 36(3), pages 585 to 594. The two figures for average inscription length are given separately by each side and are not merged here.

The prior question

A writing system, as linguists usually define it, is a symbol system used to represent language. Plenty of symbol systems are not: heraldry, mathematical notation, dance notation, road signs. They carry meaning without encoding speech.

Since the first seal impression was found at Harappa in the 1870s, the standard assumption has been that the Indus symbols are writing and that the civilisation was literate. In 2004 Steve Farmer, Richard Sproat and Michael Witzel published an argument that it is not, and that the symbols are a non-linguistic system.

That paper is contested in turn, in print, by named scholars. Rao and colleagues cite rebuttals from Parpola, Vidale and McIntosh among others, and Vidale's conclusion is blunt: he sees no acceptable demonstration of a non-scriptural nature, and therefore no collapse.

Nobody has settled this. Keep that in view, because the statistical argument below is a move inside a fight that was already running.

The 2009 measurement

In April 2009, Rajesh P. N. Rao and colleagues published a one-page paper in Science arguing from conditional entropy that the Indus system is more likely than not a writing system.

Conditional entropy measures how much uncertainty remains about the next symbol given the one before it. A completely rigid system scores near zero: if y always follows x, there is no uncertainty and no information. A completely random and equiprobable system scores at the maximum. Natural languages sit in a characteristic band between the two, because grammar constrains what can follow what without reducing it to one option.

The paper plotted the Indus sequences against samples of English, Sumerian and Old Tamil, and against two non-linguistic reference curves labelled Type 1 and Type 2. The Indus points landed with the languages.

It was received as a resolution. Wired ran it as artificial intelligence cracking an ancient mystery. The Indian press covered it heavily.

The objection, and why it is not a quibble

Richard Sproat took it apart in Computational Linguistics in 2010, in a Last Words column whose title also went after the journals that published it.

His first point is about the reference curves. Read the supplementary material rather than the page and you find that Type 1 and Type 2 were not measurements of any real symbol system. They were generated: Type 2 from a rigid model, Type 1 from a random and equiprobable one. Neither extreme has plausibly ever existed, because a rigid system carries no information and a random equiprobable one contradicts the fact that symbols stand for things that occur at different rates.

Then he does the part that makes the objection stick.

Published

Sproat reports reproducing the Indus curve from a model with 400 elements, a Zipf distribution with an exponent of 1.5, and conditional independence for the bigrams. Conditional independence means there is no sequential structure whatsoever. He reports a second reproduction from an artificial corpus matching only the single-symbol frequencies of the Mahadevan corpus, again conditionally independent.

Sproat, R. (2010), Ancient Symbols, Computational Linguistics, and the Reviewing Practices of the General Science Journals, Computational Linguistics 36(3), pages 585 to 594. DOI 10.1162/coli_a_00011.

If a corpus with no sequential structure produces the curve, the curve is not evidence of sequential structure. That is a demonstration, not a philosophical worry about necessary versus sufficient conditions.

He also explains why it works, which matters. With a Zipfian distribution and no ordering constraints, the most frequent symbols carry a lot of uncertainty. As the sample grows to include rarer symbols, each batch adds less, because each batch is rarer. The curve flattens. And it never reaches the maximum, because the symbols are not equally probable. You get the shape of language out of a system containing none of it.

What Rao and colleagues said back

They replied in the same journal, in the following issue, and Sproat answered the reply in that same issue. Four papers, two journals, all of it published and all of it readable.

The reply

Misrepresentation. They say they never claimed a conditional entropy match alone proves a system is linguistic, only that it adds weight alongside the script's other language-like properties.

The figure not discussed. Sproat reproduces Figure 1A and does not mention Figure 1B, which includes DNA, protein sequences and Fortran, none of them artificial.

A stronger measure. Block entropy generalises the bigram measure to blocks of up to six symbols, estimated with a Bayesian method built for undersampled data. On it the Indus texts still track natural languages while DNA, protein and music sit noticeably higher.

Where it leaves the objection

A weight-of-evidence claim and a diagnostic-test claim are different claims, and only the second is refuted by a false positive. So this is a real disagreement about what was asserted, not only about whether it holds.

It is also a disagreement neither side can win from outside, because the Science paper is one page and what a reader took from it depends on which sentence they weighted.

The block entropy result is a genuinely stronger measure and it is still a measure on 7,000 signs. The sample size objection survives every improvement in method.

What would settle it

We are not going to referee this. What we will say is what a resolution looks like, because both sides could agree on the test in advance.

Give the measure a false-positive rate. Sproat's objection is a claim about a rate: that corpora with no sequential structure land in the language band. Generate a large set of synthetic corpora spanning the plausible range of real non-linguistic symbol systems, at Indus corpus size, and report how often the measure puts them with the languages.

If it almost never does, Sproat's reproduction was a special case and the measure has diagnostic power. If it often does, the measure does not, and no further argument about who claimed what will change it.

Sixteen years on, nobody has published that number. It is the same missing number as the Herculaneum error rate and the recall figure on a lidar survey, in a third discipline.

Where this leaves things

  • Published: the corpus is about 4,000 inscriptions and roughly 400 symbol types, and the inscriptions are extremely short. All four papers in the entropy exchange are published with DOIs.
  • Under review: whether conditional entropy can distinguish writing from a rule-governed non-linguistic system on a corpus this size. Argued across four papers and not closed by any of them.
  • Candidate: the Indus script as a genuine writing system encoding a lost language. It is the traditional view and it is not established.
  • Refuted: the reading that a 2009 statistic settled the question. Both sides now agree it did not, for different reasons.
  • Speculative: every claimed decipherment to date.

The longer version, with the scripts machines have cracked and the conditions that made those possible, is Found in the Scrolls, book 3 of the series. Advance readers can read it before it goes on sale. How machine reading of ancient text actually works is in the method pages.