Subject · Text

The Dead Sea Scrolls

What machine methods have actually established, what they have not, and the one result here that published its own error rate.

Three machine results matter here. One of them did something almost nothing else in this series does, which is to say in public how often it is wrong.

This page is about those three, and about the difference between them. It is not a general introduction to the scrolls, and it is not an edition of them. If you want to read the texts themselves, this is the wrong page and there are better ones.

Macro photograph of banded sedimentary rock, ochre and grey layers running horizontally

Layers, in the order they were laid down. That is what dating a corpus is trying to recover, and for the scrolls almost none of it survives in the ground.

Brand imagery · generated, not survey data · Before the Record 2026

Why dating them was so hard in the first place

Almost none of the scrolls carry a date, and most have no useful stratigraphy in the archaeological record. Only some of the very oldest and the very youngest manuscripts can be pinned to a calendar at all.

So the chronology was built on palaeography: reading the shape of the script and placing it in a developmental sequence. That works when you have dated anchors to calibrate against, and for the centuries in the middle of this corpus there are barely any. The sequence was, to a real extent, resting on assumptions about how quickly a script changes.

The 2025 result, and the number that makes it different

In June 2025 a team from Groningen, Southern Denmark, Pisa and Leuven published a model called Enoch. They radiocarbon dated 30 samples of parchment, got 27 valid dates out of them, and used 24 as a training set. Enoch learns the relationship between writing style and radiocarbon date, using Bayesian ridge regression over angular and allographic features of the script.

Then they did the thing this whole series exists to ask for.

Published

They published a mean absolute error. Enoch predicts radiocarbon-based dates to within 27.9 to 30.7 years, cross-validated three separate ways: a train-validation split, a manuscript-image split, and leave-one-out over the training points. It returns a posterior distribution rather than a single number, and it carries error margins onto unseen data.

Popović, M. et al. (2025), Dating ancient manuscripts using radiocarbon and AI-based writing style analysis, PLOS ONE 20(6): e0323185. DOI 10.1371/journal.pone.0323185. CC BY 4.0.

Read that against everything else on this site. A lidar survey reports a candidate count. A muon survey reports a corridor. The Herculaneum work reports how many letters it recovered. None of them reports how often it is wrong.

Enoch does, and the figure is roughly 30 years on a corpus spanning several centuries. That is what a usable result looks like, and it is the reason the rest of the study is worth taking seriously.

What it changed

Both the radiocarbon ranges and Enoch's predictions come out older than the traditional palaeographic estimates, often by a meaningful margin, including for the emergence of the script types the old chronology was built on.

Two cases carry most of the weight. 4Q114 preserves Daniel chapters 8 to 11, which scholars date on literary and historical grounds to the 160s BCE. Its accepted two-sigma calibrated radiocarbon range is 230 to 160 BCE, which overlaps the period in which that part of Daniel is thought to have been written. And 4Q109, a copy of Ecclesiastes, gets a third century BCE prediction from Enoch, against a book scholars tentatively place at the end of the third century.

If those hold, they are the first known fragments of a biblical book from the lifetime of its presumed authors. That is a genuinely large claim and it rests on a dating method with a stated error bar, which is the only reason it can be argued with.

Both sides, at full strength

Why Enoch should be believed

It publishes an error rate. Nothing else in machine archaeology does, and a method you cannot score is a method you cannot check.

Cross-validated three ways, including leave-one-out, which is the strict test on a small dataset.

It returns a distribution, not a point estimate, so a reader can see the uncertainty rather than being handed a year.

The radiocarbon dates move in the same direction independently of the model. Two different kinds of evidence pointing the same way is harder to dismiss than either alone.

Why it is not settled

Twenty-four training manuscripts is a very small dataset for a model asked to generalise across several centuries of script.

The 79% agreement with palaeography is described as a post-hoc evaluation, and palaeography is the thing being corrected. Agreement with the old chronology cannot validate a result whose headline is that the old chronology was wrong.

The authors themselves anticipate concerns from palaeographers about the relation between the target data and training data of different origin and period. That is an unresolved objection, raised by the people who did the work.

Radiocarbon on parchment has its own problems, including contamination from historical conservation treatments, and the labels the model learns from are ranges rather than dates.

Our reading: the error rate makes this the strongest machine result in the whole series, and the small training set means the specific redatings should be treated as well-founded proposals rather than settled facts. Those two statements are compatible, and most coverage picked one.

The second result: two hands in the Great Isaiah Scroll

The Great Isaiah Scroll, 1QIsaa, is 7.34 metres long, averages 26 centimetres high and carries 54 columns of Hebrew. Whether one scribe or two wrote it is a question palaeographers could not settle by eye, because the writing is near uniform.

A 2021 study attacked it in three stages, and the order is what makes it good. Unsupervised first, with no assumption about where a break might be: columns from the first and second halves of the manuscript separated into different regions of feature space. Then an independent supervised analysis, assuming a break and looking for it, which located a transition at columns 27 to 29. Only then did they look at the letters, averaging character shapes across each half and finding consistent differences in the strokes of aleph and resh.

Under review

Two main scribes, with the break at columns 27 to 29. The complication, which the authors raise themselves, is that there is a change of sheet between columns 27 and 28, with a three-line gap at the foot of column 27. A model separating two halves that were physically two pieces might be detecting the pieces.

Popović, M., Dhali, M. A., Schomaker, L. (2021), PLOS ONE 16(4): e0249769. DOI 10.1371/journal.pone.0249769. CC BY 4.0.

The authors' answer is the third stage: a difference in surface or ink would not be expected to change the shape of an aleph in a consistent, letter-specific way. That is a real answer and not a complete one.

What would close it is the same measurement missing everywhere else. Run the method across manuscripts that have sheet joins and no suspected change of hand, and report how often it finds a break anyway. That is a false-positive rate. It has not been published.

And note that a seam is exactly where two scribes sharing a scroll would divide it. The answer being where you would expect the answer to be is not evidence against the answer.

The third: an imaging programme that never counted its yield

The Leon Levy Dead Sea Scrolls Digital Library launched in December 2012, built by the Israel Antiquities Authority with Google, and put roughly 930 manuscripts and thousands of fragments online in high resolution and multiple spectra.

The stated aim was conservation. A non-invasive way to monitor the physical condition of objects too fragile to handle. Recovering text that could not previously be read was a consequence of the imaging, not the brief.

Candidate

How much previously unreadable text did it recover? No figure has been published that we can find. The capability is real, the images are public and free, and the yield was never counted. A capability without a published yield is the same shape of claim this series keeps taking apart elsewhere.

Our own search, 9 August 2026. No published figure traced for previously unreadable text recovered by the multispectral programme.

The complaint survives, narrowed. Those images fed the 2021 scribes study, which used the collection to extract dominant character shapes across the corpus. So the programme did produce research yield. Nobody counted it, and nobody asked them to.

Where this leaves things

  • Published: Enoch predicts radiocarbon-based dates from script shape with a mean absolute error of 27.9 to 30.7 years, cross-validated three ways.
  • Under review: the specific redatings of 4Q114 and 4Q109, and the two-scribe conclusion for the Great Isaiah Scroll. Both are well-founded and neither is closed.
  • Candidate: the total quantity of text recovered by multispectral imaging of the scrolls. No published figure traced.
  • Refuted: nothing here.
  • Speculative: any reading of the new chronology that treats a 30-year error bar as a date.

The full argument, including the Herculaneum scrolls and the scripts machines still cannot read, is Found in the Scrolls, book 3 of the series. It is written and not yet on sale, and advance readers can read it first. How the imaging and machine reading actually work is in the method pages.