Method · Scoring

What would a random candidate list have scored?

Recall rises when you add candidates, whether or not they detect anything. So the number that says what a recall figure means is the one a scatter of random points scores on the same ground. On Bodmin Moor, 470 random points score 28.2%. Our detector scored 28.7% with 470 candidates.

Bodmin: detector 28.7%, random 28.2%. Salisbury: detector 33.1%, random 12.0%

A recall figure looks like a property of a detector. It is not. It is a property of a detector, a tolerance, and a candidate count, and the last of those can be raised for free.

Scatter enough points over ground with a monument every 170 metres and some will land within 35 metres of one, because that is what scattering points does. Any recall figure includes that contribution and almost nobody reports it, ourselves included until this week.

The test

Take the same monument list, the same tolerance rule and the same block. Replace the candidate list with the same NUMBER of uniformly random points. Score it 400 times and take the mean. That is what your survey would have scored with no detector in it at all.

BlockDetectorCandidatesSame count, randomDifference
Salisbury Plain33.1%16812.0%*+21.1 points*
Bodmin Moor28.7%47028.2%*+0.5 points*
Run 007, 9 August 2026. 400 rounds per block, fixed seed. Same monuments, same tolerance rule of 35 metres or the monument span capped at 120.
Refuted

On Bodmin Moor our published recall figure of 28.7% is indistinguishable from what 470 random points score on the same ground. The figure is not wrong. It is uninformative, and we published it for three days without the number that says so.

Run 007, pipeline/runs/run-007-hollows.json. Random baseline over 400 rounds with seed 20260809, scored against the same 94 monuments and the same tolerance rule as run 006.

On Salisbury the same detector beats random by 21.1 points, which is a real result. So this is not a broken method. It is a method whose recall figure means something on one block and almost nothing on another, and nothing in the figure itself tells you which.

Why Bodmin and not Salisbury

Two things compound. Bodmin produced 470 candidates against Salisbury's 168, because a moor is covered in clitter, clearance heaps and peat cuttings that all look like small mounds. And the monuments are denser and smaller, so the 35 metre tolerance covers a larger share of the block.

More candidates over more generously scored ground. The null rises to meet the detector, and on that block it arrives.

What this does to the hollow finder

It is the reason run 007 exists in this form. Adding a hollow detector raised recall to 41.3% on Salisbury and 46.8% on Bodmin, which read as large improvements.

BlockMounds onlyMounds plus hollowsMounds plus the same number of random pointsReal gain
Salisbury Plain33.1%41.3%41.0%+0.3 points
Bodmin Moor28.7%46.8%44.2%+2.6 points
On Bodmin the null's 95th percentile is 50.0%, which is above the observed 46.8%. The gain is inside the noise.

The hollow finder does find hollows. It recovered a pond barrow, which the mound detector could not do at any setting, and that is a genuine closure of a structural blind spot. What it did not do is improve the survey by as much as the headline recall says, because most of that improvement is the arithmetic of having 222 and 540 more candidates.

What we are doing about it

  • Every recall figure this project publishes from now on carries its random baseline in the same table. Not in a footnote.
  • The run logs already store it, seeded, so anybody can reproduce the baseline as well as the result.
  • Nothing published earlier is being retracted. The figures were correct. What was missing was the number that says what they mean, and the fix for that is to add it, not to remove them.

This is the third time on this site the missing number has turned out to be the same shape. A false-positive rate for the Indus entropy measure. A character error rate for Herculaneum ink detection. And now a random baseline for our own recall. We have been asking other people for it in public for a fortnight and did not have it ourselves.

The four questions, applied

The same four we put to every result on this site, turned on this method.

How much of the corpus?
Both scored blocks, all monuments, 400 random rounds each with a fixed seed.
What was recovered?
A baseline: 12.0% on Salisbury Plain and 28.2% on Bodmin Moor for random points matched to the real candidate count.
What did we say about what we could not do?
That one of our two published recall figures cannot be distinguished from noise, and that we published it without this number for three days.
Did anybody check it independently?
No. The baseline is ours, the seed is published, and the code is four lines of numpy. Anybody can rerun it and get the same figures.

Sources

Related

Where this is written up in full

Lost Under the Canopy

Six surveys, 907 candidates, one confirmed. Then the same four questions turned on the most famous survey results in the world.

All four are written and none is on sale. Advance readers can read them first, in exchange for an honest review.