Method · Scoring
Recall and precision, explained on real archaeology
Recall is the share of the real monuments your survey found. Precision is the share of your candidates that are real. They pull against each other, and almost every published lidar survey reports neither. Over 16 square kilometres of Salisbury Plain we scored 33.1% recall and 23.8% precision. Over Bodmin Moor, 28.7% and 5.7%.
Salisbury 33.1% recall, 23.8% precision. Bodmin 28.7% and 5.7%
A lidar survey announces that it found 400 previously unrecorded features. That number, by itself, tells you nothing at all. It is compatible with a superb survey and with a broken one, and there is no way to tell which from the outside.
Two numbers fix that, and they are the same two used in every other detection problem, from medical screening to spam filtering.
- Recall: of the things that are really there, what fraction did you find? Miss half the barrows and your recall is 50%, however confident your candidate list looks.
- Precision: of the things you flagged, what fraction are real? Flag 400 shapes and have 40 turn out to be archaeology and your precision is 10%.
They trade against each other. Loosen your detector and recall climbs while precision falls, because you catch more real monuments and far more field drains. Tighten it and the reverse. Either number on its own can be driven to look excellent by wrecking the other, which is exactly why a survey that reports one and not the other has told you very little.
What ours are
To get a recall figure you need ground truth: an independent list of what is actually there. In England that is the National Heritage List. We picked two blocks with a lot of scheduled monuments in them, ran the detector unchanged from run 004, and scored the output against the list.
| Salisbury Plain | Bodmin Moor | |
|---|---|---|
| Area | 16 km² | 16 km² |
| Scheduled monuments in the block | 121 | 94 |
| Candidates produced | 168 | 470 |
| Monuments recovered | 40 | 27 |
| Recall | 33.1% | 28.7% |
| Precision | 23.8% | 5.7% |
On Bodmin the detector produced 470 candidates and 27 of them fall within tolerance of a scheduled monument. That is a precision of 5.7%. Roughly 19 out of every 20 things it flagged are not on the list, and while some of those may be genuine and unrecorded, most are probably modern or natural.
Run 006, Bodmin Moor, 4 by 4 km at E222000 N72000, scored against the National Heritage List, 6 August 2026.
Compare the two columns and you can see the trade happening. Bodmin has fewer monuments and nearly three times as many candidates, because a moor is covered in rocks, spoil, peat cuttings and clearance heaps that all look like small mounds. Same settings, much worse precision. Nothing changed except the landscape.
A recall figure needs one more number
Both figures above are correct and one of them is nearly meaningless, which we did not know when we published them.
On Bodmin Moor, 470 randomly placed points score 28.2% recall. Our detector scored 28.7% with 470 candidates. On Salisbury Plain the same test gives 12.0% against our 33.1%, so the detector is doing real work there and almost none on the moor.
Run 007, 9 August 2026. 400 rounds of uniformly random points matched to the real candidate count, scored against the same monuments and the same tolerance rule.
Recall counts a monument as recovered if any candidate lands within the tolerance. More candidates means more chances, detector or no detector. Every recall figure therefore needs its random baseline printed beside it, and the null model has its own page.
The asymmetry nobody mentions
Precision is cheap to improve after the fact. You look at your top candidates, discard the obvious farmyards, and your list gets better. That is triage, and everybody does it.
Recall cannot be improved that way, because the things you missed are not on your list to inspect. You cannot triage an absence. The only way to know your recall is to compare against something outside your own survey, and the only reason to do that is to find out you did worse than you thought.
That asymmetry is the whole reason recall figures are rare. There is no route by which measuring it makes your result look better.
How to read somebody else's survey
- A candidate count with no recall figure is a statement about the settings, not about the ground.
- A recall figure with no tolerance rule is not yet a number. Ours moves by more than a factor of two depending on the rule, and that has its own page.
- A precision figure needs to say what counted as correct. Ours counts a candidate as correct only if it lands within tolerance of a monument already on the national list, which is strict, and undercounts any real find nobody has recorded.
The four questions, applied
The same four we put to every result on this site, turned on this method.
- How much of the corpus?
- Two blocks of 16 square kilometres each, fully covered, containing 121 and 94 scheduled monuments.
- What was recovered?
- 40 of 121 monuments on Salisbury Plain and 27 of 94 on Bodmin Moor, at the stated tolerance.
- What did we say about what we could not do?
- Both figures, in both directions, plus a breakdown by monument type showing three classes at zero. The precision figure is a lower bound, because an unrecorded genuine find scores as a false positive.
- Did anybody check it independently?
- The ground truth is independent, which is the important part. The scoring is ours and has not been rerun by anybody else.
Sources
- LIDAR Composite Digital Terrain Model, 1 metre. Environment Agency, Open Government Licence v3, 2026.
- National Heritage List for England. Historic England, Open Government Licence v3, 2026.
Related
Where this is written up in full
Lost Under the Canopy
Six surveys, 907 candidates, one confirmed. Then the same four questions turned on the most famous survey results in the world.
All four are written and none is on sale. Advance readers can read them first, in exchange for an honest review.