Method · Scoring
The tolerance problem: one setting, and the recall figure moves by a factor of five
To score a survey you must decide how close a candidate has to be to count as a hit. That distance is the tolerance, it is almost never published, and it decides the answer. On Bodmin Moor our recall runs from 9.6% at 25 metres to 46.8% at 100 metres. Same detections, same monuments, one setting.
Bodmin recall 9.6% at 25 m, 46.8% at 100 m. Salisbury 21.5% to 48.8%
Your detector puts a candidate at a set of coordinates. A scheduled monument sits 60 metres away. Is that a hit?
There is no answer to that question in the data. It is a choice, made by whoever is doing the scoring, and it moves the headline number more than most of the actual engineering does.
The sweep
We ran the same two scored surveys at five fixed tolerances and changed nothing else. Not the detector, not the candidates, not the monument list. Only the distance at which a match counts.
| Match distance | Salisbury Plain recall | Bodmin Moor recall |
|---|---|---|
| 25 m | 21.5% | 9.6% |
| 35 m | 24.8% | 16.0% |
| 50 m | 29.8% | 24.5% |
| 75 m | 40.5% | 35.1% |
| 100 m | 48.8% | 46.8% |
Bodmin goes from 9.6% to 46.8%. That is a factor of nearly five, and every one of those percentages is a true statement about the same survey.
Read the two columns as two ranges and keep them apart. Salisbury runs 21.5% to 48.8%. Bodmin runs 9.6% to 46.8%. Merging them into a single sweeping range is a real and easy mistake, and we made it in three separate book drafts before a check caught it. Every number in the merged version was true and the sentence was false.
Why a bigger tolerance is not simply cheating
There is a genuine argument for a loose tolerance and it deserves stating properly. A heritage list entry is a point, and the monument is not. A long barrow is 60 metres of earthwork with one grid reference attached to it. A detector that lands on the far end of the mound is 40 metres from the listed point and has plainly found the monument. Score that as a miss and you understate the survey.
The argument for a tight tolerance is just as real. At 100 metres over ground as dense as Salisbury Plain, a candidate is within reach of several monuments at once and can be credited with finding something it never resolved. Loose scoring turns a rough detector into a good one on paper.
The case for scaling tolerance with size
A monument is an extent, not a point, and the record stores a point.
A fixed 35 metres penalises exactly the largest and most obvious monuments, which is backwards.
Scaling by the monument span is defensible and computable from the record itself.
The case for a flat cap
Unbounded scaling produces absurdities on the largest features, and did.
On dense ground a wide radius credits one candidate with several monuments.
A rule with a cap is harder to tune towards a flattering answer.
The aqueduct, and the rule we changed
Our first scoring rule was 40 metres, or the monument span if that was larger, with no upper limit. It sounded principled. Then it met a Roman aqueduct.
A 2,749 metre stretch of Roman aqueduct was scored as recovered by a candidate 891 metres away, because the rule scaled tolerance with the monument span and the span was enormous. Nothing about that is a detection. We capped the tolerance at 120 metres from run 005 onward.
Run 004, Dorset downland block A, rescored 6 August 2026. Both the as-run rule and the corrected rule are recorded in pipeline/runs/run-004-dorset.json.
The uncomfortable part is not the bug. It is that we changed the measuring rule after seeing what it did to a result, which is precisely the move this project exists to catch other people making. So the run file records both rules and the reason, rather than quietly holding the new one. A rule change that improves your number is only defensible if the change is visible.
What to publish
- The tolerance rule, in full, including any cap.
- A sweep, not a single value. One number invites the reader to assume it is robust.
- Whether the rule was chosen before or after the scores were seen. Ours was changed after, and that is written down.
The four questions, applied
The same four we put to every result on this site, turned on this method.
- How much of the corpus?
- Both scored runs, all five tolerances, every candidate and every monument in each block.
- What was recovered?
- Recall figures spanning 21.5% to 48.8% on Salisbury Plain and 9.6% to 46.8% on Bodmin Moor, from identical detections.
- What did we say about what we could not do?
- That we changed the tolerance rule after seeing a result it produced, and that both rules are kept in the run file rather than one.
- Did anybody check it independently?
- The rescoring was done by a script with no model in the loop, comparing the manuscripts against the run logs. That catches drift between text and data, and it is not the same as an outside team.
Sources
- National Heritage List for England. Historic England, Open Government Licence v3, 2026.
- LIDAR Composite Digital Terrain Model, 1 metre. Environment Agency, Open Government Licence v3, 2026.
Related
Where this is written up in full
Lost Under the Canopy
Six surveys, 907 candidates, one confirmed. Then the same four questions turned on the most famous survey results in the world.
All four are written and none is on sale. Advance readers can read them first, in exchange for an honest review.