Method · Scoring

How to score your own survey against a national monument record

Choose the block by a written rule before you fetch any data. Run the detector unchanged. Pull the monument records for the block. Match candidates to monuments at a stated tolerance, sweep it, and break the result down by monument type. The type breakdown is where the useful information is, because that is where the zeros show up.

Block chosen by rule at 18:20:42, lidar fetched at 18:21:24. Both timestamps in the run file

Scoring your own survey is easy to do and easy to do dishonestly, and the dishonesty rarely feels like dishonesty at the time. It feels like picking a good test area.

Step 1: choose the ground by rule, before you look at it

This is the step that decides whether the whole exercise means anything. If you pick a block, score badly, and try somewhere else, you will eventually find ground where your detector looks good, and you will have measured nothing except your own persistence.

So write the rule down first and let it choose. Ours is: the densest window of scheduled monuments in the search region. Not the easiest ground. The densest, which on Bodmin meant a moor covered in cairns and clitter and gave us our worst precision figure by a wide margin.

Published

The block selection is timestamped to disk 42 seconds before the lidar fetch begins. That is not ceremony. It is the only evidence that a friendlier block was not substituted afterwards, and without it the recall figure is an assertion.

Run 006, Bodmin Moor. Block chosen 2026-08-06T18:20:42, survey started 18:21:24. Both timestamps stored in pipeline/runs/run-006-bodmin.json.

Step 2: do not touch the detector

Runs 005 and 006 both record the same line: detector unchanged from run 004, because changing it here would be tuning to the answer. Once you can see the score, every adjustment you make is informed by it, and the number stops being a measurement.

Tune all you like. Then rescore on ground the tuning has never seen.

Step 3: get the ground truth

In England, the National Heritage List for England covers scheduled monuments and is published under the Open Government Licence. Pull every entry whose coordinates fall inside your block. That list is your denominator, and it is deliberately not a complete inventory of everything that exists, only of what is protected. Your recall is therefore recall against the record, which is a slightly different and more honest claim.

Step 4: match, sweep, and break it down by type

Match at your stated tolerance, then run the sweep. Then do the part that most reporting skips, which is to split the result by monument class.

Monument typeRecoveredTotalRecall
Round barrow3560.0%
Bell barrow4850.0%
Bowl barrow318735.6%
Long barrow2633.3%
Disc barrow060%
Pond barrow050%
Henge020%
Cursus010%
Run 005, Salisbury Plain, by monument class. The aggregate figure was 33.1%.

The aggregate said 33.1%. The breakdown says something far more useful: the detector is reasonable on domed mounds and completely blind to four other classes. Three of those zeros are structural rather than tuning problems, and you would never see them in the headline figure.

Step 5: run the null before you believe your own number

Replace your candidate list with the same NUMBER of uniformly random points, score it a few hundred times, and take the mean. That is what your survey scores with no detector in it.

We added this step after publishing two recall figures without it. One of them, 28.7% on Bodmin Moor, turns out to be indistinguishable from what 470 random points score on the same ground. The other beats random by 21 points. Nothing in either figure told us which was which.

Step 6: publish the zeros

A zero in a class is the most informative cell in the table. It tells you the method has a shape it cannot see, which is a fact about the instrument and not about the ground, and it will hold everywhere you take it.

It is also the number nobody wants in their abstract. That is the reason to make it a rule rather than a preference.

The four questions, applied

The same four we put to every result on this site, turned on this method.

How much of the corpus?
Two blocks of 16 square kilometres, chosen by a written rule before any data was fetched, with the choice timestamped.
What was recovered?
A procedure that can be rerun by anyone with open lidar and an open heritage record, and two worked examples with their full type breakdowns.
What did we say about what we could not do?
That recall here is recall against the protected record rather than against everything that exists, and that no candidate has been checked on the ground.
Did anybody check it independently?
The heritage record is independent. The scoring code is ours, and a separate script verifies that the published text matches the run logs with no model involved.

Sources

Related

Where this is written up in full

Lost Under the Canopy

Six surveys, 907 candidates, one confirmed. Then the same four questions turned on the most famous survey results in the world.

All four are written and none is on sale. Advance readers can read them first, in exchange for an honest review.