How to set a change score threshold for field verification
Every change detection layer hands you a score per tile. The question nobody answers for you is where to cut it: send everything above 0.6 to a crew, or hold the line at 0.8 and eat a few missed edits. Set the cutoff too low and crews drive to tiles that never changed. Set it too high and sections of the basemap stay wrong because nobody flagged them for a second look.
Start from what a miss costs you
Before you touch a slider, work out what happens downstream when the layer is wrong. If your output feeds an emergency response basemap, a missed road washout is expensive in a way a false alarm isn't. If you're updating parcel boundaries on an annual cycle, a crew sent to a tile that turns out to be shadow or seasonal crop change just costs a day of field time. Agencies that get this wrong usually set one threshold for the whole program instead of setting it by what the layer feeds.
One way to work this out: pull a sample of forty or fifty tiles spread across the score range, and have a desk analyst call each one before anyone drives out. That exercise shows you the point in the score range where tiles with real change start to outnumber tiles that are just noise, which matters more than landing on a precise accuracy number from a one-time sample.
Desk review vs field check: let the threshold do the sorting
The threshold's real job is to split your worklist into two queues. Tiles below the cutoff get five minutes from someone who knows the area, checking the current imagery against the prior vintage to confirm whether the change is real. That's desk review, cheap and fast, and it catches a lot of what would otherwise look like a false positive once a crew gets there.
Tiles above the cutoff go to field verification, either because the desk analyst can't resolve them from imagery alone, or because the change type (new structure, road realignment, shoreline movement) is one your revision standard requires a ground check for regardless of score. Some agencies also carve out a middle band, scores that get a second desk opinion before anyone decides which queue they belong in. That middle band is where most of the threshold tuning happens over time.
Why the false positive rate matters more than the score
A cutoff that looked right in your pilot sample can drift once the layer is running across a full season. Construction sites get re-flagged as they progress. Agricultural tiles cycle in and out of "changed" depending on when the pass catches them. Track your false positive rate at the threshold you've set: how many field-verified tiles came back "no change," broken out by region and season, since a single program-wide figure hides where the real problem is. A threshold that holds up in a dense urban tile set can flood you with noise in farmland, and the reverse is just as common.
If the rate creeps up, you have two honest options. Raise the cutoff and accept you'll catch fewer genuine changes on the margin, or add a rule that routes certain change types straight to desk review regardless of score, since some categories (seasonal vegetation, shadow, water level) stay noisy at any threshold you pick. Lowering the bar further usually shifts the noise to a different part of the distribution. It rarely removes it.
Revisit the threshold on a set schedule
Set a date, every two quarters or at the start of each new coverage pass, to pull the current threshold against a fresh sample. Basemap vintage changes, sensor coverage changes, and the mix of urban versus rural tiles in your program changes too. A cutoff set once at program launch and never revisited is usually the reason a mapping agency ends up either under-resourced for field work or sending crews to tiles that never needed them.
National Change Alerts builds the worklist this whole process runs on, a continuously updated index of which tiles the basemap no longer matches, scored and ready for you to set your own cutoff against. If you're trying to move a field program off a fixed resurvey cycle and onto a real worklist, that's the place to start.