GeoAI & Machine Learning

Training Data and Annotation for Geospatial AI

GLOBEIR Encyclopedia 2 min readTopic 3 of 5 in GeoAI & Machine Learning

Every supervised model is a compression of its training data — and in geospatial AI, that data must be created: polygons around buildings, masks over crops, labels on change. Annotation looks like unskilled work and is anything but; consistency rules, edge-case policy and quality scoring decide whether a model learns signal or noise. Teams that treat labelling as a production discipline ship better models than teams with fancier architectures.

Specification before annotation

Good programmes begin with an annotation specification: what counts as a building (sheds? ruins? under construction?), where boundaries run (roof edge or footprint?), how occlusions and ambiguity are handled, minimum sizes, class definitions with visual examples. Without it, ten annotators produce ten datasets.

The spec evolves through pilot rounds where disagreements surface — cheap corrections at page stage, expensive ones after a hundred thousand labels.

Quality as a measured property

Professional pipelines measure labels like any product: inter-annotator agreement on overlap samples, review tiers for hard cases, gold-standard tasks seeded through the stream, error taxonomies feeding annotator coaching. Quality scores accompany the dataset, not just the promise of care.

Geospatial adds its own checks: geometric precision against imagery, topology of adjacent labels, alignment across image dates, and CRS bookkeeping so labels land on the right pixels forever.

Strategy: what to label

Naive random labelling wastes budget on easy redundancy. Active learning routes annotators to samples the current model finds uncertain; hard-negative mining targets its confident mistakes; geographic stratification buys the diversity that transfers. Synthetic augmentation stretches datasets but never replaces real edge cases.

And datasets are living assets: versioned, documented (source imagery, date, licence, spec), and extended as deployment reveals new failure modes.

Frequently asked questions

Why is annotation quality more important than quantity?

Noisy labels teach the model noise — and validation against noisy labels hides it. A smaller, consistent, spec-driven dataset routinely beats a larger careless one, in both trained accuracy and honest measurement of it.

What is active learning in annotation?

A loop where the model nominates the samples it is least certain about for human labelling — concentrating effort where it changes the model most, and typically cutting labelling cost substantially for the same accuracy.

Need this done professionally?

GLOBEIR delivers this work for government and enterprise clients.