GeoAI & Machine Learning

Deep Learning on Satellite Imagery

GLOBEIR Encyclopedia 2 min readTopic 2 of 5 in GeoAI & Machine Learning

Deep learning transformed satellite image analysis by learning features instead of hand-crafting them. Convolutional networks — and now transformer architectures — read imagery the way trained interpreters do: from texture, shape and context, not just per-pixel spectra. Semantic segmentation models delineate buildings, roads, water and crops pixel by pixel; detection models find discrete objects; change networks compare dates. At continental scale, they do in days what manual digitisation could never finish.

The architectures that matter

Semantic segmentation is the mapping workhorse: encoder–decoder networks in the U-Net lineage classify every pixel, producing mask layers that vectorise into GIS features. Instance segmentation separates touching objects (adjacent buildings); object detection boxes discrete targets (vehicles, ships, aircraft); Siamese networks compare image pairs for change.

Vision transformers and geospatial foundation models — pretrained on massive unlabelled archives — increasingly provide the backbone, cutting the labelled-data cost of each new task.

What makes imagery different from photos

Satellite scenes are huge (tiling and stitching strategies matter), multispectral (networks can ingest NIR and SWIR bands photos lack), georeferenced (outputs must be too), and viewed from above (no canonical “up”, so rotation augmentation is standard). Resolution varies by an order of magnitude across sensors, and labels must align to pixels precisely.

These specifics reward geospatially literate ML teams; generic vision pipelines stumble on exactly these seams.

The production reality

Model quality is mostly training-data quality: consistent annotation, hard examples, geographic diversity. Inference at scale is engineering: tiling, edge effects, post-processing (regularising building polygons, snapping road topology) and merging into seamless layers.

And accuracy must be proven per deployment — confusion metrics on independent areas, not leaderboard numbers — before outputs feed decisions.

Frequently asked questions

How much training data does a segmentation model need?

From a few thousand annotated image chips for fine-tuning pretrained backbones to hundreds of thousands for training from scratch. Foundation-model pretraining keeps shrinking the labelled requirement — but never to zero, and never without geographic diversity.

Can deep learning work on free Sentinel imagery?

Yes — at 10 m it excels at land cover, crops, water and large-structure mapping. Building-level extraction, however, needs sub-metre commercial imagery; resolution bounds what any model can resolve.

Need this done professionally?

GLOBEIR delivers this work for government and enterprise clients.