← Back to all Thoughts RSS Feed
Blog post icon September 12, 2026 • 2 min read • Originally published on Other

Spatial SQL, Cloud-Native Imagery, and Analysis at Scale

One of the most important shifts happening in geospatial right now is the collapse of the old “download → preprocess → stitch → analyze” workflow. Tools like SedonaDB show what happens when satellite imagery becomes cloud‑native data and the entire pipeline collapses into SQL + NumPy.

This article from the Apache Sedona team on the Gironde/Landes wildfire is a perfect illustration of this new pattern.

Cloud-native imagery as queryable tables
Planet publishes crisis imagery as Cloud Optimized GeoTIFFs with STAC‑GeoParquet metadata. SedonaDB reads that catalog directly over HTTPS, and a ST_Intersects query selects the 11 PlanetScope scenes that overlap the study area: eight pre‑fire, three post‑fire. Scene selection took about two seconds, and nothing was downloaded.

Raster alignment in SQL
In this Sedona example, each scene is opened lazily with RS_FromPath: metadata loads instantly, pixels stay remote until needed. RS_Clip and RS_ReprojectMatch then align all scenes onto a 12 m reference grid, averaging the original 3 m pixels. Six raster queries (red, NIR, and usable‑data masks for both dates) took ~11.5 minutes end‑to‑end — on a single machine. NumPy mosaics the aligned rasters, keeping the first clear pixel per cell. Coverage reached ~87% despite clouds, smoke, and scene gaps.

Vegetation loss as a SQL‑backed NumPy operation
As expected from the burn, NDVI drops between dates flag vegetation loss. Otsu’s method found a threshold of 0.225 for pixels with healthy pre‑fire NDVI. The burn mask becomes an in‑database raster via Raster.from_numpy, ready for SQL operations.

Pixels → polygons → statistics
They used RS_Polygonize to convert the 7‑million‑cell mask into 3,769 patches in about three seconds. Filtering to patches ≥1 ha yields 173 polygons covering 24,091 ha: the largest a 21,090 ha burn scar between Le Porge and Lanton. RS_ZonalStats then computes burned area by commune in ~14 seconds, producing a clear ranking:
Le Porge (6,202 ha), Saumos (4,045 ha), Lanton (3,499 ha), Arès (3,498 ha), and so on.

Why this matters
This is what “geospatial at scale” actually looks like:

  • Data stays in the cloud; SQL streams only what’s needed.
  • Raster operations become declarative, not a tangle of scripts.
  • NumPy acts as zero‑copy compute, not a separate pipeline.
  • Outputs are transparent: polygons, GeoParquet, GeoTIFF, all versionable.
  • Performance is predictable: the entire raster run took ~12 minutes on one machine.

The distance between raw imagery and actionable insight is shrinking; not because machines are bigger, but because the abstractions are finally right.

Spatial SQL + cloud‑native rasters isn’t just a convenience. It’s the foundation for environmental analysis that is auditable, scalable, and operationally sane.

Related Thoughts

Perspectives sharing related architectures, models, and domain context.

All Thoughts →
Sep 22, 2026 5 min read

The Shift to Cloud-Native Geospatial: Access, Scale, and the Open Source Ecosystem

Traditional GIS architectures were built around assumptions that no longer hold: you download data to your local...

Sep 21, 2026 4 min read

The Convergence of Spatial SQL and Ontologies: Building a Semantic Backbone for Earth Observation

The geospatial industry is undergoing a quiet but profound shift. For decades, the standard workflow for analyzing...

Oct 01, 2026 3 min read

Everyone's Thinking Spatially Even If They Don't Call It Geography

Matt Forrest shared a story that struck a chord with me. He talked about winning an atlas in a third grade contest and...