Spatial SQL, Cloud-Native Imagery, and Analysis at Scale
One of the most important shifts happening in geospatial right now is the collapse of the old “download → preprocess → stitch → analyze” workflow. Tools like SedonaDB show what happens when satellite imagery becomes cloud‑native data and the entire pipeline collapses into SQL + NumPy.
This article from the Apache Sedona team on the Gironde/Landes wildfire is a perfect illustration of this new pattern.
Cloud-native imagery as queryable tables
Planet publishes crisis imagery as Cloud Optimized GeoTIFFs with STAC‑GeoParquet metadata. SedonaDB reads that catalog directly over HTTPS, and a ST_Intersects query selects the 11 PlanetScope scenes that overlap the study area: eight pre‑fire, three post‑fire. Scene selection took about two seconds, and nothing was downloaded.
Raster alignment in SQL
In this Sedona example, each scene is opened lazily with RS_FromPath: metadata loads instantly, pixels stay remote until needed. RS_Clip and RS_ReprojectMatch then align all scenes onto a 12 m reference grid, averaging the original 3 m pixels. Six raster queries (red, NIR, and usable‑data masks for both dates) took ~11.5 minutes end‑to‑end — on a single machine. NumPy mosaics the aligned rasters, keeping the first clear pixel per cell. Coverage reached ~87% despite clouds, smoke, and scene gaps.
Vegetation loss as a SQL‑backed NumPy operation
As expected from the burn, NDVI drops between dates flag vegetation loss. Otsu’s method found a threshold of 0.225 for pixels with healthy pre‑fire NDVI. The burn mask becomes an in‑database raster via Raster.from_numpy, ready for SQL operations.
Pixels → polygons → statistics
They used RS_Polygonize to convert the 7‑million‑cell mask into 3,769 patches in about three seconds. Filtering to patches ≥1 ha yields 173 polygons covering 24,091 ha: the largest a 21,090 ha burn scar between Le Porge and Lanton. RS_ZonalStats then computes burned area by commune in ~14 seconds, producing a clear ranking:
Le Porge (6,202 ha), Saumos (4,045 ha), Lanton (3,499 ha), Arès (3,498 ha), and so on.
Why this matters
This is what “geospatial at scale” actually looks like:
- Data stays in the cloud; SQL streams only what’s needed.
- Raster operations become declarative, not a tangle of scripts.
- NumPy acts as zero‑copy compute, not a separate pipeline.
- Outputs are transparent: polygons, GeoParquet, GeoTIFF, all versionable.
- Performance is predictable: the entire raster run took ~12 minutes on one machine.
The distance between raw imagery and actionable insight is shrinking; not because machines are bigger, but because the abstractions are finally right.
Spatial SQL + cloud‑native rasters isn’t just a convenience. It’s the foundation for environmental analysis that is auditable, scalable, and operationally sane.
Related Thoughts
Perspectives sharing related architectures, models, and domain context.
The Shift to Cloud-Native Geospatial: Access, Scale, and the Open Source Ecosystem
Traditional GIS architectures were built around assumptions that no longer hold: you download data to your local...
The Convergence of Spatial SQL and Ontologies: Building a Semantic Backbone for Earth Observation
The geospatial industry is undergoing a quiet but profound shift. For decades, the standard workflow for analyzing...
Everyone's Thinking Spatially Even If They Don't Call It Geography
Matt Forrest shared a story that struck a chord with me. He talked about winning an atlas in a third grade contest and...