The Shift to Cloud-Native Geospatial: Access, Scale, and the Open Source Ecosystem
Traditional GIS architectures were built around assumptions that no longer hold: you download data to your local machine before working with it. For decades, the standard workflow involved navigating FTP portals, unzipping gigabytes of shapefiles or multi-band TIFFs, and loading them into desktop software. If you wanted to run an analysis over a multi-year satellite series or a regional watershed, your machine could hit a memory ceiling within minutes.
Cloud-native geospatial fundamentally inverts this workflow. Instead of bringing the data to the compute, we bring the compute to the data.
Through standard file specifications, catalog APIs, and modern toolchains, the geospatial ecosystem has evolved into a composable data science stack. Much of this democratization has been catalyzed by open source software, notably through the work of Qiusheng Wu and the broader open geospatial community.
Here is a breakdown of how the modern cloud-native stack fits together, why it matters, and how to build workflows on top of it.
1. The Core Formats: COG, STAC, and GeoParquet
The foundation of cloud-native geospatial relies on three open standards designed specifically for HTTP range requests and serverless querying.
- Cloud Optimized GeoTIFF (COG): A COG is a standard TIFF file organized internally with tiling and downsampled overviews (pyramids). When hosted on an S3 bucket or Google Cloud Storage, a client does not need to download the full 500 MB raster to view or analyze a single bounding box. By using HTTP GET range requests, the client requests only the precise byte offsets required for the current view extent or analysis resolution.
- SpatioTemporal Asset Catalog (STAC): While COGs solve the raster storage problem, discovering millions of scenes across global archives requires standardized metadata. STAC provides a JSON-based specification for indexing geospatial assets across space and time. Instead of maintaining proprietary catalog databases, public archives (USGS Landsat, ESA Sentinel, NOAA, Planet) expose searchable STAC APIs.
- GeoParquet: Vector workflows have historically lagged behind rasters in cloud efficiency, relying on shapefiles, GeoJSON, or bulky database exports. GeoParquet brings Apache Parquet's columnar compression, dictionary encoding, and fast partition pruning to geometries. By storing bounding box metadata in Parquet file footers, engines can skip entire files or row groups that do not intersect a query polygon.
2. Bridging the Gap: The Open Source Python Ecosystem
Open standards require accessible software to become useful. The open-source geospatial community, with contributors like Qiusheng Wu, has spent recent years building bridges between raw cloud assets and the Python data science environment.
GeoLibre and the Evolution Beyond Leafmap
While geemap opened Google Earth Engine to Jupyter and leafmap served as an essential unified mapping package for years, the cloud-native ecosystem has transitioned to GeoLibre.

Developed by Qiusheng Wu and the opengeos community, GeoLibre represents the next evolutionary step: a lightweight, cloud-native GIS platform that runs across desktop, browser, and Jupyter environments. Rather than treating the notebook as a simple tile viewer, GeoLibre’s Python package (geolibre) embeds a modern, high-performance GIS interface (built on MapLibre GL, WebAssembly, and deck.gl) directly into a notebook cell via an anywidget bridge while maintaining familiar, leafmap-style ergonomics.
GeoLibre connects directly to cloud-native formats. You can stream remote Cloud Optimized GeoTIFFs, query STAC collections, or render massive GeoParquet files without maintaining dedicated spatial middleware or downloading local raster files:
import geolibre
m = geolibre.Map()
# Stream a remote Cloud Optimized GeoTIFF directly via range requests
cog_url = "[https://opendata.digitalglobe.com/events/mauritius-oil-spill/post-event/2020-08-12/105001001A085400/105001001A085400.tif](https://opendata.digitalglobe.com/events/mauritius-oil-spill/post-event/2020-08-12/105001001A085400/105001001A085400.tif)"
m.add_cog_layer(cog_url, name="Remote Satellite Scene")
# Display the interactive map
m
In-Process Spatial SQL: DuckDB Spatial
One of the most consequential advancements in recent geospatial workflows is pairing DuckDB with its spatial extension.
You no longer need to spin up a PostGIS instance or configure an enterprise database server just to run spatial joins across vector layers. DuckDB can execute spatial predicates directly against remote Parquet files stored on object storage, streaming only the relevant byte ranges over HTTP:
INSTALL spatial;
LOAD spatial;
-- Query remote GeoParquet directly without downloading the file
SELECT
name,
ST_Area(ST_GeomFromWKB(geometry)) AS footprint_area
FROM read_parquet('s3://my-spatial-bucket/building_footprints.parquet')
WHERE ST_Intersects(
ST_GeomFromWKB(geometry),
ST_Point(-77.0369, 38.9072)
);
High-Performance Vector Visualization: Lonboard
Rendering millions of vector vertices in an interactive web browser used to freeze the DOM. Through libraries like lonboard (built on top of geoarrow and deck.gl), Python data scientists can now push millions of points, linestrings, and polygons directly to GPU memory inside Jupyter notebooks with near-zero serialization latency.
3. Why This Architectural Shift Matters
The move toward cloud-native geospatial is not merely a change in tooling; it redefines operational economics and team capabilities:
- Lower Infrastructure Costs: You eliminate the requirement for persistent, always-on spatial database clusters for exploratory data analysis. Serverless queries and object storage cost pennies compared to running dedicated compute instances.
- Reproducibility: An analysis script can reference public STAC records and remote COGs directly. Anyone with internet access can execute the code without downloading gigabytes of prerequisite assets.
- Decoupled Compute and Storage: Teams can run ad-hoc transformations using DuckDB locally, scale up to distributed Dask or Apache Sedona clusters when data volumes expand, and write results back to GeoParquet without altering underlying storage schemas.
4. The Path Forward
The cloud-native geospatial ecosystem is maturing quickly. As standards like GeoParquet reach widespread adoption and browser engines leverage WebAssembly (Wasm) and WebGPU for client-side spatial compute, the line between data engineering, spatial analysis, and desktop GIS will continue to blur.
For practitioners looking to modernize their spatial pipelines, the starting point is straightforward: stop downloading files. Start cataloging assets in STAC, convert raster archives to COG, store vector features in GeoParquet, and use modern cloud-native tools like GeoLibre and DuckDB to query data where it lives.
Related Thoughts
Perspectives sharing related architectures, models, and domain context.
The Convergence of Spatial SQL and Ontologies: Building a Semantic Backbone for Earth Observation
The geospatial industry is undergoing a quiet but profound shift. For decades, the standard workflow for analyzing...
Spatial SQL, Cloud-Native Imagery, and Analysis at Scale
One of the most important shifts happening in geospatial right now is the collapse of the old “download → preprocess →...
Everyone's Thinking Spatially Even If They Don't Call It Geography
Matt Forrest shared a story that struck a chord with me. He talked about winning an atlas in a third grade contest and...