All posts

Data engineering · 8 min read

Space Insights: two ways to see a wildfire from orbit

One satellite sees the heat of a fire while it burns. Another sees the scar it leaves, days later, in light the eye cannot register. I built a platform that does both. The interesting part is what happens where they disagree.

This is the overview. Each pipeline has its own write-up underneath it, and they are where the real detail lives:

servingpublic globePMTiles on R2, no loginDatabricks appglobe · earth · marsAI/BI dashboardsfires · burn scarsNASA FIRMSVIIRS, globalground stationGitHub Actionsbronzeraw detectionssilvercleaned, H3 indexedgoldhex aggregatesCopernicusSentinel-2 L2Aground stationbyte-range readsbronzeGeoTIFF patchessilverper-cell indicesgolddNBR, burn eventss2_firms_agreementthe cross-checkthermal: where something is burning nowoptical: what the fire actually destroyed
The whole platform on one page. Two sources with nothing in common, two ground stations outside the lakehouse, one medallion, and a single model at the bottom of both gold layers whose only job is to ask whether the two halves agree. The serving layer is shared rather than split per pipeline: one Databricks app carries both maps as tabs, the dashboards carry both pipelines as pages, and the public globe exists only because Databricks Apps cannot be opened to anonymous visitors.

The question that produced two pipelines

I wanted to build something on satellite data that was not a tutorial. The tutorial version of this project is one pipeline: pull a feed, clean it, put it on a map. It works, it teaches you the tooling, and it proves nothing, because nothing in it can be wrong in an interesting way.

What makes fire data interesting is that there are two completely different ways to observe the same event, and they fail at different times.

ThermalOptical
What it measuresEnergy radiating right nowHow much living vegetation is gone
InstrumentVIIRS on three polar orbitersSentinel-2 A and B
AnswersWhere is something hot, everywhere, todayHow badly did this ground burn
TimingDuring the fireOn the next cloud-free pass
MissesFires that start and end between overpassesAnything under cloud, for as long as the cloud lasts
Lies aboutGas flares, volcanoes, industrial heatHarvest, drought stress, cloud shadow

Look at the last two rows. Their blind spots do not overlap and their false positives have nothing in common. A gas flare is hot forever but never loses vegetation. A harvested field loses vegetation but was never hot. So where both signals fire on the same ground within a few days, that is real corroboration, not one measurement counted twice.

That is the reason to build two pipelines instead of one, and it is the reason the project has a cross-validation model joining them rather than two disconnected dashboards.

The constraint that shaped everything

The whole platform runs on Databricks Free Edition, which is serverless only, single workspace, and restricts outbound network access to an allowlist. Neither NASA nor Copernicus is on that allowlist.

There is no way to argue with this, so the architecture absorbs it: anything that touches the open internet runs outside the lakehouse as a scheduled GitHub Actions job and pushes data inward through the Files API and the SQL warehouse. I call these ground stations, and the term is used consistently throughout the codebase.

A constraint that turned out to be a feature

Separating ingestion from transformation was forced on me, but it is what I would have done anyway at scale. The network-touching code has its own failure domain, its own retry behaviour and its own schedule, and none of it can take the lakehouse down with it. A CDSE outage delays data; it does not fail a job in the warehouse.

The two ground stations have almost nothing in common beyond that shape. One downloads three CSVs and merges them in memory. The other issues HTTP range requests into gigabyte-scale rasters to extract a few megabytes of pixels, encodes them as GeoTIFFs, and uploads them to a volume. Same pattern, wildly different content, which is exactly the test of whether a pattern is real.

One architecture, two very different shapes of data

The temptation with mixed sources is to let each one invent its own architecture. Both pipelines here land in the same medallion, in the same Unity Catalog, under the same rules:

space_insights          (prd)          space_insights_dev     (dev)
  firms_bronze / silver / gold           same schemas, same models
  wildfire_bronze / silver / gold / ml
  ops

The medallion boundaries earn their keep in one specific way here. The optical pipeline's bronze layer holds raw GeoTIFF patches in a volume, and those get deleted after sixty days. They are scratch, and Copernicus can always re-serve them. Silver and gold are Delta and are kept. Being explicit about which layer is disposable is the difference between a retention policy and losing data.

The daily rhythm

06:00 UTC   ground station - Copernicus scenes into bronze
06:15 UTC   ground station - NASA detections into bronze
07:15 UTC   firms_transform      dbt: silver, hexes, gold
07:45 UTC   publish              static globe + tile archive
08:00 UTC   wildfire_transform   pixels -> cells -> dNBR -> score -> prune

The ordering is load-bearing in two places. The two-hour gap between ingest and transform absorbs a slow or retried satellite fetch without anything breaking. And the optical transform runs after the thermal one specifically so that when the cross-validation view is rebuilt, both of its sides are fresh; a table there would always be showing one run of stale agreement, which is why it is a view.

Where the machine learning actually is

Only the optical pipeline has a model, and it is worth being clear about what it does and does not do, because "ML for wildfire detection" usually means something more grandiose than this.

Burn severity itself is arithmetic: a published set of thresholds on how much the near-infrared to shortwave-infrared ratio dropped between two observations. It needs no model, no training data, and it is auditable by anyone who knows the literature. That detector shipped first and it still runs.

The model sits beside it and answers a question the thresholds structurally cannot: is this observation unusual across all its dimensions at once: the drop, the state it landed in, the vegetation index, how long since the last clear view, and how much of the cell was actually visible? An IsolationForest, trained unsupervised, because per-cell burnt/not-burnt ground truth does not exist and hand-labelling it from dNBR would just be laundering the thresholds back into the model.

On the map, the fill colour is always the measurement and the outline is the model. It annotates; it never repaints. And when no trained model exists yet, scoring is skipped and the pipeline degrades to the arithmetic instead of failing, which is only possible because the rule was never a placeholder for the model. The full argument is in the optical write-up.

What it costs

Nothing. Databricks Free Edition, free NASA and Copernicus accounts, the GitHub Actions free tier, Cloudflare Pages and R2 free tiers for the public globe.

I did not set out to build it for zero. But working inside a platform that refuses outbound network calls, cannot create catalogs through its own Terraform provider, and will not let you make an app public, forced a decision on every one of those points instead of letting me spend my way past it. The separation between ingestion and transformation, the precomputed public site with no backend behind it, the explicit retention policy on raw pixels. All of them started as workarounds and all of them are things I would keep with a budget.

Read on