Data engineering · 8 min read
Space Insights: two ways to see a wildfire from orbit
One satellite sees the heat of a fire while it burns. Another sees the scar it leaves, days later, in light the eye cannot register. I built a platform that does both. The interesting part is what happens where they disagree.
This is the overview. Each pipeline has its own write-up underneath it, and they are where the real detail lives:
- Every fire on Earth, every morning covers the thermal pipeline. A quarter of a million detections a day, hexagonal indexing, the 7% bug I nearly shipped, and a public demo with no backend at all.
- What a fire leaves behind covers the optical pipeline. Reading satellite pixels without downloading scenes, why the change matters and the reflectance does not, and a model that is not allowed to overrule the arithmetic next to it.
The question that produced two pipelines
I wanted to build something on satellite data that was not a tutorial. The tutorial version of this project is one pipeline: pull a feed, clean it, put it on a map. It works, it teaches you the tooling, and it proves nothing, because nothing in it can be wrong in an interesting way.
What makes fire data interesting is that there are two completely different ways to observe the same event, and they fail at different times.
| Thermal | Optical | |
|---|---|---|
| What it measures | Energy radiating right now | How much living vegetation is gone |
| Instrument | VIIRS on three polar orbiters | Sentinel-2 A and B |
| Answers | Where is something hot, everywhere, today | How badly did this ground burn |
| Timing | During the fire | On the next cloud-free pass |
| Misses | Fires that start and end between overpasses | Anything under cloud, for as long as the cloud lasts |
| Lies about | Gas flares, volcanoes, industrial heat | Harvest, drought stress, cloud shadow |
Look at the last two rows. Their blind spots do not overlap and their false positives have nothing in common. A gas flare is hot forever but never loses vegetation. A harvested field loses vegetation but was never hot. So where both signals fire on the same ground within a few days, that is real corroboration, not one measurement counted twice.
That is the reason to build two pipelines instead of one, and it is the reason the project has a cross-validation model joining them rather than two disconnected dashboards.
The constraint that shaped everything
The whole platform runs on Databricks Free Edition, which is serverless only, single workspace, and restricts outbound network access to an allowlist. Neither NASA nor Copernicus is on that allowlist.
There is no way to argue with this, so the architecture absorbs it: anything that touches the open internet runs outside the lakehouse as a scheduled GitHub Actions job and pushes data inward through the Files API and the SQL warehouse. I call these ground stations, and the term is used consistently throughout the codebase.
A constraint that turned out to be a feature
The two ground stations have almost nothing in common beyond that shape. One downloads three CSVs and merges them in memory. The other issues HTTP range requests into gigabyte-scale rasters to extract a few megabytes of pixels, encodes them as GeoTIFFs, and uploads them to a volume. Same pattern, wildly different content, which is exactly the test of whether a pattern is real.
One architecture, two very different shapes of data
The temptation with mixed sources is to let each one invent its own architecture. Both pipelines here land in the same medallion, in the same Unity Catalog, under the same rules:
space_insights (prd) space_insights_dev (dev) firms_bronze / silver / gold same schemas, same models wildfire_bronze / silver / gold / ml ops
- Bronze is owned by whoever wrote it: the ground stations, not dbt. Both use delete-then-insert for idempotency, keyed by whatever their natural unit of rework is: a date range for the thermal feed, a scene for the optical one.
- Everything tabular after the first Delta table is dbt, in one project shared by both pipelines and tag-isolated so each job builds only its own models. File parsing and raster maths stay in Python wheel tasks, because they are not SQL-shaped and pretending otherwise produces unreadable SQL.
- Business logic is pure and unit tested: spectral index maths, the analysis grid, the pixel-to-cell reduction, the anomaly model. No Spark, no network, no workspace. Job entry points stay thin: I/O and orchestration only.
- No click-ops. Terraform for schemas, volumes and secret scopes; Databricks Asset Bundles for jobs, experiments, registered models and apps. Push to main deploys production.
The medallion boundaries earn their keep in one specific way here. The optical pipeline's bronze layer holds raw GeoTIFF patches in a volume, and those get deleted after sixty days. They are scratch, and Copernicus can always re-serve them. Silver and gold are Delta and are kept. Being explicit about which layer is disposable is the difference between a retention policy and losing data.
The daily rhythm
06:00 UTC ground station - Copernicus scenes into bronze 06:15 UTC ground station - NASA detections into bronze 07:15 UTC firms_transform dbt: silver, hexes, gold 07:45 UTC publish static globe + tile archive 08:00 UTC wildfire_transform pixels -> cells -> dNBR -> score -> prune
The ordering is load-bearing in two places. The two-hour gap between ingest and transform absorbs a slow or retried satellite fetch without anything breaking. And the optical transform runs after the thermal one specifically so that when the cross-validation view is rebuilt, both of its sides are fresh; a table there would always be showing one run of stale agreement, which is why it is a view.
Where the machine learning actually is
Only the optical pipeline has a model, and it is worth being clear about what it does and does not do, because "ML for wildfire detection" usually means something more grandiose than this.
Burn severity itself is arithmetic: a published set of thresholds on how much the near-infrared to shortwave-infrared ratio dropped between two observations. It needs no model, no training data, and it is auditable by anyone who knows the literature. That detector shipped first and it still runs.
The model sits beside it and answers a question the thresholds structurally cannot: is this observation unusual across all its dimensions at once: the drop, the state it landed in, the vegetation index, how long since the last clear view, and how much of the cell was actually visible? An IsolationForest, trained unsupervised, because per-cell burnt/not-burnt ground truth does not exist and hand-labelling it from dNBR would just be laundering the thresholds back into the model.
On the map, the fill colour is always the measurement and the outline is the model. It annotates; it never repaints. And when no trained model exists yet, scoring is skipped and the pipeline degrades to the arithmetic instead of failing, which is only possible because the rule was never a placeholder for the model. The full argument is in the optical write-up.
What it costs
Nothing. Databricks Free Edition, free NASA and Copernicus accounts, the GitHub Actions free tier, Cloudflare Pages and R2 free tiers for the public globe.
I did not set out to build it for zero. But working inside a platform that refuses outbound network calls, cannot create catalogs through its own Terraform provider, and will not let you make an app public, forced a decision on every one of those points instead of letting me spend my way past it. The separation between ingestion and transformation, the precomputed public site with no backend behind it, the explicit retention policy on raw pixels. All of them started as workarounds and all of them are things I would keep with a budget.
Read on
- Every fire on Earth, every morning covers thermal detection at planetary scale: hexagonal indexing, a feed that rewrites its own past, and a public globe with nothing behind it.
- What a fire leaves behind covers optical burn-scar detection: byte-range reads into gigabyte rasters, cloud shadow as the real adversary, and an unsupervised model kept firmly in its place.