DataLoader is the imagery-acquisition front end for eMapR pipelines (LandTrendr, BULC-D, BugNet, foundation-model work, ...). A pipeline describes the imagery it needs in a small YAML config: provider, sensor, area, years/season, bands. DataLoader then finds it, downloads it, and writes plain GeoTIFFs plus a versioned, machine-readable manifest recording what was acquired, where it came from, and what was done to it. Downstream code reads the manifest and never needs provider-specific code or filename parsing.
config.yaml ──► DataLoader ──► provider (USGS ARD, Planetary Computer, Earth Search, ...)
│
└──► <output.dir>/ GeoTIFFs + manifest.json + items.jsonl + source metadata
│
└──► your pipeline
- One tool for small pulls and big archives. A one-month AOI and a 35-year, 23-tile Landsat archive use the same command and produce the same output layout. The difference is only in the config.
- Resumable and updatable. If a run is interrupted, run the same command again and it picks up where it stopped. Leave the end year open and re-running later adds the new imagery.
- Provider-independent. Change
provider:to switch sources. Before downloading, DataLoader checks that the provider supplies the scientific product you asked for. - Nothing hidden. Cloud masking is off unless you ask for it. QA bands can be kept, source metadata is saved verbatim, and failures are recorded rather than silently dropped.
Python 3.9 or newer. Everything installs from pip; GDAL comes bundled with
the rasterio wheel.
git clone https://github.com/eMapR/data-loader.git
cd data-loader
python3 -m venv .venv && source .venv/bin/activate
pip install -e . # add '.[gee]' for Google Earth Engine, '.[dev]' for tests
data-loader --versionexamples/quickstart.yaml pulls every Landsat scene over a small Oregon AOI
for July 2023 from Microsoft Planetary Computer. It needs no account.
version: 1
provider: planetary_computer
sensors: [landsat]
aoi:
upper_left: [-122.4327, 44.2860] # [lon, lat]
lower_right: [-122.3426, 44.2315]
time:
start_date: 2023-07-01
end_date: 2023-07-31
temporal_mode: scene # or annual_composite
output:
dir: output/quickstart
bands: [red, nir]
indices: [ndvi]
qa_band: truedata-loader validate examples/quickstart.yaml # checks the config, no network
data-loader run examples/quickstart.yaml # ~1 minute; prints the manifest path
data-loader status output/quickstartoutput/quickstart/
manifest.json what this dataset is: request, product, bands, grids, processing, counts
items.jsonl one line per acquisition: date, sensor, source ID, status, files, checksums, provenance
request.yaml the exact request that produced it
metadata/<provider>/ each source item's original metadata record, verbatim
landsat/aoi/2023/2023-07-04_LE07_L2SP_045029_20230704_02_T1_bands.tif
landsat/aoi/2023/2023-07-04_LE07_L2SP_045029_20230704_02_T1_indices.tif
...
Reading it from Python:
from data_loader import open_dataset
import rasterio
ds = open_dataset("output/quickstart")
print(ds.manifest["bandSets"]["landsat"]["bands"]) # band order, dtype, scale/offset, nodata, QA meaning
for path, item, f in ds.files(start="2023-07-10"):
with rasterio.open(path) as src:
red = src.read(1)
print(item["date"], item["itemId"], item["eo:cloud_cover"])Adapt the example: edit the corners, dates, and bands, and point
output.dir (or --output-dir) somewhere new.
| Provider | Data | Account | Use it for |
|---|---|---|---|
usgs_ard |
Landsat 4–9 C2 U.S. ARD SR, straight from USGS | USGS EROS + M2M token | Preferred for Landsat time series and archives. USGS is the authoritative source, it covers Landsat 4–9, and every observation sits on a fixed 5000×5000 px Albers tile grid. |
planetary_computer |
Landsat 5–9 C2 L2 and Sentinel-2 L2A | none | Quick AOI pulls and Sentinel-2 with no setup. |
aws_earth_search |
Sentinel-2 L2A (free); Landsat (AWS Requester Pays) | none for Sentinel-2 | Sentinel-2. Its Landsat bucket bills an AWS account. |
gee |
Landsat C2 L2, Sentinel-2 SR (harmonized) | Earth Engine project | Exploratory small AOIs. Requests are capped at ~48 MB each. |
glad_ard |
GLAD 16-day Landsat ARD, 2020+ | none | GLAD's normalized 16-day composites specifically. |
usgs_m2m |
Landsat C2 L2 scene bundles | USGS EROS + M2M token | Experimental: the bundle download path has not been verified end to end. |
The full comparison, with caveats such as DataLoader's Planetary Computer Landsat path covering Landsat 5–9 only and how Sentinel-2 processing versions are handled, is in docs/providers.md. The USGS account setup is in docs/getting-started.md.
For a persistent archive, request whole USGS ARD tiles. DataLoader then stores each observation exactly as USGS distributes it: uint16 DN with QA_PIXEL, no cloud masking, and no resampling. Downstream pipelines make their own cloud, shadow and snow decisions.
aoi:
tiles: [h003v004, h003v005]
time:
start_year: 1990 # no end_year: through today, updatable
temporal_mode: scene
output:
encoding: native
bands: [blue, green, red, nir, swir1, swir2]
qa_band: trueSee examples/usgs_ard_tile_archive.yaml and eMapR's Oregon archive
examples/oregon_landsat_ard_archive.yaml, along with
docs/large-jobs.md for runtime, storage, tmux, resume and updates. The
Oregon archive is about 89,000 observations, takes about 3–4 weeks on one USGS
account, and needs roughly 5–6 TB.
| Getting started | Install, credentials for each provider, first runs |
| Configuration | Every config key; years, seasons and dates; AOIs and tiles; migrating old configs |
| Providers | What each provider serves, accounts and costs, when to choose which |
| Outputs | Directory layout, manifest.json and items.jsonl field by field, reading data, masking with QA |
| Large jobs | Archives, runtime and storage, resume, updates, failures |
| Troubleshooting | Common errors and fixes |
| Development | Architecture, adding a provider, tests, benchmarks and development history |
Example configs live in examples/. Run the tests with
pytest (they need no network or credentials).
Version 1.0. The config schema (version: 1) and the output manifest
(dataloader-manifest 1.0) are the stable interfaces; see
CHANGELOG.md. Planned next: vector (GeoJSON) AOIs, and
Sentinel-2 directly from the Copernicus Data Space Ecosystem.