Opening this here rather than in datasets because the decisions that remain are about how this repo generates and records builds. Would value your read on the audit scope in particular.
What changed upstream
cng-datasets 0.8.0 stamps every published parquet with the version that produced it, in the file's own key-value metadata rather than a sidecar, so it survives copies and syncs:
SELECT key, value FROM parquet_kv_metadata('s3://bucket/dataset/hex/h0=*/data_0.parquet');
-- cng_datasets_version 0.8.0
-- built_at 2026-09-21T23:06:45Z
-- gdal_version 3.13.0
-- duckdb_version 1.5.4
-- image (when CNG_IMAGE is set)
Covers the hex partitions, the merged data_0.parquet, the vector repartition output and the GeoParquet.
Why it was needed, and why it is sharper than it sounds
Generated manifests pin ghcr.io/boettiger-lab/datasets:latest, and that image is rebuilt on every push to main. So the code that produced a dataset was not necessarily a released version — it was whatever main happened to be when the pod pulled the image. Nothing recorded which.
Separately: every version-tagged image build had been failing since v0.5.0 (an empty {{branch}} in the tag template produced an invalid reference, which buildx rejected before pushing anything). Until today there were no versioned images at all, so pinning was not even possible. That is fixed, and :0.8.0, :0.8 and :latest all resolve now.
The audit
Two correctness bugs were fixed in 0.7.0 (2026-09-20). Both are undetectable from the output alone — right row count, right schema, right value range, job exits 0:
-
Multi-band sources silently read band 1 (datasets#214). The hex path passed the source to exactextract with no band index and labelled the result with --value-column regardless. This published 384,922,346 rows of annual forb & grass cover documented as perennial cover — two variables in the same 0–100 domain.
-
--h0-subset takes grid positions, not H3 base cell numbers (datasets#213, #218). Both numberings run 0–121, so a list computed the obvious way — from the H3 library — is accepted, and builds a different part of the world. In the measured case that was ~80% of a western-US dataset missing, with the job reporting success.
A dataset whose parquet carries no cng_datasets_version was built before 0.8.0 and falls in scope. Concretely, worth checking:
- any hex build from a multi-band source (bug 1 applies only to those);
- any build that passed
--h0-subset / --h0-cells where the list came from the H3 library rather than the grid's own i column;
- ingests from roughly 2026-09-12 to 2026-09-20, which is when
:latest carried 0.6.x.
Everything else from that window is a performance/reliability story rather than a correctness one — evictions, OOMs and slow runs, all of which failed rather than publishing bad data.
What I'd value your view on
- Is the audit list above the right scope? You will know better than I do which recent ingests used multi-band sources or a hand-computed
--h0-subset.
- Should generated manifests pin
:0.8.0 rather than :latest? Now possible. It makes a build reproducible and a rollback real, at the cost of updates no longer arriving automatically — a build would pick up a new version only on regeneration. That tradeoff is yours more than ours.
- Should the version reach STAC as well as the parquet? The parquet stamp travels with the data, but a catalogue consumer will not open the file to find it.
- Anything else worth stamping? The completion markers are the obvious next candidate, so a partly-failed or gap-filled fan-out shows whether it was assembled from more than one build.
Happy to implement any of (2)–(4) upstream; (1) is the one I cannot do from here.
Opening this here rather than in
datasetsbecause the decisions that remain are about how this repo generates and records builds. Would value your read on the audit scope in particular.What changed upstream
cng-datasets0.8.0 stamps every published parquet with the version that produced it, in the file's own key-value metadata rather than a sidecar, so it survives copies and syncs:Covers the hex partitions, the merged
data_0.parquet, the vector repartition output and the GeoParquet.Why it was needed, and why it is sharper than it sounds
Generated manifests pin
ghcr.io/boettiger-lab/datasets:latest, and that image is rebuilt on every push tomain. So the code that produced a dataset was not necessarily a released version — it was whatevermainhappened to be when the pod pulled the image. Nothing recorded which.Separately: every version-tagged image build had been failing since v0.5.0 (an empty
{{branch}}in the tag template produced an invalid reference, which buildx rejected before pushing anything). Until today there were no versioned images at all, so pinning was not even possible. That is fixed, and:0.8.0,:0.8and:latestall resolve now.The audit
Two correctness bugs were fixed in 0.7.0 (2026-09-20). Both are undetectable from the output alone — right row count, right schema, right value range, job exits 0:
Multi-band sources silently read band 1 (datasets#214). The hex path passed the source to exactextract with no band index and labelled the result with
--value-columnregardless. This published 384,922,346 rows of annual forb & grass cover documented as perennial cover — two variables in the same 0–100 domain.--h0-subsettakes grid positions, not H3 base cell numbers (datasets#213, #218). Both numberings run 0–121, so a list computed the obvious way — from the H3 library — is accepted, and builds a different part of the world. In the measured case that was ~80% of a western-US dataset missing, with the job reporting success.A dataset whose parquet carries no
cng_datasets_versionwas built before 0.8.0 and falls in scope. Concretely, worth checking:--h0-subset/--h0-cellswhere the list came from the H3 library rather than the grid's ownicolumn;:latestcarried 0.6.x.Everything else from that window is a performance/reliability story rather than a correctness one — evictions, OOMs and slow runs, all of which failed rather than publishing bad data.
What I'd value your view on
--h0-subset.:0.8.0rather than:latest? Now possible. It makes a build reproducible and a rollback real, at the cost of updates no longer arriving automatically — a build would pick up a new version only on regeneration. That tradeoff is yours more than ours.Happy to implement any of (2)–(4) upstream; (1) is the one I cannot do from here.