Fix spatial analysis bugs from adversarial review - #9
Merged
Merged
Conversation
…est overlap assign_features_to_polygons() wrapped st_join(largest = TRUE) in a tryCatch that reran the join without `largest` on any error. The comment blamed a predicate that does not support `largest`, but sf never calls the predicate on that path; the retry only fired when GEOS threw on an invalid ring. One bow-tie parcel then reassigned the whole layer by tie_break: 95 of 200 buffered parcels changed cell, and a polygon 91% inside cell 3 went to cell 2, with no R warning. Invalid features and cells are now repaired with st_make_valid() for the join only, with a warning that counts them, and the features come back with the geometry they arrived with. If the largest-overlap join still fails the function stops and names largest = FALSE instead of silently changing the rule. The docs no longer claim `largest` is dropped for some predicates and say that `predicate` is not used when `largest` applies. Adds test-review-assignment.R: an invalid feature, an invalid cell and a join that still fails. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…with no data kriging_adequacy() divided gstat's block-kriging variance, the variance of a cell MEAN, by the point sill. With no data a cell mean's variance levels off at C(B,B), the covariance averaged over the cell, which is far below the sill once the cell is about as wide as the range. So kr_ratio never approached 1: empty cells 130-410 m beyond a 90 m effective range read 0.09-0.1, print() reported "0 cell(s) above 0.5" with half the cells pure prior, and with unequal cells (Voronoi, Delaunay) a large empty cell ranked below small populated ones. kr_ratio is now kr_var over each cell's own C(B,B), computed on the discretisation gstat block-kriges with and without the nugget, as gstat does; it agrees with gstat's own value to 1e-7. A cell the data do not reach reads 1 (ordinary kriging adds the variance of the estimated mean, so it sits at or above its prior and is clamped), and a pure-nugget model, which gives a cell mean no prior variance, gives NA. The roxygen, the Rd and the print() label say what the ratio now is. The extra cost is one pass over each cell's discretisation, about as long as the krige() call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
prep_model_data() coerced geometry to POINT but kept any Z or M coordinate, and the sf-to-sp coercion GWmodel needs keeps it too. On POINT Z data (GPS, GeoPackage, KML altitude) predict() reshaped XYZ prediction locations with matrix(, ncol = 2) and returned values for the wrong places without a warning, even for a constant altitude of 0; gwr_model_selection() ranked models on 3-D distances; fit_gwr_model() failed below 2500 rows and fitted on 3-D distances above; cv_gwr() lost every fold. prep_model_data() now drops Z and M, as make_folds() and estimate_sac_range() already did, and .to_sp() does the same for callers that skip preparation (fit_gwr_model(.already_prepped = TRUE)) and for fits made before this change. The test looks at the coordinate count, since sf does not set z_range on points built from three coordinate columns, and skips the per-feature st_zm() rebuild for 2-D layers. The random-forest and Bayesian backends already read only the first two coordinate columns and give the same results. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
GWmodel is built with OpenMP, and GNU libgomp is not fork-safe: once any GWR had run in the session (fit_gwr_model(), a sequential cv_gwr(), a bare GWmodel::bw.gwr()), every mclapply() worker blocked on a futex at its first OpenMP region and cv_gwr(parallel = n) -- or compare_models_cv() with gwr_args = list(parallel = n) -- never returned. That is the ordinary fit-then-cross-validate order. cv_gwr() now runs its folds one after another whatever `parallel` says, and warns when more than one core was asked for. A PSOCK cluster was rejected: its workers need the same spatialkit installed, which pkgload::load_all() does not give them, and they cannot see what a `metrics` closure reads from the global environment. The other cv_*() keep mclapply(). The docs of cv_gwr() and cv_spatial() and the README say so, and a regression test runs fit_gwr_model() then cv_gwr(parallel = 2) in a child R under `timeout`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…aw response resolution_profile() fitted the OLS detrend on every row first and only refit on the complete rows if that succeeded. lm.fit() refuses any NA, so one missing response or predictor value made the first fit fail and the profile fell back to the raw response, trend and all, while the variogram it was compared against had been fitted to the residuals: RSS, Cp and Moran's z described a different variable than the nugget and floor. The fit now runs on the complete rows only, the rows it leaves out are counted in a logged warning, and they stay in the geometry as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
assign_features_to_polygons() harmonised CRS towards the features, so
lon/lat features pulled the package's projected cells into lon/lat. Every
cell edge became a great-circle arc (s2) or a lon/lat straight line, and
with s2 the largest-overlap join failed on a degenerate intersection piece
and fell back to the tie-break: sf's nc counties against a 36-cell grid
put 55 of 100 counties outside their largest-overlap cell. The same path
crashed lon/lat points against some hex grids ('Loop 3 is not valid'),
needed lwgeom with s2 off, and moved points near an edge into the
neighbouring cell.
The join now runs in the polygons' CRS whenever it is projected, on a
transformed copy of the features, and the rows come back with the
geometry they arrived with instead of a transform round trip. Polygons in
lon/lat or without a CRS are joined as before. On the nc repro 0 of 100
counties are misassigned, and 50,000 lon/lat points match the projected
join exactly, with s2 on or off.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
When some folds, not all, failed to fit or predict, cv_spatial(), cv_rf(), cv_gwr() and cv_bayes() wrote a logger line and pooled `overall` over the surviving folds, so tryCatch(), expect_warning() and options(warn = 2) never saw that the headline score had left out a block -- usually the hardest one, such as the only rows of a factor level (164 of 200 rows scored, no condition raised). They now raise an R warning that names each failed fold and its status and says how many of the rows `overall` covers. Folds dropped before fitting are left out of it, since .remap_folds() already warns about those, so no fold is warned about twice. cv_spatial()'s documentation says so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
… NA in silence gstat gives two observations at the same coordinates the full sill, nugget included, as their covariance, so repeat visits to a station made every kriging system holding them singular. gstat answered with NA and, at debug.level 0, nothing else: 80 stations x 3 campaigns gave NA in all 16 cells and no cross-validation prediction, with no R warning, and print() still said every empty cell had a kriged estimate. kriging_adequacy() now kriges one observation per location, the mean of its replicates, and warns; n and mean still count every point. The replicates' pooled within-location variance, capped at the nugget, is the part of the nugget that differs between visits, so a mean of m replicates carries c0 - s_w^2 + s_w^2/m, passed to gstat as a known measurement error (weights) on the model with its nugget zeroed. That is exactly the kriging of every observation when replicates differ by measurement error alone, and of one observation when they are identical; keeping the full nugget on every mean instead pulled the CV z-score variance down to about 0.7 on simulated revisits. A location seen once is kriged as before. The cross-validation holds a location out whole and runs the folds with gstat::krige(), since krige.cv() does not subset weights; on distinct points it matches krige.cv() to 15 digits. A cell or held-out location gstat still cannot krige now raises a warning. The result carries n_locations, and print() counts the cells with no estimate, says how many empty cells have one, and gives the number of distinct locations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
coerce_to_points(mode = "auto") segfaulted the R session on an empty MULTILINESTRING. st_cast() splits one into a single empty LINESTRING, not zero parts, and st_line_sample() on an empty line crashes in sf 1.0.x. That is what a null geometry in a line layer read from a GeoPackage or a shapefile looks like, and prep_model_data() and the cross-validation helpers pointize before they drop empty rows, so the documented cleaning path took the session down. An empty LINESTRING was refused with an error instead, although every other geometry type already came back as an empty point. Empty lines, and empty parts of a MULTILINESTRING, no longer reach st_line_sample(). An empty feature becomes an empty POINT in its own row, and prep_model_data() and make_folds() drop it as they drop any other empty geometry. The regression test runs the crashing calls in a child R process, so a relapse fails one test instead of killing the run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
build_tessellation(method = "triangles") passed raw projected coordinates to qhull. qhull lifts each point onto x^2 + y^2, and at UTM magnitudes that lift loses the precision to separate points a few metres apart, so they were dropped as coplanar and never became vertices. 200 points over 100 m at (5e5, 5e6) gave 26 triangles instead of 386, and the package's own example setup 14 instead of 30. Every point still landed in some triangle, so nothing looked wrong. The coordinates are now centred on their bounding-box midpoint before triangulating, and the triangles are still built from the original coordinates. The midpoint does not depend on row order, unlike a mean, so a permuted layer gets the same triangles and cell_ids as before. qhull returns the triangles in a different order after centring, so triangle cell_id and $index values change even for inputs that were already triangulated correctly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
select_features_forward() read each set's pooled cv_spatial() score, which
covers only the folds and rows that set managed to predict. A set whose fit
or predict failed on a fold -- a factor level found in one spatial block
only, the ordinary case -- was scored on fewer, usually easier, rows and
could win for that alone: a noise land-cover factor beat the true driver
(RF RMSE 2.35 on 192 rows against 2.63 on 250), and with lm the factor was
added at step 2 because {a, lc} scored 0.94 on 192 rows against 2.33.
Every set is now scored on one reference row set: the rows the null model
predicted, or, when there is no null model (RF, GWR), the rows any step-1
set predicted. A set that leaves a reference row unpredicted is scored NA
with one warning per step naming it, and cv_spatial()'s own partial-fold
warning is not repeated for every candidate. history gains n_pred and
params gains n_scored; the documentation says how sets are scored.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
A lon/lat layer spanning more than 180 degrees of longitude with no large gap was treated as global coverage and projected to EPSG:3857 (Equal Earth for purpose = "area"). Latitude was never checked, so Antarctic stations and pan-Arctic networks, which circle a pole, got Web Mercator. It splits them at +/-180 and stretches them towards the pole: worst-case distance errors of 15,000-20,000% against about 2% for a pole-centred Lambert azimuthal, a 111 km pair measured as 445 km, and the South Pole at y = -2.4e8 m. Variogram ranges, block sizes, bandwidths and Voronoi assignment all read those distances. When every point lies on one side of the equator, the layer now gets a Lambert azimuthal equal-area projection centred on that pole. It is equal-area, so it serves purpose = "area" too. It is scored against the global fallback and used only when it measures lower, because a low-latitude belt that does not wrap is fitted better by Web Mercator. Both measured figures come back in crs_choice. Data spanning both hemispheres and data straddling the antimeridian behave as before, and the documentation no longer says that only truly global coverage falls back to EPSG:3857. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
compare_models_cv() put each backend's pooled `overall` row side by side and said the columns were comparable because the backends share folds. They share folds, not rows: a backend that loses a fold -- GWR with a fixed bandwidth across a gap in the data -- loses the hardest, extrapolation rows and looks better for it. In the repro GWR scored RMSE 1.81 on 158 rows against RF's 1.95 on 200, while RF scored 0.99 on the same 158 rows. When the backends' predicted ..row_id sets differ, the function now warns with each backend's row count and recomputes every backend's pooled metrics, user metrics included, on the rows all of them predicted, so n_pred agrees across the table. What each backend reported over its own rows stays in its *_cv element and in attr(overall, "all_rows"). by_fold and the Bayesian coverage and CRPS columns are per fold and are left as they were, as the warning and the documentation say. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
resolution_profile() subsamples large layers (sample_n, default 1500) for the k-means fits, and then took n from the subsample for everything else: the support ceiling floor(n / min_cell_n), the supported verdict and its warning, the Cp penalty and the reliability. Every layer larger than sample_n was capped at 166 cells by default, a layer that supports its range floor could be declared unsupported, and the answer moved with sample_n. The count was then handed to a tessellation of every point. The ceiling, the verdict, the area, the distinct locations and the reliability now come from the whole layer, and the print says so. The subsample only has to fit the cells: the ceiling is also held to two of its points per cell (ceiling_from = "sample_n", with a request to raise sample_n, which is not a verdict on the data). Cp is estimated for the whole-layer design, RSS/m + tau^2 L_m/m + tau^2 L/N: the subsample's RSS with the optimism of its own cell means added back, plus the variance of cell means built from all N points; with no subsample it is Mallows' RSS/n + 2 tau^2 L/n as before. bounds$n is now the layer's count and bounds$n_sample the subsample's; the test that pinned the subsample's ceiling is updated to the layer's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
predict.gwr_fit() called GWmodel::gwr.predict(), which returned every prediction as NA in three common cases: one location with an empty or singular window (a grid cell beyond a fixed bandwidth, a point inside a held-out block) made inv() throw for the whole call, so fixed-bandwidth block CV lost every fold; more than 10000 training plus new rows failed on 'DM3.given' not found, so a 100 x 100 grid predicted nothing; more than 5000 training rows failed for any newdata. It also built the n_train x n_train hat matrix for a prediction variance this method threw away, in cubic time, on every call and every cv_gwr() fold. Predictions now come from gwr.basic(regression.points = ) in chunks, with the fit's kernel, bandwidth and gw.dist() distances, and x'beta summed in the order GWmodel sums it. A chunk that throws is redone location by location with gw_reg_1(), the call gwr.predict() made per point, so only the windows that are really singular come back NA, with a warning counting them. Where gwr.predict() worked the values are identical to the bit, for all five kernels, fixed and adaptive, and predictions at the training points still equal fitted(). The design is built from the fitted terms and factor levels, so rows cannot shift against the coefficients; NA rows in newdata stay where they were. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
On brms 2.17 to 2.22, brms rebuilds the HSGP boundary L = c * max(1, pooled centred range) from the rows of every predict call. The two padding rows predict.bayesian_fit() appends could only widen that range, so a single newdata row past the training envelope widened L and shifted the prediction of every row in the call, interior ones included (by up to 0.2 on a response of SD 1.5 on brms 2.20.4), and predict_surface() depended on chunk_size whenever the grid reached past the sample bbox. Only an INFO line said so, and the docs claimed the boundary had to grow whatever was done. predict() now hands brms a copy of the fit whose gp() term has c * S_fit / S_new, so brms rebuilds the fitted L exactly whatever rows share the call. The training centre this needs is stored as $info$gp_cmeans; fits saved earlier read it off the brmsfit's data, as brms did. Rows further than L from that centre, where the basis is an odd reflection of the fitted surface, come back NA with a warning. brms >= 2.23.0 stores L itself, so there c is left alone. The docs and the stale "brms 2.x does not store L" comments now say what the code does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
With select_on = "split", determine_optimal_levels() (with model variables) and resolution_profile() ran entirely on the selection half: its coordinates, point count, hull area and clusters. The documented workflow then tessellates every point with that count. On four clusters at the corners of a square every spatial half holds two, and the split answered 2 where the layer has 4; on a three-cluster layer the profile's area, floor, ceiling and verdict described a hull a thirtieth the size of the layer's, and the count moved from 10 to 23 with the split seed. The split exists to keep the estimation half's response out of the choice, and the cells are geometry, so now only what reads the response is held to the selection half: the Moran's I pass in determine_optimal_levels(), and the OLS fit, the variogram, RSS, Cp and Moran's z in resolution_profile(). The k-means cells, the WSS curve and elbow, the bounds and the point counts behind Cp's variance term and the reliability come from every point, so the count is for the layer that gets tessellated and agrees with the whole-layer support of the previous change. The docs say which parts read which half, and the two select-on tests that pinned the half's count and bounds now pin the layer's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
GEOS 3.12.1 segfaults computing the interior point of a non-empty feature that holds an empty line: a MULTILINESTRING with an empty part beside a real one, or a GEOMETRYCOLLECTION with an empty LINESTRING among its members. .crs_distance_error() reduces every non-point lon/lat layer to such points, so plain ensure_projected() on a line layer like that took the R session down, and so did coerce_to_points(mode = "auto") at its default tmp_project = TRUE and coerce_to_points(mode = "point_on_surface"). tryCatch() cannot catch a crash. A collection with an empty POLYGON did not crash but got POINT EMPTY although it held a line. Every st_point_on_surface() call now goes through .drop_empty_parts() first. It removes the empty parts of multi-part features and collections and keeps every row. A feature with no empty part is returned exactly as it was, so the polygon-only callers (stable ids, plot labels, polygon rows in coerce_to_points) are unchanged. The regression test runs the crashing calls in a child R process. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Adds a Bug fixes entry for each fix from the adversarial review that touches code shipped in 2.0.0, and corrects the development version's kriging_adequacy() entry, whose ratio and handling of repeat visits changed before release. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Both elbows took the point furthest below the chord of the WSS curve on linear axes. Points with no cluster structure have WSS close to c/k, and on that curve the rule lands at sqrt(first x last level) whatever the data: determine_optimal_levels() answered 4, 4, 7, 9, 13 for max_levels 12 to 160 on the same uniform layer, the profile's elbow column followed min_cell_n, and build_tessellation(approx_n_cells = profile) handed that count on for any geometry-only profile. The same rule missed separated clusters: eight of them at the default max_levels = 12 came back as 3. The elbow is now read on log-log axes, where c/k and every power law is a straight line and separated clusters bend at the cluster count. It is the level whose log WSS sags furthest below the line joining k = 1 and the last level, and it counts only when the sag is at least log(1.25), the WSS a fifth below that power law: uniform layouts of many shapes stayed at or below 0.16, two to ten separated clusters at or above 0.6. With no elbow determine_optimal_levels() still returns the linear-axis answer, but with an R warning that the ladder chose it; the profile's elbow column is NA and its print says there is no elbow, so select_resolution(criterion = "elbow"), and a geometry-only profile passed to build_tessellation() or get_voronoi_seeds(), refuse with a message saying why. The help example still answers 2 1 3. An elongated extent bends at about its aspect ratio and that bend is reported; the docs say so. Existing tests that read the elbow off uniform fixtures now read it off clustered ones, or assert that there is none. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
The comment in determine_optimal_levels() named `in_sel`, the name the same mask has in resolution_profile(); here it is `moran_rows`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
The resolution vignette called the elbow defensible on a fixture whose uniform points have none, described an elbow panel the profile plot no longer draws for them, and said the split profile's ladder is shorter because the ceiling scales with n, which stopped being true once the cells are drawn on every point. The getting-started pipeline now warns that the ladder chose its count; the prose says so, since the vignette hides warnings. NEWS gets a Bug fixes entry for determine_optimal_levels()'s elbow, which shipped in 2.0.0, and corrects the development entries for resolution_profile() and select_on, which have not. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
The fold_separation() example, the README quick start, the resolution, diagnostics and spatial cross-validation vignettes and the fixture the tour scripts share gave EPSG:32632 to x and y between 0 and 1000, which is on the equator near 4.5E, outside the zone. They now add 5e5 east and 5e6 north, as the other examples already did. The package works in planar units, so every printed number in the README and the three vignettes is unchanged; only the coordinates and the map graticules move. Script 07 selected its western half with x < 500 and now uses 5e5 + 500. In script 05 the k = 16 Moran z moves from 6.08 to 6.07: one 16th/17th neighbour pair is 0.0005 apart, and the kNN tie tolerance scales with the size of the coordinates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…recheck Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
What the user saw: - fit_gwr_model() warned that local regressions were "unstable" because of where a predictor's units start: a temperature field in kelvin was flagged at 200 of 200 locations (plus a global "collinearity risk" warning, index 230) and its slope map was refused as "every location is masked", while the identical slopes in degrees C were not. The nc_demo vignette fit was flagged at 25 of 300. v2.0.0 did the same for two or more predictors. - gwr_model_selection() checked an adaptive bandwidth less strictly than fit_gwr_model() (3e9 gave a bare "missing value where TRUE/FALSE needed", 0.5 became "0 neighbours"); cv_gwr() failed the check in every fold. - gwr_model_selection() with a fixed bandwidth in the wrong units died on a bare "inv(): matrix is singular". - The undefined-AICc warning advised a larger bandwidth when the adaptive one was already every observation; the adaptive floor warning and docs claimed it always suffices, which ties at the kernel's edge (a regular grid) break. - print() showed a fixed bandwidth as "1.224e+05" with no unit. - fit_gwr_model() did not refuse a character/factor response as the README says; it ran a failing bandwidth search and stopped with an unrelated error. - Tour script 09 labelled non-finite-coefficient rows "locally singular". What changed: - The global condition index is computed on centred predictors (1 for a single predictor). The local survey keeps Belsley's uncentred index `cn` and adds `cn_slopes` (predictors centred at their weighted window mean, scaled by their study-area SD). A window's slopes are collinear when cn_slopes > 30 or singular, or cn > 1e6; n_local_collinear and the warning count those. The coefficient map masks the Intercept with cn > 30 and slopes with the slope flag. An intercept-only problem is logged. - gwr_model_selection() and cv_gwr() apply fit_gwr_model()'s adaptive bandwidth range check up front. - gwr_model_selection() raises the tiny-fixed-bandwidth warning and explains a singular window (shared helpers with fit_gwr_model()). - AICc advice at bandwidth = n names fewer predictors, more data or a gaussian/exponential kernel; the floor warning and docs mention edge ties. - print() shows "122,372 metre (fixed, ...)" / "42 neighbours (adaptive, ...)". - fit_gwr_model() refuses a non-numeric response before any fitting. - Script 09 labels match ?fit_gwr_model. Findings: S7-GWR-CLASSES-1, -3, -4, -5, -6, -7, -8; S10-DOCS-PKG-7. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
The local collinearity survey keeps Belsley's intercept-inclusive index for the intercept and adds a slope index on window-centred predictors, so a predictor's origin (kelvin against Celsius) no longer flags or masks the slopes. plot.spatial_fit()'s help now describes the per-term mask. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
ggplot2 3.5 names the layers of a plot, so the vector of geom classes the test built carried names and expect_identical() against a bare string failed on CI, while ggplot2 3.4 (no layer names) passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
… say What users saw: - determine_optimal_levels() on five stations visited thirty times each warned "no cluster structure" and returned 3 2 4; the level with one cell per station (WSS 0, or 7.8e-17 of floating-point residue) was left out of the log-log line and nothing counted the fall to zero. The warning also named max_levels when the stations ended the ladder, and a two-level ladder was told it "falls in a straight line". - criterion = "combined" still ranked geometry with the linear chord the elbow had stopped using, and unscored candidates shared an average rank that shrank as more went unscored: 8 separated clusters came back as 6 5 10 or 6 10 7. - resolution_profile() took Cp's nugget from a sac refused because its range is below the shortest lag (so the nugget is not identified; 6e-7 on a sill of 0.99 sent Cp to the ceiling), and its nugget-0 guard tested for exactly 0. Its warnings blamed a `sac` the caller never passed, it only logged dropped empty points, and its print credited "30 distinct locations" for a ceiling one short of the 30 points. - Tour script 02 stopped at 02.6 (select_resolution(prof, "elbow") on uniform points), and scripts 02/08, the README figure and troubleshooting entry, the resolution vignette and four help examples made claims the current code contradicts. What changed: - .elbow_read(): zero WSS is relative (1e-12 of the total), left out of the line, and when the rest has no elbow the fall to zero is the elbow. Used by determine_optimal_levels() and resolution_profile()'s elbow column. - No-elbow warnings name the bound that ended the ladder; a ladder of two levels gets its own "too short to read an elbow" warning. - Combined: geometric axis is the log-log sag (flat without an elbow), unscored candidates all take the last Moran place, exact ties go to the k nearest the elbow. - resolution_profile(): "shortest lag" refusals give neither cp nor reliability; nugget-0 warning at 1e-4 of the sill; messages name the internally estimated variogram when no sac was passed; dropped points raise an R warning; print/.ladder_edge say "one short of the points". - Docs: examples on set.seed(4); script 02 draws the Cp pick, prints each optimum's bound and computes the split comparison; script 08 prose and block-size comment; README figure (regenerated, label names the reliability pick), paragraph and troubleshooting entry; vignette table and intro. - Tests: test-review3-S1-resolution.R; test-review2-resolution.R now expects k = 5 on the stations with no no-elbow warning. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
summarize_by_cell():
- A rejected `sac` raised the classed fallback warning "Falling back to
deff = 1", after which a variogram estimated from `response_var` was
applied (median deff 5.2, deff_applied TRUE), so tryCatch() on the class
discarded a corrected result; with the estimate rejected too, one
fallback gave two R warnings. A replaced sac now gets a plain warning,
and the classed warning is raised once, only when the SEs really are
uncorrected, naming every reason.
- deff = "kish" attached no "deff_applied" attribute (and set the column
FALSE) when only the predictor ICC was positive, although the predictor
SEs were inflated 11x. It is attached whenever either ICC is positive.
- A residual (detrended) sac corrected the response SEs in silence while
kriging_adequacy() warns about the same object; it now warns too.
- A pure-nugget model was reported as "could not be read" with the fallback
warning; it is applied as deff = 1 (deff_applied TRUE).
- deff_max_n of 0 or 1 gave uncorrected SEs marked corrected, NA an obscure
error; it must be a number >= 2 under deff = "variogram".
- The NA-ID group (keep_unassigned = TRUE input) got cell_weight 0.
- A fixed deff with one populated cell was recorded as c(2, NA, ...).
- agg_funs = stats::median is named "median", not "agg1".
assign_features_to_polygons():
- A features column named like the polygons' fallback ID column ('id') was
dropped; the ID now travels through the join under a reserved name.
- largest = TRUE let the polygon row order decide an exact tie in overlap;
it is now decided by tie_break and counted in "ties".
- sf's "attribute variables are assumed to be spatially constant" warning
leaked from every largest-overlap join; that warning alone is muffled.
Row records: dplyr's row verbs (filter, slice, arrange, distinct) drop the
record as `[` does (dplyr_row_slice method registered in .onLoad); n_rows,
st_drop_geometry() and binding of geometry-free frames are documented.
Logging: under knitr a deff fallback, and a .log_warn() followed directly
by warning() (the cross-validation cautions), appeared twice; both are now
marked as raised, centrally in .log_warn().
Docs: summarize_by_cell's help page gives the small-sample rescaling's
derivation and coverage (the package page moves it to the "chosen, not
cited" list); lon/lat cell_area is geodesic m^2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
The README, four vignettes and ?summarize_by_cell's autocorrelation section told readers that the deff-corrected standard error is the right one for each cell mean, and called the default one anticonservative, with no estimand named. The corrected SE is the SE of a cell mean as an estimate of the population (grand) mean. For a cell's own mean, which is what a map reports, s/sqrt(n) is already right when the points are spread through the cell, and the corrected SE is sqrt(deff / (1 - rho)) times too wide (4.6 at 20 points a cell and rho = 0.5, measured with the package). Every passage now names the estimand and points per-cell uses to deff = 1, and the getting-started pipeline, which maps its cells, aggregates at deff = 1. Also: the README troubleshooting list covers "no recognised model requested" and scopes "no viable models" to uninstalled backends; the getting-started install table no longer says patchwork is needed for the plot functions; the North Carolina vignette explains the GWR local-collinearity warnings its chunks raise (an uncentred elevation, nearly collinear with the intercept; centring it gives the same fit and no warning); and its fold-map alt text describes the blocked folds as they are drawn. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
The corrected standard error is documented, in the help, README and vignettes, as the SE of a cell mean for the population mean; the plain SE is the one for a cell's own mean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
What the user saw, and what changed:
- predict() on an rf_fit stopped with ranger's "sample_fraction too small"
when every newdata row was incomplete, so predict_surface() aborted on a
chunk outside the covariates' coverage and model_metrics(newdata =)
errored. It returns all NA with a log line again, as the other backends
do; ranger's own errors are still raised. (RF-1)
- fit_bayesian_spatial_model(standardize_predictors = TRUE) could not fit
brms::categorical() or mixture(): the automatic normal(0, 5) slope prior
had no dpar and matched no slope. It is now set per distributional
parameter from brms::get_prior(), as the lscale prior is. (RF-2)
- cv_bayes() under a categorical or ordinal family sampled every fold and
then scored none (4.35 min for k = 2). It is refused before anything is
fitted, naming the family; the "Which metrics survive" section and
?cv_bayes say so. (RF-3)
- ?fit_bayesian_spatial_model's check_convergence said a too-coarse GP basis
sets convergence_ok = FALSE and is flagged by print(); it is only logged
and recorded. The doc now says that. (RF-4)
- cv_rf() raised fit_rf_model()'s out-of-bag warning once per fold, each
telling the user to score the forest with cv_rf(). Fold forests are now
silent and cv_rf() warns once per run with the number of fold forests
affected (counted through fold_info_fn, so parallel runs count right).
(RF-6)
- The per-category epred error counted the two GP-boundary padding rows
("150 x 7 x 3" for five rows); they are dropped from the array. (RF-7)
- That error pointed at posterior_epred(<fit>$engine, newdata = ), which
refuses new rows; it now points at type = "predict", draws = TRUE (share of
draws per category). type = "predict" without draws on a categorical()
fit, the mean of unordered category indices, is now an error. (RF-8)
- print() on a forest with no out-of-bag row printed an empty importance
line and no OOB line; it says both are undefined. The doc says $info
holds NA for the OOB error. (RF-9)
- The README said every backend refuses a factor response; it notes the
categorical/ordinal exception. (RF-10)
Tests: tests/testthat/test-review3-S8-bayes-rf.R (one real-sampler test,
opt-in via SPATIALKIT_TEST_BRMS).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…ls as summarize_by_cell does estimate_sac_range(detrend = "reml") refused genuine short ranges as "below the shortest lag fitted" by comparing the REML range with the first bin of a gstat variogram the REML fit never uses: 10 of 20 REML estimates of a true 30 m range (n = 400) were refused at 14.5-28.8 m, and all came back at cutoff = 0.1. The REML floor is now the distance within which 30 pairs of the points the fit used lie (.reml_trend() returns it as pair_floor), capped at that first lag. True 30 m fields: 0 of 20 refused (n = 400, same at cutoff 0.5 and 0.1), 0 of 15 at n = 300. White noise (n = 300, 30 draws): 16 of the 19 old refusals remain, finite in 13 of 30. The refusal now records range_floor, keeps the reml list, and its warning names the reference; the variogram-path warning points at a smaller cutoff. plot() captions the refusal with both numbers. The past-the-lags and non-convergence warnings no longer say "supply predictor_vars" of an estimate that was already detrended. print() says a CRS-less estimate is in the layer's own coordinate units. ?estimate_sac_range: the rejected-range bound as identification, the corrected white-noise numbers, when the directional maximum is logged, and when to rerun with a smaller cutoff. kriging_adequacy() lost the cells of double IDs from 1e5 up (n = 0: "1e+05" missed "100000") and refused cells keyed by 'id' or 'grid_id'. It now finds the ID columns and converts them as summarize_by_cell() does (.id_as_character()) and warns on unmatched IDs. A CRS-less cells layer (or points layer) is taken to be in the other's CRS before anything is moved, instead of failing inside gstat::krige(). With no variogram model the error passes on rejected_reason and does not blame `sac` when none was passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
What users saw, and what changed:
- build_tessellation(method = "hex"/"square") laid its lattice over a
near-global lon/lat boundary in Web Mercator (whole hexagons 5.75x apart in
true area) where create_grid_polygons() used Equal Earth. The lattice now
gets the same equal-area check (shared .equal_area_grid_crs()), with no
crs or a geographic one, and the points are indexed in that CRS.
- A CRS-less boundary given with lon/lat points was resolved against the UTM
zone picked for the points: a one-degree tile became a one-metre square and
every point was indexed NA. It is now read in the points' own CRS when it
fits the lon/lat envelope, and refused otherwise, in build_tessellation(),
create_voronoi_polygons() and clip_target_for(). CRS-less lon/lat points
with a CRS-less boundary in metres are refused instead of stamped and
reported as "not polygonal".
- create_voronoi_polygons() now applies the lon/lat heuristic to CRS-less
points, as build_tessellation() does (18% of locations were in the wrong
cell).
- method = "triangles": an EMPTY point no longer aborts the rank check
("NA/NaN/Inf in foreign function call"); with a geographic crs the returned
triangles' corners are snapped back onto the input points so a spatial join
reproduces `index`; params no longer records an ignored approx_n_cells.
- clip_target_for() treats a bbox whose short side is below 1e-6 of the long
side as degenerate (a noisy transect gave a sliver and a 166,536-cell grid
for 25); the max_cells error names target_cells/n/cellsize, whichever set
the size.
- The builders accept an sf/sfc layer as `crs`; a whole build_tessellation()
result passed as a boundary or cell layer is named, with the component to
pass, instead of a bare sf error; build_tessellation() on polygon or line
features says to use coerce_to_points().
- Seeding on a lon/lat boundary no longer warns about lwgeom on every call,
and the union is taken on the sphere (same seeds with s2 on or off).
- .crs_distance_error() thins a detailed outline before unique() (2.8 s ->
0.06 s per candidate on 300k vertices), interpolates outline edges the
short way round +-180, and uses the spherical centroid for a feature that
crosses it (164% -> 0.03% reported for a box around Fiji).
- The projection message names the CRS chosen; docs corrected for hex counts
by orientation, coerce_to_points(tmp_project), CRS-less pairs, and the
getting-started keep_duplicates line.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
R CMD check counts a requireNamespace() call in the tests as a use of that package and warned that lwgeom, which DESCRIPTION does not list, was used undeclared. The test only needs to know whether lwgeom is installed, so it asks system.file() instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…nge is NA
fold_separation() and kriging_adequacy() matched make_folds() splits to
rows by position and never checked the probe that cv_*() check. So folds
built on prep_model_data()'s 200 points and applied to the 195 that
assign_features_to_polygons() kept measured the wrong points: within_range
1.0 against 0.23-0.48, and a kriging CV RMSE of 1.82 against 2.11, with
nothing said. Both now run .check_fold_probe() through a small wrapper,
.check_fold_provenance(). The wrapper skips the location comparison when
the probe's geometry kind (POINT or not) differs from the layer's, so
folds built on polygons still measure their pointized copy.
estimate_sac_range(predictor_vars =) passed every row to lm(), and
na.exclude drops NA but not Inf. One Inf predictor or response aborted
the detrending, and the variogram fell back to the raw response (2950
against 2168). Under detrend = "reml" a -Inf predictor first gave a false
"did not converge" warning. Rows with a missing or non-finite response or
predictor are now left out of both fits, with a logged count.
estimate_sac_range()'s early NA returns (fewer than 30 points or finite
values, a constant response or residuals, no extent, gstat missing) carried
no reason. They now carry rejected_reason and remain unclassed with no other
attribute. make_folds(auto_range = TRUE), kriging_adequacy() and
summarize_by_cell() already read that attribute, and now quote the reason.
fold_separation()'s print told random, leave-location-out, buffered-LOO
and NNDM folds to "widen the blocks", and called NNDM optimistic although
NNDM reproduces the prediction distances. The advice now depends on the
fold method. The range line also named the CRS code where the unit belongs
("in EPSG:32617 units"); it now gives the CRS's linear unit ("in metres").
fit_rf_model()'s all-NaN out-of-bag warning now says that
area_of_applicability() cannot be weighted by the NaN importance and to
pass weights = NULL. ?fit_rf_model's See also qualifies the pmax()
recipe the same way.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
… recheck fold_separation() and kriging_adequacy() now check that the folds were built on the layer they are given, estimate_sac_range() fits on finite rows only and gives every early NA a reason, fold_separation()'s advice depends on the fold method, and plot.sac_range()'s refusal names that reason instead of saying nothing is attached. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…t a CRS-less boundary
area_of_applicability(model = fit, folds = folds) stopped with "fold 1
refers to rows outside 1:n" whenever prep_model_data() had dropped a row,
for the same folds cv_*() accepted: the fold IDs name the rows of the
layer fitted from, and were read as positions in the fit's shorter data.
The fit's "dropped" record now rebuilds the kept rows' IDs, and the
removed rows are taken out of the folds (and out of a label vector with
one label per input row), with a log line. area_of_applicability() also
makes the fold provenance check cv_*() make, so folds built on other rows
are refused; the check is skipped when the folds were built on polygons
and the training data are the points a fit reduced them to.
A boundary without a CRS got only a log line in make_folds(), cv_*() and
predict_surface(), where every other function raises an R warning, and
cv_rf() warned twice about one lon/lat boundary, naming
ensure_projected() both times. make_folds() (boundary, blocks,
prediction_points) and predict_surface() now align through
.transform_or_stamp(), and cv_*() resolve the boundary once, with one
warning naming the caller. The stamping warning names the CRS it stamps
instead of "the supplied `crs`".
resolution_profile() read sac = set_units(1.5, "km") as a range of 1.5 m
(a floor of 41 million cells); a units or character sac is now refused by
name. summarize_by_cell(deff = "variogram") ignored a sac without a
variogram model without a word; it is now set aside with a warning, the
classed fallback one when deff falls back to 1.
The cross-validation cautions that were a .log_warn() followed by
warning() are one .warn_and_log() call each, with the warning text and
class unchanged, and .next_is_warning(), which read the caller's body
under knitr to spot such pairs, is gone.
vignette("diagnostics"): the selection example that was meant not to leak
used blocks narrower than the range and a learner that could not fit the
intercept-only model; it now passes block_size = 400, fits z ~ 1 for an
empty set, and shows sel$history.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Bug fixes for code that shipped in 2.0.0 are listed under Bug fixes; the recheck's corrections to features new since 2.0.0 amend their existing entries, and documentation-only changes are listed under Documentation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…m the recheck area_of_applicability() checks that its folds match the fit's rows and handles rows prep_model_data() dropped; a CRS-less boundary raises the same R warning everywhere; resolution_profile() refuses a units sac; the diagnostics vignette's selection example no longer leaks; and the logged cautions raised as warnings go through .warn_and_log() instead of the caller-inspection that decided it under knitr. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…ir own index The vignette said fit_gwr_model() warns that 25 of the 300 windows are collinear and that compare_models_cv() repeats it; since the slope index was added neither warns. The 25 windows are those where the index with the intercept exceeds 30, so their intercepts are masked and the slopes are not; the text now says that, and that centring elevation determines the intercepts too (checked: same bandwidth, fitted values within 3e-10, no window flagged). A make_folds() comment still said a bare NA from estimate_sac_range() carries no reason, which it now always does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Adds the area_of_applicability() fold, CRS-less boundary, infinite-value and sac-without-model entries, amends the resolution_profile(), fold_separation(), kriging_adequacy(), make_folds(blocks) and fit_rf_model() entries, and updates the misspelt-variable example to the "combined" criterion's current answer (11 10 7). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
Folding each cv_*() function's log line and warning into one
.warn_and_log() call kept the warning's text, which had no count, and
dropped the log line's "all N folds failed to produce predictions". The
single message now carries both ("all folds failed (all 2 folds failed
to produce predictions); ..."), so the log still says how many folds were
attempted. test-evaluation.R caught it on CI, where brms is not
installed; it skips wherever brms is. The README quotes the new text.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…ount Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…RS-less prediction layers cv_bayes() and compare_models() scored a Bayesian fit whose sampler did not converge like any other, and only the fit's log lines said so. cv_bayes()'s fold_metrics and compare_models()'s table gain convergence_ok (TRUE/FALSE as fit_bayesian_spatial_model() judged the sampler, NA when unchecked), and each raises one R warning naming the folds or models that did not converge. predict_surface() aligned a CRS-less `grid` or `covariates` layer to the fit's CRS with a log line only; it now does so as it does `boundary`, reprojecting lon/lat-looking coordinates or stamping the fit's CRS, with an R warning naming the argument either way. NEWS also corrects the "combined" criterion entry: ten is the smallest count Moran's I can score, not the response's preference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
…en cells An elbow below ten cells is a count Moran's I cannot score, and every candidate below that floor ranks last on the Moran axis, so the rank average put ten -- the smallest count Moran's I scores -- first whatever the response did: on eight separated clusters the geometric call gave 8 7 9 and "combined" 10 11 12, for a response of noise and for one varying by cluster alike. Moran's I has nothing to say about the counts where the elbow lies, so "combined" now returns the geometric ranking there, logs why, and records it in the diagnostics (criterion = "geometric", fallback). Elbows at ten cells or more, and curves with no elbow, are ranked as before. The help page also says that "combined" is not an estimate of the number of clusters, and when to use "geometric" or resolution_profile() instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR addresses multiple regressions and edge cases discovered during adversarial review of the spatial analysis package, spanning tessellation, cross-validation, kriging, assignment, and related functionality.
Summary
Comprehensive bug fixes and improvements across multiple modules, including handling of edge cases in spatial operations, CRS management, convergence checking, and cross-validation workflows. The changes improve robustness when dealing with invalid geometries, empty features, and various CRS configurations.
Key Changes
Tessellation & CRS Handling:
coerce_to_points(mode = "auto")segfault on EMPTY MULTILINESTRING and rejection of EMPTY LINESTRINGbuild_tessellation(method = "triangles")precision loss by normalizing UTM coordinates before passing to qhull.pick_local_projected_crs()incorrectly sending circumpolar lon/lat data to Web Mercatorwith_s2_on()to ensure spherical calculations run with s2 enabled regardless of session settingSpatial Probes & Fold Validation:
sf_use_s2()settingkind = "planar_centroid"tracking for probes to distinguish legacy behaviorlegacyparameter to probe generation for checking older foldsCross-Validation:
mean_CRPSto reserved metric column names to prevent user overwritescv_gwr()hanging issuecompare_models_cv()Kriging & Resolution:
kriging_adequacy()documentation:kr_rationow represents variance ratio to cell's prior variance (not total sill)max_neighboursandmax_box_ratioparameters tokriging_adequacy().rbar_hull()for measuring domain correlation on convex hull instead of bounding boxrange_floorparameter toresolution_profile()Model Fitting & Diagnostics:
print.spatial_fit()to distinguish "not checked" (NA) from failure.validate_kernel()function from GWR modulepredict.bayesian_fit()to pin HSGP boundary at fitted valueAssignment & Aggregation:
assign_features_to_polygons()to run joins in projected CRS when available for accurate planar overlaplargest = TRUEArea of Applicability:
.aoa_equal_if_all_zero()chunk_sizeand duplicate training rowsDocumentation & Testing:
Notable Implementation Details
https://claude.ai/code/session_01AwT2vKb3g4LrJDB1CTbviN