Code and data release for:
Leveraging Semantic Maps for City-Scale Cross-View Localization Ethan Fahnestock*, Erick Fuentes*, Philip R. Osteen, Nicholas Roy arXiv:2607.25215 · project page · dataset
LOCI (Landmark-Oriented Cross-view Inference) localizes a robot against overhead imagery and OpenStreetMap by extracting semantic landmarks from panoramas with a VLM, matching them against OSM landmarks with a lightweight learned correspondence classifier, and fusing the resulting observation likelihood with an image-based likelihood in a histogram Bayes filter.
- Code (this repository): the full pipeline — dataset creation, VLM landmark extraction, correspondence-classifier training, model training (WAG / WAG+OSM / LOCI-EF), histogram-filter path evaluation, timing benchmarks, and paper figure/table generation.
- Datasets (
ethankf/loci-dataon Hugging Face, gated — accept the terms of use to download): GPS-stamped panoramas (Mapillary-derived and self-collected), OSM landmark tables, VLM-extracted landmark annotations, the landmark-correspondence dataset, evaluation paths, model checkpoints, and the per-path evaluation outputs behind the paper's results.
Most satellite imagery is NOT redistributed. Google Maps and Esri terms
do not permit republishing tiles. Instead, data_pipeline/satellite/
re-downloads the exact tiles (pinned zoom/patch grid per environment; Esri
Wayback pinned to a specific release) and assembles the same 640x640 patches
used in the paper. The two public-record exceptions ship with the dataset:
MassGIS 2025 patches (framingham) and CT 2023 orthoimagery patches
(middletown). Panoramas for the VIGOR cities (chicago, seattle, new_york)
must be obtained
from the VIGOR authors
(academic-use only, not redistributable); the paper's San Francisco
panoramas are Mapillary-sourced (only its satellite patches come from VIGOR).
| Path | Contents |
|---|---|
experimental/overhead_matching/swag/ |
Core pipeline: data loaders, models, filter, evaluation, scripts (Bazel) |
experimental/overhead_matching/baseline/ |
WAG+OSM baseline: OSM vector tiles rendered to rasterized satellite-style images |
common/, planning/, third_party/, toolchain/ |
Shared libraries (torch loading, web mercator, C++ OSM landmark extraction, A*) and build support |
data_pipeline/ |
Dataset creation: Mapillary → VIGOR (mapillary/), satellite tile download (satellite/), 360-video+GPS rig ingestion (video_to_dataset/) |
paper/ |
Ground-truth configs, sigma calibrations, and the exact pipeline scripts behind the paper's results |
docs/ |
Step-by-step guides for each pipeline stage |
website/ |
Project page (GitHub Pages; interactive results explorer built from the released evaluation outputs) |
tools/ |
Checkpoint reassembly (reassemble_checkpoints.py), released-artifact verification (verify_artifacts.py), token-usage accounting |
# One-time: installs system deps and writes the machine-local .bazelrc_ubuntu
# (selects the clang toolchain for C++ targets)
./setup.sh
# Bazel (bazelisk) picks up the pinned version from .bazeliskrc (7.6.1)
bazel build //...
bazel test //experimental/overhead_matching/...
# Download the dataset and reproduce the paper's headline results from the
# released evaluation outputs: see docs/reproducing_results.mdDataset paths are rooted at $DATA_ROOT (wherever you download/untar the
data release); see docs/reproducing_results.md for the expected layout
and download instructions.
Read in pipeline order:
docs/dataset_creation.md— panoramas (Mapillary / camera rig / VIGOR), satellite patches, OSM landmarksdocs/landmark_extraction.md— Gemini VLM landmark extraction from panoramasdocs/correspondence_model.md— correspondence dataset generation + classifier trainingdocs/training.md— WAG, WAG+OSM, and LOCI-EF model trainingdocs/evaluation.md— evaluation paths, histogram filter, convergence metricsdocs/timing_and_figures.md— latency/scaling benchmarks and paper figures/tablesdocs/reproducing_results.md— end-to-end reproduction of Table V / Figure 7
Code is released under the MIT License (see LICENSE). The dataset has
its own composite terms — ODbL for OpenStreetMap-derived tables, CC BY-SA
4.0 for street-level imagery and VLM annotations, CC BY 4.0 for checkpoints
and evaluation artifacts — recorded in the LICENSE file of the
dataset repository.
See THIRD_PARTY.md for the complete upstream inventory.
This work was supported in part by ARL under Grants W911NF-21-2-0150 and W911NF-17-2-0181.