Skip to content
efahnestockPublic

About

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

LOCI: Leveraging Semantic Maps for City-Scale Cross-View Localization

Code and data release for:

Leveraging Semantic Maps for City-Scale Cross-View Localization Ethan Fahnestock*, Erick Fuentes*, Philip R. Osteen, Nicholas Roy arXiv:2607.25215 · project page · dataset

LOCI (Landmark-Oriented Cross-view Inference) localizes a robot against overhead imagery and OpenStreetMap by extracting semantic landmarks from panoramas with a VLM, matching them against OSM landmarks with a lightweight learned correspondence classifier, and fusing the resulting observation likelihood with an image-based likelihood in a histogram Bayes filter.

What is released

  1. Code (this repository): the full pipeline — dataset creation, VLM landmark extraction, correspondence-classifier training, model training (WAG / WAG+OSM / LOCI-EF), histogram-filter path evaluation, timing benchmarks, and paper figure/table generation.
  2. Datasets (ethankf/loci-data on Hugging Face, gated — accept the terms of use to download): GPS-stamped panoramas (Mapillary-derived and self-collected), OSM landmark tables, VLM-extracted landmark annotations, the landmark-correspondence dataset, evaluation paths, model checkpoints, and the per-path evaluation outputs behind the paper's results.

Most satellite imagery is NOT redistributed. Google Maps and Esri terms do not permit republishing tiles. Instead, data_pipeline/satellite/ re-downloads the exact tiles (pinned zoom/patch grid per environment; Esri Wayback pinned to a specific release) and assembles the same 640x640 patches used in the paper. The two public-record exceptions ship with the dataset: MassGIS 2025 patches (framingham) and CT 2023 orthoimagery patches (middletown). Panoramas for the VIGOR cities (chicago, seattle, new_york) must be obtained from the VIGOR authors (academic-use only, not redistributable); the paper's San Francisco panoramas are Mapillary-sourced (only its satellite patches come from VIGOR).

Repository layout

Path Contents
experimental/overhead_matching/swag/ Core pipeline: data loaders, models, filter, evaluation, scripts (Bazel)
experimental/overhead_matching/baseline/ WAG+OSM baseline: OSM vector tiles rendered to rasterized satellite-style images
common/, planning/, third_party/, toolchain/ Shared libraries (torch loading, web mercator, C++ OSM landmark extraction, A*) and build support
data_pipeline/ Dataset creation: Mapillary → VIGOR (mapillary/), satellite tile download (satellite/), 360-video+GPS rig ingestion (video_to_dataset/)
paper/ Ground-truth configs, sigma calibrations, and the exact pipeline scripts behind the paper's results
docs/ Step-by-step guides for each pipeline stage
website/ Project page (GitHub Pages; interactive results explorer built from the released evaluation outputs)
tools/ Checkpoint reassembly (reassemble_checkpoints.py), released-artifact verification (verify_artifacts.py), token-usage accounting

Quickstart

# One-time: installs system deps and writes the machine-local .bazelrc_ubuntu
# (selects the clang toolchain for C++ targets)
./setup.sh

# Bazel (bazelisk) picks up the pinned version from .bazeliskrc (7.6.1)
bazel build //...
bazel test //experimental/overhead_matching/...

# Download the dataset and reproduce the paper's headline results from the
# released evaluation outputs: see docs/reproducing_results.md

Dataset paths are rooted at $DATA_ROOT (wherever you download/untar the data release); see docs/reproducing_results.md for the expected layout and download instructions.

Documentation

Read in pipeline order:

  1. docs/dataset_creation.md — panoramas (Mapillary / camera rig / VIGOR), satellite patches, OSM landmarks
  2. docs/landmark_extraction.md — Gemini VLM landmark extraction from panoramas
  3. docs/correspondence_model.md — correspondence dataset generation + classifier training
  4. docs/training.md — WAG, WAG+OSM, and LOCI-EF model training
  5. docs/evaluation.md — evaluation paths, histogram filter, convergence metrics
  6. docs/timing_and_figures.md — latency/scaling benchmarks and paper figures/tables
  7. docs/reproducing_results.md — end-to-end reproduction of Table V / Figure 7

Licenses

Code is released under the MIT License (see LICENSE). The dataset has its own composite terms — ODbL for OpenStreetMap-derived tables, CC BY-SA 4.0 for street-level imagery and VLM annotations, CC BY 4.0 for checkpoints and evaluation artifacts — recorded in the LICENSE file of the dataset repository. See THIRD_PARTY.md for the complete upstream inventory.

Acknowledgments

This work was supported in part by ARL under Grants W911NF-21-2-0150 and W911NF-17-2-0181.

About

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages