Skip to content

perf(db): index ingestion and shared-frame lookups - #705

Merged
MateoLostanlen merged 2 commits into
mainfrom
codex/perf-database
Oct 8, 2026
Merged

MateoLostanlen merged 2 commits into
mainfrom
codex/perf-database

Conversation

@frgfm

@frgfm frgfm commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

This PR includes the useful indexes from #664. It adds a separate index for unassigned detections: detections that are not yet grouped into a sequence.

The main gain over #664 is narrower than the gain over the baseline. For the unassigned-detection query, this PR took 1.83 ms instead of 7.97 ms with a generic query plan. That is 77% less time, or 4.4x faster. The shared-frame and other common indexes come from #664. This benchmark does not show an extra gain for those shared paths.

An index works like a sorted list. #664 finds recent unassigned rows across all cameras, then checks the camera and pose (viewing direction). #705 can go directly to the rows for one camera and pose.

flowchart TB
    Q["Find recent unassigned detections<br/>for one camera and pose"]
    Q --> A["#664: find recent unassigned rows<br/>across all cameras"]
    A --> B["Check camera and pose<br/>Remove unrelated rows"]
    B --> R["Return the matching detections"]
    Q --> C["#705: use the extra index<br/>Go to this camera, pose, and time"]
    C --> R
Loading

The saved query plans show the smaller search: #664 found 6,000 rows, then kept 120 after checking the camera and pose. #705 found those 120 rows directly. Both queries returned the same detections.

A generic plan is one query plan reused for different input values. We tested that mode and the default setting, which lets PostgreSQL choose its plan.

Find recent unassigned detections #664 #705 Measured gain over #664
Default plan setting 2.35 ms 2.08 ms 12% less time
Force a generic plan 7.97 ms 1.83 ms 77% less time; 4.4x faster

The extra index uses (camera_id, pose_id, created_at) and includes only rows where sequence_id IS NULL. It occupies about 16 MiB in this fixture. The database must also update it when data changes. We did not measure the cost of those writes.

The following table shows the larger gain from adding indexes to the baseline. #664 had lower medians on these common paths. These measurements do not establish that #705 is faster there. Both PRs use the same indexes for these queries.

Query, default plan setting Baseline without these indexes #664 #705
Latest detection with a bounding box 99.01 ms 0.69 ms 1.07 ms
Candidate sequences, 576 returned rows 9.76 ms 6.03 ms 7.63 ms
Shared-frame lookup, original query 84.31 ms 0.53 ms 0.79 ms

This PR also limits the shared-frame check to two rows. That is enough to find whether another detection refers to the file. The limited query took 0.76 ms, about 111x faster than the original baseline query. The index from #664 provides the main gain. This fixture has only two rows per shared frame, so it does not measure the benefit of limiting a larger result.

Merge one database PR. #705 includes #664's index changes. Do not merge both: their migrations create the same index names from the same parent revision.

Benchmark method, implementation, and checks

Python 3.11.15 / PostgreSQL 15.19. Fixture: 2.6 million detections, 56,000 sequences, 50 active cameras/poses. Rows from different cameras are spread through the table. Medians of 21 calls after warmup include the actual Python CRUD call and database round trip. Nonempty result IDs matched across index sets. Index sets were tested in sequence; timing differences on shared paths do not isolate the extra index's effect.

The candidate query still builds 576 result objects in Python. These synthetic query benchmarks do not measure full API latency or application peak memory.

Add four indexes and matching model definitions. Declare the existing validation-queue index in the model. Deployment stops the backend before migration, so create/drop indexes in one transaction. A failed build then rolls back earlier builds. Keep the latest-detection index unfiltered so it also supports generic plans.

Five tests check index column order and filter conditions. Real PostgreSQL checks passed for migration upgrade/downgrade/re-upgrade, rollback after a name collision, and agreement between model-created and migrated indexes. The delete endpoint kept a shared file until the last reference was deleted. CI, lint, formatting, type checks, and coverage checks pass.

Net diff: +60 production/migration lines; +85 including tests. Benchmark files stay outside the diff. The four indexes occupy about 168 MiB on this fixture, including 73 MiB for the bucket-key index. Write throughput and production migration downtime were not measured.

import sqlalchemy as sa
from alembic import op

revision: str = "e7a4b9c3d5f2"
from alembic import op

revision: str = "e7a4b9c3d5f2"
down_revision: Union[str, None] = "d6f3a8b2c4e1"
@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.88%. Comparing base (25f3d2e) to head (c6d19a7).

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #705   +/-   ##
=======================================
  Coverage   93.87%   93.88%           
=======================================
  Files          59       59           
  Lines        3218     3221    +3     
=======================================
+ Hits         3021     3024    +3     
  Misses        197      197           
Flag Coverage Δ
backend 93.99% <100.00%> (+<0.01%) ⬆️
client 91.30% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@fe51

fe51 commented Oct 3, 2026

Copy link
Copy Markdown
Member

Thanks for the PR ! Maybe could you check #664 first and share ideas?

@frgfm frgfm changed the title perf(db): index detection and sequence ingestion queries perf(db): index ingestion and shared-frame lookups Oct 3, 2026

revision: str = "e7a4b9c3d5f2"
down_revision: Union[str, None] = "d6f3a8b2c4e1"
branch_labels: Union[str, Sequence[str], None] = None
revision: str = "e7a4b9c3d5f2"
down_revision: Union[str, None] = "d6f3a8b2c4e1"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
@frgfm frgfm self-assigned this Oct 3, 2026
@frgfm frgfm added the type: improvement New feature or request label Oct 3, 2026
@frgfm

frgfm commented Oct 6, 2026

Copy link
Copy Markdown
Member Author

Thanks for the PR ! Maybe could you check #664 first and share ideas?

I forgot to reply, but I went through it, and integrated the additional benefits from that PR here so that we have a net positive PR, with minimal edits, and major performance boosts 🤓

@MateoLostanlen MateoLostanlen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @frgfm a lot for this PR, great to get these indexes in!

@MateoLostanlen
MateoLostanlen merged commit fa501a9 into main Oct 8, 2026
23 checks passed
@MateoLostanlen
MateoLostanlen deleted the codex/perf-database branch October 8, 2026 13:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants