VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

TL;DR: A 100 kB map of class-labelled 3D boxes is enough for long-term indoor relocalization once the query is an odometry-connected video clip rather than a single frame.

Qianru Li1,2, Xuyang Chen1,2, Xuqin Wang1,2, Zhenghao Zhang1, Hongyi Luo1,2, Tao Wu2, Daniel Cremers1, Lu Liu2, Yanfeng Zhang2†

1Technical University of Munich    2Huawei Hilbert Research Center (Dresden)
Corresponding author

📄 Paper (coming soon) 📜 arXiv 🎬 Video 💻 Code (coming soon)
Left: mapping visit stored as class-labelled boxes (sofa, table, chairs, cabinet, window) with no appearance. Middle: frames from a re-visit weeks later with changed lighting and a moved chair. Right: the re-visit trajectory relocalized in the box map, with gravity and Manhattan wall directions as orientation cues.
Map size versus recall at 1 m/10 degrees over all RIO10 re-scan frames: VideoReloc reaches the highest recall with the smallest map.
Video relocalization with a compact semantic map. Left: the map stores class-labelled boxes for objects, walls and floors in about 100 kB. A clip from a later RGB-D video is localized in the map's coordinate system after changes in lighting and furniture layout. Right: map size versus recall at 1 m/10° over all RIO10 re-scan frames. VideoReloc achieves the highest recall with the smallest map among these methods.

Abstract

Given a compact semantic scene graph, long-term indoor video relocalization estimates a map-frame trajectory after lighting and furniture changes. Visual methods rely on appearance and become unreliable under these changes; localizing one frame at a time from object classes and geometry instead leaves sparse, ambiguous evidence. We introduce VideoReloc, whose adaptive clips use odometry to gather spatial evidence until object and motion criteria are met, adapting query length to the observed scene. Its run-level decision rechecks conflicting placements using evidence accumulated across connected clips, stabilizing the trajectory beyond adjacent-clip tracking. Hypothesis-first registration proposes poses from object triplets and verifies each using clip-wide object centers and box surfaces. Orientation-aware refinement uses box faces, gravity and wall directions to resolve ambiguity in camera orientation and refine the full pose. This reframes sparse-map relocalization as verification of spatially extended video queries, moving discriminative support from stored appearance to temporal context and permitting a 100 kB map of class-labelled boxes. On RIO10 and ReplicaCAD, the all-frame localization success rate at 1 m/10° is 73.5 % and 61.1 % under causal evaluation, rising to 90.6 % and 74.8 % with clip closure. The evaluated per-frame scene coordinate regressors reach up to 47.6 % and 49.8 %, respectively, with maps of 12.6–42 MB.


Video

Overview video (2 min 43 s, English narration, captions burned in). A later visit to a changed room and a kilobyte-scale map of class-labelled boxes; causal localization frame by frame on RIO10; why one frame is ambiguous and how an adaptive clip fuses motion into a spatially extended query; object-triangle hypotheses verified against clip-wide geometry; gravity, wall directions and box faces resolving orientation; the run-level decision revising an earlier clip; the same frames under VideoReloc, ACE-G and R-SCoRe; illumination and furniture changes in ReplicaCAD; results on both datasets; and a failure case.

Method

VideoReloc pipeline: per-clip graph construction from RGB-D odometry, SAM 3 masks and Map-Det3D boxes; clip candidate generation by previous-clip prediction, triangle retrieval, center consensus, frame-pose anchors and orientation fallback; box-surface consistency verification; and a back-end with orientation-aware refinement, run-level decision and a clip-rigid pose graph, all against a global scene graph of text labels, 3D boxes, gravity and wall axes.
Overview of VideoReloc. Per-clip graph construction produces the local scene graph Lc from RGB-D odometry and depth-backed instance fusion. Clip candidates generation combines previous-clip prediction, triangle retrieval, frame-pose anchors and orientation fallback; box-surface consistency verifies them. The back-end applies orientation-aware refinement, the run-level decision and the clip-rigid pose graph; final frame poses satisfy Ti = ScVi.

Contributions

VideoReloc replaces mapping-visit appearance with a persistent map of class-labelled 3D boxes, about 100 kB per room, and draws discriminative support from the incoming video. Adaptive clips use odometry to turn partial frame observations into a spatially extended query, closing after sufficient object evidence and camera travel.


Qualitative Results

RIO10 qualitative results on scenes 07, 09 and 10: estimated trajectories of VideoReloc, ACE-G and R-SCoRe drawn over the map boxes, coloured by position error, with rolling median translation and rotation error curves below.
RIO10 qualitative results. Grey boxes are map objects; trajectory colour continuously encodes position error (green: 0 m; yellow: 0.5 m; red: ≥ 1 m). VideoReloc uses estimates with clip closure. Lower plots show log-scale rolling median translation/rotation errors; dashed lines mark 1 m/10°.

Quantitative Results

RIO10 contains ten rooms scanned weeks apart with furniture and lighting changes; maps use the first visit, and evaluation includes all 34,415 re-scan frames. ReplicaCAD is a controlled evaluation built from six authored layouts of one apartment, rendered along the same 4,000-frame path under daylight, warm evening lights and two warm night lamps; the map is built from a separate 6,000-frame daylight tour. Recall counts a frame as successful only when neither position nor rotation error exceeds the stated thresholds; frames without an estimate count as failures.

Comparison of recall and pose errors across relocalization methods and evaluation variants on RIO10 and ReplicaCAD. Best in bold, second best underlined.
RIO10, 10 re-scans, 34,415 frames ReplicaCAD, 18 tours, 72,000 frames
Method Map size Pos./rot. R@1 m/10° 0.25 m/5° Pos./rot. R@1 m/10° 0.25 m/5°
HLoc1–2 GB12 cm / 3.9°49.841.34 cm / 0.6°48.444.5
ACE-G12.6 MB36 cm / 11.7°45.525.752 cm / 8.2°49.837.1
R-SCoRe42 MB40 cm / 12.1°47.638.572 cm / 9.4°48.642.5
MSG-Loc1.5 MB161 cm / 49.3°1.20.1265 cm / 27.7°11.31.3
FPFH–ICP6–13 MB11 cm / 3.4°22.421.71 cm / 0.2°10.09.9
ACE-G12.6 MB10 cm / 3.6°79.866.214 cm / 1.8°67.059.4
R-SCoRe42 MB9 cm / 3.0°85.871.011 cm / 1.7°66.759.8
VideoReloc (ours), causal95 kB20 cm / 4.3°73.544.027 cm / 3.0°61.140.4
VideoReloc (ours), with clip closure95 kB18 cm / 3.3°90.664.122 cm / 2.1°74.853.0

Protocol. Recall (%) includes all frames, with missing poses counted as failures. Causal evaluation uses frames up to the query frame; evaluation with clip closure also uses the remaining frames of that clip. Median position and rotation errors are ranked separately, and ties share rank. marks clip-rigid controls: the same clip partition and within-clip odometry as VideoReloc, but each clip is localized independently without cross-clip tracking, run-level decisions, or pose-graph optimization. Map size and errors. Map size is the map each method builds on RIO10; ReplicaCAD maps occupy 134 kB for VideoReloc, 1.96 MB for MSG-Loc and 18.1 MB for FPFH–ICP. Median errors use only returned poses. Pose coverage (RIO10 / ReplicaCAD) is 100/100 % for the regressors, 75/68 % for HLoc, 29/12 % for FPFH–ICP, 21/61 % for MSG-Loc, 94/97 % for VideoReloc under causal evaluation and 100/99 % with clip closure. FPFH–ICP errors summarize scene medians on RIO10 and a separate replay on ReplicaCAD.

Recall at 1 m/10° under illumination and furniture changes in ReplicaCAD. Best in bold, second best underlined.
Illumination Rearrangement
Method dayeveningnightΔL unchangedrearrangedΔRkept
HLoc55.551.937.7−17.775.942.9−33.056 %
ACE-G53.149.646.8−6.378.444.1−34.356 %
R-SCoRe54.250.341.5−12.674.743.4−31.358 %
MSG-Loc12.212.09.6−2.633.06.9−26.121 %
FPFH–ICP10.19.18.9−1.240.13.2−36.88 %
ACE-G68.965.766.5−2.489.462.6−26.870 %
R-SCoRe68.866.864.6−4.387.862.5−25.371 %
VideoReloc (ours), causal60.064.259.0−1.080.157.2−22.971 %
VideoReloc (ours), with clip closure73.876.873.9+0.188.572.1−16.481 %

Recall (%) uses all frames; illumination averages six layouts per column, and unchanged/rearranged average three/fifteen tours. ΔL and ΔR are night–day and rearranged–unchanged changes (points); kept is rearranged/unchanged (%). Rankings use unrounded recall; ties share rank. is defined as in the table above. FPFH–ICP factors use per-tour medians across replay seeds.

Effect of the localization unit on RIO10.
Unit1 m/10°0.25 m/5°
one frame2.60.4
30-frame clip19.19.3
100-frame clip39.323.7
250-frame clip52.835.4

Each unit is registered independently, without tracking or a pose graph. Recall (%) uses all 34,415 frames and is measured when the unit closes.

Effects of refinement, run-level decisions and adaptive clips.
Mechanism RIO10 ReplicaCAD tours
orientation-aware run-level adaptive with clip closure causal with clip closure causal
coarsefinecoarsefinecoarsefinecoarsefine
76.439.370.833.464.938.656.333.6
78.544.472.538.070.445.162.041.3
81.343.970.836.278.951.669.447.1
91.061.373.041.856.339.345.429.2
90.664.173.544.074.853.061.140.4

Recall (%) includes all frames. Coarse: 1 m/10°. Fine: 0.25 m/5°. Checkmarks indicate enabled components; empty cells indicate point-to-point refinement, no run-level decision and fixed 100-frame clips, respectively. Tracking and the pose graph remain enabled. The last row is the full VideoReloc. Bold/underline: first/second in each column.


BibTeX

@misc{li2026videoreloc,
  title         = {VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph},
  author        = {Li, Qianru and Chen, Xuyang and Wang, Xuqin and Zhang, Zhenghao and Luo, Hongyi and Wu, Tao and Cremers, Daniel and Liu, Lu and Zhang, Yanfeng},
  year          = {2026},
  eprint        = {2609.21804},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.21804}
}