Visual Floorplan Localization · 2026

From uncertainty
to determinism.

Coarse-to-fine visual floorplan localization
without ray matching.

Shiyong MengBolei Chen*Ping Zhong*Yang WanRongzhi WangJiazhi XiaJianxin Wang

School of Computer Science and Engineering · Central South University

01 Global pose diffusion
02 Local pose refinement
θ
NO RAY MATCHINGMULTIMODAL POSE DIFFUSIONSUB-METER REFINEMENTREAL-TIME INFERENCE
01 / OVERVIEW

One image can correspond to many places.
CF²Loc preserves that ambiguity—then resolves it.

Repetitive corridors and symmetric rooms make visual floorplan localization fundamentally multimodal. Existing systems compress visual evidence into sparse rays, introducing an information bottleneck before localization even begins.

CF²Loc directly estimates the conditional pose distribution over the whole map. A generative global search tracks plausible hypotheses, then a lightweight local refiner delivers precise, deterministic corrections.

A

Ray-based localization

Image → sparse ray prediction → exhaustive map matching

Visual evidenceInformation bottleneck
B

Direct image-to-map

Image + floorplan → continuous multimodal pose distribution

Refine
02 / METHOD

A clean division of labor.

Probabilistic where the map is ambiguous. Deterministic where precision matters.

Stage 01Global
p(p | I, M)

Pose diffusion

Image-conditioned denoising routes stochastic pose particles toward distinct candidate modes across the full floorplan.

  • Continuous SE(2) pose distribution
  • Cross-modal image-to-map attention
  • KDE-based mode extraction
Stage 02Local
Δx, Δy, Δθ

Pose refinement

Candidate-centered 5 m × 5 m floorplan crops remove global ambiguity and enable bounded sub-meter residual correction.

  • Orientation-canonicalized local crop
  • Translation and heading residuals
  • Confidence-based re-ranking
CF2Loc training and inference pipeline
Figure 2Training and inference workflow. The pose diffusion model first generates coarse candidates; the frozen global stage is then paired with a lightweight local refiner.
03 / RESULTS

State of the art, twice over.

Evaluated on synthetic furnished scenes and real-world homes, CF²Loc improves both strict accuracy and robust global recall.

Benchmark

S3D (full)

SOTA
0.1 m15.0%12.0%
0.5 m71.0%65.2%
1.0 m79.1%71.9%
1 m / 30°78.4%71.4%
Benchmark

ZInD

SOTA
0.1 m11.5%8.1%
0.5 m51.7%43.6%
1.0 m63.2%53.5%
1 m / 30°54.7%50.2%
04 / TAKEAWAY
“Global ambiguity should be modeled,
not prematurely collapsed.”

CF²Loc bridges global multi-hypothesis tracking and local unimodal refinement in one rendering-free framework.

05 / CITATION

Build on our work.

If CF²Loc supports your research, please cite the paper. Code and pretrained models will be released in this repository.

@article{meng2026cf2loc,
  title   = {From Uncertainty to Determinism: Coarse-to-Fine
             Visual Floorplan Localization without Ray Matching},
  author  = {Meng, Shiyong and Chen, Bolei and Zhong, Ping and
             Wan, Yang and Wang, Rongzhi and Xia, Jiazhi and Wang, Jianxin},
  year    = {2026}
}