Pose diffusion
Image-conditioned denoising routes stochastic pose particles toward distinct candidate modes across the full floorplan.
- Continuous SE(2) pose distribution
- Cross-modal image-to-map attention
- KDE-based mode extraction
Coarse-to-fine visual floorplan localization
without ray matching.
School of Computer Science and Engineering · Central South University
One image can correspond to many places.
CF²Loc preserves that ambiguity—then resolves it.
Repetitive corridors and symmetric rooms make visual floorplan localization fundamentally multimodal. Existing systems compress visual evidence into sparse rays, introducing an information bottleneck before localization even begins.
CF²Loc directly estimates the conditional pose distribution over the whole map. A generative global search tracks plausible hypotheses, then a lightweight local refiner delivers precise, deterministic corrections.
Image → sparse ray prediction → exhaustive map matching
Image + floorplan → continuous multimodal pose distribution
Probabilistic where the map is ambiguous. Deterministic where precision matters.
Image-conditioned denoising routes stochastic pose particles toward distinct candidate modes across the full floorplan.
Candidate-centered 5 m × 5 m floorplan crops remove global ambiguity and enable bounded sub-meter residual correction.

Evaluated on synthetic furnished scenes and real-world homes, CF²Loc improves both strict accuracy and robust global recall.
“Global ambiguity should be modeled,
not prematurely collapsed.”
CF²Loc bridges global multi-hypothesis tracking and local unimodal refinement in one rendering-free framework.
If CF²Loc supports your research, please cite the paper. Code and pretrained models will be released in this repository.
@article{meng2026cf2loc,
title = {From Uncertainty to Determinism: Coarse-to-Fine
Visual Floorplan Localization without Ray Matching},
author = {Meng, Shiyong and Chen, Bolei and Zhong, Ping and
Wan, Yang and Wang, Rongzhi and Xia, Jiazhi and Wang, Jianxin},
year = {2026}
}