ACCV 2026 · Accepted

MuViSeg

Multi-view segment correspondences from dense geometry priors

Denis Fatykhoph, Timur Akhtyamov, Konstantin Pakulev, German Devchich, Gonzalo Ferrer

Applied AI Institute, Moscow, Russia

Asian Conference on Computer Vision (ACCV), 2026

Shared context, segment-level correspondences MuViSeg Shared context. Segment-level correspondences. View 1 View 2 View 3 View 4 Segment tokens view 1 … view 2 … view 3 … view 4 … Joint attention over one shared sequence Predicted segment matches View 1 ↔ View 4 / 104° rotation Ground-truth masks are inputs; ribbons show accepted segment matches. Token links are schematic; ribbon width follows mask extent.
Joint multi-view attention recovers segment correspondences across a 104° viewpoint change. Segments from all N views are concatenated into one sequence with per-view position embeddings, so any segment can attend to any other. Ground-truth masks are inputs; ribbons show accepted matches.
01 — Abstract

Object-level mapping and navigation require correspondences between whole instance segments rather than pixels or keypoints. We study how frozen 3D foundation features should be processed for segment matching and introduce MuViSeg: a learned pairwise attention head on MASt3R, multi-scale VGGT feature fusion, and a joint multi-view head that reasons over segments from several views to recover transitive matches.

In zero-shot evaluation on Replica and Virtual KITTI 2 across 0°–180° viewpoint changes, the MASt3R head improves over a Sinkhorn matcher by 4.85 and 25.9 AUPRC points, respectively. The multi-view head is strongest among our heads at the widest Replica baselines, although Sinkhorn remains competitive. Across 102 HM3D navigation episodes, direct segment matchers and sparse keypoint aggregation are statistically indistinguishable, while RoMa's dense matches voted into masks score 13.72–17.98 SPL points lower than direct segment matching.

Keywords   Segment matching · 3D foundation models · Multi-view correspondence · Topological navigation

02 — The question

What is the most effective way to process segments for matching?

A robot revisiting a room has to decide whether this chair seen now is the same chair seen a minute ago. Sparse and dense matchers answer at the wrong level: a downstream system must recover object identity by voting matches inside a mask, and the associations are only as reliable as that voting rule. We take the segment-level route and test three concrete pipelines, each a different answer to the question above.

  • H1Matching capacity is better spent in a learned cross-segment head than before the matcher. Supported at small-to-moderate baselines and outdoors; it reverses past 135° indoors.
  • H2Preserving multi-scale spatial detail up to the mask boundary yields sharper per-segment descriptors than late pooling. Supported: DPT fusion gives +2.8 AUPRC over single-layer pooling on ScanNet++ validation.
  • H3Matching should not be done one pair at a time: segments from several views should be scored jointly. The method we propose. Strongest among our heads exactly where the baseline is widest.
03 — Method

One template, three heads.

A frozen backbone produces per-patch features, a class-agnostic segmenter partitions each image into instance masks, masked average pooling gives one descriptor per segment, and a head refines those descriptors and predicts correspondences. The three variants differ in where capacity sits, which backbone feeds it, and how much context the head sees.

MuViSeg architecture MuViSeg Architecture Frozen Trainable 01Encode segments Input views Frozen MASt3RLocal feature head 24D Mask poolingAverage over mask ProjectionD = 128 (a) Pairwise LGv2 Frozen VGGTLayers 5 · 11 · 17 · 23 DPT fusionMulti-layer spatial Mask pooling ProjectionD = 128 (b) Pairwise VGGT (c) Joint VGGT mᵛ mᵛ 02Joint multi-view matching (c) main contribution View 1 + e1 … View 2 + e2 … View N + eN … learned view embeddings Concatenate Joint self-attention + FFN ×3 layers x x̃ One sequence · all views View i View j DoubleSoftmax + matchabilityScore any requested pair (i, j) Pair scores Pairwise alternatives (a) (b) (a) LGv2 on MASt3R (b) VGGT + DPT Shared head · two separate streams View A View B Self-attn Self-attn Cross-attn Cross-attn FFN FFN ×3 DoubleSoftmax + matchabilityOne pair at a time Colour denotes view, not object identity. mᵛ: instance masks. Attention links are schematic.
Architecture. Grey: frozen. Teal: trainable. 01 — Encode segments: both backbones produce per-patch features that are mask-pooled to one descriptor per segment and projected to a shared 128-dim space; VGGT additionally passes through DPT fusion of Aggregator layers 5 · 11 · 17 · 23 before pooling. 02 — Joint multi-view matching (our main contribution): segment sets from all N views are tagged with learned view embeddings, concatenated into one sequence and refined by three joint self-attention layers, after which the DoubleSoftmax scorer reads off any requested pair (i, j). Colour denotes view, not object identity; attention links are schematic and the score map is a real model output.
(a) Pairwise · MASt3R

SegMASt3R+LGv2

A LightGlue-style attention head on frozen 24-dim MASt3R descriptors: three blocks of weight-shared self-attention, bidirectional cross-attention and an FFN.

≈800K trainable
(b) Pairwise · VGGT

SegVGGT-DPT

DPT fusion over four Aggregator depths restores boundary precision that single-layer pooling loses, then the same attention stack as (a).

≈10.7M trainable
(c) Joint · Main contribution

SegVGGT-DPT Joint

Descriptor sets from all N views are concatenated into one sequence with learnable per-view embeddings and refined by three joint self-attention layers. Any pair can then be scored.

Trained at N=4, evaluated at N ∈ {2,4,6,8}

Why the backbone choice is forced

MASt3R's features are pair-dependent: they cannot be computed without the second image, which suits a strong pairwise head but cannot be shared across a variable set of views. VGGT's Aggregator processes all input views in one shared sequence and its features are pair-independent. Joint multi-view matching is therefore only reachable on VGGT — the backbone comparison is a consequence of studying that axis, not an independent variable.

Scoring and abstention

From the scaled affinity between refined descriptors, a row-softmax and a column-softmax give two conditional distributions; the DoubleSoftmax score is the log of their product, which is large only when each segment is the other's best match. Not every segment has a counterpart — some are occluded, out of frame, or absent — so each carries a matchability logit, and the score matrix is augmented with a learnable dustbin row and column. A segment is matched only if both its matchability passes 0.5 and its best score exceeds the dustbin.

Training

  • DataScanNet++, ground-truth correspondences from projected 3D instance annotations
  • SegmenterSAM, class-agnostic; up to Mmax = 100 segments per image
  • ObjectiveNormalised NLL over ground-truth pairs + binary matchability cross-entropy, λ = 0.3
  • EmbeddingD = 128 shared space; τ ≡ 1 for the MASt3R head, learned τ ≤ 1 for the VGGT heads
  • OptimiserAdamW, lr 10⁻⁴, weight decay 10⁻⁴, 500 warmup steps, cosine schedule, bf16
  • Hardware2 × NVIDIA RTX 5090 with DDP
  • Inference173 ms per joint window at N = 4; 290 ms at N = 6 (measured on the robot capture below)

Clamping the learned DoubleSoftmax temperature to τ ≤ 1 cuts the dominant false-dustbin errors by 36% relative, re-allocating the error budget toward harder wrong-match cases. For this head, calibrated scoring rather than richer descriptors is the binding constraint.

04 — Results

The winner depends on the viewpoint regime.

All models are trained only on ScanNet++ and evaluated zero-shot on Replica (indoor) and Virtual KITTI 2 (outdoor driving). Pairs are stratified into four bins by relative camera rotation. The headline metric is AUPRC; R@1 and R@5 are query-weighted across bins.

Model AUPRC by bin Overall
0–45°45–90°90–135°135–180° AUPRCR@1R@5
vizEnc+Radio41.9138.9941.2538.3640.6958.2079.04
LiftFeat47.1732.1325.1924.9235.6940.2167.57
vizEnc+DINOv250.7847.2449.4950.1549.4859.2980.29
DA3 (Sequential)69.3655.9138.9933.2354.2973.9673.96
SegMASt3R (Sinkhorn)80.4877.0479.8381.7279.6381.4594.21
SegVGGT-DPT (ours)81.2080.9378.0773.3679.7074.1090.11
SegVGGT-DPT Joint, N=4 (ours)81.6683.0381.8577.6981.3775.6490.42
SegMASt3R+LGv2 (ours)91.1783.5977.0970.2884.4877.4693.63

Table 1 — Replica (indoor, 3,200 pairs, 28,022 queries). Best per column in bold, second best underlined. LGv2 takes overall AUPRC with +4.85 over the Sinkhorn baseline, concentrated at narrow baselines (+10.7 at 0–45°) — but at 135–180° it drops 11.4 points below that baseline. The learned head sharpens the regime well represented in training and extrapolates worse to low-overlap, high-rotation pairs. There the Joint head takes over, beating the pairwise DPT variant by 4.3 AUPRC at the widest bin.

Model AUPRC by bin Overall
0–45°45–90°90–135°135–180° AUPRCR@1R@5
vizEnc+Radio29.0128.2923.4721.1227.5443.3866.89
vizEnc+DINOv231.8631.0525.5023.4030.2246.6870.34
LiftFeat35.4030.0027.2120.8832.3547.2168.04
DA3 (Sequential)43.6443.2935.8426.5340.6753.7475.40
SegMASt3R (Sinkhorn)57.9855.9657.2649.0656.9163.6781.57
SegVGGT-DPT (ours)46.6058.6968.7273.2952.2842.8067.27
SegVGGT-DPT Joint, N=4 (ours)47.5057.5470.1076.9252.8542.9366.75
SegMASt3R+LGv2 (ours)79.9488.5292.1678.3482.7873.2596.22

Table 2 — Virtual KITTI 2 (outdoor driving, 3,200 pairs, 10,308 queries), same protocol, also zero-shot. LGv2 reaches 82.78% overall, +25.87 over the Sinkhorn baseline and +30.50 over pairwise VGGT: MASt3R's explicit cross-view geometric coupling transfers from indoor training to outdoor driving better than VGGT's joint self-attention does. Both VGGT heads show an inverted, driving-specific angular pattern — they rise from the narrowest to the widest bin, consistent with difficulty separating near-identical segments at narrow baselines.

AUPRC by angular bin Replica (indoor) 30 40 50 60 70 80 90 0–45° 45–90° 90–135° 135–180° Relative camera rotation Virtual KITTI 2 (outdoor) 0–45° 45–90° 90–135° 135–180° Relative camera rotation AUPRC (%) best head changes here, past 90° LGv2 ahead in every bin SegMASt3R + LGv2 (ours) SegVGGT-DPT Joint, N=4 (ours) SegVGGT-DPT (ours) SegMASt3R (Sinkhorn) DA3 (Sequential)
Fig. 2 — No single head dominates. On Replica, LGv2 leads at narrow baselines but is overtaken by Joint, and by the Sinkhorn baseline, past 90°. On Virtual KITTI 2 it dominates every bin. The crossover is the central observation: the best head depends on the viewpoint regime.

In-domain check on real captured imagery

Replica and Virtual KITTI 2 are both rendered. We therefore also report ScanNet++ validation on real captured imagery from held-out scenes — indoor and in-domain rather than zero-shot.

MatcherAUPRCR@1R@5
SegMASt3R (Sinkhorn)88.983.397.1
SegVGGT-DPT (ours)89.682.297.6
SegVGGT-DPT Joint (ours)90.685.597.1
SegMASt3R+LGv2 (ours)91.085.498.5

Table 3 — ScanNet++ validation, held-out scenes. All three of our heads sit at or above the Sinkhorn baseline when the domain matches training.

05 — How many views

Extra views pay off only where overlap is scarce.

The Joint model is trained at N = 4 and evaluated at N ∈ {2, 4, 6, 8} without retraining. Position-embedding indices 4–7 are never directly supervised, yet the model extrapolates without collapse — consistent with joint attention learning view-relative rather than view-absolute patterns.

AUPRC vs number of jointly attended views note the different y-ranges Replica (indoor) 74 76 78 80 82 84 86 trained at N = 4 2 4 6 8 Number of jointly attended views N Virtual KITTI 2 (outdoor) 45 50 55 60 65 70 75 80 trained at N = 4 2 4 6 8 Number of jointly attended views N flat in N — Replica views already overlap +7.1 AUPRC, saturating by N = 6 AUPRC (%) Relative camera rotation 0–45° 45–90° 90–135° 135–180°
Fig. 3 — Effect of the number of jointly attended views. On Replica, N has almost no effect (81.10–81.57 overall AUPRC): indoor scenes have dense per-view overlap, so extra views add little. On Virtual KITTI 2, N helps specifically at wide baselines — the 135–180° bin improves from 71.05 (N=2) to 78.22 (N=6) before saturating, while narrow bins stay flat at around 47.5%. Note the different y-ranges between panels.
06 — On a real robot

Identity held across a moving camera.

The Joint model, trained on ScanNet++ at N = 4, applied without retraining to a sequence captured by a robot at Skoltech. A sliding window of consecutive frames is matched in one joint pass over 60 windows; one colour per instance, kept across all views the model matched it in. Masks are FastSAM proposals rather than the ground-truth masks used in the benchmarks above, and the capture has no ground-truth identities — so this is qualitative evidence, and it is also the only setting here where the segmenter is the one a deployed system would actually run.

Atrium sequence, N = 4. Thirteen instances held across the window, 22 links drawn. The shelving units, the yellow balustrade and the individual chairs keep their identities while the camera translates and the scale of every object changes — the case where appearance-only descriptors tend to collapse similar instances into one.
Joint window, N = 4. Four instances tracked, 11 links drawn in the window shown. 699 ms per batch of four, 173 ms per window.
The same sequence at N = 6, from the same checkpoint and with no retraining — the position-embedding indices beyond 3 were never directly supervised. Six instances, 21 links. Cost grows with the window: 1161 ms per batch, 290 ms per window. This is the Replica result of Fig. 3 seen from the other side: in a corridor with dense frame-to-frame overlap, the extra views mostly buy longer tracks rather than better matches, and the measured price of buying them is roughly 1.7× the runtime.
08 — Limitations

What this evidence does not cover.

  • Training uses indoor ScanNet++ only; an outdoor source may close the Virtual KITTI 2 gap.
  • Quantitative real-image evaluation is indoor and in-domain (ScanNet++ validation), not zero-shot. The in-house capture in Fig. 4 is qualitative, with no ground-truth identities. Real outdoor transfer remains untested.
  • Matching quality is bounded by the supplied masks. A segmenter-sensitivity study across SAM granularity and FastSAM is left to future work.
  • Closed-loop evaluation is confined to simulated indoor navigation.

Future work should test outdoor training data and extreme-rotation sampling, routing between pairwise and joint heads using deployment cues, real outdoor zero-shot transfer, and sensitivity to segmenter choice and granularity.

09 — Citation

BibTeX

@inproceedings{fatykhoph2026muviseg,
  title     = {MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors},
  author    = {Fatykhoph, Denis and Akhtyamov, Timur and Pakulev, Konstantin
               and Devchich, German and Ferrer, Gonzalo},
  booktitle = {Asian Conference on Computer Vision (ACCV)},
  year      = {2026}
}

The authors used Anthropic's Claude to assist with code development. All content was reviewed and verified by the authors, who take full responsibility for the final work.