Preprint · 2026

Matisse: Evidence-Space Reasoning
for Active 3D Reconstruction

A unified, training-free framework for active view acquisition and keyframe selection.

1 Massachusetts Institute of Technology2 National University of Singapore
Selected camera views and reconstructed object meshes in a real robot scene
Matisse in the real world. Aggregated depth with selected views and object boxes (left), the captured scene (upper right), and reconstructed meshes (lower right).

Presentation

Matisse in video

A visual overview of evidence-space reasoning for efficient active 3D reconstruction.

Matisse project video · Watch on YouTube ↗

Abstract

How can a 3D reconstruction system acquire and retain useful information to understand a scene from partial views under a limited computation budget? Existing active-view methods typically estimate uncertainty only over observed geometry, while long-horizon methods often retain redundant observations.

We introduce Matisse, a training-free framework that unifies active reconstruction and keyframe selection using evidence from a pretrained generative 3D model. Matisse estimates Evidential Uncertainty from cross-attention evidence and derives Evidential Information Gain to choose views according to their expected reduction in posterior entropy.

01

Reason beyond observed surfaces

Generative tokens represent visible and completed geometry, extending uncertainty into unseen regions.

02

One criterion, two decisions

The same information gain selects the next camera and useful frames from long sequences.

03

Plan without full reconstruction

Candidate views are scored in evidence space at an amortized 1.26 ms per candidate.

Method at a glance

How Matisse works

Matisse interprets view-to-token association weights probabilistically. Evidence at each aligned 3D token gives its posterior uncertainty; candidate cameras are ranked by how much uncertainty they are expected to remove.

01Observe
Initial observation

Initialize from a partial view

A generative model hypothesizes visible and occluded shape.

→
02Estimate
Uncertainty and initial mesh

Measure uncertainty

Weakly supported 3D tokens remain uncertain.

→
03Select
Two selected camera views

Acquire the best view

Visible uncertainty and exploration rank cameras.

→
04Update
Completed reconstruction

Refine and retain

New evidence updates shape and retained keyframes.

Evidential uncertaintyUt[q]=αα+Et[q]

Uncertainty falls as accumulated evidence grows for token q.

Approximate information gainIG^t(c)=∑q∈Ω(c)Ut[q]

A view is informative when it sees many uncertain tokens.

Experiments

Results

We evaluate sparse-view reconstruction on GSO30, YCB-V, and Replica: single objects, cluttered tabletops, and room-scale environments.

12.7%lower Chamfer distance
on GSO30
1.50×faster end-to-end
active reconstruction
7 / 48views to converge
at comparable quality

Qualitative comparison on GSO30

Matisse reconstructs geometry and appearance more faithfully than active-reconstruction baselines.

GTRandomFisherRFGauSS-MIGAVISMAGICIANMatisse
Lunch bag reconstruction comparisonSchool bus reconstruction comparisonTrain reconstruction comparisonTurtle reconstruction comparison

Quantitative evaluation

GSO30 active reconstruction

Compact comparison using the Stream3D reconstruction backend.

MethodCD ↓IoU ↑P-FID ↓Runtime ↓
Random62.0780.73648.45616.37 s
FisherRF65.8660.71653.44355.57 s
GauSS-MI88.6760.68664.83153.88 s
GAVIS87.5980.68764.39495.68 s
MAGICIAN63.6660.73350.17523.97 s
Matisse-G Ours54.1670.76545.35416.02 s

Keyframe selection

Comparable quality with one-seventh of the views

Matisse selects 7 informative keyframes and reaches the quality Stream3D obtains after processing 48 sequential views.

Chamfer distance convergence for Matisse and Stream3D

Citation

BibTeX

matisse.bib
@article{yu2026matisse,
  title={Matisse: Evidence-Space Reasoning for Active 3D Reconstruction},
  author={Yu, Xihang and Zhou, Kaichen and Shaikewitz, Lorenzo and Jambon, Cl{\'e}ment and Zhan, Xiao and Talak, Rajat and Carlone, Luca},
  journal={arXiv preprint arXiv:2609.38746},
  year={2026}
}