Current vision-language models (VLMs) demonstrate strong capabilities on 2D perception and captioning,
but remain brittle when asked to perform 3D-aware reasoning from a single image, such as estimating
distances, sizes, occlusions, and navigation feasibility.
We introduce MonoSR, a large-scale benchmark for
open-vocabulary spatial reasoning from monocular RGB images.
MonoSR spans diverse domains including indoor scenes, outdoor driving and navigation,
and object-centric views. For each image, we construct structured 3D scene graphs with
aligned 2D/3D geometry and generate hierarchical question–answer pairs that probe
(1) low-level geometric perception, (2) mid-level perspective-aware reasoning, and
(3) high-level situational understanding.
We evaluate a range of state-of-the-art VLMs and 3D-aware models, and show that even
the strongest systems fall far below human performance on challenging monocular spatial
reasoning tasks, revealing substantial headroom for 3D-grounded multimodal reasoning
from a single view.