SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

Danial Kamali1, Tanawan Premsri1, Shreya Rajpal1, Amir Zadeh2, Chuan Li2, Parisa Kordjamshidi1

1Michigan State University    2Lambda Labs

Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026

TL;DR Vision-language models often fail when a spatial question requires someone else's point of view. SATURN estimates a 3D scene from the images and answers with a short program of soft spatial predicates. Each predicate uses the appropriate frame of reference to judge left, right, front, or behind from a given position and facing direction. We also release 3D FORCE, a benchmark that controls reasoning depth, number of views, and how frames of reference are mixed.

One example, step by step

This walkthrough follows one question from 3D FORCE that SATURN answers correctly, while GPT-5.1, Gemini-3.1-Pro, Qwen3.5-9B, and Qwen3-VL-8B all answer incorrectly. Answering the question requires two frames of reference: the bus's own view and camera 0's view. Press ▶ to play or pause, use the numbered steps to jump, and drag the 3D view to rotate the scene while the animation is paused.

Loading the example…

The detections, the reconstructed scene and cameras, and the program come from one SATURN run of the released code (seed 0; Qwen3-VL-8B as the VLM, SAM3, VGGT and Orient Anything V2 for perception), and every score comes from the execution trace of that run. The ground shading shows only which side each frame calls “behind”, while the numbers are the predicate values.

The baseline answers are the models' own outputs from the paper's evaluation, scored by the paper's rule (IoU > 0.5). This question comes from the “Obj + One-Cam” REF setting, which combines one object frame with one camera frame. In that setting, SATURN reaches 73% and Gemini-3.1-Pro reaches 41% (Figure 4 of the paper).

The same scene from each perspective

To make the two frames of reference concrete, we rendered the example scene again with the 3D FORCE generator. The camera visits the three input views, rises to a bird's-eye view with the bus facing the top of the picture, and then returns to camera 0. The clip uses the benchmark's ground-truth scene to show what the question means, not what SATURN estimated. SATURN only sees the three input views.

In the bus's frame, the green tank on the left in camera 0 (the answer) lies in the lower half of the picture, behind the bus. The green tank on the right in camera 0 lies in the upper half, so that tank is not behind the bus.

Abstract

Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existing tool-augmented spatial reasoning methods make reasoning more explicit, but often rely on low-level geometric procedures and hard binary decisions over noisy perception. We propose SATURN, a neuro-symbolic framework for perspective-aware compositional spatial reasoning. SATURN reconstructs an approximate 3D scene, derives soft perspective-aware spatial predicates, and composes them with a training-free Pythonic symbolic executor, separating perception from reasoning while preserving uncertainty through multi-hop inference. We also introduce 3D FORCE, a diagnostic benchmark that controls reasoning depth, view, and perspective composition for spatial arrangement grounding (SAG) and referring expression grounding (REF). On 3D FORCE, VLMs and spatially trained models degrade sharply as depth and perspective complexity increase, whereas SATURN degrades the least and outperforms every baseline at each depth. On the real-world MindCube benchmark, SATURN achieves 78.06% overall accuracy, outperforming the strongest baseline by 14 percentage points.

How SATURN works

SATURN uses a VLM to identify which objects the question needs and a code LLM to write a short program over an estimated 3D scene. The program uses soft predicates in frames defined by a position and facing direction at an object or camera. A soft-logic engine runs the program and selects the object or option with the highest score for the claim.

SATURN overview: query-guided scene estimation, spatial engine, and program generation and execution
SATURN overview (Figure 2 of the paper).
1Query-guided scene estimationA VLM parser reads the question and lists the required objects. SAM3 grounds those objects in every view, VGGT reconstructs the objects' 3D positions and the cameras, and Orient Anything V2 estimates which way each object faces. SATURN also uses any facts the question states about the cameras to refine the camera poses, such as two views taken from the same spot or a fixed rotation between views.
2Spatial engineEvery frame of reference is a local coordinate system attached to a camera, an object, or a virtual viewer. The spatial engine computes relations such as left, front, or behind in that frame and assigns soft scores in [0, 1] rather than hard yes/no decisions.
3Program and soft executionA code LLM writes a short Python program that uses the VLM to score semantic concepts and calls the frame-aware predicates to score spatial relations. The executor combines these scores with soft logic (AND is the minimum), allowing uncertainty from perception to carry through every hop rather than being thresholded away.

The 3D FORCE benchmark

3D FORCE isolates the reasoning needed to compose spatial relations across perspectives. High-resolution rendered scenes keep perception and object grounding simple, while the benchmark controls reasoning depth, relation topology, view count, and the frame of reference for each relation. Every question comes with a formal logical form, and answers come from the scene graph used to render the images.

The benchmark has two subsets. SAG (spatial arrangement grounding) asks whether a described arrangement of objects exists in the scene. REF (referring expression grounding) asks for a box around the object identified by a multi-hop description.

2,088REF questions
1,150SAG questions
1–4views, plus partial views
0–6relation hops; chain, star, hybrid
Overview of the 3D FORCE benchmark with its SAG and REF subsets
Overview of 3D FORCE (Figure 3 of the paper). Download the benchmark from Hugging Face.

Key results

3D FORCE

MethodMicro Avg.SAGREF
Random choice18.8050.171.53
General-purpose VLMs
Qwen3-VL-8B24.3165.911.39
Qwen3-VL-235B26.4473.220.67
Qwen3.5-9B57.7574.8748.32
Qwen3.5-35B57.6376.3547.32
InternVL3.5-38B30.8258.2615.71
GPT-5.129.8071.396.90
Gemini-3.1-Pro57.6375.3047.89
Spatially trained VLMs
SpaceOm-3B17.7949.390.38
Cosmos-Reason1-7B13.4737.040.48
Tool-augmented
GCA24.2451.009.50
pySpatial44.8448.3542.91
TIGeR20.3830.4314.85
SATURN (ours)82.89 ±0.4085.88 ±0.1381.24 ±0.68
SATURN with ground-truth boxes88.85 ±0.5587.57 ±0.9789.56 ±0.67
SATURN with oracle 3D95.94 ±0.3594.40 ±0.6896.79 ±0.39

Selected rows from Table 1 in the paper, showing accuracy (%). REF counts a box as correct at IoU > 0.5, and SATURN rows report the mean and standard deviation over three runs. The full table lists all 19 VLMs.

Accuracy on SAG and REF by perspective setting for InternVL3.5-38B, Gemini-3.1-Pro, Qwen3.5-35B, and SATURN
Accuracy by perspective setting (Figure 4). Every VLM performs best with a single camera frame, with accuracy dropping when an object frame or a second camera frame is involved. Across settings, SATURN varies by about 7 points on SAG and 13 on REF, while each VLM varies by at least 22 points.
REF accuracy by number of reasoning hops
REF accuracy by number of reasoning hops (Figure 5). From zero to six hops, accuracy falls from 94% to 67% for SATURN, from 80% to 33% for Gemini-3.1-Pro, and from 74% to 29% for Qwen3.5-35B. SATURN outperforms every baseline at each hop count.

Why soft predicates

The paper compares three variants using the same scenes, built from the benchmark's ground-truth boxes. Writing low-level geometry code directly reaches 75.43% on REF, while using the declarative predicate interface with crisp 0/1 scores reaches 79.07%. Using the same predicates with continuous scores raises accuracy to 89.6%, so preserving uncertainty accounts for the larger gain (+10.5 points).

Real-world benchmarks: MindCube and MMSI

MethodMindCubeRotationAmongAroundMMSI
Qwen3-VL-8B-Instruct38.8643.5032.8349.6029.40
Qwen3-VL-235B-Thinking47.3087.0035.0047.3032.60
Gemini-2.5-Pro57.5089.5048.8054.5036.90
GCA (235B-Think)64.2082.0059.8061.8041.90
pySpatial53.1438.5050.3071.6028.20
Qwen3VL + scene state60.8649.5060.5070.8039.40
SATURN (ours)78.06 ±1.0585.67 ±2.9377.00 ±1.0274.53 ±2.3148.77 ±0.61

Selected rows from Table 2 in the paper, showing accuracy (%). The MMSI result is the overall score; the paper also reports the four MMSI categories and paired confidence intervals.

BibTeX

@misc{kamali2026saturnsymbolicspatialreasoning,
      title={SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding},
      author={Danial Kamali and Tanawan Premsri and Shreya Rajpal and Amir Zadeh and Chuan Li and Parisa Kordjamshidi},
      year={2026},
      eprint={2606.22694},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2606.22694},
}