SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding
1Michigan State University 2Lambda Labs
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
One example, step by step
This walkthrough follows one question from 3D FORCE that SATURN answers correctly, while GPT-5.1, Gemini-3.1-Pro, Qwen3.5-9B, and Qwen3-VL-8B all answer incorrectly. Answering the question requires two frames of reference: the bus's own view and camera 0's view. Press ▶ to play or pause, use the numbered steps to jump, and drag the 3D view to rotate the scene while the animation is paused.
Loading the example…
The detections, the reconstructed scene and cameras, and the program come from one SATURN run of the released code (seed 0; Qwen3-VL-8B as the VLM, SAM3, VGGT and Orient Anything V2 for perception), and every score comes from the execution trace of that run. The ground shading shows only which side each frame calls “behind”, while the numbers are the predicate values.
The baseline answers are the models' own outputs from the paper's evaluation, scored by the paper's rule (IoU > 0.5). This question comes from the “Obj + One-Cam” REF setting, which combines one object frame with one camera frame. In that setting, SATURN reaches 73% and Gemini-3.1-Pro reaches 41% (Figure 4 of the paper).
The same scene from each perspective
To make the two frames of reference concrete, we rendered the example scene again with the 3D FORCE generator. The camera visits the three input views, rises to a bird's-eye view with the bus facing the top of the picture, and then returns to camera 0. The clip uses the benchmark's ground-truth scene to show what the question means, not what SATURN estimated. SATURN only sees the three input views.
In the bus's frame, the green tank on the left in camera 0 (the answer) lies in the lower half of the picture, behind the bus. The green tank on the right in camera 0 lies in the upper half, so that tank is not behind the bus.
Abstract
Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existing tool-augmented spatial reasoning methods make reasoning more explicit, but often rely on low-level geometric procedures and hard binary decisions over noisy perception. We propose SATURN, a neuro-symbolic framework for perspective-aware compositional spatial reasoning. SATURN reconstructs an approximate 3D scene, derives soft perspective-aware spatial predicates, and composes them with a training-free Pythonic symbolic executor, separating perception from reasoning while preserving uncertainty through multi-hop inference. We also introduce 3D FORCE, a diagnostic benchmark that controls reasoning depth, view, and perspective composition for spatial arrangement grounding (SAG) and referring expression grounding (REF). On 3D FORCE, VLMs and spatially trained models degrade sharply as depth and perspective complexity increase, whereas SATURN degrades the least and outperforms every baseline at each depth. On the real-world MindCube benchmark, SATURN achieves 78.06% overall accuracy, outperforming the strongest baseline by 14 percentage points.
How SATURN works
SATURN uses a VLM to identify which objects the question needs and a code LLM to write a short program over an estimated 3D scene. The program uses soft predicates in frames defined by a position and facing direction at an object or camera. A soft-logic engine runs the program and selects the object or option with the highest score for the claim.

The 3D FORCE benchmark
3D FORCE isolates the reasoning needed to compose spatial relations across perspectives. High-resolution rendered scenes keep perception and object grounding simple, while the benchmark controls reasoning depth, relation topology, view count, and the frame of reference for each relation. Every question comes with a formal logical form, and answers come from the scene graph used to render the images.
The benchmark has two subsets. SAG (spatial arrangement grounding) asks whether a described arrangement of objects exists in the scene. REF (referring expression grounding) asks for a box around the object identified by a multi-hop description.

Key results
3D FORCE
| Method | Micro Avg. | SAG | REF |
|---|---|---|---|
| Random choice | 18.80 | 50.17 | 1.53 |
| General-purpose VLMs | |||
| Qwen3-VL-8B | 24.31 | 65.91 | 1.39 |
| Qwen3-VL-235B | 26.44 | 73.22 | 0.67 |
| Qwen3.5-9B | 57.75 | 74.87 | 48.32 |
| Qwen3.5-35B | 57.63 | 76.35 | 47.32 |
| InternVL3.5-38B | 30.82 | 58.26 | 15.71 |
| GPT-5.1 | 29.80 | 71.39 | 6.90 |
| Gemini-3.1-Pro | 57.63 | 75.30 | 47.89 |
| Spatially trained VLMs | |||
| SpaceOm-3B | 17.79 | 49.39 | 0.38 |
| Cosmos-Reason1-7B | 13.47 | 37.04 | 0.48 |
| Tool-augmented | |||
| GCA | 24.24 | 51.00 | 9.50 |
| pySpatial | 44.84 | 48.35 | 42.91 |
| TIGeR | 20.38 | 30.43 | 14.85 |
| SATURN (ours) | 82.89 ±0.40 | 85.88 ±0.13 | 81.24 ±0.68 |
| SATURN with ground-truth boxes | 88.85 ±0.55 | 87.57 ±0.97 | 89.56 ±0.67 |
| SATURN with oracle 3D | 95.94 ±0.35 | 94.40 ±0.68 | 96.79 ±0.39 |
Selected rows from Table 1 in the paper, showing accuracy (%). REF counts a box as correct at IoU > 0.5, and SATURN rows report the mean and standard deviation over three runs. The full table lists all 19 VLMs.


Why soft predicates
The paper compares three variants using the same scenes, built from the benchmark's ground-truth boxes. Writing low-level geometry code directly reaches 75.43% on REF, while using the declarative predicate interface with crisp 0/1 scores reaches 79.07%. Using the same predicates with continuous scores raises accuracy to 89.6%, so preserving uncertainty accounts for the larger gain (+10.5 points).
Real-world benchmarks: MindCube and MMSI
| Method | MindCube | Rotation | Among | Around | MMSI |
|---|---|---|---|---|---|
| Qwen3-VL-8B-Instruct | 38.86 | 43.50 | 32.83 | 49.60 | 29.40 |
| Qwen3-VL-235B-Thinking | 47.30 | 87.00 | 35.00 | 47.30 | 32.60 |
| Gemini-2.5-Pro | 57.50 | 89.50 | 48.80 | 54.50 | 36.90 |
| GCA (235B-Think) | 64.20 | 82.00 | 59.80 | 61.80 | 41.90 |
| pySpatial | 53.14 | 38.50 | 50.30 | 71.60 | 28.20 |
| Qwen3VL + scene state | 60.86 | 49.50 | 60.50 | 70.80 | 39.40 |
| SATURN (ours) | 78.06 ±1.05 | 85.67 ±2.93 | 77.00 ±1.02 | 74.53 ±2.31 | 48.77 ±0.61 |
Selected rows from Table 2 in the paper, showing accuracy (%). The MMSI result is the overall score; the paper also reports the four MMSI categories and paired confidence intervals.
BibTeX
@misc{kamali2026saturnsymbolicspatialreasoning,
title={SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding},
author={Danial Kamali and Tanawan Premsri and Shreya Rajpal and Amir Zadeh and Chuan Li and Parisa Kordjamshidi},
year={2026},
eprint={2606.22694},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.22694},
}