Exhaustive 4D instance segmentation within feed-forward visual geometry.
Segment Anything models work on 2D image or video masks and carry object identity through sequential memory. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry: from any set of RGB observations, a space-time query decoder turns shared visual-geometric features into exhaustive, class-agnostic 4D instance masks, each persistent object query binding one entity across all observations. Any instance can also be selected by a point or box prompt, with no seed mask or temporal ordering required, and an agentic harness grounds natural-language descriptions in large 4D scenes.
Top: learnable object queries decode every instance in 4D; a point or box prompt selects one. Bottom: Active Tree Search finds the relevant observations, the Grounder proposes a query match and a box, and the Critic keeps the better 4D mask.
Each queried object is highlighted in turn in the reconstructed 4D scene.
@article{zhang2026s4vy,
title={S4VY: Segment Anything in Feed-Forward 4D Visual Geometry},
author={Zhang, Jingdong and Li, Xin and Kautz, Jan and Wang, Wenping and Choy, Chris},
journal={arXiv preprint arXiv:2609.36875},
year={2026}
}