S4VY: Segment Anything in Feed-Forward 4D Visual Geometry

1Texas A&M University    2NVIDIA
TAMU Logo NVIDIA Logo
Input observations4D instance masks
Roundabout

Exhaustive 4D instance segmentation within feed-forward visual geometry.

Abstract

Segment Anything models work on 2D image or video masks and carry object identity through sequential memory. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry: from any set of RGB observations, a space-time query decoder turns shared visual-geometric features into exhaustive, class-agnostic 4D instance masks, each persistent object query binding one entity across all observations. Any instance can also be selected by a point or box prompt, with no seed mask or temporal ordering required, and an agentic harness grounds natural-language descriptions in large 4D scenes.

  • Feed-forward Segment Anything in 4D: shared visual-geometric features become exhaustive, class-agnostic 4D instances with persistent identities, and any of them can be selected by a point or box prompt.
  • Agentic language grounding in 4D: dual-stream prediction, active tree search and comparative critic selection locate a described object across large observation sets.
  • Unified evaluation: state-of-the-art 4D instance segmentation and strong language grounding across static and dynamic scenes.

Method

Architecture of S4VY: feed-forward 4D segmentation (top) and agentic language grounding (bottom) SEGMENT ANYTHING IN 4D RGB observations any number, any order point / box prompt Visual Geometry Backbone feed-forward LoRA Prompt Encoder object queries promptable query Space-Time Query Decoder 4D instance segmentation every object Prompted 4D instance from a point or box AGENTIC LANGUAGE GROUNDING “the black car moving through the roundabout” ×✓ ✓× Active Tree Search visits only relevant observations Grounder VLM, dual-stream object-query match box prediction AB Critic 4D grounding the better of the two candidates

Top: learnable object queries decode every instance in 4D; a point or box prompt selects one. Bottom: Active Tree Search finds the relevant observations, the Grounder proposes a query match and a box, and the Critic keeps the better 4D mask.

Language-Guided 4D Grounding

Tabletop
  1. a yellow bottle being picked up by a hand
  2. gray scissors on the dark wood table
  3. a mug on the light-colored small cabinet

Each queried object is highlighted in turn in the reconstructed 4D scene.

Demo Video

Coming soon

BibTeX

@article{zhang2026s4vy,
  title={S4VY: Segment Anything in Feed-Forward 4D Visual Geometry},
  author={Zhang, Jingdong and Li, Xin and Kautz, Jan and Wang, Wenping and Choy, Chris},
  journal={arXiv preprint arXiv:2609.36875},
  year={2026}
}