Beyond Thinking: Imagining in 360° for Humanoid Visual Search

1Texas A&M University    2Adobe    *Corresponding author
TAMU Logo Adobe Logo
NeurIPS 2026
Agent’s view360° scene · explored so far
“Find the pink suitcase.” · found in 3 steps

Zero-shot search in the wild: imagining the unseen part of the scene tells the agent where to turn its head.

Abstract

Humanoid Visual Search asks an agent to find a target in an immersive 360° scene while it sees only a narrow field of view at a time. Prior methods treat it as one monolithic task driven by cumulative, multi-turn chain-of-thought, which is cognitively heavy and needs expensive trajectory-level annotations. We propose Imagining in 360°, which decouples exploration into an Imaginator and an Actor. In a single step, the Imaginator infers the semantic layout of both observed and unobserved regions; sampling several hypotheses gives the Actor a distribution of spatial priors that hedges against uncertainty during the search.

  • Decoupled paradigm: spatial imagination is separated from action planning, for more active and grounded exploration.
  • Probabilistic Imaginator: learns intuitive spatial priors and guides any Actor with sampled estimates of where things are.
  • Scalable data engine: a fully automated pipeline produces 1.92M training samples without manual trajectory labels.

Method

Imagine first, then act: one model writes down the whole scene, another decides where to look.

Method: the Imaginator predicts a 360 degree semantic layout in one pass, and its sampled suggestions guide a frozen Actor step by step IMAGINATOR · ONE STATELESS PASS VIEWS SO FAR (295,6) · (267,13) · (47,42) “Where is the bed with wooden posts?” instruction Imaginator Qwen3-VL-8B no history, one pass [Observed] ◆ open white bedroom door: (25,0) ◆ rocking chair with stuffed toys: (270,-23) ◆ wall-mounted picture frames: (274,16) ◆ wall-mounted coat rack: (292,0) ◆ window seat with cushion: (227,-9) [Imagined] ◇ carpeted floor pathway: (0,-45) ◇ arched doorway to adjacent room: (356,-2) ◇ bed with wooden posts: (126,-13) ◇ window with outdoor view: (166,9) ◇ ceiling light fixture: (148,52) suggest check(126,−13) 1 2 3 open white bedroom door rocking chair with stuffed toys wall-mounted picture frames wall-mounted coat rack window seat with cushion carpeted floor pathway arched doorway to adjacent room bed with wooden posts window with outdoor view ceiling light fixture check(126,−13) UNFOLDED 360° SCENE 180°270°0°90°180° observed imagined suggested target coordinates are absolute (yaw, pitch) in degrees SEARCH LOOP · THE IMAGINATOR SUGGESTS, THE ACTOR DECIDES View at step t current direction Imaginator 3 samples: 1 greedy + 2 at T RANKED SUGGESTIONS 1 rotate(Δφ, Δγ) 2 rotate(Δφ′, Δγ′) 3 submit(φ, γ) Actor any VLM, not fine-tuned weighs them against its view rotate(Δφ, Δγ) → step t+1 submit(φ, γ) → done the next view goes back to the Imaginator STEP 1STEP 2STEP 3 “Find the pink suitcase.” hypotheses narrow as more is seen; found in 3 steps

Top: from the views seen so far, the Imaginator lists observed landmarks and imagines unseen ones at absolute (yaw, pitch), then suggests where the target is. Bottom: at every step three such hypotheses are sampled and handed to an unchanged Actor as ranked suggestions; they narrow onto the target as more of the scene is observed.

Imagination, Step by Step

Imagined target hypotheses on the 360 degree panorama
Imagined target hypotheses · explored views in green
Imaginator suggests
    Actor

    Three hypotheses are sampled at every step. They spread over plausible regions at first and converge on the target as more of the scene is observed.

    Scalable Data Engine

    Single-step layout supervision, generated from unlabeled panoramas.

    Panoramas
    ~20k
    SUN360, Matterport3D, DiT360
    Training samples
    1.92M
    observed + imagined layouts
    Landmarks per scene
    11
    median, observed and imagined
    Scene categories Outdoor urban31.6% Residential interior25.9% Other / mixed10.6% Museum & culture7.5% Hotel & hospitality6.1% Retail & commerce4.0% Food & beverage3.7% Transportation hub3.5% Office & workspace2.4% Nature & landscape2.4% Entertainment & leisure1.8% Industrial & warehouse0.5% Panorama sources SUN36050.5% Matterport3D24.8% DiT360 (generated)24.7% Labeled landmark types Object43.8% Pathway18.3% Other18.2% Exit / entrance10.8% Sign & text6.9% Stairs / escalator1.8%

    Qwen3-VL-235B filters the panoramas and proposes salient targets; search trajectories rendered around them turn every step into one sample of observed and imagined layout.

    Results

    A plug-and-play gain for every Actor on H*Bench.

    ActorParamsObject Search (HOS)Path Search (HPS)
    Actor+ ImaginatorGainActor+ ImaginatorGain
    Open-source
    Qwen2.5-VL3B21.1749.54+28.47.6925.94+18.2
    Qwen2.5-VL7B12.7957.42+44.69.8828.38+18.5
    Qwen2.5-VL72B29.5462.25+32.712.9438.19+25.2
    Qwen3-VL32B34.3857.46+23.123.1931.31+8.1
    Qwen3-VL235B23.6265.04+41.421.4438.25+16.8
    Gemma-34B18.1757.75+39.614.5035.06+20.6
    Gemma-312B22.8854.62+31.715.3133.56+18.2
    Kimi-VLA3B4.3332.46+28.15.6221.00+15.4
    Proprietary
    Gemini-2.5 Flash–43.7063.20+19.527.7037.70+10.0
    Gemini-2.5 Pro–55.0066.10+11.137.9043.10+5.2
    Fine-tuned for HVS
    HVS3B48.0462.75+14.724.1239.38+15.3

    Success rate (%) within 10 steps. The same Imaginator-8B guides every Actor, none of which is fine-tuned on its output.

    Object search (HOS) 46 48 50 52 54 56 0 0.38 0.77 1.15 1.92 imagination data (M samples) success rate (%) Actor + SFT + RL · 48.04 49.54 52.54 53.04 54.08 54.88 Actor + Imaginator Path search (HPS) 22 24 26 28 30 32 0 0.38 0.77 1.15 1.92 imagination data (M samples) success rate (%) Actor + SFT + RL · 24.12 25.94 28.75 29.44 30.94 31.31 Actor + Imaginator

    More imagination data keeps helping; 0.38M samples already beat the Actor fine-tuned with SFT and RL.

    Video

    BibTeX

    @inproceedings{zhang2026imagining,
      title={Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search},
      author={Zhang, Jingdong and Wang, Yizhou and Tu, Zhengzhong and Li, Xin and Wang, Wenping and Zhan, Xiaohang},
      booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
      year={2026}
    }