Zero-shot search in the wild: imagining the unseen part of the scene tells the agent where to turn its head.
Humanoid Visual Search asks an agent to find a target in an immersive 360° scene while it sees only a narrow field of view at a time. Prior methods treat it as one monolithic task driven by cumulative, multi-turn chain-of-thought, which is cognitively heavy and needs expensive trajectory-level annotations. We propose Imagining in 360°, which decouples exploration into an Imaginator and an Actor. In a single step, the Imaginator infers the semantic layout of both observed and unobserved regions; sampling several hypotheses gives the Actor a distribution of spatial priors that hedges against uncertainty during the search.
Imagine first, then act: one model writes down the whole scene, another decides where to look.
Top: from the views seen so far, the Imaginator lists observed landmarks and imagines unseen ones at absolute (yaw, pitch), then suggests where the target is. Bottom: at every step three such hypotheses are sampled and handed to an unchanged Actor as ranked suggestions; they narrow onto the target as more of the scene is observed.
Three hypotheses are sampled at every step. They spread over plausible regions at first and converge on the target as more of the scene is observed.
Single-step layout supervision, generated from unlabeled panoramas.
Qwen3-VL-235B filters the panoramas and proposes salient targets; search trajectories rendered around them turn every step into one sample of observed and imagined layout.
A plug-and-play gain for every Actor on H*Bench.
| Actor | Params | Object Search (HOS) | Path Search (HPS) | ||||
|---|---|---|---|---|---|---|---|
| Actor | + Imaginator | Gain | Actor | + Imaginator | Gain | ||
| Open-source | |||||||
| Qwen2.5-VL | 3B | 21.17 | 49.54 | +28.4 | 7.69 | 25.94 | +18.2 |
| Qwen2.5-VL | 7B | 12.79 | 57.42 | +44.6 | 9.88 | 28.38 | +18.5 |
| Qwen2.5-VL | 72B | 29.54 | 62.25 | +32.7 | 12.94 | 38.19 | +25.2 |
| Qwen3-VL | 32B | 34.38 | 57.46 | +23.1 | 23.19 | 31.31 | +8.1 |
| Qwen3-VL | 235B | 23.62 | 65.04 | +41.4 | 21.44 | 38.25 | +16.8 |
| Gemma-3 | 4B | 18.17 | 57.75 | +39.6 | 14.50 | 35.06 | +20.6 |
| Gemma-3 | 12B | 22.88 | 54.62 | +31.7 | 15.31 | 33.56 | +18.2 |
| Kimi-VL | A3B | 4.33 | 32.46 | +28.1 | 5.62 | 21.00 | +15.4 |
| Proprietary | |||||||
| Gemini-2.5 Flash | – | 43.70 | 63.20 | +19.5 | 27.70 | 37.70 | +10.0 |
| Gemini-2.5 Pro | – | 55.00 | 66.10 | +11.1 | 37.90 | 43.10 | +5.2 |
| Fine-tuned for HVS | |||||||
| HVS | 3B | 48.04 | 62.75 | +14.7 | 24.12 | 39.38 | +15.3 |
Success rate (%) within 10 steps. The same Imaginator-8B guides every Actor, none of which is fine-tuned on its output.
More imagination data keeps helping; 0.38M samples already beat the Actor fine-tuned with SFT and RL.
@inproceedings{zhang2026imagining,
title={Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search},
author={Zhang, Jingdong and Wang, Yizhou and Tu, Zhengzhong and Li, Xin and Wang, Wenping and Zhan, Xiaohang},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026}
}