Hear the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation
Sound should follow its visible source as the world moves.
THE RESEARCH QUESTION
When the world moves,
does sound follow?
Abstract
Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.
TWO WORLDS, ONE SCENE
What we see and hear
must occupy the same space.
A compelling audio-visual scene is more than a synchronized picture and soundtrack. Sound should be associated with the right source, at the right position, as that source moves through time.
Where is the source?
Objects, depth, scene geometry, and visible motion provide spatial cues.
Where does it sound?
Direction, balance, and acoustic change provide the corresponding auditory cues.
DATA CONSTRUCTION
Teaching sound
where to move.
StereoWorld-29K pairs stereo audio-video data with corresponding motion tracks through data curation, source discovery and verification, motion tracking, physically grounded 3D audio rendering, and diversity-driven synthesis.
FIG. 02 Construction pipeline for StereoWorld-29K. Open the figure to inspect the full-resolution diagram.
TECHNICAL OVERVIEW
From correspondence
to generation.
Motion tracks condition StereoBind at complementary levels: VMB Tokens establish entity-level audio-visual binding, STE encodes absolute source trajectories, and RT-RoPE models trajectory-induced relative spatial relations.
FIG. 03 StereoBind architecture with VMB Tokens, Spatial Track Encoder, and Residual Track RoPE.
QUALITATIVE RESULTS
Follow the source
in stereo.
Examples are organized by spatial state and motion direction. Listen with headphones and compare the visible source trajectory with the perceived stereo position.
Wear headphones for the best experience
Static — Left
The visible source remains on the left, and the generated sound should stay left-localized.
Static — Right
The visible source remains on the right, and the generated sound should stay right-localized.
Dynamic — Left to Right
The visible source moves from left to right, with stereo position following the same trajectory.
Dynamic — Right to Left
The visible source moves from right to left, with stereo position following the same trajectory.
Dynamic — Left → Right → Left
The source reverses direction after reaching the right side, testing whether the acoustic trajectory turns with it.
Dynamic — Right → Left → Right
The source reverses direction after reaching the left side, testing motion-aligned stereo across the full path.
QUALITATIVE COMPARISON
Six scenes.
Six systems.
Each row uses the same input across all six systems. Compare visual fidelity, audio quality, and whether the perceived stereo position follows the visible source motion.
Headphones recommended. The first column is StereoBind. On smaller screens, scroll horizontally to compare every model within the same sample.
QUANTITATIVE RESULTS
Measured across
quality and space.
We compare visual quality, audio quality, audio-visual synchronization, and stereo spatial fidelity. Best results are shown in bold and second-best results are underlined.
TABLE 01 Comparison results on joint video-audio generation. Open the table to inspect all metrics at full resolution.
THE PAPER
Build on
this work.
The manuscript contains the complete formulation, model details, experimental protocol, and evaluation.
