STEREOBIND · JOINT VIDEO-AUDIO GENERATION

Hear the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation

Sound should follow its visible source as the world moves.

Paper ↗ Code URL pending Explore results
01 / ABSTRACTRESEARCH OVERVIEW

THE RESEARCH QUESTION

When the world moves,
does sound follow?

Abstract

Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sound location to evolve consistently with the motion of its corresponding visual source. We refer to this property as Dynamic Spatial Correspondence and propose StereoBind, a framework that binds visual source motion to stereo sound generation. StereoBind uses motion tracks to coordinate visual motion and stereo audio through three complementary mechanisms. Visual Motion Binding establishes source-aware audiovisual correspondence, the Spatial Track Encoder captures absolute source positions, and Residual Track RoPE models relative motion. For supervision and evaluation, we construct StereoWorld-29K, a large-scale stereo audio-video dataset with paired motion tracks, and StereoWorldBench for measuring audiovisual spatial consistency. Experiments show that StereoBind substantially improves spatial alignment in stereo audio generation over existing models while preserving overall audiovisual quality.

02 / CORRESPONDENCETHE CENTRAL QUESTION

TWO WORLDS, ONE SCENE

What we see and hear
must occupy the same space.

A compelling audio-visual scene is more than a synchronized picture and soundtrack. Sound should be associated with the right source, at the right position, as that source moves through time.

01 VISUAL WORLD

Where is the source?

Objects, depth, scene geometry, and visible motion provide spatial cues.

02 ACOUSTIC WORLD

Where does it sound?

Direction, balance, and acoustic change provide the corresponding auditory cues.

SPACE
TIMETHE CONNECTION IS A TRAJECTORY, NOT A SINGLE FRAME.
WATCH & LISTENHEADPHONES RECOMMENDED
03 / TEASERDYNAMIC SPATIAL CORRESPONDENCE

THE CORE IDEA

See the motion.
Hear it move.

StereoBind uses motion tracks to bind a visible source’s trajectory to its perceived stereo position, so acoustic motion evolves with visual motion over time.

TEASER FIGURE 01
StereoBind overviewVisual motion → dynamic spatial correspondence → immersive generation
HIGH-RESOLUTION IMAGEFIGURE 01

FIG. 01 StereoBind links visible source motion with the evolving position of generated stereo audio.

04 / DATASETSTEREOWORLD-29K

DATA CONSTRUCTION

Teaching sound
where to move.

StereoWorld-29K pairs stereo audio-video data with corresponding motion tracks through data curation, source discovery and verification, motion tracking, physically grounded 3D audio rendering, and diversity-driven synthesis.

DATASET PIPELINE STEREOWORLD-29K
StereoWorld-29K constructionCuration · AVS agent · spatial data synthesis
HIGH-RESOLUTION IMAGEFIGURE 02

FIG. 02 Construction pipeline for StereoWorld-29K. Open the figure to inspect the full-resolution diagram.

05 / METHODSTEREOBIND

TECHNICAL OVERVIEW

From correspondence
to generation.

Motion tracks condition StereoBind at complementary levels: VMB Tokens establish entity-level audio-visual binding, STE encodes absolute source trajectories, and RT-RoPE models trajectory-induced relative spatial relations.

METHOD FIGURE STEREOBIND
StereoBind architectureEntity binding · absolute trajectory · relative spatial relations
HIGH-RESOLUTION IMAGEFIGURE 03

FIG. 03 StereoBind architecture with VMB Tokens, Spatial Track Encoder, and Residual Track RoPE.

06 / RESULTSSTEREO AUDIO · HEADPHONES RECOMMENDED

QUALITATIVE RESULTS

Follow the source
in stereo.

Examples are organized by spatial state and motion direction. Listen with headphones and compare the visible source trajectory with the perceived stereo position.

Wear headphones for the best experience

01 / 06

Static — Left

The visible source remains on the left, and the generated sound should stay left-localized.

02 / 06

Static — Right

The visible source remains on the right, and the generated sound should stay right-localized.

03 / 06

Dynamic — Left to Right

The visible source moves from left to right, with stereo position following the same trajectory.

04 / 06

Dynamic — Right to Left

The visible source moves from right to left, with stereo position following the same trajectory.

05 / 06

Dynamic — Left → Right → Left

The source reverses direction after reaching the right side, testing whether the acoustic trajectory turns with it.

06 / 06

Dynamic — Right → Left → Right

The source reverses direction after reaching the left side, testing motion-aligned stereo across the full path.

07 / SIDE BY SIDESAME INPUT · MATCHED LOUDNESS

QUALITATIVE COMPARISON

Six scenes.
Six systems.

Each row uses the same input across all six systems. Compare visual fidelity, audio quality, and whether the perceived stereo position follows the visible source motion.

Headphones recommended. The first column is StereoBind. On smaller screens, scroll horizontally to compare every model within the same sample.

08 / QUANTITATIVE RESULTSQUALITY · SYNC · SPATIAL FIDELITY

QUANTITATIVE RESULTS

Measured across
quality and space.

We compare visual quality, audio quality, audio-visual synchronization, and stereo spatial fidelity. Best results are shown in bold and second-best results are underlined.

QUANTITATIVE COMPARISON TABLE 01
Joint video-audio generation resultsVisual quality · audio quality · stereo spatial fidelity
HIGH-RESOLUTION IMAGETABLE 01

TABLE 01 Comparison results on joint video-audio generation. Open the table to inspect all metrics at full resolution.

09 / CITATIONRESEARCH DETAILS

THE PAPER

Build on
this work.

The manuscript contains the complete formulation, model details, experimental protocol, and evaluation.

Citation