UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map

Keke Yang1,*, Erqi Wang1,*, Sainan Guan2, Hongliang Ren1,†
1The Chinese University of Hong Kong
2The Eighth Affiliated Hospital, Sun Yat-sen University

* Equal contribution   † Corresponding author

Preprint · Under review

Slide
Rock
Depth adjustment
Sweep

Ultrasound futures predicted by UltraWorld from an initial observation and a sequence of scanning actions.

Abstract

World models can enable autonomous ultrasound scanning by predicting the outcomes of probe motions from local observations. Learning this action–observation relationship typically relies on synchronized video–pose pairs, which are costly to collect at scale and largely unavailable in routine clinical recordings. Reliable action following further requires modeling ultrasound’s cross-sectional sampling geometry.

We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations. Starting from clinical videos, we adapt a video foundation model into an ultrasound generator conditioned on reference images and anatomical masks. Anatomical masks sampled along programmable trajectories through 3D anatomy provide spatial guidance for synthesizing action–video pairs. We then use these synthetic pairs to self-distill the generator into a world model that predicts future observations from local observations and actions, without requiring anatomical masks or other 3D assets at inference time. To further improve action following, we introduce the Acoustic Sampling Map (AsMap), which represents probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths.

Experiments demonstrate improved prediction fidelity and action following. Across nine simulated closed-loop local planning episodes, UltraWorld reduces the mean final distance to the goal and orientation error by 29% and 38%, respectively, compared with visual servoing.

Overview

Figure 2: UltraWorld framework
Learning interactive ultrasound world models without tracked clinical scans. Routine clinical videos lack the synchronized probe poses needed for action supervision. UltraWorld learns visual and temporal priors from these videos, then uses anatomical masks sampled along programmable trajectories through 3D anatomy to synthesize action–video pairs. A shared video backbone transfers these priors to an AsMap-conditioned world model through self-distillation. The resulting model predicts future ultrasound observations from the current observation and prescribed scanning actions, without anatomical masks or 3D assets at inference.

Acoustic Sampling Map

Figure 3: Acoustic Sampling Map
Acquisition-geometric parameterization. AsMap resolves pose-only under-specification by jointly parameterizing probe pose Tt and intrinsic acquisition parameters κt as a seven-channel, first-frame-relative sampling field, injected through token-aligned additive conditioning.
How scanning actions change the sampled anatomy. The animation pairs the probe and B-mode cross-section with maps of sampling position, beam direction, and depth. AsMap makes the spatial effects of probe motion and imaging-depth adjustments explicit at each pixel, including sampling changes that can occur without a change in probe pose. This gives the world model geometric conditioning for predicting how ultrasound observations respond to scanning actions.

Self-distillation

Figure 4: Self-distillation recipe
Transferring clinical video priors to action-conditioned prediction. A mask-conditioned ultrasound generator learns appearance and temporal priors, but requires anatomical guidance that is unavailable during interactive scanning. UltraWorld initializes the world model from this generator and replaces mask conditioning with AsMap while retaining the observation input. On synthesized action–video trajectories, it first trains the geometry encoder with the video backbone frozen, then jointly fine-tunes both. This transfers the learned priors to prediction from local observations and scanning actions, without requiring real video–pose pairs or anatomical masks at inference.

Experiments

Slide

Rock

Sweep

Depth adjustment

Revisit

Quantitative comparisons

Table 1, original paper screenshot: comparison of action-conditioning representations

Analysis

Parameter comparison

Appendix Table 4: action encoder architectures and parameter counts

Generalization to real clinical ultrasound

Table 3, original paper screenshot: BUSI and BUV zero-action and round-trip consistency

Unseen initial frames from patient-held-out BUV and independent BUSI evaluate zero-action stability and round-trip recovery consistency. These experiments do not use paired real future sequences.

Applications

Robotic Ultrasound Planning

Figure 8: Closed-loop robotic ultrasound planning
Closed-loop planning in simulation. UltraWorld imagines candidate futures for model predictive control with the cross-entropy method. Only the first optimized action is executed before replanning from the new observation.

BibTeX

@misc{yang2026ultraworld,
  title={UltraWorld: Learning Interactive Ultrasound World Models
         from Untracked Clinical Videos with Acoustic Sampling Map},
  author={Yang, Keke and Wang, Erqi and Guan, Sainan and Ren, Hongliang},
  year={2026},
  note={Preprint. Under review.}
}

Citation for the supplied preprint.