Optical imaging underwater regularly fails: light absorption, scattering, and color distortion reduce contrast and detail so severely that object detection models that work on land experience significant performance losses when directly applied to underwater imagery. Researchers are therefore turning to multimodal approaches that combine camera data with multibeam sonar and stabilize it through foundation models.
Degradation through wavelength-dependent attenuation
The visual problems arise from the physical properties of water: wavelength-dependent attenuation filters out color information, backscatter overlays the useful signal with noise, variable water optics and lighting conditions prevent stable representation. Cameras capture rich semantic information such as color, texture, and shape for fine recognition at short range, but lack depth information and degrade rapidly in turbid water.
Multibeam sonars provide distance measurements through acoustic reflections that remain robust to turbidity and dim light. However, they produce only thin echoes without semantic content for object understanding. No single modality alone offers a complete picture of the underwater scene.
Existing fusion approaches require elaborate calibration
Kim et al. developed an optical-multibeam sonar fusion system with geometric calibration and pixel-level fusion of color and distance information. Li et al. proposed the RGB-Sonar-Tracking benchmark RGBS50 and the SCANet network, which introduced a spatial cross-attention module to mitigate spatial misalignment between modalities and generated "sonar-style" images through simulation to compensate for data scarcity.
Current multimodal fusion methods rely on strictly synchronized, well-calibrated RGB-sonar pair data, which is costly to acquire and requires complex experimental platforms. This complicates deployment for underwater robots that operate long-term in open waters. Some studies resort to "pseudo-sonar images" through simulation, but discrepancies between real and simulated domain can compromise performance.
Vision-language models generalize to unseen objects
Vision-language models offer promising solutions for underwater robotics software according to a study from March 2026, as they generalize to unseen objects and remain robust to noisy conditions through inference of contextual cues. This addresses the challenges of underwater environments, where deep learning models often rely on sparse and noisy labeled data.
The framework UnderwaterVLA, published in September 2025, integrated multimodal foundation models with embodied intelligence for autonomous underwater navigation. Underwater operations are particularly demanding due to hydrodynamic disturbances, limited communication bandwidth, and degraded sensors.
AquaJEPA improves robustness under sensor damage
The framework AquaJEPA, published in July 2026, uses an action-conditioned multimodal JEPA for underwater robot dynamics. The study shows that complete multimodal prediction improves over pure sonar or state components, while sensor dropout training provides robustness under sensor damage.
Cross-modal prediction between camera and sonar modalities remained underexplored due to limited available sonar-visual datasets. In May and June 2026, new sonar-visual datasets for cross-modal underwater robot perception were released to address this gap.
Application in marine inspection and environmental monitoring
Underwater robots, including remotely operated vehicles (ROVs) and autonomous underwater vehicles (AUVs), are increasingly deployed for marine inspection, debris removal, and environmental monitoring. Multi-modal spatial perception approaches focus primarily on vision-based perception, but enhance it through multiple other sensor modalities to mitigate visual degradation in underwater environments.
Research shows progressive development from foundational work on multimodal and 3D visual cues in January 2026 through dual-modal vision-sonar object detection methods in February 2026 to vision-language model evaluation for autonomous underwater robots in March 2026 and current dataset releases.
