Researchers from Zhejiang University, Hong Kong University of Science and Technology, Tsinghua University, and Tencent Jarvis Lab introduced MedUAG on August 19, 2026, a framework for unified understanding and generation in medical multimodal models. The project comprises the largest training dataset for medical Unified Understanding and Generation (UAG) according to the authors, featuring over 6 million instances, a standardized evaluation benchmark, and a trained baseline model.
The research work is available on arXiv under number 2608.18937. Xian Wu and Zuozhu Liu serve as Corresponding Authors. Additional authors include Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, and Jian Wu.
Three Components: Dataset, Benchmark, and Model
MedUAG consists of three core components. The training dataset MedUAGCorpus comprises over 6 million instances according to the researchers and covers 14 different imaging modalities. The dataset was designed for diverse training tasks within the UAG framework.
The benchmark MedUAGBench standardizes the evaluation of medical generation capabilities across 12 different tasks. According to the authors, MedUAGBench extends previous evaluation of medical generation beyond existing boundaries through unified protocols.
The MedUAG model itself was trained end-to-end across all understanding and generation tasks. The researchers report strong performance across a broad spectrum of tasks and position the model as a baseline for next-generation medical multimodal systems.
Problem Statement: Resource Gap in Medical UAG
The authors identify two central deficits in transferring UAG paradigms to medical applications. Despite rapid development of Multimodal Large Language Models (MLLMs), comprehensive, standardized corpora and benchmarks for the medical domain are lacking. Additionally, few validated medical models exist that combine understanding and generation in a unified system – the existing work is limited by the described resource gaps.
According to the research work, MLLMs have evolved from isolated understanding or generation systems to unified UAG frameworks. These enable learning of universal representations within a single system, reduce dependence on task-specific fine-tuning, and scale across heterogeneous modalities.
Parallel Research on Medical Multimodal Learning
MedUAG is part of a series of research works on medical multimodal learning. MedM2G, published on March 7, 2024, was the first medical generative model for unified Medical Multi-Modal Generation according to the publication and supports text-to-image, image-to-text, and modality generation (CT, MRI, X-Ray) via Cross-Guided Diffusion.
UniMedVL, released in October 2025, also pursues a unified approach to medical multimodal understanding and generation through Observation-Knowledge-Analysis. UniMod, published on August 10, 2026, addresses Multi-Modal Learning for medical diagnosis through Cross-Modality and Within-Modality Alignment. In October 2025, the MEDIQA-WV 2025 Shared Task on Medical Visual Question Answering (MedVQA) for Wound Care VQA took place with natural language queries over medical images.
Objective and Scientific Positioning
According to the researchers, MedUAG aims to contribute to the unification of medical AI systems by providing comprehensive resources – data and benchmarks – for medical UAG for the first time, establishing a competitive baseline model, and paving the way for next-generation medical multimodal systems that combine understanding and generation in an integrated framework.
The research work is freely accessible on arXiv. Details on training methodology, architecture, and quantitative results are provided in the publication.
