ONCO/RADAR Zurück zur Suche
arXivComputergestütztProstate

A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation

Yang Xing · Jiong Wu · Savas Ozdemir · Yang Zhou · Boxiao Yu · Ying Zhang · Zheren Zhu · Chenyu You · Wei Shao · Yang Lu · Kang Wang · Tinsu Pan · Yang Yang · Kuang Gong

Originalquelle öffnen Nachvollziehbarer Datensatz

01

Zusammenfassung

Accurate PSMA PET/CT interpretation is central to prostate cancer management, yet existing PET/CT AI models typically address isolated tasks. We propose a unified PSMA PET/CT vision-language model for report generation, visual question answering, and lesion segmentation. The framework adopts an LLaVA-style architecture, comprising a PET/CT vision encoder, an MLP-Mixer projection module, a LoRA-tuned large language model, and a 3D segmentation branch. Training followed a four-stage strategy: vision encoder pretraining, projection-layer alignment, VLM fine-tuning, and final multitask tuning. Language tasks used 5,747 PSMA PET/CT datasets with paired reports, while segmentation used the PSMA subset of AutoPET. The model outperformed PET2REP and a CT-based baseline across standard report-generation metrics, improved performance across VQA question types, and achieved higher Dice and lesion-level overlap F1 than SegAnyPET and nnUNet. These results support the feasibility of a unified framework for structured, interactive, interpretable PSMA PET/CT analysis with voxel-level grounding within a single multitask model architecture.

02

Indexierte Passagen

Suchergebnisse verweisen auf diese Retrieval-Einheiten und bewahren den Publikationsbezug.

Dieser Datensatz ist noch nicht im Hybridindex enthalten.