Research on Brain Tumor MRI Image Segmentation Method Based on BiomedCLIP
DOI:
https://doi.org/10.63808/ihf.v2i3.429Keywords:
Medical Image Segmentation, Brain Tumor, Vision-Language Models (VLMs), Cross-modal AttentionAbstract
To address the challenges of blurry lesion boundaries, unstable training with small samples, and insufficient cross-domain generalization in brain tumor MRI image segmentation, this paper proposes BiomedCLIP-Seg, a novel segmentation method based on the biomedical vision-language pre-training model BiomedCLIP. Taking T2-FLAIR axial slices as input, the method introduces medical text semantic priors into the pixel-level lesion localization process through structured medical prompts and zero-initialized soft prompt mechanisms. Meanwhile, a cross-layer visual feature fusion module is designed to integrate shallow boundary textures, mid-level regional contexts, and high-level lesion semantics. Furthermore, a QKV cross-modal segmentation decoder is constructed to achieve explicit interaction between visual patch representations and category text semantics. To alleviate the issues of peripheral missed segmentation and region adhesion caused by weak tumor boundaries, a cross-modal consistency loss and an L2 norm-based boundary constraint loss are further introduced. Experimental results on the BraTS 2021 dataset demonstrate that BiomedCLIP-Seg achieves Dice, IoU, and HD95 scores of 0.903±0.004, 0.823±0.005, and 5.42±0.27 mm, respectively, outperforming representative methods including U-Net, TransUNet, CLIPSeg, nnU-Net, Swin UNETR, MedSAM, and SAMed. Ablation studies and hyperparameter sensitivity analyses further validate the effectiveness of structured medical prompts, soft prompt adaptation, cross-layer feature fusion, cross-modal consistency constraints, and boundary losses in improving model performance. This study provides a new technical pathway for utilizing biomedical vision-language models to tackle weak boundary segmentation problems in medical imaging.
References
[1] Baid, U., Ghodasara, S., Mohan, S., Bilello, M., Calabrese, E., Colak, E., Farahani, K., Kalpathy-Cramer, J., Kitamura, F. C., Pati, S., Prevedello, L. M., Rudie, J. D., Sako, C., Shinohara, R. T., Bergquist, T., Chai, R., Eddy, J., … Bakas, S. (2021). The RSNA-ASNR-MICCAI BraTS 2021 benchmark on brain tumor segmentation and radiogenomic classification [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2107.02314
[2] Bakas, S., Akbari, H., Sotiras, A., Bilello, M., Rozycki, M., Kirby, J. S., Freymann, J. B., Farahani, K., & Davatzikos, C. (2017). Advancing The Cancer Genome Atlas glioma MRI collections with expert segmentation labels and radiomic features. Scientific Data, 4, Article 170117. https://doi.org/10.1038/sdata.2017.117
[3] Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., Lu, L., Yuille, A. L., & Zhou, Y. (2021). TransUNet: Transformers make strong encoders for medical image segmentation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2102.04306
[4] Dice, L. R. (1945). Measures of the amount of ecologic association between species. Ecology, 26(3), 297–302. https://doi.org/10.2307/1931617
[5] Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16×16 words: Transformers for image recognition at scale. In Proceedings of the 9th International Conference on Learning Representations. https://doi.org/10.48550/arXiv.2010.11929
[6] Hatamizadeh, A., Nath, V., Tang, Y., Yang, D., Roth, H. R., & Xu, D. (2022). Swin UNETR: Swin transformers for semantic segmentation of brain tumors in MRI images. In Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries (Vol. 12962, pp. 272–284). Springer. https://doi.org/10.1007/978-3-031-08999-2_22
[7] Hou, J., Zhou, Y., Qin, J., Zhang, J., Li, X., & Gao, S. (2025). SSCLMix: A self-supervised contrastive learning-based data mixing augmentation method. Neural Networks, 192, Article 108171. https://doi.org/10.1016/j.neunet.2025.108171
[8] Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, 18(2), 203–211. https://doi.org/10.1038/s41592-020-01008-z
[9] Jia, M., Tang, L., Chen, B. C., Cardie, C., & Belongie, S. (2022). Visual prompt tuning. In Proceedings of the European Conference on Computer Vision (Vol. 13693, pp. 709–727). Springer. https://doi.org/10.1007/978-3-031-19827-4_41
[10] Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1412.6980
[11] Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 4015–4026). IEEE. https://doi.org/10.1109/ICCV51070.2023.00344
[12] Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. In Proceedings of the 7th International Conference on Learning Representations. https://doi.org/10.48550/arXiv.1711.05101
[13] Ma, J., He, Y., Li, F., Han, L., You, C., & Wang, B. (2024). Segment anything in medical images. Nature Communications, 15(1), Article 654. https://doi.org/10.1038/s41467-024-44940-2
[14] Menze, B. H., Jakab, A., Bauer, S., Kalpathy-Cramer, J., Farahani, K., Kirby, J., Burren, Y., Porz, N., Slotboom, J., Wiest, R., Lanczi, L., Gerstner, E., Weber, M.-A., Arbel, T., Avants, B. B., Ayache, N., Buendia, P., … Van Leemput, K. (2015). The multimodal brain tumor image segmentation benchmark (BRATS). IEEE Transactions on Medical Imaging, 34(10), 1993–2024. https://doi.org/10.1109/TMI.2014.2377694
[15] Milletari, F., Navab, N., & Ahmadi, S. A. (2016). V-Net: Fully convolutional neural networks for volumetric medical image segmentation. In Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV) (pp. 565–571). IEEE. https://doi.org/10.1109/3DV.2016.79
[16] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8748–8763). PMLR. https://doi.org/10.48550/arXiv.2103.00020
[17] Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (Vol. 9351, pp. 234–241). Springer. https://doi.org/10.1007/978-3-319-24574-4_28
[18] Zhang, K., & Liu, D. (2023). Customized segment anything model for medical image segmentation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2304.13785
[19] Lüddecke, T., & Ecker, A. (2022). CLIPSeg: Image segmentation using text and image prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7086–7096). IEEE. https://doi.org/10.1109/CVPR52688.2022.00695
[20] Zhang, S., Xu, Y., Usuyama, N., Xu, H., Bagga, J., Tinn, R., Preston, S., Rao, R., Wei, M., Valluri, N., Wong, C., Tupini, A., Wang, Y., Mazzola, M., Shukla, S., Liden, L., … Poon, H. (2023). BiomedCLIP: A multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2303.00915
[21] Zhou, K., Yang, J., Loy, C. C., & Liu, Z. (2022). Learning to prompt for vision-language models. International Journal of Computer Vision, 130, 2337–2348. https://doi.org/10.1007/s11263-022-01653-1
[22] Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017). Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2881–2890). IEEE. https://doi.org/10.1109/CVPR.2017.660
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Mengyuan Cao, Jialu Zhao, Xinxin Song

This work is licensed under a Creative Commons Attribution 4.0 International License.