Research on Popular Science Image-Text Generation Method Based on Concept Constraint and Semantic Consistency Reranking
DOI:
https://doi.org/10.63808/css.v2i3.427Keywords:
Popular science education, Multimodal generation, Diffusion model, Concept constraint, Image-text consistencyAbstract
General text-to-image models suffer from key concept missing, image-text semantic deviation and insufficient educational applicability in popular science education scenarios. To address these problems, this paper proposes a popular science image-text generation method based on concept constraint and semantic consistency reranking. Firstly, the method extracts scientific concept sets from input popular science texts and constructs enhanced prompts combined with disciplinary categories and image expression requirements. Secondly, the pre-trained Chinese text-to-image generator IDEA-CCNL /Taiyi-Stable-Diffusion-1B-Chinese-v0.1 is invoked to generate multiple candidate images. Finally, candidate images are reranked and the optimal result is output by comprehensively integrating CLIP image-text similarity, scientific concept coverage and image educational applicability scores. Experiments are conducted on a self-built popular science text-image evaluation dataset to compare the proposed method with conventional diffusion model generation, prompt enhancement and single CLIP reranking methods. The experimental results show that the proposed method can effectively improve image-text consistency and concept expression completeness, and is especially suitable for lightweight application scenarios such as scientific concept explanation, classroom illustration matching and popular science resource generation.
The paper further reports exact implementation settings, concept extraction rules, an educational applicability rubric, evaluator calibration, inter-rater agreement, per-discipline results, statistical significance testing, qualitative comparisons and failure cases, thereby improving evaluation transparency and reproducibility.
References
[1] Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 10684–10695). IEEE. https://doi.org/10.1109/CVPR52688.2022.01042
[2] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8748–8763). PMLR. https://doi.org/10.48550/arXiv.2103.00020
[3] Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., & Kalyan, A. (2022). Learn to explain: Multimodal reasoning via thought chains for science question answering. In Advances in Neural Information Processing Systems (Vol. 35, pp. 2507–2521).
[4] Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (Vol. 33, pp. 6840–6851). https://doi.org/10.48550/arXiv.2006.11239
[5] Nichol, A. Q., & Dhariwal, P. (2021). Improved denoising diffusion probabilistic models. In Proceedings of the 38th International Conference on Machine Learning (Vol. 162, pp. 8162–8171). PMLR. https://doi.org/10.48550/arXiv.2102.09672
[6] Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Seyed Ghasemipour, S. K., Gontijo Lopes, R. K., Ayan, B. K., Salimans, T., Ho, J., Fleet, D. J., & Norouzi, M. (2022). Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems (Vol. 35, pp. 36479–36494). https://doi.org/10.48550/arXiv.2205.11487
[7] Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., & Sutskever, I. (2021). Zero-shot text-to-image generation. In Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8821–8831). PMLR. http://proceedings.mlr.press/v139/ramesh21a.html
[8] Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with CLIP latents [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2204.06125
[9] Hessel, J., Holtzman, A., Forbes, M., & Choi, Y. (2021). CLIPScore: A reference-free evaluation metric for image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 7514–7528). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.595
[10] Li, J., Li, D., Xiong, C., & Hoi, S. (2022). BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the 39th International Conference on Machine Learning (Vol. 162, pp. 12888–12900). PMLR. http://proceedings.mlr.press/v162/li22n.html
[11] Li, J., Li, D., Savarese, S., & Hoi, S. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 19730–19742). PMLR. http://proceedings.mlr.press/v202/li23q.html
[12] Zhang, C., Zhang, C., Zheng, S., Qiao, Y., Li, C., Zhang, M., Dam, A., Thrush, T., Kembhavi, A., & Ostendorf, M. (2023a). A survey on multimodal large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2306.13549
[13] Maharana, A., Hannan, D., & Bansal, M. (2022). StoryDALL-E: Adapting pretrained text-to-image transformers for story continuation. In Proceedings of the 17th European Conference on Computer Vision (ECCV) (pp. 70–87). Springer. https://doi.org/10.1007/978-3-031-19769-9_5
[14] Liu, V., & Chilton, L. B. (2022). Design guidelines for prompt engineering text-to-image generative models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (pp. 1–23). Association for Computing Machinery. https://doi.org/10.1145/3491102.3501824
[15] Mayer, R. E. (2020). Multimedia learning (3rd ed.). Cambridge University Press.
[16] Ainsworth, S. (2006). DeFT: A conceptual framework for considering learning with multiple representations. Learning and Instruction, 16(3), 183–198. https://doi.org/10.1016/j.learninstruc.2006.02.001
[17] Zhang, L., Rao, A., & Agrawala, M. (2023b). Adding conditional control to text-to-image diffusion models. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3836–3847). IEEE. https://doi.org/10.1109/ICCV51070.2023.00355
[18] Xue, D., Qian, S., & Xu, C. (2023a). Variational causal inference network for explanatory visual question answering. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 2515–2525). IEEE. https://doi.org/10.1109/ICCV51070.2023.00220
[19] Xue, D., Qian, S., & Xu, C. (2024). Integrating neural-symbolic reasoning with variational causal inference network for explanatory visual question answering. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 7893–7808. https://doi.org/10.1109/TPAMI.2024.3388273
[20] Xue, D., Qian, S., Fang, Q., & Xu, C. (2022). MMT: Image-guided story ending generation with multimodal memory transformer. In Proceedings of the 30th ACM International Conference on Multimedia (MM) (pp. 750–758). Association for Computing Machinery. https://doi.org/10.1145/3503161.3547916
[21] Xue, D., Qian, S., Fang, Q., & Xu, C. (2025). LININ: Logic integrated neural inference network for explanatory visual question answering. IEEE Transactions on Multimedia. https://doi.org/10.1109/TMM.2024.3521709
[22] Wang, J., Zhang, Y., Zhang, L., Yang, P., Gao, X., Wu, Z., Dong, X., He, J., Zhuo, J., Yang, Q., Huang, Y., Li, X., Wu, Y., Lu, J., Zhu, X., Chen, W., Han, T., Pan, K., Wang, R., Wang, H., Wu, X., Zeng, Z., Chen, C., Gan, R., & Zhang, J. (2022). Fengshenbang 1.0: Being the foundation of Chinese cognitive intelligence [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2209.02970
[23] Yang, A., Pan, J., Lin, J., Men, R., Zhang, Y., Zhou, J., & Zhou, C. (2022). Chinese CLIP: Contrastive vision-language pretraining in Chinese [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2211.01335
[24] Qwen. (2025). Qwen2.5 technical report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.15115
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Yuewen Cao, Xinyang Li, Chunyu Lei, Lin Zhang

This work is licensed under a Creative Commons Attribution 4.0 International License.