Research on Personalized Popular Science Q&A Generation Method Integrating Difficulty-Aware Retrieval-Augmented Generation
DOI:
https://doi.org/10.63808/css.v2i3.426Keywords:
Science education, Large language model, Retrieval-augmented generation, Personalized Q&A, Parameter-efficient fine-tuningAbstract
General large language models suffer from factual hallucinations, mismatched expression difficulty and insufficient individual adaptation in science education question-and-answer (Q&A) tasks. To address these issues, this paper proposes a personalized popular science Q&A generation method integrating difficulty-aware retrieval-augmented generation. Firstly, a scientific knowledge base with academic stage difficulty labels is constructed in this method. Secondly, learner portrait matching items and knowledge difficulty deviation penalty items are introduced into the traditional semantic retrieval score to realize evidence screening for students in different academic stages. In the generation stage, structured prompts are organized with retrieved evidence, student portraits and answer constraints, and parameter-efficient fine-tuning is adopted to enhance the model’s adaptability to popular science Q&A style and knowledge boundaries. Taking a self-built primary and secondary school science Q&A dataset as the research object, experiments are conducted from the dimensions of scientific accuracy, age appropriateness, completeness and hallucination rate. The experimental results show that the proposed method outperforms conventional retrieval-augmented methods in terms of scientific accuracy and age appropriateness, and can improve the reliability and teaching applicability of popular science Q&A with low training costs.
References
[1] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., … Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems (Vol. 33, pp. 1877–1901).
[2] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (Vol. 35, pp. 24824–24837).
[3] Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474).
[4] Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W.-t. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 6769–6781). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.550
[5] Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In Proceedings of the 10th International Conference on Learning Representations. https://openreview.net/forum?id=nZeVKeeFYf9
[6] Li, X. L., & Liang, P. (2021). Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (pp. 4582–4597). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.353
[7] Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., & Gelly, S. (2019). Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (Vol. 97, pp. 2790–2799). PMLR. http://proceedings.mlr.press/v97/houlsby19a.html
[8] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
[9] Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389. https://doi.org/10.1561/1500000019
[10] Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., & Wang, H. (2023). Retrieval-augmented generation for large language models: A survey [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2312.10997
[11] Mallen, A., Asai, A., Zhong, V., Pérez, E., & Hajishirzi, H. (2023). When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (pp. 9802–9822). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.546
[12] Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., & Winne, P. H. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274
[13] Zhai, X. (2022). ChatGPT user experience: Implications for education. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4312418
[14] Wang, S., Scells, H., Zuccon, G., & Azzopardi, L. (2021). Neural ranking models for document retrieval: A survey. Information Retrieval Journal, 24, 349–388. https://doi.org/10.1007/s10791-021-09396-7
[15] OpenAI. (2023). GPT-4 technical report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2303.08774
[16] Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1410
[17] Johnson, J., Douze, M., & Jégou, H. (2019). Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3), 535–547. https://doi.org/10.1109/TBDATA.2019.2956512
[18] Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes. Harvard University Press.
[19] Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
[20] Nogueira, R., & Cho, K. (2019). Passage re-ranking with BERT [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1901.04085
[21] Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378–382. https://doi.org/10.1037/h0031619
[22] Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. https://doi.org/10.1037/0033-2909.86.2.420
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Yuewen Cao, Xinyang Li, Chunyu Lei, Lin Zhang

This work is licensed under a Creative Commons Attribution 4.0 International License.