Trustworthy Retrieval-Augmented Generation for the Digital Preservation and Transparent Presentation of Biomedical Scientific Heritage
DOI:
https://doi.org/10.55630/dipp.2026.16.31Keywords:
Retrieval-Augmented Generation, Biomedical NLP, Trustworthy AI, Digital Preservation, Biomedical Trust IndexAbstract
Biomedical scientific heritage faces a dual challenge of exponential growth and inaccessibility. This paper presents a Trustworthy Retrieval Augmented Generation (RAG) framework designed for the digital preservation and transparent presentation of biomedical knowledge integrating ontology aware retrieval (MeSH/UMLS), hybrid BM25 + dense search via Reciprocal Rank Fusion, sentence level attribution, and a novel composite biomedical Trust Index (BTI) quantifying factual accuracy, attribution fidelity and adversarial robustness on the MedQuAD corpus, the framework achieves 89.1% top-5 recall, 85.6% attribution accuracy, F1 score of 75.2 and BTI of 0.84 on clean data.References
Aghaebrahimian, A., Anisimova, M., & Gil, M. (2022). Ontology-aware biomedical relation extraction. In P. Sojka, A. Horák, I. Kopeček, & K. Pala (Eds.), Text, speech, and dialogue: 25th International Conference, TSD 2022, Brno, Czech Republic, September 6–9, 2022, proceedings (Lecture Notes in Computer Science, Vol. 13502, pp. 160–171). Springer. https://doi.org/10.1007/978-3-031-16270-1_14
Albahri, A. S., Duhaim, A. M., Fadhel, M. A., Alnoor, A., Baqer, N. S., Alzubaidi, L., Albahri, O. S., Alamoodi, A. H., Bai, J., Salhi, A., Santamaría, J., Ouyang, C., Gupta, A., Gu, Y., & Deveci, M. (2023). A systematic review of trustworthy and explainable artificial intelligence in healthcare: Assessment of quality, bias risk, and data fusion. Information Fusion, 96, 156–191. https://doi.org/10.1016/j.inffus.2023.03.008
Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., & Weston, J. (2024). Chain-of-verification reduces hallucination in large language models. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024 (pp. 3563–3578). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.212
Fadeeva, E., Rubashevskii, A., Shelmanov, A., Petrakov, S., Li, H., Mubarak, H., Tsymbalov, E., Kuzmin, G., Panchenko, A., Baldwin, T., Nakov, P., & Panov, M. (2024). Fact-checking the output of large language models via token-level uncertainty quantification. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024 (pp. 9367–9385). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.558
Guan, J., Dodge, J., Wadden, D., Huang, M., & Peng, H. (2024). Language models hallucinate, but may excel at fact verification. In K. Duh, H. Gomez, & S. Bethard (Eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 1090–1111). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.naacl-long.62
Hang, C. N., Yu, P. D., & Tan, C. W. (2025). TrumorGPT: Graph-based retrieval-augmented large language model for fact-checking. IEEE Transactions on Artificial Intelligence. https://doi.org/10.1109/TAI.2025.3567369
Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240. https://doi.org/10.1093/bioinformatics/btz682
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, & H. Lin (Eds.), Advances in Neural Information Processing Systems 33 (pp. 9459–9474). Curran Associates. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
Li, M., Zhan, Z., Yang, H., Xiao, Y., Zhou, H., Huang, J., & Zhang, R. (2025). Benchmarking retrieval-augmented large language models in biomedical NLP: Application, robustness, and self-awareness. Science Advances, 11(47). https://doi.org/10.1126/sciadv.adr1443
Manakul, P., Liusie, A., & Gales, M. J. (2023). SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. In H. Bouamor, J. Pino, & K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 9004–9017). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.557
Sood, S., & Imran, H. (2024). Information retrieval and query expansion for biomedical data. In A. Sharan, N. Malik, H. Imran, & I. Ghosh, (Eds), Text mining approaches for biomedical data (pp. 193–235). Springer Nature Singapore. https://doi.org/10.1007/978-981-97-3962-2_11
Wang, S., Zhuang, S., & Zuccon, G. (2021). BERT-based dense retrievers require interpolation with BM25 for effective passage retrieval. In Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR '21) (pp. 317–324). Association for Computing Machinery. https://doi.org/10.1145/3471158.3472233
Zhou, Y., Liu, Y., Li, X., Jin, J., Qian, H., Liu, Z., Li, C., Dou, Z., Ho, T.-Y., & Yu, P. S. (2024). Trustworthiness in retrieval-augmented generation systems: A survey. arXiv preprint. https://doi.org/10.48550/arXiv.2409.10102
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Digital Presentation and Preservation of Cultural and Scientific Heritage

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
