Trustworthy Retrieval-Augmented Generation for the Digital Preservation and Transparent Presentation of Biomedical Scientific Heritage

Authors

  • Harshetha Murthy Keshav Murthy SRH Hochschule Berlin, 221 Sonnenallee, 12059 Berlin, Germany
  • Alexander I. Iliev SRH Hochschule Berlin, 221 Sonnenallee, 12059 Berlin, Germany , Institute of Mathematics and Informatics, Bulgarian Academy of Sciences, 8 Acad. Georgi Bonchev Street, 1113 Sofia, Bulgaria

DOI:

https://doi.org/10.55630/dipp.2026.16.31

Keywords:

Retrieval-Augmented Generation, Biomedical NLP, Trustworthy AI, Digital Preservation, Biomedical Trust Index

Abstract

Biomedical scientific heritage faces a dual challenge of exponential growth and inaccessibility. This paper presents a Trustworthy Retrieval Augmented Generation (RAG) framework designed for the digital preservation and transparent presentation of biomedical knowledge integrating ontology aware retrieval (MeSH/UMLS), hybrid BM25 + dense search via Reciprocal Rank Fusion, sentence level attribution, and a novel composite biomedical Trust Index (BTI) quantifying factual accuracy, attribution fidelity and adversarial robustness on the MedQuAD corpus, the framework achieves 89.1% top-5 recall, 85.6% attribution accuracy, F1 score of 75.2 and BTI of 0.84 on clean data.

References

Aghaebrahimian, A., Anisimova, M., & Gil, M. (2022). Ontology-aware biomedical relation extraction. In P. Sojka, A. Horák, I. Kopeček, & K. Pala (Eds.), Text, speech, and dialogue: 25th International Conference, TSD 2022, Brno, Czech Republic, September 6–9, 2022, proceedings (Lecture Notes in Computer Science, Vol. 13502, pp. 160–171). Springer. https://doi.org/10.1007/978-3-031-16270-1_14

Albahri, A. S., Duhaim, A. M., Fadhel, M. A., Alnoor, A., Baqer, N. S., Alzubaidi, L., Albahri, O. S., Alamoodi, A. H., Bai, J., Salhi, A., Santamaría, J., Ouyang, C., Gupta, A., Gu, Y., & Deveci, M. (2023). A systematic review of trustworthy and explainable artificial intelligence in healthcare: Assessment of quality, bias risk, and data fusion. Information Fusion, 96, 156–191. https://doi.org/10.1016/j.inffus.2023.03.008

Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., & Weston, J. (2024). Chain-of-verification reduces hallucination in large language models. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024 (pp. 3563–3578). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.212

Fadeeva, E., Rubashevskii, A., Shelmanov, A., Petrakov, S., Li, H., Mubarak, H., Tsymbalov, E., Kuzmin, G., Panchenko, A., Baldwin, T., Nakov, P., & Panov, M. (2024). Fact-checking the output of large language models via token-level uncertainty quantification. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024 (pp. 9367–9385). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.558

Guan, J., Dodge, J., Wadden, D., Huang, M., & Peng, H. (2024). Language models hallucinate, but may excel at fact verification. In K. Duh, H. Gomez, & S. Bethard (Eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 1090–1111). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.naacl-long.62

Hang, C. N., Yu, P. D., & Tan, C. W. (2025). TrumorGPT: Graph-based retrieval-augmented large language model for fact-checking. IEEE Transactions on Artificial Intelligence. https://doi.org/10.1109/TAI.2025.3567369

Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240. https://doi.org/10.1093/bioinformatics/btz682

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, & H. Lin (Eds.), Advances in Neural Information Processing Systems 33 (pp. 9459–9474). Curran Associates. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html

Li, M., Zhan, Z., Yang, H., Xiao, Y., Zhou, H., Huang, J., & Zhang, R. (2025). Benchmarking retrieval-augmented large language models in biomedical NLP: Application, robustness, and self-awareness. Science Advances, 11(47). https://doi.org/10.1126/sciadv.adr1443

Manakul, P., Liusie, A., & Gales, M. J. (2023). SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. In H. Bouamor, J. Pino, & K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 9004–9017). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.557

Sood, S., & Imran, H. (2024). Information retrieval and query expansion for biomedical data. In A. Sharan, N. Malik, H. Imran, & I. Ghosh, (Eds), Text mining approaches for biomedical data (pp. 193–235). Springer Nature Singapore. https://doi.org/10.1007/978-981-97-3962-2_11

Wang, S., Zhuang, S., & Zuccon, G. (2021). BERT-based dense retrievers require interpolation with BM25 for effective passage retrieval. In Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR '21) (pp. 317–324). Association for Computing Machinery. https://doi.org/10.1145/3471158.3472233

Zhou, Y., Liu, Y., Li, X., Jin, J., Qian, H., Liu, Z., Li, C., Dou, Z., Ho, T.-Y., & Yu, P. S. (2024). Trustworthiness in retrieval-augmented generation systems: A survey. arXiv preprint. https://doi.org/10.48550/arXiv.2409.10102

Downloads

Published

2026-09-05

How to Cite

Murthy Keshav Murthy, H., & I. Iliev, A. (2026). Trustworthy Retrieval-Augmented Generation for the Digital Preservation and Transparent Presentation of Biomedical Scientific Heritage. Digital Presentation and Preservation of Cultural and Scientific Heritage, 16, 377-386. https://doi.org/10.55630/dipp.2026.16.31

Most read articles by the same author(s)

1 2 > >>