AI Use Cases for the Digital Documentation and Preservation of Cultural and Scientific Heritage

Authors

  • Alexander I. Iliev SRH University, Campus Berlin, Ernst-Reuter-Platz 10, 10587 Berlin, Germany , Institute of Mathematics and Informatics, Bulgarian Academy of Sciences, 8 Acad. Georgi Bonchev Street, 1113 Sofia, Bulgaria
  • Aman Malik SRH University, Campus Berlin, Ernst-Reuter-Platz 10, 10587 Berlin, Germany
  • Gagan Anil Gowda SRH University, Campus Berlin, Ernst-Reuter-Platz 10, 10587 Berlin, Germany
  • Satyajit Samal SRH University, Campus Berlin, Ernst-Reuter-Platz 10, 10587 Berlin, Germany

DOI:

https://doi.org/10.55630/dipp.2026.16.19

Keywords:

Multimodal AI, Digital Cultural Heritage, Metadata Enrichment, Semantic Linking, Multimedia Retrieval

Abstract

This paper presents a multimodal approach to the artificial intelligence of the digital presentation and preservation of cultural and scientific heritage. Integrating the multiple modalities of indexing, automated metadata enrichment and semantic linking, the framework provides better archival discoverability. The main contribution is a non-destructive architecture that utilizes a shared semantic space to fuse text, visual, and audio modalities while employing a Semantic Anchor strategy to prevent the erasure of "thin" legacy metadata. Experimental findings demonstrate a Mean Precision @ 5 of 0.92 and a 100% Data Retention rate, proving that the framework successfully reclaims "dark data" while maintaining high retrieval consistency across heterogeneous multimedia collections.

References

Bearman, D. (2008). Representing museum knowledge. In P. F. Marty & K. Jones (Eds.), Museum Informatics: People, Information, and Technology in Museums. Routledge. https://doi.org/10.4324/9780203939147

Borgman, C. L. (2015). Big data, little data, no data: Scholarship in the networked world. MIT Press.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

Doerr, M. (2003). The CIDOC conceptual reference module: An ontological approach to semantic interoperability of metadata. AI Magazine, 24(3), 75–92. https://doi.org/10.5555/958671.958678

Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.

Hinton, G., Deng, L., Yu, D., Dahl, G., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., & Kingsbury, B. (2012). Deep neural networks for acoustic modeling in speech recognition. IEEE Signal Processing Magazine, 29(6), 82–97. https://doi.org/10.1109/MSP.2012.2205597

Hyvönen, E. (2012). Publishing and using cultural heritage linked data on the Semantic Web. Morgan & Claypool. https://doi.org/10.2200/S00452ED1V01Y201210WBE003

Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. In F. Pereira, C. J. Burges, L. Bottou, & K. Weinberger (Eds.), Advances in Neural Information Processing Systems (Vol. 25, pp. 1097–1105). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, & H. Lin (Eds.), Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf

Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press. https://doi.org/10.1017/CBO9780511809071

Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In M. Meila & T. Zhang (Eds.), Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8748–8763). PMLR. https://proceedings.mlr.press/v139/radford21a.html

Sivic, J., & Zisserman, A. (2003). Video Google: A text retrieval approach to object matching in videos. In Proceedings of the Ninth IEEE International Conference on Computer Vision (Vol. 2, pp. 1470–1477). IEEE. https://doi.org/10.1109/ICCV.2003.1238663

Smeulders, A. W. M., Worring, M., Santini, S., Gupta, A., & Jain, R. (2000). Content-based image retrieval at the end of the early years. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(12), 1349–1380. https://doi.org/10.1109/34.895972

Trant, J. (2009). Tagging, folksonomy and art museums: Early experiments and ongoing research. Journal of Digital Information, 10(1). https://jodi-ojs-tdl.tdl.org/jodi/article/view/270

Downloads

Published

2026-09-05

How to Cite

I. Iliev, A., Malik, A., Anil Gowda, G., & Samal, S. (2026). AI Use Cases for the Digital Documentation and Preservation of Cultural and Scientific Heritage. Digital Presentation and Preservation of Cultural and Scientific Heritage, 16, 231-244. https://doi.org/10.55630/dipp.2026.16.19

Most read articles by the same author(s)