AI Use Cases for the Digital Documentation and Preservation of Cultural and Scientific Heritage
DOI:
https://doi.org/10.55630/dipp.2026.16.19Keywords:
Multimodal AI, Digital Cultural Heritage, Metadata Enrichment, Semantic Linking, Multimedia RetrievalAbstract
This paper presents a multimodal approach to the artificial intelligence of the digital presentation and preservation of cultural and scientific heritage. Integrating the multiple modalities of indexing, automated metadata enrichment and semantic linking, the framework provides better archival discoverability. The main contribution is a non-destructive architecture that utilizes a shared semantic space to fuse text, visual, and audio modalities while employing a Semantic Anchor strategy to prevent the erasure of "thin" legacy metadata. Experimental findings demonstrate a Mean Precision @ 5 of 0.92 and a 100% Data Retention rate, proving that the framework successfully reclaims "dark data" while maintaining high retrieval consistency across heterogeneous multimedia collections.References
Bearman, D. (2008). Representing museum knowledge. In P. F. Marty & K. Jones (Eds.), Museum Informatics: People, Information, and Technology in Museums. Routledge. https://doi.org/10.4324/9780203939147
Borgman, C. L. (2015). Big data, little data, no data: Scholarship in the networked world. MIT Press.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
Doerr, M. (2003). The CIDOC conceptual reference module: An ontological approach to semantic interoperability of metadata. AI Magazine, 24(3), 75–92. https://doi.org/10.5555/958671.958678
Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.
Hinton, G., Deng, L., Yu, D., Dahl, G., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., & Kingsbury, B. (2012). Deep neural networks for acoustic modeling in speech recognition. IEEE Signal Processing Magazine, 29(6), 82–97. https://doi.org/10.1109/MSP.2012.2205597
Hyvönen, E. (2012). Publishing and using cultural heritage linked data on the Semantic Web. Morgan & Claypool. https://doi.org/10.2200/S00452ED1V01Y201210WBE003
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. In F. Pereira, C. J. Burges, L. Bottou, & K. Weinberger (Eds.), Advances in Neural Information Processing Systems (Vol. 25, pp. 1097–1105). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, & H. Lin (Eds.), Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf
Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press. https://doi.org/10.1017/CBO9780511809071
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In M. Meila & T. Zhang (Eds.), Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8748–8763). PMLR. https://proceedings.mlr.press/v139/radford21a.html
Sivic, J., & Zisserman, A. (2003). Video Google: A text retrieval approach to object matching in videos. In Proceedings of the Ninth IEEE International Conference on Computer Vision (Vol. 2, pp. 1470–1477). IEEE. https://doi.org/10.1109/ICCV.2003.1238663
Smeulders, A. W. M., Worring, M., Santini, S., Gupta, A., & Jain, R. (2000). Content-based image retrieval at the end of the early years. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(12), 1349–1380. https://doi.org/10.1109/34.895972
Trant, J. (2009). Tagging, folksonomy and art museums: Early experiments and ongoing research. Journal of Digital Information, 10(1). https://jodi-ojs-tdl.tdl.org/jodi/article/view/270
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Digital Presentation and Preservation of Cultural and Scientific Heritage

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
