A Survey on Metadata for Machine Learning Models and Datasets: Standards, Practices, and Harmonization Challenges

Gesese, Genet-Asefa; Chen, Zongxiong; Zoubia, Oussama; Limani, Fidan; Silva, Kanishka; Suryani, Muhammad Asif; Zapilko, Benjamin; Castro, Leyla Jael; Ekaterina, Kutafina; Solanki, Dhwani; Fliegl, Heike; Schimmler, Sonja; Boukhers, Zeyd; Sack, Harald

doi:10.24406/publica-5924

2025

Conference Paper

Abstract

The growing availability of machine learning (ML) models, datasets, and related artifacts across platforms, such as Hugging Face, GitHub, and Zenodo, has amplified the need for structured and standardized metadata. However, metadata practices remain highly heterogeneous, differing in schema design, vocabulary usage, and semantic expressiveness, posing significant challenges for tasks such as representation, extraction, alignment, and integration. This fragmentation impedes the development of infrastructures that depend on machine-actionable metadata to support discovery, provenance tracking, or cross-platform interoperability. While metadata is also foundational to enabling FAIR (Findable, Accessible, Interoperable, and Reusable) principles in ML, there is a lack of consolidated understanding of how existing standards support interoperability and alignment across platforms. In this survey, we review and compare a range of general-purpose and ML-specific metadata standards, evaluating their suitability for cross-platform alignment, discoverability, extensibility, and interoperability. We assess these standards based on defined criteria and analyze their potential to support unified, FAIR-compliant metadata infrastructures for ML, laying the groundwork for scalable and interoperable tooling in future ML ecosystems.