講演情報

[U04-P01]Toward FAIR-compliant Open Science: A Knowledge Graph Framework for Integrating Heterogeneous Research Data★Invited Papers

*張 麒1、小財 正義1、金尾 政紀1、門倉 昭1 (1.情報・システム研究機構 データサイエンス共同利用基盤施設)

キーワード:

オープンサイエンス、FAIR 原則、ナレッジグラフ

Open Science aims to accelerate scientific progress by promoting transparency, accessibility, and reuse of research outputs. Alongside policy-driven initiatives and community-based efforts, the FAIR Principles—Findable, Accessible, Interoperable, and Reusable—have become a widely accepted foundation for sustainable research data stewardship. However, despite increasing openness, the practical realization of FAIR-compliant Open Science remains challenging when research data are produced by heterogeneous platforms, methodologies, and disciplinary communities.

A central challenge lies in data heterogeneity. Even when datasets are openly available, differences in formats, metadata structures, terminologies, and implicit assumptions often hinder effective reuse and integration. In addition, contextual information such as provenance, spatial and temporal coverage, and methodological background is frequently fragmented across systems. These issues limit not only human-driven data reuse but also the application of machine learning and artificial intelligence, which rely on semantically consistent and machine-interpretable data representations.

To address these challenges, we propose a knowledge graph–based framework for integrating heterogeneous research data in support of FAIR-compliant Open Science. The framework organizes research information into two conceptual layers: a Data Layer and a Knowledge Layer. This separation clarifies the roles of structured research data and higher-level semantics, while enabling integration without enforcing a single rigid schema.

The Data Layer represents concrete research artifacts such as datasets, observations, samples, and measurements, enriched with essential metadata including identifiers, temporal and spatial information, and measurement values. Entities in this layer provide a structured and persistent representation of research data, serving as stable entry points for data discovery, access, and linkage.

The Knowledge Layer captures semantic concepts and relationships that connect data entities across datasets and domains. This layer represents shared concepts, such as observed variables or categorical entities, as well as contextual relationships including spatial, temporal, and methodological associations. By explicitly modeling these semantics, the Knowledge Layer enables integrated discovery, cross-dataset queries, and reasoning over heterogeneous research data.

The framework has been applied to the integration of multiple real-world research datasets characterized by diverse data types and observational contexts. Through semantic linking between the Data Layer and the Knowledge Layer, previously isolated datasets can be explored in an integrated manner, supporting provenance tracing, exploratory analysis, and cross-domain data interpretation. In addition, the structured representation provided by the knowledge graph offers an AI-ready foundation for downstream analytical workflows.

In summary, this work demonstrates how a two-layer knowledge graph framework can address key barriers to FAIR-compliant Open Science by enabling the semantic integration of heterogeneous research data. By making research data entities and their semantic relationships explicit, the framework enhances interoperability, reusability, and machine-readability, contributing to the development of open and extensible research data infrastructures.