Semantic Technologies for Patient-Centred Healthcare Data: A Survey of Knowledge Graph Construction, Reasoning, and Evaluation

Tracking #: 4128-5342

This paper is currently under review
Authors: 
Chenyang Ding
Ahmad Taha
João Ponciano

Responsible editor: 
Philipp Cimiano

Submission type: 
Survey Article
Abstract: 
Healthcare is among the most semantically heterogeneous domains to which knowledge graph technologies have been applied. Patient data are dispersed across electronic health records, laboratory systems, imaging archives and free-text clinical narratives, each governed by different vocabularies, granularities and update regimes. This heterogeneity makes the domain a demanding proving ground for semantic technologies, and it generates requirements that the Semantic Web community does not always encounter elsewhere: safety-critical interpretability, regulatory traceability, and institutional constraints that prohibit centralising the data in the first place. This survey examines how knowledge graphs are constructed, reasoned over and evaluated in patient-centred healthcare settings. Working from a focused corpus of twenty studies selected through a documented search of PubMed, IEEE Xplore, Web of Science, Scopus and arXiv, we examine four recurring application areas: patient similarity and cohort classification, diagnostic support, personalised therapy and drug repurposing, and operational surveillance. We organise the material around the design decisions that practitioners actually face, namely whether to anchor a graph in a curated ontology or to learn representations from data; how to align SNOMED CT, UMLS, ICD-10, LOINC, RxNorm, FHIR and OMOP CDM when their scopes only partially overlap; which declarative mapping and validation machinery (R2RML, RML, SHACL) survives contact with production clinical data; and what reasoning is realistically affordable at clinical scale. We situate these decisions within a two-dimensional frame of interpretability against scalability and use it to characterise the trade-offs among ontology-driven, embedding-based and hybrid pipelines. Across the corpus, reporting of external validation, calibration and provenance remains uneven, which limits comparability more than any single technical shortcoming. We close by identifying where semantic infrastructure is already load-bearing in the systems these studies describe, where it goes unreported, and what that asymmetry suggests for the Semantic Web research agenda.
Full PDF Version: 
Tags: 
Under Review