A Unified Benchmark for Evaluating Knowledge Graph Construction Methods and Graph Neural Networks

Tracking #: 4064-5278

Authors: 
Othmane Kabal
Mounira Harzallah
Fabrice Guillet
Hideaki Takeda
Ryutaro Ichise

Responsible editor: 
Michael Cochez

Submission type: 
Dataset Description
Abstract: 
Knowledge graphs automatically constructed from text are increasingly used in real-world applications. However, their inherent noise, fragmentation, and semantic inconsistencies significantly affect the performance of Graph Neural Networks (GNNs) on downstream tasks. Assessing their performance and robustness remains difficult, as it is often unclear whether observed results stem from the learning model or from the quality of the constructed graph itself. In this work, we introduce a dual-purpose benchmark designed to jointly evaluate (i) the performance of GNNs on noisy, text-derived graphs and (ii) the effectiveness of graph construction methods on a downstream task. The benchmark is built in the biomedical domain from a single textual corpus and includes two automatically constructed graphs generated using different extraction methods, alongside a high-quality reference graph curated by experts that serves as an upper performance bound. This design enables controlled comparison of construction methods and systematic evaluation of GNN robustness through semi-supervised node classification. We further provide a standardized, reproducible, and extensible evaluation framework, facilitating the integration of new graph extraction methods and learning models.
Full PDF Version: 
Tags: 
Reviewed

Decision/Status: 
Major Revision

Solicited Reviews:
Click to Expand/Collapse
Review #1
Anonymous submitted on 14/May/2026
Suggestion:
Reject
Review Comment:

The paper identifies that in the world of Graph Neural Networks (GNN) the source of the graph is just as important as the model itself. The authors propose a benchmark with dual purposes: a) to evaluate the performance of GNN on text-derived graphs and b) to evaluate the graph construction methods. Built on the MedMentions biomedical corpus, the study provides two automatically constructed graphs: GT2KG and KGGen alongside a curated reference graph derived from UMLS-NCI. The motivation behind the paper is the fact that existing benchmarks cannot distinguish whether performance degradation originates from the learning model or the construction method. At a first glance, this seems to be an important contribution, but the proposed approach has several issues.
Indeed, many studies that were conducted on the GNN performance considered large scale curated knowledge graphs or domain specific graphs derived from structured data sources (i.e. the graph structure is not extracted from text but is inherent to the data or collectively created). The paper's opening claim is that "knowledge graphs automatically constructed from text are increasingly used in real-world applications." This is true in the knowledge graph construction realm, but it does not follow that GNNs are routinely being trained on these noisy, automatically constructed graphs as their primary input. The benchmark is implicitly assuming a specific pipeline: extract KG from text and then train a GNN on it.
Regarding the first objective (evaluating KG construction methods) it is genuinely useful to know how much downstream performance degrades relative to a curated graph, but the proposed approach is questionable. The authors evaluate two different extraction methods: one is based on OpenIE pipeline followed by a syntactic cleaning and an LLM-based validation and the other one is a fully LLM-based pipeline.
The paper uses UMLS-NCI as the reference graph and MedMentions as the corpus from which the automated graphs are created, but these two were created independently, by different organizations, for different purposes. A proper reference graph for this benchmark would require domain experts to read the actual MedMentions abstracts and manually annotate the relations between the entities that appear in those texts.
Another issue is that the graph are adjusted by the authors (e.g. they applied a multicriteria filtering process: removing the types with less than 2000 instances in the reference graph, keeping the types that are common in all three graphs; a subgraph was selected for the reference graph through edge pruning, adding type declarations). So, there are changes that were performed either on the generated graphs or on the reference one and I am not sure whether these changes are working as intended. On the generated graphs it is not clear why the normalization ensures only node alignment (common nodes across graphs) but does not cover edge space alignment. There is quite a significant difference between the reference graph which has only 110 curated relation types and the generated graphs that contain up to 5787 relation types and I expect this to count in the learned graph representations that are used in the downstream tasks.
On the reference graph, edge pruning removed inverse relations and low-frequency relation types from the reference graph. This means that many relations that are in the UMLS (and might be captured by the GT2KG or KGGen from text) are simply absent from the reference graph. Then the random 1% sampling of concept-type edges additions is not well justified. The authors only mention that “terms are frequently introduced through definitions or contextual explanations using is-a constructions” (so the paper assumes there will be this kind of definitions in the processed abstracts, but these can be completely independent from the added triples). If the extraction pipeline operating on the medical abstracts do not produce type edges adding them to the reference graph makes it less similar to the text derived graphs not more. The structural neighborhood of a node is the reference graph is partly determined by the random process and this creates a structural mismatch.
A structural comparison among the graphs is provided in table 3.
Except a structural comparison between the generated graphs, the authors do not seem to evaluate the quality of the graphs prior to the downstream tasks. The paper does not reveal how the distinction between TBOX and ABOX is handled at all, which is essential for a KG to serve its semantic competency and reasoning purposes. Quantitative indicators seem to be favored (number of nodes, relations etc) instead of the reasoning patterns/capabilities that should be enabled by the resulting graphs, and the potentially induced inconsistencies. The KG purposes must be a factor here and in its quality consideration, not only its structure or size.
The authors claim that new Knowledge graph extraction methods that are proposed by researchers can be tested on the MedMentions corpus and if they exhibit incomplete coverage of types should be understood as a limitation in the construction method (“as it indicates misses relevant entities”). What if researchers are using different adjustment techniques? It is perfectly normal that different KG construction methods to make fundamentally different design choices and this affects coverage (what nodes are included) in ways that have nothing to do with the extraction quality. Still, the authors consider this as an evidence that a KG construction method is of low quality because it misses common nodes.
Regarding the second objective, evaluating the GNN models, the authors state that “researchers developing new GNN architectures can evaluate their methods any of the provided graphs” . I do not think this can be interchangeable. The three graphs differ not just in quality but also in their structure. As table 3 shows, they differ quite substantially regarding the number of nodes, edges, unique relations, connected components. It means that training a GNN model on them will consider these differences. A GNN will perform well on the reference graph as it is exploiting a dense, well-connected graph (every node has the potential to influence every other node given enough layers), and the results reported in table 4 demonstrate this. An additional proof is provided by the results on the RCGN model: it performs best on reference graph but it can not even be evaluated on the KGGen due to memory constraints from the 5787 relation types.
The authors admit that “for consistency, the triples introduced during the semantic type integration step in the clean reference graph can also be added to the automatically constructed graphs”. Supposing that the sampled type edges are differentiated in the reference graph and they are also added in the automatically constructed graphs, how do you ensure that the same concept nodes they connect exist in those graphs?
The authors exclusively tested the quality of the generated graphs through the lens of the downstream task, which is node classification. However, a graph that is excellent for node classification can be terrible for link prediction.
The issues identified above can be summarized on the following three dimensions that are relevant for a dataset paper, as this paper was framed:
1) Quality and stability of the dataset
The quality of the graphs in the provided benchmark is insufficient.
The reference graph benefits from the established credibility of the UMLS-NCI as an expert curated resource, but the adaptations applied to it such as edge pruning, 1% random type additions (which by default implies challenges in reproducing this reference graph) introduce instability.
It is not clear whether the authors used the errors presented in table 1 and 2 as criteria to evaluate the quality of the text-derived graphs or this has been entirely left later in the measurements done for the downstream task.
2) Usefulness of the dataset
Although the dataset is meant to address a surfacing problem, its practical usefulness is limited by several factors. First, it is not domain agnostic as it refers to only the medical domain. Secondly it addresses only a single downstream task. The extensibility claim that new KG construction methods can be submitted and compared is undermined by the adjustments and assumptions about following a specific pipeline.
3) clarity and completeness of the description
The paper is described in reasonable detail, supported by informative figures. The dual purpose framing is stated clearly in the abstract and in the introduction but the design decisions behind the applied methodology are not fully convincing.
Overall the paper provides enough information to understand what was built and why but not that this is what it should and not enough to replicate it or to extend it.

Review #2
Anonymous submitted on 14/May/2026
Suggestion:
Minor Revision
Review Comment:

Dear authors,

We consider the research interesting and relevant. We would like, nevertheless, point to the following improvement opportunities:

(1)- "Second, we excluded relation types with fewer than 50 occurrences, as their contribution is negligible compared to high-frequency relations." -> We consider it would be useful to understand the distribution of relations across nodes, as to understand what negligible means and put it in perspective w.r.t high-frequency relations.

(2)- "The splits are stratified by type to preserve class proportions across sets and avoid class imbalance." -> We understand stratified sampling does not resolve the class imbalance issue: it rather ensures that a class is represented across the train, validation, and test splits. Do the authors apply some additional processing (re-resample the corresponding splits) as to ensure no class imbalance exists? Did the authors perform experiments as to understand whether such a synthetic setting would result in positive outcomes in imbalanced/real-world settings?

(3)- "If some common nodes are missing, the evaluation is restricted to the subset of nodes present in the graph. However, incomplete coverage directly reflects a limitation of the construction method, as it indicates missed relevant entities (i.e., lower recall)." -> (a) "If some common nodes are missing, the evaluation is restricted to the subset of nodes present in the graph. However, incomplete coverage directly reflects a limitation of the construction method, as it indicates missed relevant entities (i.e., lower Recall)."; (b) while the authors mention measuring Accuracy, Macro F1-score, Macro Precision, and Recall, most of these metrics are not being reported in the result tables. We encourage them to include the metrics; (c) the metrics considered above (i) are sensitive to class imbalance - we encorage the authors explore metrics that would not be sensitive to class imbalance, (ii) do not consider that different models may issue different predictive scores distributions, therefore, choosing the cut-off threshold is a key decision that influences the outcome: the authors should be transparent regard the criteria used to determine such a threshold and ensure that the threshold is set adapting to the predictive score distributions of each model.

(4)- While the authors indicate the standard deviation for some of the scores obtained, we would appreciate it if they could indicate whether the results obtained are statistically significant. Furthermore, are there classes for which it is more important to determine whether they are correctly classified or not? Drawing attention to results in such cases would be a relevant aspect to be reported in the leaderboard.

(5)- While the authors provide a GitHub link to the project, which includes data and partitions, we miss a detailed README.md file explaining the structure of the project, dataset characteristics, and instructions on how the benchmark could be executed. While some dataset characteristics are explained in the manuscript, we would expect more technical details regarding the dataset to be described in the README.md to facilitate its use and dissemination.

We acknowledge that the manuscript was submitted as a 'Data Description' and has been reviewed along the dimensions of relevance. While the feedback provided above, in some cases, goes beyond the dataset itself, we consider it relevant to the quality of the submitted manuscript and the proposed benchmark.

Review #3
By Kushal Bose submitted on 01/Aug/2026
Suggestion:
Major Revision
Review Comment:

Summary:

The paper introduces a dual-purpose benchmark for (i) evaluating KG construction pipelines via downstream performance and (ii) testing GNN robustness on noisy, text-derived graphs. Three graphs are built over a shared entity set: an expert-curated reference graph (UMLS-NCI; 184,574 nodes, 1.26M edges, 110 relation types, with Semantic Network types injected as nodes), and two graphs automatically extracted from the same MedMentions corpus — GT2KG (OpenIE + LLM validation) and KGGen (fully LLM-based, DeepSeek-Chat). Exact string matching yields 1,032 nodes over 8 semantic types common to all three.

Authors evaluated by performing semi-supervised node classification on those nodes (10/10/80 stratified splits, 5 fixed seeds, standardized training, MiniLM node features), with six baselines spanning convolutional, attentional, relation-parameterized, and geometric-relational message passing. Accuracy falls from ~0.71 on the reference graph to ~0.57/~0.59 on the extracted graphs, with relation-aware and geometric models degrading most gracefully. The code release includes a PyG loader, evaluator API, fixed splits, GitHub repository, and a public leaderboard, which is a solid contribution.

Strengths:

1. The gap is real and precisely stated. No existing resource jointly supplies a corpus, KGs automatically built from it, and an expert-curated reference over overlapping entities such that degradation cannot currently be attributed to the extractor versus the learner. This framing is the paper's strongest asset.

2. The MedMentions/UMLS pairing is economical and non-obvious. CUI annotations give entity alignment for free; semantic types give labels at no annotation cost.

3. Tables 1 and 2 document real extraction noise concretely — spurious is-a, connectors promoted to predicates, composite entities, raw counts as entities, resolution self-loops. Rarer in this literature than it should be, and something researchers can design against.

4. Reproducibility infrastructure is above the median: fixed splits and seeds, a shared training protocol, a PyG Dataset wrapper, a documented submission format, and a public leaderboard. The contribution is coherent.

5. Baselines span meaningfully distinct paradigms (convolutional / attentional / relation-parameterized / geometric), with both convolutional and attention variants, more thorough than most benchmark papers.

Weaknesses:

1. No graph-free baseline. Node features are MiniLM embeddings of UMLS preferred terms; labels are UMLS semantic types. An embedding of "Regulatory Gene" almost certainly encodes "Gene or Genome," so an MLP with no edges may match every number in Table 4, including on the reference graph. Until this is reported, nothing in Table 4 is shown to measure graph quality rather than text quality and the dual-purpose claim is unsupported. Report the full 2x2: PLM features vs. random features, crossed with full graph vs. no edges.

2. The "upper bound" is likely label-leaking. Labels are semantic types, and the reference graph contains type nodes linked by is-a relations, reachable within two hops of 2-layer message passing. The paper never reports how many of the 1,032 labeled nodes are incident to a type node. UMLS predicates are additionally type-constrained by construction, making relation identity near-label-predictive on the reference graph but not on the open-schema graphs, which alone could explain RGCN's win and the size of the gap. Required: Table 4 with and without semantic type integration, along with disclosure of which setting the current numbers use. That this is unstated is a reproducibility failure on the headline table.

3. Quality is confounded with scale. The reference graph and GT2KG differ in noise, but also in size (184K vs 37K nodes), density (average degree roughly 6.8 vs 0.9), fragmentation (1 vs 6,987 components), and relation vocabulary. "Precise quantification of the degradation induced by noise" is therefore overstated: four factors are entangled. A size- and sparsity-matched degraded version of the reference graph is needed to separate semantic noise from structure; a synthetic noise ladder would make the benchmark genuinely diagnostic. Without these, the paper offers three data points, not an instrument.

4. The evaluation set cannot discriminate the models it ranks. Roughly 103 training nodes across 8 classes, and roughly 103 validation nodes doing double duty for early stopping and hyperparameter selection. On GT2KG: 0.508, 0.508, 0.504, 0.542, 0.565 with standard deviations of 0.017 to 0.021; on the reference graph, 0.712 vs 0.696 with standard deviation around 0.010. Bolding a "best" model under those conditions is not defensible without paired testing, and a leaderboard whose top entries are statistically indistinguishable will not function as one. The text also conflates "N = 5 random splits" with "five random seeds" using identical values different variance sources that should be crossed.

5. Homophily and labeled-node connectivity are unreported. Node classification performance tracks homophily more strongly than almost anything else, and the fraction of labeled nodes that are isolated or sitting in components smaller than 10 could dominate the GT2KG and KGGen results outright. Both are one-line additions to Table 3, and without them "noise hurts GNNs" cannot be separated from "these graphs are less homophilous and more fragmented."

6. Text2KGBench and KGrEaT are cited but never compared against. KGrEaT is also downstream-task KG evaluation, so an explicit differentiation is mandatory. Beyond that, resource-paper conventions are unmet: no archival DOI, no licenses stated separately for code and derived data, no precise UMLS version string with reconstruction checksums, no versioning or maintenance plan, and no RDF or SPARQL availability for natively triple-shaped data. A GitHub URL and a Hugging Face Space are not sufficient at this venue.

Minor issues:

1. The caption of the Figure 1 should be more elaborated. The font size of the words used in the figure is very little to read.

2. Cite published versions of Kipf & Welling,(ICLR 2017), Veličković et al. (ICLR 2018). arXiv-only citations for these two are unusual in a journal submission.

3. Table 4 caption should define the conv and attn subscripts; they are explained only obliquely in prose ("we evaluate both standard and attention-enhanced variants").

4. The conclusion substantially restates the introduction. Compress by roughly half and use the space for error analysis.

5. State licenses separately for code, for the GT2KG/KGGen graphs, and for the UMLS reconstruction scripts, given the redistribution restriction.