Review Comment:
Summary:
The paper introduces a dual-purpose benchmark for (i) evaluating KG construction pipelines via downstream performance and (ii) testing GNN robustness on noisy, text-derived graphs. Three graphs are built over a shared entity set: an expert-curated reference graph (UMLS-NCI; 184,574 nodes, 1.26M edges, 110 relation types, with Semantic Network types injected as nodes), and two graphs automatically extracted from the same MedMentions corpus — GT2KG (OpenIE + LLM validation) and KGGen (fully LLM-based, DeepSeek-Chat). Exact string matching yields 1,032 nodes over 8 semantic types common to all three.
Authors evaluated by performing semi-supervised node classification on those nodes (10/10/80 stratified splits, 5 fixed seeds, standardized training, MiniLM node features), with six baselines spanning convolutional, attentional, relation-parameterized, and geometric-relational message passing. Accuracy falls from ~0.71 on the reference graph to ~0.57/~0.59 on the extracted graphs, with relation-aware and geometric models degrading most gracefully. The code release includes a PyG loader, evaluator API, fixed splits, GitHub repository, and a public leaderboard, which is a solid contribution.
Strengths:
1. The gap is real and precisely stated. No existing resource jointly supplies a corpus, KGs automatically built from it, and an expert-curated reference over overlapping entities such that degradation cannot currently be attributed to the extractor versus the learner. This framing is the paper's strongest asset.
2. The MedMentions/UMLS pairing is economical and non-obvious. CUI annotations give entity alignment for free; semantic types give labels at no annotation cost.
3. Tables 1 and 2 document real extraction noise concretely — spurious is-a, connectors promoted to predicates, composite entities, raw counts as entities, resolution self-loops. Rarer in this literature than it should be, and something researchers can design against.
4. Reproducibility infrastructure is above the median: fixed splits and seeds, a shared training protocol, a PyG Dataset wrapper, a documented submission format, and a public leaderboard. The contribution is coherent.
5. Baselines span meaningfully distinct paradigms (convolutional / attentional / relation-parameterized / geometric), with both convolutional and attention variants, more thorough than most benchmark papers.
Weaknesses:
1. No graph-free baseline. Node features are MiniLM embeddings of UMLS preferred terms; labels are UMLS semantic types. An embedding of "Regulatory Gene" almost certainly encodes "Gene or Genome," so an MLP with no edges may match every number in Table 4, including on the reference graph. Until this is reported, nothing in Table 4 is shown to measure graph quality rather than text quality and the dual-purpose claim is unsupported. Report the full 2x2: PLM features vs. random features, crossed with full graph vs. no edges.
2. The "upper bound" is likely label-leaking. Labels are semantic types, and the reference graph contains type nodes linked by is-a relations, reachable within two hops of 2-layer message passing. The paper never reports how many of the 1,032 labeled nodes are incident to a type node. UMLS predicates are additionally type-constrained by construction, making relation identity near-label-predictive on the reference graph but not on the open-schema graphs, which alone could explain RGCN's win and the size of the gap. Required: Table 4 with and without semantic type integration, along with disclosure of which setting the current numbers use. That this is unstated is a reproducibility failure on the headline table.
3. Quality is confounded with scale. The reference graph and GT2KG differ in noise, but also in size (184K vs 37K nodes), density (average degree roughly 6.8 vs 0.9), fragmentation (1 vs 6,987 components), and relation vocabulary. "Precise quantification of the degradation induced by noise" is therefore overstated: four factors are entangled. A size- and sparsity-matched degraded version of the reference graph is needed to separate semantic noise from structure; a synthetic noise ladder would make the benchmark genuinely diagnostic. Without these, the paper offers three data points, not an instrument.
4. The evaluation set cannot discriminate the models it ranks. Roughly 103 training nodes across 8 classes, and roughly 103 validation nodes doing double duty for early stopping and hyperparameter selection. On GT2KG: 0.508, 0.508, 0.504, 0.542, 0.565 with standard deviations of 0.017 to 0.021; on the reference graph, 0.712 vs 0.696 with standard deviation around 0.010. Bolding a "best" model under those conditions is not defensible without paired testing, and a leaderboard whose top entries are statistically indistinguishable will not function as one. The text also conflates "N = 5 random splits" with "five random seeds" using identical values different variance sources that should be crossed.
5. Homophily and labeled-node connectivity are unreported. Node classification performance tracks homophily more strongly than almost anything else, and the fraction of labeled nodes that are isolated or sitting in components smaller than 10 could dominate the GT2KG and KGGen results outright. Both are one-line additions to Table 3, and without them "noise hurts GNNs" cannot be separated from "these graphs are less homophilous and more fragmented."
6. Text2KGBench and KGrEaT are cited but never compared against. KGrEaT is also downstream-task KG evaluation, so an explicit differentiation is mandatory. Beyond that, resource-paper conventions are unmet: no archival DOI, no licenses stated separately for code and derived data, no precise UMLS version string with reconstruction checksums, no versioning or maintenance plan, and no RDF or SPARQL availability for natively triple-shaped data. A GitHub URL and a Hugging Face Space are not sufficient at this venue.
Minor issues:
1. The caption of the Figure 1 should be more elaborated. The font size of the words used in the figure is very little to read.
2. Cite published versions of Kipf & Welling,(ICLR 2017), Veličković et al. (ICLR 2018). arXiv-only citations for these two are unusual in a journal submission.
3. Table 4 caption should define the conv and attn subscripts; they are explained only obliquely in prose ("we evaluate both standard and attention-enhanced variants").
4. The conclusion substantially restates the introduction. Compress by roughly half and use the space for error analysis.
5. State licenses separately for code, for the GT2KG/KGGen graphs, and for the UMLS reconstruction scripts, given the redistribution restriction.
|